Skip to content

Specify the binary encoding, so Protocol Buffers is not required - #12

Merged
strimo378 merged 1 commit into
mainfrom
claude/onnx-metanorma-standard-gloozz
Sep 21, 2026
Merged

strimo378 merged 1 commit into
mainfrom
claude/onnx-metanorma-standard-gloozz

Conversation

@strimo378

Copy link
Copy Markdown
Contributor

Clause 12 said a model "is serialized as a Protocol Buffers message" and left the rest to an editorial note, because Protocol Buffers has no citable specification with a stable identifier. That was one of the six blocking gaps: a standard cannot rest its wire format on a document it may not cite.

The clause now states the octets itself. An implementer needs this document and nothing else.

12.2 the four primitive encodings and their wire types: variable-length integer, 32-bit, 64-bit, length-delimited. Signed integers as the varint of their 64-bit two's complement, so every negative value is ten octets
12.3 the field key as the varint of (tag × 8) + wire type; order and repetition; the packed and unpacked forms of a repeated field
12.4 skipping an unrecognized field by wire type alone — which is why wire types 3 and 4 are excluded
12.5 limits
12.6 the message definitions, which stay in Annex B
12.7 the relationship to Protocol Buffers

Protocol Buffers moves from the normative references, where it could not be cited, to the bibliography. 12.7 says plainly that the encoding above is that of Protocol Buffers restricted to the constructs Annex B uses, that an implementation may therefore be built on a Protocol Buffers library, and that where the two differ this document governs.

The restriction is what makes the clause short enough to state at all: no groups, no zig-zag or fixed-width integer types, seven scalar types in all.

Annex B gains a wire type column, so the annex and the clause fit together — the tag and the wire type are what the field key is made of.

None of this is asserted

I wrote the clause from a reading of the format, not from a specification this standard is allowed to cite. That is only worth something if what the clause says is what implementations actually write. So scripts/check-encoding.rb restates the rules as a small encoder and compares its output, byte for byte, against the Protocol Buffers runtime:

  • every scalar type the schema uses — including a negative int32, the largest uint64, a float negative zero, a string beyond ASCII
  • both field key lengths, tag 1 and tag 20. Nine ONNX fields have tags above 15; I would otherwise have been stating the two-octet key untested
  • both forms of a repeated field, and both forms of one field in one message
  • the literal octets of the clause's own worked examples — an edit that makes an example wrong fails the build
  • a nested message, by bootstrap: the script builds a message descriptor by encoding a FileDescriptorProto with its own encoder and hands it to the reference library, which rejects it or produces the wrong fields if the encoding is wrong. That is also where the high tags and the packed field come from — the library's own bundled types have neither
encoding: the clause and the reference implementation agree

28 of 28 comparisons agree. make check-encoding, run by CI in both workflows, so the clause cannot drift from what implementations write.

Also

google-protobuf is declared in the Gemfile. It arrived transitively through the Metanorma toolchain, and an undeclared dependency has broken this build twice already in this repository.

What is left of the gap

Only the maximum message size, now a minor row rather than a major one: the reference implementation applies 2 GB to a single message, and this document neither fixes that nor leaves it explicitly to the implementation.

One of the six blocking rows in Annex C is gone.

Verification

All three parts compile with "Syntax Valid", clean at severity 3 only — two new informational rows, both Relaton failing to resolve the new Protocol Buffers bibliography entry, the same class as the existing ONNX and JDF entries. check-schema, check-operators, check-encoding and check-stylesheet all pass. The rendered clause was read back from the HTML.

PDF is CI's, as usual — the STIX font host returns 403 in the authoring environment.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD


Generated by Claude Code

Clause 12 stated that a model "is serialized as a Protocol Buffers message"
and left the rest to an editorial note, because Protocol Buffers has no
citable specification with a stable identifier. That was one of the blocking
gaps: a standard cannot rest its wire format on a document it may not cite.

The clause now states the octets itself. An implementer needs this document
and nothing else.

  12.2  the four primitive encodings, with their wire types: the
        variable-length integer, the 32-bit and 64-bit forms, and the
        length-delimited form; signed integers as the varint of their 64-bit
        two's complement, so every negative value is ten octets
  12.3  the field key as the varint of (tag x 8) + wire type; order and
        repetition; the packed and unpacked forms of a repeated field
  12.4  skipping a field the reader does not recognize, by wire type alone --
        which is why wire types 3 and 4 are excluded
  12.5  limits
  12.6  the message definitions, which stay in Annex B
  12.7  the relationship to Protocol Buffers

Protocol Buffers moves from the normative references, where it could not be
cited, to the bibliography. 12.7 says plainly that the encoding above is
that of Protocol Buffers restricted to the constructs Annex B uses, that an
implementation may therefore be built on a Protocol Buffers library, and
that where the two differ this document governs. The restriction is what
makes the clause short enough to state: no groups, no zig-zag or fixed-width
integer types, seven scalar types in all.

Annex B gains a wire type column, so the annex and the clause fit together:
the tag and the wire type are what the field key is made of.

None of this is asserted from a reading of the format. scripts/check-encoding.rb
restates the rules of the clause as a small encoder and compares its output,
byte for byte, against the Protocol Buffers runtime:

  - every scalar type the schema uses, including a negative int32, the
    largest uint64, a float negative zero and a string beyond ASCII
  - both field key lengths -- tag 1 and tag 20, since nine ONNX fields have
    tags above 15 and I would otherwise be stating the two-octet key untested
  - both forms of a repeated field, and both forms of one field in one
    message
  - the literal octets of the clause's own worked examples, so an edit that
    makes an example wrong fails the build
  - a nested message, by bootstrap: the script builds a message descriptor by
    encoding a FileDescriptorProto with its own encoder and hands it to the
    reference library, which rejects it or produces the wrong fields if the
    encoding is wrong. That is also where the high tags and the packed field
    come from, the library's own bundled types having neither.

All 28 comparisons agree. `make check-encoding`, run by CI in both workflows.

google-protobuf is declared in the Gemfile. It arrives transitively through
the Metanorma toolchain, and an undeclared dependency has broken this build
twice already.

What remains of the serialization gap is the maximum message size, now a
minor row rather than a major one: the reference implementation applies 2 GB
to a single message and this document neither fixes that nor leaves it
explicitly to the implementation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD
@strimo378
strimo378 merged commit f0be2af into main Sep 21, 2026
3 checks passed
@strimo378
strimo378 deleted the claude/onnx-metanorma-standard-gloozz branch September 21, 2026 21:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant