Repository navigation
Name the IR version this edition specifies, and separate the two packings - #9
Merged
Merged
Conversation
…fies Two findings from checking how new the 6-bit floating-point types are. The 6-bit types are new in ONNX 1.23.0, the vendored baseline, and they arrive under IR version 14 — which the upstream schema carries as current with its publication date given as "TBD". IR version 14 is not published. Three things this document specifies exist only in it: FLOAT6E2M3, FLOAT6E3M2, and the opaque type outside the ONNX-ML build. A standard cannot specify an unpublished version of its own subject. Clause 14 now states which IR version this edition specifies, carries the version history as a table, and says plainly that the choice — hold the three for an amendment and specify IR version 13, or publish 14 first — is the Steering Committee's and should not fall out of whichever ONNX release happened to be vendored. Annex C records it as blocking. The second finding is a defect in the packing subclause as committed. A tensor carries its elements either as an octet string or in a field typed for them, and the two pack a narrow type differently. For 4-bit and 2-bit types the packing coincides, which is why the single statement looked right; for the 6-bit types it does not. Typed storage puts one 6-bit element in the low six bits of a 32-bit integer and requires the rest to be zero, so it costs at least an octet per element and is not a packing at all. The subclause now states the two separately, with a table for typed storage, and says a producer should use raw storage for a 6-bit type. Also stated there: a floating-point element narrower than 32 bits goes into typed storage as the unsigned integer holding its bits. Twenty gaps now close rather than eighteen; Part 1 carries 25 editorial notes rather than 24, the IR version note being the third that came out of reading the source rather than skimming it. Verified: all three parts compile with "Syntax Valid", clean at severity 3 only. The rendered packing subclause was read back. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD
The previous commit claimed IR version 14 is unpublished and recorded it in
Annex C as a blocking gap. That is wrong, and the evidence against it was in
the same vendored baseline, in the file this part's versioning clause was
written from:
docs/Versioning.md, "Released Versions"
...
1.22.0 | 13 | 27 | 5 | 1
1.23.0 | 14 | 28 | 5 | 1
IR version 14 shipped in ONNX release 1.23.0, which is the release this
document is drafted against. There is no gap.
What I read instead was a comment in the schema, "IR VERSION 14 published on
TBD", and took a date field nobody had filled in at release time for a
statement that the version does not exist. A release is the publication
event; a comment is not. I asserted the opposite in a normative clause and
raised it as blocking, on a source I had not checked against the one table
that answers the question directly.
Clause 14.1.2 now states that this edition specifies IR version 14 as
published in ONNX release 1.23.0, and the version history table gives the
release each version was published in rather than the schema's date, which
is not maintained. A note records that the schema still reads TBD for 14, so
that a reader who finds it is not misled the same way.
The Annex C row is removed; nothing replaces it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD
strimo378
force-pushed
the
claude/onnx-ir-version-packing
branch
from
September 21, 2026 19:45
238c789 to
827d0de
Compare
strimo378
added a commit
that referenced
this pull request
Sep 21, 2026
…10) **2 of 3.** #9 is merged, so this now sits directly on `main` — one commit, no conflict. #11 builds on it. One of the three serialization gaps is derivable from the schema, and this is it. Annex B now states **35 messages and 166 fields** as tables of field name, wire tag, type and obligation, generated by `scripts/generate-schema.rb` from `upstream/onnx/proto/`. It was empty. ## The obligation is the whole point In the Protocol Buffers syntax this schema uses, every field is syntactically optional. Which ones a producer must supply is carried in the comments, by the convention upstream's versioning document defines: ```proto // This field MUST be present in this version of the IR. optional string name = 1; ``` The generator reads that convention, and **23 of the 166** fields come out mandatory. A schema file cannot state that, and a standard has to. The prose of the schema comments is deliberately **not** carried across. A field table states structure; where a comment carries a normative statement it belongs in the clause it concerns, and after the previous PRs the clauses are where those statements are. The schema source stays vendored as the informative aid the annex points at. Also emitted: the seven enumerations with their values including the IR version history; the `oneof` groups, as a sentence saying exactly one member shall be present; and the reserved tags and names, as a sentence saying they shall not be used. ## The parser fails loudly, and that mattered It is not a Protocol Buffers parser and does not pretend to be one. It reads the subset of proto2 these files use and **raises** on anything it does not recognize inside a message body rather than skipping it, so a schema change upstream fails the run rather than dropping a row. That caught two things while writing it. The harmless one was a top-level enum. The other was mine: fields were being attached to the *last message created* rather than to the *message currently open*, so once a nested message closed, every field that followed moved onto it. `TensorShapeProto` and `TypeProto` were **missing from the output entirely** — and the counts still looked plausible. Hence two cross-checks, independent of the parser, by grep: | | | |---|---| | field declarations | 134 + 13 + 19 = **166** ✅ | | `MUST be present` comments | 19 + 4 = **23** ✅ | Both match the generator exactly. ## Contract ```sh make schema # regenerate from upstream/onnx/proto/ make check-schema # fail if the committed file differs ``` CI runs `check-schema` in both workflows, on the same contract as the operator clauses. ## Clause changes Clause 12.3 now points at the annex and states the tag rule. The Annex C row asking for this closes; a narrower one opens in its place, because the annex states the obligations at **one** IR version and the schema carries no history of them — so a model produced against an earlier IR version cannot be checked against it, which Clause 14.4 requires it can be. ## What is still not derivable The other two thirds of the serialization gap. The **wire subset** (unknown fields, repeated field ordering, maximum message size) is not in the schema at all; the reference implementation enforces a maximum message size in its checker, whose source is not part of the vendored baseline. A **citable Protocol Buffers specification** is not a derivation question at all. Annex C says so now. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
strimo378
added a commit
that referenced
this pull request
Sep 21, 2026
**3 of 3.** #9 and #10 are merged, so this now sits directly on `main` — one commit, no conflict. Every reference to a term of Clause 3 rendered as a **clause number** in the published document. The definition of "model" read: > self-contained description of a computation, comprising a top-level **Clause 3.2**, the **Clause 3.10** versions it depends upon, and producer and metadata information That is the whole of Part 1: **65 references** across the clauses, each one the word the sentence needs replaced by a number. The published HTML, DOC and PDF have carried it since the first draft. ## Cause A bare cross-reference to a term renders as a clause number. `:xrefstyle: short` is **not** involved — I checked that before changing anything, by removing it and rebuilding; the numbers stayed. Giving the reference its text renders the word and keeps the link, which is what `<<node,nodes>>` was already doing in the two places it appeared. So all 65 now carry it, plus the two of Part 3. Metanorma's concept syntax `{{model}}` was the other candidate. It renders `model (3.1)`, which is correct ISO style for a first use and noise on the twelfth, so it is not used here. `scripts/generate-schema.rb` emits two such references and now follows the same convention — or the next `make check-schema` would have failed on a file nobody edited. ## Also Metanorma opens a terms clause with "For the purposes of this document, the following terms and definitions apply." itself, and Parts 1 and 3 said it again underneath. Part 1's copy is removed; Part 3's is reworded to say only what the automatic one does not — that the terms of Part 1 also apply. ## Verification Read back from the rendered HTML of all three parts: **no reference to Clause 3 renders as a number any more**, and the opening sentence appears once. | | before | after | |---|---|---| | `Clause 3.x` in rendered text | 65 | 0 | | duplicated opening sentence | Parts 1, 3 | none | 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
1 of 3. This is the base of a stack; #10 and #11 build on it. Merge this one first.
Two things, both from checking how new the 6-bit floating-point types are.
Clause 14 now names the IR version this edition specifies
FLOAT6E2M3andFLOAT6E3M2are new in ONNX 1.23.0, the vendored baseline, and arrive with IR version 14. Part 1 did not say anywhere which IR version it specifies — an omission for a document whose Clause 14 is about version axes.It now does, with the history as a table giving the release each version was published in:
FLOAT8E8M0UINT2,INT2FLOAT6E2M3,FLOAT6E3M2; the opaque type outside the ONNX-ML buildThe packing subclause was wrong
A defect in what the previous PR committed. A tensor carries its elements either as an octet string or in a field typed for them, and the two pack a narrow type differently. For 4-bit and 2-bit types the packing coincides, which is why the single statement looked right. For the 6-bit types it does not:
One octet per element — not a packing at all. The subclause now states the two separately, with a table for typed storage, and says a producer SHOULD use raw storage for a 6-bit type. Also stated there, which I had missed: a floating-point element narrower than 32 bits goes into typed storage as the unsigned integer holding its bits.
The correction
The first version of this PR asserted, in a normative clause, that IR version 14 is not published, and raised it in Annex C as blocking — on the strength of a comment in the schema:
I took a date field nobody had filled in at release time for a statement that the version does not exist. The evidence against it was in the same vendored baseline, in the very file Clause 14 was written from —
docs/Versioning.md, under the heading "Released Versions":A release is the publication event; a comment is not. The clause now states that this edition specifies IR version 14 as published in ONNX release 1.23.0, the history table gives releases rather than the schema's dates, and a note records that the schema still reads TBD so the next reader is not misled the same way. The Annex C row is removed and nothing replaces it.
The correction is its own commit, so the mistake and its fix are both in the record.
Counts
Nineteen gaps close; Part 1 carries 24 editorial notes, down from 34. Two new notes came out of reading the source closely, both in the narrow-type conversion rules.
Verification
All three parts compile with "Syntax Valid", clean at severity 3 only. The rendered packing subclause and the version table were read back from the HTML.
🤖 Generated with Claude Code
https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD