Skip to content

Generate Annex B of Part 1 from the vendored Protocol Buffers schema - #10

Merged
strimo378 merged 1 commit into
mainfrom
claude/onnx-annex-b-schema
Sep 21, 2026
Merged

strimo378 merged 1 commit into
mainfrom
claude/onnx-annex-b-schema

Conversation

@strimo378

@strimo378 strimo378 commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

2 of 3. #9 is merged, so this now sits directly on main — one commit, no conflict. #11 builds on it.

One of the three serialization gaps is derivable from the schema, and this is it. Annex B now states 35 messages and 166 fields as tables of field name, wire tag, type and obligation, generated by scripts/generate-schema.rb from upstream/onnx/proto/. It was empty.

The obligation is the whole point

In the Protocol Buffers syntax this schema uses, every field is syntactically optional. Which ones a producer must supply is carried in the comments, by the convention upstream's versioning document defines:

// This field MUST be present in this version of the IR.
optional string name = 1;

The generator reads that convention, and 23 of the 166 fields come out mandatory. A schema file cannot state that, and a standard has to.

The prose of the schema comments is deliberately not carried across. A field table states structure; where a comment carries a normative statement it belongs in the clause it concerns, and after the previous PRs the clauses are where those statements are. The schema source stays vendored as the informative aid the annex points at.

Also emitted: the seven enumerations with their values including the IR version history; the oneof groups, as a sentence saying exactly one member shall be present; and the reserved tags and names, as a sentence saying they shall not be used.

The parser fails loudly, and that mattered

It is not a Protocol Buffers parser and does not pretend to be one. It reads the subset of proto2 these files use and raises on anything it does not recognize inside a message body rather than skipping it, so a schema change upstream fails the run rather than dropping a row.

That caught two things while writing it. The harmless one was a top-level enum. The other was mine: fields were being attached to the last message created rather than to the message currently open, so once a nested message closed, every field that followed moved onto it. TensorShapeProto and TypeProto were missing from the output entirely — and the counts still looked plausible.

Hence two cross-checks, independent of the parser, by grep:

field declarations 134 + 13 + 19 = 166 ✅
MUST be present comments 19 + 4 = 23 ✅

Both match the generator exactly.

Contract

make schema        # regenerate from upstream/onnx/proto/
make check-schema  # fail if the committed file differs

CI runs check-schema in both workflows, on the same contract as the operator clauses.

Clause changes

Clause 12.3 now points at the annex and states the tag rule. The Annex C row asking for this closes; a narrower one opens in its place, because the annex states the obligations at one IR version and the schema carries no history of them — so a model produced against an earlier IR version cannot be checked against it, which Clause 14.4 requires it can be.

What is still not derivable

The other two thirds of the serialization gap. The wire subset (unknown fields, repeated field ordering, maximum message size) is not in the schema at all; the reference implementation enforces a maximum message size in its checker, whose source is not part of the vendored baseline. A citable Protocol Buffers specification is not a derivation question at all. Annex C says so now.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD

@strimo378
strimo378 force-pushed the claude/onnx-annex-b-schema branch from 9465b31 to 5c3e633 Compare September 21, 2026 19:43
@strimo378
strimo378 force-pushed the claude/onnx-ir-version-packing branch from 238c789 to 827d0de Compare September 21, 2026 19:45
@strimo378
strimo378 force-pushed the claude/onnx-annex-b-schema branch from 5c3e633 to d5f932e Compare September 21, 2026 19:45
strimo378 added a commit that referenced this pull request Sep 21, 2026
…ings (#9)

**1 of 3.** This is the base of a stack; #10 and #11 build on it. Merge
this one first.

> **Correction.** The first version of this PR claimed IR version 14 is
unpublished and recorded it in Annex C as blocking. That was wrong — see
the last section. The claim is gone; the clause now states the fact.

Two things, both from checking how new the 6-bit floating-point types
are.

## Clause 14 now names the IR version this edition specifies

`FLOAT6E2M3` and `FLOAT6E3M2` are new in ONNX 1.23.0, the vendored
baseline, and arrive with **IR version 14**. Part 1 did not say anywhere
which IR version it specifies — an omission for a document whose Clause
14 is about version axes.

It now does, with the history as a table giving the release each version
was published in:

| IR | ONNX release | Introduced |
|---|---|---|
| 12 | 1.19.0 | `FLOAT8E8M0` |
| 13 | 1.20.0 | `UINT2`, `INT2` |
| 14 | 1.23.0 | `FLOAT6E2M3`, `FLOAT6E3M2`; the opaque type outside the
ONNX-ML build |

## The packing subclause was wrong

A defect in what the previous PR committed. A tensor carries its
elements either as an octet string or in a field typed for them, and
**the two pack a narrow type differently**. For 4-bit and 2-bit types
the packing coincides, which is why the single statement looked right.
For the 6-bit types it does not:

```
// For FLOAT6E2M3 and FLOAT6E3M2, each `int32_data` entry stores one
//   element's unsigned 6-bit encoding in bits 0-5; bits 6-31 MUST be zero.
```

One octet per element — not a packing at all. The subclause now states
the two separately, with a table for typed storage, and says a producer
SHOULD use raw storage for a 6-bit type. Also stated there, which I had
missed: a floating-point element narrower than 32 bits goes into typed
storage as the unsigned integer holding its bits.

## The correction

The first version of this PR asserted, in a normative clause, that IR
version 14 is not published, and raised it in Annex C as blocking — on
the strength of a comment in the schema:

```
// IR VERSION 14 published on TBD
```

I took a date field nobody had filled in at release time for a statement
that the version does not exist. The evidence against it was in the same
vendored baseline, in the very file Clause 14 was written from —
`docs/Versioning.md`, under the heading **"Released Versions"**:

```
1.22.0 | 13 | 27 | 5 | 1
1.23.0 | 14 | 28 | 5 | 1
```

A release is the publication event; a comment is not. The clause now
states that this edition specifies IR version 14 as published in ONNX
release 1.23.0, the history table gives releases rather than the
schema's dates, and a note records that the schema still reads TBD so
the next reader is not misled the same way. The Annex C row is removed
and nothing replaces it.

The correction is its own commit, so the mistake and its fix are both in
the record.

## Counts

Nineteen gaps close; Part 1 carries 24 editorial notes, down from 34.
Two new notes came out of reading the source closely, both in the
narrow-type conversion rules.

## Verification

All three parts compile with "Syntax Valid", clean at severity 3 only.
The rendered packing subclause and the version table were read back from
the HTML.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Base automatically changed from claude/onnx-ir-version-packing to main September 21, 2026 20:15
One of the three serialization gaps is derivable from the schema, and this
is it: Annex B now states 35 messages and 166 fields as tables of field
name, wire tag, type and obligation, generated by
scripts/generate-schema.rb from upstream/onnx/proto/. It was empty.

The obligation is the whole point. In the Protocol Buffers syntax this
schema uses, every field is syntactically optional; which ones a producer
must supply is carried in the comments, by the convention upstream's
versioning document defines. The generator reads that convention — a field
whose comment says it MUST be present for this version of the IR — and 23
of the 166 fields come out mandatory. A schema file cannot state that, and
a standard has to.

The prose of the schema comments is deliberately not carried across. A
field table states structure; where a comment carries a normative statement
it belongs in the clause it concerns, and after the previous commits the
clauses are where those statements are. The schema source stays vendored as
the informative aid the annex points at.

Also emitted: the seven enumerations with their values, including the IR
version history; the oneof groups, as a sentence saying exactly one member
shall be present; and the reserved tags and names, as a sentence saying they
shall not be used.

The script is not a Protocol Buffers parser and does not pretend to be one.
It reads the subset of proto2 these files use and raises on anything it does
not recognize inside a message body rather than skipping it, so a schema
change upstream fails the run rather than dropping a row. That caught two
things while writing it: a top-level enum, and — the one that mattered —
fields being attached to the last message created rather than to the message
currently open, which silently moved every field of TensorShapeProto and
TypeProto onto the nested message that preceded them. Both messages were
missing from the output entirely and the counts still looked plausible.

Cross-checked against the schema independently of the parser: 134 + 13 + 19
= 166 field declarations by grep, and 19 + 4 = 23 "MUST be present"
comments. Both match the generator exactly.

`make schema` regenerates, `make check-schema` fails if the committed file
differs, and CI runs it in both workflows, on the same contract as the
operator clauses.

Clause 12.3 now points at the annex and states the tag rule. The Annex C row
asking for this closes; a narrower one opens in its place, because the annex
states the obligations at one IR version and the schema carries no history of
them, so a model produced against an earlier IR version cannot be checked
against it — which Clause 14.4 requires it can be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD
@strimo378
strimo378 force-pushed the claude/onnx-annex-b-schema branch from d5f932e to 6e439c7 Compare September 21, 2026 20:17
@strimo378
strimo378 merged commit fcdee32 into main Sep 21, 2026
3 checks passed
@strimo378
strimo378 deleted the claude/onnx-annex-b-schema branch September 21, 2026 20:22
strimo378 added a commit that referenced this pull request Sep 21, 2026
**3 of 3.** #9 and #10 are merged, so this now sits directly on `main` —
one commit, no conflict.

Every reference to a term of Clause 3 rendered as a **clause number** in
the published document. The definition of "model" read:

> self-contained description of a computation, comprising a top-level
**Clause 3.2**, the **Clause 3.10** versions it depends upon, and
producer and metadata information

That is the whole of Part 1: **65 references** across the clauses, each
one the word the sentence needs replaced by a number. The published
HTML, DOC and PDF have carried it since the first draft.

## Cause

A bare cross-reference to a term renders as a clause number.
`:xrefstyle: short` is **not** involved — I checked that before changing
anything, by removing it and rebuilding; the numbers stayed. Giving the
reference its text renders the word and keeps the link, which is what
`<<node,nodes>>` was already doing in the two places it appeared. So all
65 now carry it, plus the two of Part 3.

Metanorma's concept syntax `{{model}}` was the other candidate. It
renders `model (3.1)`, which is correct ISO style for a first use and
noise on the twelfth, so it is not used here.

`scripts/generate-schema.rb` emits two such references and now follows
the same convention — or the next `make check-schema` would have failed
on a file nobody edited.

## Also

Metanorma opens a terms clause with "For the purposes of this document,
the following terms and definitions apply." itself, and Parts 1 and 3
said it again underneath. Part 1's copy is removed; Part 3's is reworded
to say only what the automatic one does not — that the terms of Part 1
also apply.

## Verification

Read back from the rendered HTML of all three parts: **no reference to
Clause 3 renders as a number any more**, and the opening sentence
appears once.

| | before | after |
|---|---|---|
| `Clause 3.x` in rendered text | 65 | 0 |
| duplicated opening sentence | Parts 1, 3 | none |

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant