From 6e439c72c6a7dbc9b1df25bc37a5c76965a69e2f Mon Sep 17 00:00:00 2001 From: Timo Stripf Date: Mon, 21 Sep 2026 15:39:02 +0000 Subject: [PATCH] Generate Annex B of Part 1 from the vendored Protocol Buffers schema MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit One of the three serialization gaps is derivable from the schema, and this is it: Annex B now states 35 messages and 166 fields as tables of field name, wire tag, type and obligation, generated by scripts/generate-schema.rb from upstream/onnx/proto/. It was empty. The obligation is the whole point. In the Protocol Buffers syntax this schema uses, every field is syntactically optional; which ones a producer must supply is carried in the comments, by the convention upstream's versioning document defines. The generator reads that convention — a field whose comment says it MUST be present for this version of the IR — and 23 of the 166 fields come out mandatory. A schema file cannot state that, and a standard has to. The prose of the schema comments is deliberately not carried across. A field table states structure; where a comment carries a normative statement it belongs in the clause it concerns, and after the previous commits the clauses are where those statements are. The schema source stays vendored as the informative aid the annex points at. Also emitted: the seven enumerations with their values, including the IR version history; the oneof groups, as a sentence saying exactly one member shall be present; and the reserved tags and names, as a sentence saying they shall not be used. The script is not a Protocol Buffers parser and does not pretend to be one. It reads the subset of proto2 these files use and raises on anything it does not recognize inside a message body rather than skipping it, so a schema change upstream fails the run rather than dropping a row. That caught two things while writing it: a top-level enum, and — the one that mattered — fields being attached to the last message created rather than to the message currently open, which silently moved every field of TensorShapeProto and TypeProto onto the nested message that preceded them. Both messages were missing from the output entirely and the counts still looked plausible. Cross-checked against the schema independently of the parser: 134 + 13 + 19 = 166 field declarations by grep, and 19 + 4 = 23 "MUST be present" comments. Both match the generator exactly. `make schema` regenerates, `make check-schema` fails if the committed file differs, and CI runs it in both workflows, on the same contract as the operator clauses. Clause 12.3 now points at the annex and states the tag rule. The Annex C row asking for this closes; a narrower one opens in its place, because the annex states the obligations at one IR version and the schema carries no history of them, so a model produced against an earlier IR version cannot be checked against it — which Clause 14.4 requires it can be. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01DeGJcqqgBdz7kUpH6EsYxD --- .github/workflows/metanorma.yml | 6 + .github/workflows/pages.yml | 6 + CONTRIBUTING.md | 32 + Makefile | 11 +- README.md | 1 + scripts/generate-schema.rb | 322 ++++++++ sources/part1/sections/12-serialization.adoc | 18 +- sources/part1/sections/14-versioning.adoc | 1 + .../sections/annex-b-protobuf-schema.adoc | 726 +++++++++++++++++- .../part1/sections/annex-c-known-gaps.adoc | 22 +- 10 files changed, 1116 insertions(+), 29 deletions(-) create mode 100755 scripts/generate-schema.rb diff --git a/.github/workflows/metanorma.yml b/.github/workflows/metanorma.yml index 263adcc..312a178 100644 --- a/.github/workflows/metanorma.yml +++ b/.github/workflows/metanorma.yml @@ -41,6 +41,12 @@ jobs: - name: Check generated operator clauses run: bundle exec make check-operators + # Annex B of Part 1 is generated from upstream/onnx/proto/. Same + # contract: a hand edit, or a refresh of the vendored schema without + # regenerating, fails here rather than drifting. + - name: Check generated schema annex + run: bundle exec make check-schema + - name: Check PDF stylesheet run: bundle exec ruby scripts/check-stylesheet.rb diff --git a/.github/workflows/pages.yml b/.github/workflows/pages.yml index 1330fb5..f5fcb8a 100644 --- a/.github/workflows/pages.yml +++ b/.github/workflows/pages.yml @@ -47,6 +47,12 @@ jobs: - name: Check generated operator clauses run: bundle exec make check-operators + # Annex B of Part 1 is generated from upstream/onnx/proto/. Same + # contract: a hand edit, or a refresh of the vendored schema without + # regenerating, fails here rather than drifting. + - name: Check generated schema annex + run: bundle exec make check-schema + - name: Check PDF stylesheet run: bundle exec ruby scripts/check-stylesheet.rb diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index f5684d2..6d69791 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -132,6 +132,38 @@ operator. That is reported on every run and recorded in Annex C of Part 1. It is the substance of the work remaining, and it cannot be generated — it has to be written per operator and agreed. +## The generated schema annex + +Annex B of Part 1 — 35 messages and 166 fields — is **generated** from the +vendored Protocol Buffers schema: + +```sh +make schema # regenerate from upstream/onnx/proto/ +make check-schema # fail if the committed file differs +``` + +Never edit `sources/part1/sections/annex-b-protobuf-schema.adoc` by hand. CI +runs `check-schema`, on the same contract as the operator clauses. + +The annex states, per message, each field's wire tag, type and obligation. The +obligation is the point of it: in the Protocol Buffers syntax this schema uses +every field is syntactically optional, and which ones a producer must supply is +carried in the comments, by the convention upstream's versioning document +defines. The generator reads that convention — a field whose comment says it +MUST be present for this version of the IR is mandatory — and 23 of the 166 +fields come out mandatory. + +The prose of the schema comments is deliberately **not** carried across. A +field table states structure; where a comment carries a normative statement, +that statement belongs in the clause it concerns, and the clauses are where +those statements are. The schema source stays vendored as an informative aid. + +`scripts/generate-schema.rb` is not a Protocol Buffers parser and is not meant +to be one. It reads the subset of proto2 these files use and **fails loudly** +on anything it does not recognize inside a message body, rather than skipping +it — so a schema change upstream shows up as a failed run rather than as a +missing row. + ## The vendored upstream copy `upstream/onnx/` is a verbatim copy of the ONNX documentation the draft diff --git a/Makefile b/Makefile index cd3e3d5..52f6d2b 100644 --- a/Makefile +++ b/Makefile @@ -29,7 +29,7 @@ RUBY ?= ruby SEVERITY ?= 1 .PHONY: all html doc pdf site lint clean deps check-stylesheet \ - operators check-operators + operators check-operators schema check-schema all: html @@ -70,6 +70,15 @@ operators: check-operators: @$(RUBY) scripts/generate-operators.rb --check +# Annex B of Part 1 is generated from the vendored Protocol Buffers schema: +# 35 messages and 166 fields, restated as tables of wire tag, type and +# obligation, which the schema source cannot express. +schema: + @$(RUBY) scripts/generate-schema.rb + +check-schema: + @$(RUBY) scripts/generate-schema.rb --check + # mn2pdf parses the stylesheet itself, and Metanorma exits 0 when that parse # fails — producing no PDF while the build looks successful. Check it first; # the check needs no fonts, so it also runs where a full PDF render cannot. diff --git a/README.md b/README.md index ea2c6ab..3efe459 100644 --- a/README.md +++ b/README.md @@ -54,6 +54,7 @@ sources/ scripts/check-errors.rb Gates the build on Metanorma diagnostic severity scripts/vendor-onnx-docs.sh Refreshes the vendored upstream copy scripts/generate-operators.rb Generates the operator clauses of Part 2 +scripts/generate-schema.rb Generates Annex B of Part 1 from the vendored .proto upstream/onnx/ Verbatim upstream ONNX docs (complete) + .proto .github/workflows/ CI: build HTML + PDF, publish as artifacts ``` diff --git a/scripts/generate-schema.rb b/scripts/generate-schema.rb new file mode 100755 index 0000000..435e94e --- /dev/null +++ b/scripts/generate-schema.rb @@ -0,0 +1,322 @@ +#!/usr/bin/env ruby +# frozen_string_literal: true + +# Generate Annex B of Part 1 from the vendored Protocol Buffers schema. +# +# scripts/generate-schema.rb [--check] +# +# Annex B owes a table of fields per message, giving wire tag, type and +# obligation. The schema source cannot state an obligation -- every proto2 +# field is syntactically optional -- so the obligation is read from the +# comment convention the upstream versioning document defines: a field whose +# leading comment says it MUST be present for this version of the IR is +# mandatory, and every other field is optional. A repeated field is neither; +# its obligation is its cardinality. +# +# `--check` regenerates into memory and fails if the file on disk differs, +# which is what CI runs. +# +# What this script deliberately does not carry across is the prose of the +# schema comments. A field table states structure; where a comment carries a +# normative statement it belongs in a clause, and the clauses of this part are +# where those statements now are. The schema source stays vendored as the +# informative aid Annex B points at. + +require "optparse" + +ROOT = File.expand_path("..", __dir__) + +# The ONNX-ML schema is the superset: it defines the map and sequence types +# and the ai.onnx.ml operator set that Part 2 specifies alongside the default +# domain. The plain onnx.proto is the same file with those omitted. +SOURCES = [ + { file: "upstream/onnx/proto/onnx-ml.proto", + title: "Model, graph and tensor messages" }, + { file: "upstream/onnx/proto/onnx-operators-ml.proto", + title: "Operator set messages" }, + { file: "upstream/onnx/proto/onnx-data.proto", + title: "Map and sequence messages" }, +].freeze + +TARGET = "sources/part1/sections/annex-b-protobuf-schema.adoc" + +MANDATORY = /MUST be present (?:in|for) this version of the IR/i.freeze + +Field = Struct.new(:label, :type, :name, :tag, :mandatory, :oneof, :deprecated) +Enum = Struct.new(:name, :values) +Msg = Struct.new(:name, :fields, :enums, :reserved) + +# --------------------------------------------------------------------- parse + +# A hand-rolled reader for the subset of proto2 these files use: messages, +# nested messages, enums, oneofs, reserved ranges and scalar fields. It is not +# a proto parser and is not meant to be one; it fails loudly on anything it +# does not recognise inside a message body rather than skipping it, so that a +# schema change upstream shows up as a failed run rather than a missing row. +def parse(path) + messages = [] + stack = [] # enclosing message names + comment = [] # comment lines gathered since the last statement + enum = nil + oneof = nil + depth_of_oneof = nil + + # The message currently being read, by its full path. `messages.last` is not + # it: once a nested message closes, the fields that follow belong to the + # enclosing message again, and appending them to the last entry created + # silently moved them onto the nested one. + current = lambda do + path_now = stack.join(".") + messages.reverse.find { |m| m.name == path_now } + end + + File.readlines(path, encoding: "UTF-8").each_with_index do |raw, i| + line = raw.strip + lineno = i + 1 + + if line.start_with?("//") + comment << line.sub(%r{\A//\s?}, "") + next + end + if line.empty? + comment.clear if stack.empty? && enum.nil? + next + end + + case line + when /\A(?:syntax|package|option|import)\b/ + comment.clear + when /\Amessage\s+(\w+)\s*\{/ + stack << Regexp.last_match(1) + messages << Msg.new(stack.join("."), [], [], []) + comment.clear + when /\Aenum\s+(\w+)\s*\{/ + enum = Enum.new([*stack, Regexp.last_match(1)].join("."), []) + # A top-level enum has no message to hang on; give it one of its own so + # that it is emitted in source order rather than dropped. + messages << Msg.new(enum.name, [], [], []) if stack.empty? + comment.clear + when /\Aoneof\s+(\w+)\s*\{/ + oneof = Regexp.last_match(1) + depth_of_oneof = stack.length + comment.clear + when /\Areserved\s+(.+);/ + owner = current.call or raise "#{path}:#{lineno}: reserved outside a message" + owner.reserved << Regexp.last_match(1).strip + comment.clear + when /\A\}\s*;?\z/ + if enum + # Nested enums belong to the message being read; a top-level enum + # belongs to the entry opened for it above, which is the last one. + owner = stack.empty? ? messages.last : current.call + owner or raise "#{path}:#{lineno}: enum has no owner" + owner.enums << enum + enum = nil + elsif oneof && depth_of_oneof == stack.length + oneof = nil + else + stack.pop or raise "#{path}:#{lineno}: unbalanced brace" + end + comment.clear + when /\A(\w+)\s*=\s*(0[xX][0-9a-fA-F]+|\d+)\s*;/ # enum member + enum or raise "#{path}:#{lineno}: enum member outside an enum: #{line}" + enum.values << [Regexp.last_match(1), Integer(Regexp.last_match(2))] + comment.clear + when /\A(optional|repeated|required)?\s*([\w.<>, ]+?)\s+(\w+)\s*=\s*(\d+)\s*(\[[^\]]*\])?\s*;/ + label = Regexp.last_match(1) || (oneof ? "oneof" : "optional") + type = Regexp.last_match(2).strip + name = Regexp.last_match(3) + tag = Regexp.last_match(4).to_i + text = comment.join(" ") + owner = current.call or raise "#{path}:#{lineno}: field outside a message" + owner.fields << Field.new( + label, type, name, tag, + text.match?(MANDATORY), oneof, + text.match?(/\bdeprecated\b/i) + ) + comment.clear + else + raise "#{path}:#{lineno}: unrecognised: #{line}" + end + end + + stack.empty? or raise "#{path}: unterminated message #{stack.inspect}" + messages.reject { |m| m.fields.empty? && m.enums.empty? } +end + +# --------------------------------------------------------------------- emit + +def obligation(field) + return "deprecated" if field.deprecated + return "repeated" if field.label == "repeated" + return "one of the group" if field.label == "oneof" + + field.mandatory ? "mandatory" : "optional" +end + +def adoc_type(type) + "`#{type}`" +end + +def anchor(name) + "schema-#{name.downcase.tr('.', '-')}" +end + +def render_message(msg) + out = [] + out << "[[#{anchor(msg.name)}]]" + out << "==== #{msg.name}" + out << "" + + unless msg.fields.empty? + out << "[[tbl-#{anchor(msg.name)}]]" + out << ".Fields of `#{msg.name}`" + out << '[cols="3,1,3,2"]' + out << "|===" + out << "| Field | Tag | Type | Obligation" + out << "" + msg.fields.each do |f| + out << "| `#{f.name}` | #{f.tag} | #{adoc_type(f.type)} | #{obligation(f)}" + end + out << "|===" + out << "" + end + + groups = msg.fields.map(&:oneof).compact.uniq + groups.each do |g| + members = msg.fields.select { |f| f.oneof == g }.map { |f| "`#{f.name}`" } + out << "Exactly one of #{members.join(', ')} SHALL be present, being the " \ + "group `#{g}`." + out << "" + end + + unless msg.reserved.empty? + out << "Tags and names reserved in `#{msg.name}`, which SHALL NOT be used: " \ + "#{msg.reserved.map { |r| "`#{r}`" }.join('; ')}." + out << "" + end + + msg.enums.each do |e| + out << "[[tbl-#{anchor(e.name)}]]" + out << ".Values of `#{e.name}`" + out << '[cols="3,1"]' + out << "|===" + out << "| Name | Value" + out << "" + e.values.each { |(n, v)| out << "| `#{n}` | #{v}" } + out << "|===" + out << "" + end + + out.join("\n") +end + +def render(groups) + counts = groups.sum { |g| g[:messages].length } + fields = groups.sum { |g| g[:messages].sum { |m| m.fields.length } } + mandatory = groups.sum do |g| + g[:messages].sum { |m| m.fields.count(&:mandatory) } + end + + out = [] + out << "// Generated by scripts/generate-schema.rb from upstream/onnx/proto/" + out << "// at the release recorded in upstream/onnx/SOURCE.txt." + out << "// Do not edit: run the script instead." + out << "// See CONTRIBUTING.md, \"The generated schema annex\"." + out << "" + out << "[[annex-schema]]" + out << "[appendix,obligation=normative]" + out << "== Protocol Buffers schema" + out << "" + out << "=== General" + out << "" + out << "This annex states the message definitions of the serialized form: " \ + "#{counts} messages and #{fields} fields." + out << "" + out << "The tag of a field is normative; it identifies the field on the " \ + "wire and SHALL NOT be reused. A field marked mandatory SHALL be " \ + "present. A field marked optional MAY be absent, and a " \ + "<> SHALL accept a message in which it is. A field marked " \ + "repeated holds zero or more values. A field marked deprecated " \ + "SHALL NOT be written by a <>, and a consumer that reads " \ + "one SHALL ignore it." + out << "" + out << "The obligations are those in force at the IR version stated in " \ + "<>." + out << "" + out << "NOTE: #{mandatory} of the #{fields} fields are mandatory. Protocol " \ + "Buffers cannot express the distinction: in the syntax this schema " \ + "uses every field is optional, and the obligation is carried in the " \ + "comments. Stating it in a table is the reason this annex exists." + out << "" + out << "[[tbl-schema-sources]]" + out << ".Subclauses of this annex and the schema file each restates" + out << '[cols="3,4,1"]' + out << "|===" + out << "| Subclause | Source | Messages" + out << "" + groups.each_with_index do |g, i| + out << "| B.#{i + 2} #{g[:title]} | `#{g[:file]}` | #{g[:messages].length}" + end + out << "|===" + out << "" + + groups.each do |g| + out << "=== #{g[:title]}" + out << "" + g[:messages].each { |m| out << render_message(m) } + end + + out << "=== Schema source" + out << "" + out << "The schema source is vendored in the repository at " \ + "`upstream/onnx/proto/`, at the release recorded in " \ + "`upstream/onnx/SOURCE.txt`. It is an informative aid: where a " \ + "comment in it carries a normative statement, that statement is in " \ + "the clause it belongs to and not in this annex." + out << "" + + out.join("\n").gsub(/\n{3,}/, "\n\n") +end + +# ---------------------------------------------------------------------- main + +check = false +OptionParser.new do |o| + o.banner = "usage: generate-schema.rb [--check]" + o.on("--check", "fail if the generated file differs from the one on disk") do + check = true + end +end.parse! + +groups = SOURCES.map do |src| + path = File.join(ROOT, src[:file]) + File.exist?(path) or abort "missing: #{src[:file]}" + { title: src[:title], file: src[:file], messages: parse(path) } +end + +text = render(groups) +target = File.join(ROOT, TARGET) + +if check + on_disk = File.exist?(target) ? File.read(target, encoding: "UTF-8") : nil + if on_disk == text + puts "check: Annex B is up to date" + else + warn "check: #{TARGET} differs from the generated output." + warn " Run `make schema` and commit the result." + exit 1 + end +else + File.write(target, text) + total_m = groups.sum { |g| g[:messages].length } + total_f = groups.sum { |g| g[:messages].sum { |m| m.fields.length } } + mand = groups.sum { |g| g[:messages].sum { |m| m.fields.count(&:mandatory) } } + enums = groups.sum { |g| g[:messages].sum { |m| m.enums.length } } + groups.each do |g| + puts format(" %-40s %3d messages", g[:file], g[:messages].length) + end + puts + puts format(" %d messages, %d fields (%d mandatory), %d enumerations -> %s", + total_m, total_f, mand, enums, TARGET) +end diff --git a/sources/part1/sections/12-serialization.adoc b/sources/part1/sections/12-serialization.adoc index 5ddca65..1b6573d 100644 --- a/sources/part1/sections/12-serialization.adoc +++ b/sources/part1/sections/12-serialization.adoc @@ -27,16 +27,16 @@ which is why <> is a separate question. === Message definitions -The message definitions are given normatively in Annex B. +The message definitions are given normatively in <>, as a table +of fields per message giving wire tag, type and obligation. -[IMPORTANT] -.Editorial note -==== -Annex B currently reproduces the schema as source. The intended form is a table -of fields per message, giving wire tag, type and obligation, with the schema -source retained as an informative aid. A table states each field's obligation, -which the schema cannot express, and survives a change of encoding. -==== +A field's tag identifies it on the wire. A tag SHALL NOT be reused for a +different field, and a tag recorded as reserved SHALL NOT be used at all. + +NOTE: The obligation is what the schema source cannot express, and is the +reason the annex is a table rather than the source. In the Protocol Buffers +syntax this schema uses, every field is syntactically optional; which ones a +producer must supply is carried in the comments. [[canonical-form]] === Canonical form diff --git a/sources/part1/sections/14-versioning.adoc b/sources/part1/sections/14-versioning.adoc index d93ed25..de9d06b 100644 --- a/sources/part1/sections/14-versioning.adoc +++ b/sources/part1/sections/14-versioning.adoc @@ -118,6 +118,7 @@ A model binds each of its nodes to an operator definition by the operator set version the model imports; see <>. How a consumer binds that definition to an implementation is outside the scope of this document. +[[compatibility]] === Compatibility obligations A consumer that claims support for operator set version stem:[n] of a domain diff --git a/sources/part1/sections/annex-b-protobuf-schema.adoc b/sources/part1/sections/annex-b-protobuf-schema.adoc index 473cb31..3c46184 100644 --- a/sources/part1/sections/annex-b-protobuf-schema.adoc +++ b/sources/part1/sections/annex-b-protobuf-schema.adoc @@ -1,15 +1,719 @@ +// Generated by scripts/generate-schema.rb from upstream/onnx/proto/ +// at the release recorded in upstream/onnx/SOURCE.txt. +// Do not edit: run the script instead. +// See CONTRIBUTING.md, "The generated schema annex". + [[annex-schema]] [appendix,obligation=normative] == Protocol Buffers schema -[IMPORTANT] -.Editorial note -==== -This annex is to carry the message definitions normatively, as a table of -fields per message giving wire tag, type and obligation, per the editorial note -in <>. It is empty pending the decision recorded there on the -normative status of the Protocol Buffers encoding. - -The schema source this annex will restate is vendored in the repository at -`upstream/onnx/proto/`, at the release recorded in its provenance. -==== +=== General + +This annex states the message definitions of the serialized form: 35 messages and 166 fields. + +The tag of a field is normative; it identifies the field on the wire and SHALL NOT be reused. A field marked mandatory SHALL be present. A field marked optional MAY be absent, and a <> SHALL accept a message in which it is. A field marked repeated holds zero or more values. A field marked deprecated SHALL NOT be written by a <>, and a consumer that reads one SHALL ignore it. + +The obligations are those in force at the IR version stated in <>. + +NOTE: 23 of the 166 fields are mandatory. Protocol Buffers cannot express the distinction: in the syntax this schema uses every field is optional, and the obligation is carried in the comments. Stating it in a table is the reason this annex exists. + +[[tbl-schema-sources]] +.Subclauses of this annex and the schema file each restates +[cols="3,4,1"] +|=== +| Subclause | Source | Messages + +| B.2 Model, graph and tensor messages | `upstream/onnx/proto/onnx-ml.proto` | 30 +| B.3 Operator set messages | `upstream/onnx/proto/onnx-operators-ml.proto` | 2 +| B.4 Map and sequence messages | `upstream/onnx/proto/onnx-data.proto` | 3 +|=== + +=== Model, graph and tensor messages + +[[schema-version]] +==== Version + +[[tbl-schema-version]] +.Values of `Version` +[cols="3,1"] +|=== +| Name | Value + +| `_START_VERSION` | 0 +| `IR_VERSION_2017_10_10` | 1 +| `IR_VERSION_2017_10_30` | 2 +| `IR_VERSION_2017_11_3` | 3 +| `IR_VERSION_2019_1_22` | 4 +| `IR_VERSION_2019_3_18` | 5 +| `IR_VERSION_2019_9_19` | 6 +| `IR_VERSION_2020_5_8` | 7 +| `IR_VERSION_2021_7_30` | 8 +| `IR_VERSION_2023_5_5` | 9 +| `IR_VERSION_2024_3_25` | 10 +| `IR_VERSION_2025_05_12` | 11 +| `IR_VERSION_2025_08_26` | 12 +| `IR_VERSION_2025_11_06` | 13 +| `IR_VERSION` | 14 +|=== + +[[schema-attributeproto]] +==== AttributeProto + +[[tbl-schema-attributeproto]] +.Fields of `AttributeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | mandatory +| `ref_attr_name` | 21 | `string` | optional +| `doc_string` | 13 | `string` | optional +| `type` | 20 | `AttributeType` | mandatory +| `f` | 2 | `float` | mandatory +| `i` | 3 | `int64` | optional +| `s` | 4 | `bytes` | optional +| `t` | 5 | `TensorProto` | optional +| `g` | 6 | `GraphProto` | optional +| `sparse_tensor` | 22 | `SparseTensorProto` | optional +| `tp` | 14 | `TypeProto` | deprecated +| `floats` | 7 | `float` | repeated +| `ints` | 8 | `int64` | repeated +| `strings` | 9 | `bytes` | repeated +| `tensors` | 10 | `TensorProto` | repeated +| `graphs` | 11 | `GraphProto` | repeated +| `sparse_tensors` | 23 | `SparseTensorProto` | repeated +| `type_protos` | 15 | `TypeProto` | repeated +|=== + +Tags and names reserved in `AttributeProto`, which SHALL NOT be used: `12, 16 to 19`; `"v"`. + +[[tbl-schema-attributeproto-attributetype]] +.Values of `AttributeProto.AttributeType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `FLOAT` | 1 +| `INT` | 2 +| `STRING` | 3 +| `TENSOR` | 4 +| `GRAPH` | 5 +| `SPARSE_TENSOR` | 11 +| `TYPE_PROTO` | 13 +| `FLOATS` | 6 +| `INTS` | 7 +| `STRINGS` | 8 +| `TENSORS` | 9 +| `GRAPHS` | 10 +| `SPARSE_TENSORS` | 12 +| `TYPE_PROTOS` | 14 +|=== + +[[schema-valueinfoproto]] +==== ValueInfoProto + +[[tbl-schema-valueinfoproto]] +.Fields of `ValueInfoProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | mandatory +| `type` | 2 | `TypeProto` | mandatory +| `doc_string` | 3 | `string` | optional +| `metadata_props` | 4 | `StringStringEntryProto` | repeated +|=== + +[[schema-nodeproto]] +==== NodeProto + +[[tbl-schema-nodeproto]] +.Fields of `NodeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `input` | 1 | `string` | repeated +| `output` | 2 | `string` | repeated +| `name` | 3 | `string` | optional +| `op_type` | 4 | `string` | optional +| `domain` | 7 | `string` | optional +| `overload` | 8 | `string` | optional +| `attribute` | 5 | `AttributeProto` | repeated +| `doc_string` | 6 | `string` | optional +| `metadata_props` | 9 | `StringStringEntryProto` | repeated +| `device_configurations` | 10 | `NodeDeviceConfigurationProto` | repeated +|=== + +[[schema-intintlistentryproto]] +==== IntIntListEntryProto + +[[tbl-schema-intintlistentryproto]] +.Fields of `IntIntListEntryProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `key` | 1 | `int64` | optional +| `value` | 2 | `int64` | repeated +|=== + +[[schema-nodedeviceconfigurationproto]] +==== NodeDeviceConfigurationProto + +[[tbl-schema-nodedeviceconfigurationproto]] +.Fields of `NodeDeviceConfigurationProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `configuration_id` | 1 | `string` | mandatory +| `sharding_spec` | 2 | `ShardingSpecProto` | repeated +| `pipeline_stage` | 3 | `int32` | optional +|=== + +[[schema-shardingspecproto]] +==== ShardingSpecProto + +[[tbl-schema-shardingspecproto]] +.Fields of `ShardingSpecProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `tensor_name` | 1 | `string` | mandatory +| `device` | 2 | `int64` | repeated +| `index_to_device_group_map` | 3 | `IntIntListEntryProto` | repeated +| `sharded_dim` | 4 | `ShardedDimProto` | repeated +|=== + +[[schema-shardeddimproto]] +==== ShardedDimProto + +[[tbl-schema-shardeddimproto]] +.Fields of `ShardedDimProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `axis` | 1 | `int64` | mandatory +| `simple_sharding` | 2 | `SimpleShardedDimProto` | repeated +|=== + +[[schema-simpleshardeddimproto]] +==== SimpleShardedDimProto + +[[tbl-schema-simpleshardeddimproto]] +.Fields of `SimpleShardedDimProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dim_value` | 1 | `int64` | one of the group +| `dim_param` | 2 | `string` | one of the group +| `num_shards` | 3 | `int64` | mandatory +|=== + +Exactly one of `dim_value`, `dim_param` SHALL be present, being the group `dim`. + +[[schema-traininginfoproto]] +==== TrainingInfoProto + +[[tbl-schema-traininginfoproto]] +.Fields of `TrainingInfoProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `initialization` | 1 | `GraphProto` | optional +| `algorithm` | 2 | `GraphProto` | optional +| `initialization_binding` | 3 | `StringStringEntryProto` | repeated +| `update_binding` | 4 | `StringStringEntryProto` | repeated +|=== + +[[schema-modelproto]] +==== ModelProto + +[[tbl-schema-modelproto]] +.Fields of `ModelProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `ir_version` | 1 | `int64` | optional +| `opset_import` | 8 | `OperatorSetIdProto` | repeated +| `producer_name` | 2 | `string` | optional +| `producer_version` | 3 | `string` | optional +| `domain` | 4 | `string` | optional +| `model_version` | 5 | `int64` | optional +| `doc_string` | 6 | `string` | optional +| `graph` | 7 | `GraphProto` | optional +| `metadata_props` | 14 | `StringStringEntryProto` | repeated +| `training_info` | 20 | `TrainingInfoProto` | repeated +| `functions` | 25 | `FunctionProto` | repeated +| `configuration` | 26 | `DeviceConfigurationProto` | repeated +|=== + +[[schema-deviceconfigurationproto]] +==== DeviceConfigurationProto + +[[tbl-schema-deviceconfigurationproto]] +.Fields of `DeviceConfigurationProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | mandatory +| `num_devices` | 2 | `int32` | mandatory +| `device` | 3 | `string` | repeated +|=== + +[[schema-stringstringentryproto]] +==== StringStringEntryProto + +[[tbl-schema-stringstringentryproto]] +.Fields of `StringStringEntryProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `key` | 1 | `string` | optional +| `value` | 2 | `string` | optional +|=== + +[[schema-tensorannotation]] +==== TensorAnnotation + +[[tbl-schema-tensorannotation]] +.Fields of `TensorAnnotation` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `tensor_name` | 1 | `string` | optional +| `quant_parameter_tensor_names` | 2 | `StringStringEntryProto` | repeated +|=== + +[[schema-graphproto]] +==== GraphProto + +[[tbl-schema-graphproto]] +.Fields of `GraphProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `node` | 1 | `NodeProto` | repeated +| `name` | 2 | `string` | optional +| `initializer` | 5 | `TensorProto` | repeated +| `sparse_initializer` | 15 | `SparseTensorProto` | repeated +| `doc_string` | 10 | `string` | optional +| `input` | 11 | `ValueInfoProto` | repeated +| `output` | 12 | `ValueInfoProto` | repeated +| `value_info` | 13 | `ValueInfoProto` | repeated +| `quantization_annotation` | 14 | `TensorAnnotation` | repeated +| `metadata_props` | 16 | `StringStringEntryProto` | repeated +|=== + +Tags and names reserved in `GraphProto`, which SHALL NOT be used: `3, 4, 6 to 9`; `"ir_version", "producer_version", "producer_tag", "domain"`. + +[[schema-tensorproto]] +==== TensorProto + +[[tbl-schema-tensorproto]] +.Fields of `TensorProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dims` | 1 | `int64` | repeated +| `data_type` | 2 | `int32` | optional +| `segment` | 3 | `Segment` | optional +| `float_data` | 4 | `float` | repeated +| `int32_data` | 5 | `int32` | repeated +| `string_data` | 6 | `bytes` | repeated +| `int64_data` | 7 | `int64` | repeated +| `name` | 8 | `string` | optional +| `doc_string` | 12 | `string` | optional +| `raw_data` | 9 | `bytes` | optional +| `external_data` | 13 | `StringStringEntryProto` | repeated +| `data_location` | 14 | `DataLocation` | optional +| `double_data` | 10 | `double` | repeated +| `uint64_data` | 11 | `uint64` | repeated +| `metadata_props` | 16 | `StringStringEntryProto` | repeated +|=== + +[[tbl-schema-tensorproto-datatype]] +.Values of `TensorProto.DataType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `FLOAT` | 1 +| `UINT8` | 2 +| `INT8` | 3 +| `UINT16` | 4 +| `INT16` | 5 +| `INT32` | 6 +| `INT64` | 7 +| `STRING` | 8 +| `BOOL` | 9 +| `FLOAT16` | 10 +| `DOUBLE` | 11 +| `UINT32` | 12 +| `UINT64` | 13 +| `COMPLEX64` | 14 +| `COMPLEX128` | 15 +| `BFLOAT16` | 16 +| `FLOAT8E4M3FN` | 17 +| `FLOAT8E4M3FNUZ` | 18 +| `FLOAT8E5M2` | 19 +| `FLOAT8E5M2FNUZ` | 20 +| `UINT4` | 21 +| `INT4` | 22 +| `FLOAT4E2M1` | 23 +| `FLOAT8E8M0` | 24 +| `UINT2` | 25 +| `INT2` | 26 +| `FLOAT6E2M3` | 27 +| `FLOAT6E3M2` | 28 +|=== + +[[tbl-schema-tensorproto-datalocation]] +.Values of `TensorProto.DataLocation` +[cols="3,1"] +|=== +| Name | Value + +| `DEFAULT` | 0 +| `EXTERNAL` | 1 +|=== + +[[schema-tensorproto-segment]] +==== TensorProto.Segment + +[[tbl-schema-tensorproto-segment]] +.Fields of `TensorProto.Segment` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `begin` | 1 | `int64` | optional +| `end` | 2 | `int64` | optional +|=== + +[[schema-sparsetensorproto]] +==== SparseTensorProto + +[[tbl-schema-sparsetensorproto]] +.Fields of `SparseTensorProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `values` | 1 | `TensorProto` | optional +| `indices` | 2 | `TensorProto` | optional +| `dims` | 3 | `int64` | repeated +|=== + +[[schema-tensorshapeproto]] +==== TensorShapeProto + +[[tbl-schema-tensorshapeproto]] +.Fields of `TensorShapeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dim` | 1 | `Dimension` | repeated +|=== + +[[schema-tensorshapeproto-dimension]] +==== TensorShapeProto.Dimension + +[[tbl-schema-tensorshapeproto-dimension]] +.Fields of `TensorShapeProto.Dimension` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dim_value` | 1 | `int64` | one of the group +| `dim_param` | 2 | `string` | one of the group +| `denotation` | 3 | `string` | optional +|=== + +Exactly one of `dim_value`, `dim_param` SHALL be present, being the group `value`. + +[[schema-typeproto]] +==== TypeProto + +[[tbl-schema-typeproto]] +.Fields of `TypeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `tensor_type` | 1 | `Tensor` | one of the group +| `sequence_type` | 4 | `Sequence` | one of the group +| `map_type` | 5 | `Map` | one of the group +| `optional_type` | 9 | `Optional` | one of the group +| `sparse_tensor_type` | 8 | `SparseTensor` | one of the group +| `opaque_type` | 7 | `Opaque` | one of the group +| `denotation` | 6 | `string` | optional +|=== + +Exactly one of `tensor_type`, `sequence_type`, `map_type`, `optional_type`, `sparse_tensor_type`, `opaque_type` SHALL be present, being the group `value`. + +[[schema-typeproto-tensor]] +==== TypeProto.Tensor + +[[tbl-schema-typeproto-tensor]] +.Fields of `TypeProto.Tensor` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `int32` | mandatory +| `shape` | 2 | `TensorShapeProto` | optional +|=== + +[[schema-typeproto-sequence]] +==== TypeProto.Sequence + +[[tbl-schema-typeproto-sequence]] +.Fields of `TypeProto.Sequence` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `TypeProto` | mandatory +|=== + +[[schema-typeproto-map]] +==== TypeProto.Map + +[[tbl-schema-typeproto-map]] +.Fields of `TypeProto.Map` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `key_type` | 1 | `int32` | mandatory +| `value_type` | 2 | `TypeProto` | mandatory +|=== + +[[schema-typeproto-optional]] +==== TypeProto.Optional + +[[tbl-schema-typeproto-optional]] +.Fields of `TypeProto.Optional` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `TypeProto` | mandatory +|=== + +[[schema-typeproto-sparsetensor]] +==== TypeProto.SparseTensor + +[[tbl-schema-typeproto-sparsetensor]] +.Fields of `TypeProto.SparseTensor` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `int32` | mandatory +| `shape` | 2 | `TensorShapeProto` | optional +|=== + +[[schema-typeproto-opaque]] +==== TypeProto.Opaque + +[[tbl-schema-typeproto-opaque]] +.Fields of `TypeProto.Opaque` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `domain` | 1 | `string` | optional +| `name` | 2 | `string` | optional +|=== + +Tags and names reserved in `TypeProto.Opaque`, which SHALL NOT be used: `3`; `"parameters"`. + +[[schema-operatorsetidproto]] +==== OperatorSetIdProto + +[[tbl-schema-operatorsetidproto]] +.Fields of `OperatorSetIdProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `domain` | 1 | `string` | mandatory +| `version` | 2 | `int64` | mandatory +|=== + +[[schema-operatorstatus]] +==== OperatorStatus + +[[tbl-schema-operatorstatus]] +.Values of `OperatorStatus` +[cols="3,1"] +|=== +| Name | Value + +| `EXPERIMENTAL` | 0 +| `STABLE` | 1 +|=== + +[[schema-functionproto]] +==== FunctionProto + +[[tbl-schema-functionproto]] +.Fields of `FunctionProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `input` | 4 | `string` | repeated +| `output` | 5 | `string` | repeated +| `attribute` | 6 | `string` | repeated +| `attribute_proto` | 11 | `AttributeProto` | repeated +| `node` | 7 | `NodeProto` | repeated +| `doc_string` | 8 | `string` | optional +| `opset_import` | 9 | `OperatorSetIdProto` | repeated +| `domain` | 10 | `string` | optional +| `overload` | 13 | `string` | optional +| `value_info` | 12 | `ValueInfoProto` | repeated +| `metadata_props` | 14 | `StringStringEntryProto` | repeated +|=== + +Tags and names reserved in `FunctionProto`, which SHALL NOT be used: `2`; `"since_version"`; `3`; `"status"`. + +=== Operator set messages + +[[schema-operatorproto]] +==== OperatorProto + +[[tbl-schema-operatorproto]] +.Fields of `OperatorProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `op_type` | 1 | `string` | mandatory +| `since_version` | 2 | `int64` | mandatory +| `status` | 3 | `OperatorStatus` | optional +| `doc_string` | 10 | `string` | optional +|=== + +[[schema-operatorsetproto]] +==== OperatorSetProto + +[[tbl-schema-operatorsetproto]] +.Fields of `OperatorSetProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `magic` | 1 | `string` | mandatory +| `ir_version` | 2 | `int64` | mandatory +| `ir_version_prerelease` | 3 | `string` | optional +| `ir_build_metadata` | 7 | `string` | optional +| `domain` | 4 | `string` | optional +| `opset_version` | 5 | `int64` | optional +| `doc_string` | 6 | `string` | optional +| `operator` | 8 | `OperatorProto` | repeated +| `functions` | 9 | `FunctionProto` | repeated +|=== + +=== Map and sequence messages + +[[schema-sequenceproto]] +==== SequenceProto + +[[tbl-schema-sequenceproto]] +.Fields of `SequenceProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `elem_type` | 2 | `int32` | optional +| `tensor_values` | 3 | `TensorProto` | repeated +| `sparse_tensor_values` | 4 | `SparseTensorProto` | repeated +| `sequence_values` | 5 | `SequenceProto` | repeated +| `map_values` | 6 | `MapProto` | repeated +| `optional_values` | 7 | `OptionalProto` | repeated +|=== + +[[tbl-schema-sequenceproto-datatype]] +.Values of `SequenceProto.DataType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `TENSOR` | 1 +| `SPARSE_TENSOR` | 2 +| `SEQUENCE` | 3 +| `MAP` | 4 +| `OPTIONAL` | 5 +|=== + +[[schema-mapproto]] +==== MapProto + +[[tbl-schema-mapproto]] +.Fields of `MapProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `key_type` | 2 | `int32` | optional +| `keys` | 3 | `int64` | repeated +| `string_keys` | 4 | `bytes` | repeated +| `values` | 5 | `SequenceProto` | optional +|=== + +[[schema-optionalproto]] +==== OptionalProto + +[[tbl-schema-optionalproto]] +.Fields of `OptionalProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `elem_type` | 2 | `int32` | optional +| `tensor_value` | 3 | `TensorProto` | optional +| `sparse_tensor_value` | 4 | `SparseTensorProto` | optional +| `sequence_value` | 5 | `SequenceProto` | optional +| `map_value` | 6 | `MapProto` | optional +| `optional_value` | 7 | `OptionalProto` | optional +|=== + +[[tbl-schema-optionalproto-datatype]] +.Values of `OptionalProto.DataType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `TENSOR` | 1 +| `SPARSE_TENSOR` | 2 +| `SEQUENCE` | 3 +| `MAP` | 4 +| `OPTIONAL` | 5 +|=== + +=== Schema source + +The schema source is vendored in the repository at `upstream/onnx/proto/`, at the release recorded in `upstream/onnx/SOURCE.txt`. It is an informative aid: where a comment in it carries a normative statement, that statement is in the clause it belongs to and not in this annex. diff --git a/sources/part1/sections/annex-c-known-gaps.adoc b/sources/part1/sections/annex-c-known-gaps.adoc index 947ab0b..6130a35 100644 --- a/sources/part1/sections/annex-c-known-gaps.adoc +++ b/sources/part1/sections/annex-c-known-gaps.adoc @@ -38,13 +38,6 @@ form a standard cannot adopt. testable conditions with severities. | Blocking -| Annex B -| The Protocol Buffers schema is not yet restated normatively as a table of - fields with obligations. <> now requires that annex to identify - the fields each IR version makes mandatory, which the schema source cannot - express. -| Blocking - | <> Clauses 5, 6 | The generated operator clauses supply shape inference, determinism and error conditions for 0 of 222 operators: the upstream documentation states none of @@ -103,7 +96,17 @@ form a standard cannot adopt. | Clause 12 (encoding) | The admissible subset of the Protocol Buffers wire format --- unknown fields, - repeated field ordering, maximum message size --- is unstated. + repeated field ordering, maximum message size --- is unstated. It cannot be + derived from the schema, which says nothing about any of the three; the + reference implementation enforces a maximum message size in its checker, + whose source is not part of the vendored baseline. +| Major + +| <> +| The annex states which fields are mandatory at one IR version, being the one + <> names. The schema source carries no history of the + obligations, so a model produced against an earlier IR version cannot be + checked against this annex, and <> requires that it can be. | Major | <> @@ -174,6 +177,9 @@ transcription rather than take it on trust, and so that a refresh of | <> (typed storage) | `onnx-ml.proto`, `TensorProto.int32_data` +| <> (the whole annex) +| `upstream/onnx/proto/`, generated by `scripts/generate-schema.rb` + | <> (the IR version history) | `docs/Versioning.md`, "Released Versions"; `onnx-ml.proto`, `Version`