diff --git a/.github/workflows/metanorma.yml b/.github/workflows/metanorma.yml index 263adcc..312a178 100644 --- a/.github/workflows/metanorma.yml +++ b/.github/workflows/metanorma.yml @@ -41,6 +41,12 @@ jobs: - name: Check generated operator clauses run: bundle exec make check-operators + # Annex B of Part 1 is generated from upstream/onnx/proto/. Same + # contract: a hand edit, or a refresh of the vendored schema without + # regenerating, fails here rather than drifting. + - name: Check generated schema annex + run: bundle exec make check-schema + - name: Check PDF stylesheet run: bundle exec ruby scripts/check-stylesheet.rb diff --git a/.github/workflows/pages.yml b/.github/workflows/pages.yml index 1330fb5..f5fcb8a 100644 --- a/.github/workflows/pages.yml +++ b/.github/workflows/pages.yml @@ -47,6 +47,12 @@ jobs: - name: Check generated operator clauses run: bundle exec make check-operators + # Annex B of Part 1 is generated from upstream/onnx/proto/. Same + # contract: a hand edit, or a refresh of the vendored schema without + # regenerating, fails here rather than drifting. + - name: Check generated schema annex + run: bundle exec make check-schema + - name: Check PDF stylesheet run: bundle exec ruby scripts/check-stylesheet.rb diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index f5684d2..6d69791 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -132,6 +132,38 @@ operator. That is reported on every run and recorded in Annex C of Part 1. It is the substance of the work remaining, and it cannot be generated — it has to be written per operator and agreed. +## The generated schema annex + +Annex B of Part 1 — 35 messages and 166 fields — is **generated** from the +vendored Protocol Buffers schema: + +```sh +make schema # regenerate from upstream/onnx/proto/ +make check-schema # fail if the committed file differs +``` + +Never edit `sources/part1/sections/annex-b-protobuf-schema.adoc` by hand. CI +runs `check-schema`, on the same contract as the operator clauses. + +The annex states, per message, each field's wire tag, type and obligation. The +obligation is the point of it: in the Protocol Buffers syntax this schema uses +every field is syntactically optional, and which ones a producer must supply is +carried in the comments, by the convention upstream's versioning document +defines. The generator reads that convention — a field whose comment says it +MUST be present for this version of the IR is mandatory — and 23 of the 166 +fields come out mandatory. + +The prose of the schema comments is deliberately **not** carried across. A +field table states structure; where a comment carries a normative statement, +that statement belongs in the clause it concerns, and the clauses are where +those statements are. The schema source stays vendored as an informative aid. + +`scripts/generate-schema.rb` is not a Protocol Buffers parser and is not meant +to be one. It reads the subset of proto2 these files use and **fails loudly** +on anything it does not recognize inside a message body, rather than skipping +it — so a schema change upstream shows up as a failed run rather than as a +missing row. + ## The vendored upstream copy `upstream/onnx/` is a verbatim copy of the ONNX documentation the draft diff --git a/Makefile b/Makefile index cd3e3d5..52f6d2b 100644 --- a/Makefile +++ b/Makefile @@ -29,7 +29,7 @@ RUBY ?= ruby SEVERITY ?= 1 .PHONY: all html doc pdf site lint clean deps check-stylesheet \ - operators check-operators + operators check-operators schema check-schema all: html @@ -70,6 +70,15 @@ operators: check-operators: @$(RUBY) scripts/generate-operators.rb --check +# Annex B of Part 1 is generated from the vendored Protocol Buffers schema: +# 35 messages and 166 fields, restated as tables of wire tag, type and +# obligation, which the schema source cannot express. +schema: + @$(RUBY) scripts/generate-schema.rb + +check-schema: + @$(RUBY) scripts/generate-schema.rb --check + # mn2pdf parses the stylesheet itself, and Metanorma exits 0 when that parse # fails — producing no PDF while the build looks successful. Check it first; # the check needs no fonts, so it also runs where a full PDF render cannot. diff --git a/README.md b/README.md index ea2c6ab..3efe459 100644 --- a/README.md +++ b/README.md @@ -54,6 +54,7 @@ sources/ scripts/check-errors.rb Gates the build on Metanorma diagnostic severity scripts/vendor-onnx-docs.sh Refreshes the vendored upstream copy scripts/generate-operators.rb Generates the operator clauses of Part 2 +scripts/generate-schema.rb Generates Annex B of Part 1 from the vendored .proto upstream/onnx/ Verbatim upstream ONNX docs (complete) + .proto .github/workflows/ CI: build HTML + PDF, publish as artifacts ``` diff --git a/scripts/generate-schema.rb b/scripts/generate-schema.rb new file mode 100755 index 0000000..435e94e --- /dev/null +++ b/scripts/generate-schema.rb @@ -0,0 +1,322 @@ +#!/usr/bin/env ruby +# frozen_string_literal: true + +# Generate Annex B of Part 1 from the vendored Protocol Buffers schema. +# +# scripts/generate-schema.rb [--check] +# +# Annex B owes a table of fields per message, giving wire tag, type and +# obligation. The schema source cannot state an obligation -- every proto2 +# field is syntactically optional -- so the obligation is read from the +# comment convention the upstream versioning document defines: a field whose +# leading comment says it MUST be present for this version of the IR is +# mandatory, and every other field is optional. A repeated field is neither; +# its obligation is its cardinality. +# +# `--check` regenerates into memory and fails if the file on disk differs, +# which is what CI runs. +# +# What this script deliberately does not carry across is the prose of the +# schema comments. A field table states structure; where a comment carries a +# normative statement it belongs in a clause, and the clauses of this part are +# where those statements now are. The schema source stays vendored as the +# informative aid Annex B points at. + +require "optparse" + +ROOT = File.expand_path("..", __dir__) + +# The ONNX-ML schema is the superset: it defines the map and sequence types +# and the ai.onnx.ml operator set that Part 2 specifies alongside the default +# domain. The plain onnx.proto is the same file with those omitted. +SOURCES = [ + { file: "upstream/onnx/proto/onnx-ml.proto", + title: "Model, graph and tensor messages" }, + { file: "upstream/onnx/proto/onnx-operators-ml.proto", + title: "Operator set messages" }, + { file: "upstream/onnx/proto/onnx-data.proto", + title: "Map and sequence messages" }, +].freeze + +TARGET = "sources/part1/sections/annex-b-protobuf-schema.adoc" + +MANDATORY = /MUST be present (?:in|for) this version of the IR/i.freeze + +Field = Struct.new(:label, :type, :name, :tag, :mandatory, :oneof, :deprecated) +Enum = Struct.new(:name, :values) +Msg = Struct.new(:name, :fields, :enums, :reserved) + +# --------------------------------------------------------------------- parse + +# A hand-rolled reader for the subset of proto2 these files use: messages, +# nested messages, enums, oneofs, reserved ranges and scalar fields. It is not +# a proto parser and is not meant to be one; it fails loudly on anything it +# does not recognise inside a message body rather than skipping it, so that a +# schema change upstream shows up as a failed run rather than a missing row. +def parse(path) + messages = [] + stack = [] # enclosing message names + comment = [] # comment lines gathered since the last statement + enum = nil + oneof = nil + depth_of_oneof = nil + + # The message currently being read, by its full path. `messages.last` is not + # it: once a nested message closes, the fields that follow belong to the + # enclosing message again, and appending them to the last entry created + # silently moved them onto the nested one. + current = lambda do + path_now = stack.join(".") + messages.reverse.find { |m| m.name == path_now } + end + + File.readlines(path, encoding: "UTF-8").each_with_index do |raw, i| + line = raw.strip + lineno = i + 1 + + if line.start_with?("//") + comment << line.sub(%r{\A//\s?}, "") + next + end + if line.empty? + comment.clear if stack.empty? && enum.nil? + next + end + + case line + when /\A(?:syntax|package|option|import)\b/ + comment.clear + when /\Amessage\s+(\w+)\s*\{/ + stack << Regexp.last_match(1) + messages << Msg.new(stack.join("."), [], [], []) + comment.clear + when /\Aenum\s+(\w+)\s*\{/ + enum = Enum.new([*stack, Regexp.last_match(1)].join("."), []) + # A top-level enum has no message to hang on; give it one of its own so + # that it is emitted in source order rather than dropped. + messages << Msg.new(enum.name, [], [], []) if stack.empty? + comment.clear + when /\Aoneof\s+(\w+)\s*\{/ + oneof = Regexp.last_match(1) + depth_of_oneof = stack.length + comment.clear + when /\Areserved\s+(.+);/ + owner = current.call or raise "#{path}:#{lineno}: reserved outside a message" + owner.reserved << Regexp.last_match(1).strip + comment.clear + when /\A\}\s*;?\z/ + if enum + # Nested enums belong to the message being read; a top-level enum + # belongs to the entry opened for it above, which is the last one. + owner = stack.empty? ? messages.last : current.call + owner or raise "#{path}:#{lineno}: enum has no owner" + owner.enums << enum + enum = nil + elsif oneof && depth_of_oneof == stack.length + oneof = nil + else + stack.pop or raise "#{path}:#{lineno}: unbalanced brace" + end + comment.clear + when /\A(\w+)\s*=\s*(0[xX][0-9a-fA-F]+|\d+)\s*;/ # enum member + enum or raise "#{path}:#{lineno}: enum member outside an enum: #{line}" + enum.values << [Regexp.last_match(1), Integer(Regexp.last_match(2))] + comment.clear + when /\A(optional|repeated|required)?\s*([\w.<>, ]+?)\s+(\w+)\s*=\s*(\d+)\s*(\[[^\]]*\])?\s*;/ + label = Regexp.last_match(1) || (oneof ? "oneof" : "optional") + type = Regexp.last_match(2).strip + name = Regexp.last_match(3) + tag = Regexp.last_match(4).to_i + text = comment.join(" ") + owner = current.call or raise "#{path}:#{lineno}: field outside a message" + owner.fields << Field.new( + label, type, name, tag, + text.match?(MANDATORY), oneof, + text.match?(/\bdeprecated\b/i) + ) + comment.clear + else + raise "#{path}:#{lineno}: unrecognised: #{line}" + end + end + + stack.empty? or raise "#{path}: unterminated message #{stack.inspect}" + messages.reject { |m| m.fields.empty? && m.enums.empty? } +end + +# --------------------------------------------------------------------- emit + +def obligation(field) + return "deprecated" if field.deprecated + return "repeated" if field.label == "repeated" + return "one of the group" if field.label == "oneof" + + field.mandatory ? "mandatory" : "optional" +end + +def adoc_type(type) + "`#{type}`" +end + +def anchor(name) + "schema-#{name.downcase.tr('.', '-')}" +end + +def render_message(msg) + out = [] + out << "[[#{anchor(msg.name)}]]" + out << "==== #{msg.name}" + out << "" + + unless msg.fields.empty? + out << "[[tbl-#{anchor(msg.name)}]]" + out << ".Fields of `#{msg.name}`" + out << '[cols="3,1,3,2"]' + out << "|===" + out << "| Field | Tag | Type | Obligation" + out << "" + msg.fields.each do |f| + out << "| `#{f.name}` | #{f.tag} | #{adoc_type(f.type)} | #{obligation(f)}" + end + out << "|===" + out << "" + end + + groups = msg.fields.map(&:oneof).compact.uniq + groups.each do |g| + members = msg.fields.select { |f| f.oneof == g }.map { |f| "`#{f.name}`" } + out << "Exactly one of #{members.join(', ')} SHALL be present, being the " \ + "group `#{g}`." + out << "" + end + + unless msg.reserved.empty? + out << "Tags and names reserved in `#{msg.name}`, which SHALL NOT be used: " \ + "#{msg.reserved.map { |r| "`#{r}`" }.join('; ')}." + out << "" + end + + msg.enums.each do |e| + out << "[[tbl-#{anchor(e.name)}]]" + out << ".Values of `#{e.name}`" + out << '[cols="3,1"]' + out << "|===" + out << "| Name | Value" + out << "" + e.values.each { |(n, v)| out << "| `#{n}` | #{v}" } + out << "|===" + out << "" + end + + out.join("\n") +end + +def render(groups) + counts = groups.sum { |g| g[:messages].length } + fields = groups.sum { |g| g[:messages].sum { |m| m.fields.length } } + mandatory = groups.sum do |g| + g[:messages].sum { |m| m.fields.count(&:mandatory) } + end + + out = [] + out << "// Generated by scripts/generate-schema.rb from upstream/onnx/proto/" + out << "// at the release recorded in upstream/onnx/SOURCE.txt." + out << "// Do not edit: run the script instead." + out << "// See CONTRIBUTING.md, \"The generated schema annex\"." + out << "" + out << "[[annex-schema]]" + out << "[appendix,obligation=normative]" + out << "== Protocol Buffers schema" + out << "" + out << "=== General" + out << "" + out << "This annex states the message definitions of the serialized form: " \ + "#{counts} messages and #{fields} fields." + out << "" + out << "The tag of a field is normative; it identifies the field on the " \ + "wire and SHALL NOT be reused. A field marked mandatory SHALL be " \ + "present. A field marked optional MAY be absent, and a " \ + "<> SHALL accept a message in which it is. A field marked " \ + "repeated holds zero or more values. A field marked deprecated " \ + "SHALL NOT be written by a <>, and a consumer that reads " \ + "one SHALL ignore it." + out << "" + out << "The obligations are those in force at the IR version stated in " \ + "<>." + out << "" + out << "NOTE: #{mandatory} of the #{fields} fields are mandatory. Protocol " \ + "Buffers cannot express the distinction: in the syntax this schema " \ + "uses every field is optional, and the obligation is carried in the " \ + "comments. Stating it in a table is the reason this annex exists." + out << "" + out << "[[tbl-schema-sources]]" + out << ".Subclauses of this annex and the schema file each restates" + out << '[cols="3,4,1"]' + out << "|===" + out << "| Subclause | Source | Messages" + out << "" + groups.each_with_index do |g, i| + out << "| B.#{i + 2} #{g[:title]} | `#{g[:file]}` | #{g[:messages].length}" + end + out << "|===" + out << "" + + groups.each do |g| + out << "=== #{g[:title]}" + out << "" + g[:messages].each { |m| out << render_message(m) } + end + + out << "=== Schema source" + out << "" + out << "The schema source is vendored in the repository at " \ + "`upstream/onnx/proto/`, at the release recorded in " \ + "`upstream/onnx/SOURCE.txt`. It is an informative aid: where a " \ + "comment in it carries a normative statement, that statement is in " \ + "the clause it belongs to and not in this annex." + out << "" + + out.join("\n").gsub(/\n{3,}/, "\n\n") +end + +# ---------------------------------------------------------------------- main + +check = false +OptionParser.new do |o| + o.banner = "usage: generate-schema.rb [--check]" + o.on("--check", "fail if the generated file differs from the one on disk") do + check = true + end +end.parse! + +groups = SOURCES.map do |src| + path = File.join(ROOT, src[:file]) + File.exist?(path) or abort "missing: #{src[:file]}" + { title: src[:title], file: src[:file], messages: parse(path) } +end + +text = render(groups) +target = File.join(ROOT, TARGET) + +if check + on_disk = File.exist?(target) ? File.read(target, encoding: "UTF-8") : nil + if on_disk == text + puts "check: Annex B is up to date" + else + warn "check: #{TARGET} differs from the generated output." + warn " Run `make schema` and commit the result." + exit 1 + end +else + File.write(target, text) + total_m = groups.sum { |g| g[:messages].length } + total_f = groups.sum { |g| g[:messages].sum { |m| m.fields.length } } + mand = groups.sum { |g| g[:messages].sum { |m| m.fields.count(&:mandatory) } } + enums = groups.sum { |g| g[:messages].sum { |m| m.enums.length } } + groups.each do |g| + puts format(" %-40s %3d messages", g[:file], g[:messages].length) + end + puts + puts format(" %d messages, %d fields (%d mandatory), %d enumerations -> %s", + total_m, total_f, mand, enums, TARGET) +end diff --git a/sources/part1/sections/12-serialization.adoc b/sources/part1/sections/12-serialization.adoc index 5ddca65..1b6573d 100644 --- a/sources/part1/sections/12-serialization.adoc +++ b/sources/part1/sections/12-serialization.adoc @@ -27,16 +27,16 @@ which is why <> is a separate question. === Message definitions -The message definitions are given normatively in Annex B. +The message definitions are given normatively in <>, as a table +of fields per message giving wire tag, type and obligation. -[IMPORTANT] -.Editorial note -==== -Annex B currently reproduces the schema as source. The intended form is a table -of fields per message, giving wire tag, type and obligation, with the schema -source retained as an informative aid. A table states each field's obligation, -which the schema cannot express, and survives a change of encoding. -==== +A field's tag identifies it on the wire. A tag SHALL NOT be reused for a +different field, and a tag recorded as reserved SHALL NOT be used at all. + +NOTE: The obligation is what the schema source cannot express, and is the +reason the annex is a table rather than the source. In the Protocol Buffers +syntax this schema uses, every field is syntactically optional; which ones a +producer must supply is carried in the comments. [[canonical-form]] === Canonical form diff --git a/sources/part1/sections/14-versioning.adoc b/sources/part1/sections/14-versioning.adoc index d93ed25..de9d06b 100644 --- a/sources/part1/sections/14-versioning.adoc +++ b/sources/part1/sections/14-versioning.adoc @@ -118,6 +118,7 @@ A model binds each of its nodes to an operator definition by the operator set version the model imports; see <>. How a consumer binds that definition to an implementation is outside the scope of this document. +[[compatibility]] === Compatibility obligations A consumer that claims support for operator set version stem:[n] of a domain diff --git a/sources/part1/sections/annex-b-protobuf-schema.adoc b/sources/part1/sections/annex-b-protobuf-schema.adoc index 473cb31..3c46184 100644 --- a/sources/part1/sections/annex-b-protobuf-schema.adoc +++ b/sources/part1/sections/annex-b-protobuf-schema.adoc @@ -1,15 +1,719 @@ +// Generated by scripts/generate-schema.rb from upstream/onnx/proto/ +// at the release recorded in upstream/onnx/SOURCE.txt. +// Do not edit: run the script instead. +// See CONTRIBUTING.md, "The generated schema annex". + [[annex-schema]] [appendix,obligation=normative] == Protocol Buffers schema -[IMPORTANT] -.Editorial note -==== -This annex is to carry the message definitions normatively, as a table of -fields per message giving wire tag, type and obligation, per the editorial note -in <>. It is empty pending the decision recorded there on the -normative status of the Protocol Buffers encoding. - -The schema source this annex will restate is vendored in the repository at -`upstream/onnx/proto/`, at the release recorded in its provenance. -==== +=== General + +This annex states the message definitions of the serialized form: 35 messages and 166 fields. + +The tag of a field is normative; it identifies the field on the wire and SHALL NOT be reused. A field marked mandatory SHALL be present. A field marked optional MAY be absent, and a <> SHALL accept a message in which it is. A field marked repeated holds zero or more values. A field marked deprecated SHALL NOT be written by a <>, and a consumer that reads one SHALL ignore it. + +The obligations are those in force at the IR version stated in <>. + +NOTE: 23 of the 166 fields are mandatory. Protocol Buffers cannot express the distinction: in the syntax this schema uses every field is optional, and the obligation is carried in the comments. Stating it in a table is the reason this annex exists. + +[[tbl-schema-sources]] +.Subclauses of this annex and the schema file each restates +[cols="3,4,1"] +|=== +| Subclause | Source | Messages + +| B.2 Model, graph and tensor messages | `upstream/onnx/proto/onnx-ml.proto` | 30 +| B.3 Operator set messages | `upstream/onnx/proto/onnx-operators-ml.proto` | 2 +| B.4 Map and sequence messages | `upstream/onnx/proto/onnx-data.proto` | 3 +|=== + +=== Model, graph and tensor messages + +[[schema-version]] +==== Version + +[[tbl-schema-version]] +.Values of `Version` +[cols="3,1"] +|=== +| Name | Value + +| `_START_VERSION` | 0 +| `IR_VERSION_2017_10_10` | 1 +| `IR_VERSION_2017_10_30` | 2 +| `IR_VERSION_2017_11_3` | 3 +| `IR_VERSION_2019_1_22` | 4 +| `IR_VERSION_2019_3_18` | 5 +| `IR_VERSION_2019_9_19` | 6 +| `IR_VERSION_2020_5_8` | 7 +| `IR_VERSION_2021_7_30` | 8 +| `IR_VERSION_2023_5_5` | 9 +| `IR_VERSION_2024_3_25` | 10 +| `IR_VERSION_2025_05_12` | 11 +| `IR_VERSION_2025_08_26` | 12 +| `IR_VERSION_2025_11_06` | 13 +| `IR_VERSION` | 14 +|=== + +[[schema-attributeproto]] +==== AttributeProto + +[[tbl-schema-attributeproto]] +.Fields of `AttributeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | mandatory +| `ref_attr_name` | 21 | `string` | optional +| `doc_string` | 13 | `string` | optional +| `type` | 20 | `AttributeType` | mandatory +| `f` | 2 | `float` | mandatory +| `i` | 3 | `int64` | optional +| `s` | 4 | `bytes` | optional +| `t` | 5 | `TensorProto` | optional +| `g` | 6 | `GraphProto` | optional +| `sparse_tensor` | 22 | `SparseTensorProto` | optional +| `tp` | 14 | `TypeProto` | deprecated +| `floats` | 7 | `float` | repeated +| `ints` | 8 | `int64` | repeated +| `strings` | 9 | `bytes` | repeated +| `tensors` | 10 | `TensorProto` | repeated +| `graphs` | 11 | `GraphProto` | repeated +| `sparse_tensors` | 23 | `SparseTensorProto` | repeated +| `type_protos` | 15 | `TypeProto` | repeated +|=== + +Tags and names reserved in `AttributeProto`, which SHALL NOT be used: `12, 16 to 19`; `"v"`. + +[[tbl-schema-attributeproto-attributetype]] +.Values of `AttributeProto.AttributeType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `FLOAT` | 1 +| `INT` | 2 +| `STRING` | 3 +| `TENSOR` | 4 +| `GRAPH` | 5 +| `SPARSE_TENSOR` | 11 +| `TYPE_PROTO` | 13 +| `FLOATS` | 6 +| `INTS` | 7 +| `STRINGS` | 8 +| `TENSORS` | 9 +| `GRAPHS` | 10 +| `SPARSE_TENSORS` | 12 +| `TYPE_PROTOS` | 14 +|=== + +[[schema-valueinfoproto]] +==== ValueInfoProto + +[[tbl-schema-valueinfoproto]] +.Fields of `ValueInfoProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | mandatory +| `type` | 2 | `TypeProto` | mandatory +| `doc_string` | 3 | `string` | optional +| `metadata_props` | 4 | `StringStringEntryProto` | repeated +|=== + +[[schema-nodeproto]] +==== NodeProto + +[[tbl-schema-nodeproto]] +.Fields of `NodeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `input` | 1 | `string` | repeated +| `output` | 2 | `string` | repeated +| `name` | 3 | `string` | optional +| `op_type` | 4 | `string` | optional +| `domain` | 7 | `string` | optional +| `overload` | 8 | `string` | optional +| `attribute` | 5 | `AttributeProto` | repeated +| `doc_string` | 6 | `string` | optional +| `metadata_props` | 9 | `StringStringEntryProto` | repeated +| `device_configurations` | 10 | `NodeDeviceConfigurationProto` | repeated +|=== + +[[schema-intintlistentryproto]] +==== IntIntListEntryProto + +[[tbl-schema-intintlistentryproto]] +.Fields of `IntIntListEntryProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `key` | 1 | `int64` | optional +| `value` | 2 | `int64` | repeated +|=== + +[[schema-nodedeviceconfigurationproto]] +==== NodeDeviceConfigurationProto + +[[tbl-schema-nodedeviceconfigurationproto]] +.Fields of `NodeDeviceConfigurationProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `configuration_id` | 1 | `string` | mandatory +| `sharding_spec` | 2 | `ShardingSpecProto` | repeated +| `pipeline_stage` | 3 | `int32` | optional +|=== + +[[schema-shardingspecproto]] +==== ShardingSpecProto + +[[tbl-schema-shardingspecproto]] +.Fields of `ShardingSpecProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `tensor_name` | 1 | `string` | mandatory +| `device` | 2 | `int64` | repeated +| `index_to_device_group_map` | 3 | `IntIntListEntryProto` | repeated +| `sharded_dim` | 4 | `ShardedDimProto` | repeated +|=== + +[[schema-shardeddimproto]] +==== ShardedDimProto + +[[tbl-schema-shardeddimproto]] +.Fields of `ShardedDimProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `axis` | 1 | `int64` | mandatory +| `simple_sharding` | 2 | `SimpleShardedDimProto` | repeated +|=== + +[[schema-simpleshardeddimproto]] +==== SimpleShardedDimProto + +[[tbl-schema-simpleshardeddimproto]] +.Fields of `SimpleShardedDimProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dim_value` | 1 | `int64` | one of the group +| `dim_param` | 2 | `string` | one of the group +| `num_shards` | 3 | `int64` | mandatory +|=== + +Exactly one of `dim_value`, `dim_param` SHALL be present, being the group `dim`. + +[[schema-traininginfoproto]] +==== TrainingInfoProto + +[[tbl-schema-traininginfoproto]] +.Fields of `TrainingInfoProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `initialization` | 1 | `GraphProto` | optional +| `algorithm` | 2 | `GraphProto` | optional +| `initialization_binding` | 3 | `StringStringEntryProto` | repeated +| `update_binding` | 4 | `StringStringEntryProto` | repeated +|=== + +[[schema-modelproto]] +==== ModelProto + +[[tbl-schema-modelproto]] +.Fields of `ModelProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `ir_version` | 1 | `int64` | optional +| `opset_import` | 8 | `OperatorSetIdProto` | repeated +| `producer_name` | 2 | `string` | optional +| `producer_version` | 3 | `string` | optional +| `domain` | 4 | `string` | optional +| `model_version` | 5 | `int64` | optional +| `doc_string` | 6 | `string` | optional +| `graph` | 7 | `GraphProto` | optional +| `metadata_props` | 14 | `StringStringEntryProto` | repeated +| `training_info` | 20 | `TrainingInfoProto` | repeated +| `functions` | 25 | `FunctionProto` | repeated +| `configuration` | 26 | `DeviceConfigurationProto` | repeated +|=== + +[[schema-deviceconfigurationproto]] +==== DeviceConfigurationProto + +[[tbl-schema-deviceconfigurationproto]] +.Fields of `DeviceConfigurationProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | mandatory +| `num_devices` | 2 | `int32` | mandatory +| `device` | 3 | `string` | repeated +|=== + +[[schema-stringstringentryproto]] +==== StringStringEntryProto + +[[tbl-schema-stringstringentryproto]] +.Fields of `StringStringEntryProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `key` | 1 | `string` | optional +| `value` | 2 | `string` | optional +|=== + +[[schema-tensorannotation]] +==== TensorAnnotation + +[[tbl-schema-tensorannotation]] +.Fields of `TensorAnnotation` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `tensor_name` | 1 | `string` | optional +| `quant_parameter_tensor_names` | 2 | `StringStringEntryProto` | repeated +|=== + +[[schema-graphproto]] +==== GraphProto + +[[tbl-schema-graphproto]] +.Fields of `GraphProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `node` | 1 | `NodeProto` | repeated +| `name` | 2 | `string` | optional +| `initializer` | 5 | `TensorProto` | repeated +| `sparse_initializer` | 15 | `SparseTensorProto` | repeated +| `doc_string` | 10 | `string` | optional +| `input` | 11 | `ValueInfoProto` | repeated +| `output` | 12 | `ValueInfoProto` | repeated +| `value_info` | 13 | `ValueInfoProto` | repeated +| `quantization_annotation` | 14 | `TensorAnnotation` | repeated +| `metadata_props` | 16 | `StringStringEntryProto` | repeated +|=== + +Tags and names reserved in `GraphProto`, which SHALL NOT be used: `3, 4, 6 to 9`; `"ir_version", "producer_version", "producer_tag", "domain"`. + +[[schema-tensorproto]] +==== TensorProto + +[[tbl-schema-tensorproto]] +.Fields of `TensorProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dims` | 1 | `int64` | repeated +| `data_type` | 2 | `int32` | optional +| `segment` | 3 | `Segment` | optional +| `float_data` | 4 | `float` | repeated +| `int32_data` | 5 | `int32` | repeated +| `string_data` | 6 | `bytes` | repeated +| `int64_data` | 7 | `int64` | repeated +| `name` | 8 | `string` | optional +| `doc_string` | 12 | `string` | optional +| `raw_data` | 9 | `bytes` | optional +| `external_data` | 13 | `StringStringEntryProto` | repeated +| `data_location` | 14 | `DataLocation` | optional +| `double_data` | 10 | `double` | repeated +| `uint64_data` | 11 | `uint64` | repeated +| `metadata_props` | 16 | `StringStringEntryProto` | repeated +|=== + +[[tbl-schema-tensorproto-datatype]] +.Values of `TensorProto.DataType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `FLOAT` | 1 +| `UINT8` | 2 +| `INT8` | 3 +| `UINT16` | 4 +| `INT16` | 5 +| `INT32` | 6 +| `INT64` | 7 +| `STRING` | 8 +| `BOOL` | 9 +| `FLOAT16` | 10 +| `DOUBLE` | 11 +| `UINT32` | 12 +| `UINT64` | 13 +| `COMPLEX64` | 14 +| `COMPLEX128` | 15 +| `BFLOAT16` | 16 +| `FLOAT8E4M3FN` | 17 +| `FLOAT8E4M3FNUZ` | 18 +| `FLOAT8E5M2` | 19 +| `FLOAT8E5M2FNUZ` | 20 +| `UINT4` | 21 +| `INT4` | 22 +| `FLOAT4E2M1` | 23 +| `FLOAT8E8M0` | 24 +| `UINT2` | 25 +| `INT2` | 26 +| `FLOAT6E2M3` | 27 +| `FLOAT6E3M2` | 28 +|=== + +[[tbl-schema-tensorproto-datalocation]] +.Values of `TensorProto.DataLocation` +[cols="3,1"] +|=== +| Name | Value + +| `DEFAULT` | 0 +| `EXTERNAL` | 1 +|=== + +[[schema-tensorproto-segment]] +==== TensorProto.Segment + +[[tbl-schema-tensorproto-segment]] +.Fields of `TensorProto.Segment` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `begin` | 1 | `int64` | optional +| `end` | 2 | `int64` | optional +|=== + +[[schema-sparsetensorproto]] +==== SparseTensorProto + +[[tbl-schema-sparsetensorproto]] +.Fields of `SparseTensorProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `values` | 1 | `TensorProto` | optional +| `indices` | 2 | `TensorProto` | optional +| `dims` | 3 | `int64` | repeated +|=== + +[[schema-tensorshapeproto]] +==== TensorShapeProto + +[[tbl-schema-tensorshapeproto]] +.Fields of `TensorShapeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dim` | 1 | `Dimension` | repeated +|=== + +[[schema-tensorshapeproto-dimension]] +==== TensorShapeProto.Dimension + +[[tbl-schema-tensorshapeproto-dimension]] +.Fields of `TensorShapeProto.Dimension` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `dim_value` | 1 | `int64` | one of the group +| `dim_param` | 2 | `string` | one of the group +| `denotation` | 3 | `string` | optional +|=== + +Exactly one of `dim_value`, `dim_param` SHALL be present, being the group `value`. + +[[schema-typeproto]] +==== TypeProto + +[[tbl-schema-typeproto]] +.Fields of `TypeProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `tensor_type` | 1 | `Tensor` | one of the group +| `sequence_type` | 4 | `Sequence` | one of the group +| `map_type` | 5 | `Map` | one of the group +| `optional_type` | 9 | `Optional` | one of the group +| `sparse_tensor_type` | 8 | `SparseTensor` | one of the group +| `opaque_type` | 7 | `Opaque` | one of the group +| `denotation` | 6 | `string` | optional +|=== + +Exactly one of `tensor_type`, `sequence_type`, `map_type`, `optional_type`, `sparse_tensor_type`, `opaque_type` SHALL be present, being the group `value`. + +[[schema-typeproto-tensor]] +==== TypeProto.Tensor + +[[tbl-schema-typeproto-tensor]] +.Fields of `TypeProto.Tensor` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `int32` | mandatory +| `shape` | 2 | `TensorShapeProto` | optional +|=== + +[[schema-typeproto-sequence]] +==== TypeProto.Sequence + +[[tbl-schema-typeproto-sequence]] +.Fields of `TypeProto.Sequence` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `TypeProto` | mandatory +|=== + +[[schema-typeproto-map]] +==== TypeProto.Map + +[[tbl-schema-typeproto-map]] +.Fields of `TypeProto.Map` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `key_type` | 1 | `int32` | mandatory +| `value_type` | 2 | `TypeProto` | mandatory +|=== + +[[schema-typeproto-optional]] +==== TypeProto.Optional + +[[tbl-schema-typeproto-optional]] +.Fields of `TypeProto.Optional` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `TypeProto` | mandatory +|=== + +[[schema-typeproto-sparsetensor]] +==== TypeProto.SparseTensor + +[[tbl-schema-typeproto-sparsetensor]] +.Fields of `TypeProto.SparseTensor` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `elem_type` | 1 | `int32` | mandatory +| `shape` | 2 | `TensorShapeProto` | optional +|=== + +[[schema-typeproto-opaque]] +==== TypeProto.Opaque + +[[tbl-schema-typeproto-opaque]] +.Fields of `TypeProto.Opaque` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `domain` | 1 | `string` | optional +| `name` | 2 | `string` | optional +|=== + +Tags and names reserved in `TypeProto.Opaque`, which SHALL NOT be used: `3`; `"parameters"`. + +[[schema-operatorsetidproto]] +==== OperatorSetIdProto + +[[tbl-schema-operatorsetidproto]] +.Fields of `OperatorSetIdProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `domain` | 1 | `string` | mandatory +| `version` | 2 | `int64` | mandatory +|=== + +[[schema-operatorstatus]] +==== OperatorStatus + +[[tbl-schema-operatorstatus]] +.Values of `OperatorStatus` +[cols="3,1"] +|=== +| Name | Value + +| `EXPERIMENTAL` | 0 +| `STABLE` | 1 +|=== + +[[schema-functionproto]] +==== FunctionProto + +[[tbl-schema-functionproto]] +.Fields of `FunctionProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `input` | 4 | `string` | repeated +| `output` | 5 | `string` | repeated +| `attribute` | 6 | `string` | repeated +| `attribute_proto` | 11 | `AttributeProto` | repeated +| `node` | 7 | `NodeProto` | repeated +| `doc_string` | 8 | `string` | optional +| `opset_import` | 9 | `OperatorSetIdProto` | repeated +| `domain` | 10 | `string` | optional +| `overload` | 13 | `string` | optional +| `value_info` | 12 | `ValueInfoProto` | repeated +| `metadata_props` | 14 | `StringStringEntryProto` | repeated +|=== + +Tags and names reserved in `FunctionProto`, which SHALL NOT be used: `2`; `"since_version"`; `3`; `"status"`. + +=== Operator set messages + +[[schema-operatorproto]] +==== OperatorProto + +[[tbl-schema-operatorproto]] +.Fields of `OperatorProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `op_type` | 1 | `string` | mandatory +| `since_version` | 2 | `int64` | mandatory +| `status` | 3 | `OperatorStatus` | optional +| `doc_string` | 10 | `string` | optional +|=== + +[[schema-operatorsetproto]] +==== OperatorSetProto + +[[tbl-schema-operatorsetproto]] +.Fields of `OperatorSetProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `magic` | 1 | `string` | mandatory +| `ir_version` | 2 | `int64` | mandatory +| `ir_version_prerelease` | 3 | `string` | optional +| `ir_build_metadata` | 7 | `string` | optional +| `domain` | 4 | `string` | optional +| `opset_version` | 5 | `int64` | optional +| `doc_string` | 6 | `string` | optional +| `operator` | 8 | `OperatorProto` | repeated +| `functions` | 9 | `FunctionProto` | repeated +|=== + +=== Map and sequence messages + +[[schema-sequenceproto]] +==== SequenceProto + +[[tbl-schema-sequenceproto]] +.Fields of `SequenceProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `elem_type` | 2 | `int32` | optional +| `tensor_values` | 3 | `TensorProto` | repeated +| `sparse_tensor_values` | 4 | `SparseTensorProto` | repeated +| `sequence_values` | 5 | `SequenceProto` | repeated +| `map_values` | 6 | `MapProto` | repeated +| `optional_values` | 7 | `OptionalProto` | repeated +|=== + +[[tbl-schema-sequenceproto-datatype]] +.Values of `SequenceProto.DataType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `TENSOR` | 1 +| `SPARSE_TENSOR` | 2 +| `SEQUENCE` | 3 +| `MAP` | 4 +| `OPTIONAL` | 5 +|=== + +[[schema-mapproto]] +==== MapProto + +[[tbl-schema-mapproto]] +.Fields of `MapProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `key_type` | 2 | `int32` | optional +| `keys` | 3 | `int64` | repeated +| `string_keys` | 4 | `bytes` | repeated +| `values` | 5 | `SequenceProto` | optional +|=== + +[[schema-optionalproto]] +==== OptionalProto + +[[tbl-schema-optionalproto]] +.Fields of `OptionalProto` +[cols="3,1,3,2"] +|=== +| Field | Tag | Type | Obligation + +| `name` | 1 | `string` | optional +| `elem_type` | 2 | `int32` | optional +| `tensor_value` | 3 | `TensorProto` | optional +| `sparse_tensor_value` | 4 | `SparseTensorProto` | optional +| `sequence_value` | 5 | `SequenceProto` | optional +| `map_value` | 6 | `MapProto` | optional +| `optional_value` | 7 | `OptionalProto` | optional +|=== + +[[tbl-schema-optionalproto-datatype]] +.Values of `OptionalProto.DataType` +[cols="3,1"] +|=== +| Name | Value + +| `UNDEFINED` | 0 +| `TENSOR` | 1 +| `SPARSE_TENSOR` | 2 +| `SEQUENCE` | 3 +| `MAP` | 4 +| `OPTIONAL` | 5 +|=== + +=== Schema source + +The schema source is vendored in the repository at `upstream/onnx/proto/`, at the release recorded in `upstream/onnx/SOURCE.txt`. It is an informative aid: where a comment in it carries a normative statement, that statement is in the clause it belongs to and not in this annex. diff --git a/sources/part1/sections/annex-c-known-gaps.adoc b/sources/part1/sections/annex-c-known-gaps.adoc index 947ab0b..6130a35 100644 --- a/sources/part1/sections/annex-c-known-gaps.adoc +++ b/sources/part1/sections/annex-c-known-gaps.adoc @@ -38,13 +38,6 @@ form a standard cannot adopt. testable conditions with severities. | Blocking -| Annex B -| The Protocol Buffers schema is not yet restated normatively as a table of - fields with obligations. <> now requires that annex to identify - the fields each IR version makes mandatory, which the schema source cannot - express. -| Blocking - | <> Clauses 5, 6 | The generated operator clauses supply shape inference, determinism and error conditions for 0 of 222 operators: the upstream documentation states none of @@ -103,7 +96,17 @@ form a standard cannot adopt. | Clause 12 (encoding) | The admissible subset of the Protocol Buffers wire format --- unknown fields, - repeated field ordering, maximum message size --- is unstated. + repeated field ordering, maximum message size --- is unstated. It cannot be + derived from the schema, which says nothing about any of the three; the + reference implementation enforces a maximum message size in its checker, + whose source is not part of the vendored baseline. +| Major + +| <> +| The annex states which fields are mandatory at one IR version, being the one + <> names. The schema source carries no history of the + obligations, so a model produced against an earlier IR version cannot be + checked against this annex, and <> requires that it can be. | Major | <> @@ -174,6 +177,9 @@ transcription rather than take it on trust, and so that a refresh of | <> (typed storage) | `onnx-ml.proto`, `TensorProto.int32_data` +| <> (the whole annex) +| `upstream/onnx/proto/`, generated by `scripts/generate-schema.rb` + | <> (the IR version history) | `docs/Versioning.md`, "Released Versions"; `onnx-ml.proto`, `Version`