You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The following content was generated by AI, across two sessions. The body has
been rewritten rather than appended to: the first version filed this as a
follow-up to #1672, and it is a prerequisite — #1672 cannot ship without
it. The design below is new; the measurements are from the original session
and their method and caveats are with them.
Result.save writes the Record, the frames and reasons.parquet. The Record carries spec_digest and nothing about sources. So an answer read back
against a spec and its data has its spec checked by digest and its data
checked on nothing.
#1672 adds spec= / sources= to load_result and scan_result. Measured on
its branch, against examples/dispatch.yaml:
what you pass as sources=
at the load
at the first undeclared evaluate
{'nope': 1}
accepted
DataError
{}
accepted
DataError
right keys, cost missing
accepted
DataError
right keys, wrong numbers
accepted
[7920.0, 7920.0], where the truth is 80.0
Fitness is checked, late, by the data contract at the rebuild. Identity is never
checked. The last row is a 99× wrong answer from data that fits the spec
perfectly, and CONTRIBUTING is unambiguous about that shape: a permissive input
that hides a silent wrong answer gets fixed in place.
The blast radius is narrow — primal, dual, activity and every declared
expression are saved frames, and only the undeclared-expression evaluator reads
the supplied data — but narrow and silent is still wrong.
Design: lazy, cached, and free for a solve that never saves
The objection to recording digests was that hashing every source taxes everyone
to protect one flow. It does not have to be paid at solve time.
what you do
what you pay
solve() and nothing else
nothing — nothing asks, so nothing is hashed
solve() then save()
one pass over the sources, once per data set
load_result(spec=, sources=)
one pass over the caller's own copy
Model computes source digests lazily and caches them.
Model.solve hands the Result a thunk, the shape Result._evaluate
already uses, so save can ask for them at save time.
load_result / scan_result digest what they are handed and refuse a
mismatch, naming the source that moved.
The thunk must close over the sources as of the solve, not over Model._sources. Otherwise solve() → update(other) → save() digests data
the answer never saw and stamps a lie onto it. Model.solve captures dict(self._sources) — a reference copy — and the memo keys on that snapshot
rather than living on the Model.
An earlier version of this issue said a closure on a Result was barred because strategy.py pickles one. That was wrong._Answer's own docstring says "Plain data throughout — frames, strings and numbers, never a result or a
model — so it can cross a process": a Result never crosses, which is what _Answer is for. And Result._evaluate already holds a closure over a Spec
and its sources, built in api.py. The layering rule is about imports, not
runtime capture.
What to digest
The tidied tables, not the source files and not the raw frames. Model.__init__ already calls tidy_sources, so the canonical form is there.
A file digest is the wrong instrument, and measurably so — the same three-row
table written eight ways:
the same data, written differently
file digest
content digest
identical frame, written again
same
same
snappy instead of zstd
differs
same
uncompressed
differs
same
row group size 1
differs
same
statistics off
differs
same
rows in a different order
differs
differs
columns swapped
differs
differs
value as float32
differs
differs
Four encoding choices change the file digest and nothing about the model. Each
is a refusal the reader cannot act on, and a polars default that moves
invalidates every stored digest — which bites hardest where the digest is taken
on one machine and checked on another, the deployment this is for.
Order sensitivity
Source row order does reach the built model's label order: writing examples/dispatch.yaml with the generator tables reversed puts a different
variable at x0 (+1.0 x0 against +50.0 x0).
It does not reach the re-evaluated answer. A solve saved from declared order,
read back against reversed sources, evaluates sum(p * cost, over=generator) to the same values — the saved frames carry
coordinates and are re-aligned by them rather than positionally.
So the digest should be insensitive to row and column order and sensitive to
dtype. Combine the per-row hashes commutatively rather than sorting, and select
columns in declared order.
Checked on one model and one expression. Order-invariance of re-evaluation
wants a corpus test before the digest relies on it. If it does not hold
everywhere, the digest has to be order-sensitive and the false alarms come with
it.
Cost, when it is paid
Measured on examples/dispatch.yaml at 5M snapshots × 3 generators, min of 3
repeats:
cost
share of a build
digest_of_file, 47.9 MB of parquet
0.124 s
2.1%
serialize + sha256, 15M cells in memory
0.463 s
8.2%
hash_rows + sha256, same
0.244 s
4.3%
The on-disk share held across a 5× size change (2.2% at 9.6 MB, 2.1% at
47.9 MB). Caveats: a shared container rather than an idle machine, one process,
warm page cache. A cold cache makes hashing IO-bound, but a build reads the same
bytes, so the share would if anything fall. Build time only — against build plus
solve it is smaller again.
Under the design above none of this is on the solve path. It is the price of a save, once per data set.
Scope
Model holds a lazy, cached digest of its tidy tables, invalidated by update.
Result carries a thunk for them; save writes them beside the Record.
load_result / scan_result compare and name the source that moved.
One scheme, one home. The archive's sources.parquet uses digest_of_file
today; two schemes for one fact is the drift parquet.other_specs was
consolidated to prevent, so the archive moves to the same digest or the
overlap is resolved deliberately. See an archive's source digests are written, read back and never compared #1674.
A corpus test for order-invariance, before the digest depends on it.
Note
The following content was generated by AI, across two sessions. The body has
been rewritten rather than appended to: the first version filed this as a
follow-up to #1672, and it is a prerequisite — #1672 cannot ship without
it. The design below is new; the measurements are from the original session
and their method and caveats are with them.
Result.savewrites theRecord, the frames andreasons.parquet. TheRecordcarriesspec_digestand nothing about sources. So an answer read backagainst a spec and its data has its spec checked by digest and its data
checked on nothing.
Why this blocks #1672
#1672 adds
spec=/sources=toload_resultandscan_result. Measured onits branch, against
examples/dispatch.yaml:sources=evaluate{'nope': 1}DataError{}DataErrorcostmissingDataError[7920.0, 7920.0], where the truth is80.0Fitness is checked, late, by the data contract at the rebuild. Identity is never
checked. The last row is a 99× wrong answer from data that fits the spec
perfectly, and CONTRIBUTING is unambiguous about that shape: a permissive input
that hides a silent wrong answer gets fixed in place.
The blast radius is narrow —
primal,dual,activityand every declaredexpression are saved frames, and only the undeclared-expression evaluator reads
the supplied data — but narrow and silent is still wrong.
Design: lazy, cached, and free for a solve that never saves
The objection to recording digests was that hashing every source taxes everyone
to protect one flow. It does not have to be paid at solve time.
solve()and nothing elsesolve()thensave()load_result(spec=, sources=)Modelcomputes source digests lazily and caches them.Model.solvehands theResulta thunk, the shapeResult._evaluatealready uses, so
savecan ask for them at save time.load_result/scan_resultdigest what they are handed and refuse amismatch, naming the source that moved.
The thunk must close over the sources as of the solve, not over
Model._sources. Otherwisesolve()→update(other)→save()digests datathe answer never saw and stamps a lie onto it.
Model.solvecapturesdict(self._sources)— a reference copy — and the memo keys on that snapshotrather than living on the
Model.An earlier version of this issue said a closure on a
Resultwas barred becausestrategy.pypickles one. That was wrong._Answer's own docstring says"Plain data throughout — frames, strings and numbers, never a result or a
model — so it can cross a process": a
Resultnever crosses, which is what_Answeris for. AndResult._evaluatealready holds a closure over aSpecand its sources, built in
api.py. The layering rule is about imports, notruntime capture.
What to digest
The tidied tables, not the source files and not the raw frames.
Model.__init__already callstidy_sources, so the canonical form is there.A file digest is the wrong instrument, and measurably so — the same three-row
table written eight ways:
Four encoding choices change the file digest and nothing about the model. Each
is a refusal the reader cannot act on, and a polars default that moves
invalidates every stored digest — which bites hardest where the digest is taken
on one machine and checked on another, the deployment this is for.
Order sensitivity
Source row order does reach the built model's label order: writing
examples/dispatch.yamlwith the generator tables reversed puts a differentvariable at
x0(+1.0 x0against+50.0 x0).It does not reach the re-evaluated answer. A solve saved from declared order,
read back against reversed sources, evaluates
sum(p * cost, over=generator)to the same values — the saved frames carrycoordinates and are re-aligned by them rather than positionally.
So the digest should be insensitive to row and column order and sensitive to
dtype. Combine the per-row hashes commutatively rather than sorting, and select
columns in declared order.
Checked on one model and one expression. Order-invariance of re-evaluation
wants a corpus test before the digest relies on it. If it does not hold
everywhere, the digest has to be order-sensitive and the false alarms come with
it.
Cost, when it is paid
Measured on
examples/dispatch.yamlat 5M snapshots × 3 generators, min of 3repeats:
digest_of_file, 47.9 MB of parquetserialize+ sha256, 15M cells in memoryhash_rows+ sha256, sameThe on-disk share held across a 5× size change (2.2% at 9.6 MB, 2.1% at
47.9 MB). Caveats: a shared container rather than an idle machine, one process,
warm page cache. A cold cache makes hashing IO-bound, but a build reads the same
bytes, so the share would if anything fall. Build time only — against build plus
solve it is smaller again.
Under the design above none of this is on the solve path. It is the price of a
save, once per data set.Scope
Modelholds a lazy, cached digest of its tidy tables, invalidated byupdate.Resultcarries a thunk for them;savewrites them beside theRecord.load_result/scan_resultcompare and name the source that moved.sources.parquetusesdigest_of_filetoday; two schemes for one fact is the drift
parquet.other_specswasconsolidated to prevent, so the archive moves to the same digest or the
overlap is resolved deliberately. See an archive's source digests are written, read back and never compared #1674.