Skip to content

a saved answer records nothing about the data it was solved with #1673

Description

@FBumann

Note

The following content was generated by AI, across two sessions. The body has
been rewritten rather than appended to: the first version filed this as a
follow-up to #1672, and it is a prerequisite — #1672 cannot ship without
it. The design below is new; the measurements are from the original session
and their method and caveats are with them.

Result.save writes the Record, the frames and reasons.parquet. The
Record carries spec_digest and nothing about sources. So an answer read back
against a spec and its data has its spec checked by digest and its data
checked on nothing.

Why this blocks #1672

#1672 adds spec= / sources= to load_result and scan_result. Measured on
its branch, against examples/dispatch.yaml:

what you pass as sources= at the load at the first undeclared evaluate
{'nope': 1} accepted DataError
{} accepted DataError
right keys, cost missing accepted DataError
right keys, wrong numbers accepted [7920.0, 7920.0], where the truth is 80.0

Fitness is checked, late, by the data contract at the rebuild. Identity is never
checked. The last row is a 99× wrong answer from data that fits the spec
perfectly, and CONTRIBUTING is unambiguous about that shape: a permissive input
that hides a silent wrong answer gets fixed in place
.

The blast radius is narrow — primal, dual, activity and every declared
expression are saved frames, and only the undeclared-expression evaluator reads
the supplied data — but narrow and silent is still wrong.

Design: lazy, cached, and free for a solve that never saves

The objection to recording digests was that hashing every source taxes everyone
to protect one flow. It does not have to be paid at solve time.

what you do what you pay
solve() and nothing else nothing — nothing asks, so nothing is hashed
solve() then save() one pass over the sources, once per data set
load_result(spec=, sources=) one pass over the caller's own copy
  1. Model computes source digests lazily and caches them.
  2. Model.solve hands the Result a thunk, the shape Result._evaluate
    already uses, so save can ask for them at save time.
  3. load_result / scan_result digest what they are handed and refuse a
    mismatch, naming the source that moved.

The thunk must close over the sources as of the solve, not over
Model._sources. Otherwise solve() → update(other) → save() digests data
the answer never saw and stamps a lie onto it. Model.solve captures
dict(self._sources) — a reference copy — and the memo keys on that snapshot
rather than living on the Model.

An earlier version of this issue said a closure on a Result was barred because
strategy.py pickles one. That was wrong. _Answer's own docstring says
"Plain data throughout — frames, strings and numbers, never a result or a
model — so it can cross a process"
: a Result never crosses, which is what
_Answer is for. And Result._evaluate already holds a closure over a Spec
and its sources, built in api.py. The layering rule is about imports, not
runtime capture.

What to digest

The tidied tables, not the source files and not the raw frames.
Model.__init__ already calls tidy_sources, so the canonical form is there.

A file digest is the wrong instrument, and measurably so — the same three-row
table written eight ways:

the same data, written differently file digest content digest
identical frame, written again same same
snappy instead of zstd differs same
uncompressed differs same
row group size 1 differs same
statistics off differs same
rows in a different order differs differs
columns swapped differs differs
value as float32 differs differs

Four encoding choices change the file digest and nothing about the model. Each
is a refusal the reader cannot act on, and a polars default that moves
invalidates every stored digest — which bites hardest where the digest is taken
on one machine and checked on another, the deployment this is for.

Order sensitivity

Source row order does reach the built model's label order: writing
examples/dispatch.yaml with the generator tables reversed puts a different
variable at x0 (+1.0 x0 against +50.0 x0).

It does not reach the re-evaluated answer. A solve saved from declared order,
read back against reversed sources, evaluates
sum(p * cost, over=generator) to the same values — the saved frames carry
coordinates and are re-aligned by them rather than positionally.

So the digest should be insensitive to row and column order and sensitive to
dtype. Combine the per-row hashes commutatively rather than sorting, and select
columns in declared order.

Checked on one model and one expression. Order-invariance of re-evaluation
wants a corpus test before the digest relies on it. If it does not hold
everywhere, the digest has to be order-sensitive and the false alarms come with
it.

Cost, when it is paid

Measured on examples/dispatch.yaml at 5M snapshots × 3 generators, min of 3
repeats:

cost share of a build
digest_of_file, 47.9 MB of parquet 0.124 s 2.1%
serialize + sha256, 15M cells in memory 0.463 s 8.2%
hash_rows + sha256, same 0.244 s 4.3%

The on-disk share held across a 5× size change (2.2% at 9.6 MB, 2.1% at
47.9 MB). Caveats: a shared container rather than an idle machine, one process,
warm page cache. A cold cache makes hashing IO-bound, but a build reads the same
bytes, so the share would if anything fall. Build time only — against build plus
solve it is smaller again.

Under the design above none of this is on the solve path. It is the price of a
save, once per data set.

Scope

  • Model holds a lazy, cached digest of its tidy tables, invalidated by
    update.
  • Result carries a thunk for them; save writes them beside the Record.
  • load_result / scan_result compare and name the source that moved.
  • One scheme, one home. The archive's sources.parquet uses digest_of_file
    today; two schemes for one fact is the drift parquet.other_specs was
    consolidated to prevent, so the archive moves to the same digest or the
    overlap is resolved deliberately. See an archive's source digests are written, read back and never compared #1674.
  • A corpus test for order-invariance, before the digest depends on it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions