diff --git a/docs/developer/README.md b/docs/developer/README.md index 7219744..01508af 100644 --- a/docs/developer/README.md +++ b/docs/developer/README.md @@ -331,7 +331,7 @@ An author may declare that the result does not depend on how the elements are gr The backend may then compute a request over many elements in parts, on many processes; the record is the same. A reduction with a sum in the middle splits into three specs. -A package derives them from its sciline `Aggregation`, sciline's description of such a split: +A package builds CONTRIBUTE and FINALIZE from the stages of one sciline `split`, cut at the keys that PARTS_SUM sums: ```text run 611 ── CONTRIBUTE ──┐ @@ -351,6 +351,7 @@ FINALIZE = WorkflowSpec(name='sans-finalize', ..., params=FinalizeParams, output `ContributeOutputs` has the fields `numerator` and `denominator`, and may have more, such as a transmission per run, which are not summed. `FinalizeParams` has the data fields `numerator` and `denominator`; FINALIZE does not know that they are sums. +A FINALIZE may read several sums, for example one over sample runs and one over background runs. Position i of each list in a PARTS_SUM request refers to the same run. Which quantity is summed changes the result: summing counts and normalizing once is not the same as averaging normalized curves. diff --git a/docs/developer/plans/handoff.md b/docs/developer/plans/handoff.md index c328ce6..6e6b217 100644 --- a/docs/developer/plans/handoff.md +++ b/docs/developer/plans/handoff.md @@ -20,7 +20,7 @@ A living document for the next session on branch `event-log`. Read it first, the | Tests | `packages/essapps/tests/`: `backend_test.py`, `log_test.py`, `sessions_test.py`, `pipeline_test.py`, `stories/*_test.py` (one test per API-tier story), toy specs and fixtures in `stories/conftest.py` | | The previous attempt of this session's work | branch `core-old` (README with 8 terms, stories with the old vocabulary, the accumulating-inputs draft) | | The previous design and skeleton | branch `architecture-sketch`, tip `b840b1f`; see "Implementation notes" in `todo.md` | -| sciline ADR 0003 (Stage, Aggregation, Accumulator) | `/workspace/sciline`, branch `map-reduce-outside-the-graph`; `docs/developer/adr/0003-*.md` and `docs/developer/architecture-and-design/map-reduce-outside-the-graph.md` | +| sciline ADR 0003 (Stage, split, Accumulator) | `/workspace/sciline`, branch `map-reduce-outside-the-graph`; `docs/developer/adr/0003-*.md` and `docs/developer/architecture-and-design/map-reduce-outside-the-graph.md` | | ess.reduce workflow spec (ADR 0001) | `/workspace/ess`, branch `653-minimal-workflow-spec`, `packages/essreduce/src/ess/reduce/spec/` | Environment: `.venv` in the worktree, made with `python3 -m venv --system-site-packages .venv`, then `pip install -e /workspace/ess/packages/essreduce --no-deps` and `pip install -e 'packages/essapps[test]' --no-deps`. @@ -53,7 +53,7 @@ The essentials, all in the README with code: - A record holds the request with every value filled in; statuses `pending`, `completed`, `failed`, `cancelled`; a finished record never changes. Records are history, kept for a retention period, and outlive sessions; an output's value is kept only while a pending request, a client's record handle, or a holder holds it, or once saved (system.md); long-term provenance is what `publish` puts in the catalogue. Provenance stops at datasets (what lies behind a dataset belongs to its source). - Every connection between requests is a reference; a reference to a pending record is a valid input, which is the only scheduling mechanism. - Datasets are named by `dataset(run=/path=/pid=)`; the record names the identity. The backend resolves names and reads data through its dataset source; drivers and forms list, watch, and read metadata through `client.datasets`, which shows only the datasets the client's proposal may read. Selectors match raw datasets unless they name another kind. -- An accumulator spec (`AccumulatorSpec(name=, version=, element=)`) takes one list per element field and outputs the element model, so a combined value can be pushed again. A package derives CONTRIBUTE, the accumulator spec, and FINALIZE from its sciline `Aggregation`. +- An accumulator spec (`AccumulatorSpec(name=, version=, element=)`) takes one list per element field and outputs the element model, so a combined value can be pushed again. A package builds CONTRIBUTE and FINALIZE from the stages of one sciline `split`, with the accumulator spec between them. - Holders live in a session (`client.session(where=...)`): a stage holds a template; an accumulator holds pushed elements. A stage never changes what a record says; a snapshot's record names its accumulator and how many elements it covers. - A driver is code that uses the client over time (notebook, application, trigger loop in a driving server); drivers never run in the backend. A tree of partial sums over a known list is how the backend may execute one accumulator request, not a driver. - Several ways to write a sum are accepted: a spec with a list parameter, a chain of requests, the same chain through holders. diff --git a/docs/developer/plans/todo.md b/docs/developer/plans/todo.md index 9eb00f8..54eeec4 100644 --- a/docs/developer/plans/todo.md +++ b/docs/developer/plans/todo.md @@ -46,4 +46,4 @@ The old `adapter.py` also let a binding ask for a dataset as a local path or as - `AccumulatorSpec` and `combine` belong in `ess.reduce.spec`, next to `WorkflowSpec`. That needs a proposal in essreduce. - `PipelineBinding` belongs in ess.reduce, where specs meet sciline workflows; it is the only module here that imports sciline. -- Stages and accumulators build on sciline ADR 0003 (`Stage`, `Aggregation`, `Accumulator`), which is still proposed. The venv needs sciline from `/workspace/sciline`, branch `map-reduce-outside-the-graph`. +- Stages and accumulators build on sciline ADR 0003 (`Stage`, `split`, `Accumulator`), which is still proposed. The venv needs sciline from `/workspace/sciline`, branch `map-reduce-outside-the-graph`.