Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs/developer/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -331,7 +331,7 @@ An author may declare that the result does not depend on how the elements are gr
The backend may then compute a request over many elements in parts, on many processes; the record is the same.

A reduction with a sum in the middle splits into three specs.
A package derives them from its sciline `Aggregation`, sciline's description of such a split:
A package builds CONTRIBUTE and FINALIZE from the stages of one sciline `split`, cut at the keys that PARTS_SUM sums:

```text
run 611 ── CONTRIBUTE ──┐
Expand All @@ -351,6 +351,7 @@ FINALIZE = WorkflowSpec(name='sans-finalize', ..., params=FinalizeParams, output

`ContributeOutputs` has the fields `numerator` and `denominator`, and may have more, such as a transmission per run, which are not summed.
`FinalizeParams` has the data fields `numerator` and `denominator`; FINALIZE does not know that they are sums.
A FINALIZE may read several sums, for example one over sample runs and one over background runs.
Position i of each list in a PARTS_SUM request refers to the same run.

Which quantity is summed changes the result: summing counts and normalizing once is not the same as averaging normalized curves.
Expand Down
4 changes: 2 additions & 2 deletions docs/developer/plans/handoff.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ A living document for the next session on branch `event-log`. Read it first, the
| Tests | `packages/essapps/tests/`: `backend_test.py`, `log_test.py`, `sessions_test.py`, `pipeline_test.py`, `stories/*_test.py` (one test per API-tier story), toy specs and fixtures in `stories/conftest.py` |
| The previous attempt of this session's work | branch `core-old` (README with 8 terms, stories with the old vocabulary, the accumulating-inputs draft) |
| The previous design and skeleton | branch `architecture-sketch`, tip `b840b1f`; see "Implementation notes" in `todo.md` |
| sciline ADR 0003 (Stage, Aggregation, Accumulator) | `/workspace/sciline`, branch `map-reduce-outside-the-graph`; `docs/developer/adr/0003-*.md` and `docs/developer/architecture-and-design/map-reduce-outside-the-graph.md` |
| sciline ADR 0003 (Stage, split, Accumulator) | `/workspace/sciline`, branch `map-reduce-outside-the-graph`; `docs/developer/adr/0003-*.md` and `docs/developer/architecture-and-design/map-reduce-outside-the-graph.md` |
| ess.reduce workflow spec (ADR 0001) | `/workspace/ess`, branch `653-minimal-workflow-spec`, `packages/essreduce/src/ess/reduce/spec/` |

Environment: `.venv` in the worktree, made with `python3 -m venv --system-site-packages .venv`, then `pip install -e /workspace/ess/packages/essreduce --no-deps` and `pip install -e 'packages/essapps[test]' --no-deps`.
Expand Down Expand Up @@ -53,7 +53,7 @@ The essentials, all in the README with code:
- A record holds the request with every value filled in; statuses `pending`, `completed`, `failed`, `cancelled`; a finished record never changes. Records are history, kept for a retention period, and outlive sessions; an output's value is kept only while a pending request, a client's record handle, or a holder holds it, or once saved (system.md); long-term provenance is what `publish` puts in the catalogue. Provenance stops at datasets (what lies behind a dataset belongs to its source).
- Every connection between requests is a reference; a reference to a pending record is a valid input, which is the only scheduling mechanism.
- Datasets are named by `dataset(run=/path=/pid=)`; the record names the identity. The backend resolves names and reads data through its dataset source; drivers and forms list, watch, and read metadata through `client.datasets`, which shows only the datasets the client's proposal may read. Selectors match raw datasets unless they name another kind.
- An accumulator spec (`AccumulatorSpec(name=, version=, element=)`) takes one list per element field and outputs the element model, so a combined value can be pushed again. A package derives CONTRIBUTE, the accumulator spec, and FINALIZE from its sciline `Aggregation`.
- An accumulator spec (`AccumulatorSpec(name=, version=, element=)`) takes one list per element field and outputs the element model, so a combined value can be pushed again. A package builds CONTRIBUTE and FINALIZE from the stages of one sciline `split`, with the accumulator spec between them.
- Holders live in a session (`client.session(where=...)`): a stage holds a template; an accumulator holds pushed elements. A stage never changes what a record says; a snapshot's record names its accumulator and how many elements it covers.
- A driver is code that uses the client over time (notebook, application, trigger loop in a driving server); drivers never run in the backend. A tree of partial sums over a known list is how the backend may execute one accumulator request, not a driver.
- Several ways to write a sum are accepted: a spec with a list parameter, a chain of requests, the same chain through holders.
Expand Down
2 changes: 1 addition & 1 deletion docs/developer/plans/todo.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,4 +46,4 @@ The old `adapter.py` also let a binding ask for a dataset as a local path or as

- `AccumulatorSpec` and `combine` belong in `ess.reduce.spec`, next to `WorkflowSpec`. That needs a proposal in essreduce.
- `PipelineBinding` belongs in ess.reduce, where specs meet sciline workflows; it is the only module here that imports sciline.
- Stages and accumulators build on sciline ADR 0003 (`Stage`, `Aggregation`, `Accumulator`), which is still proposed. The venv needs sciline from `/workspace/sciline`, branch `map-reduce-outside-the-graph`.
- Stages and accumulators build on sciline ADR 0003 (`Stage`, `split`, `Accumulator`), which is still proposed. The venv needs sciline from `/workspace/sciline`, branch `map-reduce-outside-the-graph`.