feat(memory): generation phase spans and derived metrics — TTFT, tok/s, effective bandwidth (SKEEP-003 P4, S2.2) - #1082
Merged
Conversation
…s, effective bandwidth, per-module breakdown Closes #1035 (SKEEP-003 P4, S2.2, proposal §4.9). The trace stream carried phases, kernel runs and counters, but nothing turned them into the numbers §4.9 asks for. A traced generation loop can now report on itself. - `Phases` / `Counters`: the vocabulary both sides agree on, so a typo cannot silently produce a metric of zero. Span helpers `prefill(tokens)`, `decodeStep(step)`, `sample(step)` and nested `module(path | TensorId)`. - `GenerationMetrics.from(events)`: TTFT (prompt pass through the first sampled token), prefill and decode tok/s, per-module breakdown ordered by cost, adapter count and bytes, kernel share of decode, page faults from the counter deltas inside the decode window, and **effective memory bandwidth** — bytes a decode step actually read ÷ how long it took, plus utilization when the device peak is known. Kernels outside the decode spans are excluded: counting the prompt pass would flatter the number. Rates are null rather than infinite when a span is too short for the platform clock (JS, Wasm), and `emitTo` publishes the derived numbers as counters so they appear beside the spans in Perfetto. - The decode harness opens the real spans — a prompt pass, per-weight module spans, sampling — so the metrics are asserted end to end on every target rather than only against a synthetic stream. - Benchmark JSON: `BenchmarkRecord.generation` (optional, so a matmul microbenchmark omits it), `GenerationMetrics.toRecord()`, and `scripts/check_engine_json.sh` validates the block whenever it appears — required fields, non-negative values, utilization ≤ 100 %, a sane module breakdown — and prints decode tok/s and GB/s alongside the primary metric. Verified: the checker accepts a record with generation metrics and one without, and rejects a missing field, a negative rate, a >100 % utilization and a negative module cost. Gate: scripts/pr-gate.sh — all legs passed; :skainet-backends:benchmarks:jvm-cpu-publish:test passes (that module has no CI leg of its own; the engine-benchmarks workflow exercises the script). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
📖 Documentation Preview The documentation has been built successfully for this PR. Generated Files:
Artifacts:
This comment will be updated automatically when the PR is updated. |
aharakal
approved these changes
Aug 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1035 · Phase P4 · Milestone M2 · PRD §4.9 metrics table · Proposal §4.9
What was missing
The trace stream already carried phases, kernel runs, adapters and counters — but nothing turned them into the numbers §4.9 asks for. A run could be traced and still not tell you its tok/s. This closes that: a traced generation loop reports on itself.
Vocabulary, then metrics
PhasesandCountersare the names both sides agree on, so a typo cannot silently yield a metric of zero. On top of them:GenerationMetrics.from(events)reads them back into: TTFT (start of the prompt pass to the end of the first sampled token — falling back to the first decode step when a loop does not trace sampling), prefill and decode tok/s, the per-module breakdown ordered by cost, adapter count and bytes, kernel share of decode, page faults from the counter deltas inside the decode window, and the effective memory bandwidth: bytes a decode step actually read ÷ how long it took, with utilization when the device peak is known.Two deliberate choices worth flagging:
null, never infinite. A span too short for the platform clock (JS, Wasm) legitimately times as 0 ns; the metric is absent rather thanInfinity.emitTo(sink)publishes the derived numbers as counters, so they land in the Perfetto trace as tracks beside the spans.Where it is proven
GenerationMetricsTest(11 cases, all targets): the event stream is hand-built with explicit timestamps, so TTFT, tok/s, bandwidth, utilization, adapter share, page-fault rate and the module ordering are asserted exactly rather than depending on clock resolution. Includes the empty-stream case (nulls, not division by zero) and the Perfetto counter tracks.GenerationMetricsHarnessTest(4 cases, all targets): the M1 decode harness now opens the real spans — prompt pass, per-weight module spans, sampling — so the metrics are exercised end to end throughKernelDispatch, not only against a synthetic stream. Every weight appears in the breakdown with one call per token; the decode step count and kernel-run count match exactly.Benchmark JSON
BenchmarkRecord.generationis optional — a matmul microbenchmark has no generation loop and omits the block entirely — withGenerationMetrics.toRecord()mapping onto the published snake_case names.scripts/check_engine_json.shvalidates the block whenever it appears: required fields, non-negative values, utilization ≤ 100 %, a well-formed module breakdown, and "decode steps recorded but never timed". It now also prints decode tok/s and GB/s beside the primary metric.Verified by hand on generated fixtures — the checker accepts a record with the block and one without, and rejects each of these:
The real feed remains the decode sample in SKaiNET-transformers, which owns a model; what lands here is the vocabulary, the reader, the harness proof and the validated schema the sample will publish into.
Gate
scripts/pr-gate.sh— all legs passed:jvmTest·apiCheck·verifyNpmPins jsTest wasmJsTest wasmWasiTest·linuxX64Test·assemble·:skainet-test:skainet-test-java:test.:skainet-backends:benchmarks:jvm-cpu-publish:testpasses. Note that module has no CI leg of its own — it is a JVM-only module, so the repo-widejvmTestleg does not reach it; theengine-benchmarksworkflow exercises the script path instead.Keeps develop green by
Purely additive instrumentation: new files, one optional serialized field with a default, and the harness switching from
phase("decode", step)todecodeStep(step)— the same phase name, so the existing M1 acceptance assertions are untouched and still pass. No hot path is instrumented, and every helper is a no-op when the sink is disabled (the default).🤖 Generated with Claude Code