Skip to content

feat(memory): generation phase spans and derived metrics — TTFT, tok/s, effective bandwidth (SKEEP-003 P4, S2.2) - #1082

Merged
michalharakal merged 1 commit into
developfrom
feature/1035-phase-markers-bandwidth
Aug 24, 2026
Merged

michalharakal merged 1 commit into
developfrom
feature/1035-phase-markers-bandwidth

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

Closes #1035 · Phase P4 · Milestone M2 · PRD §4.9 metrics table · Proposal §4.9

What was missing

The trace stream already carried phases, kernel runs, adapters and counters — but nothing turned them into the numbers §4.9 asks for. A run could be traced and still not tell you its tok/s. This closes that: a traced generation loop reports on itself.

Vocabulary, then metrics

Phases and Counters are the names both sides agree on, so a typo cannot silently yield a metric of zero. On top of them:

sink.prefill(tokens = 128) { … }
sink.decodeStep(step) { sink.module("model.layers[3].attn") { … } }
sink.sample(step) { … }

GenerationMetrics.from(events) reads them back into: TTFT (start of the prompt pass to the end of the first sampled token — falling back to the first decode step when a loop does not trace sampling), prefill and decode tok/s, the per-module breakdown ordered by cost, adapter count and bytes, kernel share of decode, page faults from the counter deltas inside the decode window, and the effective memory bandwidth: bytes a decode step actually read ÷ how long it took, with utilization when the device peak is known.

Two deliberate choices worth flagging:

  • Kernels outside the decode spans do not count. The prompt pass runs the same kernels over far more tokens; folding those bytes into the bandwidth figure would flatter it. There is a test that adds a fat prefill kernel and asserts the decode number does not move.
  • Rates are null, never infinite. A span too short for the platform clock (JS, Wasm) legitimately times as 0 ns; the metric is absent rather than Infinity.

emitTo(sink) publishes the derived numbers as counters, so they land in the Perfetto trace as tracks beside the spans.

Where it is proven

  • GenerationMetricsTest (11 cases, all targets): the event stream is hand-built with explicit timestamps, so TTFT, tok/s, bandwidth, utilization, adapter share, page-fault rate and the module ordering are asserted exactly rather than depending on clock resolution. Includes the empty-stream case (nulls, not division by zero) and the Perfetto counter tracks.
  • GenerationMetricsHarnessTest (4 cases, all targets): the M1 decode harness now opens the real spans — prompt pass, per-weight module spans, sampling — so the metrics are exercised end to end through KernelDispatch, not only against a synthetic stream. Every weight appears in the breakdown with one call per token; the decode step count and kernel-run count match exactly.

Benchmark JSON

BenchmarkRecord.generation is optional — a matmul microbenchmark has no generation loop and omits the block entirely — with GenerationMetrics.toRecord() mapping onto the published snake_case names. scripts/check_engine_json.sh validates the block whenever it appears: required fields, non-negative values, utilization ≤ 100 %, a well-formed module breakdown, and "decode steps recorded but never timed". It now also prints decode tok/s and GB/s beside the primary metric.

Verified by hand on generated fixtures — the checker accepts a record with the block and one without, and rejects each of these:

FAIL missing_field.json: generation missing ['bytes_read']
FAIL neg_rate.json:      generation.decode_tokens_per_second=-3.0 must be a non-negative number or null
FAIL over_peak.json:     generation.bandwidth_utilization_percent=140.0 exceeds 100%
FAIL neg_module.json:    generation.module_breakdown_ms['model.layers[0].attn']=-1.0 must be non-negative

The real feed remains the decode sample in SKaiNET-transformers, which owns a model; what lands here is the vocabulary, the reader, the harness proof and the validated schema the sample will publish into.

Gate

scripts/pr-gate.sh — all legs passed: jvmTest · apiCheck · verifyNpmPins jsTest wasmJsTest wasmWasiTest · linuxX64Test · assemble · :skainet-test:skainet-test-java:test.
:skainet-backends:benchmarks:jvm-cpu-publish:test passes. Note that module has no CI leg of its own — it is a JVM-only module, so the repo-wide jvmTest leg does not reach it; the engine-benchmarks workflow exercises the script path instead.

Keeps develop green by

Purely additive instrumentation: new files, one optional serialized field with a default, and the harness switching from phase("decode", step) to decodeStep(step) — the same phase name, so the existing M1 acceptance assertions are untouched and still pass. No hot path is instrumented, and every helper is a no-op when the sink is disabled (the default).

🤖 Generated with Claude Code

…s, effective bandwidth, per-module breakdown

Closes #1035 (SKEEP-003 P4, S2.2, proposal §4.9).

The trace stream carried phases, kernel runs and counters, but nothing
turned them into the numbers §4.9 asks for. A traced generation loop can
now report on itself.

- `Phases` / `Counters`: the vocabulary both sides agree on, so a typo
  cannot silently produce a metric of zero. Span helpers `prefill(tokens)`,
  `decodeStep(step)`, `sample(step)` and nested `module(path | TensorId)`.
- `GenerationMetrics.from(events)`: TTFT (prompt pass through the first
  sampled token), prefill and decode tok/s, per-module breakdown ordered by
  cost, adapter count and bytes, kernel share of decode, page faults from
  the counter deltas inside the decode window, and **effective memory
  bandwidth** — bytes a decode step actually read ÷ how long it took, plus
  utilization when the device peak is known. Kernels outside the decode
  spans are excluded: counting the prompt pass would flatter the number.
  Rates are null rather than infinite when a span is too short for the
  platform clock (JS, Wasm), and `emitTo` publishes the derived numbers as
  counters so they appear beside the spans in Perfetto.
- The decode harness opens the real spans — a prompt pass, per-weight
  module spans, sampling — so the metrics are asserted end to end on every
  target rather than only against a synthetic stream.
- Benchmark JSON: `BenchmarkRecord.generation` (optional, so a matmul
  microbenchmark omits it), `GenerationMetrics.toRecord()`, and
  `scripts/check_engine_json.sh` validates the block whenever it appears —
  required fields, non-negative values, utilization ≤ 100 %, a sane module
  breakdown — and prints decode tok/s and GB/s alongside the primary
  metric.

Verified: the checker accepts a record with generation metrics and one
without, and rejects a missing field, a negative rate, a >100 % utilization
and a negative module cost.

Gate: scripts/pr-gate.sh — all legs passed;
:skainet-backends:benchmarks:jvm-cpu-publish:test passes (that module has
no CI leg of its own; the engine-benchmarks workflow exercises the script).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

📖 Documentation Preview

The documentation has been built successfully for this PR.

Generated Files:

  • Operator documentation: docs/modules/operators/_generated_/
  • JSON schema output: operators.json

Artifacts:

  • Download the documentation-preview-1082 artifact to view the complete documentation locally.

This comment will be updated automatically when the PR is updated.

@michalharakal
michalharakal requested a review from aharakal August 24, 2026 08:33
@michalharakal
michalharakal merged commit 40bae31 into develop Aug 24, 2026
19 checks passed
@michalharakal
michalharakal deleted the feature/1035-phase-markers-bandwidth branch August 24, 2026 08:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[S2.2] P4: phase markers + effective-bandwidth metric in the generation loop (core API; emitted by the decode sample)

2 participants