Utilization observables for the structural model (#209 M2b) - #212
Conversation
Adds util:{service}:{resource} as its own observable family (never folded into
sig:{service}) to recover locally silent RESOURCE faults — a CPU- or event-loop-
saturated service whose requests still look normal in traces. Trace evidence
plateaued at 4/8 (M2a); the ceiling survey found clean, attributable metric
separation for 2 of the 4 remaining misses.
Frozen semantics (posted to #209 before implementing):
- Metrics are classified by OpenTelemetry semantic-convention NAME SHAPE, not by
names picked from the corpus (util_metric_class): a utilization gauge (last
segment ends in "utilization"; resource = preceding segment, e.g. cpu,
eventloop, memory) or a cumulative CPU-time counter (*.cpu.time -> cpu).
Collector/SDK self-telemetry (otel.sdk.*, otelcol*) is excluded — it moves when
a service merely emits more spans (the product-catalog false-positive trap).
- summarize_utilization: gauges compare incident vs baseline MEAN (periodic
aggregates); a malformed sample leaves the metric unmeasured. CPU-time counters
compare incident vs baseline RATE, trusted only for a verifiably single-series,
monotonic stream — the converter flattens attribute dimensions, so interleaved
series (e.g. process.cpu.time state=user/system: 2 samples/timestamp, decreases)
are UNKNOWN rather than guessed. Unattributed samples (service=None) are not
evidence. PRESENT iff ratio >= 2x; ABSENT if measured below (a drop is not a
saturation anomaly); else UNKNOWN. The (service, resource) observable takes the
witness metric (PRESENT > ABSENT > UNKNOWN).
- build_hypotheses: a service with a PRESENT util observable joins the candidates
and process:{S} expects its util coordinates PRESENT. Soft only; without util
signals the M2a model is byte-identical.
- structural_signals now returns a StructuralInputs NamedTuple (signals, edges,
edge_signals, util_signals); both runners use it — still one model.
19 new unit tests (name classes, self-telemetry exclusion, gauge/counter rules,
interleaved + reset counters, witness, model wiring, end-to-end locally silent CPU
fault). Full unit suite 1515; structural/trigger/absence integration green.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
REQUEST CHANGES
The candidate wiring is real: a PRESENT util: coordinate joins process:{S}, UNKNOWN stays off F_usable, self-telemetry and service=None contribute nothing, and omitting util_signals leaves the M2a hypothesis set alone. That is not what util:{service}:{resource} measures. The reducer keys a bag of (timestamp, value) pairs by metric name, then emits OBSERVED ABSENT for series the semantic conventions it cites are required to split. CI is green. This is not a failing-test problem. The new tests never put two values on one timestamp except for the counter case they already special-cased.
1. A multi-mode utilization gauge is averaged into OBSERVED ABSENT
BLOCKING.
summarize_utilization says a gauge is incident mean / baseline mean, and that a bad sample is never averaged into false-normal evidence. util_metric_class then names system.cpu.utilization and system.memory.utilization as the canonical gauges. Both instruments are per-state: cpu.mode (user, system, idle, …) and system.memory.state (used, free, …). The gauge branch never looks at that. Every sample with the same name is one series.
Take system.cpu.utilization scraped twice per timestamp, baseline idle=0.9 / user=0.1, incident idle=0.1 / user=0.9. Both window means are 0.5. Ratio is 1. discretize_ratio returns NORMAL, and the metric is ABSENT. The CPU is saturated. The observable says it was measured and it was not. system.memory.utilization does the same thing with used against free: the mean does not move, so memory pressure is ABSENT too.
ABSENT is not an internal scratch value. build_observables emits it, so it enters F_usable. A service that is already a candidate then predicts that coordinate PRESENT and the packet records a mismatch against a measurement that was never a utilization. A service that is only visible through this gauge does not become a candidate at all, because only PRESENT joins. That is the locally-silent fault this PR exists to recover, reported as measured-normal.
The counter branch underneath this, in the same function, already refuses a flattened state=user / state=system stream. The gauge branch does not. ಠ_ಠ The single-series rule was not a counter quirk. It was the only thing standing between a collapsed attribute dimension and a false ABSENT. jvm.cpu.recent_utilization and nodejs.eventloop.utilization survive only because those particular instruments are one series per scrape. The abstraction does not.
The tests construct one value per timestamp, so an implementation that averages cpu.mode still passes test_semantic_convention_shapes and the end-to-end CPU case. Fix the representation: a gauge with more than one sample at a timestamp is unmeasured, same as a counter, until the series key (cpu.mode, system.memory.state, …) is actually on the sample. Do not emit ABSENT for a mean over modes.
2. The counter "single-series" check is not a series check, and the product path cannot see the collisions it requires
MAJOR.
single is len({timestamps}) == len(points) plus non-decreasing values. Two cumulative series with distinct timestamps and a merged sequence that happens to be monotonic are accepted and rated. That is not "verifiably single-series". It is "no timestamp collision, and the blend went upwards".
On the path build_structural_view actually runs, the collision clause is dead. metric_pk is scope|service|metric|ts, and persist_metric_samples inserts that id with ON CONFLICT DO NOTHING. Attributes are not part of the key — parse_otlp_metrics does not even copy datapoint attributes onto the sample. Two modes at one timestamp cannot both be rows in metric_samples. The unit test that pins UNKNOWN feeds summarize_utilization a list that persistence has already made unrepresentable. Shadow eval still reads jsonl, where both lines survive, which is why the otel-fresh number and the unit test agree with each other and not with the product view.
Whatever point was inserted first becomes the whole series. If that point is idle, saturation is a drop, and a drop is defined to be ABSENT. Insert order is not an invariant.
metric_type is already on the row (counter vs sum vs gauge; null is specified as gauge on MetricSample) and this classifier never reads it. A delta *.cpu.time usually fails the monotonic check and falls out as UNKNOWN, so that particular lie is latent. Name shape is not instrument semantics. Do not describe it as such.
The guard has to run on the samples the product view loads, or the product view must not claim it. Minimum: series identity has to exist before a measured ABSENT/PRESENT is legal. A primary key that deletes the second point, and a check that only works when the second point is still in the list, are not that.
VERDICT
This is an abstraction failure. UtilSignal promises a per-resource utilization measurement with the same UNKNOWN discipline as sig. The runtime measures a name-keyed bag of numbers. One bad float voids the metric; a second attribute series is averaged, or, after ingest, silently discarded. The M2b recoveries (jvm.cpu.recent_utilization, nodejs.eventloop.utilization) are single-series gauges and do not exercise this. The next corpus that actually contains system.cpu.utilization or system.memory.utilization will record those resources as measured-normal while they are pegged.
Do not merge until a collapsed series is UNKNOWN rather than ABSENT, on the input build_structural_view reads and not only on a hand-built list.
Sent by Cursor Automation: Code Reviewer
…212 review) Addresses both points of the #212 review. Our metric samples had no series identity anywhere in the pipeline, so M2b's "single-series" guarantee was not real: 1. BLOCKING — per-state gauges were averaged into OBSERVED ABSENT. system.cpu.utilization (per cpu.mode) / system.memory.utilization (per state): with no series key, a saturated CPU (idle 0.9->0.1, user 0.1->0.9) averaged to a flat 0.5 and was emitted as measured-normal. 2. MAJOR — the counter "single-series" check was not a series check, and the product path could not see collisions: metric_pk = scope|service|metric|ts with ON CONFLICT DO NOTHING silently DROPPED a second series at the same timestamp (whichever row landed first became "the" series); offset-timestamp series with a monotonic blend passed; metric_type was never read. Fix — series identity end to end: - Converter (parse_otlp_metrics) records each datapoint's series identity: datapoint attributes + the resource service.instance.id, always a dict for OTLP-derived samples ({} = attribute-free datapoint; None = never recorded). - metric_pk includes the canonical series identity when present, so two series at one timestamp are two rows (a general data-loss bug). Identity-less sources keep their original key, so their idempotent re-ingest is unchanged. - The jsonl corpus writer/loader round-trips attributes (omitted when None). - summarize_utilization measures a metric ONLY when every sample carries identity, they form exactly one series, no two samples share a timestamp, and the declared instrument (metric_type) is the one the reducer needs (gauge; cumulative counter for *.cpu.time — the name gives resource + required instrument, never the instrument itself). Anything else is UNKNOWN, never a measured ABSENT. A decrease inside a counter series is a reset -> UNKNOWN. Consequence (stated in the PR): the otel-fresh corpus carries no series identity, so M2b measures nothing on it and the earlier 6/8 claim is withdrawn — it rested on an unverified single-series assumption. M2b becomes measurable on M3's fresh capture with the fixed converter. Tests: util suite rewritten around series identity incl. the reviewer's per-mode CPU and offset-interleaved-counter cases, multiple instances, no-identity, instrument mismatch; converter/key/jsonl identity unit tests; product-path integration test (two series at one ts both persist -> UNKNOWN; single verified series through the DB -> measured; identity-less rows -> nothing claimed). Unit 1527, integration green. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
/review |
There was a problem hiding this comment.
REQUEST CHANGES
The previous round is fixed. A per-mode gauge is no longer averaged into OBSERVED ABSENT, the counter check is an actual series key, metric_type is read, and metric_pk keeps both rows. The integration test loads those rows and gets UNKNOWN. CI is green (run 36236534703). This is not a failing-test problem.
What the fuck is "one verified series" supposed to guarantee? summarize_utilization says a metric is measured only when every sample carries identity and the samples are exactly one series, and that missing identity must never be averaged into ABSENT — including several reporting instances. The converter then stamps every OTLP datapoint with a dict, and {} means "no attributes". single_series treats that dict as proof. Two replicas that do not carry service.instance.id, scraped on offset timestamps, are one series. A calm replica plus a pegged one comes out measured-normal.
The same key change feeds the default explain path. metric_anomaly_onsets still reduces (service, metric) and does not even load attributes. The util reducer refuses the blend. The onset reducer performs it, on every explain that has a primary cluster, not behind --structural.
1. Empty identity is a verified series
BLOCKING.
summarize_utilization documents the guard: no recorded identity, and the metric is UNKNOWN, because otherwise several instances get averaged into a number that is not one resource's utilization. _series_key implements "no identity" as attributes is None. parse_otlp_metrics never produces None. It always produces a dict, and it copies service.instance.id only when the resource happens to have one. No datapoint attributes and no instance id become {}. None not in keys and len(keys) == 1 is then true.
Concrete input, jvm.cpu.recent_utilization, two replicas, attributes {}, timestamps offset by 5s so distinct_ts passes:
- replica A: baseline 0.40, incident 0.45
- replica B: baseline 0.40, incident 0.90
Baseline mean 0.40, incident mean 0.675, ratio 1.6875. discretize_ratio is NORMAL. The metric is ABSENT. Replica B alone is 2.25× and would be PRESENT. The service is saturated and the observable says it was measured and it was not. Same timestamps would have failed distinct_ts and stayed UNKNOWN. OTel periodic readers are per-process, so the offset case is the normal one. ಠ_ಠ
test_two_reporting_instances_are_unknown gives the replicas different instance ids. test_attribute_free_datapoint_still_records_identity asserts {} and never summarizes two of them. An implementation that measures {} passes the suite.
Do not treat a missing instance id as the empty series. If the converter cannot name the instance, that sample's identity is incomplete (None, or an explicit unknown), and summarize_utilization must keep the metric UNKNOWN. A dict you synthesized because you looked and found nothing is not a verification.
2. The onset path averages the series this key just stopped dropping
MAJOR.
metric_pk now keeps every attribute-series. That is the right storage change for the util reducer. It is not local to it. _metric_onsets selects service, metric, value, ts — not attributes — and metric_anomaly_onsets keys (service, metric) and calls that a series. metric_features does the same mean. Both run on the default explain path. --structural does not gate them.
system.cpu.utilization with idle 0.9→0.1 and user 0.1→0.9: the util reducer returns UNKNOWN. The onset reducer sees both values at every timestamp, the mean stays 0.5, and the saturation never becomes an onset. Before this patch the second point was deleted and whichever mode landed could still move. The "data-loss fix" deletes that movement for every consumer that is not summarize_utilization.
Frozen otel-fresh does not show it: those rows have attributes is None, so the old key still collapses them. The converter in this PR is what populates attributes. The next capture is the regression. "Eval delta: none, because --structural is opt-in" does not cover a primary-key change on the trigger path.
The series key has to be the key those reducers group by, or they have to refuse a metric that has more than one. Leaving them on (service, metric) while the table no longer means that is how the next feature reintroduces the blend this PR just forbade.
VERDICT
This is the same abstraction failure as the last round, one layer out. Series identity is real for datapoint attributes and for an instance id that is present. It is fictional for the empty dict the converter emits when the instance id is absent, and it is ignored by the reducers that already owned metric samples. The util function's UNKNOWN discipline holds only for the shapes the new tests build.
Do not merge until a missing instance id cannot become OBSERVED ABSENT, and until metric_anomaly_onsets / metric_features do not average across the series metric_pk now retains.
Sent by Cursor Automation: Code Reviewer
…per series (#212 review r2) Both points of the second #212 review: 1. BLOCKING — an empty identity was treated as a verified series. The converter stamped {} when a datapoint had no attributes and the resource no instance id, and summarize_utilization accepted it: two unnamed replicas on offset timestamps (A 0.40->0.45, B 0.40->0.90) blended to 1.69x and read as measured-normal while B was saturated. Now a series is verified only if its instance is NAMED. New shared module src/core/rca/metric_series defines identity once: series_key, datapoint signature, and instance_identity from the semconv resource attributes that name one process (service.instance.id, k8s.pod.uid, container.id, or host.name + process.pid). The converter records whichever exist. summarize_utilization: unnamed instance -> UNKNOWN; more than one datapoint signature (modes, e.g. cpu.mode idle/user) -> UNKNOWN; under one mode each named instance is measured separately — PRESENT if any replica is saturated, ABSENT only if every replica is measured-normal. 2. MAJOR — other reducers still averaged the series metric_pk now retains. metric_anomaly_onsets (trigger path; _metric_onsets now also SELECTs attributes), metric_features (ranker features) and summarize_metrics (structural sig) now group by (service, metric, series_key): features take the max over series, onsets are detected per series, the sig branches are anomalous if ANY series is and measured-normal only if EVERY series is. Identity-less samples form one group per (service, metric), exactly as before. Eval delta (measured, not assumed): raglogs eval on trace-loc with the learned ranker + rare_event trigger path, main vs this branch: 0/24 cases differ in root cause, predicted services, top trigger or confidence; aggregates identical. All existing corpora carry no series identity, so the default path is byte-identical on them; series-aware reduction applies to identity-carrying captures. Tests: reviewer's unnamed-replica case (unit + from the DB), named replicas (calm+pegged -> PRESENT, both calm -> ABSENT, one unmeasured -> UNKNOWN), each semconv instance attribute (host.name alone does not name), per-series reducer tests with identity-less equivalence, and a DB test that _metric_onsets and compute_features see every cpu.mode series. Unit 1538, integration 52 green. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
/review |
There was a problem hiding this comment.
COMMENT
The two blockers from the last look are closed. {} is not a verified series: instance_identity is None, and the reviewer's offset replicas (0.40/0.45 and 0.40/0.90) stay UNKNOWN, including when the rows are read back from metric_samples. Idle/user is no longer one mean. metric_anomaly_onsets, metric_features, and summarize_metrics group by series_key, and _metric_onsets loads attributes. Identity-less samples still collapse to one (service, metric) group. CI is green (run 36238744554: test 3.10/3.12, openapi, docker, migrations). This is not a failing-test problem.
The utilization contract holds: one datapoint signature, every sample names an instance, PRESENT if any replica is saturated, ABSENT only if every replica is measured and below. I am not blocking the merge on the onset budget below. It does not invent an ABSENT utilization. It does throw away the series key on the default explain path the moment there are more series than max_candidates.
1. Per-series onsets are emitted as one metric, then cut at five
MAJOR.
metric_anomaly_onsets now detects one onset per (service, metric, series_key). AnomalyOnset still has only service and metric. find_trigger_candidates renders every one of them as "{metric} anomaly onset". The list is then sliced to max_candidates (default 5), a budget sized when a metric was one series.
system.cpu.utilization is the instrument this PR uses as the example. Host metrics carry on the order of eight cpu.mode values (idle, user, system, nice, iowait, irq, softirq, steal). A saturation moves several of them. Each series appends an onset. The five that win are five copies of the same metric name. A different instrument on the cause service, with a smaller relative deviation, is not in the list.
Near-zero modes make this worse. A flat baseline uses dev / |mean|. steal at 0.001→0.02 is magnitude 19. jvm.cpu.recent_utilization at 0.2→0.9 is magnitude 3.5. The small mode outranks the gauge. Same max ratio lands in metric_features as met_anom. The learned ranker was trained on blended, identity-less means. Existing corpora hide it (eval 0/24, attributes None). The converter in this PR is what starts writing cpu.mode.
The opening line of the docstring still says "earliest per-(service, metric)". The next paragraph says per series. test_per_mode_onset_is_detected builds two series and asserts that some onset exists. Two is less than five. An implementation that drops every series after the first still passes.
Carry the series key on the onset, or rank one onset per (service, metric) after the per-series detection (the max magnitude, not the mean). Do not spend the whole trigger budget on indistinguishable copies of one instrument.
The test module docstring still says a metric is measured only when every sample belongs to exactly one series. test_named_replicas_are_measured_per_instance expects two named replicas to be PRESENT. The production docstring is the one that matches the code. The test header is the previous contract.
VERDICT
The utilization observable does what it now claims. A missing instance id cannot become OBSERVED ABSENT, modes are not averaged, and the other reducers no longer blend distinct series_keys into one mean. That was the merge bar from the last review, and it is met. The remaining hole is the trigger result: series identity is computed and then discarded, and a cap of five turns one multi-mode instrument into the entire candidate list. That is a default-path defect, not a reason to call the util measurement unsound.
Merge is acceptable for M2b. Fix the onset budget before treating trigger candidates on an identity-carrying capture as meaningful. The trace-loc eval delta does not cover that capture.
Sent by Cursor Automation: Code Reviewer
…the budget (#212 review r3) Review r3 (merge acceptable for M2b; one default-path MAJOR + a stale docstring): - metric_anomaly_onsets detected per series but emitted onsets that carry only (service, metric), then cut the list at max_candidates=5. One multi-mode instrument (eight cpu.mode series of system.cpu.utilization) could spend the whole budget on indistinguishable copies and push a different instrument on the cause service off the list. Detection stays per series; the candidate is now ONE onset per (service, metric) — its strongest series (highest score; the max, never a mean) — before ranking and the cap. Docstring's first line corrected. For identity-less data each (service, metric) is already one series, so this is a no-op there: default-path eval on trace-loc (ranker + rare_event) vs main still 0/24 cases differ, aggregates identical. - test_structural_util_signals header now states the actual contract (named instance, one datapoint signature, any saturated replica PRESENT) instead of the superseded single-series refusal. New test: eight moving cpu.mode series + a different instrument on another service with max_candidates=5 -> exactly one system.cpu.utilization onset and the other instrument present (fails on the previous code). Unit 1539. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
/review |


What (M2b of #209)
Adds
util:{service}:{resource}as its own observable family, never folded intosig:{service}, to recover locally silent resource faults: a CPU- or event-loop-saturated service whose requests still look normal in traces. It also adds the prerequisite this revision found missing: metric series identity through ingestion.Series identity (added after reviews r1–r2)
A metric name is not a series: one instrument emits several series, per datapoint attribute (
cpu.mode,system.memory.state, route) and per reporting instance (replicas). The first revisions averaged across them. On the product path,metric_pk = scope|service|metric|tswithON CONFLICT DO NOTHINGalso silently dropped a second series at the same timestamp. Identity is now defined once, insrc/core/rca/metric_series, and carried end to end:service.instance.id,k8s.pod.uid,container.id, orhost.name+process.pid).None= the source never recorded identity.{}included) is recorded but not a verified series: it can't tell two replicas apart.attributes.metric_features(ranker) takes the max over series;metric_anomaly_onsets(trigger path;_metric_onsetsnow loadsattributes) detects onsets per series, then keeps one candidate per(service, metric)(its strongest series). That way several series of one instrument can't fill themax_candidatesbudget with indistinguishable copies;summarize_metrics(structural sig) is anomalous if any series is, and measured-normal only if every series is.(service, metric), exactly as before.Utilization semantics
util_metric_classmaps a name, by OpenTelemetry semantic-convention shape, to (resource, required instrument):*utilization→ gauge; the resource is the preceding segment (cpu,eventloop,memory, …);*.cpu.time→ cumulative counter.metric_typemust match. Self-telemetry (otel.sdk.*,otelcol*) is excluded.summarize_utilizationis UNKNOWN (never a measured ABSENT) unless every sample has recorded identity, names its instance, shares one datapoint signature, and matches the declared instrument.cpu.modeidle/user) whose direction differs, so there are no per-mode semantics yet.(service, resource)observable takes the witness metric.util:observable adds its service as a candidate, andprocess:{S}expects itsutil:coordinates PRESENT. Soft only.structural_signalsreturns aStructuralInputsNamedTuple used by both runners — still one model.Measurement (otel-fresh, frozen protocol) — the earlier 6/8 is withdrawn
The otel-fresh corpus carries no series identity — 0 of 37,282 utilization/CPU samples has
attributes, and only 9% even name a service. So M2b now measures nothing on it, which is the correct outcome.The first revision reported 6/8, recovering adHighCpu and loadGeneratorFlood. That rested on an unverified single-series assumption — true for those two instruments in this deployment, but not a property the pipeline could check. Validating M2b needs M3's fresh capture with this converter, which now records identity.
Tests
metric_samplesrowsbuild_structural_viewreads):Full unit suite 1539; all 52 integration tests green.
Scope / eval delta
src/core/rca/metric_series.py(new),structural_model.py,features.py,triggers.py,structural.py,src/core/explain/evidence.py,src/core/ingestion/telemetry.py,src/eval/otlp.py,src/eval/rcaeval.py,src/eval/structural_shadow.py.Eval delta, measured:
raglogs evalon trace-loc with the learned ranker and the rare_event trigger path enabled (the reducers whose grouping changed),mainvs this branch:Existing corpora carry no series identity, so the default path is byte-identical on them. Series-aware reduction applies to identity-carrying captures.
--structuralis opt-in; on otel-fresh M2b measures nothing (see above).Known follow-up for identity-carrying captures (not blocking; recorded for M3). Relative change on near-zero series is large: a flat-baseline
cpu.mode=stealat 0.001→0.02 has magnitude 19, versusjvm.cpu.recent_utilizationat 0.2→0.9 with 3.5. So such series can outrank other instruments in onset ranking, and dominatemet_anom(the max over series). The learned ranker was trained on identity-less, blended features, so its feature distribution will shift on the first identity-carrying capture. M3 should re-evaluate the ranker (and the relative-magnitude rule) on that capture before trusting ranker or trigger output there. None of this affects existing corpora.Next: M3 — a fresh frozen capture with the identity-carrying converter validates M1 + M2a + M2b out of sample, including the pre-registered edge-latency and IDENTIFIED-precision follow-ups.
🤖 Generated with Claude Code