Skip to content

feat(eval): frozen M3 evaluator for the structural model (#209) - #214

Merged
leo-aa88 merged 3 commits into
mainfrom
feat/m3-structural-eval-209
Sep 26, 2026
Merged

leo-aa88 merged 3 commits into
mainfrom
feat/m3-structural-eval-209

Conversation

@leo-aa88

@leo-aa88 leo-aa88 commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

This is step 3 of the M3 protocol frozen on #209 (protocol): the evaluator is committed and reviewed before the out-of-sample capture exists. The measurement is therefore fixed before any data can shape it. After merge, the maintainer captures the corpus on a dedicated non-prod cluster (with the #213 guard), and this runs once.

What it measures

src/eval/structural_m3.py is pure and needs no DB.

  • Three nested arms over the same inputs:

    • M1 is sig only;
    • M2a adds edge:;
    • M2a+M2b adds util: and is the frozen model.

    Each arm is the product's own resolve step on a trimmed StructuralInputs. Candidate generation from the incident call graph is M1 behaviour and stays in every arm.

  • Per-arm metrics, all of them reported whatever the outcome:

    • generation;
    • truth retained, over all positives and over those whose cause has telemetry;
    • median candidates and median candidate fraction;
    • healthy abstention;
    • the outcome distribution;
    • IDENTIFIED precision, plus IDENTIFIED on healthy windows.
  • Per-case attribution: which family witnessed the retained truth (sig / edge / util). A cause retained with no PRESENT observable about it is labelled topology.

  • Edge specificity: the healthy-window flagged-edge rate at the unchanged ≥2× cutoff, and the latency-only share.

  • Series-identity coverage: the share of utilization samples that name their instance.

  • Default path: raglogs eval scoring with the learned ranker and the rare_event trigger, i.e. the pre-registered ranker check on identity-carrying data.

Decision rules (decide())

Every threshold is a constant copied from the protocol. None is tunable. (Revised after review r1: the first version paraphrased the M2b rule, and this section now encodes it literally.)

  • Generalization, on the frozen arm:

    • generalizes requires truth retained (cause has telemetry) ≥ 0.50, median candidate fraction ≤ 0.33 and healthy abstention ≥ 0.58;
    • the result is untestable if there are no positives with cause telemetry or no healthy windows;
    • otherwise it is does_not_generalize.

    The bound is the frozen literal 0.33. otel-fresh M1's reference fraction is exactly 1/3, which fails it, and a test pins that. A 1/3 reading would be a protocol amendment, so it is not made here.

  • M2b, as frozen:

    • not_validated if any healthy window has a PRESENT util:;
    • untestable only when series-identity coverage is ~0, encoded exactly as no named utilization sample (rate in (None, 0.0); otel-fresh was 0/37282);
    • validated iff a resource-fault positive is retained via util. That is the arm delta: the frozen arm retains the cause, M2a (the same inputs without util) does not, and a util: on the cause is PRESENT;
    • otherwise not_validated.

    A co-present util on a fault that sig already retains does not validate.

  • Protocol deviations are reported: fewer than 12 healthy negatives, or a missing resource-fault scenario. A missing scenario cannot validate M2b.

The gated statistics use the otel-fresh reference definitions exactly.

  • The universe is every service in span or metric rows.
  • Generated means any hypothesis exists.
  • Selectivity is taken over generated positives, with an empty compatible set counting as 0.
  • Healthy abstention means no hypothesis generated. My first version said "no localization claim" and claimed that matched the reference. It did not, and that is corrected. "No localization claim" is reported as a secondary, ungated figure.

Reproduction check. Rerun on the spent otel-fresh corpus, the evaluator reproduces every reference number:

  • M1: generation 5/9, retained 3/8, fraction 1/3, abstention 11/12;
  • M2a: 8/9, 4/8, 0.139, 7/12, IDENTIFIED 0/8;
  • M2a+M2b equals M2a, and M2b is untestable (identity coverage 0/37282);
  • healthy edges at ≥2× latency in 5/12 windows, and 0/12 on the error branch.

Edge specificity

The ≥2× latency rate is counted on its own, including edges whose error branch also fired. The error-branch rate and the combined PRESENT rate are reported alongside it. The results post labels UNOBSERVABLE cases, meaning no telemetry from the cause.

The driver (scripts/eval/m3_structural.py)

  • It refuses a dirty working tree, because the frozen model must be an exact commit. --allow-dirty exists for smoke runs only and is stamped DIRTY.
  • It records the model commit, a sha256 of the corpus and the trigger mode.
  • It loads the ranker and calibrator with the product loaders before any DB work, and refuses (exit 2) if either fails (review r2). The product falls back silently to the volume selector, and on a one-shot run that would be a false "ranker" result. Absolute paths, sha256s and *_loaded go into provenance. It also refuses if settings did not take rare_event or the ranker path.
  • It ingests once (the harness ingest), runs the default path, then reads each case's persisted rows exactly as the product does and runs the arms.
  • It writes the JSON report plus the as-is markdown post for EPIC: Real-telemetry causal observable grounding — make OTel evidence usable by the structural engine #209.

src/core change (behaviour-preserving)

build_structural_view is split into load_structural_rows and resolve_structural, so the evaluator reads the same rows and runs the same resolve(part, observations=…) as the product. My ad-hoc M1–M2b scripts called resolve without observations. That was checked: observations only feeds integration_gaps and never the outcome or localization, so those numbers stand. The shared step prevents any future drift. tests/integration/test_structural_m3_parity.py asserts that the frozen arm equals build_structural_view's outcome and localization on the same DB rows.

Eval delta

No change. build_structural_view (outcome, localization and Phase G ranking) was run per case on main and on this branch:

corpus cases differ structural view generated
trace-loc 24 0 24
otel-fresh 21 0 13

The default explain path doesn't call the structural view, so it is untouched.

A smoke run on trace-loc (spent corpus, --allow-dirty) produced 24/24 retained in every arm, the same as before. It also correctly flagged that corpus's deviations: 0 healthy negatives and no resource faults. Runtime was 16s.

Tests

  • tests/unit/test_structural_m3.py (32 tests):
    • nested arms: an unreachable callee recovered only by edge, and a locally silent CPU fault recovered only by util;
    • an unnamed instance counts against identity coverage and is never measured;
    • healthy windows abstain;
    • every decision rule at and around its bound, including 7/12 passing at 0.58 and exact 1/3 failing 0.33;
    • the reviewer's two M2b counterexamples: a sig-retained resource fault with PRESENT util is not_validated; coverage 0.8 with unmeasured causes is not_validated, while coverage 0 is untestable;
    • edge specificity: an error-only edge is not a latency flag, and a 3× edge with errors still counts;
    • deviations, topology attribution, and rendering.
  • The integration parity test above.

make lint ✅ · make test-unit ✅ (1584) · integration (structural parity + util series identity) ✅

Refs #209.

🤖 Generated with Claude Code

The M3 protocol on #209 requires the evaluator to be committed and
reviewed before the out-of-sample capture exists, so the measurement
cannot be shaped by the data. This encodes it:

- src/eval/structural_m3.py (pure):
  - three nested arms over the same inputs: M1 (sig), M2a (+edge),
    M2a+M2b (+util, the frozen model);
  - per-case attribution of which observable family witnessed the
    retained truth (sig / edge / util / topology);
  - healthy-window edge specificity at the unchanged cutoff;
  - series-identity coverage;
  - decide(), which applies the pre-registered thresholds (truth
    retained >= 0.50, median candidate fraction <= 0.33, healthy
    abstention >= 0.58; M2b validated / not_validated / untestable)
    and reports protocol deviations.
- scripts/eval/m3_structural.py:
  - refuses a dirty tree and records the model commit + corpus
    sha256;
  - ingests once, runs the default path (ranker + rare_event) and the
    three arms on the persisted rows.
- src/core/rca/structural.py: build_structural_view is split into
  load_structural_rows + resolve_structural, so the evaluator reads
  the same rows and runs the same resolve step as the product
  (behaviour-preserving; an integration test asserts parity).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@leo-aa88

Copy link
Copy Markdown
Member Author

/review

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

REQUEST CHANGES

load_structural_rows / resolve_structural is the right cut. The frozen arm calls the product's resolve, and the parity test locks outcome and localization to that path. decide() then ignores the arm structure that cut exists to make comparable, and it does not implement the M3 rule frozen on #209.

CI is green (test 3.10, test 3.12, docker, migrations, openapi). This is not a failing-test problem. The unit tests build CaseEvals that already agree with decide().

1. decide() does not implement the frozen M2b rule

BLOCKING.

The protocol (issue #209, comment 5846202221) says M2b is validated only when a resource-fault positive is retained via a util: observable on a named-instance series, and no healthy window has a PRESENT util:. It is untestable when series-identity coverage is ~0. It is not validated when util fires on a healthy window.

decide() does two different things.

Validated means "util was also present," not "util retained the cause." retained_via is every PRESENT family about the cause under the full inputs (cause_families), or topology if none. The three nested arms are never consulted. ಠ_ಠ

adHighCpu, cause ad, span error rate already PRESENT: M1 localizes ad. A named jvm.cpu.recent_utilization series is also ≥2×, so retained_via == ("sig", "util"). decide() returns validated. Drop util — the M2a arm — and ad is still retained. Util did not retain the truth. The milestone exists for the locally silent fault M1 cannot see. A co-occurring blip is not that result, and this verdict is what gets posted as-is after the one-shot.

The check that matches the protocol is the arm delta the module already computes: retained on M2a+M2b and not retained on M2a. A measured util is already a named-instance series; summarize_utilization will not emit PRESENT/ABSENT without one. Do not re-encode that invariant as the coverage rule.

Untestable ignores identity_coverage. That function is the protocol's prerequisite ("~0 → untestable, not failed") and decide() never reads it. Untestable is instead "no resource-fault cause has sig_state is not None". The PR description calls this the same condition made exact. It is not. ¯_(ツ)_/¯

~0 coverage implies nothing was measured. The converse is false. Identity coverage 0.8, ad and recommendation unmeasured (no util samples, or the wrong instrument), no healthy PRESENT util: the code says untestable. The protocol's only untestable escape is coverage ~0. This corpus was testable. Util did not retain the fault. That is not_validated.

otel-fresh was exactly 0/37282, so the non-fuzzy form of "~0" is rate is None or rate == 0. A pre-registered epsilon is also a threshold. A different predicate is a protocol change, and it opens an escape hatch the freeze was written to close. If that predicate is actually what you want, amend the #209 protocol before this encoder merges. Do not ship the substitution as an encoding.

Healthy-window PRESENT util correctly forces not_validated and correctly outranks the other branches. The generalization triple (0.50 / 0.33 / 0.58, undefined rate → untestable) matches the written bars. This finding is the M2b verdict only.

2. The ≥2× edge follow-up is not in the report

MAJOR.

edge_specificity's docstring says it is the healthy-window flagged-edge rate at the unchanged ≥2× cutoff. edges_present is sig_state == PRESENT, and summarize_edges sets that when the error rate is ≥5% or the latency ratio is ≥2×. An error-only edge (rate 0.10, ratio 1.0) increments healthy_windows_with_present_edge and never crossed 2×.

edges_present_latency_only then drops every PRESENT edge whose error branch also fired. An edge at 3× with a 10% error rate is absent from that count. The ≥2× window rate and the ≥2× edge rate are not recoverable from the JSON. The protocol asked for that rate, reported as-is, not recalibrated. Store latency-high on its own, including edges that also tripped the error cutoff.

3. The results post does not label unobservable cases

MAJOR.

The protocol says unobservable cases (no telemetry for the cause) are listed as such, because they are not model failures to tune away. cause_has_telemetry is computed and stored on each JSON case. render_markdown drops it. A miss and an unobservable case are the same row: ✗, retained-via —. The markdown is the post this driver writes for #209. Put the flag on the row.

Tests

test_m2b_validated_by_a_resource_fault_retained_via_util hands decide() a CaseEval whose families is already ("util",) and whose every arm shares one localization. test_m2b_untestable_when_no_resource_fault_is_measured asserts the substituted predicate. Neither can fail the two counterexamples above. Add those two cases. CI staying green is not evidence the protocol is implemented.

Question, not a blocker: summarize_arm takes the median candidate fraction only over positives with a non-empty localization. I cannot prove the otel-fresh "0.33" was computed that way. If that number included abstaining positives as fraction 0, the gate and the reference statistic are not the same object. Same question for <= 0.33 versus a rounded 1/3: statistics.median of thirds is 0.333…, which fails this bound, and the unit test only checks 4/10.

VERDICT

The defect is the decision rule, not the pipeline split. The evaluator's job is to make the frozen protocol impossible to reinterpret after the capture. decide() reinterprets it in advance: "retained via util" becomes "util was present," and "coverage ~0" becomes "the resource-fault cause was never measured." Both change the one-shot verdict, and the tests lock the rewritten rule in.

Do not merge. Make decide() the protocol — arm delta for "via util", identity_coverage for untestable — and report the ≥2× edge rate and the unobservable flag the post is supposed to contain.

Open in Web View Automation 

Sent by Cursor Automation: Code Reviewer

Comment thread src/eval/structural_m3.py
Comment thread src/eval/structural_m3.py
Comment thread src/eval/structural_m3.py Outdated
Comment thread tests/unit/test_structural_m3.py Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

REQUEST CHANGES

CI is green (run 36243829694: test 3.10, test 3.12, openapi, docker, migrations). This is not a failing-test problem. The tests exercise the predicates this module invented. They do not exercise the M3 protocol on #209.

load_structural_rows / resolve_structural is a real split. build_structural_view is still load, then structural_signals, then hypotheses, then resolve(part, observations=…), then rank_classes. The parity test compares the frozen arm's outcome and localization to that product result on the same persisted rows. The three arms call that resolve step on trimmed inputs, and candidate generation from the incident call graph stays in every arm. That part of the contract holds.

decide() does not. It is the function this PR exists to freeze, and it is a different function from the one pre-registered on #209.

1. decide() is not the M2b rule

BLOCKING.

#209, M3 protocol: M2b is untestable when series-identity coverage is ~0. It is validated when at least one of adHighCpu / recommendationCpuStress is retained via a util: observable on a named-instance series, and no healthy window has a PRESENT util:. Otherwise it is not validated.

identity_coverage() computes that rate. build_m3_report prints it. decide() never reads it. ಠ_ಠ

The branch it uses instead:

elif not any(c.util_measured_for_cause for c in resource):
    m2b = "untestable"
elif any("util" in c.retained_via for c in resource):
    m2b = "validated"

util_measured_for_cause means "some UtilSignal for the cause has sig_state is not None". The PR description says this is coverage ~0 "made exact." It is not. A measured util requires a named instance (summarize_utilization refuses otherwise), so coverage 0 implies the cause was not measured. The converse is false, and this branch implements the converse.

Failure 1. Every utilization sample names an instance (coverage 1.0), but ad's cpu instrument has two datapoint signatures, so the class stays UNKNOWN and util_measured_for_cause is false. decide() returns untestable. Coverage is not ~0, the resource fault was not retained via util:, and the protocol's verdict is not_validated. This is the substitution that turns a failed M2b into an untestable one. That is the measurement being shaped before the capture exists.

Failure 2. One named sample out of 10 000, and that sample is a PRESENT util:ad:cpu that keeps ad in the localization. Coverage is 0.0001. The protocol says ~0 is untestable, not a pass. decide() returns validated. The non-fuzzy reading of "~0" is named == 0. That value is already computed. It is not what this branch tests.

Failure 3, same verdict. retained_via is which PRESENT families co-occur with a cause the full arm kept, not which arm added the cause. _case() copies one ArmResult onto every arm, and test_m2b_validated_by_a_resource_fault_retained_via_util expects validated for a cause M1 already localizes, as long as families=("util",). Between M2a and the full arm the only added input is util_signals, expectations are soft, and candidates only grow. not retained("M2a") and retained("M2a+M2b") is the counterfactual "retained via util". "util" in retained_via is not. You built the nested arms and then the verdict does not consult them.

Same function, one layer short: generalizes is assigned before protocol_deviations and does not read them. One abstaining healthy window is a 100% abstention rate, clears 0.58, and the headline is generalizes next to 1 healthy negatives < 12. The bar is 7/12. It is not a rate you can satisfy with n=1. n=0 is already untestable. n<12 is the same corpus failure.

Fix decide():

  • untestable when named utilization samples are 0, or a resource-fault scenario is absent;
  • not_validated when any healthy window has a PRESENT util:;
  • validated only when some resource-fault case is retained by the full arm and not by M2a;
  • otherwise not_validated;
  • do not emit generalizes when there are fewer than 12 healthy windows.

Add the two cases the suite cannot currently fail: coverage 1.0 with the resource-fault util UNKNOWN, and a cause retained by M2a that also has a PRESENT util.

2. The ≥2× flagged-edge rate is a different set

MAJOR.

summarize_edges marks an edge PRESENT when the error branch clears 5% or the latency ratio is ≥2×. latency_only keeps PRESENT edges whose error branch is not PRESENT, i.e. latency_high and not error_present.

edge_specificity then reports windows with any PRESENT edge, PRESENT/measured, and latency-only/PRESENT. The first two include error-only flags the 2× rule did not raise. The third drops edges that are both error-flagged and ≥2×. The pre-registered follow-up is the healthy-window rate at the unchanged ≥2× cutoff: latency_measured and discretize_ratio(ratio) == HIGH. That count is none of the three numbers, and it cannot be reconstructed from them. Report |latency_high| / |measured| and the window rate of those edges. Keep an error-branch rate beside it if you want one. Do not label the blend as the cutoff.

3. "Learned ranker" is a label, not a checked load

MAJOR.

load_ranker returns None for a missing or unparseable file, and _rank_candidates then uses the legacy log-cluster pool. The driver sets RCA_RANKER_MODEL_PATH, never checks that a model loaded, and the markdown always reads Default path (learned ranker + rare_event). Provenance stores the path string. The committed models/rca_ranker.json makes the default invocation work. A bad --ranker, a corrupt artifact, or a tree without that file still writes the one-shot #209 post, and the numbers are the volume selector. The protocol's default path is the ranker check. Exit non-zero if the ranker or the calibrator does not load, and record that they did.

VERDICT

The problem is the decision function, not the arm wiring. The product split preserves behavior and the parity test actually checks outcome and localization. decide() answers a nearby question: "was the resource-fault cause ever measured, and did a util observable co-occur with a retained cause?" #209 asked "is series-identity coverage ~0, and did a util: observable on a named-instance series retain a resource fault the trace arms did not?" Those disagree on the published M2b verdict, which is the reason this evaluator is frozen before any capture exists. Do not merge until decide() matches that comment and a test fails when coverage and the arm diff are ignored.

Open in Web View Automation 

Sent by Cursor Automation: Code Reviewer

Comment thread src/eval/structural_m3.py
Comment thread src/eval/structural_m3.py
Comment thread src/eval/structural_m3.py Outdated
Comment thread scripts/eval/m3_structural.py Outdated
Comment thread tests/unit/test_structural_m3.py Outdated
…review)

- M2b "retained via util" is now the arm delta: the frozen arm retains
  the cause, M2a (the same inputs without util) does not, and util on
  the cause is PRESENT. A co-present util on a fault that sig already
  retains no longer validates.
- M2b "untestable" now reads series-identity coverage, as the protocol
  says: exactly no named utilization sample (otel-fresh was 0/37282).
  High coverage with unmeasured resource causes is not_validated.
- The >=2x edge rate is counted on its own, including edges whose error
  branch also fired. The error branch and the combined PRESENT are
  reported separately.
- The results post labels unobservable cases (no telemetry from the
  cause).
- The gated statistics are now the otel-fresh reference definitions:
  - the universe is span OR metric services;
  - generated means any hypothesis;
  - selectivity is over generated positives, with an empty set counting
    as 0;
  - abstention means no hypothesis.
  "No localization claim" is reported as a secondary figure. Rerun on
  the spent otel-fresh corpus, the evaluator reproduces every reference
  number exactly.
- The candidate-fraction bound stays the frozen literal 0.33, pinned by
  a test that exact 1/3 (otel-fresh M1's value) fails it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@leo-aa88

Copy link
Copy Markdown
Member Author

/review

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

REQUEST CHANGES

decide() now implements the M2b rule frozen on #209. Healthy PRESENT util: is not_validated. Coverage in (None, 0.0) is untestable. validated requires the arm delta: the full arm retains the cause, M2a does not, and a util: on the cause is PRESENT. The ≥2× edge rate is counted on its own, including edges whose error branch also fired. Unobservable cases are labeled on the row. Those were the blockers. CI on 3af69ad is green (test 3.10, test 3.12, docker, migrations, openapi). This is not a failing-test problem.

The one-shot post can still title the volume selector as the learned ranker.

1. Default path claims a ranker that may not have loaded

MAJOR.

The driver sets RCA_RANKER_MODEL_PATH and RCA_CALIBRATOR_MODEL_PATH, runs run_cases, and always appends Default path (learned ranker + rare_event). Provenance stores the path strings.

The protocol's default path is raglogs eval with the learned ranker and the rare_event trigger. load_ranker returns None when the path is missing or unparseable. _rank_candidates then returns empty, and explain_window scores the legacy log-cluster pool. models/rca_ranker.json makes the default invocation work. A bad --ranker, a corrupt artifact, or a checkout without that file still exits 0 and writes the #209 post. Those numbers are not the ranker check, and this evaluator runs once.

load_calibrator fails open the same way. The headline accuracies do not use it. The JSON report's calibration block does, and it becomes the ordinal fallback while provenance still names the calibrator path.

Resolve both paths, load them, and return 2 if either load returns None. Record ranker_loaded and calibrator_loaded in provenance. Do not print the learned-ranker heading unless the ranker load succeeded.

2. The abstention constant's comment describes the statistic that is no longer gated

The gate is outcome != "no_candidates". The comment on HEALTHY_ABSTENTION_MIN still says "no localization claim". That is the secondary healthy_no_claim figure. A healthy window that generated hypotheses and then found none compatible abstains under the comment and does not under the code. The module docstring matches the runtime. The constant comment does not. Not a merge blocker by itself. It is the comment a later edit will trust.

VERDICT

The problem left is the driver, not the decision function. decide() matches the frozen M2b rule, the ≥2× edge count is the cutoff the protocol asked for, and the product split preserves outcome and localization. The default-path heading describes a ranker load the runtime does not perform. On a measurement that runs once, fail-open is a false result. Do not merge until a failed load refuses the run.

Open in Web View Automation 

Sent by Cursor Automation: Code Reviewer

Comment thread scripts/eval/m3_structural.py Outdated
Comment thread src/eval/structural_m3.py Outdated
…view r2)

The product's load_ranker/load_calibrator fail open: None falls back to
the volume selector or to ordinal confidence. The driver still labelled
the output "learned ranker", which would have been a false result on a
run that happens once.

- It now resolves both paths and loads them with the product loaders
  before any DB work. It refuses (exit 2) if either load fails.
- It records absolute paths, sha256s and *_loaded in provenance.
- It refuses if settings did not take rare_event or the ranker path.
- The heading names the ranker only when it loaded.

Also corrects the constant comments: the abstention gate is "no
hypothesis generated", not the ungated healthy_no_claim, and selectivity
is taken over generated positives.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@leo-aa88

Copy link
Copy Markdown
Member Author

/review

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Open in Web View Automation 

Sent by Cursor Automation: Code Reviewer

@leo-aa88
leo-aa88 merged commit 1ec4208 into main Sep 26, 2026
7 checks passed
@leo-aa88
leo-aa88 deleted the feat/m3-structural-eval-209 branch September 26, 2026 19:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant