You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The #177 epic (Causal diagnostic inference under partial observability) is implemented: Observable → Hypothesis → ObservationModel → structural partition → outcome → class scoring → presentation are all merged, and the structural result is exposed opt-in in the CLI/API (--structural, PR #207). Its empirical result is clear, and it moved the bottleneck:
On trace-loc the structural bridge retains the truth extremely well — ~100% candidate recall, ~2.4 candidates/case — but 0% unique identification.
On the real OTel corpus the structural adapter often cannot generate candidates at all.
So the limiting factor is no longer partitioning, scoring, or another heuristic. It is grounding trustworthy observables from real telemetry. This epic makes real OTel evidence legible to the structural engine.
This is inside the existing freeze carve-out (AGENTS.md: "metrics/traces ingestion for causal localization" is unfrozen). It is analysis-core grounding, not net-new platform surface.
The goal (the strategic shift)
Optimize for "turn a huge telemetry mess into a compact, trustworthy diagnostic packet without throwing away the true cause" — not "improve top-1." "Truth retained among 2–3 evidence-backed hypotheses" is already a useful result. Headline metrics are candidate recall + selectivity + outcome distribution, never top-1.
What the data actually shows (grounding the design)
Observables are metric-only today — the structural adapter derives sig:{service} from error_rate / latency_msmetrics, so a service with spans but no metrics is UNKNOWN → "can't generate candidates."
Spans are latency-rich, error-sparse — an incident case had 52k spans, 96% with duration_ms, only 24 with error-status. Latency is a strong per-span signal we largely ignore for observables.
We drop the richest evidence at ingest — raw OTLP carries span kind / peer.service / per-edge status; our ingestion keeps only service/operation/duration/status and leaves TraceSpan.attributes null. That per-edge evidence is exactly what Phase F needed for sound absence attribution.
Milestones
M1 (first, deliberately narrow): span-derived observables — get a selective candidate set on real OTel without inventing evidence
Derive per-service observables from real OTLP spans — latency-first, plus sparse error-status where actually present — without converting missing telemetry into ABSENT (preserve UNKNOWN). Run the existing structural bridge against the already-spent OTel dev corpus and report candidate recall + selectivity + outcome distribution. No ranking changes, no new causal heuristics, no default promotion.
Success criteria (frozen before looking):
truth retained ≥ 90%
candidate set materially smaller than service enumeration
UNKNOWN discipline preserved (missing measurement never becomes ABSENT)
invented evidence: 0
M2 (later): preserve OTLP semantic attributes through ingestion
Carry span kind / peer.service / per-edge status into TraceSpan.attributes so per-edge causal evidence exists — the signal Phase F's sound absence attribution and #82's complete linkage both need. Measured against the same frozen gate.
M3 (later): re-run the structural experiment on a fresh frozen real OTel corpus
Freeze a new real corpus before looking. Report candidate recall, median candidate-set size / fraction of services, healthy abstention, outcome distribution, causal-kind coverage. Mostly-UNCERTAIN is acceptable; the question is whether the packet narrowed the search space without deleting the truth.
Explicitly NOT in scope
No Bayesian layer, no new reranker, no internal agent loop, no MCP server, no more synthetic fault families. The evidence says the highest-value missing component is the boring, load-bearing one: make real telemetry legible to the causal model.
Why this epic exists
The #177 epic (Causal diagnostic inference under partial observability) is implemented: Observable → Hypothesis → ObservationModel → structural partition → outcome → class scoring → presentation are all merged, and the structural result is exposed opt-in in the CLI/API (
--structural, PR #207). Its empirical result is clear, and it moved the bottleneck:trace-locthe structural bridge retains the truth extremely well — ~100% candidate recall, ~2.4 candidates/case — but 0% unique identification.So the limiting factor is no longer partitioning, scoring, or another heuristic. It is grounding trustworthy observables from real telemetry. This epic makes real OTel evidence legible to the structural engine.
This is inside the existing freeze carve-out (AGENTS.md: "metrics/traces ingestion for causal localization" is unfrozen). It is analysis-core grounding, not net-new platform surface.
The goal (the strategic shift)
Optimize for "turn a huge telemetry mess into a compact, trustworthy diagnostic packet without throwing away the true cause" — not "improve top-1." "Truth retained among 2–3 evidence-backed hypotheses" is already a useful result. Headline metrics are candidate recall + selectivity + outcome distribution, never top-1.
What the data actually shows (grounding the design)
Checked the real otel-fresh spans directly:
parent_span_id;call_edges/build_service_graphderive the graph from it. The graph is incomplete (not every path is traced — this is why Redesign trigger detection: rare-event correlation and service linkage, not 12 regexes + a 30-minute window #82's linkage dropped ~44% of positives), not absent.sig:{service}fromerror_rate/latency_msmetrics, so a service with spans but no metrics isUNKNOWN→ "can't generate candidates."duration_ms, only 24 with error-status. Latency is a strong per-span signal we largely ignore for observables.peer.service/ per-edge status; our ingestion keeps onlyservice/operation/duration/statusand leavesTraceSpan.attributesnull. That per-edge evidence is exactly what Phase F needed for sound absence attribution.Milestones
M1 (first, deliberately narrow): span-derived observables — get a selective candidate set on real OTel without inventing evidence
Success criteria (frozen before looking):
UNKNOWNdiscipline preserved (missing measurement never becomesABSENT)M2 (later): preserve OTLP semantic attributes through ingestion
Carry span kind /
peer.service/ per-edge status intoTraceSpan.attributesso per-edge causal evidence exists — the signal Phase F's sound absence attribution and #82's complete linkage both need. Measured against the same frozen gate.M3 (later): re-run the structural experiment on a fresh frozen real OTel corpus
Freeze a new real corpus before looking. Report candidate recall, median candidate-set size / fraction of services, healthy abstention, outcome distribution, causal-kind coverage. Mostly-
UNCERTAINis acceptable; the question is whether the packet narrowed the search space without deleting the truth.Explicitly NOT in scope
No Bayesian layer, no new reranker, no internal agent loop, no MCP server, no more synthetic fault families. The evidence says the highest-value missing component is the boring, load-bearing one: make real telemetry legible to the causal model.
Relationship to other issues
--structuralfrom opt-in to default, and Redesign trigger detection: rare-event correlation and service linkage, not 12 regexes + a 30-minute window #82's regex-detector replacement — all were deferred pending exactly this real per-edge grounding.