Skip to content

EPIC: Real-telemetry causal observable grounding — make OTel evidence usable by the structural engine #209

Description

@leo-aa88

Why this epic exists

The #177 epic (Causal diagnostic inference under partial observability) is implemented: Observable → Hypothesis → ObservationModel → structural partition → outcome → class scoring → presentation are all merged, and the structural result is exposed opt-in in the CLI/API (--structural, PR #207). Its empirical result is clear, and it moved the bottleneck:

  • On trace-loc the structural bridge retains the truth extremely well — ~100% candidate recall, ~2.4 candidates/case — but 0% unique identification.
  • On the real OTel corpus the structural adapter often cannot generate candidates at all.

So the limiting factor is no longer partitioning, scoring, or another heuristic. It is grounding trustworthy observables from real telemetry. This epic makes real OTel evidence legible to the structural engine.

This is inside the existing freeze carve-out (AGENTS.md: "metrics/traces ingestion for causal localization" is unfrozen). It is analysis-core grounding, not net-new platform surface.

The goal (the strategic shift)

Optimize for "turn a huge telemetry mess into a compact, trustworthy diagnostic packet without throwing away the true cause" — not "improve top-1." "Truth retained among 2–3 evidence-backed hypotheses" is already a useful result. Headline metrics are candidate recall + selectivity + outcome distribution, never top-1.

What the data actually shows (grounding the design)

Checked the real otel-fresh spans directly:

  • caller→callee is already available — spans carry parent_span_id; call_edges / build_service_graph derive the graph from it. The graph is incomplete (not every path is traced — this is why Redesign trigger detection: rare-event correlation and service linkage, not 12 regexes + a 30-minute window #82's linkage dropped ~44% of positives), not absent.
  • Observables are metric-only today — the structural adapter derives sig:{service} from error_rate / latency_ms metrics, so a service with spans but no metrics is UNKNOWN → "can't generate candidates."
  • Spans are latency-rich, error-sparse — an incident case had 52k spans, 96% with duration_ms, only 24 with error-status. Latency is a strong per-span signal we largely ignore for observables.
  • We drop the richest evidence at ingest — raw OTLP carries span kind / peer.service / per-edge status; our ingestion keeps only service/operation/duration/status and leaves TraceSpan.attributes null. That per-edge evidence is exactly what Phase F needed for sound absence attribution.

Milestones

M1 (first, deliberately narrow): span-derived observables — get a selective candidate set on real OTel without inventing evidence

Derive per-service observables from real OTLP spans — latency-first, plus sparse error-status where actually present — without converting missing telemetry into ABSENT (preserve UNKNOWN). Run the existing structural bridge against the already-spent OTel dev corpus and report candidate recall + selectivity + outcome distribution. No ranking changes, no new causal heuristics, no default promotion.

Success criteria (frozen before looking):

  • truth retained ≥ 90%
  • candidate set materially smaller than service enumeration
  • UNKNOWN discipline preserved (missing measurement never becomes ABSENT)
  • invented evidence: 0

M2 (later): preserve OTLP semantic attributes through ingestion

Carry span kind / peer.service / per-edge status into TraceSpan.attributes so per-edge causal evidence exists — the signal Phase F's sound absence attribution and #82's complete linkage both need. Measured against the same frozen gate.

M3 (later): re-run the structural experiment on a fresh frozen real OTel corpus

Freeze a new real corpus before looking. Report candidate recall, median candidate-set size / fraction of services, healthy abstention, outcome distribution, causal-kind coverage. Mostly-UNCERTAIN is acceptable; the question is whether the packet narrowed the search space without deleting the truth.

Explicitly NOT in scope

No Bayesian layer, no new reranker, no internal agent loop, no MCP server, no more synthetic fault families. The evidence says the highest-value missing component is the boring, load-bearing one: make real telemetry legible to the causal model.

Relationship to other issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions