Skip to content

Epic: Explanation quality — measure it before extending it #74

Description

@leo-aa88

Context

A full-codebase review (2026-08-26) found that raglogs has built a substantial production surface — auth, rate limiting, retention, webhooks, OTel, 5 adapters, scope isolation, 39 env vars — around an explanation engine that has never been measured on anything but a single 404-line synthetic incident (sample_data/sample_incident/).

The three structural problems, in order:

  1. No way to know if the output is correct. 40 unit test files, ~11k lines of tests, and none of them assert that an explanation is right. There is no eval corpus, no accuracy metric, and no baseline to compare against.
  2. The flagship output is fitted to the demo fixture. The words checkout, Webhook queue, and the Stripe event prefix evt_ are literals inside the analysis engine. On anyone else's logs the narrative degrades to a generic template while the README advertises the polished demo output.
  3. CI enforces none of the project's own rules. The workflow is the unmodified GitHub Python starter template. make lint fails on main right now (make lint fails on main with 24 ruff errors (contract says it must pass) #68).

Ordering rationale

Measurement comes before change. Every P2 item modifies the core analysis engine, and there is currently no way to tell whether such a change improves or regresses output quality — so P1 must land first, or P2 is guesswork.

Tracked issues

P0 — unblocks everything else

P1 — measurement foundation

P2 — core quality

P3 — correctness and debt

Related bugs already filed

#64 (numeric ID normalization fragments clusters — relevant to #80) · #65 (timeline effect ordering) · #67 (SEVERITY_WEIGHT_* dead config — same pattern as #84) · #70 (compare hides dropped_triggers)

Dependency graph

#68 ──> #75 ──> #77 ──┬──> #78 ──┬──> #81
                      │          ├──> #82 ──> #83
                      ├──> #79 ──┘
                      └──> #80

#76, #84, #85, #87   (independent, any time)
#86                  (after #81)
#88                  (policy decision, any time)

Guiding principle

Every P2 change should be accompanied by a delta in the eval numbers, reported as lift over a trivial baseline ("most frequent error cluster in window"), not as an absolute score.

There is a live critique in the RCA literature (arXiv:2510.04711) that simple rule-based methods match or beat state-of-the-art on four widely used public benchmarks — so absolute scores on these corpora mean little. If raglogs' lift over GROUP BY fingerprint ORDER BY count is near zero, that is the finding, and it is far better to learn it in week one than after another six subsystems.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationepicTracking issue spanning multiple featuresevalEvaluation harness, corpora, and quality measurementqualityExplanation quality / core analysis engine

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions