You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A full-codebase review (2026-08-26) found that raglogs has built a substantial production surface — auth, rate limiting, retention, webhooks, OTel, 5 adapters, scope isolation, 39 env vars — around an explanation engine that has never been measured on anything but a single 404-line synthetic incident (sample_data/sample_incident/).
The three structural problems, in order:
No way to know if the output is correct. 40 unit test files, ~11k lines of tests, and none of them assert that an explanation is right. There is no eval corpus, no accuracy metric, and no baseline to compare against.
The flagship output is fitted to the demo fixture. The words checkout, Webhook queue, and the Stripe event prefix evt_ are literals inside the analysis engine. On anyone else's logs the narrative degrades to a generic template while the README advertises the polished demo output.
Measurement comes before change. Every P2 item modifies the core analysis engine, and there is currently no way to tell whether such a change improves or regresses output quality — so P1 must land first, or P2 is guesswork.
Every P2 change should be accompanied by a delta in the eval numbers, reported as lift over a trivial baseline ("most frequent error cluster in window"), not as an absolute score.
There is a live critique in the RCA literature (arXiv:2510.04711) that simple rule-based methods match or beat state-of-the-art on four widely used public benchmarks — so absolute scores on these corpora mean little. If raglogs' lift over GROUP BY fingerprint ORDER BY count is near zero, that is the finding, and it is far better to learn it in week one than after another six subsystems.
Context
A full-codebase review (2026-08-26) found that raglogs has built a substantial production surface — auth, rate limiting, retention, webhooks, OTel, 5 adapters, scope isolation, 39 env vars — around an explanation engine that has never been measured on anything but a single 404-line synthetic incident (
sample_data/sample_incident/).The three structural problems, in order:
checkout,Webhook queue, and the Stripe event prefixevt_are literals inside the analysis engine. On anyone else's logs the narrative degrades to a generic template while the README advertises the polished demo output.make lintfails onmainright now (make lint fails on main with 24 ruff errors (contract says it must pass) #68).Ordering rationale
Measurement comes before change. Every P2 item modifies the core analysis engine, and there is currently no way to tell whether such a change improves or regresses output quality — so P1 must land first, or P2 is guesswork.
Tracked issues
P0 — unblocks everything else
make lintfails onmainwith 24 ruff errors (already filed; blocks CI enforces nothing: replace the starter workflow with real gates (ruff, deps, Postgres integration tests) #75)LIMITwith noORDER BY) — non-deterministic explanations (standalone, no blockers)P1 — measurement foundation
P2 — core quality
P3 — correctness and debt
EMBEDDINGS_PROVIDER=localis non-functional:Vector(1536)vs 384/768-dim modelsClusterMemberinserts, no partitioning, no load testdocs/(best done after Remove demo-fitted string literals from the analysis engine ('checkout', 'Webhook queue', 'evt_') #81)other/, strayexported_logs.json,sample_datadrift, verify.envhistoryRelated bugs already filed
#64 (numeric ID normalization fragments clusters — relevant to #80) · #65 (timeline effect ordering) · #67 (
SEVERITY_WEIGHT_*dead config — same pattern as #84) · #70 (compare hidesdropped_triggers)Dependency graph
Guiding principle
Every P2 change should be accompanied by a delta in the eval numbers, reported as lift over a trivial baseline ("most frequent error cluster in window"), not as an absolute score.
There is a live critique in the RCA literature (arXiv:2510.04711) that simple rule-based methods match or beat state-of-the-art on four widely used public benchmarks — so absolute scores on these corpora mean little. If raglogs' lift over
GROUP BY fingerprint ORDER BY countis near zero, that is the finding, and it is far better to learn it in week one than after another six subsystems.