Skip to content

Eval harness: ground truth, 4-way ladder, judge grid, agent metrics, CI gates - #28

Merged
mr-j90 merged 4 commits into
masterfrom
feat/9-eval-harness
Aug 9, 2026
Merged

mr-j90 merged 4 commits into
masterfrom
feat/9-eval-harness

Conversation

@mr-j90

@mr-j90 mr-j90 commented Aug 9, 2026

Copy link
Copy Markdown
Owner

Parent

Closes #9 — SPEC §6

What this delivers

Ground truth (200 questions pinned to the snapshot, committed): 140 retrieval + 25 synthesis LLM-written and hand-checked by Jordan (recorded in the manifest); 25 analytical + 10 freshness templated with exact expected answers computed from the store. Every record carries its expected tool.

Results (committed, report):

  • Retrieval ladder: hybrid+rerank ships at 0.893 hit-rate@8 / 0.749 MRR — +17pts hit-rate over dense-only; sparse BM25 alone is strong (0.864), which is honest and interesting.
  • Agent: 0.995 routing accuracy (n=200); tool-arg exact match 0.743; execution accuracy 0.80 (strict exact-match — equivalent-but-differently-phrased date windows count as misses).
  • Judge grid: citation_strict/sonnet best (3.13/5); the citation-strict prompt is a real effect (+0.39 over baseline on haiku); validates Haiku-dev/Sonnet-demo tiering. Harsh judge, middling absolute scores — documented, judgments committed for spot-checks.

CI: free smoke gate on every push (execution 35/35, hybrid 1.0 vs 0.55 threshold, in-memory qdrant + real local embedders, no keys) + an on-demand full-eval workflow (qdrant service container, snapshot-built index, needs ANTHROPIC_API_KEY secret, ~$5/run, skip_llm_grid toggle).

Found along the way: sentence-transformers auto-picks MPS on Apple silicon and segfaults when LangGraph runs tools in executor threads — all local models now pinned to CPU (deploy targets are CPU-only anyway).

Known rough edges (documented, not hidden)

  • Runners write results only at completion — no mid-run checkpointing; a killed run forfeits progress.
  • Judge shares a model family with graded answers (caveat in the report).
  • best_mode tie-break: hit-rate primary (all top-k reaches the agent), MRR secondary.

🤖 Generated with Claude Code

mr-j90 and others added 4 commits August 9, 2026 11:52
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…PU-pinned models

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…trics, CI gates

Ground truth: 200 questions pinned to the snapshot (140 retrieval + 25
synthesis LLM-written and hand-checked by Jordan; 25 analytical + 10
freshness templated with exact expected answers), every record labeled
with its expected tool. Results (committed + report.md): hybrid+rerank
ships at 0.893 hit-rate@8 / 0.749 MRR (+17pts hit-rate over dense-only);
routing accuracy 0.995; tool-arg match 0.743; execution accuracy 0.80;
judge grid: citation_strict/sonnet best (3.13), citation-strict prompt
+0.39 over baseline on haiku. CI: free smoke gate on every push
(execution 35/35, hybrid 1.0 vs 0.55 threshold) + on-demand full-eval
workflow with qdrant service container. Local models pinned to CPU
(MPS segfaults in executor threads).

Closes #9

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…vindicated

The store treats submitted_to as inclusive-through-day; month templates
passed next-month-01, counting an extra day and penalizing the agent's
correct Jan-31 phrasing. Corrected templates + regenerated templated
sets: tool-arg match 0.743 -> 1.0, execution accuracy 0.80 -> 1.0
(routing unchanged 0.995). Also: report marks the grid winner (best)
with the SPEC tiering rationale instead of falsely claiming it ships;
smoke gate tightened to 0.8 + fusion-sanity check implemented (was
dead); tool-arg metric honestly labeled subset-match; generator can no
longer truncate committed sets on partial regeneration and preserves
the hand-check record; eval.yml env cleanup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@mr-j90
mr-j90 merged commit d1cb447 into master Aug 9, 2026
2 checks passed
@mr-j90
mr-j90 deleted the feat/9-eval-harness branch August 9, 2026 20:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval harness: ground truth, retrieval ladder, judge grid, CI smoke

1 participant