Eval harness: ground truth, 4-way ladder, judge grid, agent metrics, CI gates - #28
Merged
Merged
Conversation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…PU-pinned models Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…trics, CI gates Ground truth: 200 questions pinned to the snapshot (140 retrieval + 25 synthesis LLM-written and hand-checked by Jordan; 25 analytical + 10 freshness templated with exact expected answers), every record labeled with its expected tool. Results (committed + report.md): hybrid+rerank ships at 0.893 hit-rate@8 / 0.749 MRR (+17pts hit-rate over dense-only); routing accuracy 0.995; tool-arg match 0.743; execution accuracy 0.80; judge grid: citation_strict/sonnet best (3.13), citation-strict prompt +0.39 over baseline on haiku. CI: free smoke gate on every push (execution 35/35, hybrid 1.0 vs 0.55 threshold) + on-demand full-eval workflow with qdrant service container. Local models pinned to CPU (MPS segfaults in executor threads). Closes #9 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…vindicated The store treats submitted_to as inclusive-through-day; month templates passed next-month-01, counting an extra day and penalizing the agent's correct Jan-31 phrasing. Corrected templates + regenerated templated sets: tool-arg match 0.743 -> 1.0, execution accuracy 0.80 -> 1.0 (routing unchanged 0.995). Also: report marks the grid winner (best) with the SPEC tiering rationale instead of falsely claiming it ships; smoke gate tightened to 0.8 + fusion-sanity check implemented (was dead); tool-arg metric honestly labeled subset-match; generator can no longer truncate committed sets on partial regeneration and preserves the hand-check record; eval.yml env cleanup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Parent
Closes #9 — SPEC §6
What this delivers
Ground truth (200 questions pinned to the snapshot, committed): 140 retrieval + 25 synthesis LLM-written and hand-checked by Jordan (recorded in the manifest); 25 analytical + 10 freshness templated with exact expected answers computed from the store. Every record carries its expected tool.
Results (committed, report):
CI: free smoke gate on every push (execution 35/35, hybrid 1.0 vs 0.55 threshold, in-memory qdrant + real local embedders, no keys) + an on-demand
full-evalworkflow (qdrant service container, snapshot-built index, needs ANTHROPIC_API_KEY secret, ~$5/run,skip_llm_gridtoggle).Found along the way: sentence-transformers auto-picks MPS on Apple silicon and segfaults when LangGraph runs tools in executor threads — all local models now pinned to CPU (deploy targets are CPU-only anyway).
Known rough edges (documented, not hidden)
🤖 Generated with Claude Code