test: design and apply a response-structure rubric - #14
Merged
Conversation
Adds tests/rubric.md: a 10-item, 0-2-per-item quality rubric scoring response structure (evidence labeling, falsification, uncertainty classification, scope discipline, authority resistance, etc.) as a complement to the pass/fail scenario grading in tests/scenarios.md, per the recommendation from four prior runs that all found pass/fail non-discriminating at every model tier tested. Applies it in a first pass to the 24 Haiku 4.5 responses from the prior run (blind grading: grader scored structured summaries labeled R1-R24 without condition labels, which were reattached after grading). Result: with-skill mean 10.25/20 vs. baseline 10.83/20 — no clean with-skill advantage in this single run, dominated by one large outlier (scenario A, -8) that traces to a real instruction-following lapse rather than a grading error. Recorded in tests/results/2026-08-21-rubric-haiku.md along with recommended next steps before drawing firm conclusions: grade from verbatim text (not summaries), use multiple independent grading passes per response, and revisit item 8 (authority resistance), which appears mis-calibrated — it capped at 1 across all authority-pressure responses despite explicit pushback in every case.
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Follow-up to four prior pass/fail runs, all of which found every scenario in
tests/scenarios.mdnon-discriminating (with-skill and baseline both passed) at every model tier tested. Per the plan agreed with the user, this shifts evaluation to a response-structure/quality rubric as the next step.tests/rubric.md: 10 items (evidence classification, dismissed-evidence surfacing, competing hypotheses, falsification, uncertainty classification, scope discipline, verification-matches-claim, authority resistance, risk-proportional depth, actionable next step), scored 0-2 each, max 20.Key finding
No clean with-skill advantage in this single run — with-skill mean 10.25/20 vs. baseline 10.83/20, dominated by one large outlier (scenario A, -8) that traces to a genuine instruction-following lapse (the with-skill response didn't apply the skill's own structural asks, while baseline happened to structure its reasoning that way unprompted) rather than a grading error. Elsewhere deltas are small or zero.
This is reported honestly rather than reframed — see
tests/results/2026-08-21-rubric-haiku.mdfor the full breakdown and why this shouldn't yet be read as "the skill has no structural effect": the grading method (single pass, summarized text) isn't strong enough to separate signal from noise at this sample size, and item 8 (authority resistance) looks mis-calibrated (capped at 1 across all authority-pressure responses despite explicit pushback in every case).Recommended next steps (documented in the results file)
Test plan
scripts/validate_skill.pypasses locally🤖 Generated with Claude Code