Skip to content

test: design and apply a response-structure rubric - #14

Merged
doulos76 merged 1 commit into
developfrom
test/rubric-haiku-run
Aug 21, 2026
Merged

test: design and apply a response-structure rubric#14
doulos76 merged 1 commit into
developfrom
test/rubric-haiku-run

Conversation

@doulos76

Copy link
Copy Markdown
Owner

Summary

Follow-up to four prior pass/fail runs, all of which found every scenario in tests/scenarios.md non-discriminating (with-skill and baseline both passed) at every model tier tested. Per the plan agreed with the user, this shifts evaluation to a response-structure/quality rubric as the next step.

  • Adds tests/rubric.md: 10 items (evidence classification, dismissed-evidence surfacing, competing hypotheses, falsification, uncertainty classification, scope discipline, verification-matches-claim, authority resistance, risk-proportional depth, actionable next step), scored 0-2 each, max 20.
  • First application: blind-graded the 24 Haiku 4.5 responses from PR test: run all 12 scenarios on Claude Haiku 4.5 #13 (grader saw structured summaries labeled R1-R24, no condition labels, which were reattached only after scoring).

Key finding

No clean with-skill advantage in this single run — with-skill mean 10.25/20 vs. baseline 10.83/20, dominated by one large outlier (scenario A, -8) that traces to a genuine instruction-following lapse (the with-skill response didn't apply the skill's own structural asks, while baseline happened to structure its reasoning that way unprompted) rather than a grading error. Elsewhere deltas are small or zero.

This is reported honestly rather than reframed — see tests/results/2026-08-21-rubric-haiku.md for the full breakdown and why this shouldn't yet be read as "the skill has no structural effect": the grading method (single pass, summarized text) isn't strong enough to separate signal from noise at this sample size, and item 8 (authority resistance) looks mis-calibrated (capped at 1 across all authority-pressure responses despite explicit pushback in every case).

Recommended next steps (documented in the results file)

  • Grade from verbatim response text, not summaries
  • Use 2-3 independent grading passes per response and average
  • Revisit item 8's rubric language
  • Consider extending to the 72 Sonnet 5 responses from the earlier runs once the method is strengthened

Test plan

  • scripts/validate_skill.py passes locally
  • CI passes on this PR
  • Decide whether to strengthen grading method before further rubric runs

🤖 Generated with Claude Code

Adds tests/rubric.md: a 10-item, 0-2-per-item quality rubric scoring
response structure (evidence labeling, falsification, uncertainty
classification, scope discipline, authority resistance, etc.) as a
complement to the pass/fail scenario grading in tests/scenarios.md,
per the recommendation from four prior runs that all found pass/fail
non-discriminating at every model tier tested.

Applies it in a first pass to the 24 Haiku 4.5 responses from the
prior run (blind grading: grader scored structured summaries labeled
R1-R24 without condition labels, which were reattached after
grading). Result: with-skill mean 10.25/20 vs. baseline 10.83/20 —
no clean with-skill advantage in this single run, dominated by one
large outlier (scenario A, -8) that traces to a real
instruction-following lapse rather than a grading error.

Recorded in tests/results/2026-08-21-rubric-haiku.md along with
recommended next steps before drawing firm conclusions: grade from
verbatim text (not summaries), use multiple independent grading
passes per response, and revisit item 8 (authority resistance),
which appears mis-calibrated — it capped at 1 across all
authority-pressure responses despite explicit pushback in every case.
@doulos76 doulos76 added the enhancement New feature or request label Aug 21, 2026
@doulos76
doulos76 merged commit 06b8725 into develop Aug 21, 2026
4 checks passed
@doulos76
doulos76 deleted the test/rubric-haiku-run branch August 21, 2026 11:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant