Skip to content

Repository files navigation

llm-response-eval-kit

Pairwise evaluation of model-written code: one coding prompt, two responses, each rated 1–5 on five dimensions, then a single 7-point preference label.

The part most preference-data pipelines are missing is the third step — checking that the preference a grader chose is supported by the ratings they themselves wrote down.

$ python3 -m evalkit compare c06

c06-p07 ------------------------------------------------------------------

probe: p07-degenerate-arguments

DIMENSION                             A      B    DELTA
Instruction following              5.00   3.00    -2.00
Code correctness                   5.00   2.50    -2.50
Completeness                       4.50   1.50    -3.00
Code quality                       4.00   3.00    -1.00
Conciseness                        4.00   5.00    +1.00

grader_1   A 4.60  B 2.80  gap -1.80  ->  A much better      OK
grader_2   A 4.40  B 3.20  gap -1.20  ->  About the same     FLAG
           ratings imply band -2, grader stated +0 (2 bands apart, tolerance 1)

preference: graders disagree (3 bands apart)

grader_2 rated A higher on four of five dimensions — including completeness 4 versus 2 — and then recorded "About the same". The ratings say A better; the label says tie.

The harness cannot know which half is wrong, and does not guess. What it can say is that this row is not usable as training signal, because the stated preference is not supported by the working. It goes back for a second look before it ships.


No weighted scores, on purpose

An earlier version computed a weighted total per response. That was removed.

The deliverable of this task is the preference label. A weighted total quietly becomes a second answer to the same question, and the two drift apart — a grader who can see that one of five dimensions is a hard constraint and the other four are not will pick the right label, while an average will not. The loader now rejects any dimension carrying a weight.

Ratings exist to justify the preference. evalkit check verifies that they do.


What's here

8 coding probes Each names the mistake it expects, written before the prompt
5 dimensions Instruction following, code correctness, completeness, code quality, conciseness — anchored at every level, no weights
7-point preference A much better → About the same → B much better
6 comparisons Real code on both sides, all double-graded, rationale required on every non-tie
4 SFT corrections Fixed code, what changed, and what each example could teach wrongly
Consistency + agreement Preference-vs-ratings check, and quadratic-weighted kappa on both scales
pip install -r requirements.txt

python3 -m evalkit validate     # structural checks across rubric, probes, comparisons
python3 -m evalkit probes       # the probe set and the mistake each one targets
python3 -m evalkit compare      # side-by-side ratings and each grader's preference
python3 -m evalkit report       # aggregate across the whole set
python3 -m evalkit agreement    # do two graders reproduce each other?
python3 -m evalkit check        # exit non-zero if a preference contradicts its ratings

Writing prompts that provoke a predicted mistake

Anything makes a model fail if you push hard enough. A useful probe is one where the mistake is named in advance, so the score can be attributed:

- id: p01-tictactoe-antidiagonal
  family: multi_rule_spec
  prompt: |
    Write a Python function `check_winner(board)` for 3x3 tic-tac-toe.
    ...
    Winning lines are the three rows, the three columns, and both diagonals.
  probes: >
    Whether every rule in a multi-part spec survives. The anti-diagonal
    (top-right to bottom-left) is the line most often dropped.
  expected_failure: >
    Checks rows, columns, and the main diagonal, and silently omits the
    anti-diagonal -- so a genuine win goes unreported. Passes any test suite
    that does not specifically place three marks on that line.

It failed exactly that way. The correction records why: rows and columns came from loops that cannot miss a member, and the diagonals were the only lines written by hand. The bug lives in the asymmetry between generated and hand-written cases — which is the part that generalises, and the part a "you forgot a line" correction would never teach.

Eight families: multi-rule spec, side effect, complexity constraint, forbidden dependency, false API claim, swallowed errors, boundary, and one negative control. More in docs/PROMPT_DESIGN.md.


The rubric

Dimension The question it asks
Instruction following Did it do what it was told — signature, forbidden imports, required complexity?
Code correctness Does the code run and give the right result?
Completeness Is everything that was asked for there, edge cases included?
Code quality Would it pass review — and do the claims match the code?
Conciseness Is anything here that does not need to be?

Rules the loader enforces:

Every dimension declares what it separates_from. Most grader disagreement is not a judgement gap — it is two people scoring the same flaw on different dimensions:

Code can produce the right answer and still ignore the instructions: an O(n²) solution to a problem that required O(n) passes every test and fails here. Judge the instructions, not the output.

Every non-tie preference needs a written rationale. A preference label nobody can justify is noise wearing the shape of signal.

No weights. See above.


What the set says

which side won
  A 3   B 8   tie 1   (across 12 grader rows)

mean rating gap per dimension (B minus A)
  Instruction following            +1.42
  Code correctness                 +0.42
  Completeness                     +0.67
  Code quality                     +1.42
  Conciseness                      -0.75

Two readings worth having.

Instruction following is the widest gap (+1.42) and code correctness the narrowest (+0.42). Across these pairs the weaker responses are nearly as correct as the stronger ones — they lose on doing what they were told. c02 is the sharpest case: an O(n²) first_duplicate that is right on every input, including the subtlety most implementations miss, scored code correctness 5 by both graders and still lost the comparison.

Conciseness runs the other way (−0.75) — the weaker responses are consistently shorter. c03's losing answer is a single line, datetime.fromisoformat(text), which raises on Python 3.9 for the exact input the prompt gave. Short and wrong is a real combination, and folding conciseness into code quality would hide it.

The set is deliberately not one-sided: c05 and c06 are won by A. A set where the same side always wins measures the authors' expectations rather than the models, and graders who see B win five times start reading for it — a test enforces this.


Agreement

Every comparison is graded twice, independently:

DIMENSION                         EXACT  WITHIN1     KAPPA  MAX GAP
Instruction following               92%     100%     0.987        1
Code correctness                    92%     100%     0.981        1
Completeness                        67%     100%     0.913        1
Code quality                        75%     100%     0.935        1
Conciseness                         75%     100%     0.786        1

PREFERENCE LABEL                    67%      83%     0.873        3

mean dimension kappa: 0.920 -- almost perfect
preference kappa:     0.873 -- almost perfect
lowest agreement: conciseness -- this is where the anchors need rewriting

The preference label is the deliverable, so its agreement is the number that decides whether the data ships. Dimension kappa diagnoses where a low preference kappa comes from.

Kappa rather than raw agreement, because on a five-point scale where most ratings sit at 4 or 5 two graders who never read the rubric would still agree often. Quadratic weighting because both scales are ordinal — 4-vs-5 is a rounding difference, 1-vs-5 means the rubric failed, and the same argument applies with more force to the preference scale, where the one max_gap of 3 is c06.

Every disagreement is kept as text in a disagreement_note, never averaged away.


Layout

rubric/dimensions.yaml     5 dimensions, anchors at every level, 7-point preference scale
prompts/probes.yaml        8 coding probes with the mistake each expects
comparisons/*.yaml         two responses, two graders, ratings + preference + rationale
sft/corrections/*.md       fixed code with training-set caveats
evalkit/                   loader, compare (consistency), agreement, report, cli
docs/                      rubric design, prompt design, case study
tests/                     47 unit tests over consistency, kappa, and validation

Responses A and B are written for this repository as illustrative examples. They are not outputs from, and are not attributed to, any real product.

License

MIT.

About

Pairwise code evaluation: two responses rated on five dimensions, a 7-point preference, and a check that the preference matches the ratings.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages