Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
590c9ba
feat(fd3): re-enter build-spec at validation when handed a finished spec
grixu Sep 29, 2026
9adc319
fix(fd3): place the spec in its repository's layout, never the scratc…
grixu Sep 29, 2026
4bd217c
fix(fd3): relay write-spec questions through AskUserQuestion with num…
grixu Sep 29, 2026
fd80428
fix(fd3): re-validate every spec edit made after the final verdict
grixu Sep 29, 2026
97e23bd
fix(fd3): run validate-spec probes in a scratch worktree without touc…
grixu Sep 29, 2026
d3c50a0
fix(fd3): give a check with only non-blocking findings its own result…
grixu Sep 29, 2026
aef5bb4
fix(fd3): route an unvalidated spec to build-spec and never let the s…
grixu Sep 29, 2026
79e4c07
feat(fd3): record the commit each repository was validated at under t…
grixu Sep 29, 2026
30691ad
fix(fd3): stop the split when origin changed files the spec cites sin…
grixu Sep 29, 2026
9709c0d
feat(fd3): serialise tasks along a commit sequence the spec binds ins…
grixu Sep 29, 2026
1b6e214
fix(fd3): ask to reuse the branch the spec itself names even when it …
grixu Sep 29, 2026
4d19a66
fix(fd3): write the split report after the coverage re-run, with bare…
grixu Sep 29, 2026
aa82464
fix(fd3): re-check sized recommendations when an answer widens the gr…
grixu Sep 29, 2026
3b5034d
fix(fd3): pin every workflow agent to its own step, never the relayed…
grixu Sep 29, 2026
4e56e57
fix(fd3): cut the baseline worktree detached so a parked base branch …
grixu Sep 29, 2026
d6b984f
fix(fd3): read CI verdicts from each command's own exit status
grixu Sep 29, 2026
2d8bb8e
fix(fd3): send a CI failure that needs a behaviour change to the huma…
grixu Sep 29, 2026
921d90a
fix(fd3): commit each repair decision separately
grixu Sep 29, 2026
a3fcbfe
refactor(code-review): move the scanner and merge contracts out of st…
grixu Sep 29, 2026
b80fb0e
feat(code-review): let get_changes.py review another checkout with -C
grixu Sep 29, 2026
9f6713d
refactor(code-review): move the conventions note and lens-set rules i…
grixu Sep 29, 2026
5e1bf9b
refactor(code-review): define fix risk classes in the merge contract …
grixu Sep 29, 2026
3d3d945
feat(code-review): add headless cr-prepare, cr-scan and cr-merge skil…
grixu Sep 29, 2026
c325d73
fix(fd3): measure a root branch against origin/<default>, never the p…
grixu Sep 29, 2026
1ba876e
feat(fd3): review each branch with the code-review plugin's headless …
grixu Sep 29, 2026
4e77e3b
feat(fd3): review each repair's own commits and hand its findings to …
grixu Sep 29, 2026
bab381a
feat(fd3): ask for the code-review plugin's headless review and reope…
grixu Sep 29, 2026
b104013
docs(fd3): describe the headless code-review stage and when a reviewe…
grixu Sep 29, 2026
5e1440f
docs(code-review): document the headless cr-prepare, cr-scan and cr-m…
grixu Sep 29, 2026
672ab71
test(fd3): eval that the split stops when origin changed a file the s…
grixu Sep 29, 2026
4bd139d
test(fd3): eval that build-spec re-validates a finished spec instead …
grixu Sep 29, 2026
f3eb4ed
test(code-review): eval that the headless skills review a committed c…
grixu Sep 29, 2026
5d8b7b6
docs(fd3): changelog for the 2026-09-25 run tuning and headless review
grixu Sep 29, 2026
20c8dd7
docs(code-review): changelog for the headless review skills
grixu Sep 29, 2026
7718000
fix(code-review): let a headless skill's closing block end the skill,…
grixu Sep 29, 2026
0144f63
fix(fd3): keep spec findings out of the automatic review fixer
grixu Sep 29, 2026
98edd97
fix(fd3): hand boy-scout review findings to the human, never the auto…
grixu Sep 29, 2026
d39ba0b
fix(fd3): hold a branch on review findings the fixer left unfixed
grixu Sep 29, 2026
787d24b
docs(fd3): count check 9's result form as the fifth
grixu Sep 29, 2026
a2af1d5
test(fd3): accept a coded worker heading as closing the uncoded-eleme…
grixu Sep 29, 2026
dc32d83
test(fd3): tell the validate evals that no one answers the skill's qu…
grixu Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions plugins/code-review/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,16 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Added

- Headless skills for workflow callers — `cr-prepare` (scope, conventions, standards and the
active lens set, written to a context directory), `cr-scan` (one lens) and `cr-merge` (one
report, every finding returned with its fix-risk class). Model-only, they ask nothing and edit
nothing in the checkout, and an empty change or a missing lens returns a status, never a clean
review
- `get_changes.py -C <checkout>` reviews another checkout without changing directory
- Eval-20 runs the headless pipeline against a git sandbox

### Changed

- Author metadata now reads `Mateusz Gostański <mg@grixu.dev>` in `plugin.json` and the marketplace entry.
Expand Down
16 changes: 16 additions & 0 deletions plugins/code-review/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,22 @@ partial review, invoke `/comment-review` or `/quality-review` directly; both sta
independently available and share the same rule text as the command. The three
added lenses have no standalone skill.

### Headless skills for workflows

A workflow agent cannot answer questions or dispatch agents of its own, so `/start-cr`
cannot run there. Three model-only skills (hidden from the `/` menu) split the same
pipeline into steps a caller orchestrates, all eight lenses included:

| Skill | Arguments | Writes |
|---|---|---|
| `code-review:cr-prepare` | `--base <ref> --out <dir> [-C <checkout>] [--spec <path>]` | `scope.json`, `conventions.md`, `standards.md` |
| `code-review:cr-scan` | `--lens <lens> --context <dir>` | `<lens>.md` — run one per active lens, each as its own agent |
| `code-review:cr-merge` | `--context <dir>` | `report.md`, and returns every finding with its fix-risk class |

None of them edits the checkout or asks anything. An empty change, a lens that did not
report or a missing lens file comes back as a status, never as a clean review.
`fd3`'s implementation workflows are the first caller.

The report groups by **file**, with the two vocabularies side by side — comment
verdicts (`R1`–`R12` · KEEP/REMOVE/REWRITE/MOVE/ADD) and quality findings
(`` `family` · rule · severity `` across eleven families: `readability`, `tests`,
Expand Down
533 changes: 32 additions & 501 deletions plugins/code-review/commands/start-cr.md

Large diffs are not rendered by default.

7 changes: 7 additions & 0 deletions plugins/code-review/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,15 @@ evals/
prompts/standards.txt # quality trigger with the standards fixture dir as repo root
fixtures/ # inputs; fixtures/spec/ and fixtures/standards/ are multi-file
scope-mix/ # eval-19 input, kept out of fixtures/ so its paths classify by kind
prompts/headless.txt # cr-prepare → cr-scan × N → cr-merge, in that order — headless track
checks/ # deterministic asserts that read files on disk (headless track)
reset-sandbox.sh # rebuilds .sandbox/headless, the git checkout eval-20 reviews
```

The headless skills review committed changes only, so eval-20 runs against a git
checkout rather than a fixture path; `scripts/run-evals.sh code-review` rebuilds it
before every run.

Node dev deps (`@anthropic-ai/claude-agent-sdk` + `promptfoo`) and the run
scripts live at the **repo root** (`package.json`, single shared `node_modules`),
not per-plugin.
Expand Down
46 changes: 46 additions & 0 deletions plugins/code-review/evals/checks/headless-pipeline.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
import { execFileSync } from 'node:child_process';
import fs from 'node:fs';
import path from 'node:path';
import { fileURLToPath } from 'node:url';

const LENSES = ['comments', 'readability-tests', 'naming-module', 'objects-patterns', 'simplicity-types', 'security', 'performance', 'spec'];
const FINDING = /^- (high|medium|nit) · [\w-]+ · [\w-]+ · \S+:L\d+(?:-\d+)? · (safe|structural|report-only)\b.* — /m;

export default (output, context) => {
const { checkout, context: dir } = context.vars;
const root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..', '..', '..', '..');
const ctx = path.resolve(root, dir);
const failures = [];
const check = (ok, msg) => {
if (!ok) failures.push(msg);
return ok;
};

const asked = (context.providerResponse?.metadata?.toolCalls || []).filter((c) => c.name === 'AskUserQuestion');
check(asked.length === 0, `a headless skill asked ${asked.length} question(s)`);

const scopePath = path.join(ctx, 'scope.json');
if (check(fs.existsSync(scopePath), 'cr-prepare wrote no scope.json')) {
const scope = JSON.parse(fs.readFileSync(scopePath, 'utf8'));
const active = scope.lenses.active.map((l) => l.lens);
const named = [...active, ...scope.lenses.inactive.map((l) => l.lens)].sort();
check(JSON.stringify(named) === JSON.stringify([...LENSES].sort()), `scope.json does not account for all eight lenses once: ${named.join(', ')}`);
check(!active.includes('spec'), 'the spec lens is active with no spec named');
check(scope.files.map((f) => f.path).sort().join() === 'src/quality-recall.ts,src/security-recall.ts', `judged files are not the committed change: ${scope.files.map((f) => f.path).join(', ')}`);
for (const lens of active) {
const p = path.join(ctx, `${lens}.md`);
check(fs.existsSync(p) && /^Lens: /m.test(fs.readFileSync(p, 'utf8')), `cr-scan left no ${lens}.md with its Lens header`);
}
}

const report = path.join(ctx, 'report.md');
check(fs.existsSync(report) && /Reconciliation/.test(fs.readFileSync(report, 'utf8')), 'cr-merge wrote no report.md with a Reconciliation line');
check(/status: merged/.test(output), 'cr-merge did not return status: merged');
check(FINDING.test(output), 'no finding line in the fixed `severity · family · rule · path:L · class — fix` shape');
check(/^- high · security · /m.test(output), 'the planted security finding did not come through the merge');

const dirty = execFileSync('git', ['-C', path.resolve(root, checkout), 'status', '--porcelain'], { encoding: 'utf8' }).trim();
check(dirty === '', `the review edited the checkout: ${dirty}`);

return { pass: failures.length === 0, score: failures.length === 0 ? 1 : 0, reason: failures.join('; ') || 'all checks passed' };
};
15 changes: 15 additions & 0 deletions plugins/code-review/evals/promptfooconfig.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,10 @@ prompts:
# for the CODING_STANDARDS pair.
- id: file://prompts/standards.txt
label: standards-track
# Headless track: the three model-only skills a workflow drives, run in sequence against a
# committed change in a git sandbox that reset-sandbox.sh rebuilds.
- id: file://prompts/headless.txt
label: headless-track

providers:
- id: anthropic:claude-agent-sdk # alias: anthropic:claude-code
Expand Down Expand Up @@ -805,3 +809,14 @@ tests:
- type: regex
value: '[Ss]kipped'

- description: 'eval-20 headless-pipeline (cr-prepare → cr-scan × N → cr-merge)'
prompts: [headless-track]
vars:
checkout: plugins/code-review/evals/.sandbox/headless
context: plugins/code-review/evals/.sandbox/headless-ctx
# Eight lenses in one agent's context: the default cap would abort mid-run. The skills write
# their context files; without Write and Bash the agent narrates a review it never ran.
options: { max_budget_usd: 12, allow_all_tools: true }
assert:
- type: javascript
value: file://checks/headless-pipeline.mjs
10 changes: 10 additions & 0 deletions plugins/code-review/evals/prompts/headless.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
Review the last commit of the git checkout {{checkout}} with the code-review plugin's headless
skills, exactly in this order, through the Skill tool:

1. `code-review:cr-prepare` with `--base HEAD~1 --out {{context}} -C {{checkout}}`.
2. `code-review:cr-scan` with `--lens <lens> --context {{context}}`, once for every lens its
closing block lists as active.
3. `code-review:cr-merge` with `--context {{context}}`.

Then reply with the three skills' closing blocks, verbatim, in that order — the cr-scan blocks
one after another — and nothing else.
22 changes: 22 additions & 0 deletions plugins/code-review/evals/reset-sandbox.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
# Rebuild the git checkout the headless-track eval reviews: a base commit, then one commit that adds
# two fixtures with planted findings. The headless skills review committed changes only, so the
# file-path fixtures the other tracks read cannot serve them. scripts/run-evals.sh runs this first.
set -euo pipefail
cd "$(dirname "$0")"

rm -rf .sandbox
repo=.sandbox/headless
mkdir -p "$repo/src"
git -C "$repo" init -q -b main
printf '# orders service\n' > "$repo/README.md"
git -C "$repo" add -A
git -C "$repo" -c user.name='cr-evals' -c user.email='cr-evals@localhost' commit -q -m 'base' --no-gpg-sign

# Under src/, not fixtures/: scope.md classes anything below a fixtures/ directory as test code.
cp fixtures/security-recall.ts "$repo/src/security-recall.ts"
cp fixtures/quality-recall.ts "$repo/src/quality-recall.ts"
git -C "$repo" add -A
git -C "$repo" -c user.name='cr-evals' -c user.email='cr-evals@localhost' commit -q -m 'feat: add order lookup and pricing' --no-gpg-sign

echo "code-review sandbox reset: $(pwd)/$repo"
Loading
Loading