LLM agents play Reversi against each other through a local referee CLI. The board is just a display — the CLI holds the truth.
The ranked ladder also includes approved direct-choice API players, labeled direct-choice. Their adapters only pass referee state and submit the selected move; they do not search or choose moves themselves. See interface differences and the JEV pilot admission.
View the live standings and match results on AGI Labo
- The CLI is the source of truth: positions, legal moves, and results are all decided by
reversi.mjs. Players only pick from the legal-move list, and the referee rejects anything else - Players never touch a browser: player I/O is JSON only. The board page exists for spectators
- Records are replay-verifiable:
matches/*.jsonkeeps the full move history, so anyone can replay a game and verify its legality
node reversi.mjs new --size 8 \
--black "Opus 5 · high" --white "Grok 4.6 · medium" \
--black-model claude-opus-5 --white-model grok-4.6
node reversi.mjs serve --port 8765 # spectator board: http://127.0.0.1:8765Launch two agents and hand each one prompts/player.md plus a side (B / W). Each player then runs on its own until the game is over:
node reversi.mjs wait --as B --json # wait for your turn
node reversi.mjs play d6 --as B --jsonWith AGI Cockpit you can pin a seat (model × effort) per task:
cockpit task create --agent-type claude --effort high \
--directory /path/to/reversi-bench \
--instruction "$(cat prompts/player.md) — Your side: B"node reversi.mjs new [--size 4|6|8] [--id current] [--black Name] [--white Name]
[--black-model id] [--white-model id]
node reversi.mjs state [--id current] [--json]
node reversi.mjs play <coord|pass> --as B|W [--id current] [--json]
node reversi.mjs wait --as B|W [--timeout 120] [--id current]
node reversi.mjs thinking <B|W|clear> [--id current]
node reversi.mjs say <B|W> <text...> [--id current]
node reversi.mjs serve [--id current] [--port 8765]
node reversi.mjs selftest
Game records live in matches/<id>.json. standings.json holds the aggregated ladder standings and game index, regenerated mechanically after every game:
node standings.mjsThese matches are unranked and use a separate protocol and record directory. They never enter the ordinary season standings.
node reversi.mjs new --size 8 --confidence-after 10 --id current --json
node reversi.mjs play e3 --as B --win 70 --draw 10 --loss 20 --json
node reversi.mjs state --spectator --jsonThe play command is an example for a legal move once measurement.required is true. The first ten played moves use the ordinary command; played move 11 onward requires three percentages summing to 100. Passes do not count. Forecasts describe the acting player's eventual win/draw/loss against the current opponent after the selected move, not optimal-move confidence or perfect-play outcomes.
Players use prompts/player-confidence.md through side-bound referee wrappers. Accepted forecasts are saved with the move, but stripped from player-facing CLI output. Only the spectator state and Discord relay expose them. See experiments/README.md for recording and validation rules.
Exhibition:
| Game | Result |
|---|---|
| 8×8 · Grok 4.6 (Black) vs Grok 4.5 (White) | Black 44–20 — matches/exhibition-8x8-grok-4-6-vs-grok-4-5.json |
Optional: stream a running match to Discord as one board screenshot per move. See live/README.md. The referee CLI stays independent of it.
How wins, stone margins, and incident rates are measured — and how model × effort seats are ranked — is described in METHOD.md.
MIT