Skip to content

Repository files navigation

Reversi Bench

LLM agents play Reversi against each other through a local referee CLI. The board is just a display — the CLI holds the truth.

The ranked ladder also includes approved direct-choice API players, labeled direct-choice. Their adapters only pass referee state and submit the selected move; they do not search or choose moves themselves. See interface differences and the JEV pilot admission.

View the live standings and match results on AGI Labo

Design principles

  • The CLI is the source of truth: positions, legal moves, and results are all decided by reversi.mjs. Players only pick from the legal-move list, and the referee rejects anything else
  • Players never touch a browser: player I/O is JSON only. The board page exists for spectators
  • Records are replay-verifiable: matches/*.json keeps the full move history, so anyone can replay a game and verify its legality

Quick start

node reversi.mjs new --size 8 \
  --black "Opus 5 · high" --white "Grok 4.6 · medium" \
  --black-model claude-opus-5 --white-model grok-4.6

node reversi.mjs serve --port 8765   # spectator board: http://127.0.0.1:8765

Launch two agents and hand each one prompts/player.md plus a side (B / W). Each player then runs on its own until the game is over:

node reversi.mjs wait --as B --json   # wait for your turn
node reversi.mjs play d6 --as B --json

With AGI Cockpit you can pin a seat (model × effort) per task:

cockpit task create --agent-type claude --effort high \
  --directory /path/to/reversi-bench \
  --instruction "$(cat prompts/player.md) — Your side: B"

CLI

node reversi.mjs new [--size 4|6|8] [--id current] [--black Name] [--white Name]
                     [--black-model id] [--white-model id]
node reversi.mjs state [--id current] [--json]
node reversi.mjs play <coord|pass> --as B|W [--id current] [--json]
node reversi.mjs wait --as B|W [--timeout 120] [--id current]
node reversi.mjs thinking <B|W|clear> [--id current]
node reversi.mjs say <B|W> <text...> [--id current]
node reversi.mjs serve [--id current] [--port 8765]
node reversi.mjs selftest

Records

Game records live in matches/<id>.json. standings.json holds the aggregated ladder standings and game index, regenerated mechanically after every game:

node standings.mjs

Experimental outcome-confidence matches

These matches are unranked and use a separate protocol and record directory. They never enter the ordinary season standings.

node reversi.mjs new --size 8 --confidence-after 10 --id current --json
node reversi.mjs play e3 --as B --win 70 --draw 10 --loss 20 --json
node reversi.mjs state --spectator --json

The play command is an example for a legal move once measurement.required is true. The first ten played moves use the ordinary command; played move 11 onward requires three percentages summing to 100. Passes do not count. Forecasts describe the acting player's eventual win/draw/loss against the current opponent after the selected move, not optimal-move confidence or perfect-play outcomes.

Players use prompts/player-confidence.md through side-bound referee wrappers. Accepted forecasts are saved with the move, but stripped from player-facing CLI output. Only the spectator state and Discord relay expose them. See experiments/README.md for recording and validation rules.

Exhibition:

Game Result
8×8 · Grok 4.6 (Black) vs Grok 4.5 (White) Black 44–20 — matches/exhibition-8x8-grok-4-6-vs-grok-4-5.json

Live streaming

Optional: stream a running match to Discord as one board screenshot per move. See live/README.md. The referee CLI stays independent of it.

Methodology

How wins, stone margins, and incident rates are measured — and how model × effort seats are ranked — is described in METHOD.md.

License

MIT

About

LLM agents play Reversi through a local referee CLI — records, methodology, and a live spectator board

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages