Can LLMs read the room? — 空気読めるかベンチマーク。
KY-Bench measures social air-reading in LLM agents through cooperative games: how accurately an agent reads what others mean (decoding), how legibly it signals what it means (legibility), and whether it stands its ground when the group consensus is wrong. Every game is played fully blind — players know each other only as P1…Pn — and sealed estimates turn the conversation into per-seat quantitative metrics, scored mechanically with no LLM judging.
A cooperative number-ordering game, inspired by the party game ito. Players secretly hold numbers from 1 to 100, express them only through themed hints ("how scary", "how heavy"), and try to play their cards in ascending order. The referee CLI holds the truth.
- Legibility — how accurately other players read your hints (their error on your number)
- Decoding — how accurately you read everyone else's hints
- Bias — whether a seat systematically overstates or understates
- Consistency — whether actions (play / pass) match the seat's own sealed estimates
- Team result — mistakes per round (cards skipped by wrong ordering)
See METHOD.md for definitions.
standings.json holds the aggregated rankings and a per-game index, regenerated mechanically from matches/ after every game:
node standings.mjsnode referee.mjs new --players 3 --themes "怖いもの"
node referee.mjs serve --port 8768 # spectator page: http://127.0.0.1:8768One round = one game is the default: fresh agents and a fresh theme per game keep every data point independent (see METHOD.md).
Each player loops on its own, using only these commands:
./referee wait # what do I need to do?
./referee clue 夜中にきしむ床の音 # express your number through the theme
./referee guess P2=40 P3=75 # sealed estimates of the others
./referee play # or: ./referee passPlayers never call referee.mjs directly. node referee.mjs state returns every secret number — it is the runner's view — so any player that can reach the script can read the whole game. Blinding is enforced by handing each player a throwaway directory whose only referee is a shim with --as <player> baked in:
ARENA=$(pwd) # wherever `new` wrote matches/
for p in P1 P2 P3; do
D=$(mktemp -d)
printf '#!/bin/sh\nexec node "%s/referee.mjs" "$@" --as %s\n' "$ARENA" "$p" > "$D/referee"
chmod +x "$D/referee"
# launch one agent in $D with prompts/player.md as its instruction
doneThe shim appends --as, so a player cannot override it or act as another seat. Play in an arena outside this checkout as well, so players cannot reach matches/, past records, or standings.json.
Which model sits in which seat lives in a file that only the runner and the keyed spectator page ever see:
{"seats": {"P1": {"agentType": "claude", "model": "claude-opus-5", "effort": "medium"},
"P2": {"agentType": "codex", "model": "gpt-5.6-sol", "effort": "high"}}}serve --seats <file> --key <token> reveals them at ?key=<token> only. A display string on a seat overrides the derived model · effort label.
--discussion makes players talk between the two sealed estimates. The referee never posts and knows nothing about the chat: it holds the round at phase hold until the runner runs open, and passes the room id to players as data.room. Everything else is the runner's job:
- create the room and pass its id with
--room - post the round header, then
node referee.mjs opento release the players - mirror the transcript into a JSON file for
serve --talk
prompts/player.md speaks to the room with cockpit talk (AGI Cockpit). Swap those two lines for your own chat CLI — or run without --discussion, which needs nothing but the referee.
The referee leaves the played game at matches/<id>.json with no seats and no summary, because seats are revealed only after the game ends. record.mjs merges them in and writes the archive that standings.mjs reads:
node record.mjs --game g003 --seats seats.json --arena "$ARENA" [--talk talk.json]
node standings.mjsRun node record.mjs --help for notes, token usage, and the other optional fields.
node referee.mjs new [--players 3] [--rounds 1] [--themes a,b,c] [--id current]
[--discussion [--room talk_xxx] [--cycles 2]]
node referee.mjs state [--id current]
node referee.mjs wait --as P1 [--timeout 120]
node referee.mjs clue <text...> --as P1
node referee.mjs guess P2=40 P3=75 --as P1
node referee.mjs said --as P1
node referee.mjs open [--id current]
node referee.mjs play | pass --as P1
node referee.mjs report [--id current]
node referee.mjs serve [--port 8768] [--key token] [--talk talk.json] [--seats seats.json]
node referee.mjs selftest
node record.mjs --game g003 --seats seats.json [--arena dir] [--id current]
[--talk talk.json] [--notes notes.json] [--tokens tokens.json]
[--kind text] [--blind false] [--force]
serve shows a neutral spectator page; adding ?key=<token> (set via --key) reveals secret numbers, estimates, and — when --seats points at a seat-assignment file — each player's identity ({"seats":{"P1":{"agentType":"...","model":"...","effort":"..."}}}, or a display string per seat) to spectators only. Players never see any of it.
--talk embeds a live discussion feed at the bottom of the spectator page. Point it at a JSON file shaped {"messages":[{"seq":1,"name":"P1","text":"...","at":"ISO"}]} and keep the file updated with whatever chat system hosts the discussion — the runner mirrors the transcript in, the referee stays chat-agnostic.
MIT