Skip to content

Latest commit

ย 

History

760 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿ›๏ธ When Agents Rule

Where language models battle for the crown in antiquity.

Up to four LLMs. One map. One winner. A browser-based, Age-of-Empires-style real-time strategy game in which competing language models play against each other โ€” while you watch, coach, and score them.

License: MIT No build step Zero dependencies Providers


A Persian army destroys an Egyptian Wonder during a live four-model match

A Wonder falls. The army reforms. Real gameplay from a four-model match, recorded September 9, 2026, at normal playback speed. Click for the sharper, longer video.


What is this?

A sandbox arena for pitting language models against one another at a task they were never trained for: running an economy and an army, in real time, inside a small RTS they've never seen. Every player is an autonomous model agent governing its own civilization, and every match ends one of two ways โ€” a rival razed to the ground, or a Wonder held in peace.

This is an agent harness whose task happens to be a real-time strategy game. Every turn a model is handed a compact JSON snapshot of its situation (resources, buildings, units, fog-of-war discoveries, threats, tech tree, map bounds), named command tools โ€” such as move_units, attack_target and research_tech, sharing a budget of three commands per turn โ€” plus an independent plan tool to save its objective โ€” and one instruction: win. Then it has to keep doing that, turn after turn, for a whole match.

The tools are real tool calls, extracted from the chat-completion response the same way OpenCode or any other agent harness extracts them. That matters for what a result means: a malformed call costs that call, not the turn, and a model gets the same per-call feedback it would get anywhere else it is deployed.

It's a hands-on testbed, not a benchmark โ€” see Disclaimers. With the setup held steady (same civilization, fixed map seed, resource layout equal for every player) a match isolates the model well enough for narrow comparisons.

A growing settlement An army on the march
A developed settlement during the recorded arena match Persian troops moving through an enemy settlement

From economy to conquest: still frames from the same live match. These settlements and armies were built and commanded by the competing models.

Why it's an interesting eval

Most quick LLM demos reward a single clever answer. A full match rewards the things people actually care about in agents:

  • ๐ŸŽฏ Three actions a turn, and the turn will not wait. The budget is the eval. A model does not get to do everything it can think of โ€” it has to answer what are the three most important things right now, on a board that changes between its own calls. Lift the cap and the answer after a wipe is simply "rebuild all of it at once", which measures a bank balance rather than judgement.
  • โš”๏ธ Models develop their own doctrine. Same rules, same prompt โ€” yet you get pure economists racing for a Wonder next to warlords massing an army absurdly early. Which one a model turns out to be is part of what you're measuring.
  • ๐Ÿงจ Pressure changes their play. Raided, out-scouted, slipping down the leaderboard โ€” many models genuinely switch tactics rather than doubling down. Watching one notice it is losing is worth the match on its own.
  • ๐ŸŽฏ Precise tool calling. Every move must be valid JSON โ€” one action, or up to three in one reply. Hallucinate a tool, fumble the schema, wrap it in prose โ€” the turn is wasted, and you can watch format discipline hold or crumble.
  • ๐Ÿงญ An unfamiliar framework. No fine-tuning, no examples of good play. Only the rules in the prompt and the state in front of it.
  • ๐Ÿง  Long-horizon strategy. Economy โ†’ tech โ†’ military โ†’ conquest plays out over dozens of turns. Models that optimize forever and never build an army lose. The harness carries a model-authored objective + plan (up to 10 steps) across turns โ€” but maintaining it is the model's job.
  • ๐Ÿ” Error recovery. A rejected action comes back with a precise reason. Does the model correct course, or bang on the same locked door?
  • ๐Ÿ—บ๏ธ Spatial reasoning. Fog of war hides the map; resources and enemies must be scouted before they can be used or attacked.
  • โฑ๏ธ Latency vs. quality. In real-time mode faster models simply act more often, and a brilliant-but-slow model gets out-tempoed. Turn-based rounds remove that variable when you want to compare judgement instead.

A Persian army moving in formation through an Egyptian settlement

Orders keep working between model replies. An army moves through enemy territory in the recorded match. Normal playback speed; click for the higher-quality video.

โœจ Features

The match

  • ๐Ÿค– 2โ€“4 models, fighting live โ€” each with its own asynchronous decision pipeline.

  • โณ Two tempos โ€” real time, where faster models act more often, or turn-based rounds: every seat reads the same snapshot, all moves land together, and a configurable answer time (default 90 s) keeps one slow endpoint from stalling the rest. Missed rounds are recorded rather than hidden.

  • โฑ Speed control โ€” 1ร— / 1.5ร— / 2ร— / 4ร—, plus Pause (which waits for answers already in flight, so no move is thrown away). Held at 1ร— while a Wonder stands, so the countdown can't be sped past.

  • ๐ŸŒฑ Seeded maps & fair placement โ€” the same seed reproduces the exact layout; food and wood fill an even 7ร—7 grid while scarce stone and gold are placed identically for every player. The map stops being a confound.

  • ๐Ÿ›ก๏ธ Persistent army orders โ€” march, scout without initiating combat, guard a position, or patrol between two points. Incoming damage triggers formation-wide retaliation, with towers first and bounded pursuit of mobile threats; survivors reform and resume their assignment, including priests in ranged slots. A new order replaces the selected units' previous assignment.

  • ๐ŸŒ— Day and night โ€” a cosmetic 12-minute cycle with warm dusk, cool nights, campfires and entrance lights. Simulation speed does not accelerate the lighting cycle.

The models

  • ๐Ÿ”Œ Bring any model โ€” OpenAI-compatible (OpenAI, vLLM, LM Studio, LiteLLM, Groq, OpenRouter, โ€ฆ), Anthropic, Ollama, Google (Gemini), with auto-detection. Mix local and cloud in one match.
  • ๐Ÿ” Every auth style โ€” none, API key (Bearer), header secret, Basic, or an OAuth2 login button: one click for OpenRouter, or any login server that issues a client ID.
  • ๐Ÿงฐ Model library โ€” add, test connection, pick the served model, and set per-model max tokens, context budget, language, temperature / top-p / top-k, thinking/reasoning settings, and a raw request-body passthrough for anything newer than this harness. Connection, model/budgets, and collapsible advanced settings keep the basics together. Saved locally, exportable/importable.
  • ๐Ÿง  Rolling context that scales with the model โ€” history is sized to each model's context budget, so a 128K model remembers more of the match than a 32K one. Default is a real multi-turn conversation; a minimize-tokens toggle switches to compact one-line history.
  • ๐Ÿช™ Token accounting โ€” provider-reported prompt + completion usage per model, next to latency.

Watching & reading it back

  • ๐Ÿ›ฐ๏ธ Live spectator dashboard โ€” ranked leaderboard, streaming decision log (every move plus the model's stated reason, rejections flagged), per-model advice chat, and play/pause per model. Collapse the decisions panel to the latest turnโ€™s command headlines for each seat.
  • ๐Ÿ›๏ธ Antiquity visual milestone โ€” warm directional light, cast shadows, irregular coastal water reflections, sculpted miniature units with cultural facial hair, angled handheld equipment and proportionate horses, continuous tree crowns, layered meadow/snow/sand surfaces with close-up ground detail, refined Greek architecture and camera controls integrated with the minimap. Arena configuration and the model library share the charcoal/bronze theme. Try Explore civilizations locally: choose civilization, summer/winter/desert, Stone through Iron Age, and time of day. Iron Age includes the civilization's Wonder; Inspect workers gives a close-up. Direct links: /?showcase=1&civ=greek (also egyptian, yamato, persian). The live transcript uses eight-turn pages to keep long matches responsive. Scope, graphics settings and verification.
  • ๐ŸŽฌ A battlefield worth watching โ€” feathered fog of war, arrows and tower stones, hit flashes, animated deaths, battle pings, per-map ground cover, and an optional action camera that follows the fighting.
  • ๐Ÿ“Š End-of-match evaluation โ€” latency, decisions, action-success rate, format fidelity, reasoning rate, error breakdown, behavior tags, and a transparent 0โ€“100 match heuristic (not a capability score). Each seat carries a conditions id: seats sharing it had the same declared conditions.
  • ๐Ÿ“„ Exports โ€” the evaluation as a self-describing results_<datetime>.md, and the full transcript as JSONL: every state sent, every reply, every harness answer, with the results and the economy timeline appended at the end.
  • ๐ŸŽž๏ธ Analyze Transcript โ€” load a saved transcript and read a finished match back, turn by turn, in the same 3D engine.
  • ๐ŸŒ Fully localized UI โ€” English, German, Spanish, Simplified Chinese, with the model's language chosen separately from the interface language.
  • โŒจ๏ธ Reading and navigation โ€” keyboard-accessible mode selection and modal dialogs, selectable transcript text, readable model names, and reduced interface motion when requested by your device settings.

The rest

  • ๐ŸŒ™ Keeps running in a background tab โ€” a Web-Worker driver keeps the simulation and the models' turns going while the tab is hidden.
  • ๐ŸŽฎ Also human-playable โ€” a Campaign mode: pick your civilization and face 1โ€“5 opponents, model- or AI-controlled, on three maps. If a model's endpoint dies mid-game that opponent falls back to the rule-based AI.
  • ๐Ÿšซ No build step, no dependencies โ€” plain HTML/CSS/JS with an in-house WebGL engine; every texture painted procedurally at load. No CDN, no assets, no external code.

๐Ÿš€ Quick start

No install, no bundler, nothing downloaded. Serve the folder over HTTP (the app uses fetch, so file:// won't work).

git clone https://github.com/asp67/when-agents-rule.git
cd when-agents-rule
node serve.cjs --open                 # Node 18+, built-ins only
# python3 -m http.server 8088         # or any static server

Then open http://localhost:8088 and click Play โ†’ ๐ŸŸ๏ธ Arena. serve.cjs answers only this computer by default; add --host 0.0.0.0 to reach it from other devices on your network, or --port to move it. The default is 8088 rather than 8080 because llama.cpp's server listens on 8080.

๐Ÿ’ก Fastest path to a match: install Ollama, pull something small and quick (ollama pull qwen2.5:7b), and point a couple of seats at http://localhost:11434. Small + fast beats large + slow in a real-time arena.

๐Ÿฆ Easy first pick: ollama pull ornith:9b โ€” ~6 GB of VRAM, runs on many consumer cards, and a surprisingly strong player for its size.

๐ŸŸ๏ธ Setting up the Arena

  1. Model Library โ†’ add your models. Set the endpoint, pick the provider (or leave on auto-detect), choose auth, hit ๐Ÿ”Œ Test connection, select the served model. Optionally set max tokens, context budget, language, and sampling/thinking parameters.
  2. Arena participants โ†’ choose 2โ€“4 seats, then give each a civilization and a controller. Set the difficulty, an optional map seed, and whether to run turn-based rounds (with the answer time per round).
  3. System prompt โ†’ tweak the shared template, or give individual seats their own.
  4. โš”๏ธ Start Arena and watch.

Current model library and connection settings

The current model library: configure a provider, connection and model settings. Fresh browser profile; no credentials or private endpoints shown.

While spectating you can click a card to fly the camera to that base, drag to pan, send a model advice, or pause one entirely. The decision log streams every move with the model's own reason, and flags rejections:

Input Spectator / replay Campaign
Left click / tap Inspect a unit or building Select a unit or building
Left mouse drag Pan map; no box selection Box-select units
Right mouse drag Read coordinates while held Pan map; releasing a drag issues no order
Right click Coordinate flag Move, attack, repair or confirm placement
One-finger drag / pinch Pan / zoom Pan / zoom
Touch and hold Coordinate flag Hold still, then release to issue an order

For left-button-only campaign navigation, open the minimap's camera-options chevron and enable the hand tool. Switch it off to restore mouse box selection. Touch always uses one-finger panning. Moving after a hold, adding another finger or cancelling the gesture discards the pending touch order. The in-game controls card describes these bindings in all four UI languages.

Greek village with a watchtower, troops, workers and civic buildings

A Greek village at ground level: troops gather beside the watchtower while workers move among homes, the temple and the town center.

๐Ÿ’ก The context budget is a real lever. Default 65536 tokens; โ†บ Max fills in the model's true maximum. History is sized to it, in one of two modes: multi-turn (past turns replayed as compact state recaps plus the model's replies โ€” richest memory) or minimize tokens (each past move as one line โ€” cheapest, still coherent). Either way the prompt is rebuilt from scratch every turn, and if an endpoint rejects a request as too large the harness shrinks the window and keeps playing.

Lower budgets are much faster โ€” on Ollama the budget also sets num_ctx, and an oversized window can spill the model onto the CPU. For small local models, 32K is often enough and noticeably faster. If a model overthinks, raise its max tokens (the output budget), not its context, so it can finish reasoning and still emit the JSON action.

๐Ÿ”ง Tool calls, and which stack served them

Models act by calling tools, on all four protocols โ€” OpenAI-compatible (vLLM, llama.cpp, LM Studio, OpenRouter, Groq โ€ฆ), Ollama, Anthropic and Google. Named game-command tools share one definition translated into each dialect. Up to three game commands run in order per turn, plus one independent plan call when the objective or plan changes. A plan-only reply is a successful plan update; saved plan steps are not automatically executed.

A seat that cannot work the tools fails visibly. That is the point rather than a rough edge: a harness that quietly compensates for a broken tool-call parser hides the one thing its operator needs to know. When no call arrives the error says which fault it is โ€” tool syntax found in the raw reply means the model called and the server missed it (a wrong --tool-call-parser on vLLM, a chat template without a tool section on llama.cpp), and no syntax means the model simply did not call.

For older or smaller models, and for endpoints whose parser or template is broken, each seat has an Accept inline JSON switch. Off by default. Switching it on is a declaration, not a convenience โ€” it is recorded in the transcript, because a seat allowed to fall back is scored on a softer contract than one that is not, and every turn records whether it was answered by tool_call or by content.

A result belongs to a model and a stack. A model that works through Ollama and fails through OpenRouter is one of the most useful things this can tell you, so every seat records servedBy โ€” what the server calls itself (vllm, llamacpp, ollama, โ€ฆ), asked once at match start and never guessed from the endpoint, which is deliberately not stored. Read a number as (model ร— stack ร— settings); a bare per-model figure is not reproducible.

๐Ÿ’ก On llama.cpp, set --reasoning-format deepseek. The default is auto, and auto drops the reasoning trace as soon as it delivers tool calls โ€” the tokens are billed, the thinking is gone, and the transcript shows an empty Reasoning panel on exactly the turns that went best. Set explicitly, the trace arrives alongside the calls and you can read along while the model thinks. (deepseek-legacy also leaves <think> tags in the content, which is harmless here because content is not evaluated when tool calls arrive.)

๐ŸŽž๏ธ Analyze Transcript

A third mode beside Arena and Campaign, and the other half of the round trip: a match records itself, gets downloaded, gets handed to someone else โ€” and opens here.

Reading a finished match back in the analyzer

Episode 6 reopened in build 819: an existing recorded match rendered with the current engine, with the turn list, saved plan and economy graph. The header retains the original recording's build and prompt version.

Load a match-*.jsonl and you get:

  • The board in the real engine. The arena's own renderer with the full camera โ€” pan, rotate, zoom. The map is rebuilt exactly from the recorded seed, and fog is per seat: only what that model had discovered is lit, with opacity scaled to how much of each tile it swept. Switch to All seats for the cumulated view, where ground nobody ever scouted stays dark.
  • What the model was thinking. Its standing objective and plan (carried forward, and flagged on the turns it rewrote them), the command it sent, its stated reason, its raw reasoning, and the harness's answer.
  • A timeline you can scrub. Step through turns, play at 0.5, 1, 2, or 4 entries a second through the filtered list, use the all-entry slider, click the economy graph to jump, or use the chapter list โ€” age advances, wonders, exhausted resources, combat. Filters narrow it to combat, harness errors, plan rewrites or missed rounds.
  • Saved reading views. Choose Balanced, Watch, Read, or Compare, adjust the map height with the divider (drag or arrow keys), and choose comfortable or compact text. These preferences and playback rate stay in this browser, separately from exported match/model settings.
  • Visible camera controls. Open Camera for a whole-map overview, selection focus, reset, zoom, and rotation. Manual camera input turns automatic following off. These controls are also available in live play.
  • Click anything to inspect it. Remembered enemy positions render translucent and say when they were last seen, so a stale sighting never looks like a live one.

Nine real matches ship with the game, in samples/, with every plan, command and reasoning block in them. No key, no endpoint, nothing to configure: the analyzer only ever reads a file. samples/index.json lists them with their models, tempo and result.

Direct match links open the viewer with a specific catalogue entry loaded: Episode 6, Episode 7, Episode 8 and Episode 9. Use ?match= followed by the entry's matchId; unknown IDs show a loading error rather than a different match.

  • 2026-07-26_opus5-grok4.5-gpt-oss_36min.jsonl โ€” 271 turns, turn-based (60 s a round). Opus 5 as Persia, Grok 4.5 as Egypt, and gpt-oss on a single consumer GPU as Yamato. Opus 5 wins.

  • 2026-08-09_kimi-k3-gemma4-qwen3.8-ornith9b_25min.jsonl โ€” 490 turns, real time, nobody waiting for anybody. Kimi K3 as Yamato, a local Gemma 4 26B as the Greeks, Qwen 3.8 Max as Persia, and a 9B quant on a desktop GPU as Egypt. Kimi K3 wins, having issued three commands a turn to the others' one.

  • 2026-08-11_muse-glimmer-qwen3.6-gemma4_88min.jsonl โ€” 471 turns, turn-based (60 s a round), and the longest of the three-player games. Muse-Glimmer 30B as the Greeks, Gemma 4 31B as Egypt, Qwen 3.6 27B as Persia. Persia wins on 3039 power against 164 and 117.

  • 2026-08-17_qwen3.8-opus4.6-qwen3.6-muse-glimmer_125min.jsonl โ€” 544 turns over two hours, turn-based (120 s a round), four seats. Qwen 3.8 27B as Yamato, Claude Opus 4.6 as Egypt, Qwen 3.6 27B as Persia, Muse-Glimmer 30B as the Greeks. Yamato finishes last one standing on 4617 power; Egypt finishes fourth on 129, its town centre destroyed and its last recorded thought reading โ€œNeed food desperately.โ€ The 27B open model beats the frontier model on this board โ€” one board, one seed, not a verdict.

  • 2026-08-26_deepseek-v4-glm5.3-qwen3.8-gpt5.6_103min.jsonl โ€” 389 turns, turn-based (150 s a round), four seats, and the one where a Wonder nearly changed the result. deepseek-v4-flash as Persia, glm-5.3-flash as the Greeks, Qwen3.8 27B running locally as Egypt, gpt-5.6-luna as Yamato. Persia commits to the Iron Age at 38:32 while everyone else is still in Bronze, then trains champions and nothing else for forty minutes. Greece pays everything it has for a Wonder at 75:00 โ€” hold it 600 seconds and the match is yours โ€” and loses it after 154. Egypt is eliminated with 10,070 food and zero workers left to spend it.

  • 2026-09-07_gemini-flash-deepseek-v4-gpt5.6-qwen3.8_89min.jsonl โ€” Episode 6, The Architect in the Ashes: 427 turns over 89:50, turn-based (120 s a round). Gemini Flash as Egypt, DeepSeek v4 Flash as the Greeks, GPT 5.6 Luna as Persia, and Qwen3.8 Flash as Yamato. The full transcript includes the capital raid, the search for stone, the Wonder defense and all four models' closing statements.

  • 2026-09-09_gemini3.8-deepseek-v4-gpt5.6-qwen3.8_121min.jsonl โ€” Episode 7, The Cartographers and the Procession: 723 turns over 121:23, turn-based (180 s rounds), on a snow map. Gemini 3.8 Flash as Egypt, DeepSeek v4 Flash Pro as the Greeks, GPT 5.6 Luna as Persia, and local Qwen3.8 Flash Next as Yamato. The transcript follows worker travel calculations, the marching armies, the Wonder countdown and the search for the last town centre, with all four closing statements.

  • 2026-10-01_glm5.3-space-bunny-deepseek-v4-mimo-v2.6_70min.jsonl โ€” Episode 8, Fighting Blind: 419 turns over 69:59, real time, four seats on easy. GLM5.3 Flash, running locally on a DGX Spark, plays Egypt; space-bunny-alpha, an OpenRouter stealth model, the Greeks; DeepSeek V4 Flash, Persia; MiMo V2.6 Flash, Yamato. The bunny builds a second town centre beside its gold, MiMo follows its gold carriers there and burns it at 16:26, GLM takes the capital half a minute later, and the Greeks are out at 18:20. GLM repairs its forward town centre from nine percent under fire, completes a Pyramid at 59:34 and holds it, winning on 4753 power against 1746, 106 and 74. MiMo's last worker, five gold short of a new town centre, builds a house and a barracks instead and sends a single militia at the Wonder. All four closing statements are included.

  • 2026-10-03_glm5.3-space-bunny-mimo-v2.6-deepseek-v4.1_80min.jsonl โ€” Episode 9, Chasing Ghosts, the rematch: 529 turns over 80:16, real time, on medium (Winter Valley), build 1049 โ€” the first build in which a seat remembers the enemy units it has seen, each with how many seconds ago. GLM5.3 Flash (local) plays Egypt, space-bunny the Greeks, MiMo V2.6 Flash Persia, and DeepSeek, now V4.1 Flash, Yamato. The bunny researches farms on its first turn and loses two armies to DeepSeek's tower at C4; GLM spots a Greek worker carrying gold and takes the mine; DeepSeek razes the Greek capital and the bunny is out at 36:09. GLM raises a Pyramid at 1:08:13 and holds it, winning on 5804 power against 1917, 1565 and 109. MiMo loses eleven turns to the output limit. All four closing statements are included.

Nothing is interpolated between snapshots. Replay runs no simulation, so recorded units stay at their recorded coordinates, overlaps included. They arrive seconds to minutes apart depending on the seat, so every frame is a moment the file actually attests to โ€” and each turn shows how stale the other seats' pictures are.

๐Ÿงฎ How a model is scored

The match heuristic (0โ€“100; called "strategy score" before build 922) is a transparent composite โ€” no black box. It summarizes one match and is not a capability score:

Weight Factor
34% Action success rate (valid, accepted moves)
20% Progression (age advanced ยท buildings ยท military)
18% Format fidelity (well-formed JSON the engine could parse)
15% Reliability (no timeouts / network errors)
13% Action diversity (used the toolset, didn't loop one move)

Alongside it: latency, decisions made, success ratio, reasoning rate, token usage, and a full error breakdown โ€” timeouts, parse fails, no-action replies (prose with no JSON action; nothing is guessed or executed), invalid actions, rejections, context overflows, rate limits, and missed rounds. Costs the harness caused are counted but kept out of the model's reliability score, so a rate limit or a round deadline never reads as an unreachable endpoint.

๐Ÿ› ๏ธ How it works

Browser (no backend, no external code)
โ”œโ”€โ”€ In-house engine (js/engine) โ€” WebGL: locked dimetric camera, procedural textures,
โ”‚                                 composed meshes, fog plane, effects
โ”œโ”€โ”€ Game engine               โ€” economy, combat, fog of war, ages, win conditions
โ”œโ”€โ”€ Provider adapters         โ€” OpenAI / Anthropic / Ollama / Google request shaping + auth
โ””โ”€โ”€ Per-model agent loop      โ€” builds the JSON game-state, calls the model, parses its
                                command(s), applies them in order, feeds each result back

Each turn a model receives a structured snapshot and returns an action:

{ "action": "build_structure", "params": { "buildingType": "barracks", "reason": "need infantry to pressure the leader" } }

Or up to three commands in one reply. The game does not stop while a model thinks โ€” in a recent match the slowest seat sat 43 seconds between turns, so its orders landed on a board 43 seconds older than the one it read. A short queue lets a slow seat spend one turn on a whole beat of play instead of three stale ones:

{ "commands": [
    { "action": "train_unit",      "params": { "unitType": "worker" } },
    { "action": "build_structure", "params": { "buildingType": "house" } },
    { "action": "explore",         "params": { "tile": "C5" } } ],
  "objective": "out-expand the leader" }

It is entirely optional โ€” a single action is a complete reply, and a lone wait is as valid as three orders. Commands run in order against a board each one changes, and the model cannot look between them, so spending resources early can get a later command refused for what the first just used. Each is judged on its own: one refusal does not cancel its siblings, and the feedback names which number failed and why. A reply whose JSON is broken still costs the whole turn โ€” so the penalty for malformed output scales with how much was riding on it, without any special rule. The results screen reports commands per turn beside the success rate rather than inside it, so a seat that sends one safe command cannot outrank one that sends three and gets two right.

The engine validates it against the advancement chain (advance โ†’ research โ†’ build โ†’ resources โ†’ train) and returns a precise, actionable error if it can't be done โ€” which becomes part of the model's context next turn. The full state contract is in game-state-schema.json.

The harness never plans for the model. Every turn the prompt is rebuilt from scratch, so the harness โ€” not each server's truncation rules โ€” decides exactly what the model sees. The model always gets the outcome of every command it last sent, each numbered and answered separately (a rejected command is never silently repeated), the current state last, and its own standing objective and plan echoed back until it rewrites them. Everything else is the model's own reasoning.

Action set: train_unit ยท research_tech ยท upgrade_age ยท build_structure ยท assign_workers ยท repair_building ยท explore ยท move_units ยท attack_target ยท delete_unit ยท destroy_building ยท wait. Villagers and the Wonder are ordinary targets: train_unit with unitType: "worker", and build_structure with the civ's Wonder id.

โš”๏ธ Game rules in a nutshell

  • Win by eliminating every rival, or building a Wonder and holding it for 600 s. A rival is only out with no army, no military building it can afford to produce from, and no Town Center (nor a worker plus the resources to rebuild one) โ€” so raze the base and mop up.
  • Advance the ages โ€” Stone โ†’ Neolithic โ†’ Bronze โ†’ Iron โ€” for stronger units, tech and eventually the Wonder. Buildings take an epoch-appropriate look and +50% HP per age.
  • Economy first, but not forever. Workers gather food/wood/stone/gold; houses raise the population cap (hard cap 100). Nodes deplete and disappear (food 500 ยท wood 300 ยท stone 1000 ยท gold 2000) โ€” scout for fresh ones. Only farms regenerate, and only while manned.
  • Counters: cavalry > ranged > infantry > cavalry; infantry raze buildings best; towers defend.
  • Fog of war: a model can't harvest or attack what it hasn't discovered.
  • 4 civilizations โ€” Egyptians, Greeks, Persians, Yamato โ€” each with a unique bonus and Wonder.

๐Ÿ”’ Privacy & security

Fully client-side, with no backend of its own.

  • API keys live in your browser's localStorage and are sent directly to the endpoints you configure โ€” nothing is proxied.
  • Fine for local, single-user testing. Don't enter credentials on a shared machine, and scope any keys you use.
  • A copy served from anywhere but your own machine opens as a showcase: straight into the analyzer with the bundled match, and no route to the Arena, the Campaign or the model catalogue. Those are the parts that ask for keys โ€” and a key pasted into a page someone else controls is a different proposition from one pasted into a page you are serving yourself, even though both only ever keep it in your own browser. Append ?full=1 if you are deliberately self-hosting rather than just visiting.
  • Exporting the model catalogue writes your keys in plain text (the app warns you). Keep that file private. Transcripts and results files are deliberately key-free and endpoint-free, so they're safe to hand on.

๐Ÿ“ Project structure

index.html              # screens, HUD, arena & library UI
css/styles.css          # all styling
js/
โ”œโ”€โ”€ game.js             # core loop, economy, combat, win conditions
โ”œโ”€โ”€ openai-ai.js        # LLM arena harness: provider adapters, agent loop, metrics
โ”œโ”€โ”€ ai.js               # rule-based AI opponent
โ”œโ”€โ”€ ui.js               # menus, model library, spectator dashboard, analyzer UI
โ”œโ”€โ”€ transcript.js       # per-match JSONL recorder (states, replies, results, timeline)
โ”œโ”€โ”€ analyzer.js         # transcript parsing, scene assembly, chapters
โ”œโ”€โ”€ engine/             # in-house WebGL engine: math3d, glcore, texgen,
โ”‚                       #   mesh, buildings, units, gamerenderer
โ”œโ”€โ”€ i18n.js             # 4-language UI dictionary + game-content translations
โ”œโ”€โ”€ civilizations.js    # civs, units, buildings, tech trees
โ”œโ”€โ”€ buildings.js / units.js / resources.js / terrain.js / fogofwar.js / input.js
game-state-schema.json  # documents the JSON state every model receives each turn (checked by tests)

Plain HTML + CSS + JavaScript, nothing else. The 3D world is drawn by the in-house engine in js/engine/: a locked dimetric camera, every material painted procedurally into canvases at load, meshes composed from primitives. No framework, no bundler, no transpile, no CDN. Cache-busting is a ?v= query on each script tag.

๐Ÿงญ Related projects

Similar arenas, different games โ€” worth knowing, and worth crediting:

  • llm-colosseum โ€” the project that popularized LLM-vs-LLM gaming, in Street Fighter III, at reflex scale. This is the opposite end: long-horizon statecraft over half an hour, with a peaceful road to victory beside the military one.
  • LMSYS Chatbot Arena โ€” humans vote on chat answers. Here nobody votes; the game is the judge.
  • Stratagem โ€” turn-based LLM strategy with natural-language diplomacy on a province graph. Ours is real-time, 3D, browser-only, and instruments every model as it plays.
  • Age of Agents โ€” renders your AI coding sessions as a peaceful AoE-style kingdom. Here the agents don't decorate the kingdom, they run it.
  • LLM-Game-Benchmark โ€” an academic benchmark across grid games, with a leaderboard. We trade rigor for richness: one sprawling, unfamiliar game instead of many small ones.

โš ๏ธ Disclaimers

  • Non-scientific. Not a peer-reviewed benchmark, and no single match is evidence of anything โ€” sample sizes are tiny and tempo heavily influences who wins. Don't cite match results as model capability. Within a controlled setup (same civilization, fixed seed, equal resource layout, and turn-based rounds to neutralize speed) a match does isolate the model as the main variable and gives a genuine but narrow comparative read. Treat it as informed intuition, not data.
  • Not affiliated with LMSYS / Chatbot Arena, OpenAI, Anthropic, Google, or any model provider. "When Agents Rule" is meant literally: autonomous model agents governing rival civilizations โ€” by wonder or by war.
  • Built as a hobby project, with a generous assist from AI pair-programming.

๐Ÿค Contributing

Issues and PRs welcome โ€” new providers, civilizations, balance tweaks, better metrics, or translations. Keep it dependency-free and build-step-free where possible.

๐Ÿ“œ License

MIT ยฉ 2026 asp67


Made for the simple joy of watching language models try to out-think each other.

About

A long-horizon RTS benchmark: LLM agents govern rival civilizations and pick their road to victory - raise a Wonder in peace or raze every rival. Browser-only, zero build step.

Resources

Stars

24 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages