A terminal coding agent for OpenRouter models where a command you control determines whether completion is recorded as verified, failed, or not verified. Give it a bounded task and an acceptance command; the agent works, and before its answer is accepted the command runs. A failing check earns exactly one additional model response, then it stops with the evidence.
Coding agents are everywhere — opencode, pi, Claude Code. They all run tools, keep sessions, and can run your tests. The difference here is not more models, lower cost, or more polish: completion is a checked state, recorded honestly, and the same engine doubles as an experiment harness — contained, audited, treatment-separated — so claims about whether the completion policy helps are measured, not asserted.
It is for developers who want to try several OpenRouter models on bounded repository tasks while keeping explicit acceptance evidence — a command that must pass before the agent's answer is accepted.
What it does:
- bounded task contracts (
--task//task) with a user-owned acceptance command (--verify-command//verify//check); - honest completion states: verified / failed / not verified;
- exactly one additional model response on a failed check, then it stops;
- tool actions (
run_bash,list_dir/search_text/read_file/write_file/edit_file,discoverweb search/navigate); - interactive permission gating (
allow/deny/ask); - session persistence and context visibility with honest cache accounting (provider cache counters are reported only when the provider exposes them);
- diff review (
/diff) and undo that only claims what the tool itself tracked; - an experiment harness on the same engine: preregistered runs, a verifier-assisted policy kept separate from ordinary results, a campaign audit, and contained real-model execution.
pip install openrouter-agent-cli # from PyPI
pipx install openrouter-agent-cli # or isolated CLI installOr from a source checkout:
cd openrouter-agent-cli
python3 -m venv .venv
source .venv/bin/activate
pip install -e . # core includes pyunbrowser (real web discovery)
pip install -e ".[viz]" # + png gantt (matplotlib)
pip install -e ".[full]" # all extras (openai, dotenv, viz)Note: real web discovery depends on pyunbrowser, which currently ships
Linux/macOS wheels — on Windows the rest of the CLI works, but discover
may be unavailable.
export OPENROUTER_API_KEY=sk-or-...
openrouter-agentOr without installation:
export OPENROUTER_API_KEY=sk-or-...
python -m openrouter_agent_cli.cli--prompt (short -p) lets another process run the CLI with a single user message, emit only the assistant reply to stdout, and exit immediately. Operation logs, tool call summaries, and permission notices are written to stderr, and tool calls are automatically denied unless you disable tools with --no-tools.
Example:
openrouter-agent --prompt "Explain tail recursion" --no-toolsFor a bounded coding task, give the session an objective and a developer-owned acceptance command:
openrouter-agent --workdir ./my-repo \
--task "Fix the failing login test" \
--verify-command "pytest tests/test_auth.py"The command runs at the completion boundary (once initially, and once more
after the single permitted repair response when the first check fails). A
failed command earns exactly one additional model response; a timeout or
execution error is reported as not_verified rather than treated as success.
The workflow can be exercised without credentials or network access with
openrouter-agent-self-test (or openrouter-agent --self-test).
openrouter-agent \
--model nvidia/nemotron-3.5-lightning:free \
--session-id my-session \
--workdir ~/Projects \
--max-turns 24 \
--max-history-messages 60 \
--command-timeout 30 \
--discovery auto \
--max-concurrency 5 \
--env-file .env
# web discovery modes: auto (real if pyunbrowser installed else error), mock (synthetic example.com), real (requires pyunbrowser), off
openrouter-agent --discovery mock --allow-discovery --prompt "what is unbrowser?"
openrouter-agent --discovery real --allow-discovery # needs BRAVE_API_KEY for searchEnv auto-load: .env in cwd and repo root is loaded via python-dotenv (or allowlisted fallback: OPENROUTER_*, BRAVE_API_KEY, UNBROWSER_BINARY). --env-file overrides.
Disable tools:
openrouter-agent --no-toolsDebug logging:
openrouter-agent --debug # timestamps + idle instrumentation on stderr/help/exit/new [id](fresh session — history reset, isolates crypto contamination)/model [id]/usage/status/task [description]/verify [command](useoffto clear it)/check(run the acceptance command immediately)/diff [path](review working-tree changes;--statfor a summary)/context [n]/compact [--preview]/undo(last file-tool edit or compaction — shell changes are never rolled back)/clear(same id)/tools/tools on|off/allow <tool|*>(cached across sessions →~/.openrouter-agent-cli/policy.json)/deny <tool|*>(cached)/unallow <tool|*>/undeny <tool|*>/cwd [path]/discovery [auto|mock|real|off]/concurrency [n](1-16, cap for paralleldiscover)/inspect <call-id>/sessions/resume <id>/policy/export [path]
- history is saved in
~/.openrouter-agent-cli/sessions/<session_id>.json /usageshows a rough token estimate and process-only API counters/usagealso shows the stable context-prefix estimate and whether the provider exposed explicit cache counters;not observableis a real state, not a zero-cache claim/statusshows the active model, session, cwd, tool policy, context estimate, and current lifecycle state/compactforces summarization;/compact --previewshows what would be summarized- automatic compaction triggers around 12k estimated tokens;
--max-history-messagesalso controls persisted-history trimming when both message and token thresholds are high /undorestores the last compaction snapshot or the last file write/edit; failed summarization does not replace the original history
run_bashexecutes shell commands on your machine in--workdir(cwd only, not jailed — unlikelist_dir/search_text/read_file/write_file/edit_filewhich enforce workdir jail). It returns structured JSON, runs in a separate process group, has bounded output, kills descendants on timeout, and does not pass OpenRouter/Brave API keys to the child environment.discoverfetches web content (httpsonly, private/loopback/link-local/metadata blocked after hostname resolution and on observed final URLs; browser subrequests still depend on the underlying browser boundary). All web content is untrusted — model may be prompt-injected via page content; keep allow/deny gated.BRAVE_API_KEYused fordiscover(search),UNBROWSER_BINARYcan select browser binary — both allowlisted from.env;autono longer silently mocks (requires--discovery mockfor syntheticexample.com/mock).- independent
discoverbatches use stateless clients for real concurrency (max_concurrency, default 5); stateful calls outside a batch reuse one session and are queued.run_bash/file ops serialize in mixed batches. Eachdiscoverhas a configurable1-120stimeout (30s default). - Model outputs (tool calls) are reflected literally in
run_bash+ URLs indiscover, so treat every allowed tool call as untrusted input and keep the allow/deny policy enforced unless you deliberately want to run everything. - default policy is
askfor every tool call; approvals can be once, batch, turn, session, or explicitly persistent - use
/deny *for a fully no-tools session;/allow discoveris persistent while--allow-discoveryis scoped to the current process - default model is free-tier (
nvidia/nemotron-3.5-lightning:free); override with--modelorOPENROUTER_MODEL
When tools are enabled, each OpenRouter request includes these tool definitions (filtered when --discovery off):
[
{
"type": "function",
"function": {
"name": "run_bash",
"description": "Run a host shell command; returns structured JSON (ok/exit_code/stdout/stderr/timed_out).",
"parameters": {
"type": "object",
"properties": {
"command": {"type": "string"},
"timeout_seconds": {"type": "integer", "default": 30}
},
"required": ["command"]
}
}
},
{
"type": "function",
"function": {
"name": "list_dir",
"description": "List directory entries inside the workdir jail.",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string"},
"max_entries": {"type": "integer"}
},
"required": []
}
}
},
{
"type": "function",
"function": {
"name": "search_text",
"description": "Literal text search under a workdir path.",
"parameters": {
"type": "object",
"properties": {
"pattern": {"type": "string"},
"path": {"type": "string"},
"max_matches": {"type": "integer"},
"max_file_bytes": {"type": "integer"}
},
"required": ["pattern"]
}
}
},
{
"type": "function",
"function": {
"name": "read_file",
"description": "Read file (workdir-jail) with line range, max_lines paging, and next_cursor.",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string"},
"start_line": {"type": "integer"},
"end_line": {"type": "integer"},
"max_lines": {"type": "integer"},
"max_bytes": {"type": "integer"},
"cursor": {"type": "string"}
},
"required": ["path"]
}
}
},
{
"type": "function",
"function": {
"name": "write_file",
"description": "Atomic write (workdir-jail) with optional expected_sha256 and dry_run.",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string"},
"content": {"type": "string"},
"expected_sha256": {"type": "string"},
"dry_run": {"type": "boolean"}
},
"required": ["path", "content"]
}
}
},
{
"type": "function",
"function": {
"name": "edit_file",
"description": "Atomic unique string replace (workdir-jail) with optional expected_sha256 and dry_run.",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string"},
"old_string": {"type": "string"},
"new_string": {"type": "string"},
"expected_sha256": {"type": "string"},
"dry_run": {"type": "boolean"}
},
"required": ["path", "old_string", "new_string"]
}
}
},
{
"type": "function",
"function": {
"name": "discover",
"description": "Web discovery via pyunbrowser. Use search with query or navigate with an https URL; independent calls may run stateless and concurrently.",
"parameters": {
"type": "object",
"properties": {
"kind": {"type": "string", "enum": ["search", "navigate"]},
"query": {"type": "string"},
"url": {"type": "string"},
"goal": {"type": "string"},
"timeout_seconds": {"type": "integer"}
},
"required": ["kind", "goal"]
}
}
}
]Request body shape sent to OpenRouter (simplified):
{
"model": "nvidia/nemotron-3.5-lightning:free",
"messages": [...],
"temperature": 0,
"max_tokens": 4096,
"tools": [...],
"tool_choice": "auto"
}If tools are disabled (--no-tools or /tools off), the request sets:
{
"tool_choice": "none"
}Execution flow per user turn:
- Model returns
tool_callsin assistant message (may be 1 or parallel batch whenparallel_tool_calls=True). - CLI decodes
function.argumentsJSON into a dict. - Permission policy is applied per tool:
denylist blocks immediately.allowlist runs immediately.- otherwise prompt user with once/batch/turn/session/persistent scopes.
- For
run_bash, CLI executes:asyncio.create_subprocess_shell(command, cwd=<workdir>, stdout=PIPE, stderr=PIPE)in a separate process group with sensitive API keys removed from the child environment- waits with
asyncio.wait_for(..., timeout_seconds) - kills process on timeout
- concurrent batches: independent pure
discovercalls use stateless clients and run concurrent (max_concurrency, semaphore); anyrun_bash/file ops in batch force serialization
- For
discover, CLI executes (blocking →to_thread, configurable1-120stimeout, 30s default):kind=search:SmartClient.search(query, engine=brave)(needsBRAVE_API_KEY)kind=navigate:SmartClient.navigate_auto(url, goal=goal)(LLM extraction)mockmode: syntheticexample.com/mockhits (explicit--discovery mockonly)
- CLI records each call with a stable ID, formats a bounded result, and appends exactly one
role:toolmessage pertool_call_id; structured discovery truncation remains valid JSON. Use/inspect <call-id>to inspect a recent full result.
Example tool call from model:
{
"id": "call_123",
"type": "function",
"function": {
"name": "run_bash",
"arguments": "{\"command\":\"ls -la\",\"timeout_seconds\":30}"
}
}Example tool result message added by CLI:
{
"role": "tool",
"tool_call_id": "call_123",
"content": "total 64\n-rw-r--r-- ..."
}Note: despite the name run_bash, execution uses create_subprocess_shell (system shell), not an explicit bash binary unless the command itself invokes bash.
This repo includes a small harness for comparing system prompts:
- script:
scripts/ab_test_system_prompts.py - prompt variants:
prompts/system_prompt_control.mdprompts/system_prompt_agentic_v1.md
- sample tasks:
ab_tests/tasks_sample.txt
Run prompt-only comparison (no tools):
export OPENROUTER_API_KEY=sk-or-...
python scripts/ab_test_system_prompts.py \
--tool-mode none \
--model nvidia/nemotron-3.5-lightning:freeRun with tool execution enabled (use cautiously):
export OPENROUTER_API_KEY=sk-or-...
python scripts/ab_test_system_prompts.py \
--tool-mode execute \
--workdir "$(pwd)" \
--model nvidia/nemotron-3.5-lightning:freeArtifacts are written to ab_tests/results/<timestamp>/:
results.jsonfull transcripts and metadatasummary.csvflat comparison tablesummary.mdquick markdown summary
Run a harder repeated suite (2 prompts x 6 tasks x 3 repeats):
export OPENROUTER_API_KEY=sk-or-...
python scripts/ab_test_system_prompts.py \
--tool-mode execute \
--tasks-file ab_tests/tasks_hard_suite_v1.txt \
--repeats 3 \
--max-turns 3 \
--max-tokens 1000 \
--request-timeout 40 \
--command-timeout 20 \
--workdir "$(pwd)" \
--model nvidia/nemotron-3.5-lightning:free \
--output-dir ab_tests/results/hard_suite_v1_r3Evaluate quality and groundedness from a run:
export OPENROUTER_API_KEY=sk-or-...
python scripts/evaluate_ab_results.py \
--results ab_tests/results/hard_suite_v1_r3/results.json \
--judge-model nvidia/nemotron-3.5-lightning:free \
--output-dir ab_tests/results/hard_suite_v1_r3/evalEvaluator artifacts:
evaluation.jsonper-case raw evaluation detailsevaluation.csvtabular scoresleaderboard.mdaggregated per-prompt ranking
The repository includes a small evaluation workflow that runs the real CLI engine, real tool layer, and real task verifiers. The mock profile is provider-offline: it makes no provider calls (mock-generated commands are host shell commands, so they are not a network guarantee). Suite manifests and mock scripts are trusted operator-provided input: their commands execute on the host inside a disposable working directory, so review them like test code before running.
After installing the source checkout with pip install -e .:
openrouter-agent-eval \
--suite eval_suites/coding_smoke_v1/suite.json \
--profile worker=eval_suites/mock_worker.jsonThe command writes append-only attempt records under .agent-eval/runs/ and
prints a paired report with pass counts, shared-task outcomes, token/latency
accounting, uncertainty intervals, and the suite-specific leaderboard. To
re-render an existing run without executing attempts:
openrouter-agent-eval \
--suite eval_suites/coding_smoke_v1/suite.json \
--eval-dir .agent-eval \
--report-onlyTo exercise both ordinary and verifier-assisted treatments fully offline, provide the mock profile twice and mark one profile as assisted:
openrouter-agent-eval \
--suite eval_suites/coding_smoke_v1/suite.json \
--profile baseline=eval_suites/mock_worker.json \
--profile assisted=eval_suites/mock_worker.json \
--assisted-profile assisted \
--repeats 2The assisted rows are reported separately and are excluded from the ordinary model leaderboard. Prompt-file profiles make real OpenRouter calls and must only be used with an explicitly approved model budget and the documented execution-containment settings.
- benchmark findings:
docs/AB_FINDINGS_2026-02-21.md - public release checklist:
docs/PUBLIC_RELEASE_CHECKLIST.md - security policy:
SECURITY.md - env template:
.env.example