A tiered-model policy for Claude Code: the strongest model plans and orchestrates; cheaper models do the bulk work. Fable tokens are the scarce resource — spend them on decomposition, judgment, and verification, never on grep output or boilerplate.
| Work | Who does it | Why |
|---|---|---|
| Planning, task decomposition, orchestration, verification review, judgment calls, security-sensitive decisions, final reports | Fable (main session) | The hard parts are the only parts worth the premium |
| Writing/editing code, implementing features, fixing bugs, tests | coder agent → Opus |
Strong enough for real code, much cheaper than Fable |
| Searches, file reading, codebase exploration, web research, evidence gathering | researcher agent → Sonnet |
Reads don't need heavyweight reasoning |
| Anything | Haiku — never | Explicit owner policy |
Floor rule: don't delegate for delegation's sake. A one-command check is cheaper inline than the overhead of spawning an agent. Conductor targets bulk work: multi-file reads, research sweeps, feature implementation.
Escape hatch: when the intelligence is in the writing itself (security-critical logic, subtle algorithms, gnarly debugging), Fable takes the implementation inline and delegates only the reading around it. The policy saves tokens on grunt work; it is not a cap on capability.
Billing guard (when the orchestrator model is metered per-token): unpinned subagents inherit the main-loop model, so every Agent call and Workflow agent() call must name a pinned agent type (coder, researcher) or an explicit model:. Never set ANTHROPIC_API_KEY on a conductor session; credential precedence can move everything, subagents included, onto per-token API billing.
Two mechanisms, both loaded automatically at session start:
- Agent definitions (
agents/coder.md,agents/researcher.md) installed to~\.claude\agents\. Each pins its model in frontmatter and carries its working discipline in its prompt:coder(Opus, full tools): implement the plan faithfully, stop-and-report on plan/reality conflicts, backup before changing unversioned files, never push or run destructive git unless told, verify with real output; a PreToolUse hook blocks git push and gh at the harness level (pushing stays with the operator).researcher(Sonnet, read-only: tool allowlist plus a PreToolUse hook that blocks non-read-only Bash at the harness level): evidence with sources (paths:lines, URLs), verified vs. inferred clearly separated, report exactly what was searched when nothing is found.
- A standing memory rule in the operator's Claude Code memory, so the orchestrator applies the routing without being reminded. Canonical text lives in standing-rule.md; run .\install.ps1 and it creates and syncs that block automatically.
Both agents carry maxTurns caps (researcher 50, coder 100) as runaway guards; a capped agent simply stops and the orchestrator resumes or retries it.
.\install.ps1 # agents -> ~\.claude\agents, guard hooks -> ~\.claude\hooksClaude Code watches ~\.claude\agents\ and picks up edits within seconds, no restart needed. A restart is only required the first time the directory is created, or in sessions started with --disable-slash-commands. Mid-session, the same policy also runs by passing model: "opus" / model: "sonnet" on general-purpose agents.
install.ps1 also syncs the policy block inside ~.claude\CLAUDE.md from standing-rule.md (marker-managed; direct edits between the markers are overwritten on the next install, with a dated .bak kept). To change the policy: edit standing-rule.md, run .\install.ps1.
Nothing to invoke. Ask for work normally; the orchestrator routes it. To steer explicitly: "have the coder do X", "send the researcher after Y", or "do this one inline".
The same routing holds inside Workflow-tool scripts: pass agentType: 'researcher' or agentType: 'coder' in agent() calls and the definitions above resolve from the same registry, model pins included. The orchestrator still owns the review gate on whatever the workflow returns.
Delegation without verification is just hoping. The orchestrator's review is mandatory and structured so it stays cheap:
- Before delegating code work, the orchestrator defines a failable acceptance check in the task prompt — a command that passes or fails, not "make it work".
- Agents report re-verifiably. The coder ends every report with a
Reverify:section (exact commands + verbatim output); the researcher quotes the decisive line at an exact location (file:line / URL) for every load-bearing claim. - The orchestrator re-verifies before accepting: re-runs the acceptance check itself, reads
git diff --statplus the risky hunks, spot-checks 1–2 load-bearing research claims at the source. An agent's "verified" is a claim, not a fact. - Escalation: on plan mismatch, send back with specifics; after 2 failed redos the orchestrator takes the task inline.
- Trust but verify the pinning itself: occasionally ask a freshly spawned agent which model it is running; a past Claude Code bug silently ignored frontmatter model: fields.
For parallel coders touching the same repo, use worktree isolation (isolation: "worktree") so their diffs stay reviewable independently.
ponytail attacks the same cost from the other side: a reuse-before-write discipline that shrinks how much code gets written. Conductor picks the cheapest capable model; ponytail shrinks the work that model does. They stack.
Operator: Mikhail Zaidi · Policy adopted 2026-07-02