An independent, single-task A/B experiment on whether Codex's instruction harness can hold GPT-5.6 Sol back on coding work.
The claim is directionally supported, not universally proved.
On one matched Codex CLI task, replacing the built-in base instructions with a 156-word coding prompt preserved the observed deterministic score and materially reduced latency, tokens, and agent rounds:
| Metric | Stock base | Slim base | Change |
|---|---|---|---|
| Deterministic evaluator | 18/18 | 18/18 | observed parity |
| Elapsed time | 1,424,239 ms | 897,851 ms | 37.0% faster |
| Total tokens | 2,268,917 | 1,659,272 | 26.9% fewer |
| Output tokens | 34,322 | 23,602 | 31.2% fewer |
| Reasoning tokens | 11,769 | 6,677 | 43.3% fewer |
| LLM rounds | 34 | 28 | 17.6% fewer |
| First output | 5,205 ms | 1,633 ms | 68.6% faster |
| API-equivalent cost estimate | $2.8818 | $2.0101 | 30.3% lower |
This is one stochastic pair on one CLI workload. It does not establish causality, cross-task generalization, or Desktop parity.
Both arms used:
- GPT-5.6 Sol
- high reasoning
- the default service tier
- the same benchmark task and task prompt
- the same developer/user guidance, tools, and full-access execution boundary
The treatment added only model_instructions_file, replacing the built-in base with prompts/measured-slim-base.md. The observed built-in base was 2,826 words / 17,767 bytes; the measured treatment was 156 words / 1,095 bytes.
The complete design and caveats are in docs/experiment.md. Machine-readable numbers live in evidence/ab-results.json.
The fastest measured option is not the safest global default. model_instructions_file replaces Codex's built-in base rather than supplementing it, so applying a slim prompt globally can remove guidance needed by non-coding tools and workflows.
The four-review-cycle conclusion is:
- Keep the stock Codex base globally.
- Use the compact global guidance in
examples/lean-global-AGENTS.md. - Remove duplicative global
developer_instructions. - Use Sol with high reasoning, high Plan effort, and low verbosity.
- Select the hardened slim prompt explicitly through a CLI profile for trusted coding tasks.
- Use a trusted repository override only when automatic Desktop + CLI activation is worth the broader within-repository blast radius.
The recommended successor, prompts/hardened-slim-base.md, restores instruction hierarchy, untrusted-content handling, permission fidelity, and blocked-check reporting. It is 128 words / 903 bytes. It is a reviewed successor—not the exact prompt measured in the paid A/B.
See docs/setup.md for exact configuration and rollback steps.
The included config preserves the experiment operator's fixed settings:
approval_policy = "never"
sandbox_mode = "danger-full-access"Those settings remove Codex's normal sandbox and approval pauses. They were a user requirement and a constant across the comparison; they are not required for prompt slimming and are not a general security recommendation. Understand the consequence before copying them.
The verifier uses only the Python standard library and does not call a model:
python3 scripts/verify.pyIt checks the evidence arithmetic, prompt sizes, score parity, TOML syntax, intended active keys, and the fixed permission values.
docs/experiment.md: protocol, results, limitations, and four-cycle reviewdocs/setup.md: global, CLI-profile, and Desktop/project setupevidence/ab-results.json: sanitized measurement recordprompts/measured-slim-base.md: exact 156-word treatment used in B1prompts/hardened-slim-base.md: recommended 128-word successorexamples/recommended-config.toml: recommended user-level coreexamples/slim.config.toml: opt-in CLI profile templateexamples/lean-global-AGENTS.md: compact durable personal guidance
Raw Codex rollouts are intentionally excluded. They contain built-in instruction text, tool schemas, absolute local paths, and unrelated environment context that are unnecessary to verify the aggregate result. This repository publishes the exact treatment prompt, sanitized metrics, benchmark provenance, deterministic score, and a verifier.
This is an independent experiment, not an OpenAI publication or product recommendation.