Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

codex-harness

An independent, single-task A/B experiment on whether Codex's instruction harness can hold GPT-5.6 Sol back on coding work.

Bottom line

The claim is directionally supported, not universally proved.

On one matched Codex CLI task, replacing the built-in base instructions with a 156-word coding prompt preserved the observed deterministic score and materially reduced latency, tokens, and agent rounds:

Metric Stock base Slim base Change
Deterministic evaluator 18/18 18/18 observed parity
Elapsed time 1,424,239 ms 897,851 ms 37.0% faster
Total tokens 2,268,917 1,659,272 26.9% fewer
Output tokens 34,322 23,602 31.2% fewer
Reasoning tokens 11,769 6,677 43.3% fewer
LLM rounds 34 28 17.6% fewer
First output 5,205 ms 1,633 ms 68.6% faster
API-equivalent cost estimate $2.8818 $2.0101 30.3% lower

This is one stochastic pair on one CLI workload. It does not establish causality, cross-task generalization, or Desktop parity.

What was isolated

Both arms used:

  • GPT-5.6 Sol
  • high reasoning
  • the default service tier
  • the same benchmark task and task prompt
  • the same developer/user guidance, tools, and full-access execution boundary

The treatment added only model_instructions_file, replacing the built-in base with prompts/measured-slim-base.md. The observed built-in base was 2,826 words / 17,767 bytes; the measured treatment was 156 words / 1,095 bytes.

The complete design and caveats are in docs/experiment.md. Machine-readable numbers live in evidence/ab-results.json.

Recommended practical setup

The fastest measured option is not the safest global default. model_instructions_file replaces Codex's built-in base rather than supplementing it, so applying a slim prompt globally can remove guidance needed by non-coding tools and workflows.

The four-review-cycle conclusion is:

  1. Keep the stock Codex base globally.
  2. Use the compact global guidance in examples/lean-global-AGENTS.md.
  3. Remove duplicative global developer_instructions.
  4. Use Sol with high reasoning, high Plan effort, and low verbosity.
  5. Select the hardened slim prompt explicitly through a CLI profile for trusted coding tasks.
  6. Use a trusted repository override only when automatic Desktop + CLI activation is worth the broader within-repository blast radius.

The recommended successor, prompts/hardened-slim-base.md, restores instruction hierarchy, untrusted-content handling, permission fidelity, and blocked-check reporting. It is 128 words / 903 bytes. It is a reviewed successor—not the exact prompt measured in the paid A/B.

See docs/setup.md for exact configuration and rollback steps.

Important permission warning

The included config preserves the experiment operator's fixed settings:

approval_policy = "never"
sandbox_mode = "danger-full-access"

Those settings remove Codex's normal sandbox and approval pauses. They were a user requirement and a constant across the comparison; they are not required for prompt slimming and are not a general security recommendation. Understand the consequence before copying them.

Verify this repository

The verifier uses only the Python standard library and does not call a model:

python3 scripts/verify.py

It checks the evidence arithmetic, prompt sizes, score parity, TOML syntax, intended active keys, and the fixed permission values.

Repository map

Raw-artifact policy

Raw Codex rollouts are intentionally excluded. They contain built-in instruction text, tool schemas, absolute local paths, and unrelated environment context that are unnecessary to verify the aggregate result. This repository publishes the exact treatment prompt, sanitized metrics, benchmark provenance, deterministic score, and a verifier.

This is an independent experiment, not an OpenAI publication or product recommendation.

About

A/B evidence and practical configuration for GPT-5.6 Sol in the Codex harness

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages