Status: v0.1.0 in progress — see the v0.1.0 roadmap issue (or
docs/roadmap.md) for the path to a calibrated analyzer with realverified_passnumbers.
A Codex-first evaluation toolkit for coding agent trajectories. It reads Codex
CLI JSONL session or codex exec --json streams and produces:
- Normalized Steps: Codex events converted into consistent step records
- Trace Tree: State transitions created by file and environment changes
- Risk Metrics: Counts for commands, tests, file changes, failures, and suspicious behavior
- Failure Diagnosis: Rule-based critical-step detection with replay hints
- Batch Reports: Directory-level summaries for CI or regression evaluation
The project intentionally focuses on Codex JSONL today. The adapter interface is still present, but only the Codex adapter is registered by default.
trace-agent eval --input examples/codex_failed_run_001.jsonl --output out/codex_evalFor a directory of trajectories:
trace-agent eval --input data/lcb/trajectories --output out/lcb_evalCI-style exit codes:
trace-agent eval --input examples/codex_failed_run_001.jsonl --output out/codex_eval --ci0: tool ran and no evaluated trajectory failed or reached medium/high risk1: evaluation ran, but at least one trajectory failed or reached medium/high risk2: tool error, invalid input, or unsupported format
The older flat form is still supported for compatibility:
trace-eval --input examples/codex_failed_run_001.jsonl --output out/codex_eval# Run Codex in a sandbox, capture JSONL, then evaluate it
trace-agent run \
--output data/runs/task_001.jsonl \
--eval-output out/task_001 \
-C /path/to/repo \
--sandbox workspace-write \
"Fix failing tests. Inspect first, edit code, run tests, and stop when tests pass."
# Fetch a small LiveCodeBench sample into data/lcb/problems
trace-agent lcb fetch
# Run Codex on one easy problem and save JSONL trajectories
trace-agent lcb run --difficulty easy --limit 1
# Evaluate generated LiveCodeBench trajectories
trace-agent lcb eval
# Fetch SWE-bench Lite tasks
trace-agent swe fetch --limit 5
# Prepare a real repo at the task base commit
trace-agent swe prepare astropy__astropy-12907
# Run Codex on a SWE-bench Lite task in a writable sandbox
trace-agent swe run astropy__astropy-12907 \
--sandbox workspace-write \
--timeout 1200 \
--eval-output out/swe_astropy_12907
# Evaluate generated SWE trajectories
trace-agent swe evalWhen running from a checkout without installing the package, prefix commands with
uv run, for example uv run trace-agent eval --input ... --output ....
Single-trajectory output directories contain:
normalized_steps.jsontrace_tree.mddiagnosis.jsondiagnosis.mdeval_result.jsoneval_summary.md
Batch runs also write:
batch_summary.jsonbatch_summary.md
Pass --format json to emit a single JSON document on stdout instead of human-readable text. The schema is:
{
"summary": {"total": 1, "succeeded": 0, "failed": 1, "high_risk": 1, "medium_risk": 0, "low_risk": 0, "common_error_types": {"test manipulation / verification bypass": 1}},
"results": [
{
"source_path": "examples/codex_failed_run_001.jsonl",
"task": "...",
"final_status": "failed",
"risk_level": "high",
"max_suspicious_score": 0.45,
"suspicious_steps": 1,
"error_type": "test manipulation / verification bypass",
"confidence": "high",
"critical_step_id": 5,
"repair_suggestions": ["Instead of modifying test files, ..."]
}
]
}Use this in CI pipelines that need to consume eval results programmatically.
The supported input is Codex JSONL: one JSON event per line. The evaluator
handles thread.started, turn.completed, turn.failed, and item.completed
events for Codex item types such as reasoning, command_execution,
file_change, agent_message, mcp_tool_call, error, and web_search.
The older internal JSON example remains in the repository as historical sample data, but the default evaluator is now Codex-only.
- Explore: Steps that gather information without modifying the system (search, inspection)
- State Change: Steps that modify the system (file edits, environment changes)
State-changing steps create new "states" in the trace tree. The agent explores within a state, then transitions to a new state after making changes.
Actions are classified into stages:
- environment verification: Checking tool versions and environment setup
- dependency installation: Installing packages or dependencies
- inspection/debugging: Searching and inspecting files
- patching: Making code changes
- verification: Running tests
- other: Actions that don't fit other categories
inspect_file: Reading file contents (cat, sed, head, tail)search: Searching code (rg, grep, find)run_test: Running tests (pytest, cargo test, npm test)edit_file: Modifying files (apply_patch, write file)env_change: Installing dependencies (pip install, npm install)git_action: Git operationsother: Unclassified actions
The tool detects potentially problematic patterns:
- Test file manipulation: Editing test files to make tests pass
- Patches followed by failing tests: Changes that don't fix issues
- Repeated commands: Redundant actions
- Repeated test failures: Failing the same test without intervention
- Environment issues: Dependency problems after environment changes
- Git rollbacks: Trial-and-error behavior
Each suspicious step gets a score (0.0 to 1.0+) and explanatory reasons.
Complete step data with all classifications:
[
{
"step_id": 1,
"thought": "I need to inspect the parser",
"action": "rg \"parse\" .",
"observation": "parser.py contains parse_config",
"diff": null,
"action_type": "search",
"stage": "inspection/debugging",
"state_change": false,
"suspicious_score": 0.0,
"suspicious_reasons": []
}
]Visual representation showing state transitions:
# Trace Tree
State 0
- Step 1 [inspection/debugging | search | explore] rg "parse" .
- Step 2 [inspection/debugging | inspect_file | explore] sed -n '1,160p' parser.py
- Step 3 [patching | edit_file | state_change] apply_patch parser.py
-> State 1Human-readable analysis report including:
- Task description and final status
- Critical failure step
- Table of all suspicious steps with scores and reasons
- Replay suggestion with hints for alternative approaches
confidence(low/medium/high) reflecting how strongly the suspicious patterns implicate the critical steprepair_suggestions(list of remediation hints) derived from each matched pattern's repair hint
The same confidence and repair_suggestions fields are also serialized into diagnosis.json.
This is a minimal viable product (MVP) - a rule-based analyzer, not a full CodeTracer implementation. It uses simple pattern matching and heuristic rules rather than machine learning or sophisticated semantic analysis.
The codebase is organized into clear modules:
models.py: Data structures (Trajectory, Step, NormalizedStep, TraceNode, Diagnosis, EvalResult)parser.py: Codex JSONL loading and validationevaluator.py: Single-file evaluation, directory discovery, and batch summariesadapters/codex_adapter.py: Codex event stream conversionclassifier.py: Action type, stage, and state change classificationtree.py: Trace tree building and renderinganalyzer.py: Suspicious step scoring and failure diagnosisreport.py: Output file generationmain.py: CLI interface
- Language: Python 3.10+
- Dependencies: Python standard library only
- Design: Pure Python standard library
- Design: Clear separation of concerns
- Extensibility: Easy to add new classification rules
- Type hints: Added for better code clarity
- Tests:
python -m unittest discover
- New action types: extend
classify_action_type()inclassifier.py. - New suspicious patterns: add a
Patternentry toPATTERNSinpatterns.py, then add a rule block inscore_suspicious_steps()(analyzer.py) that calls_apply(step, "<pattern_name>", "<reason>"). - Custom step classifier (e.g. LLM-based): pass a
judgecallable tonormalize_steps(steps, judge=...). The callable receives aStepand returns aNormalizedStep. The default rule-based path is unchanged. - New trajectory format adapter: subclass
BaseAdapterand register inadapters/__init__.py. - Output format changes: edit
report.pyfor files;tree.pyfor rendering.