Skip to content

About

Codex/Claude-first agent trajectory evaluator: parses JSONL session streams, builds trace trees, scores suspicious steps, and grades SWE-bench Lite trajectories with verified FAIL_TO_PASS / PASS_TO_PASS results.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Agent Trajectory Eval

Status: v0.1.0 in progress — see the v0.1.0 roadmap issue (or docs/roadmap.md) for the path to a calibrated analyzer with real verified_pass numbers.

A Codex-first evaluation toolkit for coding agent trajectories. It reads Codex CLI JSONL session or codex exec --json streams and produces:

  1. Normalized Steps: Codex events converted into consistent step records
  2. Trace Tree: State transitions created by file and environment changes
  3. Risk Metrics: Counts for commands, tests, file changes, failures, and suspicious behavior
  4. Failure Diagnosis: Rule-based critical-step detection with replay hints
  5. Batch Reports: Directory-level summaries for CI or regression evaluation

The project intentionally focuses on Codex JSONL today. The adapter interface is still present, but only the Codex adapter is registered by default.

Quick Start

trace-agent eval --input examples/codex_failed_run_001.jsonl --output out/codex_eval

For a directory of trajectories:

trace-agent eval --input data/lcb/trajectories --output out/lcb_eval

CI-style exit codes:

trace-agent eval --input examples/codex_failed_run_001.jsonl --output out/codex_eval --ci
  • 0: tool ran and no evaluated trajectory failed or reached medium/high risk
  • 1: evaluation ran, but at least one trajectory failed or reached medium/high risk
  • 2: tool error, invalid input, or unsupported format

The older flat form is still supported for compatibility:

trace-eval --input examples/codex_failed_run_001.jsonl --output out/codex_eval

Commands

# Run Codex in a sandbox, capture JSONL, then evaluate it
trace-agent run \
  --output data/runs/task_001.jsonl \
  --eval-output out/task_001 \
  -C /path/to/repo \
  --sandbox workspace-write \
  "Fix failing tests. Inspect first, edit code, run tests, and stop when tests pass."

# Fetch a small LiveCodeBench sample into data/lcb/problems
trace-agent lcb fetch

# Run Codex on one easy problem and save JSONL trajectories
trace-agent lcb run --difficulty easy --limit 1

# Evaluate generated LiveCodeBench trajectories
trace-agent lcb eval

# Fetch SWE-bench Lite tasks
trace-agent swe fetch --limit 5

# Prepare a real repo at the task base commit
trace-agent swe prepare astropy__astropy-12907

# Run Codex on a SWE-bench Lite task in a writable sandbox
trace-agent swe run astropy__astropy-12907 \
  --sandbox workspace-write \
  --timeout 1200 \
  --eval-output out/swe_astropy_12907

# Evaluate generated SWE trajectories
trace-agent swe eval

When running from a checkout without installing the package, prefix commands with uv run, for example uv run trace-agent eval --input ... --output ....

Outputs

Single-trajectory output directories contain:

  • normalized_steps.json
  • trace_tree.md
  • diagnosis.json
  • diagnosis.md
  • eval_result.json
  • eval_summary.md

Batch runs also write:

  • batch_summary.json
  • batch_summary.md

JSON output (CI-friendly)

Pass --format json to emit a single JSON document on stdout instead of human-readable text. The schema is:

{
  "summary": {"total": 1, "succeeded": 0, "failed": 1, "high_risk": 1, "medium_risk": 0, "low_risk": 0, "common_error_types": {"test manipulation / verification bypass": 1}},
  "results": [
    {
      "source_path": "examples/codex_failed_run_001.jsonl",
      "task": "...",
      "final_status": "failed",
      "risk_level": "high",
      "max_suspicious_score": 0.45,
      "suspicious_steps": 1,
      "error_type": "test manipulation / verification bypass",
      "confidence": "high",
      "critical_step_id": 5,
      "repair_suggestions": ["Instead of modifying test files, ..."]
    }
  ]
}

Use this in CI pipelines that need to consume eval results programmatically.

Input Format

The supported input is Codex JSONL: one JSON event per line. The evaluator handles thread.started, turn.completed, turn.failed, and item.completed events for Codex item types such as reasoning, command_execution, file_change, agent_message, mcp_tool_call, error, and web_search.

The older internal JSON example remains in the repository as historical sample data, but the default evaluator is now Codex-only.

Key Concepts

Explore vs State Change

  • Explore: Steps that gather information without modifying the system (search, inspection)
  • State Change: Steps that modify the system (file edits, environment changes)

State-changing steps create new "states" in the trace tree. The agent explores within a state, then transitions to a new state after making changes.

Stages

Actions are classified into stages:

  • environment verification: Checking tool versions and environment setup
  • dependency installation: Installing packages or dependencies
  • inspection/debugging: Searching and inspecting files
  • patching: Making code changes
  • verification: Running tests
  • other: Actions that don't fit other categories

Action Types

  • inspect_file: Reading file contents (cat, sed, head, tail)
  • search: Searching code (rg, grep, find)
  • run_test: Running tests (pytest, cargo test, npm test)
  • edit_file: Modifying files (apply_patch, write file)
  • env_change: Installing dependencies (pip install, npm install)
  • git_action: Git operations
  • other: Unclassified actions

Suspicious Step Detection

The tool detects potentially problematic patterns:

  • Test file manipulation: Editing test files to make tests pass
  • Patches followed by failing tests: Changes that don't fix issues
  • Repeated commands: Redundant actions
  • Repeated test failures: Failing the same test without intervention
  • Environment issues: Dependency problems after environment changes
  • Git rollbacks: Trial-and-error behavior

Each suspicious step gets a score (0.0 to 1.0+) and explanatory reasons.

Output Files

normalized_steps.json

Complete step data with all classifications:

[
  {
    "step_id": 1,
    "thought": "I need to inspect the parser",
    "action": "rg \"parse\" .",
    "observation": "parser.py contains parse_config",
    "diff": null,
    "action_type": "search",
    "stage": "inspection/debugging",
    "state_change": false,
    "suspicious_score": 0.0,
    "suspicious_reasons": []
  }
]

trace_tree.md

Visual representation showing state transitions:

# Trace Tree

State 0
  - Step 1 [inspection/debugging | search | explore] rg "parse" .
  - Step 2 [inspection/debugging | inspect_file | explore] sed -n '1,160p' parser.py
  - Step 3 [patching | edit_file | state_change] apply_patch parser.py
    -> State 1

diagnosis.md

Human-readable analysis report including:

  • Task description and final status
  • Critical failure step
  • Table of all suspicious steps with scores and reasons
  • Replay suggestion with hints for alternative approaches
  • confidence (low/medium/high) reflecting how strongly the suspicious patterns implicate the critical step
  • repair_suggestions (list of remediation hints) derived from each matched pattern's repair hint

The same confidence and repair_suggestions fields are also serialized into diagnosis.json.

Limitations

This is a minimal viable product (MVP) - a rule-based analyzer, not a full CodeTracer implementation. It uses simple pattern matching and heuristic rules rather than machine learning or sophisticated semantic analysis.

Architecture

The codebase is organized into clear modules:

  • models.py: Data structures (Trajectory, Step, NormalizedStep, TraceNode, Diagnosis, EvalResult)
  • parser.py: Codex JSONL loading and validation
  • evaluator.py: Single-file evaluation, directory discovery, and batch summaries
  • adapters/codex_adapter.py: Codex event stream conversion
  • classifier.py: Action type, stage, and state change classification
  • tree.py: Trace tree building and rendering
  • analyzer.py: Suspicious step scoring and failure diagnosis
  • report.py: Output file generation
  • main.py: CLI interface

Technical Details

  • Language: Python 3.10+
  • Dependencies: Python standard library only
  • Design: Pure Python standard library
  • Design: Clear separation of concerns
  • Extensibility: Easy to add new classification rules
  • Type hints: Added for better code clarity
  • Tests: python -m unittest discover

Extending the Analyzer

  • New action types: extend classify_action_type() in classifier.py.
  • New suspicious patterns: add a Pattern entry to PATTERNS in patterns.py, then add a rule block in score_suspicious_steps() (analyzer.py) that calls _apply(step, "<pattern_name>", "<reason>").
  • Custom step classifier (e.g. LLM-based): pass a judge callable to normalize_steps(steps, judge=...). The callable receives a Step and returns a NormalizedStep. The default rule-based path is unchanged.
  • New trajectory format adapter: subclass BaseAdapter and register in adapters/__init__.py.
  • Output format changes: edit report.py for files; tree.py for rendering.

About

Codex/Claude-first agent trajectory evaluator: parses JSONL session streams, builds trace trees, scores suspicious steps, and grades SWE-bench Lite trajectories with verified FAIL_TO_PASS / PASS_TO_PASS results.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages