Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
-
Updated
Aug 18, 2026 - Shell
Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
Copy inflation in multi-turn search agents: 78-92% of generated tokens are copied from retrieved documents and carry inflated log-probabilities, breaking confidence-based methods. Diagnostic toolkit + Retrieval-Grounded Voting. Findings of EMNLP 2026.
VCR cassettes for agent trajectories: record agent runs as DAGs, replay the canonical path, only call the LLM for net-new paths
ATIF-native analytics over Claude Code agent trajectories: convert sessions to ATIF, materialize a corpus, query it with DuckDB.
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Open rubric and a small, hand-built gold set for judging multi-step web-research agent trajectories. The first dataset from Assayo.
Independent research on trajectory-aware AI-agent evaluation, delegated authority, control integrity, and failure-preserving reproducibility.
Capture coding-agent sessions (Claude Code / Codex) at the source — trajectory + verifiable git environment — and score each session's training value (grounded × rich × focused). Local, deterministic, no model at runtime.
To associate your repository with the agent-trajectories topic, visit your repo's landing page and select "manage topics."