Evaluate AI agents with Unix-style pipeline commands. Schema-driven adapters for any CLI agent, trajectory capture, pass@k metrics, and multi-run comparison.
-
Updated
Jun 17, 2026 - TypeScript
Evaluate AI agents with Unix-style pipeline commands. Schema-driven adapters for any CLI agent, trajectory capture, pass@k metrics, and multi-run comparison.
Qualitative benchmark suite for evaluating AI coding agents and orchestration paradigms on realistic, complex development tasks
Compare LangGraph, CrewAI, and custom orchestration using MCP. Complete working examples for contract analysis with shared tools. Production-ready patterns.
Add a description, image, and links to the agent-comparison topic page so that developers can more easily learn about it.
To associate your repository with the agent-comparison topic, visit your repo's landing page and select "manage topics."