A modular, framework-agnostic evaluation and benchmarking harness for testing, measuring, comparing, and improving AI agents.
Define → Run → Evaluate → Measure → Compare → Improve
- Why This Project?
- What is an Evaluation Harness?
- Core Workflow
- Key Features
- Architecture
- Supported Evaluators
- Supported Metrics
- Supported Adapters
- Quick Start
- Command Line Interface (CLI)
- Example Evaluation Output
- Creating a Test Case
- Creating a Custom Evaluator
- Creating a Custom Agent Adapter
- Adding a Metric
- Model Context Protocol (MCP) Extensibility
- Agent Evolution & Iterative Improvement
- Comparing Agent Versions
- Project Structure
- Technology Stack
- Testing & Quality Assurance
- Local Development
- Contributing
- Hacktoberfest
- Roadmap
- Design Principles
- Security
- License
- Acknowledgements
- Project Status
Evaluating AI agents is fundamentally harder than testing standard software applications. Unlike deterministic functions, AI agents exhibit:
- Non-deterministic responses: The exact same prompt can yield variations in phrasing that break rigid string assertions even when the underlying reasoning is correct.
- Multi-step execution trajectories: Modern agents reason, execute tools, retrieve data, and retry. Measuring only the final output obscures unnecessary tool invocations, token inflation, and faulty intermediate steps.
- Latency and operational cost: An agent that answers correctly in 15 seconds at $0.08 per request may be unusable compared to one that succeeds in 1.2 seconds at $0.003.
- Fragile reliability: Minor prompt tweaks, tool schema edits, or model version bumps can introduce unexpected regressions in previously working capabilities.
- Framework lock-in: Many existing benchmarking suites tightly couple evaluation logic to specific commercial APIs or heavyweight orchestration frameworks.
Traditional unit testing frameworks (such as unittest or raw pytest) do not natively provide trajectory accounting, multi-metric directionality, regression threshold enforcement, or cross-run delta tracking.
The Agentic Evolution Harness provides a lightweight, modular foundation that abstracts test definitions, execution, assertions, and metrics so you can evaluate any AI agent systematically.
An evaluation harness is an automated test bench designed specifically for evaluating agentic systems.
In classical software engineering, a test harness feeds fixed inputs to a function and asserts on the exact return value. In an AI agent evaluation harness:
- A Test Case specifies user inputs, criteria, and expected behaviors.
- An Agent receives the input and returns a structured output along with observable traces (e.g., tools called, token usage, latency).
- Pluggable Evaluators score individual outputs against task requirements.
- Pluggable Metrics aggregate performance indicators across the full suite.
- A Reporter produces machine-readable artifacts and terminal dashboards.
- A Comparator detects performance improvements or regressions against baseline runs.
Test Case ──► Agent ──► Agent Output / Trace ──► Evaluators ──► Metrics ──► Report
graph TD
TC[Test Cases] --> Runner[Agent Runner]
Agent[Agent Adapter] --> Runner
Runner -->|Execute Input| Agent
Agent -->|Output + Traces| Runner
Runner -->|Evaluate Case| Evaluators[Evaluator Plugins]
Evaluators -->|Normalized Scores| Runner
Runner -->|Aggregate Results| Metrics[Metric Plugins]
Metrics -->|Summary Statistics| Experiment[Experiment Record]
Experiment --> Report[Console / JSON / Markdown Report]
Experiment --> Comparator[Experiment Comparator]
Baseline[Baseline Experiment JSON] --> Comparator
Comparator --> Regression[Regression Detection]
Regression --> Gate{Pass or Fail CI?}
- Define: Author test suites in declarative YAML or JSON format.
- Execute: The runner invokes your agent through an adapter without framework lock-in.
- Trace: Observable execution metadata (tool calls, arguments, latency, errors) is recorded.
- Evaluate: Evaluators score each test case independently with explanations.
- Aggregate: Metrics calculate cross-benchmark statistics (success rate, accuracy, cost, latency).
- Compare: The comparator computes directional performance deltas between versions.
- Gate: Regression rules check whether candidate agent revisions satisfy defined thresholds.
- Framework-Agnostic Core: Connect any agent—built with LangChain, LangGraph, CrewAI, AutoGen, raw API calls, or custom Python functions.
- Protocol-Driven Extensibility: Evaluators, metrics, adapters, and reporters are defined via Python Protocols (
PEP 544). Add plugins without touching core runner logic. - Zero-LLM Default Dependencies: Built-in evaluators (
exact_match,contains,json_match,tool_usage) run locally in milliseconds without API keys or network dependencies. - Directional Metric Math: Metrics declare their direction (
higher_is_better,lower_is_better, orneutral). The comparison engine automatically identifies whether a metric delta is an improvement or a regression. - Regression Detection: Configure threshold budgets (e.g., minimum 80% success rate, maximum 2.0s latency) to enforce CI quality gates.
- Rich Terminal UX: Formatted terminal dashboards with progress tables, status badges, and summary cards powered by Rich.
- Machine-Readable Artifacts: Export runs to standardized JSON schema artifacts for archival and CI/CD pipelines.
- Strict Typing & Quality: 100% type-annotated with MyPy strict mode, linted with Ruff, and tested with pytest.
The codebase enforces strict separation of concerns across decoupled packages:
src/harness/
├── core/ # Protocols (Agent, Evaluator, Metric) & Data Models
├── runner/ # AgentRunner pipeline orchestrator & execution timer
├── evaluators/ # Single-test evaluation plugins (exact_match, contains, json_match, tool_usage)
├── metrics/ # Cross-test aggregation plugins (success_rate, accuracy, latency, cost, reliability)
├── adapters/ # Agent integration bridges (generic, mock, subprocess)
├── reporters/ # Output formatters (console, json, markdown)
├── benchmarks/ # YAML/JSON benchmark schema & loader
├── comparison/ # ExperimentComparator & RegressionDetector
└── cli/ # Typer CLI commands (evaluate, compare, validate, list-plugins)
- Core (
harness.core): Defines immutable data models (TestCase,AgentOutput,EvaluationResult,TestCaseResult,Experiment,Trace) and runtime protocols. Has zero external dependencies outside Pydantic. - Runner (
harness.runner): Dispatches inputs to the agent, records execution durations, resolves evaluators dynamically, and packages results. Contains zero evaluator-specific logic. - Evaluators (
harness.evaluators): Independent plugins evaluating a single(TestCase, AgentOutput)pair. - Metrics (
harness.metrics): Plugins computing aggregate statistics over a collection ofTestCaseResultrecords. - Adapters (
harness.adapters): Adapters standardizing external agent invocations into theAgentprotocol. - Reporters (
harness.reporters): Translators formatting experiment models into Rich console output, JSON artifacts, or Markdown tables. - Comparison (
harness.comparison): Compares baseline and candidate experiment runs, computes metric deltas, and validates regression thresholds.
All built-in evaluators are deterministic and run locally without API keys:
| Evaluator | Identifier | Purpose | Key Options |
|---|---|---|---|
| Exact Match | exact_match |
Verifies identical string equality between output and expected text. | ignore_case (bool), strip_whitespace (bool) |
| Contains | contains |
Asserts substring inclusion or regular expression pattern matching. | substring (str), is_regex (bool), case_sensitive (bool) |
| JSON Match | json_match |
Verifies structural and semantic JSON equivalence. Tolerates code blocks. | subset (bool), ignore_order (bool) |
| Tool Usage | tool_usage |
Validates tool call trajectories, required sequences, and arguments. | expected_tools (list), ordered (bool), forbidden_tools (list), validate_args (dict) |
Metrics aggregate performance across all executed test cases:
| Metric | Identifier | Direction | Purpose |
|---|---|---|---|
| Success Rate | success_rate |
higher_is_better |
Percentage of test cases that passed all assigned evaluators. |
| Accuracy | accuracy |
higher_is_better |
Mean normalized evaluation score (0.0 to 1.0) across all tests. |
| Average Latency | latency |
lower_is_better |
Mean execution duration in seconds per test case. |
| Total Cost | cost |
lower_is_better |
Summed monetary cost in USD across all executions. |
| Reliability | reliability |
higher_is_better |
Percentage of agent runs that completed without errors or crashes. |
Adapters connect external agents to the standard Agent protocol:
| Adapter | Identifier | Purpose | Configuration |
|---|---|---|---|
| Generic | generic |
Dynamically imports and executes any Python function or class callable. | module: "pkg.module:callable" |
| Mock | mock |
Returns deterministic canned responses and simulated tool calls. | responses (dict), latency (float), default_output (str) |
| Subprocess | subprocess |
Executes an external CLI tool or binary via standard input/output. | command (list/str), timeout_seconds (float) |
You can clone, install, and execute your first agent evaluation benchmark in under 3 minutes.
git clone https://github.com/CH-JASWANTH-KUMAR/agentic-evolution-harness.git
cd agentic-evolution-harness
# Install dependencies using uv
uv sync
# Or using pip in a virtual environment:
# python -m venv .venv && source .venv/bin/activate && pip install -e ".[dev]"Verify that your local environment is correctly configured:
uv run pytestEvaluate the included customer support mock agent against a test benchmark:
uv run harness evaluate examples/basic/benchmark.yamlCompare a baseline run against an improved agent version to inspect performance deltas:
uv run harness compare examples/basic/v1_baseline.json examples/basic/v2_improved.jsonThe CLI provides four core commands:
Executes an evaluation benchmark against an agent:
uv run harness evaluate <path/to/benchmark.yaml> [OPTIONS]
# Common Options:
# --agent, -a TEXT Override agent callable (e.g. "my_module:run_agent")
# --output, -o PATH Output artifact path (default: results/latest.json)
# --format, -f TEXT Report format: console, json, or markdown
# --thresholds / --no-thresholds Enforce regression thresholds (default: on)Compares two experiment JSON runs and displays metric deltas:
uv run harness compare <baseline.json> <candidate.json> [OPTIONS]
# Common Options:
# --fail-on-regression / --allow-regression Exit code 1 if metrics regressValidates a benchmark configuration syntax and schema without executing the agent:
uv run harness validate examples/basic/benchmark.yamlLists all discovered and registered evaluators, metrics, adapters, and reporters:
uv run harness list-pluginsWhen running uv run harness evaluate examples/basic/benchmark.yaml:
╭─────────────────────────────────────────────────────────╮
│ Agentic Evolution Harness │
│ Experiment ID: c14aae1f | Name: basic-support-benchmark │
╰─────────────────────────────────────────────────────────╯
Test Results (5 cases)
╭──────────┬─────────────────┬───────┬──────────┬──────────────────────────────╮
│ Status │ Test Case │ Score │ Duration │ Explanation │
├──────────┼─────────────────┼───────┼──────────┼──────────────────────────────┤
│ ✓ PASS │ refund_request │ 1.00 │ 0.00s │ Output contains 'Refund of │
│ │ │ │ │ $49.99 processed'. │
│ ✓ PASS │ order_status │ 1.00 │ 0.00s │ Output contains 'in │
│ │ │ │ │ transit'. │
│ ✓ PASS │ weather_lookup │ 1.00 │ 0.00s │ Output contains '65°F and │
│ │ │ │ │ sunny'. │
│ ✓ PASS │ greeting │ 1.00 │ 0.00s │ Output matches expected │
│ │ │ │ │ output exactly. │
│ ✗ FAIL │ unknown_request │ 0.00 │ 0.00s │ Expected 'Rocket launched │
│ │ │ │ │ successfully.', but received │
│ │ │ │ │ 'I'm sorry, I didn't │
│ │ │ │ │ understand your request.'. │
╰──────────┴─────────────────┴───────┴──────────┴──────────────────────────────╯
Summary Metrics
╭──────────────┬─────────┬────────────────────┬────────────────────────────────╮
│ Metric │ Value │ Direction │ Description │
├──────────────┼─────────┼────────────────────┼────────────────────────────────┤
│ Success Rate │ 80.0% │ ▲ higher is better │ Percentage of passed test │
│ │ │ │ cases │
│ Accuracy │ 0.80 │ ▲ higher is better │ Average normalized score │
│ │ │ │ across all test cases │
│ Latency │ 0.19s │ ▼ lower is better │ Average execution latency in │
│ │ │ │ seconds │
│ Cost │ $0.0007 │ ▼ lower is better │ Total monetary cost in USD │
│ Reliability │ 100.0% │ ▲ higher is better │ Percentage of error-free │
│ │ │ │ executions │
╰──────────────┴─────────┴────────────────────┴────────────────────────────────╯
Full report written to: results/latest.json
Test cases are defined in YAML or JSON files conforming to BenchmarkConfig:
version: "1.0"
name: "customer-support-benchmark"
description: "Evaluates order lookup and refund capabilities."
agent:
adapter: generic
module: "examples.basic.mock_agent:support_agent"
evaluators:
- name: contains
metrics:
- success_rate
- accuracy
- latency
thresholds:
success_rate:
minimum: 0.80
tests:
- id: "refund-01"
name: "refund_request"
description: "Verify refund processing and tool trajectory."
input: "My item was broken. Please process a refund."
expected: "Refund of $49.99 processed"
evaluators:
- name: contains
options:
substring: "Refund of $49.99 processed"
- name: tool_usage
options:
expected_tools: ["lookup_order", "process_refund"]
ordered: true
tags: ["billing", "refunds"]id: Unique identifier for the test case.name: Human-readable name used in reports and diffs.input: Prompt string or dictionary provided to the agent.expected: Ground truth text, structured dictionary, or expected tool list.evaluators: List of evaluator configurations specific to this test case.tags: Optional list of category labels for filtering and reporting.
Contributors can add new evaluators without modifying the core runner or existing components.
- Create a new file in
src/harness/evaluators/fuzzy.py:
from __future__ import annotations
from harness.core.agent import AgentOutput
from harness.core.result import EvaluationResult
from harness.core.testcase import TestCase
from harness.evaluators.base import EvaluatorRegistry
@EvaluatorRegistry.register("fuzzy")
class FuzzyMatchEvaluator:
"""Evaluates whether output text matches expected text above a threshold."""
name: str = "fuzzy"
def __init__(self, threshold: float = 0.8) -> None:
self.threshold = threshold
def evaluate(self, test_case: TestCase, output: AgentOutput) -> EvaluationResult:
if output.is_error:
return EvaluationResult(
evaluator=self.name,
score=0.0,
passed=False,
explanation=f"Agent produced an error: {output.error}",
)
expected = str(test_case.expected or "")
actual = output.text
# Compute simple character overlap ratio
overlap = len(set(actual) & set(expected)) / max(len(set(expected)), 1)
passed = overlap >= self.threshold
return EvaluationResult(
evaluator=self.name,
score=overlap,
passed=passed,
explanation=f"Overlap ratio was {overlap:.2f} (threshold: {self.threshold:.2f}).",
)- Expose the evaluator in
src/harness/evaluators/__init__.py:
from harness.evaluators.fuzzy import FuzzyMatchEvaluator
__all__ = [
...,
"FuzzyMatchEvaluator",
]-
Add unit tests in
tests/unit/evaluators/test_fuzzy.py. -
Verify:
uv run harness list-plugins evaluators
uv run pytest tests/unit/evaluators/test_fuzzy.pyTo evaluate an agent built with LangChain, CrewAI, AutoGen, or an internal framework, wrap it in a lightweight callable or create an adapter:
from harness.core.agent import AgentOutput
from harness.core.trace import ToolCallRecord
def my_custom_agent(user_input: str) -> AgentOutput:
# 1. Call your model or agent framework
response = call_llm(user_input)
# 2. Return standard AgentOutput
return AgentOutput(
output=response.text,
latency=response.latency_seconds,
tool_calls=[ToolCallRecord(tool_name=t.name, args=t.args) for t in response.tool_calls],
)Specify your agent in your benchmark YAML:
agent:
adapter: generic
module: "my_project.agent:my_custom_agent"Metrics aggregate results across multiple test cases.
- Create the file in
src/harness/metrics/p95_latency.py:
from __future__ import annotations
from collections.abc import Sequence
from harness.core.experiment import MetricResult
from harness.core.result import TestCaseResult
from harness.metrics.base import MetricDirection, MetricRegistry
@MetricRegistry.register("p95_latency")
class P95LatencyMetric:
"""Computes 95th percentile execution latency across test cases."""
name: str = "p95_latency"
direction: MetricDirection = MetricDirection.LOWER_IS_BETTER
def calculate(self, results: Sequence[TestCaseResult]) -> MetricResult:
if not results:
return MetricResult(
name=self.name,
value=0.0,
formatted_value="0.00s",
direction=self.direction.value,
description="95th percentile latency",
)
latencies = sorted(r.duration for r in results)
idx = int(0.95 * len(latencies))
p95 = latencies[min(idx, len(latencies) - 1)]
return MetricResult(
name=self.name,
value=p95,
formatted_value=f"{p95:.2f}s",
direction=self.direction.value,
description="95th percentile execution latency in seconds",
)- Register the metric in
src/harness/metrics/__init__.pyand add unit tests intests/unit/metrics/test_p95_latency.py.
The harness is designed to support the Model Context Protocol (MCP) ecosystem without making MCP mandatory for the core framework.
- Structured Tool Traces:
ToolCallRecordcaptures tool names, input arguments, outputs, execution duration, and statuses. - Trajectory Validation:
ToolUsageEvaluatorvalidates:- Inclusion of expected tool calls.
- Strict sequence ordering of tool calls.
- Prohibition of forbidden tools.
- Exact argument matching.
- MCP Client Adapter: Direct invocation of MCP servers to capture server-side tool calls and protocol traffic.
- MCP Trajectory Evaluator: Verification of multi-server tool selection, intermediate tool error recovery, and context resource retrieval.
In this framework, "Evolution" does not mean autonomous, self-modifying code. Rather, it refers to the evidence-based iterative improvement cycle used by engineers to improve agent performance:
Agent v1 ──► Evaluation ──► Baseline Results
│
┌───────────────────────────┘
▼
Refactor Agent Prompt / Tools
│
▼
Agent v2 ──► Evaluation ──► Candidate Results
│
▼
Compare Runs
│
├──► Verified Improvements
└──► Regression Detection
The comparison engine computes directional deltas between two runs:
uv run harness compare results/v1_baseline.json results/v2_improved.json Metric Comparison
╭──────────────┬───────────────┬───────────────┬───────────────┬───────────────╮
│ Metric │ Baseline (v1) │ Candidate(v2) │ Delta │ Evaluation │
├──────────────┼───────────────┼───────────────┼───────────────┼───────────────┤
│ Accuracy │ 0.80 │ 1.00 │ +0.2000 │ ▲ Improvement │
│ │ │ │ (+25.0%) │ │
│ Cost │ $0.0007 │ $0.0006 │ -0.0001 │ ▲ Improvement │
│ │ │ │ (-13.0%) │ │
│ Latency │ 0.19s │ 0.13s │ -0.0620 │ ▲ Improvement │
│ │ │ │ (-32.3%) │ │
│ Reliability │ 100.0% │ 100.0% │ 0.0000 │ • Neutral │
│ Success Rate │ 80.0% │ 100.0% │ +0.2000 │ ▲ Improvement │
│ │ │ │ (+25.0%) │ │
╰──────────────┴───────────────┴───────────────┴───────────────┴───────────────╯
Newly passing tests (1): unknown_request
If a candidate violates a regression threshold (e.g., latency rises above allowable limits or success rate drops), the CLI exits with status code 1 to stop deployment in CI.
.
├── benchmarks/
│ └── examples/ # Standard benchmark datasets (YAML)
├── docs/
│ ├── architecture.md # Detailed architecture specification
│ ├── quickstart.md # Quickstart walkthrough
│ └── concepts/ # Deep-dive concept guides
├── examples/
│ └── basic/ # Runnable examples, mock agent, and baseline runs
├── src/
│ └── harness/ # Core package
│ ├── adapters/ # Generic, Mock, Subprocess adapters
│ ├── benchmarks/ # Benchmark schemas and loader
│ ├── cli/ # Typer CLI commands
│ ├── comparison/ # Comparator and regression detection
│ ├── core/ # Protocols and domain models
│ ├── evaluators/ # Built-in evaluators
│ ├── metrics/ # Built-in metrics
│ ├── reporters/ # Console, JSON, Markdown reporters
│ └── runner/ # Execution orchestrator
├── tests/
│ ├── integration/ # Pipeline and CLI integration tests
│ └── unit/ # Unit tests per component
├── CHANGELOG.md # Version history
├── CODE_OF_CONDUCT.md # Contributor Covenant v2.1
├── CONTRIBUTING.md # Contributor guidelines and workflow
├── LICENSE # MIT License
├── Makefile # Automation commands
├── pyproject.toml # Project configuration and dependencies
└── SECURITY.md # Security vulnerability reporting policy
- Python: 3.11+
- Pydantic (v2): Data validation, schema definitions, and serialization.
- Typer: Type-safe CLI commands with argument and option parsing.
- Rich: Terminal tables, panels, badges, and progress rendering.
- PyYAML: Benchmark configuration parsing.
- pytest & pytest-cov: Unit and integration testing with branch coverage.
- Ruff: Fast Python formatting and linting.
- MyPy: Strict static type checking.
The codebase enforces strict quality checks:
# Run complete test suite with coverage
uv run pytest
# Check formatting and linting
uv run ruff check .
uv run ruff format --check .
# Run static type checker
uv run mypy srcYou can run all quality checks with a single command:
make checkSetting up the development environment:
# Sync all dependencies including dev tools
uv sync --all-extras
# Run tests
uv run pytest
# Auto-format and fix linter issues
make formatWe welcome community contributions! Please review our Contributing Guide and Contributor Extensibility Guide for full instructions, protocols, and examples.
| Goal | Target Directory | Tests Directory |
|---|---|---|
| Add an Evaluator | src/harness/evaluators/ |
tests/unit/evaluators/ |
| Add a Metric | src/harness/metrics/ |
tests/unit/metrics/ |
| Add an Agent Adapter | src/harness/adapters/ |
tests/unit/adapters/ |
| Add a Reporter | src/harness/reporters/ |
tests/unit/reporters/ |
| Add Benchmark Data | benchmarks/examples/ |
tests/integration/ |
| Improve CLI | src/harness/cli/ |
tests/integration/test_cli.py |
| Documentation | docs/ or README.md |
None |
This repository is designed specifically for Hacktoberfest contributors:
- Modular Isolation: Evaluators, metrics, and adapters can be implemented in a single isolated file without altering the runner.
- Fast Local Feedback: The entire test suite runs in under 1 second locally without external API keys.
- Clear Contracts: Protocols clearly define method signatures, input types, and return values.
Check out our Contributing Guide and Contributor Extensibility Guide for starter ideas, exact protocol specifications, and working examples.
- Protocol-driven
Agent,Evaluator,Metric, andReporterarchitecture. - Built-in evaluators:
exact_match,contains,json_match,tool_usage. - Built-in metrics:
success_rate,accuracy,latency,cost,reliability. - Built-in adapters:
generic,mock,subprocess. - Built-in reporters:
console,json,markdown. - Experiment persistence, cross-run comparison, and regression threshold detection.
- Typer CLI (
evaluate,compare,validate,list-plugins). - 90% branch test coverage, MyPy strict typing, Ruff linting, and GitHub Actions CI.
- Parallel Test Runner: Concurrent test case execution via worker thread pools.
- LLM-as-a-Judge: Abstract judge extension with provider agnostic connectors.
- Interactive HTML Reporter: Standalone interactive HTML report with charts.
- Native MCP Client: Live Model Context Protocol client adapter for inspecting tool invocations.
- Statistical Significance: Bootstrapping and t-tests for comparing benchmark iterations.
- Modularity: Every evaluator and metric is an independent, pluggable component.
- Framework Agnosticism: The harness treats any agent as a callable black box.
- Reproducibility: Evaluations produce deterministic, versioned JSON records.
- Minimal Dependencies: The core framework avoids heavy dependencies or vendor lock-in.
- Developer Experience: Clear protocols, automated linting, type safety, and helpful CLI diagnostics.
Please refer to our Security Policy for information on reporting security vulnerabilities.
This project is licensed under the MIT License.
Built for AI agent researchers, software engineers, and the open-source community participating in Hacktoberfest.
Active Development (v0.1.0). The core evaluation pipeline, CLI, built-in evaluators, metrics, reporters, and comparison engine are fully implemented, tested, and ready for use.