Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@

## Unreleased

## 0.17.5 — 2026-10-05

- fix: reject duplicate Inspect samples instead of counting them as extra attempts. evalarc diff now refuses an Inspect AI log, JSON or .eval archive, that contains the same (sample id, epoch) pair more than once. Before this change the copy was silently counted as an extra attempt, which inflated coverage and could hide a less_covered change. The ids are compared after str(), so an integer 7 and a string '7' in the same epoch also count as a duplicate. A duplicate in either the baseline or the current file makes load_results raise ValueError, and the CLI prints that message to stderr and exits 2 before any comparison. The message names the sample (shown with repr() and cut to 80 characters, so newlines and control characters stay escaped) and the epoch, and tells you to re-export with inspect log dump. Valid multi-epoch logs, logs without an epoch field, the recorded examples, the diff.json schema and the gate logic are unchanged. docs/ci-gate.md and docs/ci-gate.zh-CN.md add this case to the exit-2 list.

- CI tests Python 3.14 alongside 3.11–3.13, and the package declares 3.14
support. A new test reads zstd-compressed Inspect `.eval` archives directly
on 3.14 and checks they give the same diff as the JSON logs.
Expand Down
4 changes: 2 additions & 2 deletions CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@ title: "EvalArc: Auditable Evaluations for AI Agents"
type: software
authors:
- name: EvalArc contributors
version: 0.17.4
date-released: 2026-10-02
version: 0.17.5
date-released: 2026-10-05
license: MIT
repository-code: "https://github.com/noteflowai/evalarc"
url: "https://noteflowai.github.io/evalarc/"
Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,13 +66,13 @@ and the upgrade gate fails. These checks share two format/schema failures.
installs the published wheel, downloads the records and rebuilds the comparison.
Verification exits 0 for consistency; comparison exits 1 for the regression.

Install the released reviewer from [PyPI](https://pypi.org/project/evalarc/0.17.4/)
Install the released reviewer from [PyPI](https://pypi.org/project/evalarc/0.17.5/)
in a fresh virtual environment:

```bash
python3 -m venv .venv
. .venv/bin/activate
python -m pip install evalarc==0.17.4
python -m pip install evalarc==0.17.5
evalarc --version
```

Expand Down Expand Up @@ -124,7 +124,7 @@ because checks missing from it cannot be compared; rerun that evaluation to
completion and compare again.

```yaml
- uses: noteflowai/evalarc@v0.17.4 # or a full commit SHA
- uses: noteflowai/evalarc@v0.17.5 # or a full commit SHA
with:
baseline: evals/baseline.json
current: results/current.json
Expand Down
6 changes: 3 additions & 3 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,13 +52,13 @@ EvalArc 是面向 Agent 评测的 Python 命令行复核工具:逐项对比已
3. **本地复算。** [首次复核指南](docs/first-review.zh-CN.md)使用发布的 wheel 和原始记录重建报告;
复核退出 0 表示一致,对照退出 1 表示发现退步。

在新的虚拟环境中,从 [PyPI](https://pypi.org/project/evalarc/0.17.4/)
在新的虚拟环境中,从 [PyPI](https://pypi.org/project/evalarc/0.17.5/)
安装已发布的复核工具:

```bash
python3 -m venv .venv
. .venv/bin/activate
python -m pip install evalarc==0.17.4
python -m pip install evalarc==0.17.5
evalarc --version
```

Expand Down Expand Up @@ -102,7 +102,7 @@ PyPI 与 GitHub Releases 提供相同的 wheel 和源码包,也支持
准确率从 0.625 升至 0.8125,同时有三个检查项丢失通过。

```yaml
- uses: noteflowai/evalarc@v0.17.4 # 或完整的提交 SHA
- uses: noteflowai/evalarc@v0.17.5 # 或完整的提交 SHA
with:
baseline: evals/baseline.json
current: results/current.json
Expand Down
12 changes: 11 additions & 1 deletion docs/ci-gate.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,17 @@ status, such as JUnit, are never marked incomplete.

Exit codes: **0** no check lost passes, **1** at least one did, **2** the files
could not be read or compared (different formats, different Inspect task names,
no samples, malformed input). Treat 2 as a broken pipeline, not a regression.
no samples, malformed input, an Inspect sample repeated in the same epoch).
Treat 2 as a broken pipeline, not a regression.

An Inspect log that contains the same sample ID twice in one epoch (an integer
`7` and a string `"7"` count as the same ID) exits 2 and names the sample
and epoch. Counting the copy as an extra attempt would overstate coverage and
could hide a `less_covered` change. Such logs usually come from merging,
concatenating or editing results; re-export with
`inspect log dump LOG.eval > LOG.json` or rerun the evaluation. The same ID in
different epochs is a normal repeated attempt, and samples without an integer
`epoch` are not checked.

## Sampling-noise annotation

Expand Down
7 changes: 6 additions & 1 deletion docs/ci-gate.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,12 @@ epoch 中都对重复订单退款(`refund-duplicate` 退化),`cancel-pendi
门禁以 `less_covered` 失败,直到尝试次数与基线一致或提交新的基线;通过率下降时仍优先报告
`regressed` 或 `less_reliable`。删除一个失败的检查项也会失败,
否则删掉它就能让门禁变绿。退出码:**0** 无检查项丢失通过,**1** 至少一项丢失,
**2** 文件无法读取或无法比较(格式不同、Inspect 任务名不同、没有样本、格式错误)。
**2** 文件无法读取或无法比较(格式不同、Inspect 任务名不同、没有样本、格式错误、
Inspect 样本在同一 epoch 中重复)。同一 epoch 中重复出现的样本 ID(整数 `7` 与字符串
`"7"` 视为同一 ID)会以退出码 2 报告并给出样本与 epoch:若把副本算作额外尝试,会夸大覆盖并
掩盖 `less_covered`。这类日志通常来自手工合并、拼接或编辑;请用
`inspect log dump LOG.eval > LOG.json` 重新导出,或重新运行评测。同一 ID 出现在不同 epoch
是正常的重复尝试;没有整数 `epoch` 的样本不做此检查。
2 表示流水线本身有问题,不是退化。

## 采样噪声标注
Expand Down
4 changes: 2 additions & 2 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "evalarc-evidence-site",
"version": "0.17.4",
"version": "0.17.5",
"private": true,
"description": "Browser checks and the TypeScript report front end for EvalArc",
"scripts": {
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"

[project]
name = "evalarc"
version = "0.17.4"
version = "0.17.5"
description = "Find agent regressions behind better scores: compare changed checks, review recorded tool actions and verify evidence offline."
readme = "README.md"
requires-python = ">=3.11"
Expand Down
2 changes: 1 addition & 1 deletion src/evalarc/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
"""Auditable evaluations for AI agents."""

__version__ = "0.17.4"
__version__ = "0.17.5"
20 changes: 20 additions & 0 deletions src/evalarc/results_diff.py
Original file line number Diff line number Diff line change
Expand Up @@ -439,8 +439,22 @@ def _inspect(document: dict, threshold: float) -> dict:
usage: dict[str, list] = {}
outputs: dict[str, dict[str, list]] = {}
samples_meta: dict[str, list] = {}
seen: set[tuple[str, int]] = set()
for sample in document["samples"]:
case_id = str(sample["id"])
epoch = sample.get("epoch")
if isinstance(epoch, int) and not isinstance(epoch, bool):
# A repeated (id, epoch) pair would silently count as an extra attempt and
# overstate coverage. Ids 7 and "7" collide because case_id is str(id).
key = (case_id, epoch)
if key in seen:
raise ValueError(
f"Inspect log has duplicate sample {_shown_id(sample['id'])} epoch {epoch}: "
"each sample and epoch must appear once, or the copy would count as an "
"extra attempt. Re-export the log from Inspect (inspect log dump) instead "
"of merging or editing it."
)
seen.add(key)
checks = cases.setdefault(case_id, {})
_add_material(material, case_id, "input", _message_text(sample.get("input")))
usage.setdefault(case_id, []).append(inspect_usage(sample))
Expand Down Expand Up @@ -505,6 +519,12 @@ def _inspect(document: dict, threshold: float) -> dict:
}


def _shown_id(value: object, limit: int = 80) -> str:
"""Render an untrusted sample id safely: repr() escapes control characters."""
text = repr(value)
return text if len(text) <= limit else text[: limit - 3] + "..."


def _inspect_passed(value: object, threshold: float) -> bool | None:
if isinstance(value, bool):
return value
Expand Down
111 changes: 111 additions & 0 deletions tests/test_feature_e0a029ce3b1b.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
"""Duplicate Inspect samples (same id and epoch) are rejected, not counted as attempts.

All logs here are synthetic copies of the recorded examples, written by the tests.
"""

import copy
import json
import zipfile
from pathlib import Path

import pytest

from evalarc.cli import main
from evalarc.results_diff import diff, load_results

EXAMPLES = Path(__file__).resolve().parents[1] / "examples/results-diff"


def _load(name):
return json.loads((EXAMPLES / f"inspect/{name}.json").read_text())


def _write(path, document):
path.write_text(json.dumps(document))
return path


def _scored(template, sample_id, epoch):
sample = copy.deepcopy(template)
sample["id"] = sample_id
if epoch is None:
sample.pop("epoch", None)
else:
sample["epoch"] = epoch
sample.pop("error", None)
sample["scores"] = {"match": {"value": "C"}}
return sample


@pytest.mark.parametrize("side", ["baseline", "current"])
def test_duplicated_sample_exits_2_and_names_it(tmp_path, capsys, side):
document = _load(side)
duplicate = copy.deepcopy(document["samples"][0])
document["samples"].append(duplicate)
broken = _write(tmp_path / f"{side}-dup.json", document)
paths = {
"baseline": str(EXAMPLES / "inspect/baseline.json"),
"current": str(EXAMPLES / "inspect/current.json"),
}
paths[side] = str(broken)
output = tmp_path / "review"
code = main(["diff", paths["baseline"], paths["current"], "--output", str(output)])
assert code == 2
err = capsys.readouterr().err
assert "duplicate sample" in err
assert repr(duplicate["id"]) in err
assert f"epoch {duplicate['epoch']}" in err
assert "inspect log dump" in err
assert not (output / "diff.json").exists()


def test_integer_and_string_ids_collide(tmp_path, capsys):
document = _load("baseline")
template = document["samples"][0]
document["samples"] = [_scored(template, 7, 1), _scored(template, "7", 1)]
path = _write(tmp_path / "mixed.json", document)
assert main(["diff", str(EXAMPLES / "inspect/baseline.json"), str(path)]) == 2
err = capsys.readouterr().err
assert "duplicate sample" in err and "epoch 1" in err


def test_eval_archive_with_duplicate_members_is_rejected(tmp_path):
document = _load("current")
samples = document.pop("samples")
archive = tmp_path / "current.eval"
with zipfile.ZipFile(archive, "w", zipfile.ZIP_DEFLATED) as output:
output.writestr("header.json", json.dumps(document))
output.writestr("samples/a.json", json.dumps(samples[0]))
output.writestr("samples/b.json", json.dumps(samples[0]))
with pytest.raises(ValueError, match="duplicate sample"):
load_results(archive)


def test_control_characters_in_id_are_escaped(tmp_path):
document = _load("baseline")
template = document["samples"][0]
name = "bad\nid" + "x" * 200
document["samples"] = [_scored(template, name, 1), _scored(template, name, 1)]
with pytest.raises(ValueError) as caught:
load_results(_write(tmp_path / "ctrl.json", document))
message = str(caught.value)
assert "\n" not in message
assert "x" * 100 not in message


def test_valid_logs_are_unchanged(tmp_path):
result = diff(
load_results(EXAMPLES / "inspect/baseline.json"),
load_results(EXAMPLES / "inspect/current.json"),
)
assert result["blocking_changes"] == 3 and not result["gate_passed"]

document = _load("baseline")
template = document["samples"][0]
document["samples"] = [_scored(template, "repeat", epoch) for epoch in (1, 2, 3)]
loaded = load_results(_write(tmp_path / "epochs.json", document))
assert len(loaded["cases"]["repeat"]["match"]) == 3

document["samples"] = [_scored(template, "no-epoch", None) for _ in range(2)]
loaded = load_results(_write(tmp_path / "no-epoch.json", document))
assert len(loaded["cases"]["no-epoch"]["match"]) == 2
Loading