diff --git a/CHANGELOG.md b/CHANGELOG.md index 87d9851..755110f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,10 @@ ## Unreleased +## 0.17.5 — 2026-10-05 + +- fix: reject duplicate Inspect samples instead of counting them as extra attempts. evalarc diff now refuses an Inspect AI log, JSON or .eval archive, that contains the same (sample id, epoch) pair more than once. Before this change the copy was silently counted as an extra attempt, which inflated coverage and could hide a less_covered change. The ids are compared after str(), so an integer 7 and a string '7' in the same epoch also count as a duplicate. A duplicate in either the baseline or the current file makes load_results raise ValueError, and the CLI prints that message to stderr and exits 2 before any comparison. The message names the sample (shown with repr() and cut to 80 characters, so newlines and control characters stay escaped) and the epoch, and tells you to re-export with inspect log dump. Valid multi-epoch logs, logs without an epoch field, the recorded examples, the diff.json schema and the gate logic are unchanged. docs/ci-gate.md and docs/ci-gate.zh-CN.md add this case to the exit-2 list. + - CI tests Python 3.14 alongside 3.11–3.13, and the package declares 3.14 support. A new test reads zstd-compressed Inspect `.eval` archives directly on 3.14 and checks they give the same diff as the JSON logs. diff --git a/CITATION.cff b/CITATION.cff index 2f0b2c0..a5b5b70 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -4,8 +4,8 @@ title: "EvalArc: Auditable Evaluations for AI Agents" type: software authors: - name: EvalArc contributors -version: 0.17.4 -date-released: 2026-10-02 +version: 0.17.5 +date-released: 2026-10-05 license: MIT repository-code: "https://github.com/noteflowai/evalarc" url: "https://noteflowai.github.io/evalarc/" diff --git a/README.md b/README.md index 03962bf..73248d7 100644 --- a/README.md +++ b/README.md @@ -66,13 +66,13 @@ and the upgrade gate fails. These checks share two format/schema failures. installs the published wheel, downloads the records and rebuilds the comparison. Verification exits 0 for consistency; comparison exits 1 for the regression. -Install the released reviewer from [PyPI](https://pypi.org/project/evalarc/0.17.4/) +Install the released reviewer from [PyPI](https://pypi.org/project/evalarc/0.17.5/) in a fresh virtual environment: ```bash python3 -m venv .venv . .venv/bin/activate -python -m pip install evalarc==0.17.4 +python -m pip install evalarc==0.17.5 evalarc --version ``` @@ -124,7 +124,7 @@ because checks missing from it cannot be compared; rerun that evaluation to completion and compare again. ```yaml -- uses: noteflowai/evalarc@v0.17.4 # or a full commit SHA +- uses: noteflowai/evalarc@v0.17.5 # or a full commit SHA with: baseline: evals/baseline.json current: results/current.json diff --git a/README.zh-CN.md b/README.zh-CN.md index becda52..c964d81 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -52,13 +52,13 @@ EvalArc 是面向 Agent 评测的 Python 命令行复核工具:逐项对比已 3. **本地复算。** [首次复核指南](docs/first-review.zh-CN.md)使用发布的 wheel 和原始记录重建报告; 复核退出 0 表示一致,对照退出 1 表示发现退步。 -在新的虚拟环境中,从 [PyPI](https://pypi.org/project/evalarc/0.17.4/) +在新的虚拟环境中,从 [PyPI](https://pypi.org/project/evalarc/0.17.5/) 安装已发布的复核工具: ```bash python3 -m venv .venv . .venv/bin/activate -python -m pip install evalarc==0.17.4 +python -m pip install evalarc==0.17.5 evalarc --version ``` @@ -102,7 +102,7 @@ PyPI 与 GitHub Releases 提供相同的 wheel 和源码包,也支持 准确率从 0.625 升至 0.8125,同时有三个检查项丢失通过。 ```yaml -- uses: noteflowai/evalarc@v0.17.4 # 或完整的提交 SHA +- uses: noteflowai/evalarc@v0.17.5 # 或完整的提交 SHA with: baseline: evals/baseline.json current: results/current.json diff --git a/docs/ci-gate.md b/docs/ci-gate.md index ca05c30..d725bce 100644 --- a/docs/ci-gate.md +++ b/docs/ci-gate.md @@ -84,7 +84,17 @@ status, such as JUnit, are never marked incomplete. Exit codes: **0** no check lost passes, **1** at least one did, **2** the files could not be read or compared (different formats, different Inspect task names, -no samples, malformed input). Treat 2 as a broken pipeline, not a regression. +no samples, malformed input, an Inspect sample repeated in the same epoch). +Treat 2 as a broken pipeline, not a regression. + +An Inspect log that contains the same sample ID twice in one epoch (an integer +`7` and a string `"7"` count as the same ID) exits 2 and names the sample +and epoch. Counting the copy as an extra attempt would overstate coverage and +could hide a `less_covered` change. Such logs usually come from merging, +concatenating or editing results; re-export with +`inspect log dump LOG.eval > LOG.json` or rerun the evaluation. The same ID in +different epochs is a normal repeated attempt, and samples without an integer +`epoch` are not checked. ## Sampling-noise annotation diff --git a/docs/ci-gate.zh-CN.md b/docs/ci-gate.zh-CN.md index e83a40b..2e1b3e5 100644 --- a/docs/ci-gate.zh-CN.md +++ b/docs/ci-gate.zh-CN.md @@ -46,7 +46,12 @@ epoch 中都对重复订单退款(`refund-duplicate` 退化),`cancel-pendi 门禁以 `less_covered` 失败,直到尝试次数与基线一致或提交新的基线;通过率下降时仍优先报告 `regressed` 或 `less_reliable`。删除一个失败的检查项也会失败, 否则删掉它就能让门禁变绿。退出码:**0** 无检查项丢失通过,**1** 至少一项丢失, -**2** 文件无法读取或无法比较(格式不同、Inspect 任务名不同、没有样本、格式错误)。 +**2** 文件无法读取或无法比较(格式不同、Inspect 任务名不同、没有样本、格式错误、 +Inspect 样本在同一 epoch 中重复)。同一 epoch 中重复出现的样本 ID(整数 `7` 与字符串 +`"7"` 视为同一 ID)会以退出码 2 报告并给出样本与 epoch:若把副本算作额外尝试,会夸大覆盖并 +掩盖 `less_covered`。这类日志通常来自手工合并、拼接或编辑;请用 +`inspect log dump LOG.eval > LOG.json` 重新导出,或重新运行评测。同一 ID 出现在不同 epoch +是正常的重复尝试;没有整数 `epoch` 的样本不做此检查。 2 表示流水线本身有问题,不是退化。 ## 采样噪声标注 diff --git a/package-lock.json b/package-lock.json index d9adee8..b2ee151 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "evalarc-evidence-site", - "version": "0.17.4", + "version": "0.17.5", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "evalarc-evidence-site", - "version": "0.17.4", + "version": "0.17.5", "devDependencies": { "@axe-core/playwright": "4.13.0", "playwright": "1.63.0", diff --git a/package.json b/package.json index f728a75..68697e2 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "evalarc-evidence-site", - "version": "0.17.4", + "version": "0.17.5", "private": true, "description": "Browser checks and the TypeScript report front end for EvalArc", "scripts": { diff --git a/pyproject.toml b/pyproject.toml index a584a85..621a433 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "evalarc" -version = "0.17.4" +version = "0.17.5" description = "Find agent regressions behind better scores: compare changed checks, review recorded tool actions and verify evidence offline." readme = "README.md" requires-python = ">=3.11" diff --git a/src/evalarc/__init__.py b/src/evalarc/__init__.py index 13a97f3..fe84c97 100644 --- a/src/evalarc/__init__.py +++ b/src/evalarc/__init__.py @@ -1,3 +1,3 @@ """Auditable evaluations for AI agents.""" -__version__ = "0.17.4" +__version__ = "0.17.5" diff --git a/src/evalarc/results_diff.py b/src/evalarc/results_diff.py index 3135443..c7b44ac 100644 --- a/src/evalarc/results_diff.py +++ b/src/evalarc/results_diff.py @@ -439,8 +439,22 @@ def _inspect(document: dict, threshold: float) -> dict: usage: dict[str, list] = {} outputs: dict[str, dict[str, list]] = {} samples_meta: dict[str, list] = {} + seen: set[tuple[str, int]] = set() for sample in document["samples"]: case_id = str(sample["id"]) + epoch = sample.get("epoch") + if isinstance(epoch, int) and not isinstance(epoch, bool): + # A repeated (id, epoch) pair would silently count as an extra attempt and + # overstate coverage. Ids 7 and "7" collide because case_id is str(id). + key = (case_id, epoch) + if key in seen: + raise ValueError( + f"Inspect log has duplicate sample {_shown_id(sample['id'])} epoch {epoch}: " + "each sample and epoch must appear once, or the copy would count as an " + "extra attempt. Re-export the log from Inspect (inspect log dump) instead " + "of merging or editing it." + ) + seen.add(key) checks = cases.setdefault(case_id, {}) _add_material(material, case_id, "input", _message_text(sample.get("input"))) usage.setdefault(case_id, []).append(inspect_usage(sample)) @@ -505,6 +519,12 @@ def _inspect(document: dict, threshold: float) -> dict: } +def _shown_id(value: object, limit: int = 80) -> str: + """Render an untrusted sample id safely: repr() escapes control characters.""" + text = repr(value) + return text if len(text) <= limit else text[: limit - 3] + "..." + + def _inspect_passed(value: object, threshold: float) -> bool | None: if isinstance(value, bool): return value diff --git a/tests/test_feature_e0a029ce3b1b.py b/tests/test_feature_e0a029ce3b1b.py new file mode 100644 index 0000000..123f656 --- /dev/null +++ b/tests/test_feature_e0a029ce3b1b.py @@ -0,0 +1,111 @@ +"""Duplicate Inspect samples (same id and epoch) are rejected, not counted as attempts. + +All logs here are synthetic copies of the recorded examples, written by the tests. +""" + +import copy +import json +import zipfile +from pathlib import Path + +import pytest + +from evalarc.cli import main +from evalarc.results_diff import diff, load_results + +EXAMPLES = Path(__file__).resolve().parents[1] / "examples/results-diff" + + +def _load(name): + return json.loads((EXAMPLES / f"inspect/{name}.json").read_text()) + + +def _write(path, document): + path.write_text(json.dumps(document)) + return path + + +def _scored(template, sample_id, epoch): + sample = copy.deepcopy(template) + sample["id"] = sample_id + if epoch is None: + sample.pop("epoch", None) + else: + sample["epoch"] = epoch + sample.pop("error", None) + sample["scores"] = {"match": {"value": "C"}} + return sample + + +@pytest.mark.parametrize("side", ["baseline", "current"]) +def test_duplicated_sample_exits_2_and_names_it(tmp_path, capsys, side): + document = _load(side) + duplicate = copy.deepcopy(document["samples"][0]) + document["samples"].append(duplicate) + broken = _write(tmp_path / f"{side}-dup.json", document) + paths = { + "baseline": str(EXAMPLES / "inspect/baseline.json"), + "current": str(EXAMPLES / "inspect/current.json"), + } + paths[side] = str(broken) + output = tmp_path / "review" + code = main(["diff", paths["baseline"], paths["current"], "--output", str(output)]) + assert code == 2 + err = capsys.readouterr().err + assert "duplicate sample" in err + assert repr(duplicate["id"]) in err + assert f"epoch {duplicate['epoch']}" in err + assert "inspect log dump" in err + assert not (output / "diff.json").exists() + + +def test_integer_and_string_ids_collide(tmp_path, capsys): + document = _load("baseline") + template = document["samples"][0] + document["samples"] = [_scored(template, 7, 1), _scored(template, "7", 1)] + path = _write(tmp_path / "mixed.json", document) + assert main(["diff", str(EXAMPLES / "inspect/baseline.json"), str(path)]) == 2 + err = capsys.readouterr().err + assert "duplicate sample" in err and "epoch 1" in err + + +def test_eval_archive_with_duplicate_members_is_rejected(tmp_path): + document = _load("current") + samples = document.pop("samples") + archive = tmp_path / "current.eval" + with zipfile.ZipFile(archive, "w", zipfile.ZIP_DEFLATED) as output: + output.writestr("header.json", json.dumps(document)) + output.writestr("samples/a.json", json.dumps(samples[0])) + output.writestr("samples/b.json", json.dumps(samples[0])) + with pytest.raises(ValueError, match="duplicate sample"): + load_results(archive) + + +def test_control_characters_in_id_are_escaped(tmp_path): + document = _load("baseline") + template = document["samples"][0] + name = "bad\nid" + "x" * 200 + document["samples"] = [_scored(template, name, 1), _scored(template, name, 1)] + with pytest.raises(ValueError) as caught: + load_results(_write(tmp_path / "ctrl.json", document)) + message = str(caught.value) + assert "\n" not in message + assert "x" * 100 not in message + + +def test_valid_logs_are_unchanged(tmp_path): + result = diff( + load_results(EXAMPLES / "inspect/baseline.json"), + load_results(EXAMPLES / "inspect/current.json"), + ) + assert result["blocking_changes"] == 3 and not result["gate_passed"] + + document = _load("baseline") + template = document["samples"][0] + document["samples"] = [_scored(template, "repeat", epoch) for epoch in (1, 2, 3)] + loaded = load_results(_write(tmp_path / "epochs.json", document)) + assert len(loaded["cases"]["repeat"]["match"]) == 3 + + document["samples"] = [_scored(template, "no-epoch", None) for _ in range(2)] + loaded = load_results(_write(tmp_path / "no-epoch.json", document)) + assert len(loaded["cases"]["no-epoch"]["match"]) == 2