Repository navigation
fix: reject duplicate Inspect samples instead of counting them as extra attempts - #69
Merged
noteflowai merged 1 commit intoOct 5, 2026
Conversation
noteflowai
deleted the
automation/feature-2026-10-05-evalarc-134110770388
branch
October 5, 2026 15:06
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Users who gate a PR with evalarc diff on Inspect AI logs can get a misleading verdict when a log contains the same sample twice in one epoch. This happens when logs are merged by hand, concatenated by a script or edited. Today the copy becomes an extra attempt. A check recorded as 2/2 can read as 3/3, coverage can look equal or better when an attempt is actually missing, and less_covered never fires. An integer id and a string id with the same text, such as 7 and '7', collide in the same way, because case_id is str(sample['id']). Roadmap EA-02 names this missing-evidence defect explicitly (duplicate IDs, malformed records). The current code and docs have no check or message for it.
Scope: the Inspect AI reader in src/evalarc/results_diff.py (_inspect). Both JSON logs and .eval zip archives reach it through _inspect_document. Entry point: evalarc diff BASELINE CURRENT, and the same load_results() API that the tests and the Action call.
Input contract: while iterating document['samples'], the reader builds a key (case_id, epoch). case_id is the existing str(sample['id']) normalization. epoch is sample['epoch'] when it is an int and not a bool. If a key has already been seen, load_results raises ValueError with this message: 'Inspect log has duplicate sample epoch : each sample and epoch must appear once, or the copy would count as an extra attempt. Re-export the log from Inspect (inspect log dump) instead of merging or editing it.' The id comes from untrusted input, so it is shown with repr() and truncated to 80 characters. A newline or control character in it therefore cannot break terminal or summary output. Samples whose epoch is missing or not an integer are not checked, so older or minimal logs load as they do today. Ids 7 and '7' in the same epoch count as a duplicate, because they would otherwise merge into one case.
Output contract: the CLI's existing ValueError path prints the message to stderr and exits 2. Exit 2 already means a broken pipeline, not a regression. The input is rejected while loading, before any comparison, so diff.json, summary.md and index.html are not produced for that pair. A duplicate in either the baseline or the current file triggers the error. The diff.json schema, KINDS, BLOCKING, counts and the gate_passed logic do not change. The same sample id in different epochs stays valid and still forms repeated attempts. The sampling-noise annotation, less_covered and the incomplete-baseline handling keep working as before.
Error cases: a .eval archive with two samples/*.json members that share an id and epoch is caught by the same check, because archive members are loaded into document['samples'] first. The missing-header, zstd and BadZipFile messages are unchanged. Repeats in promptfoo and repeated JUnit cases are intentional, documented repeated attempts and are not touched.
Documentation: docs/ci-gate.md and docs/ci-gate.zh-CN.md add duplicate Inspect samples to the exit-2 list. They add one sentence on why counting a copy as an attempt would misstate coverage, and give the recovery step: re-export with inspect log dump, or rerun.
Acceptance:
One specialist feature. Behavioral acceptance fails on the base and passes on the implementation; project checks and independent model review passed. GPU results are included only when actually executed. Automatically developed and reviewed; no human review is claimed.