fix: fail the diff gate when the baseline run did not finish - #62
Merged
noteflowai merged 1 commit intoOct 1, 2026
Merged
Conversation
noteflowai
deleted the
automation/feature-2026-10-01-evalarc-140131148828
branch
October 1, 2026 14:41
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An engineer who gates a pull request with evalarc diff can get a passing gate from an interrupted baseline. Checks missing from the unfinished baseline are classified as 'added', which never blocks, so the report reads as extra coverage instead of missing evidence. diff() reads only current["incomplete"]. None of the three outputs (terminal, Markdown, HTML) has a message for an unfinished baseline, and no test covers the case.
Scope: one change to the existing evalarc diff comparison and its three existing renderers (CLI stdout, render_markdown, render_html). Code lives in src/evalarc/results_diff.py and src/evalarc/cli.py. There is no new command, flag, input format or mode.
Input: the two result files evalarc diff already accepts. The loader already gives each run 'incomplete' and identity.status. For Inspect, incomplete is true when status is not None and not 'success'.
Computation in diff(baseline, current): read baseline["incomplete"]. When it is true:
Classification, counts, blocking_changes, noise and grader-conflict annotations stay unchanged. Checks missing from the baseline are still kind 'added', so the schema-v1 kinds stay valid. When the baseline is complete, the key is omitted. This is the same additive convention less_covered uses, so earlier reports and verify recomputation stay byte-compatible. evidence.diff_passed already starts from gate_passed, so the exit code and the Action fail without extra wiring. evalarc verify recomputes through compute_diff and gets the same value.
Terminal (cli._results_diff, default non --json output): next to the existing current-incomplete line, print:
'The baseline run is incomplete (status error); N check(s) appear only in the current run and were not compared. Rerun the baseline to completion; the gate fails.'
N is counts['added'], and the status comes from result['baseline']['identity']['status']. The existing summary line and the blocking rows still print first. --json output gains only the new key.
Markdown (summary.md, --markdown, PR comment):
HTML (render_html, index.html in --output folders):
Exit codes: 0 when the gate passes; 1 when it fails, which now includes an unfinished baseline; 2 for unreadable, incomparable or invalid input or an existing output folder, as today.
Remedy: rerun or replace the baseline with a finished evaluation. EvalArc does not reconstruct samples, rerun anything or call a model.
Docs: docs/ci-gate.md gains an 'Unfinished baseline' paragraph with the message, exit code 1 and the remedy. The README CI-gate paragraph states that an unfinished baseline or current run fails the gate.
Acceptance:
One specialist feature. Behavioral acceptance fails on the base and passes on the implementation; project checks and independent model review passed. GPU results are included only when actually executed. Automatically developed and reviewed; no human review is claimed.