Skip to content

fix: fail the diff gate when the baseline run did not finish - #62

Merged
noteflowai merged 1 commit into
mainfrom
automation/feature-2026-10-01-evalarc-140131148828
Oct 1, 2026
Merged

noteflowai merged 1 commit into
mainfrom
automation/feature-2026-10-01-evalarc-140131148828

Conversation

@noteflowai

Copy link
Copy Markdown
Owner

An engineer who gates a pull request with evalarc diff can get a passing gate from an interrupted baseline. Checks missing from the unfinished baseline are classified as 'added', which never blocks, so the report reads as extra coverage instead of missing evidence. diff() reads only current["incomplete"]. None of the three outputs (terminal, Markdown, HTML) has a message for an unfinished baseline, and no test covers the case.

Scope: one change to the existing evalarc diff comparison and its three existing renderers (CLI stdout, render_markdown, render_html). Code lives in src/evalarc/results_diff.py and src/evalarc/cli.py. There is no new command, flag, input format or mode.

Input: the two result files evalarc diff already accepts. The loader already gives each run 'incomplete' and identity.status. For Inspect, incomplete is true when status is not None and not 'success'.

Computation in diff(baseline, current): read baseline["incomplete"]. When it is true:

  • add "baseline_incomplete": true to the result
  • set gate_passed to blocking == 0 and not current_incomplete and not baseline_incomplete

Classification, counts, blocking_changes, noise and grader-conflict annotations stay unchanged. Checks missing from the baseline are still kind 'added', so the schema-v1 kinds stay valid. When the baseline is complete, the key is omitted. This is the same additive convention less_covered uses, so earlier reports and verify recomputation stay byte-compatible. evidence.diff_passed already starts from gate_passed, so the exit code and the Action fail without extra wiring. evalarc verify recomputes through compute_diff and gets the same value.

Terminal (cli._results_diff, default non --json output): next to the existing current-incomplete line, print:
'The baseline run is incomplete (status error); N check(s) appear only in the current run and were not compared. Rerun the baseline to completion; the gate fails.'
N is counts['added'], and the status comes from result['baseline']['identity']['status']. The existing summary line and the blocking rows still print first. --json output gains only the new key.

Markdown (summary.md, --markdown, PR comment):

  • The verdict chain becomes: gate passed -> 'No check lost passes'; blocking > 0 -> the existing blocking text; both incomplete -> 'the baseline and current runs are incomplete'; baseline only -> 'the baseline run is incomplete'; otherwise the existing current text.
  • Below the table, add the line: 'The baseline run did not finish (status \u0060error\u0060); the gate fails because checks missing from it cannot be compared. N check(s) appear only in the current run. Rerun the baseline to completion.'
  • The status value goes through the existing _cell escaping. The output is plain text and does not depend on colour.

HTML (render_html, index.html in --output folders):

  • The existing reasons list gains 'the baseline run is incomplete', so the existing _verdict banner reads 'Gate failed' with that reason.
  • When baseline_incomplete is true and blocking_changes is 0, the action text is 'Rerun the baseline to completion and compare again.' instead of 'Read the blocking rows below...', which would point at rows that do not exist. When there are blocking changes, the existing action text is kept.
  • The HTML uses fixed text only, never the raw status string, so no untrusted value enters the markup through this change. The existing verdict component, styles and contrast are reused, with no layout change.

Exit codes: 0 when the gate passes; 1 when it fails, which now includes an unfinished baseline; 2 for unreadable, incomparable or invalid input or an existing output folder, as today.

Remedy: rerun or replace the baseline with a finished evaluation. EvalArc does not reconstruct samples, rerun anything or call a model.

Docs: docs/ci-gate.md gains an 'Unfinished baseline' paragraph with the message, exit code 1 and the remedy. The README CI-gate paragraph states that an unfinished baseline or current run fails the gate.

Acceptance:

  • The baseline is a copy of examples/results-diff/inspect/baseline.json with status 'error'. Current is the unmodified baseline.json. diff() returns changes == [], baseline_incomplete is True, current_incomplete is False and gate_passed is False. render_markdown contains 'EvalArc: the baseline run is incomplete' and 'The baseline run did not finish (status \u0060error\u0060)'.
  • The baseline is a copy of examples/results-diff/inspect/baseline.json with status 'error' and every refund-duplicate sample removed. Current is the unmodified baseline.json, which has status success. The only changes are ('added','refund-duplicate','includes') and ('added','refund-duplicate','match'), blocking_changes == 0, counts['added'] == 2 and gate_passed is False. The Markdown states '2 check(s) appear only in the current run'.
  • When both inputs are complete, the result has no 'baseline_incomplete' key, and gate_passed and counts match existing behavior. The existing tests in tests/test_results_diff.py, including identical inputs, removed/added and the incomplete current run, pass unchanged.
  • CLI acceptance test (tests/test_feature_f107c0a41d3a.py): main(['diff', , <baseline.json>, '--output', ]) returns 1. capsys stdout contains 'The baseline run is incomplete (status error); 2 check(s) appear only in the current run' and 'Rerun the baseline to completion'. diff.json has baseline_incomplete true. index.html contains 'Gate failed', 'the baseline run is incomplete' and 'Rerun the baseline to completion and compare again', and does not contain 'Gate passed'. evidence.verify_report on the folder reports a successful recomputation.
  • docs/ci-gate.md documents the unfinished-baseline message, exit code 1 and the remedy of rerunning the baseline. The full regression gates (pytest, ruff check, ruff format --check, site build, node site check) pass.

One specialist feature. Behavioral acceptance fails on the base and passes on the implementation; project checks and independent model review passed. GPU results are included only when actually executed. Automatically developed and reviewed; no human review is claimed.

@noteflowai
noteflowai merged commit 99a159f into main Oct 1, 2026
14 checks passed
@noteflowai
noteflowai deleted the automation/feature-2026-10-01-evalarc-140131148828 branch October 1, 2026 14:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant