Skip to content

Exercise the CUDA comparability warning - #17

Merged
emrefbulut merged 1 commit into
mainfrom
fix/cuda-warning-test
Sep 13, 2026
Merged

emrefbulut merged 1 commit into
mainfrom
fix/cuda-warning-test

Conversation

@emrefbulut

Copy link
Copy Markdown
Owner

What changed

Two tests covering the CUDA comparability warning in iqforge train, which no test had ever executed.

Why

README.md and CONTRIBUTING.md both promise it:

A CUDA run prints a warning that its numbers are not bit-comparable with CPU runs

The line exists (src/iqforge/cli.py:1055) and nothing ran it. It is the only thing that tells a CUDA user their numbers cannot be placed next to a CPU table, and a promise nothing exercises is a promise nobody has checked.

How it was verified

No GPU needed, and none wanted. torch.cuda.is_available is pinned True so resolve_device("cuda") and describe_environment run for real and produce a genuine cuda stamp; train_baseline is replaced so nothing reaches for a CUDA context this machine does not have. Mocking availability is not the same as having a device — manual_seed_all and Adam would find that out, which is why the existing test_models.py tests stop at placement too.

What is under test is the CLI branch reading result.environment["device"] — the line that decides whether a CUDA user is told anything.

Both halves of the branch are covered. A warning printed unconditionally would satisfy the positive test while telling every CPU user their numbers are not comparable, so the CPU case asserts its absence.

Mutations, each red on the right test:

mutation caught by
delete the warning (if False:) test_a_cuda_run_warns_that_its_numbers_are_not_comparable
print it unconditionally (if True:) test_a_cpu_run_prints_no_comparability_warning
ignore --device, always pass cpu test_a_cuda_run_warns_that_its_numbers_are_not_comparable
uv run pytest tests/test_cli_output.py -k comparab -q   # 2 passed
uv run ruff check .                                     # clean

One measurement worth recording. The assertions compare with all whitespace removed. rich wraps at a console width this test cannot set, and it strips the space it wrapped on — which turns "Do not mix" into "Do notmix". The existing _flat helper drops newlines but does not restore that space. Found by the assertion failing, not predicted.

Noticed, not fixed here

Two gaps in the same feature, both out of scope for a test-only PR:

  1. measure-leakage --device cuda prints no such warning. Only train has it. A CUDA leakage measurement emits environment: device cuda | ... and nothing else — and that command is the one whose output gets pasted into a table.
  2. The warning sits after the empty-test-split early return (if result.test_accuracy is None: return), so a CUDA run with an empty test split trains on the GPU and exits without warning.

Conventions it touches

  • 2 A passing test does not prove it can fail — three mutations, each confirmed red on the test that should catch it.
  • 3 Do not claim what you did not measure — the whitespace behaviour above was measured, and the two gaps are reported rather than left for someone to discover.

README and CONTRIBUTING both promise that a CUDA run prints a warning that
its numbers are not bit-comparable with CPU runs. The line existed in
`train` and no test had ever executed it. A promise nothing exercises is a
promise nobody has checked.

No GPU is needed and none is wanted. torch.cuda.is_available is pinned True
so resolve_device and describe_environment run for real and produce a
genuine cuda stamp, and train_baseline is replaced so nothing reaches for a
CUDA context this machine does not have -- mocking availability is not the
same as having a device, and manual_seed_all and Adam would find that out.
What is under test is the CLI branch that reads
result.environment["device"], which is what decides whether a CUDA user is
told anything at all.

Both halves of the branch are covered. A warning printed unconditionally
would satisfy the positive test while telling every CPU user their numbers
are not comparable, so the CPU case asserts the warning is absent.

Mutations run, each red on the test that should catch it: deleting the
warning fails the CUDA test; printing it unconditionally fails the CPU test;
ignoring --device and always passing cpu fails the CUDA test.

The assertions compare with all whitespace removed. rich wraps at a console
width this test cannot set, and it strips the space it wrapped on, which
turns "Do not mix" into "Do notmix" -- measured while writing this, not
assumed. `_flat` alone was not enough.
@emrefbulut
emrefbulut merged commit 60212f1 into main Sep 13, 2026
5 checks passed
emrefbulut added a commit that referenced this pull request Sep 14, 2026
The stacked PRs #18, #19 and #20 were merged at the same moment, so only #18
reached main -- #19 merged into fix/gate-semantics-and-provenance and #20 into
docs/refresh, their own bases. Everything from those two is therefore sitting
on this branch and main is still on 0.4.0. That was my mistake in stacking
three deep instead of retargeting; this merge is the fix.

One conflict, in CHANGELOG.md. main carried the entry describing
docs/release-notes/v0.5.0.md as an unpublished draft, which commit 6491f27 on
this branch had already rewritten into "the four places that carry a version
agree again" -- the draft label was the state this release ends. Kept this
branch's side, which also carries three entries main does not have.

Checked rather than assumed, because a merge across branches that both edited
the test files could silently drop one side: #17's CUDA comparability tests
and #16's forced-measurement width test are both present afterwards, and the
suite goes from 385 to 388 passing, which is main's three additions arriving
rather than anything being lost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant