ci: nightly MSan over the Boost suites with an instrumented libc++ - #177
Conversation
TEMP COMMIT: carries a pull_request trigger so the job can be verified before merge; the trigger is removed in a follow-up commit on this branch.
libc++abi defaults to LIBCXXABI_USE_LLVM_UNWINDER=ON, so configuring without libunwind in LLVM_ENABLE_RUNTIMES fails outright: LIBCXXABI_USE_LLVM_UNWINDER is set to ON, but libunwind is not specified in LLVM_ENABLE_RUNTIMES. Building it here also instruments it, which matters on its own: an uninstrumented unwinder linked into an MSan binary is another source of phantom reports, which is the problem this job exists to escape.
The first full run got through the instrumented libc++ build and the teeth check, then failed in unit_test/nn -- with zero MemorySanitizer reports. The two failures were Boost assertions: test_case_lstm_weight_serialization 0.023587908467455981 >= 0.02 test_case_rmsprop_fixedpoint_xor average error 25 > 4 Both are training tests whose result depends on the initial random weights. std::uniform_real_distribution is not required to generate the same sequence across implementations and libc++ and libstdc++ differ, so swapping the C++ runtime changes the weights and moves trained results outside tolerances that were tuned against libstdc++. That is a property of the runtime swap rather than a defect, and failing on it would leave this job permanently red for a question it was never asking. The criterion is now explicit: any MemorySanitizer report fails; Boost's exit 201 is noted and tolerated; any other non-zero exit still fails, so a crash or signal is not swallowed. The suite output is kept as msan-libcxx.log for the artifact upload. Worth recording the negative result: with an instrumented libc++, MSan found no uninitialized reads in unit_test/nn across its 250 cases.
The tolerance branch never fired. Boost.Test exits 201, but the suite runs through a sub-make and make reports its own failure as 2, so the 201 comparison matched nothing and unit_test/nn's numeric divergence was misreported as "exited 2 (crash or signal, not a test assertion)". Match on Boost's "failures are detected" in the captured output instead, which is what actually distinguishes an assertion failure from a crash. The rest of the run is a clean result: with the cached instrumented libc++, all ten Boost suites executed and only unit_test/nn diverged numerically. No MemorySanitizer report anywhere.
Verification is done. The job was exercised on this PR across four runs, which is what the trigger existed for: 1. libc++ configure failed -- libunwind missing from LLVM_ENABLE_RUNTIMES 2. libc++ built, teeth check passed, unit_test/nn failed on numeric tolerance 3. tolerance branch never fired -- matched Boost's 201 through a sub-make 4. green: teeth check passes, all ten suites run, no MemorySanitizer reports Merging this unverified would have left a nightly failing silently at 06:00 UTC until somebody noticed a red scheduled run. Back to schedule + workflow_dispatch only.
|
Verified and ready. The temporary Four runs on this PR, which is exactly what the trigger existed for:
Final run output: The result is a clean negative: with an instrumented libc++, MSan finds no uninitialized reads anywhere in the ten Boost suites. That is the coverage the #171 static-init bug slipped through, now watched nightly. The Merging this unverified would have left a nightly failing silently at 06:00 UTC until someone noticed a red scheduled run. |
Closes the first of the two gaps flagged after #175 — MSan coverage for the Boost suites.
Why
The blocking
MSan (uninitialized reads)job coversunit_test/embeddedonly. That scope was never a preference — MSan requires every linked object to be instrumented, and against a stock libstdc++ the Boost suites abort inside Boost.Test's own static initialization before reaching any TinyMind code. Those reports are artifacts of the uninstrumented runtime, not findings.That gap is exactly where the real bug lived. The static-init-order UB fixed in #171 — a Q-learning suite training on all-zero weights and "passing" in 17ms — sat in a Boost suite, invisible to ASan, UBSan, and the embedded MSan job alike.
How
Builds
libc++andlibc++abifrom llvm-project with-DLLVM_USE_SANITIZER=MemoryWithOrigins, caches by LLVM version, and runs the ten Boost suites against it viamake check-msan-libcxx.Only the runtimes are built, not the compiler — the distro clang does the compiling and is merely given an instrumented libc++ to link against. That keeps a cold build in minutes rather than hours.
Nightly, not a PR gate. The cold build is 10–20 minutes, which is the wrong shape for something blocking every merge and the right shape for a scheduled run; the suites themselves take seconds once the runtime exists.
Boost.Test needs no separate handling: the suites include the header-only
<boost/test/included/unit_test.hpp>, so it compiles into each TU and is instrumented with it. The C++ runtime was the only uninstrumented piece.The job fails itself if MSan isn't working
A sanitizer job that reports nothing is indistinguishable from one that is broken, and this series has already produced two such false greens (stale artifacts in #168, a
-maxdepththat silently skipped two directories in #175).So before trusting anything the real run says, the job plants an uninitialized read and requires MSan to catch it:
That encodes the
-O0finding from #172 as a gate rather than a comment, so a later flag change cannot quietly hollow it out.Verification
workflow_dispatchonly fires for workflows already on the default branch, so a temporarypull_requesttrigger is the only way to exercise this before merging. Merging it unverified would contradict the whole point of the teeth check.Plan: let this run on the PR, confirm the libc++ build, teeth check and suites all pass, then strip the trigger and merge.
🤖 Generated with Claude Code
https://claude.ai/code/session_019tVgXeXCcfbHqfMjhufWzd