Skip to content

Make the three-build comparison of decision 0010 runnable by anyone - #11

Merged
marcos-mendez merged 9 commits into
19.xfrom
feat/layer-measure
Sep 29, 2026
Merged

marcos-mendez merged 9 commits into
19.xfrom
feat/layer-measure

Conversation

@marcos-mendez

@marcos-mendez marcos-mendez commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #10. Paired with Keel-Linux/handbook#8, which amends decision 0010 to name
this command and reports what a recount of the two merged extractions found
with it.

Handbook decision 0010 accepts the component split on one condition, stated in
step 3 of its order of work: components move one at a time, "re-running this
three-build comparison for each, including the control-vs-control run".
Nothing in this organization ran it. The capture half was
/root/measure/build.sh on the build host, a scratch file in root's home
directory on one machine; the comparison half, the part that turns captures
into the number a pull request quotes, was ad hoc shell that was never
committed at all. Keel-Linux/keel-mariadb#11 and Keel-Linux/keel-postgresql#7
were both approved on a count only their author could produce.

bt-layer-measure capture   --dir DIR --layer NAME [--parent NAME] --tag TAG
bt-layer-measure compare   --dir DIR --a TAG --b TAG
bt-layer-measure attribute --dir DIR --unit TAG --control TAG
                           --control-again TAG [--control-again TAG]...
                           [--cleared FILE]

--dir is required by every subcommand and defaults to nothing, so the
directory holding the measurement is part of the command and a second person
re-runs the same command against the same directory. No state in a home
directory on one machine. capture takes /run/lock/keel-build.lock with a
bounded wait. A tag is never overwritten. The full procedure is in
docs/layer-measure.md.

Why this repository and not .github

.github/lib lends shell to the organization's reusable workflows; this is not
that. The capture half runs bt-layer and reads the rootfs it leaves in
$BT_BUILDS/layers, so it is written against bin/layer-lib's own vocabulary
of layers, parents, manifests and units. The build host already has this
repository at /turnkey/buildtasks-keel, which is where the three builds
happen, and .github is not checked out there at all. This repository also
already has the gate the work needs, tests/coverage.sh under kcov at a 99
percent floor, and decision 0010 step 1, making bt-layer unit-aware, landed
here in #8.

What the hand-run measurement got wrong

It counted paths it could not key on. The capture wrote sha256sum's
default <hash> plus two spaces plus <path> and the comparison split on
whitespace. A Debian rootfs holds
setuptools/_vendor/jaraco/text/Lorem ipsum.txt, which becomes two fields and
never matches itself, so the LAMP measurement of 2026-09-27 reported 199 and
183 differing files where the true counts were 198 and 182 (docs/traps.md,
"A path with a space in it"). Every record here is path<TAB>... with the path
in field 1, and a path carrying a tab, a newline or a backslash is refused at
capture time rather than mis-counted.

It was blind where it mattered most. For mariadb 173 of the 179 files in
the floor are ./var/lib/mysql/**, which holds mysql/global_priv and
mysql/user.*, and the component's build-time job includes deleting accounts.
A path is put inside the floor by its name; whether its difference is noise is
a question about its bytes.

It treated a file it could not read as a file that did not change.

The rework after review

Review of the first push found two ways this failed in the same direction as
the defect it exists to correct. Both are closed.

The signature compared positions and threw the content away

o<n>/n<n>, @<n> and a bare len are positions. Two unrelated changes at
one line number or one byte offset annihilated, and the evidence block was
suppressed for anything called noise, so the one mistake the scheme could make
was the one nobody could see. Four cases came back noise, exit 0, with
nothing printed:

  • an account added to /etc/shadow at the line number the control pair also
    appends at,
  • a root password set on the line the control pair rewrites for a clock reason,
  • a binary differing at the same eight offsets with different byte values,
  • a file 3584 bytes shorter against a control pair that varies by one.

The controls are now the model and the model carries content. At each position
the controls vary at, the longest common prefix and suffix of their variants is
the shape of that variation, and the unit is admitted only if its line has
that shape. # written 1790475089 against # written 1790475323 gives
# written ..., which # written 1790475511 fits and root:$y$hash:... does
not. A mask with neither a prefix nor a suffix is what two unrelated lines have
in common, which is no evidence about a third, so it never yields noise. Bytes
carry no shape at all, so a binary is never noise: the most that can be said is
that it coincides in position, in size delta and in how much differs, which is
the new verdict overlapping, and it fails until somebody reads the bytes and
says so in --cleared. The evidence is printed for every verdict, noise
included.

Verdict Meaning Run
noise every differing line is at a position the controls vary at and has the shape they vary with there passes
overlapping coincides with the controls' variation but cannot be shown to be that variation fails unless cleared
real differs where the controls do not vary, or with content or a magnitude they never show fails
not sampled no bytes kept on some side fails
unchanged mode differs, bytes identical passes

tests/compare is that list of four, one case each. They come out
overlapping, real, overlapping, real.

The blindness could be switched off at attribute time

MEASURE_STATE_PATHS was read when attribute ran, so a one-line substituted
list turned FAIL (167 not sampled) into PASS on the real captures while the
block still printed a digest. capture now copies the list into the capture,
attribute reads that copy and never the variable, and it refuses unless every
capture carries the same list and the digest each one recorded matches it.

The rest of the review

  • --control-again is repeatable and the floor is the union over every
    control pair. One pair is not a floor: mariadb differs by 179 files in one
    pair and 182 in another, and the residue is 3 against the first and 0 against
    the second. The block prints what each single pair would have said.
  • Preconditions are enforced. attribute refuses unless the captures share
    layer, parent and SOURCE_DATE_EPOCH, were built with an epoch at all,
    carry the same state path list, and are distinct. --control X --unit X is
    refused.
  • share/layer-state-paths gains ./etc/mysql/**, ./etc/postgresql/**,
    ./etc/redis/**, ./root/.my.cnf, ./root/.pgpass, ./root/.netrc,
    ./etc/machine-id and ./var/lib/dbus/machine-id. The postgresql residue
    lived in the second of those and was caught only because it fell outside the
    floor.
  • The list is treated as a failure mode, not a list. Every subtracted path
    is named rather than counted, and any whose name looks like state
    (share/layer-state-suspect) fails the run, with two committed ways out: add
    it to the list, or clear it with a reason.
  • Text is no longer "contains no NUL": a file with no newline, a line over
    8 KiB, or a control byte other than tab, newline and carriage return is a
    blob and is compared as bytes.
  • The verdict totals must add up to the paths examined, and an empty verdict
    is fatal rather than uncounted.
  • A capture directory is created 0700; retention is in the doc.
  • tests/layer, tests/aplinfo and tests/signature set pipefail, which
    the scripts they test already did. One assertion changed with it:
    tests/signature piped generate-signature --help into grep and usage
    exits non-zero.
  • docs/layer-measure.md exports BT_BUILDS, and the layers directory is
    derived from it through config/common.cfg the way bt-layer derives it.
  • A binary differing in more than 20000 bytes is overlapping without
    materialising the offsets: ib_logfile0 is 100 MiB.

The second rework, after re-review

The re-review found the first fix incomplete in the same direction, and it was
right twice.

The mask pinned only the two ends of a line

It constrained the common prefix and the common suffix and left the region
between them free: no length, no character class, no resemblance to anything a
control put there. So wherever the controls vary at random the mask degenerates
to the line's fixed framing, and any value carrying that framing was admitted
as noise:

planted in the unit framing the mask read as evidence old verdict
a hardcoded ssh_host_ed25519_key.pub ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAI / root@keel noise
a baked-in root password hash root:$y$j9T$ / :20000:0:99999:7::: noise
a free region grown by 400 characters token= / empty noise
timestamp=0001790574 OR 1=1; GRANT ALL ON *.* TO evil@localhost, on the real mysql/user.frm of the gate run timestamp=0001790574 / empty noise

The fix is the rule already written in the file applied to the other branch.
bin/layer-compare-lib argues that bytes carry no shape, so a binary
coincidence is never noise. A per-build random text field carries no more
shape than a random byte field.
So noise now requires the controls to pin
the free region: they must agree on its length, it must be no longer than
MEASURE_FREE_MAX (24), and every character must be in the class they used
there, which may only be digits and the punctuation of a date. A line a
control itself produced is admitted whatever its shape, because that is a
value the recipe makes on its own. All five cases above now fail.

This also removes the dependence you named on when the gate ran. The
user.frm mask is 21 characters of prefix only because four builds landed
inside the same thousand seconds, and across a digit carry it collapses to 9.
tests/compare runs exactly that: at either prefix length the clock is
noise and the payload is not, because the criterion is now the field's width
and class rather than how much framing happened to survive.

A cleared path printed no verdict line and no evidence

The waiver branch returned before both. A cleared path now prints its verdict,
the rule and justification that cleared it, and its bytes, always and not
subject to the abbreviation cap. It is the one place a person overrides the
instrument, so it is the last place to withhold the evidence.

The waiver was unbounded and unrecorded

A rule is now five tab separated fields, and share/layer-waivers.example is
the template, committed:

Field What it prevents
rule a rule with too little literal path is refused: * matches ./etc/shadow
count the paths it covered when written. A run where it covers a different number fails, so a new file in a waived directory cannot be swept in
max-position the furthest offset or line it was justified against. A covered path differing beyond it is not waived
justified-against the capture directory and recipe commit, so it cannot outlive its evidence
reason what was read and what it was found to be

The block prints the file's digest, every rule, how many paths each covered
against how many it was written for, and every cleared path with its bytes.

Your composition, as a regression test

tests/measure builds it: three controls with a random host key each, a unit
with a hardcoded one, and a waiver naming only ./var/lib/mysql/**. It fails,
the key is named as overlapping, and the waiver covers only the two database
paths it was written for:

  overlapping ./etc/ssh/ssh_host_ed25519_key.pub
              line 1 has the framing the controls vary within, and nothing pins what it put between: 21 characters, not a number or a timestamp; read the evidence
  cleared     ./var/lib/mysql/mysql/global_priv.MAD
              cleared by rule ./var/lib/mysql/**, justified against fixtures, recipe deadbeef: database creation stamps, read on 2026-09-28
  totals: noise 1, overlapping 1, real 0, not sampled 0, unchanged 0, cleared 2
waivers       .../db-waiver (sha256 20e8193e47ddc999)
  rule ./var/lib/mysql/**: 2 paths (written against 2, max position 99999)
verdict: FAIL

The third rework, after the third review

Two findings, one of them a line I had written to be helpful that inverted the
property the tool exists to have.

The control-produced exemption

measure_mask_admits admitted a unit line identical to any control's line
at that position, ahead of every pin. The base capture is the first control, so
a value copied from the second or third differs from the base, is flagged as a
differing position, and was then waved through:

unit ships was is now
the SSH host key control 2 produced noise overlapping
the /etc/machine-id control 3 produced noise overlapping
/etc/shadow replaying control 2's root hash noise overlapping
a fresh random host key at the same path overlapping overlapping

The last row was the finding: the safe case was the one that failed. My
rationale, "a value this recipe makes on its own", was true and did not support
the conclusion. What makes a generated secret safe is that it is different
every build
, so matching a control is evidence of pinning, not of freshness,
and docs/traps.md already records the instance: "Every appliance built from
core has the same machine-id".

The early return is deleted. It cost nothing: a genuinely pinned field still
passes on length and class, and two independent random values do not collide,
so the branch only ever admitted replays. tests/compare carries the three
replays and the fresh-key control.

Glob breadth is not a character count

./etc/ss* had exactly the eight literal characters the old check wanted, and
covered an SSH host key and a TLS private key under one sentence. ./var/lib/*
had ten and covered every state tree.

A rule is now an exact path, or a named directory and its subtree written
DIR/** with at least two components below ./, and the paths it actually
covers must sit under one immediate child of that directory. Verified
against the current tree:

  *                                      REFUSED
  ./etc/*                                REFUSED
  ./etc/ss*                              REFUSED
  ./var/lib/*                            REFUSED
  ./etc/**                               REFUSED
  ./var/lib/mysql/**                     allowed
  ./etc/ssh/**                           allowed
  ./var/lib/mysql/mysql/global_priv.MAD  allowed

A directory is a thing somebody chose; a character count is not. One sentence
now justifies one subtree.

The three MEDIUM

  • justified-against is enforced. The capture directory it names is
    compared with --dir and the run fails when they differ. On the real
    captures, rewriting the template's directory gives
    THIS RULE NAMES OTHER CAPTURES and a FAIL.
  • The template no longer teaches the value that disables the check. It
    carries real observed maxima, 16384 for the .MAD and 76 for the .frm, one
    rule per file shape, and a max-position at or past the sample cap is now
    refused outright.
  • Only overlapping is abbreviated. noise passes, so the one remaining
    route to a bad noise was landing where no bytes were printed. In the gate
    run the block now prints bytes for 145 of 167 paths, up from 52, with 22
    abbreviated and all of them overlapping.

shellcheck -S style exits 0, which it did not when I checked that box:
SC2115 and SC2086 in tests/measure.

The waiver cannot reach PASS, and the correction is yours

Confirmed by running the shipped template against the real captures. It clears
the two paths it names, with rule, justification and bytes printed, and the run
is still FAIL:

  rule ./var/lib/mysql/mysql/global_priv.MAD: 1 paths (written against 1, max position 16384)
  rule ./var/lib/mysql/mysql/global_priv.frm: 1 paths (written against 1, max position 76)
  cleared in this run: 2 paths, every one of them named above
  totals: noise 102, overlapping 31, real 32, not sampled 0, unchanged 0, cleared 2
verdict: FAIL (0 attributable, 32 real, 31 overlapping and uncleared, ...)

So it is more control captures and then a waiver, not one or the other.
docs/layer-measure.md says so now, and the earlier claim that either would do
is removed.

The fourth rework, and why this merges with its first subject failing

The waiver vocabulary is now the state list's vocabulary

Two measures of glob breadth had been tried and both were wrong the same way.
Counting literal characters let ./etc/ss* through. Counting components below
./ let ./var/lib/** through, and ./var/lib is the common parent of the six
subtrees share/layer-state-paths declares separately, so one sentence
about MariaDB accounts could clear PostgreSQL state in a later run whose count
coincided. The counts coinciding is not luck: the same component shape produces
the same number of state files.

A glob rule's directory must now be a directory the capture's own
state-paths.txt declares, or below one. ./var/lib/mysql/** allowed,
./var/lib/** refused, ./etc/ssh/** still allowed because
./etc/ssh/ssh_host_* names that directory. A glob with no wildcard in its last
component declares a file and no directory, so ./etc/shadow does not license
./etc/**, and ./etc/ssl/** is refused where ./etc/ssl/private/** is the
declaration. ./var/cache/**, ./usr/share/** and ./etc/webmin/** are
refused too. MEASURE_WAIVER_MIN_DEPTH is gone. The list is already compared
byte for byte across the captures with its digest checked, so this adds nothing
new to trust.

justified-against was a substring test and accepted mariadb-gate2 for a run
on mariadb-gate; its first comma separated component is compared for equality
now.

What this cannot see, now written down

docs/layer-measure.md gains a section on it. State identical in every capture
is invisible to a differential measurement by construction, including a secret
copied out of the base control's own build tree, and
share/layer-state-suspect does not cover that either, because such a path
never differs and so is never subtracted. A PASS means "nothing the change did
shows up as a difference from the control". It does not mean the layer ships no
secret it should not.

The 32 are an open item, not a waiver, and more captures cannot clear them

#13. The field was decoded from the captures: bytes 180 to
183 of the Aria index header are a four byte big endian wall clock Unix
second
, SOURCE_DATE_EPOCH does not touch it, and 1790574106902134 / 10^6
is exactly user.frm's microsecond timestamp in the same directory. The unit
crossed into byte 181 where three controls had not.

Byte 181 carries every 18.2 hours and builds are 160 s apart, so an expected
single crossing needs on the order of 410 consecutive control builds; byte
180 carries every 194 days. For any finite control set there is always a higher
byte that did not carry, and the spread required grows by 256 per byte. So
more control captures cannot fix this and the build host should not be asked
for them.
That reverses what an earlier revision of this description said.

The fix is the pin this tool already has for text, applied to a binary field,
and it is not in this pull request: issue #13 carries the four conditions.
Landing a new inference rule in the same change as the instrument, in order to
make the instrument's first subject pass, is how the defect this pull request
exists to fix came about. So the 32 are published as they are, recorded against
#13, with no waiver over them, which the tool refuses anyway because a waiver
only ever clears overlapping.

Merging the instrument while its first subject fails

That is deliberate and it is the point. The alternative is to hold the tool
until mariadb passes, which means the first thing anybody does with it is tune
it until it agrees with a conclusion already reached. That is precisely how the
measurement this replaces came to be trusted: a number was produced, it said
zero, and nothing that could have contradicted it was ever run. A FAIL that
names 32 paths, decodes the field they differ in and points at the issue that
would resolve it is worth more than a PASS arranged by adjusting the
instrument.

The fifth rework: the list is anchored to the repository

The licensing rule made state-paths.txt load bearing twice. It decides which
bytes a capture keeps, and, since a waiver rule's directory has to be one the
list declares, it now also decides which waivers are legal. The preconditions
checked that every capture carried a byte identical list whose recorded digest
matched, which says the operator was consistent with themselves and nothing
more: four captures taken against a list with ./var/lib/** added agree with
each other, the block says they agree, the waiver is licensed, and one sentence
about MariaDB accounts clears PostgreSQL state. share/layer-state-suspect
cannot catch it and reports 0 subtracted but state shaped throughout, because
widening the list makes more paths state paths so nothing new is subtracted.

What changed is that widening the list used to be harmless. It only ever
increased what was compared by content. The same knob now points both ways.

attribute compares the capture's copy with the committed
share/layer-state-paths and refuses on mismatch. A run against another
list declares a reason with --state-paths-override, printed in the block as
OVERRIDE, and both digests are printed in any case.

Refusal rather than a printed warning, which was the alternative offered.
Every other precondition here refuses, and the defect being closed is a field
that was recorded and never read; a warning relies on somebody noticing, which
is the failure mode itself. The digests are printed regardless, so the weaker
remedy is there too.

The anchor has no environment override. MEASURE_STATE_PATHS is a capture
time knob and attribute ignores it entirely, including as a way of moving the
anchor, so the property that pointing it at a narrower file changes nothing
after the fact survives and is still tested. tests/measure now takes its
captures against the committed list rather than a fixture list, which removes a
difference between the suite and a real run.

The two LOW

./home/**/.ssh/** declares ./home/**/.ssh, compared literally, which no
wildcard-free rule can sit below, so that declaration licenses nothing. That is
intended and the doc says so: it fails closed, an exact-path waiver is the route
for those paths, and matching a glob against a glob here would put back the
breadth ambiguity the rule exists to remove.

The sentence "a PASS does not mean the layer ships no secret it should not" is
now repeated in the section about what a pull request quotes, because the person
over-reading a PASS is reading the block.

What the coverage number does and does not say

Recorded in COVERAGE.md rather than left as a bare 100 percent. Of the six
branches added with the licensing rule, three fail the suite when removed: the
wildcard-in-last-component guard in measure_state_dirs, the equality clause in
measure_dir_at_or_below, and loosening that function's path boundary to a
prefix match. Two survive because they are redundant, the licensing loop
refusing the same inputs without them, so they are defence in depth and a
surviving mutant is the right result. The sixth, the */* guard that skips a
malformed entry with no slash, was executing untested and now has
state-dirs-noslash behind it.

The two LOW from the sixth review

The OVERRIDE line is printed only when the anchor check actually failed. A
reason given on captures that do carry the committed list is printed on an
override line as declared and not needed, so the block never asserts a
mismatch its own two digests contradict. And tests/measure asserts which
fixture paths were sampled rather than how many, so adding a state path to the
committed list no longer breaks the suite; checked by adding ./etc/keel/**
and ./var/lib/dpkg/**.

The gate block below predates the fifth rework and these two; neither changes a
verdict, and the header of a run today also prints the committed digest line.

The merge gate: mariadb, four captures

Run on the build host, SOURCE_DATE_EPOCH=1700000000, core as parent, three
controls and one unit, lock taken with a bounded wait, nothing else built and
nothing published. Captures in /mnt/builds/measure/mariadb-gate, 590 MiB.

=== bt-layer-measure attribution ===========================================
command       bt-layer-measure attribute --dir /mnt/builds/measure/mariadb-gate --unit unit --control control --control-again control-again --control-again control-third
state paths   4 captures agree on the list (sha256 b3649941adaa9570)
unit          unit  mariadb on core, epoch 1700000000, captured 2026-09-28T05:52:00Z, 438 packages, 35335 files, 3380 symlinks, 238 state paths
control       control  mariadb on core, epoch 1700000000, captured 2026-09-28T05:44:08Z, 438 packages, 35335 files, 3380 symlinks, 238 state paths
control       control-again  mariadb on core, epoch 1700000000, captured 2026-09-28T05:46:43Z, 438 packages, 35335 files, 3380 symlinks, 238 state paths
control       control-third  mariadb on core, epoch 1700000000, captured 2026-09-28T05:49:26Z, 438 packages, 35335 files, 3380 symlinks, 238 state paths
control pairs 3, and the floor below is their union

attributable file differences if a single control pair were the floor:
  control vs control-again: 0
  control vs control-third: 0
  control-again vs control-third: 0
  (this command uses the union of all of them, below)

attributable differences: 0

state paths the unit pair differs at: 167
  overlapping ./var/lib/mysql/mysql/global_priv.MAD
              13 bytes at offsets the controls also vary at, size delta 0 within 0 to 0; bytes carry no shape, so this is a coincidence nobody has ruled out
  real        ./var/lib/mysql/mysql/global_priv.MAI
              differs at bytes the controls do not vary at: byte 181, byte 197
  overlapping ./var/lib/mysql/mysql/global_priv.frm
              8 bytes at offsets the controls also vary at, size delta 0 within 0 to 0; bytes carry no shape, so this is a coincidence nobody has ruled out
  noise       ./var/lib/mysql/mysql/user.frm
              2 differing lines, every one of them matching the mask the controls establish at that line
  totals: noise 102, overlapping 33, real 32, not sampled 0, unchanged 0, cleared 0

subtracted as floor noise, and not state paths: 6
  subtracted (6):
    ./etc/webmin/mysql/config
    ./root/.wget-hsts
    ./var/cache/ldconfig/aux-cache
    ./var/log/alternatives.log
    ./var/log/webmin/webmin.log
    ./var/webmin/module.infos.cache

verdict: FAIL (0 attributable, 32 real, 33 overlapping and uncleared, 0 not sampled, 0 subtracted but state shaped)
============================================================================

Abridged: the full 676 lines are /root/gate-block.txt on the build host.

Zero attributable, which is the conclusion decision 0010 reached, now
against the union of three control pairs and also against each pair alone. What
is new is the 167 state paths, which that measurement subtracted by name
without opening.

The four, confirmed against the bytes

Path Verdict What the bytes say
global_priv.MAD overlapping differs at {8894-8896, 9093-9095, 9298-9300, 16381-16384}, exactly the set control vs control-third differs at. The strings show password_last_changed moving 1790574121 to 1790574610, a clock, and the account set is identical: admin at 127.0.0.1, ::1, localhost. So it is a clock, and bytes cannot prove that, which is why it must be cleared rather than assumed.
global_priv.MAI real bytes 181 and 197 differ and no control pair moves them. Correct, and the reason is below.
global_priv.frm overlapping differs at {67-72, 75-76}, identical to control vs control-third; no string differs. .frm create and update timestamps.
user.frm noise correct, and it really is text: mysql.user is a TYPE=VIEW over global_priv, the .frm is ASCII with newlines, and its only difference is one line, timestamp=0001790574106902134 against ...596687502, which the mask the three controls establish there admits.

Why 32 came out real

Almost all of them are *.MAI. The Aria index header carries a Unix timestamp,
and in 30 of the 32 it increases strictly in capture order. Three control
builds moved only its low bytes; the unit build, four minutes later, carried
into the byte above. The evidence now prints the bytes from every capture at
the first differing offset, in capture order, which is what makes that readable
rather than a list of offsets:

      the 8 bytes at offset 181, in capture order:
        control       b9 fe 1a 00 00 00 00 00
        control       b9 fe c6 00 00 00 00 00
        control       b9 ff 64 00 00 00 00 00
        unit          ba 00 04 00 00 00 00 00

The other 33, the overlapping ones, are almost all *.frm, whose timestamps
sit at fixed offsets every build moves together.

So this FAIL is the honest state of the measurement and not a defect in the
component. Neither a waiver nor more captures clears it: a waiver is consulted
only for overlapping and these are real, and the field is a wall clock
second whose higher bytes carry on scales of 18.2 hours and 194 days, so no
finite set of controls spans them. The resolution is the binary pin in
#13, deliberately not in this pull request. See "The 32 are
an open item" above.

Coverage

Split three ways under decision 0004: bin/layer-measure-lib (capture and the
formats), bin/layer-compare-lib (whether a difference is noise),
bin/layer-report-lib (the blocks and the preconditions), with
bt-layer-measure a thin main. Measured with tests/coverage.sh 99:

File Lines covered Cover
bin/layer-measure-lib 206 of 206 100 percent
bin/layer-compare-lib 330 of 330 100 percent
bin/layer-report-lib 273 of 273 100 percent
bt-layer-measure 78 of 78 100 percent
bin/layer-lib 331 of 332 99 percent
bt-layer 110 of 111 99 percent
bin/aplinfo-lib 186 of 186 100 percent
bt-aplinfo 47 of 47 100 percent
bin/signature-lib 76 of 76 100 percent
bin/generate-signature 87 of 87 100 percent

All four new files at 100 percent line and branch: every subcommand, every exit
code, every verdict and every error path, including the two guards that must
never fire, which tests/measure reaches by replacing measure_verdict with a
stub. The gate stays at 99, the lowest file rounded down.

shellcheck is clean at -S style, which closes the four info-level notes in
the review (SC2016, SC2015, SC2086 twice) and the ones the rework introduced.

Test plan

  • tests/measure and tests/compare pass; so do tests/layer,
    tests/aplinfo and tests/signature with pipefail added
  • tests/coverage.sh 99 exits 0, numbers above and in COVERAGE.md
  • shellcheck -S style exits 0 on all seven touched shell files
  • the four false negatives from the review, each one a case in
    tests/compare
  • the MEASURE_STATE_PATHS substitution and the identical-tag run, both
    now refused, both cases in tests/measure
  • a real four-capture run against mariadb, block above, recipe restored,
    nothing else built, nothing published
  • the four global_priv / user.frm verdicts confirmed against the bytes
    one at a time, table above
  • the five mask attacks from the re-review, each a case in tests/compare,
    and the clock they hid behind still noise at both prefix lengths
  • the waiver composition from the re-review, a case in tests/measure
  • the gate re-run with the stricter rule: same tallies, 102 / 33 / 32
  • the three replays and the fresh-key control, cases in tests/compare
  • every refused and allowed glob form, cases in tests/compare
  • the shipped template run against the real captures: clears its two
    paths, prints their bytes, and the run stays FAIL on the 32 real
  • every refused and allowed glob form against a real state path list,
    including ./var/lib/**, cases in tests/compare
  • a justification naming a directory that merely has the run's as a
    prefix, a case in tests/measure
  • captures taken against a list that is not the committed one: refused,
    and the variable cannot move the anchor, cases in tests/measure
  • a mutation test of the six branches the licensing rule added, recorded
    in COVERAGE.md
  • tests / coverage green on this pull request

Handbook decision 0010 step 3 requires a three-build comparison, control
against unit against control again, for every component taken out of the
shared tree. Nothing ran it. The capture half was /root/measure/build.sh on
the build host, a scratch file in root's home directory on one machine, and
the comparison half was never committed anywhere, so keel-mariadb#11 and
keel-postgresql#7 were approved on a number only their author could produce.

bt-layer-measure does both halves. capture builds one layer under the build
lock and records its package list, file tree, symlinks, layer manifest and
state paths under a directory named on the command line; compare reports what
two captures differ by; attribute reports what the unit build differs from
the control by that the control does not differ from itself by, and prints a
block a pull request quotes whole, so a claim names the command behind it.

Three things the hand-run measurement got wrong, and this does not:

Paths are keyed on path<TAB>hash, not on sha256sum's "<hash>  <path>". A
Debian rootfs holds setuptools/_vendor/jaraco/text/Lorem ipsum.txt, which a
whitespace split turns into two fields that never match themselves; that
phantom cost the LAMP measurement of 2026-09-27 one file in each of its two
counts (docs/traps.md, "A path with a space in it"). A path carrying a tab, a
newline or a backslash is refused rather than mis-counted.

A path inside the noise floor is no longer subtracted on the strength of its
name. For mariadb the floor is 179 files of which 173 are ./var/lib/mysql/**,
and that directory holds mysql/global_priv and mysql/user.*, so a component
whose build-time job includes deleting accounts was measured blind exactly
where the accounts live. share/layer-state-paths lists the paths that can
hold accounts, credentials, keys or database content; their bytes are kept at
capture time and compared per line for a text file and per byte offset for a
binary one, in both pairs. A clock that moves in every build touches the same
line or offset in both and cancels; a deleted account row leaves one the
control pair never produces and is reported as real, with a three-way diff.

A file that could not be read is not a file that did not change. A state path
above the sample cap is recorded as not sampled and fails the run, so no
report can say "0 attributable differences" about bytes it never saw.

Logic in bin/layer-measure-lib, thin main in bt-layer-measure, tests in
tests/measure, documented in docs/layer-measure.md. tests/coverage.sh runs
the new suite: 379 of 379 lines in the library and 64 of 64 in the
executable, every subcommand, every exit code and every error path, gate
unchanged at 99.

Closes #10
@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Review

I reproduced the three recount claims independently from /root/tree-{m,p}-*.sha256, with my own script and not with this tool, and then again through this tool. All three hold exactly: mariadb 182 / 179 / 182 (173 of them ./var/lib/mysql/**), postgresql 17 / 13 / 16 (6 of them ./var/lib/postgresql/**); the mariadb residue against the first control pair is exactly the three files named and is zero against the union; ./etc/postgresql/17/main/postgresql.conf survives every control pair and the union. All four of global_priv.MAD, .MAI, .frm and user.frm differ in the control-unit pair and sit inside all three control floors, so all four were subtracted unexamined. I also reproduced tests/measure, shellcheck and tests/coverage.sh 99 (379/379 and 64/64, exit 0).

The design is right and the diagnosis is right. Two things in the implementation fail in the same direction as the defect this exists to fix, which is why this is a Block and not a Warning. Both are cheap to fix.


CRITICAL 1. The signature cancels a real difference against a coincidental one at the same position, and prints nothing when it does

bin/layer-measure-lib:309-331 (measure_signature), :370-395 (measure_state_verdict), :721 (the evidence guard).

A token is a position with the content thrown away: o<n>/n<n> from diff at :316, @<n> from cmp -l | awk '{print "@" $1}' at :325, and a bare len at :326. measure_set_minus then cancels position against position, so two unrelated changes at one line number or one byte offset annihilate.

Run against this branch, a unit build that adds an account to /etc/shadow:

  noise      ./etc/shadow
             every difference here is one the control pair shows against itself (3 in total)
  totals: noise 1, real 0, not sampled 0

verdict: PASS

exit 0. The unit file contains mallory:$y$REALPASSWORDHASH:20000:::. It is reported as noise because the control pair happened to append a different line at the same line number. Three more, all verified by running:

  • floor rewrites line 1 for a clock reason, unit sets a real root password on line 1 → noise
  • true binary with NULs, floor differs at bytes 9-16, unit differs at the same eight offsets with different byte values → noise, all eight tokens cancelled
  • len carries no magnitude: floor +1 byte against unit -3584 bytes → noise

The bite is not hypothetical. global_priv.MAD, .MAI, .frm and user.frm differ in both pairs in the captures on the build host today, and those are exactly the files an account deletion lands in. The next mariadb extraction gets four noise lines and a count.

What makes it unrecoverable is :721: if [[ "$verdict" != noise ]] suppresses the evidence block, so the one error this scheme can make is the one error nobody can see. Minimum fix: print the three-way evidence for every state path with a non-empty control-unit signature, not only for non-noise ones, and say how many tokens cancelled. Better: give len both sizes, and make "cancelled at the same position but with different content" a verdict of its own that a person has to clear.

CRITICAL 2. unsampled can be configured away at attribute time, and the pasted block still looks provenanced

bt-layer-measure:95-96 and :174: state_paths_file() reads $MEASURE_STATE_PATHS when attribute runs, not from the captures.

$ bt-layer-measure attribute --dir D --control c2 --unit u --control-again c3
verdict: FAIL (0 attributable, 0 real inside the noise floor, 167 not sampled)

$ MEASURE_STATE_PATHS=./narrow-list bt-layer-measure attribute --dir D --control c2 --unit u --control-again c3
state paths   ./narrow-list (sha256 464279cdd6df1594)
attributable differences: 0
state paths the unit pair differs at: 0
verdict: PASS

exit 0, on the real mariadb captures, with a one-line substituted list. capture writes state_paths and state_paths_sha256 into meta.tsv at :151-152 for precisely this purpose and nothing ever reads either field: the only measure_meta_get calls in the tree are layer, parent, epoch and captured (bin/layer-measure-lib:576-579). The header prints the digest of whatever list was substituted, so the block a pull request pastes carries a digest that proves nothing.

Fix: attribute reads state_paths_sha256 from all three captures, requires they agree with each other and with the list in hand, and refuses otherwise. The field is already there.


HIGH 3. attribute implements the defect the paired handbook change writes into traps.md

bin/layer-measure-lib:542: one --control-again, no way to pass a second. The new trap entry's Fix says "capture a third control and subtract the union of both control pairs"; this command cannot do that.

Run on the real captures, changing only which pair is the floor:

floor mariadb attributable postgresql attributable
control vs control-again 3 4
control-again vs third control 0 1

So the first real run against mariadb prints FAIL (3 attributable) for a component the handbook concludes is at zero, and the reviewer is invited to wave it through by hand. That is traps.md, "A reproducibility reference that cannot match is worse than none". Make --control-again repeatable, subtract the union, and report which floors each residue survived.

HIGH 4. Nothing enforces the precondition the whole measurement rests on

docs/layer-measure.md:36-37 and the amended 0010 both say the three captures must share SOURCE_DATE_EPOCH, one common commit and one parent "or nothing the comparison prints means anything". epoch, layer and parent are recorded and printed (bin/layer-measure-lib:574-580) and never compared. epoch none is accepted. A mismatched epoch inflates both pairs, grows the floor and still reports PASS.

Related, same line of defence: --control X --unit X --control-again X gives attributable differences: 0, verdict: PASS, exit 0. Verified. Three tags, no distinctness check.

HIGH 5. share/layer-state-paths omits credential paths for the two components already in flight

share/layer-state-paths:18-46. Absent: ./etc/mysql/**, which holds debian.cnf and the debian-sys-maint password; ./etc/postgresql/**, which is where the one real postgresql residue actually lived and was caught only because it fell outside the floor; ./root/.my.cnf; ./root/.pgpass; ./etc/redis/**; ./etc/machine-id (traps.md, "Every appliance built from core has the same machine-id").

The answer to "what happens when a future component puts state somewhere not on this list" is: if the control pair also differs there, it is subtracted by name and leaves no trace. That is the original defect, relocated. It is made worse by the report never naming what it subtracted (for mariadb, 176 paths leave the report as a single count). measure_print_set already caps long lists; printing the subtracted set would let a reviewer eyeball it for anything credential-shaped.


MEDIUM 6. Any NUL-free file is treated as text

bin/layer-measure-lib:287-298. A NUL-free, newline-free blob is one line, so its whole signature degenerates to o1 n1, which cancels against any other difference at all. Verified with three mutually different 1500-byte blobs: noise.

MEDIUM 7. A state path can be examined and counted in no total

bin/layer-measure-lib:713 reads the verdict through < <(measure_state_verdict ...) and discards the function's return value; an empty verdict increments none of the three counters, and :730 fails the run on attributable + real + unsampled only. Unreachable today because the missing-file case is pre-checked at :377-383, but it is the same shape as the pipefail bug you found. One cheap invariant closes it: assert the three counters sum to the number of paths examined.

MEDIUM 8. Captures keep long-lived copies of shadow, private keys and database bytes

bin/layer-measure-lib:207-225. A tag is never overwritten and nothing is cleaned up, by design. cp does preserve the file modes (I checked: 600 and 640 survive), so this is not a plain leak, but the directories it creates are 755 where /etc/ssl/private is 710 upstream, and --dir itself is created at the ambient umask. chmod 700 on the capture root and a retention line in the doc.

MEDIUM 9. The same shell-option divergence is still live in the three older suites

tests/measure sets pipefail and tests/layer, tests/aplinfo and tests/signature do not, while bt-layer, bt-aplinfo and bin/generate-signature all do. Out of scope for this diff, but bt-layer is what produces every capture this tool reads.


LOW 10. The documented procedure does not run as written

docs/layer-measure.md:42-45 exports seven variables and omits BT_BUILDS, which bt-layer:91 treats as fatal and which is unset in root's environment on the build host (checked). bt-layer-measure:109 hardcodes /mnt/builds/layers rather than deriving it from BT_BUILDS, so the two can also disagree silently until :124 reports no rootfs. Loud rather than silent, but docs/layer-measure.md is the reproducibility claim, and this is the first thing a second person hits.

LOW 11. shellcheck

Exits 0, so "clean" holds, but there are four info-level notes: SC2016 at bin/layer-measure-lib:160, SC2015 at tests/measure:49, SC2086 at tests/measure:443 and :445.


On the unchecked box

Leaving "a real three-build run" unchecked is the honest call, and it should gate the merge rather than follow it. Suggested gate for the first real use: run it against mariadb with four captures, publish the block, and confirm by hand that the four global_priv/user.frm verdicts say what the bytes say. Those four files are the ones this tool exists for and the ones finding 1 will mis-report.

I verified by running: the recount (both families, all pairs), the three residues, the four blind files, four false-negative cases end to end, the MEASURE_STATE_PATHS substitution, the identical-tag run, the per-floor attributable swing, tests/measure, shellcheck, and tests/coverage.sh 99. I inferred, without running: the traps.md cross-references and the severity of finding 8 on a multi-user host.

Block, on findings 1 and 2 only. Everything else can follow. This is much better than what it replaces and the recount behind it is sound; it is blocked because an instrument that reports PASS on an added account, and whose blindness can be switched off by an environment variable without the printed block showing it, fails in the exact direction it was built to correct.

…tchable

Review of #11 found two ways this failed in the same direction as the defect
it exists to correct. Both are closed here.

The signature recorded where two files differ and threw the content away, so
two unrelated changes at one line number or one byte offset annihilated, and
the evidence block was suppressed for anything called noise. Four cases came
back as noise, exit 0, with nothing printed: an account added to /etc/shadow
at the line number the control pair also appends at; a root password set on
the line the control pair rewrites for a clock reason; a binary differing at
the same eight offsets with different byte values; a file 3584 bytes shorter
against a control pair that varies by one. The mariadb captures on the build
host show global_priv.MAD, .MAI, .frm and user.frm differing in both pairs
today, so the next extraction would have got four noise lines and a count.

The controls are now the model and the model carries content. At each position
the controls vary at, the longest common prefix and suffix of their variants
is the shape of that variation, and the unit is admitted only if its line has
that shape. A mask with neither a prefix nor a suffix is what two unrelated
lines have in common, which is no evidence about a third, so it never yields
noise. Bytes carry no shape at all: the most that can be said of a binary is
that it coincides in position, in size delta and in how much differs, which is
the new verdict `overlapping`, and it fails until somebody reads the bytes and
says so in --cleared. The evidence is printed for every verdict, noise
included, because the one mistake this scheme can still make is the one nobody
would otherwise see.

The state path list was read from the environment when attribute ran, so
MEASURE_STATE_PATHS pointing at a one line file turned FAIL (167 not sampled)
into PASS on the real captures while the block still printed a digest. capture
now copies the list into the capture, attribute reads that copy and never the
variable, and it refuses unless every capture carries the same list and the
digest each one recorded matches it.

Also from the review:

- --control-again is repeatable and the floor is the union over every control
  pair. One pair is not a floor: mariadb differs by 179 files in one pair and
  182 in another, and the residue is 3 against the first and 0 against the
  second. The block prints what each single pair would have said.
- attribute refuses unless the captures share layer, parent and
  SOURCE_DATE_EPOCH, were built with an epoch at all, and are distinct.
  Recording a precondition and never checking it is not recording it.
- share/layer-state-paths gains ./etc/mysql/**, ./etc/postgresql/**,
  ./etc/redis/**, ./root/.my.cnf, ./root/.pgpass and ./etc/machine-id. The
  postgresql residue lived in the second of those and was caught only because
  it fell outside the floor.
- The list is treated as a failure mode, not a list. Every subtracted path is
  named rather than counted, and any whose name looks like state
  (share/layer-state-suspect) fails the run, with two committed ways out: add
  it to the list, or clear it with a reason.
- Text is no longer "contains no NUL": a file with no newline, or with a line
  over 8 KiB, is a blob and is compared as bytes.
- The verdict totals must add up to the paths examined, and an empty verdict
  is fatal rather than uncounted.
- A capture directory is created 0700, and the retention rule is in the doc.
- tests/layer, tests/aplinfo and tests/signature set pipefail, which the
  scripts they test already did.
- docs/layer-measure.md exports BT_BUILDS, and the layers directory is derived
  from it through config/common.cfg the way bt-layer derives it.

Logic split under decision 0004: bin/layer-measure-lib keeps capture and the
formats, bin/layer-compare-lib decides whether a difference is noise,
bin/layer-report-lib prints the blocks and enforces the preconditions.
tests/compare is the four false negatives above. Coverage: 204 of 204, 189 of
189, 230 of 230 and 70 of 70, all four at 100 percent, gate unchanged at 99.
The first real use of this, four captures of mariadb on the build host, found
one thing missing from the evidence and confirmed the rest.

For a binary the block listed the offsets that differ and not the bytes at
them, which is not enough to settle anything by hand. It now prints a window
from every capture at the first differing offset, in capture order, and that
turns an unreadable list into a finding: for mysql/global_priv.MAI the value
reads b9fe1a, b9fec6, b9ff64, ba0004 down the four captures, so a reader sees
an Aria header clock that carried into the byte above while three control
builds moved only its low bytes.

measure_is_text also refuses a file carrying control bytes other than tab,
newline and carriage return. Being wrong about text is safe in one direction
only: a binary masked line by line can come out as noise, while a text file
compared as bytes comes out overlapping and has to be cleared by hand. It stays
strict rather than paranoid on purpose, because mysql/user.frm is a TYPE=VIEW
definition that really is text, and its only difference between builds is one
timestamp line that the mask admits.

The run itself, recorded in docs/layer-measure.md: 0 attributable differences,
against each single control pair and against the union of all three; 167 state
paths that differ, of which 102 noise, 33 overlapping, 32 real, 0 not sampled;
6 subtracted paths, all named, none state shaped. The four the review asked
about were checked against the bytes one at a time and all four say what the
bytes say: global_priv.MAD and .frm overlapping at the offsets every control
also moves, .MAI real because its counter carried where no control's did, and
user.frm noise on a timestamp line.

Coverage: 204, 216, 230 and 70 lines, all four files at 100 percent, gate
unchanged at 99.
@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Both CRITICALs are closed, and the gate you proposed has been run. The full
picture is in the pull request description; the short version:

1. Positions cancelled content. The controls are now the model and the
model carries content: at each position they vary at, the longest common prefix
and suffix of their variants is the shape the unit has to fit. A mask with
nothing in it proves nothing. Bytes carry no shape at all, so a binary is never
noise; it is overlapping, which fails until it is cleared in writing with a
reason the block prints. The evidence is printed for every verdict including
noise, which is the part that made this unrecoverable. Your four cases are
four cases in tests/compare; all four now come out real, real,
overlapping, real.

2. unsampled was configurable. capture copies the state path list into
the capture, attribute reads that copy and never the environment, and it
refuses unless every capture carries the same list and the digest each one
recorded matches it. The substitution you showed is now exit 3.

Also done: --control-again is repeatable and the floor is the union over
every pair, with the per-pair counts printed so the 3-against-0 swing is
visible; preconditions on layer, parent, epoch and distinctness are enforced;
the six missing credential paths are on the list and /etc/machine-id with
them; every subtracted path is named and a state-shaped one fails the run;
text now requires a newline, no over-long line and no stray control bytes; the
counter invariant is asserted; captures are 0700 with a retention line;
pipefail is in the three older suites; BT_BUILDS is in the doc and derived
through config/common.cfg; shellcheck is clean at -S style.

The gate. Four captures of mariadb, three controls, lock taken with a
bounded wait, nothing else built, nothing published. 0 attributable, against
each single pair and against the union. 167 state paths differ: 102 noise, 33
overlapping, 32 real, 0 not sampled. Your four say what the bytes say, checked
one at a time:

  • global_priv.MAD overlapping, at exactly the offsets control vs
    control-third also moves; password_last_changed 1790574121 to 1790574610
    and the account set identical.
  • global_priv.MAI real at bytes 181 and 197.
  • global_priv.frm overlapping, offsets identical to the control pair's.
  • user.frm noise, and it really is text: mysql.user is a TYPE=VIEW, and
    the only differing line is timestamp=0001790574106902134 against
    ...596687502, which the mask admits.

The 32 real are almost all *.MAI, and in 30 of them the differing value
increases strictly in capture order: the Aria index header clock, whose carry
three controls did not span. The evidence now prints the bytes from every
capture at the first differing offset so that reads off the page:

      the 8 bytes at offset 181, in capture order:
        control       b9 fe 1a 00 00 00 00 00
        control       b9 fe c6 00 00 00 00 00
        control       b9 ff 64 00 00 00 00 00
        unit          ba 00 04 00 00 00 00 00

So the FAIL is the honest state of that measurement, not a defect in the
component, and getting to PASS is a decision I have not made for you: more
control captures until the floor spans the carry, or a committed waiver saying
the Aria clock was read and accepted. I considered teaching the binary path to
recognise a run that climbs with capture time and downgrade it, and did not,
because that is the kind of unreviewed inference that got the first version
blocked.

Two things from your list I did not do. MEDIUM 9 is done rather than deferred,
since bt-layer produces every capture this reads. And I have left the
mariadb-gate captures on the build host, 590 MiB, because they are the
evidence for the block above; they hold /etc/shadow and database bytes, so
they should go once this is settled.

@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Re-review

Both blocking findings are fixed, and one of them is fixed better than I asked for. I re-ran the four false negatives, attacked the new scheme, and checked the gate run against the bytes on the build host.

Verified fixed. All four cases now fail: overlapping, real, overlapping, real. (The body says case 1 is real; it is overlapping. Both fail, nothing turns on it, but the block should say what the tool says.) The state path list is now read from the captures and MEASURE_STATE_PATHS at attribute time is not refused but ignored entirely, which is stronger than the exit 3 I asked for: my substitution changed nothing, and measure_preconditions also requires all four captures to carry byte-identical lists whose recorded digests match. Also confirmed by running: --control-again is repeatable and the floor is the union, with each single pair's answer printed; distinct tags, matching layer/parent/epoch and a refusal on epoch none; the totals invariant; measure_is_text strict; the capture directory 700 on the build host; pipefail in all five suites; BT_BUILDS documented and refused loudly; shellcheck -S style clean; tests/measure and tests/compare pass; and tests/coverage.sh 99 exits 0 with every number in the table matching, 204, 216, 230 and 70 all at 100 percent.

The share/layer-state-suspect net works. I narrowed the list at capture time for all four captures, consistently, so every digest agreed: the run still failed, 2 subtracted but state shaped, because ./etc/shadow and the host key hit the suspect names. That is a real defence and it closes the capture-time version of the old hole.


CRITICAL 1. The mask constrains the two ends of a line and nothing else, so where the controls vary at random the unit may substitute a constant and be called noise

bin/layer-compare-lib:228-235. measure_mask_admits tests four things: prefix or suffix non-empty, ${#line} >= ${#prefix} + ${#suffix}, starts-with, ends-with. There is no bound on the length of the free region, no constraint on its character class, and no requirement that it resemble anything a control put there.

So the rule is: wherever the controls legitimately vary at random, the mask degenerates to the line's fixed framing, and any unit value with that framing is admitted, including a constant. Verified by running, three controls in each case:

case mask verdict
ssh_host_ed25519_key.pub, random key per control, hardcoded key in the unit ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAI / root@keel noise
/etc/shadow, random root hash per control, baked-in hash in the unit root:$y$j9T$ / :20000:0:99999:7::: noise
a planted line sharing a clock line's prefix and suffix # written 1790475 / by build noise
the free region grown by 400 characters token= / empty noise

Those two masks are the algorithm header and the comment field, and the hash format and the date fields. They are the parts that cannot vary. The scheme reads them as evidence about the part that can.

And on the gate's own bytes: I took the real mysql/user.frm from /mnt/builds/measure/mariadb-gate/control, put timestamp=0001790574 OR 1=1; GRANT ALL ON *.* TO evil@localhost on line 10, and the tool returns noise.

The principle already in the file disposes of this. bin/layer-compare-lib:27 says bytes carry no shape, so a binary coincidence is never noise. A per-build random text field carries no more shape than a random byte field. noise should require the free region to be pinned, by the controls agreeing on its length and its character class and the unit matching both, and should fall to overlapping when it is not. That is not new policy, it is the policy already written applied to the other branch.

This is the original defect narrowed rather than closed, and it now lands on precisely the paths share/layer-state-paths exists to watch. It did not fire in the gate run, because every differing state path there was under ./var/lib/mysql/.

HIGH 2. A cleared path prints no evidence

bin/layer-report-lib:328-333: the waiver branch increments waived and continues before the verdict line is printed and before measure_evidence is called. Verified end to end: the cleared path appears zero times in the evidence section.

That contradicts the section's own header, "the evidence below is printed whatever the verdict, because the one mistake this scheme can make is to call a real difference noise", and the comment at bin/layer-compare-lib:240, "A glob waiver is only honest because the report names every path it covered". A waiver is the one place a person overrides the instrument. It is the last place to withhold the bytes.

HIGH 3. Finding 2 composed with the waiver route turns the gate green over a real defect

Verified, and this is the reason 1 and 2 are worth fixing before the waiver is granted rather than after. With a waiver naming only ./var/lib/mysql/**, which is exactly the route under discussion:

  noise       ./etc/shadow
  noise       ./etc/ssh/ssh_host_ed25519_key.pub
  totals: noise 2, overlapping 0, real 0, not sampled 0, unchanged 0, cleared 1
waived in writing (1 paths, every one of them named):
  ./var/lib/mysql/mysql/global_priv.MAD  ->  Aria index header clock, ...
verdict: PASS

exit 0, and the unit in that run ships a hardcoded SSH host key. The waiver removes the last overlapping, and what is left is the column the mask can get wrong, with its evidence abbreviated or, for the waived path, absent.

MEDIUM 4. --cleared breadth is unbounded and unrecorded

bin/layer-compare-lib:243-256. A single rule *<TAB>reason matches every path, including ./etc/shadow and ./etc/ssh/ssh_host_key (verified); ./etc/* crosses slashes and matches ./etc/ssl/private/a.key. The report lists the covered paths, which is right, but capped at MEASURE_PRINT_MAX, and it never says which rule covered each path. The header at bin/layer-report-lib:250 prints the waiver file's name and no digest of its contents, so a block quoting --cleared waivers.txt cannot be checked against the file that produced it.

Worth crediting on the same mechanism: a waiver applies only to overlapping and can never clear a real. I tried it; the real survives and the run fails.

LOW 5. Doc rot

tests/measure's header comment still enumerates measure_signature, measure_state_verdict and measure_state_evidence. None of them exist now.


The gate run

Checked independently against /root/gate-block.txt (891 lines) and the captures. The tallies are right: 102 noise, 33 overlapping, 32 real. All 167 differing state paths are under ./var/lib/mysql/, and all 102 noise are .frm files, mysql.user plus the sys schema views, every one of them a TYPE=VIEW text .frm with one timestamp= line. That is coherent.

global_priv.MAD is confirmed. I pulled all four copies and read the records. Six accounts in every capture, admin at 127.0.0.1, ::1 and localhost, plus mariadb.sys, mysql and root at localhost, with identical access masks, plugins and authentication_string values. The only thing that moves is password_last_changed, 1790574121, 1790574292, 1790574458, 1790574610, climbing with capture time. overlapping is the honest verdict and the account set claim is exactly right.

global_priv.MAI is confirmed and the new window evidence earns its place. b9 fe 1a, b9 fe c6, b9 ff 64, ba 00 04 reads as a big-endian counter carrying into byte 181, which is what real at 181 and 197 means. A reviewer can settle that by eye, which was not possible before.

user.frm is correct, and the tool did not establish it. The factual claim holds: it is genuinely ASCII, longest line 4718 against the 8192 cap, only line 10 differs in every pair, and the four timestamps climb monotonically. It is also safer than the note argues, because mysql.user in 10.4+ is a view over global_priv and holds no account rows at all, so no account file is being let through here. But the mask is timestamp=0001790574 with an empty suffix, and that 21 character prefix is an accident of four builds landing inside the same thousand seconds. I checked: across a digit carry the prefix shortens, and a planted payload that keeps it is admitted. The verdict is right about these bytes and is not established by the reasoning that produced it.

The 32, and what should gate the first PASS

The waiver is the right instrument, and for a stronger reason than the one given. More captures is not merely luck on the real side: the same luck governs the noise side, because the user.frm mask's strength depends on where the builds fell relative to a digit boundary. A criterion whose answer depends on the hour the gate ran is not a criterion. Declining to teach the tool to downgrade a monotone run was also right.

But the waiver should not be granted before finding 2 is fixed, because finding 2 plus the waiver is finding 3. Fix 1 and 2, then waive, and make the waiver name:

  1. the exact paths, or a glob plus the number of paths it covered when it was written, so it fails instead of silently stretching when a new file appears in that directory;
  2. the offsets and the observed value range it covers, so the same path turning overlapping at different offsets is not quietly included;
  3. the capture directory and recipe commit it was justified against, so it cannot outlive its evidence;
  4. and it must be a file in the repository, with the report printing its digest and its rules the way it already prints the state path list's.

Operational

The evidence directory is 590 MiB and mode 700 throughout, which settles the permissions point from the last round. The disk claim in the brief does not hold: /mnt/builds is on /dev/mapper/turnkey-root, 122G, 46 percent used, 64G free, not 95 percent. Nothing needs pruning on that account, though a retention line in the doc is still worth having since a tag is never overwritten.

The pull request body reads correctly end to end: 280 lines, every section present, the gate block intact and matching /root/gate-block.txt, the coverage table matching what I measured, and the test plan boxes consistent with what I could reproduce. Nothing appears to have been lost in the rewrite.

History is clean: three commits on the tip of 19.x, all authored and committed by the same person, all passing the repository's commit-msg and pre-commit hooks when replayed, and no forbidden strings in the diff.


Verified by running: the four original cases, four new mask attacks, the planted payload on the real user.frm, the attribute-time substitution, a consistent capture-time narrowing, the waiver composition, waiver breadth and the real waiver limit, the account-set comparison on all four global_priv.MAD copies, the gate block tallies, both suites, shellcheck -S style and the coverage gate. Inferred, without running: that no credential-shaped text state path will differ in a future component's gate run, which is what finding 1 would meet.

Block, on findings 1 and 2, which are one function and one continue apart. Everything I raised last round is genuinely settled, the gate run is sound and its FAIL is honest, and I would approve this the moment noise requires a pinned free region and a cleared path prints its bytes.

Re-review found the first fix incomplete in the same direction as the defect,
and it was right.

The mask pinned a line's common prefix and common suffix and left the region
between them free: no length, no character class, no resemblance to anything a
control put there. So wherever the controls vary at random the mask degenerates
to the line's fixed framing, and any value carrying that framing was admitted
as noise. A hardcoded ssh_host_ed25519_key.pub behind
"ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAI" and " root@keel", a baked-in root
password hash behind "root:$y$j9T$" and ":20000:0:99999:7:::", a free region
grown by 400 characters, and a SQL payload planted on the real mysql/user.frm
of the gate run all came back noise.

The fix is the rule already in the file applied to the other branch. This file
argues that bytes carry no shape, so a binary coincidence is never noise; a
per-build random text field carries no more shape than a random byte field.
So noise now requires the controls to pin the free region: they agree on its
length, it is no longer than MEASURE_FREE_MAX, and every character is in the
class they used there, which may only be digits and the punctuation of a date.
A line a control itself produced is admitted whatever its shape, because that
is a value the recipe makes on its own.

That also removes the dependence on when the gate ran. The user.frm mask is 21
characters of prefix only because four builds landed inside the same thousand
seconds, and across a digit carry it collapses to 9. tests/compare runs that
case at both lengths: the clock is noise and the payload is not, because the
criterion is now the field's width and class rather than how much framing
survived.

A cleared path printed no verdict line and no evidence; the waiver branch
returned before both. That contradicted the section's own header and it is
where the bytes matter most, because a waiver is the one place a person
overrides the instrument. A cleared path now prints its verdict, the rule and
justification that cleared it, and its bytes, always and not subject to the
abbreviation cap.

The waiver itself was unbounded and unrecorded: one "*" rule matched
everything including /etc/shadow, the report never said which rule covered
which path, and it printed no digest. A rule is now five fields, rule, count,
max-position, justified-against and reason. A rule with too little literal
path is refused, a rule that covers a different number of paths than it was
written against fails the run rather than stretching, and a rule justified to
one offset does not cover a path differing beyond it. The block prints the
file's digest and every rule with what it covered.
share/layer-waivers.example is the template and it lives in the repository.

tests/measure carries the composition the review built: with a waiver naming
only ./var/lib/mysql/**, a unit shipping a hardcoded host key must still fail.
It does; the key is overlapping and named, and the waiver covers only the two
database paths it was written for.

Also: tests/measure's header comment no longer lists three functions that were
removed. Coverage 204, 298, 253 and 71 lines, all four files at 100 percent,
gate unchanged at 99.
@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Both fixed, and the argument for the first one is yours: a per-build random
text field carries no more shape than a random byte field, and this file
already refuses to call a binary coincidence noise for exactly that reason.

1. noise now requires a pinned free region. The controls must agree on
its length, it must be no longer than 24 characters, and every character must
be in the class they used there, which may only be digits and the punctuation
of a date. A line a control itself produced is admitted whatever its shape,
because that is a value the recipe makes on its own; everything else falls to
overlapping. Your five cases are five cases in tests/compare and all five
now fail, the SSH key and the password hash on the class test, the 400
character growth on the length test, and the user.frm payload on both.

It also answers your point about the hour the gate ran, which I think is the
better version of the argument. tests/compare now runs the user.frm field
across a digit carry, where the prefix collapses from 21 characters to 9: the
clock is noise at both lengths and the payload is overlapping at both,
because the criterion is the field's width and class and no longer how much
framing happened to survive.

2. A cleared path prints its verdict, the rule and justification that
cleared it, and its bytes
, always, and not subject to the abbreviation cap.

3. Your composition is a test. Three controls with a random host key each,
a unit with a hardcoded one, a waiver naming only ./var/lib/mysql/**. FAIL,
the key named as overlapping, the waiver covering the two database paths it
was written for and nothing else.

4. The waiver names the four things. Five tab separated fields: rule,
count, max-position, justified-against, reason. A rule carrying too little
literal path is refused, so * and ./etc/* are both out. A rule that covers
a different number of paths than it was written against fails the run rather
than stretching. A rule justified to one offset does not cover a path
differing beyond it. The block prints the file's digest and every rule with
what it actually covered. share/layer-waivers.example is the template and it
is in the repository.

5. tests/measure's header no longer lists the three removed functions.

The gate re-ran with the stricter rule and the tallies are unchanged, 102
noise, 33 overlapping, 32 real, because all 102 are TYPE=VIEW .frm files
whose one differing line is a fixed width decimal timestamp=. That is the
case the new rule is meant to keep, so it is a useful confirmation that it is
not merely stricter.

Corrections to my own last comment, both yours: the four original cases come
out overlapping, real, overlapping, real, and I wrote real for the
first; the pull request body now says so too. And I repeated a stale 95
percent disk figure. /mnt/builds is 122G at 46 percent with 64G free, so the
evidence directory is under no pressure; the retention line is in the doc
anyway, since a tag is never overwritten.

I have still not granted the waiver for the 32. That is the next thing and it
is yours to accept: the rule would be ./var/lib/mysql/** with its count, its
max position, the capture directory and the recipe commit, and a reason saying
the six accounts and their authentication_string values are identical in all
four captures and only password_last_changed moves.
share/layer-waivers.example carries exactly that line as the template, which
is a template and not a waiver in force.

Coverage: 204, 298, 253 and 71 lines, all four files at 100 percent, gate
unchanged at 99, shellcheck -S style clean.

@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Third review

The mask fix works. I re-ran all seven of my attacks and none of them is noise any more: the added account, the 400 character free region and the planted clock line come back overlapping, the root password and the 3584 byte truncation come back real, the hardcoded host key comes back overlapping. The free region summary is well judged too, since it reports "43 characters, not a number or a timestamp" without printing the value.

Also verified by running: a cleared path now prints its verdict, its rule, its justification and its bytes, and bin/layer-report-lib:356 exempts it from the abbreviation cap; the composition regression behaves as described, with the waiver naming only ./var/lib/mysql/**, the host key named overlapping and the run failing; *, ./etc/* and ./etc/s* are all refused; a stretched rule fails the run and says so by name; the waiver file's digest is printed; the dead function names are gone from the suite headers; both suites pass; and tests/coverage.sh 99 exits 0 with 204, 298, 253 and 71 all at 100 percent, matching the table exactly.

The carry claim holds, and for the right reason. I built the user.frm field at both prefix lengths. With no carry the prefix is 20 characters and the free region is 9 digits; across a carry the prefix is 19 and the free region is 10. The clock is noise at both and the payload is overlapping at both. The reason is the one that matters: the verdict is decided by "all controls' free regions are the same length, all in the pinned class, and the unit's matches", which does not reference the prefix length at all. The prefix only decides where the free region starts. That is not fixture luck, and it does answer the objection.

The tallies are well aimed, not vacuous. I fetched the state samples for all four captures and recomputed the verdicts myself with this branch's code. Over what I fetched, excluding the four large binaries: 102 noise, 32 overlapping, 30 real, the three missing being the ones I left out. The noise class is homogeneous, and I checked every member of it rather than sampling: all 102 are .frm, all 102 begin TYPE=VIEW, all 102 differ at exactly one line, and in all 102 that line matches ^timestamp=[0-9]{19}$. Zero exceptions. So the rule changed no verdict because the only text difference in the whole run is a fixed width decimal field, which is exactly the shape the rule is built to admit, while my attacks show it now rejects the shapes it should. Well aimed.


CRITICAL 1. The control-produced exemption inverts the property it protects

bin/layer-compare-lib, in measure_mask_admits:

for v in "$@"; do
    ...
    [[ "$line" != "$v" ]] || return 0
done

A unit line identical to any control's line at that position is admitted unconditionally, ahead of every pin. The base capture is the first control, so a value copied from the second or third control differs from the base, is flagged as a differing position, and is then admitted by this branch.

Verified, three controls each:

unit ships verdict
the SSH host key control 2 produced noise
the /etc/machine-id control 3 produced noise
/etc/shadow replaying control 2's root hash noise
a fresh random host key at the same path overlapping

The last row is the whole finding. The fresh random key is the safe behaviour and it is flagged; the fixed key is the dangerous one and it passes. End to end, with a waiver naming only ./var/lib/mysql/**, the unit shipping control 2's host key gives:

  noise       ./etc/ssh/ssh_host_ed25519_key.pub
  totals: noise 2, overlapping 0, real 0, not sampled 0, unchanged 0, cleared 1
verdict: PASS

exit 0, and the key in that capture is byte for byte the one in control-again.

The rationale, "it is a value this recipe makes on its own", is true and does not support the conclusion. What makes a generated secret safe is that it is different every build. A unit that pins it to one previously observed instance is unsafe because it is pinned, and matching a control is evidence of pinning rather than of freshness. docs/traps.md records the instance already: the published core layer ships a populated /etc/machine-id, which was a value the recipe made on its own, and that is the entry titled "Every appliance built from core has the same machine-id". This exemption is that reasoning written into the instrument that is supposed to catch it.

Deleting the early return costs nothing I can find. Where the field is pinned, the normal path already admits a coincidence on length and class. Where it is not pinned, two independent random values do not collide, so the only lines the branch admits are replays. If it is kept for a reason I have not seen, it should at least yield overlapping rather than noise, so the bytes get printed and a person decides.

HIGH 2. A rule with exactly the minimum literal path clears two unrelated secrets

MEASURE_WAIVER_MIN_LITERAL=8, and literal=${rule//[*?]/}. The rule ./etc/ss* has a literal of exactly 8 and passes every one of the five checks. It covers ./etc/ssh/** and ./etc/ssl/private/** together. Verified:

              cleared by rule ./etc/ss*, ... : both of these are fine honest
              cleared by rule ./etc/ss*, ... : both of these are fine honest
  totals: noise 1, overlapping 0, real 0, ... cleared 4
verdict: PASS

exit 0, with a hardcoded SSH host key and a rotating TLS private key waived under one sentence. ./var/lib/* has a literal of 10 and likewise passes while covering every database and application state tree.

The count check is what saves this in a single run, and it does not save it across runs: the rule stays in the repository and authorises the whole of ./etc/ss* for every future gate whose count happens to match. A length of literal characters is the wrong measure of how broad a glob is. Requiring the literal part to end at a path separator, or confining a glob to one subtree named in share/layer-state-paths, would say what the count cannot.

MEDIUM 3. The shipped waiver template sets max-position to a value that bounds nothing

share/layer-waivers.example uses 16777216, the state sample cap. measure_max_position returns a line number for a text file and a byte offset for a binary, so that figure bounds neither: it is past the end of anything the cap let through. The gate's own maxima are about 16384 for global_priv.MAD and 197 for the .MAI. Nothing checks that the number is plausible, and this is the line the maintainer is about to copy. The field is the right idea; the template teaches the value that disables it.

MEDIUM 4. justified-against is printed and never checked

Field 4 is only echoed, at bin/layer-report-lib:351 and :443. Nothing compares it to the run's --dir or to any recipe commit, so "it cannot outlive its evidence" is documentation rather than enforcement. This is the same shape as state_paths_sha256 two rounds ago, which was recorded and unread until it was flagged. Comparing the capture directory it names against --dir, and failing when they differ, is a two line check.

MEDIUM 5. The abbreviation cap still hides most of the noise column

MEASURE_EVIDENCE_PATHS is 20. In the gate run 135 paths are noise or overlapping and 115 print no bytes. noise passes, so the one remaining route to a bad noise, finding 1, lands in the part of the report that prints nothing. Cleared and real paths are now exempt from the cap; noise should be too, or the cap should count only noise paths whose free region was pinned by more than one control.

LOW 6. The shellcheck box is checked and the command does not pass

shellcheck -S style on the seven files exits 1: SC2115 (warning) at tests/measure:521, rm -rf "$fix/$2", and SC2086 (info) at :522.


The waiver: not yet, and it cannot do what it is being asked to do

Two things, and the second is the one that changes the plan.

First, it should not be granted while finding 1 stands, because a waiver over the overlapping column leaves noise as the only column left, and finding 1 is a hole in noise.

Second, and independently: a waiver cannot reach PASS here at all. bin/layer-report-lib:329 only consults the waiver file when the verdict is overlapping. The 32 are real. Granting the example waiver changes FAIL (0 attributable, 32 real, 33 overlapping and uncleared) into FAIL (0 attributable, 32 real, 0 overlapping and uncleared). It is still a FAIL, and the 32 are untouched by it. The choice in the pull request is not between more captures and a waiver; it is more captures and then a waiver. More controls until the floor spans the Aria counter's carry is what moves those 32 from real to overlapping, and only then is there anything for a waiver to clear.

That also resolves the objection I raised last round about spanning a carry being luck. It is luck, but it is the only lever that touches the 32, because the alternative, teaching the tool to recognise a monotone run, was correctly declined.

When it is granted, after findings 1 and 2, the five fields should read:

  1. rule ./var/lib/mysql/**, the exact subtree. Not a wildcard that can reach another tree, and one rule per subtree rather than one rule spanning several.
  2. count the number of overlapping paths that rule covers in the run it was written against, taken from that run's own block rather than estimated.
  3. max-position the largest offset actually read, which in the current captures is about 16384 for the .MAD and well under 1000 for the .MAI, so one rule per file shape rather than one number covering both. Never the sample cap.
  4. justified-against /mnt/builds/measure/mariadb-gate and the recipe commit, and this only becomes a real constraint once finding 4 makes it checked.
  5. reason what was read and what it was found to be. The existing text is the right standard: it names the six accounts, states that their authentication_string values are identical in all four captures, and says only password_last_changed moves. I confirmed all of that against the bytes myself last round.

History is clean: four commits on the tip of 19.x, all by the same author, all passing the repository's commit-msg and pre-commit hooks when replayed, no forbidden strings in the diff.

Verified by running: the seven mask attacks, the three exemption replays and the fresh-key control, the carry fixture at both prefix lengths, an independent recount of all 167 gate verdicts from the fetched samples with a per-path check of every one of the 102 noise paths, the waiver breadth and refusal cases, the ./etc/ss* composition, the cleared-path evidence, both suites, shellcheck -S style and the coverage gate. Inferred, without running: that more control captures would move the 32 from real to overlapping, which follows from the offsets being outside the floor but which no run here demonstrates.

Block, on findings 1 and 2. Finding 1 is a deleted line, and finding 2 is a stricter test than counting literal characters. Everything else from the previous two rounds is settled, the pinning rule is well aimed on evidence rather than by assertion, and the carry test does what it claims.

…properly

Two findings from the third review, and the first was one line undoing the rule
it sat inside.

measure_mask_admits admitted a unit line identical to any control's line at
that position, ahead of every pin, on the grounds that it was a value the
recipe makes on its own. The base capture is the first control, so a value
copied from the second or third differs from the base, is flagged as a
differing position, and was then waved through. Against three controls: a unit
shipping the host key control 2 produced came out noise, the machine-id
control 3 produced came out noise, a replayed root hash came out noise, and a
fresh random key at the same path came out overlapping. The safe case was the
one that failed, which is the property this file exists to have, inverted.

The rationale was true and did not support the conclusion. What makes a
generated secret safe is that it is different every build, so matching a
control is evidence of pinning rather than of freshness. docs/traps.md,
"Every appliance built from core has the same machine-id", is that defect
already recorded here. The early return is deleted and it cost nothing: a
genuinely pinned field still passes on length and class, and two independent
random values do not collide, so the branch only ever admitted replays. All
four cases now come out overlapping, and tests/compare carries the three
replays and the fresh-key control.

MEASURE_WAIVER_MIN_LITERAL counted the literal characters of a glob, and
./etc/ss* has exactly the eight it required while covering an SSH host key and
a TLS private key under one sentence; ./var/lib/* has ten and covers every
state tree. A rule is now an exact path, or a named directory and its subtree
written DIR/**, with at least two components below ./, and the paths it
actually covers must sit under one immediate child of that directory. So *,
./etc/*, ./etc/ss*, ./var/lib/*, ./etc/** and ./var/lib/my*/** are all
refused, and one sentence justifies one subtree. A directory is a thing
somebody chose; a character count is not.

Three more from the same round:

justified-against was printed and never read, which is the shape
state_paths_sha256 had two rounds ago. The capture directory it names is now
compared with --dir and the run fails when they differ.

share/layer-waivers.example set max-position to the sample cap, which is past
the end of anything that was sampled and so bounds nothing, on the line a
maintainer copies. The template now carries real observed maxima, one rule per
file shape, and a value at or past the cap is refused.

The abbreviation cap hid 115 of the 135 non-real paths, and noise passes, so
the one remaining route to a bad noise landed where no bytes were printed.
Only overlapping is abbreviated now.

And the correction to the plan, which was the reviewer's and is right: a
waiver is consulted only for overlapping, so granting one over the 32 real
paths would leave them untouched and the run still FAIL. It is more control
captures and then a waiver, not one or the other. docs/layer-measure.md says
so now.

shellcheck -S style exits 0, which it did not before: SC2115 and SC2086 in
tests/measure. Coverage 204, 316, 264 and 71 lines, all four files at 100
percent, gate unchanged at 99.
@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Both fixed, and finding 1 was the more useful of the two to have had pointed
out, because the line was mine and I had written it as a kindness.

1. The exemption is deleted. You were right that the rationale was true and
did not support the conclusion: what makes a generated secret safe is that it is
different every build, so matching a control is evidence of pinning rather than
of freshness, and traps.md already had the instance. All four of your rows now
come out overlapping, including the fresh random key that was already right,
so the safe case and the dangerous one are no longer sorted the wrong way round.
Deleting it changed nothing else: a genuinely pinned field still passes on
length and class, which is why admits-pinned-digit in tests/compare still
holds while admits-replay next to it does not. The three replays and the
fresh-key control are cases there now.

2. Glob breadth is structural. A rule is an exact path, or a named directory
and its subtree as DIR/** with at least two components below ./, and the
paths it actually covers must sit under one immediate child of that directory.
So *, ./etc/*, ./etc/ss*, ./var/lib/*, ./etc/** and ./var/lib/my*/**
are all refused, and a rule that covered two subtrees of the directory it names
fails with SPANS 2 SUBTREES. Your phrasing is the one I took: a directory is
something somebody chose, a character count is not.

3. justified-against is compared with --dir and the run fails when they
differ. On the real captures, rewriting the template's directory gives
THIS RULE NAMES OTHER CAPTURES and a FAIL.

4. The template carries observed maxima, 16384 for the .MAD and 76 for the
.frm, one rule per file shape, and a max-position at or past the sample cap
is refused outright rather than merely discouraged.

5. Only overlapping is abbreviated now. The gate block prints bytes for
145 of 167 paths, up from 52; the 22 still abbreviated are all
overlapping, which is named in the totals and fails anyway.

6. shellcheck -S style exits 0. You were right that I had checked that box
against a command that did not pass.

Your correction to the plan, confirmed by running it. I ran the shipped
template against the real captures. It clears the two paths it names, with rule,
justification and bytes printed, and the run is still FAIL:

  rule ./var/lib/mysql/mysql/global_priv.MAD: 1 paths (written against 1, max position 16384)
  rule ./var/lib/mysql/mysql/global_priv.frm: 1 paths (written against 1, max position 76)
  cleared in this run: 2 paths, every one of them named above
  totals: noise 102, overlapping 31, real 32, not sampled 0, unchanged 0, cleared 2
verdict: FAIL (0 attributable, 32 real, 31 overlapping and uncleared, ...)

So it is more captures and then a waiver. The pull request and
docs/layer-measure.md both said "either, or" and now say what you established.
I have not added control captures in this round: that is a build-host decision
about how many and when, and it is the lever that moves the 32, so it should be
a deliberate run rather than something I slip in with a code fix.

The gate re-ran with all of the above and the tallies are unchanged, 102 / 33 /
32, which is the confirmation I wanted that none of this round moved a verdict
by accident.

Coverage 204, 316, 264 and 71 lines, all four files at 100 percent, gate
unchanged at 99.

@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Fourth review

Both blocks are closed and I could not reopen either. One new finding in the replacement glob rule, described below, and it is a HIGH rather than a block: nothing in the tree exercises it, and the block prints enough for a careful reader to catch it.

The exemption. Deleted outright, and the author's account of why it was wrong is correct. I re-ran the four rows: replaying control 2's host key, control 3's machine-id and control 2's root hash from /etc/shadow all come back overlapping, and a fresh random key at the same path also comes back overlapping. That last equality is the point. Under the old code the fixed value passed and the fresh one failed, which is backwards; now they are treated alike, which is the honest answer, since a differential measurement cannot tell a new draw from a pinned one and should not pretend to. admits-pinned-digit still holds: a fixed width digit clock is still noise.

I went looking for another route to the same shortcut and found one place that compares the unit to a control, bin/layer-compare-lib:521, cmp -s -- "$base" "$unit" for the none verdict. That one is sound and also unreachable in a real run, because a path whose hash matches the base is never in the differing set to begin with. It does imply a limit worth one line in docs/layer-measure.md: state copied out of the base control's build tree into a component is invisible to this measurement by construction, because it is not a difference from the control. That is not a defect in this scheme, it is what "differential" means, and the share/layer-state-suspect net does not cover it either since such a path never differs. Saying so in the doc is better than leaving a reader to assume the measurement sees all shipped state.

Glob breadth. Structural now, and much better. Verified refused: *, ./etc/*, ./etc/ss*, ./var/lib/*, ./etc/**, ./var/**, ./var/lib/my*/**, ./home/**, ./root/**. Verified allowed: ./var/lib/mysql/**, ./etc/ssh/**. A rule spanning two immediate children fails with SPANS 2 SUBTREES. The count check, the justification check and the subtree check all fire, and I confirmed the correctly scoped rule is refused when reused on a different run: THIS RULE HAS STRETCHED: written against 2, covers 0.

The three MEDIUM and the LOW. justified-against is compared against --dir and fails with THIS RULE NAMES OTHER CAPTURES. The template now carries two exact paths with observed maxima, 16384 for the .MAD and 76 for the .frm, matching what I measured from the captures myself last round, and max-position at or past the sample cap is refused at bin/layer-compare-lib:414. Only overlapping is abbreviated now, at bin/layer-report-lib:359, so noise, real and cleared always print bytes. shellcheck -S style exits 0 with no output.

Tallies unchanged. I recomputed every verdict from the fetched state samples with this revision's code: 102 noise, 32 overlapping, 30 real over the file set I hold, identical to what the previous revision gave over the same set. Nothing moved.


HIGH 1. ./var/lib/** satisfies the new structure and carries a justification across components

MEASURE_WAIVER_MIN_DEPTH=2 counts components below ./, so the directory ./var/lib has depth 2 and passes. But ./var/lib is the common parent of six subtrees that share/layer-state-paths declares separately, precisely because they are different things:

./var/lib/mysql/**   ./var/lib/postgresql/**   ./var/lib/redis/**
./var/lib/couchdb/** ./var/lib/mongodb/**      ./var/lib/nodebb/**

The "one immediate child" check only looks at what the rule covered in the run in front of it, so a rule at ./var/lib/** passes whenever a single service's state differs. Demonstrated with two runs and one rule, changing nothing but which service the differing files sit under:

run 1, mysql-shaped:
  cleared     ./var/lib/mysql/sub/a.dat
  verdict: PASS

run 2, postgresql-shaped, same rule, same count:
  cleared     ./var/lib/postgresql/sub/a.dat
              cleared by rule ./var/lib/**, ... : MariaDB Aria index header clock;
              the six MariaDB accounts are identical in all four captures.
  verdict: PASS

exit 0 both times, with a sentence about MariaDB accounts clearing PostgreSQL state. The count check does not catch it because the counts coincide, and counts coinciding is not a coincidence: the same component shape produces the same number of state files. For contrast, the correctly scoped ./var/lib/mysql/** is refused on run 2, which is the mechanism working as intended and the reason this is one rule away from closed.

The fix uses data already in the capture and already digest checked. A glob rule's directory must be at or below the directory part of some glob in that capture's state-paths.txt. Then ./var/lib/mysql/** is allowed because the list declares that directory; ./var/lib/** is refused because nothing declares ./var/lib; ./etc/ssh/** stays allowed because ./etc/ssh/ssh_host_* declares that directory. That replaces an arbitrary depth with the project's own statement of what is state, so the waiver vocabulary becomes exactly the vocabulary of the thing being waived. MEASURE_WAIVER_MIN_DEPTH then goes away.

Two others are legal at depth 2 and worth a glance while that is being done: ./etc/ssl/** is broader than the declared ./etc/ssl/private/**, and ./var/cache/**, ./usr/share/** and ./etc/webmin/** are all legal. Only ./var/lib/** actually bites today, because it is the only one sitting above sibling declared state.

LOW 2. The justified-against comparison is a substring test

bin/layer-report-lib:450, [[ "$(cut -f4 <<<"$rule")" != *"$dir"* ]]. A waiver justified against /mnt/builds/measure/mariadb-gate2 is accepted for a run whose --dir is /mnt/builds/measure/mariadb-gate, since the shorter string is a substring of the longer. It fails the other way round, so the laxness is one directional and narrow. Comparing the field's first comma separated component for equality would close it.


The 32, and whether more captures can clear them

Putting this here as well as in my report, because it bears on what the tool should promise rather than on this diff.

I decoded the field from the captures. Bytes 180 to 183, one based, are a four byte big endian integer:

control        1790574106   0x6ab9fe1a
control-again  1790574278   0x6ab9fec6
control-third  1790574436   0x6ab9ff64
unit           1790574596   0x6aba0004
steps: 172, 158, 160

and 1790574106902134 / 10^6 = 1790574106, which is user.frm's microsecond timestamp in the same directory. So it is a wall clock Unix second, unaffected by SOURCE_DATE_EPOCH, which is why it moves in every build and why pinning the epoch cannot make it reproducible.

Byte 181 is bits 16 to 23, so it carries every 2^16 seconds, 18.2 hours. Builds are about 160 seconds apart, so two consecutive controls straddle that boundary with probability about 0.24 percent, and an expected single crossing needs on the order of 410 consecutive control builds. Byte 180 is bits 24 to 31 and carries every 2^24 seconds, 194 days. So for any finite set of controls there is always a higher byte that did not carry inside the capture window, and a unit built after such a carry differs there, outside the floor, and is real.

So the number is unbounded in principle. More captures does not converge; the spread required grows by 256 per byte. Asking the build host for "a few more controls" cannot fix this and should not be asked for.

What I would do instead is the pin the tool already has for text, applied to binaries. At a run of differing offsets, read the enclosing fixed width big endian integer and admit it only when all four of these hold: the controls' values at that field increase strictly in capture order; the unit's value exceeds every control's; the unit's step over the last control lies inside the range of the controls' own steps, here 158 to 172; and the field's lower bytes are among the offsets the controls vary at. Every one of those is a measurement of the controls, not an inference about the unit, which is the same standard the free region pin already meets, and it is stricter than the "downgrade a monotone run" idea that was rightly declined. Under it this field is noise legitimately, 1790574596 beats 1790574436 by 160, and a planted value fails on the step range or on continuing the sequence.

Until that exists, the honest position is not to chase PASS on this component. Publish the block, record the 32 and this decode as an open item, and do not write a waiver that implies they were cleared, which in any case the tool correctly refuses to let you do.


History is clean: five commits on the tip of 19.x, all by the same author, all passing the repository's commit-msg and pre-commit hooks when replayed, no forbidden strings in the diff. Both suites pass and my attacks appear as cases in them.

Verified by running: the four replay rows and the pinned digit control, the none path, nine glob shapes against the shape check, the two run justification transfer, the stretched rule refusal, the template maxima and the cap refusal, the justification mismatch, the abbreviation change, an independent recount of the gate verdicts, shellcheck -S style, both suites and the coverage gate. Inferred, without running: the carry probability arithmetic above, which follows from the decode but which no build demonstrates.

Warning. No blocking issue. Finding 1 should land before the first waiver is written, because the first thing that happens after merge is somebody writes one, and ./var/lib/** is the natural shape to reach for.

Two measures of how broad a waiver rule is have now been tried and both were
wrong the same way, by being easier to write than to defend.

Counting the literal characters of the glob let ./etc/ss* through, which waives
an SSH host key and a TLS private key under one sentence. Counting its
components below ./ let ./var/lib/** through, and ./var/lib is the common parent
of the six subtrees share/layer-state-paths declares separately: mysql,
postgresql, redis, couchdb, mongodb, nodebb. The "one immediate child" check
only sees the run in front of it, so such a rule passes whenever a single
service's state differs, and a justification about MariaDB accounts clears
PostgreSQL state in the next run. The count check does not catch it because the
counts coincide, and that is not luck, because the same component shape produces
the same number of state files.

So the measure is now the project's own statement of what state is. A glob
rule's directory has to be a directory the capture's own state path list
declares, or below one. ./var/lib/mysql/** is allowed because the list names
that directory, ./var/lib/** is refused because nothing names ./var/lib, and
./etc/ssh/** stays allowed because ./etc/ssh/ssh_host_* names it. A glob with no
wildcard in its last component declares a file and not a directory, so
./etc/shadow does not license a waiver over ./etc, and ./etc/ssl/** is refused
where ./etc/ssl/private/** is the declaration. ./var/cache/**, ./usr/share/**
and ./etc/webmin/** are refused too, being declared nowhere.
MEASURE_WAIVER_MIN_DEPTH is gone. The list is already compared byte for byte
across the captures and its digest already checked, so this adds nothing new to
trust.

justified-against was a substring test, so a waiver justified against
mariadb-gate2 was accepted for a run on mariadb-gate. Its first comma separated
component is now compared for equality.

docs/layer-measure.md gains a section on what this cannot see, which the review
drew out of the one remaining comparison against a control, the cmp -s for the
`none` verdict. State identical in every capture is invisible to a differential
measurement by construction, including a secret copied out of the base
control's own build tree, and the suspect net does not cover it either because
such a path never differs and so is never subtracted. A PASS means nothing the
change did shows up as a difference from the control; it does not mean the layer
ships no secret it should not.
@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Both landed in this pull request, plus the doc line and the decision you asked
me to state in the body.

HIGH 1. Your fix, taken as proposed: a glob rule's directory must be a
directory the capture's own state-paths.txt declares, or below one. A glob
with no wildcard in its last component declares a file and no directory, so
./etc/shadow does not license ./etc/**. MEASURE_WAIVER_MIN_DEPTH is gone.
Refused now, all cases in tests/compare against a real list: ./var/lib/**,
./var/**, ./etc/ssl/** where ./etc/ssl/private/** is the declaration,
./var/cache/**, ./usr/share/**, ./etc/webmin/**, alongside the shapes that
were already refused. Allowed: ./var/lib/mysql/**, ./var/lib/mysql/mysql/**,
./etc/ssl/private/**, ./etc/ssh/**, and any exact path. You are right that
this costs nothing to trust: the list is already compared byte for byte across
the captures with its digest checked, and the waiver vocabulary is now the
vocabulary of the thing being waived.

LOW 2. The first comma separated component, compared for equality. A
justification naming a directory that merely has the run's as a prefix is a case
in tests/measure.

The doc line. docs/layer-measure.md has a "What this cannot see" section.
State identical in every capture is invisible by construction, including a secret
copied out of the base control's own build tree, and the suspect net does not
cover it either because such a path never differs and so is never subtracted. A
PASS means nothing the change did shows up as a difference from the control; it
does not mean the layer ships no secret it should not. Better said there than
left to a reader, as you put it.

The decode, and the reversal. Thank you for doing that arithmetic. I had
been proposing more control captures in the body and that was wrong, so the body
now says so: the field is a wall clock second, the higher bytes carry on scales
of 18.2 hours and 194 days, and no finite control set spans them. I have taken
out the "more captures and then a waiver" line and I am not asking the build host
for anything on this account.

The binary pin is the next pull request, #13, which
carries your four conditions and the decode. Not this one, and the reason is the
one in your third review: landing a new inference rule in the same change as the
instrument, in order to make the instrument's first subject pass, is how the
defect this whole thing exists to fix came about. So the 32 are published as
they are, recorded against #13, with no waiver over them.

And the merge-while-failing argument is in the body, in your words rather
than paraphrased, because it is the argument that decides it: holding the tool
until mariadb passes means the first thing anybody does with it is tune it until
it agrees with a conclusion already reached.

Coverage 204, 322, 264 and 71 lines, all four at 100 percent, gate unchanged at
99. shellcheck -S style exits 0. One thing I fixed after your review that you
did not raise: the new loop read its directory list through
done < <(measure_state_dirs ...), which kcov cannot attribute, so it took
layer-compare-lib to 99.69; it reads into an array now.

@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Fifth review

The licensing rule is the right one and is correctly built. I could not find a hole in it. The hole is now in the thing it licenses from, which is what you would expect and is the one finding below.

Verified as reported. Refused: ./var/lib/**, ./var/**, ./etc/ssl/** against a declaration of ./etc/ssl/private/**, ./var/cache/**, ./usr/share/**, ./etc/webmin/**, ./etc/**, ./var/lib/my*/**. Allowed: ./var/lib/mysql/**, ./var/lib/mysql/mysql/**, ./etc/ssl/private/**, ./etc/ssh/**, and exact paths. Two details worth crediting because they are where this kind of check usually leaks: measure_dir_at_or_below is [[ "$1" == "$2" ]] || [[ "$1" == "$2"/* ]], a real path boundary, so ./var/lib/mysqlevil/** is refused rather than swallowed by a prefix match; and measure_state_dirs skips a declaration whose last component has no wildcard, so ./etc/shadow licenses nothing. The refusal message naming the directories that would be legal is a good touch.

Also verified: justified-against now compares the first comma separated component for equality, and a waiver naming mariadb-gate2 is refused for a run on mariadb-gate; all five suites pass; shellcheck -S style exits 0 with no output; and the tallies are unchanged for a third revision running, 102 noise, 32 overlapping, 30 real over the file set I hold.


HIGH 1. The list now decides two things and is still anchored to nothing

The licensing rule makes state-paths.txt load bearing twice: it decides which bytes are compared, and now also which waivers are legal. measure_preconditions checks that all captures carry a byte identical list and that each capture's recorded digest matches the copy it carries. Nothing compares that list to the committed share/layer-state-paths; grepping for it finds only the default for MEASURE_STATE_PATHS, two comments and one advice string.

So the way to license a rule is to declare its directory. Four captures taken against a list with ./var/lib/** added:

state paths   4 captures agree on the list (sha256 4bb5a4cf9a98fa13)
verdict: FAIL (... 2 overlapping and uncleared, ... 0 subtracted but state shaped ...)

  -- with a ./var/lib/** waiver --
  cleared     ./var/lib/mysql/sub/a.dat
  verdict: PASS

and the same rule, with the same MariaDB sentence, on a postgresql shaped run:

  cleared     ./var/lib/postgresql/sub/a.dat
              cleared by rule ./var/lib/**, ... : MariaDB Aria header clock, accounts identical
  verdict: PASS

exit 0 both times. share/layer-state-suspect does not catch it, and cannot: that net reads the paths the floor subtracted, and widening the list makes more paths state paths, so nothing new is subtracted. The run reports 0 subtracted but state shaped throughout.

What changed with this commit is that widening the list used to be benign. It only ever increased what was compared by content, which is the safe direction. It now also widens the waiver vocabulary, so the same knob has acquired a second effect pointing the other way.

The fix is the last unfinished piece of a finding from two rounds ago, when state_paths_sha256 was recorded and never read. That was closed for consistency between captures; it is still open against the repository. attribute should compare the capture's state-paths.txt with $BT/share/layer-state-paths and refuse on mismatch, with a deliberate override named in the block so that it is on the record rather than in somebody's environment. Short of that, print both digests side by side, because at present the block prints one digest and gives a reader nothing to check it against.

Not a block: it takes an operator running captures against a list that is not the committed one, the mismatch is in principle visible in the digest, and the fix is one comparison. But it is the same shape as the defect this tool exists to correct, so I would not leave it for later.

MEDIUM 2. Three of the new guards are not pinned by a test

I mutated each of the six new branches and ran tests/compare. Three are caught: removing the wildcard-in-last-component guard in measure_state_dirs, removing the equality clause in measure_dir_at_or_below, and loosening that function to a prefix match. Three are not: the [[ "$glob" == */* ]] guard in measure_state_dirs, and the no-wildcard-in-dir and empty-globs guards in measure_waiver_rule_shape.

The last two are redundant rather than untested: remove them and the licensing loop refuses the same inputs anyway, so they are defence in depth and the mutation surviving is the correct result. The slash guard is the one with no test behind it, covering a malformed list entry with no / in it, where ${glob%/*} would otherwise declare the glob itself as a directory.

Worth saying because of what it means for the number: see below.

LOW 3. A declaration with a wildcard in a non-final component licenses nothing

./home/**/.ssh/** yields the declared directory ./home/**/.ssh, and measure_dir_at_or_below compares literally, so no wildcard free rule directory can ever be at or below it. ./home/alice/.ssh/** is refused. It fails closed, which is the right direction, but it means that one declaration cannot be waived at all, which is probably not intended. A line in the doc or a small normalisation would settle it.


The coverage question

Two separate things, and the answer differs between them.

Is the array version equivalent? Yes, and by measurement rather than by reading: I reimplemented the licensing loop in the while read form it replaced and ran both over 15 rule shapes against three different lists, 45 cases, with zero disagreements. The refactor is behaviour preserving.

Is the restored 100 percent real? Yes as line coverage, and it is not a reattribution artefact: the lines execute under the tests, and three of the six new branches fail the suite when removed, which is coverage doing work. But 100 percent line coverage is not 100 percent branch behaviour, and the mutation test above says exactly which three lines a test would not miss. Two of those three are redundant guards, so the honest summary is that one new line, the slash guard, is executed without being tested. That is a small and specific gap, and saying which line it is seems better than either claiming the number means more than it does or discounting it.

Changing the loop and saying so, rather than quietly restoring the figure, was the right way to handle it.

The "What this cannot see" section

Strong enough, and written for the reader you described. It names the case that matters rather than gesturing at it, a secret copied out of the base control's own build tree; it explains why the suspect net does not cover it, that a path which does not differ was never subtracted; and it gives the two sentences somebody hunting for a larger claim cannot get round, that a PASS means nothing the change did shows up as a difference from the control and does not mean the layer ships no secret it should not. Pointing at a single image instrument and at the machine-id trap entry is the right place to send them next.

One addition. The person who over-reads a PASS is reading the block, not this section. I would repeat the second of those two sentences in the part of the doc about what a pull request quotes, so it travels with the thing being quoted.


History is clean: six commits on the tip of 19.x, all by the same author, all passing the repository's commit-msg and pre-commit hooks when replayed, no forbidden strings in the diff. Splitting the binary pin out to a separate pull request, explicitly so that a new inference rule does not land in the change whose first subject it would make pass, is the right call and the right reason.

Verified by running: 14 rule shapes against the committed list and against a widened one, the declared-directory derivation, the path boundary case, the end to end list widening and the cross component justification transfer, the suspect net's silence on it, the justification prefix case, six mutations of the new branches, a 45 case equivalence test between the array and process substitution forms, an independent recount of the gate verdicts, five suites, shellcheck -S style and the coverage gate. Inferred, without running: that no committed list today declares a directory above sibling state, which is true of the list in this diff but is a property of the file rather than of the check.

Warning. One HIGH, which is one comparison. Everything from the previous four rounds is closed and I could not reopen any of it.

The licensing rule of the previous commit made state-paths.txt load bearing
twice: it decides which bytes a capture keeps, and, since a waiver rule's
directory has to be one the list declares, it now also decides which waivers
are legal. measure_preconditions checked that every capture carried a byte
identical list whose recorded digest matched, which says the operator was
consistent with themselves and nothing more.

So four captures taken against a list with ./var/lib/** added agree with each
other, the block reports that they agree, a ./var/lib/** waiver is licensed,
and one sentence about MariaDB accounts clears PostgreSQL state on the next
run. share/layer-state-suspect cannot catch it and reports 0 subtracted but
state shaped throughout, because widening the list makes more paths state
paths, so nothing new is subtracted. What changed is that widening the list
used to be harmless: it only ever increased what was compared by content. The
same knob now points both ways.

attribute compares the capture's copy with the committed
share/layer-state-paths and refuses on mismatch. A run against another list
declares a reason with --state-paths-override, which the block prints as
OVERRIDE, and the block prints both digests in any case, because one digest
with nothing to check it against is the shape of the finding two rounds ago
where state_paths_sha256 was recorded and never read.

Refusal rather than a printed warning, for that same reason: every other
precondition here refuses, and a warning relies on somebody noticing.

The anchor has no environment override. MEASURE_STATE_PATHS is a capture time
knob and attribute ignores it entirely, including as a way of moving the
anchor, so the property that pointing it at a narrower file changes nothing
after the fact is unchanged and is still tested. tests/measure now takes its
captures against the committed list rather than a fixture list, which is the
real configuration and removes a difference between the suite and a real run.

Two smaller things from the same review.

A declaration with a wildcard in a non-final component, ./home/**/.ssh/**,
yields the declared directory ./home/**/.ssh, which is compared literally and
which no wildcard-free rule can sit below, so it licenses nothing. That is
intended and now says so in the doc: it fails closed, the route for those
paths is an exact path waiver, and matching a glob against a glob here would
put back the breadth ambiguity the rule exists to remove.

docs/layer-measure.md repeats, in the section about what a pull request
quotes, that a PASS does not mean the layer ships no secret it should not,
because the person over-reading a PASS is reading the block and not the
section that explains it.

And the slash guard in measure_state_dirs, which skips a malformed list entry
with no / in it, was executing without a test behind it. It has one now,
state-dirs-noslash. COVERAGE.md records which of the six branches added last
commit a mutation test kills and which survive because they are redundant,
since 100 percent line coverage with an untested line in it claims more than
it has.
@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Landed, and refusal rather than a printed warning.

The argument for refusing. Every other precondition here refuses: epoch,
layer, parent, distinctness, cross-capture list consistency. And the defect
being closed is the last piece of a finding about a field that was recorded
and never read
, so a remedy that relies on a reader noticing a second digest
is the failure mode rather than the fix. The block prints both digests anyway,
so the weaker remedy is there as well; it just is not the thing standing
between a widened list and a PASS.

--state-paths-override REASON is the escape, printed in the block as
OVERRIDE beside both digests, so a deliberate divergence is on the record
rather than in somebody's environment.

The anchor has no environment override, which matters for the property you
credited two rounds ago. MEASURE_STATE_PATHS stays a capture-time knob and
attribute ignores it entirely, including as a way of moving the anchor, so
pointing it at a narrower file still changes nothing after the fact and that
test still holds. tests/measure now takes its captures against the committed
list instead of a fixture list, which removes a difference between the suite and
a real run and is how I noticed the anchor had to be unconditional: my first
attempt read it through state_paths_file() and would have let the variable
move it.

I kept both things you flagged: measure_dir_at_or_below is still
[[ "$1" == "$2" ]] || [[ "$1" == "$2"/* ]], and measure_state_dirs still
skips a declaration whose last component has no wildcard.

LOW 3, intended. ./home/**/.ssh/** licenses nothing and that is the right
direction: it fails closed, an exact-path waiver is the route for those paths,
and a glob-against-glob comparison here would put back the breadth ambiguity the
rule exists to remove. Said in the doc now rather than left to be inferred.

LOW 4. The "does not mean this layer ships no secret it should not" sentence
is repeated in the section about what a pull request quotes.

The coverage gap. Tested, not excused: state-dirs-noslash in
tests/compare covers the */* guard, with a malformed entry that has no slash
in it beside a good one. COVERAGE.md now records the whole mutation result,
the three that kill the suite and the two that survive because they are
redundant, because a bare 100 percent with a line in it that no assertion would
miss claims more than it has. Thank you for running that; it is a better check
on the number than the number.

Merge order. Keel-Linux/handbook#8 says 0010's step 3 names bt-layer-measure
and describes the repeatable --control-again and the union floor, none of which
is true until this lands. I have put a line at the top of that pull request
saying it merges second.

Coverage 206, 330, 270 and 77 lines, all four at 100 percent, gate unchanged at
99. shellcheck -S style exits 0. Five suites pass.

@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Sixth review, scoped to the delta

The anchor holds. I attacked the three things you named and could not get past any of them; two small things below, neither blocking.

The two credited details survive verbatim. bin/layer-compare-lib:437 is still [[ "$1" == "$2" ]] || [[ "$1" == "$2"/* ]], and :429 is still [[ "$last" == *[*?]* ]] || continue.

1. The override

It cannot be set without appearing in the block, and it cannot be injected. measure_parse_args zeroes MEASURE_LIST_OVERRIDE before it reads anything, so MEASURE_LIST_OVERRIDE=sneaky in the environment is discarded and the run is refused at exit 3. With the flag, the block carries all three of the things a reader needs, and they are consistent with each other:

state paths   4 captures agree on the list (sha256 b08af7187f0b223f)
committed     .../share/layer-state-paths (sha256 b3649941adaa9570)
OVERRIDE      the captures were taken against a list that is not the committed
              one, declared: widened for the redis component, see issue 41

The two digests differ on their face, so the OVERRIDE line is not the only signal, which is the right belt and braces given what the finding two rounds ago was about. A clean run prints both digests equal and no OVERRIDE line, so the two are distinguishable by someone reading the block and nothing else.

On durability: the reason lives only in the block. Nothing is written into the capture, so a later reader of the capture directory alone cannot tell an override was ever used. That is the safe direction, because re-running attribute without the flag refuses rather than quietly proceeding, and the block is the artefact a pull request quotes. Worth knowing rather than worth changing. Any non-empty string is accepted as a reason; an empty one is refused.

2. The anchor and the environment

Verified genuinely immovable. attribute never calls state_paths_file(), the only function that reads MEASURE_STATE_PATHS; grep shows that variable reaching only do_capture. Pointing MEASURE_STATE_PATHS at the widened list during attribute does not move the anchor: captures taken against a non-committed list are still refused at exit 3, and captures taken against the committed list still pass with both digests equal. Splitting committed_paths_file() out rather than parameterising state_paths_file() is what makes that true by construction rather than by care, which is the better of the two.

3. The suite against the committed list

The content coupling is gone and was removed the right way: patterns-two and patterns-first, which asserted the list had exactly two entries and that the first was ./etc/shadow, are replaced by assert_grep presence checks. Those do not rot.

A count coupling remains. tests/measure:303, assert_eq sample-count 4, counts how many fixture files match the committed list. I simulated three future additions:

add ./etc/keel/**       -> FATAL [measure]: sample-count: got '5', want '4'
add ./var/lib/dpkg/**   -> FATAL [measure]: sample-count: got '5', want '4'
add ./etc/nginx/**      -> suite still passes

So adding a state path that matches something in the fixture rootfs breaks the suite. It fails loudly and names the assertion and both numbers, so it is self-announcing rather than silent, and the fix is obvious to whoever adds the path. Asserting which paths were sampled rather than how many would make it immune. Low, and I would not hold the merge for it.

4. Refusal rather than a warning

The reasoning is right and I would have argued the same. The specific point that carries it is that the defect being closed is a field recorded and never read, so a remedy whose whole weight rests on a reader noticing a second line reproduces the defect in a new place rather than closing it. Consistency with the other preconditions is the weaker half of the argument but also true: epoch, the distinct tags, the cross-capture list comparison and the digest check all refuse, and a single precondition that only warns would be the one people learn to skim. Printing both digests unconditionally is what makes the refusal auditable afterwards rather than merely obstructive, and that is the part I would have asked for if it were missing.

LOW. The OVERRIDE line asserts a mismatch it has not checked

bin/layer-report-lib:265 prints the line whenever --state-paths-override is non-empty, not when the anchor actually failed. Passing the flag on a run whose list does match produces:

state paths   4 captures agree on the list (sha256 b3649941adaa9570)
committed     .../share/layer-state-paths (sha256 b3649941adaa9570)
OVERRIDE      the captures were taken against a list that is not the committed one, declared: not actually needed

Two identical digests and a line directly beneath them saying they differ. It errs in the safe direction, since it over-reports and can never hide a real override, but in a block whose purpose is that a reader can check a claim against the evidence printed next to it, a line contradicting the two lines above it is worth one condition. Gate the print on the anchor having actually failed, or reword it to "an override was declared" and let the digests say whether it was needed.

The rest of the delta

state-dirs-noslash genuinely pins the guard: I removed [[ "$glob" == */* ]] || continue and the suite fails, where before it passed. COVERAGE.md records the whole mutation result, including that two of the six survive because they are redundant and that a surviving mutant is the correct outcome for those. Recording which branches a mutation test kills, rather than only the line percentage, is a better claim than the number alone and I would keep the habit.

Both LOW items from last round are settled as described, and the ./home/**/.ssh/** case is now documented as intended with the exact-path route and the reason, which is the right call: a glob-against-glob comparison would put back exactly the ambiguity the licensing rule removed. The PASS caveat now appears in the section about what a pull request quotes, which is where the person over-reading a PASS is looking.

Gates: all five suites pass, shellcheck -S style exits 0 with no output, coverage 206 / 330 / 270 / 77 all at 100 percent with the gate at 99. Seven commits on the tip of 19.x, the new one by the same author, passing both hooks when replayed, no forbidden strings.

Approve. The one LOW is a printed line that over-states, not a check that under-performs, and the anchor itself is sound.

marcos-mendez and others added 2 commits September 29, 2026 03:19
The OVERRIDE line asserted "the captures were taken against a list that is
not the committed one" whenever --state-paths-override was given, whether or
not the anchor check had failed. On captures that carry the committed list the
block then printed two identical digests and, directly beneath them, a line
saying they differ. It erred in the safe direction, but the block exists so a
reader can check each claim against the evidence printed next to it.

attribute now records that the anchor check failed, and the report prints
OVERRIDE only then. A reason given on captures that do carry the committed
list is still printed, on an "override" line saying it was declared and not
needed, so nothing the operator passed is dropped from the block.

Coverage: layer-report-lib 273 of 273, bt-layer-measure 78 of 78.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfWQScmDZ94KCS5DMYrfe6
tests/measure asserted that exactly four fixture files matched the committed
state path list. Since the suite takes its captures against the committed
list, adding a state path that happens to match a fixture file, ./etc/keel/**
or ./var/lib/dpkg/** for instance, broke the suite with a count mismatch that
said nothing about sampling. It now asserts the four paths that must be
sampled and two that must not, so the list can grow without touching the
suite.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfWQScmDZ94KCS5DMYrfe6
marcos-mendez added a commit to Keel-Linux/handbook that referenced this pull request Sep 29, 2026
0010 and the trap entry described the first version of the command: one
--control-again in the worked example, and a floor subtracted per line and per
byte offset in which a clock "touches the same line or offset in both pairs
and cancels". Review of Keel-Linux/buildtasks#11 showed that scheme is the
defect, not the fix: position cancels against position with the content
thrown away, so an account added on the line a clock moves on came back as
noise. The command no longer works that way, and a note that makes its block
the standard has to describe the block it actually prints.

- The worked example takes three controls and passes every one, since the
  floor is the union over every pair.
- The preconditions attribute enforces are named, including the anchor to
  the committed state path list and --state-paths-override.
- The block's contents are listed as they now are, and it carries the
  sentence that a PASS does not mean the layer ships no secret it should not.
- The verdicts are described as they are: noise only where the controls pin
  the varying field, overlapping for anything else at a varying position and
  for every binary, real where the controls do not vary, and waivers that
  clear only overlapping.
- The first real run is recorded: 0 attributable, FAIL on 32 real state
  paths that are the Aria header wall clock, open in Keel-Linux/buildtasks#13.

The trap entry's Fix now says not to cancel by position, instead of
prescribing it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfWQScmDZ94KCS5DMYrfe6
@marcos-mendez

Copy link
Copy Markdown
Collaborator Author

Both LOW from the sixth review fixed in 0a0cea1 and fb6c1ed: OVERRIDE is printed only when the anchor check actually failed (an unneeded reason now prints as override declared and not needed), and tests/measure asserts which fixture paths were sampled instead of how many. Coverage 273/273 and 78/78, CI green.

@marcos-mendez
marcos-mendez merged commit b2b4014 into 19.x Sep 29, 2026
1 check passed
marcos-mendez added a commit that referenced this pull request Sep 29, 2026
Make the three-build comparison of decision 0010 runnable by anyone
marcos-mendez added a commit to Keel-Linux/handbook that referenced this pull request Sep 29, 2026
0010 and the trap entry described the first version of the command: one
--control-again in the worked example, and a floor subtracted per line and per
byte offset in which a clock "touches the same line or offset in both pairs
and cancels". Review of Keel-Linux/buildtasks#11 showed that scheme is the
defect, not the fix: position cancels against position with the content
thrown away, so an account added on the line a clock moves on came back as
noise. The command no longer works that way, and a note that makes its block
the standard has to describe the block it actually prints.

- The worked example takes three controls and passes every one, since the
  floor is the union over every pair.
- The preconditions attribute enforces are named, including the anchor to
  the committed state path list and --state-paths-override.
- The block's contents are listed as they now are, and it carries the
  sentence that a PASS does not mean the layer ships no secret it should not.
- The verdicts are described as they are: noise only where the controls pin
  the varying field, overlapping for anything else at a varying position and
  for every binary, real where the controls do not vary, and waivers that
  clear only overlapping.
- The first real run is recorded: 0 attributable, FAIL on 32 real state
  paths that are the Aria header wall clock, open in Keel-Linux/buildtasks#13.

The trap entry's Fix now says not to cancel by position, instead of
prescribing it.
@marcos-mendez
marcos-mendez deleted the feat/layer-measure branch September 29, 2026 06:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The three-build comparison decision 0010 requires has no committed implementation

1 participant