Skip to content

Rebuild reprepro's database from the pool before publishing - #36

Merged
guys-inc-ops[bot] merged 1 commit into
linuxfrom
fix/publish-rebuild-db-from-pool
Aug 27, 2026
Merged

guys-inc-ops[bot] merged 1 commit into
linuxfrom
fix/publish-rebuild-db-from-pool

Conversation

@guys-inc-ops

@guys-inc-ops guys-inc-ops Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Excluding db/ from the bucket was the right call — earlier runs served reprepro's working state publicly. But it reintroduced, by another route, exactly the failure #33 was written to prevent.

The problem

reprepro export writes dists/ from the database, not from the pool. db/ is the one thing the state pull cannot bring back, so every run starts with an empty database and regenerates the indexes from only the packages it included itself.

Today that is invisible, which is why it survived review — all three architectures ship in one run, and reprepro keeps one version per package anyway, so the output is correct by coincidence rather than by construction. I verified that: a persisted db and a lost db produce identical output for the current flow. Nothing is broken right now.

It stops being correct the moment a run publishes a subset — one architecture failing to build, a re-publish of a single asset, or a second package in the repo. Measured in a debian:bookworm container, publishing both arches at 3.4.9 and then amd64 alone at 3.4.10:

run 1: publish both arches at 3.4.9      index: amd64=3.4.9  arm64=3.4.9
run 2, db LOST (today):  amd64 3.4.10    index: amd64=3.4.10 arm64=<empty>
run 2, db KEPT (control): amd64 3.4.10   index: amd64=3.4.10 arm64=3.4.9

arm64 users lose the package and the run stays green.

The fix

Reconstruct the database from what the bucket already holds, before including anything new. Verified rather than assumed:

  • reprepro exits 0 and skips a package it already has (Skipping inclusion … as it has already '2.0'), so a re-publish of the same tag stays idempotent — confirmed by running the same tag twice: exit 0, index and pool unchanged.
  • Including a deb from the pool does not delete the file it is reading. I checked this specifically, because reprepro does print Deleting files just added to the pool but not used in the superseded case.
  • It also fixes a smaller live problem: superseded pool objects are currently never removed, because nothing knows they exist. With the rebuild, github-desktop_3.4.9_amd64.deb is cleaned up when 3.4.10 supersedes it.

Running the two new steps verbatim against the failure scenario:

===== WITHOUT the fix =====        ===== WITH the fix =====
amd64: 1 listed, 2 in pool         restored 3 package(s) from the pool
arm64: 0 listed, 1 in pool         amd64: 1 listed, 1 in pool
armhf: 0 listed, 1 in pool         arm64: 1 listed, 1 in pool
guard exit=1                       armhf: 1 listed, 1 in pool
                                   guard exit=0

The guard

The second step is the one that would have caught this. An empty index is a valid, correctly signed index — indistinguishable downstream from a repository that genuinely has no packages — so the check has to happen here, while the pool is still around to compare against.

Worth flagging: my first version of that guard was broken and passed the very case it exists to catch. grep -c prints 0 and exits 1 on a zero-match file, so the || echo 0 fallback produced "0\n0" and the comparison silently stopped working. Caught by running the step against the real failure rather than trusting it. Now: exit 1 with two ::error:: lines on the bug, exit 0 clean on a healthy repo.

Excluding db/ from the bucket was right - earlier runs served reprepro's
working state publicly. But `export` writes dists/ from the database, not
from the pool, and the database is the one thing the state pull cannot bring
back. Every run therefore starts empty and regenerates the indexes from only
the packages it included itself.

Today that is invisible, which is why it survived review: all three
architectures ship in one run, and reprepro keeps a single version per
package anyway, so the output is correct by coincidence rather than by
construction. It stops being correct the moment a run publishes a subset -
one architecture failing to build, a re-publish of a single asset, or a
second package in the repository.

Measured in a bookworm container, publishing both arches at 3.4.9 and then
amd64 alone at 3.4.10:

  db lost (today):  amd64=3.4.10  arm64=<empty>
  db kept:          amd64=3.4.10  arm64=3.4.9

arm64 users lose the package, and the run stays green. That is the same
failure the state-pull guard exists to prevent, arriving by another route.

So reconstruct the database from what the bucket already holds before adding
anything new. reprepro exits 0 and skips a package it already has, so a
re-publish of the same tag is still idempotent, and it does not delete the
pool file it is reading from - both verified rather than assumed. It also
fixes a smaller live problem: superseded pool objects are currently never
removed, because nothing knows they exist.

The second step is the one that would have caught this. An empty index is a
valid, correctly signed index, indistinguishable downstream from a
repository that genuinely has no packages, so the check has to happen here,
while the pool is still around to compare against.
@guys-inc-ops
guys-inc-ops Bot requested a review from Cam8863 as a code owner August 27, 2026 12:52
@guys-inc-ops
guys-inc-ops Bot merged commit 22997cc into linux Aug 27, 2026
6 checks passed
@guys-inc-ops
guys-inc-ops Bot deleted the fix/publish-rebuild-db-from-pool branch August 27, 2026 14:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant