Skip to content

Give the boot test container a systemd the first boot can rely on - #5

Merged
marcos-mendez merged 3 commits into
mainfrom
test/container-marks-and-apparmor
Sep 27, 2026
Merged

marcos-mendez merged 3 commits into
mainfrom
test/container-marks-and-apparmor

Conversation

@marcos-mendez

Copy link
Copy Markdown
Contributor

The other half of what the gate found once the layer booted. The layer's
own half, the cluster listening on the IPv4 loopback alone, is fixed in
the pull request before this one; this is the container the test builds.

Under the stock LXC container apparmor profile systemd cannot give a unit
a mount namespace, so every unit that asks for one fails with
status=226/NAMESPACE before its own first line runs. Measured on the
build host in a container booted exactly as this test boots it,
systemd-journald, systemd-logind, systemd-sysusers, systemd-sysctl and
tmp.mount had all failed that way. That is why 15regen-sslcert and
95secupdates failed beside the database, and why the container had no
journal to read when it was asked why. The cluster itself survived it,
because postgresql@.service asks for no namespace, which is exactly what
made the report confusing: the database was up and unreachable for an
unrelated reason.

Two things fix it, both ported from keel-nodebb rather than invented
again, since that repository met both first and keel-mariadb is taking
the same two:

bt_lxc_config now writes lxc.apparmor.profile = generated and
lxc.apparmor.allow_nesting = 1, which is what the appliance containers
on the build host have carried all along (keel-nodebb pull request 8).

bt_mark_container replaces the single marker line and does all three
things a container build does: the marker under /var/lib/turnkey-info,
REDIRECT_OUTPUT=true in etc/default/inithooks, and a drop-in giving
inithooks.service StandardOutput=journal. The layer ships the plain
appliance unit, which runs the hooks on /dev/tty1; nothing reads tty1
in a container nobody has attached to, so a hook that prints past the
terminal buffer blocks there forever. It hung keel-nodebb's first boot
(pull request 9). No hook here prints that much today, which is exactly
why it should not be left for the next one to find.

Measured on the build host against the published chain with both in
place, and with the layer's listen_addresses corrected: every hook from
01ipconfig to 98finalize completes, including 15regen-sslcert and
95secupdates, and the only failed unit is
inithooks-restart-getty1.service, which wants a tty no container has.

Seven tests for the two changes; boot-test-lib.sh stays at 100 percent
(154/154) and the three measured files total 100 (210/210) over 65 bats
tests. Nothing in the layer changes here, so no changelog entry.

The other half of what the gate found once the layer booted. The layer's
own half, the cluster listening on the IPv4 loopback alone, is fixed in
the pull request before this one; this is the container the test builds.

Under the stock LXC container apparmor profile systemd cannot give a unit
a mount namespace, so every unit that asks for one fails with
status=226/NAMESPACE before its own first line runs. Measured on the
build host in a container booted exactly as this test boots it,
systemd-journald, systemd-logind, systemd-sysusers, systemd-sysctl and
tmp.mount had all failed that way. That is why 15regen-sslcert and
95secupdates failed beside the database, and why the container had no
journal to read when it was asked why. The cluster itself survived it,
because postgresql@.service asks for no namespace, which is exactly what
made the report confusing: the database was up and unreachable for an
unrelated reason.

Two things fix it, both ported from keel-nodebb rather than invented
again, since that repository met both first and keel-mariadb is taking
the same two:

  bt_lxc_config now writes lxc.apparmor.profile = generated and
  lxc.apparmor.allow_nesting = 1, which is what the appliance containers
  on the build host have carried all along (keel-nodebb pull request 8).

  bt_mark_container replaces the single marker line and does all three
  things a container build does: the marker under /var/lib/turnkey-info,
  REDIRECT_OUTPUT=true in etc/default/inithooks, and a drop-in giving
  inithooks.service StandardOutput=journal. The layer ships the plain
  appliance unit, which runs the hooks on /dev/tty1; nothing reads tty1
  in a container nobody has attached to, so a hook that prints past the
  terminal buffer blocks there forever. It hung keel-nodebb's first boot
  (pull request 9). No hook here prints that much today, which is exactly
  why it should not be left for the next one to find.

Measured on the build host against the published chain with both in
place, and with the layer's listen_addresses corrected: every hook from
01ipconfig to 98finalize completes, including 15regen-sslcert and
95secupdates, and the only failed unit is
inithooks-restart-getty1.service, which wants a tty no container has.

Seven tests for the two changes; boot-test-lib.sh stays at 100 percent
(154/154) and the three measured files total 100 (210/210) over 65 bats
tests. Nothing in the layer changes here, so no changelog entry.
@marcos-mendez
marcos-mendez merged commit 0f0d205 into main Sep 27, 2026
2 of 3 checks passed
@marcos-mendez
marcos-mendez deleted the test/container-marks-and-apparmor branch September 27, 2026 07:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant