Skip to content

fix(csi): repair a controller that is connected but carries no path - #411

Closed
schmidt-scaled wants to merge 1 commit into
mainfrom
fix/csi-orphan-controller-repair
Closed

schmidt-scaled wants to merge 1 commit into
mainfrom
fix/csi-orphan-controller-repair

Conversation

@schmidt-scaled

Copy link
Copy Markdown
Contributor

Problem

A controller can be connected while contributing no path to the namespace head. Its namespace never joined, so the device-scoped nvme list-subsys does not list it and numActive stays below expected — but nvme connect refuses with already connected, because at the controller level nothing is missing. The two views never reconcile and reconcileNonOptimizedPaths spins forever.

This state is invisible to every layer: nvme connect succeeds, the target establishes qpairs, the client prints new ctrl, and no keep-alive timeout or controller reset ever fires — the path simply does not exist for I/O routing.

K8sNativeResilientFailoverTest iteration 28 (2026-08-09): 106 such retries over 11 minutes for volume 638be965, not one of which produced a single packet to the target. The volume ran at 2 of 3 paths the whole time and lost all I/O when the outage removed the other two — no available path - failing I/O, ext4 remounted read-only, FIO err=5.

Across that 42-hour run the same state affected many volumes (78726d0e 608 degraded reports, 64467bfa 437, 638be965 308), so the cluster was routinely below its configured redundancy with no operator-visible signal.

Change

Two call sites (reconcileNonOptimizedPaths, reconcileOptimizedPath) now call connectMissingPath() instead of connectViaNVMe().

connectMissingPath() tries the ordinary connect first and returns immediately unless it fails with exactly already connected — so the normal path is unchanged. Only then does it:

  1. locate the orphaned controller via a host-wide nvme list-subsys -o json — deliberately without a device argument, since the device-scoped form is precisely what omits it;
  2. nvme disconnect -d <ctrl> (same command shape as the existing disconnectDevicePath);
  3. reconnect, forcing a fresh controller and a fresh namespace scan.

A 5-minute per-(NQN, IP:port) cooldown keeps a repair that does not stick from becoming a disconnect/reconnect loop at monitor cadence.

Supporting helpers: findControllerForAddress(), isAlreadyConnectedErr(), parseTrsvcid().

Related

The control-plane side of this incident — the source of the pathless subsystems — is fixed separately in sbcli 97ef965c2.

Reviewer notes

  • NOT COMPILED. No Go toolchain was available in the environment this was written in. Every referenced symbol and signature was verified by inspection, but go build ./... and go vet ./pkg/util/... have not been run, and no test was added. Please verify before merging.
  • Step 1 assumes host-wide nvme list-subsys reports a controller whose namespace never joined the head, while the device-scoped form does not. This is consistent with the incident data (three controllers existed, numActive was 2) but was not verified against a live host. If the assumption is wrong, findControllerForAddress returns not-found and the caller emits a clearer error instead of silently spinning — degraded, not harmful.
  • On a shared subsystem (several lvols behind one NQN), disconnecting the controller also drops its sibling namespaces' paths until the reconnect completes. They retain their other paths meanwhile, so it is a brief single-path dip — this is the reason for the cooldown.
  • isAlreadyConnectedErr matches on error text, which works because execWithTimeout folds CombinedOutput into the returned error.

The path-recovery loop cannot fix the one failure mode that matters most.
A controller can be connected while contributing NO path to the namespace
head: its namespace never joined, so the device-scoped `nvme list-subsys`
does not list it and numActive stays below expected -- but `nvme connect`
refuses with "already connected", because at the controller level nothing
is missing. The two views never reconcile and reconcileNonOptimizedPaths
spins forever.

K8sNativeResilientFailoverTest iteration 28 (2026-08-09): 106 such
retries over 11 minutes for volume 638be965, not one of which produced a
single packet to the target -- the target log shows no connect attempt
after the original one. The volume ran at 2 of 3 paths the whole time and
lost all I/O when the outage removed the other two: "no available path -
failing I/O" and ext4 remounted read-only. Across that 42-hour run the
same state affected many volumes (78726d0e 608 degraded reports, 64467bfa
437, 638be965 308), so the cluster was routinely below its configured
redundancy with no operator-visible signal.

connectMissingPath() tries the ordinary connect first and returns
immediately unless it fails with exactly "already connected", so the
normal path is unchanged. Only then does it locate the orphaned
controller via a host-wide `nvme list-subsys` -- deliberately without a
device argument, since the device-scoped form is precisely what omits it
-- disconnect it, and reconnect, forcing a fresh controller and a fresh
namespace scan. A 5-minute per-(NQN, IP:port) cooldown keeps a repair
that does not stick from becoming a disconnect/reconnect loop at monitor
cadence.

The control-plane side of this incident (the source of the pathless
subsystems) is fixed separately in sbcli 97ef965c2.

NOT COMPILED: no Go toolchain was available in the environment this was
written in. Every referenced symbol and signature was verified by
inspection, but `go build ./...` and `go vet ./pkg/util/...` have not been
run, and no test was added. Please verify before merging.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noctarius

Copy link
Copy Markdown
Collaborator

Superseded by #429

@noctarius noctarius closed this Aug 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants