What was left undone:
Owner-reference garbage collection of the kinds served at two versions — a v1alpha1 CRD and the
worker's aggregated v1 API (Devices, Instance, InstanceType, ModelDeployment and the other
kinds registered on both) — pauses whenever the worker is unavailable, and resumes only on the
garbage collector's own backoff after the worker is back. #592 covered Devices alone, by having the
worker delete a node's Devices itself once the Node is confirmed gone. The other kinds have no
such cleanup, and whether each of them needs one has not been decided.
Where it came from (PR, commit, or spec section):
Found while investigating the stale Devices that #592 fixes.
How this was known (tick one — this is a different question from the line above, which asks
where the work is; this one asks who noticed it was missing):
What was read and measured:
- On a local kind cluster, with the worker pinned to the node that was then stopped and deleted, the
garbage collector's delete of the orphaned Devices first hung for about 23 s, then failed with
unable to get REST mapping for worker.gpustack.ai/v1/Devices, and retried on its exponential
backoff (5 ms up to 1000 s). After the worker came back, the object was deleted about 108 s later.
- The garbage collector resolves each kind through the discovery-preferred version, which for these
kinds is the aggregated v1, so its delete requests are served by the worker.
- With the worker healthy, deleting the
Node removed its Devices within 1-5 s in five scenarios:
the owner reference itself is correct.
What is inferred and not measured: every other kind served at both versions behaves the same way,
because the mechanism is the preferred-version routing and not anything specific to Devices.
What it needs in order to run (hardware, cluster shape, driver version, credentials, tooling):
A kind cluster with at least two workers, the operator's worker pinned to one of them, and an owner
chain for each kind under test (for example an InstanceType whose owner is deleted while the worker
is down). No accelerator is needed.
How to verify it (the exact case or command, and the figure that would prove it):
For each kind: stop the worker, delete the owner, bring the worker back, and measure the time from
the worker's first Ready to the dependent's deletion. The gap shows the pause; a per-kind cleanup, or
a change to how the kinds are registered, should bring it close to the healthy-worker 1-5 s.
What is unproven until then:
That dependents of every dual-version kind other than Devices are collected promptly after a worker
outage. Until then they may linger for up to the garbage collector's maximum backoff, and a status
view derived from them (the way the stale Devices inflated an InstanceType's capacity) can read
high for that long.
What was left undone:
Owner-reference garbage collection of the kinds served at two versions — a
v1alpha1CRD and theworker's aggregated
v1API (Devices,Instance,InstanceType,ModelDeploymentand the otherkinds registered on both) — pauses whenever the worker is unavailable, and resumes only on the
garbage collector's own backoff after the worker is back. #592 covered
Devicesalone, by having theworker delete a node's
Devicesitself once theNodeis confirmed gone. The other kinds have nosuch cleanup, and whether each of them needs one has not been decided.
Where it came from (PR, commit, or spec section):
Found while investigating the stale
Devicesthat #592 fixes.How this was known (tick one — this is a different question from the line above, which asks
where the work is; this one asks who noticed it was missing):
What was read and measured:
garbage collector's delete of the orphaned
Devicesfirst hung for about 23 s, then failed withunable to get REST mapping for worker.gpustack.ai/v1/Devices, and retried on its exponentialbackoff (5 ms up to 1000 s). After the worker came back, the object was deleted about 108 s later.
kinds is the aggregated
v1, so its delete requests are served by the worker.Noderemoved itsDeviceswithin 1-5 s in five scenarios:the owner reference itself is correct.
What is inferred and not measured: every other kind served at both versions behaves the same way,
because the mechanism is the preferred-version routing and not anything specific to
Devices.What it needs in order to run (hardware, cluster shape, driver version, credentials, tooling):
A kind cluster with at least two workers, the operator's worker pinned to one of them, and an owner
chain for each kind under test (for example an
InstanceTypewhose owner is deleted while the workeris down). No accelerator is needed.
How to verify it (the exact case or command, and the figure that would prove it):
For each kind: stop the worker, delete the owner, bring the worker back, and measure the time from
the worker's first Ready to the dependent's deletion. The gap shows the pause; a per-kind cleanup, or
a change to how the kinds are registered, should bring it close to the healthy-worker 1-5 s.
What is unproven until then:
That dependents of every dual-version kind other than
Devicesare collected promptly after a workeroutage. Until then they may linger for up to the garbage collector's maximum backoff, and a status
view derived from them (the way the stale
Devicesinflated anInstanceType's capacity) can readhigh for that long.