Skip to content

todo: collect dependents of dual-version kinds promptly after a worker outage #601

Description

@thxCode

What was left undone:

Owner-reference garbage collection of the kinds served at two versions — a v1alpha1 CRD and the
worker's aggregated v1 API (Devices, Instance, InstanceType, ModelDeployment and the other
kinds registered on both) — pauses whenever the worker is unavailable, and resumes only on the
garbage collector's own backoff after the worker is back. #592 covered Devices alone, by having the
worker delete a node's Devices itself once the Node is confirmed gone. The other kinds have no
such cleanup, and whether each of them needs one has not been decided.

Where it came from (PR, commit, or spec section):

Found while investigating the stale Devices that #592 fixes.

How this was known (tick one — this is a different question from the line above, which asks
where the work is; this one asks who noticed it was missing):

  • The change knew — the pull request deferred this deliberately and said so at the time
  • Worked out afterwards — somebody read the change later and derived that it had been left

What was read and measured:

  • On a local kind cluster, with the worker pinned to the node that was then stopped and deleted, the
    garbage collector's delete of the orphaned Devices first hung for about 23 s, then failed with
    unable to get REST mapping for worker.gpustack.ai/v1/Devices, and retried on its exponential
    backoff (5 ms up to 1000 s). After the worker came back, the object was deleted about 108 s later.
  • The garbage collector resolves each kind through the discovery-preferred version, which for these
    kinds is the aggregated v1, so its delete requests are served by the worker.
  • With the worker healthy, deleting the Node removed its Devices within 1-5 s in five scenarios:
    the owner reference itself is correct.

What is inferred and not measured: every other kind served at both versions behaves the same way,
because the mechanism is the preferred-version routing and not anything specific to Devices.

What it needs in order to run (hardware, cluster shape, driver version, credentials, tooling):

A kind cluster with at least two workers, the operator's worker pinned to one of them, and an owner
chain for each kind under test (for example an InstanceType whose owner is deleted while the worker
is down). No accelerator is needed.

How to verify it (the exact case or command, and the figure that would prove it):

For each kind: stop the worker, delete the owner, bring the worker back, and measure the time from
the worker's first Ready to the dependent's deletion. The gap shows the pause; a per-kind cleanup, or
a change to how the kinds are registered, should bring it close to the healthy-worker 1-5 s.

What is unproven until then:

That dependents of every dual-version kind other than Devices are collected promptly after a worker
outage. Until then they may linger for up to the garbage collector's maximum backoff, and a status
view derived from them (the way the stale Devices inflated an InstanceType's capacity) can read
high for that long.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    todoWork a pull request knowingly left undone

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions