[Docs] Document on-demand TPU slice provisioning for OMENative - #24
Conversation
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
|
Important Review skippedBot user detected. To trigger a single review, invoke the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
|
Documentation maintenance: ready. Unsuccessful content rounds: 0/3. Operational attempts: 0/3. Head: The diff adds exactly one new guide (guides/omenative/provision-tpu-slices.md) with its nav.ts entry and guides/index.md card, all serving the single scheduling/tpu-slice-provisioning concern; no code or generated reference/api edits, and no redirect is needed since no Hugo page covers TPU slices. Every substantive claim was verified against the pinned OME checkout: the tpuSliceProvisioning config contract (default {} off, all fields required, startup-only read, manager exit on a malformed block with the quoted log line), the exact opt-in annotation semantics on the default/leader runner template, slice-per-slot and surge behavior, slice naming/labels/owner-UID-only ownership, readyStates gating and node-selector confinement, the CapacityProvisioning hold parking InstanceReadyTimeout and its op-hold/last-failure presentation, webhook introduce-only rejection with ComponentReconcileError on the ISVC and the controller's SliceDemandInvalid repeat, all six troubleshooting reason strings verbatim from demand.go/shape.go, SliceHostUnavailable/SliceOwnershipConflict message formats, the five metrics' semantics, the CRD-absent log and restart requirement, the teardown-deadline warning text, and the chart's configmap/checksum rollout. The YAML examples tile their topologies exactly per Tile() and use real API fields. The Step 1 sentence 'the chart's with a two-host topology added' is correct: the chart example lists only ["2x2x1"]. Style, anchors, helm-command form and since: v1.3 match repo conventions. No feedback threads exist in the context.
Human review threads and CODEOWNER approval remain under repository policy. |
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
What this PR does
How do I serve an OMENative component on GKE TPU node pools whose slices are provisioned on demand?
Why we need it
Source change: ome-projects/ome@9534e7f
Commit 2 added the complete feature, present in the current checkout: pkg/controller/v1beta1/sliceprovision (provision.go, demand.go, placer.go, metrics.go) provisions cluster-scoped GKE Slice objects (accelerator.gke.io/v1beta1, pkg/tpuslice/gke/slice.go), one slice per slot (a multi-pod Instance shares one; a single-pod Instance gets one per pod ordinal), gates a slot's pod creation on slice readiness, confines pods with a node selector on the slice name, and releases ready slices no pod holds. Configuration is the required-all-fields ome.controller.tpuSliceProvisioning Helm value / tpuSliceProvisioning key of inferenceservice-config (charts/ome-resources/values.yaml:490-521, default {} = off; a malformed block stops the manager). Pods opt in with the ome.io/tpu-slice-provisioning: "true" template annotation (pkg/constants/constants.go:125-128) and declare the shape through the configured accelerator and topology node-selector labels; pkg/tpuslice/shape.go requires the pods' chip requests to fill the topology exactly, and the InferenceReplica webhook rejects runners whose demand cannot be placed (pkg/webhook/admission/inferencereplica/validator.go, validateIRTPUSlices). Observable via ome_tpu_slice_created_total, released_total, create_failures_total and provision_duration_seconds (pkg/controller/v1beta1/sliceprovision/metrics.go). No page under src/lib/content mentions TPU slices (grep for TPU/tpuSlice matches only quota resource names), and no Hugo page covers it, so no redirects.json entry is needed. This is one concern: one feature answering one task question. It falls under scheduling-and-capacity rather than CLI, model storage, runtime selection or rollout orchestration: it provisions accelerator capacity and decides where OMENative pods may run.
Scope: scheduling / tpu-slice-provisioning. Other concerns are deferred.
How to test
git diff --checkand website content/link tests, type checks, lint and production build.Checklist
git commit -s)pnpm lint && pnpm check && pnpm test && pnpm buildpasses (run by the publisher on an isolated copy)