Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 14 additions & 12 deletions .agents/skills/gpustack-operator-docs/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,17 +40,19 @@ one, not to widen the overview.
| A `KVCachePool` or `KVCachePoolBinding`: the grant, the reuse domain, a quota ceiling or grant, what a full quota does | `docs/kv-cache/pool.md` |
| Standing a cache up end to end, or which object comes first: the pasteable four-object sequence | `docs/kv-cache/walkthrough.md` |
| How a **Pod** consumes a pool: the inject label and annotations, the injected keys per engine, a refusal, the isolation record | `docs/reference/kv-cache-injection.md` |
| A `ModelArtifact`: its sources, resolution and revalidation, the manifest digest, how a `ModelDeployment` or an `Instance` mounts or downloads it, claim placement, the weight identity in KV keys | `docs/reference/model-artifact.md` |
| The `image` source of a `ModelArtifact`: the digest contract, building weights into an image, image-volume delivery, the version floors, double storage, kubelet image GC, registry mirrors | `docs/reference/model-image-source.md` |
| A `ModelStore`, `ModelStoreBinding` or `ModelPrefetch`: the grant, the budget, pinning, TTL expiry, the warm-up pod and why it is label-free | `docs/reference/model-prefetch.md` |
| The `v1` views of `ModelArtifact` and `NodeModelStore`, the `progress` subresource and who may read it, the GPUStack server capability map | `docs/reference/model-artifact-views.md` |
| A `NodeModelStore` or the `model-manager` plugin: a field and its writer, the status guard, mount authorization, materialization, a failure reason, collection, a metric | `docs/reference/node-model-store.md` |
| Running node delivery: the chart values, where the node's configuration comes from, reading a node, the watermark cap, switching delivery, where replicas land and turning the preference off, upgrading, removing the cache | `docs/operation/model-store.md` |
| A `ModelDeployment` metrics snapshot, which series each field reads per engine, role and router, cache-hit scope or Pod scrape annotation | `docs/reference/model-deployment-metrics.md` |
| How a prefill role and a decode role are paired: the connector each engine and router renders, `spec.router` and its fields, `spec.kvTransfer`, roles on different hardware, a role's own Service | `docs/reference/model-deployment-prefill-decode.md` |
| Which replica a managed router picks, its default routing policy, switching it through `spec.router.extraArgs`, the router's own per-replica series | `docs/reference/model-deployment-routing.md` |
| What a `ModelDeployment` replica does between its Pod's delete and its engine's exit: the drain hook, its timings, what it does not cover | `docs/reference/model-deployment-shutdown.md` |
| The lowest engine release a deployment shape runs on, the Mooncake client its runner image carries, the store line it needs, which transport each engine can use on each leg | `docs/reference/engine-versions.md` |
| A `ModelArtifact`: its sources, resolution and revalidation, the manifest digest, how a `ModelDeployment` or an `Instance` mounts or downloads it, claim placement, the weight identity in KV keys | `docs/model-store/artifact.md` |
| The `image` source of a `ModelArtifact`: the digest contract, building weights into an image, image-volume delivery, the version floors, double storage, kubelet image GC, registry mirrors | `docs/model-store/image-source.md` |
| A `ModelStore`, `ModelStoreBinding` or `ModelPrefetch`: the grant, the budget, pinning, TTL expiry, the warm-up pod and why it is label-free | `docs/model-store/prefetch.md` |
| The `v1` views of `ModelArtifact` and `NodeModelStore`, the `progress` subresource and who may read it, the GPUStack server capability map | `docs/model-store/views.md` |
| A `NodeModelStore` or the `model-manager` plugin: a field and its writer, the status guard, mount authorization, materialization, a failure reason, collection, a metric | `docs/model-store/node-store.md` |
| Running node delivery: the chart values, where the node's configuration comes from, reading a node, the watermark cap, switching delivery, where replicas land and turning the preference off, upgrading, removing the cache | `docs/model-store/operations.md` |
| The `ModelDeployment` contract: the inherited reuse domain, the three override tiers, the owned-key table, the runner-image formula, prefill/decode pairing, the topology-placement field contract | `docs/model-deployment/deployment.md` |
| A `ModelDeployment` metrics snapshot, which series each field reads per engine, role and router, cache-hit scope or Pod scrape annotation | `docs/model-deployment/metrics.md` |
| What a `ModelDeployment` status condition or published field means, and how to read them when a deployment misbehaves | `docs/model-deployment/status.md` |
| How a prefill role and a decode role are paired: the connector each engine and router renders, `spec.router` and its fields, `spec.kvTransfer`, roles on different hardware, a role's own Service | `docs/model-deployment/prefill-decode.md` |
| Which replica a managed router picks, its default routing policy, switching it through `spec.router.extraArgs`, the router's own per-replica series | `docs/model-deployment/routing.md` |
| What a `ModelDeployment` replica does between its Pod's delete and its engine's exit: the drain hook, its timings, what it does not cover | `docs/model-deployment/shutdown.md` |
| The lowest engine release a deployment shape runs on, the Mooncake client its runner image carries, the store line it needs, which transport each engine can use on each leg | `docs/model-deployment/engine-versions.md` |
| A resource key, a request rule, a request example | `docs/accelerator-requests.md` |
| How many RDMA endpoints a workload asks for, setting or reading the kubelet TopologyManager policy, what to do about RDMA keys no queue meters, the engine image an EFA leg needs | `docs/operation/rdma.md` |
| Enabling Topograph, publishing topology snapshots, webhook trust, requesting a level, TAS diagnosis or EKS validation | `docs/operation/topology-aware-scheduling.md` |
Expand Down Expand Up @@ -89,7 +91,7 @@ reader's trust.
| `deploy/gpustack-operator/chart/README.md`, `values.schema.json` | `values.yaml` + `README.md.gotmpl` via `make generate chart` | generated — never hand-edit; a doc path quoted in a `values.yaml` comment needs a regenerate, and `chart.yml` fails on drift |
| `README.md` accelerator matrix | `pkg/nodefeature/knowns.go` (resource names, `SharedResourceMaxSize`, `_ManufacturerPartitionKindMap`) and **whether** each `pkg/devicemanager/detector/<mfr>/device.go` sets `LogicalSliced` at all | nothing fails; the table silently lies about what a vendor can do. The matrix is deliberately ✅/— only — per-card slice counts and the per-vendor isolation mechanism live in `docs/architecture/device-discovery.md`, not on the front page |
| `README.md` Quick Start's four request shapes | `docs/accelerator-requests.md` — *The resource keys* and *Worked example per family* | nothing fails; the front page and the normative contract drift apart. This copy is the one sanctioned exception to "state a fact once" (the README is the shop window) — change both together |
| `docs/reference/model-deployment.md` owned-key table | `modelDeploymentOwnedKeys` via `TestModelDeploymentOwnedKeysDocs` | the test matches each owned key inside its engine's **table row**, not anywhere on the page. One-way: a key the code owns must be documented; a name the page merely explains is free. It exists because the code side already had an invariant and the page had none, and the page then omitted a security-relevant key while the webhook refused it |
| `docs/model-deployment/deployment.md` owned-key table | `modelDeploymentOwnedKeys` via `TestModelDeploymentOwnedKeysDocs` | the test matches each owned key inside its engine's **table row**, not anywhere on the page. One-way: a key the code owns must be documented; a name the page merely explains is free. It exists because the code side already had an invariant and the page had none, and the page then omitted a security-relevant key while the webhook refused it |
| `docs/settings.md` tables | `pkg/worker/settings` and the `GPUSTACK_*` readers | nothing fails; an operator configures something that no longer exists |
| `docs/README.md` page table | the set of files under `docs/` | `check-docs.sh` fails |
| `docs/README.md` page **labels** | each page's `#` H1, character for character | `check-docs.sh` fails; a page's file name, H1 and index label are one set of words |
Expand Down
17 changes: 12 additions & 5 deletions .agents/skills/gpustack-operator-docs/references/conventions.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,10 +65,16 @@ shaped by the directory it lives in:
| Directory | H1 form | Example |
|---|---|---|
| root, `architecture/` | `<Subject>` | `Installation Modes` |
| a domain directory — `kv-cache/`, `model-store/`, `model-deployment/` | `<Domain> <Topic>` where the domain reads naturally; a subject that names itself (`Engine Versions`, `Node-to-Node Sync`) stands without it; the domain's runbook keeps the `Operations` suffix | `KV Cache Backend`, `Model Artifact`, `Engine Versions`, `Node-to-Node Sync`, `Model Store Operations` |
| `operation/` | `<Subject> Operations` | `High Availability Operations` |
| `migration/` | `Migrating <from\|to> <what>`; a recovery page is `<Subject> Troubleshooting` | `Migrating from v0.5.x`, `Migration Troubleshooting` |
| `reference/` | `<Subject> Reference` | `Instance Metrics Reference` |

A domain directory collects every page orbiting one CR family — the contracts, the field references,
the views, the runbook — under short file names (`backend.md`, `artifact.md`, `deployment.md`) whose
H1s name the domain where it reads naturally. It exists so `reference/` stays true lookup tables
rather than becoming the dumping ground for whichever domain landed last.

`##` and `###` headings are sentence case. GitHub lowercases anchors, so re-casing a heading keeps every
inbound link; changing its *words* does not.

Expand Down Expand Up @@ -105,10 +111,10 @@ it as follows — when a page starts serving two modes at once, that is the mome

| Mode | Reader is… | Our pages |
|---|---|---|
| Tutorial | learning by doing | `README.md` Quick Start, `docs/walkthrough.md`, the MIG walkthrough |
| How-to | achieving a goal | `docs/operation/*`, `docs/migration/*`, `docs/development.md` |
| Tutorial | learning by doing | `README.md` Quick Start, `docs/walkthrough.md`, the MIG walkthrough, `docs/kv-cache/walkthrough.md` |
| How-to | achieving a goal | `docs/operation/*`, `docs/migration/*`, `docs/development.md`, a domain runbook (`docs/model-store/operations.md`) |
| Reference | looking something up | `docs/accelerator-requests.md`, `docs/settings.md`, `docs/reference/*` |
| Explanation | building understanding | `docs/architecture.md` and `docs/architecture/*` |
| Explanation | building understanding | `docs/architecture.md`, `docs/architecture/*`, and the domain pages under `docs/kv-cache/`, `docs/model-store/` and `docs/model-deployment/` (a domain page serves the mode its reader arrives in — contract pages read as reference, mechanism pages as explanation) |

Two consequences worth stating:

Expand Down Expand Up @@ -158,8 +164,9 @@ Two consequences worth stating:

## Adding a page

1. Put it under the directory of the reader it serves (`architecture/`, `operation/`, `migration/`,
`reference/`).
1. Put it under the directory that fits: the reader it serves (`architecture/`, `operation/`,
`migration/`, `reference/`), or the domain directory of the CR family it orbits
(`kv-cache/`, `model-store/`, `model-deployment/`) when the page joins a family that already has one.
2. Copy the template above; fill the header block honestly — an inflated read time is worse than none.
3. Add a row to the `docs/README.md` page table, and a step to any reading path it belongs on.
4. Add it to the routing table in the skill's `SKILL.md` and to `references/page-map.md`, saying what it
Expand Down
Loading
Loading