You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
One place to see every cross-vendor and Ascend verification this project owes, what hardware each row
needs, and — for each row — what does not answer it.
This consolidates #216 and #333, which were separate carriers for rows of one matrix. Their
"what does not close this" paragraphs are reproduced verbatim below, because those paragraphs are
what the issues were actually worth: each names the observation that looks like success and is not.
The rows are not interchangeable and two of them are not even about two vendors. Reading a green
result on one row as an answer to another is the failure this issue exists to prevent, so each row
states its own question.
The matrix
#
question
vendors
hardware
state
R1
Does an Ascend engine actually serve with the configuration this operator injects?
R1 blocks R2. A single Ascend engine that has not been shown to serve with our injected
configuration cannot be one half of a mutual-read test: a failure would not distinguish "the two
halves cannot share" from "that half was never serving".
R3 is a different layer and does not block either. It answers what this operator writes into an
engine's configuration. Whether the two halves can then share anything is not something a fixture
reaches.
R1 — an Ascend engine serving with the configuration we inject
From #333. case-60's vLLM-Ascend row asks two questions and its header separates them:
"Is the name we render in that engine's registry?" — ANSWERED, and answered better without a
cluster: one docker run --rm per image, listing the whole registry rather than probing for one
name. Verbatim: "Re-running it through a Pod adds no information."
"Does the engine actually run with the configuration we inject?" — NOT ANSWERABLE on a machine
with no accelerator, structurally. Measured: vllm-ascend aborts on torch_npu before any vLLM
module loads.
A cluster run without Ascend hardware, for the three reasons quoted above.
Building the second Ascend-transport backend and pool. That is a means; on hardware-less clusters
it cannot answer question 2, and on real hardware the ModelDeployment path answers it more
directly.
Re-running the registry listing. That is question 1, already answered, and answered better offline.
Removing the skipped row. The row is the record that question 2 is unverified.
Two pointers still have to move, and neither is done (carried over from #333, unchanged):
pkg/worker/kvcache/inject/engine.go:198-203 carries a "HOW TO EXERCISE A ROW" note saying the
vLLM-Ascend transport row is "reachable only through the ModelDeployment path", and warning that
trying it from a Pod annotation gets refused by ParseEngine "on the engine name and not on the
transport". Accurate for the vllm-ascend spelling, but it omits the manufacturer spelling
(engine: vllm + manufacturer: ascend), which pod_kv_cache_inject_test.go:90 shows is admitted
on an Ascend pool. A reader following that note concludes the annotation route does not exist at
all.
Until those move, a reader arriving through the code still lands on a closed issue.
Whether heterogeneous clients — the -npu build and the -cuda13 build — can share one master and read each other's segments. S2's own words: "Verified only as loading."
Both images starting does not count. Loading is what has already been verified; the assertion
here is mutual reads — one build writing a segment the other build can read back.
Hardware: same set of machines as the Ascend work. Stating the requirement in those same words on
purpose, so the two are recognisable as one hardware ask rather than two.
R3 — what the operator renders for a cross-manufacturer pair
Stays in #435 as its item 3, alongside the two other items of that issue. Listed here so this matrix
is complete, not to move it. It is fixture-only: no hardware, no cluster, four independent
observables (admission, entrance and credits, image backend, connector name) which can disagree with
each other, which is why they are kept apart.
What a reader should not conclude from this issue
That a green run on R1 says anything about R2. R1 is one vendor.
That R3 passing means cross-manufacturer P/D works. The fixtures answer what this operator writes
into the engine's configuration; whether the two halves can share is a property of the two engine
projects, and no fixture reaches it.
Closing any single row. The rows have separate blockers and separate hardware; a row that is answered
gets its answer recorded here and struck from the table, and the issue stays open while any row is
unanswered.
One place to see every cross-vendor and Ascend verification this project owes, what hardware each row
needs, and — for each row — what does not answer it.
This consolidates #216 and #333, which were separate carriers for rows of one matrix. Their
"what does not close this" paragraphs are reproduced verbatim below, because those paragraphs are
what the issues were actually worth: each names the observation that looks like success and is not.
The rows are not interchangeable and two of them are not even about two vendors. Reading a green
result on one row as an answer to another is the failure this issue exists to prevent, so each row
states its own question.
The matrix
-npuclient and a-cuda13client share one master and read each other's segments?R1 blocks R2. A single Ascend engine that has not been shown to serve with our injected
configuration cannot be one half of a mutual-read test: a failure would not distinguish "the two
halves cannot share" from "that half was never serving".
R3 is a different layer and does not block either. It answers what this operator writes into an
engine's configuration. Whether the two halves can then share anything is not something a fixture
reaches.
R1 — an Ascend engine serving with the configuration we inject
From #333.
case-60's vLLM-Ascend row asks two questions and its header separates them:cluster: one
docker run --rmper image, listing the whole registry rather than probing for onename. Verbatim: "Re-running it through a Pod adds no information."
with no accelerator, structurally. Measured:
vllm-ascendaborts ontorch_npubefore any vLLMmodule loads.
Only the second is the gap.
What does NOT close this (verbatim from #333):
Ascend-transport backend and pool. That is a means; on hardware-less clustersit cannot answer question 2, and on real hardware the
ModelDeploymentpath answers it moredirectly.
Two pointers still have to move, and neither is done (carried over from #333, unchanged):
case-60's header still says "Tracked with the rest of that family in issue bug: vllm-ascend needs the ascend transport and nothing checks for it #172". That pointer isdead and has to name this issue instead.
pkg/worker/kvcache/inject/engine.go:198-203carries a "HOW TO EXERCISE A ROW" note saying thevLLM-Ascend transport row is "reachable only through the ModelDeployment path", and warning that
trying it from a Pod annotation gets refused by
ParseEngine"on the engine name and not on thetransport". Accurate for the
vllm-ascendspelling, but it omits the manufacturer spelling(
engine: vllm+manufacturer: ascend), whichpod_kv_cache_inject_test.go:90shows is admittedon an
Ascendpool. A reader following that note concludes the annotation route does not exist atall.
Until those move, a reader arriving through the code still lands on a closed issue.
R2 — heterogeneous clients sharing one master
From #216. Layer: L2 · Source: S2.
Whether heterogeneous clients — the
-npubuild and the-cuda13build — can share one master andread each other's segments. S2's own words: "Verified only as loading."
What does NOT fill this (verbatim from #216):
Both images starting does not count. Loading is what has already been verified; the assertion
here is mutual reads — one build writing a segment the other build can read back.
Hardware: same set of machines as the Ascend work. Stating the requirement in those same words on
purpose, so the two are recognisable as one hardware ask rather than two.
R3 — what the operator renders for a cross-manufacturer pair
Stays in #435 as its item 3, alongside the two other items of that issue. Listed here so this matrix
is complete, not to move it. It is fixture-only: no hardware, no cluster, four independent
observables (admission, entrance and credits, image backend, connector name) which can disagree with
each other, which is why they are kept apart.
What a reader should not conclude from this issue
into the engine's configuration; whether the two halves can share is a property of the two engine
projects, and no fixture reaches it.
nodeWideDeviceGroupshas never run on a nodecarrying two accelerator manufacturers — that is todo: the cross-manufacturer merge has never run on a real two-vendor node #181, a different subsystem with a different
observable, and it is deliberately not a row here.
What does NOT close this issue
Closing any single row. The rows have separate blockers and separate hardware; a row that is answered
gets its answer recorded here and struck from the table, and the issue stays open while any row is
unanswered.
/kind cleanup
/area testing