Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23365
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 New Failure, 1 Unclassified FailureAs of commit 4bef881 with merge base c16dd0a ( NEW FAILURE - The following job has failed:
UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
This was referenced Oct 2, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack from ghstack (oldest at bottom):
Moves the byte layer of the off-graph KV cache -- per-layer K/V allocation,
geometric growth, cross-stream ordering, AOTI constant discovery and binding,
and the CUDA-graph recapture a growth forces -- out of CudaSequenceKVCache into
CudaKVPool. CudaSequenceKVCache keeps what is specific to one sequence: the
neutral SequenceCache planner, admission, and the logical length. This mirrors
MLX, where MLXSequenceCache and MLXCellCache share one Pool.
The pool is what a batched layout needs next. Besides growable and fixed
layers it holds fixed-size side buffers a layout declares by FQN; they are
allocated once, start zeroed, and never move, so a captured graph may keep
their addresses across growth.
No behavior change for CudaSequenceKVCache: same allocation sizes, growth
policy, log lines and metrics. The pool carries every runtime safeguard the
sequence cache had, including those added in review of the base stack:
drains the device rather than touching the caller's (possibly destroyed)
stream;
check_compiled: every storage constant's compiled dtype and bytes mustmatch what the pool allocates, now covering side buffers too. AOTI reports
bytes rounded up to 64 when a program also holds CPU constants
(cpp_wrapper_cpu.py), so either form matches -- the base's exact
comparison would reject e.g. an 8-byte buffer reported as 64;
to them;
captures, including one about to capture for the first time.
check_write_lengthandadmitstay in CudaSequenceKVCache.Also registers test_cuda_kv_cache in CMake and runs it in the OSS
unittest-cuda-runtime job. Until now only the Buck target ran it.
Differential Revision: D123051477