Skip to content

[executorch][cuda] Extract CudaKVPool from CudaSequenceKVCache - #23365

Open
Gasoonjia wants to merge 1 commit into
gh/Gasoonjia/32/basefrom
gh/Gasoonjia/32/head
Open

Gasoonjia wants to merge 1 commit into
gh/Gasoonjia/32/basefrom
gh/Gasoonjia/32/head

Conversation

@Gasoonjia

@Gasoonjia Gasoonjia commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Moves the byte layer of the off-graph KV cache -- per-layer K/V allocation,
geometric growth, cross-stream ordering, AOTI constant discovery and binding,
and the CUDA-graph recapture a growth forces -- out of CudaSequenceKVCache into
CudaKVPool. CudaSequenceKVCache keeps what is specific to one sequence: the
neutral SequenceCache planner, admission, and the logical length. This mirrors
MLX, where MLXSequenceCache and MLXCellCache share one Pool.

The pool is what a batched layout needs next. Besides growable and fixed
layers it holds fixed-size side buffers a layout declares by FQN; they are
allocated once, start zeroed, and never move, so a captured graph may keep
their addresses across growth.

No behavior change for CudaSequenceKVCache: same allocation sizes, growth
policy, log lines and metrics. The pool carries every runtime safeguard the
sequence cache had, including those added in review of the base stack:

  • the device guard around stream and event calls, and a destructor that
    drains the device rather than touching the caller's (possibly destroyed)
    stream;
  • check_compiled: every storage constant's compiled dtype and bytes must
    match what the pool allocates, now covering side buffers too. AOTI reports
    bytes rounded up to 64 when a program also holds CPU constants
    (cpp_wrapper_cpu.py), so either form matches -- the base's exact
    comparison would reject e.g. an 8-byte buffer reported as 64;
  • a reloaded handle replaces its descriptors and binding instead of adding
    to them;
  • growth gives every graph-enabled handle at least one eager run before it
    captures, including one about to capture for the first time.
    check_write_length and admit stay in CudaSequenceKVCache.

Also registers test_cuda_kv_cache in CMake and runs it in the OSS
unittest-cuda-runtime job. Until now only the Buck target ran it.

Differential Revision: D123051477

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23365

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Unclassified Failure

As of commit 4bef881 with merge base c16dd0a (image):

NEW FAILURE - The following job has failed:

UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

This branch was successfully deployed

1 active deployment
cadence — 4bef8818 Deployed Oct 2, 2026 by Gasoonjia via hifi-op-test / hifi4 #31141
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant