Skip to content

[executorch][cuda] Add CudaExecutor - #23370

Open
Gasoonjia wants to merge 1 commit into
gh/Gasoonjia/37/basefrom
gh/Gasoonjia/37/head
Open

Gasoonjia wants to merge 1 commit into
gh/Gasoonjia/37/basefrom
gh/Gasoonjia/37/head

Conversation

@Gasoonjia

@Gasoonjia Gasoonjia commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

The CUDA counterpart of ModuleExecutor, in backends/cuda/batching: implements
the neutral batching::Executor over a cell-layout off-graph KV cache, so the
neutral Runner and DecodeFirstScheduler drive a CUDA program unchanged.

The CUDA backend compiles static graphs, so where ModuleExecutor runs one
forward, CudaExecutor runs two methods and routes each slice of a batch:

  • decode: one token, static. A CUDA graph is captured for it by default.
  • prefill: dynamic from 2 tokens to its exported width W, eager.
    plan_slices (step_plan.h, header-only) cuts a batch into forwards of at most
    W tokens, in order; a lone tail token runs decode. Each slice declares its
    tokens to the cache, runs, and samples its selected rows.

create():

  • reads the neutral metadata (context, vocab, geometry, logits mode, which
    must be selected since decode's selector is static at one row);
  • validates both methods' inputs and outputs;
  • reads get_offgraph_kv_max_cells and builds the batched-cell cache with
    the program's max_cells and W, which fix the step buffers' shapes;
  • requires max_sessions * max_session_tokens to fit the pool;
  • sets the process-wide CUDA backend options from CudaExecutorOptions:
    weight_sharing_across_methods (default on) and
    enable_cuda_graph_for_method=decode (default on). Both are overridable.

initialize() loads both methods with the cache's registry key. Session,
clone, sampling and build_step are copied from ModuleExecutor; a follow-up
diff moves the shared parts into extension/llm/batching/util.

Builds with Buck and CMake (cuda_batching, added from the root after
extension/llm/batching, installed to ExecuTorchTargets for downstream
runners).

Differential Revision: D123051481

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23370

Note: Links to docs will display an error until the docs builds have been completed.

❌ 6 New Failures, 3 Unclassified Failures

As of commit 2548fb5 with merge base c16dd0a (image):

NEW FAILURES - The following jobs have failed:

UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

This branch was successfully deployed

1 active deployment
cadence — 2548fb59 Deployed Oct 2, 2026 by Gasoonjia via hifi-op-test / hifi4 #31147
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant