Skip to content

[executorch][muse] Move off-graph KV to the batching path and test batching in OSS CI - #23375

Open
Gasoonjia wants to merge 1 commit into
gh/Gasoonjia/42/basefrom
gh/Gasoonjia/42/head
Open

Gasoonjia wants to merge 1 commit into
gh/Gasoonjia/42/basefrom
gh/Gasoonjia/42/head

Conversation

@Gasoonjia

@Gasoonjia Gasoonjia commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

The off-graph KV cache now reaches Muse Glimmer through the batching path
only (export_solo_batching.py + run_solo_batching). This removes the earlier
single-sequence wiring from everywhere else in MG and switches the OSS CI mode
to the batching path, which now also tests batching itself.

Removed, so these match their state before the off-graph stack:

  • runtime/engine/muse_glimmer_engine.{h,cpp}, runtime/runners/solo.cpp and
    muse_glimmer_worker.cpp: the off-graph plan, cache building and install,
    reset handling, offgraph_initial_capacity, and the engine's part of CUDA
    graph capture with the off-graph cache;
  • export/export_solo.py: --use-offgraph-kv-cache and its branches. One
    piece stays: _solo_constant_methods omits get_mutable_buffer_metadata
    when there is none, because export_solo_batching.py shares it for a model
    whose cache lives off-graph.

Kept as they are, because the batching path uses them: the off-graph source
transformation in source_transformations/cuda.py (with its tests in
test_cuda_pipeline.py), export_solo_batching.py and run_solo_batching.

The CUDA backend side of the stack (lowering, CudaKVPool, CudaSequenceKVCache,
CudaCellCache, delegate-driven steps, graph recapture) is unchanged; the
batching path uses it.

CI: solo-text-offgraph becomes solo-text-batching. It exports with
export_solo_batching.py and runs run_solo_batching three times:

  • one prompt with eager decode;
  • the same prompt with the captured decode graph (these two cover what
    solo-text-offgraph checked, "Paris");
  • five prompts in one batch: France, Japan, Italy, France again, and a prompt
    longer than a forward (Germany), which prefills in slices beside the others'
    decodes.

Each run writes a report, and .ci/scripts/check_muse_glimmer_batching_report.py
checks three things:

  1. every generation finishes normally and contains its answer (Paris, Tokyo,
    Rome, Paris, Berlin), and eager and graph decode generate the same tokens.
    Tokens are not compared across batches: the France prompt prefilled alone
    and the one prefilled beside others run GEMMs of different shapes, and on
    the real model bf16 rounding steers greedy decoding differently (both
    still answer Paris);
  2. the weights load once: loading costs at most 1.25x the weights, and does
    not vary by more than 256 MiB across runs reserving 4 or 5 sessions;
  3. memory behaves like an off-graph KV cache: the pool holds the cells in use,
    grows geometrically (at most twice what was needed) instead of reserving
    every session's context, its bytes equal rows x bytes per cell, and
    generation adds no more than the pool plus a 3 GiB activation slack.

The export's mode check now accepts this mode (it listed only solo-text and
dflash-image, so the old mode could not have passed it).

Differential Revision: D123051488

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23375

Note: Links to docs will display an error until the docs builds have been completed.

❌ 10 New Failures, 3 Unclassified Failures

As of commit 59464d8 with merge base c16dd0a (image):

NEW FAILURES - The following jobs have failed:

UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 2, 2026
@Gasoonjia
Gasoonjia deployed to upload-benchmark-results October 2, 2026 21:54 — with GitHub Actions Active

This branch was successfully deployed

2 active deployments
upload-benchmark-results — 59464d83 Deployed Oct 2, 2026 by Gasoonjia via upload-benchmark-results #3075
cadence — 59464d83 Deployed Oct 2, 2026 by Gasoonjia via hifi-op-test / hifi4 #31154
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant