Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
119 commits
Select commit Hold shift + click to select a range
4cc4a56
[fix] add ep ipc fix, deepep works without nvl_bytes=0
xenshinu Jun 15, 2026
14dda18
[sglang] add DeepEP no-fabric IPC policy
devin-ai-integration[bot] Jul 22, 2026
1158dff
[sglang] add no-fabric DeepEP IPC recipe
devin-ai-integration[bot] Jul 22, 2026
bdaad16
[sglang] fail clearly on old DeepEP API
devin-ai-integration[bot] Jul 22, 2026
6501d7f
[sglang] move inspect import to module top
devin-ai-integration[bot] Jul 22, 2026
ae9d308
[sglang] enforce DeepEP use_fabric requirement at Buffer construction
devin-ai-integration[bot] Jul 22, 2026
a02b083
[sglang] add experimental dense tensor-parallel recipe
devin-ai-integration[bot] Jul 23, 2026
71bb812
[fix] fix workspace mismatch when using fa2
xenshinu Jul 8, 2026
3f1c769
test: add reproducible SGLang TP validation
cursoragent Jul 23, 2026
0244e81
fix: keep TP NCCL transport symmetric
cursoragent Jul 23, 2026
8bcb00e
fix: pin TP validation to immutable revisions
cursoragent Jul 23, 2026
4c0247e
fix: align TP graph mode across phases
cursoragent Jul 23, 2026
21bf734
docs: mark SGLang tensor parallel validated
cursoragent Jul 23, 2026
7b57b02
test: strengthen tensor parallel verification
cursoragent Jul 23, 2026
c46e6bb
fix: harden graph and IPC restoration
cursoragent Jul 23, 2026
08c94a8
test: add reproducible vLLM TP validation
cursoragent Jul 23, 2026
7b8abf8
style: format vLLM TP harness
cursoragent Jul 23, 2026
2baf318
fix: pin compatible vLLM CUDA wheel
cursoragent Jul 23, 2026
cbeb4c7
fix: load torch before vLLM extension check
cursoragent Jul 23, 2026
8ff7050
fix: verify vLLM extension without CUDA driver
cursoragent Jul 23, 2026
944222d
test: cover zero-alignment VMM reservations
cursoragent Jul 23, 2026
ef56a76
fix: preserve VMM cursor for default alignment
cursoragent Jul 23, 2026
ea48b50
fix: keep vLLM TP graphs on NCCL collectives
cursoragent Jul 23, 2026
160a8c0
fix: pass vLLM TP compilation config atomically
cursoragent Jul 23, 2026
dd76136
test: use stable vLLM TP acceptance signals
cursoragent Jul 23, 2026
d9d757d
test: bound vLLM TP output comparison
cursoragent Jul 23, 2026
96083be
docs: mark vLLM tensor parallel validated
cursoragent Jul 23, 2026
fce0ab2
fix: harden VMM IPC lifecycle
cursoragent Jul 23, 2026
9842f2f
test: verify restored DeepEP outputs
cursoragent Jul 23, 2026
87c38bc
test: require complete vLLM TP archives
cursoragent Jul 23, 2026
13eb689
test: stabilize VMM IPC GPU regressions
cursoragent Jul 23, 2026
06f00aa
docs: describe TP offset invariants
cursoragent Jul 23, 2026
0f74290
fix: version preallocated IPC chunks
cursoragent Jul 23, 2026
f7bc0cd
test: require complete TP replay evidence
cursoragent Jul 23, 2026
77fd180
fix: pass default graph pool explicitly
cursoragent Jul 23, 2026
204c852
test: compare semantic TP graph archives
cursoragent Jul 23, 2026
e43fad2
docs: make TP validation invocation portable
cursoragent Jul 23, 2026
d830e89
test: unpack loaded graph results
cursoragent Jul 23, 2026
b7552d2
fix: pin TP validation seed
cursoragent Jul 23, 2026
f85e429
docs: explain deterministic TP validation
cursoragent Jul 23, 2026
0e77e4e
test: compare deterministic TP graph structure
cursoragent Jul 23, 2026
bc0f32f
docs: align TP scope with upstream roadmap
cursoragent Jul 23, 2026
30b4716
feat: port SGLang integration to main runner API
cursoragent Jul 23, 2026
07d5759
test: add SGLang main SAVE LOAD smoke
cursoragent Jul 23, 2026
3261fec
fix: fail closed on unsupported SGLang main modes
cursoragent Jul 23, 2026
d70817f
fix: propagate immutable refs into Modal smoke
cursoragent Jul 23, 2026
e097ec7
test: run TP validation on SGLang main
cursoragent Jul 23, 2026
69b5393
[sglang] port graph-capture hook to init_forward_metadata 3-method AB…
devin-ai-integration[bot] Jul 23, 2026
1750050
[sglang] Cursor-neutral pre-capture warmup for dense SAVE (Qwen3.5 hy…
devin-ai-integration[bot] Jul 23, 2026
2f92d75
docs: record Qwen3.5 TP limitations
cursoragent Jul 23, 2026
801696a
fix: disable multimem gather under Foundry
cursoragent Jul 23, 2026
96e2a92
fix: mirror SGLang main graph warmups on load
cursoragent Jul 23, 2026
c9a7bc2
test: parameterize TP validation models
cursoragent Jul 23, 2026
6a86f4f
[sglang] Route FlashInfer prepass through hybrid linear-attn wrapper
devin-ai-integration[bot] Jul 23, 2026
5eb9c5b
docs: align TP scope with upstream roadmap
cursoragent Jul 23, 2026
ccb6ec9
test: cover no-fabric DeepEP IPC
cursoragent Jul 23, 2026
e573ca6
test: exercise SGLang torch symmetric memory TP
cursoragent Jul 23, 2026
89b71a7
fix: isolate torch symmetric memory from Foundry VMM
cursoragent Jul 23, 2026
3a6fa04
test: cover symmetric-memory TP recipe
cursoragent Jul 23, 2026
45ce719
docs: describe symmetric-memory TP experiment
cursoragent Jul 23, 2026
7a433d0
fix: reject unrestorable symmetric-memory TP
cursoragent Jul 23, 2026
16222b0
docs: record symmetric-memory replay failure
cursoragent Jul 23, 2026
da36e85
test: verify outputs in every TP phase
cursoragent Jul 23, 2026
eba20e1
[sglang] In-region warmup on both SAVE and LOAD for hybrid model grap…
devin-ai-integration[bot] Jul 23, 2026
94bb1ae
test: instrument symmetric-memory replay
cursoragent Jul 23, 2026
888219d
test: isolate NCCL graph mixing in TP replay
cursoragent Jul 23, 2026
73ab787
test: probe live symmetric-memory collectives
cursoragent Jul 23, 2026
fd89bf5
test: isolate symmetric-memory eager probe
cursoragent Jul 23, 2026
2405be3
test: compare full symmetric graph loading
cursoragent Jul 23, 2026
d7804c8
test: preserve full symmetric graph JSON
cursoragent Jul 23, 2026
d2794d7
test: reproduce symmetric-memory graph replay
cursoragent Jul 23, 2026
02cba84
test: inspect symmetric-memory signal state
cursoragent Jul 23, 2026
8510fe9
test: reset symmetric signal pads before replay
cursoragent Jul 23, 2026
2e33730
test: trace SGLang graph tensor addresses
cursoragent Jul 23, 2026
530d1d1
fix: disable split-K for SGLang graph replay
cursoragent Jul 23, 2026
0e2bc52
fix: avoid cuBLASLt in SGLang replay graphs
cursoragent Jul 23, 2026
68f23db
test: inspect live symmetric pointer tables
cursoragent Jul 23, 2026
34343cd
test: relocate symmetric graph buffer in replay probe
cursoragent Jul 23, 2026
186735d
test: replay repeated symmetric collectives
cursoragent Jul 23, 2026
2fa04ba
test: trace restored symmetric layer outputs
cursoragent Jul 23, 2026
922531e
test: relocate SGLang symmetric graph pointers
cursoragent Jul 23, 2026
27008be
test: isolate symmetric pointer relocation
cursoragent Jul 23, 2026
f8b1b5a
fix: relocate SGLang symmetric-memory graphs
cursoragent Jul 23, 2026
018d9fd
fix: validate SGLang graph relocation ABI
cursoragent Jul 23, 2026
9c98ad5
fix: restore SGLang graph warmups
cursoragent Jul 23, 2026
97c5d05
fix: load SGLang symmetric graphs sequentially
cursoragent Jul 23, 2026
ca4f637
fix: fence SGLang graph warmups
cursoragent Jul 23, 2026
81c8841
fix: constrain symmetric replay to one graph
cursoragent Jul 23, 2026
2b9f919
test: harden symmetric graph validation
cursoragent Jul 23, 2026
8383c4d
test: verify graph state after eager fallback
cursoragent Jul 23, 2026
302cea6
fix: validate all symmetric graph operands
cursoragent Jul 23, 2026
71cfeae
docs: record verified SGLang symmetric replay
cursoragent Jul 23, 2026
d43f314
docs: design multi-shape SGLang replay state
cursoragent Jul 23, 2026
5223c37
docs: plan multi-shape SGLang replay
cursoragent Jul 24, 2026
fe8ec46
feat: add SGLang shape replay state
cursoragent Jul 24, 2026
cc86714
fix: load shape replay state tests without foundry package init
cursoragent Jul 24, 2026
fd5544e
feat: persist SGLang shape state manifest
cursoragent Jul 24, 2026
6a909a0
feat: adapt SGLang FlashInfer shape state
cursoragent Jul 24, 2026
e017915
fix: harden SGLang shape activation ownership
cursoragent Jul 24, 2026
2575119
fix: harden SGLang shape activation cleanup
cursoragent Jul 24, 2026
032a6c8
feat: add shape-aware SGLang graph relocation
cursoragent Jul 24, 2026
811eda0
fix: harden FlashInfer graph relocation validation
cursoragent Jul 24, 2026
97447b9
fix: make FlashInfer relocation atomic
cursoragent Jul 24, 2026
f6646b4
feat: materialize SGLang shape replay capsules
cursoragent Jul 24, 2026
142912b
fix: harden SGLang shape capsule lifecycle
cursoragent Jul 24, 2026
af51e35
feat: gate SGLang multi-shape replay
cursoragent Jul 24, 2026
4fb8e7d
fix: require explicit batch schedule in multi-shape hook
cursoragent Jul 24, 2026
ead8a98
test: verify every SGLang replay shape
cursoragent Jul 24, 2026
b428d0e
test: harden SGLang replay shape validation harness
cursoragent Jul 24, 2026
d42a7e0
fix: adopt padded replay metadata for SGLang
cursoragent Jul 24, 2026
4e07a7e
test: parse SGLang replay from decode graph key size
cursoragent Jul 24, 2026
da1e51d
test: require eager decode evidence for max+1 fallback
cursoragent Jul 24, 2026
742970b
test: force deterministic eager fallback decode sampling
cursoragent Jul 24, 2026
90dc214
docs: normalize task report list formatting
cursoragent Jul 24, 2026
d23d54b
docs: record SGLang multi-shape milestones
cursoragent Jul 24, 2026
227d570
chore: remove tracked SDD scratch reports
cursoragent Jul 24, 2026
87c224d
test: batch graph-shape exercises in one generate request
cursoragent Jul 24, 2026
6d53218
test: exercise graph shapes largest-to-smallest
cursoragent Jul 24, 2026
0db3666
test: reject duplicate graph exercise batch keys
cursoragent Jul 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -362,4 +362,7 @@ marimo/_lsp/
__marimo__/

# Streamlit
.streamlit/secrets.toml
.streamlit/secrets.toml
tests/deepep_matrix_work/
deepep_fabric_archive/
hook_archive/
30 changes: 29 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -85,7 +85,35 @@ Foundry ships engine integrations under `foundry/python/foundry/integration/`. P
| SGLang | ✅ | ✅ | 🚧 | ✅ |
| TensorRT-LLM | 🚧 | 🚧 | 🚧 | 🚧 |

✅ validated end-to-end (SAVE → LOAD → query)  ·  🚧 not yet
✅ officially validated end-to-end  ·  🚧 not official/general

An experimental SGLang dense-TP recipe is validated with Qwen3-8B at TP=2 on
2×H100. Both ranks restore 36 graphs at their saved allocation offsets, and
deterministic LOAD responses match a non-Foundry baseline using the same pinned
settings. This is not general TP support: upstream
[Discussion #5](https://github.com/orgs/foundry-org/discussions/5) reports a
successful internal torch symmetric-memory prototype but no official
implementation; support remains tracked in
[issue #6](https://github.com/foundry-org/foundry/issues/6).

SGLang-main also has an opt-in torch symmetric-memory TP=2 path. Foundry records
the process-local communication buffer and pointer-table operands, restores the
per-shape warmups, and relocates those operands to the fresh LOAD communicator.
The checked-in 2×H100 proof validates one self-contained batch-1 CUDA graph per
rank at `final_alloc_offset=68555898880`: baseline, two SAVE passes, and fresh
LOAD return byte-identical deterministic responses with no recapture or runtime
errors. Larger batches fall back to eager execution. Multi-shape symmetric
capture remains rejected because later graphs depend on mutable FlashInfer state
owned by earlier shapes.

An experimental vLLM dense-TP recipe is validated with Qwen3-8B at TP=2 on
2×H100. Both ranks restore 64 graphs at the same per-rank offset produced by
repeated deterministic SAVE passes, without recapture. Sequential and
concurrent LOAD responses match a non-Foundry baseline using identical pinned
NCCL settings. As with SGLang, this is a pinned result rather than general TP
support; upstream reports an unpublished torch symmetric-memory prototype
([Discussion #5](https://github.com/orgs/foundry-org/discussions/5),
[issue #6](https://github.com/foundry-org/foundry/issues/6)).

The adapted vLLM / SGLang / TensorRT-LLM forks will be released alongside this repo at `foundry-org/vllm`, `foundry-org/sglang`, `foundry-org/TensorRT-LLM`.

Expand Down
73 changes: 57 additions & 16 deletions RELEASE.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,55 @@
# Foundry 0.0.2
# Release Notes

## Unreleased

- **Experimental vLLM dense tensor parallel.** Qwen3-8B TP=2 is validated
end-to-end on 2×H100 with NCCL P2P/IPC. Deterministic SAVE passes produce
identical per-rank offsets and graph inventories; fresh LOAD restores 64
graphs per rank without capture and returns byte-identical sequential and
concurrent factual completions. General TP remains open in
[upstream issue #6](https://github.com/foundry-org/foundry/issues/6);
[Discussion #5](https://github.com/orgs/foundry-org/discussions/5) reports an
internal torch symmetric-memory prototype but no official implementation.
- **Experimental SGLang dense tensor parallel.** Qwen3-8B TP=2 is validated
end-to-end on 2×H100 with NCCL P2P/IPC. Fresh LOAD restores 36 graphs per
rank at the recorded allocation offset and returns byte-identical
temperature-0 outputs versus a non-Foundry baseline using identical pinned
settings. General TP remains open in
[upstream issue #6](https://github.com/foundry-org/foundry/issues/6);
[Discussion #5](https://github.com/orgs/foundry-org/discussions/5) reports an
unpublished torch symmetric-memory prototype. SGLang-main now has a validated
TP=2 symmetric-memory option for one batch-1 CUDA graph: SAVE records the
external communication operands and full graph JSON, while LOAD rebuilds
required warmup state and relocates those operands to the live communicator.
A 2×H100 baseline/SAVE/SAVE/LOAD matrix returned byte-identical outputs at
`final_alloc_offset=68555898880`. Larger batches use eager execution;
multi-shape symmetric graph capture remains fail-closed.
- **No-fabric DeepEP IPC.** The VMM-IPC bridge transports shareable file
descriptors with `SCM_RIGHTS`, allowing vLLM and SGLang DeepEP NVL buffers to
work without fabric/IMEX. This extends
[upstream PR #3](https://github.com/foundry-org/foundry/pull/3). The verified
regression is BF16 on 2×H100; GB300/aarch64 FP8 and RDC relinking reported in
[upstream issue #1](https://github.com/foundry-org/foundry/issues/1) remain
outside that validation scope.
- **FA2 workspace-parity fix.** CUDA generator registration is skipped on LOAD
when every graph has zero RNG increment, preventing unused seed/offset tensors
from drifting the tracked workspace cursor. This incorporates
[upstream PR #4](https://github.com/foundry-org/foundry/pull/4) and the root
cause from [issue #2](https://github.com/foundry-org/foundry/issues/2);
this branch did not independently rerun the A100 hardware case.
- **CUDA default-alignment VMM reservations.** `cuMemAddressReserve` accepts
`alignment=0` to request the driver default. Foundry now preserves the current
VMM cursor for that value instead of bit-masking it to address zero.

---

## Foundry 0.0.2

SGLang graduates to a fully validated engine. This release brings the SGLang
integration to parity with vLLM across single GPU, data parallel, and expert
parallel — with a self-contained recipe and no vLLM build dependency for EP.

## Highlights
### Highlights

- **SGLang single GPU / DP / EP all validated end-to-end.**
SAVE → LOAD → query verified on single-GPU Qwen3-1.7B / 4B / 14B,
Expand All @@ -23,7 +68,7 @@ parallel — with a self-contained recipe and no vLLM build dependency for EP.
`data_parallel_controller.py`). All save/load logic lives in the integration
layer; the edits are inert unless `--foundry-graph-extension-config-path` is set.

## Engine integrations
### Engine integrations

- **SGLang** — integration for SGLang v0.5.13. Working configurations: single
GPU, data parallel (DP), expert parallel (EP, DeepEP low-latency + DP-attention with fa3). Self-contained recipes under `recipe/sglang/` (shared TOML pair +
Expand All @@ -45,7 +90,7 @@ parallel — with a self-contained recipe and no vLLM build dependency for EP.
on LOAD, with `init_nvshmem_for_loaded_modules` run once on LOAD before any
NVSHMEM-kernel graph replays.

## Fixes
### Fixes

- **Per-rank VMM device binding (DP/TP/EP).** `set_allocation_region` binds to
the current CUDA device, so the integration now calls `set_device(gpu_id)`
Expand All @@ -63,26 +108,22 @@ parallel — with a self-contained recipe and no vLLM build dependency for EP.
capture loop never sets), and the FlashInfer per-bs pre-pass gated off for fa3
while still populating `decode_cuda_graph_metadata` post-load for replay.

## Docs
### Docs

- New `docs/sglang/` set (overview, direct-edits, hooks, memory-lifecycle,
save-load-workflow, memory-consistency) and a self-contained `recipe/sglang/`
README with install, run, performance, and troubleshooting.
- Top-level README parallelism status table updated to mark SGLang single GPU,
DP, and EP as validated.

---

## Previous Releases

## Foundry 0.0.1

First public release of Foundry — a CUDA-graph persistence library that
captures an entire model's CUDA graphs (plus their device context: modules,
workspaces, VMM layout) once and replays them at startup, eliminating compile,
warmup, and capture from cold-start time.

## Highlights
### Highlights

- **Deterministic memory layout — zero patching on graph load.**
Foundry indirects memory allocation to the same reserved memory region
Expand All @@ -99,7 +140,7 @@ warmup, and capture from cold-start time.
On load, the template is rebuilt once per group and instances are reconstructed
on demand, keeping load fast and asynchronous.

## Engine integrations
### Engine integrations

- **vLLM** — compatible with vLLM v0.21. Working configurations: single GPU,
data parallel (DP), expert parallel (EP, DeepEP low-latency). End-to-end
Expand All @@ -111,31 +152,31 @@ warmup, and capture from cold-start time.
integration layer (`foundry/integration/<engine>/`); engine forks contain
only minimal hook calls.

## Verified kernel & comm support
### Verified kernel & comm support

- **cuBLAS NVJET** kernels (Hopper+).
- **torch.compile** modules.
- **NVSHMEM / DeepEP** validated.
- **DeepGEMM FP8 MoE** validated.

## Dependency
### Dependency

- **PyTorch 2.11.0** (compatible with 2.9 – 2.11).
- **CUDA 12+**, CMake 4.0+, Boost.

## Documentation
### Documentation

- Integration design notes and per-engine recipes under `docs/` and `recipe/`.
- vLLM recipe README covers save/load workflow, archive layout, and required
env settings.

## Repository hygiene
### Repository hygiene

- Open-source pre-commit hooks (ruff, ruff-format, clang-format,
markdownlint, actionlint, DCO sign-off).
- Smoke tests covering re-export imports and archive round-trip.

## Roadmap
### Roadmap

- Adapted vLLM and SGLang forks published alongside the release.
- Tensor parallel support.
Expand Down
12 changes: 9 additions & 3 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,12 +26,18 @@

- [x] Sync with latest vLLM release
- [x] EP on vLLM with DeepEP
- [ ] TP on vLLM
- [ ] Official TP on vLLM
- [x] Experimental NCCL TP=2 recipe on 2×H100
- [ ] Evaluate/publish the torch symmetric-memory prototype ([upstream discussion](https://github.com/orgs/foundry-org/discussions/5))
- [x] Quantized MoE with DeepGemm
- [x] Drop-in integration layer for CUDA graph persistence in vLLM
- [x] Sync with latest SGLang release
- [ ] EP on SGLang
- [ ] TP on SGLang
- [x] EP on SGLang
- [ ] Official TP on SGLang
- [x] Experimental NCCL TP=2 recipe on 2×H100
- [x] Reproduce torch symmetric-memory fresh-LOAD corruption on 2×H100
- [x] Relocate process-local symmetric-memory operands for a TP=2 batch-1 graph
- [ ] Generalize symmetric-memory replay to multiple graph shapes ([upstream discussion](https://github.com/orgs/foundry-org/discussions/5))

## Stage 5: Disaggregated and Large-Scale Serving

Expand Down
27 changes: 21 additions & 6 deletions csrc/CUDAGraph.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -1528,15 +1528,30 @@ GraphLoadResult CUDAGraph::load(const std::string& json_path, MempoolId_t pool)
}

const json::array& generators_array = root.at("generators").as_array();
bool has_rng_generator = false;
for (const auto& gen_val : generators_array) {
const json::object& gen_obj = gen_val.as_object();
uint64_t state_id = gen_obj.at("id").to_number<uint64_t>();
uint64_t seed = gen_obj.at("seed").to_number<uint64_t>();
uint64_t wholegraph_increment = gen_obj.at("wholegraph_increment").to_number<uint64_t>();
auto increment_it = gen_obj.find("wholegraph_increment");
// Missing metadata is treated conservatively: registering may expose an
// allocation mismatch, but skipping could silently leave RNG kernel
// arguments pointing at the SAVE process's seed/offset tensors.
if (increment_it == gen_obj.end() || increment_it->value().to_number<uint64_t>() != 0) {
has_rng_generator = true;
break;
}
}

auto state = global_generator_state_registry.get_state_from_id(state_id, seed);
state->register_graph(reinterpret_cast<at::cuda::CUDAGraph*>(graph.get()));
graph->captured_generator_states_[state] = wholegraph_increment;
if (has_rng_generator) {
for (const auto& gen_val : generators_array) {
const json::object& gen_obj = gen_val.as_object();
uint64_t state_id = gen_obj.at("id").to_number<uint64_t>();
uint64_t seed = gen_obj.at("seed").to_number<uint64_t>();
uint64_t wholegraph_increment = gen_obj.at("wholegraph_increment").to_number<uint64_t>();

auto state = global_generator_state_registry.get_state_from_id(state_id, seed);
state->register_graph(reinterpret_cast<at::cuda::CUDAGraph*>(graph.get()));
graph->captured_generator_states_[state] = wholegraph_increment;
}
}

const json::object& allocator_events = root.at("allocator_events").as_object();
Expand Down
21 changes: 18 additions & 3 deletions csrc/CUDAGraphParallel.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -1658,6 +1658,7 @@ std::shared_ptr<PendingGraphLoads> start_graph_builds_impl(
go["seed"] = gens[g].seed;
go["wholegraph_increment"] = gens[g].wholegraph_increment;
gens_array.push_back(go);
pending->has_rng_generators |= gens[g].wholegraph_increment != 0;
}
pending->entries[i].generators_meta = boost::json::value(std::move(gens_array));
}
Expand Down Expand Up @@ -1690,6 +1691,15 @@ std::shared_ptr<PendingGraphLoads> start_graph_builds_impl(

auto gen_it = pre_root.find("generators");
if (gen_it != pre_root.end()) {
const boost::json::array& generators = gen_it->value().as_array();
for (const auto& gen_val : generators) {
const boost::json::object& generator = gen_val.as_object();
auto increment_it = generator.find("wholegraph_increment");
if (increment_it == generator.end() || increment_it->value().to_number<uint64_t>() != 0) {
pending->has_rng_generators = true;
break;
}
}
pending->entries[i].generators_meta = std::move(gen_it->value());
pre_root.erase(gen_it);
}
Expand Down Expand Up @@ -2045,9 +2055,14 @@ GraphLoadResult finish_one_graph_load_impl(std::shared_ptr<PendingGraphLoads> pe

auto& entry = pending->entries[index];

// Register deferred generators (before allocator replay to match
// SAVE mode timing where generators are created before graph capture).
if (pending->registry && !entry.generators_meta.is_null()) {
// register_graph lazily creates seed/offset tensors before allocator replay.
// They must stay in the deterministic region when a captured kernel consumes
// Philox state, because their pointers are embedded in that kernel's args.
// When the entire pending set has zero RNG increment, no captured node
// references those tensors; skip registration entirely. This avoids the
// caching-allocator segment growth that caused a 2 MiB A100/FA2 cursor drift
// without replacing deterministic pointers with ordinary CUDA allocations.
if (pending->has_rng_generators && pending->registry && !entry.generators_meta.is_null()) {
const boost::json::array& gen_array = entry.generators_meta.as_array();
for (const auto& gen_val : gen_array) {
const boost::json::object& gen_obj = gen_val.as_object();
Expand Down
Loading