Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
vision_sft_nanotrains only the Generator-sidemoe_gen,time_embedder,vae2llm, andllm2vaeparameters. The Reasoner remains frozen, but theoriginal joint path still constructs, FSDP-shards, all-gathers, and executes its
8.19B parameters, while also duplicating them in the FP32 EMA model.
This PR moves the frozen Reasoner behind a provider-neutral per-layer K/V
boundary so Generator training ranks can avoid materializing it.
What changed
Shared conditioning interface
joint,inline,offline, andremoteReasoner-conditioning backends.jointas the default, so existing recipes remain unchanged.K/V remain live and trainable.
wrapping for external backends, including the EMA model.
inlineis the local two-pass numerical reference.read_throughis reservedin the configuration contract but deliberately remains unimplemented.
Offline backend
VAE, or EMA model.
and distributed completion checks.
Remote backend
token-based admission, queue limits, and backpressure.
corresponding FSDP collective.
Install the optional dependencies with:
Validation
Automated validation
git diff --check, pre-commit, anduv lock --checkpassed.The tests cover cache integrity, tensor codecs, protocol corruption, retry and
deadline handling, cancellation, admission, OOM propagation, distributed
failure synchronization, checkpoint loading, and resource cleanup.
Offline 8xH20 A/B
Across six paired steady-state steps, offline conditioning reduced mean step
time by 17.05%, increased aggregate token throughput by 20.56%, and reduced
peak allocated memory by 15.63 GiB per GPU. This is a short hot-cache benchmark,
not a full-corpus storage benchmark.
Remote correctness and training
A real 36-layer Nano Reasoner on H20 produced bitwise-identical K/V tensors
through direct extraction and localhost gRPC for every layer.
A separate run reserved GPU 0 for the Reasoner service and used GPUs 1--7 for
seven-rank Generator FSDP. The first nonzero-LR update changed all 405 live
Generator tensors. A same-job restart restored model, optimizer, scheduler, and
trainer state at iteration 2, continued at the restored LR, and changed all 405
tensors again at iteration 3.
Peak physical memory was approximately 18.6 GiB on the Reasoner GPU and
30.0--33.0 GiB per Generator GPU. No CUDA OOM, RPC, NCCL, or training error
occurred. Checkpoints contain 405 live plus 405 EMA Generator leaves and no
Reasoner parameters.
Limitations and remaining gates
and networking remains unmeasured.
configured joint/inline/offline comparison.
parity remain pending.
attention, currently fail closed.
read_through, dynamic batching, layerwise H2D staging, automatic artifactdigest derivation, and persistent service caching are not implemented.
isolated trusted network only.
Raw benchmark logs and 104 GiB checkpoints are local artifacts under
outputs/and are not included in this PR.
Rollout guidance
jointremains the default.offlinefor a fixed and enumerable SFT corpus.remotewhen prompts or augmentations make exhaustive caching impractical.ranks; a separate shared Reasoner pool is preferable for production.