Skip to content

fix(inference): load 64 GB Spark weights lazily - #12633

Merged
ericksoa merged 1 commit into
mainfrom
codex/spark64-lazy-safetensors
Oct 5, 2026
Merged

ericksoa merged 1 commit into
mainfrom
codex/spark64-lazy-safetensors

Conversation

@jyaunches

@jyaunches jyaunches commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Outcome

The 64 GB Spark recipe now completes local model startup using lazy safetensors loading. A physical 64 GB Spark passed automatic onboarding, local chat, and read, write, and exec tool calls, including after a restart.

Reason

The original fastsafetensors recipe stalled during startup on this host, with repeated NV_ERR_NO_MEMORY kernel errors. Available host memory fell to about 19 MB, and the operator stopped the container before it became ready. Docker reported OOMKilled: false.

The loader change resolved this failure in the tested configuration. The precise source of the previous allocation pressure remains unproven.

Related issues

Refs #12502

Changes

  • Use --load-format safetensors --safetensors-load-strategy lazy in the existing 64 GB Spark recipe.
  • Assert those arguments for automatic, picker, and resume installation paths.
  • Retain the model and image pins, 32K context, concurrency, memory utilization, Marlin configuration, and CUDA graphs. The larger Spark recipe is unchanged.

Verification

  • Regression test: the three new assertions failed before the recipe change.
  • npx vitest run --project cli src/lib/inference/vllm-fixed-catalog-install.test.ts src/lib/inference/vllm-runtime-selection.test.ts — 28 tests passed.
  • NODE_OPTIONS=--max-old-space-size=8192 npm run typecheck:cli — passed.
  • npm run catalog:compile, npm run catalog:check, npx oxfmt --check src/lib/inference/vllm-fixed-catalog-install.test.ts, and git diff --check — passed.
  • Physical GB10 Spark: 63,084,990,464 bytes usable RAM, Ubuntu 24.04.4 aarch64, driver 580.159.03. Source-built main 1812d698aa4c5b12db4475b878bb6d261ac79a6b plus this recipe change; no memory or profile override.
  • Automatic onboarding selected nvidia/Qwen3.6-35B-A3B-NVFP4 and the 64 GB profile. Onboarding exited zero and the local inference route was healthy.
  • Host CLI chat and actual read/write/exec calls exited zero. Independent file hashes matched. Successful tool receipts with replayInvalid: true remained successful.
  • A direct local request with 30,037 prompt tokens returned the expected response in 5.49 seconds.
  • Restarted both model and sandbox, then repeated chat, tools, file verification, and health checks successfully. Model readiness returned in 205 seconds.
  • Minimum available host RAM across startup, tests, and restart: 21.3 GB. Peak swap usage: 1.048 GB, versus 1.037 GB initially. No new NVIDIA memory errors or OOM kernel messages.
  • Normal pre-push publication validation and plugin, JavaScript, and CLI compiler checks — passed. GitHub verified the signed commit.
  • Inspected the complete diff: no secrets, API keys, or credentials.

Review notes

Both changed paths are sensitive under the repository policy. Self-review covered commit fc3b75baeae8a37d0a7236a4e687cdff6819432d and its complete two-file change in NVIDIA/NemoClaw against 1812d698aa4c5b12db4475b878bb6d261ac79a6b, including the generated launch command and the physical hardware evidence above. No independent pre-publication review exists; both paths await review.

Hardware evidence covers one physical host and one restart. It does not establish prolonged or concurrent-load stability. The hosted curl installer was not tested. Existing swap was retained; no system packages, drivers, or kernel settings were changed for this test.


Signed-off-by: Julie Yaunches jyaunches@nvidia.com

Summary by CodeRabbit

  • Bug Fixes
    • Updated the vLLM recipe to use safetensors loading with a lazy loading strategy, replacing the previous loading format.

The fastsafetensors recipe exhausted host memory during startup on a physical
64 GB Spark. Use safetensors with lazy loading for this profile while retaining
its model, context, and runtime settings.

Three automatic, picker, and resume assertions failed before the recipe change.
The focused suite now passes 28 tests. Physical hardware completed onboarding,
local chat, read/write/exec tools, a 30,037-token prompt, and a model/sandbox
restart with at least 21.3 GB host memory available.

Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

🧰 Additional context used
📚 Code guidelines (1)
test/README.md — auto-discovered

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/NemoClaw/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 4febc0ca-e43e-4c24-afdc-3e065f768676
📥 Commits

Reviewing files that changed from the base of the PR and between 1812d69 and fc3b75b.

📒 Files selected for processing (2)
  • managed-inference/recipes/vllm.qwen3-6-35b-a3b-nvfp4.spark-single-64gb.v1.yaml
  • src/lib/inference/vllm-fixed-catalog-install.test.ts

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

The bounded Spark vLLM recipe now uses safetensors with the lazy loading strategy. Its test checks for these options and verifies that the command omits fastsafetensors.

Changes

Spark recipe loading

Layer / File(s) Summary
Update Spark loading options
managed-inference/recipes/vllm.qwen3-6-35b-a3b-nvfp4.spark-single-64gb.v1.yaml, src/lib/inference/vllm-fixed-catalog-install.test.ts
The recipe uses --load-format safetensors and --safetensors-load-strategy lazy. The test expects these options and checks that the command omits --load-format fastsafetensors.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~5 minutes

Change: Bug fix

Suggested reviewers: ericksoa

Merge Risk: ⚪ Minimal · up to fc3b7

The 64 GB Spark recipe switches to lazy safetensors loading, and its pinned runtime supports both options. No actionable merge risk remains in the reviewed changes.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: loading 64 GB Spark weights lazily.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 1…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@jyaunches
jyaunches marked this pull request as ready for review October 5, 2026 21:50
@jyaunches
jyaunches requested a review from ericksoa October 5, 2026 21:51
@github-code-quality

Copy link
Copy Markdown
Contributor

Code Coverage Overview

Languages: TypeScript

TypeScript / code-coverage/plugin

The overall line coverage in commit fc3b75b in the codex/spark64-lazy-s... branch is 97%. The line coverage in commit 63002cd in the main branch is 96%.

Show a line coverage summary of the most impacted files.
File main 63002cd codex/spark64-lazy-s... fc3b75b +/-
nemoclaw/src/onboard/config.ts 98% 96% -2%
nemoclaw/src/index.ts 94% 93% -1%
nemoclaw/src/bl...t-management.ts 100% 100% 0%
nemoclaw/src/co.../config-show.ts 100% 100% 0%
nemoclaw/src/commands/slash.ts 100% 100% 0%
nemoclaw/src/on...native-route.ts 0% 100% +100%

@ericksoa ericksoa left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit fc3b75b against 1812d69. No blocking findings in either changed path.

The 64 GB Spark recipe selects safetensors with lazy loading through the existing catalog and command materialization. The added assertions exercise automatic, picker, and resume installs through real selection logic. The larger Spark recipe, model and image pins, serving limits, and security settings remain unchanged.

Independent local validation: all 28 tests in vllm-fixed-catalog-install.test.ts and vllm-runtime-selection.test.ts passed. Catalog compilation, catalog validation, Oxfmt, and git diff --check passed. The review checkout has no tracked changes. GitHub reports the commit signature as valid.

The physical Spark results are author-provided evidence; I did not rerun hardware validation. Broader CI and CodeRabbit are still running. This approval records the independent code review and focused validation; it does not claim those pending checks have completed.

@ericksoa
ericksoa merged commit 04e99d3 into main Oct 5, 2026
83 of 96 checks passed
@ericksoa
ericksoa deleted the codex/spark64-lazy-safetensors branch October 5, 2026 21:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants