fix(inference): disable MTP for Spark Qwen profile - #8248
Conversation
Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
📝 WalkthroughWalkthroughThe DGX Spark Qwen profile no longer enables MTP speculative decoding by default. Tests validate the updated command. Documentation describes the serving arguments, operator overrides, resource limits, and host-freeze limitation. Architecture budgets are reduced. ChangesQwen serving defaults
Architecture budget updates
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Security review verdictPASS — low risk. The exact remote diff at Nine-category analysis
Files reviewed
Reviewer: Codex Desktop, using the repository's |
|
🌿 Preview your docs: https://nvidia-preview-pr-8248.docs.buildwithfern.com/nemoclaw |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage in commit 04024f3 in the TypeScript / code-coverage/cliThe overall coverage in commit 04024f3 in the Show a code coverage summary of the most impacted files.
Updated |
PR Review Advisor — No blocking findings reportedAdvisor assessment: No blocking advisor findings reported Model lanes
Second-opinion terminology and E2E selections are advisory. They do not change the primary assessment or E2E / PR Gate. 2 semantic terminology decisionsTerminology decisions are advisory. They affect the assessment only when a separate finding identifies concrete semantic impact.
E2E guidanceAdvisory only. E2E / PR Gate selects and runs jobs independently. Recommended E2E: This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge. |
Physical DGX Spark follow-upI ran the exact runtime shape produced by this PR on the issue's DGX Spark/GB10 host:
For comparison, the requested This validates the PR as a memory-safety mitigation, not as a complete long-context performance fix. The MTP-off agent run still had severe tail latency (maximum turn 170s), although it completed all turns, and the original complete host freeze was not reproduced. Credit: @dfernandez365-rgb prepared and independently reviewed the earlier narrow candidate at dfernandez365-rgb@2637cb6. This PR is the sole upstream integration vehicle for the same MTP-off/async-on direction, reconciled with current |
|
The narrow MTP-off default is supported by the DGX Spark evidence in #7127, and I found no blocking defect in the four-file contribution. The PR still needs these repository gates before merge:
The DCO declaration is present, and GitHub shows both current commits as Verified. Maintainer edits are disabled for this PR, so the branch refresh must come from an authorized branch owner. |
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
<!-- markdownlint-disable MD041 --> ## Summary Add the canonical dated changelog entry for the planned NemoClaw v0.0.103 release. The new `docs/changelog/2026-08-05.mdx` entry uses the exact `## v0.0.103` heading and summarizes supported user-visible changes merged since v0.0.102. ## Changes - Add the parser-safe MDX SPDX header, three-paragraph release summary, and detailed grouped bullets to `docs/changelog/2026-08-05.mdx`. - Link each release-note group to the most specific published OpenClaw, Hermes, or Deep Agents documentation routes. - Exclude dormant MXC and Podman foundations, internal managed-inference adapters, test-only changes, and maintainer tooling from the supported product narrative. ### Source summary - [#8082](#8082) -> `docs/changelog/2026-08-05.mdx`: Document the new one-command agent launch flow. - [#8314](#8314) -> `docs/changelog/2026-08-05.mdx`: Document managed vLLM host capability validation and restart handling. - [#8248](#8248) -> `docs/changelog/2026-08-05.mdx`: Record the DGX Spark Qwen profile MTP default change. - [#8223](#8223) -> `docs/changelog/2026-08-05.mdx`: Record explicit model preservation across provider switches. - [#8209](#8209) -> `docs/changelog/2026-08-05.mdx`: Document corrected Windows WSL provider selection. - [#8316](#8316) -> `docs/changelog/2026-08-05.mdx`: Record clean managed-checkout reuse after installation. - [#8239](#8239) -> `docs/changelog/2026-08-05.mdx`: Record the packaged-service teardown fallback. - [#8247](#8247) -> `docs/changelog/2026-08-05.mdx`: Document uninstall behavior for an already-removed sandbox. - [#7998](#7998) -> `docs/changelog/2026-08-05.mdx`: Record preserved container-start diagnostics. - [#8027](#8027) -> `docs/changelog/2026-08-05.mdx`: Record journal-backed not-ready repair authority. - [#7812](#7812) -> `docs/changelog/2026-08-05.mdx`: Document actionable rebuild preflight diagnostics. - [#8222](#8222) -> `docs/changelog/2026-08-05.mdx`: Record redacted top-level CLI failures. - [#8313](#8313) -> `docs/changelog/2026-08-05.mdx`: Record structured MCP bridge destruction failures. - [#8211](#8211) -> `docs/changelog/2026-08-05.mdx`: Document cleanup of incomplete snapshot captures. - [#8212](#8212) -> `docs/changelog/2026-08-05.mdx`: Document best-effort post-restore policy reconciliation. - [#8245](#8245) -> `docs/changelog/2026-08-05.mdx`: Clarify manifest-defined OpenClaw workspace persistence. - [#8254](#8254) -> `docs/changelog/2026-08-05.mdx`: Include corrected snapshot restore selection guidance. - [#8238](#8238) -> `docs/changelog/2026-08-05.mdx`: Document preservation of managed MCP policy entries. - [#7568](#7568) -> `docs/changelog/2026-08-05.mdx`: Record mutable-default Shields rollback preservation. - [#8200](#8200) -> `docs/changelog/2026-08-05.mdx`: Record truthful Shields state after a rejected transition. - [#7895](#7895) -> `docs/changelog/2026-08-05.mdx`: Record descriptor-bound Shields lock inspection. - [#7892](#7892) -> `docs/changelog/2026-08-05.mdx`: Document the canonical Hermes dashboard profile and migration. - [#7871](#7871) -> `docs/changelog/2026-08-05.mdx`: Document fail-closed Hermes cron restore. - [#7894](#7894) -> `docs/changelog/2026-08-05.mdx`: Record the reset Hermes health budget after recovery. - [#8228](#8228) -> `docs/changelog/2026-08-05.mdx`: Document Hermes build-time corporate CA trust. - [#8206](#8206) -> `docs/changelog/2026-08-05.mdx`: Document bounded Deep Agents Code failure classification. - [#8297](#8297) -> `docs/changelog/2026-08-05.mdx`: Record reuse of the published Deep Agents Code base image. - [#8321](#8321) -> `docs/changelog/2026-08-05.mdx`: Document aligned endpoint SSRF protections and userinfo rejection. - [#8299](#8299) -> `docs/changelog/2026-08-05.mdx`: Document the fail-closed `setpriv` transition in managed images. - [#7603](#7603) -> `docs/changelog/2026-08-05.mdx`: Record corrected confidentiality-root traversal. - [#8334](#8334) -> `docs/changelog/2026-08-05.mdx`: Record removal of the unsupported logs audit example. - [#8256](#8256) -> `docs/changelog/2026-08-05.mdx`: Record reordered network-policy walkthrough prerequisites. - [#7767](#7767) -> `docs/changelog/2026-08-05.mdx`: Record platform runtime shape validation. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [x] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [ ] Tests added or updated for changed behavior - [x] Existing tests cover changed behavior — justification: `npx vitest run test/changelog-docs.test.ts` passed all 6 tests. - [ ] Tests not applicable — justification: - [x] Docs updated for user-facing behavior changes - [ ] Docs not applicable — justification: - [ ] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [ ] Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: ## Documentation Writer Review - [ ] Documentation writer subagent reviewed the completed changes - Result: `docs-updated` - Evidence: `docs/changelog/2026-08-05.mdx` follows the release-prep and documentation writing rules. The changelog contract tests passed 6/6, and `npm run docs` completed with 0 errors and the repository's 2 existing Fern warnings. - Agent: Codex Desktop <!-- docs-review-head-sha: 66fcd80 --> <!-- docs-review-agents-blob-sha: 3dd7c24 --> ## DGX Station Hardware Evidence - [ ] Tested on DGX Station - Tested commit: Not applicable. - Station profile/scenario: Not applicable. - Result: Not applicable. - Supporting evidence: Not applicable. ## Verification - [x] PR description includes a `Signed-off-by:` line and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run validate:pr` passed after refreshing `origin/main` when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — `npx vitest run test/changelog-docs.test.ts`: 1 file and 6 tests passed. - [ ] Applicable broad gate passed — `npm test` for broad runtime/test-harness changes; `npm run check` for repo-wide validation/coverage changes — command/result: Not run for this doc-only change. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) — completed with 0 errors and 2 existing Fern warnings. - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) — the native changelog uses the required parser-safe MDX SPDX comment and does not use page frontmatter. --- Signed-off-by: Charan Jagwani <cjagwani@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added release notes for v0.0.103. * Documented the new `nemoclaw launch` command. * Included updates covering onboarding, inference, installation, recovery, snapshots, security, integrations, endpoint validation, sandbox hardening, and related guidance. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
Summary
Disable multi-token prediction (MTP) speculative decoding in the managed single-node DGX Spark profile for
nvidia/Qwen3.6-35B-A3B-NVFP4. Controlled Spark results in #7127 isolate MTP as the stronger contributor to the profile's cold-start memory pressure and long-context instability; the original fatal host freeze was not reproduced, so this PR does not claim to fix it.Related Issue
Addresses #7127. The issue remains useful for the remaining physical cold-start, kernel-event, and soak-test acceptance criteria.
Changes
--speculative-configMTP argument.0.4GPU-memory utilization, the 262K context limit, batching/concurrency limits, FP8 KV cache, chunked prefill, prefix caching, parsers, backends, and fastsafetensors loading.32Kcontext, one sequence,4096batched tokens, async off) with its throughput tradeoff and no freeze-prevention guarantee.v0.0.102changelog entry.The supporting physical-Spark A/B matrix in #7127 observed the following; these are issue results, not validation of this exact PR commit:
-20.19 GiB42.04 GiB41 GiB+0.54 GiB23.29 GiB62 GiB-20.19 GiBclass of plan42 GiB41 GiBThe negative estimate is supporting evidence, not an asserted trigger. The controlled matrix reproduced a recoverable
NV_ERR_NO_MEMORYduring the MTP drafter cold-start path, but reproduced neither the original hard freeze nor host/SSH loss.Type of Change
Quality Gates
04024f3f6a0d2323875bc1478738eca028d9a9daagainst base3a39ff352f98c4630ab8e18b83f4ec43ac56a2cc, treee94fcb39978ce71549767e5efc3b59313556c9d7, and binary diff SHA-25685b387d485bb2ce37cbd48d1c2490b58c2601215bfcf99755cb68a528db06bdc. Removing MTP reduces memory pressure without weakening a security control; the architecture limits are tightened, not relaxed.Documentation Writer Review
docs-updated04024f3f6a0d2323875bc1478738eca028d9a9daagainst base3a39ff352f98c4630ab8e18b83f4ec43ac56a2cc, treee94fcb39978ce71549767e5efc3b59313556c9d7, and binary diff SHA-25685b387d485bb2ce37cbd48d1c2490b58c2601215bfcf99755cb68a528db06bdc. The documentation matches the implementation, all OpenClaw, Hermes, and Deep Agents generated variants passed review, and no findings remain.npm run validate:pr, CLI build, CLI typecheck, repository checks, and docs build passed; docs reported 0 errors and 2 existing warnings; focused tests passed 63/63;git diff --checkpassed./root/pr8073_docs_review)DGX Station Hardware Evidence
Verification
Signed-off-by:line and every commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run validate:prpassed after refreshingorigin/mainwhen hooks were skipped or unavailablenpx vitest run --project cli src/lib/inference/vllm-models.test.ts(40/40);npx vitest run --project integration test/inference-options-docs.test.ts test/changelog-docs.test.ts(29/29); exact source-architecture baseline assertion (1/1).npm testfor broad runtime/test-harness changes;npm run checkfor repo-wide validation/coverage changes — broad local check passed all pre-commit/manual checks through test budgets, then its macOS CLI/integration coverage process hung after its worker exited and was stopped. Do not treat it as passing evidence; exact-head required CI must pass before merge.npm run docsbuilds without warnings (doc changes only) — zero errors; Fern reported two existing warnings that its normal output does not printSigned-off-by: Prekshi Vyas prekshiv@nvidia.com
Summary by CodeRabbit
New Features
Documentation