Skip to content

fix(installer): avoid vLLM process false positives - #7541

Merged
cv merged 2 commits into
NVIDIA:mainfrom
senthilr-nv:codex/fix-station-vllm-process-detection
Jul 26, 2026
Merged

fix(installer): avoid vLLM process false positives#7541
cv merged 2 commits into
NVIDIA:mainfrom
senthilr-nv:codex/fix-station-vllm-process-detection

Conversation

@senthilr-nv

@senthilr-nv senthilr-nv commented Jul 25, 2026

Copy link
Copy Markdown
Collaborator

Summary

DGX Station preparation treated unrelated diagnostic commands such as grep -qi vllm as active inference workloads. Process checks now require a vLLM executable or Python module launch while preserving existing container and port conflict checks.

Changes

  • Recognize direct vLLM executables, python -m vllm modules, and docker-init vLLM commands.
  • Ignore grep, rg, and shell diagnostics that only mention vllm as data.
  • Apply the same classification to Station preparation and saved Local vLLM continuation.
  • Add regression coverage for false-positive diagnostics, real workload detection, and model-name redaction.

Type of Change

  • Code change (feature, bug fix, or refactor)
  • Code change with doc updates
  • Doc only (prose changes, no code sample modifications)
  • Doc only (includes code sample changes)

Quality Gates

  • Tests added or updated for changed behavior
  • Existing tests cover changed behavior — justification:
  • Tests not applicable — justification:
  • Docs updated for user-facing behavior changes
  • Docs not applicable — justification: Existing Station documentation already limits the ownership flow to an existing vLLM workload. This fix restores that documented contract.
  • Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging)
  • Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: Maintainer implementation review is pending. The existing-vLLM ownership contract was accepted in PR fix(installer): hand off existing Station vLLM #7285.
  • Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue:

Documentation Writer Review

  • Documentation writer subagent reviewed the completed changes
  • Result: no-docs-needed
  • Evidence: No documentation path changed. Existing DGX Station documentation already describes the intended active-vLLM workload contract, and this fix excludes unrelated diagnostic processes from that contract.
  • Agent: Codex Desktop

DGX Station Hardware Evidence

  • Tested on DGX Station
  • Tested commit: a5f7262c7ea2aae5762deb30a7695e477aec8edb
  • Station profile/scenario: pmgb300ws-0016, generic Ubuntu 24.04.4 ARM64 on a DGX Station GB300. A bounded read-only check used the exact committed script bytes against diagnostic process rows and the running managed vLLM.
  • Result: PASS. grep, rg, and shell diagnostic rows did not block. The real managed vLLM still produced exit 12, Local vLLM continuation still detected it, ECC remained 0/0, and the managed container remained running.
  • Supporting evidence: Exact-commit DGX Station classifier evidence

Verification

  • PR description includes a Signed-off-by: line and every commit appears as Verified in GitHub
  • Normal pre-commit, commit-msg, and pre-push hooks passed, or npm run check:diff passed when hooks were skipped or unavailable
  • Targeted behavior tests pass for the current change set, or tests are marked not applicable above — command/result or justification: npx vitest run test/install-station-host-preparation.test.ts test/install-station-vllm-continuation.test.ts passed 65 tests. The related six-file Station suite passed five files; after building its required CLI artifacts, test/install-preflight.test.ts passed 94 tests.
  • Applicable broad gate passed — npm test for broad runtime/test-harness changes; npm run check for repo-wide validation/coverage changes — command/result: Not run; the focused Station preparation and continuation suites cover this classifier change.
  • Quality Gates section completed with required justifications or waivers
  • No secrets, API keys, or credentials committed
  • npm run docs builds without warnings (doc changes only)
  • Doc pages follow the style guide (doc changes only)
  • New doc pages include SPDX header and frontmatter (new pages only)

Signed-off-by: Senthil Ravichandran senthilr@nvidia.com

@senthilr-nv senthilr-nv self-assigned this Jul 25, 2026
@coderabbitai

coderabbitai Bot commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

vLLM conflict detection now recognizes direct executables, Python module launches, and docker-init wrappers while ignoring diagnostic commands that only mention vLLM. Tests cover these process forms, conflict reporting, stop commands, and model-name redaction.

Changes

vLLM workload detection

Layer / File(s) Summary
Expand process matching
scripts/lib/station-vllm-conflict.sh, scripts/prepare-dgx-station-host.sh
Process scanning normalizes executable and argument fields and recognizes direct vLLM processes, Python -m vllm launches, and docker-init vLLM commands.
Validate active and diagnostic process cases
test/install-station-host-preparation.test.ts, test/install-station-vllm-continuation.test.ts
Tests verify diagnostic commands are ignored, supported active process forms are detected, conflicts return status 12, stop commands are reported, and model names are omitted.

Estimated code review effort: 3 (Moderate) | ~20 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: improving installer vLLM process detection to avoid false positives.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@senthilr-nv

Copy link
Copy Markdown
Collaborator Author

DGX Station hardware evidence

Tested commit: 486eeca5c4438ce301f25581c1770ccb1b89b855

Scenario: bounded read-only classifier check on pmgb300ws-0016, generic Ubuntu 24.04.4 ARM64, DGX Station GB300. The test used the exact committed script bytes. It did not stop or change a process, container, port, or service.

Result: PASS

  • grep, rg, and shell diagnostic rows that mention vllm did not block Station preparation.
  • The saved Local vLLM continuation classifier also ignored those diagnostic rows.
  • The real managed vLLM remained detected by both classifiers.
  • ECC remained 0 corrected / 0 uncorrected.
  • The managed vLLM container remained running.
  • Local evidence SHA-256: cfc3a28c7e165b85012d56c200126d9babca6d313d56faadaa100751ca695fd5
tested_commit=486eeca5c4438ce301f25581c1770ccb1b89b855
prepare_sha256=62dc6c120f4d8616df65b9cad64ef63a891e9dad623530f6895784075dbd9652
conflict_sha256=269cc6e090615fccf3630b3cacbde71a328f2921d42181d5bd664bd014fddc45
hostname=pmgb300ws-0016
os=Ubuntu 24.04.4 LTS
arch=aarch64
boot_id=5755e984-11c2-43d5-bac4-9c079694dcf4
product=Dell Pro Max with Station GB300
diagnostic_prepare_rc=0
[station-prepare] agent_inference_workloads=none port_8000=free
diagnostic_continuation_rc=1
actual_prepare_rc=12
[station-prepare] ERROR: vLLM inference workload is active: process=vllm. NemoClaw did not stop or modify it.
actual_continuation_rc=0
NVIDIA GB300, 610.43.02, 0, 0
managed_vllm_status=running
HARDWARE_TEST_RESULT=PASS

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@test/install-station-host-preparation.test.ts`:
- Around line 446-464: Extend the test case around
check_agent_and_inference_conflicts to include a docker-init command invoking
vLLM with a sensitive model name. Assert the same active-conflict status, PID
and stop-command reporting, and model-name redaction already checked for vLLM
and Python processes, using the scanner’s observable output boundary.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: d2021dc3-5391-4ac5-a740-4426d5f15519

📥 Commits

Reviewing files that changed from the base of the PR and between 12b9440 and 486eeca.

📒 Files selected for processing (4)
  • scripts/lib/station-vllm-conflict.sh
  • scripts/prepare-dgx-station-host.sh
  • test/install-station-host-preparation.test.ts
  • test/install-station-vllm-continuation.test.ts

Comment thread test/install-station-host-preparation.test.ts
@github-actions

github-actions Bot commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

PR Review Advisor — No blocking findings reported

Advisor assessment: No blocking advisor findings reported
Next action: No advisor follow-up needed.
Findings: 0 blockers · 0 warnings · 0 suggestions

Model lanes

  • GPT-5.6 Terra (primary): Completed · high confidence · 0 blockers · 0 warnings · 0 suggestions
  • Nemotron 3 Ultra (second opinion): Completed · high confidence · 0 blockers · 0 warnings · 0 suggestions
  • Model comparison: normalized findings match; normalized E2E selections differ; severity counts match.

Nemotron output stays in workflow artifacts and does not change the assessment above.

E2E guidance

Advisory only. E2E / PR Gate selects and runs jobs independently.

Recommended E2E: None

Workflow run details

This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge.

@senthilr-nv

Copy link
Copy Markdown
Collaborator Author

DGX Station hardware evidence

Tested commit: a5f7262c7ea2aae5762deb30a7695e477aec8edb

Scenario: bounded read-only classifier check on pmgb300ws-0016, generic Ubuntu 24.04.4 ARM64, DGX Station GB300. The test used the exact committed script bytes. It did not stop or change a process, container, port, or service.

Result: PASS

  • grep, rg, and shell diagnostic rows that mention vllm did not block Station preparation.
  • The saved Local vLLM continuation classifier also ignored those diagnostic rows.
  • The real managed vLLM remained detected by both classifiers.
  • ECC remained 0 corrected / 0 uncorrected.
  • The managed vLLM container remained running.
  • Local evidence SHA-256: d3aae8c3df689235c11ffac09844e752d437d6b1378ea329796a328ffebe9626
tested_commit=a5f7262c7ea2aae5762deb30a7695e477aec8edb
prepare_sha256=62dc6c120f4d8616df65b9cad64ef63a891e9dad623530f6895784075dbd9652
conflict_sha256=269cc6e090615fccf3630b3cacbde71a328f2921d42181d5bd664bd014fddc45
hostname=pmgb300ws-0016
os=Ubuntu 24.04.4 LTS
arch=aarch64
boot_id=5755e984-11c2-43d5-bac4-9c079694dcf4
product=Dell Pro Max with Station GB300
diagnostic_prepare_rc=0
[station-prepare] agent_inference_workloads=none port_8000=free
diagnostic_continuation_rc=1
actual_prepare_rc=12
[station-prepare] ERROR: vLLM inference workload is active: process=vllm. NemoClaw did not stop or modify it.
actual_continuation_rc=0
NVIDIA GB300, 610.43.02, 0, 0
managed_vllm_status=running
HARDWARE_TEST_RESULT=PASS

@senthilr-nv

Copy link
Copy Markdown
Collaborator Author

Reviewed PRA-1 on a5f7262c7. The Python-module regex is not anchored after the interpreter; it already finds -m vllm later in the command line. The follow-up tests now prove python3 -u -m vllm... in host preparation and python3 -u -X dev -m vllm... in Local vLLM continuation. Both classify the workload as active, and the focused suite passes 65/65. No classifier broadening was needed.

@senthilr-nv senthilr-nv added v0.0.96 bug-fix PR fixes a bug or regression area: install Install, setup, prerequisites, or uninstall flow area: local-models Local model providers, downloads, launch, or connectivity provider: vllm vLLM local or hosted provider behavior platform: dgx-station Affects DGX Station hardware or workflows labels Jul 25, 2026
@cv
cv merged commit cf4cd43 into NVIDIA:main Jul 26, 2026
91 of 97 checks passed
@cv cv mentioned this pull request Jul 26, 2026
23 tasks
apurvvkumaria pushed a commit that referenced this pull request Jul 27, 2026
<!-- markdownlint-disable MD041 -->
## Summary

Add the canonical `docs/changelog/2026-07-25.mdx` release entry with the
exact `## v0.0.96` heading.
The entry reconciles all 90 first-parent commits since v0.0.95 with all
92 merged PRs in the live `v0.0.96` label ledger and groups the
user-visible changes by operator journey.

## Changes

- Add the parser-safe dated MDX changelog entry for v0.0.96 with
root-absolute links to the focused user guides.
- Source summary:
- [#7194](#7194) ->
`docs/changelog/2026-07-25.mdx`: Document persistent baseline network
policy exclusions and their inspection, rebuild, and snapshot behavior.
- [#7188](#7188),
[#7427](#7427), and
[#7546](#7546) ->
`docs/changelog/2026-07-25.mdx`: Document DNS-backed HTTPS inference
routing, keyless loopback endpoints, and provider-marker isolation.
- [#7238](#7238) ->
`docs/changelog/2026-07-25.mdx`: Document blueprint sandbox and provider
identifier validation before state writes or OpenShell calls, with
bounded terminal-safe rejection previews.
- [#7319](#7319),
[#7274](#7274),
[#7528](#7528),
[#7353](#7353), and
[#7560](#7560) ->
`docs/changelog/2026-07-25.mdx`: Document the managed default gateway
service, onboarding readiness, and container-runtime identity
safeguards.
- [#7349](#7349),
[#7498](#7498),
[#7406](#7406),
[#7196](#7196),
[#7559](#7559),
[#7421](#7421),
[#7510](#7510),
[#7295](#7295), and
[#7565](#7565) ->
`docs/changelog/2026-07-25.mdx`: Document gateway-scoped status,
lifecycle diagnostics, managed MCP recovery, delete-edge safeguards, and
fail-closed CLI prompt and command output.
- [#7591](#7591) ->
`docs/changelog/2026-07-25.mdx`: Document opt-in authenticated MCP
tool-name discovery, its bounded and names-only contract, probe
interaction, and rebuild requirement.
- [#7305](#7305),
[#7480](#7480),
[#7471](#7471),
[#7365](#7365), and
[#7541](#7541) ->
`docs/changelog/2026-07-25.mdx`: Document installer version checks,
version-tag reporting, license guidance, WSL Ollama selection, and DGX
Station vLLM detection.
- [#7482](#7482),
[#7466](#7466),
[#7208](#7208),
[#7434](#7434), and
[#7586](#7586) ->
`docs/changelog/2026-07-25.mdx`: Document Ollama resource details,
reasoning precedence, Hermes onboarding behavior, and preserved managed
Hermes BuildKit failures.

- [#6830](#6830),
[#7492](#7492),
[#7563](#7563), and
[#7582](#7582) ->
`docs/changelog/2026-07-25.mdx`: Document the authoritative OpenClaw
production lock, fixed managed-image dependencies, immutable Hermes base
adoption, and Hermes image-size reduction.
- [#7505](#7505),
[#7530](#7530),
[#7547](#7547),
[#7508](#7508),
[#7548](#7548),
[#7549](#7549),
[#7537](#7537),
[#7534](#7534),
[#7515](#7515),
[#7511](#7511),
[#7551](#7551),
[#7562](#7562),
[#7575](#7575),
[#7496](#7496),
[#7594](#7594),
[#7595](#7595), and
[#7599](#7599) ->
`docs/changelog/2026-07-25.mdx`: Summarize release validation, transient
and bounded dispatch reconciliation, exact pre-tag qualification,
identity revalidation, npm-audit retry, sharding, image reuse, timeout,
telemetry, and workflow-hardening changes.
- Reconciled without separate changelog prose:
- [#7539](#7539),
[#7526](#7526),
[#7507](#7507),
[#7506](#7506),
[#7519](#7519),
[#7516](#7516),
[#7396](#7396),
[#7254](#7254),
[#7583](#7583),
[#7596](#7596), and
[#7598](#7598): Test-harness or
fixture-only changes.
- [#7403](#7403),
[#7161](#7161),
[#6877](#6877),
[#7531](#7531),
[#7525](#7525),
[#7522](#7522),
[#7536](#7536),
[#7552](#7552),
[#7566](#7566),
[#7553](#7553),
[#7561](#7561),
[#7577](#7577),
[#7569](#7569),
[#7585](#7585),
[#7584](#7584),
[#7592](#7592),
[#7580](#7580),
[#7571](#7571),
[#7517](#7517),
[#7589](#7589),
[#7402](#7402),
[#7558](#7558),
[#7544](#7544), and
[#7601](#7601): Dependency,
internal recovery, validation, contributor-workflow, E2E optimization,
telemetry, or CI trust changes with no separate user-facing release
claim.
- [#7556](#7556),
[#7573](#7573),
[#7576](#7576), and
[#7578](#7578): Experimental
repository-maintainer conflict automation with no canonical user
documentation surface.

## Type of Change

- [ ] Code change (feature, bug fix, or refactor)
- [ ] Code change with doc updates
- [x] Doc only (prose changes, no code sample modifications)
- [ ] Doc only (includes code sample changes)

## Quality Gates

- [ ] Tests added or updated for changed behavior
- [x] Existing tests cover changed behavior — justification:
`test/changelog-docs.test.ts` validates dated changelog structure,
version headings, and published links.
- [ ] Tests not applicable — justification:
- [x] Docs updated for user-facing behavior changes
- [ ] Docs not applicable — justification:
- [ ] Sensitive paths changed (security, policy, credentials, preflight,
onboarding, inference, runner, sandbox, or messaging)
- [ ] Sensitive-path review completed or maintainer-approved waiver
recorded — reviewer/approval link/justification:
- [ ] Non-success, skipped, or missing CI check accepted by maintainer —
check name, approval link, and follow-up issue:

## Documentation Writer Review

- [x] Documentation writer subagent reviewed the completed changes
- Result: `docs-updated`
- Evidence: Reviewed `docs/changelog/2026-07-25.mdx` at exact head
`0f5dedb47` against 90 first-parent release commits and 92 merged PRs
labeled `v0.0.96`. Verified parser-safe MDX SPDX, the exact version
heading, literal CLI names, writing style, skip terms, all 20
root-absolute published links, and the accepted #7591 opt-in
authenticated discovery bounds. #7544, #7599, and #7601 remain internal
or CI-only release-ledger entries. Changelog tests passed 6/6, the docs
build passed with 0 errors and two pre-existing Fern warnings, and `npm
run check:diff` plus the final diff check passed.
- Agent: Codex Desktop documentation-writer subagent
<!-- docs-review-head-sha: 0f5dedb -->
<!-- docs-review-agents-blob-sha: be20a09 -->

## DGX Station Hardware Evidence

- [ ] Tested on DGX Station
- Tested commit:
- Station profile/scenario:
- Result:
- Supporting evidence:

## Verification

- [x] PR description includes a `Signed-off-by:` line and every commit
appears as `Verified` in GitHub
- [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or
`npm run check:diff` passed when hooks were skipped or unavailable
- [x] Targeted behavior tests pass for the current change set, or tests
are marked not applicable above — `npx vitest run
test/changelog-docs.test.ts`: 6/6 passed.
- [ ] Applicable broad gate passed — `npm test` for broad
runtime/test-harness changes; `npm run check` for repo-wide
validation/coverage changes — command/result: Not applicable to this
prose-only changelog entry.
- [x] Quality Gates section completed with required justifications or
waivers
- [x] No secrets, API keys, or credentials committed
- [ ] `npm run docs` builds without warnings (doc changes only) — the
build passed with 0 errors and 2 existing Fern warnings; the
published-route check passed.
- [x] Doc pages follow the [style
guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md)
(doc changes only)
- [ ] New doc pages include SPDX header and frontmatter (new pages only)
— native changelog files use the required parser-safe MDX SPDX comment
and no frontmatter.

---
Signed-off-by: Carlos Villela <cvillela@nvidia.com>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Persistent network policy exclusions with consistent restore/exclusion
reporting across rebuilds/snapshots.
* Opt-in MCP tool discovery via `mcp status --tools` with bounded,
redacted authenticated traffic.
* Improved HTTPS inference switching for custom endpoints and refreshed
onboarding/model menu details.
* Refined OpenShell gateway defaults for port `8080`, including more
reliable readiness checks.
* **Bug Fixes**
* Prevent incorrect provider/model restoration after compatible-provider
update failures.
* Preserve managed MCP state after exec loss and tighten gateway/doctor
status scoping.
* **Tests**
* Stronger, fail-closed release validation with hardened
evidence/artifact handoff and bounded timeouts/retries.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com>
Co-authored-by: Prekshi Vyas <prekshiv@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: install Install, setup, prerequisites, or uninstall flow area: local-models Local model providers, downloads, launch, or connectivity bug-fix PR fixes a bug or regression platform: dgx-station Affects DGX Station hardware or workflows provider: vllm vLLM local or hosted provider behavior

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants