ci(e2e): simplify Launchable E2E dispatch - #7666
Conversation
Signed-off-by: J. Yaunches <jmyaunch@gmail.com>
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughThe PR removes readiness-variable gating from staging Brev Launchable E2E qualification, adds evidence-driven full-mode release dispatch, updates workflow validation and ledger filtering, and documents separate Ordinary, Launchable, and Full-mode procedures. ChangesLaunchable E2E qualification
Estimated code review effort: 3 (Moderate) | ~25 minutes Sequence Diagram(s)sequenceDiagram
participant Maintainer
participant ReleasePreflight
participant E2EWorkflow
participant LaunchableJob
Maintainer->>ReleasePreflight: Provide candidate SHA and runner status
ReleasePreflight-->>Maintainer: Return full-mode dispatch plan
Maintainer->>E2EWorkflow: Dispatch trusted full or Launchable mode
E2EWorkflow->>LaunchableJob: Select staging qualification
LaunchableJob-->>Maintainer: Publish run and job evidence
Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
PR Review Advisor — No blocking findings reportedAdvisor assessment: No blocking advisor findings reported Model lanes
Nemotron output stays in workflow artifacts and does not change the assessment above. E2E guidanceAdvisory only. E2E / PR Gate selects and runs jobs independently. Recommended E2E: 1 optional E2E recommendation
This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge. |
There was a problem hiding this comment.
🧹 Nitpick comments (2)
test/e2e/support/e2e-workflow.test.ts (1)
66-103: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick winCover the non-empty
targetsboundary.The implementation requires both
jobsandtargetsto be empty for full qualification, but this test only exercises a non-emptyjobsselector. Add a case withincludeStagingBrevLaunchable: true, emptyjobs, and non-emptytargets, expectingrunQualification: false.As per path instructions, this test should verify the dispatch contract through observable behavior at the public boundary.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/e2e/support/e2e-workflow.test.ts` around lines 66 - 103, Add a case to the evaluateStagingBrevLaunchableDispatch test with includeStagingBrevLaunchable true, empty jobs, and a non-empty targets selector, asserting runQualification is false. Keep the assertion at the public evaluateStagingBrevLaunchableDispatch boundary.Source: Path instructions
tools/e2e/workflow-boundary.mts (1)
4793-4797: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick winKeep an explicit regression guard for readiness-job removal.
Removing both the validator call and its drift-test mutation means a future
staging-brev-launchable-readinessjob could return without failing boundary validation.
tools/e2e/workflow-boundary.mts#L4793-L4797: reject the forbidden readiness job explicitly.test/e2e/support/e2e-workflow.test.ts#L124-L137: restore a mutation and expected validation error proving the superseded path remains removed.As per path instructions, migration tests must prove the superseded path is unreachable or removed.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tools/e2e/workflow-boundary.mts` around lines 4793 - 4797, Keep an explicit regression guard for the removed readiness job: in tools/e2e/workflow-boundary.mts lines 4793-4797, add validation that rejects staging-brev-launchable-readiness; in test/e2e/support/e2e-workflow.test.ts lines 124-137, restore the drift-test mutation and expected validation error proving that superseded path remains unreachable.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@test/e2e/support/e2e-workflow.test.ts`:
- Around line 66-103: Add a case to the evaluateStagingBrevLaunchableDispatch
test with includeStagingBrevLaunchable true, empty jobs, and a non-empty targets
selector, asserting runQualification is false. Keep the assertion at the public
evaluateStagingBrevLaunchableDispatch boundary.
In `@tools/e2e/workflow-boundary.mts`:
- Around line 4793-4797: Keep an explicit regression guard for the removed
readiness job: in tools/e2e/workflow-boundary.mts lines 4793-4797, add
validation that rejects staging-brev-launchable-readiness; in
test/e2e/support/e2e-workflow.test.ts lines 124-137, restore the drift-test
mutation and expected validation error proving that superseded path remains
unreachable.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 3254e819-34fd-4a6e-a238-5c5d455701c5
📒 Files selected for processing (6)
.agents/skills/nemoclaw-maintainer-cut-release-tag/SKILL.md.agents/skills/nemoclaw-maintainer-e2e/SKILL.md.github/workflows/e2e.yamltest/e2e/README.mdtest/e2e/support/e2e-workflow.test.tstools/e2e/workflow-boundary.mts
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage in commit a87d7b0 in the TypeScript / code-coverage/cliThe overall coverage in commit a87d7b0 in the Show a code coverage summary of the most impacted files.
Updated |
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
.agents/skills/nemoclaw-maintainer-cut-release-tag/scripts/release-e2e-evidence.mts (1)
448-470: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick winBind each job record to the enclosing workflow run.
Match onjob.run_id === runIdbefore recording attempts; otherwise ajobsJsonpayload from another run can be counted as green. Also ignore job attempts newer than the run attempt.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.agents/skills/nemoclaw-maintainer-cut-release-tag/scripts/release-e2e-evidence.mts around lines 448 - 470, Update the job-processing loop around flattenJobs so each job is recorded only when its run_id matches the enclosing runId, and skip attempts whose run_attempt is newer than the enclosing run attempt. Apply these checks before pushing values into attempts, while preserving the existing name matching and ambiguity handling.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In
@.agents/skills/nemoclaw-maintainer-cut-release-tag/scripts/release-e2e-evidence.mts:
- Around line 448-470: Update the job-processing loop around flattenJobs so each
job is recorded only when its run_id matches the enclosing runId, and skip
attempts whose run_attempt is newer than the enclosing run attempt. Apply these
checks before pushing values into attempts, while preserving the existing name
matching and ambiguity handling.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 6121c31c-08d4-44f4-8614-896a1239116c
📒 Files selected for processing (11)
.agents/skills/nemoclaw-maintainer-cut-release-tag/SKILL.md.agents/skills/nemoclaw-maintainer-cut-release-tag/scripts/release-e2e-evidence.mts.agents/skills/nemoclaw-maintainer-e2e/SKILL.md.agents/skills/nemoclaw-maintainer-evening/SKILL.md.agents/skills/nemoclaw-maintainer-policies/references/release-train.md.github/workflows/e2e.yamltest/e2e/README.mdtest/e2e/support/e2e-workflow.test.tstest/maintainer-skills-policy.test.tstest/release-e2e-evidence.test.tstools/e2e/workflow-boundary.mts
🚧 Files skipped from review as they are similar to previous changes (4)
- .github/workflows/e2e.yaml
- test/e2e/support/e2e-workflow.test.ts
- .agents/skills/nemoclaw-maintainer-e2e/SKILL.md
- tools/e2e/workflow-boundary.mts
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/release-e2e-evidence.test.ts`:
- Line 69: Update the fixture setup in the cross-run test to define a distinct
workflow run ID, use it for run_id, run.id, and workflowRunId, and use attempt
only for run-attempt fields. Preserve the test’s ledger assertions so distinct
contract values verify the observable behavior.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 58ff2ae3-ab41-448c-bfa5-439b4bcc72a5
📒 Files selected for processing (2)
.agents/skills/nemoclaw-maintainer-cut-release-tag/scripts/release-e2e-evidence.mtstest/release-e2e-evidence.test.ts
🚧 Files skipped from review as they are similar to previous changes (1)
- .agents/skills/nemoclaw-maintainer-cut-release-tag/scripts/release-e2e-evidence.mts
<!-- markdownlint-disable MD041 --> ## Summary Add the canonical dated changelog entry for NemoClaw v0.0.97 before the release plan captures `origin/main`. The entry groups the user-visible and maintainer-facing changes since v0.0.96 while preserving the Deferred dual-Station status, experimental runtime-identity boundary, and pending physical IGX validation. ## Changes - Add `docs/changelog/2026-07-28.mdx` with the parser-safe MDX SPDX comment and exact `## v0.0.97` heading. - Summarize the 43 merged PRs in the release range, omitting internal-only changes from the public entry and linking each grouped change to its most specific published documentation. - Keep the experimental Okta reference explicitly opt-in and outside normal onboarding, keep the two-Station path Deferred, and state that physical IGX Orin validation remains pending. ### Source summary - [#7440](#7440), [#7443](#7443), and [#7445](#7445) -> `docs/changelog/2026-07-28.mdx`: Document read-only host readiness reports and fail-closed platform qualification. - [#7030](#7030) -> `docs/changelog/2026-07-28.mdx`: Document the Deferred trusted two-Station vLLM evaluation. - [#7265](#7265) -> `docs/changelog/2026-07-28.mdx`: Document the bounded experimental direct-runner Okta runtime-identity reference. - [#7711](#7711) and [#7648](#7648) -> `docs/changelog/2026-07-28.mdx`: Document compatible-endpoint reasoning effort and retired NVIDIA Build model paths. - [#7746](#7746), [#7763](#7763), and [#7681](#7681) -> `docs/changelog/2026-07-28.mdx`: Document safe compatible-provider creation, replacement refusal, and narrow OpenShell bridge URL handling. - [#7641](#7641), [#7690](#7690), [#7631](#7631), and [#7710](#7710) -> `docs/changelog/2026-07-28.mdx`: Document paused-container recovery, recreation journaling, pre-mutation uninstall checks, and source-checkout OpenShell selection. - [#7624](#7624) and [#7762](#7762) -> `docs/changelog/2026-07-28.mdx`: Document Jetson release diagnostics and bounded render-device group propagation. - [#7639](#7639), [#7760](#7760), [#7721](#7721), and [#7761](#7761) -> `docs/changelog/2026-07-28.mdx`: Document Telegram, MCP media-type, Hermes image-mode, and locked-restart fixes. - [#7653](#7653) and [#7680](#7680) -> `docs/changelog/2026-07-28.mdx`: Document Deep Agents policy tasks and the bounded Claude Code OAuth path. - [#7679](#7679) -> `docs/changelog/2026-07-28.mdx`: Document the checksum-bound libssh2 and Python HTMLParser backports. - [#7655](#7655), [#7651](#7651), [#7664](#7664), [#7666](#7666), [#7670](#7670), [#7719](#7719), and [#7741](#7741) -> `docs/changelog/2026-07-28.mdx`: Document exact candidate E2E evidence, Launchable selection, diagnostic consolidation, and trusted WSL validation. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [x] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [ ] Tests added or updated for changed behavior - [x] Existing tests cover changed behavior — justification: `test/changelog-docs.test.ts` validates the dated changelog contract, MDX header, heading uniqueness, and release-entry structure. - [ ] Tests not applicable — justification: - [x] Docs updated for user-facing behavior changes - [ ] Docs not applicable — justification: - [ ] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [ ] Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: ## Documentation Writer Review - [x] Documentation writer subagent reviewed the completed changes - Result: `docs-updated` - Evidence: The committed `docs/changelog/2026-07-28.mdx` blob exactly matches the reviewed file. Completeness, factual accuracy, link shape, parser-safe MDX header, one-sentence-per-line style, `.docs-skip` compliance, and bounded product claims passed. - Agent: Codex Desktop documentation writer subagent <!-- docs-review-head-sha: da6aa27 --> <!-- docs-review-agents-blob-sha: be20a09 --> ## DGX Station Hardware Evidence - [ ] Tested on DGX Station - Tested commit: Not applicable; this PR changes only the dated changelog. - Station profile/scenario: Not applicable. - Result: Not applicable. - Supporting evidence: Not applicable. ## Verification - [x] PR description includes a `Signed-off-by:` line and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run check:diff` passed when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — `npx vitest run test/changelog-docs.test.ts` passed 6/6. - [ ] Applicable broad gate passed — `npm test` for broad runtime/test-harness changes; `npm run check` for repo-wide validation/coverage changes — not applicable to this doc-only release entry. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) — completed with 0 errors and 2 pre-existing Fern warnings. - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) — native changelog entries use the required parser-safe MDX SPDX comment and intentionally have no frontmatter. --- Signed-off-by: Charan Jagwani <cjagwani@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added improved host readiness reporting and Jetson onboarding guidance. * Added controls for reasoning effort with compatible endpoints and enhanced managed MCP discovery. * Improved Deep Agents task publication and preset support. * **Bug Fixes** * Hardened provider switching, sandbox recovery, uninstall behavior, and Telegram connectivity. * Improved container image integrity checks, media-type handling, and checksum validation. * Enhanced vLLM evaluation behavior and release diagnostics. * **Documentation** * Added the NemoClaw v0.0.97 changelog. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
Summary
Remove the persistent Launchable readiness gate that prevented one-off Launchable E2E runs. Trusted manual runs can now select only the Launchable job, while pre-tag full runs still bind the default E2E suite and the exact Launchable E2E job to one candidate SHA.
Changes
jobs=staging-brev-launchableto run one Launchable E2E run from trustedmain.include_staging_brev_launchable=trueas the full pre-tag evidence contract.Type of Change
Quality Gates
Documentation Writer Review
docs-updated.agents/skills/nemoclaw-maintainer-e2e/SKILL.md,.agents/skills/nemoclaw-maintainer-cut-release-tag/SKILL.md,.agents/skills/nemoclaw-maintainer-evening/SKILL.md,.agents/skills/nemoclaw-maintainer-policies/references/release-train.md, andtest/e2e/README.mdDGX Station Hardware Evidence
Verification
Signed-off-by:line and every commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run check:diffpassed when hooks were skipped or unavailablenpx vitest run --project e2e-support test/e2e/support/e2e-workflow.test.tspassed 45 tests;npx vitest run --project integration test/release-e2e-evidence.test.ts test/maintainer-skills-policy.test.ts test/maintainer-e2e-skill.test.ts test/brev-launchable-e2e.test.tspassed 43 tests.npm testfor broad runtime/test-harness changes;npm run checkfor repo-wide validation/coverage changes — command/result:npm run docsbuilds without warnings (doc changes only)Signed-off-by: J. Yaunches jmyaunch@gmail.com
Summary by CodeRabbit