feat(e2e): require exact pre-tag qualification - #7496
Conversation
Signed-off-by: J. Yaunches <jmyaunch@gmail.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
📝 WalkthroughWalkthroughAdds a trusted maintainer E2E skill, full-mode workflow dispatch with staging Brev Launchable qualification, strict evidence validation, workflow boundary checks, and candidate-SHA release confirmation gates. ChangesFull E2E release evidence
Estimated code review effort: 4 (Complex) | ~45 minutes Suggested labels: Suggested reviewers: Sequence Diagram(s)sequenceDiagram
participant Maintainer
participant MaintainerE2E
participant GitHubActions
participant BrevQualification
participant ReleaseTag
Maintainer->>MaintainerE2E: Request full E2E for candidate SHA
MaintainerE2E->>GitHubActions: Dispatch trusted e2e.yaml run
GitHubActions->>BrevQualification: Run Exact staging Brev Launchable
BrevQualification-->>GitHubActions: Return qualification and cleanup evidence
GitHubActions-->>MaintainerE2E: Return run, job, and receipt evidence
MaintainerE2E-->>ReleaseTag: Bind validated evidence to candidate SHA
ReleaseTag->>ReleaseTag: Require evidence or itemized exceptions
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage in commit 173ac9b in the TypeScript / code-coverage/cliThe overall coverage in commit 173ac9b in the Show a code coverage summary of the most impacted files.
Updated |
VerdictPASS. I reviewed the complete Findings TableNo findings. Detailed Analysis
Files Reviewed
|
PR Review Advisor — No blocking findings reportedAdvisor assessment: No blocking advisor findings reported Model lanes
Nemotron output stays in workflow artifacts and does not change the assessment above. E2E guidanceAdvisory only. E2E / PR Gate selects and runs jobs independently. Recommended E2E: 1 optional E2E recommendation
This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge. |
Signed-off-by: J. Yaunches <jmyaunch@gmail.com>
|
Exact-head reliability review at
That means a pre-tag full run or its qualification can still be superseded before it produces evidence—the cancellation mode this work is intended to prevent. Please give full-mode runs a non-superseding boundary and queue the serialized qualification (or add an equivalent regression-pinned design): https://docs.github.com/en/actions/how-tos/write-workflows/choose-when-workflows-run/control-workflow-concurrency No other blocker was found. Readiness/evidence validation fail closed, and the exact merge-result focused checks passed (44 workflow tests and 28 maintainer-policy tests). |
Isolate each empty-selector full dispatch with its server-assigned run ID. Queue protected Brev qualification jobs so newer runs cannot replace pending release evidence. Signed-off-by: Carlos Villela <cvillela@nvidia.com>
Signed-off-by: Carlos Villela <cvillela@nvidia.com> # Conflicts: # test/e2e/README.md
There was a problem hiding this comment.
🧹 Nitpick comments (1)
.github/workflows/e2e.yaml (1)
6141-6157: 🔒 Security & Privacy | 🔵 Trivial | 💤 Low valueConsider passing
needsviaenv:instead of direct template expansion.zizmor flags
const needs = ${{ toJSON(needs) }};as a template-injection pattern: the JSON text is spliced directly into the script source before execution rather than passed as a properly isolated value. Practical risk is low here (workflow_dispatch/schedule only, not PR-triggered), but the safer idiom isenv: NEEDS_JSON: ${{ toJSON(needs) }}thenJSON.parse(process.env.NEEDS_JSON)inside the script.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/e2e.yaml around lines 6141 - 6157, Avoid direct template expansion of needs in the script by passing toJSON(needs) through the step’s env configuration as NEEDS_JSON. In the script containing loadWorkflowRunJobs and runtime audit handling, parse process.env.NEEDS_JSON with JSON.parse and use the resulting value as needs.Source: Linters/SAST tools
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In @.github/workflows/e2e.yaml:
- Around line 6141-6157: Avoid direct template expansion of needs in the script
by passing toJSON(needs) through the step’s env configuration as NEEDS_JSON. In
the script containing loadWorkflowRunJobs and runtime audit handling, parse
process.env.NEEDS_JSON with JSON.parse and use the resulting value as needs.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 9e2efcff-30c5-4535-918e-f0a162b97fc0
📒 Files selected for processing (1)
.github/workflows/e2e.yaml
Signed-off-by: Carlos Villela <cvillela@nvidia.com>
There was a problem hiding this comment.
🧹 Nitpick comments (1)
test/hosted-runner-recovery-workflow.test.ts (1)
89-89: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick winAssert the rendered run-name behavior, not only the expression text.
toContain(...)can pass even if the correlation-id branch is unreachable or renders the wrong value. Add cases covering non-empty and emptycorrelation_idthrough the workflow’s public boundary. As per path instructions, tests should provide behavioral confidence rather than lock onto source-text fragments.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/hosted-runner-recovery-workflow.test.ts` at line 89, Replace the source-text assertion on E2E_RUN_NAME with behavioral tests through the workflow’s public boundary, covering both non-empty and empty correlation_id inputs. Verify that the rendered run name uses the correlation ID when present and the expected fallback when absent, without asserting expression fragments.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@test/hosted-runner-recovery-workflow.test.ts`:
- Line 89: Replace the source-text assertion on E2E_RUN_NAME with behavioral
tests through the workflow’s public boundary, covering both non-empty and empty
correlation_id inputs. Verify that the rendered run name uses the correlation ID
when present and the expected fallback when absent, without asserting expression
fragments.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 4ce041a7-10f1-4255-aba3-fcb1521e529e
📒 Files selected for processing (1)
test/hosted-runner-recovery-workflow.test.ts
Signed-off-by: Carlos Villela <cvillela@nvidia.com>
cv
left a comment
There was a problem hiding this comment.
Fresh exact-head review passed: product scope, nine-category security review, documentation receipt, all current CI, automated review, and selected release-qualification E2E evidence are green.
<!-- markdownlint-disable MD041 --> ## Summary Add the canonical `docs/changelog/2026-07-25.mdx` release entry with the exact `## v0.0.96` heading. The entry reconciles all 90 first-parent commits since v0.0.95 with all 92 merged PRs in the live `v0.0.96` label ledger and groups the user-visible changes by operator journey. ## Changes - Add the parser-safe dated MDX changelog entry for v0.0.96 with root-absolute links to the focused user guides. - Source summary: - [#7194](#7194) -> `docs/changelog/2026-07-25.mdx`: Document persistent baseline network policy exclusions and their inspection, rebuild, and snapshot behavior. - [#7188](#7188), [#7427](#7427), and [#7546](#7546) -> `docs/changelog/2026-07-25.mdx`: Document DNS-backed HTTPS inference routing, keyless loopback endpoints, and provider-marker isolation. - [#7238](#7238) -> `docs/changelog/2026-07-25.mdx`: Document blueprint sandbox and provider identifier validation before state writes or OpenShell calls, with bounded terminal-safe rejection previews. - [#7319](#7319), [#7274](#7274), [#7528](#7528), [#7353](#7353), and [#7560](#7560) -> `docs/changelog/2026-07-25.mdx`: Document the managed default gateway service, onboarding readiness, and container-runtime identity safeguards. - [#7349](#7349), [#7498](#7498), [#7406](#7406), [#7196](#7196), [#7559](#7559), [#7421](#7421), [#7510](#7510), [#7295](#7295), and [#7565](#7565) -> `docs/changelog/2026-07-25.mdx`: Document gateway-scoped status, lifecycle diagnostics, managed MCP recovery, delete-edge safeguards, and fail-closed CLI prompt and command output. - [#7591](#7591) -> `docs/changelog/2026-07-25.mdx`: Document opt-in authenticated MCP tool-name discovery, its bounded and names-only contract, probe interaction, and rebuild requirement. - [#7305](#7305), [#7480](#7480), [#7471](#7471), [#7365](#7365), and [#7541](#7541) -> `docs/changelog/2026-07-25.mdx`: Document installer version checks, version-tag reporting, license guidance, WSL Ollama selection, and DGX Station vLLM detection. - [#7482](#7482), [#7466](#7466), [#7208](#7208), [#7434](#7434), and [#7586](#7586) -> `docs/changelog/2026-07-25.mdx`: Document Ollama resource details, reasoning precedence, Hermes onboarding behavior, and preserved managed Hermes BuildKit failures. - [#6830](#6830), [#7492](#7492), [#7563](#7563), and [#7582](#7582) -> `docs/changelog/2026-07-25.mdx`: Document the authoritative OpenClaw production lock, fixed managed-image dependencies, immutable Hermes base adoption, and Hermes image-size reduction. - [#7505](#7505), [#7530](#7530), [#7547](#7547), [#7508](#7508), [#7548](#7548), [#7549](#7549), [#7537](#7537), [#7534](#7534), [#7515](#7515), [#7511](#7511), [#7551](#7551), [#7562](#7562), [#7575](#7575), [#7496](#7496), [#7594](#7594), [#7595](#7595), and [#7599](#7599) -> `docs/changelog/2026-07-25.mdx`: Summarize release validation, transient and bounded dispatch reconciliation, exact pre-tag qualification, identity revalidation, npm-audit retry, sharding, image reuse, timeout, telemetry, and workflow-hardening changes. - Reconciled without separate changelog prose: - [#7539](#7539), [#7526](#7526), [#7507](#7507), [#7506](#7506), [#7519](#7519), [#7516](#7516), [#7396](#7396), [#7254](#7254), [#7583](#7583), [#7596](#7596), and [#7598](#7598): Test-harness or fixture-only changes. - [#7403](#7403), [#7161](#7161), [#6877](#6877), [#7531](#7531), [#7525](#7525), [#7522](#7522), [#7536](#7536), [#7552](#7552), [#7566](#7566), [#7553](#7553), [#7561](#7561), [#7577](#7577), [#7569](#7569), [#7585](#7585), [#7584](#7584), [#7592](#7592), [#7580](#7580), [#7571](#7571), [#7517](#7517), [#7589](#7589), [#7402](#7402), [#7558](#7558), [#7544](#7544), and [#7601](#7601): Dependency, internal recovery, validation, contributor-workflow, E2E optimization, telemetry, or CI trust changes with no separate user-facing release claim. - [#7556](#7556), [#7573](#7573), [#7576](#7576), and [#7578](#7578): Experimental repository-maintainer conflict automation with no canonical user documentation surface. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [x] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [ ] Tests added or updated for changed behavior - [x] Existing tests cover changed behavior — justification: `test/changelog-docs.test.ts` validates dated changelog structure, version headings, and published links. - [ ] Tests not applicable — justification: - [x] Docs updated for user-facing behavior changes - [ ] Docs not applicable — justification: - [ ] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [ ] Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: ## Documentation Writer Review - [x] Documentation writer subagent reviewed the completed changes - Result: `docs-updated` - Evidence: Reviewed `docs/changelog/2026-07-25.mdx` at exact head `0f5dedb47` against 90 first-parent release commits and 92 merged PRs labeled `v0.0.96`. Verified parser-safe MDX SPDX, the exact version heading, literal CLI names, writing style, skip terms, all 20 root-absolute published links, and the accepted #7591 opt-in authenticated discovery bounds. #7544, #7599, and #7601 remain internal or CI-only release-ledger entries. Changelog tests passed 6/6, the docs build passed with 0 errors and two pre-existing Fern warnings, and `npm run check:diff` plus the final diff check passed. - Agent: Codex Desktop documentation-writer subagent <!-- docs-review-head-sha: 0f5dedb --> <!-- docs-review-agents-blob-sha: be20a09 --> ## DGX Station Hardware Evidence - [ ] Tested on DGX Station - Tested commit: - Station profile/scenario: - Result: - Supporting evidence: ## Verification - [x] PR description includes a `Signed-off-by:` line and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run check:diff` passed when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — `npx vitest run test/changelog-docs.test.ts`: 6/6 passed. - [ ] Applicable broad gate passed — `npm test` for broad runtime/test-harness changes; `npm run check` for repo-wide validation/coverage changes — command/result: Not applicable to this prose-only changelog entry. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) — the build passed with 0 errors and 2 existing Fern warnings; the published-route check passed. - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) — native changelog files use the required parser-safe MDX SPDX comment and no frontmatter. --- Signed-off-by: Carlos Villela <cvillela@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Persistent network policy exclusions with consistent restore/exclusion reporting across rebuilds/snapshots. * Opt-in MCP tool discovery via `mcp status --tools` with bounded, redacted authenticated traffic. * Improved HTTPS inference switching for custom endpoints and refreshed onboarding/model menu details. * Refined OpenShell gateway defaults for port `8080`, including more reliable readiness checks. * **Bug Fixes** * Prevent incorrect provider/model restoration after compatible-provider update failures. * Preserve managed MCP state after exec loss and tighten gateway/doctor status scoping. * **Tests** * Stronger, fail-closed release validation with hardened evidence/artifact handoff and bounded timeouts/retries. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Prekshi Vyas <prekshiv@nvidia.com> Co-authored-by: Prekshi Vyas <prekshiv@nvidia.com>
Summary
Pre-tag full E2E now runs the default-enabled suite and exact staging Brev Launchable qualification in one trusted workflow run. Release confirmation requires machine-verifiable evidence for the candidate SHA, qualification job, boot identity, and workspace cleanup. Full dispatches and protected qualification jobs can no longer supersede pending release evidence.
Related Issue
Fixes #7487
Changes
include_staging_brev_launchablewith ordinary, full, selective, scheduled, and fail-closed readiness boundaries.nemoclaw-maintainer-e2eas the agent entry point for trusted Actions dispatch and exact-candidate evidence collection.test/maintainer-e2e-skill.test.tsprotects the validator contract.github.run_idand queue protectedstaging-brev-launchablejobs withqueue: maxso newer runs cannot replace pending evidence.Type of Change
Quality Gates
173ac9b3fpassed all nine categories with no findings. The final delta is the verified fix(hermes): repin published sandbox base #7582 immutable Hermes digest repin; the original complete-diff review is recorded at feat(e2e): require exact pre-tag qualification #7496 (comment).Documentation Writer Review
docs-updated.agents/skills/nemoclaw-maintainer-e2e/SKILL.md, the cut-release/evening/release-train/skills-guide guidance, andtest/e2e/README.mdconsistently document non-superseding full dispatches,github.run_idisolation, queued Brev qualification, exact-SHA receipts, and cleanup evidence. The final inherited fix(hermes): repin published sandbox base #7582 Dockerfile digest repin needs no additional documentation.DGX Station Hardware Evidence
Verification
Signed-off-by:line and every published commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run check:diffpassed when hooks were skipped or unavailable2359c65dc5a017293d25b14abe1c2b1c7cbe19a4: E2E workflow boundary 46 passed; maintainer evidence/policy 28 passed; hosted-runner recovery workflow 6 passed; semantic E2E phase coverage passed for 125 tests across 82 files; focused Hermes immutable-pin/remediation contracts 45 passed; plugin type-check passed on the complete repair head.npm run checkpassed for the original complete change set atb62382ade; final head173ac9b3fpassednpm run check:diff, signed commit hooks, CLI type-check, and focused immutable-pin contracts. Fresh exact-head CI and selected E2E remain required before merge.npm run docsbuilds without warnings (not applicable: nodocs/orfern/files changed)docs/pages changed)docs/pages)Signed-off-by: J. Yaunches jmyaunch@gmail.com
Summary by CodeRabbit
New Features
Bug Fixes
Documentation