Skip to content

refactor(prover): remove legacy unused reachability constraints - #1

Closed
eviehoward wants to merge 41 commits into
mainfrom
2412-remove-legacy-reachability-constraints/eviehoward
Closed

eviehoward wants to merge 41 commits into
mainfrom
2412-remove-legacy-reachability-constraints/eviehoward

Conversation

@eviehoward

@eviehoward eviehoward commented Jul 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

Removes can_write_to_endpoint() and the solver variables and encoding
methods that exist solely to support it. The active query layer uses
can_exfil_via_endpoint() and direct credential lookups exclusively;
the removed Z3 encodings were scaffolding for query categories that were
never implemented or were later cut. Confirmed zero callers across the
monorepo; crate is not published externally.

Related Issue

Closes issue 2412

Changes

  • Removed can_write_to_endpoint() — zero callers confirmed across the monorepo
  • Removed binary_can_write, credential_has_write — only consumed by can_write_to_endpoint()
  • Removed credential_has_destructive, filesystem_readable — directly dead, previously silenced with #[allow(dead_code)]
  • Removed encode_credentials(), encode_filesystem() — exclusively populated the above fields
  • Removed two #[allow(dead_code)] allowances, HashMap::new() initialisers, and build() call sites

Note: removing encode_credentials() leaves write_actions_for_scopes()
and destructive_actions_for_scopes() in credentials.rs callerless.
Flagged for follow-up — out of scope for this change.

Testing

  • mise run pre-commit passes
  • Unit tests added/updated
  • E2E tests added/updated (if applicable)

cargo test -p openshell-prover passes. No new tests added — existing
tests (test_findings_for_github_policy, test_wildcard_endpoint_covering_credential_host_emits_credential_reach,
test_known_metadata_hostname_emits_link_local_finding, test_empty_policy_no_findings)
cover all four active finding categories end-to-end and would be
duplicated exactly by any new test.

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO)
  • Architecture docs updated (if applicable) — no architecture doc covers prover model internals

SDAChess and others added 30 commits July 16, 2026 18:09
…pendency (NVIDIA#2318)

Remove Z3 from the openshell CLI. Proving is handled by the gateway, so bundling the solver in the client duplicates functionality and complicates portable CLI builds and packaging.

Signed-off-by: Simon Scatton <sscatton@nvidia.com>
Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
…A#2303)

Previously, Docker was auto-detected when the CLI was installed or a candidate
Unix socket existed. Neither check verified that the Docker API was responsive.
A similar check was done when auto-detecting Podman in the past, but was
replaced in 1f07bf0 with a probe of candidate Podman sockets instead.

This change applies the functional API probing approach introduced for Podman
in 1f07bf0 to Docker. It also makes Docker driver initialization use the same
socket-selection mechanism as Docker auto-detection instead of Bollard’s local
defaults. This means the previously auto-detectable Docker socket paths
$HOME/.docker/run/docker.sock and $XDG_RUNTIME_DIR/docker.sock will actually be
usable.

When no working compute driver can be auto-detected, the gateway exits early
with a message saying as much:

> configuration error: no compute driver configured and auto-detection found no
> suitable driver; set --drivers or OPENSHELL_DRIVERS to kubernetes, podman,
> docker, or vm

This makes for a better user experience when installing OpenShell without an
available supported compute driver.

Signed-off-by: Kris Hicks <khicks@nvidia.com>
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 6.4.0 to 7.0.0.
- [Release notes](https://github.com/actions/setup-node/releases)
- [Commits](actions/setup-node@48b55a0...8207627)

---
updated-dependencies:
- dependency-name: actions/setup-node
  dependency-version: 7.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…IDIA#2288)

Bumps [softprops/action-gh-release](https://github.com/softprops/action-gh-release) from 3.0.1 to 3.0.2.
- [Release notes](https://github.com/softprops/action-gh-release/releases)
- [Changelog](https://github.com/softprops/action-gh-release/blob/master/CHANGELOG.md)
- [Commits](softprops/action-gh-release@718ea10...3d0d988)

---
updated-dependencies:
- dependency-name: softprops/action-gh-release
  dependency-version: 3.0.2
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…ies (NVIDIA#2287)

* feat(tui): navigate panels via Up/Down arrow overflow at list boundaries

When at the bottom of a panel's item list, pressing Down/j now moves
focus to the next panel instead of being a silent no-op. Likewise,
pressing Up/k at the top moves to the previous panel with the cursor
on its last item. Empty panels are skipped and the ring wraps around.

Closes NVIDIA#2273

Signed-off-by: Varsha Prasad Narsing <vnarsing@nvidia.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* fix(tui): guard Up handlers against stale cursor in empty panels

The Down handlers already check whether the list is non-empty before
incrementing the cursor, but the Up handlers only checked cursor > 0.
When a list becomes empty after a refresh with a nonzero cursor, Up
would decrement the stale cursor instead of overflowing to the
previous panel. Add the same non-empty guard to all four Up arms.

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

* docs(sandboxes): add dashboard keyboard navigation to manage-sandboxes

Describe Tab/Shift+Tab panel cycling, Up/Down and j/k boundary
overflow, and middle-pane tab switching in the OpenShell Terminal
section.

Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>

---------

Signed-off-by: Varsha Prasad Narsing <vnarsing@nvidia.com>
Signed-off-by: Varsha Prasad Narsing <varshaprasad96@gmail.com>
…le (NVIDIA#1782)

Add gateway-managed AWS STS credential refresh (provider-v2, NVIDIA#1576). The
gateway calls sts:AssumeRole and writes three short-lived credentials
(AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN) to the
provider record; the proxy re-signs requests with SigV4. Adds the aws and
aws-s3 provider profiles and a declarative multi-output refresh model
(additional_outputs) so one AssumeRole co-mints all three credentials.

Signed-off-by: Russell Bryant <rbryant@redhat.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
* test(e2e): run VM suite in CI

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(ci): configure KVM permissions directly

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(e2e): flush VM overlay before restart

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs: simplify VM test documentation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* test(e2e): include gateway resume in VM run

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: John Myers <johntmyers@users.noreply.github.com>
Co-authored-by: John Myers <johntmyers@users.noreply.github.com>
Signed-off-by: lr90 <qiuweimin@matrixorigin.cn>
* docs: bump stated Rust MSRV from 1.88 to 1.90

Cargo.toml sets rust-version = "1.90" (rust-toolchain.toml pins
1.95.0), so building with the previously documented 1.88 fails
Cargo's MSRV check.

* docs: bump e2e/rust MSRV to 1.90

* fix: align remaining Rust version fields to 1.90

examples/governance-interceptor/Cargo.toml still had rust-version
1.88. Also bump e2e/rust's prost dependency to 0.14 to match the
workspace, since it was on 0.13 in an otherwise standalone crate.
login-action and setup-buildx-action used a mutable version tag while
every other action in the repo is pinned to a commit SHA. Pin both,
and align login-action to the same v4 SHA already used in ci-image.yml.

Use the full resolved version in the trailing comment (v3.12.0) to
match the more common convention used elsewhere in .github/.
* ci(e2e): reuse prebuilt CLI artifacts

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): reuse prebuilt gateway artifacts

Signed-off-by: Evan Lezar <elezar@nvidia.com>

* ci(e2e): reuse prebuilt VM driver artifact

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: Evan Lezar <elezar@nvidia.com>
- README: fix github-sandbox tutorial link missing get-started segment
- README: replace dead community-sandboxes doc link with the actual repo
- README: match supported host list to support-matrix.mdx
- architecture/README: list the missing google-vertex-ai-provider doc
- SECURITY.md: fix a mis-indented list item
- standardize on NVIDIA/OpenShell-Community casing for repo links
* fix(gateway): honor tty flag for interactive exec

Pass the requested TTY mode through the interactive SSH relay.
Skip PTY allocation and resize forwarding when TTY is disabled,
and add regression coverage for both modes.

Signed-off-by: emonq <emonq@outlook.com>

* test(gateway): improve `test_sandbox_interactive_exec_honors_tty` to test streamed stdin and stdout/stderr

Signed-off-by: emonq <emonq@outlook.com>

---------

Signed-off-by: emonq <emonq@outlook.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
* chore(deps): bump actions/attest from 4.1.1 to 4.2.0

Bumps [actions/attest](https://github.com/actions/attest) from 4.1.1 to 4.2.0.
- [Release notes](https://github.com/actions/attest/releases)
- [Changelog](https://github.com/actions/attest/blob/main/RELEASE.md)
- [Commits](actions/attest@a1948c3...f7c74d2)

---
updated-dependencies:
- dependency-name: actions/attest
  dependency-version: 4.2.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* chore(deps): note actions attest release version

Signed-off-by: Evan Lezar <elezar@nvidia.com>

---------

Signed-off-by: dependabot[bot] <support@github.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Evan Lezar <elezar@nvidia.com>
…IDIA#2317)

* fix(providers): allow git clone/fetch via default GitHub provider

The github.com:443 git-transport endpoint used the read-only access
preset, which expands to GET/HEAD/OPTIONS only. Git smart HTTP requires
a POST to */git-upload-pack for clone and fetch, so the L7 proxy denied
those operations and `gh repo clone` / `git clone https://...` failed.

Replace the preset with explicit rules that permit the read-only methods
plus POST */git-upload-pack, so clone/fetch work while push
(git-receive-pack) stays blocked. Enabling push still requires an
explicit policy proposal.

Why allowing this POST is still read-only: in git's smart HTTP protocol
POST is an RPC transport, not a write. A clone/fetch does GET
*/info/refs (ref discovery) followed by POST */git-upload-pack, whose
body is only the client's want/have negotiation; the server responds
with a packfile and nothing on the server is modified (data flows
server -> client). The service names are from the server's perspective:
git-upload-pack = the server uploads a pack to the client (a read/
download), while git-receive-pack = the server receives a pack from the
client (the actual write/push). The new rule is scoped to
*/git-upload-pack only, so push (git-receive-pack) and arbitrary POSTs
to github.com remain denied.

Add a provider-profile regression test and a rego enforcement test
covering ref discovery, upload-pack (allowed), and receive-pack (denied).

Closes NVIDIA#1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): strengthen git-transport regression and add clone e2e

Pin the exact allowed rule set for the built-in github git-transport
endpoint in both the provider-profile and composed-policy tests, so a
broader or additional POST rule (e.g. POST **) that could enable push
via git-receive-pack fails the test instead of passing a substring
check. Add an e2e test that attaches the built-in github provider and
clones a public repo over HTTPS, exercising provider attachment,
effective-policy composition, TLS interception, and real git behavior.

Update the Providers V2 docs so the github.com git-transport endpoint
shows explicit clone/fetch rules instead of the stale read-only preset.

Refs NVIDIA#1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): isolate providers_v2 mutation in clone e2e

The clone e2e enables the gateway-global providers_v2_enabled setting.
Restore its exact prior value (or absence) captured via GetGatewayConfig
instead of unconditionally deleting it, and serialize the mutation
across xdist workers with an exclusive file lock on the run's shared
base temp dir, so a shared or pre-configured gateway is left untouched
and parallel workers cannot race the read-modify-restore.

Refs NVIDIA#1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

* test(providers): serialize providers_v2 mutation with a suite-wide guard

The clone e2e's per-fixture lock only coordinated fixtures that acquired
it; other xdist workers hit the same gateway without it and could
observe the transiently-enabled providers_v2_enabled global during their
own sandbox creation (CWE-362).

Add an autouse readers-writer guard in conftest: every test holds a
shared lock on the gateway config, and a test marked
exclusive_gateway_config holds an exclusive lock. Mark the clone test
exclusive so no other worker is mid-test while it enables and restores
the gateway-global setting. Exact prior-value restoration is retained.

Refs NVIDIA#1769

Signed-off-by: Russell Bryant <rbryant@redhat.com>

---------

Signed-off-by: Russell Bryant <rbryant@redhat.com>
NVIDIA#2307)

Closes NVIDIA#2112

The host `cargo zigbuild` for `*-unknown-linux-musl` opens ~333 `.rlib`
files at once during the static link, exceeding macOS's default soft
limit of 256 and failing with `ProcessFdQuotaExceeded`. This blocked the
docker/podman `mise run gateway` paths for macOS contributors; only the
VM driver path guarded against it.

Extract the VM path's `ensure_build_nofile_limit` guard into a shared
`tasks/scripts/build-env.sh` and call it from the host-staging chokepoint
(`stage-prebuilt-binaries.sh`), fixing docker, podman, and all
docker:*/multiarch host cross-compiles at once. De-duplicate the VM
script to source the shared helper. The guard is a no-op on Linux and
when cargo-zigbuild is absent, so CI and Linux dev are unaffected. The
limit is read from `OPENSHELL_BUILD_NOFILE_LIMIT` (default 8192),
honoring the legacy `OPENSHELL_VM_BUILD_NOFILE_LIMIT` for back-compat.

Also correct the stale comment in gateway-docker.sh (the cross-compile
runs on the host, not inside Linux containers) and document the guard in
architecture/build.md. Adds tasks/scripts/test-build-env.sh, wired into
`mise run test` via `test:build-env`.

Signed-off-by: Jim Meyer <jim@meyer4hire.com>
NVIDIA#2243)

* feat(workspace): implement workspace model (Phase 1 of RFC 0011)

Implements workspace and membership model providing hard isolation
boundaries for multi-player OpenShell deployments.

Workspace CRUD with Kubernetes-style Terminating phase for graceful
deletion. All resources scoped by workspace via ObjectMeta. Membership
RPCs for workspace access control. Persistence migration shifts name
uniqueness to (object_type, workspace, name). Provider profiles support
platform and workspace scoping. Service routing uses workspace-prefixed
DNS labels. Inference routes renamed and workspace-scoped with
DeleteInferenceRoute RPC. Python SDK with WorkspaceClient, two-method
list pattern (workspace-scoped and for_all_workspaces), and workspace
parameter on all methods. CLI workspace flags, TUI workspace cycling.
K8s driver filters unmanaged CRs and uses delete preconditions. Podman
driver uses immutable container IDs. Label serialization fixed across
all put_if call sites.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(cli): delegate sandbox upload command to existing upload function

The standalone `sandbox upload` command reimplemented upload logic
inline with two bugs: it used `Path::exists()` which follows symlinks
(rejecting dangling symlinks), and it ran git-aware filtering on
symlink sources. The `run::sandbox_upload()` function already handles
both cases correctly via `sandbox_upload_plan()`. Replace the inline
logic with a call to the existing function.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): shorten sandbox names and fix test compatibility

Shorten the sandbox name in initial_sparse_policy_is_acknowledged_as_loaded
from 'e2e-2159-sparse-enrich' (22 chars) to 'e2e-sparse-enrich' (17 chars)
to comply with MAX_ROUTABLE_NAME_LEN (19 chars).

Also capture stderr in create_keep_with_args so future sandbox creation
failures include the actual CLI error instead of reporting empty output.

Signed-off-by: Derek Carr <decarr@redhat.com>

* test(workspace): add test coverage for workspace CRUD and persistence isolation

Add unit tests for workspace create happy path, get round-trip, get
not-found, get empty-name rejection, already-exists error, and
resolve_workspace not-found. Add persistence test proving cross-workspace
name uniqueness (same name in different workspaces produces separate
records). Add workspace name max-length boundary tests. Fix e2e harness
to include stderr in name-parse-failure error path. Align Python e2e
test_workspace_crud with try/finally pattern. Document provider profile
catalog workspace scoping gap in RFC 0011.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(examples): update examples for workspace model compatibility

Shorten sandbox names in demo scripts to fit the 19-character
MAX_ROUTABLE_NAME_LEN limit: policy-demo prefix to pd-, multi-agent
notepad derives a short SANDBOX_TAG from the run ID, governance
interceptor uses gs-PID-RANDOM. Update vscode-remote-sandbox.md SSH
host aliases from openshell-{name} to openshell-{name}.{workspace}
format.

Signed-off-by: Derek Carr <decarr@redhat.com>

* feat(sdk): add workspace-scoped client and workspace CRUD

Add WorkspaceScopedClient modeled after kube::Api::namespaced — captures
workspace once and injects it into every sandbox request. Add workspace
CRUD methods (create, get, list, delete) and list_sandboxes_all_workspaces
on OpenShellClient. Extend SandboxRef with workspace field and add
WorkspaceRef type. Include mock tests for all new operations.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings in workspace test assertions

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(docs): convert indented code blocks to fenced in RFC 0011

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(lint): resolve clippy warnings and apply cargo fmt across workspace

Auto-format with cargo fmt and fix clippy warnings exposed by the
reformat: unnecessary qualifications, map_unwrap_or, identical match
arms, unused variable prefix, dead code annotations, and let-unit-value
in e2e harness.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): address workspace scoping issues from review

- Add workspace field to settings JSON output (CLI)
- Skip Podman containers missing workspace label instead of defaulting
  to empty string, matching K8s driver behavior
- Add resource_version to list_by_scope SELECT in both SQLite and
  Postgres backends, with regression test
- Gate PolicyLocalContext proposal/lookup routes on workspace readiness,
  returning 503 when workspace is not yet discovered
- Block sandbox and provider creation in TUI all-workspaces mode
- Clear workspace vectors in TUI reset_sandbox_state

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog workspace-aware

Thread workspace through snapshot_catalog so the
EffectiveProviderProfileCatalog enforces workspace boundaries on both
read and write paths. UserProviderProfileSource now loads platform-scoped
profiles (workspace "") plus the target workspace's profiles, preventing
cross-workspace duplicate profile ID collisions that previously caused
global catalog failures.

Update RFC 0011 to reflect catalog scoping is implemented in Phase 1
rather than deferred to future work.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(persistence): include workspace column in atomic policy revision INSERT

put_policy_revision_atomic omitted the workspace column from the INSERT
into the objects table in both SQLite and Postgres backends, causing
atomically-written policy revisions to lose their workspace association.
Add workspace field to AtomicPolicyRevisionWrite and thread it through
both backend INSERT statements, matching the non-atomic put_policy_revision
path which already included it.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(proxy): skip ancestor walk when socket owner is the entrypoint

collect_ancestor_identities walked the entire process tree above the
entrypoint when the connecting process was the entrypoint itself,
SHA256-hashing every ancestor binary (IDE, shell, container runtime).
On dev machines with large binaries in the ancestor chain this exceeded
the 30-second test timeout. When start_pid == stop_pid there are no
intermediate ancestors to verify, so return an empty list immediately.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): make provider profile catalog scope-aware

Allow the same profile ID at platform and workspace scopes by
introducing layered catalog entries where workspace profiles shadow
platform profiles. Add source and scope fields to the ProviderProfile
proto and CLI output. Migrate List/Get handlers to the catalog,
fixing divergence with runtime profile resolution.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align podman e2e labels with centralized driver constants

The podman driver moved its container labels to the centralized
openshell.ai/ prefix, but the e2e test harness and cleanup script
still referenced the old openshell.sandbox-* keys, causing the
local_driver_token_restart test to fail on container lookup.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(e2e): align python profile isolation test with scope-aware catalog

Platform profiles are now visible in workspace listings as fallbacks
per the layered catalog design. Update the assertion to match.

Signed-off-by: Derek Carr <decarr@redhat.com>

* fix(workspace): honor profile_workspace in runtime profile resolution

Runtime profile lookups now consult provider.profile_workspace via
get_type_profile_for_scope. Providers created with --global-profile
(profile_workspace="") resolve to the platform profile even when a
workspace profile shadows the same ID. All 6 runtime call sites
updated; type-only call sites remain scope-agnostic.

Signed-off-by: Derek Carr <decarr@redhat.com>

---------

Signed-off-by: Derek Carr <decarr@redhat.com>
Closes NVIDIA#2218

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
* perf(build): share sccache across worktrees

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(build): make sccache directory overridable

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

* fix(build): support pinned mise version

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>

---------

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
Bumps [actions/checkout](https://github.com/actions/checkout) from 7.0.0 to 7.0.1.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](actions/checkout@9c091bb...3d3c42e)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: 7.0.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
moezdil and others added 11 commits July 21, 2026 14:01
)

* fix(driver-podman): resolve Podman socket via auto-detection

Align Podman socket selection with the existing Docker model: explicit
config wins, otherwise probe openshell-core for a responsive socket.
This also fixes the original HOME-unset panic, since resolution no
longer hardcodes a per-OS default path.

- add detect_podman_socket() in openshell-core, mirroring
  detect_docker_socket
- PodmanComputeConfig.socket_path is now Option<PathBuf>, no default
- remove default_socket_path() (podman driver) and podman_socket_path()
  (vm driver), both replaced by the shared detector
- update server env override and CLI for the new Option type
- add tests: responsive-candidate detection in openshell-core, and
  config-error (not panic) when no socket is configured or reachable

* fix(driver-podman): address review feedback on socket resolution

Extract socket resolution into resolve_socket_path, taking the
detector as a parameter so tests do not depend on real env vars or
the host's actual Podman state. Replace the flaky env-mutating test
with three deterministic cases: explicit wins, detected is used when
absent, and neither source errors.

Fix the socket_path doc comment to describe it from a config user's
point of view, matching DockerComputeConfig's docstring. Drop a
comment that only made sense next to the Docker driver code.

Update the Podman README and gateway docs: they described a fixed
per-OS default path that no longer exists, replace with the actual
probe-then-fail behavior.

* docs(driver-podman): simplify socket default description

Previous wording was self-contradictory (says auto-detect on unset,
then lists the same var as a probed candidate) and omitted the Linux
/run/user/uid/podman/podman.sock candidate.
…2236)

The branches API endpoint cannot handle ref names with slashes (e.g.
pull-request/2223). When the API returns a 404, gh api writes the error
JSON to stdout before exiting non-zero, so the || fallback never fires.
This causes mirror_sha to contain raw JSON like {"message":"Branch not
found",...} and the comment displays `{"messa` as the SHA.

Switch to the git/ref/heads/ endpoint which handles slashes natively,
and add a regex guard to reset mirror_sha to empty when it does not
look like a valid 40-char hex SHA.

Signed-off-by: Roland Huß <rhuss@redhat.com>
* docs(agents): refresh CLI and debugging skills

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(agents): keep project skills synchronized

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

* docs(agents): cover middleware workflows

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>

---------

Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
)

* fix(proxy): include OPA deny reason in CONNECT 403 response

When a CONNECT request was denied by OPA policy, the 403 response
used a generic "not permitted by policy" message for both "endpoint
not in policy" and "endpoint matched but binary didn't match." Users
had no way to distinguish the two without reading supervisor logs.

The OPA policy already computes a detailed deny_reason (e.g.,
"binary '/usr/bin/node' not allowed in policy 'X'") but the proxy
was not including it in the HTTP response.

Now the CONNECT deny response includes a "reason" field with the
OPA deny reason when available. When the reason is empty, the field
is omitted for backward compatibility.

Fixes NVIDIA#2355

Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>

* docs(observability): document optional reason field in CONNECT 403 response

The proxy now includes a reason field in the JSON body of denied
CONNECT responses when the policy engine provides a specific denial
cause. Update the Proxy Error Responses section to show the field
and describe when it is present vs omitted.

Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>

---------

Signed-off-by: Adel Zaalouk <azaalouk@redhat.com>
…A#2391)

The scripts/bin/openshell wrapper hardcoded the binary path to
$PROJECT_ROOT/target/debug/openshell. When CARGO_TARGET_DIR is set
(e.g. via .bashrc or mise), cargo places the binary elsewhere and the
wrapper fails with 'No such file or directory'.

Use ${CARGO_TARGET_DIR:-$PROJECT_ROOT/target} so the wrapper finds
the binary regardless of where the build artifacts live.

Signed-off-by: Jesse Jaggars <jjaggars@jjaggars-kubevirt.rht.csb>
Signed-off-by: Jesse Jaggars <jjaggars@redhat.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
…IA#2399)

Signed-off-by: Matthew Grossman <mgrossman@nvidia.com>
* docs(extensibility): add gateway interceptor guide

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): list interceptable routes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): stabilize source links

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* docs(extensibility): link canonical interceptor routes

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Drew Newberry <anewberry@nvidia.com>
… NemoClaw redirect (NVIDIA#2405)

The OpenClaw community sandbox was removed from NVIDIA/OpenShell-Community
in PR NVIDIA#73 (May 16, 2026). The Docker Compose tutorial and docker-compose.yml
comment still referenced the stale --from openclaw command and GHCR image.
Replace the broken OpenClaw tab with a redirect to the NemoClaw Quickstart,
which is the supported path per docs/about/supported-agents.mdx. Remove the
stale pre-pull command for the removed image.
Fixes NVIDIA#2404

Signed-off-by: Matias Schimuneck <schimuneck.matias@gmail.com>
Signed-off-by: Evie Howard <evhoward@redhat.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Remove can_write_to_endpoint() and the solver variables and encoding
methods that exist solely to support it. No active query calls this
method; confirmed zero callers across the monorepo. Crate is not
published externally, so no public API concern.

Removed as root cause:
- can_write_to_endpoint(): zero callers in the codebase

Removed as direct consumers of can_write_to_endpoint():
- binary_can_write: only read inside can_write_to_endpoint()
- credential_has_write: only read inside can_write_to_endpoint()

Removed as directly dead (no query ever read them):
- credential_has_destructive: flagged by compiler, silenced with
  #[allow(dead_code)] rather than removed
- filesystem_readable: flagged by compiler, silenced with
  #[allow(dead_code)] rather than removed

Removed as now-empty encoding methods:
- encode_credentials(): exclusively populated credential_has_write
  and credential_has_destructive
- encode_filesystem(): exclusively populated filesystem_readable

Also removed two #[allow(dead_code)] allowances and the corresponding
HashMap::new() initialisers and build() call sites.

Active query behavior is unchanged. All existing prover tests pass
without modification and cover all four finding categories end-to-end.
No new tests added as they would duplicate existing coverage exactly.

Note: removing encode_credentials() leaves write_actions_for_scopes()
and destructive_actions_for_scopes() in credentials.rs with no callers.
Left as a follow-up outside the scope of this change.

Closes NVIDIA#2412

Signed-off-by: Evie Howard <evhoward@redhat.com>
@eviehoward eviehoward closed this Jul 22, 2026
eviehoward pushed a commit that referenced this pull request Aug 6, 2026
…VIDIA#2271)

* feat(sdk/go): add Go SDK foundation, types, and sandbox client (A)

Add the Go SDK module with the full API contract and a working sandbox
client as the first vertical slice. All other resource clients are present
as stubs returning Unimplemented errors, to be replaced with real
implementations in subsequent PRs.

Contents:
- Module setup (go.mod, Makefile, mise.toml)
- All domain types (types/ package)
- Full ClientInterface with all sub-client accessors
- Shared infrastructure (errors, auth, gRPC connection, logging)
- Sandbox client with converter and tests (fully functional)
- Stub clients for remaining resources (exec, file, health, provider,
  profile, config, refresh, policy, service, ssh, tcp)

Part of the Go SDK decomposition plan (NVIDIA#2270).
Implements NVIDIA#2044.

* fix(sdk/go): address review feedback on PR NVIDIA#2271

- Make scheme parsing drive transport selection: http:// uses plaintext
  gRPC, https:// or no scheme uses TLS. Add regression tests.
- Add Resources and DriverConfig fields to SandboxTemplate and update
  both converter directions (SandboxFromProto/SandboxSpecToProto).
- Regenerate proto bindings from current canonical proto sources to
  eliminate drift (SigV4/MCP fields, params matchers, reserved fields).
- Run gofmt/goimports on all handwritten Go files.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* fix(sdk/go): address principal engineer review findings

- Remove dead boolCount function that would fail golangci-lint (#1)
- Emit EventAdded for the first watch event instead of EventModified,
  matching k8s watch semantics (#7)
- Add mutex locking to all mock server methods that access the shared
  sandboxes map, fixing latent race conditions (NVIDIA#12)
- Skip HealthCheck integration test that calls an unimplemented stub (NVIDIA#13)
- Scope doc.go examples: mark sections for sub-clients not yet available
  in this PR with "available in a future release" (#4)
- Document Config.Timeout/RetryPolicy/Logger and WatchOptions fields
  as reserved for future use (#2, #6)

Signed-off-by: Roland Huß <rhuss@redhat.com>

* refactor(sdk/go): migrate mise config to centralized task include

Move Go SDK mise configuration from standalone sdk/go/mise.toml into
the project's centralized pattern:

- Add Go tools (go, golangci-lint, protoc-gen-go, protoc-gen-go-grpc)
  to root mise.toml [tools] section
- Create tasks/go.toml with all SDK tasks using go: namespace prefix
  and dir=sdk/go for working directory
- Update sdk/go/Makefile to reference namespaced task names
- Update proto:sync default path for monorepo layout

Addresses review feedback from drew on PR NVIDIA#2271 regarding mise
convention alignment.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* refactor(sdk/go): remove UPSTREAM_VERSION standalone repo artifact

Remove sdk/go/proto/UPSTREAM_VERSION file and its exclusion from
proto:check. This was a leftover from the standalone repo prototype.
In a monorepo, proto drift is detectable via git diff between
sdk/go/proto/ and proto/ directly.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* refactor(sdk/go): switch proto generation from protoc to buf

Replace raw protoc invocations with buf for Go SDK proto code generation,
aligning with the TS SDK approach (PR NVIDIA#2122).
- Add repo-level buf.yaml declaring proto/ as the buf module with lint
  and breaking change detection config
- Add sdk/go/buf.gen.yaml configuring buf to generate Go code directly
  from root proto/ (no more vendored .proto copies)
- Delete vendored .proto source files from sdk/go/proto/
- Rewrite go:proto:gen and go:proto:check mise tasks to use buf
- Remove go:proto:sync and go:proto:clean tasks (no longer needed)
- Add proto target to sdk/go/Makefile
- Add buf 1.72.0 to root mise.toml tool dependencies
- Include options.proto in generation (was stripped from vendored copies)
- Regenerate all .pb.go files via the new buf pipeline
Signed-off-by: Roland Huß <rhuss@redhat.com>

* test(sdk/go): add proto-converter field coverage detection

Use protobuf reflection to enumerate all fields on key proto messages
(SandboxSpec, SandboxTemplate, SandboxStatus, SandboxCondition,
SandboxPolicy) and compare against explicit handled/skipped sets in the
converter tests.

Unhandled fields produce warnings (t.Log), not failures, so proto
contributors are not forced to fix SDK converters in the same PR. Stale
entries in the handled set (removed proto fields) do fail, since they
indicate the converter references something that no longer exists.

A follow-up CI workflow will create GitHub issues when converter drift
lands on main.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* fix(sdk/go): bump Go to 1.26 and fix errcheck lint violations

The upstream go.mod now has `toolchain go1.26.4`, which requires Go 1.26
to build golangci-lint. Bump the mise.toml Go version from 1.25 to 1.26
and wrap deferred Close() calls in test helpers to satisfy errcheck.

Assisted-By: 🤖 Claude Code

* feat(sdk/go): add ObjectMeta fields (annotations, workspace, deletion_timestamp)

Add three new proto ObjectMeta fields to Sandbox and Provider domain
types: Annotations (map), Workspace (string), and DeletionTimestamp
(*time.Time). Update converters in both directions, deep-copy maps at
the proto/SDK boundary, and add TimeFromMillisPtr/MillisFromTimePtr
helper functions.

Assisted-By: 🤖 Claude Code

* chore(sdk/go): regenerate proto bindings after rebase

Pick up workspace fields from upstream PR NVIDIA#2445 (Wire authorization
into workspace model). All request messages now include workspace
parameter in the generated Go bindings.

Assisted-By: 🤖 Claude Code

* feat(sdk/go): add workspace scoping to all RPC interfaces

Add workspace parameter to every sandbox-scoped RPC method across all
interfaces (Sandbox, Exec, File, Service, SSH, TCP, Config, Policy,
Provider, Profile, Refresh). The workspace string is passed as the
second parameter after ctx, following the convention workspace then
resource-name.

Key changes:
- SandboxInterface: all 10 methods gain workspace parameter
- sandbox_client.go: passes Workspace field in every proto request
- ListOptions: add AllWorkspaces field for cross-workspace queries
- All stub interfaces updated to match new signatures
- All sandbox client tests updated with "default" workspace

Assisted-By: 🤖 Claude Code

* chore(sdk/go): remove coverage.out from tracking

Assisted-By: 🤖 Claude Code

* fix(sdk/go): address review feedback from mrunalp

- Add RefreshStrategyAWSStsAssumeRole to match proto enum value 6,
  fulfilling the "all domain types upfront" contract
- Wrap context.DeadlineExceeded and context.Canceled in StatusError
  so IsDeadlineExceeded() and IsCancelled() helpers work correctly
- Return error from mapToStruct/SandboxSpecToProto instead of silently
  discarding structpb.NewStruct failures on invalid template maps

Signed-off-by: Roland Huss <rhuss@redhat.com>

* fix(sdk/go): address remaining review items

- Wire go:ci into root ci task so SDK is tested in repository CI
- Fix gofmt formatting on converter files
- Add goimports to mise.toml tools
- Add coverage.out to .gitignore
- Add Go SDK section to AGENTS.md and CONTRIBUTING.md
- Add regression tests for context-error wrapping (IsDeadlineExceeded,
  IsCancelled) and invalid template map rejection
- Remove panic from SandboxToProto, return error instead

Signed-off-by: Roland Huss <rhuss@redhat.com>

* fix(sdk/go): pin goimports version and update lockfile

Pin goimports to 0.48.0 instead of "latest" and regenerate mise.lock
to include the new entry.

Signed-off-by: Roland Huss <rhuss@redhat.com>

* fix(sdk/go): TLS.Insecure means skip-verify, not plaintext

Align TLS.Insecure semantics with the Rust SDK: Insecure: true now
uses TLS with InsecureSkipVerify (skip cert verification) instead of
switching to plaintext. Only the http:// scheme triggers plaintext.

This fixes token auth against dev/k3d gateways: StaticToken and
RefreshableToken require transport security, which real TLS (even
with InsecureSkipVerify) satisfies, but plaintext does not.

For http:// + token auth (dev gateways without TLS), wrap the auth
provider to override RequireTransportSecurity, matching the Rust
SDK's behavior where http:// accepts any auth mode.

Transport decision table (matches Rust SDK crates/openshell-sdk):
  http://  + any TLS config  -> plaintext (TLS config ignored)
  https:// + Insecure: true  -> TLS, skip cert verify
  https:// + Insecure: false -> TLS, full verification
  no scheme                  -> same as https://

Signed-off-by: Roland Huss <rhuss@redhat.com>

* feat(sdk/go): add missing policy proto fields

Add 6 previously silently dropped fields to the network policy types
and converters, preventing security-relevant data loss on round-trip:

NetworkEndpoint fields 19-23:
- CredentialSigning: SigV4 re-signing mode
- SigningService: AWS service name for SigV4
- SigningRegion: AWS region override for SigV4
- JsonRpcMaxBodyBytes: JSON-RPC body inspection limit
- Mcp: MCP-specific policy options (new McpOptions type)

L7Allow and L7DenyRule field 9:
- Params: MCP params matcher map for tools/call filtering

New type McpOptions with StrictToolNames and AllowAllKnownMcpMethods
optional booleans matching the proto definitions.

Signed-off-by: Roland Huss <rhuss@redhat.com>

* fix(sdk/go): enforce coverage test and extend to policy messages

Change coverage_test.go from t.Logf (silent) to t.Errorf so that
unhandled proto fields fail the test immediately. Add coverage tests
for NetworkEndpoint (23 fields), L7Allow (8 fields), L7DenyRule
(8 fields), and McpOptions (2 fields).

Any new proto field that is not in the handled set or explicitly
skipped now breaks the build, closing the silent-drift gap.

Signed-off-by: Roland Huss <rhuss@redhat.com>

* ci(sdk/go): add Go SDK job to branch-checks workflow

Add a Go SDK job to branch-checks.yml that runs mise run go:ci
(lint, build, test, proto-check, docs-check) on every PR. This
ensures the SDK is tested in CI, not just locally.

Signed-off-by: Roland Huss <rhuss@redhat.com>

* fix(sdk/go): address should-fix review items

#6 Fix broken godoc examples: add workspace parameter to all method
   calls in doc.go that were broken after workspace scoping.

#7 Add Err field to Event[T]: Watch error events now carry the
   underlying error instead of discarding it.

#8 Separate Unauthenticated from PermissionDenied: add
   ErrorUnauthenticated code and IsUnauthenticated() helper. gRPC
   Unauthenticated (401) now maps to its own code instead of
   collapsing into PermissionDenied (403).

NVIDIA#9 Add Unwrap to StatusError: replace dead Details field with Cause
   error field. StatusError.Unwrap() returns Cause, enabling
   errors.Is/As unwrapping. FromGRPCError and contextError both
   populate Cause.

Signed-off-by: Roland Huss <rhuss@redhat.com>

* ci(sdk/go): add go:format:check to CI pipeline

Add gofmt format verification to go:ci. Catches unformatted Go files
before they reach the PR. Fix formatting on coverage_test.go.

Signed-off-by: Roland Huss <rhuss@redhat.com>

* chore(sdk/go): remove Makefile in favor of mise tasks

All build, lint, test, and proto-gen tasks are already defined in
tasks/go.toml and invoked via mise. The Makefile was a leftover
that duplicated this and raised questions in review.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* feat(sdk/go): sync proto bindings and add credential handle support

Regenerate Go proto bindings after rebase to pick up new
CredentialHandle message and Provider.credential_handles and
profile_workspace fields from upstream. Add domain types, converter
support, and proto field coverage tests for Provider and
CredentialHandle.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* fix(sdk/go): reject plaintext auth leak and fix watch error handling

Reject http:// addresses when the auth provider requires transport
security instead of silently stripping the requirement. Remove the
insecureAuthWrapper that overrode RequireTransportSecurity.

Fix watch stream error handling: use blocking send for terminal
errors so they are never silently dropped when the channel is full,
and wrap mid-stream errors with converter.FromGRPCError so SDK error
helpers like IsUnavailable work on watch Event.Err.

Signed-off-by: Roland Huß <rhuss@redhat.com>

* fix(sdk/go): address review findings from multi-agent code review

- WaitReady now detects SandboxDeleting phase and returns immediately
  instead of polling indefinitely
- Watch goroutine defers streamCancel() to prevent context leaks
- Fix StopOnTerminal=false test to keep stream open (was wrong-reason
  pass due to stream ending, not StopOnTerminal logic)
- Add EventDeleted test covering the Deleting phase branch
- Add provider converter unit tests for CredentialHandle round-trip,
  nil handling, and empty maps

Signed-off-by: Roland Huß <rhuss@redhat.com>

---------

Signed-off-by: Roland Huß <rhuss@redhat.com>
Signed-off-by: Roland Huss <rhuss@redhat.com>
eviehoward pushed a commit that referenced this pull request Sep 11, 2026
* feat(mxc): ETW->OCSF audit consumer + Windows OCSF JSONL parity (cp6 P1)

Add a Windows MXC ETW->OCSF audit trail in openshell-driver-mxc: a real-time
Sandboxing-provider ETW consumer that decodes events (TDH), attributes each to
an OpenShell sandbox_id, and maps them to OCSF (lifecycle 6002, config 5019,
process 1007, finding 2004).

cp6 Phase 1 - durable OCSF JSONL audit-file parity with Linux:
- openshell-ocsf: add emit_ocsf_event_routed (populates the event-bridge
  thread-local AND stamps sandbox_id+message in one dispatch) plus public
  set/clear_current_event; OS-aware device (Device::windows/for_current_os) so
  device.os.name reflects the host instead of a hardcoded Linux stub.
- etw_consumer: emit via the routed emit (previously fired a bare info! that
  never populated the bridge, so the structured event was dropped).
- openshell-server: install OcsfJsonlLayer over a synchronous daily-rotated
  appender (durable under force-kill), gated by OPENSHELL_OCSF_JSON, path via
  %PROGRAMDATA%\OpenShell\logs (override OPENSHELL_OCSF_LOG_DIR).
- device.hostname now resolves to the real gateway machine name.

Box-proven on 7F203-MXC-001: JSONL lines == shorthand OCSF rows, all valid
OCSF JSON, per-sandbox attribution intact, disabled state writes nothing.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc): map remaining Sandboxing ETW events to OCSF

Close the last three ETW->OCSF gaps so the audit trail covers the full
set of events the Sandboxing provider emits (12/12):

- ProcessLaunched -> Process Activity [1007] "Launch" (confirmed start;
  carries the real processId/threadId, the twin of CreateProcessInSandbox
  which only has the request + command line).
- SandboxProxyConfigured -> Device Config State Change [5019] (the one
  network-plane setup event; surfaces proxyPort, "no proxy" when 0).
- SandboxConsoleReferencePlumbed -> Device Config State Change [5019]
  (console-handle plumbing).

map_config_state now handles the full config/hardening/setup family and
carries proxyPort/hasConsoleReference/creationFlags as unmapped fields.
Verified on 7F203-MXC-001: 11/12 event types emit OCSF without a proxy
(SandboxProxyConfigured requires proxy config to fire).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): seed ETW attribution under registry lock + Device tests

Address CodeRabbit review on !31:

- Prevent stale ETW attribution on a delete/launch race: register the
  wxc-exec pid while holding the registry lock, and bail if the sandbox
  entry is already gone. Previously the attribution key could be seeded
  after `delete` had removed the sandbox, leaving a stale key that could
  misroute later Sandboxing ETW events to a dead sandbox_id. Lock order
  (registry -> attribution) matches the delete path, so no deadlock.
- Add unit tests for the new Device::windows and Device::for_current_os
  constructors to harden Windows/Linux OCSF device parity.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): buffer+replay racing events and harden attribution keys

Addresses two ETW->OCSF attribution review items (Shailendra #1, #2).

#2 early-event loss: ETW delivers the sandbox create/config burst the instant wxc-exec starts, which can beat the driver's register_launch (now under the registry lock post-Ready). process_event previously dropped anything unresolved, losing the racing burst. Add a bounded, time-bounded pending buffer (PENDING_MAX=4096, PENDING_TTL=5s): unresolved events are held and replayed once attribution lands, aged-out ones dropped. Consumer switched to a timed recv_timeout(200ms) so the buffer is re-driven after each event and on a tick. Emit path factored into shared emit_resolved().

#1 attribution collisions: a Windows PID is recycled after exit and a command line is commonly identical across sandboxes. register_launch now rebinds by_pid on reuse and clears the stale last_pid_sid hint (warns if the PID still pointed at a different, leaked sandbox); command line is held in by_cmd only while unique and demoted to a new ambiguous_cmds set on a second owner, so a duplicate command refuses to resolve rather than misroute.

Unit tests: buffer replay (direct + cross-link), buffer bound, PID-reuse rebind, duplicate-cmd non-resolution. Box-verified on 7F203-MXC-001 (5 sandboxes, identical cmd -> 5 isolated sandbox_ids, 50/50 OCSF/JSONL, BuffersLost=0).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* docs(mxc-etw): note cmd_line is captured raw with no privacy filtering

Review item #3 (Shailendra): add a PRIVACY NOTE on map_process_launch stating cmd_line is copied verbatim into OCSF process.cmd_line with no redaction, so secrets/PII on a command line land unredacted in the durable audit trail (deliberate audit-fidelity trade-off; treat the log as sensitive). Redaction is owned by an upstream privacy layer, not this path; no general audit-output PII scrubber exists today (openshell_core::secrets [CREDENTIAL] redaction is scoped to the proxy HTTP-target logging, a separate egress path).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): open ETW trace on caller thread so start_session reports real status

Review item #4 (Shailendra): start_session previously returned Ok(EtwSession) as soon as the pump thread was spawned, but OpenTraceW ran later inside that thread; if it failed we still handed back a live-looking session and logged 'consumer started' (silent failure = false audit coverage).

Split the two Win32 calls instead of adding a channel handshake (avoids any lost-wakeup/hang risk): the quick, synchronous OpenTraceW now runs on the caller thread (open_trace), and only the blocking ProcessTrace runs on the pump thread (run_trace). start_session returns Err if OpenTraceW fails (reclaiming the boxed Sender so the consumer disconnects, stopping the session, joining the consumer) and returns Ok/logs 'started' only once capture is genuinely open. Opened handle + LoggerName buffer + boxed Sender are carried to the pump via a Send OpenedTrace so they outlive ProcessTrace.

Box-verified on 7F203-MXC-001: consumer started=True, failed-to-start=False, 50 OCSF rows / 50 JSONL, BuffersLost=0 (no regression to capture/emit).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): guard pending-event replay against PID recycling

CodeRabbit flagged that drain_resolved() re-resolved buffered events
against the live by_pid map, so if Windows recycled a wxc-exec PID within
PENDING_TTL a stale event from the dead sandbox could be emitted under the
new owner.

Stamp each by_pid registration with its Instant and add resolve_replay(),
used only on the buffered/replay path. It (a) never falls back to the
recycle-/ambiguity-prone by_cmd or last_pid_sid keys, and (b) trusts a PID
match only when the registration is not newer than the buffered event by
more than REPLAY_PID_GRACE (2s) - a recycled PID's registration lands well
outside that window, so the stale event ages out instead of misattributing.
The legitimate #2 seed race (registration lands ~immediately) still replays.

Adds unit tests for the recycle-refusal, in-grace acceptance, and
weak-fallback exclusion.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc-etw): surface unexpected ProcessTrace termination (review #4)

start_session already returns Err on OpenTraceW failure (runs on the
caller thread since e41a770), closing the first half of Shailendra's #4.
This closes the second half: ProcessTrace's result was discarded, so if
capture died mid-run the backend had no way to know.

Add a shared CaptureHealth (stopped/stopping/exit_code) between the pump
thread and EtwSession. run_trace now records ProcessTrace's WIN32_ERROR
and, when the pump returns without a deliberate stop, logs at ERROR that
MXC OCSF capture is no longer running. EtwSession::stop() sets `stopping`
before teardown so a normal shutdown isn't misreported, and
EtwSession::is_capture_alive() exposes the state for status/diagnostics.

Box-verified on 7F203-MXC-001: 5 sandboxes, 50 attributed OCSF rows,
JSONL parity 50/50, BuffersLost=0, clean start/stop (no false failure).

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): add ETW->OCSF audit-trail example kit; fix proxy-configured message

Add a runnable OCSF audit-trail example under examples/ (run-ocsf-audit.ps1,
mxc-ocsf-audit.toml, ocsf-audit.yaml, README) that spins up sandboxes with the
in-process ETW consumer and egress proxy on, emitting a full OCSF JSONL audit
trail across all four classes (6002/5019/1007/2004).

Fix SandboxProxyConfigured mapping to log "MXC sandbox proxy configured" instead
of a misleading "(no proxy)" when the provider reports proxyPort=0; the event's
presence already indicates proxy configuration. Verified on-box: 26 events, all
mapped ETW event types present.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(mxc-ocsf): clearer audit report + client-safe run-ocsf-audit.ps1

Improve the ETW to OCSF audit-trail example output and make it safe to ship.

Report:
- Add an event-type coverage count ("N of M expected event types fired");
  the denominator auto-adjusts (8 with proxy on, 7 with -NoProxy).
- Split the checklist into expected event types vs anomaly findings
  (ActivityError/FallbackError), which are reported separately and not
  counted toward coverage (a clean run may emit none).
- Verdict is now coverage-based (all expected types must fire) instead of
  the looser "at least 3 OCSF classes".
- Call out the absolute path to the durable OCSF JSONL log prominently.

Client-safety:
- Default -ShareOut to empty (no auto-copy); pass -ShareOut a UNC path to
  opt in. Removes a hardcoded internal share path from a published example.
- Drop internal-team wording ("Hand that zip back for evaluation", "BUNDLE:")
  in favor of neutral "Results bundle:".
- Update README-ocsf-audit.txt to match the opt-in -ShareOut behavior.

Verified on both MXC boxes: 7F203-MXC-001 (base-container) -> PASS, 8 of 8
event types, 26 OCSF events across 4 classes; 7F203-MXC-003 (AppContainer
fallback) -> reduced set as expected, clean output.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): configure OCSF audit workloads per sandbox

- remove unsupported gateway-scoped workload fields from the shipped MXC audit example.
- build the command, working directory, and filesystem grant from each run's ShareDir
- pass the workload through --driver-config-json.
- preserve the host CONNECT proxy configuration and conditional audit coverage for the future host_connect_proxy merge
- require the workload output when determining the audit verdict.

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(mxc): omit command arguments from OCSF audit logs

- record only the executable basename for MXC CreateProcessInSandbox audit events
- leave process.cmd_line unset so workload arguments cannot reach shorthand or JSONL logs
- cover tokens, passwords, signed URLs, and PII with a secret-leak regression test
- update the audit example, architecture guidance, and published logging documentation
- preserve ETW attribution and future host_connect_proxy enforcement behavior

Signed-off-by: Akber Raza <akberr@nvidia.com>

* feat(etw): enhance ETW session management with distinct naming for concurrent gateways

* fix(etw): bound the audit queue during overload

- replace the unbounded ETW callback channel with count- and byte-bounded buffering

- keep the ETW callback non-blocking and count records rejected during overload

- emit immediate, rate-limited warnings that identify resulting audit coverage gaps

- make the audit example fail when queue overload causes dropped ETW records

- cover stalled consumers, oversized events, and warning throttling with unit tests

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): harden sandbox audit attribution

- remove command-line and persistent per-PID fallback keys from live and replay resolution

- retire the driver-owned wxc-exec PID before publishing child completion

- retain established identity, activity, and correlation-vector links only for the five-second late-event window

- prevent buffered records from crossing rapid PID retirement and reuse boundaries

- add resolver and lifecycle coverage and document the attribution trust boundary

Signed-off-by: Akber Raza <akberr@nvidia.com>

* chore(mxc): address rebase follow-ups

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): align OCSF audit example with driver config

- remove unsupported egress proxy settings

- stop requiring the unavailable proxy audit event

- update example documentation for supported event coverage

Signed-off-by: Akber Raza <akberr@nvidia.com>

* fix(etw): redact command-line secrets in DecodedEtwEvent summary

* fix(etw): enhance PID resolution and event attribution logic for ETW records

* fix(ocsf): restrict gateway-local JSONL sink to Windows/MXC path with opt-in configuration

* address rebase issues

* fix(mxc): address ETW audit review feedback

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): fail closed across ambiguous PID reuse

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

* fix(mxc): bind ETW attribution to process generation

Signed-off-by: Drew Newberry <anewberry@nvidia.com>

---------

Signed-off-by: Akber Raza <akberr@nvidia.com>
Signed-off-by: Drew Newberry <anewberry@nvidia.com>
Co-authored-by: Jamie King <jamiek@nvidia.com>
Co-authored-by: Drew Newberry <anewberry@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.