diff --git a/.agents/skills/build-openshell-mxc-windows/SKILL.md b/.agents/skills/build-openshell-mxc-windows/SKILL.md index 21b49f1a5d..56a86fcd37 100644 --- a/.agents/skills/build-openshell-mxc-windows/SKILL.md +++ b/.agents/skills/build-openshell-mxc-windows/SKILL.md @@ -31,7 +31,6 @@ The Windows build lane is implemented by these tracked files: | `tasks/rust.toml`, `tasks/test.toml`, and `tasks/markdown.toml` | Windows routing for compiler-bearing checks, explicit Unix-only test skips, and Markdown dependency setup. | | `tasks/scripts/windows-msvc.ps1` | PowerShell wrapper that enters the Visual Studio developer environment and invokes Cargo. | | `.github/workflows/windows-msvc.yml` | Opt-in PR lint and test plus advisory `windows` branch cache seeding and dependent binary builds on native x64 and ARM64 runners. | -| `architecture/windows-msvc-build.md` | Design notes and validation contract. | | `.agents/skills/build-openshell-mxc-windows/` | This skill and companion reference material. | Use the code that is already in the repo. Do not generate a parallel Windows diff --git a/.agents/skills/build-openshell-mxc-windows/reference.md b/.agents/skills/build-openshell-mxc-windows/reference.md index 5ee638e1a9..81aecb9ec1 100644 --- a/.agents/skills/build-openshell-mxc-windows/reference.md +++ b/.agents/skills/build-openshell-mxc-windows/reference.md @@ -10,7 +10,6 @@ maintaining the existing build-only Windows MSVC lane. | `tasks/windows.toml` | Mise task definitions for `windows:*`. | | `tasks/scripts/windows-msvc.ps1` | Visual Studio environment discovery, rustup target setup, Cargo invocation, logs, artifact report. | | `.github/workflows/windows-msvc.yml` | Opt-in PR lint and test plus `windows` branch cache seeding and dependent binary builds on native x64 and ARM64 runners. | -| `architecture/windows-msvc-build.md` | Human-readable design contract. | ## Commands diff --git a/.agents/skills/create-rfc/SKILL.md b/.agents/skills/create-rfc/SKILL.md index 6df87ef315..6e8cffd428 100644 --- a/.agents/skills/create-rfc/SKILL.md +++ b/.agents/skills/create-rfc/SKILL.md @@ -33,12 +33,12 @@ Keep the template as the source of truth for section guidance. ## Writing Standards - Prefer concrete design statements over placeholder language. -- Link to relevant issues, prior RFCs, and architecture docs when they provide +- Link to relevant issues, prior RFCs, and published docs when they provide needed context. - Keep rejected or left-out designs in Alternatives, not Proposal. - Use Mermaid diagrams for architecture or data flow when a diagram would make the proposal easier to review. -- Do not update `architecture/` or published docs just because an RFC was +- Do not update published docs just because an RFC was drafted. Those updates belong with implementation or with an accepted RFC when the user asks for them. diff --git a/.agents/skills/create-spike/SKILL.md b/.agents/skills/create-spike/SKILL.md index 7e24cc8115..5c6d16c2ec 100644 --- a/.agents/skills/create-spike/SKILL.md +++ b/.agents/skills/create-spike/SKILL.md @@ -91,7 +91,7 @@ The prompt to the reviewer **must** instruct it to: 8. **Look at relevant tests to understand test coverage expectations.** What test patterns exist? What level of coverage is expected for this area? -9. **Check architecture docs** in the `architecture/` directory for relevant documentation about the affected subsystems. +9. **Check design records** in `rfc/` and the affected crate `README.md` files for relevant decisions and constraints. 10. **Assess gateway config documentation impact.** If the change would add, remove, rename, or change defaults for gateway TOML keys or driver-specific config options, call out that `docs/how-it-works/gateways/configuration.mdx` must be updated. If the change is surfaced through Helm or compute-driver setup docs, call out the relevant deployment or compute-driver docs too. @@ -156,7 +156,7 @@ gh issue create \ ### Architecture Overview - + ### Code References @@ -278,7 +278,7 @@ User says: "Allow sandbox egress to private IP space via networking policy" - Reads OPA policy evaluation pipeline in `opa.rs` and `crates/openshell-sandbox/data/sandbox-policy.rego` - Reads proto definitions in `sandbox.proto` for `NetworkEndpoint` - Maps the 4-layer defense model: netns, seccomp, OPA, SSRF check - - Reads `architecture/security-policy.md` and `architecture/sandbox.md` + - Reads RFC 0002 and the `openshell-policy` crate README - Identifies exact insertion points: policy field addition, SSRF check bypass path, OPA rule extension - Assesses: Medium complexity, High confidence, ~6 files 3. Fetch labels — select `area:sandbox`, `area:proxy`, `area:policy`, `state:validated` diff --git a/.agents/skills/sync-agent-infra/SKILL.md b/.agents/skills/sync-agent-infra/SKILL.md index f2800afca7..68a623c392 100644 --- a/.agents/skills/sync-agent-infra/SKILL.md +++ b/.agents/skills/sync-agent-infra/SKILL.md @@ -150,8 +150,8 @@ For each file in the table above, check for the following inconsistencies: 3. **Unique names** — Parse the `name` field from every `SKILL.md` under both roots. Every name must be globally unique and match the documented inventory. 4. **Local references** — Every relative Markdown link and referenced file in a skill must resolve within that installed skill directory unless the reference is an explicit published URL. 5. **Canonical paths** — Contributor skills that name the source location of a public skill must use `skills//...`, never `.agents/skills//...`. -6. **Public portability** — Public skills must not require repository-relative files under `docs/`, `architecture/`, `crates/`, `deploy/`, or `.agents/`; source builds; `mise`; or repository E2E workflows. Use installed `openshell --help` for command syntax and Markdown endpoints under `https://docs.nvidia.com/openshell/latest/` (URLs ending in `.md`) for product documentation. -7. **No canonical documentation copies** — Review public reference files and large command/schema blocks. Remove material that merely copies CLI help, policy schemas, architecture docs, or published operational documentation; retain only skill-specific reasoning and worked interactions. +6. **Public portability** — Public skills must not require repository-relative files under `docs/`, `crates/`, `deploy/`, or `.agents/`; source builds; `mise`; or repository E2E workflows. Use installed `openshell --help` for command syntax and Markdown endpoints under `https://docs.nvidia.com/openshell/latest/` (URLs ending in `.md`) for product documentation. +7. **No canonical documentation copies** — Review public reference files and large command/schema blocks. Remove material that merely copies CLI help, policy schemas, RFCs, or published operational documentation; retain only skill-specific reasoning and worked interactions. 8. **Discovery** — Run `npx -y skills add . --list` from a clean checkout or disposable copy. It must list exactly the four public skills. Remove any generated lock file or installed directory after the check. ## Step 3: Report Drift diff --git a/.agents/skills/tui-development/SKILL.md b/.agents/skills/tui-development/SKILL.md index 055884e1ed..832bb75144 100644 --- a/.agents/skills/tui-development/SKILL.md +++ b/.agents/skills/tui-development/SKILL.md @@ -102,7 +102,7 @@ match app.screen { } ``` -Within the `Sandbox` screen, the top 20% renders sandbox metadata (`sandbox_detail`), and the bottom 80% dispatches based on focus and tab state: +Within the `Sandbox` screen, `sandbox_detail::required_height` sizes the metadata and restart status pane to its contents. The remaining area dispatches based on focus and tab state: ```rust match app.focus { @@ -262,7 +262,7 @@ The `Theme` struct has 16 `Style` fields, accessed at runtime via `app.theme`: | `border` | EVERGLADE fg | Light sage fg | Unfocused panel borders | | `border_focused` | NVIDIA_GREEN fg | NVIDIA_GREEN_DARK fg | Focused panel borders | | `status_ok` | NVIDIA_GREEN fg | NVIDIA_GREEN_DARK fg | Healthy, INFO, Ready | -| `status_warn` | Yellow fg | Dark yellow fg | Degraded, WARN, Provisioning | +| `status_warn` | Yellow fg | Dark yellow fg | Degraded, WARN, Provisioning, Starting | | `status_err` | Red fg | Dark red fg | Unhealthy, ERROR | | `key_hint` | NVIDIA_GREEN fg | NVIDIA_GREEN_DARK fg | Keyboard shortcut labels | | `log_cursor` | EVERGLADE bg | Light green bg | Selected log line highlight | @@ -297,7 +297,7 @@ fn draw_detail_popup(frame: &mut Frame<'_>, data: &MyData, area: Rect, theme: &T - **Selected row**: Green `▌` left-border marker on the selected row. Active gateway also gets a green `●` dot. - **Focused panel**: Border changes from `border` to `border_focused` style. -- **Status indicators**: Green for healthy/ready/info, yellow for degraded/provisioning/warn, red for unhealthy/error. +- **Status indicators**: Green for healthy/ready/info, yellow for degraded/provisioning/starting/warn, red for unhealthy/error. - **Separators**: Muted `│` characters between title bar segments and nav bar sections. - **Log source labels**: `"sandbox"` source renders in `accent` (green), `"gateway"` in `muted`. @@ -423,7 +423,7 @@ All actions are accessible via keyboard shortcuts displayed in the nav bar. The | `crates/openshell-tui/src/ui/providers.rs` | Provider list table with profile-aware columns: Name, Category, Type, Credentials, Workspace | | `crates/openshell-tui/src/ui/global_settings.rs` | Global settings table: Key, Type, Value. Includes edit overlay, confirm-set, and confirm-delete popups | | `crates/openshell-tui/src/ui/sandboxes.rs` | Reusable sandbox table widget with columns: Name, Status, Created, Age, Image, Workspace, Notes | -| `crates/openshell-tui/src/ui/sandbox_detail.rs` | Sandbox metadata view — name, status, image, created, age, providers, policy version | +| `crates/openshell-tui/src/ui/sandbox_detail.rs` | Sandbox metadata view — name, status, image, created, age, restart policy/status, providers, policy version | | `crates/openshell-tui/src/ui/sandbox_policy.rs` | Policy viewer — rendered policy lines with scroll support, tab title | | `crates/openshell-tui/src/ui/sandbox_settings.rs` | Sandbox settings table: Key, Type, Value, Scope. Includes edit overlay and confirm popups | | `crates/openshell-tui/src/ui/sandbox_logs.rs` | Structured log viewer — timestamp, source, level, target, message, key=value fields, scroll position, source filter, visual selection mode, clipboard copy | diff --git a/.claude/agent-memory/arch-doc-writer/MEMORY.md b/.claude/agent-memory/arch-doc-writer/MEMORY.md deleted file mode 100644 index 67844988a0..0000000000 --- a/.claude/agent-memory/arch-doc-writer/MEMORY.md +++ /dev/null @@ -1,184 +0,0 @@ -# Arch Doc Writer Memory - -## Project Structure -- Crates: `openshell-cli`, `openshell-server`, `openshell-sandbox`, `openshell-bootstrap`, `openshell-core`, `openshell-providers`, `openshell-router`, `openshell-policy` -- CLI entry: `crates/openshell-cli/src/main.rs` (clap parser + dispatch) -- CLI logic: `crates/openshell-cli/src/run.rs` (all command implementations) -- Sandbox entry: `crates/openshell-sandbox/src/lib.rs` (`run_sandbox()`) -- OPA engine: `crates/openshell-sandbox/src/opa.rs` (single file, not a directory) -- Identity cache: `crates/openshell-sandbox/src/identity.rs` (SHA256 TOFU, uses Mutex NOT DashMap) -- L7 inspection: `crates/openshell-sandbox/src/l7/` (mod.rs, tls.rs, relay.rs, rest.rs, provider.rs, inference.rs) -- Proxy: `crates/openshell-sandbox/src/proxy.rs` -- Policy crate: `crates/openshell-policy/src/lib.rs` (YAML<->proto conversion, validation, restrictive default) -- Server multiplex: `crates/openshell-server/src/multiplex.rs` -- SSH sessions: `crates/openshell-server/src/ssh_sessions.rs` (session persistence, reaper) -- Sandbox SSH server: `crates/openshell-sandbox/src/ssh.rs` -- Providers: `crates/openshell-providers/src/providers/` (per-provider modules) -- Bootstrap: `crates/openshell-bootstrap/src/lib.rs` (cluster lifecycle) -- Proto files: `proto/` directory (openshell.proto, sandbox.proto, datamodel.proto, inference.proto) - -## Architecture Docs -- Files renamed from numbered prefix format to descriptive names (e.g., `2 - server-architecture.md` -> `gateway-architecture.md`) -- Current files: README.md, sandbox-providers.md, cluster-single-node.md, build-containers.md, sandbox-connect.md, sandbox.md, security-policy.md, gateway.md, gateway-security.md, sandbox-custom-containers.md, inference-routing.md -- Cross-references use plain filenames: `[text](gateway.md)` -- Naming convention: "gateway" in prose for the control plane component; code identifiers like `openshell-server` stay unchanged - -## Key Patterns -- OPA baked-in rules: `include_str!("../data/sandbox-policy.rego")` in opa.rs -- Policy loading: gRPC mode (OPENSHELL_SANDBOX_ID + OPENSHELL_ENDPOINT) or file mode (--policy-rules + --policy-data) -- Env vars: sandbox uses OPENSHELL_* prefix (e.g., OPENSHELL_SANDBOX_ID, OPENSHELL_ENDPOINT, OPENSHELL_POLICY_RULES) -- CLI flag: `--gateway-endpoint` (direct URL to gateway); resolution: --gateway-endpoint > --gateway > OPENSHELL_GATEWAY env > active_gateway file -- Provider env injection: both entrypoint process (tokio Command) and SSH shell (std Command) -- Cluster bootstrap: `sandbox_create_with_bootstrap()` auto-deploys when no cluster exists (main.rs ~line 632) -- CLI cluster resolution: --cluster flag > OPENSHELL_CLUSTER env > active cluster file - -## Bootstrap Crate Details -- `docker.rs`: `ensure_container()` sets ~12 env vars (REGISTRY_*, IMAGE_*, PUSH_IMAGE_REFS, etc.) -- `runtime.rs`: Polling params: health 180x2s, mTLS 90x2s -- `metadata.rs`: Metadata at `gateways/{name}/metadata.json` (nested), mTLS at `gateways/{name}/mtls/` (nested) -- `push.rs`: Uses `ctr` (not `k3s ctr`) with k3s containerd socket, `k8s.io` namespace -- IMPORTANT: `ClusterHandle::destroy()` does NOT remove metadata; only CLI `cluster_admin_destroy()` in run.rs does -- `ensure_image()`: Local-only refs (no `/`) get error with build instructions, not a Docker Hub pull attempt -- Dockerfile.cluster: k3s v1.29.8-k3s1 base, manifests in `/opt/openshell/manifests/` (volume mount overwrites `/var/lib/`) -- Healthcheck: checks k8s readyz, StatefulSet ready, Gateway Programmed, conditionally mTLS secret - -## Server Crate Details -- Two gRPC services: OpenShell (grpc.rs) and Inference (inference.rs), multiplexed via GrpcRouter by URI path -- Gateway is control-plane only for inference: SetClusterInference + GetClusterInference + GetInferenceBundle -- GetInferenceBundle: resolves managed route from provider record at request time, returns ResolvedRoute list + revision hash + generated_at_ms -- SetClusterInference: takes provider_name + model_id, stores only references (endpoint/key/protocols resolved at bundle time) -- Persistence: single `objects` table, protobuf payloads, Store enum dispatches SQLite vs Postgres by URL prefix -- Persistence CRUD: upsert ON CONFLICT (id) not (object_type, id); list ORDER BY created_at_ms ASC, name ASC (not id!) -- --db-url has no code default; Helm values.yaml sets `sqlite:/var/openshell/openshell.db` -- Object types: "sandbox", "provider", "ssh_session", "inference_route", "service_endpoint", "provider_profile" -- each implements ObjectType/ObjectId/ObjectName -- Config: `openshell_core::Config` in `crates/openshell-core/src/config.rs`, all flags have env var fallbacks -- SSH transport: CLI opens ForwardTcp gRPC stream (gated by CreateSshSession short-lived token), gateway relays via DuplexStream to supervisor RelayStream, supervisor connects to sandbox russh server over root-only Unix socket (/run/openshell/ssh.sock); channel uses mTLS when https:// endpoint configured, plaintext when http:// (Podman driver does not yet inject mTLS client materials). NSSH1 appears only in openshell-ocsf examples/tests, not on any live code path. -- Phase derivation: transient reasons (ReconcilerError, DependenciesNotReady) -> Provisioning; all others -> Error -- Broadcast bus buffer sizes: SandboxWatchBus=128, TracingLogBus=1024, PlatformEventBus=1024 -- Sandbox CRD: `agents.x-k8s.io/v1alpha1/Sandbox`, labels: `openshell.ai/sandbox-id`, `openshell.ai/managed-by` -- Proto files also include: `proto/inference.proto` (openshell.inference.v1) - -## Container/Build Details -- Four runtime images: sandbox (5 stages), gateway (2 stages), cluster (k3s base), pki-job (Alpine) -- Two build-only images: python-wheels (Linux multi-arch), python-wheels-macos (osxcross cross-compile) -- CI image: Dockerfile.ci (Ubuntu 24.04, pre-installs docker/buildx/aws/kubectl/helm/mise/uv/sccache/socat) -- Cross-compilation: `deploy/docker/cross-build.sh` shared by sandbox + gateway Dockerfiles -- Sandbox image has coding-agents stage: Claude CLI (native installer), OpenCode, Codex (npm) -- Helm chart deploys a StatefulSet (NOT Deployment), PVC 1Gi at /var/openshell -- Cluster image does NOT bundle image tarballs -- components pulled at runtime from distribution registry -- PKI job generates CA + server cert + client cert for mTLS (RSA 2048, 10yr, Helm pre-install hook) -- Build tasks in `tasks/*.toml`; scripts in `tasks/scripts/` -- `cluster-deploy-fast.sh` supports both auto mode (git diff) and explicit targets (gateway/sandbox/chart/all) -- `cluster-bootstrap.sh` ensures local Docker registry on port 5000, pushes all components, then deploys -- Default values.yaml: repository is CloudFront-backed CDN, tag: "latest", pullPolicy: Always -- Envoy Gateway version: v1.5.8 (set in mise.toml) -- DNS solution in cluster-entrypoint.sh: iptables DNAT proxy (NOT host-gateway resolv.conf) - -## Sandbox Connect Details -- CLI SSH module: `crates/openshell-cli/src/ssh.rs` (sandbox_connect, sandbox_connect_editor, sandbox_forward, sandbox_exec, sandbox_sync_up_files, sandbox_sync_up, sandbox_sync_down, sandbox_ssh_proxy, sandbox_ssh_proxy_by_name) -- Re-exported from run.rs: `pub use crate::ssh::{...}` for backward compat -- ssh-proxy subcommand: `Commands::SshProxy` in main.rs (~line 139) -- Gateway loopback resolution: `resolve_ssh_gateway()` in `crates/openshell-core/src/forward.rs:439` -- overrides loopback with cluster endpoint host; imported by ssh.rs and tui -- ExecSandbox gRPC: uses single-use TCP proxy + russh client in `grpc/sandbox.rs` (handle_exec_sandbox -> stream_exec_over_relay); operates over a relay DuplexStream through the supervisor session, not a direct TCP connection -- PTY I/O: 3 std::threads (writer, reader, exit) with reader_done sync for SSH protocol ordering -- SSH daemon: russh server, ephemeral Ed25519 key, pre_exec: setsid -> TIOCSCTTY -> setns -> drop_privileges -> harden_child_process -> sandbox::linux::enforce(prepared) [Linux] / sandbox::apply [non-Linux]; sandbox::linux::prepare() runs before fork - -## Policy Reload Details -- Poll loop: `run_policy_poll_loop()` in lib.rs, spawned after child process, gRPC mode only -- `OpaEngine::reload_from_proto()`: reuses `from_proto()` pipeline, atomically swaps inner engine, LKG on failure -- `CachedOpenShellClient` in grpc_client.rs: persistent mTLS channel for poll + status report (mirrors CachedInferenceClient) -- Dynamic domains: network_policies only (inference removed from policy). Static domains: filesystem, landlock, process (pre_exec, immutable) -- Server-side: `UpdateSandboxPolicy` RPC rejects changes to static fields or network mode changes -- Server-side validation: `validate_static_fields_unchanged()` + `validate_network_mode_unchanged()` in grpc.rs -- Poll interval: `OPENSHELL_POLICY_POLL_INTERVAL_SECS` env var (default 30), no CLI flag -- Version tracking: monotonic i64 per sandbox, `GetSandboxPolicyResponse` has version + policy_hash -- Version 1 backfill: lazy on first `GetSandboxPolicy` from spec.policy if no policy_revisions row exists -- `supersede_pending_policies()`: marks older pending revisions as superseded when new version persisted -- Status reporting: `ReportPolicyStatus` RPC with `PolicyStatus` enum (PENDING, LOADED, FAILED, SUPERSEDED) -- `report_policy_status()` updates `sandbox.current_policy_version` on LOADED, notifies watch bus -- Proto files: `ReportPolicyStatusRequest`/`Response` in openshell.proto, `GetSandboxPolicyResponse` in sandbox.proto -- `Sandbox.current_policy_version` (uint32) in datamodel.proto -- tracks active loaded version -- Persistence: `PolicyRecord` in persistence/mod.rs (id, sandbox_id, version, policy_payload, policy_hash, status, load_error, timestamps) -- CLI: `PolicyCommands` enum in main.rs (~line 516): Set, Get, List subcommands -- CLI: `sandbox_policy_set()` in run.rs (~line 2901): loads YAML, calls UpdateSandboxPolicy, optionally polls for status -- CLI: `sandbox_policy_get()` in run.rs (~line 3015): supports --rev N (version=0 means latest) and --full (YAML output via policy_to_yaml) -- CLI: `sandbox_logs()` in run.rs (~line 3124): --source (all/gateway/sandbox) and --level (error/warn/info/debug/trace) filters -- Deterministic hashing: `deterministic_policy_hash()` in grpc.rs (~line 1222): sorts network_policies by key, hashes fields individually, NO inference field -- Idempotent UpdateSandboxPolicy: compares hash of new policy to latest stored hash, returns existing version if match -- `policy_to_yaml()` in run.rs: converts proto to YAML via openshell_policy::serialize_sandbox_policy (moved to openshell-policy crate) -- `policy_record_to_revision()` in grpc.rs (~line 1334): `include_policy` param controls whether full proto is included -- Server-side log filtering: `source_matches()` + `level_matches()` in grpc.rs, applied in both get_sandbox_logs and watch_sandbox -- Standalone `proxy_inference()` was removed; inference handled in-sandbox by openshell-router -- Provider types: claude, codex, opencode, generic, openai, anthropic, nvidia, gitlab, github, outlook - -## Policy System Details -- YAML data file top-level keys: filesystem_policy, landlock, process, network_policies (NO inference key -- removed) -- Proto SandboxPolicy fields: version, filesystem, landlock, process, network_policies (NO inference field) -- Proto message field `filesystem` maps to YAML key `filesystem_policy` (different names!) -- IMPORTANT: Sandbox always runs in Proxy mode. NetworkMode::Block exists as enum variant but is NEVER set. -- Both file mode and gRPC mode set NetworkMode::Proxy unconditionally (see load_policy() in lib.rs and TryFrom in policy.rs) -- Reason: proxy always needed so all egress is evaluated by OPA -- OPA two-action model: Allow, Deny (NetworkAction in opa.rs). InspectForInference was REMOVED. -- Rego network_action rule: "allow" or "deny" only (no "inspect_for_inference") -- Behavioral trigger: endpoint `protocol` field -> L7 inspection; absent -> L4 raw copy_bidirectional -- Behavioral trigger: `tls: terminate` -> MITM TLS with ephemeral CA; requires `protocol` to also be set -- Behavioral trigger: `enforcement: enforce` -> deny at proxy; `audit` (default) -> log + forward -- Access presets: read-only (GET/HEAD/OPTIONS), read-write (+POST/PUT/PATCH), full (*/*) -- Validation: rules+access mutual exclusion, protocol requires rules/access, sql+enforce blocked, empty rules rejected -- YAML policy parsing moved to openshell-policy crate (parse_sandbox_policy, serialize_sandbox_policy) -- PolicyFile uses deny_unknown_fields for strict YAML parsing -- restrictive_default_policy() in openshell-policy: no network policies, sandbox user, best_effort landlock -- CONTAINER_POLICY_PATH: /etc/openshell/policy.yaml (well-known path for container-shipped policy) -- clear_process_identity(): clears run_as_user/run_as_group for custom images -- Policy safety validation: validate_sandbox_policy() checks root identity, path traversal, relative paths, overly broad paths, max 256 paths, max 4096 chars -- Identity binding: /proc/net/tcp -> inode -> PID -> /proc/PID/exe + ancestors + cmdline, SHA256 TOFU cache -- Network namespace: 10.200.0.1 (host/proxy) <-> 10.200.0.2 (sandbox), port 3128 default -- Enforcement order in pre_exec: setns -> drop_privileges -> landlock -> seccomp -- TLS cert cache: 256 entries max, overflow clears entire map -- CA files: /etc/openshell-tls/openshell-ca.pem (standalone) + ca-bundle.pem (system CAs + sandbox CA) -- Trust env vars: NODE_EXTRA_CA_CERTS, SSL_CERT_FILE, REQUESTS_CA_BUNDLE, CURL_CA_BUNDLE - -## Proxy SSRF Protection -- `is_internal_ip()` and `resolve_and_reject_internal()` in proxy.rs -- Blocks: 127/8, 10/8, 172.16/12, 192.168/16, 169.254/16, ::1, fe80::/10, IPv4-mapped IPv6 -- Runs after OPA allow, before TcpStream::connect -- Control plane endpoints exempt (they connect via hostname, skip SSRF check) -- DNS failure also rejects the connection -- Non-CP connections use pre-resolved addrs: `TcpStream::connect(addrs.as_slice())` - -## Inference Routing Details -- Sandbox-local execution via openshell-router crate -- InferenceContext in proxy.rs: Router + patterns + `Arc>>` route cache -- Route sources: `--inference-routes` YAML file (standalone) > cluster bundle via gRPC; empty routes gracefully disable -- Cluster bundle refreshed every ROUTE_REFRESH_INTERVAL_SECS (30s) -- Patterns: POST /v1/chat/completions, /v1/completions, /v1/responses, /v1/messages; GET /v1/models, /v1/models/* -- Managed inference CONNECT traffic was intercepted before OPA evaluation in the proxy -- InferenceProviderProfile in openshell-core/src/inference.rs: centralized provider metadata -- proxy.rs: only managed inference CONNECT traffic was handled; non-CONNECT requests received 403 for all hosts -- Buffer: INITIAL_INFERENCE_BUF=64KiB, MAX_INFERENCE_BUF=10MiB; grows by doubling -- Dev sandbox: `mise run sandbox -e VAR_NAME` forwards host env vars; NVIDIA_API_KEY always passed - -## Log Streaming Details -- LogPushLayer: `crates/openshell-sandbox/src/log_push.rs` -- tracing layer + spawn_log_push_task() -- Initialized in main.rs before run_sandbox(), gRPC mode only -- mpsc channel: 1024 lines (bounded), try_send (best-effort, never blocks) -- Background task: batches up to 50 lines, flushes every 500ms via PushSandboxLogs client-streaming RPC -- Secondary channel to gRPC call: mpsc::channel::(32) wrapped in ReceiverStream -- CachedOpenShellClient.raw_client() returns clone of inner OpenShellClient for direct RPC calls -- OPENSHELL_LOG_PUSH_LEVEL env var (default INFO), parsed in LogPushLayer::new() -- Server handler: push_sandbox_logs in grpc.rs, caps 100 lines/batch, forces source="sandbox" + sandbox_id -- TracingLogBus.publish_external(): injects into same broadcast + tail buffer as SandboxLogLayer -- Tail buffer: DEFAULT_TAIL = 2000 lines per sandbox (was 200, increased with log push) -- SandboxLogLayer (server tracing layer): sets source="gateway", only publishes events with sandbox_id field -- CLI: --source (gateway/sandbox/all), --level (error/warn/info/debug/trace) on `sandbox logs` -- Source filter: "all" normalized to empty list (no filter) in run.rs -- level_matches(): numeric ranking ERROR=0..TRACE=4, unknown levels always pass -- Create-watch filter: log_sources: ["gateway"] to prevent sandbox logs from blocking stop_on_terminal -- Proto: SandboxLogLine.source (string), SandboxLogLine.fields (map) -- Proto: PushSandboxLogsRequest/Response, GetSandboxLogsRequest (sources, min_level fields) -- Proto: WatchSandboxRequest (log_sources, log_min_level fields) - -## Naming Conventions -- The project name "OpenShell" appears in code but docs should use generic terms per user preference -- CLI binary: `openshell` (aliased as `nav` in dev via mise) -- Provider types: claude, codex, opencode, generic, openai, anthropic, nvidia, gitlab, github, outlook (see ProviderRegistry::new()) diff --git a/.claude/agent-memory/principal-engineer-reviewer/MEMORY.md b/.claude/agent-memory/principal-engineer-reviewer/MEMORY.md index 505250c1d8..67cc1f18ee 100644 --- a/.claude/agent-memory/principal-engineer-reviewer/MEMORY.md +++ b/.claude/agent-memory/principal-engineer-reviewer/MEMORY.md @@ -11,7 +11,7 @@ - Sandbox gRPC client: `crates/openshell-sandbox/src/grpc_client.rs` - CLI commands: `crates/openshell-cli/src/main.rs` (clap defs), `crates/openshell-cli/src/run.rs` (impl) - Python SDK: `python/openshell/` -- Plans go in: `architecture/plans/` +- Plans go in: `plans/` ## Key Patterns - TracingLogBus: per-sandbox broadcast::channel(1024) + VecDeque tail buffer (200 lines) @@ -23,6 +23,6 @@ - Build: `mise run sandbox` for sandbox infra ## Review Preferences (observed) -- Plans stored as markdown in architecture/plans/ +- Plans stored as markdown in plans/ - Conventional commits required - No AI attribution in commits diff --git a/.claude/agents/arch-doc-writer.md b/.claude/agents/arch-doc-writer.md deleted file mode 100644 index 90fbd707ea..0000000000 --- a/.claude/agents/arch-doc-writer.md +++ /dev/null @@ -1,206 +0,0 @@ ---- -name: arch-doc-writer -description: "Use this agent when documentation in the `architecture/` directory needs to be updated or created for a specific file after implementing a feature, fix, refactor, or behavior change. Launch one instance of this agent per file that needs updating. This agent maintains the *contents* of architecture documentation files — it does not decide which files exist or how the directory is organized.\\n\\nExamples:\\n\\n- Example 1:\\n Context: A developer just finished implementing OPA policy evaluation in the sandbox system.\\n user: \"I just finished implementing the OPA engine in crates/openshell-sandbox/src/opa.rs. Update architecture/sandbox.md to reflect the new policy evaluation flow.\"\\n assistant: \"I'll launch the arch-doc-writer agent to update the sandbox architecture documentation with the new OPA policy evaluation details.\"\\n \\n\\n- Example 2:\\n Context: A refactor changed how the HTTP CONNECT proxy handles allowlists.\\n user: \"The proxy allowlist logic was refactored. Please update architecture/proxy.md.\"\\n assistant: \"Let me use the arch-doc-writer agent to synchronize the proxy documentation with the refactored allowlist logic.\"\\n \\n\\n- Example 3:\\n Context: After implementing a new CLI command, the assistant proactively updates docs.\\n user: \"Add a --rego-policy flag to the CLI.\"\\n assistant: \"Here is the implementation of the --rego-policy flag.\"\\n \\n assistant: \"Now let me launch the arch-doc-writer agent to update the CLI architecture documentation with the new flag.\"\\n \\n\\n- Example 4:\\n Context: A user wants high-level overview documentation for a non-engineering audience.\\n user: \"Update architecture/overview.md with a non-engineer-friendly explanation of the sandbox system.\"\\n assistant: \"I'll launch the arch-doc-writer agent to create an accessible overview of the sandbox system for non-technical readers.\"\\n \\n\\n- Example 5:\\n Context: Multiple files need updating after a large feature lands.\\n user: \"I just landed the network namespace isolation feature. Update architecture/sandbox.md and architecture/networking.md.\"\\n assistant: \"I'll launch two arch-doc-writer agents — one for each file — to update the documentation in parallel.\"\\n \\n " -model: opus -color: yellow -memory: project ---- - -You are a principal-level technical writer with deep expertise in systems programming, distributed systems, and developer documentation. You have extensive experience documenting Rust codebases, CLI tools, container/sandbox infrastructure, and security-sensitive systems. Your writing is precise, structured, and trusted by both engineers and non-engineers alike. - -## Your Mission - -You maintain the contents of documentation files in the `architecture/` directory of this project. Your goal is to keep documentation perfectly synchronized with the actual codebase so that humans and agents can trust it as a reliable source of truth. You do NOT decide which files to create or how the directory is organized — you are given a specific file to update and you make its contents accurate, clear, and comprehensive. - -## Project Context - -This is the OpenShell project — a sandbox/isolation system built in Rust. - -The docs in `architecture/` are structured as subsystem[-component].md. Key sub-systems are: - -- build (build system) -- cluster (the entire deployment that can run on a single node or multi-node kubernetes cluster) -- gateway (the control plane / server system that manages a cluster and sandboxes) -- inference (access to models for agents and what they produce, includes privacy aware model routing) -- sandbox (long-running agentic environments that are strictly controlled by security policies) -- security - -Markdown files document 2-tuples of subsystem + component. - -Proto definitions live in `proto/`, Rust crates in `crates/`, and docs in `architecture/`. - -## Core Workflow - -When you receive a task to update a documentation file: - -1. **Read the target file** first to understand its current state, structure, and scope. -2. **Traverse the codebase** to understand the subsystem(s) the file documents. Read the relevant source files — don't guess or rely on memory. Key places to look: - - `crates/` for Rust source code - - `proto/` for protobuf definitions - - `Cargo.toml` files for dependency relationships - - `src/` directories for module structure - - Test files for behavioral expectations - - `CONTRIBUTING.md` for build/test/run instructions - - Existing `architecture/` docs for cross-references -3. **Identify what changed** by comparing the current code against what the documentation says. Note discrepancies, missing sections, outdated descriptions, and new functionality. -4. **Write the updated documentation** following the standards below. -5. **Self-verify** by re-reading relevant source files to confirm every claim in your documentation is accurate. - -## Audience Modes - -You operate in two modes based on the caller's instructions: - -### Non-Engineer Mode -When asked to write for non-engineers: -- Lead with **what** the system does and **why** it exists -- Use analogies and plain language — avoid jargon or define it inline -- Focus on capabilities, guarantees, and user-facing behavior -- Diagrams should show high-level data flow and system boundaries -- Code examples should be CLI commands a user would actually run, with plain-English explanations of what happens -- Omit internal implementation details unless they're essential to understanding behavior -- Structure: Purpose → How It Works (conceptual) → Examples → Guarantees/Limitations - -### Engineer Mode (default) -When asked to write for engineers, or when no audience is specified: -- Be precise about implementation details: data structures, control flow, error handling, concurrency model -- Reference specific files, functions, structs, and modules by name -- Include type signatures and code paths where they clarify behavior -- Diagrams should show internal component interactions, state machines, and data flow through specific modules -- Code examples should include both CLI usage AND traces through the codebase showing what happens internally -- Document edge cases, failure modes, and security boundaries -- Structure: Overview → Architecture → Components (with file references) → Data Flow → Examples with Code Traces → Error Handling → Security Considerations - -## Documentation Standards - -### Writing Style -- **Concise and direct.** Every sentence must earn its place. No filler, no hedging, no "it should be noted that." -- **Active voice.** "The proxy validates the hostname" not "The hostname is validated by the proxy." -- **Present tense** for describing current behavior. Past tense only for historical context. -- **Consistent terminology.** Use the same term for the same concept throughout. Match the terminology used in the source code. -- **No marketing language.** Don't say "powerful" or "robust" — describe what it does and let the reader judge. - -### Structure -- Use clear hierarchical headings (##, ###, ####) -- Start each major section with a 1-2 sentence summary -- Use bullet lists for enumerations, numbered lists for sequences/steps -- Keep paragraphs short — 3-5 sentences maximum -- Use code blocks with language annotations (```rust, ```bash, ```yaml) - -### Diagrams -Create diagrams using Mermaid syntax (```mermaid code blocks). Include diagrams for: -- **Component interaction**: How subsystems connect and communicate -- **Data flow**: How a request/command flows through the system -- **State machines**: For components with distinct states (e.g., sandbox lifecycle) -- **Sequence diagrams**: For multi-step processes involving multiple components - -Diagram guidelines: -- Label all edges with what flows between components -- Keep diagrams focused — one concept per diagram -- Use consistent naming that matches source code identifiers -- Add a brief caption or description above each diagram explaining what it shows - -### Code Traces -For practical examples, follow this pattern: -1. Show the CLI command or API call a user would execute -2. Trace what that command does through the codebase, referencing specific files and functions -3. Explain key decision points and branching logic -4. Show the expected output or side effects - -Example format: -```bash -# User runs: -openshell sandbox run --policy sandbox.yaml -- /bin/ls -``` -**Trace:** -1. `crates/openshell-cli/src/main.rs` → `SandboxRunCmd::execute()` -2. Policy loaded from YAML via `crates/openshell-sandbox/src/policy.rs` → `Policy::from_yaml()` -3. ... (continue through the actual code path) - -### Cross-References -- Link to other architecture docs when referencing related subsystems: `[Proxy Architecture](proxy.md)` -- Reference source files with relative paths from repo root: `crates/openshell-sandbox/src/lib.rs` -- When referencing plans, link to `architecture/plans/` - -### What NOT to Include -- Do not include speculative future plans unless they are documented in `architecture/plans/` -- Do not include TODO items — document current behavior -- Do not copy-paste large blocks of source code — reference it and explain it -- Do not document test utilities or internal test helpers unless they are part of the public interface - -## System Architecture Diagram - -The file `architecture/system-architecture.md` contains a top-level Mermaid diagram of the entire OpenShell system — all deployable components, external systems, communication protocols, and security boundaries. It is the single source of truth for the system's visual architecture. - -**After completing any documentation update**, check whether your changes affect the system-level architecture diagram. You MUST update `architecture/system-architecture.md` if any of the following are true: - -- A new deployable component or service was added or removed -- A new external system, API, or third-party dependency was introduced or removed -- Communication protocols between components changed (new connections, changed protocols, removed paths) -- Security boundaries or isolation layers changed -- Ports, endpoints, or addressing changed -- Data stores were added, removed, or changed - -When updating the diagram: -- Keep the Mermaid syntax valid and renderable -- Use the same component names as in the rest of the documentation -- Annotate arrows with communication types and protocols -- Avoid overlapping connections — keep the diagram readable -- Update the "Key Communication Flows" section below the diagram if flows changed -- Update the "Component Legend" table if new component categories were added - -If your documentation update does NOT affect any of the above, you do not need to modify the diagram. - -## Quality Checklist - -Before finishing, verify: -- [ ] Every file path referenced actually exists in the codebase -- [ ] Every function/struct/module name referenced exists in the code -- [ ] Every behavioral claim matches what the code actually does -- [ ] Diagrams accurately reflect current component relationships -- [ ] Code traces follow actual execution paths (verified by reading the source) -- [ ] No orphaned cross-references to removed or renamed components -- [ ] The document reads coherently from top to bottom -- [ ] Terminology is consistent with the source code -- [ ] `architecture/system-architecture.md` is updated if your changes affect system-level components, connections, or boundaries - -## Update your agent memory - -As you traverse the codebase to write documentation, update your agent memory with discoveries about: -- Codebase structure: where key modules, types, and entry points live -- Subsystem boundaries: how components interact and what interfaces they expose -- Naming conventions and terminology used in the code vs. documentation -- Architectural patterns: error handling strategies, async patterns, configuration approaches -- Common cross-references between architecture docs -- File paths that have moved or been renamed since the last documentation pass - -This builds institutional knowledge that makes future documentation updates faster and more accurate. - -# Persistent Agent Memory - -You have a persistent Persistent Agent Memory directory at `.claude/agent-memory/arch-doc-writer/`. Its contents persist across conversations. - -As you work, consult your memory files to build on previous experience. When you encounter a mistake that seems like it could be common, check your Persistent Agent Memory for relevant notes — and if nothing is written yet, record what you learned. - -Guidelines: -- `MEMORY.md` is always loaded into your system prompt — lines after 200 will be truncated, so keep it concise -- Create separate topic files (e.g., `debugging.md`, `patterns.md`) for detailed notes and link to them from MEMORY.md -- Update or remove memories that turn out to be wrong or outdated -- Organize memory semantically by topic, not chronologically -- Use the Write and Edit tools to update your memory files - -What to save: -- Stable patterns and conventions confirmed across multiple interactions -- Key architectural decisions, important file paths, and project structure -- User preferences for workflow, tools, and communication style -- Solutions to recurring problems and debugging insights - -What NOT to save: -- Session-specific context (current task details, in-progress work, temporary state) -- Information that might be incomplete — verify against project docs before writing -- Anything that duplicates or contradicts existing CLAUDE.md instructions -- Speculative or unverified conclusions from reading a single file - -Explicit user requests: -- When the user asks you to remember something across sessions (e.g., "always use bun", "never auto-commit"), save it — no need to wait for multiple interactions -- When the user asks to forget or stop remembering something, find and remove the relevant entries from your memory files -- Since this memory is project-scope and shared with your team via version control, tailor your memories to this project diff --git a/.github/ISSUE_TEMPLATE/feature_request.yml b/.github/ISSUE_TEMPLATE/feature_request.yml index 10ae9cbe50..fb7907f27a 100644 --- a/.github/ISSUE_TEMPLATE/feature_request.yml +++ b/.github/ISSUE_TEMPLATE/feature_request.yml @@ -91,7 +91,7 @@ body: attributes: label: Checklist options: - - label: I've reviewed existing issues and the architecture docs + - label: I've reviewed existing issues and the published docs required: true - label: This is a design proposal, not a "please build this" request required: true diff --git a/.github/copy-pr-bot.yaml b/.github/copy-pr-bot.yaml index a4e9281742..4cfbdc7f05 100644 --- a/.github/copy-pr-bot.yaml +++ b/.github/copy-pr-bot.yaml @@ -1,21 +1,3 @@ enabled: true auto_sync_draft: false auto_sync_ready: true -vetters_override: - - alangou - - derekwaynecarr - - drew - - elezar - - johnnygreco - - johntmyers - - kirit93 - - krishicks - - matthewgrossman - - mrunalp - - pimlock - - purp - - SDAChess - - shailendra-nv - - sjenning - - TaylorMutch - - zredlined diff --git a/.github/workflows/branch-e2e.yml b/.github/workflows/branch-e2e.yml index 4b939f79d2..49cf199a3c 100644 --- a/.github/workflows/branch-e2e.yml +++ b/.github/workflows/branch-e2e.yml @@ -224,6 +224,7 @@ jobs: test-matrix: >- [ {"environment":"ubuntu-docker-rootful","installer":"binaries","testsuite":"conformance"}, + {"environment":"ubuntu-k3s","installer":"k3s","testsuite":"conformance"}, {"environment":"fedora-podman-rootful","installer":"binaries","testsuite":"conformance"}, {"environment":"fedora-podman-rootless","installer":"binaries","testsuite":"conformance"} ] diff --git a/.github/workflows/integration-runner.yml b/.github/workflows/integration-runner.yml index 1048e5abd3..768879c7b8 100644 --- a/.github/workflows/integration-runner.yml +++ b/.github/workflows/integration-runner.yml @@ -25,6 +25,7 @@ on: default: >- [ {"environment":"ubuntu-docker-rootful","installer":"deb","testsuite":"conformance"}, + {"environment":"ubuntu-k3s","installer":"k3s","testsuite":"conformance"}, {"environment":"fedora-podman-rootful","installer":"binaries","testsuite":"conformance"}, {"environment":"fedora-podman-rootless","installer":"binaries","testsuite":"conformance"} ] @@ -67,6 +68,7 @@ jobs: - uses: ./.github/actions/setup-nix - name: Cache tmachine disks + if: matrix.environment != 'ubuntu-k3s' uses: actions/cache@caa296126883cff596d87d8935842f9db880ef25 # v5.1.0 with: path: ~/.cache/tmachine diff --git a/.github/workflows/integration-test.yml b/.github/workflows/integration-test.yml index bf23de2ae5..bc91802882 100644 --- a/.github/workflows/integration-test.yml +++ b/.github/workflows/integration-test.yml @@ -23,6 +23,7 @@ on: default: >- [ {"environment":"ubuntu-docker-rootful","installer":"binaries","testsuite":"conformance"}, + {"environment":"ubuntu-k3s","installer":"k3s","testsuite":"conformance"}, {"environment":"fedora-podman-rootful","installer":"binaries","testsuite":"conformance"}, {"environment":"fedora-podman-rootless","installer":"binaries","testsuite":"conformance"} ] diff --git a/.github/workflows/prepare-integration-inputs.yml b/.github/workflows/prepare-integration-inputs.yml index 0db9dbe61e..77be3181d0 100644 --- a/.github/workflows/prepare-integration-inputs.yml +++ b/.github/workflows/prepare-integration-inputs.yml @@ -89,12 +89,16 @@ jobs: - name: Log in to GHCR run: echo "${{ github.token }}" | docker login ghcr.io -u "${GITHUB_ACTOR}" --password-stdin - - name: Export runtime images + - name: Export OpenShell images env: IMAGE_TAG: ${{ steps.artifact-run.outputs.source_sha }} run: | mkdir -p artifacts/images + docker pull "ghcr.io/nvidia/openshell/gateway:${IMAGE_TAG}" + docker tag "ghcr.io/nvidia/openshell/gateway:${IMAGE_TAG}" openshell/gateway:tmachine + docker save --output artifacts/images/openshell-gateway-tmachine.tar openshell/gateway:tmachine + docker pull "ghcr.io/nvidia/openshell/sandbox:${IMAGE_TAG}" docker tag "ghcr.io/nvidia/openshell/sandbox:${IMAGE_TAG}" openshell/sandbox:tmachine docker save --output artifacts/images/openshell-sandbox-tmachine.tar openshell/sandbox:tmachine @@ -109,6 +113,9 @@ jobs: - name: Build test workload images run: nix run .#build-artifacts-test-images + - name: Package Helm chart + run: nix run .#build-artifacts-helm + - name: Upload integration inputs id: upload-integration-inputs uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 diff --git a/.gitignore b/.gitignore index b4e8cd1631..32842afa40 100644 --- a/.gitignore +++ b/.gitignore @@ -223,6 +223,7 @@ artifacts/ mise.local.toml # Ignore plans for now +/plans architecture/plans rfc.md diff --git a/.markdownlint-cli2.jsonc b/.markdownlint-cli2.jsonc index 125df0f812..d20a8c0ad5 100644 --- a/.markdownlint-cli2.jsonc +++ b/.markdownlint-cli2.jsonc @@ -1,7 +1,6 @@ { "globs": [ "*.md", - "architecture/**/*.md", "crates/**/*.md", "deploy/**/*.md", "docs/**/*.md", @@ -16,6 +15,7 @@ ".claude/**", ".opencode/**", ".github/**", + "plans/**", "architecture/plans/**", "**/node_modules/**", "target/**", diff --git a/.opencode/agents/arch-doc-writer.md b/.opencode/agents/arch-doc-writer.md deleted file mode 100644 index ac0e2540a2..0000000000 --- a/.opencode/agents/arch-doc-writer.md +++ /dev/null @@ -1,169 +0,0 @@ ---- -description: > - Use this agent when documentation in the `architecture/` directory needs to be - updated or created after implementing a feature, fix, refactor, or behavior - change. Launch one instance per file that needs updating. Maintains the - contents of architecture documentation files - does not decide which files - exist or how the directory is organized. -mode: subagent -color: "#e6b800" -tools: - bash: false ---- - -You are a principal-level technical writer with deep expertise in systems programming, distributed systems, and developer documentation. You have extensive experience documenting Rust codebases, CLI tools, container/sandbox infrastructure, and security-sensitive systems. Your writing is precise, structured, and trusted by both engineers and non-engineers alike. - -## Your Mission - -You maintain the contents of documentation files in the `architecture/` directory of this project. Your goal is to keep documentation perfectly synchronized with the actual codebase so that humans and agents can trust it as a reliable source of truth. You do NOT decide which files to create or how the directory is organized — you are given a specific file to update and you make its contents accurate, clear, and comprehensive. - -## Project Context - -This is the OpenShell project — a sandbox/isolation system built in Rust. - -The docs in `architecture/` are structured as subsystem[-component].md. Key sub-systems are: - -- build (build system) -- cluster (the entire deployment that can run on a single node or multi-node kubernetes cluster) -- gateway (the control plane / server system that manages a cluster and sandboxes) -- inference (access to models for agents and what they produce, includes privacy aware model routing) -- sandbox (long-running agentic environments that are strictly controlled by security policies) -- security - -Markdown files document 2-tuples of subsystem + component. - -Proto definitions live in `proto/`, Rust crates in `crates/`, and docs in `architecture/`. - -## Core Workflow - -When you receive a task to update a documentation file: - -1. **Read the target file** first to understand its current state, structure, and scope. -2. **Traverse the codebase** to understand the subsystem(s) the file documents. Read the relevant source files — don't guess or rely on memory. Key places to look: - - `crates/` for Rust source code - - `proto/` for protobuf definitions - - `Cargo.toml` files for dependency relationships - - `src/` directories for module structure - - Test files for behavioral expectations - - `CONTRIBUTING.md` for build/test/run instructions - - Existing `architecture/` docs for cross-references -3. **Identify what changed** by comparing the current code against what the documentation says. Note discrepancies, missing sections, outdated descriptions, and new functionality. -4. **Write the updated documentation** following the standards below. -5. **Self-verify** by re-reading relevant source files to confirm every claim in your documentation is accurate. - -## Audience Modes - -You operate in two modes based on the caller's instructions: - -### Non-Engineer Mode -When asked to write for non-engineers: -- Lead with **what** the system does and **why** it exists -- Use analogies and plain language — avoid jargon or define it inline -- Focus on capabilities, guarantees, and user-facing behavior -- Diagrams should show high-level data flow and system boundaries -- Code examples should be CLI commands a user would actually run, with plain-English explanations of what happens -- Omit internal implementation details unless they're essential to understanding behavior -- Structure: Purpose → How It Works (conceptual) → Examples → Guarantees/Limitations - -### Engineer Mode (default) -When asked to write for engineers, or when no audience is specified: -- Be precise about implementation details: data structures, control flow, error handling, concurrency model -- Reference specific files, functions, structs, and modules by name -- Include type signatures and code paths where they clarify behavior -- Diagrams should show internal component interactions, state machines, and data flow through specific modules -- Code examples should include both CLI usage AND traces through the codebase showing what happens internally -- Document edge cases, failure modes, and security boundaries -- Structure: Overview → Architecture → Components (with file references) → Data Flow → Examples with Code Traces → Error Handling → Security Considerations - -## Documentation Standards - -### Writing Style -- **Concise and direct.** Every sentence must earn its place. No filler, no hedging, no "it should be noted that." -- **Active voice.** "The proxy validates the hostname" not "The hostname is validated by the proxy." -- **Present tense** for describing current behavior. Past tense only for historical context. -- **Consistent terminology.** Use the same term for the same concept throughout. Match the terminology used in the source code. -- **No marketing language.** Don't say "powerful" or "robust" — describe what it does and let the reader judge. - -### Structure -- Use clear hierarchical headings (##, ###, ####) -- Start each major section with a 1-2 sentence summary -- Use bullet lists for enumerations, numbered lists for sequences/steps -- Keep paragraphs short — 3-5 sentences maximum -- Use code blocks with language annotations (```rust, ```bash, ```yaml) - -### Diagrams -Create diagrams using Mermaid syntax (```mermaid code blocks). Include diagrams for: -- **Component interaction**: How subsystems connect and communicate -- **Data flow**: How a request/command flows through the system -- **State machines**: For components with distinct states (e.g., sandbox lifecycle) -- **Sequence diagrams**: For multi-step processes involving multiple components - -Diagram guidelines: -- Label all edges with what flows between components -- Keep diagrams focused — one concept per diagram -- Use consistent naming that matches source code identifiers -- Add a brief caption or description above each diagram explaining what it shows - -### Code Traces -For practical examples, follow this pattern: -1. Show the CLI command or API call a user would execute -2. Trace what that command does through the codebase, referencing specific files and functions -3. Explain key decision points and branching logic -4. Show the expected output or side effects - -Example format: -```bash -# User runs: -openshell sandbox run --policy sandbox.yaml -- /bin/ls -``` -**Trace:** -1. `crates/openshell-cli/src/main.rs` → `SandboxRunCmd::execute()` -2. Policy loaded from YAML via `crates/openshell-sandbox/src/policy.rs` → `Policy::from_yaml()` -3. ... (continue through the actual code path) - -### Cross-References -- Link to other architecture docs when referencing related subsystems: `[Proxy Architecture](proxy.md)` -- Reference source files with relative paths from repo root: `crates/openshell-sandbox/src/lib.rs` -- When referencing plans, link to `architecture/plans/` - -### What NOT to Include -- Do not include speculative future plans unless they are documented in `architecture/plans/` -- Do not include TODO items — document current behavior -- Do not copy-paste large blocks of source code — reference it and explain it -- Do not document test utilities or internal test helpers unless they are part of the public interface - -## System Architecture Diagram - -The file `architecture/system-architecture.md` contains a top-level Mermaid diagram of the entire OpenShell system — all deployable components, external systems, communication protocols, and security boundaries. It is the single source of truth for the system's visual architecture. - -**After completing any documentation update**, check whether your changes affect the system-level architecture diagram. You MUST update `architecture/system-architecture.md` if any of the following are true: - -- A new deployable component or service was added or removed -- A new external system, API, or third-party dependency was introduced or removed -- Communication protocols between components changed (new connections, changed protocols, removed paths) -- Security boundaries or isolation layers changed -- Ports, endpoints, or addressing changed -- Data stores were added, removed, or changed - -When updating the diagram: -- Keep the Mermaid syntax valid and renderable -- Use the same component names as in the rest of the documentation -- Annotate arrows with communication types and protocols -- Avoid overlapping connections — keep the diagram readable -- Update the "Key Communication Flows" section below the diagram if flows changed -- Update the "Component Legend" table if new component categories were added - -If your documentation update does NOT affect any of the above, you do not need to modify the diagram. - -## Quality Checklist - -Before finishing, verify: -- [ ] Every file path referenced actually exists in the codebase -- [ ] Every function/struct/module name referenced exists in the code -- [ ] Every behavioral claim matches what the code actually does -- [ ] Diagrams accurately reflect current component relationships -- [ ] Code traces follow actual execution paths (verified by reading the source) -- [ ] No orphaned cross-references to removed or renamed components -- [ ] The document reads coherently from top to bottom -- [ ] Terminology is consistent with the source code -- [ ] `architecture/system-architecture.md` is updated if your changes affect system-level components, connections, or boundaries diff --git a/AGENTS.md b/AGENTS.md index 0ec8027da8..a590b117ce 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -80,8 +80,7 @@ These pipelines connect skills into end-to-end workflows. Individual skill files | `fern/` | Docs site config | Fern site config, components, and theme assets | | `skills/` | Public agent skills | Installable workflows for using and operating OpenShell | | `.agents/skills/` | Contributor agent skills | Repository-aware workflows for developing OpenShell | -| `.agents/agents/` | Agent personas | Sub-agent definitions (e.g., reviewer, doc writer) | -| `architecture/` | Architecture docs | Design decisions and component documentation | +| `.agents/agents/` | Agent personas | Sub-agent definitions (e.g., reviewer) | ## Public API Conventions @@ -110,7 +109,7 @@ design, and schema evolution. ## Plans -- Store plan documents in `architecture/plans`. This is git ignored so its for easier access for humans. When asked to create Spikes or issues, you can skip to GitHub issues. Only use the plans dir when you aren't writing data somewhere else specific. +- Store plan documents in `plans/`. This is git ignored so its for easier access for humans. When asked to create Spikes or issues, you can skip to GitHub issues. Only use the plans dir when you aren't writing data somewhere else specific. - When asked to write a plan, write it there without asking for the location. ## Sandbox Logging (OCSF) @@ -264,7 +263,7 @@ When behavior, commands, or development workflows change, review the related age ## Documentation -- When making changes, update the relevant documentation in the `architecture/` directory. +- Put crate-specific implementation details in the relevant crate `README.md`, design proposals in `rfc/`, and temporary plans in the ignored `plans/` directory. - When changes affect user-facing behavior, update the relevant published docs pages under `docs/` and navigation in `docs/index.yml`. - When changing gateway TOML fields, driver-specific config options, config defaults, or Helm rendering of `gateway.toml`, update `docs/how-it-works/gateways/configuration.mdx` in the same branch. - `fern/` contains the Fern site config, components, preview workflow inputs, publish settings, and publishing documentation in `fern/README.md`. @@ -272,17 +271,6 @@ When behavior, commands, or development workflows change, review the related age - Fern PR previews run through `.github/workflows/branch-docs.yml`. Release Dev publishes `dev`, and Release Tag publishes an immutable stable version plus `latest`. Both production paths call `.github/workflows/sync-docs.yml` once. - Use the `update-docs-from-commits` skill to scan recent commits and draft doc updates. -### Architecture Docs - -- Architecture docs are short canonical subsystem overviews, not exhaustive implementation notes. -- Update one of the existing top-level architecture docs before adding a new file. -- Put useful crate-specific details in the relevant crate `README.md`. -- Add a new top-level architecture doc only when explicitly requested or when an RFC-level design needs a stable home. -- Keep architecture docs focused on stable boundaries, data/control flow, invariants, and operational constraints. -- Remove stale detail instead of preserving it by default. -- Do not include testing transcripts, historical debugging notes, long source-file inventories, or field-by-field schema references. -- Put user-facing instructions in `docs/`, broad design proposals in `rfc/`, and temporary plans in ignored `architecture/plans/`. - ## Security - Never commit secrets, API keys, or credentials. If a file looks like it contains secrets (`.env`, `credentials.json`, etc.), do not stage it. diff --git a/CI.md b/CI.md index 1eb087f2a6..24bee35696 100644 --- a/CI.md +++ b/CI.md @@ -8,11 +8,10 @@ For local test commands see [TESTING.md](TESTING.md). For PR conventions see [CO PR CI that runs on NVIDIA self-hosted runners uses NVIDIA's copy-pr-bot. The bot mirrors trusted PR commits to internal `pull-request/` branches in this repository. The gated workflows trigger on pushes to those branches, not on the original PR. -When a PR is not mirrored automatically, only the GitHub users listed in -`.github/copy-pr-bot.yaml` under `vetters_override` can admit its current -revision with `/ok to test `. The list is a snapshot of the codeowners -and selected repository maintainers; update it when those people change. -This setting does not change the bot's automatic trust policy for ready PRs. +When a PR is not mirrored automatically, anyone with Write, Maintain, or Admin +access to this repository can admit its current revision with +`/ok to test `. This includes external maintainers with repository access. +Manual admission does not change the bot's automatic trust policy for ready PRs. `Branch Checks` run automatically after copy-pr-bot mirrors the PR. `Required CI Gates` posts PR-head statuses that verify the mirror exists, is current, and ran the expected push-based workflows. E2E suites are opt-in because they are more expensive and publish temporary images. @@ -40,6 +39,13 @@ The GitHub ruleset should require the `OpenShell / ...` statuses published by `Required CI Gates` plus the direct `OpenShell / Trivy Changes` result, not the push-triggered workflow jobs themselves. +### K3s conformance version baseline + +The tmachine `ubuntu-k3s` conformance lane pins Agent Sandbox v0.5.0 as the +compatibility baseline for the v1beta1 Sandbox API. It does not track the local +K3s development default, currently v1.0.3. OpenShell also supports v0.4.6 through +its v1alpha1 fallback, so v0.5.0 is not the overall minimum supported version. + ### Run only the policy advisor conformance tests Manually dispatch `Integration Tests` on the candidate branch with an @@ -347,11 +353,11 @@ Prerequisites: Flow: 1. Open the PR. The vouch check confirms first-time external contributors are vouched (otherwise their PRs are auto-closed). -2. If copy-pr-bot does not mirror it automatically, a listed vetter reviews the diff and comments `/ok to test ` with the latest commit SHA. Fork location alone does not determine whether a PR is mirrored automatically. +2. If copy-pr-bot does not mirror it automatically, a maintainer with Write access or greater reviews the diff and comments `/ok to test ` with the latest commit SHA. Fork location alone does not determine whether a PR is mirrored automatically. 3. After `/ok to test`, copy-pr-bot mirrors to `pull-request/`. From here the flow is identical to automatically admitted PRs: `Required CI Gates` verifies the mirror and required push workflows, and maintainers apply the E2E label when the extra suites are needed. 4. When the PR is ready to merge, maintainers add it to the merge queue so the queued integration state is tested before it reaches `main`. -Important: if a PR requires manual admission, every new commit needs another `/ok to test ` from a listed vetter before push-based CI will run on it. If a label is applied while the mirror is stale, `E2E Label Help` will post a comment explaining what's needed. +Important: if a PR requires manual admission, every new commit needs another `/ok to test ` from a maintainer with Write access or greater before push-based CI will run on it. If a label is applied while the mirror is stale, `E2E Label Help` will post a comment explaining what's needed. ## Merge queue diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 60a9d3f15c..935328b2b4 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -16,7 +16,7 @@ OpenShell is agent-first, not agent-only. The distinction matters: - **Do** use the skills in `.agents/skills/` — they exist to make your agent effective. - **Do** interrogate your agent until you understand every edge case and interaction in your changes. - **Don't** submit code you can't explain without your agent open. -- **Don't** use agents as a substitute for understanding the system. Read the architecture docs. +- **Don't** use agents as a substitute for understanding the system. Read the RFCs, crate READMEs, and published docs. ## First-Time Contributors @@ -470,7 +470,7 @@ These are the primary `mise` tasks for day-to-day development: | `deploy/` | Dockerfiles, Helm chart, Kubernetes manifests | | `docs/` | Published Fern docs source, navigation, and content assets | | `fern/` | Fern site config, components, and theme assets | -| `architecture/` | Architecture docs and plans | +| `plans/` | Local plans (git-ignored) | | `rfc/` | Request for Comments proposals | | `skills/` | Public skills for using and operating OpenShell | | `.agents/` | Contributor skills and persona definitions | diff --git a/Cargo.lock b/Cargo.lock index c6330b8897..b5e21ec266 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -636,7 +636,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "31b698c5f9a010f6573133b09e0de5408834d0c82f8d7475a89fc1867a71cd90" dependencies = [ "axum-core", - "base64", + "base64 0.22.1", "bytes", "form_urlencoded", "futures-util", @@ -737,6 +737,12 @@ version = "0.22.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "72b3254f16251a8381aa12e40e3c4d2f0199f8c6508fbecb9d91f575e0fbb8c6" +[[package]] +name = "base64" +version = "0.23.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "ac07cdecf99051d9a5238b80f35af32cdeba5b336e55d957b318b50137e18da5" + [[package]] name = "base64-simd" version = "0.8.0" @@ -832,7 +838,7 @@ version = "0.20.2" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "ee04c4c84f1f811b017f2fbb7dd8815c976e7ca98593de9c1e2afad0f636bff4" dependencies = [ - "base64", + "base64 0.22.1", "bollard-stubs", "bytes", "futures-core", @@ -2387,7 +2393,7 @@ version = "0.4.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "b3314d5adb5d94bcdf56771f2e50dbbc80bb4bdf88967526706205ac9eff24eb" dependencies = [ - "base64", + "base64 0.22.1", "bytes", "headers-core", "http 1.4.0", @@ -2685,7 +2691,7 @@ version = "0.1.20" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "96547c2556ec9d12fb1578c4eaf448b04993e7fb79cbaad930a656880a6bdfa0" dependencies = [ - "base64", + "base64 0.22.1", "bytes", "futures-channel", "futures-util", @@ -3160,7 +3166,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "0529410abe238729a60b108898784df8984c87f6054c9c4fcacc47e4803c1ce1" dependencies = [ "aws-lc-rs", - "base64", + "base64 0.22.1", "ed25519-dalek 2.2.0", "getrandom 0.2.17", "hmac 0.12.1", @@ -3183,7 +3189,7 @@ version = "0.24.0" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "2c75b990324f09bef15e791606b7b7a296d02fc88a344f6eba9390970a870ad5" dependencies = [ - "base64", + "base64 0.22.1", "chrono", "serde", "serde-value", @@ -3264,7 +3270,7 @@ version = "0.99.0" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "7fc2ed952042df20d15ac2fe9614d0ec14b6118eab89633985d4b36e688dccf1" dependencies = [ - "base64", + "base64 0.22.1", "bytes", "chrono", "either", @@ -3542,7 +3548,7 @@ version = "0.18.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "3589659543c04c7dc5526ec858591015b87cd8746583b51b48ef4353f99dbcda" dependencies = [ - "base64", + "base64 0.22.1", "http-body-util", "hyper", "hyper-util", @@ -3858,7 +3864,7 @@ version = "5.0.0" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "51e219e79014df21a225b1860a479e2dcd7cbd9130f4defd4bd0e191ea31d67d" dependencies = [ - "base64", + "base64 0.22.1", "chrono", "getrandom 0.2.17", "http 1.4.0", @@ -3990,7 +3996,7 @@ name = "openshell-cli" version = "0.0.0" dependencies = [ "anyhow", - "base64", + "base64 0.22.1", "bytes", "chrono", "clap", @@ -4066,7 +4072,7 @@ name = "openshell-core" version = "0.0.0" dependencies = [ "async-trait", - "base64", + "base64 0.22.1", "chrono", "glob", "ipnet", @@ -4107,7 +4113,7 @@ version = "0.0.0" dependencies = [ "async-trait", "aws-lc-rs", - "base64", + "base64 0.22.1", "futures", "openshell-core", "serde", @@ -4219,7 +4225,7 @@ dependencies = [ name = "openshell-driver-mxc" version = "0.0.0" dependencies = [ - "base64", + "base64 0.22.1", "futures", "noyalib", "openshell-core", @@ -4264,6 +4270,7 @@ dependencies = [ "serde_json", "tar", "temp-env", + "tempfile", "thiserror 2.0.20", "tokio", "tokio-stream", @@ -4305,7 +4312,7 @@ dependencies = [ name = "openshell-driver-vm" version = "0.0.0" dependencies = [ - "base64", + "base64 0.22.1", "bollard", "clap", "flate2", @@ -4491,6 +4498,7 @@ dependencies = [ "noyalib", "serde", "serde_json", + "thiserror 2.0.20", ] [[package]] @@ -4543,7 +4551,7 @@ version = "0.0.0" dependencies = [ "anyhow", "async-trait", - "base64", + "base64 0.22.1", "bytes", "capctl", "clap", @@ -4648,7 +4656,7 @@ dependencies = [ "aws-config", "aws-sdk-sts", "axum", - "base64", + "base64 0.22.1", "bytes", "chrono", "clap", @@ -4764,6 +4772,7 @@ dependencies = [ "rustls", "serde", "serde_json", + "socket2", "temp-env", "tempfile", "tokio", @@ -4816,7 +4825,7 @@ dependencies = [ "aws-credential-types", "aws-sigv4", "aws-smithy-runtime-api", - "base64", + "base64 0.22.1", "bytes", "flate2", "futures", @@ -4833,6 +4842,7 @@ dependencies = [ "openshell-isolation-interface", "openshell-ocsf", "openshell-policy", + "openshell-policy-schema", "openshell-supervisor-middleware", "openshell-supervisor-middleware-builtins", "prost-types", @@ -4868,7 +4878,7 @@ version = "0.0.0" dependencies = [ "anyhow", "async-trait", - "base64", + "base64 0.22.1", "bytes", "hex", "libc", @@ -4894,7 +4904,7 @@ dependencies = [ name = "openshell-tui" version = "0.0.0" dependencies = [ - "base64", + "base64 0.22.1", "crossterm 0.28.1", "futures", "indexmap", @@ -5170,7 +5180,7 @@ version = "3.0.6" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "1d30c53c26bc5b31a98cd02d20f25a7c8567146caf63ed593a9d87b2775291be" dependencies = [ - "base64", + "base64 0.22.1", "serde_core", ] @@ -5579,7 +5589,7 @@ version = "0.16.5" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "01b80ea363c31af2de2b92e3c07ed1156628f7838c4afb4df75ee78a37fedbd1" dependencies = [ - "base64", + "base64 0.22.1", "prost", "prost-types", "serde", @@ -5988,7 +5998,7 @@ version = "0.12.28" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "eddd3ca559203180a307f12d114c268abf583f59b03cb906fd0b3ff8646c1147" dependencies = [ - "base64", + "base64 0.22.1", "bytes", "futures-channel", "futures-core", @@ -6029,7 +6039,7 @@ version = "0.13.2" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "ab3f43e3283ab1488b624b44b0e988d0acea0b3214e694730a055cb6b2efa801" dependencies = [ - "base64", + "base64 0.22.1", "bytes", "futures-core", "futures-util", @@ -6637,6 +6647,7 @@ version = "1.0.149" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "83fc039473c5595ace860d8c4fafa220ff474b3fc6bfdb4293327f1a37e94d86" dependencies = [ + "indexmap", "itoa", "memchr", "serde", @@ -7002,7 +7013,7 @@ version = "0.9.0" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "05b44e85bf579a8eeb4ceaa77a3a523baf2bf0e9bac7e40f405d537b5d2d5ccb" dependencies = [ - "base64", + "base64 0.22.1", "bytes", "cfg-if", "crc", @@ -7104,7 +7115,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "87a2bdd6e83f6b3ea525ca9fee568030508b58355a43d0b2c1674d5f79dcd65e" dependencies = [ "atoi", - "base64", + "base64 0.22.1", "bitflags 2.13.2", "byteorder", "crc", @@ -7702,7 +7713,7 @@ checksum = "ac2a5518c70fa84342385732db33fb3f44bc4cc748936eb5833d2df34d6445ef" dependencies = [ "async-trait", "axum", - "base64", + "base64 0.22.1", "bytes", "h2", "http 1.4.0", @@ -7800,7 +7811,7 @@ version = "0.6.8" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "d4e6559d53cc268e5031cd8429d05415bc4cb4aefc4aa5d6cc35fbf5b924a1f8" dependencies = [ - "base64", + "base64 0.22.1", "bitflags 2.13.2", "bytes", "futures-util", @@ -7824,11 +7835,12 @@ checksum = "121c2a6cda46980bb0fcd1647ffaf6cd3fc79a013de288782836f6df9c48780e" [[package]] name = "tower-mcp-types" -version = "0.12.0" +version = "0.22.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6511f1f32c7cb7fd4525edc0eb4dcf307db8f7eceb2833ab24a37b4cc10cda61" +checksum = "bafe456e348e69895f4d0f1731a8f3094075f0e5f0a3f7d5f3c2443a0412b035" dependencies = [ - "base64", + "base64 0.23.1", + "indexmap", "serde", "serde_json", "thiserror 2.0.20", @@ -8819,7 +8831,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "08db1edfb05d9b3c1542e521aea074442088292f00b5f28e435c714a98f85031" dependencies = [ "assert-json-diff", - "base64", + "base64 0.22.1", "deadpool", "futures", "http 1.4.0", diff --git a/Cargo.toml b/Cargo.toml index 2e7237190f..7cab66d291 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -88,7 +88,7 @@ serde_json = "1" serde_yml = { package = "noyalib", version = "0.0.28", default-features = false, features = ["std", "compat-serde-yaml"] } toml = "0.8" apollo-parser = "0.8.5" -tower-mcp-types = "0.12.0" +tower-mcp-types = "=0.22.2" regex = "1" # HTTP client diff --git a/TESTING.md b/TESTING.md index 59031f242f..bce0f36a62 100644 --- a/TESTING.md +++ b/TESTING.md @@ -178,6 +178,35 @@ Rust-based e2e tests that exercise the `openshell` CLI binary as a subprocess. They live in the `openshell-e2e` crate and use a shared harness for sandbox lifecycle management, output parsing, and cleanup. +Exposed service URLs use virtual hostnames for gateway routing. Host-side tests +must connect the TCP socket directly to a reachable gateway listener address, +normally loopback, and send the service URL authority in the HTTP `Host` +header. Do not resolve `*.openshell.localhost`; resolver support for arbitrary +`.localhost` subdomains varies across local and CI environments. + +Treat the advertised service URL scheme as authoritative. For HTTPS, use the +virtual service hostname for TLS SNI and the configured gateway trust roots. +When the listener requires mTLS, present the active gateway client identity; +the local e2e wrappers register these materials under +`$XDG_CONFIG_HOME/openshell/gateways/$OPENSHELL_GATEWAY/mtls/`. Do not downgrade +an HTTPS service URL to plaintext when dialing loopback. Parse the URL and load +TLS material before entering a readiness loop so permanent configuration +errors fail immediately. Retry only transient connection failures and +documented readiness responses, and include the last observation in timeout +diagnostics. + +Verify exposed-service tests in both the default local mode and the +CI-equivalent HTTPS mode: + +```shell +mise run e2e:rust +OPENSHELL_ENABLE_LOOPBACK_SERVICE_HTTP=false mise run e2e:rust +``` + +When more than one test needs this behavior, put the transport in the shared +Rust e2e harness and require callers to use it instead of duplicating DNS, +HTTP `Host`, TLS SNI, and mTLS handling. + Suites: - Common suite (`--features e2e`) - driver-neutral CLI behavior, sandbox lifecycle, sync, port forwarding, policy, and provider tests. diff --git a/architecture/README.md b/architecture/README.md deleted file mode 100644 index f8c6b24048..0000000000 --- a/architecture/README.md +++ /dev/null @@ -1,187 +0,0 @@ -# OpenShell Architecture - -OpenShell runs fleets of autonomous AI agents in sandboxed environments with explicit -policy, credential, identity, and network boundaries. The target architecture is -built around three stable runtime components: the **CLI**, the **Gateway**, and -the **Supervisor**. - -The CLI, SDK, and TUI provide user-facing access. The gateway is the -authenticated control plane: it owns API access, durable state, policy and -settings delivery, provider configuration and attachments, and relay -coordination. The supervisor runs inside every sandbox workload and is the local -security boundary. It launches the agent as a restricted child process and -enforces policy where process identity, filesystem access, network egress, and -runtime credentials are visible. - -Infrastructure-specific work sits behind integration boundaries. Compute, -credentials, control-plane identity, and sandbox identity each have a driver or -adapter boundary so OpenShell can integrate with native runtimes, secret stores, -identity providers, and workload identity systems without moving those concerns -into the core gateway or sandbox model. - -```mermaid -flowchart TB - subgraph USER["User Interfaces"] - CLI["CLI"] - SDK["SDK"] - TUI["TUI"] - end - - subgraph CP["Control Plane"] - GW["Gateway core"] - PROVER["Policy prover"] - DB[("Shared persistence")] - COMPUTE["Compute"] - CREDS["Credentials"] - CPIDENT["Control-plane identity"] - SIDENT["Sandbox identity"] - CDRV["Compute driver"] - CRDRV["Credentials driver"] - CPIDRV["Control-plane identity driver"] - SIDRV["Sandbox identity driver"] - end - - subgraph INFRA["Integrated Infrastructure"] - RUNTIME["Docker / Podman / Kubernetes / VM"] - SECRETSTORE["Eg: Keychain / Secret Service / Vault / Kubernetes Secrets"] - IDP["Eg: mTLS / OIDC / Local identity"] - WORKLOADID["Eg: SPIFFE / Gateway-issued workload identity"] - end - - subgraph DP["Sandbox Data Plane"] - SUP["Supervisor"] - PROXY["Policy proxy"] - POLICY["OPA policy engine"] - AGENT["Restricted agent process"] - end - - CLI -->|"gRPC / HTTP"| GW - SDK -->|"gRPC / HTTP"| GW - TUI -->|"gRPC / HTTP"| GW - - GW --> DB - GW -->|"proposed policy changes"| PROVER - GW --> COMPUTE - GW --> CREDS - GW --> CPIDENT - GW --> SIDENT - - COMPUTE -->|"gRPC / UDS"| CDRV - CREDS -->|"gRPC / UDS"| CRDRV - CPIDENT -->|"gRPC / UDS"| CPIDRV - SIDENT -->|"gRPC / UDS"| SIDRV - - CDRV --> RUNTIME - CRDRV --> SECRETSTORE - CPIDRV --> IDP - SIDRV --> WORKLOADID - RUNTIME -->|"provisions workload"| SUP - - SUP -->|"outbound control, config, logs, relay"| GW - SUP -->|"spawn + restrict"| AGENT - AGENT -->|"all ordinary egress"| PROXY - PROXY -->|"evaluate"| POLICY - PROXY -->|"allowed traffic"| EXT["External services"] - PROXY -->|"profile-authorized traffic"| MODEL["Model providers"] -``` - -## Core Boundaries - -| Component | Boundary | -|---|---| -| CLI, SDK, TUI | User-facing management surfaces. They talk to the gateway and do not need to know which infrastructure drivers are active. | -| Gateway | Authenticated control plane, API server, durable state, policy and settings delivery, provider config, supervisor session ownership, and relay coordination. | -| Policy prover | Gateway-linked formal verification of proposed policy changes. It reports categorical findings (new credentialed reach, new HTTP methods, L7 bypass, link-local reach); any finding blocks auto-approval. The same crate backs the standalone `openshell-prover check` boundary command, which runs on local files without a gateway. See [Security Policy](security-policy.md). | -| Compute subsystem | Sandbox lifecycle semantics: creation, deletion, watching, reconciliation, and state transitions. Platform provisioning details belong to the compute driver. | -| Credentials subsystem | Logical provider and credential resolution. Secret storage and platform-native credential access belong to credentials drivers. | -| Control-plane identity | Authentication and authorization for users, operators, and API clients. External identity verification belongs to identity drivers. | -| Sandbox identity | Workload identity for supervisors and sandbox-to-sandbox authorization. Identity issuance or verification belongs to sandbox identity drivers. | -| Supervisor | Sandbox-local security boundary. It prepares isolation, fetches config, injects credentials, runs relay endpoints, starts the proxy, and launches restricted agent processes. | -| Policy proxy | Mandatory egress path for agent traffic. It enforces destination, binary identity, SSRF, TLS/L7, and endpoint-bound provider credential injection. | - -## Integrating with the Ecosystem - -OpenShell should integrate with infrastructure ecosystems instead of replacing -them. The core value is safe, policy-enforced agent execution. Runtimes, -schedulers, secret stores, identity providers, workload identity systems, image -pipelines, storage, and GPU or device exposure should remain owned by the -platforms that already provide them. - -The gateway owns OpenShell control-plane semantics: sandbox state, lifecycle -ordering, policy and settings resolution, credential mapping, authorization, -and relay coordination. Drivers translate those semantics into platform-native -operations. They should stay thin, preserve -native behavior by default, and report platform lifecycle events back through -the shared contracts. - -The supervisor owns OpenShell sandbox semantics. Filesystem policy, process -privilege reduction, network proxying, endpoint-bound credential injection, -security logging, and gateway relay behavior should remain -consistent across runtimes. - -This keeps OpenShell usable in local single-player setups, Kubernetes -deployments, VM-backed sandboxes, and future third-party environments. A new -integration should make OpenShell feel like a well-behaved member of that -ecosystem. - -## Gateways and Sandboxes - -The gateway and sandbox split control-plane authority from runtime enforcement. -The gateway owns durable platform state: sandboxes, policy revisions, runtime -settings, provider records and profiles, session records, and -authorization decisions. A sandbox owns the local execution boundary: process -identity, filesystem access, network egress, credential injection, local logs, -and the agent child process. - -The relationship is supervisor initiated. Each sandbox supervisor connects -outbound to a known gateway endpoint, authenticates as a sandbox workload, and -keeps a live session open for control traffic and relays. This avoids requiring -every compute driver to solve gateway-to-sandbox reachability through pod IPs, -bridge networks, port mappings, NAT traversal, or bespoke tunnels. The common -runtime requirement is narrower: the supervisor must be able to reach the -gateway. - -The compute-driver capability contract identifies whether a driver reports -runtime readiness. Most drivers use the supervisor session model above. A -driver that sets `driver_reports_runtime_readiness` may self-report readiness -without a supervisor session. Every driver receives the canonical create-time -policy in `DriverSandboxSpec`; drivers without a supervisor use the existing -sandbox configuration API for later revisions. The Windows MXC driver reports -its own readiness and does not expose interactive connect or governed egress. - -The gateway delivers desired state; the sandbox applies it locally. Policy, -settings, provider attachments, and credential bindings flow from the gateway -to the supervisor. The supervisor validates and applies what can change at runtime, -keeps last-known-good config when refresh fails, and leaves static isolation -controls in place until the sandbox is recreated. - -Live operations use the same authenticated gateway-supervisor relationship. -Config refresh, policy updates, credential delivery, log push, connect, exec, -file sync, and relay setup are multiplexed over supervisor sessions. If a -session drops, the sandbox may keep running, but live operations fail or become -unreachable until the supervisor reconnects and reconciles state. - -## Architecture Docs - -Architecture docs are short subsystem overviews. User-facing how-to content -lives in `docs/`. Implementation notes that only matter to one crate belong in -that crate's `README.md`. - -| Document | Purpose | -|---|---| -| [Gateway](gateway.md) | Gateway control plane, auth, APIs, persistence, settings, and relay coordination. | -| [Sandbox](sandbox.md) | Sandbox supervisor, child process isolation, proxy, provider credentials, connect, and logs. | -| [Sandbox Limits](sandbox-limits.md) | Sandbox supervisor and egress safety ceilings, ownership rules, current enforcement, and known gaps. | -| [Security Policy](security-policy.md) | Policy model, enforcement layers, policy updates, policy advisor, and security logging. | -| [Compute Runtimes](compute-runtimes.md) | Docker, Podman, Kubernetes, VM, sandbox images, and runtime-specific responsibilities. | -| [Build](build.md) | Build artifacts, CI/E2E, docs site validation, and release packaging. | -| [Google Vertex AI Provider](google-vertex-ai-provider.md) | Implementation reference for the `google-vertex-ai` provider, from CLI through gateway to sandbox. | -| [Windows MSVC Build](windows-msvc-build.md) | Build-only native Windows MSVC lane (x64/ARM64) and unsupported-runtime behavior on Windows. | - -## `rfc/` vs `architecture/` - -Broad design proposals start as GitHub issues. If maintainers decide a proposal needs broad consensus, they assign an RFC number from the issue and the RFC lives in `rfc/`. Once an RFC is adopted, appropriate details should be written back to architecture docs. - -`architecture/` serves as the canonical reference for OpenShell's design and architecture. - -`rfc` serves to help facilitate discussion and ensure the few changes that need that level of review are appropriately designed. These are useful for understanding the context in which certain architecture designs were made. diff --git a/architecture/build.md b/architecture/build.md deleted file mode 100644 index 556a3a0094..0000000000 --- a/architecture/build.md +++ /dev/null @@ -1,610 +0,0 @@ -# Build - -This page records the stable build, CI, docs, and release architecture. It is -not a command reference. Contributor-facing workflow details live in -`CONTRIBUTING.md`, `CI.md`, and published docs. - -## Artifacts - -OpenShell builds these main artifacts: - -| Artifact | Source | -|---|---| -| Gateway binary | `crates/openshell-gateway` | -| CLI binaries and system packages | `crates/openshell-cli` plus release packaging | -| E2E conformance CLI | `crates/openshell-conformance-cli` | -| Standalone policy prover | `crates/openshell-prover-cli` | -| Python SDK wheel | `python/openshell` | -| TypeScript SDK package | `sdk/typescript` | -| Gateway container image | `deploy/docker/Dockerfile.gateway` | -| Sandbox runtime binary and container image | `crates/openshell-sandbox` and `deploy/docker/Dockerfile.sandbox` | -| Supervisor binary and container image | `crates/openshell-supervisor` and `deploy/docker/Dockerfile.supervisor` | -| Helm chart | `deploy/helm/openshell` | -| VM driver/runtime assets | `crates/openshell-driver-vm` | -| Published docs site | `docs/` rendered by Fern config in `fern/` | - -Release Tag publishes the same tagged docs commit as an immutable `vX.Y.Z` -snapshot and the mutable `latest` alias. Release Dev updates `dev`. The Fern -selector pins `latest` and `dev`, then lists versioned snapshots newest first. -The current release workflows do not request availability badges. - -Workload images are standard OCI images supplied by operators or users. - -## Build Features - -Anonymous telemetry emission is gated behind a default-on `telemetry` Cargo -feature. It is defined in `openshell-core` (where the emission code, HTTP -client, and endpoint live) and forwarded by the binary crates that emit or -collect telemetry: `openshell-gateway`, `openshell-sandbox`, -`openshell-supervisor`, and `openshell-driver-vm`. Every crate depends on -`openshell-core` with `default-features = false`, so the binary crate's feature -is the single switch that enables `openshell-core/telemetry` for its build -graph. In-process drivers (`docker`, `kubernetes`, `podman`) inherit the -gateway's setting through feature unification and carry no passthrough. - -Building a binary without the `telemetry` feature compiles out telemetry -entirely: no endpoint, no telemetry HTTP client, and no emission code. With -telemetry compiled out, `telemetry::enabled()` is always `false` and the -`emit_*` helpers are no-ops, so the data-model types stay available and -dependent crates compile unchanged. The runtime `OPENSHELL_TELEMETRY_ENABLED` -switch remains the way to disable telemetry in a default (telemetry-enabled) -build. - -Cargo cannot subtract a single default feature, so each of the three binary -crates also defines a `defaults-without-telemetry` alias listing every default -except `telemetry`. Telemetry-free builds use -`--no-default-features --features defaults-without-telemetry` and stay correct -as the default set grows. The alias is a keep-list, -not a switch: enabling it on top of the defaults would otherwise yield a -telemetry-on binary that reads as telemetry-free, so each crate root carries a -`compile_error!` for the `telemetry` + `defaults-without-telemetry` combination. -`rust:verify:defaults-without-telemetry` guards both properties — that each -alias still equals its crate's defaults minus `telemetry`, and that the -mutual-exclusion error is wired up — and `rust:verify:telemetry-off` builds -through the alias and inspects the resulting binaries for telemetry markers. - -Supervisor upstream TLS root-store selection is controlled by the -`bundled-ca-roots` Cargo feature (on by default). Default builds use Mozilla -roots through `webpki-roots` plus locally-installed CAs from the system bundle. -Building without `bundled-ca-roots` switches to the platform trust store via -`rustls-native-certs` and excludes bundled Mozilla root crates such as -`webpki-roots` and `webpki-root-certs` from the dependency graph. The -`system-ca-roots` feature alias on `openshell-supervisor` includes all other -defaults (currently `telemetry`) except `bundled-ca-roots`, so Linux -distribution builds (e.g. RPM) can use -`--no-default-features --features system-ca-roots` without manually re-adding -unrelated defaults. Other Rustls clients use native roots directly because that -already satisfies Linux distribution trust-store policy. - -The workspace uses `z3` versions whose `z3-sys` dependency keeps downloader -HTTP/TLS support behind explicit build features, so default system-Z3 builds do -not reintroduce bundled Mozilla roots. Release builds that need bundled Z3 -continue to opt in with `bundled-z3`. - -Release workflows build the standalone `openshell-prover` executable for Linux -musl x86_64 and aarch64 and macOS Apple Silicon. The standard Debian, RPM, and -Homebrew installations include it. Releases also publish one standalone archive -per target plus a dedicated SHA-256 manifest. Before publication, target-native -jobs extract each archive, reject host Z3 or Nix store linkage, and run a real -local containment check. The standalone artifact therefore requires neither an -OpenShell installation nor a separately installed Z3 runtime. - -## Linux Runtime Environments - -OpenShell uses different Linux libc environments for different host artifacts. -The standalone `openshell` CLI is built as a static musl binary so it can run on -a wide range of Linux distributions without depending on the host's glibc. Host -runtime binaries that use the GNU/Linux runtime environment are GNU-linked. -`openshell-gateway` and `openshell-driver-vm` are built with a glibc 2.28 floor. -The gateway bundles z3 into the release binary so Linux packages, standalone -tarballs, and gateway images do not depend on distro-specific z3 shared-library -SONAMEs. - -The workload-side `openshell-sandbox` binary is statically linked with musl so -drivers can stage it into an arbitrary agent image without depending on that -image's libc. The separate `openshell-supervisor` binary is dynamically linked -with GNU libc and uses the same glibc 2.28 compatibility floor as the gateway. - -## Container Builds - -Docker E2E tool-dependent workloads use a dedicated Noble-based fixture, -separate from the product's minimal default image. The fixture supplies the -test identity and tools, with Python aligned to the host test runner for -serialized callable compatibility. Default-image coverage retains the product -image. Other compute-driver test lanes retain their existing workload fixtures. - -The Docker image pipeline is a two-step flow: build the Rust binary natively -for the target architecture, then assemble the container image from the -prebuilt binary. The gateway, sandbox, and supervisor images use distinct -Dockerfiles under `deploy/docker/`. None of the Dockerfiles compile Rust; they -copy staged binaries out of -`deploy/docker/.build/prebuilt-binaries//` into the final image. - -Local binary staging is driven by `tasks/scripts/stage-prebuilt-binaries.sh`. Because -staging cross-compiles on the host, it sources `tasks/scripts/build-env.sh` and -raises the per-process open-file limit before invoking `cargo zigbuild` on -macOS — the static musl link opens hundreds of `.rlib` files at once and would -otherwise fail with `ProcessFdQuotaExceeded` under macOS's default soft limit of -256. The guard is a no-op on Linux and when `cargo-zigbuild` is absent. Gateway -binaries use `cargo zigbuild` with GNU targets pinned to glibc 2.28, including -native-architecture builds, so the gateway image, standalone tarballs, and Linux -packages share the same host portability floor. The gateway build enables -`bundled-z3`. Linux VM driver release artifacts use the same glibc floor so -package-managed VM support does not raise the package runtime requirement. -Gateway staging and release workflows set up the Zig C/C++ wrapper before -bundled Z3 builds and verify the maximum referenced `GLIBC_*` symbol version -before publishing or copying artifacts. -Supervisor staging uses the GNU build path and verifies the glibc 2.28 floor. -Sandbox staging uses the static musl build path. Local Docker image tasks infer the -target architecture from `DOCKER_PLATFORM` when set. Otherwise, they require -valid container engine host metadata and fail when the engine query is -unavailable or reports an unsupported architecture, avoiding host-kernel -fallbacks that can target the wrong architecture. CI instead compiles binaries -in platform-specific Nix development shells through reusable workflows and the -shared `build-rust-binary` action. The image build downloads each binary artifact -into the staging directory before running Buildx. - -Gateway and supervisor binaries staged into branch E2E, Release Dev, and Release -Tag images are compiled through `cargo auditable` (pinned in `mise.toml`), which -embeds a `.dep-v0` section describing the Rust dependencies actually compiled -into the binary. That section holds data rather than symbols, so it survives the -workspace's `strip = true` release profile, and Syft can catalog the crates -present in image binaries instead of inferring them from the source tree. This -is a different artifact from the source SBOM produced by `syft dir:.` in -`tasks/sbom.toml`, which describes the checkout, and from the image SBOM -attestation below, which describes a published image. - -The shared binary build action compiles release artifacts with `cargo auditable`. -The standalone prover uses this same action, while its package workflow adds -target-native extracted-archive linkage and containment smoke checks before -producing its checksum manifest. -Branch E2E, Release Dev, and Release Tag image jobs stage those same artifacts -instead of rebuilding binaries in Docker. Each binary build scans its output with -Syft and requires at least one decoded Cargo package before uploading the -artifact. Darwin builds replace Nix's `libiconv` load command with the macOS -system install name, ad-hoc sign the modified binary, and fail if `otool -L` -reports any remaining `/nix/store` dependency. Runtime and Syft verification -run after that normalization. The CI image gains the pinned `cargo-auditable` -tool through `mise install --locked` but ships no auditable OpenShell binary of -its own. - -Pushed Docker images carry minimal SLSA provenance and a per-platform SPDX SBOM -generated by BuildKit's default Syft scanner. The registry exporter uses OCI -media types and `oci-artifact=true`, so each attestation identifies its subject. -GHCR exposes these through the image index because it has no referrers API. - -Attestations require a registry-backed image index. Local builds therefore keep -`--provenance=false`, and Podman builds carry neither attestation. -`tasks/scripts/verify-image-sbom.sh` verifies the merged multi-arch tag and runs -with `--require-cargo` for auditable builds, so those attestations must also -contain Cargo packages. - -Runtime layout: - -- **Gateway**: `gcr.io/distroless/cc-debian13:nonroot` base, GNU-linked binary at - `/usr/local/bin/openshell-gateway`, runs as UID/GID `1000:1000`. Linux GNU - gateway binaries must not reference `GLIBC_*` symbols newer than - `GLIBC_2.28`; release workflows verify this before publishing artifacts. The - gateway bundles z3, so the image does not need a distro-provided z3 runtime. - The base is pinned to a multi-architecture digest; distro security updates - require refreshing that digest and rebuilding the gateway image. - Updating the container's glibc package does not raise the binary's glibc - compatibility floor. -- **VM driver**: host GNU-linked binary installed at - `/usr/libexec/openshell/openshell-driver-vm` in Linux packages and published - as a release artifact. Linux GNU VM driver binaries must not reference - `GLIBC_*` symbols newer than `GLIBC_2.28`; release workflows verify this - before publishing artifacts. Nix produces the platform-specific compressed - runtime inputs. CI combines them with the matching supervisor artifact in a - runner-temporary directory outside Cargo's `target/` before the shared Rust - cache action runs. An explicitly configured VM runtime bundle is required to - contain every non-empty embedding input; the driver build fails before - packaging when an input is absent or empty. -- **Sandbox**: Alpine-based `openshell/sandbox` image containing the static - musl `/openshell-sandbox` binary and its static VM guest-init helper. - Drivers stage this binary into the workload trust domain. -- **Supervisor**: digest-pinned `gcr.io/distroless/base-nossl-debian13` base - with the dynamically linked GNU `/openshell-supervisor` binary. The base - supplies glibc and CA roots without a shell, package manager, OpenSSL or zlib. - GNU supervisor builds must not reference `GLIBC_*` symbols newer than - `GLIBC_2.28`. Image defaults remain UID 0 and working directory `/`; compute - drivers set the runtime identity and writable mounts. Docker stages private - files with the same numeric identity as the supervisor so archive uploads - preserve access regardless of the base image's default user. Health probes execute - the supervisor binary directly. Base updates require refreshing the - multi-architecture digest and rebuilding the image. - -Gateway image builds bake the corresponding supervisor image tag into the -gateway binary so Docker sandboxes do not depend on `:latest` by default. -The Helm chart omits the supervisor image from gateway configuration unless an -operator supplies a repository or tag override, preserving that build-time -pairing for Kubernetes sandboxes as well. -Package formulas also pin Docker supervisor extraction to the matching release -image tag so standalone gateway binaries do not infer image tags from package -versions. -The Homebrew service keeps gateway TLS under the Homebrew state directory but -mirrors Docker sandbox client TLS into `$HOME/.local/state/openshell/homebrew/tls` -at service start, because Docker Desktop bind mounts must use paths visible to -the macOS user's shared home directory. -On upgrade, the formula atomically replaces only exact package-generated -schema-v1 gateway configs, leaving user-edited configs untouched. - -Local image work should use `mise` tasks rather than direct Docker commands so -the same staging and tagging assumptions are used locally and in CI. - -Container-engine selection is centralized in `tasks/scripts/container-engine.sh`. -`CONTAINER_ENGINE=docker|podman` is the only explicit override. Docker- and -Podman-backed e2e wrappers validate that override against their lane, set -`OPENSHELL_E2E_DRIVER`, and reject the removed -`OPENSHELL_E2E_CONTAINER_ENGINE` selector so build helpers and Rust e2e support -containers use the same engine. When no explicit override is present, an e2e -driver requirement wins, then a local-cluster requirement, then host -auto-detection. - -Local Kubernetes image workflows opt into cluster-aware selection with -`CONTAINER_ENGINE_TARGET=local-k8s-cluster`. The hint is intentionally scoped to -Skaffold-style `push: false` builds where the image must land in the engine -backing the active local cluster: `k3d-*` contexts require Docker, `kind-*` -contexts use `KIND_EXPERIMENTAL_PROVIDER=docker|podman` when set, and ambiguous -or unknown contexts require an explicit `CONTAINER_ENGINE`. Other image builds -do not infer from kube context. - -## Disposable Test Guests - -The Nix test guest harness under `nix/test-guest` boots native-architecture cloud images -through QEMU for package, release, and E2E validation. A prepared cache entry is -captured after the exact ordered Ansible configuration list and before -test-specific packages, copied binaries, forwarded ports, or commands. -On macOS, the test guest and tmachine paths use the same pinned QEMU and OVMF -package set so the hypervisor and firmware remain compatible. - -Prepared disks are flattened, sanitized QCOW2 images. The local cache keeps them -read-only and each test receives a fresh writable overlay and cloud-init -identity. The optional shared cache stores the compressed standalone disk and -its compatibility metadata as a custom OCI artifact. Normal test runs ensure -the exact local entry exists, invoking the cache builder automatically on a -miss before booting a disposable overlay. The separate cache app owns OCI -pulls and explicit publication. OCI pulls require a trusted manifest digest -and retain that provenance with the local entry; mutable tags are used only -for explicit publication. - -CLI conformance runs after target provisioning and operates only through the -configured OpenShell CLI. The smoke scenario verifies the black-box sandbox -lifecycle by creating, inspecting, executing in, and deleting a sandbox. The -file-transfer scenario verifies portable upload and download behavior, Git-aware -filtering, and sandbox workspace path safety. -Feature suites use the same disposable guest but may provision isolated -dependencies after installation. The Keycloak provider-refresh suite starts a -guest-local Keycloak realm and verifies a successful OAuth refresh followed by -revocation and the gateway's reauthorization-required recovery state. - -Tmachine environments define the guest machine and runtime setup, while named -installers define how OpenShell is installed. This keeps the runtime mode -independent from binary or package installation and lets multiple installers -reuse the same prepared setup disk. The `none` installer skips OpenShell -installation and boots the prepared environment directly. - -### Interactive tmachine shell - -The test command is `tmachine test `. The -`shell` testsuite prepares the selected environment and installer, then opens an -interactive SSH session in the disposable guest for manual debugging. - -Start an Ubuntu Docker guest without installing OpenShell: - -```shell -nix run .#tmachine -- test ubuntu-docker-rootful none shell -``` - -Replace `none` with `deb` to install the locally staged Debian package before -opening the shell: - -```shell -nix run .#tmachine -- test ubuntu-docker-rootful deb shell -``` - -Exit the SSH session to shut down and discard the disposable guest. - -The `tests/tmachine` setup and install caches include a digest of the -entire directory containing `ANSIBLE_CONFIG`, including local roles, task -includes, templates, inventory, and requirements. The digest uses sorted -relative paths, file contents, and executable permissions; source symlinks -are unsupported. Both keys also retain the ordered playbook paths and contents, -their base disk contents, and whether Galaxy is enabled; install keys -include named artifact inputs. The top-level `.roles` directory is excluded: -Galaxy release pins in `requirements.yaml` are treated as immutable, including -any transitive dependency pins. Cache misses with Galaxy enabled reinstall -the required roles and their dependencies before running playbooks. - -The `tests/artifacts.nix` helpers build the CLI, conformance CLI, and sandbox -with musl, and the gateway and supervisor with GNU. Image assembly stages -the gateway, sandbox, and supervisor as separate binaries for their respective -Dockerfiles. The helpers stage binaries under `artifacts/binaries` so local and -CI builds expose the same inputs to tmachine and image assembly. The Ubuntu -Docker and Fedora Podman environments import both local runtime images and -configure the gateway to use them. The Ubuntu `deb` installer consumes -`artifacts/packages/openshell.deb`; the `binaries` installer remains available -for direct executable installation on every environment. Release Dev and -Release Tag run Ubuntu conformance through the Debian package, while Fedora -continues using direct executable installation until RPM coverage is available. -The release canary separately exercises the public installer on Ubuntu. The -installer selects the OpenShell Snap only with `OPENSHELL_INSTALL_METHOD=snap` -or when the Snap is already installed; otherwise it uses the native Debian or -RPM package. The Snap requires a compatible, preinstalled non-Snap Docker -daemon. Its positive canary uses system Docker; negative preflight coverage -verifies that the installer rejects both missing Docker and the Docker Snap -before installing OpenShell. -Explicit release tags and the `pre` alias always use the native Debian or RPM -package path. The `pre` alias -checks matching Git tags in version order, then looks up the exact platform -artifact and verifies the release run instead of listing every repository -artifact. - -Snapd runs the gateway as a root-owned system service. Its generated client -certificates reside in root-owned snap state. The installer copies the client -bundle into the target user's private Snap state and registers the TLS endpoint; -direct Snap installs require the same enrollment. The install and post-refresh -hooks replace configs that explicitly enable plaintext or unauthenticated access -with the secure default. Snap refreshes -restart the gateway so the migrated config takes effect immediately. - -Debian and RPM packages instead run systemd user services with user-owned mTLS -material. Sandbox-to-gateway sessions remain authenticated with gateway-minted -JWTs. - -The Debian qualification profile keeps candidate-image overrides outside the -operator-owned gateway configuration: it writes a harness-owned file under -`/var/lib/openshell-qualification` and selects it through the packaged systemd -unit's `gateway.env` hook. Ordinary package installations continue to use the -gateway's built-in runtime-image defaults unless the operator configures an -override. - -## Python Wheel Packaging - -The generated protobuf/gRPC stubs under `python/openshell/_proto/` are gitignored -build outputs of `mise run python:proto`. Setuptools includes them through the -package-data configuration in `pyproject.toml`. Release workflows build the -wheel directly and do not produce a source distribution. Setuptools SCM derives -local versions from Git and accepts the release workflow's computed version -through its distribution-specific override. - -The build produces one platform-independent `py3-none-any` wheel. A verifier -checks its tag, metadata, version, required package files, and the absence of -native files or an `openshell` executable entry point. Release workflows build -the wheel once, install it in a clean virtual environment, import the public -package modules, and confirm that installation did not create an `openshell` -command. - -## TypeScript SDK Packaging - -The native TypeScript SDK in `sdk/typescript` uses Connect over the generated -OpenShell protobuf surface. `sdk/typescript/buf.gen.yaml` selects the client -proto closure, and `mise run sdk:ts:proto` generates gitignored sources under -`src/gen`. TypeScript compilation includes those sources in `dist`, so package -consumers do not run code generation. - -Branch checks run `mise run sdk:ts:ci`, enforce an 80% line-coverage floor, and -exercise version stamping plus `npm publish --dry-run`. Tagged releases publish -`@nvidia/openshell-sdk` to GitHub Packages. The repository keeps package version -`0.0.0`; the release task derives and temporarily stamps the npm version from -the release tag. - -## CI and E2E - -Required checks run on GitHub Actions. Pull-request workflows that use NVIDIA self-hosted runners trigger from copy-pr-bot mirror branches, so trusted PRs are mirrored into `pull-request/` branches before those workflows run. `main` also uses GitHub merge queue so the final queued integration commit is validated before it merges. - -For PRs that need manual admission, copy-pr-bot accepts `/ok to test ` -only from the explicit `vetters_override` list in `.github/copy-pr-bot.yaml`. -This list is maintained separately from the bot's automatic PR trust policy. - -The high-level CI model: - -1. PR-context gate jobs publish required statuses for the PR head commit. -2. Standard branch checks run from trusted mirror branches. -3. Label-gated Docker, Podman, VM, GPU, and Kubernetes E2E checks run from - trusted mirror branches. -4. Merge-group checks run against GitHub's temporary queue branch for the final integration state. -5. Gate jobs verify that the mirror branch matches the PR head, or that the merge-group workflow ran for the queued SHA, and that the expected non-gate workflow actually ran. -6. Release workflows rebuild and publish binaries, wheels, images, and docs. - -Repository CI keeps telemetry compiled into release-parity artifacts but -disables emission for Rust tests, E2E runs, and release canaries. This prevents -synthetic activity from contributing to product usage metrics. - -Static security checks are deliberately outside the mirror-branch path. They run -directly on GitHub-hosted runners and none of them consume NVIDIA self-hosted -capacity. The change-oriented ones receive no secrets, so they also cover fork -pull requests. Codex Security release qualification is the exception: it needs a -scoped API key, which routes its model calls to NVIDIA-hosted inference while -the job itself stays GitHub-hosted. That placement is load-bearing rather than -incidental: on the repository self-hosted runner the scan agent executes no -shell commands at all, so its preflight never scopes the diff and it seals no -draft. Scanner jobs request `security-events: write` and upload SARIF to Code -Scanning directly on every event they run on, including fork and Dependabot -pull requests, which Code Scanning permits for -`pull_request` runs despite their read-only `GITHUB_TOKEN`. No privileged -intermediate workflow relays those uploads. Manually dispatched Codex Security -runs are the one opt-in exception, described below. Report retention differs by -scanner: Actionlint, Zizmor, and CodeQL keep their reports as workflow artifacts, -and Codex Security keeps no raw report. -Triggers differ by workflow: `.github/workflows/workflow-security.yml` runs on -`pull_request`, `merge_group`, `main`, and a weekly schedule; -`.github/workflows/dependency-review.yml` runs on `pull_request` and -`merge_group` only, because it needs a base and head commit to compare; -`.github/workflows/codeql.yml` runs nightly on the default branch (`main`) via -`schedule`, with `workflow_dispatch` kept for manual diagnostics; and -`.github/workflows/codex-security.yml` is called by the aggregate release scan -for pre-release tags and is also callable through `workflow_call` and -`workflow_dispatch`. CodeQL does -not run on `pull_request`, `merge_group`, or pushes to `main`, so it reports -repository-level Code Scanning state on the default branch instead of per-PR -results, and its four-language matrix stays off the per-change critical path. -Codex Security is release-scoped rather than change-scoped, so it never runs on -a pull request or merge group. - -- **Actionlint and Zizmor** analyze the workflow definitions themselves. - Repository configuration lives in `.github/actionlint.yml` (self-hosted runner - labels, scoped per-file ignores) and `.github/zizmor.yml` (scoped rule - suppressions). Zizmor runs offline and reports only High severity, which is - its maximum level. Both publish SARIF to Code Scanning and retain report - artifacts. The Nix flake provides both scanners, so local runs use - `nix develop --command actionlint -shellcheck= -pyflakes=` and - `nix develop --command zizmor --offline --persona=regular --min-severity=high --no-exit-codes .`. -- **Dependency Review** compares the base and head dependency graphs. It - preflights the GitHub Dependency Graph compare API and neutralizes itself with - a warning while that repository feature is unavailable, so the check begins - reporting on its own once the feature is enabled. Reviews run in warn-only - mode. -- **CodeQL** analyzes product Rust code, examples, and the Go, Python, and - TypeScript SDKs, scoped by `.github/codeql/codeql-config.yml`. Rust test code - is excluded in two layers: the analyze job sets - `CODEQL_EXTRACTOR_RUST_OPTION_CARGO_CFG_OVERRIDES=-test` so the extractor skips - `#[cfg(test)]` blocks, and `paths-ignore` drops `crates/*/tests`, whose - integration targets the cfg override does not reach. Examples remain in scope, - and E2E test code stays excluded because `e2e/` is not an analyzed path. Only - Go requires a build; the other languages use build mode `none`. Analysis runs - on the nightly schedule or by manual dispatch. Results are uploaded to Code - Scanning and always retained as workflow artifacts. -- **Codex Security** qualifies release candidates rather than individual - changes. The job installs a pinned `@openai/codex-security` release into the - runner temp directory before the repository is checked out and invokes it by - absolute path, so repository-controlled files cannot shadow the scanner. Model - calls go to NVIDIA-hosted inference at `https://inference-api.nvidia.com/v1`, - declared as a custom Codex provider named `nvidia` that uses the Responses - wire API with WebSockets disabled. The scan runs `openai/openai/gpt-5.6-sol` - at `medium` reasoning effort, with the multi-agent runtime capped at eight - concurrent threads through - `features.multi_agent_v2.max_concurrent_threads_per_session`. The - `CODEX_SECURITY_API_KEY` secret holds the - NVIDIA key and is exposed to the scan step alone, as `OPENAI_API_KEY` so the - CLI selects API-key auth and as `NVIDIA_INFERENCE_API_KEY`, the provider - `env_key` read by the Codex child process. `CODEX_SECURITY_STATE_DIR` and - `SCAN_DIR` are suffixed with `github.run_id` and `github.run_attempt` and - created mode `700`, so no scanner state or result set from a previous run or - retry attempt is reused even on a runner with a reusable temp directory. - `tasks/scripts/codex_security_range.py` resolves the scan range, reusing the - tag parsers in `tasks/scripts/release.py` so both stay on one definition of a - release tag while requiring the `v` prefix that a release workflow needs. The - job stages both files out of the workspace from the workflow's own revision - and runs the resolver by absolute path, because a scanned candidate predates - them and a revision under scan must not choose its own scan range. The range - itself is resolved against the checked-out candidate: the - candidate must be a `vX.Y.Z-pre.N` tag that is an ancestor of `origin/main`, - and the base is the newest stable `vX.Y.Z` tag merged into the candidate that - is strictly older than the release train `vX.Y.Z` the candidate targets. A - full-repository scan is only possible when no such stable tag exists and the - caller passes `allow_full_bootstrap`. Each candidate scans the cumulative - stable-to-candidate diff, so later candidates re-cover earlier ones. SARIF is - uploaded against `refs/heads/main` at the candidate commit under the - train-scoped category `codex-security/vX.Y.Z`, which makes each candidate's - analysis replace the previous one for that train. Automatic pre-release tag - pushes and `workflow_call` runs always upload. `workflow_dispatch` runs still - perform the scan and the SARIF export, but skip the Code Scanning upload - unless the caller sets the `upload_sarif` input, so manual diagnostics do not - overwrite a train's published analysis by default. Codex Security 0.1.24 - cannot apply `--max-cost` to a slash-qualified model identifier, so the run - has no CLI-enforced cost ceiling. Spend is bounded instead by the 120-minute - job timeout, a single repository-wide concurrency group that serializes - qualification so starting a newer candidate cancels an in-flight one, and - NVIDIA account-side controls. No raw report is retained. -- The job clears `kernel.apparmor_restrict_unprivileged_userns` before - installing the scanner. Codex confines model-run commands with bubblewrap, - which needs unprivileged user namespaces; Ubuntu 24.04 restricts those through - AppArmor, so bubblewrap fails to configure the sandbox network namespace - (`bwrap: loopback: Failed RTM_NEWADDR`) and the agent executes no commands at - all. The failure is silent: the agent retries its shell tool, gives up, and - seals no draft, while the scanner only reports a missing or incomplete draft. - Lifting a kernel restriction on the runner is what allows the sandbox that - confines the agent to start, and the runner is ephemeral and GitHub-hosted. -- The scan sets `approval_policy="never"`. Codex Security keeps - `approvals_reviewer="auto_review"` unconditionally, and that reviewer runs on - its own model rather than the configured one. Because the workflow declares a - single provider that serves only `openai/openai/gpt-5.6-sol`, any approval - request reaches a model the endpoint does not serve, so the agent never gets a - shell command approved and seals no draft. The scan stays confined by its - `workspace-write` sandbox with network access disabled and by the scanner's - own permission profile, which grants read access to the filesystem root and - write access only to the workspace roots. -- A scan that cannot execute commands reports only a missing or incomplete - draft, so diagnosing one means reading the scanner's session rollouts under - `CODEX_SECURITY_STATE_DIR`, where every shell command the agent ran is - recorded. No command at all is the signal that the sandbox failed to start. - -Findings never fail these checks; scanner and build failures do. A scanner that -cannot run, a CodeQL analyzer that does not complete, an unexpected Dependency -Graph API error, and a Codex Security range, scan, or export failure are all -errors, which keeps an informational check from silently degrading into a no-op. -Codex Security also rejects any scan scope other than the resolved -cumulative diff or an approved full bootstrap, so a qualification run either -covers the whole stable-to-candidate range or fails; a separate no-permission -job republishes the analysis job's outcome as the -`OpenShell / Codex Security (informational)` status. None of these checks are -required statuses, so they do not gate merges. - -The workflow only reports on candidates that already exist. It incrementally -implements the qualification model from -[RFC 0014](../rfc/0014-release-stability/release-qualification.md): failed -checks do not prevent publication of an immutable pre-release candidate. The -summary explicitly records that the current profile does not yet provide the -RFC's complete qualification coverage. - -The tagged release workflow calls the aggregate Security Scan after publishing -the candidate's commit-addressed gateway, sandbox, and supervisor images. CodeQL, -Trivy, Cargo Deny, and Actionlint/Zizmor run for every release tag; Codex Security -also runs for pre-release tags. Scanner failures, Cargo Deny advisories, and -High or Critical Codex Security findings fail qualification. CodeQL, Trivy, and -Zizmor findings are temporarily informational while the findings accepted for -v0.1.0 are addressed in 0.1.x releases. - -The `Release Qualification` job aggregates security, conformance, feature, -Docker E2E, and VM E2E results. The currently implemented profile gates stable -publication, but it does not represent complete RFC 0014 qualification. For a -pre-release it records a failed result without blocking artifact assembly, -image tagging, or Helm publication; the failing underlying suite keeps the -workflow visibly red. Every attempt writes a summary to the Actions run summary -and a 90-day Actions artifact. After release assembly succeeds, the workflow -publishes the same result to -`ghcr.io/nvidia/openshell/qualification:-run--attempt-`. -`tasks/scripts/generate-qualification-summary.sh` generates qualification -metadata only. Artifact identity remains the responsibility of the separate -release manifest. Including both the run ID and attempt preserves the result of -each rerun. - -`release-auto-tag.yml` runs at 14:00 Europe/Zurich on weekdays (including daylight -saving time changes) and supports manual dispatch. Maintainers start weekday -pre-release publishing by tagging the initial `vX.Y.Z-pre.1` release candidate. -The workflow increments the highest release series' pre-release number on `main` -only when that seed exists, its stable tag does not exist, and new commits are -available. It never chooses a minor or patch version or creates the initial seed. -After pushing the tag, it explicitly dispatches `release-tag.yml` to build the -candidate. - -See `CI.md` for the contributor workflow, labels, and maintainer merge-queue workflow. - -## Docs Site - -Published docs live in `docs/`. Navigation lives in `docs/index.yml`. Fern site -configuration, components, theme assets, and publish settings live in `fern/`. - -Use `mise run docs` for Fern validation and navigation-to-file-path consistency, -and `mise run docs:serve` for local preview. The docs PR workflow also runs the -navigation check's unit tests (`mise run test:docs-nav`). -PR previews are produced by `.github/workflows/branch-docs.yml` when -Fern credentials are available. Production docs publish from the release tag -workflow. Redirect rules follow the mutable snapshot that owns their source URL -(or destination for unversioned aliases). Syncing replaces that channel's rules, -including deletions; `dev` owns shared fallback rules. Stable promotion updates -`latest` routing together with its content, while older maintenance releases -preserve both. - -## Validation Expectations - -- Run `mise run pre-commit` before committing. -- Run `mise run test` after code changes. -- Run `mise run e2e` for sandbox, policy, driver, or deployment changes when the - affected runtime can be exercised. -- Run `mise run ci` before opening a PR when practical. -- Run `mise run docs` when `docs/` or `fern/` changes. - -Architecture-only changes should still check links and references because this -directory is used by agents during implementation and review. diff --git a/architecture/compute-runtimes.md b/architecture/compute-runtimes.md deleted file mode 100644 index ec6fff115a..0000000000 --- a/architecture/compute-runtimes.md +++ /dev/null @@ -1,607 +0,0 @@ -# Compute Runtimes - -Compute runtimes create, stop, start, delete, and watch sandbox workloads for the -gateway. A supported runtime provisions `openshell-sandbox` inside the workload, -`openshell-supervisor` outside it, a protected channel between them, and an -independent outer network fence. Drivers do not implement policy evaluation. - -Podman provisions a paired workload and supervisor container using its native -libpod API. The workload uses `network=none`; the external supervisor alone joins -the configured network. A per-sandbox named volume carries their mutually -authenticated gRPC Unix socket, with supervisor credentials kept in its separate -filesystem. Both containers run as the resolved non-root identity with all -capabilities dropped. They share only a user namespace for volume ownership, -not PID, mount, or network namespaces. Podman owns paired lifecycle and health; -the common protocol owns process, identity, TCP, DNS, and forwarding semantics. - -## Driver Contract - -External resource admission is an operator-owned boundary shared by drivers. -The gateway gates caller driver JSON independently from attachment approval. -Drivers resolve the complete effective attachment inventory against authoritative -resource labels before launch and on reuse. Missing labels or an unsupported -resolver deny access; GPU attachments are an explicit temporary exception. -Fresh sandbox-private resources instead require verified provisioning ownership. -Workload metadata must not grant approval or override admission evidence. - -The shared evaluator lives in `openshell-core`; native resolution remains in -each driver. External drivers acknowledge the effective versioned policy through -capabilities, and policy mismatch prevents activation or new launch operations. -Trusted deployment configuration can explicitly disable label admission, but -that opt-out does not waive other ownership and isolation checks. This boundary -assumes operators control approval metadata and runtime resource replacement; -it does not provide atomic mount authorization or instantaneous revocation. - -Each runtime receives a sandbox spec and canonical policy from the gateway and -is responsible for: - -- Selecting the sandbox image. -- Resolving an immutable non-root sandbox identity before workload creation. -- Supplying separate sandbox and supervisor bootstrap material. -- Delivering `openshell-sandbox` to the workload and `openshell-supervisor` only - to the external supervisor placement. -- Provisioning protected control and boundary configs plus a private Unix socket, - TLS-authenticated TCP, or vsock transport when the supervisor is separated. - Runtime-specific code supplies immutable resource claims and transport - coordinates; the shared boundary protocol supplies lifecycle, exec, signaling, - forwarding, and binary identity semantics. -- Forwarding the exact canonical main-process argv and TTY mode without shell - reconstruction. The sandbox-level environment and policy workspace apply to - the main process. -- Reporting lifecycle and platform events back to the gateway. -- Cleaning up runtime-owned resources. - -Drivers report **runtime-observed state only** and must not hold references to -gateway-internal types. For supervisor-controlled runtimes, `Ready=True` means -only that the compute resource is healthy; the gateway also requires a -supervisor session before publishing `SandboxPhase::Ready`. For -drivers that report runtime readiness, `Ready=True` is authoritative because the driver -launches and monitors the policy-constrained workload itself. - -`compute_driver.proto` is the supported gateway/driver extension boundary. -At initialization the gateway snapshots the driver's identity, version, -default image, gateway-lifecycle preference, and -`driver_reports_runtime_readiness` from `GetCapabilities`. The gateway includes -the canonical `SandboxPolicy` in `DriverSandboxSpec.policy` for validation and -creation. Drivers that enforce policy outside the standard supervisor fetch -later revisions through `GetSandboxConfig` and acknowledge them through -`ReportPolicyStatus`. -Process-identity omissions are preserved across this boundary so every driver -can apply its native image or runtime defaults. Drivers connect supervisors to -the operator-configured gateway endpoint; they do not request additional -gateway listeners. - -Canonical main-process support is part of the `ComputeDriver` contract. Every -in-tree and extension driver must forward the exact specification; it is not an -optional capability that drivers can omit or negotiate. - -Drivers own runtime-specific platform event interpretation. When an event should -drive client provisioning UI, the driver attaches the shared -`openshell.progress.*` metadata defined in `openshell-core` instead of requiring -clients to parse Kubernetes reasons, VM cache states, or other driver-local -reason strings. - -## Sandbox Readiness Composition - -The gateway composes driver state with the advertised readiness behavior to -produce the public `SandboxPhase`: - -``` -backend_phase = derive_phase(driver_status) - -public_phase = - if backend_phase in {Error, Deleting}: → pass through (terminal precedence) - if driver_reports_runtime_readiness && backend_phase == Ready: → Ready - if backend_phase == Ready && session connected: → Ready - if backend_phase == Ready && no session: → Provisioning - if backend_phase in {Provisioning, Unknown} && session: → Ready - if backend_phase in {Provisioning, Unknown} && no session: → Provisioning -``` - -For a supervisor-controlled runtime, `public_phase == Ready` means both the -backend resource is healthy and a supervisor session is registered. A sandbox whose -backend reports ready but has no supervisor session yet holds `Provisioning` with a -`Ready=False`, `SupervisorNotConnected` condition and the message -`Backend ready; waiting for supervisor session`. This distinguishes it from a sandbox -whose compute resource is still provisioning without exposing contradictory public -readiness signals. When the driver reports runtime readiness, its ready condition -is published without waiting for a supervisor session. - -The supervisor keeps retrying session establishment while the gateway is unavailable. -Its control readiness socket remains absent until the gateway accepts a session and -is removed if that session disconnects. A transient gateway delay during startup -therefore leaves the sandbox provisioning without terminating the supervisor. - -**Session precedence over lagging driver snapshots:** A supervisor session can only be -established by a running workload. When `set_supervisor_session_state` promotes the -store record to `Ready` on session connect, a driver watch event may still arrive -shortly after carrying a stale `Provisioning` or `Unknown` backend phase. The -composition rule treats a connected session as the stronger signal and keeps `Ready` -in that case, preventing a lagging snapshot from undoing the session-driven promotion. - -**HA session composition:** Live relay handles remain process-local, while the -session-owning gateway publishes a short-lived owner record in shared PostgreSQL. -Driver reconciliation treats a fresh local or remote owner as connected, so a -non-owner replica cannot demote the shared sandbox phase merely because it lacks the -in-memory stream. Session-bound requests are forwarded to the owning gateway; a -supervisor reconnect publishes a higher connection epoch before stale-session cleanup -can demote readiness. Multi-replica deployments therefore require shared PostgreSQL -and the gateway peer Service configured by the Helm chart. - -**Extension point:** Driver-reported readiness is a capability, not an -operator-configurable hook. A driver may enable it only when it owns workload -readiness. Policy delivery remains independent: create-time policy is embedded -in the sandbox specification, and later revisions use the existing sandbox -configuration API. RFC-0010 lifecycle hooks may observe readiness transitions via -`post_commit`; they do not override the composition rule. - -The capability RPC reports driver identity, version, and the default sandbox -image used by the gateway. GPU availability stays driver-local and is validated -when a sandbox create request asks for GPU resources. - -The gateway sends its common extension peer metadata with the startup capability -request. The driver validates that metadata before responding, and the gateway -rejects a driver whose protocol major or capability requirements are -incompatible. It records the negotiated protocol, implementation identity and -version, capability sets, and typed resource support once. Elevated gateway info -reports that immutable snapshot instead of re-querying drivers on each request. - -## Compiled Driver Selection - -The gateway binary explicitly installs the compute drivers compiled into that -binary before entering server startup. The server selects a configured driver -by normalized registry name. When no driver is configured, it evaluates only -the installed drivers' probes and chooses the lowest registered priority. -Drivers without a probe, including VM, remain opt-in. - -Startup computes this selection once after merging configuration. The same -selection drives authentication defaults and runtime construction, so a probe -result cannot change which driver is constructed later in startup. - -This follows the same composition model as SQLx's `Any` drivers: the binary -defines the available implementation set, while the runtime consumes a generic -registry. Adding or removing a compiled driver therefore changes registration -rather than the server's selection flow. Alternate gateway binaries can install -their own `ComputeDriverFactory` registrations and hand the completed registry -to `run_cli_with_compute_drivers`; factories receive merged driver config and -return either an in-process driver or a gateway-managed remote endpoint. The -server constructs the common runtime adapter and snapshots `GetCapabilities` -for either result. A configured UDS endpoint still takes precedence over a -compiled registration with the same name. - -The `openshell-gateway` composition crate groups first-party registrations -behind the `in-tree-compute-drivers` feature. `openshell-server` has no compute -driver dependencies or backend-name dispatch. Protocol-only gateway builds -disable the composition feature and link no compute-driver crates. E2E lanes -compose that gateway with Docker, Podman, Kubernetes, and VM driver executables -over the public UDS gRPC contract so an in-tree driver cannot silently depend -on a server-only API. - -## Stop and Start Lifecycle - -The gateway persists lifecycle intent before mutating compute: - -```text -Ready -> Stopping -> Stopped -> Starting -> Ready -``` - -A canonical main process that exits successfully follows `Ready -> Completed`. -A nonzero or signal-normalized result follows `Ready -> Error` with a -`MainProcessFailed` condition. Both retained results may be started explicitly, -which creates a fresh main-process instance. Drivers must not automatically -restart a completed or failed canonical process. Before an explicit restart, -the gateway disconnects the prior supervisor session and deletes its SSH -sessions so credentials cannot cross runtime generations. - -`StopSandbox` and `StartSandbox` are idempotent driver operations. Stop -retains the driver resource and its persistent workspace boundary while making -exec, SSH, forwarding, and exposed services unavailable. Start reactivates the -same resource. The gateway requires a fresh supervisor session before a -starting sandbox returns to `Ready`; stale driver snapshots and supervisor -sessions cannot promote a `Stopped` row. - -Runtime credentials are generation-scoped and memory-only after launch. A -supervisor or Sandbox Runtime process replacement does not resume a running -generation. Planned upgrades stop the sandbox first; the following start mints -a fresh session, TLS identity, and credential pair. An unexpected replacement -leaves the old workload on the normal fail-closed disconnect path. -The gateway commits the new authorization identity and the durable `Starting` -phase in one resource-version update. A concurrent start that loses that update -reuses the winning identity, so every idempotent driver retry receives credentials -that match the persisted sandbox. - -A driver stop operation does not complete while its backend still reports an -in-progress stop. This prevents an immediate start from racing the previous -run's delayed exit event and regressing the new run to `Error`. - -The Kubernetes driver records stop as a durable two-phase transition. The -`releasing` phase releases the sandbox's runtime-control relationship while the -workload boundary remains reachable. The `suspending` phase then suspends the -Agent Sandbox workload and cleans generation bootstrap material. The current -dedicated-supervisor implementation releases control by deleting the supervisor -Pod. Periodic reconciliation resumes either phase after a gateway restart. Pod -deletion waits include the configured termination grace period plus Kubernetes -API observation headroom. - -Persisted `Stopping` and `Starting` rows are retried at startup. Stable -`Stopped` rows remain stopped. Docker and Podman retain the stopped container -and attached storage, Kubernetes retains the Sandbox CR and PVC while scaling -compute to zero, and VM retains its launch request and writable overlay beside -a stop marker. Delete remains a separate operation that removes these -resources. - -On graceful gateway shutdown, persisted running intent for Docker, Podman, and -VM is stopped through the shared `StopSandbox` RPC before any gateway-managed -driver process exits. The gateway does not persist `Stopped` for this -infrastructure event. On startup, it reconciles the retained intent through the -shared idempotent `StartSandbox` RPC before watch processing begins. Explicitly -`Stopped` sandboxes are excluded from both sweeps. Kubernetes workloads are -cluster-owned and continue running without gateway shutdown or startup -lifecycle calls. - -The driver reports this behavior through -`GetCapabilities.gateway_manages_lifecycle`. The same declaration works for -in-process and external drivers. Older drivers omit the field and retain the -conservative operator-managed behavior. - -Drivers that can verify a platform-native sandbox credential advertise -`GetCapabilities.supports_sandbox_authentication`. On the path-scoped -`IssueSandboxToken` exchange, the gateway forwards the opaque bearer credential -to that selected driver through `AuthenticateSandbox`. The driver returns the -authenticated sandbox ID and opaque runtime identity. The gateway verifies -that both match its durable sandbox record and returns a generation-bound -session JWT whose lineage is checked on every subsequent sandbox RPC. Legacy -unbound sandbox JWTs are not admitted when session authentication is enabled. -The driver socket is therefore a sandbox-identity trust boundary, but it does -not grant user or administrator authority. - -## Deletion Lifecycle - -Lifecycle requests use per-sandbox gates to serialize stop, start, and -delete attempts. A delete request -resolves the name once and remains bound to that stable ID. The only -combined lock order is lifecycle gate, then the gateway-wide state guard; external -driver calls run without the global guard. - -Lifecycle gates are process-local and do not coordinate gateway replicas. They -serialize attempts rather than share results: if one attempt fails and recovery -restores a deletable state, a request waiting on the gate may retry the driver. -Persisted resource-version checks remain the cross-replica safety boundary. - -Watcher events do not acquire lifecycle gates. Exact resource-version checks allow -them to interleave safely: status snapshots are no-ops for `Deleting` rows, -deleted events are idempotent, and snapshots for absent rows are ignored. - -An accepted delete (`deleted = true`) is finalized by the watcher. If the -backend is already absent (`deleted = false`), the request removes gateway state -synchronously. Sandbox row removal remains bound to the stable ID and resource -version. Settings retain their existing best-effort name-based cleanup; SSH -sessions, indexes, and watch/log buses are cleaned after confirmed removal. -Owned-record cleanup discovers records before mutating them and uses bounded -set-based deletes so teardown cannot amplify one sandbox into an unbounded -sequence of individual persistence writes. - -When a sandbox is instead discovered gone out-of-band — a watcher deletion -event, or the periodic prune sweep finding no matching driver resource, with -no explicit `DeleteSandbox` request involved at all — the gateway also -releases driver-owned resources (for example Podman's per-sandbox secrets and -workspace volume) by calling the driver's idempotent `DeleteSandbox`, not just -gateway state. Both paths skip that call when a request-side lifecycle -operation already holds the sandbox's gate, since that operation already owns -driver-side cleanup. The watch path defers the call itself to a background -task after a non-blocking gate check, so a slow driver call cannot stall the -sequential watch loop; the prune sweep calls the driver inline, since it -already makes a blocking `GetSandbox` call per sandbox as part of its normal -operation. - -The request acquires both locks before starting owned work, so cancellation -while queued does not leave a delete armed. After that commitment point, the -owned task prevents cancellation from stranding a mutation. A gateway restart -does not start a persisted `Deleting` operation. If the backend completed the -delete, reconciliation removes the row; otherwise it can remain `Deleting`. - -## Runtime Summary - -| Runtime | Best fit | Sandbox boundary | Notes | -|---|---|---|---| -| Docker | Local development with Docker available. | Capability-free workload container. | Uses `network_mode=none`; a separate capability-free supervisor container mediates egress and access over a private daemon-local Unix socket volume. | -| Podman | Existing rootless driver. | Container. | Not converted by this isolation stack. | -| Kubernetes | Cluster deployment through Helm. | Capability-free sandbox Pod. | Always creates a namespace-wide empty-egress workload NetworkPolicy and a separate capability-free supervisor Pod over mutually authenticated TLS. It requires an enforcing CNI and trusted sandbox namespace; the Kubernetes API does not attest policy enforcement. | -| VM | Experimental microVM isolation. | Per-sandbox libkrun or QEMU VM. | The NIC-less guest runs `openshell-sandbox` as PID 1; host `openshell-supervisor` owns gateway networking and reaches the guest over vsock. | -| Extension | Out-of-tree drivers operated alongside the gateway. | Whatever boundary the driver implements. | Selected by a custom `compute_drivers = [""]` entry with `[openshell.drivers.].socket_path`, or at launch time by pairing `--drivers ` with `--compute-driver-socket=`. A launch-time endpoint may use a canonical built-in name to preserve its driver-config key while replacing in-process construction. The gateway connects to an operator-provisioned UDS, snapshots `GetCapabilities`, and dispatches all sandbox lifecycle calls through `compute_driver.proto`. The driver process and socket lifecycle are operator-owned; the gateway does not spawn, supervise, or remove unmanaged extension drivers. The trust boundary is the socket's filesystem permissions: the operator must ensure only the gateway uid can read/write it. | - -Per-sandbox CPU and memory values currently enter the driver layer through -template resource limits. Docker and Podman apply them as runtime limits. -Kubernetes mirrors each limit into the matching request. VM accepts the fields -but currently ignores them. - -Reusable sandbox workload templates are resolved before the compute-driver -boundary. Drivers do not receive a separate template resource; the gateway -lowers the selected `SandboxWorkloadTemplate` into the existing sandbox spec -and validates that spec before calling `ValidateSandboxCreate` or -`CreateSandbox`. Template CPU and memory become the same typed resource limits -described above. Template GPU settings become `ResourceRequirements`, preserving -the driver's default GPU assignment when the count is omitted. Template -`driver_config` remains a driver-keyed envelope until the compute layer selects -the active driver block and forwards only that block to the driver. - -Docker and Podman also accept per-sandbox driver-config mounts for existing -runtime-managed named volumes and tmpfs mounts. Podman additionally accepts -image mounts through its image-volume API. User-supplied bind and volume mounts -default to read-only. Direct host bind mounts, and Docker or Podman local-driver -bind-backed named volumes, are available only when explicitly enabled in the -active local driver table of `gateway.toml`. Host bind mounts are an unsafe -operator override because they place gateway-host filesystem state inside the -sandbox and can negate OpenShell workspace isolation and filesystem-policy -controls. Driver-owned supervisor, token, and TLS bind mounts stay reserved. - -Network features follow the driver/substrate split. Drivers own only the outer -fence and protected channel. The sandbox owns seccomp notification, local DNS, -socket virtualization, process observation, and binary identity. The supervisor -owns DNS eligibility, policy authorization, destination filtering, upstream -dials, relay behavior, credential rewriting, and OCSF decisions. No supported -path requires nftables, a workload network namespace, proxy environment -variables, added capabilities, or an unconfined AppArmor profile. - -The Kubernetes deployment packaging has two ownership boundaries. The gateway -chart owns the gateway workload, configuration, Services, PKI, and -cluster-scoped gateway resources. The workspace chart is installed into a -pre-provisioned sandbox namespace and owns only the sandbox ServiceAccount, -namespaced RBAC, and sandbox ingress NetworkPolicy. Its RoleBinding names the -gateway ServiceAccount and namespace explicitly, so the two releases have -disjoint lifecycle ownership. A shared-mode gateway can target one external -namespace, while operator mode maps workspace names to multiple -platform-provisioned namespaces. - -Resource requirements enter the driver layer through `SandboxSpec.resource_requirements`. This includes a set of GPU requirements, where a user -can request a specific number of GPUs or the driver-specific default behaviour. -For all in-tree drivers, this is equivalent to selecting a single GPU. - -VM runtime state paths are derived only from driver-validated sandbox IDs -matching `[A-Za-z0-9._-]{1,128}`. The gateway-owned VM driver socket uses a -private `run/` directory plus Unix peer UID/PID checks. Standalone -unauthenticated TCP mode is disabled unless explicitly enabled for local -development. The VM image cache is owner-only. When the host assembles a -bootstrap rootfs from OCI layers, cross-layer symlinks may resolve only within -that rootfs; absolute or escaping targets reject the image before a later layer -can write through them. Rootfs traversal and mutation use opened directory -handles with no-follow file creation, so later copies, permission changes, and -whiteouts cannot be redirected by replacing a validated pathname component. -The bootstrap rootfs comes only from the operator-configured `bootstrap_image` -or `default_image`; a sandbox-requested image never becomes the VM bootstrap -image. The gateway rejects configurations without either trusted source, and -the standalone driver independently fails startup for the same condition. - -Runtime-specific implementation notes belong in the driver crate README: - -- `crates/openshell-driver-docker/README.md` -- `crates/openshell-driver-podman/README.md` -- `crates/openshell-driver-kubernetes/README.md` -- `crates/openshell-driver-vm/README.md` - -The VM guest bootstrap runs once as root to prepare mounts, loopback, and the -safe port-53 sysctl. It then drops to the resolved identity with empty -capability sets and executes `openshell-sandbox` as guest PID 1. - -## Supervisor Delivery - -Drivers deliver the two binaries to separate trust domains: - -| Runtime | Delivery model | -|---|---| -| Docker | A digest-pinned daemon-local volume supplies `openshell-sandbox`; the companion image runs `openshell-supervisor`. | -| Podman | Existing driver behavior; not converted by this stack. | -| Kubernetes | A non-root init container stages `openshell-sandbox` into a memory volume; a directly managed Pod runs `openshell-supervisor`. | -| VM | `openshell-sandbox` is embedded in the guest rootfs; a separately digest-checked native `openshell-supervisor` runs on the host. | -| Extension | Defined by the out-of-tree driver. | - -Driver-controlled sandbox bootstrap must override image or template values for -sandbox identity, command metadata, resolver configuration, and public trust -paths. Gateway endpoints, callback credentials, policy, and private TLS material -belong only to the supervisor placement. - -## Process Identity - -The gateway preserves whether each policy process field was omitted and passes -the admitted selectors to the driver. The driver resolves one exact UID, GID, -and supplementary-group set before creating the immutable workload: - -- Docker pins the image ID, resolves policy selectors against the image's - `/etc/passwd` and `/etc/group`, and validates its OCI working directory. -- Kubernetes uses platform-resolved numeric values, including OpenShift - namespace ranges. -- VM uses the configured numeric guest identity. - -UID/GID zero and `u32::MAX` are invalid. The sandbox and every child start with -the resolved identity and zero capability masks; neither process performs an -in-workload UID transition. Identity-changing policy updates require sandbox -recreation, while other policy updates remain live. - -Docker uses an absolute OCI working directory as the workspace. Empty, root, -and explicit `/sandbox` values select `/sandbox`; other paths must already -exist without symlink or reserved-mount collisions and must be usable by the -resolved identity. Kubernetes and VM use `/sandbox`. - -### Executable Identity Binding - -Every mediated network open carries a connection-bound `BinaryIdentity` from -the isolation backend. The identity contains the socket-owning executable and -each executable ancestor, nearest first, as an absolute workload path plus a -SHA-256 digest. The backend resolves these values for the accepted connection -and hashes already-open live executable objects rather than reopening their -paths. Command-line paths remain diagnostic context and cannot authorize a -request. - -Before policy evaluation, the supervisor validates and pins the complete leaf -and ancestor chain in one runtime-scoped trust-on-first-use cache. It rejects -missing digests, invalid paths, conflicting evidence within a chain, or a -digest that differs from an existing path pin. Validation and insertion are -atomic, so a rejected chain cannot leave partial pins. - -The cache is shared across authorization paths for the lifetime of the -supervisor network runtime. Policy reloads replace policy state without -clearing executable pins; restarting the runtime creates a new cache. OPA -receives executable and ancestor paths plus endpoint policy context. Digests -remain supervisor-side integrity evidence and are not policy inputs. Any -unavailable, incomplete, or conflicting executable evidence fails closed -before OPA can authorize the connection. - -The Kubernetes driver creates the namespace-wide empty-egress workload fence -before a suspended Sandbox CR, then provisions split immutable bootstrap -Secrets, the private runtime Service, and a gated supervisor Pod. A -non-root init container stages `openshell-sandbox` and one-use bootstrap files -into memory volumes. The workload Pod never mounts supervisor or gateway -credentials. The driver removes its scheduling gate only after the companions -exist; measured confirmation and supervisor-session registration gate public -readiness. - -## Images - -The gateway image and Helm chart are built from this repository. Users supply -workload images as standard OCI images. - -Custom sandbox images must include the agent runtime and any system -dependencies, but they should not need to include the gateway. GPU-capable -images must include the user-space libraries required by the workload. The -runtime still owns GPU device injection. GPU requests are explicit, and can be -refined with a driver-native device identifier or requested count; the gateway -validates the request shape and each runtime enforces the GPU allocation modes it -supports. - -## Deployment Shape - -Kubernetes deployments use the Helm chart under `deploy/helm/openshell`. The -chart deploys the gateway and sandbox runtime integration. The default gateway -workload is a StatefulSet for SQLite-backed single-replica installs. External -database-backed installs can render a Deployment with `workload.kind=deployment`; -HA deployments must point `server.externalDbSecret` at an operator-managed -PostgreSQL database. Agent Sandbox CRDs and controller lifecycle remain -operator-owned; the chart can optionally preflight for a served supported API -but does not install the cluster-scoped dependency. OpenShell's Kubernetes test -clusters install the upstream core manifest; Agent Sandbox extensions are not -required by the gateway. -Standalone local deployments start the gateway with a selected runtime such as -Docker, Podman, or VM. The CLI can register multiple gateways and switch between -them without changing the sandbox architecture. - -## Workspace Namespace Modes (Kubernetes) - -The Kubernetes driver maps workspaces to namespaces through the `workspace_mode` -configuration field (`WorkspaceMode` in `crates/openshell-driver-kubernetes/src/config.rs`). -The mode controls namespace resolution, resource naming, sandbox CR watching, SA -token authentication, and RBAC requirements. - -| Mode | Namespace resolution | Resource name | Namespace lifecycle | -|---|---|---|---| -| **Shared** (default) | Single static namespace from config | `{workspace}--{name}` | None | -| **Managed** | `openshell-{gateway_id}-{workspace}` | bare sandbox name | Driver creates and deletes | -| **Operator** | Workspace name maps 1:1 to a pre-provisioned namespace | bare sandbox name | External (platform team) | - -**Shared** renders all sandboxes into one configured namespace. Resource names -embed the workspace prefix for collision avoidance. No namespace lifecycle -management. RBAC uses a namespace-scoped Role. - -**Managed** auto-creates a K8s namespace per workspace on first sandbox create. -Each new namespace receives a ServiceAccount and the configured gateway-only -SSH ingress NetworkPolicy. Each sandbox runtime generation gets immutable copies -of the configured image-pull Secrets, read from the driver's source namespace and -named after the generation, so a sandbox picks up rotated registry credentials -on its next start. Their sources are operator-selected gateway configuration, -not caller attachments. An existing Secret with a generation name fails the -create and is never adopted. The namespace also copies -OpenShift SCC UID-range and supplemental-group annotations from the gateway -namespace when present. The driver deletes the namespace during workspace -deletion. The workspace remains durably `Terminating` until the Kubernetes API -accepts namespace cleanup, so a transient failure can be retried. Namespace -deletion uses the fetched UID as a -precondition to avoid deleting a replacement namespace. Requires a non-empty -`gateway_id` (validated as a -DNS-1123 label at startup) so the namespace prefix fits within the K8s 63-character -limit. RBAC promotes sandbox CRD permissions to a ClusterRole and adds namespace -`create`/`delete` and ServiceAccount `create`/`get` permissions. - -Outside shared mode, the gateway client TLS material is staged into each -generation's supervisor bootstrap Secret rather than mounted from a Secret in -the workspace namespace. Every Secret the driver writes into a workspace -namespace is therefore generation-scoped, immutable, and created with `create` -only. RBAC cannot constrain `create` by `resourceNames`, so managed mode grants -cluster-wide Secret `create` and `delete`; source reads use a Role in the -driver's source namespace. Recovery deletes the Secrets of the recorded and -target runtime generations by exact name; Pod owner references let garbage -collection remove any other generation. The -driver exercises these broad permissions only in gateway-owned managed -namespaces. This depends on the managed-mode ownership invariant described below; -the gateway ServiceAccount must not be shared with unrelated workloads. - -Operator mode does not create NetworkPolicies or copy image-pull Secrets. -Platform teams must apply the gateway ingress boundary and provision configured -image-pull Secrets in every operator-managed namespace. -The gateway ClusterRole grants no Secret permissions in operator mode. The -`openshell-workspace` chart Role installed in each operator-managed namespace -grants bootstrap Secret `create` and `delete`. - -**Operator** uses pre-provisioned namespaces discovered through two optional -sources: a K8s label selector (`operator_namespace_label`) and a drop-in -allowlist file (`operator_namespace_file`). Exactly one must be configured. -The compute driver and the gateway's ServiceAccount authenticator independently -watch that public config source; no in-process driver state crosses into the -server. Sandbox creation and token bootstrap fail closed if the workspace is -not in the current allowlist. Platform teams manage namespace lifecycle -externally. RBAC uses the same ClusterRole as managed mode but without namespace -`create`/`delete` or ServiceAccount permissions. - -### Watching and Querying - -Managed and operator modes set `is_multi_namespace() == true`, which switches -sandbox CR watchers from namespace-scoped `Api::namespaced` to cluster-wide -`Api::all_with`. In managed mode the driver scopes cluster-wide queries with a -`LABEL_GATEWAY_ID` label selector to support multiple gateways on the same -cluster. K8s Events are not watched in cluster-wide mode — the cluster-wide -watcher emits only sandbox CR changes, not platform events. - -### SA Token Authentication - -The Kubernetes driver's `AuthenticateSandbox` implementation applies its named -`[openshell.drivers.kubernetes]` configuration per mode: - -- **Shared:** `Exact` — accepts only the single configured namespace. -- **Managed:** `Prefix` — accepts any namespace starting with `openshell-{gateway_id}-`. -- **Operator:** `Allowlist` — accepts namespaces present in the dynamic - `BTreeSet` populated by the label/file watchers. Starts empty (fail-closed) - until the first watcher update. - -It validates the projected token with Kubernetes `TokenReview`, checks the live -pod UID, and verifies the pod's controlling Sandbox CR UID and sandbox ID. The -driver returns both the sandbox ID and an opaque runtime identity derived from -the namespace, immutable Sandbox CR UID, and authenticated supervisor Pod UID. -Advertising sandbox authentication includes the runtime-binding contract. The -gateway requires non-empty runtime identities from successful create, start, -and authentication responses. It records the runtime identity when provisioning -succeeds and requires an exact match before issuing a sandbox JWT. If binding -validation or storage fails after a lifecycle call succeeds, the gateway -compensates that call before returning the error. This correlates credential -authentication with the durable runtime record rather than authorizing from the -sandbox ID alone. - -`StartSandbox` carries the previously recorded opaque identity. Kubernetes -requires exactly one label-selected Sandbox CR and verifies that its namespace -and immutable UID match that identity before replacing the supervisor Pod. The -new Pod UID becomes the updated binding only after the continuity check passes. - -Shared and managed modes still reserve the sandbox namespace, Sandbox CRs, -sandbox pods, and configured sandbox ServiceAccount for the Kubernetes driver -and trusted Agent Sandbox controller. In operator mode, the platform operator -retains namespace lifecycle ownership and must preserve the same control of -those resources. An allowlisted namespace is a trust grant, not a tenant -isolation boundary. - -### Credential Driver Integration - -The Kubernetes Secrets credential driver (`openshell-driver-kubernetes-secrets`) -stores every provider credential in its single configured namespace, in every -workspace mode, and rejects handles that reference another namespace. The -gateway reaches those Secrets through a namespaced Role; the gateway -ClusterRole grants no credential Secret permissions. - -When runtime infrastructure changes, validate the relevant sandbox e2e path and -update the matching driver README if a maintainer-facing constraint changes. diff --git a/architecture/gateway.md b/architecture/gateway.md deleted file mode 100644 index 1e3b761d2d..0000000000 --- a/architecture/gateway.md +++ /dev/null @@ -1,1336 +0,0 @@ -# Gateway - -The gateway is the OpenShell control plane. It exposes the API used by the CLI, -SDK, and TUI; persists platform state; manages provider credentials and -attachments; and asks compute runtimes to create or delete sandbox workloads. - -## Responsibilities - -- Authenticate clients and sandbox supervisor sessions. -- Serve gRPC APIs for sandbox lifecycle, provider management, policy updates, - settings, logs, watch streams, and relay forwarding. -- Serve HTTP endpoints for health and edge-auth flows, plus the opt-in - WebSocket tunnel for edge-proxy deployments. -- Persist domain objects in SQLite or Postgres. -- Resolve endpoint-bound provider environments for sandbox supervisors. -- Coordinate supervisor relay sessions for connect, exec, file sync, and - service forwarding. -- Persist the canonical main-process instance ID and normalized exit code on - sandbox status. Exit code zero transitions the sandbox to `Completed`; - nonzero results transition it to `Error/MainProcessFailed`. Infrastructure - failures also use `Error`, with a distinct reason and no fabricated command - result. - -The gateway does not enforce agent network policy at request time. That happens -inside each sandbox, where the supervisor and proxy can observe local process -identity. - -The live supervisor session is the readiness authority for its main-process -instance. The supervisor reports its normalized result through the -sandbox-authenticated `ReportMainProcessExit` RPC, and the gateway rejects -results from stale instance IDs. Foreground creation carries a one-shot -attachment intent to the process supervisor. The supervisor durably reports the -result immediately, accepts that declared SSH attachment even when the process -has already exited, sends the retained output and exit status, and waits for the -peer's channel close before finalizing the result for ephemeral cleanup. -Detached commands carry no attachment intent, so they finalize and exit -immediately without a grace period. Finalization is persisted separately from -the exit result; the gateway deletes an ephemeral sandbox only after the -finalized supervisor session disconnects. - -Local Docker development builds the supervisor image separately from the -`openshell-sandbox` workload runtime. Cross-platform runtime extraction uses -the sandbox image, which exports `/openshell-sandbox`. - -## Configuration Boundary - -The gateway accepts exactly schema version 2. Missing, legacy, and future -versions fail before runtime construction, and driver settings belong only to -`[openshell.drivers.]`. The process does not migrate legacy files. -Package lifecycle code may replace an exact package-generated v1 default, but -it preserves edited configurations for explicit operator migration. - -Gateway listener TLS and sandbox supervisor TLS are separate inputs. A selected -local Docker, Podman, or VM driver requires a complete guest bundle whenever -the gateway listener uses TLS; package-managed local TLS can supply that bundle. -Kubernetes instead projects guest credentials through its configured Secret. -The gateway validates this requirement before constructing the selected driver. - -## Protocol and Auth - -Gateway validation and concurrency errors use the standard rich gRPC error -envelope. Shared field validators attach `google.rpc.BadRequest`, and conditional -write conflicts attach `google.rpc.ErrorInfo` with a stable reason and current -version when available. `google.rpc.RetryInfo` expresses a minimum retry delay; -it does not establish that a mutation is safe to repeat. SDKs retain the original -transport status, metadata, and unknown details alongside decoded fields. -SDK deletion waits recognize missing-resource status through typed error wrappers -without suppressing other failures. - -Ordinary user-callable unary mutations explicitly opt into durable request -admission when the client supplies a UUID. Typed adapters -check current authorization before looking up a caller/method/workspace-scoped -key. The payload fingerprint excludes that UUID and canonicalizes protobuf maps. -An atomic, quota-checked insert chooses one executor; owned execution survives -client cancellation. Success is persisted before acknowledgment. Errors or -interruption leave permanent unresolved claims, never stealable leases. - -Admission rows live outside user workspace namespaces and are bounded per caller. -Successes expire after 24 hours; cleanup uses the unique admission incarnation -and version so an old cleaner cannot delete a new attempt. Replay stores only -resource references and reviewed public scalar/diagnostic receipts, never -credential-bearing response snapshots. It checks original identities and current -authorization and never substitutes a same-name resource. Sandbox responses are -live projections of the original UUID; normal status reconciliation does not -invalidate replay. Refresh status additionally requires the original grant epoch -and no deletion timestamp, including a timestamp at the Unix epoch. -Other resource projections retain exact-version guards. Terminal delete receipts -do not require the deleted target or parent to remain present. - -Sandbox, service, provider/profile, and policy/config adapters use keyed payload -fingerprints derived from existing gateway JWT or primary TLS private material. -Replicas must share that material; missing keys or key changes fail closed without -changing admission identity. Workspace/template adapters retain their original -format. Intercepted requests carry the original decoded payload only in a private -in-memory extension. Replay reauthorizes original and current effective scopes, -requires the same effective payload, and reruns current interceptor validation. -Interceptors cannot mutate the request UUID. Server-marked replay suppresses -post-commit observation, which remains best-effort rather than an outbox. -Credential capabilities require separate contracts. - -Both exec RPCs share keyed admission but defer completion to the SSH producer. -The initial interactive Start uses the exec request schema under a distinct RPC -namespace; later stdin and resize frames are not replayable inputs. Admission -resolves the public sandbox name and workspace, durably binds both original and -effective sandbox UUIDs, and rejects same-name replacements. When the original -and effective selectors match, both authorization lookups must resolve the same -workspace and sandbox identities before admission. A different effective target -is allowed only when the selector changes. The owner checks the effective UUID -again before relay opening, then hands its CAS-only finalizer to the owned producer, -releasing the shared admission-worker permit after handoff. Only a confirmed -remote exit records a terminal marker and starts 24-hour retention. Synthetic -timeouts, disconnects without exit confirmation, and persistence failures leave -permanent unresolved claims. Duplicates never launch, attach, or replay output: -pending records report uncertainty, and terminal records report stream -unavailability. Existing transport cancellation behavior remains unchanged. -Exec timeouts retain the public duration's precision and presence: an absent -timeout is unbounded, while an explicit zero is a finite timeout. - -The gateway listens on one service port and multiplexes gRPC and HTTP traffic. -The default local single-user deployment mode is mTLS user authentication: -clients present a certificate signed by the local deployment CA, and the -gateway maps the verified certificate subject to a user principal. Kubernetes -deployments use mTLS for transport only and require OIDC or a trusted access -proxy for user authentication unless the explicit unsafe local-development -`allow_unauthenticated_users` switch is enabled. -When that service port is bound to loopback, the listener can also accept -plaintext HTTP on the same port for sandbox service subdomains only. That local -browser path is enabled by default and disabled with -`--enable-loopback-service-http=false`; it never serves gateway APIs, auth, -health, metrics, or tunnel routes. The plaintext service router also rejects -browser requests whose Fetch Metadata, Origin, or Referer headers indicate a -cross-origin or sibling-subdomain request. - -The normative public contract rules live in the -[protobuf API conventions](../proto/README.md). Public API fields follow one -entity-reference convention. `name` identifies the -primary resource targeted by an RPC. A role field such as `sandbox`, `provider`, -`service`, or `workload_template` identifies an entity referenced while -operating on another resource or relationship. Entity references never append -`_name`; their string value is already the canonical name. - -Public workspace-scoped RPCs declare `workspace_scope` first and use the typed -`WorkspaceSelector`. A request that targets one workspace selects a non-empty -canonical workspace name; `default` is an ordinary explicit name, not an -omitted-value fallback. Only sandbox, sandbox template, provider, and service -collection list RPCs accept `all_workspaces`, after Platform Admin -authorization. Provider-profile requests may omit the selector to address the -platform profile scope. The authenticated sandbox bootstrap request may also -omit it because the gateway resolves the immutable sandbox identity before the -supervisor has learned its workspace. Canonical sandbox -IDs remain internal metadata used at authentication, persistence, and -compute-driver boundaries; public callers do not use them as sandbox -references. The gateway resolves the name to the persisted sandbox record only -after authorizing the selected workspace. A -sandbox principal is instead resolved by the immutable ID in its authenticated -identity, then checked against the requested name and workspace. Missing and -unauthorized references use the same response within each principal class so -the resolver does not expose an object-existence oracle. Sandbox, sandbox -template, provider, and service collection list RPCs use the same field with an -all-workspaces marker. Platform-global policy operations omit both `sandbox` -and `workspace_scope`, while sandbox policy operations require both. - -Docker and Podman supervisors use host networking and connect through the -gateway's primary listener. On Linux, local supervisors use the primary -loopback endpoint. Sandbox JWT authentication and the generated sandbox RPC -allowlist remain the authorization boundary; the gateway does not negotiate or -bind compute-driver-specific listeners. - -The `rpc_auth` classification is the source of truth for supervisor access. -Marking an RPC as `sandbox` or `dual` makes it callable by an authenticated -sandbox principal on the primary listener. Review such changes as -authorization-surface changes. - -Operators can configure a gateway-wide gRPC request rate limit. The limit is -applied only to gRPC API traffic after protocol multiplexing; health, metrics, -and local sandbox-service HTTP routes are not rate limited by this control. - -Gateway interceptors run in one middleware layer on the `openshell.v1.OpenShell` -gRPC service after authentication and before tonic dispatches to individual -handlers. At startup the gateway calls each configured interceptor's `Describe` -RPC, validates declared bindings against the compiled OpenShell descriptor set, -and builds an immutable execution plan. Only unary OpenShell methods in the -gateway's explicit interceptable-method allowlist are decoded through the -descriptor set into protobuf JSON, evaluated through configured phases, and -re-encoded before the handler sees the request. New RPCs are non-interceptable -until deliberately added to this allowlist. Interception remains centralized: -allowlisting a unary RPC does not require method-specific gateway -instrumentation. - -Remote extension clients share `openshell-extension-core` transport and bearer -primitives. When gateway JWT signing is configured, the gateway mints -short-lived, exact-audience EdDSA credentials for middleware and interceptors, -rotates their in-memory slots without rebuilding clients, and publishes the -public verification key at `/.well-known/jwks.json` alongside OIDC-shaped -discovery metadata at `/.well-known/openid-configuration`. HTTPS extensions can -pin an operator-provided CA while retaining endpoint-hostname verification. - -Extension credentials reuse the sandbox signing key and are separated from -sandbox-to-gateway admission tokens by exact audience and by an explicit -`typ` of `openshell-ext+jwt`, so a verifier that checks either one alone -cannot confuse the two. After authenticated `Describe` succeeds, a service may -advertise `expected_audience` as a post-authentication consistency assertion; -a mismatch against operator configuration fails gateway startup. A strict -verifier may reject an incorrect audience before returning the manifest. A -registration may opt out of extension authentication entirely with -`allow_insecure_transport`, which permits a plaintext endpoint, attaches no -credential, and warns at every startup. Credential minting is bounded per -sandbox because it resolves the caller's effective policy. - -Each configured interceptor selects a binding policy. `dynamic` accepts valid -manifest declarations and preserves the compatibility behavior. `allowlist` -enables only operator-configured RPCs and phases, while `exact` requires the -configured and declared sets to match. Strict policies match by RPC rather than -manifest binding ID, so renaming a binding does not change authority. Provider -profile sources remain a separate operator-controlled capability. - -The protobuf schema marks dedicated credential, token, and refresh-material -fields with a custom secret option. The middleware recursively omits those -fields from every request and post-commit response sent to an interceptor while -retaining the complete protobuf operation for handler dispatch. JSON Patch -paths and source paths cannot select an omitted field or replace a containing -object. There is no configuration that exposes annotated fields. - -`SubmitPolicyAnalysis` is interceptable because proposed chunks can eventually -change active policy through the gateway's approval workflow. An interceptor -may therefore reject policy proposals while permitting telemetry-only requests. -Gateways without a matching binding retain the standard proposal behavior. - -The descriptor codec uses protobuf's standard `oneof` semantics. If binary -input contains multiple alternatives from one group, the last member on the -wire wins. The middleware converts that selected value to ProtoJSON and -re-encodes it before dispatch, so the interceptor and handler observe the same -canonical request. ProtoJSON input that names multiple alternatives remains -invalid. - -Modification results are atomic per binding. After applying one binding's full -JSON Patch list, the middleware re-encodes the candidate as the request's -protobuf type and decodes those accepted bytes back to canonical ProtoJSON. -Invalid candidates follow that binding's failure policy: fail-open restores the -exact pre-binding operation, while fail-closed rejects the request before -handler dispatch. Later bindings only observe the same schema-valid operation -that the handler will receive; protobuf map entry ordering is not treated as a -semantic difference. - -Each interceptor evaluation selects exactly one phase payload: -`modify_operation`, `validate`, or `post_commit`. Modification and validation -payloads carry the protobuf JSON operation entering that phase. Post-commit -payloads carry the successful committed response instead of echoing the -request. Only the `validate` payload can also carry optional read-only -`current_state`; modification and post-commit evaluations never receive it. The -gateway does not yet load method-specific state, so the field remains absent; -an absent state is distinct from an explicitly empty object. Method-specific -state schemas and persistence-version binding are deferred until a concrete -consumer requires them. - -Post-commit evaluation is strictly observational. A binding that includes -`post_commit` must resolve to `fail_open`, or interceptor initialization fails. -After a handler returns success, failures never replace the committed response. -Binding failures emit the standard fail-open warning and counter; response -observation or evaluation failures outside binding policy emit warnings and the -`openshell_gateway_interceptor_post_commit_observation_failures_total` metric. -The gateway reconstructs the original response frames, including trailers and -body errors, before evaluating the observer. - -Interceptor manifests can also vend provider profile catalogs. No profile is -compiled into the gateway: configuration selects the exact ordered source set -from the stored user source and named profile-capable interceptors. Omitting the -setting selects the user source alone, so a gateway with nothing imported serves -an empty catalog; selecting only an interceptor makes it authoritative by -omission. Every selected source uses the same snapshot, -semantic-validation, and duplicate-detection path. Duplicate normalized profile -IDs fail instead of creating source precedence. The gateway treats configured -interceptors as trusted sources and does not verify signature annotations in -their profile payloads. - -The CLI exposes reusable profile definitions through `openshell profile`, with `list` and `describe` reading the same effective catalog used by provider creation. Export, import, update, lint, and delete share that top-level command group. Import and lint accept local files, local directories, or a single HTTP or HTTPS URL; the CLI fetches and parses remote content before submitting it to the gateway. Workspace selection and explicit platform scope apply at the existing profile API boundary; `openshell provider` manages credential-bearing instances. - -Each logical gateway request captures the selected sources into one validated, -immutable effective catalog before deriving provider behavior. Policy layers, -credential scope, injected environment material, dynamic token grants, and -provider-environment revisions use that same catalog. Each configured source is -therefore fetched at most once per request, and a source revision change becomes -visible on the next request instead of partway through the current request. The -capture emits debug diagnostics with the combined catalog revision, source fetch -count, and profile count; it never logs provider credentials or profile material. - -Supported auth modes: - -| Mode | Use | -|---|---| -| mTLS user auth | Local single-user Docker, Podman, and VM gateway access. | -| Plaintext | Local development or a trusted reverse proxy boundary. | -| Unauthenticated local users | Trusted Kubernetes dev or fully trusted proxy deployments only. | -| Cloudflare JWT | Edge-authenticated deployments where Cloudflare Access supplies identity. | -| OIDC | Bearer-token auth for users, with browser or device-code PKCE and client credentials login. Discovery and JWKS retrieval require HTTPS, reject redirects, and pin JWKS to the issuer origin or an explicit origin allowlist. JWKS validation accepts RS256, RS384, RS512, PS256, PS384, PS512, ES256, ES384, and EdDSA (Ed25519) signing keys. | - -The CLI persists the scopes requested during OIDC login in gateway metadata and -reuses them when refreshing an access token. This preserves the intended API -resource selection for identity providers that bind access-token audiences to -OAuth scopes. - -Python and Go SDK client-credentials providers can use the same registered -issuer, client ID, audience, and scope metadata; the TypeScript provider accepts -those fields explicitly. All three own a separate in-memory lifecycle, repeat -the grant before expiry, and never persist the client secret or acquired access -token into the CLI token cache. They require TLS when sending renewable bearer -credentials to non-loopback gateways. This keeps non-interactive SDK -authentication independent from refresh-token rotation and shared disk state. - -Gateway health and user authentication are separate probes. `OpenShell.Health` -remains unauthenticated so deployment and load-balancer health checks do not -depend on user credentials. The CLI uses the existing, side-effect-free -`OpenShell.GetGatewayInfo` capability query as its protected authentication -probe. `Unauthenticated` means the credentials were rejected, while -`PermissionDenied` proves authentication succeeded before the caller failed -the capability query's admin authorization check. The CLI combines the health -and capability results so a reachable gateway with an expired or rejected -token is reported as connected but unauthenticated. - -Sandbox supervisor RPCs authenticate with explicit sandbox credentials; mTLS -does not grant sandbox identity. Kubernetes deployments use the -gateway-minted JWT bootstrap path: the supervisor starts with a projected -ServiceAccount token, exchanges it for a gateway-minted sandbox JWT, and uses -that JWT on subsequent gateway RPCs. -User-facing RPCs are authorized by descriptor-declared role and scope policy -when OIDC or edge identity is enabled. The OIDC admin role grants platform-wide -access and bypasses workspace membership checks. Workspace Admin and Workspace -User roles are durable membership records keyed by workspace and authenticated -subject. Handlers resolve the resource workspace and require sufficient -membership after the middleware validates the global role and optional scope. -The authenticated `GetCurrentUser` endpoint exposes the gateway's validated -user subject, display name, roles, scopes, and identity provider for CLI -identity inspection without client-side token decoding. - -Sandbox secrets are gateway-signed JWTs bound to a single sandbox ID. Docker, -Podman, and VM drivers deliver the initial token through supervisor-only -runtime material; Kubernetes supervisors exchange a projected ServiceAccount -token through `IssueSandboxToken`. The gateway delegates that opaque credential -to the selected compute driver's `AuthenticateSandbox` RPC. A capable driver -returns the authenticated sandbox ID and an opaque runtime identity. The -gateway requires both a matching durable sandbox record and the exact -driver/runtime identity recorded at provisioning before returning the current -generation-bound session JWT. Session authentication checks the durable runtime -generation and token lineage for every sandbox RPC, so a replaced runtime and -legacy unbound tokens cannot retain provider or control-plane access. The -gateway admits only explicitly typed, generation-bound session JWTs for sandbox -RPCs. It does not accept the pre-session untyped JWT format. The -Kubernetes driver uses its own named configuration to run TokenReview and -verify the live pod and controlling Sandbox CR. Its runtime identity binds the -namespace, immutable Sandbox CR UID, and supervisor Pod UID. Restart preserves -the namespace and CR UID, rejects ambiguous label matches, and rotates only the -Pod-bound portion of the identity. The bootstrap path accepts -both `agents.x-k8s.io/v1beta1` ownerReferences from newer Agent Sandbox -controllers and `agents.x-k8s.io/v1alpha1` ownerReferences from existing -deployments. Supervisors renew gateway JWTs in memory before expiry only while -the sandbox record still exists. Each successful refresh atomically stores the -new gateway-token ID in that sandbox record. The immediately consumed bearer -can recover that same successor for 30 seconds when the request matches, but it -cannot authorize ordinary RPCs or choose another successor. Advancing the -successor removes that retry path across every gateway replica. Short -`gateway_jwt.ttl_secs` lifetimes still bound the exposure of a current bearer -that has not yet been refreshed. Omitting `gateway_jwt.ttl_secs` selects -non-expiring launch-scoped gateway and Sandbox Protocol tokens for local -single-player Docker, Podman, and VM gateways; both token profiles carry -`exp = 0`, and supervisors skip periodic renewal of those session tokens. -Typed extension JWTs retain a 900-second default when the field is -omitted. Kubernetes and other shared deployments should set a positive TTL. -Explicit zero is rejected. - -Gateway JWT signing-key rotation is currently an offline operator action. The -runtime loads one active signing key and one matching public verification key -from the configured secret at startup. To rotate that key material today, -operators must delete or replace the JWT key secret, let certgen recreate it, -and restart the gateway pods. This invalidates outstanding supervisor tokens; -running supervisors recover by re-running their bootstrap path where available -or by reconnecting after sandbox restart. Online rotation with multiple -verification keys keyed by `kid` is tracked separately. - -Sandbox JWTs are not user credentials. The gRPC router accepts -`Principal::Sandbox` only on the supervisor-to-gateway RPC allowlist -(`ConnectSupervisor`, `RelayStream`, token renewal, config sync, policy status, -log push, and policy-analysis callbacks). Handlers then compare the -authenticated sandbox ID with any sandbox ID or name resolved from the request. -Supervisor control and relay streams require a matching sandbox principal before -the gateway registers the session or bridges relay bytes. - -## HA Supervisor Ownership - -In multi-replica Kubernetes deployments, every gateway pod can accept client -RPCs, but a sandbox supervisor maintains one active stream to one gateway -replica at a time. The connected replica publishes a short-lived supervisor -owner record in the shared Postgres object store with its replica id, peer DNS -endpoint, supervisor instance id, and connection epoch. Ownership does not move -because another gateway receives a client request. It changes only when the -supervisor opens a new control stream, usually after the previous owner pod is -terminated or the stream breaks. A reconnect from the same supervisor instance -with a newer epoch can supersede the previous owner before the TTL expires, and -heartbeats from the active connection renew that current owner record. -Cleanup from an older connection checks the shared owner record before and -after changing sandbox readiness. It cannot demote a sandbox after a newer -replica has published replacement ownership. - -Session-bound operations such as exec, TCP forwarding, file sync, and sandbox -service routing first check the local session registry. If the supervisor is -owned by another gateway replica, the serving gateway opens an internal -`PeerRelay` stream to that owner and asks it to open the supervisor relay. This -keeps client traffic working when a Kubernetes Service routes the client to a -non-owner gateway pod. If a peer owner is stale or unreachable during a rollout, -the serving gateway retries ownership lookup until the normal relay wait -deadline. Each retry re-reads the owner record, so a supervisor reconnect or -heartbeat can surface a new owner; if no fresh reachable owner appears before -the deadline, the client operation fails rather than electing an owner itself. -Provider-readiness reports, endpoint-status reports, and provider-status reads -also follow the durable owner record through unary peer RPCs. The owning replica -validates the current supervisor session and keeps the in-memory evidence; a -non-owner never accepts evidence from a stale local session or projects a -remote session as disconnected. - -Nothing redistributes established sessions, so after a rolling restart the last -surviving replica holds most sessions and a new replica serves none until -sandboxes reconnect. That skew decays only as sandboxes churn. Client traffic -stays correct throughout because a non-owner relays to the owner. - -File upload and download use tar-over-SSH through the same relay path. A gateway -pod termination drops the active SSH proxy byte stream, so the CLI retries the -whole sync operation with a fresh SSH session instead of attempting mid-stream -resume. - -Gateway peer RPCs authenticate with Kubernetes ServiceAccount identity rather -than a shared secret. Helm mounts a projected, pod-bound token with audience -`openshell-gateway-peer`; the receiving gateway validates it through -TokenReview, checks the live pod UID and chart selector labels, and authorizes -only the internal peer RPC methods. When gateway TLS is enabled, peer clients -also trust the chart CA, present the chart-generated client certificate for -mTLS, and verify the stable gateway Service DNS name even when connecting to a -Deployment pod IP. - -`WatchSandbox` uses the local update bus for same-replica writes. On -multi-replica backends one shared poller per gateway observes resource-version -changes made by other replicas and feeds that bus for all local watchers, -avoiding a database poll per client stream. SQLite deployments do not run the -poller because they are single-replica and the local bus already sees every -write. - -Mutations whose invariants span sandbox, provider-profile, policy, or provider -records take a process-local mutex and a shared PostgreSQL advisory lock. The -database session remains dedicated to the request and closes when the guard is -dropped, which releases the lock on normal completion, cancellation, or error. -SQLite deployments use only the local mutex because they are single-replica. - -## API Surface - -The gateway API is organized around platform objects and operational streams: - -| Area | Examples | -|---|---| -| Sandbox lifecycle | Create, list, delete, watch, exec, SSH session bootstrap, ForwardTcp service forwarding. | -| Providers | Store provider records, discover credentials, resolve runtime environment. | -| Policy and settings | Get effective sandbox config, update sandbox policy, manage global settings. | -| Observability | Push sandbox logs, stream sandbox status and logs to clients. | - -Domain objects use shared metadata: stable server-generated IDs, human-readable -names, creation timestamps, and labels. Crate-level details live in -`crates/openshell-core/README.md`. - -### Watch streams - -`WatchSandbox` merges three per-sandbox sources into one client stream: status -snapshots, server/sandbox logs, and platform events. Logs and platform events -are resumable; a shared per-sandbox allocator stamps each with a `cursor`. -Cursor-ordered delivery is guaranteed for the replay phase: on resume the -buffered events from both sources are sorted before emission. Live events are -monotonic within each source, but the two sources are read independently, so a -client should order across sources by `cursor` rather than by arrival. Status -snapshots and warnings are re-read on demand and carry an empty cursor. - -#### Cursor spaces - -A sandbox's cursors live in a **cursor space**: a `{epoch, seq}` pair, where the -epoch is a UUID minted on the first publish and `seq` counts from 1. The epoch -is dropped by `TracingLogBus::remove`, so a teardown — or a gateway restart — -retires the space, and the next publish mints a new one. A sequence number alone -cannot distinguish a caught-up client from one holding a cursor out of a space -that no longer exists, because the replacement space reuses the same numbers; -the epoch answers *which counter issued this*, which is the question resume -validation actually has to ask. Behind multiple replicas the same rule makes a -reconnect to a different replica fail loudly rather than return the wrong -events. - -On the wire a cursor is an opaque, fixed-width token. Clients may only compare -two cursors from one stream and keep the greater; the encoding zero-pads `seq` -so that byte-wise comparison matches sequence order, which is what lets every -SDK track a high-water mark without parsing. A stream only ever observes one -epoch — a reset closes both resumable broadcast receivers, ending the stream -rather than switching spaces mid-flight — so that comparison is always well -defined where clients are allowed to use it. The gateway does not rely on it: -server-side ordering runs on the raw `u64` seq carried alongside each event in -`CursoredEvent`, never on the token. - -The gateway holds a bounded in-memory tail per sandbox. Loss is reported with -two distinct, documented behaviors: - -- **Recoverable lag** — a broadcast receiver falls behind and the server skips - ahead. The stream emits a `SandboxStreamWarning` event and continues. Since - cursors are opaque, the warning is the client's only signal. -- **Unrecoverable gap** — the server sends a snapshot, then terminates with - `OUT_OF_RANGE`. Three cases reach it: the tail was trimmed past the requested - cursor, the cursor's epoch does not match the sandbox's current space, or no - space exists because nothing has been published since teardown. The status - tells the client to restart with an empty cursor; retrying the same token - fails identically. A token the gateway could not have issued is rejected - earlier, as `INVALID_ARGUMENT` on the call itself. - -Both resumable sources draw from one cursor space, so the server merges them by -seq before emitting rather than draining each in turn: on resume it replays only -events after the client's cursor, and without one it replays each bus's retained -tail. Either way the batch leaves in ascending cursor order. The two tails are -bounded independently (`log_tail_lines` and `event_tail`), so merging orders -whatever each bus kept; it does not align their depths. - -The broadcast receivers are subscribed before replay, so an event buffered during -initialization could appear in both replay and the live receiver; the producer -tracks the highest replayed seq and suppresses live events at or below it, so -each event is delivered once. That mark is per source. The two tails are read at -different instants and bounded independently, so one shared mark would let the -deeper source censor the shallower one — with `event_tail` unset the mark rises -to the newest buffered log while no platform event is replayed at all, and -platform events published during initialization are discarded as duplicates of a -replay that never ran. Subscribing never mints a cursor space, so a -resume against a torn-down sandbox cannot create the space its stale cursor is -then checked against. Clients track the highest observed `cursor` and pass it as -`resume_after_cursor` on reconnect. - -The epoch is validated twice on resume: once before reading the tails and again -once both are in hand, before anything is emitted. The check and each read take -their locks separately, so a teardown plus a republish can retire the validated -space and install a replacement in between; the reads would then apply the old -space's seq to the replacement's buffers, and a trimmed-range check that only -compares numbers would report no gap while skipping the replacement's lower -events. The second look ends the stream with `OUT_OF_RANGE` instead. - -## Persistence - -The gateway persistence layer is a protobuf object store. Domain services store -typed protobuf messages as opaque binary payloads, while the database keeps a -small set of indexed metadata columns for lookup, listing, versioning, and -workflow state. The implementation lives in the -[gateway persistence module](../crates/openshell-server/src/persistence/mod.rs); -backend-specific SQL lives in the SQLite and Postgres migration directories -under `crates/openshell-server/migrations/`. - -The storage schema is intentionally narrow: - -| Column | Purpose | -|---|---| -| `id` | Stable gateway-generated object ID and primary key. | -| `object_type` | Logical resource kind, such as `sandbox`, `provider`, `provider_profile`, `ssh_session`, `sandbox_policy`, or `draft_policy_chunk`. | -| `name` | Human-readable name, unique within an object type when present. | -| `scope` | Optional owner or namespace for scoped/versioned records, such as a sandbox ID for policy revisions. | -| `version` | Optional monotonically increasing version for scoped records. | -| `status` | Optional workflow state for records such as policy revisions or draft policy chunks. | -| `dedup_key` and `hit_count` | Optional policy-advisor fields for coalescing repeated observations. | -| `resource_version` | Monotonically increasing counter for optimistic concurrency control. Incremented atomically on each update. | -| `payload` | Prost-encoded protobuf payload for the full domain object. | -| `created_at_ms` and `updated_at_ms` | Gateway timestamps used for ordering and list output. | -| `labels` | JSON object carrying Kubernetes-style object labels for filtering and organization. | - -### Protobuf API and storage boundaries - -Public RPC contracts and durable protobuf formats have separate ownership. The `openshell.v1.OpenShell` service's request and response roots, streaming flags, and transitive message closure come from the public descriptor set generated by `openshell-core`. The `public_and_durable_schema_inventories_are_complete` test in `openshell-server` owns the counts and fingerprints and requires this inventory to be reviewed whenever it changes. Compute-driver, credential-driver, gateway-interceptor, and supervisor-middleware services are compiled contracts for internal extension boundaries. - -`ReportEndpointStatus` is a sandbox-authenticated public gateway RPC. Its request, response, and `EndpointObservation` messages belong only to the public closure. `EndpointStatus` and `EndpointResult` also belong to the durable closure because `Sandbox.status.endpoint_statuses` persists them. The repeated status field uses a new wire tag; stored sandboxes without it decode with an empty endpoint list and retain their lifecycle fields. A fixed payload encoded with the earlier sandbox schema verifies that no database rewrite is required. - -Allow and deny append requests carry `L7RuleTarget` to declare the rule, endpoint, and complete affected scope. The removed `host` and `port` fields remain reserved by number and name, and requests without a target are rejected. These mutation requests are not persisted formats. - -`GetSandboxProviderStatus` and `ReportProviderReadiness` are unary public gateway RPCs. The first lets authorized users inspect a provider change; the second accepts installation reports only from the sandbox's current authenticated supervisor session. - -`CreateSandboxRequest.service_exposures` is an additive public API field. Its -`SandboxServiceExposure` entries register service endpoints as part of sandbox -creation and are represented in the Rust, Python, TypeScript, and Go SDK create -options. The request-only exposure description is not durable; the gateway -persists the resulting `ServiceEndpoint` objects through the existing endpoint -store after it persists the sandbox. `SandboxResponse.service_urls` returns the -routed URLs keyed by service name for `CreateSandbox`; the empty key represents -the unnamed endpoint, and other sandbox operations leave the map empty. - -The removed `NetworkBinary.harness` field remains reserved by number and name, -so protobuf implementations cannot reuse its wire slot or source identifier. -The durable-policy compatibility decoder reads the former boolean before Prost -discards it and migrates advisor provenance to the rule endpoint. A fixed -pre-0.1.0 policy payload verifies that the former wire format still decodes. - -Storage-only messages live in the private, versioned -`openshell.storage.v1` package under `crates/openshell-server/proto`. The server -generates these types separately, so the public descriptor set and the Rust, -Go, Python, and TypeScript client generation inputs do not advertise them. -When a frozen scalar storage field cannot distinguish absence from its zero -value, gateway-owned object metadata annotations carry that presence bit rather -than extending the frozen message. - -| Storage classification | Protobuf messages | Durable use | -|---|---|---| -| Encoded storage roots | `StoredProviderCredentialRefreshStateV2`, `StoredProviderProfile`, `PolicyRevisionPayload`, `DraftChunkPayload` | Complete protobuf payload stored in an object row or a scoped policy row. The frozen V1 refresh state remains available only for transactional upgrade decoding. | -| Nested storage-only type | `StoredRefreshMaterialDeletion` | Repeated child records inside provider refresh state. | -| SQL materializations | `StoredPolicyRevision`, `StoredDraftChunk` | Server-only typed results assembled from indexed columns and decoded payloads; not public RPC messages. | -| Public messages used directly as encoded storage roots | `Sandbox`, `SandboxWorkloadTemplate`, `Provider`, `Workspace`, `WorkspaceMember`, `SshSession`, `ServiceEndpoint` | The generated public type is also the persisted payload. `SshSession` is not in the current public RPC message closure. | -| Embedded encoded root | `SandboxPolicy` | Stored in policy rows and inside the JSON settings envelope. | -| Configuration operation storage root | `StoredConfigUpdateOperation` | One common operation resource stores the exact provider target, receipt projection, snapshot failure reason, and historical outcome. | -| Public dependencies of an operation | `ConfigUpdateOperation`, `ProviderMutationReceipt`, `ProviderReadinessReason` | Their complete message and enum closures are durable contracts. | - -The descriptor-derived test owns the complete public, durable, and intersecting inventories and their reviewed fingerprints. The tables here record their roots and classifications. A synthetic sandbox-spec byte fixture verifies that an absent server-owned attachment epoch decodes to the valid empty initial identity; direct and template creation tests separately require the gateway to replace any caller-supplied epoch. - -Public delete, membership-removal, and SSH-revocation responses use -`DeletionOutcome`, not a transport-success boolean. `COMPLETED` establishes -logical gateway deletion or revocation; it does not guarantee that downstream -platform garbage collection has finished. Sandbox deletion returns `ACCEPTED` -while its captured object ID remains in the store, and returns that ID so callers -can distinguish the original sandbox from a same-name replacement. Identity-aware -SDK deletion waits complete on absence or a different observed ID; name-only waits -continue until the name is absent. The existing owned deletion worker continues -after request cancellation. - -Missing targets return `NOT_FOUND` unless `allow_missing` explicitly requests -`ALREADY_ABSENT`. Authorization, parent resolution, preconditions, and backend -failures remain errors. Already-revoked sessions complete without another write -after current authorization. The removed response booleans are reserved by name -and number; this coordinated pre-1.0 API change does not alter durable schemas. -The outcome alone does not provide request deduplication. Opted-in unary methods -require a request UUID for the admission contract. - -Configuration admission adds `SandboxStatus.configuration_admission` at field -11 and optional `configuration_activated` at field 12, extending the public and -durable closures. New sandboxes explicitly store `false` until first acceptance; -acceptance stores `true` permanently, including across restart. Legacy rows -have neither field and conservatively retain static-policy restrictions. No -database rewrite is required. A pre-admission byte fixture verifies that legacy -phase and policy-version fields survive without fabricated admission or activation. -`SandboxStatus.provisioning` uses field 13 for gateway-owned attempt timing and -compute reclamation progress. Its timestamps survive supervisor reconnects and -ordinary driver status updates. Older records decode with no provisioning -record; timing must be adopted once and persisted, never reconstructed from the -object's frequently changing update timestamp. The additive message requires -no rewrite of existing payloads and leaves the frozen storage-v1 schema intact. -Stored settings JSON also carries per-key change IDs and commit timestamps, -including deletion tombstones. Legacy values acquire stable source identities -on read; a subsequent write preserves them. These clocks distinguish effective -edits from no-op writes without treating status updates as configuration edits. -With timestamp types, deletion outcomes, and optional mutation request IDs, the -admission contract brings the public closure to 298 messages and 21 enums, the -durable closure to 92 messages and 16 enums, and their overlap to 80 messages -and 16 enums. Mutation request IDs extend public request fields without adding -messages to these closures or changing the durable protobuf schema. - -| Dual-purpose encoded root | Current decision | -|---|---| -| `Sandbox` | Defer a storage twin; govern its complete dependency closure as durable. | -| `SandboxWorkloadTemplate` | Defer a storage twin; govern its complete dependency closure as durable. | -| `Provider` | Defer a storage twin; govern its complete dependency closure as durable. | -| `Workspace` | Defer a storage twin; govern its complete dependency closure as durable. | -| `WorkspaceMember` | Defer a storage twin; govern its complete dependency closure as durable. | -| `SshSession` | Defer a storage twin; govern its complete dependency closure as durable. | -| `ServiceEndpoint` | Defer a storage twin; govern its complete dependency closure as durable. | -| `SandboxPolicy` | Defer a storage twin; govern its complete dependency closure as durable. | -| `ConfigUpdateOperation` | Persist the common historical outcome within `StoredConfigUpdateOperation`; govern its complete dependency closure as durable. | -| `ProviderMutationReceipt` | Persist the immutable provider projection within `StoredConfigUpdateOperation`; govern its complete dependency closure as durable. | -| `ProviderReadinessReason` | Persist only the closed snapshot failure category within the operation; govern its enum values as durable. | - -The public/storage overlap is deliberate for the current format. Storage twins -for the public roots are deferred: introducing them would require a broad -conversion boundary, and Prost does not retain unknown fields through a -decode-and-reencode conversion. Each root therefore carries a reviewed decision -to remain dual-purpose, and its complete transitive dependency closure is also -a durable format. Important embedded dependencies include `ObjectMeta`, -`ProviderProfile`, `CredentialHandle`, `SandboxPolicy`, and -`NetworkPolicyRule`. Global and sandbox settings additionally store an encoded -`SandboxPolicy` inside their JSON envelope. - -Public API compatibility and storage compatibility are reviewed independently: - -- Public compatibility is evaluated from public service descriptors and SDK - generation inputs. Storage-only packages must never enter that closure. -- `openshell.storage.v1` is frozen. Its test fingerprint covers message names, - field numbers, cardinality, scalar wire types, referenced types, map-entry - shapes, and optional presence. Keep its decoder available and introduce a - new versioned package plus an explicit migration or fallback decoder for a - format change; never reuse removed tags or names. -- Checked-in synthetic byte fixtures were encoded with the former - `openshell.v1` declarations. Current storage types must continue to decode - them semantically, which proves the package move does not require a database - rewrite. Protobuf payload bytes do not encode a message's package name. -- A change to a dual-purpose public message or any transitive durable - dependency requires both public-wire review and storage-migration review. - Wire-incompatible changes require a migration or fallback decoder and a - fixture for the earlier format. -- Mixed-version writers are unsupported. An older Prost writer can discard - fields it does not know when it reads and rewrites a record, even when the - newer field is wire-compatible. - -Common resources use generic helpers that derive `object_type`, `id`, `name`, -and labels from protobuf metadata traits before encoding the full message into -`payload`. Policy revisions and draft policy chunks use the same table but also -populate `scope`, `version`, `status`, `dedup_key`, and `hit_count` so the -gateway can efficiently fetch the latest policy, track load status, and manage -advisor drafts without creating resource-specific tables. - -Mutation admission uses a private, version-tagged JSON envelope in the same -object store. Its identity namespace stays stable across format changes, and an -unknown format fails closed. It contains explicit typed receipts, not arbitrary -public response payloads, and is not part of the protobuf storage closure. -Workspace create/delete admissions include the requested workspace name in the -key, but omit a workspace UUID guard. Different names have independent request-ID -namespaces; deletion receipts remain replayable after the target disappears. - -Each sandbox policy revision stores the complete provenance annotation map -supplied with that update. The revision payload is the authoritative immutable -record; sandbox metadata receives the same annotations only as a convenience -projection and can retain keys from earlier revisions. Policy revision creation, -optional first-policy backfill, metadata projection, and superseding older -revisions commit in one database transaction. SQLite serializes this operation -with an immediate transaction, while Postgres locks the sandbox row. A failed -resource-version check or revision insert rolls back the entire operation. - -SQLite is the default local store; Postgres is supported for deployments that -need an external database or multi-replica coordination. Both backends expose -the same `Store` API and the same logical schema. Backend differences stay -inside the adapters: for example, SQLite stores labels as JSON text and payloads -as `BLOB`, while Postgres stores labels as `JSONB` and payloads as `BYTEA`. -Domain code should depend on the object-store contract, not SQL dialect details. -This keeps the gateway data model portable across storage backends and leaves -room for future stores that can provide the same object, label, version, and -scope semantics. - -For in-memory SQLite, the adapter retains a dedicated keepalive connection for -the store lifetime. Operational connection replacement therefore preserves the -shared in-memory schema and objects instead of creating an empty database. - -Public protobuf APIs represent absolute times with `google.protobuf.Timestamp` -and elapsed time with `google.protobuf.Duration`. The integer -`created_at_ms` and `updated_at_ms` database columns are intentionally internal -bookkeeping values, not part of that public convention. On startup, both -storage backends transactionally rewrite legacy scalar time fields inside -protobuf payloads before serving requests. A malformed affected payload aborts -and rolls back startup migration. Legacy driver-provided condition strings that -cannot be represented as timestamps are dropped so an accepted historical -value cannot make the upgraded gateway unavailable. -Gateway and Sandbox Protocol token responses follow the same convention: a -present expiration timestamp carries the absolute deadline, while absence means -the issued token does not expire. - -On-disk SQLite databases run in WAL journal mode with `synchronous=FULL`. -The adapter switches the file to WAL on a single connection before the pool -opens, then applies both settings to every pooled connection. WAL lets readers -proceed while a writer commits and reduces each commit to one WAL `fsync`, -which matters because gateway hot paths such as SSH session issuance and -revocation are many small autocommit writes. `synchronous` stays at `FULL` -because some of those writes tighten authorization: under `NORMAL`, a power -loss could roll back an acknowledged SSH session revocation and make the token -valid again. Writes that are safe to lose, currently only SSH session issuance -through `Store::create_relaxed`, use a second single-connection pool with -`synchronous=NORMAL`. Losing a minted token only invalidates it, and because -both pools share one WAL, the next `FULL` commit also makes earlier relaxed -commits durable. Deployments that need multiple replicas use Postgres, where -`create_relaxed` is an ordinary durable insert. WAL requires a local filesystem with working shared -memory, so the SQLite file must not live on a network mount, and backups must -use `sqlite3 .backup` or `VACUUM INTO` rather than copying the main file alone. - -The SQLite adapter tightens the on-disk database file to mode `0o600` on every -connect so that provider API keys, SSH session tokens, and sandbox metadata are -not readable by other local users on shared hosts. The same restriction is -reapplied to the `-wal` and `-shm` sidecars that WAL mode creates, -which mirror the same sensitive contents. - -Persisted state includes sandboxes, providers, provider profiles, provider -credential refresh state, SSH sessions, policy revisions, settings, deployment -records, and reusable sandbox workload templates. Provider refresh -state is stored as a separate object scoped to the provider instance through -`objects.scope`. Its non-secret configuration remains inline, while refresh -tokens, client secrets, private keys, and other secret source material are -stored through the active credential driver and represented by opaque handles. -The provider record keeps only the current injectable credential handles and -optional per-credential expiry timestamps. A refresh normally mints one -credential, but a strategy may co-mint several (AWS STS mints the access key, -secret key, and session token in one call); the refresh state pins the resolved -set of env keys it owns so collision checks reserve all of them before the -first mint. Provider records keep inline credential values only for legacy -records created before credential driver storage. New provider and -refresh-material writes keep driver-owned credential handles. When no external -credential driver is configured, gateways use server-owned encrypted database -credential storage for defense in depth. Multi-replica deployments can use that -default with a shared database and shared key-encryption key, or opt into an -external backend such as Vault or Kubernetes Secrets. - -The Vault credential driver requires HTTPS for every non-loopback backend, -never follows HTTP redirects, and keeps standard certificate hostname -verification enabled. Operators can add private Vault trust roots with a PEM -CA bundle; the driver does not replace platform roots or expose a certificate -verification bypass. - -Sandbox workload templates are workspace-scoped gateway resources. Workspace -admins create and delete them; workspace users can read and list them. A -template owns reusable workload intent: image, environment, CPU and memory -limits, GPU request, driver-specific config, and service-level hints. A sandbox -created from a template resolves that resource once and persists an ordinary -`SandboxSpec` snapshot. The create request still owns per-sandbox governance: -name, labels, annotations, provider attachments, and policy. The sandbox stores -template provenance as the template name and resource version used for the -snapshot, so later template edits or deletes do not mutate existing sandboxes. - -OAuth refresh failures retain a gateway-owned recovery classification alongside -the refresh state. The gateway reads only a bounded error response and maps -recognized OAuth codes to retry, reauthorization, configuration repair, or -investigation without persisting issuer-controlled descriptions. Terminal -reauthorization failures remain parked until a manual retry or explicit refresh -reconfiguration. Configuration failures retry hourly so an externally repaired -clock, policy, or stored credential can recover without rapid endpoint traffic; -short-lived credentials still fail closed at their recorded expiry. - -Credential handles remain bound to the driver that created them. Before the -0.1.0 compatibility boundary, gateways do not migrate inline refresh material -or move handles between credential drivers; operators reconfigure affected -grants when upgrading or changing backends. - -### Optimistic Concurrency (CAS) - -Every object row carries a `resource_version` that the database increments -atomically on each write. Concurrent mutations use compare-and-swap (CAS): the -writer reads the current version, applies changes, and writes back with a -`WHERE resource_version = ` guard. If another writer updated the row -in between, the guard fails and the caller receives a `Conflict` error. - -This matters for HA deployments where multiple gateway replicas share the same -Postgres database, and for single-node deployments where concurrent gRPC -handlers or the reconciler mutate the same sandbox. - -**Compile-time enforcement.** The unconditional write methods `put` and -`put_message` are gated behind `#[cfg(test)]`. Production code must use -`put_if` with an explicit `WriteCondition` or `update_message_cas`. The -compiler rejects any other write path, making non-CAS writes structurally -impossible outside of tests. - -Every write goes through one of three conditions: - -- `MustCreate` -- insert-only. The database rejects the write with a - `UniqueViolation` error if a row with that ID already exists. Handlers match - on the structured `PersistenceError::UniqueViolation { .. }` variant to - distinguish creation conflicts from other failures. -- `MatchResourceVersion(v)` -- update-only. The database rejects the write - with a `Conflict` error if the current version differs from `v`. -- `Unconditional` -- test-only; not reachable in production builds. - -**Creates.** All create paths use `MustCreate` and hydrate the response -directly from the `WriteResult` returned by `put_if`, which carries the -assigned `resource_version`, `created_at_ms`, and `updated_at_ms`. This -eliminates a read-after-write round trip and the race window that would come -with it. - -**Updates.** The `update_message_cas` helper makes a single CAS attempt: it -fetches the current object, applies a mutation closure, and writes with a -`MatchResourceVersion` condition. On conflict the persistence layer returns a -`Conflict` error, which gRPC handlers map to `ABORTED` status so the client -(or the next watch/reconcile event) can retry with fresh state. Provider -attach/detach handlers retry bounded server-driven conflicts around this helper; -an explicit client version still fails on its first conflict. - -The helper accepts an `expected_version` parameter that selects between two -modes: - -- **Server-driven** (`expected_version = 0`): the helper uses the version it - just read from the database. Internal operations (reconciler, policy status - reports, compute phase transitions) use this mode because the caller does - not track versions. -- **Client-driven** (`expected_version != 0`): the helper validates that the - caller's version matches the current database version before applying the - mutation. If they diverge it returns `Conflict` without attempting the - write. Client-facing operations that carry an `expected_resource_version` - field use this mode: `AttachSandboxProvider`, `DetachSandboxProvider`, - `UpdateProvider`, `UpdateProviderProfiles`, and `UpdateConfig` (policy - backfill and sandbox annotation updates). - -**Lists.** Public list RPCs follow AIP-158: requests carry direct `page_size` -and `page_token` fields, and responses carry `next_page_token`. The gateway -clamps page sizes to 1,000 and returns opaque base64url continuation tokens. -Tokens bind the RPC and every request parameter except `page_size`, contain no -authorization grant, and use immutable keyset cursors. Each page repeats normal -authentication and authorization. Pagination is resumable and keyset-based; it -does not provide a historical snapshot while the collection is mutated -concurrently. - -Object-backed lists (`ListSandboxTemplates`, `ListSandboxes`, `ListServices`, -`ListProviders`, `ListWorkspaces`, and `ListWorkspaceMembers`) sort ascending -by `(created_at_ms, name, workspace, id)`. `ListProviderProfiles` sorts -ascending by `(id, scope)`, and `ListSandboxPolicies` sorts by descending policy -version. Provider attachments sort ascending by provider name. - -The token wire format is a private shared protobuf used only by the gateway. -Public request and response messages repeat the standard AIP fields directly -instead of wrapping them in a shared pagination message. - -The CLI returns paginated JSON and YAML as response-shaped envelopes containing -the resource collection and `next_page_token`; table output reports a non-empty -token on stderr. The TUI traverses complete workspace, provider, profile, and -sandbox collections in one cancellable background refresh task, never overlaps -periodic list refreshes, and discards results after a gateway or workspace -change. - -Curated Rust, Python, Go, and TypeScript SDK list methods return lazy pagers. -Advancing a pager issues one list RPC and exposes its continuation token; -explicit `list_all` helpers are the only curated APIs that exhaust a collection. - -**Migration.** This pagination contract is a breaking replacement for the -former `limit`/`offset` list APIs. Protocol clients must send `page_size` and -resume only with the returned `next_page_token`; an empty token is the sole -completion signal. SDK callers that need every item must use the explicit -full-iteration helper (for example, Rust's `list_all_sandboxes`) rather than -awaiting `list_sandboxes`, which now returns one-page `Pager` state. Callers -that need one page should advance that pager once and retain its token. Internal -full scans use the reusable iteration helpers from the persistence pagination -audit rather than manually advancing offsets. - -Persistence distinguishes one-page operations from exhaustive scans. -`list_object_page` and `list_message_page` return one keyset page and its next -cursor. `collect_records` and `collect_messages` exhaust those pages, fail on -database or protobuf decode errors, and hydrate `resource_version` from the -authoritative database column. Internal callers that require every matching -record use the exhaustive helpers; bounded lookups continue to use page-level -methods. - -**Deletes.** Delete operations are not yet CAS-protected -- the delete request -protos do not carry `expected_resource_version`. A `delete_if` primitive exists -in the persistence layer but is not wired into gRPC handlers. - -**Coverage.** All `ObjectMeta`-bearing message types have write-condition -coverage: - -| Type | Create | Update | List | -|---|---|---|---| -| Sandbox | `MustCreate` | `update_message_cas` | `list_messages` | -| Provider | `MustCreate` | `update_message_cas` | `list_messages` | -| ProviderProfile | `MustCreate` | `MatchResourceVersion` | `list_messages` | -| SandboxPolicy | scoped versioning | scoped versioning | scoped query | -| Settings | `Mutex`-guarded | `Mutex`-guarded | single-row | - -Global settings updates use a Tokio `Mutex` to serialize multi-step -validation within a single gateway process, with CAS on the underlying -persistence write as defense in depth. In an HA deployment with multiple -gateways, the Mutex alone would be insufficient. Sandbox-scoped settings -rely entirely on CAS without a Mutex. - -The `resource_version` is surfaced to clients through `ObjectMeta` in proto -responses. Provider profiles are the exception: custom profile get/list/export -responses copy the stored version onto the profile payload so exported YAML can -carry the expected version for safe single-profile updates. Profile update -requests also carry an explicit target profile ID; the payload ID must match the -target so an edited export cannot overwrite a different profile. Database -migrations backfill existing rows with version 1. - -Provider profile imports, updates, and deletes hold the sandbox synchronization -guard while checking attached-sandbox dynamic token grant ambiguity or in-use -state and writing the profile record. Sandbox creation with initial providers and -sandbox provider attach/detach use the same guard, so gateway replicas cannot -interleave a profile mutation with a sandbox provider-set mutation that would -leave an ambiguous final dynamic-token state or a deleted custom profile that is -still referenced by a sandbox. - -Policy and runtime settings are delivered together through the effective sandbox -config path. A gateway-global policy can override sandbox-scoped policy. The -sandbox supervisor polls for config revisions and hot-reloads dynamic policy -when the policy engine accepts the update. - -External supervisor middleware registration is operator-owned configuration -under `[[openshell.supervisor.middleware]]`. At startup the gateway connects to -each service and validates its described bindings and operator body limit. -Policies attach a complete external middleware by its operator-owned registration -name. Manifest bindings are identified by operation and phase, and each manifest -may declare at most one binding for an operation and phase pair. -Attaching a registration does not require it to advertise every supported -operation. Supervisors select only the manifest bindings that match the current -operation; policy-local config identity remains internal audit metadata. -Before persisting a policy, the gateway asks each selected implementation to -validate its config. The effective sandbox config contains only the registered -services required by that policy; supervisors invoke those services directly on -the request path. - -Provider credential expiry is enforced during gateway-to-sandbox credential -resolution and again by the sandbox placeholder resolver. This keeps expired -credentials from resolving even when a running sandbox still has retained -placeholder generations from an earlier provider credential snapshot. - -All gateway-owned extension registries negotiate the same peer metadata envelope -before accepting work. Compute drivers, credential drivers, gateway interceptors, -and supervisor middleware retain their typed family manifests. Both the gateway -and extension run the shared validator against the startup exchange, enforcing -protocol-major compatibility and mutual required-capability sets before either -peer accepts the other. The gateway aggregates immutable, non-secret startup -snapshots for the protected gateway-info API; it does not publish transport, -authentication, or backend configuration. - -Static credential delivery is capability-negotiated and endpoint-bound. The -gateway classifies each returned environment entry as either a credential or -non-secret provider configuration and associates every credential key with the -host, port, and path selectors from its effective provider profile. It withholds -static credential material from supervisors that do not advertise binding -support. If a selected provider profile has no usable endpoint, the gateway -withholds only that profile's static credential keys and their expiry and -binding metadata. It continues to return provider-generated non-secret -configuration, valid endpoint-bound static credentials from other attached -providers, and the dynamic credential snapshot. Provider environment revisions -include profile endpoint and binding changes. - -The supervisor owns provider fetching, support negotiation, and credential resolution outside the workload. The authenticated sandbox boundary receives a revision and its prepared child environment from one snapshot; it preserves the issued placeholders without receiving the secret resolver or turning those placeholders into new references. - -Provider mutations return immutable per-sandbox receipts that separate saved desired state from observed runtime installation. A receipt pins sandbox and provider identity, the attachment epoch, and the exact provider/configuration/policy fingerprints. Attachment-set changes replace the epoch in the same sandbox CAS write; credential updates stage distinct backend objects before publishing their handles and provider revision. This prevents a published revision from referring to an unfinished in-place credential write. Update receipts retain the sandbox target set selected before publication. - -Each receipt projects a common configuration operation in the `config_update_operation` store; its receipt ID is the operation ID. The provider status path evaluates current installation evidence and records historical terminal outcomes through resource-version CAS. A previously applied operation does not bypass current session, freshness, or authority checks. The live provider projection can be pending after disconnection or superseded after another change even when historical operation state remains applied. Pending operations are evaluated through provider status queries; this path adds no background delivery engine or mutation replay contract. - -When a provider mutation opts into admission with `request_id`, its replay record retains references to the original configuration operations and the original mutation ID. Replay returns those immutable receipts even if sandbox attachments have since changed; it does not select new targets or create replacement receipts. Missing operation evidence makes replay unavailable without executing the mutation again. Provider resources still require their original recorded version, and sandbox resources retain the replay contract's current-state projection. - -Provider mutation and operation-result writes are separate. A result-storage failure can follow a saved mutation and returns structured uncertainty without a rollback or safe-retry claim. A failed initial snapshot remains failed rather than acquiring a different target during a later lookup. Operations contain only identities, revisions, timestamps, and closed reason categories. - -The CLI recognizes the gateway's `CONFIG_OPERATION_STORAGE_UNCERTAIN` error reason and domain for attach, detach, and update. It reports fixed guidance to inspect and reconcile the saved change before retrying, while withholding arbitrary server messages and error metadata. An uncertain mutation never starts a readiness wait or automatic replay. - -Provider receipts, installation status, and common operations represent absolute times with protobuf `Timestamp`; report intervals and evidence lifetimes use protobuf `Duration`. Receipt identity compares the full canonical timestamp without truncating nanoseconds. An absent observation or completion time represents missing evidence or an unfinished operation, independently of the Unix epoch. - -Provider installation reports belong to the existing `ConnectSupervisor` session. Each report names that session, has an increasing sequence, and expires unless the supervisor reports again. Reconnection or disconnect invalidates prior observations; stored change records survive a gateway restart, but runtime evidence does not. Replaying an identical report cannot extend its lifetime. - -Reports and status also compare the supervisor instance with the sandbox's -persisted current instance. A different supervisor becoming current invalidates -an older connection, including one retained by another gateway replica. -Observations stay local to the gateway holding the supervisor session. A status -request or report reaching another replica follows the shared owner record to -that gateway, which remains the sole authority for accepting and projecting the -session's evidence. - -The supervisor reports success only after it installs the matching credentials, activates the effective policy, and receives an acknowledgment from the authenticated workload boundary that it installed the environment for future processes. Environment synchronization shares the process-launch lock, and its acknowledgment identifies the exact publication, including retries at the same provider revision. Failed policy installation cannot reuse evidence for a different installed policy. Ready and revoked statuses also recheck the requested sandbox, provider, attachment and configuration identities; revision fingerprints are compared only for equality. Revocation applies to future credential resolution and future processes. Requests already forwarded upstream can still finish. - -Ordinary static credentials retain revision-scoped references. After update readiness, a newly launched process receives the updated reference; an existing process keeps its original revision. Installation completion does not retarget that reference or establish that an old upstream key can be retired. - -## Provider Environment Resolution - -The gateway resolves only the providers attached to a sandbox. It combines each -provider instance with its profile, returns non-secret configuration, and marks -credentials with the profile's host, port, and path boundaries. The supervisor -uses those bindings when it replaces credential placeholders in policy-allowed -requests. - -Model selection, API protocol, request and response shapes, streaming behavior, -and endpoint URL construction remain responsibilities of the workload's native -client. The gateway does not parse or transform model API requests. - -For `google-vertex-ai` providers created with CLI `--from-gcloud-adc`, the CLI -calls gateway `ConfigureProviderRefresh` with OAuth2 refresh material from gcloud -ADC, then `RotateProviderCredential` to mint the first access token before -reporting success. ADC-backed providers mint into `GOOGLE_VERTEX_AI_TOKEN`. A -successful create therefore yields an immediately usable provider; failures roll -back the provider record. Service-account JSON and private keys remain gateway-side -refresh bootstrap material; sandboxes receive minted access tokens instead. - -## Supervisor Relay - -Sandbox workloads maintain an outbound supervisor session to the gateway. This -lets the gateway open per-request byte relays without requiring inbound network -access to the sandbox workload. - -```mermaid -sequenceDiagram - participant CLI - participant GW as Gateway - participant SUP as Sandbox supervisor - participant Target as Sandbox target - - SUP->>GW: ConnectSupervisor stream - CLI->>GW: ForwardTcp / exec / sync request - GW->>SUP: RelayOpen(channel, target) - SUP->>Target: Dial SSH socket or loopback service - SUP->>GW: RelayStream(channel) - CLI->>GW: Client bytes - GW-->>CLI: Client bytes - GW->>SUP: Relay bytes - SUP-->>GW: Relay bytes -``` - -The same relay pattern backs interactive SSH, command execution, file sync, and -local service forwarding. The gateway tracks live sessions in memory and -persists session records so tokens can expire or be revoked. - -Graceful gateway shutdown closes supervisor-session admission before stopping -local compute. It then signals the remaining control sessions to exit and waits -up to ten seconds for their cleanup, including conditional deletion of persisted -ownership. Pending connection setup and sessions already removed from the live -registry remain tracked until cleanup finishes. This lets a replacement -supervisor claim ownership immediately after restart without deleting a newer -replica's claim. An incomplete drain is reported as a shutdown error. Closing -these control sessions does not stop Kubernetes-owned workloads. - -Relay liveness has two backstops so a reset supervisor session cannot leave a -request parked forever. The gateway runs server-side HTTP/2 keepalive on -supervisor connections, and each exec relay's SSH client uses SSH keepalive: an -exec channel may be legitimately silent for a long time (e.g. an agent whose -stdout is redirected to a file), so the exec is never ended on output-idle -alone — instead an unanswered keepalive on a wedged or orphaned relay closes the -channel and returns the exec with an error. Once a command reports its exit -status, the gateway also bounds how long it waits for the trailing channel close. -An exec relay reports success only after receiving both the SSH exit status and -channel close; a missing close is a transport error even when the exit status is zero. -The supervisor derives that SSH exit status from the boundary output stream's -terminal frame after draining stdout and stderr, so an output delivery failure -cannot be masked by a successful process wait. - -Interactive exec treats normal request-stream EOF as the end of stdin and resize -input. The gateway sends SSH EOF while keeping the output channel open until -command completion. Input errors terminate the operation rather than masquerading -as normal EOF. The input and output pumps are owned by the exec operation, so -timeout or response abandonment cannot leave a detached stdin task behind. -The pumps share polling fairly, and request processing yields cooperatively even -for ignored resize messages, so sustained input cannot monopolize the operation. -The CLI keeps piped input in the unary request when the complete encoded -request fits the gateway's decoder limit, preserving compatibility with older -gateways. Input up to the CLI's 4 MiB cap uses this stream with bounded frames -when the unary message would exceed the decoder limit, with or without a PTY. -The CLI closes the input side at pipe EOF. - -Go and TypeScript interactive-exec helpers distinguish process exit from stream -completion. They consume the final gRPC status before reporting success and retain -an observed exit code if transport completion fails. Callers must drain output -concurrently with waiting for completion. - -TypeScript starts interactive exec eagerly and uses a bounded output queue between -the background receiver and the consumer. Cancellation wakes a receiver blocked -on that queue. Go exposes input closure through an optional session capability, -preserving the original interface for existing implementations. TypeScript also -preserves its original session interface; SDK-created sessions expose lifecycle -controls through an extended interface. - -`ForwardTcp` is the client-facing byte stream for SSH and service forwarding. -The first frame is a `TcpForwardInit` that carries the workspace-scoped sandbox -name, an authorization token from `CreateSshSession`, and an explicit target: -`target.ssh` for the sandbox SSH socket or `target.tcp` for a loopback service -inside the sandbox. The gateway validates the token and sandbox readiness, -sends a targeted `RelayOpen` to the supervisor, then bridges -`TcpForwardFrame::Data` to `RelayFrame::Data` until either side closes. - -Browser service URLs use the same supervisor relay path after host-based -routing resolves `sandbox--service.` to a stored -service endpoint. Accepted service routing domains are derived from wildcard -DNS SANs configured on the gateway server certificate, with -`openshell.localhost` available by default for loopback gateways. TLS-enabled -loopback gateways print `http://` URLs when loopback plaintext service HTTP is -enabled; non-loopback TLS gateways continue to print `https://` URLs. - -For `target.tcp`, the gateway only accepts loopback destinations such as -`localhost`, `127.0.0.0/8`, or `::1`. The gateway never needs to know or dial a -sandbox pod IP; supervisors connect outbound and bridge only the explicit target -requested for that relay. - -## PKI Bootstrap - -`openshell-gateway generate-certs` is the one place local mTLS materials and -sandbox JWT signing material are created. Deployment paths use it as follows: - -| Output mode | Selector | Layout | -|---|---|---| -| Kubernetes Secrets | (default) `--namespace`, `--server-secret-name`, `--client-secret-name`, `--jwt-secret-name` | Two `kubernetes.io/tls` Secrets with `tls.crt` / `tls.key` / `ca.crt` plus one Opaque sandbox JWT Secret with `signing.pem` / `public.pem` / `kid`. | -| Kubernetes JWT-only Secret | `--namespace`, `--jwt-only`, `--jwt-secret-name` | One Opaque sandbox JWT Secret with `signing.pem` / `public.pem` / `kid`. | -| Filesystem | `--output-dir ` | `/{ca.crt, ca.key, server/tls.{crt,key}, client/tls.{crt,key}, jwt/{signing.pem,public.pem,kid}}`. Also copies client materials to `$XDG_CONFIG_HOME/openshell/gateways/openshell/mtls/` for CLI auto-discovery. | - -On Kubernetes, the Helm chart runs the command via a pre-install/pre-upgrade -hook Job using the gateway image itself -- no separate cert-generation image, -no extra mirror burden in air-gapped environments. In the default built-in PKI -path the hook creates TLS and sandbox JWT Secrets. When cert-manager is enabled, -cert-manager owns TLS Secrets and the hook runs with `--jwt-only` so the -required sandbox JWT Secret still exists before the gateway workload mounts it, -even if `pkiInitJob.enabled` remains true. On package-managed local -gateways, the same command runs from the systemd -unit's `ExecStartPre` to bootstrap PKI into the configured local TLS directory -on first start. The Linux package unit defaults that directory to -`~/.local/state/openshell/tls` through `OPENSHELL_LOCAL_TLS_DIR` so certificate -generation and runtime auto-detection use the same path across systemd -versions. - -The bootstrap paths share the same idempotency contract: all requested targets -present -> skip; partial requested state -> fail with a recovery hint; nothing -requested present -> generate and write. This guards continuity across restarts -and upgrades while still recovering cleanly if an operator deletes everything -and starts over. - -When `grpcRoute.backendTLSPolicy.enabled=true`, the certgen hook also creates a -`ConfigMap` containing the CA certificate (`ca.crt`) used by the Gateway proxy -to validate the backend pod's TLS certificate. The CA is always read from the -authoritative server Secret (not the in-memory bundle) so that enabling -BackendTLSPolicy on an existing release uses the CA that actually signed the -server certificate. The ConfigMap is reconciled on every hook run: if the CA -changes (rotation, re-issue), the ConfigMap is updated in place. In built-in PKI -mode the ConfigMap is created in the same pre-install hook. In cert-manager -mode, a separate post-install/post-upgrade hook Job polls for the cert-manager- -issued server Secret and then creates or updates the ConfigMap, because -cert-manager Certificate resources are regular release objects applied after -pre-install hooks. - -The `server.tls.enableMtls` value controls whether the gateway requires client -certificates. When `enableMtls` is `false`, the gateway runs HTTPS-only without -client certificate verification (use OIDC for identity instead). BackendTLSPolicy -requires `enableMtls=false` because the ingress proxy cannot present client -certificates to the backend. - -Operators who manage TLS PKI with cert-manager enable `certManager.enabled`; -cert-manager takes precedence over built-in TLS generation and the chart still -renders the JWT-only hook. Operators who pre-create all TLS and JWT Secrets can -disable both `pkiInitJob.enabled` and `certManager.enabled`. - -## Configuration - -The gateway reads its configuration from three sources, merged in this -precedence (highest first): - -``` -Gateway CLI flag > gateway OPENSHELL_* env var > TOML file > built-in default -``` - -The TOML file is opt-in via `--config ` / `OPENSHELL_GATEWAY_CONFIG`. -Driver implementation settings live exclusively in TOML driver tables. The -selector is the singular `[openshell.gateway] compute_driver`; legacy -`compute_drivers` lists are rejected. See `docs/how-it-works/gateways/configuration.mdx` -for worked per-driver examples and RFC 0003 for the full schema. - -Each installation has an operator-assigned gateway name. Configure it with -`[openshell.gateway].name`, `--name`, or `OPENSHELL_GATEWAY_NAME`. -The built-in default is `openshell`; the Helm chart defaults it to the chart -fullname so every replica in one installation reports the same identity. -Operators must set a globally distinct name when one telemetry collector serves -installations in multiple Kubernetes namespaces or clusters. -The name identifies the gateway installation independently of client-side -aliases, network names, and the sandbox JWT issuer. - -`database_url` is env-only and rejected when present in the file -(`OPENSHELL_DB_URL` / `--db-url`). - -### Driver ownership - -`[openshell.gateway]` contains gateway process settings only. Each selected -driver reads its own configuration exclusively from -`[openshell.drivers.]`; values are never inherited from gateway scope. -Kubernetes owns `namespace`, `default_image`, `supervisor_image`, -`client_tls_secret_name`, `service_account_name`, `host_gateway_ip`, -`enable_user_namespaces`, and `sa_token_ttl_secs`. Docker uses -`sandbox_label` instead of the legacy `sandbox_namespace` name. Podman and VM -likewise own their image, endpoint, and runtime settings in their tables. - -`image_pull_policy` uses the shared canonical vocabulary -`always | if_not_present | never | newer`. Drivers translate it to their runtime -APIs; `newer` is supported only by Podman and rejected by Docker and Kubernetes. - -### OTLP export - -The gateway already uses Rust's `tracing` framework for structured log events -and request-span context consumed by stdout and the sandbox log bus. OTLP export -adds an OpenTelemetry layer to the same subscriber. That layer turns selected -`tracing` spans into distributed traces; it does not export log events or -replace the existing logging paths. - -`[openshell.gateway.otlp]` is the only enablement path for OpenTelemetry -export: the table's presence is the on-switch, and `OTEL_EXPORTER_OTLP_ENDPOINT` -is ignored so enablement has a single source. TOML decides whether and where -to export; the SDK's `OTEL_*` variables tune how. Transport is OTLP over gRPC -only. Shared provider, resource, and tracing-layer construction lives in -`openshell-otel`, along with shared HTTP/tonic trace-context propagation and -gRPC failure recording. - -The `tower_http` `TraceLayer` in `multiplex.rs` opens a span per inbound request, -and that span continues incoming W3C trace context when present or starts a new -trace otherwise. It is named for the RPC and carries the request ID that also -appears in the gateway's logs — the identifier that lets an operator pivot -between a trace and its log lines. Store and compute-driver spans become -children of the request span. Reconciliation, provider refresh, and -driver-watch loops create their own operation spans because they have no -inbound request to provide a parent. gRPC status is recorded when response -trailers arrive. Gateway spans carry resource attributes for the gateway -identity and configured compute driver. - -The gateway forwards OTLP configuration, its configured gateway name, and W3C -trace context to managed external drivers. Built-in drivers use dedicated -in-process providers that preserve the same RPC trace boundary. A shared -compute-driver tracing descriptor derives provider layers and standalone -installation from each driver's service identity, and routes both the shared -RPC boundary target and that driver's crate target prefix. A shared RPC tracer -owns unary and streaming boundary outcomes. Each driver -exports to the configured collector under its own service name and carries the -gateway name and configured compute driver as resource attributes. -Compute-driver client and server spans share the fully qualified protobuf -operation name, such as `openshell.compute.v1.ComputeDriver/CreateSandbox`, -in both the span name and `rpc.method`; the current RPC semantic conventions -integrate the service into that fully qualified method and do not emit -`rpc.service`. `service.name` and span kind distinguish the two sides. -Backend-prefixed spans describe implementation work beneath that boundary. -Streaming `WatchSandboxes` spans remain open for the stream lifetime on both -sides. A terminal stream status records its outcome; -consumer teardown without a terminal status leaves the span status unset. - -Two invariants shape the failure behavior. Telemetry is diagnostic, so no OTLP -failure stops the gateway from serving: a malformed endpoint is logged at -startup and disables export. Export is best-effort — the SDK logs runtime -failures, and a failed batch is dropped rather than retried. Buffered spans -flush after the server loop exits so `SIGTERM` does not drop in-flight traces. - -### Package-managed gateway registry - -The CLI reads its active-gateway and per-gateway metadata from -`$XDG_CONFIG_HOME/openshell/`. It also looks for a package-manager owned -system config root at `/etc/openshell`, using the same layout as the per-user -config root: `active_gateway` plus `gateways//metadata.json`. Packages -or runtimes that need a different location can override that root with a -non-empty absolute `OPENSHELL_SYSTEM_GATEWAY_DIR`; empty or relative values -fall back to `/etc/openshell` and emit a warning. The CLI falls back to this -system config when no per-user `metadata.json` exists; malformed user metadata -still shadows the system entry, but stray empty directories do not. - -System entries are read-only from the CLI, so `gateway remove` rejects a pure -system entry instead of pretending to delete package-manager owned state. - -## Operational Constraints - -- Gateway TLS and client certificate distribution are deployment concerns owned - by the operator or packaging layer. -- Compute runtimes own the mechanics of starting workloads and injecting the - gateway endpoint. Docker and Podman supervisors use host networking; local - Linux supervisors use the gateway's primary loopback endpoint. Kubernetes - uses the gateway Service rendered by Helm. VM supervisors use their - runtime-specific host route. -- Docker Desktop requires host networking to be enabled and cannot combine it - with Enhanced Container Isolation. Set an explicit remote `grpc_endpoint` - when the gateway is not reachable on the Docker daemon host. -- Podman Machine uses gvproxy's host-loopback route on macOS. Native Linux - Podman uses the primary loopback endpoint. -- Gateway restarts recover persisted objects from storage, but live relay - streams must be re-established by supervisors. -- User-facing behavior changes must update published docs in `docs/`; this file - should only record stable architecture. diff --git a/architecture/google-vertex-ai-provider.md b/architecture/google-vertex-ai-provider.md deleted file mode 100644 index 2bc1c51694..0000000000 --- a/architecture/google-vertex-ai-provider.md +++ /dev/null @@ -1,98 +0,0 @@ -# Google Vertex AI Provider - -The `google-vertex-ai` provider gives selected sandboxes direct access to -Google Vertex AI endpoints without exposing long-lived Google credentials. -The provider profile owns endpoint policy and credential metadata; the -sandbox attachment owns which workload receives access. - -## Boundaries - -| Component | Responsibility | -|---|---| -| CLI | Discovers ADC or accepts service-account bootstrap material and creates the provider record. | -| Gateway | Stores refresh material through the credential driver and rotates short-lived access tokens. | -| Provider profile | Declares Vertex hosts, credential aliases, refresh constraints, and permitted binaries. | -| Sandbox supervisor | Delivers opaque credential placeholders and resolves them only for profile-authorized Vertex requests. | -| Workload | Selects the native Vertex endpoint, model, request format, streaming mode, and timeout. | - -The provider does not select a model or transform requests. Anthropic Claude -uses Vertex's publisher-model `rawPredict` or `streamRawPredict` paths. Gemini -and other models use their documented native or OpenAI-compatible Vertex API. - -## Credential Flows - -Vertex accepts short-lived Google OAuth2 access tokens. Both supported setup -flows converge on a rotating access token stored in the provider record. - -### Service account - -The operator creates the provider with service-account bootstrap material and -configures `google-service-account-jwt` refresh. The private key remains in the -gateway credential store. The gateway mints -`GOOGLE_VERTEX_AI_SERVICE_ACCOUNT_TOKEN` and refreshes it before expiry. - -### gcloud ADC - -For local development, `--from-gcloud-adc` reads an authorized-user ADC file, -stores the refresh grant at the gateway, and mints -`GOOGLE_VERTEX_AI_TOKEN`. The ADC file and refresh token do not enter the -sandbox. - -## Runtime Data Flow - -1. The operator attaches the provider to a sandbox. -2. The effective policy includes the profile's Vertex endpoint and binary - rules. -3. A newly launched workload receives the current token as an opaque - placeholder plus non-secret project and region configuration. -4. The workload calls the native Vertex endpoint with that placeholder in the - `Authorization: Bearer` header. -5. After policy and endpoint binding pass, the proxy substitutes the current - real access token and forwards the request. -6. Token refresh updates the resolver; the workload continues using the same - placeholder. - -The raw `GOOGLE_SERVICE_ACCOUNT_KEY` credential is bootstrap-only and is never -part of sandbox runtime material. - -## Endpoint Boundary - -The example profile in `providers/google-vertex-ai.yaml` permits official Vertex hosts: - -- `-aiplatform.googleapis.com` -- `aiplatform.googleapis.com` -- `aiplatform.us.rep.googleapis.com` -- `aiplatform.eu.rep.googleapis.com` - -The workload constructs the documented project, location, publisher, and model -path. The proxy does not infer the publisher or rewrite the body. - -For Claude on Vertex, a request uses this form: - -```text -https://-aiplatform.googleapis.com/v1/projects//locations//publishers/anthropic/models/:rawPredict -``` - -The request body includes the Vertex Anthropic API version expected by Google. -Streaming uses the corresponding native streaming endpoint. - -## Invariants - -- Refresh bootstrap material remains gateway-only. -- Sandboxes receive placeholders, never real access tokens. -- A token resolves only at endpoints covered by the attached provider profile. -- Detach and expiry revoke placeholder resolution. -- Project IDs, regions, publishers, and model IDs are non-secret workload - configuration. -- Provider attachment does not grant access when a gateway global policy - override suppresses provider-derived policy. -- Provider environment keys and dynamic credential bindings remain - unambiguous across all providers attached to one sandbox. - -## Operational Notes - -Attach or detach the provider with the normal sandbox provider lifecycle. A -running sandbox observes policy and resolver changes, but a process must be -launched after attachment to receive new environment variables. Direct native -requests provide the end-to-end verification path; provider creation does not -probe a model endpoint or validate a model ID. diff --git a/architecture/sandbox-limits.md b/architecture/sandbox-limits.md deleted file mode 100644 index 4a245a5ed9..0000000000 --- a/architecture/sandbox-limits.md +++ /dev/null @@ -1,179 +0,0 @@ -# Sandbox Limits - -The sandbox supervisor processes untrusted agent traffic while sharing one -process across connections. Limits protect that process from unbounded memory, -work queues, parser effort, and waits. They are safety boundaries, not capacity -targets or compatibility guarantees. OpenShell is still early, so the values -and some ownership boundaries will change as production evidence improves. - -This document inventories durable sandbox supervisor and egress limits. It -intentionally omits retry cadence, test-only values, and ordinary buffer chunk -sizes that do not cap aggregate resource use. - -## Limit Model - -OpenShell uses three kinds of limits: - -| Kind | Purpose | Configuration | -|---|---|---| -| Platform ceiling | Protect the supervisor even when policy or an external service is hostile or mistaken. | Fixed in code by default. | -| Operator ceiling | Bound an operator-run integration below the platform maximum. | Configurable within platform validation bounds. | -| Policy inspection bound | State how much application data a policy needs the supervisor to buffer and inspect. | Configurable when the application protocol needs it, preferably below a platform ceiling. | - -The component that first allocates, queues, or waits on a resource owns its -limit. Network frame and message assembly belongs to the network supervisor; -middleware RPC and envelope limits belong to the middleware runner; application -inspection limits belong to the corresponding L7 parser. - -New limits should follow these rules: - -- Enforce the bound before allocation or admission whenever possible. -- Acquire shared capacity before buffering and retain it through the complete - buffered operation. -- Give partial-progress protocols both an idle bound and an absolute bound when - either one alone still permits resource pinning. -- Let operator or policy configuration narrow a platform ceiling, never raise - it silently. -- Define the terminal behavior: reject, close, truncate, shed, or backpressure. -- Do not let `fail_open` bypass a platform safety or protocol-integrity bound. -- Emit safe saturation or limit telemetry without request bodies, credentials, - query parameters, or external free-form diagnostics. -- Test time bounds with simulated time and test shared budgets under saturation. - -## Gateway Sandbox Resources - -Gateway-owned sandbox resources also carry admission limits before they can -produce supervisor work. Reusable workload templates are capped at 1000 per -workspace. Template payloads reuse sandbox spec validation for environment -entry count and size, image and resource field sizes, driver-config serialized -size, and GPU count. Template names use the same DNS-style resource-name rules -as other named gateway resources. - -## Middleware - -Middleware limits are process-wide per sandbox. Registry replacement preserves -the shared work, waiter, and persistent-session admission state so activity -retained by an older generation still consumes the same process-lifetime -budgets as new activity. - -| Resource | Current bound | Scope and behavior | -|---|---:|---| -| Concurrent buffered work | 32 | Shared by HTTP requests, WebSocket messages, and WebSocket preflight. One permit covers one complete unit of work. | -| Admission waiters | 64 | Additional work is shed when both the active budget and waiter budget are full. HTTP receives a complete 503 response before its body is buffered. | -| Persistent middleware sessions | 32 | Shared process-wide session budget for streaming middleware protocols. WebSocket preflight uses immediate admission before opening streams and retains one permit while any stage remains active. | -| HTTP body or WebSocket text message | 4 MiB | Platform maximum for input and replacement payloads. Service, operator, and stage limits may narrow it. | -| Middleware configs and stages | 10 | At most 10 configs in policy and 10 selected stages in one chain. | -| Selector patterns | 32 | Combined include and exclude patterns per middleware config. | -| Per-stage RPC | 500 ms default, 10 ms–30 s | An operator timeout caps a binding timeout. | -| Complete message chain | 30 s | Starts after work admission; admission backpressure does not consume the chain budget. | -| WebSocket preflight | 1 s maximum | Caps handshake delay independently of the message RPC timeout. | -| Remote service connect | 5 s | Applies while establishing a middleware gRPC channel. | - -Middleware also validates every non-body envelope component. Important examples -include 64 KiB service config, 4 KiB request context, 32 KiB target data, 128 -request headers totaling 64 KiB, 64 header mutations, 32 findings per stage, -and 64 metadata entries. The external contract lives in -`proto/supervisor_middleware.proto`, with service-author guidance in the -[supported middleware operations](../docs/extensibility/supervisor-middleware/operations.mdx). - -The work semaphore bounds aggregate buffered middleware input to approximately -`32 × 4 MiB`, plus bounded envelope and parser overhead. It is a concurrency -safety valve, not rate limiting or a promise that 32 simultaneous maximum-size -messages are inexpensive. - -The persistent session semaphore is independent from the work semaphore. One -WebSocket middleware session consumes one permit regardless of its active-stage -fan-out, which is separately capped at 10 stages. All-skip preflight releases -the permit immediately. A retained session releases it at connection end or as -soon as its last active stage is disabled. Session admission does not wait: -capacity exhaustion follows each selected config's `on_error` behavior before -any stream opens. The protocol-neutral registry ownership allows future -streaming HTTP middleware to use the same process-wide budget. - -## Egress Framing and Inspection - -| Path | Current bound | Terminal behavior | -|---|---:|---| -| Initial CONNECT request headers | 8 KiB | Reject the proxy request. | -| Inspected HTTP/1 request headers | 16 KiB | Reject the request. | -| Streamed HTTP/1 chunk framing | 16 KiB per chunk-size line; 16 KiB and 128 fields for the complete trailer block | End the relay. Chunk payloads pass through a fixed 8 KiB buffer and do not accumulate to the declared chunk size. | -| Credential-rewritten HTTP body | 256 KiB | Reject when rewriting requires a larger buffered body. | -| SigV4 body signing | 10 MiB | Reject when signing requires a larger buffered body. | -| GraphQL request body | 64 KiB default | Policy can set a positive `graphql_max_body_bytes`; there is no shared platform ceiling yet. | -| MCP or JSON-RPC request body | 64 KiB default | Policy can set a positive `max_body_bytes`; there is no shared platform ceiling yet. | -| Parsed WebSocket client text message | 4 MiB | Close with `1009` when the complete or decompressed message is larger. | -| Concurrent parsed WebSocket text assemblies | 32 active, 64 waiters | Shared process-wide across all parsed relays. Additional messages close with `1013` before payload allocation or reading when both bounds are full. | -| Parsed WebSocket fragments per message | 4,096 | Close with `1002`. | -| Parsed WebSocket text assembly | 30 s input idle, 2 min total | Close with `1002`; the total includes initial and continuation payloads, continuation headers, and interleaved control frames. These deliberately permissive initial bounds can be tightened, or made operator-tunable within platform bounds, after production behavior is understood. | -| Parsed WebSocket text forwarding | 2 min total | End the relay and release assembly capacity. A timeout does not append a close frame to a partially written data frame. | -| Parsed raw WebSocket binary frame | 16 MiB | Close with `1002`; binary messages are relayed rather than inspected. | -| HTTP relay waiting for EOF | 5 s input idle | End the relay with a timeout. | -| TLS certificate cache | 256 hosts | Clear the cache before inserting another host. | - -Ordinary allowed traffic is streamed rather than accumulated to a -connection-sized buffer. A parsing or transformation feature introduces a -buffer only when it owns an explicit bound. - -Every parsed WebSocket text message acquires network-owned assembly capacity before payload allocation or reading, including relays used only for native policy, credential rewriting, compression, or a disabled fail-open middleware session. The process-lifetime budget survives policy reloads, and the assembly retains its permit through decompression, policy and middleware evaluation, credential rewriting, and upstream forwarding. Active middleware sessions additionally acquire shared middleware work before buffering. Input progress resets only the idle deadline. Forwarding uses one total deadline across the complete frame header, payload, and flush. Every timeout and terminal parser error releases both permits through ordinary ownership. Queue exhaustion emits a payload-free network denial event. - -The operator middleware `max_payload_bytes` ceiling applies to payloads exposed -through HTTP-body and WebSocket text-message bindings. It does not replace the -raw binary frame safety bound because binary messages are never delivered to V1 -middleware. A passed binary logical message still advances the active -middleware session sequence and emits coverage telemetry, so a later text RPC -can contain a valid sequence gap. - -## Network and Upstream Proxying - -| Path | Current bound | Terminal behavior | -|---|---:|---| -| Executable identity pins | 4,096 unique paths per supervisor lifetime | Reject an entire identity chain before insertion when its new paths would exceed the bound. Existing pins remain usable and are never evicted. Staged TCP reports resource exhaustion. | -| Corporate proxy CONNECT response headers | 8 KiB | Fail the tunnel. | -| Corporate proxy CONNECT handshake | 30 s total | Fail the tunnel; validated-address attempts share the aggregate budget. | -| Token-grant HTTP request | 30 s request and connect | Fail credential resolution. | -| Response-derived token cache TTL | 5 min default; 1 h response cap; 30 s expiry margin | A positive profile `cache_ttl_seconds` override replaces the response-derived calculation. | - -## Sandbox-Local Surfaces - -| Surface | Current bound | Scope and behavior | -|---|---:|---| -| `policy.local` request body | 64 KiB and 15 s read | Reject an oversized or stalled local request. | -| Policy proposal long-poll | 60 s default, 1–300 s | Clamp the requested hold time; clients can issue another poll. | -| `policy.local` denial read | 100 records, 4 KiB per surfaced line | Bound response and log parsing work. | -| Log push reconnect buffer | 200 records | Drop new records above the local batch ceiling while disconnected. | -| Log push reconnect backoff | 30 s maximum | Cap the delay between reconnect attempts. | -| Policy status outbox | No fixed capacity | Preserve FIFO revision status without blocking policy reconciliation. | -| Policy status retry backoff | 32 s maximum | Retain a retryable update and retry independently of enforcement. | - -The bounded log batch favors supervisor health over retaining an unbounded -diagnostic backlog. The policy status outbox makes the opposite tradeoff: -revision ordering and delivery survive an extended gateway outage at the cost -of potential queue growth. - -## Known Gaps and Review Triggers - -The current limits grew with individual features and are not yet a complete -resource model. Known gaps include: - -- GraphQL, MCP, and JSON-RPC policy body limits have defaults but no common - platform maximum. -- A positive token-cache TTL override replaces the response-derived one-hour - ceiling rather than narrowing it. -- Socket read deadlines are not expressed consistently as idle plus total - budgets across every parser. -- There is no documented aggregate connection budget or per-destination - fairness policy in the supervisor. -- The policy status outbox is intentionally unbounded and can grow if policy - revisions continue while its gateway endpoint remains unavailable. -- Limit telemetry is not yet uniform enough to derive saturation trends across - all paths. - -Revisit this document when adding a parser, body transformation, persistent -stream, shared queue, cache, or external call. A change should state: - -1. What untrusted resource can grow or wait. -2. Which component owns the bound. -3. Whether the scope is per message, connection, destination, or sandbox. -4. Whether configuration can narrow the limit. -5. How saturation or timeout terminates. -6. Which telemetry and deterministic tests prove the behavior. diff --git a/architecture/sandbox.md b/architecture/sandbox.md deleted file mode 100644 index a8a09310d0..0000000000 --- a/architecture/sandbox.md +++ /dev/null @@ -1,934 +0,0 @@ -# Sandbox - -A sandbox is the runtime boundary where agent code executes. A compute driver -creates it and connects two dedicated components: `openshell-sandbox` inside -the workload boundary and `openshell-supervisor` outside it. - -## Runtime Model - -Each sandbox has three trust levels: - -| Component | Role | -|---|---| -| Supervisor | Owns gateway credentials, admitted policy, L7 proxying, SSH, and gateway relays. It never executes inside the agent workload. | -| Sandbox | Runs as the same non-root identity as the agent, installs the workload seccomp listener, applies the Landlock baseline, owns child processes, and mediates the protected supervisor channel. | -| Agent child | Inherits the sandbox network listener and runs with zero capabilities, `no_new_privs`, Landlock, and the final syscall filter. | - -The runtime grants neither trusted component nor agent child any Linux -capability inside the workload. Drivers resolve one exact non-root UID, GID, -and supplementary-group set before launch. The sandbox and all of its children -use that immutable identity, so no in-workload privilege transition is needed. -The supervisor uses its own driver-defined identity and has no workload-creation -or backend-admin authority. - -The compute driver provisions separate protected configurations and one -mutually authenticated gRPC connection over a private Unix socket, Kubernetes -TCP Service, or VM vsock channel. Independent bidirectional `Exchange` RPCs -carry lifecycle, exec, TCP, and forwarding traffic, while one persistent -bidirectional `Mediate` RPC carries multiplexed DNS traffic. General application -UDP is unsupported; UDP DNS remains mediated by the supervisor. -Both ends size HTTP/2 flow control so the connection window exceeds the -per-stream window times the concurrent-stream limit plus a reserve. Relays whose -workload stops reading therefore stall only their own streams and cannot starve -DNS, exec, or control traffic of connection-level credit. The supervisor caps -concurrent TCP relay and pending-accept streams below the stream limit, so new -workload connections queue before control exchanges lose stream slots. -The sandbox probes HTTP/2 connection liveness every five seconds and closes -connections that miss a ten-second acknowledgement deadline. Closing a -connection freezes the owned workload process tree and cancels its stream -bridges before releasing the exclusive DNS mediation lease. The supervisor has -30 seconds to reconnect, replay attach, and reconfirm the boundary. Every -supervisor process generates an ephemeral instance ID, and the sandbox pins the -first ID it accepts for its process lifetime. The same process can therefore -recover a dropped transport, but a replacement supervisor cannot reuse launch -credentials to claim the existing runtime generation. Confirmation resumes the -workload; expiration terminates it. A credential replacement does not displace -the active connection until the new connection is confirmed. Idle healthy -connections remain usable. -A failed stream alone does not trigger reconnection. The supervisor first -reconfirms on the current connection and keeps it if the boundary answers. -Otherwise it closes that transport before replaying attach, and retries while -the sandbox still reports the equal-epoch connection as active, until the -boundary observes the disconnect. The sandbox reports attach and confirm -rejections as typed errors rather than closing the stream. -TCP mediation accepts use the same authenticated transport recovery as process waits. A healthy idle accept has no timeout. An interrupted pending open fails closed, while a replacement accept waits for new workload traffic; decisions and established byte streams are not replayed. Boundary rejections and failed recovery remain terminal to the proxy. - -A renewed Sandbox Protocol bearer is authenticated even when its credential epoch is unchanged. The supervisor confirms that bearer on the active physical connection and records its fingerprint only after confirmation succeeds, preserving pending streams and the mediation session. Changing the credential epoch still requires an authenticated replacement connection. - -Local single-player Docker, Podman, and VM gateways propagate an omitted -`gateway_jwt.ttl_secs` value to both launch-scoped credential profiles. Those -credentials use `exp = 0`, so host suspension cannot strand the supervisor -after a refresh deadline passes. Shared deployments retain expiring credentials -and a durable compute-platform bootstrap identity. - -Unauthenticated TLS handshakes have a separate bounded asynchronous pool and -five-second deadline, never consuming authenticated control slots or threads. -The socket broker reserves the TCP control-listener port against workload -connections, including loopback aliases. Unix control listeners reject workload -descendants using kernel peer credentials and process ancestry, while ordinary -workload loopback and Unix services remain available. -NetworkPolicy is an outer reachability fence, not a confidentiality boundary. -Each sandbox generation receives a fresh CA and distinct server/client leaves; -both endpoints bind the same workload identity and immutable driver resource -claims. Driver crates do not appear in generic process, network, SSH, or -session code. - -The supervisor exposes readiness only after the sandbox is confirmed and the -gateway access plane is registered. Driver-owned channel directories limit -reachability, while mutual authentication and channel epochs prevent endpoint -replacement from granting authority. - -## Startup Flow - -1. The driver resolves the immutable workload identity, installs the outer - network fence, validates its native evidence, and starts `openshell-sandbox` - with one-use bootstrap state. Docker inspects container networking, - Kubernetes verifies its NetworkPolicy, and VM drivers inspect the guest - device model; those native schemas remain in their driver crates. -2. The sandbox consumes and unlinks bootstrap material, proves the admitted - runtime posture, and listens on the protected driver channel. It does not - run untrusted code yet. -3. `openshell-supervisor` loads policy and runtime settings from the gateway, - attaches to the sandbox, and verifies the driver's generation and evidence. -4. The sandbox installs its seccomp notification broker and Landlock baseline, - validates its mechanism-specific audit evidence, and reports backend-neutral - enforcement properties. The supervisor must accept those properties and - their immutable session and resource binding before it sends the launch - permit. Other isolation backends may establish the same properties with - different mechanisms and retain their detailed evidence in backend-owned - audit data. -5. The sandbox starts the canonical process through its single workload - launcher. The supervisor starts SSH and registers its gateway session. -6. Exec, signaling, PTY, DNS, TCP, and loopback-forwarding operations cross the - authenticated channel for the lifetime of the sandbox generation. - -The shared isolation contract receives only normalized outer-fence guarantees: -egress is default-deny, there is no unmanaged egress path, the evidence is bound -to the sandbox generation, revocation has been verified, and controller loss -fails closed. A digest commits those guarantees to the native -evidence without teaching the shared contract about container networks, -Kubernetes objects, VM devices, or accelerator resources. - -The component that owns the outer fence also validates its native evidence and -makes that projection explicitly. In the current Docker, Podman, Kubernetes, -and VM placements, that component is the compute driver. A delegated isolation -backend may own the fence and make the same projection instead. Non-empty native -evidence alone does not establish a guarantee: - -| Current enforcement owner | Native evidence | Guarantees projected by the owner | -|---|---|---| -| Docker | Pinned container ID, `network_mode=none`, and no unexpected network attachments | No workload route establishes default-deny, revocation, and controller-loss behavior; the attachment inspection establishes that no unmanaged route exists. | -| Podman | Pinned container ID, `--network=none`, and no unexpected network attachments | The same container-network facts establish the same four guarantees. | -| Kubernetes | NetworkPolicy UID and resource version, ingress and egress isolation, and zero workload egress rules | The persisted, selecting policy establishes default-deny and continued denial after revocation or controller loss; zero egress rules establish that no unmanaged route is permitted. | -| VM | Generation and zero guest network devices | The absent NIC establishes all four guarantees; approved traffic uses the separate supervisor-owned channel. | - -The shared contract checks that all four guarantees are present, that the -projection names the admitted generation, and that its evidence digest matches -the value passed to the workload-side runtime. It does not infer guarantees or -interpret the native fields. - -When the admitted main process exits, its status and retained terminal output -remain available. The confirmed sandbox and supervisor-owned access plane continue -to serve policy-authorized exec and loopback forwarding until explicit stop or -delete tears down the boundary and terminates any remaining workload processes. -When the sandbox is PID 1, it reaps adopted workload children after child-exit -notifications, with a low-frequency recovery sweep. Children owned by an -explicit process waiter remain registered and are never consumed by the orphan -reaper. This keeps idle sandboxes from scanning procfs continuously while -preserving wait results and eventual zombie cleanup. - -Completed exec output handles can be reclaimed, but execution request IDs remain -reserved for the boundary generation. The sandbox accepts at most 4,096 exec -attempts per generation, then rejects new attempts rather than forgetting replay -protection. A disconnected attachment does not authorize another execution. -While an exec handle is retained, independent waits return its stable exit or -signal status, whether or not an output attachment is open or the main process -has exited. Waiting never holds the exec registry lock, so other operations can -still signal or attach to the process. -Exec output uses a bounded queue that backpressures the process reader until its -attachment consumes the bytes. An attached reader can apply backpressure -without losing bytes; an absent reader that leaves the queue full eventually -fails the exec and terminates its process. The final stream status carries -delivery failures separately from the process handle's stable wait result. -After the exec parent exits, output drains until its pipes close. If a -descendant keeps a pipe open beyond 30 seconds, the exec reports an output -delivery failure and drains later writes without retaining them. This avoids -guessing which bytes belong to the parent. -Canonical main-process output retains its bounded replay log. - -## Isolation Layers - -OpenShell uses overlapping controls rather than a single sandbox primitive: - -| Layer | Purpose | -|---|---| -| Filesystem policy | Landlock restricts the paths the agent can read or write. | -| Process policy | Sandbox and children run as one immutable non-root identity with zero capabilities. | -| Seccomp notification | Virtualizes supported INET sockets and sends DNS/TCP decisions to the supervisor without nftables or proxy environment variables. | -| Outer network fence | The component that owns network enforcement prevents any missed or unsupported kernel path from escaping. Current examples are Docker `network_mode=none`, a NIC-less VM, and Kubernetes NetworkPolicy. | -| Policy proxy | Evaluates destination, binary identity, TLS/L7 rules, SSRF checks, and inference interception. | - -The supervisor may enrich baseline filesystem allowances for proxy support -files. GPU allowances are added by the workload-side sandbox only when the -immutable driver resource claims request a GPU and GPU devices are visible -inside the workload. Host supervisor device discovery must not influence these -allowances; a CPU-only VM preserves read-only `/proc` even on a GPU host. -These internal allowances must stay sandbox-scoped and avoid exposing host -secrets. For example, MXC governed egress grants the generated public CA bundle -while the ephemeral CA private key remains in the host proxy's memory. -This limitation does not prevent MXC sandboxes from launching in general. -Filesystem policies work normally, and a `process_container` sandbox with -governed egress enabled can still enforce `network_policies` through the host -proxy. The limitation applies only when the effective policy contains -`network_middlewares`, which select built-in or remote services that inspect or -transform network traffic. The MXC host proxy does not currently receive the -gateway registry that resolves those services. MXC therefore rejects such a -policy synchronously during `CreateSandbox`, before it inserts runtime state or -invokes `wxc-exec`, rather than running a chain with missing implementations. -Remove the middleware entries or use a compute driver whose sandbox supervisor -receives the gateway middleware registry. - -The mandatory self-protection baseline is separate from optional workload -filesystem policy. It requires Landlock ABI v3, including pathname truncation -protection. Rules cover individually opened root children except `/.openshell`; -the sandbox opens entries relative to a pinned root descriptor without following -symlinks. An image-provided alias cannot grant access to the protected subtree. -The reserved `/.openshell` root must itself be a real directory if present; -a symlink or non-directory aborts preparation so private child mounts cannot -redirect into an allowed subtree. - -## Network and Inference - -See [Sandbox Limits](sandbox-limits.md) for the current numeric safety ceilings, -their ownership, terminal behavior, and known gaps. - -### Standalone network proxy - -`openshell-supervisor --role=network-proxy` runs the policy proxy without an -Isolation Backend or `openshell-sandbox`. It accepts explicit HTTP proxy and -CONNECT requests on a loopback listener and applies the same local Rego rules, -YAML policy data, destination checks, and L7 enforcement used by supervised -sandboxes: - -```shell -openshell-supervisor \ - --role=network-proxy \ - --listen=127.0.0.1:3128 \ - --tls-dir=/tmp/openshell-proxy-tls \ - --policy-rules=/path/to/sandbox-policy.rego \ - --policy-data=/path/to/sandbox-policy.yaml -``` - -The standalone listener cannot observe which process opened a connection, so -this role evaluates endpoint and protocol rules without binary identity. It -does not launch a workload, attach a Sandbox Runtime, fetch gateway policy, -inject provider credentials, or provide exec and lifecycle operations. The -listener is loopback-only. TLS interception writes its generated public CA and -combined trust bundle to `--tls-dir`; when omitted, the supervisor uses a -process-specific directory under the system temporary directory. - -The sandbox installs one seccomp user-notification listener on a dedicated -launcher thread. Every canonical and exec process inherits that listener. It -virtualizes supported INET sockets before they enter the agent FD table, copies -bounded syscall inputs from the notifying task, resolves the calling binary, -and blocks external `connect` until the supervisor returns a policy decision -and relay stream. Connected data stays on ordinary kernel sockets, so the -notification path is limited to socket setup and pointer-bearing operations. -Blocking listener accepts retain native workload socket flags. A broker-owned -watchdog interrupts an accept when its seccomp notification is cancelled or the -broker stops, including when readiness disappears before the accept syscall. -The sandbox reserves `SIGUSR2` with a non-restarting no-op handler for these -broker threads; startup rejects a conflicting handler. This signal disposition -is process-global kernel state, while registrations and cancellation state are -owned by the broker. Workload exec resets the caught handler to its default. -The broker copies pointer-bearing syscall arguments from a same-UID workload -child with `process_vm_readv` / `process_vm_writev`, falling back to -`/proc//mem` when those system calls are unavailable or blocked. Runtime -qualification keeps the trusted broker non-dumpable and proves read/write -access against a dumpable child, matching the production process topology. -It never treats access to the broker's own memory as workload evidence. - -This sandbox runtime requires Landlock ABI v3 (Linux 6.2, or an equivalent -vendor backport). The seccomp listener is installed in one of two cancellation -modes, and the launch confirmation enforces the invariant -`cancellation || task_memory_writes_disabled`: - -- **Killable** (`SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV`, Linux 5.19+): the - notified workload thread waits kill-only, so a non-fatal signal cannot resume - a mediated syscall between notification validation and the broker's result - write. Full mediation, including task-memory output writes. -- **LegacyReadOnly** (kernels < 5.19, e.g. RHEL 9.x / 5.14): the flag is - unavailable (`EINVAL`), so the listener falls back to a plain notifier and the - broker refuses every task-memory *output* write to stay cancellation-safe. - Concretely, in this mode `getpeername`, `accept`/`accept4` **with a non-null - peer-address argument**, and `sendmmsg` paths that write per-message lengths - fail closed with `EOPNOTSUPP`. `accept` with a null address, and socket - creation, `connect`, `bind`, `listen`, `sendto`, and `sendmsg` continue to - work — they use copied inputs, scalar responses, or atomic `ADDFD_SEND`, none - of which write into workload memory. Some server workloads whose accept - wrappers request the peer address will therefore not run until the kernel - provides `WAIT_KILLABLE_RECV` (a distribution backport); outbound-oriented - workloads are unaffected. - -Input mediation, DNS/TCP authorization, and outer-fence enforcement are -identical in both modes. The selected mode is emitted in the sandbox -qualification output (`seccomp_listener_mode`). - -DNS uses an exact sandbox-local resolver at `127.0.0.53:53`. The driver sets that -nameserver without search domains, so names reach policy DNS as the workload -wrote them, and permits an unprivileged bind to port 53. UDP and TCP DNS requests -are forwarded through the supervisor, which applies hostname-based DNS policy. -The Podman driver supplies that resolver configuration as a driver-owned, -read-only secret mounted at `/etc/resolv.conf`; the workload remains on -`network=none` and receives no host aliases directly. For -`host.openshell.internal`, the supervisor returns the trusted concrete host -destination carried in its runtime descriptor rather than relying on Podman's -workload-side host-gateway injection. -DNS sender identity is explicitly unavailable: native writes can come from an -inheriting process or after exec, and neither the connecting binary nor a later -descriptor-owner snapshot proves who sent an already queued query. Consumers -must not use this unavailable identity to grant binary-specific access. TCP -connection authorization still uses decision-time binary identity. - -The sandbox retains only bounded DNS socket-admission records, consumes TCP -records on accept, and reclaims closed UDP records when capacity is reached. -The kernel delivers replies from the configured nameserver address, including -for strict musl and c-ares resolvers. No proxy environment variable, nftables -rule, or workload network namespace setup is part of enforcement. The supervisor -retries failed DNS accepts with backoff, preserving service across a channel -reconnect. - -External TCP opens wait at most 30 seconds for a supervisor decision, then fail -with `ETIMEDOUT` and release their worker quota. An approval is tied to the -original socket identity; replacing the descriptor during policy evaluation -cannot transfer that approval to another socket. - -The outer fence remains mandatory. If notification handling misses a syscall, -loses the supervisor, exceeds a bound, or encounters an unsupported socket -type, the request fails and the outer fence still blocks direct egress. - -CONNECT and absolute-form forward HTTP are explicit-proxy adapters over the same -egress pipeline. Each adapter normalizes its request into an egress intent, and -the shared authorization result carries the process evidence and endpoint -metadata used by destination validation and relay selection. Network action, -matched policy, endpoint configuration, and exact-host authorization are -evaluated as one atomic snapshot from one policy generation. Destination validation -returns an unopened connector so adapters retain their existing response and -upstream-dial timing. CONNECT prepares a generation-pinned relay context before -entering shared TLS-terminated or plaintext HTTP relays; non-HTTP traffic uses -the shared raw byte relay after the existing adapter gates. Forward HTTP retains -its guarded single-request relay while sharing authorization, request context, -policy-pinning, and destination boundaries. -Adapter-specific response and OCSF event shapes remain at the protocol boundary. -HTTP response framing and connection persistence are separate decisions. After -forwarding a complete closing response (explicit `Connection: close` or HTTP/1.0 -without keep-alive), the relay flushes and shuts down downstream writes before -ending the exchange, including TLS close notification. Response middleware -preserves this lifetime rule; persistent responses remain eligible for reuse. - -An explicit `protocol: tcp` endpoint with a valid DNS hostname opts into native -DNS and transparent TCP when the selected runtime advertises that substrate. -Hostless `allowed_ips` and literal-IP selectors remain available only to the -legacy explicit-proxy path when `protocol` is omitted. For an eligible DNS -name, the shared supervisor returns an epoch-scoped synthetic address and -publishes the expiring name, endpoint, ports, policy generation, and validated -real addresses as one correlation. A name absent from policy receives at most a -contract-free observation address for policy advisor proposals; see -[Security Policy](security-policy.md). A connection to that synthetic address is -captured before the bypass fence, mapped back to its workload process, authorized -through the same egress pipeline, and dialed only through the pinned addresses. -Omitted protocol endpoints retain explicit-proxy behavior. - -Policy DNS reserves `policy.local` inside each sandbox without requiring an -authored network endpoint or a trusted external lookup. The source-backed TCP -boundary routes plain HTTP on the reserved address and port 80 to the -supervisor's sandbox-local policy API; it does not open an upstream connection. -The API checks the effective agent proposal setting for every request. Policy -denials at the staged TCP authorization gate enter the supervisor's denial -aggregator with the resolved binary and destination, so the mechanistic mapper -can propose a scoped rule. Invalid mappings and destination validation failures -do not become policy proposals. - -Provider credential placeholders are resolved through the live provider state -for each HTTP request, after destination and L7 policy admission. A static -credential resolves only when the request host, port, and path match an endpoint -in that provider's effective profile. CONNECT, absolute-form forward HTTP, -request targets, headers, supported request bodies, SigV4 signing, and opted-in -WebSocket text rewriting use the same scoped resolver. Provider refresh swaps -credential values and endpoint bindings atomically. An invalid or unavailable -refresh revokes the previous static credential state instead of leaving a -partially active or last-known-good static set. Invalid metadata preserves the -supplied dynamic snapshot, while a fetch failure preserves the currently active -dynamic snapshot. - -Across the protected sandbox/supervisor channel, the provider environment revision remains -an opaque content fingerprint and has no numeric ordering semantics. The -network supervisor assigns a separate, connection-local monotonic generation -to each distinct environment it publishes. The process supervisor applies only -newer generations, which accepts descending fingerprint values while rejecting -duplicate or delayed supervisor messages. - -Gateway-managed refresh credentials use an opaque workload handle derived from -the sandbox, provider identity, credential key, refresh authorization epoch, -and canonical endpoint boundary. The handle remains stable while the gateway -rotates the short-lived value, so an already-running process keeps one -placeholder and each request resolves against the current token. Explicit -refresh reconfiguration, provider replacement or detachment, and endpoint -boundary changes produce a new handle and revoke the old one. Supervisors do -not retain old values for these handles. Public provider updates cannot replace -or delete the refresh-owned primary credential or co-minted outputs; internal -CAS rotation and explicit refresh lifecycle operations own those values. -Unmanaged static credentials retain the bounded revision-generation behavior. - -Route selection and policy evaluation use a syntax-only redacted request target; -they do not materialize real credentials. Cross-endpoint placeholder use returns -HTTP 403. After a WebSocket upgrade it closes the connection with policy -violation code 1008. Both paths emit a denied activity event and a detection -finding without logging the placeholder, environment key, secret, or query. - -For inspected HTTP traffic, the proxy can enforce REST method/path rules, -WebSocket upgrade and text-message rules, GraphQL operation rules, and -MCP method, tool, and supported params rules or generic JSON-RPC method rules -on sandbox-to-server request bodies. MCP and JSON-RPC inspection buffers -bounded request bodies. MCP `tools/call` tool names are checked against the -spec-recommended syntax by default before policy evaluation, with a per-endpoint -`mcp.strict_tool_names` compatibility opt-out. Generic JSON-RPC policies do not -support `params` matchers; generic JSON-RPC rules match only the method. -JSON-RPC responses and server-to-client MCP messages on response or SSE streams -are relayed but are not currently parsed for policy enforcement. - -Every `protocol: mcp` endpoint carries a canonical, nonempty `mcp.versions` allowlist drawn from OpenShell's exact revision registry: `2025-03-26`, `2025-06-18`, and `2025-11-25`. A policy author may omit the entire `mcp` object when using the other endpoint defaults, or omit `mcp.versions` while setting another MCP option. Both forms resolve immediately to the exact allowlist `["2025-11-25"]`; omission never means latest or all known revisions. Defaulting applies only when the corresponding YAML key is absent: `mcp: null`, `versions: null`, and an explicit `versions: []` are invalid. At protobuf ingress, an empty repeated field means omission and uses the same default because protobuf repeated fields do not preserve presence. Normalization stores and serializes the materialized allowlist in semantic order, so adding a supported revision to the registry never widens a previously normalized policy. An explicit nonempty allowlist remains available as an advanced compatibility or downgrade control. The registry is a closed set rather than a date range, so duplicate or padded values, unknown dates, and moving aliases such as `draft` or `latest` are rejected. The sessionless `2026-07-28` revision is not accepted until OpenShell supports its distinct per-request runtime contract. A version names a core protocol revision only; there is no policy syntax for layering a separately named SEP onto it. The registry owns immutable batch-shape metadata: `2025-03-26` permits nonempty same-side top-level JSON-RPC batches, which OpenShell's planned enforcement caps at 64 members, while `2025-06-18` and `2025-11-25` prohibit top-level arrays. These are declared profile facts, not current forwarding claims. The allowlist does not yet select request parsing or forwarding behavior. Later response-aware runtime state must observe the successful server response, require the selected revision to be in the allowlist, and apply that one exact profile without a union or fallback; OpenShell must not bind the client proposal in `initialize` as though it were the server-selected revision. - -For admitted HTTP requests, the proxy can run an ordered supervisor middleware -chain after L7 policy evaluation and before credential injection. Destination -host selectors choose the chain independently of the network rule that admitted -the request. Policy-local map keys identify configs, while built-in names or -operator-owned registration names identify implementations. - -Built-ins run in-process against a borrowed view of the chain's current HTTP -request state. Operator services retain the bounded protobuf/gRPC contract, and -the remote adapter materializes an owned HTTP evaluation only when a request -crosses that transport boundary. Both paths support bounded bidirectional -WebSocket sessions, so a manifest advertises capabilities independently of -transport. -When a stage ends, the remote adapter sends its terminal event, half-closes the -request stream, and briefly drains the response stream before releasing the -transport. This keeps a queued terminal event from being canceled with the -bidirectional RPC. -The runtime keeps three states distinct: host selection attaches policy configs, -manifest operation and phase bindings select the active chain, and the parsed -message type determines whether that chain can inspect an individual payload. -An attachment without a WebSocket binding is not a failed WebSocket stage. -Binary messages are outside the V1 text-message binding. Both cases pass through -with informational coverage telemetry rather than applying `on_error`. -The chain runner owns shared sequencing, deadlines, backpressure, and response -validation. `openshell-policy` validates policy-owned structure, and the active -middleware registry validates implementation-owned config. The generic -registry and chain runner live in `openshell-supervisor-middleware`; first-party -implementations live in `openshell-supervisor-middleware-builtins`. - -The selected middleware chain can also inspect the final HTTP response before -it returns to the workload. Stages select header-only, whole-body, or streaming -inspection independently. The relay owns response framing when body bytes can -change. Preflight exposes upstream `Content-Length`, `Content-Encoding`, and -`Content-Range` as read-only metadata, while the relay emits final framing -separately from middleware-visible headers. Stage failures follow policy-local -`on_error`; explicit denials always block delivery. Once delivery has started, -blocking aborts the response. - -The network supervisor represents the destination-selected request and response -pair as one `HttpMiddlewareExchange`. It retains the full chain, runner, request -identity, and policy generation while request and response bindings are selected -independently. The HTTP response adapter owns wire parsing, downstream commit -state, generation fences, framing, and transport error classification. The -generic middleware crate owns stage selection, remote stream lifecycle, ordered -body processing, limits, and result validation. - -The supervisor installs policy and middleware registry changes as one runtime -generation and preserves the last-known-good generation if preparation fails. -Policy-only updates reuse the connected registry, so an external middleware -outage cannot block unrelated policy changes. - -For authenticated operator middleware, the supervisor requests credentials by -registration name through `RefreshSandboxToken`. The gateway resolves names -against the effective policy and mints exact-audience credentials. The -supervisor keeps them in refreshable in-memory slots outside stable middleware -configuration, so rotation neither changes `config_revision` nor reconnects -the registry. Public custom-CA PEM travels with the stable registration. - -The slots live in a supervisor-owned `ExtensionCredentialStore` shared by every -gateway connection the supervisor opens, so the registry's clients and the -polling loop that rotates them observe the same credentials. Configuration -polling runs far more frequently than credentials expire, so the loop rotates -only when a credential is missing or has passed four fifths of its lifetime, -and bounds its sleep by the soonest rotation deadline. - -Middleware cannot observe injected credentials, introduce credential -placeholders, or mutate supervisor-owned credential, routing, or framing -headers. Body transformations are re-evaluated -against body-aware L7 policy before later stages or the upstream can observe -them. Requests, results, chain length, execution time, and diagnostics are -bounded; external free-form diagnostic text is not exposed in responses or -security logs. See -[Supervisor Middleware](../docs/extensibility/supervisor-middleware/index.mdx) for -an introduction, or the [configuration guide](../docs/extensibility/supervisor-middleware/configure.mdx) -for service registration and policy attachment. - -Inference providers use the same egress path as other external services. An -attached provider profile contributes endpoint and binary policy. The proxy -then resolves the provider's credential placeholder only when both policy and -the profile's endpoint binding authorize the native request. Model selection, -request shape, headers, streaming, and timeouts remain client concerns. - -In proxy-required networks, the supervisor chains upstream TLS tunnels through -a corporate forward proxy with HTTP CONNECT instead of connecting directly, -once policy and SSRF checks pass. Only TLS (CONNECT) egress is chained: -plain-HTTP requests always dial the destination directly, because forwarding -plain HTTP through a proxy requires absolute-form request forwarding rather -than CONNECT tunneling and is out of scope. The proxy configuration is an -operator-owned boundary delivered on the supervisor's command line -(`--upstream-proxy` and friends) by the compute driver; sandbox and template -environment — and `ENV` values baked into the sandbox image — cannot -influence it, since none of these can alter the argv the driver sets. The -conventional `HTTPS_PROXY`/`HTTP_PROXY`/`NO_PROXY` variables a sandbox -controls are ignored on this path. Operator `NO_PROXY` destinations and -loopback always dial directly; add driver-injected host aliases (e.g. -`host.containers.internal`) to the operator `NO_PROXY` list when the corporate -proxy cannot reach the container host. `NO_PROXY` matching is port-aware and -resolution-aware: an entry with a `:port` qualifier only bypasses that port, -and IP/CIDR entries also match hostnames through their validated resolved -addresses, with the direct dial limited to the addresses the entry contains. `http://` and `https://` proxy URLs in explicit -`scheme://host:port` form are supported — the scheme and port are both -required, and a path, query, or fragment is rejected. For an `https://` proxy -the supervisor wraps the connection to the proxy in TLS before the CONNECT -handshake, verifying the proxy certificate against the built-in and system -roots plus the optional operator CA bundle (see below). Local DNS resolution -and SSRF validation still run before the proxied dial, and the CONNECT -target sent to the corporate proxy is a validated resolved address, so the -proxy performs no DNS resolution of its own and the tunnel stays bound to -the answer that passed SSRF and `allowed_ips` validation. The hostname still -travels inside the tunnel (TLS SNI, application `Host`). In split-horizon -networks, point the gateway host at the corporate resolver so internal names -validate to their internal addresses; the `proxy_connect_by_hostname` -opt-in exists as a -last resort for proxies whose ACLs filter on hostnames and reject IP CONNECT -targets — with it, the proxy resolves the name itself and its ACLs become -the effective egress control for proxied TLS. (Resolving through the proxy's -own DNS view, e.g. DoH tunneled via CONNECT, is a possible future -enhancement and out of scope.) Workload proxy variables are removed from the -protected launch environment; transparent socket mediation does not depend on -them. - -The canonical main process receives the declared workload environment before -supervisor-only values are stripped and provider placeholders are injected. -Template environment is treated like user-provided sandbox environment. It can -shape the workload child, but it cannot override driver-controlled identity, -gateway endpoint, TLS, relay socket, proxy, provider, or supervisor coordination -variables. Drivers and the supervisor rewrite those reserved values after image -and template environment are considered. - -The configuration is fail-closed: a setting that is present but invalid — an -empty value, an unsupported or malformed proxy URL, an unreadable auth file or -CA bundle, a malformed credential, or an auth file, `NO_PROXY` list, or CA -bundle set while no proxy URL is configured — is fatal to supervisor startup -instead of being treated as unset, so a misconfiguration can never silently -degrade to direct dialing or unauthenticated proxy access. Only an omitted -argument means "no proxy". The driver validates the same rules at -sandbox-create time through validators shared with the supervisor -(`openshell_core::driver_utils::parse_upstream_proxy_url` and -`parse_upstream_proxy_credential`). - -An optional operator CA bundle (`--upstream-proxy-ca-bundle`, a supervisor-only -PEM path) extends the trust boundary for corporate proxies. A CA certificate is -not secret, but the supervisor is still its only configuration authority. It is -trusted in two places: the TLS handshake with an `https://` proxy, and — -because a TLS-intercepting proxy (mitmproxy, squid `ssl-bump`) re-signs -tunneled server certificates with the same CA — the sandbox combined trust -bundle (`write_ca_files`) and the L7 upstream re-encryption store -(`build_upstream_client_config`). Folding it into both means intercepted -upstream handshakes succeed and sandbox workload processes trust the re-signed -certificates; trusting it only for the proxy-listener handshake would leave -every intercepted upstream connection failing. The bundle is valid with either -an `http://` or `https://` proxy (an intercepting proxy can be reached over -plain HTTP) and is fail-closed: an unreadable or certificate-free file is fatal. - -How the bundle reaches the supervisor is driver-specific. Drivers that run the -supervisor locally bind-mount the operator's file. The Kubernetes driver cannot: -the supervisor Pod is scheduled remotely, and a trust anchor for every upstream -the sandbox reaches must not be sourced from the workload namespace, where it -would widen to anyone holding write access there. The gateway instead reads the -PEM from its own filesystem and stages it into the per-generation supervisor -bootstrap Secret, which is immutable, so the anchor cannot change underneath a -running sandbox. - -Proxy credentials are never embedded in the URL: an inline `user:pass@` is -rejected because it would be stored in `gateway.toml` and exposed in container -metadata. Operators supply credentials via `proxy_auth_file`; the driver -stages them as a supervisor-only secret mounted at a fixed path and passes only -that path on the supervisor's command line. The supervisor reads the -file and builds the `Proxy-Authorization: Basic` header; a credential that is -empty, contains control characters, or is not in `user:pass` form is fatal on -both sides. - -The VM driver starts `openshell-supervisor` on the host and -`openshell-sandbox` as capability-free guest PID 1. Corporate proxy arguments, -credentials, private CA keys, policy, and gateway credentials stay host-side. -Both libkrun and QEMU guests are NIC-less; intercepted workload connections -cross the authenticated vsock channel. A gateway-host proxy is addressed as -`host.openshell.internal`, which the host supervisor normalizes to `127.0.0.1`. - -The Docker driver runs `openshell-supervisor` in a separate companion container. -Its private named volume contains supervisor bootstrap and channel material. -The workload container receives only `openshell-sandbox`, public interception -CA material, and the other sandbox half of the authenticated channel. - -For Kubernetes, the operator configures a Secret name and key rather than a -gateway-host file path. Kubernetes projects that Secret only into the separate -supervisor Pod. The sandbox Pod never mounts corporate-proxy credentials -or the interception CA private key. - -The Basic header travels over the plain-TCP connection to the `http://` proxy, -so it is readable on the network path between sandbox host and proxy. -Configuring `proxy_auth_file` therefore requires the explicit opt-in -`proxy_auth_allow_insecure = true`. Both the -driver (at sandbox-create time) and the supervisor (at startup) reject an -auth file without the acknowledgement, and the acknowledgement without an -auth file, so credentials are never sent in cleartext without an explicit -operator decision. - -## Tool server connection status - -A sandbox can be `Ready` while a call to an external tool server fails. For configured endpoints that use MCP over HTTP, OpenShell records the last observed network result beside the endpoint's address. Users can identify the server and distinguish policy denial, unavailable credentials, TLS or network failure, and an upstream HTTP rejection without combining client and supervisor logs. These observations do not affect sandbox lifecycle readiness. - -```mermaid -flowchart LR - Traffic[Calls to configured tool servers] --> Observer[Observe network result] - Observer --> Reporter[Background reporter] - Reporter --> Gateway[Validate and store results] - Gateway --> Status[Endpoint address, last result, report time] -``` - -The gateway exposes one record per configured endpoint in `Sandbox.status.endpoint_statuses`. Each record contains its address, an opaque identifier, a typed result, and the time the gateway accepted that result. The address remains available when evidence resets to `NoObservedExchange`. Ordinary sandbox conditions continue to describe lifecycle and platform state. - -Network observers send only the endpoint identifier and a fixed result classification, without credentials, payloads, or raw upstream errors. Requests capture observation authority before selecting policy or credentials, then bind the selected policy hash and provider revision to that capture. Observation handles also identify the installed endpoint inventory and supervisor authority, so a concurrent update cannot attribute an old request to a new configuration, and reinstalling the same configuration cannot revive an obsolete request. The sandbox reporter coalesces observations and retries an immutable snapshot through `ReportEndpointStatus`. Bounded delivery can drop observations, so endpoint status is not a complete request history. - -The gateway validates the reporting supervisor's session, configuration revisions, and report sequence before atomically storing results. Global policy writes share the report's synchronization boundary. An effective policy change resets all endpoint evidence; repeated acknowledgements and metadata-only policy revisions preserve it. A provider environment change resets evidence for endpoints that depend on those credentials. Supervisor disconnection or replacement and gateway restart also invalidate observation authority. Identical report retries leave timestamps unchanged. After an inventory reset, still-valid pending evidence can be accepted again if an acknowledgement was lost, advancing the report time without another exchange. - -These are passive observations with no expiry. `HttpResponseReceived` means the server returned a final HTTP status below 400, including a protocol upgrade; that response can still contain an MCP error. Informational responses alone do not establish success. MCP protocol-version and request-body policy rejections produce `PolicyDenied`. A failure before the HTTP path is known updates status only when the host and port identify one distinct endpoint. Ambiguous failures remain in structured events and logs. Results combine callers and effective ports for an endpoint. Consumers that need current tool availability must verify an actual operation. The sandbox management guide explains the public fields and results. - -## Credentials - -Provider credentials are stored at the gateway and fetched by the supervisor at -runtime. The supervisor injects resolved environment variables into the initial -agent process and SSH child processes. Driver-controlled environment variables -override template values so sandbox images cannot spoof identity, callback, or -relay settings. - -Supervisor bootstrap identity and provider workload-identity sockets never -enter the sandbox workload. The authenticated channel carries only the -policy-authorized provider environment intended for child launch and public -trust material intended for TLS clients. - -Credential placeholders in mediated HTTP requests can be resolved by the proxy -when policy allows the target endpoint. Secrets must not be logged in OCSF or -plain tracing output. The supervisor uses revision-scoped -placeholders for unmanaged rotating credentials and identity-stable opaque -handles for gateway-managed refresh credentials. Provider environment keys -beginning with `v_` or `s<64 lowercase hex characters>_` are reserved -for those placeholder namespaces. - -Provider profiles can also declare dynamic token grants. For matching HTTP -endpoints, the supervisor obtains or exchanges OAuth2 access tokens, caches -them, and injects them before forwarding the request. `client_credentials` -grants use the supervisor SPIFFE JWT-SVID directly as the client assertion. -`token_exchange` grants ask the gateway to broker an intermediate token using a -stored provider subject credential and the gateway's own SPIFFE JWT-SVID; the -supervisor then exchanges that intermediate token for the final upstream token -using its own JWT-SVID. The gateway validates that its own JWT-SVID has the -requested audience, a SPIFFE subject, and a non-expired `exp` claim when -present. It also validates that the stored subject credential is declared by the -provider profile, and that the supervisor JWT-SVID is a well-formed -three-segment JWT with a SPIFFE subject in the same trust domain as the gateway -SVID. The gateway verifies the supervisor JWT-SVID signature with JWT bundles -fetched from its SPIFFE Workload API. Token grant endpoints are HTTPS-only -except for loopback and Kubernetes service DNS hosts, and returned access tokens -must be bearer-compatible before they are cached or injected. Token response -lifetimes are capped and cached with an expiry margin unless a profile supplies -an explicit cache TTL override. Cache entries are scoped by the sandbox provider -environment revision so provider credential updates miss the old token cache -without changing endpoint matching semantics. Gateway-brokered intermediate -tokens are cached separately by provider resource version, supervisor SPIFFE -subject, and gateway SPIFFE subject, and their cache lifetime is capped by the -intermediate token response, stored subject-token expiry, and supervisor SVID -expiry. - -For AWS endpoints that require request-level signing, the proxy supports SigV4 -re-signing. When `credential_signing: sigv4` is set on an L7 endpoint, the proxy -strips the client's placeholder-based AWS auth headers, re-signs with real -credentials from the provider, and forwards the request upstream. The signing -endpoint must have a credential source before the policy generation activates: -an attached endpoint-bearing AWS profile whose boundary covers the endpoint, or -an attached endpointless AWS profile explicitly named by the endpoint's -`credential_binding.provider`. Policy activation rejects missing or mismatched -sources atomically. The signing mode is auto-detected from the client SDK's -`x-amz-content-sha256` header: - -- **Signed body** (hex hash): buffers the request body, computes its SHA-256, - and includes the hash in the signature. Used by Bedrock and most AWS services. -- **Streaming unsigned** (`STREAMING-UNSIGNED-PAYLOAD-TRAILER`): signs headers - only and streams the body through without buffering. Used by S3 uploads with - `aws-chunked` encoding. -- **Unsigned payload** (`UNSIGNED-PAYLOAD`): signs headers only with no body - hash. Used by S3 over HTTPS for non-chunked requests. - -Chunk-signed streaming modes (`STREAMING-AWS4-HMAC-SHA256-PAYLOAD` and other -`STREAMING-*` variants) are rejected — the proxy cannot reproduce per-chunk -signatures. Use `sigv4:no_body` for those clients. - -Two explicit overrides are available: `credential_signing: sigv4:body` (always -buffer and hash) and `sigv4:no_body` (always unsigned). The `Expect: -100-continue` header is handled within the SigV4 path so clients like boto3 -transmit the body before the proxy forwards to upstream. - -The AWS region is extracted from the endpoint hostname. For non-standard -endpoints (VPC endpoints, custom proxies), set `signing_region` in the policy -endpoint to provide an explicit override. The proxy rejects requests when -neither hostname extraction nor `signing_region` yields a region. - -`credential_signing` and `request_body_credential_rewrite` are mutually -exclusive on the same endpoint. The policy validator rejects policies that -set both. - -## Connect and Logs - -The supervisor runs an SSH server on a Unix socket inside the sandbox. The -gateway reaches it through the outbound supervisor relay, not by dialing the -sandbox workload directly. The relay supports: - -- Attachment to the canonical main process through the `openshell-main` SSH - subsystem. The supervisor owns its retained PTY or pipes, a 1 MiB replay - buffer, and a single stdin lease across client disconnects. Ctrl-C interrupts - the foreground process. For read-only attachments, Ctrl-C only exits the - current viewer. -- Supervised CLI attachment. After an established SSH transport fails, the CLI - remains alive, requests a fresh SSH session from the gateway, and reattaches - to the same canonical main process within a bounded recovery window. It does - not stop or restart the sandbox to recover the client connection. The same - deadline bounds replacement-session RPCs. Process-targeted termination is - forwarded to the SSH child, which the CLI reaps before exiting. -- Independent interactive shell sessions. -- Command execution. Commands run through a login shell (`bash -lc`) by default, - so the first of the user's `.bash_profile`, `.bash_login`, or `.profile` is - sourced (and `.bashrc` only if that file sources it). Callers set - `ExecSandboxRequest.no_login_shell` to skip those files; the gateway signals - this to the supervisor over the SSH `OPENSHELL_NO_LOGIN_SHELL` env request, - which selects `bash -c` instead of `bash -lc`. Note `bash -c` still reads - `BASH_ENV` when the child environment sets it. -- Tar-based file sync. -- Port forwarding where supported by the CLI/TUI surface. -- Persistent HTTP and WebSocket service routing through gateway-managed - `ServiceEndpoint` records. `CreateSandboxRequest.service_exposures` registers - named or unnamed endpoints as part of sandbox creation, and the gateway - returns their routed URLs keyed by service name. The empty key identifies the - unnamed endpoint. Routing starts only while the sandbox is ready. `sandbox - create --expose PORT` uses the unnamed create-time endpoint and keeps the - sandbox. The standalone service API can add, update, or remove endpoints - later. - -Sandbox logs are emitted locally and can also be pushed back to the gateway. -Security-relevant sandbox behavior uses OCSF structured events; internal -diagnostics use ordinary tracing. -The OCSF device describes the sandbox environment, with type ID Other and type -label `Sandbox`; its operating system is a separate attribute. -HTTP Activity records contain a request or response; early rejections with only -connection context use Network Activity. Producer regression tests validate -required fields and `at_least_one` constraints against the vendored OCSF 1.8 -schemas. -Network Activity records identify at least one observed endpoint; connection -failures retain their known peer or listening endpoint. -Configuration diagnostics use Config State Change. -Unix socket relay and relay-control notifications, plus proxy and mediation -failures without an observed network endpoint, use Base Event. - -## Policy Proposals - -When an L4 CONNECT is denied, the proxy emits a `DenialEvent`. The denial -aggregator batches these events and flushes summaries to the gateway every 10 -seconds (configurable via `OPENSHELL_DENIAL_FLUSH_INTERVAL_SECS`). The gateway -runs them through the mechanistic mapper, which generates a pending -`NetworkPolicyRule` proposal visible under `openshell rule get --status pending`. - -L7 denials (HTTP 403 from method/path rules) are intentionally excluded from -mechanistic mapping. L4 denials carry only `host:port`, which a deterministic mapper can handle. -L7 denials carry method, path, query, and body context. The agent loop reads -the structured 403 and authors the narrowest rule. Mechanistically mapping L7 -would either over-broaden rules or require path-templating logic that rots -quickly. - -## Configuration Admission - -Gateway-managed supervisors reconcile configuration before launching the main -process or exposing workload services. Admission covers the effective policy, -provider layers, credential bindings, and gateway-derived provenance. Explicit -user and global policy precedence is unchanged; an image without a policy uses -the restrictive baseline. An invalid image policy does not become a launchable -default. - -The gateway tracks configuration admission independently of compute health. -A blocked startup remains `Provisioning` with a `ConfigurationInvalid` readiness -condition, even when the container backend reports readiness. Gateway management -operations remain available. The TUI summarizes configuration rejection in sandbox -NOTES alongside active port forwards; the detail view wraps the full diagnostic, -which is also available through sandbox inspection. -Replacing the policy or repairing providers allows -the same supervisor to reconcile and launch; it does not recreate the sandbox. -Startup retries continue reporting readiness, but unchanged configuration rejections -produce only one log event. A changed configuration or diagnostic emits a new -rejection event; successful repair emits a recovery event. -The gateway gives each initial provisioning attempt and explicit restart a -300-second repair window. Persisted configuration-source clocks reset the window -from the latest effective stored change, including settings deletion and provider -attachment changes. The first accepted rejection for that generation grants one -full window; repeated reports and reconnects do not extend it. Ready disarms the -timer. Failed desired updates to a running sandbox do not arm it. - -A leader-owned scan runs independently of driver inventory. Expiry records -`Error`/`ProvisioningTimedOut` before reclaiming compute; cleanup progress and -backoff survive restart. Late runtime reports cannot replace that result. The -record and restartable storage survive cleanup, including for ephemeral creates. -Explicit start is blocked while cleanup is pending, then creates a fresh attempt -using the latest configuration. Configuration edits alone never restart an -expired sandbox. Legacy provisioning records receive one persisted rollout -window. Cross-object configuration serialization uses the gateway's existing -single-writer guard; enabling concurrent configuration writers still requires -the database-backed invariant work tracked by #1255. - -Docker startup health remains unready during policy quarantine. A failed probe -does not terminate a live provisioning supervisor; the gateway deadline owns -that decision. Cleanup cancels pending driver startup before stopping compute -so a late startup failure cannot remove retained workload storage. - -Static policy fields can be replaced before the first accepted activation. -A durable first-activation marker closes this repair window permanently, including -across stop/start and later rejected configurations. Legacy records without the -marker retain static-field immutability. -Admission validates policy composition; image and host setup failures, such as -an unresolved OCI user or unavailable isolation facilities, retain their existing -startup error behavior. - -Acceptance identifies the effective policy hash/version, configuration revision, -provider-environment revision, and reporting supervisor instance. Startup captures -the matching provider environment and constructs the runtime before reporting -acceptance. Live reconciliation begins only after the main process has spawned, -so it cannot replace the configuration captured for that launch. Restart resets -admission and requires a fresh accepted configuration. Permanent gateway errors -and exhausted transient retries terminate startup; each RPC attempt has a -10-second deadline, including acceptance reports; only acknowledged configuration -rejections wait for repair within the gateway's provisioning deadline. Image discovery uses the authenticated -sandbox boundary control request deadline. - -Policy and provider refreshes are prepared before publication. Publication -invalidates prior policy guards before exposing new provider material and swaps -the policy under the same publication locks. Rejected candidates cannot install -their credentials alongside the previous policy. Existing runtime fail-closed -checks remain necessary for in-flight traffic and invalid live updates. - -The supervisor reads the workload image policy through an authenticated, -read-only `DiscoverPolicy` boundary request before attaching or launching the -workload. The boundary reads only the well-known policy paths, bounds the response, -and distinguishes missing policy from unreadable or invalid content. The supervisor -validates this candidate with gateway provider composition and obtains admission -before `Attach`, `Confirm`, networking startup, and `StartAgent`. Workload image -environment variables cannot configure the isolated supervisor. - -## Policy Revision Acknowledgement - -When the supervisor loads a sandbox-scoped policy from the gateway, it retains -the version, hash, source, and configuration revision returned with that exact -policy snapshot. After the OPA engine is built successfully, the supervisor -reports that revision as `LOADED`, which advances -`SandboxStatus.current_policy_version` and moves the revision out of `Pending`. -If policy construction fails, it reports the captured revision as `FAILED` with -the original construction error. It never infers revision identity by comparing -policy structure. - -This holds even when the initial policy is enriched with baseline paths during -startup: the enriched revision the supervisor synced back to the gateway is the -revision it acknowledges, so a successfully constructed initial policy never -remains `Pending`. If the first poll returns a different revision, the supervisor -processes it through the normal reload path instead of treating it as already -loaded. - -A newer sandbox-scoped revision can carry the same non-empty effective policy -hash as the currently loaded revision, for example when provenance changes -without changing enforcement content. The supervisor acknowledges that newer -revision without reloading identical policy. If the revision also requires -middleware or policy-runtime reconciliation, acknowledgement waits until that -reconciliation succeeds. Global policies, local overrides, equal or older -versions, and different hashes do not use this shortcut. Success telemetry is -emitted only after the gateway accepts the resulting loaded-status report. - -Policy status delivery uses a FIFO background worker. Retryable delivery -failures retain the ordered update and retry with capped exponential backoff; -terminal errors are logged and discarded. The outbox is nonblocking and does -not discard updates because of a fixed queue capacity, so status endpoint -outages cannot block policy polling, enforcement, settings, or provider -refreshes and cannot permanently lose the initial acknowledgement. - -Only sandbox-scoped revisions (`PolicySource::Sandbox`, version greater than -zero) use the policy revision acknowledgement API. Global policies use the -configuration admission contract without a sandbox policy revision acknowledgement. -Local Rego/data overrides remain available for standalone development; combining -them with a gateway-managed sandbox is rejected because the gateway cannot admit -the runtime policy it would enforce. - -## Failure Behavior - -- If gateway config polling fails, the sandbox keeps its last-known-good policy. -- If a live policy or middleware-registry update is invalid, the supervisor - rejects the update and keeps the current runtime pair. -- If an operator-run middleware call fails, the selected config's `on_error` - behavior decides whether to deny the request or continue without that stage. -- Existing raw byte streams are connection scoped. Dynamic policy changes apply - to new connections or the next parsed HTTP request where the proxy can safely - re-evaluate. -- If the supervisor relay drops, the sandbox stops the canonical agent and exec - process groups, rejects new runtime operations, and closes mediated streams. - A replacement supervisor has 30 seconds to authenticate, replay the identical - attach, and reconfirm the boundary. Successful confirmation resumes the - process tree; otherwise the sandbox sends `SIGTERM`, waits the normal stop - grace period, sends `SIGKILL` to survivors, and makes the session terminal. - Explicit supervisor shutdown uses the same terminal transition and requires - an acknowledgement before treating the boundary as stopped. -- If the canonical main process exits, the supervisor durably reports the - normalized result immediately. A foreground create declares a one-shot main - attachment, so the supervisor accepts it even after a fast process exits, - sends the retained output and SSH exit status, waits for the peer's channel - close, and then finalizes ephemeral cleanup. With no declared or active - attachment, it finalizes and exits without a grace period. The gateway waits - for that finalized supervisor session to disconnect before deleting an - ephemeral sandbox. Exit code 0 records - `Completed/MainProcessCompleted`; nonzero and signal-normalized exits record - `Error/MainProcessFailed`. Infrastructure failures also use `Error`, with a - distinct condition reason and no fabricated canonical-process result. Runtime - restart policies must not replace the canonical process. - -## Shared Boundary Primitives - -`openshell-isolation-interface` owns the common boundary protocol and Linux -mechanisms. Drivers provide the protected transport and immutable resource -identity; they do not implement their own process or network protocol. All -remote traffic uses one mutually authenticated gRPC connection. Independent -streams carry process control, exec output, and TCP bytes; a persistent -`Mediate` stream carries DNS queries and supervisor-produced answers. There is -no alternate raw-TLS application protocol or general UDP framing. - -The shared process-signal mediator resolves each positive target PID or TID to -its thread-group leader, excludes the sandbox leader, retains a pidfd, and sends -the signal through that descriptor. It never continues the original numeric-PID -syscall after inspection. This prevents TID aliases or PID reuse from turning an -agent signal into a signal to the sandbox. Ordinary mediated `kill` reports the -broker as its sender, not the original calling agent's `SI_USER` identity. -Queued signals preserve permitted application siginfo payloads; they cannot -forge kernel-generated or `SI_TKILL` codes. Programs requiring original sender -identity must account for this mediation boundary. diff --git a/architecture/security-policy.md b/architecture/security-policy.md deleted file mode 100644 index 36a919573a..0000000000 --- a/architecture/security-policy.md +++ /dev/null @@ -1,511 +0,0 @@ -# Security Policy - -OpenShell policy defines what a sandboxed agent can access. The policy is -enforced inside each sandbox by kernel controls, process setup, and the local -policy proxy. The gateway stores and delivers policy, but it does not make -per-request egress decisions. - -For the field-by-field YAML reference, use -[Policy Schema Reference](../docs/how-it-works/policies/schema.mdx). - -## Policy Areas - -| Area | Enforcement | -|---|---| -| Filesystem | Landlock restricts read-only and read-write paths. | -| Process | The supervisor launches the agent as an unprivileged user with reduced capabilities. | -| Network | The proxy evaluates destination, port, calling binary, and optional L7 rules. | -| Provider access | Attached provider profiles contribute endpoint and binary rules; credentials remain bound to profile-authorized endpoints. | -| Runtime settings | Typed settings are delivered with policy and can be global or sandbox scoped. | - -Filesystem and process policy are startup-time controls. Network policy is -dynamic and can be hot-reloaded when the new policy validates successfully. - -### Authored policy boundary - -`openshell-policy-schema` is the sole owner of the authored YAML and JSON -representation. It preserves authored distinctions such as an absent -`filesystem_policy` versus an explicitly empty object, rejects duplicate keys, -and applies parser budgets while noyalib constructs the document. It also owns -pure language semantics such as access presets, MCP revision vocabulary, -effective ports and rule names, protocol classification, and lexical policy -path normalization. - -Consumers project that syntax into purpose-specific models. `openshell-policy` -owns protobuf conversion, composition, merge behavior, raw-protobuf checks, and -validation that depends on runtime components. Both the proposal-risk prover and -standalone containment checker project the shared schema into their own models -using the same fail-closed parser as the runtime. The parser -requires `version: 1` and rejects managed annotations and every unknown field -before any consumer-specific projection runs. There is no permissive parsing -profile: unsupported policy fields always invalidate the document. Middleware `config`, query and persisted-query names, and recursive MCP -parameter names are open user-data maps rather than schema extensions. - -Before applying Landlock, the supervisor enriches baseline filesystem paths that -the runtime needs. Missing baseline paths are skipped so one absent runtime path -does not weaken the whole ruleset. When GPU devices are present, GPU baseline -enrichment adds existing GPU device nodes as read-write paths and promotes -`/proc` to read-write because CUDA workloads write thread metadata under -`/proc//task//comm`. - -Landlock rules are tailored to the inode type reported by the already-opened -path descriptor. Directories retain the requested directory and file rights; -regular files, device nodes, sockets, and other non-directories retain only -file-compatible rights. This avoids rejecting valid mixed-path policies without -weakening `hard_requirement`: genuine unsupported ABI capabilities and -preparation failures still fail sandbox startup. - -## Network Decisions - -Ordinary network traffic follows this order: - -1. Force traffic through the sandbox proxy with namespace and seccomp controls. -2. Identify the calling binary and compare its trusted identity. -3. Reject hard-blocked destinations, including unsafe internal IP ranges unless - explicitly allowed. -4. Match the destination and binary against network policy blocks. -5. Apply optional HTTP/L7 rules for endpoints that enable protocol inspection. -6. Allow, deny, audit, or log according to the matched policy. - -Explicit deny and hardening checks win over allow rules. If no rule matches, the -request is denied. - -## Host Wildcards - -Network endpoint `host` patterns accept a `*` wildcard inside the first DNS -label and as an entire middle DNS label. The OPA runtime matches with a `.` -label boundary, so a wildcard never spans dots. The validator enforces the same -boundary so that policy load fails fast instead of silently mismatching at the -proxy. - -| Pattern | Accepted | Example match | Notes | -|---|---|---|---| -| `*.example.com` | Yes | `api.example.com` | Single first label of any value. | -| `**.example.com` | Yes | `a.b.example.com` | Recursive wildcard as the entire first label. | -| `*-aiplatform.googleapis.com` | Yes | `us-central1-aiplatform.googleapis.com` | Intra-label wildcard inside the first DNS label. | -| `*.s3.*.amazonaws.com` | Yes | `bucket.s3.us-east-1.amazonaws.com` | Middle-label `*` matches exactly one DNS label. | -| `*` or `**` | No | — | Matches every host. | -| `*.com`, `**.com` | No | — | TLD wildcards (`labels <= 2`). | -| `foo.us-*.example.com` | No | — | Partial middle-label wildcards are not allowed. | -| `foo.**.example.com` | No | — | Recursive wildcard outside the first label is not allowed. | -| `foo**.example.com` | No | — | Recursive `**` mixed inside a label; allowed only as the entire first label. | - -Validation rejects the disallowed patterns at policy load time. OPA load errors omit the supplied hostname. Exact hosts and IP addresses do not use this wildcard-validation path. - -## TLS and L7 Inspection - -For HTTP endpoints that need request-level controls, the proxy can terminate TLS -with the sandbox's ephemeral CA and inspect method/path or protocol-specific -metadata before forwarding. The proxy also supports credential injection on -terminated HTTP streams when policy allows the endpoint. - -Static provider credentials have an independent endpoint-binding boundary. -Provider profile endpoints define that boundary by default. An endpointless -profile can delegate binding authority to sandbox policy through an endpoint -that names the concrete attached provider instance. The gateway rejects -unattached, profileless, endpointful, and gateway-global uses of that policy -binding. Policy endpoint changes rotate the provider-environment revision so -the supervisor installs policy and credential binding snapshots atomically. - -Raw streams and long-lived response bodies are connection scoped. Policy -generation changes close relays pinned to the previous generation instead of -allowing them to continue under stale authorization. HTTP upgrades switch to -raw relay by default. A `protocol: rest` endpoint can opt in to -`websocket_credential_rewrite` for client-to-server WebSocket text messages -after an allowed `101` upgrade; server-to-client traffic and all other upgraded -protocols remain raw passthrough. - -A `protocol: tcp` hostname is a connection-routing constraint, not an -application-authority boundary. Transparent capture validates the approved DNS -name, pinned destination address, port, and calling binary before opening the -stream, but it does not inspect TLS SNI, HTTP `Host`, or another protocol-level -destination. Compatible shared infrastructure can therefore let a client -select another tenant, virtual host, or service behind the approved front door. - -## Credentialed Endpoints - -OpenShell keeps provider credentials on paths it can inspect or rewrite by -default. The gateway derives credential provenance from the attached providers -and stamps it onto the effective policy at composition time. This provenance is -internal, contains no credential identifiers or values, and is never trusted -from user-authored policy. - -Every evaluation clears provenance across the whole policy and re-derives it -from two sources: - -- the endpoints of attached provider profiles that carry credentials, and -- the valid `credential_binding` entries of the sandbox policy that name an - attached provider whose profile is endpointless. - -A binding reduces to a host and port scope only. Dropping the path is -deliberate: a path is not observable on an L4 or `tls: skip` endpoint, so a -path-scoped derivation would omit the marker on exactly the surfaces the -uninspected-credential gate exists to catch. A malformed binding — empty -provider, missing host, or a port outside `1..=65535` — fails the evaluation -instead of contributing a scope. Bindings naming an endpointful profile or an -unattached provider contribute nothing; the gateway rejects those uses -separately. - -Both sources merge into one deduplicated scope set, and each endpoint is -stamped once per evaluation from that set, so binding-derived scopes reach the -same gates as profile-derived ones. The stamp is an assignment, not an -accumulation, so an endpoint that stops matching a credentialed scope — or -whose binding was removed — loses its marker in the same pass. This must remain -a full recomputation: a delta-based derivation would let a series of -individually valid edits reach a state no single edit would have admitted. - -Credentialed L4-only and `tls: skip` endpoints fail policy validation unless the -public `allow_uninspected_credentials` escape hatch is explicitly enabled. The -flag defaults to `false` and is security-flagged in policy approval flows. -Incremental merges only ever add the flag to a matching endpoint; clearing it -requires removing the endpoint or replacing the policy. - -Image discovery may persist a desired policy for repair, but does not authorize -workload activation. The gateway applies the credential gate after full provider -composition and provenance derivation. A rejected effective configuration keeps -startup blocked with a bounded diagnostic; the supervisor waits for management -repair instead of launching with connection-time denials or a fallback policy. -Accepted runtime state includes the matching provider-environment revision, so -policy and credential updates cannot activate independently. - -The network supervisor independently enforces the same boundary. Credentialed -WebSocket upgrades use the parsed relay, binary frames fail closed, and text -placeholders require rewrite. REST bodies continue streaming when body rewrite is disabled. The relay holds -complete placeholder candidates until a request-scoped metadata snapshot can -classify them. Authoritatively unknown keys and valid credentials with a current -binding pass unchanged, including references bound to the destination itself. -Revoked or invalid identities and unavailable classification state -fail closed. Classification never substitutes secret values and shares the request's -credential revision; stale credential or policy generations terminate forwarding. -Header rewriting scans only headers. Body bytes received in the initial proxy -read follow the same body classifier or explicit rewriter as later reads. -Candidates are limited to 4096 wire bytes, including percent encoding. Malformed -or oversized candidates fail closed; HTTP trailers retain prefix-based rejection. Explicitly opted-in -endpoints retain raw passthrough behavior. - -Body denials return `credential_placeholder_in_request_body` in a local HTTP 403 -response and discard any partially written upstream request. Safe preceding bytes -may already have reached upstream. - -Denials emit both the relevant network activity and a detection finding. Events -identify only the destination, policy, traffic surface, and controlled denial reason; they never include -credential names, placeholders, body content, or secret values. - -Credential provenance is gateway-derived and deliberately absent from the policy -YAML schema, so it does not survive a policy that never transits the gateway. -Gateway-delivered policy is the authoritative source for this control, and a -policy without provenance applies neither the raw-tunnel refusal nor the -WebSocket binary-frame refusal. The request-body backstop still applies, because -it keys off the presence of a secret resolver rather than endpoint provenance. - -Two supervisor-local paths load a policy without provenance. A supervisor -booting from an explicitly provisioned policy file has a bounded window before -that policy is resynchronized to the gateway, which then serves a stamped -effective policy. An explicit supervisor Rego and data override is permanent, -because gateway revisions are observed for settings and providers but never -replace the local policy. Workload-image files and environment variables cannot -configure the separately isolated supervisor. When a supervisor override is -combined with injected provider credentials, the supervisor emits a -high-severity detection finding at startup naming the inactive controls. - -## Policy Load Diagnostics - -`OpaEngine` bounds the error messages returned when loading or reloading policy from files, strings, or protobuf. Each message contains at most eight error items and 512 UTF-8 bytes, including its heading, separators, and any `additional violations omitted` marker. The loader reports complete, fixed categories and discards authored names, values, paths, source snippets, and nested error chains. - -Typed validation categories distinguish process identity, filesystem paths and limits, Landlock compatibility, endpoint hosts and ports, credential signing and rewriting, MCP configuration, and middleware configuration. Opaque L7 errors identify the protocol-configuration or policy-validation stage; endpoint conflicts report ambiguous selectors. YAML errors retain a fixed parser category and numeric line and column when available. File I/O, Rego loading, and internal policy-data errors use fixed messages. - -A candidate rejected during validation does not replace the active engine or advance its generation. The supervisor separately applies `policy_validation_failure_mode` and may publish a quarantine generation as described below. The diagnostic bounds cover returned OPA load errors; accepted-policy warnings, runtime request diagnostics, and gateway-authored policy parser messages have separate reporting contracts. - -## Live Updates - -The gateway stores sandbox-authored policy revisions separately from derived -effective sandbox configuration. Effective configuration can include -gateway-global policy overrides and provider-profile policy layers. The -supervisor polls for config revisions and attempts to load new dynamic policy -into the in-process OPA engine; CLI reads of the latest sandbox policy use the -same effective configuration path. - -The OPA loader checks the object and list shapes of raw policy data before injecting runtime fields, normalizing values, or expanding access presets. It rejects the first malformed container with a fixed structural error that excludes authored keys and values. This check preserves valid versionless OPA data and runtime-only fields. A rejected OPA engine reload leaves that engine's installed policy, generation, and decisions unchanged; the supervisor separately applies its configured runtime rejection mode. - -After validating L7 rules, the OPA loader converts nonempty string query and MCP parameter matchers into explicit `glob` objects, including MCP `tool` aliases and deny rules. Matchers representable in both YAML and protobuf therefore expose the same representation to endpoint configuration consumers. Already lowered `glob` and `any` matchers retain their values across reloads; normalization preserves runtime endpoint provenance. Empty scalar query matchers remain an OPA-only form because Rego gives them different behavior from empty `glob` objects. - -An explicitly supplied MCP rule `params` value must be a map in both allow and deny rules. Omit the field when using only the `tool` alias; `params: null` is rejected before alias lowering. A rejected raw policy reload preserves the active evaluator and its generation. - -The supervisor validates complete effective policy generations before -activation. Overlapping endpoint selectors may contribute request allow and -deny rules only when their connection and request-processing metadata agree; -conflicting TLS, destination, credential, parser, or enforcement metadata -rejects the complete generation. Plain L4 endpoints do not contribute -request-processing metadata, so they may overlap an L7 endpoint when their -connection metadata agrees. When request paths overlap, a path endpoint with a -higher specificity rank deterministically overrides broader request-processing -metadata. Equally specific overlapping endpoints must agree. - -Endpoint `tls`, `enforcement`, and `access` use protobuf enums and retain their -named YAML spellings. `tls` admits only an omitted value, meaning auto-detect -and terminate for inspection, or `skip`; every other value, including the -removed `terminate` and `passthrough` spellings, is rejected. `protocol` -remains a string so the supported protocol set can evolve, but every ingress -validates it before persistence or activation. -The supervisor also refuses unknown enum numbers and protocol values -defensively; an unrecognized enforcement value never falls back to audit. - -Gateway mutation paths validate the complete effective candidate before -persistence when the affected sandbox scope is known. Direct replacements, -incremental merges and approvals, provider attachment, and profile fanout reject -ambiguity atomically, without creating an invalid revision or partially -activating an update. Supervisor validation remains the defense-in-depth -boundary for startup, concurrent changes, and sources outside those mutations. - -L7 allow and deny append operations carry an explicit rule target and the complete affected binary and port scope. The merge engine resolves one non-provider endpoint within that rule, optionally by exact endpoint path, and compares both scope sets before mutation. A partial declaration, ambiguous target, or changed scope rejects the batch before revision persistence. The declaration records operator intent; it does not grant policy-writing authority or change the stored binary and port sets. - -The `[openshell.gateway] policy_validation_failure_mode` configuration controls -candidates rejected by supervisor runtime validation. Gateway preflight -rejections never become generations and leave the active policy unchanged. The -runtime mode defaults to `fail_closed`, which publishes a quarantine generation, -denies new egress, invalidates existing relays, and leaves the previous policy -inactive. Operators may explicitly select -`retain_last_valid`, which keeps the previous generation active. With no -previous valid generation, the effective mode remains `fail_closed` regardless -of the configured mode. The gateway distributes this startup configuration to -sandbox supervisors with each effective policy snapshot. OCSF configuration and finding events state the -candidate version, validation rationale, configured and effective modes, active -generation, and whether the previous policy is active. Static controls, -such as filesystem allowlists and process identity, require a new sandbox -because they are applied before the child process starts. - -Gateway-global policy can override sandbox-scoped policy. Use it sparingly -because it changes the effective access model for every sandbox on the gateway. - -## Policy Advisor - -The policy advisor pipeline turns observed denials into draft policy -recommendations. There are two proposers (sandbox-side mechanistic mapper, -agent-authored via `policy.local`); the gateway is the single referee. -For a new DNS name absent from policy, the supervisor can -publish a bounded, short-lived synthetic observation mapping without querying -an upstream resolver. It carries only the name to the subsequent TCP mediation -step, which supplies the destination port and verified process identity for a -denial summary. Observation mappings have no endpoint contracts or real pinned -addresses, cannot authorize a relay, and become stale on policy generation -change. Transparent TCP re-acquires the mapping at the generation that -authorized the connection, so a reload between DNS and authorization fails -closed instead of resolving the name again. Every unknown name still produces a -DNS denial event. Reserved names, fail-closed quarantine, and an exhausted -observation budget (a quarter of each address family's synthetic pool) keep the -plain DNS refusal. This mechanistic observation path does not depend on the -agent-authored proposal setting. -When enabled, L7 `policy_denied` responses include both structured -`next_steps` and a short `agent_guidance` string so generic agents can continue -through the proposal loop instead of treating the denial as terminal. - -1. **Submit.** Both proposers POST through the same `SubmitPolicyAnalysis` - path. Each chunk is persisted with its `analysis_mode` for audit provenance. - Agent-authored chunks cannot request `protocol: tcp` or `tls: skip`; the - sandbox-local API rejects those transport choices for immediate feedback, - and the gateway repeats the check before persistence. - Omitted-protocol endpoints remain available through the explicit proxy with - default TLS termination and HTTP authority checks. Administrators can still - author native TCP and raw TLS policy directly. -2. **Build and validate the candidate.** The gateway first canonicalizes a - mechanistic proposal against the live effective policy. If an endpoint is - already governed by an inspected or provider-owned contract, the candidate - preserves that contract and adds only the proposed sandbox binary. Provider - rules are immutable inputs; the sandbox contribution is stored as an - overlay. The gateway then performs the same merge, policy validation, - provider composition, credential preflight, and prover evaluation that the - candidate would encounter when applied. Each chunk stores the resulting - effective candidate, its hashes, any application error, and a review token - derived from the candidate and its non-secret live inputs. -3. **Auto-approval gate (proposer-agnostic, opt-in).** Auto-approval fires - only when *all three* conditions hold: (a) `proposal_approval_mode` - resolves to `"auto"` — gateway scope wins, sandbox scope is the - per-sandbox override, default is `"manual"`; (b) the prover delta is empty - (`prover: no new findings`); and (c) the security notes recomputed from - the chunk's current proposed rule are empty (see - [Security-notes gate](#security-notes-gate)). Before merging, the gateway - reloads the stored chunk and recomputes its candidate from live policy, - provider, and credential inputs. If the review token is unchanged, the - gateway reuses the persisted prover result. If it changed, the gateway - persists the refreshed candidate and requires a fresh review instead of - applying it. Decode, prover, merge, provider-composition, or credential - failures leave the chunk pending with an application error. The audit event uses `CONFIG:APPROVED` and carries - `auto=true`, `source=`, `prover_delta=empty`, and - `resolved_from=` as unmapped fields, with message text - `"auto-approved: no new prover findings"` — never `safe`. The opt-in gate - preserves OpenShell's default-deny posture: with no setting at either - scope, every proposal lands in `pending` for human review, even - when the prover sees no findings. -4. **Implicit supersede.** On any successful submission, the gateway scans - the sandbox's pending chunks for matches on `(host, port, binary)` and - auto-rejects the older ones with reason `"superseded by chunk X"`. This - gives the agent a refinement path (broad mechanistic L4 → narrow agent - L7) without an explicit `supersedes_chunk_id` field. -5. **Mechanistic dedup and self-reject.** Mechanistic submissions dedup on - `(host, port, binary)`: a repeat denial for an endpoint that already has a - draft row folds into that row (bumping `hit_count`) instead of creating a - new chunk. When the endpoint is already covered by a *different* approved - chunk, the redundant mechanistic submission self-rejects on arrival with - reason `"already covered by approved chunk X"`. This only ever acts on a - genuinely fresh, still-`pending` submission; it never rewrites the status - of an already-decided chunk. A dedup hit that resolves to an approved - row's own id leaves that chunk `approved` and merged — the self-reject - path does not un-merge a rule, so flipping an approved chunk to `rejected` - would leave the governance ledger disagreeing with the still-enforced - policy. -6. **Escalation.** Anything else lands in `pending` for human review. - -After any successful policy write, pending chunks already covered by the new -live effective policy are rejected as redundant. This keeps the review inbox -aligned with what the sandbox currently enforces. - -Endpoint advisor markers are provenance, not authorization or connection -metadata. Provider- or user-authored endpoints carry explicit provenance; -`policy.local` endpoints carry advisor provenance. A difference in endpoint -provenance alone is compatible during effective-policy ambiguity validation. -When identical endpoints merge, an explicit declaration dominates an advisor -declaration. Proposal coverage likewise ignores provenance so an approved -overlay converges when an existing explicit declaration already supplies the -same identity. - -This compatibility does not weaken SSRF classification. Exact-host trust -requires one matching rule to contain both an exact explicit endpoint and a -matching binary identity. An advisor endpoint cannot assemble that trust from -an unrelated explicit endpoint. When the advisor observes a new binary for an -existing explicit endpoint contract, canonicalization keeps the observation in -a separate rule whose endpoint retains advisor provenance. A provider or user -rule may independently establish trust for its own explicit endpoint and binary -pair, but an advisor overlay does not broaden that pair to a different binary. - -### Security-notes gate - -Separately from the prover, each chunk carries advisory `security_notes`. -Reads, bulk approval, and auto-approval regenerate them from the current -stored proposed rule instead of trusting a persisted value that may be stale -after an edit. Non-empty notes block auto-approval and make -`ApproveAllDraftChunks` skip the chunk unless `include_security_flagged` is -set. The chunk stays `pending`; an explicit human approval can still merge a -flagged chunk. - -Private/internal destinations are advisory, not blocking. A literal endpoint -IP, `allowed_ips` entry, or CIDR intersection in RFC 1918, CGNAT -`100.64.0.0/10`, IPv6 ULA `fc00::/7`, or another special-use range covered by -`openshell-core` `net::is_internal_net` produces a note. A hostless rule -carrying `allowed_ips` earns an extra note because it can match any hostname -resolving into the range. - -Always-blocked destinations are separate from this advisory classification. -Loopback, link-local, and unspecified IPs/CIDRs, plus `localhost` and known -metadata endpoint hostnames, are excluded from security notes. Submit and edit -may store such a draft, but existing merge validation rejects it when an -approval attempts to add it to policy; runtime SSRF protections remain the -final enforcement boundary. - -## Standalone boundary checks - -The standalone `openshell-prover check` command compares a fully composed local -candidate policy with an operator-supplied local boundary. It establishes -`Allowed(candidate) ⊆ Allowed(boundary)` for the policy domains reported in its -result. It does not fetch gateway state, compose provider rules, apply policy, -or decide whether an in-boundary change is eligible for automatic approval. - -The containment model covers filesystem paths, supported process identities, -Landlock compatibility requirements, L4 destinations including IP ranges, and -enforced REST method and path authority. Identity comparisons assume consistent -user and group resolution. Compatibility checks compare requested enforcement -requirements, not the actual kernel state of a running sandbox. -It returns explicit unsupported or inconclusive -results when a sound decision depends on authority or runtime context outside -the model. The result records the covered domains so callers can require the -authority relevant to their decision. The JSON `schema_version` versions the -result contract, while `prover_version` identifies the producing implementation. - -Before semantic validation, the checker observes cancellation and applies -aggregate limits across both inputs. Oversized checks therefore return -`resource_limit` without building validation indexes. Cross-protocol ambiguity -validation indexes host and port authority rather than comparing every endpoint -pair, and it checks cancellation while scanning admitted policies. - -The Rust containment API has an explicit extensibility contract: options and -modeled-domain evidence permit additive growth, while the four `CheckResult` -states remain exhaustive and authorization accepts only `Within`. This Rust -source-compatibility boundary is separate from the CLI JSON schema and the -reported modeled domains. See the `openshell-prover` crate README for the -supported construction and matching patterns. - -This containment operation is separate from the proposal-risk queries below. -See the [standalone policy prover documentation](../docs/how-it-works/policies/prover.mdx) -for installation, command behavior, model limitations, evidence, and exit codes. - -## What the proposal prover decides - -The prover answers four formal questions about each proposed policy -change. Each "yes" answer becomes its own categorical finding — there is -no severity grade. Any finding (of any category) blocks auto-approval. -The categories are intended to be (mostly) mutually exclusive per -underlying change: the gateway suppresses `capability_expansion` paths -whose `(binary, host, port)` is also in the `credential_reach_expansion` -delta, so a brand-new credentialed reach surfaces as one finding rather -than one reach + N method findings. - -| Category | The prover detects… | -|---|---| -| `link_local_reach` | The proposal grants reach to a host in `169.254.0.0/16`, `fe80::/10`, or a known metadata hostname such as `metadata.google.internal`. Unconditional — cloud-metadata endpoints serve credentials regardless of sandbox state. | -| `l7_bypass_credentialed` | The proposal lets a binary using a non-HTTP wire protocol (`git-remote-https`, `ssh`, `nc`) reach a host where a sandbox credential is in scope. The L7 proxy cannot inspect the wire protocol; the reviewer decides whether to trust the binary with the credential. | -| `credential_reach_expansion` | A binary gained credentialed reach to a (host, port) it could not reach before. New authenticated reach is a stated intent change; the reviewer confirms the binary should authenticate to the host at all. | -| `capability_expansion` | On a (binary, host, port) that already had credentialed reach, the policy adds a new HTTP method. The reviewer sees exactly which method was added (e.g., PUT) and decides if it's part of the agent's task. | - -"Credential in scope" is sandbox-coarse, not binary-fine: a credential is -considered in scope if the sandbox has a provider attached whose -`target_hosts` include the proposed endpoint's host, including runtime-like -first-label wildcard coverage such as `*.github.com` covering -`api.github.com`. v1 does not model credential scopes (read-only vs write); -presence is enough. - -Proposals intentionally omit `allowed_ips`. If a proposed rule targets a host -that resolves to a private IP, the proxy's runtime SSRF classification blocks -the connection. The operator must then add an explicit `allowed_ips` entry to -permit it — a two-step flow that keeps SSRF protection on by default. - -The advisor proposes narrow additions and preserves explicit-deny behavior. -Auto-approval is gated on prover determinism, not human judgment; an LLM-based -contextual reviewer is a deliberate future addition layered on top of the -deterministic prover gate. - -## Security Logging - -Sandbox events that represent observable behavior use OCSF structured logs: - -| Event | OCSF class | -|---|---| -| Network and proxy decisions | Network or HTTP activity | -| SSH authentication and relay activity | SSH activity | -| Process lifecycle | Process activity | -| Policy and settings changes | Configuration state change | -| Security findings | Detection finding | - -Use plain tracing for internal plumbing such as retries, debug state, and -intermediate steps where the final observable event is logged separately. -Forward-proxy success is a final observable event: emit it only after -middleware, token grants, credential rewriting, policy-generation checks, and -the HTTP relay have succeeded so a later denial cannot coexist with an allowed -record for the same request. - -Never log secrets, credentials, bearer tokens, or query parameters in OCSF -messages. OCSF JSONL output may be shipped to external systems. -The gateway-local OCSF JSONL file sink is restricted to the Windows/MXC path -and requires an explicit `OPENSHELL_OCSF_JSON=1` opt-in. Other gateway -deployments do not initialize this gateway file sink; a cross-platform gateway -sink requires its own storage and configuration integration. -MXC ETW process events record executable identity but omit command-line -arguments from structured fields and messages. Raw ETW debug summaries replace -the `commandLine` value with `[REDACTED]`, including pending-buffer eviction -diagnostics. -MXC ETW attribution never treats command text as ownership evidence. It uses the -driver-owned `wxc-exec` PID plus its kernel process start key as the initial -anchor. ETW attaches that generation key to each record, and the driver queries -the same key from its child process handle. PID attribution requires both values -to match, so reuse cannot transfer ownership between process generations. -Retired PID evidence is discarded; established identity, activity, and -correlation-vector links remain eligible during the five-second late-event -window. Records without matching generation evidence fail closed. diff --git a/architecture/windows-msvc-build.md b/architecture/windows-msvc-build.md deleted file mode 100644 index 198ecaade3..0000000000 --- a/architecture/windows-msvc-build.md +++ /dev/null @@ -1,211 +0,0 @@ -# Windows MSVC Build Design - -This page records the design decisions for the native Windows MSVC build lane. -It provides the native build lane and validates the in-process MXC compute -driver. It does not make Windows a Docker, Kubernetes, Podman, or VM runtime host. - -## Goals - -- Compile the OpenShell gateway and CLI for `x86_64-pc-windows-msvc` and `aarch64-pc-windows-msvc`. -- Keep the Linux and macOS build paths unchanged. -- Preserve gateway configuration parsing for all existing compute driver names. -- Build and test the in-process MXC driver on supported Windows hosts. -- Use the ordinary in-process compute-driver composition path; MXC receives the - canonical sandbox policy through `DriverSandboxSpec` and advertises that it - reports runtime readiness. -- Return clear unsupported errors when a Windows gateway is configured to use Docker, Kubernetes, Podman, or VM. -- Keep dedicated `windows:*` validation tasks while allowing the repository-wide - `pre-commit` task to delegate compiler-bearing Rust checks to the native - Windows MSVC environment. - -## Non-Goals - -- Do not support Docker Desktop, WSL, Hyper-V, Podman machine, Podman Desktop, Kubernetes, or VM-backed sandbox execution on Windows. -- Do not ship Windows standalone binaries for Docker, Kubernetes, Podman, or VM drivers. -- Do not implement named-pipe driver IPC, Windows services, MSI packaging, Credential Manager integration, or DPAPI integration in this lane. - -## Unsupported Driver Strategy - -The gateway composition crate installs platform-specific registration stubs on -Windows. These registrations preserve config-file selection and reject -unsupported drivers with a clear error without depending on their runtime -crates. - -Each stub follows its corresponding `compute-driver-*` Cargo feature. -`compute-driver-mxc` independently links and registers MXC, so a gateway built -with only that feature has only the MXC registration. The default -`in-tree-compute-drivers` alias enables all five features and preserves the -existing MXC plus unsupported-driver registrations. The focused Windows -contract tasks cover default, protocol-only, MXC-only, Docker-stub-only, and -MXC plus Docker-stub compositions. - -The Windows lane does not build, release, package, or smoke-test standalone -driver binaries for Docker, Kubernetes, Podman, or VM. Those binaries are Linux -or macOS deliverables only. - -The Kubernetes Secrets and Vault packages are also excluded as top-level -Windows workspace targets because their standalone driver binaries use Unix -domain sockets. Their libraries remain in the gateway dependency graph, so the -gateway's credential-driver configuration and in-process behavior still compile -on Windows. - -The standalone sandbox and supervisor runtimes are Unix-only and are excluded -as top-level Windows workspace targets. The MXC driver links only the -cross-platform supervisor network library needed by its host egress proxy. - -| Driver | Windows build behavior | Runtime behavior | -|---|---|---| -| Docker | Driver crate excluded; gateway registration stub retained. | Gateway construction returns unsupported. | -| Kubernetes | Driver crate excluded; gateway registration stub retained. | Gateway construction returns unsupported. | -| Podman | Driver crate excluded; gateway registration stub retained. | Gateway construction returns unsupported. | -| VM | Driver crate excluded; gateway registration stub retained. | Gateway construction returns unsupported. | -| MXC | Driver links into the native gateway and runs in Windows validation. | `process_container` is default-deny; grant-only `isolation_session` requires explicit configuration. | - -This keeps Windows behavior explicit without carrying runtime dependencies or -creating misleading Windows driver artifacts. - -## Mise Lane - -The GitHub Actions workflow runs Clippy for the Windows-supported workspace and -e2e crates plus Rust tests for pull-request mirror branches labeled `test:windows`. -Merge queues do not run this workflow. On -pushes to `windows`, a cache-seed job runs the same lint and test commands before a -dependent job builds the release binaries. Manual dispatches on `windows` exercise the same -seed-then-build path. The binaries remain CI validation artifacts and are not -uploaded or published. - -Each job restores and saves a dedicated Rust cache containing the Cargo -registry and dependency build artifacts, including artifacts from failed runs. -The seed job and pull-request job use the same Cargo target and sccache -namespaces, but GitHub scopes caches by branch. Pull-request mirror branches -cannot restore the `windows` branch cache. The release build waits for the seed -job, then restores its newly warmed cache rather than compiling concurrently -from a cold cache. - -Windows validation is exposed through `tasks/windows.toml`: - -| Task | Purpose | -|---|---| -| `windows:check:x64` | Check the x64 MSVC gateway/CLI build graph. | -| `windows:check:arm64` | Check the ARM64 MSVC gateway/CLI build graph. | -| `windows:lint:x64` | Run Clippy over the Windows-supported workspace for x64 MSVC. | -| `windows:lint:arm64` | Run Clippy over the Windows-supported workspace for ARM64 MSVC. | -| `windows:build:x64` | Build release x64 `openshell-gateway.exe` and `openshell.exe`. | -| `windows:build:arm64` | Build release ARM64 `openshell-gateway.exe` and `openshell.exe`. | -| `windows:test:x64` | Run native x64 workspace tests with the nextest CI profile and server test support, while excluding unsupported Windows packages as top-level test targets. | -| `windows:test:arm64` | Run the same suite natively on ARM64. | -| `windows:test:unsupported:x64` | Run focused gateway-composition tests for unsupported driver contracts. | -| `windows:test:unsupported:arm64` | Run the same focused contracts natively on ARM64. | -| `windows:ci` | Run check, build, test, unsupported-contract tests, and artifact reporting. | - -The Windows tasks call `tasks/scripts/windows-msvc.ps1`. The wrapper discovers -Visual Studio's `VsDevCmd.bat` with `vswhere` or by enumerating installed -release directories, validates the requested compiler and ARM64 Spectre -libraries, adds rustup MSVC targets, preserves an inherited `RUSTC_WRAPPER` -when the command is available, and keeps build artifacts under the normal -Cargo target tree. If the wrapper command is unavailable, it warns and clears -the setting so local builds continue without compiler caching. -On Windows, the generic `rust:check`, `rust:lint`, and `test:rust` tasks call -the same wrapper with the host-native MSVC target. The wrapper preserves the -Unix Cargo commands on Linux and macOS, excludes unsupported Windows runtime -packages, and runs the server test-support suite separately. Windows Clippy -continues to deny all warnings except unused imports, dead code, and unused -async functions caused by cfg-gated Windows stubs. Repository-wide pre-commit -skips only Linux-specific installer, build-environment shell-helper, and -packaging-asset tests; its -cross-platform Python, Markdown, license, and documentation checks still run. -Tracked Cargo lockfiles are checked natively through PowerShell. Deterministic -gateway parity uses Git for Windows Bash with temporary, checkout-scoped Python -launchers. The TypeScript SDK uses Windows protobuf plugin paths and x64 Biome -under emulation on ARM64, while its test binding follows Node's architecture -and the locked Rolldown version. Go tests retain race coverage wherever the -toolchain supports it; POSIX permission-bit checks are not Windows ACL tests. -Test tasks require the Rust target architecture to match the Windows host, so -an ARM64 test result is native coverage rather than x64 emulation coverage. -By default it enables the `z3-sys` prebuilt-release feature and pins Z3 5.1.0. -On a clean target directory, `z3-sys` downloads the official static library for -the selected Windows architecture instead of compiling Z3 through -CMake/MSBuild. GitHub Actions supplies its read-only workflow token for the -release lookup, and the Cargo target cache preserves the extracted library for -subsequent runs. When -`Z3_LIBRARY_PATH_OVERRIDE` points at a directory containing `libz3.lib`, the -wrapper uses that system Z3 instead and requires `Z3_SYS_Z3_HEADER` to point at -the full path to `z3.h`. Local clean builds use the unauthenticated GitHub API -unless `READ_ONLY_GITHUB_TOKEN` is set. - -GitHub Actions layers the Cargo target cache with sccache's GitHub Actions -backend. The target cache lets Cargo skip intact dependency builds; sccache -recovers cacheable Rust compiler outputs when source changes invalidate part of -that target tree. CI enables client-side mode and normalizes the checkout root -for stable compiler cache keys. The target-cache action runs its metadata step -with `RUSTC_WRAPPER` cleared so cache maintenance does not depend on sccache. -Hosted jobs use an isolated `RUSTUP_HOME` containing the pinned toolchain so -unused toolchains in runner images cannot change the target-cache restore key. - -The lane uses `mise run --skip-tools windows:*` because Windows Rust comes from -rustup and linking comes from Visual Studio Build Tools. Mise orchestrates the -tasks; it does not own the Windows toolchain. - -ARM64 validation requires the Visual Studio ARM64 MSVC tools, ARM64 -Spectre-mitigated libraries, host-native Clang tools, CMake tools, and an -ARM64-capable Windows SDK. Clang provides `libclang.dll` for `bindgen` and -`clang-cl.exe` for ARM64 crypto dependencies. During x64-to-ARM64 check/build, -the wrapper discovers and adds the Visual Studio-bundled Ninja to `PATH` for -native dependencies. Z3 uses the official prebuilt ARM64 static library, so it -does not inherit compiler settings from those native dependencies. Artifact -hashing uses .NET SHA256 directly because module autoloading in the -mise-launched Windows PowerShell process is not guaranteed. - -The wrapper defaults Cargo compilation to four jobs. Set -`OPENSHELL_WINDOWS_BUILD_JOBS` to a positive integer to override that limit. -A host-local mutex serializes wrapper-owned Cargo commands so concurrent -pre-commit tasks do not multiply the compiler process count. -The wrapper does not set `CL` or `_CL_`: those variables are also consumed by -`clang-cl`, where MSVC's `/MP` option can be interpreted as an input file and -break ARM64 crypto dependency builds. - -## CI Shape - -The x64 GitHub Actions jobs run on `windows-2025`; native ARM64 jobs run on -`windows-11-arm`. Pull-request mirrors labeled `test:windows` execute the matching -architecture-specific tasks: - -```powershell -mise run --skip-tools windows:lint: -mise run --skip-tools windows:test: -``` - -Pushes to `windows` and manual dispatches on that branch first seed the shared caches with those -same lint and test commands. Both seed and build jobs use job-level -`continue-on-error: true`, so Windows job failures do not fail the release-branch -workflow. Opt-in PR jobs still report failures normally. After the seed job -finishes, a separate job executes: - -```powershell -mise run --skip-tools windows:build: -``` - -The server test-support suite includes the unsupported-driver contract test, so -CI does not run the focused test task a second time. The focused task remains -available for local diagnosis. - -The hosted workflow uses architecture-specific cache namespaces and does not -cache Cargo-installed binaries. - -The local aggregate `windows:ci` task can still cross-build ARM64 on an x64 -host. Hosted tests use architecture-matched runners, so ARM64 test results are -native rather than emulated coverage. - -## Validation Contract - -A successful Windows build report should include: - -- x64 and ARM64 `cargo check` status. -- x64 and ARM64 release build status for `openshell-gateway.exe` and `openshell.exe`. -- x64 test summary. -- Native ARM64 test summary when validation runs on an ARM64 host. -- Focused unsupported-driver contract test status. -- Artifact size and SHA256 for each Windows binary. - -Warnings from Linux-only dead code are acceptable in the native Windows lane when -they come from code paths intentionally disabled on Windows. diff --git a/crates/openshell-cli/src/commands/provider.rs b/crates/openshell-cli/src/commands/provider.rs index fca1d16f3f..05ec158e42 100644 --- a/crates/openshell-cli/src/commands/provider.rs +++ b/crates/openshell-cli/src/commands/provider.rs @@ -2566,7 +2566,8 @@ fn format_provider_profile_details(profile: &ProviderProfile) -> String { let access = match network_access_preset_to_str(endpoint.access) { Some("") if !endpoint.rules.is_empty() => "custom rules".to_string(), Some("") if is_mcp && allow_all_known_mcp_methods == Some(true) => { - "all known MCP methods (subject to tool and deny rules)".to_string() + "core MCP methods for the selected revision (subject to tool and deny rules)" + .to_string() } Some("") => "not specified".to_string(), Some(access) => access.to_string(), @@ -2601,9 +2602,26 @@ fn format_provider_profile_details(profile: &ProviderProfile) -> String { || "not specified".to_string(), |options| options.versions.join(", "), ); - let _ = writeln!(rendered, " Allow all known MCP methods: {methods}"); + let _ = writeln!( + rendered, + " Allow core MCP methods for the selected revision: {methods}" + ); + if allow_all_known_mcp_methods == Some(true) { + rendered + .push_str(" Tool restrictions still apply; deny rules take precedence.\n"); + rendered.push_str( + " Without tool-specific allow rules, all tool names are allowed.\n", + ); + } + rendered.push_str(" Extension methods: require an exact allow rule\n"); let _ = writeln!(rendered, " Strict MCP tool names: {strict_names}"); let _ = writeln!(rendered, " MCP versions (declared): {versions}"); + rendered.push_str( + " Revision selection: MCP-Protocol-Version header; 2025-03-26 when absent\n", + ); + rendered.push_str( + " Legacy initialize requests negotiate their revision in the body.\n", + ); } if !endpoint.allowed_ips.is_empty() { let _ = writeln!( @@ -3143,11 +3161,39 @@ binaries: [/usr/bin/curl] let rendered = format_provider_profile_details(&proto); assert!( rendered - .contains("Access: all known MCP methods (subject to tool and deny rules)\n") + .contains("Access: core MCP methods for the selected revision (subject to tool and deny rules)\n") ); - assert!(rendered.contains("Allow all known MCP methods: true\n")); + assert!(rendered.contains("Allow core MCP methods for the selected revision: true\n")); + assert!( + rendered.contains("Tool restrictions still apply; deny rules take precedence.") + ); + assert!( + rendered.contains("Without tool-specific allow rules, all tool names are allowed.") + ); + assert!(rendered.contains("Extension methods: require an exact allow rule\n")); assert!(rendered.contains("Strict MCP tool names: true (default)\n")); assert!(rendered.contains("MCP versions (declared): 2025-11-25\n")); + assert!(rendered.contains( + "Revision selection: MCP-Protocol-Version header; 2025-03-26 when absent\n" + )); + assert!( + rendered + .contains("Legacy initialize requests negotiate their revision in the body.\n") + ); + + for output in ["json", "yaml"] { + let structured = format_provider_profile_description(&proto, output) + .expect("structured description renders"); + let roundtrip = if output == "json" { + parse_profile_json(&structured) + } else { + parse_profile_yaml(&structured) + } + .expect("structured description is a full profile document"); + assert_eq!(roundtrip, ProviderTypeProfile::from_proto(&proto)); + assert!(structured.contains("allow_all_known_mcp_methods")); + assert!(!structured.contains("Allow core MCP methods")); + } // Explicit rules remain relevant when the method default is enabled, // and independent tool-name validation must not disappear from view. @@ -3163,7 +3209,13 @@ binaries: [/usr/bin/curl] options.strict_tool_names = Some(false); let rendered = format_provider_profile_details(&proto); assert!(rendered.contains("Access: custom rules\n Rules: 1 allow, 0 deny\n")); - assert!(rendered.contains(&format!("Allow all known MCP methods: {allow_all}\n"))); + assert!(rendered.contains(&format!( + "Allow core MCP methods for the selected revision: {allow_all}\n" + ))); + assert_eq!( + rendered.contains("Tool restrictions still apply"), + allow_all + ); assert!(rendered.contains("Strict MCP tool names: false\n")); } } diff --git a/crates/openshell-cli/src/main.rs b/crates/openshell-cli/src/main.rs index e5f105919c..ef40f79936 100644 --- a/crates/openshell-cli/src/main.rs +++ b/crates/openshell-cli/src/main.rs @@ -19,7 +19,7 @@ use openshell_bootstrap::{ use openshell_cli::completers; use openshell_cli::run; use openshell_cli::tls::TlsOptions; -use openshell_core::proto::GpuResourceRequirements; +use openshell_core::proto::{GpuResourceRequirements, ServiceAuthorizationMode}; /// Resolved gateway context: name + gateway endpoint. struct GatewayContext { @@ -770,6 +770,22 @@ enum OutputFormat { Json, } +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq, ValueEnum)] +enum CliServiceAuthorizationMode { + #[default] + Strip, + BearerPassthrough, +} + +impl From for ServiceAuthorizationMode { + fn from(value: CliServiceAuthorizationMode) -> Self { + match value { + CliServiceAuthorizationMode::Strip => Self::Strip, + CliServiceAuthorizationMode::BearerPassthrough => Self::BearerPassthrough, + } + } +} + #[derive(Clone, Debug, ValueEnum)] enum CliProviderRefreshStrategy { Oauth2RefreshToken, @@ -1520,6 +1536,10 @@ enum SandboxCommands { )] expose: Option, + /// Handling for an incoming application Authorization header. + #[arg(long, value_enum, default_value_t, requires = "expose")] + expose_authorization_mode: CliServiceAuthorizationMode, + /// Allocate a pseudo-terminal for the remote command. /// Defaults to auto-detection (on when stdin and stdout are terminals). /// Use --tty to force a PTY even when auto-detection fails, or @@ -1535,6 +1555,14 @@ enum SandboxCommands { #[arg(long, conflicts_with = "editor")] detach: bool, + /// Restart behavior after the canonical main process exits. + #[arg( + long, + value_parser = ["never", "on-failure", "always"], + default_value = "never" + )] + restart_policy: String, + /// Auto-create missing providers from local credentials. /// /// Without this flag, an interactive prompt asks per-provider; @@ -1666,13 +1694,18 @@ enum SandboxCommands { /// For interactive shell sessions, use `sandbox connect` instead. /// /// Examples: + /// openshell sandbox exec my-sandbox -- ls -la /workspace /// openshell sandbox exec --name my-sandbox -- ls -la /workspace /// openshell sandbox exec -n my-sandbox --workdir /app -- python script.py /// echo "hello" | openshell sandbox exec -n my-sandbox -- cat #[command(help_template = LEAF_HELP_TEMPLATE, next_help_heading = "FLAGS")] Exec { /// Sandbox name (defaults to last-used sandbox). - #[arg(long, short = 'n', add = ArgValueCompleter::new(completers::complete_sandbox_names))] + #[arg(add = ArgValueCompleter::new(completers::complete_sandbox_names))] + sandbox: Option, + + /// Sandbox name; same as the positional argument. + #[arg(long, short = 'n', conflicts_with = "sandbox", add = ArgValueCompleter::new(completers::complete_sandbox_names))] name: Option, /// Working directory inside the sandbox. @@ -1709,15 +1742,15 @@ enum SandboxCommands { #[arg(long = "env", value_name = "KEY=VALUE")] envs: Vec, - /// Command and arguments to execute. - #[arg(required = true, trailing_var_arg = true, allow_hyphen_values = true)] + /// Command and arguments to execute, after `--`. + #[arg(required = true, last = true)] command: Vec, }, /// Connect to a sandbox. /// /// When no name is given, reconnects to the last-used sandbox. - /// Press Ctrl-P Ctrl-Q to disconnect without terminating the main process. + /// Press Ctrl-D or Ctrl-P Ctrl-Q to disconnect without terminating the main process. #[command(help_template = LEAF_HELP_TEMPLATE, next_help_heading = "FLAGS")] Connect { /// Sandbox name (defaults to last-used sandbox). @@ -2356,6 +2389,10 @@ enum ServiceCommands { /// Service name. service: Option, + + /// Handling for an incoming application Authorization header. + #[arg(long, value_enum, default_value_t)] + authorization_mode: CliServiceAuthorizationMode, }, /// List exposed sandbox service endpoints. @@ -2901,6 +2938,7 @@ async fn run_async() -> Result<()> { sandbox, service, target_port, + authorization_mode, } => { let service = service.unwrap_or_default(); run::service_expose( @@ -2908,6 +2946,7 @@ async fn run_async() -> Result<()> { &sandbox, &service, target_port, + authorization_mode.into(), &cli.workspace, &tls, ) @@ -3311,9 +3350,11 @@ async fn run_async() -> Result<()> { policy, forward, expose, + expose_authorization_mode, tty, no_tty, detach, + restart_policy, auto_providers, no_auto_providers, labels, @@ -3407,6 +3448,7 @@ async fn run_async() -> Result<()> { policy: policy.as_deref(), forward, expose, + expose_authorization_mode: expose_authorization_mode.into(), command: &command, tty_override, auto_providers_override, @@ -3416,6 +3458,7 @@ async fn run_async() -> Result<()> { output: output.as_str(), detach, suppress_credential_warnings: no_credential_warnings, + restart_policy: &restart_policy, }, &cli.workspace, &tls, @@ -3559,6 +3602,7 @@ async fn run_async() -> Result<()> { let _ = save_last_sandbox(&ctx.name, &cli.workspace, &name); } SandboxCommands::Exec { + sandbox, name, workdir, timeout, @@ -3568,7 +3612,8 @@ async fn run_async() -> Result<()> { command, no_login_shell, } => { - let name = resolve_sandbox_name(name, &ctx.name, &cli.workspace)?; + let name = + resolve_sandbox_name(name.or(sandbox), &ctx.name, &cli.workspace)?; // Resolve --tty / --no-tty into an Option override. let tty_override = if no_tty { Some(false) @@ -4398,6 +4443,89 @@ mod tests { assert_eq!(provider, "work-github"); } + #[test] + fn exec_grammar_requires_separator_before_remote_command() { + use clap::error::ErrorKind; + + // Returns (target, command, tty) or the clap error kind. + let parse = |args: &[&str]| { + let mut argv = vec!["openshell", "sandbox", "exec"]; + argv.extend(args); + let cli = Cli::try_parse_from(argv).map_err(|e| e.kind())?; + let Some(Commands::Sandbox { + command: + Some(SandboxCommands::Exec { + sandbox, + name, + command, + tty, + .. + }), + }) = cli.command + else { + panic!("expected sandbox exec command"); + }; + Ok::<_, ErrorKind>((name.or(sandbox), command, tty)) + }; + let check = |args: &[&str], target: Option<&str>, command: &[&str], tty: bool| { + let got = parse(args).unwrap_or_else(|kind| panic!("{args:?} failed: {kind:?}")); + let command = command.iter().map(ToString::to_string).collect(); + assert_eq!(got, (target.map(str::to_string), command, tty), "{args:?}"); + }; + + check( + &["a", "--", "echo", "hi"], + Some("a"), + &["echo", "hi"], + false, + ); + check(&["-n", "a", "--", "echo"], Some("a"), &["echo"], false); + check(&["--name", "a", "--", "echo"], Some("a"), &["echo"], false); + check(&["--", "echo", "hi"], None, &["echo", "hi"], false); + // Flags on either side of the positional target. + check(&["--tty", "a", "--", "echo"], Some("a"), &["echo"], true); + check(&["a", "--tty", "--", "echo"], Some("a"), &["echo"], true); + check( + &["-n", "a", "--tty", "--", "echo"], + Some("a"), + &["echo"], + true, + ); + // Hyphenated remote args are opaque. + check(&["a", "--", "ls", "-la"], Some("a"), &["ls", "-la"], false); + check(&["a", "--", "--tty"], Some("a"), &["--tty"], false); + check(&["--", "-n", "x"], None, &["-n", "x"], false); + // An inner delimiter belongs to the remote command. + let git = ["git", "log", "--", "path"]; + check( + &["a", "--", "git", "log", "--", "path"], + Some("a"), + &git, + false, + ); + check(&["--", "git", "log", "--", "path"], None, &git, false); + + let err: &[(&[&str], ErrorKind)] = &[ + // Target given twice. + (&["-n", "a", "b", "--", "echo"], ErrorKind::ArgumentConflict), + (&["b", "-n", "a", "--", "echo"], ErrorKind::ArgumentConflict), + // Missing `--`. + (&["a", "echo", "hi"], ErrorKind::UnknownArgument), + (&["a", "--tty", "echo"], ErrorKind::UnknownArgument), + (&["-n", "a", "echo", "hi"], ErrorKind::UnknownArgument), + (&["git", "log"], ErrorKind::UnknownArgument), + (&["a"], ErrorKind::MissingRequiredArgument), + (&[], ErrorKind::MissingRequiredArgument), + // Missing remote command. + (&["a", "--"], ErrorKind::MissingRequiredArgument), + (&["-n", "a", "--"], ErrorKind::MissingRequiredArgument), + (&["--"], ErrorKind::MissingRequiredArgument), + ]; + for (args, kind) in err { + assert_eq!(parse(args).map(|_| ()), Err(*kind), "{args:?}"); + } + } + #[test] fn provider_readiness_commands_have_bounded_waits_and_structured_output() { for action in ["attach", "detach", "status"] { @@ -6008,6 +6136,51 @@ mod tests { } } + #[test] + fn sandbox_create_restart_policy_defaults_to_never() { + let cli = Cli::try_parse_from(["openshell", "sandbox", "create"]).unwrap(); + match cli.command { + Some(Commands::Sandbox { + command: Some(SandboxCommands::Create { restart_policy, .. }), + .. + }) => assert_eq!(restart_policy, "never"), + other => panic!("expected SandboxCommands::Create, got: {other:?}"), + } + } + + #[test] + fn sandbox_create_restart_policy_accepts_on_failure() { + let cli = Cli::try_parse_from([ + "openshell", + "sandbox", + "create", + "--restart-policy", + "on-failure", + ]) + .unwrap(); + match cli.command { + Some(Commands::Sandbox { + command: Some(SandboxCommands::Create { restart_policy, .. }), + .. + }) => assert_eq!(restart_policy, "on-failure"), + other => panic!("expected SandboxCommands::Create, got: {other:?}"), + } + } + + #[test] + fn sandbox_create_restart_policy_rejects_unknown_value() { + assert!( + Cli::try_parse_from([ + "openshell", + "sandbox", + "create", + "--restart-policy", + "unless-stopped", + ]) + .is_err() + ); + } + #[test] fn sandbox_create_detach_parses_with_main_command() { let cli = Cli::try_parse_from([ @@ -6571,9 +6744,19 @@ mod tests { match cli.command { Some(Commands::Sandbox { - command: Some(SandboxCommands::Create { expose, detach, .. }), + command: + Some(SandboxCommands::Create { + expose, + expose_authorization_mode, + detach, + .. + }), }) => { assert_eq!(expose, Some(4500)); + assert_eq!( + expose_authorization_mode, + CliServiceAuthorizationMode::Strip + ); assert!(detach); } other => panic!("expected SandboxCommands::Create, got: {other:?}"), @@ -6601,6 +6784,44 @@ mod tests { assert!(result.is_err()); } + #[test] + fn sandbox_create_parses_bearer_passthrough_and_requires_expose() { + let cli = Cli::try_parse_from([ + "openshell", + "sandbox", + "create", + "--expose", + "4500", + "--expose-authorization-mode", + "bearer-passthrough", + ]) + .expect("create-time authorization mode should parse with --expose"); + match cli.command { + Some(Commands::Sandbox { + command: + Some(SandboxCommands::Create { + expose_authorization_mode, + .. + }), + }) => assert_eq!( + expose_authorization_mode, + CliServiceAuthorizationMode::BearerPassthrough + ), + other => panic!("expected SandboxCommands::Create, got: {other:?}"), + } + + assert!( + Cli::try_parse_from([ + "openshell", + "sandbox", + "create", + "--expose-authorization-mode", + "bearer-passthrough", + ]) + .is_err() + ); + } + #[test] fn service_expose_accepts_positional_target_port_and_service() { let cli = Cli::try_parse_from([ @@ -6620,11 +6841,13 @@ mod tests { sandbox, target_port, service, + authorization_mode, }), }) => { assert_eq!(sandbox, "my-sandbox"); assert_eq!(target_port, 8080); assert_eq!(service.as_deref(), Some("api")); + assert_eq!(authorization_mode, CliServiceAuthorizationMode::Strip); } other => panic!("expected service expose command, got: {other:?}"), } @@ -6642,16 +6865,45 @@ mod tests { sandbox, target_port, service, + authorization_mode, }), }) => { assert_eq!(sandbox, "my-sandbox"); assert_eq!(target_port, 8080); assert_eq!(service, None); + assert_eq!(authorization_mode, CliServiceAuthorizationMode::Strip); } other => panic!("expected service expose command, got: {other:?}"), } } + #[test] + fn service_expose_parses_bearer_passthrough() { + let cli = Cli::try_parse_from([ + "openshell", + "service", + "expose", + "my-sandbox", + "4500", + "--authorization-mode", + "bearer-passthrough", + ]) + .expect("service authorization mode should parse"); + + match cli.command { + Some(Commands::Service { + command: + Some(ServiceCommands::Expose { + authorization_mode, .. + }), + }) => assert_eq!( + authorization_mode, + CliServiceAuthorizationMode::BearerPassthrough + ), + other => panic!("expected service expose command, got: {other:?}"), + } + } + #[test] fn service_alias_parses_service_commands() { let cli = Cli::try_parse_from(["openshell", "svc", "expose", "my-sandbox", "8080"]) @@ -6664,6 +6916,7 @@ mod tests { sandbox, target_port, service, + .. }), }) => { assert_eq!(sandbox, "my-sandbox"); diff --git a/crates/openshell-cli/src/run.rs b/crates/openshell-cli/src/run.rs index 15d91b5891..02c806c421 100644 --- a/crates/openshell-cli/src/run.rs +++ b/crates/openshell-cli/src/run.rs @@ -56,11 +56,11 @@ use openshell_core::proto::{ ListSandboxPoliciesRequest, ListSandboxTemplatesRequest, ListSandboxesRequest, ListServicesRequest, PolicySource, PolicyStatus, RejectDraftChunkRequest, ResourceRequirements, RevokeSshSessionRequest, Sandbox, SandboxCondition, SandboxPhase, SandboxPolicy, - SandboxResources, SandboxServiceExposure, SandboxServiceLevel, SandboxSpec, SandboxStartup, - SandboxTemplate, SandboxWorkloadConfig, SandboxWorkloadTemplate, SandboxWorkloadTemplateSpec, - ServiceEndpointResponse, SettingScope, StartSandboxRequest, StopSandboxRequest, - TcpForwardFrame, TcpForwardInit, TcpRelayTarget, UpdateConfigRequest, WatchSandboxRequest, - exec_sandbox_event, tcp_forward_init, + SandboxResources, SandboxRestartPolicy, SandboxServiceExposure, SandboxServiceLevel, + SandboxSpec, SandboxStartup, SandboxTemplate, SandboxWorkloadConfig, SandboxWorkloadTemplate, + SandboxWorkloadTemplateSpec, ServiceAuthorizationMode, ServiceEndpointResponse, SettingScope, + StartSandboxRequest, StopSandboxRequest, TcpForwardFrame, TcpForwardInit, TcpRelayTarget, + UpdateConfigRequest, WatchSandboxRequest, exec_sandbox_event, tcp_forward_init, }; use openshell_core::settings; use openshell_core::{ObjectId, ObjectName, ObjectWorkspace}; @@ -443,6 +443,7 @@ pub struct SandboxCreateConfig<'a> { pub policy: Option<&'a str>, pub forward: Option, pub expose: Option, + pub expose_authorization_mode: ServiceAuthorizationMode, pub command: &'a [String], pub tty_override: Option, pub auto_providers_override: Option, @@ -452,6 +453,7 @@ pub struct SandboxCreateConfig<'a> { pub output: &'a str, pub detach: bool, pub suppress_credential_warnings: bool, + pub restart_policy: &'a str, } impl Default for SandboxCreateConfig<'_> { @@ -471,6 +473,7 @@ impl Default for SandboxCreateConfig<'_> { policy: None, forward: None, expose: None, + expose_authorization_mode: ServiceAuthorizationMode::Strip, command: &[], tty_override: None, auto_providers_override: None, @@ -480,6 +483,7 @@ impl Default for SandboxCreateConfig<'_> { output: "table", detach: false, suppress_credential_warnings: false, + restart_policy: "never", } } } @@ -507,6 +511,7 @@ pub async fn sandbox_create( policy, forward, expose, + expose_authorization_mode, command, tty_override, auto_providers_override, @@ -516,6 +521,7 @@ pub async fn sandbox_create( output, detach, suppress_credential_warnings, + restart_policy, } = config; if editor.is_some() && !command.is_empty() { @@ -670,6 +676,12 @@ pub async fn sandbox_create( template: inline_template, command: main_command, tty: main_terminal, + restart_policy: match restart_policy { + "never" => SandboxRestartPolicy::Never as i32, + "on-failure" => SandboxRestartPolicy::OnFailure as i32, + "always" => SandboxRestartPolicy::Always as i32, + value => return Err(miette::miette!("invalid restart policy '{value}'")), + }, ..SandboxSpec::default() }), name: name.unwrap_or_default().to_string(), @@ -684,6 +696,7 @@ pub async fn sandbox_create( .map(|target_port| SandboxServiceExposure { service: String::new(), target_port: u32::from(target_port), + authorization_mode: expose_authorization_mode as i32, }) .into_iter() .collect(), @@ -1697,6 +1710,41 @@ where "Resource version:".dimmed(), sandbox.metadata.as_ref().map_or(0, |m| m.resource_version) ); + println!( + " {} {}", + "Restart policy:".dimmed(), + sandbox + .spec + .as_ref() + .map_or("never", |spec| { restart_policy_name(spec.restart_policy) }) + ); + if let Some(status) = sandbox.status.as_ref() { + println!( + " {} {}", + "Main process instance:".dimmed(), + if status.main_process_instance_id.is_empty() { + "-" + } else { + &status.main_process_instance_id + } + ); + println!( + " {} {}", + "Last exit code:".dimmed(), + status + .exit_code + .map_or_else(|| "-".to_string(), |code| code.to_string()) + ); + println!(" {} {}", "Restart count:".dimmed(), status.restart_count); + println!( + " {} {}", + "Next restart:".dimmed(), + status.next_restart_time.as_ref().map_or_else( + || "-".to_string(), + |time| format_epoch_ms(proto_timestamp_ms(Some(time))), + ) + ); + } // Display labels if present if let Some(metadata) = &sandbox.metadata @@ -2846,6 +2894,14 @@ fn endpoint_status_display_lines(endpoint: &EndpointStatus) -> Vec { ] } +fn restart_policy_name(policy: i32) -> &'static str { + match SandboxRestartPolicy::try_from(policy) { + Ok(SandboxRestartPolicy::OnFailure) => "on-failure", + Ok(SandboxRestartPolicy::Always) => "always", + Ok(SandboxRestartPolicy::Unspecified | SandboxRestartPolicy::Never) | Err(_) => "never", + } +} + fn sandbox_detail_to_json( sandbox: &Sandbox, config: &GetSandboxConfigResponse, @@ -2855,6 +2911,33 @@ fn sandbox_detail_to_json( .as_object_mut() .expect("sandbox_to_json returns object"); + let restart_policy = sandbox + .spec + .as_ref() + .map_or("never", |spec| restart_policy_name(spec.restart_policy)); + obj.insert("restart_policy".into(), serde_json::json!(restart_policy)); + if let Some(status) = sandbox.status.as_ref() { + obj.insert( + "main_process_instance_id".into(), + serde_json::json!(status.main_process_instance_id), + ); + obj.insert("exit_code".into(), serde_json::json!(status.exit_code)); + obj.insert( + "restart_count".into(), + serde_json::json!(status.restart_count), + ); + obj.insert( + "next_restart_at_ms".into(), + serde_json::json!(proto_timestamp_ms(status.next_restart_time.as_ref())), + ); + obj.insert( + "main_process_started_at_ms".into(), + serde_json::json!(proto_timestamp_ms( + status.main_process_started_time.as_ref() + )), + ); + } + let policy_source = if config.policy_source == PolicySource::Global as i32 { "global" } else { @@ -3744,11 +3827,20 @@ pub async fn service_expose( sandbox: &str, service: &str, target_port: u16, + authorization_mode: ServiceAuthorizationMode, workspace: &str, tls: &TlsOptions, ) -> Result<()> { - let response = - expose_service_endpoint(server, sandbox, service, target_port, workspace, tls).await?; + let response = expose_service_endpoint( + server, + sandbox, + service, + target_port, + authorization_mode, + workspace, + tls, + ) + .await?; if service.is_empty() { println!( @@ -3778,6 +3870,7 @@ async fn expose_service_endpoint( sandbox: &str, service: &str, target_port: u16, + authorization_mode: ServiceAuthorizationMode, workspace: &str, tls: &TlsOptions, ) -> Result { @@ -3789,6 +3882,7 @@ async fn expose_service_endpoint( name: service.to_string(), target_port: u32::from(target_port), domain: true, + authorization_mode: authorization_mode as i32, workspace_scope: Some(openshell_core::proto::workspace_selector( workspace.to_string(), )), @@ -3961,6 +4055,7 @@ fn print_service_endpoint_table( .map_or("", |m| m.workspace.as_str()); let service = service_display_name(&endpoint.name).to_string(); let target = format!("127.0.0.1:{}", endpoint.target_port); + let authorization = service_authorization_mode_name(endpoint.authorization_mode); let url = if response.url.is_empty() { String::new() } else { @@ -3971,6 +4066,7 @@ fn print_service_endpoint_table( endpoint.sandbox.clone(), service, target, + authorization, url, )) }) @@ -3982,7 +4078,7 @@ fn print_service_endpoint_table( let ws_width = if all_workspaces { rows.iter() - .map(|(ws, _, _, _, _)| ws.len()) + .map(|(ws, _, _, _, _, _)| ws.len()) .max() .unwrap_or(9) .max(9) @@ -3991,50 +4087,52 @@ fn print_service_endpoint_table( }; let sandbox_width = rows .iter() - .map(|(_, sandbox, _, _, _)| sandbox.len()) + .map(|(_, sandbox, _, _, _, _)| sandbox.len()) .max() .unwrap_or(7) .max(7); let service_width = rows .iter() - .map(|(_, _, service, _, _)| service.len()) + .map(|(_, _, service, _, _, _)| service.len()) .max() .unwrap_or(7) .max(7); let target_width = rows .iter() - .map(|(_, _, _, target, _)| target.len()) + .map(|(_, _, _, target, _, _)| target.len()) .max() .unwrap_or(6) .max(6); if all_workspaces { println!( - "{: &str { if service.is_empty() { "-" } else { service } } +fn service_authorization_mode_name(mode: i32) -> &'static str { + match ServiceAuthorizationMode::try_from(mode).unwrap_or(ServiceAuthorizationMode::Strip) { + ServiceAuthorizationMode::BearerPassthrough => "bearer_passthrough", + ServiceAuthorizationMode::Unspecified | ServiceAuthorizationMode::Strip => "strip", + } +} + /// Read gcloud Application Default Credentials from disk. /// /// Returns `(client_id, client_secret, refresh_token)`. @@ -5861,7 +5967,7 @@ fn print_policy_revision_table(revisions: &[openshell_core::proto::SandboxPolicy &rev.policy_hash }; let error_short = if rev.load_error.len() > 40 { - format!("{}...", &rev.load_error[..40]) + truncate_status_field(&rev.load_error, 40) } else { rev.load_error.clone() }; @@ -6440,10 +6546,11 @@ mod tests { use openshell_core::proto::{ EndpointResult, EndpointStatus, GetSandboxConfigResponse, GpuResourceRequirements, PolicySource, PolicyStatus, ResourceRequirements, Sandbox, SandboxCondition, SandboxPhase, - SandboxPolicy, SandboxPolicyRevision, SandboxResources, SandboxStatus, - SandboxWorkloadConfig, SandboxWorkloadTemplate, SandboxWorkloadTemplateProvenance, - SandboxWorkloadTemplateSpec, ServiceEndpoint, ServiceEndpointResponse, WorkspaceMember, - WorkspaceRole, datamodel::v1::ObjectMeta, + SandboxPolicy, SandboxPolicyRevision, SandboxResources, SandboxRestartPolicy, SandboxSpec, + SandboxStatus, SandboxWorkloadConfig, SandboxWorkloadTemplate, + SandboxWorkloadTemplateProvenance, SandboxWorkloadTemplateSpec, ServiceAuthorizationMode, + ServiceEndpoint, ServiceEndpointResponse, WorkspaceMember, WorkspaceRole, + datamodel::v1::ObjectMeta, }; #[test] @@ -6526,6 +6633,29 @@ mod tests { assert!(unknown[0].get("sandbox").is_none()); } + #[test] + fn policy_revision_table_handles_unicode_load_errors() { + // Stored diagnostics can contain Unicode paths. Byte 40 splits the + // character in the 40-scalar case; longer errors must also remain safe. + let revisions = [ + String::new(), + "a".repeat(40), + "a".repeat(41), + format!("{}é", "a".repeat(39)), + format!("{}éz", "a".repeat(39)), + ] + .into_iter() + .map(|load_error| SandboxPolicyRevision { + version: 1, + status: PolicyStatus::Failed as i32, + load_error, + ..Default::default() + }) + .collect::>(); + + super::print_policy_revision_table(&revisions); + } + #[test] fn service_endpoint_json_has_raw_fields_and_normalized_url() { let response = ServiceEndpointResponse { @@ -6537,6 +6667,7 @@ mod tests { sandbox: "api".to_string(), name: String::new(), target_port: 8080, + authorization_mode: ServiceAuthorizationMode::BearerPassthrough as i32, ..Default::default() }), url: "https://api.openshell.localhost:3000/".to_string(), @@ -6551,6 +6682,7 @@ mod tests { "sandbox": "api", "service": "", "target_port": 8080, + "authorization_mode": "bearer_passthrough", "url": "https://api.openshell.localhost:17670/", }) ); @@ -7717,6 +7849,22 @@ mod tests { }; sandbox.set_phase(SandboxPhase::Ready as i32); sandbox.set_current_policy_version(2); + sandbox.spec = Some(SandboxSpec { + restart_policy: SandboxRestartPolicy::OnFailure as i32, + ..Default::default() + }); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + main_process_instance_id: "main-2".to_string(), + exit_code: Some(9), + restart_count: 2, + next_restart_time: openshell_core::time::timestamp_from_millis(1_700_000_000_000).ok(), + main_process_started_time: openshell_core::time::timestamp_from_millis( + 1_699_999_000_000, + ) + .ok(), + ..Default::default() + }); let config = GetSandboxConfigResponse { policy_source: PolicySource::Global as i32, @@ -7728,7 +7876,13 @@ mod tests { assert_eq!(json["id"], "sb-123"); assert_eq!(json["name"], "test-sb"); - assert_eq!(json["phase"], "Ready"); + assert_eq!(json["phase"], "Starting"); + assert_eq!(json["restart_policy"], "on-failure"); + assert_eq!(json["main_process_instance_id"], "main-2"); + assert_eq!(json["exit_code"], 9); + assert_eq!(json["restart_count"], 2); + assert_eq!(json["next_restart_at_ms"], 1_700_000_000_000_i64); + assert_eq!(json["main_process_started_at_ms"], 1_699_999_000_000_i64); assert_eq!(json["policy_source"], "global"); assert_eq!(json["revision"], 3); assert!(json["policy"].is_null()); diff --git a/crates/openshell-cli/src/ssh.rs b/crates/openshell-cli/src/ssh.rs index 204c8f6366..8d818062ca 100644 --- a/crates/openshell-cli/src/ssh.rs +++ b/crates/openshell-cli/src/ssh.rs @@ -640,7 +640,7 @@ async fn sandbox_connect_supervised( recovery_deadline = Some(Instant::now() + CONNECT_RECOVERY_TIMEOUT); retry_delay = CONNECT_RETRY_INITIAL_DELAY; eprintln!( - "Connection to sandbox lost; reconnecting. To disconnect, press Ctrl-P then Ctrl-Q after reattachment; press Ctrl-C while retrying to cancel." + "Connection to sandbox lost; reconnecting. To disconnect, press Ctrl-D or Ctrl-P then Ctrl-Q after reattachment; press Ctrl-C while retrying to cancel." ); } @@ -761,20 +761,11 @@ pub async fn sandbox_forward( let session = ssh_session_config(server, name, tls, workspace, None).await?; - let mut command = TokioCommand::from(ssh_base_command(&session.proxy_command)); - command - .arg("-N") - .arg("-o") - .arg("ExitOnForwardFailure=yes") - .arg("-o") - .arg(format!( - "SetEnv=OPENSHELL_FORWARD_SANDBOX_ID={}", - session.sandbox_id - )) - .arg("-L") - .arg(spec.ssh_forward_arg()); - - command.arg("sandbox"); + let mut command = TokioCommand::from(ssh_forward_command( + &session.proxy_command, + &session.sandbox_id, + spec, + )); if background { command @@ -793,34 +784,43 @@ pub async fn sandbox_forward( if background { let mut child = command.spawn().into_diagnostic()?; - let pid = child.id().ok_or_else(|| { - miette::miette!("ssh process did not expose a PID for background tracking") - })?; + let Some(pid) = child.id() else { + terminate_owned_forward_child(&mut child).await; + return Err(miette::miette!( + "ssh process did not expose a PID for background tracking" + )); + }; if let Err(err) = wait_for_forward_start(&mut child, spec) .await .wrap_err("ssh process started but local forward listener was not reachable") { - terminate_owned_forward_child(&mut child); + terminate_owned_forward_child(&mut child).await; return Err(err); } - track_background_forward_or_cleanup( + let tracked = track_background_forward_or_cleanup( workspace, name, port, pid, &session.sandbox_id, &spec.bind_addr, - || terminate_owned_forward_child(&mut child), - )?; + || { + let _ = child.start_kill(); + }, + ); + if tracked.is_err() { + let _ = child.wait().await; + } + tracked?; return Ok(()); } let status = { let mut child = command.spawn().into_diagnostic()?; if let Err(err) = wait_for_forward_start(&mut child, spec).await { - let _ = child.kill().await; + terminate_owned_forward_child(&mut child).await; return Err(err); } eprintln!("{}", foreground_forward_started_message(name, spec)); @@ -834,6 +834,30 @@ pub async fn sandbox_forward( Ok(()) } +/// A forward must stay owned by the process we spawn so it can be reaped or +/// tracked by PID. User SSH config may otherwise move it into a persistent mux. +fn ssh_forward_command(proxy_command: &str, sandbox_id: &str, spec: &ForwardSpec) -> Command { + let mut command = ssh_base_command(proxy_command); + command + .arg("-o") + .arg("ControlMaster=no") + .arg("-o") + .arg("ControlPath=none") + .arg("-o") + .arg("ControlPersist=no") + .arg("-o") + .arg("ForkAfterAuthentication=no") + .arg("-N") + .arg("-o") + .arg("ExitOnForwardFailure=yes") + .arg("-o") + .arg(format!("SetEnv=OPENSHELL_FORWARD_SANDBOX_ID={sandbox_id}")) + .arg("-L") + .arg(spec.ssh_forward_arg()) + .arg("sandbox"); + command +} + /// Wait for the local listener, racing the probe against the `ssh` child /// exiting. An early exit (e.g. `ExitOnForwardFailure=yes` tearing down the /// session) means forwarding never came up, so it errors instead of waiting @@ -859,7 +883,16 @@ async fn wait_for_forward_start(child: &mut Child, spec: &ForwardSpec) -> Result )) } } + }?; + + if let Some(status) = child.try_wait().into_diagnostic()? { + return Err(miette::miette!( + "ssh exited with status {status} after local forward listener opened on {}:{}", + forward_probe_host(spec), + spec.port, + )); } + Ok(()) } /// Poll the local endpoint until a connect succeeds or `wait_for` elapses. The @@ -935,9 +968,10 @@ fn forward_probe_host(spec: &ForwardSpec) -> &str { } } -/// Best-effort cleanup for the SSH child this process spawned. -fn terminate_owned_forward_child(child: &mut Child) { - let _ = child.start_kill(); +/// Best-effort cleanup for the SSH child this process spawned. `kill` waits for +/// the child to exit, so failed startup cannot leave an unreaped process. +async fn terminate_owned_forward_child(child: &mut Child) { + let _ = child.kill().await; } /// Track a verified background forward, cleaning it up if PID-file persistence fails. @@ -2434,6 +2468,55 @@ mod tests { assert_eq!(forward_probe_host(&loopback), "127.0.0.1"); } + #[cfg(unix)] + #[test] + fn forward_ssh_command_overrides_user_multiplexing_and_forking() { + let _guard = TEST_ENV_LOCK + .lock() + .unwrap_or_else(std::sync::PoisonError::into_inner); + let dir = tempfile::tempdir().unwrap(); + let config = dir.path().join("ssh_config"); + fs::write( + &config, + format!( + "Host sandbox\n ControlMaster auto\n ControlPath {}/mux-%C\n ControlPersist yes\n ForkAfterAuthentication yes\n", + dir.path().display() + ), + ) + .unwrap(); + + let forward = ssh_forward_command( + "openshell ssh-proxy --sandbox demo", + "sandbox-id", + &ForwardSpec::new(18080), + ); + let output = Command::new("ssh") + .arg("-F") + .arg(&config) + .arg("-G") + .args(forward.get_args()) + .output() + .unwrap(); + assert!( + output.status.success(), + "ssh -G failed: {}", + String::from_utf8_lossy(&output.stderr) + ); + let effective = String::from_utf8(output.stdout).unwrap(); + assert!(effective.lines().any(|line| line == "controlmaster false")); + assert!(effective.lines().any(|line| line == "controlpersist no")); + assert!( + effective + .lines() + .any(|line| line == "forkafterauthentication no") + ); + assert!( + !effective + .lines() + .any(|line| { line.starts_with("controlpath ") && line != "controlpath none" }) + ); + } + #[tokio::test] async fn wait_for_forward_listener_accepts_ready_listener() { let listener = tokio::net::TcpListener::bind(("127.0.0.1", 0)) diff --git a/crates/openshell-cli/tests/sandbox_create_lifecycle_integration.rs b/crates/openshell-cli/tests/sandbox_create_lifecycle_integration.rs index 0b638d5694..964559a8ad 100644 --- a/crates/openshell-cli/tests/sandbox_create_lifecycle_integration.rs +++ b/crates/openshell-cli/tests/sandbox_create_lifecycle_integration.rs @@ -2784,6 +2784,8 @@ async fn sandbox_create_exposes_service_after_ready_and_keeps_sandbox() { name: Some("sandbox"), keep: false, expose: Some(4500), + expose_authorization_mode: + openshell_core::proto::ServiceAuthorizationMode::BearerPassthrough, detach: true, ..test_config() }, @@ -2799,9 +2801,38 @@ async fn sandbox_create_exposes_service_after_ready_and_keeps_sandbox() { assert_eq!(create_requests[0].service_exposures.len(), 1); assert_eq!(create_requests[0].service_exposures[0].service, ""); assert_eq!(create_requests[0].service_exposures[0].target_port, 4500); + assert_eq!( + create_requests[0].service_exposures[0].authorization_mode(), + openshell_core::proto::ServiceAuthorizationMode::BearerPassthrough + ); assert!(expose_service_requests(&server).await.is_empty()); } +#[tokio::test] +async fn service_expose_forwards_bearer_passthrough_mode() { + let server = run_server().await; + let tls = test_tls(&server); + + run::service_expose( + &server.endpoint, + "sandbox", + "codex", + 4500, + openshell_core::proto::ServiceAuthorizationMode::BearerPassthrough, + "default", + &tls, + ) + .await + .expect("service expose should succeed"); + + let requests = expose_service_requests(&server).await; + assert_eq!(requests.len(), 1); + assert_eq!( + requests[0].authorization_mode(), + openshell_core::proto::ServiceAuthorizationMode::BearerPassthrough + ); +} + #[tokio::test] async fn sandbox_forward_background_tracks_owned_child_when_pid_discovery_fails() { let server = run_server().await; diff --git a/crates/openshell-conformance/src/scenarios/sandbox_lifecycle.rs b/crates/openshell-conformance/src/scenarios/sandbox_lifecycle.rs index 61a26558ff..0b3890c458 100644 --- a/crates/openshell-conformance/src/scenarios/sandbox_lifecycle.rs +++ b/crates/openshell-conformance/src/scenarios/sandbox_lifecycle.rs @@ -3,7 +3,8 @@ //! Portable sandbox lifecycle conformance scenarios. -use std::time::Duration; +use std::collections::HashSet; +use std::time::{Duration, Instant}; use serde::Deserialize; @@ -16,10 +17,17 @@ const TRANSITION_INTERVAL: Duration = Duration::from_secs(2); #[derive(Debug, Deserialize)] struct SandboxState { + id: String, name: String, phase: String, } +#[derive(Debug, Deserialize)] +struct SandboxListPage { + sandboxes: Vec, + next_page_token: String, +} + /// Certify sandbox stop, start, and deletion lifecycle behavior. pub const SANDBOX_LIFECYCLE_SCENARIO: Scenario = Scenario { name: "sandbox-lifecycle", @@ -118,9 +126,10 @@ async fn stopped_can_be_deleted(runner: &mut OpenShellRunner) -> Result<(), Stri .await?; run_lifecycle_command(runner, "stop", &sandbox_name, "stopped-delete/stop").await?; - wait_for_phase(runner, &sandbox_name, "Stopped", "stopped-delete/stopped").await?; + let sandbox = + wait_for_phase(runner, &sandbox_name, "Stopped", "stopped-delete/stopped").await?; run_lifecycle_command(runner, "delete", &sandbox_name, "stopped-delete/delete").await?; - wait_for_absence(runner, &sandbox_name, "stopped-delete/deleted").await?; + wait_for_absence(runner, &sandbox.id, &sandbox_name, "stopped-delete/deleted").await?; runner.forget_sandbox(&sandbox_name); Ok(()) } @@ -151,7 +160,9 @@ async fn create_running_sandbox( .await .map_err(|error| error.to_string())?; create.require_success()?; - wait_for_phase(runner, sandbox_name, "Ready", &format!("{step}/ready")).await + wait_for_phase(runner, sandbox_name, "Ready", &format!("{step}/ready")) + .await + .map(|_| ()) } async fn run_lifecycle_command( @@ -199,7 +210,7 @@ async fn wait_for_phase( sandbox_name: &str, expected_phase: &str, step: &str, -) -> Result<(), String> { +) -> Result { let sandbox_name = sandbox_name.to_string(); let expected_phase = expected_phase.to_string(); let step = step.to_string(); @@ -229,7 +240,7 @@ async fn wait_for_phase( "sandbox get returned {:?}; expected '{sandbox_name}'", state.name )), - Ok(state) if state.phase == expected_phase => Poll::Ready(()), + Ok(state) if state.phase == expected_phase => Poll::Ready(state), Ok(state) => Poll::Pending(format!( "sandbox '{sandbox_name}' phase is {:?}; expected {expected_phase:?}", state.phase @@ -246,9 +257,11 @@ async fn wait_for_phase( async fn wait_for_absence( runner: &mut OpenShellRunner, + sandbox_id: &str, sandbox_name: &str, step: &str, ) -> Result<(), String> { + let sandbox_id = sandbox_id.to_string(); let sandbox_name = sandbox_name.to_string(); let step = step.to_string(); let poll_step = step.clone(); @@ -257,22 +270,79 @@ async fn wait_for_absence( &poll_step, TRANSITION_TIMEOUT, TRANSITION_INTERVAL, - async move |runner| { - let result = runner - .step(format!("{step}/get")) - .description(format!("sandbox '{sandbox_name}' is no longer retrievable")) - .with_timeout(COMMAND_TIMEOUT) - .run(&["sandbox", "get", &sandbox_name, "--output", "json"]) - .await; - match result { - Ok(result) if !result.success() => Poll::Ready(()), - Ok(_) => { - Poll::Pending(format!("sandbox '{sandbox_name}' is still retrievable")) - } - Err(error) => Poll::Pending(error.to_string()), - } + async move |runner| match sandbox_is_listed(runner, &sandbox_id, &sandbox_name, &step) + .await + { + Ok(false) => Poll::Ready(()), + Ok(true) => Poll::Pending(format!( + "sandbox '{sandbox_name}' with ID '{sandbox_id}' is still listed" + )), + Err(error) => Poll::Pending(error), }, ) .await .map_err(|error| error.to_string()) } + +async fn sandbox_is_listed( + runner: &OpenShellRunner, + sandbox_id: &str, + sandbox_name: &str, + step: &str, +) -> Result { + let deadline = Instant::now() + COMMAND_TIMEOUT; + let mut seen_page_tokens = HashSet::new(); + let mut page_token = String::new(); + let mut page = 0u32; + loop { + let remaining = deadline.saturating_duration_since(Instant::now()); + if remaining.is_zero() { + return Err(format!( + "sandbox list observation for '{sandbox_name}' exceeded its {COMMAND_TIMEOUT:?} deadline" + )); + } + + let result = runner + .step(format!("{step}/list/{page}")) + .description(format!( + "sandbox list confirms whether '{sandbox_name}' with ID '{sandbox_id}' still exists" + )) + .with_timeout(remaining) + .run(&[ + "sandbox", + "list", + "--page-size", + "1000", + "--page-token", + &page_token, + "--output", + "json", + ]) + .await + .map_err(|error| error.to_string())?; + result.require_success()?; + + let response = result + .json::() + .map_err(|error| error.to_string())?; + if response + .sandboxes + .iter() + .any(|sandbox| sandbox.id == sandbox_id) + { + return Ok(true); + } + if response.next_page_token.is_empty() { + return Ok(false); + } + if !seen_page_tokens.insert(response.next_page_token.clone()) { + return Err(format!( + "sandbox list returned a repeated page token while looking for '{sandbox_name}'" + )); + } + page_token = response.next_page_token; + page = page + .checked_add(1) + .ok_or_else(|| "sandbox list page counter overflowed".to_string())?; + } +} diff --git a/crates/openshell-core/src/grpc_client.rs b/crates/openshell-core/src/grpc_client.rs index f49466b0e2..9808bd16e0 100644 --- a/crates/openshell-core/src/grpc_client.rs +++ b/crates/openshell-core/src/grpc_client.rs @@ -1116,6 +1116,7 @@ fn provider_environment_result( .collect::>>()?; Ok(ProviderEnvironmentResult { environment: inner.environment, + files: inner.files, provider_env_revision: inner.provider_env_revision, provider_attachment_epoch: inner.provider_attachment_epoch, policy_hash: inner.policy_hash, @@ -1380,6 +1381,7 @@ mod settings_poll_tests { /// Credential material and the authority snapshot that produced its bindings. pub struct ProviderEnvironmentResult { pub environment: HashMap, + pub files: HashMap, pub provider_env_revision: u64, /// Attachment identity captured with the delivered credential records. pub provider_attachment_epoch: String, diff --git a/crates/openshell-core/src/mcp.rs b/crates/openshell-core/src/mcp.rs index 4300b713c6..e110f76e9d 100644 --- a/crates/openshell-core/src/mcp.rs +++ b/crates/openshell-core/src/mcp.rs @@ -1,15 +1,14 @@ // SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. // SPDX-License-Identifier: Apache-2.0 -//! `OpenShell`-owned MCP protocol revisions and immutable batch-shape metadata. - -use std::collections::BTreeSet; +//! `OpenShell`-owned MCP policy revisions, allowlist parsing, and resource limits. use crate::proto::{McpOptions, ProviderProfile}; pub use openshell_policy_schema::{ DEFAULT_MCP_PROTOCOL_VERSION, MAX_MCP_LEGACY_BATCH_MESSAGES, McpProtocolVersion, - McpWireProfile, ParseMcpProtocolVersionError, + ParseMcpProtocolVersionError, ParseMcpVersionsError, canonicalize_mcp_versions, + parse_mcp_versions, }; /// Return whether a policy protocol name denotes MCP. @@ -48,20 +47,11 @@ pub fn normalize_provider_profile_mcp_fields(profile: &mut ProviderProfile) { continue; } - // Parse into the shared version type before mutation. Comparing the - // set size with the input length detects duplicates without erasing - // the duplicate values that a fail-closed validator must report. - let Ok(versions) = options - .versions - .iter() - .map(|version| version.parse::()) - .collect::, _>>() - else { + // Validate the explicit list before mutation so malformed values retain + // their original order and spelling for the checked policy boundary. + let Ok(versions) = parse_mcp_versions(&options.versions) else { continue; }; - if versions.len() != options.versions.len() { - continue; - } options.versions = versions .into_iter() @@ -85,6 +75,7 @@ mod tests { McpProtocolVersion::V2025_03_26, McpProtocolVersion::V2025_06_18, McpProtocolVersion::V2025_11_25, + McpProtocolVersion::V2026_07_28, ] ); assert!( @@ -95,17 +86,47 @@ mod tests { } #[test] - fn default_mcp_protocol_version_is_pinned_to_the_2025_11_25_profile() { + fn default_mcp_protocol_version_is_pinned_to_the_2025_11_25_revision() { assert_eq!( DEFAULT_MCP_PROTOCOL_VERSION, McpProtocolVersion::V2025_11_25 ); assert_eq!(DEFAULT_MCP_PROTOCOL_VERSION.as_str(), "2025-11-25"); + } + + #[test] + fn parse_mcp_versions_returns_canonical_order_without_mutating_the_authored_list() { + let values = ["2026-07-28", "2025-03-26", "2025-11-25", "2025-06-18"].map(str::to_string); + let original = values.clone(); + + let versions = parse_mcp_versions(&values).expect("explicit supported revisions"); + + assert_eq!( + versions.into_iter().collect::>(), + McpProtocolVersion::ALL + ); + assert_eq!(values, original); + } + + #[test] + fn parse_mcp_versions_reports_the_first_error_without_repairing_input() { + let duplicate = ParseMcpVersionsError::Duplicate(McpProtocolVersion::V2025_03_26); + let unsupported = ParseMcpVersionsError::Unsupported( + " 2025-11-25" + .parse::() + .expect_err("revision is unsupported"), + ); + for (values, expected) in [ + (vec![], ParseMcpVersionsError::Empty), + (vec!["2025-03-26", "2025-03-26", " 2025-11-25"], duplicate), + (vec![" 2025-11-25", "2025-03-26", "2025-03-26"], unsupported), + ] { + let values = values.into_iter().map(str::to_string).collect::>(); + let original = values.clone(); - let profile = DEFAULT_MCP_PROTOCOL_VERSION.wire_profile(); - assert_eq!(profile.version(), McpProtocolVersion::V2025_11_25); - assert!(!profile.allows_json_rpc_batches()); - assert_eq!(profile.max_batch_messages(), None); + assert_eq!(parse_mcp_versions(&values), Err(expected)); + assert_eq!(values, original); + } } fn provider_profile_with_mcp(protocol: &str, options: Option) -> ProviderProfile { @@ -166,6 +187,7 @@ mod tests { Some(McpOptions { strict_tool_names: Some(true), versions: vec![ + "2026-07-28".to_string(), "2025-11-25".to_string(), "2025-03-26".to_string(), "2025-06-18".to_string(), @@ -197,6 +219,7 @@ mod tests { "2025-03-26".to_string(), "2025-06-18".to_string(), "2025-11-25".to_string(), + "2026-07-28".to_string(), ], ..McpOptions::default() } @@ -207,9 +230,11 @@ mod tests { fn provider_profile_mcp_normalization_preserves_every_malformed_explicit_list() { for versions in [ vec!["2025-11-25", "2025-11-25"], + vec!["2026-07-28", "2025-03-26", "2025-03-26", "2025-06-18"], + vec!["2026-07-28", "latest", "2025-03-26"], vec![" 2025-11-25"], vec!["2025-11-25 "], - vec!["2026-07-28"], + vec!["2026-07-29"], vec!["latest"], vec!["draft"], ] { @@ -249,6 +274,7 @@ mod tests { assert_eq!("2025-03-26".parse(), Ok(McpProtocolVersion::V2025_03_26)); assert_eq!("2025-06-18".parse(), Ok(McpProtocolVersion::V2025_06_18)); assert_eq!("2025-11-25".parse(), Ok(McpProtocolVersion::V2025_11_25)); + assert_eq!("2026-07-28".parse(), Ok(McpProtocolVersion::V2026_07_28)); for unsupported in [ "", @@ -256,13 +282,14 @@ mod tests { " 2025-06-18", "2025-11-25\n", "2025-11-24", - "2026-07-28", + "2026-07-28 ", + "2026-07-29", "draft", "latest", ] { let error = unsupported .parse::() - .expect_err("unsupported MCP revision must be rejected"); + .expect_err("unsupported revision must fail"); assert_eq!(error.value(), unsupported); } } @@ -289,29 +316,6 @@ mod tests { assert_eq!(error.value(), rejected); } - #[test] - fn mcp_protocol_version_wire_profiles_define_batch_metadata() { - assert_eq!(MAX_MCP_LEGACY_BATCH_MESSAGES, 64); - - let legacy = McpProtocolVersion::V2025_03_26.wire_profile(); - assert_eq!(legacy.version(), McpProtocolVersion::V2025_03_26); - assert!(legacy.allows_json_rpc_batches()); - assert_eq!( - legacy.max_batch_messages(), - Some(MAX_MCP_LEGACY_BATCH_MESSAGES) - ); - - for version in [ - McpProtocolVersion::V2025_06_18, - McpProtocolVersion::V2025_11_25, - ] { - let profile = version.wire_profile(); - assert_eq!(profile.version(), version); - assert!(!profile.allows_json_rpc_batches()); - assert_eq!(profile.max_batch_messages(), None); - } - } - #[test] fn mcp_options_versions_field_number_is_stable() { let descriptor_set = FileDescriptorSet::decode(crate::FILE_DESCRIPTOR_SET) diff --git a/crates/openshell-core/src/provider_credentials.rs b/crates/openshell-core/src/provider_credentials.rs index fee82e0d3f..6988d1f4be 100644 --- a/crates/openshell-core/src/provider_credentials.rs +++ b/crates/openshell-core/src/provider_credentials.rs @@ -36,11 +36,14 @@ pub struct ChildEnvironmentSnapshot { pub revision: u64, /// Prepared environment used by future workload processes. pub environment: HashMap, + /// Complete desired set of non-secret files for this installation. + pub files: HashMap, } #[derive(Debug)] struct ProviderCredentialStateInner { current: Arc, + files: HashMap, generations: VecDeque>, current_resolver: Option>, combined_resolver: Option>, @@ -117,6 +120,7 @@ impl ProviderCredentialState { Self { inner: Arc::new(RwLock::new(ProviderCredentialStateInner { current: snapshot.clone(), + files: HashMap::new(), generations, current_resolver, combined_resolver, @@ -172,6 +176,7 @@ impl ProviderCredentialState { Ok(Self { inner: Arc::new(RwLock::new(ProviderCredentialStateInner { current: snapshot.clone(), + files: HashMap::new(), generations, current_resolver, combined_resolver, @@ -205,6 +210,7 @@ impl ProviderCredentialState { Self { inner: Arc::new(RwLock::new(ProviderCredentialStateInner { current: snapshot.clone(), + files: HashMap::new(), generations: VecDeque::new(), current_resolver: None, combined_resolver: None, @@ -410,22 +416,19 @@ impl ProviderCredentialState { inner.current = Arc::new(env); } - /// Return `child_env` with GCP static config vars resolved to real values. + /// Return `child_env` with explicitly non-secret config vars resolved. /// - /// The credential pipeline placeholderizes ALL env values, but GCP SDKs - /// and coding agents read certain vars (project ID, region, metadata host) - /// at process startup before any HTTP request flows through the proxy. - /// This method overrides those vars with resolved real values while - /// keeping secret credentials (like `GCP_ACCESS_TOKEN`) as placeholders. + /// The credential pipeline placeholderizes all env values. Workloads need + /// non-secret configuration, including provider file paths, at process + /// startup before any HTTP request flows through the proxy. Credential + /// values remain placeholders. /// /// Three layers of env var injection: /// 1. **Synthetic vars** (`GCE_METADATA_IP`, `METADATA_SERVER_DETECTION`) /// — sandbox-internal config not from user /// input, inserted directly here with real values. - /// 2. **`google_cloud::STATIC_CONFIG_KEYS`** — user-provided non-secret config - /// (project ID, region, SA email) that was placeholderized by - /// `ProviderPlugin::inject_env` → `SecretResolver`; un-placeholderized - /// here so SDKs can read them at startup. + /// 2. **Classified non-secret keys** — user-provided config and provider + /// file paths un-placeholderized so workloads can read them at startup. /// 3. Everything else stays as placeholders for proxy-time resolution. pub fn child_env_with_gcp_resolved(&self) -> HashMap { let inner = self @@ -463,9 +466,18 @@ impl ProviderCredentialState { installation_id: inner.current.installation_id.clone(), revision, environment, + files: inner.files.clone(), }) } + /// Attach the complete file set to a prepared provider snapshot. + pub fn set_managed_files(&self, files: HashMap) { + self.inner + .write() + .expect("provider credential state poisoned") + .files = files; + } + fn resolve_child_env_snapshot( inner: &ProviderCredentialStateInner, ) -> (u64, HashMap) { @@ -477,11 +489,7 @@ impl ProviderCredentialState { && inner .non_secret_environment_keys .contains("GCE_METADATA_HOST"); - let has_gcp_config = google_cloud::STATIC_CONFIG_KEYS - .iter() - .any(|key| env.contains_key(*key) && inner.non_secret_environment_keys.contains(*key)); - - if !has_gcp_metadata && !has_gcp_config { + if !has_gcp_metadata && inner.non_secret_environment_keys.is_empty() { return (inner.current.revision, env); } @@ -506,16 +514,15 @@ impl ProviderCredentialState { ); } - // Un-placeholderize non-secret config vars so SDKs can read them - // at process startup before any HTTP flows through the proxy. + // Only explicitly classified non-secret values may be unwrapped. if let Some(ref resolver) = inner.combined_resolver { - for key in google_cloud::STATIC_CONFIG_KEYS - .iter() - .filter(|key| inner.non_secret_environment_keys.contains(**key)) - { + for key in &inner.non_secret_environment_keys { + if !env.contains_key(key) || (has_gcp_metadata && key == "GCE_METADATA_HOST") { + continue; + } let placeholder = crate::secrets::placeholder_for_env_key(key); if let Some(value) = resolver.resolve_placeholder(&placeholder) { - env.insert(key.to_string(), value.to_string()); + env.insert(key.clone(), value.to_string()); } } } @@ -738,7 +745,7 @@ impl ProviderCredentialState { pub fn install_prepared(&self, prepared: &Self) -> usize { // Release the candidate lock before taking the live lock, including // when a caller passes another handle to the same state. - let (snapshot, generations, current_resolver, bindings, non_secret_keys) = { + let (snapshot, generations, current_resolver, bindings, non_secret_keys, files) = { let candidate = prepared .inner .read() @@ -749,6 +756,7 @@ impl ProviderCredentialState { candidate.current_resolver.clone(), candidate.static_credential_bindings.clone(), candidate.non_secret_environment_keys.clone(), + candidate.files.clone(), ) }; let mut inner = self @@ -781,6 +789,7 @@ impl ProviderCredentialState { ); inner.static_credential_bindings = bindings; inner.non_secret_environment_keys = non_secret_keys; + inner.files = files; inner.body_inventory_available = true; inner .known_body_keys @@ -2138,6 +2147,33 @@ mod tests { assert!(!env.contains_key("CLAUDE_CODE_USE_VERTEX")); } + #[test] + fn child_env_resolves_provider_file_path_but_keeps_credentials_placeholderized() { + let path = "/run/openshell/providers/acme/client.toml"; + let state = ProviderCredentialState::from_bound_environment( + 1, + HashMap::from([ + ("ACME_CONFIG_FILE".to_string(), path.to_string()), + ("ACME_TOKEN".to_string(), "secret-token".to_string()), + ]), + HashMap::new(), + HashMap::new(), + HashMap::from([( + "ACME_TOKEN".to_string(), + binding("api.acme.example", 443, "/**"), + )]), + vec!["ACME_CONFIG_FILE".to_string()], + ) + .expect("classified provider environment"); + + let env = state.child_env_with_gcp_resolved(); + assert_eq!(env.get("ACME_CONFIG_FILE").map(String::as_str), Some(path)); + assert_eq!( + env.get("ACME_TOKEN").map(String::as_str), + Some("openshell:resolve:env:v1_ACME_TOKEN") + ); + } + #[test] fn child_env_with_gcp_resolved_overrides_gcp_static_vars() { let state = ProviderCredentialState::from_bound_environment( diff --git a/crates/openshell-driver-docker/README.md b/crates/openshell-driver-docker/README.md index a25c455eea..ee329f8dea 100644 --- a/crates/openshell-driver-docker/README.md +++ b/crates/openshell-driver-docker/README.md @@ -50,6 +50,32 @@ through the Docker archive API. No workload launch depends on a host bind mount or a tool supplied by the workload image, so the same path works with local, remote, and VM-backed Docker daemons. +## Corporate Proxy Egress + +`[openshell.drivers.docker]` accepts the operator-owned corporate proxy fields +`https_proxy`, `no_proxy`, `proxy_auth_file`, +`proxy_auth_allow_insecure`, `proxy_connect_by_hostname`, and +`proxy_ca_bundle`. Sandbox environment, image contents, and per-sandbox driver +configuration cannot select or override them. + +The driver validates the proxy URL and cross-field relationships at gateway +startup. It bounded-reads proxy credentials and the PEM CA bundle before use. +Missing, unreadable, empty, oversized, malformed, or certificate-free CA files +fail closed. A CA bundle requires `https_proxy`, although the proxy URL may be +`http://` when the proxy intercepts destination TLS. + +For each supervisor launch, Docker copies the credential and CA contents into +the existing supervisor-only named volume. It passes only the fixed container +paths `/.openshell/supervisor/upstream-proxy-auth` and +`/.openshell/supervisor/upstream-proxy-ca-bundle.pem` to the supervisor. It +does not bind-mount the gateway-host files, which preserves remote-daemon +support and keeps host paths out of workload container metadata. + +The supervisor trusts the configured CA for its TLS connection to an HTTPS +proxy and for destination certificates re-signed by a TLS-intercepting proxy. +It also includes the corporate root in the generated combined trust bundle +used by workload processes. + ## Identity and Workspace Before creating the workload, the driver pins the image ID and reads its diff --git a/crates/openshell-driver-docker/src/lib.rs b/crates/openshell-driver-docker/src/lib.rs index 1f4c91f5e4..d4a1f2f71d 100644 --- a/crates/openshell-driver-docker/src/lib.rs +++ b/crates/openshell-driver-docker/src/lib.rs @@ -107,6 +107,8 @@ const BOUNDARY_CERTIFICATE_MOUNT_PATH: &str = "/.openshell/channel/sandbox/serve const BOUNDARY_PRIVATE_KEY_MOUNT_PATH: &str = "/.openshell/channel/sandbox/server.key"; const SUPERVISOR_STATE_MOUNT_PATH: &str = "/.openshell/supervisor"; const SUPERVISOR_PROXY_AUTH_MOUNT_PATH: &str = "/.openshell/supervisor/upstream-proxy-auth"; +const SUPERVISOR_PROXY_CA_BUNDLE_MOUNT_PATH: &str = + "/.openshell/supervisor/upstream-proxy-ca-bundle.pem"; const PROVIDER_SPIFFE_WORKLOAD_API_SOCKET_MOUNT_DIR: &str = openshell_core::driver_utils::PROVIDER_SPIFFE_WORKLOAD_API_SOCKET_MOUNT_DIR; const DRIVER_ADMITTED_BACKEND: &str = openshell_sandbox_backend::BACKEND_NAME; @@ -218,6 +220,10 @@ pub struct DockerComputeConfig { #[serde(flatten)] pub upstream_proxy: UpstreamProxyConfig, + /// Gateway-host PEM CA bundle trusted for the corporate proxy and for + /// server certificates re-signed by a TLS-intercepting proxy. + pub proxy_ca_bundle: Option, + /// Host UNIX socket projected into the supervisor for provider identity. pub provider_spiffe_workload_api_socket: Option, @@ -242,6 +248,7 @@ impl DockerComputeConfig { validate_image_pull_policy(self.image_pull_policy)?; self.upstream_proxy.validate().map_err(Error::config)?; validate_docker_proxy_auth_file(&self.upstream_proxy)?; + validate_docker_proxy_ca_bundle(self)?; if let Some(socket) = self.provider_spiffe_workload_api_socket.as_deref() { openshell_core::driver_utils::validate_provider_spiffe_unix_socket(socket) .map_err(Error::config)?; @@ -276,6 +283,7 @@ impl Default for DockerComputeConfig { sandbox_pids_limit: openshell_core::config::default_sandbox_pids_limit(), enable_bind_mounts: false, upstream_proxy: UpstreamProxyConfig::default(), + proxy_ca_bundle: None, provider_spiffe_workload_api_socket: None, app_armor_profile: None, } @@ -307,6 +315,7 @@ struct DockerDriverRuntimeConfig { sandbox_pids_limit: Option, enable_bind_mounts: bool, upstream_proxy: UpstreamProxyConfig, + proxy_ca_bundle: Option, provider_spiffe_workload_api_socket: Option, app_armor_profile: Option, } @@ -837,10 +846,7 @@ impl DockerComputeDriver { gateway_log_level: &str, docker_config: &DockerComputeConfig, ) -> CoreResult { - docker_config - .resource_admission - .validate() - .map_err(Error::config)?; + docker_config.validate_configuration(gateway_bind_address)?; let socket_path = docker_config .socket_path .clone() @@ -873,15 +879,8 @@ impl DockerComputeDriver { cdi_supported, wsl_all_gpu_fallback_enabled, }; - validate_sandbox_pids_limit(docker_config.sandbox_pids_limit)?; - validate_image_pull_policy(docker_config.image_pull_policy)?; validate_docker_app_armor_profile(docker_config.app_armor_profile.as_ref(), &info)?; let gateway_port = gateway_bind_address.port(); - if gateway_port == 0 { - return Err(Error::config( - "docker compute driver requires a fixed non-zero gateway bind port", - )); - } let mut docker_config = docker_config.clone(); if docker_config.grpc_endpoint.trim().is_empty() { docker_config.grpc_endpoint = default_docker_supervisor_grpc_endpoint( @@ -947,6 +946,7 @@ impl DockerComputeDriver { allow_driver_config: docker_config.allow_driver_config, resource_admission: docker_config.resource_admission.clone(), upstream_proxy: docker_config.upstream_proxy.clone(), + proxy_ca_bundle: docker_config.proxy_ca_bundle.clone(), provider_spiffe_workload_api_socket: docker_config .provider_spiffe_workload_api_socket .clone(), @@ -4713,11 +4713,35 @@ async fn docker_supervisor_bundle_archive( &contents, )?; } + append_docker_proxy_ca_bundle(&mut archive, config.proxy_ca_bundle.as_deref())?; archive .into_inner() .map_err(|error| Status::internal(format!("finish Docker supervisor archive: {error}"))) } +fn append_docker_proxy_ca_bundle( + archive: &mut tar::Builder>, + path: Option<&Path>, +) -> Result<(), Status> { + let Some(path) = path else { + return Ok(()); + }; + let path = path + .to_str() + .ok_or_else(|| Status::failed_precondition("proxy_ca_bundle must be valid UTF-8"))?; + let contents = + openshell_core::driver_utils::read_upstream_proxy_ca_bundle_file(path, "proxy_ca_bundle") + .map_err(Status::failed_precondition)?; + append_docker_archive_file( + archive, + "upstream-proxy-ca-bundle.pem", + 0o644, + SUPERVISOR_UID, + SUPERVISOR_GID, + contents.as_bytes(), + ) +} + async fn refresh_docker_boundary_authentication( sandbox_id: &str, config: &DockerDriverRuntimeConfig, @@ -5120,7 +5144,10 @@ async fn spawn_docker_control_process( workspace_root, format!("--health-socket-path={SUPERVISOR_HEALTH_SOCKET_PATH}"), ]; - command.extend(docker_upstream_proxy_cli_args(&config.upstream_proxy)); + command.extend(docker_upstream_proxy_cli_args( + &config.upstream_proxy, + config.proxy_ca_bundle.is_some(), + )); let mut supervisor_mounts = vec![ Mount { target: Some(BOUNDARY_MOUNT_PATH.to_string()), @@ -5871,7 +5898,30 @@ fn validate_docker_proxy_auth_file(config: &UpstreamProxyConfig) -> CoreResult<( Ok(()) } -fn docker_upstream_proxy_cli_args(config: &UpstreamProxyConfig) -> Vec { +fn validate_docker_proxy_ca_bundle(config: &DockerComputeConfig) -> CoreResult<()> { + let Some(path) = config.proxy_ca_bundle.as_ref() else { + return Ok(()); + }; + if path.as_os_str().is_empty() { + return Err(Error::config("proxy_ca_bundle must not be empty when set")); + } + if config.upstream_proxy.https_proxy.is_none() { + return Err(Error::config( + "proxy_ca_bundle is set but no https_proxy is configured", + )); + } + let path = path + .to_str() + .ok_or_else(|| Error::config("proxy_ca_bundle must be valid UTF-8"))?; + openshell_core::driver_utils::read_upstream_proxy_ca_bundle_file(path, "proxy_ca_bundle") + .map_err(Error::config)?; + Ok(()) +} + +fn docker_upstream_proxy_cli_args( + config: &UpstreamProxyConfig, + proxy_ca_bundle_configured: bool, +) -> Vec { let mut args = Vec::new(); if let Some(url) = config.https_proxy.as_ref() { args.extend(["--upstream-proxy".to_string(), url.clone()]); @@ -5891,6 +5941,12 @@ fn docker_upstream_proxy_cli_args(config: &UpstreamProxyConfig) -> Vec { if config.proxy_connect_by_hostname == Some(true) { args.push("--upstream-proxy-connect-by-hostname".to_string()); } + if proxy_ca_bundle_configured { + args.extend([ + "--upstream-proxy-ca-bundle".to_string(), + SUPERVISOR_PROXY_CA_BUNDLE_MOUNT_PATH.to_string(), + ]); + } args } diff --git a/crates/openshell-driver-docker/src/tests.rs b/crates/openshell-driver-docker/src/tests.rs index c5036f4547..bc747312c3 100644 --- a/crates/openshell-driver-docker/src/tests.rs +++ b/crates/openshell-driver-docker/src/tests.rs @@ -20,6 +20,7 @@ use openshell_core::proto::compute::v1::{ ResourceRequirements, WorkloadIdentityRequest, }; use std::fs; +use std::io::Read as _; use std::sync::Arc; use tempfile::TempDir; @@ -197,11 +198,174 @@ fn runtime_config() -> DockerDriverRuntimeConfig { sandbox_pids_limit: openshell_core::config::default_sandbox_pids_limit(), enable_bind_mounts: false, upstream_proxy: UpstreamProxyConfig::default(), + proxy_ca_bundle: None, provider_spiffe_workload_api_socket: None, app_armor_profile: Some(AppArmorProfile::Unconfined), } } +fn write_test_proxy_ca_bundle(directory: &TempDir) -> PathBuf { + let tls = generate_sandbox_tls_material(openshell_core::SandboxSessionId::new()) + .expect("generate test proxy CA"); + let path = directory.path().join("proxy-ca.pem"); + fs::write(&path, tls.trust_anchor_pem).expect("write test proxy CA"); + path +} + +#[test] +fn docker_config_parses_operator_proxy_ca_bundle() { + let config: DockerComputeConfig = toml::from_str( + r#" +https_proxy = "https://proxy.corp.example:8443" +proxy_ca_bundle = "/etc/openshell/tls/proxy-ca.pem" +"#, + ) + .expect("parse Docker proxy CA configuration"); + + assert_eq!( + config.proxy_ca_bundle, + Some(PathBuf::from("/etc/openshell/tls/proxy-ca.pem")) + ); +} + +#[test] +fn docker_proxy_ca_bundle_validation_is_fail_closed() { + let directory = TempDir::new().expect("create CA directory"); + let ca_bundle = write_test_proxy_ca_bundle(&directory); + let gateway_bind_address = "127.0.0.1:17670".parse().unwrap(); + + let mut valid = DockerComputeConfig::default(); + valid.upstream_proxy.https_proxy = Some("http://proxy.corp.example:8080".to_string()); + valid.proxy_ca_bundle = Some(ca_bundle); + valid + .validate_configuration(gateway_bind_address) + .expect("a valid CA bundle is accepted with an HTTP interception proxy"); + + let mut without_proxy = valid.clone(); + without_proxy.upstream_proxy.https_proxy = None; + let error = without_proxy + .validate_configuration(gateway_bind_address) + .expect_err("a CA bundle without a proxy must fail"); + assert!(error.to_string().contains("proxy_ca_bundle"), "{error}"); + assert!(error.to_string().contains("https_proxy"), "{error}"); + + let mut empty_path = valid.clone(); + empty_path.proxy_ca_bundle = Some(PathBuf::new()); + let error = empty_path + .validate_configuration(gateway_bind_address) + .expect_err("an empty CA bundle path must fail"); + assert!(error.to_string().contains("must not be empty"), "{error}"); + + let mut missing = valid.clone(); + missing.proxy_ca_bundle = Some(directory.path().join("missing.pem")); + let error = missing + .validate_configuration(gateway_bind_address) + .expect_err("a missing CA bundle must fail"); + assert!(error.to_string().contains("could not be read"), "{error}"); + + let malformed_path = directory.path().join("malformed.pem"); + fs::write(&malformed_path, "not a certificate\n").unwrap(); + let mut malformed = valid; + malformed.proxy_ca_bundle = Some(malformed_path); + let error = malformed + .validate_configuration(gateway_bind_address) + .expect_err("a certificate-free CA bundle must fail"); + assert!(error.to_string().contains("no PEM certificate"), "{error}"); +} + +#[tokio::test] +async fn docker_constructor_rejects_proxy_ca_bundle_without_proxy() { + let directory = TempDir::new().expect("create CA directory"); + let mut config = DockerComputeConfig { + socket_path: Some(directory.path().join("unused-docker.sock")), + proxy_ca_bundle: Some(write_test_proxy_ca_bundle(&directory)), + ..DockerComputeConfig::default() + }; + config.upstream_proxy.https_proxy = None; + + let Err(error) = + DockerComputeDriver::new("127.0.0.1:17670".parse().unwrap(), "info", &config).await + else { + panic!("constructor must reject incoherent proxy CA configuration before Docker I/O"); + }; + + assert!(error.to_string().contains("proxy_ca_bundle"), "{error}"); + assert!(error.to_string().contains("https_proxy"), "{error}"); +} + +#[test] +fn docker_proxy_ca_bundle_uses_fixed_supervisor_path() { + let proxy = UpstreamProxyConfig { + https_proxy: Some("https://proxy.corp.example:8443".to_string()), + ..UpstreamProxyConfig::default() + }; + + let args = docker_upstream_proxy_cli_args(&proxy, true); + let option = args + .iter() + .position(|arg| arg == "--upstream-proxy-ca-bundle") + .expect("proxy CA option"); + assert_eq!( + args.get(option + 1).map(String::as_str), + Some(SUPERVISOR_PROXY_CA_BUNDLE_MOUNT_PATH) + ); + assert!( + !args.iter().any(|arg| arg.contains("/etc/openshell/tls")), + "gateway-host paths must not appear in supervisor argv: {args:?}" + ); + + let args = docker_upstream_proxy_cli_args(&proxy, false); + assert!(!args.iter().any(|arg| arg == "--upstream-proxy-ca-bundle")); +} + +#[test] +fn docker_proxy_ca_bundle_is_staged_in_supervisor_archive() { + let directory = TempDir::new().expect("create CA directory"); + let ca_bundle = write_test_proxy_ca_bundle(&directory); + let expected = fs::read_to_string(&ca_bundle).unwrap(); + let mut builder = tar::Builder::new(Vec::new()); + + append_docker_proxy_ca_bundle(&mut builder, Some(&ca_bundle)).expect("append proxy CA bundle"); + let archive = builder.into_inner().expect("finish proxy CA archive"); + let mut archive = tar::Archive::new(archive.as_slice()); + let mut entries = archive.entries().unwrap(); + let mut entry = entries.next().expect("proxy CA entry").unwrap(); + + assert_eq!( + entry.path().unwrap().as_ref(), + Path::new("upstream-proxy-ca-bundle.pem") + ); + assert_eq!(entry.header().uid().unwrap(), u64::from(SUPERVISOR_UID)); + assert_eq!(entry.header().gid().unwrap(), u64::from(SUPERVISOR_GID)); + assert_eq!(entry.header().mode().unwrap(), 0o644); + let mut actual = String::new(); + entry.read_to_string(&mut actual).unwrap(); + assert_eq!(actual, expected); + assert!(entries.next().is_none()); +} + +#[test] +fn sandbox_driver_config_cannot_override_proxy_ca_bundle() { + let config = runtime_config(); + let mut sandbox = test_sandbox(); + sandbox + .spec + .as_mut() + .unwrap() + .template + .as_mut() + .unwrap() + .driver_config = Some(json_struct(serde_json::json!({ + "proxy_ca_bundle": "/workload/controlled-ca.pem" + }))); + + let error = DockerComputeDriver::validate_sandbox(&sandbox, &config) + .expect_err("sandbox driver config must not accept proxy_ca_bundle"); + assert_eq!(error.code(), tonic::Code::InvalidArgument); + assert!(error.message().contains("unknown field"), "{error}"); + assert!(error.message().contains("proxy_ca_bundle"), "{error}"); +} + fn test_workload_identity() -> ResolvedWorkloadIdentity { ResolvedWorkloadIdentity::new( 1234, diff --git a/crates/openshell-driver-kubernetes/src/driver.rs b/crates/openshell-driver-kubernetes/src/driver.rs index 5852db3ea3..425f6daaca 100644 --- a/crates/openshell-driver-kubernetes/src/driver.rs +++ b/crates/openshell-driver-kubernetes/src/driver.rs @@ -7332,7 +7332,7 @@ mod tests { assert_eq!(error.code(), tonic::Code::FailedPrecondition); assert!(error.message().contains("allow_driver_config")); assert!( - matches!(driver.create_sandbox_inner(&sandbox).await, Err(KubernetesDriverError::Precondition(message)) if message.contains("allow_driver_config")) + matches!(Box::pin(driver.create_sandbox_inner(&sandbox)).await, Err(KubernetesDriverError::Precondition(message)) if message.contains("allow_driver_config")) ); } } diff --git a/crates/openshell-driver-kubernetes/src/sandbox_runtime.rs b/crates/openshell-driver-kubernetes/src/sandbox_runtime.rs index 0588bbb95c..abee5eaf03 100644 --- a/crates/openshell-driver-kubernetes/src/sandbox_runtime.rs +++ b/crates/openshell-driver-kubernetes/src/sandbox_runtime.rs @@ -8,10 +8,11 @@ use std::path::Path; use k8s_openapi::ByteString; use k8s_openapi::api::core::v1::{ - CSIVolumeSource, Capabilities, Container, EmptyDirVolumeSource, EnvVar, ExecAction, KeyToPath, + CSIVolumeSource, Capabilities, Container, EmptyDirVolumeSource, EnvVar, KeyToPath, LocalObjectReference, Pod, PodSchedulingGate, PodSecurityContext, PodSpec, Probe, ProjectedVolumeSource, Secret, SecretVolumeSource, SecurityContext, Service, - ServiceAccountTokenProjection, ServicePort, ServiceSpec, Volume, VolumeMount, VolumeProjection, + ServiceAccountTokenProjection, ServicePort, ServiceSpec, TCPSocketAction, Volume, VolumeMount, + VolumeProjection, }; use k8s_openapi::apimachinery::pkg::apis::meta::v1::OwnerReference; use k8s_openapi::apimachinery::pkg::util::intstr::IntOrString; @@ -58,6 +59,9 @@ pub const CLIENT_TLS_CA_PATH: &str = "/.openshell/supervisor/client-ca.crt"; pub const CLIENT_TLS_CERTIFICATE_PATH: &str = "/.openshell/supervisor/client-tls.crt"; pub const CLIENT_TLS_PRIVATE_KEY_PATH: &str = "/.openshell/supervisor/client-tls.key"; pub const CONTROL_HEALTH_SOCKET_PATH: &str = "/run/openshell/health.sock"; +/// Kubelet `tcpSocket` readiness port. An exec probe would start a supervisor +/// process in every sandbox on every period. +pub const CONTROL_HEALTH_PORT: u16 = 5501; pub const NAMESPACE_WORKLOAD_POLICY_NAME: &str = "openshell-sandbox-workloads"; pub const NAMESPACE_SUPERVISOR_EGRESS_POLICY_NAME: &str = "openshell-sandbox-supervisors"; pub const SUPERVISOR_TERMINATION_GRACE_PERIOD_SECONDS: i64 = 30; @@ -334,6 +338,8 @@ pub fn supervisor_pod( "/sandbox".to_string(), "--health-socket-path".to_string(), CONTROL_HEALTH_SOCKET_PATH.to_string(), + "--health-port".to_string(), + CONTROL_HEALTH_PORT.to_string(), ]; if let Some(url) = https_proxy { command.extend(["--upstream-proxy".to_string(), url.to_string()]); @@ -411,13 +417,9 @@ pub fn supervisor_pod( termination_message_policy: Some("FallbackToLogsOnError".to_string()), env: Some(environment), readiness_probe: Some(Probe { - exec: Some(ExecAction { - command: Some(vec![ - "/openshell-supervisor".to_string(), - "health".to_string(), - "--socket".to_string(), - CONTROL_HEALTH_SOCKET_PATH.to_string(), - ]), + tcp_socket: Some(TCPSocketAction { + port: IntOrString::Int(i32::from(CONTROL_HEALTH_PORT)), + ..Default::default() }), period_seconds: Some(1), failure_threshold: Some(3), @@ -984,18 +986,14 @@ mod tests { .and_then(|capabilities| capabilities.drop.as_ref()), Some(&vec!["ALL".to_string()]) ); + let probe = container.readiness_probe.as_ref().expect("readiness probe"); + assert!( + probe.exec.is_none(), + "exec probes spawn a process per period" + ); assert_eq!( - container - .readiness_probe - .as_ref() - .and_then(|probe| probe.exec.as_ref()) - .and_then(|exec| exec.command.as_ref()), - Some(&vec![ - "/openshell-supervisor".to_string(), - "health".to_string(), - "--socket".to_string(), - CONTROL_HEALTH_SOCKET_PATH.to_string(), - ]) + probe.tcp_socket.as_ref().map(|tcp| &tcp.port), + Some(&IntOrString::Int(i32::from(CONTROL_HEALTH_PORT))) ); let command = container.command.as_ref().unwrap(); assert!( @@ -1003,6 +1001,12 @@ mod tests { .windows(2) .any(|args| args == ["--health-socket-path", CONTROL_HEALTH_SOCKET_PATH]) ); + let health_port = CONTROL_HEALTH_PORT.to_string(); + assert!( + command + .windows(2) + .any(|args| args == ["--health-port", health_port.as_str()]) + ); let env = container.env.as_ref().unwrap(); let env_value = |name: &str| { env.iter() diff --git a/crates/openshell-driver-mxc/src/driver.rs b/crates/openshell-driver-mxc/src/driver.rs index 31692d659f..84c36945b3 100644 --- a/crates/openshell-driver-mxc/src/driver.rs +++ b/crates/openshell-driver-mxc/src/driver.rs @@ -414,9 +414,11 @@ fn append_tls_readwrite_grant( } impl MxcComputeBackend { - pub fn new(config: MxcComputeConfig) -> Self { + /// Create an MXC backend whose audit records identify the configured gateway. + pub fn new(gateway_name: impl Into, config: MxcComputeConfig) -> Self { let invoker = WxcExecInvoker::new(&config.wxc_exec_path, config.debug); let (watch_tx, _) = broadcast::channel(256); + let gateway_name = gateway_name.into(); // Start the Plane-A ETW → OCSF consumer if enabled. The consumer thread // attributes each event to a `sandbox_id` via `attribution` (seeded by @@ -426,7 +428,7 @@ impl MxcComputeBackend { crate::etw_consumer::AttributionIndex::new(), )); let etw_session = if config.etw_audit { - match crate::etw_consumer::start_session(attribution.clone()) { + match crate::etw_consumer::start_session(attribution.clone(), gateway_name) { Ok(session) => Some(session), Err(e) => { warn!(error = %e, "MXC ETW audit consumer failed to start; continuing without it"); diff --git a/crates/openshell-driver-mxc/src/etw_consumer.rs b/crates/openshell-driver-mxc/src/etw_consumer.rs index 09b8079868..8117d2e70a 100644 --- a/crates/openshell-driver-mxc/src/etw_consumer.rs +++ b/crates/openshell-driver-mxc/src/etw_consumer.rs @@ -414,7 +414,10 @@ unsafe impl Send for OpenedTrace {} /// thread. The driver seeds `index` (pid → sandbox_id) as it launches sandboxes. /// /// Returns an [`EtwSession`] that must be kept alive; dropping it stops capture. -pub(crate) fn start_session(index: Arc>) -> Result { +pub(crate) fn start_session( + index: Arc>, + gateway_name: String, +) -> Result { let session_name = new_session_name(); let handle = start_trace_session(&session_name)?; @@ -423,6 +426,7 @@ pub(crate) fn start_session(index: Arc>) -> Result(EVENT_QUEUE_CAPACITY); + let processor = EtwEventProcessor::new(index, gateway_name); let consumer_thread = std::thread::Builder::new() .name("etw-ocsf-consumer".into()) @@ -440,7 +444,7 @@ pub(crate) fn start_session(index: Arc>) -> Result { release_queue_bytes(&consumer_health, raw.queued_bytes); if let Some(ev) = decode_raw(&mut raw) { - process_event(&index, ev); + processor.process_event(ev); } else { tracing::debug!( target: "mxc_etw", @@ -450,11 +454,11 @@ pub(crate) fn start_session(index: Arc>) -> Result { - drain_and_emit(&index); + processor.drain_and_emit(); overload_reporter.report_if_due(&consumer_health, false); } Err(mpsc::RecvTimeoutError::Disconnected) => { @@ -464,7 +468,7 @@ pub(crate) fn start_session(index: Arc>) -> Result, ev: DecodedEtwEvent) { - // Activity STOP is the empty twin of START — never a distinct OCSF row. - if ev.opcode == OPCODE_STOP { - return; +/// Consumer-thread state for attributing ETW events and emitting gateway OCSF. +/// The gateway identity is stable for the processor lifetime, while each event +/// is associated with a potentially different sandbox through `index`. +struct EtwEventProcessor { + index: Arc>, + gateway_name: String, +} + +impl EtwEventProcessor { + fn new(index: Arc>, gateway_name: String) -> Self { + Self { + index, + gateway_name, + } } - let resolved = { - let mut idx = index - .lock() - .unwrap_or_else(std::sync::PoisonError::into_inner); - if let Some(sid) = idx.resolve(&ev) { - let name = idx.name_of(&sid); - Some((sid, name, ev)) - } else { - // Not attributable yet: ETW delivers the create/config burst the - // instant `wxc-exec` starts, which can beat the driver's - // `register_launch`. Hold the event for replay instead of dropping - // it (see `drain_and_emit`). - idx.buffer_unresolved(ev); - None + /// Attribute one decoded event and emit an OCSF row for mapped classes. + /// Unresolved events are buffered for replay; attributed but unmapped events + /// are debug-logged by [`Self::emit_resolved`]. + fn process_event(&self, ev: DecodedEtwEvent) { + // Activity STOP is the empty twin of START — never a distinct OCSF row. + if ev.opcode == OPCODE_STOP { + return; } - }; - if let Some((sandbox_id, sandbox_name, ev)) = resolved { - emit_resolved(index, &sandbox_id, &sandbox_name, &ev); - } -} + let resolved = { + let mut idx = self + .index + .lock() + .unwrap_or_else(std::sync::PoisonError::into_inner); + if let Some(sid) = idx.resolve(&ev) { + let name = idx.name_of(&sid); + Some((sid, name, ev)) + } else { + // Not attributable yet: ETW delivers the create/config burst the + // instant `wxc-exec` starts, which can beat the driver's + // `register_launch`. Hold the event for replay instead of dropping + // it (see `drain_and_emit`). + idx.buffer_unresolved(ev); + None + } + }; -/// Re-resolve and emit any buffered events that have since become attributable. -/// Called by the consumer thread after each incoming event and on a periodic -/// tick, so a create/config burst that raced `register_launch` still lands in -/// the trail (and aged-out unresolvable events are dropped, bounded). -fn drain_and_emit(index: &Mutex) { - let ready = { - let mut idx = index - .lock() - .unwrap_or_else(std::sync::PoisonError::into_inner); - idx.drain_resolved() - }; - for (sandbox_id, sandbox_name, ev) in ready { - emit_resolved(index, &sandbox_id, &sandbox_name, &ev); + if let Some((sandbox_id, sandbox_name, ev)) = resolved { + self.emit_resolved(&sandbox_id, &sandbox_name, &ev); + } } -} -/// Map one attributed event to its OCSF class and emit it into the gateway trail. -fn emit_resolved( - index: &Mutex, - sandbox_id: &str, - sandbox_name: &str, - ev: &DecodedEtwEvent, -) { - // STOP twins are already filtered before buffering, so activity events - // reaching here are STARTs. - match ev.event_name.as_deref().unwrap_or("") { - // Lifecycle [6002]: MXC emits two create events per sandbox — - // `SandboxEngineCreate` and `SandboxCreateWithPolicyEnforcement` — and - // ETW drops them interchangeably under buffer pressure (observed: one run - // keeps the former, the next keeps the latter). Anchor on whichever - // arrives first and dedupe so the row is emitted exactly once. - "SandboxEngineCreate" | "SandboxCreateWithPolicyEnforcement" - if ev.opcode == OPCODE_START => - { - let first = index + /// Re-resolve and emit any buffered events that have since become attributable. + /// Called by the consumer thread after each incoming event and on a periodic + /// tick, so a create/config burst that raced `register_launch` still lands in + /// the trail (and aged-out unresolvable events are dropped, bounded). + fn drain_and_emit(&self) { + let ready = { + let mut idx = self + .index .lock() - .unwrap_or_else(std::sync::PoisonError::into_inner) - .take_lifecycle_once(sandbox_id); - if first { - let ctx = etw_ctx(sandbox_id, sandbox_name); - emit_ocsf(sandbox_id, map_lifecycle_create(&ctx, sandbox_name)); - } + .unwrap_or_else(std::sync::PoisonError::into_inner); + idx.drain_resolved() + }; + for (sandbox_id, sandbox_name, ev) in ready { + self.emit_resolved(&sandbox_id, &sandbox_name, &ev); } - // Process [1007]: `CreateProcessInSandbox` carries the real agent command - // line + working directory. The activity fires once empty (probe) and - // once with the command. Emit only the populated event, and map the - // command to a safe executable identity rather than durable arguments. - "CreateProcessInSandbox" if ev.opcode == OPCODE_START => { - if let Some(cmd) = ev.get_unquoted("commandLine") { - let ctx = etw_ctx(sandbox_id, sandbox_name); - emit_ocsf(sandbox_id, map_process_launch(&ctx, ev, &cmd)); - } else { + } + + /// Map one attributed event to its OCSF class and emit it into the gateway trail. + fn emit_resolved(&self, sandbox_id: &str, sandbox_name: &str, ev: &DecodedEtwEvent) { + // STOP twins are already filtered before buffering, so activity events + // reaching here are STARTs. + match ev.event_name.as_deref().unwrap_or("") { + // Lifecycle [6002]: MXC emits two create events per sandbox — + // `SandboxEngineCreate` and `SandboxCreateWithPolicyEnforcement` — and + // ETW drops them interchangeably under buffer pressure (observed: one run + // keeps the former, the next keeps the latter). Anchor on whichever + // arrives first and dedupe so the row is emitted exactly once. + "SandboxEngineCreate" | "SandboxCreateWithPolicyEnforcement" + if ev.opcode == OPCODE_START => + { + let first = self + .index + .lock() + .unwrap_or_else(std::sync::PoisonError::into_inner) + .take_lifecycle_once(sandbox_id); + if first { + let ctx = self.event_context(sandbox_id, sandbox_name); + emit_ocsf(sandbox_id, map_lifecycle_create(&ctx, sandbox_name)); + } + } + // Process [1007]: `CreateProcessInSandbox` carries the real agent command + // line + working directory. The activity fires once empty (probe) and + // once with the command. Emit only the populated event, and map the + // command to a safe executable identity rather than durable arguments. + "CreateProcessInSandbox" if ev.opcode == OPCODE_START => { + if let Some(cmd) = ev.get_unquoted("commandLine") { + let ctx = self.event_context(sandbox_id, sandbox_name); + emit_ocsf(sandbox_id, map_process_launch(&ctx, ev, &cmd)); + } else { + tracing::debug!(target: "mxc_etw", pid = ev.process_id, sandbox_id = %sandbox_id, "{}", ev.summary()); + } + } + // Process [1007]: `ProcessLaunched` is the confirmation twin of + // `CreateProcessInSandbox` — it carries the *actual* `processId`/`threadId` + // of the started in-sandbox process (the create event only has the request + + // command line). We emit it as a distinct PROC row so the trail records both + // the launch request (with executable identity) and confirmed start (with pid). + "ProcessLaunched" => { + let ctx = self.event_context(sandbox_id, sandbox_name); + emit_ocsf(sandbox_id, map_process_started(&ctx, ev)); + } + // Config [5019]: several distinct config/hardening/setup state changes. Each + // is a genuine audit-worthy config event; `SandboxConfig` is the richest but + // drops intermittently, so the reliably-captured hardening events + // (`Win32kLockdownApplied`, `ApplyUILimits`, `EnforceOsPolicy`) guarantee + // coverage. `SandboxProxyConfigured` (network/proxy setup — the one + // network-plane event the provider emits) and `SandboxConsoleReferencePlumbed` + // (console-handle plumbing) are additional per-sandbox setup state changes. + "SandboxConfig" + | "Win32kLockdownApplied" + | "ApplyUILimits" + | "EnforceOsPolicy" + | "SandboxProxyConfigured" + | "SandboxConsoleReferencePlumbed" => { + // Dump the raw decoded field set for config-family events at debug so we + // can confirm the exact property names MXC emits (e.g. which key carries + // the proxy port on `SandboxProxyConfigured`). Guarded by `debug=true`. + tracing::debug!(target: "mxc_etw", pid = ev.process_id, sandbox_id = %sandbox_id, "{}", ev.summary()); + let ctx = self.event_context(sandbox_id, sandbox_name); + emit_ocsf(sandbox_id, map_config_state(&ctx, ev)); + } + // Finding [2004]: MXC surfaces WIL error/fallback activities during + // sandbox setup. Captured as informational (non-alert) findings so the + // audit trail records setup anomalies without crying wolf. + "ActivityError" | "FallbackError" => { + let ctx = self.event_context(sandbox_id, sandbox_name); + emit_ocsf(sandbox_id, map_finding(&ctx, ev)); + } + _ => { tracing::debug!(target: "mxc_etw", pid = ev.process_id, sandbox_id = %sandbox_id, "{}", ev.summary()); } } - // Process [1007]: `ProcessLaunched` is the confirmation twin of - // `CreateProcessInSandbox` — it carries the *actual* `processId`/`threadId` - // of the started in-sandbox process (the create event only has the request + - // command line). We emit it as a distinct PROC row so the trail records both - // the launch request (with executable identity) and confirmed start (with pid). - "ProcessLaunched" => { - let ctx = etw_ctx(sandbox_id, sandbox_name); - emit_ocsf(sandbox_id, map_process_started(&ctx, ev)); - } - // Config [5019]: several distinct config/hardening/setup state changes. Each - // is a genuine audit-worthy config event; `SandboxConfig` is the richest but - // drops intermittently, so the reliably-captured hardening events - // (`Win32kLockdownApplied`, `ApplyUILimits`, `EnforceOsPolicy`) guarantee - // coverage. `SandboxProxyConfigured` (network/proxy setup — the one - // network-plane event the provider emits) and `SandboxConsoleReferencePlumbed` - // (console-handle plumbing) are additional per-sandbox setup state changes. - "SandboxConfig" - | "Win32kLockdownApplied" - | "ApplyUILimits" - | "EnforceOsPolicy" - | "SandboxProxyConfigured" - | "SandboxConsoleReferencePlumbed" => { - // Dump the raw decoded field set for config-family events at debug so we - // can confirm the exact property names MXC emits (e.g. which key carries - // the proxy port on `SandboxProxyConfigured`). Guarded by `debug=true`. - tracing::debug!(target: "mxc_etw", pid = ev.process_id, sandbox_id = %sandbox_id, "{}", ev.summary()); - let ctx = etw_ctx(sandbox_id, sandbox_name); - emit_ocsf(sandbox_id, map_config_state(&ctx, ev)); - } - // Finding [2004]: MXC surfaces WIL error/fallback activities during - // sandbox setup. Captured as informational (non-alert) findings so the - // audit trail records setup anomalies without crying wolf. - "ActivityError" | "FallbackError" => { - let ctx = etw_ctx(sandbox_id, sandbox_name); - emit_ocsf(sandbox_id, map_finding(&ctx, ev)); - } - _ => { - tracing::debug!(target: "mxc_etw", pid = ev.process_id, sandbox_id = %sandbox_id, "{}", ev.summary()); + } + + /// Build a per-event OCSF context rather than using the process-wide `ctx()` + /// singleton: one gateway process hosts many sandboxes. The gateway identity + /// comes from this processor, while the affected sandbox varies per event. + fn event_context(&self, sandbox_id: &str, sandbox_name: &str) -> EventContext { + EventContext { + sandbox_id: sandbox_id.to_string(), + sandbox_name: sandbox_name.to_string(), + container_image: "mxc/appcontainer".to_string(), + origin: openshell_ocsf::EventOrigin::Gateway { + name: self.gateway_name.clone(), + }, + hostname: gateway_hostname().to_string(), + product_version: env!("CARGO_PKG_VERSION").to_string(), + proxy_ip: std::net::IpAddr::V4(std::net::Ipv4Addr::LOCALHOST), + proxy_port: 0, } } } @@ -1878,20 +1915,6 @@ fn gateway_hostname() -> &'static str { }) } -/// Build a per-event OCSF context (not the process-wide `ctx()` singleton, since -/// one gateway process hosts many sandboxes — wrinkle #1). -fn etw_ctx(sandbox_id: &str, sandbox_name: &str) -> EventContext { - EventContext { - sandbox_id: sandbox_id.to_string(), - sandbox_name: sandbox_name.to_string(), - container_image: "mxc/appcontainer".to_string(), - hostname: gateway_hostname().to_string(), - product_version: env!("CARGO_PKG_VERSION").to_string(), - proxy_ip: std::net::IpAddr::V4(std::net::Ipv4Addr::LOCALHOST), - proxy_port: 0, - } -} - // --------------------------------------------------------------------------- // Helpers // --------------------------------------------------------------------------- @@ -1946,6 +1969,13 @@ fn wide_str_at(buf: &[u8], offset: u32) -> Option { mod tests { use super::*; + fn event_processor(gateway_name: &str) -> EtwEventProcessor { + EtwEventProcessor::new( + Arc::new(Mutex::new(AttributionIndex::new())), + gateway_name.to_string(), + ) + } + fn empty_raw_event(queued_bytes: usize) -> RawEtwEvent { RawEtwEvent { header: EVENT_HEADER::default(), @@ -2065,10 +2095,26 @@ mod tests { ); } + #[test] + fn mxc_events_identify_the_gateway_and_affected_sandbox_separately() { + let processor = event_processor("production"); + let ctx = processor.event_context("sbx-123", "agent-01"); + let event = map_lifecycle_create(&ctx, "agent-01"); + let json = event.to_json().expect("lifecycle event should serialize"); + + assert_eq!(json["metadata"]["product"]["name"], "OpenShell Gateway"); + assert_eq!(json["device"]["name"], "production"); + assert_eq!(json["device"]["uid"], "production"); + assert_eq!(json["device"]["type"], "Server"); + assert_eq!(json["container"]["name"], "agent-01"); + assert_eq!(json["container"]["uid"], "sbx-123"); + } + #[test] fn process_launch_omits_command_arguments_from_audit_output() { const SECRET: &str = "secret-value"; - let ctx = etw_ctx("sbx-secret-test", "secret-test"); + let processor = event_processor("production"); + let ctx = processor.event_context("sbx-secret-test", "secret-test"); let mut ev = mk_event(4242, "CreateProcessInSandbox"); ev.props .push(("currentDirectory".into(), r#""C:\work\openshell""#.into())); diff --git a/crates/openshell-driver-mxc/src/grpc.rs b/crates/openshell-driver-mxc/src/grpc.rs index f82a8fd85f..4c21172c47 100644 --- a/crates/openshell-driver-mxc/src/grpc.rs +++ b/crates/openshell-driver-mxc/src/grpc.rs @@ -186,8 +186,10 @@ mod tests { #[tokio::test] async fn start_sandbox_reports_one_shot_lifecycle() { - let service = - ComputeDriverService::new(MxcComputeBackend::new(MxcComputeConfig::default())); + let service = ComputeDriverService::new(MxcComputeBackend::new( + openshell_core::config::DEFAULT_GATEWAY_NAME, + MxcComputeConfig::default(), + )); let error = service .start_sandbox(Request::new(StartSandboxRequest::default())) @@ -200,8 +202,10 @@ mod tests { #[tokio::test] async fn sandbox_authentication_is_not_supported() { - let service = - ComputeDriverService::new(MxcComputeBackend::new(MxcComputeConfig::default())); + let service = ComputeDriverService::new(MxcComputeBackend::new( + openshell_core::config::DEFAULT_GATEWAY_NAME, + MxcComputeConfig::default(), + )); let error = service .authenticate_sandbox(Request::new(AuthenticateSandboxRequest::default())) @@ -214,8 +218,10 @@ mod tests { #[tokio::test] async fn workspace_lifecycle_is_an_idempotent_no_op() { - let service = - ComputeDriverService::new(MxcComputeBackend::new(MxcComputeConfig::default())); + let service = ComputeDriverService::new(MxcComputeBackend::new( + openshell_core::config::DEFAULT_GATEWAY_NAME, + MxcComputeConfig::default(), + )); service .ensure_workspace(Request::new(EnsureWorkspaceRequest::default())) diff --git a/crates/openshell-driver-mxc/tests/wxc_exec_real.rs b/crates/openshell-driver-mxc/tests/wxc_exec_real.rs index 960067b784..df4c3a5dba 100644 --- a/crates/openshell-driver-mxc/tests/wxc_exec_real.rs +++ b/crates/openshell-driver-mxc/tests/wxc_exec_real.rs @@ -767,7 +767,7 @@ async fn pc_https_egress_reads_injected_ca_bundle() { egress_proxy_addr: "127.0.0.1:18080".to_string(), ..Default::default() }; - let backend = MxcComputeBackend::new(config); + let backend = MxcComputeBackend::new(openshell_core::config::DEFAULT_GATEWAY_NAME, config); backend .create_sandbox(&sandbox) .await diff --git a/crates/openshell-driver-podman/Cargo.toml b/crates/openshell-driver-podman/Cargo.toml index c49c309f65..18b1928c33 100644 --- a/crates/openshell-driver-podman/Cargo.toml +++ b/crates/openshell-driver-podman/Cargo.toml @@ -51,6 +51,7 @@ openshell-otel-test-support = { path = "../openshell-otel-test-support" } opentelemetry_sdk = { workspace = true, features = ["testing"] } prost-types = { workspace = true } temp-env = "0.3" +tempfile = "3" tokio = { workspace = true, features = ["test-util"] } [lints] diff --git a/crates/openshell-driver-podman/src/container.rs b/crates/openshell-driver-podman/src/container.rs index bb4ee88c9f..3092eb73b3 100644 --- a/crates/openshell-driver-podman/src/container.rs +++ b/crates/openshell-driver-podman/src/container.rs @@ -263,6 +263,11 @@ pub struct ContainerSpec { stop_timeout: u32, /// Extra /etc/hosts entries for the networked supervisor container. /// The isolated workload resolves host aliases through policy DNS. + /// Native restart stays disabled; the gateway owns sandbox restart policy. + restart_policy: String, + /// Extra /etc/hosts entries. Used to inject `host.containers.internal` + /// via Podman's `host-gateway` magic so sandbox containers can reach + /// the gateway server running on the host in rootless mode. hostadd: Vec, /// Search domains written to `/etc/resolv.conf` by Podman. dns_search: Vec, @@ -1283,6 +1288,10 @@ fn build_base_spec( // Inject stable host aliases into the networked supervisor container. // The workload clears these entries and resolves the driver-neutral // alias through the policy-DNS relay instead. + restart_policy: "no".to_string(), + // Inject stable host aliases into /etc/hosts so sandbox containers can + // reach services on the host. `host.openshell.internal` is the driver- + // neutral alias used by policies and e2e tests. hostadd: hostadd_entries(config), // Preserve Podman's resolver defaults for both policy-DNS and ordinary // sandboxes. Namespace-local capture supports UDP and TCP, so it must @@ -3345,6 +3354,12 @@ mod tests { config.host_gateway_ip = "192.168.127.254".to_string(); let spec = build_container_spec(&sandbox, &config); + assert_eq!( + spec["restart_policy"].as_str(), + Some("no"), + "the gateway owns sandbox restart policy" + ); + let hostadd: Vec<&str> = spec["hostadd"] .as_array() .expect("hostadd should be an array") diff --git a/e2e/rust/tests/podman_preflight.rs b/crates/openshell-driver-podman/tests/podman_preflight.rs similarity index 60% rename from e2e/rust/tests/podman_preflight.rs rename to crates/openshell-driver-podman/tests/podman_preflight.rs index afb3bc1c38..439cb95d98 100644 --- a/e2e/rust/tests/podman_preflight.rs +++ b/crates/openshell-driver-podman/tests/podman_preflight.rs @@ -1,51 +1,22 @@ // SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. // SPDX-License-Identifier: Apache-2.0 -#![cfg(feature = "e2e-podman")] - -//! Podman driver daemon-unavailable e2e tests. +//! Podman driver daemon-unavailable integration tests. //! //! These tests verify that `openshell-driver-podman` fails fast with an //! actionable error when it cannot reach a Podman API socket, instead of //! hanging or silently serving gRPC against a dead connection. //! -//! The tests do NOT require a running Podman daemon or gateway — they point +//! They do NOT require a running Podman daemon or gateway — they point //! `--podman-socket` at a path that is guaranteed not to exist to simulate -//! the daemon being unavailable. +//! the daemon being unavailable. As a plain Cargo integration test in this +//! crate, this runs via the normal `cargo test -p openshell-driver-podman` +//! lane with no special CI wiring: Cargo provides `CARGO_BIN_EXE_` for +//! this crate's own `[[bin]]` target automatically. -use std::path::{Path, PathBuf}; +use std::path::PathBuf; use std::time::{Duration, Instant}; -use openshell_e2e::harness::output::strip_ansi; - -/// Locate the workspace root by walking up from this crate's manifest directory. -fn workspace_root() -> PathBuf { - Path::new(env!("CARGO_MANIFEST_DIR")) - .ancestors() - .nth(2) - .expect("failed to resolve workspace root from CARGO_MANIFEST_DIR") - .to_path_buf() -} - -/// Return the path to the `openshell-driver-podman` binary. -/// -/// Uses `OPENSHELL_EXTERNAL_DRIVER_BIN` when set (the same env var the shell -/// e2e harness uses for prebuilt standalone driver artifacts), otherwise -/// expects the binary at `/target/debug/openshell-driver-podman`. -fn driver_podman_bin() -> PathBuf { - let bin = std::env::var_os("OPENSHELL_EXTERNAL_DRIVER_BIN").map_or_else( - || workspace_root().join("target/debug/openshell-driver-podman"), - PathBuf::from, - ); - assert!( - bin.is_file(), - "openshell-driver-podman binary not found at {} — set OPENSHELL_EXTERNAL_DRIVER_BIN \ - or run `cargo build -p openshell-driver-podman` first", - bin.display() - ); - bin -} - /// Run `openshell-driver-podman` pointed at a Podman socket that does not /// exist, and wait for it to exit. /// @@ -53,17 +24,20 @@ fn driver_podman_bin() -> PathBuf { /// socket briefly re-activating), so this can take several seconds. async fn run_with_unreachable_podman_socket() -> (String, i32, Duration, PathBuf) { let tmpdir = tempfile::tempdir().expect("create isolated socket dir"); - let missing_socket = tmpdir.path().join("openshell-e2e-nonexistent-podman.sock"); + // Use a short relative path so miette cannot insert a line-wrap gutter + // inside it on platforms with long temporary-directory paths. + let missing_socket = PathBuf::from("missing-podman.sock"); let start = Instant::now(); - let mut cmd = tokio::process::Command::new(driver_podman_bin()); + let mut cmd = tokio::process::Command::new(env!("CARGO_BIN_EXE_openshell-driver-podman")); cmd.arg("--podman-socket") .arg(&missing_socket) + .current_dir(tmpdir.path()) .kill_on_drop(true) .stdout(std::process::Stdio::piped()) .stderr(std::process::Stdio::piped()); - let output = tokio::time::timeout(Duration::from_secs(60), cmd.output()) + let output = tokio::time::timeout(Duration::from_mins(1), cmd.output()) .await .expect("openshell-driver-podman should exit instead of hanging") .expect("spawn openshell-driver-podman"); @@ -102,15 +76,14 @@ async fn driver_error_names_unreachable_socket() { let (output, code, _, missing_socket) = run_with_unreachable_podman_socket().await; assert_ne!(code, 0); - let clean = strip_ansi(&output); assert!( - clean.contains("connection error"), - "driver error should describe a connection failure:\n{clean}" + output.contains("connection error"), + "driver error should describe a connection failure:\n{output}" ); assert!( - clean.contains(missing_socket.to_str().expect("socket path is utf-8")), - "driver error should name the unreachable socket path {}:\n{clean}", + output.contains(missing_socket.to_str().expect("socket path is utf-8")), + "driver error should name the unreachable socket path {}:\n{output}", missing_socket.display() ); } diff --git a/crates/openshell-driver-vm/src/driver.rs b/crates/openshell-driver-vm/src/driver.rs index 8cdabb22c5..58ad21892e 100644 --- a/crates/openshell-driver-vm/src/driver.rs +++ b/crates/openshell-driver-vm/src/driver.rs @@ -2315,21 +2315,15 @@ impl VmDriver { ); } - if clear_stop_marker { - match tokio::fs::remove_file(state_dir.join(SANDBOX_STOPPED_FILE)).await { - Ok(()) => {} - Err(err) if err.kind() == std::io::ErrorKind::NotFound => {} - Err(err) => { - self.registry.lock().await.remove(&sandbox.id); - warn!( - sandbox_id = %sandbox.id, - state_dir = %state_dir.display(), - error = %err, - "vm driver: cannot clear stop marker for persisted sandbox restore" - ); - return false; - } - } + if clear_stop_marker && let Err(err) = clear_explicit_start_markers(&state_dir).await { + self.registry.lock().await.remove(&sandbox.id); + warn!( + sandbox_id = %sandbox.id, + state_dir = %state_dir.display(), + error = %err, + "vm driver: cannot clear lifecycle markers for persisted sandbox restore" + ); + return false; } self.publish_platform_event( @@ -6214,6 +6208,19 @@ async fn remove_runtime_generation_material(state_dir: &Path) -> Result<(), Stri Ok(()) } +async fn clear_explicit_start_markers(state_dir: &Path) -> Result<(), std::io::Error> { + // Remove the terminal marker first. If clearing the durable stop marker + // fails, the sandbox remains stopped and a later start can safely retry. + for marker in [MAIN_PROCESS_EXITED_FILE, SANDBOX_STOPPED_FILE] { + match tokio::fs::remove_file(state_dir.join(marker)).await { + Ok(()) => {} + Err(err) if err.kind() == std::io::ErrorKind::NotFound => {} + Err(err) => return Err(err), + } + } + Ok(()) +} + #[cfg(unix)] async fn restrict_owner_read_write(path: &Path) -> Result<(), std::io::Error> { tokio::fs::set_permissions(path, fs::Permissions::from_mode(0o600)).await @@ -8735,6 +8742,9 @@ mod tests { tokio::fs::write(state_dir.join(SANDBOX_STOPPED_FILE), b"stopped\n") .await .unwrap(); + tokio::fs::write(state_dir.join(MAIN_PROCESS_EXITED_FILE), b"terminal\n") + .await + .unwrap(); let snapshot = sandbox_snapshot(&sandbox, stopped_condition(), false); driver.registry.lock().await.insert( sandbox.id.clone(), @@ -8766,6 +8776,12 @@ mod tests { .is_ok(), "failed start must retain its durable stop marker" ); + assert!( + tokio::fs::metadata(state_dir.join(MAIN_PROCESS_EXITED_FILE)) + .await + .is_ok(), + "failed start must retain its terminal marker" + ); let restored = driver .get_sandbox(&sandbox.id, &sandbox.name) .await @@ -8872,6 +8888,32 @@ mod tests { ) } + #[tokio::test] + async fn explicit_start_clears_stop_and_terminal_markers() { + let temp = tempfile::tempdir().unwrap(); + let state_dir = temp.path().join("sandbox"); + tokio::fs::create_dir_all(&state_dir).await.unwrap(); + tokio::fs::write(state_dir.join(SANDBOX_STOPPED_FILE), b"stopped\n") + .await + .unwrap(); + tokio::fs::write(state_dir.join(MAIN_PROCESS_EXITED_FILE), b"terminal\n") + .await + .unwrap(); + + clear_explicit_start_markers(&state_dir).await.unwrap(); + + assert!( + !tokio::fs::try_exists(state_dir.join(SANDBOX_STOPPED_FILE)) + .await + .unwrap() + ); + assert!( + !tokio::fs::try_exists(state_dir.join(MAIN_PROCESS_EXITED_FILE)) + .await + .unwrap() + ); + } + #[test] fn prepare_sandbox_overlay_preserves_existing_overlay_on_start() { let base = unique_temp_dir(); diff --git a/crates/openshell-gateway/src/lib.rs b/crates/openshell-gateway/src/lib.rs index b67e93c102..8864e85ede 100644 --- a/crates/openshell-gateway/src/lib.rs +++ b/crates/openshell-gateway/src/lib.rs @@ -158,7 +158,7 @@ impl openshell_server::ComputeDriverFactory for MxcFactory { context: openshell_server::ComputeDriverBuildContext<'_>, ) -> openshell_core::Result { let config: openshell_driver_mxc::MxcComputeConfig = context.driver_config()?; - let backend = openshell_driver_mxc::MxcComputeBackend::new(config); + let backend = openshell_driver_mxc::MxcComputeBackend::new(context.gateway_name(), config); let driver = openshell_driver_mxc::ComputeDriverService::new(backend); Ok(openshell_server::ComputeDriverInstance::InProcess( std::sync::Arc::new(driver), diff --git a/crates/openshell-isolation-interface/src/linux/seccomp_notify.rs b/crates/openshell-isolation-interface/src/linux/seccomp_notify.rs index c9bbb0fafa..a87dc080d5 100644 --- a/crates/openshell-isolation-interface/src/linux/seccomp_notify.rs +++ b/crates/openshell-isolation-interface/src/linux/seccomp_notify.rs @@ -432,7 +432,8 @@ pub fn install_listener(syscalls: &[i64]) -> io::Result { /// reconfigure an INET endpoint. Connected `send()`/null-destination /// `sendto()` retains the audited cBPF fast path. pub fn install_workload_listener() -> io::Result { - install_listener(&[ + #[allow(unused_mut)] // SYS_open is unavailable on some architectures. + let mut syscalls = vec![ libc::SYS_socket, libc::SYS_connect, libc::SYS_bind, @@ -447,7 +448,12 @@ pub fn install_workload_listener() -> io::Result { libc::SYS_kill, libc::SYS_tkill, libc::SYS_rt_sigqueueinfo, - ]) + libc::SYS_openat, + libc::SYS_openat2, + ]; + #[cfg(target_arch = "x86_64")] + syscalls.push(libc::SYS_open); + install_listener(&syscalls) } /// Run a no-capability conformance probe. diff --git a/crates/openshell-isolation-interface/src/linux/workload_launcher.rs b/crates/openshell-isolation-interface/src/linux/workload_launcher.rs index 8602d9df80..71fc7f336e 100644 --- a/crates/openshell-isolation-interface/src/linux/workload_launcher.rs +++ b/crates/openshell-isolation-interface/src/linux/workload_launcher.rs @@ -137,7 +137,27 @@ mod tests { }) .expect("launcher result") .expect("spawn child"); - let notification = listener.receive().expect("receive child socket"); + let notification = loop { + let notification = listener.receive().expect("receive child syscall"); + let syscall = i64::from(notification.syscall); + if syscall == libc::SYS_socket { + break notification; + } + if syscall == libc::SYS_openat || syscall == libc::SYS_openat2 { + listener + .respond_continue(notification.id) + .expect("continue ordinary file open"); + continue; + } + #[cfg(target_arch = "x86_64")] + if syscall == libc::SYS_open { + listener + .respond_continue(notification.id) + .expect("continue ordinary file open"); + continue; + } + panic!("unexpected child syscall: {syscall}"); + }; assert_eq!(i64::from(notification.syscall), libc::SYS_socket); assert!( std::path::Path::new(&format!("/proc/{}/task/{}", child.id(), notification.tid)) diff --git a/crates/openshell-ocsf/src/builders/mod.rs b/crates/openshell-ocsf/src/builders/mod.rs index bf041e5540..6d5b83c13f 100644 --- a/crates/openshell-ocsf/src/builders/mod.rs +++ b/crates/openshell-ocsf/src/builders/mod.rs @@ -161,6 +161,18 @@ use crate::enums::StatusId; use crate::events::base_event::BaseEventData; use crate::objects::{Container, Device, Endpoint, Image, Metadata, Product}; +/// Which `OpenShell` component produced an event. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum EventOrigin { + /// The supervisor process enforcing sandbox policy. + Supervisor, + /// The gateway process, which has no sandbox or container of its own. + Gateway { + /// Operator-assigned gateway name (`[openshell.gateway] name`). + name: String, + }, +} + /// Immutable context created once at sandbox startup. /// /// Passed to every event builder to populate shared OCSF fields @@ -181,15 +193,21 @@ pub struct EventContext { pub proxy_ip: IpAddr, /// Proxy listen port. pub proxy_port: u16, + /// Which component is emitting. + pub origin: EventOrigin, } impl EventContext { /// Build the OCSF `Metadata` object for any event. #[must_use] pub fn metadata(&self, profiles: &[&str]) -> Metadata { + let product = match self.origin { + EventOrigin::Supervisor => Product::openshell_sandbox(&self.product_version), + EventOrigin::Gateway { .. } => Product::openshell_gateway(&self.product_version), + }; Metadata { version: OCSF_VERSION.to_string(), - product: Product::openshell_sandbox(&self.product_version), + product, profiles: profiles.iter().map(|s| (*s).to_string()).collect(), uid: Some(uuid::Uuid::new_v4().to_string()), log_source: None, @@ -215,7 +233,10 @@ impl EventContext { /// on (Linux for the in-sandbox supervisor, Windows for the MXC gateway). #[must_use] pub fn device(&self) -> Device { - Device::for_current_os(&self.hostname) + match &self.origin { + EventOrigin::Supervisor => Device::for_current_os(&self.hostname), + EventOrigin::Gateway { name } => Device::gateway(&self.hostname, name), + } } /// Build the `proxy_endpoint` object for the Network Proxy profile. @@ -255,6 +276,7 @@ pub(crate) fn test_sandbox_context() -> EventContext { product_version: "0.1.0".to_string(), proxy_ip: "10.42.0.1".parse().unwrap(), proxy_port: 3128, + origin: EventOrigin::Supervisor, } } diff --git a/crates/openshell-ocsf/src/ctx.rs b/crates/openshell-ocsf/src/ctx.rs index f714b77ab2..7e1c4c7796 100644 --- a/crates/openshell-ocsf/src/ctx.rs +++ b/crates/openshell-ocsf/src/ctx.rs @@ -8,7 +8,7 @@ //! not been set (e.g. unit tests that exercise builders without booting the //! sandbox). -use crate::EventContext; +use crate::{EventContext, EventOrigin}; use std::sync::{LazyLock, OnceLock}; static OCSF_CTX: OnceLock = OnceLock::new(); @@ -21,6 +21,7 @@ static OCSF_CTX_FALLBACK: LazyLock = LazyLock::new(|| EventContext product_version: env!("CARGO_PKG_VERSION").to_string(), proxy_ip: std::net::IpAddr::from([127, 0, 0, 1]), proxy_port: 3128, + origin: EventOrigin::Supervisor, }); /// Initialise the process-wide OCSF sandbox context. diff --git a/crates/openshell-ocsf/src/enums/device_type.rs b/crates/openshell-ocsf/src/enums/device_type.rs index b47596e9d0..369189390f 100644 --- a/crates/openshell-ocsf/src/enums/device_type.rs +++ b/crates/openshell-ocsf/src/enums/device_type.rs @@ -15,6 +15,8 @@ use serde_repr::{Deserialize_repr, Serialize_repr}; pub enum DeviceTypeId { /// 0 — Unknown Unknown = 0, + /// 1 — Server + Server = 1, /// 99 — Other Other = 99, } @@ -30,6 +32,7 @@ impl std::fmt::Display for DeviceTypeId { fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { formatter.write_str(match self { Self::Unknown => "Unknown", + Self::Server => "Server", Self::Other => "Other", }) } @@ -43,6 +46,7 @@ mod tests { fn device_type_display_uses_schema_labels() { for (device_type, expected) in [ (DeviceTypeId::Unknown, "Unknown"), + (DeviceTypeId::Server, "Server"), (DeviceTypeId::Other, "Other"), ] { assert_eq!(device_type.to_string(), expected); @@ -52,7 +56,11 @@ mod tests { #[test] fn device_type_json_roundtrip() { - for (device_type, expected) in [(DeviceTypeId::Unknown, 0), (DeviceTypeId::Other, 99)] { + for (device_type, expected) in [ + (DeviceTypeId::Unknown, 0), + (DeviceTypeId::Server, 1), + (DeviceTypeId::Other, 99), + ] { let json = serde_json::to_value(device_type).unwrap(); assert_eq!(json, serde_json::json!(expected)); let decoded: DeviceTypeId = serde_json::from_value(json).unwrap(); diff --git a/crates/openshell-ocsf/src/format/shorthand.rs b/crates/openshell-ocsf/src/format/shorthand.rs index 77e6751b10..66e1f1c3d1 100644 --- a/crates/openshell-ocsf/src/format/shorthand.rs +++ b/crates/openshell-ocsf/src/format/shorthand.rs @@ -134,12 +134,24 @@ fn reason_tag(base: &BaseEventData) -> String { .map_or_else(String::new, |text| format!(" [reason:{text}]")) } -fn unmapped_fields(base: &BaseEventData) -> Vec { - base.unmapped +fn sorted_unmapped_fields(base: &BaseEventData) -> Vec<(&str, &serde_json::Value)> { + let mut fields: Vec<_> = base + .unmapped .as_ref() .and_then(serde_json::Value::as_object) .into_iter() .flatten() + .map(|(key, value)| (key.as_str(), value)) + .collect(); + // Cargo can enable insertion-ordered JSON maps through another dependency. + // Keep shorthand ordering and truncated field selection stable either way. + fields.sort_unstable_by_key(|(key, _)| *key); + fields +} + +fn unmapped_fields(base: &BaseEventData) -> Vec { + sorted_unmapped_fields(base) + .into_iter() .filter_map(|(key, value)| { let value = match value { serde_json::Value::Bool(value) => value.to_string(), @@ -502,9 +514,9 @@ impl OcsfEvent { if obj.is_empty() { return None; } - let fields: Vec = obj - .iter() - .take(3) // Limit to 3 most important fields + let fields: Vec = sorted_unmapped_fields(&e.base) + .into_iter() + .take(3) .map(|(k, v)| { let val = v.as_str().map_or_else(|| v.to_string(), String::from); format!("{k}:{val}") @@ -676,8 +688,8 @@ mod tests { #[test] fn test_http_activity_shorthand_includes_unmapped_attributes() { let mut base = base(4002, "HTTP Activity", 4, "Network Activity", 99, "Other"); - base.add_unmapped("attempt", serde_json::json!(2)); base.add_unmapped("cached", serde_json::json!(true)); + base.add_unmapped("attempt", serde_json::json!(2)); let event = OcsfEvent::HttpActivity(HttpActivityEvent { base, http_request: Some(HttpRequest::new( @@ -1246,4 +1258,19 @@ mod tests { "EVENT [INFO] Network namespace created [ns:openshell-sandbox-abc123]" ); } + + #[test] + fn test_base_event_selects_unmapped_fields_in_key_order() { + let mut b = base(0, "Base Event", 0, "Uncategorized", 99, "Other"); + b.set_message("Context"); + for (key, value) in [("z", 4), ("c", 3), ("a", 1), ("b", 2)] { + b.add_unmapped(key, serde_json::json!(value)); + } + + let event = OcsfEvent::Base(BaseEvent { base: b }); + assert_eq!( + event.format_shorthand(), + "EVENT [INFO] Context [a:1 b:2 c:3]" + ); + } } diff --git a/crates/openshell-ocsf/src/lib.rs b/crates/openshell-ocsf/src/lib.rs index e65301e0c1..ba8eb70f43 100644 --- a/crates/openshell-ocsf/src/lib.rs +++ b/crates/openshell-ocsf/src/lib.rs @@ -58,8 +58,8 @@ pub use objects::{ // --- Builders --- pub use builders::{ ApiActivityBuilder, AppLifecycleBuilder, BaseEventBuilder, ConfigStateChangeBuilder, - DetectionFindingBuilder, EventContext, HttpActivityBuilder, NetworkActivityBuilder, - ProcessActivityBuilder, SshActivityBuilder, + DetectionFindingBuilder, EventContext, EventOrigin, HttpActivityBuilder, + NetworkActivityBuilder, ProcessActivityBuilder, SshActivityBuilder, }; // --- Tracing layers --- diff --git a/crates/openshell-ocsf/src/objects/device.rs b/crates/openshell-ocsf/src/objects/device.rs index 89c52cdc85..0984c1b7cf 100644 --- a/crates/openshell-ocsf/src/objects/device.rs +++ b/crates/openshell-ocsf/src/objects/device.rs @@ -40,6 +40,19 @@ pub struct OsInfo { pub name: String, } +impl OsInfo { + /// Display name for a `std::env::consts::OS` value. + #[must_use] + pub fn pretty_name(os: &str) -> &str { + match os { + "linux" => "Linux", + "windows" => "Windows", + "macos" => "macOS", + other => other, + } + } +} + impl Device { /// Create a Linux sandbox device with the given hostname. #[must_use] @@ -87,6 +100,22 @@ impl Device { Self::linux(hostname) } } + + /// Create the device for a gateway replica. + #[must_use] + pub fn gateway(hostname: &str, name: &str) -> Self { + Self { + hostname: hostname.to_string(), + type_id: DeviceTypeId::Server, + type_label: DeviceTypeId::Server.to_string(), + name: Some(name.to_string()), + // Operators assign a unique name per installation; replicas share this UID. + uid: Some(name.to_string()), + os: Some(OsInfo { + name: OsInfo::pretty_name(std::env::consts::OS).to_string(), + }), + } + } } #[cfg(test)] @@ -140,4 +169,45 @@ mod tests { assert_eq!(decoded, device); assert_eq!(serde_json::to_value(&decoded).unwrap(), json); } + + #[test] + fn os_pretty_names_capitalize_known_platforms() { + for (os, expected) in [ + ("linux", "Linux"), + ("windows", "Windows"), + ("macos", "macOS"), + ("freebsd", "freebsd"), + ] { + assert_eq!(OsInfo::pretty_name(os), expected); + } + } + + #[test] + fn gateway_device_does_not_inherit_the_sandbox_type() { + let json = serde_json::to_value(Device::gateway("gateway-0", "production")).unwrap(); + + assert_eq!(json["type_id"], 1); + assert_eq!(json["type"], "Server"); + } + + #[test] + fn gateway_installations_with_identical_hostnames_have_distinct_uids() { + let first = Device::gateway("openshell-gateway-0", "production"); + let second = Device::gateway("openshell-gateway-0", "staging"); + + assert_eq!(first.uid.as_deref(), Some("production")); + assert_eq!(second.uid.as_deref(), Some("staging")); + assert_ne!(first.uid, second.uid); + assert_eq!(first.hostname, second.hostname); + } + + #[test] + fn gateway_replicas_share_the_installation_uid() { + let first = Device::gateway("openshell-gateway-0", "production"); + let second = Device::gateway("openshell-gateway-1", "production"); + + assert_eq!(first.uid.as_deref(), Some("production")); + assert_eq!(first.uid, second.uid); + assert_ne!(first.hostname, second.hostname); + } } diff --git a/crates/openshell-ocsf/src/objects/metadata.rs b/crates/openshell-ocsf/src/objects/metadata.rs index 14f5f8af61..c580d0d840 100644 --- a/crates/openshell-ocsf/src/objects/metadata.rs +++ b/crates/openshell-ocsf/src/objects/metadata.rs @@ -51,6 +51,16 @@ impl Product { version: Some(version.to_string()), } } + + /// Create the `OpenShell` Gateway product, for control-plane events. + #[must_use] + pub fn openshell_gateway(version: &str) -> Self { + Self { + name: "OpenShell Gateway".to_string(), + vendor_name: "OpenShell".to_string(), + version: Some(version.to_string()), + } + } } #[cfg(test)] diff --git a/crates/openshell-ocsf/tests/event_identity.rs b/crates/openshell-ocsf/tests/event_identity.rs index 6b5ad32e05..f876d3c2ee 100644 --- a/crates/openshell-ocsf/tests/event_identity.rs +++ b/crates/openshell-ocsf/tests/event_identity.rs @@ -6,7 +6,7 @@ use std::net::{IpAddr, Ipv4Addr}; use openshell_ocsf::{ - ActivityId, Endpoint, EventContext, NetworkActivityBuilder, OcsfEvent, SeverityId, + ActivityId, Endpoint, EventContext, EventOrigin, NetworkActivityBuilder, OcsfEvent, SeverityId, }; fn sandbox_ctx(container_image: &str) -> EventContext { @@ -18,6 +18,7 @@ fn sandbox_ctx(container_image: &str) -> EventContext { product_version: "0.42.1".to_string(), proxy_ip: IpAddr::V4(Ipv4Addr::LOCALHOST), proxy_port: 8888, + origin: EventOrigin::Supervisor, } } diff --git a/crates/openshell-ocsf/tests/gateway_context.rs b/crates/openshell-ocsf/tests/gateway_context.rs new file mode 100644 index 0000000000..f96f4f3cef --- /dev/null +++ b/crates/openshell-ocsf/tests/gateway_context.rs @@ -0,0 +1,148 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +//! Gateway-origin events carry gateway identity, not sandbox identity. + +use std::net::{IpAddr, Ipv4Addr}; + +use openshell_ocsf::{ + ActivityId, AppLifecycleBuilder, ConfigStateChangeBuilder, EventContext, EventOrigin, OsInfo, + SeverityId, StateId, StatusId, +}; + +fn gateway_ctx() -> EventContext { + EventContext { + sandbox_id: String::new(), + sandbox_name: String::new(), + container_image: String::new(), + hostname: "openshell-gateway-0".to_string(), + product_version: "0.42.1".to_string(), + proxy_ip: IpAddr::V4(Ipv4Addr::LOCALHOST), + proxy_port: 0, + origin: EventOrigin::Gateway { + name: "production-us-west".to_string(), + }, + } +} + +fn supervisor_ctx() -> EventContext { + EventContext { + sandbox_id: "sb-1".to_string(), + sandbox_name: "agent-01".to_string(), + container_image: "ghcr.io/nvidia/openshell/sandbox:0.42.1".to_string(), + hostname: "openshell-sb-1".to_string(), + product_version: "0.42.1".to_string(), + proxy_ip: IpAddr::V4(Ipv4Addr::LOCALHOST), + proxy_port: 8888, + origin: EventOrigin::Supervisor, + } +} + +#[test] +fn gateway_events_report_the_gateway_os() { + let expected_os = OsInfo::pretty_name(std::env::consts::OS); + + for sandbox_id in ["", "sb-1"] { + let mut context = gateway_ctx(); + context.sandbox_id = sandbox_id.to_string(); + let json = AppLifecycleBuilder::new(&context) + .activity(ActivityId::Open) + .build() + .to_json() + .unwrap(); + + assert_eq!(json["device"]["os"]["name"], expected_os); + } +} + +#[test] +fn supervisor_events_preserve_platform_os() { + let json = AppLifecycleBuilder::new(&supervisor_ctx()) + .activity(ActivityId::Open) + .build() + .to_json() + .unwrap(); + + let expected_os = if cfg!(target_os = "windows") { + "Windows" + } else { + "Linux" + }; + assert_eq!(json["device"]["os"]["name"], expected_os); +} + +#[test] +fn gateway_events_report_the_gateway_product() { + let event = AppLifecycleBuilder::new(&gateway_ctx()) + .activity(ActivityId::Open) + .severity(SeverityId::Informational) + .message("gateway started") + .build(); + let json = event.to_json().unwrap(); + + assert_eq!(json["metadata"]["product"]["name"], "OpenShell Gateway"); + assert_eq!(json["metadata"]["product"]["vendor_name"], "OpenShell"); + assert_eq!(json["metadata"]["version"], "1.8.0"); + assert_eq!(json["metadata"]["product"]["version"], "0.42.1"); +} + +#[test] +fn supervisor_events_still_report_the_supervisor_product() { + let event = AppLifecycleBuilder::new(&supervisor_ctx()) + .activity(ActivityId::Open) + .severity(SeverityId::Informational) + .message("supervisor started") + .build(); + let json = event.to_json().unwrap(); + + assert_eq!( + json["metadata"]["product"]["name"], + "OpenShell Sandbox Supervisor" + ); +} + +#[test] +fn gateway_events_identify_the_device_by_operator_assigned_name() { + let event = ConfigStateChangeBuilder::new(&gateway_ctx()) + .state(StateId::Enabled, "reloaded") + .severity(SeverityId::Informational) + .status(StatusId::Success) + .message("TLS certificate config reloaded") + .build(); + let json = event.to_json().unwrap(); + + assert_eq!(json["device"]["name"], "production-us-west"); + assert_eq!(json["device"]["uid"], "production-us-west"); + assert_eq!(json["device"]["hostname"], "openshell-gateway-0"); +} + +#[test] +fn gateway_events_omit_the_container_object() { + let event = ConfigStateChangeBuilder::new(&gateway_ctx()) + .state(StateId::Enabled, "reloaded") + .severity(SeverityId::Informational) + .status(StatusId::Success) + .message("TLS certificate config reloaded") + .build(); + let json = event.to_json().unwrap(); + + assert!( + json.get("container").is_none(), + "a gateway event without a sandbox association should omit container: {json}" + ); +} + +#[test] +fn supervisor_events_still_carry_their_sandbox_container() { + let event = ConfigStateChangeBuilder::new(&supervisor_ctx()) + .state(StateId::Enabled, "loaded") + .severity(SeverityId::Informational) + .status(StatusId::Success) + .message("policy loaded") + .build(); + let json = event.to_json().unwrap(); + + assert_eq!(json["container"]["name"], "agent-01"); + assert_eq!(json["container"]["uid"], "sb-1"); + assert!(json["device"].get("name").is_none()); +} diff --git a/crates/openshell-ocsf/tests/roundtrip.rs b/crates/openshell-ocsf/tests/roundtrip.rs index 4617cbfe21..bfc651be25 100644 --- a/crates/openshell-ocsf/tests/roundtrip.rs +++ b/crates/openshell-ocsf/tests/roundtrip.rs @@ -12,7 +12,7 @@ use std::net::{IpAddr, Ipv4Addr}; use openshell_ocsf::{ ActionId, ActivityId, AiModel, ApiActivityBuilder, AppLifecycleBuilder, Attack, AuthTypeId, BaseEventBuilder, ConfidenceId, ConfigStateChangeBuilder, ConnectionInfo, - DetectionFindingBuilder, DispositionId, Endpoint, EventContext, FindingInfo, + DetectionFindingBuilder, DispositionId, Endpoint, EventContext, EventOrigin, FindingInfo, HttpActivityBuilder, HttpMethod, HttpRequest, HttpResponse, LaunchTypeId, NetworkActivityBuilder, OcsfEvent, Process, ProcessActivityBuilder, RiskLevelId, SecurityLevelId, SeverityId, SshActivityBuilder, StateId, StatusId, Url, @@ -27,6 +27,7 @@ fn ctx() -> EventContext { product_version: "0.42.1".to_string(), proxy_ip: IpAddr::V4(Ipv4Addr::LOCALHOST), proxy_port: 8888, + origin: EventOrigin::Supervisor, } } diff --git a/crates/openshell-policy-schema/Cargo.toml b/crates/openshell-policy-schema/Cargo.toml index 24808d396b..6352090c26 100644 --- a/crates/openshell-policy-schema/Cargo.toml +++ b/crates/openshell-policy-schema/Cargo.toml @@ -15,6 +15,7 @@ miette = { workspace = true } serde = { workspace = true } serde_json = { workspace = true } serde_yml = { workspace = true } +thiserror = { workspace = true } [lints] workspace = true diff --git a/crates/openshell-policy-schema/src/lib.rs b/crates/openshell-policy-schema/src/lib.rs index 62c2b06dfe..0adef3ac40 100644 --- a/crates/openshell-policy-schema/src/lib.rs +++ b/crates/openshell-policy-schema/src/lib.rs @@ -7,7 +7,7 @@ //! pure schema validation, and lexical policy-path normalization. Runtime and //! protobuf adaptation intentionally live in `openshell-policy`. -use std::collections::BTreeMap; +use std::collections::{BTreeMap, BTreeSet}; use std::fmt; use std::fs::File; use std::io::Read; @@ -17,16 +17,24 @@ use std::str::FromStr; use miette::{IntoDiagnostic, Result, WrapErr}; use serde::{Deserialize, Deserializer, Serialize}; -/// Fixed batch-member bound for the MCP 2025-03-26 wire profile. +/// Fixed resource bound for legacy MCP request batches inspected by Tower. pub const MAX_MCP_LEGACY_BATCH_MESSAGES: usize = 64; /// Stable MCP protocol revisions accepted in authored policy. +/// +/// This closed vocabulary pins product support and ordering independently of +/// the protocol inspector. Wire semantics remain owned by the inspector. #[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Hash)] #[non_exhaustive] pub enum McpProtocolVersion { + /// MCP protocol revision `2025-03-26`. V2025_03_26, + /// MCP protocol revision `2025-06-18`. V2025_06_18, + /// MCP protocol revision `2025-11-25`. V2025_11_25, + /// Sessionless MCP protocol revision `2026-07-28`. + V2026_07_28, } /// Pinned revision used when authored policy omits MCP versions. @@ -35,30 +43,22 @@ pub const DEFAULT_MCP_PROTOCOL_VERSION: McpProtocolVersion = McpProtocolVersion: pub const MCP_VERSION_REMEDIATION: &str = "omit mcp.versions to use the pinned default revision, use an exact supported revision, or omit protocol and mcp for deliberate uninspected L4 passthrough only when that weaker boundary is acceptable"; impl McpProtocolVersion { - pub const ALL: &'static [Self] = &[Self::V2025_03_26, Self::V2025_06_18, Self::V2025_11_25]; + /// Every supported policy revision in canonical semantic order. + pub const ALL: &'static [Self] = &[ + Self::V2025_03_26, + Self::V2025_06_18, + Self::V2025_11_25, + Self::V2026_07_28, + ]; + /// Return the exact MCP protocol identifier accepted in policy. #[must_use] pub const fn as_str(self) -> &'static str { match self { Self::V2025_03_26 => "2025-03-26", Self::V2025_06_18 => "2025-06-18", Self::V2025_11_25 => "2025-11-25", - } - } - - #[must_use] - pub const fn wire_profile(self) -> McpWireProfile { - match self { - Self::V2025_03_26 => McpWireProfile { - version: self, - allows_json_rpc_batches: true, - max_batch_messages: Some(MAX_MCP_LEGACY_BATCH_MESSAGES), - }, - Self::V2025_06_18 | Self::V2025_11_25 => McpWireProfile { - version: self, - allows_json_rpc_batches: false, - max_batch_messages: None, - }, + Self::V2026_07_28 => "2026-07-28", } } } @@ -77,6 +77,7 @@ impl FromStr for McpProtocolVersion { "2025-03-26" => Ok(Self::V2025_03_26), "2025-06-18" => Ok(Self::V2025_06_18), "2025-11-25" => Ok(Self::V2025_11_25), + "2026-07-28" => Ok(Self::V2026_07_28), _ => Err(ParseMcpProtocolVersionError { value: value.to_owned(), }), @@ -84,12 +85,14 @@ impl FromStr for McpProtocolVersion { } } +/// Error returned for a revision outside the exact policy vocabulary. #[derive(Debug, Clone, PartialEq, Eq)] pub struct ParseMcpProtocolVersionError { value: String, } impl ParseMcpProtocolVersionError { + /// Return the original rejected identifier without normalization. #[must_use] pub fn value(&self) -> &str { &self.value @@ -108,29 +111,64 @@ impl fmt::Display for ParseMcpProtocolVersionError { impl std::error::Error for ParseMcpProtocolVersionError {} -/// Immutable batch-shape metadata for an exact MCP revision. -#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)] -pub struct McpWireProfile { - version: McpProtocolVersion, - allows_json_rpc_batches: bool, - max_batch_messages: Option, -} - -impl McpWireProfile { - #[must_use] - pub const fn version(self) -> McpProtocolVersion { - self.version +/// Sort MCP policy revisions without hiding invalid input. +/// +/// Supported revisions use [`McpProtocolVersion::ALL`] semantic order, followed +/// by unsupported identifiers in lexical order. Duplicate values and the exact +/// spelling of every identifier remain available to subsequent validation. +/// Empty lists remain empty; the caller owns omission and default handling. +pub fn canonicalize_mcp_versions(versions: &mut [String]) { + versions.sort_by(|left, right| { + match ( + left.parse::(), + right.parse::(), + ) { + (Ok(left), Ok(right)) => left.cmp(&right), + (Ok(_), Err(_)) => std::cmp::Ordering::Less, + (Err(_), Ok(_)) => std::cmp::Ordering::Greater, + (Err(_), Err(_)) => left.cmp(right), + } + }); +} + +/// Parse an explicit MCP policy allowlist into canonical semantic order. +/// +/// This does not choose a default or modify the input. Callers must handle +/// omitted fields before passing an explicit list to this function. +/// +/// # Errors +/// +/// Returns [`ParseMcpVersionsError::Empty`] for an empty list, or the first +/// unsupported or duplicate revision in input order. +pub fn parse_mcp_versions( + values: &[String], +) -> std::result::Result, ParseMcpVersionsError> { + if values.is_empty() { + return Err(ParseMcpVersionsError::Empty); } - #[must_use] - pub const fn allows_json_rpc_batches(self) -> bool { - self.allows_json_rpc_batches + let mut versions = BTreeSet::new(); + for value in values { + let version = value.parse::()?; + if !versions.insert(version) { + return Err(ParseMcpVersionsError::Duplicate(version)); + } } + Ok(versions) +} - #[must_use] - pub const fn max_batch_messages(self) -> Option { - self.max_batch_messages - } +/// Error returned when an explicit MCP policy allowlist is invalid. +#[derive(Debug, Clone, PartialEq, Eq, thiserror::Error)] +pub enum ParseMcpVersionsError { + /// The explicitly supplied allowlist contains no revisions. + #[error("mcp.versions must contain at least one supported protocol version")] + Empty, + /// An identifier does not exactly match a supported revision. + #[error(transparent)] + Unsupported(#[from] ParseMcpProtocolVersionError), + /// A supported revision occurs more than once in the list. + #[error("duplicate MCP protocol version '{0}'")] + Duplicate(McpProtocolVersion), } #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] @@ -960,21 +998,19 @@ pub fn validate_mcp_config(config: &McpConfig, context: &str) -> Result<()> { let Some(versions) = config.versions.as_deref() else { return Ok(()); }; - if versions.is_empty() { - miette::bail!( - "{context} has an empty mcp.versions list; omit it to use the pinned default revision" - ); - } - let mut seen = std::collections::BTreeSet::new(); - for value in versions { - let version = value - .parse::() - .map_err(|error| miette::miette!("{context}: {error}; {MCP_VERSION_REMEDIATION}"))?; - if !seen.insert(version) { - miette::bail!("{context} has duplicate protocol version '{value}'"); - } - } - Ok(()) + parse_mcp_versions(versions) + .map(|_| ()) + .map_err(|error| match error { + ParseMcpVersionsError::Empty => miette::miette!( + "{context} has an empty mcp.versions list; omit it to use the pinned default revision" + ), + ParseMcpVersionsError::Unsupported(error) => { + miette::miette!("{context}: {error}; {MCP_VERSION_REMEDIATION}") + } + ParseMcpVersionsError::Duplicate(version) => { + miette::miette!("{context} has duplicate protocol version '{version}'") + } + }) } fn reject_json_unknown_fields( @@ -1385,6 +1421,37 @@ network_policies: assert!(parse_policy_bytes(&[0xff]).is_err()); } + #[test] + fn mcp_config_preserves_explicit_sessionless_versions_and_omission() { + let config = parse_mcp_config(serde_json::json!({ + "versions": ["2026-07-28", "2025-03-26"], + })) + .expect("explicit supported revisions"); + assert_eq!( + config.versions.as_deref(), + Some(["2026-07-28".to_string(), "2025-03-26".to_string()].as_slice()) + ); + assert!( + parse_mcp_config(serde_json::json!({})) + .unwrap() + .versions + .is_none() + ); + assert_eq!(DEFAULT_MCP_PROTOCOL_VERSION.as_str(), "2025-11-25"); + } + + #[test] + fn mcp_config_rejects_empty_duplicate_and_unknown_versions() { + for versions in [ + serde_json::json!([]), + serde_json::json!(["2026-07-28", "2026-07-28"]), + serde_json::json!(["2026-07-29"]), + serde_json::json!([" 2026-07-28"]), + ] { + assert!(parse_mcp_config(serde_json::json!({ "versions": versions })).is_err()); + } + } + #[test] fn access_presets_expand_consistently() { assert_eq!( diff --git a/crates/openshell-policy/src/lib.rs b/crates/openshell-policy/src/lib.rs index f6ba71fed7..6d924b1f1f 100644 --- a/crates/openshell-policy/src/lib.rs +++ b/crates/openshell-policy/src/lib.rs @@ -25,7 +25,9 @@ pub use ambiguity::{EndpointAmbiguity, find_endpoint_ambiguities}; use hickory_proto::rr::Name; use miette::{IntoDiagnostic, Result, WrapErr}; -use openshell_core::mcp::{DEFAULT_MCP_PROTOCOL_VERSION, McpProtocolVersion}; +use openshell_core::mcp::{ + DEFAULT_MCP_PROTOCOL_VERSION, McpProtocolVersion, canonicalize_mcp_versions, +}; use openshell_core::proto::{ FilesystemPolicy, GraphqlOperation, L7Allow, L7DenyRule, L7QueryMatcher, L7Rule, LandlockPolicy, McpOptions, NetworkBinary, NetworkEndpoint, NetworkPolicyRule, ProcessPolicy, @@ -360,22 +362,6 @@ fn default_mcp_versions() -> Vec { vec![DEFAULT_MCP_PROTOCOL_VERSION.as_str().to_string()] } -fn canonicalize_mcp_versions(versions: &mut [String]) { - // Unknown and duplicate values remain present so canonicalization cannot - // erase evidence that the raw policy was invalid. - versions.sort_by(|left, right| { - match ( - left.parse::(), - right.parse::(), - ) { - (Ok(left), Ok(right)) => left.cmp(&right), - (Ok(_), Err(_)) => std::cmp::Ordering::Less, - (Err(_), Ok(_)) => std::cmp::Ordering::Greater, - (Err(_), Err(_)) => left.cmp(right), - } - }); -} - /// Sort one protobuf MCP contract without hiding invalid input. /// /// Exact supported revisions use semantic catalog order. Duplicate and @@ -1362,6 +1348,61 @@ pub fn validate_sandbox_policy( validate_sandbox_policy_with_mcp_presence(policy, McpVersionPresence::RequireMaterialized) } +/// Validate filesystem paths shared by typed policies and raw OPA data. +/// +/// Paths must be absolute, contain no parent traversal, and fit the path count +/// and byte limits. Read-write access to the filesystem root is forbidden. +/// Callers that expose errors outside trusted authoring tools must redact the +/// path values carried by the returned violations. +pub fn validate_filesystem_paths( + read_only: &[String], + read_write: &[String], +) -> std::result::Result<(), Vec> { + let mut violations = Vec::new(); + let total_paths = read_only.len() + read_write.len(); + if total_paths > MAX_FILESYSTEM_PATHS { + violations.push(PolicyViolation::TooManyPaths { count: total_paths }); + } + + for path_str in read_only.iter().chain(read_write.iter()) { + if path_str.len() > MAX_PATH_LENGTH { + violations.push(PolicyViolation::FieldTooLong { + path: truncate_for_display(path_str), + length: path_str.len(), + }); + continue; + } + let path = Path::new(path_str); + if !path.has_root() { + violations.push(PolicyViolation::RelativePath { + path: path_str.clone(), + }); + } + if path + .components() + .any(|c| matches!(c, std::path::Component::ParentDir)) + { + violations.push(PolicyViolation::PathTraversal { + path: path_str.clone(), + }); + } + } + + for path_str in read_write { + // Repeated separators still designate the filesystem root. + if path_str.trim_end_matches('/').is_empty() { + violations.push(PolicyViolation::OverlyBroadPath { + path: path_str.clone(), + }); + } + } + if violations.is_empty() { + Ok(()) + } else { + Err(violations) + } +} + #[derive(Debug, Clone, Copy, PartialEq, Eq)] enum McpVersionPresence { RequireMaterialized, @@ -1403,50 +1444,10 @@ fn validate_sandbox_policy_with_mcp_presence( }); } - // Check filesystem paths - if let Some(ref fs) = policy.filesystem { - let total_paths = fs.read_only.len() + fs.read_write.len(); - if total_paths > MAX_FILESYSTEM_PATHS { - violations.push(PolicyViolation::TooManyPaths { count: total_paths }); - } - - for path_str in fs.read_only.iter().chain(fs.read_write.iter()) { - if path_str.len() > MAX_PATH_LENGTH { - violations.push(PolicyViolation::FieldTooLong { - path: truncate_for_display(path_str), - length: path_str.len(), - }); - continue; - } - - let path = Path::new(path_str); - - if !path.has_root() { - violations.push(PolicyViolation::RelativePath { - path: path_str.clone(), - }); - } - - if path - .components() - .any(|c| matches!(c, std::path::Component::ParentDir)) - { - violations.push(PolicyViolation::PathTraversal { - path: path_str.clone(), - }); - } - } - - // Only reject "/" as read-write (overly broad) - for path_str in &fs.read_write { - let normalized = path_str.trim_end_matches('/'); - if normalized.is_empty() { - // Path is "/" or "///" etc. - violations.push(PolicyViolation::OverlyBroadPath { - path: path_str.clone(), - }); - } - } + if let Some(ref fs) = policy.filesystem + && let Err(errors) = validate_filesystem_paths(&fs.read_only, &fs.read_write) + { + violations.extend(errors); } // Protobuf maps do not preserve iteration order. Sort rule keys so callers @@ -2508,7 +2509,7 @@ network_policies: } } - const MCP_VERSIONS: [&str; 3] = ["2025-03-26", "2025-06-18", "2025-11-25"]; + const MCP_VERSIONS: [&str; 4] = ["2025-03-26", "2025-06-18", "2025-11-25", "2026-07-28"]; fn mcp_version_options(versions: &[&str]) -> McpOptions { McpOptions { @@ -2609,6 +2610,30 @@ network_policies: assert!(McpProtocolVersion::ALL.len() > default_mcp_versions().len()); } + #[test] + fn sessionless_mcp_version_requires_an_explicit_policy_opt_in() { + let mut authored = + mcp_version_endpoint_yaml("mcp", Some(" versions: [\"2026-07-28\"]\n")); + authored.push_str(" rules:\n - allow:\n method: tools/list\n"); + let explicit = parse_sandbox_policy(&authored) + .expect("the sessionless revision must be accepted when explicitly allowed"); + let options = explicit.network_policies["versioned"].endpoints[0] + .mcp + .as_ref() + .expect("MCP options must be materialized"); + assert_eq!(options.versions, ["2026-07-28"]); + validate_sandbox_policy(&explicit) + .expect("a sessionless endpoint with an explicit method rule must validate"); + + let yaml = serialize_sandbox_policy(&explicit) + .expect("an explicitly allowed sessionless revision must serialize"); + assert_eq!( + parse_sandbox_policy(&yaml).expect("the sessionless policy must round-trip"), + explicit + ); + assert!(!default_mcp_versions().contains(&"2026-07-28".to_string())); + } + #[test] fn mcp_version_yaml_rejects_explicit_empty_duplicate_unknown_and_misplaced_values() { let cases = [ @@ -2778,7 +2803,7 @@ network_policies: "rendered diagnostic omitted remediation choices: {authored_diagnostic}" ); - let policy = mcp_version_policy("mcp", Some(mcp_version_options(&["2026-07-28"]))); + let policy = mcp_version_policy("mcp", Some(mcp_version_options(&["2026-07-29"]))); let violations = validate_sandbox_policy(&policy) .expect_err("unsupported protobuf revisions must fail closed"); assert!(violations.iter().map(ToString::to_string).any(|message| { @@ -2875,6 +2900,7 @@ network_policies: let policy = mcp_version_policy( "mcp", Some(mcp_version_options(&[ + "2026-07-28", "2025-11-25", "2025-03-26", "2025-06-18", diff --git a/crates/openshell-policy/src/merge.rs b/crates/openshell-policy/src/merge.rs index 2f81b18f3d..ebd147db2a 100644 --- a/crates/openshell-policy/src/merge.rs +++ b/crates/openshell-policy/src/merge.rs @@ -3119,6 +3119,8 @@ mod tests { for versions in [ &["2025-03-26"][..], &["2025-03-26", DEFAULT_MCP_VERSION][..], + &["2026-07-28"][..], + &[DEFAULT_MCP_VERSION, "2026-07-28"][..], ] { let existing = rule_with_authorizations( "existing", @@ -3164,6 +3166,7 @@ mod tests { let existing = rule_with_authorizations( "existing", vec![mcp_endpoint_with_versions(&[ + "2026-07-28", "2025-11-25", "2025-03-26", "2025-06-18", @@ -3175,6 +3178,7 @@ mod tests { vec![mcp_endpoint_with_versions(&[ "2025-06-18", "2025-11-25", + "2026-07-28", "2025-03-26", ])], &["/usr/bin/new"], @@ -3197,7 +3201,7 @@ mod tests { .as_ref() .expect("MCP endpoint must retain options") .versions, - ["2025-03-26", "2025-06-18", "2025-11-25"] + ["2025-03-26", "2025-06-18", "2025-11-25", "2026-07-28"] ); } @@ -3208,6 +3212,7 @@ mod tests { rule_with_authorizations( "mcp", vec![mcp_endpoint_with_versions(&[ + "2026-07-28", "2025-11-25", "2025-03-26", "2025-06-18", @@ -3225,7 +3230,7 @@ mod tests { .as_ref() .expect("MCP endpoint must retain options") .versions, - ["2025-03-26", "2025-06-18", "2025-11-25"] + ["2025-03-26", "2025-06-18", "2025-11-25", "2026-07-28"] ); } @@ -3237,6 +3242,7 @@ mod tests { "2025-03-26", "2025-06-18", "2025-11-25", + "2026-07-28", ])], &["/usr/bin/client"], ); @@ -3251,6 +3257,7 @@ mod tests { let equivalent = rule_with_authorizations( "proposed", vec![mcp_endpoint_with_versions(&[ + "2026-07-28", "2025-11-25", "2025-03-26", "2025-06-18", diff --git a/crates/openshell-policy/testdata/mcp-version-profiles.yaml b/crates/openshell-policy/testdata/mcp-version-profiles.yaml index 1d6106029d..6860d865ba 100644 --- a/crates/openshell-policy/testdata/mcp-version-profiles.yaml +++ b/crates/openshell-policy/testdata/mcp-version-profiles.yaml @@ -13,6 +13,7 @@ network_policies: enforcement: enforce mcp: versions: + - "2026-07-28" - "2025-11-25" - "2025-03-26" - "2025-06-18" diff --git a/crates/openshell-providers/src/discovery.rs b/crates/openshell-providers/src/discovery.rs index 00e8da1890..2249adaf4a 100644 --- a/crates/openshell-providers/src/discovery.rs +++ b/crates/openshell-providers/src/discovery.rs @@ -92,6 +92,7 @@ mod tests { fn profile() -> ProviderTypeProfile { ProviderTypeProfile { + files: Vec::new(), id: "custom".to_string(), resource_version: 0, annotations: std::collections::HashMap::new(), diff --git a/crates/openshell-providers/src/profiles.rs b/crates/openshell-providers/src/profiles.rs index 01991d712e..74d7e44234 100644 --- a/crates/openshell-providers/src/profiles.rs +++ b/crates/openshell-providers/src/profiles.rs @@ -3,14 +3,17 @@ //! Declarative provider type profiles. -use openshell_core::mcp::{DEFAULT_MCP_PROTOCOL_VERSION, McpProtocolVersion}; +use openshell_core::mcp::{ + DEFAULT_MCP_PROTOCOL_VERSION, McpProtocolVersion, ParseMcpVersionsError, + canonicalize_mcp_versions, parse_mcp_versions, +}; use openshell_core::proto::{ GraphqlOperation, L7Allow, L7DenyRule, L7QueryMatcher, L7Rule, McpOptions, NetworkBinary, NetworkEndpoint, NetworkPolicyRule, ProviderCredentialRefresh, ProviderCredentialRefreshMaterial, ProviderCredentialRefreshOutput, ProviderCredentialRefreshStrategy, ProviderCredentialTokenGrantSubjectToken, ProviderCredentialTokenGrantType, ProviderProfile, ProviderProfileCategory, - ProviderProfileCredential, ProviderProfileDiscovery, + ProviderProfileCredential, ProviderProfileDiscovery, ProviderProfileFile, }; use openshell_core::secrets::uses_reserved_revision_namespace; use openshell_policy::{ @@ -499,23 +502,12 @@ fn default_mcp_profile_versions() -> Vec { fn validate_mcp_profile_versions( values: &[String], ) -> Result, String> { - if values.is_empty() { - return Err( - "mcp.versions must contain at least one supported protocol version".to_string(), - ); - } - - let mut versions = BTreeSet::new(); - for value in values { - let version = value - .parse::() - .map_err(|error| format!("{error}; {MCP_VERSION_REMEDIATION}"))?; - if !versions.insert(version) { - return Err(format!("duplicate MCP protocol version '{value}'")); + parse_mcp_versions(values).map_err(|error| match error { + ParseMcpVersionsError::Unsupported(error) => { + format!("{error}; {MCP_VERSION_REMEDIATION}") } - } - - Ok(versions) + error => error.to_string(), + }) } fn deserialize_mcp_profile_versions<'de, D>(deserializer: D) -> Result, D::Error> @@ -705,6 +697,10 @@ pub struct ProviderTypeProfile { pub category: ProviderProfileCategory, #[serde(default)] pub credentials: Vec, + /// EXPERIMENTAL: Non-secret files served to sandbox workloads on demand. + /// This API and its behavior may change or be removed. + #[serde(default, skip_serializing_if = "Vec::is_empty")] + pub files: Vec, #[serde(default)] pub endpoints: Vec, #[serde(default)] @@ -719,6 +715,81 @@ pub struct ProviderTypeProfile { pub scope: String, } +/// EXPERIMENTAL: A non-secret provider file template. This API may change or be removed. +#[derive(Debug, Clone, Deserialize, Serialize, PartialEq, Eq)] +pub struct ProviderFileProfile { + pub path: String, + pub content: String, + #[serde(default, skip_serializing_if = "String::is_empty")] + pub env_var: String, +} + +impl ProviderFileProfile { + fn validate_template(&self) -> Result<(), String> { + let mut remaining = self.content.as_str(); + while let Some(start) = remaining.find("{{") { + if remaining[..start].contains("}}") { + return Err("unexpected provider file placeholder close".to_string()); + } + let after = &remaining[start + 2..]; + let end = after + .find("}}") + .ok_or("unclosed provider file placeholder")?; + let key = after[..end] + .strip_prefix("config.") + .ok_or("provider files may reference only config.KEY")?; + if !valid_file_config_key(key) { + return Err("invalid provider file config key".to_string()); + } + remaining = &after[end + 2..]; + } + if remaining.contains("}}") { + return Err("unexpected provider file placeholder close".to_string()); + } + Ok(()) + } + + /// Render only explicitly referenced non-secret provider config values. + pub fn render(&self, config: &HashMap) -> Result { + self.validate_template()?; + let mut remaining = self.content.as_str(); + let mut rendered = String::new(); + while let Some(start) = remaining.find("{{") { + rendered.push_str(&remaining[..start]); + let after = &remaining[start + 2..]; + let end = after + .find("}}") + .ok_or("unclosed provider file placeholder")?; + let key = after[..end] + .strip_prefix("config.") + .ok_or("provider files may reference only config.KEY")?; + if !valid_file_config_key(key) { + return Err("invalid provider file config key".to_string()); + } + let value = config + .get(key) + .ok_or_else(|| format!("missing provider config key '{key}'"))?; + rendered.push_str(value); + remaining = &after[end + 2..]; + if rendered.len() > 65_536 { + return Err("rendered provider file exceeds 64 KiB".to_string()); + } + } + rendered.push_str(remaining); + if rendered.len() > 65_536 { + return Err("rendered provider file exceeds 64 KiB".to_string()); + } + Ok(rendered) + } +} + +fn valid_file_config_key(key: &str) -> bool { + !key.is_empty() + && key + .bytes() + .all(|c| c.is_ascii_alphanumeric() || c == b'_' || c == b'-') +} + // Provider profile import/export is expected to be lossless for the network // policy fields exposed by the protobuf API. Do not collapse these DTOs into a // narrower shape; direct gRPC imports and CLI YAML imports must preserve the @@ -753,6 +824,15 @@ impl ProviderTypeProfile { token_grant: credential.token_grant.as_ref().map(token_grant_from_proto), }) .collect(), + files: profile + .files + .iter() + .map(|file| ProviderFileProfile { + path: file.path.clone(), + content: file.content.clone(), + env_var: file.env_var.clone(), + }) + .collect(), endpoints: profile.endpoints.iter().map(endpoint_from_proto).collect(), binaries: profile.binaries.iter().map(binary_from_proto).collect(), inference_capable: profile.inference_capable, @@ -919,6 +999,15 @@ impl ProviderTypeProfile { token_grant: credential.token_grant.as_ref().map(token_grant_to_proto), }) .collect(), + files: self + .files + .iter() + .map(|file| ProviderProfileFile { + path: file.path.clone(), + content: file.content.clone(), + env_var: file.env_var.clone(), + }) + .collect(), endpoints: self.endpoints.iter().map(endpoint_to_proto).collect(), binaries: self.binaries.iter().map(binary_to_proto).collect(), inference_capable: self.inference_capable, @@ -1678,24 +1767,8 @@ fn materialize_and_canonicalize_mcp_profile_versions(versions: &mut Vec) if versions.is_empty() { *versions = default_mcp_profile_versions(); } else { - canonicalize_mcp_profile_versions(versions); - } -} - -fn canonicalize_mcp_profile_versions(versions: &mut [String]) { - // Preserve unsupported and duplicate values so subsequent validation can - // reject them; sorting must never repair malformed protobuf input. - versions.sort_by(|left, right| { - match ( - left.parse::(), - right.parse::(), - ) { - (Ok(left), Ok(right)) => left.cmp(&right), - (Ok(_), Err(_)) => std::cmp::Ordering::Less, - (Err(_), Ok(_)) => std::cmp::Ordering::Greater, - (Err(_), Err(_)) => left.cmp(right), - } - }); + canonicalize_mcp_versions(versions); + } } fn binary_to_proto(binary: &BinaryProfile) -> NetworkBinary { @@ -2061,8 +2134,80 @@ pub fn validate_profile_set( )); } + if profile.files.len() > 16 { + diagnostics.push(ProfileValidationDiagnostic::error( + source, + profile_id, + "files", + "at most 16 provider files are allowed", + )); + } + let mut file_paths = HashSet::new(); + let mut file_env_vars = HashSet::new(); + for file in &profile.files { + if file.path.is_empty() + || file.path == "." + || file.path == ".." + || !file + .path + .bytes() + .all(|c| c.is_ascii_alphanumeric() || matches!(c, b'.' | b'_' | b'-')) + || !file_paths.insert(file.path.as_str()) + { + diagnostics.push(ProfileValidationDiagnostic::error( + source, + profile_id, + "files.path", + "file path must be a unique, safe file name", + )); + } + if file.content.len() > 65_536 { + diagnostics.push(ProfileValidationDiagnostic::error( + source, + profile_id, + "files.content", + "file template exceeds 64 KiB", + )); + } + if let Err(error) = file.validate_template() { + diagnostics.push(ProfileValidationDiagnostic::error( + source, + profile_id, + "files.content", + error, + )); + } + if !file.env_var.is_empty() + && (!file + .env_var + .starts_with(|c: char| c.is_ascii_alphabetic() || c == '_') + || !file + .env_var + .bytes() + .all(|c| c.is_ascii_alphanumeric() || c == b'_') + || !file_env_vars.insert(file.env_var.as_str())) + { + diagnostics.push(ProfileValidationDiagnostic::error( + source, + profile_id, + "files.env_var", + "file environment key must be unique and use letters, digits, or underscores", + )); + } + } + let mut credential_names = HashSet::new(); for credential in &profile.credentials { + for env_var in &credential.env_vars { + if file_env_vars.contains(env_var.as_str()) { + diagnostics.push(ProfileValidationDiagnostic::error( + source, + profile_id, + "files.env_var", + format!("file environment key '{env_var}' conflicts with a credential"), + )); + } + } let credential_name = credential.name.trim(); if credential_name.is_empty() { diagnostics.push(ProfileValidationDiagnostic::error( @@ -4100,7 +4245,7 @@ endpoints: path: /mcp protocol: mcp mcp: - versions: ["2025-11-25", "2025-03-26", "2025-06-18"] + versions: ["2026-07-28", "2025-11-25", "2025-03-26", "2025-06-18"] strict_tool_names: false binaries: - /usr/bin/example-agent @@ -4108,7 +4253,7 @@ binaries: ) .expect("profile should parse"); - let expected_versions = ["2025-03-26", "2025-06-18", "2025-11-25"]; + let expected_versions = ["2025-03-26", "2025-06-18", "2025-11-25", "2026-07-28"]; assert_eq!( profile.endpoints[0] .mcp @@ -4559,7 +4704,7 @@ endpoints: "versions: [\"2025-03-26\", \"2025-03-26\"]", "versions: [latest]", "versions: [draft]", - "versions: ['2026-07-28']", + "versions: ['2026-07-29']", "versions: [\"2025-03-26 \"]", ] { let yaml = format!( @@ -5762,6 +5907,7 @@ binaries: ["", /usr/bin/broken] ( "space.yaml".to_string(), ProviderTypeProfile { + files: Vec::new(), id: " alex-api ".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -5780,6 +5926,7 @@ binaries: ["", /usr/bin/broken] ( "underscore.yaml".to_string(), ProviderTypeProfile { + files: Vec::new(), id: "alex_api".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -5798,6 +5945,7 @@ binaries: ["", /usr/bin/broken] ( "case.yaml".to_string(), ProviderTypeProfile { + files: Vec::new(), id: "Alex-API".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -7819,4 +7967,29 @@ binaries: "expected mcp-options-on-non-mcp rejection: {diagnostics:?}" ); } + + #[test] + fn managed_file_example_validates_and_renders_only_config_fields() { + let yaml = include_str!("../../../examples/provider-managed-files/acme-config.yaml"); + let profile = parse_profile_yaml(yaml).expect("example profile parses"); + let diagnostics = + validate_profile_set(&[("acme-config.yaml".to_string(), profile.clone())]); + assert!(diagnostics.is_empty(), "{diagnostics:?}"); + let rendered = profile.files[0] + .render(&HashMap::from([ + ( + "endpoint".to_string(), + "https://api.acme.example".to_string(), + ), + ("project".to_string(), "production".to_string()), + ])) + .expect("render provider config"); + assert!(rendered.contains("project = \"production\"")); + assert!(!rendered.contains("{{")); + + let mut invalid = profile; + invalid.files[0].content = "token = \"{{credential.API_KEY}}\"".to_string(); + let diagnostics = validate_profile_set(&[("invalid.yaml".to_string(), invalid)]); + assert!(diagnostics.iter().any(|d| d.field == "files.content")); + } } diff --git a/crates/openshell-sandbox-backend/src/boundary_protocol.rs b/crates/openshell-sandbox-backend/src/boundary_protocol.rs index cb83bcf090..6c6c66e76c 100644 --- a/crates/openshell-sandbox-backend/src/boundary_protocol.rs +++ b/crates/openshell-sandbox-backend/src/boundary_protocol.rs @@ -653,10 +653,11 @@ impl RequestEnvelope { } fn request_payload_digest(request: &Request) -> Result { - // Round-tripping through Value canonicalizes every JSON object by key. In - // particular, this makes HashMap-backed provider environments stable - // across process restarts and independently serialized retries. - let normalized = serde_json::to_value(request).map_err(FrameError::Serialize)?; + // Sort every object explicitly: dependency features may make Value retain + // insertion order. Provider environments must hash identically after + // deserialization and across independently serialized retries. + let mut normalized = serde_json::to_value(request).map_err(FrameError::Serialize)?; + normalized.sort_all_objects(); let payload = serde_json::to_vec(&normalized).map_err(FrameError::Serialize)?; let digest = Sha256::digest(payload); Ok(format!("{digest:x}")) @@ -683,6 +684,8 @@ pub enum Request { resource_claims: std::collections::BTreeMap, }, Confirm, + /// Verify file-open mediation before sending a file-bearing snapshot. + ProbeProviderFiles, StartAgent { sandbox_id: String, spec: AgentSpecWire, @@ -691,12 +694,16 @@ pub enum Request { ca_bundle: Option, provider_env_revision: u64, provider_env: std::collections::HashMap, + #[serde(default)] + provider_files: std::collections::HashMap, }, UpdateProviderEnvironment { /// Ordered publication within this authenticated boundary session. generation: u64, revision: u64, provider_env: std::collections::HashMap, + #[serde(default)] + provider_files: std::collections::HashMap, }, AttachProcess { process_id: String, @@ -771,6 +778,7 @@ impl fmt::Debug for Request { .field("resource_claims", resource_claims) .finish(), Self::Confirm => formatter.write_str("Confirm"), + Self::ProbeProviderFiles => formatter.write_str("ProbeProviderFiles"), Self::StartAgent { sandbox_id, spec, @@ -779,6 +787,7 @@ impl fmt::Debug for Request { ca_bundle, provider_env_revision, provider_env, + provider_files, } => formatter .debug_struct("StartAgent") .field("sandbox_id", sandbox_id) @@ -787,6 +796,7 @@ impl fmt::Debug for Request { .field("ca_cert_present", &ca_cert.is_some()) .field("ca_bundle_present", &ca_bundle.is_some()) .field("provider_env_revision", provider_env_revision) + .field("provider_file_count", &provider_files.len()) .field( "provider_env_keys", &provider_env.keys().collect::>(), @@ -796,10 +806,12 @@ impl fmt::Debug for Request { generation, revision, provider_env, + provider_files, } => formatter .debug_struct("UpdateProviderEnvironment") .field("generation", generation) .field("revision", revision) + .field("provider_file_count", &provider_files.len()) .field( "provider_env_keys", &provider_env.keys().collect::>(), @@ -871,6 +883,7 @@ pub enum Response { /// before workload launch. confirmation: Box, }, + ProviderFilesSupported, Started { process_id: String, provider_env_revision: u64, @@ -1562,6 +1575,7 @@ mod tests { request_id: "4e94636d-54f8-4d85-8e4e-58954fb5af0a".to_string(), payload_digest: String::new(), request: Request::StartAgent { + provider_files: std::collections::HashMap::new(), sandbox_id: "sandbox-1".to_string(), spec: AgentSpecWire { program: "/bin/true".to_string(), @@ -1611,14 +1625,44 @@ mod tests { second.insert("A".to_string(), "1".to_string()); second.insert("B".to_string(), "2".to_string()); let build = |provider_env| Request::UpdateProviderEnvironment { + provider_files: std::collections::HashMap::new(), generation: 1, revision: 2, provider_env, }; - assert_eq!( - request_payload_digest(&build(first)).expect("first digest"), - request_payload_digest(&build(second)).expect("second digest") + // Pin the canonical bytes, including the nested environment object. + // Two randomized HashMaps can otherwise happen to iterate identically + // and conceal a serializer that preserves insertion order. + let expected = format!( + "{:x}", + Sha256::digest( + br#"{"generation":1,"operation":"update_provider_environment","provider_env":{"A":"1","B":"2"},"provider_files":{},"revision":2}"# + ) ); + for provider_env in [first, second] { + let request = build(provider_env); + assert_eq!(request_payload_digest(&request).expect("digest"), expected); + + // Deserialization reconstructs the map with an independent hash + // seed; validation must retain the sender's canonical digest. + let envelope = RequestEnvelope::new(request).expect("request envelope"); + let frame = encode_frame(&envelope).expect("encode envelope"); + let mut decoded: RequestEnvelope = decode_frame(&frame).expect("decode envelope"); + assert_eq!(decoded.payload_digest, expected); + decoded + .validate_payload_digest() + .expect("round-trip digest"); + + let Request::UpdateProviderEnvironment { provider_env, .. } = &mut decoded.request + else { + panic!("decoded the wrong request variant"); + }; + provider_env.insert("A".to_string(), "changed".to_string()); + assert!(matches!( + decoded.validate_payload_digest(), + Err(FrameError::PayloadDigestMismatch) + )); + } let mut envelope = RequestEnvelope::new(build(std::collections::HashMap::new())) .expect("request envelope"); @@ -1629,11 +1673,34 @@ mod tests { envelope.validate_payload_digest(), Err(FrameError::PayloadDigestMismatch) )); + + let mut request = build(std::collections::HashMap::new()); + let Request::UpdateProviderEnvironment { provider_files, .. } = &mut request else { + unreachable!(); + }; + provider_files.insert( + "/run/openshell/providers/acme/client.toml".to_string(), + "version = 1".to_string(), + ); + let mut envelope = RequestEnvelope::new(request).expect("file-bearing request envelope"); + let Request::UpdateProviderEnvironment { provider_files, .. } = &mut envelope.request + else { + unreachable!(); + }; + provider_files.insert( + "/run/openshell/providers/acme/client.toml".to_string(), + "version = 2".to_string(), + ); + assert!(matches!( + envelope.validate_payload_digest(), + Err(FrameError::PayloadDigestMismatch) + )); } #[test] fn start_agent_with_large_ca_bundle_fits_in_frame_limit() { let request = RequestEnvelope::new(Request::StartAgent { + provider_files: std::collections::HashMap::new(), sandbox_id: "sandbox-1".to_string(), spec: AgentSpecWire { program: "/bin/true".to_string(), diff --git a/crates/openshell-sandbox-backend/src/runtime.rs b/crates/openshell-sandbox-backend/src/runtime.rs index f073c6d7ce..641d50cd24 100644 --- a/crates/openshell-sandbox-backend/src/runtime.rs +++ b/crates/openshell-sandbox-backend/src/runtime.rs @@ -396,12 +396,26 @@ impl ReadyBoundary for RemoteReady { } else { (None, None) }; - let (provider_env_revision, provider_env) = self + let provider_snapshot = self .provider_credentials - .child_env_snapshot_with_gcp_resolved() + .child_environment_snapshot() .map_err(|error| { BackendError::Process(format!("snapshot provider environment: {error}")) })?; + if !provider_snapshot.files.is_empty() { + let response = self + .client + .call_idempotent(Request::ProbeProviderFiles) + .await + .map_err(|error| { + BackendError::Process(format!("provider file capability probe failed: {error}")) + })?; + if !matches!(response, Response::ProviderFilesSupported) { + return Err(BackendError::Process( + "sandbox boundary does not support provider files".to_string(), + )); + } + } let response = self .client .call_idempotent(Request::StartAgent { @@ -410,8 +424,9 @@ impl ReadyBoundary for RemoteReady { policy: Box::new(SandboxPolicyWire::from(self.policy)), ca_cert, ca_bundle, - provider_env_revision, - provider_env, + provider_env_revision: provider_snapshot.revision, + provider_env: provider_snapshot.environment, + provider_files: provider_snapshot.files, }) .await?; let Response::Started { @@ -621,6 +636,22 @@ impl RemoteExec { .map_err(|error| { BackendError::Process(format!("snapshot provider environment: {error}")) })?; + if !snapshot.files.is_empty() { + let response = self + .client + .call_idempotent(Request::ProbeProviderFiles) + .await + .map_err(|error| { + BackendError::Process(format!( + "provider file capability probe failed: {error}" + )) + })?; + if !matches!(response, Response::ProviderFilesSupported) { + return Err(BackendError::Process( + "sandbox boundary does not support provider files".to_string(), + )); + } + } *generation = generation.checked_add(1).ok_or_else(|| { BackendError::Process("provider environment publication exhausted".to_string()) })?; @@ -631,6 +662,7 @@ impl RemoteExec { generation: requested_generation, revision: snapshot.revision, provider_env: snapshot.environment, + provider_files: snapshot.files, }) .await?; let Response::ProviderEnvironmentUpdated { @@ -2286,6 +2318,7 @@ mod tests { Request::Confirm => Response::Confirmed { confirmation: Box::new(confirmation), }, + Request::ProbeProviderFiles => Response::ProviderFilesSupported, Request::OpenMediation if mediation_ready => Response::MediationReady, Request::OpenMediation => Response::Error { kind: crate::boundary_protocol::BoundaryErrorKind::Denied, @@ -3452,6 +3485,7 @@ mod tests { assert!(matches!( client .exchange(Request::StartAgent { + provider_files: HashMap::new(), sandbox_id: context.sandbox_id, spec: AgentSpecWire::from(context.agent), policy: Box::new(SandboxPolicyWire::from(context.policy)), @@ -3511,6 +3545,7 @@ mod tests { tokio::time::timeout( Duration::from_secs(2), client.exchange(Request::StartAgent { + provider_files: HashMap::new(), sandbox_id: context.sandbox_id, spec: AgentSpecWire::from(context.agent), policy: Box::new(SandboxPolicyWire::from(context.policy)), diff --git a/crates/openshell-sandbox/src/boundary_exec.rs b/crates/openshell-sandbox/src/boundary_exec.rs index 8d0a3326ed..da35294c06 100644 --- a/crates/openshell-sandbox/src/boundary_exec.rs +++ b/crates/openshell-sandbox/src/boundary_exec.rs @@ -688,7 +688,17 @@ mod tests { .expect("start test workload launcher"); std::thread::spawn(move || { while let Ok(notification) = listener.receive() { - let _ = listener.respond_errno(notification.id, libc::EPERM); + let syscall = i64::from(notification.syscall); + if syscall == libc::SYS_openat || syscall == libc::SYS_openat2 { + let _ = listener.respond_continue(notification.id); + } else { + #[cfg(target_arch = "x86_64")] + if syscall == libc::SYS_open { + let _ = listener.respond_continue(notification.id); + continue; + } + let _ = listener.respond_errno(notification.id, libc::EPERM); + } } }); LocalBoundaryExec::new( diff --git a/crates/openshell-sandbox/src/boundary_server.rs b/crates/openshell-sandbox/src/boundary_server.rs index 7a2a73f731..6e7dd23ca3 100644 --- a/crates/openshell-sandbox/src/boundary_server.rs +++ b/crates/openshell-sandbox/src/boundary_server.rs @@ -1420,6 +1420,7 @@ mod linux { ca_bundle: Option, provider_env_revision: u64, provider_env: std::collections::HashMap, + provider_files: std::collections::HashMap, } impl StartedAgent { @@ -1945,6 +1946,10 @@ mod linux { } } Request::Confirm => self.confirm(), + Request::ProbeProviderFiles => self.network_broker.confirm_healthy().map_or_else( + |error| guest_error(BoundaryErrorKind::Unavailable, error.to_string()), + |()| Response::ProviderFilesSupported, + ), Request::StartAgent { sandbox_id, spec, @@ -1953,6 +1958,7 @@ mod linux { ca_bundle, provider_env_revision, provider_env, + provider_files, } => self.start_agent( sandbox_id, spec, @@ -1961,12 +1967,19 @@ mod linux { ca_bundle, provider_env_revision, provider_env, + provider_files, ), Request::UpdateProviderEnvironment { generation, revision, provider_env, - } => self.update_provider_environment(generation, revision, provider_env), + provider_files, + } => self.update_provider_environment( + generation, + revision, + provider_env, + provider_files, + ), Request::Wait { process_id } => self.wait(&process_id), Request::Signal { process_id, signal } => self.signal(&process_id, signal), Request::Terminate { process_id } => self.terminate(&process_id), @@ -2425,6 +2438,7 @@ mod linux { ca_bundle: Option, provider_env_revision: u64, provider_env: std::collections::HashMap, + provider_files: std::collections::HashMap, ) -> Response { let spec = match resolve_agent_spec(spec) { Ok(spec) => spec, @@ -2439,6 +2453,7 @@ mod linux { ca_bundle: ca_bundle.clone(), provider_env_revision, provider_env: provider_env.clone(), + provider_files: provider_files.clone(), }; if let RuntimeState::Running(process) = &*state { return if lock(&self.started_agent) @@ -2498,6 +2513,13 @@ mod linux { provider_env, ca_file_paths, }; + let provider_file_count = provider_files.len(); + if let Err(error) = self.network_broker.provider_files().replace(provider_files) { + return guest_error( + BoundaryErrorKind::Process, + format!("install provider files: {error}"), + ); + } let process = match ManagedProcess::spawn( &self.process_runtime, &self.workload_launcher, @@ -2510,6 +2532,15 @@ mod linux { let process_id = process.process_id(); *lock(&self.started_agent) = Some(requested); *state = RuntimeState::Running(process); + openshell_ocsf::ocsf_emit!( + openshell_ocsf::ConfigStateChangeBuilder::new(openshell_ocsf::ctx::ctx()) + .severity(openshell_ocsf::SeverityId::Informational) + .status(openshell_ocsf::StatusId::Success) + .message(format!( + "Provider file snapshot loaded [file_count:{provider_file_count}]" + )) + .build() + ); Response::Started { process_id, provider_env_revision, @@ -2522,6 +2553,7 @@ mod linux { generation: u64, revision: u64, provider_env: std::collections::HashMap, + provider_files: std::collections::HashMap, ) -> Response { let process = { let state = lock(&self.state); @@ -2548,6 +2580,13 @@ mod linux { applied: false, }; } + if let Err(error) = crate::provider_files::ProviderFiles::validate(&provider_files) { + return guest_error( + BoundaryErrorKind::Invalid, + format!("invalid provider files: {error}"), + ); + } + let requested_revision = revision; let revision = match process .provider_credentials .compare_and_install_child_env_snapshot(current.revision, revision, provider_env) @@ -2555,7 +2594,29 @@ mod linux { Ok(revision) => revision, Err(error) => return guest_error(BoundaryErrorKind::Process, error.to_string()), }; + if revision != requested_revision { + return guest_error( + BoundaryErrorKind::Invalid, + "provider environment changed during update", + ); + } + let provider_file_count = provider_files.len(); + if let Err(error) = self.network_broker.provider_files().replace(provider_files) { + return guest_error( + BoundaryErrorKind::Process, + format!("install provider files: {error}"), + ); + } *installed_generation = generation; + openshell_ocsf::ocsf_emit!( + openshell_ocsf::ConfigStateChangeBuilder::new(openshell_ocsf::ctx::ctx()) + .severity(openshell_ocsf::SeverityId::Informational) + .status(openshell_ocsf::StatusId::Success) + .message(format!( + "Provider file snapshot updated [file_count:{provider_file_count}]" + )) + .build() + ); Response::ProviderEnvironmentUpdated { revision, generation, @@ -4871,6 +4932,7 @@ mod linux { None, 0, std::collections::HashMap::new(), + std::collections::HashMap::new(), ) }; let Response::Started { @@ -4883,6 +4945,7 @@ mod linux { }; let update = RequestEnvelope::new(Request::UpdateProviderEnvironment { + provider_files: std::collections::HashMap::new(), generation: 1, revision: 7, provider_env: std::collections::HashMap::from([( @@ -4968,6 +5031,7 @@ mod linux { None, 0, std::collections::HashMap::new(), + std::collections::HashMap::new(), ), Response::Error { kind, .. } if kind == BoundaryErrorKind::Denied )); @@ -4976,7 +5040,12 @@ mod linux { // fingerprint. Distinct publications must still replace the map, // while a delayed older clear must never undo the repair. assert_eq!( - boundary.update_provider_environment(2, 7, std::collections::HashMap::default()), + boundary.update_provider_environment( + 2, + 7, + std::collections::HashMap::default(), + std::collections::HashMap::new() + ), Response::ProviderEnvironmentUpdated { revision: 7, generation: 2, @@ -4990,7 +5059,8 @@ mod linux { std::collections::HashMap::from([( "REPLAY_TEST".to_string(), "reconnected".to_string() - ),]) + ),]), + std::collections::HashMap::new(), ), Response::ProviderEnvironmentUpdated { revision: 7, @@ -4999,7 +5069,12 @@ mod linux { } ); assert_eq!( - boundary.update_provider_environment(2, 7, std::collections::HashMap::default()), + boundary.update_provider_environment( + 2, + 7, + std::collections::HashMap::default(), + std::collections::HashMap::new() + ), Response::ProviderEnvironmentUpdated { revision: 7, generation: 3, @@ -5165,6 +5240,7 @@ mod linux { *lock(&boundary.state) = RuntimeState::Running(process.clone()); *lock(&boundary.attached_policy) = Some(wire_policy.clone()); *lock(&boundary.started_agent) = Some(StartedAgent { + provider_files: std::collections::HashMap::new(), sandbox_id: "sandbox-retained".to_string(), spec: agent_spec.clone(), policy: wire_policy.clone(), @@ -5190,6 +5266,7 @@ mod linux { None, 0, std::collections::HashMap::new(), + std::collections::HashMap::new(), ), Response::Started { process_id: process.process_id(), @@ -5206,6 +5283,7 @@ mod linux { "ROTATED_TOKEN".to_string(), "refreshed".to_string(), )]), + std::collections::HashMap::new(), ), Response::ProviderEnvironmentUpdated { revision: 2, @@ -5221,6 +5299,7 @@ mod linux { "ROTATED_TOKEN".to_string(), "stale".to_string(), )]), + std::collections::HashMap::new(), ), Response::ProviderEnvironmentUpdated { revision: 2, @@ -5229,7 +5308,12 @@ mod linux { } ); assert_eq!( - boundary.update_provider_environment(2, 1, std::collections::HashMap::new()), + boundary.update_provider_environment( + 2, + 1, + std::collections::HashMap::new(), + std::collections::HashMap::new() + ), Response::ProviderEnvironmentUpdated { revision: 1, generation: 2, @@ -5245,6 +5329,7 @@ mod linux { "ROTATED_TOKEN".to_string(), "out-of-order".to_string(), )]), + std::collections::HashMap::new(), ), Response::ProviderEnvironmentUpdated { revision: 1, @@ -5254,7 +5339,12 @@ mod linux { "a stale publication must not overwrite current state" ); assert_eq!( - boundary.update_provider_environment(1, 1, std::collections::HashMap::new()), + boundary.update_provider_environment( + 1, + 1, + std::collections::HashMap::new(), + std::collections::HashMap::new() + ), Response::ProviderEnvironmentUpdated { revision: 1, generation: 2, @@ -5280,6 +5370,7 @@ mod linux { "ROTATED_TOKEN".to_string(), "replacement-control-snapshot".to_string(), )]), + std::collections::HashMap::new(), ), Response::Started { process_id: process.process_id(), diff --git a/crates/openshell-sandbox/src/lib.rs b/crates/openshell-sandbox/src/lib.rs index 957b55c664..5ba996181a 100644 --- a/crates/openshell-sandbox/src/lib.rs +++ b/crates/openshell-sandbox/src/lib.rs @@ -22,6 +22,8 @@ mod network_broker; pub mod perf; #[cfg(unix)] pub mod process; +#[cfg(target_os = "linux")] +mod provider_files; mod pty; pub mod sandbox; #[cfg(target_os = "linux")] diff --git a/crates/openshell-sandbox/src/network_broker.rs b/crates/openshell-sandbox/src/network_broker.rs index 744ce89677..56c857ce9c 100644 --- a/crates/openshell-sandbox/src/network_broker.rs +++ b/crates/openshell-sandbox/src/network_broker.rs @@ -190,6 +190,7 @@ fn register_dns_socket( #[derive(Clone)] struct NotificationQueues { + provider_files: crate::provider_files::ProviderFiles, protected_control_port: Option, accept_registrar: crate::accept_interrupt::AcceptRegistrar, identity_resolver: ProcfsIdentityResolver, @@ -204,6 +205,7 @@ struct NotificationQueues { /// Live broker handle retained by the sandbox boundary. #[derive(Clone)] pub struct NetworkBroker { + provider_files: crate::provider_files::ProviderFiles, _accept_monitor: Arc, pending: Arc>>, pending_dns: Arc>>, @@ -260,7 +262,9 @@ impl NetworkBroker { let dns_address = dns_relay.address; let retained_socket_capacity = retained_socket_capacity()?; let registry = Arc::new(Mutex::new(SocketRegistry::new(1, SOCKET_CAPACITY)?)); + let provider_files = crate::provider_files::ProviderFiles::default(); let queues = NotificationQueues { + provider_files: provider_files.clone(), protected_control_port, accept_registrar: accept_monitor.registrar(), identity_resolver: ProcfsIdentityResolver::for_pid_namespace(), @@ -311,6 +315,7 @@ impl NetworkBroker { }) .map_err(|error| io::Error::other(format!("start network broker: {error}")))?; Ok(Self { + provider_files, _accept_monitor: accept_monitor, pending: Arc::new(tokio::sync::Mutex::new(pending_rx)), pending_dns: Arc::new(tokio::sync::Mutex::new(pending_dns_rx)), @@ -352,6 +357,10 @@ impl NetworkBroker { )) } } + + pub(crate) fn provider_files(&self) -> &crate::provider_files::ProviderFiles { + &self.provider_files + } } fn start_dns_relay( @@ -528,6 +537,13 @@ fn dispatch_notification( queues: NotificationQueues, ) -> io::Result<()> { let syscall = i64::from(notification.syscall); + if syscall == libc::SYS_openat || syscall == libc::SYS_openat2 { + return queues.provider_files.handle_open(&listener, notification); + } + #[cfg(target_arch = "x86_64")] + if syscall == libc::SYS_open { + return queues.provider_files.handle_open(&listener, notification); + } if matches!(syscall, libc::SYS_kill | libc::SYS_rt_sigqueueinfo) { return openshell_isolation_interface::linux::process_signal::mediate_process_signal( &listener, @@ -1827,6 +1843,87 @@ fn error_to_errno(error: &io::Error) -> i32 { mod tests { use super::*; use openshell_isolation_interface::linux::seccomp_notify::ListenerMode; + + #[test] + fn provider_files_are_opened_on_demand_and_replaced() { + let (launcher, listener) = openshell_isolation_interface::linux::workload_launcher::start() + .expect("start workload launcher"); + let broker = NetworkBroker::start_for_test(listener).expect("start broker"); + let path = "/run/openshell/providers/acme/client.toml".to_string(); + broker + .provider_files() + .replace(HashMap::from([(path.clone(), "version = 1\n".into())])) + .unwrap(); + let first = launcher + .execute({ + let path = path.clone(); + move || std::fs::read_to_string(path) + }) + .unwrap() + .expect("first open"); + assert_eq!(first, "version = 1\n"); + let mut old_descriptor = launcher + .execute({ + let path = path.clone(); + move || std::fs::File::open(path) + }) + .unwrap() + .expect("open old version"); + let denied_write = launcher + .execute({ + let path = path.clone(); + move || std::fs::OpenOptions::new().write(true).open(path) + }) + .unwrap() + .expect_err("provider file is read only"); + assert_eq!(denied_write.raw_os_error(), Some(libc::EACCES)); + broker + .provider_files() + .replace(HashMap::from([(path.clone(), "version = 2\n".into())])) + .unwrap(); + let mut old_content = String::new(); + io::Read::read_to_string(&mut old_descriptor, &mut old_content).unwrap(); + assert_eq!(old_content, "version = 1\n"); + let second = launcher + .execute({ + let path = path.clone(); + move || std::fs::read_to_string(path) + }) + .unwrap() + .expect("second open"); + assert_eq!(second, "version = 2\n"); + let via_openat2 = launcher + .execute({ + let path = path.clone(); + move || -> io::Result { + let path = std::ffi::CString::new(path).unwrap(); + let how = [libc::O_CLOEXEC as u64, 0, 0]; + let fd = unsafe { + libc::syscall( + libc::SYS_openat2, + libc::AT_FDCWD, + path.as_ptr(), + how.as_ptr(), + 24_usize, + ) + }; + if fd < 0 { + return Err(io::Error::last_os_error()); + } + let mut file = unsafe { + std::fs::File::from_raw_fd(i32::try_from(fd).expect("open fd fits")) + }; + let mut content = String::new(); + io::Read::read_to_string(&mut file, &mut content)?; + Ok(content) + } + }) + .unwrap(); + assert_eq!(via_openat2.unwrap(), "version = 2\n"); + broker.provider_files().replace(HashMap::new()).unwrap(); + let detached = launcher.execute(move || std::fs::read(path)).unwrap(); + assert_eq!(detached.unwrap_err().raw_os_error(), Some(libc::ENOENT)); + } use std::io::{Read as _, Write as _}; use std::os::unix::net::{UnixListener, UnixStream}; @@ -2214,10 +2311,6 @@ mod tests { #[test] fn accepted_loopback_stream_is_registered_for_notified_operations() { - let reservation = TcpListener::bind("127.0.0.1:0").expect("reserve loopback port"); - let address = reservation.local_addr().expect("reserved address"); - drop(reservation); - let (launcher, listener) = openshell_isolation_interface::linux::workload_launcher::start() .expect("start workload launcher"); let _broker = NetworkBroker::start_for_test(listener).expect("start network broker"); @@ -2225,9 +2318,9 @@ mod tests { let workload = std::thread::spawn(move || { launcher .execute(move || -> io::Result { - let listener = TcpListener::bind(address)?; + let listener = TcpListener::bind("127.0.0.1:0")?; ready_tx - .send(()) + .send(listener.local_addr()?) .map_err(|_| io::Error::other("test client disappeared"))?; let (stream, _) = listener.accept()?; let peer = stream.peer_addr()?; @@ -2256,7 +2349,13 @@ mod tests { .expect("launcher result") }); - ready_rx.recv().expect("listener ready"); + let Ok(address) = ready_rx.recv() else { + let error = workload + .join() + .expect("join workload") + .expect_err("listener not ready"); + panic!("workload listener failed: {error}"); + }; let mut client = TcpStream::connect(address).expect("connect loopback client"); let mut payload = [0_u8; 8]; client diff --git a/crates/openshell-sandbox/src/provider_files.rs b/crates/openshell-sandbox/src/provider_files.rs new file mode 100644 index 0000000000..5efd4270f2 --- /dev/null +++ b/crates/openshell-sandbox/src/provider_files.rs @@ -0,0 +1,259 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +//! Read-only provider files served on demand through seccomp FD injection. + +#![allow(unsafe_code)] + +use std::collections::HashMap; +use std::ffi::CString; +use std::fs::File; +use std::io::{self, Seek as _, SeekFrom, Write as _}; +use std::os::fd::{AsRawFd as _, FromRawFd as _}; +use std::sync::{Arc, RwLock}; + +use openshell_isolation_interface::linux::seccomp_notify::{Notification, NotificationListener}; +use openshell_isolation_interface::linux::task_memory; + +const PREFIX: &str = "/run/openshell/providers/"; +const MAX_FILE_BYTES: usize = 65_536; +const MAX_TOTAL_BYTES: usize = 262_144; +const MAX_PATH_BYTES: usize = 4_096; + +type Snapshot = HashMap>; + +/// A complete provider-file generation. Open handlers clone the selected +/// content before dropping the lock, so updates never block on a child open. +#[derive(Clone, Default)] +pub struct ProviderFiles { + current: Arc>>, +} + +impl ProviderFiles { + pub(crate) fn validate(desired: &HashMap) -> io::Result<()> { + if desired.len() > 64 || desired.values().map(String::len).sum::() > MAX_TOTAL_BYTES + { + return Err(io::Error::new( + io::ErrorKind::InvalidInput, + "provider file set is too large", + )); + } + for (path, content) in desired { + validate_path(path)?; + if content.len() > MAX_FILE_BYTES { + return Err(io::Error::new( + io::ErrorKind::InvalidInput, + "provider file exceeds 64 KiB", + )); + } + } + Ok(()) + } + + pub(crate) fn replace(&self, desired: HashMap) -> io::Result<()> { + Self::validate(&desired)?; + let next = desired + .into_iter() + .map(|(path, content)| (path, Arc::<[u8]>::from(content.into_bytes()))) + .collect(); + *self + .current + .write() + .unwrap_or_else(std::sync::PoisonError::into_inner) = Arc::new(next); + Ok(()) + } + + pub(crate) fn handle_open( + &self, + listener: &NotificationListener, + notification: Notification, + ) -> io::Result<()> { + let syscall = i64::from(notification.syscall); + let path_address = if syscall == libc::SYS_openat || syscall == libc::SYS_openat2 { + notification.args[1] + } else { + notification.args[0] + }; + // Every workload open reaches the listener. Copy only the reserved + // prefix for ordinary paths; full path reads are rare. + let mut prefix = [0_u8; PREFIX.len()]; + if task_memory::read_exact(notification.tid, path_address, &mut prefix).is_err() + || prefix != PREFIX.as_bytes() + { + return listener.respond_continue(notification.id); + } + // A failed or non-absolute lookup is left to the kernel. In particular, + // this preserves its normal EFAULT result for an invalid path pointer. + let Ok(path) = read_path(notification.tid, path_address) else { + return listener.respond_continue(notification.id); + }; + if !path.starts_with(PREFIX) { + return listener.respond_continue(notification.id); + } + let content = self + .current + .read() + .unwrap_or_else(std::sync::PoisonError::into_inner) + .get(&path) + .cloned(); + let Some(content) = content else { + return listener.respond_errno(notification.id, libc::ENOENT); + }; + let flags = match open_flags(¬ification) { + Ok(flags) => flags, + Err(error) => { + return listener.respond_errno( + notification.id, + error.raw_os_error().unwrap_or(libc::EINVAL), + ); + } + }; + if flags & libc::O_ACCMODE != libc::O_RDONLY + || flags + & (libc::O_CREAT | libc::O_EXCL | libc::O_TRUNC | libc::O_TMPFILE | libc::O_APPEND) + != 0 + { + return listener.respond_errno(notification.id, libc::EACCES); + } + if flags & (libc::O_DIRECTORY | libc::O_PATH | libc::O_DIRECT) != 0 { + return listener.respond_errno(notification.id, libc::EINVAL); + } + let file = sealed_memfd(&content)?; + listener.add_fd_and_send( + notification.id, + file.as_raw_fd(), + flags & libc::O_CLOEXEC != 0, + )?; + Ok(()) + } +} + +fn safe_component(value: &str) -> bool { + !value.is_empty() + && value != "." + && value != ".." + && value + .bytes() + .all(|c| c.is_ascii_alphanumeric() || matches!(c, b'.' | b'_' | b'-')) +} + +fn validate_path(path: &str) -> io::Result<()> { + let suffix = path.strip_prefix(PREFIX).ok_or_else(|| { + io::Error::new( + io::ErrorKind::InvalidInput, + "provider file escapes managed root", + ) + })?; + let (provider, name) = suffix.split_once('/').ok_or_else(|| { + io::Error::new( + io::ErrorKind::InvalidInput, + "provider file must name a provider and file", + ) + })?; + if !safe_component(provider) || !safe_component(name) { + return Err(io::Error::new( + io::ErrorKind::InvalidInput, + "invalid provider file path", + )); + } + Ok(()) +} + +fn read_path(tid: u32, mut address: u64) -> io::Result { + if address == 0 { + return Err(io::Error::from_raw_os_error(libc::EFAULT)); + } + let mut path = Vec::with_capacity(128); + while path.len() < MAX_PATH_BYTES { + // VMAs are page aligned. Never read across a page boundary before + // finding NUL, because the following page may be unmapped. + let page_remaining = 4096 - usize::try_from(address & 4095).expect("page offset fits"); + let length = page_remaining.min(MAX_PATH_BYTES - path.len()); + let mut chunk = vec![0; length]; + task_memory::read_exact(tid, address, &mut chunk)?; + if let Some(end) = chunk.iter().position(|byte| *byte == 0) { + path.extend_from_slice(&chunk[..end]); + return String::from_utf8(path).map_err(|_| io::Error::from_raw_os_error(libc::EINVAL)); + } + path.extend_from_slice(&chunk); + address += length as u64; + } + Err(io::Error::from_raw_os_error(libc::ENAMETOOLONG)) +} + +fn open_flags(notification: &Notification) -> io::Result { + let syscall = i64::from(notification.syscall); + if syscall == libc::SYS_openat2 { + if notification.args[3] < 24 { + return Err(io::Error::from_raw_os_error(libc::EINVAL)); + } + let mut how = [0_u8; 24]; + task_memory::read_exact(notification.tid, notification.args[2], &mut how)?; + let flags = u64::from_ne_bytes(how[0..8].try_into().expect("eight bytes")); + let mode = u64::from_ne_bytes(how[8..16].try_into().expect("eight bytes")); + let resolve = u64::from_ne_bytes(how[16..24].try_into().expect("eight bytes")); + if mode != 0 || resolve != 0 { + return Err(io::Error::from_raw_os_error(libc::EINVAL)); + } + i32::try_from(flags).map_err(|_| io::Error::from_raw_os_error(libc::EINVAL)) + } else if syscall == libc::SYS_openat { + i32::try_from(notification.args[2]).map_err(|_| io::Error::from_raw_os_error(libc::EINVAL)) + } else { + i32::try_from(notification.args[1]).map_err(|_| io::Error::from_raw_os_error(libc::EINVAL)) + } +} + +fn sealed_memfd(content: &[u8]) -> io::Result { + let name = CString::new("openshell-provider").expect("static name"); + let fd = + unsafe { libc::memfd_create(name.as_ptr(), libc::MFD_CLOEXEC | libc::MFD_ALLOW_SEALING) }; + if fd < 0 { + return Err(io::Error::last_os_error()); + } + let mut file = unsafe { File::from_raw_fd(fd) }; + file.write_all(content)?; + file.seek(SeekFrom::Start(0))?; + let seals = libc::F_SEAL_SEAL | libc::F_SEAL_WRITE | libc::F_SEAL_GROW | libc::F_SEAL_SHRINK; + if unsafe { libc::fcntl(file.as_raw_fd(), libc::F_ADD_SEALS, seals) } < 0 { + return Err(io::Error::last_os_error()); + } + // memfd_create returns O_RDWR. Reopen the sealed object through our own + // procfs descriptor so the child receives an actual O_RDONLY description. + File::open(format!("/proc/self/fd/{}", file.as_raw_fd())) +} + +#[cfg(test)] +mod tests { + use super::{ProviderFiles, sealed_memfd}; + use std::collections::HashMap; + use std::io::Read as _; + use std::os::fd::AsRawFd as _; + + #[test] + fn paths_cannot_escape_the_managed_tree() { + let valid = "/run/openshell/providers/acme/client.toml"; + assert!(ProviderFiles::validate(&HashMap::from([(valid.into(), "ok".into())])).is_ok()); + for path in [ + "/etc/passwd", + "/run/openshell/providers/acme/../passwd", + "/run/openshell/providers/acme/sub/file", + "/run/openshell/providers/../client.toml", + ] { + assert!( + ProviderFiles::validate(&HashMap::from([(path.into(), "ok".into())])).is_err(), + "{path}" + ); + } + } + + #[test] + fn memfd_is_read_only_and_positioned_at_start() { + let mut file = sealed_memfd(b"version = 1\n").unwrap(); + let flags = unsafe { libc::fcntl(file.as_raw_fd(), libc::F_GETFL) }; + assert_eq!(flags & libc::O_ACCMODE, libc::O_RDONLY); + let mut read = String::new(); + file.read_to_string(&mut read).unwrap(); + assert_eq!(read, "version = 1\n"); + assert!(std::io::Write::write_all(&mut file, b"changed").is_err()); + } +} diff --git a/crates/openshell-sdk/README.md b/crates/openshell-sdk/README.md index 4f721cbece..68c95e0c5b 100644 --- a/crates/openshell-sdk/README.md +++ b/crates/openshell-sdk/README.md @@ -56,8 +56,10 @@ portable workload shape and driver config. Failures map to a typed `SdkError` with a discriminable kind. Set `SandboxSpec::service_exposures` to register named or unnamed loopback HTTP -services during creation. Each `ServiceExposure` contains a service name and a -target port; an empty name selects the unnamed endpoint. The returned +services during creation. Each `ServiceExposure` contains a service name, a +target port, and an authorization mode; an empty name selects the unnamed +endpoint. Authorization is stripped by default. Select `BearerPassthrough` only +when the sandbox application validates its own bearer credential. The returned `SandboxRef::service_urls` map contains each routed URL under the same name. Curated calls without a workspace argument explicitly select the `default` diff --git a/crates/openshell-sdk/src/client.rs b/crates/openshell-sdk/src/client.rs index 0bf2412282..69b75ffbd2 100644 --- a/crates/openshell-sdk/src/client.rs +++ b/crates/openshell-sdk/src/client.rs @@ -1327,6 +1327,7 @@ fn create_sandbox_request(spec: SandboxSpec) -> proto::CreateSandboxRequest { command, tty, service_exposures, + restart_policy, } = spec; let template = image.map(|image| proto::SandboxTemplate { image, @@ -1344,6 +1345,7 @@ fn create_sandbox_request(spec: SandboxSpec) -> proto::CreateSandboxRequest { resource_requirements, command, tty, + restart_policy: proto::SandboxRestartPolicy::from(restart_policy) as i32, ..proto::SandboxSpec::default() }), name: name.unwrap_or_default(), @@ -1357,6 +1359,9 @@ fn create_sandbox_request(spec: SandboxSpec) -> proto::CreateSandboxRequest { .map(|exposure| proto::SandboxServiceExposure { service: exposure.service, target_port: u32::from(exposure.target_port), + authorization_mode: proto::ServiceAuthorizationMode::from( + exposure.authorization_mode, + ) as i32, }) .collect(), } @@ -1395,6 +1400,9 @@ fn create_sandbox_from_template_request( .map(|exposure| proto::SandboxServiceExposure { service: exposure.service, target_port: u32::from(exposure.target_port), + authorization_mode: proto::ServiceAuthorizationMode::from( + exposure.authorization_mode, + ) as i32, }) .collect(), } @@ -1634,11 +1642,16 @@ mod tests { let request = create_sandbox_request(SandboxSpec { command: vec!["/opt/agent binary".into(), "--serve exactly".into()], tty: false, + restart_policy: crate::types::SandboxRestartPolicy::OnFailure, ..SandboxSpec::default() }); let spec = request.spec.expect("sandbox spec should be present"); assert_eq!(spec.command, ["/opt/agent binary", "--serve exactly"]); assert!(!spec.tty); + assert_eq!( + spec.restart_policy(), + proto::SandboxRestartPolicy::OnFailure + ); } } diff --git a/crates/openshell-sdk/src/lib.rs b/crates/openshell-sdk/src/lib.rs index 269104ae2d..5459fa7532 100644 --- a/crates/openshell-sdk/src/lib.rs +++ b/crates/openshell-sdk/src/lib.rs @@ -51,9 +51,9 @@ pub use pagination::{Page, Pager}; pub use refresh::{Refresh, RefreshError, RefreshedToken, TokenSource}; pub use types::{ DeleteOptions, DeletionOutcome, DeletionResult, ExecOptions, ExecResult, Health, ListOptions, - LogLine, PlatformEvent, SandboxPhase, SandboxRef, SandboxResources, SandboxServiceLevel, - SandboxSpec, SandboxStartup, SandboxTemplateCreateSpec, SandboxTemplateListOptions, - SandboxWorkloadConfig, SandboxWorkloadTemplate, SandboxWorkloadTemplateProvenance, - SandboxWorkloadTemplateSpec, ServiceExposure, ServiceStatus, WatchEvent, WatchOptions, - WorkspaceRef, + LogLine, PlatformEvent, SandboxPhase, SandboxRef, SandboxResources, SandboxRestartPolicy, + SandboxServiceLevel, SandboxSpec, SandboxStartup, SandboxTemplateCreateSpec, + SandboxTemplateListOptions, SandboxWorkloadConfig, SandboxWorkloadTemplate, + SandboxWorkloadTemplateProvenance, SandboxWorkloadTemplateSpec, ServiceAuthorizationMode, + ServiceExposure, ServiceStatus, WatchEvent, WatchOptions, WorkspaceRef, }; diff --git a/crates/openshell-sdk/src/types.rs b/crates/openshell-sdk/src/types.rs index 57bed05e08..aa2bc71bf9 100644 --- a/crates/openshell-sdk/src/types.rs +++ b/crates/openshell-sdk/src/types.rs @@ -240,6 +240,25 @@ impl From for SandboxPhase { } } +/// Gateway policy for replacing the canonical main process after it exits. +#[derive(Clone, Copy, Debug, Default, Eq, PartialEq)] +pub enum SandboxRestartPolicy { + #[default] + Never, + OnFailure, + Always, +} + +impl From for proto::SandboxRestartPolicy { + fn from(value: SandboxRestartPolicy) -> Self { + match value { + SandboxRestartPolicy::Never => Self::Never, + SandboxRestartPolicy::OnFailure => Self::OnFailure, + SandboxRestartPolicy::Always => Self::Always, + } + } +} + /// Caller intent for a new sandbox. /// /// Only the most commonly used fields are exposed. Callers that need the @@ -266,6 +285,8 @@ pub struct SandboxSpec { pub tty: bool, /// Loopback HTTP services to expose when the sandbox is created. pub service_exposures: Vec, + /// Restart behavior after the canonical main process exits. + pub restart_policy: SandboxRestartPolicy, } /// A loopback HTTP service to expose during sandbox creation. @@ -275,6 +296,27 @@ pub struct ServiceExposure { pub service: String, /// Loopback TCP port inside the sandbox. pub target_port: u16, + /// Whether the gateway strips or forwards an application bearer credential. + pub authorization_mode: ServiceAuthorizationMode, +} + +/// Handling for an incoming application `Authorization` header. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub enum ServiceAuthorizationMode { + /// Remove the header before proxying to the sandbox service. + #[default] + Strip, + /// Forward one syntactically valid bearer credential unchanged. + BearerPassthrough, +} + +impl From for proto::ServiceAuthorizationMode { + fn from(value: ServiceAuthorizationMode) -> Self { + match value { + ServiceAuthorizationMode::Strip => Self::Strip, + ServiceAuthorizationMode::BearerPassthrough => Self::BearerPassthrough, + } + } } /// Caller intent for creating a sandbox from a named workload template. @@ -346,6 +388,9 @@ pub struct SandboxRef { /// Service URLs returned by sandbox creation, keyed by service name. The /// empty key identifies the unnamed service. Non-create reads leave this empty. pub service_urls: HashMap, + pub restart_count: u32, + pub next_restart_at_ms: Option, + pub main_process_started_at_ms: Option, } /// Reusable workload template revision used to create a sandbox. @@ -359,7 +404,6 @@ pub struct SandboxWorkloadTemplateProvenance { impl SandboxRef { pub(crate) fn from_proto(sandbox: proto::Sandbox) -> Self { let phase = sandbox.phase().into(); - let exit_code = sandbox.status.as_ref().and_then(|status| status.exit_code); let created_from_workload_template = sandbox .created_from_workload_template @@ -367,6 +411,23 @@ impl SandboxRef { name: p.name, resource_version: p.resource_version, }); + let (exit_code, restart_count, next_restart_at_ms, main_process_started_at_ms) = sandbox + .status + .as_ref() + .map_or((None, 0, None, None), |status| { + ( + status.exit_code, + status.restart_count, + status + .next_restart_time + .as_ref() + .and_then(|time| openshell_core::time::timestamp_to_millis(time).ok()), + status + .main_process_started_time + .as_ref() + .and_then(|time| openshell_core::time::timestamp_to_millis(time).ok()), + ) + }); let meta = sandbox.metadata.unwrap_or_default(); Self { id: meta.id, @@ -378,6 +439,9 @@ impl SandboxRef { exit_code, created_from_workload_template, service_urls: HashMap::new(), + restart_count, + next_restart_at_ms, + main_process_started_at_ms, } } } diff --git a/crates/openshell-sdk/tests/client_mock.rs b/crates/openshell-sdk/tests/client_mock.rs index 1c86ada7dd..9506160f07 100644 --- a/crates/openshell-sdk/tests/client_mock.rs +++ b/crates/openshell-sdk/tests/client_mock.rs @@ -13,8 +13,8 @@ use openshell_core::proto::open_shell_server::{OpenShell, OpenShellServer}; use openshell_sdk::{ AuthConfig, ClientConfig, ExecOptions, ListOptions, OpenShellClient, Refresh, RefreshError, RefreshedToken, SandboxPhase, SandboxSpec, SandboxTemplateCreateSpec, - SandboxTemplateListOptions, ServiceExposure, ServiceStatus as SdkServiceStatus, WatchEvent, - WatchOptions, + SandboxTemplateListOptions, ServiceAuthorizationMode, ServiceExposure, + ServiceStatus as SdkServiceStatus, WatchEvent, WatchOptions, }; use std::collections::HashMap; use std::sync::Arc; @@ -1139,6 +1139,7 @@ async fn create_sandbox_passes_spec_through() { service_exposures: vec![ServiceExposure { service: "web".to_string(), target_port: 8080, + authorization_mode: ServiceAuthorizationMode::BearerPassthrough, }], ..Default::default() }; @@ -1158,6 +1159,10 @@ async fn create_sandbox_passes_spec_through() { assert_eq!(observed.service_exposures.len(), 1); assert_eq!(observed.service_exposures[0].service, "web"); assert_eq!(observed.service_exposures[0].target_port, 8080); + assert_eq!( + observed.service_exposures[0].authorization_mode(), + proto::ServiceAuthorizationMode::BearerPassthrough + ); let observed_spec = observed.spec.unwrap(); assert!( observed_spec diff --git a/crates/openshell-server-macros/src/lib.rs b/crates/openshell-server-macros/src/lib.rs index e85d316293..167ebb970b 100644 --- a/crates/openshell-server-macros/src/lib.rs +++ b/crates/openshell-server-macros/src/lib.rs @@ -14,8 +14,6 @@ //! Generated code references `crate::auth::method_authz::{MethodAuth, //! AuthMode, Role}`, so the macro is only intended for use inside the //! `openshell-server` crate. -//! -//! See `architecture/plans/scope-annotations.md` for the design. use proc_macro::TokenStream; use quote::quote; diff --git a/crates/openshell-server/src/cli.rs b/crates/openshell-server/src/cli.rs index f531d88c5d..9000ab1324 100644 --- a/crates/openshell-server/src/cli.rs +++ b/crates/openshell-server/src/cli.rs @@ -603,7 +603,23 @@ async fn run_from_args( ) -> Result<()> { let prepared = prepare_server_config_with_drivers(&mut args, &matches, &compute_drivers)?; + // Initialize OCSF identity before tracing can emit gateway events. + let gateway_identity = crate::gateway_ocsf::GatewayIdentity { + name: prepared.config.name.clone(), + hostname: crate::compute::lease::replica_id(), + }; + if !crate::gateway_ocsf::set_identity(gateway_identity) { + tracing::debug!("gateway OCSF identity already initialized, keeping existing"); + } + let tracing_log_bus = TracingLogBus::new(); + let ocsf_log = prepared + .config_file + .as_ref() + .and_then(|file| file.openshell.gateway.ocsf_log.clone()) + .map(crate::ocsf_log::OcsfLog::start) + .transpose() + .into_diagnostic()?; let otlp_config = prepared .config_file .as_ref() @@ -617,9 +633,11 @@ async fn run_from_args( &prepared.config.compute_driver_endpoints, ); let (tracing_handle, setup_error) = crate::tracing_setup::install( - EnvFilter::try_from_default_env() - .unwrap_or_else(|_| EnvFilter::new(&prepared.config.log_level)), + &EnvFilter::try_from_default_env() + .unwrap_or_else(|_| EnvFilter::new(&prepared.config.log_level)) + .to_string(), &tracing_log_bus, + ocsf_log.as_ref(), otlp_config, compute_driver_tracing, gateway_resource, @@ -687,6 +705,10 @@ async fn run_from_args( tracing_handle.shutdown(); + if let Some(log) = ocsf_log { + log.shutdown().await; + } + result.into_diagnostic() } diff --git a/crates/openshell-server/src/compute/mod.rs b/crates/openshell-server/src/compute/mod.rs index 4a38361e1b..d2e9642263 100644 --- a/crates/openshell-server/src/compute/mod.rs +++ b/crates/openshell-server/src/compute/mod.rs @@ -25,6 +25,8 @@ use hyper_util::rt::TokioIo; use openshell_core::extension_protocol::{ ExtensionFamily, NegotiatedExtension, gateway_metadata, negotiate, }; +#[cfg(test)] +use openshell_core::proto::ServiceAuthorizationMode; use openshell_core::proto::compute::v1::{ AuthenticateSandboxRequest, CreateSandboxRequest, DeleteSandboxRequest, DeleteWorkspaceRequest, DeleteWorkspaceResponse, DriverCondition, DriverPlatformEvent, DriverResourceRequirements, @@ -38,8 +40,8 @@ use openshell_core::proto::compute::v1::{ compute_driver_server::ComputeDriver, watch_sandboxes_event, }; use openshell_core::proto::{ - PlatformEvent, Sandbox, SandboxCondition, SandboxPhase, SandboxSpec, SandboxStatus, - SandboxTemplate, SandboxWorkloadTemplate, ServiceEndpoint, SshSession, + PlatformEvent, Sandbox, SandboxCondition, SandboxPhase, SandboxRestartPolicy, SandboxSpec, + SandboxStatus, SandboxTemplate, SandboxWorkloadTemplate, ServiceEndpoint, SshSession, }; use openshell_core::telemetry::TelemetryComputeDriver; use openshell_core::{ObjectLabels, ObjectWorkspace}; @@ -49,11 +51,11 @@ use std::fmt; use std::future::Future; use std::path::{Path, PathBuf}; use std::pin::Pin; -use std::sync::{Arc, Mutex as StdMutex, Weak}; +use std::sync::{Arc, Mutex as StdMutex, OnceLock, Weak}; use std::time::{Duration, Instant}; #[cfg(unix)] use tokio::net::UnixStream; -use tokio::sync::{Mutex, watch}; +use tokio::sync::{Mutex, Notify, watch}; use tonic::transport::Channel; #[cfg(unix)] use tonic::transport::Endpoint; @@ -319,6 +321,17 @@ async fn wait_for_startup( /// Interval between store-vs-backend reconciliation sweeps. const RECONCILE_INTERVAL: Duration = Duration::from_mins(1); +/// Restart intents live in sandbox rows. A lightweight holder-only scan makes +/// cross-replica writes and lost wakeups recover without changing the general +/// backend reconciliation cadence. +const RESTART_SCAN_INTERVAL: Duration = Duration::from_millis(500); +const RESTART_BACKOFF_BASE_DELAY_MS: i64 = 1_000; +const RESTART_MAX_DELAY_MS: i64 = 180_000; +const RESTART_STABILITY_WINDOW_MS: i64 = 10_000; +const RESTART_DRIVER_RETRY_DELAY_MS: i64 = 10_000; +const RESTART_TERMINAL_DELIVERY_GRACE_MS: i64 = 10_000; +const RESTART_READINESS_TIMEOUT_MS: i64 = 120_000; + /// How long a sandbox can remain provisioning in the store without a /// corresponding backend resource before it is considered orphaned. const ORPHAN_GRACE_PERIOD: Duration = Duration::from_mins(5); @@ -643,6 +656,9 @@ pub struct ComputeRuntime { /// clones: `ServerState` holds `ComputeRuntime` by value, so a per-clone /// table would make a token minted on one clone invisible to another. rootfs_tar_staging: Arc, + restart_authority: + Arc>>>, + restart_notify: Arc, } pub struct SandboxSyncGuard { @@ -738,6 +754,8 @@ impl ComputeRuntime { lifecycle_gates: Arc::new(LifecycleGateRegistry::default()), replica_id: lease::replica_id(), rootfs_tar_staging, + restart_authority: Arc::new(OnceLock::new()), + restart_notify: Arc::new(Notify::new()), }) } @@ -1341,9 +1359,10 @@ impl ComputeRuntime { .map_err(Status::internal)?; return Ok(current); } - if !matches!(phase, SandboxPhase::Ready | SandboxPhase::Stopping) { + let automatic_restart = is_automatic_restart_transition(¤t); + if !matches!(phase, SandboxPhase::Ready | SandboxPhase::Stopping) && !automatic_restart { return Err(Status::failed_precondition(format!( - "sandbox must be Ready to stop (current phase: {phase:?})" + "sandbox must be Ready or in an automatic restart to stop (current phase: {phase:?})" ))); } @@ -2749,7 +2768,9 @@ impl ComputeRuntime { &self, shutdown_rx: watch::Receiver, startup_rx: watch::Receiver, + authority: Option>, ) { + let _ = self.restart_authority.set(authority); let runtime = Arc::new(self.clone()); if self.store.is_single_replica() { let deadline_runtime = runtime.clone(); @@ -2766,11 +2787,16 @@ impl ComputeRuntime { } Box::pin(watch_runtime.watch_loop(watch_shutdown)).await; }); + let reconcile_runtime = runtime.clone(); + let reconcile_shutdown = shutdown_rx.clone(); tokio::spawn(async move { - if !wait_for_startup(startup_rx, shutdown_rx.clone()).await { + if !wait_for_startup(startup_rx, reconcile_shutdown.clone()).await { return; } - runtime.reconcile_loop(shutdown_rx).await; + reconcile_runtime.reconcile_loop(reconcile_shutdown).await; + }); + tokio::spawn(async move { + runtime.restart_loop(shutdown_rx).await; }); } else { tokio::spawn(async move { @@ -2989,6 +3015,12 @@ impl ComputeRuntime { failed += 1; continue; } + if is_automatic_restart_transition(&sandbox) { + // The restart loop resumes both scheduled and claimed + // attempts. It always stops the old runtime before starting + // its replacement, including after a gateway crash. + continue; + } let sandbox_name = sandbox.object_name().to_string(); let generation_id = match sandbox_runtime_generation(&sandbox) { @@ -3470,6 +3502,12 @@ impl ComputeRuntime { runtime.reconcile_loop(cancel_rx).await; }); + let runtime = self.clone(); + let cancel_rx = cancel_tx.subscribe(); + let restart_handle = tokio::spawn(async move { + runtime.restart_loop(cancel_rx).await; + }); + loop { tokio::select! { () = tokio::time::sleep(LEASE_RENEWAL_INTERVAL) => { @@ -3497,6 +3535,7 @@ impl ComputeRuntime { let _ = watch_handle.await; let _ = reconcile_handle.await; let _ = deadline_handle.await; + let _ = restart_handle.await; return; } } @@ -3507,6 +3546,7 @@ impl ComputeRuntime { let _ = watch_handle.await; let _ = reconcile_handle.await; let _ = deadline_handle.await; + let _ = restart_handle.await; info!(replica = %lease.replica_id(), "reconciler lease lost — returning to standby"); } @@ -3568,125 +3608,553 @@ impl ComputeRuntime { } } - #[tracing::instrument( - name = "reconcile", - skip_all, - fields( - otel.name = "reconcile.sandboxes", - driver.name = %self.driver_info.name, - backend_count = tracing::field::Empty, - store_count = tracing::field::Empty, - ) - )] - async fn reconcile_store_with_backend(&self, grace_period: Duration) -> Result<(), String> { - let sweep_started_at_ms = openshell_core::time::now_ms(); - // Reclaims staging directories whose driver failed before its own - // cleanup ran, which the token table cannot see once consumed. - self.rootfs_tar_staging.sweep_orphans(); - let backend_sandboxes = self - .driver - .call( - openshell_otel::rpc::LIST_SANDBOXES, - None, - |driver| async move { - driver - .list_sandboxes(Request::new(ListSandboxesRequest {})) - .await - }, - ) - .await - .map_err(|e| e.to_string()) - .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))? - .into_inner() - .sandboxes; - let backend_ids = backend_sandboxes - .iter() - .map(|sandbox| sandbox.id.clone()) - .collect::>(); - tracing::Span::current().record("backend_count", backend_sandboxes.len()); + async fn restart_loop(self: Arc, mut cancel: watch::Receiver) { + loop { + if let Err(err) = self.restart_due_sandboxes().await { + warn!(error = %err, "Sandbox restart sweep failed"); + } + tokio::select! { + () = tokio::time::sleep(RESTART_SCAN_INTERVAL) => {} + () = self.restart_notify.notified() => {} + _ = cancel.changed() => return, + } + } + } - for sandbox in backend_sandboxes { - self.reconcile_snapshot_sandbox(sandbox, sweep_started_at_ms) + async fn restart_due_sandboxes(&self) -> Result<(), String> { + let now_ms = openshell_core::time::now_ms(); + let mut offset = 0; + loop { + let records = self + .store + .list_by_type(Sandbox::object_type(), 1000, offset) .await - .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))?; + .map_err(|err| err.to_string())?; + let page_len = records.len(); + for record in records { + let sandbox = match Sandbox::decode(record.payload.as_slice()) { + Ok(sandbox) => sandbox, + Err(err) => { + warn!(error = %err, "Failed to decode sandbox during restart sweep"); + continue; + } + }; + let phase = + SandboxPhase::try_from(sandbox.phase()).unwrap_or(SandboxPhase::Unknown); + let due = sandbox + .status + .as_ref() + .is_some_and(|status| restart_is_due(status, now_ms)); + if phase == SandboxPhase::Starting + && is_automatic_restart_transition(&sandbox) + && due + && let Err(err) = self.restart_sandbox_runtime(sandbox.object_id()).await + { + warn!( + sandbox_id = %sandbox.object_id(), + sandbox_name = %sandbox.object_name(), + error = %err, + "Automatic sandbox restart attempt failed" + ); + } + } + if page_len < 1000 { + break; + } + offset += u32::try_from(page_len).expect("restart scan page length is capped at 1000"); } + Ok(()) + } - let records = self + async fn restart_sandbox_runtime(&self, sandbox_id: &str) -> Result<(), String> { + let lifecycle_guard = self.lifecycle_gates.lock_for(sandbox_id).await; + let global_guard = self.lock_global_for_lifecycle(&lifecycle_guard).await; + let Some(current) = self .store - .collect_records(Sandbox::object_type(), ObjectListQuery::AllWorkspaces) + .get_message::(sandbox_id) .await - .map_err(|e| e.to_string()) - .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))?; - tracing::Span::current().record("store_count", records.len()); + .map_err(|err| err.to_string())? + else { + return Ok(()); + }; + let phase = SandboxPhase::try_from(current.phase()).unwrap_or(SandboxPhase::Unknown); + let now_ms = openshell_core::time::now_ms(); + let due = current + .status + .as_ref() + .is_some_and(|status| restart_is_due(status, now_ms)); + if phase != SandboxPhase::Starting || !is_automatic_restart_transition(¤t) || !due { + return Ok(()); + } - let grace_ms = grace_period.as_millis().try_into().unwrap_or(i64::MAX); + let sandbox_name = current.object_name().to_string(); + let authority = self.restart_authority.get().and_then(Option::as_deref); + let already_claimed = current + .status + .as_ref() + .is_some_and(|status| next_restart_at_ms(status) == 0); + let next_identity = if already_claimed { + None + } else { + authority + .map(|_| next_runtime_identity(¤t)) + .transpose() + .map_err(|status| status.to_string())? + }; + let claimed = if already_claimed { + current.clone() + } else { + self.store + .update_message_cas::( + sandbox_id, + sandbox_resource_version(¤t), + |sandbox| { + if let (Some(identity), Some(metadata)) = + (next_identity.as_ref(), sandbox.metadata.as_mut()) + { + identity.write(&mut metadata.annotations); + } + let status = sandbox.status.get_or_insert_with(Default::default); + // Zero durably claims this due attempt across gateway + // replicas. The readiness watchdog starts only after the + // old runtime has stopped. + set_next_restart_at_ms(status, 0); + upsert_ready_condition( + &mut sandbox.status, + SandboxCondition { + r#type: "Ready".to_string(), + status: "False".to_string(), + reason: "SandboxRestarting".to_string(), + message: + "Replacing sandbox runtime after canonical main process exit" + .to_string(), + transition_time: None, + }, + ); + }, + ) + .await + .map_err(|err| err.to_string())? + }; + self.sandbox_index.update_from_sandbox(&claimed); + self.sandbox_watch_bus.notify(sandbox_id); + drop(global_guard); - for record in records { - let sandbox = match Sandbox::decode(record.payload.as_slice()) { - Ok(sandbox) => sandbox, - Err(err) => { - warn!(error = %err, "Failed to decode sandbox record during reconciliation"); - continue; - } - }; + let stop_result = self + .driver + .call( + openshell_otel::rpc::STOP_SANDBOX, + Some(sandbox_id), + |driver| { + let sandbox_id = sandbox_id.to_string(); + let sandbox_name = sandbox_name.clone(); + async move { + driver + .stop_sandbox(Request::new(StopSandboxRequest { + sandbox_id, + name: sandbox_name, + })) + .await + } + }, + ) + .await; + if let Err(status) = stop_result { + return self + .record_restart_driver_failure(&lifecycle_guard, sandbox_id, "stop", status) + .await; + } - if backend_ids.contains(sandbox.object_id()) { - continue; - } + self.cleanup_stopped_sandbox_sessions(&claimed).await?; - self.prune_missing_sandbox(record, sweep_started_at_ms, grace_ms) - .await - .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))?; + let Some(armed) = self.arm_restart_readiness_watchdog(&claimed).await? else { + // A concurrent stop or delete changed durable intent while the + // old runtime was stopping. Do not recreate it. + return Ok(()); + }; + + let launch_authentication = serialize_persisted_launch_authentication(authority, &armed) + .map_err(|status| status.to_string())?; + let generation_id = sandbox_runtime_generation(&armed)?.into_string(); + let expected_runtime_identity = sandbox_compute_runtime_identity(¤t); + + let start_result = self + .driver + .call( + openshell_otel::rpc::START_SANDBOX, + Some(sandbox_id), + |driver| { + let sandbox_id = sandbox_id.to_string(); + let sandbox_name = sandbox_name.clone(); + async move { + driver + .start_sandbox(Request::new(StartSandboxRequest { + sandbox_id, + name: sandbox_name, + launch_authentication, + generation_id, + expected_runtime_identity, + })) + .await + } + }, + ) + .await; + let response = match start_result { + Ok(response) => response.into_inner(), + Err(status) => { + return self + .record_restart_driver_failure(&lifecycle_guard, sandbox_id, "start", status) + .await; + } + }; + if self.supports_sandbox_authentication() { + if response.runtime_identity.is_empty() { + return Err( + "compute driver returned an empty runtime identity after restart".into(), + ); + } + self.persist_runtime_binding( + sandbox_id, + &armed, + self.configured_driver_name(), + &response.runtime_identity, + &[SandboxPhase::Starting, SandboxPhase::Ready], + ) + .await?; } + self.enforce_lifecycle_after_restart_start(&armed).await?; + + info!( + sandbox_id, + sandbox_name, "Sandbox runtime restarted; waiting for replacement supervisor" + ); Ok(()) } - async fn apply_watch_event(&self, event: WatchSandboxesEvent) -> Result<(), String> { - let (operation, sandbox_id) = match &event.payload { - Some(watch_sandboxes_event::Payload::Sandbox(update)) => ( - "driver_watch.sandbox_updated", - update - .sandbox + /// Arm the replacement readiness watchdog immediately before starting + /// compute. The zero deadline written by the restart claim fences other + /// gateway replicas without charging slow driver shutdown time against + /// the replacement supervisor's readiness window. + async fn arm_restart_readiness_watchdog( + &self, + claimed: &Sandbox, + ) -> Result, String> { + let sandbox_id = claimed.object_id().to_string(); + let mut current = claimed.clone(); + for _ in 0..DELETE_PHASE_CAS_RETRY_LIMIT { + let phase = SandboxPhase::try_from(current.phase()).unwrap_or(SandboxPhase::Unknown); + let claim_owned = phase == SandboxPhase::Starting + && is_automatic_restart_transition(¤t) + && current + .status .as_ref() - .map(|sandbox| sandbox.id.as_str()) - .unwrap_or_default(), - ), - Some(watch_sandboxes_event::Payload::Deleted(deleted)) => { - ("driver_watch.sandbox_deleted", deleted.sandbox_id.as_str()) + .is_some_and(|status| next_restart_at_ms(status) == 0); + if !claim_owned { + return Ok(None); } - Some(watch_sandboxes_event::Payload::PlatformEvent(platform_event)) => ( - "driver_watch.platform_event", - platform_event.sandbox_id.as_str(), - ), - None => return Ok(()), - }; - let span = tracing::info_span!( - "driver_watch", - otel.name = operation, - otel.status_code = tracing::field::Empty, - sandbox.id = %sandbox_id, - ); - async { - let result = self.apply_watch_event_inner(event).await; - if result.is_err() { - crate::otel_tracing::mark_error(&tracing::Span::current()); + + let readiness_deadline = + openshell_core::time::now_ms().saturating_add(RESTART_READINESS_TIMEOUT_MS); + match self + .store + .update_message_cas::( + &sandbox_id, + sandbox_resource_version(¤t), + |sandbox| { + let status = sandbox.status.get_or_insert_with(Default::default); + set_next_restart_at_ms(status, readiness_deadline); + }, + ) + .await + { + Ok(armed) => { + self.sandbox_index.update_from_sandbox(&armed); + self.sandbox_watch_bus.notify(&sandbox_id); + return Ok(Some(armed)); + } + Err(crate::persistence::PersistenceError::Conflict { .. }) => { + let Some(latest) = self + .store + .get_message::(&sandbox_id) + .await + .map_err(|err| err.to_string())? + else { + return Ok(None); + }; + current = latest; + } + Err(err) => return Err(err.to_string()), } - result } - .instrument(span) - .await + + Ok(None) } - async fn apply_watch_event_inner(&self, event: WatchSandboxesEvent) -> Result<(), String> { - validate_driver_watch_event_timestamps(&event)?; - match event.payload { - Some(watch_sandboxes_event::Payload::Sandbox(sandbox)) => { - if let Some(sandbox) = sandbox.sandbox { - Box::pin(self.apply_sandbox_update(sandbox)).await?; - } - } - Some(watch_sandboxes_event::Payload::Deleted(deleted)) => { + /// Reassert a concurrent stop/delete after the replacement start call. + /// The durable phase is authoritative across gateway replicas; the + /// compensating driver operation prevents a late start from leaving + /// running compute behind a stopped row or an orphan after deletion. + async fn enforce_lifecycle_after_restart_start(&self, armed: &Sandbox) -> Result<(), String> { + let sandbox_id = armed.object_id().to_string(); + let sandbox_name = armed.object_name().to_string(); + let latest = self + .store + .get_message::(&sandbox_id) + .await + .map_err(|err| err.to_string())?; + let phase = latest.as_ref().map(|sandbox| { + SandboxPhase::try_from(sandbox.phase()).unwrap_or(SandboxPhase::Unknown) + }); + + match phase { + None | Some(SandboxPhase::Deleting) => { + let result = self + .driver + .call( + openshell_otel::rpc::DELETE_SANDBOX, + Some(&sandbox_id), + |driver| { + let sandbox_id = sandbox_id.clone(); + let sandbox_name = sandbox_name.clone(); + async move { + driver + .delete_sandbox(Request::new(DeleteSandboxRequest { + sandbox_id, + name: sandbox_name, + })) + .await + } + }, + ) + .await; + if let Err(status) = result + && status.code() != Code::NotFound + { + return Err(format!( + "failed to reassert concurrent sandbox delete after restart: {status}" + )); + } + } + Some(SandboxPhase::Stopping | SandboxPhase::Stopped) => { + let result = self + .driver + .call( + openshell_otel::rpc::STOP_SANDBOX, + Some(&sandbox_id), + |driver| { + let sandbox_id = sandbox_id.clone(); + let sandbox_name = sandbox_name.clone(); + async move { + driver + .stop_sandbox(Request::new(StopSandboxRequest { + sandbox_id, + name: sandbox_name, + })) + .await + } + }, + ) + .await; + if let Err(status) = result + && status.code() != Code::NotFound + { + return Err(format!( + "failed to reassert concurrent sandbox stop after restart: {status}" + )); + } + } + _ => {} + } + + Ok(()) + } + + async fn record_restart_driver_failure( + &self, + lifecycle_guard: &SandboxLifecycleGuard, + sandbox_id: &str, + operation: &str, + driver_status: Status, + ) -> Result<(), String> { + let _global_guard = self.lock_global_for_lifecycle(lifecycle_guard).await; + let Some(current) = self + .store + .get_message::(sandbox_id) + .await + .map_err(|err| err.to_string())? + else { + return Ok(()); + }; + if !is_automatic_restart_transition(¤t) { + return Ok(()); + } + let terminal = driver_status.code() == Code::NotFound; + let operation = operation.to_string(); + let driver_message = driver_status.message().to_string(); + let now_ms = openshell_core::time::now_ms(); + let updated = self + .store + .update_message_cas::( + sandbox_id, + sandbox_resource_version(¤t), + |sandbox| { + let status = sandbox.status.get_or_insert_with(Default::default); + if terminal { + set_next_restart_at_ms(status, 0); + sandbox.set_phase(SandboxPhase::Error as i32); + } else { + set_next_restart_at_ms( + status, + now_ms.saturating_add( + restart_delay_ms(status.restart_count.max(1)) + .max(RESTART_DRIVER_RETRY_DELAY_MS), + ), + ); + } + upsert_ready_condition( + &mut sandbox.status, + SandboxCondition { + r#type: "Ready".to_string(), + status: "False".to_string(), + reason: "SandboxRestartFailed".to_string(), + message: format!( + "Sandbox restart {operation} failed: {driver_message}" + ), + transition_time: None, + }, + ); + }, + ) + .await + .map_err(|err| err.to_string())?; + self.sandbox_index.update_from_sandbox(&updated); + self.sandbox_watch_bus.notify(sandbox_id); + Err(format!( + "driver {operation} failed during sandbox restart: {driver_status}" + )) + } + + #[tracing::instrument( + name = "reconcile", + skip_all, + fields( + otel.name = "reconcile.sandboxes", + driver.name = %self.driver_info.name, + backend_count = tracing::field::Empty, + store_count = tracing::field::Empty, + ) + )] + async fn reconcile_store_with_backend(&self, grace_period: Duration) -> Result<(), String> { + let sweep_started_at_ms = openshell_core::time::now_ms(); + // Reclaims staging directories whose driver failed before its own + // cleanup ran, which the token table cannot see once consumed. + self.rootfs_tar_staging.sweep_orphans(); + let backend_sandboxes = self + .driver + .call( + openshell_otel::rpc::LIST_SANDBOXES, + None, + |driver| async move { + driver + .list_sandboxes(Request::new(ListSandboxesRequest {})) + .await + }, + ) + .await + .map_err(|e| e.to_string()) + .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))? + .into_inner() + .sandboxes; + let backend_ids = backend_sandboxes + .iter() + .map(|sandbox| sandbox.id.clone()) + .collect::>(); + tracing::Span::current().record("backend_count", backend_sandboxes.len()); + + for sandbox in backend_sandboxes { + self.reconcile_snapshot_sandbox(sandbox, sweep_started_at_ms) + .await + .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))?; + } + + let records = self + .store + .collect_records(Sandbox::object_type(), ObjectListQuery::AllWorkspaces) + .await + .map_err(|e| e.to_string()) + .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))?; + tracing::Span::current().record("store_count", records.len()); + + let grace_ms = grace_period.as_millis().try_into().unwrap_or(i64::MAX); + + for record in records { + let sandbox = match Sandbox::decode(record.payload.as_slice()) { + Ok(sandbox) => sandbox, + Err(err) => { + warn!(error = %err, "Failed to decode sandbox record during reconciliation"); + continue; + } + }; + + if backend_ids.contains(sandbox.object_id()) { + continue; + } + + self.prune_missing_sandbox(record, sweep_started_at_ms, grace_ms) + .await + .inspect_err(|_| crate::otel_tracing::mark_error(&tracing::Span::current()))?; + } + + Ok(()) + } + + async fn apply_watch_event(&self, event: WatchSandboxesEvent) -> Result<(), String> { + let (operation, sandbox_id) = match &event.payload { + Some(watch_sandboxes_event::Payload::Sandbox(update)) => ( + "driver_watch.sandbox_updated", + update + .sandbox + .as_ref() + .map(|sandbox| sandbox.id.as_str()) + .unwrap_or_default(), + ), + Some(watch_sandboxes_event::Payload::Deleted(deleted)) => { + ("driver_watch.sandbox_deleted", deleted.sandbox_id.as_str()) + } + Some(watch_sandboxes_event::Payload::PlatformEvent(platform_event)) => ( + "driver_watch.platform_event", + platform_event.sandbox_id.as_str(), + ), + None => return Ok(()), + }; + let span = tracing::info_span!( + "driver_watch", + otel.name = operation, + otel.status_code = tracing::field::Empty, + sandbox.id = %sandbox_id, + ); + async { + let result = self.apply_watch_event_inner(event).await; + if result.is_err() { + crate::otel_tracing::mark_error(&tracing::Span::current()); + } + result + } + .instrument(span) + .await + } + + async fn apply_watch_event_inner(&self, event: WatchSandboxesEvent) -> Result<(), String> { + validate_driver_watch_event_timestamps(&event)?; + match event.payload { + Some(watch_sandboxes_event::Payload::Sandbox(sandbox)) => { + if let Some(sandbox) = sandbox.sandbox { + Box::pin(self.apply_sandbox_update(sandbox)).await?; + } + } + Some(watch_sandboxes_event::Payload::Deleted(deleted)) => { self.apply_deleted(&deleted.sandbox_id).await?; } Some(watch_sandboxes_event::Payload::PlatformEvent(platform_event)) => { @@ -4007,6 +4475,15 @@ impl ComputeRuntime { if !connected && current_phase != SandboxPhase::Ready { return Ok(()); } + if connected + && is_automatic_restart_transition(¤t) + && current.status.as_ref().is_some_and(|status| { + !status.main_process_instance_id.is_empty() + && Some(status.main_process_instance_id.as_str()) == instance_id + }) + { + return Err("restarting sandbox requires a fresh supervisor instance".into()); + } let expected_resource_version = sandbox_resource_version(¤t); let result = self .store @@ -4017,9 +4494,18 @@ impl ComputeRuntime { if connected { ensure_supervisor_ready_status(&mut sandbox.status); let status = sandbox.status.get_or_insert_with(Default::default); + let same_instance = + Some(status.main_process_instance_id.as_str()) == instance_id; status.main_process_instance_id = instance_id.unwrap_or_default().to_string(); status.exit_code = None; + set_next_restart_at_ms(status, 0); + if !same_instance || main_process_started_at_ms(status) == 0 { + set_main_process_started_at_ms( + status, + openshell_core::time::now_ms(), + ); + } sandbox.set_phase(SandboxPhase::Ready as i32); } else { ensure_supervisor_not_ready_status(&mut sandbox.status); @@ -4109,6 +4595,11 @@ impl ComputeRuntime { ) { return Ok(()); } + let accepting_replacement_exit = existing.status.as_ref().is_some_and(|status| { + phase == SandboxPhase::Starting + && !status.main_process_instance_id.is_empty() + && status.main_process_instance_id != instance_id + }); if let Some(status) = existing.status.as_ref() { if !status.main_process_instance_id.is_empty() { // While Starting, the stored id belongs to the stopped @@ -4134,7 +4625,9 @@ impl ComputeRuntime { return Ok(()); } } - if let Some(current_exit_code) = status.exit_code { + if let Some(current_exit_code) = status.exit_code + && !accepting_replacement_exit + { if current_exit_code != exit_code { tracing::warn!( sandbox_id, @@ -4148,11 +4641,21 @@ impl ComputeRuntime { return Ok(()); } } + let restart = !has_specific_infrastructure_error(&existing) + && existing + .spec + .as_ref() + .is_some_and(|spec| should_restart_main_process(spec.restart_policy, exit_code)); + let now_ms = openshell_core::time::now_ms(); let expected_resource_version = sandbox_resource_version(&existing); let sandbox = self .store .update_message_cas::(sandbox_id, expected_resource_version, |sandbox| { - apply_main_process_exit(sandbox, instance_id, exit_code); + if restart { + apply_main_process_restart(sandbox, instance_id, exit_code, now_ms); + } else { + apply_main_process_exit(sandbox, instance_id, exit_code); + } }) .await .map_err(|error| error.to_string())?; @@ -4198,6 +4701,27 @@ impl ComputeRuntime { { return Err("main-process instance does not match the terminal result".to_string()); } + if is_terminal_delivery_pending(status) { + let restart_count = status.restart_count; + let delay_ms = restart_delay_ms(restart_count); + let updated = self + .store + .update_message_cas::( + sandbox_id, + sandbox_resource_version(&sandbox), + |sandbox| { + upsert_ready_condition( + &mut sandbox.status, + restart_ready_condition(restart_count, delay_ms), + ); + }, + ) + .await + .map_err(|error| error.to_string())?; + self.sandbox_index.update_from_sandbox(&updated); + self.sandbox_watch_bus.notify(sandbox_id); + self.restart_notify.notify_one(); + } Ok(()) } @@ -4772,12 +5296,7 @@ fn apply_main_process_exit(sandbox: &mut Sandbox, instance_id: &str, exit_code: // ContainerExited is only a provisional classification: replace it with // the canonical process result once its exit code is known. Preserve all // other infrastructure errors. - let preserve_infrastructure_error = sandbox.phase() == SandboxPhase::Error as i32 - && !sandbox.status.as_ref().is_some_and(|status| { - status.conditions.iter().any(|condition| { - condition.r#type == "Ready" && condition.reason == "ContainerExited" - }) - }); + let preserve_infrastructure_error = has_specific_infrastructure_error(sandbox); let status = sandbox.status.get_or_insert_with(SandboxStatus::default); status.main_process_instance_id = instance_id.to_string(); status.exit_code = Some(exit_code); @@ -4797,6 +5316,8 @@ fn apply_main_process_exit(sandbox: &mut Sandbox, instance_id: &str, exit_code: format!("Canonical main process exited with status {exit_code}"), ) }; + set_next_restart_at_ms(status, 0); + set_main_process_started_at_ms(status, 0); upsert_ready_condition( &mut sandbox.status, SandboxCondition { @@ -4810,6 +5331,15 @@ fn apply_main_process_exit(sandbox: &mut Sandbox, instance_id: &str, exit_code: sandbox.set_phase(phase as i32); } +fn has_specific_infrastructure_error(sandbox: &Sandbox) -> bool { + sandbox.phase() == SandboxPhase::Error as i32 + && !sandbox.status.as_ref().is_some_and(|status| { + status.conditions.iter().any(|condition| { + condition.r#type == "Ready" && condition.reason == "ContainerExited" + }) + }) +} + fn is_failed_main_process_result(sandbox: &Sandbox) -> bool { sandbox.phase() == SandboxPhase::Error as i32 && sandbox.status.as_ref().is_some_and(|status| { @@ -4822,12 +5352,157 @@ fn is_failed_main_process_result(sandbox: &Sandbox) -> bool { }) } -/// Connect to an unmanaged remote compute driver that is already listening on -/// `socket_path` and return the acquired endpoint. -/// -/// The gateway does not spawn or own the driver process — the operator is -/// responsible for placing the driver alongside the gateway and granting the -/// gateway uid read/write on the socket. The host portion of the URL is +fn should_restart_main_process(policy: i32, exit_code: i32) -> bool { + match SandboxRestartPolicy::try_from(policy).unwrap_or(SandboxRestartPolicy::Never) { + SandboxRestartPolicy::Always => true, + SandboxRestartPolicy::OnFailure => exit_code != 0, + SandboxRestartPolicy::Unspecified | SandboxRestartPolicy::Never => false, + } +} + +fn is_automatic_restart_status(status: &SandboxStatus) -> bool { + status.restart_count > 0 && status.exit_code.is_some() +} + +fn is_automatic_restart_transition(sandbox: &Sandbox) -> bool { + SandboxPhase::try_from(sandbox.phase()) == Ok(SandboxPhase::Starting) + && sandbox + .status + .as_ref() + .is_some_and(is_automatic_restart_status) +} + +fn restart_delay_ms(restart_count: u32) -> i64 { + if restart_count <= 1 { + return 0; + } + let shift = restart_count.saturating_sub(2).min(31); + RESTART_BACKOFF_BASE_DELAY_MS + .saturating_mul(1_i64.checked_shl(shift).unwrap_or(i64::MAX)) + .min(RESTART_MAX_DELAY_MS) +} + +fn is_terminal_delivery_pending(status: &SandboxStatus) -> bool { + status.conditions.iter().any(|condition| { + condition.r#type == "Ready" && condition.reason == "MainProcessExitDraining" + }) +} + +fn restart_is_due(status: &SandboxStatus, now_ms: i64) -> bool { + let deadline_ms = next_restart_at_ms(status); + if deadline_ms > now_ms { + return false; + } + if !is_terminal_delivery_pending(status) { + return true; + } + let exit_time_ms = deadline_ms.saturating_sub(restart_delay_ms(status.restart_count)); + now_ms >= exit_time_ms.saturating_add(RESTART_TERMINAL_DELIVERY_GRACE_MS) +} + +fn restart_ready_condition(restart_count: u32, delay_ms: i64) -> SandboxCondition { + let (reason, message) = if delay_ms == 0 { + ( + "MainProcessRestartScheduled", + format!("Canonical main process exited; restart {restart_count} scheduled immediately"), + ) + } else { + let delay_seconds = delay_ms / 1_000; + let delay_unit = if delay_seconds == 1 { + "second" + } else { + "seconds" + }; + ( + "MainProcessRestartBackoff", + format!( + "Canonical main process exited; restart {restart_count} scheduled in {delay_seconds} {delay_unit}" + ), + ) + }; + SandboxCondition { + r#type: "Ready".to_string(), + status: "False".to_string(), + reason: reason.to_string(), + message, + transition_time: None, + } +} + +fn next_restart_at_ms(status: &SandboxStatus) -> i64 { + status + .next_restart_time + .as_ref() + .and_then(|time| openshell_core::time::timestamp_to_millis(time).ok()) + .unwrap_or_default() +} + +fn set_next_restart_at_ms(status: &mut SandboxStatus, millis: i64) { + status.next_restart_time = (millis > 0) + .then(|| openshell_core::time::timestamp_from_millis(millis).ok()) + .flatten(); +} + +fn main_process_started_at_ms(status: &SandboxStatus) -> i64 { + status + .main_process_started_time + .as_ref() + .and_then(|time| openshell_core::time::timestamp_to_millis(time).ok()) + .unwrap_or_default() +} + +fn set_main_process_started_at_ms(status: &mut SandboxStatus, millis: i64) { + status.main_process_started_time = (millis > 0) + .then(|| openshell_core::time::timestamp_from_millis(millis).ok()) + .flatten(); +} + +#[cfg(test)] +fn restart_timestamp(millis: i64) -> Option { + (millis > 0) + .then(|| openshell_core::time::timestamp_from_millis(millis).ok()) + .flatten() +} + +fn apply_main_process_restart( + sandbox: &mut Sandbox, + instance_id: &str, + exit_code: i32, + now_ms: i64, +) { + let status = sandbox.status.get_or_insert_with(SandboxStatus::default); + let stable_run = main_process_started_at_ms(status) > 0 + && now_ms.saturating_sub(main_process_started_at_ms(status)) >= RESTART_STABILITY_WINDOW_MS; + status.restart_count = if stable_run || status.restart_count == 0 { + 1 + } else { + status.restart_count.saturating_add(1) + }; + let delay_ms = restart_delay_ms(status.restart_count); + status.main_process_instance_id = instance_id.to_string(); + status.exit_code = Some(exit_code); + set_main_process_started_at_ms(status, 0); + set_next_restart_at_ms(status, now_ms.saturating_add(delay_ms)); + upsert_ready_condition( + &mut sandbox.status, + SandboxCondition { + r#type: "Ready".to_string(), + status: "False".to_string(), + reason: "MainProcessExitDraining".to_string(), + message: "Canonical main process exited; waiting for terminal delivery before restart" + .to_string(), + transition_time: None, + }, + ); + sandbox.set_phase(SandboxPhase::Starting as i32); +} + +/// Connect to an unmanaged remote compute driver that is already listening on +/// `socket_path` and return the acquired endpoint. +/// +/// The gateway does not spawn or own the driver process — the operator is +/// responsible for placing the driver alongside the gateway and granting the +/// gateway uid read/write on the socket. The host portion of the URL is /// ignored because the connector resolves to the UDS rather than DNS. #[cfg(unix)] pub async fn connect_remote_compute_driver( @@ -5300,6 +5975,9 @@ fn public_status_from_driver( configuration_admission: None, configuration_activated: None, provisioning: None, + restart_count: 0, + next_restart_time: None, + main_process_started_time: None, } } @@ -5365,7 +6043,24 @@ fn apply_driver_snapshot( }, ); + // Driver status has no canonical-main fields. Preserve the gateway-owned + // process generation and restart controller state across every snapshot. + if let (Some(old_status), Some(new_status)) = (sandbox.status.as_ref(), status.as_mut()) { + new_status + .main_process_instance_id + .clone_from(&old_status.main_process_instance_id); + new_status.exit_code = old_status.exit_code; + new_status.restart_count = old_status.restart_count; + new_status.next_restart_time = old_status.next_restart_time; + new_status.main_process_started_time = old_status.main_process_started_time; + } + phase = match old_phase { + // Backoff and replacement are gateway-owned. Driver snapshots can + // describe the old runtime stopping but cannot cancel restart intent. + SandboxPhase::Starting if is_automatic_restart_transition(sandbox) => { + SandboxPhase::Starting + } // SIGTERM-driven runtime exits are reported as a runtime restart by // Docker and Podman. While an explicit stop owns this durable // transition, preserve Stopping so the stop result chooses whether @@ -5776,9 +6471,9 @@ fn rewrite_user_facing_conditions(status: &mut Option, spec: Opti } /// Phases for which a sandbox should have a running compute resource. -/// `Deleting` and `Error` are intentionally excluded: deletion is in -/// progress, or the sandbox has already failed and should not be -/// silently revived. `Unspecified` is included because it is the proto +/// Terminal and lifecycle-controller-owned phases are intentionally excluded: +/// they must not be silently revived or bypass a persisted transition/backoff. +/// `Unspecified` is included because it is the proto /// default value; persisted rows with that value should be reconciled /// from the live driver state rather than skipped forever. fn sandbox_phase_should_be_running(phase: SandboxPhase) -> bool { @@ -5815,6 +6510,9 @@ fn apply_lifecycle_phase(sandbox: &mut Sandbox, phase: SandboxPhase, reason: &st // Retain the previous instance id as a tombstone until the restarted // supervisor registers its new id. status.exit_code = None; + status.restart_count = 0; + set_next_restart_at_ms(status, 0); + set_main_process_started_at_ms(status, 0); if phase == SandboxPhase::Starting { status.provisioning = Some(provisioning_deadline::new_record( openshell_core::time::now_ms(), @@ -6201,6 +6899,8 @@ pub fn new_test_runtime_with_driver( lifecycle_gates: Arc::new(LifecycleGateRegistry::default()), replica_id: "test-replica".to_string(), rootfs_tar_staging: Arc::new(rootfs_tar::RootfsTarStagingRegistry::disabled()), + restart_authority: Arc::new(OnceLock::new()), + restart_notify: Arc::new(Notify::new()), } } @@ -7248,6 +7948,8 @@ mod tests { lifecycle_gates: Arc::new(LifecycleGateRegistry::default()), replica_id: "test-replica".to_string(), rootfs_tar_staging: Arc::new(rootfs_tar::RootfsTarStagingRegistry::disabled()), + restart_authority: Arc::new(OnceLock::new()), + restart_notify: Arc::new(Notify::new()), } } @@ -7893,6 +8595,7 @@ mod tests { #[test] fn main_process_exit_preserves_specific_infrastructure_error() { let mut sandbox = error_sandbox_record("sb-1", "sandbox-a", "BackendResourceMissing"); + assert!(has_specific_infrastructure_error(&sandbox)); apply_main_process_exit(&mut sandbox, "instance-1", 0); assert_eq!( @@ -7907,6 +8610,595 @@ mod tests { })); } + #[test] + fn restart_policy_matches_kubernetes_exit_semantics() { + assert!(!should_restart_main_process( + SandboxRestartPolicy::Never as i32, + 1 + )); + assert!(!should_restart_main_process( + SandboxRestartPolicy::OnFailure as i32, + 0 + )); + assert!(should_restart_main_process( + SandboxRestartPolicy::OnFailure as i32, + 1 + )); + assert!(should_restart_main_process( + SandboxRestartPolicy::Always as i32, + 0 + )); + } + + #[test] + fn restart_backoff_starts_after_first_exit_and_caps_at_three_minutes() { + assert_eq!(restart_delay_ms(0), 0); + assert_eq!(restart_delay_ms(1), 0); + assert_eq!(restart_delay_ms(2), 1_000); + assert_eq!(restart_delay_ms(3), 2_000); + assert_eq!(restart_delay_ms(4), 4_000); + assert_eq!(restart_delay_ms(8), 64_000); + assert_eq!(restart_delay_ms(9), 128_000); + assert_eq!(restart_delay_ms(10), 180_000); + assert_eq!(restart_delay_ms(20), 180_000); + } + + #[test] + fn unfinished_terminal_delivery_gets_bounded_restart_fallback() { + let now_ms = openshell_core::time::now_ms(); + let status = SandboxStatus { + restart_count: 2, + next_restart_time: restart_timestamp(now_ms + restart_delay_ms(2)), + conditions: vec![SandboxCondition { + r#type: "Ready".into(), + reason: "MainProcessExitDraining".into(), + ..Default::default() + }], + ..Default::default() + }; + + assert!(!restart_is_due(&status, now_ms + restart_delay_ms(2))); + assert!(!restart_is_due( + &status, + now_ms + RESTART_TERMINAL_DELIVERY_GRACE_MS - 1 + )); + assert!(restart_is_due( + &status, + now_ms + RESTART_TERMINAL_DELIVERY_GRACE_MS + )); + } + + #[tokio::test] + async fn always_policy_schedules_zero_exit_once_and_new_supervisor_becomes_ready() { + let runtime = test_runtime(Arc::new(TestDriver::default())).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Ready); + sandbox.spec = Some(SandboxSpec { + restart_policy: SandboxRestartPolicy::Always as i32, + ..Default::default() + }); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Ready as i32, + main_process_instance_id: "instance-1".into(), + main_process_started_time: restart_timestamp(openshell_core::time::now_ms()), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + let before_exit_ms = openshell_core::time::now_ms(); + runtime + .main_process_exited("sb-1", "instance-1", 0) + .await + .unwrap(); + runtime + .main_process_exited("sb-1", "instance-1", 0) + .await + .unwrap(); + + let restarting = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(restarting.phase(), SandboxPhase::Starting as i32); + let status = restarting.status.unwrap(); + assert_eq!(status.restart_count, 1); + assert_eq!(status.exit_code, Some(0)); + assert!(next_restart_at_ms(&status) >= before_exit_ms); + assert!(next_restart_at_ms(&status) <= openshell_core::time::now_ms()); + assert!(!restart_is_due(&status, openshell_core::time::now_ms())); + assert!(status.conditions.iter().any(|condition| { + condition.r#type == "Ready" && condition.reason == "MainProcessExitDraining" + })); + + runtime + .finalize_main_process_exit("sb-1", "instance-1") + .await + .unwrap(); + let finalized = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + let status = finalized.status.unwrap(); + assert!(restart_is_due(&status, openshell_core::time::now_ms())); + assert_eq!( + status + .conditions + .iter() + .find(|condition| condition.r#type == "Ready") + .map(|condition| condition.message.as_str()), + Some("Canonical main process exited; restart 1 scheduled immediately") + ); + assert!(status.conditions.iter().any(|condition| { + condition.r#type == "Ready" && condition.reason == "MainProcessRestartScheduled" + })); + tokio::time::timeout( + Duration::from_millis(100), + runtime.restart_notify.notified(), + ) + .await + .expect("finalized restart should wake the local restart worker"); + + runtime + .supervisor_session_connected("sb-1", "instance-2") + .await + .unwrap(); + let ready = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(ready.phase(), SandboxPhase::Ready as i32); + let status = ready.status.unwrap(); + assert_eq!(status.main_process_instance_id, "instance-2"); + assert_eq!(status.exit_code, None); + assert_eq!(next_restart_at_ms(&status), 0); + assert!(main_process_started_at_ms(&status) > 0); + } + + #[tokio::test] + async fn replacement_exit_before_supervisor_connect_uses_new_instance_result() { + let runtime = test_runtime(Arc::new(TestDriver::default())).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.spec = Some(SandboxSpec { + restart_policy: SandboxRestartPolicy::OnFailure as i32, + ..Default::default() + }); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + main_process_instance_id: "instance-1".into(), + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() + 120_000), + conditions: vec![SandboxCondition { + r#type: "Ready".into(), + status: "False".into(), + reason: "SandboxRestarting".into(), + ..Default::default() + }], + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + runtime + .main_process_exited("sb-1", "instance-2", 0) + .await + .unwrap(); + + let stored = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(stored.phase(), SandboxPhase::Completed as i32); + let status = stored.status.unwrap(); + assert_eq!(status.main_process_instance_id, "instance-2"); + assert_eq!(status.exit_code, Some(0)); + assert_eq!(next_restart_at_ms(&status), 0); + } + + #[tokio::test] + async fn restarting_sandbox_rejects_previous_supervisor_instance() { + let runtime = test_runtime(Arc::new(TestDriver::default())).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + main_process_instance_id: "instance-1".into(), + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() + 10_000), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + let error = runtime + .supervisor_session_connected("sb-1", "instance-1") + .await + .unwrap_err(); + + assert!(error.contains("requires a fresh supervisor instance")); + let stored = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(stored.phase(), SandboxPhase::Starting as i32); + assert_eq!( + stored.status.unwrap().main_process_instance_id, + "instance-1" + ); + } + + #[tokio::test] + async fn same_main_reconnect_preserves_process_start_time() { + let runtime = test_runtime(Arc::new(TestDriver::default())).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Ready); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Ready as i32, + main_process_instance_id: "instance-1".into(), + main_process_started_time: restart_timestamp(12_345), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + runtime + .supervisor_session_connected("sb-1", "instance-1") + .await + .unwrap(); + + let reconnected = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!( + main_process_started_at_ms(&reconnected.status.unwrap()), + 12_345 + ); + } + + #[tokio::test] + async fn on_failure_policy_keeps_zero_exit_terminal() { + let runtime = test_runtime(Arc::new(TestDriver::default())).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Ready); + sandbox.spec = Some(SandboxSpec { + restart_policy: SandboxRestartPolicy::OnFailure as i32, + ..Default::default() + }); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Ready as i32, + main_process_instance_id: "instance-1".into(), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + runtime + .main_process_exited("sb-1", "instance-1", 0) + .await + .unwrap(); + + let stored = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(stored.phase(), SandboxPhase::Completed as i32); + assert_eq!(stored.status.unwrap().restart_count, 0); + } + + #[test] + fn stable_main_run_resets_restart_backoff_count() { + let now_ms = openshell_core::time::now_ms(); + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Ready); + sandbox.status = Some(SandboxStatus { + restart_count: 5, + main_process_started_time: restart_timestamp(now_ms - RESTART_STABILITY_WINDOW_MS), + ..Default::default() + }); + + apply_main_process_restart(&mut sandbox, "instance-6", 9, now_ms); + + let status = sandbox.status.unwrap(); + assert_eq!(status.restart_count, 1); + assert_eq!(next_restart_at_ms(&status), now_ms); + } + + #[test] + fn short_main_run_keeps_restart_backoff_count() { + let now_ms = openshell_core::time::now_ms(); + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Ready); + sandbox.status = Some(SandboxStatus { + restart_count: 5, + main_process_started_time: restart_timestamp(now_ms - RESTART_STABILITY_WINDOW_MS + 1), + ..Default::default() + }); + + apply_main_process_restart(&mut sandbox, "instance-6", 9, now_ms); + + let status = sandbox.status.unwrap(); + assert_eq!(status.restart_count, 6); + assert_eq!(next_restart_at_ms(&status), now_ms + 16_000); + } + + #[tokio::test] + async fn due_restart_stops_and_starts_driver_without_exposing_stopped_phase() { + let driver = ControlledDriver::new(); + let runtime = test_runtime(driver.clone()).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.spec = Some(SandboxSpec { + restart_policy: SandboxRestartPolicy::OnFailure as i32, + ..Default::default() + }); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + main_process_instance_id: "instance-1".into(), + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() - 1), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + runtime.restart_sandbox_runtime("sb-1").await.unwrap(); + + assert_eq!(driver.stop_calls(), 1); + assert_eq!(driver.start_calls(), 1); + let stored = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(stored.phase(), SandboxPhase::Starting as i32); + let status = stored.status.unwrap(); + assert_eq!(status.exit_code, Some(9)); + assert!(next_restart_at_ms(&status) > openshell_core::time::now_ms()); + assert!(status.conditions.iter().any(|condition| { + condition.r#type == "Ready" && condition.reason == "SandboxRestarting" + })); + } + + #[tokio::test] + async fn restart_readiness_deadline_starts_after_driver_stop() { + let driver = ControlledDriver::new(); + driver.block_stop(); + let runtime = test_runtime(driver.clone()).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() - 1), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + let restart_runtime = runtime.clone(); + let restart = + tokio::spawn(async move { restart_runtime.restart_sandbox_runtime("sb-1").await }); + driver.stop_started.notified().await; + + let claimed = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(next_restart_at_ms(&claimed.status.unwrap()), 0); + + driver.release_stop(); + restart.await.unwrap().unwrap(); + let started = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert!(next_restart_at_ms(&started.status.unwrap()) > openshell_core::time::now_ms()); + } + + #[tokio::test] + async fn claimed_restart_replays_stop_before_start_after_gateway_recovery() { + let driver = ControlledDriver::new(); + let runtime = test_runtime(driver.clone()).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + main_process_instance_id: "previous-instance".into(), + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(0), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + runtime.restart_due_sandboxes().await.unwrap(); + + assert_eq!(driver.stop_calls(), 1); + assert_eq!(driver.start_calls(), 1); + let stored = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert!(next_restart_at_ms(&stored.status.unwrap()) > openshell_core::time::now_ms()); + } + + #[tokio::test] + async fn concurrent_stop_transition_prevents_replacement_start() { + let driver = ControlledDriver::new(); + driver.block_stop(); + let runtime = test_runtime(driver.clone()).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() - 1), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + let restart_runtime = runtime.clone(); + let restart = + tokio::spawn(async move { restart_runtime.restart_sandbox_runtime("sb-1").await }); + driver.stop_started.notified().await; + let claimed = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + runtime + .store + .update_message_cas::( + "sb-1", + sandbox_resource_version(&claimed), + |sandbox| sandbox.set_phase(SandboxPhase::Stopping as i32), + ) + .await + .unwrap(); + + driver.release_stop(); + restart.await.unwrap().unwrap(); + assert_eq!(driver.start_calls(), 0); + } + + #[tokio::test] + async fn concurrent_delete_after_replacement_start_is_reasserted() { + let driver = ControlledDriver::new(); + driver.block_start(); + let runtime = test_runtime(driver.clone()).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() - 1), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + let restart_runtime = runtime.clone(); + let restart = + tokio::spawn(async move { restart_runtime.restart_sandbox_runtime("sb-1").await }); + driver.start_started.notified().await; + let armed = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + runtime + .store + .update_message_cas::("sb-1", sandbox_resource_version(&armed), |sandbox| { + sandbox.set_phase(SandboxPhase::Deleting as i32); + }) + .await + .unwrap(); + + driver.release_start(); + restart.await.unwrap().unwrap(); + assert_eq!(driver.delete_calls(), 1); + } + + #[tokio::test] + async fn retryable_restart_driver_failure_reschedules_attempt() { + let driver = ControlledDriver::new(); + driver.set_stop_outcome(ControlledLifecycleOutcome::Error("transport unavailable")); + let runtime = test_runtime(driver).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() - 1), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + let before_failure_ms = openshell_core::time::now_ms(); + let error = runtime.restart_sandbox_runtime("sb-1").await.unwrap_err(); + + assert!(error.contains("transport unavailable")); + let stored = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(stored.phase(), SandboxPhase::Starting as i32); + let status = stored.status.unwrap(); + assert!( + next_restart_at_ms(&status) + >= before_failure_ms.saturating_add(RESTART_DRIVER_RETRY_DELAY_MS) + ); + assert!(status.conditions.iter().any(|condition| { + condition.r#type == "Ready" && condition.reason == "SandboxRestartFailed" + })); + } + + #[tokio::test] + async fn missing_runtime_during_restart_is_terminal() { + let driver = ControlledDriver::new(); + driver.set_stop_outcome(ControlledLifecycleOutcome::NotFound); + let runtime = test_runtime(driver).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + exit_code: Some(9), + restart_count: 1, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() - 1), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + runtime.restart_sandbox_runtime("sb-1").await.unwrap_err(); + + let stored = runtime + .store + .get_message::("sb-1") + .await + .unwrap() + .unwrap(); + assert_eq!(stored.phase(), SandboxPhase::Error as i32); + let status = stored.status.unwrap(); + assert_eq!(next_restart_at_ms(&status), 0); + assert_eq!(status.exit_code, Some(9)); + } + + #[tokio::test] + async fn manual_stop_cancels_pending_restart() { + let driver = ControlledDriver::new(); + let runtime = test_runtime(driver.clone()).await; + let mut sandbox = sandbox_record("sb-1", "sandbox-a", SandboxPhase::Starting); + sandbox.status = Some(SandboxStatus { + phase: SandboxPhase::Starting as i32, + exit_code: Some(9), + restart_count: 2, + next_restart_time: restart_timestamp(openshell_core::time::now_ms() + 60_000), + ..Default::default() + }); + runtime.store.put_message(&sandbox).await.unwrap(); + + let stopped = runtime.stop_sandbox("default", "sandbox-a").await.unwrap(); + + assert_eq!(stopped.phase(), SandboxPhase::Stopped as i32); + let status = stopped.status.unwrap(); + assert_eq!(status.exit_code, None); + assert_eq!(status.restart_count, 0); + assert_eq!(next_restart_at_ms(&status), 0); + assert_eq!(driver.stop_calls(), 1); + } + #[tokio::test] async fn stale_main_process_exit_is_acknowledged_without_replacing_active_instance() { let runtime = test_runtime(Arc::new(TestDriver::default())).await; @@ -8219,6 +9511,7 @@ mod tests { name: "web".to_string(), target_port: 8080, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, } } diff --git a/crates/openshell-server/src/config_file.rs b/crates/openshell-server/src/config_file.rs index 87180766ab..de3f3df6bd 100644 --- a/crates/openshell-server/src/config_file.rs +++ b/crates/openshell-server/src/config_file.rs @@ -167,6 +167,8 @@ pub struct GatewayFileSection { pub gateway_jwt: Option, #[serde(default)] pub otlp: Option, + #[serde(default)] + pub ocsf_log: Option, // ── Disallowed-in-file fields ──────────────────────────────────────── // @@ -198,6 +200,83 @@ pub struct GatewayTlsFileConfig { pub external_server_names: Vec, } +#[derive(Debug, Clone, Copy, Default, PartialEq, Eq, Serialize, Deserialize)] +#[serde(rename_all = "snake_case")] +pub enum OcsfLogRotation { + Never, + #[default] + Daily, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)] +pub enum OcsfSchemaVersion { + #[serde(rename = "1.1")] + V1_1, + #[serde(rename = "1.3")] + V1_3, +} + +impl OcsfSchemaVersion { + pub const fn as_str(self) -> &'static str { + match self { + Self::V1_1 => "1.1", + Self::V1_3 => "1.3", + } + } +} + +#[derive(Debug, Clone, Serialize, Deserialize)] +#[serde(try_from = "RawOcsfLogConfig")] +pub struct OcsfLogConfig { + pub path: PathBuf, + #[serde(skip_serializing_if = "Option::is_none")] + pub schema_version: Option, + pub rotation: OcsfLogRotation, + #[serde(skip_serializing_if = "Option::is_none")] + pub max_files: Option, + pub queue_capacity: std::num::NonZeroUsize, + pub queue_max_bytes: std::num::NonZeroUsize, +} + +#[derive(Deserialize)] +#[serde(deny_unknown_fields)] +struct RawOcsfLogConfig { + path: PathBuf, + schema_version: Option, + #[serde(default)] + rotation: OcsfLogRotation, + max_files: Option, + queue_capacity: Option, + queue_max_bytes: Option, +} + +impl TryFrom for OcsfLogConfig { + type Error = &'static str; + + fn try_from(raw: RawOcsfLogConfig) -> Result { + if raw.path.as_os_str().is_empty() { + return Err("ocsf_log.path must not be empty"); + } + if raw.rotation == OcsfLogRotation::Never && raw.max_files.is_some() { + return Err("ocsf_log.max_files requires daily rotation"); + } + Ok(Self { + path: raw.path, + schema_version: raw.schema_version, + rotation: raw.rotation, + max_files: (raw.rotation == OcsfLogRotation::Daily).then(|| { + raw.max_files + .unwrap_or(std::num::NonZeroUsize::new(7).unwrap()) + }), + queue_capacity: raw + .queue_capacity + .unwrap_or(std::num::NonZeroUsize::new(10_000).unwrap()), + queue_max_bytes: raw + .queue_max_bytes + .unwrap_or(std::num::NonZeroUsize::new(16 * 1024 * 1024).unwrap()), + }) + } +} /// `[openshell.gateway.otlp]` section. /// /// Presence of this table enables OTLP export; there is no `enabled` flag. @@ -848,6 +927,60 @@ service_name = "openshell-gateway-dev" assert_eq!(otlp.service_name.as_deref(), Some("openshell-gateway-dev")); } + #[test] + fn gateway_accepts_a_single_ocsf_log_destination() { + let tmp = write_tmp("[openshell.gateway.ocsf_log]\npath = 'events.jsonl'\n"); + let config = load(tmp.path()) + .unwrap() + .openshell + .gateway + .ocsf_log + .unwrap(); + assert_eq!(config.rotation, OcsfLogRotation::Daily); + assert_eq!(config.schema_version, None); + assert_eq!(config.max_files.unwrap().get(), 7); + assert_eq!(config.queue_capacity.get(), 10_000); + assert_eq!(config.queue_max_bytes.get(), 16 * 1024 * 1024); + assert!(ConfigFile::default().openshell.gateway.ocsf_log.is_none()); + + let tmp = write_tmp( + "[openshell.gateway.ocsf_log]\npath = 'events.jsonl'\nschema_version = '1.3'\n", + ); + assert_eq!( + load(tmp.path()) + .unwrap() + .openshell + .gateway + .ocsf_log + .unwrap() + .schema_version, + Some(OcsfSchemaVersion::V1_3) + ); + } + + #[test] + fn ocsf_log_rejects_invalid_and_unshipped_options() { + for settings in [ + "", + "path = ''", + "path = 'log'\nqueue_capacity = 0", + "path = 'log'\nqueue_max_bytes = 0", + "path = 'log'\nmax_files = 0", + "path = 'log'\nrotation = 'hourly'", + "path = 'log'\nschema_version = ''", + "path = 'log'\nschema_version = '1.8'", + "path = 'log'\nkind = 'jsonl'", + "path = 'log'\nrotation = 'never'\nmax_files = 7", + ] { + let tmp = write_tmp(&format!("[openshell.gateway.ocsf_log]\n{settings}\n")); + assert!(load(tmp.path()).is_err(), "accepted {settings}"); + } + let tmp = write_tmp("[openshell.gateway.ocsf_log]\npath = 'log'\nrotation = 'never'\n"); + let config = load(tmp.path()).unwrap(); + let encoded = toml::to_string(&config).unwrap(); + assert!(toml::from_str::(&encoded).is_ok()); + } + #[test] fn otlp_config_requires_only_endpoint() { let toml = r#" diff --git a/crates/openshell-server/src/gateway_ocsf.rs b/crates/openshell-server/src/gateway_ocsf.rs new file mode 100644 index 0000000000..abc823f415 --- /dev/null +++ b/crates/openshell-server/src/gateway_ocsf.rs @@ -0,0 +1,90 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +//! Process-wide OCSF identity for gateway-origin events. +//! +//! Gateway events are emitted from places with no access to the server config +//! (the TLS reload watcher, the service router), so the identity is resolved +//! once at startup rather than threaded through all of them. + +use std::sync::OnceLock; + +use openshell_ocsf::{EventContext, EventOrigin}; + +/// Identity shared by every gateway-origin OCSF event. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct GatewayIdentity { + /// Operator-assigned gateway name, shared across replicas of one install. + pub name: String, + /// Per-replica hostname (the pod name under Kubernetes). + pub hostname: String, +} + +static IDENTITY: OnceLock = OnceLock::new(); + +/// Initialise the process-wide gateway identity. +/// +/// Returns `false` if it was already set; the caller may log and continue. +pub fn set_identity(identity: GatewayIdentity) -> bool { + IDENTITY.set(identity).is_ok() +} + +/// Return the gateway identity, falling back to placeholders when unset (in +/// tests, and in any code path that runs before startup completes). +#[must_use] +pub fn identity() -> GatewayIdentity { + IDENTITY.get().cloned().unwrap_or_else(|| GatewayIdentity { + name: openshell_core::config::DEFAULT_GATEWAY_NAME.to_string(), + hostname: "openshell-gateway".to_string(), + }) +} + +/// Build the OCSF context for a gateway-origin event. +/// +/// `sandbox_id` and `sandbox_name` describe the sandbox the event is *about*, +/// and may be empty. The emitting device is always the gateway. +#[must_use] +pub fn context(sandbox_id: &str, sandbox_name: &str) -> EventContext { + let identity = identity(); + EventContext { + sandbox_id: sandbox_id.to_string(), + sandbox_name: sandbox_name.to_string(), + container_image: String::new(), + hostname: identity.hostname, + product_version: openshell_core::VERSION.to_string(), + proxy_ip: std::net::IpAddr::V4(std::net::Ipv4Addr::LOCALHOST), + proxy_port: 0, + origin: EventOrigin::Gateway { + name: identity.name, + }, + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn context_marks_events_as_gateway_origin() { + let ctx = context("sb-1", "agent-01"); + + assert!(matches!(ctx.origin, EventOrigin::Gateway { .. })); + assert_eq!(ctx.sandbox_id, "sb-1"); + assert_eq!(ctx.sandbox_name, "agent-01"); + } + + #[test] + fn gateway_context_produces_gateway_product_and_no_container() { + let ctx = context("", ""); + + assert_eq!(ctx.metadata(&[]).product.name, "OpenShell Gateway"); + assert!(ctx.container().is_none()); + } + + #[test] + fn identity_falls_back_when_unset() { + let identity = identity(); + assert!(!identity.name.is_empty()); + assert!(!identity.hostname.is_empty()); + } +} diff --git a/crates/openshell-server/src/grpc/mutation_replay/tests.rs b/crates/openshell-server/src/grpc/mutation_replay/tests.rs index 28dfc688f2..7b784213b6 100644 --- a/crates/openshell-server/src/grpc/mutation_replay/tests.rs +++ b/crates/openshell-server/src/grpc/mutation_replay/tests.rs @@ -7,7 +7,8 @@ use openshell_core::proto::datamodel::v1::ObjectMeta; use openshell_core::proto::open_shell_server::OpenShell; use openshell_core::proto::{ CreateSandboxRequest, SandboxServiceExposure, SandboxSpec, SandboxWorkloadConfig, - SandboxWorkloadTemplate, SandboxWorkloadTemplateSpec, WorkspaceMember, WorkspaceRole, + SandboxWorkloadTemplate, SandboxWorkloadTemplateSpec, ServiceAuthorizationMode, + WorkspaceMember, WorkspaceRole, }; use openshell_core::rpc_error::StatusExt; use std::collections::HashMap; @@ -113,6 +114,7 @@ async fn create_sandbox_replay_preserves_service_urls() { service_exposures: vec![SandboxServiceExposure { service: "web".into(), target_port: 8080, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }], request_id: uuid::Uuid::new_v4().to_string(), ..Default::default() @@ -124,13 +126,26 @@ async fn create_sandbox_replay_preserves_service_urls() { .unwrap() .into_inner(); let replay = service - .create_sandbox(authed_request(request)) + .create_sandbox(authed_request(request.clone())) .await .unwrap(); assert_eq!(replay.metadata().get("openshell-replayed").unwrap(), "true"); assert_eq!(replay.get_ref().service_urls, original.service_urls); assert_eq!(replay.into_inner(), original); + + let mut changed_authorization = request; + changed_authorization.service_exposures[0].authorization_mode = + ServiceAuthorizationMode::BearerPassthrough as i32; + assert_eq!( + reason( + &service + .create_sandbox(authed_request(changed_authorization)) + .await + .unwrap_err() + ), + "REQUEST_ID_PAYLOAD_MISMATCH" + ); } async fn exercise_backend(url: &str) { diff --git a/crates/openshell-server/src/grpc/policy.rs b/crates/openshell-server/src/grpc/policy.rs index 74fac49de8..2d92a64ddc 100644 --- a/crates/openshell-server/src/grpc/policy.rs +++ b/crates/openshell-server/src/grpc/policy.rs @@ -67,14 +67,12 @@ use openshell_core::telemetry::{ LifecycleOperation, LifecycleResource, PolicyDecisionOperation, TelemetryOutcome, }; use openshell_core::{ - GetResourceVersion, VERSION, + GetResourceVersion, endpoint_path::EndpointPathPattern, host_pattern::{host_matches, host_patterns_overlap}, settings::{self, SettingValueKind}, }; -use openshell_ocsf::{ - ConfigStateChangeBuilder, EventContext, OCSF_TARGET, OcsfEvent, SeverityId, StateId, StatusId, -}; +use openshell_ocsf::{ConfigStateChangeBuilder, OcsfEvent, SeverityId, StateId, StatusId}; use openshell_policy::{ L7BinaryScope, L7RuleTarget, PolicyMergeOp, ProviderPolicyLayer, canonicalize_advisor_add_rule, compose_effective_policy, merge_policy, policy_covers_rule, serialize_sandbox_policy, @@ -92,7 +90,7 @@ use openshell_prover::{ use prost::Message; use sha2::{Digest, Sha256}; use std::collections::{BTreeMap, HashMap, HashSet}; -use std::net::{IpAddr, Ipv4Addr}; +use std::net::IpAddr; use std::sync::Arc; use tonic::{Request, Response, Status}; use tracing::{debug, info, warn}; @@ -269,7 +267,7 @@ fn emit_gateway_policy_audit_log( version: i64, policy_hash: &str, ) { - let message = build_gateway_policy_audit_message( + openshell_ocsf::ocsf_emit!(build_gateway_policy_audit_event( sandbox_id, sandbox_name, state_label, @@ -277,12 +275,7 @@ fn emit_gateway_policy_audit_log( version, policy_hash, &[], - ); - info!( - target: OCSF_TARGET, - sandbox_id = %sandbox_id, - message = %message - ); + )); } /// Emit a `CONFIG:APPROVED` audit event for an auto-approval — same event @@ -307,7 +300,7 @@ fn emit_gateway_policy_auto_approve_audit_log( ("prover_delta", "empty".to_string()), ("resolved_from", resolved_from.to_string()), ]; - let message = build_gateway_policy_audit_message( + openshell_ocsf::ocsf_emit!(build_gateway_policy_audit_event( sandbox_id, sandbox_name, "approved", @@ -315,15 +308,10 @@ fn emit_gateway_policy_auto_approve_audit_log( version, policy_hash, &extra, - ); - info!( - target: OCSF_TARGET, - sandbox_id = %sandbox_id, - message = %message - ); + )); } -fn build_gateway_policy_audit_message( +fn build_gateway_policy_audit_event( sandbox_id: &str, sandbox_name: &str, state_label: &str, @@ -331,16 +319,8 @@ fn build_gateway_policy_audit_message( version: i64, policy_hash: &str, extra_fields: &[(&str, String)], -) -> String { - let ctx = EventContext { - sandbox_id: sandbox_id.to_string(), - sandbox_name: sandbox_name.to_string(), - container_image: "openshell/gateway".to_string(), - hostname: "openshell-gateway".to_string(), - product_version: VERSION.to_string(), - proxy_ip: IpAddr::V4(Ipv4Addr::LOCALHOST), - proxy_port: 0, - }; +) -> OcsfEvent { + let ctx = crate::gateway_ocsf::context(sandbox_id, sandbox_name); let mut builder = ConfigStateChangeBuilder::new(&ctx) .state(StateId::Other, state_label) .severity(SeverityId::Informational) @@ -355,8 +335,7 @@ fn build_gateway_policy_audit_message( for (key, value) in extra_fields { builder = builder.unmapped(key, value.clone()); } - let event: OcsfEvent = builder.build(); - event.format_shorthand() + builder.build() } fn summarize_cli_policy_merge_op(operation: &PolicyMergeOp) -> String { @@ -3367,10 +3346,10 @@ pub(super) async fn handle_get_sandbox_provider_environment( .await .map_err(|e| Status::internal(format!("fetch sandbox failed: {e}")))? .ok_or_else(|| Status::not_found("sandbox not found"))?; - Ok(Response::new( + let environment = load_sandbox_provider_environment(state, &sandbox, supports_static_credential_bindings) - .await?, - )) + .await?; + Ok(Response::new(environment)) } /// Materialize a privileged provider snapshot after the caller has authorized @@ -3497,6 +3476,7 @@ pub(super) async fn load_sandbox_provider_environment( .collect(); Ok(GetSandboxProviderEnvironmentResponse { environment: provider_environment.environment, + files: provider_environment.files, provider_env_revision, credential_expiration_times, dynamic_credentials: provider_environment.dynamic_credentials, @@ -11292,6 +11272,7 @@ mod tests { deletion_time: None, }), profile: Some(openshell_core::proto::ProviderProfile { + files: Vec::new(), id: "generic".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -11385,6 +11366,7 @@ mod tests { deletion_time: None, }), profile: Some(openshell_core::proto::ProviderProfile { + files: Vec::new(), id: "custom-api".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -11456,6 +11438,7 @@ mod tests { deletion_time: None, }), profile: Some(openshell_core::proto::ProviderProfile { + files: Vec::new(), id: "custom-api".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -12613,6 +12596,7 @@ mod tests { deletion_time: None, }), profile: Some(ProviderProfile { + files: Vec::new(), id: "custom-policy".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -12836,6 +12820,89 @@ mod tests { assert_eq!(v2_env.get("GITHUB_TOKEN"), Some(&"ghp-test".to_string())); } + #[tokio::test] + async fn provider_files_do_not_block_legacy_provider_environment_requests() { + use openshell_core::proto::{ + GetSandboxProviderEnvironmentRequest, ProviderProfile, ProviderProfileCategory, + ProviderProfileFile, + }; + + let state = test_server_state().await; + state + .store + .put_message(&StoredProviderProfile { + metadata: Some(openshell_core::proto::datamodel::v1::ObjectMeta { + id: "profile-config-only".to_string(), + name: "config-only".to_string(), + workspace: "default".to_string(), + ..Default::default() + }), + profile: Some(ProviderProfile { + id: "config-only".to_string(), + display_name: "Config only".to_string(), + category: ProviderProfileCategory::Other as i32, + files: vec![ProviderProfileFile { + path: "client.toml".to_string(), + content: "endpoint = '{{config.endpoint}}'".to_string(), + env_var: "CLIENT_CONFIG_FILE".to_string(), + }], + ..Default::default() + }), + }) + .await + .unwrap(); + + let mut file_provider = test_provider("work-config", "config-only"); + file_provider.credentials.clear(); + file_provider + .config + .insert("endpoint".to_string(), "https://config.example".to_string()); + state.store.put_message(&file_provider).await.unwrap(); + state + .store + .put_message(&test_provider("work-github", "github")) + .await + .unwrap(); + state + .store + .put_message(&test_sandbox( + "sb-files-and-credentials", + "files-and-credentials", + test_policy_with_rule("sandbox_only", "sandbox.example.com"), + vec!["work-config".to_string(), "work-github".to_string()], + )) + .await + .unwrap(); + + // An older supervisor sends this request without a provider-file + // capability field and ignores the additive files response field. + let response = handle_get_sandbox_provider_environment( + &state, + with_user(Request::new(GetSandboxProviderEnvironmentRequest { + sandbox_id: "sb-files-and-credentials".to_string(), + supports_static_credential_bindings: true, + })), + ) + .await + .unwrap() + .into_inner(); + + assert_eq!( + response.environment.get("GITHUB_TOKEN"), + Some(&"ghp-test".to_string()) + ); + assert_eq!( + response.environment.get("CLIENT_CONFIG_FILE"), + Some(&"/run/openshell/providers/work-config/client.toml".to_string()) + ); + assert_eq!( + response + .files + .get("/run/openshell/providers/work-config/client.toml"), + Some(&"endpoint = 'https://config.example'".to_string()) + ); + } + #[tokio::test] async fn provider_readiness_snapshot_uses_baseline_static_binding_contract() { let state = test_server_state().await; @@ -13733,6 +13800,7 @@ mod tests { deletion_time: None, }), profile: Some(ProviderProfile { + files: Vec::new(), id: "custom-token".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -14063,6 +14131,7 @@ mod tests { profiles: vec![ProviderProfileImportItem { source: "custom-api.yaml".to_string(), profile: Some(ProviderProfile { + files: Vec::new(), id: "custom-api".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -17736,6 +17805,7 @@ mod tests { deletion_time: None, }), profile: Some(ProviderProfile { + files: Vec::new(), id: "custom-api".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -18980,8 +19050,8 @@ mod tests { } #[test] - fn build_gateway_policy_audit_message_formats_ocsf_config_line() { - let message = build_gateway_policy_audit_message( + fn build_gateway_policy_audit_event_formats_ocsf_config_line() { + let message = build_gateway_policy_audit_event( "sb-123", "demo-sandbox", "merged", @@ -18989,7 +19059,8 @@ mod tests { 7, "sha256:testhash", &[], - ); + ) + .format_shorthand(); assert_eq!( message, @@ -19004,13 +19075,13 @@ mod tests { /// findings" — never "safe" — because the claim is about the prover's /// reasoning, not the world. #[test] - fn build_gateway_policy_audit_message_carries_auto_approve_provenance() { + fn build_gateway_policy_audit_event_carries_auto_approve_provenance() { let extra = [ ("auto", "true".to_string()), ("source", "agent_authored".to_string()), ("prover_delta", "empty".to_string()), ]; - let message = build_gateway_policy_audit_message( + let message = build_gateway_policy_audit_event( "sb-123", "demo-sandbox", "approved", @@ -19018,7 +19089,8 @@ mod tests { 12, "sha256:autohash", &extra, - ); + ) + .format_shorthand(); assert!( message.contains("CONFIG:APPROVED"), "auto-approval reuses CONFIG:APPROVED; got: {message}" @@ -19041,6 +19113,58 @@ mod tests { ); } + #[tokio::test] + async fn gateway_policy_audit_events_reach_ocsf_jsonl() { + use tracing_subscriber::prelude::*; + + let directory = tempfile::tempdir().unwrap(); + let path = directory.path().join("events.jsonl"); + let config = toml::from_str(&format!( + "path = {:?}\nrotation = 'never'\n", + path.display().to_string() + )) + .unwrap(); + let log = crate::ocsf_log::OcsfLog::start(config).unwrap(); + let subscriber = tracing_subscriber::registry().with(log.layer()); + tracing::subscriber::with_default(subscriber, || { + emit_gateway_policy_audit_log( + "sb-123", + "demo-sandbox", + "approved", + "approved chunk abc", + 7, + "sha256:manual", + ); + emit_gateway_policy_auto_approve_audit_log( + "sb-123", + "demo-sandbox", + "auto-approved: no new prover findings", + 8, + "sha256:auto", + "mechanistic", + "gateway", + ); + }); + log.shutdown().await; + + let events: Vec = std::fs::read_to_string(path) + .unwrap() + .lines() + .map(|line| serde_json::from_str(line).unwrap()) + .collect(); + assert_eq!(events.len(), 2); + for event in &events { + assert_eq!(event["metadata"]["product"]["name"], "OpenShell Gateway"); + assert_eq!(event["container"]["uid"], "sb-123"); + assert!(event["container"].get("image").is_none()); + } + assert_eq!(events[0]["unmapped"]["policy_hash"], "sha256:manual"); + assert!(events[0]["unmapped"].get("auto").is_none()); + assert_eq!(events[1]["unmapped"]["policy_hash"], "sha256:auto"); + assert_eq!(events[1]["unmapped"]["auto"], "true"); + assert_eq!(events[1]["unmapped"]["resolved_from"], "gateway"); + } + #[test] fn summarize_cli_policy_merge_op_formats_rest_allow_rules() { let operation = PolicyMergeOp::AddAllowRules { diff --git a/crates/openshell-server/src/grpc/provider.rs b/crates/openshell-server/src/grpc/provider.rs index 2141ecb4c2..cc2395fe7d 100644 --- a/crates/openshell-server/src/grpc/provider.rs +++ b/crates/openshell-server/src/grpc/provider.rs @@ -74,6 +74,7 @@ pub(super) struct ProviderEnvironment { pub dynamic_credentials: HashMap, pub static_credential_bindings: HashMap, pub static_credential_keys: HashSet, + pub files: HashMap, } /// Immutable provider records used to build one provider-environment response. @@ -1127,6 +1128,8 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin let mut expires = HashMap::new(); let mut static_credential_bindings = HashMap::new(); let mut static_credential_keys = HashSet::new(); + let mut files = HashMap::new(); + let mut file_env_keys = HashSet::new(); let mut readiness_reason = openshell_core::proto::ProviderReadinessReason::Unspecified; let now_ms = crate::persistence::current_time_ms(); validate_provider_environment_records_unique_at(store, catalog, records, now_ms).await?; @@ -1369,11 +1372,59 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin // or populates its own keys. Cross-provider credential/config // collisions have already been rejected by the validation above. inject_provider_plugin_environment(catalog, provider, ®istry, &mut provider_env); + if let Some(profile) = profile.as_ref() { + if !profile.files.is_empty() + && (name.is_empty() + || !name + .bytes() + .all(|c| c.is_ascii_alphanumeric() || c == b'-' || c == b'_')) + { + return Err(Status::failed_precondition( + "provider name cannot be used in a managed file path", + )); + } + for file in &profile.files { + let path = format!("/run/openshell/providers/{name}/{}", file.path); + let content = file.render(&provider.config).map_err(|error| { + Status::failed_precondition(format!( + "provider '{name}' file '{}': {error}", + file.path + )) + })?; + if files.insert(path.clone(), content).is_some() { + return Err(Status::failed_precondition( + "duplicate provider file destination", + )); + } + if !file.env_var.is_empty() { + if provider_env.insert(file.env_var.clone(), path).is_some() + || env.contains_key(&file.env_var) + { + return Err(Status::failed_precondition(format!( + "provider file environment key '{}' conflicts with another provider output", + file.env_var + ))); + } + file_env_keys.insert(file.env_var.clone()); + } + } + } for (key, value) in provider_env { + if env.contains_key(&key) && file_env_keys.contains(&key) { + return Err(Status::failed_precondition(format!( + "provider file environment key '{key}' conflicts with another provider output" + ))); + } env.entry(key).or_insert(value); } } + if files.len() > 64 || files.values().map(String::len).sum::() > 262_144 { + return Err(Status::failed_precondition( + "provider file set exceeds sandbox limits", + )); + } + Ok(ProviderEnvironment { readiness_reason, environment: env, @@ -1381,6 +1432,7 @@ pub(super) async fn resolve_provider_environment_from_records_with_policy_bindin dynamic_credentials: resolve_dynamic_credentials_from_records(catalog, records), static_credential_bindings, static_credential_keys, + files, }) } @@ -3334,6 +3386,26 @@ fn validate_provider_credentials( provider: &Provider, pending_credentials: &HashMap, ) -> Result<(), Status> { + if !profile.files.is_empty() { + let name = provider.object_name(); + if name.is_empty() + || !name + .bytes() + .all(|c| c.is_ascii_alphanumeric() || c == b'-' || c == b'_') + { + return Err(Status::invalid_argument( + "provider name cannot be used in a managed file path", + )); + } + for file in &profile.files { + file.render(&provider.config).map_err(|error| { + Status::invalid_argument(format!( + "provider file '{}' cannot be rendered: {error}", + file.path + )) + })?; + } + } let declared_keys = profile .credentials .iter() @@ -5342,6 +5414,7 @@ mod tests { }), }; let profile = ProviderProfile { + files: Vec::new(), id: "keycloak-sso".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -6123,6 +6196,7 @@ mod tests { fn custom_profile(id: &str) -> ProviderProfile { ProviderProfile { + files: Vec::new(), id: id.to_string(), resource_version: 0, annotations: HashMap::new(), @@ -6301,6 +6375,7 @@ mod tests { "google-cloud", "google-vertex-ai", "nvidia", + "oci-genai", "openai", "openrouter", "pypi" @@ -6950,6 +7025,7 @@ mod tests { request_id: String::new(), profiles: vec![ProviderProfileImportItem { profile: Some(ProviderProfile { + files: Vec::new(), id: "advanced-api".to_string(), resource_version: 0, annotations: HashMap::new(), @@ -10335,6 +10411,7 @@ mod tests { request_id: String::new(), profiles: vec![ProviderProfileImportItem { profile: Some(ProviderProfile { + files: Vec::new(), id: "delegated-refresh-api".to_string(), resource_version: 0, annotations: HashMap::new(), diff --git a/crates/openshell-server/src/grpc/sandbox.rs b/crates/openshell-server/src/grpc/sandbox.rs index 82bc5d1c40..3bd34b449b 100644 --- a/crates/openshell-server/src/grpc/sandbox.rs +++ b/crates/openshell-server/src/grpc/sandbox.rs @@ -18,7 +18,7 @@ use crate::pagination::Pagination; use crate::persistence::{ ObjectLabels, ObjectListQuery, ObjectType, WriteCondition, generate_name, }; -use crate::tracing_bus::CursoredEvent; +use crate::tracing_bus::{CursoredEvent, ResumeSnapshot}; use crate::watch_cursor::WatchCursor; use futures::future; use openshell_core::net::set_tcp_nodelay_best_effort; @@ -40,7 +40,7 @@ use openshell_core::proto::{ }; use openshell_core::proto::{ BeginRootfsTarStagingRequest, BeginRootfsTarStagingResponse, Sandbox, SandboxPhase, - SandboxTemplate, SshSession, + SandboxRestartPolicy, SandboxTemplate, SshSession, }; use openshell_core::telemetry::{ LifecycleOperation, LifecycleResource, SandboxTemplateSource, TelemetryOutcome, @@ -97,19 +97,6 @@ const RESUME_SPACE_GONE: &str = "resume_after_cursor belongs to a cursor space t const RESUME_CURSOR_AHEAD: &str = "resume_after_cursor is ahead of every cursor this sandbox has \ issued. Restart the watch with an empty resume_after_cursor."; -/// Whether `sandbox_id`'s cursor space is still the one that issued `epoch`. -/// -/// A teardown retires the space and the next publish mints a replacement that -/// renumbers from 1, so a surviving epoch is the only proof that a seq validated -/// earlier still addresses the same numbering. Absent counts as changed: there -/// is nothing left for the cursor to point into. -fn cursor_space_is(state: &ServerState, sandbox_id: &str, epoch: uuid::Uuid) -> bool { - state - .tracing_log_bus - .cursor_space(sandbox_id) - .is_some_and(|space| space.epoch == epoch) -} - #[derive(Debug)] pub struct WatchSandboxStream { receiver: ReceiverStream>, @@ -238,6 +225,13 @@ pub(super) async fn resolve_and_authorize_sandbox_name( Ok(sandbox) } +fn require_ready_sandbox(sandbox: &Sandbox) -> Result<(), Status> { + match SandboxPhase::try_from(sandbox.phase()).ok() { + Some(SandboxPhase::Ready) => Ok(()), + _ => Err(Status::failed_precondition("sandbox is not ready")), + } +} + fn generate_routable_name() -> String { let name = petname::petname(2, "-").unwrap_or_else(generate_name); let mut truncated = &name[..name.len().min(MAX_ROUTABLE_NAME_LEN)]; @@ -420,6 +414,13 @@ async fn handle_create_sandbox_inner( validate_create_sandbox_request_pre_io(&request, &workload_template_name)?; + // Validate labels (keys and values must meet Kubernetes requirements). + for (key, value) in &request.labels { + crate::grpc::validation::validate_label_key(key)?; + crate::grpc::validation::validate_label_value(value)?; + } + crate::grpc::validation::validate_annotations(&request.annotations, "annotations")?; + let authz = authorize_workspace( &state.store, &state.admin_role, @@ -454,12 +455,16 @@ async fn handle_create_sandbox_inner( resolved.providers = governance_spec.providers; resolved.command = governance_spec.command; resolved.tty = governance_spec.tty; + resolved.restart_policy = governance_spec.restart_policy; (resolved, Some(provenance)) }; // Attachment identity belongs to the gateway. Accepting an epoch from a // create request or workload template could revive stale installation proof. spec.provider_attachment_epoch = uuid::Uuid::new_v4().to_string(); + if spec.restart_policy == SandboxRestartPolicy::Unspecified as i32 { + spec.restart_policy = SandboxRestartPolicy::Never as i32; + } // Leave an omitted command empty rather than persisting a concrete shell: // the sandbox boundary resolves the default login shell against the agent image @@ -662,6 +667,11 @@ async fn handle_create_sandbox_inner( &sandbox, &exposure.service, exposure.target_port, + super::service::validate_service_exposure_request( + &exposure.service, + exposure.target_port, + exposure.authorization_mode, + )?, ) .await { @@ -723,7 +733,11 @@ fn validate_create_sandbox_request_pre_io( } let mut service_names = HashSet::with_capacity(request.service_exposures.len()); for exposure in &request.service_exposures { - super::service::validate_service_exposure_request(&exposure.service, exposure.target_port)?; + super::service::validate_service_exposure_request( + &exposure.service, + exposure.target_port, + exposure.authorization_mode, + )?; if !service_names.insert(exposure.service.as_str()) { return Err(Status::invalid_argument(format!( "duplicate service exposure name: '{}'", @@ -1939,24 +1953,22 @@ pub(super) async fn handle_watch_sandbox( // replay buffer and a live receiver; the live loop suppresses events // at or below its source's mark so each is delivered exactly once. // - // The two marks must stay separate. Both buses number from one - // shared cursor space, but they are read at different instants and - // bounded independently (`log_tail_lines` vs `event_tail`, which has - // no default and so replays nothing unless the client asks). A - // single shared mark therefore lets the deeper source censor the - // shallower one: with the default `event_tail` of 0 the mark rises - // to the newest buffered log while no platform event was replayed at - // all, and every platform event published in the initialization - // window is dropped as a duplicate of something never sent. Keyed by - // source, an event is suppressed only if its own source's replay - // actually covered it. + // Both reads happen under one cursor-space lock hold, so every + // event at or below the snapshot's high-water mark was buffered + // when they ran. A resume replays everything after the cursor from + // each followed source, so each source's mark is the last seq its + // own replay returned. The initial tail is depth-bounded, so both + // marks move to the snapshot's high-water mark instead: anything at + // or below it was either in the batch or deliberately left out as + // pre-snapshot history, and its live copy must not arrive after a + // higher cursor from the batch. // - // No unit test pins this. The only reachable window is between the - // subscribe above and the log tail read below -- an event published - // earlier is replayed rather than live, and one published later - // outranks the mark -- and the producer crosses that window with no - // await a test can wedge open. Reproducing it needs a seam in the - // producer, which is not worth adding to production code. + // No gateway-level test reaches the window between the subscribe + // above and the snapshot below: the producer crosses it with no + // await a test can wedge open. The bus-level + // `snapshot_tail_blocks_while_cursor_space_is_locked` and + // `snapshot_after_blocks_while_cursor_space_is_locked` pin the + // atomicity this relies on. let resume_seq = resume_after.map_or(0, |resume| resume.seq); let mut log_cutoff: u64 = resume_seq; let mut platform_cutoff: u64 = resume_seq; @@ -1976,62 +1988,39 @@ pub(super) async fn handle_watch_sandbox( // and silently swallow every live event beneath it. Comparing // epochs answers "did this cursor come from *this* space?", // which no numeric bound can. - match state.tracing_log_bus.cursor_space(&sandbox_id) { - None => { - let _ = tx.send(Err(Status::out_of_range(RESUME_SPACE_GONE))).await; - return; - } - Some(space) if space.epoch != resume.epoch => { + // + // Validation and both bus reads happen inside one call, under + // one lock hold (see `TracingLogBus::snapshot_after`). A + // separate validate-then-read pair -- even with the two reads + // themselves atomic against each other -- still leaves a + // window between the validation's lock and the reads' lock + // for a teardown plus a republish to retire the validated + // space and install a replacement in between; the reads would + // then apply the old space's seq to the replacement's + // buffers and find no gap, since `tail_after` only compares + // numbers. Folding validation into the same hold as the reads + // closes that window entirely. + let (log_replay, platform_replay) = match state.tracing_log_bus.snapshot_after( + &sandbox_id, + resume.epoch, + resume.seq, + resume.seq, + follow_logs, + follow_events, + ) { + ResumeSnapshot::SpaceGone => { let _ = tx.send(Err(Status::out_of_range(RESUME_SPACE_GONE))).await; return; } - // Right space, but ahead of anything it issued: only a - // fabricated token gets here. Reject rather than accept a - // cutoff no event can ever exceed. - Some(space) if resume.seq > space.highest_seq => { + ResumeSnapshot::CursorAhead => { let _ = tx .send(Err(Status::out_of_range(RESUME_CURSOR_AHEAD))) .await; return; } - Some(_) => {} - } - - let log_replay = if follow_logs { - Some(state.tracing_log_bus.tail_after(&sandbox_id, resume.seq)) - } else { - None + ResumeSnapshot::Read(log, platform) => (log, platform), }; - let platform_replay = if follow_events { - Some( - state - .tracing_log_bus - .platform_event_bus - .tail_after(&sandbox_id, resume.seq), - ) - } else { - None - }; - - // Re-check the epoch now that both tails are in hand. The check - // above and each `tail_after` take their locks independently, so - // a teardown plus a republish can retire the validated space and - // install a replacement in between. The reads would then have - // applied the old space's seq to the new space's buffers, and - // `tail_after` -- which only knows numbers -- would report no gap - // while skipping every replacement event at or below it. The - // second look is cheap and runs before anything is emitted, so a - // space that moved under us ends the stream instead of serving a - // truncated replay. - // - // Ordering is unchanged: this takes only the allocator lock and - // releases it, never held across a bus lock. - if !cursor_space_is(&state, &sandbox_id, resume.epoch) { - let _ = tx.send(Err(Status::out_of_range(RESUME_SPACE_GONE))).await; - return; - } - // Gap check FIRST (borrows), before the merge moves the vecs. for replay in [&log_replay, &platform_replay] { if let Some(Err(gap)) = replay { @@ -2104,29 +2093,97 @@ pub(super) async fn handle_watch_sandbox( // cursors expects. // // The two windows are truncated independently (log_tail vs - // event_tail), so this merges whatever each bus retained; it - // does not align their depths. + // event_tail). A client tracks one scalar cursor across both + // sources -- whatever the highest one it saw was -- so a + // cursor this batch hands out is only safe to resume from if + // *every* followed source can vouch it covered everything up + // to that point. + // + // `tail_with_floor` reports each source's own coverage floor: + // the *newest* event its window excluded. The source's + // excluded set is a prefix of its own tail, so it delivered + // everything it retains above the floor. The batch as a whole + // is therefore safe above the highest floor across followed + // sources: every event from either source past that point is + // in the batch, and everything at or below it is outside the + // window, the same as ordinary tail truncation. A source never + // withholds its own window this way; only a sibling's floor + // can cut a source's window back. + // + // That can mean fewer events than a source's depth asked for + // whenever a followed sibling left out an event newer than + // part of its window. This depends on where the histories + // fall in the cursor order, not on the depths being unequal: + // logs at 1-2 and platform at 3-4 with both depths at 1 + // leaves out platform 3, so log 2 is withheld. + // That's the trade-off of a single shared scalar cursor: an + // event handed out below the sibling's floor would let a + // resume skip the sibling's withheld events past it. + // + // Both windows and the space's high-water mark come from one + // lock hold (see `TracingLogBus::snapshot_tail`), so no + // publish can land between the two reads. Read separately, a + // log published after the log read and a platform event + // published after it could let the platform window hand out + // the higher cursor while the log arrives live afterwards. + let crate::tracing_bus::TailSnapshot { + mut logs, + log_floor, + mut events, + platform_floor, + highest_seq, + } = state.tracing_log_bus.snapshot_tail( + &sandbox_id, + follow_logs.then_some(log_tail as usize), + follow_events.then_some(event_tail as usize), + ); + + // 0 means "nothing excluded" for a source, so a fully-covered + // source never raises the maximum. + let critical_floor = log_floor.max(platform_floor); + + let before = logs.len(); + logs.retain(|cursored| cursored.seq > critical_floor); + let log_withheld = before - logs.len(); + let before = events.len(); + events.retain(|cursored| cursored.seq > critical_floor); + let platform_withheld = before - events.len(); + + // Both cutoffs move to the snapshot's high-water mark. The + // receivers were subscribed before the snapshot, so an event + // published in between is both buffered and queued live. At + // or below the mark, it was either delivered above or left + // out of this batch as pre-snapshot history; delivering its + // live copy would emit a lower cursor after a higher one + // from the batch. Above the mark, it can only arrive live. + log_cutoff = log_cutoff.max(highest_seq); + platform_cutoff = platform_cutoff.max(highest_seq); + let mut tail: Vec = Vec::new(); - if follow_logs { - let logs = state.tracing_log_bus.tail(&sandbox_id, log_tail as usize); - if let Some(last) = logs.last() { - log_cutoff = log_cutoff.max(last.seq); - } - tail.extend(logs); - } - if follow_events { - let events = state - .tracing_log_bus - .platform_event_bus - .tail(&sandbox_id, event_tail as usize); - if let Some(last) = events.last() { - platform_cutoff = platform_cutoff.max(last.seq); - } - tail.extend(events); - } + tail.extend(logs); + tail.extend(events); tail.sort_by_key(|cursored| cursored.seq); + // Disclose the gap before anything else in this batch. The + // withheld events sit at or below the snapshot's mark, so the + // cutoffs above keep their live copies (if any) from being + // delivered out of order, while later live events past the + // mark still reach the client. The client has to learn the + // gap exists from this warning; a resume from any cursor this + // batch hands out cannot recover it, and `tail_after` reports + // no gap when asked, since nothing was evicted, only never + // sent. + if log_withheld > 0 || platform_withheld > 0 { + let warning = crate::sandbox_watch::coverage_gap_warning_event( + log_withheld, + platform_withheld, + ); + if tx.send(Ok(warning)).await.is_err() { + return; + } + } + for cursored in tail { // Log filters; platform events carry no log fields and pass. if let Some(openshell_core::proto::sandbox_stream_event::Payload::Log( @@ -2416,9 +2473,7 @@ pub(super) async fn handle_exec_sandbox( completion.ensure_target(sandbox.object_id())?; } - if SandboxPhase::try_from(sandbox.phase()).ok() != Some(SandboxPhase::Ready) { - return Err(Status::failed_precondition("sandbox is not ready")); - } + require_ready_sandbox(&sandbox)?; // Open a relay channel through the supervisor session. Use a 15s // session-wait timeout, enough to cover a transient supervisor reconnect @@ -2931,9 +2986,7 @@ pub(super) async fn handle_exec_sandbox_interactive_start( completion.ensure_target(sandbox.object_id())?; } - if SandboxPhase::try_from(sandbox.phase()).ok() != Some(SandboxPhase::Ready) { - return Err(Status::failed_precondition("sandbox is not ready")); - } + require_ready_sandbox(&sandbox)?; let (channel_id, relay_rx) = crate::supervisor_session::open_routed_relay_with_target( state, @@ -3910,7 +3963,9 @@ mod tests { test_server_state_with_driver, }; use openshell_core::proto::datamodel::v1::ObjectMeta; - use openshell_core::proto::{GpuResourceRequirements, SandboxServiceExposure, ServiceEndpoint}; + use openshell_core::proto::{ + GpuResourceRequirements, SandboxServiceExposure, ServiceAuthorizationMode, ServiceEndpoint, + }; // ---- shell_escape ---- @@ -4658,6 +4713,420 @@ mod tests { assert_eq!(got, vec![1, 2, 3, 4]); } + /// Regression for the inverted coverage-floor clamp: a single followed + /// source with more buffered events than `log_tail_lines` must still + /// deliver its newest window. A source never withholds its own window, + /// and the older lines it excluded are ordinary tail truncation, so no + /// coverage-gap warning is due either. + #[tokio::test] + async fn initial_tail_delivers_a_truncated_single_source_window() { + use tokio_stream::StreamExt as _; + + let state = test_server_state().await; + let sandbox = test_sandbox("truncated", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_log_lines(&state, &id, 5); // cursors 1..=5 + + let response = handle_watch_sandbox( + &state, + authed_request(WatchSandboxRequest { + sandbox: sandbox.object_name().to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + follow_logs: true, + log_tail_lines: 2, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stream = response.into_inner(); + let snap = stream.next().await.unwrap().unwrap(); + assert!(snap.cursor.is_empty()); + + let mut got = Vec::new(); + for _ in 0..2 { + got.push(seq_of(&stream.next().await.unwrap().unwrap())); + } + assert_eq!(got, vec![4, 5]); + + let next = tokio::time::timeout(std::time::Duration::from_millis(200), stream.next()).await; + assert!(next.is_err(), "expected nothing further, got {next:?}"); + } + + /// A sibling's excluded backlog that is older than everything the deeper + /// source delivers imposes nothing: `event_tail: 0` leaves out platform + /// seq 1, but logs 2 and 3 both sit above it, so a resume from 3 skips + /// only an event outside the requested window. + #[tokio::test] + async fn initial_tail_delivers_logs_newer_than_a_sibling_sources_unreplayed_backlog() { + use tokio_stream::StreamExt as _; + + let state = test_server_state().await; + let sandbox = test_sandbox("floor", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_platform_event(&state, &id, "e1"); // cursor 1, never replayed (event_tail: 0) + seed_log_lines(&state, &id, 2); // cursors 2, 3 + + let response = handle_watch_sandbox( + &state, + authed_request(WatchSandboxRequest { + sandbox: sandbox.object_name().to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + follow_logs: true, + follow_events: true, + // Default: no backlog requested from the platform bus. + event_tail: 0, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stream = response.into_inner(); + let snap = stream.next().await.unwrap().unwrap(); + assert!(snap.cursor.is_empty()); + + let mut got = Vec::new(); + for _ in 0..2 { + got.push(seq_of(&stream.next().await.unwrap().unwrap())); + } + assert_eq!(got, vec![2, 3]); + + let next = tokio::time::timeout(std::time::Duration::from_millis(200), stream.next()).await; + assert!( + next.is_err(), + "expected no warning or further events, got {next:?}" + ); + } + + /// The A2 repro: a platform event (seq 2) that `event_tail: 0` leaves + /// out sits between two logs. Handing out log 3 would let the client + /// remember cursor 3, and a resume from 3 would exclude the platform + /// event forever -- `tail_after` reports no gap, since nothing was + /// evicted, it was just never sent. Platform's floor (2) instead cuts the + /// log window back to seq 3 and withholds log 1 with a warning. + #[tokio::test] + async fn initial_tail_withholds_logs_behind_a_sibling_sources_unreplayed_backlog() { + use openshell_core::proto::sandbox_stream_event::Payload; + use tokio_stream::StreamExt as _; + + let state = test_server_state().await; + let sandbox = test_sandbox("floorcut", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_log_lines(&state, &id, 1); // cursor 1, withheld (at or below floor) + seed_platform_event(&state, &id, "e2"); // cursor 2, never replayed (event_tail: 0) + seed_log_lines(&state, &id, 1); // cursor 3, delivered + + let response = handle_watch_sandbox( + &state, + authed_request(WatchSandboxRequest { + sandbox: sandbox.object_name().to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + follow_logs: true, + follow_events: true, + event_tail: 0, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stream = response.into_inner(); + let snap = stream.next().await.unwrap().unwrap(); + assert!(snap.cursor.is_empty()); + + // The client learns about the withheld log from an explicit warning + // rather than discovering it silently on a later reconnect. + let warning = stream.next().await.unwrap().unwrap(); + match warning.payload { + Some(Payload::Warning(w)) => { + assert!( + w.message + .contains("withheld 1 log line(s) and 0 platform event(s)"), + "message: {}", + w.message + ); + } + other => panic!("expected a coverage-gap warning, got {other:?}"), + } + assert!(warning.cursor.is_empty()); + + let delivered = stream.next().await.unwrap().unwrap(); + assert_eq!(seq_of(&delivered), 3); + + let next = tokio::time::timeout(std::time::Duration::from_millis(200), stream.next()).await; + assert!(next.is_err(), "expected nothing further, got {next:?}"); + } + + /// The coverage-gap warning discloses the withheld backlog once, up + /// front, rather than gating every later event: a live event published + /// after connect still reaches the client, even though its cursor sits + /// past the withheld platform backlog. The withheld events were + /// published before subscribe and can never arrive live, so refusing to + /// deliver anything past them would starve the stream indefinitely + /// instead of just disclosing the one, bounded gap. + #[tokio::test] + async fn live_delivery_proceeds_normally_after_the_coverage_gap_warning() { + use openshell_core::proto::sandbox_stream_event::Payload; + use tokio_stream::StreamExt as _; + + let state = test_server_state().await; + let sandbox = test_sandbox("floorlive", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_log_lines(&state, &id, 1); // cursor 1, withheld (below critical_floor) + seed_platform_event(&state, &id, "e2"); // cursor 2, withheld (event_tail: 0) + + let response = handle_watch_sandbox( + &state, + authed_request(WatchSandboxRequest { + sandbox: sandbox.object_name().to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + follow_logs: true, + follow_events: true, + event_tail: 0, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stream = response.into_inner(); + let snap = stream.next().await.unwrap().unwrap(); + assert!(snap.cursor.is_empty()); + + let warning = stream.next().await.unwrap().unwrap(); + assert!(matches!(warning.payload, Some(Payload::Warning(_)))); + + // Published after connect: a genuinely new live event, not part of + // the withheld backlog. It must still be delivered. + seed_log_lines(&state, &id, 1); // cursor 3, live + + let live = tokio::time::timeout(std::time::Duration::from_secs(2), stream.next()) + .await + .expect("live event should arrive") + .unwrap() + .unwrap(); + assert_eq!(seq_of(&live), 3); + } + + /// `event_tail: 0` excludes *both* of the platform bus's own events + /// (seqs 1 and 3); a log event (seq 2) sits interleaved between them and + /// is otherwise fully covered by `log_tail_lines`. Delivering it would let + /// the client remember cursor 2 while platform's seq 3 was never sent, so + /// the floor (3, the newest excluded platform seq) withholds it too. + #[tokio::test] + async fn initial_tail_withholds_an_interleaved_sibling_event_between_two_excluded_seqs() { + use tokio_stream::StreamExt as _; + + let state = test_server_state().await; + let sandbox = test_sandbox("interleavedfloor", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_platform_event(&state, &id, "e1"); // cursor 1, withheld + seed_log_lines(&state, &id, 1); // cursor 2, must also be withheld + seed_platform_event(&state, &id, "e3"); // cursor 3, withheld + + let response = handle_watch_sandbox( + &state, + authed_request(WatchSandboxRequest { + sandbox: sandbox.object_name().to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + follow_logs: true, + follow_events: true, + event_tail: 0, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stream = response.into_inner(); + let snap = stream.next().await.unwrap().unwrap(); + assert!(snap.cursor.is_empty()); + + let warning = stream.next().await.unwrap().unwrap(); + match warning.payload { + Some(openshell_core::proto::sandbox_stream_event::Payload::Warning(w)) => { + assert!(w.message.contains('1'), "message: {}", w.message); + } + other => panic!("expected a coverage-gap warning, got {other:?}"), + } + + // Nothing else: the log event must not slip through between + // platform's two excluded seqs. + let next = tokio::time::timeout(std::time::Duration::from_millis(200), stream.next()).await; + assert!( + next.is_err(), + "expected the interleaved log event to be withheld too, got {next:?}" + ); + } + + /// Windows that leave nothing out impose no floor: every buffered event + /// is delivered, as before A2. + #[tokio::test] + async fn initial_tail_unclamped_when_no_source_has_unreplayed_backlog() { + use tokio_stream::StreamExt as _; + + let state = test_server_state().await; + let sandbox = test_sandbox("nofloor", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_platform_event(&state, &id, "e1"); // cursor 1 + seed_log_lines(&state, &id, 2); // cursors 2, 3 + + let response = handle_watch_sandbox( + &state, + authed_request(WatchSandboxRequest { + sandbox: sandbox.object_name().to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + follow_logs: true, + follow_events: true, + // Deep enough to cover the one platform event: no floor. + event_tail: 10, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stream = response.into_inner(); + let snap = stream.next().await.unwrap().unwrap(); + assert!(snap.cursor.is_empty()); + + let mut got = Vec::new(); + for _ in 0..3 { + got.push(seq_of(&stream.next().await.unwrap().unwrap())); + } + assert_eq!(got, vec![1, 2, 3]); + } + + /// Open an initial watch following both sources with the given depths and + /// collect what the initial batch delivers: the coverage-gap warning + /// message, if any, and the delivered cursors in order. + async fn initial_tail_following_both( + state: &Arc, + sandbox: &Sandbox, + log_tail_lines: u32, + event_tail: u32, + ) -> (Option, Vec) { + use openshell_core::proto::sandbox_stream_event::Payload; + use tokio_stream::StreamExt as _; + + let response = handle_watch_sandbox( + state, + authed_request(WatchSandboxRequest { + sandbox: sandbox.object_name().to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + follow_logs: true, + follow_events: true, + log_tail_lines, + event_tail, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stream = response.into_inner(); + let snap = stream.next().await.unwrap().unwrap(); + assert!(snap.cursor.is_empty()); + + let mut warning = None; + let mut delivered = Vec::new(); + while let Ok(Some(item)) = + tokio::time::timeout(std::time::Duration::from_millis(200), stream.next()).await + { + let event = item.unwrap(); + match event.payload { + Some(Payload::Warning(w)) => { + assert!(delivered.is_empty(), "warning must precede the batch"); + assert!(warning.replace(w.message).is_none(), "one warning at most"); + } + _ => delivered.push(seq_of(&event)), + } + } + (warning, delivered) + } + + /// Equal depths do not rule out trimming. Logs sit at cursors 1-2 and + /// platform events at 3-4; with both depths at 1 the platform window + /// leaves out 3, which is newer than log 2, so log 2 is withheld even + /// though `log_tail_lines` alone would have delivered it. + #[tokio::test] + async fn initial_tail_equal_depths_withhold_an_older_log_window() { + let state = test_server_state().await; + let sandbox = test_sandbox("equallogsfirst", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_log_lines(&state, &id, 2); // cursors 1, 2 + seed_platform_event(&state, &id, "e3"); // cursor 3, left out by event_tail + seed_platform_event(&state, &id, "e4"); // cursor 4 + + let (warning, delivered) = initial_tail_following_both(&state, &sandbox, 1, 1).await; + let warning = warning.expect("withheld log 2 must be disclosed"); + assert!( + warning.contains("withheld 1 log line(s) and 0 platform event(s)"), + "message: {warning}" + ); + assert_eq!(delivered, vec![4]); + } + + /// The mirror case: platform events at cursors 1-2 and logs at 3-4. The + /// log window leaves out 3, which is newer than platform 2, so this time + /// the platform event is withheld. + #[tokio::test] + async fn initial_tail_equal_depths_withhold_an_older_platform_window() { + let state = test_server_state().await; + let sandbox = test_sandbox("equalplatformfirst", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_platform_event(&state, &id, "e1"); // cursor 1 + seed_platform_event(&state, &id, "e2"); // cursor 2 + seed_log_lines(&state, &id, 2); // cursors 3, 4; 3 left out by log_tail_lines + + let (warning, delivered) = initial_tail_following_both(&state, &sandbox, 1, 1).await; + let warning = warning.expect("withheld platform event 2 must be disclosed"); + assert!( + warning.contains("withheld 0 log line(s) and 1 platform event(s)"), + "message: {warning}" + ); + assert_eq!(delivered, vec![4]); + } + + /// Equal depths over interleaved histories withhold nothing when neither + /// window leaves out an event newer than part of the other: logs at 1 and + /// 3, platform at 2 and 4, both depths 1. Each window's excluded event is + /// older than everything delivered, so both 3 and 4 arrive, no warning. + #[tokio::test] + async fn initial_tail_equal_depths_over_interleaved_histories_withhold_nothing() { + let state = test_server_state().await; + let sandbox = test_sandbox("equalinterleaved", Vec::new()); + state.store.put_message(&sandbox).await.unwrap(); + let id = sandbox.object_id().to_string(); + + seed_log_lines(&state, &id, 1); // cursor 1, left out by log_tail_lines + seed_platform_event(&state, &id, "e2"); // cursor 2, left out by event_tail + seed_log_lines(&state, &id, 1); // cursor 3 + seed_platform_event(&state, &id, "e4"); // cursor 4 + + let (warning, delivered) = initial_tail_following_both(&state, &sandbox, 1, 1).await; + assert_eq!(warning, None); + assert_eq!(delivered, vec![3, 4]); + } + #[tokio::test] async fn live_delivery_orders_events_across_sources_by_cursor() { use tokio_stream::StreamExt as _; @@ -5008,45 +5477,6 @@ mod tests { assert!(stream.next().await.is_none()); } - /// The guard behind the producer's post-replay epoch re-check. - /// - /// Validation and the two `tail_after` reads take their locks separately, so - /// a teardown plus a republish can swap the space in between and leave the - /// reads applying an old seq to a replacement's buffers -- `tail_after` only - /// compares numbers, so it reports no gap while skipping every replacement - /// event at or below that seq. The producer re-checks the epoch once both - /// tails are in hand and before emitting anything; this pins what that check - /// must answer. - /// - /// The interleaving itself is not reachable from a test: the producer runs - /// validation, both reads, and the re-check with no await in between, so - /// there is nothing to suspend it on. - #[tokio::test] - async fn cursor_space_is_rejects_a_replacement_space() { - let state = test_server_state().await; - let sandbox = test_sandbox("respace", Vec::new()); - state.store.put_message(&sandbox).await.unwrap(); - let id = sandbox.object_id().to_string(); - - seed_log_lines(&state, &id, 1); - let original = state.tracing_log_bus.cursor_space(&id).unwrap().epoch; - assert!(cursor_space_is(&state, &id, original)); - - // Teardown alone leaves no space to point into. - state.tracing_log_bus.remove(&id); - assert!(state.tracing_log_bus.cursor_space(&id).is_none()); - assert!(!cursor_space_is(&state, &id, original)); - - // The republish installs a replacement renumbered from 1. Its seqs - // overlap the retired space's, so only the epoch separates them. - seed_log_lines(&state, &id, 1); - let replacement = state.tracing_log_bus.cursor_space(&id).unwrap(); - assert_ne!(replacement.epoch, original); - assert_eq!(replacement.highest_seq, 1); - assert!(!cursor_space_is(&state, &id, original)); - assert!(cursor_space_is(&state, &id, replacement.epoch)); - } - #[tokio::test] async fn resume_from_reset_cursor_space_terminates_out_of_range() { use tokio_stream::StreamExt as _; @@ -6369,7 +6799,11 @@ mod tests { authed_request(CreateSandboxRequest { name: "mcp-canonical".to_string(), spec: Some(SandboxSpec { - policy: Some(mcp_policy_with_versions(&["2025-11-25", "2025-03-26"])), + policy: Some(mcp_policy_with_versions(&[ + "2026-07-28", + "2025-11-25", + "2025-03-26", + ])), ..Default::default() }), labels: HashMap::new(), @@ -6399,7 +6833,7 @@ mod tests { .as_ref() .expect("MCP options") .versions; - assert_eq!(versions, &["2025-03-26", "2025-11-25"]); + assert_eq!(versions, &["2025-03-26", "2025-11-25", "2026-07-28"]); } #[tokio::test] @@ -6590,7 +7024,7 @@ mod tests { let state = test_server_state().await; let cases: &[(&str, &[&str])] = &[ ("mcp-duplicate-versions", &["2025-11-25", "2025-11-25"]), - ("mcp-unsupported-version", &["2026-07-28"]), + ("mcp-unsupported-version", &["2026-07-29"]), ]; for &(sandbox_name, versions) in cases { @@ -6653,6 +7087,10 @@ mod tests { .into_inner(); let created = response.sandbox.expect("created sandbox"); + assert_eq!( + created.spec.as_ref().unwrap().restart_policy(), + SandboxRestartPolicy::Never + ); assert_eq!( created .metadata @@ -6698,10 +7136,12 @@ mod tests { SandboxServiceExposure { service: String::new(), target_port: 4500, + authorization_mode: ServiceAuthorizationMode::Unspecified as i32, }, SandboxServiceExposure { service: "metrics".to_string(), target_port: 9090, + authorization_mode: ServiceAuthorizationMode::BearerPassthrough as i32, }, ], ..Default::default() @@ -6734,9 +7174,46 @@ mod tests { assert_eq!(endpoint.name, service); assert_eq!(endpoint.target_port, target_port); assert!(endpoint.domain); + let expected_mode = if service.is_empty() { + ServiceAuthorizationMode::Strip + } else { + ServiceAuthorizationMode::BearerPassthrough + }; + assert_eq!(endpoint.authorization_mode(), expected_mode); } } + #[tokio::test] + async fn create_sandbox_rejects_unknown_service_authorization_mode_before_persisting() { + let state = test_server_state().await; + let error = handle_create_sandbox( + &state, + authed_request(CreateSandboxRequest { + name: "invalid-service-authorization".to_string(), + spec: Some(SandboxSpec::default()), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + service_exposures: vec![SandboxServiceExposure { + service: String::new(), + target_port: 4500, + authorization_mode: 99, + }], + ..Default::default() + }), + ) + .await + .expect_err("unknown service authorization mode should be rejected"); + + assert_eq!(error.code(), tonic::Code::InvalidArgument); + assert!( + state + .store + .get_message_by_name::("default", "invalid-service-authorization") + .await + .expect("sandbox lookup should succeed") + .is_none() + ); + } + #[tokio::test] async fn create_sandbox_begins_rollback_when_service_exposure_fails() { let state = test_server_state().await; @@ -6766,10 +7243,12 @@ mod tests { SandboxServiceExposure { service: "web".to_string(), target_port: 8080, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }, SandboxServiceExposure { service: "metrics".to_string(), target_port: 9090, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }, ], ..Default::default() @@ -6853,10 +7332,12 @@ mod tests { SandboxServiceExposure { service: "web".to_string(), target_port: 8080, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }, SandboxServiceExposure { service: "web".to_string(), target_port: 8081, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }, ], ..Default::default() @@ -7545,7 +8026,7 @@ mod tests { fn template_create_sandbox_spec_field_policy_is_exhaustive() { assert_proto_fields_classified( "openshell.v1.SandboxSpec", - &["policy", "providers", "command", "tty"], + &["policy", "providers", "command", "tty", "restart_policy"], &[ "log_level", "environment", diff --git a/crates/openshell-server/src/grpc/service.rs b/crates/openshell-server/src/grpc/service.rs index 675b695656..bb44f395b8 100644 --- a/crates/openshell-server/src/grpc/service.rs +++ b/crates/openshell-server/src/grpc/service.rs @@ -7,7 +7,8 @@ use std::sync::Arc; use openshell_core::proto::datamodel::v1::ObjectMeta; use openshell_core::proto::{ DeleteServiceRequest, DeleteServiceResponse, ExposeServiceRequest, GetServiceRequest, - ListServicesRequest, ListServicesResponse, Sandbox, ServiceEndpoint, ServiceEndpointResponse, + ListServicesRequest, ListServicesResponse, Sandbox, ServiceAuthorizationMode, ServiceEndpoint, + ServiceEndpointResponse, }; use openshell_core::{GetResourceVersion, ObjectId, ObjectName, ObjectWorkspace}; use prost::Message as _; @@ -49,19 +50,37 @@ pub(super) async fn handle_expose_service( super::workspace::resolve_workspace(state.store.as_ref(), sandbox.object_workspace()) .await? .ensure_active()?; - validate_service_exposure_request(&req.name, req.target_port)?; - expose_service_endpoint(state, &workspace, &sandbox, &req.name, req.target_port).await + let authorization_mode = + validate_service_exposure_request(&req.name, req.target_port, req.authorization_mode)?; + expose_service_endpoint( + state, + &workspace, + &sandbox, + &req.name, + req.target_port, + authorization_mode, + ) + .await } pub(super) fn validate_service_exposure_request( service: &str, target_port: u32, -) -> Result<(), Status> { + authorization_mode: i32, +) -> Result { validate_optional_endpoint_name("service", service, MAX_SERVICE_NAME_LEN)?; if target_port == 0 || target_port > u32::from(u16::MAX) { return Err(Status::invalid_argument("target_port must be in 1..=65535")); } - Ok(()) + match ServiceAuthorizationMode::try_from(authorization_mode) { + Ok(ServiceAuthorizationMode::Unspecified | ServiceAuthorizationMode::Strip) => { + Ok(ServiceAuthorizationMode::Strip) + } + Ok(ServiceAuthorizationMode::BearerPassthrough) => { + Ok(ServiceAuthorizationMode::BearerPassthrough) + } + Err(_) => Err(Status::invalid_argument("authorization_mode is invalid")), + } } pub(super) async fn expose_service_endpoint( @@ -70,6 +89,7 @@ pub(super) async fn expose_service_endpoint( sandbox: &Sandbox, service: &str, target_port: u32, + authorization_mode: ServiceAuthorizationMode, ) -> Result, Status> { let sandbox_name = sandbox.object_name(); @@ -132,6 +152,7 @@ pub(super) async fn expose_service_endpoint( name: service.to_string(), target_port, domain: true, + authorization_mode: authorization_mode as i32, }; // Single-attempt CAS write: fails with ABORTED on concurrent modification @@ -344,8 +365,16 @@ async fn get_service_endpoint( fn service_endpoint_response( state: &Arc, - endpoint: ServiceEndpoint, + mut endpoint: ServiceEndpoint, ) -> ServiceEndpointResponse { + endpoint.authorization_mode = + match ServiceAuthorizationMode::try_from(endpoint.authorization_mode) { + Ok(ServiceAuthorizationMode::BearerPassthrough) => { + ServiceAuthorizationMode::BearerPassthrough as i32 + } + Ok(ServiceAuthorizationMode::Unspecified | ServiceAuthorizationMode::Strip) + | Err(_) => ServiceAuthorizationMode::Strip as i32, + }; let workspace = endpoint.object_workspace(); let url = service_routing::endpoint_url(&state.config, workspace, &endpoint.sandbox, &endpoint.name) @@ -456,6 +485,100 @@ mod tests { assert!(validate_endpoint_name("service", "Web", 28).is_err()); } + #[test] + fn authorization_mode_defaults_to_strip_and_rejects_unknown_values() { + assert_eq!( + validate_service_exposure_request("web", 8080, 0).unwrap(), + ServiceAuthorizationMode::Strip + ); + assert_eq!( + validate_service_exposure_request( + "web", + 8080, + ServiceAuthorizationMode::BearerPassthrough as i32, + ) + .unwrap(), + ServiceAuthorizationMode::BearerPassthrough + ); + assert_eq!( + validate_service_exposure_request("web", 8080, 99) + .unwrap_err() + .code(), + tonic::Code::InvalidArgument + ); + } + + #[tokio::test] + async fn unknown_authorization_mode_is_rejected_before_persistence() { + let state = test_server_state().await; + seed_sandbox(&state, "my-sandbox").await; + + let error = handle_expose_service( + &state, + authed_request(ExposeServiceRequest { + sandbox: "my-sandbox".to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + name: "web".to_string(), + target_port: 8080, + authorization_mode: 99, + ..Default::default() + }), + ) + .await + .unwrap_err(); + + assert_eq!(error.code(), tonic::Code::InvalidArgument); + assert!( + get_service_endpoint(&state, "default", "my-sandbox", "web") + .await + .unwrap() + .is_none() + ); + } + + #[tokio::test] + async fn legacy_unspecified_authorization_mode_is_reported_as_strip() { + let state = test_server_state().await; + seed_sandbox(&state, "my-sandbox").await; + + handle_expose_service( + &state, + authed_request(ExposeServiceRequest { + sandbox: "my-sandbox".to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + name: "web".to_string(), + target_port: 8080, + ..Default::default() + }), + ) + .await + .unwrap(); + + let mut stored = get_service_endpoint(&state, "default", "my-sandbox", "web") + .await + .unwrap() + .unwrap(); + stored.authorization_mode = ServiceAuthorizationMode::Unspecified as i32; + state.store.put_message(&stored).await.unwrap(); + + let response = handle_get_service( + &state, + authed_request(GetServiceRequest { + sandbox: "my-sandbox".to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + name: "web".to_string(), + }), + ) + .await + .unwrap() + .into_inner(); + + assert_eq!( + response.endpoint.unwrap().authorization_mode(), + ServiceAuthorizationMode::Strip + ); + } + #[tokio::test] async fn endpoint_lifecycle_round_trip() { let state = test_server_state().await; @@ -472,12 +595,17 @@ mod tests { name: "web".to_string(), target_port: 8080, domain: true, + authorization_mode: ServiceAuthorizationMode::Unspecified as i32, }), ) .await .unwrap() .into_inner(); assert_eq!(exposed.endpoint.as_ref().unwrap().target_port, 8080); + assert_eq!( + exposed.endpoint.as_ref().unwrap().authorization_mode(), + ServiceAuthorizationMode::Strip + ); let listed = handle_list_services( &state, @@ -511,6 +639,26 @@ mod tests { .into_inner(); assert_eq!(fetched.endpoint.as_ref().unwrap().target_port, 8080); + let updated = handle_expose_service( + &state, + authed_request(ExposeServiceRequest { + sandbox: "my-sandbox".to_string(), + workspace_scope: Some(openshell_core::proto::workspace_selector("default")), + name: "web".to_string(), + target_port: 9090, + authorization_mode: ServiceAuthorizationMode::BearerPassthrough as i32, + ..Default::default() + }), + ) + .await + .unwrap() + .into_inner(); + assert_eq!(updated.endpoint.as_ref().unwrap().target_port, 9090); + assert_eq!( + updated.endpoint.as_ref().unwrap().authorization_mode(), + ServiceAuthorizationMode::BearerPassthrough + ); + let deleted = handle_delete_service( &state, authed_request(DeleteServiceRequest { @@ -643,6 +791,7 @@ mod tests { name: "web".to_string(), target_port: 8080, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await @@ -661,6 +810,7 @@ mod tests { name: "web".to_string(), target_port: 9090, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await @@ -713,6 +863,7 @@ mod tests { name: "web".to_string(), target_port: 7070, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await @@ -732,6 +883,7 @@ mod tests { name: "web".to_string(), target_port: 8080, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await @@ -750,6 +902,7 @@ mod tests { name: "web".to_string(), target_port: 9090, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await @@ -840,6 +993,7 @@ mod tests { name: "web".to_string(), target_port: 8080, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await @@ -856,6 +1010,7 @@ mod tests { name: "web".to_string(), target_port: 9090, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await @@ -999,6 +1154,7 @@ mod tests { name: "api".to_string(), target_port: 3000, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }), ) .await diff --git a/crates/openshell-server/src/grpc/validation.rs b/crates/openshell-server/src/grpc/validation.rs index 4218f296da..f72c571ba5 100644 --- a/crates/openshell-server/src/grpc/validation.rs +++ b/crates/openshell-server/src/grpc/validation.rs @@ -10,7 +10,7 @@ use openshell_core::proto::{ CredentialHandle, ExecSandboxRequest, Provider, SandboxPolicy as ProtoSandboxPolicy, - SandboxSpec, SandboxTemplate, + SandboxRestartPolicy, SandboxSpec, SandboxTemplate, }; use openshell_core::rpc_error::invalid_argument; use prost::Message; @@ -206,6 +206,7 @@ pub(super) fn validate_sandbox_spec(name: &str, spec: &SandboxSpec) -> Result<() if !spec.command.is_empty() { validate_main_process_command(&spec.command)?; } + validate_restart_policy(spec)?; // --- spec.policy serialized size --- validate_sandbox_policy_size(spec)?; @@ -222,10 +223,21 @@ pub(super) fn validate_sandbox_governance_spec( if !spec.command.is_empty() { validate_main_process_command(&spec.command)?; } + validate_restart_policy(spec)?; validate_sandbox_policy_size(spec)?; Ok(()) } +fn validate_restart_policy(spec: &SandboxSpec) -> Result<(), Status> { + SandboxRestartPolicy::try_from(spec.restart_policy).map_err(|_| { + invalid_argument( + "spec.restart_policy", + format!("unknown restart_policy value: {}", spec.restart_policy), + ) + })?; + Ok(()) +} + fn validate_sandbox_name(name: &str) -> Result<(), Status> { if !name.is_empty() && name.len() > MAX_ROUTABLE_NAME_LEN { return Err(invalid_argument( @@ -1227,6 +1239,17 @@ mod tests { assert!(validate_sandbox_spec("", &default_spec()).is_ok()); } + #[test] + fn validate_sandbox_spec_rejects_unknown_restart_policy() { + let spec = SandboxSpec { + restart_policy: 99, + ..Default::default() + }; + let err = validate_sandbox_spec("", &spec).unwrap_err(); + assert_eq!(err.code(), Code::InvalidArgument); + assert!(err.message().contains("restart_policy")); + } + #[test] fn validate_sandbox_spec_accepts_exact_main_process_argv() { let spec = SandboxSpec { diff --git a/crates/openshell-server/src/lib.rs b/crates/openshell-server/src/lib.rs index ea1f74c0f4..d6d35f6ba3 100644 --- a/crates/openshell-server/src/lib.rs +++ b/crates/openshell-server/src/lib.rs @@ -22,10 +22,12 @@ mod config_update_operation; mod credentials; mod defaults; mod gateway_listener; +mod gateway_ocsf; mod grpc; mod http; mod middleware; mod multiplex; +mod ocsf_log; mod otel_tracing; mod pagination; mod persistence; @@ -949,9 +951,11 @@ pub(crate) async fn run_server( // Deadlines must run while restored supervisors wait for policy repair. let (startup_tx, startup_rx) = watch::channel(false); - state - .compute - .spawn_watchers(shutdown_rx.clone(), startup_rx); + state.compute.spawn_watchers( + shutdown_rx.clone(), + startup_rx, + state.sandbox_session_jwt_authority.clone(), + ); // Serve the gateway before reconciling persisted sandboxes so restored // supervisors can fetch policy and register their sessions. diff --git a/crates/openshell-server/src/multiplex.rs b/crates/openshell-server/src/multiplex.rs index 8988542b1a..a909ff2bb7 100644 --- a/crates/openshell-server/src/multiplex.rs +++ b/crates/openshell-server/src/multiplex.rs @@ -63,6 +63,16 @@ impl MakeRequestId for UuidRequestId { } } +/// Paths called on a timer rather than by someone waiting on the result. +const POLLED_PATHS: &[&str] = &[ + "/health", + "/healthz", + "/readyz", + "/openshell.v1.OpenShell/GetSandboxConfig", + "/openshell.v1.OpenShell/ReportProviderReadiness", + "/openshell.v1.OpenShell/PeerReportProviderReadiness", +]; + /// Build a tracing span for an inbound request, recording the `request_id` /// header (set by [`UuidRequestId`] or supplied by the client). fn make_request_span(req: &Request) -> Span { @@ -79,7 +89,7 @@ fn make_request_span(req: &Request) -> Span { // the callsite name. let otel_name = otel_span_name(req.method(), path); - let span = if matches!(path, "/health" | "/healthz" | "/readyz") { + let span = if POLLED_PATHS.contains(&path) { tracing::debug_span!( "request", method = %req.method(), @@ -2223,6 +2233,34 @@ mod tests { ); } + #[test] + fn polled_paths_get_debug_request_spans() { + let _traced = crate::otel_tracing::test_exporter::install_traced(); + let level = |path: &str| { + let req = Request::builder() + .uri(path) + .body(Empty::::new()) + .unwrap(); + *make_request_span(&req) + .metadata() + .expect("span enabled") + .level() + }; + + for path in [ + "/healthz", + "/openshell.v1.OpenShell/GetSandboxConfig", + "/openshell.v1.OpenShell/ReportProviderReadiness", + "/openshell.v1.OpenShell/PeerReportProviderReadiness", + ] { + assert_eq!(level(path), tracing::Level::DEBUG, "{path}"); + } + assert_eq!( + level("/openshell.v1.OpenShell/CreateSandbox"), + tracing::Level::INFO + ); + } + /// The `TraceLayer` creates the server span, so no gRPC handler needs /// `#[instrument]`. The request ID carries into it so a trace can be /// correlated with the gateway's logs. diff --git a/crates/openshell-server/src/ocsf_log.rs b/crates/openshell-server/src/ocsf_log.rs new file mode 100644 index 0000000000..8858c9c587 --- /dev/null +++ b/crates/openshell-server/src/ocsf_log.rs @@ -0,0 +1,712 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +use std::collections::{BTreeMap, VecDeque}; +use std::fs::{File, OpenOptions}; +use std::io::{self, Read, Seek, SeekFrom, Write}; +use std::path::Path; +use std::sync::{Arc, Condvar, Mutex}; +use std::time::{Duration, Instant}; + +use chrono::{DateTime, NaiveDate, Utc}; +use openshell_ocsf::{OcsfEvent, format::downgrade::downgrade_event}; +use tracing::{Event, Subscriber}; +use tracing_subscriber::Layer; +use tracing_subscriber::layer::Context; + +use crate::config_file::{OcsfLogConfig, OcsfLogRotation, OcsfSchemaVersion}; + +const BATCH_SIZE: usize = 100; +const FLUSH_INTERVAL: Duration = Duration::from_millis(500); +const SHUTDOWN_BUDGET: Duration = Duration::from_secs(5); + +#[derive(Default)] +struct State { + queue: VecDeque>, + queued_bytes: usize, + in_flight: usize, + closed: bool, + abandoned: bool, + losses: BTreeMap<&'static str, u64>, +} + +struct Shared { + state: Mutex, + ready: Condvar, + config: OcsfLogConfig, +} + +#[derive(Clone)] +pub struct Collector(Arc); + +impl Collector { + fn new(config: OcsfLogConfig) -> Self { + Self(Arc::new(Shared { + state: Mutex::new(State::default()), + ready: Condvar::new(), + config, + })) + } + + pub(crate) fn collect(&self, event: &OcsfEvent) { + match serialize_event(event, self.0.config.schema_version) { + Ok(line) => { + self.enqueue(line); + } + Err(_) => self.record_loss("serialization", 1), + } + } + + fn enqueue(&self, line: Vec) -> bool { + let mut state = self.0.state.lock().unwrap(); + let reason = if state.closed { + Some("closed") + } else if line.len() > self.0.config.queue_max_bytes.get() { + Some("oversized") + } else if state.queue.len() >= self.0.config.queue_capacity.get() { + Some("queue_capacity") + } else if line.len() > self.0.config.queue_max_bytes.get() - state.queued_bytes { + Some("queue_bytes") + } else { + None + }; + if let Some(reason) = reason { + Self::loss(&mut state, reason, 1); + return false; + } + state.queued_bytes += line.len(); + state.queue.push_back(line); + Self::queue_metrics(&state); + metrics::counter!("openshell_ocsf_log_queued_total").increment(1); + self.0.ready.notify_one(); + true + } + + #[allow(clippy::cast_precision_loss)] + fn queue_metrics(state: &State) { + metrics::gauge!("openshell_ocsf_log_queue_records").set(state.queue.len() as f64); + metrics::gauge!("openshell_ocsf_log_queue_bytes").set(state.queued_bytes as f64); + } + + fn loss(state: &mut State, reason: &'static str, count: u64) { + if count > 0 { + *state.losses.entry(reason).or_default() += count; + metrics::counter!("openshell_ocsf_log_dropped_total", "reason" => reason) + .increment(count); + } + } + + pub(crate) fn record_loss(&self, reason: &'static str, count: u64) { + Self::loss(&mut self.0.state.lock().unwrap(), reason, count); + } + + fn close(&self) { + self.0.state.lock().unwrap().closed = true; + self.0.ready.notify_all(); + } + + fn abandon(&self) { + let mut state = self.0.state.lock().unwrap(); + state.abandoned = true; + let queued = state.queue.len() as u64; + let uncertain = state.in_flight as u64; + state.queue.clear(); + state.queued_bytes = 0; + state.in_flight = 0; + Self::loss(&mut state, "shutdown", queued); + Self::loss(&mut state, "shutdown_uncertain", uncertain); + Self::queue_metrics(&state); + self.0.ready.notify_all(); + } + + fn batch(&self) -> Option>> { + let mut state = self.0.state.lock().unwrap(); + while state.queue.is_empty() && !state.closed { + state = self.0.ready.wait(state).unwrap(); + } + let deadline = Instant::now() + FLUSH_INTERVAL; + while !state.closed && state.queue.len() < BATCH_SIZE && Instant::now() < deadline { + state = self + .0 + .ready + .wait_timeout(state, deadline.saturating_duration_since(Instant::now())) + .unwrap() + .0; + } + if state.queue.is_empty() || state.abandoned { + return None; + } + let count = BATCH_SIZE.min(state.queue.len()); + let batch: Vec<_> = state.queue.drain(..count).collect(); + state.queued_bytes -= batch.iter().map(Vec::len).sum::(); + state.in_flight = count; + Self::queue_metrics(&state); + Some(batch) + } + + fn complete(&self, failure: Option<&'static str>) { + let mut state = self.0.state.lock().unwrap(); + if state.abandoned { + return; + } + state.in_flight -= 1; + if let Some(reason) = failure { + Self::loss(&mut state, reason, 1); + } else { + metrics::counter!("openshell_ocsf_log_written_total").increment(1); + } + } +} + +fn serialize_event( + event: &OcsfEvent, + schema_version: Option, +) -> Result, serde_json::Error> { + let Some(schema_version) = schema_version else { + return event.to_json_line().map(String::into_bytes); + }; + let mut event = serde_json::to_value(event)?; + downgrade_event(&mut event, schema_version.as_str()); + let mut line = serde_json::to_vec(&event)?; + line.push(b'\n'); + Ok(line) +} + +pub struct OcsfLog { + collector: Collector, + done: tokio::sync::oneshot::Receiver<()>, +} + +impl OcsfLog { + pub(crate) fn start(config: OcsfLogConfig) -> io::Result { + let collector = Collector::new(config); + let worker = collector.clone(); + let (done_tx, done) = tokio::sync::oneshot::channel(); + std::thread::Builder::new() + .name("ocsf-jsonl".into()) + .spawn(move || { + run_writer(&worker); + let _ = done_tx.send(()); + })?; + Ok(Self { collector, done }) + } + + pub(crate) fn collector(&self) -> Collector { + self.collector.clone() + } + + pub(crate) fn layer(&self) -> impl Layer + use { + CaptureLayer(self.collector()) + } + + pub(crate) async fn shutdown(mut self) { + self.shutdown_with_budget(SHUTDOWN_BUDGET).await; + } + + async fn shutdown_with_budget(&mut self, budget: Duration) { + self.collector.close(); + if !matches!( + tokio::time::timeout(budget, &mut self.done).await, + Ok(Ok(())) + ) { + self.collector.abandon(); + tracing::warn!( + "OCSF JSONL shutdown did not finish; queued records lost and outstanding writes uncertain" + ); + } + } +} + +impl Drop for OcsfLog { + fn drop(&mut self) { + self.collector.close(); + } +} + +struct CaptureLayer(Collector); + +impl Layer for CaptureLayer { + fn on_event(&self, event: &Event<'_>, _context: Context<'_, S>) { + if event.metadata().target() == openshell_ocsf::OCSF_TARGET + && let Some(event) = openshell_ocsf::clone_current_event() + { + self.0.collect(&event); + } + } +} + +fn run_writer(collector: &Collector) { + let mut writer = FileWriter::new(collector.0.config.clone()); + let mut retry_at = Instant::now(); + let mut backoff = Duration::from_millis(500); + while let Some(batch) = collector.batch() { + for line in batch { + if collector.0.state.lock().unwrap().abandoned { + return; + } + if Instant::now() < retry_at { + collector.complete(Some("unavailable")); + continue; + } + match writer.append(&line, Utc::now().date_naive()) { + Ok(()) => { + backoff = Duration::from_millis(500); + collector.complete(None); + } + Err(error) => { + writer.file = None; + metrics::counter!("openshell_ocsf_log_writer_errors_total").increment(1); + tracing::warn!(%error, "OCSF JSONL write failed; record may be incomplete, later records discarded until reopen"); + collector.complete(Some("write_uncertain")); + retry_at = Instant::now() + backoff; + backoff = (backoff * 2).min(Duration::from_secs(30)); + } + } + } + } +} + +struct FileWriter { + config: OcsfLogConfig, + file: Option, + day: Option, +} + +impl FileWriter { + fn new(config: OcsfLogConfig) -> Self { + Self { + config, + file: None, + day: None, + } + } + + fn open(&mut self) -> io::Result<()> { + if let Some(parent) = self + .config + .path + .parent() + .filter(|path| !path.as_os_str().is_empty()) + { + std::fs::create_dir_all(parent)?; + } + if let Ok(metadata) = std::fs::metadata(&self.config.path) + && !metadata.is_file() + { + return Err(io::Error::other( + "OCSF log destination must be a regular file", + )); + } + let mut options = OpenOptions::new(); + options.create(true).read(true).write(true).truncate(false); + #[cfg(unix)] + { + use std::os::unix::fs::OpenOptionsExt; + options.mode(0o600); + } + let mut file = options.open(&self.config.path)?; + self.day = Some(DateTime::::from(file.metadata()?.modified()?).date_naive()); + let discarded = recover_tail(&mut file)?; + if discarded > 0 { + metrics::counter!("openshell_ocsf_log_recovery_discarded_bytes_total") + .increment(discarded); + tracing::warn!( + discarded_bytes = discarded, + "OCSF JSONL incomplete tail removed; number of lost records unknown" + ); + } + self.file = Some(file); + Ok(()) + } + + fn append(&mut self, line: &[u8], today: NaiveDate) -> io::Result<()> { + if self.file.is_none() { + self.open()?; + } + if self.config.rotation == OcsfLogRotation::Daily && self.day != Some(today) { + self.file = None; + rotate(&self.config.path, self.day.unwrap())?; + if let Err(error) = + prune_rotated(&self.config.path, self.config.max_files.unwrap().get()) + { + metrics::counter!("openshell_ocsf_log_writer_errors_total").increment(1); + tracing::warn!(%error, "OCSF JSONL retention cleanup failed; rotated files may exceed max_files until the next rotation"); + } + self.open()?; + } + self.day = Some(today); + append_record(self.file.as_mut().unwrap(), line) + } +} + +trait RecordFile: Write + Seek { + fn truncate(&mut self, length: u64) -> io::Result<()>; +} + +impl RecordFile for File { + fn truncate(&mut self, length: u64) -> io::Result<()> { + self.set_len(length) + } +} + +fn append_record(file: &mut impl RecordFile, line: &[u8]) -> io::Result<()> { + let boundary = file.stream_position()?; + if let Err(error) = file.write_all(line).and_then(|()| file.flush()) { + file.truncate(boundary)?; + file.seek(SeekFrom::Start(boundary))?; + return Err(error); + } + Ok(()) +} + +fn recover_tail(file: &mut File) -> io::Result { + let length = file.metadata()?.len(); + let mut end = length; + let mut buffer = [0u8; 8192]; + while end > 0 { + let start = end.saturating_sub(buffer.len() as u64); + let count = usize::try_from(end - start).unwrap(); + file.seek(SeekFrom::Start(start))?; + file.read_exact(&mut buffer[..count])?; + if let Some(index) = buffer[..count].iter().rposition(|byte| *byte == b'\n') { + let boundary = start + index as u64 + 1; + file.set_len(boundary)?; + file.seek(SeekFrom::Start(boundary))?; + return Ok(length - boundary); + } + end = start; + } + file.set_len(0)?; + file.seek(SeekFrom::Start(0))?; + Ok(length) +} + +fn rotate(path: &Path, day: NaiveDate) -> io::Result<()> { + loop { + let mut name = path.as_os_str().to_os_string(); + name.push(format!(".{day}.{}", uuid::Uuid::new_v4())); + let archive = std::path::PathBuf::from(name); + match OpenOptions::new() + .write(true) + .create_new(true) + .open(&archive) + { + Ok(reservation) => { + drop(reservation); + if let Err(error) = std::fs::rename(path, &archive) { + let _ = std::fs::remove_file(&archive); + return Err(error); + } + return Ok(()); + } + Err(error) if error.kind() == io::ErrorKind::AlreadyExists => {} + Err(error) => return Err(error), + } + } +} + +fn prune_rotated(path: &Path, max_files: usize) -> io::Result<()> { + let parent = path + .parent() + .filter(|parent| !parent.as_os_str().is_empty()) + .unwrap_or_else(|| Path::new(".")); + let Some(stem) = path.file_name().and_then(|name| name.to_str()) else { + return Ok(()); + }; + let prefix = format!("{stem}."); + let mut rotated = Vec::new(); + for entry in std::fs::read_dir(parent)? { + let entry = entry?; + let name = entry.file_name(); + let Some(suffix) = name.to_str().and_then(|name| name.strip_prefix(&prefix)) else { + continue; + }; + let Some((day, unique)) = suffix.split_once('.') else { + continue; + }; + if let Ok(day) = NaiveDate::parse_from_str(day, "%Y-%m-%d") + && uuid::Uuid::parse_str(unique).is_ok() + && entry.file_type()?.is_file() + { + rotated.push((day, entry.metadata()?.modified()?, entry.path())); + } + } + rotated.sort(); + let excess = rotated.len().saturating_sub(max_files); + for (_, _, stale) in rotated.into_iter().take(excess) { + std::fs::remove_file(stale)?; + } + Ok(()) +} + +#[cfg(test)] +mod tests { + use super::*; + use openshell_ocsf::{ConfigStateChangeBuilder, ocsf_emit}; + use tracing_subscriber::prelude::*; + + fn config(path: &Path) -> OcsfLogConfig { + toml::from_str(&format!( + "path = {:?}\nrotation = 'never'\n", + path.display().to_string() + )) + .unwrap() + } + + #[tokio::test] + async fn gateway_native_events_are_written_with_console_off() { + let directory = tempfile::tempdir().unwrap(); + let path = directory.path().join("events.jsonl"); + let log = OcsfLog::start(config(&path)).unwrap(); + let event = ConfigStateChangeBuilder::new(&crate::gateway_ocsf::context("", "")) + .message("Gateway TLS configuration changed") + .build(); + let expected: serde_json::Value = + serde_json::from_str(&event.to_json_line().unwrap()).unwrap(); + let subscriber = tracing_subscriber::registry() + .with( + tracing_subscriber::fmt::layer() + .with_filter(tracing_subscriber::EnvFilter::new("off")), + ) + .with(log.layer()); + tracing::subscriber::with_default(subscriber, || { + tracing::info!("ordinary diagnostics must not enter the file"); + ocsf_emit!(event); + }); + log.shutdown().await; + let contents = std::fs::read_to_string(path).unwrap(); + assert_eq!(contents.lines().count(), 1); + assert_eq!( + serde_json::from_str::(&contents).unwrap(), + expected + ); + } + + #[tokio::test] + async fn gateway_events_use_the_configured_schema_version() { + let directory = tempfile::tempdir().unwrap(); + let path = directory.path().join("events.jsonl"); + let mut settings = config(&path); + settings.schema_version = Some(OcsfSchemaVersion::V1_3); + let log = OcsfLog::start(settings).unwrap(); + let event = ConfigStateChangeBuilder::new(&crate::gateway_ocsf::context("", "")) + .message("Gateway configuration changed") + .build(); + let subscriber = tracing_subscriber::registry().with(log.layer()); + tracing::subscriber::with_default(subscriber, || ocsf_emit!(event)); + log.shutdown().await; + + let event: serde_json::Value = + serde_json::from_str(&std::fs::read_to_string(path).unwrap()).unwrap(); + assert_eq!(event["metadata"]["version"], "1.3"); + assert_eq!( + event["unmapped"]["downgraded_from"], + openshell_ocsf::OCSF_VERSION + ); + } + + #[tokio::test] + async fn gateway_certificate_reload_reaches_jsonl_without_a_sandbox() { + let directory = tempfile::tempdir().unwrap(); + crate::tls_test_utils::generate_test_certs_with_ca(directory.path()); + let acceptor = crate::tls::TlsAcceptor::from_files( + &directory.path().join("server-cert.pem"), + &directory.path().join("server-key.pem"), + None, + false, + None, + None, + Vec::new(), + ) + .unwrap(); + let path = directory.path().join("events.jsonl"); + let log = OcsfLog::start(config(&path)).unwrap(); + let subscriber = tracing_subscriber::registry().with(log.layer()); + tracing::subscriber::with_default(subscriber, || acceptor.reload().unwrap()); + log.shutdown().await; + let contents = std::fs::read_to_string(path).unwrap(); + let event: serde_json::Value = serde_json::from_str(&contents).unwrap(); + assert_eq!( + event["message"], + "TLS certificate config reloaded successfully" + ); + assert_eq!(event["metadata"]["product"]["name"], "OpenShell Gateway"); + assert!(event.get("container").is_none()); + } + + #[test] + fn queue_limits_drop_incoming_records_without_evicting_history() { + let mut config = config(Path::new("unused")); + config.queue_capacity = std::num::NonZeroUsize::new(2).unwrap(); + config.queue_max_bytes = std::num::NonZeroUsize::new(8).unwrap(); + let collector = Collector::new(config); + assert!(collector.enqueue(b"1234\n".to_vec())); + assert!(!collector.enqueue(b"5678\n".to_vec())); + assert!(!collector.enqueue(b"oversized\n".to_vec())); + assert!(collector.enqueue(b"0\n".to_vec())); + assert!(!collector.enqueue(b"\n".to_vec())); + collector.close(); + assert!(!collector.enqueue(b"\n".to_vec())); + let state = collector.0.state.lock().unwrap(); + assert_eq!(state.queued_bytes, 7); + assert_eq!(state.queue.front().unwrap(), b"1234\n"); + for reason in ["queue_bytes", "oversized", "queue_capacity", "closed"] { + assert_eq!(state.losses[reason], 1); + } + } + + #[test] + fn reopening_removes_only_the_incomplete_tail() { + let directory = tempfile::tempdir().unwrap(); + let path = directory.path().join("events.jsonl"); + for prefix in [Vec::new(), b"{}\n".to_vec()] { + let mut damaged = prefix.clone(); + damaged.extend(vec![b'x'; 20_000]); + std::fs::write(&path, &damaged).unwrap(); + let mut file = OpenOptions::new() + .read(true) + .write(true) + .open(&path) + .unwrap(); + assert_eq!(recover_tail(&mut file).unwrap(), 20_000); + append_record(&mut file, b"{\"new\":true}\n").unwrap(); + let mut expected = prefix; + expected.extend(b"{\"new\":true}\n"); + assert_eq!(std::fs::read(&path).unwrap(), expected); + } + } + + struct FailingFile { + bytes: io::Cursor>, + remaining: usize, + fail_truncate: bool, + fail_flush: bool, + } + + impl Write for FailingFile { + fn write(&mut self, bytes: &[u8]) -> io::Result { + if self.remaining == 0 { + return Err(io::Error::other("disk full")); + } + let count = bytes.len().min(self.remaining); + self.remaining -= count; + self.bytes.write(&bytes[..count]) + } + fn flush(&mut self) -> io::Result<()> { + if self.fail_flush { + Err(io::Error::other("flush failed")) + } else { + Ok(()) + } + } + } + + impl Seek for FailingFile { + fn seek(&mut self, position: SeekFrom) -> io::Result { + self.bytes.seek(position) + } + } + + impl RecordFile for FailingFile { + fn truncate(&mut self, length: u64) -> io::Result<()> { + if self.fail_truncate { + return Err(io::Error::other("truncate failed")); + } + self.bytes + .get_mut() + .truncate(usize::try_from(length).unwrap()); + Ok(()) + } + } + + #[test] + fn partial_writes_do_not_replay_completed_records() { + let mut file = FailingFile { + bytes: io::Cursor::new(Vec::new()), + remaining: 6, + fail_truncate: false, + fail_flush: false, + }; + append_record(&mut file, b"{}\n").unwrap(); + assert!(append_record(&mut file, b"{\"partial\":true}\n").is_err()); + assert_eq!(file.bytes.get_ref(), b"{}\n"); + file.remaining = 100; + append_record(&mut file, b"{\"later\":true}\n").unwrap(); + assert_eq!(file.bytes.get_ref(), b"{}\n{\"later\":true}\n"); + file.fail_flush = true; + assert!(append_record(&mut file, b"{}\n").is_err()); + assert_eq!(file.bytes.get_ref(), b"{}\n{\"later\":true}\n"); + file.fail_truncate = true; + assert!(append_record(&mut file, b"{}\n").is_err()); + } + + #[test] + fn same_day_archives_never_overwrite_and_pruning_ignores_unrelated_files() { + let directory = tempfile::tempdir().unwrap(); + let path = directory.path().join("events.jsonl"); + let today = Utc::now().date_naive(); + for contents in ["first\n", "second\n", "third\n"] { + std::fs::write(&path, contents).unwrap(); + rotate(&path, today).unwrap(); + } + let mut contents: Vec<_> = std::fs::read_dir(directory.path()) + .unwrap() + .map(|entry| std::fs::read_to_string(entry.unwrap().path()).unwrap()) + .collect(); + contents.sort(); + assert_eq!(contents, ["first\n", "second\n", "third\n"]); + let unrelated = directory.path().join("events.jsonl.notes"); + std::fs::write(&unrelated, "keep").unwrap(); + prune_rotated(&path, 1).unwrap(); + assert_eq!(std::fs::read_dir(directory.path()).unwrap().count(), 2); + assert_eq!(std::fs::read_to_string(unrelated).unwrap(), "keep"); + } + + #[test] + fn daily_rotation_keeps_previous_day_records_out_of_the_active_file() { + let directory = tempfile::tempdir().unwrap(); + let path = directory.path().join("events.jsonl"); + let mut settings = config(&path); + settings.rotation = OcsfLogRotation::Daily; + settings.max_files = std::num::NonZeroUsize::new(1); + let mut writer = FileWriter::new(settings); + let today = Utc::now().date_naive(); + writer.append(b"{}\n", today).unwrap(); + writer + .append(b"{\"next\":true}\n", today.succ_opt().unwrap()) + .unwrap(); + assert_eq!(std::fs::read_to_string(&path).unwrap(), "{\"next\":true}\n"); + assert_eq!(std::fs::read_dir(directory.path()).unwrap().count(), 2); + } + + #[tokio::test] + async fn invalid_destination_is_visible_outside_the_output_file() { + let directory = tempfile::tempdir().unwrap(); + let log = OcsfLog::start(config(directory.path())).unwrap(); + let collector = log.collector(); + collector.enqueue(b"{}\n".to_vec()); + log.shutdown().await; + let state = collector.0.state.lock().unwrap(); + assert_eq!(state.in_flight, 0); + assert_eq!(state.losses["write_uncertain"], 1); + } + + #[tokio::test] + async fn shutdown_budget_accounts_queued_and_uncertain_records() { + let collector = Collector::new(config(Path::new("unused"))); + collector.enqueue(b"{}\n".to_vec()); + collector.0.state.lock().unwrap().in_flight = 2; + let (_sender, done) = tokio::sync::oneshot::channel(); + let mut log = OcsfLog { + collector: collector.clone(), + done, + }; + log.shutdown_with_budget(Duration::from_millis(1)).await; + collector.complete(None); + let state = collector.0.state.lock().unwrap(); + assert!(state.queue.is_empty()); + assert_eq!(state.losses["shutdown"], 1); + assert_eq!(state.losses["shutdown_uncertain"], 2); + assert_eq!(state.in_flight, 0); + } +} diff --git a/crates/openshell-server/src/persistence/mod.rs b/crates/openshell-server/src/persistence/mod.rs index c3a6649252..2513e80c10 100644 --- a/crates/openshell-server/src/persistence/mod.rs +++ b/crates/openshell-server/src/persistence/mod.rs @@ -360,6 +360,7 @@ impl Store { #[allow(clippy::too_many_arguments)] #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.put_if", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id, object.name = %name, workspace = %workspace) )] @@ -394,6 +395,7 @@ impl Store { /// anything must use [`Self::put_if`], which is always durable. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.create_relaxed", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id, object.name = %name, workspace = %workspace) )] @@ -429,6 +431,7 @@ impl Store { /// * `Err(Conflict)` - Resource version mismatch #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.delete_if", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id) )] @@ -445,6 +448,7 @@ impl Store { #[allow(clippy::too_many_arguments)] #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.put_scoped", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id, object.name = %name, workspace = %workspace, scope = %scope) )] @@ -477,6 +481,7 @@ impl Store { #[allow(clippy::too_many_arguments)] #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.create_scoped", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id, object.name = %name, workspace = %workspace, scope = %scope) )] @@ -510,6 +515,7 @@ impl Store { #[allow(clippy::too_many_arguments)] #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.create_if_workspace_count_below", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id, object.name = %name, workspace = %workspace, max_count = max_count) )] @@ -537,6 +543,7 @@ impl Store { /// Fetch an object by id. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.get", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id) )] @@ -551,6 +558,7 @@ impl Store { /// Fetch an object by name within an object type and workspace. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields( otel.name = "store.get_by_name", otel.status_code = tracing::field::Empty, @@ -571,6 +579,7 @@ impl Store { /// Delete an object by id. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.delete", otel.status_code = tracing::field::Empty, object_type = %object_type, object.id = %id) )] @@ -581,6 +590,7 @@ impl Store { /// Delete objects of one type by id in bounded, set-based statements. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields( otel.name = "store.delete_many", @@ -597,6 +607,7 @@ impl Store { /// Count objects of a given type within a workspace. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.count_in_workspace", otel.status_code = tracing::field::Empty, object_type = %object_type, workspace = %workspace) )] @@ -611,6 +622,7 @@ impl Store { /// Delete all objects of a given type within a workspace. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.delete_all_in_workspace", otel.status_code = tracing::field::Empty, object_type = %object_type, workspace = %workspace) )] @@ -625,6 +637,7 @@ impl Store { /// Delete all objects of a given type with a matching scope. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.delete_by_scope", otel.status_code = tracing::field::Empty, object_type = %object_type, scope = %scope) )] @@ -635,6 +648,7 @@ impl Store { /// Delete an object by name within an object type and workspace. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.delete_by_name", otel.status_code = tracing::field::Empty, object_type = %object_type, workspace = %workspace, object.name = %name) )] @@ -650,6 +664,7 @@ impl Store { /// List objects by type and workspace. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.list", otel.status_code = tracing::field::Empty, object_type = %object_type, workspace = %workspace) )] @@ -666,6 +681,7 @@ impl Store { /// List objects by type across all workspaces. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.list_by_type", otel.status_code = tracing::field::Empty, object_type = %object_type) )] @@ -681,6 +697,7 @@ impl Store { /// List workspace objects after a stable cursor, without offset drift. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields( otel.name = "store.list_after", @@ -702,6 +719,7 @@ impl Store { /// List objects across workspaces after a stable cursor, without offset drift. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields( otel.name = "store.list_by_type_after", @@ -827,6 +845,7 @@ impl Store { /// UUIDs which are globally unique. Revisit if non-UUID scopes are introduced. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.list_by_scope", otel.status_code = tracing::field::Empty, object_type = %object_type, scope = %scope) )] @@ -844,6 +863,7 @@ impl Store { /// Label selector format: "key1=value1,key2=value2" (comma-separated equality matches). #[tracing::instrument( name = "store", + level = "debug", skip_all, fields( otel.name = "store.list_with_selector", otel.status_code = tracing::field::Empty, @@ -912,6 +932,7 @@ impl Store { /// List objects by type across all workspaces with label selector filtering. #[tracing::instrument( name = "store", + level = "debug", skip_all, fields(otel.name = "store.list_all_with_selector", otel.status_code = tracing::field::Empty, object_type = %object_type, label_selector = %label_selector) )] @@ -1310,6 +1331,7 @@ pub fn parse_label_selector(selector: &str) -> PersistenceResult, interval: }); } -#[tracing::instrument( - name = "refresh", - skip_all, - fields( - otel.name = "refresh.provider_credentials", - watched_count = tracing::field::Empty, - due_count = tracing::field::Empty, - ) -)] async fn run_refresh_worker_tick( store: &Store, credentials: Option<&crate::credentials::CredentialRuntime>, compute: Option<&crate::compute::ComputeRuntime>, ) -> Result<(), Status> { let now_ms = current_time_ms(); - let states = list_all_refresh_states(store).await.inspect_err(|_| { - crate::otel_tracing::mark_error(&tracing::Span::current()); - })?; + let states = list_all_refresh_states(store).await?; let watched_count = states.len(); let due_count = states .iter() @@ -1958,13 +1947,44 @@ async fn run_refresh_worker_tick( .iter() .filter(|state| state.status == "rotation_requested") .count(); - let span = tracing::Span::current(); - span.record("watched_count", watched_count); - span.record("due_count", due_count); info!( watched_count, due_count, rotation_requested_count, "provider credential refresh worker sweep" ); + if !states + .iter() + .any(|state| refresh_state_has_work(state, now_ms)) + { + return Ok(()); + } + let span = tracing::info_span!( + "refresh", + otel.name = "refresh.provider_credentials", + watched_count, + due_count, + ); + Box::pin(refresh_states(store, credentials, compute, states, now_ms).instrument(span)).await; + Ok(()) +} + +fn refresh_state_has_work(state: &StoredProviderCredentialRefreshState, now_ms: i64) -> bool { + state + .metadata + .as_ref() + .is_some_and(|metadata| metadata.deletion_time.is_some()) + || !state.pending_secret_deletions.is_empty() + || state.next_refresh_at_ms <= 0 + || state.next_refresh_at_ms <= now_ms + || state.status == "rotation_requested" +} + +async fn refresh_states( + store: &Store, + credentials: Option<&crate::credentials::CredentialRuntime>, + compute: Option<&crate::compute::ComputeRuntime>, + states: Vec, + now_ms: i64, +) { for state in states { if state .metadata @@ -2079,7 +2099,6 @@ async fn run_refresh_worker_tick( ); } } - Ok(()) } #[cfg(test)] @@ -3560,11 +3579,9 @@ mod tests { assert_eq!(credentials.stored_credential_count(), Some(0)); } - /// The worker ticks on a timer with no inbound request, so without a span - /// of its own its store reads export as anonymous single-span traces. #[tokio::test] #[ignore = "flaky under concurrent test execution"] - async fn refresh_worker_ticks_are_roots_and_store_operations_have_parents() { + async fn refresh_worker_records_a_root_span_only_when_a_state_has_work() { use crate::otel_tracing::test_exporter; let store = test_store().await; @@ -3573,27 +3590,38 @@ mod tests { Box::pin(run_refresh_worker_tick(&store, None, None)) .await .unwrap(); + assert!( + traced + .spans_named("refresh.provider_credentials") + .is_empty(), + "an idle tick records no refresh span" + ); - let spans = traced.finished_spans(); - let root = spans - .iter() - .find(|s| s.name == "refresh.provider_credentials") - .unwrap_or_else(|| { - panic!( - "the tick records a span of its own, got {:?}", - spans.iter().map(|s| &s.name).collect::>() - ) - }); - - test_exporter::assert_is_root(root); - let store_span = spans - .iter() - .find(|span| { - span.name.starts_with("store.") - && span.span_context.trace_id() == root.span_context.trace_id() - }) - .expect("the tick records its store operation"); - test_exporter::assert_has_parent(store_span); + let provider = provider("my-external", "outlook"); + store.put_message(&provider).await.unwrap(); + let state = new_refresh_state( + &provider, + "default", + "MS_GRAPH_ACCESS_TOKEN", + NewRefreshStateConfig { + additional_output_keys: HashMap::new(), + strategy: ProviderCredentialRefreshStrategy::External, + material: HashMap::new(), + secret_material_keys: Vec::new(), + expires_at_ms: 0, + token_url: String::new(), + scopes: Vec::new(), + refresh_before: None, + max_lifetime: None, + }, + ) + .unwrap(); + put_refresh_state(&store, &state).await.unwrap(); + + Box::pin(run_refresh_worker_tick(&store, None, None)) + .await + .unwrap(); + test_exporter::assert_is_root(&traced.span_named("refresh.provider_credentials")); } #[test] diff --git a/crates/openshell-server/src/sandbox_watch.rs b/crates/openshell-server/src/sandbox_watch.rs index 679912d352..7bc3ce8728 100644 --- a/crates/openshell-server/src/sandbox_watch.rs +++ b/crates/openshell-server/src/sandbox_watch.rs @@ -154,6 +154,43 @@ pub fn lag_warning_event(n: u64) -> openshell_core::proto::SandboxStreamEvent { } } +/// Build the warning payload emitted when the cross-source coverage floor +/// withholds events on connect. +/// +/// Unlike broadcast lag, nothing was evicted here: the withheld events are +/// still sitting in a bus's tail, excluded only because a sibling source's +/// shallower replay depth means handing them out would let the client's +/// single shared cursor imply coverage that sibling can't back. They won't +/// be resent later in this stream -- a resume replays past the cursor this +/// connect settles on, and these are at or below it by construction -- so +/// the client needs to know now, not discover it silently after a reconnect. +pub fn coverage_gap_warning(log_withheld: usize, platform_withheld: usize) -> SandboxStreamWarning { + SandboxStreamWarning { + message: format!( + "resume cursor cannot cover every followed source's backlog; withheld {log_withheld} \ + log line(s) and {platform_withheld} platform event(s) that a shallower sibling source \ + had not replayed. Restart the watch with a deeper log_tail_lines/event_tail if you need them; \ + they will not be recoverable from this cursor." + ), + } +} + +/// Wrap [`coverage_gap_warning`] in a `SandboxStreamEvent` ready to send. +pub fn coverage_gap_warning_event( + log_withheld: usize, + platform_withheld: usize, +) -> openshell_core::proto::SandboxStreamEvent { + use openshell_core::proto::sandbox_stream_event::Payload; + openshell_core::proto::SandboxStreamEvent { + payload: Some(Payload::Warning(coverage_gap_warning( + log_withheld, + platform_withheld, + ))), + // Warnings are not part of the resumable log/platform sequence. + cursor: String::new(), + } +} + #[cfg(test)] mod tests { use super::*; @@ -226,6 +263,35 @@ mod tests { } } + #[test] + fn coverage_gap_warning_reports_withheld_counts() { + let warning = coverage_gap_warning(2, 1); + assert!( + warning.message.contains('2'), + "message: {}", + warning.message + ); + assert!( + warning.message.contains('1'), + "message: {}", + warning.message + ); + } + + #[test] + fn coverage_gap_warning_event_wraps_warning_payload_with_empty_cursor() { + use openshell_core::proto::sandbox_stream_event::Payload; + let evt = coverage_gap_warning_event(2, 1); + assert!(evt.cursor.is_empty()); + match evt.payload { + Some(Payload::Warning(w)) => { + assert!(w.message.contains('2')); + assert!(w.message.contains('1')); + } + other => panic!("expected Warning payload, got {other:?}"), + } + } + // Broadcast lag is recoverable at the tokio layer: after `Lagged`, the same // receiver keeps yielding the oldest surviving messages instead of closing. #[tokio::test] diff --git a/crates/openshell-server/src/service_routing.rs b/crates/openshell-server/src/service_routing.rs index 3d49d74d70..c139d282ab 100644 --- a/crates/openshell-server/src/service_routing.rs +++ b/crates/openshell-server/src/service_routing.rs @@ -9,20 +9,21 @@ use axum::{ }; use http::{HeaderMap, HeaderValue, Method, Request, Response, StatusCode, header}; use hyper_util::rt::TokioIo; +use openshell_core::ObjectId; use openshell_core::config::ServiceRoutingConfig; -use openshell_core::proto::{Sandbox, SandboxPhase, ServiceEndpoint, TcpRelayTarget, relay_open}; -use openshell_core::{ObjectId, VERSION}; +use openshell_core::proto::{ + Sandbox, SandboxPhase, ServiceAuthorizationMode, ServiceEndpoint, TcpRelayTarget, relay_open, +}; use openshell_ocsf::{ ActionId, ActivityId, ConfigStateChangeBuilder, DispositionId, Endpoint, EventContext, HttpActivityBuilder, HttpRequest, HttpResponse as OcsfHttpResponse, NetworkActivityBuilder, - OCSF_TARGET, OcsfEvent, SeverityId, StateId, StatusId, Url as OcsfUrl, + OcsfEvent, SeverityId, StateId, StatusId, Url as OcsfUrl, }; use std::collections::HashMap; -use std::net::{IpAddr, Ipv4Addr}; use std::sync::{Arc, Mutex}; use std::time::{Duration, Instant}; use tokio::io::AsyncWriteExt; -use tracing::{info, warn}; +use tracing::warn; use crate::ServerState; use crate::persistence::{ObjectType, Store}; @@ -429,6 +430,19 @@ async fn proxy_to_endpoint( ); return Err(err); } + let authorization_mode = effective_authorization_mode(endpoint.authorization_mode); + if validate_application_authorization(&req, authorization_mode).is_err() { + let err = ServiceRouteError::invalid_request(); + emit_service_http_failure( + &state, + &req, + &sandbox_name, + &service_name, + Some(&endpoint), + &err, + ); + return Err(err); + } let websocket_upgrade = is_websocket_upgrade(&req); let downstream_upgrade = websocket_upgrade.then(|| hyper::upgrade::on(&mut req)); @@ -447,7 +461,7 @@ async fn proxy_to_endpoint( None => open_upstream(&state, &sandbox, &endpoint, target_port, websocket_upgrade).await?, }; - let upstream = build_upstream_request(req, target_port, websocket_upgrade)?; + let upstream = build_upstream_request(req, target_port, websocket_upgrade, authorization_mode)?; let replay = reused.then(|| replayable_request(&upstream)).flatten(); let first_attempt = if reused { sender.try_send_request(upstream).await.map_err(|mut err| { @@ -627,6 +641,7 @@ fn build_upstream_request( req: Request, target_port: u16, preserve_upgrade_headers: bool, + authorization_mode: ServiceAuthorizationMode, ) -> Result, ServiceRouteError> { let (parts, body) = req.into_parts(); let path = parts.uri.path_and_query().map_or("/", |path| path.as_str()); @@ -645,7 +660,7 @@ fn build_upstream_request( for (name, value) in &parts.headers { if (is_hop_by_hop_header(name) && !(preserve_upgrade_headers && is_websocket_hop_by_hop_header(name))) - || is_gateway_auth_header(name) + || is_gateway_auth_header(name, authorization_mode) { continue; } @@ -725,15 +740,83 @@ fn is_websocket_hop_by_hop_header(name: &header::HeaderName) -> bool { matches!(name.as_str(), "connection" | "upgrade") } -fn is_gateway_auth_header(name: &header::HeaderName) -> bool { - matches!( - name.as_str(), - "authorization" - | "cf-access-jwt-assertion" - | "x-forwarded-client-cert" - | "x-ssl-client-cert" - | "x-client-cert" - ) +pub fn effective_authorization_mode(value: i32) -> ServiceAuthorizationMode { + match ServiceAuthorizationMode::try_from(value) { + Ok(ServiceAuthorizationMode::BearerPassthrough) => { + ServiceAuthorizationMode::BearerPassthrough + } + Ok(ServiceAuthorizationMode::Unspecified | ServiceAuthorizationMode::Strip) | Err(_) => { + ServiceAuthorizationMode::Strip + } + } +} + +fn authorization_mode_label(value: i32) -> &'static str { + match effective_authorization_mode(value) { + ServiceAuthorizationMode::BearerPassthrough => "bearer_passthrough", + ServiceAuthorizationMode::Unspecified | ServiceAuthorizationMode::Strip => "strip", + } +} + +fn validate_application_authorization( + req: &Request, + authorization_mode: ServiceAuthorizationMode, +) -> Result<(), ServiceRouteError> { + if authorization_mode != ServiceAuthorizationMode::BearerPassthrough { + return Ok(()); + } + + let mut values = req.headers().get_all(header::AUTHORIZATION).iter(); + let Some(value) = values.next() else { + return Ok(()); + }; + if values.next().is_some() { + return Err(ServiceRouteError::invalid_request()); + } + + let value = value + .to_str() + .map_err(|_| ServiceRouteError::invalid_request())?; + let Some((scheme, credential)) = value.split_once(' ') else { + return Err(ServiceRouteError::invalid_request()); + }; + let credential = credential.trim_start_matches(' '); + if !scheme.eq_ignore_ascii_case("bearer") || !is_bearer_token68(credential) { + return Err(ServiceRouteError::invalid_request()); + } + Ok(()) +} + +fn is_bearer_token68(value: &str) -> bool { + let mut saw_data = false; + let mut saw_padding = false; + for byte in value.bytes() { + if byte == b'=' { + saw_padding = true; + } else if !saw_padding + && (byte.is_ascii_alphanumeric() + || matches!(byte, b'-' | b'.' | b'_' | b'~' | b'+' | b'/')) + { + saw_data = true; + } else { + return false; + } + } + saw_data +} + +fn is_gateway_auth_header( + name: &header::HeaderName, + authorization_mode: ServiceAuthorizationMode, +) -> bool { + match name.as_str() { + "authorization" => authorization_mode != ServiceAuthorizationMode::BearerPassthrough, + "cf-access-jwt-assertion" + | "x-forwarded-client-cert" + | "x-ssl-client-cert" + | "x-client-cert" => true, + _ => false, + } } fn sanitize_cookie_header(value: &HeaderValue) -> Option { @@ -760,12 +843,12 @@ fn is_gateway_auth_cookie(name: &str) -> bool { pub fn emit_service_endpoint_config_event(endpoint: &ServiceEndpoint, url: &str, created: bool) { let event = build_service_endpoint_config_event(endpoint, url, created); - emit_gateway_ocsf_event(&endpoint.sandbox_id, event); + openshell_ocsf::ocsf_emit!(event); } pub fn emit_service_endpoint_delete_event(endpoint: &ServiceEndpoint) { let event = build_service_endpoint_delete_event(endpoint); - emit_gateway_ocsf_event(&endpoint.sandbox_id, event); + openshell_ocsf::ocsf_emit!(event); } pub fn emit_cross_origin_service_http_rejection(state: &ServerState, req: &Request) { @@ -801,13 +884,12 @@ fn emit_service_http_failure( endpoint, err, ); - let sandbox_id = endpoint.map_or("", |endpoint| endpoint.sandbox_id.as_str()); - emit_gateway_ocsf_event(sandbox_id, event); + openshell_ocsf::ocsf_emit!(event); } fn emit_service_relay_failure(endpoint: &ServiceEndpoint, target_port: u16, reason: &str) { let event = build_service_relay_failure_event(endpoint, target_port, reason); - emit_gateway_ocsf_event(&endpoint.sandbox_id, event); + openshell_ocsf::ocsf_emit!(event); } fn build_service_endpoint_config_event( @@ -832,7 +914,11 @@ fn build_service_endpoint_config_event( )) .unmapped("endpoint_name", endpoint_name(endpoint)) .unmapped("service_name", endpoint.name.clone()) - .unmapped("target_port", u64::from(endpoint.target_port)); + .unmapped("target_port", u64::from(endpoint.target_port)) + .unmapped( + "authorization_mode", + authorization_mode_label(endpoint.authorization_mode), + ); if !url.is_empty() { builder = builder.unmapped("url", url.to_string()); @@ -851,6 +937,10 @@ fn build_service_endpoint_delete_event(endpoint: &ServiceEndpoint) -> OcsfEvent .unmapped("endpoint_name", endpoint_name(endpoint)) .unmapped("service_name", endpoint.name.clone()) .unmapped("target_port", u64::from(endpoint.target_port)) + .unmapped( + "authorization_mode", + authorization_mode_label(endpoint.authorization_mode), + ) .build() } @@ -923,25 +1013,8 @@ fn build_service_relay_failure_event( .build() } -fn emit_gateway_ocsf_event(sandbox_id: &str, event: OcsfEvent) { - let message = event.format_shorthand(); - info!( - target: OCSF_TARGET, - sandbox_id = %sandbox_id, - message = %message - ); -} - fn gateway_ocsf_ctx(sandbox_id: &str, sandbox_name: &str) -> EventContext { - EventContext { - sandbox_id: sandbox_id.to_string(), - sandbox_name: sandbox_name.to_string(), - container_image: "openshell/gateway".to_string(), - hostname: "openshell-gateway".to_string(), - product_version: VERSION.to_string(), - proxy_ip: IpAddr::V4(Ipv4Addr::LOCALHOST), - proxy_port: 0, - } + crate::gateway_ocsf::context(sandbox_id, sandbox_name) } fn endpoint_name(endpoint: &ServiceEndpoint) -> String { @@ -1009,6 +1082,7 @@ mod tests { name: "web".to_string(), target_port: 8080, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, } } @@ -1252,6 +1326,7 @@ mod tests { assert_eq!(json["unmapped"]["endpoint_name"], "my-sandbox--web"); assert_eq!(json["unmapped"]["service_name"], "web"); assert_eq!(json["unmapped"]["target_port"], 8080); + assert_eq!(json["unmapped"]["authorization_mode"], "strip"); assert!( event .format_shorthand() @@ -1268,6 +1343,7 @@ mod tests { assert_eq!(json["unmapped"]["endpoint_name"], "my-sandbox--web"); assert_eq!(json["unmapped"]["service_name"], "web"); assert_eq!(json["unmapped"]["target_port"], 8080); + assert_eq!(json["unmapped"]["authorization_mode"], "strip"); assert!( event .format_shorthand() @@ -1328,7 +1404,8 @@ mod tests { .body(Body::empty()) .unwrap(); - let upstream = build_upstream_request(request, 8080, false).unwrap(); + let upstream = + build_upstream_request(request, 8080, false, ServiceAuthorizationMode::Strip).unwrap(); assert_eq!(upstream.uri(), "/path"); assert!(!upstream.headers().contains_key(header::AUTHORIZATION)); @@ -1341,6 +1418,145 @@ mod tests { assert_eq!(upstream.headers()["x-app-header"], "kept"); } + #[test] + fn unspecified_endpoint_authorization_mode_strips_authorization() { + let request = Request::builder() + .uri("/path") + .header(header::AUTHORIZATION, "Bearer application-token") + .body(Body::empty()) + .unwrap(); + + let mode = effective_authorization_mode(ServiceAuthorizationMode::Unspecified as i32); + validate_application_authorization(&request, mode).unwrap(); + let upstream = build_upstream_request(request, 8080, false, mode).unwrap(); + + assert_eq!(mode, ServiceAuthorizationMode::Strip); + assert!(!upstream.headers().contains_key(header::AUTHORIZATION)); + } + + #[test] + fn bearer_passthrough_preserves_valid_authorization_and_strips_gateway_identity() { + let request = Request::builder() + .uri("/path") + .header(header::AUTHORIZATION, "bEaReR application-token") + .header("cf-access-jwt-assertion", "edge-token") + .header("x-forwarded-client-cert", "cert") + .header(header::PROXY_AUTHORIZATION, "Basic proxy-secret") + .header( + header::COOKIE, + "theme=dark; CF_Authorization=edge-cookie; app=session", + ) + .body(Body::empty()) + .unwrap(); + + let mode = ServiceAuthorizationMode::BearerPassthrough; + validate_application_authorization(&request, mode).unwrap(); + let upstream = build_upstream_request(request, 8080, false, mode).unwrap(); + + assert_eq!( + upstream.headers()[header::AUTHORIZATION], + "bEaReR application-token" + ); + assert!(!upstream.headers().contains_key("cf-access-jwt-assertion")); + assert!(!upstream.headers().contains_key("x-forwarded-client-cert")); + assert!(!upstream.headers().contains_key(header::PROXY_AUTHORIZATION)); + assert_eq!( + upstream.headers()[header::COOKIE], + "theme=dark; app=session" + ); + } + + #[test] + fn bearer_passthrough_allows_missing_authorization() { + let request = Request::builder().uri("/path").body(Body::empty()).unwrap(); + + let mode = ServiceAuthorizationMode::BearerPassthrough; + validate_application_authorization(&request, mode).unwrap(); + let upstream = build_upstream_request(request, 8080, false, mode).unwrap(); + + assert!(!upstream.headers().contains_key(header::AUTHORIZATION)); + } + + #[test] + fn bearer_passthrough_rejects_ambiguous_or_malformed_authorization() { + for value in [ + "", + "Bearer", + "Bearer ", + "Basic abc", + "Bearer abc extra", + "Bearer abc,def", + "Bearer abc=def", + "Bearer\tabc", + " Bearer abc", + ] { + let request = Request::builder() + .uri("/path") + .header(header::AUTHORIZATION, value) + .body(Body::empty()) + .unwrap(); + assert!( + validate_application_authorization( + &request, + ServiceAuthorizationMode::BearerPassthrough, + ) + .is_err() + ); + } + + for value in ["Bearer abc", "bearer abc-._~+/==", "Bearer abc"] { + let request = Request::builder() + .uri("/path") + .header(header::AUTHORIZATION, value) + .body(Body::empty()) + .unwrap(); + validate_application_authorization( + &request, + ServiceAuthorizationMode::BearerPassthrough, + ) + .unwrap(); + } + + let mut request = Request::builder().uri("/path").body(Body::empty()).unwrap(); + request.headers_mut().append( + header::AUTHORIZATION, + HeaderValue::from_static("Bearer first"), + ); + request.headers_mut().append( + header::AUTHORIZATION, + HeaderValue::from_static("Bearer second"), + ); + assert!( + validate_application_authorization( + &request, + ServiceAuthorizationMode::BearerPassthrough, + ) + .is_err() + ); + } + + #[test] + fn authorization_value_is_not_in_service_routing_events() { + const SENTINEL: &str = "never-log-this-capability"; + let request = Request::builder() + .uri("/private") + .header(header::AUTHORIZATION, format!("Bearer {SENTINEL}")) + .body(Body::empty()) + .unwrap(); + let err = ServiceRouteError::invalid_request(); + let event = build_service_http_failure_event( + 18080, + &request, + "my-sandbox", + "web", + Some(&endpoint()), + &err, + ); + + assert!(!event.to_json().unwrap().to_string().contains(SENTINEL)); + assert!(!event.format_shorthand().contains(SENTINEL)); + } + #[test] fn detects_websocket_upgrade_request() { let request = Request::builder() @@ -1365,7 +1581,8 @@ mod tests { .body(Body::empty()) .unwrap(); - let upstream = build_upstream_request(request, 8080, true).unwrap(); + let upstream = + build_upstream_request(request, 8080, true, ServiceAuthorizationMode::Strip).unwrap(); assert_eq!(upstream.uri(), "/chat?session=main"); assert_eq!(upstream.headers()[header::CONNECTION], "Upgrade"); @@ -1374,6 +1591,30 @@ mod tests { assert_eq!(upstream.headers()[header::HOST], "127.0.0.1:8080"); } + #[test] + fn bearer_passthrough_preserves_authorization_on_websocket_upgrade() { + let request = Request::builder() + .method(Method::GET) + .uri("/chat") + .header(header::CONNECTION, "Upgrade") + .header(header::UPGRADE, "websocket") + .header("sec-websocket-key", "abc") + .header(header::AUTHORIZATION, "Bearer application-token") + .body(Body::empty()) + .unwrap(); + + let mode = ServiceAuthorizationMode::BearerPassthrough; + validate_application_authorization(&request, mode).unwrap(); + let upstream = build_upstream_request(request, 8080, true, mode).unwrap(); + + assert_eq!( + upstream.headers()[header::AUTHORIZATION], + "Bearer application-token" + ); + assert_eq!(upstream.headers()[header::CONNECTION], "Upgrade"); + assert_eq!(upstream.headers()[header::UPGRADE], "websocket"); + } + #[tokio::test] async fn load_endpoint_uses_workspace_for_lookup() { let store = crate::persistence::test_store().await; @@ -1394,6 +1635,7 @@ mod tests { name: "web".to_string(), target_port: 8080, domain: true, + authorization_mode: ServiceAuthorizationMode::Strip as i32, }; store.put_message(&ep).await.unwrap(); @@ -1406,7 +1648,6 @@ mod tests { "should not find endpoint in wrong workspace" ); } - /// Returns a live upstream plus the sandbox end of the connection. Hold /// the returned half: dropping it closes the upstream. async fn test_upstream() -> (UpstreamSender, tokio::io::DuplexStream) { @@ -1592,4 +1833,48 @@ mod tests { "eviction must be scoped to one endpoint" ); } + + /// Captures structured OCSF events during tracing dispatch. + #[derive(Clone, Default)] + struct ProbeLayer { + seen: Arc>>>, + } + + impl tracing_subscriber::Layer for ProbeLayer { + fn on_event( + &self, + event: &tracing::Event<'_>, + _ctx: tracing_subscriber::layer::Context<'_, S>, + ) { + if event.metadata().target() == openshell_ocsf::OCSF_TARGET { + self.seen + .lock() + .unwrap() + .push(openshell_ocsf::clone_current_event()); + } + } + } + + #[test] + fn gateway_ocsf_events_expose_the_structured_event_to_layers() { + use tracing_subscriber::layer::SubscriberExt; + + let probe = ProbeLayer::default(); + let subscriber = tracing_subscriber::registry().with(probe.clone()); + + let endpoint = endpoint(); + let expected = build_service_endpoint_config_event(&endpoint, "https://example.test", true); + let expected_shorthand = expected.format_shorthand(); + + tracing::subscriber::with_default(subscriber, || { + emit_service_endpoint_config_event(&endpoint, "https://example.test", true); + }); + + let seen = probe.seen.lock().unwrap(); + assert_eq!(seen.len(), 1, "expected exactly one OCSF tracing event"); + let event = seen[0] + .as_ref() + .expect("structured OCSF event should be reachable from the layer"); + assert_eq!(event.format_shorthand(), expected_shorthand); + } } diff --git a/crates/openshell-server/src/storage_proto.rs b/crates/openshell-server/src/storage_proto.rs index 16a733415c..e6d8730aa0 100644 --- a/crates/openshell-server/src/storage_proto.rs +++ b/crates/openshell-server/src/storage_proto.rs @@ -118,19 +118,27 @@ mod tests { const STORAGE_V1_SCHEMA_SHA256: &str = "d68401809d8cea445c35233ef32412bbd041cb2ac5acaf368a0d0bf74d2ddf17"; - // Carries this branch's exec request IDs together with main's opaque watch - // cursor and well-known time types. These unreleased public-only fields add - // no messages or enums and touch no stored type, so the durable and overlap - // fingerprints below remain unchanged. + // Restart policy is stored in SandboxSpec, and the count and well-known + // timestamps are stored in SandboxStatus. Legacy payloads decode with + // Unspecified (treated as Never), zero count, and absent timestamps. + // ProviderProfileFile is reachable from stored provider profiles. Its + // additive declaration changes the durable and public/durable overlap + // inventories; the provider-environment file map is public-only. The + // request has no provider-file capability field: older supervisors ignore + // the additive file map while retaining the rest of the response. + // Service authorization also extends both schemas additively. Legacy + // payloads retain the safe Strip default. const PUBLIC_RPC_SCHEMA_SHA256: &str = - "8fb59b0932ec2f227fdec2d46b6204925e79595695810a247bef731ddd632594"; + "2e156c6ad3c8eb51bcd30dc13b173fe339b38207a1b1f98f7be2e0cad8e3bd45"; const DURABLE_SCHEMA_SHA256: &str = - "9eeaa29dfba187bff69fb7bc4f9a13a0f1d7be3f7049a38c8f0e20ce77ec7d8b"; + "38165d9d76f49fcfe98a12f241e032838a2376c1d1a87ea2796fd33b9b1a3541"; const PUBLIC_DURABLE_OVERLAP_SHA256: &str = - "a6e97fdde30c439ffaa03c2952a43033f8ea338fed6b1456ebe2d7d8af14e834"; + "761dea31a521b0650840fe2a823ad6e36a265ed323ba4506889781d630df0ee3"; // A persisted Sandbox without endpoint status retains its lifecycle fields; // the absent repeated field decodes empty and needs no database rewrite. const SANDBOX_WITHOUT_ENDPOINT_STATUS: &str = "0a1e0a0a73616e64626f782d6964120773616e64626f783a0764656661756c741a2b0a0773616e64626f782a0d0a05526561647912045472756530023807420d73757065727669736f722d6964"; + // ServiceEndpoint encoded before authorization_mode field 7 existed. + const SERVICE_ENDPOINT_WITHOUT_AUTHORIZATION_MODE: &str = "0a260a0b656e64706f696e742d6964120c73616e64626f782d2d77656228073a0764656661756c74120a73616e64626f782d69641a0773616e64626f78220377656228903f3001"; // Synthetic payloads generated with the public declarations at v0.0.116, // before their relocation into openshell.storage.v1. Values are deliberately // non-secret and the ordinary protobuf bytes contain no package names. @@ -590,9 +598,9 @@ mod tests { overlap_hash.as_str(), ), ( - (304, 25), - (92, 19), - (80, 19), + (306, 27), + (93, 21), + (81, 21), PUBLIC_RPC_SCHEMA_SHA256, DURABLE_SCHEMA_SHA256, PUBLIC_DURABLE_OVERLAP_SHA256 @@ -601,6 +609,30 @@ mod tests { ); } + #[test] + fn service_endpoint_without_authorization_mode_decodes_as_strip() { + use openshell_core::proto::{ServiceAuthorizationMode, ServiceEndpoint}; + + let endpoint = ServiceEndpoint::decode( + legacy_bytes(SERVICE_ENDPOINT_WITHOUT_AUTHORIZATION_MODE).as_slice(), + ) + .expect("stored service endpoint without authorization mode must decode"); + + assert_eq!(endpoint.sandbox_id, "sandbox-id"); + assert_eq!(endpoint.sandbox, "sandbox"); + assert_eq!(endpoint.name, "web"); + assert_eq!(endpoint.target_port, 8080); + assert!(endpoint.domain); + assert_eq!( + endpoint.authorization_mode(), + ServiceAuthorizationMode::Unspecified + ); + assert_eq!( + crate::service_routing::effective_authorization_mode(endpoint.authorization_mode), + ServiceAuthorizationMode::Strip + ); + } + #[test] fn pre_readiness_sandbox_spec_decodes_with_initial_attachment_epoch() { let spec = openshell_core::proto::SandboxSpec::decode( @@ -651,6 +683,9 @@ mod tests { assert_eq!(status.current_policy_version, 7); assert_eq!(status.main_process_instance_id, "supervisor-id"); assert_eq!(status.exit_code, None); + assert_eq!(status.restart_count, 0); + assert!(status.next_restart_time.is_none()); + assert!(status.main_process_started_time.is_none()); assert_eq!(status.conditions.len(), 1); assert_eq!(status.conditions[0].r#type, "Ready"); assert_eq!(status.conditions[0].status, "True"); diff --git a/crates/openshell-server/src/tls.rs b/crates/openshell-server/src/tls.rs index 8ce64e3f25..90372b9dad 100644 --- a/crates/openshell-server/src/tls.rs +++ b/crates/openshell-server/src/tls.rs @@ -14,9 +14,7 @@ use arc_swap::ArcSwap; use notify::event::EventKind; use notify::{Event, RecursiveMode, Watcher}; use openshell_core::{Error, Result}; -use openshell_ocsf::{ - ConfigStateChangeBuilder, EventContext, OCSF_TARGET, SeverityId, StateId, StatusId, -}; +use openshell_ocsf::{ConfigStateChangeBuilder, EventContext, SeverityId, StateId, StatusId}; use rustls::ServerConfig; use rustls::crypto::aws_lc_rs::sign; use rustls::pki_types::{CertificateDer, PrivateKeyDer}; @@ -117,11 +115,7 @@ impl TlsAcceptor { .state(StateId::Enabled, "reloaded") .message("TLS certificate config reloaded successfully") .build(); - info!( - target: OCSF_TARGET, - sandbox_id = "", - message = %event.format_shorthand() - ); + openshell_ocsf::ocsf_emit!(event); Ok(()) } @@ -247,11 +241,7 @@ impl TlsAcceptor { "TLS certificate reload failed: {e}" )) .build(); - info!( - target: OCSF_TARGET, - sandbox_id = "", - message = %event.format_shorthand() - ); + openshell_ocsf::ocsf_emit!(event); warn!(error = %e, "TLS certificate reload failed, keeping existing config"); } break; @@ -482,15 +472,7 @@ fn load_key(path: &Path) -> Result> { /// Build an OCSF context for gateway-level (non-sandbox) events. fn tls_ocsf_ctx() -> EventContext { - EventContext { - sandbox_id: String::new(), - sandbox_name: String::new(), - container_image: "openshell/gateway".to_string(), - hostname: "openshell-gateway".to_string(), - product_version: openshell_core::VERSION.to_string(), - proxy_ip: std::net::IpAddr::V4(std::net::Ipv4Addr::LOCALHOST), - proxy_port: 0, - } + crate::gateway_ocsf::context("", "") } #[cfg(test)] diff --git a/crates/openshell-server/src/tracing_bus.rs b/crates/openshell-server/src/tracing_bus.rs index fea7abad5c..ef8cc22874 100644 --- a/crates/openshell-server/src/tracing_bus.rs +++ b/crates/openshell-server/src/tracing_bus.rs @@ -29,6 +29,10 @@ struct Inner { per_id: HashMap, } +/// Result of a resumable-source replay read: the events after the requested +/// cursor, or the gap that made them unreplayable. +pub(crate) type ReplayResult = Result, ResumeGap>; + /// A buffered or broadcast stream event paired with its raw sequence number. /// /// The wire `SandboxStreamEvent.cursor` is an opaque token; ordering decisions @@ -182,6 +186,36 @@ fn tail_after_impl( Ok(res) } +/// Split a tail into what a bounded request returns and the coverage floor +/// that bound leaves behind. +/// +/// `max` may return fewer events than the tail currently retains; the ones +/// excluded are the oldest, at the front. The floor is the *newest* excluded +/// event's seq: this bus returned every event it retains above the floor and +/// none at or below it. `0` means `max` covered everything currently +/// retained, so there's nothing this call withheld. +/// +/// Because the excluded set is a prefix of this bus's retained deque, the +/// floor is the only boundary a caller needs: an event from any followed +/// source with a seq above the floor is safe to hand out as far as this bus +/// is concerned, and one at or below it may sit next to an event this bus +/// withheld. Excluding events below the delivered window is ordinary tail +/// truncation, the same as a single-source `log_tail_lines`. +/// +/// This is what lets a caller answer "is a cursor from this batch safe to +/// resume from," which `tail_after`'s eviction-only gap check can't: a +/// shallow request and a bus eviction both withhold events, but only the +/// eviction leaves a trace in `last_trimmed_seq`. +fn tail_with_floor_impl(tail: &VecDeque, max: usize) -> (Vec, u64) { + let excluded = tail.len().saturating_sub(max); + let floor = excluded + .checked_sub(1) + .and_then(|newest_excluded| tail.get(newest_excluded)) + .map_or(0, |cursored| cursored.seq); + let events = tail.iter().skip(excluded).cloned().collect(); + (events, floor) +} + impl Default for TracingLogBus { fn default() -> Self { Self::new() @@ -274,18 +308,6 @@ impl TracingLogBus { self.seq.space(sandbox_id) } - pub(crate) fn tail_after( - &self, - sandbox_id: &str, - after_seq: u64, - ) -> Result, ResumeGap> { - let inner = self.inner.lock().expect("tracing bus lock poisoned"); - inner.per_id.get(sandbox_id).map_or_else( - || Ok(Vec::new()), - |per| tail_after_impl(&per.tail, per.last_trimmed_seq, after_seq), - ) - } - /// Publish a log line from an external source (e.g., sandbox push). /// /// Injects the line into the same broadcast channel and tail buffer @@ -329,6 +351,149 @@ impl TracingLogBus { } } } + + /// Validate a resume cursor's epoch and read both resumable buses, all + /// under one lock hold. + /// + /// Folding validation into the same hold as the reads is what makes this + /// atomic against a teardown: checking the epoch first and reading after + /// -- even via a single call that itself locks once internally -- still + /// leaves a window between the two lock acquisitions for a `remove` plus + /// a republish to retire the validated space and install a replacement + /// numbered from 1. A read landing after that would apply the old + /// space's seq to the replacement's buffers and find no gap, since + /// `tail_after` only compares numbers. Locking the cursor space before + /// checking the epoch -- the same lock every publish path holds across + /// its own tail insert -- closes that window entirely, and also makes + /// the two bus reads atomic against any publish on either bus, for the + /// same reason. + pub(crate) fn snapshot_after( + &self, + sandbox_id: &str, + expected_epoch: Uuid, + log_after: u64, + platform_after: u64, + want_log: bool, + want_platform: bool, + ) -> ResumeSnapshot { + let spaces = self.seq.lock(); + let Some(space) = spaces.get(sandbox_id) else { + return ResumeSnapshot::SpaceGone; + }; + if space.epoch != expected_epoch { + return ResumeSnapshot::SpaceGone; + } + // Right space, but ahead of anything it has issued: only a + // fabricated cursor gets here. Reject rather than accept a cutoff no + // event can ever exceed. + if log_after.max(platform_after) > space.next.saturating_sub(1) { + return ResumeSnapshot::CursorAhead; + } + + let log = want_log.then(|| { + let inner = self.inner.lock().expect("tracing bus lock poisoned"); + inner.per_id.get(sandbox_id).map_or_else( + || Ok(Vec::new()), + |per| tail_after_impl(&per.tail, per.last_trimmed_seq, log_after), + ) + }); + let platform = want_platform.then(|| { + let inner = self + .platform_event_bus + .inner + .lock() + .expect("platform event bus lock poisoned"); + inner.per_id.get(sandbox_id).map_or_else( + || Ok(Vec::new()), + |per| tail_after_impl(&per.tail, per.last_trimmed_seq, platform_after), + ) + }); + ResumeSnapshot::Read(log, platform) + } + + /// Read both bounded initial tails and the space's high-water mark, all + /// under one lock hold. + /// + /// `None` skips a source. Holding the cursor-space lock -- the one every + /// publish holds across its tail insert and broadcast send -- means no + /// publish can land between the two reads, and `highest_seq` is exactly + /// the boundary between what these tails could see and what can only + /// arrive live. Every event at or below it was already buffered when the + /// tails were read, so the caller can suppress live copies at or below it + /// without dropping anything the tails could not have covered. + /// + /// Reads never mint a space: with none, both tails are empty and + /// `highest_seq` is `0`. + pub(crate) fn snapshot_tail( + &self, + sandbox_id: &str, + log_max: Option, + platform_max: Option, + ) -> TailSnapshot { + let spaces = self.seq.lock(); + let highest_seq = spaces + .get(sandbox_id) + .map_or(0, |space| space.next.saturating_sub(1)); + + let (logs, log_floor) = log_max.map_or_else( + || (Vec::new(), 0), + |max| { + let inner = self.inner.lock().expect("tracing bus lock poisoned"); + inner.per_id.get(sandbox_id).map_or_else( + || (Vec::new(), 0), + |per| tail_with_floor_impl(&per.tail, max), + ) + }, + ); + let (events, platform_floor) = platform_max.map_or_else( + || (Vec::new(), 0), + |max| { + let inner = self + .platform_event_bus + .inner + .lock() + .expect("platform event bus lock poisoned"); + inner.per_id.get(sandbox_id).map_or_else( + || (Vec::new(), 0), + |per| tail_with_floor_impl(&per.tail, max), + ) + }, + ); + drop(spaces); + + TailSnapshot { + logs, + log_floor, + events, + platform_floor, + highest_seq, + } + } +} + +/// Outcome of [`TracingLogBus::snapshot_tail`]. Each floor is the one +/// `tail_with_floor_impl` reports for that source; an unread source has no +/// events and a floor of `0`. +pub(crate) struct TailSnapshot { + pub(crate) logs: Vec, + pub(crate) log_floor: u64, + pub(crate) events: Vec, + pub(crate) platform_floor: u64, + /// Highest seq the space had issued at the lock hold; `0` with no space. + pub(crate) highest_seq: u64, +} + +/// Outcome of [`TracingLogBus::snapshot_after`]. +pub(crate) enum ResumeSnapshot { + /// No cursor space exists for this sandbox, or its epoch no longer + /// matches the resume cursor's -- the space that issued it is gone. + SpaceGone, + /// The right space, but the cursor is ahead of anything it has issued. + /// Only a fabricated cursor reaches this. + CursorAhead, + /// Both requested reads succeeded at the lock hold; each is + /// independently either the replay or the gap that made it unreplayable. + Read(Option, Option), } #[derive(Debug, Clone)] @@ -346,6 +511,20 @@ where let mut visitor = LogVisitor::default(); event.record(&mut visitor); + if meta.target() == OCSF_TARGET + && let Some(ocsf) = openshell_ocsf::clone_current_event() + { + // The structured bridge carries no tracing fields. Route by the + // affected sandbox, not by the gateway that produced the event. + visitor.sandbox_id = ocsf + .base() + .container + .as_ref() + .and_then(|container| container.uid.clone()) + .filter(|id| !id.is_empty()); + visitor.message = Some(ocsf.format_shorthand()); + } + let Some(sandbox_id) = visitor.sandbox_id else { return; }; @@ -411,6 +590,45 @@ fn display_level(target: &str, level: &str) -> String { mod tests { use super::*; + #[test] + fn gateway_ocsf_reaches_sandbox_tail_and_live_stream() { + use tracing_subscriber::prelude::*; + let bus = TracingLogBus::new(); + let mut receiver = bus.subscribe("sb-audit"); + let event = openshell_ocsf::ConfigStateChangeBuilder::new(&crate::gateway_ocsf::context( + "sb-audit", "audit", + )) + .message("policy approved") + .build(); + let expected = event.format_shorthand(); + let subscriber = tracing_subscriber::registry().with(bus.layer()); + tracing::subscriber::with_default(subscriber, || { + openshell_ocsf::ocsf_emit!(event); + openshell_ocsf::ocsf_emit!( + openshell_ocsf::ConfigStateChangeBuilder::new(&crate::gateway_ocsf::context( + "", "" + ),) + .message("gateway-wide event") + .build() + ); + }); + let live = receiver + .try_recv() + .expect("sandbox audit event must reach live stream"); + let Some(openshell_core::proto::sandbox_stream_event::Payload::Log(log)) = + live.event.payload + else { + panic!("expected log payload"); + }; + assert_eq!(log.sandbox_id, "sb-audit"); + assert_eq!(log.source, "gateway"); + assert_eq!(log.level, "OCSF"); + assert_eq!(log.message, expected); + assert_eq!(bus.tail("sb-audit", 10).len(), 1); + assert!(bus.tail("", 10).is_empty()); + assert!(receiver.try_recv().is_err()); + } + fn make_log_event(sandbox_id: &str, message: &str) -> SandboxLogLine { SandboxLogLine { sandbox_id: sandbox_id.to_string(), @@ -493,6 +711,182 @@ mod tests { assert!(tail_after_impl(&tail, 0, 99).expect("ok").is_empty()); } + #[test] + fn tail_with_floor_impl_empty_tail_returns_zero_floor() { + let tail = VecDeque::new(); + let (events, floor) = tail_with_floor_impl(&tail, 10); + assert!(events.is_empty()); + assert_eq!(floor, 0); + } + + #[test] + fn tail_with_floor_impl_max_covers_everything_returns_zero_floor() { + let tail = tail_of(1, 5); + let (events, floor) = tail_with_floor_impl(&tail, 5); + assert_eq!(cursors(&events), vec![1, 2, 3, 4, 5]); + assert_eq!(floor, 0); + + // Asking for more than exists is the same as asking for everything. + let (events, floor) = tail_with_floor_impl(&tail, 99); + assert_eq!(cursors(&events), vec![1, 2, 3, 4, 5]); + assert_eq!(floor, 0); + } + + #[test] + fn tail_with_floor_impl_truncated_reports_newest_excluded_seq() { + let tail = tail_of(1, 5); + // Newest 2 returned (4, 5); 1..=3 excluded, newest of those is 3. + // Every returned event sits above the floor, so a bus never + // withholds its own delivered window. + let (events, floor) = tail_with_floor_impl(&tail, 2); + assert_eq!(cursors(&events), vec![4, 5]); + assert_eq!(floor, 3); + assert!(events.iter().all(|cursored| cursored.seq > floor)); + } + + #[test] + fn tail_with_floor_impl_zero_max_excludes_everything() { + let tail = tail_of(1, 5); + // Nothing returned; every event is excluded, so the floor is the + // newest of them, 5. + let (events, floor) = tail_with_floor_impl(&tail, 0); + assert!(events.is_empty()); + assert_eq!(floor, 5); + } + + /// `snapshot_after` must take the same lock `publish` holds across its own + /// tail insert, so a publish can never land between the log read and the + /// platform read. Proving this directly (rather than racing threads and + /// hoping to catch a bad interleaving) is what makes this test reliable: + /// hold the lock ourselves and show `snapshot_after` cannot proceed until + /// it is released. + #[test] + fn snapshot_after_blocks_while_cursor_space_is_locked() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-lock"; + bus.publish_external(make_log_event(sandbox_id, "line")); + let epoch = bus.cursor_space(sandbox_id).expect("space exists").epoch; + + let guard = bus.seq.lock(); + + let bus2 = bus.clone(); + let (done_tx, done_rx) = std::sync::mpsc::channel(); + let handle = std::thread::spawn(move || { + bus2.snapshot_after(sandbox_id, epoch, 0, 0, true, true); + done_tx.send(()).unwrap(); + }); + + // Give the worker time to reach (and block on) the lock. + std::thread::sleep(std::time::Duration::from_millis(50)); + assert_eq!( + done_rx.try_recv(), + Err(std::sync::mpsc::TryRecvError::Empty), + "snapshot_after returned without waiting for the cursor-space lock" + ); + + drop(guard); + done_rx + .recv_timeout(std::time::Duration::from_secs(1)) + .expect("snapshot_after should complete once the lock is released"); + handle.join().unwrap(); + } + + /// Baseline correctness: interleaved log/platform publishes draw from the + /// same cursor space, and `snapshot_after` returns each source's own + /// events under its own key, not merged or cross-contaminated. + #[test] + fn snapshot_after_matches_independent_tail_after_calls() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-snap"; + bus.publish_external(make_log_event(sandbox_id, "l1")); // seq 1 + bus.platform_event_bus.publish( + sandbox_id, + SandboxStreamEvent { + payload: None, + cursor: String::new(), + }, + ); // seq 2 + bus.publish_external(make_log_event(sandbox_id, "l2")); // seq 3 + let epoch = bus.cursor_space(sandbox_id).expect("space exists").epoch; + + let ResumeSnapshot::Read(log, platform) = + bus.snapshot_after(sandbox_id, epoch, 0, 0, true, true) + else { + panic!("expected a successful read"); + }; + assert_eq!( + cursors(&log.expect("followed").expect("no gap")), + vec![1, 3] + ); + assert_eq!( + cursors(&platform.expect("followed").expect("no gap")), + vec![2] + ); + } + + /// A source that isn't followed isn't read at all. + #[test] + fn snapshot_after_skips_unfollowed_sources() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-skip"; + bus.publish_external(make_log_event(sandbox_id, "l1")); + let epoch = bus.cursor_space(sandbox_id).expect("space exists").epoch; + + let ResumeSnapshot::Read(log, platform) = + bus.snapshot_after(sandbox_id, epoch, 0, 0, true, false) + else { + panic!("expected a successful read"); + }; + assert_eq!(log.expect("followed").expect("no gap").len(), 1); + assert!(platform.is_none()); + } + + /// The epoch check and both bus reads happen under one lock hold, so a + /// teardown plus republish between them can't apply a stale epoch's + /// cursor to the replacement space's buffers. + #[test] + fn snapshot_after_rejects_a_cursor_from_a_retired_space() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-retired"; + bus.publish_external(make_log_event(sandbox_id, "l1")); + let original_epoch = bus.cursor_space(sandbox_id).expect("space exists").epoch; + + bus.remove(sandbox_id); + bus.publish_external(make_log_event(sandbox_id, "l1-again")); // seq 1 again, new epoch + + assert!(matches!( + bus.snapshot_after(sandbox_id, original_epoch, 0, 0, true, false), + ResumeSnapshot::SpaceGone + )); + } + + /// No space at all -- an unpublished or never-existent sandbox -- is the + /// same outcome as a retired one: there's nothing for the cursor to + /// address. + #[test] + fn snapshot_after_rejects_when_no_space_exists() { + let bus = TracingLogBus::new(); + assert!(matches!( + bus.snapshot_after("nope", Uuid::nil(), 0, 0, true, false), + ResumeSnapshot::SpaceGone + )); + } + + /// A cursor ahead of anything the space has issued -- only reachable with + /// a fabricated token -- is rejected rather than treated as caught up. + #[test] + fn snapshot_after_rejects_a_cursor_ahead_of_the_space() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-ahead"; + bus.publish_external(make_log_event(sandbox_id, "l1")); + let epoch = bus.cursor_space(sandbox_id).expect("space exists").epoch; + + assert!(matches!( + bus.snapshot_after(sandbox_id, epoch, 99, 0, true, false), + ResumeSnapshot::CursorAhead + )); + } + #[test] fn tail_after_impl_boundary_at_last_trimmed_is_serviceable() { // Bus trimmed up to seq 2, retains 3..=5. Client saw exactly 2, so @@ -528,56 +922,166 @@ mod tests { } #[test] - fn tracing_log_bus_tail_after_serviceable_and_missing() { + fn tracing_log_bus_resume_replay_serviceable_and_missing() { let bus = TracingLogBus::new(); let sandbox_id = "sb-ta"; for _ in 0..3 { bus.publish_external(make_log_event(sandbox_id, "line")); } + let epoch = bus.cursor_space(sandbox_id).expect("space exists").epoch; + // Cursors start at 1, so three publishes are seqs 1,2,3. + let ResumeSnapshot::Read(log, _) = bus.snapshot_after(sandbox_id, epoch, 0, 0, true, false) + else { + panic!("expected a successful read"); + }; assert_eq!( - cursors(&bus.tail_after(sandbox_id, 0).unwrap()), + cursors(&log.expect("followed").expect("no gap")), vec![1, 2, 3] ); - assert_eq!(cursors(&bus.tail_after(sandbox_id, 2).unwrap()), vec![3]); - // Unknown sandbox: no entry, nothing buffered, no gap. - assert!(bus.tail_after("nope", 5).unwrap().is_empty()); + let ResumeSnapshot::Read(log, _) = bus.snapshot_after(sandbox_id, epoch, 2, 0, true, false) + else { + panic!("expected a successful read"); + }; + assert_eq!(cursors(&log.expect("followed").expect("no gap")), vec![3]); } #[test] - fn platform_event_bus_tail_after_serviceable() { + fn platform_event_bus_resume_replay_serviceable() { let bus = TracingLogBus::new(); let platform = &bus.platform_event_bus; let sandbox_id = "sb-pe"; for _ in 0..3 { platform.publish(sandbox_id, stream_event(0)); } + let epoch = bus.cursor_space(sandbox_id).expect("space exists").epoch; + // Shared allocator, but only the platform bus published here, so its // seqs are 1,2,3. + let ResumeSnapshot::Read(_, platform_replay) = + bus.snapshot_after(sandbox_id, epoch, 0, 0, false, true) + else { + panic!("expected a successful read"); + }; assert_eq!( - cursors(&platform.tail_after(sandbox_id, 0).unwrap()), + cursors(&platform_replay.expect("followed").expect("no gap")), vec![1, 2, 3] ); + let ResumeSnapshot::Read(_, platform_replay) = + bus.snapshot_after(sandbox_id, epoch, 0, 1, false, true) + else { + panic!("expected a successful read"); + }; assert_eq!( - cursors(&platform.tail_after(sandbox_id, 1).unwrap()), + cursors(&platform_replay.expect("followed").expect("no gap")), vec![2, 3] ); } #[test] - fn shared_allocator_interleaves_cursors_across_buses() { + fn snapshot_tail_reports_truncation() { let bus = TracingLogBus::new(); - let sandbox_id = "sb-mix"; - // Interleave log and platform publishes; the shared allocator gives - // each a unique, increasing cursor in one merged space. - bus.publish_external(make_log_event(sandbox_id, "a")); // seq 1 - bus.platform_event_bus.publish(sandbox_id, stream_event(0)); // seq 2 - bus.publish_external(make_log_event(sandbox_id, "b")); // seq 3 + let sandbox_id = "sb-floor"; + for _ in 0..5 { + bus.publish_external(make_log_event(sandbox_id, "line")); + } + let snap = bus.snapshot_tail(sandbox_id, Some(2), None); + assert_eq!(cursors(&snap.logs), vec![4, 5]); + assert_eq!(snap.log_floor, 3); + + // Wide enough to cover everything: no floor. + let snap = bus.snapshot_tail(sandbox_id, Some(10), None); + assert_eq!(cursors(&snap.logs), vec![1, 2, 3, 4, 5]); + assert_eq!(snap.log_floor, 0); + } - let logs = cursors(&bus.tail_after(sandbox_id, 0).unwrap()); - let events = cursors(&bus.platform_event_bus.tail_after(sandbox_id, 0).unwrap()); - assert_eq!(logs, vec![1, 3]); - assert_eq!(events, vec![2]); + /// `highest_seq` is the space's high-water mark at the lock hold, not the + /// newest event either window returned: events a depth excluded still + /// count, since they were buffered when the tails were read and can only + /// reach a live receiver as a duplicate of pre-snapshot history. + #[test] + fn snapshot_tail_reports_the_high_water_mark_past_excluded_events() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-hwm"; + bus.publish_external(make_log_event(sandbox_id, "l1")); // seq 1 + bus.publish_external(make_log_event(sandbox_id, "l2")); // seq 2 + bus.platform_event_bus.publish(sandbox_id, stream_event(0)); // seq 3 + + let snap = bus.snapshot_tail(sandbox_id, Some(10), Some(0)); + assert_eq!(cursors(&snap.logs), vec![1, 2]); + assert_eq!(snap.log_floor, 0); + assert!(snap.events.is_empty()); + assert_eq!(snap.platform_floor, 3); + assert_eq!(snap.highest_seq, 3); + } + + /// An unread source is skipped outright; a sandbox with nothing published + /// yields empty tails and a zero mark without minting a cursor space. + #[test] + fn snapshot_tail_skips_unread_sources_and_never_mints_a_space() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-tail-skip"; + bus.publish_external(make_log_event(sandbox_id, "l1")); + bus.platform_event_bus.publish(sandbox_id, stream_event(0)); + + let snap = bus.snapshot_tail(sandbox_id, Some(10), None); + assert_eq!(cursors(&snap.logs), vec![1]); + assert!(snap.events.is_empty()); + assert_eq!(snap.platform_floor, 0); + assert_eq!(snap.highest_seq, 2); + + let snap = bus.snapshot_tail("nope", Some(10), Some(10)); + assert!(snap.logs.is_empty() && snap.events.is_empty()); + assert_eq!(snap.highest_seq, 0); + assert_eq!(bus.cursor_space("nope"), None); + } + + /// Same guarantee as `snapshot_after_blocks_while_cursor_space_is_locked` + /// for the initial read: `snapshot_tail` takes the lock every publish holds + /// across its tail insert and broadcast send, so no publish can land + /// between the log read and the platform read. + #[test] + fn snapshot_tail_blocks_while_cursor_space_is_locked() { + let bus = TracingLogBus::new(); + let sandbox_id = "sb-tail-lock"; + bus.publish_external(make_log_event(sandbox_id, "line")); + + let guard = bus.seq.lock(); + + let bus2 = bus.clone(); + let (done_tx, done_rx) = std::sync::mpsc::channel(); + let handle = std::thread::spawn(move || { + bus2.snapshot_tail(sandbox_id, Some(10), Some(10)); + done_tx.send(()).unwrap(); + }); + + // Give the worker time to reach (and block on) the lock. + std::thread::sleep(std::time::Duration::from_millis(50)); + assert_eq!( + done_rx.try_recv(), + Err(std::sync::mpsc::TryRecvError::Empty), + "snapshot_tail returned without waiting for the cursor-space lock" + ); + + drop(guard); + done_rx + .recv_timeout(std::time::Duration::from_secs(1)) + .expect("snapshot_tail should complete once the lock is released"); + handle.join().unwrap(); + } + + #[test] + fn platform_event_bus_tail_with_floor_reports_truncation() { + let bus = TracingLogBus::new(); + let platform = &bus.platform_event_bus; + let sandbox_id = "sb-pe-floor"; + for _ in 0..5 { + platform.publish(sandbox_id, stream_event(0)); + } + let (events, floor) = platform.tail_with_floor(sandbox_id, 0); + assert!(events.is_empty()); + // Nothing returned: floor is the newest of everything excluded, 5. + assert_eq!(floor, 5); } #[test] @@ -661,7 +1165,9 @@ mod tests { assert_eq!(after_platform.highest_seq, 2); let log_cursor = &bus.tail(sandbox_id, 10)[0].event.cursor; - let platform_cursor = &bus.platform_event_bus.tail(sandbox_id, 10)[0].event.cursor; + let platform_cursor = &bus.platform_event_bus.tail_with_floor(sandbox_id, 10).0[0] + .event + .cursor; assert_eq!( WatchCursor::parse(log_cursor).expect("valid").epoch, WatchCursor::parse(platform_cursor).expect("valid").epoch, @@ -848,8 +1354,9 @@ mod tests { } // Tail should return all events in order - let events = bus.tail(sandbox_id, 10); + let (events, floor) = bus.tail_with_floor(sandbox_id, 10); assert_eq!(events.len(), 5); + assert_eq!(floor, 0, "max covered everything retained"); // Verify order (oldest first) for (i, cursored) in events.iter().enumerate() { @@ -861,8 +1368,9 @@ mod tests { } // Tail with smaller max should return most recent events - let events = bus.tail(sandbox_id, 2); + let (events, floor) = bus.tail_with_floor(sandbox_id, 2); assert_eq!(events.len(), 2); + assert_eq!(floor, 3, "newest excluded event is Event2 (seq 3)"); if let Some(sandbox_stream_event::Payload::Event(ref e)) = events[0].event.payload { assert_eq!(e.reason, "Event3"); } @@ -874,8 +1382,9 @@ mod tests { #[test] fn platform_event_bus_tail_empty_sandbox() { let bus = PlatformEventBus::new(SeqAllocator::default()); - let events = bus.tail("nonexistent", 10); + let (events, floor) = bus.tail_with_floor("nonexistent", 10); assert!(events.is_empty()); + assert_eq!(floor, 0); } #[test] @@ -888,10 +1397,10 @@ mod tests { cursor: String::new(), }; bus.publish(sandbox_id, evt); - assert_eq!(bus.tail(sandbox_id, 10).len(), 1); + assert_eq!(bus.tail_with_floor(sandbox_id, 10).0.len(), 1); bus.remove(sandbox_id); - assert!(bus.tail(sandbox_id, 10).is_empty()); + assert!(bus.tail_with_floor(sandbox_id, 10).0.is_empty()); } } @@ -958,28 +1467,19 @@ impl PlatformEventBus { } } - /// Return buffered platform events for replay to late subscribers. - pub(crate) fn tail(&self, sandbox_id: &str, max: usize) -> Vec { - let inner = self.inner.lock().expect("platform event bus lock poisoned"); - inner - .per_id - .get(sandbox_id) - .map(|d| d.tail.iter().rev().take(max).cloned().collect::>()) - .unwrap_or_default() - .into_iter() - .rev() - .collect() - } - - pub(crate) fn tail_after( + /// Return buffered platform events and the coverage floor `max` leaves + /// behind. See `tail_with_floor_impl`. Production reads go through + /// `TracingLogBus::snapshot_tail`, which holds the cursor-space lock. + #[cfg(test)] + pub(crate) fn tail_with_floor( &self, sandbox_id: &str, - after_seq: u64, - ) -> Result, ResumeGap> { + max: usize, + ) -> (Vec, u64) { let inner = self.inner.lock().expect("platform event bus lock poisoned"); inner.per_id.get(sandbox_id).map_or_else( - || Ok(Vec::new()), - |per| tail_after_impl(&per.tail, per.last_trimmed_seq, after_seq), + || (Vec::new(), 0), + |per| tail_with_floor_impl(&per.tail, max), ) } diff --git a/crates/openshell-server/src/tracing_setup.rs b/crates/openshell-server/src/tracing_setup.rs index d6b1ab0ecf..3a74edd72f 100644 --- a/crates/openshell-server/src/tracing_setup.rs +++ b/crates/openshell-server/src/tracing_setup.rs @@ -10,6 +10,7 @@ use openshell_ocsf::OcsfJsonlLayer; use opentelemetry_sdk::trace::SdkTracerProvider; use tracing_subscriber::EnvFilter; +use tracing_subscriber::filter::{FilterExt, filter_fn}; use tracing_subscriber::prelude::*; use crate::config_file::OtlpConfig; @@ -36,9 +37,41 @@ impl TracingHandle { } } +fn filter_from(directives: &str) -> EnvFilter { + EnvFilter::try_new(directives).unwrap_or_else(|_| EnvFilter::new("info")) +} + +struct GatewayEventFormat; + +impl tracing_subscriber::fmt::FormatEvent for GatewayEventFormat +where + S: tracing::Subscriber + for<'a> tracing_subscriber::registry::LookupSpan<'a>, + N: for<'a> tracing_subscriber::fmt::FormatFields<'a> + 'static, +{ + fn format_event( + &self, + context: &tracing_subscriber::fmt::FmtContext<'_, S, N>, + mut writer: tracing_subscriber::fmt::format::Writer<'_>, + event: &tracing::Event<'_>, + ) -> std::fmt::Result { + if event.metadata().target() == openshell_ocsf::OCSF_TARGET + && let Some(ocsf) = openshell_ocsf::clone_current_event() + { + return writeln!( + writer, + "{} OCSF {}", + chrono::Utc::now().format("%Y-%m-%dT%H:%M:%S%.3fZ"), + ocsf.format_shorthand() + ); + } + tracing_subscriber::fmt::format().format_event(context, writer, event) + } +} + pub fn install( - env_filter: EnvFilter, + filter_directives: &str, tracing_log_bus: &TracingLogBus, + ocsf_log: Option<&crate::ocsf_log::OcsfLog>, otlp_config: Option<&OtlpConfig>, driver: Option, gateway: GatewayResourceAttributes<'_>, @@ -66,14 +99,31 @@ pub fn install( // level. An explicit JSONL opt-in must keep every OCSF event even when the // console and routed diagnostic logs are restricted to `warn` or `error`. tracing_subscriber::registry() - .with(tracing_subscriber::fmt::layer().with_filter(env_filter.clone())) - .with(tracing_log_bus.layer().with_filter(env_filter.clone())) + .with( + ocsf_log + .map(crate::ocsf_log::OcsfLog::layer) + .with_filter(filter_fn(|metadata| { + metadata.target() == openshell_ocsf::OCSF_TARGET + })), + ) + .with( + tracing_subscriber::fmt::layer() + .event_format(GatewayEventFormat) + .with_filter(filter_from(filter_directives)), + ) + .with( + tracing_log_bus + .layer() + .with_filter(filter_from(filter_directives).or(filter_fn(|metadata| { + metadata.target() == openshell_ocsf::OCSF_TARGET + }))), + ) .with(jsonl_layer) .with( tracer_provider .as_ref() .map(|provider| crate::otel_tracing::layer(provider, driver)) - .with_filter(env_filter.clone()), + .with_filter(filter_from(filter_directives)), ) .with( driver_tracer_provider @@ -83,7 +133,7 @@ pub fn install( .expect("a driver provider requires a selected driver") .in_process_layer(provider) }) - .with_filter(env_filter), + .with_filter(filter_from(filter_directives)), ) .init(); @@ -248,6 +298,7 @@ mod tests { product_version: openshell_core::VERSION.into(), proxy_ip: "127.0.0.1".parse().unwrap(), proxy_port: 0, + origin: openshell_ocsf::EventOrigin::Supervisor, }; tracing::subscriber::with_default(subscriber, || { @@ -293,3 +344,41 @@ mod tests { } } } + +#[cfg(test)] +mod gateway_format_tests { + use std::io::{Read, Seek}; + + use super::*; + + #[test] + fn gateway_ocsf_console_preserves_details_without_jsonl() { + let file = tempfile::tempfile().unwrap(); + let reader = file.try_clone().unwrap(); + let subscriber = tracing_subscriber::registry().with( + tracing_subscriber::fmt::layer() + .event_format(GatewayEventFormat) + .with_ansi(false) + .with_writer(std::sync::Arc::new(file)) + .with_filter(filter_from("info")), + ); + let event = + openshell_ocsf::ConfigStateChangeBuilder::new(&crate::gateway_ocsf::context("", "")) + .message("TLS certificate config reloaded successfully") + .build(); + let expected = event.format_shorthand(); + tracing::subscriber::with_default(subscriber, || { + openshell_ocsf::ocsf_emit!(event); + tracing::info!(answer = 42, "ordinary diagnostic"); + }); + let mut reader = reader; + reader.rewind().unwrap(); + let mut output = String::new(); + reader.read_to_string(&mut output).unwrap(); + assert!(output.contains(&expected), "missing OCSF details: {output}"); + assert!(!output.contains("ocsf_event")); + assert!(output.contains("ordinary diagnostic")); + assert!(output.contains("answer=42")); + assert_eq!(output.lines().count(), 2); + } +} diff --git a/crates/openshell-supervisor-network/Cargo.toml b/crates/openshell-supervisor-network/Cargo.toml index d4aa9d1ee9..300db03acd 100644 --- a/crates/openshell-supervisor-network/Cargo.toml +++ b/crates/openshell-supervisor-network/Cargo.toml @@ -16,6 +16,7 @@ openshell-core = { path = "../openshell-core", features = ["oauth"] } openshell-isolation-interface = { path = "../openshell-isolation-interface" } openshell-ocsf = { path = "../openshell-ocsf" } openshell-policy = { path = "../openshell-policy" } +openshell-policy-schema = { path = "../openshell-policy-schema" } openshell-supervisor-middleware = { path = "../openshell-supervisor-middleware" } async-trait = "0.1" diff --git a/crates/openshell-supervisor-network/data/sandbox-policy.rego b/crates/openshell-supervisor-network/data/sandbox-policy.rego index 4053b23c9a..a477e1253f 100644 --- a/crates/openshell-supervisor-network/data/sandbox-policy.rego +++ b/crates/openshell-supervisor-network/data/sandbox-policy.rego @@ -424,11 +424,66 @@ request_deny_reason := reason if { reason := "JSON-RPC response frames are not permitted from client to server" } +# Explain only parsed MCP calls on an endpoint that matches this request path. +# The relay evaluates batch members separately. Response frames and protocol +# errors keep their own diagnostics, and a sibling endpoint must not select +# the explanation merely because it shares the connection's host and port. +mcp_policy_request if { + input.request.method == "POST" + not is_object(object.get(input.request, "graphql", null)) + not jsonrpc_response_frame_present(input.request) + jsonrpc := object.get(input.request, "jsonrpc", null) + is_object(jsonrpc) + jsonrpc_no_parse_error(jsonrpc) + method := object.get(jsonrpc, "method", "") + is_string(method) + method != "" + object.get(jsonrpc, "mcp_method_classification", "") in {"available", "extension"} + endpoint := _matching_endpoint_configs[_] + endpoint.protocol == "mcp" + endpoint_path_matches_request(endpoint, input.request) +} + +# These reasons use fixed text because method names and tool parameters can +# contain caller data. Deny rules take precedence over missing allow rules. +request_deny_reason := reason if { + mcp_policy_request + deny_request + reason := "MCP request blocked by a deny rule; ask the policy owner to review deny_rules and tool selectors" +} + +request_deny_reason := reason if { + mcp_policy_request + not deny_request + not allow_request + input.request.jsonrpc.mcp_method_classification == "extension" + reason := "MCP extension method has no matching exact allow rule; ask the policy owner to review rules with an exact method name and any parameter restrictions; allow_all_known_mcp_methods does not allow extensions" +} + +request_deny_reason := reason if { + mcp_policy_request + not deny_request + not allow_request + input.request.jsonrpc.mcp_method_classification == "available" + input.request.jsonrpc.method == "tools/call" + reason := "MCP tool call has no matching allow rule; ask the policy owner to review rules and tool selectors" +} + +request_deny_reason := reason if { + mcp_policy_request + not deny_request + not allow_request + input.request.jsonrpc.mcp_method_classification == "available" + input.request.jsonrpc.method != "tools/call" + reason := "MCP core method is not permitted by policy; ask the policy owner to review rules for this method in the selected MCP revision" +} + request_deny_reason := reason if { input.request deny_request not graphql_request_has_operations(input.request) not jsonrpc_response_frame_present(input.request) + not mcp_policy_request reason := sprintf("%s %s blocked by deny rule", [input.request.method, input.request.path]) } @@ -438,6 +493,7 @@ request_deny_reason := reason if { not allow_request not graphql_request_has_operations(input.request) not jsonrpc_response_frame_present(input.request) + not mcp_policy_request reason := sprintf("%s %s not permitted by policy", [input.request.method, input.request.path]) } @@ -472,6 +528,10 @@ request_allowed_for_endpoint(request, endpoint) if { rule.allow.method not jsonrpc_response_frame_present(request) jsonrpc_rule_matches(request, endpoint, rule.allow) + jsonrpc := object.get(request, "jsonrpc", null) + method := object.get(jsonrpc, "method", "") + rule_method := object.get(rule.allow, "method", "") + jsonrpc_allow_rule_classification_allowed(jsonrpc, endpoint, method, rule_method) } # MCP can allow the method layer by endpoint option while still using @@ -487,6 +547,7 @@ request_allowed_for_endpoint(request, endpoint) if { method := object.get(jsonrpc, "method", "") is_string(method) method != "" + object.get(jsonrpc, "mcp_method_classification", "") == "available" not mcp_tool_call_narrowed_by_policy(endpoint, method) } @@ -774,6 +835,23 @@ jsonrpc_rule_matches(request, endpoint, rule) if { jsonrpc_rule_params_match_for_protocol(jsonrpc, endpoint, rule) } +jsonrpc_allow_rule_classification_allowed(_, endpoint, _, _) if { + endpoint.protocol == "json-rpc" +} + +jsonrpc_allow_rule_classification_allowed(jsonrpc, endpoint, _, _) if { + endpoint.protocol == "mcp" + object.get(jsonrpc, "mcp_method_classification", "") == "available" +} + +# Extension methods remain addressable, but only by an exact policy literal. +# A wildcard must not silently authorize methods outside the selected core profile. +jsonrpc_allow_rule_classification_allowed(jsonrpc, endpoint, method, rule_method) if { + endpoint.protocol == "mcp" + object.get(jsonrpc, "mcp_method_classification", "") == "extension" + rule_method == method +} + jsonrpc_rule_method_matches(endpoint, _, rule_method) if { endpoint.protocol == "json-rpc" rule_method == "*" diff --git a/crates/openshell-supervisor-network/src/l7/graphql.rs b/crates/openshell-supervisor-network/src/l7/graphql.rs index 994d101314..db322766b4 100644 --- a/crates/openshell-supervisor-network/src/l7/graphql.rs +++ b/crates/openshell-supervisor-network/src/l7/graphql.rs @@ -3,14 +3,14 @@ //! GraphQL-over-HTTP L7 inspection. -use crate::l7::provider::{BodyLength, L7Provider, L7Request}; +use crate::l7::provider::{L7Provider, L7Request}; use apollo_parser::Parser; use apollo_parser::cst; -use miette::{IntoDiagnostic, Result, miette}; +use miette::{Result, miette}; use serde::Serialize; use serde_json::Value; use std::collections::{HashMap, HashSet}; -use tokio::io::{AsyncRead, AsyncReadExt, AsyncWrite}; +use tokio::io::{AsyncRead, AsyncWrite}; pub const DEFAULT_MAX_BODY_BYTES: usize = 64 * 1024; @@ -85,7 +85,7 @@ pub(crate) async fn inspect_graphql_request( ) -> Result { let header_str = header_str(request)?; reject_unsupported_headers(header_str)?; - let body = read_body_for_inspection(client, request, max_body_bytes).await?; + let body = crate::l7::http::read_body_for_inspection(client, request, max_body_bytes).await?; Ok(classify_request(request, &body)) } @@ -379,195 +379,6 @@ fn unique_persisted_query_id( Ok(selected.map(|(_, value)| value)) } -async fn read_body_for_inspection( - client: &mut C, - request: &mut L7Request, - max_body_bytes: usize, -) -> Result> { - let header_end = request - .raw_header - .windows(4) - .position(|w| w == b"\r\n\r\n") - .map_or(request.raw_header.len(), |p| p + 4); - let overflow = request.raw_header[header_end..].to_vec(); - - match request.body_length { - BodyLength::None => Ok(Vec::new()), - BodyLength::ContentLength(len) => { - let len = usize::try_from(len) - .map_err(|_| miette!("GraphQL request body length exceeds platform limit"))?; - if len > max_body_bytes { - return Err(miette!( - "GraphQL request body exceeds {max_body_bytes} byte inspection limit" - )); - } - if overflow.len() > len { - return Err(miette!( - "GraphQL request contains more body bytes than Content-Length" - )); - } - let remaining = len - overflow.len(); - let mut body = overflow; - if remaining > 0 { - let start = body.len(); - body.resize(len, 0); - client - .read_exact(&mut body[start..]) - .await - .into_diagnostic()?; - } - request.raw_header.truncate(header_end); - request.raw_header.extend_from_slice(&body); - Ok(body) - } - BodyLength::Chunked => { - let body = read_chunked_body_for_inspection( - client, - request, - header_end, - overflow, - max_body_bytes, - ) - .await?; - normalize_chunked_request_to_content_length(request, header_end, &body)?; - Ok(body) - } - } -} - -fn normalize_chunked_request_to_content_length( - request: &mut L7Request, - header_end: usize, - body: &[u8], -) -> Result<()> { - let header_str = std::str::from_utf8(&request.raw_header[..header_end]) - .map_err(|_| miette!("GraphQL HTTP headers contain invalid UTF-8"))?; - let header_str = header_str - .strip_suffix("\r\n\r\n") - .ok_or_else(|| miette!("GraphQL HTTP headers missing terminator"))?; - - let mut normalized = Vec::with_capacity(header_str.len() + body.len() + 32); - for (idx, line) in header_str.split("\r\n").enumerate() { - if idx > 0 { - let name = line - .split_once(':') - .map(|(name, _)| name.trim().to_ascii_lowercase()); - if matches!( - name.as_deref(), - Some("transfer-encoding" | "content-length" | "trailer") - ) { - continue; - } - } - normalized.extend_from_slice(line.as_bytes()); - normalized.extend_from_slice(b"\r\n"); - } - normalized.extend_from_slice(format!("Content-Length: {}\r\n\r\n", body.len()).as_bytes()); - normalized.extend_from_slice(body); - - request.raw_header = normalized; - request.body_length = BodyLength::ContentLength(body.len() as u64); - Ok(()) -} - -async fn read_chunked_body_for_inspection( - client: &mut C, - request: &mut L7Request, - header_end: usize, - overflow: Vec, - max_body_bytes: usize, -) -> Result> { - let mut raw = overflow; - let mut decoded = Vec::new(); - let mut pos = 0usize; - - loop { - let size_line_end = loop { - if let Some(end) = find_crlf(&raw, pos) { - break end; - } - read_more(client, &mut raw, max_body_bytes).await?; - }; - let size_line = std::str::from_utf8(&raw[pos..size_line_end]) - .into_diagnostic() - .map_err(|_| miette!("Invalid UTF-8 in GraphQL chunk-size line"))?; - let size_token = size_line - .split(';') - .next() - .map(str::trim) - .unwrap_or_default(); - let chunk_size = usize::from_str_radix(size_token, 16) - .into_diagnostic() - .map_err(|_| miette!("Invalid GraphQL chunk size token: {size_token:?}"))?; - pos = size_line_end + 2; - - if decoded.len().saturating_add(chunk_size) > max_body_bytes { - return Err(miette!( - "GraphQL request body exceeds {max_body_bytes} byte inspection limit" - )); - } - - if chunk_size == 0 { - loop { - let trailer_end = loop { - if let Some(end) = find_crlf(&raw, pos) { - break end; - } - read_more(client, &mut raw, max_body_bytes).await?; - }; - let trailer_line = &raw[pos..trailer_end]; - pos = trailer_end + 2; - if trailer_line.is_empty() { - request.raw_header.truncate(header_end); - request.raw_header.extend_from_slice(&raw[..pos]); - return Ok(decoded); - } - } - } - - let chunk_end = pos - .checked_add(chunk_size) - .ok_or_else(|| miette!("GraphQL chunk size overflow"))?; - let chunk_with_crlf_end = chunk_end - .checked_add(2) - .ok_or_else(|| miette!("GraphQL chunk size overflow"))?; - while raw.len() < chunk_with_crlf_end { - read_more(client, &mut raw, max_body_bytes).await?; - } - decoded.extend_from_slice(&raw[pos..chunk_end]); - if raw.get(chunk_end..chunk_with_crlf_end) != Some(&b"\r\n"[..]) { - return Err(miette!("GraphQL chunk payload missing terminating CRLF")); - } - pos = chunk_with_crlf_end; - } -} - -async fn read_more( - client: &mut C, - raw: &mut Vec, - max_body_bytes: usize, -) -> Result<()> { - if raw.len() > max_body_bytes.saturating_mul(2).max(max_body_bytes) { - return Err(miette!( - "GraphQL chunked request body exceeds inspection framing limit" - )); - } - let mut buf = [0u8; 8192]; - let n = client.read(&mut buf).await.into_diagnostic()?; - if n == 0 { - return Err(miette!("GraphQL chunked body ended before terminator")); - } - raw.extend_from_slice(&buf[..n]); - Ok(()) -} - -fn find_crlf(buf: &[u8], start: usize) -> Option { - buf.get(start..)? - .windows(2) - .position(|w| w == b"\r\n") - .map(|p| start + p) -} - fn header_str(request: &L7Request) -> Result<&str> { let header_end = request .raw_header @@ -602,6 +413,8 @@ fn reject_unsupported_headers(headers: &str) -> Result<()> { #[cfg(test)] mod tests { use super::*; + use crate::l7::provider::BodyLength; + use tokio::io::AsyncReadExt; fn request(method: &str, target: &str) -> L7Request { L7Request { @@ -698,6 +511,68 @@ mod tests { ); } + #[tokio::test] + async fn chunked_graphql_preserves_pipelined_request() { + use tokio::io::{AsyncWriteExt, BufReader}; + + let body = r#"{"query":"query Viewer { viewer { login } }"}"#; + let next = "GET /next HTTP/1.1\r\nHost: example.com\r\n\r\n"; + for capacity in [1, 8192] { + for trailers in ["", "X-Sig: ignored\r\n"] { + let wire = format!( + "POST /graphql HTTP/1.1\r\nHost: example.com\r\nContent-Type: application/json\r\nTransfer-Encoding: chunked\r\n\r\n{:x};ext=yes\r\n{body}\r\n0\r\n{trailers}\r\n{next}", + body.len() + ); + let (mut sender, receiver) = tokio::io::duplex(8192); + sender.write_all(wire.as_bytes()).await.unwrap(); + sender.shutdown().await.unwrap(); + // A one-byte buffer splits every CRLF; a large buffer makes + // the complete next request available during body inspection. + let mut reader = BufReader::with_capacity(capacity, receiver); + let parsed = parse_graphql_http_request( + &mut reader, + DEFAULT_MAX_BODY_BYTES, + crate::l7::path::CanonicalizeOptions::default(), + ) + .await + .unwrap() + .unwrap(); + assert_eq!(parsed.info.error, None); + assert_eq!(parsed.info.operations[0].fields, ["viewer"]); + let mut remaining = String::new(); + reader.read_to_string(&mut remaining).await.unwrap(); + assert_eq!( + remaining, next, + "GraphQL inspection consumed the next request" + ); + } + } + } + + #[tokio::test] + async fn graphql_inspection_preserves_configured_body_limit() { + let body = br#"{"query":"query { viewer }"}"#; + for chunked in [false, true] { + let mut req = request("POST", "/graphql"); + let wire = if chunked { + req.body_length = BodyLength::Chunked; + format!( + "{:x}\r\n{}\r\n0\r\n\r\n", + body.len(), + std::str::from_utf8(body).unwrap() + ) + .into_bytes() + } else { + req.body_length = BodyLength::ContentLength(body.len() as u64); + body.to_vec() + }; + let error = inspect_graphql_request(&mut wire.as_slice(), &mut req, body.len() - 1) + .await + .expect_err("the configured body limit must apply to either HTTP framing"); + assert!(error.to_string().contains("inspection limit")); + } + } + #[tokio::test] async fn chunked_graphql_post_is_normalized_after_inspection() { let body = br#"{"query":"query Viewer { viewer { login } }"}"#; diff --git a/crates/openshell-supervisor-network/src/l7/http.rs b/crates/openshell-supervisor-network/src/l7/http.rs index 66269f6ba2..9be6e7d891 100644 --- a/crates/openshell-supervisor-network/src/l7/http.rs +++ b/crates/openshell-supervisor-network/src/l7/http.rs @@ -112,12 +112,7 @@ async fn read_chunked_body_for_inspection( let mut pos = 0usize; loop { - let size_line_end = loop { - if let Some(end) = find_crlf(&raw, pos) { - break end; - } - read_more(client, &mut raw, max_body_bytes).await?; - }; + let size_line_end = read_chunked_line(client, &mut raw, pos, max_body_bytes).await?; let size_line = std::str::from_utf8(&raw[pos..size_line_end]) .into_diagnostic() .map_err(|_| miette!("Invalid UTF-8 in HTTP chunk-size line"))?; @@ -139,12 +134,7 @@ async fn read_chunked_body_for_inspection( if chunk_size == 0 { loop { - let trailer_end = loop { - if let Some(end) = find_crlf(&raw, pos) { - break end; - } - read_more(client, &mut raw, max_body_bytes).await?; - }; + let trailer_end = read_chunked_line(client, &mut raw, pos, max_body_bytes).await?; let trailer_line = &raw[pos..trailer_end]; pos = trailer_end + 2; if trailer_line.is_empty() { @@ -162,7 +152,8 @@ async fn read_chunked_body_for_inspection( .checked_add(2) .ok_or_else(|| miette!("HTTP chunk size overflow"))?; while raw.len() < chunk_with_crlf_end { - read_more(client, &mut raw, max_body_bytes).await?; + let remaining = chunk_with_crlf_end - raw.len(); + read_more(client, &mut raw, max_body_bytes, remaining).await?; } decoded.extend_from_slice(&raw[pos..chunk_end]); if raw.get(chunk_end..chunk_with_crlf_end) != Some(&b"\r\n"[..]) { @@ -172,10 +163,30 @@ async fn read_chunked_body_for_inspection( } } +async fn read_chunked_line( + client: &mut C, + raw: &mut Vec, + start: usize, + max_body_bytes: usize, +) -> Result { + let mut scan_start = start; + loop { + if let Some(end) = find_crlf(raw, scan_start) { + return Ok(end); + } + // Retain a possible trailing CR without rescanning the entire line. + scan_start = raw.len().saturating_sub(1).max(start); + // Framing has no known length. Stop at CRLF so the connection's reader + // retains any following request for its own inspection and policy check. + read_more(client, raw, max_body_bytes, 1).await?; + } +} + async fn read_more( client: &mut C, raw: &mut Vec, max_body_bytes: usize, + remaining: usize, ) -> Result<()> { if raw.len() > max_body_bytes.saturating_mul(2).max(max_body_bytes) { return Err(miette!( @@ -183,7 +194,9 @@ async fn read_more( )); } let mut buf = [0u8; READ_BUF_SIZE]; - let n = client.read(&mut buf).await.into_diagnostic()?; + // Payload reads stop at the declared chunk boundary, including its CRLF. + let to_read = remaining.min(buf.len()); + let n = client.read(&mut buf[..to_read]).await.into_diagnostic()?; if n == 0 { return Err(miette!("HTTP chunked body ended before terminator")); } @@ -197,3 +210,129 @@ fn find_crlf(buf: &[u8], start: usize) -> Option { .position(|w| w == b"\r\n") .map(|p| start + p) } + +#[cfg(test)] +mod tests { + use std::collections::HashMap; + + use super::*; + use tokio::io::{AsyncWriteExt, BufReader}; + + fn chunked_request() -> L7Request { + L7Request { + action: "POST".into(), + target: "/mcp".into(), + query_params: HashMap::new(), + raw_header: + b"POST /mcp HTTP/1.1\r\nHost: example.test\r\nTransfer-Encoding: chunked\r\n\r\n" + .to_vec(), + body_length: BodyLength::Chunked, + } + } + + #[tokio::test] + async fn chunked_inspection_preserves_pipelined_request() { + let next = b"DELETE /next HTTP/1.1\r\nHost: example.test\r\n\r\n"; + let mut wire = b"5\r\nhello\r\n0\r\n\r\n".to_vec(); + wire.extend_from_slice(next); + let (mut sender, receiver) = tokio::io::duplex(8192); + sender.write_all(&wire).await.unwrap(); + sender.shutdown().await.unwrap(); + let mut reader = BufReader::with_capacity(8192, receiver); + let mut request = chunked_request(); + + let body = read_body_for_inspection(&mut reader, &mut request, 1024) + .await + .unwrap(); + assert_eq!(body, b"hello"); + let mut remaining = Vec::new(); + reader.read_to_end(&mut remaining).await.unwrap(); + assert_eq!(remaining, next, "inspection consumed the next request"); + } + + #[tokio::test] + async fn chunked_inspection_preserves_boundary_with_fragmentation_and_buffered_prefix() { + let next = b"GET /next HTTP/1.1\r\nHost: example.test\r\n\r\n"; + let encoded = b"2;ext=value\r\nhe\r\n3\r\nllo\r\n0\r\nX-Checksum: ignored\r\n\r\n"; + for capacity in [1, 2, 8192] { + for prefix_len in 0..=encoded.len() { + let mut request = chunked_request(); + request.raw_header.extend_from_slice(&encoded[..prefix_len]); + let mut wire = encoded[prefix_len..].to_vec(); + wire.extend_from_slice(next); + let mut reader = BufReader::with_capacity(capacity, wire.as_slice()); + let body = read_body_for_inspection(&mut reader, &mut request, 1024) + .await + .unwrap(); + assert_eq!(body, b"hello"); + assert!(matches!(request.body_length, BodyLength::ContentLength(5))); + assert_eq!( + request.raw_header, + b"POST /mcp HTTP/1.1\r\nHost: example.test\r\nContent-Length: 5\r\n\r\nhello" + ); + let mut remaining = Vec::new(); + reader.read_to_end(&mut remaining).await.unwrap(); + assert_eq!(remaining, next, "capacity={capacity}, prefix={prefix_len}"); + } + } + } + + #[tokio::test] + async fn chunked_inspection_preserves_request_after_empty_body() { + let next = b"GET /next HTTP/1.1\r\nHost: example.test\r\n\r\n"; + let mut wire = b"0\r\n\r\n".to_vec(); + wire.extend_from_slice(next); + let mut reader = BufReader::new(wire.as_slice()); + let mut request = chunked_request(); + assert!( + read_body_for_inspection(&mut reader, &mut request, 1024) + .await + .unwrap() + .is_empty() + ); + let mut remaining = Vec::new(); + reader.read_to_end(&mut remaining).await.unwrap(); + assert_eq!(remaining, next); + } + + #[tokio::test] + async fn chunked_inspection_rejects_malformed_or_incomplete_framing() { + for (encoded, expected) in [ + (b"z\r\n".as_slice(), "Invalid HTTP chunk size token"), + (b"\xff\r\n", "Invalid UTF-8 in HTTP chunk-size line"), + (b"1\r\nx!!", "HTTP chunk payload missing terminating CRLF"), + (b"1\r", "HTTP chunked body ended before terminator"), + (b"5\r\nhe", "HTTP chunked body ended before terminator"), + ( + b"0\r\nX-Trailer: unfinished", + "HTTP chunked body ended before terminator", + ), + ] { + let mut reader = BufReader::with_capacity(2, encoded); + let error = read_body_for_inspection(&mut reader, &mut chunked_request(), 1024) + .await + .unwrap_err(); + assert!(error.to_string().contains(expected), "{error}"); + } + } + + #[tokio::test] + async fn chunked_inspection_enforces_body_and_framing_limits() { + let mut reader = b"11\r\n".as_slice(); + let error = read_body_for_inspection(&mut reader, &mut chunked_request(), 16) + .await + .unwrap_err(); + assert!(error.to_string().contains("16 byte inspection limit")); + + for encoded in [ + format!("0;{}\r\n\r\n", "x".repeat(64)), + format!("0\r\nX-Trailer: {}\r\n\r\n", "x".repeat(64)), + ] { + let mut reader = encoded.as_bytes(); + let error = read_body_for_inspection(&mut reader, &mut chunked_request(), 16) + .await + .unwrap_err(); + assert!(error.to_string().contains("inspection framing limit")); + } + } +} diff --git a/crates/openshell-supervisor-network/src/l7/jsonrpc.rs b/crates/openshell-supervisor-network/src/l7/jsonrpc.rs index aea53edc08..6d767191fd 100644 --- a/crates/openshell-supervisor-network/src/l7/jsonrpc.rs +++ b/crates/openshell-supervisor-network/src/l7/jsonrpc.rs @@ -3,12 +3,19 @@ //! JSON-RPC 2.0 over HTTP L7 inspection. -use std::collections::HashMap; +use std::{collections::HashMap, fmt}; use miette::Result; +use openshell_core::mcp::{MAX_MCP_LEGACY_BATCH_MESSAGES, McpProtocolVersion}; +use serde::de::{DeserializeSeed, MapAccess, SeqAccess, Visitor}; use tokio::io::{AsyncRead, AsyncWrite}; -use tower_mcp_types::protocol::{ - JSONRPC_VERSION, JsonRpcNotification, JsonRpcRequest, McpNotification, McpRequest, +use tower_mcp_types::{ + inspection::{ + JsonRpcEnvelope, JsonRpcPayload, McpCallKind, McpDirection, McpInspection, + McpInspectionError, McpInspectionErrorKind, McpInspector, + McpMethodClassification as TowerMcpMethodClassification, + }, + protocol::{InitializeParams, JSONRPC_VERSION, McpRequest, RequestMeta, validate_meta_object}, }; use crate::l7::provider::{L7Provider, L7Request}; @@ -35,8 +42,13 @@ impl JsonRpcInspectionMode { /// Endpoint-specific JSON-RPC-family parser settings. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub struct JsonRpcInspectionOptions { + /// Select generic JSON-RPC or MCP semantic inspection. pub mode: JsonRpcInspectionMode, + /// Enforce the MCP recommended tool-name character and length boundary. pub mcp_strict_tool_names: bool, + /// Exact revision selected by the HTTP transport for non-initialize MCP + /// traffic. `None` permits only a bootstrap `initialize` request. + pub mcp_revision: Option, } impl JsonRpcInspectionOptions { @@ -44,8 +56,42 @@ impl JsonRpcInspectionOptions { Self { mode: JsonRpcInspectionMode::for_protocol(config.protocol), mcp_strict_tool_names: config.mcp_strict_tool_names, + // Policy declares an allowlist, not an effective wire revision. + // The transport must select one from the request version header + // or the declared legacy fallback. + mcp_revision: None, } } + + /// Return MCP inspection options for an initial `initialize` request. + /// + /// Any non-initialize MCP payload inspected with these options fails + /// closed because no effective revision has been selected yet. + #[must_use] + pub const fn mcp_bootstrap(strict_tool_names: bool) -> Self { + Self { + mode: JsonRpcInspectionMode::Mcp, + mcp_strict_tool_names: strict_tool_names, + mcp_revision: None, + } + } + + /// Return MCP inspection options for one exact effective wire revision. + #[must_use] + pub const fn mcp_selected(revision: McpProtocolVersion, strict_tool_names: bool) -> Self { + Self { + mode: JsonRpcInspectionMode::Mcp, + mcp_strict_tool_names: strict_tool_names, + mcp_revision: Some(revision), + } + } + + /// Bind an exact transport-selected MCP revision to existing options. + #[must_use] + pub const fn with_mcp_revision(mut self, revision: McpProtocolVersion) -> Self { + self.mcp_revision = Some(revision); + self + } } impl From for JsonRpcInspectionOptions { @@ -53,6 +99,10 @@ impl From for JsonRpcInspectionOptions { Self { mode, mcp_strict_tool_names: true, + // Mode alone cannot select an MCP wire revision. The resulting + // options accept bootstrap initialize only until the transport + // supplies exact per-request revision evidence. + mcp_revision: None, } } } @@ -142,10 +192,25 @@ pub struct JsonRpcRequestInfo { pub is_batch: bool, pub receive_stream: bool, pub has_response: bool, + /// Exact MCP revision used for semantic inspection. Bootstrap initialize, + /// receive-stream requests, and generic JSON-RPC leave this unset. + pub mcp_revision: Option, + /// Body fields mirrored into HTTP headers by sessionless MCP requests. + /// Populated only after the request passes its selected inspection profile. + pub mcp_http_metadata: Option, /// Typed inspection failure discovered before policy evaluation. pub error: Option, } +/// Inspected body fields that the MCP HTTP transport must compare with headers. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct McpHttpRequestMetadata { + /// The JSON-RPC method mirrored in `Mcp-Method`. + pub method: String, + /// Tool or prompt name, or resource URI, mirrored in `Mcp-Name` when required. + pub name: Option, +} + /// Stable kind of failure found while inspecting a JSON-RPC-family message. #[derive(Debug, Clone, Copy, PartialEq, Eq)] #[non_exhaustive] @@ -154,6 +219,13 @@ pub enum JsonRpcInspectionErrorKind { InvalidJson, /// The decoded JSON fails JSON-RPC or protocol-specific message validation. InvalidMessage, + /// MCP traffic lacked the transport-selected revision required for exact + /// semantic inspection. + RevisionNotSelected, + /// The selected MCP revision's immutable wire profile rejected the body. + McpProfileViolation, + /// An MCP lifecycle message appeared in a forbidden wire shape. + McpLifecycleViolation, } /// Typed JSON-RPC inspection failure with an inseparable kind and diagnostic. @@ -171,6 +243,13 @@ impl JsonRpcInspectionError { } } + fn invalid_json_with_detail(detail: impl Into) -> Self { + Self { + kind: JsonRpcInspectionErrorKind::InvalidJson, + detail: detail.into(), + } + } + fn invalid_message(detail: impl Into) -> Self { Self { kind: JsonRpcInspectionErrorKind::InvalidMessage, @@ -178,6 +257,27 @@ impl JsonRpcInspectionError { } } + fn revision_not_selected() -> Self { + Self { + kind: JsonRpcInspectionErrorKind::RevisionNotSelected, + detail: "MCP effective protocol revision has not been selected".to_string(), + } + } + + fn mcp_profile_violation(detail: impl Into) -> Self { + Self { + kind: JsonRpcInspectionErrorKind::McpProfileViolation, + detail: detail.into(), + } + } + + fn mcp_lifecycle_violation(detail: impl Into) -> Self { + Self { + kind: JsonRpcInspectionErrorKind::McpLifecycleViolation, + detail: detail.into(), + } + } + /// Return the stable failure kind used by typed control flow. #[must_use] pub const fn kind(&self) -> JsonRpcInspectionErrorKind { @@ -198,14 +298,29 @@ impl JsonRpcInspectionError { } } -impl std::fmt::Display for JsonRpcInspectionError { - fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { +impl fmt::Display for JsonRpcInspectionError { + fn fmt(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result { formatter.write_str(&self.detail) } } impl std::error::Error for JsonRpcInspectionError {} +/// Policy-relevant classification of one accepted MCP method. +/// +/// Known methods unavailable in the selected revision are rejected before a +/// call is produced, so the policy boundary represents only core methods that +/// are available and unknown extension methods that need an exact allow rule. +#[derive(Debug, Clone, Copy, PartialEq, Eq, serde::Serialize)] +#[serde(rename_all = "snake_case")] +#[non_exhaustive] +pub enum McpMethodClassification { + /// A core method defined by the selected MCP revision. + Available, + /// A method unknown to every supported core profile. + Extension, +} + #[derive(Debug, Clone, PartialEq, Eq)] pub struct JsonRpcCallInfo { /// JSON-RPC method, or the MCP method name after typed MCP parsing. @@ -217,6 +332,9 @@ pub struct JsonRpcCallInfo { /// MCP `tools/call` tool name when known. Generic JSON-RPC leaves this /// unset because params are not inspected. pub tool: Option, + /// Exact-profile MCP method classification retained for policy evaluation. + /// Generic JSON-RPC calls leave this unset. + pub mcp_classification: Option, /// Whether this call is a JSON-RPC notification without an `id`. /// /// MCP initialization is a request, so transport code uses this bit to @@ -232,6 +350,8 @@ impl JsonRpcRequestInfo { is_batch, receive_stream: false, has_response: false, + mcp_revision: None, + mcp_http_metadata: None, error: Some(error), } } @@ -244,6 +364,8 @@ impl JsonRpcRequestInfo { is_batch: false, receive_stream: true, has_response: false, + mcp_revision: None, + mcp_http_metadata: None, error: None, } } @@ -279,6 +401,10 @@ fn request_accepts_sse(request: &L7Request) -> bool { }) } /// Parse a JSON-RPC-family body using the endpoint's inspection mode. +/// +/// Mode-only MCP inspection accepts the bootstrap `initialize` request but +/// rejects later requests until the transport-selected request revision is supplied through +/// [`parse_jsonrpc_body_with_options`]. pub fn parse_jsonrpc_body( body: &[u8], inspection_mode: JsonRpcInspectionMode, @@ -291,10 +417,15 @@ pub fn parse_jsonrpc_body_with_options( body: &[u8], inspection_options: JsonRpcInspectionOptions, ) -> JsonRpcRequestInfo { - let Ok(value) = serde_json::from_slice::(body) else { - return JsonRpcRequestInfo::rejected(false, JsonRpcInspectionError::invalid_json()); + let value = match parse_unique_json_value(body) { + Ok(value) => value, + Err(error) => return JsonRpcRequestInfo::rejected(false, error), }; + if inspection_options.mode == JsonRpcInspectionMode::Mcp { + return parse_mcp_payload(&value, inspection_options); + } + if let serde_json::Value::Array(items) = value { if items.is_empty() { return JsonRpcRequestInfo::rejected( @@ -305,24 +436,8 @@ pub fn parse_jsonrpc_body_with_options( let mut calls = Vec::new(); let mut has_response = false; for item in &items { - match parse_jsonrpc_message(item, inspection_options) { - Ok(JsonRpcMessageInfo::Call(call)) => { - // Initialization establishes the MCP exchange before any - // profile-specific request rules apply. A batch has no - // single bootstrap exchange, so fail closed before policy - // evaluation can forward any member. - if inspection_options.mode == JsonRpcInspectionMode::Mcp - && call.method == "initialize" - { - return JsonRpcRequestInfo::rejected( - true, - JsonRpcInspectionError::invalid_message( - "MCP initialize must not appear in a batch", - ), - ); - } - calls.push(call); - } + match parse_jsonrpc_message(item) { + Ok(JsonRpcMessageInfo::Call(call)) => calls.push(call), Ok(JsonRpcMessageInfo::Response) => has_response = true, Err(error) => { return JsonRpcRequestInfo::rejected( @@ -339,16 +454,20 @@ pub fn parse_jsonrpc_body_with_options( is_batch: true, receive_stream: false, has_response, + mcp_revision: None, + mcp_http_metadata: None, error: None, }; } - match parse_jsonrpc_message(&value, inspection_options) { + match parse_jsonrpc_message(&value) { Ok(JsonRpcMessageInfo::Call(call)) => JsonRpcRequestInfo { calls: vec![call], is_batch: false, receive_stream: false, has_response: false, + mcp_revision: None, + mcp_http_metadata: None, error: None, }, Ok(JsonRpcMessageInfo::Response) => JsonRpcRequestInfo { @@ -356,6 +475,8 @@ pub fn parse_jsonrpc_body_with_options( is_batch: false, receive_stream: false, has_response: true, + mcp_revision: None, + mcp_http_metadata: None, error: None, }, Err(error) => { @@ -364,6 +485,161 @@ pub fn parse_jsonrpc_body_with_options( } } +#[derive(Debug)] +struct DuplicateJsonObjectKey { + key: String, + object_path: String, +} + +struct UniqueJsonValueSeed<'a> { + duplicate: &'a mut Option, + path: String, +} + +impl<'de> DeserializeSeed<'de> for UniqueJsonValueSeed<'_> { + type Value = serde_json::Value; + + fn deserialize(self, deserializer: D) -> std::result::Result + where + D: serde::Deserializer<'de>, + { + deserializer.deserialize_any(UniqueJsonValueVisitor { + duplicate: self.duplicate, + path: self.path, + }) + } +} + +struct UniqueJsonValueVisitor<'a> { + duplicate: &'a mut Option, + path: String, +} + +impl<'de> Visitor<'de> for UniqueJsonValueVisitor<'_> { + type Value = serde_json::Value; + + fn expecting(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result { + formatter.write_str("a JSON value without duplicate object keys") + } + + fn visit_bool(self, value: bool) -> std::result::Result { + Ok(serde_json::Value::Bool(value)) + } + + fn visit_i64(self, value: i64) -> std::result::Result { + Ok(serde_json::Value::Number(value.into())) + } + + fn visit_u64(self, value: u64) -> std::result::Result { + Ok(serde_json::Value::Number(value.into())) + } + + fn visit_f64(self, value: f64) -> std::result::Result + where + E: serde::de::Error, + { + serde_json::Number::from_f64(value) + .map(serde_json::Value::Number) + .ok_or_else(|| E::custom("JSON number must be finite")) + } + + fn visit_str(self, value: &str) -> std::result::Result + where + E: serde::de::Error, + { + Ok(serde_json::Value::String(value.to_string())) + } + + fn visit_string(self, value: String) -> std::result::Result { + Ok(serde_json::Value::String(value)) + } + + fn visit_none(self) -> std::result::Result { + Ok(serde_json::Value::Null) + } + + fn visit_some(self, deserializer: D) -> std::result::Result + where + D: serde::Deserializer<'de>, + { + UniqueJsonValueSeed { + duplicate: self.duplicate, + path: self.path, + } + .deserialize(deserializer) + } + + fn visit_unit(self) -> std::result::Result { + Ok(serde_json::Value::Null) + } + + fn visit_seq(self, mut sequence: A) -> std::result::Result + where + A: SeqAccess<'de>, + { + let mut values = Vec::new(); + while let Some(value) = sequence.next_element_seed(UniqueJsonValueSeed { + duplicate: &mut *self.duplicate, + path: format!("{}[{}]", self.path, values.len()), + })? { + values.push(value); + } + Ok(serde_json::Value::Array(values)) + } + + fn visit_map(self, mut object: A) -> std::result::Result + where + A: MapAccess<'de>, + { + let mut values = serde_json::Map::new(); + while let Some(key) = object.next_key::()? { + // Record the first duplicate but continue decoding so syntax and + // trailing-data failures still use serde_json's normal handling. + if values.contains_key(&key) && self.duplicate.is_none() { + *self.duplicate = Some(DuplicateJsonObjectKey { + key: key.clone(), + object_path: self.path.clone(), + }); + } + let value = object.next_value_seed(UniqueJsonValueSeed { + duplicate: &mut *self.duplicate, + path: format!("{}.{}", self.path, key), + })?; + values.insert(key, value); + } + Ok(serde_json::Value::Object(values)) + } +} + +/// Decode one JSON value while rejecting duplicate object keys recursively. +/// +/// Callers that inspect and then forward the original bytes must use this +/// boundary so `OpenShell` and the upstream cannot select different values for +/// the same object key. +pub(crate) fn parse_unique_json_value( + body: &[u8], +) -> std::result::Result { + let mut duplicate = None; + let mut deserializer = serde_json::Deserializer::from_slice(body); + let value = UniqueJsonValueSeed { + duplicate: &mut duplicate, + path: "$".to_string(), + } + .deserialize(&mut deserializer) + .map_err(|_| JsonRpcInspectionError::invalid_json())?; + deserializer + .end() + .map_err(|_| JsonRpcInspectionError::invalid_json())?; + + if let Some(duplicate) = duplicate { + return Err(JsonRpcInspectionError::invalid_json_with_detail(format!( + "duplicate JSON object key '{}' at {}", + duplicate.key, duplicate.object_path + ))); + } + Ok(value) +} + enum JsonRpcMessageInfo { Call(JsonRpcCallInfo), Response, @@ -373,7 +649,6 @@ enum JsonRpcMessageInfo { // only after the common JSON-RPC version/method/response checks. fn parse_jsonrpc_message( value: &serde_json::Value, - inspection_options: JsonRpcInspectionOptions, ) -> std::result::Result { let version = value .get("jsonrpc") @@ -395,22 +670,13 @@ fn parse_jsonrpc_message( } if has_method { - return parse_jsonrpc_call(value, inspection_options).map(JsonRpcMessageInfo::Call); + return parse_jsonrpc_call(value).map(JsonRpcMessageInfo::Call); } Err("missing or non-string 'method' field".to_string()) } -fn parse_jsonrpc_call( - value: &serde_json::Value, - inspection_options: JsonRpcInspectionOptions, -) -> std::result::Result { - // MCP mode delegates method-specific validation to tower-mcp-types. The - // generic mode intentionally remains looser for non-MCP JSON-RPC servers. - if inspection_options.mode == JsonRpcInspectionMode::Mcp { - return parse_mcp_call(value, inspection_options.mcp_strict_tool_names); - } - +fn parse_jsonrpc_call(value: &serde_json::Value) -> std::result::Result { let method = value .get("method") .and_then(|m| m.as_str()) @@ -419,6 +685,7 @@ fn parse_jsonrpc_call( method: method.to_string(), params: HashMap::new(), tool: None, + mcp_classification: None, is_notification: value.get("id").is_none(), }) } @@ -452,60 +719,395 @@ fn parse_jsonrpc_response(value: &serde_json::Value) -> std::result::Result<(), Ok(()) } -fn parse_mcp_call( +fn parse_mcp_payload( value: &serde_json::Value, - strict_tool_names: bool, -) -> std::result::Result { - if value.get("id").is_some() { - // Typed parsing validates known MCP params, but policy method profiles - // stay OpenShell-owned; see McpOptions in proto/sandbox.proto. - let request: JsonRpcRequest = serde_json::from_value(value.clone()) - .map_err(|error| format!("invalid MCP request: {error}"))?; - request - .validate() - .map_err(|error| format!("invalid MCP request: {error:?}"))?; - let mcp_request = McpRequest::from_jsonrpc(&request) - .map_err(|error| format!("invalid MCP request params: {error}"))?; - let tool = mcp_tool_name(&mcp_request); - if strict_tool_names && let Some(tool_name) = tool.as_deref() { - validate_mcp_tool_name(tool_name)?; - } - - let params = mcp_policy_params(tool.as_deref()); - return Ok(JsonRpcCallInfo { - method: mcp_request.method_name().to_string(), - params, + inspection_options: JsonRpcInspectionOptions, +) -> JsonRpcRequestInfo { + let payload = match JsonRpcPayload::inspect(value) { + Ok(payload) => payload, + Err(error) => { + return JsonRpcRequestInfo::rejected( + value.is_array(), + JsonRpcInspectionError::invalid_message(error.to_string()), + ); + } + }; + + // Legacy initialization proposes a revision rather than selecting one. + // Sessionless requests carry their revision in _meta and must never use + // this exemption, even if the transport has not selected a profile yet. + if payload_contains_method(&payload, "initialize") + && inspection_options.mcp_revision != Some(McpProtocolVersion::V2026_07_28) + && value + .pointer("/params/_meta/io.modelcontextprotocol~1protocolVersion") + .and_then(serde_json::Value::as_str) + != Some("2026-07-28") + { + return parse_mcp_initialize(payload); + } + + let Some(revision) = inspection_options.mcp_revision else { + return JsonRpcRequestInfo::rejected( + payload.is_batch(), + JsonRpcInspectionError::revision_not_selected(), + ); + }; + + let mut mcp_http_metadata = match inspect_mcp_request_metadata(value, revision) { + Ok(metadata) => metadata, + Err(error) => return JsonRpcRequestInfo::rejected(payload.is_batch(), error), + }; + + let inspection = + match inspect_mcp_payload_for_revision(&payload, revision, McpDirection::ClientToServer) { + Ok(inspection) => inspection, + Err(error) => return JsonRpcRequestInfo::rejected(payload.is_batch(), error), + }; + + let mut calls = Vec::with_capacity(inspection.methods().len()); + for method in inspection.methods() { + let named_request = match mcp_named_request_for_inspected_method( + inspection.payload(), + method.method(), + method.batch_index(), + ) { + Ok(request) => request, + Err(error) => { + return JsonRpcRequestInfo::rejected( + inspection.payload().is_batch(), + JsonRpcInspectionError::invalid_message(error), + ); + } + }; + // Use the same typed tool name for authorization and its HTTP mirror. + // Prompt names and resource URIs participate only in HTTP matching. + let (tool, http_name) = match named_request { + Some(McpRequest::CallTool(params)) => (Some(params.name.clone()), Some(params.name)), + Some(McpRequest::GetPrompt(params)) => (None, Some(params.name)), + Some(McpRequest::ReadResource(params)) => (None, Some(params.uri)), + None => (None, None), + Some(_) => { + return JsonRpcRequestInfo::rejected( + inspection.payload().is_batch(), + JsonRpcInspectionError::invalid_message( + "inspected MCP name-bearing method has an unexpected typed request", + ), + ); + } + }; + if let Some(metadata) = mcp_http_metadata.as_mut() { + metadata.name = http_name; + } + if inspection_options.mcp_strict_tool_names + && let Some(tool_name) = tool.as_deref() + && let Err(error) = validate_mcp_tool_name(tool_name) + { + return JsonRpcRequestInfo::rejected( + inspection.payload().is_batch(), + JsonRpcInspectionError::invalid_message(error), + ); + } + + let mcp_classification = match method.classification() { + TowerMcpMethodClassification::Available => McpMethodClassification::Available, + TowerMcpMethodClassification::Extension => McpMethodClassification::Extension, + TowerMcpMethodClassification::Unavailable => { + return JsonRpcRequestInfo::rejected( + inspection.payload().is_batch(), + unavailable_mcp_method_error(revision, method.method()), + ); + } + _ => { + return JsonRpcRequestInfo::rejected( + inspection.payload().is_batch(), + JsonRpcInspectionError::mcp_profile_violation(format!( + "MCP {revision} returned an unsupported classification for `{}`", + method.method() + )), + ); + } + }; + + calls.push(JsonRpcCallInfo { + method: method.method().to_string(), + params: mcp_policy_params(tool.as_deref()), tool, - is_notification: false, + mcp_classification: Some(mcp_classification), + is_notification: method.kind() == McpCallKind::Notification, }); } - // Notifications have no id and no response expectation. Validate them as - // MCP notifications but keep extension notifications addressable. - let notification: JsonRpcNotification = serde_json::from_value(value.clone()) - .map_err(|error| format!("invalid MCP notification: {error}"))?; - if notification.jsonrpc != JSONRPC_VERSION { - return Err(format!( - "unsupported JSON-RPC version '{}'", - notification.jsonrpc + JsonRpcRequestInfo { + calls, + is_batch: inspection.payload().is_batch(), + receive_stream: false, + has_response: payload_has_response(inspection.payload()), + mcp_revision: Some(revision), + mcp_http_metadata, + error: None, + } +} + +/// Validate the common request contract before method-specific inspection. +/// Extension requests need the same metadata as core requests, even when the +/// upstream inspector has no typed parameter schema for their method. +fn inspect_mcp_request_metadata( + value: &serde_json::Value, + revision: McpProtocolVersion, +) -> std::result::Result, JsonRpcInspectionError> { + // A legacy header or missing-header fallback cannot reinterpret a body + // that explicitly declares a different per-request protocol revision. + let messages = value + .as_array() + .map_or_else(|| std::slice::from_ref(value), Vec::as_slice); + for message in messages { + if let Some(declared) = + message.pointer("/params/_meta/io.modelcontextprotocol~1protocolVersion") + && declared.as_str() != Some(revision.as_str()) + { + return Err(JsonRpcInspectionError::mcp_profile_violation( + "MCP request metadata protocol version must match the selected revision", + )); + } + } + + if revision != McpProtocolVersion::V2026_07_28 { + return Ok(None); + } + let Some(method) = value.get("method").and_then(serde_json::Value::as_str) else { + return Err(JsonRpcInspectionError::mcp_profile_violation( + "MCP 2026-07-28 HTTP bodies must contain one request or extension notification", )); + }; + if value.get("id").is_none() { + // Cancellation for HTTP is closing the response stream. The core + // notification exists only for stdio; extension notifications retain + // their explicit policy gate and have no standardized HTTP metadata. + if method == "notifications/cancelled" { + return Err(JsonRpcInspectionError::mcp_profile_violation( + "MCP 2026-07-28 notifications/cancelled is not available over HTTP", + )); + } + return Ok(None); } - McpNotification::from_jsonrpc(¬ification) - .map_err(|error| format!("invalid MCP notification params: {error}"))?; - Ok(JsonRpcCallInfo { - method: notification.method, - params: HashMap::new(), - tool: None, - is_notification: true, + let meta = value.pointer("/params/_meta").ok_or_else(|| { + JsonRpcInspectionError::invalid_message("MCP 2026-07-28 requests require params._meta") + })?; + validate_meta_object(meta) + .map_err(|error| JsonRpcInspectionError::invalid_message(error.to_string()))?; + // Tower owns the capability and implementation schemas. Apply its common + // metadata contract here so extension requests receive the same checks + // even though the method inspector has no extension parameter schema. + let meta: RequestMeta = serde_json::from_value(meta.clone()) + .map_err(|error| JsonRpcInspectionError::invalid_message(error.to_string()))?; + meta.validate_for_version(revision.as_str()) + .map_err(|error| JsonRpcInspectionError::invalid_message(error.to_string()))?; + + Ok(Some(McpHttpRequestMetadata { + method: method.to_string(), + // Name-bearing parameters are projected from Tower's typed requests + // only after exact-revision method inspection succeeds. + name: None, + })) +} + +/// Inspect one complete JSON-RPC payload under an exact MCP wire profile. +/// +/// The caller supplies the peer direction because it is transport context, not +/// JSON-RPC data. This boundary enforces the revision's batch shape and bound, +/// typed method parameters, peer direction, and known-method availability. +/// Unknown extension methods remain valid for explicit policy authorization. +pub(crate) fn inspect_mcp_payload_for_revision( + payload: &JsonRpcPayload, + revision: McpProtocolVersion, + direction: McpDirection, +) -> std::result::Result { + // Bound work before typed inspection. Tower owns whether the selected + // revision permits batches; this limit applies to every attempted batch. + if payload.is_batch() && payload.len() > MAX_MCP_LEGACY_BATCH_MESSAGES { + return Err(JsonRpcInspectionError::mcp_profile_violation(format!( + "MCP {revision} batch contains {} messages; the maximum is {MAX_MCP_LEGACY_BATCH_MESSAGES}", + payload.len() + ))); + } + + let inspector = McpInspector::new(revision.as_str()) + .map_err(|error| JsonRpcInspectionError::mcp_profile_violation(error.to_string()))?; + let inspection = inspector + .inspect_payload(payload.clone(), Some(direction)) + .map_err(map_mcp_inspection_error)?; + + for method in inspection.methods() { + match method.classification() { + TowerMcpMethodClassification::Unavailable => { + return Err(unavailable_mcp_method_error(revision, method.method())); + } + TowerMcpMethodClassification::Available | TowerMcpMethodClassification::Extension => {} + _ => { + return Err(JsonRpcInspectionError::mcp_profile_violation(format!( + "MCP {revision} returned an unsupported classification for `{}`", + method.method() + ))); + } + } + } + + Ok(inspection) +} + +// Only a method recognized by Tower's closed core-method registry reaches +// this diagnostic. Policy permission cannot make it valid in another revision. +fn unavailable_mcp_method_error( + revision: McpProtocolVersion, + method: &str, +) -> JsonRpcInspectionError { + JsonRpcInspectionError::mcp_profile_violation(format!( + "MCP method `{method}` is unavailable in revision {revision}; use a method defined by this core revision, or check client/server support and mcp.versions before selecting another revision; allow rules and allow_all_known_mcp_methods cannot enable an unavailable method" + )) +} + +fn parse_mcp_initialize(payload: JsonRpcPayload) -> JsonRpcRequestInfo { + if payload.is_batch() { + return JsonRpcRequestInfo::rejected( + true, + JsonRpcInspectionError::mcp_lifecycle_violation( + "MCP `initialize` must be exactly one non-batched request", + ), + ); + } + + let Some(JsonRpcEnvelope::Request(request)) = payload.as_single() else { + return JsonRpcRequestInfo::rejected( + false, + JsonRpcInspectionError::mcp_lifecycle_violation( + "MCP `initialize` must be a request with an id", + ), + ); + }; + let Some(params) = request.params.as_ref() else { + return JsonRpcRequestInfo::rejected( + false, + JsonRpcInspectionError::invalid_message("MCP `initialize` params are required"), + ); + }; + if let Err(error) = serde_json::from_value::(params.clone()) { + return JsonRpcRequestInfo::rejected( + false, + JsonRpcInspectionError::invalid_message(format!( + "invalid MCP `initialize` params: {error}" + )), + ); + } + + JsonRpcRequestInfo { + calls: vec![JsonRpcCallInfo { + method: "initialize".to_string(), + params: HashMap::new(), + tool: None, + mcp_classification: Some(McpMethodClassification::Available), + is_notification: false, + }], + is_batch: false, + receive_stream: false, + has_response: false, + mcp_revision: None, + mcp_http_metadata: None, + error: None, + } +} + +fn payload_contains_method(payload: &JsonRpcPayload, expected: &str) -> bool { + if let Some(envelope) = payload.as_single() { + return envelope_method(envelope) == Some(expected); + } + payload.as_batch().is_some_and(|batch| { + batch + .messages() + .iter() + .any(|envelope| envelope_method(envelope) == Some(expected)) }) } -fn mcp_tool_name(request: &McpRequest) -> Option { - if let McpRequest::CallTool(params) = request { - Some(params.name.clone()) - } else { - None +fn envelope_method(envelope: &JsonRpcEnvelope) -> Option<&str> { + match envelope { + JsonRpcEnvelope::Request(request) => Some(request.method.as_str()), + JsonRpcEnvelope::Notification(notification) => Some(notification.method.as_str()), + _ => None, + } +} + +fn payload_has_response(payload: &JsonRpcPayload) -> bool { + if let Some(envelope) = payload.as_single() { + return matches!( + envelope, + JsonRpcEnvelope::Result(_) | JsonRpcEnvelope::Error(_) + ); + } + payload.as_batch().is_some_and(|batch| { + batch.messages().iter().any(|envelope| { + matches!( + envelope, + JsonRpcEnvelope::Result(_) | JsonRpcEnvelope::Error(_) + ) + }) + }) +} + +// The inspector retains untyped envelopes. Decode name-bearing requests with +// Tower so policy and HTTP metadata share its parameter schemas. +fn mcp_named_request_for_inspected_method( + payload: &JsonRpcPayload, + method: &str, + batch_index: Option, +) -> std::result::Result, String> { + if !matches!(method, "tools/call" | "prompts/get" | "resources/read") { + return Ok(None); + } + + let envelope = batch_index + .map_or_else( + || payload.as_single(), + |index| { + payload + .as_batch() + .and_then(|batch| batch.messages().get(index)) + }, + ) + .ok_or_else(|| "inspected MCP method is missing its JSON-RPC envelope".to_string())?; + let JsonRpcEnvelope::Request(request) = envelope else { + return Err(format!("MCP `{method}` must be a request with an id")); + }; + let request = McpRequest::from_jsonrpc(request) + .map_err(|error| format!("invalid MCP `{method}` params: {error}"))?; + Ok(Some(request)) +} + +fn map_mcp_inspection_error(error: McpInspectionError) -> JsonRpcInspectionError { + match error.kind() { + McpInspectionErrorKind::BatchUnavailable => JsonRpcInspectionError::mcp_profile_violation( + format!("{error}; send each JSON-RPC message in a separate request for this revision"), + ), + McpInspectionErrorKind::DirectionMismatch => { + JsonRpcInspectionError::mcp_profile_violation(format!( + "{error}; send the method from the peer role defined by this revision; an allow rule cannot change its direction" + )) + } + McpInspectionErrorKind::UnsupportedProfile => { + JsonRpcInspectionError::mcp_profile_violation(format!( + "{error}; use an MCP revision supported by this OpenShell build" + )) + } + McpInspectionErrorKind::InitializeInBatch => { + JsonRpcInspectionError::mcp_lifecycle_violation(error.to_string()) + } + McpInspectionErrorKind::JsonRpc + | McpInspectionErrorKind::MessageKindMismatch + | McpInspectionErrorKind::MissingParams + | McpInspectionErrorKind::InvalidParams => { + JsonRpcInspectionError::invalid_message(error.to_string()) + } + _ => JsonRpcInspectionError::mcp_profile_violation(error.to_string()), } } @@ -538,6 +1140,21 @@ mod tests { use super::*; use std::collections::HashMap; + fn mcp_options(revision: McpProtocolVersion) -> JsonRpcInspectionOptions { + JsonRpcInspectionOptions::mcp_selected(revision, true) + } + + #[test] + fn every_policy_revision_selects_the_matching_tower_profile() { + // Policy support is deliberately independent of Tower's vocabulary; + // every enabled revision must still have an exact inspector profile. + for revision in McpProtocolVersion::ALL { + let inspector = McpInspector::new(revision.as_str()) + .expect("supported policy revision must have an inspector"); + assert_eq!(inspector.revision().as_str(), revision.as_str()); + } + } + #[test] fn parses_method_from_request_body() { let body = br#"{"jsonrpc":"2.0","id":1,"method":"initialize","params":{}}"#; @@ -574,6 +1191,216 @@ mod tests { assert!(info.error.is_none()); } + fn sessionless_request(method: &str, mut params: serde_json::Value) -> serde_json::Value { + params["_meta"] = serde_json::json!({ + "io.modelcontextprotocol/protocolVersion": "2026-07-28", + "io.modelcontextprotocol/clientCapabilities": {} + }); + serde_json::json!({"jsonrpc": "2.0", "id": 1, "method": method, "params": params}) + } + + #[test] + fn mcp_sessionless_inspects_core_and_extension_requests_without_initialize() { + for (method, params, name, classification) in [ + ( + "server/discover", + serde_json::json!({}), + None, + McpMethodClassification::Available, + ), + ( + "tools/list", + serde_json::json!({}), + None, + McpMethodClassification::Available, + ), + ( + "tools/call", + serde_json::json!({"name": "echo", "arguments": {}}), + Some("echo"), + McpMethodClassification::Available, + ), + ( + "prompts/get", + serde_json::json!({"name": "summarize"}), + Some("summarize"), + McpMethodClassification::Available, + ), + ( + "resources/read", + serde_json::json!({"uri": "file:///guide.txt"}), + Some("file:///guide.txt"), + McpMethodClassification::Available, + ), + ( + "subscriptions/listen", + serde_json::json!({"notifications": {"toolsListChanged": true}}), + None, + McpMethodClassification::Available, + ), + ( + "vendor/inspect", + serde_json::json!({}), + None, + McpMethodClassification::Extension, + ), + ] { + let body = serde_json::to_vec(&sessionless_request(method, params)).unwrap(); + let info = parse_jsonrpc_body_with_options( + &body, + mcp_options(McpProtocolVersion::V2026_07_28), + ); + assert!(info.error.is_none(), "{method}: {:?}", info.error); + assert_eq!(info.calls.len(), 1); + assert_eq!(info.calls[0].mcp_classification, Some(classification)); + let expected_tool = if method == "tools/call" { name } else { None }; + assert_eq!(info.calls[0].tool.as_deref(), expected_tool); + assert_eq!( + info.calls[0].params.get("name").map(String::as_str), + expected_tool + ); + assert_eq!( + info.mcp_http_metadata, + Some(McpHttpRequestMetadata { + method: method.to_string(), + name: name.map(str::to_string), + }) + ); + } + } + + #[test] + fn mcp_sessionless_requires_valid_common_metadata_for_extension_requests_too() { + for method in ["server/discover", "vendor/inspect"] { + for meta in [ + serde_json::Value::Null, + serde_json::json!({}), + serde_json::json!({"io.modelcontextprotocol/protocolVersion": "2026-07-28"}), + serde_json::json!({"io.modelcontextprotocol/protocolVersion": "2025-11-25", "io.modelcontextprotocol/clientCapabilities": {}}), + serde_json::json!({"io.modelcontextprotocol/protocolVersion": "2026-07-28", "io.modelcontextprotocol/clientCapabilities": "invalid"}), + serde_json::json!({"io.modelcontextprotocol/protocolVersion": "2026-07-28", "io.modelcontextprotocol/clientCapabilities": {}, "invalid key": true}), + ] { + let mut request = sessionless_request(method, serde_json::json!({})); + request["params"]["_meta"] = meta; + let info = parse_jsonrpc_body_with_options( + &serde_json::to_vec(&request).unwrap(), + mcp_options(McpProtocolVersion::V2026_07_28), + ); + assert!(info.error.is_some(), "accepted {request}"); + assert!(info.mcp_http_metadata.is_none()); + } + } + } + + #[test] + fn mcp_sessionless_preserves_arbitrary_extension_and_experimental_settings() { + for method in ["server/discover", "vendor/inspect"] { + let mut request = sessionless_request(method, serde_json::json!({})); + request["params"]["_meta"]["io.modelcontextprotocol/clientCapabilities"] = serde_json::json!({ + "roots": {"listChanged": true, "deprecated": {}}, + "sampling": {"tools": {}, "context": {}, "deprecated": {}}, + "elicitation": {"form": {}, "url": {}}, + "tasks": { + "list": {}, "cancel": {}, + "requests": {"sampling": {"createMessage": {}}, "elicitation": {"create": {}}} + }, + "experimental": {"list": [1, true, null], "label": "custom", "optional": null}, + "extensions": {"com.example/custom": {"list": [1, true, null], "optional": null}}, + "vendorData": [] + }); + let info = parse_jsonrpc_body_with_options( + &serde_json::to_vec(&request).unwrap(), + mcp_options(McpProtocolVersion::V2026_07_28), + ); + assert!(info.error.is_none(), "{method}: {:?}", info.error); + assert!(info.mcp_http_metadata.is_some()); + } + } + + #[test] + fn mcp_sessionless_validates_subscription_parameters() { + for (notifications, valid) in [ + (serde_json::json!({}), true), + (serde_json::json!({"toolsListChanged": true}), true), + (serde_json::json!(null), false), + (serde_json::json!({"toolsListChanged": []}), false), + ] { + let request = sessionless_request( + "subscriptions/listen", + serde_json::json!({"notifications": notifications}), + ); + let info = parse_jsonrpc_body_with_options( + &serde_json::to_vec(&request).unwrap(), + mcp_options(McpProtocolVersion::V2026_07_28), + ); + assert_eq!(info.error.is_none(), valid, "{request}: {:?}", info.error); + assert_eq!(info.mcp_http_metadata.is_some(), valid); + } + } + + #[test] + fn mcp_sessionless_metadata_cannot_be_inspected_as_a_legacy_revision() { + let body = + serde_json::to_vec(&sessionless_request("tools/list", serde_json::json!({}))).unwrap(); + for revision in [ + McpProtocolVersion::V2025_03_26, + McpProtocolVersion::V2025_06_18, + McpProtocolVersion::V2025_11_25, + ] { + let info = parse_jsonrpc_body_with_options(&body, mcp_options(revision)); + assert!( + info.error.is_some(), + "sessionless body accepted as {revision}" + ); + } + } + + #[test] + fn mcp_sessionless_rejects_legacy_lifecycle_and_non_http_message_shapes() { + let options = mcp_options(McpProtocolVersion::V2026_07_28); + let initialize = sessionless_request( + "initialize", + serde_json::json!({ + "protocolVersion": "2026-07-28", "capabilities": {}, "clientInfo": {"name": "test", "version": "1"} + }), + ); + let bootstrap = parse_jsonrpc_body_with_options( + &serde_json::to_vec(&initialize).unwrap(), + JsonRpcInspectionOptions::mcp_bootstrap(true), + ); + assert!( + bootstrap.error.is_some(), + "sessionless initialize used legacy bootstrap" + ); + for request in [ + initialize, + sessionless_request("logging/setLevel", serde_json::json!({"level": "info"})), + sessionless_request( + "resources/subscribe", + serde_json::json!({"uri": "file:///guide.txt"}), + ), + serde_json::json!({"jsonrpc": "2.0", "method": "notifications/initialized"}), + serde_json::json!({"jsonrpc": "2.0", "method": "notifications/cancelled", "params": {"requestId": 1}}), + serde_json::json!({"jsonrpc": "2.0", "id": 1, "result": {}}), + serde_json::json!([sessionless_request("tools/list", serde_json::json!({}))]), + sessionless_request("subscriptions/listen", serde_json::json!({})), + ] { + let info = + parse_jsonrpc_body_with_options(&serde_json::to_vec(&request).unwrap(), options); + assert!(info.error.is_some(), "accepted {request}"); + } + let extension = parse_jsonrpc_body_with_options( + br#"{"jsonrpc":"2.0","method":"vendor/notice","params":{}}"#, + options, + ); + assert!(extension.error.is_none(), "{:?}", extension.error); + assert_eq!( + extension.calls[0].mcp_classification, + Some(McpMethodClassification::Extension) + ); + assert!(extension.mcp_http_metadata.is_none()); + } + #[test] fn inspection_error_kinds_distinguish_invalid_json_and_messages() { fn assert_inspection_error( @@ -615,6 +1442,85 @@ mod tests { ); } + #[test] + fn rejects_duplicate_jsonrpc_envelope_keys_before_semantic_inspection() { + let fixtures: &[(&str, &[u8])] = &[ + ( + "jsonrpc", + br#"{"jsonrpc":"2.0","jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}"#, + ), + ( + "id", + br#"{"jsonrpc":"2.0","id":1,"id":2,"method":"tools/list","params":{}}"#, + ), + ( + "method", + br#"{"jsonrpc":"2.0","id":1,"method":"tools/list","method":"vendor/other","params":{}}"#, + ), + ( + "params", + br#"{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{},"params":{"cursor":"next"}}"#, + ), + ]; + + for (key, body) in fixtures { + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_11_25)); + assert!(info.calls.is_empty()); + assert_eq!( + info.error.as_ref().map(JsonRpcInspectionError::kind), + Some(JsonRpcInspectionErrorKind::InvalidJson), + "duplicate {key} must fail before MCP inspection: {info:?}" + ); + assert!( + info.error + .as_ref() + .is_some_and(|error| error.detail().contains(key)), + "duplicate-key detail should identify {key}: {info:?}" + ); + } + } + + #[test] + fn rejects_duplicate_keys_recursively_but_allows_same_key_in_distinct_objects() { + let duplicate = br#"{ + "jsonrpc":"2.0", + "id":1, + "method":"tools/call", + "params":{ + "name":"search_web", + "arguments":{"query":"allowed","query":"different"} + } + }"#; + let rejected = parse_jsonrpc_body_with_options( + duplicate, + mcp_options(McpProtocolVersion::V2025_11_25), + ); + assert!(rejected.calls.is_empty()); + assert_eq!( + rejected.error.as_ref().map(JsonRpcInspectionError::kind), + Some(JsonRpcInspectionErrorKind::InvalidJson) + ); + + let distinct_objects = br#"{ + "jsonrpc":"2.0", + "id":1, + "method":"tools/call", + "params":{ + "name":"search_web", + "arguments":{"left":{"value":1},"right":{"value":2}} + } + }"#; + let accepted = parse_jsonrpc_body_with_options( + distinct_objects, + mcp_options(McpProtocolVersion::V2025_11_25), + ); + assert!( + accepted.error.is_none(), + "keys may repeat in different objects: {accepted:?}" + ); + } + #[test] fn ignores_params_when_extracting_method() { let body = br#"{"jsonrpc":"2.0","id":1,"method":"reports.search","params":{"query":"quarterly","filters":{"scope":"workspace/main"}}}"#; @@ -649,7 +1555,8 @@ mod tests { #[test] fn mcp_mode_validates_known_methods_and_extracts_tool() { let body = br#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"search_web","arguments":{"query":"openshell"}}}"#; - let info = parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp); + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_11_25)); assert!(info.error.is_none(), "expected valid MCP call: {info:?}"); let call = info.calls.first().expect("single MCP call"); @@ -660,12 +1567,18 @@ mod tests { Some("search_web") ); assert_eq!(call.params.len(), 1); + assert_eq!( + call.mcp_classification, + Some(McpMethodClassification::Available) + ); + assert_eq!(info.mcp_revision, Some(McpProtocolVersion::V2025_11_25)); } #[test] fn mcp_mode_rejects_non_recommended_tool_names_by_default() { let body = br#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"read status","arguments":{}}}"#; - let info = parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp); + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_11_25)); assert!(info.calls.is_empty()); assert!( @@ -682,10 +1595,7 @@ mod tests { let body = br#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"read status","arguments":{}}}"#; let info = parse_jsonrpc_body_with_options( body, - JsonRpcInspectionOptions { - mode: JsonRpcInspectionMode::Mcp, - mcp_strict_tool_names: false, - }, + JsonRpcInspectionOptions::mcp_selected(McpProtocolVersion::V2025_11_25, false), ); let call = info @@ -703,7 +1613,8 @@ mod tests { #[test] fn mcp_mode_rejects_invalid_known_method_params() { let body = br#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"arguments":{"query":"openshell"}}}"#; - let info = parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp); + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_11_25)); assert!(info.calls.is_empty()); assert_eq!( @@ -714,7 +1625,7 @@ mod tests { info.error .as_ref() .map(JsonRpcInspectionError::detail) - .is_some_and(|error| error.contains("invalid MCP request params")), + .is_some_and(|error| error.contains("invalid `tools/call` params")), "expected MCP params validation error, got {info:?}" ); } @@ -723,7 +1634,8 @@ mod tests { fn mcp_mode_allows_unknown_extension_methods() { let body = br#"{"jsonrpc":"2.0","id":1,"method":"vendor/extension","params":{"name":"custom"}}"#; - let info = parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp); + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_06_18)); assert!( info.error.is_none(), @@ -733,17 +1645,16 @@ mod tests { info.calls.first().map(|call| call.method.as_str()), Some("vendor/extension") ); - assert!( - info.calls - .first() - .is_some_and(|call| call.params.is_empty() && call.tool.is_none()) - ); + assert!(info.calls.first().is_some_and(|call| call.params.is_empty() + && call.tool.is_none() + && call.mcp_classification == Some(McpMethodClassification::Extension))); } #[test] fn mcp_mode_ignores_tool_arguments_when_extracting_policy_params() { let body = br#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"read_status","arguments":{"scope.key":"literal","scope":{"key":"nested"}}}}"#; - let info = parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp); + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_11_25)); let call = info.calls.first().expect("single MCP call"); assert!(info.error.is_none(), "expected valid MCP call: {info:?}"); @@ -862,21 +1773,176 @@ mod tests { } #[test] - fn preserves_current_mcp_batch_acceptance() { + fn mcp_batch_acceptance_follows_the_selected_revision() { let body = br#"[ {"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"read_status","arguments":{}}}, {"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"search_web","arguments":{"query":"openshell"}}} ]"#; - let info = parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp); + let march = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_03_26)); assert!( - info.error.is_none(), - "MCP batch should remain valid: {info:?}" + march.error.is_none(), + "March MCP batch should be valid: {march:?}" ); + assert!(march.is_batch); + assert_eq!(march.calls.len(), 2); + assert_eq!(march.calls[0].tool.as_deref(), Some("read_status")); + assert_eq!(march.calls[1].tool.as_deref(), Some("search_web")); + + for revision in [ + McpProtocolVersion::V2025_06_18, + McpProtocolVersion::V2025_11_25, + ] { + let info = parse_jsonrpc_body_with_options(body, mcp_options(revision)); + assert!(info.calls.is_empty()); + assert!(info.is_batch); + assert_eq!( + info.error.as_ref().map(JsonRpcInspectionError::kind), + Some(JsonRpcInspectionErrorKind::McpProfileViolation), + "revision {revision} must reject batches: {info:?}" + ); + } + } + + #[test] + fn march_batch_limit_counts_all_top_level_messages() { + let mut messages = (0..MAX_MCP_LEGACY_BATCH_MESSAGES) + .map(|index| { + serde_json::json!({ + "jsonrpc": "2.0", + "id": index, + "result": {"ok": true} + }) + }) + .collect::>(); + let at_limit = serde_json::to_vec(&serde_json::Value::Array(messages.clone())) + .expect("serialize in-bound MCP batch fixture"); + let accepted = parse_jsonrpc_body_with_options( + &at_limit, + mcp_options(McpProtocolVersion::V2025_03_26), + ); + assert!( + accepted.error.is_none(), + "64-member March batch should be accepted: {accepted:?}" + ); + + messages.push(serde_json::json!({ + "jsonrpc": "2.0", + "id": MAX_MCP_LEGACY_BATCH_MESSAGES, + "result": {"ok": true} + })); + let over_limit = serde_json::to_vec(&serde_json::Value::Array(messages)) + .expect("serialize over-limit MCP batch fixture"); + let info = parse_jsonrpc_body_with_options( + &over_limit, + mcp_options(McpProtocolVersion::V2025_03_26), + ); + + assert!(info.calls.is_empty()); assert!(info.is_batch); - assert_eq!(info.calls.len(), 2); - assert_eq!(info.calls[0].tool.as_deref(), Some("read_status")); - assert_eq!(info.calls[1].tool.as_deref(), Some("search_web")); + assert_eq!( + info.error.as_ref().map(JsonRpcInspectionError::kind), + Some(JsonRpcInspectionErrorKind::McpProfileViolation) + ); + assert!( + info.error + .as_ref() + .is_some_and(|error| error.detail().contains("maximum is 64")), + "expected total-message batch bound, got {info:?}" + ); + } + + #[test] + fn initialize_proposal_is_not_treated_as_an_effective_revision() { + let body = br#"{ + "jsonrpc":"2.0", + "id":"bootstrap-1", + "method":"initialize", + "params":{ + "protocolVersion":"2099-12-31", + "capabilities":{}, + "clientInfo":{"name":"test-client","version":"1.0"} + } + }"#; + let info = + parse_jsonrpc_body_with_options(body, JsonRpcInspectionOptions::mcp_bootstrap(true)); + + assert!(info.error.is_none(), "initialize should parse: {info:?}"); + assert_eq!(info.calls.len(), 1); + assert_eq!(info.calls[0].method, "initialize"); + assert_eq!( + info.calls[0].mcp_classification, + Some(McpMethodClassification::Available) + ); + assert_eq!(info.mcp_revision, None); + } + + #[test] + fn non_initialize_mcp_requires_a_transport_selected_revision() { + let body = br#"{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}"#; + let info = + parse_jsonrpc_body_with_options(body, JsonRpcInspectionOptions::mcp_bootstrap(true)); + + assert!(info.calls.is_empty()); + assert_eq!( + info.error.as_ref().map(JsonRpcInspectionError::kind), + Some(JsonRpcInspectionErrorKind::RevisionNotSelected) + ); + } + + #[test] + fn known_method_unavailable_in_selected_revision_is_rejected() { + let body = br#"{"jsonrpc":"2.0","id":1,"method":"tasks/get","params":{"taskId":"task-1"}}"#; + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_06_18)); + + assert!(info.calls.is_empty()); + assert_eq!( + info.error.as_ref().map(JsonRpcInspectionError::kind), + Some(JsonRpcInspectionErrorKind::McpProfileViolation) + ); + let detail = info.error.as_ref().expect("unavailable method").detail(); + assert!(detail.contains("`tasks/get` is unavailable in revision 2025-06-18")); + assert!(detail.contains("check client/server support and mcp.versions")); + assert!(detail.contains("allow_all_known_mcp_methods cannot enable")); + assert!( + !detail.contains("task-1"), + "task params must not be reflected" + ); + + let november = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_11_25)); + assert!(november.error.is_none(), "November task remains valid"); + } + + #[test] + fn server_originated_method_is_rejected_in_client_to_server_direction() { + let body = br#"{ + "jsonrpc":"2.0", + "id":1, + "method":"sampling/createMessage", + "params":{"messages":[],"maxTokens":1} + }"#; + let info = + parse_jsonrpc_body_with_options(body, mcp_options(McpProtocolVersion::V2025_03_26)); + + assert!(info.calls.is_empty()); + assert_eq!( + info.error.as_ref().map(JsonRpcInspectionError::kind), + Some(JsonRpcInspectionErrorKind::McpProfileViolation) + ); + assert!( + info.error + .as_ref() + .is_some_and(|error| error.detail().contains("client-to-server")), + "expected direction mismatch evidence, got {info:?}" + ); + assert!(info.error.as_ref().is_some_and(|error| { + error + .detail() + .contains("an allow rule cannot change its direction") + })); } #[test] @@ -885,14 +1951,15 @@ mod tests { {"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"test","version":"1"}}}, {"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"read_status","arguments":{}}} ]"#; - let info = parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp); + let info = + parse_jsonrpc_body_with_options(body, JsonRpcInspectionOptions::mcp_bootstrap(true)); assert!(info.is_batch); assert!(info.calls.is_empty()); assert!(!info.has_response); assert_eq!( info.error.as_ref().map(JsonRpcInspectionError::detail), - Some("MCP initialize must not appear in a batch") + Some("MCP `initialize` must be exactly one non-batched request") ); } diff --git a/crates/openshell-supervisor-network/src/l7/mcp.rs b/crates/openshell-supervisor-network/src/l7/mcp.rs index d5fbaed8bc..26538c9467 100644 --- a/crates/openshell-supervisor-network/src/l7/mcp.rs +++ b/crates/openshell-supervisor-network/src/l7/mcp.rs @@ -1,21 +1,25 @@ // SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. // SPDX-License-Identifier: Apache-2.0 -//! MCP Streamable HTTP request-version selection. +//! MCP Streamable HTTP revision selection and request metadata validation. +use base64::Engine; +use base64::engine::general_purpose::STANDARD; use openshell_core::mcp::McpProtocolVersion; use crate::l7::jsonrpc::JsonRpcRequestInfo; use crate::l7::provider::L7Request; const MCP_PROTOCOL_VERSION_HEADER: &str = "mcp-protocol-version"; +const MCP_METHOD_HEADER: &str = "mcp-method"; +const MCP_NAME_HEADER: &str = "mcp-name"; /// Protocol revision selected for one MCP HTTP request. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub(super) enum McpRequestProtocolVersion { - /// A valid standalone initialize request selects its revision in the JSON-RPC body. + /// A valid standalone legacy initialize request negotiates its revision in the body. Initialization, - /// A subsequent request selected an exact revision from its header or the legacy fallback. + /// A request selected an exact revision from its header or the legacy fallback. Selected(McpProtocolVersion), } @@ -26,16 +30,37 @@ pub(super) enum McpProtocolVersionError { InvalidHeader, /// The header value is not an MCP revision supported by this `OpenShell` build. UnsupportedHeaderValue, + /// Required HTTP metadata is missing, malformed, or differs from the inspected body. + InvalidRequestMetadata, + /// The selected revision does not support this HTTP method. + MethodNotAllowed, /// The selected supported revision is absent from the endpoint allowlist. NotAllowed(McpProtocolVersion), } impl McpProtocolVersionError { + /// Explain a rejected revision using validated header metadata. + /// + /// Missing headers select the 2025-03-26 fallback. Only the validated parser + /// may identify that case; duplicate or hop-by-hop headers must keep their + /// own rejection reason without being described as absent. + pub(super) fn rejection_detail(self, request: &L7Request) -> String { + match self { + Self::NotAllowed(version) => { + format!("{self}; {}", selected_revision_context(request, version)) + } + _ => self.to_string(), + } + } + /// Return the HTTP status for this transport or policy rejection. #[must_use] pub(super) const fn http_status(self) -> &'static str { match self { - Self::InvalidHeader | Self::UnsupportedHeaderValue => "400 Bad Request", + Self::InvalidHeader | Self::UnsupportedHeaderValue | Self::InvalidRequestMetadata => { + "400 Bad Request" + } + Self::MethodNotAllowed => "405 Method Not Allowed", Self::NotAllowed(_) => "403 Forbidden", } } @@ -46,6 +71,8 @@ impl McpProtocolVersionError { match self { Self::InvalidHeader => "invalid_mcp_protocol_version_header", Self::UnsupportedHeaderValue => "unsupported_mcp_protocol_version", + Self::InvalidRequestMetadata => "invalid_mcp_request_metadata", + Self::MethodNotAllowed => "mcp_http_method_not_allowed", Self::NotAllowed(_) => "mcp_protocol_version_not_allowed", } } @@ -55,14 +82,20 @@ impl std::fmt::Display for McpProtocolVersionError { fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { match self { Self::InvalidHeader => formatter.write_str( - "MCP-Protocol-Version must contain one non-empty end-to-end header value", + "MCP-Protocol-Version must contain one non-empty end-to-end header value; send exactly one revision, without duplicates, comma-separated values, or Connection nomination", ), Self::UnsupportedHeaderValue => { - formatter.write_str("MCP-Protocol-Version names an unsupported protocol version") + formatter.write_str("MCP-Protocol-Version names a revision unsupported by this OpenShell build; use a supported client/server revision permitted by mcp.versions") + } + Self::InvalidRequestMetadata => { + formatter.write_str("MCP request headers must match the inspected request metadata; send MCP-Protocol-Version, Mcp-Method, and any required Mcp-Name consistently with the 2026-07-28 message") + } + Self::MethodNotAllowed => { + formatter.write_str("MCP protocol version 2026-07-28 requires HTTP POST; send a POST message instead of a legacy GET stream or DELETE session request") } Self::NotAllowed(version) => write!( formatter, - "MCP protocol version {version} is not allowed by endpoint policy" + "MCP protocol version {version} is not allowed by endpoint policy; use a client/server revision permitted by mcp.versions" ), } } @@ -70,34 +103,126 @@ impl std::fmt::Display for McpProtocolVersionError { impl std::error::Error for McpProtocolVersionError {} +/// Describe how the request selected its MCP revision. +/// +/// Initial and post-middleware inspection supply their current request so the +/// explanation describes the headers used for that selection. +pub(super) fn selected_revision_context( + request: &L7Request, + version: McpProtocolVersion, +) -> String { + match request_protocol_version_header(&request.raw_header) { + Ok(None) => format!( + "selected MCP revision {version} from the missing MCP-Protocol-Version header fallback; send the client/server revision explicitly in that header" + ), + Ok(Some(_)) => format!("selected MCP revision {version} from MCP-Protocol-Version"), + Err(_) => format!("selected MCP revision {version}"), + } +} + /// Select and authorize the protocol revision for one MCP HTTP request. /// -/// MCP initialization negotiates its revision in the JSON-RPC body and is the -/// only request exempt from the header. Every other request is self-contained: -/// an absent header selects the specification-defined `2025-03-26` fallback, -/// and the resulting supported revision must appear in the endpoint allowlist. +/// Legacy initialization negotiates its revision in the body. It cannot exempt +/// an explicit `2026-07-28` request or an endpoint allowing only that revision. +/// An absent header otherwise selects the `2025-03-26` fallback; the selected +/// revision must appear in the endpoint allowlist. pub(super) fn select_request_protocol_version( request: &L7Request, info: &JsonRpcRequestInfo, allowed_versions: &[McpProtocolVersion], ) -> Result { - if is_standalone_initialize(info) { - return Ok(McpRequestProtocolVersion::Initialization); - } - - let version = match request_protocol_version_header(&request.raw_header)? { + let header = request_protocol_version_header(&request.raw_header)?; + let version = match header { Some(value) => value .parse::() .map_err(|_| McpProtocolVersionError::UnsupportedHeaderValue)?, None => McpProtocolVersion::V2025_03_26, }; + let modern_only = !allowed_versions.is_empty() + && allowed_versions + .iter() + .all(|version| *version == McpProtocolVersion::V2026_07_28); + if is_standalone_initialize(info) && version != McpProtocolVersion::V2026_07_28 && !modern_only + { + return Ok(McpRequestProtocolVersion::Initialization); + } + if !allowed_versions.contains(&version) { return Err(McpProtocolVersionError::NotAllowed(version)); } + // Sessionless MCP carries each message in its own POST. Standalone GET + // streams and DELETE session termination belong to the legacy revisions. + if version == McpProtocolVersion::V2026_07_28 && request.action != "POST" { + return Err(McpProtocolVersionError::MethodNotAllowed); + } Ok(McpRequestProtocolVersion::Selected(version)) } +/// Match the sessionless revision's mandatory HTTP mirrors to the inspected body. +/// +/// Only validated request bodies carry this metadata. Extension notifications +/// have no specified mirrored-header contract; unknown `Mcp-Param-*` fields are +/// also left to the server that owns the tool schema. +pub(super) fn validate_request_metadata( + request: &L7Request, + info: &JsonRpcRequestInfo, +) -> Result<(), McpProtocolVersionError> { + if info.mcp_revision != Some(McpProtocolVersion::V2026_07_28) { + return Ok(()); + } + let Some(metadata) = &info.mcp_http_metadata else { + return Ok(()); + }; + let invalid = McpProtocolVersionError::InvalidRequestMetadata; + let version = request_header_value(&request.raw_header, MCP_PROTOCOL_VERSION_HEADER) + .map_err(|_| invalid)? + .ok_or(invalid)?; + if version != McpProtocolVersion::V2026_07_28.as_str() { + return Err(invalid); + } + let method = request_header_value(&request.raw_header, MCP_METHOD_HEADER) + .map_err(|_| invalid)? + .ok_or(invalid)?; + if !is_plain_header_value(method) || method != metadata.method { + return Err(invalid); + } + if let Some(name) = &metadata.name { + let value = request_header_value(&request.raw_header, MCP_NAME_HEADER) + .map_err(|_| invalid)? + .ok_or(invalid)?; + // Decoding happens before equality so header-based routing and + // body-based policy authorize the same tool, prompt, or resource. + if decode_name_header(value)? != *name { + return Err(invalid); + } + } + Ok(()) +} + +fn is_plain_header_value(value: &str) -> bool { + value + .bytes() + .all(|byte| byte == b'\t' || (b' '..=b'~').contains(&byte)) +} + +fn decode_name_header(value: &str) -> Result { + if let Some(encoded) = value + .strip_prefix("=?base64?") + .and_then(|value| value.strip_suffix("?=")) + { + let bytes = STANDARD + .decode(encoded) + .map_err(|_| McpProtocolVersionError::InvalidRequestMetadata)?; + return String::from_utf8(bytes) + .map_err(|_| McpProtocolVersionError::InvalidRequestMetadata); + } + if !is_plain_header_value(value) { + return Err(McpProtocolVersionError::InvalidRequestMetadata); + } + Ok(value.to_string()) +} + fn is_standalone_initialize(info: &JsonRpcRequestInfo) -> bool { !info.is_batch && !info.has_response @@ -111,6 +236,13 @@ fn is_standalone_initialize(info: &JsonRpcRequestInfo) -> bool { fn request_protocol_version_header( raw_header: &[u8], ) -> Result, McpProtocolVersionError> { + request_header_value(raw_header, MCP_PROTOCOL_VERSION_HEADER) +} + +fn request_header_value<'a>( + raw_header: &'a [u8], + header_name: &str, +) -> Result, McpProtocolVersionError> { let header_end = raw_header .windows(4) .position(|window| window == b"\r\n\r\n") @@ -118,12 +250,11 @@ fn request_protocol_version_header( + 4; let headers = std::str::from_utf8(&raw_header[..header_end]) .map_err(|_| McpProtocolVersionError::InvalidHeader)?; - // Forwarding removes Connection-nominated fields. A revision field must - // survive that cleanup. Use the forwarding parser's canonical field names - // so authorization and removal agree, including after middleware rebuilds. + // Forwarding removes Connection-nominated fields. Required MCP metadata + // must survive that cleanup, including after middleware rebuilds. let nominated = crate::l7::rest::connection_nominated_header_names(&raw_header[..header_end]) .map_err(|_| McpProtocolVersionError::InvalidHeader)?; - if nominated.contains(MCP_PROTOCOL_VERSION_HEADER) { + if nominated.contains(header_name) { return Err(McpProtocolVersionError::InvalidHeader); } let mut values = headers.split("\r\n").skip(1).filter_map(|line| { @@ -131,7 +262,7 @@ fn request_protocol_version_header( // HTTP field-value optional whitespace is only SP or HTAB. Using // Unicode whitespace trimming here would accept bytes that are part // of the protocol-version value rather than HTTP framing. - name.eq_ignore_ascii_case(MCP_PROTOCOL_VERSION_HEADER) + name.eq_ignore_ascii_case(header_name) .then_some(value.trim_matches([' ', '\t'])) }); let Some(value) = values.next() else { @@ -146,8 +277,9 @@ fn request_protocol_version_header( #[cfg(test)] mod tests { use super::*; - use crate::l7::jsonrpc::{JsonRpcInspectionMode, parse_jsonrpc_body}; + use crate::l7::jsonrpc::{JsonRpcInspectionMode, McpHttpRequestMetadata, parse_jsonrpc_body}; use crate::l7::provider::BodyLength; + use std::fmt::Write as _; fn request(method: &str, headers: &str) -> L7Request { L7Request { @@ -164,6 +296,21 @@ mod tests { parse_jsonrpc_body(body, JsonRpcInspectionMode::Mcp) } + fn modern_request_info(method: &str, name: Option<&str>) -> JsonRpcRequestInfo { + JsonRpcRequestInfo { + calls: Vec::new(), + is_batch: false, + receive_stream: false, + has_response: false, + mcp_revision: Some(McpProtocolVersion::V2026_07_28), + mcp_http_metadata: Some(McpHttpRequestMetadata { + method: method.to_string(), + name: name.map(str::to_string), + }), + error: None, + } + } + #[test] fn standalone_initialize_uses_body_negotiation_only() { let info = request_info( @@ -283,7 +430,7 @@ mod tests { let info = request_info(br#"{"jsonrpc":"2.0","id":1,"method":"tools/list"}"#); for value in [ - "2026-07-28", + "2099-01-01", "2025-11-25, 2025-03-26", "2025-11-25x", "\u{00a0}2025-11-25", @@ -315,4 +462,189 @@ mod tests { ); assert_eq!(error.http_status(), "403 Forbidden"); } + + #[test] + fn sessionless_revision_requires_explicit_policy_opt_in() { + let info = modern_request_info("server/discover", None); + let request = request("POST", "MCP-Protocol-Version: 2026-07-28\r\n"); + + assert_eq!( + select_request_protocol_version(&request, &info, &[McpProtocolVersion::V2025_11_25],), + Err(McpProtocolVersionError::NotAllowed( + McpProtocolVersion::V2026_07_28 + )) + ); + assert_eq!( + select_request_protocol_version(&request, &info, &[McpProtocolVersion::V2026_07_28],), + Ok(McpRequestProtocolVersion::Selected( + McpProtocolVersion::V2026_07_28 + )) + ); + } + + #[test] + fn initialize_cannot_exempt_sessionless_requests_from_revision_selection() { + let info = request_info( + br#"{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"test","version":"1"}}}"#, + ); + assert_eq!( + select_request_protocol_version( + &request("POST", "MCP-Protocol-Version: 2026-07-28\r\n"), + &info, + &[McpProtocolVersion::V2026_07_28], + ), + Ok(McpRequestProtocolVersion::Selected( + McpProtocolVersion::V2026_07_28 + )) + ); + assert_eq!( + select_request_protocol_version( + &request("POST", ""), + &info, + &[McpProtocolVersion::V2026_07_28], + ), + Err(McpProtocolVersionError::NotAllowed( + McpProtocolVersion::V2025_03_26 + )) + ); + assert_eq!( + select_request_protocol_version( + &request("POST", "MCP-Protocol-Version: \r\n"), + &info, + &[McpProtocolVersion::V2025_11_25], + ), + Err(McpProtocolVersionError::InvalidHeader) + ); + } + + #[test] + fn sessionless_revision_accepts_only_post() { + let info = modern_request_info("subscriptions/listen", None); + for method in ["GET", "DELETE", "PUT", "post"] { + let error = select_request_protocol_version( + &request(method, "MCP-Protocol-Version: 2026-07-28\r\n"), + &info, + &[McpProtocolVersion::V2026_07_28], + ) + .expect_err("sessionless messages require POST"); + assert_eq!(error, McpProtocolVersionError::MethodNotAllowed); + assert_eq!(error.http_status(), "405 Method Not Allowed"); + } + } + + #[test] + fn sessionless_standard_headers_match_each_request_shape() { + for (method, name) in [ + ("server/discover", None), + ("subscriptions/listen", None), + ("example/extension", None), + ("tools/call", Some("get_weather")), + ("prompts/get", Some("summarize")), + ("resources/read", Some("file:///documents/readme.txt")), + ] { + let mut headers = + format!("mCp-PrOtOcOl-VeRsIoN:\t2026-07-28 \r\nMcP-MeThOd: {method}\t\r\n"); + if let Some(name) = name { + write!(headers, "mCp-NaMe: {name}\r\n").unwrap(); + } + assert_eq!( + validate_request_metadata( + &request("POST", &headers), + &modern_request_info(method, name), + ), + Ok(()), + "standard metadata for {method} must match its body fields" + ); + } + } + + #[test] + fn sessionless_metadata_requires_unambiguous_matching_headers() { + let info = modern_request_info("tools/call", Some("get_weather")); + let valid = "MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: tools/call\r\nMcp-Name: get_weather\r\n"; + for headers in [ + valid.replace("MCP-Protocol-Version: 2026-07-28\r\n", ""), + valid.replace("2026-07-28", "2025-11-25"), + valid.replace("Mcp-Method: tools/call\r\n", ""), + valid.replace("tools/call", "Tools/Call"), + valid.replace("Mcp-Name: get_weather\r\n", ""), + valid.replace("get_weather", ""), + valid.replace("get_weather", "Get_Weather"), + format!("{valid}mcp-protocol-version: 2026-07-28\r\n"), + format!("{valid}mcp-method: tools/call\r\n"), + format!("{valid}mcp-name: get_weather\r\n"), + format!("{valid}Connection: keep-alive, MCP-Protocol-Version\r\n"), + format!("{valid}Connection: keep-alive, Mcp-Method\r\n"), + format!("{valid}Connection: keep-alive, Mcp-Name\r\n"), + ] { + assert_eq!( + validate_request_metadata(&request("POST", &headers), &info), + Err(McpProtocolVersionError::InvalidRequestMetadata), + "invalid headers: {headers:?}" + ); + } + } + + #[test] + fn sessionless_names_decode_utf8_base64_before_matching() { + for name in ["天気", " padded ", "=?base64?literal?=", "line1\nline2"] { + let encoded = STANDARD.encode(name); + let headers = format!( + "MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: prompts/get\r\nMcp-Name: =?base64?{encoded}?=\r\n" + ); + assert_eq!( + validate_request_metadata( + &request("POST", &headers), + &modern_request_info("prompts/get", Some(name)), + ), + Ok(()) + ); + } + } + + #[test] + fn sessionless_names_reject_invalid_encoding_and_unsafe_plain_values() { + for (name, value) in [ + ("天気", "天気"), + (" padded ", " padded "), + ("=?base64?literal?=", "=?base64?literal?="), + ("weather", "=?base64?%%%?="), + ("weather", "=?base64?/w==?="), + ("weather", "=?BASE64?d2VhdGhlcg==?="), + ("bad\u{7f}name", "bad\u{7f}name"), + ] { + let headers = format!( + "MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: prompts/get\r\nMcp-Name: {value}\r\n" + ); + assert_eq!( + validate_request_metadata( + &request("POST", &headers), + &modern_request_info("prompts/get", Some(name)), + ), + Err(McpProtocolVersionError::InvalidRequestMetadata) + ); + } + } + + #[test] + fn sessionless_mirrors_leave_unknown_headers_and_removed_session_fields_alone() { + let info = modern_request_info("tools/call", Some("get_weather")); + let headers = "MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: tools/call\r\nMcp-Name: get_weather\r\nMcp-Param-Region: west\r\nMcp-Session-Id: unused\r\nLast-Event-ID: unused\r\n"; + assert_eq!( + validate_request_metadata(&request("POST", headers), &info), + Ok(()) + ); + } + + #[test] + fn mirrored_headers_are_not_required_for_legacy_or_extension_notifications() { + let request = request("POST", ""); + let mut info = modern_request_info("example/notification", None); + info.mcp_http_metadata = None; + assert_eq!(validate_request_metadata(&request, &info), Ok(())); + + let mut info = modern_request_info("tools/call", Some("weather")); + info.mcp_revision = Some(McpProtocolVersion::V2025_11_25); + assert_eq!(validate_request_metadata(&request, &info), Ok(())); + } } diff --git a/crates/openshell-supervisor-network/src/l7/middleware.rs b/crates/openshell-supervisor-network/src/l7/middleware.rs index 6ea9132902..a55ecdba95 100644 --- a/crates/openshell-supervisor-network/src/l7/middleware.rs +++ b/crates/openshell-supervisor-network/src/l7/middleware.rs @@ -1213,6 +1213,7 @@ mod tests { product_version: "0".into(), proxy_ip: [127, 0, 0, 1].into(), proxy_port: 3128, + origin: openshell_ocsf::EventOrigin::Supervisor, }; let eval = L7EvalContext { diff --git a/crates/openshell-supervisor-network/src/l7/relay.rs b/crates/openshell-supervisor-network/src/l7/relay.rs index 574ad3be8f..d0351e2153 100644 --- a/crates/openshell-supervisor-network/src/l7/relay.rs +++ b/crates/openshell-supervisor-network/src/l7/relay.rs @@ -183,79 +183,172 @@ where Ok(()) } -/// Enforce MCP request-version policy and emit a transport or policy rejection. -/// Non-MCP adapters share this entry point without changing their behavior. -/// Record endpoint version-policy denials before client response delivery. +/// Return the selected revision's inspection for policy evaluation, or emit a +/// rejection and return `None`. Non-MCP adapters retain their original inspection. +/// Record local rejections before client response delivery can fail. pub(crate) async fn enforce_mcp_protocol_version( config: &L7EndpointConfig, request: &crate::l7::provider::L7Request, - info: &crate::l7::jsonrpc::JsonRpcRequestInfo, + mut info: crate::l7::jsonrpc::JsonRpcRequestInfo, client: &mut W, ctx: &L7EvalContext, redacted_target: &str, observer: Option<&EndpointObserver>, -) -> Result +) -> Result> where W: AsyncWrite + Unpin, { if config.protocol != L7Protocol::Mcp { - return Ok(true); + return Ok(Some(info)); } - match crate::l7::mcp::select_request_protocol_version(request, info, &config.mcp_versions) { - Ok(crate::l7::mcp::McpRequestProtocolVersion::Initialization) => Ok(true), + match crate::l7::mcp::select_request_protocol_version(request, &info, &config.mcp_versions) { + Ok(crate::l7::mcp::McpRequestProtocolVersion::Initialization) => Ok(Some(info)), Ok(crate::l7::mcp::McpRequestProtocolVersion::Selected(version)) => { debug!(mcp_protocol_version = %version, "Selected MCP request protocol version"); - Ok(true) + info = crate::l7::jsonrpc::inspect_buffered_jsonrpc_http_request( + request, + crate::l7::jsonrpc::JsonRpcInspectionOptions::mcp_selected( + version, + config.mcp_strict_tool_names, + ), + )?; + // A bodyless receive stream has nothing for the JSON-RPC parser to + // classify, but later middleware re-evaluation must still retain + // the exact transport-selected revision. + info.mcp_revision = Some(version); + + if let Some(error) = info.error.as_ref() { + if let Some(observer) = observer { + // A selected revision's schema rejection is a local denial, + // even when the caller disconnects before receiving it. + observer.observe(EndpointResult::PolicyDenied); + } + let reason = format!( + "{error}; {}", + crate::l7::mcp::selected_revision_context(request, version) + ); + let summary = l7_protocol_log_summary(None, Some(&info)); + ocsf_emit!(build_l7_request_event( + ctx, + &request.action, + redacted_target, + "deny", + "l7-mcp", + &reason, + summary.as_deref(), + )); + emit_activity(ctx, true, "l7_parse_rejection"); + let body = serde_json::json!({ + "error": "invalid_mcp_request", + "detail": reason, + "policy": ctx.policy_name, + "layer": "l7", + "protocol": "mcp", + "method": request.action, + "path": redacted_target, + }); + crate::l7::rest::send_json_response( + &ctx.policy_name, + body, + client, + "400 Bad Request", + ) + .await?; + return Ok(None); + } + + if let Err(error) = crate::l7::mcp::validate_request_metadata(request, &info) { + reject_mcp_protocol_version( + error, + request, + &info, + client, + ctx, + redacted_target, + observer, + ) + .await?; + return Ok(None); + } + + Ok(Some(info)) } Err(error) => { - if let Some(observer) = observer { - // Every explicit version-gate rejection is local to this MCP - // endpoint. Record it before delivery can fail; an invalid client - // version is not an upstream transport failure. - observer.observe(EndpointResult::PolicyDenied); - } - let reason = error.to_string(); - let summary = l7_protocol_log_summary(None, Some(info)); - ocsf_emit!(build_l7_request_event( + reject_mcp_protocol_version( + error, + request, + &info, + client, ctx, - &request.action, redacted_target, - "deny", - "l7-mcp", - &reason, - summary.as_deref(), - )); - let deny_group = match error { - crate::l7::mcp::McpProtocolVersionError::NotAllowed(_) => "l7_policy", - crate::l7::mcp::McpProtocolVersionError::InvalidHeader - | crate::l7::mcp::McpProtocolVersionError::UnsupportedHeaderValue => { - "l7_parse_rejection" - } - }; - emit_activity(ctx, true, deny_group); - - let body = serde_json::json!({ - "error": error.response_code(), - "detail": reason, - "policy": ctx.policy_name, - "layer": "l7", - "protocol": "mcp", - "method": request.action, - "path": redacted_target, - }); - crate::l7::rest::send_json_response( - &ctx.policy_name, - body, - client, - error.http_status(), + observer, ) .await?; - Ok(false) + Ok(None) } } } +/// Emit the same transport or policy rejection at initial and final inspection. +async fn reject_mcp_protocol_version( + error: crate::l7::mcp::McpProtocolVersionError, + request: &crate::l7::provider::L7Request, + info: &crate::l7::jsonrpc::JsonRpcRequestInfo, + client: &mut W, + ctx: &L7EvalContext, + redacted_target: &str, + observer: Option<&EndpointObserver>, +) -> Result<()> +where + W: AsyncWrite + Unpin, +{ + if let Some(observer) = observer { + // Version, metadata, and method rejections are local to this endpoint. + // Record the denial before client delivery can fail. + observer.observe(EndpointResult::PolicyDenied); + } + let reason = error.rejection_detail(request); + let summary = l7_protocol_log_summary(None, Some(info)); + ocsf_emit!(build_l7_request_event( + ctx, + &request.action, + redacted_target, + "deny", + "l7-mcp", + &reason, + summary.as_deref(), + )); + let deny_group = match error { + crate::l7::mcp::McpProtocolVersionError::NotAllowed(_) => "l7_policy", + crate::l7::mcp::McpProtocolVersionError::InvalidHeader + | crate::l7::mcp::McpProtocolVersionError::UnsupportedHeaderValue + | crate::l7::mcp::McpProtocolVersionError::InvalidRequestMetadata + | crate::l7::mcp::McpProtocolVersionError::MethodNotAllowed => "l7_parse_rejection", + }; + emit_activity(ctx, true, deny_group); + let body = serde_json::json!({ + "error": error.response_code(), + "detail": reason, + "policy": ctx.policy_name, + "layer": "l7", + "protocol": "mcp", + "method": request.action, + "path": redacted_target, + }); + let allowed_methods = + (error == crate::l7::mcp::McpProtocolVersionError::MethodNotAllowed).then_some("POST"); + crate::l7::rest::send_json_response_with_allow( + &ctx.policy_name, + body, + client, + error.http_status(), + allowed_methods, + ) + .await?; + Ok(()) +} + /// Reinspect the buffered outgoing MCP request after request transformations. /// The forwarding adapter must call this before any upstream request write. pub(crate) async fn enforce_final_mcp_protocol_version( @@ -276,16 +369,17 @@ where request, crate::l7::jsonrpc::JsonRpcInspectionOptions::for_config(config), )?; - enforce_mcp_protocol_version( + Ok(enforce_mcp_protocol_version( config, request, - &info, + info, client, ctx, redacted_target, observer, ) - .await + .await? + .is_some()) } fn build_request_authority_mismatch_event( @@ -707,34 +801,45 @@ fn engine_type_for_protocol(protocol: L7Protocol) -> &'static str { } } -async fn deny_h2c_upgrade_if_requested( +/// Refuses an upgrade the endpoint cannot inspect and reports whether the +/// request was answered. +/// +/// Every L7 request loop calls this before the L7 policy decision. A refusal +/// records a policy denial for the endpoint, emits a parse-rejection event, +/// and answers `403` regardless of enforcement mode; see +/// `unsupported_upgrade_detail` for which upgrades each protocol refuses. The +/// response carries the `unsupported_l7_protocol` error rather than the +/// policy-denial body, because no policy rule can allow the request. +async fn deny_unsupported_upgrade_if_requested( req: &crate::l7::provider::L7Request, config: &L7EndpointConfig, ctx: &L7EvalContext, + observer: Option<&EndpointObserver>, client: &mut C, ) -> Result where C: AsyncRead + AsyncWrite + Unpin + Send, { - if !crate::l7::rest::request_is_h2c_upgrade(&req.raw_header) { + let Some(detail) = + crate::l7::rest::unsupported_upgrade_detail(&req.raw_header, config.protocol) + else { return Ok(false); - } + }; - emit_parse_rejection( - ctx, - crate::l7::rest::UNSUPPORTED_H2C_UPGRADE_DETAIL, - engine_type_for_protocol(config.protocol), - ); - crate::l7::rest::RestProvider::default() - .deny_with_redacted_target( - req, - &ctx.policy_name, - crate::l7::rest::UNSUPPORTED_H2C_UPGRADE_DETAIL, - client, - None, - Some(crate::l7::rest::DenyResponseContext::from_l7_context(ctx)), - ) - .await?; + if let Some(observer) = observer { + observer.observe(EndpointResult::PolicyDenied); + } + emit_parse_rejection(ctx, detail, engine_type_for_protocol(config.protocol)); + crate::l7::rest::send_json_response( + &ctx.policy_name, + serde_json::json!({ + "error": "unsupported_l7_protocol", + "detail": detail, + }), + client, + "403 Forbidden", + ) + .await?; Ok(true) } @@ -934,7 +1039,9 @@ where .await?; return Ok(()); } - if deny_h2c_upgrade_if_requested(&req, config, ctx, client).await? { + if deny_unsupported_upgrade_if_requested(&req, config, ctx, observer.as_ref(), client) + .await? + { return Ok(()); } @@ -966,7 +1073,7 @@ where } else { None }; - let jsonrpc_info = if config.protocol.is_jsonrpc_family() { + let mut jsonrpc_info = if config.protocol.is_jsonrpc_family() { if crate::l7::jsonrpc::jsonrpc_receive_stream_request(&req) { Some(crate::l7::jsonrpc::JsonRpcRequestInfo::receive_stream()) } else { @@ -1019,15 +1126,8 @@ where } }; - let request_info = L7RequestInfo { - action: req.action.clone(), - target: redacted_target.clone(), - query_params: req.query_params.clone(), - graphql: graphql_info.clone(), - jsonrpc: jsonrpc_info.clone(), - }; - if let Some(info) = jsonrpc_info.as_ref() - && !enforce_mcp_protocol_version( + if let Some(info) = jsonrpc_info.take() { + let Some(inspected) = enforce_mcp_protocol_version( config, &req, info, @@ -1037,9 +1137,18 @@ where observer.as_ref(), ) .await? - { - return Ok(()); + else { + return Ok(()); + }; + jsonrpc_info = Some(inspected); } + let request_info = L7RequestInfo { + action: req.action.clone(), + target: redacted_target.clone(), + query_params: req.query_params.clone(), + graphql: graphql_info.clone(), + jsonrpc: jsonrpc_info.clone(), + }; let websocket_request = crate::l7::rest::request_is_websocket_upgrade(&req.raw_header); if config.protocol == L7Protocol::Websocket && !websocket_request { crate::l7::rest::RestProvider::default() @@ -1148,18 +1257,6 @@ where return Ok(()); } }; - if !enforce_final_mcp_protocol_version( - config, - &req, - client, - ctx, - &redacted_target, - observer.as_ref(), - ) - .await? - { - return Ok(()); - } let scoped_ctx = scoped_context_for_request(ctx, &req); let ctx = scoped_ctx.as_ref().unwrap_or(ctx); // Credential scoping can acquire a newer revision after middleware. @@ -1212,6 +1309,13 @@ where crate::l7::rest::RelayRequestOptions { resolver: ctx.secret_resolver.as_deref(), body_classifier: ctx.body_classifier.as_deref(), + mcp_request_validation: (config.protocol == L7Protocol::Mcp).then_some( + crate::l7::rest::McpRequestValidation { + config, + ctx, + redacted_target: &redacted_target, + }, + ), credential_generation: credential_generation_guard(ctx), generation_guard: Some(engine.generation_guard()), websocket_extensions: websocket_extension_mode( @@ -1283,6 +1387,27 @@ where websocket_permessage_deflate, websocket_subprotocol, } => { + // Protocols whose rules apply to individual HTTP requests + // never upgrade (see `upgrade_refusal_for_protocol`). No + // current path forwards upgrade headers for them: the + // request-side refusal rejects them, and request + // middleware cannot add upgrade or connection headers. If + // a later change lets such a request reach an upstream + // that answers `101`, close instead of relaying frames + // that no rule would inspect. + if crate::l7::rest::upgrade_refusal_for_protocol(config.protocol).is_some() { + warn!( + host = %ctx.host, + port = ctx.port, + "closing per-request L7 connection after unexpected protocol upgrade" + ); + if let Some(session) = middleware_session.take() { + session + .end(openshell_core::proto::MiddlewareSessionEndReason::ProtocolError) + .await; + } + return Ok(()); + } let mut options = upgrade_options( config, ctx, @@ -1751,7 +1876,7 @@ where reject_request_authority_mismatch(client, ctx, &req.action).await?; return Ok(()); } - if deny_h2c_upgrade_if_requested(&req, config, ctx, client).await? { + if deny_unsupported_upgrade_if_requested(&req, config, ctx, None, client).await? { return Ok(()); } @@ -1968,6 +2093,7 @@ where crate::l7::rest::RelayRequestOptions { resolver: ctx.secret_resolver.as_deref(), body_classifier: ctx.body_classifier.as_deref(), + mcp_request_validation: None, credential_generation: credential_generation_guard(ctx), generation_guard: Some(engine.generation_guard()), websocket_extensions: websocket_extension_mode( @@ -2218,6 +2344,11 @@ where reject_request_authority_mismatch(client, ctx, &req.action).await?; return Ok(()); } + if deny_unsupported_upgrade_if_requested(&req, config, ctx, observer.as_ref(), client) + .await? + { + return Ok(()); + } if close_if_stale(engine.generation_guard(), ctx) { return Ok(()); } @@ -2233,26 +2364,27 @@ where } }; - let request_info = L7RequestInfo { - action: req.action.clone(), - target: redacted_target.clone(), - query_params: req.query_params.clone(), - graphql: None, - jsonrpc: Some(jsonrpc_info.clone()), - }; - if !enforce_mcp_protocol_version( + let Some(jsonrpc_info) = enforce_mcp_protocol_version( config, &req, - &jsonrpc_info, + jsonrpc_info, client, ctx, &redacted_target, observer.as_ref(), ) .await? - { + else { return Ok(()); - } + }; + + let request_info = L7RequestInfo { + action: req.action.clone(), + target: redacted_target.clone(), + query_params: req.query_params.clone(), + graphql: None, + jsonrpc: Some(jsonrpc_info.clone()), + }; let hard_deny_reason = l7_request_hard_deny_reason(config.protocol, &request_info); let force_deny = hard_deny_reason.is_some(); @@ -2370,18 +2502,6 @@ where return Ok(()); } }; - if !enforce_final_mcp_protocol_version( - config, - &req, - client, - ctx, - &redacted_target, - observer.as_ref(), - ) - .await? - { - return Ok(()); - } let scoped_ctx = scoped_context_for_request(ctx, &req); let ctx = scoped_ctx.as_ref().unwrap_or(ctx); // The outgoing resolver and revision are one snapshot. Rebind using @@ -2394,11 +2514,8 @@ where ctx.provider_credential_revision, Some(engine.generation_guard()), ); - // Future MCP response/SSE introspection or rewrite would hook here - // before returning upstream bytes. The current policy schema has no - // trusted-annotations or version-profile field, so MCP responses and - // SSE streams are relayed unchanged; see McpOptions in - // proto/sandbox.proto for planned policy extensions. + // Policy inspects client request bodies. Response bodies and SSE + // messages remain opaque and are relayed without MCP inspection. let Some(outcome) = relay_http_request_with_credential_rejection_observed( &req, client, @@ -2406,6 +2523,13 @@ where crate::l7::rest::RelayRequestOptions { resolver: ctx.secret_resolver.as_deref(), body_classifier: ctx.body_classifier.as_deref(), + mcp_request_validation: (config.protocol == L7Protocol::Mcp).then_some( + crate::l7::rest::McpRequestValidation { + config, + ctx, + redacted_target: &redacted_target, + }, + ), credential_generation: credential_generation_guard(ctx), generation_guard: Some(engine.generation_guard()), ..Default::default() @@ -2470,23 +2594,56 @@ where C: AsyncRead + AsyncWrite + Unpin + Send, U: AsyncRead + AsyncWrite + Unpin + Send, { + let provider = + crate::l7::rest::RestProvider::with_options(crate::l7::path::CanonicalizeOptions { + allow_encoded_slash: config.allow_encoded_slash, + ..Default::default() + }); + loop { if close_if_stale(engine.generation_guard(), ctx) { return Ok(()); } - let parsed = match crate::l7::graphql::parse_graphql_http_request( + // Validate the head, including body framing, before deciding whether + // this endpoint can inspect the requested protocol. Upgrade refusal + // must not wait for a body or depend on its inspection size limit. + let mut req = match provider.parse_request(client).await { + Ok(Some(req)) => req, + Ok(None) => return Ok(()), + Err(e) => { + if is_benign_connection_error(&e) { + debug!( + host = %ctx.host, + port = ctx.port, + error = %e, + "GraphQL L7 connection closed" + ); + } else { + let detail = + parse_rejection_detail(&e.to_string(), ParseRejectionMode::L7Endpoint); + emit_parse_rejection(ctx, &detail, "l7-graphql"); + } + return Ok(()); + } + }; + + if !request_authority_matches_endpoint(&req, ctx) { + reject_request_authority_mismatch(client, ctx, &req.action).await?; + return Ok(()); + } + if deny_unsupported_upgrade_if_requested(&req, config, ctx, None, client).await? { + return Ok(()); + } + + let graphql_info = match crate::l7::graphql::inspect_graphql_request( client, + &mut req, config.graphql_max_body_bytes, - crate::l7::path::CanonicalizeOptions { - allow_encoded_slash: config.allow_encoded_slash, - ..Default::default() - }, ) .await { - Ok(Some(parsed)) => parsed, - Ok(None) => return Ok(()), + Ok(info) => info, Err(e) => { if is_benign_connection_error(&e) { debug!( @@ -2504,15 +2661,13 @@ where } }; - let req = parsed.request; - let graphql_info = parsed.info; + // Inspection appends the body to raw_header. An HTTP/1.0 request + // without Host must still be denied if that body contains reserved + // credential markers, so repeat the authority check on the full request. if !request_authority_matches_endpoint(&req, ctx) { reject_request_authority_mismatch(client, ctx, &req.action).await?; return Ok(()); } - if deny_h2c_upgrade_if_requested(&req, config, ctx, client).await? { - return Ok(()); - } if close_if_stale(engine.generation_guard(), ctx) { return Ok(()); @@ -2687,25 +2842,20 @@ where ); return Ok(()); } - RelayOutcome::Upgraded { - overflow, - websocket_permessage_deflate, - .. - } => { - let options = UpgradeRelayOptions { - assembly_budget: Some( - crate::l7::websocket::WebSocketAssemblyBudget::default(), - ), - websocket: WebSocketUpgradeBehavior { - permessage_deflate: websocket_permessage_deflate, - ..Default::default() - }, - ..Default::default() - }; - return handle_upgrade( - client, upstream, overflow, &ctx.host, ctx.port, options, - ) - .await; + RelayOutcome::Upgraded { .. } => { + // GraphQL rules apply to individual HTTP requests. No + // current path forwards upgrade headers here: the + // request-side refusal rejects them, and request + // middleware cannot add upgrade or connection headers. If + // a later change lets such a request reach an upstream + // that answers `101`, close instead of relaying frames + // that no GraphQL rule would inspect. + warn!( + host = %ctx.host, + port = ctx.port, + "closing GraphQL connection after unexpected protocol upgrade" + ); + return Ok(()); } } } else { @@ -2909,6 +3059,8 @@ fn evaluate_jsonrpc_l7_request_for_log( is_batch: true, receive_stream: false, has_response: false, + mcp_revision: jsonrpc.mcp_revision, + mcp_http_metadata: None, error: None, }, }); @@ -2932,6 +3084,8 @@ fn jsonrpc_request_for_call( is_batch: false, receive_stream: false, has_response: false, + mcp_revision: request.jsonrpc.as_ref().and_then(|info| info.mcp_revision), + mcp_http_metadata: None, error: None, }); item_request @@ -2967,10 +3121,20 @@ fn reevaluate_transformed_body( // unimplemented SQL relay. L7Protocol::Rest | L7Protocol::Websocket | L7Protocol::Sql => return Ok(None), L7Protocol::JsonRpc | L7Protocol::Mcp => { - let info = crate::l7::jsonrpc::parse_jsonrpc_body_with_options( - body, - crate::l7::jsonrpc::JsonRpcInspectionOptions::for_config(config), - ); + let mut inspection_options = + crate::l7::jsonrpc::JsonRpcInspectionOptions::for_config(config); + if let Some(revision) = request_info + .jsonrpc + .as_ref() + .and_then(|info| info.mcp_revision) + { + // Inspect the replacement under the revision authorized on entry. + // The final forwarding check validates the resulting header-selected + // profile and body/header mirrors after any header mutations. + inspection_options = inspection_options.with_mcp_revision(revision); + } + let info = + crate::l7::jsonrpc::parse_jsonrpc_body_with_options(body, inspection_options); let mut transformed_info = request_info.clone(); transformed_info.jsonrpc = Some(info); (jsonrpc_engine_type(config.protocol), transformed_info) @@ -3071,6 +3235,8 @@ fn jsonrpc_policy_input(info: &crate::l7::jsonrpc::JsonRpcRequestInfo) -> serde_ "method": call.map(|call| call.method.as_str()), "params": call.map(|call| &call.params), "tool": call.and_then(|call| call.tool.as_deref()), + "mcp_method_classification": call + .and_then(|call| call.mcp_classification), "receive_stream": info.receive_stream, "has_response": info.has_response, // Rust keeps the inspection failure kind typed. Rego's stable boundary is @@ -3380,6 +3546,7 @@ mod tests { use openshell_core::proto::{StaticCredentialBinding, StaticCredentialEndpointBinding}; use openshell_core::provider_credentials::ProviderCredentialState; use std::collections::HashMap as TestHashMap; + use std::fmt::Write as _; use std::path::PathBuf; use tokio::io::{AsyncRead, AsyncReadExt, AsyncWriteExt}; @@ -3639,6 +3806,7 @@ mod tests { product_version: "0".into(), proxy_ip: [127, 0, 0, 1].into(), proxy_port: 3128, + origin: openshell_ocsf::EventOrigin::Supervisor, }; let eval = L7EvalContext { @@ -4610,7 +4778,8 @@ network_policies: } fn mcp_test_relay_context() -> (L7EndpointConfig, TunnelPolicyEngine, L7EvalContext) { - let data = r" + mcp_relay_context_from_data( + r" network_policies: mcp_api: name: mcp_api @@ -4625,8 +4794,48 @@ network_policies: method: initialize binaries: - { path: /usr/bin/python3 } -"; - let engine = OpaEngine::from_strings(TEST_POLICY, data).unwrap(); +", + ) + } + + fn mcp_sessionless_test_relay_context() -> (L7EndpointConfig, TunnelPolicyEngine, L7EvalContext) + { + mcp_relay_context_from_data( + r#" +network_policies: + mcp_api: + name: mcp_api + endpoints: + - host: mcp.example.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: enforce + mcp: + versions: ["2026-07-28"] + allow_all_known_mcp_methods: true + rules: + - allow: {} + - allow: + method: vendor/inspect + deny_rules: + - method: tools/call + tool: blocked + binaries: + - { path: /usr/bin/python3 } +"#, + ) + } + + fn mcp_relay_context_from_data( + data: &str, + ) -> (L7EndpointConfig, TunnelPolicyEngine, L7EvalContext) { + mcp_relay_context_from_engine(OpaEngine::from_strings(TEST_POLICY, data).unwrap()) + } + + fn mcp_relay_context_from_engine( + engine: OpaEngine, + ) -> (L7EndpointConfig, TunnelPolicyEngine, L7EvalContext) { let input = NetworkInput { host: "mcp.example.test".into(), port: 8000, @@ -7649,7 +7858,7 @@ network_policies: } #[test] - fn jsonrpc_inspection_error_opa_projection_remains_string_or_null() { + fn jsonrpc_inspection_opa_projection_uses_stable_values() { let invalid_json = crate::l7::jsonrpc::parse_jsonrpc_body( b"{", crate::l7::jsonrpc::JsonRpcInspectionMode::JsonRpc, @@ -7672,6 +7881,29 @@ network_policies: serde_json::json!("missing or non-string 'jsonrpc' field") ); assert!(jsonrpc_policy_input(&accepted)["error"].is_null()); + + let available = crate::l7::jsonrpc::parse_jsonrpc_body_with_options( + br#"{"jsonrpc":"2.0","id":1,"method":"tools/list"}"#, + crate::l7::jsonrpc::JsonRpcInspectionOptions::mcp_selected( + openshell_core::mcp::McpProtocolVersion::V2025_11_25, + true, + ), + ); + let extension = crate::l7::jsonrpc::parse_jsonrpc_body_with_options( + br#"{"jsonrpc":"2.0","id":1,"method":"tools/vendor"}"#, + crate::l7::jsonrpc::JsonRpcInspectionOptions::mcp_selected( + openshell_core::mcp::McpProtocolVersion::V2025_11_25, + true, + ), + ); + assert_eq!( + jsonrpc_policy_input(&available)["mcp_method_classification"], + serde_json::json!("available") + ); + assert_eq!( + jsonrpc_policy_input(&extension)["mcp_method_classification"], + serde_json::json!("extension") + ); } #[test] @@ -7901,9 +8133,12 @@ network_policies: target: "/mcp".into(), query_params: std::collections::HashMap::new(), graphql: None, - jsonrpc: Some(crate::l7::jsonrpc::parse_jsonrpc_body( + jsonrpc: Some(crate::l7::jsonrpc::parse_jsonrpc_body_with_options( br#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"read_status","arguments":{}}}"#, - crate::l7::jsonrpc::JsonRpcInspectionMode::Mcp, + crate::l7::jsonrpc::JsonRpcInspectionOptions::mcp_selected( + openshell_core::mcp::McpProtocolVersion::V2025_11_25, + true, + ), )), }; @@ -7921,9 +8156,12 @@ network_policies: assert!(allowed_message.contains("rule_methods=tools/call")); assert!(allowed_message.contains("tools=read_status")); - request.jsonrpc = Some(crate::l7::jsonrpc::parse_jsonrpc_body( + request.jsonrpc = Some(crate::l7::jsonrpc::parse_jsonrpc_body_with_options( br#"{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"delete_resource","arguments":{"scope":"workspace/main"}}}"#, - crate::l7::jsonrpc::JsonRpcInspectionMode::Mcp, + crate::l7::jsonrpc::JsonRpcInspectionOptions::mcp_selected( + openshell_core::mcp::McpProtocolVersion::V2025_11_25, + true, + ), )); let parsed = request.jsonrpc.as_ref().expect("parsed MCP request"); assert!( @@ -9028,124 +9266,411 @@ network_policies: let _ = tokio::time::timeout(std::time::Duration::from_secs(1), relay).await; } - fn masked_text_frame(payload: &[u8]) -> Vec { - let mask = [0x11, 0x22, 0x33, 0x44]; - assert!( - payload.len() <= 125, - "test helper only supports small frames" - ); - let payload_len = u8::try_from(payload.len()).expect("small frame length"); - let mut frame = vec![0x81, 0x80 | payload_len]; - frame.extend_from_slice(&mask); - frame.extend( - payload - .iter() - .enumerate() - .map(|(idx, byte)| byte ^ mask[idx % 4]), - ); - frame - } - - async fn read_http_headers(reader: &mut R) -> Vec { - let mut bytes = Vec::new(); - let mut byte = [0u8; 1]; - loop { - reader.read_exact(&mut byte).await.unwrap(); - bytes.push(byte[0]); - if bytes.ends_with(b"\r\n\r\n") { - return bytes; - } - } - } + /// A `tools/call` that no JSON-RPC-family fixture below allows. + const UNALLOWED_TOOL_CALL: &[u8] = + br#"{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"delete_resource","arguments":{}}}"#; - async fn read_text_frame( - reader: &mut R, - ) -> std::io::Result<(bool, String)> { - let mut header = [0u8; 2]; - reader.read_exact(&mut header).await?; - assert_eq!(header[0] & 0x0f, 0x1, "expected text frame"); - let masked = header[1] & 0x80 != 0; - let payload_len = usize::from(header[1] & 0x7f); - assert!(payload_len <= 125, "test helper only supports small frames"); - let mut mask = [0u8; 4]; - if masked { - reader.read_exact(&mut mask).await?; - } - let mut payload = vec![0u8; payload_len]; - reader.read_exact(&mut payload).await?; - if masked { - for (idx, byte) in payload.iter_mut().enumerate() { - *byte ^= mask[idx % 4]; - } - } - Ok((masked, String::from_utf8(payload).expect("text payload"))) - } + /// A receive-stream GET that also asks to switch to WebSocket. + const JSONRPC_WEBSOCKET_UPGRADE_REQUEST: &[u8] = b"GET /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nAccept: text/event-stream\r\nMCP-Protocol-Version: 2025-11-25\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==\r\nSec-WebSocket-Version: 13\r\n\r\n"; - #[tokio::test] - async fn l7_relay_closes_keep_alive_tunnel_after_policy_generation_change() { - let initial_data = r#" + /// Builds two endpoints on one host and port so every request goes + /// through per-request route selection: a JSON-RPC-family endpoint at + /// `/mcp` that allows only `initialize`, and a REST endpoint at `/api/**`. + fn jsonrpc_and_rest_route_configs( + protocol: &str, + enforcement: &str, + ) -> (Vec, TunnelPolicyEngine, L7EvalContext) { + let data = format!( + r#" network_policies: - rest_api: - name: rest_api + shared_api: + name: shared_api endpoints: - - host: api.example.test - port: 8080 + - host: mcp.example.test + port: 8000 + path: "/mcp" + protocol: {protocol} + enforcement: {enforcement} + rules: + - allow: + method: initialize + - host: mcp.example.test + port: 8000 + path: "/api/**" protocol: rest enforcement: enforce rules: - allow: - method: POST - path: "/write" + method: GET + path: "/api/**" binaries: - - { path: /usr/bin/curl } -"#; - let reloaded_data = r#" + - {{ path: /usr/bin/python3 }} +"# + ); + two_endpoint_route_configs(&data, "mcp.example.test", "shared_api") + } + + /// Builds per-request route selection for a GraphQL endpoint at + /// `/graphql` that allows only `query { viewer }`, and a REST endpoint at + /// `/api/**`, on the same host and port. + fn graphql_and_rest_route_configs( + enforcement: &str, + ) -> (Vec, TunnelPolicyEngine, L7EvalContext) { + let data = format!( + r#" network_policies: - rest_api: - name: rest_api + shared_graphql: + name: shared_graphql endpoints: - - host: api.example.test - port: 8080 + - host: graphql.example.test + port: 8000 + path: "/graphql" + protocol: graphql + enforcement: {enforcement} + rules: + - allow: + operation_type: query + fields: [viewer] + - host: graphql.example.test + port: 8000 + path: "/api/**" protocol: rest enforcement: enforce rules: - allow: method: GET - path: "/write" + path: "/api/**" binaries: - - { path: /usr/bin/curl } -"#; - let engine = OpaEngine::from_strings(TEST_POLICY, initial_data).unwrap(); + - {{ path: /usr/bin/python3 }} +"# + ); + two_endpoint_route_configs(&data, "graphql.example.test", "shared_graphql") + } + + /// Loads `data` and returns the two L7 configs that share `host:8000` + /// for `/usr/bin/python3`, with a matching tunnel engine and context. + fn two_endpoint_route_configs( + data: &str, + host: &str, + policy_name: &str, + ) -> (Vec, TunnelPolicyEngine, L7EvalContext) { + let engine = OpaEngine::from_strings(TEST_POLICY, data).unwrap(); let input = NetworkInput { - host: "api.example.test".into(), - port: 8080, - binary_path: PathBuf::from("/usr/bin/curl"), + host: host.into(), + port: 8000, + binary_path: PathBuf::from("/usr/bin/python3"), binary_sha256: "unused".into(), ancestors: vec![], cmdline_paths: vec![], }; - let (endpoint_config, generation) = engine - .query_endpoint_config_with_generation(&input) + let (endpoint_configs, generation) = engine + .query_endpoint_configs_with_generation(&input) .unwrap(); - let config = crate::l7::parse_l7_config(&endpoint_config.unwrap()).unwrap(); + let configs: Vec = endpoint_configs + .iter() + .map(|config| crate::l7::parse_l7_config(config).unwrap()) + .collect(); + assert_eq!(configs.len(), 2, "both endpoints must share the route"); let tunnel_engine = engine.clone_engine_for_tunnel(generation).unwrap(); let ctx = L7EvalContext { - host: "api.example.test".into(), - port: 8080, - request_default_port: Some(8080), - policy_name: "rest_api".into(), - binary_path: "/usr/bin/curl".into(), + host: host.into(), + port: 8000, + request_default_port: Some(8000), + policy_name: policy_name.into(), + binary_path: "/usr/bin/python3".into(), ancestors: vec![], cmdline_paths: vec![], secret_resolver: None, ..Default::default() }; + (configs, tunnel_engine, ctx) + } + + /// Result of sending one upgrade request through a relay whose upstream + /// accepts every upgrade it receives. + struct UpgradeScenario { + /// The response head the client received, or empty if none arrived. + response: String, + /// The response body of a non-`101` response. + body: String, + /// Every byte the upstream received. + upstream_seen: Vec, + } + + /// Sends `request`, answers any forwarded upgrade with a valid `101`, and + /// after a `101` writes `frame` as a WebSocket text message. The upstream + /// never refuses, so a relay that forwards the upgrade and then copies + /// bytes delivers `frame` to it. + async fn run_upgrade_scenario(request: &[u8], frame: &[u8], relay: F) -> UpgradeScenario + where + F: FnOnce( + tokio::io::DuplexStream, + tokio::io::DuplexStream, + ) -> tokio::task::JoinHandle>, + { + let (mut app, relay_client) = tokio::io::duplex(8192); + let (relay_upstream, mut upstream) = tokio::io::duplex(8192); + let relay = relay(relay_client, relay_upstream); + let upstream_task = tokio::spawn(async move { + let mut seen = Vec::new(); + let mut buf = [0u8; 4096]; + let mut answered = false; + loop { + let read = tokio::time::timeout( + std::time::Duration::from_secs(2), + upstream.read(&mut buf), + ) + .await; + let Ok(Ok(n)) = read else { break }; + if n == 0 { + break; + } + seen.extend_from_slice(&buf[..n]); + if !answered && seen.windows(4).any(|w| w == b"\r\n\r\n") { + answered = true; + let _ = upstream + .write_all( + b"HTTP/1.1 101 Switching Protocols\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=\r\n\r\n", + ) + .await; + } + } + seen + }); + + app.write_all(request).await.unwrap(); + let mut response = Vec::new(); + let mut byte = [0u8; 1]; + while !response.ends_with(b"\r\n\r\n") { + let read = + tokio::time::timeout(std::time::Duration::from_secs(2), app.read(&mut byte)).await; + let Ok(Ok(1)) = read else { break }; + response.push(byte[0]); + } + let response = String::from_utf8_lossy(&response).into_owned(); + let mut body = Vec::new(); + if response.starts_with("HTTP/1.1 101") { + let _ = app.write_all(&masked_text_frame(frame)).await; + } else { + // Refusals close the connection after the body. + let _ = tokio::time::timeout( + std::time::Duration::from_secs(2), + app.read_to_end(&mut body), + ) + .await; + } + drop(app); + let _ = tokio::time::timeout(std::time::Duration::from_secs(2), relay).await; + let upstream_seen = upstream_task.await.unwrap(); + UpgradeScenario { + response, + body: String::from_utf8_lossy(&body).into_owned(), + upstream_seen, + } + } + + fn contains_bytes(haystack: &[u8], needle: &[u8]) -> bool { + haystack + .windows(needle.len()) + .any(|window| window == needle) + } + + fn assert_upgrade_denied_before_forwarding(scenario: &UpgradeScenario) { + assert_upgrade_refused_before_forwarding( + scenario, + UNALLOWED_TOOL_CALL, + crate::l7::rest::UNSUPPORTED_JSONRPC_UPGRADE_DETAIL, + ); + } + + /// Asserts that the relay answered the upgrade with the `detail` refusal + /// and that neither the request nor `frame` reached the upstream. + fn assert_upgrade_refused_before_forwarding( + scenario: &UpgradeScenario, + frame: &[u8], + detail: &str, + ) { + assert!( + !contains_bytes(&scenario.upstream_seen, &masked_text_frame(frame)), + "an uninspected frame reached the upstream" + ); + assert!( + scenario.upstream_seen.is_empty(), + "the upgrade request must not reach the upstream, got: {}", + String::from_utf8_lossy(&scenario.upstream_seen) + ); + assert!( + scenario.response.starts_with("HTTP/1.1 403"), + "expected a 403 denial, got: {}", + scenario.response + ); + assert!( + scenario.body.contains("\"unsupported_l7_protocol\"") && scenario.body.contains(detail), + "expected the upgrade refusal, got: {}", + scenario.body + ); + } + + #[tokio::test] + async fn mcp_websocket_upgrade_refusal_records_policy_denied() { + use openshell_core::endpoint_status::EndpointStatusCommand; + + for route_selected in [false, true] { + let (mut config, tunnel_engine, mut ctx) = mcp_test_relay_context(); + let mut receiver = install_mcp_test_observation(&mut config, &mut ctx).await; + let scenario = run_upgrade_scenario( + JSONRPC_WEBSOCKET_UPGRADE_REQUEST, + UNALLOWED_TOOL_CALL, + move |mut client, mut upstream| { + tokio::spawn(async move { + if route_selected { + relay_with_route_selection( + &[config], + tunnel_engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + } else { + relay_with_inspection( + &config, + tunnel_engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + } + }) + }, + ) + .await; + assert_upgrade_denied_before_forwarding(&scenario); + assert!( + matches!( + receiver.try_recv(), + Ok(EndpointStatusCommand::Observe { + result: EndpointResult::PolicyDenied, + .. + }) + ), + "route_selected={route_selected}: refusal must record a policy denial" + ); + assert!( + receiver.try_recv().is_err(), + "route_selected={route_selected}: one result per exchange" + ); + } + } + + #[tokio::test] + async fn route_selected_mcp_websocket_upgrade_is_denied_before_forwarding() { + let (configs, tunnel_engine, ctx) = jsonrpc_and_rest_route_configs("mcp", "enforce"); + let scenario = run_upgrade_scenario( + JSONRPC_WEBSOCKET_UPGRADE_REQUEST, + UNALLOWED_TOOL_CALL, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_route_selection( + &configs, + tunnel_engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + }) + }, + ) + .await; + assert_upgrade_denied_before_forwarding(&scenario); + } + + #[tokio::test] + async fn route_selected_audit_jsonrpc_websocket_upgrade_is_denied_before_forwarding() { + // Audit mode forwards requests that policy would deny, so the upgrade + // refusal must not depend on the policy decision. + let (configs, tunnel_engine, ctx) = jsonrpc_and_rest_route_configs("json-rpc", "audit"); + let scenario = run_upgrade_scenario( + JSONRPC_WEBSOCKET_UPGRADE_REQUEST, + UNALLOWED_TOOL_CALL, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_route_selection( + &configs, + tunnel_engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + }) + }, + ) + .await; + assert_upgrade_denied_before_forwarding(&scenario); + } + + #[tokio::test] + async fn single_endpoint_mcp_websocket_upgrade_is_denied_before_forwarding() { + let (config, tunnel_engine, ctx) = mcp_test_relay_context(); + let scenario = run_upgrade_scenario( + JSONRPC_WEBSOCKET_UPGRADE_REQUEST, + UNALLOWED_TOOL_CALL, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_inspection(&config, tunnel_engine, &mut client, &mut upstream, &ctx) + .await + }) + }, + ) + .await; + assert_upgrade_denied_before_forwarding(&scenario); + } + + #[tokio::test] + async fn route_selected_rest_websocket_upgrade_still_relays_beside_mcp() { + // The refusal follows the selected endpoint's protocol: a REST + // upgrade on the same host and port keeps its documented raw relay. + let (configs, tunnel_engine, ctx) = jsonrpc_and_rest_route_configs("mcp", "enforce"); + let frame = br#"{"type":"ping"}"#; + let scenario = run_upgrade_scenario( + b"GET /api/ws HTTP/1.1\r\nHost: mcp.example.test:8000\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==\r\nSec-WebSocket-Version: 13\r\n\r\n", + frame, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_route_selection( + &configs, + tunnel_engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + }) + }, + ) + .await; + assert!( + scenario.response.starts_with("HTTP/1.1 101"), + "REST upgrade should still switch protocols, got: {}", + scenario.response + ); + assert!(contains_bytes( + &scenario.upstream_seen, + &masked_text_frame(frame) + )); + } + #[tokio::test] + async fn route_selected_mcp_receive_stream_without_upgrade_is_still_forwarded() { + let (configs, tunnel_engine, ctx) = jsonrpc_and_rest_route_configs("mcp", "enforce"); let (mut app, mut relay_client) = tokio::io::duplex(8192); let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); let relay = tokio::spawn(async move { - relay_with_inspection( - &config, + relay_with_route_selection( + &configs, tunnel_engine, &mut relay_client, &mut relay_upstream, @@ -9155,72 +9680,413 @@ network_policies: }); app.write_all( - b"POST /write HTTP/1.1\r\nHost: api.example.test\r\nContent-Length: 0\r\nConnection: keep-alive\r\n\r\n", + b"GET /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nAccept: text/event-stream\r\nMCP-Protocol-Version: 2025-11-25\r\n\r\n", ) .await .unwrap(); - - let mut first_upstream = [0u8; 512]; - let n = tokio::time::timeout( - std::time::Duration::from_secs(1), - upstream.read(&mut first_upstream), + let forwarded = tokio::time::timeout( + std::time::Duration::from_secs(2), + read_http_headers(&mut upstream), ) .await - .expect("first request should reach upstream") - .unwrap(); - let first_upstream = String::from_utf8_lossy(&first_upstream[..n]); - assert!( - first_upstream.starts_with("POST /write HTTP/1.1"), - "unexpected upstream request: {first_upstream:?}" + .expect("receive-stream GET should reach the upstream"); + let forwarded = String::from_utf8_lossy(&forwarded); + assert!(forwarded.starts_with("GET /mcp HTTP/1.1\r\n")); + assert!(!forwarded.to_ascii_lowercase().contains("upgrade")); + relay.abort(); + let _ = relay.await; + } + + /// A GraphQL-over-WebSocket message that no GraphQL fixture allows. + const UNALLOWED_GRAPHQL_MUTATION: &[u8] = + br#"{"id":"1","type":"subscribe","payload":{"query":"mutation { deleteRepository }"}}"#; + + /// A GET whose query the GraphQL fixtures allow, plus WebSocket upgrade + /// headers. + const GRAPHQL_WEBSOCKET_UPGRADE_REQUEST: &[u8] = b"GET /graphql?query=%7Bviewer%7D HTTP/1.1\r\nHost: graphql.example.test:8000\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==\r\nSec-WebSocket-Version: 13\r\n\r\n"; + + fn assert_graphql_upgrade_refused_before_forwarding(scenario: &UpgradeScenario) { + assert_upgrade_refused_before_forwarding( + scenario, + UNALLOWED_GRAPHQL_MUTATION, + crate::l7::rest::UNSUPPORTED_GRAPHQL_UPGRADE_DETAIL, ); + } - upstream - .write_all(b"HTTP/1.1 200 OK\r\nContent-Length: 2\r\nConnection: keep-alive\r\n\r\nOK") - .await - .unwrap(); + #[tokio::test] + async fn single_endpoint_graphql_websocket_upgrade_is_denied_before_forwarding() { + let (config, tunnel_engine, ctx) = graphql_test_relay_context(); + let scenario = run_upgrade_scenario( + GRAPHQL_WEBSOCKET_UPGRADE_REQUEST, + UNALLOWED_GRAPHQL_MUTATION, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_inspection(&config, tunnel_engine, &mut client, &mut upstream, &ctx) + .await + }) + }, + ) + .await; + assert_graphql_upgrade_refused_before_forwarding(&scenario); + } + + #[tokio::test(start_paused = true)] + async fn single_endpoint_graphql_upgrade_is_refused_before_body_read() { + // Send only the complete head. Neither an oversized declaration nor + // an incomplete allowed-size body may delay the upgrade refusal. + for content_length in [65537, 16] { + for enforcement in [EnforcementMode::Enforce, EnforcementMode::Audit] { + for upgrade in ["websocket", "custom"] { + let (mut config, engine, ctx) = graphql_test_relay_context(); + config.enforcement = enforcement; + assert_eq!(config.graphql_max_body_bytes, 65536); + let request = format!( + "POST /graphql HTTP/1.1\r\nHost: graphql.example.test:8000\r\nContent-Type: application/json\r\nConnection: Upgrade\r\nUpgrade: {upgrade}\r\nContent-Length: {content_length}\r\n\r\n" + ); + let started = tokio::time::Instant::now(); + let scenario = run_upgrade_scenario( + request.as_bytes(), + UNALLOWED_GRAPHQL_MUTATION, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_inspection( + &config, + engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + }) + }, + ) + .await; + assert_graphql_upgrade_refused_before_forwarding(&scenario); + assert_eq!(started.elapsed(), std::time::Duration::ZERO); + } + } + } + } - let mut first_response = [0u8; 512]; - let n = tokio::time::timeout( - std::time::Duration::from_secs(1), - app.read(&mut first_response), + #[tokio::test(start_paused = true)] + async fn single_endpoint_graphql_upgrade_preserves_head_validation() { + // Framing errors are rejected by the HTTP parser, and authority + // mismatches retain their specific denial before upgrade handling. + for (host, framing, expected) in [ + ( + "graphql.example.test:8000", + "Content-Length: 16\r\nTransfer-Encoding: chunked\r\n", + "", + ), + ( + "graphql.example.test:8000", + "Content-Length: 16\r\nContent-Length: 17\r\n", + "", + ), + ( + "other.example.test:8000", + "Content-Length: 16\r\n", + "request_authority_mismatch", + ), + ] { + let (config, engine, ctx) = graphql_test_relay_context(); + let request = format!( + "POST /graphql HTTP/1.1\r\nHost: {host}\r\nConnection: Upgrade\r\nUpgrade: custom\r\n{framing}\r\n" + ); + let scenario = run_upgrade_scenario( + request.as_bytes(), + UNALLOWED_GRAPHQL_MUTATION, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_inspection(&config, engine, &mut client, &mut upstream, &ctx) + .await + }) + }, + ) + .await; + assert!(scenario.upstream_seen.is_empty()); + if expected.is_empty() { + assert!(scenario.response.is_empty()); + } else { + assert!(scenario.response.starts_with("HTTP/1.1 403")); + assert!(scenario.body.contains(expected)); + } + } + } + + #[tokio::test(start_paused = true)] + async fn single_endpoint_graphql_post_still_inspects_body() { + for (field, host, token) in [ + ("viewer", "graphql.example.test:8000", ""), + ("admin", "graphql.example.test:8000", ""), + ("viewer", "", ""), + ("viewer", "", "openshell:resolve:env:v1_API_TOKEN"), + ] { + let (config, engine, ctx) = graphql_test_relay_context(); + let (mut app, mut client) = tokio::io::duplex(8192); + let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); + let relay = tokio::spawn(async move { + relay_with_inspection(&config, engine, &mut client, &mut relay_upstream, &ctx).await + }); + let body = format!(r#"{{"query":"{{{field}}}","variables":{{"token":"{token}"}}}}"#); + let (version, authority) = if host.is_empty() { + ("HTTP/1.0", String::new()) + } else { + ("HTTP/1.1", format!("Host: {host}\r\n")) + }; + let request = format!( + "POST /graphql {version}\r\n{authority}Content-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{body}", + body.len() + ); + app.write_all(request.as_bytes()).await.unwrap(); + let mut forwarded = Vec::new(); + if field == "viewer" && token.is_empty() { + let headers = tokio::time::timeout( + std::time::Duration::from_secs(1), + read_http_headers(&mut upstream), + ) + .await + .expect("allowed POST head must reach upstream"); + assert!(headers.starts_with(format!("POST /graphql {version}\r\n").as_bytes())); + let mut bytes = vec![0; body.len()]; + tokio::time::timeout( + std::time::Duration::from_secs(1), + upstream.read_exact(&mut bytes), + ) + .await + .expect("allowed POST must reach upstream") + .unwrap(); + assert_eq!(bytes, body.as_bytes()); + } else { + let mut response = Vec::new(); + tokio::time::timeout( + std::time::Duration::from_secs(1), + app.read_to_end(&mut response), + ) + .await + .expect("unlisted field or credential marker without Host must be denied") + .unwrap(); + assert!(response.starts_with(b"HTTP/1.1 403")); + if !token.is_empty() { + assert!( + String::from_utf8_lossy(&response).contains("request_authority_mismatch") + ); + } + upstream.read_to_end(&mut forwarded).await.unwrap(); + assert!(forwarded.is_empty()); + } + relay.abort(); + let _ = relay.await; + } + } + + #[tokio::test] + async fn route_selected_graphql_websocket_upgrade_is_denied_before_forwarding() { + // Audit mode forwards requests that policy would deny, so the upgrade + // refusal must not depend on the policy decision. + for enforcement in ["enforce", "audit"] { + let (configs, tunnel_engine, ctx) = graphql_and_rest_route_configs(enforcement); + let scenario = run_upgrade_scenario( + GRAPHQL_WEBSOCKET_UPGRADE_REQUEST, + UNALLOWED_GRAPHQL_MUTATION, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_route_selection( + &configs, + tunnel_engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + }) + }, + ) + .await; + assert_graphql_upgrade_refused_before_forwarding(&scenario); + } + } + + #[tokio::test] + async fn route_selected_audit_graphql_upgrade_with_denied_query_is_refused() { + // Audit mode forwards a query the policy denies. The refusal must + // still fire, because it runs before the policy decision. + let (configs, tunnel_engine, ctx) = graphql_and_rest_route_configs("audit"); + let scenario = run_upgrade_scenario( + b"GET /graphql?query=%7Badmin%7D HTTP/1.1\r\nHost: graphql.example.test:8000\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==\r\nSec-WebSocket-Version: 13\r\n\r\n", + UNALLOWED_GRAPHQL_MUTATION, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_route_selection( + &configs, + tunnel_engine, + &mut client, + &mut upstream, + &ctx, + ) + .await + }) + }, ) - .await - .expect("first response should reach client") - .unwrap(); - let first_response = String::from_utf8_lossy(&first_response[..n]); - assert!(first_response.contains("200 OK")); + .await; + assert_graphql_upgrade_refused_before_forwarding(&scenario); + } - engine.reload(TEST_POLICY, reloaded_data).unwrap(); - app.write_all( - b"POST /write HTTP/1.1\r\nHost: api.example.test\r\nContent-Length: 0\r\nConnection: keep-alive\r\n\r\n", + #[tokio::test] + async fn single_endpoint_graphql_subscription_handshake_gets_upgrade_refusal() { + // A standard GraphQL-over-WebSocket handshake carries no query. It + // must receive the refusal that names the supported alternative, not + // a policy denial that suggests adding a rule. + let (config, tunnel_engine, ctx) = graphql_test_relay_context(); + let scenario = run_upgrade_scenario( + b"GET /graphql HTTP/1.1\r\nHost: graphql.example.test:8000\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==\r\nSec-WebSocket-Version: 13\r\nSec-WebSocket-Protocol: graphql-transport-ws\r\n\r\n", + UNALLOWED_GRAPHQL_MUTATION, + move |mut client, mut upstream| { + tokio::spawn(async move { + relay_with_inspection(&config, tunnel_engine, &mut client, &mut upstream, &ctx) + .await + }) + }, ) - .await - .unwrap(); + .await; + assert_graphql_upgrade_refused_before_forwarding(&scenario); + } - tokio::time::timeout(std::time::Duration::from_secs(1), relay) + #[tokio::test] + async fn route_selected_graphql_query_without_upgrade_is_still_forwarded() { + let (configs, tunnel_engine, ctx) = graphql_and_rest_route_configs("enforce"); + let (mut app, mut relay_client) = tokio::io::duplex(8192); + let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); + let relay = tokio::spawn(async move { + relay_with_route_selection( + &configs, + tunnel_engine, + &mut relay_client, + &mut relay_upstream, + &ctx, + ) .await - .expect("relay should close stale tunnel") - .unwrap() - .unwrap(); + }); - let mut second_upstream = [0u8; 128]; - let n = tokio::time::timeout( - std::time::Duration::from_secs(1), - upstream.read(&mut second_upstream), + app.write_all( + b"GET /graphql?query=%7Bviewer%7D HTTP/1.1\r\nHost: graphql.example.test:8000\r\n\r\n", ) .await - .expect("upstream side should close") .unwrap(); - assert_eq!(n, 0, "stale request must not be forwarded upstream"); + let forwarded = tokio::time::timeout( + std::time::Duration::from_secs(2), + read_http_headers(&mut upstream), + ) + .await + .expect("allowed GraphQL GET should reach the upstream"); + let forwarded = String::from_utf8_lossy(&forwarded); + assert!(forwarded.starts_with("GET /graphql?query=%7Bviewer%7D HTTP/1.1\r\n")); + assert!(!forwarded.to_ascii_lowercase().contains("upgrade")); + relay.abort(); + let _ = relay.await; + } + + fn masked_text_frame(payload: &[u8]) -> Vec { + let mask = [0x11, 0x22, 0x33, 0x44]; + assert!( + payload.len() <= 125, + "test helper only supports small frames" + ); + let payload_len = u8::try_from(payload.len()).expect("small frame length"); + let mut frame = vec![0x81, 0x80 | payload_len]; + frame.extend_from_slice(&mask); + frame.extend( + payload + .iter() + .enumerate() + .map(|(idx, byte)| byte ^ mask[idx % 4]), + ); + frame + } + + async fn read_http_headers(reader: &mut R) -> Vec { + let mut bytes = Vec::new(); + let mut byte = [0u8; 1]; + loop { + reader.read_exact(&mut byte).await.unwrap(); + bytes.push(byte[0]); + if bytes.ends_with(b"\r\n\r\n") { + return bytes; + } + } + } + + async fn read_text_frame( + reader: &mut R, + ) -> std::io::Result<(bool, String)> { + let mut header = [0u8; 2]; + reader.read_exact(&mut header).await?; + assert_eq!(header[0] & 0x0f, 0x1, "expected text frame"); + let masked = header[1] & 0x80 != 0; + let payload_len = usize::from(header[1] & 0x7f); + assert!(payload_len <= 125, "test helper only supports small frames"); + let mut mask = [0u8; 4]; + if masked { + reader.read_exact(&mut mask).await?; + } + let mut payload = vec![0u8; payload_len]; + reader.read_exact(&mut payload).await?; + if masked { + for (idx, byte) in payload.iter_mut().enumerate() { + *byte ^= mask[idx % 4]; + } + } + Ok((masked, String::from_utf8(payload).expect("text payload"))) } #[tokio::test] - async fn passthrough_relay_closes_keep_alive_tunnel_after_policy_generation_change() { - let policy_data = "network_policies: {}\n"; - let engine = OpaEngine::from_strings(TEST_POLICY, policy_data).unwrap(); - let generation_guard = engine - .generation_guard(engine.current_generation()) + async fn l7_relay_closes_keep_alive_tunnel_after_policy_generation_change() { + let initial_data = r#" +network_policies: + rest_api: + name: rest_api + endpoints: + - host: api.example.test + port: 8080 + protocol: rest + enforcement: enforce + rules: + - allow: + method: POST + path: "/write" + binaries: + - { path: /usr/bin/curl } +"#; + let reloaded_data = r#" +network_policies: + rest_api: + name: rest_api + endpoints: + - host: api.example.test + port: 8080 + protocol: rest + enforcement: enforce + rules: + - allow: + method: GET + path: "/write" + binaries: + - { path: /usr/bin/curl } +"#; + let engine = OpaEngine::from_strings(TEST_POLICY, initial_data).unwrap(); + let input = NetworkInput { + host: "api.example.test".into(), + port: 8080, + binary_path: PathBuf::from("/usr/bin/curl"), + binary_sha256: "unused".into(), + ancestors: vec![], + cmdline_paths: vec![], + }; + let (endpoint_config, generation) = engine + .query_endpoint_config_with_generation(&input) .unwrap(); + let config = crate::l7::parse_l7_config(&endpoint_config.unwrap()).unwrap(); + let tunnel_engine = engine.clone_engine_for_tunnel(generation).unwrap(); let ctx = L7EvalContext { host: "api.example.test".into(), port: 8080, @@ -9236,18 +10102,18 @@ network_policies: let (mut app, mut relay_client) = tokio::io::duplex(8192); let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); let relay = tokio::spawn(async move { - relay_passthrough_with_credentials( + relay_with_inspection( + &config, + tunnel_engine, &mut relay_client, &mut relay_upstream, &ctx, - &generation_guard, - None, ) .await }); app.write_all( - b"GET /first HTTP/1.1\r\nHost: api.example.test\r\nConnection: keep-alive\r\n\r\n", + b"POST /write HTTP/1.1\r\nHost: api.example.test\r\nContent-Length: 0\r\nConnection: keep-alive\r\n\r\n", ) .await .unwrap(); @@ -9258,7 +10124,99 @@ network_policies: upstream.read(&mut first_upstream), ) .await - .expect("first passthrough request should reach upstream") + .expect("first request should reach upstream") + .unwrap(); + let first_upstream = String::from_utf8_lossy(&first_upstream[..n]); + assert!( + first_upstream.starts_with("POST /write HTTP/1.1"), + "unexpected upstream request: {first_upstream:?}" + ); + + upstream + .write_all(b"HTTP/1.1 200 OK\r\nContent-Length: 2\r\nConnection: keep-alive\r\n\r\nOK") + .await + .unwrap(); + + let mut first_response = [0u8; 512]; + let n = tokio::time::timeout( + std::time::Duration::from_secs(1), + app.read(&mut first_response), + ) + .await + .expect("first response should reach client") + .unwrap(); + let first_response = String::from_utf8_lossy(&first_response[..n]); + assert!(first_response.contains("200 OK")); + + engine.reload(TEST_POLICY, reloaded_data).unwrap(); + app.write_all( + b"POST /write HTTP/1.1\r\nHost: api.example.test\r\nContent-Length: 0\r\nConnection: keep-alive\r\n\r\n", + ) + .await + .unwrap(); + + tokio::time::timeout(std::time::Duration::from_secs(1), relay) + .await + .expect("relay should close stale tunnel") + .unwrap() + .unwrap(); + + let mut second_upstream = [0u8; 128]; + let n = tokio::time::timeout( + std::time::Duration::from_secs(1), + upstream.read(&mut second_upstream), + ) + .await + .expect("upstream side should close") + .unwrap(); + assert_eq!(n, 0, "stale request must not be forwarded upstream"); + } + + #[tokio::test] + async fn passthrough_relay_closes_keep_alive_tunnel_after_policy_generation_change() { + let policy_data = "network_policies: {}\n"; + let engine = OpaEngine::from_strings(TEST_POLICY, policy_data).unwrap(); + let generation_guard = engine + .generation_guard(engine.current_generation()) + .unwrap(); + let ctx = L7EvalContext { + host: "api.example.test".into(), + port: 8080, + request_default_port: Some(8080), + policy_name: "rest_api".into(), + binary_path: "/usr/bin/curl".into(), + ancestors: vec![], + cmdline_paths: vec![], + secret_resolver: None, + ..Default::default() + }; + + let (mut app, mut relay_client) = tokio::io::duplex(8192); + let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); + let relay = tokio::spawn(async move { + relay_passthrough_with_credentials( + &mut relay_client, + &mut relay_upstream, + &ctx, + &generation_guard, + None, + ) + .await + }); + + app.write_all( + b"GET /first HTTP/1.1\r\nHost: api.example.test\r\nConnection: keep-alive\r\n\r\n", + ) + .await + .unwrap(); + + let mut first_upstream = [0u8; 512]; + let n = tokio::time::timeout( + std::time::Duration::from_secs(1), + upstream.read(&mut first_upstream), + ) + .await + .expect("first passthrough request should reach upstream") .unwrap(); let first_upstream = String::from_utf8_lossy(&first_upstream[..n]); assert!(first_upstream.starts_with("GET /first HTTP/1.1")); @@ -9307,131 +10265,1044 @@ network_policies: } #[tokio::test] - async fn jsonrpc_relay_forwards_allowed_method() { - let (config, tunnel_engine, ctx) = jsonrpc_test_relay_context(); - let (mut app, mut relay_client) = tokio::io::duplex(8192); - let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); - let relay = tokio::spawn(async move { - relay_with_inspection( + async fn jsonrpc_relay_forwards_allowed_method() { + let (config, tunnel_engine, ctx) = jsonrpc_test_relay_context(); + let (mut app, mut relay_client) = tokio::io::duplex(8192); + let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); + let relay = tokio::spawn(async move { + relay_with_inspection( + &config, + tunnel_engine, + &mut relay_client, + &mut relay_upstream, + &ctx, + ) + .await + }); + + let body = br#"{"jsonrpc":"2.0","id":1,"method":"initialize"}"#; + let request = format!( + "POST /rpc HTTP/1.1\r\nHost: jsonrpc.example.test:8000\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n\r\n", + body.len() + ); + app.write_all(request.as_bytes()).await.unwrap(); + app.write_all(body).await.unwrap(); + + let mut upstream_bytes = Vec::new(); + let mut upstream_buf = [0u8; 1024]; + loop { + let n = tokio::time::timeout( + std::time::Duration::from_secs(1), + upstream.read(&mut upstream_buf), + ) + .await + .expect("allowed JSON-RPC request should reach upstream") + .unwrap(); + assert_ne!(n, 0, "upstream closed before JSON-RPC body arrived"); + upstream_bytes.extend_from_slice(&upstream_buf[..n]); + if String::from_utf8_lossy(&upstream_bytes).contains(r#""method":"initialize""#) { + break; + } + } + let upstream_request = String::from_utf8_lossy(&upstream_bytes); + assert!(upstream_request.starts_with("POST /rpc HTTP/1.1")); + assert!(upstream_request.contains(r#""method":"initialize""#)); + + upstream + .write_all( + b"HTTP/1.1 200 OK\r\nContent-Length: 36\r\nConnection: close\r\n\r\n{\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{}}", + ) + .await + .unwrap(); + + let mut response = [0u8; 512]; + let n = tokio::time::timeout(std::time::Duration::from_secs(1), app.read(&mut response)) + .await + .expect("upstream response should reach client") + .unwrap(); + assert!(String::from_utf8_lossy(&response[..n]).contains("200 OK")); + + drop(app); + tokio::time::timeout(std::time::Duration::from_secs(1), relay) + .await + .expect("relay should complete") + .unwrap() + .unwrap(); + } + + #[tokio::test] + async fn mcp_relay_forwards_standalone_initialize_without_version_header() { + let (config, tunnel_engine, ctx) = mcp_test_relay_context(); + let (mut app, mut relay_client) = tokio::io::duplex(8192); + let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); + let relay = tokio::spawn(async move { + relay_with_inspection( + &config, + tunnel_engine, + &mut relay_client, + &mut relay_upstream, + &ctx, + ) + .await + }); + + let body = br#"{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"test","version":"1"}}}"#; + let request = format!( + "POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n\r\n", + body.len() + ); + app.write_all(request.as_bytes()).await.unwrap(); + app.write_all(body).await.unwrap(); + + let mut upstream_bytes = vec![0; 2048]; + let count = tokio::time::timeout( + std::time::Duration::from_secs(1), + upstream.read(&mut upstream_bytes), + ) + .await + .expect("standalone initialize should reach upstream") + .unwrap(); + let upstream_request = String::from_utf8_lossy(&upstream_bytes[..count]); + assert!(upstream_request.contains(r#""method":"initialize""#)); + + upstream + .write_all( + b"HTTP/1.1 200 OK\r\nContent-Length: 36\r\nConnection: close\r\n\r\n{\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{}}", + ) + .await + .unwrap(); + let mut response = [0; 512]; + let count = + tokio::time::timeout(std::time::Duration::from_secs(1), app.read(&mut response)) + .await + .expect("initialize response should reach client") + .unwrap(); + assert!(String::from_utf8_lossy(&response[..count]).contains("200 OK")); + + drop(app); + tokio::time::timeout(std::time::Duration::from_secs(1), relay) + .await + .expect("relay should complete") + .unwrap() + .unwrap(); + } + + fn sessionless_mcp_body(method: &str, mut params: serde_json::Value) -> String { + params["_meta"] = serde_json::json!({ + "io.modelcontextprotocol/protocolVersion": "2026-07-28", + "io.modelcontextprotocol/clientCapabilities": {} + }); + serde_json::json!({"jsonrpc":"2.0", "id":1, "method":method, "params":params}).to_string() + } + + async fn run_sessionless_mcp_relay( + route_selected: bool, + headers: &str, + body: &str, + upstream_response: &str, + ) -> (String, Vec) { + run_mcp_relay_case( + mcp_sessionless_test_relay_context(), + route_selected, + headers, + body, + upstream_response, + ) + .await + } + + #[tokio::test] + async fn chunked_http_pipeline_authorizes_each_request() { + for protocol in ["mcp", "json-rpc", "graphql", "rest"] { + for route_selected in [false, true] { + for second_allowed in [false, true] { + let rules = match protocol { + "mcp" => "method: tools/call, tool: echo", + "json-rpc" => "method: echo", + "graphql" => "operation_type: query, fields: [echo]", + "rest" => "method: POST, path: /mcp/allowed", + _ => unreachable!(), + }; + let endpoint_path = if protocol == "rest" { + "/mcp/**" + } else { + "/mcp" + }; + let data = format!( + r" +network_policies: + mcp_api: + name: mcp_api + endpoints: + - host: mcp.example.test + port: 8000 + path: {endpoint_path} + protocol: {protocol} + enforcement: enforce + rules: + - allow: {{ {rules} }} + binaries: + - {{ path: /usr/bin/python3 }} +" + ); + let (config, tunnel_engine, ctx) = mcp_relay_context_from_data(&data); + let body = |id, name| { + match protocol { + "mcp" => serde_json::json!({ + "jsonrpc": "2.0", "id": id, "method": "tools/call", + "params": {"name": name, "arguments": {}} + }), + "json-rpc" => { + serde_json::json!({"jsonrpc": "2.0", "id": id, "method": name}) + } + "graphql" => { + serde_json::json!({"query": format!("query {{ {name} }}")}) + } + "rest" => serde_json::json!({"id": id, "value": name}), + _ => unreachable!(), + } + .to_string() + }; + let first = body(1, "echo"); + let second = body(2, if second_allowed { "echo" } else { "blocked" }); + let mut wire = String::new(); + for (index, body) in [&first, &second].into_iter().enumerate() { + let target = if protocol == "rest" { + if index == 0 || second_allowed { + "/mcp/allowed" + } else { + "/mcp/blocked" + } + } else { + "/mcp" + }; + write!( + wire, + "POST {target} HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\nMCP-Protocol-Version: 2025-11-25\r\nTransfer-Encoding: chunked\r\n\r\n{:x}\r\n{body}\r\n0\r\n\r\n", + body.len() + ) + .unwrap(); + } + let (mut app, mut relay_client) = tokio::io::duplex(8192); + let (mut relay_upstream, upstream) = tokio::io::duplex(8192); + // Queue both requests before the relay reads. A normal + // sequential exchange cannot expose chunked read-ahead loss. + app.write_all(wire.as_bytes()).await.unwrap(); + app.shutdown().await.unwrap(); + let relay = async move { + if route_selected { + relay_with_route_selection( + &[config], + tunnel_engine, + &mut relay_client, + &mut relay_upstream, + &ctx, + ) + .await + } else { + relay_with_inspection( + &config, + tunnel_engine, + &mut relay_client, + &mut relay_upstream, + &ctx, + ) + .await + } + }; + let server = async move { + let mut upstream = tokio::io::BufReader::new(upstream); + let provider = crate::l7::rest::RestProvider::with_options( + crate::l7::path::CanonicalizeOptions::default(), + ); + let mut forwarded = Vec::new(); + while let Some(mut request) = + provider.parse_request(&mut upstream).await.unwrap() + { + // REST streams chunked framing; the body inspectors + // normalize the same message to Content-Length. + if protocol == "rest" { + assert!(matches!( + request.body_length, + crate::l7::provider::BodyLength::Chunked + )); + } else { + assert!(matches!( + request.body_length, + crate::l7::provider::BodyLength::ContentLength(_) + )); + } + let body = crate::l7::http::read_body_for_inspection( + &mut upstream, + &mut request, + 1024, + ) + .await + .unwrap(); + forwarded.push(String::from_utf8(body).unwrap()); + upstream + .write_all(b"HTTP/1.1 204 No Content\r\nContent-Length: 0\r\n\r\n") + .await + .unwrap(); + } + forwarded + }; + let client = async move { + let mut response = String::new(); + app.read_to_string(&mut response).await.unwrap(); + response + }; + let (result, forwarded, response) = Box::pin(tokio::time::timeout( + std::time::Duration::from_secs(5), + async { tokio::join!(relay, server, client) }, + )) + .await + .expect("pipelined requests must finish without losing a request"); + result.unwrap(); + let expected = if second_allowed { + vec![first, second] + } else { + vec![first] + }; + assert_eq!( + forwarded, expected, + "{protocol}, route_selected={route_selected}" + ); + assert_eq!( + response.matches("HTTP/1.1 204 No Content").count(), + if second_allowed { 2 } else { 1 }, + "{response}" + ); + assert_eq!( + response.contains("403 Forbidden"), + !second_allowed, + "{response}" + ); + } + } + } + } + + async fn run_mcp_relay_case( + context: (L7EndpointConfig, TunnelPolicyEngine, L7EvalContext), + route_selected: bool, + headers: &str, + body: &str, + upstream_response: &str, + ) -> (String, Vec) { + run_mcp_method_relay_case( + context, + route_selected, + "POST", + headers, + body, + upstream_response, + ) + .await + } + + async fn run_mcp_method_relay_case( + (config, tunnel_engine, ctx): (L7EndpointConfig, TunnelPolicyEngine, L7EvalContext), + route_selected: bool, + method: &str, + headers: &str, + body: &str, + upstream_response: &str, + ) -> (String, Vec) { + let (mut app, mut relay_client) = tokio::io::duplex(8192); + let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); + let relay = tokio::spawn(async move { + if route_selected { + relay_with_route_selection( + &[config], + tunnel_engine, + &mut relay_client, + &mut relay_upstream, + &ctx, + ) + .await + } else { + relay_with_inspection( + &config, + tunnel_engine, + &mut relay_client, + &mut relay_upstream, + &ctx, + ) + .await + } + }); + let upstream_response = upstream_response.to_string(); + let server = tokio::spawn(async move { + let mut forwarded = Vec::new(); + let mut bytes = [0; 2048]; + loop { + let count = upstream.read(&mut bytes).await.unwrap(); + if count == 0 { + return forwarded; + } + forwarded.extend_from_slice(&bytes[..count]); + if let Some(header_end) = forwarded + .windows(4) + .position(|window| window == b"\r\n\r\n") + { + // Middleware can change the body length. Wait for the + // complete forwarded representation, not the input size. + let header = std::str::from_utf8(&forwarded[..header_end]).unwrap(); + let body_len = header + .lines() + .find_map(|line| { + let (name, value) = line.split_once(':')?; + name.eq_ignore_ascii_case("content-length") + .then(|| value.trim().parse::().unwrap()) + }) + .expect("MCP fixture requests include Content-Length"); + if forwarded.len() < header_end + 4 + body_len { + continue; + } + upstream + .write_all(upstream_response.as_bytes()) + .await + .unwrap(); + return forwarded; + } + } + }); + let request = format!( + "{method} /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\nAccept: application/json, text/event-stream\r\n{headers}Content-Length: {}\r\nConnection: close\r\n\r\n{body}", + body.len() + ); + app.write_all(request.as_bytes()).await.unwrap(); + let mut response = Vec::new(); + // A relayed SSE response can leave the client half of the duplex + // open. Read this fixture's complete HTTP response, then close it. + tokio::time::timeout(std::time::Duration::from_secs(5), async { + let mut bytes = [0; 2048]; + loop { + let count = app.read(&mut bytes).await.unwrap(); + if count == 0 { + break; + } + response.extend_from_slice(&bytes[..count]); + if let Some(header_end) = + response.windows(4).position(|window| window == b"\r\n\r\n") + { + let header = std::str::from_utf8(&response[..header_end]).unwrap(); + let length = header + .lines() + .find_map(|line| { + let (name, value) = line.split_once(':')?; + name.eq_ignore_ascii_case("content-length") + .then(|| value.trim().parse::().unwrap()) + }) + .expect("fixture responses include Content-Length"); + if response.len() >= header_end + 4 + length { + break; + } + } + } + }) + .await + .unwrap_or_else(|_| { + panic!("timed out receiving response for {body}; received {response:?}") + }); + drop(app); + relay.await.unwrap().unwrap(); + (String::from_utf8(response).unwrap(), server.await.unwrap()) + } + + #[tokio::test] + async fn mcp_legacy_receive_stream_get_does_not_admit_tool_bodies_or_delete() { + let data = r#" +network_policies: + mcp_api: + name: mcp_api + endpoints: + - host: mcp.example.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: enforce + mcp: + versions: ["2025-11-25"] + rules: + - allow: + method: tools/call + tool: read_status + deny_rules: + - method: tools/call + tool: delete_resource + binaries: + - { path: /usr/bin/python3 } +"#; + let tool_body = r#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"delete_resource","arguments":{}}}"#; + let allowed_tool_body = r#"{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"read_status","arguments":{}}}"#; + let event = "event: message\ndata: {\"jsonrpc\":\"2.0\",\"method\":\"notifications/tools/list_changed\"}\n\n"; + let upstream_response = format!( + "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{event}", + event.len() + ); + for route_selected in [false, true] { + for (method, body, status) in [ + ("GET", "", "200 OK"), + ("GET", tool_body, "403 Forbidden"), + ("GET", allowed_tool_body, "403 Forbidden"), + ("DELETE", "", "400 Bad Request"), + ] { + let (response, forwarded) = run_mcp_method_relay_case( + mcp_relay_context_from_data(data), + route_selected, + method, + "MCP-Protocol-Version: 2025-11-25\r\n", + body, + &upstream_response, + ) + .await; + assert!( + response.starts_with(&format!("HTTP/1.1 {status}")), + "{method}, body={body}, route_selected={route_selected}: {response}" + ); + if status == "200 OK" { + // The GET receive-stream exception applies only without a + // client operation body. Preserve its complete SSE response. + let forwarded = String::from_utf8(forwarded).unwrap(); + let (headers, forwarded_body) = forwarded.split_once("\r\n\r\n").unwrap(); + assert!(headers.starts_with("GET /mcp HTTP/1.1\r\n")); + assert!(forwarded_body.is_empty()); + assert!( + response.ends_with(event), + "receive-stream event changed: {response}" + ); + } else { + assert!( + forwarded.is_empty(), + "{method}: rejected request reached upstream" + ); + if method == "DELETE" { + // Legacy cleanup is unsupported: an empty DELETE is + // rejected as an invalid MCP body, not as sessionless 405. + assert!(response.contains("invalid_mcp_request"), "{response}"); + } + } + } + } + } + + /// Replaces a tool call and its sessionless name mirror in one real stage. + struct McpToolReplacingService { + replacement: Vec, + tool_name: &'static str, + sessionless: bool, + invocations: Arc, + } + + #[tonic::async_trait] + impl openshell_core::middleware::InProcessMiddleware for McpToolReplacingService { + async fn describe(&self) -> openshell_core::proto::MiddlewareManifest { + openshell_core::middleware::InProcessMiddleware::describe(&BodyReplacingService { + replacement: b"", + }) + .await + } + + async fn validate_config( + &self, + _middleware_name: &str, + _config: &prost_types::Struct, + ) -> Result<()> { + Ok(()) + } + + async fn evaluate_http_request( + &self, + request: openshell_core::middleware::HttpRequestView<'_>, + ) -> Result { + use openshell_core::proto::{ + Decision, ExistingHeaderAction, HeaderMutation, HttpRequestResult, WriteHeader, + header_mutation, + }; + let original: serde_json::Value = serde_json::from_slice(request.body()).unwrap(); + assert_eq!(original["params"]["name"], "read_status"); + self.invocations + .fetch_add(1, std::sync::atomic::Ordering::SeqCst); + let header_mutations = if self.sessionless { + vec![HeaderMutation { + operation: Some(header_mutation::Operation::Write(WriteHeader { + name: "Mcp-Name".into(), + value: self.tool_name.into(), + on_existing: ExistingHeaderAction::Overwrite as i32, + })), + }] + } else { + Vec::new() + }; + Ok(HttpRequestResult { + decision: Decision::Allow as i32, + body: self.replacement.clone(), + has_body: true, + header_mutations, + ..Default::default() + }) + } + } + + #[tokio::test] + async fn mcp_middleware_tool_rewrites_obey_policy_with_matching_metadata() { + use std::sync::atomic::{AtomicUsize, Ordering}; + + for route_selected in [false, true] { + // A shared allowlist must preserve the selected revision's request + // profile as well as its membership in the permitted revisions. + for (version, configured_versions) in [ + ("2025-06-18", &["2025-06-18"][..]), + ("2025-11-25", &["2025-11-25"][..]), + ("2026-07-28", &["2026-07-28"][..]), + ("2025-11-25", &["2025-11-25", "2026-07-28"][..]), + ("2026-07-28", &["2025-11-25", "2026-07-28"][..]), + ] { + let configured_versions = serde_json::to_string(configured_versions).unwrap(); + let sessionless = version == "2026-07-28"; + let body_for = |name, arguments| { + let params = serde_json::json!({"name": name, "arguments": arguments}); + if sessionless { + sessionless_mcp_body("tools/call", params) + } else { + serde_json::json!({ + "jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": params + }) + .to_string() + } + }; + let original = body_for("read_status", serde_json::json!({})); + let mut headers = format!("MCP-Protocol-Version: {version}\r\n"); + if sessionless { + headers.push_str("Mcp-Method: tools/call\r\nMcp-Name: read_status\r\n"); + } + for enforcement in ["enforce", "audit"] { + for tool_name in ["read_status", "delete_resource"] { + // A changed argument marker makes the allowed control + // prove that the replacement, not the original, arrived. + let replacement = + body_for(tool_name, serde_json::json!({"rewritten": true})); + assert_ne!(original.len(), replacement.len()); + let invocations = Arc::new(AtomicUsize::new(0)); + let data = format!( + r#" +network_middlewares: + rewriter: + middleware: test/rewriter + on_error: fail_closed + endpoints: + include: ["mcp.example.test"] +network_policies: + mcp_api: + name: mcp_api + endpoints: + - host: mcp.example.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: {enforcement} + mcp: + versions: {configured_versions} + rules: + - allow: + method: tools/call + tool: read_status + deny_rules: + - method: tools/call + tool: delete_resource + binaries: + - {{ path: /usr/bin/python3 }} +"# + ); + let engine = OpaEngine::from_strings(TEST_POLICY, &data).unwrap(); + engine.set_middleware_runner_for_tests( + openshell_supervisor_middleware::ChainRunner::new(Arc::new( + McpToolReplacingService { + replacement: replacement.as_bytes().to_vec(), + tool_name, + sessionless, + invocations: Arc::clone(&invocations), + }, + )), + ); + let (response, forwarded) = run_mcp_relay_case( + mcp_relay_context_from_engine(engine), + route_selected, + &headers, + &original, + "HTTP/1.1 204 No Content\r\nContent-Length: 0\r\nConnection: close\r\n\r\n", + ) + .await; + assert_eq!(invocations.load(Ordering::SeqCst), 1); + if tool_name == "delete_resource" && enforcement == "enforce" { + assert_middleware_failure_response(&response, "mcp_api"); + assert!( + forwarded.is_empty(), + "rewritten denied tool reached upstream" + ); + continue; + } + // Audit permits policy denials, while final revision and + // metadata checks still apply to the rewritten request. + assert!( + response.starts_with("HTTP/1.1 204 No Content"), + "{response}" + ); + let forwarded = String::from_utf8(forwarded).unwrap(); + let (header, body) = forwarded.split_once("\r\n\r\n").unwrap(); + assert_eq!(body, replacement); + let content_length = header + .lines() + .find_map(|line| { + let (name, value) = line.split_once(':')?; + name.eq_ignore_ascii_case("content-length") + .then(|| value.trim().parse::().unwrap()) + }) + .expect("rewritten request includes Content-Length"); + assert_eq!(content_length, replacement.len()); + if sessionless { + let names = header + .lines() + .filter_map(|line| { + let (name, value) = line.split_once(':')?; + name.eq_ignore_ascii_case("mcp-name").then(|| value.trim()) + }) + .collect::>(); + assert_eq!(names, [tool_name]); + } + } + } + } + } + } + + #[tokio::test] + async fn mcp_march_batches_authorize_every_member_before_forwarding() { + let call = |id, name| { + serde_json::json!({ + "jsonrpc": "2.0", "id": id, "method": "tools/call", + "params": {"name": name, "arguments": {}} + }) + }; + let allowed = call(1, "read_status"); + let denied = call(2, "delete_resource"); + let malformed = serde_json::json!({ + "jsonrpc": "2.0", "id": 2, "method": "tools/call", + "params": {"name": 7, "arguments": {}} + }); + let cases = [ + ( + "allowed", + serde_json::json!([allowed, call(2, "read_status")]), + false, + false, + ), + ( + "deny last", + serde_json::json!([allowed, denied]), + true, + false, + ), + ( + "deny first", + serde_json::json!([denied, allowed]), + true, + false, + ), + ( + "malformed last", + serde_json::json!([allowed, malformed]), + false, + true, + ), + ]; + for route_selected in [false, true] { + for enforcement in ["enforce", "audit"] { + let data = format!( + r#" +network_policies: + mcp_api: + name: mcp_api + endpoints: + - host: mcp.example.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: {enforcement} + mcp: + versions: ["2025-03-26"] + rules: + - allow: + method: tools/call + tool: read_status + deny_rules: + - method: tools/call + tool: delete_resource + binaries: + - {{ path: /usr/bin/python3 }} +"# + ); + for (case, members, policy_denied, malformed) in &cases { + let body = members.to_string(); + let (response, forwarded) = run_mcp_relay_case( + mcp_relay_context_from_data(&data), + route_selected, + "MCP-Protocol-Version: 2025-03-26\r\n", + &body, + "HTTP/1.1 204 No Content\r\nContent-Length: 0\r\nConnection: close\r\n\r\n", + ) + .await; + // Audit forwards policy denials, but never malformed MCP. + // Capturing the whole upstream exchange also catches partial + // forwarding of an allowed prefix before a later denial. + let should_forward = !*malformed && (!*policy_denied || enforcement == "audit"); + let status = if *malformed { + "400 Bad Request" + } else if should_forward { + "204 No Content" + } else { + "403 Forbidden" + }; + assert!( + response.starts_with(&format!("HTTP/1.1 {status}")), + "{case}, route_selected={route_selected}, {enforcement}: {response}" + ); + if should_forward { + assert!( + forwarded.ends_with(body.as_bytes()), + "{case}: batch changed" + ); + } else { + assert!( + forwarded.is_empty(), + "{case}: rejected batch reached upstream" + ); + } + if *malformed { + assert!(response.contains("invalid_mcp_request"), "{response}"); + } + } + } + } + } + + #[tokio::test] + async fn mcp_sessionless_relays_discovery_tools_extensions_and_subscription_sse() { + for route_selected in [false, true] { + for (method, params, name, sse) in [ + ("server/discover", serde_json::json!({}), None, false), + ( + "tools/call", + serde_json::json!({"name":"echo", "arguments":{}}), + Some("echo"), + false, + ), + ("vendor/inspect", serde_json::json!({}), None, false), + ( + "subscriptions/listen", + serde_json::json!({"notifications":{"toolsListChanged":true}}), + None, + true, + ), + ] { + let mut headers = + format!("MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: {method}\r\n"); + if let Some(name) = name { + write!(headers, "Mcp-Name: {name}\r\n").unwrap(); + } + let body = sessionless_mcp_body(method, params); + let content_type = if sse { + "text/event-stream" + } else { + "application/json" + }; + let response_body = if sse { + "event: message\ndata: {\"jsonrpc\":\"2.0\",\"method\":\"notifications/tools/list_changed\"}\n\n" + } else { + "{\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{}}" + }; + let upstream_response = format!( + "HTTP/1.1 200 OK\r\nContent-Type: {content_type}\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{response_body}", + response_body.len() + ); + let (response, forwarded) = + run_sessionless_mcp_relay(route_selected, &headers, &body, &upstream_response) + .await; + assert!( + response.starts_with("HTTP/1.1 200 OK"), + "{method}: {response}" + ); + assert!( + response.ends_with(response_body), + "{method}: response changed" + ); + let forwarded = String::from_utf8(forwarded).unwrap(); + assert!(forwarded.contains(&headers), "{method}: headers changed"); + assert!(forwarded.ends_with(&body), "{method}: body changed"); + } + } + } + + #[tokio::test] + async fn mcp_sessionless_relays_apply_metadata_and_method_policy() { + for route_selected in [false, true] { + for (method, params, header_method, name, status) in [ + ( + "tools/list", + serde_json::json!({}), + "server/discover", + None, + "400 Bad Request", + ), + ( + "tools/call", + serde_json::json!({"name":"echo"}), + "tools/call", + None, + "400 Bad Request", + ), + ( + "tools/call", + serde_json::json!({"name":"blocked"}), + "tools/call", + Some("blocked"), + "403 Forbidden", + ), + ( + "vendor/unlisted", + serde_json::json!({}), + "vendor/unlisted", + None, + "403 Forbidden", + ), + ] { + let mut headers = + format!("MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: {header_method}\r\n"); + if let Some(name) = name { + write!(headers, "Mcp-Name: {name}\r\n").unwrap(); + } + let body = sessionless_mcp_body(method, params); + let (response, forwarded) = + run_sessionless_mcp_relay(route_selected, &headers, &body, "").await; + assert!( + response.starts_with(&format!("HTTP/1.1 {status}")), + "{method}: {response}" + ); + assert!( + forwarded.is_empty(), + "{method}: rejected request was forwarded" + ); + } + } + } + + #[tokio::test] + async fn final_mcp_sessionless_check_validates_transformed_headers_and_body() { + let (config, _, ctx) = mcp_sessionless_test_relay_context(); + for (header_method, valid) in [("tools/list", true), ("server/discover", false)] { + let body = sessionless_mcp_body("tools/list", serde_json::json!({})); + let request = crate::l7::provider::L7Request { + action: "POST".to_string(), + target: "/mcp".to_string(), + query_params: std::collections::HashMap::new(), + raw_header: format!("POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nMCP-Protocol-Version: 2026-07-28\r\nMcp-Method: {header_method}\r\nContent-Length: {}\r\n\r\n{body}", body.len()).into_bytes(), + body_length: crate::l7::provider::BodyLength::ContentLength(body.len() as u64), + }; + let mut response = Vec::new(); + let allowed = enforce_final_mcp_protocol_version( &config, - tunnel_engine, - &mut relay_client, - &mut relay_upstream, + &request, + &mut response, &ctx, + "/mcp", + None, ) .await - }); - - let body = br#"{"jsonrpc":"2.0","id":1,"method":"initialize"}"#; - let request = format!( - "POST /rpc HTTP/1.1\r\nHost: jsonrpc.example.test:8000\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n\r\n", - body.len() - ); - app.write_all(request.as_bytes()).await.unwrap(); - app.write_all(body).await.unwrap(); - - let mut upstream_bytes = Vec::new(); - let mut upstream_buf = [0u8; 1024]; - loop { - let n = tokio::time::timeout( - std::time::Duration::from_secs(1), - upstream.read(&mut upstream_buf), - ) - .await - .expect("allowed JSON-RPC request should reach upstream") .unwrap(); - assert_ne!(n, 0, "upstream closed before JSON-RPC body arrived"); - upstream_bytes.extend_from_slice(&upstream_buf[..n]); - if String::from_utf8_lossy(&upstream_bytes).contains(r#""method":"initialize""#) { - break; + assert_eq!(allowed, valid); + if !valid { + assert!( + String::from_utf8(response) + .unwrap() + .contains("invalid_mcp_request_metadata") + ); } } - let upstream_request = String::from_utf8_lossy(&upstream_bytes); - assert!(upstream_request.starts_with("POST /rpc HTTP/1.1")); - assert!(upstream_request.contains(r#""method":"initialize""#)); - - upstream - .write_all( - b"HTTP/1.1 200 OK\r\nContent-Length: 36\r\nConnection: close\r\n\r\n{\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{}}", - ) - .await - .unwrap(); - - let mut response = [0u8; 512]; - let n = tokio::time::timeout(std::time::Duration::from_secs(1), app.read(&mut response)) - .await - .expect("upstream response should reach client") - .unwrap(); - assert!(String::from_utf8_lossy(&response[..n]).contains("200 OK")); - - drop(app); - tokio::time::timeout(std::time::Duration::from_secs(1), relay) - .await - .expect("relay should complete") - .unwrap() - .unwrap(); } #[tokio::test] - async fn mcp_relay_forwards_standalone_initialize_without_version_header() { - let (config, tunnel_engine, ctx) = mcp_test_relay_context(); - let (mut app, mut relay_client) = tokio::io::duplex(8192); - let (mut relay_upstream, mut upstream) = tokio::io::duplex(8192); - let relay = tokio::spawn(async move { - relay_with_inspection( + async fn mcp_sessionless_get_advertises_post_in_method_rejection() { + let (config, _, ctx) = mcp_sessionless_test_relay_context(); + let request = crate::l7::provider::L7Request { + action: "GET".to_string(), + target: "/mcp".to_string(), + query_params: std::collections::HashMap::new(), + raw_header: b"GET /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nMCP-Protocol-Version: 2026-07-28\r\nAccept: text/event-stream\r\n\r\n".to_vec(), + body_length: crate::l7::provider::BodyLength::None, + }; + let mut response = Vec::new(); + assert!( + !enforce_final_mcp_protocol_version( &config, - tunnel_engine, - &mut relay_client, - &mut relay_upstream, + &request, + &mut response, &ctx, + "/mcp", + None ) .await - }); - - let body = br#"{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"test","version":"1"}}}"#; - let request = format!( - "POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\nContent-Length: {}\r\nConnection: close\r\n\r\n", - body.len() + .unwrap() ); - app.write_all(request.as_bytes()).await.unwrap(); - app.write_all(body).await.unwrap(); + let response = String::from_utf8(response).unwrap(); + assert!(response.starts_with("HTTP/1.1 405 Method Not Allowed\r\n")); + assert!(response.contains("\r\nAllow: POST\r\n")); + } - let mut upstream_bytes = vec![0; 2048]; - let count = tokio::time::timeout( - std::time::Duration::from_secs(1), - upstream.read(&mut upstream_bytes), + #[tokio::test] + async fn mcp_sessionless_rest_revalidates_outgoing_metadata_before_write() { + let (config, _, ctx) = mcp_sessionless_test_relay_context(); + let body = sessionless_mcp_body("tools/list", serde_json::json!({})); + let request = crate::l7::provider::L7Request { + action: "POST".to_string(), + target: "/mcp".to_string(), + query_params: std::collections::HashMap::new(), + raw_header: format!("POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nMCP-Protocol-Version: 2026-07-28\r\nMcp-Method: server/discover\r\nAuthorization: Bearer fixture\r\nContent-Length: {}\r\n\r\n{body}", body.len()).into_bytes(), + body_length: crate::l7::provider::BodyLength::ContentLength(body.len() as u64), + }; + let (mut client, mut relay_client) = tokio::io::duplex(2048); + let (mut relay_upstream, mut upstream) = tokio::io::duplex(2048); + let outcome = crate::l7::rest::relay_http_request_with_options_guarded( + &request, + &mut relay_client, + &mut relay_upstream, + crate::l7::rest::RelayRequestOptions { + mcp_request_validation: Some(crate::l7::rest::McpRequestValidation { + config: &config, + ctx: &ctx, + redacted_target: "/mcp", + }), + ..Default::default() + }, ) .await - .expect("standalone initialize should reach upstream") .unwrap(); - let upstream_request = String::from_utf8_lossy(&upstream_bytes[..count]); - assert!(upstream_request.contains(r#""method":"initialize""#)); - - upstream - .write_all( - b"HTTP/1.1 200 OK\r\nContent-Length: 36\r\nConnection: close\r\n\r\n{\"jsonrpc\":\"2.0\",\"id\":1,\"result\":{}}", - ) - .await - .unwrap(); - let mut response = [0; 512]; - let count = - tokio::time::timeout(std::time::Duration::from_secs(1), app.read(&mut response)) - .await - .expect("initialize response should reach client") - .unwrap(); - assert!(String::from_utf8_lossy(&response[..count]).contains("200 OK")); - - drop(app); - tokio::time::timeout(std::time::Duration::from_secs(1), relay) - .await - .expect("relay should complete") - .unwrap() - .unwrap(); + assert!(matches!(outcome, RelayOutcome::Consumed)); + drop(relay_client); + drop(relay_upstream); + let mut response = String::new(); + client.read_to_string(&mut response).await.unwrap(); + assert!(response.contains("invalid_mcp_request_metadata")); + let mut forwarded = Vec::new(); + upstream.read_to_end(&mut forwarded).await.unwrap(); + assert!(forwarded.is_empty()); } - async fn run_rejected_mcp_version_request( + async fn run_rejected_mcp_request( route_selected: bool, version_headers: &str, + body: &[u8], ) -> (String, Vec) { let (config, tunnel_engine, ctx) = mcp_test_relay_context(); let (mut app, mut relay_client) = tokio::io::duplex(8192); @@ -9458,7 +11329,6 @@ network_policies: } }); - let body = br#"{"jsonrpc":"2.0","id":7,"result":{"ok":true}}"#; let request = format!( "POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\n{version_headers}Content-Length: {}\r\nConnection: close\r\n\r\n", body.len() @@ -9496,38 +11366,56 @@ network_policies: #[tokio::test] async fn mcp_relay_rejects_invalid_disallowed_and_missing_versions_without_forwarding() { - for (headers, status, code) in [ + for (headers, status, code, remedy) in [ ( - "MCP-Protocol-Version: 2026-07-28\r\n", + "MCP-Protocol-Version: 2026-07-29\r\n", "400 Bad Request", "unsupported_mcp_protocol_version", + "use a supported client/server revision permitted by mcp.versions", ), ( "MCP-Protocol-Version: 2025-11-25\r\nMCP-Protocol-Version: 2025-11-25\r\n", "400 Bad Request", "invalid_mcp_protocol_version_header", + "send exactly one revision", ), ( "MCP-Protocol-Version: 2025-06-18\r\n", "403 Forbidden", "mcp_protocol_version_not_allowed", + "selected MCP revision 2025-06-18 from MCP-Protocol-Version", + ), + ( + "", + "403 Forbidden", + "mcp_protocol_version_not_allowed", + "missing MCP-Protocol-Version header fallback; send the client/server revision explicitly", ), - ("", "403 Forbidden", "mcp_protocol_version_not_allowed"), ] { - let (response, forwarded) = run_rejected_mcp_version_request(false, headers).await; + let (response, forwarded) = run_rejected_mcp_request( + false, + headers, + br#"{"jsonrpc":"2.0","id":7,"result":{"ok":true}}"#, + ) + .await; assert!( response.starts_with(&format!("HTTP/1.1 {status}")), "{response}" ); assert!(response.contains(code), "{response}"); + assert!(response.contains(remedy), "{response}"); assert!(forwarded.is_empty(), "rejected request reached upstream"); } } #[tokio::test] async fn route_selected_mcp_relay_enforces_request_version_before_forwarding() { - let (response, forwarded) = - run_rejected_mcp_version_request(true, "MCP-Protocol-Version: 2026-07-28\r\n").await; + let (response, forwarded) = run_rejected_mcp_request( + true, + "MCP-Protocol-Version: 2026-07-29\r\n", + br#"{"jsonrpc":"2.0","id":7,"result":{"ok":true}}"#, + ) + .await; assert!( response.starts_with("HTTP/1.1 400 Bad Request"), @@ -9553,10 +11441,15 @@ network_policies: "invalid_mcp_protocol_version_header", ), ( - "MCP-Protocol-Version: 2026-07-28\r\n", + "MCP-Protocol-Version: 2099-01-01\r\n", "400 Bad Request", "unsupported_mcp_protocol_version", ), + ( + "MCP-Protocol-Version: 2026-07-28\r\n", + "403 Forbidden", + "mcp_protocol_version_not_allowed", + ), ( "MCP-Protocol-Version: 2025-06-18\r\n", "403 Forbidden", @@ -9750,6 +11643,99 @@ network_policies: } } + #[tokio::test] + async fn endpoint_observation_records_typed_mcp_rejections_before_delivery() { + use openshell_core::endpoint_status::EndpointStatusCommand; + + for (method, header_method, body, status, response_code) in [ + ( + "POST", + "tools/call", + sessionless_mcp_body("tools/call", serde_json::json!({})), + "400 Bad Request", + "invalid_mcp_request", + ), + ( + "POST", + "server/discover", + sessionless_mcp_body("tools/list", serde_json::json!({})), + "400 Bad Request", + "invalid_mcp_request_metadata", + ), + ( + "GET", + "tools/list", + String::new(), + "405 Method Not Allowed", + "mcp_http_method_not_allowed", + ), + ] { + for disconnect_client in [false, true] { + let (mut config, _, mut ctx) = mcp_sessionless_test_relay_context(); + let mut receiver = install_mcp_test_observation(&mut config, &mut ctx).await; + config.mcp_versions = vec![openshell_core::mcp::McpProtocolVersion::V2026_07_28]; + let observer = + EndpointObserver::begin(ctx.endpoint_observation_tx.as_ref(), &config) + .expect("begin typed rejection observation"); + let request = crate::l7::provider::L7Request { + action: method.into(), + target: "/mcp".into(), + query_params: TestHashMap::new(), + raw_header: format!( + "{method} /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nMCP-Protocol-Version: 2026-07-28\r\nMcp-Method: {header_method}\r\nContent-Type: application/json\r\nContent-Length: {}\r\n\r\n{body}", + body.len() + ) + .into_bytes(), + body_length: crate::l7::provider::BodyLength::ContentLength(body.len() as u64), + }; + let (mut client, app) = tokio::io::duplex(2048); + let mut app = Some(app); + if disconnect_client { + // Observation must survive a failed write of the denial response. + drop(app.take()); + } + let result = enforce_final_mcp_protocol_version( + &config, + &request, + &mut client, + &ctx, + "/mcp", + Some(&observer), + ) + .await; + if disconnect_client { + assert!( + result.is_err(), + "{response_code}: denial delivery must fail" + ); + } else { + assert!(!result.expect("typed request rejection")); + } + drop(client); + if let Some(mut app) = app { + let mut response = String::new(); + app.read_to_string(&mut response).await.unwrap(); + assert!( + response.starts_with(&format!("HTTP/1.1 {status}")), + "{response}" + ); + assert!( + response.contains(&format!("\"{response_code}\"")), + "{response}" + ); + } + assert!(matches!( + receiver.try_recv().expect("typed rejection observation"), + EndpointStatusCommand::Observe { + result: EndpointResult::PolicyDenied, + .. + } + )); + assert!(receiver.try_recv().is_err(), "one result per exchange"); + } + } + } + async fn install_mcp_test_observation( config: &mut L7EndpointConfig, ctx: &mut L7EvalContext, @@ -9797,9 +11783,9 @@ network_policies: let observer = EndpointObserver::begin(ctx.endpoint_observation_tx.as_ref(), &config) .expect("begin transformed request observation"); - let buffered_request = |body: &[u8]| { + let buffered_request = |body: &[u8], headers: &str| { let mut raw_header = format!( - "POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\n{version_headers}Content-Length: {}\r\nConnection: close\r\n\r\n", + "POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\n{headers}Content-Length: {}\r\nConnection: close\r\n\r\n", body.len() ) .into_bytes(); @@ -9816,6 +11802,7 @@ network_policies: }; let initialize = buffered_request( br#"{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"test","version":"1"}}}"#, + "", ); let (mut client, app) = tokio::io::duplex(2048); assert!( @@ -9834,8 +11821,10 @@ network_policies: receiver.try_recv().is_err(), "allowed request is not a terminal result" ); - let rewritten = - buffered_request(br#"{"jsonrpc":"2.0","id":2,"method":"tools/list"}"#); + let rewritten = buffered_request( + br#"{"jsonrpc":"2.0","id":2,"method":"tools/list"}"#, + version_headers, + ); let mut app = Some(app); if disconnect_client { // Dropping the peer makes the final rejection undeliverable. @@ -9882,6 +11871,58 @@ network_policies: } } + #[tokio::test] + async fn mcp_relay_rejects_body_outside_selected_profile_before_policy() { + let body = br#"[ + {"jsonrpc":"2.0","id":1,"method":"tools/list"}, + {"jsonrpc":"2.0","id":2,"method":"tools/list"} + ]"#; + let (response, forwarded) = + run_rejected_mcp_request(false, "MCP-Protocol-Version: 2025-11-25\r\n", body).await; + + assert!( + response.starts_with("HTTP/1.1 400 Bad Request"), + "{response}" + ); + assert!(response.contains("invalid_mcp_request"), "{response}"); + assert!( + response.contains("does not permit top-level JSON-RPC batches"), + "{response}" + ); + assert!(response.contains("send each JSON-RPC message in a separate request")); + assert!(response.contains("selected MCP revision 2025-11-25 from MCP-Protocol-Version")); + assert!( + forwarded.is_empty(), + "profile-invalid request reached upstream" + ); + } + + #[tokio::test] + async fn mcp_relay_explains_unavailable_method_without_reflecting_params() { + // tasks/update is known to the parser but absent from the selected + // core revision. The diagnostic must not suggest adding an allow rule. + let body = br#"{"jsonrpc":"2.0","id":1,"method":"tasks/update","params":{"taskId":"private-task-marker","inputResponses":{"secret":"private-argument-marker"}}}"#; + for route_selected in [false, true] { + let (response, forwarded) = run_rejected_mcp_request( + route_selected, + "MCP-Protocol-Version: 2025-11-25\r\n", + body, + ) + .await; + assert!( + response.starts_with("HTTP/1.1 400 Bad Request"), + "{response}" + ); + assert!(response.contains("invalid_mcp_request"), "{response}"); + assert!(response.contains("`tasks/update` is unavailable in revision 2025-11-25")); + assert!(response.contains("allow_all_known_mcp_methods cannot enable")); + assert!(response.contains("from MCP-Protocol-Version")); + assert!(!response.contains("private-task-marker")); + assert!(!response.contains("private-argument-marker")); + assert!(forwarded.is_empty(), "unavailable method reached upstream"); + } + } + #[tokio::test] async fn final_mcp_version_check_reclassifies_a_rewritten_initialize_body() { let (config, _, ctx) = mcp_test_relay_context(); diff --git a/crates/openshell-supervisor-network/src/l7/rest.rs b/crates/openshell-supervisor-network/src/l7/rest.rs index 4592600231..cb7ed28e78 100644 --- a/crates/openshell-supervisor-network/src/l7/rest.rs +++ b/crates/openshell-supervisor-network/src/l7/rest.rs @@ -84,6 +84,9 @@ const HTTP_METHOD_PREFIXES: &[&[u8]] = &[ pub(crate) const HTTP2_PRIOR_KNOWLEDGE_PREFACE: &[u8] = b"PRI * HTTP/2.0\r\n\r\nSM\r\n\r\n"; pub(crate) const UNSUPPORTED_H2C_UPGRADE_DETAIL: &str = "HTTP/2 cleartext upgrade (h2c) is not supported for L7-inspected endpoints"; +pub(crate) const UNSUPPORTED_JSONRPC_UPGRADE_DETAIL: &str = + "HTTP upgrade is not supported for JSON-RPC or MCP endpoints"; +pub(crate) const UNSUPPORTED_GRAPHQL_UPGRADE_DETAIL: &str = "HTTP upgrade is not supported for GraphQL endpoints; serve GraphQL over WebSocket from a separate protocol: websocket endpoint on another path or port"; const MIN_HTTP2_PREFACE_DETECTION_BYTES: usize = 8; /// Idle timeout for `relay_until_eof`. If no data arrives within this window @@ -745,6 +748,7 @@ where RelayRequestOptions { resolver, body_classifier: None, + mcp_request_validation: None, credential_generation: None, generation_guard, websocket_extensions: WebSocketExtensionMode::Preserve, @@ -771,6 +775,8 @@ pub(crate) enum WebSocketExtensionMode { pub(crate) struct RelayRequestOptions<'a> { pub(crate) resolver: Option<&'a SecretResolver>, pub(crate) body_classifier: Option<&'a openshell_core::secrets::body::BodyCredentialClassifier>, + /// Revalidate buffered MCP requests after header transformations. + pub(crate) mcp_request_validation: Option>, pub(crate) credential_generation: Option>, pub(crate) generation_guard: Option<&'a PolicyGenerationGuard>, pub(crate) websocket_extensions: WebSocketExtensionMode, @@ -783,6 +789,14 @@ pub(crate) struct RelayRequestOptions<'a> { pub(crate) port: u16, } +/// Policy and logging context for checking the MCP request sent upstream. +#[derive(Clone, Copy)] +pub(crate) struct McpRequestValidation<'a> { + pub(crate) config: &'a crate::l7::L7EndpointConfig, + pub(crate) ctx: &'a crate::l7::relay::L7EvalContext, + pub(crate) redacted_target: &'a str, +} + #[derive(Clone, Copy)] pub(crate) struct CredentialGenerationGuard<'a> { state: &'a openshell_core::provider_credentials::ProviderCredentialState, @@ -915,6 +929,33 @@ where let rewrite_result = rewrite_http_header_block(&header_bytes, options.resolver).map_err(miette::Report::new)?; + if let Some(validation) = options.mcp_request_validation { + // Header credential resolution and hop-by-hop cleanup must finish + // before checking MCP mirrors. MCP bodies are already fully buffered + // and are not eligible for credential body rewriting. + let mut raw_header = rewrite_result.rewritten.clone(); + raw_header.extend_from_slice(&req.raw_header[header_end..]); + let outgoing = L7Request { + action: req.action.clone(), + target: req.target.clone(), + query_params: req.query_params.clone(), + raw_header, + body_length: req.body_length, + }; + if !crate::l7::relay::enforce_final_mcp_protocol_version( + validation.config, + &outgoing, + client, + validation.ctx, + validation.redacted_target, + observer, + ) + .await? + { + return Ok(RelayOutcome::Consumed); + } + } + if let Some(guard) = options.generation_guard { guard.ensure_current()?; } @@ -2393,6 +2434,71 @@ pub(crate) fn request_is_h2c_upgrade(raw_header: &[u8]) -> bool { upgrade_h2c && connection_upgrade } +/// Returns why an L7 endpoint using `protocol` must refuse this request's +/// upgrade, or `None` when the request may continue. +/// +/// Every inspected protocol refuses h2c. Protocols named by +/// `upgrade_refusal_for_protocol` refuse every request that carries an +/// `Upgrade` header. Callers apply this before the L7 policy decision and +/// regardless of enforcement mode, because an upgrade would end inspection +/// rather than break a rule that audit mode could log. +pub(crate) fn unsupported_upgrade_detail( + raw_header: &[u8], + protocol: crate::l7::L7Protocol, +) -> Option<&'static str> { + if request_is_h2c_upgrade(raw_header) { + return Some(UNSUPPORTED_H2C_UPGRADE_DETAIL); + } + let refusal = upgrade_refusal_for_protocol(protocol)?; + request_has_upgrade_header(raw_header).then_some(refusal) +} + +/// Returns the refusal detail for a protocol whose policy applies only to +/// individual HTTP requests, or `None` for a protocol that may relay an +/// allowed upgrade. +/// +/// JSON-RPC, MCP and GraphQL rules inspect each HTTP request body or query. +/// After an upgrade the relay would copy frames that no rule of these +/// protocols evaluates, so they never upgrade; GraphQL over WebSocket is +/// served by separate `protocol: websocket` endpoints with GraphQL operation +/// rules. REST and WebSocket endpoints relay allowed upgrades, and SQL +/// endpoints keep their existing upgrade behavior. The match is exhaustive so +/// a new protocol must choose. +pub(crate) fn upgrade_refusal_for_protocol( + protocol: crate::l7::L7Protocol, +) -> Option<&'static str> { + use crate::l7::L7Protocol; + match protocol { + L7Protocol::JsonRpc | L7Protocol::Mcp => Some(UNSUPPORTED_JSONRPC_UPGRADE_DETAIL), + L7Protocol::Graphql => Some(UNSUPPORTED_GRAPHQL_UPGRADE_DETAIL), + L7Protocol::Rest | L7Protocol::Websocket | L7Protocol::Sql => None, + } +} + +/// Returns true when a request carries an `Upgrade` header, whatever its +/// value. +/// +/// Both relay checks that can lead to a protocol switch require that header: +/// `request_is_websocket_upgrade`, which decides whether upgrade headers are +/// forwarded, and `client_requested_upgrade`, which decides whether an +/// upstream `101` may reach the client. A refusal based on this test therefore +/// covers every request either one treats as an upgrade. Headers that are not +/// UTF-8 return false; the shared relay rejects them before forwarding. +fn request_has_upgrade_header(raw_header: &[u8]) -> bool { + let header_end = raw_header + .windows(4) + .position(|w| w == b"\r\n\r\n") + .map_or(raw_header.len(), |p| p + 4); + let Ok(header_str) = std::str::from_utf8(&raw_header[..header_end]) else { + return false; + }; + + header_str.lines().skip(1).any(|line| { + line.split_once(':') + .is_some_and(|(name, _)| name.trim().eq_ignore_ascii_case("upgrade")) + }) +} + fn rewrite_websocket_extensions_for_mode( raw_header: &[u8], mode: WebSocketExtensionMode, @@ -2784,13 +2890,27 @@ pub(crate) async fn send_json_response( body: serde_json::Value, client: &mut C, status: &str, +) -> Result<()> { + send_json_response_with_allow(policy_name, body, client, status, None).await +} + +/// Send a JSON response, including the required `Allow` field for an HTTP 405. +/// The allowed methods are supplied by the protocol adapter, never the peer. +pub(crate) async fn send_json_response_with_allow( + policy_name: &str, + body: serde_json::Value, + client: &mut C, + status: &str, + allowed_methods: Option<&'static str>, ) -> Result<()> { let body_bytes = body.to_string(); + let allow = allowed_methods.map_or_else(String::new, |methods| format!("Allow: {methods}\r\n")); let response = format!( "HTTP/1.1 {status}\r\n\ Content-Type: application/json\r\n\ Content-Length: {}\r\n\ X-OpenShell-Policy: {}\r\n\ + {allow}\ Connection: close\r\n\ \r\n\ {}", @@ -6027,6 +6147,95 @@ mod tests { (observer, receiver) } + #[tokio::test] + async fn endpoint_observation_records_mcp_denial_after_header_cleanup() { + for disconnect_client in [false, true] { + let (observer, mut receiver) = test_endpoint_observer().await; + let config = crate::l7::parse_l7_config( + ®orus::Value::from_json_str( + r#"{"protocol":"mcp","mcp_versions":["2025-11-25"]}"#, + ) + .expect("parse MCP config JSON"), + ) + .expect("parse MCP config"); + let ctx = crate::l7::relay::L7EvalContext { + host: "mcp.example.test".into(), + port: 8000, + policy_name: "mcp-policy".into(), + ..Default::default() + }; + let body = r#"{"jsonrpc":"2.0","id":1,"method":"tools/list"}"#; + // The incoming version is allowed, but Connection nominates it for + // removal. The final gate must observe the outgoing request's denial. + let raw_request = format!( + "POST /mcp HTTP/1.1\r\nHost: mcp.example.test:8000\r\nContent-Type: application/json\r\nMCP-Protocol-Version: 2025-11-25\r\nConnection: close, MCP-Protocol-Version\r\nContent-Length: {}\r\n\r\n{body}", + body.len(), + ); + let request = + request_from_buffered_http("POST", "/mcp", "/mcp", raw_request.into_bytes()) + .expect("parse buffered MCP request"); + let (mut client, peer) = tokio::io::duplex(4096); + let mut peer = Some(peer); + if disconnect_client { + drop(peer.take()); + } + let (mut upstream, mut upstream_peer) = tokio::io::duplex(4096); + let result = tokio::time::timeout( + std::time::Duration::from_secs(1), + relay_http_request_with_response_middleware_guarded_observed( + &request, + &mut client, + &mut upstream, + RelayRequestOptions { + mcp_request_validation: Some(McpRequestValidation { + config: &config, + ctx: &ctx, + redacted_target: "/mcp", + }), + ..Default::default() + }, + None, + Some(&observer), + ), + ) + .await + .expect("reject before upstream I/O"); + if disconnect_client { + assert!( + result.is_err(), + "denial delivery must fail after disconnect" + ); + } else { + assert!(matches!( + result.expect("deliver denial"), + RelayOutcome::Consumed + )); + } + assert!(matches!( + receiver.try_recv().expect("final MCP denial observation"), + EndpointStatusCommand::Observe { + result: EndpointResult::PolicyDenied, + .. + } + )); + assert!(receiver.try_recv().is_err(), "one result per exchange"); + drop(client); + drop(upstream); + let mut sent = Vec::new(); + upstream_peer.read_to_end(&mut sent).await.unwrap(); + assert!(sent.is_empty(), "rejected request reached upstream"); + if let Some(mut peer) = peer { + let mut response = String::new(); + peer.read_to_string(&mut response).await.unwrap(); + assert!(response.starts_with("HTTP/1.1 403 Forbidden"), "{response}"); + assert!( + response.contains("mcp_protocol_version_not_allowed"), + "{response}" + ); + } + } + } + async fn assert_observed_response_result(response: &'static [u8], expected: EndpointResult) { let (observer, mut receiver) = test_endpoint_observer().await; let (mut upstream, mut upstream_peer) = tokio::io::duplex(4096); @@ -8897,6 +9106,175 @@ mod tests { assert!(!client_requested_upgrade(headers)); } + #[test] + fn unsupported_upgrade_detail_refuses_h2c_for_every_protocol() { + let raw = b"GET /api HTTP/1.1\r\nHost: example.com\r\nConnection: Upgrade, HTTP2-Settings\r\nUpgrade: h2c\r\nHTTP2-Settings: AAMAAABkAAQAAP__\r\n\r\n"; + for protocol in [ + crate::l7::L7Protocol::Rest, + crate::l7::L7Protocol::Websocket, + crate::l7::L7Protocol::Graphql, + crate::l7::L7Protocol::JsonRpc, + crate::l7::L7Protocol::Mcp, + ] { + assert_eq!( + unsupported_upgrade_detail(raw, protocol), + Some(UNSUPPORTED_H2C_UPGRADE_DETAIL), + "{protocol:?}" + ); + } + } + + #[test] + fn unsupported_upgrade_detail_refuses_any_upgrade_on_jsonrpc_family() { + let websocket = format!( + "GET /mcp HTTP/1.1\r\nHost: example.com\r\nAccept: text/event-stream\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: {VALID_WS_KEY}\r\nSec-WebSocket-Version: 13\r\n\r\n" + ); + // Any request that carries an `Upgrade` header is refused, including + // one without `Connection: upgrade` that the relay would not upgrade. + let requests: [&[u8]; 3] = [ + websocket.as_bytes(), + b"POST /mcp HTTP/1.1\r\nHost: example.com\r\nUpgrade: websocket\r\nConnection: keep-alive, upgrade\r\nContent-Length: 0\r\n\r\n", + b"GET /mcp HTTP/1.1\r\nHost: example.com\r\nUpgrade: custom\r\n\r\n", + ]; + for raw in requests { + for protocol in [crate::l7::L7Protocol::JsonRpc, crate::l7::L7Protocol::Mcp] { + assert_eq!( + unsupported_upgrade_detail(raw, protocol), + Some(UNSUPPORTED_JSONRPC_UPGRADE_DETAIL), + "{protocol:?}: {}", + String::from_utf8_lossy(raw) + ); + } + } + } + + #[test] + fn unsupported_upgrade_detail_allows_ordinary_requests_and_relaying_protocols() { + let websocket = format!( + "GET /ws HTTP/1.1\r\nHost: example.com\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: {VALID_WS_KEY}\r\nSec-WebSocket-Version: 13\r\n\r\n" + ); + for protocol in [ + crate::l7::L7Protocol::Rest, + crate::l7::L7Protocol::Websocket, + ] { + assert_eq!( + unsupported_upgrade_detail(websocket.as_bytes(), protocol), + None, + "{protocol:?}" + ); + } + + // Streamable HTTP requests, a look-alike header name, and a stray + // `Connection: upgrade` without an `Upgrade` header stay allowed; + // none of them can switch protocols. + let ordinary: [&[u8]; 4] = [ + b"POST /mcp HTTP/1.1\r\nHost: example.com\r\nContent-Type: application/json\r\nConnection: keep-alive\r\nContent-Length: 2\r\n\r\n{}", + b"GET /mcp HTTP/1.1\r\nHost: example.com\r\nAccept: text/event-stream\r\n\r\n", + b"GET /mcp HTTP/1.1\r\nHost: example.com\r\nUpgrade-Insecure-Requests: 1\r\n\r\n", + b"GET /mcp HTTP/1.1\r\nHost: example.com\r\nConnection: Upgrade\r\n\r\n", + ]; + for raw in ordinary { + for protocol in [ + crate::l7::L7Protocol::JsonRpc, + crate::l7::L7Protocol::Mcp, + crate::l7::L7Protocol::Graphql, + ] { + assert_eq!( + unsupported_upgrade_detail(raw, protocol), + None, + "{protocol:?}: {}", + String::from_utf8_lossy(raw) + ); + } + } + } + + #[test] + fn unsupported_upgrade_detail_refuses_upgrades_on_graphql() { + let requests = [ + format!( + "GET /graphql?query=%7Bviewer%7D HTTP/1.1\r\nHost: example.com\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: {VALID_WS_KEY}\r\nSec-WebSocket-Version: 13\r\n\r\n" + ), + format!( + "GET /graphql HTTP/1.1\r\nHost: example.com\r\nUpgrade: websocket\r\nConnection: Upgrade\r\nSec-WebSocket-Key: {VALID_WS_KEY}\r\nSec-WebSocket-Version: 13\r\nSec-WebSocket-Protocol: graphql-transport-ws\r\n\r\n" + ), + "POST /graphql HTTP/1.1\r\nHost: example.com\r\nUpgrade: custom\r\nConnection: upgrade\r\nContent-Length: 0\r\n\r\n" + .to_string(), + ]; + for raw in &requests { + assert_eq!( + unsupported_upgrade_detail(raw.as_bytes(), crate::l7::L7Protocol::Graphql), + Some(UNSUPPORTED_GRAPHQL_UPGRADE_DETAIL), + "{raw}" + ); + } + } + + #[test] + fn upgrade_refusal_for_protocol_names_every_per_request_protocol() { + use crate::l7::L7Protocol; + assert_eq!( + upgrade_refusal_for_protocol(L7Protocol::JsonRpc), + Some(UNSUPPORTED_JSONRPC_UPGRADE_DETAIL) + ); + assert_eq!( + upgrade_refusal_for_protocol(L7Protocol::Mcp), + Some(UNSUPPORTED_JSONRPC_UPGRADE_DETAIL) + ); + assert_eq!( + upgrade_refusal_for_protocol(L7Protocol::Graphql), + Some(UNSUPPORTED_GRAPHQL_UPGRADE_DETAIL) + ); + for protocol in [L7Protocol::Rest, L7Protocol::Websocket, L7Protocol::Sql] { + assert_eq!(upgrade_refusal_for_protocol(protocol), None, "{protocol:?}"); + } + } + + #[test] + fn unsupported_upgrade_detail_covers_every_relay_upgrade_check() { + // The refusal must be at least as broad as both relay checks that can + // lead to a protocol switch, across header spellings that pass + // ingress validation. + let valid_websocket = |head: &str| { + format!("{head}Sec-WebSocket-Key: {VALID_WS_KEY}\r\nSec-WebSocket-Version: 13\r\n\r\n") + }; + let requests = [ + valid_websocket( + "GET /mcp HTTP/1.1\r\nHost: example.com\r\nUPGRADE: WebSocket\r\nCONNECTION: UPGRADE\r\n", + ), + valid_websocket( + "GET /mcp HTTP/1.1\r\nHost: example.com\r\nConnection: keep-alive\r\nConnection: upgrade\r\nUpgrade: websocket\r\n", + ), + "GET /mcp HTTP/1.1\r\nHost: example.com\r\nUpgrade:\r\nConnection: upgrade\r\n\r\n" + .to_string(), + "GET /mcp HTTP/1.0\r\nUpgrade: websocket\r\nConnection:upgrade\r\n\r\n".to_string(), + "POST /mcp HTTP/1.1\r\nHost: example.com\r\nTransfer-Encoding: chunked\r\nUpgrade: custom\r\nConnection: close, Upgrade\r\n\r\n" + .to_string(), + "GET /mcp HTTP/1.1\r\nHost: example.com\r\nUpgrade: websocket\r\nUpgrade: h2c\r\nConnection: upgrade\r\n\r\n" + .to_string(), + ]; + for raw in &requests { + assert!( + validate_http_request_header_block(raw.as_bytes()).is_ok(), + "fixture must pass ingress validation: {raw}" + ); + assert!( + client_requested_upgrade(raw) || request_is_websocket_upgrade(raw.as_bytes()), + "fixture must be an upgrade to the relay: {raw}" + ); + for protocol in [ + crate::l7::L7Protocol::JsonRpc, + crate::l7::L7Protocol::Mcp, + crate::l7::L7Protocol::Graphql, + ] { + assert!( + unsupported_upgrade_detail(raw.as_bytes(), protocol).is_some(), + "{protocol:?} must refuse: {raw}" + ); + } + } + } + #[test] fn client_requested_upgrade_handles_comma_separated_connection() { let headers = "GET /ws HTTP/1.1\r\nHost: example.com\r\nUpgrade: websocket\r\nConnection: keep-alive, Upgrade\r\n\r\n"; diff --git a/crates/openshell-supervisor-network/src/opa.rs b/crates/openshell-supervisor-network/src/opa.rs index c883651aaf..74a14466da 100644 --- a/crates/openshell-supervisor-network/src/opa.rs +++ b/crates/openshell-supervisor-network/src/opa.rs @@ -16,6 +16,10 @@ use openshell_core::policy::{ use openshell_core::policy_identity::deterministic_policy_hash; use openshell_core::proto::SandboxPolicy as ProtoSandboxPolicy; use openshell_policy::{L7ConfigStanza, L7Protocol as PolicyL7Protocol, PolicyViolation}; +use openshell_policy_schema::{ + FilesystemPolicy as AuthoredFilesystemPolicy, LandlockPolicy as AuthoredLandlockPolicy, + ProcessPolicy as AuthoredProcessPolicy, +}; use openshell_supervisor_middleware::{ChainEntry, ChainRunner, MiddlewareRegistry}; use std::path::{Path, PathBuf}; use std::sync::{ @@ -25,6 +29,8 @@ use std::sync::{ use tokio::sync::watch; use tracing::info; +mod raw_schema; + /// Baked-in rego rules for OPA policy evaluation. /// These rules define the network access decision logic and static config /// passthroughs. They reference `data.sandbox.*` for policy data. @@ -925,22 +931,24 @@ impl OpaEngine { .lock() .map_err(|_| miette::miette!("OPA engine lock poisoned"))?; - // Query filesystem policy + // Evaluation errors can include authored Rego source. Regorus may + // evaluate other rules while resolving any one query, so report the + // static-settings operation without attributing it to a single rule. let fs_val = engine .eval_rule("data.openshell.sandbox.filesystem_policy".into()) - .map_err(|e| miette::miette!("{e}"))?; + .map_err(|_| miette::miette!("failed to evaluate static sandbox settings"))?; let filesystem = parse_filesystem_policy(&fs_val); // Query landlock policy let ll_val = engine .eval_rule("data.openshell.sandbox.landlock_policy".into()) - .map_err(|e| miette::miette!("{e}"))?; + .map_err(|_| miette::miette!("failed to evaluate static sandbox settings"))?; let landlock = parse_landlock_policy(&ll_val); // Query process policy let proc_val = engine .eval_rule("data.openshell.sandbox.process_policy".into()) - .map_err(|e| miette::miette!("{e}"))?; + .map_err(|_| miette::miette!("failed to evaluate static sandbox settings"))?; let process = parse_process_policy(&proc_val); Ok(SandboxConfig { @@ -1482,6 +1490,48 @@ fn validate_opa_object_array<'a>( Ok(entries) } +/// Validate static sections with the authored types while retaining OPA-only data. +/// Keep schema errors and their source chains out of load diagnostics because +/// they can contain authored keys, paths, or values. +fn validate_opa_static_settings(data: &mut serde_json::Value) -> Result<()> { + if let Some(filesystem) = data.get_mut("filesystem_policy") { + let settings: AuthoredFilesystemPolicy = serde_json::from_value(filesystem.clone()) + .map_err(|_| miette::miette!("invalid filesystem policy settings"))?; + openshell_policy::validate_filesystem_paths(&settings.read_only, &settings.read_write) + .map_err(|violations| { + miette::miette!(render_bounded_validation_diagnostics( + "invalid filesystem policy settings", + violations.iter().map(redacted_policy_violation_category), + )) + })?; + // A present stanza defaults to false; an absent stanza must stay absent + // so Rego's undefined result retains the runtime workdir default. + filesystem["include_workdir"] = settings.include_workdir.into(); + } + if let Some(landlock) = data.get_mut("landlock") { + let settings: AuthoredLandlockPolicy = serde_json::from_value(landlock.clone()) + .map_err(|_| miette::miette!("invalid Landlock policy settings"))?; + // Serde also accepts a map for a unit enum variant. Runtime consumers + // read a string, so retain the validated meaning in canonical form. + *landlock = serde_json::to_value(settings) + .map_err(|_| miette::miette!("failed to serialize Landlock policy settings"))?; + } + if let Some(process) = data.get("process") { + let settings = serde_json::from_value::(process.clone()) + .map_err(|_| miette::miette!("invalid process policy settings"))?; + // Omitted identities are resolved by the compute runtime. Explicit + // values follow the same non-root identity contract as typed policy. + for identity in [&settings.run_as_user, &settings.run_as_group] { + if !identity.is_empty() && !openshell_policy::is_valid_sandbox_identity(identity) { + return Err(miette::miette!( + "invalid process policy settings: invalid process identity" + )); + } + } + } + Ok(()) +} + /// Select a fixed category without formatting authored fields or nested reasons. /// Keep this match exhaustive so new validator variants require an explicit /// decision before their diagnostics can cross the supervisor load boundary. @@ -1597,6 +1647,8 @@ fn preprocess_yaml_data( ) })?; validate_opa_data_structure(&data)?; + raw_schema::validate_network_settings(&data)?; + validate_opa_static_settings(&mut data)?; inject_runtime_policy_data(&mut data, require_binary_identity); normalize_endpoint_protocols(&mut data); @@ -1811,13 +1863,8 @@ fn normalize_l7_config_alias( let Some(config) = ep.get(key).cloned() else { return; }; - if config.is_null() { - ep.remove(key); - if stanza == L7ConfigStanza::Mcp { - errors.push(format!("{loc}.{key}: mcp config must be an object")); - } - return; - } + // Explicit null must reach the canonical parser: it is invalid authored + // configuration, while an omitted stanza may select protocol defaults. match openshell_policy::l7_config_alias_runtime_fields(stanza, config) { Ok(fields) => { ep.remove(key); @@ -4256,6 +4303,442 @@ process: ); } + /// The typed schema gives an absent `filesystem_policy` the platform + /// default, `include_workdir: true`, and gives a present but empty stanza + /// `include_workdir: false`. Loading the same YAML directly into OPA must + /// agree in both cases, and versionless OPA data must follow the same rule + /// for a present empty stanza. + #[test] + fn yaml_and_proto_filesystem_policy_have_include_workdir_parity() { + for (data, expected) in [ + ("version: 1\n", true), + ("version: 1\nfilesystem_policy: {}\n", false), + ] { + let proto = openshell_policy::parse_sandbox_policy(data) + .expect("fixture must parse into the typed schema"); + let proto_config = OpaEngine::from_proto(&proto) + .expect("engine from protobuf") + .query_sandbox_config() + .expect("config from protobuf"); + let yaml_config = OpaEngine::from_strings(TEST_POLICY, data) + .expect("engine from YAML") + .query_sandbox_config() + .expect("config from YAML"); + assert_eq!( + proto_config.filesystem.include_workdir, expected, + "typed contract for {data:?}" + ); + assert_eq!( + yaml_config.filesystem.include_workdir, expected, + "raw OPA loading for {data:?}" + ); + } + + let versionless = OpaEngine::from_strings(TEST_POLICY, "filesystem_policy: {}\n") + .expect("versionless OPA data must load") + .query_sandbox_config() + .expect("config from versionless OPA data"); + assert!( + !versionless.filesystem.include_workdir, + "raw OPA loading of a versionless empty filesystem stanza" + ); + } + + #[test] + fn yaml_and_proto_landlock_policy_have_compatibility_parity() { + for (stanza, hard_requirement) in [ + ("{}", false), + ("{compatibility: best_effort}", false), + ("{compatibility: hard_requirement}", true), + ("{compatibility: {best_effort: null}}", false), + ("{compatibility: {hard_requirement: null}}", true), + ] { + let versionless = format!("landlock: {stanza}\n"); + let versioned = format!("version: 1\n{versionless}"); + let proto = openshell_policy::parse_sandbox_policy(&versioned) + .expect("valid typed Landlock representation"); + let typed = OpaEngine::from_proto(&proto) + .expect("typed engine") + .query_sandbox_config() + .expect("typed sandbox config"); + assert_eq!( + matches!( + typed.landlock.compatibility, + LandlockCompatibility::HardRequirement + ), + hard_requirement, + "typed contract for {stanza}" + ); + for data in [&versioned, &versionless] { + let raw = OpaEngine::from_strings(TEST_POLICY, data) + .expect("valid raw Landlock representation") + .query_sandbox_config() + .expect("raw sandbox config"); + assert_eq!( + matches!( + raw.landlock.compatibility, + LandlockCompatibility::HardRequirement + ), + hard_requirement, + "raw OPA loading for {data}" + ); + } + } + } + + /// Nested values that the typed schema rejects must also be rejected when + /// the same policy is loaded directly into OPA, with or without a + /// `version` key, instead of being replaced by defaults + /// (`include_workdir: true`, best-effort Landlock) or dropped. Each case + /// has a valid twin that both paths accept, so the test cannot pass by + /// rejecting valid input. + #[test] + fn raw_opa_loading_rejects_nested_values_the_typed_schema_rejects() { + const JSON_RPC_ENDPOINT: &str = r"network_policies: + rpc: + name: rpc + endpoints: + - host: jsonrpc.parity.test + port: 443 + path: /rpc + protocol: json-rpc + enforcement: enforce + json_rpc: JSON_RPC_OPTIONS + rules: + - allow: { method: status.get } + binaries: + - { path: /usr/bin/curl } +"; + // Each case is (name, invalid body, valid twin body). Bodies omit + // `version` so each invalid body is also loaded as versionless OPA data. + let cases = [ + ( + "string include_workdir", + "filesystem_policy: {include_workdir: \"false\"}\n".to_owned(), + "filesystem_policy: {include_workdir: false}\n".to_owned(), + ), + ( + "non-string read_only entry", + "filesystem_policy: {read_only: [7]}\n".to_owned(), + "filesystem_policy: {read_only: [\"/usr\"]}\n".to_owned(), + ), + ( + "non-string read_write entry", + "filesystem_policy: {read_write: [false]}\n".to_owned(), + "filesystem_policy: {read_write: [\"/tmp\"]}\n".to_owned(), + ), + ( + "unknown filesystem field", + "filesystem_policy: {private_typo: false}\n".to_owned(), + "filesystem_policy: {}\n".to_owned(), + ), + ( + "unknown Landlock compatibility", + "landlock: {compatibility: required}\n".to_owned(), + "landlock: {compatibility: hard_requirement}\n".to_owned(), + ), + ( + "unknown Landlock field", + "landlock: {private_typo: true}\n".to_owned(), + "landlock: {}\n".to_owned(), + ), + ( + "non-string process identity", + "process: {run_as_user: 7}\n".to_owned(), + "process: {run_as_user: sandbox}\n".to_owned(), + ), + ( + "null process group", + "process: {run_as_group: null}\n".to_owned(), + "process: {run_as_group: sandbox}\n".to_owned(), + ), + ( + "unknown process field", + "process: {private_typo: sandbox}\n".to_owned(), + "process: {}\n".to_owned(), + ), + ( + "explicit null json_rpc options", + JSON_RPC_ENDPOINT.replace("JSON_RPC_OPTIONS", "null"), + JSON_RPC_ENDPOINT.replace("JSON_RPC_OPTIONS", "{ max_body_bytes: 32768 }"), + ), + ]; + + // Check every fixture before failing, so one run reports each input the + // raw loader accepts instead of stopping at the first. + let mut accepted_by_raw_opa = Vec::new(); + for (case, invalid, valid) in &cases { + let valid = format!("version: 1\n{valid}"); + assert!( + openshell_policy::parse_sandbox_policy(&valid).is_ok(), + "typed schema must accept the valid {case} twin" + ); + assert!( + OpaEngine::from_strings(TEST_POLICY, &valid).is_ok(), + "raw OPA loading must accept the valid {case} twin" + ); + + let versioned = format!("version: 1\n{invalid}"); + assert!( + openshell_policy::parse_sandbox_policy(&versioned).is_err(), + "typed schema must reject the {case} fixture" + ); + for (form, data) in [("versioned", versioned.as_str()), ("versionless", invalid)] { + match OpaEngine::from_strings(TEST_POLICY, data) { + Ok(_) => accepted_by_raw_opa.push(format!("{case} ({form})")), + Err(error) => assert_safe_load_error(&error, &["private_typo"]), + } + } + } + + let hard_requirement = OpaEngine::from_strings( + TEST_POLICY, + "version: 1\nlandlock: {compatibility: hard_requirement}\n", + ) + .expect("valid Landlock twin must load") + .query_sandbox_config() + .expect("config from valid Landlock twin"); + assert!(matches!( + hard_requirement.landlock.compatibility, + LandlockCompatibility::HardRequirement + )); + + assert!( + accepted_by_raw_opa.is_empty(), + "raw OPA loading accepted fixtures that the typed schema rejects: {accepted_by_raw_opa:?}" + ); + } + + #[test] + fn raw_opa_network_leaves_reject_typed_schema_errors() { + let valid = opa_container_policy(); + let mut accepted = Vec::new(); + for (field, invalid, control) in [ + ( + "allow_encoded_slash", + serde_json::json!("true"), + serde_json::json!(true), + ), + ( + "websocket_credential_rewrite", + serde_json::json!("true"), + serde_json::json!(true), + ), + ("port", serde_json::json!("443"), serde_json::json!(443)), + ( + "deny_rules", + serde_json::json!([{"method": ["DELETE"], "path": "/admin/**"}]), + serde_json::json!([{"method": "DELETE", "path": "/admin/**"}]), + ), + ] { + let mut policy = valid.clone(); + policy["version"] = 1.into(); + policy["network_policies"]["admin"]["endpoints"][0][field] = control; + openshell_policy::parse_sandbox_policy(&policy.to_string()).expect("typed valid twin"); + OpaEngine::from_strings(TEST_POLICY, &policy.to_string()).expect("raw valid twin"); + policy["network_policies"]["admin"]["endpoints"][0][field] = invalid; + assert!( + openshell_policy::parse_sandbox_policy(&policy.to_string()).is_err(), + "typed {field}" + ); + for versioned in [true, false] { + if !versioned { + policy.as_object_mut().expect("object").remove("version"); + } + match OpaEngine::from_strings(TEST_POLICY, &policy.to_string()) { + Ok(_) => accepted.push(format!("{field} (versioned={versioned})")), + Err(error) => { + assert_safe_load_error(&error, &["admin.example.test", "/admin/**"]); + } + } + } + } + assert!( + accepted.is_empty(), + "raw loader accepted invalid network fields: {accepted:?}" + ); + } + + #[test] + fn raw_opa_static_semantics_match_typed_validation() { + let mut accepted = Vec::new(); + for invalid in [ + serde_json::json!({"process": {"run_as_user": "root"}}), + serde_json::json!({"process": {"run_as_group": "0"}}), + serde_json::json!({"process": {"run_as_user": "4294967295"}}), + serde_json::json!({"filesystem_policy": {"read_write": ["/"]}}), + serde_json::json!({"filesystem_policy": {"read_write": ["///"]}}), + serde_json::json!({"filesystem_policy": {"read_only": ["relative-private-path"]}}), + serde_json::json!({"filesystem_policy": {"read_only": ["/tmp/../private-path"]}}), + serde_json::json!({"filesystem_policy": {"read_only": [format!("/{}", "p".repeat(4096))]}}), + serde_json::json!({"filesystem_policy": {"read_only": vec!["/usr"; 257]}}), + ] { + let mut versioned = invalid.clone(); + versioned["version"] = 1.into(); + assert!( + openshell_policy::parse_sandbox_policy(&versioned.to_string()).map_or( + true, + |policy| openshell_policy::validate_sandbox_policy(&policy).is_err() + ), + "typed invalid static settings" + ); + for policy in [&invalid, &versioned] { + match OpaEngine::from_strings(TEST_POLICY, &policy.to_string()) { + Ok(_) => accepted.push(invalid.clone()), + Err(error) => { + assert_safe_load_error(&error, &["root", "private-path", "4294967295"]); + } + } + } + } + for valid in [ + serde_json::json!({"process": {}}), + serde_json::json!({"process": {"run_as_user": "sandbox", "run_as_group": "1"}}), + serde_json::json!({"process": {"run_as_user": "4294967294", "run_as_group": "sandbox"}}), + serde_json::json!({"filesystem_policy": {"read_only": ["/"], "read_write": ["/tmp"]}}), + serde_json::json!({"filesystem_policy": {"read_only": vec!["/usr"; 256]}}), + serde_json::json!({"filesystem_policy": {"read_only": [format!("/{}", "p".repeat(4095))]}}), + ] { + let mut versioned = valid.clone(); + versioned["version"] = 1.into(); + let typed = openshell_policy::parse_sandbox_policy(&versioned.to_string()) + .expect("typed static control"); + openshell_policy::validate_sandbox_policy(&typed).expect("valid typed static settings"); + for policy in [&valid, &versioned] { + OpaEngine::from_strings(TEST_POLICY, &policy.to_string()) + .expect("raw static control") + .query_sandbox_config() + .expect("static config"); + } + } + assert!( + accepted.is_empty(), + "raw loader accepted {} invalid static policies", + accepted.len() + ); + } + + #[test] + fn startup_evaluation_errors_discard_authored_rego_and_sources() { + let directory = tempfile::tempdir().expect("temporary policy directory"); + let rules_path = directory.path().join("private-rules.rego"); + let data_path = directory.path().join("private-data.yaml"); + std::fs::write(&data_path, "{}").expect("write data"); + for rule in ["filesystem_policy", "landlock_policy", "process_policy"] { + for repetitions in [1, 20] { + let marker = "private-eval-marker-".repeat(repetitions); + let rules = format!( + "package openshell.sandbox\n{rule} := {{\"value\": \"{marker}one\"}}\n{rule} := {{\"value\": \"{marker}two\"}}\n" + ); + std::fs::write(&rules_path, &rules).expect("write conflicting rules"); + for engine in [ + OpaEngine::from_strings(&rules, "{}").expect("conflict is an evaluation error"), + OpaEngine::from_files(&rules_path, &data_path).expect("file rules compile"), + ] { + let error = engine + .query_sandbox_config() + .err() + .expect("conflicting complete rules must fail"); + assert_safe_load_error( + &error, + &["private-eval-marker", "private-rules.rego", "value"], + ); + assert_eq!( + error.to_string(), + "failed to evaluate static sandbox settings" + ); + } + } + } + } + + #[test] + fn raw_opa_nested_validation_preserves_file_loads_and_rejected_reload() { + let directory = tempfile::tempdir().expect("temporary policy directory"); + let rules_path = directory.path().join("policy.rego"); + let data_path = directory.path().join("policy.yaml"); + std::fs::write(&rules_path, TEST_POLICY).expect("write policy rules"); + let valid = opa_container_policy(); + std::fs::write(&data_path, valid.to_string()).expect("write valid policy"); + let engine = OpaEngine::from_files(&rules_path, &data_path).expect("valid file policy"); + let proxy = OpaEngine::from_files_for_endpoint_only_proxy(&rules_path, &data_path, None) + .expect("valid endpoint-only policy"); + for loaded in [&engine, &proxy] { + assert!( + !loaded + .query_sandbox_config() + .expect("filesystem config") + .filesystem + .include_workdir, + "file loaders must retain the typed default for an empty filesystem stanza" + ); + } + let allowed = l7_input("admin.example.test", 443, "GET", "/public"); + let denied = l7_input("admin.example.test", 443, "DELETE", "/admin/users"); + let generation = engine.current_generation(); + assert!(eval_l7(&engine, &allowed)); + assert!(!eval_l7(&engine, &denied)); + + // Malformed nested settings must fail before replacing any installed + // decisions or notifying consumers of a new generation. + for (pointer, replacement) in [ + ( + "/filesystem_policy", + serde_json::json!({"include_workdir": "private-value-秘密".repeat(1024)}), + ), + ( + "/landlock", + serde_json::json!({"compatibility": "private-value-秘密".repeat(1024)}), + ), + ("/process", serde_json::json!({"run_as_user": 7})), + ("/process", serde_json::json!({"run_as_user": "root"})), + ( + "/filesystem_policy", + serde_json::json!({"read_write": ["/"]}), + ), + ( + "/network_policies/admin/endpoints/0/deny_rules/0/method", + serde_json::json!(["DELETE"]), + ), + ( + "/network_policies/admin/binaries/0/path", + serde_json::json!(["/usr/bin/curl"]), + ), + ] { + let mut malformed = valid.clone(); + *malformed + .pointer_mut(pointer) + .expect("existing static section") = replacement; + let data = malformed.to_string(); + std::fs::write(&data_path, &data).expect("write malformed policy"); + for error in [ + OpaEngine::from_files(&rules_path, &data_path) + .err() + .expect("file loader must reject malformed nested values"), + OpaEngine::from_files_for_endpoint_only_proxy(&rules_path, &data_path, None) + .err() + .expect("endpoint-only file loader must reject malformed nested values"), + engine + .reload(TEST_POLICY, &data) + .expect_err("reload must reject malformed nested values"), + ] { + assert_safe_load_error(&error, &["private-value", "秘密", "admin.example.test"]); + } + assert_eq!(engine.current_generation(), generation); + assert!(eval_l7(&engine, &allowed)); + assert!(!eval_l7(&engine, &denied)); + } + + let mut repaired = valid; + repaired["network_policies"]["admin"]["endpoints"][0]["deny_rules"][0]["path"] = + "/other/**".into(); + engine + .reload(TEST_POLICY, &repaired.to_string()) + .expect("valid replacement after rejected reloads"); + assert_eq!(engine.current_generation(), generation + 1); + assert!(eval_l7(&engine, &denied)); + } + #[test] fn query_sandbox_config_extracts_filesystem() { let engine = test_engine(); @@ -4818,7 +5301,8 @@ process: "path": path, "query_params": {}, "jsonrpc": { - "method": method + "method": method, + "mcp_method_classification": "available" } } }) @@ -4924,6 +5408,31 @@ process: val == regorus::Value::from(true) } + fn assert_l7_denial( + engine: &OpaEngine, + input: &serde_json::Value, + deny_rule_matches: bool, + expected_reason: &str, + ) { + let mut eng = engine.engine.lock().unwrap(); + set_regorus_input(&mut eng, input.clone()).unwrap(); + assert_eq!( + eng.eval_rule("data.openshell.sandbox.allow_request".into()) + .expect("evaluate allow_request"), + regorus::Value::from(false), + ); + assert_eq!( + eng.eval_rule("data.openshell.sandbox.deny_request".into()) + .expect("evaluate deny_request"), + regorus::Value::from(deny_rule_matches), + ); + assert_eq!( + eng.eval_rule("data.openshell.sandbox.request_deny_reason".into()) + .expect("evaluate one unambiguous denial reason"), + regorus::Value::from(expected_reason), + ); + } + fn eval_l7_raw_data(data: serde_json::Value, input: serde_json::Value) -> bool { let mut engine = regorus::Engine::new(); engine @@ -6270,10 +6779,34 @@ network_policies: "tools/call", serde_json::json!({"name": "blocked_action"}), ); - assert!(!eval_l7(&engine, &blocked)); + assert_l7_denial( + &engine, + &blocked, + true, + "MCP request blocked by a deny rule; ask the policy owner to review deny_rules and tool selectors", + ); + + let unmatched_tool = l7_jsonrpc_input_with_params( + "mcp.params.test", + 8000, + "/mcp", + "tools/call", + serde_json::json!({"name": "private_tool_name", "arguments.secret": "private-value"}), + ); + assert_l7_denial( + &engine, + &unmatched_tool, + false, + "MCP tool call has no matching allow rule; ask the policy owner to review rules and tool selectors", + ); let list_tools = l7_jsonrpc_input("mcp.params.test", 8000, "/mcp", "tools/list"); - assert!(!eval_l7(&engine, &list_tools)); + assert_l7_denial( + &engine, + &list_tools, + false, + "MCP core method is not permitted by policy; ask the policy owner to review rules for this method in the selected MCP revision", + ); } #[test] @@ -6309,6 +6842,185 @@ network_policies: let list_tools = l7_jsonrpc_input("mcp.default.test", 8000, "/mcp", "tools/list"); assert!(eval_l7(&engine, &list_tools)); + + let mut extension = l7_jsonrpc_input( + "mcp.default.test", + 8000, + "/mcp", + "vendor/private_method_name", + ); + extension["request"]["jsonrpc"]["mcp_method_classification"] = + serde_json::json!("extension"); + assert_l7_denial( + &engine, + &extension, + false, + "MCP extension method has no matching exact allow rule; ask the policy owner to review rules with an exact method name and any parameter restrictions; allow_all_known_mcp_methods does not allow extensions", + ); + } + + #[test] + fn l7_mcp_extension_requires_an_exact_method_literal() { + let data = r#" +network_policies: + exact_extension: + name: exact_extension + endpoints: + - host: mcp.extension-exact.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: enforce + rules: + - allow: + method: tools/vendor + binaries: + - { path: /usr/bin/curl } + wildcard_extension: + name: wildcard_extension + endpoints: + - host: mcp.extension-wildcard.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: enforce + rules: + - allow: + method: tools/* + binaries: + - { path: /usr/bin/curl } + denied_extension: + name: denied_extension + endpoints: + - host: mcp.extension-denied.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: enforce + rules: + - allow: + method: tools/vendor + deny_rules: + - method: tools/* + binaries: + - { path: /usr/bin/curl } +"#; + let engine = OpaEngine::from_strings(TEST_POLICY, data).expect("engine from yaml"); + + let mut exact = l7_jsonrpc_input("mcp.extension-exact.test", 8000, "/mcp", "tools/vendor"); + exact["request"]["jsonrpc"]["mcp_method_classification"] = serde_json::json!("extension"); + assert!(eval_l7(&engine, &exact)); + + exact["request"]["jsonrpc"]["mcp_method_classification"] = serde_json::json!("unavailable"); + assert!( + !eval_l7(&engine, &exact), + "known methods unavailable in the selected profile must fail closed" + ); + + let mut wildcard = + l7_jsonrpc_input("mcp.extension-wildcard.test", 8000, "/mcp", "tools/vendor"); + wildcard["request"]["jsonrpc"]["mcp_method_classification"] = + serde_json::json!("extension"); + assert_l7_denial( + &engine, + &wildcard, + false, + "MCP extension method has no matching exact allow rule; ask the policy owner to review rules with an exact method name and any parameter restrictions; allow_all_known_mcp_methods does not allow extensions", + ); + + let mut denied = + l7_jsonrpc_input("mcp.extension-denied.test", 8000, "/mcp", "tools/vendor"); + denied["request"]["jsonrpc"]["mcp_method_classification"] = serde_json::json!("extension"); + assert_l7_denial( + &engine, + &denied, + true, + "MCP request blocked by a deny rule; ask the policy owner to review deny_rules and tool selectors", + ); + } + + #[test] + fn l7_mcp_denial_reason_matches_request_path_and_protocol() { + let data = r#" +network_policies: + mixed: + name: mixed + endpoints: + - host: mcp.reasons.test + port: 8000 + path: /rpc + protocol: json-rpc + enforcement: enforce + rules: [{allow: {method: reports.list}}] + - host: mcp.reasons.test + port: 8000 + path: /rest + protocol: rest + enforcement: enforce + rules: [{allow: {method: GET, path: /rest}}] + - host: mcp.reasons.test + port: 8000 + path: /mcp + protocol: mcp + enforcement: enforce + mcp: + allow_all_known_mcp_methods: true + rules: [{allow: {tool: read_status}}] + binaries: + - {path: /usr/bin/curl} +"#; + let engine = OpaEngine::from_strings(TEST_POLICY, data).expect("engine from yaml"); + + // Select the explanation by the current request path, not the first + // configured endpoint for this host and port. + let unmatched_tool = l7_jsonrpc_input_with_params( + "mcp.reasons.test", + 8000, + "/mcp", + "tools/call", + serde_json::json!({"name": "other_tool"}), + ); + assert_l7_denial( + &engine, + &unmatched_tool, + false, + "MCP tool call has no matching allow rule; ask the policy owner to review rules and tool selectors", + ); + for path in ["/rest", "/rpc", "/other"] { + assert_l7_denial( + &engine, + &l7_jsonrpc_input("mcp.reasons.test", 8000, path, "tools/call"), + false, + &format!("POST {path} not permitted by policy"), + ); + } + + // Batch envelopes have no selected call until the relay evaluates + // each member. Protocol failures also remain outside policy hints. + for field in ["method", "error", "mcp_method_classification"] { + let mut input = unmatched_tool.clone(); + input["request"]["jsonrpc"][field] = match field { + "method" => serde_json::Value::Null, + "error" => serde_json::json!("invalid MCP request"), + _ => serde_json::json!("unavailable"), + }; + assert_l7_denial(&engine, &input, false, "POST /mcp not permitted by policy"); + } + + assert_l7_denial( + &engine, + &l7_jsonrpc_response_input("mcp.reasons.test", 8000, "/rpc"), + true, + "JSON-RPC response frames are not permitted from client to server", + ); + assert!(eval_l7( + &engine, + &l7_jsonrpc_response_input("mcp.reasons.test", 8000, "/mcp") + )); + assert!(eval_l7( + &engine, + &l7_jsonrpc_input("mcp.reasons.test", 8000, "/mcp", "tools/list") + )); } #[test] diff --git a/crates/openshell-supervisor-network/src/opa/raw_schema.rs b/crates/openshell-supervisor-network/src/opa/raw_schema.rs new file mode 100644 index 0000000000..9090265507 --- /dev/null +++ b/crates/openshell-supervisor-network/src/opa/raw_schema.rs @@ -0,0 +1,438 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +//! Validate raw network fields before normalization can erase malformed values. +//! +//! Raw OPA data can contain application data and lowered protocol options that +//! are not authored-policy fields. Validation uses the authored DTOs on copies; +//! the original objects retain their extra fields for Rego evaluation. + +use miette::Result; +use openshell_policy_schema::{ + JsonRpcConfig, L7Allow, L7DenyRule, McpConfig, NetworkBinary, NetworkEndpoint, + NetworkPolicyRule, +}; +use serde::{Deserializer, de::DeserializeOwned}; +use serde_json::{Map, Value}; + +/// Validate consumed fields after the parent has checked collection shapes. +/// +/// Shared fields use the authored DTO types. Only the raw runtime field names +/// and matcher representation need adaptation; existing L7 validators still +/// own protocol combinations and matcher semantics. +pub(super) fn validate_network_settings(data: &Value) -> Result<()> { + let Some(policies) = data.get("network_policies").and_then(Value::as_object) else { + return Ok(()); + }; + for policy in policies.values() { + let mut fields = object_copy(policy)?; + fields.remove("endpoints"); + fields.remove("binaries"); + validate_known::(fields, "invalid network policy settings")?; + + for binary in array_entries(policy, "binaries") { + validate_known::( + object_copy(binary)?, + "invalid network binary settings", + )?; + } + for endpoint in array_entries(policy, "endpoints") { + validate_endpoint(endpoint)?; + } + } + Ok(()) +} + +fn validate_endpoint(endpoint: &Value) -> Result<()> { + let diagnostic = "L7 policy validation failed: invalid L7 policy configuration"; + let mut fields = object_copy(endpoint)?; + // Rules are validated separately because raw matchers accept an explicit + // `glob` object and allow rules also have a flattened OPA representation. + fields.remove("rules"); + fields.remove("deny_rules"); + // The following alias-normalization stage already parses these stanzas + // through the canonical config DTOs and owns their diagnostic category. + fields.remove("mcp"); + fields.remove("json_rpc"); + validate_known::(fields, diagnostic)?; + + validate_renamed::( + endpoint, + &[("json_rpc_max_body_bytes", "max_body_bytes")], + diagnostic, + )?; + validate_renamed::( + endpoint, + &[ + ("mcp_versions", "versions"), + ("mcp_strict_tool_names", "strict_tool_names"), + ( + "mcp_allow_all_known_mcp_methods", + "allow_all_known_mcp_methods", + ), + ], + diagnostic, + )?; + + // These values are emitted by protobuf lowering and read by network/L7 + // consumers. They have no authored DTO field, so check their wire types + // explicitly rather than letting failed reads become runtime defaults. + for field in ["provider_credentialed", "advisor_proposed"] { + if endpoint.get(field).is_some_and(|value| !value.is_boolean()) { + return Err(miette::miette!(diagnostic)); + } + } + for field in ["endpoint_id", "policy_hash"] { + if endpoint.get(field).is_some_and(|value| !value.is_string()) { + return Err(miette::miette!(diagnostic)); + } + } + + for rule in array_entries(endpoint, "rules") { + let mut fields = object_copy(rule.get("allow").unwrap_or(rule))?; + adapt_matchers(&mut fields); + validate_known::(fields, diagnostic)?; + } + for rule in array_entries(endpoint, "deny_rules") { + let mut fields = object_copy(rule)?; + adapt_matchers(&mut fields); + validate_known::(fields, diagnostic)?; + } + Ok(()) +} + +fn array_entries<'a>(object: &'a Value, field: &str) -> &'a [Value] { + // The parent shape validator rejects malformed present arrays before this + // traversal. Missing arrays retain the raw loader's empty default. + object + .get(field) + .and_then(Value::as_array) + .map_or(&[], Vec::as_slice) +} + +fn object_copy(value: &Value) -> Result> { + value.as_object().cloned().ok_or_else(|| { + miette::miette!("L7 policy validation failed: invalid L7 policy configuration") + }) +} + +fn validate_renamed( + value: &Value, + names: &[(&str, &str)], + diagnostic: &'static str, +) -> Result<()> { + let fields = names + .iter() + .filter_map(|(raw, authored)| { + value + .get(*raw) + .map(|value| ((*authored).to_string(), value.clone())) + }) + .collect(); + validate_known::(fields, diagnostic) +} + +/// Adapt only validation copies of runtime matcher leaves. +/// +/// An explicit `{glob: string}` has the same type as the authored scalar +/// matcher. Leave malformed objects intact so the canonical DTO rejects them. +/// In particular, a mixed `glob`/`any` object must not lose either selector. +fn adapt_matchers(rule: &mut Map) { + // The raw query validator treats null as omission. Preserve that OPA-only + // form in installed data while omitting it from this authored-type check. + if rule.get("query").is_some_and(Value::is_null) { + rule.remove("query"); + } + for field in ["query", "params"] { + if let Some(matchers) = rule.get_mut(field).and_then(Value::as_object_mut) { + for matcher in matchers.values_mut() { + adapt_matcher(matcher); + } + } + } + if let Some(tool) = rule.get_mut("tool") { + adapt_matcher(tool); + } +} + +fn adapt_matcher(matcher: &mut Value) { + if let Some(fields) = matcher.as_object() + && fields.len() == 1 + && let Some(glob) = fields.get("glob").filter(|value| value.is_string()) + { + *matcher = glob.clone(); + } +} + +fn validate_known( + fields: Map, + diagnostic: &'static str, +) -> Result<()> { + T::deserialize(KnownFields(Value::Object(fields))) + .map(|_| ()) + // Serde errors include authored values and keys. Do not retain their + // Display, Debug, or source chain at the policy loading boundary. + .map_err(|_| miette::miette!(diagnostic)) +} + +/// Use the canonical DTO's declared field list instead of duplicating it. +/// +/// Unknown fields at this object belong to raw OPA data and remain untouched +/// in the installed document. Nested authored configuration still uses its +/// normal deserializer, including its own unknown-field restrictions. +struct KnownFields(Value); + +impl<'de> Deserializer<'de> for KnownFields { + type Error = serde_json::Error; + + fn deserialize_any>( + self, + visitor: V, + ) -> std::result::Result { + self.0.deserialize_any(visitor) + } + + fn deserialize_struct>( + mut self, + _name: &'static str, + fields: &'static [&'static str], + visitor: V, + ) -> std::result::Result { + if let Value::Object(object) = &mut self.0 { + object.retain(|name, _| fields.contains(&name.as_str())); + } + self.0.deserialize_any(visitor) + } + + serde::forward_to_deserialize_any! { + bool i8 i16 i32 i64 u8 u16 u32 u64 f32 f64 char str string bytes + byte_buf option unit unit_struct newtype_struct seq tuple tuple_struct + map enum identifier ignored_any + } +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::opa::{BAKED_POLICY_RULES, NetworkInput, OpaEngine}; + + fn valid_policy() -> Value { + serde_json::json!({ + "version": 1, + "network_policies": {"private-policy-name": { + "name": "private-policy-name", + "endpoints": [{ + "host": "private-policy-host.test", "port": 443, + "protocol": "rest", "enforcement": "enforce", "access": "full", + "deny_rules": [{"method": "DELETE", "path": "/**"}] + }], + "binaries": [{"path": "/usr/bin/curl"}] + }} + }) + } + + fn input(host: &str) -> NetworkInput { + NetworkInput { + host: host.to_string(), + port: 443, + binary_path: "/usr/bin/curl".into(), + binary_sha256: String::new(), + ancestors: Vec::new(), + cmdline_paths: Vec::new(), + } + } + + #[test] + fn raw_network_leaf_types_reject_before_startup_file_load_and_reload() { + let valid = valid_policy(); + let source = valid.to_string(); + openshell_policy::parse_sandbox_policy(&source).expect("valid typed control"); + let active = OpaEngine::from_strings(BAKED_POLICY_RULES, &source).unwrap(); + let allowed = input("private-policy-host.test"); + let denied = input("unlisted.test"); + let generation = active.current_generation(); + let directory = tempfile::tempdir().unwrap(); + let rego_path = directory.path().join("policy.rego"); + let data_path = directory.path().join("policy.yaml"); + std::fs::write(®o_path, BAKED_POLICY_RULES).unwrap(); + + for (pointer, value) in [ + ( + "/endpoints/0/allow_encoded_slash", + serde_json::json!("true"), + ), + ( + "/endpoints/0/websocket_credential_rewrite", + serde_json::json!("true"), + ), + ( + "/endpoints/0/request_body_credential_rewrite", + serde_json::json!("true"), + ), + ( + "/endpoints/0/allow_uninspected_credentials", + serde_json::json!("true"), + ), + ("/endpoints/0/host", serde_json::json!(["private-value"])), + ("/endpoints/0/port", serde_json::json!("443")), + ("/endpoints/0/ports", serde_json::json!([443, "8443"])), + ("/endpoints/0/tls", serde_json::json!(true)), + ("/endpoints/0/protocol", serde_json::json!(false)), + ("/endpoints/0/allowed_ips", serde_json::json!([7])), + ( + "/endpoints/0/deny_rules/0/method", + serde_json::json!(["DELETE"]), + ), + ( + "/endpoints/0/deny_rules/0/path", + serde_json::json!(["/private-value"]), + ), + ("/endpoints/0/graphql_max_body_bytes", serde_json::json!(-1)), + ("/binaries/0/path", serde_json::json!(["/private-value"])), + ("/name", serde_json::json!(["private-value"])), + ] { + let mut candidate = valid.clone(); + let policy = &mut candidate["network_policies"]["private-policy-name"]; + let (parent, field) = pointer.rsplit_once('/').unwrap(); + policy.pointer_mut(parent).unwrap()[field] = value; + let source = candidate.to_string(); + assert!( + openshell_policy::parse_sandbox_policy(&source).is_err(), + "{pointer}" + ); + let startup = OpaEngine::from_strings(BAKED_POLICY_RULES, &source) + .err() + .expect("malformed leaf must reject at startup"); + std::fs::write(&data_path, &source).unwrap(); + let file = OpaEngine::from_files(®o_path, &data_path) + .err() + .expect("malformed leaf must reject from files"); + let reload = active + .reload(BAKED_POLICY_RULES, &source) + .expect_err("malformed leaf must reject on reload"); + for error in [startup, file, reload] { + for rendered in [ + error.to_string(), + format!("{error:?}"), + format!("{error:#}"), + ] { + assert!(rendered.len() <= 512, "{pointer}: {rendered}"); + assert!(!rendered.contains("private-"), "{pointer}: {rendered}"); + } + } + assert_eq!(active.current_generation(), generation, "{pointer}"); + assert!( + active.evaluate_network(&allowed).unwrap().allowed, + "{pointer}" + ); + assert!( + !active.evaluate_network(&denied).unwrap().allowed, + "{pointer}" + ); + } + } + + #[test] + fn raw_network_validation_retains_runtime_forms_and_custom_data() { + let mut data = valid_policy(); + data.as_object_mut().unwrap().remove("version"); + data["custom_rego_data"] = serde_json::json!({"private": [1, 2]}); + let endpoint = &mut data["network_policies"]["private-policy-name"]["endpoints"][0]; + endpoint["allow_encoded_slash"] = true.into(); + endpoint["provider_credentialed"] = true.into(); + endpoint["advisor_proposed"] = true.into(); + endpoint["policy_hash"] = "private-hash".into(); + endpoint["endpoint_id"] = "private-endpoint".into(); + endpoint["custom_metadata"] = serde_json::json!([null, {"extra": true}]); + endpoint["deny_rules"][0]["query"] = serde_json::json!({ + "repo": {"glob": "private/*"}, "name": "", "org": {"any": ["one", "two"]} + }); + let original = data.clone(); + validate_network_settings(&data).unwrap(); + assert_eq!(data, original); + let engine = OpaEngine::from_strings(BAKED_POLICY_RULES, &data.to_string()).unwrap(); + let endpoint = engine + .query_endpoint_config(&input("private-policy-host.test")) + .unwrap() + .unwrap(); + let endpoint: Value = serde_json::from_str(&endpoint.to_string()).unwrap(); + assert_eq!(endpoint["allow_encoded_slash"], true); + assert_eq!(endpoint["provider_credentialed"], true); + assert_eq!(endpoint["policy_hash"], "private-hash"); + assert_eq!( + endpoint["custom_metadata"], + original["network_policies"]["private-policy-name"]["endpoints"][0]["custom_metadata"] + ); + + let mut null_query = valid_policy(); + null_query["network_policies"]["private-policy-name"]["endpoints"][0]["deny_rules"][0]["query"] = + Value::Null; + OpaEngine::from_strings(BAKED_POLICY_RULES, &null_query.to_string()).unwrap(); + + for protocol in ["json-rpc", "mcp"] { + let mut data = valid_policy(); + let endpoint = &mut data["network_policies"]["private-policy-name"]["endpoints"][0]; + endpoint["protocol"] = protocol.into(); + endpoint.as_object_mut().unwrap().remove("access"); + endpoint.as_object_mut().unwrap().remove("deny_rules"); + endpoint["json_rpc_max_body_bytes"] = 4096.into(); + endpoint["rules"] = serde_json::json!([{"allow": {"method": "tools/call"}}]); + if protocol == "mcp" { + endpoint["mcp_versions"] = serde_json::json!(["2025-11-25"]); + endpoint["mcp_strict_tool_names"] = false.into(); + endpoint["mcp_allow_all_known_mcp_methods"] = true.into(); + // Non-strict MCP names require an exact tool selector; glob + // syntax is permitted only when strict names are enabled. + endpoint["rules"][0]["allow"]["params"] = + serde_json::json!({"name": {"glob": "read_tool"}}); + } + OpaEngine::from_strings(BAKED_POLICY_RULES, &data.to_string()).unwrap(); + } + } + + #[test] + fn raw_network_runtime_fields_and_rule_scalars_reject_malformed_values() { + for (field, value) in [ + ("json_rpc_max_body_bytes", serde_json::json!("4096")), + ("mcp_strict_tool_names", serde_json::json!("false")), + ("mcp_allow_all_known_mcp_methods", Value::Null), + ("mcp_versions", serde_json::json!([42])), + ("provider_credentialed", serde_json::json!("false")), + ("advisor_proposed", serde_json::json!(1)), + ("endpoint_id", serde_json::json!(["id"])), + ("policy_hash", serde_json::json!(false)), + ] { + let mut data = valid_policy(); + data["network_policies"]["private-policy-name"]["endpoints"][0][field] = value; + assert!( + OpaEngine::from_strings(BAKED_POLICY_RULES, &data.to_string()).is_err(), + "{field}" + ); + } + for nested in [true, false] { + for (field, value) in [ + ("method", serde_json::json!(["GET"])), + ("path", serde_json::json!(["/**"])), + ("command", serde_json::json!(true)), + ("operation_name", serde_json::json!(42)), + ("fields", serde_json::json!([42])), + ("query", serde_json::json!({"name": {"glob": 1}})), + ] { + let mut data = valid_policy(); + let endpoint = &mut data["network_policies"]["private-policy-name"]["endpoints"][0]; + endpoint.as_object_mut().unwrap().remove("access"); + let mut rule = serde_json::json!({"method": "GET", "path": "/**"}); + rule[field] = value; + endpoint["rules"] = if nested { + serde_json::json!([{"allow": rule}]) + } else { + serde_json::json!([rule]) + }; + assert!( + OpaEngine::from_strings(BAKED_POLICY_RULES, &data.to_string()).is_err(), + "{nested}: {field}" + ); + } + } + } +} diff --git a/crates/openshell-supervisor-network/src/proxy.rs b/crates/openshell-supervisor-network/src/proxy.rs index 6ac6628cc4..ec567bc47b 100644 --- a/crates/openshell-supervisor-network/src/proxy.rs +++ b/crates/openshell-supervisor-network/src/proxy.rs @@ -5115,6 +5115,7 @@ where crate::l7::rest::RelayRequestOptions { resolver: options.secret_resolver, body_classifier: options.body_classifier, + mcp_request_validation: None, credential_generation: options.credential_generation, generation_guard: Some(options.generation_guard), websocket_extensions: options.websocket_extensions, @@ -5647,7 +5648,15 @@ async fn handle_forward_proxy( .await?; return Ok(()); } - if crate::l7::rest::request_is_h2c_upgrade(&forward_request_bytes) { + // Refuse upgrades this endpoint cannot inspect before the L7 policy + // decision; see `unsupported_upgrade_detail` for the per-protocol rule. + if let Some(upgrade_detail) = crate::l7::rest::unsupported_upgrade_detail( + &forward_request_bytes, + l7_config.config.protocol, + ) { + if let Some(observer) = endpoint_observer.as_ref() { + observer.observe(EndpointResult::PolicyDenied); + } let event = HttpActivityBuilder::new(openshell_ocsf::ctx::ctx()) .activity(ActivityId::Other) .action(ActionId::Denied) @@ -5666,9 +5675,9 @@ async fn handle_forward_proxy( ) .firewall_rule(policy_str, "l7") .message(format!( - "FORWARD_L7 denied unsupported h2c upgrade for {method} {host_lc}:{port}{telemetry_path}" + "FORWARD_L7 denied unsupported upgrade for {method} {host_lc}:{port}{telemetry_path}" )) - .status_detail(crate::l7::rest::UNSUPPORTED_H2C_UPGRADE_DETAIL) + .status_detail(upgrade_detail) .build(); ocsf_emit!(event); emit_activity_simple(activity_tx, true, "l7_parse_rejection"); @@ -5678,7 +5687,7 @@ async fn handle_forward_proxy( port, &binary_str, &decision, - crate::l7::rest::UNSUPPORTED_H2C_UPGRADE_DETAIL, + upgrade_detail, "forward-l7-parse-rejection", ); respond( @@ -5687,7 +5696,7 @@ async fn handle_forward_proxy( 403, "Forbidden", "unsupported_l7_protocol", - crate::l7::rest::UNSUPPORTED_H2C_UPGRADE_DETAIL, + upgrade_detail, ), ) .await?; @@ -5809,22 +5818,21 @@ async fn handle_forward_proxy( crate::l7::jsonrpc::JsonRpcInspectionOptions::for_config(&l7_config.config), ) }; - // Forward HTTP shares the MCP transport gate with CONNECT before - // method authorization. Borrow the buffered request so checking - // the version does not copy the inspected body. - if !crate::l7::relay::enforce_mcp_protocol_version( + // Policy evaluation must use the selected revision's inspection, + // including method classification, just as the CONNECT relays do. + let Some(info) = crate::l7::relay::enforce_mcp_protocol_version( &l7_config.config, &jsonrpc_request, - &info, + info, client, &l7_ctx, &telemetry_path, endpoint_observer.as_ref(), ) .await? - { + else { return Ok(()); - } + }; forward_request_bytes = jsonrpc_request.raw_header; Some(info) } else { @@ -6559,6 +6567,28 @@ async fn handle_forward_proxy( websocket_permessage_deflate, websocket_subprotocol, } => { + // Protocols whose rules apply to individual HTTP requests never + // upgrade (see `upgrade_refusal_for_protocol`). No current path + // forwards upgrade headers for them: the request-side refusal + // rejects them, and request middleware cannot add upgrade or + // connection headers. If a later change lets such a request reach + // an upstream that answers `101`, close instead of relaying frames + // that no rule would inspect. + if forward_upgrade_config.as_ref().is_some_and(|config| { + crate::l7::rest::upgrade_refusal_for_protocol(config.protocol).is_some() + }) { + warn!( + host = %host_lc, + port, + "closing forwarded per-request L7 connection after unexpected protocol upgrade" + ); + if let Some(session) = middleware_session.take() { + session + .end(openshell_core::proto::MiddlewareSessionEndReason::ProtocolError) + .await; + } + return Ok(()); + } let mut upgrade_options = if let (Some(config), Some(engine)) = ( forward_upgrade_config.as_ref(), forward_tunnel_engine.as_ref(), @@ -8426,14 +8456,41 @@ network_policies: {} return; }; - for (body, version_header) in [ + // Every case uses the complete forwarding and middleware path. Sessionless + // requests carry their own metadata, and subscription responses retain SSE bytes. + // Share the proxy's identity cache across requests from this client binary + // so each profile does not hash the entire test executable again. + let identity_cache = Arc::new(BinaryIdentityCache::new()); + for (body, mcp_headers, response_content_type, response_body) in [ ( r#"{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"test","version":"1"}}}"#, "", + "application/json", + r#"{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-11-25","capabilities":{},"serverInfo":{"name":"test","version":"1"}}}"#, ), ( r#"{"jsonrpc":"2.0","id":2,"method":"tools/list"}"#, "MCP-Protocol-Version: 2025-11-25\r\n", + "application/json", + r#"{"jsonrpc":"2.0","id":2,"result":{"tools":[]}}"#, + ), + ( + r#"{"jsonrpc":"2.0","id":3,"method":"server/discover","params":{"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28","io.modelcontextprotocol/clientCapabilities":{}}}}"#, + "MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: server/discover\r\n", + "application/json", + r#"{"jsonrpc":"2.0","id":3,"result":{"supportedVersions":["2026-07-28"],"capabilities":{},"ttlMs":0,"cacheScope":"private"}}"#, + ), + ( + r#"{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"echo","arguments":{"text":"hello"},"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28","io.modelcontextprotocol/clientCapabilities":{}}}}"#, + "MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: tools/call\r\nMcp-Name: echo\r\n", + "application/json", + r#"{"jsonrpc":"2.0","id":4,"result":{"content":[{"type":"text","text":"hello"}]}}"#, + ), + ( + r#"{"jsonrpc":"2.0","id":5,"method":"subscriptions/listen","params":{"notifications":{"toolsListChanged":true},"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28","io.modelcontextprotocol/clientCapabilities":{}}}}"#, + "MCP-Protocol-Version: 2026-07-28\r\nMcp-Method: subscriptions/listen\r\n", + "text/event-stream", + "event: message\ndata: {\"jsonrpc\":\"2.0\",\"method\":\"notifications/tools/list_changed\"}\n\n", ), ] { let upstream_listener = TcpListener::bind((upstream_ip, 0)) @@ -8457,12 +8514,20 @@ network_policies: port: {upstream_port} path: /mcp protocol: mcp + mcp_versions: ["2025-11-25", "2026-07-28"] enforcement: enforce rules: - allow: method: initialize - allow: method: tools/list + - allow: + method: server/discover + - allow: + method: tools/call + tool: echo + - allow: + method: subscriptions/listen binaries: - {{ path: "{executable}" }} "#, @@ -8500,12 +8565,11 @@ network_policies: break; } } - socket - .write_all( - b"HTTP/1.1 200 OK\r\nContent-Length: 2\r\nConnection: close\r\n\r\nok", - ) - .await - .unwrap(); + let response = format!( + "HTTP/1.1 200 OK\r\nContent-Type: {response_content_type}\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{response_body}", + response_body.len(), + ); + socket.write_all(response.as_bytes()).await.unwrap(); String::from_utf8(request).expect("UTF-8 MCP request") }); let proxy_listener = TcpListener::bind("127.0.0.1:0") @@ -8514,7 +8578,7 @@ network_policies: let proxy_address = proxy_listener.local_addr().unwrap(); let target = format!("http://{upstream_ip}:{upstream_port}/mcp"); let request = format!( - "POST {target} HTTP/1.1\r\nHost: {upstream_ip}:{upstream_port}\r\nContent-Type: application/json\r\n{version_header}Content-Length: {}\r\nConnection: close\r\n\r\n{body}", + "POST {target} HTTP/1.1\r\nHost: {upstream_ip}:{upstream_port}\r\nContent-Type: application/json\r\nAccept: application/json, text/event-stream\r\n{mcp_headers}Content-Length: {}\r\nConnection: close\r\n\r\n{body}", body.len(), ); let client = tokio::spawn(async move { @@ -8542,7 +8606,7 @@ network_policies: None, socket_addrs, engine, - Arc::new(BinaryIdentityCache::new()), + Arc::clone(&identity_cache), Arc::new(AtomicU32::new(std::process::id())), None, AgentProposals::default(), @@ -8563,20 +8627,275 @@ network_policies: let response = client.await.expect("join MCP client"); assert!(response.starts_with(b"HTTP/1.1 200 OK")); + let response_header_end = response + .windows(4) + .position(|window| window == b"\r\n\r\n") + .map(|end| end + 4) + .expect("response has HTTP headers"); + assert_eq!( + &response[response_header_end..], + response_body.as_bytes(), + "MCP response bytes must reach the client unchanged" + ); + let response_headers = std::str::from_utf8(&response[..response_header_end]) + .expect("UTF-8 response headers"); + assert!( + response_headers.contains(&format!("Content-Type: {response_content_type}\r\n")) + ); let forwarded = upstream.await.expect("join MCP upstream"); assert!(forwarded.starts_with("POST /mcp HTTP/1.1\r\n")); - if version_header.is_empty() { + if mcp_headers.is_empty() { assert!( !forwarded .to_ascii_lowercase() .contains("mcp-protocol-version:") ); } else { - assert!(forwarded.contains(version_header)); + for header in mcp_headers.lines() { + assert!(forwarded.contains(header), "missing forwarded {header}"); + } } } } + #[tokio::test] + async fn forward_mcp_websocket_upgrade_is_denied_before_connecting_upstream() { + if !cfg!(target_os = "linux") { + eprintln!("skipping: handler identity binding requires /proc (Linux)"); + return; + } + let Some(upstream_ip) = non_loopback_test_ipv4() else { + eprintln!("skipping: no routable non-loopback IPv4 test address"); + return; + }; + + let upstream_listener = TcpListener::bind((upstream_ip, 0)) + .await + .expect("bind MCP upstream listener"); + let upstream_port = upstream_listener.local_addr().unwrap().port(); + let executable = std::env::current_exe().expect("current executable"); + let data = format!( + r#" +network_policies: + mcp-upstream: + name: mcp-upstream + endpoints: + - host: "{upstream_ip}" + port: {upstream_port} + path: /mcp + protocol: mcp + enforcement: enforce + rules: + - allow: + method: initialize + binaries: + - {{ path: "{executable}" }} +"#, + executable = executable.display(), + ); + let engine = Arc::new( + OpaEngine::from_strings(include_str!("../data/sandbox-policy.rego"), &data) + .expect("load MCP policy"), + ); + + let proxy_listener = TcpListener::bind("127.0.0.1:0") + .await + .expect("bind proxy listener"); + let proxy_address = proxy_listener.local_addr().unwrap(); + let target = format!("http://{upstream_ip}:{upstream_port}/mcp"); + // A receive-stream GET that the MCP policy allows, plus WebSocket + // upgrade headers. + let request = format!( + "GET {target} HTTP/1.1\r\nHost: {upstream_ip}:{upstream_port}\r\nAccept: text/event-stream\r\nMCP-Protocol-Version: 2025-11-25\r\nConnection: Upgrade\r\nUpgrade: websocket\r\nSec-WebSocket-Version: 13\r\nSec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==\r\n\r\n" + ); + let client = tokio::spawn(async move { + let mut socket = TcpStream::connect(proxy_address) + .await + .expect("connect proxy"); + let mut response = Vec::new(); + socket + .read_to_end(&mut response) + .await + .expect("read proxy response"); + response + }); + let (proxy_connection, _) = proxy_listener.accept().await.unwrap(); + let socket_addrs = proxy_connection + .peer_addr() + .ok() + .zip(proxy_connection.local_addr().ok()); + let stream: BoundaryDuplexStream = Box::new(proxy_connection); + let mut proxy_connection = tokio::io::BufReader::new(stream); + + tokio::time::timeout( + std::time::Duration::from_secs(30), + Box::pin(handle_forward_proxy( + "GET", + &target, + request.as_bytes(), + request.len(), + &mut proxy_connection, + None, + socket_addrs, + engine, + Arc::new(BinaryIdentityCache::new()), + Arc::new(AtomicU32::new(std::process::id())), + None, + AgentProposals::default(), + Arc::new(None), + Arc::new(None), + None, + None, + None, + None, + None, + None, + )), + ) + .await + .expect("refused upgrade must complete without an upstream response") + .expect("handle refused MCP WebSocket upgrade"); + drop(proxy_connection); + + let response = String::from_utf8(client.await.unwrap()).expect("UTF-8 response"); + assert!( + response.starts_with("HTTP/1.1 403"), + "the upgrade must be refused: {response}" + ); + assert!( + response.contains(crate::l7::rest::UNSUPPORTED_JSONRPC_UPGRADE_DETAIL), + "the refusal must name the unsupported upgrade: {response}" + ); + assert!( + tokio::time::timeout( + std::time::Duration::from_millis(100), + upstream_listener.accept() + ) + .await + .is_err(), + "a refused upgrade must not establish an upstream connection" + ); + } + + #[tokio::test] + async fn forward_graphql_websocket_upgrade_is_denied_before_connecting_upstream() { + if !cfg!(target_os = "linux") { + eprintln!("skipping: handler identity binding requires /proc (Linux)"); + return; + } + let Some(upstream_ip) = non_loopback_test_ipv4() else { + eprintln!("skipping: no routable non-loopback IPv4 test address"); + return; + }; + + let upstream_listener = TcpListener::bind((upstream_ip, 0)) + .await + .expect("bind GraphQL upstream listener"); + let upstream_port = upstream_listener.local_addr().unwrap().port(); + let executable = std::env::current_exe().expect("current executable"); + let data = format!( + r#" +network_policies: + graphql-upstream: + name: graphql-upstream + endpoints: + - host: "{upstream_ip}" + port: {upstream_port} + path: /graphql + protocol: graphql + enforcement: enforce + rules: + - allow: + operation_type: query + fields: [viewer] + binaries: + - {{ path: "{executable}" }} +"#, + executable = executable.display(), + ); + let engine = Arc::new( + OpaEngine::from_strings(include_str!("../data/sandbox-policy.rego"), &data) + .expect("load GraphQL policy"), + ); + + let proxy_listener = TcpListener::bind("127.0.0.1:0") + .await + .expect("bind proxy listener"); + let proxy_address = proxy_listener.local_addr().unwrap(); + let target = format!("http://{upstream_ip}:{upstream_port}/graphql?query=%7Bviewer%7D"); + // A GET whose query the policy allows, plus WebSocket upgrade headers. + let request = format!( + "GET {target} HTTP/1.1\r\nHost: {upstream_ip}:{upstream_port}\r\nConnection: Upgrade\r\nUpgrade: websocket\r\nSec-WebSocket-Version: 13\r\nSec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==\r\n\r\n" + ); + let client = tokio::spawn(async move { + let mut socket = TcpStream::connect(proxy_address) + .await + .expect("connect proxy"); + let mut response = Vec::new(); + socket + .read_to_end(&mut response) + .await + .expect("read proxy response"); + response + }); + let (proxy_connection, _) = proxy_listener.accept().await.unwrap(); + let socket_addrs = proxy_connection + .peer_addr() + .ok() + .zip(proxy_connection.local_addr().ok()); + let stream: BoundaryDuplexStream = Box::new(proxy_connection); + let mut proxy_connection = tokio::io::BufReader::new(stream); + + tokio::time::timeout( + std::time::Duration::from_secs(30), + Box::pin(handle_forward_proxy( + "GET", + &target, + request.as_bytes(), + request.len(), + &mut proxy_connection, + None, + socket_addrs, + engine, + Arc::new(BinaryIdentityCache::new()), + Arc::new(AtomicU32::new(std::process::id())), + None, + AgentProposals::default(), + Arc::new(None), + Arc::new(None), + None, + None, + None, + None, + None, + None, + )), + ) + .await + .expect("refused upgrade must complete without an upstream response") + .expect("handle refused GraphQL WebSocket upgrade"); + drop(proxy_connection); + + let response = String::from_utf8(client.await.unwrap()).expect("UTF-8 response"); + assert!( + response.starts_with("HTTP/1.1 403"), + "the upgrade must be refused: {response}" + ); + assert!( + response.contains(crate::l7::rest::UNSUPPORTED_GRAPHQL_UPGRADE_DETAIL), + "the refusal must name the unsupported upgrade: {response}" + ); + assert!( + tokio::time::timeout( + std::time::Duration::from_millis(100), + upstream_listener.accept() + ) + .await + .is_err(), + "a refused upgrade must not establish an upstream connection" + ); + } + #[tokio::test] async fn plaintext_websocket_preflight_denial_does_not_connect_upstream() { if !cfg!(target_os = "linux") { @@ -9702,6 +10021,8 @@ network_policies: is_batch: false, receive_stream: false, has_response: true, + mcp_revision: None, + mcp_http_metadata: None, error: None, }), }; diff --git a/crates/openshell-supervisor-process/Cargo.toml b/crates/openshell-supervisor-process/Cargo.toml index db56a1fd48..39d48a96e5 100644 --- a/crates/openshell-supervisor-process/Cargo.toml +++ b/crates/openshell-supervisor-process/Cargo.toml @@ -39,6 +39,7 @@ libc = "0.2" [dev-dependencies] openshell-ocsf = { path = "../openshell-ocsf", features = ["test-support"] } tempfile = "3" +tokio = { workspace = true, features = ["test-util"] } [lints] workspace = true diff --git a/crates/openshell-supervisor-process/src/log_push.rs b/crates/openshell-supervisor-process/src/log_push.rs index 470ed8e416..d12061e82e 100644 --- a/crates/openshell-supervisor-process/src/log_push.rs +++ b/crates/openshell-supervisor-process/src/log_push.rs @@ -317,8 +317,8 @@ impl tracing::field::Visit for LogVisitor { mod tests { use super::*; use openshell_ocsf::{ - ActionId, ActivityId, DispositionId, Endpoint, EventContext, NetworkActivityBuilder, - SeverityId, StatusId, ocsf_emit, + ActionId, ActivityId, DispositionId, Endpoint, EventContext, EventOrigin, + NetworkActivityBuilder, SeverityId, StatusId, ocsf_emit, }; use tracing_subscriber::layer::SubscriberExt; @@ -326,6 +326,7 @@ mod tests { EventContext { sandbox_id: "sb-test".to_string(), sandbox_name: "test-sandbox".to_string(), + origin: EventOrigin::Supervisor, container_image: "openshell/sandbox:test".to_string(), hostname: "test-host".to_string(), product_version: "0.0.0".to_string(), diff --git a/crates/openshell-supervisor-process/src/ssh.rs b/crates/openshell-supervisor-process/src/ssh.rs index 01be34bf75..40e0db1668 100644 --- a/crates/openshell-supervisor-process/src/ssh.rs +++ b/crates/openshell-supervisor-process/src/ssh.rs @@ -25,6 +25,15 @@ use tracing::warn; const NO_LOGIN_SHELL_ENV: (&str, &str) = ("OPENSHELL_NO_LOGIN_SHELL", "1"); const MAIN_DETACH_PREFIX: u8 = 0x10; const MAIN_DETACH_KEY: u8 = 0x11; +const MAIN_DETACH_EOF: u8 = 0x04; +const SSH_KEEPALIVE_INTERVAL: Duration = Duration::from_secs(15); +const SSH_PEER_TIMEOUT: Duration = Duration::from_mins(1); + +mod peer_stream; + +#[cfg(test)] +#[path = "ssh/reconnect_tests.rs"] +mod reconnect_tests; fn filter_main_detach_sequence(prefix_pending: &mut bool, data: &[u8]) -> (Vec, bool) { let mut forward = Vec::with_capacity(data.len() + usize::from(*prefix_pending)); @@ -37,6 +46,9 @@ fn filter_main_detach_sequence(prefix_pending: &mut bool, data: &[u8]) -> (Vec, boundary_exec: Arc, main_session: Option>, + peer_timeout: Duration, ) -> Result<()> { // Access is gated by the Unix-socket filesystem permissions (root-only), // not by an application-level preface. The supervisor bridges the @@ -312,9 +334,12 @@ async fn handle_connection( ); let handler = SshHandler::new(port_forward, boundary_exec, main_session); + let stream = peer_stream::PeerStream::new(stream, peer_timeout); russh::server::run_stream(config, stream, handler) .await - .map_err(|err| miette::miette!("ssh stream error: {err}"))?; + .map_err(|err| miette::miette!("ssh stream error: {err}"))? + .await + .map_err(|err| miette::miette!("ssh session error: {err}"))?; Ok(()) } @@ -335,6 +360,9 @@ struct ChannelState { main_input_owner: Option, main_attached: bool, main_read_only: bool, + /// A writable attachment denied stdin may retry on later input. Explicit + /// viewers and channels that sent EOF must never acquire a new lease. + main_input_pending: bool, main_detach_prefix_pending: bool, main_output_task: Option, } @@ -700,6 +728,7 @@ impl russh::server::Handler for SshHandler { Err(error) => (None, Some(error)), } }; + state.main_input_pending = warning.is_some(); state.input_sender = input; state.main_detach_prefix_pending = false; let mut output = main_session.subscribe(); @@ -711,7 +740,7 @@ impl russh::server::Handler for SshHandler { .extended_data( channel, 1, - format!("openshell: {error}; attached read-only; press Ctrl-C to exit{line_ending}").into_bytes(), + format!("openshell: {error}; attached read-only; retry input after the owner disconnects; press Ctrl-C or Ctrl-D to exit{line_ending}").into_bytes(), ) .await; } @@ -827,11 +856,35 @@ impl russh::server::Handler for SshHandler { .await; return Ok(()); } - let (forward, detach) = if state.main_attached { + let denied_prefix = state.main_input_pending && state.main_detach_prefix_pending; + let (mut forward, detach) = if state.main_attached { filter_main_detach_sequence(&mut state.main_detach_prefix_pending, data) } else { (data.to_vec(), false) }; + // Remember a denied viewer's prefix only to recognize split detach + // sequences. It must not become process input when a later frame wins + // the lease: that earlier keystroke was sent without write ownership. + if denied_prefix && forward.first() == Some(&MAIN_DETACH_PREFIX) { + forward.remove(0); + } + // A reconnect can precede detection of the old connection's failure. + // Retry only on new input, using the same exclusive acquisition as a + // fresh attachment. Detach keys retain their read-only behavior. + if state.main_attached + && state.main_input_pending + && !state.main_read_only + && !detach + && !forward.is_empty() + && let Some(main_session) = self.main_session.as_ref() + && !main_session.finished() + && let Ok((owner, input)) = main_session.acquire_input() + { + state.main_input_owner = Some(owner); + state.input_sender = Some(InputSender::Main(input)); + state.main_input_pending = false; + session.extended_data(channel, 1, b"openshell: input enabled\r\n".to_vec())?; + } let error = (!forward.is_empty()) .then(|| state.input_sender.as_ref()?.send(forward).err()) .flatten(); @@ -862,6 +915,7 @@ impl russh::server::Handler for SshHandler { main_session.release_input(owner); } state.input_sender.take(); + state.main_input_pending = false; state.main_detach_prefix_pending = false; } else { warn!("channel_eof on unknown channel {channel:?}"); @@ -1082,6 +1136,7 @@ impl SshHandler { main_session.release_input(owner); } state.input_sender.take(); + state.main_input_pending = false; state.main_detach_prefix_pending = false; if let Some(task) = state.main_output_task.take() { task.abort(); @@ -1217,7 +1272,7 @@ mod tests { use std::io::Write as _; use std::process::{Command, Stdio}; - struct AcceptAnyServerKey; + pub(super) struct AcceptAnyServerKey; impl russh::client::Handler for AcceptAnyServerKey { type Error = russh::Error; @@ -1230,7 +1285,7 @@ mod tests { } } - struct TestLoopbackConnector; + pub(super) struct TestLoopbackConnector; #[async_trait::async_trait] impl openshell_isolation_interface::contract::BoundaryLoopbackConnector for TestLoopbackConnector { @@ -1252,7 +1307,7 @@ mod tests { } } - struct RejectingExec; + pub(super) struct RejectingExec; #[async_trait::async_trait] impl openshell_isolation_interface::contract::BoundaryExec for RejectingExec { @@ -1501,15 +1556,20 @@ mod tests { #[tokio::test] async fn main_attachment_occupied_stdin_still_attaches_read_only() { - assert_read_only_attachment(false, "\n").await; + assert_read_only_attachment(false, "\n", 0x03).await; } #[tokio::test] async fn main_attachment_read_only_pty_warning_returns_cursor_to_start_of_line() { - assert_read_only_attachment(true, "\r\n").await; + assert_read_only_attachment(true, "\r\n", 0x03).await; + } + + #[tokio::test] + async fn main_attachment_read_only_ctrl_d_detaches_without_releasing_owner_lease() { + assert_read_only_attachment(false, "\n", MAIN_DETACH_EOF).await; } - async fn assert_read_only_attachment(terminal: bool, line_ending: &str) { + async fn assert_read_only_attachment(terminal: bool, line_ending: &str, detach_key: u8) { let main_session = MainSession::inert(); let (owner, _input) = main_session.acquire_input().unwrap(); let client = main_test_client(Some(main_session.clone())).await; @@ -1537,14 +1597,17 @@ mod tests { assert_eq!( String::from_utf8_lossy(&data), format!( - "openshell: canonical main process already has an input owner; attached read-only; press Ctrl-C to exit{line_ending}" + "openshell: canonical main process already has an input owner; attached read-only; retry input after the owner disconnects; press Ctrl-C or Ctrl-D to exit{line_ending}" ) ); } event => panic!("expected read-only warning, got {event:?}"), } assert!(main_session.acquire_input().is_err()); - channel.data(&b"ignored\x03also ignored"[..]).await.unwrap(); + let mut data = b"ignored".to_vec(); + data.push(detach_key); + data.extend_from_slice(b"also ignored"); + channel.data(&data[..]).await.unwrap(); assert_viewer_closed(&mut channel).await; assert!(!main_session.finished()); assert!( @@ -1603,6 +1666,52 @@ mod tests { main_session.release_input(owner); } + #[tokio::test] + async fn main_attachment_ctrl_d_detaches_without_closing_main_stdin() { + let (main_session, mut input) = MainSession::inert_with_input(); + let client = main_test_client(Some(main_session.clone())).await; + let mut channel = client.channel_open_session().await.unwrap(); + channel + .request_subsystem(true, "openshell-main") + .await + .unwrap(); + assert!(matches!( + next_main_event(&mut channel).await, + russh::ChannelMsg::Success + )); + + channel.data(&b"hello\x04ignored"[..]).await.unwrap(); + assert_eq!(input.recv().await.unwrap(), b"hello"); + assert_viewer_closed(&mut channel).await; + assert!(!main_session.finished()); + assert!(matches!( + input.try_recv(), + Err(tokio::sync::mpsc::error::TryRecvError::Empty) + )); + let (owner, _) = main_session + .acquire_input() + .expect("detaching must release the input lease"); + main_session.release_input(owner); + } + + #[test] + fn main_detach_filter_preserves_prefix_before_ctrl_d() { + let mut prefix_pending = false; + assert_eq!( + filter_main_detach_sequence(&mut prefix_pending, b"\x10"), + (Vec::new(), false) + ); + assert_eq!( + filter_main_detach_sequence(&mut prefix_pending, b"\x04ignored"), + (vec![0x10], true) + ); + assert!(!prefix_pending); + assert_eq!( + filter_main_detach_sequence(&mut prefix_pending, b"\x10\x11"), + (Vec::new(), true) + ); + } + #[tokio::test] async fn main_attachment_input_owner_ctrl_c_reaches_main_stdin() { let (main_session, mut input) = MainSession::inert_with_input(); diff --git a/crates/openshell-supervisor-process/src/ssh/peer_stream.rs b/crates/openshell-supervisor-process/src/ssh/peer_stream.rs new file mode 100644 index 0000000000..92af3ebfe8 --- /dev/null +++ b/crates/openshell-supervisor-process/src/ssh/peer_stream.rs @@ -0,0 +1,118 @@ +// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +//! Bound SSH peer silence even while the session is blocked writing to its relay. + +use std::future::Future; +use std::io; +use std::pin::Pin; +use std::task::{Context, Poll}; +use std::time::Duration; +use tokio::io::{AsyncRead, AsyncWrite, ReadBuf}; +use tokio::time::{Instant, Sleep}; + +/// Only inbound bytes extend the deadline. SSH keepalive replies allow healthy +/// idle clients to retain their connections without sending application input. +pub(super) struct PeerStream { + stream: S, + timeout: Duration, + deadline: Pin>, +} + +impl PeerStream { + pub(super) fn new(stream: S, timeout: Duration) -> Self { + Self { + stream, + timeout, + deadline: Box::pin(tokio::time::sleep(timeout)), + } + } + + fn poll_deadline(&mut self, cx: &mut Context<'_>) -> io::Result<()> { + if self.deadline.as_mut().poll(cx).is_ready() { + return Err(io::Error::new( + io::ErrorKind::TimedOut, + "SSH peer did not respond before the receive deadline", + )); + } + Ok(()) + } +} + +impl AsyncRead for PeerStream { + fn poll_read( + mut self: Pin<&mut Self>, + cx: &mut Context<'_>, + buf: &mut ReadBuf<'_>, + ) -> Poll> { + self.poll_deadline(cx)?; + let before = buf.filled().len(); + match Pin::new(&mut self.stream).poll_read(cx, buf) { + Poll::Ready(Ok(())) => { + if buf.filled().len() > before { + let next = Instant::now() + self.timeout; + self.deadline.as_mut().reset(next); + } + Poll::Ready(Ok(())) + } + other => other, + } + } +} + +impl AsyncWrite for PeerStream { + fn poll_write( + mut self: Pin<&mut Self>, + cx: &mut Context<'_>, + buf: &[u8], + ) -> Poll> { + // Russh awaits packet writes outside its keepalive select. Polling here + // makes an expired peer close even when the relay stops draining output. + self.poll_deadline(cx)?; + Pin::new(&mut self.stream).poll_write(cx, buf) + } + + fn poll_flush(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll> { + self.poll_deadline(cx)?; + Pin::new(&mut self.stream).poll_flush(cx) + } + + fn poll_shutdown(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll> { + self.poll_deadline(cx)?; + Pin::new(&mut self.stream).poll_shutdown(cx) + } +} + +#[cfg(test)] +mod tests { + use super::*; + use tokio::io::{AsyncReadExt, AsyncWriteExt}; + + #[tokio::test(start_paused = true)] + async fn blocked_write_expires_at_receive_deadline() { + let (stream, _peer) = tokio::io::duplex(1); + let mut stream = PeerStream::new(stream, Duration::from_mins(1)); + let started = Instant::now(); + let error = stream.write_all(b"ab").await.unwrap_err(); + assert_eq!(error.kind(), io::ErrorKind::TimedOut); + assert_eq!(started.elapsed(), Duration::from_mins(1)); + } + + #[tokio::test(start_paused = true)] + async fn only_received_bytes_extend_the_deadline() { + let (stream, mut peer) = tokio::io::duplex(1); + let mut stream = PeerStream::new(stream, Duration::from_mins(1)); + let started = Instant::now(); + tokio::time::advance(Duration::from_secs(40)).await; + peer.write_all(b"r").await.unwrap(); + let mut reply = [0]; + stream.read_exact(&mut reply).await.unwrap(); + assert_eq!(&reply, b"r"); + tokio::time::advance(Duration::from_secs(40)).await; + // A successful write must not renew peer liveness either. + stream.write_all(b"a").await.unwrap(); + let error = stream.write_all(b"b").await.unwrap_err(); + assert_eq!(error.kind(), io::ErrorKind::TimedOut); + assert_eq!(started.elapsed(), Duration::from_secs(100)); + } +} diff --git a/crates/openshell-supervisor-process/src/ssh/reconnect_tests.rs b/crates/openshell-supervisor-process/src/ssh/reconnect_tests.rs new file mode 100644 index 0000000000..f326c63cf2 --- /dev/null +++ b/crates/openshell-supervisor-process/src/ssh/reconnect_tests.rs @@ -0,0 +1,568 @@ +// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +//! Exercise a half-open gateway path through the production supervisor frame bridge. + +use super::tests::{AcceptAnyServerKey, RejectingExec, TestLoopbackConnector}; +use super::*; +use openshell_core::proto::{RelayFrame, relay_frame}; +use openshell_isolation_interface::contract::{ + BackendError, BoundaryExitStatus, BoundaryProcess, BoundarySignal, ProcessAttachment, +}; +use std::sync::atomic::{AtomicUsize, Ordering}; +use tokio::io::{AsyncReadExt, AsyncWriteExt}; +use tokio::sync::{mpsc, watch}; + +const TEST_INTERVAL: Duration = Duration::from_millis(50); +const WAIT: Duration = Duration::from_secs(3); + +#[derive(Clone, Copy, PartialEq)] +enum RelayMode { + Forward, + Discard, + Stalled, +} + +#[derive(Default)] +struct RetainedProcess(AtomicUsize); + +#[async_trait::async_trait] +impl BoundaryProcess for RetainedProcess { + async fn wait(&self) -> Result { + std::future::pending().await + } + + async fn signal(&self, _: BoundarySignal) -> Result<(), BackendError> { + self.0.fetch_add(1, Ordering::SeqCst); + Ok(()) + } + + async fn terminate(&self) -> Result<(), BackendError> { + self.0.fetch_add(1, Ordering::SeqCst); + Ok(()) + } +} + +struct CanonicalProcess { + session: Arc, + process: Arc, + input: tokio::io::DuplexStream, + output: Option, +} + +impl CanonicalProcess { + fn new() -> Self { + let (stdin, input) = tokio::io::duplex(64 * 1024); + let (stdout, output) = tokio::io::duplex(64 * 1024); + let process = Arc::new(RetainedProcess::default()); + let session = MainSession::from_boundary( + ProcessAttachment { + stdin: Box::new(stdin), + stdout: Box::new(stdout), + stderr: None, + terminal: None, + }, + process.clone(), + ); + Self { + session, + process, + input, + output: Some(output), + } + } + + async fn assert_input(&mut self, channel: &russh::Channel) { + channel.data(&b"typing restored\n"[..]).await.unwrap(); + self.expect_input().await; + } + + async fn expect_input(&mut self) { + let mut received = [0; 16]; + tokio::time::timeout(WAIT, self.input.read_exact(&mut received)) + .await + .expect("attachment must deliver input to the retained process") + .unwrap(); + assert_eq!(&received, b"typing restored\n"); + assert!(!self.session.finished()); + assert_eq!(self.process.0.load(Ordering::SeqCst), 0); + } +} + +struct Connection { + client: russh::client::Handle, + mode: watch::Sender, + tasks: Vec>, +} + +impl Drop for Connection { + fn drop(&mut self) { + for task in &self.tasks { + task.abort(); + } + } +} + +impl Connection { + async fn new(main: Arc) -> Self { + Self::with_window(main, None).await + } + + async fn with_window(main: Arc, window: Option) -> Self { + let dir = tempfile::tempdir().unwrap(); + let socket = dir.path().join("ssh.sock"); + let (listener, mut config, _) = ssh_server_init(&socket, &None, false).unwrap(); + // Preserve whether production enables probes. Only shorten the interval. + let config_mut = Arc::get_mut(&mut config).unwrap(); + config_mut.keepalive_interval = config_mut.keepalive_interval.map(|_| TEST_INTERVAL); + let server_task = tokio::spawn(async move { + let (stream, _) = listener.accept().await.unwrap(); + let _ = handle_connection( + stream, + config, + Arc::new(TestLoopbackConnector), + Arc::new(RejectingExec), + Some(main), + TEST_INTERVAL * 4, + ) + .await; + }); + let target = tokio::net::UnixStream::connect(&socket).await.unwrap(); + let (to_supervisor, inbound) = mpsc::channel(16); + let (outbound, mut from_supervisor) = mpsc::channel::(16); + let relay = tokio::spawn(crate::supervisor_session::test_bridge_ssh_relay( + target, inbound, outbound, + )); + let (client_stream, gateway_stream) = tokio::io::duplex(64 * 1024); + let (mut gateway_read, mut gateway_write) = tokio::io::split(gateway_stream); + let (mode, mut input_mode) = watch::channel(RelayMode::Forward); + let mut output_mode = input_mode.clone(); + let inbound_pump = tokio::spawn(async move { + let mut bytes = [0; 16 * 1024]; + loop { + if *input_mode.borrow() == RelayMode::Stalled { + if input_mode.changed().await.is_err() { + break; + } + continue; + } + tokio::select! { + changed = input_mode.changed() => { + if changed.is_err() { break; } + } + read = gateway_read.read(&mut bytes) => { + let Ok(size) = read else { break; }; + if size == 0 { break; } + if *input_mode.borrow() == RelayMode::Forward { + let frame = RelayFrame { + payload: Some(relay_frame::Payload::Data(bytes[..size].to_vec())), + }; + if to_supervisor.send(Ok(frame)).await.is_err() { break; } + } + } + } + } + }); + let outbound_pump = tokio::spawn(async move { + loop { + if *output_mode.borrow() == RelayMode::Stalled { + if output_mode.changed().await.is_err() { + break; + } + continue; + } + tokio::select! { + changed = output_mode.changed() => { + if changed.is_err() { break; } + } + frame = from_supervisor.recv() => { + let Some(frame) = frame else { break; }; + if let Some(relay_frame::Payload::Data(data)) = frame.payload + && *output_mode.borrow() == RelayMode::Forward + && gateway_write.write_all(&data).await.is_err() + { + break; + } + } + } + } + }); + let mut client_config = russh::client::Config::default(); + if let Some(window) = window { + client_config.window_size = window; + } + let mut client = russh::client::connect_stream( + Arc::new(client_config), + client_stream, + AcceptAnyServerKey, + ) + .await + .unwrap(); + assert!(matches!( + client.authenticate_none("sandbox").await.unwrap(), + russh::client::AuthResult::Success + )); + Self { + client, + mode, + tasks: vec![server_task, relay, inbound_pump, outbound_pump], + } + } + + async fn attach(&self) -> russh::Channel { + let mut channel = self.client.channel_open_session().await.unwrap(); + channel + .request_subsystem(true, "openshell-main") + .await + .unwrap(); + assert!(matches!( + next_event(&mut channel).await, + russh::ChannelMsg::Success + )); + channel + } +} + +async fn next_event(channel: &mut russh::Channel) -> russh::ChannelMsg { + tokio::time::timeout(WAIT, channel.wait()) + .await + .unwrap() + .unwrap() +} + +async fn assert_read_only(channel: &mut russh::Channel) { + loop { + match next_event(channel).await { + russh::ChannelMsg::ExtendedData { data, ext: 1 } => { + assert!(String::from_utf8_lossy(&data).contains("attached read-only")); + break; + } + russh::ChannelMsg::Data { .. } => {} + other => panic!("expected read-only diagnostic, got {other:?}"), + } + } +} + +async fn wait_for_release(main: &MainSession) { + tokio::time::timeout(WAIT, async { + loop { + if let Ok((owner, _)) = main.acquire_input() { + main.release_input(owner); + break; + } + tokio::time::sleep(Duration::from_millis(10)).await; + } + }) + .await + .expect("dead relay must release stdin within the liveness bound"); +} + +async fn reconnect_after_cut(mode: RelayMode, producing: bool) { + let mut main = CanonicalProcess::new(); + let original = Connection::new(main.session.clone()).await; + let original_channel = original.attach().await; + main.assert_input(&original_channel).await; + let mut early = Connection::new(main.session.clone()).await; + let mut early_channel = early.attach().await; + assert_read_only(&mut early_channel).await; + let (mut early_read, early_write) = early_channel.split(); + let (notice_tx, mut notice_rx) = mpsc::channel(1); + early.tasks.push(tokio::spawn(async move { + // The replacement remains a healthy output consumer while the old + // path is cut. Its own SSH receive window must not become the fault. + while let Some(event) = early_read.wait().await { + if let russh::ChannelMsg::ExtendedData { data, ext: 1 } = event { + let _ = notice_tx.send(data).await; + } + } + })); + original.mode.send_replace(mode); + + let output_task = producing.then(|| { + let mut output = main.output.take().unwrap(); + tokio::spawn(async move { + // Faster than the stalled relay can drain, but below the retained + // output budget until the SSH event queue has applied backpressure. + let chunk = [b'x'; 16 * 1024]; + loop { + if output.write_all(&chunk).await.is_err() { + break; + } + tokio::time::sleep(Duration::from_millis(2)).await; + } + }) + }); + wait_for_release(&main.session).await; + if let Some(task) = output_task { + task.abort(); + } + // The already reconnected writer can acquire the freed lease on new input. + // Keystrokes ignored while the previous owner was alive are never replayed. + early_write.data(&b"typing restored\n"[..]).await.unwrap(); + main.expect_input().await; + let notice = tokio::time::timeout(WAIT, notice_rx.recv()) + .await + .unwrap() + .unwrap(); + assert_eq!(¬ice[..], b"openshell: input enabled\r\n"); + early_write.close().await.unwrap(); + wait_for_release(&main.session).await; + // A fresh connection after cleanup can also attach to that same process. + let mut recovered = Connection::new(main.session.clone()).await; + let recovered_channel = recovered.attach().await; + let (mut recovered_read, recovered_write) = recovered_channel.split(); + recovered.tasks.push(tokio::spawn(async move { + while recovered_read.wait().await.is_some() {} + })); + recovered_write + .data(&b"typing restored\n"[..]) + .await + .unwrap(); + main.expect_input().await; + assert!(main.session.acquire_input().is_err()); +} + +#[tokio::test] +async fn silent_half_open_relay_releases_canonical_input() { + reconnect_after_cut(RelayMode::Discard, false).await; +} + +#[tokio::test] +async fn output_does_not_keep_a_half_open_relay_alive() { + reconnect_after_cut(RelayMode::Discard, true).await; +} + +#[tokio::test] +async fn stalled_relay_writes_release_canonical_input() { + reconnect_after_cut(RelayMode::Stalled, true).await; +} + +#[tokio::test] +async fn healthy_idle_owner_survives_probes_and_cannot_be_displaced() { + let mut main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let channel = owner.attach().await; + main.assert_input(&channel).await; + tokio::time::sleep(TEST_INTERVAL * 12).await; + let viewer = Connection::new(main.session.clone()).await; + let mut viewer_channel = viewer.attach().await; + assert_read_only(&mut viewer_channel).await; + viewer_channel.data(&b"ignored\n"[..]).await.unwrap(); + viewer_channel.eof().await.unwrap(); + viewer_channel.close().await.unwrap(); + main.assert_input(&channel).await; + assert!(main.session.acquire_input().is_err()); +} + +#[tokio::test] +async fn explicit_viewer_does_not_acquire_released_input() { + let mut main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let owner_channel = owner.attach().await; + let viewer = Connection::new(main.session.clone()).await; + let mut channel = viewer.client.channel_open_session().await.unwrap(); + channel + .set_env(true, "OPENSHELL_MAIN_READ_ONLY", "1") + .await + .unwrap(); + assert!(matches!( + next_event(&mut channel).await, + russh::ChannelMsg::Success + )); + channel + .request_subsystem(true, "openshell-main") + .await + .unwrap(); + assert!(matches!( + next_event(&mut channel).await, + russh::ChannelMsg::Success + )); + owner_channel.close().await.unwrap(); + wait_for_release(&main.session).await; + channel.data(&b"ignored\n"[..]).await.unwrap(); + tokio::time::sleep(Duration::from_millis(30)).await; + let fresh = Connection::new(main.session.clone()).await; + let fresh_channel = fresh.attach().await; + main.assert_input(&fresh_channel).await; +} + +#[tokio::test] +async fn waiting_writer_eof_cancels_acquisition() { + let mut main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let owner_channel = owner.attach().await; + let waiting = Connection::new(main.session.clone()).await; + let mut channel = waiting.attach().await; + assert_read_only(&mut channel).await; + channel.eof().await.unwrap(); + owner_channel.close().await.unwrap(); + wait_for_release(&main.session).await; + channel.data(&b"ignored\n"[..]).await.unwrap(); + tokio::time::sleep(Duration::from_millis(30)).await; + let fresh = Connection::new(main.session.clone()).await; + let fresh_channel = fresh.attach().await; + main.assert_input(&fresh_channel).await; +} + +#[tokio::test] +async fn production_config_probes_before_the_receive_deadline() { + let dir = tempfile::tempdir().unwrap(); + let (_, config, _) = ssh_server_init(&dir.path().join("ssh.sock"), &None, false).unwrap(); + assert_eq!(config.keepalive_interval, Some(Duration::from_secs(15))); + assert_eq!(config.keepalive_max, 3); + assert_eq!(SSH_PEER_TIMEOUT, Duration::from_mins(1)); +} + +#[tokio::test] +async fn waiting_writer_detach_keys_do_not_acquire_released_input() { + for key in [b"\x03".as_slice(), b"\x04", b"\x10\x11"] { + let main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let owner_channel = owner.attach().await; + let waiting = Connection::new(main.session.clone()).await; + let mut channel = waiting.attach().await; + assert_read_only(&mut channel).await; + owner_channel.close().await.unwrap(); + wait_for_release(&main.session).await; + channel.data(key).await.unwrap(); + loop { + if matches!(next_event(&mut channel).await, russh::ChannelMsg::Close) { + break; + } + } + wait_for_release(&main.session).await; + assert_eq!(main.process.0.load(Ordering::SeqCst), 0); + } +} + +#[tokio::test] +async fn competing_waiting_writers_preserve_exclusive_ownership() { + let mut main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let owner_channel = owner.attach().await; + let first = Connection::new(main.session.clone()).await; + let mut first_channel = first.attach().await; + assert_read_only(&mut first_channel).await; + let second = Connection::new(main.session.clone()).await; + let mut second_channel = second.attach().await; + assert_read_only(&mut second_channel).await; + owner_channel.close().await.unwrap(); + wait_for_release(&main.session).await; + let (a, b) = tokio::join!( + first_channel.data(&b"a"[..]), + second_channel.data(&b"b"[..]) + ); + a.unwrap(); + b.unwrap(); + let mut winner = [0]; + tokio::time::timeout(WAIT, main.input.read_exact(&mut winner)) + .await + .unwrap() + .unwrap(); + let mut extra = [0]; + assert!( + tokio::time::timeout(Duration::from_millis(30), main.input.read(&mut extra)) + .await + .is_err() + ); + let winning_channel = match winner[0] { + b'a' => { + second_channel.close().await.unwrap(); + first_channel + } + b'b' => { + first_channel.close().await.unwrap(); + second_channel + } + other => panic!("unexpected input: {other}"), + }; + tokio::time::sleep(Duration::from_millis(30)).await; + assert!(main.session.acquire_input().is_err()); + main.assert_input(&winning_channel).await; +} + +#[tokio::test] +async fn detached_waiting_writer_cannot_reacquire_with_pending_output() { + let mut main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let owner_channel = owner.attach().await; + // The read-only diagnostic exhausts this channel's output window, so + // russh retains the channel while the detach close waits behind output. + let waiting = Connection::with_window(main.session.clone(), Some(0)).await; + let channel = waiting.attach().await; + owner_channel.close().await.unwrap(); + wait_for_release(&main.session).await; + channel.data(&b"\x04"[..]).await.unwrap(); + channel.data(&b"after detach"[..]).await.unwrap(); + let mut received = [0]; + assert!( + tokio::time::timeout(Duration::from_millis(100), main.input.read(&mut received)) + .await + .is_err(), + "detached channel forwarded input: {received:?}" + ); + let (lease, _) = main + .session + .acquire_input() + .expect("detach must cancel pending ownership"); + main.session.release_input(lease); + channel.close().await.unwrap(); + wait_for_release(&main.session).await; +} + +#[tokio::test] +async fn retry_does_not_forward_prefix_typed_before_owner_release() { + let mut main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let owner_channel = owner.attach().await; + let waiting = Connection::new(main.session.clone()).await; + let mut channel = waiting.attach().await; + assert_read_only(&mut channel).await; + channel.data(&b"\x10"[..]).await.unwrap(); + // A request response on the same SSH channel proves the prefix was + // processed before the original owner releases its lease. + channel.set_env(true, "REVIEW_BARRIER", "1").await.unwrap(); + assert!(matches!( + next_event(&mut channel).await, + russh::ChannelMsg::Success + )); + owner_channel.close().await.unwrap(); + wait_for_release(&main.session).await; + channel.data(&b"z"[..]).await.unwrap(); + let mut received = [0]; + tokio::time::timeout(WAIT, main.input.read_exact(&mut received)) + .await + .unwrap() + .unwrap(); + assert_eq!( + received, + [b'z'], + "retry replayed a prefix typed while stdin was denied" + ); +} + +#[tokio::test] +async fn waiting_writer_split_detach_survives_owner_release() { + let main = CanonicalProcess::new(); + let owner = Connection::new(main.session.clone()).await; + let owner_channel = owner.attach().await; + let waiting = Connection::new(main.session.clone()).await; + let mut channel = waiting.attach().await; + assert_read_only(&mut channel).await; + channel.data(&b"\x10"[..]).await.unwrap(); + channel.set_env(true, "TEST_BARRIER", "1").await.unwrap(); + assert!(matches!( + next_event(&mut channel).await, + russh::ChannelMsg::Success + )); + owner_channel.close().await.unwrap(); + wait_for_release(&main.session).await; + channel.data(&b"\x11"[..]).await.unwrap(); + loop { + if matches!(next_event(&mut channel).await, russh::ChannelMsg::Close) { + break; + } + } + wait_for_release(&main.session).await; + assert_eq!(main.process.0.load(Ordering::SeqCst), 0); +} diff --git a/crates/openshell-supervisor-process/src/supervisor_session.rs b/crates/openshell-supervisor-process/src/supervisor_session.rs index 0b713f9921..5be01017eb 100644 --- a/crates/openshell-supervisor-process/src/supervisor_session.rs +++ b/crates/openshell-supervisor-process/src/supervisor_session.rs @@ -522,6 +522,22 @@ pub async fn report_main_process_exit( Ok(()) } +#[cfg(test)] +pub(crate) async fn test_bridge_ssh_relay( + target: tokio::net::UnixStream, + inbound: mpsc::Receiver>, + out_tx: mpsc::Sender, +) { + let _ = bridge_relay( + Box::new(target), + tokio_stream::wrappers::ReceiverStream::new(inbound), + out_tx, + "half-open-test".into(), + Arc::new(AtomicBool::new(false)), + ) + .await; +} + /// Confirm terminal delivery and permit ephemeral cleanup. pub async fn finalize_main_process_exit( endpoint: &str, @@ -687,8 +703,24 @@ async fn handle_relay_open( } Err(e) => return Err(format!("relay_stream RPC failed: {e}").into()), }; - let mut inbound = response.into_inner(); + bridge_relay( + target, + response.into_inner(), + out_tx, + channel_id, + terminating, + ) + .await +} +/// Forward the relay's data frames without interpreting the target protocol. +async fn bridge_relay( + target: Box, + mut inbound: impl tokio_stream::Stream> + Unpin, + out_tx: mpsc::Sender, + channel_id: String, + terminating: Arc, +) -> Result<(), Box> { // Connect to the local SSH daemon on its Unix socket. let (mut target_r, mut target_w) = tokio::io::split(target); @@ -939,6 +971,7 @@ mod ocsf_event_tests { product_version: "0.0.1".into(), proxy_ip: "127.0.0.1".parse().unwrap(), proxy_port: 3128, + origin: openshell_ocsf::EventOrigin::Supervisor, } } diff --git a/crates/openshell-supervisor/Cargo.toml b/crates/openshell-supervisor/Cargo.toml index d4fb9cee68..483dd50bff 100644 --- a/crates/openshell-supervisor/Cargo.toml +++ b/crates/openshell-supervisor/Cargo.toml @@ -35,6 +35,7 @@ rustls = { workspace = true } rustix = { workspace = true } serde = { workspace = true } serde_json = { workspace = true } +socket2 = { workspace = true } tempfile = "3" tokio = { workspace = true } tonic = { workspace = true, features = ["channel", "tls-native-roots"] } diff --git a/crates/openshell-supervisor/src/lib.rs b/crates/openshell-supervisor/src/lib.rs index 8dc3f97600..63c5302673 100644 --- a/crates/openshell-supervisor/src/lib.rs +++ b/crates/openshell-supervisor/src/lib.rs @@ -90,32 +90,97 @@ where } } +/// Where the supervisor publishes readiness. The listener exists only while +/// the supervisor session is ready, so a successful connect means ready. +#[derive(Clone, Debug)] +enum ReadinessEndpoint { + Unix(std::path::PathBuf), + /// Wildcard TCP port reachable from the node, for kubelet `tcpSocket` probes. + Tcp(u16), +} + +enum ReadinessListener { + Unix(tokio::net::UnixListener), + Tcp(tokio::net::TcpListener), +} + +impl ReadinessEndpoint { + fn bind(&self) -> Result { + match self { + Self::Unix(path) => { + prepare_control_readiness_path(path)?; + tokio::net::UnixListener::bind(path) + .map(ReadinessListener::Unix) + .into_diagnostic() + .wrap_err_with(|| { + format!("bind supervisor readiness socket on {}", path.display()) + }) + } + Self::Tcp(port) => bind_readiness_tcp(*port) + .and_then(tokio::net::TcpListener::from_std) + .map(ReadinessListener::Tcp) + .into_diagnostic() + .wrap_err_with(|| format!("bind supervisor readiness listener on port {port}")), + } + } + + fn remove(&self) { + if let Self::Unix(path) = self { + let _ = std::fs::remove_file(path); + } + } +} + +/// Falls back to IPv4 when the network namespace has IPv6 disabled. +fn bind_readiness_tcp(port: u16) -> std::io::Result { + use socket2::{Domain, Socket, Type}; + + let bind = |domain: Domain, address: std::net::SocketAddr| { + let socket = Socket::new(domain, Type::STREAM, None)?; + if domain == Domain::IPV6 { + socket.set_only_v6(false)?; + } + socket.set_reuse_address(true)?; + socket.bind(&address.into())?; + socket.listen(128)?; + socket.set_nonblocking(true)?; + Ok::<_, std::io::Error>(std::net::TcpListener::from(socket)) + }; + bind(Domain::IPV6, (std::net::Ipv6Addr::UNSPECIFIED, port).into()) + .or_else(|_| bind(Domain::IPV4, (std::net::Ipv4Addr::UNSPECIFIED, port).into())) +} + +impl ReadinessListener { + async fn accept(&self) -> std::io::Result<()> { + match self { + Self::Unix(listener) => listener.accept().await.map(drop), + Self::Tcp(listener) => listener.accept().await.map(drop), + } + } +} + struct ControlReadiness { task: tokio::task::JoinHandle<()>, - path: std::path::PathBuf, + endpoint: ReadinessEndpoint, } impl ControlReadiness { fn start( - path: std::path::PathBuf, + endpoint: ReadinessEndpoint, mut session_readiness: Option>, ) -> Result { - prepare_control_readiness_path(&path)?; + if let ReadinessEndpoint::Unix(path) = &endpoint { + prepare_control_readiness_path(path)?; + } let listener = if session_readiness .as_ref() .is_some_and(|readiness| !*readiness.borrow()) { None } else { - Some( - tokio::net::UnixListener::bind(&path) - .into_diagnostic() - .wrap_err_with(|| { - format!("bind supervisor readiness socket on {}", path.display()) - })?, - ) + Some(endpoint.bind()?) }; - let task_path = path.clone(); + let task_endpoint = endpoint.clone(); let task = tokio::spawn(async move { let mut listener = listener; loop { @@ -125,7 +190,7 @@ impl ControlReadiness { if listener.is_none() || session_unready { if session_unready { listener.take(); - let _ = std::fs::remove_file(&task_path); + task_endpoint.remove(); let Some(readiness) = session_readiness.as_mut() else { break; }; @@ -133,16 +198,7 @@ impl ControlReadiness { break; } } - match prepare_control_readiness_path(&task_path).and_then(|()| { - tokio::net::UnixListener::bind(&task_path) - .into_diagnostic() - .wrap_err_with(|| { - format!( - "rebind supervisor readiness socket on {}", - task_path.display() - ) - }) - }) { + match task_endpoint.bind() { Ok(rebound) => listener = Some(rebound), Err(error) => { tracing::warn!(%error, "control-mode readiness rebind failed; retrying"); @@ -159,7 +215,7 @@ impl ControlReadiness { if let Some(readiness) = session_readiness.as_mut() { tokio::select! { accepted = active_listener.accept() => match accepted { - Ok((stream, _)) => drop(stream), + Ok(()) => {} Err(error) => { tracing::warn!(%error, "control-mode readiness accept failed; retrying"); tokio::time::sleep(Duration::from_millis(100)).await; @@ -173,7 +229,7 @@ impl ControlReadiness { } } else { match active_listener.accept().await { - Ok((stream, _)) => drop(stream), + Ok(()) => {} Err(error) => { tracing::warn!(%error, "control-mode readiness accept failed; retrying"); tokio::time::sleep(Duration::from_millis(100)).await; @@ -181,9 +237,9 @@ impl ControlReadiness { } } } - let _ = std::fs::remove_file(&task_path); + task_endpoint.remove(); }); - Ok(Self { task, path }) + Ok(Self { task, endpoint }) } } @@ -228,7 +284,7 @@ fn prepare_control_readiness_path(path: &std::path::Path) -> Result<()> { impl Drop for ControlReadiness { fn drop(&mut self) { self.task.abort(); - let _ = std::fs::remove_file(&self.path); + self.endpoint.remove(); } } @@ -463,6 +519,7 @@ pub async fn run_network_proxy( product_version: openshell_core::VERSION.to_string(), proxy_ip: listen.ip(), proxy_port: listen.port(), + origin: openshell_ocsf::EventOrigin::Supervisor, }) { debug!("OCSF context already initialized, keeping existing"); } @@ -571,6 +628,7 @@ pub async fn run_sandbox( policy_data: Option, ssh_socket_path: Option, health_socket_path: Option, + health_port: Option, ocsf_enabled: Arc, ocsf_schema_version: Arc>, upstream_proxy_args: openshell_supervisor_network::upstream_proxy::UpstreamProxyArgs, @@ -605,6 +663,7 @@ pub async fn run_sandbox( product_version: openshell_core::VERSION.to_string(), proxy_ip: std::net::IpAddr::from([127, 0, 0, 1]), proxy_port: 3128, + origin: openshell_ocsf::EventOrigin::Supervisor, }) { debug!("OCSF context already initialized, keeping existing"); } @@ -653,7 +712,7 @@ pub async fn run_sandbox( loaded_policy_origin, initial_agent_proposals_enabled, initial_extension_authentication_enabled, - captured_provider_credentials, + captured_provider_environment, ) = load_policy_with_gateway( sandbox_id.clone(), sandbox.clone(), @@ -678,8 +737,8 @@ pub async fn run_sandbox( let workspace = workdir; let provider_readiness = ProviderReadinessTracker::new(); - let provider_credentials = if let Some(credentials) = captured_provider_credentials { - credentials + let provider_credentials = if let Some(environment) = captured_provider_environment { + environment.install(&provider_readiness) } else { // Fetch provider environment variables from the server. // This is done after loading the policy so the sandbox can still start @@ -1125,14 +1184,12 @@ pub async fn run_sandbox( running.exec(), ) }); - let mut control_readiness = if let Some(path) = health_socket_path { - Some(ControlReadiness::start( - path, - boundary_access.session_readiness(), - )?) - } else { - None - }; + let mut control_readiness = health_socket_path + .map(ReadinessEndpoint::Unix) + .into_iter() + .chain(health_port.map(ReadinessEndpoint::Tcp)) + .map(|endpoint| ControlReadiness::start(endpoint, boundary_access.session_readiness())) + .collect::>>()?; let instance_id = boundary_access.instance_id().to_string(); let wait_agent = agent.clone(); let shutdown_requested = wait_for_control_shutdown_signal(); @@ -1186,7 +1243,7 @@ pub async fn run_sandbox( } }; if !retain_access { - control_readiness.take(); + control_readiness.clear(); } boundary_access .publish_main_exit(exit_code, await_main_process_attachment) @@ -1231,7 +1288,7 @@ pub async fn run_sandbox( } if completion_cancelled { retain_access = false; - control_readiness.take(); + control_readiness.clear(); } if retain_access { info!(backend = %backend_name, "Canonical process exited; retaining control-mode access plane"); @@ -1987,6 +2044,35 @@ enum LocalPolicyIdentity { EndpointOnly, } +struct CapturedProviderEnvironment { + credentials: ProviderCredentialState, + expires_at_ms: Option, + identity: EnvironmentIdentity, +} + +impl CapturedProviderEnvironment { + fn install(self, readiness: &ProviderReadinessTracker) -> ProviderCredentialState { + readiness.credentials_installed(self.identity, &self.credentials, self.expires_at_ms); + self.credentials + } + + fn new( + credentials: ProviderCredentialState, + provider: &openshell_core::grpc_client::ProviderEnvironmentResult, + ) -> Self { + Self { + credentials, + expires_at_ms: provider + .credential_expires_at_ms + .values() + .copied() + .filter(|expiry| *expiry > 0) + .min(), + identity: EnvironmentIdentity::from_environment(provider), + } + } +} + async fn load_policy( sandbox_id: Option, sandbox: Option, @@ -2003,7 +2089,7 @@ async fn load_policy( LoadedPolicyOrigin, bool, bool, - Option, + Option, )> { load_policy_with_gateway( sandbox_id, @@ -2043,7 +2129,7 @@ async fn load_policy_with_gateway( LoadedPolicyOrigin, bool, bool, - Option, + Option, )> { use openshell_core::proto::ConfigurationAdmissionState; // File mode: load OPA engine from rego rules + YAML data (dev override) @@ -2406,7 +2492,10 @@ async fn load_policy_with_gateway( }, agent_proposals_enabled_from_settings(&snapshot.settings), snapshot.extension_authentication_enabled, - Some(captured_provider_credentials), + Some(CapturedProviderEnvironment::new( + captured_provider_credentials, + &provider, + )), )); } } @@ -2493,7 +2582,7 @@ fn provider_environment_is_installable(reason: ProviderReadinessReason) -> bool fn prepare_provider_environment( provider: &openshell_core::grpc_client::ProviderEnvironmentResult, ) -> Result { - ProviderCredentialState::from_bound_environment( + let prepared = ProviderCredentialState::from_bound_environment( provider.provider_env_revision, provider.environment.clone(), provider.credential_expires_at_ms.clone(), @@ -2501,7 +2590,9 @@ fn prepare_provider_environment( provider.static_credential_bindings.clone(), provider.non_secret_environment_keys.clone(), ) - .map_err(|_| miette::miette!("Provider credential bindings are invalid")) + .map_err(|_| miette::miette!("Provider credential bindings are invalid"))?; + prepared.set_managed_files(provider.files.clone()); + Ok(prepared) } // Retain only the most recent rejection, so A -> B -> A emits all transitions. @@ -3036,6 +3127,7 @@ fn initial_provider_credentials( result.non_secret_environment_keys, ) { Ok(credentials) => { + credentials.set_managed_files(result.files); readiness.credentials_installed(identity, &credentials, expires_at_ms); credentials } @@ -4744,8 +4836,8 @@ mod tests { async fn control_readiness_exists_only_while_guard_is_live() { let root = tempfile::tempdir().unwrap(); let path = root.path().join("health.sock"); - let readiness = - ControlReadiness::start(path.clone(), None).expect("start readiness listener"); + let readiness = ControlReadiness::start(ReadinessEndpoint::Unix(path.clone()), None) + .expect("start readiness listener"); check_control_readiness(&path).expect("running supervisor accepts readiness probes"); drop(readiness); @@ -4758,8 +4850,9 @@ mod tests { let root = tempfile::tempdir().unwrap(); let path = root.path().join("health.sock"); let (session_tx, session_rx) = tokio::sync::watch::channel(false); - let _readiness = ControlReadiness::start(path.clone(), Some(session_rx)) - .expect("start readiness listener"); + let _readiness = + ControlReadiness::start(ReadinessEndpoint::Unix(path.clone()), Some(session_rx)) + .expect("start readiness listener"); assert!(check_control_readiness(&path).is_err()); session_tx.send_replace(true); @@ -4790,6 +4883,65 @@ mod tests { .expect("replacement session restores readiness socket"); } + #[test] + fn tcp_readiness_listener_accepts_ipv4_regardless_of_bindv6only() { + let port = std::net::TcpListener::bind("127.0.0.1:0") + .and_then(|reserved| reserved.local_addr()) + .expect("reserve loopback port") + .port(); + let listener = bind_readiness_tcp(port).expect("bind readiness listener"); + if listener.local_addr().expect("local address").is_ipv6() { + assert!( + !socket2::SockRef::from(&listener) + .only_v6() + .expect("read IPV6_V6ONLY"), + "net.ipv6.bindv6only=1 must not make the wildcard listener IPv6-only" + ); + } + std::net::TcpStream::connect(("127.0.0.1", port)) + .expect("IPv4 kubelet probe reaches the readiness listener"); + } + + #[tokio::test] + async fn tcp_control_readiness_tracks_supervisor_session() { + let port = std::net::TcpListener::bind("127.0.0.1:0") + .and_then(|reserved| reserved.local_addr()) + .expect("reserve loopback port") + .port(); + let (session_tx, session_rx) = tokio::sync::watch::channel(true); + let readiness = ControlReadiness::start(ReadinessEndpoint::Tcp(port), Some(session_rx)) + .expect("start TCP readiness listener"); + let connects = || std::net::TcpStream::connect(("127.0.0.1", port)).is_ok(); + assert!(connects(), "accepted session is ready"); + + session_tx.send_replace(false); + timeout(Duration::from_secs(1), async { + while connects() { + tokio::task::yield_now().await; + } + }) + .await + .expect("lost session closes readiness listener"); + + session_tx.send_replace(true); + timeout(Duration::from_secs(1), async { + while !connects() { + tokio::task::yield_now().await; + } + }) + .await + .expect("replacement session reopens readiness listener"); + + drop(readiness); + timeout(Duration::from_secs(1), async { + while connects() { + tokio::task::yield_now().await; + } + }) + .await + .expect("dropped guard closes readiness listener"); + } + #[test] fn control_readiness_rejects_relative_path() { let error = prepare_control_readiness_path(std::path::Path::new("health.sock")) @@ -5415,6 +5567,7 @@ network_policies: fn startup_provider(revision: u64) -> openshell_core::grpc_client::ProviderEnvironmentResult { openshell_core::grpc_client::ProviderEnvironmentResult { + files: std::collections::HashMap::new(), provider_env_revision: revision, provider_attachment_epoch: String::new(), policy_hash: String::new(), @@ -5443,6 +5596,22 @@ network_policies: assert_eq!(credentials.revision(), 10); } + #[test] + fn startup_environment_seeds_provider_readiness() { + let mut provider = startup_provider(10); + provider.provider_attachment_epoch = "epoch".to_string(); + provider.policy_hash = "policy".to_string(); + let identity = EnvironmentIdentity::from_environment(&provider); + let credentials = prepare_provider_environment(&provider).unwrap(); + let readiness = ProviderReadinessTracker::new(); + + let credentials = + CapturedProviderEnvironment::new(credentials, &provider).install(&readiness); + + assert_eq!(credentials.revision(), 10); + assert!(!readiness.needs_environment(&identity)); + } + #[test] fn startup_configuration_accepts_fail_closed_provider_environment() { let policy = proto_policy_fixture(); @@ -5671,6 +5840,7 @@ network_policies: use std::collections::HashMap; let mut result = openshell_core::grpc_client::ProviderEnvironmentResult { + files: HashMap::new(), environment: HashMap::new(), provider_env_revision: revision, provider_attachment_epoch: String::new(), @@ -6080,6 +6250,65 @@ network_policies: let _ = task.await; } + #[tokio::test] + async fn provider_readiness_startup_environment_avoids_unchanged_refresh() { + let policy = proto_policy_fixture(); + let mut settings = settings_poll_result( + Some(policy.clone()), + 1, + openshell_core::proto::PolicySource::Sandbox, + ); + settings.provider_env_revision = 6; + let engine = Arc::new(OpaEngine::from_proto(&policy).unwrap()); + let mut ctx = policy_poll_test_context( + engine.clone(), + LoadedPolicyOrigin::Gateway { + revision: Some(LoadedPolicyRevision::from_snapshot(&settings)), + has_last_valid_policy: true, + }, + default_middleware_connector(), + ); + let mut provider = static_provider_environment(6, Some("initial")); + provider.policy_hash.clone_from(&settings.policy_hash); + let credentials = prepare_provider_environment(&provider).unwrap(); + ctx.provider_credentials = CapturedProviderEnvironment::new(credentials, &provider) + .install(&ctx.provider_readiness); + let generation = engine.current_generation(); + let guard = engine.generation_guard(generation).unwrap(); + let (policy_gateway, polls, mut reports) = scripted_policy_gateway(); + let observed_polls = policy_gateway.polled_sandboxes.clone(); + let (requests, mut received) = tokio::sync::mpsc::unbounded_channel(); + let task = tokio::spawn(run_policy_poll_loop_with_client( + ctx, + ScriptedProviderGateway { + policy: policy_gateway, + requests, + }, + )); + + polls.send(settings.clone()).unwrap(); + expect_policy_report(&mut reports, 1).await; + polls.send(settings).unwrap(); + timeout(Duration::from_secs(1), async { + while observed_polls.lock().await.len() < 2 { + tokio::task::yield_now().await; + } + }) + .await + .unwrap(); + + assert!( + timeout(Duration::from_millis(50), received.recv()) + .await + .is_err(), + "unchanged settings must not refetch the startup environment" + ); + assert_eq!(engine.current_generation(), generation); + assert!(!guard.is_stale()); + task.abort(); + let _ = task.await; + } + #[tokio::test] async fn provider_poll_installs_fail_closed_environment_and_acknowledges_policy() { let policy = proto_policy_fixture(); diff --git a/crates/openshell-supervisor/src/main.rs b/crates/openshell-supervisor/src/main.rs index df10d6b32a..4bfb469891 100644 --- a/crates/openshell-supervisor/src/main.rs +++ b/crates/openshell-supervisor/src/main.rs @@ -84,6 +84,10 @@ struct Args { #[arg(long, env = "OPENSHELL_HEALTH_SOCKET_PATH")] health_socket_path: Option, + /// TCP port that accepts connections only while the supervisor is ready. + #[arg(long, env = "OPENSHELL_HEALTH_PORT")] + health_port: Option, + #[arg(long)] upstream_proxy: Option, @@ -234,6 +238,7 @@ fn validate_role_arguments(args: &Args) -> Result<()> { || args.openshell_endpoint.is_some() || args.ssh_socket_path.is_some() || args.health_socket_path.is_some() + || args.health_port.is_some() || args.main_exit_marker.is_some() || args.parent_liveness_fd.is_some() { @@ -439,6 +444,7 @@ fn main() -> Result<()> { args.policy_data, args.ssh_socket_path, args.health_socket_path, + args.health_port, ocsf_enabled, ocsf_schema_version, upstream_proxy_args, diff --git a/crates/openshell-tui/src/app.rs b/crates/openshell-tui/src/app.rs index 7a53343142..4b218a37cf 100644 --- a/crates/openshell-tui/src/app.rs +++ b/crates/openshell-tui/src/app.rs @@ -420,6 +420,26 @@ fn apply_discovered_provider(form: &mut CreateProviderForm, discovered: Discover // Provider detail view (Get) // --------------------------------------------------------------------------- +/// Availability of strictly serialized provider profile YAML in the detail view. +pub enum ProviderProfileYaml { + /// The provider has no associated profile. + Absent, + /// The profile serialized successfully. + Valid(String), + /// A profile exists, but its YAML serialization failed. + Invalid, +} + +impl ProviderProfileYaml { + /// Return YAML only when strict profile serialization succeeded. + pub fn yaml(&self) -> Option<&str> { + match self { + Self::Valid(yaml) => Some(yaml), + Self::Absent | Self::Invalid => None, + } + } +} + pub struct ProviderDetailView { pub name: String, pub provider_id: String, @@ -430,7 +450,7 @@ pub struct ProviderDetailView { pub show_raw_provider: bool, pub raw_profile_scroll: usize, pub raw_provider_scroll: usize, - pub raw_profile_yaml: Option, + pub raw_profile_yaml: ProviderProfileYaml, pub raw_provider_yaml: String, pub profile_name: Option, pub profile_category: Option, @@ -656,6 +676,10 @@ pub struct App { pub sandbox_ages: Vec, pub sandbox_created: Vec, pub sandbox_images: Vec, + pub sandbox_restart_policies: Vec, + pub sandbox_restart_counts: Vec, + pub sandbox_exit_codes: Vec>, + pub sandbox_next_restart_at: Vec, pub sandbox_notes: Vec, pub sandbox_detail_notes: Vec, /// Formatted labels for each sandbox (e.g., "env=prod,team=platform" or empty string). @@ -1019,6 +1043,10 @@ impl App { sandbox_ages: Vec::new(), sandbox_created: Vec::new(), sandbox_images: Vec::new(), + sandbox_restart_policies: Vec::new(), + sandbox_restart_counts: Vec::new(), + sandbox_exit_codes: Vec::new(), + sandbox_next_restart_at: Vec::new(), sandbox_notes: Vec::new(), sandbox_detail_notes: Vec::new(), sandbox_labels: Vec::new(), @@ -2925,7 +2953,7 @@ impl App { KeyCode::Esc | KeyCode::Enter => { self.provider_detail = None; } - KeyCode::Char('y') if detail.raw_profile_yaml.is_some() => { + KeyCode::Char('y') if detail.raw_profile_yaml.yaml().is_some() => { detail.show_raw_profile = !detail.show_raw_profile; detail.show_raw_provider = false; detail.raw_profile_scroll = 0; @@ -2938,7 +2966,7 @@ impl App { KeyCode::Char('j') | KeyCode::Down if detail.show_raw_profile => { let max_scroll = detail .raw_profile_yaml - .as_ref() + .yaml() .map_or(0, |raw| raw.lines().count().saturating_sub(1)); detail.raw_profile_scroll = (detail.raw_profile_scroll + 1).min(max_scroll); } @@ -3433,9 +3461,12 @@ impl App { }, ); - let raw_profile_yaml = profile.and_then(|profile| { + let raw_profile_yaml = profile.map_or(ProviderProfileYaml::Absent, |profile| { let dto = ProviderTypeProfile::from_proto(profile); - openshell_providers::profile_to_yaml(&dto).ok() + // Serializer errors may echo authored values. Keep the failure + // distinct from absence without retaining potentially secret text. + openshell_providers::profile_to_yaml(&dto) + .map_or(ProviderProfileYaml::Invalid, ProviderProfileYaml::Valid) }); ProviderDetailView { @@ -3523,6 +3554,10 @@ impl App { self.sandbox_ages.clear(); self.sandbox_created.clear(); self.sandbox_images.clear(); + self.sandbox_restart_policies.clear(); + self.sandbox_restart_counts.clear(); + self.sandbox_exit_codes.clear(); + self.sandbox_next_restart_at.clear(); self.sandbox_notes.clear(); self.sandbox_detail_notes.clear(); self.sandbox_labels.clear(); @@ -3648,6 +3683,201 @@ mod tests { KeyEvent::new(code, KeyModifiers::NONE) } + fn detail_mcp_profile(versions: &[&str]) -> openshell_core::proto::ProviderProfile { + openshell_core::proto::ProviderProfile { + id: "mcp-example".to_string(), + display_name: "MCP Example".to_string(), + endpoints: vec![openshell_core::proto::NetworkEndpoint { + host: "mcp.example.com".to_string(), + port: 443, + protocol: "mcp".to_string(), + mcp: Some(openshell_core::proto::McpOptions { + versions: versions.iter().map(ToString::to_string).collect(), + ..Default::default() + }), + ..Default::default() + }], + ..Default::default() + } + } + + fn detail_app(profile: Option) -> App { + let mut app = test_app(); + let provider = openshell_core::proto::Provider { + metadata: Some(openshell_core::proto::ObjectMeta { + id: "provider-id".to_string(), + name: "example-provider".to_string(), + ..Default::default() + }), + r#type: "mcp-example".to_string(), + credentials: HashMap::from([("API_KEY".to_string(), "test-only-secret".to_string())]), + ..Default::default() + }; + app.provider_entries.push(ProviderListEntry { + provider: provider.clone(), + profile, + }); + app.provider_detail = Some(app.provider_detail_from_provider(&provider)); + app + } + + fn render_provider_detail(app: &App) -> String { + let mut terminal = ratatui::Terminal::new(ratatui::backend::TestBackend::new(100, 36)) + .expect("test terminal"); + terminal + .draw(|frame| crate::ui::create_provider::draw_detail(frame, app, frame.size())) + .expect("provider detail renders"); + terminal + .backend() + .buffer() + .content() + .iter() + .map(ratatui::buffer::Cell::symbol) + .collect() + } + + #[tokio::test] + async fn provider_detail_absent_profile_keeps_object_yaml_available() { + let mut app = detail_app(None); + let detail = app.provider_detail.as_ref().expect("detail view"); + assert!(matches!( + detail.raw_profile_yaml, + ProviderProfileYaml::Absent + )); + let text = render_provider_detail(&app); + assert!(text.contains("Profile: (legacy/unprofiled provider)")); + assert!(!text.contains("Profile YAML unavailable")); + assert!(!text.contains("[y]")); + + app.handle_provider_detail_key(key(KeyCode::Char('y'))); + assert!( + !app.provider_detail + .as_ref() + .expect("detail view") + .show_raw_profile + ); + app.handle_provider_detail_key(key(KeyCode::Char('o'))); + let text = render_provider_detail(&app); + assert!(text.contains("Provider Object YAML")); + assert!(text.contains("")); + assert!(!text.contains("test-only-secret")); + } + + #[tokio::test] + async fn provider_detail_valid_profile_preserves_yaml_navigation() { + let profile = detail_mcp_profile(&["2025-11-25"]); + let expected_yaml = + openshell_providers::profile_to_yaml(&ProviderTypeProfile::from_proto(&profile)) + .expect("valid profile YAML"); + let mut app = detail_app(Some(profile)); + assert_eq!( + app.provider_detail + .as_ref() + .expect("detail view") + .raw_profile_yaml + .yaml(), + Some(expected_yaml.as_str()) + ); + assert!(render_provider_detail(&app).contains("[y] Profile YAML")); + app.handle_provider_detail_key(key(KeyCode::Char('y'))); + let text = render_provider_detail(&app); + assert!(text.contains("Provider Profile YAML")); + assert!(text.contains("mcp-example")); + assert!(!text.contains("test-only-secret")); + app.handle_provider_detail_key(key(KeyCode::Char('j'))); + assert_eq!( + app.provider_detail + .as_ref() + .expect("detail view") + .raw_profile_scroll, + 1 + ); + app.handle_provider_detail_key(key(KeyCode::Char('k'))); + assert_eq!( + app.provider_detail + .as_ref() + .expect("detail view") + .raw_profile_scroll, + 0 + ); + app.handle_provider_detail_key(key(KeyCode::Char('o'))); + assert!( + app.provider_detail + .as_ref() + .expect("detail view") + .show_raw_provider + ); + app.handle_provider_detail_key(key(KeyCode::Char('y'))); + let detail = app.provider_detail.as_ref().expect("detail view"); + assert!(detail.show_raw_profile); + assert!(!detail.show_raw_provider); + app.handle_provider_detail_key(key(KeyCode::Esc)); + assert!( + !app.provider_detail + .as_ref() + .expect("detail view") + .show_raw_profile + ); + app.handle_provider_detail_key(key(KeyCode::Esc)); + assert!(app.provider_detail.is_none()); + } + + #[tokio::test] + async fn provider_detail_invalid_profile_remains_visible_without_yaml_action() { + for versions in [vec!["2025-11-25", "2025-11-25"], vec!["latest"]] { + let mut app = detail_app(Some(detail_mcp_profile(&versions))); + let detail = app.provider_detail.as_ref().expect("detail view"); + assert!(matches!( + detail.raw_profile_yaml, + ProviderProfileYaml::Invalid + )); + assert_eq!(detail.profile_name.as_deref(), Some("MCP Example")); + let text = render_provider_detail(&app); + assert!(text.contains("example-provider")); + assert!(text.contains("Profile YAML unavailable: serialization failed.")); + assert!(text.contains("Correct the profile at its source, then reopen this view.")); + assert!(!text.contains("legacy/unprofiled")); + assert!(!text.contains("[y]")); + app.handle_provider_detail_key(key(KeyCode::Char('y'))); + assert!( + !app.provider_detail + .as_ref() + .expect("detail view") + .show_raw_profile + ); + app.handle_provider_detail_key(key(KeyCode::Char('o'))); + let text = render_provider_detail(&app); + assert!(text.contains("Provider Object YAML")); + assert!(text.contains("")); + assert!(!text.contains("test-only-secret")); + } + } + + #[tokio::test] + async fn provider_detail_multibyte_profile_error_is_bounded_and_redacted() { + let malformed = "秘密é🚧\nprivate-profile-value".repeat(2048); + let profile = detail_mcp_profile(&[malformed.as_str()]); + let error = + openshell_providers::profile_to_yaml(&ProviderTypeProfile::from_proto(&profile)) + .expect_err("malformed revision must fail strict serialization"); + assert!(error.to_string().contains("秘密")); + let app = detail_app(Some(profile)); + assert!(matches!( + app.provider_detail + .as_ref() + .expect("detail view") + .raw_profile_yaml, + ProviderProfileYaml::Invalid + )); + let text = render_provider_detail(&app); + assert!(text.contains("Profile YAML unavailable: serialization failed.")); + assert!(!text.contains("秘密")); + assert!(!text.contains("private-profile-value")); + assert!(!text.contains("test-only-secret")); + let ordinary_failure = detail_app(Some(detail_mcp_profile(&["latest"]))); + assert_eq!(text, render_provider_detail(&ordinary_failure)); + } + #[tokio::test] async fn create_provider_enter_with_empty_profile_catalog_is_recoverable() { let mut app = test_app(); diff --git a/crates/openshell-tui/src/lib.rs b/crates/openshell-tui/src/lib.rs index 0ede05d87b..f235c4e003 100644 --- a/crates/openshell-tui/src/lib.rs +++ b/crates/openshell-tui/src/lib.rs @@ -23,8 +23,8 @@ use miette::{IntoDiagnostic, Result}; use openshell_bootstrap::list_gateways_with_source; use openshell_core::auth::EdgeAuthInterceptor; use openshell_core::metadata::{ObjectId, ObjectLabels, ObjectName, ObjectWorkspace}; -use openshell_core::proto::SandboxPhase; use openshell_core::proto::open_shell_client::OpenShellClient; +use openshell_core::proto::{SandboxPhase, SandboxRestartPolicy}; use ratatui::Terminal; use ratatui::backend::CrosstermBackend; use tokio::sync::mpsc; @@ -2767,6 +2767,41 @@ fn apply_sandbox_refresh(app: &mut App, sandboxes: Vec "on-failure", + Some(SandboxRestartPolicy::Always) => "always", + _ => "never", + } + .to_string() + }) + .collect(); + app.sandbox_restart_counts = sandboxes + .iter() + .map(|s| s.status.as_ref().map_or(0, |status| status.restart_count)) + .collect(); + app.sandbox_exit_codes = sandboxes + .iter() + .map(|s| s.status.as_ref().and_then(|status| status.exit_code)) + .collect(); + app.sandbox_next_restart_at = sandboxes + .iter() + .map(|s| { + format_timestamp( + s.status + .as_ref() + .and_then(|status| status.next_restart_time.as_ref()) + .and_then(|time| openshell_core::time::timestamp_to_millis(time).ok()) + .unwrap_or_default(), + ) + }) + .collect(); app.sandbox_images = sandboxes .iter() .map(|s| { diff --git a/crates/openshell-tui/src/ui/create_provider.rs b/crates/openshell-tui/src/ui/create_provider.rs index b68ddf5ea9..fcd7cdcc1a 100644 --- a/crates/openshell-tui/src/ui/create_provider.rs +++ b/crates/openshell-tui/src/ui/create_provider.rs @@ -6,7 +6,9 @@ use ratatui::layout::{Constraint, Direction, Layout, Rect}; use ratatui::text::{Line, Span}; use ratatui::widgets::{Block, Borders, Clear, Padding, Paragraph}; -use crate::app::{App, CreateProviderPhase, ProviderKeyField, UpdateProviderField}; +use crate::app::{ + App, CreateProviderPhase, ProviderKeyField, ProviderProfileYaml, UpdateProviderField, +}; use indexmap::IndexMap; use std::ops::Range; @@ -602,7 +604,7 @@ pub fn draw_detail(frame: &mut Frame<'_>, app: &App, area: Rect) { let title = if detail.show_raw_provider { " Provider Object YAML " - } else if detail.show_raw_profile { + } else if detail.show_raw_profile && detail.raw_profile_yaml.yaml().is_some() { " Provider Profile YAML " } else { " Provider Detail " @@ -629,13 +631,12 @@ pub fn draw_detail(frame: &mut Frame<'_>, app: &App, area: Rect) { return; } - if detail.show_raw_profile { + if detail.show_raw_profile + && let Some(raw) = detail.raw_profile_yaml.yaml() + { draw_raw_yaml( frame, - detail - .raw_profile_yaml - .as_deref() - .unwrap_or("No provider profile is available for this provider."), + raw, detail.raw_profile_scroll, "Summary", "y", @@ -676,6 +677,18 @@ pub fn draw_detail(frame: &mut Frame<'_>, app: &App, area: Rect) { t.status_warn, ))); } + if matches!(detail.raw_profile_yaml, ProviderProfileYaml::Invalid) { + // Fixed diagnostics bound the display and never expose values echoed by + // the serializer, even when malformed profile fields contain secrets. + lines.push(Line::from(Span::styled( + "Profile YAML unavailable: serialization failed.", + t.status_warn, + ))); + lines.push(Line::from(Span::styled( + "Correct the profile at its source, then reopen this view.", + t.status_warn, + ))); + } if let Some(description) = &detail.profile_description { lines.push(Line::from(vec![ Span::styled("Description: ", t.muted), @@ -704,7 +717,7 @@ pub fn draw_detail(frame: &mut Frame<'_>, app: &App, area: Rect) { Span::styled("[o]", t.key_hint), Span::styled(" Object YAML ", t.muted), ]; - if detail.raw_profile_yaml.is_some() { + if detail.raw_profile_yaml.yaml().is_some() { hint_spans.extend([ Span::styled("[y]", t.key_hint), Span::styled(" Profile YAML ", t.muted), diff --git a/crates/openshell-tui/src/ui/sandbox_detail.rs b/crates/openshell-tui/src/ui/sandbox_detail.rs index 341e46c71a..552bda6a59 100644 --- a/crates/openshell-tui/src/ui/sandbox_detail.rs +++ b/crates/openshell-tui/src/ui/sandbox_detail.rs @@ -8,7 +8,7 @@ use ratatui::widgets::{Block, Borders, Padding, Paragraph}; use crate::app::App; -const BASE_CONTENT_ROWS: u16 = 6; +const BASE_CONTENT_ROWS: u16 = 7; const BORDER_ROWS: u16 = 2; fn pending_draft_count(app: &App) -> usize { @@ -113,6 +113,32 @@ pub fn draw(frame: &mut Frame<'_>, app: &App, area: Rect) { Span::styled(age, t.text), ]); + let restart_policy = app + .sandbox_restart_policies + .get(idx) + .map_or("never", String::as_str); + let restart_count = app.sandbox_restart_counts.get(idx).copied().unwrap_or(0); + let exit_code = app + .sandbox_exit_codes + .get(idx) + .copied() + .flatten() + .map_or_else(|| "-".to_string(), |code| code.to_string()); + let next_restart = app + .sandbox_next_restart_at + .get(idx) + .map_or("-", String::as_str); + let row_restart = Line::from(vec![ + Span::styled(" Restart: ", t.muted), + Span::styled(restart_policy, t.text), + Span::styled(" Count: ", t.muted), + Span::styled(restart_count.to_string(), t.text), + Span::styled(" Exit: ", t.muted), + Span::styled(exit_code, t.text), + Span::styled(" Next: ", t.muted), + Span::styled(next_restart, t.text), + ]); + // Row 3: Labels let labels_str = app .sandbox_labels @@ -146,7 +172,7 @@ pub fn draw(frame: &mut Frame<'_>, app: &App, area: Rect) { Span::styled(providers_str, t.text), ]); - let mut lines = vec![row1, row2, row3, row4, row5]; + let mut lines = vec![row1, row2, row_restart, row3, row4, row5]; lines.extend( note_lines(app, area.width) .into_iter() diff --git a/deploy/helm/openshell/README.md b/deploy/helm/openshell/README.md index bfb60ddb80..df3ea173ad 100644 --- a/deploy/helm/openshell/README.md +++ b/deploy/helm/openshell/README.md @@ -355,7 +355,7 @@ discovery endpoint or its TLS CA. | pkiInitJob.timeoutSeconds | int | `120` | Maximum time in seconds for the certgen hook to poll for cert-manager certificates. When using cert-manager with BackendTLSPolicy, the hook polls for this many seconds waiting for the certificate to be issued, then creates the backend CA ConfigMap. The Job deadline is set to (timeoutSeconds + 30) to allow time for ConfigMap creation and cleanup. Increase this if cert-manager takes longer than 120 seconds to issue certificates. | | podAnnotations | object | `{}` | Extra annotations to add to the gateway pod. | | podLabels | object | `{}` | Extra labels to add to the gateway pod. | -| podLifecycle.terminationGracePeriodSeconds | int | `5` | Grace period, in seconds, before Kubernetes terminates the gateway pod. | +| podLifecycle.terminationGracePeriodSeconds | int | `30` | Maximum time, in seconds, Kubernetes waits for the gateway to exit before killing it. The gateway exits as soon as shutdown completes; the limit covers supervisor session cleanup and draining queued OCSF records. | | podSecurityContext.fsGroup | int | `1000` | fsGroup assigned to the gateway pod. | | probes.liveness.failureThreshold | int | `3` | Liveness probe failure threshold before the container is restarted. | | probes.liveness.initialDelaySeconds | int | `2` | Liveness probe initial delay, in seconds. | @@ -420,12 +420,21 @@ discovery endpoint or its TLS CA. | server.enableUserNamespaces | bool | `false` | Enable Kubernetes user namespace isolation (hostUsers: false) for sandbox pods. Requires Kubernetes 1.33+ with user namespace support available (beta through 1.35, GA in 1.36+), plus a supporting container runtime and Linux 5.12+. When enabled, container UID 0 maps to an unprivileged host UID and capabilities become namespaced. | | server.enableWebsocketTunnel | bool | `false` | Enable the WebSocket tunnel used by CLI/SDK clients behind an authenticated edge proxy. Leave disabled for direct gateway installs. | | server.externalDbSecret | string | `""` | Name of a pre-existing Opaque Secret containing a PostgreSQL connection URI (key: uri). When set, the gateway reads OPENSHELL_DB_URL from this Secret instead of using dbUrl. The Secret must contain a `uri` key, e.g. postgresql://user:pass@host:5432/dbname. | +| server.extraVolumeMounts | list | `[]` | Additional volume mounts for the gateway container. | +| server.extraVolumes | list | `[]` | Additional volumes for the gateway pod. | | server.grpcEndpoint | string | `""` | gRPC endpoint sandboxes call back into the gateway. Leave empty to derive it from the chart fullname, release namespace, service port, and disableTls flag, for example https://openshell.openshell.svc.cluster.local:8080. Override only when sandboxes must reach the gateway via a different hostname (e.g. an external ingress or a host alias). | | server.grpcRateLimit.requests | int | `0` | Maximum gRPC requests allowed per window. Must be positive (alongside windowSeconds) to enable rate limiting; 0 (default) disables it. | | server.grpcRateLimit.windowSeconds | int | `0` | gRPC rate-limit window length in seconds. Must be positive (alongside requests) to enable rate limiting; 0 (default) disables it. | | server.hostGatewayIP | string | `""` | Host gateway IP for sandbox pod hostAliases. When set, sandbox pods get hostAliases entries mapping host.docker.internal and host.openshell.internal to this IP, allowing them to reach services running on the Docker host. Auto-detected by the cluster entrypoint script. | | server.logLevel | string | `"info"` | Gateway log level. | | server.name | string | `""` | Operator-facing gateway name. Defaults to the chart fullname so all replicas in one installation share an identity. Set explicitly when one telemetry collector receives spans from multiple namespaces or clusters. | +| server.ocsfLog.enabled | bool | `false` | Write gateway OCSF events as JSONL. | +| server.ocsfLog.maxFiles | int | `7` | Rotated files retained when rotation is daily. | +| server.ocsfLog.path | string | `"/tmp/gateway-ocsf.jsonl"` | OCSF JSONL path. The default is writable in the gateway container but does not persist across restarts. To keep records, mount a volume with server.extraVolumes and server.extraVolumeMounts and set a path on it. When replicas share the volume, set subPathExpr: $(OPENSHELL_POD_NAME) on the mount so each replica writes its own file. | +| server.ocsfLog.queueCapacity | int | `10000` | Maximum records waiting for the file writer. | +| server.ocsfLog.queueMaxBytes | int | `16777216` | Maximum encoded bytes waiting for the file writer. | +| server.ocsfLog.rotation | string | `"daily"` | Rotate the active file daily in UTC, or never. | +| server.ocsfLog.schemaVersion | string | `""` | Optional OCSF downgrade target. Empty emits native OCSF 1.8.0. Supported values: "1.1", "1.3". | | server.oidc.adminRole | string | `""` | Role name for admin access. Leave empty (with userRole also empty) for authentication-only mode. Both must be set or both empty. | | server.oidc.audience | string | `"openshell-cli"` | Expected audience claim for the API resource server. This should match the server's --oidc-audience, NOT the CLI client ID. | | server.oidc.caConfigMapName | string | `""` | Name of a ConfigMap containing a CA certificate bundle (key: ca.crt) for verifying the OIDC issuer's TLS certificate. Required when the issuer uses a non-public CA (e.g. OpenShift ingress, private PKI). | diff --git a/deploy/helm/openshell/templates/_gateway-workload.tpl b/deploy/helm/openshell/templates/_gateway-workload.tpl index 0f2a8d6149..919d3a4313 100644 --- a/deploy/helm/openshell/templates/_gateway-workload.tpl +++ b/deploy/helm/openshell/templates/_gateway-workload.tpl @@ -184,6 +184,9 @@ spec: mountPath: {{ dir .Values.server.providerTokenGrants.spiffe.workloadApiSocketPath | quote }} readOnly: true {{- end }} + {{- with .Values.server.extraVolumeMounts }} + {{- toYaml . | nindent 8 }} + {{- end }} ports: - name: grpc containerPort: {{ .Values.service.port }} @@ -291,6 +294,9 @@ spec: driver: csi.spiffe.io readOnly: true {{- end }} + {{- with .Values.server.extraVolumes }} + {{- toYaml . | nindent 4 }} + {{- end }} {{- with .Values.nodeSelector }} nodeSelector: {{- toYaml . | nindent 4 }} diff --git a/deploy/helm/openshell/templates/gateway-config.yaml b/deploy/helm/openshell/templates/gateway-config.yaml index 0449871f7d..a3e0210a35 100644 --- a/deploy/helm/openshell/templates/gateway-config.yaml +++ b/deploy/helm/openshell/templates/gateway-config.yaml @@ -14,6 +14,7 @@ One value is intentionally NOT rendered here: */}} {{- $credentialDrivers := list -}} {{- $otlp := .Values.server.otlp | default dict -}} +{{- $ocsfLog := .Values.server.ocsfLog | default dict -}} {{- if .Values.server.credentialDrivers.kubernetesSecrets.enabled -}} {{- $credentialDrivers = append $credentialDrivers "kubernetes-secrets" -}} {{- end -}} @@ -88,6 +89,41 @@ data: {{- end }} {{- end }} + {{- if $ocsfLog.enabled }} + {{- if not $ocsfLog.path -}} + {{- fail "server.ocsfLog.path must be set when server.ocsfLog.enabled is true" -}} + {{- end }} + + {{- $ocsfRotation := $ocsfLog.rotation | default "daily" -}} + {{- if not (has $ocsfRotation (list "daily" "never")) -}} + {{- fail "server.ocsfLog.rotation must be daily or never" -}} + {{- end }} + {{- $ocsfQueueCapacity := int (ternary 10000 $ocsfLog.queueCapacity (kindIs "invalid" $ocsfLog.queueCapacity)) -}} + {{- $ocsfQueueMaxBytes := int (ternary 16777216 $ocsfLog.queueMaxBytes (kindIs "invalid" $ocsfLog.queueMaxBytes)) -}} + {{- if or (lt $ocsfQueueCapacity 1) (lt $ocsfQueueMaxBytes 1) -}} + {{- fail "server.ocsfLog.queueCapacity and queueMaxBytes must be positive" -}} + {{- end }} + [openshell.gateway.ocsf_log] + path = {{ $ocsfLog.path | quote }} + {{- $ocsfSchemaVersion := $ocsfLog.schemaVersion | default "" -}} + {{- if not (has $ocsfSchemaVersion (list "" "1.1" "1.3")) -}} + {{- fail "server.ocsfLog.schemaVersion must be empty, 1.1, or 1.3" -}} + {{- end }} + {{- if $ocsfSchemaVersion }} + schema_version = {{ $ocsfSchemaVersion | quote }} + {{- end }} + rotation = {{ $ocsfRotation | quote }} + {{- if eq $ocsfRotation "daily" }} + {{- $ocsfMaxFiles := int (ternary 7 $ocsfLog.maxFiles (kindIs "invalid" $ocsfLog.maxFiles)) -}} + {{- if lt $ocsfMaxFiles 1 -}} + {{- fail "server.ocsfLog.maxFiles must be positive when rotation is daily" -}} + {{- end }} + max_files = {{ $ocsfMaxFiles }} + {{- end }} + queue_capacity = {{ $ocsfQueueCapacity }} + queue_max_bytes = {{ $ocsfQueueMaxBytes }} + {{- end }} + {{- if not .Values.server.disableTls }} [openshell.gateway.tls] diff --git a/deploy/helm/openshell/tests/gateway_config_test.yaml b/deploy/helm/openshell/tests/gateway_config_test.yaml index eacfdfa2c4..16ccfaee6b 100644 --- a/deploy/helm/openshell/tests/gateway_config_test.yaml +++ b/deploy/helm/openshell/tests/gateway_config_test.yaml @@ -303,6 +303,166 @@ tests: path: data["gateway.toml"] pattern: '\[openshell\.gateway\.otlp\]' + - it: omits OCSF JSONL output by default + template: templates/gateway-config.yaml + asserts: + - notMatchRegex: + path: data["gateway.toml"] + pattern: '\[openshell\.gateway\.ocsf_log\]' + + - it: treats a null OCSF log map as disabled + template: templates/gateway-config.yaml + set: + server.ocsfLog: null + asserts: + - notMatchRegex: + path: data["gateway.toml"] + pattern: '\[openshell\.gateway\.ocsf_log\]' + + - it: omits OCSF JSONL output when a path is set without enabling it + template: templates/gateway-config.yaml + set: + server.ocsfLog.path: /var/openshell/gateway-ocsf.jsonl + asserts: + - notMatchRegex: + path: data["gateway.toml"] + pattern: '\[openshell\.gateway\.ocsf_log\]' + + - it: writes OCSF JSONL to the default container path when enabled + template: templates/gateway-config.yaml + set: + server.ocsfLog.enabled: true + asserts: + - matchRegex: + path: data["gateway.toml"] + pattern: '(?ms)\[openshell\.gateway\.ocsf_log\].*?path\s*=\s*"/tmp/gateway-ocsf\.jsonl"' + + - it: rejects enabled OCSF JSONL output without a path + template: templates/statefulset.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.path: "" + asserts: + - failedTemplate: + errorMessage: "server.ocsfLog.path must be set when server.ocsfLog.enabled is true" + + - it: renders OCSF JSONL output when configured + template: templates/gateway-config.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.path: /var/openshell/gateway-ocsf.jsonl + server.ocsfLog.schemaVersion: "1.3" + server.ocsfLog.rotation: daily + server.ocsfLog.maxFiles: 14 + server.ocsfLog.queueCapacity: 20000 + server.ocsfLog.queueMaxBytes: 33554432 + asserts: + - matchRegex: + path: data["gateway.toml"] + pattern: '(?ms)\[openshell\.gateway\.ocsf_log\].*?path\s*=\s*"/var/openshell/gateway-ocsf\.jsonl".*?schema_version\s*=\s*"1\.3".*?rotation\s*=\s*"daily".*?max_files\s*=\s*14.*?queue_capacity\s*=\s*20000.*?queue_max_bytes\s*=\s*33554432' + + - it: omits OCSF retention when rotation is disabled + template: templates/gateway-config.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.path: /var/openshell/gateway-ocsf.jsonl + server.ocsfLog.rotation: never + asserts: + - matchRegex: + path: data["gateway.toml"] + pattern: '(?ms)\[openshell\.gateway\.ocsf_log\].*?rotation\s*=\s*"never"' + - notMatchRegex: + path: data["gateway.toml"] + pattern: 'max_files\s*=' + + - it: rejects an unsupported OCSF rotation mode + template: templates/statefulset.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.path: /var/openshell/gateway-ocsf.jsonl + server.ocsfLog.rotation: hourly + asserts: + - failedTemplate: + errorMessage: "server.ocsfLog.rotation must be daily or never" + + - it: rejects a zero OCSF queueCapacity + template: templates/statefulset.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.queueCapacity: 0 + asserts: + - failedTemplate: + errorMessage: "server.ocsfLog.queueCapacity and queueMaxBytes must be positive" + + - it: rejects a zero OCSF queueMaxBytes + template: templates/statefulset.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.queueMaxBytes: 0 + asserts: + - failedTemplate: + errorMessage: "server.ocsfLog.queueCapacity and queueMaxBytes must be positive" + + - it: rejects a zero OCSF maxFiles + template: templates/statefulset.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.maxFiles: 0 + asserts: + - failedTemplate: + errorMessage: "server.ocsfLog.maxFiles must be positive when rotation is daily" + + - it: uses default OCSF limits when they are null + template: templates/gateway-config.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.queueCapacity: null + server.ocsfLog.queueMaxBytes: null + server.ocsfLog.maxFiles: null + asserts: + - matchRegex: + path: data["gateway.toml"] + pattern: '(?ms)\[openshell\.gateway\.ocsf_log\].*?max_files\s*=\s*7.*?queue_capacity\s*=\s*10000.*?queue_max_bytes\s*=\s*16777216' + + - it: rejects an unsupported OCSF schema version + template: templates/statefulset.yaml + set: + server.ocsfLog.enabled: true + server.ocsfLog.path: /var/openshell/gateway-ocsf.jsonl + server.ocsfLog.schemaVersion: "1.8" + asserts: + - failedTemplate: + errorMessage: "server.ocsfLog.schemaVersion must be empty, 1.1, or 1.3" + + - it: mounts extra gateway volumes for OCSF output in a Deployment + template: templates/deployment.yaml + set: + workload.kind: deployment + server.externalDbSecret: my-pg-secret + server.ocsfLog.enabled: true + server.ocsfLog.path: /var/log/openshell/gateway-ocsf.jsonl + server.extraVolumes: + - name: ocsf-log + persistentVolumeClaim: + claimName: openshell-ocsf-log + server.extraVolumeMounts: + - name: ocsf-log + mountPath: /var/log/openshell + subPathExpr: $(OPENSHELL_POD_NAME) + asserts: + - contains: + path: spec.template.spec.volumes + content: + name: ocsf-log + persistentVolumeClaim: + claimName: openshell-ocsf-log + - contains: + path: spec.template.spec.containers[0].volumeMounts + content: + name: ocsf-log + mountPath: /var/log/openshell + subPathExpr: $(OPENSHELL_POD_NAME) + - it: mounts the OIDC CA bundle when TLS is disabled template: templates/statefulset.yaml set: @@ -1242,3 +1402,19 @@ tests: - matchRegex: path: data["gateway.toml"] pattern: "(?ms)\\[openshell\\.drivers\\.kubernetes\\].*?image_pull_policy\\s*=\\s*\\\"never\\\".*?supervisor_image_pull_policy\\s*=\\s*\\\"never\\\"" + + - it: gives the gateway time to finish shutdown by default + template: templates/statefulset.yaml + asserts: + - equal: + path: spec.template.spec.terminationGracePeriodSeconds + value: 30 + + - it: keeps an explicit termination grace period + template: templates/statefulset.yaml + set: + podLifecycle.terminationGracePeriodSeconds: 3 + asserts: + - equal: + path: spec.template.spec.terminationGracePeriodSeconds + value: 3 diff --git a/deploy/helm/openshell/values.yaml b/deploy/helm/openshell/values.yaml index 4e7e11142d..4bf21b88a8 100644 --- a/deploy/helm/openshell/values.yaml +++ b/deploy/helm/openshell/values.yaml @@ -212,8 +212,10 @@ agentSandbox: # Pod restart behavior and health probe tuning. podLifecycle: - # -- Grace period, in seconds, before Kubernetes terminates the gateway pod. - terminationGracePeriodSeconds: 5 + # -- Maximum time, in seconds, Kubernetes waits for the gateway to exit before + # killing it. The gateway exits as soon as shutdown completes; the limit + # covers supervisor session cleanup and draining queued OCSF records. + terminationGracePeriodSeconds: 30 probes: startup: @@ -268,6 +270,31 @@ server: endpoint: "" # -- Gateway OpenTelemetry service name. Empty uses openshell-gateway. serviceName: "" + # Gateway-native OCSF JSONL output. + ocsfLog: + # -- Write gateway OCSF events as JSONL. + enabled: false + # -- OCSF JSONL path. The default is writable in the gateway container but + # does not persist across restarts. To keep records, mount a volume with + # server.extraVolumes and server.extraVolumeMounts and set a path on it. + # When replicas share the volume, set subPathExpr: $(OPENSHELL_POD_NAME) + # on the mount so each replica writes its own file. + path: /tmp/gateway-ocsf.jsonl + # -- Optional OCSF downgrade target. Empty emits native OCSF 1.8.0. + # Supported values: "1.1", "1.3". + schemaVersion: "" + # -- Rotate the active file daily in UTC, or never. + rotation: daily + # -- Rotated files retained when rotation is daily. + maxFiles: 7 + # -- Maximum records waiting for the file writer. + queueCapacity: 10000 + # -- Maximum encoded bytes waiting for the file writer. + queueMaxBytes: 16777216 + # -- Additional volumes for the gateway pod. + extraVolumes: [] + # -- Additional volume mounts for the gateway container. + extraVolumeMounts: [] # -- Enable anonymous OpenShell telemetry from the gateway and the sandbox # supervisors it launches. telemetryEnabled: true diff --git a/docs/how-it-works/gateways/configuration.mdx b/docs/how-it-works/gateways/configuration.mdx index fc9d5aa8c2..3b3c3aaaf0 100644 --- a/docs/how-it-works/gateways/configuration.mdx +++ b/docs/how-it-works/gateways/configuration.mdx @@ -276,6 +276,43 @@ The client-certificate handshake policy is derived and has no `require_client_au `[openshell.gateway.auth] allow_unauthenticated_users = true` is an unsafe local-development and trusted-proxy escape hatch. It accepts user-facing CLI/API calls without OIDC or mTLS credentials while sandbox supervisors still authenticate with gateway-minted sandbox JWTs. Leave it false for shared and production gateways. +## OCSF JSONL Output + +`[openshell.gateway.ocsf_log]` enables a gateway-local file containing native OCSF events emitted by the gateway, including events without a sandbox association. Omit the table to disable it. Restart the gateway after changing this configuration. + +Gateway records use `[openshell.gateway] name` as `device.uid` and `device.name`, shared across replicas; `device.hostname` identifies the replica. Assign a distinct name to each installation collected together. The default name does not guarantee uniqueness, and renaming an installation changes its audit identity. + +```toml +[openshell.gateway.ocsf_log] +path = "/var/log/openshell/gateway-ocsf.jsonl" +schema_version = "1.3" +rotation = "daily" +max_files = 7 +queue_capacity = 10000 +queue_max_bytes = 16777216 +``` + +| Field | Default | Behavior | +|---|---|---| +| `path` | Required | Nonempty file path; relative paths use the gateway's working directory. | +| `schema_version` | Native OCSF 1.8.0 | Optional downgrade target. Accepted values are `1.1` and `1.3`. The configured version applies to every record in the file. | +| `rotation` | `daily` | Rotate on the first write of a new UTC day. `never` allows unbounded file growth. | +| `max_files` | `7` | Positive number of rotated segments to retain, in addition to the active file. Rejects explicit use with `rotation = "never"`. | +| `queue_capacity` | `10000` | Positive queued-record limit. Overflow discards incoming records. | +| `queue_max_bytes` | `16777216` | Positive limit on queued encoded bytes. Records exceeding this limit are discarded. | + +Unknown fields are rejected. This is one file destination, not an exporter registry. The existing `ocsf_json_enabled` sandbox setting remains independent. + +For Helm installations, set `server.ocsfLog.enabled` to render this table and optionally set `server.ocsfLog.schemaVersion`. The default `server.ocsfLog.path`, `/tmp/gateway-ocsf.jsonl`, is writable in the gateway container with either the StatefulSet or Deployment workload but does not persist across pod restarts. To keep records, mount a volume with `server.extraVolumes` and `server.extraVolumeMounts` and set the path on it. When replicas share a volume, add `subPathExpr: $(OPENSHELL_POD_NAME)` to the mount so each replica writes its own file. + +Give each gateway replica its own writable file and an external shipper read access to the containing directory. New files use owner-only permissions on Unix; arrange shipper access deliberately. Rotation renames the active file to a date- and UUID-suffixed sibling. The shipper must follow rotated files and checkpoint its progress. Do not use multiple writers or external copy-truncate rotation on this path. + +The gateway queues OCSF independently of `RUST_LOG`, excludes ordinary diagnostics, and preserves native event IDs. It writes up to 100 records per batch with a 500 ms batching interval. One additional bounded batch can be in flight. File errors do not stop sandbox execution: failed writes are not replayed, and subsequent records are discarded during reopen backoff. A restarted writer removes an incomplete trailing line before appending. + +Metrics include `openshell_ocsf_log_queued_total`, `openshell_ocsf_log_written_total`, `openshell_ocsf_log_queue_records`, `openshell_ocsf_log_queue_bytes`, `openshell_ocsf_log_dropped_total` (with a bounded `reason` label), and `openshell_ocsf_log_writer_errors_total`. Recovery reports discarded bytes through `openshell_ocsf_log_recovery_discarded_bytes_total`, not an invented lost-record count. Warnings report writer failures independently of this file. + +Allow enough retention and disk space for shipper outages. Graceful shutdown allows five seconds for draining; outstanding OS writes remain uncertain if they outlast that budget. Flush does not mean fsync. The file is a best-effort collection boundary, not durable acceptance, a replay protocol, or proof of a complete audit history. See [OCSF JSON export](/observability/ocsf-json-export) for the separate sandbox-local output. + ## OTLP Export `[openshell.gateway.otlp]` enables OpenTelemetry export over OTLP/gRPC. Omit the table to disable export; there is no separate `enabled` flag. @@ -296,7 +333,7 @@ The OpenTelemetry SDK logs export failures after startup. Spans in a failed batc `service_name` sets the gateway's `service.name` resource attribute and defaults to `openshell-gateway`. The gateway also reports `service.version`, `openshell.gateway.name` from the gateway's configured `name`, and `openshell.gateway.compute_driver`. -Only OpenTelemetry traces are exported. Inbound gRPC and HTTP requests produce server spans named for the RPC or HTTP method. Store and compute-driver operations appear as child spans. Internal reconciliation, credential-refresh, and driver-watch loops create operation roots for their store work because no inbound request supplies a parent. The gateway continues valid W3C `traceparent` context and starts a new trace when none is supplied. Request spans carry `method`, `path`, and the `request_id` that also appears in gateway logs. Health endpoint spans use DEBUG level and are not exported by the default INFO filter. +Only OpenTelemetry traces are exported. Inbound gRPC and HTTP requests produce server spans named for the RPC or HTTP method. Compute-driver operations appear as child spans. Internal reconciliation and driver-watch loops create operation roots because no inbound request supplies a parent, and credential refresh creates one only when a credential is due or needs cleanup. The gateway continues valid W3C `traceparent` context and starts a new trace when none is supplied. Request spans carry `method`, `path`, and the `request_id` that also appears in gateway logs. Store operation spans, health endpoint spans, and spans for RPCs that sandbox supervisors call on a timer (`GetSandboxConfig` and `ReportProviderReadiness`, including its peer-forwarded form) use DEBUG level. The default INFO filter does not export them. Set the gateway log level to `debug` to include them. The gateway forwards the OTLP configuration, configured gateway name, configured compute driver, and W3C trace context to managed external drivers. Built-in drivers also export their spans to the same collector through dedicated in-process providers. Driver spans retain the gateway trace context, use a distinct service name such as `openshell-driver-docker` or `openshell-driver-podman`, and carry the same `openshell.gateway.name` and `openshell.gateway.compute_driver` resource attributes as gateway spans. Compute-driver client and server spans use the same fully qualified protobuf operation name, such as `openshell.compute.v1.ComputeDriver/CreateSandbox`, in both the span name and `rpc.method`. The service name and span kind distinguish each side. Backend-prefixed child spans identify implementation work. A streaming watch records a terminal status when observed; consumer teardown without a terminal status leaves the span status unset. Operator-run external drivers own their own telemetry configuration. @@ -896,6 +933,13 @@ no_proxy = ".svc.cluster.local,10.0.0.0/8" # Optional root-owned host file containing user:pass. An http:// proxy also # requires proxy_auth_allow_insecure = true as an explicit acknowledgement. proxy_auth_file = "/etc/openshell/secrets/proxy-auth" +proxy_auth_allow_insecure = false +# Last resort for proxy ACLs that require hostname CONNECT targets. The proxy +# then performs DNS resolution and becomes part of the effective egress boundary. +proxy_connect_by_hostname = false +# Operator-owned gateway-host PEM bundle trusted for an HTTPS proxy and for +# destination certificates re-signed by a TLS-intercepting proxy. +proxy_ca_bundle = "/etc/openshell/tls/proxy-ca.pem" # Project a host Unix Workload API socket into the supervisor for provider # token exchange. The socket parent must be a dedicated absolute directory. provider_spiffe_workload_api_socket = "/run/spire/agent.sock" @@ -904,6 +948,30 @@ provider_spiffe_workload_api_socket = "/run/spire/agent.sock" Use `sandbox_label` for Docker configurations. The legacy `sandbox_namespace` key is rejected. +Docker accepts `http://` and `https://` proxy URLs in explicit +`scheme://host:port` form. `no_proxy` bypasses only the corporate proxy; +OpenShell policy still applies. Proxy URLs cannot embed credentials. Supply a +`user:pass` file with `proxy_auth_file`, and set +`proxy_auth_allow_insecure = true` only to acknowledge that Basic credentials +travel in cleartext to an `http://` proxy. + +`proxy_ca_bundle` requires `https_proxy`, but the configured proxy URL may use +either scheme because a plain HTTP proxy can still intercept destination TLS. +The file must be a readable, bounded PEM bundle containing a usable certificate +authority. Docker validates it at gateway startup and again before each +supervisor launch. Missing, unreadable, empty, oversized, malformed, and +certificate-free bundles fail closed rather than falling back to default trust +or direct egress. + +The gateway-host path is operator-owned and cannot be selected through sandbox +environment, image contents, or `template.driver_config.docker`. The driver +copies the validated contents into its supervisor-only named volume and passes +the fixed path `/.openshell/supervisor/upstream-proxy-ca-bundle.pem` to the +companion supervisor. The bundle extends trust for the TLS connection to an +`https://` proxy, supervisor connections to re-signed upstream certificates, +and the combined trust bundle exposed to workload processes. The gateway-host +path is never mounted into or exposed to the workload container. + ### Podman Each Podman sandbox uses two containers. The workload container runs `openshell-sandbox` with `network=none`; the supervisor container runs on the host network and initiates policy-approved upstream connections. A private volume carries their authenticated Unix-domain socket. Configure guest mTLS paths once under `[openshell.gateway]`; the gateway validates and injects the bundle into the selected local driver. diff --git a/docs/how-it-works/policies/manage-policies.mdx b/docs/how-it-works/policies/manage-policies.mdx index 681f7f0f29..151c10414c 100644 --- a/docs/how-it-works/policies/manage-policies.mdx +++ b/docs/how-it-works/policies/manage-policies.mdx @@ -340,6 +340,7 @@ code in the response to find the cause: | `request_authority_mismatch` | The HTTP request's host or port differs from the connection's destination. | The client's `Host` header, including any non-default port, matches the connection. | | `credential_endpoint_mismatch` | A network rule allowed the request, but the provider credential is not bound to this destination. | The provider's profile endpoints or credential binding. Do not widen the network rule. | | `credential_placeholder_in_request_body` | The request body contains an invalid or revoked credential placeholder. | Remove the stale placeholder or restore the provider. | +| `unsupported_l7_protocol` | The request used a protocol or upgrade that the endpoint cannot inspect, such as h2c or an `Upgrade` header sent to a GraphQL, MCP, or JSON-RPC endpoint. No rule can allow it. | Send the request without the upgrade, or allow WebSocket traffic through a separate `protocol: websocket` endpoint. | ### A Change Fails to Load diff --git a/docs/how-it-works/policies/network-rules.mdx b/docs/how-it-works/policies/network-rules.mdx index bcaba59218..811d920922 100644 --- a/docs/how-it-works/policies/network-rules.mdx +++ b/docs/how-it-works/policies/network-rules.mdx @@ -490,6 +490,8 @@ are application-specific, so review them against the service's schema before you rely on them. For deny rules, persisted queries, and GraphQL over WebSocket, refer to [GraphQL Rules](/how-it-works/policies/schema#graphql-rules). +GraphQL endpoints refuse requests that carry an `Upgrade` header with `403 Forbidden` in both `enforce` and `audit` mode. Serve GraphQL over WebSocket from a separate `protocol: websocket` endpoint on another path or port, with GraphQL operation rules. One path on a host and port can use only one of the two protocols. + With a GitHub provider attached, verify a REST read with `gh api zen` and a query with `gh api graphql -f query='{ viewer { login } }'`, then confirm that a mutation other than `createIssue` returns an OpenShell denial. @@ -531,13 +533,59 @@ Verify initialization and `read_status` before confirming that `delete_resource` returns a policy denial. Tool argument matching is not supported, so an allowed tool can receive any arguments accepted by the server. -Omitting `mcp.versions` allows only the `2025-11-25` revision. To support an -older server, list the exact revisions it needs, as described in [MCP Version +Omitting `mcp.versions` allows only the `2025-11-25` revision. To support a +server that uses another revision, list the exact revisions it needs, as +described in [MCP Version Selection](/how-it-works/policies/schema#mcp-version-selection). Server responses and SSE messages are relayed without MCP policy parsing. Do not put an MCP endpoint on the same host and port as an endpoint that uses a different protocol. `openshell policy update` rejects this combination. +A server that uses the sessionless `2026-07-28` revision has no initialization +step. Its clients send their protocol version and capabilities with every +request, and can call `server/discover` to learn what the server supports. This +rule allows discovery and `read_status` on such a server: + +```yaml +mcp_server: + endpoints: + - host: mcp.example.com + port: 443 + path: /mcp + protocol: mcp + enforcement: enforce + mcp: + versions: ["2026-07-28"] + rules: + - allow: + method: server/discover + - allow: + method: tools/list + - allow: + method: tools/call + tool: read_status + binaries: + - path: /usr/bin/python3.12 +``` + +OpenShell rejects a `2026-07-28` request that lacks a required +`MCP-Protocol-Version`, `Mcp-Method`, or `Mcp-Name` header, or whose headers do +not match its body. Refer to [Sessionless MCP +Requests](/how-it-works/policies/schema#sessionless-mcp-requests) for the +complete request checks. To serve clients of both revisions on one endpoint, +list both revisions and combine the rules from the two examples. OpenShell +inspects each request under the revision that the request selects. + +MCP and JSON-RPC endpoints carry only HTTP requests, because their rules apply +to each request. OpenShell answers a request to these endpoints that carries an +`Upgrade` header, such as a WebSocket upgrade, with `403 Forbidden` before it +reaches the server, in both `enforce` and `audit` mode. If the server also +accepts WebSocket connections, allow them with a separate +`protocol: websocket` endpoint, which applies `GET` and `WEBSOCKET_TEXT` rules, +not MCP method or tool rules. For an MCP server, put that endpoint on a +different host or port. A JSON-RPC endpoint can share its host and port with a +WebSocket endpoint that uses a different `path`. + ### Allow Native TCP Use `protocol: tcp` for a client that speaks a protocol other than HTTP, such diff --git a/docs/how-it-works/policies/schema.mdx b/docs/how-it-works/policies/schema.mdx index a4801b0b0a..a190cd5a91 100644 --- a/docs/how-it-works/policies/schema.mdx +++ b/docs/how-it-works/policies/schema.mdx @@ -50,6 +50,8 @@ that are not listed are inaccessible. When the effective policy has at least one network rule, OpenShell also adds the baseline paths described in [Default Policy](/how-it-works/policies/default-policy#baseline-filesystem-paths). +These defaults and the filesystem path restrictions below also apply when the supervisor or standalone proxy loads local policy files directly. Local loading enforces the same non-root process identities as typed policy. It rejects invalid filesystem, Landlock, and process field types, unknown fields in those sections, and explicit `null` JSON-RPC or MCP options. Network settings must also use the declared types: for example, `allow_encoded_slash: "true"` and an array-valued deny-rule `method` are rejected. A rejected reload keeps the active OPA policy and its generation unchanged. + Each path must be absolute, must not contain `..`, and must not exceed 4096 bytes. `read_write` cannot contain `/`. A policy can list at most 256 paths. @@ -152,7 +154,7 @@ the gateway host. | Field | Type | Default | Description | |---|---|---|---| -| `protocol` | string | None | `rest`, `websocket`, `graphql`, `mcp`, or `json-rpc` for request inspection, or `tcp` for a native TCP connection. Refer to [Connection and Request Checks](/how-it-works/policies/network-rules#connection-and-request-checks). | +| `protocol` | string | None | `rest`, `websocket`, `graphql`, `mcp`, or `json-rpc` for request inspection, or `tcp` for a native TCP connection. `graphql`, `mcp`, and `json-rpc` endpoints refuse requests that carry an `Upgrade` header with `403`. Refer to [Connection and Request Checks](/how-it-works/policies/network-rules#connection-and-request-checks). | | `tls` | string | Automatic | `skip` relays traffic without terminating TLS, so OpenShell cannot inspect it. Do not use it with a request protocol. | | `enforcement` | string | `audit` | `enforce` blocks requests that break the endpoint's rules. `audit` logs them and allows the request. | | `access` | string | None | Access preset: `read-only`, `read-write`, or `full`. Refer to [Access Presets](#access-presets). | @@ -203,7 +205,7 @@ network_policies: | `mcp.versions` | list of strings | `["2025-11-25"]` | Allowed MCP revisions. Refer to [MCP Version Selection](#mcp-version-selection). | | `mcp.max_body_bytes` | integer | `65536` | Maximum MCP request body size for inspection. | | `mcp.strict_tool_names` | bool | `true` | Requires tool names to match `^[A-Za-z0-9_.-]{1,128}$`. | -| `mcp.allow_all_known_mcp_methods` | bool | `false` | When `true`, the endpoint allows every MCP method except those that deny rules match. If rules name specific tools, `tools/call` is limited to those tools. Rules can omit `method`. Refer to [MCP Rules](#mcp-rules). | +| `mcp.allow_all_known_mcp_methods` | bool | `false` | Allow core MCP methods for the selected request revision. Tool restrictions still apply, and deny rules take precedence. Refer to [Allow Core MCP Methods](#allow-core-mcp-methods) for the behavior and [Core MCP Methods](#core-mcp-methods) for the method list. | | `json_rpc.max_body_bytes` | integer | `65536` | Maximum JSON-RPC request body size for inspection. | #### Endpoint Constraints @@ -347,9 +349,11 @@ matcher fields in one rule. - Tool arguments are not matched, so an allowed tool accepts any arguments. - One denied call denies an entire batched request. - Server responses and server-to-client messages are not inspected. +- Only an allow rule that names an extension method exactly can allow it. + Method globs and `mcp.allow_all_known_mcp_methods` do not. An extension + method is one that no supported MCP revision defines. -A client sends `initialize` and `notifications/initialized` before calling -tools, so allow both: +Under the 2025 revisions, a client sends `initialize` and `notifications/initialized` before calling tools. When `mcp.allow_all_known_mcp_methods` is omitted or `false`, allow both explicitly: ```yaml showLineNumbers={false} rules: @@ -366,23 +370,136 @@ deny_rules: tool: send_email ``` +OpenShell uses the `tower-mcp-types` library to inspect each request under its +MCP revision. It checks the JSON-RPC structure, whether a client can send the +method in that revision, and the parameter types of known methods. It rejects a +body that repeats a key within any JSON object, and a method that MCP defines +but the request's revision does not, such as `server/discover` under +`2025-11-25`. A request that fails these checks returns `400` with the error +`invalid_mcp_request`, even on an endpoint that uses `enforcement: audit`. These +type checks are not complete JSON Schema validation. + +#### Allow Core MCP Methods + +Set `mcp.allow_all_known_mcp_methods: true` to allow the [core methods](#core-mcp-methods) available in each request's [selected MCP revision](#mcp-version-selection). A policy that accepts several revisions still checks each request against its own revision. The option does not combine their method sets. + +Tool-specific allow rules continue to restrict `tools/call` to matching names, and deny rules take precedence over every allow. Without tool-specific allow rules, this option allows calls to every tool name, subject to deny rules and tool-name validation. Extension methods require an allow rule that names the method exactly. An explicit allow rule cannot enable a core method that is unavailable in the selected revision. + +For example, this endpoint allows core methods for `2025-11-25`, while limiting tool calls to `search_web` and `read_document`: + +```yaml showLineNumbers={false} +endpoints: + - host: mcp.example.com + port: 443 + protocol: mcp + enforcement: enforce + mcp: + versions: ["2025-11-25"] + allow_all_known_mcp_methods: true + rules: + - allow: + tool: + any: [search_web, read_document] +``` + +Here, `tools/call` is the JSON-RPC method, and `search_web` is the tool name in `params.name`. Allowing a method does not establish that the server supports it or that a tool is read-only. + +#### Core MCP Methods + +The following list covers client-to-server requests and notifications that OpenShell inspects over HTTP. Server-to-client methods, such as `sampling/createMessage`, `roots/list`, and `elicitation/create`, are outside this inspection. Each row lists the revisions that accept those methods over HTTP in the client-to-server direction. + +| Methods or notifications | MCP revisions | +|---|---| +| `tools/list`, `tools/call` | All supported revisions. | +| `resources/list`, `resources/templates/list`, `resources/read` | All supported revisions. | +| `prompts/list`, `prompts/get`, `completion/complete` | All supported revisions. | +| `notifications/cancelled` | `2025-03-26`, `2025-06-18`, `2025-11-25`. | +| `initialize`, `ping`, `logging/setLevel` | `2025-03-26`, `2025-06-18`, `2025-11-25`. | +| `resources/subscribe`, `resources/unsubscribe` | `2025-03-26`, `2025-06-18`, `2025-11-25`. | +| `notifications/initialized`, `notifications/progress`, `notifications/roots/list_changed` | `2025-03-26`, `2025-06-18`, `2025-11-25`. | +| `tasks/get`, `tasks/result`, `tasks/list`, `tasks/cancel`, `notifications/tasks/status` | `2025-11-25` only; experimental. | +| `server/discover`, `subscriptions/listen` | `2026-07-28` only. | + +For `2026-07-28` HTTP requests, cancel a request by closing its response stream. OpenShell rejects `notifications/cancelled` over HTTP in that revision. + +MCP introduced [tasks](https://modelcontextprotocol.io/specification/2025-11-25/basic/utilities/tasks) as an experimental feature in `2025-11-25`, so the March and June 2025 revisions do not include the task methods or `notifications/tasks/status`. The default policy revision includes them, but using tasks also requires the peer's task capabilities and, for task-augmented tool calls, the tool's task-support declaration. The policy option does not negotiate these capabilities. + #### MCP Version Selection `mcp.versions` lists the MCP revisions an endpoint accepts: `2025-03-26`, -`2025-06-18`, or `2025-11-25`. When omitted, only `2025-11-25` is allowed. +`2025-06-18`, `2025-11-25`, or `2026-07-28`. When omitted, the list is exactly +`["2025-11-25"]`, so the endpoint does not accept `2026-07-28` until you list +it. ```yaml showLineNumbers={false} mcp: versions: ["2025-03-26", "2025-11-25"] ``` -OpenShell does not check the revision of a single `initialize` request, because -the client negotiates the revision in that request. For other requests, -OpenShell reads the revision from the `MCP-Protocol-Version` header, or uses -`2025-03-26` when the header is absent. A duplicate, empty, or -unsupported header value returns `400`, and a supported revision that the -endpoint does not allow returns `403`. For a client that requires an unsupported -revision, omit `protocol` and `mcp` to allow its traffic without MCP inspection. +OpenShell selects each request's revision from its `MCP-Protocol-Version` +header. When the header is absent, OpenShell selects `2025-03-26`, so an +endpoint whose clients omit the header must allow `2025-03-26`. A duplicate, +empty, or unsupported header value returns `400`, and a supported revision that +the endpoint does not allow returns `403`. + +OpenShell inspects each request under its selected revision only. For example, +only `2025-03-26` accepts JSON-RPC batches, with at most 64 messages each, and +later revisions reject a request body that is a JSON array. + +Under the 2025 revisions, a client starts with a single `initialize` request +that proposes a revision in its body, so OpenShell does not check that request's +revision against `mcp.versions`. An endpoint that allows only `2026-07-28` +rejects `initialize`. + +For a client that requires an unsupported revision, omit `protocol` and `mcp` to +allow its traffic without MCP inspection. + +#### Understand MCP Rejections + +Use the response's error code and detail to distinguish a protocol problem from a policy denial: + +- `unsupported_mcp_protocol_version` (`400`) means this OpenShell build does not support the header's revision. Configure the client to use a revision that OpenShell and the server support and that `mcp.versions` permits. +- `mcp_protocol_version_not_allowed` (`403`) means the selected revision is supported but excluded by the endpoint policy. Use a compatible revision the policy permits, or ask the policy owner to review `mcp.versions`. +- A detail that names the missing-header fallback means OpenShell selected `2025-03-26`, even if the policy defaults to `2025-11-25`. Send the client/server revision explicitly in `MCP-Protocol-Version`; the policy must permit it. +- `invalid_mcp_request` (`400`) can mean the method is unavailable in the selected revision. For example, `tasks/get` is unavailable in `2025-06-18`. Use a method available in that revision, or configure the client and server to use a compatible revision that includes it and is permitted by policy. Adding an allow rule or enabling `allow_all_known_mcp_methods` cannot override this protocol check. +- A rule denial (`403` under `enforcement: enforce`) requires a policy review. An extension method needs an exact method allow rule; a denied tool call may need a matching tool selector. An explicit deny rule takes precedence, so adding an allow rule cannot override it. Ask the policy owner to review the relevant allow or deny rule. + +#### Sessionless MCP Requests + +The `2026-07-28` revision has no `initialize` handshake or session. Each request +carries the client's protocol version and capabilities. A client can call +`server/discover` to learn what the server supports, and can call +`subscriptions/listen` to receive notifications on that request's response +stream. Each HTTP `POST` carries one JSON-RPC request or extension notification. + +OpenShell checks each `2026-07-28` request for the following: + +- `params._meta` contains `io.modelcontextprotocol/protocolVersion` set to + `2026-07-28`, and `io.modelcontextprotocol/clientCapabilities`. Extension + requests need this metadata too. +- The `MCP-Protocol-Version` header is `2026-07-28`. Without it, OpenShell + selects `2025-03-26` and rejects the request. +- The `Mcp-Method` header equals the JSON-RPC `method`. +- For `tools/call` and `prompts/get`, the `Mcp-Name` header equals + `params.name`. For `resources/read`, it equals `params.uri`. OpenShell decodes + a `=?base64?...?=` value before comparing it. + +A request with missing or invalid metadata, or with a missing, duplicate, or +mismatched `Mcp-Method` or `Mcp-Name` header, returns `400`. OpenShell also +rejects `initialize`, `notifications/initialized`, `notifications/cancelled`, +JSON-RPC responses from the client, and batches. A request that uses an HTTP +method other than `POST` returns `405`. + +OpenShell repeats these checks after middleware changes a request. A middleware +that replaces the body of a `2026-07-28` request must also update `Mcp-Method` +and `Mcp-Name` to match the new body. + +An extension notification needs `MCP-Protocol-Version: 2026-07-28` and an allow +rule that names its `method` exactly, such as `method: vendor/notice`. It does +not need `params._meta`, `Mcp-Method`, or `Mcp-Name`. + +OpenShell forwards `Mcp-Param-*` headers without comparing them with tool +arguments. ### JSON-RPC Rules @@ -391,6 +508,7 @@ revision, omit `protocol` and `mcp` to allow its traffic without MCP inspection. | `method` | string | Yes | Exact method name, or `*` for all methods. Other globs are rejected. | Parameters are not matched. One denied call denies an entire batched request. +OpenShell rejects a body that repeats a key within any JSON object. ```yaml showLineNumbers={false} rules: diff --git a/docs/how-it-works/providers/profiles.mdx b/docs/how-it-works/providers/profiles.mdx index f28124c82c..11c5438d3c 100644 --- a/docs/how-it-works/providers/profiles.mdx +++ b/docs/how-it-works/providers/profiles.mdx @@ -238,6 +238,8 @@ openshell profile describe github -o yaml The description shows metadata, credential names and authentication settings, endpoints with their protocols and policy rule counts, allowed binaries, source, and scope. Endpoint details include TLS handling, permission for uninspected credential traffic, and MCP method and tool-name settings. A `tls: skip` endpoint is shown as a raw tunnel without L7 inspection or credential rewrite, even when it declares an L7 protocol and rules. Use JSON or YAML output to inspect the complete rule definitions. The description reads the reusable definition without reading credential values from provider instances. A missing profile ID returns an error. +For MCP endpoints, the text output labels `allow_all_known_mcp_methods` as `Allow core MCP methods for the selected revision` and shows the profile's declared revisions. These are configuration values; `describe` does not observe a connection or its negotiated capabilities. At request time, OpenShell selects the revision from `MCP-Protocol-Version`, or uses `2025-03-26` when the header is absent. Legacy standalone `initialize` requests negotiate their revision in the body, as described in [MCP version selection](/how-it-works/policies/schema#mcp-version-selection). The option allows [core methods for that revision](/how-it-works/policies/schema#core-mcp-methods), subject to tool restrictions and deny rules. Without tool-specific allow rules, it allows every tool name, subject to deny rules and tool-name validation. Extension methods still require an exact allow rule. JSON and YAML output retain the `allow_all_known_mcp_methods` field name. + List and describe use the selected workspace's effective catalog. Pass the inherited `--workspace` flag to select another workspace, or `--global` to target platform scope. Platform operations require Platform Admin access. ```shell diff --git a/docs/how-it-works/sandboxes/overview.mdx b/docs/how-it-works/sandboxes/overview.mdx index 4fd762ce91..8194b0df5b 100644 --- a/docs/how-it-works/sandboxes/overview.mdx +++ b/docs/how-it-works/sandboxes/overview.mdx @@ -20,8 +20,8 @@ Create a sandbox with a single command. For example, to create a sandbox with Cl openshell sandbox create --from registry.example.com/your-org/claude-agent:latest -- claude ``` -The trailing command is the sandbox's canonical main process. OpenShell starts -it once, streams its output, and returns its exit status. Exit code 0 leaves a +The trailing command is the sandbox's canonical main process. OpenShell streams +its output and returns its exit status. By default, exit code 0 leaves a retained sandbox in `Completed`; a nonzero exit leaves it in `Error` with a `MainProcessFailed` condition. With no trailing command, OpenShell starts a login shell in a retained pseudo-terminal: `/bin/bash -l` when the image @@ -33,6 +33,20 @@ attaching: openshell sandbox create --name worker --detach -- ./worker ``` +Use `--restart-policy on-failure` to replace the runtime after a nonzero main +process exit, or `--restart-policy always` to replace it after every exit: + +```shell +openshell sandbox create --name worker --detach --restart-policy on-failure -- ./worker +``` + +The default policy is `never`. The gateway starts the first replacement as soon +as terminal output has been delivered. Repeated quick exits wait 1 second, then +double the delay up to 3 minutes. Running for 10 seconds resets the delay. A +replacement starts a new supervisor and main process. It retains the sandbox +identity and configuration; filesystem persistence follows the compute driver's stop and +start behavior. It does not restore process memory or terminal history. + Detached commands have no attachment grace period. When the command exits, OpenShell records its terminal phase immediately. For a foreground command, the create request declares one expected SSH attachment. OpenShell retains the @@ -252,11 +266,11 @@ Attach to the canonical main process in a running sandbox: openshell sandbox connect my-sandbox ``` -Press `Ctrl-P`, then `Ctrl-Q` in sequence to disconnect. The main process keeps -running and its stdin stays open. A later `connect` attaches to the same process -instance and replays up to 1 MiB of recent output. One attachment owns stdin -at a time. Use `sandbox exec --tty -- /bin/bash -l` when you want a new -independent shell instead. +Press `Ctrl-D` or `Ctrl-P`, then `Ctrl-Q` in sequence to disconnect. The main +process keeps running and its stdin stays open. A later `connect` attaches to +the same process instance and replays up to 1 MiB of recent output. One +attachment owns stdin at a time. Use `sandbox exec --tty -- /bin/bash -l` when +you want a new independent shell instead. If an established connection is interrupted, for example when a laptop sleeps and wakes, the CLI obtains a new SSH session and reattaches to the same main @@ -264,12 +278,17 @@ process. It retries transient transport failures for up to 60 seconds. Initial authentication failures, sandbox lifecycle changes, and clean SSH exits are not retried. +The supervisor sends SSH keepalive probes after 15 seconds without inbound traffic and closes a connection after 60 seconds without receiving peer bytes, including when relay writes are stalled. Healthy idle clients answer these probes and stay attached. The timeout starts when the supervisor last reads peer bytes; relay buffering and scheduling can delay detection after physical network loss. The canonical process keeps running when its dead attachment closes. + +If recovery reaches the sandbox before the old connection releases stdin, the replacement reports `attached read-only`. After the old connection times out, type again: the replacement acquires stdin if it is available, reports `input enabled`, and forwards that new input. Keystrokes sent while another attachment owned stdin are discarded. An explicitly read-only attachment stays read-only, and recovery never displaces a healthy input owner. + `Ctrl-C` retains its normal terminal behavior and interrupts the foreground -process. For read-only attachments, `Ctrl-C` only exits the current viewer. +process. For read-only attachments, `Ctrl-C` or `Ctrl-D` exits only the current +viewer. OpenSSH's `~.` escape reports the same status as a broken transport, so it -starts automatic recovery instead of exiting. After `~.`, use `Ctrl-P`, then -`Ctrl-Q` once OpenShell reattaches, or press `Ctrl-C` while the CLI is between -retry attempts to cancel recovery. +starts automatic recovery instead of exiting. After `~.`, use `Ctrl-D` or +`Ctrl-P`, then `Ctrl-Q` once OpenShell reattaches, or press `Ctrl-C` while the CLI +is between retry attempts to cancel recovery. Launch VS Code or Cursor directly into the sandbox workspace: @@ -459,13 +478,16 @@ name selects the unnamed endpoint. The returned sandbox includes a service URL map keyed by those names; use the empty key for the unnamed endpoint: ```python -from openshell import SandboxClient, ServiceExposure +from openshell import SandboxClient, ServiceAuthorizationMode, ServiceExposure with SandboxClient.from_active_cluster() as client: sandbox = client.create( workspace="default", name="app-server", - service_exposures=[ServiceExposure(target_port=4500)], + service_exposures=[ServiceExposure( + target_port=4500, + authorization_mode=ServiceAuthorizationMode.BEARER_PASSTHROUGH, + )], ) print(sandbox.service_urls[""]) ``` @@ -474,7 +496,10 @@ with SandboxClient.from_active_cluster() as client: const sandbox = await client.sandbox.create({ name: 'app-server', image: 'base', - serviceExposures: [{ targetPort: 4500 }], + serviceExposures: [{ + targetPort: 4500, + authorizationMode: ServiceAuthorizationMode.BearerPassthrough, + }], }) console.log(sandbox.serviceUrls['']) ``` @@ -483,7 +508,10 @@ console.log(sandbox.serviceUrls['']) sandbox, err := client.Sandboxes().Create( ctx, "default", "app-server", spec, nil, v1.CreateOptions{ServiceExposures: []v1.ServiceExposure{ - {TargetPort: 4500}, + { + TargetPort: 4500, + AuthorizationMode: v1.ServiceAuthorizationModeBearerPassthrough, + }, }}, ) fmt.Println(sandbox.ServiceURLs[""]) @@ -495,6 +523,7 @@ let sandbox = client.create_sandbox(openshell_sdk::SandboxSpec { service_exposures: vec![openshell_sdk::ServiceExposure { service: String::new(), target_port: 4500, + authorization_mode: openshell_sdk::ServiceAuthorizationMode::BearerPassthrough, }], ..Default::default() }).await?; @@ -513,6 +542,63 @@ Pass an optional service name to create a named service URL: openshell service expose my-sandbox 8080 web ``` +OpenShell strips the incoming `Authorization` header by default. Opt a service +into forwarding one valid bearer credential unchanged when the application +performs its own authentication: + +```shell +openshell service expose my-sandbox 4500 \ + --authorization-mode bearer-passthrough +``` + +For a create-time exposure, use both flags: + +```shell +openshell sandbox create \ + --expose 4500 \ + --expose-authorization-mode bearer-passthrough \ + --detach \ + -- ./authenticated-server +``` + +In `bearer-passthrough` mode, OpenShell accepts zero or one application +`Authorization` header. When present, it must contain a nonempty Bearer +credential. Missing credentials reach the application so it can return its own +authentication response. Duplicate, Basic, and malformed credentials fail with +`400 Bad Request`. Gateway and edge identity headers, proxy authorization, and +edge authentication cookies remain stripped. + + +Bearer passthrough delivers the caller's credential to the sandbox service. +Enable it only when that service is trusted to receive the credential. Exposed +services still use the gateway listener and its TLS configuration in this +release. + + +For example, Codex App Server can keep the raw capability outside the sandbox +and receive only its SHA-256 verifier: + +```shell +APP_SERVER_TOKEN="$(openssl rand -hex 32)" +APP_SERVER_TOKEN_SHA256="$(printf %s "$APP_SERVER_TOKEN" | openssl dgst -sha256 -hex | awk '{print $2}')" + +openshell sandbox create \ + --name codex-server \ + --env "APP_SERVER_TOKEN_SHA256=$APP_SERVER_TOKEN_SHA256" \ + --expose 4500 \ + --expose-authorization-mode bearer-passthrough \ + --detach \ + -- codex app-server \ + --listen ws://127.0.0.1:4500 \ + --ws-auth capability-token \ + --ws-token-sha256 "$APP_SERVER_TOKEN_SHA256" +``` + +Clients send `Authorization: Bearer $APP_SERVER_TOKEN` in the WebSocket +handshake. Codex's WebSocket transport is experimental and unsupported for +production workloads. It authenticates the handshake before the app-server +`initialize` request. + List exposed endpoints: ```shell @@ -533,7 +619,8 @@ openshell service list my-sandbox --output yaml ``` Structured list output contains `services` and `next_page_token` fields. Each -record contains `workspace`, `sandbox`, `service`, `target_port`, and `url`. +record contains `workspace`, `sandbox`, `service`, `target_port`, +`authorization_mode`, and `url`. The unnamed service uses an empty `service` string. Pass the returned token to `--page-token` to continue. An empty result has an empty `services` collection. @@ -552,7 +639,14 @@ openshell service delete my-sandbox ``` -Loopback gateways return local `openshell.localhost` URLs. Remote gateways return HTTPS URLs that require normal gateway authentication. For gateway service-domain configuration, refer to [Manage Gateways](/how-it-works/gateways/overview#configure-service-forwarding). +Loopback gateways return local `openshell.localhost` URLs. Remote gateways +return HTTPS URLs on the gateway listener. Service routes bypass control-plane +RPC authorization and do not use OIDC or CLI login. Because they share the +listener in this release, its TLS configuration—including any required client +certificate—still applies. The application is responsible for authentication +enabled on its endpoint, and an upstream edge proxy may still apply its own +access policy. For gateway service-domain configuration, refer to +[Manage Gateways](/how-it-works/gateways/overview#configure-service-forwarding). ## Monitor and Debug @@ -617,7 +711,7 @@ Read `last_result` to choose the next check: - `NoObservedExchange`: no result has been reported for the current configuration and supervisor session. Try the operation and inspect its logs if no result appears. - `HttpResponseReceived`: the server returned a final HTTP status below 400, including a protocol upgrade. Informational responses alone do not establish success. Check the client's response for protocol or tool errors; an HTTP 200 response can still contain an error. -- `PolicyDenied`: OpenShell denied the request, including MCP protocol-version or request-body policy checks. Check the sandbox policy and denial logs. +- `PolicyDenied`: OpenShell denied the request, including MCP protocol-version, request-header, request-body, and HTTP-method checks. Check the sandbox policy and denial logs. - `CredentialUnavailable`: required credentials were unavailable. Check the endpoint's attached provider. - `TlsFailed`: TLS setup or the handshake failed. Check certificates and TLS configuration. - `TransportFailed`: the network exchange failed. Check name resolution, connectivity, and the server process. @@ -669,6 +763,8 @@ openshell forward start 8000 my-sandbox -d # run in background ``` OpenShell prints the local URL only after the forward listener is reachable. Background forwards must be tracked locally so `openshell forward list` and `openshell forward stop` can manage them. +CLI forwards run independently of SSH multiplexing settings in your SSH config. +Foreground forwards end when the command exits; only background forwards appear in `forward list`. List and stop active forwards: @@ -850,7 +946,7 @@ remain in force. Timed-out records are retained even for ephemeral creates; use | Ready | The sandbox is running and its supervisor control session is connected. You can connect, execute commands, sync files, and view logs. | | Stopping | The gateway accepted a stop request and is stopping compute while retaining persistent state. | | Stopped | Compute was stopped explicitly and access is unavailable. | -| Starting | Compute is starting. The sandbox becomes usable only after a fresh supervisor session connects. | +| Starting | Compute is starting, or a selected main-process restart is waiting for backoff or a fresh supervisor session. Access resumes after the new session connects. | | Completed | The canonical main process exited with code 0. Its normalized result is available in `status.exit_code`. | | Error | The canonical main process failed, or sandbox infrastructure failed. Inspect the condition reason and `status.exit_code`. | | Deleting | The sandbox is being torn down. The system releases resources and purges credentials. | @@ -865,8 +961,12 @@ temporarily while its supervisor reconnects. Wait for the phase to return to The gateway records a successful canonical main-process exit as `Ready=False` with reason `MainProcessCompleted`. Nonzero and signal-normalized results use `MainProcessFailed` and the `Error` phase. It also sets `status.exit_code`; -signal exits use the standard `128 + signal` convention. Compute runtimes do -not automatically restart that process. +signal exits use the standard `128 + signal` convention. When restart policy +selects replacement, `status.restart_count` and `status.next_restart_time` +show the attempt number and next deadline. Inspect them with +`openshell sandbox get `. `sandbox stop` cancels a pending replacement; +an explicit later start resets the restart count. Compute runtimes do not +apply their own restart policy. ## Sandbox Runtimes diff --git a/docs/how-it-works/sandboxes/runtimes.mdx b/docs/how-it-works/sandboxes/runtimes.mdx index cd3d576b52..f52a571bd2 100644 --- a/docs/how-it-works/sandboxes/runtimes.mdx +++ b/docs/how-it-works/sandboxes/runtimes.mdx @@ -10,6 +10,11 @@ position: 4 Each gateway runs agent workloads on one compute runtime, selected by its compute driver. Pick the runtime that matches the isolation boundary and infrastructure you need. The CLI workflow stays the same across runtimes: you create, connect to, stop, start, and delete sandboxes through the gateway. +The gateway owns main-process restart policy. Docker, Podman, Kubernetes, and +MicroVM native restart mechanisms remain disabled. When policy selects a +replacement, the gateway stops and starts the compute resource while retaining +the sandbox record and driver-supported persistent storage. + | Runtime | Driver | Use when | |---|---|---| | [Docker](#docker-driver) | `docker` | Local development and single-machine gateways. | @@ -91,6 +96,26 @@ Common options in `[openshell.drivers.docker]` are `socket_path`, `grpc_endpoint Docker Desktop must have host networking enabled, and it cannot use Enhanced Container Isolation. Set `grpc_endpoint` when sandboxes cannot reach the gateway on host loopback. For GPU sandboxes, configure Docker CDI before starting the gateway. +### Docker Corporate Proxy Egress + +For proxy-required networks, the Docker driver accepts `https_proxy`, +`no_proxy`, `proxy_auth_file`, `proxy_auth_allow_insecure`, +`proxy_connect_by_hostname`, and `proxy_ca_bundle`. The companion supervisor +chains policy-approved TLS tunnels through the proxy with HTTP CONNECT. + +`proxy_ca_bundle` is an operator-owned gateway-host PEM path. Docker validates +and copies the bundle into its supervisor-only named volume, so this also works +with remote Docker daemons and does not expose the host path to the workload. +The supervisor uses the bundle for an HTTPS proxy connection and for upstream +certificates re-signed by a TLS-intercepting proxy. Workload processes receive +the same corporate roots through their generated combined trust bundle. + +Sandbox environment, image contents, and `template.driver_config.docker` +cannot alter these settings. Invalid proxy relationships or CA files prevent +the gateway or sandbox from starting and never degrade to direct egress. See +the [Gateway Configuration File](/how-it-works/gateways/configuration) for URL, +authentication, `NO_PROXY`, hostname CONNECT, and CA validation details. + ### Docker Mounts Mount existing named volumes or `tmpfs` through driver config. Label volumes so resource admission accepts them: diff --git a/docs/kubernetes/setup.mdx b/docs/kubernetes/setup.mdx index b16065aa61..2b39ddf223 100644 --- a/docs/kubernetes/setup.mdx +++ b/docs/kubernetes/setup.mdx @@ -485,6 +485,8 @@ The gateway exposes `/healthz` for process liveness and `/readyz` for dependency - `startupProbe` and `livenessProbe` use `/healthz`. - `readinessProbe` uses `/readyz`, which reflects the latest result of an in-process background database check. +Each sandbox supervisor Pod uses a `tcpSocket` readiness probe on port 5501. The supervisor accepts connections on that port only while its gateway session is ready. If you add a default-deny ingress policy to sandbox namespaces, allow TCP port 5501 from the node so kubelet can reach supervisor Pods. + ## Next Steps - To run multiple gateway replicas, refer to [High Availability](/kubernetes/high-availability). diff --git a/docs/observability/accessing-logs.mdx b/docs/observability/accessing-logs.mdx index e5981c9583..2f4c7a15ae 100644 --- a/docs/observability/accessing-logs.mdx +++ b/docs/observability/accessing-logs.mdx @@ -58,6 +58,8 @@ On reconnect, a client passes the highest cursor it processed as the resume poin The gateway merges the log and platform event sources before emitting, so events normally arrive in ascending cursor order. That applies to the buffered tail you receive when the stream opens, to a replay after a resume, and to live delivery. The gateway does not delay an event to wait for a lower cursor that has not been published yet, so a cursor can still arrive late under concurrent publication. Track the highest cursor seen as the resume point rather than the last one received. +A cursor is one position shared by the log and platform sources; you track a single highest value, not one per source. If you request both logs and platform events, either source can deliver fewer events on connect than its replay depth asked for. This happens when the other source's replay leaves out an event that is newer than part of this source's history, even when both depths are equal. The gateway withholds those events rather than hand out a cursor that one of the two sources cannot fully vouch for, and reports what it withheld with a warning event rather than dropping it silently. Watches that follow only logs, such as `openshell logs`, are not affected. + ## Direct Filesystem Access Start an independent shell with `sandbox exec` to read log files directly: diff --git a/docs/observability/logging.mdx b/docs/observability/logging.mdx index e9b6a03845..a277992117 100644 --- a/docs/observability/logging.mdx +++ b/docs/observability/logging.mdx @@ -228,7 +228,7 @@ An upstream that the proxy cannot reach returns `502 Bad Gateway`: } ``` -The `error` field is a short machine-readable code (`policy_denied`, `middleware_denied`, `middleware_failed`, `ssrf_denied`, `upstream_unreachable`). The `detail` field is a human-readable explanation suitable for display in an agent transcript. The optional `reason` field, when present, provides the specific denial cause from the policy engine (for example, which binary was not allowed or which rule was missing). +The `error` field is a short machine-readable code (`policy_denied`, `middleware_denied`, `middleware_failed`, `ssrf_denied`, `upstream_unreachable`, `unsupported_l7_protocol`). The `detail` field is a human-readable explanation suitable for display in an agent transcript. The optional `reason` field, when present, provides the specific denial cause from the policy engine (for example, which binary was not allowed or which rule was missing). For L7 REST policy denials, the body also includes structured policy fields such as `method`, `path`, `rule_missing`, and `next_steps`. When the policy advisor is enabled, the body also includes `agent_guidance`, a short plain-language instruction telling the agent to read `/etc/openshell/skills/policy_advisor.md`, propose the narrowest rule through `http://policy.local/v1/proposals`, wait for `policy_reloaded: true`, and retry. A middleware denial instead identifies the policy-local config in `middleware` and can include a validated `reason_code`. A fail-closed runtime failure uses `middleware_failed` with platform-owned text. Both middleware responses omit `rule_missing`, `next_steps`, and `agent_guidance` because no policy rule is missing. diff --git a/docs/observability/ocsf-json-export.mdx b/docs/observability/ocsf-json-export.mdx index 81136e2e31..9a5b970168 100644 --- a/docs/observability/ocsf-json-export.mdx +++ b/docs/observability/ocsf-json-export.mdx @@ -7,9 +7,21 @@ description: "How to enable full OCSF JSON logging for SIEM integration, complia keywords: "Generative AI, Cybersecurity, OCSF, JSON, SIEM, Compliance, Observability" --- -The [shorthand log format](/observability/logging) is optimized for humans and agents reading logs in real time. For machine consumption, compliance archival, or SIEM integration, you can enable full OCSF JSON export. This writes every OCSF event as a complete JSON record in JSONL format, one JSON object per line. +The [shorthand log format](/observability/logging) is optimized for humans and agents reading logs in real time. For machine consumption or SIEM integration, enable OCSF JSON output: one native JSON object per line. Collection is best-effort, not a guarantee of complete audit history. -## Enable JSON Export +## Gateway Output + +Configure `[openshell.gateway.ocsf_log]` in `gateway.toml` to collect gateway-origin OCSF events into one JSONL file. This includes gateway-wide events such as TLS certificate reloads, even when no sandbox log stream applies. Collection preserves native OCSF fields and event IDs and does not depend on `RUST_LOG`. + +The gateway writes OCSF events to a local JSONL file. Give each replica its own path and configure the shipper to follow renamed rotation segments. See [OCSF JSONL configuration](/how-it-works/gateways/configuration#ocsf-jsonl-output) for queue bounds, retention, loss metrics, and best-effort delivery limits. + +This output is separate from the sandbox-local file described below. Enabling one does not enable the other. + +Gateway records use the configured gateway name for both `device.uid` and `device.name`. Replicas of one installation share that identity; `device.hostname` identifies the emitting replica. Assign distinct gateway names to installations whose events are collected together. Reusing a name, including the default name, groups those installations under the same identity. Renaming a gateway changes its audit identity. + +Without JSONL enabled, gateway OCSF events still appear as shorthand in console output, subject to the diagnostic log filter. Events associated with a sandbox also appear in its gateway log stream; gateway-wide events have no sandbox association. + +## Enable Sandbox JSON Export Use the `ocsf_json_enabled` setting to toggle JSON export. The setting can be applied globally, for all sandboxes, or per-sandbox. @@ -64,6 +76,10 @@ sinks rotate daily and retain the three most recent files. ## JSON Record Structure +Gateway-produced records use the `OpenShell Gateway` product identity; supervisor records use `OpenShell Sandbox Supervisor`. The producing device and the affected sandbox are separate identities, so a gateway event can still identify an affected sandbox through its container fields. + +For gateway-produced records, `device.os.name` identifies the OS running the gateway process: Linux, Windows, or macOS. Linux supervisor records continue to report Linux, regardless of the gateway's OS. A native macOS gateway therefore reports macOS while its Linux sandboxes report Linux; a gateway running inside a Linux container reports Linux, even on a Mac host. + `metadata.uid` uniquely identifies an event and stays unchanged when that record is serialized again. Use `container.uid` to associate the event with its sandbox, not `metadata.uid`. Records without a sandbox association omit the container; unknown images are omitted rather than represented by an empty image name. Each line is a complete OCSF v1.8.0 JSON object. Here is an example of a network connection event: @@ -186,6 +202,16 @@ records that previously appeared as `NET:*` now appear as `EVENT`. OpenShell emits OCSF v1.8.0 events internally, but many SIEMs only support older schema versions. The `ocsf_schema_version` setting tells the JSONL layer to downgrade events before writing, stripping fields and profiles that don't exist in the target version. +For gateway JSONL, set `schema_version` on the output destination. The target applies uniformly to every record written to that file and requires a gateway restart: + +```toml +[openshell.gateway.ocsf_log] +path = "/var/log/openshell/gateway-ocsf.jsonl" +schema_version = "1.3" +``` + +The sandbox-scoped setting below controls the legacy supervisor-local JSONL file only. + Set the target version globally: ```shell diff --git a/e2e/mcp-conformance/README.md b/e2e/mcp-conformance/README.md index 4d545281c6..ca95cc3a17 100644 --- a/e2e/mcp-conformance/README.md +++ b/e2e/mcp-conformance/README.md @@ -35,7 +35,11 @@ bridge at `host.openshell.internal` (the alias `e2e/with-docker-gateway.sh` attaches to the CI job container on the e2e network), at `host.docker.internal` on local Docker Desktop, or via `--add-host ...:host-gateway` on local Linux. -The generated policy uses `protocol: mcp`, inserts the conformance runner's spec revision into the endpoint allowlist, and sets `mcp.allow_all_known_mcp_methods: true` so omitted rule methods use the endpoint MCP method profile. OpenShell enforces that allowlist on each non-initialize request using `MCP-Protocol-Version`, with `2025-03-26` as the missing-header fallback. The conformance runner selects the revision used by its client and server; OpenShell's request-version check does not yet provide complete revision-specific message parsing or response validation. The policy keeps OpenShell deny-by-default at the network boundary while allowing the upstream scenarios to exercise MCP behavior. The policy body lives in `policy-template.yaml`; the wrapper renders its MCP revision, host, port, and path placeholders from the upstream server URL. +The generated policy uses `protocol: mcp`, inserts the conformance runner's spec revision into the endpoint allowlist, and sets `mcp.allow_all_known_mcp_methods: true` so omitted rule methods use the selected MCP method profile. The renderer accepts OpenShell's supported revisions, `2025-03-26`, `2025-06-18`, `2025-11-25`, and `2026-07-28`. The policy body lives in `policy-template.yaml`; the wrapper renders its MCP revision, host, port, and path placeholders from the upstream server URL. + +OpenShell checks each request against the policy revision and delegates JSON-RPC structure and MCP method, direction, message-kind, parameter, and metadata type checks to `tower-mcp-types`, pinned to `0.22.2`. These checks follow Tower's deserialization and inspection APIs and do not establish complete JSON-schema conformance. OpenShell owns revision allowlisting, HTTP/body consistency, request limits, and policy enforcement. For the 2025 revisions, a valid standalone `initialize` proposes a version; later requests select their revision through `MCP-Protocol-Version`, with `2025-03-26` as the missing-header fallback. The sessionless `2026-07-28` profile carries one JSON-RPC request or explicitly allowed extension notification per POST. Requests require per-request metadata and matching protocol-version, method, and applicable name headers. Extension notifications require an exact method allow rule and the version header, but no request metadata or method/name mirrors. Responses and SSE payloads are relayed without policy parsing. The conformance runner and its reference client exercise behavior beyond these request inspection checks. + +`OPENSHELL_MCP_CONFORMANCE_SPEC_VERSION` defaults to `2025-11-25`. The default scenarios in `e2e/mcp-conformance.sh` are `initialize`, `tools_call`, and `elicitation-sep1034-client-defaults`, selected for the pinned upstream fixture and this default revision. A passing default run does not establish `2026-07-28` conformance coverage. To exercise that revision through this harness, select an upstream fixture and scenario handlers that implement its sessionless request contract, then set the spec version and scenario list together. For local runs, the wrapper builds `openshell/supervisor:dev` automatically when no supervisor image override is set. Set `SUPERVISOR_IMAGE` to use a @@ -65,14 +69,6 @@ docker run --rm openshell-mcp-conformance-client:local \ ./node_modules/.bin/tsx src/index.ts list --client --spec-version 2025-11-25 ``` -Then confirm each scenario has a compatible handler in the pinned -`examples/clients/typescript/everything-client.ts`. The default list skips -opt-in scenarios, including auth/OAuth flows and the slow `sse-retry` scenario. -Set `OPENSHELL_MCP_CONFORMANCE_SCENARIOS=sse-retry` or pass `sse-retry` as an -argument to run it explicitly. +Then confirm each scenario has a compatible handler in the pinned `examples/clients/typescript/everything-client.ts`. The default list skips opt-in scenarios, including auth/OAuth flows and the slow `sse-retry` scenario. Set `OPENSHELL_MCP_CONFORMANCE_SCENARIOS` to `sse-retry` or pass `sse-retry` as an argument to run it explicitly. -The wrapper caches the pinned upstream checkout, the local conformance runner -build, and the Docker client image. Set -`OPENSHELL_MCP_CONFORMANCE_FORCE_REBUILD=1` to refresh those build artifacts, or -`OPENSHELL_MCP_CONFORMANCE_DOCKER_PULL=1` to pull the client image base during a -rebuild. +The wrapper caches the pinned upstream checkout, the local conformance runner build, and the Docker client image. Set `OPENSHELL_MCP_CONFORMANCE_FORCE_REBUILD` to `1` to refresh those build artifacts, or `OPENSHELL_MCP_CONFORMANCE_DOCKER_PULL` to `1` to pull the client image base during a rebuild. diff --git a/e2e/mcp-conformance/render-policy.py b/e2e/mcp-conformance/render-policy.py index 0d1b9e5c3a..acd8e4a087 100644 --- a/e2e/mcp-conformance/render-policy.py +++ b/e2e/mcp-conformance/render-policy.py @@ -30,7 +30,7 @@ raw_url, policy_file, policy_template, mcp_version = sys.argv[1:5] # Keep this boundary check synchronized with McpProtocolVersion::ALL. The # renderer is standalone Python, so it cannot import the Rust registry. -supported_mcp_versions = {"2025-03-26", "2025-06-18", "2025-11-25"} +supported_mcp_versions = {"2025-03-26", "2025-06-18", "2025-11-25", "2026-07-28"} if mcp_version not in supported_mcp_versions: raise SystemExit(f"unsupported MCP protocol version: {mcp_version!r}") parsed = urlparse(raw_url) diff --git a/e2e/rust/Cargo.lock b/e2e/rust/Cargo.lock index c5235dff8b..018ae54912 100644 --- a/e2e/rust/Cargo.lock +++ b/e2e/rust/Cargo.lock @@ -38,7 +38,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "b281d307588d634de920874890732659e2e7672f72b5e10e81badc1a8a83621e" dependencies = [ "aws-lc-sys", - "untrusted", + "untrusted 0.7.1", "zeroize", ] @@ -278,7 +278,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "39cab71617ae0d63f51a36d69f866391735b51691dbda63cf6f96d042b63efeb" dependencies = [ "libc", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -814,7 +814,7 @@ checksum = "30d65c71f1ce40ab09135ce117d742b9f8a19ff91a41a8b57ed50bc2de59c427" dependencies = [ "libc", "wasi", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -902,6 +902,8 @@ dependencies = [ "noyalib", "prost", "rand", + "rustls", + "rustls-pemfile", "serde", "serde_json", "serial_test", @@ -909,6 +911,7 @@ dependencies = [ "sha2", "tempfile", "tokio", + "tokio-rustls", "tokio-stream", "tonic", "tonic-prost", @@ -1111,6 +1114,20 @@ dependencies = [ "bitflags", ] +[[package]] +name = "ring" +version = "0.17.14" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "a4689e6c2294d81e88dc6261c768b63bc4fcdb852be6d1352498b114f61383b7" +dependencies = [ + "cc", + "cfg-if", + "getrandom 0.2.17", + "libc", + "untrusted 0.9.0", + "windows-sys 0.52.0", +] + [[package]] name = "rustc-hash" version = "2.1.3" @@ -1127,7 +1144,52 @@ dependencies = [ "errno", "libc", "linux-raw-sys", - "windows-sys", + "windows-sys 0.61.2", +] + +[[package]] +name = "rustls" +version = "0.23.45" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "0d41d731c7d2f962d1ccc364cec258de3c0e93b38c2fb3ba97ac74513048d634" +dependencies = [ + "aws-lc-rs", + "log", + "once_cell", + "rustls-pki-types", + "rustls-webpki", + "subtle", + "zeroize", +] + +[[package]] +name = "rustls-pemfile" +version = "2.2.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "dce314e5fee3f39953d46bb63bb8a46d40c2f8fb7cc5a3b6cab2bde9721d6e50" +dependencies = [ + "rustls-pki-types", +] + +[[package]] +name = "rustls-pki-types" +version = "1.15.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "2f4925028c7eb5d1fcdaf196971378ed9d2c1c4efc7dc5d011256f76c99c0a96" +dependencies = [ + "zeroize", +] + +[[package]] +name = "rustls-webpki" +version = "0.103.15" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "f3c3cf1d8b1e7d4927e2d154c3fcb02979afb9939629c62cd9048d4f07b60ac2" +dependencies = [ + "aws-lc-rs", + "ring", + "rustls-pki-types", + "untrusted 0.9.0", ] [[package]] @@ -1317,7 +1379,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "c3d1e2c7f27f8d4cb10542a02c49005dbd6e93095799d6f3be745fae9f8fedd4" dependencies = [ "libc", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -1326,6 +1388,12 @@ version = "1.2.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "6ce2be8dc25455e1f91df71bfa12ad37d7af1092ae736f3a6cd0e37bc7810596" +[[package]] +name = "subtle" +version = "2.6.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292" + [[package]] name = "syn" version = "2.0.119" @@ -1375,7 +1443,7 @@ dependencies = [ "getrandom 0.4.3", "once_cell", "rustix", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -1452,7 +1520,7 @@ dependencies = [ "signal-hook-registry", "socket2", "tokio-macros", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -1466,6 +1534,16 @@ dependencies = [ "syn 3.0.3", ] +[[package]] +name = "tokio-rustls" +version = "0.26.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "c9cc2678c2cdd569ef8215e2afd7954ada2ae20b4fdd2c5fe6139a3b02d105db" +dependencies = [ + "rustls", + "tokio", +] + [[package]] name = "tokio-stream" version = "0.1.19" @@ -1617,6 +1695,12 @@ version = "0.7.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "a156c684c91ea7d62626509bce3cb4e1d9ed5c4d978f7b4352658f96a4c26b4a" +[[package]] +name = "untrusted" +version = "0.9.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1" + [[package]] name = "url" version = "2.5.8" @@ -1738,6 +1822,15 @@ version = "0.2.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5" +[[package]] +name = "windows-sys" +version = "0.52.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "282be5f36a8ce781fad8c8ae18fa3f9beff57ec1b52cb3de0789201425d9a33d" +dependencies = [ + "windows-targets", +] + [[package]] name = "windows-sys" version = "0.61.2" @@ -1747,6 +1840,70 @@ dependencies = [ "windows-link", ] +[[package]] +name = "windows-targets" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "9b724f72796e036ab90c1021d4780d4d3d648aca59e491e6b98e725b84e99973" +dependencies = [ + "windows_aarch64_gnullvm", + "windows_aarch64_msvc", + "windows_i686_gnu", + "windows_i686_gnullvm", + "windows_i686_msvc", + "windows_x86_64_gnu", + "windows_x86_64_gnullvm", + "windows_x86_64_msvc", +] + +[[package]] +name = "windows_aarch64_gnullvm" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "32a4622180e7a0ec044bb555404c800bc9fd9ec262ec147edd5989ccd0c02cd3" + +[[package]] +name = "windows_aarch64_msvc" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "09ec2a7bb152e2252b53fa7803150007879548bc709c039df7627cabbd05d469" + +[[package]] +name = "windows_i686_gnu" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "8e9b5ad5ab802e97eb8e295ac6720e509ee4c243f69d781394014ebfe8bbfa0b" + +[[package]] +name = "windows_i686_gnullvm" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "0eee52d38c090b3caa76c563b86c3a4bd71ef1a819287c19d586d7334ae8ed66" + +[[package]] +name = "windows_i686_msvc" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "240948bc05c5e7c6dabba28bf89d89ffce3e303022809e73deaefe4f6ec56c66" + +[[package]] +name = "windows_x86_64_gnu" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "147a5c80aabfbf0c7d901cb5895d1de30ef2907eb21fbbab29ca94c5b08b1a78" + +[[package]] +name = "windows_x86_64_gnullvm" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "24d5b23dc417412679681396f2b49f3de8c1473deb516bd34410872eff51ed0d" + +[[package]] +name = "windows_x86_64_msvc" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "589f6da84c646204747d1270a2a5661ea66ed1cced2631d546fdfb155959f9ec" + [[package]] name = "wit-bindgen" version = "0.57.1" diff --git a/e2e/rust/Cargo.toml b/e2e/rust/Cargo.toml index b492c8ac86..a2a552400f 100644 --- a/e2e/rust/Cargo.toml +++ b/e2e/rust/Cargo.toml @@ -43,6 +43,11 @@ name = "policy_activation" path = "tests/policy_activation.rs" required-features = ["e2e-docker"] +[[test]] +name = "provider_files" +path = "tests/provider_files.rs" +required-features = ["e2e-docker"] + [[test]] name = "oidc_pkce" path = "tests/oidc_pkce.rs" @@ -58,6 +63,11 @@ name = "custom_image" path = "tests/custom_image.rs" required-features = ["e2e-docker"] +[[test]] +name = "service_bearer_passthrough" +path = "tests/service_bearer_passthrough.rs" +required-features = ["e2e-docker"] + [[test]] name = "rootfs_tar" path = "tests/rootfs_tar.rs" @@ -68,6 +78,11 @@ name = "docker_preflight" path = "tests/docker_preflight.rs" required-features = ["e2e-docker"] +[[test]] +name = "docker_corporate_proxy" +path = "tests/docker_corporate_proxy.rs" +required-features = ["e2e-docker"] + [[test]] name = "driver_config_volume" path = "tests/driver_config_volume.rs" @@ -93,11 +108,6 @@ name = "podman_host_gateway" path = "tests/podman_host_gateway.rs" required-features = ["e2e-podman"] -[[test]] -name = "podman_preflight" -path = "tests/podman_preflight.rs" -required-features = ["e2e-podman"] - [[test]] name = "podman_corporate_proxy" path = "tests/podman_corporate_proxy.rs" @@ -235,11 +245,14 @@ sha1 = "0.10" sha2 = "0.10" hex = "0.4" rand = "0.9" +rustls = { version = "0.23", default-features = false, features = ["std", "logging", "tls12", "aws_lc_rs"] } +rustls-pemfile = "2" serde = { version = "1", features = ["derive"] } serde_json = "1" serde_yml = { package = "noyalib", version = "0.0.28", default-features = false, features = ["std", "compat-serde-yaml"] } tonic = { version = "0.14", features = ["transport"] } tonic-prost = "0.14" +tokio-rustls = { version = "0.26", default-features = false, features = ["logging", "tls12", "aws_lc_rs"] } tower = "0.5" url = "2" nix = { version = "0.29", features = ["process", "signal", "term", "user"] } diff --git a/e2e/rust/src/harness/sandbox.rs b/e2e/rust/src/harness/sandbox.rs index 0d742792e3..9922db02b4 100644 --- a/e2e/rust/src/harness/sandbox.rs +++ b/e2e/rust/src/harness/sandbox.rs @@ -14,7 +14,7 @@ use std::time::Duration; use tokio::io::{AsyncBufReadExt, BufReader}; use tokio::time::timeout; -use super::binary::openshell_cmd; +use super::binary::{openshell_bin, openshell_cmd}; use super::output::{extract_field, strip_ansi}; /// Tool-capable workload image used by the E2E harness. @@ -67,13 +67,18 @@ fn add_test_image_if_missing(command: &mut tokio::process::Command, args: &[&str } } +/// Generate a sandbox name that is unique within and across test processes. +pub fn unique_sandbox_name() -> String { + format!( + "e2e-{}-{}", + std::process::id(), + NEXT_SANDBOX_NAME.fetch_add(1, Ordering::Relaxed) + ) +} + fn add_unique_name_if_missing(command: &mut tokio::process::Command, args: &[&str]) { if !has_explicit_sandbox_name(args) { - command.arg("--name").arg(format!( - "e2e-{}-{}", - std::process::id(), - NEXT_SANDBOX_NAME.fetch_add(1, Ordering::Relaxed) - )); + command.arg("--name").arg(unique_sandbox_name()); } } @@ -723,27 +728,22 @@ impl Drop for SandboxGuard { return; } - // We need to run async cleanup in a sync Drop. Use block_in_place to - // avoid blocking the tokio runtime. This is acceptable for test code. - let name = self.name.clone(); - let mut child = self.child.take(); - - // Attempt cleanup with a new runtime if we're not inside one, or - // block_in_place if we are. - std::thread::spawn(move || { - let rt = tokio::runtime::Runtime::new().expect("create cleanup runtime"); - rt.block_on(async { - if let Some(ref mut child) = child { - let _: Result<(), _> = child.kill().await; - let _ = child.wait().await; - } + // A detached thread here would get killed along with the test + // process before the delete command finishes, leaking the sandbox. + // Use a blocking std::process::Command instead, matching the + // ManagedCleanup pattern in workspace_namespace_managed.rs, so + // cleanup completes before this function returns. + if let Some(mut child) = self.child.take() { + let _ = child.start_kill(); + } - let mut cmd = openshell_cmd(); - cmd.arg("sandbox").arg("delete").arg(&name); - cmd.stdout(Stdio::null()).stderr(Stdio::null()); - let _ = cmd.status().await; - }); - }); + let _ = std::process::Command::new(openshell_bin()) + .arg("sandbox") + .arg("delete") + .arg(&self.name) + .stdout(Stdio::null()) + .stderr(Stdio::null()) + .status(); } } diff --git a/e2e/rust/tests/docker_corporate_proxy.rs b/e2e/rust/tests/docker_corporate_proxy.rs new file mode 100644 index 0000000000..85584b966f --- /dev/null +++ b/e2e/rust/tests/docker_corporate_proxy.rs @@ -0,0 +1,1180 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +#![cfg(feature = "e2e-docker")] + +//! Cross-layer E2E coverage for corporate forward-proxy egress from Docker +//! sandboxes (issue #3545). +//! +//! The Docker counterpart of `podman_corporate_proxy.rs`. It drives the whole +//! chain end to end: +//! +//! gateway TOML → Docker driver config → companion supervisor archive → +//! supervisor CLI parsing → policy evaluation → proxied CONNECT +//! +//! and asserts the properties only a real run can establish: +//! +//! 1. A policy-approved HTTPS request reaches its destination *through* the +//! proxy, with a validated IP as the CONNECT target. +//! 2. A policy-denied destination never reaches the proxy at all. +//! 3. Credentials arrive through the supervisor-volume file — the proxy answers +//! 407 to an unauthenticated CONNECT, so a 200 proves delivery. +//! 4. A port-qualified `no_proxy` entry bypasses the proxy for that port only. +//! 5. An `https://` proxy works when its CA is supplied via `proxy_ca_bundle`. +//! 6. A TLS-intercepting proxy CA is trusted by both inspected supervisor +//! upstream TLS and a raw workload TLS connection. +//! 7. An incoherent setting is fatal at gateway startup rather than +//! degrading to a direct dial. +//! +//! Fixtures run as host processes and are reached by the host-networked +//! supervisor. `host.openshell.internal` is normalized to the gateway host +//! address, including the CI job-container address when Docker runs on the +//! host daemon. + +use std::fmt::Write as _; +use std::io::Write as _; +use std::path::PathBuf; +use std::time::Duration; + +use openshell_e2e::harness::cli::wait_for_healthy; +use openshell_e2e::harness::gateway::ManagedGateway; +use openshell_e2e::harness::host_process::HostPythonFixture; +use openshell_e2e::harness::port::find_free_port; +use openshell_e2e::harness::sandbox::SandboxGuard; +use serial_test::serial; +use tempfile::NamedTempFile; + +/// The OpenShell alias for the gateway host. +const HOST_ALIAS: &str = "host.openshell.internal"; +const PROXY_USER: &str = "proxyuser"; +const PROXY_PASS: &str = "proxypass"; + +const ALLOWED_MARKER: &str = "docker-corp-proxy-e2e-allowed-upstream"; +const DENIED_MARKER: &str = "docker-corp-proxy-e2e-denied-upstream"; +const BYPASS_MARKER: &str = "docker-corp-proxy-e2e-bypass-upstream"; +const READY_MARKER: &str = "docker-corp-proxy-e2e-workload-done"; + +/// Ports the fixtures bind and the workload addresses them by. +struct FixturePorts { + proxy: u16, + allowed: u16, + denied: u16, + bypass: u16, +} + +impl FixturePorts { + fn pick() -> Self { + Self { + proxy: find_free_port(), + allowed: find_free_port(), + denied: find_free_port(), + bypass: find_free_port(), + } + } +} + +/// Python helper retained by the shared proxy fixture shape. +const HOST_REWRITE: &str = " +def dial_host(host): + return host +"; + +/// A forward proxy that requires Basic auth and logs every CONNECT it sees. +/// +/// The log lines are the test's evidence: they record the exact CONNECT target +/// (proving validated-IP form, and which destination ports were proxied at +/// all) and whether credentials arrived. +fn proxy_script(port: u16) -> String { + format!( + r#" +import base64, select, socket, threading +{HOST_REWRITE} +EXPECTED = 'Basic ' + base64.b64encode(b'{PROXY_USER}:{PROXY_PASS}').decode() + +def log(msg): + print(msg, flush=True) + +def read_head(conn): + data = b'' + while b'\r\n\r\n' not in data: + chunk = conn.recv(4096) + if not chunk: + return None + data += chunk + if len(data) > 65536: + return None + return data + +def pipe(a, b): + try: + while True: + ready, _, _ = select.select([a, b], [], []) + for sock in ready: + chunk = sock.recv(65536) + if not chunk: + return + (b if sock is a else a).sendall(chunk) + except OSError: + return + +def handle(conn): + try: + head = read_head(conn) + if head is None: + # Readiness probes connect and close without sending a request. + return + lines = head.decode('latin-1').split('\r\n') + parts = lines[0].split() + if len(parts) < 2 or parts[0].upper() != 'CONNECT': + log('NON_CONNECT %s' % lines[0]) + conn.sendall(b'HTTP/1.1 405 Method Not Allowed\r\nContent-Length: 0\r\n\r\n') + return + target = parts[1] + auth = None + for line in lines[1:]: + if line.lower().startswith('proxy-authorization:'): + auth = line.split(':', 1)[1].strip() + if auth != EXPECTED: + log('CONNECT %s auth=fail' % target) + conn.sendall(b'HTTP/1.1 407 Proxy Authentication Required\r\n' + b'Proxy-Authenticate: Basic realm="corp"\r\n' + b'Content-Length: 0\r\n\r\n') + return + host, _, port = target.rpartition(':') + host = host.strip('[]') + try: + upstream = socket.create_connection((dial_host(host), int(port)), timeout=10) + except OSError: + log('CONNECT %s auth=ok dial=fail' % target) + conn.sendall(b'HTTP/1.1 502 Bad Gateway\r\nContent-Length: 0\r\n\r\n') + return + log('CONNECT %s auth=ok' % target) + conn.sendall(b'HTTP/1.1 200 Connection Established\r\n\r\n') + pipe(conn, upstream) + upstream.close() + except OSError: + pass + finally: + conn.close() + +server = socket.socket(socket.AF_INET, socket.SOCK_STREAM) +server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) +server.bind(('0.0.0.0', {port})) +server.listen(64) +log('proxy-listening') +while True: + client, _ = server.accept() + threading.Thread(target=handle, args=(client,), daemon=True).start() +"# + ) +} + +// Delimits the proxy's CA certificate in its output, so the test can recover it +// and hand it back as the corporate CA bundle. +const CA_BEGIN: &str = "---PROXY-CA-BEGIN---"; +const CA_END: &str = "---PROXY-CA-END---"; + +/// A forward proxy that terminates TLS with a CA-signed certificate and logs +/// every CONNECT it sees. +/// +/// It mints a corporate CA and a listener leaf signed by it (SAN = the host +/// alias, which is the name the supervisor uses for SNI), serves the leaf, and +/// prints the CA between [`CA_BEGIN`]/[`CA_END`] so the test can trust it via +/// `proxy_ca_bundle`. The listener certificate must be a leaf: rustls rejects a +/// `CA:TRUE` certificate presented as an end-entity certificate. No Basic auth +/// here — this test isolates the `https://` proxy plus corporate-CA path. +/// Mint a corporate CA plus a listener leaf signed by it, and print the CA. +/// +/// Split out of [`tls_proxy_script`] to keep that function within clippy's +/// line budget. The listener certificate must be a leaf: rustls rejects a +/// `CA:TRUE` certificate presented as an end-entity certificate, so a single +/// self-signed `openssl req -x509` certificate cannot serve as both the +/// anchor and the listener identity. Signing a leaf also matches what a real +/// intercepting proxy does. +fn tls_proxy_pki_preamble() -> String { + format!( + r" +workdir = tempfile.mkdtemp() +ca_key = os.path.join(workdir, 'ca-key.pem') +ca_crt = os.path.join(workdir, 'ca.pem') +leaf_key = os.path.join(workdir, 'leaf-key.pem') +leaf_csr = os.path.join(workdir, 'leaf.csr') +leaf_crt = os.path.join(workdir, 'leaf.pem') +chain = os.path.join(workdir, 'chain.pem') +ext = os.path.join(workdir, 'leaf.ext') + +def run(*args): + subprocess.run(args, check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) + +run('openssl', 'req', '-x509', '-newkey', 'rsa:2048', '-nodes', + '-keyout', ca_key, '-out', ca_crt, '-days', '1', + '-subj', '/CN=docker-corp-proxy-e2e-ca', + '-addext', 'basicConstraints=critical,CA:TRUE', + '-addext', 'keyUsage=critical,keyCertSign,cRLSign') + +with open(ext, 'w') as fh: + fh.write('basicConstraints=critical,CA:FALSE\n' + 'subjectAltName=DNS:{HOST_ALIAS}\n' + 'extendedKeyUsage=serverAuth\n' + 'keyUsage=critical,digitalSignature,keyEncipherment\n') +run('openssl', 'req', '-newkey', 'rsa:2048', '-nodes', + '-keyout', leaf_key, '-out', leaf_csr, '-subj', '/CN={HOST_ALIAS}') +run('openssl', 'x509', '-req', '-in', leaf_csr, '-CA', ca_crt, '-CAkey', ca_key, + '-CAcreateserial', '-out', leaf_crt, '-days', '1', '-extfile', ext) + +with open(chain, 'w') as out: + for part in (leaf_crt, ca_crt): + with open(part) as fh: + out.write(fh.read()) + +with open(ca_crt) as fh: + print('{CA_BEGIN}\n' + fh.read() + '{CA_END}', flush=True) +" + ) +} + +fn tls_proxy_script(port: u16) -> String { + let pki = tls_proxy_pki_preamble(); + format!( + r" +import os, select, socket, ssl, subprocess, tempfile, threading +{HOST_REWRITE} +{pki} + +def log(msg): + print(msg, flush=True) + +def read_head(conn): + data = b'' + while b'\r\n\r\n' not in data: + chunk = conn.recv(4096) + if not chunk: + return None + data += chunk + if len(data) > 65536: + return None + return data + +def pipe(a, b): + try: + while True: + ready, _, _ = select.select([a, b], [], []) + for sock in ready: + chunk = sock.recv(65536) + if not chunk: + return + (b if sock is a else a).sendall(chunk) + except OSError: + return + +def handle(conn): + try: + head = read_head(conn) + if head is None: + return + lines = head.decode('latin-1').split('\r\n') + parts = lines[0].split() + if len(parts) < 2 or parts[0].upper() != 'CONNECT': + log('NON_CONNECT %s' % lines[0]) + conn.sendall(b'HTTP/1.1 405 Method Not Allowed\r\nContent-Length: 0\r\n\r\n') + return + target = parts[1] + host, _, port = target.rpartition(':') + host = host.strip('[]') + try: + upstream = socket.create_connection((dial_host(host), int(port)), timeout=10) + except OSError: + log('CONNECT %s dial=fail' % target) + conn.sendall(b'HTTP/1.1 502 Bad Gateway\r\nContent-Length: 0\r\n\r\n') + return + log('CONNECT %s ok' % target) + conn.sendall(b'HTTP/1.1 200 Connection Established\r\n\r\n') + pipe(conn, upstream) + upstream.close() + except OSError: + pass + finally: + conn.close() + +ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER) +ctx.load_cert_chain(chain, leaf_key) +server = socket.socket(socket.AF_INET, socket.SOCK_STREAM) +server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) +server.bind(('0.0.0.0', {port})) +server.listen(64) +log('tls-proxy-listening') +while True: + raw, _ = server.accept() + try: + conn = ctx.wrap_socket(raw, server_side=True) + except OSError: + # Readiness probes open a bare TCP connection; the failed handshake is + # expected and must not spam the log. + raw.close() + continue + threading.Thread(target=handle, args=(conn,), daemon=True).start() +" + ) +} + +/// A plain HTTP CONNECT proxy that intercepts tunneled destination TLS. +/// +/// The proxy presents a corporate-CA-signed certificate for +/// `host.openshell.internal`, then opens an unverified TLS connection to the +/// real fixture. This lets one test prove both consumers of the configured CA: +/// inspected traffic is verified by the supervisor, while `tls: skip` traffic +/// is verified directly by the workload through its combined trust bundle. +fn intercepting_proxy_script(port: u16) -> String { + let pki = tls_proxy_pki_preamble(); + format!( + r" +import select, socket, ssl, subprocess, tempfile, threading, os +{HOST_REWRITE} +{pki} + +def log(msg): + print(msg, flush=True) + +def read_head(conn): + data = b'' + while b'\r\n\r\n' not in data: + chunk = conn.recv(4096) + if not chunk: + return None + data += chunk + if len(data) > 65536: + return None + return data + +def relay_http(client, upstream): + # select() cannot safely drive SSLSocket application data: TLS control + # records can make the fd readable while recv() still waits for payload. + # This fixture serves one bodyless HTTP request per intercepted tunnel, so + # relay that exchange directly and let the upstream HTTP/1.0 close delimit + # the response. + request = read_head(client) + if request is None: + return + upstream.sendall(request) + while True: + chunk = upstream.recv(65536) + if not chunk: + return + client.sendall(chunk) + +server_context = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER) +server_context.load_cert_chain(chain, leaf_key) +upstream_context = ssl._create_unverified_context() + +def handle(conn): + target = '' + try: + head = read_head(conn) + if head is None: + return + lines = head.decode('latin-1').split('\r\n') + parts = lines[0].split() + if len(parts) < 2 or parts[0].upper() != 'CONNECT': + conn.sendall(b'HTTP/1.1 405 Method Not Allowed\r\nContent-Length: 0\r\n\r\n') + return + target = parts[1] + host, _, target_port = target.rpartition(':') + host = host.strip('[]') + try: + upstream_raw = socket.create_connection( + (dial_host(host), int(target_port)), timeout=10) + upstream = upstream_context.wrap_socket( + upstream_raw, server_hostname='{HOST_ALIAS}') + except OSError as err: + log('CONNECT %s mitm=dial-fail error=%r' % (target, err)) + conn.sendall(b'HTTP/1.1 502 Bad Gateway\r\nContent-Length: 0\r\n\r\n') + return + conn.sendall(b'HTTP/1.1 200 Connection Established\r\n\r\n') + try: + client = server_context.wrap_socket(conn, server_side=True) + except OSError as err: + log('CONNECT %s mitm=client-tls-fail error=%r' % (target, err)) + return + log('CONNECT %s mitm=ok' % target) + relay_http(client, upstream) + upstream.close() + client.close() + except OSError as err: + log('CONNECT %s mitm=error error=%r' % (target, err)) + finally: + try: + conn.close() + except OSError: + pass + +server = socket.socket(socket.AF_INET, socket.SOCK_STREAM) +server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) +server.bind(('0.0.0.0', {port})) +server.listen(64) +log('intercepting-proxy-listening') +while True: + conn, _ = server.accept() + threading.Thread(target=handle, args=(conn,), daemon=True).start() +" + ) +} + +/// Extract the proxy's CA certificate PEM from its output. +fn ca_cert_from_logs(logs: &str) -> Result { + let start = logs + .find(CA_BEGIN) + .ok_or_else(|| format!("proxy CA begin marker not found in logs:\n{logs}"))? + + CA_BEGIN.len(); + let end = logs[start..] + .find(CA_END) + .ok_or_else(|| format!("proxy CA end marker not found in logs:\n{logs}"))?; + Ok(logs[start..start + end].trim().to_string()) +} + +/// A TLS server with a self-signed certificate, serving one identifying marker. +/// +/// The workload uses an unverified TLS context, so the certificate only needs +/// to exist — but the handshake itself must be real, because it is what proves +/// bytes flowed end to end through the tunnel. +fn tls_upstream_script(marker: &str, port: u16) -> String { + format!( + r#" +import http.server, os, ssl, subprocess, tempfile + +workdir = tempfile.mkdtemp() +key = os.path.join(workdir, 'key.pem') +crt = os.path.join(workdir, 'cert.pem') +subprocess.run( + ['openssl', 'req', '-x509', '-newkey', 'rsa:2048', '-nodes', + '-keyout', key, '-out', crt, '-days', '1', '-subj', '/CN={HOST_ALIAS}'], + check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) + +class Handler(http.server.BaseHTTPRequestHandler): + def do_GET(self): + body = b'{{"upstream":"{marker}"}}' + self.send_response(200) + self.send_header('Content-Type', 'application/json') + self.send_header('Content-Length', str(len(body))) + self.end_headers() + self.wfile.write(body) + + def log_message(self, fmt, *args): + pass + +class Server(http.server.HTTPServer): + # Readiness probes open a bare TCP connection and close it; the failed + # handshake is expected and must not spam the log. + def handle_error(self, request, client_address): + pass + +ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER) +ctx.load_cert_chain(crt, key) +server = Server(('0.0.0.0', {port}), Handler) +server.socket = ctx.wrap_socket(server.socket, server_side=True) +print('tls-upstream-listening', flush=True) +server.serve_forever() +"# + ) +} + +/// Workload: one approved HTTPS request through the proxy, one policy-denied, +/// and one that `no_proxy` should send direct. +/// +fn workload_script(ports: &FixturePorts) -> String { + let (allowed, denied, bypass) = (ports.allowed, ports.denied, ports.bypass); + format!( + r" +import json, ssl, time, urllib.request + +ctx = ssl._create_unverified_context() + +def fetch(url, retries): + last = {{'status': -1, 'error': 'not attempted'}} + for attempt in range(retries): + try: + with urllib.request.urlopen(url, timeout=30, context=ctx) as resp: + return {{'status': resp.status, 'body': resp.read().decode()}} + except Exception as err: + last = {{'status': -1, 'error': str(err)}} + time.sleep(1) + return last + +# The approved requests are retried: policy reload during sandbox startup can +# transiently surface as a 403 in the forward proxy. +print('ALLOWED_RESULT ' + json.dumps( + fetch('https://{HOST_ALIAS}:{allowed}/', 6)), flush=True) +print('BYPASS_RESULT ' + json.dumps( + fetch('https://{HOST_ALIAS}:{bypass}/', 6)), flush=True) +# The denied request must fail, so a single attempt is enough. +print('DENIED_RESULT ' + json.dumps( + fetch('https://{HOST_ALIAS}:{denied}/', 1)), flush=True) +print('{READY_MARKER}', flush=True) +" + ) +} + +/// Workload variant that uses the system/default trust store. +/// +/// The inspected endpoint proves the supervisor accepts the corporate +/// re-signing CA for upstream TLS. The `tls: skip` endpoint proves the same CA +/// reached the workload trust bundle because Python verifies it directly. +fn trusted_workload_script(ports: &FixturePorts) -> String { + let (inspected, raw, denied) = (ports.allowed, ports.bypass, ports.denied); + format!( + r" +import json, ssl, time, urllib.request + +ctx = ssl.create_default_context() + +def fetch(url, retries): + last = {{'status': -1, 'error': 'not attempted'}} + for attempt in range(retries): + try: + with urllib.request.urlopen(url, timeout=30, context=ctx) as resp: + return {{'status': resp.status, 'body': resp.read().decode()}} + except Exception as err: + last = {{'status': -1, 'error': str(err)}} + time.sleep(1) + return last + +print('INSPECTED_RESULT ' + json.dumps( + fetch('https://{HOST_ALIAS}:{inspected}/', 6)), flush=True) +print('RAW_RESULT ' + json.dumps( + fetch('https://{HOST_ALIAS}:{raw}/', 6)), flush=True) +print('DENIED_RESULT ' + json.dumps( + fetch('https://{HOST_ALIAS}:{denied}/', 1)), flush=True) +print('{READY_MARKER}', flush=True) +" + ) +} + +/// Policy allowing the proxied and the bypassed upstream, but not the denied +/// one. `tls: skip` keeps the tunnel raw so the workload's TLS session runs end +/// to end, which is what makes the proxied CONNECT path observable. +fn policy_yaml(ports: &FixturePorts) -> String { + let (allowed, bypass) = (ports.allowed, ports.bypass); + format!( + r#"version: 1 + +filesystem_policy: + include_workdir: true + read_only: + - /usr + - /lib + - /proc + - /dev/urandom + - /app + - /etc + - /var/log + read_write: + - /sandbox + - /tmp + - /dev/null + +landlock: + compatibility: best_effort + +process: + run_as_user: sandbox + run_as_group: sandbox + +network_policies: + docker_corporate_proxy_e2e: + name: docker_corporate_proxy_e2e + endpoints: + - host: {HOST_ALIAS} + port: {allowed} + tls: skip + enforcement: enforce + allowed_ips: + - "10.0.0.0/8" + - "172.0.0.0/8" + - "192.168.0.0/16" + - "fc00::/7" + - host: {HOST_ALIAS} + port: {bypass} + tls: skip + enforcement: enforce + allowed_ips: + - "10.0.0.0/8" + - "172.0.0.0/8" + - "192.168.0.0/16" + - "fc00::/7" + binaries: + - path: /usr/bin/curl + - path: /usr/bin/python* + - path: /usr/local/bin/python* + - path: /sandbox/.venv/bin/python* + - path: /sandbox/.uv/python/*/bin/python* +"# + ) +} + +/// Allow one auto-inspected TLS endpoint and one raw `tls: skip` endpoint. +fn interception_policy_yaml(ports: &FixturePorts) -> String { + let (inspected, raw) = (ports.allowed, ports.bypass); + format!( + r#"version: 1 + +filesystem_policy: + include_workdir: true + read_only: [/usr, /lib, /proc, /dev/urandom, /app, /etc, /var/log] + read_write: [/sandbox, /tmp, /dev/null] + +landlock: + compatibility: best_effort + +process: + run_as_user: sandbox + run_as_group: sandbox + +network_policies: + docker_corporate_proxy_interception_e2e: + name: docker_corporate_proxy_interception_e2e + endpoints: + - host: {HOST_ALIAS} + port: {inspected} + enforcement: enforce + allowed_ips: ["10.0.0.0/8", "172.0.0.0/8", "192.168.0.0/16", "fc00::/7"] + - host: {HOST_ALIAS} + port: {raw} + tls: skip + enforcement: enforce + allowed_ips: ["10.0.0.0/8", "172.0.0.0/8", "192.168.0.0/16", "fc00::/7"] + binaries: + - path: /usr/bin/python* + - path: /usr/local/bin/python* + - path: /sandbox/.venv/bin/python* + - path: /sandbox/.uv/python/*/bin/python* +"# + ) +} + +/// Inserts corporate-proxy keys into the harness-generated Docker table and +/// restores the original file when dropped. +struct GatewayProxyConfig { + config_path: PathBuf, + original: Vec, + restored: bool, +} + +fn insert_docker_driver_settings(config: &str, extra: &str) -> Result { + let table_header = "[openshell.drivers.docker]"; + let table_start = config + .find(table_header) + .ok_or_else(|| format!("gateway config has no {table_header} table"))?; + let table_body = table_start + table_header.len(); + let insert_at = config[table_body..] + .find("\n[") + .map_or(config.len(), |offset| table_body + offset); + + let mut updated = String::with_capacity(config.len() + extra.len() + 2); + updated.push_str(&config[..insert_at]); + if !updated.ends_with('\n') { + updated.push('\n'); + } + updated.push_str(extra); + if !extra.ends_with('\n') { + updated.push('\n'); + } + updated.push_str(&config[insert_at..]); + Ok(updated) +} + +#[test] +fn proxy_settings_are_inserted_before_nested_driver_tables() { + let config = r#"[openshell.drivers.docker] +socket_path = "/var/run/docker.sock" + +[openshell.drivers.docker.resource_admission] +enabled = true +"#; + + let updated = insert_docker_driver_settings(config, "https_proxy = \"http://proxy\"\n") + .expect("insert proxy settings"); + let proxy_position = updated + .find("https_proxy = \"http://proxy\"") + .expect("proxy setting is present"); + let nested_table_position = updated + .find("[openshell.drivers.docker.resource_admission]") + .expect("nested table is preserved"); + + assert!(proxy_position < nested_table_position); + assert!(updated.contains("enabled = true")); +} + +impl GatewayProxyConfig { + /// Locate the gateway's `--config` path from the wrapper's args file. + fn config_path_from_args() -> Result { + let args_file = std::env::var("OPENSHELL_E2E_GATEWAY_ARGS_FILE") + .map_err(|_| "OPENSHELL_E2E_GATEWAY_ARGS_FILE must be set".to_string())?; + let raw = std::fs::read(&args_file) + .map_err(|err| format!("read gateway args file '{args_file}': {err}"))?; + let args: Vec = raw + .split(|byte| *byte == 0) + .filter(|arg| !arg.is_empty()) + .map(|arg| String::from_utf8_lossy(arg).into_owned()) + .collect(); + args.iter() + .position(|arg| arg == "--config") + .and_then(|index| args.get(index + 1)) + .map(PathBuf::from) + .ok_or_else(|| format!("no --config argument in gateway args file '{args_file}'")) + } + + /// Insert raw TOML lines into `[openshell.drivers.docker]` and restart the + /// gateway, without waiting for it to become healthy. + /// + /// Used directly by the fail-closed case, which expects the gateway *not* + /// to come up. + fn apply_raw(extra: &str) -> Result { + let config_path = Self::config_path_from_args()?; + let original = std::fs::read(&config_path) + .map_err(|err| format!("read gateway config '{}': {err}", config_path.display()))?; + + let config = std::str::from_utf8(&original).map_err(|err| { + format!( + "gateway config '{}' is not UTF-8: {err}", + config_path.display() + ) + })?; + let updated = insert_docker_driver_settings(config, extra) + .map_err(|err| format!("gateway config '{}': {err}", config_path.display()))?; + std::fs::write(&config_path, updated.as_bytes()) + .map_err(|err| format!("write gateway config '{}': {err}", config_path.display()))?; + + let guard = Self { + config_path, + original, + restored: false, + }; + let gateway = ManagedGateway::from_env()? + .ok_or_else(|| "managed gateway metadata disappeared".to_string())?; + gateway.stop()?; + gateway.start()?; + Ok(guard) + } + + /// Point the Docker driver at a corporate proxy and wait for the gateway to + /// come back healthy. + async fn apply( + proxy_url: &str, + auth_file: Option<&str>, + ca_bundle: Option<&str>, + no_proxy: Option<&str>, + ) -> Result { + let mut extra = format!("https_proxy = \"{proxy_url}\"\n"); + if let Some(auth_file) = auth_file { + let _ = write!( + extra, + "proxy_auth_file = \"{auth_file}\"\nproxy_auth_allow_insecure = true\n" + ); + } + if let Some(ca_bundle) = ca_bundle { + let _ = writeln!(extra, "proxy_ca_bundle = \"{ca_bundle}\""); + } + if let Some(no_proxy) = no_proxy { + let _ = writeln!(extra, "no_proxy = \"{no_proxy}\""); + } + let guard = Self::apply_raw(&extra)?; + wait_for_healthy(Duration::from_secs(120)).await?; + Ok(guard) + } + + /// Restore the original config and restart the gateway. + async fn restore(&mut self) -> Result<(), String> { + if self.restored { + return Ok(()); + } + std::fs::write(&self.config_path, &self.original).map_err(|err| { + format!( + "restore gateway config '{}': {err}", + self.config_path.display() + ) + })?; + let gateway = ManagedGateway::from_env()? + .ok_or_else(|| "managed gateway metadata disappeared".to_string())?; + gateway.stop()?; + gateway.start()?; + wait_for_healthy(Duration::from_secs(120)).await?; + self.restored = true; + Ok(()) + } +} + +impl Drop for GatewayProxyConfig { + fn drop(&mut self) { + if self.restored { + return; + } + // Panic path: put the original config back and synchronously restart + // the gateway so later test binaries in this run inherit neither the + // proxy settings on disk nor the configuration still loaded in the + // running process. + let _ = std::fs::write(&self.config_path, &self.original); + if let Ok(Some(gateway)) = ManagedGateway::from_env() { + let _ = gateway.stop(); + let _ = gateway.start(); + } + } +} + +/// Skip unless this run owns a Docker gateway it can reconfigure. +/// +/// The external-driver lane launches `openshell-driver-docker` from the shell +/// wrapper rather than from the gateway, so gateway config changes never reach +/// the driver and the test would assert against a driver that has no proxy +/// settings at all. +fn should_run(label: &str) -> bool { + if std::env::var("OPENSHELL_E2E_DRIVER").as_deref() != Ok("docker") { + eprintln!("Skipping {label}: e2e driver is not docker"); + return false; + } + if std::env::var("OPENSHELL_E2E_EXTERNAL_COMPUTE_DRIVER").as_deref() == Ok("1") { + eprintln!("Skipping {label}: the external Docker driver is not configured by the gateway"); + return false; + } + match ManagedGateway::from_env() { + Ok(Some(_)) => true, + Ok(None) => { + eprintln!("Skipping {label}: e2e gateway is not managed by this test run"); + false + } + Err(err) => panic!("load managed e2e gateway metadata: {err}"), + } +} + +/// Write a temp file and return its path as an owned `String`. +fn temp_file_with(contents: &str, label: &str) -> (NamedTempFile, String) { + let mut file = NamedTempFile::new().unwrap_or_else(|err| panic!("create {label}: {err}")); + file.write_all(contents.as_bytes()) + .unwrap_or_else(|err| panic!("write {label}: {err}")); + file.flush() + .unwrap_or_else(|err| panic!("flush {label}: {err}")); + let path = file + .path() + .to_str() + .unwrap_or_else(|| panic!("{label} path should be utf-8")) + .to_string(); + (file, path) +} + +/// Assert the workload's results and the proxy's own record of what it saw. +fn assert_proxied_egress(output: &str, proxy_logs: &str, ports: &FixturePorts) { + assert!( + output.contains(READY_MARKER), + "workload did not finish; output:\n{output}" + ); + + // The approved request succeeded and its body came from the upstream, + // proving bytes traversed the tunnel rather than the proxy short-circuiting. + assert!( + output.contains("ALLOWED_RESULT") && output.contains(ALLOWED_MARKER), + "approved HTTPS request should have reached the upstream through the proxy:\n{output}" + ); + + // The CONNECT target is a validated resolved address, not the hostname, + // and it arrived authenticated -- which is only possible if the credential + // staged into the overlay reached the supervisor. + assert!( + proxy_logs.lines().any(|line| { + line.starts_with("CONNECT ") + && line.contains(&format!(":{} auth=ok", ports.allowed)) + && !line.contains(HOST_ALIAS) + }), + "proxy should have seen an authenticated validated-IP CONNECT to the approved upstream:\n{proxy_logs}" + ); + assert!( + !proxy_logs.contains("auth=fail"), + "the proxy credential should have been sent on the first CONNECT:\n{proxy_logs}" + ); + + // The denied destination failed inside the sandbox, and never reached the + // proxy: policy stops it before the upstream dial. + assert!( + !proxy_logs.contains(&format!(":{}", ports.denied)), + "policy-denied destination must never reach the proxy:\n{proxy_logs}" + ); + assert!( + !output.contains(DENIED_MARKER), + "policy-denied upstream body must never reach the workload:\n{output}" + ); + + // The no_proxy entry is port-qualified, so this destination was dialed + // directly while the approved one above still went through the proxy. + assert!( + output.contains("BYPASS_RESULT") && output.contains(BYPASS_MARKER), + "no_proxy destination should have been reached by a direct dial:\n{output}" + ); + assert!( + !proxy_logs.contains(&format!(":{}", ports.bypass)), + "no_proxy destination must never reach the proxy:\n{proxy_logs}" + ); +} + +#[tokio::test] +#[serial(docker_corporate_proxy)] +async fn docker_corporate_proxy_routes_approved_tls_egress() { + if !should_run("corporate proxy test") { + return; + } + + let ports = FixturePorts::pick(); + + // ── Host fixtures, reached by the host supervisor ── + let proxy = HostPythonFixture::start(&proxy_script(ports.proxy), ports.proxy) + .await + .expect("start fake corporate proxy"); + let _allowed = HostPythonFixture::start( + &tls_upstream_script(ALLOWED_MARKER, ports.allowed), + ports.allowed, + ) + .await + .expect("start approved TLS upstream"); + let _denied = HostPythonFixture::start( + &tls_upstream_script(DENIED_MARKER, ports.denied), + ports.denied, + ) + .await + .expect("start denied TLS upstream"); + let _bypass = HostPythonFixture::start( + &tls_upstream_script(BYPASS_MARKER, ports.bypass), + ports.bypass, + ) + .await + .expect("start no_proxy TLS upstream"); + + // ── Point the Docker driver at the corporate proxy ──────────────────── + let (_auth_file, auth_path) = + temp_file_with(&format!("{PROXY_USER}:{PROXY_PASS}\n"), "proxy auth file"); + + let mut gateway_config = GatewayProxyConfig::apply( + &format!("http://{HOST_ALIAS}:{}", ports.proxy), + Some(&auth_path), + None, + // Port-qualified, so only this destination port bypasses the proxy. + Some(&format!("{HOST_ALIAS}:{}", ports.bypass)), + ) + .await + .expect("apply corporate proxy gateway config"); + + // ── Run the workload ────────────────────────────────────────────── + let (_policy, policy_path) = temp_file_with(&policy_yaml(&ports), "policy file"); + let script = workload_script(&ports); + let mut sandbox = + SandboxGuard::create(&["--policy", &policy_path, "--", "python3", "-c", &script]) + .await + .expect("create Docker sandbox behind the corporate proxy"); + + assert_proxied_egress( + &sandbox.create_output, + &proxy.logs().expect("read fake proxy logs"), + &ports, + ); + + sandbox.cleanup().await; + + gateway_config + .restore() + .await + .expect("restore gateway config"); +} + +#[tokio::test] +#[serial(docker_corporate_proxy)] +async fn docker_corporate_proxy_trusts_ca_bundle_for_https_proxy() { + if !should_run("https corporate proxy test") { + return; + } + + let ports = FixturePorts::pick(); + + let proxy = HostPythonFixture::start(&tls_proxy_script(ports.proxy), ports.proxy) + .await + .expect("start fake https corporate proxy"); + let _allowed = HostPythonFixture::start( + &tls_upstream_script(ALLOWED_MARKER, ports.allowed), + ports.allowed, + ) + .await + .expect("start approved TLS upstream"); + let _denied = HostPythonFixture::start( + &tls_upstream_script(DENIED_MARKER, ports.denied), + ports.denied, + ) + .await + .expect("start denied TLS upstream"); + let _bypass = HostPythonFixture::start( + &tls_upstream_script(BYPASS_MARKER, ports.bypass), + ports.bypass, + ) + .await + .expect("start second approved TLS upstream"); + + // The proxy prints its CA on startup; the supervisor must trust it to + // complete the TLS handshake with the proxy at all. + let ca_pem = ca_cert_from_logs(&proxy.logs().expect("read https proxy logs")) + .expect("recover corporate CA from proxy output"); + let (_ca_file, ca_path) = temp_file_with(&ca_pem, "corporate CA bundle"); + + let mut gateway_config = GatewayProxyConfig::apply( + &format!("https://{HOST_ALIAS}:{}", ports.proxy), + None, + Some(&ca_path), + None, + ) + .await + .expect("apply https corporate proxy gateway config"); + + let (_policy, policy_path) = temp_file_with(&policy_yaml(&ports), "policy file"); + let script = workload_script(&ports); + let mut sandbox = + SandboxGuard::create(&["--policy", &policy_path, "--", "python3", "-c", &script]) + .await + .expect("create Docker sandbox behind the https corporate proxy"); + + let proxy_logs = proxy.logs().expect("read https proxy logs"); + assert!( + sandbox.create_output.contains(ALLOWED_MARKER), + "approved upstream body missing -- egress did not complete through the https proxy:\n{}", + sandbox.create_output + ); + assert!( + proxy_logs.lines().any(|line| { + line.starts_with("CONNECT ") + && line.contains(&format!(":{} ok", ports.allowed)) + && !line.contains(HOST_ALIAS) + }), + "https proxy should have seen a validated-IP CONNECT to the approved upstream:\n{proxy_logs}" + ); + assert!( + !proxy_logs.contains(&format!(":{}", ports.denied)), + "policy-denied destination must never reach the https proxy:\n{proxy_logs}" + ); + + sandbox.cleanup().await; + + gateway_config + .restore() + .await + .expect("restore gateway config"); +} + +#[tokio::test] +#[serial(docker_corporate_proxy)] +async fn docker_corporate_proxy_trusts_intercepted_destination_tls() { + if !should_run("TLS-intercepting corporate proxy test") { + return; + } + + let ports = FixturePorts::pick(); + let proxy = HostPythonFixture::start(&intercepting_proxy_script(ports.proxy), ports.proxy) + .await + .expect("start TLS-intercepting corporate proxy"); + let _inspected = HostPythonFixture::start( + &tls_upstream_script(ALLOWED_MARKER, ports.allowed), + ports.allowed, + ) + .await + .expect("start inspected TLS upstream"); + let _denied = HostPythonFixture::start( + &tls_upstream_script(DENIED_MARKER, ports.denied), + ports.denied, + ) + .await + .expect("start denied TLS upstream"); + let _raw = HostPythonFixture::start( + &tls_upstream_script(BYPASS_MARKER, ports.bypass), + ports.bypass, + ) + .await + .expect("start raw workload TLS upstream"); + + let ca_pem = ca_cert_from_logs(&proxy.logs().expect("read intercepting proxy logs")) + .expect("recover intercepting proxy CA"); + let (_ca_file, ca_path) = temp_file_with(&ca_pem, "intercepting proxy CA bundle"); + let mut gateway_config = GatewayProxyConfig::apply( + &format!("http://{HOST_ALIAS}:{}", ports.proxy), + None, + Some(&ca_path), + None, + ) + .await + .expect("apply TLS-intercepting proxy configuration"); + + let (_policy, policy_path) = temp_file_with( + &interception_policy_yaml(&ports), + "interception policy file", + ); + let script = trusted_workload_script(&ports); + let mut sandbox = + SandboxGuard::create(&["--policy", &policy_path, "--", "python3", "-c", &script]) + .await + .expect("create Docker sandbox behind TLS-intercepting proxy"); + + let output = &sandbox.create_output; + let proxy_logs = proxy.logs().expect("read TLS-intercepting proxy logs"); + assert!( + output.contains("INSPECTED_RESULT") && output.contains(ALLOWED_MARKER), + "supervisor-inspected TLS should trust the corporate re-signing CA:\n\ + workload output:\n{output}\nproxy output:\n{proxy_logs}" + ); + assert!( + output.contains("RAW_RESULT") && output.contains(BYPASS_MARKER), + "raw workload TLS should trust the corporate CA through the combined bundle:\n{output}" + ); + assert!( + proxy_logs.contains(&format!(":{} mitm=ok", ports.allowed)), + "intercepting proxy should see inspected upstream TLS:\n{proxy_logs}" + ); + assert!( + proxy_logs.contains(&format!(":{} mitm=ok", ports.bypass)), + "intercepting proxy should see raw workload TLS:\n{proxy_logs}" + ); + assert!( + !proxy_logs.contains(&format!(":{}", ports.denied)) && !output.contains(DENIED_MARKER), + "policy-denied destination must not reach the intercepting proxy:\n{proxy_logs}\n{output}" + ); + + sandbox.cleanup().await; + gateway_config + .restore() + .await + .expect("restore gateway config"); +} + +#[tokio::test] +#[serial(docker_corporate_proxy)] +async fn docker_corporate_proxy_rejects_incoherent_configuration() { + if !should_run("corporate proxy fail-closed test") { + return; + } + + // A bypass list without a proxy URL means the operator believed proxying + // was in effect. Accepting it would leave every dial direct while looking + // configured, so the gateway must refuse to start instead. + let mut gateway_config = GatewayProxyConfig::apply_raw("no_proxy = \"10.0.0.0/8\"\n") + .expect("apply incoherent corporate proxy gateway config"); + + let health = wait_for_healthy(Duration::from_secs(20)).await; + assert!( + health.is_err(), + "gateway must not serve traffic with an incoherent [openshell.drivers.docker] proxy table" + ); + + let log_path = std::env::var("OPENSHELL_E2E_GATEWAY_LOG").expect("gateway log path"); + let log = std::fs::read_to_string(&log_path).expect("read gateway log"); + // The message must name the offending key rather than surfacing as an + // opaque driver-readiness timeout. Matched loosely on the key names so + // this does not break on error-wrapping changes. + assert!( + log.contains("no_proxy") && log.contains("https_proxy"), + "the startup error must name the offending key; gateway log tail:\n{}", + log.lines().rev().take(40).collect::>().join("\n") + ); + + gateway_config + .restore() + .await + .expect("restore gateway config"); +} diff --git a/e2e/rust/tests/mcp_sessionless.rs b/e2e/rust/tests/mcp_sessionless.rs new file mode 100644 index 0000000000..bf2261efa1 --- /dev/null +++ b/e2e/rust/tests/mcp_sessionless.rs @@ -0,0 +1,381 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +//! MCP request profiles through the sandbox's transparent network interception. +//! +//! A shared fixture serves legacy initialization, sessionless discovery, named +//! tools, and a bounded subscription stream. Upstream receipts distinguish a +//! policy denial from a tool request accepted by the fixture. + +#![cfg(feature = "e2e-host-gateway")] + +use std::io::Write; + +use openshell_e2e::harness::container::ContainerHttpServer; +use openshell_e2e::harness::sandbox::SandboxGuard; +use tempfile::NamedTempFile; + +const SERVER_ALIAS: &str = "mcp-sessionless.openshell.test"; + +const SERVER_SCRIPT: &str = r#" +import json +from http.server import BaseHTTPRequestHandler, HTTPServer + +SESSIONLESS_VERSION = "2026-07-28" +received = [] + +class Handler(BaseHTTPRequestHandler): + def reply(self, status, payload, content_type="application/json"): + self.send_response(status) + self.send_header("Content-Type", content_type) + self.send_header("Content-Length", str(len(payload))) + self.end_headers() + self.wfile.write(payload) + + def do_GET(self): + # The container fixture's readiness probe uses the root URL. + self.reply(200 if self.path == "/" else 405, b"") + + def read_body(self): + if self.headers.get("Transfer-Encoding", "").lower() != "chunked": + return self.rfile.read(int(self.headers.get("Content-Length", "0"))) + body = bytearray() + while True: + size = int(self.rfile.readline().split(b";", 1)[0].strip(), 16) + if size == 0: + while self.rfile.readline().strip(): + pass + return bytes(body) + body.extend(self.rfile.read(size)) + self.rfile.read(2) + + def do_POST(self): + message = json.loads(self.read_body()) + method = message["method"] + params = message.get("params", {}) + version = (params.get("protocolVersion") if method == "initialize" + else self.headers.get("MCP-Protocol-Version")) + # Record parsed requests before fixture admission so rejected revisions + # and tool attempts stay visible. HTTPServer handles requests serially. + received.append([version, method, params.get("name")]) + if self.path != "/mcp" or version not in SUPPORTED_VERSIONS: + self.reply(400, b"request revision did not reach the fixture intact") + return + if version == SESSIONLESS_VERSION: + meta = params.get("_meta", {}) + if (self.headers.get("MCP-Protocol-Version") != version + or self.headers.get("Mcp-Method") != method + or meta.get("io.modelcontextprotocol/protocolVersion") != version + or meta.get("io.modelcontextprotocol/clientCapabilities") != {} + or (method == "tools/call" + and self.headers.get("Mcp-Name") != params["name"])): + self.reply(400, b"request metadata did not reach the fixture intact") + return + + if method == "initialize" and version != SESSIONLESS_VERSION: + result = { + "protocolVersion": version, + "capabilities": {"tools": {}}, + "serverInfo": {"name": "openshell-profile-fixture", "version": "1"}, + } + elif method == "notifications/initialized" and version != SESSIONLESS_VERSION: + self.reply(202, b"") + return + elif method == "server/discover" and version == SESSIONLESS_VERSION: + result = { + "supportedVersions": [version], + "capabilities": {"tools": {"listChanged": True}}, + "ttlMs": 0, + "cacheScope": "private", + "_meta": { + "io.modelcontextprotocol/serverInfo": { + "name": "openshell-sessionless-fixture", "version": "1" + } + }, + } + elif method == "tools/call": + # Both tool names work upstream; OpenShell owns the policy denial. + result = { + "content": [{"type": "text", "text": params["name"]}], + "isError": False, + "_meta": {"fixtureRequests": list(received)}, + } + if version == SESSIONLESS_VERSION: + result["resultType"] = "complete" + elif method == "subscriptions/listen" and version == SESSIONLESS_VERSION: + subscription_meta = {"io.modelcontextprotocol/subscriptionId": message["id"]} + events = [ + { + "jsonrpc": "2.0", + "method": "notifications/subscriptions/acknowledged", + "params": {"notifications": params["notifications"], "_meta": subscription_meta}, + }, + { + "jsonrpc": "2.0", + "method": "notifications/tools/list_changed", + "params": {"_meta": subscription_meta}, + }, + { + "jsonrpc": "2.0", + "id": message["id"], + "result": {"resultType": "complete", "_meta": subscription_meta}, + }, + ] + body = "".join("event: message\ndata: " + json.dumps(event) + "\n\n" for event in events) + self.reply(200, body.encode(), "text/event-stream") + return + else: + self.reply(400, b"unexpected method for the selected fixture revision") + return + + self.reply(200, json.dumps({"jsonrpc": "2.0", "id": message["id"], "result": result}).encode()) + + def log_message(self, format, *args): + pass + +HTTPServer(("0.0.0.0", 8000), Handler).serve_forever() +"#; + +const CLIENT_HELPERS: &str = r#" +import json +import urllib.error +import urllib.request +# Direct client connections pass through the sandbox's transparent interception. +opener = urllib.request.build_opener(urllib.request.ProxyHandler({})) + +def post(request_id, method, params, version): + params = dict(params) + headers = { + "Content-Type": "application/json", + "Accept": "application/json, text/event-stream", + } + if method != "initialize": + headers["MCP-Protocol-Version"] = version + if version == "2026-07-28": + params["_meta"] = { + "io.modelcontextprotocol/protocolVersion": version, + "io.modelcontextprotocol/clientCapabilities": {}, + "io.modelcontextprotocol/clientInfo": {"name": "openshell-e2e", "version": "1"}, + } + headers["Mcp-Method"] = method + if method == "tools/call": + headers["Mcp-Name"] = params["name"] + message = {"jsonrpc": "2.0", "method": method, "params": params} + if request_id is not None: + message["id"] = request_id + request = urllib.request.Request( + f"http://{HOST}:{PORT}/mcp", + data=json.dumps(message).encode(), + headers=headers, + method="POST", + ) + try: + with opener.open(request, timeout=15) as response: + return response.status, response.headers.get_content_type(), response.read() + except urllib.error.HTTPError as error: + return error.code, error.headers.get_content_type(), error.read() +"#; + +const CLIENT_SCRIPT: &str = r#" +VERSION = "2026-07-28" + +status, content_type, body = post(1, "server/discover", {}, VERSION) +assert status == 200, ("discovery", status, body) +assert content_type == "application/json", content_type +discovery = json.loads(body) +assert discovery["id"] == 1, discovery +assert discovery["result"]["supportedVersions"] == [VERSION], discovery + +status, _, body = post(2, "tools/call", {"name": "read_status", "arguments": {}}, VERSION) +assert status == 200, ("allowed tool", status, body) +tool = json.loads(body) +assert tool["id"] == 2, tool +assert tool["result"]["content"] == [{"type": "text", "text": "read_status"}], tool + +status, _, body = post(3, "tools/call", {"name": "read_details", "arguments": {}}, VERSION) +assert status == 403, ("denied tool", status, body) + +status, content_type, body = post(4, "subscriptions/listen", {"notifications": {"toolsListChanged": True}}, VERSION) +assert status == 200, ("subscription", status, body) +assert content_type == "text/event-stream", (content_type, body) +events = [json.loads(line[6:]) for line in body.decode().splitlines() if line.startswith("data: ")] +assert len(events) == 3, events +assert events[0]["method"] == "notifications/subscriptions/acknowledged", events +assert events[0]["params"]["notifications"] == {"toolsListChanged": True}, events +assert events[1]["method"] == "notifications/tools/list_changed", events +assert events[1]["params"]["_meta"]["io.modelcontextprotocol/subscriptionId"] == 4, events +assert events[2]["id"] == 4 and events[2]["result"]["resultType"] == "complete", events + +print("MCP_SESSIONLESS_OK discovery=200 allowed_tool=200 denied_tool=403 subscription=200") +"#; + +const PROFILE_CLIENT_SCRIPT: &str = r#" +expected_receipts = [] +for version in SELECTED_VERSIONS: + if version == "2026-07-28": + status, _, body = post(1, "server/discover", {}, version) + assert status == 200, (version, "discovery", status, body) + assert json.loads(body)["result"]["supportedVersions"] == [version], body + expected_receipts.append([version, "server/discover", None]) + else: + status, _, body = post(1, "initialize", { + "protocolVersion": version, + "capabilities": {}, + "clientInfo": {"name": "openshell-e2e", "version": "1"}, + }, version) + assert status == 200, (version, "initialize", status, body) + assert json.loads(body)["result"]["protocolVersion"] == version, body + expected_receipts.append([version, "initialize", None]) + status, _, body = post(None, "notifications/initialized", {}, version) + assert status == 202, (version, "initialized", status, body) + expected_receipts.append([version, "notifications/initialized", None]) + + status, _, body = post(2, "tools/call", {"name": "read_status", "arguments": {}}, version) + assert status == 200, (version, "allowed tool", status, body) + tool = json.loads(body) + assert tool["id"] == 2, tool + assert tool["result"]["content"] == [{"type": "text", "text": "read_status"}], tool + expected_receipts.append([version, "tools/call", "read_status"]) + assert tool["result"]["_meta"]["fixtureRequests"] == expected_receipts, tool + + status, _, body = post(3, "tools/call", {"name": "read_details", "arguments": {}}, version) + assert status == 403, (version, "denied tool", status, body) + + # The fixture accepts both tools. A later allowed call proves the denial + # came from the proxy and no denied operation reached the upstream. + status, _, body = post(4, "tools/call", {"name": "read_status", "arguments": {}}, version) + assert status == 200, (version, "receipt tool", status, body) + receipt = json.loads(body) + assert receipt["id"] == 4, receipt + expected_receipts.append([version, "tools/call", "read_status"]) + assert receipt["result"]["_meta"]["fixtureRequests"] == expected_receipts, receipt + print(f"MCP_PROFILE_OK version={version} allowed_tool=200 denied_tool=403 receipts=verified") +"#; + +async fn start_server(alias: &str, versions: &[&str]) -> Result { + let versions = serde_json::to_string(versions).map_err(|err| err.to_string())?; + let script = format!("SUPPORTED_VERSIONS = {versions}\n{SERVER_SCRIPT}"); + ContainerHttpServer::start_python(alias, &script).await +} + +fn write_policy(host: &str, port: u16, versions: &[&str]) -> Result { + let mut file = NamedTempFile::new().map_err(|err| format!("create temp policy: {err}"))?; + let legacy_rules = if versions.iter().any(|version| *version != "2026-07-28") { + " - allow:\n method: initialize\n - allow:\n method: notifications/initialized\n" + } else { + "" + }; + let versions = serde_json::to_string(versions).map_err(|err| err.to_string())?; + let policy = format!( + r#"version: 1 +filesystem_policy: + include_workdir: true + read_only: [/usr, /lib, /proc, /dev/urandom, /app, /etc, /var/log] + read_write: [/sandbox, /tmp, /dev/null] +landlock: + compatibility: best_effort +process: + run_as_user: sandbox + run_as_group: sandbox +network_policies: + mcp_sessionless: + name: mcp_sessionless + endpoints: + - host: {host} + port: {port} + path: /mcp + protocol: mcp + enforcement: enforce + allowed_ips: ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "fc00::/7"] + mcp: + versions: {versions} + max_body_bytes: 65536 + rules: +{legacy_rules} - allow: + method: server/discover + - allow: + method: tools/call + tool: read_status + - allow: + method: subscriptions/listen + deny_rules: + - method: tools/call + tool: read_details + binaries: + - path: /usr/bin/python* + - path: /usr/local/bin/python* + - path: /sandbox/.uv/python/*/bin/python* +"# + ); + file.write_all(policy.as_bytes()) + .map_err(|err| format!("write temp policy: {err}"))?; + file.flush() + .map_err(|err| format!("flush temp policy: {err}"))?; + Ok(file) +} + +async fn run_client( + server: &ContainerHttpServer, + versions: &[&str], + client: &str, +) -> Result { + let policy = write_policy(&server.host, server.port, versions)?; + let policy_path = policy + .path() + .to_str() + .ok_or("temp policy path is not UTF-8")?; + let selected_versions = serde_json::to_string(versions).map_err(|err| err.to_string())?; + let script = format!( + "HOST = {:?}\nPORT = {}\nSELECTED_VERSIONS = {selected_versions}\n{CLIENT_HELPERS}\n{client}", + server.host, server.port + ); + SandboxGuard::create(&["--policy", policy_path, "--", "python3", "-c", &script]).await +} + +#[tokio::test] +async fn sessionless_discovery_tools_and_subscription_use_request_metadata() { + let versions = ["2026-07-28"]; + let server = start_server(SERVER_ALIAS, &versions) + .await + .expect("start sessionless MCP fixture"); + let sandbox = run_client(&server, &versions, CLIENT_SCRIPT) + .await + .expect("run sessionless MCP client in sandbox"); + + assert!( + sandbox.create_output.contains( + "MCP_SESSIONLESS_OK discovery=200 allowed_tool=200 denied_tool=403 subscription=200" + ), + "expected completed sessionless MCP assertions, got:\n{}", + sandbox.create_output + ); +} + +#[tokio::test] +async fn legacy_and_multi_version_profiles_authorize_tools_through_sandbox() { + for versions in [ + &["2025-03-26"][..], + &["2025-06-18"][..], + &["2025-11-25", "2026-07-28"][..], + ] { + // Each scenario starts with fresh upstream receipts. Its distinct alias + // avoids the sessionless test's fixture, and cleanup precedes alias reuse. + let server = start_server("mcp-profiles.openshell.test", versions) + .await + .unwrap_or_else(|err| panic!("{versions:?}: start MCP fixture: {err}")); + let mut sandbox = run_client(&server, versions, PROFILE_CLIENT_SCRIPT) + .await + .unwrap_or_else(|err| panic!("{versions:?}: run MCP sandbox client: {err}")); + for version in versions { + let marker = format!( + "MCP_PROFILE_OK version={version} allowed_tool=200 denied_tool=403 receipts=verified" + ); + assert!( + sandbox.create_output.contains(&marker), + "{versions:?}: expected completed {version} assertions, got:\n{}", + sandbox.create_output + ); + } + sandbox.cleanup().await; + } +} diff --git a/e2e/rust/tests/provider_files.rs b/e2e/rust/tests/provider_files.rs new file mode 100644 index 0000000000..c2779fe097 --- /dev/null +++ b/e2e/rust/tests/provider_files.rs @@ -0,0 +1,159 @@ +// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +#![cfg(feature = "e2e-docker")] + +//! A provider profile serves read-only config on each workload open. + +use std::io::Write as _; +use std::process::Stdio; +use std::time::Duration; + +use openshell_e2e::harness::binary::openshell_bin; +use openshell_e2e::harness::cli::run_cli; +use openshell_e2e::harness::sandbox::SandboxGuard; +use tokio::time::sleep; + +struct ProviderGuard { + profile: String, + provider: String, +} + +impl Drop for ProviderGuard { + fn drop(&mut self) { + let binary = openshell_bin(); + for _ in 0..20 { + let deleted = std::process::Command::new(&binary) + .args(["provider", "delete", &self.provider]) + .stdout(Stdio::null()) + .stderr(Stdio::null()) + .status(); + if deleted.is_ok_and(|status| status.success()) { + break; + } + std::thread::sleep(Duration::from_millis(250)); + } + let _ = std::process::Command::new(&binary) + .args(["profile", "delete", &self.profile]) + .stdout(Stdio::null()) + .stderr(Stdio::null()) + .status(); + } +} + +async fn cli_ok(args: &[&str]) -> Result<(), String> { + let (output, code) = run_cli(args).await; + if code == 0 { + Ok(()) + } else { + Err(format!( + "{} failed (exit {code}):\n{output}", + args.join(" ") + )) + } +} + +#[tokio::test] +async fn provider_file_open_update_and_detach() -> Result<(), String> { + let suffix = format!("{}-{:08x}", std::process::id(), rand::random::()); + let profile_id = format!("pf-{suffix}"); + let provider = format!("pf-{suffix}"); + let profile_yaml = include_str!("../../../examples/provider-managed-files/acme-config.yaml") + .replace("id: acme-config", &format!("id: {profile_id}")); + let mut profile_file = tempfile::Builder::new() + .suffix(".yaml") + .tempfile() + .map_err(|error| error.to_string())?; + profile_file + .write_all(profile_yaml.as_bytes()) + .map_err(|error| error.to_string())?; + cli_ok(&[ + "profile", + "import", + "--file", + profile_file + .path() + .to_str() + .ok_or("profile path is not UTF-8")?, + ]) + .await?; + let _provider_guard = ProviderGuard { + profile: profile_id.clone(), + provider: provider.clone(), + }; + cli_ok(&[ + "provider", + "create", + "--name", + &provider, + "--type", + &profile_id, + "--config", + "endpoint=https://api.acme.example", + "--config", + "project=production", + ]) + .await?; + + let mut sandbox = SandboxGuard::create(&["--provider", &provider, "--no-tty"]).await?; + let path = format!("/run/openshell/providers/{provider}/client.toml"); + let environment_path = sandbox.exec(&["printenv", "ACME_CONFIG_FILE"]).await?; + if !environment_path.contains(&path) { + return Err(format!( + "provider path environment variable missing: {environment_path}" + )); + } + let before = sandbox.exec(&["cat", &path]).await?; + if !before.contains("project = \"production\"") { + return Err(format!("initial provider file content missing:\n{before}")); + } + + cli_ok(&[ + "provider", + "update", + &provider, + "--config", + "project=staging", + "--wait", + "--timeout", + "90", + ]) + .await?; + let after = sandbox.exec(&["cat", &path]).await?; + if !after.contains("project = \"staging\"") { + return Err(format!("updated provider file content missing:\n{after}")); + } + + cli_ok(&[ + "sandbox", + "provider", + "detach", + &sandbox.name, + &provider, + "--wait", + "--timeout", + "90", + ]) + .await?; + let (detached, code) = run_cli(&[ + "sandbox", + "exec", + "--name", + &sandbox.name, + "--no-tty", + "--", + "cat", + &path, + ]) + .await; + if code == 0 || !detached.contains("No such file or directory") { + return Err(format!( + "detached provider path remained readable: {detached}" + )); + } + sandbox.cleanup().await; + // Let the asynchronous sandbox deletion release the provider before + // ProviderGuard removes the provider record and its profile. + sleep(Duration::from_millis(250)).await; + Ok(()) +} diff --git a/e2e/rust/tests/provider_readiness.rs b/e2e/rust/tests/provider_readiness.rs index 700c4880da..8ec6398a2b 100644 --- a/e2e/rust/tests/provider_readiness.rs +++ b/e2e/rust/tests/provider_readiness.rs @@ -477,9 +477,12 @@ impl Backend { async fn spawn(&mut self, base: &str, tls_directory: &Path) -> Result { let tls_directory = tls_directory .to_str() - .filter(|path| !path.contains([',', '\n', '\r'])) + .filter(|path| !path.contains([':', '\n', '\r'])) .ok_or("fixture TLS mount path is invalid")?; - let mount = format!("type=bind,src={tls_directory},dst=/fixture-tls,readonly"); + // Docker's structured `--mount` syntax cannot request SELinux + // relabeling. Both fixture backends mount this ephemeral directory, so + // use the shared `z` label rather than the single-container `Z` label. + let mount = format!("{tls_directory}:/fixture-tls:ro,z"); let namespace_label = format!("openshell.ai/sandbox-namespace={}", self.namespace); let mut command = Command::from(self.engine.command()); command @@ -501,7 +504,7 @@ impl Backend { "--read-only", "--cap-drop=ALL", "--security-opt=no-new-privileges:true", - "--mount", + "--volume", &mount, "--entrypoint", "/usr/bin/python3", diff --git a/e2e/rust/tests/sandbox_labels.rs b/e2e/rust/tests/sandbox_labels.rs index 90351266c5..937171ead3 100644 --- a/e2e/rust/tests/sandbox_labels.rs +++ b/e2e/rust/tests/sandbox_labels.rs @@ -5,6 +5,7 @@ use std::process::Stdio; use openshell_e2e::harness::binary::openshell_cmd; use openshell_e2e::harness::output::{extract_field, strip_ansi}; +use openshell_e2e::harness::sandbox::SandboxGuard; fn normalize_output(output: &str) -> String { let stripped = strip_ansi(output).replace('\r', ""); @@ -104,14 +105,6 @@ async fn get_sandbox_details(name: &str) -> String { combined } -async fn delete_sandbox(name: &str) { - let mut cmd = openshell_cmd(); - cmd.args(["sandbox", "delete", name]) - .stdout(Stdio::null()) - .stderr(Stdio::null()); - let _ = cmd.status().await; -} - #[tokio::test] #[allow(clippy::too_many_lines)] // end-to-end test exercises full label lifecycle async fn sandbox_labels_are_stored_and_filterable() { @@ -123,6 +116,12 @@ async fn sandbox_labels_are_stored_and_filterable() { let prod_frontend = format!("lbl-pf-{suffix}"); let dev_data = format!("lbl-dd-{suffix}"); + // Arm guards before create so a partial create still cleans up. + let _cleanup: Vec<_> = [&dev_backend, &staging_backend, &prod_frontend, &dev_data] + .into_iter() + .map(|name| SandboxGuard::manage_existing(name.clone())) + .collect(); + // Create sandboxes with different labels let name1 = create_sandbox_with_labels(&dev_backend, &[("env", "dev"), ("team", "backend")]).await; @@ -228,10 +227,4 @@ async fn sandbox_labels_are_stored_and_filterable() { all_sandboxes.contains(&name4), "list without filter should include all test sandboxes" ); - - // Cleanup - delete_sandbox(&name1).await; - delete_sandbox(&name2).await; - delete_sandbox(&name3).await; - delete_sandbox(&name4).await; } diff --git a/e2e/rust/tests/sandbox_lifecycle.rs b/e2e/rust/tests/sandbox_lifecycle.rs index 4df9f24eba..a7c306fec9 100644 --- a/e2e/rust/tests/sandbox_lifecycle.rs +++ b/e2e/rust/tests/sandbox_lifecycle.rs @@ -11,13 +11,14 @@ use std::time::Duration; use openshell_e2e::harness::binary::{openshell_cmd, openshell_tty_cmd}; use openshell_e2e::harness::cli::{run_cli, wait_for_sandbox_phase}; use openshell_e2e::harness::output::{extract_field, strip_ansi}; -use openshell_e2e::harness::sandbox::SandboxGuard; +use openshell_e2e::harness::sandbox::{SandboxGuard, unique_sandbox_name}; use serial_test::serial; use tokio::io::{AsyncBufReadExt, AsyncWriteExt, BufReader}; use tokio::time::{Instant, sleep}; const SANDBOX_PRESENCE_TIMEOUT: Duration = Duration::from_secs(30); const SANDBOX_LIST_POLL_INTERVAL: Duration = Duration::from_millis(500); +const SANDBOX_RESTART_TIMEOUT: Duration = Duration::from_secs(60); fn normalize_output(output: &str) -> String { let stripped = strip_ansi(output).replace('\r', ""); @@ -644,7 +645,19 @@ async fn sandbox_can_be_deleted_while_stopped() { #[tokio::test] #[serial(sandbox_lifecycle)] async fn canonical_main_exit_zero_completes_persistent_sandbox() { - let mut cmd = openshell_tty_cmd(&["sandbox", "create", "--", "echo", "OK"]); + // Armed before create so a failed create or parse still cleans up. + let sandbox_name = unique_sandbox_name(); + let _cleanup = SandboxGuard::manage_existing(sandbox_name.clone()); + + let mut cmd = openshell_tty_cmd(&[ + "sandbox", + "create", + "--name", + &sandbox_name, + "--", + "echo", + "OK", + ]); cmd.stdout(Stdio::piped()).stderr(Stdio::piped()); let output = cmd.output().await.expect("spawn openshell sandbox create"); @@ -657,11 +670,8 @@ async fn canonical_main_exit_zero_completes_persistent_sandbox() { combined.contains("OK"), "main output was not streamed:\n{combined}" ); - let sandbox_name = - extract_sandbox_name(&combined).expect("sandbox name should be present in output"); if let Err(last_sandbox_list) = assert_sandbox_presence_eventually(&sandbox_name, true).await { - delete_sandbox(&sandbox_name).await; panic!( "sandbox {sandbox_name} should still exist by default after {SANDBOX_PRESENCE_TIMEOUT:?}; \ last observed sandbox list: {last_sandbox_list:?}" @@ -687,16 +697,20 @@ async fn canonical_main_exit_zero_completes_persistent_sandbox() { details.contains("Phase: Completed"), "expected terminal sandbox phase:\n{details}" ); - - delete_sandbox(&sandbox_name).await; } #[tokio::test] #[serial(sandbox_lifecycle)] async fn canonical_main_nonzero_exit_preserves_status() { + // Armed before create so a failed create or parse still cleans up. + let sandbox_name = unique_sandbox_name(); + let _cleanup = SandboxGuard::manage_existing(sandbox_name.clone()); + let mut cmd = openshell_tty_cmd(&[ "sandbox", "create", + "--name", + &sandbox_name, "--", "sh", "-c", @@ -719,8 +733,6 @@ async fn canonical_main_nonzero_exit_preserves_status() { combined.contains("failed-main"), "main output was not streamed:\n{combined}" ); - let sandbox_name = - extract_sandbox_name(&combined).expect("sandbox name should be present in output"); let mut get_cmd = openshell_cmd(); get_cmd @@ -741,7 +753,6 @@ async fn canonical_main_nonzero_exit_preserves_status() { details.contains("Exit Code: 7"), "missing exit code:\n{details}" ); - delete_sandbox(&sandbox_name).await; } #[tokio::test] @@ -1266,6 +1277,71 @@ async fn canonical_main_exit_255_is_not_retried_as_transport_failure() { #[tokio::test] #[serial(sandbox_lifecycle)] +async fn on_failure_policy_replaces_runtime_and_preserves_workspace() { + const FIRST_MARKER: &str = "initial-main-ready"; + const SCRIPT: &str = r#" +marker=/sandbox/.openshell-restart-e2e +if [ -e "$marker" ]; then + printf 'replacement-%s\n' "$(cat /proc/sys/kernel/random/uuid)" > /sandbox/replacement-run + printf 'replacement-main-ready\n' + sleep 300 +else + touch "$marker" + printf 'initial-%s\n' "$(cat /proc/sys/kernel/random/uuid)" > /sandbox/initial-run + printf 'initial-main-ready\n' + sleep 2 + exit 17 +fi +"#; + + let mut sandbox = SandboxGuard::create_keep_with_args( + &["--restart-policy", "on-failure"], + &["sh", "-lc", SCRIPT], + FIRST_MARKER, + ) + .await + .expect("create sandbox with OnFailure restart policy"); + + let deadline = Instant::now() + SANDBOX_RESTART_TIMEOUT; + let mut observed_replacement_starting = false; + let final_details = loop { + let details = sandbox_details(&sandbox.name).await; + observed_replacement_starting |= details.contains("Phase: Starting"); + if observed_replacement_starting + && details.contains("Phase: Ready") + && details.contains("Restart count: 1") + { + break details; + } + assert!( + Instant::now() < deadline, + "sandbox did not complete its policy-driven restart within \ + {SANDBOX_RESTART_TIMEOUT:?}; last details:\n{details}" + ); + sleep(Duration::from_millis(250)).await; + }; + + assert!( + final_details.contains("Restart policy: on-failure"), + "restart policy should remain visible after replacement:\n{final_details}" + ); + let runs = sandbox + .exec(&["cat", "/sandbox/initial-run", "/sandbox/replacement-run"]) + .await + .expect("replacement sandbox should retain the first run's workspace"); + assert!( + runs.contains("initial-"), + "missing initial run marker:\n{runs}" + ); + assert!( + runs.contains("replacement-"), + "missing replacement run marker:\n{runs}" + ); + + sandbox.cleanup().await; +} + +#[tokio::test] async fn sandbox_create_with_no_keep_cleans_up_after_tty_command() { let name = format!("tty-{:015x}", rand::random::() & 0x0fff_ffff_ffff_ffff); // Capture startup diagnostics before --no-keep removes a failed container. diff --git a/e2e/rust/tests/service_bearer_passthrough.rs b/e2e/rust/tests/service_bearer_passthrough.rs new file mode 100644 index 0000000000..af2fddb76d --- /dev/null +++ b/e2e/rust/tests/service_bearer_passthrough.rs @@ -0,0 +1,344 @@ +// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +// SPDX-License-Identifier: Apache-2.0 + +#![cfg(feature = "e2e-docker")] + +//! Verifies application bearer authorization behavior through an exposed +//! `OpenShell` service. + +use std::fs::File; +use std::io::BufReader; +use std::net::Ipv4Addr; +use std::path::{Path, PathBuf}; +use std::process::Stdio; +use std::sync::Arc; +use std::time::Duration; + +use bytes::Bytes; +use http_body_util::{BodyExt as _, Empty}; +use hyper::client::conn::http1; +use hyper::{Request, StatusCode, header}; +use hyper_util::rt::TokioIo; +use openshell_e2e::harness::binary::openshell_cmd; +use openshell_e2e::harness::sandbox::{E2E_WORKLOAD_IMAGE, SandboxGuard}; +use rustls::pki_types::{CertificateDer, PrivateKeyDer, ServerName}; +use rustls::{ClientConfig, RootCertStore}; +use serde_json::Value; +use tokio::io::{AsyncRead, AsyncWrite}; +use tokio::net::TcpStream; +use tokio::time::{sleep, timeout}; +use tokio_rustls::TlsConnector; +use url::Position; + +const SERVICE_PORT: &str = "4500"; +const BEARER_TOKEN: &str = "Bearer openshell-e2e-application-token"; +const READY_TIMEOUT: Duration = Duration::from_secs(60); +const HEADER_ECHO_SERVER: &str = r#" +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer + +class Handler(BaseHTTPRequestHandler): + def do_GET(self): + body = self.headers.get("Authorization", "").encode() + self.send_response(200) + self.send_header("Content-Type", "text/plain") + self.send_header("Content-Length", str(len(body))) + self.end_headers() + self.wfile.write(body) + + def log_message(self, _format, *_args): + pass + +ThreadingHTTPServer(("127.0.0.1", 4500), Handler).serve_forever() +"#; + +async fn run_cli(args: &[&str]) -> Result { + let mut command = openshell_cmd(); + command + .args(args) + .stdout(Stdio::piped()) + .stderr(Stdio::piped()); + command + .output() + .await + .map_err(|error| format!("failed to run openshell: {error}")) +} + +enum ServiceTransport { + Http, + Https { + connector: TlsConnector, + server_name: ServerName<'static>, + }, +} + +struct ServiceTarget { + port: u16, + authority: String, + path: String, + transport: ServiceTransport, +} + +impl ServiceTarget { + fn from_url(url: &str) -> Result { + let url = url::Url::parse(url).map_err(|error| format!("invalid service URL: {error}"))?; + let host = url + .host_str() + .ok_or_else(|| "service URL omitted its host".to_string())?; + let port = url + .port_or_known_default() + .ok_or_else(|| "service URL omitted its port".to_string())?; + let transport = match url.scheme() { + "http" => ServiceTransport::Http, + "https" => ServiceTransport::Https { + connector: e2e_tls_connector()?, + server_name: ServerName::try_from(host.to_string()) + .map_err(|error| format!("invalid service TLS server name: {error}"))?, + }, + scheme => return Err(format!("unsupported service URL scheme {scheme:?}")), + }; + let authority = url[Position::BeforeHost..Position::AfterPort].to_string(); + let path = url.query().map_or_else( + || url.path().to_string(), + |query| format!("{}?{query}", url.path()), + ); + + Ok(Self { + port, + authority, + path, + transport, + }) + } +} + +fn e2e_mtls_dir() -> Result { + let config_home = std::env::var_os("XDG_CONFIG_HOME") + .ok_or_else(|| "XDG_CONFIG_HOME is required for an HTTPS service URL".to_string())?; + let gateway = std::env::var("OPENSHELL_GATEWAY") + .map_err(|_| "OPENSHELL_GATEWAY is required for an HTTPS service URL".to_string())?; + Ok(PathBuf::from(config_home) + .join("openshell/gateways") + .join(gateway) + .join("mtls")) +} + +fn load_certificates( + path: &Path, + description: &str, +) -> Result>, String> { + let file = File::open(path) + .map_err(|error| format!("open {description} '{}': {error}", path.display()))?; + let mut reader = BufReader::new(file); + rustls_pemfile::certs(&mut reader) + .collect::, _>>() + .map_err(|error| format!("parse {description} '{}': {error}", path.display())) +} + +fn load_private_key(path: &Path) -> Result, String> { + let file = File::open(path) + .map_err(|error| format!("open client TLS key '{}': {error}", path.display()))?; + let mut reader = BufReader::new(file); + rustls_pemfile::private_key(&mut reader) + .map_err(|error| format!("parse client TLS key '{}': {error}", path.display()))? + .ok_or_else(|| format!("client TLS key '{}' is empty", path.display())) +} + +fn e2e_tls_connector() -> Result { + let mtls_dir = e2e_mtls_dir()?; + let ca_path = mtls_dir.join("ca.crt"); + let mut roots = RootCertStore::empty(); + for certificate in load_certificates(&ca_path, "gateway CA certificate")? { + roots.add(certificate).map_err(|error| { + format!( + "add gateway CA certificate '{}': {error}", + ca_path.display() + ) + })?; + } + let client_certificates = + load_certificates(&mtls_dir.join("tls.crt"), "client TLS certificate")?; + let client_key = load_private_key(&mtls_dir.join("tls.key"))?; + let config = ClientConfig::builder() + .with_root_certificates(roots) + .with_client_auth_cert(client_certificates, client_key) + .map_err(|error| format!("build e2e mTLS client configuration: {error}"))?; + Ok(TlsConnector::from(Arc::new(config))) +} + +async fn request_over_stream( + stream: S, + target: &ServiceTarget, + authorization: &str, +) -> Result<(StatusCode, String), String> +where + S: AsyncRead + AsyncWrite + Unpin + Send + 'static, +{ + let (mut sender, connection) = http1::handshake(TokioIo::new(stream)) + .await + .map_err(|error| format!("start HTTP connection: {error}"))?; + tokio::spawn(async move { + let _ = connection.await; + }); + + let request = Request::builder() + .uri(&target.path) + .header(header::HOST, &target.authority) + .header(header::AUTHORIZATION, authorization) + .body(Empty::::new()) + .map_err(|error| format!("build service request: {error}"))?; + let response = sender + .send_request(request) + .await + .map_err(|error| format!("send service request: {error}"))?; + let status = response.status(); + let body = response + .into_body() + .collect() + .await + .map_err(|error| format!("read service response: {error}"))? + .to_bytes(); + let body = String::from_utf8(body.to_vec()) + .map_err(|error| format!("service returned non-UTF-8 data: {error}"))?; + Ok((status, body)) +} + +async fn request_service( + target: &ServiceTarget, + authorization: &str, +) -> Result<(StatusCode, String), String> { + // The service hostname is a virtual routing authority. Dial the loopback + // gateway directly so the test does not depend on the host resolver + // recognizing arbitrary subdomains of `.localhost`. + let stream = TcpStream::connect((Ipv4Addr::LOCALHOST, target.port)) + .await + .map_err(|error| format!("connect to loopback service gateway: {error}"))?; + let _ = stream.set_nodelay(true); + match &target.transport { + ServiceTransport::Http => request_over_stream(stream, target, authorization).await, + ServiceTransport::Https { + connector, + server_name, + } => { + let stream = connector + .connect(server_name.clone(), stream) + .await + .map_err(|error| format!("start service TLS connection: {error}"))?; + request_over_stream(stream, target, authorization).await + } + } +} + +async fn wait_for_authorization(url: &str, expected: &str) -> Result<(), String> { + // Parse the URL and load TLS material before the retry loop so permanent + // test-configuration errors fail immediately instead of looking like + // service-readiness timeouts. + let target = ServiceTarget::from_url(url)?; + let mut last_observation = "no request attempted".to_string(); + let result = timeout(READY_TIMEOUT, async { + loop { + match request_service(&target, BEARER_TOKEN).await { + Ok((StatusCode::OK, body)) if body == expected => return Ok(()), + Ok((StatusCode::OK, body)) => { + return Err(format!( + "service received unexpected Authorization value: {body:?}" + )); + } + Ok((status, body)) + if matches!( + status, + StatusCode::BAD_GATEWAY + | StatusCode::PRECONDITION_FAILED + | StatusCode::SERVICE_UNAVAILABLE + ) => + { + last_observation = + format!("service returned retryable status {status} with body {body:?}"); + sleep(Duration::from_millis(250)).await; + } + Err(error) => { + last_observation = error; + sleep(Duration::from_millis(250)).await; + } + Ok((status, body)) => { + return Err(format!( + "service returned unexpected status {status} with body {body:?}" + )); + } + } + } + }) + .await; + + match result { + Ok(result) => result, + Err(_) => Err(format!( + "timed out waiting for the exposed service; last observation: {last_observation}" + )), + } +} + +#[tokio::test] +async fn service_bearer_passthrough_preserves_authorization_header() { + let sandbox_name = format!("service-auth-{}", std::process::id()); + let create = run_cli(&[ + "sandbox", + "create", + "--name", + &sandbox_name, + "--from", + E2E_WORKLOAD_IMAGE, + "--expose", + SERVICE_PORT, + "--output", + "json", + "--detach", + "--no-tty", + "--", + "python3", + "-c", + HEADER_ECHO_SERVER, + ]) + .await + .expect("run sandbox create"); + assert!( + create.status.success(), + "sandbox create failed with exit {:?}: {}", + create.status.code(), + String::from_utf8_lossy(&create.stderr) + ); + let mut sandbox = SandboxGuard::manage_existing(sandbox_name.clone()); + + let created: Value = serde_json::from_slice(&create.stdout).expect("parse sandbox create JSON"); + let service_url = created + .get("service_urls") + .and_then(|urls| urls.get("")) + .and_then(Value::as_str) + .expect("unnamed service URL in create response"); + + wait_for_authorization(service_url, "") + .await + .expect("default mode should strip Authorization"); + + let expose = run_cli(&[ + "service", + "expose", + &sandbox_name, + SERVICE_PORT, + "--authorization-mode", + "bearer-passthrough", + ]) + .await + .expect("run service re-expose"); + assert!( + expose.status.success(), + "service re-expose failed with exit {:?}: {}", + expose.status.code(), + String::from_utf8_lossy(&expose.stderr) + ); + + wait_for_authorization(service_url, BEARER_TOKEN) + .await + .expect("passthrough mode should preserve Authorization"); + + sandbox.cleanup().await; +} diff --git a/examples/governance-interceptor/Cargo.lock b/examples/governance-interceptor/Cargo.lock index 9aeef0fec9..d329018546 100644 --- a/examples/governance-interceptor/Cargo.lock +++ b/examples/governance-interceptor/Cargo.lock @@ -1142,6 +1142,7 @@ dependencies = [ "noyalib", "serde", "serde_json", + "thiserror", ] [[package]] diff --git a/examples/provider-managed-files/acme-config.yaml b/examples/provider-managed-files/acme-config.yaml new file mode 100644 index 0000000000..59919575f5 --- /dev/null +++ b/examples/provider-managed-files/acme-config.yaml @@ -0,0 +1,13 @@ +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +id: acme-config +display_name: Acme client configuration +description: Non-secret settings installed as a managed sandbox file +category: other +files: + - path: client.toml + env_var: ACME_CONFIG_FILE + content: | + endpoint = "{{config.endpoint}}" + project = "{{config.project}}" diff --git a/examples/sandbox-policy-quickstart/README.md b/examples/sandbox-policy-quickstart/README.md index 156348dddc..11da5f9eaa 100644 --- a/examples/sandbox-policy-quickstart/README.md +++ b/examples/sandbox-policy-quickstart/README.md @@ -222,7 +222,7 @@ without `--dangerously-skip-permissions`. - **Customize the policy**: Change `access: read-only` to `read-write` or add explicit `rules` for specific paths. See the - [security policy reference](../../architecture/security-policy.md). + [policy schema reference](../../docs/how-it-works/policies/schema.mdx). - **Scope to an agent**: Replace the `binaries` section with your agent's binary (e.g., `/usr/local/bin/claude`) instead of `curl`. - **Add more endpoints**: Stack multiple policies in the same file diff --git a/examples/supervisor-middleware-content-guard/Cargo.lock b/examples/supervisor-middleware-content-guard/Cargo.lock index 381cb63c40..0927d784b0 100644 --- a/examples/supervisor-middleware-content-guard/Cargo.lock +++ b/examples/supervisor-middleware-content-guard/Cargo.lock @@ -1211,6 +1211,7 @@ dependencies = [ "noyalib", "serde", "serde_json", + "thiserror", ] [[package]] diff --git a/proto/README.md b/proto/README.md index 61e431c5b8..725fb21129 100644 --- a/proto/README.md +++ b/proto/README.md @@ -80,7 +80,8 @@ identifies the service being exposed. - When removing or renaming a field, reserve its old number and source name. Do not reuse either for a different meaning. - Review changes against both the public descriptor closure and durable stored - protobuf closure described in [the gateway architecture](../architecture/gateway.md#protobuf-api-and-storage-boundaries). + protobuf closure. The `public_and_durable_schema_inventories_are_complete` + test in `openshell-server` owns both inventories. - Regenerate Rust, Python, Go, and TypeScript bindings after contract changes. Run `mise run pre-commit`, the affected SDK checks, and relevant server tests before submitting the change. diff --git a/proto/openshell.proto b/proto/openshell.proto index a836e169bf..621306002d 100644 --- a/proto/openshell.proto +++ b/proto/openshell.proto @@ -1045,6 +1045,9 @@ message SandboxSpec { // Gateway-owned attachment identity, changed atomically with the provider set. // Equality only: detach and reattach must not revive an older receipt. string provider_attachment_epoch = 14; + // Gateway-owned policy for replacing the sandbox runtime after the canonical + // main process exits. Unspecified is normalized to Never before persistence. + SandboxRestartPolicy restart_policy = 15; } message ResourceRequirements { @@ -1170,9 +1173,8 @@ message SandboxStatus { // Supervisor instance currently associated with the canonical main process. // The gateway uses this to reject stale exit reports after a restart. string main_process_instance_id = 8; - // Normalized main process result. Signal exits use 128 + signal number. - // Presence indicates that the canonical main process exited. Exit code 0 - // produces Completed; nonzero and signal-normalized exits produce Error. + // Most recent normalized main process result. Signal exits use 128 + signal number. + // Cleared when a replacement main process becomes ready. optional int32 exit_code = 9; // Last accepted network result for each configured tool server endpoint. // Currently populated for MCP-over-HTTP endpoints. These passive results @@ -1184,6 +1186,12 @@ message SandboxStatus { optional bool configuration_activated = 12; // Gateway-owned repair window. Retained after timeout for inspection and retry. SandboxProvisioning provisioning = 13; + // Consecutive policy-driven restart number in the current crash loop. + uint32 restart_count = 14; + // Deadline for the next restart attempt or recovery check. + google.protobuf.Timestamp next_restart_time = 15; + // Time when the current main process became ready. + google.protobuf.Timestamp main_process_started_time = 16; } // User-facing sandbox condition derived from platform or gateway observations. @@ -1218,6 +1226,8 @@ enum SandboxPhase { SANDBOX_PHASE_STARTING = 8; // The canonical main process exited successfully and its result is final. SANDBOX_PHASE_COMPLETED = 9; + reserved 10; + reserved "SANDBOX_PHASE_RESTARTING"; } // Public platform event exposed on the sandbox watch stream. @@ -1764,6 +1774,8 @@ message ExposeServiceRequest { // Optional nonzero UUID for durable at-most-once admission. Successful results // can be replayed for 24 hours; see the API errors and retries reference. string request_id = 6; + // Application authorization behavior. Omission resolves to STRIP. + ServiceAuthorizationMode authorization_mode = 7; } // Request to fetch an exposed sandbox service endpoint. @@ -1830,6 +1842,8 @@ message ServiceEndpoint { uint32 target_port = 5; // Whether browser-facing service routing is enabled for this endpoint. bool domain = 6; + // Effective application authorization behavior for ingress requests. + ServiceAuthorizationMode authorization_mode = 7; } // Response containing a service endpoint and, when available, its local URL. @@ -2006,9 +2020,14 @@ message WatchSandboxRequest { bool follow_events = 4; // Replay the last N log lines (best-effort) before following. + // + // When following both sources, the replay can return fewer lines: events + // older than one the other source's replay left out are withheld and + // reported with a warning, even if both depths are equal. uint32 log_tail_lines = 5; // Replay the last N platform events (best-effort) before following. + // Defaults to 0. Can return fewer events; see log_tail_lines. uint32 event_tail = 6; // Stop streaming once the sandbox reaches READY or a terminal result phase @@ -2042,6 +2061,8 @@ message WatchSandboxRequest { // empty resume_after_cursor, because retrying the same token fails // identically. A cursor this server could not have issued is rejected with // INVALID_ARGUMENT. + // + // One cursor covers both log and platform events. string resume_after_cursor = 12; } @@ -2470,6 +2491,19 @@ message ProviderProfile { // Server-set visibility: "platform", "workspace", or empty for // non-scoped sources. Ignored on import/update payloads. string scope = 13; + // EXPERIMENTAL: Non-secret files rendered from provider configuration for the + // workload. This API and its behavior may change or be removed. + repeated ProviderProfileFile files = 14; +} + +// EXPERIMENTAL: Provider file templates may change or be removed. +message ProviderProfileFile { + // One virtual file name below /run/openshell/providers//. + string path = 1; + // UTF-8 template. Only {{config.KEY}} references are supported. + string content = 2; + // Optional environment variable containing the virtual absolute path. + string env_var = 3; } // Provider profile response. @@ -2623,6 +2657,8 @@ message GetSandboxProviderEnvironmentResponse { string policy_hash = 8; // Nonzero when material was withheld; installing an empty map is not readiness. ProviderReadinessReason readiness_reason = 9; + // Complete desired set of non-secret managed files, keyed by absolute path. + map files = 10; } message ExchangeProviderSubjectTokenRequest { @@ -3552,6 +3588,16 @@ enum ProviderCredentialRefreshRecoveryAction { PROVIDER_CREDENTIAL_REFRESH_RECOVERY_ACTION_INVESTIGATE = 4; } +// Policy applied when the canonical main process exits outside an intentional +// stop or delete operation. Kept after existing enums so generated enum +// descriptors remain stable. +enum SandboxRestartPolicy { + SANDBOX_RESTART_POLICY_UNSPECIFIED = 0; + SANDBOX_RESTART_POLICY_NEVER = 1; + SANDBOX_RESTART_POLICY_ON_FAILURE = 2; + SANDBOX_RESTART_POLICY_ALWAYS = 3; +} + // Workspace membership record. message WorkspaceMember { openshell.datamodel.v1.ObjectMeta metadata = 1; @@ -3750,4 +3796,17 @@ message SandboxServiceExposure { string service = 1; // Loopback TCP port inside the sandbox. uint32 target_port = 2; + // Application authorization behavior. Omission resolves to STRIP. + ServiceAuthorizationMode authorization_mode = 3; +} + +// Controls whether an exposed service receives the request's application +// Authorization header. +enum ServiceAuthorizationMode { + // Omission preserves the secure legacy behavior and resolves to STRIP. + SERVICE_AUTHORIZATION_MODE_UNSPECIFIED = 0; + // Remove Authorization before forwarding to the sandbox service. + SERVICE_AUTHORIZATION_MODE_STRIP = 1; + // Forward one syntactically valid bearer Authorization header unchanged. + SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH = 2; } diff --git a/providers/oci-genai.yaml b/providers/oci-genai.yaml new file mode 100644 index 0000000000..fc48abb203 --- /dev/null +++ b/providers/oci-genai.yaml @@ -0,0 +1,93 @@ +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +# Example provider profile. OpenShell does not load it; import it explicitly: +# openshell provider profile lint -f providers/oci-genai.yaml +# openshell provider profile import -f providers/oci-genai.yaml --global +# +# Copy and edit this file rather than importing it unchanged. `binaries` is the +# least-privilege control that decides which processes may reach the endpoints +# below, so it has to name the paths in *your* image. +# +# Client binaries: curl. Add the interpreter or agent CLI that calls the +# OpenAI-compatible endpoint (python, node, ...) before +# relying on binary-scoped attribution for this profile. +# Reference layout: any image with curl or an OpenAI SDK. Point the SDK at +# https://inference.generativeai..oci.oraclecloud.com/openai/v1 +# and pass OCI_GENAI_API_KEY as the API key; the sandbox +# proxy resolves the placeholder only at the hosts below. +# Credential scope: one OCI Generative AI API key (sk-...). The key is scoped +# to the compartment it was created in, so requests need no +# opc-compartment-id header and the sandbox never learns a +# compartment OCID. +# Endpoint access: the regional OCI Generative AI inference hosts in the +# commercial realm (oraclecloud.com), TLS terminated, L7 +# enforced, limited to the OpenAI-compatible /openai/v1 +# surface with GET, POST, and DELETE. Generative AI API keys +# are only valid there; the native /20231130 API needs OCI +# request signing, which this profile does not grant. +# Smoke test: openshell sandbox create --provider -- \ +# curl -sS --compressed \ +# https://inference.generativeai.us-chicago-1.oci.oraclecloud.com/openai/v1/chat/completions \ +# -H "Authorization: Bearer $OCI_GENAI_API_KEY" \ +# -H 'Content-Type: application/json' \ +# -d '{"model":"meta.llama-3.3-70b-instruct","messages":[{"role":"user","content":"Reply with OK"}]}' +# +# OCI-side setup, in this order: +# 1. Create the IAM policy BEFORE the key. A key minted before its policy can +# stay unauthorized long after the policy lands, and OCI returns the same +# 401 for an unknown key and an unauthorized one. Least privilege: +# allow any-user to use generative-ai-family in compartment id +# where ALL {request.principal.type='generativeaiapikey'} +# Pin it to one key later by adding request.principal.id=''. +# 2. Create the key in that compartment and region, with an expiry: +# oci generative-ai api-key create --compartment-id \ +# --region us-chicago-1 --display-name openshell-agents \ +# --key-details '[{"keyName":"primary","timeExpiry":"2027-01-01T00:00:00Z"}]' +# The sk-... secret is returned once. Rotate with `api-key renew`, +# disable with `api-key set-api-key-state`; sandboxes are unaffected +# because the gateway, not the sandbox, holds the value. +# 3. openshell provider create --name my-oci-genai --type oci-genai \ +# --credential OCI_GENAI_API_KEY="$OCI_GENAI_API_KEY" +# +# Verified through the sandbox proxy with this profile: chat completions with +# stream=true, tool calling, vision input via data: URLs (for example +# meta.llama-4-maverick-17b-128e-instruct-fp8 and google.gemini-2.5-flash, +# including a 263 KB request body), embeddings with openai.text-embedding-3-small, +# and the Responses API with stream=true. Model ids use the OCI form, such as +# meta.llama-3.3-70b-instruct or openai.gpt-oss-120b. Expect from OCI, not from +# the proxy: 404 "Entity with key not found" for models not served on +# this endpoint, 404 on GET /openai/v1/models (not a health check), and +# 400 "Unsupported OpenAI operation" for Cohere embeddings. +# +# Other realms use their realm domain instead of oraclecloud.com: copy the +# profile and add a matching endpoint entry. Signed OCI transports (API key +# signing, security token, instance/resource principal, OKE workload identity) +# need proxy-side request signing; see https://github.com/NVIDIA/OpenShell/issues/3879. + +id: oci-genai +display_name: OCI Generative AI +description: OCI Generative AI inference through the OpenAI-compatible endpoint with a compartment-scoped Generative AI API key +category: inference +inference_capable: true +credentials: + - name: api_key + description: OCI Generative AI API key (compartment-scoped, sent as a bearer token) + env_vars: [OCI_GENAI_API_KEY] + required: true + auth_style: bearer + header_name: authorization +discovery: + credentials: [api_key] +endpoints: + # `*` matches exactly one DNS label, which is the region identifier, for + # example us-chicago-1, eu-frankfurt-1, or ap-osaka-1. + - host: "inference.generativeai.*.oci.oraclecloud.com" + port: 443 + protocol: rest + enforcement: enforce + rules: + - allow: { method: POST, path: "/openai/v1/**" } + - allow: { method: GET, path: "/openai/v1/**" } + - allow: { method: DELETE, path: "/openai/v1/**" } +binaries: [/usr/bin/curl, /usr/local/bin/curl] diff --git a/python/openshell/__init__.py b/python/openshell/__init__.py index 8d0d20f578..109fe09521 100644 --- a/python/openshell/__init__.py +++ b/python/openshell/__init__.py @@ -21,6 +21,7 @@ SandboxStatusRef, SandboxTemplateClient, SandboxWorkloadTemplateProvenanceRef, + ServiceAuthorizationMode, ServiceExposure, TlsConfig, WorkspaceClient, @@ -53,6 +54,7 @@ "SandboxStatusRef", "SandboxTemplateClient", "SandboxWorkloadTemplateProvenanceRef", + "ServiceAuthorizationMode", "ServiceExposure", "TlsConfig", "WorkspaceClient", diff --git a/python/openshell/sandbox.py b/python/openshell/sandbox.py index d85d4c70fa..485af39d04 100644 --- a/python/openshell/sandbox.py +++ b/python/openshell/sandbox.py @@ -17,6 +17,7 @@ import time from collections import namedtuple from dataclasses import dataclass, field +from enum import IntEnum from typing import TYPE_CHECKING, Any, Generic, Never, SupportsIndex, TypeVar, cast from urllib.parse import urlparse @@ -115,6 +116,12 @@ def _service_exposure_messages( openshell_pb2.SandboxServiceExposure( service=exposure.service, target_port=exposure.target_port, + authorization_mode=( + openshell_pb2.SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH + if exposure.authorization_mode + == ServiceAuthorizationMode.BEARER_PASSTHROUGH + else openshell_pb2.SERVICE_AUTHORIZATION_MODE_STRIP + ), ) for exposure in exposures or () ] @@ -448,6 +455,16 @@ class SandboxStatusRef: phase: int current_policy_version: int exit_code: int | None = None + restart_count: int = 0 + next_restart_at_ms: int | None = None + main_process_started_at_ms: int | None = None + + +class ServiceAuthorizationMode(IntEnum): + """Handling for an incoming application Authorization header.""" + + STRIP = 1 + BEARER_PASSTHROUGH = 2 @dataclass(frozen=True) @@ -456,6 +473,7 @@ class ServiceExposure: target_port: int service: str = "" + authorization_mode: ServiceAuthorizationMode = ServiceAuthorizationMode.STRIP class _ImmutableLabels(dict[str, str]): @@ -1767,6 +1785,13 @@ def _sandbox_ref( exit_code=status.exit_code if status is not None and status.HasField("exit_code") else None, + restart_count=status.restart_count if status is not None else 0, + next_restart_at_ms=status.next_restart_time.ToMilliseconds() + if status is not None and status.HasField("next_restart_time") + else None, + main_process_started_at_ms=status.main_process_started_time.ToMilliseconds() + if status is not None and status.HasField("main_process_started_time") + else None, ), labels=sandbox.metadata.labels if sandbox.metadata else {}, created_from_workload_template=provenance, diff --git a/python/openshell/sandbox_test.py b/python/openshell/sandbox_test.py index 1d32ea3a80..3b744451cb 100644 --- a/python/openshell/sandbox_test.py +++ b/python/openshell/sandbox_test.py @@ -32,6 +32,7 @@ SandboxRef, SandboxStatusRef, SandboxTemplateClient, + ServiceAuthorizationMode, ServiceExposure, TlsConfig, _atomic_replace, @@ -2219,15 +2220,26 @@ def test_create_forwards_service_exposures() -> None: name="app-server", service_exposures=[ ServiceExposure(target_port=4500), - ServiceExposure(service="metrics", target_port=9090), + ServiceExposure( + service="metrics", + target_port=9090, + authorization_mode=ServiceAuthorizationMode.BEARER_PASSTHROUGH, + ), ], ) assert stub.create_request is not None assert [ - (exposure.service, exposure.target_port) + (exposure.service, exposure.target_port, exposure.authorization_mode) for exposure in stub.create_request.service_exposures - ] == [("", 4500), ("metrics", 9090)] + ] == [ + ("", 4500, openshell_pb2.SERVICE_AUTHORIZATION_MODE_STRIP), + ( + "metrics", + 9090, + openshell_pb2.SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH, + ), + ] assert dict(ref.service_urls) == { "": "https://.example.test/", "metrics": "https://metrics.example.test/", @@ -2719,10 +2731,16 @@ def test_sandbox_ref_retains_gateway_labels() -> None: def test_sandbox_ref_includes_main_process_result() -> None: proto = _make_sandbox_proto("sandbox-1", "job-1") proto.status.exit_code = 0 + proto.status.restart_count = 2 + proto.status.next_restart_time.FromMilliseconds(1_700_000_000_000) + proto.status.main_process_started_time.FromMilliseconds(1_699_999_000_000) status = _sandbox_ref(proto).status assert status.exit_code == 0 + assert status.restart_count == 2 + assert status.next_restart_at_ms == 1_700_000_000_000 + assert status.main_process_started_at_ms == 1_699_999_000_000 def test_returned_labels_are_immutable() -> None: diff --git a/rfc/0012-isolation-backend/README.md b/rfc/0012-isolation-backend/README.md index 0aec01e244..b384c13e5e 100644 --- a/rfc/0012-isolation-backend/README.md +++ b/rfc/0012-isolation-backend/README.md @@ -198,7 +198,7 @@ trait RunningBoundary: Send + Sync { `AgentSpec` carries the complete admitted agent launch specification, including command, arguments, working directory, timeout, and interactive mode. -`SandboxContext` carries the admitted create-time policy and the identity of this launch. [RFC 0002](../0002-agent-driven-policy-management/README.md) defines how network-policy revisions are proposed and approved. Approved revisions reach the supervisor through the existing [`GetSandboxConfig`](../../proto/sandbox.proto) gateway-supervisor contract, described in the [gateway](../../architecture/gateway.md) and [sandbox](../../architecture/sandbox.md#policy-revision-acknowledgement) architecture. The supervisor makes approved network-policy revisions effective through network mediation. If an approved network-policy revision cannot be loaded, it never becomes effective; the configured rejection posture retains the last valid generation or denies network access until a valid generation is loaded. +`SandboxContext` carries the admitted create-time policy and the identity of this launch. [RFC 0002](../0002-agent-driven-policy-management/README.md) defines how network-policy revisions are proposed and approved. Approved revisions reach the supervisor through the existing [`GetSandboxConfig`](../../proto/sandbox.proto) gateway-supervisor contract. The supervisor makes approved network-policy revisions effective through network mediation. If an approved network-policy revision cannot be loaded, it never becomes effective; the configured rejection posture retains the last valid generation or denies network access until a valid generation is loaded. The states have normative meanings: diff --git a/rfc/README.md b/rfc/README.md index 4a0072f5c2..2c4260141d 100644 --- a/rfc/README.md +++ b/rfc/README.md @@ -24,9 +24,8 @@ OpenShell has several places where design information lives. Use this guide to p | **GitHub issue** | Track and scope a bug, feature request, or rough idea | Always start here; new features use the feature request template | | **Spike issue** (`create-spike`) | Investigate implementation feasibility for a scoped change | You need to explore the codebase and produce a buildable issue for a specific component or feature | | **RFC** | Propose a cross-cutting decision that needs broad consensus | Maintainers requested an RFC from an existing issue and assigned an RFC number | -| **Architecture doc** (`architecture/`) | Document how things work today | Living reference material — updated as the system evolves | -The key distinction: **spikes investigate whether and how something can be done; RFCs propose that we should do it and seek agreement on the approach.** A spike may precede an RFC (to gather data) or follow one (to flesh out implementation details). When an RFC reaches `implemented`, its relevant content should be folded into the appropriate `architecture/` docs so the living reference stays current. +The key distinction: **spikes investigate whether and how something can be done; RFCs propose that we should do it and seek agreement on the approach.** A spike may precede an RFC (to gather data) or follow one (to flesh out implementation details). ## When to use an RFC diff --git a/scripts/baseline_workflow_metrics.py b/scripts/baseline_workflow_metrics.py index 430fc93e23..02d1448f8d 100755 --- a/scripts/baseline_workflow_metrics.py +++ b/scripts/baseline_workflow_metrics.py @@ -13,7 +13,7 @@ Usage: uv run python scripts/baseline_workflow_metrics.py - uv run python scripts/baseline_workflow_metrics.py --days 30 --out architecture/plans/OS-49-baseline.json + uv run python scripts/baseline_workflow_metrics.py --days 30 --out plans/OS-49-baseline.json Auth: Relies on `gh auth login` — the script shells out to `gh api` so no token @@ -360,14 +360,14 @@ def parse_args() -> argparse.Namespace: parser.add_argument( "--out", type=pathlib.Path, - default=pathlib.Path("architecture/plans/OS-49-baseline.json"), - help="Where to write the JSON report (default: architecture/plans/OS-49-baseline.json)", + default=pathlib.Path("plans/OS-49-baseline.json"), + help="Where to write the JSON report (default: plans/OS-49-baseline.json)", ) parser.add_argument( "--md", type=pathlib.Path, - default=pathlib.Path("architecture/plans/OS-49-baseline.md"), - help="Where to write the Markdown report (default: architecture/plans/OS-49-baseline.md)", + default=pathlib.Path("plans/OS-49-baseline.md"), + help="Where to write the Markdown report (default: plans/OS-49-baseline.md)", ) return parser.parse_args() diff --git a/scripts/update_license_headers.py b/scripts/update_license_headers.py index 43746f2495..48d7810171 100755 --- a/scripts/update_license_headers.py +++ b/scripts/update_license_headers.py @@ -75,6 +75,7 @@ EXCLUDE_DIRS: set[str] = { "target", "e2e/rust/target", + "plans", "architecture/plans", "scripts/lint-mermaid/node_modules", ".venv", diff --git a/sdk/go/openshell/v1/fake/service.go b/sdk/go/openshell/v1/fake/service.go index 464d8f7fc9..3d7a1e0331 100644 --- a/sdk/go/openshell/v1/fake/service.go +++ b/sdk/go/openshell/v1/fake/service.go @@ -22,7 +22,7 @@ func newFakeServiceClient(closedFunc func() bool) *fakeServiceClient { } // Expose returns Unimplemented. -func (c *fakeServiceClient) Expose(_ context.Context, _, _, _ string, _ uint32, _ bool) (*types.ServiceEndpoint, error) { +func (c *fakeServiceClient) Expose(_ context.Context, _, _, _ string, _ uint32, _ bool, _ ...v1.ExposeServiceOptions) (*types.ServiceEndpoint, error) { if c.closedFunc() { return nil, &types.StatusError{Code: types.ErrorUnavailable, Message: "client is closed"} } diff --git a/sdk/go/openshell/v1/internal/converter/coverage_test.go b/sdk/go/openshell/v1/internal/converter/coverage_test.go index 2f90b73a5d..43577d61d0 100644 --- a/sdk/go/openshell/v1/internal/converter/coverage_test.go +++ b/sdk/go/openshell/v1/internal/converter/coverage_test.go @@ -31,6 +31,7 @@ func TestConverterCoversAllProtoFields_SandboxSpec(t *testing.T) { "resource_requirements": true, "command": true, "tty": true, + "restart_policy": true, } // The gateway owns this identity. Provider status exposes it through the @@ -115,15 +116,18 @@ func TestConverterCoversAllProtoFields_SandboxStartup(t *testing.T) { func TestConverterCoversAllProtoFields_SandboxStatus(t *testing.T) { handled := fieldSet{ - "agent_pod": true, - "agent_fd": true, - "sandbox_fd": true, - "phase": true, - "conditions": true, - "endpoint_statuses": true, - "current_policy_version": true, - "exit_code": true, - "configuration_admission": true, + "agent_pod": true, + "agent_fd": true, + "sandbox_fd": true, + "phase": true, + "conditions": true, + "endpoint_statuses": true, + "current_policy_version": true, + "exit_code": true, + "configuration_admission": true, + "restart_count": true, + "next_restart_time": true, + "main_process_started_time": true, } // The instance ID coordinates internal gateway/supervisor lifecycle // fencing. The first-activation marker governs static policy repair. @@ -304,6 +308,7 @@ func TestConverterCoversAllProtoFields_ProviderProfile(t *testing.T) { "description": true, "category": true, "credentials": true, + "files": true, "endpoints": true, "binaries": true, "inference_capable": true, diff --git a/sdk/go/openshell/v1/internal/converter/profile.go b/sdk/go/openshell/v1/internal/converter/profile.go index 12be9618b8..9ef58e8057 100644 --- a/sdk/go/openshell/v1/internal/converter/profile.go +++ b/sdk/go/openshell/v1/internal/converter/profile.go @@ -363,6 +363,18 @@ func ProviderProfileFromProto(p *pb.ProviderProfile) *types.ProviderProfile { } } + // Files + if files := p.GetFiles(); len(files) > 0 { + result.Files = make([]types.ProfileFile, len(files)) + for i, file := range files { + if file != nil { + result.Files[i] = types.ProfileFile{ + Path: file.GetPath(), Content: file.GetContent(), EnvVar: file.GetEnvVar(), + } + } + } + } + // Endpoints if eps := p.GetEndpoints(); len(eps) > 0 { result.Endpoints = make([]types.NetworkEndpoint, len(eps)) @@ -419,6 +431,16 @@ func ProviderProfileToProto(p *types.ProviderProfile) *pb.ProviderProfile { } } + // Files + if len(p.Files) > 0 { + result.Files = make([]*pb.ProviderProfileFile, len(p.Files)) + for i, file := range p.Files { + result.Files[i] = &pb.ProviderProfileFile{ + Path: file.Path, Content: file.Content, EnvVar: file.EnvVar, + } + } + } + // Endpoints if len(p.Endpoints) > 0 { result.Endpoints = make([]*sbv1.NetworkEndpoint, len(p.Endpoints)) diff --git a/sdk/go/openshell/v1/internal/converter/profile_test.go b/sdk/go/openshell/v1/internal/converter/profile_test.go index 162542d1a0..e1aaf0af75 100644 --- a/sdk/go/openshell/v1/internal/converter/profile_test.go +++ b/sdk/go/openshell/v1/internal/converter/profile_test.go @@ -470,6 +470,9 @@ func TestProviderProfileFromProto(t *testing.T) { Credentials: []*pb.ProviderProfileCredential{ {Name: "API_KEY", Description: "key", Required: true}, }, + Files: []*pb.ProviderProfileFile{ + {Path: "client.toml", Content: "project = \"{{config.project}}\"", EnvVar: "CLIENT_CONFIG"}, + }, Endpoints: []*sbv1.NetworkEndpoint{ {Host: "api.anthropic.com", Port: 443, Protocol: "rest"}, }, @@ -502,6 +505,9 @@ func TestProviderProfileFromProto(t *testing.T) { require.Len(t, profile.Credentials, 1) assert.Equal(t, "API_KEY", profile.Credentials[0].Name) assert.True(t, profile.Credentials[0].Required) + require.Len(t, profile.Files, 1) + assert.Equal(t, "client.toml", profile.Files[0].Path) + assert.Equal(t, "CLIENT_CONFIG", profile.Files[0].EnvVar) require.Len(t, profile.Endpoints, 1) assert.Equal(t, "api.anthropic.com", profile.Endpoints[0].Host) @@ -514,6 +520,8 @@ func TestProviderProfileFromProto(t *testing.T) { proto.Annotations["env"] = "MUTATED" assert.Equal(t, "prod", profile.Annotations["env"], "annotations must be deep copied") + proto.Files[0].Content = "changed" + assert.Equal(t, "project = \"{{config.project}}\"", profile.Files[0].Content) } func TestProviderProfileFromProto_NilDiscovery(t *testing.T) { @@ -541,6 +549,9 @@ func TestProviderProfileToProto(t *testing.T) { Credentials: []v1.ProfileCredential{ {Name: "API_KEY", Description: "key", Required: true, Secret: true}, }, + Files: []v1.ProfileFile{ + {Path: "client.toml", Content: "project = \"{{config.project}}\"", EnvVar: "CLIENT_CONFIG"}, + }, Endpoints: []v1.NetworkEndpoint{ {Host: "api.anthropic.com", Port: 443, Protocol: "rest"}, }, @@ -572,6 +583,8 @@ func TestProviderProfileToProto(t *testing.T) { require.Len(t, proto.Credentials, 1) assert.Equal(t, "API_KEY", proto.Credentials[0].Name) + require.Len(t, proto.Files, 1) + assert.Equal(t, "client.toml", proto.Files[0].Path) require.Len(t, proto.Endpoints, 1) assert.Equal(t, "api.anthropic.com", proto.Endpoints[0].Host) @@ -584,6 +597,8 @@ func TestProviderProfileToProto(t *testing.T) { profile.Annotations["env"] = "MUTATED" assert.Equal(t, "prod", proto.Annotations["env"], "annotations must be deep copied") + profile.Files[0].Content = "changed" + assert.Equal(t, "project = \"{{config.project}}\"", proto.Files[0].Content) } func TestProviderProfileToProto_Nil(t *testing.T) { @@ -670,6 +685,9 @@ func TestProviderProfileRoundTrip(t *testing.T) { }, }, }, + Files: []v1.ProfileFile{ + {Path: "client.toml", Content: "project = \"{{config.project}}\"", EnvVar: "CLIENT_CONFIG"}, + }, Endpoints: []v1.NetworkEndpoint{ {Host: "agent.example.com", Port: 8080, Protocol: "websocket"}, }, @@ -699,6 +717,7 @@ func TestProviderProfileRoundTrip(t *testing.T) { assert.Equal(t, original.Annotations, back.Annotations) assert.Equal(t, original.Source, back.Source) assert.Equal(t, original.Scope, back.Scope) + assert.Equal(t, original.Files, back.Files) require.Len(t, back.Credentials, 1) c := back.Credentials[0] diff --git a/sdk/go/openshell/v1/internal/converter/sandbox.go b/sdk/go/openshell/v1/internal/converter/sandbox.go index 1f773c100e..7d81a7c828 100644 --- a/sdk/go/openshell/v1/internal/converter/sandbox.go +++ b/sdk/go/openshell/v1/internal/converter/sandbox.go @@ -56,10 +56,11 @@ func SandboxFromProto(s *pb.Sandbox) *types.Sandbox { func sandboxSpecFromProto(spec *pb.SandboxSpec) types.SandboxSpec { result := types.SandboxSpec{ - LogLevel: spec.GetLogLevel(), - Environment: CopyStringMap(spec.GetEnvironment()), - Providers: CopyStringSlice(spec.GetProviders()), - Policy: SandboxPolicyFromProto(spec.GetPolicy()), + LogLevel: spec.GetLogLevel(), + Environment: CopyStringMap(spec.GetEnvironment()), + Providers: CopyStringSlice(spec.GetProviders()), + Policy: SandboxPolicyFromProto(spec.GetPolicy()), + RestartPolicy: SandboxRestartPolicyFromProto(spec.GetRestartPolicy()), } if tmpl := spec.GetTemplate(); tmpl != nil { @@ -143,6 +144,9 @@ func sandboxStatusFromProto(status *pb.SandboxStatus) types.SandboxStatus { Error: admission.GetError(), } } + result.RestartCount = status.GetRestartCount() + result.NextRestartAtMs = MillisFromProto(status.GetNextRestartTime()) + result.MainProcessStartedAtMs = MillisFromProto(status.GetMainProcessStartedTime()) return result } @@ -292,10 +296,35 @@ func SandboxSpecToProto(spec *types.SandboxSpec) *pb.SandboxSpec { result.Command = CopyStringSlice(spec.Command) result.Tty = spec.TTY + result.RestartPolicy = SandboxRestartPolicyToProto(spec.RestartPolicy) return result } +// SandboxRestartPolicyFromProto converts the restart policy to the curated SDK type. +func SandboxRestartPolicyFromProto(policy pb.SandboxRestartPolicy) types.SandboxRestartPolicy { + switch policy { + case pb.SandboxRestartPolicy_SANDBOX_RESTART_POLICY_ON_FAILURE: + return types.SandboxRestartOnFailure + case pb.SandboxRestartPolicy_SANDBOX_RESTART_POLICY_ALWAYS: + return types.SandboxRestartAlways + default: + return types.SandboxRestartNever + } +} + +// SandboxRestartPolicyToProto converts the curated SDK restart policy to protobuf. +func SandboxRestartPolicyToProto(policy types.SandboxRestartPolicy) pb.SandboxRestartPolicy { + switch policy { + case types.SandboxRestartOnFailure: + return pb.SandboxRestartPolicy_SANDBOX_RESTART_POLICY_ON_FAILURE + case types.SandboxRestartAlways: + return pb.SandboxRestartPolicy_SANDBOX_RESTART_POLICY_ALWAYS + default: + return pb.SandboxRestartPolicy_SANDBOX_RESTART_POLICY_NEVER + } +} + // SandboxSpecToProtoChecked converts an SDK SandboxSpec and reports values // that protobuf Struct cannot represent instead of silently dropping them. func SandboxSpecToProtoChecked(spec *types.SandboxSpec) (*pb.SandboxSpec, error) { diff --git a/sdk/go/openshell/v1/internal/converter/sandbox_test.go b/sdk/go/openshell/v1/internal/converter/sandbox_test.go index 598029031a..2099b5b65b 100644 --- a/sdk/go/openshell/v1/internal/converter/sandbox_test.go +++ b/sdk/go/openshell/v1/internal/converter/sandbox_test.go @@ -87,21 +87,25 @@ func TestSandboxFromProto(t *testing.T) { Count: &gpuCount, }, }, - Command: []string{"/opt/agent", "--serve"}, - Tty: false, + Command: []string{"/opt/agent", "--serve"}, + Tty: false, + RestartPolicy: pb.SandboxRestartPolicy_SANDBOX_RESTART_POLICY_ON_FAILURE, }, CreatedFromWorkloadTemplate: &pb.SandboxWorkloadTemplateProvenance{ Name: "gpu-kata", ResourceVersion: "7", }, Status: &pb.SandboxStatus{ - AgentPod: "agent-pod-xyz", - AgentFd: "fd-agent", - SandboxFd: "fd-sandbox", - Phase: pb.SandboxPhase_SANDBOX_PHASE_READY, - CurrentPolicyVersion: 7, - MainProcessInstanceId: "instance-1", - ExitCode: &exitCode, + AgentPod: "agent-pod-xyz", + AgentFd: "fd-agent", + SandboxFd: "fd-sandbox", + Phase: pb.SandboxPhase_SANDBOX_PHASE_READY, + CurrentPolicyVersion: 7, + MainProcessInstanceId: "instance-1", + ExitCode: &exitCode, + RestartCount: 3, + NextRestartTime: TimestampFromMillis(1700000070000), + MainProcessStartedTime: TimestampFromMillis(1700000010000), Conditions: []*pb.SandboxCondition{ { Type: "Ready", @@ -139,6 +143,7 @@ func TestSandboxFromProto(t *testing.T) { assert.Equal(t, uint32(2), *s.Spec.GPUCount) assert.Equal(t, []string{"/opt/agent", "--serve"}, s.Spec.Command) assert.False(t, s.Spec.TTY) + assert.Equal(t, v1.SandboxRestartOnFailure, s.Spec.RestartPolicy) // Template require.NotNil(t, s.Spec.Template) @@ -170,6 +175,9 @@ func TestSandboxFromProto(t *testing.T) { assert.Equal(t, "2024-01-01T00:00:00Z", s.Status.Conditions[0].LastTransitionTime) require.NotNil(t, s.Status.ExitCode) assert.Equal(t, int32(0), *s.Status.ExitCode) + assert.Equal(t, uint32(3), s.Status.RestartCount) + assert.Equal(t, int64(1700000070000), s.Status.NextRestartAtMs) + assert.Equal(t, int64(1700000010000), s.Status.MainProcessStartedAtMs) } func TestSandboxFromProto_TemplateResourcesDeepCopy(t *testing.T) { diff --git a/sdk/go/openshell/v1/internal/converter/service.go b/sdk/go/openshell/v1/internal/converter/service.go index be50519613..bb2803c6bb 100644 --- a/sdk/go/openshell/v1/internal/converter/service.go +++ b/sdk/go/openshell/v1/internal/converter/service.go @@ -26,6 +26,7 @@ func ServiceEndpointFromProto(resp *pb.ServiceEndpointResponse) *types.ServiceEn result.Name = ep.GetName() result.TargetPort = ep.GetTargetPort() result.Domain = ep.GetDomain() + result.AuthorizationMode = serviceAuthorizationModeFromProto(ep.GetAuthorizationMode()) if m := ep.GetMetadata(); m != nil { result.ID = m.GetId() @@ -48,12 +49,27 @@ func ServiceEndpointToProto(se *types.ServiceEndpoint) *pb.ServiceEndpointRespon Id: se.ID, Workspace: se.Workspace, }, - SandboxId: se.SandboxID, - Sandbox: se.Sandbox, - Name: se.Name, - TargetPort: se.TargetPort, - Domain: se.Domain, + SandboxId: se.SandboxID, + Sandbox: se.Sandbox, + Name: se.Name, + TargetPort: se.TargetPort, + Domain: se.Domain, + AuthorizationMode: serviceAuthorizationModeToProto(se.AuthorizationMode), }, Url: se.URL, } } + +func serviceAuthorizationModeToProto(mode types.ServiceAuthorizationMode) pb.ServiceAuthorizationMode { + if mode == types.ServiceAuthorizationModeBearerPassthrough { + return pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH + } + return pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_STRIP +} + +func serviceAuthorizationModeFromProto(mode pb.ServiceAuthorizationMode) types.ServiceAuthorizationMode { + if mode == pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH { + return types.ServiceAuthorizationModeBearerPassthrough + } + return types.ServiceAuthorizationModeStrip +} diff --git a/sdk/go/openshell/v1/internal/converter/service_test.go b/sdk/go/openshell/v1/internal/converter/service_test.go index ec5fa540a4..6db09701e3 100644 --- a/sdk/go/openshell/v1/internal/converter/service_test.go +++ b/sdk/go/openshell/v1/internal/converter/service_test.go @@ -19,11 +19,12 @@ func TestServiceEndpointFromProto(t *testing.T) { Metadata: &dm.ObjectMeta{ Id: "svc-1", }, - SandboxId: "sb-1", - Sandbox: "my-sandbox", - Name: "http-server", - TargetPort: 8080, - Domain: true, + SandboxId: "sb-1", + Sandbox: "my-sandbox", + Name: "http-server", + TargetPort: 8080, + Domain: true, + AuthorizationMode: pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH, }, Url: "https://svc-1.example.com", } @@ -37,6 +38,7 @@ func TestServiceEndpointFromProto(t *testing.T) { assert.Equal(t, "http-server", se.Name) assert.Equal(t, uint32(8080), se.TargetPort) assert.True(t, se.Domain) + assert.Equal(t, v1.ServiceAuthorizationModeBearerPassthrough, se.AuthorizationMode) assert.Equal(t, "https://svc-1.example.com", se.URL) } @@ -78,13 +80,14 @@ func TestServiceEndpointFromProto_Nil(t *testing.T) { func TestServiceEndpointToProto(t *testing.T) { se := &v1.ServiceEndpoint{ - ID: "svc-1", - SandboxID: "sb-1", - Sandbox: "my-sandbox", - Name: "http-server", - TargetPort: 8080, - Domain: true, - URL: "https://svc-1.example.com", + ID: "svc-1", + SandboxID: "sb-1", + Sandbox: "my-sandbox", + Name: "http-server", + TargetPort: 8080, + Domain: true, + AuthorizationMode: v1.ServiceAuthorizationModeBearerPassthrough, + URL: "https://svc-1.example.com", } resp := ServiceEndpointToProto(se) @@ -98,6 +101,7 @@ func TestServiceEndpointToProto(t *testing.T) { assert.Equal(t, "http-server", resp.Endpoint.Name) assert.Equal(t, uint32(8080), resp.Endpoint.TargetPort) assert.True(t, resp.Endpoint.Domain) + assert.Equal(t, pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH, resp.Endpoint.AuthorizationMode) assert.Equal(t, "https://svc-1.example.com", resp.Url) } @@ -108,13 +112,14 @@ func TestServiceEndpointToProto_Nil(t *testing.T) { func TestServiceEndpointRoundTrip(t *testing.T) { original := &v1.ServiceEndpoint{ - ID: "svc-rt", - SandboxID: "sb-rt", - Sandbox: "round-trip", - Name: "web", - TargetPort: 9090, - Domain: false, - URL: "http://localhost:9090", + ID: "svc-rt", + SandboxID: "sb-rt", + Sandbox: "round-trip", + Name: "web", + TargetPort: 9090, + Domain: false, + AuthorizationMode: v1.ServiceAuthorizationModeBearerPassthrough, + URL: "http://localhost:9090", } proto := ServiceEndpointToProto(original) @@ -127,5 +132,6 @@ func TestServiceEndpointRoundTrip(t *testing.T) { assert.Equal(t, original.Name, back.Name) assert.Equal(t, original.TargetPort, back.TargetPort) assert.Equal(t, original.Domain, back.Domain) + assert.Equal(t, original.AuthorizationMode, back.AuthorizationMode) assert.Equal(t, original.URL, back.URL) } diff --git a/sdk/go/openshell/v1/sandbox_client.go b/sdk/go/openshell/v1/sandbox_client.go index 9868d11b93..f8b8290ca1 100644 --- a/sdk/go/openshell/v1/sandbox_client.go +++ b/sdk/go/openshell/v1/sandbox_client.go @@ -91,13 +91,21 @@ func serviceExposuresToProto(exposures []types.ServiceExposure) []*pb.SandboxSer result := make([]*pb.SandboxServiceExposure, 0, len(exposures)) for _, exposure := range exposures { result = append(result, &pb.SandboxServiceExposure{ - Service: exposure.Service, - TargetPort: exposure.TargetPort, + Service: exposure.Service, + TargetPort: exposure.TargetPort, + AuthorizationMode: serviceAuthorizationModeToProto(exposure.AuthorizationMode), }) } return result } +func serviceAuthorizationModeToProto(mode types.ServiceAuthorizationMode) pb.ServiceAuthorizationMode { + if mode == 0 { + return pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_STRIP + } + return pb.ServiceAuthorizationMode(mode) +} + func validateTemplateCreateSpec(spec *SandboxSpec) error { if spec == nil { return nil diff --git a/sdk/go/openshell/v1/sandbox_client_test.go b/sdk/go/openshell/v1/sandbox_client_test.go index 8ae398f9e0..70c3b6fab2 100644 --- a/sdk/go/openshell/v1/sandbox_client_test.go +++ b/sdk/go/openshell/v1/sandbox_client_test.go @@ -318,7 +318,11 @@ func TestSandboxCreate(t *testing.T) { labels, CreateOptions{ServiceExposures: []ServiceExposure{ {TargetPort: 4500}, - {Service: "metrics", TargetPort: 9090}, + { + Service: "metrics", + TargetPort: 9090, + AuthorizationMode: ServiceAuthorizationModeBearerPassthrough, + }, }}, ) @@ -334,7 +338,9 @@ func TestSandboxCreate(t *testing.T) { }, result.ServiceURLs) require.Len(t, mock.createRequest.GetServiceExposures(), 2) assert.Equal(t, uint32(4500), mock.createRequest.GetServiceExposures()[0].GetTargetPort()) + assert.Equal(t, pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_STRIP, mock.createRequest.GetServiceExposures()[0].GetAuthorizationMode()) assert.Equal(t, "metrics", mock.createRequest.GetServiceExposures()[1].GetService()) + assert.Equal(t, pb.ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH, mock.createRequest.GetServiceExposures()[1].GetAuthorizationMode()) } func TestSandboxCreate_DefaultGPURequest(t *testing.T) { diff --git a/sdk/go/openshell/v1/service.go b/sdk/go/openshell/v1/service.go index 3ef3bea2fc..ccaf93ee4b 100644 --- a/sdk/go/openshell/v1/service.go +++ b/sdk/go/openshell/v1/service.go @@ -15,9 +15,24 @@ type ServiceEndpoint = types.ServiceEndpoint // ServiceExposure describes a loopback HTTP service to expose during sandbox creation. type ServiceExposure = types.ServiceExposure +// ServiceAuthorizationMode controls handling of an incoming application Authorization header. +type ServiceAuthorizationMode = types.ServiceAuthorizationMode + +const ( + // ServiceAuthorizationModeStrip removes Authorization before proxying to the service. + ServiceAuthorizationModeStrip = types.ServiceAuthorizationModeStrip + // ServiceAuthorizationModeBearerPassthrough forwards one valid bearer credential unchanged. + ServiceAuthorizationModeBearerPassthrough = types.ServiceAuthorizationModeBearerPassthrough +) + +// ExposeServiceOptions configures service exposure behavior. +type ExposeServiceOptions struct { + AuthorizationMode ServiceAuthorizationMode +} + // ServiceInterface defines operations for managing sandbox service endpoints. type ServiceInterface interface { - Expose(ctx context.Context, workspace, sandboxName, serviceName string, targetPort uint32, domain bool) (*ServiceEndpoint, error) + Expose(ctx context.Context, workspace, sandboxName, serviceName string, targetPort uint32, domain bool, opts ...ExposeServiceOptions) (*ServiceEndpoint, error) Get(ctx context.Context, workspace, sandboxName, serviceName string) (*ServiceEndpoint, error) List(workspace, sandboxName string, opts ...ListOptions) (*Pager[*ServiceEndpoint], error) ListAll(ctx context.Context, workspace, sandboxName string, opts ...ListOptions) ([]*ServiceEndpoint, error) diff --git a/sdk/go/openshell/v1/service_client.go b/sdk/go/openshell/v1/service_client.go index 80eb2bbab7..5606ebf2a0 100644 --- a/sdk/go/openshell/v1/service_client.go +++ b/sdk/go/openshell/v1/service_client.go @@ -19,13 +19,18 @@ func newServiceClient(conn grpc.ClientConnInterface) *serviceClient { return &serviceClient{client: pb.NewOpenShellClient(conn)} } -func (s *serviceClient) Expose(ctx context.Context, workspace, sandboxName, serviceName string, targetPort uint32, domain bool) (*ServiceEndpoint, error) { +func (s *serviceClient) Expose(ctx context.Context, workspace, sandboxName, serviceName string, targetPort uint32, domain bool, opts ...ExposeServiceOptions) (*ServiceEndpoint, error) { + authorizationMode := ServiceAuthorizationModeStrip + if len(opts) > 0 && opts[0].AuthorizationMode != 0 { + authorizationMode = opts[0].AuthorizationMode + } resp, err := s.client.ExposeService(ctx, &pb.ExposeServiceRequest{ - Sandbox: sandboxName, - WorkspaceScope: namedWorkspaceScope(workspace), - Name: serviceName, - TargetPort: targetPort, - Domain: domain, + Sandbox: sandboxName, + WorkspaceScope: namedWorkspaceScope(workspace), + Name: serviceName, + TargetPort: targetPort, + Domain: domain, + AuthorizationMode: pb.ServiceAuthorizationMode(authorizationMode), }) if err != nil { return nil, converter.FromGRPCError(err) diff --git a/sdk/go/openshell/v1/service_client_test.go b/sdk/go/openshell/v1/service_client_test.go index 9b702dbde6..3ba2e5e69a 100644 --- a/sdk/go/openshell/v1/service_client_test.go +++ b/sdk/go/openshell/v1/service_client_test.go @@ -55,10 +55,11 @@ func (s *mockServiceServer) ExposeService(_ context.Context, req *pb.ExposeServi Metadata: &dm.ObjectMeta{ Id: "ep-" + req.GetName(), }, - Sandbox: req.GetSandbox(), - Name: req.GetName(), - TargetPort: req.GetTargetPort(), - Domain: req.GetDomain(), + Sandbox: req.GetSandbox(), + Name: req.GetName(), + TargetPort: req.GetTargetPort(), + Domain: req.GetDomain(), + AuthorizationMode: req.GetAuthorizationMode(), }, } if req.GetDomain() { @@ -150,7 +151,15 @@ func TestServiceExpose(t *testing.T) { client, cleanup := setupServiceTest(t, mock) defer cleanup() - ep, err := client.Expose(context.Background(), "default", "web-app", "api", 8080, true) + ep, err := client.Expose( + context.Background(), + "default", + "web-app", + "api", + 8080, + true, + ExposeServiceOptions{AuthorizationMode: ServiceAuthorizationModeBearerPassthrough}, + ) require.NoError(t, err) require.NotNil(t, ep) @@ -159,6 +168,7 @@ func TestServiceExpose(t *testing.T) { assert.Equal(t, "api", ep.Name) assert.Equal(t, uint32(8080), ep.TargetPort) assert.True(t, ep.Domain) + assert.Equal(t, ServiceAuthorizationModeBearerPassthrough, ep.AuthorizationMode) assert.Equal(t, "https://api.example.com", ep.URL) } diff --git a/sdk/go/openshell/v1/types.go b/sdk/go/openshell/v1/types.go index 6226aa0707..e565d966cb 100644 --- a/sdk/go/openshell/v1/types.go +++ b/sdk/go/openshell/v1/types.go @@ -23,6 +23,16 @@ const ( SandboxCompleted = types.SandboxCompleted ) +// SandboxRestartPolicy controls replacement after the canonical main process exits. +type SandboxRestartPolicy = types.SandboxRestartPolicy + +// Sandbox restart policy values. +const ( + SandboxRestartNever = types.SandboxRestartNever + SandboxRestartOnFailure = types.SandboxRestartOnFailure + SandboxRestartAlways = types.SandboxRestartAlways +) + // EventType classifies watch events. type EventType = types.EventType diff --git a/sdk/go/openshell/v1/types/profile.go b/sdk/go/openshell/v1/types/profile.go index b7834d2ab0..fdffd2bfe6 100644 --- a/sdk/go/openshell/v1/types/profile.go +++ b/sdk/go/openshell/v1/types/profile.go @@ -18,13 +18,15 @@ const ( ) // ProviderProfile defines a provider type template with credentials schema, -// endpoints, binaries, and discovery configuration. +// files, endpoints, binaries, and discovery configuration. type ProviderProfile struct { - ID string - DisplayName string - Description string - Category ProfileCategory - Credentials []ProfileCredential + ID string + DisplayName string + Description string + Category ProfileCategory + Credentials []ProfileCredential + // Files is EXPERIMENTAL. This API and its behavior may change or be removed. + Files []ProfileFile Endpoints []NetworkEndpoint Binaries []NetworkBinary InferenceCapable bool @@ -35,6 +37,14 @@ type ProviderProfile struct { Scope string } +// ProfileFile declares non-secret content served at a virtual sandbox path. +// EXPERIMENTAL: This API and its behavior may change or be removed. +type ProfileFile struct { + Path string + Content string + EnvVar string +} + // ProfileCredential defines a single credential required by a provider profile. type ProfileCredential struct { Name string diff --git a/sdk/go/openshell/v1/types/sandbox.go b/sdk/go/openshell/v1/types/sandbox.go index 1218bb6224..c8e7d43ef3 100644 --- a/sdk/go/openshell/v1/types/sandbox.go +++ b/sdk/go/openshell/v1/types/sandbox.go @@ -34,9 +34,10 @@ type SandboxSpec struct { GPU bool GPUCount *uint32 // Policy is the security policy for the sandbox. Nil means no policy specified. - Policy *SandboxPolicy - Command []string - TTY bool + Policy *SandboxPolicy + Command []string + TTY bool + RestartPolicy SandboxRestartPolicy } // SandboxTemplate defines the container template for a sandbox. @@ -123,6 +124,9 @@ type SandboxStatus struct { // last accepted network results, independently of sandbox readiness. EndpointStatuses []EndpointStatus ConfigurationAdmission *SandboxConfigurationAdmission + RestartCount uint32 + NextRestartAtMs int64 + MainProcessStartedAtMs int64 } // ConfigurationAdmissionState describes validation of an effective configuration. diff --git a/sdk/go/openshell/v1/types/service.go b/sdk/go/openshell/v1/types/service.go index 9cccff8258..5ae7aa0b9e 100644 --- a/sdk/go/openshell/v1/types/service.go +++ b/sdk/go/openshell/v1/types/service.go @@ -3,20 +3,32 @@ package types +// ServiceAuthorizationMode controls handling of an incoming application Authorization header. +type ServiceAuthorizationMode int32 + +const ( + // ServiceAuthorizationModeStrip removes Authorization before proxying to the sandbox service. + ServiceAuthorizationModeStrip ServiceAuthorizationMode = 1 + // ServiceAuthorizationModeBearerPassthrough forwards one valid bearer credential unchanged. + ServiceAuthorizationModeBearerPassthrough ServiceAuthorizationMode = 2 +) + // ServiceExposure describes a loopback HTTP service to expose during sandbox creation. type ServiceExposure struct { - Service string - TargetPort uint32 + Service string + TargetPort uint32 + AuthorizationMode ServiceAuthorizationMode } // ServiceEndpoint represents an exposed HTTP service on a sandbox. type ServiceEndpoint struct { - ID string - SandboxID string - Sandbox string - Name string - TargetPort uint32 - Domain bool - URL string - Workspace string + ID string + SandboxID string + Sandbox string + Name string + TargetPort uint32 + Domain bool + URL string + Workspace string + AuthorizationMode ServiceAuthorizationMode } diff --git a/sdk/go/openshell/v1/types/types.go b/sdk/go/openshell/v1/types/types.go index 5b32ec3685..99d8e24bf9 100644 --- a/sdk/go/openshell/v1/types/types.go +++ b/sdk/go/openshell/v1/types/types.go @@ -21,6 +21,16 @@ const ( SandboxCompleted SandboxPhase = "Completed" ) +// SandboxRestartPolicy controls replacement after the canonical main process exits. +type SandboxRestartPolicy string + +// Sandbox restart policy values. +const ( + SandboxRestartNever SandboxRestartPolicy = "Never" + SandboxRestartOnFailure SandboxRestartPolicy = "OnFailure" + SandboxRestartAlways SandboxRestartPolicy = "Always" +) + // EventType classifies watch events. type EventType string diff --git a/sdk/go/proto/openshellv1/openshell.pb.go b/sdk/go/proto/openshellv1/openshell.pb.go index 18d0b9bcde..6c9b4e144c 100644 --- a/sdk/go/proto/openshellv1/openshell.pb.go +++ b/sdk/go/proto/openshellv1/openshell.pb.go @@ -1005,6 +1005,61 @@ func (ProviderCredentialRefreshRecoveryAction) EnumDescriptor() ([]byte, []int) return file_openshell_proto_rawDescGZIP(), []int{15} } +// Policy applied when the canonical main process exits outside an intentional +// stop or delete operation. Kept after existing enums so generated enum +// descriptors remain stable. +type SandboxRestartPolicy int32 + +const ( + SandboxRestartPolicy_SANDBOX_RESTART_POLICY_UNSPECIFIED SandboxRestartPolicy = 0 + SandboxRestartPolicy_SANDBOX_RESTART_POLICY_NEVER SandboxRestartPolicy = 1 + SandboxRestartPolicy_SANDBOX_RESTART_POLICY_ON_FAILURE SandboxRestartPolicy = 2 + SandboxRestartPolicy_SANDBOX_RESTART_POLICY_ALWAYS SandboxRestartPolicy = 3 +) + +// Enum value maps for SandboxRestartPolicy. +var ( + SandboxRestartPolicy_name = map[int32]string{ + 0: "SANDBOX_RESTART_POLICY_UNSPECIFIED", + 1: "SANDBOX_RESTART_POLICY_NEVER", + 2: "SANDBOX_RESTART_POLICY_ON_FAILURE", + 3: "SANDBOX_RESTART_POLICY_ALWAYS", + } + SandboxRestartPolicy_value = map[string]int32{ + "SANDBOX_RESTART_POLICY_UNSPECIFIED": 0, + "SANDBOX_RESTART_POLICY_NEVER": 1, + "SANDBOX_RESTART_POLICY_ON_FAILURE": 2, + "SANDBOX_RESTART_POLICY_ALWAYS": 3, + } +) + +func (x SandboxRestartPolicy) Enum() *SandboxRestartPolicy { + p := new(SandboxRestartPolicy) + *p = x + return p +} + +func (x SandboxRestartPolicy) String() string { + return protoimpl.X.EnumStringOf(x.Descriptor(), protoreflect.EnumNumber(x)) +} + +func (SandboxRestartPolicy) Descriptor() protoreflect.EnumDescriptor { + return file_openshell_proto_enumTypes[16].Descriptor() +} + +func (SandboxRestartPolicy) Type() protoreflect.EnumType { + return &file_openshell_proto_enumTypes[16] +} + +func (x SandboxRestartPolicy) Number() protoreflect.EnumNumber { + return protoreflect.EnumNumber(x) +} + +// Deprecated: Use SandboxRestartPolicy.Descriptor instead. +func (SandboxRestartPolicy) EnumDescriptor() ([]byte, []int) { + return file_openshell_proto_rawDescGZIP(), []int{16} +} + // Result of a public delete, membership removal, or session revocation. // Default requests return NOT_FOUND for a missing target. With allow_missing, // only a missing target becomes ALREADY_ABSENT; parent lookup, authorization, @@ -1052,11 +1107,11 @@ func (x DeletionOutcome) String() string { } func (DeletionOutcome) Descriptor() protoreflect.EnumDescriptor { - return file_openshell_proto_enumTypes[16].Descriptor() + return file_openshell_proto_enumTypes[17].Descriptor() } func (DeletionOutcome) Type() protoreflect.EnumType { - return &file_openshell_proto_enumTypes[16] + return &file_openshell_proto_enumTypes[17] } func (x DeletionOutcome) Number() protoreflect.EnumNumber { @@ -1065,7 +1120,7 @@ func (x DeletionOutcome) Number() protoreflect.EnumNumber { // Deprecated: Use DeletionOutcome.Descriptor instead. func (DeletionOutcome) EnumDescriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{16} + return file_openshell_proto_rawDescGZIP(), []int{17} } // Last observed network result for a configured external tool endpoint. @@ -1126,11 +1181,11 @@ func (x EndpointResult) String() string { } func (EndpointResult) Descriptor() protoreflect.EnumDescriptor { - return file_openshell_proto_enumTypes[17].Descriptor() + return file_openshell_proto_enumTypes[18].Descriptor() } func (EndpointResult) Type() protoreflect.EnumType { - return &file_openshell_proto_enumTypes[17] + return &file_openshell_proto_enumTypes[18] } func (x EndpointResult) Number() protoreflect.EnumNumber { @@ -1139,7 +1194,61 @@ func (x EndpointResult) Number() protoreflect.EnumNumber { // Deprecated: Use EndpointResult.Descriptor instead. func (EndpointResult) EnumDescriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{17} + return file_openshell_proto_rawDescGZIP(), []int{18} +} + +// Controls whether an exposed service receives the request's application +// Authorization header. +type ServiceAuthorizationMode int32 + +const ( + // Omission preserves the secure legacy behavior and resolves to STRIP. + ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_UNSPECIFIED ServiceAuthorizationMode = 0 + // Remove Authorization before forwarding to the sandbox service. + ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_STRIP ServiceAuthorizationMode = 1 + // Forward one syntactically valid bearer Authorization header unchanged. + ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH ServiceAuthorizationMode = 2 +) + +// Enum value maps for ServiceAuthorizationMode. +var ( + ServiceAuthorizationMode_name = map[int32]string{ + 0: "SERVICE_AUTHORIZATION_MODE_UNSPECIFIED", + 1: "SERVICE_AUTHORIZATION_MODE_STRIP", + 2: "SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH", + } + ServiceAuthorizationMode_value = map[string]int32{ + "SERVICE_AUTHORIZATION_MODE_UNSPECIFIED": 0, + "SERVICE_AUTHORIZATION_MODE_STRIP": 1, + "SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH": 2, + } +) + +func (x ServiceAuthorizationMode) Enum() *ServiceAuthorizationMode { + p := new(ServiceAuthorizationMode) + *p = x + return p +} + +func (x ServiceAuthorizationMode) String() string { + return protoimpl.X.EnumStringOf(x.Descriptor(), protoreflect.EnumNumber(x)) +} + +func (ServiceAuthorizationMode) Descriptor() protoreflect.EnumDescriptor { + return file_openshell_proto_enumTypes[19].Descriptor() +} + +func (ServiceAuthorizationMode) Type() protoreflect.EnumType { + return &file_openshell_proto_enumTypes[19] +} + +func (x ServiceAuthorizationMode) Number() protoreflect.EnumNumber { + return protoreflect.EnumNumber(x) +} + +// Deprecated: Use ServiceAuthorizationMode.Descriptor instead. +func (ServiceAuthorizationMode) EnumDescriptor() ([]byte, []int) { + return file_openshell_proto_rawDescGZIP(), []int{19} } // IssueSandboxToken request. Empty body; identity is established by the @@ -2250,8 +2359,11 @@ type SandboxSpec struct { // Gateway-owned attachment identity, changed atomically with the provider set. // Equality only: detach and reattach must not revive an older receipt. ProviderAttachmentEpoch string `protobuf:"bytes,14,opt,name=provider_attachment_epoch,json=providerAttachmentEpoch,proto3" json:"provider_attachment_epoch,omitempty"` - unknownFields protoimpl.UnknownFields - sizeCache protoimpl.SizeCache + // Gateway-owned policy for replacing the sandbox runtime after the canonical + // main process exits. Unspecified is normalized to Never before persistence. + RestartPolicy SandboxRestartPolicy `protobuf:"varint,15,opt,name=restart_policy,json=restartPolicy,proto3,enum=openshell.v1.SandboxRestartPolicy" json:"restart_policy,omitempty"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache } func (x *SandboxSpec) Reset() { @@ -2347,6 +2459,13 @@ func (x *SandboxSpec) GetProviderAttachmentEpoch() string { return "" } +func (x *SandboxSpec) GetRestartPolicy() SandboxRestartPolicy { + if x != nil { + return x.RestartPolicy + } + return SandboxRestartPolicy_SANDBOX_RESTART_POLICY_UNSPECIFIED +} + type ResourceRequirements struct { state protoimpl.MessageState `protogen:"open.v1"` // GPU requirements for the sandbox. Presence indicates a GPU request. @@ -2988,9 +3107,8 @@ type SandboxStatus struct { // Supervisor instance currently associated with the canonical main process. // The gateway uses this to reject stale exit reports after a restart. MainProcessInstanceId string `protobuf:"bytes,8,opt,name=main_process_instance_id,json=mainProcessInstanceId,proto3" json:"main_process_instance_id,omitempty"` - // Normalized main process result. Signal exits use 128 + signal number. - // Presence indicates that the canonical main process exited. Exit code 0 - // produces Completed; nonzero and signal-normalized exits produce Error. + // Most recent normalized main process result. Signal exits use 128 + signal number. + // Cleared when a replacement main process becomes ready. ExitCode *int32 `protobuf:"varint,9,opt,name=exit_code,json=exitCode,proto3,oneof" json:"exit_code,omitempty"` // Last accepted network result for each configured tool server endpoint. // Currently populated for MCP-over-HTTP endpoints. These passive results @@ -3001,9 +3119,15 @@ type SandboxStatus struct { // Durable first-acceptance marker. Absent on legacy records; never reset by restart. ConfigurationActivated *bool `protobuf:"varint,12,opt,name=configuration_activated,json=configurationActivated,proto3,oneof" json:"configuration_activated,omitempty"` // Gateway-owned repair window. Retained after timeout for inspection and retry. - Provisioning *SandboxProvisioning `protobuf:"bytes,13,opt,name=provisioning,proto3" json:"provisioning,omitempty"` - unknownFields protoimpl.UnknownFields - sizeCache protoimpl.SizeCache + Provisioning *SandboxProvisioning `protobuf:"bytes,13,opt,name=provisioning,proto3" json:"provisioning,omitempty"` + // Consecutive policy-driven restart number in the current crash loop. + RestartCount uint32 `protobuf:"varint,14,opt,name=restart_count,json=restartCount,proto3" json:"restart_count,omitempty"` + // Deadline for the next restart attempt or recovery check. + NextRestartTime *timestamppb.Timestamp `protobuf:"bytes,15,opt,name=next_restart_time,json=nextRestartTime,proto3" json:"next_restart_time,omitempty"` + // Time when the current main process became ready. + MainProcessStartedTime *timestamppb.Timestamp `protobuf:"bytes,16,opt,name=main_process_started_time,json=mainProcessStartedTime,proto3" json:"main_process_started_time,omitempty"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache } func (x *SandboxStatus) Reset() { @@ -3120,6 +3244,27 @@ func (x *SandboxStatus) GetProvisioning() *SandboxProvisioning { return nil } +func (x *SandboxStatus) GetRestartCount() uint32 { + if x != nil { + return x.RestartCount + } + return 0 +} + +func (x *SandboxStatus) GetNextRestartTime() *timestamppb.Timestamp { + if x != nil { + return x.NextRestartTime + } + return nil +} + +func (x *SandboxStatus) GetMainProcessStartedTime() *timestamppb.Timestamp { + if x != nil { + return x.MainProcessStartedTime + } + return nil +} + // User-facing sandbox condition derived from platform or gateway observations. type SandboxCondition struct { state protoimpl.MessageState `protogen:"open.v1"` @@ -6016,9 +6161,11 @@ type ExposeServiceRequest struct { Sandbox string `protobuf:"bytes,1,opt,name=sandbox,proto3" json:"sandbox,omitempty"` // Optional nonzero UUID for durable at-most-once admission. Successful results // can be replayed for 24 hours; see the API errors and retries reference. - RequestId string `protobuf:"bytes,6,opt,name=request_id,json=requestId,proto3" json:"request_id,omitempty"` - unknownFields protoimpl.UnknownFields - sizeCache protoimpl.SizeCache + RequestId string `protobuf:"bytes,6,opt,name=request_id,json=requestId,proto3" json:"request_id,omitempty"` + // Application authorization behavior. Omission resolves to STRIP. + AuthorizationMode ServiceAuthorizationMode `protobuf:"varint,7,opt,name=authorization_mode,json=authorizationMode,proto3,enum=openshell.v1.ServiceAuthorizationMode" json:"authorization_mode,omitempty"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache } func (x *ExposeServiceRequest) Reset() { @@ -6093,6 +6240,13 @@ func (x *ExposeServiceRequest) GetRequestId() string { return "" } +func (x *ExposeServiceRequest) GetAuthorizationMode() ServiceAuthorizationMode { + if x != nil { + return x.AuthorizationMode + } + return ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_UNSPECIFIED +} + // Request to fetch an exposed sandbox service endpoint. type GetServiceRequest struct { state protoimpl.MessageState `protogen:"open.v1"` @@ -6425,9 +6579,11 @@ type ServiceEndpoint struct { // Loopback TCP port inside the sandbox. TargetPort uint32 `protobuf:"varint,5,opt,name=target_port,json=targetPort,proto3" json:"target_port,omitempty"` // Whether browser-facing service routing is enabled for this endpoint. - Domain bool `protobuf:"varint,6,opt,name=domain,proto3" json:"domain,omitempty"` - unknownFields protoimpl.UnknownFields - sizeCache protoimpl.SizeCache + Domain bool `protobuf:"varint,6,opt,name=domain,proto3" json:"domain,omitempty"` + // Effective application authorization behavior for ingress requests. + AuthorizationMode ServiceAuthorizationMode `protobuf:"varint,7,opt,name=authorization_mode,json=authorizationMode,proto3,enum=openshell.v1.ServiceAuthorizationMode" json:"authorization_mode,omitempty"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache } func (x *ServiceEndpoint) Reset() { @@ -6502,6 +6658,13 @@ func (x *ServiceEndpoint) GetDomain() bool { return false } +func (x *ServiceEndpoint) GetAuthorizationMode() ServiceAuthorizationMode { + if x != nil { + return x.AuthorizationMode + } + return ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_UNSPECIFIED +} + // Response containing a service endpoint and, when available, its local URL. type ServiceEndpointResponse struct { state protoimpl.MessageState `protogen:"open.v1"` @@ -7496,8 +7659,13 @@ type WatchSandboxRequest struct { // Stream platform events correlated to this sandbox. FollowEvents bool `protobuf:"varint,4,opt,name=follow_events,json=followEvents,proto3" json:"follow_events,omitempty"` // Replay the last N log lines (best-effort) before following. + // + // When following both sources, the replay can return fewer lines: events + // older than one the other source's replay left out are withheld and + // reported with a warning, even if both depths are equal. LogTailLines uint32 `protobuf:"varint,5,opt,name=log_tail_lines,json=logTailLines,proto3" json:"log_tail_lines,omitempty"` // Replay the last N platform events (best-effort) before following. + // Defaults to 0. Can return fewer events; see log_tail_lines. EventTail uint32 `protobuf:"varint,6,opt,name=event_tail,json=eventTail,proto3" json:"event_tail,omitempty"` // Stop streaming once the sandbox reaches READY or a terminal result phase // (COMPLETED, STOPPED, or ERROR). @@ -7526,6 +7694,8 @@ type WatchSandboxRequest struct { // empty resume_after_cursor, because retrying the same token fails // identically. A cursor this server could not have issued is rejected with // INVALID_ARGUMENT. + // + // One cursor covers both log and platform events. ResumeAfterCursor string `protobuf:"bytes,12,opt,name=resume_after_cursor,json=resumeAfterCursor,proto3" json:"resume_after_cursor,omitempty"` unknownFields protoimpl.UnknownFields sizeCache protoimpl.SizeCache @@ -9975,7 +10145,10 @@ type ProviderProfile struct { Source string `protobuf:"bytes,12,opt,name=source,proto3" json:"source,omitempty"` // Server-set visibility: "platform", "workspace", or empty for // non-scoped sources. Ignored on import/update payloads. - Scope string `protobuf:"bytes,13,opt,name=scope,proto3" json:"scope,omitempty"` + Scope string `protobuf:"bytes,13,opt,name=scope,proto3" json:"scope,omitempty"` + // EXPERIMENTAL: Non-secret files rendered from provider configuration for the + // workload. This API and its behavior may change or be removed. + Files []*ProviderProfileFile `protobuf:"bytes,14,rep,name=files,proto3" json:"files,omitempty"` unknownFields protoimpl.UnknownFields sizeCache protoimpl.SizeCache } @@ -10101,6 +10274,77 @@ func (x *ProviderProfile) GetScope() string { return "" } +func (x *ProviderProfile) GetFiles() []*ProviderProfileFile { + if x != nil { + return x.Files + } + return nil +} + +// EXPERIMENTAL: Provider file templates may change or be removed. +type ProviderProfileFile struct { + state protoimpl.MessageState `protogen:"open.v1"` + // One virtual file name below /run/openshell/providers//. + Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"` + // UTF-8 template. Only {{config.KEY}} references are supported. + Content string `protobuf:"bytes,2,opt,name=content,proto3" json:"content,omitempty"` + // Optional environment variable containing the virtual absolute path. + EnvVar string `protobuf:"bytes,3,opt,name=env_var,json=envVar,proto3" json:"env_var,omitempty"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache +} + +func (x *ProviderProfileFile) Reset() { + *x = ProviderProfileFile{} + mi := &file_openshell_proto_msgTypes[122] + ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) + ms.StoreMessageInfo(mi) +} + +func (x *ProviderProfileFile) String() string { + return protoimpl.X.MessageStringOf(x) +} + +func (*ProviderProfileFile) ProtoMessage() {} + +func (x *ProviderProfileFile) ProtoReflect() protoreflect.Message { + mi := &file_openshell_proto_msgTypes[122] + if x != nil { + ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) + if ms.LoadMessageInfo() == nil { + ms.StoreMessageInfo(mi) + } + return ms + } + return mi.MessageOf(x) +} + +// Deprecated: Use ProviderProfileFile.ProtoReflect.Descriptor instead. +func (*ProviderProfileFile) Descriptor() ([]byte, []int) { + return file_openshell_proto_rawDescGZIP(), []int{122} +} + +func (x *ProviderProfileFile) GetPath() string { + if x != nil { + return x.Path + } + return "" +} + +func (x *ProviderProfileFile) GetContent() string { + if x != nil { + return x.Content + } + return "" +} + +func (x *ProviderProfileFile) GetEnvVar() string { + if x != nil { + return x.EnvVar + } + return "" +} + // Provider profile response. type ProviderProfileResponse struct { state protoimpl.MessageState `protogen:"open.v1"` @@ -10111,7 +10355,7 @@ type ProviderProfileResponse struct { func (x *ProviderProfileResponse) Reset() { *x = ProviderProfileResponse{} - mi := &file_openshell_proto_msgTypes[122] + mi := &file_openshell_proto_msgTypes[123] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10123,7 +10367,7 @@ func (x *ProviderProfileResponse) String() string { func (*ProviderProfileResponse) ProtoMessage() {} func (x *ProviderProfileResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[122] + mi := &file_openshell_proto_msgTypes[123] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10136,7 +10380,7 @@ func (x *ProviderProfileResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ProviderProfileResponse.ProtoReflect.Descriptor instead. func (*ProviderProfileResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{122} + return file_openshell_proto_rawDescGZIP(), []int{123} } func (x *ProviderProfileResponse) GetProfile() *ProviderProfile { @@ -10158,7 +10402,7 @@ type ListProviderProfilesResponse struct { func (x *ListProviderProfilesResponse) Reset() { *x = ListProviderProfilesResponse{} - mi := &file_openshell_proto_msgTypes[123] + mi := &file_openshell_proto_msgTypes[124] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10170,7 +10414,7 @@ func (x *ListProviderProfilesResponse) String() string { func (*ListProviderProfilesResponse) ProtoMessage() {} func (x *ListProviderProfilesResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[123] + mi := &file_openshell_proto_msgTypes[124] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10183,7 +10427,7 @@ func (x *ListProviderProfilesResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ListProviderProfilesResponse.ProtoReflect.Descriptor instead. func (*ListProviderProfilesResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{123} + return file_openshell_proto_rawDescGZIP(), []int{124} } func (x *ListProviderProfilesResponse) GetProfiles() []*ProviderProfile { @@ -10215,7 +10459,7 @@ type ImportProviderProfilesRequest struct { func (x *ImportProviderProfilesRequest) Reset() { *x = ImportProviderProfilesRequest{} - mi := &file_openshell_proto_msgTypes[124] + mi := &file_openshell_proto_msgTypes[125] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10227,7 +10471,7 @@ func (x *ImportProviderProfilesRequest) String() string { func (*ImportProviderProfilesRequest) ProtoMessage() {} func (x *ImportProviderProfilesRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[124] + mi := &file_openshell_proto_msgTypes[125] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10240,7 +10484,7 @@ func (x *ImportProviderProfilesRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ImportProviderProfilesRequest.ProtoReflect.Descriptor instead. func (*ImportProviderProfilesRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{124} + return file_openshell_proto_rawDescGZIP(), []int{125} } func (x *ImportProviderProfilesRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -10276,7 +10520,7 @@ type ImportProviderProfilesResponse struct { func (x *ImportProviderProfilesResponse) Reset() { *x = ImportProviderProfilesResponse{} - mi := &file_openshell_proto_msgTypes[125] + mi := &file_openshell_proto_msgTypes[126] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10288,7 +10532,7 @@ func (x *ImportProviderProfilesResponse) String() string { func (*ImportProviderProfilesResponse) ProtoMessage() {} func (x *ImportProviderProfilesResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[125] + mi := &file_openshell_proto_msgTypes[126] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10301,7 +10545,7 @@ func (x *ImportProviderProfilesResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ImportProviderProfilesResponse.ProtoReflect.Descriptor instead. func (*ImportProviderProfilesResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{125} + return file_openshell_proto_rawDescGZIP(), []int{126} } func (x *ImportProviderProfilesResponse) GetDiagnostics() []*ProviderProfileDiagnostic { @@ -10347,7 +10591,7 @@ type UpdateProviderProfilesRequest struct { func (x *UpdateProviderProfilesRequest) Reset() { *x = UpdateProviderProfilesRequest{} - mi := &file_openshell_proto_msgTypes[126] + mi := &file_openshell_proto_msgTypes[127] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10359,7 +10603,7 @@ func (x *UpdateProviderProfilesRequest) String() string { func (*UpdateProviderProfilesRequest) ProtoMessage() {} func (x *UpdateProviderProfilesRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[126] + mi := &file_openshell_proto_msgTypes[127] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10372,7 +10616,7 @@ func (x *UpdateProviderProfilesRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use UpdateProviderProfilesRequest.ProtoReflect.Descriptor instead. func (*UpdateProviderProfilesRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{126} + return file_openshell_proto_rawDescGZIP(), []int{127} } func (x *UpdateProviderProfilesRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -10422,7 +10666,7 @@ type UpdateProviderProfilesResponse struct { func (x *UpdateProviderProfilesResponse) Reset() { *x = UpdateProviderProfilesResponse{} - mi := &file_openshell_proto_msgTypes[127] + mi := &file_openshell_proto_msgTypes[128] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10434,7 +10678,7 @@ func (x *UpdateProviderProfilesResponse) String() string { func (*UpdateProviderProfilesResponse) ProtoMessage() {} func (x *UpdateProviderProfilesResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[127] + mi := &file_openshell_proto_msgTypes[128] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10447,7 +10691,7 @@ func (x *UpdateProviderProfilesResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use UpdateProviderProfilesResponse.ProtoReflect.Descriptor instead. func (*UpdateProviderProfilesResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{127} + return file_openshell_proto_rawDescGZIP(), []int{128} } func (x *UpdateProviderProfilesResponse) GetDiagnostics() []*ProviderProfileDiagnostic { @@ -10483,7 +10727,7 @@ type LintProviderProfilesRequest struct { func (x *LintProviderProfilesRequest) Reset() { *x = LintProviderProfilesRequest{} - mi := &file_openshell_proto_msgTypes[128] + mi := &file_openshell_proto_msgTypes[129] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10495,7 +10739,7 @@ func (x *LintProviderProfilesRequest) String() string { func (*LintProviderProfilesRequest) ProtoMessage() {} func (x *LintProviderProfilesRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[128] + mi := &file_openshell_proto_msgTypes[129] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10508,7 +10752,7 @@ func (x *LintProviderProfilesRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use LintProviderProfilesRequest.ProtoReflect.Descriptor instead. func (*LintProviderProfilesRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{128} + return file_openshell_proto_rawDescGZIP(), []int{129} } func (x *LintProviderProfilesRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -10536,7 +10780,7 @@ type LintProviderProfilesResponse struct { func (x *LintProviderProfilesResponse) Reset() { *x = LintProviderProfilesResponse{} - mi := &file_openshell_proto_msgTypes[129] + mi := &file_openshell_proto_msgTypes[130] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10548,7 +10792,7 @@ func (x *LintProviderProfilesResponse) String() string { func (*LintProviderProfilesResponse) ProtoMessage() {} func (x *LintProviderProfilesResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[129] + mi := &file_openshell_proto_msgTypes[130] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10561,7 +10805,7 @@ func (x *LintProviderProfilesResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use LintProviderProfilesResponse.ProtoReflect.Descriptor instead. func (*LintProviderProfilesResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{129} + return file_openshell_proto_rawDescGZIP(), []int{130} } func (x *LintProviderProfilesResponse) GetDiagnostics() []*ProviderProfileDiagnostic { @@ -10588,7 +10832,7 @@ type DeleteProviderResponse struct { func (x *DeleteProviderResponse) Reset() { *x = DeleteProviderResponse{} - mi := &file_openshell_proto_msgTypes[130] + mi := &file_openshell_proto_msgTypes[131] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10600,7 +10844,7 @@ func (x *DeleteProviderResponse) String() string { func (*DeleteProviderResponse) ProtoMessage() {} func (x *DeleteProviderResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[130] + mi := &file_openshell_proto_msgTypes[131] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10613,7 +10857,7 @@ func (x *DeleteProviderResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use DeleteProviderResponse.ProtoReflect.Descriptor instead. func (*DeleteProviderResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{130} + return file_openshell_proto_rawDescGZIP(), []int{131} } func (x *DeleteProviderResponse) GetOutcome() DeletionOutcome { @@ -10639,7 +10883,7 @@ type DeleteProviderProfileRequest struct { func (x *DeleteProviderProfileRequest) Reset() { *x = DeleteProviderProfileRequest{} - mi := &file_openshell_proto_msgTypes[131] + mi := &file_openshell_proto_msgTypes[132] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10651,7 +10895,7 @@ func (x *DeleteProviderProfileRequest) String() string { func (*DeleteProviderProfileRequest) ProtoMessage() {} func (x *DeleteProviderProfileRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[131] + mi := &file_openshell_proto_msgTypes[132] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10664,7 +10908,7 @@ func (x *DeleteProviderProfileRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use DeleteProviderProfileRequest.ProtoReflect.Descriptor instead. func (*DeleteProviderProfileRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{131} + return file_openshell_proto_rawDescGZIP(), []int{132} } func (x *DeleteProviderProfileRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -10705,7 +10949,7 @@ type DeleteProviderProfileResponse struct { func (x *DeleteProviderProfileResponse) Reset() { *x = DeleteProviderProfileResponse{} - mi := &file_openshell_proto_msgTypes[132] + mi := &file_openshell_proto_msgTypes[133] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10717,7 +10961,7 @@ func (x *DeleteProviderProfileResponse) String() string { func (*DeleteProviderProfileResponse) ProtoMessage() {} func (x *DeleteProviderProfileResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[132] + mi := &file_openshell_proto_msgTypes[133] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10730,7 +10974,7 @@ func (x *DeleteProviderProfileResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use DeleteProviderProfileResponse.ProtoReflect.Descriptor instead. func (*DeleteProviderProfileResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{132} + return file_openshell_proto_rawDescGZIP(), []int{133} } func (x *DeleteProviderProfileResponse) GetOutcome() DeletionOutcome { @@ -10755,7 +10999,7 @@ type GetSandboxProviderEnvironmentRequest struct { func (x *GetSandboxProviderEnvironmentRequest) Reset() { *x = GetSandboxProviderEnvironmentRequest{} - mi := &file_openshell_proto_msgTypes[133] + mi := &file_openshell_proto_msgTypes[134] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10767,7 +11011,7 @@ func (x *GetSandboxProviderEnvironmentRequest) String() string { func (*GetSandboxProviderEnvironmentRequest) ProtoMessage() {} func (x *GetSandboxProviderEnvironmentRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[133] + mi := &file_openshell_proto_msgTypes[134] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10780,7 +11024,7 @@ func (x *GetSandboxProviderEnvironmentRequest) ProtoReflect() protoreflect.Messa // Deprecated: Use GetSandboxProviderEnvironmentRequest.ProtoReflect.Descriptor instead. func (*GetSandboxProviderEnvironmentRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{133} + return file_openshell_proto_rawDescGZIP(), []int{134} } func (x *GetSandboxProviderEnvironmentRequest) GetSandboxId() string { @@ -10809,7 +11053,7 @@ type StaticCredentialEndpointBinding struct { func (x *StaticCredentialEndpointBinding) Reset() { *x = StaticCredentialEndpointBinding{} - mi := &file_openshell_proto_msgTypes[134] + mi := &file_openshell_proto_msgTypes[135] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10821,7 +11065,7 @@ func (x *StaticCredentialEndpointBinding) String() string { func (*StaticCredentialEndpointBinding) ProtoMessage() {} func (x *StaticCredentialEndpointBinding) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[134] + mi := &file_openshell_proto_msgTypes[135] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10834,7 +11078,7 @@ func (x *StaticCredentialEndpointBinding) ProtoReflect() protoreflect.Message { // Deprecated: Use StaticCredentialEndpointBinding.ProtoReflect.Descriptor instead. func (*StaticCredentialEndpointBinding) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{134} + return file_openshell_proto_rawDescGZIP(), []int{135} } func (x *StaticCredentialEndpointBinding) GetHost() string { @@ -10878,7 +11122,7 @@ type StaticCredentialBinding struct { func (x *StaticCredentialBinding) Reset() { *x = StaticCredentialBinding{} - mi := &file_openshell_proto_msgTypes[135] + mi := &file_openshell_proto_msgTypes[136] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10890,7 +11134,7 @@ func (x *StaticCredentialBinding) String() string { func (*StaticCredentialBinding) ProtoMessage() {} func (x *StaticCredentialBinding) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[135] + mi := &file_openshell_proto_msgTypes[136] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10903,7 +11147,7 @@ func (x *StaticCredentialBinding) ProtoReflect() protoreflect.Message { // Deprecated: Use StaticCredentialBinding.ProtoReflect.Descriptor instead. func (*StaticCredentialBinding) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{135} + return file_openshell_proto_rawDescGZIP(), []int{136} } func (x *StaticCredentialBinding) GetEndpoints() []*StaticCredentialEndpointBinding { @@ -10954,13 +11198,15 @@ type GetSandboxProviderEnvironmentResponse struct { PolicyHash string `protobuf:"bytes,8,opt,name=policy_hash,json=policyHash,proto3" json:"policy_hash,omitempty"` // Nonzero when material was withheld; installing an empty map is not readiness. ReadinessReason ProviderReadinessReason `protobuf:"varint,9,opt,name=readiness_reason,json=readinessReason,proto3,enum=openshell.v1.ProviderReadinessReason" json:"readiness_reason,omitempty"` - unknownFields protoimpl.UnknownFields - sizeCache protoimpl.SizeCache + // Complete desired set of non-secret managed files, keyed by absolute path. + Files map[string]string `protobuf:"bytes,10,rep,name=files,proto3" json:"files,omitempty" protobuf_key:"bytes,1,opt,name=key" protobuf_val:"bytes,2,opt,name=value"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache } func (x *GetSandboxProviderEnvironmentResponse) Reset() { *x = GetSandboxProviderEnvironmentResponse{} - mi := &file_openshell_proto_msgTypes[136] + mi := &file_openshell_proto_msgTypes[137] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -10972,7 +11218,7 @@ func (x *GetSandboxProviderEnvironmentResponse) String() string { func (*GetSandboxProviderEnvironmentResponse) ProtoMessage() {} func (x *GetSandboxProviderEnvironmentResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[136] + mi := &file_openshell_proto_msgTypes[137] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -10985,7 +11231,7 @@ func (x *GetSandboxProviderEnvironmentResponse) ProtoReflect() protoreflect.Mess // Deprecated: Use GetSandboxProviderEnvironmentResponse.ProtoReflect.Descriptor instead. func (*GetSandboxProviderEnvironmentResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{136} + return file_openshell_proto_rawDescGZIP(), []int{137} } func (x *GetSandboxProviderEnvironmentResponse) GetEnvironment() map[string]string { @@ -11051,6 +11297,13 @@ func (x *GetSandboxProviderEnvironmentResponse) GetReadinessReason() ProviderRea return ProviderReadinessReason_PROVIDER_READINESS_REASON_UNSPECIFIED } +func (x *GetSandboxProviderEnvironmentResponse) GetFiles() map[string]string { + if x != nil { + return x.Files + } + return nil +} + type ExchangeProviderSubjectTokenRequest struct { state protoimpl.MessageState `protogen:"open.v1"` // The sandbox ID. Must match the authenticated sandbox principal. @@ -11068,7 +11321,7 @@ type ExchangeProviderSubjectTokenRequest struct { func (x *ExchangeProviderSubjectTokenRequest) Reset() { *x = ExchangeProviderSubjectTokenRequest{} - mi := &file_openshell_proto_msgTypes[137] + mi := &file_openshell_proto_msgTypes[138] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11080,7 +11333,7 @@ func (x *ExchangeProviderSubjectTokenRequest) String() string { func (*ExchangeProviderSubjectTokenRequest) ProtoMessage() {} func (x *ExchangeProviderSubjectTokenRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[137] + mi := &file_openshell_proto_msgTypes[138] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11093,7 +11346,7 @@ func (x *ExchangeProviderSubjectTokenRequest) ProtoReflect() protoreflect.Messag // Deprecated: Use ExchangeProviderSubjectTokenRequest.ProtoReflect.Descriptor instead. func (*ExchangeProviderSubjectTokenRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{137} + return file_openshell_proto_rawDescGZIP(), []int{138} } func (x *ExchangeProviderSubjectTokenRequest) GetSandboxId() string { @@ -11135,7 +11388,7 @@ type ExchangeProviderSubjectTokenResponse struct { func (x *ExchangeProviderSubjectTokenResponse) Reset() { *x = ExchangeProviderSubjectTokenResponse{} - mi := &file_openshell_proto_msgTypes[138] + mi := &file_openshell_proto_msgTypes[139] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11147,7 +11400,7 @@ func (x *ExchangeProviderSubjectTokenResponse) String() string { func (*ExchangeProviderSubjectTokenResponse) ProtoMessage() {} func (x *ExchangeProviderSubjectTokenResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[138] + mi := &file_openshell_proto_msgTypes[139] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11160,7 +11413,7 @@ func (x *ExchangeProviderSubjectTokenResponse) ProtoReflect() protoreflect.Messa // Deprecated: Use ExchangeProviderSubjectTokenResponse.ProtoReflect.Descriptor instead. func (*ExchangeProviderSubjectTokenResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{138} + return file_openshell_proto_rawDescGZIP(), []int{139} } func (x *ExchangeProviderSubjectTokenResponse) GetAccessToken() string { @@ -11233,7 +11486,7 @@ type UpdateConfigRequest struct { func (x *UpdateConfigRequest) Reset() { *x = UpdateConfigRequest{} - mi := &file_openshell_proto_msgTypes[139] + mi := &file_openshell_proto_msgTypes[140] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11245,7 +11498,7 @@ func (x *UpdateConfigRequest) String() string { func (*UpdateConfigRequest) ProtoMessage() {} func (x *UpdateConfigRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[139] + mi := &file_openshell_proto_msgTypes[140] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11258,7 +11511,7 @@ func (x *UpdateConfigRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use UpdateConfigRequest.ProtoReflect.Descriptor instead. func (*UpdateConfigRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{139} + return file_openshell_proto_rawDescGZIP(), []int{140} } func (x *UpdateConfigRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -11355,7 +11608,7 @@ type PolicyMergeOperation struct { func (x *PolicyMergeOperation) Reset() { *x = PolicyMergeOperation{} - mi := &file_openshell_proto_msgTypes[140] + mi := &file_openshell_proto_msgTypes[141] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11367,7 +11620,7 @@ func (x *PolicyMergeOperation) String() string { func (*PolicyMergeOperation) ProtoMessage() {} func (x *PolicyMergeOperation) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[140] + mi := &file_openshell_proto_msgTypes[141] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11380,7 +11633,7 @@ func (x *PolicyMergeOperation) ProtoReflect() protoreflect.Message { // Deprecated: Use PolicyMergeOperation.ProtoReflect.Descriptor instead. func (*PolicyMergeOperation) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{140} + return file_openshell_proto_rawDescGZIP(), []int{141} } func (x *PolicyMergeOperation) GetOperation() isPolicyMergeOperation_Operation { @@ -11494,7 +11747,7 @@ type AddNetworkRule struct { func (x *AddNetworkRule) Reset() { *x = AddNetworkRule{} - mi := &file_openshell_proto_msgTypes[141] + mi := &file_openshell_proto_msgTypes[142] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11506,7 +11759,7 @@ func (x *AddNetworkRule) String() string { func (*AddNetworkRule) ProtoMessage() {} func (x *AddNetworkRule) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[141] + mi := &file_openshell_proto_msgTypes[142] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11519,7 +11772,7 @@ func (x *AddNetworkRule) ProtoReflect() protoreflect.Message { // Deprecated: Use AddNetworkRule.ProtoReflect.Descriptor instead. func (*AddNetworkRule) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{141} + return file_openshell_proto_rawDescGZIP(), []int{142} } func (x *AddNetworkRule) GetRuleName() string { @@ -11547,7 +11800,7 @@ type RemoveNetworkEndpoint struct { func (x *RemoveNetworkEndpoint) Reset() { *x = RemoveNetworkEndpoint{} - mi := &file_openshell_proto_msgTypes[142] + mi := &file_openshell_proto_msgTypes[143] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11559,7 +11812,7 @@ func (x *RemoveNetworkEndpoint) String() string { func (*RemoveNetworkEndpoint) ProtoMessage() {} func (x *RemoveNetworkEndpoint) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[142] + mi := &file_openshell_proto_msgTypes[143] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11572,7 +11825,7 @@ func (x *RemoveNetworkEndpoint) ProtoReflect() protoreflect.Message { // Deprecated: Use RemoveNetworkEndpoint.ProtoReflect.Descriptor instead. func (*RemoveNetworkEndpoint) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{142} + return file_openshell_proto_rawDescGZIP(), []int{143} } func (x *RemoveNetworkEndpoint) GetRuleName() string { @@ -11605,7 +11858,7 @@ type RemoveNetworkRule struct { func (x *RemoveNetworkRule) Reset() { *x = RemoveNetworkRule{} - mi := &file_openshell_proto_msgTypes[143] + mi := &file_openshell_proto_msgTypes[144] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11617,7 +11870,7 @@ func (x *RemoveNetworkRule) String() string { func (*RemoveNetworkRule) ProtoMessage() {} func (x *RemoveNetworkRule) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[143] + mi := &file_openshell_proto_msgTypes[144] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11630,7 +11883,7 @@ func (x *RemoveNetworkRule) ProtoReflect() protoreflect.Message { // Deprecated: Use RemoveNetworkRule.ProtoReflect.Descriptor instead. func (*RemoveNetworkRule) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{143} + return file_openshell_proto_rawDescGZIP(), []int{144} } func (x *RemoveNetworkRule) GetRuleName() string { @@ -11659,7 +11912,7 @@ type L7RuleTarget struct { func (x *L7RuleTarget) Reset() { *x = L7RuleTarget{} - mi := &file_openshell_proto_msgTypes[144] + mi := &file_openshell_proto_msgTypes[145] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11671,7 +11924,7 @@ func (x *L7RuleTarget) String() string { func (*L7RuleTarget) ProtoMessage() {} func (x *L7RuleTarget) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[144] + mi := &file_openshell_proto_msgTypes[145] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11684,7 +11937,7 @@ func (x *L7RuleTarget) ProtoReflect() protoreflect.Message { // Deprecated: Use L7RuleTarget.ProtoReflect.Descriptor instead. func (*L7RuleTarget) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{144} + return file_openshell_proto_rawDescGZIP(), []int{145} } func (x *L7RuleTarget) GetRuleName() string { @@ -11739,7 +11992,7 @@ type AddDenyRules struct { func (x *AddDenyRules) Reset() { *x = AddDenyRules{} - mi := &file_openshell_proto_msgTypes[145] + mi := &file_openshell_proto_msgTypes[146] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11751,7 +12004,7 @@ func (x *AddDenyRules) String() string { func (*AddDenyRules) ProtoMessage() {} func (x *AddDenyRules) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[145] + mi := &file_openshell_proto_msgTypes[146] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11764,7 +12017,7 @@ func (x *AddDenyRules) ProtoReflect() protoreflect.Message { // Deprecated: Use AddDenyRules.ProtoReflect.Descriptor instead. func (*AddDenyRules) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{145} + return file_openshell_proto_rawDescGZIP(), []int{146} } func (x *AddDenyRules) GetDenyRules() []*sandboxv1.L7DenyRule { @@ -11791,7 +12044,7 @@ type AddAllowRules struct { func (x *AddAllowRules) Reset() { *x = AddAllowRules{} - mi := &file_openshell_proto_msgTypes[146] + mi := &file_openshell_proto_msgTypes[147] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11803,7 +12056,7 @@ func (x *AddAllowRules) String() string { func (*AddAllowRules) ProtoMessage() {} func (x *AddAllowRules) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[146] + mi := &file_openshell_proto_msgTypes[147] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11816,7 +12069,7 @@ func (x *AddAllowRules) ProtoReflect() protoreflect.Message { // Deprecated: Use AddAllowRules.ProtoReflect.Descriptor instead. func (*AddAllowRules) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{146} + return file_openshell_proto_rawDescGZIP(), []int{147} } func (x *AddAllowRules) GetRules() []*sandboxv1.L7Rule { @@ -11843,7 +12096,7 @@ type RemoveNetworkBinary struct { func (x *RemoveNetworkBinary) Reset() { *x = RemoveNetworkBinary{} - mi := &file_openshell_proto_msgTypes[147] + mi := &file_openshell_proto_msgTypes[148] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11855,7 +12108,7 @@ func (x *RemoveNetworkBinary) String() string { func (*RemoveNetworkBinary) ProtoMessage() {} func (x *RemoveNetworkBinary) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[147] + mi := &file_openshell_proto_msgTypes[148] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11868,7 +12121,7 @@ func (x *RemoveNetworkBinary) ProtoReflect() protoreflect.Message { // Deprecated: Use RemoveNetworkBinary.ProtoReflect.Descriptor instead. func (*RemoveNetworkBinary) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{147} + return file_openshell_proto_rawDescGZIP(), []int{148} } func (x *RemoveNetworkBinary) GetRuleName() string { @@ -11904,7 +12157,7 @@ type UpdateConfigResponse struct { func (x *UpdateConfigResponse) Reset() { *x = UpdateConfigResponse{} - mi := &file_openshell_proto_msgTypes[148] + mi := &file_openshell_proto_msgTypes[149] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11916,7 +12169,7 @@ func (x *UpdateConfigResponse) String() string { func (*UpdateConfigResponse) ProtoMessage() {} func (x *UpdateConfigResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[148] + mi := &file_openshell_proto_msgTypes[149] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -11929,7 +12182,7 @@ func (x *UpdateConfigResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use UpdateConfigResponse.ProtoReflect.Descriptor instead. func (*UpdateConfigResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{148} + return file_openshell_proto_rawDescGZIP(), []int{149} } func (x *UpdateConfigResponse) GetVersion() uint32 { @@ -11983,7 +12236,7 @@ type GetSandboxPolicyStatusRequest struct { func (x *GetSandboxPolicyStatusRequest) Reset() { *x = GetSandboxPolicyStatusRequest{} - mi := &file_openshell_proto_msgTypes[149] + mi := &file_openshell_proto_msgTypes[150] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -11995,7 +12248,7 @@ func (x *GetSandboxPolicyStatusRequest) String() string { func (*GetSandboxPolicyStatusRequest) ProtoMessage() {} func (x *GetSandboxPolicyStatusRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[149] + mi := &file_openshell_proto_msgTypes[150] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12008,7 +12261,7 @@ func (x *GetSandboxPolicyStatusRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use GetSandboxPolicyStatusRequest.ProtoReflect.Descriptor instead. func (*GetSandboxPolicyStatusRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{149} + return file_openshell_proto_rawDescGZIP(), []int{150} } func (x *GetSandboxPolicyStatusRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -12052,7 +12305,7 @@ type GetSandboxPolicyStatusResponse struct { func (x *GetSandboxPolicyStatusResponse) Reset() { *x = GetSandboxPolicyStatusResponse{} - mi := &file_openshell_proto_msgTypes[150] + mi := &file_openshell_proto_msgTypes[151] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12064,7 +12317,7 @@ func (x *GetSandboxPolicyStatusResponse) String() string { func (*GetSandboxPolicyStatusResponse) ProtoMessage() {} func (x *GetSandboxPolicyStatusResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[150] + mi := &file_openshell_proto_msgTypes[151] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12077,7 +12330,7 @@ func (x *GetSandboxPolicyStatusResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use GetSandboxPolicyStatusResponse.ProtoReflect.Descriptor instead. func (*GetSandboxPolicyStatusResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{150} + return file_openshell_proto_rawDescGZIP(), []int{151} } func (x *GetSandboxPolicyStatusResponse) GetRevision() *SandboxPolicyRevision { @@ -12114,7 +12367,7 @@ type ListSandboxPoliciesRequest struct { func (x *ListSandboxPoliciesRequest) Reset() { *x = ListSandboxPoliciesRequest{} - mi := &file_openshell_proto_msgTypes[151] + mi := &file_openshell_proto_msgTypes[152] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12126,7 +12379,7 @@ func (x *ListSandboxPoliciesRequest) String() string { func (*ListSandboxPoliciesRequest) ProtoMessage() {} func (x *ListSandboxPoliciesRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[151] + mi := &file_openshell_proto_msgTypes[152] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12139,7 +12392,7 @@ func (x *ListSandboxPoliciesRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ListSandboxPoliciesRequest.ProtoReflect.Descriptor instead. func (*ListSandboxPoliciesRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{151} + return file_openshell_proto_rawDescGZIP(), []int{152} } func (x *ListSandboxPoliciesRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -12191,7 +12444,7 @@ type ListSandboxPoliciesResponse struct { func (x *ListSandboxPoliciesResponse) Reset() { *x = ListSandboxPoliciesResponse{} - mi := &file_openshell_proto_msgTypes[152] + mi := &file_openshell_proto_msgTypes[153] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12203,7 +12456,7 @@ func (x *ListSandboxPoliciesResponse) String() string { func (*ListSandboxPoliciesResponse) ProtoMessage() {} func (x *ListSandboxPoliciesResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[152] + mi := &file_openshell_proto_msgTypes[153] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12216,7 +12469,7 @@ func (x *ListSandboxPoliciesResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ListSandboxPoliciesResponse.ProtoReflect.Descriptor instead. func (*ListSandboxPoliciesResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{152} + return file_openshell_proto_rawDescGZIP(), []int{153} } func (x *ListSandboxPoliciesResponse) GetRevisions() []*SandboxPolicyRevision { @@ -12250,7 +12503,7 @@ type ReportPolicyStatusRequest struct { func (x *ReportPolicyStatusRequest) Reset() { *x = ReportPolicyStatusRequest{} - mi := &file_openshell_proto_msgTypes[153] + mi := &file_openshell_proto_msgTypes[154] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12262,7 +12515,7 @@ func (x *ReportPolicyStatusRequest) String() string { func (*ReportPolicyStatusRequest) ProtoMessage() {} func (x *ReportPolicyStatusRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[153] + mi := &file_openshell_proto_msgTypes[154] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12275,7 +12528,7 @@ func (x *ReportPolicyStatusRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ReportPolicyStatusRequest.ProtoReflect.Descriptor instead. func (*ReportPolicyStatusRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{153} + return file_openshell_proto_rawDescGZIP(), []int{154} } func (x *ReportPolicyStatusRequest) GetSandboxId() string { @@ -12315,7 +12568,7 @@ type ReportPolicyStatusResponse struct { func (x *ReportPolicyStatusResponse) Reset() { *x = ReportPolicyStatusResponse{} - mi := &file_openshell_proto_msgTypes[154] + mi := &file_openshell_proto_msgTypes[155] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12327,7 +12580,7 @@ func (x *ReportPolicyStatusResponse) String() string { func (*ReportPolicyStatusResponse) ProtoMessage() {} func (x *ReportPolicyStatusResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[154] + mi := &file_openshell_proto_msgTypes[155] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12340,7 +12593,7 @@ func (x *ReportPolicyStatusResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ReportPolicyStatusResponse.ProtoReflect.Descriptor instead. func (*ReportPolicyStatusResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{154} + return file_openshell_proto_rawDescGZIP(), []int{155} } type SandboxConfigurationAdmission struct { @@ -12358,7 +12611,7 @@ type SandboxConfigurationAdmission struct { func (x *SandboxConfigurationAdmission) Reset() { *x = SandboxConfigurationAdmission{} - mi := &file_openshell_proto_msgTypes[155] + mi := &file_openshell_proto_msgTypes[156] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12370,7 +12623,7 @@ func (x *SandboxConfigurationAdmission) String() string { func (*SandboxConfigurationAdmission) ProtoMessage() {} func (x *SandboxConfigurationAdmission) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[155] + mi := &file_openshell_proto_msgTypes[156] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12383,7 +12636,7 @@ func (x *SandboxConfigurationAdmission) ProtoReflect() protoreflect.Message { // Deprecated: Use SandboxConfigurationAdmission.ProtoReflect.Descriptor instead. func (*SandboxConfigurationAdmission) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{155} + return file_openshell_proto_rawDescGZIP(), []int{156} } func (x *SandboxConfigurationAdmission) GetInstanceId() string { @@ -12447,7 +12700,7 @@ type ReportSandboxConfigurationRequest struct { func (x *ReportSandboxConfigurationRequest) Reset() { *x = ReportSandboxConfigurationRequest{} - mi := &file_openshell_proto_msgTypes[156] + mi := &file_openshell_proto_msgTypes[157] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12459,7 +12712,7 @@ func (x *ReportSandboxConfigurationRequest) String() string { func (*ReportSandboxConfigurationRequest) ProtoMessage() {} func (x *ReportSandboxConfigurationRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[156] + mi := &file_openshell_proto_msgTypes[157] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12472,7 +12725,7 @@ func (x *ReportSandboxConfigurationRequest) ProtoReflect() protoreflect.Message // Deprecated: Use ReportSandboxConfigurationRequest.ProtoReflect.Descriptor instead. func (*ReportSandboxConfigurationRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{156} + return file_openshell_proto_rawDescGZIP(), []int{157} } func (x *ReportSandboxConfigurationRequest) GetSandboxId() string { @@ -12504,7 +12757,7 @@ type ReportSandboxConfigurationResponse struct { func (x *ReportSandboxConfigurationResponse) Reset() { *x = ReportSandboxConfigurationResponse{} - mi := &file_openshell_proto_msgTypes[157] + mi := &file_openshell_proto_msgTypes[158] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12516,7 +12769,7 @@ func (x *ReportSandboxConfigurationResponse) String() string { func (*ReportSandboxConfigurationResponse) ProtoMessage() {} func (x *ReportSandboxConfigurationResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[157] + mi := &file_openshell_proto_msgTypes[158] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12529,7 +12782,7 @@ func (x *ReportSandboxConfigurationResponse) ProtoReflect() protoreflect.Message // Deprecated: Use ReportSandboxConfigurationResponse.ProtoReflect.Descriptor instead. func (*ReportSandboxConfigurationResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{157} + return file_openshell_proto_rawDescGZIP(), []int{158} } // A versioned policy revision with metadata. @@ -12562,7 +12815,7 @@ type SandboxPolicyRevision struct { func (x *SandboxPolicyRevision) Reset() { *x = SandboxPolicyRevision{} - mi := &file_openshell_proto_msgTypes[158] + mi := &file_openshell_proto_msgTypes[159] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12574,7 +12827,7 @@ func (x *SandboxPolicyRevision) String() string { func (*SandboxPolicyRevision) ProtoMessage() {} func (x *SandboxPolicyRevision) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[158] + mi := &file_openshell_proto_msgTypes[159] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12587,7 +12840,7 @@ func (x *SandboxPolicyRevision) ProtoReflect() protoreflect.Message { // Deprecated: Use SandboxPolicyRevision.ProtoReflect.Descriptor instead. func (*SandboxPolicyRevision) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{158} + return file_openshell_proto_rawDescGZIP(), []int{159} } func (x *SandboxPolicyRevision) GetVersion() uint32 { @@ -12667,7 +12920,7 @@ type GetSandboxLogsRequest struct { func (x *GetSandboxLogsRequest) Reset() { *x = GetSandboxLogsRequest{} - mi := &file_openshell_proto_msgTypes[159] + mi := &file_openshell_proto_msgTypes[160] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12679,7 +12932,7 @@ func (x *GetSandboxLogsRequest) String() string { func (*GetSandboxLogsRequest) ProtoMessage() {} func (x *GetSandboxLogsRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[159] + mi := &file_openshell_proto_msgTypes[160] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12692,7 +12945,7 @@ func (x *GetSandboxLogsRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use GetSandboxLogsRequest.ProtoReflect.Descriptor instead. func (*GetSandboxLogsRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{159} + return file_openshell_proto_rawDescGZIP(), []int{160} } func (x *GetSandboxLogsRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -12750,7 +13003,7 @@ type PushSandboxLogsRequest struct { func (x *PushSandboxLogsRequest) Reset() { *x = PushSandboxLogsRequest{} - mi := &file_openshell_proto_msgTypes[160] + mi := &file_openshell_proto_msgTypes[161] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12762,7 +13015,7 @@ func (x *PushSandboxLogsRequest) String() string { func (*PushSandboxLogsRequest) ProtoMessage() {} func (x *PushSandboxLogsRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[160] + mi := &file_openshell_proto_msgTypes[161] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12775,7 +13028,7 @@ func (x *PushSandboxLogsRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use PushSandboxLogsRequest.ProtoReflect.Descriptor instead. func (*PushSandboxLogsRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{160} + return file_openshell_proto_rawDescGZIP(), []int{161} } func (x *PushSandboxLogsRequest) GetSandboxId() string { @@ -12801,7 +13054,7 @@ type PushSandboxLogsResponse struct { func (x *PushSandboxLogsResponse) Reset() { *x = PushSandboxLogsResponse{} - mi := &file_openshell_proto_msgTypes[161] + mi := &file_openshell_proto_msgTypes[162] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12813,7 +13066,7 @@ func (x *PushSandboxLogsResponse) String() string { func (*PushSandboxLogsResponse) ProtoMessage() {} func (x *PushSandboxLogsResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[161] + mi := &file_openshell_proto_msgTypes[162] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12826,7 +13079,7 @@ func (x *PushSandboxLogsResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use PushSandboxLogsResponse.ProtoReflect.Descriptor instead. func (*PushSandboxLogsResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{161} + return file_openshell_proto_rawDescGZIP(), []int{162} } // Get sandbox logs response. @@ -12842,7 +13095,7 @@ type GetSandboxLogsResponse struct { func (x *GetSandboxLogsResponse) Reset() { *x = GetSandboxLogsResponse{} - mi := &file_openshell_proto_msgTypes[162] + mi := &file_openshell_proto_msgTypes[163] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12854,7 +13107,7 @@ func (x *GetSandboxLogsResponse) String() string { func (*GetSandboxLogsResponse) ProtoMessage() {} func (x *GetSandboxLogsResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[162] + mi := &file_openshell_proto_msgTypes[163] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12867,7 +13120,7 @@ func (x *GetSandboxLogsResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use GetSandboxLogsResponse.ProtoReflect.Descriptor instead. func (*GetSandboxLogsResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{162} + return file_openshell_proto_rawDescGZIP(), []int{163} } func (x *GetSandboxLogsResponse) GetLogs() []*SandboxLogLine { @@ -12900,7 +13153,7 @@ type SupervisorMessage struct { func (x *SupervisorMessage) Reset() { *x = SupervisorMessage{} - mi := &file_openshell_proto_msgTypes[163] + mi := &file_openshell_proto_msgTypes[164] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -12912,7 +13165,7 @@ func (x *SupervisorMessage) String() string { func (*SupervisorMessage) ProtoMessage() {} func (x *SupervisorMessage) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[163] + mi := &file_openshell_proto_msgTypes[164] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -12925,7 +13178,7 @@ func (x *SupervisorMessage) ProtoReflect() protoreflect.Message { // Deprecated: Use SupervisorMessage.ProtoReflect.Descriptor instead. func (*SupervisorMessage) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{163} + return file_openshell_proto_rawDescGZIP(), []int{164} } func (x *SupervisorMessage) GetPayload() isSupervisorMessage_Payload { @@ -13016,7 +13269,7 @@ type GatewayMessage struct { func (x *GatewayMessage) Reset() { *x = GatewayMessage{} - mi := &file_openshell_proto_msgTypes[164] + mi := &file_openshell_proto_msgTypes[165] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13028,7 +13281,7 @@ func (x *GatewayMessage) String() string { func (*GatewayMessage) ProtoMessage() {} func (x *GatewayMessage) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[164] + mi := &file_openshell_proto_msgTypes[165] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13041,7 +13294,7 @@ func (x *GatewayMessage) ProtoReflect() protoreflect.Message { // Deprecated: Use GatewayMessage.ProtoReflect.Descriptor instead. func (*GatewayMessage) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{164} + return file_openshell_proto_rawDescGZIP(), []int{165} } func (x *GatewayMessage) GetPayload() isGatewayMessage_Payload { @@ -13148,7 +13401,7 @@ type SupervisorHello struct { func (x *SupervisorHello) Reset() { *x = SupervisorHello{} - mi := &file_openshell_proto_msgTypes[165] + mi := &file_openshell_proto_msgTypes[166] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13160,7 +13413,7 @@ func (x *SupervisorHello) String() string { func (*SupervisorHello) ProtoMessage() {} func (x *SupervisorHello) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[165] + mi := &file_openshell_proto_msgTypes[166] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13173,7 +13426,7 @@ func (x *SupervisorHello) ProtoReflect() protoreflect.Message { // Deprecated: Use SupervisorHello.ProtoReflect.Descriptor instead. func (*SupervisorHello) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{165} + return file_openshell_proto_rawDescGZIP(), []int{166} } func (x *SupervisorHello) GetSandboxId() string { @@ -13217,7 +13470,7 @@ type SessionAccepted struct { func (x *SessionAccepted) Reset() { *x = SessionAccepted{} - mi := &file_openshell_proto_msgTypes[166] + mi := &file_openshell_proto_msgTypes[167] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13229,7 +13482,7 @@ func (x *SessionAccepted) String() string { func (*SessionAccepted) ProtoMessage() {} func (x *SessionAccepted) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[166] + mi := &file_openshell_proto_msgTypes[167] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13242,7 +13495,7 @@ func (x *SessionAccepted) ProtoReflect() protoreflect.Message { // Deprecated: Use SessionAccepted.ProtoReflect.Descriptor instead. func (*SessionAccepted) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{166} + return file_openshell_proto_rawDescGZIP(), []int{167} } func (x *SessionAccepted) GetSessionId() string { @@ -13270,7 +13523,7 @@ type SessionRejected struct { func (x *SessionRejected) Reset() { *x = SessionRejected{} - mi := &file_openshell_proto_msgTypes[167] + mi := &file_openshell_proto_msgTypes[168] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13282,7 +13535,7 @@ func (x *SessionRejected) String() string { func (*SessionRejected) ProtoMessage() {} func (x *SessionRejected) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[167] + mi := &file_openshell_proto_msgTypes[168] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13295,7 +13548,7 @@ func (x *SessionRejected) ProtoReflect() protoreflect.Message { // Deprecated: Use SessionRejected.ProtoReflect.Descriptor instead. func (*SessionRejected) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{167} + return file_openshell_proto_rawDescGZIP(), []int{168} } func (x *SessionRejected) GetReason() string { @@ -13314,7 +13567,7 @@ type SupervisorHeartbeat struct { func (x *SupervisorHeartbeat) Reset() { *x = SupervisorHeartbeat{} - mi := &file_openshell_proto_msgTypes[168] + mi := &file_openshell_proto_msgTypes[169] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13326,7 +13579,7 @@ func (x *SupervisorHeartbeat) String() string { func (*SupervisorHeartbeat) ProtoMessage() {} func (x *SupervisorHeartbeat) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[168] + mi := &file_openshell_proto_msgTypes[169] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13339,7 +13592,7 @@ func (x *SupervisorHeartbeat) ProtoReflect() protoreflect.Message { // Deprecated: Use SupervisorHeartbeat.ProtoReflect.Descriptor instead. func (*SupervisorHeartbeat) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{168} + return file_openshell_proto_rawDescGZIP(), []int{169} } // Gateway heartbeat. @@ -13351,7 +13604,7 @@ type GatewayHeartbeat struct { func (x *GatewayHeartbeat) Reset() { *x = GatewayHeartbeat{} - mi := &file_openshell_proto_msgTypes[169] + mi := &file_openshell_proto_msgTypes[170] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13363,7 +13616,7 @@ func (x *GatewayHeartbeat) String() string { func (*GatewayHeartbeat) ProtoMessage() {} func (x *GatewayHeartbeat) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[169] + mi := &file_openshell_proto_msgTypes[170] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13376,7 +13629,7 @@ func (x *GatewayHeartbeat) ProtoReflect() protoreflect.Message { // Deprecated: Use GatewayHeartbeat.ProtoReflect.Descriptor instead. func (*GatewayHeartbeat) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{169} + return file_openshell_proto_rawDescGZIP(), []int{170} } // Terminal result reported before the supervisor shuts down. A successful RPC @@ -13393,7 +13646,7 @@ type ReportMainProcessExitRequest struct { func (x *ReportMainProcessExitRequest) Reset() { *x = ReportMainProcessExitRequest{} - mi := &file_openshell_proto_msgTypes[170] + mi := &file_openshell_proto_msgTypes[171] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13405,7 +13658,7 @@ func (x *ReportMainProcessExitRequest) String() string { func (*ReportMainProcessExitRequest) ProtoMessage() {} func (x *ReportMainProcessExitRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[170] + mi := &file_openshell_proto_msgTypes[171] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13418,7 +13671,7 @@ func (x *ReportMainProcessExitRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ReportMainProcessExitRequest.ProtoReflect.Descriptor instead. func (*ReportMainProcessExitRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{170} + return file_openshell_proto_rawDescGZIP(), []int{171} } func (x *ReportMainProcessExitRequest) GetSandboxId() string { @@ -13450,7 +13703,7 @@ type ReportMainProcessExitResponse struct { func (x *ReportMainProcessExitResponse) Reset() { *x = ReportMainProcessExitResponse{} - mi := &file_openshell_proto_msgTypes[171] + mi := &file_openshell_proto_msgTypes[172] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13462,7 +13715,7 @@ func (x *ReportMainProcessExitResponse) String() string { func (*ReportMainProcessExitResponse) ProtoMessage() {} func (x *ReportMainProcessExitResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[171] + mi := &file_openshell_proto_msgTypes[172] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13475,7 +13728,7 @@ func (x *ReportMainProcessExitResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ReportMainProcessExitResponse.ProtoReflect.Descriptor instead. func (*ReportMainProcessExitResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{171} + return file_openshell_proto_rawDescGZIP(), []int{172} } // Terminal-delivery completion reported after all expected foreground SSH @@ -13490,7 +13743,7 @@ type FinalizeMainProcessExitRequest struct { func (x *FinalizeMainProcessExitRequest) Reset() { *x = FinalizeMainProcessExitRequest{} - mi := &file_openshell_proto_msgTypes[172] + mi := &file_openshell_proto_msgTypes[173] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13502,7 +13755,7 @@ func (x *FinalizeMainProcessExitRequest) String() string { func (*FinalizeMainProcessExitRequest) ProtoMessage() {} func (x *FinalizeMainProcessExitRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[172] + mi := &file_openshell_proto_msgTypes[173] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13515,7 +13768,7 @@ func (x *FinalizeMainProcessExitRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use FinalizeMainProcessExitRequest.ProtoReflect.Descriptor instead. func (*FinalizeMainProcessExitRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{172} + return file_openshell_proto_rawDescGZIP(), []int{173} } func (x *FinalizeMainProcessExitRequest) GetSandboxId() string { @@ -13540,7 +13793,7 @@ type FinalizeMainProcessExitResponse struct { func (x *FinalizeMainProcessExitResponse) Reset() { *x = FinalizeMainProcessExitResponse{} - mi := &file_openshell_proto_msgTypes[173] + mi := &file_openshell_proto_msgTypes[174] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13552,7 +13805,7 @@ func (x *FinalizeMainProcessExitResponse) String() string { func (*FinalizeMainProcessExitResponse) ProtoMessage() {} func (x *FinalizeMainProcessExitResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[173] + mi := &file_openshell_proto_msgTypes[174] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13565,7 +13818,7 @@ func (x *FinalizeMainProcessExitResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use FinalizeMainProcessExitResponse.ProtoReflect.Descriptor instead. func (*FinalizeMainProcessExitResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{173} + return file_openshell_proto_rawDescGZIP(), []int{174} } // Gateway requests the supervisor to open a relay channel. @@ -13594,7 +13847,7 @@ type RelayOpen struct { func (x *RelayOpen) Reset() { *x = RelayOpen{} - mi := &file_openshell_proto_msgTypes[174] + mi := &file_openshell_proto_msgTypes[175] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13606,7 +13859,7 @@ func (x *RelayOpen) String() string { func (*RelayOpen) ProtoMessage() {} func (x *RelayOpen) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[174] + mi := &file_openshell_proto_msgTypes[175] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13619,7 +13872,7 @@ func (x *RelayOpen) ProtoReflect() protoreflect.Message { // Deprecated: Use RelayOpen.ProtoReflect.Descriptor instead. func (*RelayOpen) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{174} + return file_openshell_proto_rawDescGZIP(), []int{175} } func (x *RelayOpen) GetChannelId() string { @@ -13686,7 +13939,7 @@ type SshRelayTarget struct { func (x *SshRelayTarget) Reset() { *x = SshRelayTarget{} - mi := &file_openshell_proto_msgTypes[175] + mi := &file_openshell_proto_msgTypes[176] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13698,7 +13951,7 @@ func (x *SshRelayTarget) String() string { func (*SshRelayTarget) ProtoMessage() {} func (x *SshRelayTarget) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[175] + mi := &file_openshell_proto_msgTypes[176] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13711,7 +13964,7 @@ func (x *SshRelayTarget) ProtoReflect() protoreflect.Message { // Deprecated: Use SshRelayTarget.ProtoReflect.Descriptor instead. func (*SshRelayTarget) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{175} + return file_openshell_proto_rawDescGZIP(), []int{176} } // TCP target dialed by the supervisor from inside the sandbox. @@ -13727,7 +13980,7 @@ type TcpRelayTarget struct { func (x *TcpRelayTarget) Reset() { *x = TcpRelayTarget{} - mi := &file_openshell_proto_msgTypes[176] + mi := &file_openshell_proto_msgTypes[177] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13739,7 +13992,7 @@ func (x *TcpRelayTarget) String() string { func (*TcpRelayTarget) ProtoMessage() {} func (x *TcpRelayTarget) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[176] + mi := &file_openshell_proto_msgTypes[177] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13752,7 +14005,7 @@ func (x *TcpRelayTarget) ProtoReflect() protoreflect.Message { // Deprecated: Use TcpRelayTarget.ProtoReflect.Descriptor instead. func (*TcpRelayTarget) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{176} + return file_openshell_proto_rawDescGZIP(), []int{177} } func (x *TcpRelayTarget) GetHost() string { @@ -13780,7 +14033,7 @@ type RelayInit struct { func (x *RelayInit) Reset() { *x = RelayInit{} - mi := &file_openshell_proto_msgTypes[177] + mi := &file_openshell_proto_msgTypes[178] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13792,7 +14045,7 @@ func (x *RelayInit) String() string { func (*RelayInit) ProtoMessage() {} func (x *RelayInit) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[177] + mi := &file_openshell_proto_msgTypes[178] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13805,7 +14058,7 @@ func (x *RelayInit) ProtoReflect() protoreflect.Message { // Deprecated: Use RelayInit.ProtoReflect.Descriptor instead. func (*RelayInit) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{177} + return file_openshell_proto_rawDescGZIP(), []int{178} } func (x *RelayInit) GetChannelId() string { @@ -13832,7 +14085,7 @@ type RelayFrame struct { func (x *RelayFrame) Reset() { *x = RelayFrame{} - mi := &file_openshell_proto_msgTypes[178] + mi := &file_openshell_proto_msgTypes[179] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13844,7 +14097,7 @@ func (x *RelayFrame) String() string { func (*RelayFrame) ProtoMessage() {} func (x *RelayFrame) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[178] + mi := &file_openshell_proto_msgTypes[179] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13857,7 +14110,7 @@ func (x *RelayFrame) ProtoReflect() protoreflect.Message { // Deprecated: Use RelayFrame.ProtoReflect.Descriptor instead. func (*RelayFrame) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{178} + return file_openshell_proto_rawDescGZIP(), []int{179} } func (x *RelayFrame) GetPayload() isRelayFrame_Payload { @@ -13917,7 +14170,7 @@ type PeerRelayInit struct { func (x *PeerRelayInit) Reset() { *x = PeerRelayInit{} - mi := &file_openshell_proto_msgTypes[179] + mi := &file_openshell_proto_msgTypes[180] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13929,7 +14182,7 @@ func (x *PeerRelayInit) String() string { func (*PeerRelayInit) ProtoMessage() {} func (x *PeerRelayInit) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[179] + mi := &file_openshell_proto_msgTypes[180] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -13942,7 +14195,7 @@ func (x *PeerRelayInit) ProtoReflect() protoreflect.Message { // Deprecated: Use PeerRelayInit.ProtoReflect.Descriptor instead. func (*PeerRelayInit) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{179} + return file_openshell_proto_rawDescGZIP(), []int{180} } func (x *PeerRelayInit) GetSandboxId() string { @@ -13980,7 +14233,7 @@ type PeerRelayFrame struct { func (x *PeerRelayFrame) Reset() { *x = PeerRelayFrame{} - mi := &file_openshell_proto_msgTypes[180] + mi := &file_openshell_proto_msgTypes[181] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -13992,7 +14245,7 @@ func (x *PeerRelayFrame) String() string { func (*PeerRelayFrame) ProtoMessage() {} func (x *PeerRelayFrame) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[180] + mi := &file_openshell_proto_msgTypes[181] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14005,7 +14258,7 @@ func (x *PeerRelayFrame) ProtoReflect() protoreflect.Message { // Deprecated: Use PeerRelayFrame.ProtoReflect.Descriptor instead. func (*PeerRelayFrame) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{180} + return file_openshell_proto_rawDescGZIP(), []int{181} } func (x *PeerRelayFrame) GetPayload() isPeerRelayFrame_Payload { @@ -14064,7 +14317,7 @@ type RelayOpenResult struct { func (x *RelayOpenResult) Reset() { *x = RelayOpenResult{} - mi := &file_openshell_proto_msgTypes[181] + mi := &file_openshell_proto_msgTypes[182] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14076,7 +14329,7 @@ func (x *RelayOpenResult) String() string { func (*RelayOpenResult) ProtoMessage() {} func (x *RelayOpenResult) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[181] + mi := &file_openshell_proto_msgTypes[182] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14089,7 +14342,7 @@ func (x *RelayOpenResult) ProtoReflect() protoreflect.Message { // Deprecated: Use RelayOpenResult.ProtoReflect.Descriptor instead. func (*RelayOpenResult) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{181} + return file_openshell_proto_rawDescGZIP(), []int{182} } func (x *RelayOpenResult) GetChannelId() string { @@ -14126,7 +14379,7 @@ type RelayClose struct { func (x *RelayClose) Reset() { *x = RelayClose{} - mi := &file_openshell_proto_msgTypes[182] + mi := &file_openshell_proto_msgTypes[183] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14138,7 +14391,7 @@ func (x *RelayClose) String() string { func (*RelayClose) ProtoMessage() {} func (x *RelayClose) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[182] + mi := &file_openshell_proto_msgTypes[183] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14151,7 +14404,7 @@ func (x *RelayClose) ProtoReflect() protoreflect.Message { // Deprecated: Use RelayClose.ProtoReflect.Descriptor instead. func (*RelayClose) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{182} + return file_openshell_proto_rawDescGZIP(), []int{183} } func (x *RelayClose) GetChannelId() string { @@ -14185,7 +14438,7 @@ type L7RequestSample struct { func (x *L7RequestSample) Reset() { *x = L7RequestSample{} - mi := &file_openshell_proto_msgTypes[183] + mi := &file_openshell_proto_msgTypes[184] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14197,7 +14450,7 @@ func (x *L7RequestSample) String() string { func (*L7RequestSample) ProtoMessage() {} func (x *L7RequestSample) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[183] + mi := &file_openshell_proto_msgTypes[184] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14210,7 +14463,7 @@ func (x *L7RequestSample) ProtoReflect() protoreflect.Message { // Deprecated: Use L7RequestSample.ProtoReflect.Descriptor instead. func (*L7RequestSample) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{183} + return file_openshell_proto_rawDescGZIP(), []int{184} } func (x *L7RequestSample) GetMethod() string { @@ -14284,7 +14537,7 @@ type DenialSummary struct { func (x *DenialSummary) Reset() { *x = DenialSummary{} - mi := &file_openshell_proto_msgTypes[184] + mi := &file_openshell_proto_msgTypes[185] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14296,7 +14549,7 @@ func (x *DenialSummary) String() string { func (*DenialSummary) ProtoMessage() {} func (x *DenialSummary) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[184] + mi := &file_openshell_proto_msgTypes[185] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14309,7 +14562,7 @@ func (x *DenialSummary) ProtoReflect() protoreflect.Message { // Deprecated: Use DenialSummary.ProtoReflect.Descriptor instead. func (*DenialSummary) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{184} + return file_openshell_proto_rawDescGZIP(), []int{185} } func (x *DenialSummary) GetSandboxId() string { @@ -14444,7 +14697,7 @@ type DenialGroupCount struct { func (x *DenialGroupCount) Reset() { *x = DenialGroupCount{} - mi := &file_openshell_proto_msgTypes[185] + mi := &file_openshell_proto_msgTypes[186] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14456,7 +14709,7 @@ func (x *DenialGroupCount) String() string { func (*DenialGroupCount) ProtoMessage() {} func (x *DenialGroupCount) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[185] + mi := &file_openshell_proto_msgTypes[186] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14469,7 +14722,7 @@ func (x *DenialGroupCount) ProtoReflect() protoreflect.Message { // Deprecated: Use DenialGroupCount.ProtoReflect.Descriptor instead. func (*DenialGroupCount) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{185} + return file_openshell_proto_rawDescGZIP(), []int{186} } func (x *DenialGroupCount) GetDenyGroup() string { @@ -14502,7 +14755,7 @@ type NetworkActivitySummary struct { func (x *NetworkActivitySummary) Reset() { *x = NetworkActivitySummary{} - mi := &file_openshell_proto_msgTypes[186] + mi := &file_openshell_proto_msgTypes[187] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14514,7 +14767,7 @@ func (x *NetworkActivitySummary) String() string { func (*NetworkActivitySummary) ProtoMessage() {} func (x *NetworkActivitySummary) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[186] + mi := &file_openshell_proto_msgTypes[187] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14527,7 +14780,7 @@ func (x *NetworkActivitySummary) ProtoReflect() protoreflect.Message { // Deprecated: Use NetworkActivitySummary.ProtoReflect.Descriptor instead. func (*NetworkActivitySummary) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{186} + return file_openshell_proto_rawDescGZIP(), []int{187} } func (x *NetworkActivitySummary) GetNetworkActivityCount() uint32 { @@ -14615,7 +14868,7 @@ type PolicyChunk struct { func (x *PolicyChunk) Reset() { *x = PolicyChunk{} - mi := &file_openshell_proto_msgTypes[187] + mi := &file_openshell_proto_msgTypes[188] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14627,7 +14880,7 @@ func (x *PolicyChunk) String() string { func (*PolicyChunk) ProtoMessage() {} func (x *PolicyChunk) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[187] + mi := &file_openshell_proto_msgTypes[188] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14640,7 +14893,7 @@ func (x *PolicyChunk) ProtoReflect() protoreflect.Message { // Deprecated: Use PolicyChunk.ProtoReflect.Descriptor instead. func (*PolicyChunk) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{187} + return file_openshell_proto_rawDescGZIP(), []int{188} } func (x *PolicyChunk) GetId() string { @@ -14828,7 +15081,7 @@ type DraftPolicyUpdate struct { func (x *DraftPolicyUpdate) Reset() { *x = DraftPolicyUpdate{} - mi := &file_openshell_proto_msgTypes[188] + mi := &file_openshell_proto_msgTypes[189] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14840,7 +15093,7 @@ func (x *DraftPolicyUpdate) String() string { func (*DraftPolicyUpdate) ProtoMessage() {} func (x *DraftPolicyUpdate) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[188] + mi := &file_openshell_proto_msgTypes[189] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14853,7 +15106,7 @@ func (x *DraftPolicyUpdate) ProtoReflect() protoreflect.Message { // Deprecated: Use DraftPolicyUpdate.ProtoReflect.Descriptor instead. func (*DraftPolicyUpdate) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{188} + return file_openshell_proto_rawDescGZIP(), []int{189} } func (x *DraftPolicyUpdate) GetDraftVersion() uint64 { @@ -14912,7 +15165,7 @@ type SubmitPolicyAnalysisRequest struct { func (x *SubmitPolicyAnalysisRequest) Reset() { *x = SubmitPolicyAnalysisRequest{} - mi := &file_openshell_proto_msgTypes[189] + mi := &file_openshell_proto_msgTypes[190] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -14924,7 +15177,7 @@ func (x *SubmitPolicyAnalysisRequest) String() string { func (*SubmitPolicyAnalysisRequest) ProtoMessage() {} func (x *SubmitPolicyAnalysisRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[189] + mi := &file_openshell_proto_msgTypes[190] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -14937,7 +15190,7 @@ func (x *SubmitPolicyAnalysisRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use SubmitPolicyAnalysisRequest.ProtoReflect.Descriptor instead. func (*SubmitPolicyAnalysisRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{189} + return file_openshell_proto_rawDescGZIP(), []int{190} } func (x *SubmitPolicyAnalysisRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15000,7 +15253,7 @@ type SubmitPolicyAnalysisResponse struct { func (x *SubmitPolicyAnalysisResponse) Reset() { *x = SubmitPolicyAnalysisResponse{} - mi := &file_openshell_proto_msgTypes[190] + mi := &file_openshell_proto_msgTypes[191] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15012,7 +15265,7 @@ func (x *SubmitPolicyAnalysisResponse) String() string { func (*SubmitPolicyAnalysisResponse) ProtoMessage() {} func (x *SubmitPolicyAnalysisResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[190] + mi := &file_openshell_proto_msgTypes[191] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15025,7 +15278,7 @@ func (x *SubmitPolicyAnalysisResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use SubmitPolicyAnalysisResponse.ProtoReflect.Descriptor instead. func (*SubmitPolicyAnalysisResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{190} + return file_openshell_proto_rawDescGZIP(), []int{191} } func (x *SubmitPolicyAnalysisResponse) GetAcceptedChunks() uint32 { @@ -15070,7 +15323,7 @@ type GetDraftPolicyRequest struct { func (x *GetDraftPolicyRequest) Reset() { *x = GetDraftPolicyRequest{} - mi := &file_openshell_proto_msgTypes[191] + mi := &file_openshell_proto_msgTypes[192] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15082,7 +15335,7 @@ func (x *GetDraftPolicyRequest) String() string { func (*GetDraftPolicyRequest) ProtoMessage() {} func (x *GetDraftPolicyRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[191] + mi := &file_openshell_proto_msgTypes[192] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15095,7 +15348,7 @@ func (x *GetDraftPolicyRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use GetDraftPolicyRequest.ProtoReflect.Descriptor instead. func (*GetDraftPolicyRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{191} + return file_openshell_proto_rawDescGZIP(), []int{192} } func (x *GetDraftPolicyRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15135,7 +15388,7 @@ type GetDraftPolicyResponse struct { func (x *GetDraftPolicyResponse) Reset() { *x = GetDraftPolicyResponse{} - mi := &file_openshell_proto_msgTypes[192] + mi := &file_openshell_proto_msgTypes[193] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15147,7 +15400,7 @@ func (x *GetDraftPolicyResponse) String() string { func (*GetDraftPolicyResponse) ProtoMessage() {} func (x *GetDraftPolicyResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[192] + mi := &file_openshell_proto_msgTypes[193] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15160,7 +15413,7 @@ func (x *GetDraftPolicyResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use GetDraftPolicyResponse.ProtoReflect.Descriptor instead. func (*GetDraftPolicyResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{192} + return file_openshell_proto_rawDescGZIP(), []int{193} } func (x *GetDraftPolicyResponse) GetChunks() []*PolicyChunk { @@ -15211,7 +15464,7 @@ type ApproveDraftChunkRequest struct { func (x *ApproveDraftChunkRequest) Reset() { *x = ApproveDraftChunkRequest{} - mi := &file_openshell_proto_msgTypes[193] + mi := &file_openshell_proto_msgTypes[194] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15223,7 +15476,7 @@ func (x *ApproveDraftChunkRequest) String() string { func (*ApproveDraftChunkRequest) ProtoMessage() {} func (x *ApproveDraftChunkRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[193] + mi := &file_openshell_proto_msgTypes[194] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15236,7 +15489,7 @@ func (x *ApproveDraftChunkRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ApproveDraftChunkRequest.ProtoReflect.Descriptor instead. func (*ApproveDraftChunkRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{193} + return file_openshell_proto_rawDescGZIP(), []int{194} } func (x *ApproveDraftChunkRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15286,7 +15539,7 @@ type ApproveDraftChunkResponse struct { func (x *ApproveDraftChunkResponse) Reset() { *x = ApproveDraftChunkResponse{} - mi := &file_openshell_proto_msgTypes[194] + mi := &file_openshell_proto_msgTypes[195] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15298,7 +15551,7 @@ func (x *ApproveDraftChunkResponse) String() string { func (*ApproveDraftChunkResponse) ProtoMessage() {} func (x *ApproveDraftChunkResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[194] + mi := &file_openshell_proto_msgTypes[195] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15311,7 +15564,7 @@ func (x *ApproveDraftChunkResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ApproveDraftChunkResponse.ProtoReflect.Descriptor instead. func (*ApproveDraftChunkResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{194} + return file_openshell_proto_rawDescGZIP(), []int{195} } func (x *ApproveDraftChunkResponse) GetPolicyVersion() uint32 { @@ -15347,7 +15600,7 @@ type RejectDraftChunkRequest struct { func (x *RejectDraftChunkRequest) Reset() { *x = RejectDraftChunkRequest{} - mi := &file_openshell_proto_msgTypes[195] + mi := &file_openshell_proto_msgTypes[196] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15359,7 +15612,7 @@ func (x *RejectDraftChunkRequest) String() string { func (*RejectDraftChunkRequest) ProtoMessage() {} func (x *RejectDraftChunkRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[195] + mi := &file_openshell_proto_msgTypes[196] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15372,7 +15625,7 @@ func (x *RejectDraftChunkRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use RejectDraftChunkRequest.ProtoReflect.Descriptor instead. func (*RejectDraftChunkRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{195} + return file_openshell_proto_rawDescGZIP(), []int{196} } func (x *RejectDraftChunkRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15418,7 +15671,7 @@ type RejectDraftChunkResponse struct { func (x *RejectDraftChunkResponse) Reset() { *x = RejectDraftChunkResponse{} - mi := &file_openshell_proto_msgTypes[196] + mi := &file_openshell_proto_msgTypes[197] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15430,7 +15683,7 @@ func (x *RejectDraftChunkResponse) String() string { func (*RejectDraftChunkResponse) ProtoMessage() {} func (x *RejectDraftChunkResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[196] + mi := &file_openshell_proto_msgTypes[197] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15443,7 +15696,7 @@ func (x *RejectDraftChunkResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use RejectDraftChunkResponse.ProtoReflect.Descriptor instead. func (*RejectDraftChunkResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{196} + return file_openshell_proto_rawDescGZIP(), []int{197} } // Approve all pending chunks. @@ -15457,7 +15710,7 @@ type DraftChunkApproval struct { func (x *DraftChunkApproval) Reset() { *x = DraftChunkApproval{} - mi := &file_openshell_proto_msgTypes[197] + mi := &file_openshell_proto_msgTypes[198] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15469,7 +15722,7 @@ func (x *DraftChunkApproval) String() string { func (*DraftChunkApproval) ProtoMessage() {} func (x *DraftChunkApproval) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[197] + mi := &file_openshell_proto_msgTypes[198] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15482,7 +15735,7 @@ func (x *DraftChunkApproval) ProtoReflect() protoreflect.Message { // Deprecated: Use DraftChunkApproval.ProtoReflect.Descriptor instead. func (*DraftChunkApproval) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{197} + return file_openshell_proto_rawDescGZIP(), []int{198} } func (x *DraftChunkApproval) GetChunkId() string { @@ -15518,7 +15771,7 @@ type ApproveAllDraftChunksRequest struct { func (x *ApproveAllDraftChunksRequest) Reset() { *x = ApproveAllDraftChunksRequest{} - mi := &file_openshell_proto_msgTypes[198] + mi := &file_openshell_proto_msgTypes[199] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15530,7 +15783,7 @@ func (x *ApproveAllDraftChunksRequest) String() string { func (*ApproveAllDraftChunksRequest) ProtoMessage() {} func (x *ApproveAllDraftChunksRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[198] + mi := &file_openshell_proto_msgTypes[199] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15543,7 +15796,7 @@ func (x *ApproveAllDraftChunksRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ApproveAllDraftChunksRequest.ProtoReflect.Descriptor instead. func (*ApproveAllDraftChunksRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{198} + return file_openshell_proto_rawDescGZIP(), []int{199} } func (x *ApproveAllDraftChunksRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15598,7 +15851,7 @@ type ApproveAllDraftChunksResponse struct { func (x *ApproveAllDraftChunksResponse) Reset() { *x = ApproveAllDraftChunksResponse{} - mi := &file_openshell_proto_msgTypes[199] + mi := &file_openshell_proto_msgTypes[200] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15610,7 +15863,7 @@ func (x *ApproveAllDraftChunksResponse) String() string { func (*ApproveAllDraftChunksResponse) ProtoMessage() {} func (x *ApproveAllDraftChunksResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[199] + mi := &file_openshell_proto_msgTypes[200] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15623,7 +15876,7 @@ func (x *ApproveAllDraftChunksResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ApproveAllDraftChunksResponse.ProtoReflect.Descriptor instead. func (*ApproveAllDraftChunksResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{199} + return file_openshell_proto_rawDescGZIP(), []int{200} } func (x *ApproveAllDraftChunksResponse) GetPolicyVersion() uint32 { @@ -15673,7 +15926,7 @@ type EditDraftChunkRequest struct { func (x *EditDraftChunkRequest) Reset() { *x = EditDraftChunkRequest{} - mi := &file_openshell_proto_msgTypes[200] + mi := &file_openshell_proto_msgTypes[201] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15685,7 +15938,7 @@ func (x *EditDraftChunkRequest) String() string { func (*EditDraftChunkRequest) ProtoMessage() {} func (x *EditDraftChunkRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[200] + mi := &file_openshell_proto_msgTypes[201] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15698,7 +15951,7 @@ func (x *EditDraftChunkRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use EditDraftChunkRequest.ProtoReflect.Descriptor instead. func (*EditDraftChunkRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{200} + return file_openshell_proto_rawDescGZIP(), []int{201} } func (x *EditDraftChunkRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15744,7 +15997,7 @@ type EditDraftChunkResponse struct { func (x *EditDraftChunkResponse) Reset() { *x = EditDraftChunkResponse{} - mi := &file_openshell_proto_msgTypes[201] + mi := &file_openshell_proto_msgTypes[202] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15756,7 +16009,7 @@ func (x *EditDraftChunkResponse) String() string { func (*EditDraftChunkResponse) ProtoMessage() {} func (x *EditDraftChunkResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[201] + mi := &file_openshell_proto_msgTypes[202] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15769,7 +16022,7 @@ func (x *EditDraftChunkResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use EditDraftChunkResponse.ProtoReflect.Descriptor instead. func (*EditDraftChunkResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{201} + return file_openshell_proto_rawDescGZIP(), []int{202} } // Reverse an approval (remove merged rule from active policy). @@ -15789,7 +16042,7 @@ type UndoDraftChunkRequest struct { func (x *UndoDraftChunkRequest) Reset() { *x = UndoDraftChunkRequest{} - mi := &file_openshell_proto_msgTypes[202] + mi := &file_openshell_proto_msgTypes[203] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15801,7 +16054,7 @@ func (x *UndoDraftChunkRequest) String() string { func (*UndoDraftChunkRequest) ProtoMessage() {} func (x *UndoDraftChunkRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[202] + mi := &file_openshell_proto_msgTypes[203] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15814,7 +16067,7 @@ func (x *UndoDraftChunkRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use UndoDraftChunkRequest.ProtoReflect.Descriptor instead. func (*UndoDraftChunkRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{202} + return file_openshell_proto_rawDescGZIP(), []int{203} } func (x *UndoDraftChunkRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15857,7 +16110,7 @@ type UndoDraftChunkResponse struct { func (x *UndoDraftChunkResponse) Reset() { *x = UndoDraftChunkResponse{} - mi := &file_openshell_proto_msgTypes[203] + mi := &file_openshell_proto_msgTypes[204] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15869,7 +16122,7 @@ func (x *UndoDraftChunkResponse) String() string { func (*UndoDraftChunkResponse) ProtoMessage() {} func (x *UndoDraftChunkResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[203] + mi := &file_openshell_proto_msgTypes[204] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15882,7 +16135,7 @@ func (x *UndoDraftChunkResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use UndoDraftChunkResponse.ProtoReflect.Descriptor instead. func (*UndoDraftChunkResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{203} + return file_openshell_proto_rawDescGZIP(), []int{204} } func (x *UndoDraftChunkResponse) GetPolicyVersion() uint32 { @@ -15914,7 +16167,7 @@ type ClearDraftChunksRequest struct { func (x *ClearDraftChunksRequest) Reset() { *x = ClearDraftChunksRequest{} - mi := &file_openshell_proto_msgTypes[204] + mi := &file_openshell_proto_msgTypes[205] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15926,7 +16179,7 @@ func (x *ClearDraftChunksRequest) String() string { func (*ClearDraftChunksRequest) ProtoMessage() {} func (x *ClearDraftChunksRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[204] + mi := &file_openshell_proto_msgTypes[205] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15939,7 +16192,7 @@ func (x *ClearDraftChunksRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ClearDraftChunksRequest.ProtoReflect.Descriptor instead. func (*ClearDraftChunksRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{204} + return file_openshell_proto_rawDescGZIP(), []int{205} } func (x *ClearDraftChunksRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -15973,7 +16226,7 @@ type ClearDraftChunksResponse struct { func (x *ClearDraftChunksResponse) Reset() { *x = ClearDraftChunksResponse{} - mi := &file_openshell_proto_msgTypes[205] + mi := &file_openshell_proto_msgTypes[206] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -15985,7 +16238,7 @@ func (x *ClearDraftChunksResponse) String() string { func (*ClearDraftChunksResponse) ProtoMessage() {} func (x *ClearDraftChunksResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[205] + mi := &file_openshell_proto_msgTypes[206] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -15998,7 +16251,7 @@ func (x *ClearDraftChunksResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ClearDraftChunksResponse.ProtoReflect.Descriptor instead. func (*ClearDraftChunksResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{205} + return file_openshell_proto_rawDescGZIP(), []int{206} } func (x *ClearDraftChunksResponse) GetChunksCleared() uint32 { @@ -16020,7 +16273,7 @@ type GetDraftHistoryRequest struct { func (x *GetDraftHistoryRequest) Reset() { *x = GetDraftHistoryRequest{} - mi := &file_openshell_proto_msgTypes[206] + mi := &file_openshell_proto_msgTypes[207] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16032,7 +16285,7 @@ func (x *GetDraftHistoryRequest) String() string { func (*GetDraftHistoryRequest) ProtoMessage() {} func (x *GetDraftHistoryRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[206] + mi := &file_openshell_proto_msgTypes[207] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16045,7 +16298,7 @@ func (x *GetDraftHistoryRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use GetDraftHistoryRequest.ProtoReflect.Descriptor instead. func (*GetDraftHistoryRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{206} + return file_openshell_proto_rawDescGZIP(), []int{207} } func (x *GetDraftHistoryRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -16079,7 +16332,7 @@ type DraftHistoryEntry struct { func (x *DraftHistoryEntry) Reset() { *x = DraftHistoryEntry{} - mi := &file_openshell_proto_msgTypes[207] + mi := &file_openshell_proto_msgTypes[208] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16091,7 +16344,7 @@ func (x *DraftHistoryEntry) String() string { func (*DraftHistoryEntry) ProtoMessage() {} func (x *DraftHistoryEntry) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[207] + mi := &file_openshell_proto_msgTypes[208] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16104,7 +16357,7 @@ func (x *DraftHistoryEntry) ProtoReflect() protoreflect.Message { // Deprecated: Use DraftHistoryEntry.ProtoReflect.Descriptor instead. func (*DraftHistoryEntry) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{207} + return file_openshell_proto_rawDescGZIP(), []int{208} } func (x *DraftHistoryEntry) GetEventTime() *timestamppb.Timestamp { @@ -16145,7 +16398,7 @@ type GetDraftHistoryResponse struct { func (x *GetDraftHistoryResponse) Reset() { *x = GetDraftHistoryResponse{} - mi := &file_openshell_proto_msgTypes[208] + mi := &file_openshell_proto_msgTypes[209] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16157,7 +16410,7 @@ func (x *GetDraftHistoryResponse) String() string { func (*GetDraftHistoryResponse) ProtoMessage() {} func (x *GetDraftHistoryResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[208] + mi := &file_openshell_proto_msgTypes[209] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16170,7 +16423,7 @@ func (x *GetDraftHistoryResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use GetDraftHistoryResponse.ProtoReflect.Descriptor instead. func (*GetDraftHistoryResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{208} + return file_openshell_proto_rawDescGZIP(), []int{209} } func (x *GetDraftHistoryResponse) GetEntries() []*DraftHistoryEntry { @@ -16195,7 +16448,7 @@ type CreateWorkspaceRequest struct { func (x *CreateWorkspaceRequest) Reset() { *x = CreateWorkspaceRequest{} - mi := &file_openshell_proto_msgTypes[209] + mi := &file_openshell_proto_msgTypes[210] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16207,7 +16460,7 @@ func (x *CreateWorkspaceRequest) String() string { func (*CreateWorkspaceRequest) ProtoMessage() {} func (x *CreateWorkspaceRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[209] + mi := &file_openshell_proto_msgTypes[210] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16220,7 +16473,7 @@ func (x *CreateWorkspaceRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use CreateWorkspaceRequest.ProtoReflect.Descriptor instead. func (*CreateWorkspaceRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{209} + return file_openshell_proto_rawDescGZIP(), []int{210} } func (x *CreateWorkspaceRequest) GetName() string { @@ -16254,7 +16507,7 @@ type CreateWorkspaceResponse struct { func (x *CreateWorkspaceResponse) Reset() { *x = CreateWorkspaceResponse{} - mi := &file_openshell_proto_msgTypes[210] + mi := &file_openshell_proto_msgTypes[211] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16266,7 +16519,7 @@ func (x *CreateWorkspaceResponse) String() string { func (*CreateWorkspaceResponse) ProtoMessage() {} func (x *CreateWorkspaceResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[210] + mi := &file_openshell_proto_msgTypes[211] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16279,7 +16532,7 @@ func (x *CreateWorkspaceResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use CreateWorkspaceResponse.ProtoReflect.Descriptor instead. func (*CreateWorkspaceResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{210} + return file_openshell_proto_rawDescGZIP(), []int{211} } func (x *CreateWorkspaceResponse) GetWorkspace() *datamodelv1.Workspace { @@ -16300,7 +16553,7 @@ type GetWorkspaceRequest struct { func (x *GetWorkspaceRequest) Reset() { *x = GetWorkspaceRequest{} - mi := &file_openshell_proto_msgTypes[211] + mi := &file_openshell_proto_msgTypes[212] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16312,7 +16565,7 @@ func (x *GetWorkspaceRequest) String() string { func (*GetWorkspaceRequest) ProtoMessage() {} func (x *GetWorkspaceRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[211] + mi := &file_openshell_proto_msgTypes[212] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16325,7 +16578,7 @@ func (x *GetWorkspaceRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use GetWorkspaceRequest.ProtoReflect.Descriptor instead. func (*GetWorkspaceRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{211} + return file_openshell_proto_rawDescGZIP(), []int{212} } func (x *GetWorkspaceRequest) GetName() string { @@ -16345,7 +16598,7 @@ type GetWorkspaceResponse struct { func (x *GetWorkspaceResponse) Reset() { *x = GetWorkspaceResponse{} - mi := &file_openshell_proto_msgTypes[212] + mi := &file_openshell_proto_msgTypes[213] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16357,7 +16610,7 @@ func (x *GetWorkspaceResponse) String() string { func (*GetWorkspaceResponse) ProtoMessage() {} func (x *GetWorkspaceResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[212] + mi := &file_openshell_proto_msgTypes[213] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16370,7 +16623,7 @@ func (x *GetWorkspaceResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use GetWorkspaceResponse.ProtoReflect.Descriptor instead. func (*GetWorkspaceResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{212} + return file_openshell_proto_rawDescGZIP(), []int{213} } func (x *GetWorkspaceResponse) GetWorkspace() *datamodelv1.Workspace { @@ -16397,7 +16650,7 @@ type ListWorkspacesRequest struct { func (x *ListWorkspacesRequest) Reset() { *x = ListWorkspacesRequest{} - mi := &file_openshell_proto_msgTypes[213] + mi := &file_openshell_proto_msgTypes[214] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16409,7 +16662,7 @@ func (x *ListWorkspacesRequest) String() string { func (*ListWorkspacesRequest) ProtoMessage() {} func (x *ListWorkspacesRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[213] + mi := &file_openshell_proto_msgTypes[214] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16422,7 +16675,7 @@ func (x *ListWorkspacesRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ListWorkspacesRequest.ProtoReflect.Descriptor instead. func (*ListWorkspacesRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{213} + return file_openshell_proto_rawDescGZIP(), []int{214} } func (x *ListWorkspacesRequest) GetPageSize() int32 { @@ -16458,7 +16711,7 @@ type ListWorkspacesResponse struct { func (x *ListWorkspacesResponse) Reset() { *x = ListWorkspacesResponse{} - mi := &file_openshell_proto_msgTypes[214] + mi := &file_openshell_proto_msgTypes[215] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16470,7 +16723,7 @@ func (x *ListWorkspacesResponse) String() string { func (*ListWorkspacesResponse) ProtoMessage() {} func (x *ListWorkspacesResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[214] + mi := &file_openshell_proto_msgTypes[215] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16483,7 +16736,7 @@ func (x *ListWorkspacesResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ListWorkspacesResponse.ProtoReflect.Descriptor instead. func (*ListWorkspacesResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{214} + return file_openshell_proto_rawDescGZIP(), []int{215} } func (x *ListWorkspacesResponse) GetWorkspaces() []*datamodelv1.Workspace { @@ -16514,7 +16767,7 @@ type DeleteWorkspaceRequest struct { func (x *DeleteWorkspaceRequest) Reset() { *x = DeleteWorkspaceRequest{} - mi := &file_openshell_proto_msgTypes[215] + mi := &file_openshell_proto_msgTypes[216] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16526,7 +16779,7 @@ func (x *DeleteWorkspaceRequest) String() string { func (*DeleteWorkspaceRequest) ProtoMessage() {} func (x *DeleteWorkspaceRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[215] + mi := &file_openshell_proto_msgTypes[216] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16539,7 +16792,7 @@ func (x *DeleteWorkspaceRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use DeleteWorkspaceRequest.ProtoReflect.Descriptor instead. func (*DeleteWorkspaceRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{215} + return file_openshell_proto_rawDescGZIP(), []int{216} } func (x *DeleteWorkspaceRequest) GetName() string { @@ -16573,7 +16826,7 @@ type DeleteWorkspaceResponse struct { func (x *DeleteWorkspaceResponse) Reset() { *x = DeleteWorkspaceResponse{} - mi := &file_openshell_proto_msgTypes[216] + mi := &file_openshell_proto_msgTypes[217] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16585,7 +16838,7 @@ func (x *DeleteWorkspaceResponse) String() string { func (*DeleteWorkspaceResponse) ProtoMessage() {} func (x *DeleteWorkspaceResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[216] + mi := &file_openshell_proto_msgTypes[217] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16598,7 +16851,7 @@ func (x *DeleteWorkspaceResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use DeleteWorkspaceResponse.ProtoReflect.Descriptor instead. func (*DeleteWorkspaceResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{216} + return file_openshell_proto_rawDescGZIP(), []int{217} } func (x *DeleteWorkspaceResponse) GetOutcome() DeletionOutcome { @@ -16622,7 +16875,7 @@ type WorkspaceMember struct { func (x *WorkspaceMember) Reset() { *x = WorkspaceMember{} - mi := &file_openshell_proto_msgTypes[217] + mi := &file_openshell_proto_msgTypes[218] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16634,7 +16887,7 @@ func (x *WorkspaceMember) String() string { func (*WorkspaceMember) ProtoMessage() {} func (x *WorkspaceMember) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[217] + mi := &file_openshell_proto_msgTypes[218] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16647,7 +16900,7 @@ func (x *WorkspaceMember) ProtoReflect() protoreflect.Message { // Deprecated: Use WorkspaceMember.ProtoReflect.Descriptor instead. func (*WorkspaceMember) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{217} + return file_openshell_proto_rawDescGZIP(), []int{218} } func (x *WorkspaceMember) GetMetadata() *datamodelv1.ObjectMeta { @@ -16688,7 +16941,7 @@ type AddWorkspaceMemberRequest struct { func (x *AddWorkspaceMemberRequest) Reset() { *x = AddWorkspaceMemberRequest{} - mi := &file_openshell_proto_msgTypes[218] + mi := &file_openshell_proto_msgTypes[219] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16700,7 +16953,7 @@ func (x *AddWorkspaceMemberRequest) String() string { func (*AddWorkspaceMemberRequest) ProtoMessage() {} func (x *AddWorkspaceMemberRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[218] + mi := &file_openshell_proto_msgTypes[219] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16713,7 +16966,7 @@ func (x *AddWorkspaceMemberRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use AddWorkspaceMemberRequest.ProtoReflect.Descriptor instead. func (*AddWorkspaceMemberRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{218} + return file_openshell_proto_rawDescGZIP(), []int{219} } func (x *AddWorkspaceMemberRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -16754,7 +17007,7 @@ type AddWorkspaceMemberResponse struct { func (x *AddWorkspaceMemberResponse) Reset() { *x = AddWorkspaceMemberResponse{} - mi := &file_openshell_proto_msgTypes[219] + mi := &file_openshell_proto_msgTypes[220] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16766,7 +17019,7 @@ func (x *AddWorkspaceMemberResponse) String() string { func (*AddWorkspaceMemberResponse) ProtoMessage() {} func (x *AddWorkspaceMemberResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[219] + mi := &file_openshell_proto_msgTypes[220] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16779,7 +17032,7 @@ func (x *AddWorkspaceMemberResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use AddWorkspaceMemberResponse.ProtoReflect.Descriptor instead. func (*AddWorkspaceMemberResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{219} + return file_openshell_proto_rawDescGZIP(), []int{220} } func (x *AddWorkspaceMemberResponse) GetMember() *WorkspaceMember { @@ -16805,7 +17058,7 @@ type RemoveWorkspaceMemberRequest struct { func (x *RemoveWorkspaceMemberRequest) Reset() { *x = RemoveWorkspaceMemberRequest{} - mi := &file_openshell_proto_msgTypes[220] + mi := &file_openshell_proto_msgTypes[221] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16817,7 +17070,7 @@ func (x *RemoveWorkspaceMemberRequest) String() string { func (*RemoveWorkspaceMemberRequest) ProtoMessage() {} func (x *RemoveWorkspaceMemberRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[220] + mi := &file_openshell_proto_msgTypes[221] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16830,7 +17083,7 @@ func (x *RemoveWorkspaceMemberRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use RemoveWorkspaceMemberRequest.ProtoReflect.Descriptor instead. func (*RemoveWorkspaceMemberRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{220} + return file_openshell_proto_rawDescGZIP(), []int{221} } func (x *RemoveWorkspaceMemberRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -16871,7 +17124,7 @@ type RemoveWorkspaceMemberResponse struct { func (x *RemoveWorkspaceMemberResponse) Reset() { *x = RemoveWorkspaceMemberResponse{} - mi := &file_openshell_proto_msgTypes[221] + mi := &file_openshell_proto_msgTypes[222] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16883,7 +17136,7 @@ func (x *RemoveWorkspaceMemberResponse) String() string { func (*RemoveWorkspaceMemberResponse) ProtoMessage() {} func (x *RemoveWorkspaceMemberResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[221] + mi := &file_openshell_proto_msgTypes[222] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16896,7 +17149,7 @@ func (x *RemoveWorkspaceMemberResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use RemoveWorkspaceMemberResponse.ProtoReflect.Descriptor instead. func (*RemoveWorkspaceMemberResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{221} + return file_openshell_proto_rawDescGZIP(), []int{222} } func (x *RemoveWorkspaceMemberResponse) GetOutcome() DeletionOutcome { @@ -16923,7 +17176,7 @@ type ListWorkspaceMembersRequest struct { func (x *ListWorkspaceMembersRequest) Reset() { *x = ListWorkspaceMembersRequest{} - mi := &file_openshell_proto_msgTypes[222] + mi := &file_openshell_proto_msgTypes[223] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16935,7 +17188,7 @@ func (x *ListWorkspaceMembersRequest) String() string { func (*ListWorkspaceMembersRequest) ProtoMessage() {} func (x *ListWorkspaceMembersRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[222] + mi := &file_openshell_proto_msgTypes[223] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -16948,7 +17201,7 @@ func (x *ListWorkspaceMembersRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ListWorkspaceMembersRequest.ProtoReflect.Descriptor instead. func (*ListWorkspaceMembersRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{222} + return file_openshell_proto_rawDescGZIP(), []int{223} } func (x *ListWorkspaceMembersRequest) GetWorkspaceScope() *datamodelv1.WorkspaceSelector { @@ -16984,7 +17237,7 @@ type ListWorkspaceMembersResponse struct { func (x *ListWorkspaceMembersResponse) Reset() { *x = ListWorkspaceMembersResponse{} - mi := &file_openshell_proto_msgTypes[223] + mi := &file_openshell_proto_msgTypes[224] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -16996,7 +17249,7 @@ func (x *ListWorkspaceMembersResponse) String() string { func (*ListWorkspaceMembersResponse) ProtoMessage() {} func (x *ListWorkspaceMembersResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[223] + mi := &file_openshell_proto_msgTypes[224] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17009,7 +17262,7 @@ func (x *ListWorkspaceMembersResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ListWorkspaceMembersResponse.ProtoReflect.Descriptor instead. func (*ListWorkspaceMembersResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{223} + return file_openshell_proto_rawDescGZIP(), []int{224} } func (x *ListWorkspaceMembersResponse) GetMembers() []*WorkspaceMember { @@ -17044,7 +17297,7 @@ type ExtensionServiceCredential struct { func (x *ExtensionServiceCredential) Reset() { *x = ExtensionServiceCredential{} - mi := &file_openshell_proto_msgTypes[224] + mi := &file_openshell_proto_msgTypes[225] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -17056,7 +17309,7 @@ func (x *ExtensionServiceCredential) String() string { func (*ExtensionServiceCredential) ProtoMessage() {} func (x *ExtensionServiceCredential) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[224] + mi := &file_openshell_proto_msgTypes[225] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17069,7 +17322,7 @@ func (x *ExtensionServiceCredential) ProtoReflect() protoreflect.Message { // Deprecated: Use ExtensionServiceCredential.ProtoReflect.Descriptor instead. func (*ExtensionServiceCredential) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{224} + return file_openshell_proto_rawDescGZIP(), []int{225} } func (x *ExtensionServiceCredential) GetServiceName() string { @@ -17106,7 +17359,7 @@ type EndpointObservation struct { func (x *EndpointObservation) Reset() { *x = EndpointObservation{} - mi := &file_openshell_proto_msgTypes[225] + mi := &file_openshell_proto_msgTypes[226] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -17118,7 +17371,7 @@ func (x *EndpointObservation) String() string { func (*EndpointObservation) ProtoMessage() {} func (x *EndpointObservation) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[225] + mi := &file_openshell_proto_msgTypes[226] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17131,7 +17384,7 @@ func (x *EndpointObservation) ProtoReflect() protoreflect.Message { // Deprecated: Use EndpointObservation.ProtoReflect.Descriptor instead. func (*EndpointObservation) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{225} + return file_openshell_proto_rawDescGZIP(), []int{226} } func (x *EndpointObservation) GetEndpointId() string { @@ -17174,7 +17427,7 @@ type ReportEndpointStatusRequest struct { func (x *ReportEndpointStatusRequest) Reset() { *x = ReportEndpointStatusRequest{} - mi := &file_openshell_proto_msgTypes[226] + mi := &file_openshell_proto_msgTypes[227] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -17186,7 +17439,7 @@ func (x *ReportEndpointStatusRequest) String() string { func (*ReportEndpointStatusRequest) ProtoMessage() {} func (x *ReportEndpointStatusRequest) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[226] + mi := &file_openshell_proto_msgTypes[227] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17199,7 +17452,7 @@ func (x *ReportEndpointStatusRequest) ProtoReflect() protoreflect.Message { // Deprecated: Use ReportEndpointStatusRequest.ProtoReflect.Descriptor instead. func (*ReportEndpointStatusRequest) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{226} + return file_openshell_proto_rawDescGZIP(), []int{227} } func (x *ReportEndpointStatusRequest) GetSandboxId() string { @@ -17260,7 +17513,7 @@ type ReportEndpointStatusResponse struct { func (x *ReportEndpointStatusResponse) Reset() { *x = ReportEndpointStatusResponse{} - mi := &file_openshell_proto_msgTypes[227] + mi := &file_openshell_proto_msgTypes[228] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -17272,7 +17525,7 @@ func (x *ReportEndpointStatusResponse) String() string { func (*ReportEndpointStatusResponse) ProtoMessage() {} func (x *ReportEndpointStatusResponse) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[227] + mi := &file_openshell_proto_msgTypes[228] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17285,7 +17538,7 @@ func (x *ReportEndpointStatusResponse) ProtoReflect() protoreflect.Message { // Deprecated: Use ReportEndpointStatusResponse.ProtoReflect.Descriptor instead. func (*ReportEndpointStatusResponse) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{227} + return file_openshell_proto_rawDescGZIP(), []int{228} } // A configured endpoint and its last accepted network result in one record. @@ -17314,7 +17567,7 @@ type EndpointStatus struct { func (x *EndpointStatus) Reset() { *x = EndpointStatus{} - mi := &file_openshell_proto_msgTypes[228] + mi := &file_openshell_proto_msgTypes[229] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -17326,7 +17579,7 @@ func (x *EndpointStatus) String() string { func (*EndpointStatus) ProtoMessage() {} func (x *EndpointStatus) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[228] + mi := &file_openshell_proto_msgTypes[229] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17339,7 +17592,7 @@ func (x *EndpointStatus) ProtoReflect() protoreflect.Message { // Deprecated: Use EndpointStatus.ProtoReflect.Descriptor instead. func (*EndpointStatus) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{228} + return file_openshell_proto_rawDescGZIP(), []int{229} } func (x *EndpointStatus) GetEndpointId() string { @@ -17409,7 +17662,7 @@ type SandboxProvisioning struct { func (x *SandboxProvisioning) Reset() { *x = SandboxProvisioning{} - mi := &file_openshell_proto_msgTypes[229] + mi := &file_openshell_proto_msgTypes[230] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -17421,7 +17674,7 @@ func (x *SandboxProvisioning) String() string { func (*SandboxProvisioning) ProtoMessage() {} func (x *SandboxProvisioning) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[229] + mi := &file_openshell_proto_msgTypes[230] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17434,7 +17687,7 @@ func (x *SandboxProvisioning) ProtoReflect() protoreflect.Message { // Deprecated: Use SandboxProvisioning.ProtoReflect.Descriptor instead. func (*SandboxProvisioning) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{229} + return file_openshell_proto_rawDescGZIP(), []int{230} } func (x *SandboxProvisioning) GetAttemptId() string { @@ -17520,14 +17773,16 @@ type SandboxServiceExposure struct { // Service name within the sandbox. Empty selects the unnamed endpoint. Service string `protobuf:"bytes,1,opt,name=service,proto3" json:"service,omitempty"` // Loopback TCP port inside the sandbox. - TargetPort uint32 `protobuf:"varint,2,opt,name=target_port,json=targetPort,proto3" json:"target_port,omitempty"` - unknownFields protoimpl.UnknownFields - sizeCache protoimpl.SizeCache + TargetPort uint32 `protobuf:"varint,2,opt,name=target_port,json=targetPort,proto3" json:"target_port,omitempty"` + // Application authorization behavior. Omission resolves to STRIP. + AuthorizationMode ServiceAuthorizationMode `protobuf:"varint,3,opt,name=authorization_mode,json=authorizationMode,proto3,enum=openshell.v1.ServiceAuthorizationMode" json:"authorization_mode,omitempty"` + unknownFields protoimpl.UnknownFields + sizeCache protoimpl.SizeCache } func (x *SandboxServiceExposure) Reset() { *x = SandboxServiceExposure{} - mi := &file_openshell_proto_msgTypes[230] + mi := &file_openshell_proto_msgTypes[231] ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) ms.StoreMessageInfo(mi) } @@ -17539,7 +17794,7 @@ func (x *SandboxServiceExposure) String() string { func (*SandboxServiceExposure) ProtoMessage() {} func (x *SandboxServiceExposure) ProtoReflect() protoreflect.Message { - mi := &file_openshell_proto_msgTypes[230] + mi := &file_openshell_proto_msgTypes[231] if x != nil { ms := protoimpl.X.MessageStateOf(protoimpl.Pointer(x)) if ms.LoadMessageInfo() == nil { @@ -17552,7 +17807,7 @@ func (x *SandboxServiceExposure) ProtoReflect() protoreflect.Message { // Deprecated: Use SandboxServiceExposure.ProtoReflect.Descriptor instead. func (*SandboxServiceExposure) Descriptor() ([]byte, []int) { - return file_openshell_proto_rawDescGZIP(), []int{230} + return file_openshell_proto_rawDescGZIP(), []int{231} } func (x *SandboxServiceExposure) GetService() string { @@ -17569,6 +17824,13 @@ func (x *SandboxServiceExposure) GetTargetPort() uint32 { return 0 } +func (x *SandboxServiceExposure) GetAuthorizationMode() ServiceAuthorizationMode { + if x != nil { + return x.AuthorizationMode + } + return ServiceAuthorizationMode_SERVICE_AUTHORIZATION_MODE_UNSPECIFIED +} + var File_openshell_proto protoreflect.FileDescriptor const file_openshell_proto_rawDesc = "" + @@ -17640,7 +17902,7 @@ const file_openshell_proto_rawDesc = "" + "\bmetadata\x18\x01 \x01(\v2\".openshell.datamodel.v1.ObjectMetaR\bmetadata\x12-\n" + "\x04spec\x18\x02 \x01(\v2\x19.openshell.v1.SandboxSpecR\x04spec\x123\n" + "\x06status\x18\x03 \x01(\v2\x1b.openshell.v1.SandboxStatusR\x06status\x12t\n" + - "\x1ecreated_from_workload_template\x18\x14 \x01(\v2/.openshell.v1.SandboxWorkloadTemplateProvenanceR\x1bcreatedFromWorkloadTemplateJ\x04\b\x04\x10\x05J\x04\b\x05\x10\x06R\x05phaseR\x16current_policy_version\"\xbf\x04\n" + + "\x1ecreated_from_workload_template\x18\x14 \x01(\v2/.openshell.v1.SandboxWorkloadTemplateProvenanceR\x1bcreatedFromWorkloadTemplateJ\x04\b\x04\x10\x05J\x04\b\x05\x10\x06R\x05phaseR\x16current_policy_version\"\x8a\x05\n" + "\vSandboxSpec\x12\x1b\n" + "\tlog_level\x18\x01 \x01(\tR\blogLevel\x12L\n" + "\venvironment\x18\x05 \x03(\v2*.openshell.v1.SandboxSpec.EnvironmentEntryR\venvironment\x129\n" + @@ -17650,7 +17912,8 @@ const file_openshell_proto_rawDesc = "" + "\x15resource_requirements\x18\t \x01(\v2\".openshell.v1.ResourceRequirementsR\x14resourceRequirements\x12\x18\n" + "\acommand\x18\f \x03(\tR\acommand\x12\x10\n" + "\x03tty\x18\r \x01(\bR\x03tty\x12:\n" + - "\x19provider_attachment_epoch\x18\x0e \x01(\tR\x17providerAttachmentEpoch\x1a>\n" + + "\x19provider_attachment_epoch\x18\x0e \x01(\tR\x17providerAttachmentEpoch\x12I\n" + + "\x0erestart_policy\x18\x0f \x01(\x0e2\".openshell.v1.SandboxRestartPolicyR\rrestartPolicy\x1a>\n" + "\x10EnvironmentEntry\x12\x10\n" + "\x03key\x18\x01 \x01(\tR\x03key\x12\x14\n" + "\x05value\x18\x02 \x01(\tR\x05value:\x028\x01J\x04\b\n" + @@ -17708,7 +17971,7 @@ const file_openshell_proto_rawDesc = "" + "\tmax_burst\x18\x02 \x01(\rR\bmaxBurst\"b\n" + "!SandboxWorkloadTemplateProvenance\x12\x12\n" + "\x04name\x18\x01 \x01(\tR\x04name\x12)\n" + - "\x10resource_version\x18\x02 \x01(\tR\x0fresourceVersion\"\xc9\x05\n" + + "\x10resource_version\x18\x02 \x01(\tR\x0fresourceVersion\"\x8d\a\n" + "\rSandboxStatus\x12\x1b\n" + "\tagent_pod\x18\x02 \x01(\tR\bagentPod\x12\x19\n" + "\bagent_fd\x18\x03 \x01(\tR\aagentFd\x12\x1d\n" + @@ -17725,7 +17988,10 @@ const file_openshell_proto_rawDesc = "" + " \x03(\v2\x1c.openshell.v1.EndpointStatusR\x10endpointStatuses\x12d\n" + "\x17configuration_admission\x18\v \x01(\v2+.openshell.v1.SandboxConfigurationAdmissionR\x16configurationAdmission\x12<\n" + "\x17configuration_activated\x18\f \x01(\bH\x01R\x16configurationActivated\x88\x01\x01\x12E\n" + - "\fprovisioning\x18\r \x01(\v2!.openshell.v1.SandboxProvisioningR\fprovisioningB\f\n" + + "\fprovisioning\x18\r \x01(\v2!.openshell.v1.SandboxProvisioningR\fprovisioning\x12#\n" + + "\rrestart_count\x18\x0e \x01(\rR\frestartCount\x12F\n" + + "\x11next_restart_time\x18\x0f \x01(\v2\x1a.google.protobuf.TimestampR\x0fnextRestartTime\x12U\n" + + "\x19main_process_started_time\x18\x10 \x01(\v2\x1a.google.protobuf.TimestampR\x16mainProcessStartedTimeB\f\n" + "\n" + "_exit_codeB\x1a\n" + "\x18_configuration_activated\"\xd1\x01\n" + @@ -17968,7 +18234,7 @@ const file_openshell_proto_rawDesc = "" + "\fgateway_port\x18\x04 \x01(\rR\vgatewayPort\x12%\n" + "\x0egateway_scheme\x18\x05 \x01(\tR\rgatewayScheme\x120\n" + "\x14host_key_fingerprint\x18\a \x01(\tR\x12hostKeyFingerprint\x12C\n" + - "\x0fexpiration_time\x18l \x01(\v2\x1a.google.protobuf.TimestampR\x0eexpirationTimeJ\x04\b\b\x10\tR\rexpires_at_ms\"\xf0\x01\n" + + "\x0fexpiration_time\x18l \x01(\v2\x1a.google.protobuf.TimestampR\x0eexpirationTimeJ\x04\b\b\x10\tR\rexpires_at_ms\"\xc7\x02\n" + "\x14ExposeServiceRequest\x12R\n" + "\x0fworkspace_scope\x18\x05 \x01(\v2).openshell.datamodel.v1.WorkspaceSelectorR\x0eworkspaceScope\x12\x12\n" + "\x04name\x18\x02 \x01(\tR\x04name\x12\x1f\n" + @@ -17977,7 +18243,8 @@ const file_openshell_proto_rawDesc = "" + "\x06domain\x18\x04 \x01(\bR\x06domain\x12\x18\n" + "\asandbox\x18\x01 \x01(\tR\asandbox\x12\x1d\n" + "\n" + - "request_id\x18\x06 \x01(\tR\trequestId\"\x95\x01\n" + + "request_id\x18\x06 \x01(\tR\trequestId\x12U\n" + + "\x12authorization_mode\x18\a \x01(\x0e2&.openshell.v1.ServiceAuthorizationModeR\x11authorizationMode\"\x95\x01\n" + "\x11GetServiceRequest\x12R\n" + "\x0fworkspace_scope\x18\x03 \x01(\v2).openshell.datamodel.v1.WorkspaceSelectorR\x0eworkspaceScope\x12\x12\n" + "\x04name\x18\x02 \x01(\tR\x04name\x12\x18\n" + @@ -17999,7 +18266,7 @@ const file_openshell_proto_rawDesc = "" + "\n" + "request_id\x18\x05 \x01(\tR\trequestId\"_\n" + "\x15DeleteServiceResponse\x127\n" + - "\aoutcome\x18\x02 \x01(\x0e2\x1d.openshell.v1.DeletionOutcomeR\aoutcomeJ\x04\b\x01\x10\x02R\adeleted\"\xd7\x01\n" + + "\aoutcome\x18\x02 \x01(\x0e2\x1d.openshell.v1.DeletionOutcomeR\aoutcomeJ\x04\b\x01\x10\x02R\adeleted\"\xae\x02\n" + "\x0fServiceEndpoint\x12>\n" + "\bmetadata\x18\x01 \x01(\v2\".openshell.datamodel.v1.ObjectMetaR\bmetadata\x12\x1d\n" + "\n" + @@ -18008,7 +18275,8 @@ const file_openshell_proto_rawDesc = "" + "\x04name\x18\x04 \x01(\tR\x04name\x12\x1f\n" + "\vtarget_port\x18\x05 \x01(\rR\n" + "targetPort\x12\x16\n" + - "\x06domain\x18\x06 \x01(\bR\x06domain\"f\n" + + "\x06domain\x18\x06 \x01(\bR\x06domain\x12U\n" + + "\x12authorization_mode\x18\a \x01(\x0e2&.openshell.v1.ServiceAuthorizationModeR\x11authorizationMode\"f\n" + "\x17ServiceEndpointResponse\x129\n" + "\bendpoint\x18\x01 \x01(\v2\x1d.openshell.v1.ServiceEndpointR\bendpoint\x12\x10\n" + "\x03url\x18\x02 \x01(\tR\x03url\"Z\n" + @@ -18286,7 +18554,7 @@ const file_openshell_proto_rawDesc = "" + "\n" + "request_id\x18\x06 \x01(\tR\trequestId\"g\n" + "\x1dDeleteProviderRefreshResponse\x127\n" + - "\aoutcome\x18\x02 \x01(\x0e2\x1d.openshell.v1.DeletionOutcomeR\aoutcomeJ\x04\b\x01\x10\x02R\adeleted\"\xd8\x05\n" + + "\aoutcome\x18\x02 \x01(\x0e2\x1d.openshell.v1.DeletionOutcomeR\aoutcomeJ\x04\b\x01\x10\x02R\adeleted\"\x91\x06\n" + "\x0fProviderProfile\x12\x0e\n" + "\x02id\x18\x01 \x01(\tR\x02id\x12!\n" + "\fdisplay_name\x18\x02 \x01(\tR\vdisplayName\x12 \n" + @@ -18301,10 +18569,15 @@ const file_openshell_proto_rawDesc = "" + " \x01(\x04R\x0fresourceVersion\x12P\n" + "\vannotations\x18\v \x03(\v2..openshell.v1.ProviderProfile.AnnotationsEntryR\vannotations\x12\x16\n" + "\x06source\x18\f \x01(\tR\x06source\x12\x14\n" + - "\x05scope\x18\r \x01(\tR\x05scope\x1a>\n" + + "\x05scope\x18\r \x01(\tR\x05scope\x127\n" + + "\x05files\x18\x0e \x03(\v2!.openshell.v1.ProviderProfileFileR\x05files\x1a>\n" + "\x10AnnotationsEntry\x12\x10\n" + "\x03key\x18\x01 \x01(\tR\x03key\x12\x14\n" + - "\x05value\x18\x02 \x01(\tR\x05value:\x028\x01\"R\n" + + "\x05value\x18\x02 \x01(\tR\x05value:\x028\x01\"\\\n" + + "\x13ProviderProfileFile\x12\x12\n" + + "\x04path\x18\x01 \x01(\tR\x04path\x12\x18\n" + + "\acontent\x18\x02 \x01(\tR\acontent\x12\x17\n" + + "\aenv_var\x18\x03 \x01(\tR\x06envVar\"R\n" + "\x17ProviderProfileResponse\x127\n" + "\aprofile\x18\x01 \x01(\v2\x1d.openshell.v1.ProviderProfileR\aprofile\"\x81\x01\n" + "\x1cListProviderProfilesResponse\x129\n" + @@ -18357,8 +18630,7 @@ const file_openshell_proto_rawDesc = "" + "\x17StaticCredentialBinding\x12K\n" + "\tendpoints\x18\x01 \x03(\v2-.openshell.v1.StaticCredentialEndpointBindingR\tendpoints\x12/\n" + "\x13credential_identity\x18\x02 \x01(\tR\x12credentialIdentity\x12<\n" + - "\x1aworkload_credential_handle\x18\x03 \x01(\tR\x18workloadCredentialHandle\"\x8a\n" + - "\n" + + "\x1aworkload_credential_handle\x18\x03 \x01(\tR\x18workloadCredentialHandle\"\x9a\v\n" + "%GetSandboxProviderEnvironmentResponse\x12l\n" + "\venvironment\x18\x01 \x03(\v2D.openshell.v1.GetSandboxProviderEnvironmentResponse.EnvironmentEntryB\x04\x88\xb5\x18\x01R\venvironment\x122\n" + "\x15provider_env_revision\x18\x02 \x01(\x04R\x13providerEnvRevision\x12\x92\x01\n" + @@ -18369,7 +18641,9 @@ const file_openshell_proto_rawDesc = "" + "\x19provider_attachment_epoch\x18\a \x01(\tR\x17providerAttachmentEpoch\x12\x1f\n" + "\vpolicy_hash\x18\b \x01(\tR\n" + "policyHash\x12P\n" + - "\x10readiness_reason\x18\t \x01(\x0e2%.openshell.v1.ProviderReadinessReasonR\x0freadinessReason\x1a>\n" + + "\x10readiness_reason\x18\t \x01(\x0e2%.openshell.v1.ProviderReadinessReasonR\x0freadinessReason\x12T\n" + + "\x05files\x18\n" + + " \x03(\v2>.openshell.v1.GetSandboxProviderEnvironmentResponse.FilesEntryR\x05files\x1a>\n" + "\x10EnvironmentEntry\x12\x10\n" + "\x03key\x18\x01 \x01(\tR\x03key\x12\x14\n" + "\x05value\x18\x02 \x01(\tR\x05value:\x028\x01\x1ah\n" + @@ -18381,7 +18655,11 @@ const file_openshell_proto_rawDesc = "" + "\x05value\x18\x02 \x01(\v2'.openshell.v1.ProviderProfileCredentialR\x05value:\x028\x01\x1ar\n" + "\x1dStaticCredentialBindingsEntry\x12\x10\n" + "\x03key\x18\x01 \x01(\tR\x03key\x12;\n" + - "\x05value\x18\x02 \x01(\v2%.openshell.v1.StaticCredentialBindingR\x05value:\x028\x01J\x04\b\x03\x10\x04R\x18credential_expires_at_ms\"\xbd\x01\n" + + "\x05value\x18\x02 \x01(\v2%.openshell.v1.StaticCredentialBindingR\x05value:\x028\x01\x1a8\n" + + "\n" + + "FilesEntry\x12\x10\n" + + "\x03key\x18\x01 \x01(\tR\x03key\x12\x14\n" + + "\x05value\x18\x02 \x01(\tR\x05value:\x028\x01J\x04\b\x03\x10\x04R\x18credential_expires_at_ms\"\xbd\x01\n" + "#ExchangeProviderSubjectTokenRequest\x12\x1d\n" + "\n" + "sandbox_id\x18\x01 \x01(\tR\tsandboxId\x12\x1a\n" + @@ -18884,17 +19162,18 @@ const file_openshell_proto_rawDesc = "" + "\x12cleanup_retry_time\x18\t \x01(\v2\x1a.google.protobuf.TimestampR\x10cleanupRetryTime\x120\n" + "\x14attachment_change_id\x18\n" + " \x01(\tR\x12attachmentChangeId\x12P\n" + - "\x16attachment_change_time\x18\v \x01(\v2\x1a.google.protobuf.TimestampR\x14attachmentChangeTime\"S\n" + + "\x16attachment_change_time\x18\v \x01(\v2\x1a.google.protobuf.TimestampR\x14attachmentChangeTime\"\xaa\x01\n" + "\x16SandboxServiceExposure\x12\x18\n" + "\aservice\x18\x01 \x01(\tR\aservice\x12\x1f\n" + "\vtarget_port\x18\x02 \x01(\rR\n" + - "targetPort*\xca\x01\n" + + "targetPort\x12U\n" + + "\x12authorization_mode\x18\x03 \x01(\x0e2&.openshell.v1.ServiceAuthorizationModeR\x11authorizationMode*\xca\x01\n" + "\rExtensionKind\x12\x1e\n" + "\x1aEXTENSION_KIND_UNSPECIFIED\x10\x00\x12!\n" + "\x1dEXTENSION_KIND_COMPUTE_DRIVER\x10\x01\x12$\n" + " EXTENSION_KIND_CREDENTIAL_DRIVER\x10\x02\x12&\n" + "\"EXTENSION_KIND_GATEWAY_INTERCEPTOR\x10\x03\x12(\n" + - "$EXTENSION_KIND_SUPERVISOR_MIDDLEWARE\x10\x04*\xa6\x02\n" + + "$EXTENSION_KIND_SUPERVISOR_MIDDLEWARE\x10\x04*\xc6\x02\n" + "\fSandboxPhase\x12\x1d\n" + "\x19SANDBOX_PHASE_UNSPECIFIED\x10\x00\x12\x1e\n" + "\x1aSANDBOX_PHASE_PROVISIONING\x10\x01\x12\x17\n" + @@ -18905,7 +19184,9 @@ const file_openshell_proto_rawDesc = "" + "\x16SANDBOX_PHASE_STOPPING\x10\x06\x12\x19\n" + "\x15SANDBOX_PHASE_STOPPED\x10\a\x12\x1a\n" + "\x16SANDBOX_PHASE_STARTING\x10\b\x12\x1b\n" + - "\x17SANDBOX_PHASE_COMPLETED\x10\t*\xcb\x01\n" + + "\x17SANDBOX_PHASE_COMPLETED\x10\t\"\x04\b\n" + + "\x10\n" + + "*\x18SANDBOX_PHASE_RESTARTING*\xcb\x01\n" + "\x14ProviderMutationKind\x12&\n" + "\"PROVIDER_MUTATION_KIND_UNSPECIFIED\x10\x00\x12!\n" + "\x1dPROVIDER_MUTATION_KIND_ATTACH\x10\x01\x12!\n" + @@ -19007,7 +19288,12 @@ const file_openshell_proto_rawDesc = "" + "1PROVIDER_CREDENTIAL_REFRESH_RECOVERY_ACTION_RETRY\x10\x01\x12;\n" + "7PROVIDER_CREDENTIAL_REFRESH_RECOVERY_ACTION_REAUTHORIZE\x10\x02\x12A\n" + "=PROVIDER_CREDENTIAL_REFRESH_RECOVERY_ACTION_FIX_CONFIGURATION\x10\x03\x12;\n" + - "7PROVIDER_CREDENTIAL_REFRESH_RECOVERY_ACTION_INVESTIGATE\x10\x04*\x97\x01\n" + + "7PROVIDER_CREDENTIAL_REFRESH_RECOVERY_ACTION_INVESTIGATE\x10\x04*\xaa\x01\n" + + "\x14SandboxRestartPolicy\x12&\n" + + "\"SANDBOX_RESTART_POLICY_UNSPECIFIED\x10\x00\x12 \n" + + "\x1cSANDBOX_RESTART_POLICY_NEVER\x10\x01\x12%\n" + + "!SANDBOX_RESTART_POLICY_ON_FAILURE\x10\x02\x12!\n" + + "\x1dSANDBOX_RESTART_POLICY_ALWAYS\x10\x03*\x97\x01\n" + "\x0fDeletionOutcome\x12 \n" + "\x1cDELETION_OUTCOME_UNSPECIFIED\x10\x00\x12\x1e\n" + "\x1aDELETION_OUTCOME_COMPLETED\x10\x01\x12\x1d\n" + @@ -19021,7 +19307,11 @@ const file_openshell_proto_rawDesc = "" + "&ENDPOINT_RESULT_CREDENTIAL_UNAVAILABLE\x10\x04\x12\x1e\n" + "\x1aENDPOINT_RESULT_TLS_FAILED\x10\x05\x12$\n" + " ENDPOINT_RESULT_TRANSPORT_FAILED\x10\x06\x12%\n" + - "!ENDPOINT_RESULT_UPSTREAM_REJECTED\x10\a2\xafU\n" + + "!ENDPOINT_RESULT_UPSTREAM_REJECTED\x10\a*\x9f\x01\n" + + "\x18ServiceAuthorizationMode\x12*\n" + + "&SERVICE_AUTHORIZATION_MODE_UNSPECIFIED\x10\x00\x12$\n" + + " SERVICE_AUTHORIZATION_MODE_STRIP\x10\x01\x121\n" + + "-SERVICE_AUTHORIZATION_MODE_BEARER_PASSTHROUGH\x10\x022\xafU\n" + "\tOpenShell\x12Z\n" + "\x06Health\x12\x1b.openshell.v1.HealthRequest\x1a\x1c.openshell.v1.HealthResponse\"\x15\x82\xb5\x18\x11\n" + "\x0funauthenticated\x12i\n" + @@ -19207,8 +19497,8 @@ func file_openshell_proto_rawDescGZIP() []byte { return file_openshell_proto_rawDescData } -var file_openshell_proto_enumTypes = make([]protoimpl.EnumInfo, 18) -var file_openshell_proto_msgTypes = make([]protoimpl.MessageInfo, 253) +var file_openshell_proto_enumTypes = make([]protoimpl.EnumInfo, 20) +var file_openshell_proto_msgTypes = make([]protoimpl.MessageInfo, 255) var file_openshell_proto_goTypes = []any{ (ExtensionKind)(0), // 0: openshell.v1.ExtensionKind (SandboxPhase)(0), // 1: openshell.v1.SandboxPhase @@ -19226,767 +19516,779 @@ var file_openshell_proto_goTypes = []any{ (ServiceStatus)(0), // 13: openshell.v1.ServiceStatus (WorkspaceRole)(0), // 14: openshell.v1.WorkspaceRole (ProviderCredentialRefreshRecoveryAction)(0), // 15: openshell.v1.ProviderCredentialRefreshRecoveryAction - (DeletionOutcome)(0), // 16: openshell.v1.DeletionOutcome - (EndpointResult)(0), // 17: openshell.v1.EndpointResult - (*IssueSandboxTokenRequest)(nil), // 18: openshell.v1.IssueSandboxTokenRequest - (*IssueSandboxTokenResponse)(nil), // 19: openshell.v1.IssueSandboxTokenResponse - (*RefreshSandboxTokenRequest)(nil), // 20: openshell.v1.RefreshSandboxTokenRequest - (*RefreshSandboxTokenResponse)(nil), // 21: openshell.v1.RefreshSandboxTokenResponse - (*HealthRequest)(nil), // 22: openshell.v1.HealthRequest - (*HealthResponse)(nil), // 23: openshell.v1.HealthResponse - (*GetCurrentUserRequest)(nil), // 24: openshell.v1.GetCurrentUserRequest - (*GetCurrentUserResponse)(nil), // 25: openshell.v1.GetCurrentUserResponse - (*GetGatewayInfoRequest)(nil), // 26: openshell.v1.GetGatewayInfoRequest - (*GetGatewayInfoResponse)(nil), // 27: openshell.v1.GetGatewayInfoResponse - (*NegotiatedExtensionInfo)(nil), // 28: openshell.v1.NegotiatedExtensionInfo - (*ComputeDriverInfo)(nil), // 29: openshell.v1.ComputeDriverInfo - (*ComputeDriverCapabilities)(nil), // 30: openshell.v1.ComputeDriverCapabilities - (*ResourceCapabilities)(nil), // 31: openshell.v1.ResourceCapabilities - (*CpuResourceCapabilities)(nil), // 32: openshell.v1.CpuResourceCapabilities - (*MemoryResourceCapabilities)(nil), // 33: openshell.v1.MemoryResourceCapabilities - (*GpuResourceCapabilities)(nil), // 34: openshell.v1.GpuResourceCapabilities - (*Sandbox)(nil), // 35: openshell.v1.Sandbox - (*SandboxSpec)(nil), // 36: openshell.v1.SandboxSpec - (*ResourceRequirements)(nil), // 37: openshell.v1.ResourceRequirements - (*GpuResourceRequirements)(nil), // 38: openshell.v1.GpuResourceRequirements - (*SandboxTemplate)(nil), // 39: openshell.v1.SandboxTemplate - (*SandboxWorkloadTemplate)(nil), // 40: openshell.v1.SandboxWorkloadTemplate - (*SandboxWorkloadTemplateSpec)(nil), // 41: openshell.v1.SandboxWorkloadTemplateSpec - (*SandboxWorkloadConfig)(nil), // 42: openshell.v1.SandboxWorkloadConfig - (*SandboxResources)(nil), // 43: openshell.v1.SandboxResources - (*SandboxServiceLevel)(nil), // 44: openshell.v1.SandboxServiceLevel - (*SandboxStartup)(nil), // 45: openshell.v1.SandboxStartup - (*SandboxWorkloadTemplateProvenance)(nil), // 46: openshell.v1.SandboxWorkloadTemplateProvenance - (*SandboxStatus)(nil), // 47: openshell.v1.SandboxStatus - (*SandboxCondition)(nil), // 48: openshell.v1.SandboxCondition - (*PlatformEvent)(nil), // 49: openshell.v1.PlatformEvent - (*CreateSandboxRequest)(nil), // 50: openshell.v1.CreateSandboxRequest - (*CreateSandboxTemplateRequest)(nil), // 51: openshell.v1.CreateSandboxTemplateRequest - (*GetSandboxTemplateRequest)(nil), // 52: openshell.v1.GetSandboxTemplateRequest - (*ListSandboxTemplatesRequest)(nil), // 53: openshell.v1.ListSandboxTemplatesRequest - (*DeleteSandboxTemplateRequest)(nil), // 54: openshell.v1.DeleteSandboxTemplateRequest - (*SandboxTemplateResponse)(nil), // 55: openshell.v1.SandboxTemplateResponse - (*ListSandboxTemplatesResponse)(nil), // 56: openshell.v1.ListSandboxTemplatesResponse - (*DeleteSandboxTemplateResponse)(nil), // 57: openshell.v1.DeleteSandboxTemplateResponse - (*BeginRootfsTarStagingRequest)(nil), // 58: openshell.v1.BeginRootfsTarStagingRequest - (*BeginRootfsTarStagingResponse)(nil), // 59: openshell.v1.BeginRootfsTarStagingResponse - (*GetSandboxRequest)(nil), // 60: openshell.v1.GetSandboxRequest - (*ListSandboxesRequest)(nil), // 61: openshell.v1.ListSandboxesRequest - (*ListSandboxProvidersRequest)(nil), // 62: openshell.v1.ListSandboxProvidersRequest - (*AttachSandboxProviderRequest)(nil), // 63: openshell.v1.AttachSandboxProviderRequest - (*DetachSandboxProviderRequest)(nil), // 64: openshell.v1.DetachSandboxProviderRequest - (*DeleteSandboxRequest)(nil), // 65: openshell.v1.DeleteSandboxRequest - (*StopSandboxRequest)(nil), // 66: openshell.v1.StopSandboxRequest - (*StartSandboxRequest)(nil), // 67: openshell.v1.StartSandboxRequest - (*SandboxResponse)(nil), // 68: openshell.v1.SandboxResponse - (*ListSandboxesResponse)(nil), // 69: openshell.v1.ListSandboxesResponse - (*ListSandboxProvidersResponse)(nil), // 70: openshell.v1.ListSandboxProvidersResponse - (*AttachSandboxProviderResponse)(nil), // 71: openshell.v1.AttachSandboxProviderResponse - (*DetachSandboxProviderResponse)(nil), // 72: openshell.v1.DetachSandboxProviderResponse - (*ProviderDesiredIdentity)(nil), // 73: openshell.v1.ProviderDesiredIdentity - (*ConfigSnapshotRevision)(nil), // 74: openshell.v1.ConfigSnapshotRevision - (*SandboxConfigRevision)(nil), // 75: openshell.v1.SandboxConfigRevision - (*ConfigUpdateOperation)(nil), // 76: openshell.v1.ConfigUpdateOperation - (*ProviderMutationReceipt)(nil), // 77: openshell.v1.ProviderMutationReceipt - (*ProviderReadinessObservation)(nil), // 78: openshell.v1.ProviderReadinessObservation - (*ProviderReadinessStatus)(nil), // 79: openshell.v1.ProviderReadinessStatus - (*GetSandboxProviderStatusRequest)(nil), // 80: openshell.v1.GetSandboxProviderStatusRequest - (*GetSandboxProviderStatusResponse)(nil), // 81: openshell.v1.GetSandboxProviderStatusResponse - (*ReportProviderReadinessRequest)(nil), // 82: openshell.v1.ReportProviderReadinessRequest - (*ReportProviderReadinessResponse)(nil), // 83: openshell.v1.ReportProviderReadinessResponse - (*DeleteSandboxResponse)(nil), // 84: openshell.v1.DeleteSandboxResponse - (*CreateSshSessionRequest)(nil), // 85: openshell.v1.CreateSshSessionRequest - (*CreateSshSessionResponse)(nil), // 86: openshell.v1.CreateSshSessionResponse - (*ExposeServiceRequest)(nil), // 87: openshell.v1.ExposeServiceRequest - (*GetServiceRequest)(nil), // 88: openshell.v1.GetServiceRequest - (*ListServicesRequest)(nil), // 89: openshell.v1.ListServicesRequest - (*ListServicesResponse)(nil), // 90: openshell.v1.ListServicesResponse - (*DeleteServiceRequest)(nil), // 91: openshell.v1.DeleteServiceRequest - (*DeleteServiceResponse)(nil), // 92: openshell.v1.DeleteServiceResponse - (*ServiceEndpoint)(nil), // 93: openshell.v1.ServiceEndpoint - (*ServiceEndpointResponse)(nil), // 94: openshell.v1.ServiceEndpointResponse - (*RevokeSshSessionRequest)(nil), // 95: openshell.v1.RevokeSshSessionRequest - (*RevokeSshSessionResponse)(nil), // 96: openshell.v1.RevokeSshSessionResponse - (*ExecSandboxRequest)(nil), // 97: openshell.v1.ExecSandboxRequest - (*ExecSandboxStdout)(nil), // 98: openshell.v1.ExecSandboxStdout - (*ExecSandboxStderr)(nil), // 99: openshell.v1.ExecSandboxStderr - (*ExecSandboxExit)(nil), // 100: openshell.v1.ExecSandboxExit - (*ExecSandboxEvent)(nil), // 101: openshell.v1.ExecSandboxEvent - (*TcpForwardInit)(nil), // 102: openshell.v1.TcpForwardInit - (*TcpForwardFrame)(nil), // 103: openshell.v1.TcpForwardFrame - (*ExecSandboxInput)(nil), // 104: openshell.v1.ExecSandboxInput - (*ExecSandboxWindowResize)(nil), // 105: openshell.v1.ExecSandboxWindowResize - (*SshSession)(nil), // 106: openshell.v1.SshSession - (*WatchSandboxRequest)(nil), // 107: openshell.v1.WatchSandboxRequest - (*SandboxStreamEvent)(nil), // 108: openshell.v1.SandboxStreamEvent - (*SandboxLogLine)(nil), // 109: openshell.v1.SandboxLogLine - (*SandboxStreamWarning)(nil), // 110: openshell.v1.SandboxStreamWarning - (*CreateProviderRequest)(nil), // 111: openshell.v1.CreateProviderRequest - (*GetProviderRequest)(nil), // 112: openshell.v1.GetProviderRequest - (*ListProvidersRequest)(nil), // 113: openshell.v1.ListProvidersRequest - (*UpdateProviderRequest)(nil), // 114: openshell.v1.UpdateProviderRequest - (*DeleteProviderRequest)(nil), // 115: openshell.v1.DeleteProviderRequest - (*ProviderResponse)(nil), // 116: openshell.v1.ProviderResponse - (*ListProvidersResponse)(nil), // 117: openshell.v1.ListProvidersResponse - (*ListProviderProfilesRequest)(nil), // 118: openshell.v1.ListProviderProfilesRequest - (*GetProviderProfileRequest)(nil), // 119: openshell.v1.GetProviderProfileRequest - (*ProviderProfileImportItem)(nil), // 120: openshell.v1.ProviderProfileImportItem - (*ProviderProfileDiagnostic)(nil), // 121: openshell.v1.ProviderProfileDiagnostic - (*ProviderCredentialTokenGrantAudienceOverride)(nil), // 122: openshell.v1.ProviderCredentialTokenGrantAudienceOverride - (*ProviderCredentialTokenGrantSubjectToken)(nil), // 123: openshell.v1.ProviderCredentialTokenGrantSubjectToken - (*ProviderCredentialTokenGrant)(nil), // 124: openshell.v1.ProviderCredentialTokenGrant - (*ProviderProfileCredential)(nil), // 125: openshell.v1.ProviderProfileCredential - (*ProviderCredentialRefreshMaterial)(nil), // 126: openshell.v1.ProviderCredentialRefreshMaterial - (*ProviderCredentialRefreshOutput)(nil), // 127: openshell.v1.ProviderCredentialRefreshOutput - (*ProviderCredentialRefresh)(nil), // 128: openshell.v1.ProviderCredentialRefresh - (*ProviderCredentialRefreshStatus)(nil), // 129: openshell.v1.ProviderCredentialRefreshStatus - (*ProviderProfileDiscovery)(nil), // 130: openshell.v1.ProviderProfileDiscovery - (*GetProviderRefreshStatusRequest)(nil), // 131: openshell.v1.GetProviderRefreshStatusRequest - (*GetProviderRefreshStatusResponse)(nil), // 132: openshell.v1.GetProviderRefreshStatusResponse - (*ConfigureProviderRefreshRequest)(nil), // 133: openshell.v1.ConfigureProviderRefreshRequest - (*ConfigureProviderRefreshResponse)(nil), // 134: openshell.v1.ConfigureProviderRefreshResponse - (*RotateProviderCredentialRequest)(nil), // 135: openshell.v1.RotateProviderCredentialRequest - (*RotateProviderCredentialResponse)(nil), // 136: openshell.v1.RotateProviderCredentialResponse - (*DeleteProviderRefreshRequest)(nil), // 137: openshell.v1.DeleteProviderRefreshRequest - (*DeleteProviderRefreshResponse)(nil), // 138: openshell.v1.DeleteProviderRefreshResponse - (*ProviderProfile)(nil), // 139: openshell.v1.ProviderProfile - (*ProviderProfileResponse)(nil), // 140: openshell.v1.ProviderProfileResponse - (*ListProviderProfilesResponse)(nil), // 141: openshell.v1.ListProviderProfilesResponse - (*ImportProviderProfilesRequest)(nil), // 142: openshell.v1.ImportProviderProfilesRequest - (*ImportProviderProfilesResponse)(nil), // 143: openshell.v1.ImportProviderProfilesResponse - (*UpdateProviderProfilesRequest)(nil), // 144: openshell.v1.UpdateProviderProfilesRequest - (*UpdateProviderProfilesResponse)(nil), // 145: openshell.v1.UpdateProviderProfilesResponse - (*LintProviderProfilesRequest)(nil), // 146: openshell.v1.LintProviderProfilesRequest - (*LintProviderProfilesResponse)(nil), // 147: openshell.v1.LintProviderProfilesResponse - (*DeleteProviderResponse)(nil), // 148: openshell.v1.DeleteProviderResponse - (*DeleteProviderProfileRequest)(nil), // 149: openshell.v1.DeleteProviderProfileRequest - (*DeleteProviderProfileResponse)(nil), // 150: openshell.v1.DeleteProviderProfileResponse - (*GetSandboxProviderEnvironmentRequest)(nil), // 151: openshell.v1.GetSandboxProviderEnvironmentRequest - (*StaticCredentialEndpointBinding)(nil), // 152: openshell.v1.StaticCredentialEndpointBinding - (*StaticCredentialBinding)(nil), // 153: openshell.v1.StaticCredentialBinding - (*GetSandboxProviderEnvironmentResponse)(nil), // 154: openshell.v1.GetSandboxProviderEnvironmentResponse - (*ExchangeProviderSubjectTokenRequest)(nil), // 155: openshell.v1.ExchangeProviderSubjectTokenRequest - (*ExchangeProviderSubjectTokenResponse)(nil), // 156: openshell.v1.ExchangeProviderSubjectTokenResponse - (*UpdateConfigRequest)(nil), // 157: openshell.v1.UpdateConfigRequest - (*PolicyMergeOperation)(nil), // 158: openshell.v1.PolicyMergeOperation - (*AddNetworkRule)(nil), // 159: openshell.v1.AddNetworkRule - (*RemoveNetworkEndpoint)(nil), // 160: openshell.v1.RemoveNetworkEndpoint - (*RemoveNetworkRule)(nil), // 161: openshell.v1.RemoveNetworkRule - (*L7RuleTarget)(nil), // 162: openshell.v1.L7RuleTarget - (*AddDenyRules)(nil), // 163: openshell.v1.AddDenyRules - (*AddAllowRules)(nil), // 164: openshell.v1.AddAllowRules - (*RemoveNetworkBinary)(nil), // 165: openshell.v1.RemoveNetworkBinary - (*UpdateConfigResponse)(nil), // 166: openshell.v1.UpdateConfigResponse - (*GetSandboxPolicyStatusRequest)(nil), // 167: openshell.v1.GetSandboxPolicyStatusRequest - (*GetSandboxPolicyStatusResponse)(nil), // 168: openshell.v1.GetSandboxPolicyStatusResponse - (*ListSandboxPoliciesRequest)(nil), // 169: openshell.v1.ListSandboxPoliciesRequest - (*ListSandboxPoliciesResponse)(nil), // 170: openshell.v1.ListSandboxPoliciesResponse - (*ReportPolicyStatusRequest)(nil), // 171: openshell.v1.ReportPolicyStatusRequest - (*ReportPolicyStatusResponse)(nil), // 172: openshell.v1.ReportPolicyStatusResponse - (*SandboxConfigurationAdmission)(nil), // 173: openshell.v1.SandboxConfigurationAdmission - (*ReportSandboxConfigurationRequest)(nil), // 174: openshell.v1.ReportSandboxConfigurationRequest - (*ReportSandboxConfigurationResponse)(nil), // 175: openshell.v1.ReportSandboxConfigurationResponse - (*SandboxPolicyRevision)(nil), // 176: openshell.v1.SandboxPolicyRevision - (*GetSandboxLogsRequest)(nil), // 177: openshell.v1.GetSandboxLogsRequest - (*PushSandboxLogsRequest)(nil), // 178: openshell.v1.PushSandboxLogsRequest - (*PushSandboxLogsResponse)(nil), // 179: openshell.v1.PushSandboxLogsResponse - (*GetSandboxLogsResponse)(nil), // 180: openshell.v1.GetSandboxLogsResponse - (*SupervisorMessage)(nil), // 181: openshell.v1.SupervisorMessage - (*GatewayMessage)(nil), // 182: openshell.v1.GatewayMessage - (*SupervisorHello)(nil), // 183: openshell.v1.SupervisorHello - (*SessionAccepted)(nil), // 184: openshell.v1.SessionAccepted - (*SessionRejected)(nil), // 185: openshell.v1.SessionRejected - (*SupervisorHeartbeat)(nil), // 186: openshell.v1.SupervisorHeartbeat - (*GatewayHeartbeat)(nil), // 187: openshell.v1.GatewayHeartbeat - (*ReportMainProcessExitRequest)(nil), // 188: openshell.v1.ReportMainProcessExitRequest - (*ReportMainProcessExitResponse)(nil), // 189: openshell.v1.ReportMainProcessExitResponse - (*FinalizeMainProcessExitRequest)(nil), // 190: openshell.v1.FinalizeMainProcessExitRequest - (*FinalizeMainProcessExitResponse)(nil), // 191: openshell.v1.FinalizeMainProcessExitResponse - (*RelayOpen)(nil), // 192: openshell.v1.RelayOpen - (*SshRelayTarget)(nil), // 193: openshell.v1.SshRelayTarget - (*TcpRelayTarget)(nil), // 194: openshell.v1.TcpRelayTarget - (*RelayInit)(nil), // 195: openshell.v1.RelayInit - (*RelayFrame)(nil), // 196: openshell.v1.RelayFrame - (*PeerRelayInit)(nil), // 197: openshell.v1.PeerRelayInit - (*PeerRelayFrame)(nil), // 198: openshell.v1.PeerRelayFrame - (*RelayOpenResult)(nil), // 199: openshell.v1.RelayOpenResult - (*RelayClose)(nil), // 200: openshell.v1.RelayClose - (*L7RequestSample)(nil), // 201: openshell.v1.L7RequestSample - (*DenialSummary)(nil), // 202: openshell.v1.DenialSummary - (*DenialGroupCount)(nil), // 203: openshell.v1.DenialGroupCount - (*NetworkActivitySummary)(nil), // 204: openshell.v1.NetworkActivitySummary - (*PolicyChunk)(nil), // 205: openshell.v1.PolicyChunk - (*DraftPolicyUpdate)(nil), // 206: openshell.v1.DraftPolicyUpdate - (*SubmitPolicyAnalysisRequest)(nil), // 207: openshell.v1.SubmitPolicyAnalysisRequest - (*SubmitPolicyAnalysisResponse)(nil), // 208: openshell.v1.SubmitPolicyAnalysisResponse - (*GetDraftPolicyRequest)(nil), // 209: openshell.v1.GetDraftPolicyRequest - (*GetDraftPolicyResponse)(nil), // 210: openshell.v1.GetDraftPolicyResponse - (*ApproveDraftChunkRequest)(nil), // 211: openshell.v1.ApproveDraftChunkRequest - (*ApproveDraftChunkResponse)(nil), // 212: openshell.v1.ApproveDraftChunkResponse - (*RejectDraftChunkRequest)(nil), // 213: openshell.v1.RejectDraftChunkRequest - (*RejectDraftChunkResponse)(nil), // 214: openshell.v1.RejectDraftChunkResponse - (*DraftChunkApproval)(nil), // 215: openshell.v1.DraftChunkApproval - (*ApproveAllDraftChunksRequest)(nil), // 216: openshell.v1.ApproveAllDraftChunksRequest - (*ApproveAllDraftChunksResponse)(nil), // 217: openshell.v1.ApproveAllDraftChunksResponse - (*EditDraftChunkRequest)(nil), // 218: openshell.v1.EditDraftChunkRequest - (*EditDraftChunkResponse)(nil), // 219: openshell.v1.EditDraftChunkResponse - (*UndoDraftChunkRequest)(nil), // 220: openshell.v1.UndoDraftChunkRequest - (*UndoDraftChunkResponse)(nil), // 221: openshell.v1.UndoDraftChunkResponse - (*ClearDraftChunksRequest)(nil), // 222: openshell.v1.ClearDraftChunksRequest - (*ClearDraftChunksResponse)(nil), // 223: openshell.v1.ClearDraftChunksResponse - (*GetDraftHistoryRequest)(nil), // 224: openshell.v1.GetDraftHistoryRequest - (*DraftHistoryEntry)(nil), // 225: openshell.v1.DraftHistoryEntry - (*GetDraftHistoryResponse)(nil), // 226: openshell.v1.GetDraftHistoryResponse - (*CreateWorkspaceRequest)(nil), // 227: openshell.v1.CreateWorkspaceRequest - (*CreateWorkspaceResponse)(nil), // 228: openshell.v1.CreateWorkspaceResponse - (*GetWorkspaceRequest)(nil), // 229: openshell.v1.GetWorkspaceRequest - (*GetWorkspaceResponse)(nil), // 230: openshell.v1.GetWorkspaceResponse - (*ListWorkspacesRequest)(nil), // 231: openshell.v1.ListWorkspacesRequest - (*ListWorkspacesResponse)(nil), // 232: openshell.v1.ListWorkspacesResponse - (*DeleteWorkspaceRequest)(nil), // 233: openshell.v1.DeleteWorkspaceRequest - (*DeleteWorkspaceResponse)(nil), // 234: openshell.v1.DeleteWorkspaceResponse - (*WorkspaceMember)(nil), // 235: openshell.v1.WorkspaceMember - (*AddWorkspaceMemberRequest)(nil), // 236: openshell.v1.AddWorkspaceMemberRequest - (*AddWorkspaceMemberResponse)(nil), // 237: openshell.v1.AddWorkspaceMemberResponse - (*RemoveWorkspaceMemberRequest)(nil), // 238: openshell.v1.RemoveWorkspaceMemberRequest - (*RemoveWorkspaceMemberResponse)(nil), // 239: openshell.v1.RemoveWorkspaceMemberResponse - (*ListWorkspaceMembersRequest)(nil), // 240: openshell.v1.ListWorkspaceMembersRequest - (*ListWorkspaceMembersResponse)(nil), // 241: openshell.v1.ListWorkspaceMembersResponse - (*ExtensionServiceCredential)(nil), // 242: openshell.v1.ExtensionServiceCredential - (*EndpointObservation)(nil), // 243: openshell.v1.EndpointObservation - (*ReportEndpointStatusRequest)(nil), // 244: openshell.v1.ReportEndpointStatusRequest - (*ReportEndpointStatusResponse)(nil), // 245: openshell.v1.ReportEndpointStatusResponse - (*EndpointStatus)(nil), // 246: openshell.v1.EndpointStatus - (*SandboxProvisioning)(nil), // 247: openshell.v1.SandboxProvisioning - (*SandboxServiceExposure)(nil), // 248: openshell.v1.SandboxServiceExposure - nil, // 249: openshell.v1.SandboxSpec.EnvironmentEntry - nil, // 250: openshell.v1.SandboxTemplate.LabelsEntry - nil, // 251: openshell.v1.SandboxTemplate.AnnotationsEntry - nil, // 252: openshell.v1.SandboxTemplate.EnvironmentEntry - nil, // 253: openshell.v1.SandboxWorkloadConfig.EnvironmentEntry - nil, // 254: openshell.v1.PlatformEvent.MetadataEntry - nil, // 255: openshell.v1.CreateSandboxRequest.LabelsEntry - nil, // 256: openshell.v1.CreateSandboxRequest.AnnotationsEntry - nil, // 257: openshell.v1.SandboxResponse.ServiceUrlsEntry - nil, // 258: openshell.v1.ExecSandboxRequest.EnvironmentEntry - nil, // 259: openshell.v1.SandboxLogLine.FieldsEntry - nil, // 260: openshell.v1.UpdateProviderRequest.CredentialExpirationTimesEntry - nil, // 261: openshell.v1.ConfigureProviderRefreshRequest.MaterialEntry - nil, // 262: openshell.v1.ProviderProfile.AnnotationsEntry - nil, // 263: openshell.v1.GetSandboxProviderEnvironmentResponse.EnvironmentEntry - nil, // 264: openshell.v1.GetSandboxProviderEnvironmentResponse.CredentialExpirationTimesEntry - nil, // 265: openshell.v1.GetSandboxProviderEnvironmentResponse.DynamicCredentialsEntry - nil, // 266: openshell.v1.GetSandboxProviderEnvironmentResponse.StaticCredentialBindingsEntry - nil, // 267: openshell.v1.UpdateConfigRequest.AnnotationsEntry - nil, // 268: openshell.v1.UpdateConfigResponse.AnnotationsEntry - nil, // 269: openshell.v1.SandboxPolicyRevision.ProvenanceEntry - nil, // 270: openshell.v1.CreateWorkspaceRequest.LabelsEntry - (*timestamppb.Timestamp)(nil), // 271: google.protobuf.Timestamp - (*datamodelv1.ObjectMeta)(nil), // 272: openshell.datamodel.v1.ObjectMeta - (*sandboxv1.SandboxPolicy)(nil), // 273: openshell.sandbox.v1.SandboxPolicy - (*structpb.Struct)(nil), // 274: google.protobuf.Struct - (*durationpb.Duration)(nil), // 275: google.protobuf.Duration - (*datamodelv1.WorkspaceSelector)(nil), // 276: openshell.datamodel.v1.WorkspaceSelector - (*datamodelv1.Provider)(nil), // 277: openshell.datamodel.v1.Provider - (sandboxv1.PolicySource)(0), // 278: openshell.sandbox.v1.PolicySource - (*sandboxv1.NetworkEndpoint)(nil), // 279: openshell.sandbox.v1.NetworkEndpoint - (*sandboxv1.NetworkBinary)(nil), // 280: openshell.sandbox.v1.NetworkBinary - (*sandboxv1.SettingValue)(nil), // 281: openshell.sandbox.v1.SettingValue - (*sandboxv1.NetworkPolicyRule)(nil), // 282: openshell.sandbox.v1.NetworkPolicyRule - (*sandboxv1.L7DenyRule)(nil), // 283: openshell.sandbox.v1.L7DenyRule - (*sandboxv1.L7Rule)(nil), // 284: openshell.sandbox.v1.L7Rule - (*datamodelv1.Workspace)(nil), // 285: openshell.datamodel.v1.Workspace - (*sandboxv1.GetSandboxConfigRequest)(nil), // 286: openshell.sandbox.v1.GetSandboxConfigRequest - (*sandboxv1.GetGatewayConfigRequest)(nil), // 287: openshell.sandbox.v1.GetGatewayConfigRequest - (*sandboxv1.GetSandboxConfigResponse)(nil), // 288: openshell.sandbox.v1.GetSandboxConfigResponse - (*sandboxv1.GetGatewayConfigResponse)(nil), // 289: openshell.sandbox.v1.GetGatewayConfigResponse + (SandboxRestartPolicy)(0), // 16: openshell.v1.SandboxRestartPolicy + (DeletionOutcome)(0), // 17: openshell.v1.DeletionOutcome + (EndpointResult)(0), // 18: openshell.v1.EndpointResult + (ServiceAuthorizationMode)(0), // 19: openshell.v1.ServiceAuthorizationMode + (*IssueSandboxTokenRequest)(nil), // 20: openshell.v1.IssueSandboxTokenRequest + (*IssueSandboxTokenResponse)(nil), // 21: openshell.v1.IssueSandboxTokenResponse + (*RefreshSandboxTokenRequest)(nil), // 22: openshell.v1.RefreshSandboxTokenRequest + (*RefreshSandboxTokenResponse)(nil), // 23: openshell.v1.RefreshSandboxTokenResponse + (*HealthRequest)(nil), // 24: openshell.v1.HealthRequest + (*HealthResponse)(nil), // 25: openshell.v1.HealthResponse + (*GetCurrentUserRequest)(nil), // 26: openshell.v1.GetCurrentUserRequest + (*GetCurrentUserResponse)(nil), // 27: openshell.v1.GetCurrentUserResponse + (*GetGatewayInfoRequest)(nil), // 28: openshell.v1.GetGatewayInfoRequest + (*GetGatewayInfoResponse)(nil), // 29: openshell.v1.GetGatewayInfoResponse + (*NegotiatedExtensionInfo)(nil), // 30: openshell.v1.NegotiatedExtensionInfo + (*ComputeDriverInfo)(nil), // 31: openshell.v1.ComputeDriverInfo + (*ComputeDriverCapabilities)(nil), // 32: openshell.v1.ComputeDriverCapabilities + (*ResourceCapabilities)(nil), // 33: openshell.v1.ResourceCapabilities + (*CpuResourceCapabilities)(nil), // 34: openshell.v1.CpuResourceCapabilities + (*MemoryResourceCapabilities)(nil), // 35: openshell.v1.MemoryResourceCapabilities + (*GpuResourceCapabilities)(nil), // 36: openshell.v1.GpuResourceCapabilities + (*Sandbox)(nil), // 37: openshell.v1.Sandbox + (*SandboxSpec)(nil), // 38: openshell.v1.SandboxSpec + (*ResourceRequirements)(nil), // 39: openshell.v1.ResourceRequirements + (*GpuResourceRequirements)(nil), // 40: openshell.v1.GpuResourceRequirements + (*SandboxTemplate)(nil), // 41: openshell.v1.SandboxTemplate + (*SandboxWorkloadTemplate)(nil), // 42: openshell.v1.SandboxWorkloadTemplate + (*SandboxWorkloadTemplateSpec)(nil), // 43: openshell.v1.SandboxWorkloadTemplateSpec + (*SandboxWorkloadConfig)(nil), // 44: openshell.v1.SandboxWorkloadConfig + (*SandboxResources)(nil), // 45: openshell.v1.SandboxResources + (*SandboxServiceLevel)(nil), // 46: openshell.v1.SandboxServiceLevel + (*SandboxStartup)(nil), // 47: openshell.v1.SandboxStartup + (*SandboxWorkloadTemplateProvenance)(nil), // 48: openshell.v1.SandboxWorkloadTemplateProvenance + (*SandboxStatus)(nil), // 49: openshell.v1.SandboxStatus + (*SandboxCondition)(nil), // 50: openshell.v1.SandboxCondition + (*PlatformEvent)(nil), // 51: openshell.v1.PlatformEvent + (*CreateSandboxRequest)(nil), // 52: openshell.v1.CreateSandboxRequest + (*CreateSandboxTemplateRequest)(nil), // 53: openshell.v1.CreateSandboxTemplateRequest + (*GetSandboxTemplateRequest)(nil), // 54: openshell.v1.GetSandboxTemplateRequest + (*ListSandboxTemplatesRequest)(nil), // 55: openshell.v1.ListSandboxTemplatesRequest + (*DeleteSandboxTemplateRequest)(nil), // 56: openshell.v1.DeleteSandboxTemplateRequest + (*SandboxTemplateResponse)(nil), // 57: openshell.v1.SandboxTemplateResponse + (*ListSandboxTemplatesResponse)(nil), // 58: openshell.v1.ListSandboxTemplatesResponse + (*DeleteSandboxTemplateResponse)(nil), // 59: openshell.v1.DeleteSandboxTemplateResponse + (*BeginRootfsTarStagingRequest)(nil), // 60: openshell.v1.BeginRootfsTarStagingRequest + (*BeginRootfsTarStagingResponse)(nil), // 61: openshell.v1.BeginRootfsTarStagingResponse + (*GetSandboxRequest)(nil), // 62: openshell.v1.GetSandboxRequest + (*ListSandboxesRequest)(nil), // 63: openshell.v1.ListSandboxesRequest + (*ListSandboxProvidersRequest)(nil), // 64: openshell.v1.ListSandboxProvidersRequest + (*AttachSandboxProviderRequest)(nil), // 65: openshell.v1.AttachSandboxProviderRequest + (*DetachSandboxProviderRequest)(nil), // 66: openshell.v1.DetachSandboxProviderRequest + (*DeleteSandboxRequest)(nil), // 67: openshell.v1.DeleteSandboxRequest + (*StopSandboxRequest)(nil), // 68: openshell.v1.StopSandboxRequest + (*StartSandboxRequest)(nil), // 69: openshell.v1.StartSandboxRequest + (*SandboxResponse)(nil), // 70: openshell.v1.SandboxResponse + (*ListSandboxesResponse)(nil), // 71: openshell.v1.ListSandboxesResponse + (*ListSandboxProvidersResponse)(nil), // 72: openshell.v1.ListSandboxProvidersResponse + (*AttachSandboxProviderResponse)(nil), // 73: openshell.v1.AttachSandboxProviderResponse + (*DetachSandboxProviderResponse)(nil), // 74: openshell.v1.DetachSandboxProviderResponse + (*ProviderDesiredIdentity)(nil), // 75: openshell.v1.ProviderDesiredIdentity + (*ConfigSnapshotRevision)(nil), // 76: openshell.v1.ConfigSnapshotRevision + (*SandboxConfigRevision)(nil), // 77: openshell.v1.SandboxConfigRevision + (*ConfigUpdateOperation)(nil), // 78: openshell.v1.ConfigUpdateOperation + (*ProviderMutationReceipt)(nil), // 79: openshell.v1.ProviderMutationReceipt + (*ProviderReadinessObservation)(nil), // 80: openshell.v1.ProviderReadinessObservation + (*ProviderReadinessStatus)(nil), // 81: openshell.v1.ProviderReadinessStatus + (*GetSandboxProviderStatusRequest)(nil), // 82: openshell.v1.GetSandboxProviderStatusRequest + (*GetSandboxProviderStatusResponse)(nil), // 83: openshell.v1.GetSandboxProviderStatusResponse + (*ReportProviderReadinessRequest)(nil), // 84: openshell.v1.ReportProviderReadinessRequest + (*ReportProviderReadinessResponse)(nil), // 85: openshell.v1.ReportProviderReadinessResponse + (*DeleteSandboxResponse)(nil), // 86: openshell.v1.DeleteSandboxResponse + (*CreateSshSessionRequest)(nil), // 87: openshell.v1.CreateSshSessionRequest + (*CreateSshSessionResponse)(nil), // 88: openshell.v1.CreateSshSessionResponse + (*ExposeServiceRequest)(nil), // 89: openshell.v1.ExposeServiceRequest + (*GetServiceRequest)(nil), // 90: openshell.v1.GetServiceRequest + (*ListServicesRequest)(nil), // 91: openshell.v1.ListServicesRequest + (*ListServicesResponse)(nil), // 92: openshell.v1.ListServicesResponse + (*DeleteServiceRequest)(nil), // 93: openshell.v1.DeleteServiceRequest + (*DeleteServiceResponse)(nil), // 94: openshell.v1.DeleteServiceResponse + (*ServiceEndpoint)(nil), // 95: openshell.v1.ServiceEndpoint + (*ServiceEndpointResponse)(nil), // 96: openshell.v1.ServiceEndpointResponse + (*RevokeSshSessionRequest)(nil), // 97: openshell.v1.RevokeSshSessionRequest + (*RevokeSshSessionResponse)(nil), // 98: openshell.v1.RevokeSshSessionResponse + (*ExecSandboxRequest)(nil), // 99: openshell.v1.ExecSandboxRequest + (*ExecSandboxStdout)(nil), // 100: openshell.v1.ExecSandboxStdout + (*ExecSandboxStderr)(nil), // 101: openshell.v1.ExecSandboxStderr + (*ExecSandboxExit)(nil), // 102: openshell.v1.ExecSandboxExit + (*ExecSandboxEvent)(nil), // 103: openshell.v1.ExecSandboxEvent + (*TcpForwardInit)(nil), // 104: openshell.v1.TcpForwardInit + (*TcpForwardFrame)(nil), // 105: openshell.v1.TcpForwardFrame + (*ExecSandboxInput)(nil), // 106: openshell.v1.ExecSandboxInput + (*ExecSandboxWindowResize)(nil), // 107: openshell.v1.ExecSandboxWindowResize + (*SshSession)(nil), // 108: openshell.v1.SshSession + (*WatchSandboxRequest)(nil), // 109: openshell.v1.WatchSandboxRequest + (*SandboxStreamEvent)(nil), // 110: openshell.v1.SandboxStreamEvent + (*SandboxLogLine)(nil), // 111: openshell.v1.SandboxLogLine + (*SandboxStreamWarning)(nil), // 112: openshell.v1.SandboxStreamWarning + (*CreateProviderRequest)(nil), // 113: openshell.v1.CreateProviderRequest + (*GetProviderRequest)(nil), // 114: openshell.v1.GetProviderRequest + (*ListProvidersRequest)(nil), // 115: openshell.v1.ListProvidersRequest + (*UpdateProviderRequest)(nil), // 116: openshell.v1.UpdateProviderRequest + (*DeleteProviderRequest)(nil), // 117: openshell.v1.DeleteProviderRequest + (*ProviderResponse)(nil), // 118: openshell.v1.ProviderResponse + (*ListProvidersResponse)(nil), // 119: openshell.v1.ListProvidersResponse + (*ListProviderProfilesRequest)(nil), // 120: openshell.v1.ListProviderProfilesRequest + (*GetProviderProfileRequest)(nil), // 121: openshell.v1.GetProviderProfileRequest + (*ProviderProfileImportItem)(nil), // 122: openshell.v1.ProviderProfileImportItem + (*ProviderProfileDiagnostic)(nil), // 123: openshell.v1.ProviderProfileDiagnostic + (*ProviderCredentialTokenGrantAudienceOverride)(nil), // 124: openshell.v1.ProviderCredentialTokenGrantAudienceOverride + (*ProviderCredentialTokenGrantSubjectToken)(nil), // 125: openshell.v1.ProviderCredentialTokenGrantSubjectToken + (*ProviderCredentialTokenGrant)(nil), // 126: openshell.v1.ProviderCredentialTokenGrant + (*ProviderProfileCredential)(nil), // 127: openshell.v1.ProviderProfileCredential + (*ProviderCredentialRefreshMaterial)(nil), // 128: openshell.v1.ProviderCredentialRefreshMaterial + (*ProviderCredentialRefreshOutput)(nil), // 129: openshell.v1.ProviderCredentialRefreshOutput + (*ProviderCredentialRefresh)(nil), // 130: openshell.v1.ProviderCredentialRefresh + (*ProviderCredentialRefreshStatus)(nil), // 131: openshell.v1.ProviderCredentialRefreshStatus + (*ProviderProfileDiscovery)(nil), // 132: openshell.v1.ProviderProfileDiscovery + (*GetProviderRefreshStatusRequest)(nil), // 133: openshell.v1.GetProviderRefreshStatusRequest + (*GetProviderRefreshStatusResponse)(nil), // 134: openshell.v1.GetProviderRefreshStatusResponse + (*ConfigureProviderRefreshRequest)(nil), // 135: openshell.v1.ConfigureProviderRefreshRequest + (*ConfigureProviderRefreshResponse)(nil), // 136: openshell.v1.ConfigureProviderRefreshResponse + (*RotateProviderCredentialRequest)(nil), // 137: openshell.v1.RotateProviderCredentialRequest + (*RotateProviderCredentialResponse)(nil), // 138: openshell.v1.RotateProviderCredentialResponse + (*DeleteProviderRefreshRequest)(nil), // 139: openshell.v1.DeleteProviderRefreshRequest + (*DeleteProviderRefreshResponse)(nil), // 140: openshell.v1.DeleteProviderRefreshResponse + (*ProviderProfile)(nil), // 141: openshell.v1.ProviderProfile + (*ProviderProfileFile)(nil), // 142: openshell.v1.ProviderProfileFile + (*ProviderProfileResponse)(nil), // 143: openshell.v1.ProviderProfileResponse + (*ListProviderProfilesResponse)(nil), // 144: openshell.v1.ListProviderProfilesResponse + (*ImportProviderProfilesRequest)(nil), // 145: openshell.v1.ImportProviderProfilesRequest + (*ImportProviderProfilesResponse)(nil), // 146: openshell.v1.ImportProviderProfilesResponse + (*UpdateProviderProfilesRequest)(nil), // 147: openshell.v1.UpdateProviderProfilesRequest + (*UpdateProviderProfilesResponse)(nil), // 148: openshell.v1.UpdateProviderProfilesResponse + (*LintProviderProfilesRequest)(nil), // 149: openshell.v1.LintProviderProfilesRequest + (*LintProviderProfilesResponse)(nil), // 150: openshell.v1.LintProviderProfilesResponse + (*DeleteProviderResponse)(nil), // 151: openshell.v1.DeleteProviderResponse + (*DeleteProviderProfileRequest)(nil), // 152: openshell.v1.DeleteProviderProfileRequest + (*DeleteProviderProfileResponse)(nil), // 153: openshell.v1.DeleteProviderProfileResponse + (*GetSandboxProviderEnvironmentRequest)(nil), // 154: openshell.v1.GetSandboxProviderEnvironmentRequest + (*StaticCredentialEndpointBinding)(nil), // 155: openshell.v1.StaticCredentialEndpointBinding + (*StaticCredentialBinding)(nil), // 156: openshell.v1.StaticCredentialBinding + (*GetSandboxProviderEnvironmentResponse)(nil), // 157: openshell.v1.GetSandboxProviderEnvironmentResponse + (*ExchangeProviderSubjectTokenRequest)(nil), // 158: openshell.v1.ExchangeProviderSubjectTokenRequest + (*ExchangeProviderSubjectTokenResponse)(nil), // 159: openshell.v1.ExchangeProviderSubjectTokenResponse + (*UpdateConfigRequest)(nil), // 160: openshell.v1.UpdateConfigRequest + (*PolicyMergeOperation)(nil), // 161: openshell.v1.PolicyMergeOperation + (*AddNetworkRule)(nil), // 162: openshell.v1.AddNetworkRule + (*RemoveNetworkEndpoint)(nil), // 163: openshell.v1.RemoveNetworkEndpoint + (*RemoveNetworkRule)(nil), // 164: openshell.v1.RemoveNetworkRule + (*L7RuleTarget)(nil), // 165: openshell.v1.L7RuleTarget + (*AddDenyRules)(nil), // 166: openshell.v1.AddDenyRules + (*AddAllowRules)(nil), // 167: openshell.v1.AddAllowRules + (*RemoveNetworkBinary)(nil), // 168: openshell.v1.RemoveNetworkBinary + (*UpdateConfigResponse)(nil), // 169: openshell.v1.UpdateConfigResponse + (*GetSandboxPolicyStatusRequest)(nil), // 170: openshell.v1.GetSandboxPolicyStatusRequest + (*GetSandboxPolicyStatusResponse)(nil), // 171: openshell.v1.GetSandboxPolicyStatusResponse + (*ListSandboxPoliciesRequest)(nil), // 172: openshell.v1.ListSandboxPoliciesRequest + (*ListSandboxPoliciesResponse)(nil), // 173: openshell.v1.ListSandboxPoliciesResponse + (*ReportPolicyStatusRequest)(nil), // 174: openshell.v1.ReportPolicyStatusRequest + (*ReportPolicyStatusResponse)(nil), // 175: openshell.v1.ReportPolicyStatusResponse + (*SandboxConfigurationAdmission)(nil), // 176: openshell.v1.SandboxConfigurationAdmission + (*ReportSandboxConfigurationRequest)(nil), // 177: openshell.v1.ReportSandboxConfigurationRequest + (*ReportSandboxConfigurationResponse)(nil), // 178: openshell.v1.ReportSandboxConfigurationResponse + (*SandboxPolicyRevision)(nil), // 179: openshell.v1.SandboxPolicyRevision + (*GetSandboxLogsRequest)(nil), // 180: openshell.v1.GetSandboxLogsRequest + (*PushSandboxLogsRequest)(nil), // 181: openshell.v1.PushSandboxLogsRequest + (*PushSandboxLogsResponse)(nil), // 182: openshell.v1.PushSandboxLogsResponse + (*GetSandboxLogsResponse)(nil), // 183: openshell.v1.GetSandboxLogsResponse + (*SupervisorMessage)(nil), // 184: openshell.v1.SupervisorMessage + (*GatewayMessage)(nil), // 185: openshell.v1.GatewayMessage + (*SupervisorHello)(nil), // 186: openshell.v1.SupervisorHello + (*SessionAccepted)(nil), // 187: openshell.v1.SessionAccepted + (*SessionRejected)(nil), // 188: openshell.v1.SessionRejected + (*SupervisorHeartbeat)(nil), // 189: openshell.v1.SupervisorHeartbeat + (*GatewayHeartbeat)(nil), // 190: openshell.v1.GatewayHeartbeat + (*ReportMainProcessExitRequest)(nil), // 191: openshell.v1.ReportMainProcessExitRequest + (*ReportMainProcessExitResponse)(nil), // 192: openshell.v1.ReportMainProcessExitResponse + (*FinalizeMainProcessExitRequest)(nil), // 193: openshell.v1.FinalizeMainProcessExitRequest + (*FinalizeMainProcessExitResponse)(nil), // 194: openshell.v1.FinalizeMainProcessExitResponse + (*RelayOpen)(nil), // 195: openshell.v1.RelayOpen + (*SshRelayTarget)(nil), // 196: openshell.v1.SshRelayTarget + (*TcpRelayTarget)(nil), // 197: openshell.v1.TcpRelayTarget + (*RelayInit)(nil), // 198: openshell.v1.RelayInit + (*RelayFrame)(nil), // 199: openshell.v1.RelayFrame + (*PeerRelayInit)(nil), // 200: openshell.v1.PeerRelayInit + (*PeerRelayFrame)(nil), // 201: openshell.v1.PeerRelayFrame + (*RelayOpenResult)(nil), // 202: openshell.v1.RelayOpenResult + (*RelayClose)(nil), // 203: openshell.v1.RelayClose + (*L7RequestSample)(nil), // 204: openshell.v1.L7RequestSample + (*DenialSummary)(nil), // 205: openshell.v1.DenialSummary + (*DenialGroupCount)(nil), // 206: openshell.v1.DenialGroupCount + (*NetworkActivitySummary)(nil), // 207: openshell.v1.NetworkActivitySummary + (*PolicyChunk)(nil), // 208: openshell.v1.PolicyChunk + (*DraftPolicyUpdate)(nil), // 209: openshell.v1.DraftPolicyUpdate + (*SubmitPolicyAnalysisRequest)(nil), // 210: openshell.v1.SubmitPolicyAnalysisRequest + (*SubmitPolicyAnalysisResponse)(nil), // 211: openshell.v1.SubmitPolicyAnalysisResponse + (*GetDraftPolicyRequest)(nil), // 212: openshell.v1.GetDraftPolicyRequest + (*GetDraftPolicyResponse)(nil), // 213: openshell.v1.GetDraftPolicyResponse + (*ApproveDraftChunkRequest)(nil), // 214: openshell.v1.ApproveDraftChunkRequest + (*ApproveDraftChunkResponse)(nil), // 215: openshell.v1.ApproveDraftChunkResponse + (*RejectDraftChunkRequest)(nil), // 216: openshell.v1.RejectDraftChunkRequest + (*RejectDraftChunkResponse)(nil), // 217: openshell.v1.RejectDraftChunkResponse + (*DraftChunkApproval)(nil), // 218: openshell.v1.DraftChunkApproval + (*ApproveAllDraftChunksRequest)(nil), // 219: openshell.v1.ApproveAllDraftChunksRequest + (*ApproveAllDraftChunksResponse)(nil), // 220: openshell.v1.ApproveAllDraftChunksResponse + (*EditDraftChunkRequest)(nil), // 221: openshell.v1.EditDraftChunkRequest + (*EditDraftChunkResponse)(nil), // 222: openshell.v1.EditDraftChunkResponse + (*UndoDraftChunkRequest)(nil), // 223: openshell.v1.UndoDraftChunkRequest + (*UndoDraftChunkResponse)(nil), // 224: openshell.v1.UndoDraftChunkResponse + (*ClearDraftChunksRequest)(nil), // 225: openshell.v1.ClearDraftChunksRequest + (*ClearDraftChunksResponse)(nil), // 226: openshell.v1.ClearDraftChunksResponse + (*GetDraftHistoryRequest)(nil), // 227: openshell.v1.GetDraftHistoryRequest + (*DraftHistoryEntry)(nil), // 228: openshell.v1.DraftHistoryEntry + (*GetDraftHistoryResponse)(nil), // 229: openshell.v1.GetDraftHistoryResponse + (*CreateWorkspaceRequest)(nil), // 230: openshell.v1.CreateWorkspaceRequest + (*CreateWorkspaceResponse)(nil), // 231: openshell.v1.CreateWorkspaceResponse + (*GetWorkspaceRequest)(nil), // 232: openshell.v1.GetWorkspaceRequest + (*GetWorkspaceResponse)(nil), // 233: openshell.v1.GetWorkspaceResponse + (*ListWorkspacesRequest)(nil), // 234: openshell.v1.ListWorkspacesRequest + (*ListWorkspacesResponse)(nil), // 235: openshell.v1.ListWorkspacesResponse + (*DeleteWorkspaceRequest)(nil), // 236: openshell.v1.DeleteWorkspaceRequest + (*DeleteWorkspaceResponse)(nil), // 237: openshell.v1.DeleteWorkspaceResponse + (*WorkspaceMember)(nil), // 238: openshell.v1.WorkspaceMember + (*AddWorkspaceMemberRequest)(nil), // 239: openshell.v1.AddWorkspaceMemberRequest + (*AddWorkspaceMemberResponse)(nil), // 240: openshell.v1.AddWorkspaceMemberResponse + (*RemoveWorkspaceMemberRequest)(nil), // 241: openshell.v1.RemoveWorkspaceMemberRequest + (*RemoveWorkspaceMemberResponse)(nil), // 242: openshell.v1.RemoveWorkspaceMemberResponse + (*ListWorkspaceMembersRequest)(nil), // 243: openshell.v1.ListWorkspaceMembersRequest + (*ListWorkspaceMembersResponse)(nil), // 244: openshell.v1.ListWorkspaceMembersResponse + (*ExtensionServiceCredential)(nil), // 245: openshell.v1.ExtensionServiceCredential + (*EndpointObservation)(nil), // 246: openshell.v1.EndpointObservation + (*ReportEndpointStatusRequest)(nil), // 247: openshell.v1.ReportEndpointStatusRequest + (*ReportEndpointStatusResponse)(nil), // 248: openshell.v1.ReportEndpointStatusResponse + (*EndpointStatus)(nil), // 249: openshell.v1.EndpointStatus + (*SandboxProvisioning)(nil), // 250: openshell.v1.SandboxProvisioning + (*SandboxServiceExposure)(nil), // 251: openshell.v1.SandboxServiceExposure + nil, // 252: openshell.v1.SandboxSpec.EnvironmentEntry + nil, // 253: openshell.v1.SandboxTemplate.LabelsEntry + nil, // 254: openshell.v1.SandboxTemplate.AnnotationsEntry + nil, // 255: openshell.v1.SandboxTemplate.EnvironmentEntry + nil, // 256: openshell.v1.SandboxWorkloadConfig.EnvironmentEntry + nil, // 257: openshell.v1.PlatformEvent.MetadataEntry + nil, // 258: openshell.v1.CreateSandboxRequest.LabelsEntry + nil, // 259: openshell.v1.CreateSandboxRequest.AnnotationsEntry + nil, // 260: openshell.v1.SandboxResponse.ServiceUrlsEntry + nil, // 261: openshell.v1.ExecSandboxRequest.EnvironmentEntry + nil, // 262: openshell.v1.SandboxLogLine.FieldsEntry + nil, // 263: openshell.v1.UpdateProviderRequest.CredentialExpirationTimesEntry + nil, // 264: openshell.v1.ConfigureProviderRefreshRequest.MaterialEntry + nil, // 265: openshell.v1.ProviderProfile.AnnotationsEntry + nil, // 266: openshell.v1.GetSandboxProviderEnvironmentResponse.EnvironmentEntry + nil, // 267: openshell.v1.GetSandboxProviderEnvironmentResponse.CredentialExpirationTimesEntry + nil, // 268: openshell.v1.GetSandboxProviderEnvironmentResponse.DynamicCredentialsEntry + nil, // 269: openshell.v1.GetSandboxProviderEnvironmentResponse.StaticCredentialBindingsEntry + nil, // 270: openshell.v1.GetSandboxProviderEnvironmentResponse.FilesEntry + nil, // 271: openshell.v1.UpdateConfigRequest.AnnotationsEntry + nil, // 272: openshell.v1.UpdateConfigResponse.AnnotationsEntry + nil, // 273: openshell.v1.SandboxPolicyRevision.ProvenanceEntry + nil, // 274: openshell.v1.CreateWorkspaceRequest.LabelsEntry + (*timestamppb.Timestamp)(nil), // 275: google.protobuf.Timestamp + (*datamodelv1.ObjectMeta)(nil), // 276: openshell.datamodel.v1.ObjectMeta + (*sandboxv1.SandboxPolicy)(nil), // 277: openshell.sandbox.v1.SandboxPolicy + (*structpb.Struct)(nil), // 278: google.protobuf.Struct + (*durationpb.Duration)(nil), // 279: google.protobuf.Duration + (*datamodelv1.WorkspaceSelector)(nil), // 280: openshell.datamodel.v1.WorkspaceSelector + (*datamodelv1.Provider)(nil), // 281: openshell.datamodel.v1.Provider + (sandboxv1.PolicySource)(0), // 282: openshell.sandbox.v1.PolicySource + (*sandboxv1.NetworkEndpoint)(nil), // 283: openshell.sandbox.v1.NetworkEndpoint + (*sandboxv1.NetworkBinary)(nil), // 284: openshell.sandbox.v1.NetworkBinary + (*sandboxv1.SettingValue)(nil), // 285: openshell.sandbox.v1.SettingValue + (*sandboxv1.NetworkPolicyRule)(nil), // 286: openshell.sandbox.v1.NetworkPolicyRule + (*sandboxv1.L7DenyRule)(nil), // 287: openshell.sandbox.v1.L7DenyRule + (*sandboxv1.L7Rule)(nil), // 288: openshell.sandbox.v1.L7Rule + (*datamodelv1.Workspace)(nil), // 289: openshell.datamodel.v1.Workspace + (*sandboxv1.GetSandboxConfigRequest)(nil), // 290: openshell.sandbox.v1.GetSandboxConfigRequest + (*sandboxv1.GetGatewayConfigRequest)(nil), // 291: openshell.sandbox.v1.GetGatewayConfigRequest + (*sandboxv1.GetSandboxConfigResponse)(nil), // 292: openshell.sandbox.v1.GetSandboxConfigResponse + (*sandboxv1.GetGatewayConfigResponse)(nil), // 293: openshell.sandbox.v1.GetGatewayConfigResponse } var file_openshell_proto_depIdxs = []int32{ - 271, // 0: openshell.v1.IssueSandboxTokenResponse.expiration_time:type_name -> google.protobuf.Timestamp - 271, // 1: openshell.v1.RefreshSandboxTokenResponse.expiration_time:type_name -> google.protobuf.Timestamp - 242, // 2: openshell.v1.RefreshSandboxTokenResponse.extension_credentials:type_name -> openshell.v1.ExtensionServiceCredential - 271, // 3: openshell.v1.RefreshSandboxTokenResponse.sandbox_expiration_time:type_name -> google.protobuf.Timestamp + 275, // 0: openshell.v1.IssueSandboxTokenResponse.expiration_time:type_name -> google.protobuf.Timestamp + 275, // 1: openshell.v1.RefreshSandboxTokenResponse.expiration_time:type_name -> google.protobuf.Timestamp + 245, // 2: openshell.v1.RefreshSandboxTokenResponse.extension_credentials:type_name -> openshell.v1.ExtensionServiceCredential + 275, // 3: openshell.v1.RefreshSandboxTokenResponse.sandbox_expiration_time:type_name -> google.protobuf.Timestamp 13, // 4: openshell.v1.HealthResponse.status:type_name -> openshell.v1.ServiceStatus 13, // 5: openshell.v1.GetGatewayInfoResponse.status:type_name -> openshell.v1.ServiceStatus - 29, // 6: openshell.v1.GetGatewayInfoResponse.compute_drivers:type_name -> openshell.v1.ComputeDriverInfo - 28, // 7: openshell.v1.GetGatewayInfoResponse.extensions:type_name -> openshell.v1.NegotiatedExtensionInfo + 31, // 6: openshell.v1.GetGatewayInfoResponse.compute_drivers:type_name -> openshell.v1.ComputeDriverInfo + 30, // 7: openshell.v1.GetGatewayInfoResponse.extensions:type_name -> openshell.v1.NegotiatedExtensionInfo 0, // 8: openshell.v1.NegotiatedExtensionInfo.kind:type_name -> openshell.v1.ExtensionKind - 30, // 9: openshell.v1.ComputeDriverInfo.capabilities:type_name -> openshell.v1.ComputeDriverCapabilities - 31, // 10: openshell.v1.ComputeDriverCapabilities.resource_capabilities:type_name -> openshell.v1.ResourceCapabilities - 32, // 11: openshell.v1.ResourceCapabilities.cpu:type_name -> openshell.v1.CpuResourceCapabilities - 33, // 12: openshell.v1.ResourceCapabilities.memory:type_name -> openshell.v1.MemoryResourceCapabilities - 34, // 13: openshell.v1.ResourceCapabilities.gpu:type_name -> openshell.v1.GpuResourceCapabilities - 272, // 14: openshell.v1.Sandbox.metadata:type_name -> openshell.datamodel.v1.ObjectMeta - 36, // 15: openshell.v1.Sandbox.spec:type_name -> openshell.v1.SandboxSpec - 47, // 16: openshell.v1.Sandbox.status:type_name -> openshell.v1.SandboxStatus - 46, // 17: openshell.v1.Sandbox.created_from_workload_template:type_name -> openshell.v1.SandboxWorkloadTemplateProvenance - 249, // 18: openshell.v1.SandboxSpec.environment:type_name -> openshell.v1.SandboxSpec.EnvironmentEntry - 39, // 19: openshell.v1.SandboxSpec.template:type_name -> openshell.v1.SandboxTemplate - 273, // 20: openshell.v1.SandboxSpec.policy:type_name -> openshell.sandbox.v1.SandboxPolicy - 37, // 21: openshell.v1.SandboxSpec.resource_requirements:type_name -> openshell.v1.ResourceRequirements - 38, // 22: openshell.v1.ResourceRequirements.gpu:type_name -> openshell.v1.GpuResourceRequirements - 250, // 23: openshell.v1.SandboxTemplate.labels:type_name -> openshell.v1.SandboxTemplate.LabelsEntry - 251, // 24: openshell.v1.SandboxTemplate.annotations:type_name -> openshell.v1.SandboxTemplate.AnnotationsEntry - 252, // 25: openshell.v1.SandboxTemplate.environment:type_name -> openshell.v1.SandboxTemplate.EnvironmentEntry - 274, // 26: openshell.v1.SandboxTemplate.resources:type_name -> google.protobuf.Struct - 274, // 27: openshell.v1.SandboxTemplate.driver_config:type_name -> google.protobuf.Struct - 272, // 28: openshell.v1.SandboxWorkloadTemplate.metadata:type_name -> openshell.datamodel.v1.ObjectMeta - 41, // 29: openshell.v1.SandboxWorkloadTemplate.spec:type_name -> openshell.v1.SandboxWorkloadTemplateSpec - 42, // 30: openshell.v1.SandboxWorkloadTemplateSpec.workload:type_name -> openshell.v1.SandboxWorkloadConfig - 274, // 31: openshell.v1.SandboxWorkloadTemplateSpec.driver_config:type_name -> google.protobuf.Struct - 44, // 32: openshell.v1.SandboxWorkloadTemplateSpec.desired_service_level:type_name -> openshell.v1.SandboxServiceLevel - 253, // 33: openshell.v1.SandboxWorkloadConfig.environment:type_name -> openshell.v1.SandboxWorkloadConfig.EnvironmentEntry - 43, // 34: openshell.v1.SandboxWorkloadConfig.resources:type_name -> openshell.v1.SandboxResources - 38, // 35: openshell.v1.SandboxResources.gpu:type_name -> openshell.v1.GpuResourceRequirements - 45, // 36: openshell.v1.SandboxServiceLevel.startup:type_name -> openshell.v1.SandboxStartup - 275, // 37: openshell.v1.SandboxStartup.ready_within:type_name -> google.protobuf.Duration - 48, // 38: openshell.v1.SandboxStatus.conditions:type_name -> openshell.v1.SandboxCondition - 1, // 39: openshell.v1.SandboxStatus.phase:type_name -> openshell.v1.SandboxPhase - 246, // 40: openshell.v1.SandboxStatus.endpoint_statuses:type_name -> openshell.v1.EndpointStatus - 173, // 41: openshell.v1.SandboxStatus.configuration_admission:type_name -> openshell.v1.SandboxConfigurationAdmission - 247, // 42: openshell.v1.SandboxStatus.provisioning:type_name -> openshell.v1.SandboxProvisioning - 271, // 43: openshell.v1.SandboxCondition.transition_time:type_name -> google.protobuf.Timestamp - 271, // 44: openshell.v1.PlatformEvent.event_time:type_name -> google.protobuf.Timestamp - 254, // 45: openshell.v1.PlatformEvent.metadata:type_name -> openshell.v1.PlatformEvent.MetadataEntry - 276, // 46: openshell.v1.CreateSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 36, // 47: openshell.v1.CreateSandboxRequest.spec:type_name -> openshell.v1.SandboxSpec - 255, // 48: openshell.v1.CreateSandboxRequest.labels:type_name -> openshell.v1.CreateSandboxRequest.LabelsEntry - 256, // 49: openshell.v1.CreateSandboxRequest.annotations:type_name -> openshell.v1.CreateSandboxRequest.AnnotationsEntry - 248, // 50: openshell.v1.CreateSandboxRequest.service_exposures:type_name -> openshell.v1.SandboxServiceExposure - 276, // 51: openshell.v1.CreateSandboxTemplateRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 40, // 52: openshell.v1.CreateSandboxTemplateRequest.template:type_name -> openshell.v1.SandboxWorkloadTemplate - 276, // 53: openshell.v1.GetSandboxTemplateRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 54: openshell.v1.ListSandboxTemplatesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 55: openshell.v1.DeleteSandboxTemplateRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 40, // 56: openshell.v1.SandboxTemplateResponse.template:type_name -> openshell.v1.SandboxWorkloadTemplate - 40, // 57: openshell.v1.ListSandboxTemplatesResponse.templates:type_name -> openshell.v1.SandboxWorkloadTemplate - 16, // 58: openshell.v1.DeleteSandboxTemplateResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 276, // 59: openshell.v1.BeginRootfsTarStagingRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 271, // 60: openshell.v1.BeginRootfsTarStagingResponse.expiration_time:type_name -> google.protobuf.Timestamp - 276, // 61: openshell.v1.GetSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 62: openshell.v1.ListSandboxesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 63: openshell.v1.ListSandboxProvidersRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 64: openshell.v1.AttachSandboxProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 65: openshell.v1.DetachSandboxProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 66: openshell.v1.DeleteSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 67: openshell.v1.StopSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 68: openshell.v1.StartSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 35, // 69: openshell.v1.SandboxResponse.sandbox:type_name -> openshell.v1.Sandbox - 257, // 70: openshell.v1.SandboxResponse.service_urls:type_name -> openshell.v1.SandboxResponse.ServiceUrlsEntry - 35, // 71: openshell.v1.ListSandboxesResponse.sandboxes:type_name -> openshell.v1.Sandbox - 277, // 72: openshell.v1.ListSandboxProvidersResponse.providers:type_name -> openshell.datamodel.v1.Provider - 35, // 73: openshell.v1.AttachSandboxProviderResponse.sandbox:type_name -> openshell.v1.Sandbox - 77, // 74: openshell.v1.AttachSandboxProviderResponse.receipt:type_name -> openshell.v1.ProviderMutationReceipt - 35, // 75: openshell.v1.DetachSandboxProviderResponse.sandbox:type_name -> openshell.v1.Sandbox - 77, // 76: openshell.v1.DetachSandboxProviderResponse.receipt:type_name -> openshell.v1.ProviderMutationReceipt - 75, // 77: openshell.v1.ConfigSnapshotRevision.sandbox_config:type_name -> openshell.v1.SandboxConfigRevision - 73, // 78: openshell.v1.ConfigSnapshotRevision.provider_target:type_name -> openshell.v1.ProviderDesiredIdentity - 278, // 79: openshell.v1.SandboxConfigRevision.policy_source:type_name -> openshell.sandbox.v1.PolicySource - 5, // 80: openshell.v1.ConfigUpdateOperation.component:type_name -> openshell.v1.ConfigComponent - 74, // 81: openshell.v1.ConfigUpdateOperation.target_revision:type_name -> openshell.v1.ConfigSnapshotRevision - 7, // 82: openshell.v1.ConfigUpdateOperation.state:type_name -> openshell.v1.ConfigUpdateOperationState - 6, // 83: openshell.v1.ConfigUpdateOperation.outcome:type_name -> openshell.v1.ConfigApplyOutcome - 271, // 84: openshell.v1.ConfigUpdateOperation.created_time:type_name -> google.protobuf.Timestamp - 271, // 85: openshell.v1.ConfigUpdateOperation.updated_time:type_name -> google.protobuf.Timestamp - 271, // 86: openshell.v1.ConfigUpdateOperation.completed_time:type_name -> google.protobuf.Timestamp - 2, // 87: openshell.v1.ProviderMutationReceipt.kind:type_name -> openshell.v1.ProviderMutationKind - 73, // 88: openshell.v1.ProviderMutationReceipt.desired:type_name -> openshell.v1.ProviderDesiredIdentity - 271, // 89: openshell.v1.ProviderMutationReceipt.persisted_time:type_name -> google.protobuf.Timestamp - 4, // 90: openshell.v1.ProviderReadinessObservation.reason:type_name -> openshell.v1.ProviderReadinessReason - 77, // 91: openshell.v1.ProviderReadinessStatus.receipt:type_name -> openshell.v1.ProviderMutationReceipt - 3, // 92: openshell.v1.ProviderReadinessStatus.state:type_name -> openshell.v1.ProviderReadinessState - 4, // 93: openshell.v1.ProviderReadinessStatus.reason:type_name -> openshell.v1.ProviderReadinessReason - 78, // 94: openshell.v1.ProviderReadinessStatus.observed:type_name -> openshell.v1.ProviderReadinessObservation - 271, // 95: openshell.v1.ProviderReadinessStatus.observed_time:type_name -> google.protobuf.Timestamp - 271, // 96: openshell.v1.ProviderReadinessStatus.evaluated_time:type_name -> google.protobuf.Timestamp - 76, // 97: openshell.v1.ProviderReadinessStatus.operation:type_name -> openshell.v1.ConfigUpdateOperation - 276, // 98: openshell.v1.GetSandboxProviderStatusRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 79, // 99: openshell.v1.GetSandboxProviderStatusResponse.status:type_name -> openshell.v1.ProviderReadinessStatus - 78, // 100: openshell.v1.ReportProviderReadinessRequest.observation:type_name -> openshell.v1.ProviderReadinessObservation - 275, // 101: openshell.v1.ReportProviderReadinessResponse.report_interval:type_name -> google.protobuf.Duration - 275, // 102: openshell.v1.ReportProviderReadinessResponse.observation_ttl:type_name -> google.protobuf.Duration - 16, // 103: openshell.v1.DeleteSandboxResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 276, // 104: openshell.v1.CreateSshSessionRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 271, // 105: openshell.v1.CreateSshSessionResponse.expiration_time:type_name -> google.protobuf.Timestamp - 276, // 106: openshell.v1.ExposeServiceRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 107: openshell.v1.GetServiceRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 108: openshell.v1.ListServicesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 94, // 109: openshell.v1.ListServicesResponse.services:type_name -> openshell.v1.ServiceEndpointResponse - 276, // 110: openshell.v1.DeleteServiceRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 16, // 111: openshell.v1.DeleteServiceResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 272, // 112: openshell.v1.ServiceEndpoint.metadata:type_name -> openshell.datamodel.v1.ObjectMeta - 93, // 113: openshell.v1.ServiceEndpointResponse.endpoint:type_name -> openshell.v1.ServiceEndpoint - 16, // 114: openshell.v1.RevokeSshSessionResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 276, // 115: openshell.v1.ExecSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 258, // 116: openshell.v1.ExecSandboxRequest.environment:type_name -> openshell.v1.ExecSandboxRequest.EnvironmentEntry - 275, // 117: openshell.v1.ExecSandboxRequest.execution_timeout:type_name -> google.protobuf.Duration - 98, // 118: openshell.v1.ExecSandboxEvent.stdout:type_name -> openshell.v1.ExecSandboxStdout - 99, // 119: openshell.v1.ExecSandboxEvent.stderr:type_name -> openshell.v1.ExecSandboxStderr - 100, // 120: openshell.v1.ExecSandboxEvent.exit:type_name -> openshell.v1.ExecSandboxExit - 193, // 121: openshell.v1.TcpForwardInit.ssh:type_name -> openshell.v1.SshRelayTarget - 194, // 122: openshell.v1.TcpForwardInit.tcp:type_name -> openshell.v1.TcpRelayTarget - 102, // 123: openshell.v1.TcpForwardFrame.init:type_name -> openshell.v1.TcpForwardInit - 97, // 124: openshell.v1.ExecSandboxInput.start:type_name -> openshell.v1.ExecSandboxRequest - 105, // 125: openshell.v1.ExecSandboxInput.resize:type_name -> openshell.v1.ExecSandboxWindowResize - 272, // 126: openshell.v1.SshSession.metadata:type_name -> openshell.datamodel.v1.ObjectMeta - 271, // 127: openshell.v1.SshSession.expiration_time:type_name -> google.protobuf.Timestamp - 276, // 128: openshell.v1.WatchSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 271, // 129: openshell.v1.WatchSandboxRequest.since_time:type_name -> google.protobuf.Timestamp - 35, // 130: openshell.v1.SandboxStreamEvent.sandbox:type_name -> openshell.v1.Sandbox - 109, // 131: openshell.v1.SandboxStreamEvent.log:type_name -> openshell.v1.SandboxLogLine - 49, // 132: openshell.v1.SandboxStreamEvent.event:type_name -> openshell.v1.PlatformEvent - 110, // 133: openshell.v1.SandboxStreamEvent.warning:type_name -> openshell.v1.SandboxStreamWarning - 206, // 134: openshell.v1.SandboxStreamEvent.draft_policy_update:type_name -> openshell.v1.DraftPolicyUpdate - 271, // 135: openshell.v1.SandboxLogLine.event_time:type_name -> google.protobuf.Timestamp - 259, // 136: openshell.v1.SandboxLogLine.fields:type_name -> openshell.v1.SandboxLogLine.FieldsEntry - 276, // 137: openshell.v1.CreateProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 277, // 138: openshell.v1.CreateProviderRequest.provider:type_name -> openshell.datamodel.v1.Provider - 276, // 139: openshell.v1.GetProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 140: openshell.v1.ListProvidersRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 141: openshell.v1.UpdateProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 277, // 142: openshell.v1.UpdateProviderRequest.provider:type_name -> openshell.datamodel.v1.Provider - 260, // 143: openshell.v1.UpdateProviderRequest.credential_expiration_times:type_name -> openshell.v1.UpdateProviderRequest.CredentialExpirationTimesEntry - 276, // 144: openshell.v1.DeleteProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 277, // 145: openshell.v1.ProviderResponse.provider:type_name -> openshell.datamodel.v1.Provider - 77, // 146: openshell.v1.ProviderResponse.target_receipts:type_name -> openshell.v1.ProviderMutationReceipt - 277, // 147: openshell.v1.ListProvidersResponse.providers:type_name -> openshell.datamodel.v1.Provider - 276, // 148: openshell.v1.ListProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 149: openshell.v1.GetProviderProfileRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 139, // 150: openshell.v1.ProviderProfileImportItem.profile:type_name -> openshell.v1.ProviderProfile - 275, // 151: openshell.v1.ProviderCredentialTokenGrant.cache_ttl:type_name -> google.protobuf.Duration - 122, // 152: openshell.v1.ProviderCredentialTokenGrant.audience_overrides:type_name -> openshell.v1.ProviderCredentialTokenGrantAudienceOverride - 8, // 153: openshell.v1.ProviderCredentialTokenGrant.grant_type:type_name -> openshell.v1.ProviderCredentialTokenGrantType - 123, // 154: openshell.v1.ProviderCredentialTokenGrant.subject_token:type_name -> openshell.v1.ProviderCredentialTokenGrantSubjectToken - 128, // 155: openshell.v1.ProviderProfileCredential.refresh:type_name -> openshell.v1.ProviderCredentialRefresh - 124, // 156: openshell.v1.ProviderProfileCredential.token_grant:type_name -> openshell.v1.ProviderCredentialTokenGrant - 9, // 157: openshell.v1.ProviderCredentialRefresh.strategy:type_name -> openshell.v1.ProviderCredentialRefreshStrategy - 275, // 158: openshell.v1.ProviderCredentialRefresh.refresh_before:type_name -> google.protobuf.Duration - 275, // 159: openshell.v1.ProviderCredentialRefresh.max_lifetime:type_name -> google.protobuf.Duration - 126, // 160: openshell.v1.ProviderCredentialRefresh.material:type_name -> openshell.v1.ProviderCredentialRefreshMaterial - 127, // 161: openshell.v1.ProviderCredentialRefresh.additional_outputs:type_name -> openshell.v1.ProviderCredentialRefreshOutput - 9, // 162: openshell.v1.ProviderCredentialRefreshStatus.strategy:type_name -> openshell.v1.ProviderCredentialRefreshStrategy - 271, // 163: openshell.v1.ProviderCredentialRefreshStatus.expiration_time:type_name -> google.protobuf.Timestamp - 271, // 164: openshell.v1.ProviderCredentialRefreshStatus.next_refresh_time:type_name -> google.protobuf.Timestamp - 271, // 165: openshell.v1.ProviderCredentialRefreshStatus.last_refresh_time:type_name -> google.protobuf.Timestamp - 15, // 166: openshell.v1.ProviderCredentialRefreshStatus.recovery_action:type_name -> openshell.v1.ProviderCredentialRefreshRecoveryAction - 271, // 167: openshell.v1.ProviderCredentialRefreshStatus.last_error_time:type_name -> google.protobuf.Timestamp - 276, // 168: openshell.v1.GetProviderRefreshStatusRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 129, // 169: openshell.v1.GetProviderRefreshStatusResponse.credentials:type_name -> openshell.v1.ProviderCredentialRefreshStatus - 276, // 170: openshell.v1.ConfigureProviderRefreshRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 9, // 171: openshell.v1.ConfigureProviderRefreshRequest.strategy:type_name -> openshell.v1.ProviderCredentialRefreshStrategy - 261, // 172: openshell.v1.ConfigureProviderRefreshRequest.material:type_name -> openshell.v1.ConfigureProviderRefreshRequest.MaterialEntry - 271, // 173: openshell.v1.ConfigureProviderRefreshRequest.expiration_time:type_name -> google.protobuf.Timestamp - 129, // 174: openshell.v1.ConfigureProviderRefreshResponse.status:type_name -> openshell.v1.ProviderCredentialRefreshStatus - 276, // 175: openshell.v1.RotateProviderCredentialRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 129, // 176: openshell.v1.RotateProviderCredentialResponse.status:type_name -> openshell.v1.ProviderCredentialRefreshStatus - 276, // 177: openshell.v1.DeleteProviderRefreshRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 16, // 178: openshell.v1.DeleteProviderRefreshResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 10, // 179: openshell.v1.ProviderProfile.category:type_name -> openshell.v1.ProviderProfileCategory - 125, // 180: openshell.v1.ProviderProfile.credentials:type_name -> openshell.v1.ProviderProfileCredential - 279, // 181: openshell.v1.ProviderProfile.endpoints:type_name -> openshell.sandbox.v1.NetworkEndpoint - 280, // 182: openshell.v1.ProviderProfile.binaries:type_name -> openshell.sandbox.v1.NetworkBinary - 130, // 183: openshell.v1.ProviderProfile.discovery:type_name -> openshell.v1.ProviderProfileDiscovery - 262, // 184: openshell.v1.ProviderProfile.annotations:type_name -> openshell.v1.ProviderProfile.AnnotationsEntry - 139, // 185: openshell.v1.ProviderProfileResponse.profile:type_name -> openshell.v1.ProviderProfile - 139, // 186: openshell.v1.ListProviderProfilesResponse.profiles:type_name -> openshell.v1.ProviderProfile - 276, // 187: openshell.v1.ImportProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 120, // 188: openshell.v1.ImportProviderProfilesRequest.profiles:type_name -> openshell.v1.ProviderProfileImportItem - 121, // 189: openshell.v1.ImportProviderProfilesResponse.diagnostics:type_name -> openshell.v1.ProviderProfileDiagnostic - 139, // 190: openshell.v1.ImportProviderProfilesResponse.profiles:type_name -> openshell.v1.ProviderProfile - 276, // 191: openshell.v1.UpdateProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 120, // 192: openshell.v1.UpdateProviderProfilesRequest.profile:type_name -> openshell.v1.ProviderProfileImportItem - 121, // 193: openshell.v1.UpdateProviderProfilesResponse.diagnostics:type_name -> openshell.v1.ProviderProfileDiagnostic - 139, // 194: openshell.v1.UpdateProviderProfilesResponse.profile:type_name -> openshell.v1.ProviderProfile - 276, // 195: openshell.v1.LintProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 120, // 196: openshell.v1.LintProviderProfilesRequest.profiles:type_name -> openshell.v1.ProviderProfileImportItem - 121, // 197: openshell.v1.LintProviderProfilesResponse.diagnostics:type_name -> openshell.v1.ProviderProfileDiagnostic - 16, // 198: openshell.v1.DeleteProviderResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 276, // 199: openshell.v1.DeleteProviderProfileRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 16, // 200: openshell.v1.DeleteProviderProfileResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 152, // 201: openshell.v1.StaticCredentialBinding.endpoints:type_name -> openshell.v1.StaticCredentialEndpointBinding - 263, // 202: openshell.v1.GetSandboxProviderEnvironmentResponse.environment:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.EnvironmentEntry - 264, // 203: openshell.v1.GetSandboxProviderEnvironmentResponse.credential_expiration_times:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.CredentialExpirationTimesEntry - 265, // 204: openshell.v1.GetSandboxProviderEnvironmentResponse.dynamic_credentials:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.DynamicCredentialsEntry - 266, // 205: openshell.v1.GetSandboxProviderEnvironmentResponse.static_credential_bindings:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.StaticCredentialBindingsEntry - 4, // 206: openshell.v1.GetSandboxProviderEnvironmentResponse.readiness_reason:type_name -> openshell.v1.ProviderReadinessReason - 275, // 207: openshell.v1.ExchangeProviderSubjectTokenResponse.expires_after:type_name -> google.protobuf.Duration - 276, // 208: openshell.v1.UpdateConfigRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 273, // 209: openshell.v1.UpdateConfigRequest.policy:type_name -> openshell.sandbox.v1.SandboxPolicy - 281, // 210: openshell.v1.UpdateConfigRequest.setting_value:type_name -> openshell.sandbox.v1.SettingValue - 158, // 211: openshell.v1.UpdateConfigRequest.merge_operations:type_name -> openshell.v1.PolicyMergeOperation - 267, // 212: openshell.v1.UpdateConfigRequest.annotations:type_name -> openshell.v1.UpdateConfigRequest.AnnotationsEntry - 159, // 213: openshell.v1.PolicyMergeOperation.add_rule:type_name -> openshell.v1.AddNetworkRule - 160, // 214: openshell.v1.PolicyMergeOperation.remove_endpoint:type_name -> openshell.v1.RemoveNetworkEndpoint - 161, // 215: openshell.v1.PolicyMergeOperation.remove_rule:type_name -> openshell.v1.RemoveNetworkRule - 163, // 216: openshell.v1.PolicyMergeOperation.add_deny_rules:type_name -> openshell.v1.AddDenyRules - 164, // 217: openshell.v1.PolicyMergeOperation.add_allow_rules:type_name -> openshell.v1.AddAllowRules - 165, // 218: openshell.v1.PolicyMergeOperation.remove_binary:type_name -> openshell.v1.RemoveNetworkBinary - 282, // 219: openshell.v1.AddNetworkRule.rule:type_name -> openshell.sandbox.v1.NetworkPolicyRule - 280, // 220: openshell.v1.L7RuleTarget.binaries:type_name -> openshell.sandbox.v1.NetworkBinary - 283, // 221: openshell.v1.AddDenyRules.deny_rules:type_name -> openshell.sandbox.v1.L7DenyRule - 162, // 222: openshell.v1.AddDenyRules.target:type_name -> openshell.v1.L7RuleTarget - 284, // 223: openshell.v1.AddAllowRules.rules:type_name -> openshell.sandbox.v1.L7Rule - 162, // 224: openshell.v1.AddAllowRules.target:type_name -> openshell.v1.L7RuleTarget - 268, // 225: openshell.v1.UpdateConfigResponse.annotations:type_name -> openshell.v1.UpdateConfigResponse.AnnotationsEntry - 276, // 226: openshell.v1.GetSandboxPolicyStatusRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 176, // 227: openshell.v1.GetSandboxPolicyStatusResponse.revision:type_name -> openshell.v1.SandboxPolicyRevision - 276, // 228: openshell.v1.ListSandboxPoliciesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 176, // 229: openshell.v1.ListSandboxPoliciesResponse.revisions:type_name -> openshell.v1.SandboxPolicyRevision - 12, // 230: openshell.v1.ReportPolicyStatusRequest.status:type_name -> openshell.v1.PolicyStatus - 11, // 231: openshell.v1.SandboxConfigurationAdmission.state:type_name -> openshell.v1.ConfigurationAdmissionState - 173, // 232: openshell.v1.ReportSandboxConfigurationRequest.admission:type_name -> openshell.v1.SandboxConfigurationAdmission - 12, // 233: openshell.v1.SandboxPolicyRevision.status:type_name -> openshell.v1.PolicyStatus - 271, // 234: openshell.v1.SandboxPolicyRevision.created_time:type_name -> google.protobuf.Timestamp - 271, // 235: openshell.v1.SandboxPolicyRevision.loaded_time:type_name -> google.protobuf.Timestamp - 273, // 236: openshell.v1.SandboxPolicyRevision.policy:type_name -> openshell.sandbox.v1.SandboxPolicy - 269, // 237: openshell.v1.SandboxPolicyRevision.provenance:type_name -> openshell.v1.SandboxPolicyRevision.ProvenanceEntry - 276, // 238: openshell.v1.GetSandboxLogsRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 271, // 239: openshell.v1.GetSandboxLogsRequest.since_time:type_name -> google.protobuf.Timestamp - 109, // 240: openshell.v1.PushSandboxLogsRequest.logs:type_name -> openshell.v1.SandboxLogLine - 109, // 241: openshell.v1.GetSandboxLogsResponse.logs:type_name -> openshell.v1.SandboxLogLine - 183, // 242: openshell.v1.SupervisorMessage.hello:type_name -> openshell.v1.SupervisorHello - 186, // 243: openshell.v1.SupervisorMessage.heartbeat:type_name -> openshell.v1.SupervisorHeartbeat - 199, // 244: openshell.v1.SupervisorMessage.relay_open_result:type_name -> openshell.v1.RelayOpenResult - 200, // 245: openshell.v1.SupervisorMessage.relay_close:type_name -> openshell.v1.RelayClose - 184, // 246: openshell.v1.GatewayMessage.session_accepted:type_name -> openshell.v1.SessionAccepted - 185, // 247: openshell.v1.GatewayMessage.session_rejected:type_name -> openshell.v1.SessionRejected - 187, // 248: openshell.v1.GatewayMessage.heartbeat:type_name -> openshell.v1.GatewayHeartbeat - 192, // 249: openshell.v1.GatewayMessage.relay_open:type_name -> openshell.v1.RelayOpen - 200, // 250: openshell.v1.GatewayMessage.relay_close:type_name -> openshell.v1.RelayClose - 275, // 251: openshell.v1.SessionAccepted.heartbeat_interval:type_name -> google.protobuf.Duration - 193, // 252: openshell.v1.RelayOpen.ssh:type_name -> openshell.v1.SshRelayTarget - 194, // 253: openshell.v1.RelayOpen.tcp:type_name -> openshell.v1.TcpRelayTarget - 195, // 254: openshell.v1.RelayFrame.init:type_name -> openshell.v1.RelayInit - 192, // 255: openshell.v1.PeerRelayInit.relay_open:type_name -> openshell.v1.RelayOpen - 197, // 256: openshell.v1.PeerRelayFrame.init:type_name -> openshell.v1.PeerRelayInit - 271, // 257: openshell.v1.DenialSummary.first_seen_time:type_name -> google.protobuf.Timestamp - 271, // 258: openshell.v1.DenialSummary.last_seen_time:type_name -> google.protobuf.Timestamp - 201, // 259: openshell.v1.DenialSummary.l7_request_samples:type_name -> openshell.v1.L7RequestSample - 203, // 260: openshell.v1.NetworkActivitySummary.denials_by_group:type_name -> openshell.v1.DenialGroupCount - 282, // 261: openshell.v1.PolicyChunk.proposed_rule:type_name -> openshell.sandbox.v1.NetworkPolicyRule - 271, // 262: openshell.v1.PolicyChunk.created_time:type_name -> google.protobuf.Timestamp - 271, // 263: openshell.v1.PolicyChunk.decided_time:type_name -> google.protobuf.Timestamp - 271, // 264: openshell.v1.PolicyChunk.first_seen_time:type_name -> google.protobuf.Timestamp - 271, // 265: openshell.v1.PolicyChunk.last_seen_time:type_name -> google.protobuf.Timestamp - 273, // 266: openshell.v1.PolicyChunk.current_effective_policy:type_name -> openshell.sandbox.v1.SandboxPolicy - 273, // 267: openshell.v1.PolicyChunk.candidate_effective_policy:type_name -> openshell.sandbox.v1.SandboxPolicy - 276, // 268: openshell.v1.SubmitPolicyAnalysisRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 202, // 269: openshell.v1.SubmitPolicyAnalysisRequest.summaries:type_name -> openshell.v1.DenialSummary - 205, // 270: openshell.v1.SubmitPolicyAnalysisRequest.proposed_chunks:type_name -> openshell.v1.PolicyChunk - 204, // 271: openshell.v1.SubmitPolicyAnalysisRequest.network_activity_summaries:type_name -> openshell.v1.NetworkActivitySummary - 276, // 272: openshell.v1.GetDraftPolicyRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 205, // 273: openshell.v1.GetDraftPolicyResponse.chunks:type_name -> openshell.v1.PolicyChunk - 271, // 274: openshell.v1.GetDraftPolicyResponse.last_analyzed_time:type_name -> google.protobuf.Timestamp - 276, // 275: openshell.v1.ApproveDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 276: openshell.v1.RejectDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 277: openshell.v1.ApproveAllDraftChunksRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 215, // 278: openshell.v1.ApproveAllDraftChunksRequest.approvals:type_name -> openshell.v1.DraftChunkApproval - 276, // 279: openshell.v1.EditDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 282, // 280: openshell.v1.EditDraftChunkRequest.proposed_rule:type_name -> openshell.sandbox.v1.NetworkPolicyRule - 276, // 281: openshell.v1.UndoDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 282: openshell.v1.ClearDraftChunksRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 276, // 283: openshell.v1.GetDraftHistoryRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 271, // 284: openshell.v1.DraftHistoryEntry.event_time:type_name -> google.protobuf.Timestamp - 225, // 285: openshell.v1.GetDraftHistoryResponse.entries:type_name -> openshell.v1.DraftHistoryEntry - 270, // 286: openshell.v1.CreateWorkspaceRequest.labels:type_name -> openshell.v1.CreateWorkspaceRequest.LabelsEntry - 285, // 287: openshell.v1.CreateWorkspaceResponse.workspace:type_name -> openshell.datamodel.v1.Workspace - 285, // 288: openshell.v1.GetWorkspaceResponse.workspace:type_name -> openshell.datamodel.v1.Workspace - 285, // 289: openshell.v1.ListWorkspacesResponse.workspaces:type_name -> openshell.datamodel.v1.Workspace - 16, // 290: openshell.v1.DeleteWorkspaceResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 272, // 291: openshell.v1.WorkspaceMember.metadata:type_name -> openshell.datamodel.v1.ObjectMeta - 14, // 292: openshell.v1.WorkspaceMember.role:type_name -> openshell.v1.WorkspaceRole - 276, // 293: openshell.v1.AddWorkspaceMemberRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 14, // 294: openshell.v1.AddWorkspaceMemberRequest.role:type_name -> openshell.v1.WorkspaceRole - 235, // 295: openshell.v1.AddWorkspaceMemberResponse.member:type_name -> openshell.v1.WorkspaceMember - 276, // 296: openshell.v1.RemoveWorkspaceMemberRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 16, // 297: openshell.v1.RemoveWorkspaceMemberResponse.outcome:type_name -> openshell.v1.DeletionOutcome - 276, // 298: openshell.v1.ListWorkspaceMembersRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector - 235, // 299: openshell.v1.ListWorkspaceMembersResponse.members:type_name -> openshell.v1.WorkspaceMember - 271, // 300: openshell.v1.ExtensionServiceCredential.expiration_time:type_name -> google.protobuf.Timestamp - 17, // 301: openshell.v1.EndpointObservation.result:type_name -> openshell.v1.EndpointResult - 243, // 302: openshell.v1.ReportEndpointStatusRequest.observations:type_name -> openshell.v1.EndpointObservation - 17, // 303: openshell.v1.EndpointStatus.last_result:type_name -> openshell.v1.EndpointResult - 271, // 304: openshell.v1.EndpointStatus.last_reported_time:type_name -> google.protobuf.Timestamp - 271, // 305: openshell.v1.SandboxProvisioning.configuration_change_time:type_name -> google.protobuf.Timestamp - 271, // 306: openshell.v1.SandboxProvisioning.first_rejection_time:type_name -> google.protobuf.Timestamp - 271, // 307: openshell.v1.SandboxProvisioning.deadline:type_name -> google.protobuf.Timestamp - 271, // 308: openshell.v1.SandboxProvisioning.timeout_time:type_name -> google.protobuf.Timestamp - 271, // 309: openshell.v1.SandboxProvisioning.cleanup_completed_time:type_name -> google.protobuf.Timestamp - 271, // 310: openshell.v1.SandboxProvisioning.cleanup_retry_time:type_name -> google.protobuf.Timestamp - 271, // 311: openshell.v1.SandboxProvisioning.attachment_change_time:type_name -> google.protobuf.Timestamp - 271, // 312: openshell.v1.UpdateProviderRequest.CredentialExpirationTimesEntry.value:type_name -> google.protobuf.Timestamp - 271, // 313: openshell.v1.GetSandboxProviderEnvironmentResponse.CredentialExpirationTimesEntry.value:type_name -> google.protobuf.Timestamp - 125, // 314: openshell.v1.GetSandboxProviderEnvironmentResponse.DynamicCredentialsEntry.value:type_name -> openshell.v1.ProviderProfileCredential - 153, // 315: openshell.v1.GetSandboxProviderEnvironmentResponse.StaticCredentialBindingsEntry.value:type_name -> openshell.v1.StaticCredentialBinding - 22, // 316: openshell.v1.OpenShell.Health:input_type -> openshell.v1.HealthRequest - 24, // 317: openshell.v1.OpenShell.GetCurrentUser:input_type -> openshell.v1.GetCurrentUserRequest - 26, // 318: openshell.v1.OpenShell.GetGatewayInfo:input_type -> openshell.v1.GetGatewayInfoRequest - 50, // 319: openshell.v1.OpenShell.CreateSandbox:input_type -> openshell.v1.CreateSandboxRequest - 58, // 320: openshell.v1.OpenShell.BeginRootfsTarStaging:input_type -> openshell.v1.BeginRootfsTarStagingRequest - 60, // 321: openshell.v1.OpenShell.GetSandbox:input_type -> openshell.v1.GetSandboxRequest - 61, // 322: openshell.v1.OpenShell.ListSandboxes:input_type -> openshell.v1.ListSandboxesRequest - 51, // 323: openshell.v1.OpenShell.CreateSandboxTemplate:input_type -> openshell.v1.CreateSandboxTemplateRequest - 52, // 324: openshell.v1.OpenShell.GetSandboxTemplate:input_type -> openshell.v1.GetSandboxTemplateRequest - 53, // 325: openshell.v1.OpenShell.ListSandboxTemplates:input_type -> openshell.v1.ListSandboxTemplatesRequest - 54, // 326: openshell.v1.OpenShell.DeleteSandboxTemplate:input_type -> openshell.v1.DeleteSandboxTemplateRequest - 62, // 327: openshell.v1.OpenShell.ListSandboxProviders:input_type -> openshell.v1.ListSandboxProvidersRequest - 63, // 328: openshell.v1.OpenShell.AttachSandboxProvider:input_type -> openshell.v1.AttachSandboxProviderRequest - 64, // 329: openshell.v1.OpenShell.DetachSandboxProvider:input_type -> openshell.v1.DetachSandboxProviderRequest - 80, // 330: openshell.v1.OpenShell.GetSandboxProviderStatus:input_type -> openshell.v1.GetSandboxProviderStatusRequest - 65, // 331: openshell.v1.OpenShell.DeleteSandbox:input_type -> openshell.v1.DeleteSandboxRequest - 66, // 332: openshell.v1.OpenShell.StopSandbox:input_type -> openshell.v1.StopSandboxRequest - 67, // 333: openshell.v1.OpenShell.StartSandbox:input_type -> openshell.v1.StartSandboxRequest - 85, // 334: openshell.v1.OpenShell.CreateSshSession:input_type -> openshell.v1.CreateSshSessionRequest - 87, // 335: openshell.v1.OpenShell.ExposeService:input_type -> openshell.v1.ExposeServiceRequest - 88, // 336: openshell.v1.OpenShell.GetService:input_type -> openshell.v1.GetServiceRequest - 89, // 337: openshell.v1.OpenShell.ListServices:input_type -> openshell.v1.ListServicesRequest - 91, // 338: openshell.v1.OpenShell.DeleteService:input_type -> openshell.v1.DeleteServiceRequest - 95, // 339: openshell.v1.OpenShell.RevokeSshSession:input_type -> openshell.v1.RevokeSshSessionRequest - 97, // 340: openshell.v1.OpenShell.ExecSandbox:input_type -> openshell.v1.ExecSandboxRequest - 103, // 341: openshell.v1.OpenShell.ForwardTcp:input_type -> openshell.v1.TcpForwardFrame - 104, // 342: openshell.v1.OpenShell.ExecSandboxInteractive:input_type -> openshell.v1.ExecSandboxInput - 111, // 343: openshell.v1.OpenShell.CreateProvider:input_type -> openshell.v1.CreateProviderRequest - 112, // 344: openshell.v1.OpenShell.GetProvider:input_type -> openshell.v1.GetProviderRequest - 113, // 345: openshell.v1.OpenShell.ListProviders:input_type -> openshell.v1.ListProvidersRequest - 118, // 346: openshell.v1.OpenShell.ListProviderProfiles:input_type -> openshell.v1.ListProviderProfilesRequest - 119, // 347: openshell.v1.OpenShell.GetProviderProfile:input_type -> openshell.v1.GetProviderProfileRequest - 142, // 348: openshell.v1.OpenShell.ImportProviderProfiles:input_type -> openshell.v1.ImportProviderProfilesRequest - 144, // 349: openshell.v1.OpenShell.UpdateProviderProfiles:input_type -> openshell.v1.UpdateProviderProfilesRequest - 146, // 350: openshell.v1.OpenShell.LintProviderProfiles:input_type -> openshell.v1.LintProviderProfilesRequest - 114, // 351: openshell.v1.OpenShell.UpdateProvider:input_type -> openshell.v1.UpdateProviderRequest - 131, // 352: openshell.v1.OpenShell.GetProviderRefreshStatus:input_type -> openshell.v1.GetProviderRefreshStatusRequest - 133, // 353: openshell.v1.OpenShell.ConfigureProviderRefresh:input_type -> openshell.v1.ConfigureProviderRefreshRequest - 135, // 354: openshell.v1.OpenShell.RotateProviderCredential:input_type -> openshell.v1.RotateProviderCredentialRequest - 137, // 355: openshell.v1.OpenShell.DeleteProviderRefresh:input_type -> openshell.v1.DeleteProviderRefreshRequest - 115, // 356: openshell.v1.OpenShell.DeleteProvider:input_type -> openshell.v1.DeleteProviderRequest - 149, // 357: openshell.v1.OpenShell.DeleteProviderProfile:input_type -> openshell.v1.DeleteProviderProfileRequest - 286, // 358: openshell.v1.OpenShell.GetSandboxConfig:input_type -> openshell.sandbox.v1.GetSandboxConfigRequest - 287, // 359: openshell.v1.OpenShell.GetGatewayConfig:input_type -> openshell.sandbox.v1.GetGatewayConfigRequest - 157, // 360: openshell.v1.OpenShell.UpdateConfig:input_type -> openshell.v1.UpdateConfigRequest - 167, // 361: openshell.v1.OpenShell.GetSandboxPolicyStatus:input_type -> openshell.v1.GetSandboxPolicyStatusRequest - 169, // 362: openshell.v1.OpenShell.ListSandboxPolicies:input_type -> openshell.v1.ListSandboxPoliciesRequest - 171, // 363: openshell.v1.OpenShell.ReportPolicyStatus:input_type -> openshell.v1.ReportPolicyStatusRequest - 244, // 364: openshell.v1.OpenShell.ReportEndpointStatus:input_type -> openshell.v1.ReportEndpointStatusRequest - 82, // 365: openshell.v1.OpenShell.ReportProviderReadiness:input_type -> openshell.v1.ReportProviderReadinessRequest - 174, // 366: openshell.v1.OpenShell.ReportSandboxConfiguration:input_type -> openshell.v1.ReportSandboxConfigurationRequest - 151, // 367: openshell.v1.OpenShell.GetSandboxProviderEnvironment:input_type -> openshell.v1.GetSandboxProviderEnvironmentRequest - 155, // 368: openshell.v1.OpenShell.ExchangeProviderSubjectToken:input_type -> openshell.v1.ExchangeProviderSubjectTokenRequest - 177, // 369: openshell.v1.OpenShell.GetSandboxLogs:input_type -> openshell.v1.GetSandboxLogsRequest - 178, // 370: openshell.v1.OpenShell.PushSandboxLogs:input_type -> openshell.v1.PushSandboxLogsRequest - 181, // 371: openshell.v1.OpenShell.ConnectSupervisor:input_type -> openshell.v1.SupervisorMessage - 188, // 372: openshell.v1.OpenShell.ReportMainProcessExit:input_type -> openshell.v1.ReportMainProcessExitRequest - 190, // 373: openshell.v1.OpenShell.FinalizeMainProcessExit:input_type -> openshell.v1.FinalizeMainProcessExitRequest - 196, // 374: openshell.v1.OpenShell.RelayStream:input_type -> openshell.v1.RelayFrame - 198, // 375: openshell.v1.OpenShell.PeerRelay:input_type -> openshell.v1.PeerRelayFrame - 82, // 376: openshell.v1.OpenShell.PeerReportProviderReadiness:input_type -> openshell.v1.ReportProviderReadinessRequest - 244, // 377: openshell.v1.OpenShell.PeerReportEndpointStatus:input_type -> openshell.v1.ReportEndpointStatusRequest - 80, // 378: openshell.v1.OpenShell.PeerGetSandboxProviderStatus:input_type -> openshell.v1.GetSandboxProviderStatusRequest - 107, // 379: openshell.v1.OpenShell.WatchSandbox:input_type -> openshell.v1.WatchSandboxRequest - 207, // 380: openshell.v1.OpenShell.SubmitPolicyAnalysis:input_type -> openshell.v1.SubmitPolicyAnalysisRequest - 209, // 381: openshell.v1.OpenShell.GetDraftPolicy:input_type -> openshell.v1.GetDraftPolicyRequest - 211, // 382: openshell.v1.OpenShell.ApproveDraftChunk:input_type -> openshell.v1.ApproveDraftChunkRequest - 213, // 383: openshell.v1.OpenShell.RejectDraftChunk:input_type -> openshell.v1.RejectDraftChunkRequest - 216, // 384: openshell.v1.OpenShell.ApproveAllDraftChunks:input_type -> openshell.v1.ApproveAllDraftChunksRequest - 218, // 385: openshell.v1.OpenShell.EditDraftChunk:input_type -> openshell.v1.EditDraftChunkRequest - 220, // 386: openshell.v1.OpenShell.UndoDraftChunk:input_type -> openshell.v1.UndoDraftChunkRequest - 222, // 387: openshell.v1.OpenShell.ClearDraftChunks:input_type -> openshell.v1.ClearDraftChunksRequest - 224, // 388: openshell.v1.OpenShell.GetDraftHistory:input_type -> openshell.v1.GetDraftHistoryRequest - 18, // 389: openshell.v1.OpenShell.IssueSandboxToken:input_type -> openshell.v1.IssueSandboxTokenRequest - 20, // 390: openshell.v1.OpenShell.RefreshSandboxToken:input_type -> openshell.v1.RefreshSandboxTokenRequest - 227, // 391: openshell.v1.OpenShell.CreateWorkspace:input_type -> openshell.v1.CreateWorkspaceRequest - 229, // 392: openshell.v1.OpenShell.GetWorkspace:input_type -> openshell.v1.GetWorkspaceRequest - 231, // 393: openshell.v1.OpenShell.ListWorkspaces:input_type -> openshell.v1.ListWorkspacesRequest - 233, // 394: openshell.v1.OpenShell.DeleteWorkspace:input_type -> openshell.v1.DeleteWorkspaceRequest - 236, // 395: openshell.v1.OpenShell.AddWorkspaceMember:input_type -> openshell.v1.AddWorkspaceMemberRequest - 238, // 396: openshell.v1.OpenShell.RemoveWorkspaceMember:input_type -> openshell.v1.RemoveWorkspaceMemberRequest - 240, // 397: openshell.v1.OpenShell.ListWorkspaceMembers:input_type -> openshell.v1.ListWorkspaceMembersRequest - 23, // 398: openshell.v1.OpenShell.Health:output_type -> openshell.v1.HealthResponse - 25, // 399: openshell.v1.OpenShell.GetCurrentUser:output_type -> openshell.v1.GetCurrentUserResponse - 27, // 400: openshell.v1.OpenShell.GetGatewayInfo:output_type -> openshell.v1.GetGatewayInfoResponse - 68, // 401: openshell.v1.OpenShell.CreateSandbox:output_type -> openshell.v1.SandboxResponse - 59, // 402: openshell.v1.OpenShell.BeginRootfsTarStaging:output_type -> openshell.v1.BeginRootfsTarStagingResponse - 68, // 403: openshell.v1.OpenShell.GetSandbox:output_type -> openshell.v1.SandboxResponse - 69, // 404: openshell.v1.OpenShell.ListSandboxes:output_type -> openshell.v1.ListSandboxesResponse - 55, // 405: openshell.v1.OpenShell.CreateSandboxTemplate:output_type -> openshell.v1.SandboxTemplateResponse - 55, // 406: openshell.v1.OpenShell.GetSandboxTemplate:output_type -> openshell.v1.SandboxTemplateResponse - 56, // 407: openshell.v1.OpenShell.ListSandboxTemplates:output_type -> openshell.v1.ListSandboxTemplatesResponse - 57, // 408: openshell.v1.OpenShell.DeleteSandboxTemplate:output_type -> openshell.v1.DeleteSandboxTemplateResponse - 70, // 409: openshell.v1.OpenShell.ListSandboxProviders:output_type -> openshell.v1.ListSandboxProvidersResponse - 71, // 410: openshell.v1.OpenShell.AttachSandboxProvider:output_type -> openshell.v1.AttachSandboxProviderResponse - 72, // 411: openshell.v1.OpenShell.DetachSandboxProvider:output_type -> openshell.v1.DetachSandboxProviderResponse - 81, // 412: openshell.v1.OpenShell.GetSandboxProviderStatus:output_type -> openshell.v1.GetSandboxProviderStatusResponse - 84, // 413: openshell.v1.OpenShell.DeleteSandbox:output_type -> openshell.v1.DeleteSandboxResponse - 68, // 414: openshell.v1.OpenShell.StopSandbox:output_type -> openshell.v1.SandboxResponse - 68, // 415: openshell.v1.OpenShell.StartSandbox:output_type -> openshell.v1.SandboxResponse - 86, // 416: openshell.v1.OpenShell.CreateSshSession:output_type -> openshell.v1.CreateSshSessionResponse - 94, // 417: openshell.v1.OpenShell.ExposeService:output_type -> openshell.v1.ServiceEndpointResponse - 94, // 418: openshell.v1.OpenShell.GetService:output_type -> openshell.v1.ServiceEndpointResponse - 90, // 419: openshell.v1.OpenShell.ListServices:output_type -> openshell.v1.ListServicesResponse - 92, // 420: openshell.v1.OpenShell.DeleteService:output_type -> openshell.v1.DeleteServiceResponse - 96, // 421: openshell.v1.OpenShell.RevokeSshSession:output_type -> openshell.v1.RevokeSshSessionResponse - 101, // 422: openshell.v1.OpenShell.ExecSandbox:output_type -> openshell.v1.ExecSandboxEvent - 103, // 423: openshell.v1.OpenShell.ForwardTcp:output_type -> openshell.v1.TcpForwardFrame - 101, // 424: openshell.v1.OpenShell.ExecSandboxInteractive:output_type -> openshell.v1.ExecSandboxEvent - 116, // 425: openshell.v1.OpenShell.CreateProvider:output_type -> openshell.v1.ProviderResponse - 116, // 426: openshell.v1.OpenShell.GetProvider:output_type -> openshell.v1.ProviderResponse - 117, // 427: openshell.v1.OpenShell.ListProviders:output_type -> openshell.v1.ListProvidersResponse - 141, // 428: openshell.v1.OpenShell.ListProviderProfiles:output_type -> openshell.v1.ListProviderProfilesResponse - 140, // 429: openshell.v1.OpenShell.GetProviderProfile:output_type -> openshell.v1.ProviderProfileResponse - 143, // 430: openshell.v1.OpenShell.ImportProviderProfiles:output_type -> openshell.v1.ImportProviderProfilesResponse - 145, // 431: openshell.v1.OpenShell.UpdateProviderProfiles:output_type -> openshell.v1.UpdateProviderProfilesResponse - 147, // 432: openshell.v1.OpenShell.LintProviderProfiles:output_type -> openshell.v1.LintProviderProfilesResponse - 116, // 433: openshell.v1.OpenShell.UpdateProvider:output_type -> openshell.v1.ProviderResponse - 132, // 434: openshell.v1.OpenShell.GetProviderRefreshStatus:output_type -> openshell.v1.GetProviderRefreshStatusResponse - 134, // 435: openshell.v1.OpenShell.ConfigureProviderRefresh:output_type -> openshell.v1.ConfigureProviderRefreshResponse - 136, // 436: openshell.v1.OpenShell.RotateProviderCredential:output_type -> openshell.v1.RotateProviderCredentialResponse - 138, // 437: openshell.v1.OpenShell.DeleteProviderRefresh:output_type -> openshell.v1.DeleteProviderRefreshResponse - 148, // 438: openshell.v1.OpenShell.DeleteProvider:output_type -> openshell.v1.DeleteProviderResponse - 150, // 439: openshell.v1.OpenShell.DeleteProviderProfile:output_type -> openshell.v1.DeleteProviderProfileResponse - 288, // 440: openshell.v1.OpenShell.GetSandboxConfig:output_type -> openshell.sandbox.v1.GetSandboxConfigResponse - 289, // 441: openshell.v1.OpenShell.GetGatewayConfig:output_type -> openshell.sandbox.v1.GetGatewayConfigResponse - 166, // 442: openshell.v1.OpenShell.UpdateConfig:output_type -> openshell.v1.UpdateConfigResponse - 168, // 443: openshell.v1.OpenShell.GetSandboxPolicyStatus:output_type -> openshell.v1.GetSandboxPolicyStatusResponse - 170, // 444: openshell.v1.OpenShell.ListSandboxPolicies:output_type -> openshell.v1.ListSandboxPoliciesResponse - 172, // 445: openshell.v1.OpenShell.ReportPolicyStatus:output_type -> openshell.v1.ReportPolicyStatusResponse - 245, // 446: openshell.v1.OpenShell.ReportEndpointStatus:output_type -> openshell.v1.ReportEndpointStatusResponse - 83, // 447: openshell.v1.OpenShell.ReportProviderReadiness:output_type -> openshell.v1.ReportProviderReadinessResponse - 175, // 448: openshell.v1.OpenShell.ReportSandboxConfiguration:output_type -> openshell.v1.ReportSandboxConfigurationResponse - 154, // 449: openshell.v1.OpenShell.GetSandboxProviderEnvironment:output_type -> openshell.v1.GetSandboxProviderEnvironmentResponse - 156, // 450: openshell.v1.OpenShell.ExchangeProviderSubjectToken:output_type -> openshell.v1.ExchangeProviderSubjectTokenResponse - 180, // 451: openshell.v1.OpenShell.GetSandboxLogs:output_type -> openshell.v1.GetSandboxLogsResponse - 179, // 452: openshell.v1.OpenShell.PushSandboxLogs:output_type -> openshell.v1.PushSandboxLogsResponse - 182, // 453: openshell.v1.OpenShell.ConnectSupervisor:output_type -> openshell.v1.GatewayMessage - 189, // 454: openshell.v1.OpenShell.ReportMainProcessExit:output_type -> openshell.v1.ReportMainProcessExitResponse - 191, // 455: openshell.v1.OpenShell.FinalizeMainProcessExit:output_type -> openshell.v1.FinalizeMainProcessExitResponse - 196, // 456: openshell.v1.OpenShell.RelayStream:output_type -> openshell.v1.RelayFrame - 198, // 457: openshell.v1.OpenShell.PeerRelay:output_type -> openshell.v1.PeerRelayFrame - 83, // 458: openshell.v1.OpenShell.PeerReportProviderReadiness:output_type -> openshell.v1.ReportProviderReadinessResponse - 245, // 459: openshell.v1.OpenShell.PeerReportEndpointStatus:output_type -> openshell.v1.ReportEndpointStatusResponse - 81, // 460: openshell.v1.OpenShell.PeerGetSandboxProviderStatus:output_type -> openshell.v1.GetSandboxProviderStatusResponse - 108, // 461: openshell.v1.OpenShell.WatchSandbox:output_type -> openshell.v1.SandboxStreamEvent - 208, // 462: openshell.v1.OpenShell.SubmitPolicyAnalysis:output_type -> openshell.v1.SubmitPolicyAnalysisResponse - 210, // 463: openshell.v1.OpenShell.GetDraftPolicy:output_type -> openshell.v1.GetDraftPolicyResponse - 212, // 464: openshell.v1.OpenShell.ApproveDraftChunk:output_type -> openshell.v1.ApproveDraftChunkResponse - 214, // 465: openshell.v1.OpenShell.RejectDraftChunk:output_type -> openshell.v1.RejectDraftChunkResponse - 217, // 466: openshell.v1.OpenShell.ApproveAllDraftChunks:output_type -> openshell.v1.ApproveAllDraftChunksResponse - 219, // 467: openshell.v1.OpenShell.EditDraftChunk:output_type -> openshell.v1.EditDraftChunkResponse - 221, // 468: openshell.v1.OpenShell.UndoDraftChunk:output_type -> openshell.v1.UndoDraftChunkResponse - 223, // 469: openshell.v1.OpenShell.ClearDraftChunks:output_type -> openshell.v1.ClearDraftChunksResponse - 226, // 470: openshell.v1.OpenShell.GetDraftHistory:output_type -> openshell.v1.GetDraftHistoryResponse - 19, // 471: openshell.v1.OpenShell.IssueSandboxToken:output_type -> openshell.v1.IssueSandboxTokenResponse - 21, // 472: openshell.v1.OpenShell.RefreshSandboxToken:output_type -> openshell.v1.RefreshSandboxTokenResponse - 228, // 473: openshell.v1.OpenShell.CreateWorkspace:output_type -> openshell.v1.CreateWorkspaceResponse - 230, // 474: openshell.v1.OpenShell.GetWorkspace:output_type -> openshell.v1.GetWorkspaceResponse - 232, // 475: openshell.v1.OpenShell.ListWorkspaces:output_type -> openshell.v1.ListWorkspacesResponse - 234, // 476: openshell.v1.OpenShell.DeleteWorkspace:output_type -> openshell.v1.DeleteWorkspaceResponse - 237, // 477: openshell.v1.OpenShell.AddWorkspaceMember:output_type -> openshell.v1.AddWorkspaceMemberResponse - 239, // 478: openshell.v1.OpenShell.RemoveWorkspaceMember:output_type -> openshell.v1.RemoveWorkspaceMemberResponse - 241, // 479: openshell.v1.OpenShell.ListWorkspaceMembers:output_type -> openshell.v1.ListWorkspaceMembersResponse - 398, // [398:480] is the sub-list for method output_type - 316, // [316:398] is the sub-list for method input_type - 316, // [316:316] is the sub-list for extension type_name - 316, // [316:316] is the sub-list for extension extendee - 0, // [0:316] is the sub-list for field type_name + 32, // 9: openshell.v1.ComputeDriverInfo.capabilities:type_name -> openshell.v1.ComputeDriverCapabilities + 33, // 10: openshell.v1.ComputeDriverCapabilities.resource_capabilities:type_name -> openshell.v1.ResourceCapabilities + 34, // 11: openshell.v1.ResourceCapabilities.cpu:type_name -> openshell.v1.CpuResourceCapabilities + 35, // 12: openshell.v1.ResourceCapabilities.memory:type_name -> openshell.v1.MemoryResourceCapabilities + 36, // 13: openshell.v1.ResourceCapabilities.gpu:type_name -> openshell.v1.GpuResourceCapabilities + 276, // 14: openshell.v1.Sandbox.metadata:type_name -> openshell.datamodel.v1.ObjectMeta + 38, // 15: openshell.v1.Sandbox.spec:type_name -> openshell.v1.SandboxSpec + 49, // 16: openshell.v1.Sandbox.status:type_name -> openshell.v1.SandboxStatus + 48, // 17: openshell.v1.Sandbox.created_from_workload_template:type_name -> openshell.v1.SandboxWorkloadTemplateProvenance + 252, // 18: openshell.v1.SandboxSpec.environment:type_name -> openshell.v1.SandboxSpec.EnvironmentEntry + 41, // 19: openshell.v1.SandboxSpec.template:type_name -> openshell.v1.SandboxTemplate + 277, // 20: openshell.v1.SandboxSpec.policy:type_name -> openshell.sandbox.v1.SandboxPolicy + 39, // 21: openshell.v1.SandboxSpec.resource_requirements:type_name -> openshell.v1.ResourceRequirements + 16, // 22: openshell.v1.SandboxSpec.restart_policy:type_name -> openshell.v1.SandboxRestartPolicy + 40, // 23: openshell.v1.ResourceRequirements.gpu:type_name -> openshell.v1.GpuResourceRequirements + 253, // 24: openshell.v1.SandboxTemplate.labels:type_name -> openshell.v1.SandboxTemplate.LabelsEntry + 254, // 25: openshell.v1.SandboxTemplate.annotations:type_name -> openshell.v1.SandboxTemplate.AnnotationsEntry + 255, // 26: openshell.v1.SandboxTemplate.environment:type_name -> openshell.v1.SandboxTemplate.EnvironmentEntry + 278, // 27: openshell.v1.SandboxTemplate.resources:type_name -> google.protobuf.Struct + 278, // 28: openshell.v1.SandboxTemplate.driver_config:type_name -> google.protobuf.Struct + 276, // 29: openshell.v1.SandboxWorkloadTemplate.metadata:type_name -> openshell.datamodel.v1.ObjectMeta + 43, // 30: openshell.v1.SandboxWorkloadTemplate.spec:type_name -> openshell.v1.SandboxWorkloadTemplateSpec + 44, // 31: openshell.v1.SandboxWorkloadTemplateSpec.workload:type_name -> openshell.v1.SandboxWorkloadConfig + 278, // 32: openshell.v1.SandboxWorkloadTemplateSpec.driver_config:type_name -> google.protobuf.Struct + 46, // 33: openshell.v1.SandboxWorkloadTemplateSpec.desired_service_level:type_name -> openshell.v1.SandboxServiceLevel + 256, // 34: openshell.v1.SandboxWorkloadConfig.environment:type_name -> openshell.v1.SandboxWorkloadConfig.EnvironmentEntry + 45, // 35: openshell.v1.SandboxWorkloadConfig.resources:type_name -> openshell.v1.SandboxResources + 40, // 36: openshell.v1.SandboxResources.gpu:type_name -> openshell.v1.GpuResourceRequirements + 47, // 37: openshell.v1.SandboxServiceLevel.startup:type_name -> openshell.v1.SandboxStartup + 279, // 38: openshell.v1.SandboxStartup.ready_within:type_name -> google.protobuf.Duration + 50, // 39: openshell.v1.SandboxStatus.conditions:type_name -> openshell.v1.SandboxCondition + 1, // 40: openshell.v1.SandboxStatus.phase:type_name -> openshell.v1.SandboxPhase + 249, // 41: openshell.v1.SandboxStatus.endpoint_statuses:type_name -> openshell.v1.EndpointStatus + 176, // 42: openshell.v1.SandboxStatus.configuration_admission:type_name -> openshell.v1.SandboxConfigurationAdmission + 250, // 43: openshell.v1.SandboxStatus.provisioning:type_name -> openshell.v1.SandboxProvisioning + 275, // 44: openshell.v1.SandboxStatus.next_restart_time:type_name -> google.protobuf.Timestamp + 275, // 45: openshell.v1.SandboxStatus.main_process_started_time:type_name -> google.protobuf.Timestamp + 275, // 46: openshell.v1.SandboxCondition.transition_time:type_name -> google.protobuf.Timestamp + 275, // 47: openshell.v1.PlatformEvent.event_time:type_name -> google.protobuf.Timestamp + 257, // 48: openshell.v1.PlatformEvent.metadata:type_name -> openshell.v1.PlatformEvent.MetadataEntry + 280, // 49: openshell.v1.CreateSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 38, // 50: openshell.v1.CreateSandboxRequest.spec:type_name -> openshell.v1.SandboxSpec + 258, // 51: openshell.v1.CreateSandboxRequest.labels:type_name -> openshell.v1.CreateSandboxRequest.LabelsEntry + 259, // 52: openshell.v1.CreateSandboxRequest.annotations:type_name -> openshell.v1.CreateSandboxRequest.AnnotationsEntry + 251, // 53: openshell.v1.CreateSandboxRequest.service_exposures:type_name -> openshell.v1.SandboxServiceExposure + 280, // 54: openshell.v1.CreateSandboxTemplateRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 42, // 55: openshell.v1.CreateSandboxTemplateRequest.template:type_name -> openshell.v1.SandboxWorkloadTemplate + 280, // 56: openshell.v1.GetSandboxTemplateRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 57: openshell.v1.ListSandboxTemplatesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 58: openshell.v1.DeleteSandboxTemplateRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 42, // 59: openshell.v1.SandboxTemplateResponse.template:type_name -> openshell.v1.SandboxWorkloadTemplate + 42, // 60: openshell.v1.ListSandboxTemplatesResponse.templates:type_name -> openshell.v1.SandboxWorkloadTemplate + 17, // 61: openshell.v1.DeleteSandboxTemplateResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 280, // 62: openshell.v1.BeginRootfsTarStagingRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 275, // 63: openshell.v1.BeginRootfsTarStagingResponse.expiration_time:type_name -> google.protobuf.Timestamp + 280, // 64: openshell.v1.GetSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 65: openshell.v1.ListSandboxesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 66: openshell.v1.ListSandboxProvidersRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 67: openshell.v1.AttachSandboxProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 68: openshell.v1.DetachSandboxProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 69: openshell.v1.DeleteSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 70: openshell.v1.StopSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 71: openshell.v1.StartSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 37, // 72: openshell.v1.SandboxResponse.sandbox:type_name -> openshell.v1.Sandbox + 260, // 73: openshell.v1.SandboxResponse.service_urls:type_name -> openshell.v1.SandboxResponse.ServiceUrlsEntry + 37, // 74: openshell.v1.ListSandboxesResponse.sandboxes:type_name -> openshell.v1.Sandbox + 281, // 75: openshell.v1.ListSandboxProvidersResponse.providers:type_name -> openshell.datamodel.v1.Provider + 37, // 76: openshell.v1.AttachSandboxProviderResponse.sandbox:type_name -> openshell.v1.Sandbox + 79, // 77: openshell.v1.AttachSandboxProviderResponse.receipt:type_name -> openshell.v1.ProviderMutationReceipt + 37, // 78: openshell.v1.DetachSandboxProviderResponse.sandbox:type_name -> openshell.v1.Sandbox + 79, // 79: openshell.v1.DetachSandboxProviderResponse.receipt:type_name -> openshell.v1.ProviderMutationReceipt + 77, // 80: openshell.v1.ConfigSnapshotRevision.sandbox_config:type_name -> openshell.v1.SandboxConfigRevision + 75, // 81: openshell.v1.ConfigSnapshotRevision.provider_target:type_name -> openshell.v1.ProviderDesiredIdentity + 282, // 82: openshell.v1.SandboxConfigRevision.policy_source:type_name -> openshell.sandbox.v1.PolicySource + 5, // 83: openshell.v1.ConfigUpdateOperation.component:type_name -> openshell.v1.ConfigComponent + 76, // 84: openshell.v1.ConfigUpdateOperation.target_revision:type_name -> openshell.v1.ConfigSnapshotRevision + 7, // 85: openshell.v1.ConfigUpdateOperation.state:type_name -> openshell.v1.ConfigUpdateOperationState + 6, // 86: openshell.v1.ConfigUpdateOperation.outcome:type_name -> openshell.v1.ConfigApplyOutcome + 275, // 87: openshell.v1.ConfigUpdateOperation.created_time:type_name -> google.protobuf.Timestamp + 275, // 88: openshell.v1.ConfigUpdateOperation.updated_time:type_name -> google.protobuf.Timestamp + 275, // 89: openshell.v1.ConfigUpdateOperation.completed_time:type_name -> google.protobuf.Timestamp + 2, // 90: openshell.v1.ProviderMutationReceipt.kind:type_name -> openshell.v1.ProviderMutationKind + 75, // 91: openshell.v1.ProviderMutationReceipt.desired:type_name -> openshell.v1.ProviderDesiredIdentity + 275, // 92: openshell.v1.ProviderMutationReceipt.persisted_time:type_name -> google.protobuf.Timestamp + 4, // 93: openshell.v1.ProviderReadinessObservation.reason:type_name -> openshell.v1.ProviderReadinessReason + 79, // 94: openshell.v1.ProviderReadinessStatus.receipt:type_name -> openshell.v1.ProviderMutationReceipt + 3, // 95: openshell.v1.ProviderReadinessStatus.state:type_name -> openshell.v1.ProviderReadinessState + 4, // 96: openshell.v1.ProviderReadinessStatus.reason:type_name -> openshell.v1.ProviderReadinessReason + 80, // 97: openshell.v1.ProviderReadinessStatus.observed:type_name -> openshell.v1.ProviderReadinessObservation + 275, // 98: openshell.v1.ProviderReadinessStatus.observed_time:type_name -> google.protobuf.Timestamp + 275, // 99: openshell.v1.ProviderReadinessStatus.evaluated_time:type_name -> google.protobuf.Timestamp + 78, // 100: openshell.v1.ProviderReadinessStatus.operation:type_name -> openshell.v1.ConfigUpdateOperation + 280, // 101: openshell.v1.GetSandboxProviderStatusRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 81, // 102: openshell.v1.GetSandboxProviderStatusResponse.status:type_name -> openshell.v1.ProviderReadinessStatus + 80, // 103: openshell.v1.ReportProviderReadinessRequest.observation:type_name -> openshell.v1.ProviderReadinessObservation + 279, // 104: openshell.v1.ReportProviderReadinessResponse.report_interval:type_name -> google.protobuf.Duration + 279, // 105: openshell.v1.ReportProviderReadinessResponse.observation_ttl:type_name -> google.protobuf.Duration + 17, // 106: openshell.v1.DeleteSandboxResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 280, // 107: openshell.v1.CreateSshSessionRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 275, // 108: openshell.v1.CreateSshSessionResponse.expiration_time:type_name -> google.protobuf.Timestamp + 280, // 109: openshell.v1.ExposeServiceRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 19, // 110: openshell.v1.ExposeServiceRequest.authorization_mode:type_name -> openshell.v1.ServiceAuthorizationMode + 280, // 111: openshell.v1.GetServiceRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 112: openshell.v1.ListServicesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 96, // 113: openshell.v1.ListServicesResponse.services:type_name -> openshell.v1.ServiceEndpointResponse + 280, // 114: openshell.v1.DeleteServiceRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 17, // 115: openshell.v1.DeleteServiceResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 276, // 116: openshell.v1.ServiceEndpoint.metadata:type_name -> openshell.datamodel.v1.ObjectMeta + 19, // 117: openshell.v1.ServiceEndpoint.authorization_mode:type_name -> openshell.v1.ServiceAuthorizationMode + 95, // 118: openshell.v1.ServiceEndpointResponse.endpoint:type_name -> openshell.v1.ServiceEndpoint + 17, // 119: openshell.v1.RevokeSshSessionResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 280, // 120: openshell.v1.ExecSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 261, // 121: openshell.v1.ExecSandboxRequest.environment:type_name -> openshell.v1.ExecSandboxRequest.EnvironmentEntry + 279, // 122: openshell.v1.ExecSandboxRequest.execution_timeout:type_name -> google.protobuf.Duration + 100, // 123: openshell.v1.ExecSandboxEvent.stdout:type_name -> openshell.v1.ExecSandboxStdout + 101, // 124: openshell.v1.ExecSandboxEvent.stderr:type_name -> openshell.v1.ExecSandboxStderr + 102, // 125: openshell.v1.ExecSandboxEvent.exit:type_name -> openshell.v1.ExecSandboxExit + 196, // 126: openshell.v1.TcpForwardInit.ssh:type_name -> openshell.v1.SshRelayTarget + 197, // 127: openshell.v1.TcpForwardInit.tcp:type_name -> openshell.v1.TcpRelayTarget + 104, // 128: openshell.v1.TcpForwardFrame.init:type_name -> openshell.v1.TcpForwardInit + 99, // 129: openshell.v1.ExecSandboxInput.start:type_name -> openshell.v1.ExecSandboxRequest + 107, // 130: openshell.v1.ExecSandboxInput.resize:type_name -> openshell.v1.ExecSandboxWindowResize + 276, // 131: openshell.v1.SshSession.metadata:type_name -> openshell.datamodel.v1.ObjectMeta + 275, // 132: openshell.v1.SshSession.expiration_time:type_name -> google.protobuf.Timestamp + 280, // 133: openshell.v1.WatchSandboxRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 275, // 134: openshell.v1.WatchSandboxRequest.since_time:type_name -> google.protobuf.Timestamp + 37, // 135: openshell.v1.SandboxStreamEvent.sandbox:type_name -> openshell.v1.Sandbox + 111, // 136: openshell.v1.SandboxStreamEvent.log:type_name -> openshell.v1.SandboxLogLine + 51, // 137: openshell.v1.SandboxStreamEvent.event:type_name -> openshell.v1.PlatformEvent + 112, // 138: openshell.v1.SandboxStreamEvent.warning:type_name -> openshell.v1.SandboxStreamWarning + 209, // 139: openshell.v1.SandboxStreamEvent.draft_policy_update:type_name -> openshell.v1.DraftPolicyUpdate + 275, // 140: openshell.v1.SandboxLogLine.event_time:type_name -> google.protobuf.Timestamp + 262, // 141: openshell.v1.SandboxLogLine.fields:type_name -> openshell.v1.SandboxLogLine.FieldsEntry + 280, // 142: openshell.v1.CreateProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 281, // 143: openshell.v1.CreateProviderRequest.provider:type_name -> openshell.datamodel.v1.Provider + 280, // 144: openshell.v1.GetProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 145: openshell.v1.ListProvidersRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 146: openshell.v1.UpdateProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 281, // 147: openshell.v1.UpdateProviderRequest.provider:type_name -> openshell.datamodel.v1.Provider + 263, // 148: openshell.v1.UpdateProviderRequest.credential_expiration_times:type_name -> openshell.v1.UpdateProviderRequest.CredentialExpirationTimesEntry + 280, // 149: openshell.v1.DeleteProviderRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 281, // 150: openshell.v1.ProviderResponse.provider:type_name -> openshell.datamodel.v1.Provider + 79, // 151: openshell.v1.ProviderResponse.target_receipts:type_name -> openshell.v1.ProviderMutationReceipt + 281, // 152: openshell.v1.ListProvidersResponse.providers:type_name -> openshell.datamodel.v1.Provider + 280, // 153: openshell.v1.ListProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 154: openshell.v1.GetProviderProfileRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 141, // 155: openshell.v1.ProviderProfileImportItem.profile:type_name -> openshell.v1.ProviderProfile + 279, // 156: openshell.v1.ProviderCredentialTokenGrant.cache_ttl:type_name -> google.protobuf.Duration + 124, // 157: openshell.v1.ProviderCredentialTokenGrant.audience_overrides:type_name -> openshell.v1.ProviderCredentialTokenGrantAudienceOverride + 8, // 158: openshell.v1.ProviderCredentialTokenGrant.grant_type:type_name -> openshell.v1.ProviderCredentialTokenGrantType + 125, // 159: openshell.v1.ProviderCredentialTokenGrant.subject_token:type_name -> openshell.v1.ProviderCredentialTokenGrantSubjectToken + 130, // 160: openshell.v1.ProviderProfileCredential.refresh:type_name -> openshell.v1.ProviderCredentialRefresh + 126, // 161: openshell.v1.ProviderProfileCredential.token_grant:type_name -> openshell.v1.ProviderCredentialTokenGrant + 9, // 162: openshell.v1.ProviderCredentialRefresh.strategy:type_name -> openshell.v1.ProviderCredentialRefreshStrategy + 279, // 163: openshell.v1.ProviderCredentialRefresh.refresh_before:type_name -> google.protobuf.Duration + 279, // 164: openshell.v1.ProviderCredentialRefresh.max_lifetime:type_name -> google.protobuf.Duration + 128, // 165: openshell.v1.ProviderCredentialRefresh.material:type_name -> openshell.v1.ProviderCredentialRefreshMaterial + 129, // 166: openshell.v1.ProviderCredentialRefresh.additional_outputs:type_name -> openshell.v1.ProviderCredentialRefreshOutput + 9, // 167: openshell.v1.ProviderCredentialRefreshStatus.strategy:type_name -> openshell.v1.ProviderCredentialRefreshStrategy + 275, // 168: openshell.v1.ProviderCredentialRefreshStatus.expiration_time:type_name -> google.protobuf.Timestamp + 275, // 169: openshell.v1.ProviderCredentialRefreshStatus.next_refresh_time:type_name -> google.protobuf.Timestamp + 275, // 170: openshell.v1.ProviderCredentialRefreshStatus.last_refresh_time:type_name -> google.protobuf.Timestamp + 15, // 171: openshell.v1.ProviderCredentialRefreshStatus.recovery_action:type_name -> openshell.v1.ProviderCredentialRefreshRecoveryAction + 275, // 172: openshell.v1.ProviderCredentialRefreshStatus.last_error_time:type_name -> google.protobuf.Timestamp + 280, // 173: openshell.v1.GetProviderRefreshStatusRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 131, // 174: openshell.v1.GetProviderRefreshStatusResponse.credentials:type_name -> openshell.v1.ProviderCredentialRefreshStatus + 280, // 175: openshell.v1.ConfigureProviderRefreshRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 9, // 176: openshell.v1.ConfigureProviderRefreshRequest.strategy:type_name -> openshell.v1.ProviderCredentialRefreshStrategy + 264, // 177: openshell.v1.ConfigureProviderRefreshRequest.material:type_name -> openshell.v1.ConfigureProviderRefreshRequest.MaterialEntry + 275, // 178: openshell.v1.ConfigureProviderRefreshRequest.expiration_time:type_name -> google.protobuf.Timestamp + 131, // 179: openshell.v1.ConfigureProviderRefreshResponse.status:type_name -> openshell.v1.ProviderCredentialRefreshStatus + 280, // 180: openshell.v1.RotateProviderCredentialRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 131, // 181: openshell.v1.RotateProviderCredentialResponse.status:type_name -> openshell.v1.ProviderCredentialRefreshStatus + 280, // 182: openshell.v1.DeleteProviderRefreshRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 17, // 183: openshell.v1.DeleteProviderRefreshResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 10, // 184: openshell.v1.ProviderProfile.category:type_name -> openshell.v1.ProviderProfileCategory + 127, // 185: openshell.v1.ProviderProfile.credentials:type_name -> openshell.v1.ProviderProfileCredential + 283, // 186: openshell.v1.ProviderProfile.endpoints:type_name -> openshell.sandbox.v1.NetworkEndpoint + 284, // 187: openshell.v1.ProviderProfile.binaries:type_name -> openshell.sandbox.v1.NetworkBinary + 132, // 188: openshell.v1.ProviderProfile.discovery:type_name -> openshell.v1.ProviderProfileDiscovery + 265, // 189: openshell.v1.ProviderProfile.annotations:type_name -> openshell.v1.ProviderProfile.AnnotationsEntry + 142, // 190: openshell.v1.ProviderProfile.files:type_name -> openshell.v1.ProviderProfileFile + 141, // 191: openshell.v1.ProviderProfileResponse.profile:type_name -> openshell.v1.ProviderProfile + 141, // 192: openshell.v1.ListProviderProfilesResponse.profiles:type_name -> openshell.v1.ProviderProfile + 280, // 193: openshell.v1.ImportProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 122, // 194: openshell.v1.ImportProviderProfilesRequest.profiles:type_name -> openshell.v1.ProviderProfileImportItem + 123, // 195: openshell.v1.ImportProviderProfilesResponse.diagnostics:type_name -> openshell.v1.ProviderProfileDiagnostic + 141, // 196: openshell.v1.ImportProviderProfilesResponse.profiles:type_name -> openshell.v1.ProviderProfile + 280, // 197: openshell.v1.UpdateProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 122, // 198: openshell.v1.UpdateProviderProfilesRequest.profile:type_name -> openshell.v1.ProviderProfileImportItem + 123, // 199: openshell.v1.UpdateProviderProfilesResponse.diagnostics:type_name -> openshell.v1.ProviderProfileDiagnostic + 141, // 200: openshell.v1.UpdateProviderProfilesResponse.profile:type_name -> openshell.v1.ProviderProfile + 280, // 201: openshell.v1.LintProviderProfilesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 122, // 202: openshell.v1.LintProviderProfilesRequest.profiles:type_name -> openshell.v1.ProviderProfileImportItem + 123, // 203: openshell.v1.LintProviderProfilesResponse.diagnostics:type_name -> openshell.v1.ProviderProfileDiagnostic + 17, // 204: openshell.v1.DeleteProviderResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 280, // 205: openshell.v1.DeleteProviderProfileRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 17, // 206: openshell.v1.DeleteProviderProfileResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 155, // 207: openshell.v1.StaticCredentialBinding.endpoints:type_name -> openshell.v1.StaticCredentialEndpointBinding + 266, // 208: openshell.v1.GetSandboxProviderEnvironmentResponse.environment:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.EnvironmentEntry + 267, // 209: openshell.v1.GetSandboxProviderEnvironmentResponse.credential_expiration_times:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.CredentialExpirationTimesEntry + 268, // 210: openshell.v1.GetSandboxProviderEnvironmentResponse.dynamic_credentials:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.DynamicCredentialsEntry + 269, // 211: openshell.v1.GetSandboxProviderEnvironmentResponse.static_credential_bindings:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.StaticCredentialBindingsEntry + 4, // 212: openshell.v1.GetSandboxProviderEnvironmentResponse.readiness_reason:type_name -> openshell.v1.ProviderReadinessReason + 270, // 213: openshell.v1.GetSandboxProviderEnvironmentResponse.files:type_name -> openshell.v1.GetSandboxProviderEnvironmentResponse.FilesEntry + 279, // 214: openshell.v1.ExchangeProviderSubjectTokenResponse.expires_after:type_name -> google.protobuf.Duration + 280, // 215: openshell.v1.UpdateConfigRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 277, // 216: openshell.v1.UpdateConfigRequest.policy:type_name -> openshell.sandbox.v1.SandboxPolicy + 285, // 217: openshell.v1.UpdateConfigRequest.setting_value:type_name -> openshell.sandbox.v1.SettingValue + 161, // 218: openshell.v1.UpdateConfigRequest.merge_operations:type_name -> openshell.v1.PolicyMergeOperation + 271, // 219: openshell.v1.UpdateConfigRequest.annotations:type_name -> openshell.v1.UpdateConfigRequest.AnnotationsEntry + 162, // 220: openshell.v1.PolicyMergeOperation.add_rule:type_name -> openshell.v1.AddNetworkRule + 163, // 221: openshell.v1.PolicyMergeOperation.remove_endpoint:type_name -> openshell.v1.RemoveNetworkEndpoint + 164, // 222: openshell.v1.PolicyMergeOperation.remove_rule:type_name -> openshell.v1.RemoveNetworkRule + 166, // 223: openshell.v1.PolicyMergeOperation.add_deny_rules:type_name -> openshell.v1.AddDenyRules + 167, // 224: openshell.v1.PolicyMergeOperation.add_allow_rules:type_name -> openshell.v1.AddAllowRules + 168, // 225: openshell.v1.PolicyMergeOperation.remove_binary:type_name -> openshell.v1.RemoveNetworkBinary + 286, // 226: openshell.v1.AddNetworkRule.rule:type_name -> openshell.sandbox.v1.NetworkPolicyRule + 284, // 227: openshell.v1.L7RuleTarget.binaries:type_name -> openshell.sandbox.v1.NetworkBinary + 287, // 228: openshell.v1.AddDenyRules.deny_rules:type_name -> openshell.sandbox.v1.L7DenyRule + 165, // 229: openshell.v1.AddDenyRules.target:type_name -> openshell.v1.L7RuleTarget + 288, // 230: openshell.v1.AddAllowRules.rules:type_name -> openshell.sandbox.v1.L7Rule + 165, // 231: openshell.v1.AddAllowRules.target:type_name -> openshell.v1.L7RuleTarget + 272, // 232: openshell.v1.UpdateConfigResponse.annotations:type_name -> openshell.v1.UpdateConfigResponse.AnnotationsEntry + 280, // 233: openshell.v1.GetSandboxPolicyStatusRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 179, // 234: openshell.v1.GetSandboxPolicyStatusResponse.revision:type_name -> openshell.v1.SandboxPolicyRevision + 280, // 235: openshell.v1.ListSandboxPoliciesRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 179, // 236: openshell.v1.ListSandboxPoliciesResponse.revisions:type_name -> openshell.v1.SandboxPolicyRevision + 12, // 237: openshell.v1.ReportPolicyStatusRequest.status:type_name -> openshell.v1.PolicyStatus + 11, // 238: openshell.v1.SandboxConfigurationAdmission.state:type_name -> openshell.v1.ConfigurationAdmissionState + 176, // 239: openshell.v1.ReportSandboxConfigurationRequest.admission:type_name -> openshell.v1.SandboxConfigurationAdmission + 12, // 240: openshell.v1.SandboxPolicyRevision.status:type_name -> openshell.v1.PolicyStatus + 275, // 241: openshell.v1.SandboxPolicyRevision.created_time:type_name -> google.protobuf.Timestamp + 275, // 242: openshell.v1.SandboxPolicyRevision.loaded_time:type_name -> google.protobuf.Timestamp + 277, // 243: openshell.v1.SandboxPolicyRevision.policy:type_name -> openshell.sandbox.v1.SandboxPolicy + 273, // 244: openshell.v1.SandboxPolicyRevision.provenance:type_name -> openshell.v1.SandboxPolicyRevision.ProvenanceEntry + 280, // 245: openshell.v1.GetSandboxLogsRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 275, // 246: openshell.v1.GetSandboxLogsRequest.since_time:type_name -> google.protobuf.Timestamp + 111, // 247: openshell.v1.PushSandboxLogsRequest.logs:type_name -> openshell.v1.SandboxLogLine + 111, // 248: openshell.v1.GetSandboxLogsResponse.logs:type_name -> openshell.v1.SandboxLogLine + 186, // 249: openshell.v1.SupervisorMessage.hello:type_name -> openshell.v1.SupervisorHello + 189, // 250: openshell.v1.SupervisorMessage.heartbeat:type_name -> openshell.v1.SupervisorHeartbeat + 202, // 251: openshell.v1.SupervisorMessage.relay_open_result:type_name -> openshell.v1.RelayOpenResult + 203, // 252: openshell.v1.SupervisorMessage.relay_close:type_name -> openshell.v1.RelayClose + 187, // 253: openshell.v1.GatewayMessage.session_accepted:type_name -> openshell.v1.SessionAccepted + 188, // 254: openshell.v1.GatewayMessage.session_rejected:type_name -> openshell.v1.SessionRejected + 190, // 255: openshell.v1.GatewayMessage.heartbeat:type_name -> openshell.v1.GatewayHeartbeat + 195, // 256: openshell.v1.GatewayMessage.relay_open:type_name -> openshell.v1.RelayOpen + 203, // 257: openshell.v1.GatewayMessage.relay_close:type_name -> openshell.v1.RelayClose + 279, // 258: openshell.v1.SessionAccepted.heartbeat_interval:type_name -> google.protobuf.Duration + 196, // 259: openshell.v1.RelayOpen.ssh:type_name -> openshell.v1.SshRelayTarget + 197, // 260: openshell.v1.RelayOpen.tcp:type_name -> openshell.v1.TcpRelayTarget + 198, // 261: openshell.v1.RelayFrame.init:type_name -> openshell.v1.RelayInit + 195, // 262: openshell.v1.PeerRelayInit.relay_open:type_name -> openshell.v1.RelayOpen + 200, // 263: openshell.v1.PeerRelayFrame.init:type_name -> openshell.v1.PeerRelayInit + 275, // 264: openshell.v1.DenialSummary.first_seen_time:type_name -> google.protobuf.Timestamp + 275, // 265: openshell.v1.DenialSummary.last_seen_time:type_name -> google.protobuf.Timestamp + 204, // 266: openshell.v1.DenialSummary.l7_request_samples:type_name -> openshell.v1.L7RequestSample + 206, // 267: openshell.v1.NetworkActivitySummary.denials_by_group:type_name -> openshell.v1.DenialGroupCount + 286, // 268: openshell.v1.PolicyChunk.proposed_rule:type_name -> openshell.sandbox.v1.NetworkPolicyRule + 275, // 269: openshell.v1.PolicyChunk.created_time:type_name -> google.protobuf.Timestamp + 275, // 270: openshell.v1.PolicyChunk.decided_time:type_name -> google.protobuf.Timestamp + 275, // 271: openshell.v1.PolicyChunk.first_seen_time:type_name -> google.protobuf.Timestamp + 275, // 272: openshell.v1.PolicyChunk.last_seen_time:type_name -> google.protobuf.Timestamp + 277, // 273: openshell.v1.PolicyChunk.current_effective_policy:type_name -> openshell.sandbox.v1.SandboxPolicy + 277, // 274: openshell.v1.PolicyChunk.candidate_effective_policy:type_name -> openshell.sandbox.v1.SandboxPolicy + 280, // 275: openshell.v1.SubmitPolicyAnalysisRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 205, // 276: openshell.v1.SubmitPolicyAnalysisRequest.summaries:type_name -> openshell.v1.DenialSummary + 208, // 277: openshell.v1.SubmitPolicyAnalysisRequest.proposed_chunks:type_name -> openshell.v1.PolicyChunk + 207, // 278: openshell.v1.SubmitPolicyAnalysisRequest.network_activity_summaries:type_name -> openshell.v1.NetworkActivitySummary + 280, // 279: openshell.v1.GetDraftPolicyRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 208, // 280: openshell.v1.GetDraftPolicyResponse.chunks:type_name -> openshell.v1.PolicyChunk + 275, // 281: openshell.v1.GetDraftPolicyResponse.last_analyzed_time:type_name -> google.protobuf.Timestamp + 280, // 282: openshell.v1.ApproveDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 283: openshell.v1.RejectDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 284: openshell.v1.ApproveAllDraftChunksRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 218, // 285: openshell.v1.ApproveAllDraftChunksRequest.approvals:type_name -> openshell.v1.DraftChunkApproval + 280, // 286: openshell.v1.EditDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 286, // 287: openshell.v1.EditDraftChunkRequest.proposed_rule:type_name -> openshell.sandbox.v1.NetworkPolicyRule + 280, // 288: openshell.v1.UndoDraftChunkRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 289: openshell.v1.ClearDraftChunksRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 280, // 290: openshell.v1.GetDraftHistoryRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 275, // 291: openshell.v1.DraftHistoryEntry.event_time:type_name -> google.protobuf.Timestamp + 228, // 292: openshell.v1.GetDraftHistoryResponse.entries:type_name -> openshell.v1.DraftHistoryEntry + 274, // 293: openshell.v1.CreateWorkspaceRequest.labels:type_name -> openshell.v1.CreateWorkspaceRequest.LabelsEntry + 289, // 294: openshell.v1.CreateWorkspaceResponse.workspace:type_name -> openshell.datamodel.v1.Workspace + 289, // 295: openshell.v1.GetWorkspaceResponse.workspace:type_name -> openshell.datamodel.v1.Workspace + 289, // 296: openshell.v1.ListWorkspacesResponse.workspaces:type_name -> openshell.datamodel.v1.Workspace + 17, // 297: openshell.v1.DeleteWorkspaceResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 276, // 298: openshell.v1.WorkspaceMember.metadata:type_name -> openshell.datamodel.v1.ObjectMeta + 14, // 299: openshell.v1.WorkspaceMember.role:type_name -> openshell.v1.WorkspaceRole + 280, // 300: openshell.v1.AddWorkspaceMemberRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 14, // 301: openshell.v1.AddWorkspaceMemberRequest.role:type_name -> openshell.v1.WorkspaceRole + 238, // 302: openshell.v1.AddWorkspaceMemberResponse.member:type_name -> openshell.v1.WorkspaceMember + 280, // 303: openshell.v1.RemoveWorkspaceMemberRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 17, // 304: openshell.v1.RemoveWorkspaceMemberResponse.outcome:type_name -> openshell.v1.DeletionOutcome + 280, // 305: openshell.v1.ListWorkspaceMembersRequest.workspace_scope:type_name -> openshell.datamodel.v1.WorkspaceSelector + 238, // 306: openshell.v1.ListWorkspaceMembersResponse.members:type_name -> openshell.v1.WorkspaceMember + 275, // 307: openshell.v1.ExtensionServiceCredential.expiration_time:type_name -> google.protobuf.Timestamp + 18, // 308: openshell.v1.EndpointObservation.result:type_name -> openshell.v1.EndpointResult + 246, // 309: openshell.v1.ReportEndpointStatusRequest.observations:type_name -> openshell.v1.EndpointObservation + 18, // 310: openshell.v1.EndpointStatus.last_result:type_name -> openshell.v1.EndpointResult + 275, // 311: openshell.v1.EndpointStatus.last_reported_time:type_name -> google.protobuf.Timestamp + 275, // 312: openshell.v1.SandboxProvisioning.configuration_change_time:type_name -> google.protobuf.Timestamp + 275, // 313: openshell.v1.SandboxProvisioning.first_rejection_time:type_name -> google.protobuf.Timestamp + 275, // 314: openshell.v1.SandboxProvisioning.deadline:type_name -> google.protobuf.Timestamp + 275, // 315: openshell.v1.SandboxProvisioning.timeout_time:type_name -> google.protobuf.Timestamp + 275, // 316: openshell.v1.SandboxProvisioning.cleanup_completed_time:type_name -> google.protobuf.Timestamp + 275, // 317: openshell.v1.SandboxProvisioning.cleanup_retry_time:type_name -> google.protobuf.Timestamp + 275, // 318: openshell.v1.SandboxProvisioning.attachment_change_time:type_name -> google.protobuf.Timestamp + 19, // 319: openshell.v1.SandboxServiceExposure.authorization_mode:type_name -> openshell.v1.ServiceAuthorizationMode + 275, // 320: openshell.v1.UpdateProviderRequest.CredentialExpirationTimesEntry.value:type_name -> google.protobuf.Timestamp + 275, // 321: openshell.v1.GetSandboxProviderEnvironmentResponse.CredentialExpirationTimesEntry.value:type_name -> google.protobuf.Timestamp + 127, // 322: openshell.v1.GetSandboxProviderEnvironmentResponse.DynamicCredentialsEntry.value:type_name -> openshell.v1.ProviderProfileCredential + 156, // 323: openshell.v1.GetSandboxProviderEnvironmentResponse.StaticCredentialBindingsEntry.value:type_name -> openshell.v1.StaticCredentialBinding + 24, // 324: openshell.v1.OpenShell.Health:input_type -> openshell.v1.HealthRequest + 26, // 325: openshell.v1.OpenShell.GetCurrentUser:input_type -> openshell.v1.GetCurrentUserRequest + 28, // 326: openshell.v1.OpenShell.GetGatewayInfo:input_type -> openshell.v1.GetGatewayInfoRequest + 52, // 327: openshell.v1.OpenShell.CreateSandbox:input_type -> openshell.v1.CreateSandboxRequest + 60, // 328: openshell.v1.OpenShell.BeginRootfsTarStaging:input_type -> openshell.v1.BeginRootfsTarStagingRequest + 62, // 329: openshell.v1.OpenShell.GetSandbox:input_type -> openshell.v1.GetSandboxRequest + 63, // 330: openshell.v1.OpenShell.ListSandboxes:input_type -> openshell.v1.ListSandboxesRequest + 53, // 331: openshell.v1.OpenShell.CreateSandboxTemplate:input_type -> openshell.v1.CreateSandboxTemplateRequest + 54, // 332: openshell.v1.OpenShell.GetSandboxTemplate:input_type -> openshell.v1.GetSandboxTemplateRequest + 55, // 333: openshell.v1.OpenShell.ListSandboxTemplates:input_type -> openshell.v1.ListSandboxTemplatesRequest + 56, // 334: openshell.v1.OpenShell.DeleteSandboxTemplate:input_type -> openshell.v1.DeleteSandboxTemplateRequest + 64, // 335: openshell.v1.OpenShell.ListSandboxProviders:input_type -> openshell.v1.ListSandboxProvidersRequest + 65, // 336: openshell.v1.OpenShell.AttachSandboxProvider:input_type -> openshell.v1.AttachSandboxProviderRequest + 66, // 337: openshell.v1.OpenShell.DetachSandboxProvider:input_type -> openshell.v1.DetachSandboxProviderRequest + 82, // 338: openshell.v1.OpenShell.GetSandboxProviderStatus:input_type -> openshell.v1.GetSandboxProviderStatusRequest + 67, // 339: openshell.v1.OpenShell.DeleteSandbox:input_type -> openshell.v1.DeleteSandboxRequest + 68, // 340: openshell.v1.OpenShell.StopSandbox:input_type -> openshell.v1.StopSandboxRequest + 69, // 341: openshell.v1.OpenShell.StartSandbox:input_type -> openshell.v1.StartSandboxRequest + 87, // 342: openshell.v1.OpenShell.CreateSshSession:input_type -> openshell.v1.CreateSshSessionRequest + 89, // 343: openshell.v1.OpenShell.ExposeService:input_type -> openshell.v1.ExposeServiceRequest + 90, // 344: openshell.v1.OpenShell.GetService:input_type -> openshell.v1.GetServiceRequest + 91, // 345: openshell.v1.OpenShell.ListServices:input_type -> openshell.v1.ListServicesRequest + 93, // 346: openshell.v1.OpenShell.DeleteService:input_type -> openshell.v1.DeleteServiceRequest + 97, // 347: openshell.v1.OpenShell.RevokeSshSession:input_type -> openshell.v1.RevokeSshSessionRequest + 99, // 348: openshell.v1.OpenShell.ExecSandbox:input_type -> openshell.v1.ExecSandboxRequest + 105, // 349: openshell.v1.OpenShell.ForwardTcp:input_type -> openshell.v1.TcpForwardFrame + 106, // 350: openshell.v1.OpenShell.ExecSandboxInteractive:input_type -> openshell.v1.ExecSandboxInput + 113, // 351: openshell.v1.OpenShell.CreateProvider:input_type -> openshell.v1.CreateProviderRequest + 114, // 352: openshell.v1.OpenShell.GetProvider:input_type -> openshell.v1.GetProviderRequest + 115, // 353: openshell.v1.OpenShell.ListProviders:input_type -> openshell.v1.ListProvidersRequest + 120, // 354: openshell.v1.OpenShell.ListProviderProfiles:input_type -> openshell.v1.ListProviderProfilesRequest + 121, // 355: openshell.v1.OpenShell.GetProviderProfile:input_type -> openshell.v1.GetProviderProfileRequest + 145, // 356: openshell.v1.OpenShell.ImportProviderProfiles:input_type -> openshell.v1.ImportProviderProfilesRequest + 147, // 357: openshell.v1.OpenShell.UpdateProviderProfiles:input_type -> openshell.v1.UpdateProviderProfilesRequest + 149, // 358: openshell.v1.OpenShell.LintProviderProfiles:input_type -> openshell.v1.LintProviderProfilesRequest + 116, // 359: openshell.v1.OpenShell.UpdateProvider:input_type -> openshell.v1.UpdateProviderRequest + 133, // 360: openshell.v1.OpenShell.GetProviderRefreshStatus:input_type -> openshell.v1.GetProviderRefreshStatusRequest + 135, // 361: openshell.v1.OpenShell.ConfigureProviderRefresh:input_type -> openshell.v1.ConfigureProviderRefreshRequest + 137, // 362: openshell.v1.OpenShell.RotateProviderCredential:input_type -> openshell.v1.RotateProviderCredentialRequest + 139, // 363: openshell.v1.OpenShell.DeleteProviderRefresh:input_type -> openshell.v1.DeleteProviderRefreshRequest + 117, // 364: openshell.v1.OpenShell.DeleteProvider:input_type -> openshell.v1.DeleteProviderRequest + 152, // 365: openshell.v1.OpenShell.DeleteProviderProfile:input_type -> openshell.v1.DeleteProviderProfileRequest + 290, // 366: openshell.v1.OpenShell.GetSandboxConfig:input_type -> openshell.sandbox.v1.GetSandboxConfigRequest + 291, // 367: openshell.v1.OpenShell.GetGatewayConfig:input_type -> openshell.sandbox.v1.GetGatewayConfigRequest + 160, // 368: openshell.v1.OpenShell.UpdateConfig:input_type -> openshell.v1.UpdateConfigRequest + 170, // 369: openshell.v1.OpenShell.GetSandboxPolicyStatus:input_type -> openshell.v1.GetSandboxPolicyStatusRequest + 172, // 370: openshell.v1.OpenShell.ListSandboxPolicies:input_type -> openshell.v1.ListSandboxPoliciesRequest + 174, // 371: openshell.v1.OpenShell.ReportPolicyStatus:input_type -> openshell.v1.ReportPolicyStatusRequest + 247, // 372: openshell.v1.OpenShell.ReportEndpointStatus:input_type -> openshell.v1.ReportEndpointStatusRequest + 84, // 373: openshell.v1.OpenShell.ReportProviderReadiness:input_type -> openshell.v1.ReportProviderReadinessRequest + 177, // 374: openshell.v1.OpenShell.ReportSandboxConfiguration:input_type -> openshell.v1.ReportSandboxConfigurationRequest + 154, // 375: openshell.v1.OpenShell.GetSandboxProviderEnvironment:input_type -> openshell.v1.GetSandboxProviderEnvironmentRequest + 158, // 376: openshell.v1.OpenShell.ExchangeProviderSubjectToken:input_type -> openshell.v1.ExchangeProviderSubjectTokenRequest + 180, // 377: openshell.v1.OpenShell.GetSandboxLogs:input_type -> openshell.v1.GetSandboxLogsRequest + 181, // 378: openshell.v1.OpenShell.PushSandboxLogs:input_type -> openshell.v1.PushSandboxLogsRequest + 184, // 379: openshell.v1.OpenShell.ConnectSupervisor:input_type -> openshell.v1.SupervisorMessage + 191, // 380: openshell.v1.OpenShell.ReportMainProcessExit:input_type -> openshell.v1.ReportMainProcessExitRequest + 193, // 381: openshell.v1.OpenShell.FinalizeMainProcessExit:input_type -> openshell.v1.FinalizeMainProcessExitRequest + 199, // 382: openshell.v1.OpenShell.RelayStream:input_type -> openshell.v1.RelayFrame + 201, // 383: openshell.v1.OpenShell.PeerRelay:input_type -> openshell.v1.PeerRelayFrame + 84, // 384: openshell.v1.OpenShell.PeerReportProviderReadiness:input_type -> openshell.v1.ReportProviderReadinessRequest + 247, // 385: openshell.v1.OpenShell.PeerReportEndpointStatus:input_type -> openshell.v1.ReportEndpointStatusRequest + 82, // 386: openshell.v1.OpenShell.PeerGetSandboxProviderStatus:input_type -> openshell.v1.GetSandboxProviderStatusRequest + 109, // 387: openshell.v1.OpenShell.WatchSandbox:input_type -> openshell.v1.WatchSandboxRequest + 210, // 388: openshell.v1.OpenShell.SubmitPolicyAnalysis:input_type -> openshell.v1.SubmitPolicyAnalysisRequest + 212, // 389: openshell.v1.OpenShell.GetDraftPolicy:input_type -> openshell.v1.GetDraftPolicyRequest + 214, // 390: openshell.v1.OpenShell.ApproveDraftChunk:input_type -> openshell.v1.ApproveDraftChunkRequest + 216, // 391: openshell.v1.OpenShell.RejectDraftChunk:input_type -> openshell.v1.RejectDraftChunkRequest + 219, // 392: openshell.v1.OpenShell.ApproveAllDraftChunks:input_type -> openshell.v1.ApproveAllDraftChunksRequest + 221, // 393: openshell.v1.OpenShell.EditDraftChunk:input_type -> openshell.v1.EditDraftChunkRequest + 223, // 394: openshell.v1.OpenShell.UndoDraftChunk:input_type -> openshell.v1.UndoDraftChunkRequest + 225, // 395: openshell.v1.OpenShell.ClearDraftChunks:input_type -> openshell.v1.ClearDraftChunksRequest + 227, // 396: openshell.v1.OpenShell.GetDraftHistory:input_type -> openshell.v1.GetDraftHistoryRequest + 20, // 397: openshell.v1.OpenShell.IssueSandboxToken:input_type -> openshell.v1.IssueSandboxTokenRequest + 22, // 398: openshell.v1.OpenShell.RefreshSandboxToken:input_type -> openshell.v1.RefreshSandboxTokenRequest + 230, // 399: openshell.v1.OpenShell.CreateWorkspace:input_type -> openshell.v1.CreateWorkspaceRequest + 232, // 400: openshell.v1.OpenShell.GetWorkspace:input_type -> openshell.v1.GetWorkspaceRequest + 234, // 401: openshell.v1.OpenShell.ListWorkspaces:input_type -> openshell.v1.ListWorkspacesRequest + 236, // 402: openshell.v1.OpenShell.DeleteWorkspace:input_type -> openshell.v1.DeleteWorkspaceRequest + 239, // 403: openshell.v1.OpenShell.AddWorkspaceMember:input_type -> openshell.v1.AddWorkspaceMemberRequest + 241, // 404: openshell.v1.OpenShell.RemoveWorkspaceMember:input_type -> openshell.v1.RemoveWorkspaceMemberRequest + 243, // 405: openshell.v1.OpenShell.ListWorkspaceMembers:input_type -> openshell.v1.ListWorkspaceMembersRequest + 25, // 406: openshell.v1.OpenShell.Health:output_type -> openshell.v1.HealthResponse + 27, // 407: openshell.v1.OpenShell.GetCurrentUser:output_type -> openshell.v1.GetCurrentUserResponse + 29, // 408: openshell.v1.OpenShell.GetGatewayInfo:output_type -> openshell.v1.GetGatewayInfoResponse + 70, // 409: openshell.v1.OpenShell.CreateSandbox:output_type -> openshell.v1.SandboxResponse + 61, // 410: openshell.v1.OpenShell.BeginRootfsTarStaging:output_type -> openshell.v1.BeginRootfsTarStagingResponse + 70, // 411: openshell.v1.OpenShell.GetSandbox:output_type -> openshell.v1.SandboxResponse + 71, // 412: openshell.v1.OpenShell.ListSandboxes:output_type -> openshell.v1.ListSandboxesResponse + 57, // 413: openshell.v1.OpenShell.CreateSandboxTemplate:output_type -> openshell.v1.SandboxTemplateResponse + 57, // 414: openshell.v1.OpenShell.GetSandboxTemplate:output_type -> openshell.v1.SandboxTemplateResponse + 58, // 415: openshell.v1.OpenShell.ListSandboxTemplates:output_type -> openshell.v1.ListSandboxTemplatesResponse + 59, // 416: openshell.v1.OpenShell.DeleteSandboxTemplate:output_type -> openshell.v1.DeleteSandboxTemplateResponse + 72, // 417: openshell.v1.OpenShell.ListSandboxProviders:output_type -> openshell.v1.ListSandboxProvidersResponse + 73, // 418: openshell.v1.OpenShell.AttachSandboxProvider:output_type -> openshell.v1.AttachSandboxProviderResponse + 74, // 419: openshell.v1.OpenShell.DetachSandboxProvider:output_type -> openshell.v1.DetachSandboxProviderResponse + 83, // 420: openshell.v1.OpenShell.GetSandboxProviderStatus:output_type -> openshell.v1.GetSandboxProviderStatusResponse + 86, // 421: openshell.v1.OpenShell.DeleteSandbox:output_type -> openshell.v1.DeleteSandboxResponse + 70, // 422: openshell.v1.OpenShell.StopSandbox:output_type -> openshell.v1.SandboxResponse + 70, // 423: openshell.v1.OpenShell.StartSandbox:output_type -> openshell.v1.SandboxResponse + 88, // 424: openshell.v1.OpenShell.CreateSshSession:output_type -> openshell.v1.CreateSshSessionResponse + 96, // 425: openshell.v1.OpenShell.ExposeService:output_type -> openshell.v1.ServiceEndpointResponse + 96, // 426: openshell.v1.OpenShell.GetService:output_type -> openshell.v1.ServiceEndpointResponse + 92, // 427: openshell.v1.OpenShell.ListServices:output_type -> openshell.v1.ListServicesResponse + 94, // 428: openshell.v1.OpenShell.DeleteService:output_type -> openshell.v1.DeleteServiceResponse + 98, // 429: openshell.v1.OpenShell.RevokeSshSession:output_type -> openshell.v1.RevokeSshSessionResponse + 103, // 430: openshell.v1.OpenShell.ExecSandbox:output_type -> openshell.v1.ExecSandboxEvent + 105, // 431: openshell.v1.OpenShell.ForwardTcp:output_type -> openshell.v1.TcpForwardFrame + 103, // 432: openshell.v1.OpenShell.ExecSandboxInteractive:output_type -> openshell.v1.ExecSandboxEvent + 118, // 433: openshell.v1.OpenShell.CreateProvider:output_type -> openshell.v1.ProviderResponse + 118, // 434: openshell.v1.OpenShell.GetProvider:output_type -> openshell.v1.ProviderResponse + 119, // 435: openshell.v1.OpenShell.ListProviders:output_type -> openshell.v1.ListProvidersResponse + 144, // 436: openshell.v1.OpenShell.ListProviderProfiles:output_type -> openshell.v1.ListProviderProfilesResponse + 143, // 437: openshell.v1.OpenShell.GetProviderProfile:output_type -> openshell.v1.ProviderProfileResponse + 146, // 438: openshell.v1.OpenShell.ImportProviderProfiles:output_type -> openshell.v1.ImportProviderProfilesResponse + 148, // 439: openshell.v1.OpenShell.UpdateProviderProfiles:output_type -> openshell.v1.UpdateProviderProfilesResponse + 150, // 440: openshell.v1.OpenShell.LintProviderProfiles:output_type -> openshell.v1.LintProviderProfilesResponse + 118, // 441: openshell.v1.OpenShell.UpdateProvider:output_type -> openshell.v1.ProviderResponse + 134, // 442: openshell.v1.OpenShell.GetProviderRefreshStatus:output_type -> openshell.v1.GetProviderRefreshStatusResponse + 136, // 443: openshell.v1.OpenShell.ConfigureProviderRefresh:output_type -> openshell.v1.ConfigureProviderRefreshResponse + 138, // 444: openshell.v1.OpenShell.RotateProviderCredential:output_type -> openshell.v1.RotateProviderCredentialResponse + 140, // 445: openshell.v1.OpenShell.DeleteProviderRefresh:output_type -> openshell.v1.DeleteProviderRefreshResponse + 151, // 446: openshell.v1.OpenShell.DeleteProvider:output_type -> openshell.v1.DeleteProviderResponse + 153, // 447: openshell.v1.OpenShell.DeleteProviderProfile:output_type -> openshell.v1.DeleteProviderProfileResponse + 292, // 448: openshell.v1.OpenShell.GetSandboxConfig:output_type -> openshell.sandbox.v1.GetSandboxConfigResponse + 293, // 449: openshell.v1.OpenShell.GetGatewayConfig:output_type -> openshell.sandbox.v1.GetGatewayConfigResponse + 169, // 450: openshell.v1.OpenShell.UpdateConfig:output_type -> openshell.v1.UpdateConfigResponse + 171, // 451: openshell.v1.OpenShell.GetSandboxPolicyStatus:output_type -> openshell.v1.GetSandboxPolicyStatusResponse + 173, // 452: openshell.v1.OpenShell.ListSandboxPolicies:output_type -> openshell.v1.ListSandboxPoliciesResponse + 175, // 453: openshell.v1.OpenShell.ReportPolicyStatus:output_type -> openshell.v1.ReportPolicyStatusResponse + 248, // 454: openshell.v1.OpenShell.ReportEndpointStatus:output_type -> openshell.v1.ReportEndpointStatusResponse + 85, // 455: openshell.v1.OpenShell.ReportProviderReadiness:output_type -> openshell.v1.ReportProviderReadinessResponse + 178, // 456: openshell.v1.OpenShell.ReportSandboxConfiguration:output_type -> openshell.v1.ReportSandboxConfigurationResponse + 157, // 457: openshell.v1.OpenShell.GetSandboxProviderEnvironment:output_type -> openshell.v1.GetSandboxProviderEnvironmentResponse + 159, // 458: openshell.v1.OpenShell.ExchangeProviderSubjectToken:output_type -> openshell.v1.ExchangeProviderSubjectTokenResponse + 183, // 459: openshell.v1.OpenShell.GetSandboxLogs:output_type -> openshell.v1.GetSandboxLogsResponse + 182, // 460: openshell.v1.OpenShell.PushSandboxLogs:output_type -> openshell.v1.PushSandboxLogsResponse + 185, // 461: openshell.v1.OpenShell.ConnectSupervisor:output_type -> openshell.v1.GatewayMessage + 192, // 462: openshell.v1.OpenShell.ReportMainProcessExit:output_type -> openshell.v1.ReportMainProcessExitResponse + 194, // 463: openshell.v1.OpenShell.FinalizeMainProcessExit:output_type -> openshell.v1.FinalizeMainProcessExitResponse + 199, // 464: openshell.v1.OpenShell.RelayStream:output_type -> openshell.v1.RelayFrame + 201, // 465: openshell.v1.OpenShell.PeerRelay:output_type -> openshell.v1.PeerRelayFrame + 85, // 466: openshell.v1.OpenShell.PeerReportProviderReadiness:output_type -> openshell.v1.ReportProviderReadinessResponse + 248, // 467: openshell.v1.OpenShell.PeerReportEndpointStatus:output_type -> openshell.v1.ReportEndpointStatusResponse + 83, // 468: openshell.v1.OpenShell.PeerGetSandboxProviderStatus:output_type -> openshell.v1.GetSandboxProviderStatusResponse + 110, // 469: openshell.v1.OpenShell.WatchSandbox:output_type -> openshell.v1.SandboxStreamEvent + 211, // 470: openshell.v1.OpenShell.SubmitPolicyAnalysis:output_type -> openshell.v1.SubmitPolicyAnalysisResponse + 213, // 471: openshell.v1.OpenShell.GetDraftPolicy:output_type -> openshell.v1.GetDraftPolicyResponse + 215, // 472: openshell.v1.OpenShell.ApproveDraftChunk:output_type -> openshell.v1.ApproveDraftChunkResponse + 217, // 473: openshell.v1.OpenShell.RejectDraftChunk:output_type -> openshell.v1.RejectDraftChunkResponse + 220, // 474: openshell.v1.OpenShell.ApproveAllDraftChunks:output_type -> openshell.v1.ApproveAllDraftChunksResponse + 222, // 475: openshell.v1.OpenShell.EditDraftChunk:output_type -> openshell.v1.EditDraftChunkResponse + 224, // 476: openshell.v1.OpenShell.UndoDraftChunk:output_type -> openshell.v1.UndoDraftChunkResponse + 226, // 477: openshell.v1.OpenShell.ClearDraftChunks:output_type -> openshell.v1.ClearDraftChunksResponse + 229, // 478: openshell.v1.OpenShell.GetDraftHistory:output_type -> openshell.v1.GetDraftHistoryResponse + 21, // 479: openshell.v1.OpenShell.IssueSandboxToken:output_type -> openshell.v1.IssueSandboxTokenResponse + 23, // 480: openshell.v1.OpenShell.RefreshSandboxToken:output_type -> openshell.v1.RefreshSandboxTokenResponse + 231, // 481: openshell.v1.OpenShell.CreateWorkspace:output_type -> openshell.v1.CreateWorkspaceResponse + 233, // 482: openshell.v1.OpenShell.GetWorkspace:output_type -> openshell.v1.GetWorkspaceResponse + 235, // 483: openshell.v1.OpenShell.ListWorkspaces:output_type -> openshell.v1.ListWorkspacesResponse + 237, // 484: openshell.v1.OpenShell.DeleteWorkspace:output_type -> openshell.v1.DeleteWorkspaceResponse + 240, // 485: openshell.v1.OpenShell.AddWorkspaceMember:output_type -> openshell.v1.AddWorkspaceMemberResponse + 242, // 486: openshell.v1.OpenShell.RemoveWorkspaceMember:output_type -> openshell.v1.RemoveWorkspaceMemberResponse + 244, // 487: openshell.v1.OpenShell.ListWorkspaceMembers:output_type -> openshell.v1.ListWorkspaceMembersResponse + 406, // [406:488] is the sub-list for method output_type + 324, // [324:406] is the sub-list for method input_type + 324, // [324:324] is the sub-list for extension type_name + 324, // [324:324] is the sub-list for extension extendee + 0, // [0:324] is the sub-list for field type_name } func init() { file_openshell_proto_init() } @@ -20027,7 +20329,7 @@ func file_openshell_proto_init() { (*SandboxStreamEvent_Warning)(nil), (*SandboxStreamEvent_DraftPolicyUpdate)(nil), } - file_openshell_proto_msgTypes[140].OneofWrappers = []any{ + file_openshell_proto_msgTypes[141].OneofWrappers = []any{ (*PolicyMergeOperation_AddRule)(nil), (*PolicyMergeOperation_RemoveEndpoint)(nil), (*PolicyMergeOperation_RemoveRule)(nil), @@ -20035,29 +20337,29 @@ func file_openshell_proto_init() { (*PolicyMergeOperation_AddAllowRules)(nil), (*PolicyMergeOperation_RemoveBinary)(nil), } - file_openshell_proto_msgTypes[144].OneofWrappers = []any{} - file_openshell_proto_msgTypes[163].OneofWrappers = []any{ + file_openshell_proto_msgTypes[145].OneofWrappers = []any{} + file_openshell_proto_msgTypes[164].OneofWrappers = []any{ (*SupervisorMessage_Hello)(nil), (*SupervisorMessage_Heartbeat)(nil), (*SupervisorMessage_RelayOpenResult)(nil), (*SupervisorMessage_RelayClose)(nil), } - file_openshell_proto_msgTypes[164].OneofWrappers = []any{ + file_openshell_proto_msgTypes[165].OneofWrappers = []any{ (*GatewayMessage_SessionAccepted)(nil), (*GatewayMessage_SessionRejected)(nil), (*GatewayMessage_Heartbeat)(nil), (*GatewayMessage_RelayOpen)(nil), (*GatewayMessage_RelayClose)(nil), } - file_openshell_proto_msgTypes[174].OneofWrappers = []any{ + file_openshell_proto_msgTypes[175].OneofWrappers = []any{ (*RelayOpen_Ssh)(nil), (*RelayOpen_Tcp)(nil), } - file_openshell_proto_msgTypes[178].OneofWrappers = []any{ + file_openshell_proto_msgTypes[179].OneofWrappers = []any{ (*RelayFrame_Init)(nil), (*RelayFrame_Data)(nil), } - file_openshell_proto_msgTypes[180].OneofWrappers = []any{ + file_openshell_proto_msgTypes[181].OneofWrappers = []any{ (*PeerRelayFrame_Init)(nil), (*PeerRelayFrame_Data)(nil), } @@ -20066,8 +20368,8 @@ func file_openshell_proto_init() { File: protoimpl.DescBuilder{ GoPackagePath: reflect.TypeOf(x{}).PkgPath(), RawDescriptor: unsafe.Slice(unsafe.StringData(file_openshell_proto_rawDesc), len(file_openshell_proto_rawDesc)), - NumEnums: 18, - NumMessages: 253, + NumEnums: 20, + NumMessages: 255, NumExtensions: 0, NumServices: 1, }, diff --git a/sdk/typescript/src/client.test.ts b/sdk/typescript/src/client.test.ts index 7fef9b0db9..178a73cc6e 100644 --- a/sdk/typescript/src/client.test.ts +++ b/sdk/typescript/src/client.test.ts @@ -20,9 +20,16 @@ import { SandboxClient, SandboxTemplateClient, SCOPE_NAMES, + ServiceAuthorizationMode, STATUS_NAMES, } from './client.js'; -import { OpenShell, SandboxPhase, ServiceStatus } from './gen/openshell_pb.js'; +import { + OpenShell, + ServiceAuthorizationMode as ProtoServiceAuthorizationMode, + SandboxPhase, + SandboxRestartPolicy, + ServiceStatus, +} from './gen/openshell_pb.js'; import { PolicySource, SettingScope } from './gen/sandbox_pb.js'; import type { ExecInteractiveSession, ExecInteractiveSessionControl } from './index.js'; @@ -276,7 +283,9 @@ describe('exec / execStream', () => { describe('create', () => { it('sends create-time service exposures', async () => { - let created: { serviceExposures?: Array<{ service?: string; targetPort?: number }> } = {}; + let created: { + serviceExposures?: Array<{ service?: string; targetPort?: number; authorizationMode?: number }>; + } = {}; const sandbox = client({ createSandbox: (req) => { created = req; @@ -292,12 +301,29 @@ describe('create', () => { const result = await sandbox.create({ image: 'img', - serviceExposures: [{ targetPort: 4500 }, { service: 'metrics', targetPort: 9090 }], + serviceExposures: [ + { targetPort: 4500 }, + { + service: 'metrics', + targetPort: 9090, + authorizationMode: ServiceAuthorizationMode.BearerPassthrough, + }, + ], }); - expect(created.serviceExposures?.map(({ service, targetPort }) => ({ service, targetPort }))).toEqual([ - { service: '', targetPort: 4500 }, - { service: 'metrics', targetPort: 9090 }, + expect( + created.serviceExposures?.map(({ service, targetPort, authorizationMode }) => ({ + service, + targetPort, + authorizationMode, + })), + ).toEqual([ + { service: '', targetPort: 4500, authorizationMode: ProtoServiceAuthorizationMode.STRIP }, + { + service: 'metrics', + targetPort: 9090, + authorizationMode: ProtoServiceAuthorizationMode.BEARER_PASSTHROUGH, + }, ]); expect(result.serviceUrls).toEqual({ '': 'https://sb.example.test/', @@ -336,6 +362,20 @@ describe('create', () => { expect(created.spec?.tty).toBe(true); }); + it('sends the restart policy', async () => { + let created: { spec?: { restartPolicy?: SandboxRestartPolicy } } = {}; + const sandbox = client({ + createSandbox: (req) => { + created = req; + return readySandbox('sb', 'sb-id'); + }, + }); + + await sandbox.create({ image: 'img', restartPolicy: 'on-failure' }); + + expect(created.spec?.restartPolicy).toBe(SandboxRestartPolicy.ON_FAILURE); + }); + it('rawSpec reaches an ungated field and overrides a curated one', async () => { let created: { spec?: { @@ -842,6 +882,29 @@ describe('sandbox templates', () => { await expect(templates.delete(' ')).rejects.toMatchObject({ code: 'invalid_config' }); await expect(templates.get('missing-response')).rejects.toMatchObject({ code: 'invalid_config' }); }); + + it('maps restart controller status', async () => { + const sandbox = client({ + getSandbox: () => ({ + sandbox: { + metadata: { id: 'sb-id', name: 'sb', resourceVersion: 8n }, + status: { + phase: SandboxPhase.STARTING, + restartCount: 3, + nextRestartTime: { seconds: 1_700_000_000n, nanos: 0 }, + mainProcessStartedTime: { seconds: 1_699_999_000n, nanos: 0 }, + }, + }, + }), + }); + + await expect(sandbox.get('sb')).resolves.toMatchObject({ + phase: 'starting', + restartCount: 3, + nextRestartAtMs: 1_700_000_000_000, + mainProcessStartedAtMs: 1_699_999_000_000, + }); + }); }); describe('waits', () => { diff --git a/sdk/typescript/src/client.ts b/sdk/typescript/src/client.ts index 93ad8ba6d1..e73bccd314 100644 --- a/sdk/typescript/src/client.ts +++ b/sdk/typescript/src/client.ts @@ -22,7 +22,9 @@ import type { Sandbox, SandboxWorkloadTemplate, UpdateConfigResponse } from './g import { type ExecSandboxInputSchema, OpenShell, + ServiceAuthorizationMode as ProtoServiceAuthorizationMode, SandboxPhase, + SandboxRestartPolicy, type SandboxSpecSchema, type SandboxWorkloadTemplateSchema, ServiceStatus, @@ -84,6 +86,9 @@ export type SandboxPhaseName = | 'starting' | 'completed'; +/** Restart behavior after the canonical main process exits. */ +export type SandboxRestartPolicyName = 'never' | 'on-failure' | 'always'; + /** Lowercase mirror of the generated `ServiceStatus` enum. Hand-maintained. */ export type HealthStatus = 'unspecified' | 'healthy' | 'degraded' | 'unhealthy'; @@ -141,6 +146,8 @@ export interface SandboxSpec { tty?: boolean; /** Loopback HTTP services to expose when the sandbox is created. */ serviceExposures?: ServiceExposure[]; + /** Restart behavior after the canonical main process exits. */ + restartPolicy?: SandboxRestartPolicyName; /** * Create-time sandbox policy (the safety boundary). Sandbox-scoped * `setPolicy` cannot introduce static fields later, so express filesystem, @@ -162,6 +169,23 @@ export interface ServiceExposure { service?: string; /** Loopback TCP port inside the sandbox. */ targetPort: number; + /** Handling for an incoming application Authorization header. */ + authorizationMode?: ServiceAuthorizationMode; +} + +export enum ServiceAuthorizationMode { + Strip = 'strip', + BearerPassthrough = 'bearer_passthrough', +} + +function serviceAuthorizationModeToProto(mode: ServiceAuthorizationMode | undefined): ProtoServiceAuthorizationMode { + switch (mode) { + case ServiceAuthorizationMode.BearerPassthrough: + return ProtoServiceAuthorizationMode.BEARER_PASSTHROUGH; + case ServiceAuthorizationMode.Strip: + case undefined: + return ProtoServiceAuthorizationMode.STRIP; + } } export interface SandboxFromTemplateSpec { @@ -197,6 +221,9 @@ export interface SandboxRef { createdFromWorkloadTemplate?: SandboxWorkloadTemplateProvenance; /** Service URLs returned by creation, keyed by service name. */ serviceUrls: Record; + restartCount: number; + nextRestartAtMs?: number; + mainProcessStartedAtMs?: number; } export interface SandboxWorkloadTemplateProvenance { @@ -472,6 +499,17 @@ export const POLICY_SOURCE_NAMES: Record = { function phaseName(p: SandboxPhase): SandboxPhaseName { return PHASE_NAMES[p] ?? 'unspecified'; } + +function restartPolicyValue(policy: SandboxRestartPolicyName | undefined): SandboxRestartPolicy { + switch (policy) { + case 'on-failure': + return SandboxRestartPolicy.ON_FAILURE; + case 'always': + return SandboxRestartPolicy.ALWAYS; + default: + return SandboxRestartPolicy.NEVER; + } +} function statusName(s: ServiceStatus): HealthStatus { return STATUS_NAMES[s] ?? 'unspecified'; } @@ -488,6 +526,8 @@ function sandboxRef(sandbox: Sandbox | undefined, serviceUrls: Record ({ service: exposure.service ?? '', targetPort: exposure.targetPort, + authorizationMode: serviceAuthorizationModeToProto(exposure.authorizationMode), })) ?? [], }); return sandboxRef(resp.sandbox, resp.serviceUrls); @@ -1023,6 +1068,7 @@ export class SandboxClient { spec.serviceExposures?.map((exposure) => ({ service: exposure.service ?? '', targetPort: exposure.targetPort, + authorizationMode: serviceAuthorizationModeToProto(exposure.authorizationMode), })) ?? [], }); return sandboxRef(resp.sandbox, resp.serviceUrls); diff --git a/sdk/typescript/src/index.ts b/sdk/typescript/src/index.ts index e0802fc8f7..9229f8bcbe 100644 --- a/sdk/typescript/src/index.ts +++ b/sdk/typescript/src/index.ts @@ -33,6 +33,7 @@ export type { SandboxPolicy, SandboxRef, SandboxResources, + SandboxRestartPolicyName, SandboxServiceLevel, SandboxSpec, SandboxStartup, @@ -52,7 +53,14 @@ export type { WaitOptions, WorkspaceListScope, } from './client.js'; -export { errorCode, OpenShellClient, Pager, SandboxClient, SandboxTemplateClient } from './client.js'; +export { + errorCode, + OpenShellClient, + Pager, + SandboxClient, + SandboxTemplateClient, + ServiceAuthorizationMode, +} from './client.js'; export type { ErrorInfo, FieldViolation, SdkErrorCode } from './errors.js'; export { fromConnect, SdkError } from './errors.js'; export type { ClientCredentialsOptions, OidcTokenProvider } from './oidc.js'; diff --git a/skills/debug-openshell-cluster/SKILL.md b/skills/debug-openshell-cluster/SKILL.md index 85360a77ad..6382428946 100644 --- a/skills/debug-openshell-cluster/SKILL.md +++ b/skills/debug-openshell-cluster/SKILL.md @@ -296,6 +296,34 @@ Common findings: - The sandbox fails its enforcement probe: inspect the sandbox log for the exact nested seccomp user-notification, task-memory, Landlock, loopback DNS, or socket-injection check that failed. A runtime may return `ENOSYS` for `process_vm_readv` and `process_vm_writev` while satisfying the production parent-to-workload-child task-memory probe through `/proc//mem`; only failure of both backends is fatal. Do not add capabilities or switch to an unconfined seccomp profile; use a runtime whose default profile permits the unprivileged probe. - A GPU sandbox fails because Docker reports no discovered NVIDIA CDI devices: verify `.DiscoveredDevices` contains entries such as `nvidia.com/gpu=all`, verify `/etc/cdi` or `/var/run/cdi` contains a generated NVIDIA spec, and check that `nvidia-cdi-refresh.service` and `nvidia-cdi-refresh.path` from NVIDIA Container Toolkit are enabled and healthy. The service is a one-shot unit, so `inactive (dead)` can be normal after a successful run; use `systemctl status` and `journalctl` to distinguish success from a skipped or failed refresh. Restart `nvidia-cdi-refresh.service` to regenerate missing or stale CDI specs, then restart or reload Docker and re-check `docker info`. +#### Corporate upstream proxy + +Docker corporate proxy settings are operator-owned fields under +`[openshell.drivers.docker]`. Confirm the complete proxy table and inspect the +companion supervisor command and logs: + +```bash +grep -A20 '^\[openshell.drivers.docker\]' | grep -E 'https_proxy|no_proxy|proxy_auth_file|proxy_auth_allow_insecure|proxy_connect_by_hostname|proxy_ca_bundle' +docker ps --filter label=openshell.ai/isolation-role=supervisor +docker inspect --format '{{json .Config.Cmd}} {{json .Mounts}}' +docker logs --tail=200 | grep -Ei 'upstream|connect|proxy|certificate' +``` + +`proxy_ca_bundle` names a gateway-host PEM file and requires `https_proxy`. +The proxy URL may use `http://` or `https://`; a plain HTTP proxy may still +re-sign destination TLS. Missing, unreadable, empty, oversized, malformed, or +certificate-free bundles fail closed. The Docker driver copies a validated +bundle into its supervisor-only named volume and passes the fixed path +`/.openshell/supervisor/upstream-proxy-ca-bundle.pem`. The gateway-host path +must not appear in container arguments, mounts, workload environment, or +`template.driver_config.docker`. + +An HTTPS proxy certificate error usually means the bundle lacks the proxy +listener issuer or its certificate does not match the proxy hostname. A +TLS-intercepted destination error means the re-signing issuer is missing or the +supervisor did not receive the bundle. Keep verification enabled and correct +the operator bundle. + During a graceful gateway restart, Docker, Podman, and VM sandboxes with running intent should stop before the gateway exits and restart after it returns. Check for `Stopped sandbox during gateway shutdown` and `Started @@ -311,6 +339,14 @@ the associated persistence errors: a stopped supervisor's owner record may remain until its lease expires and temporarily block reconnection. Successful compute stop alone does not confirm that session cleanup finished. +For a sandbox with `--restart-policy on-failure` or `always`, inspect +`openshell sandbox get ` for the policy, exit code, restart count, and +next restart time. `Starting` can mean the gateway is waiting for backoff or +for a replacement supervisor. Check gateway logs for stop or start errors if +the deadline passes without a transition to `Ready`. Docker, Podman, +Kubernetes, and VM native restart policies remain disabled; the gateway owns +the replacement. + ### Step 5: Check Podman-Backed Gateways ```bash @@ -904,6 +940,7 @@ credential failures. | Symptom | Likely cause | Check | |---|---|---| | `openshell status` fails | Gateway endpoint unreachable or auth mismatch | `openshell gateway info`, gateway logs | +| Gateway OCSF JSONL stops growing, has gaps, or is rejected by a SIEM | File errors, queue pressure, shipper falling behind retention, or a schema mismatch | Inspect gateway warnings and `openshell_ocsf_log_*` metrics; check `[openshell.gateway.ocsf_log]`, `schema_version`, directory permissions, free space, per-replica paths, and shipper rotation checkpoints. See the published [gateway configuration reference](https://docs.nvidia.com/openshell/latest/how-it-works/gateways/configuration). | | `BatchSpanProcessor.ExportError` repeatedly reports connection refused on `127.0.0.1:4317` | The local gateway started with OTLP configured but the collector forwarding task later stopped, or the config was created manually | Restart `gateway:docker`, `gateway:podman`, or `gateway:vm` so it re-detects the listener; inspect the generated `gateway.toml` for `[openshell.gateway.otlp]` | | Gateway starts but sandbox create fails | Compute driver cannot reach runtime | Docker/Podman/Kubernetes/VM driver logs | | Docker or Podman sandbox never registers | Wrong gateway endpoint, unavailable host networking, or supervisor startup failure | Gateway logs and supervisor container logs | diff --git a/skills/openshell-cli/SKILL.md b/skills/openshell-cli/SKILL.md index a2c548c251..3977018a01 100644 --- a/skills/openshell-cli/SKILL.md +++ b/skills/openshell-cli/SKILL.md @@ -108,7 +108,7 @@ The agent will be prompted interactively if credentials are missing. ### Step 4: Exit and clean up -Exit the sandbox shell (`exit` or Ctrl-D), then: +Exit the sandbox shell with `exit`, or detach with Ctrl-D, then: ```bash openshell sandbox delete @@ -277,6 +277,7 @@ Key flags: - `--label KEY=VALUE`: Add labels for later selection (repeatable) - `--env KEY=VALUE`: Set non-secret sandbox environment variables (repeatable); use `--provider` for credentials - `--tty`: Allocate a retained PTY for the canonical main process +- `--restart-policy never|on-failure|always`: Select gateway-owned main-process restart behavior; `never` is the default - `--approval-mode manual|auto`: Control handling of agent-authored policy proposals; `manual` is the default - `--upload [:]`: Upload local files into the container working directory or an explicit destination - `--no-git-ignore`: Disable `.gitignore` filtering for uploads @@ -364,22 +365,28 @@ that process running; reconnecting targets the same process instance and replays recent output. If an established SSH transport is interrupted, such as when a laptop sleeps and wakes, the CLI retries transient failures for up to 60 seconds and reattaches to that same process. Use `sandbox exec --tty -- /bin/bash -l` -for a new shell. Press `Ctrl-P`, then `Ctrl-Q` to disconnect without terminating -main. OpenSSH's `~.` escape looks like transport loss and therefore starts -automatic recovery; after it reattaches, use `Ctrl-P`, then `Ctrl-Q` to exit, or -press `Ctrl-C` between retry attempts to cancel recovery. When you own stdin, -`Ctrl-C` interrupts the foreground process. In a read-only attachment, `Ctrl-C` -exits the viewer and leaves main and other attachments running. Configure VS -Code Remote-SSH with: +for a new shell. Press `Ctrl-D` or `Ctrl-P`, then `Ctrl-Q` to disconnect without +terminating main. OpenSSH's `~.` escape looks like transport loss and therefore +starts automatic recovery; after it reattaches, use `Ctrl-D` or `Ctrl-P`, then +`Ctrl-Q` to exit, or press `Ctrl-C` between retry attempts to cancel recovery. +When you own stdin, `Ctrl-C` interrupts the foreground process. In a read-only +attachment, `Ctrl-C` or `Ctrl-D` exits the viewer and leaves main and other +attachments running. Configure VS Code Remote-SSH with: ```bash openshell sandbox ssh-config my-sandbox >> ~/.ssh/config ``` +A writable attachment that finds another attachment holding stdin reports `attached read-only; retry input after the owner disconnects`. Automatic recovery can hit this when it reattaches before the supervisor closes the dead connection. The supervisor closes a connection 60 seconds after it last received bytes from it, which can be later than 60 seconds after the network failed if the relay buffered data. After the old owner disconnects or times out, send the input you meant to type next. If stdin is free, the attachment prints `input enabled` and forwards that input to the process, so do not probe with Enter or a prompt answer such as `y`. If nothing prints, the old connection still holds stdin. Input sent while the attachment was read-only never reaches the process, so send it again later. `Ctrl-C`, `Ctrl-D`, and `Ctrl-P` then `Ctrl-Q` still exit a read-only attachment instead of enabling input; enable input first if you need `Ctrl-C` to interrupt the process. Recovery never takes stdin from a healthy owner, and an explicitly read-only attachment stays read-only. + If `connect` reports `canonical main process already finished`, inspect the -result with `sandbox get`. A pending -foreground attachment can still retrieve retained output in `Completed` or -`Error`; phase alone does not determine whether attachment is available. +result with `sandbox get`. A pending foreground attachment can still retrieve +retained output in `Completed` or `Error`; phase alone does not determine +whether attachment is available. A nonzero main-process exit under +`on-failure`, or any exit under `always`, +moves the sandbox to `Starting` during backoff and resource replacement. Connect +and exec commands resume after the new supervisor session makes it `Ready`. +An explicit `sandbox stop` cancels a pending restart. ### Upload and download files @@ -404,10 +411,14 @@ within it. ### Execute a non-interactive command ```bash -openshell sandbox exec --name my-sandbox --workdir /workspace -- ls -la +openshell sandbox exec my-sandbox --workdir /workspace -- ls -la openshell sandbox exec --name my-sandbox --env MODE=test -- cargo test ``` +The sandbox is a positional name or `--name`, not both; omit it to use the +last-used sandbox. `--` is required and everything after it is the remote +command, so put options such as `--tty` before it. + `sandbox exec` starts an independent sibling process and streams output. After stdout and stderr drain, it returns the remote command's exit code if delivery succeeds. Output delivery failure instead returns exit code 74, even when the @@ -689,6 +700,9 @@ openshell forward start 8080 my-app -d ``` The service is now reachable at `localhost:8080`. +CLI forwards ignore SSH multiplexing and automatic backgrounding settings in the +user's SSH config. Only background forwards are tracked by `forward list` and +managed by `forward stop`; foreground forwards end when the command exits. Manage or iterate on the sandbox: @@ -872,11 +886,13 @@ openshell sandbox create \ --name my-app \ --from my-app:latest \ --expose 8080 \ + --expose-authorization-mode bearer-passthrough \ --detach \ -- ./start-server.sh # Expose and manage an HTTP service through the gateway. -openshell service expose my-app 8080 web +openshell service expose my-app 8080 web \ + --authorization-mode bearer-passthrough openshell service list my-app openshell service list my-app --output json openshell service get my-app web @@ -891,6 +907,19 @@ request and keeps the sandbox running. Add `--output json` for automation; the result contains a `service_urls` map whose empty key is the unnamed endpoint. Use `openshell service expose` after creation to add or update named endpoints. +Exposed services strip `Authorization` by default. Select +`bearer-passthrough` only when the application inside the sandbox authenticates +its own clients. This mode accepts either no `Authorization` header or exactly +one non-empty Bearer credential and forwards that value unchanged. It rejects +duplicate, malformed, or non-Bearer authorization before contacting the +application. The application remains responsible for validating the token, and +the raw token reaches the sandbox process, so never log it. Service routes +bypass control-plane RPC authorization, but they still use the gateway's +existing listener, domain routing, and TLS configuration, including any client +certificate requirement. See the published +[sandbox service documentation](https://docs.nvidia.com/openshell/latest/how-it-works/sandboxes/overview.md) +for the complete security contract. + Prefer loopback binds unless the user explicitly needs LAN-visible local access. --- diff --git a/tasks/scripts/gateway-docker.sh b/tasks/scripts/gateway-docker.sh index b85c9a6638..9974cb8a0d 100644 --- a/tasks/scripts/gateway-docker.sh +++ b/tasks/scripts/gateway-docker.sh @@ -214,6 +214,9 @@ fi if [[ -n "${OPENSHELL_SANDBOX_PROXY_CONNECT_BY_HOSTNAME+x}" ]]; then printf 'proxy_connect_by_hostname = %s\n' "${OPENSHELL_SANDBOX_PROXY_CONNECT_BY_HOSTNAME}" >>"${CONFIG_PATH}" fi +if [[ -n "${OPENSHELL_SANDBOX_PROXY_CA_BUNDLE+x}" ]]; then + printf 'proxy_ca_bundle = "%s"\n' "$(toml_escape "${OPENSHELL_SANDBOX_PROXY_CA_BUNDLE}")" >>"${CONFIG_PATH}" +fi if [[ -n "${OPENSHELL_PROVIDER_SPIFFE_WORKLOAD_API_SOCKET+x}" ]]; then printf 'provider_spiffe_workload_api_socket = "%s"\n' "$(toml_escape "${OPENSHELL_PROVIDER_SPIFFE_WORKLOAD_API_SOCKET}")" >>"${CONFIG_PATH}" fi diff --git a/tasks/scripts/test-install-sh.sh b/tasks/scripts/test-install-sh.sh index 286fa3f1ff..581251891c 100755 --- a/tasks/scripts/test-install-sh.sh +++ b/tasks/scripts/test-install-sh.sh @@ -112,7 +112,7 @@ assert_linux_package_method() { local expected=$7 local actual - actual="$( + ( export OPENSHELL_INSTALL_METHOD="$install_method" export OPENSHELL_VERSION="$requested_version" has_cmd() { @@ -125,7 +125,8 @@ assert_linux_package_method() { } snap() { [ "$*" = "list openshell" ] && [ "$snap_state" = "installed" ]; } linux_package_method - )" + ) >"$out" + actual="$(cat "$out")" if [ "$actual" != "$expected" ]; then echo "FAIL: ${name}: expected ${expected}, got ${actual}" >&2 exit 1 @@ -270,7 +271,7 @@ assert_snap_install_flow() { local expected=$5 local calls - calls="$( + ( has_cmd() { case "$1" in snap) return 0 ;; @@ -295,7 +296,8 @@ assert_snap_install_flow() { export TARGET_USER=test-user export OPENSHELL_VERSION="$requested_version" install_linux_snap - )" + ) >"$out" + calls="$(cat "$out")" if [ "$calls" != "$expected" ]; then echo "FAIL: ${name}: unexpected command sequence" >&2 echo "Expected:" >&2 diff --git a/tests/ansible/playbooks/k3s.yaml b/tests/ansible/playbooks/k3s.yaml new file mode 100644 index 0000000000..e1a2aa3397 --- /dev/null +++ b/tests/ansible/playbooks/k3s.yaml @@ -0,0 +1,67 @@ +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +--- +- name: Install K3s and Helm + hosts: all + become: true + gather_facts: false + vars: + helm_version: 4.2.0 + k3s_version: 1.36.3+k3s1 + tasks: + - name: Wait for SSH + ansible.builtin.wait_for_connection: + + - name: Install download prerequisites + ansible.builtin.apt: + name: + - ca-certificates + - curl + state: present + update_cache: true + + - name: Download K3s installer + ansible.builtin.get_url: + url: https://get.k3s.io + checksum: sha256:e5cc3b3d9dfc1662c2d9be6da5abc9a4cd317d6abc3a5ffc02e3dd3248207fee + dest: /tmp/install-k3s.sh + mode: "0755" + + - name: Install K3s + ansible.builtin.command: + argv: + - /tmp/install-k3s.sh + - server + - --disable=servicelb + - --disable=traefik + creates: /usr/local/bin/k3s + environment: + INSTALL_K3S_VERSION: "v{{ k3s_version }}" + + - name: Wait for K3s + ansible.builtin.command: + argv: + - /usr/local/bin/k3s + - kubectl + - wait + - --for=condition=Ready + - node + - --all + - --timeout=120s + changed_when: false + + - name: Download Helm installer + ansible.builtin.get_url: + url: https://raw.githubusercontent.com/helm/helm/06468084e85c244c712834933d25ea232a4c2093/scripts/get-helm-4 + checksum: sha256:b68c5f694cff19f14ee8a5784ffd3de27fa7034ec8f973d703fc6fb85496ced7 + dest: /tmp/install-helm.sh + mode: "0755" + + - name: Install Helm + ansible.builtin.command: + argv: + - /tmp/install-helm.sh + creates: /usr/local/bin/helm + environment: + DESIRED_VERSION: "v{{ helm_version }}" diff --git a/tests/ansible/playbooks/openshell-k3s.yaml b/tests/ansible/playbooks/openshell-k3s.yaml new file mode 100644 index 0000000000..dd00aa4bda --- /dev/null +++ b/tests/ansible/playbooks/openshell-k3s.yaml @@ -0,0 +1,175 @@ +# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 + +--- +- name: Install OpenShell on K3s + hosts: all + become: true + gather_facts: false + tasks: + - name: Wait for SSH + ansible.builtin.wait_for_connection: + + - name: Install OpenShell CLI + ansible.builtin.copy: + src: "{{ openshell_cli_binary }}" + dest: /usr/local/bin/openshell + mode: "0755" + + - name: Create OpenShell artifact directory + ansible.builtin.file: + path: /var/lib/openshell/artifacts + state: directory + mode: "0700" + + - name: Copy OpenShell images + ansible.builtin.copy: + src: "{{ item.src }}" + dest: "/var/lib/openshell/artifacts/{{ item.name }}.tar" + mode: "0600" + loop: + - name: gateway + src: "{{ openshell_gateway_image }}" + - name: sandbox + src: "{{ openshell_sandbox_image }}" + - name: supervisor + src: "{{ openshell_supervisor_image }}" + + - name: Import OpenShell images into K3s + ansible.builtin.command: + argv: + - /usr/local/bin/k3s + - ctr + - --namespace + - k8s.io + - images + - import + - "/var/lib/openshell/artifacts/{{ item }}.tar" + loop: + - gateway + - sandbox + - supervisor + + - name: Copy OpenShell Helm chart + ansible.builtin.copy: + src: "{{ openshell_helm_chart }}" + dest: /var/lib/openshell/artifacts/helm-chart.tgz + mode: "0600" + + - name: Write OpenShell Helm values + ansible.builtin.copy: + dest: /var/lib/openshell/artifacts/values.yaml + mode: "0600" + content: | + global: + image: + registry: "" + gateway: + image: + repository: openshell/gateway + tag: tmachine + pullPolicy: Never + sandboxRuntime: + image: + repository: openshell/sandbox + tag: tmachine + pullPolicy: never + supervisor: + image: + repository: openshell/supervisor + tag: tmachine + pullPolicy: never + networkPolicy: + enabled: true + server: + auth: + allowUnauthenticatedUsers: true + disableTls: true + telemetryEnabled: false + + - name: Install Agent Sandbox + ansible.builtin.command: + argv: + - /usr/local/bin/k3s + - kubectl + - apply + - --filename + - "https://github.com/kubernetes-sigs/agent-sandbox/releases/download/v{{ agent_sandbox_version }}/manifest.yaml" + + - name: Wait for Agent Sandbox CRD + ansible.builtin.command: + argv: + - /usr/local/bin/k3s + - kubectl + - wait + - --for=condition=Established + - crd/sandboxes.agents.x-k8s.io + - --timeout=120s + changed_when: false + + - name: Wait for Agent Sandbox controller + ansible.builtin.command: + argv: + - /usr/local/bin/k3s + - kubectl + - --namespace + - agent-sandbox-system + - rollout + - status + - deployment/agent-sandbox-controller + - --timeout=300s + changed_when: false + + - name: Install OpenShell Helm chart + ansible.builtin.command: + argv: + - /usr/local/bin/helm + - install + - openshell + - /var/lib/openshell/artifacts/helm-chart.tgz + - --namespace + - openshell + - --create-namespace + - --values + - /var/lib/openshell/artifacts/values.yaml + - --wait + - --timeout=5m + environment: + KUBECONFIG: /etc/rancher/k3s/k3s.yaml + + - name: Install gateway port-forward service + ansible.builtin.copy: + dest: /etc/systemd/system/openshell-k3s-port-forward.service + mode: "0644" + content: | + [Unit] + Description=OpenShell K3s gateway port forward + After=k3s.service + Requires=k3s.service + + [Service] + ExecStart=/usr/local/bin/k3s kubectl --namespace openshell port-forward --address 127.0.0.1 service/openshell 17670:8080 + Restart=always + RestartSec=1 + + [Install] + WantedBy=multi-user.target + + - name: Start gateway port-forward service + ansible.builtin.systemd_service: + name: openshell-k3s-port-forward.service + daemon_reload: true + enabled: true + state: started + + - name: Wait for OpenShell gateway + ansible.builtin.wait_for: + host: 127.0.0.1 + port: 17670 + timeout: 60 + +- name: Register OpenShell gateway for test client + hosts: all + gather_facts: false + roles: + - openshell_client diff --git a/tests/artifacts.nix b/tests/artifacts.nix index c3017a5da5..c86d8f6d30 100644 --- a/tests/artifacts.nix +++ b/tests/artifacts.nix @@ -116,10 +116,6 @@ let "podman_corporate_proxy" "podman_gateway_start" "podman_oci_identity" - # This validates the standalone driver binary's daemon-unavailable path and - # belongs in the Podman driver crate's integration tests. The E2E archive - # does not contain `openshell-driver-podman`. - "podman_preflight" "provider_auto_create" # The provider-refresh feature suite covers revoked Keycloak grants. This # binary instead covers stable workload handles across repeated rotations diff --git a/tests/config.nix b/tests/config.nix index a19b81cee8..4bf7488ff8 100644 --- a/tests/config.nix +++ b/tests/config.nix @@ -43,6 +43,17 @@ let ]; }; } + { + name = "ubuntu-k3s"; + machine = "ubuntu"; + setup = { + use_galaxy = false; + playbooks = [ + "ansible/playbooks/nextest.yaml" + "ansible/playbooks/k3s.yaml" + ]; + }; + } { name = "fedora-podman-rootful"; machine = "fedora"; @@ -70,6 +81,19 @@ let ]; installers = [ + { + name = "k3s"; + use_galaxy = false; + playbooks = [ "ansible/playbooks/openshell-k3s.yaml" ]; + inputs = { + agent_sandbox_version = "0.5.0"; + openshell_cli_binary = "../artifacts/binaries/${muslTarget}/openshell"; + openshell_gateway_image = "../artifacts/images/openshell-gateway-tmachine.tar"; + openshell_helm_chart = "../artifacts/helm/helm-chart-0.0.0.tgz"; + openshell_sandbox_image = "../artifacts/images/openshell-sandbox-tmachine.tar"; + openshell_supervisor_image = "../artifacts/images/openshell-supervisor-tmachine.tar"; + }; + } { name = "none"; use_galaxy = false; diff --git a/tests/suites/features/Cargo.lock b/tests/suites/features/Cargo.lock index 268f8bac92..8561dc97df 100644 --- a/tests/suites/features/Cargo.lock +++ b/tests/suites/features/Cargo.lock @@ -38,7 +38,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "b281d307588d634de920874890732659e2e7672f72b5e10e81badc1a8a83621e" dependencies = [ "aws-lc-sys", - "untrusted", + "untrusted 0.7.1", "zeroize", ] @@ -278,7 +278,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "39cab71617ae0d63f51a36d69f866391735b51691dbda63cf6f96d042b63efeb" dependencies = [ "libc", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -804,7 +804,7 @@ checksum = "4b18443e9c262bfe8fa82f51666e2642c53393f7e5c27b3e1aeab922cff5b9d8" dependencies = [ "libc", "wasi", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -903,12 +903,15 @@ dependencies = [ "noyalib", "prost", "rand", + "rustls", + "rustls-pemfile", "serde", "serde_json", "sha1", "sha2", "tempfile", "tokio", + "tokio-rustls", "tokio-stream", "tonic", "tonic-prost", @@ -1122,6 +1125,20 @@ dependencies = [ "bitflags", ] +[[package]] +name = "ring" +version = "0.17.14" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "a4689e6c2294d81e88dc6261c768b63bc4fcdb852be6d1352498b114f61383b7" +dependencies = [ + "cc", + "cfg-if", + "getrandom 0.2.17", + "libc", + "untrusted 0.9.0", + "windows-sys 0.52.0", +] + [[package]] name = "rustc-hash" version = "2.1.3" @@ -1138,7 +1155,52 @@ dependencies = [ "errno", "libc", "linux-raw-sys", - "windows-sys", + "windows-sys 0.61.2", +] + +[[package]] +name = "rustls" +version = "0.23.45" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "0d41d731c7d2f962d1ccc364cec258de3c0e93b38c2fb3ba97ac74513048d634" +dependencies = [ + "aws-lc-rs", + "log", + "once_cell", + "rustls-pki-types", + "rustls-webpki", + "subtle", + "zeroize", +] + +[[package]] +name = "rustls-pemfile" +version = "2.2.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "dce314e5fee3f39953d46bb63bb8a46d40c2f8fb7cc5a3b6cab2bde9721d6e50" +dependencies = [ + "rustls-pki-types", +] + +[[package]] +name = "rustls-pki-types" +version = "1.15.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "2f4925028c7eb5d1fcdaf196971378ed9d2c1c4efc7dc5d011256f76c99c0a96" +dependencies = [ + "zeroize", +] + +[[package]] +name = "rustls-webpki" +version = "0.103.15" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "f3c3cf1d8b1e7d4927e2d154c3fcb02979afb9939629c62cd9048d4f07b60ac2" +dependencies = [ + "aws-lc-rs", + "ring", + "rustls-pki-types", + "untrusted 0.9.0", ] [[package]] @@ -1303,7 +1365,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "c3d1e2c7f27f8d4cb10542a02c49005dbd6e93095799d6f3be745fae9f8fedd4" dependencies = [ "libc", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -1312,6 +1374,12 @@ version = "1.2.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "6ce2be8dc25455e1f91df71bfa12ad37d7af1092ae736f3a6cd0e37bc7810596" +[[package]] +name = "subtle" +version = "2.6.1" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292" + [[package]] name = "syn" version = "2.0.119" @@ -1361,7 +1429,7 @@ dependencies = [ "getrandom 0.4.3", "once_cell", "rustix", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -1438,7 +1506,7 @@ dependencies = [ "signal-hook-registry", "socket2", "tokio-macros", - "windows-sys", + "windows-sys 0.61.2", ] [[package]] @@ -1452,6 +1520,16 @@ dependencies = [ "syn 3.0.5", ] +[[package]] +name = "tokio-rustls" +version = "0.26.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "c9cc2678c2cdd569ef8215e2afd7954ada2ae20b4fdd2c5fe6139a3b02d105db" +dependencies = [ + "rustls", + "tokio", +] + [[package]] name = "tokio-stream" version = "0.1.19" @@ -1603,6 +1681,12 @@ version = "0.7.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "a156c684c91ea7d62626509bce3cb4e1d9ed5c4d978f7b4352658f96a4c26b4a" +[[package]] +name = "untrusted" +version = "0.9.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "8ecb6da28b8a351d773b68d5825ac39017e680750f980f3a1a85cd8dd28a47c1" + [[package]] name = "url" version = "2.5.8" @@ -1724,6 +1808,15 @@ version = "0.2.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5" +[[package]] +name = "windows-sys" +version = "0.52.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "282be5f36a8ce781fad8c8ae18fa3f9beff57ec1b52cb3de0789201425d9a33d" +dependencies = [ + "windows-targets", +] + [[package]] name = "windows-sys" version = "0.61.2" @@ -1733,6 +1826,70 @@ dependencies = [ "windows-link", ] +[[package]] +name = "windows-targets" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "9b724f72796e036ab90c1021d4780d4d3d648aca59e491e6b98e725b84e99973" +dependencies = [ + "windows_aarch64_gnullvm", + "windows_aarch64_msvc", + "windows_i686_gnu", + "windows_i686_gnullvm", + "windows_i686_msvc", + "windows_x86_64_gnu", + "windows_x86_64_gnullvm", + "windows_x86_64_msvc", +] + +[[package]] +name = "windows_aarch64_gnullvm" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "32a4622180e7a0ec044bb555404c800bc9fd9ec262ec147edd5989ccd0c02cd3" + +[[package]] +name = "windows_aarch64_msvc" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "09ec2a7bb152e2252b53fa7803150007879548bc709c039df7627cabbd05d469" + +[[package]] +name = "windows_i686_gnu" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "8e9b5ad5ab802e97eb8e295ac6720e509ee4c243f69d781394014ebfe8bbfa0b" + +[[package]] +name = "windows_i686_gnullvm" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "0eee52d38c090b3caa76c563b86c3a4bd71ef1a819287c19d586d7334ae8ed66" + +[[package]] +name = "windows_i686_msvc" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "240948bc05c5e7c6dabba28bf89d89ffce3e303022809e73deaefe4f6ec56c66" + +[[package]] +name = "windows_x86_64_gnu" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "147a5c80aabfbf0c7d901cb5895d1de30ef2907eb21fbbab29ca94c5b08b1a78" + +[[package]] +name = "windows_x86_64_gnullvm" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "24d5b23dc417412679681396f2b49f3de8c1473deb516bd34410872eff51ed0d" + +[[package]] +name = "windows_x86_64_msvc" +version = "0.52.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "589f6da84c646204747d1270a2a5661ea66ed1cced2631d546fdfb155959f9ec" + [[package]] name = "wit-bindgen" version = "0.57.1"