From 3ba31e92d27684de18e8c1df4957c621ecf27e42 Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 08:27:24 +0000 Subject: [PATCH 01/10] docs: merge Runtime telemetry contracts into one API and one internal contract Runtime observations, Session history and node host history share one Core-key API document; collection, storage, OTLP export and provider mapping share one internal contract. The design document, the separate history and node host history documents are removed. Session diagnostics keep their own contract without restated rules. --- contracts/agents-api/node-host-history.md | 79 --- contracts/agents-api/runtime-history-api.md | 160 ----- .../agents-api/runtime-observability-api.md | 379 +++++------ .../runtime-observability-design.md | 594 ------------------ contracts/agents-api/runtime-observability.md | 282 ++++----- contracts/agents-api/session-diagnostics.md | 95 +-- 6 files changed, 323 insertions(+), 1266 deletions(-) delete mode 100644 contracts/agents-api/node-host-history.md delete mode 100644 contracts/agents-api/runtime-history-api.md delete mode 100644 contracts/agents-api/runtime-observability-design.md diff --git a/contracts/agents-api/node-host-history.md b/contracts/agents-api/node-host-history.md deleted file mode 100644 index 6c4a03af1..000000000 --- a/contracts/agents-api/node-host-history.md +++ /dev/null @@ -1,79 +0,0 @@ -# Administrator node host observations - -`GET /core/v1/sandbox/nodes/{node_id}?range=1h|6h|24h` requires the Core key. -Core Web forwards it with the Core key after console sign-in; Project, enrollment -and node credentials do not authorize this read. -The omitted range defaults to `1h`. Invalid, repeated or unknown query parameters -return the existing invalid-input response; missing or removed nodes return 404. - -The response contains the existing node fields plus these two objects. The node -list response is unchanged. This endpoint does not alter enrollment, administrator -capacity approval, scheduling or any public Agents API contract. - -```json -{ - "host": { - "effective_cpu_cores": 4, - "cpu_utilization": 0.35, - "total_memory_bytes": 17179869184, - "available_memory_bytes": 8589934592, - "available_disk_bytes": 107374182400, - "observed_at": "2026-09-25T09:00:00Z" - }, - "history": { - "resolution_seconds": 60, - "points": [{ - "start": "2026-09-25T08:59:00Z", - "cpu_utilization_max": 0.4, - "memory_used_bytes_max": 8589934592, - "available_disk_bytes_min": 107374182400 - }] - } -} -``` - -Every unavailable measurement is `null`, including unobserved timestamps. `host` -is the last received observation; an offline node retains its original timestamp -and last-known measurements. Clients must use `online` and `observed_at` when -presenting freshness. Reading the endpoint never samples or writes history. - -On Linux, CPU utilization is the increase in aggregate busy `/proc/stat` ticks -divided by total ticks between heartbeats, across the whole visible host. Idle -and I/O-wait ticks are not busy; guest counters are not counted twice. The first -observation and a reset or unavailable baseline have null utilization. This is -not node-process CPU or the sum of sandbox utilization. Total memory and available -memory come from `MemTotal` and `MemAvailable`. Effective CPU capacity accounts for -observable process affinity and cgroup limits; it is null if those limits cannot -be established, rather than an invented scheduling limit. Available disk refers -to the node's existing state filesystem, not a sandbox quota. The node should run -on the machine being monitored; namespace visibility limits what it can observe. - -The existing Runtime sampling sweep copies fresh, authenticated node heartbeats -into `node_host_history_samples` in the same PostgreSQL database. It uses the -Runtime history cadence (30 seconds by default) and seven-day retention cleanup. -A node observation is keyed by node and its original timestamp, so repeated sweeps -do not manufacture additional observations. Disconnected, stale, future-dated or -removed observations are not copied. The sampler is best effort with a bounded -query timeout; failures do not change execution or admission authority. - -Ranges use complete UTC buckets: `1h` has 60-second buckets, `6h` has 300-second -buckets and `24h` has 900-second buckets. Each response includes every bucket in -the range. CPU and memory are maxima of recorded observations; used memory is -`total_memory_bytes - available_memory_bytes` from the same observation. Disk is -the minimum recorded available space. Missing measurements and offline buckets -stay null, independently for each metric. The active partial bucket is excluded. -No interpolation, offline backfill or whole-host process inventory is performed. - -History survives Core and node restarts until normal retention expiry. Core -process history remains a separate in-memory series and resets when Core restarts. -Neither series contains credentials, model data, paths or user content. - -## Implementation rules - -Administrator node detail adds host observations and history as documented in -[node-host-history.md](node-host-history.md). Keep the node -list unchanged apart from the address each node enrolled with (`core_url`). Reuse authenticated heartbeat ownership, the Runtime sampling -sweep and PostgreSQL retention cleanup; node observations have their own table -because they do not belong to a Project, Session or Environment. History is -best-effort telemetry, never scheduling truth. No read-triggered sampling or -offline backfill is allowed. diff --git a/contracts/agents-api/runtime-history-api.md b/contracts/agents-api/runtime-history-api.md deleted file mode 100644 index b44716154..000000000 --- a/contracts/agents-api/runtime-history-api.md +++ /dev/null @@ -1,160 +0,0 @@ -# Runtime history API - -Status: Core-key contract, strict TypeScript client, PostgreSQL history and Core Web -History ranges implemented. Core uses its existing database; the execution owner -samples every 30 seconds by default. API-only processes without an execution worker -collect only on read; their Session history read returns -`503 runtime_history_unavailable` rather than claiming Durable coverage. - -This is a read-only, backend-neutral administrator read under `/core/v1`, -authenticated by the Core key. The former project routes -`GET /v1/agents/runtime-history/capabilities` and -`GET /v1/agents/sessions/{session_id}/runtime-history`, and the administrator -capability read, are removed. History is available only with a validated Reader, -retention and query bounds and qualified periodic collection; otherwise, or when -that configuration is malformed, the Session read fails closed with 503. Backend -names, URLs, credentials, table names and tenant data are never returned. The -browser never receives a storage endpoint, OTLP credential, provider-native -identity or tenant selector. - -## Session history - -```http -GET /core/v1/projects/{project_id}/sessions/{session_id}/runtime-history?start=1789951200&end=1789954800&max_points=120 -Authorization: Bearer ... -``` - -`start` is inclusive and `end` is exclusive, both in whole Unix seconds. -`max_points` defaults to the lower of 120 and the configured maximum. Core -selects an effective whole-second resolution. Unknown parameters, duplicate -parameters, negative timestamps, invalid ranges, and invalid point limits are -rejected before a Reader query. - -The caller supplies a Project ID and a Session ID. Core resolves the Project's -tenant, then the Session and Environment from its store, and only then calls the -Reader. Missing Sessions and other Projects' Sessions remain indistinguishable. Allocation IDs, -provider IDs, and backend labels are result identity, never authority-bearing -query inputs. - -```json -{ - "object": "agent.runtime_history", - "source": "durable", - "session_id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", - "requested_range": { "start": 1789951200, "end": 1789954800 }, - "resolution_seconds": 60, - "generated_at": 1789954801, - "coverage": { - "retained_start": 1789951200, - "first_sample_at": 1789951210, - "last_sample_at": 1789954750, - "sample_count": 118, - "expected_sample_count": 120, - "buckets": [] - }, - "series": [], - "token_usage": [] -} -``` - -`coverage` describes all resolved samples, including unavailable observations that -cannot safely be attached to one allocation series. `retained_start` is the -latest of the requested start, configured retention boundary, and backend-reported -retention boundary. Expected coverage is calculated only over that retained -period and only from qualified periodic cadence. - -Resource `series` are keyed only by `allocation_id`. This keeps one continuous -Dashboard lifecycle when a provider pauses, restores, restarts, or replaces its -underlying compute without replacing the durable allocation. `started_at` remains -the earliest retained provider start estimate; it is -not series identity: - -```json -{ - "started_at": { - "seconds": 1789951210, - "nanoseconds": 123456789 - } -} -``` - -The two integers are JSON-safe, nonnegative Unix seconds and a 0–999,999,999 -nanosecond remainder. CPU utilization is derived from ordered cumulative counters -inside the allocation. A counter regression resets the baseline, so no interval -is derived across a compute replacement. Successive valid counter intervals are -assigned to the bucket containing their right endpoint and combined by CPU-capacity -time. E2B reports no cumulative CPU time: each periodic E2B sample stores its -reported utilization ratio, and the bucket's `utilization_ratio` is the mean of -the ratios sampled in it, in the same field. Disk observations are not retained. Memory values are the last observed values in a bucket. Every point contains -observation and contributor counts. Missing values are null and gaps remain gaps. -Numeric zero is retained as an observed value. - -`token_usage` is Session-scoped rather than allocation-scoped. Each point is the -last cumulative measured Session usage sampled in that bucket and -contains `input_tokens`, `output_tokens`, and `sampled_at`. Web derives throughput -only from adjacent nondecreasing cumulative points. A missing measurement or a -counter regression produces a gap; it is never filled with zero. The sampled -value is Core's measured Session usage, a Core extension that sums every -recorded root Turn snapshot, active Turns included. It differs by design from -public Session usage, which is null while a root Turn runs or after one ends -unmeasured ([usage](sessions-events.md#usage)). These counters -are measured model usage, not price, cost, or billing records. - -Core Web queries each current managed Session through the administrator Session -route with bounded concurrency and an all-or-nothing target budget, offering 1h, -6h, and 24h History ranges. Reloading Web reconstructs the -charts from the backend; Live browser samples and Durable buckets remain explicit -separate sources and are never silently merged. - -## Errors and bounds - -| HTTP | Code | Meaning | -| --- | --- | --- | -| 400 | `unsupported_parameter` or `invalid_request` | Invalid query shape or range. | -| 401 | `invalid_admin_key` | Missing or invalid Core key. | -| 404 | `not_found` | Missing Project, or a Session outside it. | -| 409 | `runtime_history_unsupported` | The Session has no supported managed Runtime history scope. | -| 503 | `runtime_history_unavailable` | Durable history is unconfigured, timed out, unavailable, or returned malformed data. | - -Reader failures and malformed results are sanitized. Raw backend errors are not -logged or returned. Results are bounded by per-series points, series count, and -total response points. Coverage and series points must be ordered, non-overlapping, -inside the half-open request range, and contain valid safe JSON values. - -## Client contract - -`packages/agents-client` exposes: - -```ts -class AdminClient { - retrieveRuntimeHistory(projectId: string, sessionId: string, query: { - start: number; - end: number; - maxPoints?: number; - signal?: AbortSignal; - }): Promise; -} -``` - -The client validates exact fields, requested-range echo, -Session identity, half-open bucket ordering, coverage totals, allocation identity, -contributor counts, token usage ordering, nullability, finite numbers, and response size. Unknown fields -or malformed data reject the entire response with a 502 `invalid_admin_response` error. - -## Explicit boundaries - -- The routes never sample a live provider, provision compute, or mutate lifecycle. -- The contract does not expose a storage backend. -- History availability does not imply current Runtime readiness. -- Current observations and Durable history have separate freshness and retention - semantics and must remain separately labelled in Web. - -Compute uptime is available from current observations only. Retained allocation -series can span compute restarts and unavailable intervals; their earliest start -is not a per-bucket compute start and must not be used to draw an uptime history. -Clients may project a Dashboard active-Sandbox count by counting distinct -allocation identities with `observed_count > 0` in each bucket and deduplicating -the same allocation across Session histories. The single-Session presentation -collapses any positive count to `1`; a missing or unavailable bucket is currently -rendered as `0`. This temporary zero-fill policy does not distinguish a sleeping -Runtime from missing collection coverage. diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md index 507bc07ab..dafaaee76 100644 --- a/contracts/agents-api/runtime-observability-api.md +++ b/contracts/agents-api/runtime-observability-api.md @@ -1,40 +1,25 @@ -# Runtime observation API +# Runtime telemetry API -Status: Phase 2 and initial Core Web consumption implemented. The current-snapshot routes, strict -`packages/agents-client` projection, and generated `core.openapi.yaml` contract are -implemented and consumed by the Dashboard through complete Session/observation -identity joins. Durable history uses the separate optional -[Runtime history API](runtime-history-api.md); lifecycle controls remain outside -this phase. +Core reports what hosted Runtimes and sandbox nodes consume through read-only administrator routes under `/core/v1`: current Runtime observations, the stored Runtime history of one Session, and the host observations and history of a sandbox node. Reads never create, wake, renew or change compute and never add samples to history. [Runtime observability](runtime-observability.md) defines how Core collects and keeps these values; [Console API usage](../../docs/web/console-api-usage.md) lists the Web pages that read them. -These are administrator reads under `/core/v1`, authenticated by the Core key, -not upstream OpenAI Agents resources. The -former project routes `GET /v1/agents/runtime-observations` and -`GET /v1/agents/sessions/{session_id}/runtime-observation` are removed. +Every route requires the Core key as the bearer credential; a missing or invalid key returns 401 `invalid_admin_key`. A Project ID in a path selects the target Project and does not authenticate. Responses carry `Cache-Control: no-store`, use the Core error envelope and never contain provider responses, native identifiers, paths or credentials. -## Routes +## Current Runtime observations -### List current Runtime observations +### List observations of every Project ```http GET /core/v1/sandbox/runtime-observations?after={session_id}&limit=20&order=desc -Authorization: Bearer ... +Authorization: Bearer ``` -| Field | Rules | +| Parameter | Rules | | --- | --- | -| `after` | Observation ID from the previous page. Optional, supplied once. | -| `limit` | Integer 1–100, default 20. | -| `order` | `asc` or `desc`, default `desc`. | - -The list contains one current Runtime context for every Session of every managed -Project, labelled with its owning `project_id`, including explicit `none`, -unsupported `self_hosted`, and released managed contexts. Ordering uses the same Session creation-time and ID -keyset as the Session list. An observation ID is the Session UUID, so pagination -does not change when the underlying Runtime incarnation changes. Pages are not an -atomic telemetry snapshot; every row has its own `resolved_at`, and a successful -provider sample has its own `observed_at`. A client completes the entire page chain -before publishing a new Dashboard snapshot. +| `after` | Observation ID (a Session ID) that ended the previous page. | +| `limit` | 1 to 100, default 20. | +| `order` | `asc` or `desc` by Session creation time, default `desc`. | + +The list has one row for every Session of every Project that is not deleted, including `none`, `self_hosted` and released managed Sessions. Each row carries the owning `project_id` and an `observation`: the [`RuntimeObservation`](#runtimeobservation) plus [`disk`](#disk). Pages use the Session list's creation-time and ID keyset. The observation ID is the Session ID, so page boundaries do not move when the Runtime behind a Session changes. A page is not an atomic snapshot: each row has its own `resolved_at` and, when sampled, `observed_at`. Unknown query keys are ignored. ```json { @@ -55,6 +40,7 @@ before publishing a new Dashboard snapshot. "device_id": "2e434f4f-76aa-4e54-a707-4757036d90ef", "connection_generation": null }, + "lifecycle_state": "active", "status": "observed", "reason": null, "allocation_created_at": 1789951200, @@ -81,197 +67,218 @@ before publishing a new Dashboard snapshot. } ``` -### Retrieve one Session's current Runtime observation +### Retrieve one Session's observation ```http GET /core/v1/projects/{project_id}/sessions/{session_id}/runtime-observation -Authorization: Bearer ... +Authorization: Bearer ``` -This returns the same object shape as a list item's `observation`, without the -administrator list's `disk` (see the [administrator contract](admin-api.md)). It never starts a Turn, creates -an Environment, provisions compute, renews a lease, or changes lifecycle state. - -A valid `environment:none` Session returns `200` with status `unsupported`; the -Session exists but has no attributable Runtime instance. A missing Session, or one -outside the Project, returns the existing indistinguishable not-found error. - -## Resource schema +This returns one `RuntimeObservation`, without `disk`. It accepts no query parameters. An `environment:none` Session returns 200 with status `unsupported`. ### `RuntimeObservation` -| Field | Type | Required | Semantics | -| --- | --- | --- | --- | -| `id` | string | yes | Session UUID; stable identity of this current-observation resource and its list cursor. | -| `object` | literal | yes | `agent.runtime_observation`. | -| `session_id` | string | yes | Authorized Core Session. | -| `environment_id` | string or null | yes | Null only for mode `none`. | -| `mode` | enum | yes | `none`, `self_hosted`, `openai_hosted`. | -| `provider_type` | string or null | yes | Forward-compatible safe source kind such as `docker` or `microsandbox`; null when no provider applies. Clients must not treat an unknown nonempty value as an error. | -| `instance` | object | yes | Provider-neutral current incarnation identity; explicit `kind=none` when no compute applies. | -| `status` | enum | yes | `observed`, `unsupported`, `unavailable`. | -| `reason` | enum or null | yes | Safe reason when status is not `observed`. | -| `allocation_created_at` | integer or null | yes | Unix seconds for managed allocation age. | -| `resolved_at` | integer | yes | Unix seconds when Core resolved identity and status for this row. | -| `observed_at` | integer or null | yes | Provider sample time; null without a sample. | -| `started_at` | integer or null | yes | Current compute incarnation start time. | -| `cpu` | object or null | yes | Null when no CPU fields were observed. | -| `memory` | object or null | yes | Null when no memory fields were observed. | +| Field | Type | Meaning | +| --- | --- | --- | +| `id` | string | The Session ID; the stable identity of this resource and its list cursor. | +| `object` | string | `agent.runtime_observation`. | +| `session_id` | string | The Session. | +| `environment_id` | string or null | Null only for mode `none`. | +| `mode` | enum | `none`, `self_hosted` or `openai_hosted`. | +| `provider_type` | string or null | Source kind, such as `docker`, `microsandbox` or `e2b`; null when no provider was read. Treat an unknown value as a new kind, not an error. | +| `instance` | object | The current compute identity; see [`RuntimeInstance`](#runtimeinstance). | +| `lifecycle_state` | enum or null | Core's own lifecycle view of a managed allocation; null for `none` and `self_hosted`. See below. | +| `status` | enum | `observed`, `unsupported` or `unavailable`. | +| `reason` | enum or null | Why the row has no sample; see [Status and reason](#status-and-reason). | +| `allocation_created_at` | integer or null | Unix seconds when the managed allocation was created. | +| `resolved_at` | integer | Unix seconds when Core resolved this row. | +| `observed_at` | integer or null | Unix seconds of the provider sample; null without a sample. | +| `started_at` | integer or null | Unix seconds when the current compute incarnation started. | +| `cpu` | object or null | Null when no CPU value was observed. | +| `memory` | object or null | Null when no memory value was observed. | + +`lifecycle_state` comes from Core's allocation records, never from the sample: + +| Value | Allocation | +| --- | --- | +| `pending` | Not created yet, or being created | +| `active` | Running | +| `sleeping` | Suspended | +| `transitioning` | Quiescing, suspending, restoring or waking | +| `stopped` | Cleanup pending, or released | ### `RuntimeInstance` -```json -{ - "kind": "managed_allocation", - "allocation_id": "alloc_...", - "device_id": "device_...", - "connection_generation": null -} -``` +| Field | Meaning | +| --- | --- | +| `kind` | `managed_allocation`, `self_hosted_connection` or `none`. | +| `allocation_id` | The managed allocation, which identifies the compute of a managed Session; null otherwise. | +| `device_id` | The Runtime device bound to the managed allocation, when there is one; null otherwise. | +| `connection_generation` | Always null: Core does not observe self-hosted connections. | -`kind` is `managed_allocation`, `self_hosted_connection`, or `none`. For a managed -context, `allocation_id` is the incarnation key; for self-hosted, the current -`connection_generation` is the incarnation key. Fields that do not apply are -explicit nulls. Provider-native container IDs, pod names, host paths, credentials, -and raw labels are not public fields. +### `cpu` -### `RuntimeCPUObservation` +All fields are finite nonnegative numbers or null. Zero is an observed zero; null is unavailable. -```json -{ - "usage_seconds_total": 482.75, - "capacity_cores": 2.0, - "usage_cores": 1.42, - "utilization_ratio": 0.71 -} +| Field | Meaning | +| --- | --- | +| `usage_seconds_total` | Cumulative CPU seconds of the current compute incarnation (Docker, microsandbox). | +| `capacity_cores` | Configured CPU capacity, greater than zero. | +| `usage_cores` | Always null. | +| `utilization_ratio` | Provider-reported share of `capacity_cores` (E2B), not clamped; null for providers that report cumulative CPU time. | + +### `memory` + +`usage_bytes` and `limit_bytes` are safe JSON integers or null. Zero usage is observed zero; `limit_bytes` is at least 1, and an unknown or unlimited limit is null. + +### `disk` + +Only list rows carry `disk`: null, or `{usage_bytes, limit_bytes}` with the rules of `memory`. E2B fills it when the sandbox reports both its disk usage and a nonzero capacity. Docker and microsandbox return null. A non-null `disk` appears only on an `observed` row. + +### Status and reason + +| Status | Reason | When | +| --- | --- | --- | +| `observed` | null | The provider returned a sample. | +| `unsupported` | `runtime_mode_not_observable` | `none` and `self_hosted` Sessions. | +| `unavailable` | `allocation_pending` | The managed allocation does not exist yet or is being created. | +| `unavailable` | `runtime_not_running` | The allocation is being cleaned up or is released, or the provider reports the Runtime absent, stopped or suspended. | +| `unavailable` | `source_not_configured` | No observation source serves the allocation's provider. | +| `unavailable` | `sample_timeout` | The provider read exceeded its deadline. | +| `unavailable` | `sample_unavailable` | The provider could not produce a current sample. | + +An ownership mismatch, malformed durable identity or invalid provider evidence fails the request instead of becoming an `unavailable` row. The generated `core.openapi.yaml` records each field's type, nullability and enum but cannot express which combinations of status, mode and fields are valid; this table and the field rules above are normative. + +### Errors + +| HTTP | Code | When | +| --- | --- | --- | +| 400 | `invalid_request_error` | List: a repeated `after`, `limit` or `order`, or an invalid `limit` or `order`. | +| 400 | `unsupported_parameter` | Single read: any query parameter. | +| 404 | `not_found_error` | A missing Project; a missing, malformed or foreign Session or list cursor. | +| 500 | `internal_error` | Inconsistent identity or invalid provider evidence. | +| 503 | `execution_unavailable` | The list exceeded its collection budget. | + +### Client + +`packages/agents-client` exposes `AdminClient.listRuntimeObservations({after, limit, order})` and `AdminClient.retrieveRuntimeObservation(projectId, sessionId, options)`. `RuntimeObservation` in `src/types.ts` is a union discriminated by `status` and `mode`; `AdminRuntimeObservation` adds `disk`. The client checks every field, enum, nullability rule, timestamp and number and rejects unknown fields. A malformed observation rejects with a 502 `invalid_runtime_observation` error and a malformed page with `invalid_admin_response`; one bad row rejects the whole page. + +## Session Runtime history + +```http +GET /core/v1/projects/{project_id}/sessions/{session_id}/runtime-history?start=1789951200&end=1789954800&max_points=120 +Authorization: Bearer ``` -All fields are `number | null`. Values are finite and nonnegative; -`capacity_cores`, when present, is greater than zero. Numeric zero is observed -zero. Null is unavailable. `usage_cores` is the cumulative CPU delta divided by -the observation-time delta for two ordered samples of the same incarnation. -`utilization_ratio` is `usage_cores / capacity_cores`. It is not clamped: a value -above 1 is retained as provider/accounting evidence and is not interpreted as a -lifecycle signal. Both derived fields are null after a cache restart or whenever -either source sample is absent or invalid. The API never derives CPU rate from a -single sample of cumulative time. A provider that reports only a current share -of its CPU capacity (E2B) fills `utilization_ratio` with that report and leaves -`usage_seconds_total` and `usage_cores` null. +| Parameter | Rules | +| --- | --- | +| `start` | Required. Inclusive Unix second, 0 or more. | +| `end` | Required. Exclusive Unix second, after `start`, at most 24 hours after it and at most one second in the future. | +| `max_points` | Optional. Buckets per array, 2 to 1000; default 120. | + +Each parameter may appear once. Core chooses the bucket width: the range divided by `max_points`, rounded up to whole seconds, and at least 30 seconds or the sampling interval, whichever is longer. Buckets start at `start`; the last one ends at `end`. -### `RuntimeMemoryObservation` +Core resolves the Project, then the Session and its Environment, before it reads storage; allocation and provider identities are results, never query inputs. History exists only for `openai_hosted` Sessions. ```json { - "usage_bytes": 805306368, - "limit_bytes": 2147483648 + "object": "agent.runtime_history", + "source": "durable", + "session_id": "6c77d3a2-71d6-4ed5-884f-687aecda02a3", + "requested_range": { "start": 1789951200, "end": 1789954800 }, + "resolution_seconds": 60, + "generated_at": 1789954801, + "coverage": { + "retained_start": 1789951200, + "first_sample_at": 1789951210, + "last_sample_at": 1789954750, + "sample_count": 118, + "expected_sample_count": 120, + "buckets": [] + }, + "series": [], + "token_usage": [] } ``` -Both fields are `integer | null`. Values are nonnegative and safe JSON integers. -Zero usage is observed zero. A missing or unlimited provider limit is null. +`source` is always `durable`. `resolution_seconds` is the bucket width and `generated_at` the read time. Only buckets that hold at least one sample appear in `coverage.buckets`, `series[].points` and `token_usage`; a gap stays a gap, never a zero. -## Status and reason matrix +### Coverage -| Status | Allowed reason | -| --- | --- | -| `observed` | null | -| `unsupported` | `runtime_mode_not_observable` | -| `unavailable` | `allocation_pending`, `runtime_not_running`, `source_not_configured`, `sample_timeout`, `sample_unavailable` | +`coverage` counts every stored sample of the Session in the range, including unavailable ones that belong to no allocation. `retained_start` is the later of `start` and seven days before `generated_at`. `expected_sample_count` is the number of sampling intervals between `retained_start` and `end`, rounded up. Each bucket has `start`, `end`, `first_observed_at`, `last_observed_at`, `observation_count`, `observed_count` and `unavailable_count`. -Ownership mismatch, malformed durable identity, corrupt provider evidence, and -authorization failure are not downgraded to unavailable rows. +### Series -## Error responses +There is one series per managed allocation, keyed by `allocation_id`, so a provider that pauses, restores or replaces compute under the same allocation keeps one series. `environment_id` and `provider_type` identify its source. `started_at` is the earliest retained start of the allocation's compute, as JSON-safe `{seconds, nanoseconds}` with nanoseconds 0 to 999,999,999; it is not a per-bucket start, and compute uptime comes only from current observations. -Use the existing Agents API error envelope. +Each point has the bucket bounds, the coverage counts of the allocation's samples, and nullable `cpu` and `memory` objects with a `contributor_count` of at least 1: -| HTTP | Type / code | When | +- `cpu.utilization_ratio` comes from consecutive cumulative CPU counters of one compute incarnation: the CPU seconds consumed divided by the elapsed time multiplied by the capacity, over the intervals that end in the bucket. The baseline resets when the incarnation changes or a counter decreases. E2B reports no cumulative CPU time; its bucket value is the mean of the ratios sampled in it. `cpu.capacity_cores` is the last capacity in the bucket. +- `memory.usage_bytes` and `memory.limit_bytes` are the last values observed in the bucket. + +Disk is not kept in history. + +### Token usage + +`token_usage` belongs to the Session, not to an allocation. Each point holds the last cumulative measured Session usage sampled in its bucket: `start`, `end`, `sampled_at`, `input_tokens` and `output_tokens`. Measured Session usage is a Core extension that sums every recorded root Turn snapshot, active Turns included. It differs from [public Session usage](history-events-usage.md), which is null while a root Turn runs or after one ends unmeasured. These counters are measured model tokens, not prices or billing records. + +### Errors and bounds + +| HTTP | Code | When | | --- | --- | --- | -| 400 | `invalid_request_error` / `invalid_request_error` | List: a repeated supported query key, or an empty or invalid limit or order, with the shared Beta list messages. Unknown list query keys are ignored. | -| 400 | `invalid_request_error` / `unsupported_parameter` | Single-Session retrieval with any query parameter. | -| 401 | `invalid_request_error` / `invalid_admin_key` | Missing or invalid Core key. | -| 404 | `not_found_error` / `not_found_error` | Missing, malformed or foreign Session/cursor, indistinguishably, as for the [Session list cursor](wire-semantics.md#cursors). | -| 500 | `server_error` / `internal_error` | Integrity, ownership, or invalid provider evidence. | -| 503 | `server_error` / `execution_unavailable` | Required Runtime observation service is not configured, or list collection exceeded its request budget. | - -Errors never include provider raw responses or credentials. - -## Freshness and caching - -- Return `Cache-Control: no-store`. -- The Phase 2 implementation performs bounded direct reads and has no observation - cache. A later internal cache may coalesce reads for at most five seconds. -- `observed_at` is authoritative for freshness; HTTP response time is not. -- Clients mark samples stale according to their own explicit threshold. -- `ETag` is not proposed because observations change independently. - -## Client contract - -`packages/agents-client` exposes: - -```ts -type RuntimeObservationStatus = "observed" | "unsupported" | "unavailable"; -type RuntimeObservationReason = - | "runtime_mode_not_observable" - | "allocation_pending" - | "runtime_not_running" - | "source_not_configured" - | "sample_timeout" - | "sample_unavailable"; - -type RuntimeObservation = - | RuntimeObservedObservation - | RuntimeUnavailableObservation - | RuntimeNoneObservation - | RuntimeSelfHostedObservation; - -interface AdminRuntimeObservation { - project_id: string; - observation: RuntimeObservation & { disk: RuntimeDiskObservation | null }; -} +| 400 | `unsupported_parameter` | A parameter other than `start`, `end` and `max_points`, or one supplied twice. | +| 400 | `invalid_request` | An invalid range or `max_points`. | +| 404 | `not_found_error` | A missing Project, or a Session missing from it. | +| 409 | `runtime_history_unsupported` | The Session is not `openai_hosted`. | +| 500 | `internal_error` | Inconsistent stored identity. | +| 503 | `runtime_history_unavailable` | Core collects no periodic history (it runs without the execution worker), or the read failed, timed out or produced a result outside the bounds. | -class AdminClient { - listRuntimeObservations(options?: { - after?: string; - limit?: number; - order?: "asc" | "desc"; - }): Promise>; +A response holds at most `max_points` buckets per array, 64 series and 10,000 coverage and series points in total. Storage error text is neither returned nor logged. - retrieveRuntimeObservation(projectId: string, sessionId: string): Promise; +### Client + +`AdminClient.retrieveRuntimeHistory(projectId, sessionId, {start, end, maxPoints, signal})` validates the query before sending it. It then checks the exact fields, the echoed range and Session, bucket order within the range, coverage totals, allocation identity, contributor counts, token usage order, nullability, numbers and response size. Any violation rejects the whole response with a 502 `invalid_admin_response` error. + +## Node host observations and history + +```http +GET /core/v1/sandbox/nodes/{node_id}?range=1h +Authorization: Bearer +``` + +`range` is `1h` (the default), `6h` or `24h`. Another parameter, a repeated or invalid `range`, or a malformed node ID returns 400 `invalid_request`; a missing or removed node returns 404 `not_found_error`. The response is the node object of the [node list](sandbox-deployment.md) plus `host` and `history`: + +```json +{ + "host": { + "effective_cpu_cores": 4, + "cpu_utilization": 0.35, + "total_memory_bytes": 17179869184, + "available_memory_bytes": 8589934592, + "available_disk_bytes": 107374182400, + "observed_at": "2026-09-25T09:00:00Z" + }, + "history": { + "resolution_seconds": 60, + "points": [{ + "start": "2026-09-25T08:59:00Z", + "cpu_utilization_max": 0.4, + "memory_used_bytes_max": 8589934592, + "available_disk_bytes_min": 107374182400 + }] + } } ``` -These exported variants discriminate on `status` and `mode`; their instance, -reason, timestamps, CPU, and memory fields narrow accordingly. The exact variant -definitions live in `packages/agents-client/src/types.ts` and mirror the status -and reason matrix above. - -The client validates every required field, enum, nullability rule, timestamp, and -finite number. The current pinned contract rejects unknown additive fields so an -unreviewed server expansion cannot silently cross the browser boundary. Malformed -data rejects the whole page; Web does not publish a partial snapshot. - -The generated OpenAPI 2 schema records field-level required/nullability rules, -UUID formats, reason enums, and numeric minima. OpenAPI 2 -cannot encode the complete cross-field discriminated union. The matrix above is -normative for wire consumers; the server projection and strict TypeScript -projector enforce it, and the exported TypeScript type prevents invalid -status/mode combinations in typed consumers. - -Web also applies a configured whole-refresh budget. If `has_more` remains true -when that budget is exhausted, it retains the prior complete snapshot and marks -the refresh incomplete; it does not publish partial values as global totals. -After both Runtime-observation and Session traversals complete, Web also requires -their Session ID sets to be identical. A mismatch caused by concurrent creation or -deletion makes the candidate incomplete and prevents publication. - -## Deliberately excluded - -- Token usage: use existing Session/Turn Usage. -- Billing and cost: product/backend concern. -- Historical series in these routes: the optional capability uses the separate - [Runtime history API](runtime-history-api.md). -- Container logs and command output. -- Provider credentials or native configuration. -- Start, stop, pause, resume, restart, renew, or delete operations. -- Idle classification and automatic shutdown. +`host` is the node's last received heartbeat observation; every unavailable value, including an unobserved `observed_at`, is null. An offline node keeps its last values and their original time, so judge freshness by the node's `online` and `host.observed_at`. + +- `cpu_utilization` is the share of busy ticks in the host's aggregate `/proc/stat` counters between two heartbeats, 0 to 1. Idle and I/O-wait ticks are not busy, and guest time is not counted twice. The first heartbeat of a connection, a counter reset and an unreadable baseline give null. It measures the whole visible host, not the node process or its sandboxes. +- `effective_cpu_cores` accounts for the node process's CPU affinity and cgroup limits; null when those cannot be established. +- `total_memory_bytes` and `available_memory_bytes` are `MemTotal` and `MemAvailable`. +- `available_disk_bytes` is the free space of the node's state filesystem, not a sandbox quota. + +The node measures what its namespaces can see, so run it on the host it reports on. + +`history` covers the range in complete UTC buckets: 60 seconds for `1h`, 300 for `6h` and 900 for `24h`. Every bucket of the range is present, and the bucket in progress is left out. `cpu_utilization_max` and `memory_used_bytes_max` are the maxima of the recorded observations, where used memory is total minus available memory of the same observation; `available_disk_bytes_min` is the minimum. Each metric is null for a bucket without a recorded value, including offline periods. Core never interpolates or backfills. + +`SandboxAdminClient.retrieveNode(nodeId, range, options)` in `packages/agents-client` reads this route and validates the response. diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md deleted file mode 100644 index dec05fa85..000000000 --- a/contracts/agents-api/runtime-observability-design.md +++ /dev/null @@ -1,594 +0,0 @@ -# Runtime observability and Dashboard design - -Status: provider abstraction with Docker and microsandbox sampling, the -current-snapshot API/client contract, and Core Web Live and capability-gated -Durable Dashboard sources are implemented. Phase 4 includes the bounded sanitized -exporter seam, optional OTLP/HTTP transport, execution-owner singleton background -sampling, PostgreSQL history, the Core-key history read, and 1h/6h/24h Web ranges. -History uses the existing Core database by default. Other provider sources are not implemented. -Microsandbox idle suspension is a separate durable lifecycle feature; it does -not consume this telemetry as authority. - -## 1. Problem statement - -Operators need one Dashboard that answers four separate questions without -confusing their sources of truth: - -1. Which Runtime instances currently belong to which tenant, Session, and - Environment? -2. What compute is allocated and what is it consuming now? -3. How long has allocation, compute, and model work been active? -4. How many model tokens have been reported for the corresponding Sessions? - -The design must work across managed Docker now and later managed Kubernetes, -E2B, and authenticated self-hosted Runtime deployments. Metrics are operational -evidence. They must not become execution or lifecycle authority. - -## 2. Goals - -- Resolve every sample through durable Core identity before provider access. -- Keep provider-specific collection behind one source interface. -- Preserve observed zero, unavailable measurements, and unsupported modes as - different states. -- Provide a bounded read-only API suitable for Core Web and other operators. -- Let Web combine Runtime observations with existing Session and Turn usage - without copying execution truth into the browser. -- Keep current snapshots independent from an optional history backend. -- Define a safe path to future idle shutdown without implementing it implicitly. - -## 3. Non-goals - -- Redefining the pinned OpenAI Agents resources. -- Adding product users, organizations, billing, or authorization tables to Core. -- Treating a Session, daemon socket, container, pod, native harness Session, or - Turn as the same identity. -- Estimating missing CPU, memory, token, or duration values. -- Using telemetry, heartbeat age, low CPU, or Prometheus state to stop compute. -- Adding lifecycle actions to the first Dashboard release. -- Storing time-series samples in PostgreSQL. - -## 4. Source-of-truth model - -| Concern | Authority | Notes | -| --- | --- | --- | -| Tenant and Session ownership | Core database | Every read is scoped to the tenant of the Project in its `/core/v1` path. | -| Environment placement | Session configuration and Environment row | `none`, `self_hosted`, or `openai_hosted`. | -| Observation resource identity | Session ID | One current observation resource exists per tenant-owned Session. | -| Managed Runtime identity | `runtime_allocations` | Allocation and provider key identify the compute incarnation. | -| Self-hosted Runtime identity | Environment connection generation | Future telemetry must be generation-fenced. | -| Container/pod resource values | Selected provider source | Read-only, point-in-time evidence. | -| Turn state and busy duration | Core Turns | Never inferred from CPU. | -| Token usage | Existing Session/Turn usage | Missing native usage remains unknown. | -| Historical resource series | Optional telemetry backend | Not execution or lifecycle authority. | -| Idle shutdown decision | Future durable Core control state | Separate design and migration. | - -## 5. Identity chain - -```text -managed -tenant_id -> session_id -> environment_id -> runtime_allocation_id - -> provider_key -> provider-owned container/pod/instance - -self-hosted (future) -tenant_id -> session_id -> environment_id - -> device_id + connection_generation -> authenticated Runtime report - -none -tenant_id -> session_id - -> no Session-owned Runtime instance -``` - -Provider-native identifiers are never accepted from browser input. The resolver -starts from the authorized tenant and Session, loads the committed Environment and -allocation, and only then selects the configured source by persisted provider key. -The provider independently verifies its labels or equivalent ownership metadata. - -## 6. Component architecture - -```mermaid -flowchart LR - Web[Core Web Dashboard] --> Client[packages/agents-client] - Client --> API[Agents API read handlers] - API --> Service[runtimeobs.Service] - Service --> Resolver[durable identity resolver] - Resolver --> DB[(Core PostgreSQL)] - Service --> Registry[provider source registry] - Registry --> Docker[Docker Inspect and one-shot Stats] - Registry --> Micro[microsandbox exact-compute one-shot Metrics] - Registry -. future .-> K8s[Kubernetes Metrics API or cAdvisor] - Registry -. future .-> E2B[E2B metrics adapter] - Registry -. future .-> Self[authenticated daemon telemetry] - Service -. optional export .-> Telemetry[OTLP or Prometheus pipeline] - Telemetry -. future history reads .-> History[operator history adapter] -``` - -### 6.1 `runtimeobs` - -Owns provider-neutral identity, mode resolution, source selection, sample -validation, and observation status. It must not import provider SDKs or mutate -Runtime lifecycle. - -### 6.2 Provider sources - -Each source receives a fully resolved target and returns one normalized sample. -A source must verify target ownership, make only bounded read calls, preserve -missing fields, return cumulative CPU seconds, and never create, renew, restart, -pause, or stop compute. - -Docker uses Inspect followed by non-streaming one-shot Stats. Kubernetes should -retain pod UID, container identity, and restart boundaries. E2B must use an -API-supported instance identity rather than display names. Self-hosted metrics -require authenticated daemon messages fenced by the current connection generation. - -Microsandbox uses the persisted allocation's opaque compute receipt to select the -exact current generation. The pure-Go Core adapter sends a read-only request to the -existing one-shot Linux helper; the helper verifies allocation labels, compute ID, -generation and snapshot provenance before calling `SandboxHandle.Metrics`. It maps -`VCPUTimeNs`, memory usage/limit and uptime into the common sample. Instantaneous -`CPUPercent`, host RSS, disk/network counters and overlay usage remain unprojected. - -### 6.3 API composition - -The API resolves durable rows first, samples sources with bounded concurrency, -maps only ordinary absence/timeouts to safe unavailable reasons, and fails closed -on ownership or integrity errors. It returns current observations only. - -### 6.4 Web composition - -Web loads the complete paginated Runtime observation collection before publishing -a new Dashboard snapshot. The collection follows the same Session creation-time -and ID keyset as the Session list, so a Runtime incarnation change cannot invalidate -pagination. It separately uses existing Session/Turn reads for status and tokens -and joins only by exact Session identity. Before publication, the set of Session -IDs from both complete traversals must be identical. Concurrent Session creation -or deletion can make the sets differ because the APIs have no shared snapshot -token; Web then discards the candidate, marks the refresh incomplete, and keeps -the previous successful snapshot visibly stale. - -The browser enforces a configured refresh budget for total pages, targets, and -elapsed time. Exhausting that budget is an incomplete refresh: Web retains the -previous complete snapshot and does not relabel partial aggregates as tenant-wide. - -## 7. Normalized sample - -```go -type Sample struct { - ObservedAt time.Time - StartedAt *time.Time - - CPUUsageSecondsTotal *float64 - CPUCapacityCores *float64 - MemoryUsageBytes *uint64 - MemoryLimitBytes *uint64 -} -``` - -Pointer presence is semantic. `0` means observed zero; `nil` means unavailable. -CPU percentage is derived from the delta between two cumulative samples and their -observation times. A single sample cannot truthfully supply CPU percentage. - -The API projection may additionally expose `usage_cores` and `utilization_ratio` -only when the service has two ordered samples for the same Runtime allocation. -A future process-local observation cache may keep the previous cumulative value -for this calculation. Phase 2 intentionally leaves both derived fields null -because it has only one provider sample per request. The browser-local live window -therefore derives interval utilization from adjacent cumulative samples only when -Session, allocation, and provider observation order still match. Counter regression, -missing capacity, or cache loss creates a gap; none changes the cumulative source -measurement or lifecycle state. - -## 8. Duration semantics - -| UI label | Calculation | Meaning | -| --- | --- | --- | -| Allocation age | allocation `created_at` to `released_at` or now | Age of Core's allocation record. | -| Compute uptime | provider `started_at` to sample `observed_at` | Age of the current compute incarnation. | -| Active sandboxes | distinct active allocations in the selected snapshot or bucket | Count across Sessions on the Dashboard; naturally 0/1 in a single-Session view. | -| Busy duration | Turn `started_at` to `completed_at` or now | Time model work has been active. | -| Idle duration | future durable `idle_since` | Not available in the current design. | - -Container restart resets compute uptime but not allocation age. Live CPU deltas -require the same known compute start as well as the same allocation. Trend charts -show CPU, memory, active Sandbox count, and tokens. Live active count uses the -provider-neutral lifecycle state and deduplicates allocation identities; retained -history counts observed allocation identities in each bucket because lifecycle -state is not retained yet. A single-Session view therefore remains binary while -the Dashboard shows the sum across Sessions. Compute uptime stays in current -target details because the history contract does not supply each bucket's compute -start. Dashboard labels must not collapse these values into one generic Runtime -duration. - -## 9. Collection behavior - -### 9.1 Current snapshot path - -- List one current target context for each tenant-owned Session in the same stable - Session creation-time and ID order used by the Session list. A released managed - allocation remains attributable but reports `runtime_not_running`; `none` and - unsupported `self_hosted` remain explicit rows rather than disappearing. -- Default page size 20, maximum 100. -- Sample at most eight providers concurrently. -- Default per-source budget two seconds; one batch read of up to 100 sandboxes (E2B) gets at least five seconds. Whole-request budget ten seconds. -- Do not retry a source call inside the HTTP request. -- An optional process-local singleflight/cache may coalesce identical reads for up - to five seconds and retain the previous cumulative sample for CPU-rate - calculation. It is an optimization only and may be lost on restart. -- Do not write samples to the Core database. - -### 9.2 Error classification - -| Condition | API result | -| --- | --- | -| Mode `none` or unsupported `self_hosted` | Row status `unsupported`. | -| Managed allocation not created yet | `unavailable`, reason `allocation_pending`. | -| Owned Runtime absent or stopped | `unavailable`, reason `runtime_not_running`. | -| Source not configured | `unavailable`, reason `source_not_configured`. | -| Source deadline | `unavailable`, reason `sample_timeout`. | -| Ownership mismatch or invalid durable identity | Fail the request and log a sanitized integrity error. | -| Database/authentication failure | Existing safe API error mapping. | - -Raw Docker, Kubernetes, E2B, daemon, host, credential, or network diagnostics are -never returned to the browser. - -## 10. Historical metrics - -Current API reads and history are separate capabilities. The initial API does not -provide charts over time. A later operator-configured adapter may query an -OTLP/Prometheus-compatible backend. Core must not make that backend mandatory for -Session execution or current snapshot reads. - -Recommended instruments are: - -- `agents.runtime.cpu.usage` cumulative seconds; -- `agents.runtime.cpu.capacity` cores; -- `agents.runtime.memory.usage` bytes; -- `agents.runtime.memory.limit` bytes; -- `agents.session.tokens.input` cumulative measured tokens; -- `agents.session.tokens.output` cumulative measured tokens; -- `agents.runtime.sample` success/unavailable count; and -- `agents.runtime.sample.duration` seconds. - -Provider type, Runtime mode, and coarse status are safe low-cardinality labels. -High-cardinality identities require tenant-scoped access and retention policies; -they are not global Prometheus labels by default. - -### 10.1 Lightweight deployment - -Core reuses its PostgreSQL database for bounded recent Runtime history. The -provider-neutral observation service hands each sanitized periodic result to a -bounded asynchronous writer. One typed row contains the observation and optional -measured Session usage snapshot (recorded root Turn snapshots, active Turns -included, not the public Session usage rule). The `/core/v1` read resolves the Project's -Session before issuing bounded queries; Web never queries storage directly. - -External OTLP export remains optional. Each destination has an independent queue, -so a Collector outage cannot delay local persistence. Neither history nor export -is execution or lifecycle authority. No additional metrics service is deployed. - -### 10.2 OTLP transport configuration and instruments - -Core enables external export only when `OAC_HISTORY_SETTINGS_FILE` contains -an OTLP endpoint. With the variable unset, local history and 30-second sampling -remain enabled. An optional server-only configuration is: - -```json -{ - "transport": "otlp_http", - "endpoint": "https://collector.example.com/v1/metrics", - "headers": {"Authorization": "Bearer operator-managed-secret"}, - "queue_capacity": 256, - "timeout_seconds": 2, - "sample_interval_seconds": 30 -} -``` - -The file may contain transport credentials and must never be served to Web or -committed. Plain HTTP requires the explicit combination of an `http` endpoint -and `"insecure": true`; HTTPS rejects that flag. Endpoint userinfo, query -strings, fragments, invalid headers, reserved transport headers, queues above -4096 records, and timeouts above 30 seconds fail startup without echoing config -contents. `sample_interval_seconds` accepts 5 through 300; omission uses 30. -Periodic sampling requires the execution Worker and its existing database lease. -API-only processes without that worker advertise on-read collection. External -export is optional and best effort; PostgreSQL writes use a separate bounded queue. - -The OTLP request uses standard protobuf metrics and these instruments: - -| Instrument | OTLP aggregation | Source | -| --- | --- | --- | -| `agents.runtime.cpu.usage` | monotonic cumulative sum, seconds | provider cumulative CPU counter | -| `agents.runtime.cpu.capacity` | gauge, cores | configured provider capacity | -| `agents.runtime.cpu.utilization` | gauge, ratio | provider-reported share of capacity (E2B), when there is no cumulative counter | -| `agents.runtime.memory.usage` | gauge, bytes | provider memory usage | -| `agents.runtime.memory.limit` | gauge, bytes | configured provider limit | -| `agents.session.tokens.input` | gauge, tokens | measured cumulative Session usage (`MeasuredSessionUsage`) | -| `agents.session.tokens.output` | gauge, tokens | measured cumulative Session usage (`MeasuredSessionUsage`) | -| `agents.runtime.sample` | monotonic delta sum | one validated result, including unavailable/unsupported | -| `agents.runtime.sample.duration` | delta histogram, seconds | bounded provider read duration; a batch read is recorded once | - -Core-owned tenant, Session, Environment, allocation, mode, provider type, -status, safe reason, collection source (`on_read` or `periodic`), and Core -resolved/observed timestamps are metric attributes. The explicit nanosecond -timestamps preserve the record join key when a backend's generic OTLP tables -store metric event time at lower precision. CPU, capacity, and memory points are -exported only when the sample also carries a provider start estimate; it is -included for compatible display but is not series identity. Provider keys, -provider receipts, native container/pod/instance -identifiers, raw errors, paths, and credentials are not attributes. Missing -measurements produce no value point; they are represented only by the explicit -sample status and reason. - -When periodic sampling is enabled, the execution-owner service performs one -immediate, non-overlapping full keyset scan and repeats it after the configured -interval. The read-only scan covers nondeleted managed Sessions across tenants, -uses bounded pages and provider concurrency, gives each source an independent -deadline, and reuses the same resolver and observation service as current reads. -It never keeps, wakes, pauses, stops, or otherwise mutates compute. Failed rows -do not prevent later rows from being attempted, and a failed sweep is retried on -the next interval. The sampler monitors the execution lease during a sweep, -cancels in-flight provider reads on detected ownership loss, and rechecks the -lease before every periodic export handoff. `on_read` remains distinct from -`periodic`, so ad hoc API -traffic cannot be counted as qualified cadence coverage. - -A history API must expose actual sample coverage. Core Web must not -advertise a durable range until the operator backend, query adapter, and a -qualified periodic collection cadence are all configured. - -The PostgreSQL Reader shares Core database access and SQLC ownership. It limits -queries to seven-day retention, a 24-hour range, 1,000 buckets per series, 64 series -and 10,000 output points. Raw input has a separate bounded budget, so dense sampling -does not consume the output budget before downsampling. A bounded periodic cleanup -removes expired rows even when no Runtime is active. Expired rows are excluded from -reads immediately; physical removal is incremental. - -### 10.3 Backend-neutral history query boundary - -`services/core/internal/runtimehistory` defines the server-side query -contract independently from SQL, OTLP, and the public HTTP shape. Its -service resolves the authenticated tenant and Session to durable Core identity -before calling a Reader. Reader queries always carry tenant, Session, and -Environment scope plus a bounded start, exclusive end, server-selected step, -and total point budget. Provider-native identity is never a query input. - -Reader results remain divided by allocation. Provider `started_at` values are -retained as metadata and never split one durable allocation -into multiple Dashboard series. -Every bucket reports explicit observation coverage and nullable CPU/memory -values. CPU utilization may be derived only from ordered cumulative counters -inside one fence; successive intervals are assigned to the bucket containing -their right endpoint and combined by CPU-capacity time. Memory uses the final -observed value in the bucket. Dashboard memory totals aggregate only allocations with -a complete observed usage/limit pair in that bucket; an unavailable or released -target does not erase measurements from active targets. Empty -buckets remain gaps. The service rejects cross-scope rows, duplicate series, -overlapping or out-of-range buckets, unsafe provider labels, invalid numeric -values, and results exceeding the total point budget. - -Capabilities contain only safe backend-neutral limits: collection mode, -qualified sample interval, retention, minimum step, maximum range, point budget, -and supported metrics. A configured Reader without qualified periodic sampling -is not sufficient to advertise a Durable Dashboard source. Backend identity, -URLs, credentials, and tenant data are never capability fields. - -This internal boundary, the Session-scoped public extension, capability discovery, -strict client, and PostgreSQL Reader are implemented. The Reader includes tenant, -Session, Environment and bounded-time predicates. Only periodic records are stored. -It aggregates resource points by allocation, resetting CPU derivation across compute -incarnations, missing counters and cumulative-counter regressions. Durable -Web ranges remain gated on real retention, isolation, restart, and incarnation -acceptance. - -## 11. Dashboard information architecture - -### 11.1 Overview - -- Active managed Runtime count. -- Observed CPU usage and known configured capacity. -- Observed memory usage and known limits. -- Reported Session tokens, together with the reporting Session count. -- Confirmed active Sandbox count over time. Each bucket counts managed allocations - with an observed provider sample; unavailable or timed-out samples are not - presented as confirmed active. A bucket with collection coverage but no observed - allocation is zero. The history model retains missing coverage as null; the - current Dashboard presentation renders that null as zero until sleeping and - collection-failure history are represented separately. -- Data freshness and source coverage. - -Aggregates include only present measurements. Each total states its denominator, -for example, `6.4 / 12 cores across 6 of 8 active Runtimes`. Unknown is never added -as zero. - -### 11.2 Runtime table - -Each row shows Session, Agent/harness when already available from the Session -snapshot, mode, observation status, cumulative CPU time, memory, compute uptime, -Session state, and reported tokens. The semantic table supports local search, -status and mode filters, sortable columns, and bounded pagination over the last -complete snapshot. Rows navigate to the existing Session view. No stop, restart, -pause, or delete actions appear in the first release. - -### 11.3 Detail view - -The detail surface shows exact Session/Environment/allocation identity, provider -type, observation timestamps, allocation age, compute uptime, and safe unavailable -reason. It displays only the Core-owned identifiers explicitly present in the -public contract. Provider-native container IDs, pod names, instance names, host -paths, and raw labels are never displayed. - -### 11.4 States - -- **Loading:** no previous complete Runtime snapshot. -- **Fresh:** every page loaded and each row carries its own resolution time; a - provider sample also carries its independent observation time. -- **Stale:** refresh failed; previous complete snapshot retained. -- **Unavailable row:** identity is valid, measurement is temporarily absent. -- **Unsupported row:** mode is recognized but has no qualified source. -- **Integrity failure:** do not publish a partial replacement snapshot. - -### 11.5 Web implementation shape - -The Web implementation belongs in Core Web, not the Core service layer. It uses -`packages/agents-client` as the only Runtime-observation transport and keeps four -seams separate: - -1. A client/parser module validates one page and exposes list and Session-scoped - retrieval methods. -2. A refresh coordinator loads all observation pages plus the canonical Session - collection, applies page/target/time budgets, requires exact equality of their - Session ID sets, and atomically swaps only a complete joined snapshot. A set - mismatch is an incomplete refresh, not a partial success. -3. A feature-local state model retains `last_complete`, current refresh status, - local filters, and the selected time range. It aborts an overlapping refresh - and marks old data stale after a failed or incomplete refresh. -4. Presentational components render summary metrics, the Runtime table, an identity - detail surface, and capability-selected trends. A qualified periodic Reader is - the primary trend source. Without one, the fallback live window contains only - complete snapshots collected while this Dashboard instance is mounted; it is - bounded, ephemeral, and never presented as durable operator history. - -The live window retains the latest complete cumulative CPU counters only long -enough to calculate the next interval. Historical chart points contain the bounded -derived series, not all Runtime identities. This makes real Docker and microsandbox -CPU charts work without moving lifecycle authority or durable history into Web. - -The initial implementation uses a 30-second Web cadence plus up to five seconds -of jitter and a 15-second whole-refresh budget. These are Web configuration, not -API guarantees. Web pauses periodic reads when hidden, refreshes when visibility -returns, and adds jitter so multiple browsers do not synchronize. Filtering is -local to the last complete snapshot and never changes tenant authorization or -provider selection. The same complete snapshots feed the one-hour, 120-sample -browser-local live window. Operators can select a 15-minute or one-hour view -without discarding the retained buffer. The range is measured back from the newest -complete snapshot rather than browser wall-clock time. Reload, navigation, or -connection replacement may reset the window, and no point is interpolated or -persisted by Core. - -## 12. Token usage boundary - -Provider Runtime sources do not own or report token usage. During a periodic -history sweep, the Core resolver reads the cumulative measured Session usage -(`MeasuredSessionUsage`: every recorded root Turn snapshot, active Turns -included) from the execution store alongside Runtime identity. Public Session -usage follows the stricter official rule and can be null meanwhile -([usage](sessions-events.md#usage)). The exporter -emits Session-scoped input/output token gauges with the same Session and sampling -time, independently of Docker, microsandbox, Kubernetes, or another provider. -PostgreSQL retains those cumulative snapshots alongside the sample; query results -keep Session token points separate from allocation series. Web derives throughput from adjacent nondecreasing points. Missing or -incomplete native usage, counter regressions and intervals where the live trend -holds a Session's last public total remain gaps, never zero. - -The current snapshot API still does not duplicate Usage fields: Web joins its -existing Session collection by exact `session_id`. The Dashboard reports measured -coverage and never estimates missing usage. - -Cost and billing stay outside this Core API. A product may join billing in its own -authorized backend, never by exposing product credentials to Core Web. - -## 13. Security and tenancy - -- Authenticate with the existing Agents API mechanism. -- Scope resolution to the authenticated tenant before provider access. -- Do not accept provider key, allocation ID, container ID, pod UID, or device ID - as an authority-bearing query parameter. -- Bound per-request list size, concurrency, response bytes, and source deadlines; - Web separately bounds a complete multi-page refresh. -- Sanitize logs through `internal/obs/log`. -- Never return credentials, environment variables, Docker raw JSON, daemon status - payloads, host paths, image registry credentials, or backend credentials. -- Rate-limit collection separately from ordinary Session reads. - -## 14. Data model impact - -Current snapshots reuse existing control-plane resources. Retained history adds -one bounded observation table to the Core database. It does not change canonical -Sessions, Turns or Usage and cannot authorize execution or lifecycle changes. - -Automatic idle shutdown is a separate feature. It requires durable fields such as -`activity_revision`, `idle_since`, and `shutdown_requested_at` with fenced state -transitions. That migration cannot read a monitoring backend as authority. - -## 15. Delivery plan - -### Phase 1: provider-neutral foundation - -Implemented for Docker and microsandbox. - -- `runtimeobs` identity, resolver, source, sample, and service. -- Managed Docker Inspect/Stats source. -- Managed microsandbox exact-compute point-in-time Metrics source. -- CPU, memory, and current compute start time. -- Explicit unsupported and unavailable states. - -### Phase 2: current snapshot API - -Implemented by the Runtime Observation extension routes and -`packages/agents-client`. The generated OpenAPI contract records the extension; -this does not add an upstream OpenAI operation. - -- Add extension types under `contracts/agents-api/v1`. -- Add collection and Session-scoped handlers. -- Add `packages/agents-client` methods and raw HTTP/client coverage. -- Add bounded concurrency, timeout, authorization, and error tests. -- Regenerate the public OpenAPI contract. - -### Phase 3: Web Dashboard - -Implemented for the browser-local current-snapshot live window. - -- Add Runtime observations as a third independent Dashboard collection. -- Publish only complete traversals and retain the previous snapshot on failure. -- Join existing Session Usage and Turn status by exact Session ID. -- Add responsive, keyboard-accessible current-resource views. -- Build an explicitly ephemeral live window from complete Web snapshots. -- Expose 15-minute and one-hour views with an explicit browser-local source label. - -### Phase 4: retained history - -- Bounded asynchronous writes of sanitized periodic observations to Core PostgreSQL. -- Optional OTLP/HTTP export with server-only credentials and a separate queue. -- Execution-owner sampling with bounded pages, concurrency and source deadlines. -- Tenant-scoped history queries with input/output limits, explicit gaps and CPU fences. -- Seven-day retention and bounded cleanup independent of active Runtime count. -- Core Web history restoration after refresh, including measured token snapshots. -- Real Docker/microsandbox and PostgreSQL acceptance remains mandatory for delivery. -- Exporter queue/drop/error coverage counters remain outside this batch. - -### Phase 5: additional sources - -- Kubernetes, E2B, and generation-fenced self-hosted telemetry. -- Each source requires independent mechanism and deployment acceptance. - -### Phase 6: idle policy - -- Separate durable activity and shutdown state machine. -- No automatic action until race, fencing, recovery, and operator-control - acceptance is complete. - -## 16. Acceptance criteria - -- Every observation proves tenant, Session, Environment, and Runtime-instance - association. -- Managed Docker emits correct present/absent semantics and never mutates compute. -- Managed microsandbox verifies the exact current generation and never mutates, - resumes, pauses, snapshots, stops, or removes compute while observing it. -- Unsupported modes never look like zero usage. -- Collection calls are bounded and one ordinary unavailable source invents no data. -- Ownership/integrity mismatch fails closed. -- Web never publishes a partial page traversal as a current snapshot. -- Web rejects cross-collection Session membership skew, including concurrent - Session create/delete cases, before publishing tenant-wide aggregates. -- Token totals report coverage and do not estimate missing usage. -- No lifecycle action is reachable from the first Dashboard. -- No new database table is required for current snapshots or history export. - -## 17. Recorded design decisions - -1. The API is a documented Core extension rather than an upstream OpenAI resource. -2. Current snapshots have no atomic cross-row time semantics; every row - exposes its own `observed_at`. -3. History is optional and external, not a PostgreSQL sample table. -4. The first Web release has no lifecycle controls. -5. `self_hosted` remains visibly unsupported until authenticated, - generation-fenced telemetry is qualified. diff --git a/contracts/agents-api/runtime-observability.md b/contracts/agents-api/runtime-observability.md index cc1a4cb19..078bccd2d 100644 --- a/contracts/agents-api/runtime-observability.md +++ b/contracts/agents-api/runtime-observability.md @@ -1,11 +1,12 @@ -# Runtime observability contract +# Runtime observability -This document defines the internal Runtime observation boundary. It does not add -an Agents API resource or change the pinned public protocol. +This is the contributor contract for how Core observes Runtimes and keeps their history. The routes and response fields are in the [Runtime telemetry API](runtime-observability-api.md). The code lives in `services/core/internal/runtimeobs` (resolution, sources, sampler and export), `internal/runtimehistory` (history queries and the PostgreSQL store) and `internal/runtimeobs/otlpexporter`. -## Ownership and identity +Observations are telemetry. They never create, renew, wake, restore or stop compute, never touch Session activity, and never decide idle time, suspension, admission or execution outcomes. -Runtime telemetry is attributed to durable Core identity before it is sampled: +## Identity and source selection + +Core attributes every observation to durable Core identity before it reads a provider: ```text managed: tenant_id -> session_id -> environment_id -> runtime_allocation_id @@ -13,180 +14,111 @@ self-hosted: tenant_id -> session_id -> environment_id -> device_id + connection none: tenant_id -> session_id (no Session-owned Runtime instance) ``` -The managed allocation's persisted `provider_key` selects exactly one configured -observation source. A provider must independently verify the allocation labels or -equivalent ownership data. A Session, daemon connection, process, container, and -native harness Session are different identities and must not be substituted for -one another. +The resolver (`internal/runtimeobs/storeresolver`) reads the Session, its Environment, the current allocation and the Session's measured usage from the store. A Session, daemon connection, process, container and native Harness Session are different identities and never stand in for one another. + +Managed Docker, microsandbox and E2B allocations are observed. `none` and `self_hosted` Sessions are `unsupported`; Core never attributes shared host statistics to an `environment:none` Session. + +The allocation's persisted `provider_key` selects exactly one configured source, which verifies the allocation's labels or equivalent ownership data before it returns values. Before any provider read, the allocation state decides some rows: `creating` or no allocation yet gives `allocation_pending`, `cleanup_pending` or `released` gives `runtime_not_running`, and a provider key without a source gives `source_not_configured`. A provider read that exceeds its deadline gives `sample_timeout`, a not-running result `runtime_not_running`, and an unavailable result `sample_unavailable`. Any other error, an ownership mismatch or an invalid sample fails the read. -The current implementation supports managed Docker, microsandbox and E2B allocations. -`self_hosted` and `none` are recognized but explicitly unsupported. A future self-hosted source must -use authenticated daemon telemetry fenced by the current connection generation. -Core must not attribute shared host statistics to an `environment:none` Session. +A source implements `Observe` and declares `ObserveBatch` in its provider operations. When `ObserveBatch` is declared supported, one call reads up to 100 targets of that provider; when it is declared unsupported, Core reads each target with `Observe`. A failed batch read is never retried target by target. The [Sandbox Provider guide](../../docs/sandbox-provider.md) describes the operation declarations. ## Sample semantics -One sample contains: - -- `observed_at`, the provider observation time; -- `started_at`, the current compute incarnation start time; -- cumulative CPU usage in seconds; -- configured CPU capacity in cores, when known; -- current memory usage in bytes; -- configured memory limit in bytes, when known; -- a provider-reported CPU utilization ratio, only for providers without - cumulative CPU time (E2B); and -- current disk usage and capacity in bytes, only where the provider reports - them (E2B). Only the administrator list exposes disk. - -Measurements are optional. A present pointer with value zero means the provider -observed zero. An absent measurement means it was unavailable and must never be -rendered or aggregated as zero. A whole observation has one of three states: -`observed`, `unsupported`, or `unavailable`. Provider and permission failures are -errors, not ordinary unavailability. - -Managed observations also expose a provider-neutral `lifecycle_state` derived -from Core's allocation and compute lifecycle: `active`, `sleeping`, -`transitioning`, `pending`, or `stopped`. Non-managed modes return `null`. -This field is current control-plane state; it is not inferred from a failed -provider sample. - -Docker reports cumulative cgroup CPU time and current cgroup memory usage. CPU and -memory capacity come from the inspected container configuration. Inspect and Stats -are read-only; observation must not renew, restart, create, or stop the container. -The Docker `StartedAt` value defines current compute uptime and resets after a -container restart. - -Microsandbox reports cumulative vCPU time, current guest memory usage, its effective -memory limit, and compute uptime through the pinned SDK's point-in-time metrics -operation. Core invokes that SDK only through the existing one-shot Linux helper. -The helper first verifies the allocation labels and exact persisted compute -generation/ID, then reads metrics under the allocation lock. Restored generations -therefore reset compute uptime without resetting allocation age. Paused, stopped, -suspended, metrics-disabled, and no-current-sample states are unavailable, never -observed zero. The SDK also supplies instantaneous CPU percent, host RSS, disk, -network, and overlay values; those are intentionally outside this public sample -until their cross-provider semantics and API fields are designed. -Legacy suspension-disabled allocations without a persisted exact compute receipt -are also unavailable. A deterministic sandbox name is not an incarnation identity -and is never used as a sampling fallback. - -E2B reports only the latest point of a current CPU percentage, memory and disk, -through `GET /sandboxes/metrics?sandbox_ids=...`. One helper request reads a whole -page: at most 100 allocations, one metrics request and, concurrently, one listing -of this installation's running sandboxes by their allocation labels. The helper -takes each sandbox ID from its private receipt without the allocation lock; the -listing confirms that exactly that sandbox is running with the allocation's labels -and supplies its `started_at`; it stops paging once every requested sandbox has -been listed, so duplicate-label detection covers only the pages read. A listed -sandbox without a metrics point or with a malformed point, an ambiguous listing -or an E2B API failure (including a rejected key) is unavailable; a malformed -point affects only its own row. A sandbox absent from the running listing is -`runtime_not_running`. Observation -never connects to, renews or changes a sandbox and never writes receipts. - -The E2B mapping is: `cpuUsedPct / 100` to `cpu.utilization_ratio`, `cpuCount` to -`cpu.capacity_cores`, `memUsed` and `memTotal` to `memory.usage_bytes` and -`memory.limit_bytes`, and `diskUsed` and `diskTotal` to the administrator -`disk.usage_bytes` and `disk.limit_bytes`. E2B has no cumulative CPU time, so -`usage_seconds_total` and `usage_cores` stay null. A template whose envd predates -E2B disk metrics reports no disk capacity; disk is then null. `observed_at` is the -point's E2B timestamp; a point up to 30 seconds ahead of Core's clock is recorded -at Core's time, and a larger lead is unavailable. Docker disk is null; -microsandbox disk is null until its disk semantics are designed. - -## Duration boundaries - -These durations answer different questions and must remain separate: - -- allocation age: `runtime_allocations.created_at` through `released_at` or now; -- compute uptime: provider `started_at` through `observed_at`; and -- busy Turn duration: `turns.started_at` through `completed_at` or now. - -This phase supplies compute uptime evidence and retains the existing durable -allocation and Turn timestamps. It does not infer idle time. CPU quietness, -heartbeat age, connection status, and `kept_at` are not authoritative idle state. - -Web projects active Runtime state differently by scope. The Dashboard shows one -summed series of distinct allocation identities: live snapshots count -`lifecycle_state: active`, while retained buckets count successfully observed -allocations because lifecycle state is not retained yet. The single-Session view -collapses the same value to `1` or `0`. Missing or unavailable retained values are -currently rendered as zero, so this presentation intentionally does not yet -distinguish sleeping from collection failure. - -Future automatic suspension requires a separate durable control model, including -an activity revision and timestamps such as `idle_since` and -`shutdown_requested_at`. Metrics, an in-memory cache, or a monitoring backend must -not become the lifecycle authority. +A sample carries: + +- `observed_at`, the provider's observation time, and `started_at`, the start of the current compute incarnation; +- cumulative CPU seconds and the configured CPU capacity in cores; +- a provider-reported CPU utilization ratio, only from providers without cumulative CPU time (E2B); +- current memory usage and the memory limit in bytes; +- current disk usage and capacity in bytes, only where the provider reports them (E2B). + +Every measurement is optional. A present zero is an observed zero; an absent value is unavailable and is never shown or aggregated as zero. Core rejects a sample whose `observed_at` is later than its own clock, whose `started_at` is later than `observed_at`, or whose values are negative, non-finite, a zero capacity or limit, or beyond the JSON safe-integer range. + +`lifecycle_state` is derived from the allocation state and compute phase in Core's records, never from a sample: no allocation or `creating` is `pending`; `running` with compute phase `running` or `disabled` is `active`, `suspended` is `sleeping`, and `quiescing`, `suspending`, `restoring` or `waking` is `transitioning`; `cleanup_pending` and `released` are `stopped`. + +## Provider mapping + +### Docker + +One non-streaming Inspect and Stats read of the owned container. Cumulative CPU time comes from the cgroup counter, memory usage from the current cgroup usage, and CPU and memory capacity from the container's configured limits. The container's `StartedAt` is the incarnation start, so a container restart resets compute uptime. A missing or stopped container is `runtime_not_running`. Disk is null. + +### microsandbox + +A suspended allocation is `runtime_not_running` without calling the helper. Otherwise Core sends the helper a `metrics` request for the exact compute recorded on the allocation. Under the allocation lock, the helper checks the sandbox's identity through the pinned SDK, runs the pinned `msb metrics --format json` CLI read, and checks the identity again; the sandbox must be running or draining. The [microsandbox helper](../../services/core/tools/microsandbox-provider/README.md) owns that request. + +Cumulative vCPU time, guest memory usage and the effective memory limit come from that one CLI sample. CPU capacity is the deployment's configured CPU count. `started_at` is the sample's own timestamp minus its millisecond uptime, never rounded uptime subtracted from a later clock reading, so a restored generation restarts compute uptime while the allocation age continues. Instantaneous CPU percent, host memory, disk and network values are not used, and disk is null. A suspension-disabled allocation without a recorded compute identity is `sample_unavailable`, and any other allocation without one fails the read; a deterministic sandbox name is never used instead. + +### E2B + +One helper `observe` request reads a page of at most 100 allocations: E2B's batch metrics for the sandboxes named in the private receipts, and a labelled listing of the installation's running sandboxes that confirms each one. The [E2B helper](../../services/core/tools/e2b-provider/README.md) owns that request. It never connects to, renews or changes a sandbox and writes no receipts. + +| E2B value | Sample field | +| --- | --- | +| `cpuUsedPct / 100` | CPU utilization ratio | +| `cpuCount` | CPU capacity | +| `memUsed`, `memTotal` | Memory usage and limit | +| `diskUsed`, `diskTotal` | Disk usage and capacity, kept only when both are present and the total is nonzero | + +E2B reports no cumulative CPU time, so CPU seconds stay null. `observed_at` is E2B's point time; a point up to 30 seconds ahead of Core's clock is recorded at Core's time, and a larger lead is `sample_unavailable`. A sandbox missing from the running listing is `runtime_not_running`. A missing or malformed point, an ambiguous listing and an E2B API failure, a rejected key included, are `sample_unavailable`; a malformed point affects only its own row. + +## Read budgets + +A current list read handles one page of up to 100 Sessions (default 20) with at most eight concurrent provider reads. Each provider read has two seconds, a batch read at least five, and the whole list request ten; beyond that the list returns 503. A single-Session read has two seconds. No provider call is retried within a request, and Core keeps no observation cache. + +## Durations + +These durations answer different questions and stay separate: + +- allocation age: `runtime_allocations.created_at` to `released_at`, or now; +- compute uptime: the sample's `started_at` to `observed_at`; +- busy Turn duration: `turns.started_at` to `completed_at`, or now. + +CPU quietness, heartbeat age, connection state and keepalive time are not idle time. ## Retained history and optional export -Periodic Runtime observations are persisted asynchronously in the existing Core -PostgreSQL database. The execution owner samples every 30 seconds by default, -using bounded pages, concurrency and source deadlines. Collection never wakes or -mutates compute. The worker lease is checked during the sweep and before each -handoff. Only periodic samples populate durable history; current API reads cannot -inflate cadence coverage. Retention is seven days; history reads span at most -24 hours and have explicit input and output limits. - -`OAC_HISTORY_SETTINGS_FILE` optionally changes sampling and adds OTLP/HTTP -export. The local database and external exporter have independent bounded queues. -No Collector is required for the Dashboard. Failures and queue saturation remain -missing observations rather than fabricated zeroes or failed executions. Transport -credentials stay server-only; native identifiers, receipts, raw errors, paths and -credentials are excluded from observations. - -The internal `runtimehistory` boundary validates Core scope, bucket coverage, -nullability, time bounds and point limits. One chart series represents an allocation; -CPU deltas reset across compute incarnations or counter regressions. E2B samples -store their reported utilization ratio instead, and a bucket holds the mean of the -ratios sampled in it. Disk is not retained in history. Core's -measured Session usage (recorded root Turn snapshots, active Turns included) -supplies independently sampled token counters. History queries survive -Core restart and browser reload without replaying execution. - -See the [design](runtime-observability-design.md), -[current API](runtime-observability-api.md), [history API](runtime-history-api.md) -and [configuration](../../docs/configuration.md#settings). -Additional provider telemetry and idle-policy authority remain separate work. - -The OTLP resource identifies Core with `service.name=oac-core` and -`service.namespace=oac`. Metric names retain the `agents.*` namespace. - -## Implementation rules - -Runtime telemetry uses a separate read-only service boundary documented in -[`contracts/agents-api/runtime-observability.md`](runtime-observability.md). -Resolve durable Session, Environment and Runtime-instance identity before selecting -a provider source. Observation never extends a lease or changes compute lifecycle. -Keep observed zero, unavailable data and unsupported Runtime modes distinct. Metrics -may inform operators, but automatic suspension requires durable Core-owned activity -state and must not use a monitoring backend as lifecycle authority. -Managed Docker observes one non-streaming Inspect/Stats sample. Managed microsandbox -observes the exact persisted compute generation through the existing one-shot helper -and pinned native CLI metrics report, with SDK identity checks before and after -observation. Derive compute start from the same native sample timestamp and precise -uptime; never subtract rounded uptime from a new wall-clock timestamp. Preserve -cumulative CPU seconds, memory usage/limit and -compute uptime semantics across both. Do not use microsandbox's instantaneous CPU -percent, wake suspended compute, or expose provider-native identifiers to fill a -common field. E2B has no cumulative CPU time: one read-only helper request per -page of at most 100 allocations reads E2B's batch metrics and confirms each -receipt's sandbox in a labelled running listing. Its reported CPU share fills -`utilization_ratio`; disk appears only in the administrator list. Page reads go -through the Service's optional batch source; do not add a second collector. - -Runtime history uses the existing Core PostgreSQL database: one sanitized row per -periodic observation, seven-day retention and bounded reads. It is best-effort -operational evidence, not execution or Usage authority. The execution owner samples -by default every 30 seconds. Core's internal measured Session usage (every -recorded root Turn snapshot, active Turns included) supplies token snapshots, not -the public Session usage rule; never aggregate provider counters as model tokens. Preserve missing data -and reset CPU derivation across compute incarnations or counter regressions. E2B -history stores the reported utilization ratio; a bucket holds their mean. -The bounded asynchronous database writer and optional OTLP exporter have independent -queues; external telemetry outages must not stall local history or execution. -Retention cleanup also runs without active Runtimes. The browser queries only Core, -never storage or a Collector, and stays a lightweight administrator console. -No additional metrics database or Collector is required for retained charts. +### Periodic sampling + +Periodic collection runs only with the execution worker (Core started with `OAC_PUBLIC_URL`; see the [Core environment](../../docs/configuration.md#appendix-core-environment-without-the-installer)) and under the worker's database lease. A Core without it stores no history and answers every history read with 503; current reads work on either. + +The sampler sweeps once at startup and again each sampling interval after the previous sweep ends. A sweep is a keyset scan, in Session ID order, of the Sessions that are not deleted, are `openai_hosted` and have no released allocation. It reads pages of 32 Sessions through the same resolver and sources as current reads, with eight concurrent reads and two seconds per source. The sampler checks the lease before each page and every 100 ms during a sweep, cancels in-flight reads when ownership is lost, and checks it again before handing each record to export. A failed row does not stop the sweep, and an incomplete sweep is repeated at the next interval. + +Every observation, current or periodic, is marked with its collection source, `on_read` or `periodic`, and handed to each exporter's bounded queue. A full queue drops the record, which becomes a missing sample, never a zero. The PostgreSQL history store and the optional OTLP exporter have independent queues, so an exporter outage cannot delay local history or execution. The [`core.runtime_history` settings](../../docs/configuration.md#settings) set the interval, queue capacity, timeout and OTLP destination. + +### Stored history + +The PostgreSQL store keeps only periodic `openai_hosted` records, so API reads cannot inflate coverage. Each goes into one `runtime_history_samples` row keyed by tenant, Session, Environment and resolution time: the allocation, provider type, status, observation and start times, CPU seconds, capacity and utilization ratio, memory usage and limit, and the measured token counters. Disk is not stored. A sample without `started_at` keeps its coverage and drops its resource values, since they cannot be tied to an incarnation. + +The history service resolves the Project, Session and Environment before it queries; the query always carries that scope and bounded times, never provider identity. The store keeps seven days. A read covers at most 24 hours, reads at most 20,000 raw samples, starts two sampling intervals before the range to find CPU baselines, and returns at most 1,000 buckets per array, 64 series and 10,000 points in total. Results outside the requested scope, range or limits fail the read. The API's [Series](runtime-observability-api.md#series) section describes the aggregation. + +`runtimehistory.Capabilities` states the collection mode, interval, seven-day retention, minimum bucket width (30 seconds or the interval, whichever is longer), 24-hour range and point limits; the history route answers 503 unless they are valid and periodic. + +A cleanup loop runs every minute, even without active Runtimes. Each pass has at most two seconds and deletes expired Runtime and node host rows in batches of 256 per table, at most 16 batches. Reads never return rows older than the retention. + +### Node host history + +After each sweep, within two seconds and after a lease check, Core copies every fresh node heartbeat into `node_host_history_samples`: the host CPU utilization, used memory and available disk, keyed by node and the heartbeat's own time, so repeated sweeps add nothing. A heartbeat is fresh when its node is not removed, is connected under the current owner epoch, was seen within 45 seconds and reported a host time within the last 45 seconds and not in the future. The same seven-day cleanup applies. Node host history is telemetry, never scheduling or capacity truth, and reads never sample or backfill it. + +### Token usage + +Each observation carries the Session's measured usage from `MeasuredSessionUsage`: the sum of every recorded root Turn snapshot, active Turns included. It is not the public Session usage rule, and Core never derives tokens from provider counters, context occupancy or costs. + +### OTLP export + +With an OTLP endpoint configured, Core exports every record, both `on_read` and `periodic`, as OTLP/HTTP metrics. The resource has `service.name=oac-core` and `service.namespace=oac`. + +| Instrument | Aggregation | Source | +| --- | --- | --- | +| `agents.runtime.cpu.usage` | Monotonic cumulative sum, seconds | Cumulative CPU counter | +| `agents.runtime.cpu.capacity` | Gauge, cores | Configured CPU capacity | +| `agents.runtime.cpu.utilization` | Gauge, ratio | Provider-reported CPU share (E2B) | +| `agents.runtime.memory.usage` | Gauge, bytes | Memory usage | +| `agents.runtime.memory.limit` | Gauge, bytes | Memory limit | +| `agents.session.tokens.input` | Gauge, tokens | Measured Session input tokens | +| `agents.session.tokens.output` | Gauge, tokens | Measured Session output tokens | +| `agents.runtime.sample` | Monotonic delta sum | One per validated result, unavailable and unsupported included | +| `agents.runtime.sample.duration` | Delta histogram, seconds | Provider read duration; a batch read counts once | + +CPU and memory points are exported only when the sample has `started_at`; a missing measurement produces no point. Attributes are `agents.tenant.id`, `agents.session.id`, `agents.environment.id`, `agents.runtime.allocation.id`, `agents.runtime.mode`, `agents.runtime.provider.type`, `agents.runtime.status`, `agents.runtime.reason`, `agents.runtime.collection.source` and nanosecond `agents.runtime.resolved_at_unix_nano`, `agents.runtime.observed_at_unix_nano` and `agents.runtime.compute.started_at_unix_nano`. The nanosecond times keep records joinable when a backend stores event time at lower precision. Provider keys, receipts, native identifiers, raw errors, paths and credentials are never attributes. + +Web reads history only through Core; neither a Collector nor another metrics store is needed for its charts. diff --git a/contracts/agents-api/session-diagnostics.md b/contracts/agents-api/session-diagnostics.md index 9be3c76f3..8eeda576c 100644 --- a/contracts/agents-api/session-diagnostics.md +++ b/contracts/agents-api/session-diagnostics.md @@ -1,23 +1,15 @@ # Root Session and Turn diagnostics -These read-only routes require a Core key. The Project ID selects the target -space; it does not authenticate. Both return `Cache-Control: no-store`: +These read-only routes require the Core key. The Project ID selects the target Project and does not authenticate. Both return `Cache-Control: no-store`: - `GET /core/v1/projects/{project_id}/sessions/{session_id}/diagnostics` - `GET /core/v1/projects/{project_id}/sessions/{session_id}/turns/{turn_id}/diagnostics` -Missing, malformed, foreign and deleted resources use the existing Session/Turn -not-found rules. A Subagent Turn is not a root Turn and returns 404. Reads never -contact an executor, provision an Environment, repair history or change execution. +Missing, malformed, foreign and deleted resources follow the Session and Turn not-found rules. A Subagent Turn is not a root Turn and returns 404. Reads never contact an executor, provision an Environment, repair history or change execution. ## Session snapshot -The object is `core.session_diagnostics`, with `session_id`, official Session -`status`, and nullable `failure`. Each read uses one repeatable-read database -snapshot and the existing public status projection. Each Session or Turn snapshot -read has a five-second budget, shortened by any earlier caller deadline. -The transaction is closed after the read, including cancellation. Failure is null -unless that projection says `failed`. Its fields are: +The object is `core.session_diagnostics`, with `session_id`, the official Session `status`, and nullable `failure`. Each Session or Turn read uses one repeatable-read database snapshot and the public status projection, within a five-second budget that an earlier caller deadline shortens. The transaction is closed after the read, including on cancellation. `failure` is null unless the projection says `failed`. Its fields are: | Field | Value | | --- | --- | @@ -25,69 +17,28 @@ unless that projection says `failed`. Its fields are: | `turn_id` | Present only for `source: turn` | | `code` | Fixed category from [the diagnostics catalog](core-errors.md#diagnostic-failure-categories) | | `params` | An object of fixed safe values; `{}` for categories without parameters | -| `failed_at` | RFC3339 timestamp, or null when historical time is unknown | +| `failed_at` | RFC 3339 timestamp, or null when the time is unknown | -Hosted provisioning failure overrides input activity; input activity overrides -the latest root Turn. The same precedence drives the public Session response. -Private outcome text, native messages, provider bodies, command text, paths and -credentials are never included. Unknown outcome codes become `internal_error`; -no prefix matching or historical reason-text parsing is used. +A hosted provisioning failure takes precedence over input activity, and input activity over the latest root Turn; the public Session response uses the same precedence. Private outcome text, native messages, provider bodies, command text, paths and credentials are never included. Unknown outcome codes become `internal_error`; no prefix matching or reason-text parsing is used. -Provisioning parameters come from the confirmed structured receipt, persisted in -the existing failure transaction with Environment/input settlement and events. -`step` is `setup`, `python`, `npm`, `system`, `file`, `skill` or null. `index` is a -nonnegative JSON-safe integer for setup only, otherwise null. `exit_code` is 1–255 -for script steps, otherwise null. Unknown or historical details remain null. -Changing these private fields does not change the public failure reason or SSE. +Provisioning parameters come from the confirmed structured receipt, persisted in the failure transaction that settles the Environment and its input and records the events. `step` is `setup`, `python`, `npm`, `system`, `file`, `skill` or null. `index` is a nonnegative JSON-safe integer for setup only, otherwise null. `exit_code` is 1 to 255 for script steps, otherwise null. Details that were not recorded remain null. These private fields do not change the public failure reason or SSE. ## Turn snapshot and Item receipt timing -The object is `core.turn_diagnostics`, with `session_id`, `turn_id`, the official -root Turn `status`, nullable `failure`, `items` and `items_truncated`. Turn failure -has `code`, `params` and nullable `failed_at`, without a source or Turn ID. Only -failed Turns have failure details; cancelled/completed Turns never inherit an -error classification from a private outcome. - -`items` contains at most 1000 root Items, ordered by `(created_at, position, id)` -ascending, matching public Items order. Storage reads at most 1001 rows to detect -truncation. Each entry has `item_id`, `started_at`, nullable `completed_at`, and -nullable integer `observed_duration_ms`. - -- `started_at` is the first persisted Core input/event receipt. -- `completed_at` is the first terminal input/event receipt. An Item first observed - terminal settles at that same receipt (zero observed duration). -- An Item still in progress when its root Turn terminates settles with one database - `clock_timestamp()` sampled after the existing Session lock and terminal - projection. It does not use PostgreSQL transaction-start `now()` or the native - Turn completion timestamp. All such Items in that transaction share that time. -- Repeated terminal projections preserve settlement. Already-terminal historical - Items with no settlement retain null; no backfill or duration estimate is made. -- `observed_duration_ms` is the integer millisecond difference when both receipts - are known, otherwise null. It is not native execution duration. Journal batching, - transport, persistence and database wall-clock behavior affect this interval; - the event journal normally flushes about every 100 ms. Values are not clamped. - -The public successful Turn completion time may originate at the native executor; -that public timestamp remains unchanged. Public Session, Turn, Item, SSE and -`duration_ms` contracts are unchanged. Subagent diagnostics are separate work. -Native failure classification uses only -the finite top-level outcome metadata of a failed `engine_failed` Turn; see -[native classification](native-error-classification.md). Connection failure params -contain `http_status` (100–599 or null); other native categories have empty params. - -The typed client exposes `retrieveSessionDiagnostics` and -`retrieveTurnDiagnostics` on `AdminClient`, retaining request cancellation and -validating scope, categories, nullability and safe parameter values. No browser -credential or Core Web implementation is added by this contract. - -## Implementation rules - -Root Session/Turn diagnostics are Core-key reads from one committed database -snapshot, reusing public status projection and precedence. Keep their finite -failure catalog separate from private outcome text. Persist safe provisioning -details in the existing failure transaction; never reconstruct historical details -from reason strings. Item settlement belongs to the existing Session lock and -terminal transaction: event/input receipt for normal terminal projections, one -post-lock database wall-clock sample for forced incomplete Items. Preserve -historical terminal nulls and public native completion times. See the -[diagnostics contract](session-diagnostics.md). +The object is `core.turn_diagnostics`, with `session_id`, `turn_id`, the official root Turn `status`, nullable `failure`, `items` and `items_truncated`. A Turn failure has `code`, `params` and nullable `failed_at`, without a source or Turn ID. Only failed Turns have failure details; cancelled and completed Turns never inherit a classification from a private outcome. + +Native categories apply only to a failed Turn whose outcome is `engine_failed`, and come from the finite outcome metadata described in [native failure classification](../../docs/runtime-protocol.md#native-failure-classification). `connection_failed` params contain `http_status`, 100 to 599 or null; other native categories have empty params. + +`items` contains at most 1000 root Items, ordered by `(created_at, position, id)` ascending like the public Items list. Storage reads at most 1001 rows to detect truncation. Each entry has `item_id`, `started_at`, nullable `completed_at`, and nullable integer `observed_duration_ms`. + +- `started_at` is Core's first persisted input or event receipt of the Item. +- `completed_at` is the first terminal input or event receipt. An Item first observed terminal settles at that same receipt, with zero observed duration. +- An Item still in progress when its root Turn terminates settles with one database `clock_timestamp()` sampled after the Session lock and the terminal projection, shared by all such Items of that transaction. It uses neither the transaction-start `now()` nor the native Turn completion time. +- Repeated terminal projections keep the first settlement. Terminal Items stored without a settlement keep null; Core does not backfill or estimate them. +- `observed_duration_ms` is the integer millisecond difference when both receipts are known, otherwise null. It is not native execution time: journal batching (the event journal flushes about every 100 ms), transport, persistence and database clock behavior all affect it. Values are not clamped. + +The public completion time of a successful Turn may come from the native executor and is unchanged by these reads. + +## Client + +`AdminClient.retrieveSessionDiagnostics(projectId, sessionId, options)` and `AdminClient.retrieveTurnDiagnostics(projectId, sessionId, turnId, options)` in `packages/agents-client` support request cancellation and validate scope, categories, nullability and safe parameter values. From e694a227fbfd5e437e28047a70ddd3422cab0a0f Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 08:35:14 +0000 Subject: [PATCH 02/10] docs: reduce deploy and helper READMEs to adapter rules Each Harness page holds its native pin, adapter rules, failure classification and Runtime image contents. The E2B deploy page covers the template build and the application-managed launch; the E2B and microsandbox helper READMEs are reference Provider adapter docs. The microsandbox deploy page is merged into its helper README. Operator copies, stale model-provider routes, acceptance history and restated contracts are removed. --- services/core/deploy/claude/README.md | 141 +++------ services/core/deploy/codex/README.md | 211 ++++--------- services/core/deploy/e2b/README.md | 180 +++-------- services/core/deploy/mcode/README.md | 264 +++------------- services/core/deploy/microsandbox/README.md | 173 ---------- services/core/tools/e2b-provider/README.md | 199 ++++-------- .../tools/microsandbox-provider/README.md | 298 ++++-------------- 7 files changed, 335 insertions(+), 1131 deletions(-) delete mode 100644 services/core/deploy/microsandbox/README.md diff --git a/services/core/deploy/claude/README.md b/services/core/deploy/claude/README.md index 2cbd3d301..797f52750 100644 --- a/services/core/deploy/claude/README.md +++ b/services/core/deploy/claude/README.md @@ -1,92 +1,49 @@ -# Claude Code on the dedicated Runtime - -Core is independently deployed. Each managed Session receives one Docker Runtime -containing the daemon, pinned Claude Agent SDK/native harness and local workspace. -The SDK owns the model/tool loop. Public clients use the same Agents API contract. - -## Build and configure - -The [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) -builds the image. The bundle pins SDK `0.3.269` and native Claude Code `2.1.269`. The Dockerfile pins -Node's Linux amd64 manifest. Keep the exported SDK bundle immutable. Configure -Core's database-owned managed deployment with the resulting immutable Runtime -image. Existing Docker outer security settings are unchanged; the daemon and -native adapter do not add an inner sandbox. Tools run as UID/GID 1000 and can -access files available to that user. - -Install required system dependencies while building the image/template. Runtime -has no automatic apt installation, sudo or elevated daemon permissions; -`system_packages` is unsupported and missing dependencies cause operation failure. -npm/Python package and setup commands run directly with the existing user's -permissions. The packaged image no longer needs bubblewrap or socat for an inner -sandbox. For native self-hosting and platform limits, see the -[self-hosted guide](../../../../docs/getting-started/self-hosted.md#platforms). - -Set `OAC_DEFAULT_HARNESS=claude_sdk`. Configure the deployment default model provider -with the Core key, in Web or through Core's API: - -```sh -curl -fsS -X PUT http://127.0.0.1:8091/core/v1/harnesses/claude_sdk/model-provider \ - -H "Authorization: Bearer $CORE_KEY" -H "Content-Type: application/json" \ - -d '{"protocol":"anthropic","base_url":"https://api.anthropic.com","api_key":"REPLACE_WITH_OPERATOR_SECRET"}' -``` - -A Session may instead supply its own `x_agents_core.model_provider`. Use the -endpoint and model supported by your actual provider. Do not bake keys into the -image. Core freezes the bundle in the Session's encrypted snapshot and delivers it -to the adapter; it is not a public Agent field. Ordinary environment projection -does not isolate locally stored credentials from tools running as the same user. -Follow the existing Core setup for independent PostgreSQL credentials, migrations, -API authentication and managed Runtime enrollment. Parsar is not a dependency. - -The supported hosted profile accepts text execution with medium verbosity and -native Bash/Read/Edit and declared public functions with text results. Files and Artifacts use the shared public interfaces. -Check the current qualified engine profile for MCP, subagents and structured -output combinations. Historical hosted-profile acceptance does not qualify native -self-hosted platforms or every feature combination. The existing -`environment:none` function/MCP profile is separate. This is not complete upstream -protocol compatibility. - -## Runtime and adapter rules - -Native self-hosted installations use the same bundle and daemon protocol. -The shared local binding selects the actual workspace. SDK history, home and -scratch remain under `OAC_RUNTIME_HOME/runtime/claude-sdk` for session ownership, -without restricting tools. Authentication remains under `OAC_RUNTIME_HOME/daemon`. -The daemon requires neither nested sandbox privileges nor host security changes. - -The `local_runtime_v2` capability identifies the current workspace and direct MCP -launcher contract. Reject earlier workspace bundles; matching SDK versions alone -do not establish adapter compatibility. -The adapter advertises `local_runtime_v2` after checking the installed bridge and -native binary. Windows additionally needs Git Bash for its Bash tool. Registration -combines that check with the local binding; Core consumes the same readiness -contract on every platform. The profile supports Bash/Read/Edit, preparation, -shared Files/Artifacts, cancellation and history continuation. Other capabilities -remain subject to their actual engine qualification; bypass execution does not -silently qualify a new MCP or function combination. - -Recovery uses the SDK's history APIs. An explicitly supplied native identity must -exist. If Core requires existing history without having recorded an identity, the -adapter accepts only one nonempty native history for the exact bound cwd. Missing, -foreign, ambiguous or metadata-only history rejects before model input. The Runtime -volume and shared Environment/Session binding establish ownership; this lookup -cannot select another Session's home or infer ownership from a model response. - -Claude hosted functions compose the existing SDK function bridge with the common -workspace execution profile. Only declared function tools and the verified native tool -inventory are available. The bundle advertises this combination separately from -basic workspace execution; function preparations require that verified combination. -Service-origin hosted MCP remains unqualified; Environment Plugin declarations -use the separately qualified initialization path above. Function callbacks do not change file, -credential, history, subagent or network authority. - -## Adding another native engine - -Follow [Add a native Harness](../../../../contracts/agents-api/harness-onboarding.md). -The Claude adapter is one of its [native references](../../../../contracts/agents-api/harness-onboarding.md#native-references): -it implements the shared `agent.ExecutorFactory`, `Executor` and `Turn` -interfaces and registers them in -[`cli/claude_sdk.go`](../../../../apps/daemon/internal/cli/claude_sdk.go). -Acceptance requirements are in -[Harness onboarding](../../../../contracts/agents-api/harness-onboarding.md#qualify-the-adapter). +# Claude Code Runtime + +The Claude adapter runs Claude Code through the pinned Claude Agent SDK. The SDK owns the model and tool loop. Two parts make up the adapter: the private TypeScript bridge in [`packages/claude-sdk-adapter`](../../../../packages/claude-sdk-adapter/README.md), which owns the bridge protocol and native SDK configuration, and the Go adapter in [`agent/claudesdk`](../../../../apps/daemon/internal/agent/claudesdk), which owns the bridge process. This page holds the Runtime-level rules and the Claude Runtime image. [Harness onboarding](../../../../contracts/agents-api/harness-onboarding.md) owns the obligations shared by all adapters. + +Native tools run with the daemon user's permissions; the outer sandbox provides isolation ([Runtime and outer isolation](../../../../docs/design-principles.md#runtime-and-outer-isolation)). + +## Native pin and readiness + +The bundle pins Claude Agent SDK `0.3.269` ([`package.json`](../../../../packages/claude-sdk-adapter/package.json)), which reports native Claude Code `2.1.269`. The bridge uses protocol 3 for a prepared Executor with separately identified Turns; the Go adapter's readiness and Executor checks reject any other protocol version ([`readiness.go`](../../../../apps/daemon/internal/agent/claudesdk/readiness.go), [`executor.go`](../../../../apps/daemon/internal/agent/claudesdk/executor.go)). + +Workspace execution requires the bundle to report the `workspace_tools`, `workspace_prepare`, `workspace_command_observations` and `local_runtime_v2` features; matching SDK versions alone do not establish compatibility. Function tools in workspace execution additionally require `workspace_functions`. The native installer rejects a bundle that lacks the workspace features ([`installation.go`](../../../../apps/daemon/internal/agent/claudesdk/installation.go)). + +Claude accepts only the `anthropic` model protocol ([`harnessconfig/claudesdk`](../../../../internal/harnessconfig/claudesdk/configuration.go)). [Model execution](../../../../contracts/agents-api/model-execution.md#deployment-defaults) owns provider selection. + +## Native state and recovery + +With a workspace binding, the SDK's history, home and scratch directories are `history`, `home` and `scratch` under `$OAC_RUNTIME_HOME/runtime/claude-sdk/` (mode 0700; `OAC_RUNTIME_HOME` defaults to `~/.oac`). The daemon's authentication stays under `$OAC_RUNTIME_HOME/daemon/`. These locations keep Session state apart; they do not restrict the tools. + +Recovery uses the SDK's history APIs ([`recovery.ts`](../../../../packages/claude-sdk-adapter/src/recovery.ts)). An explicitly supplied native Session ID must exist. When Core requires existing history but has no recorded ID, the adapter accepts only a single native Session whose recorded cwd equals the bound workspace and which has at least one message. Missing, foreign, ambiguous or empty history is rejected before any model input. + +## Native failure classification + +The bridge ([`native_failure.ts`](../../../../packages/claude-sdk-adapter/src/native_failure.ts)) classifies a failure only from root assistant messages (no parent tool use) of the same native Session that answer Core-submitted inputs not yet completed; replayed and synthetic messages are ignored. A classification is committed only when the matching native result is an error; a successful result clears it. The [Runtime protocol](../../../../docs/runtime-protocol.md#native-failure-classification) defines the codes. + +| SDK assistant error | Code | +| --- | --- | +| `authentication_failed`, `oauth_org_not_allowed`, `account_on_hold`, `verification_required`, `cloud_credential_error` | `authentication_error` | +| `billing_error` | `usage_limit_exceeded` | +| `rate_limit` | `rate_limit_exceeded` | +| `overloaded` | `server_overloaded` | +| `invalid_request` | `invalid_request` | +| `model_not_found` | `resource_not_found` | +| `server_error` | `server_error` | + +`unknown`, `max_output_tokens` and other results stay unclassified. The bridge drops the classification when any other error ends the Turn, when cleanup fails or when no native process ran. The Go adapter keeps it only for the reported failed result of the same native Session, and only after it received the Turn settlement ([`executor_turn.go`](../../../../apps/daemon/internal/agent/claudesdk/executor_turn.go)). + +## Runtime image + +[`Dockerfile`](Dockerfile) builds the Claude Runtime image from a prepared context that holds only `oac-daemon` and the exported SDK bundle; the [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) builds it. Keep the exported bundle unchanged. + +| Item | Value | +| --- | --- | +| Base | Digest-pinned `node:22.23.1-bookworm-slim` with `ca-certificates`, `bash`, `git`, `python3`, `python3-pip` and `ripgrep` | +| Programs | `/usr/local/bin/oac-daemon` and the SDK bundle at `/opt/claude-sdk` | +| User | UID/GID 1000 with `HOME=/home/runtime` | +| Environment | `OAC_RUNTIME_HOME=/home/runtime/.oac`, `OAC_RUNTIME_CLAUDE_SDK_NODE=/usr/local/bin/node`, `OAC_RUNTIME_CLAUDE_SDK_ENTRYPOINT=/opt/claude-sdk/dist/main.js`, `OAC_RUNTIME_WORKSPACE=/environment/workspace`, `OAC_RUNTIME_INITIALIZATION_DIRECTORY=/environment/initialization`, `OAC_RUNTIME_PACKAGE_DIRECTORY=/environment/packages` | +| Entry point | `oac-daemon connect --profile default`, working directory `/environment/workspace` | + +The build runs the bundle's `runtime_check.js` against its entry point. The distribution copies `/opt/claude-sdk` into the combined Runtime image. Sandboxes run the image with the [Docker sandbox settings](../codex/README.md#docker-sandbox-settings). diff --git a/services/core/deploy/codex/README.md b/services/core/deploy/codex/README.md index 6917d71c3..e490b2370 100644 --- a/services/core/deploy/codex/README.md +++ b/services/core/deploy/codex/README.md @@ -1,145 +1,66 @@ -# Codex colocated Runtime - -A managed Session runs the daemon, Codex and workspace in a dedicated outer -Environment. Core stays outside it. Use Codex 0.153.4 and its matching -`codex-resources`. The same daemon supports native self-hosted installations; see -[the self-hosted guide](../../../../docs/getting-started/self-hosted.md#platforms) for platform status. - -Codex tools run with the daemon user's existing permissions. There is no inner -filesystem, permission or network sandbox, no immutable Codex requirements file, -and no managed Bash wrapper. Tools may read Runtime state that the user can read. -Model credentials still belong in authenticated Runtime configuration, never image -layers. Do not mount another Session's history, host home, Docker socket or Core -credentials into the Environment. - -The existing outer Docker configuration is unchanged: non-root user, dropped -capabilities, no-new-privileges, read-only root, private PID namespace, dedicated -bridge networking, bounded tmpfs and process/memory/CPU limits. Its existing -`nested_sandbox` setting and seccomp file are deployment settings, not evidence -that the daemon adds inner isolation. The container-specific AppArmor exception -in that configuration removes an outer LSM layer; do not change host-wide policy. - -Seccomp source: [Moby profiles, revision 65adc7e](https://github.com/moby/profiles/blob/65adc7e022c97f55e45c054ff012988027733b87/seccomp/default.json), Apache-2.0. -Unmodified source SHA-256: -`785b2429264afba4d594320337cb17f144f3c7d51585f9805eef72e28f4f9334`. -The retained extra syscall rule is unchanged by removal of the inner sandbox. -Historical qualification of private-file denial under the former native sandbox -does not describe current tool permissions or qualify the new deployment. - -The packaged image retains `/environment/workspace`, initialization, package and -staging directories plus daemon/native state under `/home/runtime`. Docker also -mounts the workspace at `/workspace`. Runtime binds public Files/Artifacts to its -configured workspace; native tools execute with ordinary user permissions. The -shared helper artifacts remain in the matched Runtime bundle. The Docker engine -must support volume subpath mounts; its recorded qualification used 29.1.3. - -Published Artifacts are immutable Core-owned copies of successful Turn outputs; -listing and downloading them requires no running Environment. Validate actual -execution, file operations, cancellation, process settlement and retained native -history using the current binary and outer deployment. An alive container alone -is not acceptance. - -## Managed Runtime image and Docker adapter - -The [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) -builds the Linux amd64 image from the official `@openai/codex@0.153.4-linux-x64` -package. It contains the unmodified native executable and matching resources, and -no product server, product CLI, credentials or workspace data. - -The service's `internal/sandbox` interface has five operations. Its Docker adapter -uses the official Moby Go client and an operator-selected immutable image digest, -network, installation UUID and the contents of this directory's `seccomp.json`. -The Runtime authenticates outward through the ordinary daemon bootstrap path; -`CoreURL` includes the existing `/api/v1` gateway prefix. The image's default -entrypoint is the same daemon connect command used by a user-managed Runtime. -No socket, host home or product configuration is mounted inside the Runtime. - -The caller persists a fresh allocation UUID with the authorized tenant and -Environment before Create, and serializes lifecycle operations for that allocation. -Create returns the reference even on failure. Duplicate allocation creation does -not rewrite credentials or restart the container. After a lost response, inspect -the allocation and reconcile its actual state; do not blindly replay Create. -Core owns durable allocation reconciliation through its internal managed Runtime -coordinator. Public admission requires the explicit operator configuration below. - -Two labelled named volumes retain native state and the workspace/staging pair. -The trusted daemon auth profile is copied with restrictive permissions before -startup; it does not enter image layers, environment variables, labels or arguments. -GetInfo describes observed compute state, not daemon or native readiness. Docker -has no renewable provider lease: Renew verifies the allocation, while Core owns -keepalives and expiry. Kill verifies allocation ownership, removes -the container, explicitly removes its named volumes and confirms absence. Keep the -reference and retry cleanup when an operation fails; an HTTP timeout is not proof -that a resource disappeared. Never use broad container or volume pruning. - -Provider commands are limited to bootstrap and resource lifecycle. Runtime owns -initial files, configuration, npm/Python packages, setup and capabilities over its -authenticated protocol. An unconfirmed bootstrap result requires allocation -cleanup before reuse; compute readiness does not prove Runtime readiness. - -## Standalone operator configuration - -Use the [installation guide](../../../../docs/getting-started/install.md) and -[deployment configuration API](../../../../contracts/agents-api/sandbox-deployment.md). -Web or the deployment administrator API selects Docker, uniform per-sandbox CPU -and memory, and the matched immutable Runtime release in PostgreSQL. Docker does -not accept independent hard root or workspace disk quotas through this contract. -The saved public Core origin must be reachable from sandbox guests. - -Enroll an ordinary sandbox node with Web's one-command [Add node](../../../../docs/getting-started/nodes.md) -flow. It fetches and validates the deployment configuration, imports the matched -Runtime image and actively connects to Core. Core can run independently without a -Docker socket; the node owns its local Docker access. Node files contain the installed -selection and host-specific paths, never a separate provider or resource choice. The -Core host joins through the same command as any other host. - -Same-provider resource and Runtime edits currently require no active reset, -the current generation and verified zero held allocations or pending Environments. -A backend change requires explicit reset and confirmed cleanup, then a new setup. -See the [reset procedure](../../../../docs/getting-started/nodes.md#change-the-sandbox-configuration). -Stopped compute, snapshots and unknown operations remain blockers. Keep the original -node identity, paths and credentials until cleanup is confirmed. Explicit archive -preserves history and persisted Files/Artifacts but discards unpersisted workspace; -ordinary Session deletion has different retention behavior. Existing Sessions never -migrate to another backend. - -Node providers use explicit local Unix Docker sockets, ignoring ambient -`DOCKER_HOST`. Optional `extra_hosts` is trusted node configuration. No Docker -socket is mounted in a Runtime. Caller-managed `self_hosted` enrollment retains -its separate public lifecycle. - -Hosted Codex needs a model provider: either the Session's -`x_agents_core.model_provider` or the deployment default for `codex`, set with the -Core key in Web or through `PUT /core/v1/harnesses/codex/model-provider`. Core -freezes the bundle in the Session's encrypted snapshot and delivers it as the -common `model_provider` bundle. Runtime automatically adapts Chat Completions or -Anthropic to Codex Responses; a hosted Session without a provider is rejected with 400 -`model_provider_required`. Do not place credentials in images. - -With the qualified Codex image and Docker provider configured, create an idle or -initial-text Session using `environment: {"type": "openai_hosted"}`. Core commits -its identity before automatically provisioning it. Queries expose durable -connection status; execution separately prepares the native harness. Session -deletion revokes authority before owned container/volume cleanup. The daemon does -not enforce `disabled` or `restricted` networking; combinations without matching -outer enforcement are unsupported. Templates and inline configuration share -initial files, env, npm/Python packages, ordered setup and capabilities; see the -[supported fields and limits](../../../../contracts/agents-api/environments.md#templates). - -### Environment initialization - -Templates and inline env/setup/npm/Python configuration share the daemon's Go -Runtime preparation implementation on all platforms. The packaged image sets -`OAC_RUNTIME_INITIALIZATION_DIRECTORY=/environment/initialization` and -`OAC_RUNTIME_PACKAGE_DIRECTORY=/environment/packages` to retain its resource layout. -No Python initialization wrapper is shipped. It runs commands directly as UID/GID 1000 using configured tool -variables; it does not inject those variables into the daemon or native harness -launcher. Cancellation and finite receipts remain Runtime responsibilities. - -System dependencies must be installed during image/template construction. The -daemon never runs apt, sudo or another privilege escalation, and -`system_packages` is unsupported. A missing executable or library fails the -operation requiring it. The image no longer contains a system-root seed or -system-package launcher. Native self-hosted users prepare their own dependencies -before starting the daemon. See -[initialization contract and limits](../../../../contracts/agents-api/environments.md#runtime-capability-preparation). +# Codex Runtime + +The Codex adapter ([`agent/codex`](../../../../apps/daemon/internal/agent/codex)) runs the native Codex CLI through its app-server protocol inside the daemon. This page holds the Codex-specific adapter rules, the Codex Runtime image and the Docker sandbox settings that every Docker Runtime uses. [Harness onboarding](../../../../contracts/agents-api/harness-onboarding.md) owns the obligations shared by all adapters. + +Native tools run with the daemon user's permissions; the outer sandbox provides isolation ([Runtime and outer isolation](../../../../docs/design-principles.md#runtime-and-outer-isolation)). + +## Native pin + +The adapter accepts only Codex `0.153.4`: installation and recovery checks require `codex --version` to report `codex-cli 0.153.4` ([`installation.go`](../../../../apps/daemon/internal/agent/codex/installation.go), [`recovery.go`](../../../../apps/daemon/internal/agent/codex/recovery.go)). The Runtime image carries the official Linux amd64 package of that version with its matching `codex-resources`; the [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) builds it. + +## Model provider and native state + +Codex accepts only the `responses` model protocol ([`harnessconfig/codex`](../../../../internal/harnessconfig/codex/configuration.go)). A Chat Completions or Anthropic provider is rejected; nothing converts between protocols. [Model execution](../../../../contracts/agents-api/model-execution.md#deployment-defaults) owns provider selection, including the per-Harness deployment default. + +Each Session has its own `CODEX_HOME` at `$OAC_RUNTIME_HOME/daemon/agent-sessions//` (`OAC_RUNTIME_HOME` defaults to `~/.oac`). Native history stays there. The adapter regenerates that directory's `config.toml` on every prompt: the Session's frozen provider bundle becomes the `[model_providers.oac]` block, with `wire_api` `responses`, and the thread is pinned to that provider, so Codex never falls back to its built-in `openai` provider. + +## Native failure classification + +The adapter classifies a failed Turn only from the exact root terminal Turn's `codexErrorInfo` ([`error_classification.go`](../../../../apps/daemon/internal/agent/codex/error_classification.go)); error notifications, including retry notifications, never classify it. The [Runtime protocol](../../../../docs/runtime-protocol.md#native-failure-classification) defines the codes. + +| `codexErrorInfo` | Code | +| --- | --- | +| `unauthorized` | `authentication_error` | +| `usageLimitExceeded` | `usage_limit_exceeded` | +| `rateLimitExceeded` | `rate_limit_exceeded` | +| `contextWindowExceeded` | `context_length_exceeded` | +| `serverOverloaded` | `server_overloaded` | +| `internalServerError` | `server_error` | +| `badRequest` | `invalid_request` | +| `cyberPolicy` | `cyber_policy` | +| `httpConnectionFailed`, `responseStreamConnectionFailed`, `responseStreamDisconnected` | `connection_failed`, with the upstream `httpStatusCode` | +| `responseTooManyFailedAttempts` | `rate_limit_exceeded` when `httpStatusCode` is 429, otherwise `connection_failed` | + +Every other variant stays unclassified. + +## Runtime image + +[`Dockerfile`](Dockerfile) builds the Codex Runtime image from a prepared context that holds only `oac-daemon`, the unmodified `codex` executable and `codex-resources`. + +| Item | Value | +| --- | --- | +| Base | Digest-pinned `debian:bookworm-slim` with `ca-certificates`, `bash`, `git`, `python3`, `python3-pip`, `nodejs`, `npm` and `ripgrep` | +| Programs | `/usr/local/bin/oac-daemon`, `/usr/local/bin/codex` (mode 0555) and `/usr/local/codex-resources` | +| User | `runtime`, UID/GID 1000, home `/home/runtime` | +| Environment | `OAC_RUNTIME_HOME=/home/runtime/.oac`, `OAC_RUNTIME_CODEX_BIN=/usr/local/bin/codex`, `OAC_RUNTIME_WORKSPACE=/environment/workspace`, `OAC_RUNTIME_INITIALIZATION_DIRECTORY=/environment/initialization`, `OAC_RUNTIME_PACKAGE_DIRECTORY=/environment/packages` | +| Entry point | `oac-daemon connect --profile default`, working directory `/environment/workspace` | + +The build fails unless `codex --version` reports the pinned version. The image holds no credentials, workspace data or product software. The distribution copies the Codex executable and resources into the combined Runtime image; see the [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers). + +## Docker sandbox settings + +The Docker Sandbox Provider ([`sandbox/docker`](../../../../services/core/internal/sandbox/docker)) runs every Runtime image, whichever Harness it serves, with the same container settings ([`container_options.go`](../../../../services/core/internal/sandbox/docker/container_options.go)): + +- user 1000:1000, read-only root filesystem, all capabilities dropped, `no-new-privileges`, this directory's `seccomp.json` and AppArmor `unconfined`; +- the Docker network from the node's provider configuration (`oac-node-` from the node installer), plus any `extra_hosts` there; +- CPU and memory from the deployment specification, a 128-process limit and a 128 MiB `/tmp` tmpfs; +- two named volumes labelled with the installation, tenant, Environment and allocation: `-home` at `/home` and `-environment` at `/environment`, whose `workspace` subdirectory is also mounted at `/workspace`. The Docker Engine must support volume subpath mounts; +- with `nested_sandbox`, which the node installer sets, Docker's `/proc` masks are lifted (`/sys/firmware` and `/sys/devices/virtual/powercap` stay masked) and the container runs an init process. + +Create refuses to reuse retained volumes that have no container. It copies the [Runtime bootstrap](../../../../docs/runtime-bootstrap.md) file to `/home/runtime/runtime-bootstrap.json` (mode 0600, UID 1000) and the `/environment` workspace, staging, initialization and package directories into the container, then starts `oac-daemon connect --profile default --bootstrap-file /home/runtime/runtime-bootstrap.json`. When the created container does not have the configured CPU, memory and exact image, Create returns the error with `CreateSettled`. Docker has no lease, so Renew only reads the container state. Kill checks the ownership labels of the container and both volumes before removing any of them, then confirms that all three are gone. + +The node uses the explicit Unix socket in its provider configuration (`unix:///var/run/docker.sock` from the node installer) and ignores `DOCKER_HOST`. No Docker socket, host home or Core credential is mounted into a Runtime. + +### Seccomp profile + +`seccomp.json` is the Moby default profile at [revision 65adc7e](https://github.com/moby/profiles/blob/65adc7e022c97f55e45c054ff012988027733b87/seccomp/default.json) (Apache-2.0, see `seccomp.LICENSE`; upstream file SHA-256 `785b2429264afba4d594320337cb17f144f3c7d51585f9805eef72e28f4f9334`) with one appended rule that allows `clone`, `unshare`, `setns`, `mount`, `umount2` and `pivot_root`. The distribution ships this file to every Docker node as `runtime/seccomp.json`. diff --git a/services/core/deploy/e2b/README.md b/services/core/deploy/e2b/README.md index 27842da91..aadf2393c 100644 --- a/services/core/deploy/e2b/README.md +++ b/services/core/deploy/e2b/README.md @@ -1,54 +1,15 @@ -# E2B Runtime deployment - -E2B supports two separate ownership choices. Core-managed `openai_hosted` uses -the deployment-wide E2B Provider. Application-managed `self_hosted` uses the -startup example below; in that path the application creates, renews and destroys -its own sandbox, and Core receives neither the E2B account key nor an allocation -request. The sandbox runs the existing V1 daemon, selected native harness, tools -and workspace together. Codex, Claude Code and MiniMax Code use the same startup -contract and their respective qualified Runtime images. - -This deployment uses OpenAgentCore daemon enrollment, not Codex `exec-server` or Noise. -Public execution and Files/Artifacts continue through Core and the daemon; -E2B commands/files are used only for deployment, initialization and inspection. -A public Session deletion does not destroy an application-owned VM; managed -compute cleanup remains Core's responsibility. - -## Core-managed hosted deployment - -Select E2B in Core's hosted deployment setup and supply the account key and an -immutable `templateID:build_UUID`. One deployment uses one managed Provider; -E2B needs no physical node enrollment. Changing backend requires explicit reset -and confirmed cleanup of every owned allocation, pending creation and snapshot. -Do not infer execution readiness from saved configuration or running compute. - -The pure-Go adapter invokes the packaged official Python SDK helper. The Core -image and native installation include its runtime dependencies; configuring E2B -does not require installing Python or pip on the server. The account key is -write-only server configuration, passed to the helper over stdin. It never enters -a template, command argument, inherited environment, receipt or public response. - -Keep the helper's private receipt directory on durable storage, owned by the Core -service user with mode 0700. Manual deployments using the default non-root Core -image must give its service UID ownership of that directory. SDK connection -credentials and one-shot allocation claims are stored there; they are not Core -execution state. Preserve them across upgrades and failures. An empty cloud -lookup cannot settle an unknown Create, and inspection never creates, resumes -or reboots compute. Initialization uncertainty requires reclaiming the original -allocation rather than replaying startup. See the [helper contract](../../tools/e2b-provider/README.md). - -Rebuild old templates with this directory's `build-template.py` before managed -use: it includes protected `init.py` and `managed_init.py`. The latter consumes -Core's existing managed daemon auth profile, while application-managed startup -continues to use the executor credential flow below. Both reuse the same daemon, -native harnesses and colocated workspace. A qualified combined Runtime image can -support several harnesses; provider selection is independent of harness choice. - -## Build the packaged Runtime - -Use Python 3.12+, Docker and a qualified Linux amd64 Runtime image containing the -environment-aware daemon `connect` command. Keep keys and build outputs outside -the checkout, in private directories. Install the pinned SDK from this directory: +# E2B Runtime template and application-managed launch + +An E2B template packages a qualified Runtime image so that its daemon, native Harnesses and workspace run together in one E2B sandbox. The template serves two ownership models: + +- **Core-managed** (`openai_hosted`): the deployment selects E2B as its Sandbox Provider and Core creates, renews and destroys sandboxes from the template. [Sandbox deployment](../../../../contracts/agents-api/sandbox-deployment.md) owns the selection and the [E2B helper](../../tools/e2b-provider/README.md) owns the adapter. +- **Application-managed** (`self_hosted`): the application creates, renews and destroys its own sandbox and enrolls the daemon into a `self_hosted` Environment. Core receives neither the E2B account key nor an allocation request, and deleting the Session leaves the sandbox running. + +This page covers building the template and the application-managed launch. Execution and Files always go through Core and the daemon; E2B commands and files are used only to start and inspect the sandbox. + +## Build a template + +Use Python 3.12 or newer, Docker and a qualified Linux amd64 Runtime image. Keep keys and build outputs outside the checkout, in private directories. Install the pinned SDK (`e2b` 2.51.0, [`requirements.txt`](requirements.txt)) and build: ```sh python -m venv "$HOME/.oac/build/e2b-sdk" @@ -60,38 +21,15 @@ python -m venv "$HOME/.oac/build/e2b-sdk" --output "$HOME/.oac/build/e2b-template.json" ``` -The builder preserves the existing image's binaries, native configuration and -private workspace layout. Its `template` output is an immutable -`templateID:build_UUID`; use that exact value. The build must qualify every harness it advertises; a combined Runtime image -can include several harnesses. No E2B account key, executor key or model credential belongs in a -build, template environment, metadata, command argument or log. The builder gives -traversable modes only to the public archive ancestors it creates (`usr`, -`usr/local`, `etc`). Runtime file and directory modes, private build contexts, -key inputs and the umask of its output stay unchanged, including under umask 077. - -System dependencies must be installed when building the template. The builder may -use root during image construction, but the running daemon remains UID/GID 1000 -with no automatic apt, sudo or privilege escalation. Runtime `system_packages` -is unsupported; missing dependencies fail their consuming operation. The template -no longer installs bubblewrap or socat for inner sandboxing. - -## Application-managed: start an existing self-hosted Environment - -Create a public `self_hosted` Environment through Core and retain its ID and exact -returned `remote_url`. Obtain an authorized connect-only executor key scoped to -that Environment (or its owning principal) through the operator credential flow. -The key JSON is `{"key_id":"UUID","executor_token":"SECRET"}` with an optional -`environment_id` restriction. Store both this JSON and the separate E2B API key -in private files with mode `0600`. - -The current packaged profile uses public `/workspace`, backed by -`/environment/workspace`. The remote endpoint must be reachable from the VM; -use the returned `wss://.../api/v1/agent-daemon/ws` unchanged. The native profile -comes from the image, and model credentials arrive through authenticated Core -execution. Managed startup uses the separate [Runtime bootstrap contract](../../../../docs/runtime-bootstrap.md). - -Generate and retain an application launch UUID once. `launch.py` is a thin SDK -example, not a service or a replacement lifecycle owner: +[`build-template.py`](build-template.py) requires a `sha256:` image ID, a Linux amd64 image and the Runtime layout (`OAC_RUNTIME_WORKSPACE=/environment/workspace`). It copies the image's `/usr/local/bin`, `/opt` and, when present, `/usr/local/codex-resources` and `/etc/codex`, together with its `HOME` and `OAC_*` environment, onto a digest-pinned `node:22.23.1-bookworm-slim` base with `ca-certificates`, `bash`, `git`, `python3`, `python3-pip`, `ripgrep` and `util-linux`. It installs [`init.py`](init.py) and [`managed_init.py`](managed_init.py) read-only under `/opt/oac-e2b`, runs the daemon as UID/GID 1000 (`runtime`, home `/home/runtime`) and builds with 2 vCPUs and 2048 MiB. Only the `usr`, `usr/local` and `etc` archive ancestors it creates get traversable modes; Runtime file modes, private build contexts, key inputs and the output's umask stay unchanged. + +The output file records `template` (the immutable `templateID:build_UUID`), the source `image`, the packaged `runtime_sha256` and the `base`. Use that exact `template` value. The template must qualify every Harness its image advertises. Install system dependencies into the image; the daemon runs as UID/GID 1000. No E2B account key, executor key or model credential belongs in a build, template environment, metadata, command argument or log. + +## Launch an application-managed Runtime + +1. Create a `self_hosted` Environment through Core and keep its ID and the exact returned `remote_url` (`wss://…/api/v1/agent-daemon/ws`). The VM must reach it. Set `workspace_directory` to `/workspace`, which the template binds to `/environment/workspace`. +2. Obtain a connect-only executor key for that Environment; [Environment executor credentials](../../../../contracts/agents-api/environment-executor-credentials.md) owns issuance. The key file is `{"key_id":"UUID","executor_token":"SECRET"}`, with an optional `environment_id` restriction. Store it and the E2B API key in private files with mode `0600`. +3. Generate an application launch UUID once and keep it. Run [`launch.py`](launch.py), a thin SDK example rather than a service: ```sh "$HOME/.oac/build/e2b-sdk/bin/python" services/core/deploy/e2b/launch.py \ @@ -105,19 +43,11 @@ example, not a service or a replacement lifecycle owner: --timeout 7200 ``` -Choose a lease supported by your E2B account and renew it before expiry. The -example creates an exclusive private launch record before Create, stores the -returned sandbox ID before startup, and sets `on_timeout=kill` with auto-resume -disabled. Metadata contains only the application launch ID and Environment ID. -It never repeats Create/start, replaces a sandbox or deletes failure evidence. -Reusing the record path rejects before any cloud call. Do not bypass that guard -by supplying a new path after an uncertain result. +The example creates the private launch record exclusively before Create, so reusing a record path fails before any cloud call. It saves the sandbox ID before startup and creates the sandbox with `on_timeout=kill` and auto-resume disabled; its metadata holds only `oac_launch_id` and `oac_environment_id`. It never repeats Create or startup, replaces a sandbox or deletes failure evidence. After an uncertain result, inspect the record and the sandbox; do not start again with a new record path. Choose a lease your E2B account supports and renew it before it expires. ## Inspect, renew and destroy -The application remains responsible for the lease and cleanup, including after -Session deletion or daemon failure. These SDK calls use the retained exact ID -and do not connect to, resume or recreate a sandbox: +The application owns the lease and cleanup, including after Session deletion or daemon failure. These SDK calls use the recorded sandbox ID and never connect to, resume or recreate a sandbox: ```python import json @@ -135,59 +65,29 @@ Sandbox.set_timeout(sandbox_id, 7200, api_key=api_key) # When renewing the live Sandbox.kill(sandbox_id, api_key=api_key) ``` -If Create's response was lost before its ID was saved, discover candidates using -`Sandbox.list(query=SandboxQuery(metadata={'oac_launch_id': launch_id}), -api_key=api_key)`, importing `SandboxQuery` from `e2b`. Consume pages while -`paginator.has_next` via `paginator.next_items()`. Verify both metadata fields -against the private record, retain every matching provider ID, and explicitly -inspect or destroy those allocations. An empty lookup is not permission to retry -an uncertain Create. Do not select an arbitrary candidate or rotate its identity. - -An uncertain startup result requires inspection of the same VM or explicit -cleanup. Root-only `/root/.oac/e2b/launch.json` records the startup claim; -`ready.json` records only successful process handoff (`daemon_started`), even if -the daemon subsequently exits. Neither proves enrollment, native readiness or a -successful Turn. Check public Core Environment status and the private daemon log -at `/home/runtime/.oac/daemon/default/daemon.log`. The daemon's separate -`/home/runtime/.oac/daemon/environment.json` records the verified -Environment/Session binding. Keep these records and native history on failure. -Do not rerun `init.py`; it refuses any claimed attempt, including interrupted ones. -The SDK's `Sandbox.connect` can resume paused sandboxes, so it is not used as an -automatic recovery/inspection step here. VM expiry destroys volatile history; -never claim a replacement VM recovered the original Session. +If Create's response was lost before its ID was saved, list candidates with `Sandbox.list(query=SandboxQuery(metadata={'oac_launch_id': launch_id}), api_key=api_key)` (`SandboxQuery` comes from `e2b`) and read every page while `paginator.has_next`, using `paginator.next_items()`. Check both metadata fields against the record, keep every matching sandbox ID and inspect or destroy each of those sandboxes. An empty listing does not permit another Create. Do not pick one candidate arbitrarily. -## Startup and security boundary +After an uncertain startup, inspect the same VM or destroy it. The SDK's `Sandbox.connect` can resume a paused sandbox, so do not use it to inspect. The VM keeps these records: -The protected image initializer uses the image's explicit environment, restores -E2B-finalized executable/service permissions, locks the unused privileged `user` -account, and binds `/workspace`. It writes the executor key to a mode-`0600` file -inside the protected daemon directory, then starts the existing daemon as UID -1000 with that file path. The input is removed before startup. No bearer enters -the daemon's argv or inherited environment. Enrollment and immutable local binding -remain daemon responsibilities; startup does not invent device/Session IDs. +| Path | Content | +| --- | --- | +| `/root/.oac/e2b/launch.json` | The startup claim, written before any other startup step | +| `/root/.oac/e2b/ready.json` | Process handoff only (`daemon_started`, with the daemon PID), even if the daemon exits later | +| `/home/runtime/.oac/daemon/default/daemon.log` | The daemon log | +| `/home/runtime/.oac/daemon/environment.json` | The daemon's verified Environment and Session binding | + +Neither record proves enrollment, native readiness or a successful Turn; check the Environment status in Core. Keep the records and native history after a failure. `init.py` refuses to run again once any launch record exists. VM expiry destroys the history in it; a replacement VM never recovers the original Session. + +## Startup and security boundary -E2B clears `/run` at boot, so startup records live under `/root/.oac/e2b`. -The E2B VM is the outer isolation boundary. The daemon and native harness add no -inner filesystem, permission or network sandbox; tools have UID 1000's access to -Runtime state. Verify the actual outer boundary, process cleanup and native -execution for each template. Prior private-file denial results describe the old -inner sandbox and do not qualify the current behavior. -The one-shot receipt cannot be reused to replace the daemon or overwrite history. +`init.py` runs once as root. It restores the ownership and modes that E2B finalization changes under `/usr/local` and on `envd`, `/etc/inittab` and `/etc/init.d/rcS`, locks E2B's passwordless `user` account, checks that the image environment is not already bound to an Environment or Session and bind-mounts `/environment/workspace` at `/workspace`. It writes the executor key to `/home/runtime/.oac/daemon/executor-key.json` (mode 0600, owned by UID 1000), deletes the startup input and starts `oac-daemon connect --profile default --remote … --environment-id … --credential-file …` as UID/GID 1000. No credential enters the daemon's arguments or inherited environment. The daemon owns enrollment and the local binding. E2B clears `/run` at boot, so the records live under `/root/.oac/e2b`. -## Verification scope +The E2B VM is the isolation boundary; tools have UID 1000's access to Runtime state ([Runtime and outer isolation](../../../../docs/design-principles.md#runtime-and-outer-isolation)). -Run the focused local tests with the pinned SDK environment: +## Tests ```sh -python -m unittest discover -s services/core/deploy/e2b -p '*_test.py' -v +make check-e2b-provider ``` -These are controlled startup-contract tests: input binding, protected key output, -unchanged URL, no secret in argv/environment/record, one-shot claim, and retained -provider ID on unknown outcomes. They do not create billable resources or qualify -E2B security, enrollment or model execution. The [qualification record](../../../../contracts/agents-api/harness-capabilities.md) -records historical three-harness deployment results and their verification limits, -including shared Core credential lifecycle checks and explicit application-owned -cleanup. `tests/official_user_runtime.py` supplies the shared -public execution checks. The former Core-managed `official_e2b_v1.py` fixture is -retired; its original source and evidence remain in Git history. +With `OAC_TEST_E2B_SDK_PYTHON` pointing at the pinned SDK environment, this runs this directory's tests and the helper's. They cover startup input binding, the protected key file, the unchanged URL, credentials kept out of arguments, environment and records, the one-shot claim, and retained sandbox IDs after unknown outcomes. They create no billable resources. `make check-core` runs `managed_init_test.py` with only the standard library. diff --git a/services/core/deploy/mcode/README.md b/services/core/deploy/mcode/README.md index 0539f74b7..42e784646 100644 --- a/services/core/deploy/mcode/README.md +++ b/services/core/deploy/mcode/README.md @@ -1,215 +1,53 @@ # MiniMax Code Runtime -MiniMax Code uses the same Core/Runtime execution contract as the other engines. -The text profile supports `environment:none`; the dedicated Docker profile adds -workspace execution and the shared Files/Artifacts path. The historical -workspace qualification (evidence under `~/.parsar/remediation/20260919/mcode-workspace/`) -does not describe the current bypass profile. -Public MCP/functions, image input and native Subagent execution remain outside -this batch. - -## Pinned prerequisite - -Use the official [`@minimax-ai/code`](https://github.com/MiniMax-AI/minimax-code) -package version **0.4.12**, with Node.js 22.x and its native SQLite dependency. The Docker image pins Node.js -22.23.1; the text fixture used 22.22.0. -The inspected upstream source is `33b259bbbeb1c16433390869938191d09bdb0680`. -Install outside the checkout, under a private operator directory in `~/.oac/`. -Check the native install succeeds and `mcode --version` reports exactly 0.4.12. -This profile runs on a trusted execution host. - -Set `OAC_RUNTIME_MCODE_BIN` to that absolute executable and `OAC_RUNTIME_MCODE_AGENTS_API=1` -for the daemon. The opt-in only advertises the profile for the qualified version. -Use the existing authenticated daemon connection and operator device enrollment; -native self-hosted installation uses the same Runtime protocol. See the -[self-hosted guide](../../../../docs/getting-started/self-hosted.md#platforms); MiniMax on Windows remains -unsupported. -Set `OAC_DEFAULT_HARNESS=mcode` in the independent Core deployment. Existing Sessions -retain their engine. Do not expose a new public harness selector. - -## Docker workspace - -The [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) -builds the companion from the pinned native source and the Runtime image. - -Configure Core's existing managed Docker provider with the immutable image ID, -`deploy/codex/seccomp.json` and `nested_sandbox: true`. Core, database ownership, -enrollment and the public protocol remain shared. The image supplies the private -`OAC_RUNTIME_MCODE_WORKSPACE=managed` and companion paths; caller Agent options cannot -change them. Public Files and Artifacts use the common bound workspace helpers. - -The native process, ACP Session and six original native tools use the declared -Environment workspace (`/workspace` in the Docker image). The tools connect -through one trusted MCP bridge with the daemon user's ordinary permissions and -no inner sandbox. MCP is internal transport here; its presence alone does not -enable caller-supplied public MCP servers. Native project discovery uses that -same workspace; adapter-owned configuration and history remain in the private -Session data directory. Exact native-ID continuation remains required; recovery -without a recorded ID fails closed. Native Bash observations become public -`command_execution` items after command arguments arrive. Their text output and -status are retained; absent native exit code/duration stay unknown. The private -MCP server and file/skill/task utilities are not invented public MCP or function -calls. - -## Provider configuration - -Supply an Anthropic-compatible bundle with the model's context and output limits, -either per Session as `x_agents_core.model_provider` or as the deployment default, -set with the Core key in Web or through Core's API: - -```sh -curl -fsS -X PUT http://127.0.0.1:8091/core/v1/harnesses/mcode/model-provider \ - -H "Authorization: Bearer $CORE_KEY" -H "Content-Type: application/json" \ - -d '{"protocol":"anthropic","base_url":"https://api.moonshot.cn/anthropic","api_key":"","context_window":64000,"max_output_tokens":4096}' -``` - -Core maps the bundle to the native custom provider for the Session's exact -`agent.model`, with these limits. The adapter selects it explicitly; it does not -use a native account fallback. -Use the real provider endpoint for the selected credential. Test proxies may -forward requests unchanged and capture tool names/status, but must not synthesize -model responses or log keys. - -MiniMax M2 Chat Completions may return inline `` content. The pinned native -Harness does not expose the provider's `reasoning_split` option; the adapter does -not guess which response text to remove. Anthropic Messages supplies distinct -thinking blocks. File read/write utilities also have no qualified public Item -mapping; only actual Bash calls become `command_execution`. - -## Execution boundaries - -Each API Session uses a separate native state directory. Workspace execution -uses the declared Environment directory; text-only execution uses a private cwd. -The execution child inherits only process and model-network essentials; it does -not inherit Core/daemon tokens or arbitrary Node startup configuration. The -adapter owns its native config, instructions and home and disables external -skills, delegated work, web search, builtin file/shell tools, browser tools, -mcode-tools and native goals. The workspace profile uses its tool bridge without -sandbox isolation. Tools can read any local state available to the launching -user. The outer Environment owns managed isolation; daemon network modes do not -add another network boundary. - -Provide system dependencies at image/template build time or on the self-hosted -machine before execution. Runtime does not run apt or sudo or elevate daemon -permissions. `system_packages` is unsupported; missing dependencies fail the -operation requiring them. npm/Python packages and setup retain direct execution -with the user's existing permissions. - -Workspace Skills use the frozen installation's exact names and native `skill` -loader. Session-local links resolve to the installed package directories; -external discovery stays disabled. The text-only profile has no selected Skills. -Task utilities such as `task_query`, `task_output` and `task_stop` cannot create -work and enforce native Session ownership. Auxiliary native title requests are -internal bookkeeping. Do not equate these with public function or workspace tools. - -ACP `mcode/session/steer` acknowledges acceptance for the active native Turn, -separately from the transport write. It is not a promise that the model consumed -that input before cancellation. Unknown outcomes are never automatically replayed. -Cancellation waits for process-group exit and output settlement. Public slash -text remains a model message instead of invoking native ACP operator commands. - -Cold continuation requires the exact persisted native Session ID and matching -workspace cwd (or private cwd for text-only execution). Missing/foreign history -fails; recovery by guessing an ID from native session listings is not qualified. -Public usage breakdown is unavailable because native ACP context occupancy and -cumulative cost are not per-Turn usage. - -### Adapter rules - -MiniMax Code's opt-in Agents API profile qualifies native ACP 0.4.12 for -`environment:none` text execution. It reuses the shared lifecycle without public -workspace, functions or MCP. Enabled Subagents use the separately qualified -common observation path below. Native configuration disables -file/shell authority and external -capability discovery; the child receives a private Session home and a restricted -environment. Active-input application requires a native ACP receipt. Normal -Turns retain one ACP connection and native Session. Cancellation requires native -root and child completion evidence before reuse. Without a qualified root-state -reader, the ACP cancellation response is insufficient: stop the invalid native -process and wait for its exit before settling the Turn as non-reusable. Common -Runtime recovery then loads the exact owned native history in a new Executor. Do not infer history IDs or qualify hosted execution from this text -profile. See [Acceptance](#acceptance). - -MiniMax companion readiness uses private protocol 2; old companions are rejected -even when the upstream version matches. Cancellation retires the executor and -settles its native workers and detached Bash groups before acknowledgement; -later work recovers the same native history in a new owner without replay. - -The MiniMax workspace profile builds one CLI from the fixed upstream source and -lockfile through the existing companion packaging path, and connects native -workspace tools through its standard MCP client. The process and native Session -share the declared workspace directory, including for cold continuation. A -trusted adapter-owned bridge runs the original six tool implementations with the -launching user's ordinary permissions. It adds no inner sandbox on any platform. -Keep native history in the Session data directory and native execution and -Files/Artifacts bound to the same Environment workspace. Core and shared file -helpers remain engine neutral. This internal MCP transport does not admit public -MCP configuration. Record the upstream revision, native admission patch hashes -and worker-source provenance. The bounded patch checks the shared -descendant-task limit inside the existing native SQLite admission transaction -before start, without another scheduler. ACP initialization must acknowledge the -applied limit before input. The native tool catalog selects the same admitted -workspace tools for every child, not only the root's configured profile; this is -tool selection, not filesystem isolation. Only the Session's authorized internal -workspace MCP entry crosses the native child selector; this does not grant -external MCP access. Initialization must acknowledge the admitted tool inventory -before input as well. Subagent reads use the Session's native database; the -daemon does not protect it from other tools running as the same user. -Multi-agent workspace execution installs only the existing authorized workspace -MCP entry in that private native configuration so children inherit the same -tools; public MCP and Environment-origin MCP combinations remain separately -qualified. Complete the hosted -[acceptance checklist](../../../../contracts/agents-api/harness-onboarding.md#qualify-the-adapter) -before enabling hosted execution. The standalone companion uses its own npm -lock; `make check` runs its lifecycle tests and script checks, while its -exact-source Linux build and native qualification for the actual supported -platform and outer deployment are required when the companion changes. -Historical tests of both inner network policies do not qualify the current -bypass implementation. - -## Acceptance - -Recorded results below are historical evidence for their exact binaries and -profiles, not qualification of the current native platforms or bypass behavior. - -The opt-in `TestNativeMCodePublicExecution` uses the fixed official Python SDK, -raw HTTP, actual daemon/gateway/Worker and a dedicated PostgreSQL test database. -Provide private `OAC_TEST_MCODE_REAL_OPTIONS` (the provider object above plus the -`model` string), `OAC_RUNTIME_MCODE_BIN`, `OAC_TEST_NATIVE_DAEMON_BIN`, -`OAC_TEST_NATIVE_PROOF_DIR`, `OAC_TEST_OFFICIAL_SDK_PYTHON` and -`OAC_TEST_DATABASE_URL`, then run: - -```sh -go test ./services/core/internal/store \ - -run '^TestNativeMCodePublicExecution$' -count=1 -v -timeout=15m -``` - -It verifies real text execution, same-Turn steering and durable input receipt, -daemon cold restart with the same native ID, ordinary slash text, foreign-tenant -rejection, cancellation and subsequent continuation. Run independently with real -Kimi and MiniMax options. This fixture creates only operator device credentials -privately; all tested Sessions and inputs enter through public HTTP. - -`TestNativeMCodeHistoryIsolation` additionally takes -`OAC_TEST_MCODE_FOREIGN_NATIVE_ID` from a successful public run and verifies rejection -in another private native home, plus rejection of a nonexistent history ID. Run it -in the daemon's `internal/agent/mcode` package with the same private native/provider -options. It must fail before model input, without a replacement native Session. - -On 2026-09-19, the public fixture passed independently with real Kimi K3 and -MiniMax M2.7 APIs. Both exercised execution, native steering receipts, daemon -restart, history continuation, tenant rejection and cancel/continue. Native -missing/foreign history rejection passed separately. Credential scans found the -provider key only in each private mode-0600 native config, not in logs, public -results or native history. Product ACP regression uses the same 0.4.12 package -and covers new/resume, model/instruction refresh, Skill discovery and MCP using -its existing synthetic provider fixture. - -That text acceptance used an in-process Core HTTP server and a separate real -daemon. The subsequent Docker workspace qualification -(`~/.parsar/remediation/20260919/mcode-workspace/`) used independently built -Core binaries, its own database and managed Runtime. -It covers public Files/Artifacts, cancellation, reconnect, exact-history recovery -and isolation with real Kimi and MiniMax APIs. Read its recorded Kimi continuation -limitation: new model-issued commands are not automatic recovery replay. Neither -acceptance establishes complete official protocol compatibility. +The MiniMax Code adapter ([`agent/mcode`](../../../../apps/daemon/internal/agent/mcode)) runs the native MiniMax Code CLI over ACP inside the daemon. MiniMax Code keeps its own ACP Session, model loop and history. The companion package [`packages/mcode-harness`](../../../../packages/mcode-harness/README.md) owns the workspace tool bridge, the native patch and the Subagent history reader. This page holds the Runtime-level adapter rules and the MiniMax Code Runtime image. [Harness onboarding](../../../../contracts/agents-api/harness-onboarding.md) owns the obligations shared by all adapters. + +Native tools run with the daemon user's permissions; the outer sandbox provides isolation ([Runtime and outer isolation](../../../../docs/design-principles.md#runtime-and-outer-isolation)). + +## Native pin and readiness + +The adapter accepts only `@minimax-ai/code` `0.4.12` ([`version.go`](../../../../apps/daemon/internal/agent/mcode/version.go)), built from the upstream source revision in [`source.json`](../../../../packages/mcode-harness/source.json) and run with Node.js 22. The daemon advertises MiniMax execution only when `OAC_RUNTIME_MCODE_AGENTS_API=1` and the version matches ([`execution.go`](../../../../apps/daemon/internal/agent/mcode/execution.go)); the native installer and the Runtime image set it. MiniMax Code runs on Linux and macOS. + +Workspace execution also requires the companion's readiness report: private protocol 2, the pinned native version and the pinned source revision ([`workspace_readiness.go`](../../../../apps/daemon/internal/agent/mcode/workspace_readiness.go)). An older companion is rejected even when the upstream version matches. + +## Model provider + +MiniMax Code accepts the `anthropic`, `responses` and `chat_completions` protocols, and each requires the provider's context window and maximum output tokens ([`harnessconfig/mcode`](../../../../internal/harnessconfig/mcode/configuration.go)). The adapter writes the frozen bundle as the native custom provider for the Session's exact `agent.model`, with those limits, using the native `anthropic-messages`, `openai-responses` or `openai-completions` API ([`model_provider.go`](../../../../apps/daemon/internal/agent/mcode/model_provider.go)). It selects that provider explicitly; there is no native account fallback. [Model execution](../../../../contracts/agents-api/model-execution.md#deployment-defaults) owns provider selection. + +## Execution profiles + +Each Session has its own native data directory, `$OAC_RUNTIME_HOME/runtime/mcode/state//` (`OAC_RUNTIME_HOME` defaults to `~/.oac`), holding the generated native configuration, instructions (`AGENTS.md`, at most 32 KiB), MCP configuration and history. The native process gets `MINIMAX_DATA_DIR`, `HOME` and `USERPROFILE` set to that directory on top of the daemon user's environment. + +| Profile | Where it runs | Native configuration | +| --- | --- | --- | +| Text (`environment: none`) | A private working directory inside the Session's data directory | No native tools, Skills, web search, browser tools, `mcode-tools`, thread goals or user questions. With Subagents enabled, the default agent gets only the task tools (`task`, `task_append`, `task_query`, `task_output`, `task_stop`) and delegation | +| Workspace | The Environment's workspace directory, for the native process, the ACP Session and the workspace tools | The text configuration plus permission mode `bypassPermissions`, the native sandbox off, the selected workspace Skills and the `oac_workspace` MCP bridge for file and shell tools | + +The text profile rejects workspace, function tools and MCP servers. Workspace execution requires Linux or macOS and an unrestricted Environment network policy ([`workspace.go`](../../../../apps/daemon/internal/agent/mcode/workspace.go)); the adapter enforces no network restriction. + +In the workspace profile, the frozen installation's Skills are linked into the Session's `skills` directory under their exact names and loaded through the native `skill` tool; external Skill discovery stays off. Public MCP servers use the Environment origin only, as the [MCP origin contract](../../../../contracts/agents-api/environments.md#public-mcp-connection-origin) describes. With Subagents enabled, the Session's native `mcp.json` holds only the `oac_workspace` server, so child agents get the same workspace tools. + +## Turns, cancellation and continuation + +- Normal Turns reuse one ACP connection and native Session. +- Active input uses ACP `mcode/session/steer`. Its receipt confirms that the active native Turn accepted the input, not that the model consumed it. Unknown outcomes are never replayed. +- Public input stays a model message: the adapter adds an empty text block so that ACP does not treat slash text as an operator command. +- A Turn that ends with an error, a cancellation, an unknown input outcome, an unfinished tool call or an unanswered permission or question makes the Executor non-reusable. The adapter stops the native process group and waits for it to exit before the Turn settles; later work loads the same native history in a new Executor. +- Continuation requires the exact recorded native Session ID and the same working directory. Recovery without a recorded ID is rejected. +- MiniMax Code reports no usage: ACP context occupancy and cumulative cost are not per-Turn usage. +- Workspace `workspace_bash` calls become command observations with the command text, output text and status; exit code and duration stay unknown. Native task and Skill utilities produce no public Items. +- The adapter attaches no [native failure classification](../../../../docs/runtime-protocol.md#native-failure-classification) to its errors. + +## Runtime image + +[`Dockerfile`](Dockerfile) builds the MiniMax Code Runtime image from a prepared context that holds `oac-daemon` and the companion artifact; the [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) builds both. + +| Item | Value | +| --- | --- | +| Base | Digest-pinned `node:22.23.1-bookworm-slim` with `ca-certificates`, `bash`, `git`, `python3`, `python3-pip` and `ripgrep` | +| Programs | `/usr/local/bin/oac-daemon` and the companion at `/opt/mcode-harness` (native CLI at `native/cli.js`, bridge at `bridge.mjs`) | +| User | UID/GID 1000 with `HOME=/home/runtime` | +| Environment | `OAC_RUNTIME_HOME=/home/runtime/.oac`, `OAC_RUNTIME_MCODE_NODE`, `OAC_RUNTIME_MCODE_BIN`, `OAC_RUNTIME_MCODE_WORKSPACE_BRIDGE`, `OAC_RUNTIME_MCODE_AGENTS_API=1`, `OAC_RUNTIME_WORKSPACE=/environment/workspace`, `OAC_RUNTIME_INITIALIZATION_DIRECTORY=/environment/initialization`, `OAC_RUNTIME_PACKAGE_DIRECTORY=/environment/packages` | +| Entry point | `oac-daemon connect --profile default`, working directory `/environment/workspace` | + +The build runs the companion's `check.mjs` and the native `--version`. The combined Runtime image uses this image as its base. Sandboxes run it with the [Docker sandbox settings](../codex/README.md#docker-sandbox-settings). diff --git a/services/core/deploy/microsandbox/README.md b/services/core/deploy/microsandbox/README.md deleted file mode 100644 index b18621910..000000000 --- a/services/core/deploy/microsandbox/README.md +++ /dev/null @@ -1,173 +0,0 @@ -# Single-host idle microVM suspension - -Run the ordinary standalone sandbox node and its microsandbox helper natively on -Linux amd64 with KVM access. Core owns scheduling and durable recovery over the -node protocol; it can run independently in a container or on another host. -PostgreSQL can remain in Docker. Core packaging, container or native, is independent -of the provider. - -A hosted Session gets a dedicated microVM with the existing daemon, native harness -and workspace. Core suspends it only after a completed Turn has remained idle and -all admitted work has settled. The next Turn restores the same Session history, -files and configuration. Native harness startup and shutdown retain their existing -behavior. This profile uses microsandbox v0.7.2; the helper performs one finite -operation and exits. - -Start with the [installation guide](../../../../docs/getting-started/install.md) -and the [nodes guide](../../../../docs/getting-started/nodes.md). The -[deployment configuration contract](../../../../contracts/agents-api/sandbox-deployment.md) -defines the saved selection, resources, Runtime identity and node authorization. -The [provider contract](../../tools/microsandbox-provider/README.md) describes -snapshot identity, uncertain operations and cleanup. Real-model continuation and -isolation still require deployment acceptance. - -## Select and change the deployment - -Web or the deployment administrator API selects one provider for the installation. -PostgreSQL owns the provider, per-sandbox CPU/memory/disk limits and immutable Runtime -release. Microsandbox nodes install that selection and must match its generation -and specification digest. Node files hold host paths and an installed copy of the -specification; they cannot select different resources or a different Runtime. -Harness selection is independent. Snapshot suspension is available only on -microsandbox; Docker and E2B retain their own supported lifecycle. - -Same-provider resource and Runtime edits currently require no active reset, -the current generation and verified zero held allocations or pending Environments. -A backend change requires explicit reset and confirmed cleanup, then a new setup. -See the [reset procedure](../../../../docs/getting-started/nodes.md#change-the-sandbox-configuration). -Stopped compute, snapshots and unknown operations remain blockers. Keep the original -node identity, paths and credentials until cleanup is confirmed. Explicit archive -preserves history and persisted Files/Artifacts but discards unpersisted workspace; -ordinary Session deletion has different retention behavior. Existing Sessions never -migrate to another backend. - - -Core does not automatically adopt an -older file-managed database, even after its resources are drained. Keep the -previous release and original backend available to resolve that deployment; -removing an environment variable does not migrate its configuration ownership. - -## Host and binaries - -Use a dedicated service account with access to `/dev/kvm`. The node's helper -requires glibc and the standard Linux dynamic libraries. Core's container image -does not run the helper; the standalone native node does. Use the matched -distribution's ordinary node installer for an operator installation. The -[maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) -builds the helper and names the checksum-verified `msb` and `libkrunfw.so.5.6.1` -release archive. `make check` runs the pure-Go provider tests on every supported -host. On Linux it also runs the SDK helper module; other hosts print an explicit -skip for that Linux-only module. A full Linux check is required before publishing -this profile. - -For a manual node installation, create a private, short runtime state path, for -example `~/.oac/msb`, with mode 0700. The ordinary installer instead selects -`~/.oac/m//` and stores node identity/configuration -under `~/.oac/nodes//`, both in the home of the account that runs -the node (`/var/lib/oac-node` for a node added with sudo). Keep these on persistent local storage -reserved for this installation. Unix socket path limits apply. Runtime storage -contains confidential disks, memory snapshots and SDK state; preserve it with the -node identity and Core database when recovering the host. - -## Runtime image - -Use the immutable Runtime release saved in the deployment specification. The -ordinary node installer verifies the matched distribution manifest and imports -its microsandbox image under the declared digest reference. The image must be -available to the local microsandbox installation before provisioning. For source -builds, the [Runtime image build](../../../../docs/maintainers.md#runtime-images-and-helpers) -remains the image source; produce and select a matching distribution rather than -substituting a local image for an already saved release. - -The provider performs the existing Runtime bootstrap, starts the daemon as -uid/gid 1000 and creates `/run/oac` as a private control directory. No model, -Core or tenant credential belongs in the image. The Runtime initializer executes -with that user's existing permissions; the microVM is the isolation boundary. -System dependencies belong in the image: Runtime does not run apt or sudo and -rejects `system_packages`. Snapshot restore does not rerun setup or initial files. - -## Deployment and node configuration - -Initialize microsandbox through Web, `POST /core/v1/sandbox/deployment` (including -its required `resources` and immutable `runtime` fields) or `install.sh --sandbox -microsandbox`. The standalone node installer reads -`GET /api/v1/sandbox-node/configuration` using an enrollment token, or its retained -node credential on a registered reinstall. It verifies the saved release and -resources before registration. The Core host is added the same way, with Add node, -and needs a guest-reachable, non-loopback HTTPS public URL like any other node. - -Set VM CPU, memory and disk limits in the database-owned specification. -`root_disk_mib` bounds the managed root disk; `environment_disk_mib` separately -bounds the owned ext4 disk at `/environment`, including workspace, staging and -initialization data. Both are required. Budget for both disks and retained full -snapshots. Node `max_active` and `max_retained` are separate reservation limits; -retained allocations include suspended snapshots and uncertain cleanup. The -managed microsandbox policy uses a five-minute idle interval and one-day snapshot -retention. - -The private node provider file supplies its absolute helper/runtime/firmware paths, -short Runtime home and explicit host network policy. Permit the required Core, -model and package-registry endpoints. Creation and restore apply the same host -policy. The daemon does not enforce `disabled` or `restricted` native network -modes; combinations without required outer enforcement are unsupported. Never repoint a retained backend namespace or overwrite node identity -to bypass a configuration mismatch. - -Use the existing [standalone Core setup](../../README.md) for the database, -migrations, administrator credentials and encryption key. Model providers come from -the Session request, a saved Agent or the deployment default stored in Core; see -[model execution](../../../../contracts/agents-api/model-execution.md#deployment-defaults). The saved public -Core origin must be reachable from the guest; `localhost` in a microVM refers to -the guest itself. Do not configure a Core-local managed-runtimes file. - -## Runtime observations - -The configured microsandbox provider uses the same public Runtime observation rows -and Core Web Dashboard as Docker. Each request resolves Session, Environment, -allocation, installation and the persisted current compute generation before -invoking the helper. The helper performs one read-only `SandboxHandle.Metrics` -call; it does not wake suspended compute or extend retention. - -The current projection includes cumulative vCPU seconds, configured vCPU capacity, -guest memory usage/limit and compute uptime. Suspended or otherwise non-running -compute reports `unavailable` with `runtime_not_running`; metrics that are disabled -or have no current SDK sample report `sample_unavailable`. The SDK's instantaneous -CPU percentage, host RSS, disk, network and overlay measurements are not yet -exposed. Token usage continues to come from Session/Turn usage, not this provider. -An allocation without a persisted exact compute receipt also reports -`sample_unavailable`; the deterministic sandbox name is not sufficient to identify -one compute incarnation safely. - -Observation is operational evidence only. The idle suspension state machine uses -its durable activity and compute-phase records and never consults Dashboard samples. - -## Park, restore and verification - -Core first reserves the idle allocation. The daemon rejects suspension while -Turns, preparations, input receipts or file operations remain unsettled. After -its quiesce acknowledgement, it closes the old socket and parks. Core captures -and verifies the full snapshot before stopping the source VM and reclaiming RAM. - -Restore creates the precommitted next compute generation. Core invokes the -existing daemon binary inside that exact VM: - -```sh -oac-daemon resume --control-file /run/oac/daemon-suspend.json \ - --environment-id ENVIRONMENT_ID --suspend-id SUSPENSION_ID -``` - -This is Core's recovery action, not a second service. The command validates the -protected Environment/suspension identity and the daemon's PID/start time before -signalling through a Linux pidfd. The daemon reauthenticates and requires Core's -matching resume confirmation before admitting work. Each reconnect attempt is -limited to 30 seconds; transient failures retry the same suspension with backoff. -Authentication or protocol rejection terminates recovery. Ordinary disconnects -after resume retain the existing cleanup behavior. A failed capture may roll back the still-running -source through the same controlled wake path. - -For acceptance, complete a real-model Turn that writes a known file, wait for -suspension and confirm the source VM has stopped and released RAM. Submit the next -Turn in the same Session and verify the prior history, exact file bytes and frozen -configuration. Confirm setup ran once and no prior input or tool operation was -replayed. Exercise Core restart during capture/restore, pending-work exclusion, -credential revocation and Session deletion with both live and suspended compute. -Keep the resulting evidence separate from unit-test results. diff --git a/services/core/tools/e2b-provider/README.md b/services/core/tools/e2b-provider/README.md index d9beba248..75dae4699 100644 --- a/services/core/tools/e2b-provider/README.md +++ b/services/core/tools/e2b-provider/README.md @@ -1,138 +1,69 @@ -# Managed E2B SDK helper - -Core's Go adapter invokes this one-shot helper for Create, GetInfo, Renew, Kill -and initialization RunCommand. Daily execution and Files remain on the existing -Runtime connection. The helper uses the official E2B Python SDK 2.51.0; it does -not implement provider HTTP, envd RPC, a scheduler or a network service. - -Managed deployment validation uses a separate read-only helper request, bounded -to 30 seconds. The pinned SDK reads the selected template's build inventory and -requires the exact build UUID to be ready with the configured CPU and memory; a -selection without resources adopts the ready build's CPU and memory. It returns -the build's status, CPU, memory and reported disk size for Core to record with -the selection, and creates neither compute nor allocation receipts. - -The console's Core-key-only setup flow uses two more read-only helper requests. -`list_templates` pages the credential's visible templates through the official -`GET /v2/templates` SDK operation; `list_builds` -pages one selected template and returns ready exact build IDs and resources. -Both operations use the same explicit endpoint selectors and a transient API -key. They cap results at 200, never write a receipt, and cannot replace the -deployment write's exact-build validation. -Compatible endpoints must return the E2B SDK 2.51.0 template-list and -template-build response models. The helper does not adapt provider-specific -catalog shapes. - -Runtime observation uses a third read-only request, `observe`, for at most 100 -Runtime observation uses a separate read-only request, `observe`, for at most 100 -allocations. It reads each allocation's sandbox ID from its receipt without the -allocation lock, then runs one `GET /sandboxes/metrics` request and one labelled -listing of this installation's running sandboxes concurrently, within the -caller's deadline; the listing stops once every requested sandbox has appeared. -Only a sandbox that the listing confirms for exactly that allocation is -reported, with the listing's start time. A malformed metrics point makes only -its row unavailable. It never connects to, -renews or changes a sandbox and never writes receipts. See -[Runtime observability](../../../../contracts/agents-api/runtime-observability.md). Actual sandbox information -is checked before writing bootstrap credentials and on subsequent inspection; -resource drift still permits ownership-based cleanup. E2B disk capacity is not -an independently configurable limit. Sandbox inspection does not expose a build -UUID: build provenance comes from the validated immutable create selector. - -Credential replacement uses `verify_credential`, a read-only request with at most -32 Core allocation references. It verifies the fixed build, the SDK's paginated -team-owned template listing, and each settled live receipt against the labelled -sandbox listing. Template and sandbox scans each stop at 100 pages or the caller's -30-second deadline. Missing/unsettled receipts, repeated template cursors and -unconfirmed reads never authorize replacement. No receipt is changed. Core scans -retained generations and allocation pages under one bounded verification context, -then repeats verification while provider calls are fenced before credential commit. -Authentication rejection, ownership mismatch and uncertainty remain distinct fixed -codes. A readable public template is not proof that the key owns it. - -The private JSON boundary has version 1. Requests and credentials enter stdin; -stdout contains one bounded response with sanitized error codes. API keys never -enter arguments, inherited environment or receipts. The process retains its -allocation lock when the Core caller times out, until the bounded SDK operation -returns. Core must serialize lifecycle requests and never replay Create. -The request carries the deployment's explicit API origin and sandbox domain. -Every SDK call uses these selectors after ambient `E2B_*` variables are removed. -Receipts bind an allocation to those selectors; receipts written before this -feature belong to official E2B. Endpoint changes retain earlier generations on their original API and data-plane domain. The candidate credential must verify all retained ownership before an online switch. -Core tracks actual child exit even after caller timeout. Credential fencing -waits for those children without killing them; child exit itself never proves remote -Create settled. Core must serialize lifecycle requests and never replay Create. - -`StateDir` must already exist, be owned by the service user and have mode 0700. -Keep it on durable private storage through Core upgrades/restarts. Its receipts -contain provider connection credentials, ownership identities and one-shot -creation claims, not Core execution state. Losing this directory cannot authorize -recreation or successful cleanup. Do not delete receipts after an uncertain call. - -GetInfo uses SDK metadata/ID reads. SDK `connect` is never used because it can -resume paused compute. The SDK's version-pinned constructor restores clients -from private connection material; this deprecated constructor is intentionally -contained in `sdk.py` and covered by a no-connect/no-create test. An unknown -Create without connection material can be discovered and reclaimed, but cannot -resume bootstrap. Empty lookup cannot settle an unknown Create. Cleanup retains -every matching candidate and confirms exact-ID absence before writing a tombstone. - -`CreateSettled` proves the original Create/bootstrap can no longer mutate. It is -independent of `BootstrapComplete`, which acknowledges the protected initializer's -last step, not enrollment or native readiness. Explicitly settled absence returns -successful Info with `State=absent`; ordinary missing compute has no such proof. -An explicitly rejected Create with a settled receipt and no provider IDs proves -absence without another cloud request. GetInfo and Kill retain that rejection -receipt, so repeated recovery remains possible even when the API key is invalid. -Other receipts still require cloud discovery and ownership checks. -Unconfirmed initialization commands require reclaiming the whole allocation. - -## Build - -The [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) -builds the helper. The artifact contains only regular files and directories, with -executable permissions preserved, including the native `pyqwest` and -`protobuf-py-ext` wheels. `--check` needs no account credential. The installation -owns the durable receipt path independently of this immutable helper payload. - -`deploy/e2b/build-template.py` packages `init.py` and `managed_init.py` with the -qualified Runtime image. Existing templates without these files must be rebuilt. -The shared protected image preparation is used by both managed and self-hosted -startup; their credential formats and one-shot receipts remain separate. - -## Verification +# E2B Sandbox Provider helper + +Core's E2B Sandbox Provider ([`sandbox/e2b`](../../internal/sandbox/e2b)) is a pure-Go adapter that runs this one-shot Python helper for each lifecycle and read operation. The helper uses the official E2B Python SDK 2.51.0 ([`requirements.lock`](requirements.lock)); it implements no provider HTTP, envd RPC, scheduler or network service. E2B uses direct placement: there is no node, and each sandbox's daemon connects to Core over the public URL. Runtime execution and Files use that daemon connection. [Add a Sandbox Provider](../../../../docs/sandbox-provider.md) owns the provider contract this adapter implements. + +## Deployment + +The deployment selects E2B with an account key and an immutable `templateID:build_UUID`; [Sandbox deployment](../../../../contracts/agents-api/sandbox-deployment.md) owns the selection, key replacement and reset rules. Build the template with [`build-template.py`](../../deploy/e2b/README.md#build-a-template); it carries the `managed_init.py` startup script that Create runs, and a template without it is rejected as `template_invalid`. The template is deployment configuration, not a public Environment Template. + +The account key is stored encrypted in Core's database and is write-only. It reaches the helper only on standard input, never in a template, command argument, inherited environment, receipt or response. The helper, its dependency closure and licenses ship in the Core image and native Core; the [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) builds it. The server needs no Python. + +## Operations + +| Helper operation | Core use | Behavior | +| --- | --- | --- | +| `create` | `Create` | Creates the sandbox once and runs `managed_init.py`; see [Create](#create) | +| `inspect` | `GetInfo` | Reads the sandbox by recorded ID, or by ownership metadata when no ID is recorded, and checks ownership, domain, template and resources | +| `renew` | `Renew` | Extends the lease of the running sandbox to the configured timeout, then rereads it | +| `kill` | `Kill` | Destroys every matching sandbox and confirms that none remains | +| `command` | `RunCommand` | Runs one bounded command as the Runtime user on a running sandbox whose bootstrap completed; output is limited to 1 MiB per stream | +| `validate_deployment` | Deployment setup | Reads the template's builds and requires the exact build to be ready with the configured CPU and memory. Without configured resources the selection adopts the build's CPU and memory. Returns the build's status, CPU, memory and reported disk for Core to record; bounded to 30 seconds | +| `list_templates`, `list_builds` | Web's setup wizard | Pages the key's visible templates (`GET /v2/templates`) or one template's ready builds, with a transient key. Results are capped at 200 and write no receipt | +| `observe` | Runtime observations | Up to 100 allocations; see [Observations](#observations) | +| `verify_credential` | E2B key replacement | Up to 32 allocation references; see [Credential verification](#credential-verification) | + +Compatible endpoints must return the SDK 2.51.0 template-list and template-build response models; the helper does not adapt other catalog shapes. E2B has no independently configurable disk limit, and sandbox inspection does not expose a build ID: build provenance comes from the validated create selector. + +## Private JSON boundary + +The boundary has version 1. The request and credentials arrive on standard input; standard output carries one bounded response with a sanitized error code. The helper removes ambient `E2B_*` and `PYTHON*` variables and calls the SDK only with the request's explicit API origin and sandbox domain. Each receipt is bound to those selectors, and a receipt without them belongs to the official endpoints (`https://api.e2b.app`, `e2b.app`). An endpoint change keeps earlier generations on their original API and sandbox domain; the candidate key must verify all retained ownership before an online switch. + +## Receipts and state directory + +`OAC_E2B_STATE_DIR` ([configuration](../../../../docs/configuration.md#appendix-core-environment-without-the-installer)) must already exist, be owned by the helper's user and grant no group or other access. Keep it on durable private storage across Core upgrades and restarts. Its receipts hold SDK connection material, ownership identities and one-shot creation claims; they are not Core execution state. Losing the directory cannot authorize recreation or successful cleanup; never delete receipts after an uncertain call. + +A helper holds its allocation's lock until the SDK operation returns, even after Core's caller times out. Core tracks the helper's actual exit, and a credential change waits for running helpers without killing them. Helper exit never proves that a remote Create settled. Core serializes lifecycle requests per allocation and never replays Create. + +## Create + +1. Record `create_pending` with the bootstrap identity, then call `Sandbox.create` with the template, the configured timeout, the ownership metadata, `on_timeout=kill` and auto-resume disabled. A definite rejection records a settled `rejected` receipt with no sandbox IDs. +2. Record the sandbox ID and connection material, check the sandbox domain, then read the sandbox by ID and check its ownership metadata, template and resources before writing any credential. A mismatch records a settled rejection and returns `CreateSettled` with the error. +3. Check that `/opt/oac-e2b/managed_init.py` is readable, write the managed bootstrap input to `/root/.oac/e2b/managed-bootstrap.json` and run `managed_init.py` as root. + +`managed_init.py` prepares the image as the [application-managed startup](../../deploy/e2b/README.md#startup-and-security-boundary) does, writes the [Runtime bootstrap](../../../../docs/runtime-bootstrap.md) file to `/home/runtime/runtime-bootstrap.json` (mode 0600, owned by UID 1000), sets the Environment, Session and network variables and starts `oac-daemon connect --profile default --bootstrap-file /home/runtime/runtime-bootstrap.json` as UID/GID 1000. It records process handoff in `/root/.oac/e2b/managed-ready.json` and refuses to run again once any launch record exists. `BootstrapComplete` becomes true when a later inspection reads that record with the expected identity; it does not prove enrollment or native readiness. + +An unknown Create is never repeated. A Create whose connection material was lost can be discovered and destroyed but cannot resume bootstrap, and an unconfirmed startup requires reclaiming the whole allocation. + +## Inspection and cleanup + +Inspection uses SDK metadata and ID reads only. SDK `connect` is never used because it can resume paused compute. The version-pinned constructor that restores a client from saved connection material is confined to [`sdk.py`](sdk.py) and covered by a no-connect, no-create test. Resource drift fails inspection but still permits ownership-based cleanup. + +`CreateSettled` proves that the original Create and bootstrap can no longer mutate; it is independent of `BootstrapComplete`. An empty lookup never settles an unknown Create. Kill destroys every matching sandbox, confirms that none remains and only then records a settled tombstone; it returns `State=absent` with `CreateSettled`. A settled rejected Create with no sandbox IDs proves absence without a cloud request, so GetInfo and Kill still succeed when the key is invalid. Ordinary missing compute has no such proof. + +## Observations + +`observe` reads each allocation's sandbox ID from its receipt without taking the allocation lock. It then runs one `GET /sandboxes/metrics` request and one labelled listing of the installation's running sandboxes concurrently, within the caller's deadline; the listing stops once every requested sandbox has appeared. Only a sandbox that the listing confirms for exactly that allocation is reported, with the listing's start time. A malformed metrics point makes only its row unavailable. Observation never connects to, renews or changes a sandbox and writes no receipt. [Runtime observability](../../../../contracts/agents-api/runtime-observability.md) owns the field mapping. + +## Credential verification + +`verify_credential` checks the fixed build, the key's paginated team-owned template listing and each settled live receipt against the labelled sandbox listing. Template and sandbox scans each stop at 100 pages or the 30-second deadline. A missing or unsettled receipt, a repeated template cursor or an unconfirmed read never authorizes replacement, and no receipt changes. A readable public template does not prove that the key owns it. Core scans retained generations and allocation pages under one bounded verification context, then repeats the verification while provider calls are fenced, before it commits the new key. Authentication rejection, ownership mismatch and uncertainty return distinct fixed codes. + +## Build and tests + +The [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) builds the helper. The artifact contains only regular files and directories with executable modes preserved, including the native `pyqwest` and `protobuf-py-ext` wheels. `--check` needs no account credential. ```sh -python -m unittest discover -s services/core/tools/e2b-provider -p '*_test.py' -v -python -m unittest discover -s services/core/deploy/e2b -p '*_test.py' -v +make check-e2b-provider ``` -The build runs the first suite and validates the relocated helper's `--check` -report. These checks do not establish cloud authentication, native isolation or -real execution. Live acceptance must use owned E2B compute and the same Runtime, -with actual lease renewal, restart/unknown outcome reconciliation and confirmed -cleanup. Initializer and three-harness qualification remain deployment checks. - -## Adapter rules - -Core-managed E2B is a separate hosted deployment choice, using the official pinned -Python SDK through a packaged private helper. Do not restore the retired custom -HTTP/Connect or envd implementation. The adapter implements the same five operations; -cloud allocations use direct placement with no synthetic node, while Runtime execution -and file access keep the shared daemon contract. Its immutable Runtime template build -is deployment configuration, not a public Environment Template. Keep the account API -key encrypted in the database, write-only through admin input and absent from helper -arguments, logs, metadata and receipts. SDK connection materials and attempted-create -receipts belong in the private durable provider state directory; never replace missing -state to make cleanup appear successful. Create runs once. Unknown control-plane -outcomes remain blockers even if a listing is empty. Explicit matching-reference -CreateSettled evidence proves that the original initialization cannot mutate further; -confirmed absent compute may then be released. Ordinary 404 responses do not prove it. -The helper's pinned SDK, dependencies and licenses ship with Core; users do not install -Python packages after selecting E2B in Web. Application-managed self_hosted tooling -remains independent and uses the same Runtime. Qualify each changed path using actual -provider and model execution before claiming acceptance. Run `make check-e2b-provider` -with `OAC_TEST_E2B_SDK_PYTHON` pointing to the pinned SDK environment; the packaged -helper build runs the provider tests as well. The SDK gate also covers the -application-managed launch tests. `make check` covers shared initialization and -managed initialization using only the Python standard library. +With `OAC_TEST_E2B_SDK_PYTHON` pointing at the pinned SDK environment, this runs this directory's tests and the [template scripts' tests](../../deploy/e2b/README.md#tests). The helper build runs this directory's suite and checks the relocated helper's `--check` report. These tests create no cloud resources. diff --git a/services/core/tools/microsandbox-provider/README.md b/services/core/tools/microsandbox-provider/README.md index 27e544d82..554b4fad8 100644 --- a/services/core/tools/microsandbox-provider/README.md +++ b/services/core/tools/microsandbox-provider/README.md @@ -1,234 +1,64 @@ -# microsandbox provider helper - -This Linux-only, one-operation helper links the maintained microsandbox Go SDK -v0.7.2. Core stays a pure-Go binary and uses the provider-neutral -`sandbox.CheckpointProvider` interface. There is no helper daemon, local lifecycle -database, native agent adapter, or additional scheduler. - -Managed creation checks the native CPU, memory, root disk and owned Environment -disk configuration before bootstrap. Inspection and restore reject resource or -image drift while ownership-based compute and snapshot deletion remain available. -Full snapshots record a resource proof only after the source's limits match. -The pinned native restore leaves managed root size unspecified in its config; -the verified snapshot ancestry and matching proof establish inherited capacity. -The restored target keeps that proof after its CPU, memory and Environment disk -are checked. This does not claim a new direct root-capacity measurement on restore. - -The ordinary standalone sandbox node runs this helper natively on Linux amd64, -under a dedicated service user with KVM access. Core owns lifecycle intent through -the node protocol and can run in a container or on another host. Core's -container image does not execute this helper; only the node does. Core packaging, -container or native, is independent of the provider. - -PostgreSQL owns the provider, per-sandbox resources and immutable Runtime release. -The node installs that specification and retains its generation and digest; local -provider files cannot override it. See the -[deployment contract](../../../../contracts/agents-api/sandbox-deployment.md) -and [nodes guide](../../../../docs/getting-started/nodes.md). Core no longer accepts a -file-managed startup selection or automatically adopts an older file-managed -database. - -## Build and installation - -The [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) -builds and tests the helper. The relative replacement for the parent Core module -refers to this checkout. The SDK is the real published module, pinned in go.mod -and go.sum, with no local-source replacement. - -Install the matching v0.7.2 msb runtime and firmware from checksum-verified release -artifacts. Supply absolute helper/runtime/firmware paths and expected SHA256 -values in the trusted provider Config. The helper checks the runtime and firmware -hashes, the ELF runtime version section, the SDK module version, and the resolved -local backend. It never auto-installs or upgrades these artifacts. The installed -paths must remain immutable for the lifetime of the provider key. - -Use a dedicated, private (0700), short MSB_HOME on local persistent storage. -The upstream runtime uses Unix sockets, so a short path such as -`/var/lib/oac-msb` avoids pathname limits. Never share that home with another -installation, cloud profile, or manual lifecycle controller. Its VM disks, -snapshots, SDK state and credentials are confidential execution-service data. -All managed sandbox lifecycle mutations must go through the provider. - -An immutable OCI image must already be available as `repository@sha256:<64 lowercase hex>`. Bare Docker image IDs and mutable tags are rejected. -For offline archives, `msb image load --tag repository@sha256:` must register -the digest reference explicitly; loading a mutable tag alone does not create it. -Use the manifest digest from `image inspect`, not the Docker image config ID. -The qualified image contains our existing daemon, Python 3, native harness and -shared Runtime helpers. Provider Config carries the saved VM resources and the -node's explicit host network policy. Its Runtime and firmware hashes must match -the database-owned release. This adapter does not install registry credentials. - -## Bootstrap and network - -The SDK creates a private owned ext4 disk at `/environment`, with explicit -`environment_disk_mib` capacity. Workspace, staging and generated outputs share -that filesystem, preserving the existing cross-device and link checks. The -layered root filesystem can report different device IDs for directories and -upper-layer files and is not used for workspace storage. Native full snapshots -and sandbox removal capture, restore and reclaim the owned disk; no host path or -external volume lifecycle is introduced. Creation starts in `/` until bootstrap -creates the workspace directories. - -VM creation does not implicitly run the OCI ENTRYPOINT. Before admitting native -work, the helper delivers the [Runtime bootstrap input](../../../../docs/runtime-bootstrap.md) -through confidential stdin to a protected file, creates Runtime directories, binds -the workspace at `/workspace`, and launches `connect --bootstrap-file` in -background mode as uid/gid 1000. The final provider-owned bootstrap label confirms only -completion of these writes and launch, not authentication or native readiness. -The receipt label uses the supported next-start modification policy to update -persisted metadata without restarting the guest. v0.7.2 cannot update active labels. - -A private helper response can carry `CreateSettled` with a configuration rejection. -This proof is emitted only after native Create has positively completed, the first -inspection has verified the exact created compute ID and ownership, and resource -qualification rejects it before bootstrap begins. The adapter preserves the -original error and validates that initial compute identity before passing the -proof to Core. It does not mark the sandbox ready. Ordinary inspection, uncertain -Create outcomes, timeouts and ownership failures cannot acquire this proof. Core -still requires authorized, ownership-checked cleanup before releasing the -allocation; missing compute alone never proves creation settled. - -The private daemon control directory is `/run/oac` (0700, uid/gid 1000). -`OAC_RUNTIME_DAEMON_SUSPEND_PID_FILE=/run/oac/daemon-suspend.json` enables the -daemon's separately owned idle park/wake control. RunCommandCompute can execute -the exact daemon resume command authorized by Core; it never uses pkill. - -Creation and full restore both receive the same explicit, trusted host policy. -Upstream restore defaults to Public rather than inheriting source host access. -Native tool-network policy remains the existing Runtime responsibility. No -default external mount inheritance, missing-resource allowance, or policy widening -is used. - -## Lifecycle contract - -Core persists operation IDs, source/target generations, exact identities and -snapshot evidence before depending on them. `Initial` and `NewCompute` only -construct references; they allocate nothing. Names are installation/allocation/ -generation-derived and must never be reused for another incarnation. - -Suspend pauses the exact VM, captures a full snapshot under the persisted -operation's derived group/member, verifies its complete checkpoint closure, then -force-stops the source. Pausing alone is not suspension or memory reclamation. -A completed matching artifact is inspected rather than captured again. - -After a lost response, Core uses `ObserveOnly`. This never starts capture or -restore, and a suspended-operation observation never kills its source. Core can -persist recovered snapshot evidence before KillCompute. Observation verifies -artifact integrity and source ownership independently of execution qualification, -so resource drift cannot hide a retained artifact from cleanup. If the artifact -is absent but the exact source is still running or paused with settled bootstrap, -observation returns that source with no snapshot. Core can then abort suspension; -thawing and subsequent execution still require resource qualification. Missing -state never authorizes replay of the original operation. - -Restore verifies the exact artifact and creates the precommitted target name. -Existing targets are adopted only when their immutable ID (if known) and -persisted `snapshot_parent` agree. Fresh restore and retry share completion: -verify the original artifact and resource proof, inspect the running target's -actual resources and ancestry, persist its missing derived resource-proof label, -then strictly reread the same native ID. `ObserveOnly` may finish this receipt -after an interrupted restore; it cannot restart a stopped target, change resources -or issue another Restore. A conflicting proof remains an error. Native restore -does not inherit source ownership labels; ancestry supplies that evidence. -Unfinished restore intent is not bootstrap completion. There is no ordinary Start, -replacement, disk-only restore, or cold-boot fallback. - -`ResumeCompute` only thaws the same resident source after an aborted suspension. -The pinned Go handle method is name-based; the allocation flock plus ID checks -before and after the call fence every managed replacement. External manual -lifecycle changes in the managed namespace are unsupported. - -A helper holds an allocation flock until its lifecycle SDK call actually settles. -Core's response deadline does not kill that helper or cancel its FFI wait, since -cancelling the wait does not prove the native mutation stopped. On timeout Core -retains an unknown operation and observes it; a subsequent helper cannot race -past the surviving lock holder. A stuck owner needs operator investigation, -not automatic lock deletion or another create. - -KillCompute checks the precise incarnation before stopping and removing its writable -disks. Core calls it after persisting the verified snapshot, even if Suspend -already stopped the source. GetCompute, commands, cleanup and the next suspension -verify restored provenance from persisted VM config after the consumed artifact -has been deleted. -DeleteSnapshot accepts only the derived operation selector and matching full -artifact identity, not an arbitrary path. Core owns retention, consumed snapshot -generations and cleanup ordering. A checkpoint must never roll back work admitted -after its first restore. - -RunCommand/RunCommandCompute carry stdin on anonymous pipes, fix the guest user -to 1000, impose a deadline and 1 MiB per-stream output limits, and require both -successful stdin completion and an explicit guest exit before returning a result. -Timeouts, output overflow and missing receipts return ErrCommandUnconfirmed. -Closing an SDK exec handle alone does not prove the guest process exited. - -The read-only metrics operation holds the allocation lock and verifies the exact -compute ID through the SDK before and after `msb metrics NAME --format json`. -The CLI is the same checksum/version-pinned runtime binary. Its registry report -preserves the native sample timestamp and fractional-second uptime; parsing both -at native millisecond precision reconstructs the real run start consistently -across polls and Core restarts, while new runs have a new start. The Go SDK v0.7.2 -projection drops that timestamp and truncates uptime to whole seconds, so it -cannot supply this fence. Sandbox creation time is not a run start time. - -The helper returns only the native observation time, exact uptime, cumulative -vCPU time and guest memory usage/limit. It rejects stale/exited reports and -malformed or missing fields. CLI output is bounded; command failures expose only -an unavailable code, never native diagnostics. Observation does not connect to -the guest, renew activity, resume paused compute or mutate lifecycle state. - -The pinned source evidence is tag `v0.7.2`, commit -`1c59b8dbf0ad47dda2f807c0214b529aceb81c74`: `crates/metrics/lib/registry.rs` -reads `sampled_at_unix_ms` and `started_at_unix_ms` coherently and subtracts them -for uptime; `crates/cli/lib/commands/metrics.rs` serializes the timestamp and -`uptime.as_secs_f64()`. The private native fields and identifiers do not cross -the helper boundary. - -## Acceptance boundary - -The feature is idle-only: Core must reserve a terminal Session with no pending -work before parking its daemon and suspending compute. The next Turn uses the -same Session history, files and configuration without replaying initialization. -This adapter does not promise that a native agent process persists across Turns. - -The following evidence describes earlier bounded single-host qualifications. It -does not qualify the current database-managed node installation or every resource -profile; those require their own acceptance evidence. - -SDK feasibility separately qualified exact IDs, full snapshot verification, -snapshot_parent, source memory release, tmpfs/RAM restoration and stale-handle -pause rejection. The published Go module's SDK sources and Linux amd64 FFI were -compared byte-for-byte with that qualification source/runtime. The FFI SHA256 was -`ed04ca4788c1c400e1b67040e964fbfcb3743d31d969d1b426dfb95afe7271d8`. - -Production helper qualification completed two full capture/restore cycles, removing -each source before restore and deleting each consumed artifact before the next -cycle. Workspace bind identity, uid 1000, private auth permissions and file content -survived. The real Core/Codex synthetic-model test separately completed two Turns -across idle suspension and a Core restart. Synthetic responses establish the -control flow, not real-model acceptance. - -The 2026-09-22 Linux amd64 live acceptance used microsandbox/Go SDK v0.7.2, -libkrunfw 5.6.1, Codex CLI 0.153.4 and the official Kimi K3 Responses API. The -unchanged native harness completed two real Turns in one Session across automatic -idle suspension, source removal and a Core restart. Generation 1 restored the -exact full snapshot; the consumed artifact was removed. The model's shell tool -read the original random file marker, retained environment configuration and one -initialization record, then correctly recalled the first request from history. -Retrying the second public input with the same idempotency key created no extra -Turn or tool side effect. A subsequent public file upload/list woke generation 2 -without starting another Turn. Public Session deletion released its allocation; -all qualification VMs and snapshots and temporary credentials were removed. -That run qualified its historical model/profile, not the current installation -path, every provider or broader isolation guarantees. - -The source VM was observed at 334304 KiB RSS before suspension and with zero RSS -and no executable after exit. The qualification container's PID 1 left a zombie -PID; it retained no VM memory. The separate backend RAM probe also checked live -anonymous memory and tmpfs contents after full restore. A paused resident VM alone -would not satisfy either memory-release check. - -Private qualification evidence is grouped as `ram-03` (backend), `provider-01` -(production helper), `core-01` (synthetic Core) and `live-core-03` (real model). -These are bounded single-host checks. They do not establish cold-image download -latency, tail latency, production capacity, or general native-process persistence -across Turns. +# microsandbox Sandbox Provider helper + +microsandbox runs each hosted Session in its own microVM on a Linux amd64 node with KVM, and it is the Sandbox Provider that supports idle suspension. Core forwards provider operations to the node over the [node protocol](../../../../contracts/agents-api/node-generation-protocol.md); the node's adapter ([`sandbox/microsandbox`](../../internal/sandbox/microsandbox)) runs this helper once per operation. The helper links the microsandbox Go SDK v0.7.2 with its FFI library, so Core and the node program stay CGO-free Go binaries. It implements the provider-neutral `sandbox.CheckpointProvider` with full snapshots. It has no daemon, lifecycle database, scheduler or network control plane. + +[Add a Sandbox Provider](../../../../docs/sandbox-provider.md) owns the provider contract. [Sandbox deployment](../../../../contracts/agents-api/sandbox-deployment.md) owns the resources, Runtime release and suspension policy; the [nodes guide](../../../../docs/getting-started/nodes.md) owns node installation, host requirements, the node's directories and its network policy. + +## Installation checks + +The [maintainer guide](../../../../docs/maintainers.md#runtime-images-and-helpers) builds the helper and packages the checksum-verified `msb` runtime and `libkrunfw` firmware. The node's provider configuration, written by the node installer, supplies the absolute helper, runtime and firmware paths with their SHA-256 values, the runtime home, the Runtime image reference, the saved resources and the host network policy. + +Before every operation the helper checks that it was built with the published SDK module v0.7.2 without a replacement, that the runtime and firmware match their hashes, that the runtime's `.msbver` ELF section reports the SDK version and that the SDK resolves exactly those paths with the local backend ([`main.go`](main.go)). It never installs or upgrades these files; keep them unchanged for the lifetime of the provider's backend. Ambient SDK profiles are ignored. + +The runtime home must be private (mode 0700), short, on local persistent storage and used by no other installation, profile or manual lifecycle tool. microsandbox uses Unix sockets there, so the node installer refuses a home whose path would exceed their limit. The home holds confidential VM disks, memory snapshots and SDK state; preserve it with the node identity and Core's database when recovering a host. Every managed lifecycle change goes through the provider. + +The Runtime image is an immutable `repository@sha256:<64 lowercase hex>` reference that matches the saved Runtime release; bare image IDs and mutable tags are rejected. The node installer imports the distribution's image under that reference. To load an image by hand, `msb image load --tag repository@sha256:` must register the digest reference explicitly, with the manifest digest from `image inspect`, not the Docker image config ID. The image carries the daemon, Python 3, the native Harnesses and the shared Runtime helpers. The provider installs no registry credentials. + +## Create and bootstrap + +Create names the VM from a hash of the installation and allocation reference plus the compute generation; a name never serves another incarnation. It creates the VM with the saved CPUs and memory as both initial and maximum, a managed root disk of `root_disk_mib`, an owned ext4 disk of `environment_disk_mib` mounted at `/environment`, user 1000:1000, working directory `/` and the node's network policy ([`bootstrap.go`](bootstrap.go)). The resource checks run before bootstrap. + +Workspace, staging and outputs share the `/environment` filesystem, which keeps the Runtime's cross-device and link checks intact; the layered root filesystem can report different device IDs for a directory and its upper-layer files, so it holds no workspace data. Full snapshots and sandbox removal capture, restore and reclaim this disk; there is no host path, external mount or separate storage lifecycle. + +VM creation does not run the image's entry point. The bootstrap runs as root through confidential standard input, with a two-minute limit. It creates the Runtime directories and the private control directory `/run/oac` (mode 0700, owned by UID 1000), writes the [Runtime bootstrap](../../../../docs/runtime-bootstrap.md) file to `/home/runtime/runtime-bootstrap.json`, bind-mounts `/environment/workspace` at `/workspace` and starts `oac-daemon connect --bootstrap-file` in the background as UID/GID 1000. `OAC_RUNTIME_DAEMON_SUSPEND_PID_FILE=/run/oac/daemon-suspend.json` enables the daemon's park and wake control. The `io.oac.bootstrap` label then changes from `pending` to `complete` through the SDK's next-start modification policy, because v0.7.2 cannot update the labels of a running VM. That label confirms only these writes and the launch, not authentication or native readiness. + +A helper response carries `CreateSettled` with a configuration rejection only after native Create has completed, the first inspection has verified the exact compute ID and ownership, and the resource check has rejected the VM before bootstrap started. The adapter keeps the original error and validates the compute identity before passing the proof to Core. Ordinary inspection, uncertain Create outcomes, timeouts and ownership failures never produce it, and missing compute alone never proves that Create settled. + +## Checkpoint lifecycle + +Core persists operation IDs, source and target generations, exact identities and snapshot evidence before it depends on them. `Initial` and `NewCompute` only construct references and allocate nothing. + +- **Suspend** pauses the exact VM, captures a full snapshot under the persisted operation's derived group and member, verifies the complete checkpoint closure, then force-stops the source. Pausing alone does not release memory. A completed matching artifact is inspected instead of captured again. A full snapshot records a resource proof only after the source's limits match. +- **KillCompute** checks the precise incarnation before it stops the VM and removes its writable disks. Core calls it after it has persisted the verified snapshot, even when Suspend already stopped the source, so no chain of old writable disks grows across suspension cycles. +- **Restore** verifies the exact artifact and creates the precommitted target name. An existing target is adopted only when its immutable ID, if known, and its persisted `snapshot_parent` agree. Upstream restore defaults to public networking, so restore passes the same explicit host policy as creation, and no undeclared host resource or mount is inherited. +- Native restore leaves the managed root size unset because the target inherits the verified full snapshot. The helper accepts that only with a matching snapshot resource proof and the exact source and target identities, and it checks the target's CPU, memory and Environment disk before keeping the inherited proof. A missing root size never counts as unlimited capacity, and retained state is never resized. +- Fresh restore and retry share one completion: verify the original artifact and resource proof, inspect the running target's resources and ancestry, persist its missing derived resource-proof label, then strictly reread the same native ID ([`restore_completion.go`](restore_completion.go)). A conflicting proof is an error. Native restore does not copy the source's ownership labels; ancestry supplies that evidence. There is no ordinary Start, replacement, disk-only restore or cold boot. +- **ResumeCompute** thaws the same resident source after an aborted suspension. The pinned SDK handle method is name-based; the allocation lock and ID checks before and after the call fence every managed replacement. Manual lifecycle changes in the managed namespace are unsupported. +- **DeleteSnapshot** accepts only the derived operation selector and the matching full artifact identity, never an arbitrary path. Core owns retention, consumed snapshot generations and cleanup order. A checkpoint never rolls back work admitted after its first restore. + +After a lost response Core uses `ObserveOnly`. It never starts a capture or restore, and observing a suspend operation never kills its source. Observation checks artifact integrity and source ownership independently of resource checks, so resource drift cannot hide a retained artifact from cleanup; Core can persist recovered snapshot evidence before KillCompute. If the artifact is absent but the exact source is still running or paused with a settled bootstrap, observation returns the source without a snapshot and Core can abort the suspension; thawing and further execution still require the resource checks. For an interrupted restore, `ObserveOnly` may finish the missing resource proof on the exact target but never restarts a stopped target, changes resources or restores again. Missing state never authorizes a replay. + +GetCompute, commands, cleanup and the next suspension verify restored provenance from the persisted VM configuration after the consumed artifact is deleted. + +## Locks and commands + +A helper holds a per-allocation lock, under `oac-locks/` in the runtime home, until its SDK call actually settles. Core's response deadline neither kills the helper nor cancels its FFI wait, because cancelling the wait does not prove that the native mutation stopped. On a timeout Core keeps an unknown operation and observes it; a later helper cannot pass the surviving lock holder. A stuck owner needs operator investigation, not lock deletion or another Create. + +`RunCommand` and `RunCommandCompute` pass standard input on an anonymous pipe, run as UID 1000 with a deadline and a 1 MiB limit per output stream, and return a result only after the input is fully written and the guest reports its exit. Timeouts, output overflow and missing receipts return `ErrCommandUnconfirmed`; closing an SDK exec handle does not prove that the guest process exited ([`command.go`](command.go)). Core uses `RunCommandCompute` only to run the exact wake command it authorizes: + +```sh +oac-daemon resume --control-file /run/oac/daemon-suspend.json \ + --environment-id ENVIRONMENT_ID --suspend-id SUSPENSION_ID +``` + +The command checks the protected Environment and suspension identities and the parked daemon's PID and start time, then signals it through a Linux pidfd. The daemon reauthenticates and waits for Core's resume confirmation before it admits work. + +## Metrics + +The read-only metrics operation holds the allocation lock and verifies the exact compute ID through the SDK before and after it runs `msb metrics NAME --format json` with the same pinned runtime binary ([`metrics.go`](metrics.go)). The CLI report keeps the native sample timestamp and fractional-second uptime, so the helper reconstructs one run start consistently across polls and Core restarts, and a new run gets a new start. The Go SDK's projection drops the timestamp and truncates uptime to whole seconds, so it cannot provide this; sandbox creation time is not a run start time. In the pinned source (`v0.7.2`, commit `1c59b8dbf0ad47dda2f807c0214b529aceb81c74`), `crates/metrics/lib/registry.rs` reads `sampled_at_unix_ms` and `started_at_unix_ms` together and subtracts them for uptime, and `crates/cli/lib/commands/metrics.rs` serializes the timestamp and `uptime.as_secs_f64()`. + +The helper returns only the native observation time, exact uptime, cumulative vCPU time and guest memory usage and limit. It rejects stale or exited reports and missing or malformed fields, bounds the CLI output and reports failures only as an unavailable code, never native diagnostics. Metrics never connect to the guest, renew activity, resume paused compute or change lifecycle state. [Runtime observability](../../../../contracts/agents-api/runtime-observability.md) owns the mapping to observations. + +## Tests + +`make check-microsandbox-provider`, part of `make check`, runs the pure-Go adapter tests everywhere and this module's tests on Linux; other hosts print an explicit skip for the Linux-only module. From b7599864b6f8d2c37c818c93fa72688aaeb55389 Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 08:43:15 +0000 Subject: [PATCH 03/10] docs: keep the Core-Runtime protocol to wire rules The protocol document gains the capability declarations with Core's admission rule for each, the opt-ins Core sets, the active-input receipt timers and native failure classification (from the deleted classification contract). The Executor and Turn adapter lifecycle moves to Harness onboarding; preparation, image and MCP restatements become links. Runtime bootstrap names its Go projection and links every credential source. --- .../agents-api/native-error-classification.md | 50 -- docs/design-principles.md | 2 +- docs/runtime-bootstrap.md | 44 +- docs/runtime-protocol.md | 678 ++++-------------- 4 files changed, 171 insertions(+), 603 deletions(-) delete mode 100644 contracts/agents-api/native-error-classification.md diff --git a/contracts/agents-api/native-error-classification.md b/contracts/agents-api/native-error-classification.md deleted file mode 100644 index 25a43f000..000000000 --- a/contracts/agents-api/native-error-classification.md +++ /dev/null @@ -1,50 +0,0 @@ -# Native failure classification - -Runtime adapters may add `code` and `http_status` to a prompt-level `error` frame. -The fields are optional neutral metadata, not a new terminal event or a public -Agents API error contract. Existing error text, Usage, Done, native Session identity -and cancellation receipts retain their existing ownership and order. - -Core stores accepted values in the Turn outcome as `engine_error_code` and -`engine_http_status`. Classification is subordinate to terminal status and Core's -`error_code`: it cannot turn a completed or cancelled Turn into a failure, hide an -incomplete event stream, or override persistence and cancellation-receipt failures. -Normal delivery and terminal journal draining use the same extraction rule. - -Accepted codes are `authentication_error`, `rate_limit_exceeded`, -`usage_limit_exceeded`, `server_overloaded`, `server_error`, `invalid_request`, -`resource_not_found`, `request_timeout`, `context_length_exceeded`, `cyber_policy` -and `connection_failed`. Core diagnostics expose these categories only for failed -`engine_failed` Turns. Their params are empty except `connection_failed`, whose -`http_status` is a valid integer or null. Only `connection_failed` retains an integer HTTP status -in 100..599; all other status metadata is discarded. Missing, malformed and unknown -optional values degrade to unclassified metadata without discarding Usage or Done. -Old Runtime frames remain generic harness failures. Old readers ignore new fields. -The catalog is not a claim that every harness can produce every classification. - -Codex 0.153.4 uses only the exact root terminal Turn's `codexErrorInfo`, never -notification text or a retry notification. Its finite string mappings cover -unauthorized, usageLimitExceeded, rateLimitExceeded, contextWindowExceeded, -serverOverloaded, internalServerError, badRequest and cyberPolicy. The connection -object variants preserve valid upstream status. responseTooManyFailedAttempts maps -to rate_limit_exceeded only for 429; otherwise it is connection_failed. Session -budget, misalignment policy, rollback, sandbox and steering-state errors remain -unclassified. A successful terminal Turn does not inherit an earlier error. - -Claude SDK 0.3.269 records finite root assistant errors against current submitted -input UUIDs and the same native Session. Only a matching unsuccessful native result -commits that candidate. Replayed, synthetic, child, foreign and completed-input -messages cannot classify the root failure; recovered results clear candidates. -Authentication variants map to authentication_error, billing_error to -usage_limit_exceeded, rate_limit to rate_limit_exceeded, overloaded to -server_overloaded, invalid_request to invalid_request, model_not_found to -resource_not_found and server_error to server_error. Unknown/max-output-token, -max-turn/budget/structured-retry results and unstructured exceptions remain generic. -The bridge retains native usage and identity and rejects stale classification after -malformed output or an unsuccessful bridge process exit. A validated native failure -may itself have a nonzero native child exit; that is distinct from bridge failure. - -Neither adapter classifies error prose, exposes provider messages through safe -metadata, nor synthesizes tool timeouts. request_timeout has no current producer. -MiniMax remains unclassified. Real provider rejection evidence is distinct from -controlled protocol/SDK-stream tests and must be qualified against a matched Runtime. diff --git a/docs/design-principles.md b/docs/design-principles.md index 1452a2753..4df9ed503 100644 --- a/docs/design-principles.md +++ b/docs/design-principles.md @@ -87,4 +87,4 @@ execution or compatibility with old private protocols. Implementers follow the Native failure classification is adapter-owned and uses finite, structured native values. Optional Runtime error metadata is normalized once; it never replaces Core's terminal authority, cancellation receipts, Usage or native identity. See -[native failure classification](../contracts/agents-api/native-error-classification.md). +[native failure classification](runtime-protocol.md#native-failure-classification). diff --git a/docs/runtime-bootstrap.md b/docs/runtime-bootstrap.md index e65810f2a..9a1ca5e96 100644 --- a/docs/runtime-bootstrap.md +++ b/docs/runtime-bootstrap.md @@ -1,13 +1,10 @@ # Runtime bootstrap -This document owns the Provider-to-Runtime startup boundary. The input type and -validator live in [runtimebootstrap](../internal/runtimebootstrap/bootstrap.go). -Providers must not read or write Runtime's private authentication store. +A Sandbox Provider starts a managed Runtime by handing it one bootstrap file. This document owns that Provider-to-Runtime startup input. The type and validator live in [`internal/runtimebootstrap`](../internal/runtimebootstrap/bootstrap.go); Go providers build it with `sandbox.Bootstrap.RuntimeConnection()` in [`runtime_bootstrap.go`](../services/core/internal/sandbox/runtime_bootstrap.go), and SDK helpers forward the serialized object unchanged. A provider never reads or writes the Runtime's private authentication store. ## Launch input -Deliver one JSON object in a regular file accessible only to the Runtime account -and trusted provisioning processes (0600 on managed Linux). Pass its absolute path: +Deliver one JSON object in a regular file that only the Runtime account and trusted provisioning processes can read (mode 0600 on managed Linux), and pass its absolute path: ```sh oac-daemon connect --bootstrap-file /home/runtime/runtime-bootstrap.json @@ -15,42 +12,23 @@ oac-daemon connect --bootstrap-file /home/runtime/runtime-bootstrap.json | Field | Meaning | | --- | --- | -| `version` | Exact bootstrap version declared by `runtimebootstrap.Version` | +| `version` | The exact bootstrap version, `runtimebootstrap.Version` | | `core_url` | HTTP(S) machine API base ending in `/api/v1`, without credentials, query or fragment | -| `device_id` | Canonical nonzero UUID for the daemon identity issued by Core | -| `credential` | Nonempty daemon credential issued by Core, without whitespace or NUL | +| `device_id` | Canonical nonzero UUID of the daemon identity Core issued | +| `credential` | Nonempty daemon credential Core issued, without whitespace or NUL | -The decoder rejects unknown, duplicate, missing and case-aliased fields, other -versions and documents exceeding `runtimebootstrap.MaxBytes`. Errors exclude -submitted values. Go providers use `Bootstrap.RuntimeConnection()`; SDK helpers -forward the serialized object without defining their own authentication format. +The decoder rejects unknown, duplicate, missing and case-aliased fields, other versions and documents larger than `runtimebootstrap.MaxBytes` (16 KiB). Errors never include submitted values. A missing or malformed file fails before the daemon connects. -The file is the sole authentication input for this launch. It cannot be combined -with pairing or self-hosted enrollment flags. Credentials never go in command -arguments, environment variables or receipts. The provider retains the protected -file for process restarts and removes it only with explicit owned-resource cleanup. -Runtime reads it into memory and neither overwrites nor falls back to a private -auth profile. A missing or malformed file fails before connecting. +The file is the only authentication input for this launch: the daemon refuses to combine it with pairing or self-hosted enrollment options, and reads the credential into memory without saving it to a stored profile. Credentials never go in command arguments, environment variables or receipts. The provider keeps the file for process restarts and removes it only during explicit cleanup of the resources it owns. ## Responsibilities and readiness -The Provider provisions the account, mounts and workspace, delivers this input, -sets the existing Runtime resource and Environment binding settings, and starts -the daemon as the unprivileged Runtime account. Docker supplies a file in its -owned home volume; microsandbox and E2B deliver it before launching the same -command. These are delivery mechanisms, not different bootstrap protocols. +The provider creates the account, mounts and workspace, delivers this file, sets the Runtime's resource and Environment binding settings, and starts the daemon as the unprivileged Runtime account. Docker writes the file into the Runtime's owned home volume; microsandbox and E2B deliver it before launching the same command. -Runtime validates the input and owns authentication and connection establishment. -A successful process launch proves only handoff; authenticated connection, -capability preparation and execution readiness remain separate observations under -the [Core–Runtime protocol](runtime-protocol.md). +The Runtime validates the input and owns authentication and connection. A successful launch proves only the handoff: an authenticated connection, prepared capabilities and execution readiness are separate observations under the [Core–Runtime protocol](runtime-protocol.md), and the [Sandbox Provider guide](sandbox-provider.md#four-distinct-readiness-facts) lists what each one proves. -Self-hosted enrollment exchanges its executor credential for a daemon identity -through the machine API; it is a different source of authority, not a managed -bootstrap-file fallback. Both paths enter the same Runtime execution loop. +Self-hosted executors and operator-provisioned devices get their daemon identity in other ways; the [machine connection API](../contracts/agents-api/machine-api.md#credentials) lists every credential source. All of them enter the same Runtime execution loop. ## Verification -`go test ./internal/runtimebootstrap ./apps/daemon/internal/cli` covers the -input contract, credential-source exclusivity and restart behavior. Provider tests -verify delivery and permissions without relying on private Runtime storage. +`go test ./internal/runtimebootstrap ./apps/daemon/internal/cli` covers the input contract, the exclusivity of credential sources and restart behavior. Provider tests verify delivery and file permissions without relying on the Runtime's private storage. diff --git a/docs/runtime-protocol.md b/docs/runtime-protocol.md index 96681656b..797e66e98 100644 --- a/docs/runtime-protocol.md +++ b/docs/runtime-protocol.md @@ -1,247 +1,94 @@ # Core–Runtime protocol -This is the integration entry point for a Runtime that executes work for Core. -The wire definitions live once in -[`internal/agentdaemon/proto`](../internal/agentdaemon/proto); -Core's [gateway](../internal/agentdaemon/gateway) and the reference Runtime's -[dispatcher](../apps/daemon/internal/dispatch) both use them. -The [machine HTTP API](../contracts/agents-api/runtime.openapi.yaml) describes -registration and connection endpoints. This document defines the meaning and -ordering of the messages after connection; it does not replace the typed payloads. - -Use this protocol for hosted and self-hosted Runtime implementations. A new -Harness implements the [adapter contract](../contracts/agents-api/harness-onboarding.md) -behind the Runtime registry. Do not add a Core orchestration branch named after -the Harness, operating system or Sandbox Provider. +This protocol connects Core to a Runtime daemon after the daemon has its machine credential. It defines the meaning and order of the messages on the daemon connection. The wire types, limits and validators live once in [`internal/agentdaemon/proto`](../internal/agentdaemon/proto); Core's [gateway](../internal/agentdaemon/gateway) and the reference Runtime's [dispatcher](../apps/daemon/internal/dispatch) both use them, so there is no second payload schema to keep in sync. The HTTP routes that issue credentials and open the connection are in the [machine connection API](../contracts/agents-api/machine-api.md). + +Hosted and self-hosted Runtimes use the same protocol. A Harness joins through the [Harness adapter contract](../contracts/agents-api/harness-onboarding.md), which owns the Executor and Turn lifecycle obligations behind the Runtime registry. ## Ownership and connection -Core owns durable Session, Turn, input and Environment records, scheduling and -reconciliation. Runtime owns native Executors, active Turns, transfer state and -cleanup until settlement. A Sandbox Provider owns placement and the surrounding -compute lifecycle. Releasing an execution admission or closing an Executor does -not delete, suspend or reclaim a sandbox. - -An installed daemon runs with the authority of the user who starts it. The -protocol does not make the daemon a tool, filesystem or network isolation -boundary. Core-managed Docker, E2B or other sandboxes provide the outer isolation. -A logical workspace, native tool configuration, advertised capability or successful -preparation is not proof of containment. - -1. Obtain the appropriate machine credential using the documented registration - flow. Applications use Project API keys; operators use Core keys. Neither is - a Runtime connection credential. -2. Bootstrap through `/api/v1/agent-daemon/bootstrap` with an Authorization Bearer - header and device identity. Use its authenticated connection URL. -3. Dial the reverse WebSocket at `/api/v1/agent-daemon/ws`, supplying `device_id` - and `version` query parameters and the Bearer header. Never put credentials - in a URL, payload log or trace. -4. Send an immediate heartbeat, then continue at the configured interval. - Declare `supported_agent_kinds`, availability and capabilities explicitly. - An absent first heartbeat means capabilities are unknown; an omitted kind in - a received heartbeat means it is not advertised. Neither permits inference. -5. Dispatch ordered JSON envelopes over the connection. Heartbeats establish - liveness only, not execution progress or a receipt for earlier messages. - -The wire version is [`proto.Version`](../internal/agentdaemon/proto/version.go), -independent of the Runtime build version reported in heartbeats. Core accepts -only an exact match, including the patch component. A mismatch returns HTTP 426 -`incompatible_version` before dispatch; the daemon treats it as permanent and -stops reconnecting. Deploy matching peers together. Removed fields, inferred -Claude availability and old interaction shapes have no compatibility path. - -Each physical connection has fresh routing, admission handles and transfer state. -A newer connection replaces the previous device connection. Core fences owner -leases and evicts old Run/interaction routes; the new connection does not inherit -them. A valid credential and connection are not authority to choose another -Session or Environment binding. - -### Explicit capability declarations - -`AgentKindCapabilities` describes the composed Runtime and Harness, independently -of `Available` and the Core model profile. Every field uses `CapabilitySupport`: -`CapabilitySupported` or `CapabilityUnsupported`. Zero means unspecified and is -invalid even for an unavailable Harness. Registration validates the complete -struct before changing the registry; there is no implicit basic descriptor. - -The wire still uses JSON booleans and includes every field, including `false`. -Encoding incomplete declarations fails; decoding rejects omitted, null, invalid -or unknown capability fields, including a missing capability object. An invalid -heartbeat clears the connection's admission snapshot and closes its transport. -That establishes no native completion or cancellation result. Both peers use the -same exact wire version; no historical declaration format is accepted. - -Each admitted Executor and Turn retains its declaration. Rediscovery cannot add -operations to an existing owner. Optional operations check this snapshot before -native calls; interface presence alone never grants support. A declared operation -returning `agent.ErrUnsupportedOperation` is a contract violation, distinct from -unavailability, a failed native call or an uncertain write. Uncertain operations -keep their existing receipts and ownership; they are never automatically replayed. -Workspace support includes the common Runtime workspace implementation, so a -native adapter's unsupported workspace method does not disable that composition. - -New fields require an explicit decision in each production declaration. Contract -tests enumerate every field for registration, wire round trips and the persisted -boolean projection. The shared test fixture lists current fields individually; -it does not supply defaults for future fields. The Harness interface inventory -also requires a role decision and compile assertions for every public adapter; -see [Harness onboarding](../contracts/agents-api/harness-onboarding.md). - -### Executor and Turn lifetimes - -This section owns the separation of execution and resource lifetimes. - -| Lifetime | Owner | Ends when | -| --- | --- | --- | -| Environment allocation | Resource management (Sandbox Provider) | Explicit reclamation, coordinated with Runtime execution | -| Runtime connection | Runtime transport | Disconnection or replacement by a newer connection | -| Installed capability snapshot | Runtime | Its Environment is reclaimed; never on Executor close | -| Session Executor | Runtime | `Executor.Close` on idle expiry, shutdown or confirmed invalidation | -| Turn | Adapter `Turn`, tracked by Runtime | `AwaitSettlement` confirms settlement | - -A Session owns one reusable Executor in its connected Runtime. A Turn owns one -input execution, its output stream and its cancellation. `agent.ExecutorFactory` -prepares the fixed configuration; `Executor.StartTurn` creates a new `agent.Turn` -without replacing healthy native resources. Normal completion settles only the -Turn. `Executor.Close` releases native resources on idle expiry, environment -shutdown or confirmed invalidation. Core does not keep a second Executor cache. -The same lifecycle applies after managed or user-managed environments connect, -and to the qualified no-environment profiles. Resource management owns machine -selection, allocation and Environment creation/reclamation. Closing an Executor -does not release the Environment allocation or delete its workspace. Environment -reclamation explicitly coordinates with Runtime execution. Connection, installed -capability snapshot, Session Executor and Turn retain separate lifetimes. - -The Runtime binds its Executor record to Session, Environment, connection and -immutable execution configuration. Resume identity and prior-Turn recovery flags -are continuity assertions, not configuration changes. A supplied native identity -must match the retained owner; exact-history recovery never starts a new root -when existing history is required. A configuration conflict is an error, not a -hot switch. Lost connections retire their owners and handles. Old timers, output -and cancellation cannot affect replacements. - -Each Turn receives a fresh wrapper, output channel and receipt state. Optional -steering, functions, permissions and user-choice interfaces belong to that fixed -Turn. Native callbacks capture the originating Turn before asynchronous work; -late events cannot be assigned to whichever Turn happens to be active. Native -processes, query/transport connections, fixed capability configuration and native -Session identity belong to the Executor. Do not reset completed `sync.Once` -values or repurpose an old Turn object. - -`StartTurn` returning nil guarantees that no native input was submitted and the -output channel was not retained. The Runtime then closes that channel. Once input -may have been submitted, return a non-nil Turn even with an error: the Turn owns -exactly-once output closure and remains tracked until settlement. Unknown input -is never replayed. A definite `executor_unavailable` Start rejection permits one -common recovery attempt only after the previous Executor has been closed and no -input was submitted. Recheck the same physical peer and current authorization. - -`Turn.Cancel` targets only that Turn and does not close a healthy Executor. -`AwaitSettlement` applies after both natural completion and cancellation. Success -means output can no longer be written and the Turn's native events, input, -functions, interactions and child work have settled. Native completion or -cancellation confirmation is independent of resource retirement: closing a -transport cannot supply missing native terminal or operation receipts. -`Reusable=true` additionally -confirms that the native owner can accept the next Turn. `Reusable=false` requires -a reason and subsequent confirmed Executor close. An error means settlement is -unconfirmed; it cannot free ownership or capacity. Caller deadlines stop waiting, -not tracked cleanup. Retry the same cleanup target serially. Failed cleanup -blocks replacement and retains its resource slot. Executor Close confirms resource -retirement independently of the Turn outcome: an immutable Turn error must not -prevent closing the native transport and releasing resources once their work and -output have stopped. - -One output consumer starts before native Start, drains the bounded 64-frame -channel, and retains the terminal observation until Start publication, Turn -settlement and admitted operation receipts finish. Natural Done never calls -Cancel. Input, function and interaction admission close before settlement, and -operations already admitted hold their barrier through native receipts and -outbound acknowledgement. Send cancellation to the fixed Turn before waiting -for that barrier: a written input may need native interruption to produce its -receipt. Join native settlement, any required confirmed Executor close, output -drain and all admitted operations before an applied acknowledgement or reuse. -A failed Close may report failure while retaining the same Run and outstanding -operations for retry. Closing a caller wait cannot manufacture an applied input receipt. Only then forward Done or an applied cancellation -receipt. Commit native continuity and release the old Run admission before -publishing Done, since the receiver may immediately start another Turn. A late -terminal-send failure belongs to the old Run; it cannot invalidate a successor -that already owns the Executor. Connection shutdown owns transport-loss cleanup. -Preserve the ten-second settlement wait and separate five-second receipt -send budget; timeout is not proof of quiescence. The observed cancellation outcome -retains native identity, Usage and output without fabricating missing evidence. -Adapters must include owned background work in settlement and retain the exact -native cleanup target after failure. Native termination mechanisms belong to the -adapter; a bulk cleanup acknowledgement alone cannot establish quiescence. - -Private preparation controls reserve a per-Turn admission, not a new Executor. -They carry an explicit Session identity and immutable configuration without model -input or Run identity. A fresh request returns a connection-local admission handle -and the owning Executor ID. Start supplies both identities and its actual Run ID -and ordered MessageInput. Per-admission revisions order status observations; -rejections describe control errors without inventing Run events. A reused healthy -Executor returns ready without native preparation. An admission release abandons -that admission; it does not close the Session's healthy idle Executor or cancel a -later Turn. Cancellation uses the exact Run identity. - -Preparations and Start execute outside the receive loop and Router lock. Admission -expires after five minutes; retries do not extend that deadline. Bound active -preparation and execution separately from idle retained resources, and count -closing or uncertain resources until cleanup succeeds. A definite -`execution_prepare` rejection with `preparation_capacity` leaves an unclaimed -queued Turn for the existing Worker scheduler to retry, including capacity held -by cleanup. Other errors and uncertain input delivery do not authorize replay. -At most 64 admission records are retained; old handles never consume replacement -admissions. Idle expiry -is a Runtime resource policy, not Core active-Turn concurrency. Shutdown tracks and -closes active and idle Executors, retains failed close targets, and allows a later -serialized retry. Ordinary disconnection closes the failed transport and keeps -the exact Router until shutdown succeeds. A wait timeout or failed cleanup cannot -authorize reconnect; process shutdown also keeps waiting rather than silently -discarding owned native resources. These records are connection-local, not durable -input replay. - -Read-only workspace preparations remain separate bounded filesystem operations; -they cannot start model work. Workspace operations retain exact binding and -settlement rules across Turn boundaries and Executor closure. +Core owns durable Session, Turn, input and Environment records, scheduling and reconciliation. Runtime owns native Executors, active Turns, transfer state and cleanup until settlement. A Sandbox Provider owns placement and the surrounding compute. Releasing an execution admission or closing an Executor never deletes, suspends or reclaims a sandbox. The daemon is not an isolation boundary; see [Runtime and outer isolation](design-principles.md#runtime-and-outer-isolation). + +A Runtime connects in this order: + +1. Obtain a daemon credential and device ID. The [machine connection API](../contracts/agents-api/machine-api.md#credentials) lists the credential kinds; Project API keys and the Core key are never Runtime credentials. +2. Call `POST /api/v1/agent-daemon/bootstrap` with the credential as a Bearer header and the device ID. Use the connection URL it returns. +3. Dial the WebSocket at `/api/v1/agent-daemon/ws` with `device_id` and `version` query parameters and the Bearer header. Never put a credential in a URL, a payload log or a trace. +4. Send a heartbeat at once, then at the interval bootstrap returned. Each heartbeat declares `supported_agent_kinds`, their availability and their [capabilities](#capability-declarations). Before the first heartbeat, capabilities are unknown; a kind missing from a heartbeat is not advertised. Neither permits inference. +5. Exchange ordered JSON [envelopes](#envelope-and-identity). Heartbeats establish liveness only, never execution progress or a receipt for an earlier message. + +The wire version is [`proto.Version`](../internal/agentdaemon/proto/version.go), independent of the Runtime build version that heartbeats report. Core accepts only an exact match, including the patch component. A mismatch returns HTTP 426 `incompatible_version` before any dispatch; the daemon treats it as permanent and stops reconnecting. Deploy matching peers together. + +Each physical connection has fresh routing, admission handles and transfer state. A newer connection for the same device replaces the previous one: Core fences the owner lease and evicts the old Run and interaction routes, and the new connection inherits none of them. A valid credential and connection are never authority to choose another Session or Environment binding. + +## Capability declarations + +`AgentKindCapabilities` in [`inbound.go`](../internal/agentdaemon/proto/inbound.go) describes one composed Runtime and Harness, independently of `available` and of Core's engine profile. Every field is a `CapabilitySupport`: supported or unsupported. The zero value is unspecified and invalid, even for an unavailable Harness. Registration validates the complete declaration before changing the registry; there is no implicit basic descriptor. + +On the wire each field is a JSON boolean, and every field is present, including `false`. Encoding an incomplete declaration fails. Decoding rejects omitted, null, invalid and unknown fields, and a missing capability object. An invalid heartbeat clears the connection's admission snapshot and closes its transport; that establishes no native completion or cancellation result. + +Each admitted Executor and Turn keeps the declaration it was admitted with. A later heartbeat cannot add operations to an existing owner. Optional operations check this snapshot before any native call; the presence of a Go interface never grants support. A declared operation that returns `agent.ErrUnsupportedOperation` is a contract violation, distinct from unavailability, a failed native call or an uncertain write. Uncertain operations keep their receipts and ownership and are never replayed automatically. Workspace support includes the common Runtime workspace implementation, so a native adapter's unsupported workspace method does not disable that composition. + +A new field requires an explicit decision in every production declaration. Contract tests enumerate every field for registration, wire round trips and the persisted boolean projection; the shared test fixture lists fields individually and supplies no defaults for future ones. [Harness onboarding](../contracts/agents-api/harness-onboarding.md) owns the adapter side of each declaration. + +A declaration describes what the Runtime can do. Core admits a public feature only when the Harness's engine profile also qualifies it, and checks the declaration of the selected device during device selection and again at the final check before it claims a Turn: + +| Capability | Core requires it when | +| --- | --- | +| `streaming`, `steering`, `durable_turns`, `durable_input_receipts`, `preparation`, `execution_controls`, `tool_observations` | Always, for every execution on that Harness (with `available` true) | +| `environment_none` | The Environment type is `none` | +| `local_environment`, `workspace_read_preparation`, `workspace_output_export` | The Environment type is `openai_hosted` or `self_hosted` | +| `workspace_read_preparation` | An idle Files directory read needs a read-only preparation | +| `native_session_recovery` | A Session with a started Turn has no recorded native Session ID | +| `web_search_control`, `text_verbosity` | The Harness's engine profile declares that control | +| `structured_output` and `message_items` | The Agent requests `json_schema` output | +| `subagent_observations` | `multi_agent.enabled` is true | +| `subagent_control` | `multi_agent.enabled` is false | +| `tool_search` | The Agent enables tool search or defers function loading | +| `programmatic_tool_calling_disable` | The Agent explicitly disables programmatic tool calling | +| `function_tools` | The Agent declares function tools | +| `message_images`, `function_result_images` | A message, or a function result, carries an image | +| `mcp_http_tools`, `mcp_http_required`, `mcp_http_bearer_auth` | The Agent declares HTTP MCP servers; one is `required`; a Vault credential is selected for one | + +`permissions` gates permission decisions inside the Runtime, and `workspace_authoring` gates the daemon's authoring command. Core has no admission rule for `usage` and `resume`. + +The prompt request (`prompt_request`, or the configuration of `execution_prepare`) carries the opt-ins Core sets for each Run: + +| Field | Set by Core | +| --- | --- | +| `execution_controls` | Always: web search `disabled`, the resolved text verbosity (default `medium`), an explicit programmatic-tool-calling disable and any `json_schema` output format. Native option names belong to the adapter | +| `observe_tool_observations` | Always. Tool-call frames then carry the engine-neutral `observation` | +| `observe_messages` | When the Runtime declares `message_items`. Text deltas then carry the native item ID, and `output_message` frames report message start, completion, phase and the completion text | +| `observe_subagent_identities`, `disable_subagents` | From the Agent's `multi_agent.enabled` | +| `disable_execution_environment` | For an Environment of type `none` | +| `local_environment` | For `openai_hosted` and `self_hosted`, with the exact Environment binding | +| `strict_resume`, `require_existing_native_session` | Always strict; the second when a native Session must be recovered | +| `durable_receipt` on `prompt_steer` | For every active input Core delivers | + +Requests without an opt-in keep the frames and fields they had without it. ## Envelope and identity -Every data frame is one JSON -[`Envelope`](../internal/agentdaemon/proto/envelope.go): -`type`, type-dependent `id`, typed `payload`, and optional W3C `trace`. -Trace is diagnostic correlation only; missing or invalid trace data creates a -local trace and never changes ownership. Do not use a trace ID as a request ID. +Every data frame is one JSON [`Envelope`](../internal/agentdaemon/proto/envelope.go): `type`, a type-dependent `id`, a typed `payload` and an optional W3C `trace`. The trace is diagnostic correlation only; missing or invalid trace data creates a local trace and never changes ownership. Never use a trace ID as a request ID. | Identity | Scope and meaning | | --- | --- | | Device ID and connection | Authenticated Runtime routing and connection ownership | -| Session ID / Environment ID | Core-owned configuration and workspace binding; canonical UUIDs where required by the payload validator | -| Executor ID | Runtime-owned native resource, potentially retained across settled Turns with identical configuration | -| Preparation request ID | `Envelope.id` for prepare/start/release/status; distinct from a Run | -| Admission handle | Runtime-generated reservation, valid only on its accepting connection | -| Run ID | Execution attempt; `Envelope.id` for output, cancellation, active input and functions | +| Session ID / Environment ID | Core-owned configuration and workspace binding; canonical UUIDs where the payload validator requires them | +| Executor ID | Runtime-owned native resource, possibly retained across settled Turns with identical configuration | +| Preparation request ID | `Envelope.id` for prepare, start, release and status; distinct from a Run | +| Admission handle | Runtime-generated reservation, valid only on the connection that accepted it | +| Run ID | One execution attempt; `Envelope.id` for output, cancellation, active input and functions | | Interaction ID | `permission_request.payload.request_id` or `prompt_for_user_choice.payload.ask_id`; these request envelopes still carry the Run ID | -| Delivery ID / input ID / call ID | Resolve attempt, active-input receipt and native function identity respectively; never interchangeable | +| Delivery ID / input ID / call ID | Resolve attempt, active-input receipt and native function identity; never interchangeable | | Transfer ID / suspension ID | Connection-local transfer correlation / persisted suspension-attempt fencing | -Decision and permission-cancel envelopes use the interaction ID. Cancellation and -function-result acknowledgements use the Run ID. All application decision -receipts additionally match the delivery ID. A reply without the required -correlation cannot establish acceptance. +Decision and permission-cancel envelopes use the interaction ID. Cancellation and function-result acknowledgements use the Run ID. Every application decision receipt also matches the delivery ID. A reply without the required correlation cannot establish acceptance. -User-choice decisions carry `question_answers` with an explicit `question_id` -and an `answers` array for each provided answer. The IDs must belong to the emitted -questions and cannot repeat. Question order and display headers do not identify -answers; omitted questions remain unanswered and an empty array is an explicit -non-answer. Cancellation carries `cancelled: true` without answers. Shared -validation rejects other shapes before native submission. Adapters translate the -identified values into their native response without changing question identity. +User-choice decisions carry `question_answers`: an explicit `question_id` and an `answers` array for each provided answer. The IDs must belong to the emitted questions and cannot repeat. Question order and display headers do not identify answers; an omitted question stays unanswered, and an empty array is an explicit non-answer. Cancellation carries `cancelled: true` without answers. Shared validation rejects other shapes before native submission. ## Message families -Use the linked source definitions for required fields, validators, limits and -finite error categories. There is no parallel payload schema to keep in sync. +The linked source files define the required fields, validators, limits and finite error categories. | Core → Runtime | Runtime → Core | Definition | | --- | --- | --- | @@ -254,313 +101,106 @@ finite error categories. There is no parallel payload schema to keep in sync. | `workspace_read`, `workspace_write`, `workspace_export` | Matching `*_result` | [Read](../internal/agentdaemon/proto/workspace_read.go), [write](../internal/agentdaemon/proto/workspace_write.go), [export](../internal/agentdaemon/proto/workspace_export.go) | | `environment_quiesce`, `environment_resume` | `environment_quiesced`, `environment_resumed` | [Suspension fencing](../internal/agentdaemon/proto/suspend.go) | -Initial and active input share [ordered MessageInput](../internal/agentdaemon/proto/message_input.go). -Adapters preserve message/content order and explicitly reject unsupported content. -Usage frames and the final usage snapshot replace earlier cumulative snapshots; -do not add them. An absent measurement is unknown, not zero. - -The wire version is `proto.Version` (see above). Initial, prepared and active input use -the same ordered MessageInput contract, replacing scalar prompts and attachments. -User-message boundaries and text/image order remain intact through Core and the -Runtime wire; adapters own native conversion and receipt aggregation. Text-only -transports reject image content rather than dropping it. Text is never trimmed: -engine profiles declare whether whitespace-only messages are qualified, and -unqualified harnesses (Claude SDK, MiniMax Code) reject them at admission rather than having -their input rewritten. Codex has a flat native -input list and uses blank-line separators between messages; this does not preserve -independent native user-message boundaries. -The independently packaged Claude bridge uses protocol 3 for a prepared Executor -and separately identified Turns; readiness rejects other protocol versions. -Image-bearing messages require a qualified operation profile before persistence -and image support from the selected Runtime before native delivery. These checks -apply to that operation only; ordinary text retains offline queueing. Initial, -prepared and active paths use the same content and preserve receipt ownership. -The qualified public profile is inline PNG/JPEG on Codex/Claude `none` and -Core-managed Docker `openai_hosted` and user-managed `self_hosted`. MiniMax -images and remote URLs remain explicit implementation gaps. Workspace images reuse -the existing preparation, active-input and workspace authority; they do not add -a downloader, a mount or a separate execution lifecycle. Core -does not fetch or transform media. See [message input coverage](../contracts/agents-api/message-content.md#images). +Initial, prepared and active input use the same [ordered MessageInput](../internal/agentdaemon/proto/message_input.go). Adapters keep message and content order and reject unsupported content explicitly; a text-only transport rejects image content rather than dropping it. The [message input contract](../contracts/agents-api/message-input.md) owns the public image profile, whitespace rules and each Harness's native conversion. + +Usage frames and the final usage snapshot each carry the cumulative measurement of the current execution and replace the previous snapshot; never add them. An absent measurement is unknown, not zero. ## Preparation and execution order -Hosted and self-hosted execution use the same preparation semantics. Placement -selects a connection and Runtime-owned workspace; Harness adapters perform native -configuration. The existing `prompt_request` message remains a direct execution -operation, not a fallback after failed prepared execution. - -Environment initialization uses `runtime_prepare`. For files or archives, send -`begin`, wait for `ready`, send ordered chunks and await matching `received` -offsets, then `commit` and await `completed`. Initialization and finalization -have typed headers without file data. Validate the expected outcome, offset, -size and finite error code with the shared validator. Only one transfer is -allowed per connection. A chunk receipt confirms staged bytes, not installation. -A completed commit confirms that operation, not that a future Turn has executed. - -For an execution Turn: - -1. Subscribe to preparation status before sending `execution_prepare` with - immutable Session configuration and no Run input. -2. `preparing` means the Runtime owns preparation; `ready` supplies the Executor - ID, admission handle, revision and expiry. Neither submits user input. -3. Subscribe to the Run before sending `execution_start` with that Executor, - handle, Run ID and ordered input. A valid start transfers the reservation - once. `started` confirms transfer to the Turn, not Turn completion; output - can race status delivery and must already have a subscriber. -4. Consume Run events until a native terminal outcome or loss of observation. - An execution error is followed by `done` to close that execution stream. - `done` closes the stream; preceding errors remain part of its outcome. -5. On abandonment before start, send `execution_release`. After ownership - transfers to a Run, use `prompt_cancel`; releasing the old handle cannot - cancel its successor. - -Preparation observations use monotonically increasing revisions per handle. -Ignore older or repeated revisions; do not apply a status for a different handle. -Typical transitions are `preparing → ready → starting → started`, or termination -by `released`, `expired`, or `failed`. A `rejected` control operation has an -`operation` and error code but does not replace the resource's current revision. -Read-only workspace preparation cannot start a Turn. Expiry does not remove the -Runtime's obligation to settle cleanup. - -Core and Runtime use common preparation, start, input-receipt, cancellation, -release and recovery semantics for Codex, Claude Code and MiniMax Code. Retain each -harness's native implementation behind its adapter. Core acts on verified capabilities and runtime -conditions; a capability declaration alone never grants public feature admission. -Extend existing interfaces during related functional work without introducing a -second framework or a broad rewrite. Codex, Claude Code and MiniMax Code have -qualified dedicated Docker profiles and historical Core-managed E2B evidence. -New user-managed enrollment requires separate real acceptance. Each harness has equal standing; -qualify each image/template with the common full-loop acceptance before deploying. -Additional engines remain separate work; V1 has no separate remote executor. -Later engines must satisfy the same applicable acceptance contract while keeping -their suitable native deployment layout. - -Workspace reads may request the private `workspace_read_only` preparation profile -through the existing preparation factory and verified `workspace_read_preparation` -capability. It accepts only the bound Environment and resource identity; execution -options, model/MCP credentials, native Session continuation and model/tool input are excluded. -A read-only owner rejects Start and excludes model input, execution configuration -and plugin startup. A Runtime may satisfy this contract through its bound local -filesystem implementation; starting a native Harness process is not required. -Adapters that use native filesystem controls retain their own initialization rules. -For this read profile, `released` is published only after local Close succeeds; -cleanup errors retain ownership and report `cleanup_unconfirmed`. A failed factory -must return its resource with the error if cleanup remains unconfirmed; wrappers -must preserve both values. Successful cleanup retries publish confirmed release, -and stale status snapshots cannot publish success. Failed terminal status delivery -does not retry cleanup; ownership remains until an explicit release or shutdown retry. -A release request, -HTTP disconnect or remote socket closure alone is not cleanup confirmation. This -profile does not establish remote mutation quiescence or public Files admission. - -Core directory reads reuse the Worker's Session scheduling reservation for idle -preparation and target the exact Run for active execution. Device selection uses -operation-specific capabilities; reading files never resolves model/MCP options -or creates a Turn. HTTP cancellation ends observation, not an admitted native read. -Keep the idle reservation through the bounded read and release attempt. Return -directory data only after confirmed Close; incomplete reads or uncertain cleanup -return unavailable without data. Release the Worker's scheduling reservation before -delivering the result so the caller can immediately request the next page. -Revoke the scoped read transport credential on -completion or failure. Runtime retains uncertain cleanup ownership and capacity; -this does not require a second durable Core owner registry or establish remote -write retirement. Public Files.list delegates workspace access to this reader; -the API owns tenant authorization, path validation and protocol pagination. Only -the reader's distinct `not_directory` result (a missing path, a regular file or an -unfollowed symlink) becomes an empty page; root, permission, transport and -uncertain failures keep their errors. Keep partial directory coverage and -unverified defaults explicit in the Files contract. - -Local inline file delivery uses the same authenticated daemon connection and exact -Environment/Session binding. All platforms use the daemon's Go implementation -for bounded file reads, directory listing, file creation and output export. -Files operations require no external helper executable or staging directory. -The Files API keeps workspace-relative paths and no-overwrite creation semantics; -it does not restrict native Harness tools' host permissions. - -Transfer a complete bounded body in acknowledged 64 KiB frames before invoking -the native file writer, verify the declared digest, and run no model for upload. -Keep the private 50 MiB transfer bound distinct from the official 5 MiB decoded -inline bound, which the API checks before any Runtime work. Files.create uses the -native writer's no-overwrite operation. Initial Session files retain their separate -atomic replacement behavior; do not change one caller's semantics for another. -The dedicated Runtime excludes execution while receiving or applying a write; -malformed, incomplete or expired transfers cannot reach the installer. Exact -commit/rejection receipts release the mutation owner. Missing or ambiguous -receipts retain uncertainty; observer cancellation and local process exit cannot -prove non-mutation. Before public admission, Core must durably reserve the write -under the Session lock and prevent successor mutation across restart until exact -settlement. Do not replay the request or introduce general replacement machinery. -Read-only operations retain their own authority and bounded ownership requirements. +Environment initialization uses `runtime_prepare` on every connection, managed or user-owned; the [Environment contract](../contracts/agents-api/environments.md#runtime-capability-preparation) owns what is prepared and when. For a file or archive, send `begin`, wait for `ready`, send ordered chunks and await each matching `received` offset, then send `commit` and await `completed`. Initialization and finalization have typed headers without file data. Validate the expected outcome, offset, size and finite error code with the shared validator. One transfer is allowed per connection. A chunk receipt confirms staged bytes, not installation; a completed commit confirms that operation, not that a later Turn ran. + +An execution Turn runs in five steps: + +1. Subscribe to preparation status, then send `execution_prepare` with the immutable Session configuration and no Run input. +2. `preparing` means the Runtime owns preparation. `ready` supplies the Executor ID, admission handle, revision and expiry. Neither submits user input. +3. Subscribe to the Run, then send `execution_start` with that Executor ID, handle, Run ID and ordered input. A valid start transfers the reservation once. `started` confirms the transfer to the Turn, not its completion; output can race the status, so it must already have a subscriber. +4. Consume Run events until a native terminal outcome or loss of observation. An execution error is followed by `done`, which closes the stream; the preceding errors remain part of its outcome. +5. To abandon before start, send `execution_release`. After ownership passes to a Run, use `prompt_cancel`; releasing the old handle cannot cancel its successor. + +A preparation reserves a per-Turn admission, not a new Executor. It carries an explicit Session identity and immutable configuration without model input or a Run ID. A fresh request returns a connection-local handle and the owning Executor ID; a reused healthy Executor returns `ready` without native preparation. Per-handle revisions order status observations: ignore older or repeated revisions and never apply a status to another handle. Typical transitions are `preparing → ready → starting → started`, or termination by `released`, `expired` or `failed`. A `rejected` control operation carries an `operation` and error code and does not replace the handle's current revision. Releasing an admission abandons only that admission; it does not close the Session's idle Executor or cancel a later Turn. + +Preparation and start run outside the receive loop and router lock. An admission expires five minutes after it is granted, and retries do not extend that deadline; expiry does not remove the Runtime's obligation to settle cleanup. The Runtime bounds active preparation and execution separately from idle retained resources and counts closing or uncertain resources until their cleanup succeeds. A definite `execution_prepare` rejection with `preparation_capacity` leaves the queued Turn unclaimed for the Worker to retry, including when cleanup holds the capacity; any other error or uncertain delivery authorizes no replay. The Runtime retains at most 64 admission records, and an old handle never consumes a replacement's admission. These records are connection-local, not durable input replay. + +Idle expiry of an Executor is a Runtime resource policy, separate from Core's active-Turn concurrency. On shutdown the Runtime closes active and idle Executors, keeps any target whose close failed and allows a later serialized retry. An ordinary disconnection closes the failed transport and keeps the exact router until shutdown succeeds; a wait timeout or failed cleanup never authorizes reconnection, and process shutdown keeps waiting rather than discarding owned native resources. Workspace operations keep their binding and settlement rules across Turn boundaries and Executor closure. + +`prompt_request` starts a Run directly, without an admission handle. It is not a fallback after a failed prepared start. + +## Active input receipts + +Core delivers active input as `prompt_steer` with `durable_receipt: true`, one input at a time per Run, and waits for its receipt before sending the next: + +| Phase | Timer | +| --- | --- | +| Native write | The Runtime bounds the adapter's native write at 10 seconds; the timer stops once the write is complete | +| `written` acknowledgement | Core waits at most 30 seconds from delivery for `written`; otherwise the input outcome is unknown | +| Receipt send | Each receipt send has its own 5-second, shutdown-aware budget | +| Native acceptance | `accepted` arrives under the Turn lifetime, with no automatic redelivery | +| Done | Before `done`, the Runtime waits at most 15 seconds (the write and send budgets) for an in-flight input | + +Neither `written` nor a send failure advances Core's input cursor. Once cancellation is sent, its receipt owns the terminal outcome even if an input becomes unknown first; Core records `cancel_unconfirmed` when no cancellation confirmation arrives within 15 seconds. A cancellation receipt carries the stopped Turn's confirmed continuity snapshot when no `done` is emitted. ## What each acknowledgement proves | Observation | Proven fact | | --- | --- | -| Core gateway `Send` returns nil | Envelope entered the local send queue | -| Runtime transport `Send` returns nil | WebSocket write completed locally | -| Transfer `received`, active-input `written` | Defined receive/write phase occurred; native execution or consumption is unconfirmed | -| Preparation `preparing` / `ready` | Preparation accepted / reservation ready, with no submitted Run input | +| Core gateway `Send` returns nil | The envelope entered the local send queue | +| Runtime transport `Send` returns nil | The WebSocket write completed locally | +| Transfer `received`, active-input `written` | The defined receive or write phase occurred; native execution or consumption is unconfirmed | +| Preparation `preparing` / `ready` | Preparation accepted / reservation ready, with no Run input submitted | | Preparation `started` | Admission transferred to the identified Turn | -| Active-input `accepted` | Adapter confirmed native consumption under its declared receipt semantics | -| `interaction_decision_ack.applied=true` | Identified operation settled; cancellation additionally requires native settlement | -| `done` and preceding execution events | Execution stream completed with its observed outcome | - -There is no generic receipt for every envelope. A successful send is not proof -that the peer received, accepted or completed a request. Process exit, a stop -signal and a canceled local context do not prove successful cancellation. -A `done` frame may also close a settled cancellation stream; it does not -override the cancellation receipt or imply successful execution. Cancellation -receipts may retain partial content, native identity and usage in `outcome` -even when no `done` is published. Failure to obtain settlement must -remain failed or unknown; it cannot become `applied=true`. +| Active-input `accepted` | The adapter confirmed native consumption under its declared receipt semantics | +| `interaction_decision_ack.applied=true` | The identified operation settled; a cancellation also requires native settlement | +| `done` and the preceding execution events | The execution stream completed with its observed outcome | + +No generic receipt exists for every envelope. A successful send does not prove that the peer received, accepted or completed a request. Process exit, a stop signal or a canceled local context does not prove cancellation. A `done` frame may also close a settled cancellation stream; it does not override the cancellation receipt or imply success. A cancellation receipt may keep partial content, native identity and usage in `outcome` even when no `done` is published. A failure to obtain settlement stays failed or unknown; it never becomes `applied=true`. ## Failures, retries and cleanup -Transport and execution outcomes are separate. Core's only Run subscription -entry point is `SubscribeDurable`; inspect `Subscription.Err()` when its event -channel closes. Disconnection and subscriber overflow close it with an explicit -observation error, without fabricating `error` or `done`. Core retains durable -truth and reconciles from confirmed facts. Runtime retains cleanup ownership -until native work, input receipts, interactions and child work have settled. - -The public Turn status is a separate, existing projection: -[`execution/delivery.go`](../services/core/internal/execution/delivery.go) -records an unsuccessful orchestration attempt as `failed`, including -`delivery_unknown` after an unconfirmed send and `event_stream_incomplete` -after subscription failure. A closed subscription can replace the send reason -with `event_stream_incomplete`; both retain an unknown native effect. -This public `failed` status is not proof that the Harness failed, that no side -effect occurred or that cleanup completed. Native error classification comes -only from observed Runtime error frames. Consumers must keep the observation -reason and any native evidence distinct; this contract does not add a public -`unknown` status or change the existing Turn state machine. - -| Condition | Required responsibility | +Transport and execution outcomes are separate. Core's only Run subscription entry point is `SubscribeDurable`; inspect `Subscription.Err()` when its event channel closes. Disconnection and subscriber overflow close it with an explicit observation error and fabricate no `error` or `done`. Core keeps the durable truth and reconciles from confirmed facts. The Runtime keeps cleanup ownership until native work, input receipts, interactions and child work have settled. + +The public Turn status is a separate projection. [`execution/delivery.go`](../services/core/internal/execution/delivery.go) records an unsuccessful orchestration attempt as `failed`, including `delivery_unknown` after an unconfirmed send and `event_stream_incomplete` after a subscription failure; a closed subscription can replace the send reason with `event_stream_incomplete`. Both mean the native effect is unknown: a public `failed` status does not prove that the Harness failed, that no side effect occurred or that cleanup completed. Keep the observation reason and any native evidence distinct. + +| Condition | Responsibility | | --- | --- | -| Unsupported capability or invalid binding | Reject explicitly before starting the unsupported operation; do not select another Harness | -| Confirmed preparation/execution failure | Preserve the finite error category and any observed result; Runtime settles its resources | -| Deadline or connection loss after dispatch | Caller has an unknown effect unless an application receipt proves otherwise; do not convert it to execution failure | -| Reconnection | Reestablish transport and capability advertisement; do not replay input, initialization, transfers or unresolved mutations | -| Duplicate preparation/start | Existing connection-local identity and fingerprint rules apply; a conflicting request rejects and an old handle cannot start replacement work | -| Duplicate input/function/decision | Use that family's existing receipt identity and conflict rules; no transport-wide deduplication or exactly-once promise exists | -| Cleanup failure | Retain resource ownership and report unconfirmed cleanup; a resource is not reusable merely because a waiter timed out | -| Lost cancellation receipt | Cancellation may have happened; lack of receipt cannot establish success or authorize another execution | - -Only retry operations whose own contract establishes that retry is safe. Retired -preparation request IDs may eventually allocate a new handle, so they are not -durable idempotency keys. Native-session resume is an explicit operation with -verified identity, not a response to a socket failure. A transient connection -error permits reconnecting the channel; authentication/version rejection requires -operator correction. Cleanup of an Executor remains separate from the Sandbox -Provider's confirmed reclamation of compute. - -## Workspace operations and connection observations - -The private daemon `workspace_read` control targets an existing preparation handle -or its transferred active Run on the same authenticated device connection. Require -the exact frozen Environment identity; callers cannot supply sockets, credentials -or workspace roots. Shared routing uses the optional `agent.WorkspaceReader` -interface, without selecting an engine by name. -The same control accepts `operation: directory` through the optional -`agent.WorkspaceDirectoryLister`, with mutually exclusive byte/entry limits and -typed directory metadata. Directory responses carry at most 1024 single-component -UTF-8 names of at most 255 bytes, so escaped metadata stays below the existing -frame bound. These are private transport limits, not public Files parameters. -Byte and directory operations share target checks, correlation, capacity and -retained operation waits; neither creates a Run or selects an engine by name. -This control does not itself authorize a public Files endpoint or placement. - -The separate optional `agent.WorkspaceDirectoryLister` observes one workspace-relative -directory on the existing Prepared/Session owner; empty path selects its root. -Return single-component names, entry kind, regular-file byte size and explicit -truncation only after directory/metadata access and handle cleanup settle. Reuse -byte-read admission, uncertainty and caller-detach ownership where applicable. -Do not promise a snapshot, recursive traversal or public pagination through this -private interface. The Runtime binds access to the frozen local workspace, -prevents path escape, bounds enumeration and requires a complete validated result. -Filesystem readiness alone does not enable public Files admission. - -Bound encoded request payloads to 8 KiB and correlation IDs to 128 bytes before -admission. Do not echo oversized IDs; omit oversized trace metadata in replies. -Bound raw control results to 1 MiB within the existing 4 MiB transport frame; the -limits are local policies, not pinned public protocol limits. Successful reads require complete bytes/truncation and acknowledged native -close. Safe native rejections carry no bytes; interrupted or ambiguous reads remain -unknown and stop further reads on that owner. Local RPC reap never establishes file -settlement. Retain a dispatched read's original bounded waiter across observer -cancellation and resource transfer/release; stop new admission on resource closure. -The gateway bounds subscriptions and never retries or replays on reconnect. Duplicate -pending operation IDs cannot start another read; this control does not promise durable -idempotency or result recovery. Preparation/Run ownership, public path authorization, -public file authorization and native cleanup retain their separate requirements. - -Connection observations use the existing execution lease and Session lock. A -separate `environment_connections` row retains the current generation and revision; -`environments.status` and its Session Environment-event snapshot commit together. -The producer serializes replacements, then numbers socket observations within each -generation. Duplicate or older revisions and superseded generations are inert. -Replacement retires a previously connected observation before publishing its new -registration. Registration alone creates no connected event. Event payloads contain -only public Environment identity/type/status and nullable error, never configuration, -credentials, registration IDs or revisions. Transport observations have no asserted -Turn association. `connected`/`disconnected` are distinct from native preparation -readiness; do not cast resource `expired` into the event vocabulary or emit `ready` -for a self-hosted connection. - -The existing Worker observes authenticated daemon peers for enrolled Environments, -using durable generation/revision fencing under its execution lease. On restart it -reconciles old connection observations before admitting new ones. A connection or -heartbeat does not prove native readiness or process quiescence. Failed or stale -observations cannot establish a current connection. +| Unsupported capability or invalid binding | Reject before starting the operation; never select another Harness | +| Confirmed preparation or execution failure | Keep the finite error category and any observed result; the Runtime settles its resources | +| Deadline or connection loss after dispatch | The effect is unknown unless an application receipt proves otherwise; do not convert it to an execution failure | +| Reconnection | Reestablish the transport and the capability declaration; never replay input, initialization, transfers or unresolved mutations | +| Duplicate preparation or start | Connection-local identity and fingerprint rules apply; a conflicting request rejects, and an old handle cannot start replacement work | +| Duplicate input, function result or decision | That family's receipt identity and conflict rules apply; there is no transport-wide deduplication or exactly-once promise | +| Cleanup failure | Keep resource ownership and report unconfirmed cleanup; a waiter's timeout does not make a resource reusable | +| Lost cancellation receipt | Cancellation may have happened; the missing receipt neither establishes success nor authorizes another execution | + +Retry only an operation whose own contract makes retry safe. A retired preparation request ID may eventually allocate a new handle, so request IDs are not durable idempotency keys. Resuming a native Session is an explicit operation with verified identity, never a response to a socket failure. A transient connection error permits reconnecting; an authentication or version rejection requires operator correction. Executor cleanup is separate from the Sandbox Provider's confirmed reclamation of compute. + +## Native failure classification + +An adapter may add `code` and `http_status` to a Run's `error` frame. They are optional neutral metadata, not a terminal event or a public error contract; the frame's text, Usage, `done`, native Session identity and cancellation receipts keep their order and meaning. + +The accepted codes are `authentication_error`, `rate_limit_exceeded`, `usage_limit_exceeded`, `server_overloaded`, `server_error`, `invalid_request`, `resource_not_found`, `request_timeout`, `context_length_exceeded`, `cyber_policy` and `connection_failed` ([`engine_failure.go`](../internal/agentdaemon/proto/engine_failure.go)). Only `connection_failed` keeps `http_status`, and only an integer from 100 to 599; every other status is discarded. A missing, malformed or unknown value leaves the error unclassified without discarding Usage or `done`. + +Core stores accepted values in the Turn outcome as `engine_error_code` and `engine_http_status`. The classification is subordinate to the terminal status and Core's `error_code`: it cannot turn a completed or cancelled Turn into a failure, hide an incomplete event stream, or override a persistence or cancellation-receipt failure. Normal delivery and terminal journal draining use the same extraction. [Session diagnostics](../contracts/agents-api/session-diagnostics.md) expose the category only for a failed Turn whose Core error is `engine_failed`. The [Codex](../services/core/deploy/codex/README.md) and [Claude Code](../services/core/deploy/claude/README.md) adapter guides give each Harness's mapping; an adapter never classifies error prose. + +## Workspace operations + +A workspace read that needs no running Turn uses the read-only preparation profile: `execution_prepare` with `workspace_read_only`, which requires the `workspace_read_preparation` capability. It accepts only the bound Environment and resource identity; execution options, model and MCP credentials, native Session continuation and model or tool input are excluded, and the owner rejects `execution_start`. A Runtime may serve it from its bound local filesystem without starting a Harness process. The profile publishes `released` only after local close succeeds; a cleanup error keeps ownership and reports `cleanup_unconfirmed`. A failed factory returns its resource with the error while cleanup is unconfirmed, and wrappers keep both values. A successful cleanup retry publishes the confirmed release; a stale status snapshot never publishes success. A release request, HTTP disconnect or remote socket closure alone does not confirm cleanup. + +`workspace_read` targets an existing preparation handle, or the Run it was transferred to, on the same authenticated device connection, with the exact frozen Environment identity; callers cannot supply sockets, credentials or workspace roots. `operation: directory` lists one workspace-relative directory (an empty path selects the root) with mutually exclusive byte and entry limits. A result carries at most 1024 single-component UTF-8 names of at most 255 bytes each, the entry kind, regular-file sizes and explicit truncation, and is returned only after directory access and handle cleanup settle. There is no snapshot, recursion or pagination at this layer. Byte and directory reads share target checks, correlation, capacity and retained operation waits. + +Request payloads are bounded at 8 KiB and correlation IDs at 128 bytes before admission; oversized IDs are not echoed, and oversized trace metadata is omitted from replies. Raw read results are bounded at 1 MiB within the 4 MiB transport frame. These are private transport limits, not public Files parameters. A successful read requires complete bytes or explicit truncation and an acknowledged native close. A safe native rejection carries no bytes; an interrupted or ambiguous read stays unknown and stops further reads on that owner. A dispatched read keeps its bounded waiter across observer cancellation and resource transfer or release, and resource closure stops new admission. The gateway bounds subscriptions and never retries or replays a read on reconnect; a duplicate pending operation ID cannot start another read. + +Core runs an idle directory read on the Worker's Session scheduling reservation and targets the exact Run during active execution. It keeps the reservation through the bounded read and release, returns data only after a confirmed close (an incomplete read or uncertain cleanup returns unavailable without data), releases the reservation before delivering the result, and revokes the scoped read credential on completion or failure. The Runtime keeps uncertain cleanup ownership and capacity. The [Environment Files contract](../contracts/agents-api/environment-files.md) owns public authorization, paths and pagination. + +`workspace_write` transfers a complete bounded body in acknowledged 64 KiB frames before the native writer runs, verifies the declared digest and runs no model. The private transfer bound is 50 MiB, separate from the public 5 MiB decoded inline bound that the API checks before any Runtime work. The Runtime excludes execution while it receives or applies a write; a malformed, incomplete or expired transfer never reaches the installer. An exact commit or rejection receipt releases the mutation owner. A missing or ambiguous receipt keeps the uncertainty: observer cancellation and local process exit cannot prove that nothing changed. Before public admission Core durably reserves the write under the Session lock and blocks successor mutations across restarts until exact settlement; the request is never replayed. Every platform uses the daemon's Go implementation for bounded reads, directory listing, file creation and output export, with no external helper or staging directory. + +## MCP connection authority + +Every public `MCPHTTPServer` in a prompt request carries an explicit `connection_origin`; a missing or unknown value rejects rather than selecting a default, and Core freezes the public default before dispatch. The Runtime validates the origin with the common validator before selecting a factory and resolves public and installed MCP into transient effective bindings. The [Environment contract](../contracts/agents-api/environments.md#public-mcp-connection-origin) owns the supported combinations, native limits and failure ownership. ## Contract verification -Run `make check-runtime-contract` from the repository root. It exercises the -shared wire validators, gateway, transport and dispatcher, plus -[real WebSocket contract scenarios](../apps/daemon/internal/contracttest/wire_test.go) -using a controlled Harness adapter, plus the [observation-result regression](../services/core/internal/execution/runtime_protocol_test.go). It requires no model credentials or external -sandbox. These tests are also included in `make check` through `check-go` and -`check-core`. - -The suite checks incompatible versions, preparation failure, cancellation -settlement, connection loss without invented terminal events, reconnect without -replay, stale/duplicate handles and receipts, cleanup failures, bounded transfer -validation and resource ownership after timeout. Existing detailed fault -injection remains next to the owning gateway/dispatcher implementation. - -For another Runtime or Harness, reuse these protocol sequences and assertions, -then run its own native acceptance for the capabilities it advertises. A passing -controlled-adapter test establishes the transport contract, not native Harness -behavior, OS support, provider authentication or sandbox isolation. Update the -shared types, this guide and the contract checks together when semantics change. - -The reusable [Harness text assertions](../apps/daemon/internal/agent/contracttest/text.go) -accept any prepared Executor and a small fixture supplying deterministic normal, -active and steering inputs. They check independent Turn streams, native owner and -history continuity, durable write/application receipts, stale cancellation and -healthy continuation after cancellation. The Claude Go adapter runs them against -its controlled native subprocess fixture. New adapters can call the same assertions; -no optional feature is implied. Native failures, uncertain cleanup and exact-history -recovery still require the adapter's fault and real-provider acceptance tests. - -### Environment preparation ownership - -Core tracks initialization on the Environment, independently of a managed -allocation. After authentication, both managed and user-owned Runtime connections -receive the same `runtime_prepare` operations and resource snapshots. Core never -replays a running initialization whose owner or confirmation was lost. Preparation -failure settles Environment input without destroying the machine or workspace. -The public `connected` state describes transport; initialization completion and -native executor readiness remain separate prerequisites for execution. Input -sources and frozen metadata follow the [Environment contract](../contracts/agents-api/environments.md#runtime-capability-preparation). - - -### MCP connection authority - -Public `MCPHTTPServer` messages carry an explicit `connection_origin`; missing or -unknown values reject rather than selecting a default. Core freezes the public -default before dispatch. Both peers require the exact wire version. Runtime uses -the common origin validator before selecting a factory and resolves public and -installed MCP into transient effective bindings. See the -[origin and credential contract](../contracts/agents-api/environments.md#public-mcp-connection-origin) -for supported combinations, native limits and failure ownership. +Run `make check-runtime-contract` from the repository root. It exercises the shared wire validators, gateway, transport and dispatcher, the [real WebSocket contract scenarios](../apps/daemon/internal/contracttest/wire_test.go) with a controlled Harness adapter, the [observation-result regression](../services/core/internal/execution/runtime_protocol_test.go) and the Harness declaration tests. It needs no model credentials or external sandbox, and `make check` runs the same tests through `check-go` and `check-core`. + +The suite covers incompatible versions, preparation failure, cancellation settlement, connection loss without invented terminal events, reconnection without replay, stale or duplicate handles and receipts, cleanup failures, bounded transfer validation and resource ownership after a timeout. Detailed fault injection stays next to the gateway and dispatcher code it tests. + +Another Runtime reuses these protocol sequences and assertions, then runs native acceptance for every capability it declares. A controlled-adapter test establishes the transport contract, not native Harness behavior, operating-system support, provider authentication or sandbox isolation. Change the shared types, this document and the contract checks together. Harness adapters also run the shared text assertions described in [Harness onboarding](../contracts/agents-api/harness-onboarding.md). From 60cbd79d691cb1889b3c74c1be41b893157965cf Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 08:43:15 +0000 Subject: [PATCH 04/10] docs: one owner each for machine routes, node protocol and deployment machine-api.md is the single /api/v1 document: routes, credential kinds including the operator device profile, node and daemon payloads and errors. The node protocol document covers the whole Core-to-node connection. The deployment contract lists every sandbox route and drops placement, history and qualification prose; node diagnostics follow the generation-managed rollout. The Sandbox Provider guide documents RunCommandCompute, marks the shared-code registration steps as known gaps and gains the managed lifecycle Core runs around every provider. --- AGENTS.md | 2 +- contracts/agents-api/machine-api.md | 120 ++++ .../agents-api/node-generation-protocol.md | 319 +++------ contracts/agents-api/sandbox-deployment.md | 644 +++++------------- docs/api/README.md | 14 +- docs/sandbox-provider.md | 430 ++++-------- 6 files changed, 526 insertions(+), 1003 deletions(-) create mode 100644 contracts/agents-api/machine-api.md diff --git a/AGENTS.md b/AGENTS.md index 853a2837a..5e9ee808f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -17,7 +17,7 @@ OpenAgentCore is protocol-first and modular. Core orchestrates operations that p | --- | --- | --- | | Application–Core (`/v1`) | Types in `contracts/agents-api/v1/` and route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/openapi.yaml` | [Agents API guide](docs/api/public-agent-api.md) | | Web and operators–Core (`/core/v1`) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/core.openapi.yaml` | [Core administration API](contracts/agents-api/admin-api.md) | -| Nodes and daemons–Core (`/api/v1` HTTP routes; the node and daemon wire protocols are separate rows) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/runtime.openapi.yaml` | [Machine connection API](docs/api/README.md#machine-connection-api) | +| Nodes and daemons–Core (`/api/v1` HTTP routes; the node and daemon wire protocols are separate rows) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/runtime.openapi.yaml` | [Machine connection API](contracts/agents-api/machine-api.md) | | Core–Sandbox Provider | `services/core/internal/sandbox/sandbox_provider.go` | [Sandbox Provider guide](docs/sandbox-provider.md) | | Core–sandbox node | `services/core/internal/sandbox/node/wire.go` | [Node generation protocol](contracts/agents-api/node-generation-protocol.md) | | Provider–Runtime startup | `internal/runtimebootstrap/bootstrap.go` | [Runtime bootstrap](docs/runtime-bootstrap.md) | diff --git a/contracts/agents-api/machine-api.md b/contracts/agents-api/machine-api.md new file mode 100644 index 000000000..2aca551b7 --- /dev/null +++ b/contracts/agents-api/machine-api.md @@ -0,0 +1,120 @@ +# Machine connection API + +Machines call Core under `/api/v1`: sandbox nodes, Runtime daemons and the self-hosted installer. Each route accepts only the credential listed for it, never the Core key or a Project API key, and a console sign-in grants nothing here. The reverse proxy sends `/api/v1` directly to Core; Web never serves these routes. + +## Routes + +| Route | Caller | Credential | Contract | +| --- | --- | --- | --- | +| `GET sandbox-node/configuration` | Node installer and node | Enrollment token, or node credential with `X-OAC-Node-ID` | [Read the node configuration](#read-the-node-configuration) | +| `POST sandbox-node/enroll` | Node installer | Enrollment token | [Enroll a node](#enroll-a-node) | +| `GET sandbox-node/identity?node_id=` | Node | Node credential | [Recover a node's identity](#recover-a-nodes-identity) | +| WebSocket `GET sandbox-node/connect?node_id=` | Node | Node credential | [Node generation protocol](node-generation-protocol.md) | +| `GET agent-daemon/install/{version}/…` | Self-hosted installer | None | [Installation grant](environment-executor-credentials.md#installation-grant) | +| `POST agent-daemon/installation`, `POST agent-daemon/installation/claim` | Self-hosted installer | Installation grant | [Installation grant](environment-executor-credentials.md#installation-grant) | +| `POST agent-daemon/enroll` | Self-hosted daemon | Executor credential | [Enroll a self-hosted daemon](#enroll-a-self-hosted-daemon) | +| `GET agent-daemon/connection?environment_id=` | Self-hosted installer | Executor credential | [Private connection confirmation](environment-executor-credentials.md#private-connection-confirmation) | +| `POST agent-daemon/bootstrap` | Runtime daemon | Daemon credential | [Daemon bootstrap](#daemon-bootstrap) | +| `GET agent-daemon/device-status?device_id=` | Runtime daemon | Daemon credential | [Device status](#device-status) | +| WebSocket `GET agent-daemon/ws?device_id=&version=` | Runtime daemon | Daemon credential | [Core–Runtime protocol](../../docs/runtime-protocol.md) | + +Every credential travels in an `Authorization: Bearer` header, never in a URL. + +The generated [`runtime.openapi.yaml`](runtime.openapi.yaml) describes only the sandbox-node configuration, enroll and identity routes and the two installation routes. The two WebSockets and the daemon bootstrap, device-status, enroll and connection routes are served outside the API router and have no generated schema; this document and the linked contracts are their only definition. + +## Credentials + +| Credential | Issued by | Accepted on | +| --- | --- | --- | +| Enrollment token | `POST /core/v1/sandbox/enrollment-tokens` (Web **Add node**), with the node's approved capacity. One use; it expires at the response's `expires_at` | `sandbox-node/configuration` without a node ID, `sandbox-node/enroll` | +| Node credential | The node itself: it generates a secret of 32 to 256 characters without whitespace and registers it at enrollment | `sandbox-node/configuration` with `X-OAC-Node-ID`, `sandbox-node/identity`, `sandbox-node/connect` | +| Installation grant | The `x_agents_core.installation` command of a `self_hosted` Session; short-lived | `agent-daemon/installation` and its `claim` | +| Executor credential | The installation claim, or the Core-key [executor credential routes](environment-executor-credentials.md) | `agent-daemon/enroll` and `agent-daemon/connection`; after enrollment it is also the daemon credential of the bound device | +| Daemon credential of a hosted sandbox | Core, for each managed allocation, delivered in the [bootstrap file](../../docs/runtime-bootstrap.md) | `agent-daemon/bootstrap`, `device-status` and `ws` | +| Operator device profile | `oac-core-device`, run by an operator with database access | `agent-daemon/bootstrap`, `device-status` and `ws` | + +Core keeps only a SHA-256 digest of each token and credential it stores; installation grants are signed and not stored. Credentials are not interchangeable: each works only on its own routes. + +### Operator device profile + +An engine host for `environment: none` Sessions connects with a device profile that an operator provisions directly in the database: + +```sh +umask 077 +mkdir -p ~/.oac/daemon/default +OAC_DATABASE_URL=... oac-core-device --tenant --name 'engine host' --url https://core.example > ~/.oac/daemon/default/auth.json +oac-daemon connect --profile default +``` + +`--tenant` is the Project's execution tenant UUID and `--url` Core's origin without a path. The command prints the profile once: `server_url` (the origin plus `/api/v1`), `runtime_id` (the device ID), `runner_credential` and `device_name`. Use a new profile rather than overwriting another device's file, and copy it privately to the same path on a remote host. `oac-core-device --tenant --revoke ` revokes the device: new connections are refused at once, and an open connection closes at its next heartbeat. The Worker binds each `none` Session to one connected device of its tenant that declares the required capabilities and keeps that binding across retries and restarts; self-hosted Sessions never use this path. + +## Node routes + +### Read the node configuration + +`GET /api/v1/sandbox-node/configuration` returns the active deployment for node installation and recovery. It never consumes an enrollment token. + +- A new node sends its enrollment token without `X-OAC-Node-ID`. The token must be valid, unexpired, unconsumed and issued by this installation. An active reset refuses this read. +- A registered node sends its node credential and its UUID in `X-OAC-Node-ID`. Without a query it reads the current target. `?generation=N` reads only a generation this node may still need: the current target, its serving pin, or one held by an unreleased allocation or placement on it; any other generation is refused. This read stays available during a reset, for owned recovery. + +The response has `installation_id`, `provider`, `core_url` (the installation public URL), `generation`, `specification`, `specification_digest`, `max_active` and `max_retained`. It never contains an administrator, Project or E2B credential, and exists only for node-backed providers. The [sandbox deployment contract](sandbox-deployment.md#canonical-node-specification) defines the specification and its digest. + +### Enroll a node + +`POST /api/v1/sandbox-node/enroll` registers a node and consumes the token. The body has exactly these fields: + +| Field | Value | +| --- | --- | +| `node_id` | A canonical UUID the node chose | +| `credential` | The node's secret, 32 to 256 characters without whitespace | +| `name` | Display name | +| `provider` | The deployment's provider | +| `backend_fingerprint` | The node's backend namespace digest | +| `deployment_generation`, `specification_digest` | The configuration the node read | +| `core_url` | The Core origin the node stores and connects to | + +Core checks, in one transaction, that the token is valid, the deployment is initialized, node-backed and not resetting, the generation and digest match the current specification, `core_url` equals the installation public URL and the node ID is new. Only then does it register the node, with the capacity approved in the token, and consume the token. The 201 response is the node identity: `node_id`, `installation_id`, `provider`, `deployment_generation`, `specification_digest`, `max_active` and `max_retained`. The node cannot submit capacity; the enrolled generation and digest stay the node's immutable identity, and later generations use separate configurations. + +### Recover a node's identity + +`GET /api/v1/sandbox-node/identity?node_id=` returns the same identity plus `connected` and `provider_ready`, as Core currently sees them. + +### Node route errors + +| HTTP | Code | When | +| --- | --- | --- | +| 400 | `invalid_request_error`, `param: "core_url"` | Enrollment without `core_url` | +| 400 | `invalid_request` | A malformed body, node ID or generation, or a provider other than the deployment's | +| 401 | `invalid_node_credential` | A missing, invalid, expired, consumed or foreign token or node credential | +| 409 | `sandbox_specification_mismatch` | The node's generation or digest does not match | +| 409 | `sandbox_node_address_mismatch` | `core_url` is not the installation public URL; the token stays unused | +| 409 | `idempotency_conflict` | The node ID is already registered | +| 409 | `sandbox_reset_in_progress` | Enrollment or a new node's configuration read during a reset | +| 503 | `runtime_node_unavailable` | The deployment is not initialized, or storage is unavailable | + +The credential is checked before any deployment state, so a rejected credential, including one issued for another installation, gets 401 even before initialization or under E2B. Until the deployment is initialized, the configuration read and enrollment answer an otherwise valid token with 503, and the identity read and node connection answer 401. + +## Daemon routes + +### Daemon bootstrap + +`POST /api/v1/agent-daemon/bootstrap` with the daemon credential and `{"device_id": "…"}` returns `device_id`, `workspace_id`, `ws_url` (derived from `OAC_PUBLIC_URL`, never from request headers), `heartbeat_seconds` and `protocol_version`. The daemon then dials `ws_url` as the [Core–Runtime protocol](../../docs/runtime-protocol.md#ownership-and-connection) describes. + +### Device status + +`GET /api/v1/agent-daemon/device-status?device_id=` with the daemon credential returns `device_id`, `online` and `owner`: the current connection owner's `owner_pod_id`, `owner_url`, `generation`, `status` and `lease_expires_at`, or null. + +The bootstrap, device-status and WebSocket routes share one error body, `{"error": code, "detail": text}`: 400 `missing_params`, `missing_device_id` or `bad_json`; 401 `missing_bearer`, `unknown_device` or `bad_credential`; 403 `wrong_runtime_type`; 500 `internal`; and on the WebSocket 426 `incompatible_version` when `version` is not Core's exact Runtime protocol version. + +### Enroll a self-hosted daemon + +`POST /api/v1/agent-daemon/enroll` with the executor credential and exactly `{"environment_id": "…"}` (no query) binds one dedicated device to the Environment's Session and returns `device_id`, `session_id`, `environment_id` and `workspace_directory`. It never returns another credential: the executor credential becomes the daemon credential of that device. A retry with the same credential returns the same binding. A successful response carries `Cache-Control: no-store`. + +| HTTP | When | +| --- | --- | +| 400 | A malformed body or any query | +| 401 | An invalid, revoked or foreign credential, a deleted Session, or an Environment without current executor authority | +| 409 | The Environment is already bound to a different key or device | +| 503 | Storage is unavailable | + +Enrollment creates no managed allocation and grants no Session API access. The daemon keeps the binding beside its credential and refuses another Environment's native history. The gateway and Worker recheck the credential's authority on every connection and dispatch, so rotation, revocation and Session deletion end further use. The [self-hosted guide](../../docs/getting-started/self-hosted.md) gives the operator steps, and the [executor credential contract](environment-executor-credentials.md#revoked-or-rotated-credential) describes how the daemon handles a permanent rejection. diff --git a/contracts/agents-api/node-generation-protocol.md b/contracts/agents-api/node-generation-protocol.md index 1cebfedec..debc38eff 100644 --- a/contracts/agents-api/node-generation-protocol.md +++ b/contracts/agents-api/node-generation-protocol.md @@ -1,231 +1,112 @@ -# Node generation protocol - -This is the internal, authenticated Core-to-node Provider protocol. It does not -change the pinned public Agents API. `sandbox/node.ProtocolVersion` is the only -accepted wire version; both peers reject historical versions. The current hello -explicitly advertises `generation_management` when the node can prepare and retain -multiple deployment generations. Fixed-configuration manual nodes omit that -capability and serve only their enrolled generation using the same wire protocol. -Every Provider request carries its exact allocation-owned deployment generation; -Core never strips it for an older peer. Automatic preparation and retention frames -are sent only to nodes advertising generation management. - -## Provider operation outcomes - -Node startup and generation loading validate complete Provider operation -declarations before accepting work. Proxies use the same registered declaration -for admission; unsupported operations reject before node resolution or native I/O. -The Provider [operation contract](../../docs/sandbox-provider.md#explicit-operation-contracts) -owns the inventory and support rules. - -A Provider response with `error_code: unsupported` carries an `unsupported` object -containing the exact method `operation` and an authored safe `reason` code. The -proxy checks both against the request. Missing, malformed or mismatched evidence -is an unconfirmed result, never proof that a mutation was rejected. Unsupported -remains distinct from observation unavailability and unknown compute/command -results; it does not settle resource ownership or authorize replay. The current -private wire version requires both peers to understand this outcome. - -## Bounded control - -A generation-managing node's hello or heartbeat contains at most eight generation observations. -Each names a positive signed-64-bit generation, its lowercase SHA-256 specification -digest, a `ready`, `preparing` or `failed` state, and an optional fixed diagnostic. -The target and serving generation take priority; other records rotate fairly. -Eight bounds one message, not the number of generations a node may retain. -Omitted observations never authorize deletion or imply absence. - -Welcome and heartbeat acknowledgements carry the Core-owned target generation, -its specification digest, and an explicitly nullable durable serving generation. -Target preparation is independent of the readiness of a retained serving provider. -A provider request carries its exact immutable `deployment_generation` separately -from the allocation's compute generation. - -Retention exchanges use an independent bounded control path. A request identifies -at most eight local `(generation, specification_digest)` references, a UUID request -ID, a monotonically increasing sequence, the current connection UUID and owner -epoch. Its acknowledgement must match the complete pending request, including -entry order and identity, and must explicitly supply a boolean `keep` for every -entry. Only one pending exchange exists per connection. A disconnect discards it; -an unsolicited, replayed, stale, partial or mixed acknowledgement deletes nothing. - -Control envelopes are at most 32 KiB. Their member names are exact and unique; -unknown members, case aliases, duplicate members and unexpected nulls are rejected. -The nullable serving pin and host measurements preserve unknown values. Provider -request/response frames retain the existing global size bound; control traffic -does not increase that bound or consume the provider-operation queue. +# Sandbox node protocol + +A sandbox node runs the Docker or microsandbox Provider on its host and connects to Core over one WebSocket. Core sends Provider operations over that connection; the node runs them against its local provider and reports readiness, host measurements and the deployment generations it holds. Core stays the only lifecycle owner: the node never retries a mutation or schedules work. The frames and validators live in [`services/core/internal/sandbox/node`](../../services/core/internal/sandbox/node) (`wire.go`, `generation_wire.go`); the HTTP routes a node uses to enroll and read its configuration are in the [machine connection API](machine-api.md#node-routes). + +## Frames and version + +Every frame is one JSON text message whose `version` equals `node.ProtocolVersion`; both peers reject any other version, and there is no fallback decoder. Member names are exact and unique: unknown members, case aliases, duplicates and unexpected nulls are rejected. Control frames (`hello`, `welcome`, `heartbeat`, `heartbeat_ack`, `retention`, `retention_ack`) are at most 32 KiB; `request` and `response` frames at most 72 MiB. An invalid frame closes the connection. + +## Connection + +1. The node dials `/api/v1/sandbox-node/connect?node_id=` on its stored Core origin (`wss` for `https`) with its node credential as a Bearer header. Core answers 401 to a rejected credential, which the node treats as permanent; any other failure, including a 403 from a proxy, is retried with bounded backoff. Core refuses a second connection for a node identity while one is opening, live or closing, with 409. +2. Within 15 seconds the node sends `hello` with its identity (`node_id`, `installation_id`, `provider`, `backend_fingerprint`, the enrolled `deployment_generation` and `specification_digest`, `max_active`, `max_retained`), its first health report and, when it can prepare and retain several deployment generations, `generation_management: true`. Core closes the connection unless the identity matches the authenticated node. +3. Core records the node's presence, then replies `welcome` with a new `connection_id`, the current `owner_epoch` and, for a generation-managing node, `deployment`: the target `generation`, its `specification_digest` and a nullable `serving_generation`. The node stores a higher owner epoch and refuses a lower one. +4. Every 10 seconds the node sends `heartbeat` with `connection_id`, `owner_epoch` and health. On each heartbeat Core authenticates the node credential again and checks the owner epoch, records the health and replies `heartbeat_ack`, with `deployment` for a generation-managing node. Either peer closes the connection after 35 seconds without a frame. + +Core counts a node as online while it is connected under the current owner epoch and its last heartbeat is less than 45 seconds old. A heartbeat establishes provider readiness and the last host measurements, never Session activity. + +Health carries `provider_ready`, an optional fixed `diagnostic`, `observed_at`, `active_operations` (at most 32) and the host measurements that the [Runtime telemetry API](runtime-observability-api.md#node-host-observations-and-history) reports. A node without generation management probes its provider for every report; an unready provider reports one fixed diagnostic code, classified from typed probe errors, and the probe text and host paths stay on the node. Core stores an unknown code as `provider_unavailable`. The [nodes guide](../../docs/getting-started/nodes.md#readiness-codes) lists the codes and their causes. A generation-managing node reports readiness per generation instead, as described below. + +## Provider requests + +Core sends `request` frames with: + +| Field | Meaning | +| --- | --- | +| `id` | New UUID per request | +| `sequence` | Increases by exactly one per request on this connection | +| `connection_id`, `owner_epoch` | The values from `welcome` | +| `deployment_generation` | The generation the allocation belongs to, separate from its compute generation | +| `operation` | One of the operations below | +| `timeout_ms` | Remaining budget, 1 to 120000 | +| `reference` | The exact `(tenant_id, environment_id, allocation_id)` | + +Each operation carries exactly its own argument: + +| `operation` | Provider method | Argument | +| --- | --- | --- | +| `create` | `Create` | `bootstrap` | +| `info`, `renew`, `kill`, `command` | `GetInfo`, `Renew`, `Kill`, `RunCommand` | none, or `command` | +| `observe` | `Observe` | `observation` | +| `initial`, `new_compute`, `compute`, `kill_compute`, `resume_compute`, `command_compute` | Checkpoint compute operations | none, an optional `snapshot`, `compute`, or `compute` and `command` | +| `suspend`, `resume`, `delete_snapshot` | `Suspend`, `Resume`, `DeleteSnapshot` | `suspend`, `resume` or `snapshot` | + +A request whose `connection_id`, `owner_epoch` or `sequence` does not match closes the connection. A malformed request gets an `invalid` response. A node without generation management accepts only its enrolled `deployment_generation`; a generation-managing node runs the request on that generation's provider and answers `unconfirmed` when it cannot. Core sends `create` and a `resume` that is not observe-only only to a generation that is ready on that node, and keeps at most 32 requests pending per connection. + +The budget is relative: the node anchors `timeout_ms` to its own clock on receipt and consumes it while the request waits in its queue, so the hosts' clocks need not agree. Core still bounds its own wait. A full node queue closes the connection. + +The `response` frame carries `id`, `connection_id`, an `error_code` when the call failed, and on success exactly one result (`info`, `compute`, `state`, `command` or `sample`): + +| `error_code` | Meaning | +| --- | --- | +| `invalid`, `ownership`, `exists`, `not_found` | `ErrInvalid`, `ErrOwnership`, `ErrExists`, `ErrNotFound` | +| `command_unconfirmed` | `ErrCommandUnconfirmed` | +| `observation_unavailable`, `runtime_not_running` | The observation outcomes | +| `unsupported` | The operation is declared unsupported; see below | +| `unconfirmed`, or any other value | The outcome is unknown | + +A failed response carries no result, except an `info` that is an exact-reference `CreateSettled` receipt: a confirmed native Create that failed a later check can still prove that the attempt settled. A timeout, a lost response or a disconnect is unavailable or uncertain, never evidence of absence, and Core never replays a mutation after one; it observes the original operation instead. The [Sandbox Provider guide](../../docs/sandbox-provider.md#operation-outcomes-and-retries) defines each outcome. + +Node startup and generation loading validate complete Provider operation declarations before accepting work, and the Core proxy uses the same registered declaration, so an unsupported operation rejects before node resolution or native I/O. The [operation contract](../../docs/sandbox-provider.md#explicit-operation-contracts) owns the inventory. An `unsupported` response carries an `unsupported` object with the exact method `operation` and an authored safe `reason`; the proxy checks both against the request. Missing, malformed or mismatched evidence is an unconfirmed result, never proof that a mutation was rejected. Unsupported stays distinct from observation unavailability and unknown compute or command results, and it neither settles resource ownership nor authorizes a replay. + +## Generation control + +A node without generation management serves only its enrolled generation, with fixed configuration, and receives no preparation or retention frames. A generation-managing node prepares the target generation Core announces in `welcome` and `heartbeat_ack` and keeps serving its durable serving generation while it does; target preparation is independent of the serving provider's readiness. + +Its `hello` and heartbeats carry at most eight generation observations. Each names a positive signed-64-bit generation, its lowercase SHA-256 specification digest, a `ready`, `preparing` or `failed` state and an optional fixed diagnostic. The target and serving generations come first; other records rotate fairly. Eight bounds one message, not the number of generations a node may keep. An omitted observation never authorizes deletion or implies absence. + +Retention uses its own bounded exchange. A `retention` request names at most eight local `(generation, specification_digest)` references, a UUID, a sequence that increases by one, the current `connection_id` and the owner epoch. The `retention_ack` must match the complete pending request, entry order and identity included, and give an explicit boolean `keep` for every entry. One exchange is pending per connection, and a disconnect discards it. An unsolicited, replayed, stale, partial or mixed acknowledgement deletes nothing. Retention traffic never uses the Provider request queue. ## Local retention and helper lifetime -A Core drop grant is necessary but insufficient for collection. Queued and running -provider calls, preparation, the local target and serving pin retain references. -Collection rechecks those references and refuses an already canceled connection. -A canceled provider caller does not establish that its native helper has stopped: -the parent counts that helper through its actual `Wait` completion. - -Every generation also owns a permanent private lease file at -`state/node/generations/.lease`. Before starting a native helper, the -node acquires its shared flock and validates the durable lease identity under -that lock. It also requires the matching published final provider configuration -and absence of preparing, collecting and dropped journals. The helper inherits the descriptor. The node closes its own descriptor -only after `Wait`; it never explicitly unlocks the shared open-file description. -Thus caller cancellation or node exit does not release a live helper's reference. -New native helpers set this descriptor close-on-exec before calling the SDK so VM -and daemon descendants do not inherit a helper reference. - -Collection acquires the exclusive nonblocking flock before inspecting references, -removing shared images or release files, or publishing the dropped tombstone. It -keeps the lock through those changes. Lock files belong to stable node state and -are never removed or atomically replaced during collection. Symlinks, multiply -linked files, foreign ownership, unsafe permissions and replaced lock paths are -refused. Before the first helper can start, the installer exclusively creates the lease and -fsyncs it, then atomically persists and fsyncs a private `.lease-identity` record and -its parent directory. The record binds installation, generation, specification -digest, device and inode. Python exclusive collection/repair and Go shared helper -openers validate the same record on every open, including after node restart. -Neither opener adopts a missing identity, replaces its inode, or erases -it after GC. Initialization interrupted before the identity is durable refuses -re-adoption; preserve the installation for inspection. A removed identity or an -owned replacement 0600 lease still refuses, even when its current fstat/lstat agree. - -A dropped generation cannot be prepared or used again; a future rollback -would require a new generation and a separate policy. - -An immutable older native helper may pass its inherited descriptor to descendants. -That conservatively retains its generation's bytes until those descriptors close; -collection must not kill historical VMs or a host daemon merely to reclaim disk. -A helper's exit is local file-lifetime evidence, not proof that a remote mutation -or an uncertain provider receipt has been released. Core's durable allocation and -placement retention requirements remain independent. +A Core drop grant is necessary but not sufficient for collection. Queued and running provider calls, preparation, the local target and the serving pin all keep references; collection rechecks them and refuses on an already canceled connection. A canceled provider caller does not prove that its native helper stopped: the node counts a helper until its actual `Wait` returns. + +Every generation owns a permanent private lease file, `state/node/generations/.lease`. Before starting a native helper, the node takes a shared flock on it, validates the durable lease identity under that lock, and requires the published final provider configuration and no preparing, collecting or dropped journal. The helper inherits the descriptor; the node closes its own copy only after `Wait` and never unlocks the shared open-file description, so a canceled caller or a node exit does not release a live helper's reference. Native helpers set the descriptor close-on-exec before calling the SDK, so VM and daemon descendants do not inherit it. + +Collection takes the exclusive nonblocking flock before it inspects references, removes shared images or release files, or publishes the dropped marker, and holds it through those changes. Lease files belong to stable node state and are never removed or replaced during collection; symlinks, multiply linked files, foreign ownership, unsafe permissions and replaced lock paths are refused. Before the first helper can start, the installer creates the lease exclusively and fsyncs it, then atomically persists and fsyncs a private `.lease-identity` record and its directory. The record binds installation, generation, specification digest, device and inode. The Python collector and the Go helper opener validate the same record on every open, including after a restart; neither adopts a missing identity, replaces its inode or erases it after collection. An installation interrupted before the identity is durable refuses re-adoption and is kept for inspection. A removed identity or a replaced lease refuses even when its current metadata agree. + +A dropped generation can never be prepared or used again. A helper's exit is evidence about local files only, not proof that a remote mutation or an uncertain provider receipt has been released; Core's durable allocation and placement retention stays independent. ## Matched fresh installation -The host program release and Core's selected Runtime release are independent. -A fresh node gets its executable and private preparer from the current console -release. It reads the exact Runtime source, image identities and native runtime / -firmware digests from the authenticated Core configuration. If that Runtime is -older, the console must still serve its immutable `releases//` manifest, -checksums and allowlisted artifacts. Runtime helper, firmware, seccomp and image -bytes come from that selected release; the enrolled specification records it. -Artifact URLs are pinned to their verified manifest source even if the console's -current release changes during download. - -A missing retained release refuses installation rather than substituting the -current Runtime. A local bundle that contains only a different Runtime also -refuses with guidance to use the console origin retaining the selected release. -These refusals occur before writing the installation identity, importing the -Runtime, registering the node or starting its service. +The host program release and Core's selected Runtime release are independent. A fresh node gets its executable and private preparer from the console's current release, and reads the exact Runtime source, image identities and native runtime and firmware digests from its authenticated configuration. When that Runtime is older, the console still serves its immutable `releases//` manifest, checksums and allowlisted artifacts: the Runtime helper, firmware, seccomp profile and image bytes come from the selected release, and artifact URLs stay pinned to their verified manifest even if the console's current release changes during a download. + +A missing retained release refuses the installation rather than substituting the current Runtime, and so does a local bundle that holds only a different Runtime. These refusals happen before the installer writes the node identity, imports the Runtime, registers the node or starts its service. + +A published console release keeps its metadata and artifact bytes. Publishing it again first validates all metadata and every existing declared artifact, then may atomically add only missing, checksum-matched declared artifacts; any conflict prevents every addition, and nothing is overwritten. ## Restart recovery -Missing retained Runtime bytes do not switch a pinned placement to the current -Runtime. The node retains the original generation and specification digest as -unready state, and asks Core's authenticated configuration endpoint for that exact -kept generation before recovery. Missing seccomp bytes may leave an unready -provider placeholder; missing image or native artifacts discovered by a provider -probe queue repair without advertising readiness. - -Preparation and repair are serialized. Target and serving generations take -priority, with bounded progress over other retained generations. Each attempt has -a 30-minute deadline; failures back off for 1, 2, 5, 10 and then at most 30 minutes. -No connection-established deployment facts means no preparation starts. Repair -preserves existing configurations and paths, verifies the selected release and all -existing sibling checksums, and downloads only absent immutable files. Conflicting -bytes or a different retained specification refuse repair. A missing complete -provider configuration without an exact durable preparation plan remains a refusal rather than a guessed reconstruction. - -Repair takes the same exclusive generation lease and installation lock used by -collection. A live helper or concurrent collector therefore retains ownership; -repair retries later without replacing in-use files. Once bytes are restored, the -node still runs the actual provider readiness probe. File presence and executable -capability alone never establish serving readiness. - -## Interrupted local collection - -Before native or release deletion, the node persists a private collection journal -bound to its installation, generation and specification digest. Restart loads an -unfinished journal only as a retention-exchange candidate: it cannot prepare, -probe, acquire or advertise that generation. A fresh correlated Core drop grant -is required to resume; the journal itself never authorizes deletion. - -For an installation-private microsandbox store, a successful complete native image -inventory distinguishes absence from a CLI failure. Failed queries, malformed inventories, native in-use refusals and -unknown ownership retain the local bytes. Native completion is persisted before -release-file cleanup, so a retry can finish a partly removed release without -executing an already removed helper. Shared microsandbox images and private releases remain until their last local -reference. Docker imported images belong to the shared host daemon and are retained, -including when another installation has only an idle serving pin. Automatic node -GC never runs Docker image removal or pruning. The host administrator may remove -those images only after confirming that no installation on the host needs them; -local generation collection does not claim physical Docker image GC. The final digest-bound dropped marker follows durable file -cleanup and permanently prevents re-adoption. The small immutable configuration -and ownership journals remain as local identity records. - -Fresh installations also persist the verified Runtime file checksums separately -from the host program. Collection of the original generation removes only those -exact private Runtime helper, executable, firmware, seccomp and import-cache files -that no retained configuration references. All remaining files are checked before -the first deletion; unknown hashes, changed bytes, links or missing ownership -metadata refuse cleanup. Interruption resumes under the same collection journal. -The current node executable, preparer, identity, base provider configuration and -manifests remain, so restart can read its enrolled identity and construct a newer -retained provider after the original Runtime bytes have gone. Shared native paths -are compared across all retained configurations before removal. - -New preparation has two distinct records. Before downloads, `.preparing` holds the -immutable installation/generation/specification identity, private provider paths -and `import_started:false`; it is a recovery/collection plan, not a published -provider. Before invoking the importer, the same plan records `import_started:true`. -Both Python retention discovery and Go restart recovery recognize pending-only -plans, but never build, probe or acquire a provider from them. Current-connection -Core authorization is still required for recovery or collection. - -Only successful preparation publishes the final `.json` provider configuration. -Docker records the actual immutable local ID returned by the resolver; either -builder-proven config or manifest identity can be valid for the same specification. -The final configuration is write-once. Publication is durable before clearing the -preparation journal. Interruption between those steps revalidates the same plan -and existing final identity; it does not permit editing a published configuration. -Plan/spec/path drift refuses. A canceled or failed import remains visible to fresh -Core retention exchange without becoming a serving generation. - -An interrupted download repairs only missing bytes at the original paths. If -collection precedes any import attempt, the preparation journal proves that this -generation has no imported native image. An older or interrupted generation whose -native executable is missing and whose import may have started remains retained; -missing files do not prove native absence. Receipt/store history is never -erased using an empty native inventory. +Missing Runtime bytes never move a pinned placement to the current Runtime. The node keeps the original generation and specification digest as unready and asks Core's authenticated configuration route for that exact generation before recovery. Missing seccomp bytes may leave an unready provider placeholder; missing image or native artifacts found by a provider probe queue a repair without advertising readiness. + +Preparation and repair are serialized. Target and serving generations come first, with bounded progress on the others. Each attempt has a 30-minute deadline, and failures back off for 1, 2, 5 and 10 minutes, then at most 30. Nothing is prepared before the connection has delivered deployment facts. Repair keeps existing configurations and paths, verifies the selected release and every existing sibling checksum, and downloads only missing immutable files. Conflicting bytes or a different retained specification refuse repair, and a missing provider configuration without an exact preparation plan is refused rather than reconstructed. + +Repair takes the same exclusive generation lease and installation lock as collection, so a live helper or a running collector keeps ownership and repair retries later. After the bytes are restored, the node still runs the provider's readiness probe; file presence and executable capability never establish readiness. + +## Preparation and collection records + +New preparation writes two distinct records. Before downloads, `.preparing` holds the installation, generation and specification identity, the private provider paths and `import_started: false`; it is a recovery and collection plan, not a provider. Before the importer runs, the plan records `import_started: true`. Python retention discovery and Go restart recovery both recognize pending-only plans but never build, probe or acquire a provider from them, and recovery or collection still needs authorization on the current connection. + +Only a successful preparation publishes the final `.json` provider configuration, which is write-once. Docker records the immutable local image ID its resolver returns; either the builder's config identity or the manifest identity can be valid for the same specification. Publication is durable before the preparation journal is cleared. An interruption between the two revalidates the same plan and final identity; plan, specification or path drift refuses. A canceled or failed import stays visible to Core's retention exchange without becoming a serving generation. + +Before any native or release deletion, the node persists a private collection journal bound to its installation, generation and specification digest. After a restart an unfinished journal is only a retention-exchange candidate: it cannot prepare, probe, acquire or advertise that generation, and a fresh correlated Core drop grant is needed to resume. Native completion is persisted before release files are removed, so a retry can finish a partly removed release without running an already removed helper. The digest-bound dropped marker follows durable file cleanup and prevents re-adoption; the small configuration and ownership journals stay as local identity records. + +For the installation-private microsandbox store, a successful, complete native image inventory distinguishes absence from a CLI failure; a failed query, malformed inventory, native in-use refusal or unknown ownership keeps the bytes. Shared microsandbox images and private releases stay until their last local reference. Docker images belong to the host's shared daemon: automatic collection never removes or prunes them, and only the host administrator can remove them after confirming that no installation on the host needs them. + +A fresh installation also records the verified checksums of its Runtime files separately from the host program. Collecting the original generation removes only the exact private Runtime helper, executable, firmware, seccomp and import-cache files that no retained configuration references. Every remaining file is checked before the first deletion; an unknown hash, changed bytes, links or missing ownership metadata refuse cleanup. The node executable, preparer, identity, base provider configuration and manifests stay, so a restarted node can still read its enrolled identity and construct a newer retained provider. Shared native paths are compared across all retained configurations before removal. + +An interrupted download repairs only missing bytes at the original paths. When collection comes before any import attempt, the preparation journal proves that the generation has no imported native image. A generation whose native executable is missing and whose import may have started stays retained: missing files never prove native absence, and an empty native inventory never erases receipt or store history. The diagnostic codes are authored in `services/core/internal/sandbox/node_diagnostic.go`. The shared `services/core/internal/sandbox/testdata/node-diagnostics.json` fixture checks the Go mapping, OpenAPI source annotations and generated enums, and the TypeScript client declaration. Web uses the client normalizer and checks localized messages for every declared code. Update these projections with a code change; unknown codes normalize to `provider_unavailable`. -Preparation diagnostics preserve fixed typed causes. Only artifact transfer, -checksum or release-provenance failures report `runtime_download_failed`. A private -preparer exit category communicates that class without parsing stderr; provider, -ownership, cancellation and unclassified failures remain their existing typed -code or `provider_unavailable`. No raw provider text crosses the node protocol. - -Published console releases keep immutable metadata and existing artifact bytes. -A rerun first validates all published metadata and every existing declared artifact, -then may atomically add only absent, checksum-matched declared artifacts. This -supports thin-release completion and retained HTTP artifact repair. Any existing -conflict prevents all additions; repair never overwrites a conflicting artifact. - -## Node program version policy - -Node program upgrades, historical conversion and adoption are not supported. -`node-install.pyz --update` refuses before changing the installation. Preserve old -installation files, credentials, Runtime stores, provider resources and history; -install the current program separately through the ordinary fresh-node enrollment -flow. Reinstallation does not automatically delete or migrate existing data. - -Current Runtime generation changes operate within an installation of the current -node program. They do not upgrade that program or establish compatibility with -historical node installations. - -## Qualification boundary - -Protocol and process tests do not qualify Runtime readiness, VM coexistence, -or native image/store locking. Current-version Runtime behavior requires the -separate exact-artifact KVM and Docker acceptance matrix. Historical helper -behavior remains historical evidence, not a supported upgrade workflow; absence -of local test coverage is not evidence of successful garbage collection. +Preparation diagnostics keep fixed typed causes. Only artifact transfer, checksum or release-provenance failures report `runtime_download_failed`; the private preparer signals that class through its exit category, without Core or the node parsing stderr. Provider, ownership, cancellation and unclassified failures keep their typed code or `provider_unavailable`. No raw provider text crosses the protocol. diff --git a/contracts/agents-api/sandbox-deployment.md b/contracts/agents-api/sandbox-deployment.md index f5677a145..9cd204156 100644 --- a/contracts/agents-api/sandbox-deployment.md +++ b/contracts/agents-api/sandbox-deployment.md @@ -1,118 +1,43 @@ -# Managed sandbox deployment configuration +# Sandbox deployment -This deployment-administrator contract selects the provider, per-sandbox resources -and immutable Runtime for Core-managed `openai_hosted` execution. PostgreSQL owns -one active selection per installation. Web and the administrator API write the -same configuration. Node files contain its installed copy and host-specific paths; -they cannot override its resources or Runtime. +The sandbox deployment selects the Sandbox Provider, the per-sandbox resources and the immutable Runtime release for Core-managed `openai_hosted` execution. PostgreSQL holds one active selection per installation; Web and the Core API write the same configuration. A node's files hold an installed copy of it plus host-specific paths and cannot override its resources or Runtime. The selection is independent of the Harness, and a deployment can stay unconfigured, with no nodes and no hosted admission. -The public `/v1` Agent API, Environment Templates and caller-managed `self_hosted` -provisioning are unchanged. A provider selection is independent of the harness. -A deployment can remain unconfigured, with no execution nodes or hosted admission. +This contract owns the Core API routes below and their semantics. The [nodes guide](../../docs/getting-started/nodes.md) owns the operator workflow, the [machine connection API](machine-api.md#node-routes) the routes nodes call, and the [sandbox node protocol](node-generation-protocol.md) the node connection. -See [Nodes](../../docs/getting-started/nodes.md) -for the operator workflow. Generated schemas cover the -[administrator routes](core.openapi.yaml) and the -[node machine connection routes](runtime.openapi.yaml). +## Routes -## Authority and routes +Every route requires the Core key. [Web's console server](../../docs/web/console-server.md) adds it on the server side for signed-in requests. -| Method and route | Authority | Effect | -| --- | --- | --- | -| `GET /core/v1/sandbox/deployment` | Core key | Read the safe active configuration and retained-resource counts | -| `POST /core/v1/sandbox/deployment` | Core key | Select the initial provider, resources and Runtime | -| `PUT /core/v1/sandbox/deployment` | Core key | Advance the same-provider target online while retaining existing ownership | -| `POST /core/v1/sandbox/deployment/reset` | Core key | Start or escalate a durable hosted clear | -| `DELETE /core/v1/sandbox/deployment/reset?expected_generation=N` | Core key | Cancel the remaining clear without restoring archived work | -| `POST /core/v1/sandbox/providers/{provider}/discovery` | Core Web server or operator script with Core key | Query the registered provider's configuration catalog without saving credentials or allocating compute | -| `GET /api/v1/sandbox-node/configuration` | Enrollment token or retained node credential | Read the active node installation configuration without consuming enrollment | -| `POST /api/v1/sandbox-node/enroll` | One-use enrollment token | Register a node; consumes the token | -| `GET /api/v1/sandbox-node/identity?node_id=` | Node credential | Recover the node's registered identity, approved capacity and readiness | -| WebSocket `GET /api/v1/sandbox-node/connect?node_id=` | Node credential | The node's connection to Core | - -E2B discovery posts `{ "configuration": { "api_url": "https://sandbox.sandbase.ai", "domain": "sandbox.sandbase.ai" }, "credential": { "api_key": "..." }, "query": {} }` -to `/core/v1/sandbox/providers/e2b/discovery`. Use `query: {"template":"template-id"}` for builds. -Requests are limited to 64 KiB and 30 seconds. Docker and microsandbox explicitly -reject discovery with `400 sandbox_operation_unsupported`. Unknown input members -and null objects reject. A discovery result never proves deployment admission. -The official E2B endpoint may omit both endpoint fields. The first response is -`{ "templates": [{ "id": "...", "names": ["..."] }] }`; -the second is `{ "builds": [{ "id": "build-uuid", "cpus": 2, "memory_mib": 2048 }] }`. -Both lists may be empty. The Core key authenticates the caller; the E2B key is -used only for this request, is never stored by discovery, and is never returned. -Discovery uses the pinned SDK helper and its `GET /v2/templates` operation, -makes no allocation, and is capped at 200 -results. A limit or provider failure returns 503 with a generic message. The -deployment write separately validates the selected exact ready build. - -The paired console injects the Core key server-side on every signed-in `/core/v1` -request. The browser never receives that key. Node configuration -and enrollment are machine connection routes under `/api/v1`, which the reverse -proxy sends directly to Core; the console does not serve them. They use their own -Bearer credential; a console login, the Core key or a Project key does not grant -node enrollment authority. - -Node capacity is separate from the deployment specification. The administrator's -`POST /core/v1/sandbox/enrollment-tokens` accepts optional `max_active` and -`max_retained`, defaulting to 2 and 8. Microsandbox uses both limits. Docker never -suspends, so its `max_retained` always equals `max_active`; Core replaces any -submitted value, here and in `PATCH /core/v1/sandbox/nodes/{node_id}`. Core stores -that approval with the token and copies it to the registered node. The response -is `{token, expires_at, enrollment_id}`: `enrollment_id` is a public, non-secret -handle of that command and never a credential. Node enrollment -cannot submit capacity overrides. The existing node configuration/identity reads -expose approved limits for that credential; node updates remain administrator -operations. Local host capacity checks may reject a deployment that cannot run -safely, but never raise its limits. - -Each node in `GET /core/v1/sandbox/nodes` and its detail reports `enrollment_id`, -the handle of the command that registered it (null for nodes enrolled before Core -recorded it), and `core_url`: the installation public URL when the node enrolled. A node whose `core_url` differs from -the current public URL receives no new placements. Work already placed on it -finishes there: a hosted Environment that was placed but not yet allocated before -the change is still allocated on that node, and its retained sandboxes can still -resume, while the old address reaches Core. Remove it and add it again. - -`GET /core/v1/sandbox/nodes/{node_id}/allocations` lists the node's unreleased -allocations. Each item's `compute_phase_changed_at` is the time the allocation -entered its current `compute_phase`, or null when unknown; an allocation that -existed before Core recorded it reports null until its next phase change. A -suspended microsandbox allocation's age, combined with the deployment's snapshot -retention, tells roughly when Core reclaims it. - -### Hosted and application-managed placement - -The official `openai_hosted` discriminator means hosting by this independent Core -service, using its configured hosted Provider. Keep the public value unchanged; a product-named hosted value is not -a new API type. Both placements resolve the same Environment Templates: hosted -requests use the official field, and self-hosted requests use -`x_agents_core.environment.environment_template_id`. See -[Environment preparation](environments.md) for the shared snapshot contract. -E2B onboarding follows the application-managed `self_hosted` resource workflow: -the application owns sandbox provisioning and cleanup, and our daemon connects -with the returned Environment ID, unchanged `remote_url` and scoped environment -authorization. Reuse the same Runtime and thin provider components. No OpenAI -executor process or additional execution architecture is required. Qualify -principal/tenant ownership, credentials and connection lifecycle using the pinned -client and actual execution; document our transport boundary explicitly. +| Route | Effect | +| --- | --- | +| `GET /core/v1/sandbox/deployment` | Read the safe active configuration, rollout, reset and resource counts | +| `POST /core/v1/sandbox/deployment` | Select the initial provider, resources and Runtime | +| `PUT /core/v1/sandbox/deployment` | Change the same provider's target online | +| `POST /core/v1/sandbox/deployment/reset` | Start or escalate a durable clear of hosted resources | +| `DELETE /core/v1/sandbox/deployment/reset?expected_generation=N` | Cancel the remaining clear | +| `POST /core/v1/sandbox/providers/{provider}/discovery` | Query a provider's configuration catalog with a transient credential | +| `POST /core/v1/sandbox/enrollment-tokens` | Issue a one-use node enrollment token with approved capacity | +| `GET /core/v1/sandbox/nodes` | List registered nodes | +| `GET /core/v1/sandbox/nodes/{node_id}` | Read one node with its [host observations and history](runtime-observability-api.md#node-host-observations-and-history) | +| `PATCH /core/v1/sandbox/nodes/{node_id}` | Change a node's name and capacity | +| `DELETE /core/v1/sandbox/nodes/{node_id}` | Remove a node | +| `GET /core/v1/sandbox/nodes/{node_id}/allocations` | List a node's unreleased allocations | +| `GET /core/v1/sandbox/runtime-observations` | Current Runtime observations; see the [Runtime telemetry API](runtime-observability-api.md) | ## Selection request -POST and PUT take the same complete selection and require `expected_generation` -from a preceding GET. Zero is valid for the initial unconfigured deployment; -omitted or null is invalid. A stale generation is checked before reset, provider, -resource and same-selection conditions, including an identical old request body. +POST and PUT take the same complete selection and require `expected_generation` from a preceding GET. Zero is valid for the initial unconfigured deployment; omitted or null is invalid. A stale generation is checked before reset, provider, resource and same-selection conditions, even for a body identical to an earlier request. | Field | Meaning | | --- | --- | -| `expected_generation` | Required nonnegative integer from GET; never automatically refresh and replay | +| `expected_generation` | Required nonnegative integer from GET; never refresh and replay it automatically | | `provider` | Exactly one of `docker`, `microsandbox`, `e2b` | -| `resources` | Per-sandbox resource limits described below; required for Docker/microsandbox, optional for E2B | -| `runtime` | Required immutable distribution identity for Docker/microsandbox; absent for E2B | -| `configuration` | Provider-owned public selectors. E2B accepts immutable `template` and optional paired `api_url`/`domain`; node providers accept only `{}` or omission. | -| `credential` | Write-only provider credential object. E2B accepts `{api_key}`; required at first setup, omitted on PUT to preserve the key. Null and empty keys reject. Node providers reject this object. | +| `resources` | Per-sandbox limits, below; required for Docker and microsandbox, optional for E2B | +| `runtime` | The immutable [Runtime release](#runtime-release); required for Docker and microsandbox, absent for E2B | +| `configuration` | The provider's public selectors. E2B: the immutable `template` build and the optional paired `api_url` and `domain`. Docker and microsandbox accept only `{}` or omission | +| `credential` | The provider's write-only credential. E2B: `{api_key}`, required at first setup and omitted on PUT to keep the current key; a null or empty key is invalid. Docker and microsandbox reject it | -The request has no Core address. Core derives the deployment's `core_url` from the installation public URL (`public_url` in `config.json`, `OAC_PUBLIC_URL` for Core): the HTTPS origin nodes and sandbox guests use to reach Core. A request that contains `core_url` is rejected with 400 `invalid_request` like any other unknown member. E2B guests reach Core from E2B's cloud, so an E2B selection is rejected with 409 `sandbox_configuration_error` while the public URL is loopback. Docker and microsandbox selections do not depend on the address; a loopback public URL serves local development only, because a guest's loopback address does not reach its host. Changing the public URL is an installation change, not this operation: nodes enrolled with the old address receive no new sandboxes and must be removed and added again. +The request has no Core address. Core derives the deployment's `core_url` from the installation public URL (`public_url` in `config.json`, `OAC_PUBLIC_URL` for Core): the origin nodes and sandbox guests use to reach Core. A request that contains `core_url` is rejected with 400 `invalid_request` like any other unknown member. E2B guests reach Core from E2B's cloud, so an E2B selection is rejected with 409 `sandbox_configuration_error` while the public URL is loopback. Docker and microsandbox selections accept a loopback public URL, which serves only local development because a guest's loopback address does not reach its host. Changing the public URL is an installation change: nodes enrolled with the old address receive no new sandboxes and must be removed and added again. ### Resources @@ -120,432 +45,183 @@ The request has no Core address. Core derives the deployment's `core_url` from t | --- | --- | | `cpus` | Integer, 1 through 255 | | `memory_mib` | Integer, 512 through 1048576 MiB | -| `root_disk_mib` | Microsandbox: at least 1024 MiB; Docker/E2B: omitted or zero | -| `environment_disk_mib` | Microsandbox: at least 1024 MiB; Docker/E2B: omitted or zero | - -These limits describe each sandbox. Node `max_active` and `max_retained` remain -separate reservation limits; host telemetry is an observation, not permission to -exceed either limit. Native providers may reject values that pass these structural -bounds. - -Docker applies CPU and memory limits and checks the running container's limits -and exact image identity. It does not provide independent hard root or workspace -disk quotas through this contract. E2B CPU and memory must match the exact ready -template build; Core validates that build through the pinned SDK before saving. -E2B disk capacity remains part of its native template. Neither provider silently -accepts a requested disk quota that it cannot enforce. An E2B selection may omit -`resources`: Core then validates that the build is ready and stores its CPU count -and memory as `cpus` and `memory_mib`, returned in `specification.resources` -without disk fields. The stored specification is complete and validated either way. On restart, Core loads the -committed E2B selection without repeating template-build validation. Existing -resource inspection and cleanup use the original credentials and provider -receipts; new selections still require successful template validation. - -Microsandbox configures CPU, memory, managed root disk and a separate owned disk -at `/environment`. Restored root capacity can be inherited from the verified full -snapshot when native restore metadata omits a configured size. That exception -requires the original resource proof and exact snapshot/target identity; it does -not resize a restored root or treat a missing size as arbitrary capacity. Other -native limits must still match. A failed readiness or configuration check retains -ownership and cleanup records. -An interrupted restore can finish the derived proof on its existing, verified -target; it cannot create another instance or change limits. Snapshot observation -for cleanup verifies ownership and artifact identity without requiring that the -source still qualify for execution. +| `root_disk_mib` | microsandbox: at least 1024 MiB; Docker and E2B: omitted or zero | +| `environment_disk_mib` | microsandbox: at least 1024 MiB; Docker and E2B: omitted or zero | + +These limits describe each sandbox. A node's `max_active` and `max_retained` are separate reservation limits, and host measurements never permit exceeding either. Native providers may reject values that pass these bounds. + +Docker applies the CPU and memory limits and checks the running container's limits and exact image; it has no hard root or workspace disk quota. E2B CPU and memory must equal the exact ready template build, which Core validates through the pinned SDK before saving; disk capacity stays part of the template. Neither provider accepts a disk quota it cannot enforce. An E2B selection may omit `resources`: Core then stores the build's CPU count and memory as `cpus` and `memory_mib`, returned in `specification.resources` without disk fields. On restart Core loads the committed E2B selection without validating the template build again, so an E2B outage never blocks inspection or cleanup; new selections still require validation. + +microsandbox configures the CPUs, memory, a managed root disk and a separate owned disk at `/environment`. The [microsandbox helper](../../services/core/tools/microsandbox-provider/README.md) describes how restore handles these limits. ### Runtime release -Docker/microsandbox use all fields of one verified distribution: +Docker and microsandbox use every field of one verified distribution: | Field | Identity | | --- | --- | | `source_commit` | Lowercase 40-character commit SHA | -| `image_id` | Docker image configuration ID, `sha256:` followed by 64 lowercase hex characters | -| `image_manifest_digest` | OCI image manifest digest, in the same `sha256:` form | -| `microsandbox_ref` | `oac-runtime@sha256:` followed by 64 lowercase hex characters | -| `runtime_sha256` | SHA-256 of the native microsandbox Runtime binary | -| `firmware_sha256` | SHA-256 of the matched firmware | - -Copy these identities from the matched distribution manifest. Image configuration -IDs and OCI manifest digests identify different objects; do not substitute one -for the other. The node installer verifies the saved release against its payload -before registration and retains the exact local image identity it imports. - -E2B instead uses `configuration.template` in `template-id:build-uuid` form. The build UUID must -be canonical and nonzero; a mutable template alias alone is insufficient. Omit -`runtime`. The API key is encrypted in PostgreSQL and never returned in a safe -view, bootstrap configuration, command argument or log. Same-team key, build or endpoint -changes apply online while old sandboxes retain their original specification. -By default Core uses `https://api.e2b.app` and `e2b.app`. For a compatible -service, set both `configuration.api_url` (HTTPS API origin, with no path, port, query, -fragment or credentials) and `configuration.domain` (sandbox data-plane DNS suffix). -The API host must equal the data-plane domain or be its subdomain. Core rejects -a sandbox response whose data-plane domain lies outside the selected suffix -before sending daemon credentials or using envd. Existing sandboxes retain -their original endpoint and credential across online changes. - -The operator contract uses these fields directly; the retired `e2b` request/response -member and vendor-specific discovery routes have no fallback. Migration 91 moves -existing selectors and observations into the generic objects without changing -ciphertext, generation or immutable retained ownership. Downgrade refuses unknown -provider configurations that the old schema cannot represent. The pinned `/v1` -Agents API is unchanged. +| `image_id` | Docker image configuration ID: `sha256:` and 64 lowercase hex characters | +| `image_manifest_digest` | OCI image manifest digest, in the same form | +| `microsandbox_ref` | `oac-runtime@sha256:` and 64 lowercase hex characters | +| `runtime_sha256` | SHA-256 of the native microsandbox runtime binary | +| `firmware_sha256` | SHA-256 of the matching firmware | + +Copy these identities from the matching distribution manifest. An image configuration ID and an OCI manifest digest identify different objects and never substitute for each other. The node installer verifies the saved release against its payload before registration and keeps the exact local image identity it imports. + +### E2B configuration + +E2B uses `configuration.template` in `template-id:build-uuid` form; the build UUID must be canonical and nonzero, and a mutable template alias alone is refused. The API key is encrypted in PostgreSQL and never appears in a response, bootstrap configuration, command argument or log. By default Core uses `https://api.e2b.app` and `e2b.app`. For a compatible service, set both `configuration.api_url` (an HTTPS origin without path, port, query, fragment or credentials) and `configuration.domain` (the sandbox data-plane DNS suffix); the API host must equal the domain or be a subdomain of it. Core rejects a sandbox whose data-plane domain lies outside the selected suffix before sending daemon credentials or using envd. + +### Configuration discovery + +`POST /core/v1/sandbox/providers/{provider}/discovery` takes exactly `configuration`, `credential` and `query` objects, at most 64 KiB, and runs for at most 30 seconds. Unknown members and null objects are rejected. The credential is used only for that request and is never stored or returned; discovery saves nothing, changes no deployment, allocates nothing and never proves that a selection will be admitted. Docker and microsandbox reject it with 400 `sandbox_operation_unsupported`. + +For E2B, post `{"configuration": {"api_url": "…", "domain": "…"}, "credential": {"api_key": "…"}, "query": {}}` to list templates, and add `"query": {"template": "template-id"}` to list that template's ready builds; the official endpoint may omit both endpoint fields. The results are `{"templates": [{"id": "…", "names": ["…"]}]}` and `{"builds": [{"id": "build-uuid", "cpus": 2, "memory_mib": 2048}]}`, either possibly empty. The pinned SDK helper reads `GET /v2/templates` and returns at most 200 results; a limit or provider failure returns 503 with a generic message. The deployment write validates the selected build separately. ## Safe response -GET and successful mutations return `installation_id`, `provider`, `core_url` -(read-only: the installation public URL, present before configuration), `mode`, -`generation`, `owner_epoch`, `reset`, `rollout`, `suspension` and resource -accounting. A configured deployment also returns `specification` and -`specification_digest`, `configuration` and `metadata`. Every response includes -`credential_configured`. E2B returns `configuration.template`, `configuration.api_url`, -`configuration.domain` and optional `metadata.template_build`. Docker and microsandbox -return empty configuration/metadata objects and `credential_configured: false`. -Unconfigured deployments omit both objects. Native configuration values and secrets -are never directly serialized; only the adapter's public projection is returned. - -`metadata.template_build` is `{status, resources: {cpus, memory_mib, root_disk_mib}}`: -the fixed build as Core read it through the pinned SDK when the selection was -saved. GET does not call E2B, so it stays cheap and cannot fail on an E2B outage; -the values describe the immutable build at selection time. Validation admits only -a `ready` build whose CPU count and memory equal the selected `cpus` and -`memory_mib`. `root_disk_mib` is the build's native disk size, which Core does not -enforce separately. Unknown observed values are null. `metadata: {}` means no build observation -was recorded. A verified write records the observation. Existing partial observations -are preserved across schema upgrades and read projections. An omitted-key -identical PUT is a no-op and does not refresh provider metadata. - -`suspension` is `{idle_seconds, retention_seconds}` for microsandbox, the only -provider Core suspends (currently 300 and 86400). Docker, E2B and unconfigured -deployments return null. - -Two similarly named fields have different purposes: - -- Request `resources` and response `specification.resources` contain per-sandbox - CPU, memory and supported disk limits. -- Response `resources.allocations` and `resources.pending` count unreleased - allocations and pending hosted Environments without an allocation. - -An unconfigured deployment has an empty provider and no specification. Docker and -microsandbox use `mode: nodes`; E2B uses `mode: direct` without a synthetic node. -`generation` identifies the saved selection. `owner_epoch` fences execution-owner -and node connections; it is not a replacement for `expected_generation`. -Treat `specification_digest` as the server-provided identity of the provider, -resource limits and Runtime release. Enrollment echoes it unchanged. - -The typed `SandboxAdminClient` checks the deployment, node list, node detail and -allocation responses against exactly these shapes. An unknown or missing member, -or a wrong type, rejects the whole response with a 502 `invalid_admin_response` -error. A node's `diagnostic` is absent or a code, never empty; the client reads -an unknown code as `provider_unavailable`. - -## Initialization, same-provider changes and reset - -POST validates a candidate before persistence and creates no compute, Session or -model request. At the current generation an identical selection is a no-op; an -obsolete generation returns 409 `sandbox_generation_stale`, even for the same body. -Missing prerequisites or failed preparation leave the committed provider intact. -A different backend, or an old node selection without a specification, requires -reset first. POST also initializes after a completed reset, using its new generation. - -PUT accepts the same provider and no active reset. Send the observed generation -once; never replay an uncertain mutation automatically. E2B changes apply online: -new allocations use the newly committed specification and existing allocations -retain their immutable deployment generation. No node, token or owner epoch is -retired by a same-provider update. Docker/microsandbox advance only the target; -each node prepares it independently while continuing to serve its qualified old -pin. No execution drain or reenrollment accompanies a target change. - -For E2B PUT, omit the `credential` object to preserve the current key. Omitted-key identical -selection is a no-op. Explicit nonempty key submission, including the same key, -always verifies and advances generation. Null or empty keys are invalid. A key-only -change uses the same full DTO: provider, existing template, optional resources and -expected_generation, plus the new api_key. It has no separate route or implicit reset. - -Initial setup requires the selected template to appear in the credential's team-owned -template listing; public readability alone is insufficient. Before an online change, -Core verifies that the committed key owns the current template, then requires the -candidate key to own that exact template as a shared ownership anchor. It also reads -the candidate and every retained build at its original endpoint with the candidate key and confirms each -settled live receipt in the installation-labelled sandbox listing. - -A legacy public-template selection without this ownership anchor, or a committed key -that no longer authenticates, requires `409 sandbox_reset_required`; Core cannot -establish a safe online replacement from that state. Keep the old key valid until the -successful response. A candidate outside the verified team gives `409 sandbox_credential_ownership`; -explicitly reset before initializing another team. Candidate-key 401/403 gives -`400 sandbox_credential_invalid`; an invalid candidate build gives -`400 sandbox_configuration_invalid`. Missing or unsettled receipts and unconfirmed reads -fail closed with `503 sandbox_verification_unconfirmed`. No provider text or credential is -returned. The write and `change` or `replace_credential` audit share one transaction. - -A credential replacement briefly fences provider calls, waits for actual helper -process completion even after caller cancellation, and repeats verification before -commit. Helper exit is not evidence that remote Create settled. The fence and reads -are bounded; failure preserves the old key and lifecycles. After the successful -response, all retained-generation management uses the committed key. Only then -revoke the previous key in E2B. Template/resource changes do not drain lifecycles. - -### Generation ownership and rollout - -`runtime_deployment` owns the current specification. Superseded rows contain only -immutable specification/build/endpoint metadata, never another E2B credential. An E2B -allocation binds its generation at reservation. Node placements bind at Session -admission and allocations copy that generation, even after repeated updates. -Inspection, renewal, command execution and cleanup route the allocation's original -specification and endpoint with the current credential; a missing generation never falls back -to the current specification. Released historical generation identifiers remain. - -Retain a generation while it is current, referenced by an unreleased allocation or -placement, or pinned by a nonremoved node. The durable node serving pin survives -offline state and zero resources; it is separate from current connection readiness. -Collection shares the deployment lock with updates and admission and deletes at -most 32 eligible generation rows per pass. Reset retires pins and clears superseded -rows only after confirmed resource release. Downgrade refuses ownership or serving -pins that still need generation routing. - -Every deployment response includes: +GET and successful writes return `installation_id`, `provider`, `core_url` (read-only: the installation public URL, present even before configuration), `mode`, `generation`, `owner_epoch`, `reset`, `rollout`, `suspension`, `resources` and `credential_configured`. A configured deployment also returns `specification`, `specification_digest`, `configuration` and `metadata`: the adapter's public projection of its selectors and of the observations it recorded, never raw stored values or secrets. E2B returns `configuration.template`, `configuration.api_url`, `configuration.domain` and, once recorded, `metadata.template_build`. Docker and microsandbox return empty `configuration` and `metadata` objects and `credential_configured: false`; an unconfigured deployment has neither object. + +- `metadata.template_build` is `{status, resources: {cpus, memory_mib, root_disk_mib}}`: the build as Core read it through the pinned SDK when the selection was saved. GET never calls E2B, so it stays cheap during an E2B outage. Validation admits only a `ready` build whose CPU count and memory equal the selected values; `root_disk_mib` is the build's native disk size, which Core does not enforce. Unknown values are null, `metadata: {}` means no observation was recorded, and an identical PUT without a credential does not refresh it. +- `suspension` is `{idle_seconds, retention_seconds}` for microsandbox, the only provider Core suspends (currently 300 and 86400); Docker, E2B and unconfigured deployments return null. +- Request `resources` and response `specification.resources` are per-sandbox limits. Response `resources.allocations` and `resources.pending` count unreleased allocations and pending hosted Environments without an allocation. +- An unconfigured deployment has an empty provider and no specification. Docker and microsandbox use `mode: nodes`; E2B uses `mode: direct`, without a synthetic node. +- `generation` identifies the saved selection. `owner_epoch` fences the execution owner and node connections; it does not replace `expected_generation`. +- `specification_digest` is the server's identity of the provider, limits and Runtime release; enrollment echoes it unchanged. + +The typed `SandboxAdminClient` in `packages/agents-client` checks the deployment, node list, node detail and allocation responses against exactly these shapes. An unknown or missing member, or a wrong type, rejects the whole response with a 502 `invalid_admin_response` error. A node's `diagnostic` is absent or a code, never empty, and the client reads an unknown code as `provider_unavailable`. + +## Initial setup and same-provider changes + +POST validates a candidate before persisting it and creates no compute, Session or model request. At the current generation an identical selection is a no-op; an old generation returns 409 `sandbox_generation_stale`, even for the same body. Missing prerequisites or a failed preparation leave the committed provider in place. A different backend requires a reset first. POST also initializes after a completed reset, using the reset's new generation. + +PUT accepts the same provider while no reset is active. Send the observed generation once and never replay an uncertain write automatically. New allocations use the newly committed specification, and existing allocations keep their immutable deployment generation. A same-provider change retires no node, token or owner epoch and drains no execution: Docker and microsandbox nodes prepare the new target independently while serving their old pin, and E2B changes apply at once. + +### E2B key replacement + +Omit `credential` on PUT to keep the current key; an identical selection without it is a no-op. Submitting a key, even the same one, always verifies it and advances the generation; a null or empty key is invalid. A key-only change uses the same full body (provider, existing configuration, optional resources, `expected_generation` and the new `credential`), with no separate route or implicit reset. + +Initial setup requires the selected template to appear in the key's team-owned template listing; public readability is not enough. Before an online change, Core verifies that the committed key owns the current template, then requires the candidate key to own that exact template, which anchors team ownership. It also reads the candidate build and every retained build at its original endpoint with the candidate key, and confirms each settled live receipt in the installation-labelled sandbox listing. + +| Result | Meaning | +| --- | --- | +| `409 sandbox_reset_required` | The current selection has no team ownership anchor, or the committed key no longer authenticates | +| `409 sandbox_credential_ownership` | The candidate key belongs to another team; reset before initializing another team | +| `400 sandbox_credential_invalid` | The candidate key gets 401 or 403 | +| `400 sandbox_configuration_invalid` | The candidate build is invalid or does not match the resources | +| `503 sandbox_verification_unconfirmed` | A receipt is missing or unsettled, or a read is unconfirmed | + +No provider text or credential is returned. The write and its `change` or `replace_credential` audit entry share one transaction. A replacement briefly fences provider calls, waits for helper processes to actually exit even after the caller cancelled, and verifies again before committing; a helper's exit does not prove that a remote Create settled. The fence and reads are bounded, and a failure keeps the old key and lifecycles. After the successful response, all retained-generation management uses the committed key; only then revoke the old key in E2B, never before cleanup. Template and resource changes do not drain lifecycles. + +## Generation ownership and rollout + +`runtime_deployment` holds the current specification. Superseded rows keep only immutable specification, build and endpoint metadata, never another E2B key. An E2B allocation binds its generation at reservation; a node placement binds at Session admission, and its allocation copies that generation, even across later updates. Inspection, renewal, commands and cleanup use the allocation's original specification and endpoint with the current key; a missing generation never falls back to the current specification. Released generation identifiers stay reserved. + +A generation is retained while it is current, referenced by an unreleased allocation or placement, or pinned by a node that is not removed. A node's durable serving pin survives offline periods and zero resources and is separate from its current readiness. Collection shares the deployment lock with updates and admission and deletes at most 32 eligible generation rows per pass. Reset retires pins and clears superseded rows only after confirmed release. + +Every deployment response includes `rollout`: ```json -"rollout": {"state":"settled", "previous_generation_sandboxes":3, - "nodes":null} +"rollout": {"state": "settled", "previous_generation_sandboxes": 3, "nodes": null} ``` -The previous count uses the same snapshot as resource totals: old-generation -unreleased allocations plus old-generation pending placements without allocations, -never counting one resource twice. E2B and unconfigured deployments return null -nodes. A node deployment returns counts `{ready,preparing,failed,update_required,unknown}`. -Every nonremoved node belongs to exactly one bucket. Offline current connections -are `unknown` first; an online v1 node enrolled at an older target is -`update_required`. Otherwise the exact target observation on the current -connection and owner epoch supplies `ready`, `preparing` or `failed`; missing -observations are `unknown`. The existing 45-second heartbeat predicate defines -online. Pins alone never imply readiness. Target rollout is independent of serving -readiness: unknown/preparing/failed/update_required does not erase an independently -confirmed old serving provider. `provider_ready` requires online presence and an -exact serving-generation observation on that connection and epoch. A v1 heartbeat -qualifies only its enrolled provider. - -Each node adds `rollout: {state, ready_generation, diagnostic?}`. ready_generation is -the nullable durable serving pin. Diagnostic is a fixed node reason; unknown values -project as `provider_unavailable`. Allocation items add `deployment_generation`. -`rollout.state` is the authoritative high-frequency polling signal: poll every five -seconds only while it is `preparing` or reset is nonnull. Old Sessions, failed, -update-required and offline nodes alone do not keep polling active. New admission -filters online, exact serving-generation readiness, address and shared capacity -before choosing the newest qualifying pin. A full newest node does not hide a free -older node. No candidate creates no provisional Session or placement. An online -node actually preparing with available capacity returns 503 `sandbox_nodes_preparing`; -a full or offline fleet returns `runtime_node_unavailable`. - -To change backend, explicitly start reset: +`previous_generation_sandboxes` counts, from the same snapshot as the resource totals, unreleased allocations and allocation-less pending placements of older generations, never one resource twice. E2B and unconfigured deployments return null `nodes`. A node deployment returns `nodes` as counts `{ready, preparing, failed, update_required, unknown}`, and every node that is not removed falls in exactly one bucket: + +- a node that is offline on its current connection is `unknown`; +- an online node without generation management that enrolled at an older target is `update_required`; +- otherwise the exact target observation on the current connection and owner epoch gives `ready`, `preparing` or `failed`, and a missing observation gives `unknown`. + +A node is online while it is connected under the current owner epoch with a heartbeat in the last 45 seconds. Target rollout is independent of serving readiness: `unknown`, `preparing`, `failed` or `update_required` does not remove an independently confirmed older serving generation. `provider_ready` requires online presence and an exact observation of the serving generation on that connection and epoch; a node without generation management qualifies only its enrolled generation. Pins alone never imply readiness. + +Each node adds `rollout: {state, ready_generation, diagnostic?}`, where `ready_generation` is the nullable durable serving pin and `diagnostic` a fixed code for the target generation; allocation items add `deployment_generation`. Poll every five seconds only while `rollout.state` is `preparing` or `reset` is not null; old Sessions and failed, update-required or offline nodes alone do not keep polling active. + +New admission filters nodes by online presence, exact serving-generation readiness, address and shared capacity before it prefers the newest qualifying pin, so a full newest node never hides a free older one. Without a candidate, admission creates no provisional Session or placement: an online node that is actually preparing with free capacity gives 503 `sandbox_nodes_preparing`, and a full or offline fleet gives `runtime_node_unavailable`. + +## Reset + +To change backend, start a reset: ```json {"expected_generation": 7, "clear": "auto", "deadline_seconds": 3600} ``` -`clear` is required: `auto` or `force`. Auto's `deadline_seconds` defaults to 3600 -and accepts 300–86400; force must omit it. Core persists an absolute deadline and -requester audit provenance before closing fresh hosted admission. The existing -execution owner advances the clear after the request ends and across restarts. -Auto archives idle hosted Sessions, including queued or pending work and suspended -sandboxes, but waits for root/subagent Turns in progress or waiting and pending -file writes. It rechecks this condition under the Session lock. At the persisted -deadline it durably escalates to force; force uses the ordinary archive cancellation -and cleanup path for all eligible hosted Sessions. Self-hosted Sessions are excluded. +`clear` is required: `auto` or `force`. `deadline_seconds` applies to `auto`, defaults to 3600 and accepts 300 to 86400; `force` must omit it. Core persists an absolute deadline and the requester's audit provenance before it closes fresh hosted admission, and the execution owner advances the clear after the request ends and across restarts. `auto` archives idle hosted Sessions, including queued or pending work and suspended sandboxes, and waits for root and Subagent Turns that are in progress or waiting and for pending file writes, rechecking under the Session lock. At the persisted deadline it escalates to `force` durably. `force` uses the ordinary archive cancellation and cleanup path for every eligible hosted Session. Self-hosted Sessions are never touched. -Repeated start at the current generation/mode keeps the original deadline. Auto -can escalate to force, but force cannot downgrade to auto. DELETE with the current -`expected_generation` cancels remaining reset work and reopens admission; it does -not undo archives, revive expired Environments or cancel cleanup already requested. -DELETE with no active reset is idempotent. Stale requests still return 409. +Starting again at the same generation and mode keeps the original deadline. `auto` can escalate to `force`, never the reverse. DELETE with the current `expected_generation` cancels the remaining work and reopens admission; it does not undo archives, revive expired Environments or cancel cleanup already requested. DELETE without an active reset is a no-op, and a stale request returns 409. -Every deployment response includes `reset: null` when inactive, or: +`reset` is null when inactive, otherwise: ```json -{"clear":"auto","requested_at":"2026-09-27T12:00:00Z", - "deadline_at":"2026-09-27T13:00:00Z","forced_at":null, - "remaining":{"busy":2,"idle":1,"cleanup":3,"on_offline_nodes":2, - "offline_nodes":[{"node_id":"node-uuid","name":"worker","resources":2}]}} +{"clear": "auto", "requested_at": "2026-09-27T12:00:00Z", + "deadline_at": "2026-09-27T13:00:00Z", "forced_at": null, + "remaining": {"busy": 2, "idle": 1, "cleanup": 3, "on_offline_nodes": 2, + "offline_nodes": [{"node_id": "node-uuid", "name": "worker", "resources": 2}]}} ``` -Timestamps are explicit nullable fields where applicable. A single database snapshot -partitions every unreleased allocation and every pending hosted Environment with -no allocation into cleanup first, then busy or idle. -`busy + idle + cleanup == resources.allocations + resources.pending`. Deleted or -expired Sessions with unreleased receipts still count as cleanup. `offline_nodes` -is the untruncated, ID-sorted subset attributed to receipt or active-placement node -identity; its resource sum equals `on_offline_nodes`. Presence uses the current -owner epoch, connection and a heartbeat within 45 seconds, not provider readiness. -Direct E2B resources have no node and do not enter that subset. Offline resources -remain blockers until actual cleanup confirms release. - -Only after both held counts reach zero does the owner drain and atomically clear -provider/mode/specification, E2B credential/template/build metadata and provider -policy, retire nodes and unused enrollment tokens, increment generation and owner -epoch, and record `reset_complete`. Installation identity and history survive. -The manager publishes an unconfigured state immediately; its cache keeps the new -generation even with no provider, preventing delayed old loads from reviving it. -Configure again with POST using the returned generation; no restart is necessary. - -Fresh hosted admission returns 503 `sandbox_reset_in_progress` and leaves no -provisional Session rows. Existing live input, receipt retries, restoration and -cleanup continue. Management writes and new enrollment/configuration return 409 -`sandbox_reset_in_progress`; retained matching nodes can recover for cleanup. -Explicit per-Session archive requires the current generation and remains available -without reset. It preserves history and persisted Files/Artifacts but discards -unpersisted workspace and prevents that Session from resuming. Poll its archive -GET for actual release. Reset never manufactures a release receipt. - -A force archive fences credentials and new work immediately. If its original -hosted delivery is still connected, Core preserves only that delivery's native -cancellation/terminal receipt path until terminal commit or a fixed 20-second -bound from the original cancellation request. Done does not end this bound while -a cancellation acknowledgment or terminal commit is pending. This internal drain -never authorizes reconnect, workspace/MCP access or renewed execution. Explicit -credential revocation ends the exception; repeats do not extend it. Missing or -failed receipts retain honest failure outcomes, and disconnected, expired or -restarted owners fall back to ordinary provider cleanup. There is no new public -state, request field or model/tool timeout. - -One mutation gate serializes setup, PUT, reset, cancel and finalization. Archive -locks Session before deployment; finalization never reverses that order or waits -for itself inside counted manager work. Candidates bind to the reset request time, -so cancel followed by a new reset at the same generation cannot reuse old work. -Never automatically replay a rejected or uncertain write: read current state and -make a new explicit decision. Do not revoke old E2B credentials before cleanup. - -## Node configuration and enrollment - -For a new node, send `Authorization: Bearer ` to the configuration -GET without `X-OAC-Node-ID`. The token must be valid, unexpired, unconsumed and -belong to this installation. This read does not consume it. An active reset prevents -new enrollment configuration reads. - -An already registered node sends its durable node credential as Bearer and its -UUID in `X-OAC-Node-ID`. Its original enrollment identity and installation remain -valid across same-provider target changes. Omitted `generation` reads the current -target; `?generation=N` reads only that node's exact current, serving-pinned or -unreleased-allocation/placement generation. Unknown or unkept history is refused. -This read remains available during reset for owned recovery. The old -enrollment token cannot replace a registered node's credential. - -The response contains `installation_id`, `provider`, `core_url` (the installation -public URL), `generation`, `specification` and `specification_digest`. It contains no administrator, Project -or E2B credential. It is available only for node-backed providers. - -The installer reads this configuration before preparing local assets. Its provider -file records `generation` and `specification`, alongside the host's socket, paths -and network policy. Enrollment at `POST /api/v1/sandbox-node/enroll` includes -`deployment_generation`, `specification_digest` and `core_url`, the Core origin the -node stores. A specification mismatch rejects with 409 -`sandbox_specification_mismatch` and a `core_url` other than the installation public -URL with 409 `sandbox_node_address_mismatch`, both before token consumption. A -missing `core_url` gets 400 `invalid_request_error` with `param: "core_url"`. -The original enrollment identity remains immutable. New generation preparation -uses separate exact configurations, never rewrites that identity or silently -substitutes a local default. See [node generation protocol](node-generation-protocol.md). +One database snapshot partitions every unreleased allocation and every pending hosted Environment without an allocation into `cleanup` first, then `busy` or `idle`, so `busy + idle + cleanup == resources.allocations + resources.pending`. Deleted or expired Sessions with unreleased receipts count as cleanup. `offline_nodes` is the complete, ID-sorted list of offline nodes that hold resources by allocation or active placement, and its sum equals `on_offline_nodes`; presence uses the current owner epoch, connection and a heartbeat within 45 seconds, not provider readiness. Direct E2B resources have no node. Offline resources stay blockers until cleanup confirms their release. + +When both held counts reach zero, the owner drains and atomically clears the provider, mode, specification, provider configuration, credential and metadata and provider policy, retires nodes and unused enrollment tokens, increments the generation and owner epoch, and records `reset_complete`. The installation identity and history remain. Core immediately publishes the unconfigured state and keeps the new generation even with no provider, so a delayed load cannot revive the old one. Configure again with POST and the returned generation; no restart is needed. + +During a reset, fresh hosted admission returns 503 `sandbox_reset_in_progress` and leaves no provisional Session rows; live input, receipt retries, restoration and cleanup continue. Management writes and new enrollment return 409 `sandbox_reset_in_progress`, while registered nodes can still read their configuration to recover for cleanup. Per-Session [archive](admin-api.md#administrative-session-archive) needs only the current generation and works with or without a reset. It keeps history and persisted Files and Artifacts, discards the unpersisted workspace and prevents the Session from resuming; poll the archive GET for the actual release. A reset never fabricates a release receipt. + +A force archive fences credentials and new work at once. When the hosted delivery of the Turn is still connected, Core keeps only that delivery's native cancellation and terminal receipt path open until the terminal commit, for at most 20 seconds from the original cancellation request; `done` does not end the bound while a cancellation acknowledgement or terminal commit is pending. This drain never authorizes reconnection, workspace or MCP access or further execution, and an explicit credential revocation ends it. Missing or failed receipts keep honest failure outcomes, and a disconnected, expired or restarted owner falls back to ordinary provider cleanup. + +One mutation gate serializes setup, PUT, reset, cancellation and finalization. Archive locks the Session before the deployment, and finalization never reverses that order. Candidates bind to the reset request time, so a cancel followed by a new reset at the same generation cannot reuse old work. Never replay a rejected or uncertain write: read the current state and decide again. + +## Nodes and allocations + +Node capacity is approved by the administrator, separately from the deployment specification. `POST /core/v1/sandbox/enrollment-tokens` accepts optional `max_active` and `max_retained`, default 2 and 8, and returns `{token, expires_at, enrollment_id}`; `enrollment_id` is a public handle of that command, never a credential. microsandbox uses both limits. Docker never suspends, so Core replaces its `max_retained` with `max_active`, here and in PATCH. E2B has no nodes and answers 409 `sandbox_deployment_conflict`. Core stores the approval with the token and copies it to the node it registers; a node cannot submit capacity, and its local checks may refuse a deployment it cannot run but never raise the limits. [Node capacity](../../docs/configuration.md#node-capacity) explains the limits for operators. + +`GET /core/v1/sandbox/nodes` returns `{data: [...]}` with, for each node: `id`, `name`, `provider`, `online`, `last_seen_at`, `created_at`, `max_active`, `max_retained`, the counts `active`, `reserved`, `running`, `retained`, `snapshots` and `cleanup_pending`, `provider_ready`, `diagnostic`, the host measurements `cpu_count`, `available_memory_bytes` and `available_disk_bytes`, `rollout`, `enrollment_id` and `core_url`. `enrollment_id` is the handle of the command that registered the node, or null when Core has none. `core_url` is the installation public URL at enrollment; a node whose `core_url` differs from the current public URL receives no new placements. Work already placed on it finishes there, including a placed Environment that has no allocation yet, and its retained sandboxes can still resume while the old address reaches Core. Remove it and add it again. + +A node without generation management reports its provider's readiness itself: `provider_ready`, and when it is unready one fixed `diagnostic` code. A node added with Web's command manages generations, so its readiness follows its serving generation and its fixed code for the target generation appears in `rollout.diagnostic`. The node classifies the first failed readiness check and sends only the code; Core stores any other value as `provider_unavailable` and never stores or returns probe text or host paths. The codes are `docker_unavailable`, `docker_limits_unsupported`, `runtime_download_failed`, `runtime_image_unavailable`, `kvm_unavailable`, `microsandbox_artifacts_unavailable`, `capacity_insufficient` and `provider_unavailable`. `runtime_download_failed` means the exact Runtime artifacts could not be transferred or verified; it never contains artifact URLs, credentials or transport output. [Readiness codes](../../docs/getting-started/nodes.md#readiness-codes) gives causes and operator actions. Core and nodes must come from the same distribution. + +`PATCH /core/v1/sandbox/nodes/{node_id}` takes `{name, max_active, max_retained}`; lowering a limit stops no running sandbox. `DELETE /core/v1/sandbox/nodes/{node_id}` refuses with 409 `runtime_node_in_use` while the node holds allocations, snapshots, reservations or pending cleanup, including while it is offline. Removal deletes no compute and retires the node's identity; the host can come back only as a new node. There is no node drain. + +`GET /core/v1/sandbox/nodes/{node_id}/allocations` lists the node's unreleased allocations. Each item's `compute_phase_changed_at` is when the allocation entered its current `compute_phase`, or null when unknown. For a suspended microsandbox allocation, that time plus `suspension.retention_seconds` tells roughly when Core reclaims it. ## What each field means per sandbox provider -Some fields keep one name across providers but differ in meaning, or do not apply. -Deployment fields come from `GET /core/v1/sandbox/deployment`; node and allocation -fields from the administrator node routes; runtime fields from the -[Runtime observation API](runtime-observability-api.md), with `disk` only in the -`GET /core/v1/sandbox/runtime-observations`, and the -[Runtime history API](runtime-history-api.md). +Some fields keep one name across providers but differ in meaning, or do not apply. Deployment fields come from `GET /core/v1/sandbox/deployment`, node and allocation fields from the node routes, and Runtime fields from the [Runtime telemetry API](runtime-observability-api.md), where `disk` appears only in the observation list and history is described under [Session Runtime history](runtime-observability-api.md#session-runtime-history). | Field | E2B | Docker | microsandbox | | --- | --- | --- | --- | | Deployment `specification.resources` | `cpus` and `memory_mib`, equal to the ready template build's and taken from it when omitted; no disk fields | `cpus` and `memory_mib`; no disk quota | `cpus`, `memory_mib`, `root_disk_mib` and `environment_disk_mib` | -| Deployment `specification.runtime` | Absent; the build is selected by `configuration.template` | The full [release](#runtime-release); nodes match `image_id` or `image_manifest_digest` | The full [release](#runtime-release); nodes match `microsandbox_ref`, `runtime_sha256` and `firmware_sha256` | -| Deployment `metadata.template_build` | The build as Core read it when the selection was saved | Absent (`metadata` is empty) | Absent (`metadata` is empty) | +| Deployment `specification.runtime` | Absent; `configuration.template` selects the build | The full [release](#runtime-release); nodes match `image_id` or `image_manifest_digest` | The full [release](#runtime-release); nodes match `microsandbox_ref`, `runtime_sha256` and `firmware_sha256` | +| Deployment `metadata.template_build` | The build as Core read it when the selection was saved | Absent: `metadata` is empty | Absent: `metadata` is empty | | Deployment `suspension` | `null`; Core does not suspend E2B sandboxes | `null` | `{idle_seconds, retention_seconds}` | | Deployment `resources.allocations`, `resources.pending` | Core's unreleased E2B sandboxes, and hosted Environments waiting for one | Totals across all nodes | Totals across all nodes | | Enrollment-token `max_active`, `max_retained` | 409 `sandbox_deployment_conflict`, after the 400 capacity checks; E2B has no nodes | `max_retained` always equals `max_active` | Both limits apply | | Node list and detail | Empty list; detail returns 404 | Enrolled nodes | Enrolled nodes | | Node `retained`, `snapshots`, `max_retained` | Not applicable | Docker never suspends: `retained` equals `active`, `snapshots` is 0 and `max_retained` equals `max_active` | Suspended sandboxes are `retained` minus `active` | -| Node `diagnostic` | Not applicable | `docker_unavailable`, `docker_limits_unsupported`, `runtime_image_unavailable`, `capacity_insufficient` or `provider_unavailable` | `kvm_unavailable`, `microsandbox_artifacts_unavailable`, `capacity_insufficient` or `provider_unavailable` | +| Node `diagnostic` codes | Not applicable | `docker_unavailable`, `docker_limits_unsupported`, `runtime_download_failed`, `runtime_image_unavailable`, `capacity_insufficient` or `provider_unavailable` | `kvm_unavailable`, `microsandbox_artifacts_unavailable`, `runtime_download_failed`, `capacity_insufficient` or `provider_unavailable` | | Node `host.available_disk_bytes` | Not applicable | Free space on the filesystem of the node state directory, not a container's disk | Free space on the filesystem of the node state directory; sandbox disks have their own quotas | | Allocation `compute_phase`, `compute_phase_changed_at` | Not applicable: no node allocations | Always `disabled`, counted as running until release; the time is the allocation's creation | Includes `suspended`; its time plus `suspension.retention_seconds` tells roughly when Core reclaims the snapshot | | Runtime observation `cpu`, `memory` | From E2B metrics: `cpu.utilization_ratio` and `capacity_cores`, memory usage and limit; no cumulative CPU time | From Docker stats: `cpu.usage_seconds_total`, CPU and memory limits, memory usage | From the VM: `cpu.usage_seconds_total`, CPU and memory limits, memory usage | -| Runtime observation `disk` | E2B `diskUsed` and `diskTotal`; `null` when the template does not report them | `null`: no disk quota | `null` for now | +| Runtime observation `disk` | E2B `diskUsed` and `diskTotal`; `null` when the template does not report them | `null`: no disk quota | `null` | | Runtime observation `lifecycle_state: sleeping` | Never | Never | While suspended | | Runtime history CPU | Mean of the utilization ratios E2B reported in each bucket | Derived from cumulative CPU time | Derived from cumulative CPU time | -## Failure and version boundaries - -Malformed selections return 400; validated configuration diagnostics use -`invalid_sandbox_configuration`. Stale generation returns 409 `sandbox_generation_stale` with safe `current_generation`; -A different backend returns 409 `sandbox_reset_required` with provider names. -Unconfigured mutations return 409 `sandbox_not_configured`. Other incompatible -deployments return 409 `sandbox_deployment_conflict`. A node -configuration mismatch returns 409 `sandbox_specification_mismatch`; rejected -node credentials return 401 `invalid_node_credential`. Node machine routes check -the credential before any deployment state, so a missing or rejected credential, -including one issued for another installation, gets that 401 even before -initialization or under E2B. Until the deployment is initialized, -`GET /api/v1/sandbox-node/configuration` and `POST /api/v1/sandbox-node/enroll` -answer an otherwise accepted credential with 503 `runtime_node_unavailable`; -`GET /api/v1/sandbox-node/identity` and the node connection answer 401 for any -credential. Unavailable provider preparation returns 503 `execution_unavailable`. -Storage and credential failures remain errors; an empty or failed read is not -evidence of cleanup. - -The administrator node list and node detail report an unready provider with one -fixed `diagnostic` code: `docker_unavailable`, `docker_limits_unsupported`, -`runtime_download_failed`, `runtime_image_unavailable`, `kvm_unavailable`, -`microsandbox_artifacts_unavailable`, `capacity_insufficient` or -`provider_unavailable`. It is absent while the provider is ready. The node -classifies the first failed readiness check and sends only the code; Core stores -any other value as `provider_unavailable` and never stores or returns probe error -text or host paths. Core and nodes must use the same distribution; unknown-code -handling does not establish cross-version compatibility. See -[Readiness codes](../../docs/getting-started/nodes.md#readiness-codes) -for causes, precedence and operator actions. - -A node provider file is an installed copy of the database selection, not a Core startup configuration source. Historical file-managed deployments and selections without a complete specification are unsupported. Preserve their database, identities, provider receipts, Runtime resources and history; install the current release separately. Do not clear state, run a historical service to convert it, or treat removal of an environment variable as a transfer of ownership. See the [installation version policy](../../docs/getting-started/operations.md#installation-version-policy). - -Landed migrations and their refusal conditions remain historical schema evidence. -They do not establish an operator upgrade, downgrade or conversion procedure. -Fresh database initialization uses the ordinary migration runner. - -Unit tests, database tests and provider inspection are separate from live -execution acceptance. This contract does not assert that every resource profile, -provider deployment or host-reboot recovery path has been qualified. - - -### Current Runtime identity - -The immutable microsandbox reference is `oac-runtime@sha256:<64 lowercase hex>`. -Build new E2B templates with the matching current Runtime release. Historical -Runtime names and installations have no supported upgrade, downgrade or conversion -path; preserve their data and resources and install separately. Current Runtime -generation rollout, coexistence and ownership-scoped garbage collection remain -supported and are independent of installed program version changes. See -[generation ownership and rollout](#generation-ownership-and-rollout). - -`runtime_download_failed` means the exact Runtime artifacts could not be transferred -or verified. It is distinct from provider probe and image availability failures; -the diagnostic never contains artifact URLs, credentials or transport output. +## Errors +| HTTP | Code | When | +| --- | --- | --- | +| 400 | `invalid_request_error` or `invalid_request` | A malformed request | +| 400 | `invalid_sandbox_configuration` | A validated configuration diagnostic | +| 409 | `sandbox_generation_stale` | `expected_generation` is not current; `current_generation` gives the current one | +| 409 | `sandbox_reset_required` | A different backend; the details name the current and requested providers | +| 409 | `sandbox_not_configured` | A change to an unconfigured deployment | +| 409 | `sandbox_reset_in_progress` | A management write or new enrollment during a reset | +| 409 | `sandbox_deployment_conflict` | Another state the change cannot apply to | +| 409 | `runtime_node_in_use` | Node removal while it holds resources | +| 503 | `execution_unavailable` | Provider preparation is unavailable | +| 503 | `sandbox_credential_unavailable` | The credential encryption key is unavailable | + +Storage and credential failures stay errors: an empty or failed read never proves cleanup. The [machine connection API](machine-api.md#node-route-errors) lists the errors of the node routes. ## Canonical node specification -`sandbox/deployment_contract.go` owns resource bounds, provider requirements, -release patterns and canonical field order. `sandbox/deployment.go` applies those -rules in Core. The installer consumes the generated declaration in -`deploy/install/node_spec.py`; do not maintain a second set of limits or patterns. -Regenerate it from the repository root with -`go run ./services/core/cmd/specification-contract -write`. -The sandbox Go tests, included in `make check`, reject a stale projection. - -The specification digest is SHA-256 of UTF-8 compact JSON, with `provider` first, -then `resources`, then `runtime` when required by the provider. Resource and -Runtime fields follow the contract declaration order. Zero optional disk fields -are omitted; required fields remain present. Release identities are lowercase -ASCII; the digest never hashes the incoming JSON field order or whitespace. -`internal/sandbox/testdata/deployment-contract.json` (under `services/core/`) -contains shared acceptance cases, exact canonical bytes and digests consumed by -both Go and Python tests. Cross-language validation is required; distinct peers -must not invent distinct rules. +`sandbox/deployment_contract.go` owns the resource bounds, provider requirements, release patterns and canonical field order; `sandbox/deployment.go` applies them in Core. The installer consumes the generated declaration in `deploy/install/node_spec.py`, so there is no second set of limits or patterns. Regenerate it from the repository root with `go run ./services/core/cmd/specification-contract -write`; the sandbox Go tests, part of `make check`, reject a stale projection. + +The specification digest is the SHA-256 of compact UTF-8 JSON with `provider` first, then `resources`, then `runtime` when the provider requires it. Resource and Runtime fields follow the contract's declaration order; zero optional disk fields are omitted and required fields stay present. Release identities are lowercase ASCII, and the digest never depends on the incoming field order or whitespace. `services/core/internal/sandbox/testdata/deployment-contract.json` holds shared acceptance cases, exact canonical bytes and digests that both the Go and Python tests consume. diff --git a/docs/api/README.md b/docs/api/README.md index 7bcef1d52..06947eba1 100644 --- a/docs/api/README.md +++ b/docs/api/README.md @@ -6,7 +6,7 @@ Core serves three namespaces. Each has one kind of caller and its own credential | --- | --- | --- | --- | --- | | `/v1` | Applications: business systems and the official OpenAI SDK | Project API key | Exactly the 58 method and path pairs of the pinned official Agents API, listed in [upstream-routes.json](../../contracts/agents-api/upstream-routes.json). Core-only fields sit inside `x_agents_core`: `harness`, `model_provider`, `harness_config`, `environment`, and the read-only Session `installation` | [Agents API guide](public-agent-api.md) | | `/core/v1` | Web's console server and operator scripts | [Core key](../getting-started/operations.md#core-key) | Installation facts, Projects and keys, resource reads and deletion, Session archive, executor credentials, default models, metrics, audit, sandbox deployment and nodes | [Core administration API](../../contracts/agents-api/admin-api.md) | -| `/api/v1` | Nodes, Runtime daemons, self-hosted executors and their installers | Machine credentials: node enrollment tokens and node credentials, installation grants, executor credentials, and daemon credentials. Each works only on its own routes | Machine bootstrap and connections under `/api/v1/sandbox-node/*` and `/api/v1/agent-daemon/*`, including WebSockets, and the public native installer downloads | [Machine connection API](#machine-connection-api) | +| `/api/v1` | Nodes, Runtime daemons, self-hosted executors and their installers | Machine credentials: node enrollment tokens and node credentials, installation grants, executor credentials, and daemon credentials. Each works only on its own routes | Machine bootstrap and connections under `/api/v1/sandbox-node/*` and `/api/v1/agent-daemon/*`, including WebSockets, and the public native installer downloads | [Machine connection API](../../contracts/agents-api/machine-api.md) | A credential used in another namespace gets 401: a Project API key on `/core/v1` or `/api/v1`, the Core key on `/v1` or `/api/v1`. How Projects and keys behave is in [Projects own assets](../design-principles.md#projects-own-assets). @@ -14,14 +14,4 @@ A credential used in another namespace gets 401: a Project API key on `/core/v1` ## Machine connection API -These routes are under `/api/v1`. Each accepts only the credential listed, never the Core key or a Project API key. The generated [machine OpenAPI](../../contracts/agents-api/runtime.openapi.yaml) covers the annotated node HTTP and installation grant routes; the node WebSocket `connect` route is described in the node generation protocol. - -| Routes | Caller | Credential | Contract | -| --- | --- | --- | --- | -| `POST sandbox-node/enroll` | Node installer | One-use enrollment token from `POST /core/v1/sandbox/enrollment-tokens` | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes) | -| `GET sandbox-node/configuration` | Node installer and node | Enrollment token, or node credential with `X-OAC-Node-ID` | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes) | -| `GET sandbox-node/identity`, WebSocket `GET sandbox-node/connect` | Node | Node credential registered at enrollment | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes), [node generation protocol](../../contracts/agents-api/node-generation-protocol.md) | -| `GET agent-daemon/install/{version}/*` | Self-hosted installer | None: public, immutable release content | [Installation grant](../../contracts/agents-api/environment-executor-credentials.md#installation-grant) | -| `POST agent-daemon/installation`, `POST agent-daemon/installation/claim` | Self-hosted installer | Installation grant: the short-lived authorization in a `self_hosted` Session's `x_agents_core.installation` commands | [Installation grant](../../contracts/agents-api/environment-executor-credentials.md#installation-grant) | -| `POST agent-daemon/enroll`, `GET agent-daemon/connection` | Self-hosted executor and its installer | Executor credential | [Executor credentials](../../contracts/agents-api/environment-executor-credentials.md) | -| WebSocket `GET agent-daemon/ws`, `POST agent-daemon/bootstrap`, `GET agent-daemon/device-status` | Runtime daemons | Daemon credential: Core issues one to each hosted sandbox through the [bootstrap file](../runtime-bootstrap.md); a self-hosted executor uses its executor credential; a device for `none` Sessions uses the credential an operator provisions with `oac-core-device` | [Core–Runtime protocol](../runtime-protocol.md), [device provisioning](../../services/core/README.md#internal-execution-device-connection) | +Nodes, Runtime daemons and the self-hosted installer call `/api/v1` with their own credentials. The [machine connection API](../../contracts/agents-api/machine-api.md) lists every route, caller and credential. diff --git a/docs/sandbox-provider.md b/docs/sandbox-provider.md index ba34ed313..c83d7a1c7 100644 --- a/docs/sandbox-provider.md +++ b/docs/sandbox-provider.md @@ -1,11 +1,6 @@ # Add a Sandbox Provider -A **Sandbox Provider** supplies the outer compute (an *Environment*) that a Core -Runtime daemon runs in, plus the bounded bootstrap that starts it. This guide is -the single start-to-finish path for adding one. The canonical interface is -[`SandboxProvider`](../services/core/internal/sandbox/sandbox_provider.go). - -## Before you start +A **Sandbox Provider** supplies the outer compute that a Runtime daemon runs in for a Core-managed Environment, and the bounded bootstrap that starts that daemon. This guide is the path for adding one and the reference for how Core drives it. The interface is [`SandboxProvider`](../services/core/internal/sandbox/sandbox_provider.go). | Term | Meaning | | --- | --- | @@ -14,334 +9,195 @@ the single start-to-finish path for adding one. The canonical interface is | Runtime | The daemon inside the Environment; it prepares capabilities and executes Turns | | Deployment | The single deployment-wide provider selection; see [Sandbox deployment](../contracts/agents-api/sandbox-deployment.md) | -Core owns durable Environment, allocation, placement and cleanup state. The -Provider supplies compute and bootstrap only. Runtime prepares capabilities and -workspaces through the common [daemon protocol](runtime-protocol.md); Harness -adapters translate execution. A user-owned machine uses the same Runtime -contract but has no Core-owned allocation to create or destroy. - -Use maintained provider SDKs behind thin adapters. Hosted deployments select one -deployment-wide Provider: E2B cloud, or Docker/microsandbox on -administrator-owned nodes. Providers never execute Core initialization commands; -initialization, daily execution and Files use the daemon's typed Runtime operations. - -Isolation belongs to the outer infrastructure. The daemon runs with its -launching user's authority and adds no filesystem, tool or network sandbox. -Qualify provider isolation and network enforcement independently of daemon -connectivity; see [Runtime and outer isolation](design-principles.md#runtime-and-outer-isolation). +Core owns durable Environment, allocation, placement and cleanup state; the Provider owns compute and bootstrap only. The Runtime prepares capabilities and runs Turns over the [Core–Runtime protocol](runtime-protocol.md), and the provider hands it its identity through the [Runtime bootstrap](runtime-bootstrap.md) file. A provider never runs Environment initialization, Skills, Plugins, MCP setup, initial files, execution or Files; those use the Runtime. Isolation belongs to the provider's infrastructure, not the daemon; see [Runtime and outer isolation](design-principles.md#runtime-and-outer-isolation). Use the vendor's maintained SDK behind a thin adapter. ## Steps -1. **Read the contract.** Implement the five required operations and explicitly handle all - extension interfaces in [Implement the interface](#implement-the-interface). -2. **Write the adapter package** under `services/core/internal/sandbox/` - (native SDK calls, ownership checks, identity translation, private config). - Assert `var _ sandbox.SandboxProvider = (*YourAdapter)(nil)` at compile time. - Out-of-process helpers live in `services/core/tools/-provider`. -3. **Register the kind** once using - [Register the provider kind](#register-the-provider-kind). Registration is - explicit construction, not an init-time plugin registry. -4. **Label owned resources** with `io.oac.*` labels, or `oac_*` metadata keys - where the vendor rejects dotted keys. See [Label naming](#label-naming). -5. **Run the contract suite** with `make check-sandbox-provider-contract`. See - [Validate the integration](#validate-the-integration). -6. **Run native acceptance** against real compute, then `make check`. Record - evidence and limits in [Sandbox deployment](../contracts/agents-api/sandbox-deployment.md). -7. **Document operator setup** next to the adapter, like the - [reference adapters](#reference-adapters). +1. **Read the contract.** Implement the five required operations and give an explicit decision for every extension interface in [Implement the interface](#implement-the-interface). +2. **Write the adapter package** under `services/core/internal/sandbox/`: native SDK calls, ownership checks, identity translation and private configuration. Assert `var _ sandbox.SandboxProvider = (*YourAdapter)(nil)`. An out-of-process helper lives in `services/core/tools/-provider`. +3. **Register the kind** once, following [Register the provider kind](#register-the-provider-kind). Registration is explicit construction, not an init-time plugin registry. +4. **Label owned resources** with the provider ownership labels in the [Runtime names](../CONTRIBUTING.md#openagentcore-runtime-names) table, and never accept older label names as a fallback. +5. **Run the contract suite** with `make check-sandbox-provider-contract`; see [Validate the integration](#validate-the-integration). +6. **Run native acceptance** against real compute, then `make check`. +7. **Document the adapter** next to it, like the [reference adapters](#reference-adapters). ## Implement the interface -`SandboxProvider` has **five required operations**: +`SandboxProvider` has five required operations: | Operation | Purpose | | --- | --- | | `Create` | Create compute for a `Reference` and run the bounded daemon bootstrap | -| `GetInfo` | Observe current compute without mutation | -| `Renew` | Extend a native lease, or observe when the backend has none | +| `GetInfo` | Observe current compute without changing it | +| `Renew` | Extend a native lease, or only observe when the backend has none | | `Kill` | Reclaim the allocation's compute and retained resources | -| `RunCommand` | Bounded administrative bootstrap/diagnostic command | +| `RunCommand` | Run a bounded command in the allocation's compute | -Use the existing types; do not introduce another lifecycle protocol or a -vendor-specific execution path. A backend without a native renewable lease -(Docker) still keeps service-owned hosted expiry and cleanup requirements. +A backend without a native renewable lease, such as Docker, still keeps Core's hosted expiry and cleanup requirements. Every adapter implements `RunCommand` and its tests exercise it, but Core's orchestration does not call it; only the node transport forwards it. Confidential command input travels in `Command.Stdin`, never in arguments or logs, and the result keeps byte order, bounded output and the actual exit status. ### Explicit operation contracts -Keep the existing small interfaces. Every Provider must implement their methods -and return a complete `ProviderOperations()` declaration. The interface methods -are the operation inventory; `sandbox.ValidateOperations` checks it without a -second hand-maintained list. +Every provider implements the methods of each interface below and returns a complete `ProviderOperations()` declaration. The interface methods are the operation inventory; `sandbox.ValidateOperations` checks the declaration against it without a second hand-kept list. | Contract | Requirement | Responsibility | | --- | --- | --- | -| `sandbox.SandboxProvider` | Five operations must be supported | Allocation lifecycle and bounded administrative commands | -| `sandbox.CheckpointProvider` | Explicit supported or unsupported decision for every method | Exact compute incarnations, capture/restore, retained-source resume and cleanup | +| `sandbox.SandboxProvider` | All five operations supported | Allocation lifecycle and bounded commands | +| `sandbox.CheckpointProvider` | Explicit decision for every method, the same for all of them | Exact compute incarnations, capture and restore, retained-source resume and cleanup | | `runtimeobs.Source` | Explicit decision | Ownership-checked read-only observations | -| `runtimeobs.BatchSource` | Explicit decision | Bounded observations in input order, with per-target errors | +| `runtimeobs.BatchSource` | Explicit decision; requires `Source` | Bounded observations in input order, with per-target errors | | `sandbox.SelectionDiscoverer` | Explicit decision | Read-only native configuration discovery before commit | | `sandbox.CredentialVerifier` | Explicit decision | Verify access to owned resources without mutation | -Each declaration entry has `state: supported` with no reason, or -`state: unsupported` with an authored reason code. Missing entries, zero states, -unknown entries, missing methods and unsafe reasons fail validation. The entire -checkpoint lifecycle must agree on support; batch observation requires single -observation. Adding a method to an existing interface requires an explicit -decision and implementation in every adapter. Never supply a default base class -or generate blanket unsupported implementations for future methods. - -Unsupported methods return `providercontract.UnsupportedError` before native -I/O. The error identifies the exact operation and a safe code, not a native -message, resource identity, endpoint or credential. Empty results, nil errors, -`Unavailable`, and unknown mutation outcomes cannot substitute for unsupported. -The five required methods cannot return unsupported; a backend without a native -lease preserves the existing read-only `Renew` semantics. - -Each adapter owns one `Operations()` function, shared by its concrete instance -and registration. `providers.ValidateBinding` checks both against the existing -interfaces and against each other. Runtime admission and node generation loading -also reject incomplete providers. Interface assertions establish method shape -only; callers use the declaration to decide whether an operation is supported. - -`ObserveBatch` returns an error instead of an ambiguous boolean. Only a typed, -safe `UnsupportedError` for `ObserveBatch` permits per-target `Observe` calls. -An unavailable service, timeout or other failure never triggers that fallback. -Observation never renews, starts, prepares or stops compute; see the -[observation contract](../contracts/agents-api/runtime-observability.md). - -Contract tests call every declared unsupported native method with no configured -native client, require its matching error and zero result, and reject incomplete -or contradictory declarations. Supported behavior still requires native and -lifecycle tests; declaration validation alone cannot prove SDK semantics. - -`RunCommand` is an existing administrative bootstrap/diagnostic facility, not an -alternate route for Skills, Plugins, MCP setup, initial files, Session execution or -Files. Those use Runtime. Confidential command input travels in `Command.Stdin`, -not argv or logs. Preserve its byte order, bounded output and actual exit status. +`CheckpointProvider` adds `Initial` and `NewCompute` (construct compute references without allocating), `GetCompute`, `Suspend`, `Resume`, `ResumeCompute` (thaw only the same resident instance after an aborted pause), `KillCompute`, `DeleteSnapshot` and `RunCommandCompute`, which runs a bounded command in one exact compute incarnation. Core uses `RunCommandCompute` to wake a parked daemon after a restore ([`runtime_compute_wake.go`](../services/core/internal/execution/runtime_compute_wake.go)). + +Each declaration entry is `state: supported` with no reason, or `state: unsupported` with an authored reason code. Missing, zero, unknown or unsafe entries and missing methods fail validation. Adding a method to an interface requires an explicit decision and implementation in every adapter; never supply a base type or generate blanket unsupported implementations. + +An unsupported method returns `providercontract.UnsupportedError` before any native I/O. The error names the exact operation and a safe code, never a native message, resource identity, endpoint or credential. An empty result, a nil error, `Unavailable` or an unknown mutation outcome never stands in for unsupported, and the five required methods can never return it. + +Each adapter owns one `Operations()` function, shared by its instance and its registration. `providers.ValidateBinding` checks both against the interfaces and each other, and Runtime admission and node generation loading also reject incomplete providers. An interface assertion establishes only method shape; callers use the declaration to decide support. + +`ObserveBatch` returns an error, not an ambiguous boolean. Only a typed `UnsupportedError` for `ObserveBatch` permits per-target `Observe` calls; an unavailable service, timeout or other failure never does. Observation never renews, starts, prepares or stops compute; see the [observation contract](../contracts/agents-api/runtime-observability.md). + +Contract tests call every method declared unsupported with no native client configured, require its matching error and zero result, and reject incomplete or contradictory declarations. Supported behavior still needs native and lifecycle tests. ### Identity and resource ownership -`Reference` is the exact `(TenantID, EnvironmentID, AllocationID)` tuple. Core -persists a fresh allocation ID **before** `Create`; it is not the Environment ID. -The adapter also binds resources to its installation. Do not locate or authorize -resources by a bare native ID, display name, guessed path or an unverified label. -All mutations, reads and cleanup must verify the same ownership. - -Core serializes lifecycle operations and retains the allocation after any -uncertain mutation. An adapter must preserve enough native identity/receipts for -observation and cleanup. A failed call may return both `Info` and an error: -preserve reference-bound settlement evidence without converting failure to success. -Adapters must not silently create replacement resources, overwrite credentials or -switch to a new allocation after a conflict. - -`Info.ProviderID` and `Info.State` describe compute. Core recognizes `running` -only with the matching Reference and a native identity. Other native states can -be observed without claiming readiness. `BootstrapComplete` says the bootstrap -reached its final mutating step. `CreateSettled` proves the original attempt can -no longer mutate resources; it does not mean success. In particular: - -- `State="absent"` plus matching Reference and `CreateSettled=true` is an explicit - creation-absence receipt, with no native ID or completed bootstrap. -- `ErrNotFound`, an empty listing, a timeout or a successful `Kill` alone is not - proof that an in-flight create cannot appear later. -- Core can also settle creation from a matching running resource with completed - bootstrap. It must retain unknown creation until it has such evidence or an - explicit receipt, even if a cleanup attempt currently sees no resources. - -`Kill` owns cleanup of the allocation's compute and retained resources, including -partial bootstrap storage. It must not remove another tenant's resource on a -name collision. Core releases durable ownership only after confirmed cleanup -**and** settled creation. Executor release or Harness cancellation does not delete -an Environment, its workspace or its Sandbox Provider allocation. +`Reference` is the exact `(TenantID, EnvironmentID, AllocationID)` tuple. Core persists a fresh allocation ID before `Create`; it is not the Environment ID. The adapter also binds resources to its installation, and never locates or authorizes a resource by a bare native ID, display name, guessed path or unverified label. Every mutation, read and cleanup verifies the same ownership. + +Core serializes lifecycle operations and keeps the allocation after any uncertain mutation, so the adapter must keep enough native identity and receipts for observation and cleanup. A failed call may return both `Info` and an error: keep reference-bound settlement evidence without turning the failure into success. Never silently create a replacement resource, overwrite credentials or switch to a new allocation after a conflict. + +`Info.ProviderID` and `Info.State` describe compute. Core treats compute as `running` only with the matching `Reference` and a native identity; other native states can be observed without claiming readiness. `BootstrapComplete` says the bootstrap reached its final mutating step. `CreateSettled` proves that the original attempt can no longer change resources; it does not mean success: + +- `State="absent"` with the matching `Reference` and `CreateSettled=true` is an explicit creation-absence receipt, with no native ID or completed bootstrap. +- `ErrNotFound`, an empty listing, a timeout or a successful `Kill` alone never proves that an in-flight Create cannot appear later. +- Core can also settle creation from a matching running resource with a completed bootstrap. Until it has such evidence or an explicit receipt it keeps the creation unknown, even when a cleanup attempt sees no resources. + +`Kill` owns cleanup of the allocation's compute and retained resources, including partial bootstrap storage, and never removes another tenant's resource on a name collision. Core releases durable ownership only after confirmed cleanup and settled creation. Closing an Executor or cancelling a Harness never deletes an Environment, its workspace or its allocation. ### Operation outcomes and retries -All calls receive bounded contexts. Expiry or cancellation ends the caller's -wait; it does not prove rollback, native stop, cleanup or absence. An adapter or -transport must not detach untracked mutations or replay a timed-out command. +Every call receives a bounded context. Expiry or cancellation ends the caller's wait; it proves no rollback, stop, cleanup or absence. An adapter or transport never detaches untracked mutations or replays a timed-out command. | Operation | Confirmed result | Failure or unknown result | Recovery | | --- | --- | --- | --- | -| `Create` | Matching compute and bootstrap evidence; execution still needs Runtime preparation | Invalid/foreign configuration rejects; duplicate returns `ErrExists`; transport failure may hide created resources | Observe the original Reference. Never replay `Create`, even with a new credential. Preserve partial resources for owned cleanup. | -| `GetInfo` | Current compute observation without mutation | `ErrNotFound` is only a missing observation; an error is not absence proof | Repeat a bounded read; never turn it into create/start/renew. | -| `Renew` | Existing native lease extended, or an observation for a provider without leases | Timeout may hide a lease extension; stopped/missing compute stays stopped/missing | Observe then let the normal reconciler renew the same allocation. Never revive it or fabricate a lease expiry. | -| `Kill` | Owned compute and retained storage removed; repeated confirmed absence succeeds | Error retains ownership and cleanup intent; ownership mismatch must not delete foreign resources | Retry cleanup of the same Reference after outstanding creation/mutation is fenced. Do not release the durable owner early. | -| `RunCommand` | Collected output and actual exit code; nonzero exit is a settled command failure | Missing native completion is `ErrCommandUnconfirmed`; partial output is not success | Do not replay. Preserve the owner and reclaim before reuse when completion cannot be proved. | - -Use `ErrInvalid`, `ErrOwnership`, `ErrExists`, `ErrNotFound`, -`ErrComputeUnconfirmed` and `ErrCommandUnconfirmed` for their existing meanings. -An unclassified native/transport error is conservatively unknown, not permission -to retry a mutation. Core must not interpret provider diagnostics as lifecycle -truth or expose native error text/credentials. The node boundary maps errors to -fixed codes; direct SDK details stay private. - -Checkpoint support adds `Compute` generation/name/ID and `SnapshotIdentity`. -Persist operation IDs and provider-returned snapshot provenance unchanged. -`ObserveOnly` on suspend/resume observes the previous attempt and must not start -another capture or restore. `ResumeCompute` only thaws the retained source; it -must not cold-start a stopped one. Cleanup targets the exact compute incarnation -and snapshot, not whichever instance currently has the same display name. See -[the lifecycle implementation](../services/core/internal/execution/runtime_compute.go) -and its failure tests before advertising this capability. +| `Create` | Matching compute and bootstrap evidence; execution still needs Runtime preparation | Invalid or foreign configuration rejects; a duplicate returns `ErrExists`; a transport failure may hide created resources | Observe the original `Reference`. Never replay `Create`, even with a new credential. Keep partial resources for owned cleanup | +| `GetInfo` | Current compute observation without change | `ErrNotFound` is only a missing observation; an error is not proof of absence | Repeat a bounded read; never turn it into create, start or renew | +| `Renew` | The native lease extended, or an observation for a provider without leases | A timeout may hide an extension; stopped or missing compute stays so | Observe, then let the reconciler renew the same allocation. Never revive compute or fabricate a lease expiry | +| `Kill` | Owned compute and retained storage removed; repeated confirmed absence succeeds | An error keeps ownership and cleanup intent; an ownership mismatch never deletes foreign resources | Retry cleanup of the same `Reference` after outstanding creation or mutation is fenced; never release the owner early | +| `RunCommand`, `RunCommandCompute` | Collected output and actual exit code; a nonzero exit is a settled command failure | Missing native completion is `ErrCommandUnconfirmed`; partial output is not success | Never replay. Keep the owner and reclaim before reuse when completion cannot be proved | + +`ErrInvalid`, `ErrOwnership`, `ErrExists`, `ErrNotFound`, `ErrComputeUnconfirmed` and `ErrCommandUnconfirmed` keep their defined meanings. An unclassified native or transport error is unknown, never permission to retry a mutation. Core never reads provider diagnostics as lifecycle truth or exposes native error text or credentials; the node transport maps errors to fixed codes, and direct SDK details stay private. + +Checkpoint support adds `Compute` generation, name and ID and `SnapshotIdentity`; persist operation IDs and the provider's snapshot provenance unchanged. `ObserveOnly` on suspend or resume observes the previous attempt and never starts another capture or restore. `ResumeCompute` only thaws the retained source and never cold-starts a stopped one. Cleanup targets the exact compute incarnation and snapshot, not whatever instance now has the same name. Read [`runtime_compute.go`](../services/core/internal/execution/runtime_compute.go) and its failure tests before declaring checkpoint support. ### Four distinct readiness facts | Fact | Evidence | Does not establish | | --- | --- | --- | -| Compute available | Provider observation for the owned allocation | Authenticated Runtime connection or prepared capabilities | -| Runtime connected | Gateway authentication and exact Environment/device binding | Completed Runtime preparation or a usable Harness | -| Capabilities prepared | Successful common Runtime preparation with the fixed configuration/snapshot | Acceptance or completion of a Turn | -| Execution admitted | Qualified Harness capabilities and the existing executor/Turn acceptance path | A completed input, cancellation or reclaimed compute | +| Compute available | Provider observation for the owned allocation | An authenticated Runtime connection or prepared capabilities | +| Runtime connected | Gateway authentication and the exact Environment and device binding | Completed preparation or a usable Harness | +| Capabilities prepared | Successful common Runtime preparation with the fixed configuration | Acceptance or completion of a Turn | +| Execution admitted | Qualified Harness capabilities and the executor and Turn acceptance path | A completed input, cancellation or reclaimed compute | -For both hosted and self-hosted environments, Core uses the daemon preparation -and execution protocol. Core resolves configuration and supplies resources; -Runtime installs/reads local capabilities and freezes the result for the Session. -Reconnection can reuse that prepared content; a new Session takes a new snapshot. -A provider must not implement a competing preparation path. +Hosted and self-hosted Environments use the same Runtime preparation; a provider never implements a competing one. ## Register the provider kind -`sandbox/providers/registry.go` is the sole registration table. Each entry binds -an adapter's specification/resource validators, the required `ConfigurationAdapter`, deployment -mode, defaults, the adapter-owned operation declaration, and local or direct constructor. -`providers.Build` constructs node-local adapters; `providers.BuildDirect` constructs -direct adapters. Neither allocates compute. There is no init-time registration or -runtime plugin loading. +`sandbox/providers/registry.go` is the only registration table. Each entry binds the adapter's specification and resource validators, its `sandbox.ConfigurationAdapter`, the deployment mode (`nodes` or `direct`), suspension defaults, the operation declaration and a node-local (`BuildLocal`) or direct (`BuildDirect`) constructor. `providers.Build` and `providers.BuildDirect` construct adapters without allocating compute. There is no init-time registration or plugin loading. + +A new provider takes these steps: + +1. Implement the operation contracts in the adapter package, with native contract tests. +2. Add its specification and resource validators and, for native resource discovery before commit, an optional read-only `SelectionDiscoverer`. Put native credential verification behind `CredentialVerifier`. +3. Implement `sandbox.ConfigurationAdapter` over a typed native configuration. `DecodeInput` strictly parses the separate public `configuration` and write-only `credential` objects of a request. `Encode` produces whitelisted public selectors, read-only observations and separate secret bytes, and never passes request JSON through. `Decode` restores stored selectors, and keeps access to owned resources, without remote admission or new template validation. `Normalize` copies its input before changing it. `ResolveChange`, `Equal` and `WithCredential` own inheritance, identity and credential composition. `Requirements` declares whether a credential and a public Core origin are required and whether configuration discovery is supported. Also implement `ConfigurationDiscoverer`, even when discovery is unsupported: it validates the query and returns a safe catalog, never a mutation or an admission decision, while Core keeps authorization, input limits and deadlines. A node provider accepts only an empty public object, rejects credentials and returns Unsupported for discovery and credential replacement. +4. Register its constructor, policies, configuration adapter, operation declaration and defaults in `providers/registry.go`. Node proxy identity and checkpoint support read this entry. The installer's projection combines the registered policies with the shared field bounds in `sandbox/deployment_contract.go`; regenerate it with `go run ./services/core/cmd/specification-contract -write`. +5. Supply the distribution's artifacts for the adapter and its helper. **Known gap:** offering the provider to operators still edits shared installer and Web code with vendor-specific options, such as E2B's installer flags and Web views; this is not the pattern to follow. Never add a Session or Turn scheduling path, a vendor column or API field, or a vendor switch in the store. ### Registration validation -`providers.ValidateRegistration` is the single wiring check. Lookup, constructor binding and the installer projection use it before configuration callbacks or constructors can run. Unknown provider names remain invalid input; malformed registered adapters return a safe `providercontract.ErrContract`, without including submitted configuration or native diagnostics. - -- A `nodes` registration requires only `BuildLocal`; a `direct` registration requires only `BuildDirect`. Missing, mixed or unknown modes are rejected. -- Specification/resource validators, the configuration adapter and the complete operation declaration are mandatory. An incomplete registration cannot publish a partial installer projection. -- The Runtime input policy must either accept the pinned Runtime or give the adapter's fixed reason for rejecting it; it cannot do both. -- Current checkpoint suspension requires node mode and positive idle/retention defaults that fit Runtime durations. The two durations are independent. Providers without checkpoint support must not configure suspension defaults. - -The configuration adapter must be non-nil, including its concrete value. Each `ConfigurationRequirements` field needs an explicit valid decision: `Credential` and `PublicOrigin` use `Required` or `NotRequired`; `Discovery` uses the shared supported/Unsupported declaration with its safe reason. `ConfigurationDiscoverer` must be implemented even when discovery is unsupported. New requirement fields or discovery methods require an explicit validation update; they cannot inherit an existing decision. Configuration discovery is distinct from resource selection discovery. Requiring a credential does not itself promise the resource operation `VerifyCredential`. - -These checks establish registration completeness, not the correctness of native SDK behavior. Constructor and adapter contract tests still apply. - -For a new implementation: - -1. Implement the operation contracts above in the adapter package and add native contract tests. -2. Add its configuration validators and optional read-only `SelectionDiscoverer` - for native resource discovery. Normalization must copy input before changing it. - `ConfigurationAdapter.Decode` must retain access to owned resources without requiring new - template validation. Put native credential verification behind - `CredentialVerifier` when needed. -3. Register its constructor, policies, operation declaration and defaults in `providers/registry.go`. - Node proxy identity and checkpoint advertisement consume this same entry. - The installer projection uses those registered policies and the common field - bounds in `sandbox/deployment_contract.go`; regenerate it with - `go run ./services/core/cmd/specification-contract -write`. -4. Implement `sandbox.ConfigurationAdapter` with a typed native configuration. - `DecodeInput` strictly parses separate `configuration` and write-only `credential` - objects. `Encode` creates whitelisted public selectors, read-only observations - and separate secret bytes; it must never pass request JSON through. `Decode` - restores stored selectors without remote admission. `ResolveChange`, `Equal` - and `WithCredential` own inheritance, identity and credential composition. - `Requirements` explicitly declares credential and public-origin needs plus - discovery support. Node providers accept only an empty public object, reject - credentials and explicitly return Unsupported for discovery and replacement. - Implement the separate `ConfigurationDiscoverer` interface even when unsupported. - Discovery owns query validation and safe catalog results, never mutations or - deployment admission. Core keeps admin authorization, input limits and deadlines. - A new kind requires no vendor column, API field or Store branch. -5. Supply required installer/distribution artifacts and operator labels. A new - provider must not add a Session/Turn scheduling path or a Store vendor switch. - -Preview and persistence use `providers.Normalize` and `providers.Describe`. -`SelectionDiscoverer` resolves omitted native resource values before commit; the -complete specification is validated again at persistence. Store owns transactions, -credential encryption, generation fencing, resource ownership and generic object storage. The adapter alone interprets -`provider_config` and `provider_metadata`; `provider_credential` holds ciphertext -bound to the installation and generation. Retained generations store their original -public configuration and metadata, and compose the current credential through the -adapter. No retained selector is rewritten by a credential replacement. Database -constraints validate object structure, not the registration list. -`providers.ResolveChange` owns configuration inheritance and comparison uses -normalized selectors, so preview, retry and commit share the same defaults. - -A direct adapter with credentials verifies all retained generations and allocation -references before replacing a key. The common `sandbox.CallFence` excludes native -calls and waits for helper completion, including calls whose callers timed out. -Execution invokes prepared verification/fencing callbacks without branching on a -vendor. A transport wrapper advertises only capabilities that its adapter supports; -new optional capabilities need forwarding and qualification before registration. - -Keep vendor-specific deployment validation and SDK setup at the construction -boundary. Construction must not create an Environment. For node-local adapters -the `providers.Built` result returns the provider, probe, installation identity, backend -fingerprint and specification digest; the factory also returns its close function. - -Preserve the `execution.RuntimeProvider` deployment binding: `ProviderKind`, -installation ID, backend fingerprint, generation, mode and node ownership -identify the backend. The database owns the selection; the in-memory selection -is not an alternate authority. Register optional interfaces consistently on -direct adapters and their transport wrappers. A new public capability needs its -own protocol change. - -Today Docker and microsandbox use nodes; E2B is constructed directly. The node -proxy exposes checkpoint operations only for its registered checkpoint-capable -backend. A new node backend must register both construction and the corresponding -proxy capability; otherwise that capability is unavailable. Provider names belong -in this adapter/configuration wiring, not Session scheduling, capability -preparation or Turn execution. Common lifecycle code uses `CheckpointProvider` -to admit suspension, independently of a provider name. - -The backend fingerprint identifies a native resource namespace, not mutable -capacity. Core retains deployment generations so old owned allocations continue -to resolve to their original backend. This is resource ownership, not historical -binary compatibility. Do not repoint retained allocations at a replacement backend. - -## Label naming - -Every owned native resource carries ownership labels so mutations, reads and -cleanup can verify the `Reference` and installation. - -- Use the `io.oac.` prefix, for example `io.oac.installation` and `io.oac.tenant`. -- Where the vendor rejects dotted keys, use `oac_` metadata keys (E2B). -- Never accept older label names as a fallback; see - [Runtime names](../CONTRIBUTING.md#openagentcore-runtime-names). +`providers.ValidateRegistration` is the single wiring check. Lookup, constructor binding and the installer projection run it before any configuration callback or constructor. An unknown provider name stays invalid input; a malformed registration returns a safe `providercontract.ErrContract` that includes no submitted configuration or native diagnostics. + +- A `nodes` registration has only `BuildLocal`, and a `direct` registration only `BuildDirect`; missing, mixed or unknown modes are rejected. +- The specification and resource validators, the configuration adapter and a complete operation declaration are mandatory, so an incomplete registration cannot publish a partial installer projection. +- The Runtime input policy either accepts the pinned Runtime or gives the adapter's fixed reason for rejecting it, never both. +- Checkpoint support requires node mode and positive idle and retention defaults that fit Runtime durations; a provider without checkpoint support configures no suspension defaults. + +The configuration adapter must be non-nil, including its concrete value. Every `ConfigurationRequirements` field needs an explicit valid decision: `Credential` and `PublicOrigin` are `Required` or `NotRequired`, and `Discovery` uses the shared supported or unsupported declaration with a safe reason. A new requirement field or discovery method needs an explicit validation update and never inherits an existing decision. Configuration discovery is distinct from resource selection discovery, and requiring a credential does not promise the `VerifyCredential` operation. These checks establish complete registration, not correct native SDK behavior; constructor and adapter contract tests still apply. + +### Configuration storage and construction + +Preview and persistence use `providers.Normalize` and `providers.Describe`. `SelectionDiscoverer` resolves omitted native values before commit, and the complete specification is validated again at persistence. `providers.ResolveChange` owns configuration inheritance, and comparisons use normalized selectors, so preview, retry and commit share the same defaults. The store owns transactions, credential encryption, generation fencing, resource ownership and generic object storage: only the adapter interprets `provider_config` and `provider_metadata`, and `provider_credential` holds ciphertext bound to the installation and generation. Retained generations keep their original public configuration and metadata and compose the current credential through the adapter, so a credential replacement never rewrites a retained selector. Database constraints check object structure, not the registration list. + +A direct adapter with a credential verifies all retained generations and allocation references before a key is replaced. The common `sandbox.CallFence` excludes native calls and waits for helper completion, including calls whose callers timed out; execution invokes the prepared verification and fencing callbacks without branching on a vendor. + +Vendor deployment validation and SDK setup stay at the construction boundary, and construction never creates an Environment. For node-local adapters `providers.Built` returns the provider, probe, installation identity, backend fingerprint and specification digest, and the factory also returns its close function. `execution.RuntimeProvider` binds the adapter to its kind, installation ID, backend fingerprint, generation, mode and node ownership; the database owns the selection, and the in-memory copy is never another authority. Docker and microsandbox run on nodes, and E2B is constructed directly. The node proxy exposes checkpoint operations only for a backend whose registered declaration supports them, and common lifecycle code admits suspension through `CheckpointProvider`, never through a provider name. + +The backend fingerprint identifies a native resource namespace, not capacity. Core keeps deployment generations so that owned allocations keep resolving to their original backend; never repoint retained allocations at a replacement backend. + +## Managed lifecycle + +This is what Core does around every provider. Adapters implement none of it, but they rely on it. + +### Deployment publication + +At startup Core claims the stable installation identity and a new owner epoch before it selects a provider. The runtime manager keeps generation-aware provider facades. Initial setup and replacement prepare and validate candidates before any database write, and a rejected candidate leaves the active configuration and workers unchanged. A backend replacement uses the deployment mutation gate: it pauses manager admission, drains old calls and loops, then repeats the resource and generation guards in the commit transaction, where the new selection, its generation and the retirement of old nodes and unused enrollment tokens commit together. After commit, Core publishes the prevalidated configuration and the shared observation and bootstrap cache under the manager mutex, with no further external work or fallible step, so a request cancelled after commit cannot discard it. An interrupted drain stays a barrier for retries. Provider I/O and draining never hold a database transaction or the manager's map mutex. + +A locally unavailable provider dependency keeps hosted admission closed while the existing scan waits for repair; administrator recovery stays available, also after a restart. Database and ownership errors stay failures, and unconfigured hosted admission creates no Session state. Core derives the Runtime bootstrap and daemon WebSocket addresses from the installation public URL, never from request headers, and reads the current selection from the database, never from a startup file. + +Node readiness binds to the exact generation, the current connection and the owner epoch. A durable serving pin is promoted only for readiness of the then-current target, under deployment serialization, so a late report for a superseded target never acquires a pin. + +### Allocation lifecycle + +The allocation, its dedicated daemon credential digest and the exact Session binding commit atomically before `Create`, under the execution lease and the Session lock. Only a fresh allocation receipt permits `Create`; retries and a Core restart observe the same reference without replaying it or rotating the credential. An allocation is private compute ownership, separate from public Environment connection and native readiness; adapters qualify bootstrap completion, and Core never infers it from an engine or provider name. + +With a configured provider, the Worker scans committed pending hosted Environments that have no allocation, which covers idle Session creation and recovery after an interruption between commit and bootstrap; an existing allocation never re-enters that path. The scan is bounded and serialized by the lifecycle owner and needs no caller action. An initial reservation without a Turn leaves its Session idle, and a daemon connection is never treated as native readiness. The same scan publishes authenticated connection observations with durable generations, after verifying the exact Session and device binding and a settled bootstrap. + +Connected, observed compute receives service keepalives between Turns. Keepalives never revive a lapse of one hour or a cleanup request. A node allocation never expires only because its keepalive is an hour old; explicit deletion and the snapshot retention still authorize its cleanup. A stopped or missing container never authorizes discarding retained workspace or history. Disabling the provider stops new hosted admission and bootstrap but never blocks cancellation, function results or input retry outcomes of existing Sessions. + +Terminal cleanup atomically revokes the device's authority, records the Environment's failure or expiry, settles pending input and requests cancellation, and only then calls `Kill`; original input deadlines and retry outcomes are kept. Temporary provider outages, unknown Create results and stopped compute never prove a permanent failure. After public Session deletion Core keeps the allocation and marks it released only after owned compute and volume cleanup and proof that the original Create settled; an unknown creation keeps cleanup ownership even after an absence observation, and bounded scans continue to catch late resources without another `Create`. + +### Per-node lifecycle workers + +Each registered node has one serial lifecycle worker that owns its gate, allocation and pending cursors, connections and wake hints; E2B allocations share one serial lifecycle without a node. A thin coordinator discovers nodes and shuts workers down, and never holds its map mutex during database, provider or wait operations. Workers advance independently, so a stuck provider on one online node never stalls another: lifecycle concurrency is one operation per node and grows with the node count. Offline workers stay, so their retained resources remain observable after reconnection. + +Allocation scans filter by node before their 32-row page limit, and pending scans join the unreleased committed placement. Each node advances its own cursor, including past failed observations, and wraps once at the end. Direct provisioning resolves the tenant-scoped placement before entering that node's gate, and an existing allocation must agree with it; Core never chooses another node. + +Before releasing the execution lease, the coordinator stops accepting work and cancels and drains every node worker and direct caller. Lease loss affects everything; ordinary provider failures stay within their node. A planned deployment drain or node retirement cancels lifecycle contexts synchronously between leased operations, through the lease gate with its five-second bound, including an active manual reconcile, and never cancels an in-flight leased query just to change configuration. A failed cancellation fence closes manager admission and reports owner failure. A failed retirement keeps the original lifecycle identity and gate until owner shutdown, and the drain barrier stays closed. Session locks, deployment capacity transactions and revision-checked receipts stay authoritative, and no external operation holds a database lock. + +### Placement and capacity + +Placement is automatic: the environment-to-node placement commits with Session creation and its retry identity, callers cannot choose a node, and a retry keeps its original node even while it is offline. Node capacity counts pending reservations and unresolved resources, and new placement and a suspended-to-restoring transition share a database lock. Unknown operations keep their reservations, source teardown must be confirmed before active capacity is released, and confirmed cleanup releases placement capacity. Retained ownership needs exact provider evidence: a socket path, a missing instance or an empty listing never proves cleanup or authorizes a replacement. + +The deployment's CPU, memory and disk settings, `max_active`, `max_retained` and the snapshot retention bound each node. There is no node-level drain, cross-node Session migration, multi-active Core, autoscaling or snapshot replication. Node removal is refused while the node holds allocations, snapshots, reservations, unknown results or cleanup, and offline ownership is kept. + +### Suspension + +A provider with checkpoint support can suspend idle work; the deployment's [`suspension`](../contracts/agents-api/sandbox-deployment.md#safe-response) policy sets the idle time and snapshot retention. Core suspends only after at least one Turn is terminal, when no root or Subagent Turn is queued, in progress or waiting, no input, file operation or initialization is pending, and real activity has been idle for the configured interval. For node allocations Core records the first root or child terminal transition with the database clock in the same transaction. Candidate filtering and the Session-locked recheck compare elapsed database time with the idle duration, and the initial snapshot retention deadline is anchored to the same database observation, so Core and database host clocks need not agree. Native completion timestamps stay unchanged in public history but never drive idle admission, and heartbeats never reset activity. Before acknowledging a planned suspension, the daemon closes admission and drains native cleanup, output receipts and file work. + +The Worker lease, the Session lock and the per-node gates own suspension for every provider. New Turn claims, file-write intents and capture admission serialize under the Session lock and share one compute-phase check; new pending work cancels a capture and wakes the same source. Normal preparation waits for the compute phase to be running, after the authenticated resume handshake, and pending input stays pending when its promotion conflicts with a lifecycle transition. Compute phases and revision-checked receipts live on the allocation. Core persists quiesce, capture and restore intent before the effect, only a fresh receipt performs a capture or restore, and recovery observes the exact attempt without retrying an unknown creation, capture or restore. A consumed snapshot never rolls a running generation back. Deletion, revocation and retention expiry win over wake, up to the final database compare-and-swap, and unknown cleanup identities are kept until owned resources are confirmed absent. Consumed artifacts and old compute are deleted, so suspension cycles never build a chain of writable disks. + +Queued work and live Environment file access wake a suspended Environment; history and published Artifact reads do not. Planned suspension uses an Environment and suspension token on the daemon connection. A PID and start-time fenced local control signal (`RunCommandCompute`) wakes the parked daemon, which authenticates again before admitting work. A transient disconnect before confirmation retries the same armed suspension with bounded attempts and backoff; a permanent authentication or protocol rejection closes it. Core owns the snapshot's retention deadline, and the daemon has no timer for it. A lost quiesce acknowledgement may thaw the same source through explicit rollback but never authorizes capturing it. + +### Reset and archive + +A [reset](../contracts/agents-api/sandbox-deployment.md#reset) is durable execution state that the runtime manager advances outside its counted work. Start, escalation, cancellation, setup, update and finalization serialize through the mutation gate. Idleness is rechecked with the Session lock and then the deployment lock, never in the reverse order during finalization. Work is keyset-paged and bound to the reset's request time and generation, with the absolute deadline and validated audit provenance persisted. + +One snapshot and timestamp partition the held resources. Offline ownership comes from the allocation's or active placement's node, with the same 45-second connection and owner-epoch predicate as online presence, independently of provider readiness; no cleanup failure, offline state or empty read authorizes a synthetic release. At zero held resources, Core drains outside database transactions, rechecks under the deployment lock and clears the deployment in one transaction, then publishes a generation-bearing empty provider without fallible work. If the final write or drain fails, it restores the committed provider with a bounded owner context before releasing the mutation gate; if that recovery fails, admission stays fenced and the owner stops. + +An administrator [Session archive](../contracts/agents-api/admin-api.md#administrative-session-archive) keeps its Project scope and hosted eligibility checks in a Session-first transaction, together with Environment expiry, cancellation, Runtime authority revocation and audit; a reset's background archive reconstructs the actual Project scope from trusted records and keeps the requester's provenance. The ordinary provider lifecycle releases compute and snapshots. Archive cancellation keeps a healthy receipt path until the terminal commit: only the archive that first revokes a device records its exact `archive_cancel_turn_id` (ordinary revocation clears it, and repeated cleanup keeps it), and the existing authenticated delivery may drain that cancellation for at most 20 seconds from the Turn's original `cancel_requested_at`. Core tracks the delivery through `done`, the cancellation acknowledgement and the terminal commit, independently of subscription removal, and grants no new connection, input, file or MCP authority or lease renewal. No transaction or lifecycle gate waits for the receipt, and a lost peer, expiry or restart falls back to ordinary failure and cleanup, never a fabricated cancelled outcome. ## Validate the integration -Run `make check-sandbox-provider-contract` while developing. It runs the shared -[`contracttest`](../services/core/internal/sandbox/contracttest) suite through -real adapter boundaries using controlled native failures, plus existing adapter -and node transport tests. The same packages are included in `make check`. -New adapters should call the public failure runner with native-side fixtures, -not substitute a fake implementation of `SandboxProvider` for the adapter under -test. Preserve tests for foreign ownership, unknown mutation results, cancellation, -no automatic replay, failed cleanup and reference-bound settlement. - -Node tests separately exercise disconnect/reconnect fencing and cleanup after a -lost create response. Provider helper protocols and the node protocol require an -exact version match and reject mismatches; do not add fallback decoders or old -binary migration. Direct in-process interfaces have no independent wire version. -Node generation management is an explicit current hello capability, not another -wire version. Fixed-configuration manual nodes use the same protocol and only -serve their enrolled deployment generation. See the -[node contract](../contracts/agents-api/node-generation-protocol.md). - -Run `make check` with its dedicated database before completion. Retain native -acceptance for SDK behavior that fixtures cannot prove: creation, lease behavior, -owned partial cleanup, declared isolation/limits, and snapshots where supported. -The opt-in Docker lifecycle/recovery tests use `AGENTS_RUNTIME_DOCKER_TEST_IMAGE`; -SDK helper tests use `make check-e2b-provider` and -`make check-microsandbox-provider`. Passing controlled contract tests is not a -claim of live cloud or model acceptance. - -Mocked compute cannot prove that resources were reclaimed or that the outer -Environment provides isolation. +Run `make check-sandbox-provider-contract` while developing. It runs the shared [`contracttest`](../services/core/internal/sandbox/contracttest) suite through real adapter boundaries with controlled native failures, plus the adapter and node transport tests; `make check` includes the same packages. Call the public failure runner with native-side fixtures instead of a fake `SandboxProvider`, and keep tests for foreign ownership, unknown mutation results, cancellation, no automatic replay, failed cleanup and reference-bound settlement. + +Node tests separately cover disconnect and reconnect fencing and cleanup after a lost Create response. Helper protocols and the [sandbox node protocol](../contracts/agents-api/node-generation-protocol.md) require an exact version match; direct in-process interfaces have no separate wire version. + +Native acceptance proves what fixtures cannot: creation, lease behavior, owned partial cleanup, declared isolation and limits, and snapshots where supported. The opt-in Docker lifecycle and recovery tests use `AGENTS_RUNTIME_DOCKER_TEST_IMAGE`; the SDK helpers use `make check-e2b-provider` and `make check-microsandbox-provider`. Mocked compute proves neither reclamation nor isolation. ## Reference adapters -| Kind | Adapter | Helper | Operator guide | +| Kind | Adapter | Helper and adapter rules | Operator guide | | --- | --- | --- | --- | -| Docker (node) | [`sandbox/docker`](../services/core/internal/sandbox/docker) | Node proxy in [`sandbox/node`](../services/core/internal/sandbox/node) | [Nodes](getting-started/nodes.md) | -| microsandbox (node) | [`sandbox/microsandbox`](../services/core/internal/sandbox/microsandbox) | [`tools/microsandbox-provider`](../services/core/tools/microsandbox-provider) | [`deploy/microsandbox`](../services/core/deploy/microsandbox/README.md) | -| E2B (direct) | [`sandbox/e2b`](../services/core/internal/sandbox/e2b) | [`tools/e2b-provider`](../services/core/tools/e2b-provider/README.md) | [`deploy/e2b`](../services/core/deploy/e2b/README.md) | - -Provider selection and E2B setup are owned by -[Sandbox deployment](../contracts/agents-api/sandbox-deployment.md). +| Docker (node) | [`sandbox/docker`](../services/core/internal/sandbox/docker) | Node proxy in [`sandbox/node`](../services/core/internal/sandbox/node); [Docker sandbox settings](../services/core/deploy/codex/README.md#docker-sandbox-settings) | [Nodes](getting-started/nodes.md) | +| microsandbox (node) | [`sandbox/microsandbox`](../services/core/internal/sandbox/microsandbox) | [`tools/microsandbox-provider`](../services/core/tools/microsandbox-provider/README.md) | [Nodes](getting-started/nodes.md) | +| E2B (direct) | [`sandbox/e2b`](../services/core/internal/sandbox/e2b) | [`tools/e2b-provider`](../services/core/tools/e2b-provider/README.md) | [Sandbox deployment](../contracts/agents-api/sandbox-deployment.md#e2b-configuration); application-managed templates in [`deploy/e2b`](../services/core/deploy/e2b/README.md) | From 224d3e6665aeb671a9cb18a1125ef4fb0d923983 Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 08:54:05 +0000 Subject: [PATCH 05/10] docs: describe Provider registration without known-gap labels --- docs/sandbox-provider.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/sandbox-provider.md b/docs/sandbox-provider.md index c83d7a1c7..af8cec2f4 100644 --- a/docs/sandbox-provider.md +++ b/docs/sandbox-provider.md @@ -111,7 +111,7 @@ A new provider takes these steps: 2. Add its specification and resource validators and, for native resource discovery before commit, an optional read-only `SelectionDiscoverer`. Put native credential verification behind `CredentialVerifier`. 3. Implement `sandbox.ConfigurationAdapter` over a typed native configuration. `DecodeInput` strictly parses the separate public `configuration` and write-only `credential` objects of a request. `Encode` produces whitelisted public selectors, read-only observations and separate secret bytes, and never passes request JSON through. `Decode` restores stored selectors, and keeps access to owned resources, without remote admission or new template validation. `Normalize` copies its input before changing it. `ResolveChange`, `Equal` and `WithCredential` own inheritance, identity and credential composition. `Requirements` declares whether a credential and a public Core origin are required and whether configuration discovery is supported. Also implement `ConfigurationDiscoverer`, even when discovery is unsupported: it validates the query and returns a safe catalog, never a mutation or an admission decision, while Core keeps authorization, input limits and deadlines. A node provider accepts only an empty public object, rejects credentials and returns Unsupported for discovery and credential replacement. 4. Register its constructor, policies, configuration adapter, operation declaration and defaults in `providers/registry.go`. Node proxy identity and checkpoint support read this entry. The installer's projection combines the registered policies with the shared field bounds in `sandbox/deployment_contract.go`; regenerate it with `go run ./services/core/cmd/specification-contract -write`. -5. Supply the distribution's artifacts for the adapter and its helper. **Known gap:** offering the provider to operators still edits shared installer and Web code with vendor-specific options, such as E2B's installer flags and Web views; this is not the pattern to follow. Never add a Session or Turn scheduling path, a vendor column or API field, or a vendor switch in the store. +5. Supply the distribution artifacts for the adapter and its helper, and offer the provider to operators. Today the installer's `--sandbox` choices and Web's setup views carry provider-specific options, such as E2B's installer flags and Web views, so this step edits the shared installer and Web code. Never add a Session or Turn scheduling path, a vendor column or API field, or a vendor switch in the store. ### Registration validation From 0962de6800948bd8220efbbd62f021e1da4c9fba Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 08:54:05 +0000 Subject: [PATCH 06/10] docs: reduce the Core service guides to contributor material The service README becomes a contributor guide: commands, database, run from source, tests and generated contracts. Operator, API and enrollment material leaves for its owners. IMPLEMENTATION.md keeps only code-level constraints; restated contracts, history and the sandbox chapter go to their owners, Vault rules link to the credential guide, and Codex adapter constraints move to the Codex Runtime page. --- docs/development.md | 4 +- packages/claude-sdk-adapter/README.md | 2 +- services/core/IMPLEMENTATION.md | 1731 ++----------------------- services/core/README.md | 687 +--------- services/core/deploy/codex/README.md | 24 + 5 files changed, 199 insertions(+), 2249 deletions(-) diff --git a/docs/development.md b/docs/development.md index 1fb0ba1a7..7160feba1 100644 --- a/docs/development.md +++ b/docs/development.md @@ -64,7 +64,7 @@ under `~/.oac/build/daemon/` by default. `OAC_DEV_HOME` selects another build ro `OAC_DEV_CORE_BUILD_DIR` selects an absolute Core output directory. Building does not configure a database, start a deployment or qualify native execution. -Use the [service guide](../services/core/README.md#build-standalone-binaries) +Use the [service guide](../services/core/README.md#run-from-source) to run the Core migrator and server with a separate development database. The [configuration appendix](configuration.md#appendix-core-environment-without-the-installer) owns standalone process settings. For a complete operator installation, use the @@ -82,7 +82,7 @@ site; `pnpm dev:docs` starts its development server. | Location | Responsibility | Read next | | --- | --- | --- | | `services/core/internal/api` | Public, administrator and machine HTTP boundaries | [API index](api/README.md) | -| `services/core/internal/store` and `internal/db` | Core persistence, transactions, queries and migrations | [Service guide](../services/core/README.md#database-ownership) | +| `services/core/internal/store` and `internal/db` | Core persistence, transactions, queries and migrations | [Service guide](../services/core/README.md#database) | | `services/core/internal/execution` | Durable Turn dispatch and scheduling | [Runtime protocol](runtime-protocol.md) | | `services/core/internal/engine` | Pure qualification of harness operations and placements | [Harness onboarding](../contracts/agents-api/harness-onboarding.md) | | `internal/agentdaemon/proto` and `gateway` | Shared wire types, validators and authenticated Runtime connections | [Runtime protocol](runtime-protocol.md) | diff --git a/packages/claude-sdk-adapter/README.md b/packages/claude-sdk-adapter/README.md index 2eed904d1..ec054fef7 100644 --- a/packages/claude-sdk-adapter/README.md +++ b/packages/claude-sdk-adapter/README.md @@ -52,7 +52,7 @@ With `ObserveToolObservations`, private workspace execution requires the package ### HTTP MCP -The private adapter accepts typed anonymous HTTP and static-bearer HTTPS MCP declarations on the trusted `environment:none` harness host. The packaged readiness report must include `mcp_http_tools`; discovery advertises that feature only when present, and execution rechecks the installed bundle before dispatching an MCP request. An unchanged SDK version alone cannot qualify an older bridge. Authenticated private requests also require the packaged `mcp_http_bearer_auth` feature at discovery and dispatch. The daemon generates a separate environment reference for each server and launch; only those references enter the bridge request and native SDK configuration. The native HTTP client expands them from its owned process environment. Literal bearers must never enter SDK MCP headers because that configuration enters argv. Readiness probes receive no per-request bearer environment. Token validation is shared with the Codex adapter; credential storage remains an opaque-string contract. Public MCP admission, Vault credential selection and their failure rules belong to [public MCP connection origin](../../contracts/agents-api/environments.md#public-mcp-connection-origin) and [HTTP MCP execution](../../services/core/README.md#http-mcp-execution). +The private adapter accepts typed anonymous HTTP and static-bearer HTTPS MCP declarations on the trusted `environment:none` harness host. The packaged readiness report must include `mcp_http_tools`; discovery advertises that feature only when present, and execution rechecks the installed bundle before dispatching an MCP request. An unchanged SDK version alone cannot qualify an older bridge. Authenticated private requests also require the packaged `mcp_http_bearer_auth` feature at discovery and dispatch. The daemon generates a separate environment reference for each server and launch; only those references enter the bridge request and native SDK configuration. The native HTTP client expands them from its owned process environment. Literal bearers must never enter SDK MCP headers because that configuration enters argv. Readiness probes receive no per-request bearer environment. Token validation is shared with the Codex adapter; credential storage remains an opaque-string contract. Public MCP admission, Vault credential selection and their failure rules belong to [public MCP connection origin](../../contracts/agents-api/environments.md#public-mcp-connection-origin) and [credential selection](../../services/core/credentials.md#use-a-credential-in-a-session). MCP queries use the SDK's main-thread Agent definition to restrict model-visible tools, in addition to empty built-ins, strict MCP configuration, empty setting sources and default-deny permissions. Permission allowlists alone do not restrict the native model inventory. Null selects all tools from a declared server; an empty list selects none. Host functions compose with those selections. Native server status supplies original tool identities; map their normalized native aliases while preserving the original names in observations. Native status deduplicates aliases, so it does not prove a complete original server inventory. A native PreToolUse hook waits for inventory verification before admitting root calls and denies unverified, mismatched or cancelled calls. The native Agent restriction controls model-visible tools; inventory verification is not a barrier before the model request. Anonymous HTTP declarations explicitly set an empty Authorization header to disable native OAuth and automatic credential injection. Preserve that header; do not erase native history or credentials to enforce this boundary. Servers that reject a blank Authorization header, normalized name collisions and inventory changes during a query require separate validation; this profile covers static inventories. Private SDK status/control objects can contain expanded authentication headers. Read only connection and tool identity fields; never retain, log or publish raw status/configuration or control responses. Diagnostic projections must whitelist safe fields; filtering actual model or tool output does not fix a leak. The adapter profile requires connected servers, reserves the `functions` label, accepts alphanumeric/underscore/hyphen server labels and alphanumeric/underscore/hyphen/dot selected tool names, and excludes remote environments. Required startup is separately qualified by `mcp_http_required`. All HTTP MCP queries use native SDK startup and an empty input iterator to confirm initialization hooks. Required declarations additionally check connected server status before the initial prompt is released exactly once. Pending, failed, missing or ambiguous required status rejects before input; native startup timeouts are retained without an adapter retry loop. Normal system/init still verifies Session identity and the complete inventory before input readiness/tool authority. Optional servers keep their inventory checks without a pre-input connection requirement. A Runtime must advertise the concrete required-initialization capability; there is no fallback to another execution path. diff --git a/services/core/IMPLEMENTATION.md b/services/core/IMPLEMENTATION.md index 22dad4f2a..ca0f1a5d6 100644 --- a/services/core/IMPLEMENTATION.md +++ b/services/core/IMPLEMENTATION.md @@ -1,1607 +1,134 @@ -# Agents API implementation constraints - -These rules describe how the Core service (`services/core`) currently -implements the public, administrator and machine contracts. They are contributor -constraints, not user documentation. Wire behavior and qualification evidence -stay in the linked [contracts](../../contracts/agents-api/README.md); Runtime -message semantics stay in the [Core–Runtime protocol](../../docs/runtime-protocol.md). - -## Public request handling - -Session creation requires initial input for `none`, and for streaming creation -outside `self_hosted`. Report inline agent protocol errors first, then check these -conditions before creation retry lookup or -resource resolution. The parser remains shared with subsequent message admission; -non-streaming hosted and self-hosted requests may omit input. Do not retain an -idle-none creation compatibility exception. Valid requests retain their documented -local idempotency behavior; clients may use the same request/key with stream=false -to recover a lost creation response. Session metadata updates require a supplied -metadata field, with null/empty clearing it. Validate an empty update before any -resource lookup, after authentication. - -List order parsing distinguishes omission from an explicit empty value. Lists and -single-resource routes ignore unknown query keys; a repeated supported list key -still rejects. The Environment Files list keeps its own path and cursor parsing but -uses the same unknown-key and duplicate-key rules, except that it still rejects -malformed query encoding (such as `%GG` or `;` separators) that the shared lists -drop. Reuse the shared parser and error serializer, preserving the observed Beta, Files and Skills error fields and -per-family limit bounds rather than applying one policy to every resource. Change -page bounds, cursor ownership or parent lookup order only with owned evidence for -that family. Record uncertain range/lookup behavior separately; do not reproduce -observed upstream server failures as compatibility behavior. See -[list rules](../../contracts/agents-api/wire-semantics.md#lists). - -Every Agents API JSON route reads its body through the shared gate -(`readJSONObject`) before route decoding, validation or lookup. It requires a JSON -Content-Type, applies the route's body limit and rejects invalid UTF-8, malformed -JSON (including unpaired surrogate escapes), repeated keys and non-object roots -with the official messages; an empty body or null becomes `{}`. DELETE, multipart, -Core extension and internal routes keep their own readers. Member names match -exactly: decode request objects with `decodeInputObject`, or check -`inexactMember` before another decoder, so that encoding/json never matches -a case variant to a field. See -[request bodies](../../contracts/agents-api/wire-semantics.md#request-bodies). -Report validation failures with official evidence through the typed field error, -which emits `invalid_request_error` with the observed param and message; keep -other local codes until their official fields are sampled. Every 409 has type -`conflict_error`. Session input conflicts and changed tool results also use code -`conflict_error`; documented Core-only conflicts, such as Idempotency-Key reuse, -sandbox administration and Environment input states, keep their local codes. Agent configuration -(saved create/update and the inline Session agent) uses one path-tracking -validator of the pinned shapes before its parsers and harness admission, which -keep their local codes; do not grow it into a JSON Schema engine. A malformed path -identifier must produce exactly the response of a well-formed missing one on that -route, including invalid bodies, queries and storage availability: resolve it to -the never-assigned maximum UUID and let the missing path run, or reject it -directly only where the lookup is the next check. Request-body references keep -their own errors. An `after` cursor that does not resolve inside its already -resolved parent, malformed ones included, returns that list family's observed -error: the missing-resource 404 on lookup lists, otherwise the typed store cursor -error. Foreign and missing cursors stay identical; see -[list rules](../../contracts/agents-api/wire-semantics.md#lists). Reject U+0000 in metadata -explicitly with its `metadata.` param; other stored strings rely on the -PostgreSQL error mapping, so keep each request's writes in one transaction. See -[validation errors](../../contracts/agents-api/wire-semantics.md#validation-errors). - -Serve requests on their canonical path and never redirect. `api.CanonicalPaths` -wraps the complete server handler in both configurations (the daemon ServeMux and -the API router alone), so every route group, middleware, authentication check and -handler sees one path. It starts from the request's own spelling, never a path -re-escaped from its decoded form: invalid bytes are percent-encoded, unreserved -escapes decoded, empty and dot segments resolved with ServeMux semantics, the -trailing slash kept, and other escapes such as `%2F` and `%5C` left encoded; -`Path` and `RawPath` are set consistently for chi and the ServeMux. Do not route -or authorize on a path outside that wrapper. On the Beta -group the constant OpenAI-Beta check (exactly one `agents=v1` value) runs before -authentication, and authentication still precedes every Beta handler, 404 and 405. -Every Agents API 401 has type `invalid_request_error`: null code on Beta routes; on Files -and Skills `invalid_api_key` only for a rejected Bearer credential. Agents API responses carry a fresh `X-Request-Id` (also in the log -context), `OpenAI-Version`, `OpenAI-Processing-Ms` and nosniff through the API -router's own middleware, not the shared log middleware. HEAD runs GET routes; -streaming, content-download, live directory, Runtime observation and Runtime -history routes register an explicit HEAD 405 instead. Every 405 of the API router, unknown methods included, has the JSON -body and lists the route's methods in `Allow`. +# Core implementation constraints + +These are the code-level rules of `services/core` that no contract states. Contracts own behavior: the [coverage ledger](../../contracts/agents-api/README.md) and its linked contracts own the public and administrator APIs, the [machine connection API](../../contracts/agents-api/machine-api.md) the `/api/v1` routes, the [Core–Runtime protocol](../../docs/runtime-protocol.md) the daemon wire, and the [Sandbox Provider guide](../../docs/sandbox-provider.md#managed-lifecycle) the managed compute lifecycle. When code changes one of these rules, change the rule here in the same branch. + +## Request handling + +Every Agents API JSON route reads its body through `readJSONObject` before decoding, validation or lookup. The gate requires a JSON Content-Type, applies the route's body limit and rejects invalid UTF-8, malformed JSON (including unpaired surrogate escapes), repeated keys and non-object roots with the official messages; an empty body or `null` becomes `{}`. DELETE, multipart, Core extension and internal routes keep their own readers. Member names match exactly: decode request objects with `decodeInputObject`, or check `inexactMember` before another decoder, so `encoding/json` never matches a case variant. + +Report a validation failure that has official evidence through the typed field error, which emits `invalid_request_error` with the observed param and message; keep other local codes until their official fields are sampled. Saved and inline Agent configuration pass one path-tracking validator of the pinned shapes before their parsers and harness admission; do not grow it into a JSON Schema engine. A malformed path identifier must produce exactly the response of a well-formed missing one on that route, including for invalid bodies, queries and storage availability: resolve it to the never-assigned maximum UUID and let the missing path run, or reject it directly only where the lookup is the next check. An `after` cursor that does not resolve inside its already resolved parent, malformed ones included, returns that list family's observed error, and foreign and missing cursors stay identical. U+0000 is rejected explicitly only in metadata (`metadata.`); other stored strings rely on the PostgreSQL error mapping, so keep each request's writes in one transaction. + +List queries reuse the shared parser and error serializer while keeping each family's limit bounds and error fields. The Environment Files list keeps its own path and cursor parsing but follows the same unknown-key and duplicate-key rules, and still rejects malformed query encoding that the shared lists drop. Change page bounds, cursor ownership or parent lookup order only with evidence for that family, and never reproduce an observed upstream server failure as compatibility behavior. + +Requests are served on their canonical path and never redirected. `api.CanonicalPaths` wraps the complete server handler in both configurations (the daemon ServeMux and the API router alone), so every route group, middleware, authentication check and handler sees one path. It starts from the request's own spelling, never a path re-escaped from its decoded form: invalid bytes are percent-encoded, unreserved escapes decoded, empty and dot segments resolved with ServeMux semantics and the trailing slash kept, while other escapes such as `%2F` and `%5C` stay encoded; `Path` and `RawPath` stay consistent for chi and the ServeMux. Never route or authorize on a path outside that wrapper. + +On the Beta group, the OpenAI-Beta check (exactly one `agents=v1` value) runs before authentication, and authentication precedes every Beta handler, 404 and 405. The API router's own middleware, not the shared log middleware, sets a fresh `X-Request-Id` (also in the log context), `OpenAI-Version`, `OpenAI-Processing-Ms` and `nosniff`. HEAD runs GET routes, except that streaming, content-download, live directory, Runtime observation and Runtime history routes register an explicit HEAD 405. Every 405 of the API router, unknown methods included, has the JSON body and lists the route's methods in `Allow`. ## Source Files and Artifacts -Source Files belong to the execution project and have an independent lifecycle -from copied workspace files. Store immutable source metadata and PostgreSQL large -objects in the execution database with the pinned pgx driver. Upload validation, -metadata insertion and bytes commit atomically; deletion removes metadata and -unlinks the object in one transaction. Keep OIDs private and authorize every -metadata/content/delete lookup by tenant before opening a body. Stream bounded -chunks; never hold an entire general Files upload in memory or use filenames as -filesystem paths. Public source download admission is separate from internal byte -consumption: reject direct downloads of the supported `user_data` purpose after -tenant-scoped metadata lookup; initialization and workspace copies keep their -authorized Store read. Core Web must not offer that unavailable download action. -A read-only repeatable-read transaction preserves an admitted -source snapshot across concurrent deletion. Resolve that snapshot before entering -the existing Environment write path; deleting a source does not undo a completed -workspace copy. Bound request/transaction lifetimes, roll back incomplete bodies, -and never automatically retry ambiguous commits. Backups must include PostgreSQL -large objects; live deletion does not erase WAL or historical backups. Schema -rollback must not orphan existing source objects. Do not reuse product capability -tables or introduce a second destination writer for file_id. - -Session Artifacts are immutable published output copies, separate from live -workspace files and general source Files. A private output exporter must reuse -the authorized workspace path boundary and stream bounded bytes. Require complete -capture and confirmed helper/transport success before publication; valid archive -syntax alone is insufficient. The daemon owns and drains the exporter stdout pipe -separately from child reaping, so pull-transport backpressure cannot consume a -process-exit I/O deadline. After helper exit, bound each actual pipe read to one -second to reject inherited pipes that never close; reset that allowance after -consumer delays. Cancellation closes the owned reader and the dispatch consumer, -then waits for child settlement. An arbitrary blocked writer cannot be interrupted -by the exporter itself. Never extract an output archive into Core's -filesystem or hold the global execution lease through a large transfer. Keep -publication ordered with Turn completion, and authorize stored reads independently -of Environment availability so published outputs can survive its expiration. -Capture bytes into private PostgreSQL large objects without a Session admission -lock; after confirmed export, lock and recheck the live Turn before staging metadata. -Before capture, seal native input under that lock using the private capture marker. -Later messages reuse the existing Environment input reservation and await the next -Turn; the public Turn stays in progress until publication settles. Directory reads -during capture use an independent authorized read-only preparation, not the released -native Run; they do not request model credentials or mutate the workspace. Cancellation -retains the existing Turn/reservation semantics. Do not introduce a second queue. -Publish metadata in the same transaction as Turn completion. Failed/cancelled Turns -discard private objects, and Session deletion removes both private and published -copies. The exporter skips output symlinks by their `lstat` type without following, -opening or resolving them; hard links, other special files, device crossings and -concurrent changes still reject the capture. In that completion transaction, drop -staged paths whose sha256 equals the newest remaining published Artifact for the -path in the Session, so later Turns publish only new, changed or no-longer-published -paths and never modify existing Artifacts. Reuse the source-file snapshot reader -pattern and common content response; artifact deletion does not alter workspace -files. Hosted execution requires the -Runtime's bounded output-export capability and exact read-only preparation binding; -capability advertisement alone does not qualify an operator's deployment. -Exporter component checks do not establish public Artifact compatibility. - -## Dispatch and pending input - -Environment initialization is owned by the leased Worker's common preparation -scheduler, independently of any RuntimeAllocation. Both managed and enrolled -connections use the same frozen files/setup/capability snapshots and typed Runtime -operations. Preparation has bounded concurrency separate from Turn scheduling; -a blocked Runtime must not block resource cleanup or other connections. Check -Harness availability before installing. A persisted running initialization whose -owner is lost fails without replaying side effects. Completion requires the same -authorized Environment/device binding. Failure settles pending input and preserves -compute ownership and user files. Connection observations do not imply completion. - - -Successful input commits send a coalesced hint to the existing Worker scheduler. -The scheduler keeps lease, capacity, cursor fairness and per-Session ownership -checks; a hint does not admit work itself. If capacity is occupied, preserve one -rescan for completion without turning failed preparation into a busy retry loop. -HTTP readiness waits and active input delivery subscribe before reading relevant -state and wake after committed promotion or input. They recheck storage after each -hint. Polling remains the fallback for external writers, expiry and lost hints; -notifications contain no execution authority and no durable input data. - -Core readiness, Start acknowledgement and input-to-first-text logs use a single -process monotonic clock. Start acknowledgement confirms adapter ownership, not -model input consumption. -Runtime logs identify Executor creation, reuse, idle and close independently of -Turn completion. Do not call these durations model-only latency or subtract clocks -from different machines. Native process creation and same-owner successive Turns, -real provider results and same-condition timings must substantiate reuse claims. - -The Dispatcher prepares pending Environment input only on its exact enrolled or -managed device. Preserve the same physical peer and preparation handle through -readiness, atomic promotion/claim and the first non-replay Start. Derive workspace -and Environment identity from Store ownership, never caller-selected private paths. -Initial prompt/cursor come from the reserved batch; later messages use ordinary -steering. Never hold a database lock during native preparation. - -Observe the original pending deadline, cancellation, deletion and peer loss while -waiting for readiness. Preparation failure leaves pending input and its deadline -intact unless storage has settled it; it creates no failed Turn or input history. -During Start, consume preparation controls alongside the ordinary Run stream so a -control-only rejection or pending-start cancellation can settle promptly. Reuse -ordinary journal, receipt and completion/native-history persistence. Once cancellation -is sent, preparation errors/closure cannot replace its receipt or timeout path. -Pending-start cancellation uses the adapter's observed outcome after daemon handoff -and output forwarding. Missing or unconfirmed outcomes still fail conservatively; -preparation control errors cannot substitute for the cancellation receipt. -Do not fabricate an empty cancellation outcome or infer native quiescence. -The connection owner spans preparation and the transferred Run without a reservation-derived Run -deadline; every exit releases it. - -The existing Worker scans pending inputs using the same bounded scheduling slots, -Session locks, durable deadlines and engine capability checks. Keep the one-second -Environment-input scan cadence and at most 100 candidates per scan. At EOF after a -nonempty cursor, refill the first page once in the same scan; an empty queue must -not spin. Advance the cursor before readiness checks so an unavailable Runtime -cannot starve later candidates. Preserve the configured execution concurrency (default four) and alternation -between ordinary Turns and Environment inputs. A self-hosted Session -waits for its dedicated enrolled device; it cannot select an arbitrary same-tenant -device or migrate an existing binding. Preparation failure can retry while still -pending without extending the deadline. Preparation retries share this scan cadence; -execution concurrency bounds simultaneous work, not attempt frequency. Managed-provider lifecycle -polling retains its separate five-second interval. Unknown promotion results or -errors after admission retain the existing no-replay settlement rules. - -`OAC_PUBLIC_URL` enables the private gateway; Core derives the public -`remote_url` from it. `OAC_HARNESSES` explicitly adds deployment-supported -engines to the default engine and configured managed profiles; advertising a -heartbeat alone does not enable an engine. The three native profiles share enrollment -at `/workspace`. Their new user-managed public chain requires fixed-client/raw HTTP, -real-model, recovery, cancellation and credential-lifecycle acceptance separately -from prior Docker or retired remote-executor evidence. - -## Current implementation constraints - -The constraints below describe existing code, not requirements to preserve legacy -design. The [protocol assessment](../../contracts/agents-api/README.md#known-gaps) -identifies replacements and gaps. Update these rules when their implementation is -replaced; do not carry obsolete compatibility code forward to satisfy this section. - -### Gateway, Session deletion and schema ownership - -- `internal/agentdaemon/gateway` is the shared daemon connection implementation. - Its persistence interfaces use `internal/agentdaemon/device`, never product Store - types. Shared protocol frames and validators live in `internal/agentdaemon/proto`. -- Session deletion uses a durable `sessions.deleted_at` marker. Public deletion - accepts only a durably idle or failed Session without required actions: no - queued, in-progress or waiting root Turn and no pending input reservation, the - same settlement rule as the creation stream. Subagent child Turns and pending - Environment file writes are not checked, as before this rule; their official - behavior is unobserved. Take that decision and commit the marker - under the tenant Session lock that orders Turn and input admission, so either - admission commits first and deletion conflicts, or admission observes the - deletion. A busy Session returns 409 `conflict_error` with the observed official - message and nothing changes: no cancellation, marker, event or cleanup. Callers - cancel first (`agent.session.input.cancel`), wait until the Session is idle and - delete it. Core admits a Turn synchronously, so it also conflicts right after an - `events.create` 202, where the official service was observed to return 200. - Deleting a provisioning hosted Session with reserved input used to release its - sandbox node placement at once; it now conflicts, and the placement counts - toward node capacity until the input is admitted or its five-minute deadline - expires. A later allowed deletion releases an unallocated placement. - The owner's repeated deletion returns the same 200 confirmation without writing; - foreign, missing and malformed identifiers keep the byte-identical 404. Public - reads, metadata changes, event streams and input admission exclude deleted - Sessions; admission checks visibility under that lock before retry lookup. - Creation keys remain reserved and cannot resurrect deleted Sessions; reuse of a - deleted creation identity returns 409. Existing streams close when removal is - observed without a fabricated deletion event; overlapping stream timing remains - unverified. Earlier releases also deleted busy Sessions after requesting - cancellation, so upgraded databases can hold markers with hidden work. - Internal Turn/receipt/finalization and restart reconciliation retain access so - that work settles under the existing execution lease; queued work cannot be - claimed, and Runtime cleanup still cancels pending work. Confirmation does not - guarantee native quiescence. Never revoke a shared device, - remove a saved Agent or touch product data as part of Session deletion. Physical - SQL/native history cleanup remains a separate required implementation gap; these - records are retained, not claimed purged, and purging may end repeat idempotency. - Do not deploy a pre-deletion service - against a database with deletion markers; migration rollback refuses to remove - the column while deleted records exist, preventing public resurrection. -- `services/core` owns its SQL schema, sqlc queries and embedded goose - migrations. `OAC_DATABASE_URL` is required; never fall back to the product - database URL. It stores tenant-scoped Sessions with a stable engine and durable - creation retry identity. - -### Agents, Vaults and credentials - -- Reusable Agents have their own tenant-scoped `agents` records, independent of - Session snapshots, engine bindings and product Agent definitions. The Store - persists caller-validated non-secret configuration and metadata without applying - harness capability restrictions or model defaults. Resource identity and equal - initial creation/update timestamps come from the persistence boundary. The - create primitive creates a fresh resource; public retry conformance remains - unverified. Internal storage admission is 512 KiB for configuration and 64 KiB - for metadata, not a claim about upstream limits. -- Vaults have their own tenant-scoped records in the execution database. Initial - `POST /v1/vaults` and `GET /v1/vaults/{vault_id}` operations persist and read the - resource without creating Sessions or contacting an engine. Omitted name is - null; a supplied non-null string is trimmed and limited to 1–256 UTF-8 bytes. - Omitted/null metadata becomes an empty object. Reuse the 64 KiB encoded metadata - storage bound, without applying Session-specific pair/character limits. This is - a local bound, not hosted parity. `GET /v1/vaults` lists the authenticated project's - records using creation-time/ID keysets, default descending order and a default - limit of 20 clamped to 1–100. Other resource limit policies are unchanged. - Status accepts `active`/`archived` as a scalar or SDK `status[]` array, with both - included by default. Private stored classification defaults existing/new rows - to active; it is never exposed in the Vault response. Listing reads no Credentials - and needs no encryption key or execution service connection. A scalar status - combined with `status[]` filters by their union; a repeated scalar parameter is - rejected. Exact hosted errors, equal-time - ordering and changes between pages remain unverified. Private archived fixtures - prove filtering only: there is no public archive writer, archive timestamp or - inferred delete-to-archive behavior. Retrieval, Session binding and dispatch retain - their existing rules. Archive/revocation lifecycle remains a separate gap; - do not introduce product roles or speculative lifecycle fields. Migration rollback - refuses to discard classification while archived rows exist. -- Vault `DELETE /v1/vaults/{vault_id}` removes the project-owned parent and all - Credentials through the existing foreign-key cascade in one SQL mutation. Do not - loop through child deletions, decrypt secrets, require the storage key or call - providers. Deletion applies to both stored classifications. Local missing/repeated - deletion returns not-found; subsequent parent/child reads and new attachments - cannot use the removed resources. Preserve Session snapshots, frozen choices, - history and recorded retries. Later secret lookups fail without selecting another - attached Vault or anonymous MCP. Already-resolved tokens and running Sessions - are not revoked. Exact hosted archive, visibility, overlapping-mutation and error - semantics remain unverified; row removal does not prove physical storage erasure. -- Static-bearer and OAuth Credentials are children of tenant-owned Vaults in the execution - database. Creation admits the owner in the same SQL statement as the insert; - retrieval joins the owning Vault and selects public metadata only. Public reads do not decrypt or return a token. OAuth replacement authenticates - the stored grant before applying a partial update. Encrypt before passing secret values to - SQL, using the execution service's separately configured random 32-byte key and - the standard library's random-nonce AES-GCM. The versioned authenticated binding - includes tenant, Vault, Credential, authentication purpose and exact destination. - Never reuse product master-key conventions or daemon transport encryption for - this storage boundary. Missing key configuration disables credential writes; - malformed explicit configuration fails startup. See - [Vaults and Credentials](../../contracts/agents-api/vaults.md#storage-key) for - key persistence and current limits. Storage-key rotation remains - separate work; resource creation never contacts the destination. -- `GET /v1/vaults/{vault_id}/credentials` lists safe metadata only, with both - project and Vault ownership enforced on the parent, cursor and row query. An - inaccessible parent returns not-found, even when the collection would be empty. - Reuse the Vault status/limit parser and Credential metadata mapping. SQL must - never select ciphertext for listing; no encryption key or execution is needed. - Credential status is a separate private active/archived classification, defaulting - historical/new records to active and never derived from the parent Vault's status. - Synthetic archived fixtures prove filtering only. There is no public archive writer, - timestamp or delete-to-archive inference; existing create/retrieve/token replacement, - Session bindings and dispatch keep their rules. Migration rollback refuses to lose - archived classification. Archive lifecycle and hosted query/concurrency - semantics remain gaps. -- For static auth, Credential `POST /v1/vaults/{vault_id}/credentials/{credential_id}` - replaces only the token and update time. Require `auth.type=static_bearer` and a - string `auth.token`, preserving opaque bytes; reject extra mutation fields before - writing. Reuse safe metadata for the immutable encryption binding, then scope the - atomic SQL mutation independently by tenant, Vault, Credential, static auth type - and exact destination. Never decrypt the previous token or send plaintext to SQL. - Missing encryption configuration or a failed write preserves the old row. Return - the existing safe metadata projection; identity, name, destination, creation time - and Session snapshots stay unchanged. Subsequent dispatch reads use the committed - replacement through existing scoped lookup; already-resolved requests may retain - the old token. This is not storage-key rotation, in-flight revocation or hot reload. - Exact hosted concurrent-update/retry/timestamp semantics remain gaps. -- OAuth grant ownership stays in Core. The application performs authorization and - provider revocation; do not add public login/callback/refresh/revoke routes. - Store access/refresh/client secrets together under existing authenticated tenant, - Vault, Credential, auth-type and destination encryption. Authenticate refresh - metadata against its encrypted copy before using an endpoint or grant. Read/list - queries still select safe metadata only. Shared MCP selection admits both auth - types and freezes one identity without changing native adapter contracts. - At dispatch, a known-expired grant is refreshed through the declared endpoint - auth method, with stored scope/resource, then persisted before returning access. - Serialize refresh and replacement with the same PostgreSQL Credential row lock; - deletion and Vault cascade cannot be undone by a stale refresh. Network exchanges - are bounded and fail closed; never return provider error bodies or claim an - uncertain grant exchange was committed. No background scheduler, 401 retry, - hot replacement, output repair or harness-specific OAuth path is introduced. - Refresh uses verified HTTPS, rejects redirects, and checks resolved addresses - before dialing them. Private issuer origins need explicit operator configuration - in `OAC_OAUTH_TRUSTED_ORIGINS`; tenants cannot relax that boundary and TLS - verification remains mandatory. Keycloak is acceptance infrastructure only. - Preserve the pinned update omission/null and immutable-field rules described in - [OAuth credentials](../../contracts/agents-api/vaults.md#oauth); record unspecified - hosted semantics. Native processes receive only access tokens. Provider revocation, - withdrawal of already-dispatched tokens and Session cancellation remain distinct. -- Credential `DELETE /v1/vaults/{vault_id}/credentials/{credential_id}` removes one - owned row, including ciphertext, with tenant/Vault/ID checked in the same SQL - mutation. It needs no encryption key, secret read or network call. Local reads, - updates, listings and subsequent dispatch lookups cannot use that ID; missing - and repeated deletion return not-found. Preserve frozen Session choices, retry - identities and history without fallback to another credential or anonymous MCP. - A token already read before deletion may remain in a dispatched request. Deletion - does not revoke provider tokens, cancel Sessions or prove physical erasure from - native history, WAL or backups. Do not infer a delete-to-archive mapping; exact - hosted archive, post-delete visibility and repeat/error semantics remain unverified. -- Public reusable Agent create/retrieve uses `/v1/agents` and the same authenticated - tenant/Beta-header boundary as Sessions. The resource envelope owns identity, - timestamps and metadata, separately from saved configuration and Session state. - Resolve known defaults and validate supported schema before writing. Preserve - model/name/instructions verbatim, nullable fields and structured JSON numbers. - Stored reasoning/service tiers, enabled multi-agent settings, JSON Schema output, - enabled web-search modes and deferred/tool-search/programmatic tools do not imply - execution support. Reuse function wire validation, keeping Session execution - restrictions separate. Model-derived reasoning effort is unresolved when omitted; - do not infer it from the selected harness. Omitted/null service tier currently - uses `auto`; complete upstream default/error/retry conformance and remaining MCP - variants remain gaps. Unknown/unsupported variants fail explicitly. No product - lookup is permitted. - -### Service-origin MCP - -- Service-origin public HTTP MCP uses the native harness client and tool loop on - trusted service-owned `environment:none` compute, with Codex or Claude SDK. - The V1 colocated `self_hosted` profile rejects it: user-owned compute cannot be - relabeled service-origin or receive its attached Vault credentials. Hosted - service-origin MCP also remains unsupported. Environment-origin Plugin MCP uses - its separate qualified local Runtime transport and isolation contract; do not - disable that path or infer optional-feature equality across engines. - Admission, device selection and preclaim still require the exact supported MCP - capabilities. Native declarations do not widen public placement authorization. - -- Keep accepted public MCP credential profiles separate from private adapter - capabilities. The execution service declares the verified public bearer profiles - centrally; a daemon capability alone cannot open a public profile. Reuse frozen - binding validation at admission and later input, then the same MCP capability and - placement checks at device selection, final preclaim and request construction, - before scoped decryption. Native configuration and token injection stay in the - adapters. Add bounded shared checks while changing the related execution path; - do not defer known duplication to a general engine or plugin framework. -- The shared MCP resolver preserves omitted/null `allowed_tools` as unrestricted - and an explicit empty list as deny-all. Saved HTTP transport output includes - `headers:{}`; the effective Session transport omits headers, matching the two - pinned resource types. Saved-Agent updates never change existing Session - snapshots; per-Session tools replace the whole field. The initial profile admits - HTTP(S), boolean `required` (default false), empty/null metadata and empty/null headers. - Static and OAuth bearer authentication require HTTPS and the attached-Vault rules below. - An omitted/null `connection_origin` on HTTP transport is stored as `service` - before the other checks, identical to an explicit declaration, as observed - officially. Inline authorization, URL userinfo/query/fragment, the `environment` - origin and stdio remain explicitly unsupported. -- Codex required MCP initialization additionally needs `mcp_http_required`, advertised - only for the verified native pin and checked during selection, final preclaim - and daemon dispatch/preparation. Preserve the boolean through typed messages, - native rendering and exact configuration preflight. Reuse native required-server - initialization during root thread creation and cold resume; send no native Turn - until it succeeds, and never replace a failed strict resume with a new thread. - Public work may already be accepted/queued/in progress while native initialization - waits. This is not a continuing health monitor or a new public readiness state; - exact hosted Session creation timing and initialization errors remain unverified. -- Send MCP declarations through typed daemon fields, independently of function - callbacks. In Codex, a non-nil declaration replaces operator MCP options; use the existing - native renderer and original tool names for `enabled_tools`, including `[]`. - Before thread creation/resume, query native `config/read` with the exact cwd and - reject additional servers or effective configuration differences. Disable native - plugins/apps and select file-only MCP credentials; reject existing credentials - in the private native home without deleting them or native history. Native - reserved labels are an adapter restriction, not a saved-resource schema rule. - Requests without this typed field keep the existing product behavior. The check - is a snapshot on trusted service compute, not an atomic barrier against concurrent - operator configuration changes. Discovery of a declared deny-all server can still - contact it; deny-all governs tool exposure. Reuse neutral tool observations and - the existing public `mcp_call` projection, never add a second MCP/model loop. -- The private Codex adapter's HTTPS MCP bearer authentication requires - `mcp_http_bearer_auth` and the existing MCP/environment capabilities, checked - before the factory. It is restricted to trusted service-side Codex with - `environment:none`. A transient - per-server `bearer_token` becomes a fresh daemon-owned `bearer_token_env_var` - reference for each native process. Put the exact secret only in that app-server - child's environment, after auxiliary launch probes; never in global environment, - arguments, configuration/history, public snapshots or logs. Preflight accepts - only the expected server/reference pairing and retains the existing rejection - of ambient credential sources. Service-side bearer variables must not enter - generated commands, files or public history/snapshots. Use the native HTTP client - with TLS verification. - This execution profile rejects empty values and bytes outside RFC 6750 b64token - syntax with generic errors; it never trims tokens or narrows opaque Credential - storage. Core-managed OAuth uses this same access-token path; native OAuth - login/refresh and hosted redirect/error equivalence remain separate work. - -### Session resources and model providers - -- Session `vault_ids` omission/null/empty means `[]`; nonempty attachments must all - belong to the authenticated tenant. Preserve caller order and the stored caller - `credential_id`. Saved Agents may store a nullable/nonempty credential reference - without authorizing its use. Session admission resolves an explicit credential - only inside attached Vaults for the exact declared URL, or selects the unique - matching static or OAuth credential when the ID is omitted/null. No match remains - anonymous. Selection errors use the observed official messages with a null param: - a reference without attachments, one outside the attached Vaults and one for - another URL are 400 `invalid_request_error`; several implicit matches are 409 - `conflict_error`. Missing, foreign-tenant, unattached and malformed references - share one message; echo caller values only within `internal/echotext`. Unknown - or foreign Vaults keep the same 404. Resolve after the input requirement and - before any Session, initial input or event write. Freeze safe bindings, - including anonymous decisions, in private Session configuration. Session - projections, never the stored configuration, show an implicitly selected ID in a - null/omitted public `credential_id`, also after deletion. At actual dispatch, recheck - tenant, attached Vault, selected ID, frozen auth type and exact URL before scoped - decryption. Metadata queries select no ciphertext; tokens enter only the existing - transient daemon request. Selected authentication requires `mcp_http_bearer_auth` - during device selection and the final preclaim check. Missing/wrong keys or - binding failures never fall back to anonymous execution. Exact URL equality, - immutable selection timing and hosted redirect semantics remain local decisions - or unverified gaps. No new MCP loop is permitted. -- Saved Agent execution defaults use separate input and safe-output Core extensions. - Keep model-provider bundles whole at every replacement boundary: endpoint, key, - protocol and limits must never be independently inherited. Ordinary Agent JSON - contains only safe provider fields and an output-only configured flag; encrypt the - complete bundle separately with tenant/Agent binding and a distinct purpose. - Commit configuration and secret changes together under the Agent row lock. Merge - only the extension's supplied members; omission preserves, provider null clears - its bundle, and extension null clears both defaults and secret. Model-only edits - require no key. Validate the merged harness/protocol/limits without reading keys. - Read safe defaults and ciphertext in one database snapshot for Session creation; - a complete Session override need not decrypt the inherited bundle. -- Session provider selection resolves explicit bundle, saved bundle, then the - deployment default of the resolved harness, and freezes it in the existing - encrypted Session-owned row. The deployment default is a runtime setting: one - complete bundle per harness in PostgreSQL, encrypted with its own - harness-bound purpose, managed only with the Core key through - `/core/v1/harnesses/{harness}/model-provider`, audited as a deployment-wide - write without the key, and never read back. It applies to `openai_hosted` and - operator-registered `none` Sessions, never to `self_hosted`, whose executor host - belongs to the application. Caller bundles apply to `openai_hosted` and - `self_hosted`. Hosted and self-hosted Sessions with no bundle are rejected at - creation with `model_provider_required`; legacy rows without a snapshot are - rejected at message admission and dispatch, never sent to a harness that would - fall back to a built-in endpoint. There is no operator options file: its - retirement fails startup. Sessions retain only the common frozen provider - bundle; historical native-option snapshots are unsupported. Keep runtime dispatch on the common adapter path and fail closed for - missing/decryption-failed snapshots. Agent edits/deletion, default changes, - restart and idle suspend/resume never resolve defaults again. Record caller - intent for every new hosted Session before resolving defaults; other inline - Sessions keep the resolved-request rule, whose hash leaves out the deployment - default. Provider keys enter retry hashes only as keyed fingerprints. Matching - retries return committed state without replay. No Turn-level overrides or provider catalog is - included. Public input/null semantics and examples live in - `contracts/agents-api/model-execution.md`. -- Deployment-default observations use a private UUID generated on every PUT, - including identical replacements. Read its ciphertext/revision together and freeze - the pair only at new Session creation; retries and historical Sessions never gain - or replace revision metadata. After a successful root terminal commit, one - independent pool operation has at most one second to update the matching current - revision. SQL verifies tenant, root Turn and committed outcome; only completed - Turns and fixed native provider failures with authoritative `engine_failed` qualify. - Never observe cancelled work, Core/runtime errors or input-policy classifications. - The metadata-only transaction sets server statement/lock timeouts within the - remaining client budget and issues one UPDATE, so a client timeout cannot leave - an indefinitely waiting statement. The update locks only the default and samples - DB receipt time after the lock. - Errors throttle for 30 seconds regardless of code; ordinary successes throttle - for 30 seconds, with one immediate recovery write after each accepted error. - An unchanged revision has at most three effective writes in any half-open - 30-second interval of nondecreasing DB time. Clock rollback may suppress ordinary - writes; private recovery state permits only one recovery without clamping time. - Observations never change `updated_at`, readiness or execution truth. They can be - lost or stay stale indefinitely without another eligible Turn; there is no queue, - retry, probe, history backfill or provider text in the safe fields. -- Session execution-configuration reads use a separate immutable safe projection, - written with provenance in the same creation transaction as Session resources. - Read no credential ciphertext and never recompute sources from current Agents - or operator defaults. Model/harness sources are independent; provider bundles - retain one source. The read is Core-key only, under `/core/v1`; `/v1` has no - execution-configuration read. Old Sessions expose persisted model/harness with - unknown sources and unavailable provider metadata, without backfill. Projection metadata does not alter retry - identity; retries cannot replace it. Keep this administrator query separate from - runtime observations and do not touch activity or wake sandboxes. The versioned - contract is [execution configuration](../../contracts/agents-api/admin-api.md#execution-configuration). -- Model communication uses native direct connections only. Core sends one frozen - confidential `model_provider` bundle, independent of engine and placement; - adapters apply it through their native provider configuration. The shared - `internal/harnessconfig` descriptor declares one ordered `protocols` list, - with the first entry as the default. Core admission and Runtime enforce it. - There is no built-in model API proxy, passthrough gateway or cross-protocol - conversion, including inside individual Harnesses. Unsupported combinations - fail explicitly; saved configurations and immutable Session snapshots are - never rewritten, aliased or migrated to another protocol. Provider validation, - credential encryption, capability checks and native lifecycle ownership remain - mandatory. See `contracts/agents-api/model-execution.md` for the current - protocol matrix. Go 1.26.8 is the pinned build toolchain. -- Provider input validation uses the adapter-owned rules in `internal/harnessconfig`. - Keep one internal registry for protocol and token-limit validation; Core owns - credential environment and endpoint admission policy. These rules are not a - public discovery API or Runtime registration descriptor. Operation qualification - and live readiness retain their existing owners. There is no startup - configuration read; Session frozen execution-configuration reads remain a - separate administrator read. - -### Agent updates, listing and tenancy - -- Public Agent updates use `POST /v1/agents/{agent_id}` with the same tenant/Beta - boundary and shared saved-field validation. Preserve omission separately from - null; only supplied fields replace saved values. Metadata is a separate whole-map - replacement, with null/empty clearing it. Lock the tenant-owned Agent row while - merging validated fields and enforcing the complete configuration bound, then - commit configuration, metadata and update timestamp together. Never write a stale - full snapshot over another update. No-field updates read without changing timestamps. - Except for the Core execution-default extension described above, supplied nested - fields replace the whole field and explicit null uses - existing saved defaults; exact hosted nested/null and no-op timestamp semantics - remain unverified. Model-derived reasoning defaults remain a separate gap. - Neither updates nor retries modify existing Session snapshots or execution state. -- Public Agent deletion uses `DELETE /v1/agents/{agent_id}` and one tenant-scoped - `DELETE RETURNING id` statement. Return the stored canonical ID with - `object=agent.deleted` and `deleted=true`; missing/repeated deletion locally - returns not found. It never deletes Sessions, history or runtime state and does - not cancel accepted execution. Recorded creation identities still recover the - accepted Session; new references cannot resolve an absent source. Historical - identities retain their documented limitation. Exact hosted errors and ordering - of overlapping source creation/deletion remain unverified; no tombstone or - successful result is fabricated for an absent resource. Reject body data; unknown - query keys are ignored. -- Reusable Agent listing uses the same tenant/Beta-header and response mapping as - create/retrieve. Page by `(created_at, id)` with a same-tenant saved-Agent cursor; - listing never resolves Sessions, product objects or execution capabilities. - Reuse shared list-query parsing and its per-family limit policy. Agent, Session, - Item, Subagent Item and Template lists treat limit 0 as 1 and larger limits as - 100; Vault and Credential lists also clamp negative limits; Turn, Subagent, - Subagent Turn and Artifact lists reject limits outside 1–100; Skill lists accept - 0–100, where 0 returns an empty page; Files accept 1–10000. Pages hold at most - 100 records (Files 10000) with accurate continuation. The local default is 20 - (Files 10000). Return the list envelope with data/has_more and first/last IDs - (null for empty pages). - Exact pinned upstream default/cap, empty-envelope and error semantics remain - unverified; do not present local limits or generic SDK parsing as full conformance. -- Session `agent_id` lookup uses the authenticated tenant. Copy the saved resource - ID and effective configuration into the immutable Session snapshot; saved metadata - does not become Session metadata. Never look up the source Agent when reading or - executing an existing Session. Omitted override fields inherit; supplied objects - and arrays replace whole fields before defaults and execution admission apply. - Reuse saved configuration validation and keep native capability restrictions at - Session admission. Reject unsupported effective options instead of dropping them; - an explicit supported replacement may make a saved configuration executable. - Inline Sessions use the same admission path. New saved-reference Sessions and - inline requests with Vault attachments or credential references record a separate - caller-intent hash: source ID (empty for inline), - supplied overrides (including field presence), environment/vaults, original metadata - and normalized initial input. Exclude response streaming and resolved source values. - Compare that same-tenant retry identity before looking up the source. A matching - retry returns the existing Session without input admission or source revalidation; - ordinary reads include current activity. Stream retries use the row's committed - event cursor and emit no created event. Recheck after source resolution failure - for a concurrently committed creator; do not hold a lock across resolution. - The unique creation upsert remains authoritative when concurrent resolutions differ. - Unrelated inline requests keep their existing resolved/default equivalences. - Recorded credential-bound retries recover before reading mutable Vault contents, - so another same-URL Credential cannot change an accepted selection. Rows with a - known creator but without recorded request intent retain resolved-hash behavior; - original overrides cannot be reconstructed, so no backfill is permitted. Source - mutation-independent retries apply only to recorded identities. Exact hosted - retry/error semantics remain separate work. -- Tenant scope must come from authenticated service identity before calling the - execution Store. Product workspace/user references in metadata grant no access. - Keep credentials and effective execution options out of Session metadata. - Store resolved, non-secret Agent/environment configuration in the Session's - immutable configuration snapshot. Inline and historical creation identities include - the resolved configuration; new saved references use the separate caller intent. - Session metadata updates replace only metadata under the authenticated tenant: - a supplied metadata field is required, null/empty clears, and a nonempty object - replaces all pairs. - Keep execution state, timestamps and the original creation request hash unchanged; - creation retries return the current resource without restoring its old metadata. - Public schema validation belongs to the API; the Store validates JSON structure. -- Session listing optionally filters by the immutable root `configuration.agent.id` - within the authenticated tenant. IDs are opaque and include inline Agents; never - require a surviving saved Agent or resolve product ownership. Apply filtering - before pagination and activity projection, using the tenant/Agent/creation index. - Omission retains unfiltered listing; a supplied empty string remains a filter. - Preserve the existing tenant-owned cursor and creation-time/ID ordering rules. - Hosted empty-filter, mismatched-filter cursor and exact error semantics remain - unverified. Other resource lists do not accept this parameter. -- `make sqlc-generate` and the drift gate cover this service. Run - `make check-core` with `OAC_TEST_DATABASE_URL` pointing to a - dedicated `oac_*_tests` database for Session integration tests. - CI provides a separate PostgreSQL service. Migration immutability and ordering - apply independently to each service directory. - -### Turns, input, schema and clients - -- Turn writes serialize on the tenant-scoped Session row. An idle message starts - a Turn; messages during queued/running/waiting work belong to that same Turn. - Store input retry identities and order durably. Cancellation retains its first - target, including an idle no-op, so retries cannot stop later work. Queued work - can cancel before dispatch; active work needs an executor outcome. Terminal - states and outcomes cannot be overwritten. Public event admission uses these primitives; live output streams remain separate. -- Input requests are ordered batches committed under the same Session lock. A - retry key identifies the complete ordered batch; changed length/order/content - conflicts and a failed transaction leaves no partial inputs or cancellation. - Existing single-event requests retain their identities at batch position zero. - Internal admission limits are 64 events and 512 KiB of payload per request; - the public API must still validate the upstream event schema. -- The external protocol reference is `openai/openai-python`'s `beta/agents`, pinned - in `contracts/agents-api/upstream.json`. Follow its Session/Turn/event semantics - and verify supported behavior using the official client. Track current coverage - in `contracts/agents-api/README.md`; SDK workflow objects are not this contract. -- Shared supported wire types live in `contracts/agents-api/v1`. `make openapi` - writes `openapi.yaml` (`/v1`), `core.openapi.yaml` (`/core/v1`) and - `runtime.openapi.yaml` (`/api/v1`); never mix their routes or authentication - schemes. Review the generated diffs; contract tests hold `openapi.yaml` to the - pinned routes and fields. -- The standalone service uses `OAC_DATABASE_URL`. PostgreSQL stores Projects, - their immutable execution scopes and API-key digests. Keys in the same Project - resolve to one shared service-account principal and tenant. Project/key writes - go through Core-key `/core/v1` routes, not configuration files. - Optional `OpenAI-Organization` and `OpenAI-Project` headers must match the key; - repeated/conflicting values fail authentication. Metadata, forwarded identities - and product session cookies grant no access. Issue/revoke operations take effect - without restarting Core. Deployment credentials cannot authenticate public calls. - Every new Session requires an explicit typed creator at the Store boundary, - including internal callers. Public creation derives it only from the authenticated - principal. Persist creator kind/ID in the creation transaction and never rewrite - them on retry, update or source mutation. The tenant remains the project partition; - do not duplicate project identifiers or create a product identity dependency. - Both early saved-reference recovery and the authoritative creation upsert require - matching creator kind/ID before returning a Session or event cursor. Different - credentials for the same principal can retry; another principal using the same - project/key receives the local idempotency conflict. This does not introduce - creator-only resource reads or mutations, or claim verified hosted retry parity. - Pre-migration Sessions retain null creator columns and remain project-readable; - creation retries cannot claim them. Missing creator and missing request intent - are distinct. Never infer historical ownership from keys, metadata or product - records. Retire older API writers before serving the creator-enforced deployment; - mixed-version writers are not supported. Tests must supply explicit synthetic - creators; only controlled historical fixtures may seed unknown ownership. - The operator-selected - `OAC_DEFAULT_HARNESS` is separate from the requested model. - Public execution supports the enabled `none` profiles, the three colocated - self-hosted profiles and qualified three-harness Docker hosted profiles; reject unsupported - input/environment/agent options explicitly. -- `packages/agents-client/v1` configures the pinned official `openai-go` Session - service. Use SDK request/response types, pagination and errors directly rather - than reimplementing transport or copying wire types. Supply an explicit service - base URL/key and creation retry key; SDK retries are disabled by default. Product - integration is a later cutover, not a side effect of constructing this client. -- `services/core/tests/official_client.py` verifies the actual server with - the pinned SDK and strict response validation. It requires a dedicated test DB - prepared by the Store tests and `OAC_TEST_SERVER_BIN`; it never starts Docker. - The same harness runs the official Go client with fresh execution tenants and - checks its created Sessions through the Python SDK. - -### Execution dispatch and Runtime capabilities - -- Execution devices are operator-provisioned in the Agents API database with - tenant ownership and a credential digest. Their internal daemon gateway uses - `/api/v1/agent-daemon/*`, separately from the official `/v1/agents/*` surface; - device credentials grant no Session API or product permissions. The optional - `OAC_PUBLIC_URL` enables that gateway. It is a single-process registry, - not a claim of multi-pod execution or stock `exec-server` interoperability. - Self-hosted enrollment uses this gateway with an exact Environment binding. - Session/device bindings are tenant-scoped and immutable. Revocation denies new - connections and binding reads; an existing connection closes on its next - heartbeat. Connectivity comes from the live registry, not a persisted online - flag. `last_seen_at` is diagnostic only. Product gateway behavior is unchanged. -- `services/core/internal/execution` dispatches internally resolved Turns - through that gateway. Claim `queued` to `in_progress` before subscribing/sending; - never automatically replay a claimed or interrupted Turn. Ordered extra inputs - require native steering receipts. Commit terminal outcome and native Session ID - together under the admission lock; unapplied messages prevent successful completion. - Resolve credentials separately from the immutable non-secret snapshot. -- Internal execution requires advertised durable Turns, strict resume, preparation, - applied input receipts and the qualified operation capabilities. Reject - unadvertised peers before claiming; failed strict resume cannot start unrelated - history. A completed Turn closes steering admission, settles existing native - operations and receipt sends, and closes its output before the retained Executor - can accept another Turn. Normal completion does not cancel the Executor. - Cached input identities and conflicts remain readable while completing; - queue/write success is not consumption. The receipt worker stays busy through - its send, and router shutdown cancels its native and transport waits. - The per-input `durable_receipt` opt-in requires a phased adapter. Its ten-second - transport timer stops only after a complete native write; a separate `written` - acknowledgement stops the API's thirty-second delivery timer. Neither phase - advances the input cursor. Await final native acceptance/consumption under the - Turn lifetime without automatic redelivery. Receipt sends retain a separate - five-second shutdown-aware context, and Done retains a fifteen-second final - settlement bound. Once cancellation is sent, its receipt owns the terminal - outcome even if an input becomes unknown first. Calls without the opt-in retain - their existing response deadlines. Native history still requires the device's - persisted engine files; IDs alone cannot restore deleted history. Cancellation - receipts carry the stopped Turn's confirmed continuity snapshot when no Done is - emitted. Preserve separately reported usage on failure; do not add the same - counters again when Done also includes them. - Usage frames carry cumulative snapshots for the current execution, not deltas. - Adapters publish observed snapshots promptly through the same ordered stream; - waiting for Done unnecessarily loses known measurements if the Runtime stops. - Core replaces complete valid token breakdowns and preserves the last committed - measurement on interruption. Missing measurements remain unknown. Do not infer - token consumption from context occupancy or estimated costs, or parse native - Raw payloads in Core. Public Session totals cover recorded root Turns and are - null while any root Turn has not ended or once one ends with unknown usage; - Runtime telemetry uses the separate measured sum of recorded snapshots. Subagent - Turn listings are not a summable accounting ledger. A cumulative native total that has not advanced - past the Turn's baseline is not a measurement of that Turn. Native measurement coverage - and exact provider/model attribution remain explicit qualification boundaries. - No separate public usage event or historical SSE replay is introduced. -- The dispatcher is an internal entry point used by the standalone service worker. - Legacy internal `daemon` configuration is not a public Environment type. - Self-hosted enrollment uses exact local binding; further pending interactions - remain separate slices. Unexpected interaction requests fail explicitly until supported. -- `environment_none` advertises an adapter's explicit environment-disable - path. Execution snapshots with public `environment.type=none` require that - capability and set `disable_execution_environment` on the internal prompt. - For Codex, the daemon forces `CODEX_EXEC_SERVER_URL=none` after caller environment options - and confirms native `local` and `remote` environments are unknown before starting - or resuming a thread. Unsupported binaries fail closed. The bound device hosts - the engine process; it is not a user execution environment. This is not an OS - isolation guarantee, and engine state still lives on that host. Ordinary product - requests retain their existing environment. The public worker selects an authenticated same-tenant engine host for this mode. -- `execution_controls` advertises the typed search/verbosity block on the daemon - prompt. Agents API requires it in addition to the selected engine's required capabilities - before binding/claiming work. Older peers with only option-based capabilities - must not receive controls they would ignore. The API sends resolved search and - text verbosity values; native option names belong to adapters. Codex translates - them using its existing validation/catalog path after cloning adapter options, - so explicit controls take precedence without mutating those options. Omitting - the entire block preserves ordinary product behavior; a supplied block requires - both valid fields. This internal contract does not add public configuration or - engine support. Future native adapters must verify the same semantics before - advertising the capability. -- Explicit public `programmatic_tool_calling.enabled=false` uses the common - `ExecutionControls.DisableProgrammaticToolCalling` field and the operation-specific - `programmatic_tool_calling_disable` capability. Both public qualification and - Runtime support are required for that request; omission creates no prerequisite. - New and resumed executions retain the frozen setting. Codex disables native - code-mode features and checks managed requirements before starting/resuming a - thread, rejecting a conflicting requirement. Claude and MiniMax retain their - restricted native inventories, which exclude programmatic execution. This does - not remove unrelated native utilities or claim enabled programmatic support. - Explicit `web_search.mode=disabled` reuses the existing disabled search control. - Search remains off when omitted. Optional search settings are resource data and - do not cause execution while disabled. Saved Agents keep every pinned search mode; - enabled search remains unqualified and rejects at Session admission. -- `web_search_control` advertises the Codex adapter's explicit `web_search` option - (`disabled`, `cached`, or `live`). Agents API requires this capability before Codex dispatch; - the typed execution controls force search off on new and resumed Turns. Native configuration translation stays in the - adapter. Product requests that omit the option inherit their existing defaults. - This is tool selection, not a network isolation guarantee. -- Inline Agent `text.verbosity` accepts `low`, `medium` and `high`; omitted or - null values resolve to `medium` in the immutable configuration snapshot. The - Codex dispatcher requires `text_verbosity` support and sends the effective value in - typed execution controls through the Codex adapter for both new and resumed Turns. The adapter queries - the native active catalog with `codex debug models`, checks model support and - pins that catalog snapshot for execution. The probe requires Unix process-group - cancellation; other daemon hosts do not advertise this capability. For models - without declared verbosity support, including the native unknown-model fallback, - `medium` selects native default text generation by omitting the override. The - pinned protocol defines `medium` as the default text amount. Supported models - still receive explicit `medium`, even when their catalog default differs. - Unsupported `low`/`high` and unreadable catalogs fail before model execution; - unsupported non-default levels remain an explicit implementation gap. - Product requests that omit the native option retain their existing defaults. - Structured output has a separately qualified profile described in - [Structured output execution](../../contracts/agents-api/execution-tools.md#structured-output). -- `subagent_control` advertises native subagent tool control. Agents API requires - it when resolved `multi_agent.enabled` is false and sends the typed internal - `disable_subagents` policy on both new and resumed Turns. Native translation - stays in the adapter: Codex disables both multi-agent feature generations, - overriding operator feature preferences. Product prompts that omit the policy - retain their defaults. Enabled multi-agent observations require separate operation qualification; the Agent tools list is not - proven to enumerate every harness-internal utility. -- Subagent resources use the common observations in - `internal/agentdaemon/proto/subagents.go`: verified identity, successful lifecycle - effects, native-owned Turns/Items and neutral coordination operations. Core - assigns public IDs and projects them under the existing Session lock and leased - execution journal. Native names, history parsing and outcome proof stay in - adapters. Public GETs read persisted resources without starting native work. - Child Turns have a native writer and a separate table from the Core queue. - Session Turn reads and the Session event stream carry root work only: read child - Turns and Items through the Subagent routes, and never publish child Turn or Item - events on the Session stream. A child Turn's `agent_id` is the Session's Agent - ID; `subagent_id` names the child. The migration-defined `public_execution_turns` - view has no public reader; do not reintroduce mixed Session Turn pages. - Session Items stay root-owned; copied parent transcripts never become child work. - Repeated effects are idempotent. Active includes idle; task completion, process - release and cancellation cannot fabricate public closure. Native timestamps - retain their actual precision and unknown Usage stays null. - Reuse the existing native owner for child settlement and cancellation, freeze - root output first, and deliver child Items before their terminal Turn snapshot. - Do not add another scheduler or a broad recovery framework. Capability - advertisements do not qualify unsupported native facts. The exact read contract, - admission limits and remaining evidence are in [Subagents](../../contracts/agents-api/subagents.md). -- `function_tools` advertises the optional native function-call bridge. Explicit - prompt definitions become Codex dynamic tools; unchanged prompts carry none. - Requests and ordered text/image results are scoped by Run and native call ID. - Normalize string results into one text part at the public execution boundary; - the internal result carries a typed content array, and adapters translate it - to native content without fetching images or dropping parts. Validate content - before consuming a pending call. Retry identity includes the complete ordered - content and success flag. Reuse - application receipts and conflict detection; a receipt confirms the native - reply was written, not that an external side effect succeeded. Pending calls - end with their Run; the execution service owns persistence and recovery, while - the product client retains business approval and credential-owner authorization. Do not - map native approval requests to invented official protocol resources. -- Function-call storage is scoped by authenticated tenant, Session and Turn, with - immutable public/executor call identities and arguments. Result admission and - application receipts serialize on the same Session lock as cancellation and - terminal transitions. Store the complete caller-validated result object; wire - validation and native translation belong to their API and execution boundaries. - Identical retries return the saved decision; changed results conflict. Pending - reads exclude applied calls and cancelling/terminal Turns while history remains - readable. Persistence does not imply transparent native-process recovery. - Recording a call moves the Turn to `waiting`; the last application receipt - resumes it. Session reads use one database snapshot for Turn, actions and usage. - Session state events retain their action snapshot, without private executor IDs - or results. Cancelling/terminal Turns expose no actionable calls. A waiting Turn - can fail or cancel before a result arrives; successful execution requires resume. - Actions remain visible until native application is acknowledged. This timing is - an implementation choice, not verified upstream event sequencing. -- Function results can join internal message/cancel input batches. Their explicit - Turn/call identity selects an existing call; admission never creates a Turn for - a result. Save the complete result and its input retry record in the same Session - transaction. Any invalid target, conflicting result or later batch error rolls - back the whole request. Resolve targets only after the tenant Session lookup: - an unknown call or a call of another Turn is 400 `invalid_request_error`, and - missing or foreign Sessions keep one 404. Identical saved results remain retryable after termination - without applying them again. The execution input cursor skips function results; - their separate native receipts still determine application. Public result events - validate variant-specific fields and required values before admission; store - omitted versus null error/output and ordered text/image parts. Inline function - tools resolve into the immutable configuration with explicit - `defer_loading=false`. Validate required strings and parameter objects before - persistence; reject unsupported deferred discovery. Omitted/null/empty tool - lists resolve to no tools. The public worker selects or waits for a same-tenant - device advertising `function_tools` when the Session has functions. - Function results are Session input Items: emit `item.added` with a null output - index, and never emit `item.done`, whose upstream union only allows agent output. - Project their public output/error from the saved submission; the wire always - carries both, null when not submitted, while stored payloads keep the submitted - presence. Native content normalization must not change public history. -- Codex function application requires a matching live native dynamic-tool completion, - including root thread/Turn/call identity, function, success and ordered content. - Writing its JSON-RPC response is not application. The adapter owns pending - receipts without holding their state lock across IO or waiting; terminal state, - cancellation and native loss settle unconfirmed submissions before release. - Uncertain receipt timeout ends that native execution without resending the result. - The common Runtime interface and router continue to own delivery identity, - retry/conflict and terminal ordering; Core never parses native tool events. - Record native confirmation before potentially blocking observation publication; - output backpressure cannot turn a known application into an unknown outcome. - The function submission owns raw write completion and receipt failure together; - its writer never applies an independent timeout/close decision. A confirmed - receipt releases submission even if writer completion has not yet been scheduled. -- Internal function execution requires an advertised `function_tools` capability - before claiming a Turn. Translate resolved definitions in the execution adapter, - persist declared callbacks before exposing actions, and deliver each saved result - once per live dispatch. Keep its success flag and ordered text/image output; - append a non-null error as a final text part because the native result has no - separate error field. Retain the original complete result in storage. Do not - treat transport delivery as application or automatically replay an uncertain - result. The adapter waits for outstanding application receipts even when Done - arrives first. Waiting Turns still accept execution observations and cancellation. - The public function workflow is verified with the pinned SDK and a real daemon - and Codex process against a synthetic model endpoint. This does not verify - other tool types, deferred discovery or upstream service timing. -- `message_items` advertises native assistant-message observations. Agents API - opts in with `observe_messages` only for advertised peers; ordinary product - requests retain their existing frame sequence. Opted-in text deltas carry their - native item ID, and `output_message` records start/completion, phase and the - completion text snapshot. A snapshot is not another delta; uncompleted messages - remain partial when their Turn ends. Keep these observations in the journal - before projecting public Items. This does not promise daemon event replay. -- `tool_observations` advertises engine-neutral tool snapshots. The opt-in - `observe_tool_observations` attaches the typed `observation` to tool-call frames. - Native adapters own discriminator/status/action translation and preserve raw - structured values; reuse the shared function-result content type. Kinds are - `command`, `mcp`, `function` and `web_search`; observation status is - `in_progress`, `completed`, `failed` or `incomplete`. A present empty function - content array remains distinct from missing content. These are - execution facts, not public Items: the API owns public IDs, schema projection, - lifecycle events and persistence. Product requests that omit the opt-in keep - their frame sequence and fields. Agents API requires this capability before - claiming work and always requests neutral observations. Its Item projector - validates this shared contract and never decodes engine-native tool snapshots. -- Codex callers opting into neutral tool observations also receive `command_output` - fragments with the existing native command identity. The adapter filters the root - Thread/Turn; the service requires an already indexed command in the same Turn. - Journal and Item updates commit with `agent.output.command_execution_output.delta` - events, retaining original fragments and the command's stable output index. - Completion output replaces accumulated drafts; absent completion output retains - observed text. Terminal Items ignore late fragments, and cancellation preserves - partial output without inventing successful command completion. Native text - conversion and output quotas still apply; this is not a byte-complete stdout/stderr - guarantee. Pinned native 0.153.4 also has an early-output subscription window; - missing native notifications/aggregate bytes remain a separate execution gap, - not output to reconstruct from model tool-result prose. Older peers may supply - only completion snapshots. Product requests - without the observation opt-in retain their existing frames. - -### Observations and streaming - -- Execution observations are written to tenant-scoped `turn_events` in ordered, - idempotent batches before they can back recovery or publication. Keep daemon - payloads intact; this internal journal is not the public SSE protocol. Flush at - least every 100 ms while consuming events and before terminal persistence; - uncommitted observations can be lost on a hard process crash. Terminal outcome, - journal entry and native continuity commit together. Preserve partial text on - cancellation, including frames queued before a separate cancellation receipt. - Do not infer successful completion after a persistence error or stream overflow. -- Agents API uses the gateway's durable subscription; overflow or disconnection - closes it with an explicit error. Product subscriptions retain their existing - best-effort behavior. Journal limits are 512 KiB per payload, 1 MiB per batch, - 65,536 observations and 32 MiB per Turn; terminal persistence reserves one - additional outcome entry. These are internal admission limits, not promises - about upstream API limits or durable daemon-to-service replay. - -- Live Session SSE reads execution-owned `session_events`, committed with the - corresponding input, Item or lifecycle transition under the Session lock. - Store immutable transition snapshots; never render an old event from a later - Turn state. Reuse the API's response mapping and keep internal snapshots out of - wire payloads. Historical index rebuilding emits no live events. -- The notification buffer retains at most 256 events and 64 MiB per Session - after each transaction, retaining a single oversized event if necessary. - Read batches are bounded to 32 events / 1 MiB, with the same single-event - exception. This buffer is not a public replay log: GET begins at the committed - high-water mark, ignores Last-Event-ID, and polls committed events every 100 ms. - Missing sequence positions produce a safe stream error and close; recover via - Session/Turn/Items queries. Socket writes have a five-second deadline and hold - no database connection. Client disconnect releases the handler; comments keep - idle connections alive. GET SSE does not close merely because one Turn finishes; - only creation responses end on settlement (below). - -- Session creation with `stream=true` reuses atomic input admission and the live - event loop. The upsert returns its cursor under the Session lock, before initial - inputs; never replace it with a post-commit cursor lookup. A new response emits - one request-local `agent.session.created` with the committed Session projection - that the JSON 201 response returns (read after the commit), then committed - changes from that cursor exactly once. A fresh creation stream ends right - after the first `agent.session.idle` recorded when a Turn ends or an input - reservation stops being pending (expired, cancelled or failed), or any - `agent.session.failed`, and never sends the events after it. A self-hosted - connection clearing pending input to idle, `requires_action`, function results - and resumed work keep it open. A creation that admitted nothing (no Turn or - reservation) ends right after `created`. Settlements that record no event use - a fallback: after an empty drain the stream reads the JSON-path projection and - the event cursor in one database snapshot and, if the Session is idle or failed - with no queued, running or waiting Turn and no pending reservation, sends only - events up to that cursor, then ends. Accepted follow-ups: another client's work - drained before that read can still be sent, and idles recorded by an older - binary during a rolling deploy carry no settled marker and rely on the - fallback. An input reservation made while the ending Turn captured Artifacts - can start a later Turn that the stream does not follow. The settled marker and - pending-input flag are Store-internal, never wire fields, and add no events. - Re-read the projection after a sent Session status event and otherwise at most - once a second. The local creation retry key excludes response mode; a same-key - `stream=true` retry of an existing creation returns 201 with only the - connection comment and ends at once, admitting nothing and following no work, - because official same-key requests create distinct Sessions. Retry the same - request/key with `stream=false`, or use the GET events stream, to recover. GET - event streams keep their live-only start and never end on settlement or a Turn - failure; the only server-side end is the terminal `agent.session.failed` of a - hosted provisioning failure (and Session deletion), since that Session can - never run again. - Disconnect never cancels admitted work. Official observations cover `none` - creation; self-hosted, hosted and no-input stream lifetimes and the retry - behavior are local choices, and the separate SDK one-Turn helper does not - define this endpoint. Do - not present local retry behavior as replay. A new Turn records `turn.created`, - its user input Items, then Session activity in one transaction. Terminal Turn - events carry top-level `usage` copied from their Turn snapshot, null when - unknown; never derive or sum it. - -### Public Turns, Items and events - -- Public Turn retrieve/list project persisted execution state and the immutable - Session Agent identity. Scope both resources and pagination cursors to the - authenticated tenant and Session, ordering by creation time then ID. Do not - expose adapter outcomes, native IDs or raw errors. Failure uses a customer-safe - category; usage is nullable when a complete upstream breakdown is unavailable. Submission uses the separate Session events endpoint. - - -- Public Items list reads a persisted execution-owned projection, updated in the - same Session transaction as admitted messages and journal batches. IDs derive - from the Turn and source identity; the first-observation timestamp and stable - tie breakers never change when content or status changes. Allocate each new - Item's Session position under the Session lock, preserving observation order - for equal timestamps. Allocate a separate zero-based `output_index` per Turn; - inputs do not consume output indexes. Updates and retries retain both values. - Persisted indexed history keeps its deterministic order; missing original - ordering cannot be reconstructed. Cursors are scoped to the authenticated Session. - Terminal Turns expose unfinished Items as `incomplete`, preserving completed - message/tool states independently of the Turn outcome. -- Public history reads use the persisted index; the private pre-Items journal - backfill is retired. Migration 15 and its historical validation remain evidence, - not a supported upgrade procedure. An old installation is unsupported and must - be retained separately from a fresh installation. Never mark unprepared history - indexed by hand or replay engine execution to convert it. -- Project only the declared public Item variants; native adapter metadata is not - a response schema. Preserve structured tool JSON without float conversion. - Completion text replaces accumulated deltas. Item merging must not mutate the - incoming observation or the previous snapshot: public text delta events read the - original fragment after merging, while Items retain the accumulated text. - Copy the content slice before replacing its text pointer. Assistant text - Items follow the official event sequence: `item.added` in progress with empty - content, an empty `content_part.added`, deltas, then the done events. A first - observation without its own fragment (a non-streamed native final) carries its - unchanged text in one delta; this frames the text and never alters it. Wire-only - explicit nulls (`phase`, function result `output`/`error`, Item event - `output_index`, Agent `reasoning` keys) come from response marshalling. Stored - Item payloads keep their original encoding through `Item.MarshalStored`, so - replayed child Items still compare equal, and stored configuration keeps - omitting unset reasoning keys. Keep partial output on termination; - do not turn an unfinished call into a successful result. Thinking fragments are - internal observations, not a claim of upstream reasoning-item support. - -- Public `POST /v1/agents/sessions/{session_id}/events` accepts ordered text-message - and cancellation batches through the pinned official client. Preserve individual - input messages in the Item index while deriving text for native dispatch. Batch - idempotency and cancellation targets remain durable; unsupported variants fail - before admission. Session creation accepts initial text as a string - or user-message array through the same parser and admission path. Commit the - Session, initial input, first Turn and Item/event projections in one transaction. - A creation retry returns the existing Session without re-admitting initial work, - including after terminal or later Turns. Omitted/null input is permitted only - for non-streaming hosted creation and self-hosted creation. - Creation streaming uses the shared live path above. Image support requires the - qualification in [the message-input contract](../../contracts/agents-api/message-content.md#images). - -### Worker ownership - -- Enabling daemon transport with `OAC_PUBLIC_URL` also starts a bounded execution worker. Select - only connected, capable devices owned by the authenticated tenant; bind once and - preserve native continuity. Metadata cannot select a device. Offline work stays - queued and can be cancelled. An engine host is not a self-hosted environment. -- One worker service owns an execution database through a dedicated PostgreSQL - advisory-lock connection. Its execution Store view uses that same connection for - every Session transaction: binding, claim/reconciliation, journal/Items/Usage, - function callbacks/application receipts and terminal/native continuity. Serialize - these short transactions and lease pings; execution transactions have a five-second - deadline including gate and Session-lock waits. Sandbox reset snapshots bound the - deployment relation explicitly to its legal singleton row before resource joins. - Keep this cardinality visible even on fresh databases without statistics: inflated - join estimates can trigger expensive JIT compilation inside the lease deadline. - Never hold a transaction across daemon/model work, reconnect the writer or fall - back to the pool after lease loss. - The original Store handles public admission and device/auth maintenance on pooled - connections. Execution reads may also use the pool; a read grants no write authority. - At startup, reconcile previously claimed work as failed, preserve queued inputs - and never replay uncertain execution. Shutdown cancels active dispatch and attempts - terminal persistence before releasing the lease; a lost owner cannot commit it. - Lease Close invalidates its writer and waits for pgx connection cleanup within - the caller deadline. A later Close can resume that wait after a timeout. This - drains client resources; it does not acknowledge remote advisory-lock release. - Tests that immediately transfer ownership must observe the previous owner's - exact database advisory lock disappearing before starting its successor. Bound - that wait and fail on query errors; do not retry Worker startup to mask competing - owners or change production lease behavior for a test's timing assumption. - Worker shutdown retains its existing bounded best-effort close policy. - This fences database writes, not already queued daemon commands or native effects. - Native quiescence/reconnect and recovery of unreported outcomes remain separate - gaps; this is not distributed exactly-once side-effect execution. -- Session state and last activity derive from its latest persisted Turn. Queued or - active work is `in_progress`, successful/cancelled work is `idle`, and failures - use a safe public error. The worker does not replace product dispatch, business - authorization, or the separate approval/environment lifecycle work. - -## Hosted sandbox nodes and optional suspension - -Default installation includes Core, Web and PostgreSQL but no execution node. -It always creates the Core key. Core receives only its digest; the paired console -server receives the private key, uses it for sign-in and injects it only on -`/core/v1` requests after console login and same-origin checks. The browser -never receives that key. Node and daemon connections use `/api/v1` -with their own credentials; the reverse proxy sends them directly to Core, never -through Web. Zero-node Core receives neither the Docker socket nor KVM. -The Web and deployment administrator API select one provider, per-sandbox -resources and immutable Runtime release; the Core address comes from the -installation public URL. PostgreSQL owns this complete, -generation-tagged selection under the existing execution lease and deployment lock. -The shared `sandbox.DeploymentSpec` defines required CPU/memory and supported disk -limits plus Runtime provenance; neither a node file nor the installer owns another -selection. Request `resources` describes limits; response `specification.resources` -contains those limits, while response `resources` counts retained allocations and -pending hosted Environments. Keep these meanings distinct in clients and UI. -See the [deployment contract](../../contracts/agents-api/sandbox-deployment.md). - -Docker and E2B accept CPU/memory but reject independent nonzero disk capacities; -do not claim hard root/workspace disk quotas for them. Docker creation and native -inspection enforce the declared CPU/memory and exact image. E2B setup verifies the -exact ready template build and matching CPU/memory through the pinned SDK before -saving its encrypted account key, and records the build as read for the safe view; -an omitted E2B `resources` adopts that build's CPU and memory. Loading a committed selection reconstructs its -provider from its owned generation, current committed credential and receipts without repeating candidate -template validation; a template endpoint outage must not block cleanup of existing -sandboxes. Creation and instance inspection still enforce the saved resources. -E2B uses direct placement without a node. Node -providers require one immutable distribution with source commit, Docker image ID, -OCI manifest digest, microsandbox image reference, Runtime and firmware hashes. -These identities are distinct and cannot substitute for each other. - -Startup claims the stable installation identity and a new owner epoch before -provider selection. The runtime manager retains generation-aware provider facades. -Initial setup and replacement prepare and validate candidates before database -writes. Rejected candidates preserve the active configuration and workers. A node -replacement uses the existing mutation gate, pauses manager admission, drains old -calls and loops, then repeats the resource/generation guards in the commit -transaction. A changed selection, its generation and retirement of old nodes and -unused enrollment tokens commit together. Publish the prevalidated configuration -and shared observation/bootstrap cache under the manager mutex without further -external work or a fallible activation step. Request cancellation after commit -cannot discard that publication. Interrupted drains remain barriers for retries -and resume. Keep the Worker and runtime manager as single owners; provider I/O and -draining hold no database transaction or manager map mutex. - -A locally unavailable provider dependency keeps hosted admission closed while the -existing scan waits for repair; administrator recovery remains available, including -on restart. Database and ownership errors remain failures. Unconfigured hosted -admission creates no Session state. Derive Runtime bootstrap and daemon WebSocket -addresses from the validated installation public URL, never inbound Host headers. Read the -current selection from the live deployment API; there is no startup configuration -read. - -All sandbox writes require the observed generation, including initial POST at -zero; reject stale state before provider preparation, reset or no-op checks, and -repeat it under the committing row lock. Same-provider PUT advances the target -without draining execution or retiring nodes, tokens or the owner epoch. Node -providers prepare independently and keep their old qualified serving pin. A different backend or E2B team requires reset. Historical selections without a -valid specification are unsupported and rejected at startup; keep their original -Core responsible for retained resources and install the current release separately. -Preserve historical allocation ownership and placement, never migrate a Session. - -Persist immutable allocation and placement deployment generations, distinct from -compute generations. Node allocations copy their placement; E2B binds at reservation. -Keep superseded specification/build rows without credentials while current, owned by -an unreleased allocation/placement, or pinned by any nonremoved node. A durable node -serving pin survives offline state and zero resources. GC uses the deployment lock -and bounded pages; reset clears pins only after release. Never downgrade facts needed -for generation routing. Bind node readiness to exact generation, current connection -and owner epoch. Promote a durable pin only for readiness of the then-current target -under deployment serialization. Late superseded readiness cannot acquire a pin. -Filter online/readiness/address/capacity before preferring the newest eligible pin; -newest-full must not mask older-free. Route Create, restore and cleanup through the -immutable allocation/placement generation. Fixed-configuration manual nodes serve -only their enrolled generation; generation management is an explicit hello capability -on the same current wire protocol, never a historical-version fallback. - -Use sparse generation control batches of at most eight entries with no lifetime cap. Omitted -facts never authorize deletion. Correlate whole retention grants to connection, -epoch, sequence, generation and digest; recheck queued/inflight/helper references. -Permanent generation flock files survive updates and GC. Before any helper starts, -the installer durably binds the original inode to installation/generation/specification -identity; Python and Go openers verify it after every restart and never adopt a -replacement or missing record. Preparation plans stay distinct from write-once final -provider configurations: pending-only entries may recover or collect under Core -authority, never serve. Resolve Docker's supported local image ID before publication. -Only transfer/checksum/provenance errors report runtime_download_failed; preserve -other fixed causes without parsing native error text. Retained Runtime bytes and -the console's exact-release HTTP allowlist must agree. See -`contracts/agents-api/node-generation-protocol.md` for recovery, immutable artifacts -and conservative retention of unproven historical helper ownership. -Interrupted node collection keeps an exact private generation journal and immutable -configuration. Restart treats it only as a candidate for a fresh correlated Core -drop grant, never as preparation or serving readiness. Persist native cleanup -before removing its executable and durable file cleanup before the dropped marker; -CLI errors and unknown native ownership cannot establish absence. Docker daemon -images are shared host content, so automatic node GC retains them; only the host -administrator can establish whole-host authority to remove them. Microsandbox -image GC remains scoped to the installation-private store. Fresh nodes persist -verified original Runtime file ownership separately from the host program; collect -only exact unreferenced Runtime paths, keeping program, identity, base configuration -and manifests readable across interrupted cleanup and restart. - -Use one E2B classifier. Omitted key preserves the current key; identical selection -with omitted key is a no-op. Explicit key submission, including identical bytes, -verifies and increments generation. Check the candidate and retained builds, paginated -team-owned template membership and every settled live receipt under the candidate key. -Initial E2B selection must also belong to the submitted key's team. Before online -replacement, verify the current exact template under the committed key, then under -the candidate key: shared membership is the ownership anchor, including for legacy -installations with no retained resources. The pinned SDK exposes team-owned template -listing, not an authenticated team ID; never invent or persist an inferred team ID. -Public readability cannot anchor ownership. A legacy selection outside the committed -key's team, or a revoked committed key, requires reset without blaming the candidate -key. Transport uncertainty remains unconfirmed. Keep the old key valid until the PUT -returns 200; revoke it only afterward. Missing/unsettled receipts and -unconfirmed reads reject. Fence credential commits against all provider calls and -actual helper subprocess completion after cancellation, then reverify. Helper exit -never settles unknown remote Create. Bounded failure keeps the original key and -lifecycles. Resolve each allocation's immutable specification and current key in one -snapshot; no stale-key cache, fallback to current spec, or provider-map unloading. -Audit the committed change without secret fields. - -Project rollout and reset from the same database snapshot. Old retained resources -alone never imply preparation or high-frequency polling. Rollout state is settled -unless actual target preparation is active. Offline/unconfirmed nodes are unknown; -only exact current connection/epoch/generation readiness is ready. A durable pin is -not connectivity. Fixed diagnostics alone may explain failed preparation. Report -preparing only from an actual target observation; an online fixed-configuration node can be -update_required while its valid pin remains available for admission. - -Reset is durable execution state, advanced by the existing manager outside its -counted work. Serialize start, escalation, cancel, setup, update and finalization -through the mutation gate. Auto waits only for in-progress/waiting root/subagent -Turns and pending file writes; queued work, idle and suspended Sessions can archive. -Recheck idleness with Session then deployment locks; never invert that order during -finalization. Keyset-page bounded work, bind candidates to the reset request time -as well as generation, and persist the original absolute deadline and validated -audit provenance. Cancellation cannot revive archived work. Auto's deadline -escalates to force durably; self-hosted Sessions are excluded. - -Use one snapshot and timestamp for the held-resource partition: unreleased -allocations plus pending hosted Environments without any allocation row. Cleanup -has precedence, then busy/idle. Count deleted/expired receipts until release. -Offline ownership comes from allocation or active placement node identity, with -the existing 45-second connection/owner-epoch predicate; provider readiness is -independent. Project every offline node and make its sum equal on_offline_nodes. -No cleanup failure, offline state or empty read authorizes a synthetic release. - -At zero held resources, drain outside database transactions, recheck under the -deployment lock, and atomically clear all provider policy, E2B credential/build -metadata and specification, retire nodes/tokens, advance generation/owner epoch -and audit completion. Publish a generation-bearing nil provider tombstone without -fallible work after commit, preventing delayed loads from reviving the old provider. -On a failed final write or interrupted drain, restore the committed provider using -a bounded owner context before releasing the mutation gate; recovery failure keeps -admission fenced and stops the owner. Ordinary migration resumes old Web-managed -maintenance with reset null. Refuse downgrade during reset or for an unconfigured -completed reset with generation above zero; never discard the generation. - -Explicit administrator Session archive requires a Web-managed deployment and its -current generation, without a reset precondition. Keep Project scope and hosted -eligibility checks in its Session-first transaction, alongside Environment expiry, -cancellation, Runtime authority revocation and audit. Background reset reconstructs -the actual Project scope from trusted records while retaining requester provenance. -The ordinary provider lifecycle owns compute/snapshot release. Retain history and -persisted Files/Artifacts; archived Sessions cannot resume unpersisted workspace. -The archive GET reports resource disposition, not provenance or Turn finalization. - -Archive cancellation must not destroy a healthy receipt path before its terminal -commit. Only the archive that first revokes a device may persist its exact -`archive_cancel_turn_id`; ordinary revocation clears it, and repeated cleanup -preserves rather than recreates it. The existing authenticated delivery can drain -its cancellation for at most 20 seconds from the Turn's original -`cancel_requested_at`. Track that delivery through Done, cancellation ACK and -terminal commit, independently of subscription removal. Heartbeats and both normal -and checkpoint cleanup use the same identity/deadline check; no new connection, -input, file/MCP authority or lease renewal is granted. Do not hold a transaction or -lifecycle gate waiting for the receipt. Lost peers, expiry and restart retain the -ordinary failure/cleanup fallback, never a fabricated cancelled outcome. Preserve -unknown Create ownership and reject schema downgrade with unsettled markers. - -There is no file-managed startup path or embedded local node. An older file-managed database is not automatically adopted after its environment variable is removed. Settle and drain that deployment with its previous release and original backend, preserving business data, private receipts, identities and storage. The legacy file-managed path has no automatic adoption or force conversion. The supported current path is a database-managed deployment. Harness selection and public/self-hosted contracts remain unchanged. - -The paired console serves only matched, non-secret distribution artifacts for node -installation. Never serve private installation files or arbitrary paths. Installation -reads `GET /api/v1/sandbox-node/configuration` using an unconsumed enrollment token, -or a retained node credential with `X-OAC-Node-ID`. Reads never consume enrollment; -registered nodes can read their matching configuration during reset. Validate -installation, generation, specification digest and release before writing node files, -registering or reconnecting. Reject drift rather than overwriting retained identity -or using local resource defaults. Registration consumes a token only after these -checks. The node installer requires root or sudo, verifies downloaded files and -starts a root-owned system service for the dedicated `oac-node` user it prepares. -It performs no SSH installation, Session creation or model call. Ordinary-user -installation and removal are rejected before reading credentials or mutating state. -Internal generation preparation and collection still run as the service account. -Web exposes one root/sudo command; do not retain a user-service alternative. This -boundary is specific to Sandbox Provider nodes, not native self-hosted daemons. - -Node management (Web's **Nodes** page; see the -[nodes guide](../../docs/getting-started/nodes.md)) is a -deployment-level admin surface, separate from Project credentials. The paired -console's Core key stays on its server. Enrollment credentials authorize -initial node configuration reads and registration; durable node credentials authorize -retained configuration reads and node transport. Project keys cannot read nodes, -placement or allocations. - -One execution owner manages every node through the same finite Provider protocol. -The Core host joins through the same standalone node installer and service as any -other host; every node needs a guest-reachable, non-loopback HTTPS origin. -All managed nodes actively connect over authenticated TLS. Persist private node -identity and highest owner epoch; refuse another process using the same identity -or a changed backend namespace. Reserve each NodeID before transport upgrade and -retain that reservation through disconnect cleanup; a duplicate connection must -not replace a live or opening connection. Keep one private state directory per -node and never copy its identity to another host. This is connection exclusion, -not host attestation. The Hub's global mutex protects only in-memory connection -state. Authentication, ownership and Store callbacks run synchronously outside -that mutex, respect cancellation and have a five-second limit; never detach -database writes. Closing the Hub cancels opening and live connections without -waiting for database callbacks. Keep each node reservation until its fenced -disconnect cleanup finishes. Register database presence in an explicit transaction: -a canceled statement must not later publish presence through autocommit. Disconnect -cleanup first locks the node row by identity, then applies the connection/epoch -fence with a fresh READ COMMITTED statement so an in-flight commit cannot be -missed. These transactions must not acquire the deployment-wide manager lock. -Heartbeats establish provider readiness and -last-observed host metrics, never Session activity. An unready provider reports -one fixed diagnostic code, classified by typed probe errors where the node detects -the cause; probe text and host paths stay on the node, and Core stores unknown -codes as `provider_unavailable`. Keep that set closed. Transport reconnects use -bounded backoff. Send relative operation budgets, anchored to the node clock at -receipt and consumed while queued; clocks on different hosts need not agree. Core -still bounds its own response wait. Do not replay mutations after a timeout or -lost response. Retain allocation -and checkpoint operation receipts and observe the original operation instead. -Disconnects and read timeouts are unavailable/uncertain, never resource absence. -Node transport preserves an exact-reference, explicit `CreateSettled` receipt -alongside its original provider error. Settlement authorizes eventual allocation -release, not execution. A confirmed native Create rejected by the subsequent -read-only configuration check, before bootstrap starts, can return that receipt. -Do not infer settlement from timeout, missing compute or successful Kill. Strict -configuration rejection must not prevent already-authorized cleanup: Kill still -checks ownership independently, and release still requires creation settlement. -Runtime resource observation uses the same immutable node placement through one -bounded read-only Provider operation. Preserve main's Runtime observation/history -service and authorization boundaries. The node delegates only to a provider-owned -observation source; absent capability or transport returns unavailable, never a -Core-local fallback. Observation must not create, renew, restore, or touch Session -activity. Preserve the durable compute receipt and provider timestamps; existing -observation clock validation can reject skewed samples without changing idle policy. -Online-state writes compare the handshake epoch atomically in PostgreSQL so a -stale Core cannot publish readiness for a new owner. Node-managed allocations -do not expire merely because the internal observation keepalive is an hour old; -explicit deletion and configured snapshot retention still authorize cleanup. -Persist only bounded, sanitized observation codes for offline, missing or -unconfirmed resources; keep these separate from the lifecycle and do not invent -a successful running observation after a host restart. - -Managed lifecycle state is owned by one serial worker per registered node: -its gate, allocation and pending cursors, connections and -wake hints are not shared with other nodes. A thin coordinator discovers nodes -and owns worker shutdown; it never holds its map mutex during database, provider -or wait operations. Each worker advances independently, including when another -node is online but its provider is stuck. Do not add a shared scan barrier or -global provider pool: lifecycle concurrency is at most one operation per node, -and grows with the registered node count. This is not a fixed global limit. -Keep offline workers so retained resources remain observable after reconnect. - -Allocation scans filter by the fixed node before their 32-row page limit; pending -scans join the unreleased committed placement. Each node advances its own cursor, -including failed observations, and wraps once at EOF. Direct provisioning resolves -the tenant-scoped existing placement before entering that same node's gate; an -existing allocation must agree with the placement. Never choose another node. -E2B direct allocations use one serial lifecycle without a node identity. -The coordinator stops accepting work and cancels and drains all node workers and -direct callers before releasing the sole execution lease. Lease loss is global; -ordinary provider failures stay within their node. A planned deployment drain or -inventory retirement cancels lifecycle contexts synchronously between leased -operations, using the existing lease gate with its five-second bound. Include -an active manual reconcile's cancel function: an asynchronous `AfterFunc` alone -can cancel a later leased query after the gate reopens. Never cancel an in-flight -leased query merely to change deployment configuration. A failed cancellation -fence synchronously closes manager admission and reports owner failure, even -before its coordinator starts. Failed inventory retirement retains the original -lifecycle identity and gate through owner shutdown; gate availability is not -proof that cancellation succeeded. The drain barrier stays closed and cannot -activate a replacement. Release the lease gate before waiting for provider -settlement or lifecycle accounting. Ordinary caller deadlines and owner shutdown retain their -existing cancellation and fail-closed lease-loss behavior. Session locks, deployment -capacity transactions and revision/one-shot receipts remain authoritative, with -no external operation holding a database lock. - -Commit environment-to-node placement with Session creation and its creation retry -identity. Placement is automatic: Core chooses an eligible node, and callers cannot -select one. Existing retries keep their original node even when it is offline. Node capacity counts pending reservations and -unresolved resources; new placement and suspended-to-restoring admission share a -database lock. Confirmed cleanup releases placement capacity. Retained ownership requires exact -provider evidence; a socket path, missing instance or empty listing cannot prove -cleanup or authorize replacement. Historical file-managed ownership must be resolved -using its previous release and original backend before retiring that configuration. -Do not add a startup adoption path to bypass the database-managed selection. - -Do not add node-level drain controls. Refuse node removal with pending allocations, -instances, snapshots, unknown results or cleanup resources. Offline ownership is -retained. Removing a node does not delete compute. Local and remote nodes share -the same resource guard. Keep the deployment-wide reset and configuration -change guard. This boundary does not add cross-node Session -migration, Core multi-active, autoscaling, Kubernetes or harness residency. - -The common -`services/core/internal/sandbox` contract owns the five required operations and optional capabilities documented -in [the Provider guide](../../docs/sandbox-provider.md). Core orchestration must not import an adapter or SDK. Exact compute -identity, generation construction, inspection, full snapshot capture, restore, -thaw and owned artifact cleanup use that common capability. Initialization and -execution use the authenticated Runtime peer. -Provider-specific names and snapshot identities are opaque to Core. Self-hosted -compute and providers without checkpoint support keep their existing behavior. - -The [microsandbox deployment profile](deploy/microsandbox/README.md) -pins the SDK, runtime, firmware and image. Core remains a pure-Go binary. The -one-shot native Linux helper contains the SDK/FFI and runs under the same private -service namespace; it is not another scheduler or network control plane. Ordinary -pause does not release RAM. Suspension captures and verifies a full snapshot, -stops the exact source, and removes its writable compute closure only after the -artifact is durably identified. Explicit network policy applies on create and -restore. Do not inherit undeclared host resources. -The native SDK owns a dedicated ext4 disk mounted at `/environment`, separately -bounded by `environment_disk_mib` alongside `root_disk_mib`. Workspace, staging -and outputs must share that filesystem; do not weaken cross-device or link -checks to accommodate the layered root. Creation uses `/` until bootstrap creates -the workspace. Existing full snapshots and sandbox cleanup own the disk, with -no external mount or separate storage lifecycle. Native restore may omit a configured -root-disk size because it inherits the verified full snapshot. Accept that omission -only with matching snapshot resource proof and exact source/target identity; inspect -other native limits before retaining the inherited proof on the restored target. -Never treat a missing root size as unlimited capacity or resize retained state. -Snapshot receipt observation uses ownership and artifact integrity checks, so a -resource mismatch cannot hide an existing snapshot from authorized cleanup. -Restore recovery may finish a missing derived resource proof on the exact target -of the verified snapshot, after checking its native limits. It must not repeat -Restore, start stopped compute or alter resource limits; reread the same target -strictly before returning success. - -Suspend only after at least one Turn is terminal, no queued/in-progress/waiting -root or subagent Turn, pending input/file operation or initialization remains, -and real activity has been idle for the configured interval. For node-managed -allocations, record the first root or child terminal transition in the same -transaction using Core's database clock and the existing compute activity field. -For every Environment source, positive native completion timestamps remain -unchanged in public history, even when host clock skew places them before Core's -Turn creation time. They cannot drive idle admission across hosts; repeated terminal projections never reset that timer. -Read activity together with the database observation time. Candidate filtering and -the Session-locked phase recheck compare elapsed database time with the configured -idle duration; callers must not supply a Core-wall-clock cutoff. Anchor the initial -snapshot retention deadline to that same database observation. Core and database -host clocks need not be synchronized for these decisions. -Heartbeats do not reset activity. The daemon must close admission and drain native -cleanup, output receipts and file work before acknowledging planned suspension. Never change a -harness or keep an agent process alive across Turns solely to meet this feature. -The acceptance boundary is a next Turn in the same Session with history, files -and configuration intact, without replaying an earlier request. - -The existing Worker lease, Session lock and per-node lifecycle gates own both providers. -New Turn claims, file-write intents and capture admission serialize under the -Session lock. Turn and file-write admission share the same compute-phase check; -existing receipts remain readable. New pending work cancels capture and wakes -the same source. Normal preparation waits for the -compute phase to be running, after the authenticated resume handshake; a pending -input remains pending if its promotion conflicts with a lifecycle transition. -Private compute phases and revision-checked JSON receipts live on the existing -allocation. Persist quiesce/capture/restore intent before effects; only the fresh -receipt performs a capture or restore. Recovery observes the exact attempt and -never retries an unknown creation, capture or restore. A consumed snapshot cannot -roll a running generation back. Deletion, revocation and retention expiry take -precedence over wake, including at the final database compare-and-swap. Retain -unknown cleanup identities until owned resources are confirmed absent. - -Database-owned guest CPU/memory and supported disk settings, max_active reservations, max_retained -allocation count and snapshot retention bound each assigned node. Unknown operations -retain capacity reservations. Source teardown must be confirmed before releasing -active capacity. Delete consumed artifacts and old compute closures; do not grow -a chain of old writable disks across suspension cycles. No Kubernetes, distributed -scheduler or snapshot replication belongs in this V1 profile. - -Queued work and live Environment file access request wake. History and published -artifact reads do not. Planned suspension uses private daemon wire 0.8.0 with an -Environment and suspension token; a PID/start-time fenced local control signal -wakes the parked daemon, which reauthenticates before admitting new work. A -transient disconnect before confirmation retries the same armed suspension with -bounded attempts and backoff; permanent authentication or protocol rejection -still closes it. Snapshot -lifetime has no daemon wall-clock timer: Core owns its retention deadline. A lost -quiesce acknowledgement may thaw the same source using explicit rollback control; -it does not authorize capturing it. Ordinary disconnect keeps the existing -conservative shutdown behavior. Authentication rejection cannot create a new -Runtime or replay a request. - -All normal `make check` gates still apply. Linux qualification additionally runs -the pinned helper module tests/build through `check-microsandbox-provider`, a real -KVM full-snapshot/reclamation probe, and idle-to-next-Turn integration acceptance. -Synthetic process-memory probes support the backend claim only; they do not prove -agent continuity. Independent blind review uses the clarified idle-only scope. - -Managed Runtime allocation, dedicated daemon credential hash and exact Session -binding commit atomically before Provider.Create, using the existing execution -lease and Session lock. Only the fresh allocation receipt permits Create; retries -and Core restart observe that same reference without replay or credential rotation. -The operator's stable provider key identifies one backend/installation; retain its -adapter for cleanup, and use a different key when changing the target. Never treat -absence on another backend as successful reclamation. -Disabling the default provider stops new hosted admission/bootstrap; it must not -block existing Session cancellation, tool results or input retry outcomes. -Input HTTP response budgets follow the persisted Environment type, covering the -admission wait for both hosted and self-hosted Sessions independently of operator -creation switches or remote executor configuration. - -With an explicitly configured default managed provider, the same Worker scans -committed pending hosted Environments that have no allocation. This includes idle -Session creation and recovery after commit-before-bootstrap interruption; an -existing allocation never enters that startup path. Keep the scan bounded and -serialized by the existing lifecycle owner. Hosted provisioning requires no caller -connection action. An initial reservation without a Turn leaves its Session idle, -as allowed by the pinned contract; do not emit an in-progress event before a Turn -starts or treat a daemon connection as native readiness. - -The same serialized scan publishes authenticated connection observations using -the existing durable generations after verifying the exact Session/device binding -and settled bootstrap. Socket loss remains observable during a provider outage; -Core restart fences old observations. Do not create a separate connection owner. - -Terminal managed cleanup atomically revokes authority, persists Environment failure -or expiry, settles pending input and requests cancellation before external cleanup. -Preserve original input deadlines and retry outcomes. Temporary provider outages, -unknown Create results and stopped compute do not prove permanent failure. The -pinned stream has no Environment expired event; do not invent one. Exact hosted -failure codes and ordering remain explicitly unverified. - -Allocation state is private compute ownership, separate from public Environment -connection/native readiness. Adapters qualify bootstrap completion; Core does not -infer it from an engine or provider name. Connected, observed compute receives -service keepalives between Turns. Keepalives cannot revive a one-hour lapse or a -cleanup request. The Docker provider keeps its current idle behavior; only an -explicit checkpoint policy may suspend completed, idle work as described below. A stopped/missing container -does not authorize discarding retained workspace or history. Session deletion or -expiry requests cleanup, revokes the scoped device and cancels pending work before -Provider.Kill; the existing Worker serializes these lifecycle operations and drains -them before releasing its execution lease. - -Keep the allocation after public Session deletion. Mark it released only after -owned compute/volume cleanup and evidence that its original Create has settled. -An unknown creation retains cleanup ownership even after an absence observation; -continue bounded scans for late resources without issuing another Create. This -conservative internal lifecycle does not define user-managed enrollment or prove -complete upstream expiry/error semantics. - -Qualify the actual Runtime before cutover: execution, file access, owned -cancellation, restart with retained history and files, and clear failure when -required history is missing. Daemon tools run with the launching account's full -permissions. Managed isolation is provided by the outer Environment. A user who -installs on a host does not receive a sandbox or protection from their own tools. -Keep failed probes and unverified platform combinations explicit. +Source Files are Project resources with a lifecycle independent of copied workspace files. Store immutable source metadata and PostgreSQL large objects in Core's database with the pinned pgx driver. Upload validation, metadata insertion and the bytes commit atomically; deletion removes the metadata and unlinks the object in one transaction. Keep OIDs private and authorize every metadata, content and delete lookup by tenant before opening a body. Stream bounded chunks; never hold an entire upload in memory or use a filename as a filesystem path. A direct download of the `user_data` purpose is rejected after the tenant-scoped metadata lookup, while initialization and workspace copies keep their authorized store read. A read-only repeatable-read transaction preserves an admitted source across concurrent deletion; resolve that snapshot before entering the Environment write path, and a later deletion never undoes a completed workspace copy. Bound request and transaction lifetimes, roll back incomplete bodies and never retry an ambiguous commit automatically. Backups must include PostgreSQL large objects, and a schema rollback must not orphan them. + +Session Artifacts are immutable published copies, separate from live workspace files and source Files. The private output exporter reuses the authorized workspace path boundary and streams bounded bytes; publication requires complete capture and confirmed helper and transport success, not merely valid archive syntax. The daemon owns and drains the exporter's stdout pipe separately from child reaping, so pull-transport backpressure cannot consume the process-exit I/O deadline; after helper exit, each pipe read has one second, reset after consumer delays, which rejects inherited pipes that never close. Cancellation closes the owned reader and the dispatch consumer, then waits for the child. Never extract an output archive into Core's filesystem or hold the execution lease through a large transfer. + +Capture bytes into private large objects without a Session admission lock. Before capture, seal native input under that lock with the private capture marker; later messages reuse the Environment input reservation and wait for the next Turn, and the public Turn stays in progress until publication settles. Directory reads during capture use an independent read-only preparation, not the released native Run. After a confirmed export, lock and recheck the live Turn, then publish the metadata in the same transaction as the Turn's completion. In that transaction, drop staged paths whose sha256 equals the newest published Artifact for the path in the Session, so later Turns publish only new, changed or no-longer-published paths and never modify existing Artifacts. Failed and cancelled Turns discard private objects, and Session deletion removes private and published copies. The exporter skips output symlinks by their `lstat` type without following them; hard links, other special files, device crossings and concurrent changes reject the capture. Stored reads are authorized independently of Environment availability, so published Artifacts outlive the Environment. + +## Scheduling, preparation and pending input + +The leased Worker's common preparation scheduler owns Environment initialization, independently of any allocation; managed and enrolled connections use the same frozen snapshots and typed Runtime operations. Preparation has bounded concurrency separate from Turn scheduling, and a blocked Runtime never blocks cleanup or other connections. Check Harness availability before installing. A persisted running initialization whose owner is lost fails without replaying side effects; completion requires the same authorized Environment and device binding, and failure settles pending input while keeping compute ownership and user files. A connection observation never implies completion. + +A committed input sends a coalesced hint to the Worker scheduler, which keeps its lease, capacity, cursor fairness and per-Session ownership checks; a hint admits nothing by itself. If capacity is occupied, keep one rescan for completion without turning a failed preparation into a busy retry loop. HTTP readiness waits and active input delivery subscribe before reading state, wake after committed promotion or input and recheck storage after each hint. Polling remains the fallback for external writers, expiry and lost hints; notifications carry no execution authority and no input data. + +Core readiness, Start acknowledgement and input-to-first-text logs use one process monotonic clock; a Start acknowledgement confirms adapter ownership, not model input consumption. Runtime logs report Executor creation, reuse, idle and close separately from Turn completion. Never call these durations model latency or subtract clocks of different machines. + +The Dispatcher prepares pending Environment input only on its exact enrolled or managed device, and keeps the same physical peer and preparation handle through readiness, atomic promotion and claim, and the first non-replay Start. Workspace and Environment identity come from store ownership, never caller-selected paths. The initial prompt and cursor come from the reserved batch; later messages use ordinary steering. Never hold a database lock during native preparation. While waiting for readiness, observe the original deadline, cancellation, deletion and peer loss. A preparation failure leaves pending input and its deadline intact unless storage settled it, and creates no failed Turn or input history. During Start, consume preparation controls alongside the Run stream so a control-only rejection or a pending-start cancellation settles promptly; once cancellation is sent, preparation errors cannot replace its receipt or timeout path, and a missing or unconfirmed outcome fails conservatively without a fabricated cancellation outcome. The connection owner spans preparation and the transferred Run, and every exit releases it. + +The Worker scans pending inputs with the same scheduling slots, Session locks, durable deadlines and engine capability checks: once a second, at most 100 candidates per scan. At the end of the queue after a nonempty cursor it refills the first page once in the same scan, and an empty queue never spins. The cursor advances before readiness checks so an unavailable Runtime cannot starve later candidates. Execution concurrency (`core.execution_concurrency`) bounds simultaneous work, not attempt frequency, and the Worker alternates between ordinary Turns and Environment inputs. A self-hosted Session waits for its dedicated enrolled device; it never selects another device of the tenant or migrates a binding. Managed lifecycle polling keeps its separate five-second interval. Input HTTP response budgets follow the persisted Environment type, independently of operator switches. + +`OAC_HARNESSES` adds deployment-supported engines to the default engine and the managed profiles; a daemon's heartbeat alone never enables an engine. + +## Runtime connections + +`internal/agentdaemon/gateway` is the shared daemon connection implementation; its persistence interfaces use `internal/agentdaemon/device`, and the frames and validators live in `internal/agentdaemon/proto`. It is a single-process registry: connectivity comes from the live registry, never a persisted online flag, and `last_seen_at` is diagnostic only. Session-to-device bindings are tenant-scoped and immutable. Revocation denies new connections and binding reads at once, and an open connection closes at its next heartbeat. + +A dedicated self-hosted device is bound to exactly one Environment's Session and is excluded from general device selection, even within the tenant. Enrollment creates or recovers the device and binding atomically under the Session lock; the frozen workspace and capability directories come from the Session configuration and must match the local binding. Core rechecks the persisted Environment and device binding for preparation and active reads; capability discovery never selects or authorizes a device for this placement. + +Connection observations use the execution lease and the Session lock. A separate `environment_connections` row holds the current generation and revision, and `environments.status` commits together with its Session Environment event. The producer serializes replacements and numbers socket observations within each generation; duplicate or older revisions and superseded generations are inert, and a replacement retires the previous connected observation before registering the new one. Registration alone creates no `connected` event. Event payloads carry only public Environment identity, type, status and nullable error, never configuration, credentials, registration IDs or revisions, and have no Turn association. `connected` and `disconnected` are distinct from native readiness: never cast resource `expired` into this vocabulary or emit `ready` for a self-hosted connection. On restart the Worker reconciles old observations before admitting new ones, and a failed or stale observation never establishes a connection. + +The sandbox node Hub's global mutex protects only in-memory connection state. Authentication, ownership and store callbacks run synchronously outside it, respect cancellation and have five seconds; database writes are never detached. Closing the Hub cancels opening and live connections without waiting for database callbacks, and each node reservation lasts until its fenced disconnect cleanup finishes. Presence is registered in an explicit transaction, so a canceled statement cannot publish it later through autocommit. Disconnect cleanup first locks the node row, then applies the connection and epoch fence with a fresh READ COMMITTED statement so an in-flight commit is not missed. These transactions never take the deployment-wide manager lock, and online-state writes compare the handshake epoch atomically so a stale Core cannot publish readiness for a new owner. + +## Sessions, Turns and input + +Turn writes serialize on the tenant-scoped Session row. Input requests are ordered batches committed under the same lock: a retry key identifies the complete batch, a changed length, order or content conflicts, and a failed transaction leaves no partial input or cancellation. A single-event request keeps its identity at batch position zero. Internal admission limits are 64 events and 512 KiB of payload per request; the API still validates the upstream event schema. Cancellation keeps its first target, including an idle no-op, so a retry never stops later work. Terminal states and outcomes are never overwritten. + +Session deletion takes its decision and commits the `deleted_at` marker under the tenant Session lock that orders Turn and input admission, so either an admission commits first and deletion conflicts, or admission observes the deletion. Admission checks visibility under that lock before its retry lookup. Internal Turn, receipt, finalization and restart reconciliation keep access to deleted Sessions so that work settles under the execution lease; queued work of a deleted Session is never claimed, and Runtime cleanup still cancels pending work. Deletion never revokes a shared device or removes a saved Agent. Deleted Session records are retained, and a migration rollback refuses to drop the column while any exist. + +Every new Session needs an explicit typed creator at the store boundary, internal callers included; public creation derives it from the authenticated principal only. Creator kind and ID persist in the creation transaction and never change on retry, update or source mutation. Both the early saved-reference recovery and the authoritative creation upsert require a matching creator before returning a Session or event cursor. Tenant scope always comes from the authenticated identity before any store call; metadata grants nothing. Tests supply explicit synthetic creators. + +New saved-reference Sessions, and inline requests with Vault attachments or credential references, record a separate caller-intent hash: the source ID (empty for inline), supplied overrides with field presence, Environment and Vaults, original metadata and normalized initial input, excluding response streaming and resolved source values. Compare that same-tenant retry identity before looking up the source; a match returns the existing Session without input admission or source revalidation. Recheck after a source resolution failure for a concurrently committed creator, without holding a lock across resolution; the unique creation upsert stays authoritative. Credential-bound retries recover before reading mutable Vault contents. Every new hosted Session also records caller intent, before deployment defaults are resolved; the resolved-configuration hash of other inline Sessions leaves the deployment default out. Provider keys enter retry hashes only as keyed fingerprints. + +Session listing by `agent_id` filters on the immutable root `configuration.agent.id` with the tenant, Agent and creation index, before pagination and activity projection. + +Session state and last activity derive from the latest persisted Turn: queued or active work is `in_progress`, successful or cancelled work `idle`, and failures carry a safe public error. + +## Agents and model providers + +Reusable Agents are tenant-scoped rows independent of Session snapshots and engine bindings. The store persists caller-validated configuration without applying harness restrictions or model defaults, with internal limits of 512 KiB for configuration and 64 KiB for metadata. An update locks the Agent row while merging the supplied fields and enforcing the configuration bound, then commits configuration, metadata and update time together, so a stale full snapshot never overwrites another update; an update without fields changes no timestamp. Deletion is one tenant-scoped `DELETE … RETURNING id`. A Session copies the saved configuration into its immutable snapshot and never looks up its source again. + +Saved execution defaults keep a model-provider bundle whole at every replacement boundary: endpoint, key, protocol and limits are never inherited separately. Agent JSON holds only the safe provider fields and an output-only configured flag; the complete bundle is encrypted separately with a tenant and Agent binding and its own purpose, and configuration and secret changes commit together under the Agent row lock. Model-only edits need no key. Merged harness, protocol and limits are validated without reading keys. Session creation reads safe defaults and ciphertext in one snapshot, and a complete Session override does not decrypt the inherited bundle. The resolved bundle is frozen in an encrypted Session-owned row, and dispatch fails closed when that snapshot is missing or cannot be decrypted; later Agent edits, default changes, restarts and suspension never resolve it again. + +Deployment-default observations use a private revision UUID generated on every PUT, identical replacements included. Session creation reads the ciphertext and revision together and freezes them; retries and older Sessions never gain or replace revision metadata. After a successful root terminal commit, one independent pool operation has at most one second to update the matching current revision. SQL verifies the tenant, root Turn and committed outcome, and only completed Turns and native provider failures with `engine_failed` count; cancelled work, Core or Runtime errors and input-policy classifications never do. The metadata-only transaction sets statement and lock timeouts within the remaining budget and issues one UPDATE that locks only the default and samples database time after the lock. Errors throttle for 30 seconds, ordinary successes throttle for 30 seconds with one immediate recovery write after each accepted error, and an unchanged revision has at most three effective writes in any 30-second window of nondecreasing database time. Observations never change `updated_at`, readiness or execution truth, and can be lost or stale; there is no queue, retry, probe or backfill. + +Session execution-configuration reads use a separate immutable safe projection written with its provenance in the Session's creation transaction. It reads no ciphertext, never recomputes sources from current Agents or defaults, never touches activity or wakes a sandbox, and does not affect retry identity. + +Provider input validation uses the adapter rules in `internal/harnessconfig`: one internal registry for protocol and token-limit validation, while Core owns credential environment and endpoint admission policy. + +## Vaults and credentials + +[Vault credentials](credentials.md) and [OAuth credentials](oauth-credentials.md) describe the resources, selection rules, refresh and deletion. The store implements them under these rules: + +- Credentials are children of tenant-owned Vaults. Creation admits the owner in the same SQL statement as the insert; retrieval joins the owning Vault; listing enforces Project and Vault ownership on the parent, cursor and row query. Metadata queries never select ciphertext and need no encryption key. +- Secret values are encrypted before they reach SQL, with Core's separately configured random 32-byte key and the standard library's random-nonce AES-GCM. The versioned authenticated binding covers tenant, Vault, Credential, authentication purpose and exact destination. Never reuse daemon transport encryption for this storage. A missing key disables credential writes; a malformed configured key fails startup. +- A static replacement is one SQL mutation scoped by tenant, Vault, Credential, auth type and destination, reusing the safe metadata for the immutable binding; it never decrypts the previous token, and a failed write keeps the old row. +- OAuth refresh and replacement serialize on the Credential row lock, authenticate the stored grant metadata against its encrypted copy before using an endpoint, and persist the refreshed grant before returning an access token, so a stale refresh cannot undo a deletion or Vault cascade. +- Credential deletion is one mutation checked by tenant, Vault and ID; Vault deletion removes the parent and its Credentials through the foreign-key cascade in one SQL statement, without decrypting, needing the key or calling providers. +- At dispatch, Core rechecks the tenant, attached Vault, selected ID, frozen auth type and exact URL before scoped decryption, and the token enters only the transient daemon request. A missing key or binding failure never falls back to anonymous execution. + +## MCP + +Accepted public MCP credential profiles are declared centrally by the service, separately from private adapter capabilities; a daemon capability alone never opens a profile. Frozen binding validation runs at admission and later input, and the same capability and placement checks run at device selection, the final preclaim check and request construction, before scoped decryption. Native configuration and token injection stay in the adapters; the [Codex adapter](deploy/codex/README.md#mcp-servers) and the [Claude SDK adapter](../../packages/claude-sdk-adapter/README.md#http-mcp) describe theirs. The shared resolver keeps omitted or null `allowed_tools` as unrestricted and an explicit empty list as deny-all. Saved HTTP transport output includes `headers: {}` while the effective Session transport omits headers, matching the two pinned resource types. An omitted or null `connection_origin` on HTTP transport is stored as `service` before the other checks. + +## Dispatch and execution + +`services/core/internal/execution` claims a Turn from `queued` to `in_progress` before subscribing or sending, and never replays a claimed or interrupted Turn. Extra inputs require native receipts. The terminal outcome and the native Session ID commit together under the admission lock, and unapplied messages prevent a successful completion. Credentials are resolved separately from the immutable non-secret snapshot. A peer that lacks a required capability is rejected before the claim, and a failed strict resume never starts unrelated history. Native continuation needs the device's persisted engine files; a native Session ID alone cannot restore deleted history. + +Core replaces a Turn's usage with each complete valid token breakdown and keeps the last committed measurement when execution is interrupted. It never infers tokens from context occupancy or costs and never parses native raw payloads. When `done` also carries usage that was already reported, the counters are not added again. + +Subagent observations use the common types in `internal/agentdaemon/proto/subagents.go`. Core assigns public IDs and projects them under the Session lock and the leased execution journal; native names, history parsing and outcome proof stay in the adapters. Child Turns have their own table, separate from Core's queue, and Session Turn reads and the Session event stream carry root work only. The migration-defined `public_execution_turns` view has no public reader; do not reintroduce mixed Session Turn pages. Repeated effects are idempotent, root output freezes first, and child Items are delivered before their terminal Turn snapshot. + +Function calls are stored per tenant, Session and Turn with immutable public and executor call identities and arguments. Result admission and application receipts serialize on the Session lock with cancellation and terminal transitions. The store keeps the complete caller-validated result, including omitted versus null error and output and ordered text and image parts; identical retries return the saved decision and changed results conflict. Recording a call moves the Turn to `waiting`, and the last application receipt resumes it. A function result may join a message or cancel batch: its explicit Turn and call identity select an existing call, admission never creates a Turn for it, and the result and its input retry record commit in one transaction. Targets resolve only after the tenant Session lookup; an unknown call or a call of another Turn is 400 `invalid_request_error`. The input cursor skips function results, whose own receipts decide application. A Session with functions waits for a device that declares `function_tools`, and Core delivers each saved result once per live dispatch, appending a non-null error as a final text part because the native result has no error field. + +Neutral tool and message observations go into the journal before Core projects public Items; the Item projector validates the shared observation contract and never decodes engine-native snapshots. `command_output` fragments update Items only for a command already indexed in the same Turn, and commit with their `agent.output.command_execution_output.delta` events. Completion output replaces accumulated drafts, terminal Items ignore late fragments, and cancellation keeps partial output without inventing a successful completion. + +## Journal and live events + +Execution observations are written to tenant-scoped `turn_events` in ordered, idempotent batches before they back recovery or publication, with daemon payloads intact. Flush at least every 100 ms while consuming events and before terminal persistence; the terminal outcome, its journal entry and native continuity commit together. Journal limits are 512 KiB per payload, 1 MiB per batch, 65,536 observations and 32 MiB per Turn, with one extra entry reserved for the terminal outcome. Never infer a successful completion after a persistence error or stream overflow. + +Live Session SSE reads `session_events` committed with the corresponding input, Item or lifecycle change under the Session lock, as immutable transition snapshots. The notification buffer keeps at most 256 events and 64 MiB per Session after each transaction, and read batches at most 32 events or 1 MiB, each keeping a single oversized event. GET polls committed events every 100 ms from the committed high-water mark; a missing sequence position ends the stream with a safe error. Socket writes have a five-second deadline and hold no database connection. Rebuilding historical indexes emits no live events. + +A streaming Session creation reuses atomic input admission and the live event loop. The creation upsert returns its event cursor under the Session lock, before the initial inputs; never replace it with a post-commit cursor lookup. The settled marker and pending-input flag the stream uses to end are store-internal, never wire fields. The stream rereads the Session projection after it sends a Session status event and otherwise at most once a second. + +## Items + +Public Items read a projection updated in the same Session transaction as admitted messages and journal batches. Item IDs derive from the Turn and source identity, and the first-observation timestamp and tie breakers never change when content or status does. Each new Item's Session position is allocated under the Session lock, preserving observation order for equal timestamps, and each Turn allocates its own zero-based `output_index`, which inputs do not consume; updates and retries keep both. + +Item merging never mutates the incoming observation or the previous snapshot: public text delta events read the original fragment after merging, while the Item keeps the accumulated text, and the content slice is copied before its text pointer is replaced. A first observation without its own fragment carries its unchanged text in one delta. Wire-only explicit nulls come from response marshalling, while stored Item payloads keep their original encoding through `Item.MarshalStored`, so replayed child Items compare equal. Structured tool JSON is kept without float conversion, and an unfinished call never becomes a successful result. Function results are Session input Items: they emit `item.added` with a null `output_index` and never `item.done`, whose upstream union allows only agent output, and their public output and error come from the saved submission. + +## Worker ownership + +Enabling the daemon gateway with `OAC_PUBLIC_URL` also starts the execution Worker. One Worker owns an execution database through a dedicated PostgreSQL advisory-lock connection, and its store view uses that connection for every Session transaction: binding, claim and reconciliation, journal, Items and usage, function callbacks and receipts, and terminal state. These short transactions and the lease pings serialize, with a five-second deadline that includes gate and Session-lock waits. Never hold a transaction across daemon or model work, reconnect the writer or fall back to the pool after losing the lease. Public admission and device maintenance use pooled connections, and pooled reads grant no write authority. + +Sandbox reset snapshots bind the deployment relation explicitly to its single row before joining resources, so that even on a fresh database without statistics an inflated join estimate cannot trigger JIT compilation inside the lease deadline. + +At startup the Worker fails previously claimed work, keeps queued input and never replays uncertain execution. Shutdown cancels active dispatch and attempts terminal persistence before releasing the lease; a lost owner cannot commit. Closing the lease invalidates its writer and waits for pgx cleanup within the caller's deadline; a later close can resume that wait. Tests that transfer ownership immediately must observe the previous owner's advisory lock disappear before starting the next, with a bounded wait that fails on query errors. The lease fences database writes, not already queued daemon commands or native effects. + +## Hosted sandboxes + +The [managed lifecycle](../../docs/sandbox-provider.md#managed-lifecycle) describes publication, allocation, per-node workers, placement, suspension, reset and archive; the [sandbox node protocol](../../contracts/agents-api/node-generation-protocol.md) the node connection. In code: + +- Each allocation's immutable specification and the current credential resolve in one snapshot, with no stale-credential cache, no fallback to the current specification and no unloading of a provider map that retained allocations need. +- Credential replacement fences provider calls and waits for helper subprocesses to exit, even after cancellation, before verifying again and committing; the audit entry never carries secret fields. +- Observations of offline, missing or unconfirmed resources persist only bounded, sanitized codes, separately from the lifecycle, and a host restart never fabricates a running observation. +- Runtime observation of a node allocation goes through its immutable placement as one bounded read-only Provider operation; an absent capability or transport returns unavailable, never a Core-local fallback. ## Public and administration API constraints @@ -1610,7 +137,7 @@ These constraints implement the rules in the [Agents API contracts](../../contra ### Wire validation mechanics - Stored strings other than metadata rely on PostgreSQL rejecting U+0000 and invalid UTF-8: map SQLSTATE `22021` (text parameter) and `22P05` (`\u0000` in jsonb) to the 400 unstorable-text error, including query filters such as `agent_id`. The failing statement aborts its transaction, so keep each request's writes in one transaction. -- An `after` cursor that cannot name a resource on a lookup list (Agents, Sessions, Turns, Templates, Vaults, Credentials, the Core Runtime observation list) resolves to the never-assigned maximum UUID and runs the normal lookup, so storage failures and missing rows behave as for a well-formed cursor. Resolve every cursor only inside its already resolved parent and tenant. +- An `after` cursor that cannot name a resource on a lookup list (Agents, Sessions, Turns, Templates, Vaults, Credentials) resolves to the never-assigned maximum UUID and runs the normal lookup, so storage failures and missing rows behave as for a well-formed cursor. Resolve every cursor only inside its already resolved parent and tenant. - Lists whose parent and cursor lookups are separate statements (Artifacts, Skill versions) re-check the parent before reporting a cursor 400, so a parent deleted in between still returns its 404. Item and Subagent lists read both inside one locked Session transaction. The Skill version cursor lookup is tenant-wide so another Skill's version can be told apart from a missing one; another tenant's version stays missing. - Saved Agents parse tools with the saved-form parser, which keeps every pinned `web_search` mode. Session admission re-resolves the effective tools with the execution parser, which admits only disabled search; Worker device selection and the final preclaim also refuse a non-disabled search control. diff --git a/services/core/README.md b/services/core/README.md index 1720819af..b46a0e729 100644 --- a/services/core/README.md +++ b/services/core/README.md @@ -1,671 +1,70 @@ -# Agents API +# Core service -See [Configuration](../../docs/configuration.md) for process, deployment, node and daemon configuration ownership. +`services/core` is Core: one Go service that serves the Agents API (`/v1`), the Core API (`/core/v1`) and the machine connection routes (`/api/v1`), owns its PostgreSQL schema, and runs the execution Worker that dispatches Turns to Runtimes. [Architecture](../../docs/architecture.md) describes its role, and the [API index](../../docs/api/README.md) its namespaces and credentials. This guide is for contributors who build, run and test Core from source; operators install it with the [installer](../../docs/getting-started/install.md). The [implementation constraints](IMPLEMENTATION.md) hold the code-level rules. -Independent execution service implementing part of the pinned OpenAI Agents API. -It owns reusable Agents, durable Sessions/Turns/Items, live events, function actions -and a daemon execution worker. Public execution supports qualified Codex, Claude Code -(`claude_sdk`) and MiniMax Code (`mcode`) profiles through the shared Runtime contract. -The three-harness Linux amd64 Docker V1 MVP has accepted evidence. V1 user-managed -Runtime enrollment has [recorded real acceptance](../../contracts/agents-api/harness-capabilities.md) -with explicit deployment coverage. [E2B](deploy/e2b/README.md) can back Core-managed -hosted sandboxes, selected in Web's sandbox setup, or application-managed -`self_hosted` Environments. -It builds and runs with its own PostgreSQL database and credentials; -Parsar's product service, frontend and database are not required. +## Commands -Use this guide to build, configure and connect a client. The -[protocol coverage](../../contracts/agents-api/README.md) lists supported operations, -engine limits, acceptance evidence and missing resources. The target remains the -complete pinned protocol; current workflows do not establish full compatibility. -Parsar product execution and its eventual public-client cutover are separate. +| Package | Executable | Purpose | +| --- | --- | --- | +| `cmd/server` | `oac-core` | The HTTP service and execution Worker | +| `cmd/migrate` | `oac-core-migrate` | Applies the embedded database migrations | +| `cmd/device` | `oac-core-device` | Provisions or revokes an [operator device profile](../../contracts/agents-api/machine-api.md#operator-device-profile) for `environment: none` engine hosts | +| `cmd/environment-key` | `oac-core-environment-key` | The [break-glass executor credential command](../../contracts/agents-api/environment-executor-credentials.md#break-glass-command) | +| `cmd/sandbox-node` | `oac-node` | The sandbox node program; see the [nodes guide](../../docs/getting-started/nodes.md) | +| `cmd/specification-contract` | None | Regenerates the installer's node specification projection | -## Reusable Agents +`make build-core` builds the five executables into `~/.oac/build/oac-core`; [Standalone Core builds](../../docs/maintainers.md#standalone-core-builds) describes the build and its options. `make build-daemon` builds `oac-daemon`. -Static-bearer and OAuth Vault Credentials support creation, replacement, deletion -and safe metadata retrieval/listing. Vault deletion atomically removes its -Credentials. Configure their independent encryption key and authenticated Session -use through the [credential guide](../../contracts/agents-api/vaults.md); see [OAuth credentials](../../contracts/agents-api/vaults.md#oauth) -for application authorization, dispatch-time refresh and revocation boundaries. +## Database -The pinned Python client saves an Agent independently, then starts a Session with initial input: +Core uses its own PostgreSQL database and account and shares no tables with an application. Its migrations are embedded goose migrations in [`migrations/`](migrations), tracked in `agents_api_schema_version`; a migration is never edited after it lands. Queries live in `internal/db/queries` and generate typed code with sqlc: run `make sqlc-generate` after changing a query and commit the generated files, which `make check-sqlc` compares. -```python -agent = client.beta.agents.create(model="your-model", name="Example") -session = client.beta.agents.sessions.create( - agent_id=agent.id, environment={"type": "none"}, input="Hello", -) -``` - -These resources belong to the authenticated execution tenant. Saving configuration -does not launch an engine. Optional Session `agent` fields override the saved -configuration: omitted fields inherit, supplied objects/arrays replace whole fields. -Source and Session metadata stay separate; execution never looks up the source again. - -- Retrieve with `client.beta.agents.retrieve(agent.id)`; list saved resources with - `client.beta.agents.list(limit=20, order="desc")` and SDK auto-pagination. -- Update with `client.beta.agents.update(agent.id, instructions="New instructions")`. - Omitted fields remain unchanged. Metadata replaces all pairs; null/empty clears it. - Existing Sessions retain their configuration; new Sessions resolve the update. -- Delete with `client.beta.agents.delete(agent.id)`. Existing Sessions and history - remain available. New references fail; recorded creation retries recover their - accepted snapshot without consulting the deleted source. -- List Sessions with `client.beta.agents.sessions.list(agent_id=agent.id)`. Filtering - uses the immutable root ID, including inline Agents and history after source changes. - -See the [configuration and retry limits](../../contracts/agents-api/wire-semantics.md#sessions) -before relying on optional settings or hosted error/default equivalence. - -## Build standalone binaries - -```bash -make build-core -# Optional absolute output directory: -OAC_DEV_CORE_BUILD_DIR="$HOME/.oac/build/oac-core-test" make build-core -``` - -The default output is `${OAC_DEV_HOME:-$HOME/.oac}/build/oac-core`: - -- `oac-core`: HTTP service and execution worker. -- `oac-core-migrate`: this service's embedded database migrations. -- `oac-core-device`: operator device provisioning and revocation. -- `oac-core-environment-key`: principal executor key issuance, rotation and revocation. - -Use these executables in place of the corresponding `go run` commands below. -The build needs Go and access to its pinned module dependencies; it does not need -Node, Docker, the product service or frontend. An isolated source context enforces -that boundary on every build. [Contributor rules](../../docs/maintainers.md#standalone-core-builds) -define the allowed shared packages and required checks. Runtime database/key -configuration and a separately installed execution daemon are still required; -these binaries do not establish full protocol coverage. - -## Database ownership - -Use a dedicated PostgreSQL database and account, separate from the Parsar product. -This service does not import `server/internal` or apply product migrations. -Migrations are embedded and tracked in `agents_api_schema_version`. - -```bash -OAC_DATABASE_URL='postgres://.../agents_api' \ - go run ./services/core/cmd/migrate -``` - -The public `Idempotency-Key` creation header is optional: omission creates a new -Session. A supplied key identifies the request within its authenticated tenant. -Inline retries use normalized effective configuration; new saved-Agent references -record caller intent independently of later source updates/deletion. Retries do -not admit initial input again. Every retry must match the original typed creator, -including across key rotation. Keys in the same Project share the creator principal, -so rotating a key preserves that retry identity. Records with a known creator but -no recorded request intent retain resolved-snapshot behavior; records without a creator cannot be retried. -These retry policies are not verified hosted semantics. See the -[retry boundary](../../contracts/agents-api/wire-semantics.md#creation-retries). - -The Store uses internal creation keys and preserves immutable engine/configuration, -native continuity and same-tenant device bindings. The public API applies schema -validation/defaults before storage. Internal bounds are 64 KiB for metadata and -512 KiB for configuration. Keep credentials out of both. Public metadata permits -at most 16 string pairs, 64-character keys and 512-character values; storage bounds -do not replace those rules. Violations and non-string values return -`invalid_request_error` with a `metadata` or `metadata.` param. PostgreSQL -cannot store U+0000, so requests containing it in any stored string return 400 -before anything is written; this is a local limit, not hosted parity. Tenant identity comes from authenticated credentials, -never metadata or a caller-supplied business identity. - -## Internal Turn persistence - -Validated message/cancel/function-result batches commit under a tenant-scoped -Session lock. Messages start a Turn when idle and steer active work. Retry keys -identify the whole ordered batch; an invalid event does not partially admit it. -Queued cancellation needs no live engine. Active cancellation awaits a native -outcome, and completion may win the race. Terminal states cannot be overwritten. - -A Session keeps its effective configuration, engine and device across Turns; -product Conversations, native Sessions, connections, processes and sandboxes are -different objects. Strict native resume requires retained history on that device. -The API owns durable public history and pending function decisions; adapters own -native translation and their harness owns the model/tool loop. - -The worker uses its database lease connection for execution writes. Lease loss -fences those writes; it does not prove native commands or side effects have stopped. -Restart conservatively fails previously claimed work and retains queued work. -Uncertain delivery is never blindly replayed. Function actions and live SSE are -available within the [current coverage](../../contracts/agents-api/README.md); -other pending interactions, process-loss recovery and environment lifecycle remain -incomplete. Durable acceptance is not an exactly-once side-effect guarantee. - -## Standalone HTTP service - -Run migrations first, then `go run ./services/core/cmd/server`. The service -uses `OAC_DATABASE_URL` for its dedicated database; it does not read the -product database or accept product login cookies. It requires the Core key digest -file named by `OAC_CORE_KEY_DIGESTS_FILE` at startup; the Core key -authenticates `/core/v1`, where Projects and keys are managed. The Core key cannot -authenticate `/v1`, and application API keys cannot authenticate `/core/v1`. +## Run from source -Projects and API keys live only in the database. Configuration files contain -infrastructure settings and deployment credentials, not business identities. -Installation creates no Project or application key. Using the -[administrator API](../../contracts/agents-api/admin-api.md), create a Project with -`POST /core/v1/projects` and issue a named key with -`POST /core/v1/projects/{project_id}/keys`, or use Core Web's **Projects and -keys** page. Both requests accept a JSON object containing `name`; Core generates -the identifiers. +1. Create a development database and apply the migrations: -A Project owns one tenant and one execution principal. All keys in it have equal -access to its assets and share that principal; write provenance records the actual -key separately. Issuance returns plaintext once, and the database stores its digest. -Deliver the secret only to authorized applications. Rotate by issuing another key -in the same Project and revoking the old one. No secret-reset endpoint or service -restart is needed. Renaming a Project preserves its ID, principal and assets. -Archiving disables all its keys but retains assets and already accepted execution. -Administrators can read or delete retained resources. + ```sh + OAC_DATABASE_URL='postgres://oac:…@127.0.0.1:5432/oac_dev' go run ./services/core/cmd/migrate + ``` -Optional `OpenAI-Organization` and `OpenAI-Project` headers must match the Project's -execution scope; repeated or conflicting values fail authentication. The catalog -Project UUID and its external execution-scope identifier are distinct; see the -administrator contract. These identities grant no product-user rights. New Sessions -persist the Project principal as creator atomically and never change it on retry. -Historical Sessions keep unknown creators and remain readable; no key, metadata or -product record can assign their ownership through a retry. Retire older API writers -before starting this deployment; mixed-version creation is unsupported. Executor -credentials separately match this recorded creator before authorizing an Environment -connection; they do not inherit general caller API permissions. Native Runtime and -node transport contracts are unchanged. +2. Create a Core key of at least 32 characters and a digest file holding its SHA-256, which Core uses to authenticate `/core/v1`: -`OAC_ADDR` defaults to `127.0.0.1:8091`; use a TLS reverse proxy for remote -access. `OAC_DEFAULT_HARNESS` defaults to `codex`; use `claude_sdk` for Claude Code -or `mcode` for MiniMax Code. Configure the corresponding qualified Runtime through -its deployment guide: [Codex](deploy/codex/README.md), [Claude Code](deploy/claude/README.md) or [MiniMax Code](deploy/mcode/README.md). -It selects new Sessions independently of the requested -model. Existing Sessions retain their stored engine. Set -`OAC_HARNESSES=codex,claude_sdk,mcode` to explicitly enable installed profiles -for user-managed enrollment without a managed Provider. This list supplements the -default engine and any managed engine profiles; unknown names fail startup. -Enabling a profile does not install its harness or qualify its deployment. + ```sh + umask 077; mkdir -p ~/.oac/dev + openssl rand -hex 32 > ~/.oac/dev/core.key + printf '["%s"]\n' "$(tr -d '\n' < ~/.oac/dev/core.key | sha256sum | cut -d' ' -f1)" > ~/.oac/dev/core-key-digests.json + ``` -The SDK base URL is `http://127.0.0.1:8091/v1`. Requests require a bearer key. -Agents and Vault routes also require `OpenAI-Beta: agents=v1` (set by their SDK -resources); general Files routes do not. Supported operations include: +3. Start Core. `OAC_PUBLIC_URL` enables the Runtime gateway and the Worker; without it Core executes nothing. The [Core environment table](../../docs/configuration.md#appendix-core-environment-without-the-installer) lists every variable. -- Saved Agent create/retrieve/update/list/delete. -- Session create/retrieve/list/delete and metadata-only update. Creation supports inline - configuration or a saved `agent_id`, field replacements, initial text - and ordinary or streaming responses. Initial input is required for `none` and - streamed creation outside `self_hosted`. -- Session event submission and live streaming, Turn retrieve/list and Items list. -- Environment retrieve for three-harness colocated self-hosted and Docker - profiles, bounded live file listing, and inline/source copies into qualified - local workspaces. Shared Artifacts support capture, list/retrieve/content and - deletion independently of the live Runtime after publication. -- Project-owned `user_data` source file upload/list, metadata/content retrieval and - deletion; see [source Files](../../contracts/agents-api/source-files.md). -- Vault create/retrieve/list/delete, project-scoped pagination and stored status - filtering; static-bearer and OAuth Credential create/retrieve/list/replacement/delete, - plus [dispatch-time OAuth refresh](../../contracts/agents-api/vaults.md#oauth-refresh). Public archive semantics remain gaps. Already-delivered credentials - are not withdrawn by local deletion. Session attachments support - [authenticated HTTPS MCP](../../contracts/agents-api/vaults.md#credential-selection-in-a-session). + ```sh + OAC_DATABASE_URL='postgres://oac:…@127.0.0.1:5432/oac_dev' \ + OAC_CORE_KEY_DIGESTS_FILE="$HOME/.oac/dev/core-key-digests.json" \ + OAC_PUBLIC_URL=http://127.0.0.1:8091 \ + go run ./services/core/cmd/server + ``` -Execution uses the selected -[engine profile](../../contracts/agents-api/harness-capabilities.md), -including `none` and the colocated self-hosted profile described below. -Ordinary JSON requests have a 1 MiB body limit; file transfers use the separate -bounds in the Files contracts. Session lists support `after`, `limit` (0 is treated -as 1 and values above 100 as 100), -`order` (`asc`/`desc`) and optional immutable root `agent_id`. The local defaults -are 20 and descending order; exact hosted limits/error semantics remain unverified. -Session updates require the metadata field; null/empty clears it and an object -replaces supplied pairs. An empty update body rejects before resource lookup. +4. Create a Project and an API key with the Core key, as in [Script the Core API](../../docs/getting-started/operations.md#script-the-core-api), and call `/v1` with the key as in the [quickstart](../../docs/getting-started/quickstart.md). +5. To run `environment: none` Sessions, provision an [operator device profile](../../contracts/agents-api/machine-api.md#operator-device-profile) for the Project's tenant and start `oac-daemon connect --profile default` on a host with a Harness installed. Self-hosted Sessions use the [self-hosted guide](../../docs/getting-started/self-hosted.md) instead. -Delete with `client.beta.agents.sessions.delete(session.id)`. Only a durably idle -or failed Session without required actions or pending input can be deleted; a -queued, running or waiting root Turn or a pending input reservation returns 409 -`conflict_error` and leaves the Session unchanged. Subagent child Turns and -pending Environment file writes do not block deletion. Cancel its work -with an `agent.session.input.cancel` event, wait until it is idle, then delete it. -Confirmation means public removal: Session/history reads and new input become -unavailable, and existing streams close on observing removal. Creation keys stay -reserved; deletion never affects other Sessions, saved Agents or their shared -device. Internal records are retained for execution settlement. Managed Docker -deletion separately revokes authority and reclaims owned compute/workspace/history; -caller-managed compute is not reclaimed by this service. Repeating the deletion of -your own deleted Session returns the same confirmation, missing and foreign -Sessions return 404, and creation-key reuse returns 409; overlapping stream timing -is unverified. +## Tests -Codex command Items support live `agent.output.command_execution_output.delta` -events when emitted by the connected daemon. Queries retain accumulated drafts and -authoritative completion snapshots, including observed partial output after -cancellation. Native text conversion/output quotas apply; older peers may provide -only completion snapshots. Pinned native 0.153.4 may also omit early process -output from both notifications and its final aggregate; this remains an upstream -execution gap. Recover missed output with Items queries, not SSE replay. +`make check-core` builds Core and runs the Go tests of `services/core` and `packages/agents-client`. Point `OAC_TEST_DATABASE_URL` at a dedicated database named `oac_*_tests` that holds no other tables; the tests apply only Core's migrations and use fresh tenants without truncating anything. Without it, database tests skip locally; CI provides its own PostgreSQL. Set `OAC_TEST_OFFICIAL_SDK_PYTHON` to the interpreter with the pinned SDK for the store's client fixtures. [Validate a change](../../docs/development.md#validate-a-change) lists the focused checks for other areas. -Non-text message input, Subagents and installations beyond hosted initial files, -env, ordered setup and npm/Python packages remain unsupported. Saving optional Agent configuration does not make -it executable. Unsupported requests fail explicitly. `/healthz` reports liveness only. +### Official client verification -## Managed hosted execution +The official-client suite runs the built server against the pinned OpenAI SDK from [`upstream.json`](../../contracts/agents-api/upstream.json), with the same test database, temporary keys and fresh tenants, and needs no model provider: -The basic `openai_hosted` profiles for Codex, Claude Code and MiniMax Code require -explicit operator configuration. Select the qualified native image using the -[Codex](deploy/codex/README.md), [Claude Code](deploy/claude/README.md) or [MiniMax Code](deploy/mcode/README.md) guides, -then follow the [Docker setup](deploy/codex/README.md#standalone-operator-configuration). -Core-managed hosting supports deployment-selected E2B, Docker or microsandbox; -see the [nodes guide](../../docs/getting-started/nodes.md). For the separate -user-managed E2B path, see -[E2B Runtime packaging](deploy/e2b/README.md). -Core remains independently deployed with its own database. Public idle and initial -text Sessions share the existing preparation, execution, Files and recovery paths. -Networking defaults to enabled; disabled and exact-domain restricted policy are -supported after setup. System/npm/Python packages use the shared initializer; -remaining unsupported combinations are explicit gaps. Initial inline/file_id files, confidential env, npm/Python packages, -ordered setup and [public Environment Templates](../../contracts/agents-api/environments.md#templates) -resolve to the same immutable hosted configuration, independently of provider templates. Additional harnesses -require separate integration and qualification. -Connected describes the authenticated Runtime connection, not native readiness. -A provisioning failure fails the Session with a safe step and exit-status reason -([initialization failure](../../contracts/agents-api/environments.md#initialization-state-and-failure)); -other exact hosted failure and expiry semantics remain unverified. - -The independent Docker Provider consumes an immutable Runtime image and retains -one caller-owned allocation reference through partial creation and cleanup. Persist -that reference before Create and serialize its lifecycle; resolve a lost response -with observed state, without rewriting bootstrap credentials or replaying startup. -Provider state is compute state, not public Environment readiness. Named Runtime -volumes need explicit owned cleanup after container removal. A second mount of the -same workspace volume subdirectory provides the public `/workspace` path to native -tools; trusted staging and atomic rename retain the original parent mount. Do not -copy files or widen private-path reads to preserve an alias. Initialization command -timeouts can leave processes alive and require allocation cleanup before reuse. -The [managed Runtime build and operator configuration](deploy/codex/README.md#managed-runtime-image-and-docker-adapter) -defines the explicit opt-in for basic hosted admission. Building an image alone -does not qualify its isolation or enable public creation. - -## Internal execution device connection - -The standalone service can accept existing daemon connections without a Parsar -workspace or product database. Enable its internal gateway by setting -`OAC_PUBLIC_URL=https://your-service` (HTTP only for a loopback host during -local development). Core returns the derived -`wss://your-service/api/v1/agent-daemon/ws` as `self_hosted.remote_url`. It names -our private daemon transport, not stock OpenAI `exec-server` interoperability. - -After migrations, an operator can provision a device for an execution tenant: - -```bash -umask 077 -mkdir -p ~/.oac/daemon/default -go run ./services/core/cmd/device \ - --tenant '' --name 'local executor' \ - --url 'http://127.0.0.1:8091' \ - > ~/.oac/daemon/default/auth.json -oac-daemon connect --profile default -``` - -The command requires `OAC_DATABASE_URL` and emits a secret profile once. -Use a new profile rather than overwriting an existing device's credentials. When -provisioning remote compute, securely transfer this file to the same profile path -on the executor. The database stores only the credential digest. API keys and -device credentials are not interchangeable. To revoke a device: - -```bash -go run ./services/core/cmd/device \ - --tenant '' --revoke '' -``` - -An existing connection is retired on its next heartbeat; new connections are -rejected immediately. The internal Store binds each Session to one same-tenant -device, preserves that assignment across retries/restarts, and refuses a silent -move to another device. Revoked bindings cannot be used for dispatch. Device -connections alone do not start a Turn. Submit text, cancellation or function results -through the official Session events endpoint; the worker assigns a same-tenant host and preserves that -binding. Managed Docker has three-harness evidence. This generic device provisioning path -is for `none`; self-hosted Sessions require the dedicated enrollment below. -User-managed enrollment has a separate [qualification record](../../contracts/agents-api/harness-capabilities.md); complete protocol semantics remain partial. See the [ownership rules](../../AGENTS.md#public-api). - -### Enable Claude SDK execution - -Build and extract the runtime archive into a fresh managed directory on a matching -executor host. The archive contains the compiled bridge and pinned production -SDK/MCP/native dependencies; Node is installed separately. Linux x64/glibc with -Node22 is the accepted platform. See the -[runtime artifact contract](../../packages/claude-sdk-adapter/README.md#runtime-artifact) -for build outputs, version checks and platform restrictions. - -```bash -make build-claude-sdk-runtime -# After extracting the matching archive into this operator-chosen directory: -export OAC_RUNTIME_CLAUDE_SDK_ENTRYPOINT="$HOME/.oac/runtimes/claude-sdk/dist/main.js" -export OAC_RUNTIME_CLAUDE_SDK_NODE="/absolute/path/to/node" -oac-daemon connect --profile default -``` - -Set `OAC_DEFAULT_HARNESS=claude_sdk` on the API service. Configure provider access in -the daemon's private native SDK environment. Runtime readiness checks versions and -startup before daemon registration; it does not validate provider credentials. -SDK state stays under the daemon profile, independently of the replaceable bundle. -A ready SDK can start the daemon without a legacy CLI. Product `claude_code` remains -separate. Managed Node installation, runtime activation and registry publication -are not supplied by these commands. - -## Official client verification - -After preparing a dedicated test database, build the server and verify it with -the official client installed from the commit in `contracts/agents-api/upstream.json`: - -```bash +```sh python -m pip install -r services/core/tests/requirements.txt make build-core +OAC_TEST_DATABASE_URL='postgres://…/oac_local_tests' \ OAC_TEST_SERVER_BIN="${OAC_DEV_HOME:-$HOME/.oac}/build/oac-core/oac-core" \ python services/core/tests/official_client.py ``` -The test uses `OAC_TEST_DATABASE_URL`, temporary service keys and -fresh tenant IDs. The suite checks upstream and generated response schemas, retries, -ordering, tenant isolation, unsupported options and reads after a process restart, -without a model provider. It also runs the -[official Go client integration](../../packages/agents-client/README.md#go-client), using two -fresh tenants, and validates its created Sessions through the Python SDK. - -## Checks - -```bash -OAC_TEST_DATABASE_URL='postgres://.../oac_local_tests' \ - make check-core -``` - -Set `OAC_TEST_OFFICIAL_SDK_PYTHON` to the fixed SDK interpreter for the Store client -fixtures, and run the separate official-client command above as well. -`TestEnvironmentRetrievalOfficialClient` verifies public creation, scoped safe -Environment reads and retrieval after reopening without execution configuration. -`TestEnvironmentInitialFailureOfficialClient` verifies failed Session reads and -matching SDK/raw live failure events after privately provisioned initial input -expires, without fabricating a Turn or changing the Environment status. It is a -persistence prerequisite test. `TestSelfHostedInitialCreationOfficialClient` -separately exercises ordinary/streamed public initial creation through the Worker, -retry identity, disconnect survival and explicitly controlled deadline failure. -The test database must be named `oac_*_tests` and contain no product -workspace tables. Tests apply only this service's migrations and use new tenant -IDs without truncating tables. Missing test configuration skips DB tests locally; -the `OpenAgentCore checks` CI workflow always supplies its own PostgreSQL service. Run the -full `make check` before review as well. Product OpenAPI generation excludes this -service; its supported HTTP contract is generated separately. - -Run `make openapi` after handler annotation changes. It reuses the original -Core-only swaggo v1.16.4 generator, then splits the result by namespace: `/v1` -into `openapi.yaml`, `/core/v1` into `core.openapi.yaml` and `/api/v1` into -`runtime.openapi.yaml`; the last two use base path `/`, and each document keeps -only the security schemes its operations use. All generated schemas remain free -of product routes. - -## Public text execution - -With the daemon gateway enabled, the service owns one worker per execution database -and processes up to four Turns concurrently. Other workers are rejected by a -PostgreSQL advisory lock. A disconnected host leaves unsent work queued; clients -may cancel it. Session/Turn/Items queries expose durable results. - -```python -from openai import OpenAI - -client = OpenAI(base_url="http://127.0.0.1:8091/v1", api_key="") -session = client.beta.agents.sessions.create( - agent={"model": ""}, - environment={"type": "none"}, - input="Hello", - extra_headers={"Idempotency-Key": "first-message"}, -) -``` - -Configure model access in the engine host's native configuration. API tenant keys -and daemon credentials authenticate this service, not a model provider. Never put -provider secrets in Session metadata. This path does not enable Parsar Skill/SP -callbacks or bypass the pending product authorization work. - -Session creation admits initial text atomically; add `stream=True` for -created/live events. Open a GET event stream before submitting -later work, or use the official `sessions.stream` helper for one Turn. Function -handlers return results through the same public events endpoint. Recover missed -output with Session/Turn/Items queries; reconnecting SSE does not replay history. -See [creation streaming](../../contracts/agents-api/sessions-events.md#creation-streaming) -for retry behavior and unverified hosted timing. - -The [accepted workflows](../../contracts/agents-api/harness-capabilities.md) -include real MiniMax execution through built API/daemon/Codex and Claude SDK, -function success/error, cancellation and native continuation. Controlled fixtures -remain useful but do not replace real-provider acceptance for execution changes. - -## Historical Item storage - -Migration 15 records the retirement of private journal-to-Item backfilling. It -preserves public Items and source journals and refuses unindexed historical Turns. -This is historical schema evidence, not a supported upgrade procedure. Do not set -index markers manually, replay native execution, or run an older service to convert -a database for the current release. - -Historical installations are unsupported. Preserve their database and execution -history, and install the current release separately. Fresh database initialization -continues through the ordinary migration runner. Current recovery reads use -Session/Turn/Items; they do not provide SSE replay. - -## User-managed Runtime enrollment - -V1 uses our daemon as the user-side executor. Deploy daemon, selected harness, -local tools and protected `/workspace` together using the shared Runtime. For this caller-managed path, the user owns local or E2B allocation, renewal -and destruction; use the official E2B SDK through the -[E2B guide](deploy/e2b/README.md). Deployment-managed E2B, Docker and microsandbox -are separate hosted choices in the [nodes guide](../../docs/getting-started/nodes.md). - -Create a Session with `environment={"type":"self_hosted", -"workspace_directory":"/workspace"}` and empty/default capability directories. -Retain the returned Environment ID and unchanged `remote_url`. Initial text may -wait for connection; an idle Session creates no Turn. Codex, Claude SDK and MiniMax -reuse the same exact binding and `LocalEnvironment` preparation path. Service-origin -HTTP MCP is rejected on this placement; `none` MCP and separately qualified hosted -Environment Plugin MCP remain available within their own limits. - -An operator issues the Environment's connect-only executor credential with the -Core key, through Web or -`POST /core/v1/projects/{project_id}/environments/{environment_id}/executor-credentials` -([executor credentials](../../contracts/agents-api/environment-executor-credentials.md)), -and gives the returned credential file to the executor host. With direct database -access, the operator CLI can instead issue a principal executor key: - -```bash -umask 077 -oac-core-environment-key --tenant "$TENANT_ID" \ - --organization "$ORGANIZATION_ID" --project "$PROJECT_ID" \ - --subject-kind service_account --subject-id "$SUBJECT_ID" --key-id "$KEY_ID" \ - > "$HOME/.oac/executor-key.json" -``` - -The immutable principal must match a verified project mapping and the Session's -recorded creator. Add `--environment "$ENVIRONMENT_ID"` to restrict a key to one -live Environment. Keep the one-time `executor_token` output private. Only its digest -is stored; there is no secret read-back. `--rotate` and `--revoke` require the same -management ID and full principal; neither changes the key's restriction. Unknown -historical creators cannot enroll. API bearer keys and executor keys are separate. - -On the executor host, save that JSON as a private -`$OAC_RUNTIME_HOME/daemon/executor-key.json`, outside the tool workspace. Use the [native installer](../../docs/getting-started/self-hosted.md) to prepare the selected -Harnesses. Managed images preinstall them. A preconfigured Runtime can connect -with the values returned by Session creation: - -```bash -oac-daemon connect --remote "$REMOTE_URL" \ - --environment-id "$ENVIRONMENT_ID" \ - --credential-file "$OAC_RUNTIME_HOME/daemon/executor-key.json" -``` - -Use TLS outside loopback. Enrollment supplies the trusted Session identity; do -not inject an unrelated `OAC_RUNTIME_SESSION_ID`. The daemon persists its -immutable Environment binding beside the key and refuses to adopt another -Environment's existing native history. Rotation keeps the same key ID: update the -protected file, then restart the daemon to authenticate with the new token. -WebSocket authentication uses the `Authorization` header, never a URL token. -On a permanent rejection (enrollment 401 or 409, a permanent WebSocket rejection -or close) `connect --environment-id` prints one message naming the fix, makes no -further requests and exits 0 on SIGTERM or SIGINT; transient failures exit 1. - -The private `POST /api/v1/agent-daemon/enroll` endpoint accepts that executor bearer -and `{"environment_id":"..."}`. It returns `device_id`, `session_id`, -`environment_id` and `workspace_directory`, never another credential. Enrollment -atomically binds one dedicated device/key to the Session; retries retain it and -conflicting bindings reject. No managed `runtime_allocation` is created. Runtime -onboarding connects to the returned `remote_url` using that exact binding and key. -The endpoint is outside the pinned public Agents API, which remains unchanged. -It is not stock `exec-server`, `/cloud/environment/*` or Noise interoperability; -there is no service-side harness or remote tool forwarding compatibility path. - -The gateway and Worker recheck current credential authority. Rotation/revocation, -Session deletion and lost ownership deny further use; a replacement connection -cannot overwrite a successor's observations. Connection events report authenticated -connectivity, not native preparation readiness. Retained native history and workspace -must survive a Runtime restart for continuation; missing history fails closed. -Neither disconnect nor deletion promises immediate native process quiescence or -reclaims user-owned E2B/local compute. The user must stop and destroy it explicitly. - -Pending input retains its durable identity/deadline through HTTP disconnects. Later -idle input returns 202 after preparation/admission, not after model completion; -use client/proxy timeouts above five minutes and recover progress through events -and reads. Exact upstream failure/error timing remains unverified. - -The Runtime's shared Go workspace implementation owns directory reads, writes and -output export; Harness adapters use the same authorized workspace binding. -The [qualification record](../../contracts/agents-api/harness-capabilities.md) -identifies fixed-SDK/raw HTTP, real-model, Files/Artifacts, cancellation, -restart/history and credential lifecycle evidence, with its recorded revision limits. - - -### HTTP MCP execution - -Public `agent.tools` declares an explicit MCP connection origin. Omitted/null -origin remains `service`. The [origin, credential and Harness matrix](../../contracts/agents-api/environments.md#public-mcp-connection-origin) -owns the supported combinations: Codex/Claude service HTTP on `none`, and -Codex/Claude/MiniMax Environment HTTP on managed or user-owned workspaces, subject -to declared native limits. Public declarations and installed Plugin MCP converge -on the same Runtime effective bindings; they retain distinct credential authority. - -Service-origin MCP runs on trusted service-side compute. Codex supports `environment:{"type":"none"}`; Claude SDK supports HTTP MCP with -`environment:{"type":"none"}`. Inline or saved Agent tools may declare: - -```json -{ - "type": "mcp", - "server_label": "tickets", - "transport": {"type": "http", "server_url": "https://mcp.example.com/mcp"}, - "connection_origin": "service", - "allowed_tools": ["lookup_ticket"], - "required": false -} -``` - -An omitted/null `allowed_tools` permits all tools from that server; `[]` permits -none. Native discovery may still connect to a declared server. The daemon must -advertise `mcp_http_tools`; selection waits for a capable device. The native -harness owns MCP discovery, calls and results. Public `mcp_call` Items use original -server/tool names; recover missed live events through Session, Turn and Items reads. - -With Codex, set `required:true` to require initialization before the first native Turn. It -defaults to false and additionally requires the pinned daemon's `mcp_http_required` -capability. Native root thread creation and cold resume wait for required servers; -initialization failure stops execution without replacing retained history. Public -work can already be accepted or queued during this wait. Exact hosted creation -timing/errors and continuing MCP health monitoring remain unverified. - -On `environment:none`, both Codex and Claude SDK support tenant-owned `vault_ids` -for static-bearer HTTPS MCP. `self_hosted` is explicitly unsupported for -service-origin MCP in V1, including anonymous requests. -An explicit -`credential_id` selects an attached credential for the exact HTTPS URL; omission/null -selects a unique matching credential, or stays anonymous if none matches. Ambiguity -fails before Session creation with 409 `conflict_error`. Selection is frozen -privately; Session reads and events show an implicitly selected credential ID in the -public tool, while the stored request keeps the caller's field. See -[credential setup and limits](../../contracts/agents-api/vaults.md). -Authenticated execution additionally requires `mcp_http_bearer_auth`; missing keys -or failed authorization/decryption never fall back to anonymous execution. - -The colocated V1 `self_hosted` profile does not admit service-origin MCP. The old -separate executor/service-side MCP combination and its remote capability gates are -retired. Do not forward Vault credentials to user-owned Runtime compute. - -Claude SDK requires a packaged runtime that reports `mcp_http_tools`; the SDK -version alone does not qualify an older bundle. Its current profile -accepts either `required` value, requires connected servers and static inventories. Server labels -accept ASCII letters, digits, underscore and hyphen, except reserved `functions`; -selected tool names additionally accept dots. Declared HTTP MCP tools compose with -host functions; undeclared servers, built-ins and subagents remain disabled. -Authenticated requests also require `mcp_http_bearer_auth`. The existing Vault -selection rules apply, including implicit exact-URL matches and immutable private -bindings. Per-server/per-launch environment references keep bearer values out of -native argv and stored state. A Vault with no matching credential may stay anonymous. -Anonymous requests suppress native OAuth/credential injection with a blank -Authorization header, without deleting native state. Servers rejecting that header, normalized -name collisions, changing inventories and original MCP metadata fidelity remain -gaps. Items retain the observed native JSON, which may differ from the original -MCP envelope. See the [Claude SDK profile](../../packages/claude-sdk-adapter/README.md#http-mcp). - -The current subset rejects native OAuth login, inline authorization, nonempty headers or -request metadata, URL userinfo/query/fragment, the `environment` origin, stdio -and engines other than Codex/Claude SDK. The Codex adapter also -rejects reserved native labels and stored native MCP credentials. It verifies -exact effective MCP configuration before starting/resuming a native thread, -excludes undeclared servers and disables native apps/plugins. This runs on trusted -service compute; it does not provide filesystem isolation or guard against -concurrent operator configuration mutation. These limits are implementation gaps, -not changes to the pinned official protocol. - -A dedicated local Runtime uses one Environment-scoped device credential and an -immutable binding to that Environment's Session. It is excluded from general -device selection; another Session cannot claim it, including within the same -tenant. Deleting its Session invalidates credential lookup and heartbeat renewal. -Provision a new scoped device atomically rather than widening an existing shared -device credential. Revocation does not authorize silent placement replacement. - -The private local Environment reference contains its identity and, for policy-aware -execution, its immutable network policy. Trusted Runtime deployment configuration -freezes the Environment, Session and workspace root; -requests cannot supply a replacement root. The V1 path uses only the exact local -Environment reference. Use the same preparation/start lifecycle for native execution and the -existing bounded workspace controls for directory access. Local idle directory -reads use the common Go filesystem implementation without a temporary Harness. These private capabilities alone do not authorize public requests or establish -Provider lifecycle. Self-hosted enrollment supplies the exact local binding. -Core rechecks the persisted Environment/device binding for preparation and active -reads; capability discovery cannot select or authorize a general device for this -placement. Local work uses the existing pending-input reservation and Worker -ownership without a remote connection resolver. The hosted profile supports `network.access: enabled`, `disabled` and exact-host -`restricted`; each image must qualify the supported policies before public deployment. Omitted -network settings mean enabled upstream and must not be silently treated as disabled. - -User-managed Runtime enrollment authenticates the existing principal executor key -against the exact live Session/Environment and its recorded creator. Keys retain -stable management IDs, immutable principals, optional exact-Environment restrictions, -rotation/revocation and digest-only storage. They grant connection authority, never -Session API access. No raw-token import or secret read-back is added. - -Enrollment atomically creates or recovers one dedicated device and immutable Session -binding under the Session lock. The frozen workspace and capability directory selection come from the Session -configuration and must match the local binding. Runtime snapshots declared local -Skill or Plugin directories before creating the native executor. No `runtime_allocation` is -created for user-owned compute. A retry cannot replace a device, change its bound -key or adopt another native history. Gateway authentication and dispatch recheck -current key authority; rotation/revocation and deletion deny further use. - -The daemon, selected harness, local tools and workspace run together. The Dispatcher -passes the existing typed `LocalEnvironment` after exact tenant/Session/Environment/ -device checks. Native preparation, Files and Artifacts reuse the same protected -local workspace and existing lifecycle owners. There is no registry/Noise relay, -transient harness credential, service-side harness or remote tool forwarding path. -Keep model credentials, daemon authorization and native histories private; connection -success alone establishes neither native readiness nor filesystem isolation. - -## Hosted sandbox nodes +It checks the upstream and generated response schemas, retries, ordering, tenant isolation, unsupported options and reads after a restart. It also runs the [official Go client](../../packages/agents-client/README.md#go-client) against two fresh tenants and validates the Sessions it created through the Python SDK. -The release includes `oac-node` for local and remote hosts. Add nodes with -Web's one-command flow in the [nodes guide](../../docs/getting-started/nodes.md), which -also covers provider selection, manual registration and reset. +## Generated contracts -Implementation rules for hosted sandbox nodes and optional suspension are in -[Agents API implementation constraints](IMPLEMENTATION.md#hosted-sandbox-nodes-and-optional-suspension). +Handler annotations generate the OpenAPI documents. Run `make openapi` after changing them: it runs the pinned swaggo generator and splits the result by namespace into [`openapi.yaml`](../../contracts/agents-api/openapi.yaml) (`/v1`), [`core.openapi.yaml`](../../contracts/agents-api/core.openapi.yaml) (`/core/v1`) and [`runtime.openapi.yaml`](../../contracts/agents-api/runtime.openapi.yaml) (`/api/v1`), each with only the security schemes its operations use. Review and commit the three diffs. Contract tests hold `openapi.yaml` to the pinned routes and fields, and hold all three to exactly the routes the server registers. diff --git a/services/core/deploy/codex/README.md b/services/core/deploy/codex/README.md index e490b2370..b4044e418 100644 --- a/services/core/deploy/codex/README.md +++ b/services/core/deploy/codex/README.md @@ -14,6 +14,30 @@ Codex accepts only the `responses` model protocol ([`harnessconfig/codex`](../.. Each Session has its own `CODEX_HOME` at `$OAC_RUNTIME_HOME/daemon/agent-sessions//` (`OAC_RUNTIME_HOME` defaults to `~/.oac`). Native history stays there. The adapter regenerates that directory's `config.toml` on every prompt: the Session's frozen provider bundle becomes the `[model_providers.oac]` block, with `wire_api` `responses`, and the thread is pinned to that provider, so Codex never falls back to its built-in `openai` provider. +## Execution controls + +The adapter applies Core's typed controls and opt-ins, described in the [Runtime protocol](../../../../docs/runtime-protocol.md#capability-declarations), through native Codex settings on both new and resumed Turns. Explicit controls take precedence over adapter options without changing them. + +- **`environment: none`.** The daemon forces `CODEX_EXEC_SERVER_URL=none` after the caller's environment options and confirms that Codex reports its `local` and `remote` environments as unknown before it starts or resumes a thread; a binary that cannot do this fails closed. The engine process runs on the bound device, which is not a user execution environment. +- **Web search.** The control becomes Codex's `web_search` option (`disabled`, `cached` or `live`); Core sends `disabled`. +- **Programmatic tool calling.** An explicit disable turns off the native `code_mode`, `code_mode_only` and `code_mode_prewarm` features and checks managed requirements before a thread starts or resumes, rejecting a conflicting requirement. +- **Text verbosity.** The adapter reads the model catalog with `codex debug models`, checks the model's support and pins that catalog snapshot for the execution. The probe needs Unix process-group cancellation, so other hosts do not declare `text_verbosity`. For a model without declared verbosity support, including Codex's unknown-model fallback, `medium` omits the override and keeps native default text; a supported model receives an explicit `medium`. Unsupported `low` or `high` and an unreadable catalog fail before model execution. +- **Subagents.** Disabling Subagents turns off both native multi-agent feature generations, `multi_agent` and `multi_agent_v2`, overriding the operator's feature preferences. + +## MCP servers + +A typed MCP declaration replaces operator MCP options. The adapter renders it with the native renderer and the original tool names for `enabled_tools`, including `[]`. Before creating or resuming a thread it reads native `config/read` with the exact cwd and rejects additional servers or any difference in the effective configuration. It disables native plugins and apps, selects file-only MCP credentials, and rejects existing credentials in the private native home without deleting them or native history. Reserved native labels are an adapter restriction, not a rule of the saved resource. The check is a snapshot, not a barrier against concurrent operator configuration changes, and discovery of a deny-all server can still contact it. + +With `required: true`, which needs `mcp_http_required`, root thread creation and cold resume wait for required servers to initialize and send no native Turn until they do; a failed strict resume is never replaced with a new thread. Public work may already be accepted while initialization waits. + +A selected bearer credential (`mcp_http_bearer_auth`) becomes a fresh daemon-owned `bearer_token_env_var` reference for each server and native process. The secret enters only that app-server child's environment, after auxiliary launch probes, and never global environment, arguments, configuration, history, snapshots or logs. Preflight accepts only the expected server and reference pairing and keeps rejecting ambient credential sources. Empty values and bytes outside RFC 6750 `b64token` syntax are rejected with generic errors, and tokens are never trimmed. Core-managed OAuth delivers access tokens through the same path. + +## Function results and command output + +A function result counts as applied only when Codex reports a matching live dynamic-tool completion: root thread, Turn and call identity, function, success and ordered content. Writing the JSON-RPC response is not application. The adapter owns pending receipts without holding their lock across I/O or waits; terminal state, cancellation and native loss settle unconfirmed submissions before release, and an uncertain receipt timeout ends that native execution without resending the result. It records native confirmation before any potentially blocking observation publication, so output backpressure cannot turn a known application into an unknown one. + +With neutral tool observations, the adapter emits `command_output` fragments carrying the native command identity, filtered to the root thread and Turn. Codex `0.153.4` can miss early process output in both its notifications and its final aggregate; that output is lost, and Core never reconstructs it from model text. + ## Native failure classification The adapter classifies a failed Turn only from the exact root terminal Turn's `codexErrorInfo` ([`error_classification.go`](../../../../apps/daemon/internal/agent/codex/error_classification.go)); error notifications, including retry notifications, never classify it. The [Runtime protocol](../../../../docs/runtime-protocol.md#native-failure-classification) defines the codes. From 76fac78c6ad4598299ddb88c16a1c9a73b5ba225 Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 08:57:34 +0000 Subject: [PATCH 07/10] docs: point E2B catalog operations at generic configuration discovery --- services/core/tools/e2b-provider/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/services/core/tools/e2b-provider/README.md b/services/core/tools/e2b-provider/README.md index 75dae4699..d04abcfd7 100644 --- a/services/core/tools/e2b-provider/README.md +++ b/services/core/tools/e2b-provider/README.md @@ -18,7 +18,7 @@ The account key is stored encrypted in Core's database and is write-only. It rea | `kill` | `Kill` | Destroys every matching sandbox and confirms that none remains | | `command` | `RunCommand` | Runs one bounded command as the Runtime user on a running sandbox whose bootstrap completed; output is limited to 1 MiB per stream | | `validate_deployment` | Deployment setup | Reads the template's builds and requires the exact build to be ready with the configured CPU and memory. Without configured resources the selection adopts the build's CPU and memory. Returns the build's status, CPU, memory and reported disk for Core to record; bounded to 30 seconds | -| `list_templates`, `list_builds` | Web's setup wizard | Pages the key's visible templates (`GET /v2/templates`) or one template's ready builds, with a transient key. Results are capped at 200 and write no receipt | +| `list_templates`, `list_builds` | [Configuration discovery](../../../../contracts/agents-api/sandbox-deployment.md#configuration-discovery) | Pages the key's visible templates (`GET /v2/templates`) or one template's ready builds, with a transient key. Results are capped at 200 and write no receipt | | `observe` | Runtime observations | Up to 100 allocations; see [Observations](#observations) | | `verify_credential` | E2B key replacement | Up to 32 allocation references; see [Credential verification](#credential-verification) | From dd87c1f4d71f202d8a72657f5a174e76327e22bd Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 09:12:39 +0000 Subject: [PATCH 08/10] docs: correct node operation shapes and status constraints --- .../agents-api/node-generation-protocol.md | 30 ++++++++++++------- docs/sandbox-provider.md | 4 ++- services/core/IMPLEMENTATION.md | 4 +-- 3 files changed, 25 insertions(+), 13 deletions(-) diff --git a/contracts/agents-api/node-generation-protocol.md b/contracts/agents-api/node-generation-protocol.md index debc38eff..09d84fd2a 100644 --- a/contracts/agents-api/node-generation-protocol.md +++ b/contracts/agents-api/node-generation-protocol.md @@ -31,21 +31,31 @@ Core sends `request` frames with: | `timeout_ms` | Remaining budget, 1 to 120000 | | `reference` | The exact `(tenant_id, environment_id, allocation_id)` | -Each operation carries exactly its own argument: - -| `operation` | Provider method | Argument | -| --- | --- | --- | -| `create` | `Create` | `bootstrap` | -| `info`, `renew`, `kill`, `command` | `GetInfo`, `Renew`, `Kill`, `RunCommand` | none, or `command` | -| `observe` | `Observe` | `observation` | -| `initial`, `new_compute`, `compute`, `kill_compute`, `resume_compute`, `command_compute` | Checkpoint compute operations | none, an optional `snapshot`, `compute`, or `compute` and `command` | -| `suspend`, `resume`, `delete_snapshot` | `Suspend`, `Resume`, `DeleteSnapshot` | `suspend`, `resume` or `snapshot` | +Each operation carries its own arguments and returns the following result on success: + +| `operation` | Provider method | Arguments | Successful result | +| --- | --- | --- | --- | +| `create` | `Create` | `bootstrap` | `info` | +| `info` | `GetInfo` | None | `info` | +| `renew` | `Renew` | None | `info` | +| `kill` | `Kill` | None | None | +| `command` | `RunCommand` | `command` | `command` | +| `observe` | `Observe` | `observation` | `sample` | +| `initial` | `Initial` | None | `compute` | +| `new_compute` | `NewCompute` | Positive compute `generation` and optional `snapshot` | `compute` | +| `compute` | `GetCompute` | `compute` | `state` | +| `kill_compute` | `KillCompute` | `compute` | None | +| `resume_compute` | `ResumeCompute` | `compute` | `state` | +| `command_compute` | `RunCommandCompute` | `compute` and `command` | `command` | +| `suspend` | `Suspend` | `suspend` | `state` | +| `resume` | `Resume` | `resume` | `state` | +| `delete_snapshot` | `DeleteSnapshot` | `snapshot` | None | A request whose `connection_id`, `owner_epoch` or `sequence` does not match closes the connection. A malformed request gets an `invalid` response. A node without generation management accepts only its enrolled `deployment_generation`; a generation-managing node runs the request on that generation's provider and answers `unconfirmed` when it cannot. Core sends `create` and a `resume` that is not observe-only only to a generation that is ready on that node, and keeps at most 32 requests pending per connection. The budget is relative: the node anchors `timeout_ms` to its own clock on receipt and consumes it while the request waits in its queue, so the hosts' clocks need not agree. Core still bounds its own wait. A full node queue closes the connection. -The `response` frame carries `id`, `connection_id`, an `error_code` when the call failed, and on success exactly one result (`info`, `compute`, `state`, `command` or `sample`): +The `response` frame carries `id` and `connection_id`. A successful response carries the result named in the operation table, with no result field for `kill`, `kill_compute` or `delete_snapshot`. A failed response carries an `error_code`: | `error_code` | Meaning | | --- | --- | diff --git a/docs/sandbox-provider.md b/docs/sandbox-provider.md index af8cec2f4..9b32fe2b5 100644 --- a/docs/sandbox-provider.md +++ b/docs/sandbox-provider.md @@ -111,7 +111,9 @@ A new provider takes these steps: 2. Add its specification and resource validators and, for native resource discovery before commit, an optional read-only `SelectionDiscoverer`. Put native credential verification behind `CredentialVerifier`. 3. Implement `sandbox.ConfigurationAdapter` over a typed native configuration. `DecodeInput` strictly parses the separate public `configuration` and write-only `credential` objects of a request. `Encode` produces whitelisted public selectors, read-only observations and separate secret bytes, and never passes request JSON through. `Decode` restores stored selectors, and keeps access to owned resources, without remote admission or new template validation. `Normalize` copies its input before changing it. `ResolveChange`, `Equal` and `WithCredential` own inheritance, identity and credential composition. `Requirements` declares whether a credential and a public Core origin are required and whether configuration discovery is supported. Also implement `ConfigurationDiscoverer`, even when discovery is unsupported: it validates the query and returns a safe catalog, never a mutation or an admission decision, while Core keeps authorization, input limits and deadlines. A node provider accepts only an empty public object, rejects credentials and returns Unsupported for discovery and credential replacement. 4. Register its constructor, policies, configuration adapter, operation declaration and defaults in `providers/registry.go`. Node proxy identity and checkpoint support read this entry. The installer's projection combines the registered policies with the shared field bounds in `sandbox/deployment_contract.go`; regenerate it with `go run ./services/core/cmd/specification-contract -write`. -5. Supply the distribution artifacts for the adapter and its helper, and offer the provider to operators. Today the installer's `--sandbox` choices and Web's setup views carry provider-specific options, such as E2B's installer flags and Web views, so this step edits the shared installer and Web code. Never add a Session or Turn scheduling path, a vendor column or API field, or a vendor switch in the store. +5. Supply the distribution artifacts for the adapter and its helper, and offer the provider to operators through the registered configuration contract. + +**Known design gap:** the installer's `--sandbox` choices and Web's setup views carry provider-specific options, such as E2B's installer flags and Web views. Exposing another provider through these surfaces currently requires shared installer and Web edits. This coupling does not meet [Complexity stays in the adapter](../AGENTS.md#complexity-stays-in-the-adapter); new integrations must express their configuration through the protocol and keep vendor-specific behavior in the adapter. Never add a Session or Turn scheduling path, a vendor column or API field, or a vendor switch in the store. ### Registration validation diff --git a/services/core/IMPLEMENTATION.md b/services/core/IMPLEMENTATION.md index ca0f1a5d6..8daf9ab3a 100644 --- a/services/core/IMPLEMENTATION.md +++ b/services/core/IMPLEMENTATION.md @@ -58,11 +58,11 @@ New saved-reference Sessions, and inline requests with Vault attachments or cred Session listing by `agent_id` filters on the immutable root `configuration.agent.id` with the tenant, Agent and creation index, before pagination and activity projection. -Session state and last activity derive from the latest persisted Turn: queued or active work is `in_progress`, successful or cancelled work `idle`, and failures carry a safe public error. +Session status and last activity use the public projection in [`internal/api/session_response.go`](internal/api/session_response.go); [Session diagnostics](../../contracts/agents-api/session-diagnostics.md#session-snapshot) describes its failure precedence. ## Agents and model providers -Reusable Agents are tenant-scoped rows independent of Session snapshots and engine bindings. The store persists caller-validated configuration without applying harness restrictions or model defaults, with internal limits of 512 KiB for configuration and 64 KiB for metadata. An update locks the Agent row while merging the supplied fields and enforcing the configuration bound, then commits configuration, metadata and update time together, so a stale full snapshot never overwrites another update; an update without fields changes no timestamp. Deletion is one tenant-scoped `DELETE … RETURNING id`. A Session copies the saved configuration into its immutable snapshot and never looks up its source again. +Reusable Agents are tenant-scoped rows independent of Session snapshots and engine bindings. The store persists caller-validated configuration without applying harness restrictions or model defaults, with internal limits of 512 KiB for configuration and 64 KiB for metadata. An update locks the Agent row while merging the supplied fields and enforcing the configuration bound, then commits configuration, metadata and update time together, so a stale full snapshot never overwrites another update. An empty update preserves the saved fields and advances `updated_at` through the same SQL update. Deletion is one tenant-scoped `DELETE … RETURNING id`. A Session copies the saved configuration into its immutable snapshot and never looks up its source again. Saved execution defaults keep a model-provider bundle whole at every replacement boundary: endpoint, key, protocol and limits are never inherited separately. Agent JSON holds only the safe provider fields and an output-only configured flag; the complete bundle is encrypted separately with a tenant and Agent binding and its own purpose, and configuration and secret changes commit together under the Agent row lock. Model-only edits need no key. Merged harness, protocol and limits are validated without reading keys. Session creation reads safe defaults and ciphertext in one snapshot, and a complete Session override does not decrypt the inherited bundle. The resolved bundle is frozen in an encrypted Session-owned row, and dispatch fails closed when that snapshot is missing or cannot be decrypted; later Agent edits, default changes, restarts and suspension never resolve it again. From b02b52f4050f1d5907f797f53dde526edcdf39f9 Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 09:17:51 +0000 Subject: [PATCH 09/10] docs: reconcile runtime links with execution contract owners --- contracts/agents-api/harness-onboarding.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/contracts/agents-api/harness-onboarding.md b/contracts/agents-api/harness-onboarding.md index 34f94ff68..19b42f6da 100644 --- a/contracts/agents-api/harness-onboarding.md +++ b/contracts/agents-api/harness-onboarding.md @@ -154,7 +154,7 @@ Registration is static and requires a build. The methods live in `agent/harness. The direct-call `agent.Factory` delegates to the same Executor implementation. -Every `proto.AgentKindCapabilities` field must be explicitly `proto.CapabilitySupported` or `proto.CapabilityUnsupported`, even for an unavailable Harness. `proto.CapabilityUnspecified` is invalid: zero values and omitted fields never mean Unsupported. An installation probe may set an individual field with `proto.CapabilityFromBool`; it must not populate unmentioned or future fields. Availability stays separate in `SupportedAgentKind.Available`. Registration validates the complete declaration before changing the registry, and the wire carries an explicit boolean for every field, so omitted and null fields are invalid. A new field requires a decision in every production declaration. Runtime consumers use `IsSupported()` and reject unsupported requests before native operations; an interface assertion verifies implementation, never support. Every declaration must match the behavior verified for that installation; the [Core–Runtime protocol](../../docs/runtime-protocol.md#explicit-capability-declarations) owns how declarations travel and are frozen. +Every `proto.AgentKindCapabilities` field must be explicitly `proto.CapabilitySupported` or `proto.CapabilityUnsupported`, even for an unavailable Harness. `proto.CapabilityUnspecified` is invalid: zero values and omitted fields never mean Unsupported. An installation probe may set an individual field with `proto.CapabilityFromBool`; it must not populate unmentioned or future fields. Availability stays separate in `SupportedAgentKind.Available`. Registration validates the complete declaration before changing the registry, and the wire carries an explicit boolean for every field, so omitted and null fields are invalid. A new field requires a decision in every production declaration. Runtime consumers use `IsSupported()` and reject unsupported requests before native operations; an interface assertion verifies implementation, never support. Every declaration must match the behavior verified for that installation; the [Core–Runtime protocol](../../docs/runtime-protocol.md#capability-declarations) owns how declarations travel and are frozen. The admission mapping is explicit. `Steering` controls non-durable `Steerer` input. `DurableInputReceipts` controls `DurableSteerer` input and also requires the Turn settlement contract; neither implies the other, and Core's public text profile requires both. `Permissions` qualifies permission and user-choice responses together and requires both native response paths. Workspace declarations describe the authorized resource owner, including the common Runtime workspace implementation. Runtime registration does not grant Core qualification; the service profile does. From 3bb2f375057d46804e3f0a5b3fb628f1c7528747 Mon Sep 17 00:00:00 2001 From: SaladDay <1203511142@qq.com> Date: Wed, 30 Sep 2026 09:22:46 +0000 Subject: [PATCH 10/10] docs: integrate API owners and retain implementation mechanisms --- contracts/agents-api/admin-api.md | 10 +-- contracts/agents-api/core-errors.md | 4 +- .../agents-api/runtime-observability-api.md | 2 +- contracts/agents-api/sandbox-deployment.md | 2 +- docs/runtime-protocol.md | 2 +- docs/sandbox-provider.md | 2 +- packages/claude-sdk-adapter/README.md | 2 +- scripts/name-allowlist.json | 15 ---- services/core/IMPLEMENTATION.md | 69 +++++++++---------- 9 files changed, 43 insertions(+), 65 deletions(-) diff --git a/contracts/agents-api/admin-api.md b/contracts/agents-api/admin-api.md index 573db0ff7..c9965cdf9 100644 --- a/contracts/agents-api/admin-api.md +++ b/contracts/agents-api/admin-api.md @@ -26,13 +26,13 @@ Paths are relative to `/core/v1`. | `projects/{project_id}/sessions/{session_id}/execution-configuration` | The Session's frozen model, Harness and provider selection | [Execution configuration](#execution-configuration) | | `projects/{project_id}/sessions/{session_id}/diagnostics`, `…/turns/{turn_id}/diagnostics` | Failure categories and Item receipt timing | [Session diagnostics](session-diagnostics.md) | | `projects/{project_id}/sessions/{session_id}/runtime-observation`, `sandbox/runtime-observations` | Current Runtime observations | [Runtime observations](runtime-observability-api.md), [the list's disk field](#runtime-observations) | -| `projects/{project_id}/sessions/{session_id}/runtime-history` | Stored Runtime history | [Runtime history](runtime-history-api.md) | +| `projects/{project_id}/sessions/{session_id}/runtime-history` | Stored Runtime history | [Runtime history](runtime-observability-api.md#session-runtime-history) | | `projects/{project_id}/environments/{environment_id}/installation` | Install commands of a `self_hosted` Environment | [Installation grant](environment-executor-credentials.md#installation-grant) | | `projects/{project_id}/environments/{environment_id}/executor-credentials[/{key_id}]` | Executor credentials of a `self_hosted` Environment | [Executor credentials](environment-executor-credentials.md#core-key-routes) | | `projects/{project_id}/resource-owners`, `projects/{project_id}/write-operations` | Which API key created a resource and each key's writes | [Write provenance](#write-provenance) | | `harnesses`, `harnesses/{harness}/model-configuration` | Enabled Harnesses and each Harness's deployment default model | [Deployment defaults](model-execution.md#deployment-defaults) | -| `sandbox/deployment`, `sandbox/deployment/reset`, `sandbox/providers/{provider}/discovery` | The sandbox provider, resources and Runtime, reset, and provider configuration discovery such as E2B templates | [Sandbox deployment](sandbox-deployment.md#authority-and-routes) | -| `sandbox/enrollment-tokens`, `sandbox/nodes[/{node_id}[/allocations]]` | Node enrollment tokens, nodes and their allocations and host history | [Nodes guide](../../docs/getting-started/nodes.md), [sandbox deployment](sandbox-deployment.md), [node host history](node-host-history.md) | +| `sandbox/deployment`, `sandbox/deployment/reset`, `sandbox/providers/{provider}/discovery` | The sandbox provider, resources and Runtime, reset, and provider configuration discovery such as E2B templates | [Sandbox deployment](sandbox-deployment.md#routes) | +| `sandbox/enrollment-tokens`, `sandbox/nodes[/{node_id}[/allocations]]` | Node enrollment tokens, nodes and their allocations and host history | [Nodes guide](../../docs/getting-started/nodes.md), [sandbox deployment](sandbox-deployment.md), [node host history](runtime-observability-api.md#node-host-observations-and-history) | | `summary` | Session counts and usage by Project, Agent or key | [Summary](#summary) | | `metrics` | Core's own process, execution, database and job metrics | [Core metrics](core-metrics.md) | | `audit-log` | Administrator writes | [Audit log](#audit-log) | @@ -91,7 +91,7 @@ One transaction marks the Environment expired (a failed Environment stays failed `POST` and `GET /projects/{project_id}/sessions/{session_id}/archive` return `{session_id, environment_id, state}`. `GET` is read-only and needs no generation. `state` is the resource's current disposition: `active`, `cleanup_pending` or `released`, whatever released it. `released` does not mean an active Turn has finished cancelling; read the Turn for that. -After an uncertain `POST` response, `GET` the archive before writing again. Repeating the `POST` has the same effect and records one audit entry per accepted request. To clear every hosted Session before changing the deployment, use the [deployment reset](sandbox-deployment.md#initialization-same-provider-changes-and-reset). +After an uncertain `POST` response, `GET` the archive before writing again. Repeating the `POST` has the same effect and records one audit entry per accepted request. To clear every hosted Session before changing the deployment, use the [deployment reset](sandbox-deployment.md#reset). ## Execution configuration @@ -214,7 +214,7 @@ The response is `{data, has_more, next_cursor}`. Each row has `project_id`, null ## Runtime observations -`GET /sandbox/runtime-observations` lists the current Runtime observation of every Session in every Project, as `{object: "list", data: [{project_id, observation}], has_more, first_id, last_id}` ([Runtime observations](runtime-observability-api.md)). It samples read-only and never provisions compute; a provider with a batch metrics read, such as E2B, samples the page in one bounded request. Each list `observation` adds `disk: {usage_bytes, limit_bytes}`, with the null rules of `memory`: E2B reports its sandbox disk usage and capacity, and Docker and microsandbox return null. The per-Session read has no `disk`. +The [Runtime telemetry API](runtime-observability-api.md) owns current observations, the [list-only disk field](runtime-observability-api.md#disk), Session history and node host observations and history. ## Audit log diff --git a/contracts/agents-api/core-errors.md b/contracts/agents-api/core-errors.md index 12ee8db41..a07923361 100644 --- a/contracts/agents-api/core-errors.md +++ b/contracts/agents-api/core-errors.md @@ -36,7 +36,7 @@ These have null `param` and no `details`. A Core `401 invalid_admin_key` therefo ## Sandbox provider verification -A `POST` or `PUT /core/v1/sandbox/deployment` ([sandbox deployment](sandbox-deployment.md#initialization-same-provider-changes-and-reset)) whose provider verifies a credential or configuration, as E2B does, fails with these fixed errors. None returns provider text, a template name, a key or a resource count. +A `POST` or `PUT /core/v1/sandbox/deployment` ([sandbox deployment](sandbox-deployment.md#reset)) whose provider verifies a credential or configuration, as E2B does, fails with these fixed errors. None returns provider text, a template name, a key or a resource count. | HTTP | Code | Meaning | `param` | | --- | --- | --- | --- | @@ -103,4 +103,4 @@ The [Session and Turn diagnostics reads](session-diagnostics.md) return these ca When the diagnostics reader is not configured, the reads return 503 `diagnostics_unavailable` without details. A database failure is an error, never an empty or healthy snapshot. Provisioning reasons and native messages are never parsed for categories or parameters. -Native categories apply only to a failed Turn whose outcome has `error_code: engine_failed`. Core accepts only the listed `engine_error_code` values; an unknown, malformed or absent value stays `harness_error`. Only `connection_failed` uses `engine_http_status`. Nested metadata and provider text never classify a failure. Core storage, incomplete-stream and cancellation failures take precedence, and cancelled or completed Turns have no failure. [Native error classification](native-error-classification.md) lists which adapters report each category. +Native categories apply only to a failed Turn whose outcome has `error_code: engine_failed`. Core accepts only the listed `engine_error_code` values; an unknown, malformed or absent value stays `harness_error`. Only `connection_failed` uses `engine_http_status`. Nested metadata and provider text never classify a failure. Core storage, incomplete-stream and cancellation failures take precedence, and cancelled or completed Turns have no failure. [Native error classification](../../docs/runtime-protocol.md#native-failure-classification) lists which adapters report each category. diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md index dafaaee76..a2e88f04e 100644 --- a/contracts/agents-api/runtime-observability-api.md +++ b/contracts/agents-api/runtime-observability-api.md @@ -220,7 +220,7 @@ Disk is not kept in history. ### Token usage -`token_usage` belongs to the Session, not to an allocation. Each point holds the last cumulative measured Session usage sampled in its bucket: `start`, `end`, `sampled_at`, `input_tokens` and `output_tokens`. Measured Session usage is a Core extension that sums every recorded root Turn snapshot, active Turns included. It differs from [public Session usage](history-events-usage.md), which is null while a root Turn runs or after one ends unmeasured. These counters are measured model tokens, not prices or billing records. +`token_usage` belongs to the Session, not to an allocation. Each point holds the last cumulative measured Session usage sampled in its bucket: `start`, `end`, `sampled_at`, `input_tokens` and `output_tokens`. Measured Session usage is a Core extension that sums every recorded root Turn snapshot, active Turns included. It differs from [public Session usage](sessions-events.md#usage), which is null while a root Turn runs or after one ends unmeasured. These counters are measured model tokens, not prices or billing records. ### Errors and bounds diff --git a/contracts/agents-api/sandbox-deployment.md b/contracts/agents-api/sandbox-deployment.md index 9cd204156..360c27fd3 100644 --- a/contracts/agents-api/sandbox-deployment.md +++ b/contracts/agents-api/sandbox-deployment.md @@ -163,7 +163,7 @@ One database snapshot partitions every unreleased allocation and every pending h When both held counts reach zero, the owner drains and atomically clears the provider, mode, specification, provider configuration, credential and metadata and provider policy, retires nodes and unused enrollment tokens, increments the generation and owner epoch, and records `reset_complete`. The installation identity and history remain. Core immediately publishes the unconfigured state and keeps the new generation even with no provider, so a delayed load cannot revive the old one. Configure again with POST and the returned generation; no restart is needed. -During a reset, fresh hosted admission returns 503 `sandbox_reset_in_progress` and leaves no provisional Session rows; live input, receipt retries, restoration and cleanup continue. Management writes and new enrollment return 409 `sandbox_reset_in_progress`, while registered nodes can still read their configuration to recover for cleanup. Per-Session [archive](admin-api.md#administrative-session-archive) needs only the current generation and works with or without a reset. It keeps history and persisted Files and Artifacts, discards the unpersisted workspace and prevents the Session from resuming; poll the archive GET for the actual release. A reset never fabricates a release receipt. +During a reset, fresh hosted admission returns 503 `sandbox_reset_in_progress` and leaves no provisional Session rows; live input, receipt retries, restoration and cleanup continue. Management writes and new enrollment return 409 `sandbox_reset_in_progress`, while registered nodes can still read their configuration to recover for cleanup. Per-Session [archive](admin-api.md#session-archive) needs only the current generation and works with or without a reset. It keeps history and persisted Files and Artifacts, discards the unpersisted workspace and prevents the Session from resuming; poll the archive GET for the actual release. A reset never fabricates a release receipt. A force archive fences credentials and new work at once. When the hosted delivery of the Turn is still connected, Core keeps only that delivery's native cancellation and terminal receipt path open until the terminal commit, for at most 20 seconds from the original cancellation request; `done` does not end the bound while a cancellation acknowledgement or terminal commit is pending. This drain never authorizes reconnection, workspace or MCP access or further execution, and an explicit credential revocation ends it. Missing or failed receipts keep honest failure outcomes, and a disconnected, expired or restarted owner falls back to ordinary provider cleanup. diff --git a/docs/runtime-protocol.md b/docs/runtime-protocol.md index 797e66e98..d6b4e8b56 100644 --- a/docs/runtime-protocol.md +++ b/docs/runtime-protocol.md @@ -101,7 +101,7 @@ The linked source files define the required fields, validators, limits and finit | `workspace_read`, `workspace_write`, `workspace_export` | Matching `*_result` | [Read](../internal/agentdaemon/proto/workspace_read.go), [write](../internal/agentdaemon/proto/workspace_write.go), [export](../internal/agentdaemon/proto/workspace_export.go) | | `environment_quiesce`, `environment_resume` | `environment_quiesced`, `environment_resumed` | [Suspension fencing](../internal/agentdaemon/proto/suspend.go) | -Initial, prepared and active input use the same [ordered MessageInput](../internal/agentdaemon/proto/message_input.go). Adapters keep message and content order and reject unsupported content explicitly; a text-only transport rejects image content rather than dropping it. The [message input contract](../contracts/agents-api/message-input.md) owns the public image profile, whitespace rules and each Harness's native conversion. +Initial, prepared and active input use the same [ordered MessageInput](../internal/agentdaemon/proto/message_input.go). Adapters keep message and content order and reject unsupported content explicitly; a text-only transport rejects image content rather than dropping it. The [message input contract](../contracts/agents-api/message-content.md) owns the public image profile, whitespace rules and each Harness's native conversion. Usage frames and the final usage snapshot each carry the cumulative measurement of the current execution and replace the previous snapshot; never add them. An absent measurement is unknown, not zero. diff --git a/docs/sandbox-provider.md b/docs/sandbox-provider.md index 9b32fe2b5..52facfed7 100644 --- a/docs/sandbox-provider.md +++ b/docs/sandbox-provider.md @@ -186,7 +186,7 @@ A [reset](../contracts/agents-api/sandbox-deployment.md#reset) is durable execut One snapshot and timestamp partition the held resources. Offline ownership comes from the allocation's or active placement's node, with the same 45-second connection and owner-epoch predicate as online presence, independently of provider readiness; no cleanup failure, offline state or empty read authorizes a synthetic release. At zero held resources, Core drains outside database transactions, rechecks under the deployment lock and clears the deployment in one transaction, then publishes a generation-bearing empty provider without fallible work. If the final write or drain fails, it restores the committed provider with a bounded owner context before releasing the mutation gate; if that recovery fails, admission stays fenced and the owner stops. -An administrator [Session archive](../contracts/agents-api/admin-api.md#administrative-session-archive) keeps its Project scope and hosted eligibility checks in a Session-first transaction, together with Environment expiry, cancellation, Runtime authority revocation and audit; a reset's background archive reconstructs the actual Project scope from trusted records and keeps the requester's provenance. The ordinary provider lifecycle releases compute and snapshots. Archive cancellation keeps a healthy receipt path until the terminal commit: only the archive that first revokes a device records its exact `archive_cancel_turn_id` (ordinary revocation clears it, and repeated cleanup keeps it), and the existing authenticated delivery may drain that cancellation for at most 20 seconds from the Turn's original `cancel_requested_at`. Core tracks the delivery through `done`, the cancellation acknowledgement and the terminal commit, independently of subscription removal, and grants no new connection, input, file or MCP authority or lease renewal. No transaction or lifecycle gate waits for the receipt, and a lost peer, expiry or restart falls back to ordinary failure and cleanup, never a fabricated cancelled outcome. +An administrator [Session archive](../contracts/agents-api/admin-api.md#session-archive) keeps its Project scope and hosted eligibility checks in a Session-first transaction, together with Environment expiry, cancellation, Runtime authority revocation and audit; a reset's background archive reconstructs the actual Project scope from trusted records and keeps the requester's provenance. The ordinary provider lifecycle releases compute and snapshots. Archive cancellation keeps a healthy receipt path until the terminal commit: only the archive that first revokes a device records its exact `archive_cancel_turn_id` (ordinary revocation clears it, and repeated cleanup keeps it), and the existing authenticated delivery may drain that cancellation for at most 20 seconds from the Turn's original `cancel_requested_at`. Core tracks the delivery through `done`, the cancellation acknowledgement and the terminal commit, independently of subscription removal, and grants no new connection, input, file or MCP authority or lease renewal. No transaction or lifecycle gate waits for the receipt, and a lost peer, expiry or restart falls back to ordinary failure and cleanup, never a fabricated cancelled outcome. ## Validate the integration diff --git a/packages/claude-sdk-adapter/README.md b/packages/claude-sdk-adapter/README.md index ec054fef7..2a4b826cf 100644 --- a/packages/claude-sdk-adapter/README.md +++ b/packages/claude-sdk-adapter/README.md @@ -52,7 +52,7 @@ With `ObserveToolObservations`, private workspace execution requires the package ### HTTP MCP -The private adapter accepts typed anonymous HTTP and static-bearer HTTPS MCP declarations on the trusted `environment:none` harness host. The packaged readiness report must include `mcp_http_tools`; discovery advertises that feature only when present, and execution rechecks the installed bundle before dispatching an MCP request. An unchanged SDK version alone cannot qualify an older bridge. Authenticated private requests also require the packaged `mcp_http_bearer_auth` feature at discovery and dispatch. The daemon generates a separate environment reference for each server and launch; only those references enter the bridge request and native SDK configuration. The native HTTP client expands them from its owned process environment. Literal bearers must never enter SDK MCP headers because that configuration enters argv. Readiness probes receive no per-request bearer environment. Token validation is shared with the Codex adapter; credential storage remains an opaque-string contract. Public MCP admission, Vault credential selection and their failure rules belong to [public MCP connection origin](../../contracts/agents-api/environments.md#public-mcp-connection-origin) and [credential selection](../../services/core/credentials.md#use-a-credential-in-a-session). +The private adapter accepts typed anonymous HTTP and static-bearer HTTPS MCP declarations on the trusted `environment:none` harness host. The packaged readiness report must include `mcp_http_tools`; discovery advertises that feature only when present, and execution rechecks the installed bundle before dispatching an MCP request. An unchanged SDK version alone cannot qualify an older bridge. Authenticated private requests also require the packaged `mcp_http_bearer_auth` feature at discovery and dispatch. The daemon generates a separate environment reference for each server and launch; only those references enter the bridge request and native SDK configuration. The native HTTP client expands them from its owned process environment. Literal bearers must never enter SDK MCP headers because that configuration enters argv. Readiness probes receive no per-request bearer environment. Token validation is shared with the Codex adapter; credential storage remains an opaque-string contract. Public MCP admission, Vault credential selection and their failure rules belong to [public MCP connection origin](../../contracts/agents-api/environments.md#public-mcp-connection-origin) and [credential selection](../../contracts/agents-api/vaults.md#credential-selection-in-a-session). MCP queries use the SDK's main-thread Agent definition to restrict model-visible tools, in addition to empty built-ins, strict MCP configuration, empty setting sources and default-deny permissions. Permission allowlists alone do not restrict the native model inventory. Null selects all tools from a declared server; an empty list selects none. Host functions compose with those selections. Native server status supplies original tool identities; map their normalized native aliases while preserving the original names in observations. Native status deduplicates aliases, so it does not prove a complete original server inventory. A native PreToolUse hook waits for inventory verification before admitting root calls and denies unverified, mismatched or cancelled calls. The native Agent restriction controls model-visible tools; inventory verification is not a barrier before the model request. Anonymous HTTP declarations explicitly set an empty Authorization header to disable native OAuth and automatic credential injection. Preserve that header; do not erase native history or credentials to enforce this boundary. Servers that reject a blank Authorization header, normalized name collisions and inventory changes during a query require separate validation; this profile covers static inventories. Private SDK status/control objects can contain expanded authentication headers. Read only connection and tool identity fields; never retain, log or publish raw status/configuration or control responses. Diagnostic projections must whitelist safe fields; filtering actual model or tool output does not fix a leak. The adapter profile requires connected servers, reserves the `functions` label, accepts alphanumeric/underscore/hyphen server labels and alphanumeric/underscore/hyphen/dot selected tool names, and excludes remote environments. Required startup is separately qualified by `mcp_http_required`. All HTTP MCP queries use native SDK startup and an empty input iterator to confirm initialization hooks. Required declarations additionally check connected server status before the initial prompt is released exactly once. Pending, failed, missing or ambiguous required status rejects before input; native startup timeouts are retained without an adapter retry loop. Normal system/init still verifies Session identity and the complete inventory before input readiness/tool authority. Optional servers keep their inventory checks without a pre-input connection requirement. A Runtime must advertise the concrete required-initialization capability; there is no fallback to another execution path. diff --git a/scripts/name-allowlist.json b/scripts/name-allowlist.json index 9fa95fb93..2810a1f89 100644 --- a/scripts/name-allowlist.json +++ b/scripts/name-allowlist.json @@ -59,11 +59,6 @@ "regex": "https://developers\\.openai\\.com/api/docs/guides/agents-api(?:/[A-Za-z0-9_./-]*)?", "reason": "Official upstream documentation URLs are external protocol references." }, - { - "path": "*", - "regex": "~?/\\.parsar/remediation/", - "reason": "Historical remediation evidence paths are retained." - }, { "path": ".gitignore", "regex": "(?m)^/\\.parsar/$", @@ -244,16 +239,6 @@ "regex": "and Parsar itself", "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." }, - { - "path": "services/core/README.md", - "regex": "Parsar's product|Parsar\\sworkspace|Parsar Skill/SP", - "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." - }, - { - "path": "services/core/deploy/claude/README.md", - "regex": "Parsar is not a dependency", - "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." - }, { "path": "docs/getting-started/operations.md", "regex": "~/\\.parsar/core", diff --git a/services/core/IMPLEMENTATION.md b/services/core/IMPLEMENTATION.md index 8daf9ab3a..b14d66664 100644 --- a/services/core/IMPLEMENTATION.md +++ b/services/core/IMPLEMENTATION.md @@ -14,13 +14,27 @@ Requests are served on their canonical path and never redirected. `api.Canonical On the Beta group, the OpenAI-Beta check (exactly one `agents=v1` value) runs before authentication, and authentication precedes every Beta handler, 404 and 405. The API router's own middleware, not the shared log middleware, sets a fresh `X-Request-Id` (also in the log context), `OpenAI-Version`, `OpenAI-Processing-Ms` and `nosniff`. HEAD runs GET routes, except that streaming, content-download, live directory, Runtime observation and Runtime history routes register an explicit HEAD 405. Every 405 of the API router, unknown methods included, has the JSON body and lists the route's methods in `Allow`. +- Stored strings other than metadata rely on PostgreSQL rejecting U+0000 and invalid UTF-8: map SQLSTATE `22021` (text parameter) and `22P05` (`\u0000` in jsonb) to the 400 unstorable-text error, including query filters such as `agent_id`. The failing statement aborts its transaction, so keep each request's writes in one transaction. +- An `after` cursor that cannot name a resource on a lookup list (Agents, Sessions, Turns, Templates, Vaults, Credentials) resolves to the never-assigned maximum UUID and runs the normal lookup, so storage failures and missing rows behave as for a well-formed cursor. Resolve every cursor only inside its already resolved parent and tenant. +- Lists whose parent and cursor lookups are separate statements (Artifacts, Skill versions) re-check the parent before reporting a cursor 400, so a parent deleted in between still returns its 404. Item and Subagent lists read both inside one locked Session transaction. The Skill version cursor lookup is tenant-wide so another Skill's version can be told apart from a missing one; another tenant's version stays missing. + ## Source Files and Artifacts Source Files are Project resources with a lifecycle independent of copied workspace files. Store immutable source metadata and PostgreSQL large objects in Core's database with the pinned pgx driver. Upload validation, metadata insertion and the bytes commit atomically; deletion removes the metadata and unlinks the object in one transaction. Keep OIDs private and authorize every metadata, content and delete lookup by tenant before opening a body. Stream bounded chunks; never hold an entire upload in memory or use a filename as a filesystem path. A direct download of the `user_data` purpose is rejected after the tenant-scoped metadata lookup, while initialization and workspace copies keep their authorized store read. A read-only repeatable-read transaction preserves an admitted source across concurrent deletion; resolve that snapshot before entering the Environment write path, and a later deletion never undoes a completed workspace copy. Bound request and transaction lifetimes, roll back incomplete bodies and never retry an ambiguous commit automatically. Backups must include PostgreSQL large objects, and a schema rollback must not orphan them. Session Artifacts are immutable published copies, separate from live workspace files and source Files. The private output exporter reuses the authorized workspace path boundary and streams bounded bytes; publication requires complete capture and confirmed helper and transport success, not merely valid archive syntax. The daemon owns and drains the exporter's stdout pipe separately from child reaping, so pull-transport backpressure cannot consume the process-exit I/O deadline; after helper exit, each pipe read has one second, reset after consumer delays, which rejects inherited pipes that never close. Cancellation closes the owned reader and the dispatch consumer, then waits for the child. Never extract an output archive into Core's filesystem or hold the execution lease through a large transfer. -Capture bytes into private large objects without a Session admission lock. Before capture, seal native input under that lock with the private capture marker; later messages reuse the Environment input reservation and wait for the next Turn, and the public Turn stays in progress until publication settles. Directory reads during capture use an independent read-only preparation, not the released native Run. After a confirmed export, lock and recheck the live Turn, then publish the metadata in the same transaction as the Turn's completion. In that transaction, drop staged paths whose sha256 equals the newest published Artifact for the path in the Session, so later Turns publish only new, changed or no-longer-published paths and never modify existing Artifacts. Failed and cancelled Turns discard private objects, and Session deletion removes private and published copies. The exporter skips output symlinks by their `lstat` type without following them; hard links, other special files, device crossings and concurrent changes reject the capture. Stored reads are authorized independently of Environment availability, so published Artifacts outlive the Environment. +Capture bytes into private large objects without a Session admission lock. Before capture, seal native input under that lock with the private capture marker; later messages reuse the Environment input reservation and wait for the next Turn, and the public Turn stays in progress until publication settles. Directory reads during capture use an independent read-only preparation, not the released native Run. After a confirmed export, lock and recheck the live Turn, then publish the metadata in the same transaction as the Turn's completion. Decide Artifact republication in the Turn's terminal transaction, not the capture transaction: capture commits and releases the Session lock before the Turn completes, and the terminal transaction holds the Session lock that also orders Artifact deletion. Drop staged rows whose SHA-256 equals the newest remaining published Artifact for the path, ordered by the producing Turn's creation time and then ID (publication time can come from the Runtime and does not order Turns), and unlink their large objects in the same transaction. Published rows are never modified. Failed and cancelled Turns discard private objects, and Session deletion removes private and published copies. The exporter skips output symlinks by their `lstat` type without following them; hard links, other special files, device crossings and concurrent changes reject the capture. Stored reads are authorized independently of Environment availability, so published Artifacts outlive the Environment. + +- A malformed `environment_id` Artifact filter resolves to the never-assigned maximum UUID, so it matches nothing without a text comparison. + +## Environment files and Skills + +- Environment file list tokens bind a digest of the tenant, Environment, requested directory, effective order and limit, plus a fingerprint of the full sorted regular-file path and size list and an offset that is a multiple of the limit. Every page rereads the directory; there is no cursor registry, cache or snapshot. Reject every token mismatch with the single official token message. +- The directory helper checks each requested path component with `Root.Lstat` below `os.OpenRoot(workspace)` and opens the final directory with `O_NOFOLLOW`. A missing component, a regular file or a symbolic link maps to the distinct `not_directory` result, which the daemon and gateway carry only for directory reads; Core turns it into an empty page. `not_found` (Claude SDK adapter reader), permission, transport and uncertain results keep their errors. +- A Files.create write intent stores a digest of the path, size and content, not a path ledger, so Core cannot tell a file an earlier Files.create wrote from any other file; an existing regular file therefore gets the untracked-file message. Reserve the intent under the Session lock before dispatch. The daemon verifies the complete body's SHA-256 before calling the writer, and the writer creates parents with `Root.MkdirAll(0700)`, writes `.oac-write-` in the workspace root and publishes it with `Root.Link`, which never replaces an existing entry. Known refusals return `write_rejected` with `reason` `destination_directory` or `unsafe_destination`; Core settles the intent as `rejected`, which leaves no committed receipt and releases the mutation owner. Only an exact committed or rejected receipt settles an intent; nothing settles an unknown one automatically. +- Check the 5 MiB inline bound after path validation and Base64 decoding, and before the pending-hosted check, the source File lookup and execution. The JSON body limit still admits the Base64 form of 50 MiB so that oversized inline bodies up to that size get the official message. +- Serialize Skill version uploads, default changes and version deletion on the owning Skill row lock. Deleting the default version deletes the Skill only when no other version row exists, through the same cascade as Skill deletion, so every encrypted version row goes in the same commit. `next_version` only increases. ## Scheduling, preparation and pending input @@ -72,9 +86,11 @@ Session execution-configuration reads use a separate immutable safe projection w Provider input validation uses the adapter rules in `internal/harnessconfig`: one internal registry for protocol and token-limit validation, while Core owns credential environment and endpoint admission policy. +- Saved Agents parse tools with the saved-form parser, which keeps every pinned `web_search` mode. Session admission re-resolves the effective tools with the execution parser, which admits only disabled search; Worker device selection and the final preclaim also refuse a non-disabled search control. + ## Vaults and credentials -[Vault credentials](credentials.md) and [OAuth credentials](oauth-credentials.md) describe the resources, selection rules, refresh and deletion. The store implements them under these rules: +[Vaults and credentials](../../contracts/agents-api/vaults.md) describes the resources, selection rules, refresh and deletion. The store implements them under these rules: - Credentials are children of tenant-owned Vaults. Creation admits the owner in the same SQL statement as the insert; retrieval joins the owning Vault; listing enforces Project and Vault ownership on the parent, cursor and row query. Metadata queries never select ciphertext and need no encryption key. - Secret values are encrypted before they reach SQL, with Core's separately configured random 32-byte key and the standard library's random-nonce AES-GCM. The versioned authenticated binding covers tenant, Vault, Credential, authentication purpose and exact destination. Never reuse daemon transport encryption for this storage. A missing key disables credential writes; a malformed configured key fails startup. @@ -83,10 +99,22 @@ Provider input validation uses the adapter rules in `internal/harnessconfig`: on - Credential deletion is one mutation checked by tenant, Vault and ID; Vault deletion removes the parent and its Credentials through the foreign-key cascade in one SQL statement, without decrypting, needing the key or calling providers. - At dispatch, Core rechecks the tenant, attached Vault, selected ID, frozen auth type and exact URL before scoped decryption, and the token enters only the transient daemon request. A missing key or binding failure never falls back to anonymous execution. +- `credentialcrypto` ciphertext is a format version byte followed by the standard AEAD nonce, ciphertext and tag. The authenticated data holds a fixed domain and version plus the binding (tenant, Vault, Credential, auth type, exact destination). Keep the domain string unchanged: existing rows must still decrypt. +- Random-nonce GCM allows at most 2^32 encryptions per key. `secrets/credential.key` also seals model providers, the E2B key, Skills, initial files and environment setup, so every sealed write counts toward that bound; there is no rotation or re-encryption path. +- OAuth dispatch refresh holds the Credential row lock and the external exchange under one 20-second context (`store.oauthRefreshTimeout`). The refresh HTTP client has a 10-second overall timeout and 5-second TLS handshake and response-header timeouts, uses no proxy and treats any redirect as failure. + ## MCP Accepted public MCP credential profiles are declared centrally by the service, separately from private adapter capabilities; a daemon capability alone never opens a profile. Frozen binding validation runs at admission and later input, and the same capability and placement checks run at device selection, the final preclaim check and request construction, before scoped decryption. Native configuration and token injection stay in the adapters; the [Codex adapter](deploy/codex/README.md#mcp-servers) and the [Claude SDK adapter](../../packages/claude-sdk-adapter/README.md#http-mcp) describe theirs. The shared resolver keeps omitted or null `allowed_tools` as unrestricted and an explicit empty list as deny-all. Saved HTTP transport output includes `headers: {}` while the effective Session transport omits headers, matching the two pinned resource types. An omitted or null `connection_origin` on HTTP transport is stored as `service` before the other checks. +## Message text and result targets + +- Admission and the Claude bridge share one whitespace set, the union of Go `unicode.IsSpace` and ECMAScript `String.prototype.trim` (`blankTextRune` in `services/core/internal/execution/message_support.go`). A shared table test keeps both sides equal; change them together. +- Whitespace-only admission is a field of the engine profile (`WhitespaceOnlyText`) checked with the image profile during Worker admission. Never branch on the harness name in handlers. +- `MessageInput.Validate` treats any non-empty text part as content and never trims. It runs in Core admission, Worker delivery, daemon steering and prepared start, and in the Codex and MiniMax adapters. +- Resolve a tool result's target under the tenant Session lock, after the Session lookup: a well-formed `turn_id` is looked up in that Session and must own the call; otherwise the Session's own calls decide between "unknown call" and "different Turn". Never reject a malformed `turn_id` before the Session lookup, so missing, malformed and foreign Sessions keep one 404, and read only the caller's Session. +- The recorded Environment failure reason is composed only from a fixed step label and integers (setup command index, exit status 1-255), so commands, environment values, package names, paths and process output cannot reach the reason, events, logs or responses. + ## Dispatch and execution `services/core/internal/execution` claims a Turn from `queued` to `in_progress` before subscribing or sending, and never replays a claimed or interrupted Turn. Extra inputs require native receipts. The terminal outcome and the native Session ID commit together under the admission lock, and unapplied messages prevent a successful completion. Credentials are resolved separately from the immutable non-secret snapshot. A peer that lacks a required capability is rejected before the claim, and a failed strict resume never starts unrelated history. Native continuation needs the device's persisted engine files; a native Session ID alone cannot restore deleted history. @@ -130,42 +158,7 @@ The [managed lifecycle](../../docs/sandbox-provider.md#managed-lifecycle) descri - Observations of offline, missing or unconfirmed resources persist only bounded, sanitized codes, separately from the lifecycle, and a host restart never fabricates a running observation. - Runtime observation of a node allocation goes through its immutable placement as one bounded read-only Provider operation; an absent capability or transport returns unavailable, never a Core-local fallback. -## Public and administration API constraints - -These constraints implement the rules in the [Agents API contracts](../../contracts/agents-api/README.md) and the [Core administration API](../../contracts/agents-api/admin-api.md). - -### Wire validation mechanics - -- Stored strings other than metadata rely on PostgreSQL rejecting U+0000 and invalid UTF-8: map SQLSTATE `22021` (text parameter) and `22P05` (`\u0000` in jsonb) to the 400 unstorable-text error, including query filters such as `agent_id`. The failing statement aborts its transaction, so keep each request's writes in one transaction. -- An `after` cursor that cannot name a resource on a lookup list (Agents, Sessions, Turns, Templates, Vaults, Credentials) resolves to the never-assigned maximum UUID and runs the normal lookup, so storage failures and missing rows behave as for a well-formed cursor. Resolve every cursor only inside its already resolved parent and tenant. -- Lists whose parent and cursor lookups are separate statements (Artifacts, Skill versions) re-check the parent before reporting a cursor 400, so a parent deleted in between still returns its 404. Item and Subagent lists read both inside one locked Session transaction. The Skill version cursor lookup is tenant-wide so another Skill's version can be told apart from a missing one; another tenant's version stays missing. -- Saved Agents parse tools with the saved-form parser, which keeps every pinned `web_search` mode. Session admission re-resolves the effective tools with the execution parser, which admits only disabled search; Worker device selection and the final preclaim also refuse a non-disabled search control. - -### Message text and result targets - -- Admission and the Claude bridge share one whitespace set, the union of Go `unicode.IsSpace` and ECMAScript `String.prototype.trim` (`blankTextRune` in `services/core/internal/execution/message_support.go`). A shared table test keeps both sides equal; change them together. -- Whitespace-only admission is a field of the engine profile (`WhitespaceOnlyText`) checked with the image profile during Worker admission. Never branch on the harness name in handlers. -- `MessageInput.Validate` treats any non-empty text part as content and never trims. It runs in Core admission, Worker delivery, daemon steering and prepared start, and in the Codex and MiniMax adapters. -- Resolve a tool result's target under the tenant Session lock, after the Session lookup: a well-formed `turn_id` is looked up in that Session and must own the call; otherwise the Session's own calls decide between "unknown call" and "different Turn". Never reject a malformed `turn_id` before the Session lookup, so missing, malformed and foreign Sessions keep one 404, and read only the caller's Session. -- The recorded Environment failure reason is composed only from a fixed step label and integers (setup command index, exit status 1-255), so commands, environment values, package names, paths and process output cannot reach the reason, events, logs or responses. - -### Environment files, Skills and Artifacts - -- Environment file list tokens bind a digest of the tenant, Environment, requested directory, effective order and limit, plus a fingerprint of the full sorted regular-file path and size list and an offset that is a multiple of the limit. Every page rereads the directory; there is no cursor registry, cache or snapshot. Reject every token mismatch with the single official token message. -- The directory helper checks each requested path component with `Root.Lstat` below `os.OpenRoot(workspace)` and opens the final directory with `O_NOFOLLOW`. A missing component, a regular file or a symbolic link maps to the distinct `not_directory` result, which the daemon and gateway carry only for directory reads; Core turns it into an empty page. `not_found` (Claude SDK adapter reader), permission, transport and uncertain results keep their errors. -- A Files.create write intent stores a digest of the path, size and content, not a path ledger, so Core cannot tell a file an earlier Files.create wrote from any other file; an existing regular file therefore gets the untracked-file message. Reserve the intent under the Session lock before dispatch. The daemon verifies the complete body's SHA-256 before calling the writer, and the writer creates parents with `Root.MkdirAll(0700)`, writes `.oac-write-` in the workspace root and publishes it with `Root.Link`, which never replaces an existing entry. Known refusals return `write_rejected` with `reason` `destination_directory` or `unsafe_destination`; Core settles the intent as `rejected`, which leaves no committed receipt and releases the mutation owner. Only an exact committed or rejected receipt settles an intent; nothing settles an unknown one automatically. -- Check the 5 MiB inline bound after path validation and Base64 decoding, and before the pending-hosted check, the source File lookup and execution. The JSON body limit still admits the Base64 form of 50 MiB so that oversized inline bodies up to that size get the official message. -- Serialize Skill version uploads, default changes and version deletion on the owning Skill row lock. Deleting the default version deletes the Skill only when no other version row exists, through the same cascade as Skill deletion, so every encrypted version row goes in the same commit. `next_version` only increases. -- Decide Artifact republication in the Turn's terminal transaction, not the capture transaction: capture commits and releases the Session lock before the Turn completes, and the terminal transaction holds the Session lock that also orders Artifact deletion. Drop staged rows whose SHA-256 equals the newest remaining published Artifact for the path, ordered by the producing Turn's creation time and then ID (publication time can come from the Runtime and does not order Turns), and unlink their large objects in the same transaction. Published rows are never modified. -- A malformed `environment_id` Artifact filter resolves to the never-assigned maximum UUID, so it matches nothing without a text comparison. - -### Credential encryption and OAuth refresh bounds - -- `credentialcrypto` ciphertext is a format version byte followed by the standard AEAD nonce, ciphertext and tag. The authenticated data holds a fixed domain and version plus the binding (tenant, Vault, Credential, auth type, exact destination). Keep the domain string unchanged: existing rows must still decrypt. -- Random-nonce GCM allows at most 2^32 encryptions per key. `secrets/credential.key` also seals model providers, the E2B key, Skills, initial files and environment setup, so every sealed write counts toward that bound; there is no rotation or re-encryption path. -- OAuth dispatch refresh holds the Credential row lock and the external exchange under one 20-second context (`store.oauthRefreshTimeout`). The refresh HTTP client has a 10-second overall timeout and 5-second TLS handshake and response-header timeouts, uses no proxy and treats any redirect as failure. - -### Core administration errors, metrics and write provenance +## Core administration errors, metrics and write provenance - Core error details are scoped by the `/core/v1` router's writer mark, never by a request path test. Write Core errors with `writeCoreError` and typed `CoreErrorDetails` values (string, number, boolean, null and string-array constructors); invalid or empty details are omitted as a whole. The mark preserves error observation, flushing and `http.ResponseController` access. A shared handler or a Core-looking path alone never changes a public or machine error envelope. Core authentication runs before operation configuration checks, and unknown paths keep their status and admission rules. When adding a code with details, document its fixed keys in `contracts/agents-api/core-errors.md`, and pass only safe Core-owned facts: never submitted values, secrets, native text or provider bodies. - Operation validators keep their original error text, sentinel identity and validation precedence. Package-owned typed errors carry fixed field metadata; only the marked Core error mapper translates it into operation codes and safe bound or catalog details. Keep Project and key rune limits separate from node byte limits. Sandbox validation metadata travels through its store wrapper without changing transaction or provider authority. Public Session provider validation stays byte-for-byte unchanged; cover it with handler-level golden responses. The Core clients ignore malformed optional details and never retry a write.