Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions contracts/agents-api/sandbox-deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,12 +84,13 @@ For E2B, post `{"configuration": {"api_url": "…", "domain": "…"}, "credentia
GET and successful writes return `installation_id`, `provider`, `core_url` (read-only: the installation public URL, present even before configuration), `mode`, `generation`, `owner_epoch`, `reset`, `rollout`, `suspension`, `resources` and `credential_configured`. A configured deployment also returns `specification`, `specification_digest`, `configuration` and `metadata`: the adapter's public projection of its selectors and of the observations it recorded, never raw stored values or secrets. E2B returns `configuration.template`, `configuration.api_url`, `configuration.domain` and, once recorded, `metadata.template_build`. Docker and microsandbox return empty `configuration` and `metadata` objects and `credential_configured: false`; an unconfigured deployment has neither object.

- `metadata.template_build` is `{status, resources: {cpus, memory_mib, root_disk_mib}}`: the build as Core read it through the pinned SDK when the selection was saved. GET never calls E2B, so it stays cheap during an E2B outage. Validation admits only a `ready` build whose CPU count and memory equal the selected values; `root_disk_mib` is the build's native disk size, which Core does not enforce. Unknown values are null, `metadata: {}` means no observation was recorded, and an identical PUT without a credential does not refresh it.
- `suspension` is `{idle_seconds, retention_seconds}` for microsandbox, the only provider Core suspends (currently 300 and 86400); Docker, E2B and unconfigured deployments return null.
- Request `resources` and response `specification.resources` are per-sandbox limits. Response `resources.allocations` and `resources.pending` count unreleased allocations and pending hosted Environments without an allocation.
- An unconfigured deployment has an empty provider and no specification. Docker and microsandbox use `mode: nodes`; E2B uses `mode: direct`, without a synthetic node.
- `generation` identifies the saved selection. `owner_epoch` fences the execution owner and node connections; it does not replace `expected_generation`.
- `specification_digest` is the server's identity of the provider, limits and Runtime release; enrollment echoes it unchanged.

`suspension` is `{idle_seconds, retention_seconds}` for microsandbox and new E2B selections (currently 300 and 86400). Docker and unconfigured deployments return null. E2B deployments saved before this policy retain `null` until an operator submits the same E2B selection with an updated `expected_generation` through `PUT /core/v1/sandbox/deployment`; that explicit update enables the policy and increments the deployment generation. After initialization, Core pauses an idle E2B sandbox with its memory retained, including a Session that has not yet received a Turn. The next input resumes the same sandbox ID and waits for the Runtime daemon to reconnect before admission. The 86400-second retention bounds the paused allocation; cleanup after expiry removes the sandbox, so a later request cannot resume that original compute. Paused provider storage may still incur charges. An uncertain pause result is observed under the allocation's private receipt, never replayed against a potentially late native request.

The typed `SandboxAdminClient` in `packages/agents-client` checks the deployment, node list, node detail and allocation responses against exactly these shapes. An unknown or missing member, or a wrong type, rejects the whole response with a 502 `invalid_admin_response` error. A node's `diagnostic` is absent or a code, never empty, and the client reads an unknown code as `provider_unavailable`.

## Initial setup and same-provider changes
Expand Down Expand Up @@ -190,7 +191,7 @@ Some fields keep one name across providers but differ in meaning, or do not appl
| Deployment `specification.resources` | `cpus` and `memory_mib`, equal to the ready template build's and taken from it when omitted; no disk fields | `cpus` and `memory_mib`; no disk quota | `cpus`, `memory_mib`, `root_disk_mib` and `environment_disk_mib` |
| Deployment `specification.runtime` | Absent; `configuration.template` selects the build | The full [release](#runtime-release); nodes match `image_id` or `image_manifest_digest` | The full [release](#runtime-release); nodes match `microsandbox_ref`, `runtime_sha256` and `firmware_sha256` |
| Deployment `metadata.template_build` | The build as Core read it when the selection was saved | Absent: `metadata` is empty | Absent: `metadata` is empty |
| Deployment `suspension` | `null`; Core does not suspend E2B sandboxes | `null` | `{idle_seconds, retention_seconds}` |
| Deployment `suspension` | `{idle_seconds, retention_seconds}` | `null` | `{idle_seconds, retention_seconds}` |
| Deployment `resources.allocations`, `resources.pending` | Core's unreleased E2B sandboxes, and hosted Environments waiting for one | Totals across all nodes | Totals across all nodes |
| Enrollment-token `max_active`, `max_retained` | 409 `sandbox_deployment_conflict`, after the 400 capacity checks; E2B has no nodes | `max_retained` always equals `max_active` | Both limits apply |
| Node list and detail | Empty list; detail returns 404 | Enrolled nodes | Enrolled nodes |
Expand All @@ -200,7 +201,7 @@ Some fields keep one name across providers but differ in meaning, or do not appl
| Allocation `compute_phase`, `compute_phase_changed_at` | Not applicable: no node allocations | Always `disabled`, counted as running until release; the time is the allocation's creation | Includes `suspended`; its time plus `suspension.retention_seconds` tells roughly when Core reclaims the snapshot |
| Runtime observation `cpu`, `memory` | From E2B metrics: `cpu.utilization_ratio` and `capacity_cores`, memory usage and limit; no cumulative CPU time | From Docker stats: `cpu.usage_seconds_total`, CPU and memory limits, memory usage | From the VM: `cpu.usage_seconds_total`, CPU and memory limits, memory usage |
| Runtime observation `disk` | E2B `diskUsed` and `diskTotal`; `null` when the template does not report them | `null`: no disk quota | `null` |
| Runtime observation `lifecycle_state: sleeping` | Never | Never | While suspended |
| Runtime observation `lifecycle_state: sleeping` | No running sample while paused | Never | While suspended |
| Runtime history CPU | Mean of the utilization ratios E2B reported in each bucket | Derived from cumulative CPU time | Derived from cumulative CPU time |

## Errors
Expand Down
13 changes: 13 additions & 0 deletions docs/runtime-bootstrap.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,19 @@ The Runtime validates the input and owns authentication and connection. A succes

Self-hosted executors and operator-provisioned devices get their daemon identity in other ways; the [machine connection API](../contracts/agents-api/machine-api.md#credentials) lists every credential source. All of them enter the same Runtime execution loop.

## Hosted suspension control

The private control-file path is defined once in
[`suspension.json`](../internal/runtimebootstrap/suspension.json). Core recovery
commands and Go providers read it through `runtimebootstrap.SuspendControlFile()`.
The E2B template packager writes the same value into its protected Runtime
environment configuration; managed startup reads that value and prepares its
parent directory with mode 0700 and Runtime ownership. The daemon enables the
existing suspension protocol through `OAC_RUNTIME_DAEMON_SUSPEND_PID_FILE`.
This is a packaged internal protocol setting, not an operator-editable idle policy
or a Provider lease timeout. The control file is not part of the bootstrap
credential document above.

## Verification

`go test ./internal/runtimebootstrap ./apps/daemon/internal/cli` covers the input contract, the exclusivity of credential sources and restart behavior. Provider tests verify delivery and file permissions without relying on the Runtime's private storage.
17 changes: 14 additions & 3 deletions docs/sandbox-provider.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@ A **Sandbox Provider** supplies the outer compute that a Runtime daemon runs in

Core owns durable Environment, allocation, placement and cleanup state; the Provider owns compute and bootstrap only. The Runtime prepares capabilities and runs Turns over the [Core–Runtime protocol](runtime-protocol.md), and the provider hands it its identity through the [Runtime bootstrap](runtime-bootstrap.md) file. A provider never runs Environment initialization, Skills, Plugins, MCP setup, initial files, execution or Files; those use the Runtime. Isolation belongs to the provider's infrastructure, not the daemon; see [Runtime and outer isolation](concepts.md#runtime-and-outer-isolation). Use the vendor's maintained SDK behind a thin adapter.

A thin adapter lets a Provider be replaced without changing the common execution flow. Core owns shared scheduling, persistence and recovery through declared capability contracts; provider SDK calls and native behavior stay inside the adapter. Shared policy decisions use registered policy rather than vendor names.

## Steps

1. **Read the contract.** Implement the five required operations and give an explicit decision for every extension interface in [Implement the interface](#implement-the-interface).
Expand Down Expand Up @@ -47,9 +49,12 @@ Every provider implements the methods of each interface below and returns a comp
| `runtimeobs.BatchSource` | Explicit decision; requires `Source` | Bounded observations in input order, with per-target errors |
| `sandbox.SelectionDiscoverer` | Explicit decision | Read-only native configuration discovery before commit |
| `sandbox.CredentialVerifier` | Explicit decision | Verify access to owned resources without mutation |
| `sandbox.ResidentPauseProvider` | Explicit decision for both operations | Memory-preserving pause and resume of the same owned compute ID |

`CheckpointProvider` adds `Initial` and `NewCompute` (construct compute references without allocating), `GetCompute`, `Suspend`, `Resume`, `ResumeCompute` (thaw only the same resident instance after an aborted pause), `KillCompute`, `DeleteSnapshot` and `RunCommandCompute`, which runs a bounded command in one exact compute incarnation. Core uses `RunCommandCompute` to wake a parked daemon after a restore ([`runtime_compute_wake.go`](../services/core/internal/execution/runtime_compute_wake.go)).

The checkpoint lifecycle and resident pause/resume pair must each agree on support.

Each declaration entry is `state: supported` with no reason, or `state: unsupported` with an authored reason code. Missing, zero, unknown or unsafe entries and missing methods fail validation. Adding a method to an interface requires an explicit decision and implementation in every adapter; never supply a base type or generate blanket unsupported implementations.

An unsupported method returns `providercontract.UnsupportedError` before any native I/O. The error names the exact operation and a safe code, never a native message, resource identity, endpoint or credential. An empty result, a nil error, `Unavailable` or an unknown mutation outcome never stands in for unsupported, and the five required methods can never return it.
Expand Down Expand Up @@ -90,6 +95,12 @@ Every call receives a bounded context. Expiry or cancellation ends the caller's

Checkpoint support adds `Compute` generation, name and ID and `SnapshotIdentity`; persist operation IDs and the provider's snapshot provenance unchanged. `ObserveOnly` on suspend or resume observes the previous attempt and never starts another capture or restore. `ResumeCompute` only thaws the retained source and never cold-starts a stopped one. Cleanup targets the exact compute incarnation and snapshot, not whatever instance now has the same name. Read [`runtime_compute.go`](../services/core/internal/execution/runtime_compute.go) and its failure tests before declaring checkpoint support.

Resident pause retains the original provider ID and has no separate snapshot
identity. Core quiesces the Runtime before calling `PauseResident`, persists the phase
before the native call, and admits work only after `ResumeResident` and Runtime
reconnection. An unknown pause result must be observed without replaying a
late mutation. See [the resident lifecycle](../services/core/internal/execution/runtime_compute_resident.go).

### Four distinct readiness facts

| Fact | Evidence | Does not establish |
Expand Down Expand Up @@ -124,7 +135,7 @@ A new provider takes these steps:
- A `nodes` registration has only `BuildLocal`, and a `direct` registration only `BuildDirect`; missing, mixed or unknown modes are rejected.
- The specification and resource validators, the configuration adapter and a complete operation declaration are mandatory, so an incomplete registration cannot publish a partial installer projection.
- The Runtime input policy either accepts the pinned Runtime or gives the adapter's fixed reason for rejecting it, never both.
- Checkpoint support requires node mode and positive idle and retention defaults that fit Runtime durations; a provider without checkpoint support configures no suspension defaults.
- Checkpoint suspension requires node mode; resident pause requires direct mode. Either lifecycle requires positive idle/retention defaults that fit Runtime durations. The two durations are independent. Providers supporting neither lifecycle must not configure suspension defaults.

The configuration adapter must be non-nil, including its concrete value. Every `ConfigurationRequirements` field needs an explicit valid decision: `Credential` and `PublicOrigin` are `Required` or `NotRequired`, and `Discovery` uses the shared supported or unsupported declaration with a safe reason. A new requirement field or discovery method needs an explicit validation update and never inherits an existing decision. Configuration discovery is distinct from resource selection discovery, and requiring a credential does not promise the `VerifyCredential` operation. These checks establish complete registration, not correct native SDK behavior; constructor and adapter contract tests still apply.

Expand All @@ -134,7 +145,7 @@ Preview and persistence use `providers.Normalize` and `providers.Describe`. `Sel

A direct adapter with a credential verifies all retained generations and allocation references before a key is replaced. The common `sandbox.CallFence` excludes native calls and waits for helper completion, including calls whose callers timed out; execution invokes the prepared verification and fencing callbacks without branching on a vendor.

Vendor deployment validation and SDK setup stay at the construction boundary, and construction never creates an Environment. For node-local adapters `providers.Built` returns the provider, probe, installation identity, backend fingerprint and specification digest, and the factory also returns its close function. `execution.RuntimeProvider` binds the adapter to its kind, installation ID, backend fingerprint, generation, mode and node ownership; the database owns the selection, and the in-memory copy is never another authority. Docker and microsandbox run on nodes, and E2B is constructed directly. The node proxy exposes checkpoint operations only for a backend whose registered declaration supports them, and common lifecycle code admits suspension through `CheckpointProvider`, never through a provider name.
Vendor deployment validation and SDK setup stay at the construction boundary, and construction never creates an Environment. For node-local adapters `providers.Built` returns the provider, probe, installation identity, backend fingerprint and specification digest, and the factory also returns its close function. `execution.RuntimeProvider` binds the adapter to its kind, installation ID, backend fingerprint, generation, mode and node ownership; the database owns the selection, and the in-memory copy is never another authority. Docker and microsandbox run on nodes, and E2B is constructed directly. The node proxy exposes checkpoint operations only for a backend whose registered declaration supports them, and common lifecycle code admits suspension through the declared `CheckpointProvider` or `ResidentPauseProvider` capability, never through a provider name.

The backend fingerprint identifies a native resource namespace, not capacity. Core keeps deployment generations so that owned allocations keep resolving to their original backend; never repoint retained allocations at a replacement backend.

Expand Down Expand Up @@ -182,7 +193,7 @@ The deployment's CPU, memory and disk settings, `max_active`, `max_retained` and

### Suspension

A provider with checkpoint support can suspend idle work; the deployment's [`suspension`](../contracts/agents-api/sandbox-deployment.md#safe-response) policy sets the idle time and snapshot retention. Core suspends only after at least one Turn is terminal, when no root or Subagent Turn is queued, in progress or waiting, no input, file operation or initialization is pending, and real activity has been idle for the configured interval. For node allocations Core records the first root or child terminal transition with the database clock in the same transaction. Candidate filtering and the Session-locked recheck compare elapsed database time with the idle duration, and the initial snapshot retention deadline is anchored to the same database observation, so Core and database host clocks need not agree. Native completion timestamps stay unchanged in public history but never drive idle admission, and heartbeats never reset activity. Before acknowledging a planned suspension, the daemon closes admission and drains native cleanup, output receipts and file work.
A provider with checkpoint or resident pause support can suspend idle work; the deployment's [`suspension`](../contracts/agents-api/sandbox-deployment.md#safe-response) policy sets the idle time and snapshot retention. Checkpoint suspension requires at least one terminal Turn; resident pause also admits initialized Sessions that have never received a Turn. Core suspends when no root or Subagent Turn is queued, in progress or waiting, no input, file operation or initialization is pending, and real activity has been idle for the configured interval. For node allocations Core records the first root or child terminal transition with the database clock in the same transaction. Candidate filtering and the Session-locked recheck compare elapsed database time with the idle duration, and the initial snapshot retention deadline is anchored to the same database observation, so Core and database host clocks need not agree. Native completion timestamps stay unchanged in public history but never drive idle admission, and heartbeats never reset activity. Before acknowledging a planned suspension, the daemon closes admission and drains native cleanup, output receipts and file work.

The Worker lease, the Session lock and the per-node gates own suspension for every provider. New Turn claims, file-write intents and capture admission serialize under the Session lock and share one compute-phase check; new pending work cancels a capture and wakes the same source. Normal preparation waits for the compute phase to be running, after the authenticated resume handshake, and pending input stays pending when its promotion conflicts with a lifecycle transition. Compute phases and revision-checked receipts live on the allocation. Core persists quiesce, capture and restore intent before the effect, only a fresh receipt performs a capture or restore, and recovery observes the exact attempt without retrying an unknown creation, capture or restore. A consumed snapshot never rolls a running generation back. Deletion, revocation and retention expiry win over wake, up to the final database compare-and-swap, and unknown cleanup identities are kept until owned resources are confirmed absent. Consumed artifacts and old compute are deleted, so suspension cycles never build a chain of writable disks.

Expand Down
Loading
Loading