diff --git a/AGENTS.md b/AGENTS.md index 2467c0966..853a2837a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -16,7 +16,7 @@ OpenAgentCore is protocol-first and modular. Core orchestrates operations that p | Boundary | Protocol code | Protocol doc | | --- | --- | --- | | Application–Core (`/v1`) | Types in `contracts/agents-api/v1/` and route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/openapi.yaml` | [Agents API guide](docs/api/public-agent-api.md) | -| Web and operators–Core (`/core/v1`) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/core.openapi.yaml` | [Core API](docs/api/README.md#core-api) | +| Web and operators–Core (`/core/v1`) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/core.openapi.yaml` | [Core administration API](contracts/agents-api/admin-api.md) | | Nodes and daemons–Core (`/api/v1` HTTP routes; the node and daemon wire protocols are separate rows) | Route annotations in `services/core/internal/api/`; `make openapi` generates `contracts/agents-api/runtime.openapi.yaml` | [Machine connection API](docs/api/README.md#machine-connection-api) | | Core–Sandbox Provider | `services/core/internal/sandbox/sandbox_provider.go` | [Sandbox Provider guide](docs/sandbox-provider.md) | | Core–sandbox node | `services/core/internal/sandbox/node/wire.go` | [Node generation protocol](contracts/agents-api/node-generation-protocol.md) | diff --git a/README.md b/README.md index 3ab763d42..7a6693a7a 100644 --- a/README.md +++ b/README.md @@ -58,7 +58,7 @@ Core exposes two APIs: | API | Path | Used by | | --- | --- | --- | | **[Agents API](docs/api/public-agent-api.md)** | `/v1` | Your applications. Same protocol as [OpenAI's Agents API](https://developers.openai.com/api/docs/guides/agents-api/overview) | -| **[Core API](docs/api/README.md#core-api)** | `/core/v1` | Operators, through Web | +| **[Core API](contracts/agents-api/admin-api.md)** | `/core/v1` | Operators, through Web | Core keeps all state. The Runtime runs the chosen harness inside the Environment. Each connection is a defined protocol, so any part can be replaced on its own. See diff --git a/README.zh-CN.md b/README.zh-CN.md index 7474308ea..3a9c867a7 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -55,7 +55,7 @@ Core 对外提供两组 API: | API | 路径 | 调用方 | | --- | --- | --- | | **[Agents API](docs/api/public-agent-api.md)** | `/v1` | 你的应用,与 [OpenAI 的 Agents API](https://developers.openai.com/api/docs/guides/agents-api/overview) 协议一致 | -| **[Core API](docs/api/README.md#core-api)** | `/core/v1` | 管理员,通过 Web 调用 | +| **[Core API](contracts/agents-api/admin-api.md)** | `/core/v1` | 管理员,通过 Web 调用 | 所有状态都由 Core 保存;Runtime 在 Environment 中运行所选 Harness。各部件之间都通过既定协议连接, 任何一个都可以单独替换。详见[架构说明](docs/architecture.md)。 diff --git a/contracts/agents-api/README.md b/contracts/agents-api/README.md index e4a59ded1..ebc926eff 100644 --- a/contracts/agents-api/README.md +++ b/contracts/agents-api/README.md @@ -1,921 +1,149 @@ -# Agents API contract - -Start with the [API surface index](../../docs/api/README.md) to distinguish public -application APIs, Web administration and Runtime transport. This document is the -protocol coverage/evidence record, not the Web integration guide. - -See the [58-operation evidence inventory](operation-evidence.md) for observed official behavior, local verification and remaining unknowns. -The [list-query comparison](list-query-semantics.md) distinguishes measured order -errors from unresolved range, cursor and lookup semantics. -The external reference is [openai-python beta/agents](https://github.com/openai/openai-python/tree/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents), -pinned in `upstream.json`. Its resource methods, corresponding types, pagination -and streaming helpers define the compatibility target. This directory records -the boundary; it does not imply that every upstream feature is implemented. -`upstream-routes.json` and `upstream-fields.json` are extracted from that SDK by -`scripts/extract-agents-api-upstream.py` (run it with the pinned SDK installed). -Contract tests require `openapi.yaml` and the live router to have exactly those -method and path pairs, every query parameter to be official, and every other -field to sit inside `x_agents_core` on Agents and Sessions. The -[API index](../../docs/api/README.md) identifies the extension fields; request and -response types live in [`v1/`](v1/). - -Parsar owns product Agents and Teams. This service owns upstream execution -resources, including reusable Agents and protocol subagents. The OpenAI Agents -Python **SDK** is a separate future dependency for business Team orchestration in -Parsar, not the HTTP contract. Design rules live in -[AGENTS.md](../../AGENTS.md#public-api). - -The [resource selector and error qualification](resource-selector-semantics.md) -records nullable Skill references and source Files not-found parameter fields, -with official observations separated from Core acceptance. - -## Implementation direction - -Adapter and persistence design follows the -[design rules](../../AGENTS.md#complexity-stays-in-the-adapter). -The [harness contract and parity baseline](harness-onboarding.md) describes equal-engine -registration, qualification and shared acceptance. -Verify configuration against actual execution: response defaults must not merely -describe values the adapter never applied. - -The Worker persists fenced authenticated daemon connection observations and pinned -Environment-event snapshots through the existing execution owner. Session reads -and live SSE also expose safe `self_hosted` output and reservation-owned connection -actions. Public self-hosted creation accepts initial text or empty Sessions on the three enabled harness profiles; -initial input reserves work while returning the connection target promptly. -Later idle text submissions wait for preparation/admission. Cancellation-only events -reuse durable admission without creating work or retargeting retries; pending -pre-Turn input still blocks new cancellation. HTTP acceptance does not establish -native completion or process quiescence. Environment retrieval -exposes durable status and safe empty installation metadata for that profile. -Non-deferred functions and homogeneous result-only batches reuse the existing -callback/application path, with explicit call identity and no new Turn on results. -These callbacks are not installed Environment resources. -Message-only batches append to an active Turn under the same Session lock that -reserves idle work; retries retain their original target through completion and later work. -Populated installation metadata, mixed input and full lifecycle conformance remain -unimplemented; see the [Environment scope](environments.md). -Initial messages commit with creation and a connection action; an initial deadline -failure is queryable before a Turn exists. Ordinary and streamed creation share this path. - -The three-harness Docker V1 MVP is accepted: Codex, Claude Code and MiniMax Code -share the execution/workspace contract, with independent Core/database deployment, -Files/Artifacts, cancellation and owned-history continuation. Optional features -still differ. See the [accepted scope and evidence](#accepted-milestone-and-evidence). -The same three harnesses passed historical Core-managed E2B V1 qualification in -PR #705. That original route is historical evidence; it does not qualify the later -user-managed enrollment or current deployment-level E2B configuration. See the -[current hosted provider contract](sandbox-deployment.md). The [user-managed qualification](harness-capabilities.md) -records separate real deployment acceptance and its exact scope. -Select further work only within current user authorization. Parsar cutover and -business Team orchestration are separate from protocol coverage. - -## Upstream resource inventory - -This inventory is based on the pinned Python source, not our generated OpenAPI. -It contains 42 distinct HTTP operations in 15 resource classes, excluding async -duplicates, overloads and client-side helpers. There are 42 handler entries; the six -[Subagent reads](subagents.md) have three-harness Docker workflow qualification. The separate -general `/v1/files` source-file API and `/v1/skills` resource/version operations -are outside this 42-operation count. - -An implemented route is not complete semantic compatibility. **Accepted** below -means a recorded workflow passed under a specific profile; **partial** means some -variants work; **missing** means no implementation; **unverified** means behavior -has not been shown to match upstream. Do not convert the route count into a -compatibility percentage or treat a Docker result as E2B qualification. - -The [execution and tools matrix](execution-tools.md) records message input, -structured output, tool configuration/results, discovery and required-action -recovery by operation, with exact qualified profiles and remaining gaps. - -Paths below are SDK resource paths beneath `client.beta.agents`, except Skills -and Versions under `client.skills`. Method names use the Python SDK. Vault HTTP -paths start at `/vaults`, not `/agents/vaults`. - -| Resource | Upstream operations | Current coverage | -| --- | --- | --- | -| Root reusable Agents | create, retrieve, update, list, delete | Partial create/retrieve/update/list/delete and Session references; configuration/error gaps remain | -| Skills and Versions | create, retrieve, update default, list, delete, content | [Tenant-owned encrypted bundles and hosted references](environments.md#skills); [default metadata/content and deletion evidence](file-resource-semantics.md), qualified upload limits and unresolved semantics | -| sessions | create, retrieve, update, list, delete | Create (ordinary/live), retrieve, list with root-Agent filter, metadata-only update, [idle-only public deletion](official-semantics-alignment.md#session-deletion-lifecycle--september-23) with idempotent owner repeat and owned Docker cleanup; user-managed compute stays caller-owned; general physical cleanup and exact hosted semantics remain open | -| sessions.events | create, stream | Text/cancel/function-result admission and live events; function-action state snapshots supported | -| sessions.turns | retrieve, list | Implemented reads; lifecycle conformance still partial | -| sessions.items | list | Partial Item variants | -| sessions.artifacts | retrieve, list, delete, content | Shared output capture and immutable stored reads/deletion on accepted Docker profiles and [qualified user-managed workflows](harness-capabilities.md) (prior Core-managed E2B evidence remains historical), including retained downloads after Runtime loss. [Aligned](official-semantics-alignment.md#artifact-capture-and-listing--september-23) output symlink skipping, unchanged-path non-republication, the list envelope and malformed filters; exact upstream defaults/errors, hard-link/special-file capture and cancellation-edge parity remain unverified | -| sessions.subagents | retrieve, list | [Three-harness Docker reads, native lifecycle limits and real evidence](subagents.md); full multi-agent semantics remain partial | -| sessions.subagents.items | list | Qualified own-child history reads; [limit clamping and the list envelope](subagents.md#subagent-visibility) aligned; full Item variants remain partial. Child work is not streamed on the Session, as observed officially | -| sessions.subagents.turns | retrieve, list | Implemented; child Turns carry the Session's Agent ID and are not Session Turns | -| sessions.subagents.turns.items | list | Implemented; scoped persisted reads | -| environments | retrieve | Three-harness colocated self-hosted implementation and qualified Docker hosted profiles: durable status and safe initial-file metadata; other installation inventory and full lifecycle parity remain gaps | -| environments.files | create, list | [Bounded live listing and inline/source-file creation](environment-files.md) on qualified Docker workspaces; [user-managed enrollment](harness-capabilities.md) reuses the local implementation with separate real public acceptance. [Aligned](environment-files.md#wire-alignment--september-23-2026) the 201 status, page envelope, query keys, empty pages for non-directory paths on local workspace readers, sampled path/token errors and pending hosted rejection; [aligned](environment-files.md#write-semantics--september-23-2026) parent creation, no-replacement and the 5 MiB inline bound; recursion and other errors remain partial | -| environments.templates | create, retrieve, update, list, delete | [Reusable network, files, env/setup/packages, inline/referenced Skills and Session snapshots](environments.md#templates); other initialization and full semantics remain gaps | -| vaults | create, retrieve, list, delete | Create/retrieve/list/delete with independent tenant persistence, stored status filtering, atomic Credential cascade and frozen Session attachments; archive semantics and full hosted lifecycle parity remain missing | -| vaults.credentials | create, retrieve, update, list, delete | Static-bearer and OAuth create/retrieve/list/replacement/deletion with scoped encrypted storage and dispatch-time refresh; Session attachment and exact-URL HTTPS MCP binding; archive semantics and full hosted lifecycle parity remain missing | - -## Core extension inventory - -The operations below are implemented Core extensions. They are excluded -from the 42-operation upstream inventory and must not be counted as OpenAI Agents -compatibility. None is under `/v1`: [upstream-routes.json](upstream-routes.json) -pins the exact `/v1` route set, and `/v1` objects carry Core fields only inside -`x_agents_core`. - -| Extension | Operations | Current coverage | -| --- | --- | --- | -| Administration | `/core/v1/**` | Core-key Projects with shared keys, read/delete projections, executor credentials, summary, metrics, runtime observations, audit and sandbox deployment. Separate from caller authority; see [Administrator API](admin-api.md). | -| API-key provenance | Project-scoped `/core/v1/projects/{project_id}/resource-owners` and `/write-operations` | Batch ownership and retained cursor-paginated write history; see [write provenance](write-audit.md). | -| Runtime observations | `GET /core/v1/sandbox/runtime-observations`; `GET /core/v1/projects/{project_id}/sessions/{session_id}/runtime-observation` | Administrator only; the `/v1` forms are removed. Current, read-only Session contexts labelled by Project, with stable Session-keyset pagination, bounded concurrent sampling, Docker and microsandbox metrics, explicit unsupported/unavailable states, strict `AdminClient` projection, and no lifecycle mutation. Kubernetes, E2B, self-hosted telemetry, and automatic idle policy remain unimplemented. See [Runtime observation API](runtime-observability-api.md). | -| Runtime history | `GET /core/v1/projects/{project_id}/sessions/{session_id}/runtime-history` | Administrator only; the `/v1` forms and the capability read are removed. Optional backend-neutral bounded Project/Session-scoped history contract with allocation/incarnation fencing, explicit coverage and strict client projection. Disabled by default until a production Reader and qualified periodic collection are configured; Durable Web rendering remains pending. See [Runtime history API](runtime-history-api.md). | - -For each resource, verify the referenced request/response unions and observable -behavior, not just the route. Non-text initial input, configuration -options, text/image content, function results, environment variants, full Item/SSE -variants, defaults, field omission/nullability and errors need their own cases. -Use strict official-client tests plus raw HTTP assertions; SDKs can accept extra -fields and cannot prove that reported configuration matches the running engine. -Where SDK types or public documentation do not establish behavior, record the -uncertainty and obtain upstream evidence before marking it conformant. Temporary -unsupported errors are implementation gaps, never evidence of full compatibility. - -## Accepted milestone and evidence - -The three-harness Docker V1 milestone was accepted on 2026-09-20 after -[PR #703](https://github.com/MiniMax-AI-Dev/parsar/pull/703). A fresh source-free Core -package and main-built daemon passed fixed Python SDK 3.13.0/raw HTTP/real Kimi K3 -regressions for Codex, Claude Code and MiniMax Code. The profiles use independent -execution databases and no Parsar services or product database. - -These final regressions supplement, rather than repeat, every earlier check. -Codex's full 26-check deployment evidence, Claude's hosted/Artifact/crash/security -evidence, and MiniMax's real MiniMax Artifact and Core/Runtime SIGKILL evidence -retain their exact tested scope. The full gate, focused race checks and fresh -Astra high review passed for the candidate; no additional live run is claimed by -this documentation update. Evidence on `zju_a100_2`: - -- `~/.parsar/remediation/20260920/three-harness-mvp/REPORT.md` and - `acceptance-results.json`: final baseline, runs, reused evidence and cleanup. -- `~/.parsar/remediation/20260919/docker-mvp/` and `claude-v1/`: - preceding deployment and safety acceptance. -- `~/.parsar/remediation/20260919/mcode-workspace/`: MiniMax workspace native - isolation and recovery (historical). - -Qualification is Linux amd64 Docker V1, not arbitrary host isolation, production -HA, E2B, or Anthropic-model acceptance for Claude Code. No exactly-once guarantee -is made for future model choices: the recorded Kimi continuation limitation is a -new model-issued command after a recovery prompt, not automatic API replay. - -### E2B V1 qualification - -This is historical evidence for the original Core-managed E2B route. Current -deployment-level E2B configuration is documented in the -[sandbox deployment contract](sandbox-deployment.md). -This older run does not qualify later enrollment/configuration changes or transfer -caller-owned compute to Core. - -PR #705 (`9cd1c46c7fab6eeef8cb35ce71f1ea0ca2cf8bc1`) separately qualified -Codex, Claude Code and MiniMax Code with actual E2B and real Kimi/MiniMax APIs. -Fixed SDK/raw HTTP acceptance covered independent Core/database deployment, -Files/Artifacts, auth/tenant and native credential/history isolation, cancellation, -Core/Runtime crashes, exact-history continuation without replay and native network -restrictions. All three runs completed cleanup without fallback. The five Provider -operations passed real lifecycle/race acceptance; `make check` and fresh independent -Astra high review passed. The merged tree matches the accepted candidate. - -Evidence: `~/.parsar/remediation/20260920/e2b-runtime-v1/` on `zju_a100_2`, including -`acceptance-results.json`, `provider-real-final.log`, `make-check.log`, -`blind-review.md` and exact image/template build pins. This uses the existing -colocated Runtime contract, with no public resource or protocol expansion. -The qualified Linux amd64 templates require a reachable HTTPS/WSS Core and a -minimum two-hour renewable E2B lease. Expiry destroys volatile workspace/history; -unknown effects cannot authorize recreation or replay. Pools, migration and -user-managed enrollment remain outside this qualification. - -### Remaining protocol work - -| Area | Missing or unverified scope | -| --- | --- | -| Subagents / multi_agent | Six reads and same-child recovery have three-harness Docker evidence; optional native operations, live child progress, full lifecycle/interactions and tool combinations remain explicit gaps | -| Environment Templates | Unsupported restricted hostname forms and exact hosted errors remain gaps. Referenced null network and capability-list selection follow [qualified inheritance rules](environments.md#inheritance). Template-reference env/files/commands/packages composition follows [qualified field rules](environments.md#inheritance). CRUD/list, files, env/setup/npm/Python, inline/referenced Skills, Plugins, workspace capability directories and Session references have recorded coverage. System-package and inner-isolation evidence is historical; current `packages.system` rejects and the daemon adds no sandbox. Environment Plugin MCP transport and placement limits are [listed separately](environments.md#plugin-mcp) | -| Input and configuration | Non-text initial input, broader content/configuration unions and reasoning/verbosity combinations; [structured output](execution-tools.md#structured-output) has qualified Claude function profiles on none and Core-managed Docker openai_hosted, with other combinations remaining gaps | -| Tools and interactions | [Deferred discovery qualification](execution-tools.md#deferred-function-discovery), other tool types, effective tool-set enforcement and result/cancel publication ordering; MiniMax public functions and service-origin MCP remain unsupported | -| Vault and Credentials | Archive semantics, in-flight token withdrawal and exact hosted selection/error behavior; static/OAuth CRUD, replacement and scoped dispatch-time refresh are implemented (see credential guide for qualification) | -| Existing resources | Full Item/SSE/Usage variants, omitted/null/default/error semantics, pagination and overlapping lifecycle behavior beyond recorded cases | - -An implementation gap and an unknown upstream behavior require different follow-up -work. Retain both explicitly; a restrictive local policy or successful SDK parse -cannot establish upstream equivalence. This inventory describes merged behavior. - -## Public semantics - -The [September 22 wire comparison](official-semantics-alignment.md) records the -bounded official-service observations, aligned responses and remaining differences. -It supplements the fixed SDK baseline; current documentation does not silently -upgrade the protocol. - -- Credentials use `POST /vaults/{vault_id}/credentials` and - `GET /vaults/{vault_id}/credentials/{credential_id}`. The static profile accepts - `static_bearer` with required string token and HTTPS destination, plus a - required name trimmed to 1–256 UTF-8 bytes. Tokens remain opaque and nonempty; explicitly empty tokens are rejected before - mutation, following the sampled official create/update behavior. The local URL profile - excludes userinfo/fragments and preserves queries without normalization or network - contact. Public metadata contains identity, owning Vault, name, timestamps and - auth type/destination; it never returns tokens or ciphertext and can be read - without the encryption key. Missing encryption configuration locally rejects - creation/replacement with 503. Attached Sessions can use static credentials for - exact-URL HTTPS MCP. OAuth grants use the same binding plus - [scoped refresh and replacement](../../services/core/oauth-credentials.md); - storage-key rotation, archive behavior and key scopes remain gaps. See the [credential storage guide](../../services/core/credentials.md) - for encryption and operational limits; this does not establish complete Credential - or hosted error/retry compatibility. -- `GET /vaults/{vault_id}/credentials` lists only safe metadata, with parent and - cursor ownership checked within the authenticated project and requested Vault. - It uses the same paging/filter grammar as Vault listing below. Credential status - is stored separately from Vault status, defaults active, and never appears in - the public response. Both active and archived Credentials are included by default; - synthetic archived fixtures establish read/filter behavior, not archive lifecycle. - No token/ciphertext column, decryption, execution or encryption key is needed. - An inaccessible parent is not returned as an authorized empty collection. Existing - create/retrieve/token replacement and dispatch rules are unchanged; exact hosted - errors and pagination under concurrent mutation remain unverified. -- `POST /vaults/{vault_id}/credentials/{credential_id}` replaces a static token using - only required `auth.type=static_bearer` and string `auth.token`. It preserves opaque - strings, rejects missing/null/type/extra-field mutations and returns safe metadata. - Ciphertext/update time change atomically within the same tenant/Vault/ID/type/URL; - name, destination, identity, creation time and Session bindings are unchanged. No - old-token decryption or MCP call occurs. Subsequent dispatch reads use the committed - replacement; already-resolved requests may retain the old token. OAuth partial - replacement follows its pinned union and authenticates the stored grant. - Storage-key rotation, hot reload/revocation and exact hosted concurrent-update, - timestamp and retry semantics remain gaps. -- `DELETE /vaults/{vault_id}/credentials/{credential_id}` returns only `id`, - `deleted: true` and `object: vault.credential.deleted`. It removes one owned row - and its ciphertext without an encryption key. Local retrieval/update/repeated - deletion then return 404; lists omit it. Frozen Session choices and history remain - intact, while subsequent secret lookups fail without credential reselection or - anonymous fallback. Already-resolved tokens and running Sessions are not revoked. - Archive relationships, exact hosted post-delete visibility and repeat/error - semantics are unverified; physical storage erasure is not established. -- Vaults use `POST /vaults`, `GET /vaults` and `GET /vaults/{vault_id}` with the same project/tenant - authentication and Beta header as other resources. The response contains only - `id`, `object: vault`, `created_at`, `name` and `metadata`. Omitted name stays null; - explicit null is rejected. Supplied strings are trimmed and must contain 1–256 - UTF-8 bytes. Omitted/null metadata becomes `{}`; a non-string value returns - `invalid_request_error` with param `metadata.`. - Session-specific metadata pair/character limits do not apply. The existing - 64 KiB encoded metadata and 1 MiB HTTP body bounds are local implementation - limits. Creation does not start execution. Retrieval maps missing, malformed and - foreign IDs to the same local not-found response. Exact hosted error/retry semantics, - restricted-key scopes and archive lifecycle remain unverified or unimplemented; - this is not complete Vault compatibility. Listing accepts `after`, creation order - (default `desc`), a default limit of 20 clamped to 1–100, and scalar or SDK bracket-array - `status` filters. Both `active` and `archived` are included by default. The private - classification is stored, never returned; existing/new Vaults default active. - Synthetic archived fixtures prove read/filter behavior only. No public archive - writer or delete-to-archive mapping is implemented. Equal creation times use ID - order locally. A repeated scalar status is rejected; a scalar combined with - `status[]` filters by their union. Other hosted query errors and pagination over - changing data remain unverified. -- `DELETE /vaults/{vault_id}` returns `id`, `deleted: true` and `object: vault.deleted` - after project-scoped parent removal and atomic cascade of all stored Credentials. - It needs no encryption key or execution connection. Local parent/child reads, - repeated deletion and new references return 404; lists omit the removed resources. - Existing Session snapshots and recorded retries retain their IDs and selections. - Subsequent secret lookup fails without reselection; already-dispatched tokens are - not withdrawn. Archive relationships, exact hosted visibility/concurrent errors, - provider revocation and physical erasure remain separate gaps. -- Reusable Agents use `POST /agents` and `GET /agents/{agent_id}`. Keep their own - identity, timestamps and metadata separate from Session effective configuration. - On creation, omitted/null name and instructions resolve to null, metadata to `{}`, tools to - `[]`, text to ordinary/medium, and multi-agent settings to disabled/null. Enabled - multi-agent settings default to six concurrent subagents. Function defer-loading - defaults to false and programmatic tool calling to true. Saving these values - does not itself admit a native execution. Session references are admitted separately. -- [Explicit disabled tools](execution-tools.md#web-search-and-programmatic-tool-calling) can be saved, used inline or resolved from saved Agents: - `web_search.mode=disabled` and `programmatic_tool_calling.enabled=false`. Search - responses include `context_size=medium` for omitted/null size, nullable domains - and location; an empty domain list stays empty. Saved Agents keep every pinned - search mode, saving omitted/null mode as `live`; only explicit disabled mode is - qualified, so Session admission rejects enabled search unless the Session - replaces the saved tools. Sessions reject enabled programmatic execution, - including the default true on a supplied PTC declaration. An omitted PTC declaration preserves native behavior: this is an - approved difference from the official default-on behavior, not full compatibility. - Core carries the frozen disabled intent through the common Runtime contract; - native translation and inventory restrictions stay in adapters. Codex checks - managed requirements before new/resumed execution; conflicting forced features - reject before model input. Claude and MiniMax use their restricted tool profiles. - No independent executor or model/tool loop is introduced. Resource defaults and - hosted error parity beyond this supported subset remain unverified. -- `POST /agents/{agent_id}` updates only supplied fields. Omitted fields remain - unchanged; metadata replaces all pairs and null/empty clears it. Name/instructions - null clears them. Concurrent updates preserve unrelated fields. Existing Session - snapshots and their recorded creation-retry identity remain unchanged; new Sessions - resolve the latest saved configuration. No-field updates leave timestamps unchanged. - Nested fields currently replace whole values and null uses the saved defaults; - hosted nested/null behavior, no-op timestamp policy and exact errors remain - unverified. This operation shares the existing saved-configuration coverage gaps. -- `DELETE /agents/sessions/{session_id}` returns the canonical `id`, - `object=agent.session.deleted` and `deleted=true` after durable public removal - of a durably idle or failed Session without required actions or pending input. - A queued, running or waiting root Turn or pending input returns 409 - `conflict_error` without any change; callers cancel first and delete once idle. - Subagent child Turns and pending Environment file writes do not block deletion. Session/Turn/Items - reads, live streams, metadata updates and new input exclude the resource. - Existing streams close on observing removal without an invented deletion event. - Creation keys remain reserved (local 409); the owner's repeated deletion returns - the same confirmation and missing or foreign deletion returns 404 - ([batch record](official-semantics-alignment.md#session-deletion-lifecycle--september-23)). - Qualified managed Docker deletion also reclaims its owned Runtime; - broader physical SQL/native history cleanup, immediate native quiescence and - exact hosted error/retry/overlapping-stream semantics remain unverified or - unimplemented. Shared devices, saved Agents and other Sessions are independent. -- `DELETE /agents/{agent_id}` removes the tenant-owned saved configuration and - returns `id`, `object=agent.deleted`, and `deleted=true`. Existing Sessions and - history are retained; recorded creation retries recover their frozen snapshot, - while new references to the source fail. Local missing/repeated deletion returns - 404. Exact hosted errors and overlapping creation/deletion ordering are unverified. -- `GET /agents` lists tenant-owned reusable resources with `after`, `limit` and - `order` (default `desc`). It uses creation-time/ID keysets and the same resource - mapping as retrieval. Limit 0 is treated as 1 and larger limits as 100; negative - and non-integer limits reject. Pages contain up to 100 resources, with `has_more` - and the final resource ID guiding continuation. - The local default is 20. The list envelope includes `object`, `data`, `has_more`, - `first_id` and `last_id`; empty pages use null IDs. The pinned SDK omits null - limits and empty cursors. Exact upstream default/cap, empty-envelope nullability - and error taxonomy remain unverified; SDK auto-pagination does not prove them. -- Saved Agent model-default reasoning resolution remains missing: an omitted effort - stays unresolved rather than being populated from a guessed model default. An - explicit effort/summary is retained. Omitted/null service tier currently follows - the service's `auto` policy; complete upstream-default/error/retry conformance is - unverified. HTTP MCP with `service` origin (omitted or null on HTTP transport is - saved as `service`) and - boolean `required` (default false) supports saved configuration and Codex `none` execution, - with Claude SDK - also supporting its qualified `none` subset. V1 `self_hosted` explicitly rejects - service-origin MCP; the old remote combination is retired. Required initialization - additionally needs `mcp_http_required` on the pinned native profile. Native root - thread creation/cold resume must initialize required servers before a native - Turn starts; failure cannot silently replace retained history. Public acceptance - or queued work does not prove native readiness. Hosted creation timing/error - parity and continuing server health remain unverified. - The saved HTTP transport includes `headers:{}`; effective Session transport omits - headers. Omitted/null `allowed_tools` is unrestricted; `[]` denies all tools. - Session `vault_ids` attaches tenant-owned Vaults. Explicit `credential_id` must - belong to an attached Vault and match the exact HTTPS URL; omission/null selects - one matching static or OAuth credential, zero stays anonymous and ambiguity is a - 409 `conflict_error`. Selection errors use the observed official messages, with - one message for missing, foreign and unattached references. Session projections - show an implicitly selected credential ID in the public field; the immutable - stored selection and caller intent are unchanged. Scope is - rechecked before dispatch-only decryption; authenticated execution requires the - separate bearer capability and never downgrades on failure. Exact URL/selection - timing and hosted redirect behavior remain local or unverified. - Other MCP variants and enabled web-search execution remain gaps, not changes to the pinned target - or claims of complete resource coverage. - -- Use `/agents/sessions` beneath the configured API base URL, bearer authentication - and `OpenAI-Beta: agents=v1`. Do not introduce a competing `/sessions` surface. -- Session creation takes an environment and inline agent configuration or a saved - agent reference. The saved ID and effective configuration are copied into an - immutable Session snapshot. Omitted fields inherit; supplied objects and arrays - replace the entire field ([configuration guide](https://developers.openai.com/api/docs/guides/agents-api/configuration)). - Tools null clears the list as specified by pinned `session_create_params.py`. - Saved metadata never becomes Session metadata. A model name is not a daemon engine name. - Current admission uses the qualified harness profile for multi-agent execution, - implicit reasoning, tier `auto`, ordinary text and non-deferred functions. - Enabled multi-agent/function combinations remain unsupported. Unsupported saved settings fail before - Session persistence, unless replaced by supported overrides. Other native options - remain implementation gaps, not excluded protocol variants. - Fixed SDK/raw HTTP checks cover inherited/overridden configuration, tenant ownership, - source preservation, independent Session snapshots, retries and service restart. - New saved-reference Sessions record caller intent before source lookup. Matching - creation retries recover their accepted snapshot even after source update/deletion; - changed overrides conflict. Retries also require the original typed creator. - Known creators without recorded request intent retain resolved-hash behavior; - records without creator identity reject retries. Neither identity is backfilled. These local - retry rules are not verified hosted semantics. - In the pinned `session_create_params.py`, `stream` defaults to false and neither - `stream` nor `agent_id` permits null. Metadata omission/null defaults to an empty - map; individual values must be strings, including valid empty strings. Validate - these distinctions before persistence rather than coercing null to Go zero values. -- `POST /agents/sessions/{id}` updates metadata only: an empty update body - rejects; `metadata: null` or `metadata: {}` clears it, and an object replaces all pairs. Apply the same string - and character limits as creation. Preserve execution state, effective configuration - and the original creation retry identity. Fixed SDK/raw HTTP checks cover these - distinctions, tenant isolation, active Session reads and restart persistence. -- `GET /agents/sessions` accepts `after`, `limit` (default 20; 0 is treated as 1 and - values above 100 as 100), `order` - (default `desc`) and optional `agent_id`. The filter matches the immutable root - Agent ID, including inline IDs and Sessions whose saved source was changed or - deleted. Filter before pagination within the authenticated tenant; no source - lookup is required. Omission lists all Agents. Empty filters, same-tenant cursors - outside the filter and exact hosted errors/defaults remain unverified. -- `AgentSession` includes the effective agent/environment, Unix-second timestamps, - `object: agent.session`, metadata, required actions, status, usage and vault IDs. - A Session remains reusable after its current Turn completes. -- Input, cancellation and function results are submitted through session events. - Turns are queried through `/agents/sessions/{id}/turns`; do not invent turn-create - endpoints. Event submissions support the `Idempotency-Key` header. -- Per the [official Session guide](https://developers.openai.com/api/docs/guides/agents-api/sessions), - input steers an active Turn and starts a new Turn when idle. Streams are live-only; - recover missed work through persisted Session/Turn/Items queries, not assumed SSE - replay. Internal input ordering is not a public event-stream cursor. -- List operations use the upstream `after`, `limit`, `order` and resource-specific - filters. Stream events preserve the upstream discriminators and payload shapes. -- The upstream self-hosted environment includes an exec-server `remote_url`. - V1 retains that public resource field while explicitly selecting our private - daemon transport. It does not claim stock `exec-server` wire interoperability; - public resource semantics require independent acceptance. -- Environment retrieval returns `object: agent.environment`, its ID/type, durable - resource status and required non-null `files`, `plugins` and `skills` arrays. - Hosted initial files report safe frozen metadata; empty arrays do not - describe native discovery or workspace files created by commands. Unknown - installation configurations are rejected, not reported as empty. Reads use the - owning live Session's project partition and do not require execution setup. - Remaining unsupported installation configuration, full hosted lifecycle and - exact hosted error semantics remain gaps. - -[Environment Templates](environments.md#templates) provide tenant-owned CRUD/list -and immutable Session resolution through common Environment preparation. -The `x_agents_core.environment` extension supplies that configuration to either -placement; self-hosted machines never need a managed allocation. See -[shared preparation qualification](harness-capabilities.md#environment-preparation). They do not -select an E2B image or make unsupported initialization executable. -[Harness capabilities](harness-capabilities.md#environment-preparation) records composed preparation qualification by Harness and placement. - -## Delivery and verification - -| Capability | Current state | -| --- | --- | -| Independent deployment | Source-free Core package and separate execution PostgreSQL ownership; Docker-hosted and user-managed Runtime colocate daemon, selected harness and workspace; Core owns Docker only; no Parsar dependency | -| Saved Agents and Sessions | Saved Agent routes, immutable inline/referenced Session configuration, metadata updates, root-Agent filtering and scoped cursor pagination | -| Public execution | Initial/later text, active input and cancellation through Codex, Claude Code or MiniMax Code; Codex/Claude additionally support qualified public functions; see profile limits below | -| Required actions | Persisted function calls/results/application receipts and pending Environment connection actions; Session reads and live snapshots expose the two pinned variants. See the [operation matrix](execution-tools.md) for qualification and unresolved timing | -| Public recovery and SSE | Persisted Turn/Items queries and partial Usage; live lifecycle/Item/text events, creation streaming and the official one-Turn tool-handler helper | -| Execution ownership | Immutable Session engine/device, durable input receipts and database writer fencing; uncertain claimed work fails on restart, without blind replay | -| Files and Artifacts | Bounded Environment listing and inline/file_id copies into qualified V1 workspaces; source-file lifecycle and immutable output capture/download/deletion; [Files limits](environment-files.md), [source limits](source-files.md) | -| Clients | Fixed Python SDK 3.13.0 and official Go SDK v3.61.0; raw HTTP and real provider acceptance supplement controlled tests | -| Release and product | Registry publication and Parsar cutover remain open; Docker hosting and user-side E2B provisioning are explicit opt-ins; business Team orchestration is deferred | - -### Public engine profiles - -`OAC_DEFAULT_HARNESS` supplies the default for new Sessions. The optional -[Core harness extension](model-execution.md#harness-selection) explicitly selects an enabled engine; -existing Sessions retain their immutable choice. `OAC_HARNESSES` explicitly -adds installed deployment profiles without requiring a managed Provider. Model -identity is independent. -All three profiles require implicit reasoning and service tier `auto`. Ordinary -text output is the baseline; [structured output](execution-tools.md#structured-output) has a -separately qualified Claude profile. Enabled `multi_agent` qualification is tracked separately in -[Subagents](subagents.md); other profiles continue to reject unsupported execution. Optional tools/configuration are qualified per -operation and placement; native support is not public admission by itself. - -| Engine | Qualified placements and limits | -| --- | --- | -| `codex` (default) | Qualified `none` and Docker `openai_hosted`; public functions with ordered text/image results; service-origin HTTP MCP on `none` only; Environment-origin HTTP uses the common workspace path; verbosity follows native policy | -| `claude_sdk` | Qualified `none` and Docker `openai_hosted`; medium verbosity, object-root function schemas and text or successful inline PNG/JPEG results; anonymous/static-bearer HTTP MCP on `none` (service origin) or a workspace (Environment origin), subject to qualification | -| `mcode` | Qualified `none` text and Docker `openai_hosted`; medium verbosity; Environment-origin HTTP MCP with null/omitted allowlist and optional initialization; public functions/service-origin MCP, image input and complete public usage breakdown remain unsupported | - -All three profiles implement user-managed `self_hosted` enrollment at `/workspace` -through our private daemon transport; [separate real acceptance](harness-capabilities.md) -records qualified deployments and limits. A `self_hosted` Session supplies its own -model provider in the request or through a saved Agent; deployment defaults apply -to `openai_hosted` and `none`, never to `self_hosted` ([model execution](model-execution.md)). Service-origin HTTP MCP is rejected on `self_hosted` and hosted local -placements. Explicit Environment-origin HTTP declarations use the same Runtime -binding path as Plugin MCP; see the [origin matrix](environments.md#public-mcp-connection-origin) -and [public qualification](harness-capabilities.md#tools). -The [Docker lifecycle](environments.md#hosted-openai_hosted) retains -workspace Files/Artifacts, cancellation and recovery. Managed isolation belongs to -the outer Environment; native tools use the starting account's permissions. -Configure immutable Runtime images explicitly: [Codex](../../services/core/deploy/codex/README.md), -[Claude](../../services/core/deploy/claude/README.md), -[MiniMax](../../services/core/deploy/mcode/README.md). -[E2B packaging](../../services/core/deploy/e2b/README.md) reuses the Runtime -with the official SDK; the user owns provisioning, renewal and destruction. - -The shared initialization path supports env/setup and user-directory npm/Python packages; -`packages.system` rejects explicitly and system dependencies must be preinstalled. -See the [evidence and limits](harness-capabilities.md#environment-preparation). Remaining -unsupported startup installations, unqualified restricted hostname forms and hosted -service-origin HTTP MCP remain outside these accepted profiles. Environment-origin -MCP Plugin qualification and transport limits are recorded in [Harness capabilities](harness-capabilities.md#environment-preparation). - -The [self-hosted profile](environments.md#self-hosted-self_hosted) uses -user-managed Runtime enrollment and remains distinct from Core-managed Docker. Product `claude_code` -is likewise a separate integration from the API's `claude_sdk` engine key. -Unsupported configurations fail before Session creation; unsupported results fail -before a batch write. Native capability claims cannot replace service profile -qualification, tenant authority or exact binding checks. An existing Session -never silently changes engine/device. See the -[HTTP MCP limits](../../services/core/README.md#http-mcp-execution). - -The Store's internal DTO is not the upstream response model. The API layer must -validate and resolve the upstream schema before persistence, and report only -supported options. For example, upstream metadata is limited to 16 pairs with -64-character keys and 512-character values; a storage byte limit is not a -replacement for that public validation. Violations return `invalid_request_error` -with the official `metadata` or `metadata.` param, and Agent configuration -protocol errors report their JSON path; see the -[configuration validation batch](official-semantics-alignment.md#agent-configuration-validation--september-23). -U+0000 in stored strings -is a local PostgreSQL limit and returns 400 without writing; see the -[validation error batch](official-semantics-alignment.md#validation-error-fields--september-23). - -Use the pinned official Python client against the actual service, with response -validation enabled, for supported Session/Turn/Items operations, pagination, streaming, -errors, idempotency and tenant isolation. A client import or permissive parsing -alone is not evidence of compatibility. Unsupported capabilities must be explicit -errors, not successful placeholder resources. Add any provider or engine-specific -extension separately from upstream fields and document it here when implemented. - -`openapi.yaml` is our generated supported surface; it is not the full upstream -specification. The shared Go wire types are in `v1`. Physical Session cleanup, broader image -message profiles, broader structured-output combinations, broader options/tools, remaining Vault lifecycle, -Subagents and environment/provider resources remain incomplete. Reject unsupported -requests explicitly; persisted saved configuration is not execution admission. - -### Native subagent control - -The pinned `types/beta/multi_agent_config.py` defines `enabled=false` as disabling -subagent tools. The dispatcher enforces that effective value with a typed daemon -policy and capability admission; the Codex adapter applies native feature controls -on fresh and resumed Turns. Operator feature preferences cannot re-enable them. -Controlled model-boundary tests check absence of direct/deferred subagent tools -while the official function workflow continues to run. - -Disabled `multi_agent` continues to remove native child tools. The six public reads -and their neutral identity/lifecycle/Turn/Item observations are described in -[Subagents](subagents.md); native qualification is tracked there. The service does -not infer closure from idle/unload, and public capability declarations alone do -not qualify an adapter. Environment and child tools have separate configuration; -`Agent.tools` is not assumed to enumerate every native utility. - -### Turn recovery reads - -`GET /v1/agents/sessions/{session_id}/turns` and retrieval by `turn_id` -return persisted Turn states using the pinned official client contract. Lists -support `after`, `limit` (1..100, default 20), and `order` (default `desc`). -The cursor is a Turn ID in the same tenant and Session. Failed turns expose a -generic `internal_error`, never raw engine diagnostics. `usage` exposes the latest persisted complete token breakdown, including cached input -and reasoning output. Missing measurements remain null; Session usage sums recorded -Turn measurements as best-effort usage, without estimating missing history; it is -null while any root Turn has not ended or once one ends with unknown usage. Session runtime state derives from the latest Turn. - -### Item recovery reads - -`GET /v1/agents/sessions/{session_id}/items` supports the same list controls, -with a stable Item ID cursor and first-observation ordering. Messages preserve -text, phase (null when none) and completion snapshots. Commands preserve reported output, exit -code, duration and working directory. MCP calls preserve server/tool identity, -arguments and structured results/errors. Dynamic functions have linked call and -result Items. Native file changes appear as `apply_patch` function calls with -reported changes as arguments; no result is invented when the engine reports none. -Web search exposes its supported action fields. - -Terminal Turns make unfinished Items `incomplete`; a failed tool does not imply -that the Turn failed. Native start/completion snapshots and Codex command-output fragments are -available when the daemon emits them; other -tool-output deltas remain unsupported. Tool output is visible to the Session's -authenticated tenant and may include the command's or tool's own diagnostic text. - -Reads use the durable index without reconstructing native journals. Existing -indexed history is preserved. Migration 15 refuses unindexed historical Turns. -Historical database conversion is unsupported; preserve the old data and install -separately. See [historical Item storage](../../services/core/README.md#historical-item-storage). -The retired archive format could not recover unrecorded message boundaries or -outcomes; those limitations remain in already indexed historical Items. -Other native variants, full reasoning coverage and Items mutation remain gaps. -Qualified child Items are described in [Subagents](subagents.md). Public submission -supports text messages, cancellation and function results. - -Legacy Done frames alone do not complete assistant Items. Aggregate answer text -is confirmed by successful Turn termination; failed Turns retain observed deltas -instead of treating adapter diagnostics as assistant output. - -Pagination orders by first-observation timestamp, then the Item's immutable -Session position and public ID. New Items retain observation order even when -timestamps match. The index also stores a zero-based output index per Turn for -streaming; inputs do not consume it. Updates and retries do not move Items -or change output indexes. Existing indexed history retains its recorded -deterministic order rather than guessing an unavailable original source order. - -### No-environment execution - -The dispatcher executes public `environment.type=none` on an authenticated, -bound host advertising `environment_none`. Codex uses `CODEX_EXEC_SERVER_URL=none` -and verifies native environment state before starting/resuming. Claude SDK uses -its restrictive profile with no built-in tools and only declared function callbacks. -A missing capability or unsupported native method fails rather than silently -allocating a local execution environment. Native state still lives on the host; -function callbacks may access their own resources. This is not filesystem isolation. -User-managed `self_hosted` uses an exact enrolled local Runtime instead. The -former native registry/Noise transport is retired. The pinned native source remains -a dependency reference, not a requirement to expose its executor protocol. -The [Environment assessment](environments.md) separates current boundaries from -historical native transport evidence. - -### Public execution admission - -`POST /v1/agents/sessions/{session_id}/events` accepts `agent.session.input.message` -with ordered user `input_text` content and [qualified image content](message-input.md), `agent.session.input.cancel` and -`agent.session.input.tool_result`. Successful atomic -admission returns 202 with no body, as observed from the official service. -An empty event array is an authenticated no-op: it creates no Turn or Item and -does not reserve an execution retry key. A retry -key identifies the entire ordered request; conflict does not partially admit it. -Messages start queued work or steer the active Turn. Individual input messages -remain distinct Items even when their text shares one native prompt. - -Enable the standalone daemon gateway to run the worker; without it, admission -returns 503. The worker selects capable same-tenant engine hosts, binds each Session -once, and runs at most four Turns concurrently. Queued cancellation needs no live -engine. Active cancellation waits for a native receipt; terminal completion can -win that race. Query Turn/Items to recover results after a stream interruption. - -The worker takes a database advisory lease; a second worker cannot start on the -same database. All execution writes use that lease connection and stop after loss -of ownership. Startup marks previously claimed Turns failed without replaying them -and retains queued work. Database fencing does not stop already queued native -commands, recover missing daemon frames or guarantee exactly-once external effects. -Session status reflects the latest persisted Turn; usage reports recorded measurements. - -Native verification uses `OAC_TEST_NATIVE_DAEMON_BIN`, `OAC_TEST_NATIVE_PROOF_DIR` under -`~/.oac/`, and `OAC_TEST_OFFICIAL_SDK_PYTHON` pointing to the pinned SDK environment. -The Store native integration test runs `tests/official_execution.py` against a real -HTTP handler, PostgreSQL, daemon and Codex with a synthetic model provider. - -Token measurements use the pinned SDK's `TokenUsage` fields. The optional daemon -`usage.tokens` supplies complete per-Turn counters; journal and terminal writes -replace that Turn's snapshot atomically. Repeated snapshots do not increase totals. -Codex publishes observed active-Turn snapshots before completion through this same -contract. Persisted measurements remain available after cancellation or worker -restart; measurements never received by Core cannot be recovered this way. -Unknown historical breakdowns are not backfilled, and a Session total includes only -recorded root-Turn measurements. Session Turn pages hold root Turns only; Subagent -Turn pages are not an additional accounting ledger. Costs and prices are outside -this execution contract. See [history, events and usage](history-events-usage.md) for client -recovery rules, native measurement limits and bounded official-service evidence. - -### Live events - -`GET /v1/agents/sessions/{session_id}/events` implements the official live-only -stream. Open it before submitting input. Session in-progress/idle/failed and Turn -created/in-progress/completed/failed/cancelled events carry transition snapshots. -Supported Items emit added/done events; assistant text emits content-part and -text-delta/done events. Inputs, including function results, have no output index. Function results emit -`item.added` and remain queryable; `item.done` only carries agent output. Public -function-result output/error retain the saved submission and field presence; -native error-to-text translation does not rewrite those fields. Completed text replaces -accumulated deltas; cancelled unfinished Items retain their partial content and -`incomplete` status. Codex command-output fragments use the pinned -`agent.output.command_execution_output.delta` event with the command Item ID and -stable output index. Draft Item output accumulates fragments; a supplied final -snapshot replaces it and is not emitted as another delta. Native output quotas and -text conversion apply, so the stream is not a byte-complete stdout/stderr capture. -Pinned native 0.153.4 can omit output emitted before its streaming subscription, -including from the eventual aggregate; this bridge cannot recover unobserved bytes. -That native gap remains open. Older peers may provide completion snapshots only. -Reasoning summaries and other -interim tool-output variants remain outside the supported surface. - -Events publish only after their transaction commits. An idle Session keeps its -stream open for later Turns. Reconnection starts at the latest committed position, -including when Last-Event-ID is sent; it does not replay missed work. Connect, -buffer new events, then retrieve saved Session/Turn/Items state to recover. Deduplicate -by Item ID and retain finalized Items when applying buffered updates. - -The internal buffer is limited to 256 events / 64 MiB per Session, with a single -oversized-event exception. A lagging reader receives a customer-safe `error` and -disconnects rather than silently skipping output. Slow socket writes time out -without blocking execution. Unsupported event variants are not implied by this endpoint. - -Internal function execution uses the same native daemon harness, with resolved -non-deferred definitions and Store result admission. It verifies ordered text/image -results, error text, application receipts, matching action/Item call IDs, cancellation -and native Session continuity. Codex supplies the model transport's default image -detail. This proof uses a synthetic model responder and the real daemon/Codex; -the public workflow below exercises the same native bridge through HTTP. Deferred functions, -other tool kinds and the native 64-definition limit remain compatibility gaps. - -Function-action read coverage uses persisted-call fixtures with the real service -handler, PostgreSQL and pinned official client. Call insertion and application -receipts update Turn/Session state atomically; duplicate notifications emit no new -state. Actions remain visible until the execution adapter acknowledges application, -or cancellation/terminal state removes them. This acknowledgement timing and the -exact sequence of repeated `requires_action` notifications are implementation -choices: the pinned source defines their shape but not that precise ordering. -Session state events contain `event_id`, `type` and `session`; Turn events retain -`session_id` and `turn_id`. There is no invented Turn `waiting` event. - -Public function-result admission is verified with the pinned Python client and -raw HTTP against a dedicated PostgreSQL fixture: required fields, nullable output -and error, ordered text/image output, variant rejection, atomic batches, scoped -access and retries after terminal state. This admission verification complements the native public workflow below. -The generated Swagger 2.0 document leaves the output union unconstrained because -it cannot express string-or-content-array unions; the pinned upstream types and -server validation define the supported alternatives. - -### Public function configuration - -Inline `agent.tools` accepts `function` definitions with the upstream -required name, description and JSON Schema parameter object. Missing -`defer_loading` resolves to `false`; null and other types are rejected. Omitted, -null and empty tool lists resolve to an empty list. The resolved tools are part of -the immutable Session configuration and creation retry identity. Saved-Agent -inheritance uses the same resolved tools. The bounded [deferred discovery path](execution-tools.md#deferred-function-discovery) -adds type-only `tool_search` for its qualified profile. Other discovery combinations, other tool kinds, -the native 64-definition cap and nonblank names of at most 512 bytes remain -compatibility gaps; repeated names and explicit non-object root types reject as -officially. Claude SDK additionally requires an explicit object root. It accepts text and -successful inline PNG/JPEG function results on `none`, Docker `openai_hosted` and `self_hosted`; -failed images and remote references -remain gaps. See [function image coverage](function-result-images.md). Codex internal Goal/Skills/user-input/discovery semantics need -upstream evidence; their presence alone does not prove a tool-set mismatch. - -The worker selects a same-tenant host advertising `function_tools` for configured -Sessions. Work remains queued when no compatible host is available, including -when a previously bound host no longer advertises that capability. It does not -silently discard the definitions or move an existing native Session. - -`TestNativePublicFunctionExecution` runs `tests/official_functions.py` using the -pinned official SDK against the actual HTTP handler, worker, PostgreSQL, daemon -and Codex. A synthetic model requests a configured function; the client reads -`required_actions`, submits ordered text/image results through public events, -retries the same result, receives completion, and reuses the Session. A subsequent -Turn verifies error output, and a third verifies cancellation while waiting. -The test checks native result receipts, retained function Items and no duplicate -native continuation. This proves the implemented workflow, not compatibility -with every tool variant or the upstream service's exact event timing. - -Accepted results currently enter public Items through native execution observations. -If cancellation prevents native application (for example, a result followed by -cancel in one admitted batch), the submission remains saved internally but has no -public result Item or `item.added`. Admission-time result indexing and unapplied -result recovery remain a separate compatibility gap; retries do not repair it. - -`TestNativePublicFunctionStreamHelper` runs `tests/official_function_stream.py` -with the same native fixture and pinned SDK. `sessions.stream(tool_handlers=...)` -submits a mapping returned by a handler and a generic failure when the handler -raises. It verifies one invocation per call, retained output/error field presence, -native application, and termination after the matching Turn completes and Session -returns idle. These are controlled tests with synthetic model responses. Live -execution acceptance additionally requires a real model API; provider connectivity -alone does not prove the Agents API/daemon/harness workflow. - -### Default verbosity on native models - -The pinned `AgentTextParam` defines `medium` as the default text amount. Omitted, -null and explicit `medium` keep the same effective Session configuration and retry -identity. Supported native models receive the explicit requested level. For -unsupported or unknown models, the Codex adapter removes a `medium` override and -uses native defaults while preserving the requested model and catalog snapshot. -It still rejects unsupported `low`/`high` and unreadable catalogs. - -This follows [Codex 0.153.4 request selection](https://github.com/openai/codex/blob/rust-v0.153.4/codex-rs/core/src/client.rs#L951) -and its [unknown-model fallback](https://github.com/openai/codex/blob/rust-v0.153.4/codex-rs/models-manager/src/model_info.rs#L134). -Controlled native verification checks explicit levels on supported models, absent -verbosity on an unknown model, initial/resumed Turns, and default retry equivalence. -This does not imply support for non-default verbosity on every model. - -Live MiniMax-M3 verification used the pinned SDK, actual service/worker/PostgreSQL, -daemon and Codex with MiniMax's real Responses API. Two Turns verified a successful -function result, handler failure, retained result fields, stream termination and -native history continuity by recalling a random value returned only by the first -tool invocation. Omitted, null and explicit medium reused the same creation -identity. The tool data was synthetic; model responses were live. This does not -establish non-default verbosity, tool-set enforcement or full protocol conformance. - -### Initial input at Session creation - -Session creation accepts the pinned string and user-message-array input -forms. It shares message validation and admission with the events endpoint, -including the [qualified image profile](message-input.md). The -Session and its initial work commit atomically; an identical -creation retry never re-admits the input, including after later or terminal Turns. -With `none`, this includes the first Turn and input Items. With `self_hosted`, it -includes the initial reservation and connection action; preparation and Turn -admission belong to the existing Worker. Creation returns while the executor is -offline, and an initial deadline failure leaves a failed Session without a Turn. -Initial input is required for `none`, and for streamed creation outside -`self_hosted`. Non-streaming hosted and self-hosted creation may omit input or -supply null. These conditions apply before creation retry lookup; valid retries -retain the same Session and never duplicate initial work. Existing Session reads -and subsequent events are unaffected. Execution must be enabled and the -configured engine must support admission before any initial work is persisted. - -Fixed SDK/raw HTTP and PostgreSQL tests cover the accepted forms, saved and inline -configuration, ordering, tenant isolation, retries, rollback and persistence. -Image support is bounded as documented above. Empty arrays, empty content and -empty text fail the shared message validator, as official probes also did. -Whitespace-only text is admitted and stored verbatim, as officially observed -([text content](message-input.md#text-content)). Full local size-limit and error-detail parity remains unverified. Swagger 2 cannot -express the string/array union, so input is unconstrained with a type description. - -### Session creation streaming - -`POST /v1/agents/sessions` also accepts `stream=true` for the supported creation -inputs. Fresh creation sends `agent.session.created` with the committed Session, -the same projection as the JSON 201 body (`in_progress` after `none` initial -input), then its committed activity/Turn/Item/output events. A new Turn publishes -`turn.created`, its user input `item.added`, `agent.session.in_progress`, then -`turn.in_progress`. Self-hosted initial creation shows and then emits the -Environment connection request, before native readiness and a Turn. The cursor -comes from the atomic creation upsert, so rapid initial execution cannot move the -start past its own events. The ordinary bounded-buffer/gap policy still applies. - -A fresh creation stream ends right after the first `agent.session.idle` recorded -when a Turn ends or an input reservation stops being pending (expired, cancelled -or failed), or any `agent.session.failed`, and never sends the events after it. A -pinned-SDK loop over `sessions.create(..., stream=True)` therefore ends right -after the initial Turn's idle. `requires_action`, function results, resumed work -and a self-hosted connection clearing pending input keep it open, and a -provisioning or offline reservation keeps it open until a Turn settles or the -reservation expires or fails. A creation that admitted nothing ends right after -`created`. A settlement that records no event ends the stream after the events -up to the cursor read with a settled projection in one snapshot; another client's -work drained before that read can still be sent. Observe later Turns with the GET -event stream, which does not end on settlement or a Turn failure; it ends only -after the terminal `agent.session.failed` of a hosted provisioning failure, as -officially observed ([initialization failure](environments.md#initialization-state-and-failure)), -or when the Session is deleted. Terminal Turn events carry the Turn -snapshot's `usage` at the top level, null when unknown. - -The local `Idempotency-Key` creation extension shares identity across response -modes. A same-key `stream=true` retry of an existing creation returns 201 with -only the connection comment and ends at once: it replays nothing, resubmits no -input and follows no work, since official same-key requests create distinct -Sessions. Recover a lost -Session ID by repeating the same request/key with `stream=false`, then use -Session/Turn/Items reads. Disconnect only stops the HTTP observer; committed -reservations and admitted execution continue. September 23 official `none` -observations match the created snapshot, the end at idle, the start order and the -terminal usage field ([evidence](history-events-usage.md#creation-stream-settlement-2026-09-23)). -Self-hosted, hosted and no-input creation stream lifetimes and the stream retry -behavior are local choices. -These choices are not full conformance. - -`official_session_creation_stream.py` covers the pinned client and raw HTTP on -real PostgreSQL: initial text and saved Agents, created snapshots equal to the JSON -201 body, Turn start order, terminal usage, the end at idle, GET continuation, -immediately ending stream retries and JSON retries, disconnect recovery, isolation -and errors before stream headers. Store tests cover concurrent upsert ownership, -pre-admission cursors, post-admission projections and observers draining after -execution has completed; API tests cover the stream lifetimes. - -## Acceptance evidence and remaining scope - -These accepted changes have distinct evidence levels. The associated PR records -include validation and limitations; later acceptance does not upgrade an earlier -controlled fixture into a real-provider test. - -| Area | Evidence | +# Agents API coverage ledger + +Core targets the complete OpenAI Agents API as pinned below ([public API rule](../../AGENTS.md#public-api)). This ledger records how much of each resource Core implements and which contract holds its details, then every known difference from the OpenAI service and every open gap. [API namespaces and credentials](../../docs/api/README.md) says who calls which API; the [Agents API guide](../../docs/api/public-agent-api.md) shows how to use it. + +## Pinned baseline + +| File | Contents | | --- | --- | -| Codex function stream/default text | [#544](https://github.com/MiniMax-AI-Dev/parsar/pull/544), [#545](https://github.com/MiniMax-AI-Dev/parsar/pull/545): fixed SDK/raw HTTP, actual PostgreSQL/daemon/native harness; #545 adds real MiniMax success/error and native history continuity | -| Independent build/container | [#552](https://github.com/MiniMax-AI-Dev/parsar/pull/552), [#563](https://github.com/MiniMax-AI-Dev/parsar/pull/563): isolated binaries/container, official Python/Go clients and real MiniMax execution across API restart | -| Session creation and source identity | [#564](https://github.com/MiniMax-AI-Dev/parsar/pull/564), [#567](https://github.com/MiniMax-AI-Dev/parsar/pull/567), [#572](https://github.com/MiniMax-AI-Dev/parsar/pull/572): atomic initial text, creation streaming and mutation-independent saved-reference retries | -| Saved resource lifecycle and Session filtering | [#573](https://github.com/MiniMax-AI-Dev/parsar/pull/573), [#574](https://github.com/MiniMax-AI-Dev/parsar/pull/574), [#581](https://github.com/MiniMax-AI-Dev/parsar/pull/581): official client/raw HTTP, PostgreSQL, tenant isolation and source mutation/deletion; #581 also filters completed real MiniMax Sessions | -| Claude SDK public execution | [#580](https://github.com/MiniMax-AI-Dev/parsar/pull/580): built API/registered daemon/packaged SDK with real MiniMax text, function success/error, active input, SSE/Items, pending-call cancellation and daemon cold continuation with retained native identity/history | - -Principal workflows above are accepted within their profiles. Missing resources, -broader configuration/content, complete Usage provenance, unapplied result -visibility, exact hosted errors/event timing and crash-window reconciliation remain -open. Claude SDK raw usage is retained internally; public usage stays null without -a complete token breakdown. Neither successful cold continuation nor database -writer fencing proves recovery of interrupted native side effects. Full protocol -compatibility, other harnesses/platforms and Parsar cutover are not established. - -### Caller principal foundation - -Caller keys resolve their database-owned Project and its shared service-account -principal on every request. Projects own distinct tenant scopes; keys within one -Project share the same scope and Session creator. Deployment configuration defines -no Projects or business keys. Administrator credentials cannot authenticate `/v1`. Optional official organization/project headers -must match the key's authorized scope; ambiguous or conflicting headers use the -existing 401 response. Every Agents API 401 has type `invalid_request_error`, as observed -officially; Beta routes report a null code, while Files and Skills report -`invalid_api_key` for a rejected Bearer credential and a null code without one. This scope-header policy is an implementation choice, not -verified hosted error parity. Project resource access remains shared -within the authorized project. New Sessions persist immutable creator kind/ID from -the authenticated principal; ordinary and streaming creation retries require the -same typed subject, including when recovering before saved-Agent lookup. Rotated -keys for that subject share retry identity. Unknown historical creators cannot be -claimed by retry. This local 409 policy is not verified hosted retry parity. -Creator fields remain internal and do not extend the public Session schema. -Executor keys now require the target Session's verified project and typed creator, -with optional exact-Environment restriction. Key issuance can precede Session -creation; rotation/revocation and current authorization reuse the durable ledger -and exact Runtime enrollment/gateway binding. Historical keys remain revoked and unclaimed. This -executor-specific prerequisite does not open public Environment admission or -establish complete ownership, hosted key lifecycle or error compatibility. See the -[standalone configuration](../../services/core/README.md#standalone-http-service). - -Core documents its optional [harness selection extension](model-execution.md#harness-selection) separately from the pinned upstream contract. - -The former Core startup configuration read is removed; the [installation read](installation.md) reports the public URL and installer process settings. - -Model endpoints and credentials may be supplied at Session creation through the -[write-only execution extension](model-execution.md). Provider catalogs and their -business permissions remain client/product responsibilities. - -API-key ownership and write history are documented in -[write-audit.md](write-audit.md). These are Core-key `/core/v1` reads; the pinned -public `/v1` schema is unchanged. +| [upstream.json](upstream.json) | The pin: [openai-python](https://github.com/openai/openai-python/tree/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents) 3.13.0 at commit `d7c41ef`, resources under `beta/agents`, Beta header `agents=v1` | +| [upstream-routes.json](upstream-routes.json), [upstream-fields.json](upstream-fields.json) | The 58 method and path pairs and their official fields: 42 operations under `beta/agents`, 5 Files and 11 Skills operations. `scripts/extract-agents-api-upstream.py` extracts them from the pinned SDK; run it with that SDK installed | +| [openapi.yaml](openapi.yaml) | Core's public schema, generated by `make openapi` from the route annotations in `services/core/internal/api/` and the wire types in [`v1/`](v1/) | + +Contract tests hold Core to the pin: the router and `openapi.yaml` serve exactly the pinned routes (`services/core/internal/api/routing_test.go`, `v1/upstream_contract_test.go`), every query parameter and field is official, and Core-only fields sit only inside `x_agents_core` on Agents and Sessions. Swagger 2.0 cannot express string-or-array unions, so `openapi.yaml` leaves Session `input` and function-result `output` unconstrained; the pinned types and Core's validation define them. Operations and fields newer than the pin wait for a protocol upgrade. + +Evidence for a status comes from the pinned official SDK and raw HTTP against the running service, as [CONTRIBUTING](../../CONTRIBUTING.md#compatibility-evidence) requires. + +## Coverage by resource + +**Implemented** means every operation serves the pinned shapes; limits that remain are listed under [known gaps](#known-gaps). **Partial** names what is missing. + +| Resource | Operations | Status | Contract | +| --- | --- | --- | --- | +| Agents | create, retrieve, update, list, delete | Implemented. Every pinned setting is saved; Session admission runs a subset | [Agents](wire-semantics.md#agents) | +| Sessions | create (JSON or stream), retrieve, update, list, delete | Implemented. Update takes `metadata` only; deletion requires an idle or failed Session | [Sessions](wire-semantics.md#sessions), [creation streaming](sessions-events.md#creation-streaming) | +| Session events | create, stream | Partial: messages with text and inline images, cancellation, function results; the stream is live only | [Sessions, events and history](sessions-events.md), [message content](message-content.md) | +| Turns | retrieve, list | Implemented; Session Turn routes hold root Turns only | [Turns and Items](sessions-events.md#turns-and-items) | +| Items | list | Partial: messages, commands, MCP calls, functions, web search, reasoning and Subagent coordination Items; other native variants are not projected | [Turns and Items](sessions-events.md#turns-and-items) | +| Artifacts | retrieve, list, delete, content | Implemented | [Environment files and Artifacts](environment-files.md) | +| Subagents | retrieve, list; Items; Turns retrieve and list; Turn Items | Partial: read-only child work; no live child progress or optional native operations | [Subagents](subagents.md) | +| Environments | retrieve | Implemented | [Environments](environments.md) | +| Environment files | create, list | Implemented; the list is not recursive | [Environment files and Artifacts](environment-files.md) | +| Environment Templates | create, retrieve, update, list, delete | Implemented; execution limits are listed under [known gaps](#known-gaps) | [Environment Templates](environments.md#templates) | +| Vaults | create, retrieve, list, delete | Implemented; no archive operation | [Vaults and Credentials](vaults.md) | +| Vault Credentials | create, retrieve, update, list, delete | Implemented for `static_bearer` and `mcp_oauth` | [Vaults and Credentials](vaults.md) | +| Files | create, retrieve, list, delete, content | Implemented for `purpose=user_data`; content download is rejected | [Files and Skills](source-files.md) | +| Skills and Skill versions | create, retrieve, update, list, delete, content | Implemented | [Files and Skills](source-files.md) | + +Which operation each Harness supports on each placement is in the [Harness capabilities](harness-capabilities.md). [Core wire behavior](wire-semantics.md) holds the rules that apply across resources: requests, errors and lists. + +Core's own fields sit inside `x_agents_core` ([Core extensions](../../docs/api/public-agent-api.md#core-extensions-x_agents_core)). The Core administration API (`/core/v1`) and the machine API (`/api/v1`) are not part of the Agents API. + +## Differences from OpenAI + +Each item is Core's deliberate or native behavior where the official service behaves otherwise. The linked rule states the exact behavior. + +**Requests and errors** ([Core wire behavior](wire-semantics.md)) + +- Core sends no `OpenAI-Organization` or `OpenAI-Project` response headers. +- `HEAD` on the event stream, on content downloads and on the Environment files list returns 405. +- A JSON array body is rejected; the official service reads `[]` as `{}`. +- U+0000 in a stored string returns 400; the official service stores it. +- Not-found messages never name the resource; Core quotes a full metadata key where the official message abbreviates it. +- UUID identifiers also resolve in other spellings, such as uppercase or braces. +- The Files routes keep a local `unsupported_parameter` code for a repeated query key. + +**Lists** ([lists](wire-semantics.md#lists)) + +- A deleted Agent or Session used as a cursor returns 404; the official service still pages from it. +- A Turn cursor from another Session, a Credential cursor equal to its Vault ID, and a Skills cursor that is not a Skill ID return 404. +- Vault and Credential lists clamp a negative `limit`, as the pinned SDK describes; the official service returns 400. + +**Agents and Sessions** ([Agents](wire-semantics.md#agents), [Sessions](wire-semantics.md#sessions)) + +- A repeated Session creation with the same `Idempotency-Key` returns the original Session; the official service creates a new one. +- Omitted programmatic tool calling keeps the harness's native behavior; the official default is on. +- An omitted reasoning effort stays null instead of taking the model's default. +- A Session's `agent.tools` omits `tool_search` declarations. +- Deleting a Session right after an events 202 returns 409, because Core admits the Turn in the same transaction; the official service returned 200. + +**Input, events and history** ([Sessions, events and history](sessions-events.md), [message content](message-content.md)) + +- Input to a `none` Session is admitted synchronously; Core does not emulate the official asynchronous admission window. +- A function result that resumes a waiting Turn emits `turn.in_progress`, and cancelling a Turn that waits on a function result emits an interim `agent.session.in_progress`. +- Attaching to the stream mid-Turn sends no catch-up Item snapshots. +- The Items list includes in-progress and incomplete output Items, and keeps a failed function result's submitted `output`. +- Error messages omit the call and executor IDs that official messages include. +- An empty text part beside other text is accepted and stored. +- Empty input returns 400 `invalid_request` with a generic message and a null param; the official response is `invalid_request_error` with param `input`. +- Session usage is available as soon as every root Turn has settled; official reads lag by seconds. + +**Files, Skills, Environment files and Artifacts** ([Files and Skills](source-files.md), [Environment files and Artifacts](environment-files.md)) + +- File uploads hold up to 512 MiB; the official limit is 512 MB. The Files list returns up to 10,000 Files by default, and a `purpose` filter other than `user_data` returns an empty page. +- Skill version numbers are never reused, and uploads and deletions of one Skill run one at a time. +- Environment files work on `self_hosted` Environments, which the official service refuses. +- Creating an Environment file over an existing regular file returns the "must not traverse symlinks or overwrite existing files" message. A parent symbolic link that stays inside the workspace is followed, and a parent that escapes the workspace or is a regular file gets a generic 400; the official service rejects symbolic-link parents. +- Artifact IDs are UUIDs. + +**Vaults and Credentials** ([Vaults and Credentials](vaults.md)) + +- Vault and Credential status is stored privately and defaults to `active`; with no archive operation, lists without a filter include both statuses. +- An unknown or foreign Vault in `vault_ids` returns 404 "Resource not found."; the official message names the ID. +- An explicit `null` for OAuth `access_token`, `refresh` or `token_endpoint_auth` in an update keeps the stored value. +- A static token must be an RFC 6750 `b64token` to run; other stored tokens fail at dispatch. +- Vault metadata has a 64 KiB bound and no pair or length limits; names are 1–256 bytes after trimming. + +## Known gaps + +**Configuration and tools** + +- Explicit reasoning effort or summary, service tiers other than `auto`, enabled `web_search` and enabled programmatic tool calling are saved but rejected at Session admission. +- Harness support for tools, structured output, deferred discovery, subagents and MCP differs by Harness and placement; see the [Harness capabilities](harness-capabilities.md). MiniMax Code has no public functions, no service-origin MCP and no image input. +- Model-derived reasoning defaults are not resolved. + +**Execution and history** + +- The stream does not emit reasoning-summary events, Environment `pending` or `ready` events, or every pinned interim tool-output variant. +- Native Item variants beyond those listed under [Turns and Items](sessions-events.md#turns-and-items) are not projected, and Items cannot be modified. +- A function result that cancellation prevents from being applied never appears as an Item. +- Pinned Codex can lose command output emitted before its stream subscription. +- Claude Code and MiniMax Code report no public usage. +- Core gives no crash-safe or exactly-once guarantee for native side effects; claimed work fails on restart without replay. +- Images must be inline PNG or JPEG data URIs; remote URLs, `file_id` and `detail` are rejected. + +**Environments and Templates** + +- Runtimes do not enforce `disabled` or `restricted` networks, so Sessions that need them are rejected ([restricted network policy](environments.md#restricted-network)). +- `packages.system` is rejected; system packages must be preinstalled. + +**Files and Environment files** + +- Files accept only `purpose=user_data`; other purposes, `expires_after` and the Uploads API are not supported. +- An Environment file write whose outcome is uncertain is never retried or recovered automatically; it blocks further writes and messages to the Session. + +**Vaults and Credentials** + +- There is no archive lifecycle, storage-key rotation or re-encryption. +- OAuth refresh happens only at dispatch: there is no refresh on a provider 401, no mid-Turn replacement and no withdrawal of a token already sent to a Runtime. +- Input sent to a Session whose selected Credential was deleted is admitted and then fails at dispatch. +- Credential creation takes no `Idempotency-Key`. + +**Sessions** + +- Session deletion does not purge the stored history physically. +- Stream lifetimes for `self_hosted`, hosted and no-input creation, and the creation-stream retry, are Core's own choices. + +### Unverified against the official service + +- The order of errors when one request has several faults, and error, default and payload-limit parity in general. +- Whether Subagent child Turns and pending Environment file writes block Session deletion. +- The official order between the pending-input error and an unknown result target. +- Failure reasons for npm, initial-file and Skill installation have no official sample. +- Codex behavior for an empty text part beside text, and for images in failed function results or as remote references; Claude behavior for a whitespace-only or empty text block inside a mixed message. +- Official `HEAD` behavior on the routes where Core returns 405. +- Environment Template hostname forms beyond exact hosts, and `disabled` combined with domains. +- Files purpose filtering and pagination while Files change; official Skill upload limits and error timing. +- Environment file list defaults (limit 20, the workspace root as the default path, no recursion, page-token invalidation), the 50 MiB `file_id` copy limit and the check order on create. +- Artifact capture of hard links, special files and a linked `outputs` directory, republication after changed bytes, and content headers and ranges. +- Vault and Credential error and retry semantics, pagination under concurrent writes, visibility after deletion, exact-URL matching against the official normalization, OAuth refresh timing and errors, and restricted-key scopes. diff --git a/contracts/agents-api/admin-api.md b/contracts/agents-api/admin-api.md index 87b6b6b5c..573db0ff7 100644 --- a/contracts/agents-api/admin-api.md +++ b/contracts/agents-api/admin-api.md @@ -1,213 +1,235 @@ -# Administrator API +# Core administration API -These routes live under `/core/v1` and are called by Core Web's server and -operator scripts. Every `/core/v1` route requires the Core key as a Bearer -credential; an unknown `/core/v1` path returns 404 only after authentication. -They do not change `/v1`, the fixed Python SDK, or native Runtime interfaces. -Project API keys and machine credentials cannot authenticate these routes; the -Core key cannot authenticate `/v1` or `/api/v1`. Paths below are relative to -`/core/v1`. `X-Core-Console-Actor` is a caller-asserted, display-only label that -Core records without verifying. Web sends `console`; direct Core key scripts -normally send none but could set any label. Never use it for authorization or as -proof of origin. +The Core administration API (`/core/v1`) manages an installation: Projects and their API keys, reads and deletion of Project resources, executor credentials, deployment default models, the sandbox deployment and its nodes, monitoring and audit. Web's console server calls it for the signed-in administrator ([console server](../../docs/web/console-server.md#forwarding-to-core)); operators call it from scripts on the Core host ([script the Core API](../../docs/getting-started/operations.md#script-the-core-api)). The generated schema is [core.openapi.yaml](core.openapi.yaml), and every error uses the [Core error envelope](core-errors.md). -## Error envelope +Applications never call `/core/v1`. It has no operation that creates or edits Agents, Sessions, templates, Files, Skills or Vaults, starts or cancels work, reads Source File content or streams events; applications do those through the [Agents API](../../docs/api/public-agent-api.md). -See [Core administration errors](core-errors.md) for the optional flat `details` -object and the distinct console proxy failure codes. Public and machine response -shapes remain unchanged. +## Authentication + +Every `/core/v1` request, including one for an unknown path, must send `Authorization: Bearer `. Without it Core answers 401 `invalid_admin_key` with `WWW-Authenticate: Bearer`; an unknown path answers 404 `not_found` only after authentication. + +- Core compares the SHA-256 of the bearer with the Core key digest it reads at startup from `generated/core-key-digests.json`, which `oac apply` derives from `secrets/core.key` ([Core key](../../docs/getting-started/operations.md#core-key)). A rotated Core key takes effect when Core restarts. Core serves no `/core/v1` route when no digest is configured. +- The Core key authenticates only `/core/v1`. Project API keys and machine credentials get 401 here, and the Core key gets 401 on `/v1` and `/api/v1` ([API namespaces and credentials](../../docs/api/README.md)). +- `X-Core-Console-Actor` is a label the caller asserts. Core records it as the audit `actor_label` without checking it. The console server sends `console`; scripts normally send none, which records an empty label. Never use it for authorization or as proof of origin. +- A `{project_id}` in a path selects the target Project, including an archived one; it grants nothing. + +## Routes + +Paths are relative to `/core/v1`. + +| Routes | Purpose | Contract | +| --- | --- | --- | +| `installation` | Public URL, API base URL, source commit, the installer's process settings and what is bound to the public URL | [Installation facts](#installation-facts) | +| `projects`, `projects/{project_id}`, `projects/{project_id}/archive`, `projects/{project_id}/keys[/{key_id}]` | Projects and their API keys | [Projects and keys](#projects-and-keys) | +| `projects/{project_id}/{agents,environment-templates,skills,files,vaults,sessions}/**` | Resource reads and deletion, Session history and Artifacts | [Resource reads and deletion](#resource-reads-and-deletion) | +| `projects/{project_id}/sessions/{session_id}/archive` | Archive one hosted Session | [Session archive](#session-archive) | +| `projects/{project_id}/sessions/{session_id}/execution-configuration` | The Session's frozen model, Harness and provider selection | [Execution configuration](#execution-configuration) | +| `projects/{project_id}/sessions/{session_id}/diagnostics`, `…/turns/{turn_id}/diagnostics` | Failure categories and Item receipt timing | [Session diagnostics](session-diagnostics.md) | +| `projects/{project_id}/sessions/{session_id}/runtime-observation`, `sandbox/runtime-observations` | Current Runtime observations | [Runtime observations](runtime-observability-api.md), [the list's disk field](#runtime-observations) | +| `projects/{project_id}/sessions/{session_id}/runtime-history` | Stored Runtime history | [Runtime history](runtime-history-api.md) | +| `projects/{project_id}/environments/{environment_id}/installation` | Install commands of a `self_hosted` Environment | [Installation grant](environment-executor-credentials.md#installation-grant) | +| `projects/{project_id}/environments/{environment_id}/executor-credentials[/{key_id}]` | Executor credentials of a `self_hosted` Environment | [Executor credentials](environment-executor-credentials.md#core-key-routes) | +| `projects/{project_id}/resource-owners`, `projects/{project_id}/write-operations` | Which API key created a resource and each key's writes | [Write provenance](#write-provenance) | +| `harnesses`, `harnesses/{harness}/model-configuration` | Enabled Harnesses and each Harness's deployment default model | [Deployment defaults](model-execution.md#deployment-defaults) | +| `sandbox/deployment`, `sandbox/deployment/reset`, `sandbox/providers/{provider}/discovery` | The sandbox provider, resources and Runtime, reset, and provider configuration discovery such as E2B templates | [Sandbox deployment](sandbox-deployment.md#authority-and-routes) | +| `sandbox/enrollment-tokens`, `sandbox/nodes[/{node_id}[/allocations]]` | Node enrollment tokens, nodes and their allocations and host history | [Nodes guide](../../docs/getting-started/nodes.md), [sandbox deployment](sandbox-deployment.md), [node host history](node-host-history.md) | +| `summary` | Session counts and usage by Project, Agent or key | [Summary](#summary) | +| `metrics` | Core's own process, execution, database and job metrics | [Core metrics](core-metrics.md) | +| `audit-log` | Administrator writes | [Audit log](#audit-log) | ## Projects and keys -A Project owns one tenant and shared principal. Its keys have equal access to all -its assets. Projects and keys are database-owned; deployment configuration defines -neither. There are no API users, roles or configuration-managed business keys. -Core requires the Core key digest file (`OAC_CORE_KEY_DIGESTS_FILE`) at -startup for bootstrap and management. +A Project owns one execution tenant; its keys share its principal and assets ([Projects own assets](../../docs/design-principles.md#projects-own-assets)). Web's **Projects and keys** page uses these routes. -| Operation | Path | Result | +| Operation | Route | Result | | --- | --- | --- | | List Projects | `GET /projects` | `{data, has_more}` | -| Create Project | `POST /projects` with `{name}` | Project metadata; HTTP 201 | -| Rename Project | `POST /projects/{project_id}` with `{name}` | Project metadata | -| Archive Project | `POST /projects/{project_id}/archive` | Project metadata | +| Create a Project | `POST /projects` with `{name}` | 201 and the Project | +| Rename a Project | `POST /projects/{project_id}` with `{name}` | The Project | +| Archive a Project | `POST /projects/{project_id}/archive` | The Project | | List keys | `GET /projects/{project_id}/keys` | `{data, has_more}` | -| Issue key | `POST /projects/{project_id}/keys` with `{name}` | Key metadata and one-time `key`; HTTP 201 | -| Revoke key | `DELETE /projects/{project_id}/keys/{key_id}` | `{id, deleted: true}` | -| List executor credentials | `GET /projects/{project_id}/environments/{environment_id}/executor-credentials` | `{data}` metadata only | -| Issue or rotate executor credential | `POST /projects/{project_id}/environments/{environment_id}/executor-credentials` with `{key_id, rotate}` | One-time credential; HTTP 201 | -| Revoke executor credential | `DELETE /projects/{project_id}/environments/{environment_id}/executor-credentials/{key_id}` | HTTP 204 | - -Executor credentials apply only to a `self_hosted` Environment of the Project -whose Session exists. An archived Project returns 409 `project_archived` for -issuance and rotation but still lists and revokes; see -[executor credentials](environment-executor-credentials.md). -Project and key IDs are server-generated UUIDs. Project metadata contains `id`, `name`, -`created_at`, nullable `archived_at`, and `active_key_count`. Key metadata contains -`id`, `project_id`, `name`, `prefix`, `created_at`, and nullable `revoked_at`. -Project names contain 1–128 characters; key names contain 1–80. Names are display -labels and may repeat; control characters are rejected. Lists use lexical ID -ordering, `order=asc|desc` (default `desc`), `limit=1..100` (default 20), and `after`. -No response includes the stored digest or an existing credential's plaintext. -The catalog UUID identifies management paths. Its public authentication scope uses -organization `core` and project `proj_`; optional OpenAI scope headers -must match those values. All keys share subject `service_account/project:`. - -Rotate by issuing a new key in the same Project and revoking the old one. Revoking -one key leaves other keys, assets and admitted work intact. Archive atomically -marks the Project archived, revokes all its keys and records audit. Archived -Projects cannot issue keys; their assets remain available for administrator -inspection and deletion. There is no Project deletion, unarchive, key reset or -automatic write retry operation. -After an uncertain issuance response, inspect metadata and explicitly revoke any -unusable key before issuing another; plaintext cannot be recovered. +| Issue a key | `POST /projects/{project_id}/keys` with `{name}` | 201, the key metadata and the plaintext `key` | +| Revoke a key | `DELETE /projects/{project_id}/keys/{key_id}` | `{id, deleted: true}` | + +- A Project has `id`, `name`, `created_at`, nullable `archived_at` and `active_key_count`. A key has `id`, `project_id`, `name`, `prefix`, `created_at` and nullable `revoked_at`. IDs are server-generated UUIDs. +- Project names have 1–128 characters and key names 1–80; names are labels and may repeat, and control characters are rejected. +- Lists order by ID with `order=asc|desc` (default `desc`), `limit=1..100` (default 20) and `after`. +- Only the issuance response contains the key's plaintext, in `key`; Core stores its digest. Show it once and never cache it. After an uncertain issuance response, list the keys and revoke any you cannot use before issuing another. +- Archive marks the Project archived, revokes all its keys and writes the audit entry in one transaction. Issuing a key in an archived Project returns 409 `project_archived`. There is no Project deletion, unarchive or key reset. + +The [`/v1` authentication rules](wire-semantics.md#authentication) define key lookup, revocation visibility, scope headers and authentication errors. ## Resource reads and deletion -Paths below are relative to `/projects/{project_id}`. The Project selects a tenant, including -an archived Project; it does not authenticate. Shared resource handlers preserve their -public object serialization, pagination, errors and deletion preconditions. They -receive an explicit target tenant, not a fabricated caller identity. +Paths are relative to `/core/v1/projects/{project_id}`. Each read returns the same object, pagination and errors as the matching `/v1` operation, and each deletion has the same preconditions. -| Resource | GET routes | DELETE routes | +| Resource | Reads | Deletion | | --- | --- | --- | | Agents | `/agents`, `/agents/{agent_id}` | `/agents/{agent_id}` | -| Templates | `/environment-templates`, `/environment-templates/{environment_template_id}` | Item route | -| Skills | `/skills`, `/skills/{skill_id}`, item `/content`, item `/versions`, `/versions/{version}`, version `/content` | Skill and version item routes | -| Files | `/files`, `/files/{file_id}` | Item route | -| Vaults | `/vaults`, `/vaults/{vault_id}`, item `/credentials`, `/credentials/{credential_id}` | Vault and Credential item routes | -| Sessions | `/sessions`, `/sessions/{session_id}`, item `/turns`, `/turns/{turn_id}`, `/items`, `/artifacts`, `/artifacts/{artifact_id}`, Artifact `/content`, `/execution-configuration`, `/runtime-observation`, `/runtime-history` | Session and Artifact item routes | -| Provenance | `/resource-owners`, `/write-operations` | None | - -Source File content, Session events/SSE, arbitrary creation/update and execution -operations are deliberately absent. Session deletion still requires idle state; -management deletion never cancels implicitly. Deleting a Credential does not revoke -its provider authorization. Deleting a default Skill version retains the public -constraint. Skill and Artifact downloads reject HEAD like the corresponding project -operations; the Runtime observation (single and list) and Runtime history reads also -reject HEAD with 405, so HEAD never samples a provider or queries telemetry. - -## Administrative Session archive - -`POST /projects/{project_id}/sessions/{session_id}/archive` takes -`{"expected_generation": N}`, where N is a positive uint64 from the current -sandbox deployment. It requires the Core key, a -Web-managed configured deployment and the current generation. A stale generation -returns 409 `sandbox_generation_stale`; reset is not a precondition. Only Core-managed -`openai_hosted` Sessions are eligible; other environment types return 400. A -Session outside the selected Project returns the same 404 as a missing Session. - -One transaction marks the Environment expired (retaining an existing failed -state), requests cancellation, revokes Runtime authority and records the -administrator audit. The existing lifecycle owns compute and snapshot cleanup; -no provider operation runs inside that transaction. Unknown outcomes retain -ownership until matching provider receipts confirm release. The Session itself -is not deleted. Public history and persisted Files/Artifacts remain available; -unpersisted workspace contents are lost and the original Session cannot resume. - -Both POST and `GET /projects/{project_id}/sessions/{session_id}/archive` return -`{session_id, environment_id, state}`. GET is read-only and does not require -an active reset or an expected generation. `state` describes current resource -disposition: `active`, `cleanup_pending`, or `released`. It is not archive -provenance: resources may already have expired through their normal lifecycle. -`released` does not prove that an active Turn has finished cancellation or -terminal publication; inspect that Turn separately when needed. - -After an uncertain POST response, GET this resource before choosing another -write. A repeated POST has an idempotent state effect, with a separate audit -record for each accepted request. Clients never automatically retry the mutation. -`AdminClient.archiveSession` sends the generation and -`AdminClient.retrieveSessionArchive` reads the disposition; the TypeScript client -accepts only positive safe integer generations to avoid rounding JSON numbers. - -Archive each retained hosted Session explicitly, then verify deployment allocation -and pending counts are zero before changing provider, resources or Runtime. -Snapshots and uncertain cleanup remain blockers. This is not a deployment-wide -bulk operation, Session migration or public Session deletion. - -## Historical copy provenance - -Cross-Project asset copying has been removed; no route creates copies. Resources -copied before the removal keep `api_key:null`, `source:"admin_copy"` and their -`admin_audit_id` in resource ownership, and their `copy` audit entries keep their -`result_ids` mappings. Historical unknown resources have null source and audit ID. - -## Monitoring and audit - -`GET /summary` supports optional `project_id`, `group_by=project|agent|key` -(default `project`), inclusive `created_after` and exclusive `created_before` -RFC3339 Session-creation bounds. Agent grouping requires `project_id`. `after`, -`limit`, `order` paginate Projects. Response `{data, has_more, next_cursor}` rows -contain `project_id`, nullable `agent_id` and `key_id`, current `assets` counts -(null for Agent/key groups), `sessions` counts (`total`, `idle`, `in_progress`, -`requires_action`, `failed`), cumulative `usage`, `coverage` (`measured_sessions`, -`total_sessions`, nullable `ratio`), and nullable Unix `last_active_at`. -Key groups attribute the entire Session to its recorded creation key, even if a -different key later sends input. Missing provenance becomes a null-key group. -Null public Session usage contributes no tokens but counts in the coverage -denominator. Each Project is read from one database snapshot; the page is not a -simultaneous deployment-wide snapshot. Totals are not billing records. - -`GET /sandbox/runtime-observations` uses existing Session creation-order pagination and -returns `{object:"list", data:[{project_id, observation}], has_more, first_id, last_id}`. -It reuses the bounded read-only Runtime sampler and never provisions compute; a -provider with a batch metrics read (E2B) samples the page in one bounded request. -Each `observation` is the Runtime observation plus `disk: -{usage_bytes, limit_bytes}` with memory's null rules: E2B fills it from its -reported disk usage and capacity, Docker returns null, and microsandbox returns -null until its disk semantics are designed. The per-Session administrator -observation read keeps the shape without `disk`. The per-Session execution -configuration, Runtime observation and Runtime history exist only here; `/v1` has -no equivalents. - -`GET /metrics?range=1h|6h|24h|7d` returns Core's own process metrics; see -[Core metrics](core-metrics.md). - -`GET /audit-log` lists administrator writes newest first with `project_id`, -`resource_type`, `resource_id`, `action`, inclusive `created_after`, exclusive -`created_before`, `limit=1..100` (default 50), and opaque `after` filters. Response -is `{data, has_more, next_cursor}`. Each row has `id`, `created_at`, -`admin_credential_id` (credential digest prefix), `actor_label` (the caller-asserted -display label: normally `console` from Web and empty from direct Core key requests), `action`, -`project_id`, `resource_type`, `resource_id`, `result_ids`, `request_id`, -`trace_id`. `result_ids` is an empty array except on historical `copy` entries. -Executor credential writes appear with `resource_type:"executor_credential"`, -the key ID as `resource_id` and action `issue`, `rotate` or `revoke`. Deployment -default model provider writes are deployment-wide: `project_id` is null, -`resource_type:"deployment_model_provider"`, the harness as `resource_id` and -action `set` or `delete`; a `project_id` filter excludes them. -No credential values or request bodies are recorded. Logs and historical copy -ownership do not cascade away on resource removal or key revocation. - -Administrator writes and their audit record share one PostgreSQL transaction. -Reuse existing resource deletion and serialization code. Cross-Project copying was -removed; only the historical `source:"admin_copy"` provenance read remains, fed by -`admin_resource_owners`. Unknown historical provenance remains unknown. No secrets -or request bodies enter logs. The console's fixed actor label (`console`) is only -an audit display label, never an authorization input. - -## Private installation transition - -The old configured business keys, inherited-binding issuer and key-space -management routes are removed. The Core key remains separately configured. The -migration refuses an installation containing old issued key records instead of -silently changing their ownership or deleting data. Use a clean private deployment, -or explicitly retire old key records after preserving the assets and evidence you -need. There is no automatic data migration or historical ownership backfill. - -The typed `AdminClient`, `SandboxAdminClient` and `CoreMetricsClient` in -`packages/agents-client` use `/core/v1`. -Public SDK applications continue to use the existing Agents API client and their -own API key. See [design principles](../../docs/design-principles.md). - -Sandbox deployment reset records `reset_start`, explicit `reset_force`, automatic -`reset_deadline`, `reset_cancel` and `reset_complete` in administrator audit. -Background archives retain the reset requester's credential/actor/request/trace -provenance and recover each Session's real Project scope. Audit failure rolls back -the corresponding state transition. Cancel does not undo an archive already committed. - -Online E2B deployment updates audit `change`, or `replace_credential` when an API -key is explicitly supplied (including the existing key), under resource type -`sandbox_deployment` and the installation ID. The audit and generation/credential -write commit together; verification failures and omitted-key no-ops produce no -mutation audit. No credential, request body or provider response text is recorded. +| Environment Templates | `/environment-templates`, `/environment-templates/{environment_template_id}` | The item route | +| Skills | `/skills`, `/skills/{skill_id}`, `/skills/{skill_id}/content`, `/skills/{skill_id}/versions`, `/skills/{skill_id}/versions/{version}` and its `/content` | Skill and version item routes | +| Files | `/files`, `/files/{file_id}` | The item route | +| Vaults | `/vaults`, `/vaults/{vault_id}`, `/vaults/{vault_id}/credentials`, `/vaults/{vault_id}/credentials/{credential_id}` | Vault and Credential item routes | +| Sessions | `/sessions`, `/sessions/{session_id}`, and under it `/turns`, `/turns/{turn_id}`, `/items`, `/artifacts`, `/artifacts/{artifact_id}` and its `/content` | Session and Artifact item routes | + +- Administrator deletion never cancels work: a Session that `/v1` could not delete, because a Turn or input is pending, returns the same 409. +- Deleting a Credential removes Core's copy only; it does not revoke the authorization at the provider. +- `HEAD` on Skill and Artifact content, a Runtime observation, the Runtime observation list and Runtime history returns 405, so it never samples a provider or queries telemetry. +- Administrator deletions appear in the [audit log](#audit-log), not in key write history. + +## Session archive + +`POST /projects/{project_id}/sessions/{session_id}/archive` with `{"expected_generation": N}` releases one Core-managed `openai_hosted` Session's sandbox without a deployment reset. N is the current generation of the [sandbox deployment](sandbox-deployment.md), a positive integer. The deployment must be configured in Web. + +| Case | Result | +| --- | --- | +| Stale generation | 409 `sandbox_generation_stale` | +| Environment type other than `openai_hosted` | 400 | +| Session missing or in another Project | 404 | + +One transaction marks the Environment expired (a failed Environment stays failed), requests cancellation of running work, revokes the Runtime's authority and writes the audit entry. Cleanup of the sandbox and its snapshot follows through the normal lifecycle; a resource whose release is uncertain stays owned until the provider confirms it. The Session is not deleted: its history and persisted Files and Artifacts stay readable, unpersisted workspace contents are lost and the Session cannot resume. + +`POST` and `GET /projects/{project_id}/sessions/{session_id}/archive` return `{session_id, environment_id, state}`. `GET` is read-only and needs no generation. `state` is the resource's current disposition: `active`, `cleanup_pending` or `released`, whatever released it. `released` does not mean an active Turn has finished cancelling; read the Turn for that. + +After an uncertain `POST` response, `GET` the archive before writing again. Repeating the `POST` has the same effect and records one audit entry per accepted request. To clear every hosted Session before changing the deployment, use the [deployment reset](sandbox-deployment.md#initialization-same-provider-changes-and-reset). + +## Execution configuration + +`GET /projects/{project_id}/sessions/{session_id}/execution-configuration` reports the model, Harness, native parameters and model provider a Session froze at creation. It reads only stored configuration: it never contacts a provider, starts a Turn or wakes a sandbox. Responses carry `Cache-Control: no-store`. + +```json +{ + "object": "agent.session.execution_configuration", + "schema_version": 1, + "session_id": "013773a9-44b9-4f84-baca-b51c04a01201", + "model": {"value": "requested-model", "source": "session"}, + "harness": {"value": "codex", "source": "agent"}, + "harness_config": {"value": {"model_reasoning_effort": "high"}, "source": "agent"}, + "model_provider": { + "source": "agent", + "status": "available", + "configuration": {"protocol": "responses", "base_url": "https://model.example/v1", "api_key_configured": true} + } +} +``` + +Each `source` is `session`, `agent`, `deployment` or `unknown`, recorded independently: a Session can override the model and keep its Agent's Harness and provider. An explicit inline Harness is `session`; an inline `agent.x_agents_core: null` resets the Harness to `deployment` and keeps the inherited provider; a null Session provider inherits normally. `harness_config.value` is `{}` when no native parameters apply. [Model execution](model-execution.md) owns how each value is resolved. + +| `model_provider.status` | Meaning | `configuration` | +| --- | --- | --- | +| `available` | Core recorded a safe view of the frozen provider, including a deployment default (source `deployment`) | `protocol`, `base_url`, `api_key_configured` and, when set, `context_window` and `max_output_tokens` | +| `redacted` | A deployment selection recorded without a safe view | null | +| `unavailable` | Core has no trustworthy record of the provider (source `unknown`). Execution may still have succeeded | null | + +Core writes this record in the same transaction that creates the Session. Later Agent edits or deletion, deployment default changes, restarts and same-key creation retries never change it. A Session without the record reports its stored model and Harness with source `unknown`, a null value where none is stored, and an `unavailable` provider. A missing Session and one in another Project return the same 404. The response never contains keys, ciphertext, secret references, native headers or query parameters. + +## Installation facts + +`GET /installation` reports what an administrator needs to call and change this installation. It answers before any sandbox deployment exists and calls no provider or model. + +| Field | Meaning | +| --- | --- | +| `object` | `core.installation` | +| `installation_id` | The installation ID from `state.json` ([installation directory](../../docs/configuration.md#installation-directory)); null when Core runs without the sandbox manager | +| `public_url` | The [`public_url`](../../docs/configuration.md#settings) setting: the origin applications, nodes, sandboxes and self-hosted executors use. Null when unset | +| `api_base_url` | `public_url` followed by `/v1`, the `OPENAI_BASE_URL` for Project API keys. Null when `public_url` is null | +| `local_only` | True when `public_url` names a loopback host, which only the Core host reaches | +| `source_commit` | The full source commit Core was built from; null for development builds | +| `configuration` | The installer's snapshot of `config.json`; null when the installer did not start Core | +| `address_bindings` | What a change of `public_url` affects, counted on each read | + +`configuration` has: + +- `path`: the absolute host path of `config.json`, by default `~/.oac/core/config.json`; +- `apply_command`: the command that applies changes, by default `~/.oac/core/oac apply`; +- `applied_at`: when the snapshot was last applied; +- `settings`: one entry per setting, with its dotted `key`, applied `value`, `default`, whether it is `changeable` after installation, whether it is `sensitive`, and the services it `restarts` (`core`, `web`, `database`). + +A sensitive setting has null `value` and `default` and a boolean `configured` instead; only sensitive settings have `configured`. Core refuses to start when the snapshot breaks this rule, repeats a key or has an unknown member. Core only reports the snapshot; [configuration](../../docs/configuration.md) describes each setting. + +| `address_bindings` field | Meaning | +| --- | --- | +| `nodes` | Enrolled nodes that are not removed | +| `nodes_on_other_address` | Nodes enrolled with an address other than `public_url`. They receive no new sandboxes; remove and add them again. At most `nodes` | +| `hosted_sandboxes` | Retained and pending hosted sandboxes, which run with the address current when they started | +| `self_hosted_executors` | Unrevoked executor credentials, whose executors were installed with the `remote_url` then advertised | + +## Write provenance + +Core records which Project API key made each successful public write, in the same transaction as the write; if the record fails, the write fails. Reads, rejected requests and Core's own maintenance, such as OAuth token refresh and cleanup, are not recorded. + +- A write is recorded when it commits. Session input counts once admitted, including input reserved for an Environment that is still preparing; a later execution failure or a lost response does not remove the record. An explicit empty input batch and a repeated creation or deletion are recorded without changing ownership. +- An Environment file upload is recorded when the Runtime confirms the write. Core saves the key, request and trace before sending the file, and records only confirmed uploads. +- Creating a resource also records its creation owner. Updates, retries and no-op writes never change it. A new Skill's first version and a new Session's Environment share the creating operation. Artifacts come from the Runtime and have no creation owner; deleting one is recorded. +- Revoking a key stops new writes but keeps its history. Deleting a resource keeps its creation owner and operation history. +- Each record holds its ID, time, key metadata, action, resource type and ID, parent ID, `request_id` and `trace_id`. `request_id` is server-generated per request; a shared `trace_id` is not an idempotency key. Records never hold request or response bodies, secrets, model credentials, tokens, file paths or file contents. + +| Public write | `action` | `resource_type` (parent) | +| --- | --- | --- | +| Agent create, update, delete | `create`, `update`, `delete` | `agent` | +| Session create, update, delete | `create`, `update`, `delete` | `session` | +| Session events | `send_events` | `session` | +| Artifact delete | `delete` | `artifact` (`session`) | +| Environment file upload | `upload_file` | `environment` (`session`) | +| Environment Template create, update, delete | `create`, `update`, `delete` | `environment_template` | +| Skill create, default version change, delete | `create`, `update_default_version`, `delete` | `skill` | +| Skill version upload, delete | `upload_version`, `delete` | `skill_version` (`skill`) | +| Source File upload, delete | `create`, `delete` | `file` | +| Vault create, delete | `create`, `delete` | `vault` | +| Credential create, replace, delete (static and OAuth) | `create`, `update`, `delete` | `credential` (`vault`) | + +Both routes accept only the parameters listed; an unknown, repeated or empty parameter returns 400, and a missing Project 404. + +`GET /projects/{project_id}/resource-owners?resource_type=agent&resource_ids=id1,id2` returns the creating key of 1–100 resources of one type, in request order. `resource_type` is `agent`, `session`, `environment`, `environment_template`, `skill`, `skill_version`, `file`, `vault`, `credential` or `artifact`. + +```json +{"data":[ + {"resource_id":"id1","api_key":{"id":"key-uuid","name":"SDK","prefix":"pc_example","kind":"issued","revoked_at":null},"source":"api_key","admin_audit_id":null}, + {"resource_id":"id2","api_key":null,"source":null,"admin_audit_id":null} +]} +``` + +`api_key` and `source` are null when Core has no creation record, including for resources in another Project. `source: "admin_copy"` with an `admin_audit_id` marks a resource recorded by a `copy` entry in the audit log; no current route writes one. + +`GET /projects/{project_id}/write-operations` lists writes newest first by `(created_at, id)`. Filters: `key_id`, `resource_type`, `resource_id`, inclusive `created_after` and exclusive `created_before` (RFC 3339). `limit` is 1–100, default 50. Pass the previous `next_cursor` as `after` with unchanged filters. The response is `{data, has_more, next_cursor}`; each entry has `id`, `created_at`, `api_key`, `action`, `resource_type`, `resource_id`, `parent_id` (empty when absent), `request_id` and `trace_id`. + +Creation records are kept for good, including after the resource is deleted. Other records are kept for [`core.write_audit_retention`](../../docs/configuration.md#settings), 90 days by default; every minute Core deletes up to 1,000 expired records, so a backlog drains over several passes. Revoking a key or deleting a resource never deletes records. + +## Summary + +`GET /summary` counts Sessions and usage. + +| Parameter | Meaning | +| --- | --- | +| `project_id` | One Project; required for `group_by=agent` | +| `group_by` | `project` (default), `agent` or `key` | +| `created_after`, `created_before` | Inclusive and exclusive RFC 3339 bounds on Session creation | +| `after`, `limit`, `order` | Paginate Projects | + +The response is `{data, has_more, next_cursor}`. Each row has `project_id`, nullable `agent_id` and `key_id`, current `assets` counts (null for Agent and key groups), `sessions` counts (`total`, `idle`, `in_progress`, `requires_action`, `failed`), summed `usage`, `coverage` (`measured_sessions`, `total_sessions`, nullable `ratio`) and nullable Unix `last_active_at`. + +- A key group counts each Session under the key that created it, even when another key later sends input. Sessions without a recorded creator form a group with a null `key_id`. +- A Session whose public usage is null adds no tokens but counts in the coverage denominator. +- Each Project is read in one database snapshot; a page is not one snapshot of the whole deployment. Totals are operational counts, not billing records. + +## Runtime observations + +`GET /sandbox/runtime-observations` lists the current Runtime observation of every Session in every Project, as `{object: "list", data: [{project_id, observation}], has_more, first_id, last_id}` ([Runtime observations](runtime-observability-api.md)). It samples read-only and never provisions compute; a provider with a batch metrics read, such as E2B, samples the page in one bounded request. Each list `observation` adds `disk: {usage_bytes, limit_bytes}`, with the null rules of `memory`: E2B reports its sandbox disk usage and capacity, and Docker and microsandbox return null. The per-Session read has no `disk`. + +## Audit log + +`GET /audit-log` lists administrator writes newest first. Filters: `project_id`, `resource_type`, `resource_id`, `action`, inclusive `created_after` and exclusive `created_before` (RFC 3339). `limit` is 1–100, default 50, with the opaque `after` cursor. The response is `{data, has_more, next_cursor}`. + +Each entry has `id`, `created_at`, `admin_credential_id` (the first 8 hex characters of the Core key digest), `actor_label`, `action`, `project_id`, `resource_type`, `resource_id`, `result_ids`, `request_id` and `trace_id`. `result_ids` is an empty array except on `copy` entries. Deployment-wide entries have `project_id: null`, and a `project_id` filter excludes them. + +| `resource_type` | `action` | `resource_id` | +| --- | --- | --- | +| `project` | `create`, `rename`, `archive` | Project ID | +| `api_key` | `create`, `revoke` | Key ID | +| `executor_credential` | `issue`, `rotate`, `revoke` | Key ID | +| `session` | `archive` | Session ID | +| `agent`, `environment_template`, `skill`, `skill_version`, `file`, `vault`, `credential`, `session`, `artifact` | `delete` | Resource ID | +| `deployment_model_provider` (deployment-wide) | `set`, `delete` | Harness | +| `sandbox_deployment` (deployment-wide) | `change`, `replace_credential` (a provider credential was submitted, even the same one), `reset_start`, `reset_force`, `reset_deadline`, `reset_cancel`, `reset_complete` | Installation ID | + +An administrator write and its audit entry commit in one transaction; if the entry fails, the write fails. A reset's background archives and its `reset_deadline` and `reset_complete` entries carry the requester's credential, actor label, request and trace, and each archive keeps its Session's Project. Cancelling a reset does not undo archives already committed. Rejected provider verifications and no-op updates write no entry. Entries never contain credential values, request bodies or provider response text, and they survive the deletion of their resource and the revocation of keys. diff --git a/contracts/agents-api/core-errors.md b/contracts/agents-api/core-errors.md index f0868a7ec..12ee8db41 100644 --- a/contracts/agents-api/core-errors.md +++ b/contracts/agents-api/core-errors.md @@ -1,143 +1,77 @@ # Core administration errors -Errors on `/core/v1` retain the envelope below. `message` is safe English text; -`code` and `param` are nullable. Clients use stable `code` values and the optional -field `param`, and fall back to `message` for an unknown code. They do not parse -messages or retry rejected mutations automatically. +Errors on `/core/v1` use this envelope. `message` is safe English text; `code` and `param` are nullable. Clients act on the stable `code` and the optional `param`, show `message` for an unknown code, never parse messages and never retry a rejected write automatically. ```json {"error":{"message":"A valid Core key is required as the bearer credential.","type":"invalid_request_error","code":"invalid_admin_key","param":null}} ``` +Errors on `/v1` and `/api/v1` keep their own envelopes and never carry `details`. + ## Optional details -`error.details`, when present, is a nonempty flat object. Its values may only be -strings, finite numbers, booleans, null or arrays of strings (including empty -arrays). It contains documented Core-owned facts, never submitted names, URLs, -keys, echoed request values, native error text or provider response bodies. Each -operation that adds details must document its exact keys alongside its error -code. Operation-specific validation details are listed below; this contract does not -add diagnostics endpoints. - -The typed Go `CoreErrorDetails` values have string, number, boolean, null and -string-array constructors. `writeCoreError` emits them only through the marked -Core router; empty or invalid details are omitted as a whole. The mark preserves -error observation, flushing and `http.ResponseController` access. A shared -handler or a Core-looking request path alone cannot change a public or machine -error envelope. Core authentication still runs before operation configuration -checks, and unknown paths retain their existing status and admission rules. - -`AgentCoreError.details` is optional `CoreErrorDetails` in the TypeScript client. -The Core clients accept only the flat value types above, snapshot string arrays, -and ignore malformed or empty details without changing the error's message, -status, code, param or type. The public `OpenAIAgentsClient` does not read this -Core-only field. `/v1` and `/api/v1` response shapes remain unchanged. +`error.details`, when present, is a nonempty flat object. Its values are strings, finite numbers, booleans, null or arrays of strings (possibly empty). It holds only Core-owned facts: never submitted names, URLs or keys, echoed request values, native error text or provider response bodies. Each code that has details lists its exact keys below. + +| Code | Details | +| --- | --- | +| `sandbox_generation_stale` | `current_generation` | +| `sandbox_in_use` | `allocations`, `pending` | +| `sandbox_reset_required` | `current_provider`, `requested_provider` | +| Operation validation codes | See [operation validation](#operation-validation) | + +In the TypeScript client, `AgentCoreError.details` is the optional `CoreErrorDetails`. The Core clients accept only the value types above, copy string arrays, and ignore malformed or empty details without changing the error's message, status, code, param or type. The public `OpenAIAgentsClient` does not read `details`. ## Console-generated failures -The console uses the same envelope for its sign-in, origin, unsafe-request and -transport failures in the `/core` namespace. It does not expose request values -or transport exceptions. Core responses pass -through the proxy; the console does not reinterpret their codes or details. +Web's console server uses this envelope for its own failures on `/core` paths ([request boundary](../../docs/web/console-server.md#request-boundary)). It never exposes request values or transport exceptions, and it passes Core's responses through unchanged. -| HTTP status | Code | Meaning | Param / details | +| HTTP status | Code | Meaning | `type` | | --- | --- | --- | --- | -| 401 | `console_sign_in_required` | Console session is missing or expired | null / omitted | -| 403 | `console_origin_rejected` | Host, Origin or Fetch Metadata checks failed | null / omitted | -| 400 | `console_request_invalid` | Request path, method or upgrade is unsafe | null / omitted | -| 502 | `core_unreachable` | Core transport failed or Core tried to redirect | null / omitted | +| 401 | `console_sign_in_required` | The console session is missing or expired | `invalid_request_error` | +| 403 | `console_origin_rejected` | Host, Origin or Fetch Metadata checks failed | `invalid_request_error` | +| 400 | `console_request_invalid` | The path, method or upgrade is unsafe | `invalid_request_error` | +| 502 | `core_unreachable` | Core could not be reached, or Core answered with a redirect | `server_error` | -The first three use `type: "invalid_request_error"`; the last uses -`type: "server_error"`. A Core `401 invalid_admin_key` remains distinguishable -from a missing console sign-in. `/console/auth` keeps its existing -`{"error":"…"}` errors. Bare, retired and direct public/machine paths do not -become proxyable operations. Host, origin, authentication, credential stripping, -path checks and the no-retry rule are unchanged. +These have null `param` and no `details`. A Core `401 invalid_admin_key` therefore stays distinguishable from a missing console sign-in. Console sign-in routes keep their `{"error":"…"}` errors ([sign-in](../../docs/web/console-server.md#sign-in)). -Existing operation-specific codes remain documented in the -[administrator contract](admin-api.md), [sandbox deployment contract](sandbox-deployment.md), -[executor credential contract](environment-executor-credentials.md) and related -resource contracts. The following validators refine Core operation failures only. +## Sandbox provider verification -## E2B online changes +A `POST` or `PUT /core/v1/sandbox/deployment` ([sandbox deployment](sandbox-deployment.md#initialization-same-provider-changes-and-reset)) whose provider verifies a credential or configuration, as E2B does, fails with these fixed errors. None returns provider text, a template name, a key or a resource count. -The same-provider deployment PUT uses fixed safe errors. No provider response -message, template name, key or unlisted-resource count is returned in details. - -| HTTP | Code | Meaning | Param | +| HTTP | Code | Meaning | `param` | | --- | --- | --- | --- | -| 400 | `sandbox_credential_invalid` | Provider explicitly rejected authentication | `credential` | -| 400 | `sandbox_configuration_invalid` | Candidate immutable build is invalid or does not match resources | `configuration` | -| 409 | `sandbox_credential_ownership` | Candidate key does not prove ownership/manageability of the retained deployment | `credential` | -| 503 | `sandbox_verification_unconfirmed` | Verification, receipt settlement or bounded credential fencing could not be confirmed | null | - -Missing or unsettled Create receipts are uncertainty, never evidence of a different -team or released compute. The typed client projects these codes to fixed local -messages. It preserves the three nonnull fields above only when the response -status, code and param exactly match the table; all other params on these fixed -errors become null. Numeric details remain limited to `current_generation` for -`sandbox_generation_stale` and `allocations`/`pending` for `sandbox_in_use`. Unknown -credential-bearing errors become `sandbox_configuration_unconfirmed` without replay. -It also projects `409 sandbox_configuration_error` to fixed public-URL -guidance, even when a PUT omitted its key; arbitrary upstream text is never echoed. +| 400 | `sandbox_credential_invalid` | The provider rejected the candidate credential | `credential` | +| 400 | `sandbox_configuration_invalid` | The candidate configuration, such as an E2B template build, is not ready and immutable or does not match the resources | `configuration` | +| 409 | `sandbox_credential_ownership` | The candidate credential cannot manage the retained deployment; reset before changing accounts | `credential` | +| 503 | `sandbox_verification_unconfirmed` | Verification, receipt settlement or the credential fence could not be confirmed | null | + +On every deployment write, the typed client replaces the message of these codes and of the other `sandbox_*` deployment codes with fixed local text. It keeps only the `current_generation`, `allocations`, `pending`, `min` and `max` details, and keeps `param` only when status, code and param match the table or the `invalid_sandbox_configuration` rows below exactly. `409 sandbox_configuration_error` becomes fixed public-URL guidance with a null `param`, even for a `PUT` without a key. Any other error becomes `sandbox_configuration_unconfirmed` and is not resent, because a rejection could echo the key. ## Operation validation -All entries below return HTTP 400 with type `invalid_request_error`. Missing, -malformed or incorrectly typed model-provider bundles use `invalid_model_provider` -before field validation. JSON body parsing retains its existing errors. Other -malformed administration requests retain `invalid_request`. +Each code returns HTTP 400 with `type: "invalid_request_error"`. A missing, malformed or wrongly typed model-provider bundle returns `invalid_model_provider` before any field check. JSON body parsing keeps its own errors, and other malformed administration requests return `invalid_request`. | Code | Param | Details | Meaning | | --- | --- | --- | --- | -| `invalid_name` | `name` | `max_length`: 128 for Projects/nodes, 80 for Project keys | Name failed the resource's existing validator | +| `invalid_name` | `name` | `max_length`: 128 for Projects and nodes, 80 for Project keys | The name failed the resource's validator | | `invalid_node_capacity` | `max_active` or `max_retained` | `min`: 1, `max`: 1000000 | Capacity is invalid; retained capacity must also be at least active capacity | | `invalid_model_provider` | null | omitted | A complete model-provider bundle is required | | `model_provider_base_url_invalid` | `base_url` | omitted | Requires HTTPS without credentials, query or fragment | -| `model_provider_protocol_unsupported` | `protocol` | `harness` and `allowed_protocols`, from the build's adapter catalog | Protocol is unknown or unsupported by the selected harness | -| `model_provider_api_key_invalid` | `api_key` | `max_length`: 16384 | Key is empty, too long or contains a prohibited character | -| `model_provider_token_limits_invalid` | `context_window` or `max_output_tokens` | omitted | Limits are invalid or required positive limits are missing | -| `invalid_sandbox_configuration` | `resources.cpus` | `min`: 1, `max`: 255 | CPU count is outside the supported bounds | +| `model_provider_protocol_unsupported` | `protocol` | `harness` and `allowed_protocols`, from the build's adapter catalog | The protocol is unknown or unsupported by the selected Harness | +| `model_provider_api_key_invalid` | `api_key` | `max_length`: 16384 | The key is empty, too long or contains a prohibited character | +| `model_provider_token_limits_invalid` | `context_window` or `max_output_tokens` | omitted | Limits are invalid, or the Harness requires positive limits that are missing | +| `model_configuration_model_invalid` | `model` | omitted | The deployment default's model is not a nonempty model identifier | +| `harness_config_invalid` | `harness_config` | omitted | The deployment default's native parameters are unsupported or invalid | +| `invalid_sandbox_configuration` | `resources.cpus` | `min`: 1, `max`: 255 | The CPU count is outside the supported bounds | | `invalid_sandbox_configuration` | `resources.memory_mib` | `min`: 512, `max`: 1048576 | Memory is outside the supported bounds | -| `invalid_sandbox_configuration` | `resources.root_disk_mib` or `resources.environment_disk_mib` | `min`: 1024 for microsandbox; `min`: 0, `max`: 0 for Docker/E2B | Disk capacity is missing or unsupported by the provider | -| `invalid_sandbox_configuration` | `runtime` | omitted | Runtime release is missing, mutable, invalid or prohibited for E2B | - -Bounds describe validation constants, never submitted values. Node names retain -their existing byte limit and character rules; Project/key names retain their -trimmed Unicode character limit and control-character rules. No name rules are -widened or unified. Numeric checks retain their existing order, including -provider-dependent retained capacity validation inside the existing transaction. -Model-provider checks retain URL, protocol, key, general limits, harness protocol, -then required harness limits precedence. Resource checks retain CPU, memory, disk, -then Runtime precedence. - -Shared validators preserve their original error strings and sentinel identity. -Only the marked Core router maps their typed field metadata to this catalog; -public `/v1` and machine routes retain their previous complete error bodies. -Unknown sandbox providers retain the existing untyped error. E2B provider errors -above retain their fixed redaction contract with no echoed template or key. - -Core administration error details are scoped by the `/core/v1` router writer mark, -not a request path test. Use `writeCoreError` with typed `CoreErrorDetails` values -and document fixed keys in `contracts/agents-api/core-errors.md` when adding a -code. Include only safe Core-owned facts; never pass submitted values, secrets, -native text or provider bodies. Invalid/empty details are omitted. Preserve the -public and machine error serializers, observer callbacks and streaming interfaces. -The Core client ignores malformed optional details and never retries a mutation. - -Core operation validators preserve the original error text, sentinel identity, -and validation precedence. Package-owned typed errors carry fixed field metadata; -only the marked Core error mapper translates it to operation codes and safe -bounds/catalog details. Preserve Project/key rune limits and node byte limits -separately. Keep sandbox validation metadata through its existing store wrapper -without changing transaction or provider authority. Public Session provider -validation remains byte-compatible; cover it with handler-level golden responses. +| `invalid_sandbox_configuration` | `resources.root_disk_mib` or `resources.environment_disk_mib` | `min`: 1024 for microsandbox; `min`: 0, `max`: 0 for Docker and E2B | Disk capacity is missing or unsupported by the provider | +| `invalid_sandbox_configuration` | `runtime` | omitted | The Runtime release is missing, mutable, invalid or not allowed for E2B | + +Bounds are validation constants, never submitted values. Node names are limited in bytes; Project and key names in trimmed Unicode characters without control characters. Only the first failure is reported, in this order: model provider URL, protocol, key, general limits, the Harness's protocol, then the Harness's required limits; sandbox resources CPU, memory, disk, then Runtime. Model-provider field errors inside a `model_provider` object keep that object's field as `param`. An unknown sandbox provider returns an error without these fields. ## Diagnostic failure categories -The [root diagnostics reads](session-diagnostics.md) return these categories -inside a successful HTTP 200 snapshot, not the operation error envelope. Public -`/v1` Turn errors remain unchanged. `params` is `{}` unless specified. +The [Session and Turn diagnostics reads](session-diagnostics.md) return these categories inside a successful 200 snapshot, not as an error envelope. Public `/v1` Turn errors do not change. `params` is `{}` unless the table says otherwise. | Code | Stored cause or safe meaning | | --- | --- | @@ -167,21 +101,6 @@ inside a successful HTTP 200 snapshot, not the operation error envelope. Public | `environment_unavailable` | Environment unavailable for initial input | | `environment_provisioning_failed` | Hosted provisioning failure; params contain nullable `step`, `index`, `exit_code` from a sanitized receipt | -`diagnostics_unavailable` is the HTTP 503 operation error when the diagnostic -reader is not configured; it has no details. A database failure remains an error, -never a healthy or empty diagnostic snapshot. Historical provisioning reasons and -private native messages are not parsed for categories or parameters. - -Native categories apply only to a failed Turn with top-level -`error_code: engine_failed`. Only the finite `engine_error_code` allowlist is -accepted; unknown, -malformed and absent metadata retains `harness_error`. Only `connection_failed` -uses `engine_http_status`. Nested metadata and provider prose never classify a -failure. Core persistence, incomplete-stream and cancellation failures retain -priority, and cancelled/completed Turns have no failure. See -[native classification](native-error-classification.md) for adapter coverage. - -Model configuration writes additionally return `model_configuration_model_invalid` -with `param: model`, or `harness_config_invalid` with `param: harness_config`. -Both carry fixed messages without submitted values. Existing provider field errors -retain their field params within the `model_provider` object. +When the diagnostics reader is not configured, the reads return 503 `diagnostics_unavailable` without details. A database failure is an error, never an empty or healthy snapshot. Provisioning reasons and native messages are never parsed for categories or parameters. + +Native categories apply only to a failed Turn whose outcome has `error_code: engine_failed`. Core accepts only the listed `engine_error_code` values; an unknown, malformed or absent value stays `harness_error`. Only `connection_failed` uses `engine_http_status`. Nested metadata and provider text never classify a failure. Core storage, incomplete-stream and cancellation failures take precedence, and cancelled or completed Turns have no failure. [Native error classification](native-error-classification.md) lists which adapters report each category. diff --git a/contracts/agents-api/core-metrics.md b/contracts/agents-api/core-metrics.md index 3059e34b9..f5cf8cb85 100644 --- a/contracts/agents-api/core-metrics.md +++ b/contracts/agents-api/core-metrics.md @@ -1,21 +1,12 @@ # Core operational metrics -`GET /core/v1/metrics?range=1h|6h|24h|7d` is a Core-key read. It implements the response shape agreed with Core Web PR #96 -(`53dc9d646bc6e3cc2cd9b8bbb353bc53c3513ecf`). It changes neither public `/v1` -resources nor Agent or Sandbox metrics. Project API keys cannot call it. +`GET /core/v1/metrics?range=1h|6h|24h|7d` reports Core's own health: its process, execution queue and slots, PostgreSQL and background jobs. It requires the Core key ([Core administration API](admin-api.md)). -Only `range` is accepted, once; omission defaults to `1h`. Empty, repeated, -unsupported or other query parameters return `400 invalid_request`. A missing -metrics service returns `503 core_metrics_unavailable`. Partial measurement -failures return the usual `200 core.metrics` envelope with `service.status` set -to `degraded` and unavailable fields set to JSON null. No database or native -error text, credentials, bodies, resource IDs or tenant labels are exposed. +`range` is the only parameter, sent at most once; it defaults to `1h`. An empty, repeated or unsupported value, or any other parameter, returns 400 `invalid_request`. When Core has no metrics service, or cannot read it, the route returns 503 `core_metrics_unavailable`. When only some measurements fail, the response is still `200` with `service.status` set to `degraded` and each missing value set to null. The response never contains database or native error text, credentials, bodies, resource IDs or tenant labels. ## Time and missing data -The response's `range` uses UTC RFC 3339 boundaries. Its exclusive `end` is the -most recent complete bucket boundary. The included interval is `[start,end)`; -the current partial bucket is excluded from range aggregates and series. +`range` in the response has UTC RFC 3339 `start` and `end` and `resolution_seconds`. `end` is the most recent complete bucket boundary; the interval is `[start, end)`, so the current partial bucket is never included. | Range | Bucket size | Buckets | | --- | --- | --- | @@ -24,114 +15,49 @@ the current partial bucket is excluded from range aggregates and series. | 24h | 900 seconds | 96 | | 7d | 7200 seconds | 84 | -Current gauges and range aggregates are intentionally different: current worker, -connection pool, Go heap and goroutine values are read when requested; process -CPU, RSS and limits, queue gauges, database size are -sampled every 30 seconds with bounded I/O. -Samples older than 60 seconds are not reported as current. Queue, running and -pool and process series report the highest **observed** value in each bucket, not a claim -that all intermediate peaks were captured. Missing observations and the process's -partial first bucket stay null. Successful periodic ping samples produce linear -interpolated p50/p95; there is no request-triggered ping. - -A fixed-size in-process ring retains seven days of 30-second samples and -rejection counts, plus two hours of padding for complete bucket alignment. Restart loses those measurements: no synthetic backfill occurs. -The `execution.unavailable` count is null if the requested interval starts before -this process's observation began; an entirely observed interval with no rejections -is zero. PostgreSQL Turn history remains queryable across process restarts. -An empty queue has a measured count of zero but no oldest age. No started Turns -or successful ping samples means null percentiles, not zero latency. - -## Sources - -The envelope contains `object`, `range`, `service`, `execution`, `database`, -`jobs` and `process`, matching the typed client from PR #96. All numeric values -and `service.execution_owner` are nullable; lists of complete buckets are always -present. - -- `service.revision` is a full source commit injected into `main.buildRevision` - by the standalone builder's `-ldflags`. Manual builds without a valid revision - report null. `started_at` records process initialization. `execution_owner` - reflects the execution worker's existing lease checks, with unknown ownership - represented as null. Measurement or job failures report `degraded`; sandbox reset - is reported separately by the deployment contract, not a service status. -- `execution.slots_in_use` is the worker's active Session reservation set. Its - configured capacity is four; environment input, Turns and file work share it. - It does not count native harness subprocesses. Worker-disabled installations - have zero configured execution slots. -- `queued_turns` and `in_progress_turns` count root rows in `turns`, including - operational state retained for deleted Sessions. Native Subagent views and - pending Environment input reservations are not extra queued root Turns. - `waiting_for_daemon` is the queued subset whose Session device binding is - absent from the actual connected-device registry. `oldest_queued_seconds` - measures the oldest queued row's `created_at`. -- `queue_wait_ms` uses `started_at - created_at`, in milliseconds, for Turns - started in the interval. Each bucket uses its own started Turns, with PostgreSQL - `percentile_cont`. `interrupted` counts failed Turns whose outcome error code is - `execution_interrupted`, using `completed_at` in the interval. -- `unavailable` counts actual HTTP errors emitted with code - `execution_unavailable`, once per rejected response. Other 503 codes and errors - occurring after a stream has already started are not counted. The existing error - writer reports the code; no response/request body capture is involved. -- `database.ping_ms` measures a periodic pool ping, including connection acquisition. - Pool `in_use`, `idle` and `max` come from `pgxpool.Stat()`. `size_bytes` is - `pg_database_size(current_database())`, not host disk usage. Failure to measure - one value does not turn it into zero. -- `process.memory_bytes` is Go `runtime.MemStats.Alloc` (allocated heap bytes), - not RSS or container memory. `goroutines` is `runtime.NumGoroutine()`. -- `process.cpu_cores` is the increase in this process's user plus system CPU - time divided by elapsed sampling time. Linux uses `getrusage(RUSAGE_SELF)`; - subprocess and whole-host CPU are excluded. The first interval is null. - Missing/invalid counters, counter resets, nonpositive elapsed time and gaps - longer than 60 seconds reset the baseline; they never manufacture a zero. -- `process.rss_bytes` is Linux `/proc/self/status` `VmRSS`, converted from KiB - to bytes. `cpu_limit_cores` is the process's cgroup v2 `cpu.max` quota/period, - or `GOMAXPROCS` when its quota is `max` or the actual cgroup root has no quota interface. `memory_limit_bytes` is that cgroup's - finite `memory.max`; `max` is null. Membership and mount information resolve - the process's cgroup, including nested and subtree mounts. These are that - cgroup's configured limits, not whole-host metrics or ancestor-limit discovery. - Unreadable or malformed values remain null. Non-Linux builds report null CPU, - RSS and memory limit, with `GOMAXPROCS` as CPU capacity. -- `process.series` always contains the range's complete buckets with `start`, - `cpu_cores` and `rss_bytes`. Each measurement uses its own observed maximum; - missing samples and the partial first bucket stay null. These samples share - the existing bounded ring and restart gaps. Unsupported process measurements - do not by themselves change execution or database health. +- Execution slots, connected daemons, the connection pool, Go heap and goroutines are read when the request arrives. Process CPU, RSS and limits, queue counts and database size come from a sample Core takes every 30 seconds; a sample older than 60 seconds is not reported as current. +- Each series bucket reports the highest value observed in it, not every intermediate peak. Missing observations and the process's partial first bucket are null. +- Samples and rejection counts live in memory for seven days, plus two hours of padding for bucket alignment. A restart loses them; Core does not backfill. Turn history comes from PostgreSQL and survives restarts. +- `execution.unavailable` is null when the interval starts before this process began observing, and zero for a fully observed interval without rejections. +- An empty queue has a count of zero and a null oldest age. No started Turns or no successful ping samples give null percentiles, not zero latency. Percentiles of periodic pings (p50, p95) are linearly interpolated; a request never triggers a ping. + +## Fields + +The response has `object: "core.metrics"`, `range`, `service`, `execution`, `database`, `jobs` and `process`. Every numeric value and `service.execution_owner` is nullable; each `series` always lists every complete bucket of the range. + +| Field | Meaning | +| --- | --- | +| `service.status` | `running`, or `degraded` when a measurement or job fails, the latest sample is missing or stale, or execution ownership is unknown or Core has execution slots but does not hold the execution lease. A sandbox reset is reported by the [deployment](sandbox-deployment.md), not here | +| `service.revision` | The full source commit injected at build time; null for builds without one | +| `service.started_at` | When the process initialized | +| `service.execution_owner` | Whether this process holds the execution worker's database lease | +| `execution.slots_in_use`, `execution.slots_total` | Active Session reservations of the execution worker, and its capacity: [`core.execution_concurrency`](../../docs/configuration.md#settings), 4 by default. Environment input, Turns and file work share the slots; native Harness subprocesses are not counted. Without a worker both are 0 | +| `execution.queued_turns`, `execution.in_progress_turns` | Root Turns in those states, including Turns of deleted Sessions. Subagent Turns and input reserved for a preparing Environment are not counted | +| `execution.waiting_for_daemon` | Queued Turns whose Session's device is not connected; null without a gateway | +| `execution.oldest_queued_seconds` | Age of the oldest queued Turn, from its `created_at` | +| `execution.connected_daemons` | Runtime daemons connected to Core's gateway; null without a gateway | +| `execution.queue_wait_ms` | p50 and p95 of `started_at - created_at` for Turns started in the interval, by PostgreSQL `percentile_cont` | +| `execution.interrupted` | Failed Turns with error code `execution_interrupted`, by `completed_at` in the interval | +| `execution.unavailable` | HTTP responses sent with error code `execution_unavailable`, counted once each. Other 503 codes and errors after a stream started are not counted | +| `execution.series` | Per bucket: `queued`, `in_progress` and `queue_wait_p95_ms` | +| `database.ping_ms` | p50 and p95 of the periodic pool ping, including connection acquisition | +| `database.pool` | `in_use`, `idle` and `max` connections of the pool | +| `database.size_bytes` | `pg_database_size(current_database())`, not host disk usage | +| `database.series` | Per bucket: `ping_p95_ms` and `pool_in_use` | +| `process.memory_bytes`, `process.goroutines` | Go `runtime.MemStats.Alloc` (allocated heap, not RSS) and `runtime.NumGoroutine()` | +| `process.cpu_cores` | Increase in the process's user plus system CPU time divided by the elapsed sampling time (Linux `getrusage(RUSAGE_SELF)`), excluding subprocesses. Null for the first interval; a missing or reset counter, a nonpositive interval or a gap over 60 seconds restarts the baseline | +| `process.rss_bytes` | Linux `/proc/self/status` `VmRSS`, in bytes | +| `process.cpu_limit_cores` | The process's cgroup v2 `cpu.max` quota divided by its period, or `GOMAXPROCS` when the quota is `max` or the cgroup has no quota interface | +| `process.memory_limit_bytes` | The process's cgroup `memory.max`; null when it is `max` | +| `process.series` | Per bucket: `cpu_cores` and `rss_bytes` | + +Core resolves its own cgroup, including nested and subtree mounts, and reports that cgroup's limits, not the host's or an ancestor's. Unreadable or malformed values are null. Non-Linux builds report null CPU, RSS and memory limit, with `GOMAXPROCS` as the CPU limit. Missing process measurements alone do not make the service `degraded`. ## Background jobs -The four bounded IDs are `scheduler`, `runtime_sampler`, `history_cleanup` and -`audit_cleanup`. Each reports `status`, `last_run_at`, `processed` and `failed`. -A not-yet-observed run is unknown; a disabled or stopped loop is stopped. The -last-run time is the completion/observation of the last pass, not its next deadline. - -Scheduler processed counts selected Turn/environment work in that poll. Runtime -sampling counts observed and failed targets from its existing sweep result. -Cleanup processed counts confirmed removed rows; an unsuccessful cleanup reports -unknown processed count. Cleanup/scheduler failure counts identify failed passes, -not guessed numbers of lost rows or failed Turns. The actual scheduling, sampling, -retention and execution lifecycles keep their existing owners and timing. - -No additional telemetry database, monitoring server, model-provider probe, -message queue, object store, host disk measurement or scheduling mechanism is -introduced. Metrics cannot authorize execution or change resource ownership. - -## Implementation rules - -The Core-key-only `/core/v1/metrics` contract is documented in -[core-metrics.md](core-metrics.md). Keep this separate from -Agent outcome and Sandbox capacity views. Instrument existing worker and job -owners without changing scheduling, lease or retention behavior. Periodic pool -pings and bounded in-process samples have explicit restart gaps; unknown values -must remain null. Complete UTC buckets exclude the active partial bucket. Root -Turn history is queried read-only from PostgreSQL with native timestamps. -Count `execution_unavailable` at the existing HTTP error writer, once per rejected -response; never record request/response bodies or infer this count from every -503 or failed Turn. Builds inject the source commit with ldflags. No new monitoring -service or storage system is required. Keep the frontend response shape aligned -with the paired console contract. +`jobs` lists `scheduler`, `runtime_sampler`, `history_cleanup` and `audit_cleanup`. Each has `status` (`unknown` before its first run, `ok`, `failing`, or `stopped` when its loop is disabled or ended), `last_run_at` (when the last pass finished), `processed` and `failed`, both describing the last pass. -Core process CPU, RSS and cgroup limits are sampled by the existing 30-second -Core metrics loop; Go heap and goroutine reads retain their meaning. Process -series use the same bounded ring and complete buckets, with null first CPU -intervals and restart gaps. Do not substitute host usage for process usage. +- The scheduler's `processed` counts the Turn and Environment work it selected in its last poll. +- The Runtime sampler's `processed` and `failed` count the targets its last sweep observed and failed to observe. +- A cleanup job's `processed` counts the rows it removed. +- For the scheduler and the cleanup jobs, a failed pass sets `processed` to null and `failed` to 1; `failed` never estimates lost rows or failed Turns. diff --git a/contracts/agents-api/environment-files.md b/contracts/agents-api/environment-files.md index 13ff4c5c4..8e193e9fc 100644 --- a/contracts/agents-api/environment-files.md +++ b/contracts/agents-api/environment-files.md @@ -1,328 +1,112 @@ -# Environment files +# Environment files and Artifacts -The complete protocol target remains the SDK pinned in [upstream.json](upstream.json). -The public GET and POST `/agents/environments/{id}/files` have partial coverage. -Inline and source-file (`file_id`) creation target a qualified V1 local Environment, -including the Core-managed Docker profiles for Codex, Claude Code and MiniMax Code. -[Source Files](source-files.md) have their own project-owned lifecycle. Managed -hosted provisioning and shared Artifacts are accepted within the -[recorded Docker MVP scope](README.md#accepted-milestone-and-evidence) and separate -historical Core-managed [E2B qualification](README.md#e2b-v1-qualification). The -new user-managed enrollment chain reuses local Files with separate -[real public acceptance](harness-capabilities.md); complete Files/Environment -semantics and other providers are not implied. +A Session's workspace holds live files that the agent and its tools change. `/agents/environments/{environment_id}/files` lists one workspace directory and creates files in it. When a Turn completes, Core copies the files under the workspace's `outputs/` directory into immutable Artifacts, read through `/agents/sessions/{session_id}/artifacts`. Artifacts outlive the Environment; workspace files do not. -## Pinned contract +The routes follow the SDK pinned in [upstream.json](upstream.json): the [Environment files resource](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/files.py), its [list parameters](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agents/environments/file_list_params.py), [EnvironmentFile](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agents/environments/environment_file.py) and [TokenPage](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/pagination.py). They need the `OpenAI-Beta: agents=v1` header. -The [Files resource](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/files.py) -and [list parameters](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agents/environments/file_list_params.py) -specify optional absolute-directory filtering, limit 1–100, case-sensitive -path-component ordering (default descending), and an opaque `page` token with -unchanged path/order/limit across pages. Limit and path are nullable SDK inputs; -order and page are not nullable when supplied. +## Where files work -Each [EnvironmentFile](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agents/environments/environment_file.py) -has `environment_id`, `object: agent.environment.file`, absolute `path`, and integer -`size_bytes`. The pinned [TokenPage](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/pagination.py) -requires `data`; `has_more` and `next` are optional and nullable. This implementation -returns the observed official envelope `object: page`, `data`, `next` (null on the -final page) and `has_more`, which is true exactly when `next` is set. +- Environment files work on `openai_hosted` and `self_hosted` Environments. A `none` Environment returns 503 `execution_unavailable`. +- Public paths start at `/workspace`, which stands for the Environment's workspace directory wherever it is on the machine. +- The Environment is looked up in the caller's Project first; a missing or foreign ID returns 404. +- On an `openai_hosted` Environment that is still `pending`, both operations return 400 `the hosted environment is still provisioning; wait until it is connected before accessing files`. This check runs after request validation and before any source File lookup. Any other state proceeds, and an unreachable Runtime returns 503. +- Listing and writing never start a Turn or send input to the model. -## Current scope and local policies +## List files -- Read direct regular files in one authorized self-hosted or qualified local workspace - directory. Local public paths are rooted at `/workspace`, independently of the - physical path frozen into its dedicated Runtime. - Omitted path selects the workspace root. Do not recurse or follow symlinks; - directory, symlink and other non-regular entries are omitted. A path that is - missing, a regular file or a symlink lists an empty page; the link is not followed. - Daemons without a local workspace binding use the Claude SDK adapter reader, - which keeps 404 for a missing path and 503 for a regular file or symlink. -- Omitted limit uses 20. Well-formed unknown query keys are ignored; a repeated - `path`, `limit`, `order` or `page` returns the Beta duplicate-field error. Unlike - the shared lists, which drop malformed pairs, malformed query encoding (such as - `?foo=%GG` or a `;` separator) still rejects the request locally. Empty values are - rejected by their own rule. Limit errors use the Beta `invalid_request_error` code. The pinned SDK's - [query serializer](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/_qs.py) - omits scalar `None` values, so nullable limit/path follow omission behavior. - Literal `null` and empty scalar query values are not accepted. -- Require an absolute UTF-8 POSIX directory of at most 4096 bytes within the - workspace, in cleaned form. Backslash, NUL, CR and LF are rejected. Redundant or - trailing separators, `.` and `..` components are rejected rather than normalized; - the exact directory binds the cursor and is passed workspace-relative to execution. -- Authorize through the existing tenant-scoped Environment lookup before inspecting - directory paths, cursors or runtime availability. Existing project-shared reads - remain permitted. The reader receives that exact Environment and rechecks its - execution ownership; the API never selects a daemon or a local filesystem path. -- Accept only a complete validated native directory, bounded by the shared - 1024-entry limit. Truncation, unknown kinds, missing sizes, duplicate or unsafe - names, and uncertain output return safe 503 without `data` or `next`. The bound - applies before filtering non-regular entries, sorting or public pagination. -- Cursors are bounded base64url tokens tied to tenant, Environment, canonical - directory, effective order/limit, and the full sorted regular-file path/size - result. Every page rereads the directory. Changed files or parameters invalidate - continuation with safe 400. No cache, durable cursor registry or snapshot is - promised; unchanged path/size metadata does not prove unchanged contents. -- Reader errors reuse the existing safe error mapping: invalid input 400 and - unavailable execution 503. Only the native reader's distinct `not_directory` - result becomes an empty page; a missing workspace root or other native `not_found` - keeps 404, and permission, transport and uncertain failures keep 503. Native error - text never enters the response. - Listing does not create a Turn or admit model input. Idle reads use temporary - read-only preparation; actual transport disconnect/reconnect events remain visible. +`GET /agents/environments/{environment_id}/files` lists the regular files directly in one directory. It does not recurse. Directories, symbolic links and other entries are left out. -Default limit, omitted-path scope, recursion, non-regular entries and cursor -invalidation behavior are local policies or remaining gaps, not verified hosted -semantics. The pinned source does not establish them. The sampled envelope, query, -path and empty-page rows are recorded in [wire alignment](#wire-alignment--september-23-2026). -Do not interpret the bounded direct-file implementation as complete Files.list -compatibility. +| Parameter | Rule | +| --- | --- | +| `path` | An absolute directory in clean form at or below `/workspace`; default `/workspace` | +| `limit` | 1–100; default 20 | +| `order` | `desc` (default) or `asc`, by path in byte order | +| `page` | The `next` token of the previous page, with the same `path`, `order` and `limit` | -## Inline and source-file creation +The response is `{"object": "page", "data": [...], "next": …, "has_more": …}`. `next` is null on the last page, and `has_more` is true exactly when `next` is set. Each file has `environment_id`, `object: "agent.environment.file"`, the absolute `path` and `size_bytes`. -The pinned create union requires `type: inline`, standard Base64 `data` and an -absolute destination `path` under `/workspace`, or `type: file_id`, `file_id` and -that path. Source IDs resolve only within the authenticated execution project; -filenames, URLs and filesystem paths cannot substitute for an ID. Both members -use the same destination writer. Required null/omitted fields, extra fields and -invalid Base64 are rejected; an unknown top-level field names itself as `param` -when its name is short and printable. -Unknown query keys are ignored. Empty bytes are valid. Paths must -be canonical and cannot name the workspace root; missing parent directories are -created. Inline data is limited to 5 MiB decoded, the official bound; a `file_id` -copy keeps the local 50 MiB destination limit, with bounded JSON and 64 KiB daemon -frames. See the [write semantics](#write-semantics--september-23-2026). +- A `path` that does not exist, names a regular file, or passes through a symbolic link returns an empty page. Links are never followed. +- Each page reads the directory again; there is no snapshot. If the directory's regular files (their names or sizes) or the request's parameters changed since the token was issued, the token is rejected. Unchanged names and sizes do not prove unchanged contents. +- A directory with more than 1,024 entries of any kind returns 503 and no partial page. Permission errors, a missing workspace root and transport failures also return 503. +- When the daemon has no local workspace binding, the Claude Code adapter answers the read instead: a missing path returns 404, and a regular file or symbolic link returns 503. -Creation uses the same tenant Environment lookup as listing. The Worker checks the -stored local profile, immutable exact device/Environment binding and live capability. -It never starts a model for upload or supplies a filesystem root from the request. -The deployment must qualify the protected sibling workspace/staging layout and -its selected native adapter. The [engine profile guides](README.md#public-engine-profiles) -describe accepted Docker configurations; the [E2B operator guide](../../services/core/deploy/e2b/README.md) -covers user-managed E2B Runtime packaging and links its separate real acceptance. -A capability or path declaration alone -does not establish isolation or public hosted admission. +Query errors, all with type and code `invalid_request_error` and a null `param` unless noted: -Before sending any bytes, persist the mutation identity and request digest under -the Session lock. Pending input/execution and another unresolved upload exclude a -new mutation. The Runtime receives the complete body, verifies its digest and uses -the shared installer in its create mode: it creates missing parents and installs a -fresh mode-0600 file without replacing or following any existing path. Later -independent tool writes can change the installed file; the response does not -promise a snapshot. +| Case | Message | +| --- | --- | +| `path` relative, outside `/workspace`, longer than 4,096 bytes, not UTF-8, or containing a backslash, NUL, CR or LF | `path must be an absolute directory inside /workspace` | +| `path` not in clean form: a trailing or repeated `/`, `.` or `..` | `path must identify a non-reserved directory inside /workspace` | +| A malformed token, another request's token, or a token whose listing changed | `Invalid file page token for this request` | +| `limit` outside 1–100 | `limit must be between 1 and 100`; malformed integers and invalid `order` use the shared [Beta list errors](wire-semantics.md#lists) | +| A repeated `path`, `limit`, `order` or `page` | The shared duplicate-field error ([list rules](wire-semantics.md)) | -Only an exact committed/rejected receipt settles durable ownership. Caller detach, -connection loss, timeout or missing output cannot be treated as rejection. Unknown -writes remain pending across Core restart and block successor mutation without -replay; read-only recovery remains available. Automatic uncertain-write recovery -and placement replacement are outside this batch. Controlled failures preserve -the destination only when the installer proves rejection, and only against this -operation, not independent workspace writers. Temporary-file cleanup is best effort. +Unknown query keys are ignored. An explicit empty value is invalid for every key. Malformed query encoding, such as `%GG` or a `;` separator, returns 400 `invalid_request`. -Successful creation returns 201 with only the four EnvironmentFile fields. Reuse the -common safe error mapper; apart from the sampled rows below, current 400/409/413/503 -policies and error timing are not evidence of exact upstream parity. This referenced-source milestone cannot close the complete Files -resource or Environment lifecycle requirements. +## Create a file -Source resolution reads an immutable snapshot before destination admission. A -source deleted before that lookup is unavailable; an already-resolved copy may -finish. Deleting a source never deletes a copied workspace file. A larger general -Files upload can be downloaded but is rejected before Environment dispatch when -it exceeds the destination limit. Exact hosted delete/copy timing is unverified. +`POST /agents/environments/{environment_id}/files` takes either form: -## Acceptance boundary +```json +{"type": "inline", "data": "", "path": "/workspace/data/input.csv"} +{"type": "file_id", "file_id": "file-…", "path": "/workspace/data/input.csv"} +``` -API tests cover raw response fields, complete-result validation, filters, sorting, -pagination, local cursor policies, authorization order and safe errors. The separate -official-client fixture exercises flat directories generated through a real model, -raw HTTP and pinned SDK pagination, sizes, and two-tenant isolation. It does not -establish unspecified recursive, symlink or snapshot behavior. Runtime availability -and each engine's isolated placement require their own native and service checks. +Empty inline `data` is valid and creates an empty file. It returns 201 with the four EnvironmentFile fields. `file_id` names a [File](source-files.md) of the same Project; Core reads its bytes before contacting the Runtime. -The opt-in `services/core/tests/official_environment_files_create.py` reuses -the pinned SDK and raw HTTP listing assertions. Its stdin supplies the base URL, -preconfigured Environment ID, model-input text, and two private caller token -sources (`token_env` or `token_file`). The invoking native fixture supplies an -existing `uploads` directory and a staging symlink rejection probe, verifies exact -installed hashes, and has a real model consume source-copied text after source -deletion. Its source fixture also streams a 512 MiB upload, checks public download -denial, and verifies that it cannot bypass the smaller destination bound. Internal -large-object streaming integrity has separate PostgreSQL tests. This distinction -keeps private setup separate from public hosted creation acceptance. Mechanism tests -exercise detached/unknown outcomes and durable gates independently of model output. +| Case | Result | +| --- | --- | +| Missing parent directories | Created with mode 0700; the file gets mode 0600 | +| Parent directory is a symbolic link to a directory inside the workspace | The link is followed and the file is created at its target | +| Destination exists as a file, a symbolic link or another non-directory entry | 400 `environment.files paths must not traverse symlinks or overwrite existing files`. Nothing changes | +| Destination is an existing directory | 400 `file path conflicts with an existing environment file` | +| A parent component is a regular file, or a symbolic link leading outside the workspace | 400 `invalid_request`, `Invalid resource identifier or request limits.` | +| `inline` data above 5 MiB after decoding | 400 `environment.files[0].data exceeds the 5 MiB decoded limit`, before any Runtime work. Exactly 5 MiB is accepted | +| `file_id` File above 50 MiB | 413 `request_too_large` | +| Request body larger than the Base64 form of 50 MiB plus 16 KiB | 413 `request_too_large` | +| `path` relative, the root itself, outside `/workspace`, longer than 4,096 bytes, not UTF-8, or containing a backslash, NUL, CR or LF | 400 `environment.files[0].path must be an absolute POSIX path inside /workspace` | +| `path` with an empty, `.` or `..` component | 400 `environment.files[0].path cannot contain empty, . or .. path components` | +| An unknown top-level field | 400 `Unknown parameter: ''.` with `param` set to the field. A name longer than 256 bytes or with unprintable characters gets `Unknown parameter.` and a null `param`. Only the first unknown field in the body is reported | +| A missing or null required field, the other form's field, an unknown `type`, or invalid Base64 | 400 `invalid_request` | +| Missing or foreign `file_id` | 404 `not_found_error` | -## Wire alignment — September 23, 2026 +Unless the table names a code, the 400 errors have type and code `invalid_request_error` and a null `param`. An existing destination is never replaced: the Runtime writes a temporary file and publishes it with a hard link that fails if the destination exists. Later tool writes can still change the created file. -The pin is unchanged: SDK 3.13.0, commit `d7c41ef`, `agents=v1`. This batch starts -from main `e4b124c` and aligns Files.create and Files.list wire behavior with the -hosted-environment campaign scan, recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/` (`findings.json` -HE-10, 16, 18, 32, 34–39; raw records under `official/` and `run1/`). The probe -used three owned hosted Sessions, all deleted. The batch plan is -`~/.parsar/remediation/20260923/environment-files-wire/PLAN.md`. +### Write ordering and uncertain outcomes -| Row | Case | Core behavior | Evidence (finding: request ID) | -| --- | --- | --- | --- | -| F1 | Successful Files.create | 201 with the same four fields | HE-10: `req_cce5244cdd76471d85575b7bbcc3fdac`, `req_69a44a96866a4b52b637a5a5530dcd3c` | -| F2 | List envelope | `object: page`, `data`, `next`, `has_more`; `has_more` is true exactly when `next` is set. Paging and ordering are unchanged | HE-32: `req_c8c247cfd13145a2b92be5cad412508f`, `req_a311bf80ff0643db945aa5db0a78e95b` | -| F3 | Unknown list query key | Ignored; the page equals the request without it. Malformed query encoding (such as `?foo=%GG` or `;` separators) is still rejected locally, while the shared lists drop those pairs | HE-34: `req_c4cd4618f7264840ad7454e80ac7f83c` | -| F4 | Repeated `path`, `limit`, `order` or `page` | 400 `invalid_request_error`, param null, ``Failed to deserialize query string: duplicate field `` `` | HE-35: `req_30b3c8bfced64383bf13c9d5bbc44a74` | -| F5 | `path` names a missing directory | 200 `{object: page, data: [], next: null, has_more: false}` with a local workspace reader (Claude SDK adapter exception below) | HE-36: `req_a311bf80ff0643db945aa5db0a78e95b` | -| F6 | `path` names a regular file or a symlink to a directory | The same empty page with a local workspace reader; the link is never followed | HE-37: `req_25fb0220fa7d4157a783711622b5d580`, `req_58b6784dc49640c194910c719d5c8102` | -| F7 | `path` not in cleaned form (trailing or repeated separator, `.` or `..`) | 400 `invalid_request_error`, param null, `path must identify a non-reserved directory inside /workspace` | HE-38: `req_a7ee5a49d1d24529b9de877c7ec5bc58` | -| F8 | Other validation errors | Code `invalid_request_error`, param null: list relative or outside path `path must be an absolute directory inside /workspace`; malformed, foreign or stale page token `Invalid file page token for this request`; create relative, root, outside or NUL path `environment.files[0].path must be an absolute POSIX path inside /workspace`; create empty, `.` or `..` components `environment.files[0].path cannot contain empty, . or .. path components`; unknown create body field `Unknown parameter: ''.` with param `` | HE-16: `req_3fb9feef630c442ba1c136506458362d`, `req_a41eb84c48594daf8fb55b95d8b118bd`, `req_f592639cfe28416395187ae76ad18228`; HE-39: `req_cbfed6a0336e47a395f42b82b45b20d7`, `req_5c2494e082714500bccbadbf60c6195d` | -| F9 | Files.create or list on an `openai_hosted` Environment still `pending` | 400 `invalid_request_error`, param null, `the hosted environment is still provisioning; wait until it is connected before accessing files` | HE-18: `req_938b99e48d3c4f4ab685dcf9b29baa6f`, `req_ff81d9155bb841618d5d1bbc1c987fdf` | +- Before sending any bytes, Core records the write under the Session lock. While input is pending, a Turn is running or an earlier write is unsettled, a new write returns 409 `turn_conflict`. An unsettled write also makes new messages to the Session return 409. +- The Runtime checks the complete body against its digest before creating anything, so incomplete input creates nothing. A write that fails later can leave newly created empty parent directories. +- Only a definite receipt from the Runtime settles a write, as committed or rejected. A rejected write changes nothing and releases the Session. If the connection drops, the request times out or no receipt arrives, the request returns 503 and the write stays unsettled, across Core restarts. Core never resends it and has no automatic recovery, so the Session accepts no further writes or messages. Reads still work. +- Deleting the source File after its bytes were read does not affect the copy. -Decisions: +## Artifacts -- F5/F6 are result mapping only. The native directory helper walks the requested - path below the anchored root with `O_PATH | O_DIRECTORY | O_NOFOLLOW`. When that - walk fails with `ENOENT` or `ENOTDIR`, it reports the distinct result - `not_directory`: a missing component, a regular file, a symlink to a directory, - a dangling or external symlink, and a path through any of them. A symlink fails - at its own component, so no target is followed, opened or listed. Every - component is validated before any is opened, so an invalid request cannot become - an empty page. The daemon and gateway carry `not_directory` only for directory - reads. Core lists nothing for it only after the same confirmed release that a - listing needs. -- Real failures keep their errors. The existing `invalid_path` code also covers - unsafe entry names inside a directory, and `not_found` also covers a missing - workspace root and an entry removed during observation; mapping either would hide - a failure as an empty listing. A missing or replaced root keeps 404 or 503. A - root removed or replaced after it was opened fails lookups or reads as empty. - So before reporting `not_directory` or an empty listing, the helper reopens the - root path without following links and compares device and inode with the held - root; a removed or replaced root keeps 404. Link counts are not used, because - overlayfs can keep a nonzero count for a removed lower-layer directory. The - check relies on local filesystem inode pinning; network filesystems such as NFS - are not qualified. - Permission denial, observation errors, transport loss and uncertain output keep - 503, and tenant and Environment authorization run first. The shared workspace - path code and the write installer are unchanged. -- The list path is rejected instead of normalized. `..` and other non-clean forms - use the F7 message. Outside paths, backslash, control characters, invalid UTF-8 - and paths over 4096 bytes use the F8 absolute-directory message; only the - relative form was sampled. -- The create messages keep the observed `environment.files[0].path` field name, - which comes from the official service's shared file validation. Backslash, CR, - LF, invalid UTF-8 and overlong paths are unsampled and use the absolute-path - message. The accepted path set is unchanged. -- The first unknown create field in document order is reported. Its name is - repeated in the message and param only when it is at most 256 bytes of - printable UTF-8; otherwise the error keeps code `invalid_request_error` with - param null and the message `Unknown parameter.`, so the response stays bounded. - A field of the other union member (such as `file_id` on `inline`), missing or null fields, an - unknown `type` and invalid Base64 keep the local `invalid_request` code; none was - sampled. -- Every rejected page token uses the sampled token message, including a valid token - for other parameters or a changed directory. -- F9 runs after the tenant-scoped lookup and request validation, and before - source-file resolution or execution. It applies only when the stored type is - `openai_hosted` and its status is `pending`; a disconnected hosted Environment - and every `self_hosted` status keep the existing execution path. The official - order between validation and this check was not sampled. -- Core Web sends the cleaned form of a directory entered with a trailing slash. - The client and Web accept only the new envelope and 201. +### Capture -Deferred and unchanged: recursive listing (HE-30); creating parent directories, -overwrite and the inline limit (HE-11/12/15), since aligned by the -[write semantics](#write-semantics--september-23-2026) batch; the Environment -retrieve `files[]` projection (HE-03); Environment -events (HE-02); the `self_hosted` status value (HE-04); `self_hosted` Files -support, which the official service refuses and Core's daemon keeps; limit bounds, -the default path and ordering. Malformed query encoding (such as `?foo=%GG` or `;` -separators) stays a local rejection on this list, while the shared lists drop those -pairs; there is no official sample. -The deleted-Session 404 message is not adopted; the write semantics batch adopts -the conflict and size-limit messages. The Claude SDK adapter directory reader, used -only when a daemon has no local workspace binding, is unchanged: F5 and F6 there keep 404 for a missing -path and 503 for a regular file or symlink. A Runtime image built before this batch -keeps the same 404/503 results until it is rebuilt. +When a Turn completes on an `openai_hosted` or `self_hosted` Environment, Core copies every regular file under `/workspace/outputs/` and publishes the copies in the same transaction that completes the Turn. Each Artifact's `path` is the file's absolute workspace path, such as `/workspace/outputs/report.md`. Failed and cancelled Turns publish nothing. Without an `outputs/` directory there is nothing to capture. -Go handler tests cover F1–F9 with foreign-equals-missing checks; real-PostgreSQL -Worker tests cover the `not_directory` mapping, its release confirmation and the -unchanged rejections; gateway, daemon and Rust helper tests cover the native -classification, including links to a directory and outside the workspace, dangling -links and a replaced root. The TS client and Web unit tests cover the envelope and -201. The opt-in official scripts replay the list and create rows through raw HTTP -and the pinned SDK. Live acceptance, the server gate and independent review are -recorded separately by the coordinator. +- Symbolic links anywhere below `outputs/` are skipped by their own type: never followed, opened or resolved, and never an Artifact. The remaining files are still captured. +- A later Turn in the same Session publishes a path only when no Artifact remains for it in the Session, or when its bytes (SHA-256) differ from the newest remaining Artifact for that path. Newest follows the order of the producing Turns. An unchanged path keeps its existing Artifact and ID. Published Artifacts are never modified. +- The capture fails, and the Turn fails with `artifact_capture_failed` (or ends `cancelled` when cancellation was requested), when `outputs` is not a directory (including a symbolic link to one), an entry is a FIFO, socket or device, a file changes size or modification time while it is copied, or a limit is exceeded: 4,096 entries, directory depth 64, a 4,096-byte path, 200 MiB per file or 500 MiB per Turn. -## Write semantics — September 23, 2026 +### Read and delete Artifacts -The pin is unchanged: SDK 3.13.0, commit `d7c41ef`, `agents=v1`. This batch starts -from main `1eb60c27` and aligns Files.create writes with the hosted-environment -campaign scan, recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/` (`findings.json` -HE-11, 12, 13 and 15; raw records under `official/` labelled -`fc02-nested-missing-parents`, `fc03-overwrite`, `fc06-5mib-plus-1`, -`fc15-onto-directory`, `fc20-overwrite-untracked`, `fc21-through-symlink-dir` and -`fc22-symlink-outside`), and with the "File limits" of the official Environment -files guide saved beside them. The batch plan is -`~/.parsar/remediation/20260923/workspace-file-writes/PLAN.md`. +| Operation | Behavior | +| --- | --- | +| `GET /agents/sessions/{session_id}/artifacts` | Lists the Session's Artifacts | +| `GET /agents/sessions/{session_id}/artifacts/{artifact_id}` | Returns `id`, `object: "agent.session.artifact"`, `session_id`, `turn_id`, `environment_id`, `path`, `size_bytes` and `created_at` (publication time, Unix seconds) | +| `GET /agents/sessions/{session_id}/artifacts/{artifact_id}/content` | Streams the bytes as `application/octet-stream`, with the file's base name as the attachment filename | +| `DELETE /agents/sessions/{session_id}/artifacts/{artifact_id}` | Returns `{"id": …, "object": "agent.session.artifact.deleted", "deleted": true}` | -| Row | Case | Core behavior | Evidence (finding: request ID) | -| --- | --- | --- | --- | -| FW1 | Missing parent directories | 201. Each missing component is created with mode 0700, the initial-file installer's convention, below the held no-follow walk | HE-11: `req_69a44a96866a4b52b637a5a5530dcd3c` | -| FW2 | The destination is a file that an earlier Files.create wrote | 400 `invalid_request_error`, param null, `environment.files paths must not traverse symlinks or overwrite existing files` (official: `file path conflicts with an existing environment file`). No bytes change | HE-12: `req_f661f17c5c714d829364e13da32db6b1` | -| FW3 | The destination is another existing file, such as one created by setup, a native tool or the model, or any other non-directory entry | 400 with the FW2 message. No bytes change | HE-12 `fc20`: `req_1242efc5adaa4ac5a854d2973ec81cd2` | -| FW4 | The destination is an existing directory | 400 `invalid_request_error`, param null, `file path conflicts with an existing environment file` | HE-12 `fc15`: `req_c835a06aa7d949e0ae7883ef83c3145b` | -| FW5 | A symlink anywhere in the parent chain or as the destination, or a non-directory parent component | 400 with the FW2 message. No link is followed | HE-13: `req_0c2c3cbf9e9947d79921cd198f42e9a6`, `req_02f00ab8a6c348aca0bfe47ac3704b18` | -| FW6 | Inline data above 5 MiB decoded | 400 `invalid_request_error`, param null, `environment.files[0].data exceeds the 5 MiB decoded limit`, before any Runtime work. Exactly 5 MiB is accepted; a `file_id` copy keeps the 50 MiB bound | HE-15: `req_088ba87e6e3e4436a7713015e37615b5`; guide file limits | -| FW7 | The destination or a link appears between the checks and the install | Never replaced or followed; the FW2 message | Rust race tests; not sampled officially | -| FW8 | Unknown outcome | The durable mutation gate and unknown handling are unchanged | — | -| FW9 | `self_hosted` | The daemon's local workspace writer serves Files.create for Core-managed Docker and user-managed Runtimes alike, so the same rules apply | — | -| FW10 | Initial Session files, Skills, Plugins and cold resume | Unchanged | — | -| FW11 | Tenant isolation, missing or foreign Environment | Unchanged 404 | — | +- Reads work whether or not the Environment still exists, including after it expires. +- Deleting an Artifact leaves the workspace file alone. A content read already in progress can finish; later reads return 404. Removing a file from the workspace leaves its Artifacts in place. +- Deleting the Session deletes its Artifacts. -Decisions: +List parameters: -- The daemon's Go workspace writer now serves Files.create on all platforms. - It verifies the complete body before creating parents, writes a temporary file, - and publishes it with a no-overwrite hard link. Initial Session files use a - separate atomic replacement operation. Skills and Plugins use - [shared Runtime capability preparation](environments.md#runtime-capability-preparation). - Completed recovery, including cold resume, never reinstalls files. -- Parents are created only after the complete body is verified, so incomplete or - corrupt input creates nothing. A write failure can leave newly created empty - parent directories, but never replaces an existing destination. -- Core cannot tell a file that an earlier Files.create wrote from any other file. - Its durable write intent stores a digest of the path, size and content, not a - path ledger, and a tool can remove and recreate a file afterwards; no cheap, - race-safe lookup exists. FW2 therefore uses the untracked-file message, a - message-only difference. A directory destination always uses the conflict - message; the official sample was a directory created as a parent by an earlier - write, and other directories are unsampled. Non-directory parent components and - non-regular destinations are unsampled and use the FW2 message. -- The installer reports a fixed code, the daemon sends `write_rejected` with an - optional `reason`, and Core settles the durable intent as `rejected`. A rejected - write therefore leaves no committed receipt, releases the mutation owner, - changes no bytes and reads but does not consume a Source File. -- The inline bound is checked after path validation and Base64 decoding, before - the F9 provisioning check, Source File lookup and execution. The JSON body limit - still admits Base64 for up to 50 MiB, so larger inline bodies up to that size get - the official message; beyond it the local 413 remains. -- A Runtime image built before this batch keeps replacement and the existing-parent - requirement, with the generic local 400, until it is rebuilt. An older Core - ignores the new `reason` and keeps the generic 400. A new daemon with an older - helper is refused as invalid input, without mutation. -- The TypeScript client and Core Web reject inline data above 5 MiB before sending. - The client's strict Base64 check now scans linearly, because the previous pattern - overflowed the regular-expression stack on inputs near 4 MiB and blocked an exact - 5 MiB upload. +| Parameter | Rule | +| --- | --- | +| `order` | `desc` (default) or `asc`, by publication time, then ID | +| `after` | An Artifact ID of this Session | +| `environment_id` | Only Artifacts produced in that Environment. A malformed or unknown ID returns an empty page; an empty value means no filter | -Rust tests cover parent creation, existing files, directories, hard-link aliases, -FIFOs, symlink leaves and parents, dangling and inside links, non-directory parents, -races for the destination and a missing parent, a replaced held ancestor, the -`linkat` fallback and the explicit mode selection. Daemon tests cover the local -workspace writer's create mode and codes with a scripted helper; its native-helper -case runs only when `OAC_TEST_LOCAL_WRITE_HELPER` names a built helper, so it is -opt-in and was run for this batch by hand and through live Docker acceptance. -Dispatch and gateway tests cover the rejection reason. Go handler tests cover FW6 and the error mapping. A real-PostgreSQL -HTTP and Worker test covers the conflict messages, settled `rejected` intents with no -committed receipt, an unconsumed Source File, the inline bound before any intent, -tenant B isolation and an admitted successor. The opt-in -`official_environment_files_create.py` adds the nested, repeated, directory and -5 MiB + 1 rows, with an exact 5 MiB inline case while source copies keep 50 MiB. -Live acceptance, the server gate and independent review are recorded separately by -the coordinator. +`limit` and cursor errors follow the shared [list rules](wire-semantics.md#lists). The response is `{"object": "list", "data": [...], "first_id", "last_id", "has_more"}`; an empty page has null IDs. The Session is looked up first, so a missing or foreign Session returns 404 whatever the query. diff --git a/contracts/agents-api/environments.md b/contracts/agents-api/environments.md index a6a59b674..cb20c1f01 100644 --- a/contracts/agents-api/environments.md +++ b/contracts/agents-api/environments.md @@ -324,7 +324,7 @@ template = client.beta.agents.environments.templates.create( ) ``` -The pinned `/v1/skills` resource, version and content routes use the Project API key without the Agents beta header. ZIP uploads use `files` and directory uploads repeated `files[]`. The pinned SDK 3.13.0 drops a single file tuple during multipart extraction, so upload a single ZIP with raw HTTP. An upload holds at most 500 regular files and exactly one `SKILL.md`, 5 MiB compressed and 20 MiB expanded. [File resource semantics](file-resource-semantics.md) owns default-version and deletion rules. +The pinned `/v1/skills` resource, version and content routes use the Project API key without the Agents beta header. ZIP uploads use `files` and directory uploads repeated `files[]`. The pinned SDK 3.13.0 drops a single file tuple during multipart extraction, so upload a single ZIP with raw HTTP. An upload holds at most 500 regular files and exactly one `SKILL.md`, 5 MiB compressed and 20 MiB expanded. [File resource semantics](source-files.md#versions-and-metadata) owns default-version and deletion rules. A reference with an omitted or null version selects the default at Session creation, `"latest"` the latest version, and a positive version string that version. Template responses keep the unresolved selector (`version: null` for the default); Session metadata shows `{type, skill_id, version, name, description}` with a concrete version. A Session freezes the selected version's bytes and metadata in its creation transaction; later default changes, source deletion or Template updates cannot change it. diff --git a/contracts/agents-api/execution-configuration.md b/contracts/agents-api/execution-configuration.md deleted file mode 100644 index 72e0a1a4f..000000000 --- a/contracts/agents-api/execution-configuration.md +++ /dev/null @@ -1,81 +0,0 @@ -# Execution configuration queries - -This read-only administrator read describes configuration, not execution health. -It requires the Core key. They never contact a model -provider, start a Turn or wake a sandbox. The former project route -`GET /v1/agents/sessions/{session_id}/execution-configuration` is removed. A -Session read includes `agent.x_agents_core.harness` only when its Agent selected a -harness, inline or saved; Sessions on the deployment default keep the official -Agent shape, and no read returns the provider selection. - -## Frozen Session selections - -`GET /core/v1/projects/{project_id}/sessions/{session_id}/execution-configuration` -returns: - -```json -{ - "object": "agent.session.execution_configuration", - "schema_version": 1, - "session_id": "013773a9-44b9-4f84-baca-b51c04a01201", - "model": {"value": "requested-model", "source": "session"}, - "harness": {"value": "codex", "source": "agent"}, - "model_provider": { - "source": "agent", - "status": "available", - "configuration": { - "protocol": "responses", - "base_url": "https://model.example/v1", - "api_key_configured": true - } - } -} -``` - -`source` is `session`, `agent`, `deployment` or `unknown`. Model and harness -sources are independent; a model-only override can retain an Agent's harness and -whole provider bundle. Explicit inline harness selection is `session`; a null -inline Agent extension resets the harness to `deployment` without clearing the -inherited provider. A null Session provider inherits normally. The model may come -from the Session, a saved Agent or the deployment default; which Sessions may omit -it is owned by [model execution](model-execution.md#deployment-defaults). - -The provider's `status` describes visibility: - -- `available`: a safe view of the frozen bundle. Configuration contains protocol, - endpoint, configured-key flag and optional token limits. -- `available` with source `deployment`: the harness's deployment default as it was - frozen at creation. Deployment defaults are readable with the same Core key - through `/core/v1/harnesses`, so the safe view is recorded too. -- `redacted`: a deployment selection frozen before deployment defaults moved into - Core, from the retired operator options file. Source is `deployment` and - configuration is null; this read has no endpoint-reveal override. -- `unavailable`: no trustworthy safe provider projection was recorded. Source is - `unknown` and configuration is null. This does not mean that execution failed. - -New public Session creations save a minimal safe projection and provenance in the -same transaction as the Session and its private credential snapshot. Reads select -only that safe row and immutable Session configuration, without decrypting secrets -or re-resolving an Agent or deployment defaults. Agent edits/deletion, deployment -changes, Core restart and suspend/resume cannot change the view. Same-key retries -return the original Session and do not repair or overwrite its provenance. - -Historical Sessions without the projection report their persisted model/harness -when present, with `unknown` sources; a missing value is null. Their provider view -is `unavailable`, even if a private credential row exists. No backfill guesses -provenance. Missing, deleted and other Projects' Session IDs share the existing -not-found response. Responses are `Cache-Control: no-store`. Keys, ciphertext, -secret references, native headers, query parameters and permissions are excluded. - -## Discovery boundary - -Provider configuration discovery is not exposed, and the former Core startup -configuration read is removed. Provider inputs are still validated against internal adapter-owned rules and Core -admission policy. Removing discovery does not change the supported inputs or -create/update/execute behavior. - -This Session query is an administrator read, not an OpenAI Agents API -operation. Ordinary Agent and Session operations retain their pinned upstream -contracts. The query never mutates configuration, rotates keys, migrates Sessions, -exposes a model catalog or accepts arbitrary native options. See -[model-execution.md](model-execution.md) for write and inheritance semantics. diff --git a/contracts/agents-api/execution-tools.md b/contracts/agents-api/execution-tools.md index 991bb02af..6091f8175 100644 --- a/contracts/agents-api/execution-tools.md +++ b/contracts/agents-api/execution-tools.md @@ -5,7 +5,7 @@ An Agent declares application functions, controls and MCP servers in `tools`, an ## Admission - Saved Agents keep every pinned tool declaration as resource data. Saving never qualifies execution. -- Session creation resolves saved references and inline declarations with the execution parser into the immutable Session snapshot, then checks the combination against the selected Harness's profile in `services/core/internal/engine`. An unsupported combination returns 400 `unsupported_or_invalid_configuration` before anything is written. Protocol errors, such as a repeated `web_search` or `tool_search` or a non-object schema root, use the official error fields ([validation](official-semantics-alignment.md#agent-configuration-validation--september-23)). +- Session creation resolves saved references and inline declarations with the execution parser into the immutable Session snapshot, then checks the combination against the selected Harness's profile in `services/core/internal/engine`. An unsupported combination returns 400 `unsupported_or_invalid_configuration` before anything is written. Protocol errors, such as a repeated `web_search` or `tool_search` or a non-object schema root, use the official error fields ([validation](wire-semantics.md#configuration-validation)). - Before dispatch, the selected Runtime must also advertise the operation's capability. An advertisement alone never enables an operation. - The native Harness runs the model and tool loop. Core adds no second loop, output repair, schema coercion or prompt wrapper, and selects no native tool names. @@ -23,7 +23,7 @@ A function declaration requires `name`, `description` and a JSON Schema in `para | Unknown call, or a call of another Turn, in the caller's Session | 400 `invalid_request_error`; the pending action is unchanged | | Missing or foreign Session | 404 | -[Session input conflicts](official-semantics-alignment.md#session-input-conflicts-and-result-targets--september-23) records the exact messages. Invalid or unsupported content cannot consume a pending call. Admission is separate from application: the adapter confirms a result only when the matching native tool result appears in the live root Turn ([receipt contract](function-result-images.md)). A transport write alone confirms nothing, and a confirmation says nothing about provider consumption or exactly-once external effects. Core never replays a result automatically. +[Session input conflicts](sessions-events.md#input-errors) records the exact messages. Invalid or unsupported content cannot consume a pending call. Admission is separate from application: the adapter confirms a result only when the matching native tool result appears in the live root Turn ([receipt contract](message-content.md#function-results)). A transport write alone confirms nothing, and a confirmation says nothing about provider consumption or exactly-once external effects. Core never replays a result automatically. ### Required actions and recovery @@ -56,7 +56,7 @@ Core sends `PromptRequestPayload.ToolSearch` and each `FunctionTool.DeferLoading ] ``` -Saved Agents keep every pinned `web_search` mode: omitted or null is saved as `live`, and `cached` and `live` as sent ([saved modes](official-semantics-alignment.md#saved-web_search-modes--september-23)). Search settings are resource data: omitted or null `context_size` resolves to `medium`; omitted domains and location resolve to null; an empty domain list stays empty; a supplied location, including `{}`, has `city`, `country`, `region` and `timezone`, null where omitted. +Saved Agents keep every pinned `web_search` mode: omitted or null is saved as `live`, and `cached` and `live` as sent ([saved modes](wire-semantics.md#saved-configuration)). Search settings are resource data: omitted or null `context_size` resolves to `medium`; omitted domains and location resolve to null; an empty domain list stays empty; a supplied location, including `{}`, has `city`, `country`, `region` and `timezone`, null where omitted. Execution admits only `mode: "disabled"` and `enabled: false`. Enabled or omitted-mode search and enabled or omitted-`enabled` programmatic calling are rejected at Session admission unless the Session replaces the saved tools. Omitting programmatic configuration keeps each Harness's native behavior, which differs from the official default-on behavior. Unrelated native utility tools are not removed. @@ -85,7 +85,7 @@ Execution admits only `mode: "disabled"` and `enabled: false`. Enabled or omitte - [Public MCP connection origin](environments.md#public-mcp-connection-origin) owns origin defaults, placement and credential authority; [Harness capabilities](harness-capabilities.md#tools) owns per-Harness support. - Omitted or null `allowed_tools` permits every server tool; `[]` permits none. - `required: true` makes native thread creation and cold resume wait for the server to initialize; a failure stops execution without replacing retained history. It needs the Runtime's `mcp_http_required` capability. Public work can be accepted or queued during the wait. -- Bearer authentication uses an attached static or OAuth Vault credential. [Vault credentials](../../services/core/credentials.md) owns selection, and [MCP credential authority](environments.md#public-mcp-connection-origin) owns the frozen Runtime binding. Authenticated execution requires `mcp_http_bearer_auth`. +- Bearer authentication uses an attached static or OAuth Vault credential. [Vault credentials](vaults.md) owns selection, and [MCP credential authority](environments.md#public-mcp-connection-origin) owns the frozen Runtime binding. Authenticated execution requires `mcp_http_bearer_auth`. - The Runtime must advertise `mcp_http_tools`. The native Harness owns discovery, calls and results; public `mcp_call` Items use the original server and tool names and keep the observed native result. Codex verifies the exact effective MCP configuration before starting or resuming a thread, excludes undeclared servers, disables native apps and plugins, and rejects reserved native labels and stored native MCP credentials. Claude accepts labels of ASCII letters, digits, underscore and hyphen except `functions`, tool names that may also contain dots, and requires connected servers with static inventories. Anonymous Claude requests send a blank Authorization header to suppress native OAuth injection. Native OAuth login is not supported. diff --git a/contracts/agents-api/file-resource-semantics.md b/contracts/agents-api/file-resource-semantics.md deleted file mode 100644 index 550dc7384..000000000 --- a/contracts/agents-api/file-resource-semantics.md +++ /dev/null @@ -1,147 +0,0 @@ -# File and Skill resource semantics - -The fixed baseline remains SDK 3.13.0, commit `d7c41ef`, and `agents=v1` for Beta -resources. General Files and Skills do not require the Beta header. This batch -aligns specific resource operations; it does not establish full compatibility. - -## Official observations - -The 2026-09-23 probes used the fixed client and raw HTTP against the official API. -Evidence is retained under `~/.parsar/remediation/20260923/file-resource-semantics/` -in `official-files/` and `official-skills/`. They made 56 requests in total, owned -three tiny Files and two Skills with three ZIP uploads, and created no Sessions or -model calls. Every owned resource received a successful public delete response; -physical erasure was not independently observed. Requests were not automatically -retried. - -| Operation | Observed behavior and Core rule | -| --- | --- | -| Files upload/retrieve | `user_data` returns 200 with processed status, null expires_at and status_details; the existing Core shape matches | -| Files content | Both sampled `user_data` and `assistants` uploads reject direct download with 400, invalid_request_error, null code and param; Core applies this to its supported user_data uploads after tenant lookup | -| Files list purpose | Nine candidate values reach missing-cursor lookup; unknown and uppercase values reject first with 400 and param purpose | -| Skill content | Unversioned content follows default, even when latest differs; exact-version ZIPs retain their respective manifests and bytes | -| Skill descriptive metadata | Switching default changes top-level name and description, preserving ID and created_at; Core updates all three fields atomically | -| Skill default deletion | With two versions, deleting default rejects with 400, invalid_request_error, invalid_value and param version | -| Skill latest deletion | Nondefault latest deletion returns 200 and parent latest falls back to the surviving version | - -The qualified list values are `user_data`, `assistants`, `batch`, `fine-tune`, -`vision`, `evals`, `assistants_output`, `batch_output` and `fine-tune-results`. -These observations establish validation before a missing cursor, not successful -filtering for every purpose. Omitted and explicitly empty purpose also reached -cursor lookup. Core then retained its exact empty filter; the later [list query tolerance](list-query-semantics.md#list-query-tolerance--september-23-2026) -batch treats an explicit empty purpose as omitted, as a successful official page showed. Accepting a list filter does not enable -uploads, processing or jobs for that purpose. Core still uploads user_data only. -The fixed request/response `evals` union discrepancy is unchanged. - -The Files probe stopped its first phase when the fixed SDK raised BadRequestError -for content that the runner expected to download. Raw denials were preserved, -owned resources were cleaned, and an independent missing-cursor query phase used -the remaining request budget without recreating files. Positive official filtering -was therefore not observed. The extra official `detail` error member, complete -error messages, other purposes, expiration and Uploads remain separate gaps. - -The Skill probe observed an acknowledged version deletion followed about two -seconds later by a list and exact GET that still exposed that version, while the -parent latest pointer had already changed. It stopped and cleaned up. Core does -not emulate this inconsistent visibility. Fresh/reduced sole-version deletion was -not reached in this probe; the later [sole-version deletion](#sole-version-deletion--september-23-2026) -batch records that observation. -Upload `default:true` was not separately probed: updating descriptive metadata -there is the same pointer-consistency rule, covered by Core tests rather than a -new official wire claim. No newer schema or integer selector form was adopted. - -## Implementation and acceptance boundary - -Tenant authorization precedes public source download rejection. Missing and foreign -IDs retain indistinguishable 404 responses; the denial never opens source bytes. -Environment initialization and file_id copies still consume their authorized -internal source snapshot. Skill and Artifact downloads keep their existing rules. -Core Web no longer offers a source download action that the API rejects; the SDK -retains the official content operation and propagates its error. - -Skill metadata updates use the owning row lock and validated version metadata in -the existing transaction. A data-only migration corrects existing stale parent -metadata from the matching tenant and default version before the updated service -serves requests. It changes no identifiers, pointers, encrypted contents or Session -snapshots. No decrypt-on-read, cache, Runtime interface or lifecycle owner is added. -Nondefault uploads leave parent -metadata unchanged. Exact version bytes remain encrypted and immutable; previously -frozen Session metadata/content and retry intent remain independent of later source -changes. - -`official_file_resource_semantics.py` is the joint strict-client/raw-HTTP acceptance -against actual Core and isolated PostgreSQL. Existing source snapshot/copy and -Skill reference tests retain the internal read and frozen-input regressions. -The resource acceptance does not run daemon or model execution; historical native -qualification is separate. - -### Batch validation (2026-09-23) - -The final service source `be47803` passed targeted API and actual PostgreSQL -resource/migration/reference regressions (15.977 seconds for the Store package), -including strict SDK 3.13.0/raw HTTP. The data correction preserved version rows, -other resource fields and frozen Session configuration/setup bytes; reapplying it -did not rewrite correct rows. Reconstructed Store/handler reads validate persisted -data, not a full service process restart. - -All required `make check` targets passed: the server ran `make -o check-web check` -on `zju_a100_2`, paired with a fresh local `make check-web` on `13139f5`. The latter -adds only a browser-test focus synchronization to the service source. Client -287, Web unit 583, Core doctor 63 and browser 76 checks passed. sqlc generation -and byte comparison, OpenAPI generation, Go/build/native adapter/Rust gates passed. -The optional 512 MiB source streaming and packaged MiniMax scratch/large-output -profiles were not enabled. No new native model/Provider combination was qualified. - -Failures remain evidence: an initial Web type check caught a stale callback after -removing the download action; it was removed. One Chrome context setup timed out. -The existing Vault lifecycle browser test twice raced the dialog's scheduled -initial focus and filled the URL into Name, before any credential was created. -A bounded isolated run passed, but the full-suite repeat reproduced it. Reusing -the neighboring test's initial-focus wait corrected that test synchronization; -the final full Web gate passed without weakening assertions or changing credential -business behavior. Private logs and original artifacts remain under the evidence -root above. These results do not close the remaining protocol gaps. - -## Sole-version deletion — September 23, 2026 - -Evidence: campaign scan 1, `~/.parsar/remediation/20260923/campaign-scan-1/skills-files-templates/findings.json` -SFT-01 to SFT-04, with raw records in the adjacent `official-ledger.jsonl` (labels -`s1-*`, `s2-*`, `s3-*`). The scan owned three Skills and created no Sessions or -model calls. - -| # | Case | Official observation | Core rule | -| --- | --- | --- | --- | -| V1 | Delete the only remaining version, which is also the default | 200 `{"id": "skillver_…", "object": "skill.version.deleted", "deleted": true, "version": "1"}`; retrieve and versions.list then return 404 (SFT-01) | Same body; the Skill is deleted in the same transaction. The reduced case (delete v2, then v1 is the only version) applies the same rule; official evidence covers only a fresh single-version Skill | -| V2 | Delete the default while another version is visible | 400 invalid_request_error, invalid_value, param version (SFT-04) | Unchanged | -| V3 | Delete a nondefault or latest version | 200; latest falls back | Unchanged | -| V4 | Foreign or missing Skill or version | 404 | Unchanged, indistinguishable | - -`DeleteSkillVersion` keeps the owning Skill row lock. When the target is the -default, it deletes the Skill only if no other version row exists, through the -same cascade as `skills.delete`, so every encrypted version row is removed in the -same commit; otherwise the 400 remains. Uploads take the same lock: an upload -committed first makes the default undeletable, and a deletion committed first -makes the later upload return 404. Frozen Session installations keep their own -snapshot, and Templates keep their stored reference intent, exactly as after -`skills.delete`. No schema, query or numbering change is involved. - -Recorded decisions: - -- **SFT-02, number reuse: intentional difference.** After the latest nondefault - version 2 was deleted, the next official upload was numbered "2" again (one - sample, so max+1 and latest+1 are indistinguishable). Core keeps immutable, - monotonically increasing numbers because exact Template and Session selectors - reference numbers; reusing one could re-point a stored exact selector to other - bytes and make a frozen Session's concrete version ambiguous. A sole-version - deletion removes the Skill, so numbering never restarts within a Skill. -- **SFT-03, upstream anomaly: never emulated.** Deleting default version 1 about - four seconds after an acknowledged version 2 upload returned 200 and removed the - whole Skill, including version 2. The same request with version 2 visible - returned 400 (SFT-04). Core serializes both operations on the Skill row, so a - version deletion never removes an acknowledged upload. - -Core acceptance: real-PostgreSQL store tests prove atomic Skill removal without -orphaned version rows, unchanged frozen Session contents, creation retry and -Template intent, and both lock orders of a concurrent upload; a real HTTP test -covers V1 to V4 across two tenants, and `official_skills.py` checks V1 with the -pinned SDK and raw HTTP. None of these run a model. diff --git a/contracts/agents-api/function-result-images.md b/contracts/agents-api/function-result-images.md deleted file mode 100644 index a1c7af0b6..000000000 --- a/contracts/agents-api/function-result-images.md +++ /dev/null @@ -1,72 +0,0 @@ -# Function-result image coverage - -The pinned official function result accepts a string or ordered text/image content, -independently of success. Our qualified Claude subset is successful inline PNG/JPEG -on `environment:none` and Core-managed Docker `openai_hosted`. Error images, -unqualified placements and remote URLs reject before -batch persistence, without consuming the pending call. These are implementation -gaps, not narrower official types. MiniMax public functions remain unqualified. - -Core retains the original ordered output, error field presence and retry identity. -It passes the existing neutral `FunctionResultPayload` through Runtime. Profile -validation sees placement and success; adapters own native conversion. Only actual -image-result delivery requires `function_result_images` from the selected Runtime. -Ordinary function declarations and text results do not acquire an image requirement. -A Runtime refusal after durable admission fails execution without fabricating -application; the original result remains available for recovery queries. - -Claude uses native MCP text/image blocks. A live root tool result acknowledges the -once-only pending call only with matching Session/call identity, success, exact text, -block count/order and a native base64 image at every image position. Replay, -synthetic and subagent records cannot acknowledge it. Native image resizing or -re-encoding is allowed; this is incorporation into native history, not unchanged -bytes/pixels or a guarantee that the provider has already consumed the image. -Public Items preserve the caller's bytes. Native decode failure, text fallback or -missing images fails receipt validation. The existing uncertain-delivery timeout -and cancellation behavior remain unchanged; no replay mechanism is added. - -Codex waits for the live root `item/completed` dynamic-tool observation with -matching thread, Turn, call, function name, completion status, success flag and -exact ordered text/image content. A transport write alone does not acknowledge -application. This confirms the native handler's result, not completed provider -consumption or crash recovery. The adapter owns one pending receipt; native exit, -terminal settlement or cancellation releases it without confirming application. -A missing receipt times out after ten seconds and ends the uncertain execution; -there is no automatic result replay. The original submission remains durable. -The existing router retains receipt retry/conflict and terminal publication ownership. - -## Validation - -`TestNativeFunctionImagePublicExecution` and `tests/official_function_images.py` -exercise the same independent Core/PostgreSQL/daemon/native path with each selected -real model and the pinned SDK 3.13.0 plus raw HTTP. The workflow covers mixed -text/PNG/text, a 6000x2100 PNG requiring native preprocessing, image-only JPEG, -failed text, pending-call cancellation and cold daemon continuation with unchanged -native Session identity. The real answer must read visual information absent from -the tool description and text content. Recovery reads must retain original content. -Retries admit one result; changed retries conflict; foreign tenants cannot read or -submit it. Claude invalid/remote/error image batches leave the call and history -untouched, then a valid result succeeds on that same pending call. - -Direct native feasibility probes separately establish Claude's successful image -preprocessing and lossy error-image path. They do not replace public acceptance. -Controlled tests cover malformed/missing/reordered receipts, wrong identity, -unsupported placement, operation-specific Runtime support and batch atomicity. -The bundled JPEG fixture has yellow, blue, red and green vertical bands; it contains -no metadata or credentials. PNG markers are generated with randomized band order. - -The [Docker workspace workflow](message-input.md#docker-workspace-acceptance) -uses the same function-result contract alongside native file tools, public -Files/Artifacts and cold Core/Runtime continuation. It checks Claude's rejected -error/remote result directly on an outstanding call before accepting a valid -image on that same call; no mixed-message rejection substitutes for this check. - -Run evidence is retained outside the repository. This -coverage does not qualify all native image limits, provider -parity, arbitrary managed output rewrites, crash recovery or full Agents API -compatibility. No downloader, image converter or second tool loop belongs in Core. - -User-managed Linux image execution follows the same native workspace path. See -[current qualification](harness-capabilities.md) for the tested -Harnesses, formats, continuation and remaining boundaries. Historical evidence -above retains its original deployment scope. diff --git a/contracts/agents-api/harness-capabilities.md b/contracts/agents-api/harness-capabilities.md index fd0350449..03e333d21 100644 --- a/contracts/agents-api/harness-capabilities.md +++ b/contracts/agents-api/harness-capabilities.md @@ -16,12 +16,12 @@ Placements are `none` (no Environment), hosted (`openai_hosted`) and self-hosted | --- | --- | --- | --- | | Text Turns, active input, cancellation, restart and continuation | Verified on all placements | Verified on all placements | Verified on all placements | | [Files and Artifacts](environment-files.md) | Verified: hosted, self-hosted | Verified: hosted, self-hosted | Verified: hosted, self-hosted | -| [Whitespace-only message text](message-input.md) | Admitted; delivered unchanged | Rejected | Rejected | -| [Inline PNG and JPEG message images](message-input.md) | Verified on all placements | Verified on all placements | Rejected | +| [Whitespace-only message text](message-content.md) | Admitted; delivered unchanged | Rejected | Rejected | +| [Inline PNG and JPEG message images](message-content.md) | Verified on all placements | Verified on all placements | Rejected | | Remote image URLs | Rejected | Rejected | Rejected | | Explicit `reasoning`; `service_tier` other than `auto` | Rejected | Rejected | Rejected | | `text.verbosity` other than `medium` | Admitted; the native model decides | Rejected | Rejected | -| [Public token usage](history-events-usage.md) | Measured counters | Null | Null | +| [Public token usage](sessions-events.md) | Measured counters | Null | Null | Native model parameters and provider protocols per Harness are in [model execution](model-execution.md). @@ -30,7 +30,7 @@ Native model parameters and provider protocols per Harness are in [model executi | Operation | Codex | Claude SDK | MiniMax Code | | --- | --- | --- | --- | | [Public functions](execution-tools.md#functions) with text results | Verified: `none`, hosted; admitted: self-hosted | Verified on all placements; object-root schemas only | Rejected | -| [Function results with images](function-result-images.md) | Verified: `none`, hosted; admitted: self-hosted | Verified on all placements; successful inline PNG or JPEG results only | Rejected | +| [Function results with images](message-content.md#function-results) | Verified: `none`, hosted; admitted: self-hosted | Verified on all placements; successful inline PNG or JPEG results only | Rejected | | [Structured output](execution-tools.md#structured-output) | Rejected | Verified on all placements | Rejected | | [Deferred function discovery](execution-tools.md#deferred-function-discovery) | Rejected | Verified: `none`, self-hosted; admitted: hosted | Rejected | | [Disabled web search and programmatic tool calling](execution-tools.md#web-search-and-programmatic-tool-calling) | Verified: `none`; admitted: hosted, self-hosted | Verified: `none`; admitted: hosted, self-hosted | Verified: `none`; admitted: hosted, self-hosted | diff --git a/contracts/agents-api/harness-onboarding.md b/contracts/agents-api/harness-onboarding.md index 1ad908c0c..34f94ff68 100644 --- a/contracts/agents-api/harness-onboarding.md +++ b/contracts/agents-api/harness-onboarding.md @@ -129,7 +129,7 @@ The public text path requires durable Turns, applied input receipts, ordered obs MCP, public functions, deferred function discovery, structured output, image input, verbosity controls and other optional operations need not match another engine. Reject an unqualified combination with Unsupported and record the gap; never advertise a capability to bypass selection. - Structured output: consume `ExecutionControls.OutputFormat` and publish confirmed native output through the Message contract ([execution tools](execution-tools.md#structured-output)). Register the public qualification separately from the Runtime capability. -- Images: register the Runtime's `MessageImages` and qualify the profile's `MessageImages` separately ([message input](message-input.md)). +- Images: register the Runtime's `MessageImages` and qualify the profile's `MessageImages` separately ([message input](message-content.md)). - Workspace placements additionally need verified preparation, workspace reads and output export and the dedicated Runtime binding with the shared Files helpers. Enable a placement only after its lifecycle behavior is demonstrated. ### MCP origin and native limits diff --git a/contracts/agents-api/history-events-usage.md b/contracts/agents-api/history-events-usage.md deleted file mode 100644 index 79c36d6c5..000000000 --- a/contracts/agents-api/history-events-usage.md +++ /dev/null @@ -1,374 +0,0 @@ -# History, events and usage - -This milestone covers existing execution resources and client recovery. It does -not establish complete Agents API compatibility. The protocol baseline remains -the SDK and source recorded in [upstream.json](upstream.json). - -## Query and live-stream contract - -Open the live Session events stream before submitting input. After disconnection, -subscribe again and retrieve Session, Turns and Items to recover persisted state. -An SSE connection is an observer, not the owner of execution. `Last-Event-ID` does -not introduce replay. Query responses and transition events are observations at -different times; a completed Turn can acquire a measured usage snapshot later. -Do not permanently cache `usage: null` as zero or as a final accounting result. - -The TypeScript client shares Turn and Item projections between reads and SSE. -It preserves root/child identity and supports the currently implemented reasoning -and coordination Items. `agent_message` has no status; reasoning status can be -absent or null. A child Turn's `agent_id` is the Session's Agent ID and its -`subagent_id` names the child. The Session stream carries root work only: child -Turns and child Items publish no Session events, while `agent.session.subagent.*` -events and root coordination Items remain. Earlier Core releases streamed child -Turns, which could first appear as an already terminal `turn.created` snapshot; -the client still accepts that and must not reinterpret it as newly queued work. -This does not widen the server's Item or interim reasoning-event coverage. -Claude and MiniMax retain their settlement-based child-history boundary. Recover -child state through the Subagent queries. - -Use response cursors to page history in the requested direction. Session Turn -lists contain root Turns only; a child Turn ID on the Session Turn routes is not -found. Root Items and child Items have separate query resources; use the Subagent -resources for child Turns and history. See -[Subagent visibility](subagents.md#subagent-visibility). -Tenant ownership is enforced by Core for both queries and streams. - -## Measurement boundary - -Adapters publish cumulative measurements for the current execution through the -existing neutral Usage contract. Core replaces a Turn's snapshot atomically; -repeated snapshots, including terminal repeats, do not add consumption. The -Session total sums recorded root-Turn measurements, not Subagent Turn lists. It -is null while any root Turn has not ended and once any root Turn ends with -unknown usage ([item serialization](#item-serialization-2026-09-23)). -It is best-effort accounting, not an invoice or an estimate of missing work. - -- Codex publishes observed snapshots while the Turn is active. Exact native - resume excludes the previous Session total from the new Turn. A complete public - breakdown requires valid input, cached, output, reasoning and total counters. - A thread total that has not advanced past the Turn's baseline measures nothing - for that Turn, so a Turn interrupted before any response reported usage stays - null rather than zero. -- Claude retains native result evidence internally. Its per-turn and cumulative - query/model counters have different scopes, and reasoning attribution can be - incomplete. Complete public TokenUsage remains unqualified and null. -- MiniMax ACP context occupancy describes context capacity, not measured resource - consumption. Complete public TokenUsage remains unqualified and null. - -Persisted snapshots survive cancellation and Core worker restart. This does not -recover observations lost before journal commit, stop external tool side effects, -or replay interrupted input. Provider/model reroute attribution, billed costs, -unknown historical counters and unreported partial usage remain outside this -milestone. Native differences must not be hidden with guessed zero counters. - -## Core acceptance, 2026-09-22 - -Three independent PostgreSQL/Core deployments used Docker V1 Runtime images with -the integrated daemon, fixed official SDK 3.13.0, raw HTTP and the final TypeScript -client. Codex and Claude called Kimi K3; MiniMax Code called MiniMax M2.7. Each -performed real native child delegation and an ordinary root Turn. All three passed -history identity/content agreement, both paging directions at limits one and two, -SSE disconnect/reconnect, wrong-tenant denial and stable history after idle Core -restart without another Turn. Captured coordination events passed the TS parser. -Claude's first run had coordination events but no child Turn SSE snapshot; the -final run also captured and parsed child Turn snapshots. Both observations retain -the native publication boundary above, without a continuous-progress guarantee. - -Codex additionally exposed measured usage while a tool-separated Turn was still -in progress, then retained it after cancellation, repeated reads and Core restart. -Controlled adapter tests cover ordered publication and cancellation under output -backpressure. PostgreSQL tests cover cumulative replacement, duplicate snapshots -and known usage retained across worker reconciliation after process loss. The -idle real restart does not substitute for active native crash qualification. - -Validation included the complete standalone gate split between server -`make -o check-web check` and local `make check-web`, Codex adapter race tests, -286 client tests, 573 Core Web tests and 74 fixture Playwright tests. SQL generation -matched byte-for-byte. No handler annotation, schema, DB query or migration changed. -Real proof and first-failure records are retained privately under -`~/.parsar/remediation/20260922/history-events-usage/live/`. The first Claude probe -incorrectly required a live child Turn event; the bounded follow-up checked the -documented coordination stream and child queries. Invalid historical patch fixtures -were corrected to separate function calls from their result Items, without widening -production validation. E2B and unrelated product flows were not rerun. - -## Official-service observation, 2026-09-22 - -A bounded probe used the fixed Python SDK 3.13.0, raw HTTP, `agents=v1`, and one -owned `environment:none` Session with two tiny real-model text Turns. It did not -read unrelated resources or use tools, workspace execution or Subagents. The -owned Session was deleted after the probe. - -- Creation returned HTTP 201 SSE. Input submission returned HTTP 202. -- First-turn events followed queued/in-progress, Item/text output and terminal - transitions. Its terminal usage was null. Later Session and Turn reads reported - input 7226, cached 0, output 7, reasoning 0, total 7233. -- Reconnection with an old `Last-Event-ID` produced no historical frames during - three idle seconds, then delivered new second-Turn frames. -- After closing the second observer stream, queries recovered both completed - Turns and all four Items. Ascending and descending limit-one pages reversed - exactly, with exclusive sampled cursors. The Turn completed before the client - disconnected, so this alone does not prove survival of an active disconnection. -- Session usage after the second Turn was null, as was that Turn's usage in the - observed pages. No later aggregation behavior was inferred. Twenty-two captured - events and sampled resources passed strict pinned-type validation. - -These are bounded observations, not complete timing or accounting guarantees. -Same-timestamp paging, failures/cancellation and Subagent accounting were not -probed against the official service. The current documentation describes an -Items `turn_id` filter absent from the pinned SDK and requires initial input for -`none` where the pin describes it as optional. Neither difference changes the -repository baseline; the probe supplied initial input and sent no unpinned filter. - -Sources: [Sessions](https://developers.openai.com/api/docs/guides/agents-api/sessions), -[Events](https://developers.openai.com/api/docs/guides/agents-api/sessions/events), -[Observability](https://developers.openai.com/api/docs/guides/agents-api/observability). -Sanitized request evidence is retained privately under -`~/.parsar/remediation/20260922/history-events-usage/official/`; credentials are -excluded from source and evidence. - -## Creation stream settlement, 2026-09-23 - -Evidence: the second official-semantics campaign scan compared four owned -`environment:none` official Sessions with Core (private -`~/.parsar/remediation/20260923/campaign-scan-2/events-tools/`, `findings.json` -EVT-01..24 with raw frames under `official/`). All 1091 official events passed -strict validation against the pinned types. Two independent official creation -streams (structured output, and a function call with its result) were closed by -the server; they, a third creation stream that the client closed while a call -was pending, and the 2026-09-22 probe above agree on the snapshot, order and -terminal-event observations below. This batch changes only the four Core-owned -stream differences EVT-01..04; the plan is -`~/.parsar/remediation/20260923/creation-stream-settlement/PLAN.md`. - -- **Creation stream lifetime (EVT-01).** The official service closed both creation - streams right after the first `agent.session.idle`, without `[DONE]` or an error - frame, and kept the function stream open through `requires_action` and the - result. A fresh Core creation stream ends right after the first - `agent.session.idle` recorded when a Turn ends or an input reservation stops - being pending (expired, cancelled or failed), or any `agent.session.failed`, - and never sends the events after it. A self-hosted connection that clears - pending input to idle does not end it. A creation that admitted nothing ends - right after `created`. Settlements that record no event, such as a reservation - cancelled while its Session is already idle, use a fallback: after an empty - drain the stream reads the JSON-path projection and the event cursor in one - database snapshot and, if settled, sends only events up to that cursor, then - ends. The settled marker and pending-input flag are Store-internal and add no - events. A same-key `stream=true` retry of an existing creation returns 201 with - only the connection comment and ends at once; official same-key requests create - distinct Sessions, so there is no retry stream to follow, and recovery uses - `stream=false` or GET. The TypeScript client reports that empty creation stream - as `CreationStreamRetryError`. Accepted follow-ups: another client's work - drained before the fallback read can still be sent after a silent settlement; - idles recorded by an older binary during a rolling deploy carry no settled - marker and rely on the fallback; and an input reservation made while the ending - Turn captured Artifacts can start a later Turn that the stream does not follow. - Four review rounds replaced event-only, projection-only and retry-following - designs before merge. GET event streams are unchanged: live-only, no replay, - and they never end on their own. Session deletion still ends both. -- **Created snapshot (EVT-02).** Official `agent.session.created` carried the - post-admission Session (`in_progress`, no actions, null usage), like the JSON - 201 body. Core now sends the committed projection that JSON 201 returns, read - after the creation commit, while the stream still starts at the creation - upsert cursor, so every initial Turn and Item event follows exactly once. The - snapshot is read after the commit, so it can already show a later state than - the events that follow it; the JSON 201 body has the same race. For - self-hosted input the snapshot already requests the Environment connection and - the committed `requires_action` event follows. Hosted initial input remains - `idle` while it provisions. -- **Terminal usage (EVT-03).** Official `agent.session.turn.completed` and - `.cancelled` (8/8) carried a top-level `usage`, null at emission even when later - reads were measured. Core terminal Turn events (`completed`, `failed`, - `cancelled`) now carry `usage` copied from the rendered Turn snapshot, with - explicit null when unknown; other events omit it. This batch applied it to - root and child Turn events; the - [Subagent visibility batch](subagents.md#subagent-visibility) - later stopped publishing child Turn events on the Session stream. Codex can - therefore publish measured counters at settlement, while Claude and MiniMax - stay null; no counter is derived or summed. The TypeScript client accepts the - field on terminal Turn events only and still accepts older events without it. -- **Turn start order (EVT-04).** Official new Turns published `turn.created`, the - user `item.added` (`output_index` null), `agent.session.in_progress`, then - `turn.in_progress`. Core now records the Session activity after the admitting - input's Items in the same transaction, for creation, events.create and - reservation promotion. When one batch holds several message events, later - messages follow that activity; their official order was not observed. - -Deferred, with evidence retained in `findings.json`: - -- EVT-05: Core emits `turn.in_progress` when a function result resumes a waiting - Turn and publishes the result Item at native application; the latter is the - known INTERACTION-PUBLICATION-001 receipt boundary. -- EVT-06: Core emits an interim `agent.session.in_progress` when cancelling a - Turn that waits on a function result. Statuses match. -- EVT-07: official mid-Turn attach sent catch-up Item snapshots, but not - deterministically; one more official sample is needed before designing. -- EVT-08: official Items omit in-progress and incomplete output Items; Core keeps - them under native history ownership. -- EVT-09 and EVT-10: null-valued `output_index`, `phase` and `error` fields and - the initial assistant content were since aligned by the - [item serialization batch](#item-serialization-2026-09-23). -- EVT-11 and EVT-12: unknown call/Turn result and conflict error codes were - since aligned by the [input conflict batch](official-semantics-alignment.md#session-input-conflicts-and-result-targets--september-23). -- EVT-13: official Session usage became null when any root Turn usage was - unknown; Core summed the known Turns. The - [item serialization batch](#item-serialization-2026-09-23) adopted the official - rule, including the non-terminal case later observed as ST-03. -- EVT-19: the Core terminal sequence for a Turn cancelled mid-text is recorded by - this batch's live acceptance, not changed by it. - -Batch validation used targeted Go API, store and contract tests on a dedicated -PostgreSQL database, including the pinned Python SDK 3.13.0 creation-stream, -initial-input, self-hosted and initial-failure scripts, plus the TypeScript -client and Core Web unit tests. Resource-level replay, live model acceptance and -the full gate are recorded separately; retry, self-hosted, hosted and no-input -creation stream lifetimes have no official observation. - -## Item serialization, 2026-09-23 - -Evidence: the second campaign scan's owned official `environment:none` Sessions -(private `~/.parsar/remediation/20260923/campaign-scan-2/events-tools/`, -`findings.json` EVT-09, EVT-10 and EVT-13, raw frames `official/streams-s1..s4.json` -and Items pages `official/calls-s2.json` `s2-items-after-t1`, `calls-s4.json` -`s4-items`), the first scan's SES-23 and SES-25 (`campaign-scan-1/sessions/`) and -VA-11 (`campaign-scan-1/vaults-agents/`), the fifth scan's ST-03 -(`campaign-scan-5/sessions-turns/findings.json`, raw `official/calls.json`), and -the live-kit observation EVT-24 -(`creation-stream-settlement/acceptance/candidate-evidence/codex-kimi/attempt-1/`, -`r1-events.json`, `r1-reads.json`, `daemon.log`). The plan is -`~/.parsar/remediation/20260923/item-serialization/PLAN.md`. This batch changes -field presence, event framing and the Session usage rule only; it never alters -model output text and never fills a model-derived default or counter. - -- **Null output index (S1, EVT-09).** Official `item.added` events for input Items - (user messages, function results) carried `output_index: null`. Core Item events - (`item.added`, `item.done`) now always carry `output_index`, null for input - Items; Session and Turn events keep omitting it. -- **Message phase (S2, SES-25, EVT-09).** Official message Items always carried - `phase`, null for user messages. Core messages in Items pages and events now - carry `phase`: the native phase when the adapter reports one (Codex message - observations, Claude structured output), otherwise null. Native phase values are - unchanged. -- **Function result fields (S3, EVT-09).** Official `function_call_output` Items - always carried both `output` and `error` (`error` null when not submitted; every - sampled submission had an output). Core Items and events - now always carry both, null when the submission omitted them. The stored payload - and the saved submission keep the submitted presence; the wire no longer - distinguishes an omitted field from null. The official Items page reported a - failed result's `output` as null although its event carried the submitted array; - Core keeps the submitted content in both. -- **Assistant message sequence (S4, EVT-10).** Official streams added an assistant - message in progress with `content: []`, then `content_part.added` with empty - text, deltas, `output_text.done`, `content_part.done` and `item.done`. Core no - longer pre-fills the part in `item.added`: it sends the same sequence, with the - Item in progress and empty content. The stored Item and later reads are - unchanged. Clients that append a part on `content_part.added` no longer see a - duplicate. -- **Non-streamed finals (S5, EVT-10).** Official structured output streamed like - any message. A Core message first observed complete, such as Claude structured - output or a legacy final answer, now follows the same sequence with its full - text in exactly one `output_text.delta`. The delta is the native text byte for - byte; this is event framing only. -- **Reasoning keys (S6, SES-23, VA-11).** Official Agent and Session responses - always carried `reasoning.effort` and `reasoning.summary`. Core saved Agent, - Session, Session list and Session event responses now serialize both keys, null - when unset, instead of `{}`. Official responses fill the model-derived default - effort (for example `medium`); Core does not, which remains a documented native - difference. Requests and stored configuration keep their encoding, so creation - retry identity is unchanged. -- **Session usage (S7, EVT-13, ST-03).** Official Session usage stayed null after - a root Turn ended with unknown usage, and was the exact sum when every Turn was - known (EVT-13). In three more samples it was null in every read while a root - Turn was in progress or waiting for a function result, even after earlier - Turns were measured, and returned to the sum once every Turn had settled with - known usage (ST-03: retrieve, list and the metadata update response). Core - Session usage (retrieve, list, update and Session event snapshots) is now the - sum of recorded root Turn usage only when every root Turn has ended (completed, - failed or cancelled) with known usage; otherwise it is null. A later measured - Turn does not restore the sum after an unmeasured one. The official samples did - not read a queued Turn; Core treats queued Turns like active ones, since their - consumption is not yet known. A root Turn cancelled while still queued ends - without usage, so public Session usage stays null afterwards. That follows - from the terminal rule; ST-03 has no official sample of the case, so it is an - inference. A usage snapshot that an active Codex Turn - records stays readable on that Turn but does not count in the Session until the - Turn ends. Claude and MiniMax Turns remain unmeasured, so their Sessions stay - null. Official reads also lagged settlement by seconds; Core does not copy that - timing. -- **Measured telemetry usage (Core extension).** Runtime observation telemetry, - runtime history token points and their OTLP export are Core extensions with no - official counterpart. They read a separate internal measured usage: the sum of - every recorded root Turn snapshot, active Turns included, null only when - nothing is recorded. It differs from public Session usage by design, so the - token series stays continuous while Turns run and after an unmeasured Turn. - Public Session usage (retrieve, list, update and every Session event snapshot, - including the function-action and Environment-input snapshots) keeps the rule - above. Core Web reads public Session usage. Its Runtime summary's reported - token total keeps a listed Session's last reported total while the usage is - null, so it does not drop; rows still show the current public value. Its live - token trend keeps that Session in the series but treats the held total as - unknown: those intervals are gaps, never zero, and the rate once usage is - reported again is spread over the time since its last report. -- **Cancelled Codex usage (S8, EVT-24).** The all-zero counters on a cancelled - Codex Turn came from the Codex adapter, not from Core or storage. The Turn was - the second of its Session, on a resumed native thread, and was cancelled - mid-text. Native reports a cumulative thread total, and the adapter publishes - its difference from the Turn's baseline. The daemon forwarded no usage frame - for the Turn until one right after the cancel, whose total had not advanced - past the baseline. The adapter published the difference, all zeros, as a - complete measurement, the cancellation outcome repeated it, and Core stored - what it received. Native did not report zero usage for the Turn, it reported no - new usage. The adapter now ignores a total that has not advanced past the - Turn's baseline, so the Turn stays null, as official terminal usage does. A - total that later advances is still published. An explicit native zero in a - per-Turn usage payload would still be kept. - -Unchanged: the resume/cancel event sets (EVT-05/06), result publication timing -(INTERACTION-PUBLICATION-001), reconnect catch-up (EVT-07), Items list contents -(EVT-08), native phase values, Core-only extension fields and the stored -payloads. The TypeScript client accepts `output_index: null`, `phase: null` and -the explicit function result nulls, and still accepts older Cores that omit them; -Core Web accepts a null message phase. - -Batch validation used Go contract tests for the wire shapes, API tests for the -Session and stream rendering, real-PostgreSQL store tests for the event sequences, -stored-payload presence and the Session usage rule, the Codex adapter usage tests -(with `-race`), the pinned-SDK official client suite against a local server and -the TypeScript client and Web unit tests. `make openapi` adds only `x-nullable` to -Item `phase` and event `output_index`; `make sqlc-generate` changes -`SessionTokenUsage` and adds the internal `SessionMeasuredTokenUsage`. The native pinned-SDK scripts updated for these shapes run -only with a native daemon, and live model acceptance is recorded separately. - -## Hosted initialization failure events, 2026-09-23 - -Evidence: campaign scan 6 HI-01..04 (private -`~/.parsar/remediation/20260923/campaign-scan-6/hosted-init/`, raw frames -`official/007-S2-events.json` and `009-S3-events.json`); rows H1–H8 are in -[official semantics](official-semantics-alignment.md#hosted-initialization-failure--september-23). - -- **Order.** A hosted Environment that fails to provision records - `agent.session.environment.failed`, `error` and `agent.session.failed` in one - transaction, as officially observed. The official streams showed no - `environment.pending` event; Core records none either. -- **Payloads.** `environment.error` is `{type: environment_error, code: - environment_connection_failed, message: "The environment failed to connect."}`. - The `error` event carries the pinned `SessionError`: `{type: environment_error, - code: sandbox_error, message: , param: null}`. Core's own - `stream_interrupted` frame keeps its three-field error without `param`, which - released clients validate exactly. The `agent.session.failed` snapshot has `status: failed`, - the reason as `error`, `required_actions: []` and the failure time as - `last_active_at`, identical to later retrieve and list reads. Pending input - settled by the failure is captured in the same snapshot. -- **Stream lifetime.** GET and creation streams end right after that - `agent.session.failed`, as the official GET stream did. This changes the - earlier rule that GET streams never end on their own, for this terminal case - only: a Turn failure leaves GET streams open because the Session can continue, - and a GET stream opened after the failure stays open (not observed officially). -- **Client.** The TypeScript client still raises `stream_interrupted` as an - `AgentCoreError`, and now delivers other `error` events to `onEvent` as - `AgentSessionErrorEvent`, before the failed snapshot. It accepts an optional - nullable `param` on stream errors. Core Web renders the failed Session and its - error from the snapshot and ignores the error event. - -Environments that failed before migration `000062` have no recorded reason; they -keep their earlier projection and events. diff --git a/contracts/agents-api/installation.md b/contracts/agents-api/installation.md deleted file mode 100644 index fd2042a97..000000000 --- a/contracts/agents-api/installation.md +++ /dev/null @@ -1,43 +0,0 @@ -# Installation facts - -`GET /core/v1/installation` reports what an administrator needs to call and -change this installation. Core key only, like every `/core/v1` route; Web forwards -it after sign-in. It is available before any sandbox deployment exists and makes no -provider or model call. See the [Core OpenAPI](core.openapi.yaml) for the schema. - -| Field | Source | -| --- | --- | -| `object` | Always `core.installation` | -| `installation_id` | `OAC_INSTALLATION_ID`; null when Core runs without the sandbox manager | -| `public_url` | `OAC_PUBLIC_URL`: the origin applications, nodes, sandbox guests and self-hosted executors use. Null when unset | -| `api_base_url` | `public_url` followed by `/v1`: the base URL for Project API keys (`OPENAI_BASE_URL`). Null when `public_url` is null | -| `local_only` | True when `public_url` names a loopback host, which only the Core host reaches | -| `source_commit` | The full source commit Core was built from; null for development builds | -| `configuration` | The installer's settings snapshot (`OAC_SETTINGS_FILE`); null when the installer did not start Core | -| `address_bindings` | What a change of `public_url` affects, counted on each read | - -`configuration` has: - -- `path`: the absolute host path of the installation's `config.json`, where every - process setting is changed (by default `~/.oac/core/config.json`); -- `apply_command`: the command that applies `config.json` changes, by default - `~/.oac/core/oac apply`; -- `applied_at`: when the snapshot was last applied; -- `settings`: one item per `config.json` setting, with its dotted `key`, applied - `value`, `default`, whether it is `changeable` after installation, whether it is - `sensitive`, and the services it `restarts` (`core`, `web`, `database`). - -A sensitive setting always has a null `value` and `default`, and a boolean -`configured` instead; only sensitive settings have `configured`. Core refuses to -start when the snapshot breaks this rule, has a duplicate key or has an unknown -member. Core reports the snapshot and never acts on it; runtime settings, such as -the sandbox deployment, have their own routes. - -`address_bindings` has: - -| Field | Meaning | -| --- | --- | -| `nodes` | Enrolled nodes that are not removed | -| `nodes_on_other_address` | Nodes enrolled with an address other than `public_url`. They receive no new sandboxes; remove and add them again. At most `nodes` | -| `hosted_sandboxes` | Retained and pending hosted sandboxes, started with the address current at the time | -| `self_hosted_executors` | Unrevoked executor credentials; their executors were installed with the advertised `remote_url` | diff --git a/contracts/agents-api/list-query-semantics.md b/contracts/agents-api/list-query-semantics.md deleted file mode 100644 index 69363ac85..000000000 --- a/contracts/agents-api/list-query-semantics.md +++ /dev/null @@ -1,271 +0,0 @@ -# List query semantics — September 23, 2026 - -The fixed target remains openai-python 3.13.0, commit -`d7c41efee1b0802b79f3f88a678ef2052b06e9ce`, with `agents=v1` on beta resources. -This batch starts from Core main `c86b5bb` and addresses list order validation and -the error fields actually observed for those requests. It does not change page -limits, cursor lookup order, execution or native harness behavior. - -## Official observations - -Exactly 35 new read-only official GET requests were made, with no retries, resource -writes or model execution. Nested requests used random missing Session/Vault IDs; -top-level requests used missing cursors or a nonexistent Agent filter. Every -successful list was empty. Evidence, request IDs, fixed type copies and the full -matrix are retained under -`~/.parsar/remediation/20260923/error-query-survey/`. Earlier owned Session/Turn -observations are separately identified in that directory, not counted as new calls. - -The nine sampled collections rejected explicit `order=`. The fixed enum permits -only `asc` or `desc`; omission and an explicitly empty string are different inputs. - -| Measured request | Status | Error type | Code | Param | -| --- | --- | --- | --- | --- | -| Empty order on Agents, Sessions, Turns, Items, Templates, Vaults, Credentials | 400 | invalid_request_error | invalid_request_error | null | -| Empty order on Files | 400 | invalid_request_error | null | null | -| Empty order on Skills | 400 | invalid_request_error | invalid_value | order | -| Invalid status on Vaults/Credentials | 400 | invalid_request_error | invalid_request_error | null | - -These are operation-specific observations. Do not apply the Skills code to all -validation errors or make every Files error share one parameter. Other endpoints -using the same order enum may share the parser, but were not independently probed -here. Authentication and ownership checks retain their existing boundaries. - -The fixed Python SDK's query serializer drops empty string values. Consequently, -`list(order="")` does not send `order=` and follows omission/default behavior. -Explicit-empty rejection requires raw HTTP evidence; nonempty invalid SDK order -values exercise the rejected path normally. Do not alter Core or the SDK to hide -that request-serialization distinction. - -## Deferred differences and uncertainty - -- Query limits require separate qualification. Missing cursors or parents can mask - a later numeric validation error. A successful filtered-empty request does not - establish useful zero-sized pagination or an actual maximum page capacity. -- An owned official Item list accepted 101 although its pinned type documents - 1–100; an owned Turn list rejected 101. Keep the fixed range until the conflict - is resolved explicitly. No general range change follows from this survey. -- Vault/Credential negative limits rejected on the official service, while the - pinned description broadly says values clamp to 1–100. Core's clamping policy - is unchanged in this batch. -- Files invalid purpose and missing cursor expose different `param` values; - purpose typing and cursor behavior need a separate bounded decision. The - verbose Files validation message and additional `detail` object are not a - reason to recreate an upstream validation framework. -- Two malformed Skills cursor observations returned 500, while a shaped missing - cursor returned 404. Retain the evidence; do not deliberately reproduce an - upstream failure or weaken safe missing-resource handling. -- Unknown/repeated query keys, whitespace cursors, numeric overflow, all status - combinations and concurrent-page mutation were not newly qualified. - -Current documentation was consulted alongside the pin, including the -[Session list reference](https://developers.openai.com/api/reference/go/resources/beta/subresources/agents/subresources/sessions/methods/list) -and [Vault list reference](https://developers.openai.com/api/reference/typescript/resources/beta/subresources/agents/subresources/vaults/methods/list). -New documentation and sampled tolerance do not silently replace the fixed SDK. - -## Acceptance boundary - -Acceptance must exercise fixed-SDK and raw HTTP requests against the actual Core -service and a dedicated PostgreSQL database: rejected queries, omitted/valid order, -scoped history and pagination, authentication, tenant isolation and no mutation -from rejected GETs. Controlled tests support the same contract. This read-only -change requires no new native-model capability qualification; existing real-model -evidence retains its original scope. Record completed checks and independent review -before merging; neither a deserializable response nor a route inventory proves -complete compatibility. - -## Completed acceptance - -The nine-family fixed-SDK/raw-HTTP regression passed against actual Core HTTP, -PostgreSQL and Worker admission with dispatch paused. It checked 35 primary -rejections, nine SDK empty-query omission cases, successful ascending/descending -and default pagination, authentication and foreign-tenant masking. Resource -snapshots and the original three cancelled Turns/three Items remained unchanged. -The baseline failed at raw empty-order admission; the implementation passed. -This is resource/query acceptance, not native or model execution. Logs, exact -SDK serializer source and hashes are retained in the survey directory's -`ACCEPTANCE.md`; its owned database was removed. - -Server `make -o check-web check` passed at `e4cb8a5`, including the dedicated -PostgreSQL regression, sqlc regeneration, service/adapter Go tests and builds, -Claude/MiniMax checks and Rust tests/format/Clippy. Web/client source and dependencies -are unchanged from PR #35: its 287 client tests, 583 Web tests and 76 browser cases -remain the applicable exact-source evidence; no new Web run is claimed. The -optional packaged MiniMax native-tools probe was skipped. This batch changes no -Runtime/provider/model behavior and does not requalify their combinations. - -A fresh independent GPT-6 Astra high reviewer found no grounded in-scope findings -after inspecting the full diff and evidence; API tests and `git diff --check` -passed independently. Server logs are retained under -`~/.parsar/remediation/20260923/list-query-alignment/`. The remaining limits, cursor, -lookup and Files verbose error differences above are not declared compatible. - -## List query tolerance — September 23, 2026 - -The pin is unchanged: SDK 3.13.0, commit `d7c41ef`, `agents=v1`. This batch starts -from main `284cbcf` and aligns query-string handling with owned official observations -from campaign scan 1. Findings with request IDs are retained in -`~/.parsar/remediation/20260923/campaign-scan-1/{vaults-agents,sessions,skills-files-templates}/findings.json`; -the batch plan is `~/.parsar/remediation/20260923/list-query-tolerance/PLAN.md`. -Parts of the deferred list above are now qualified; the remaining parts stay -deferred below. - -| Row | Case and families | Core behavior | Evidence (finding: request ID) | -| --- | --- | --- | --- | -| A1 | Unknown list key: Agents, Sessions, Turns, Items, Templates, Vaults, Credentials, Skills, Skill versions, Files, Subagent and Artifact lists | Ignored; the page equals the request without the key | VA-02: `req_03f2f965efd54bfab3d22d079121e604`, `req_00357eb8b919413483dba9b59fb18174`; SES-17: `req_8186446b7f68412b8e80e6ec03f8d6d0`; SFT-11: `req_89660f3432624f57afe2c103356fec45` | -| A2 | Unknown key on single-resource GET/POST/DELETE routes | Ignored, including `tenant_id` and `include`; missing and foreign resources still return the same 404 | VA-02: `req_29b59b9bf58a482ca00222b86b542c8a` (deleted Vault read), `req_1aa8acb5602f4a09b7043885d42fca68` (deleted Agent delete) | -| B1 | Repeated supported key on Beta lists, including scalar `status` | 400, code `invalid_request_error`, param null, ``Failed to deserialize query string: duplicate field `` `` | VA-03: `req_83ee26a0b9ba4e84af8996a9346f261b`, `req_de6a02d9a9c547e490e6e53aeb45544d`, `req_92b924b186cf48488cf625638130bb16`; SES-16: `req_86ab2f32451a426399e76b21debeadba`; SFT-24: `req_794a86f4e72946dfb688a2e6f23fff32` | -| B2 | Repeated supported key on Skills and Skill versions | 400, code `duplicate_parameter`, param ``, with the observed message | SFT-10: `req_36e64628c2a44b1198d00949c4ed6cc8` | -| B3 | Repeated key on Files | Unchanged local `unsupported_parameter` rejection | SFT-18: `req_9a38e0628b4e40bd8e7202ac0880209d` accepted one identical repeated `purpose`; a single sample does not define which differing value wins | -| C1, C2 | `limit=0` and `limit>100` on Agents, Sessions, Items, Templates; Subagent Items and Subagent Turn Items since the [Subagent visibility batch](subagents.md#subagent-visibility) | Clamped to 1 and 100 | VA-01: `req_4954e90768b54c588167333e83f0a958`; VA-18: `req_d43d918872cb40b5b6e545582ea94205`, `req_2efee54ffd6f41179e294870b0e62e02`; SES-10: `req_4f1fb6cb479a4d1cb64486c72d250e9f`; SES-11: `req_5af4e60768b14c9eab284bce8638348a`; SES-12: `req_2f7cab87400348408ab8bb6e68789dc9`; SES-13: `req_b5ebab5691314bfba6ad59c84c42d04e`; SFT-23: `req_f63fef0e08c1462c8158c0ee8523a88d`; SAT-02: `req_6179ae6c1d1640d899ee4798e7f9fa57`, `req_7436104afbae4e73a0eb43b00ec9e660`, `req_32899313414b4031849a22cd2927f0ad` | -| C3 | `limit` 0 or above 100 on Turns, Subagents and Subagent Turns | 400, code `invalid_request_error`, param null, `limit must be between 1 and 100` | SES-14: `req_2793f57b6a454c399a031277b6a02e45`; SAT-03 (campaign scan 3): `req_7df58d9579be4ee3ab7fdab55286aa05`, `req_b4321de4480c4a8e96b9ea285ff63a46`, `req_0f437ad4713d47a8af1f61a88636bf79`, `req_e9d476dd2a69472694cffc0851d0574c` | -| C4 | Negative or non-integer `limit` on Beta lists other than Vaults and Credentials | 400, code `invalid_request_error`, param null, `Failed to deserialize query string: limit: invalid digit found in string` | VA-04: `req_1362046e9d69497da9c23ca69517a026`, `req_c384455192dc4a99a032299a91416a74`; SES-15: `req_918128738f3a47b69203ca091f9cb9ca` | -| C5 | Vault and Credential `limit` | 0, negative and above-100 values keep the pinned clamp; a non-integer uses the C4 error | VA-04: `req_f08cda4e1b9048808be9465b9a56c4a4`, `req_e53033a2e01e4b87aa0bb8d32c932a44`; VA-06 (pin conflict): `req_2ca22663c9414afa920b5509d3574812`, `req_1d088639a3614f6ba545cd36497ff35c`; VA-18: `req_bc47269acf8143f7869908567272e433`, `req_62eaf8fd15e8483a98a9a09ebc4d5e51` | -| C6 | Skills and Skill versions `limit` | `0` returns 200 with empty `data`, null first/last IDs and `has_more` true only if a resource follows the cursor; above 100 is `integer_above_max_value`, below 0 is `integer_below_min_value`, both with param `limit` | SFT-08: `req_b71c126adaf3437d9b01aee3ec913863`, `req_218a5d425b7047a190e2c5dc2e047e81`; SFT-09: `req_0ed2668ecab54253ab40d2655954de2c`, `req_4dd2e453717f4f82b5a7921a34d68b0f` | -| C7 | Files `limit` 0 or 10001 | 400, code null, param null; the local message stays | SFT-17: `req_05ec0b8f86b445b2beb2fdfc595d0879` | -| D1 | Scalar `status` plus `status[]` on Vaults and Credentials | 200, filtering by the union; invalid values still reject | VA-05: `req_1e89a6740ecd49b5b3b42f514a144e71`, `req_6223c9b384a246df848cfb94ffe2144d` | -| D2 | Explicit empty `purpose=` on Files | Same as omission, no filter | SFT-14: `req_af6df1ba19fd428fba6b1a66745ce37f` | - -### Decisions - -The pinned Turn, Item and Template docstrings say "between 1 and 100". A server -that clamps out-of-range values keeps each effective page inside that documented -range, so clamping Items and Templates to match the live service does not contradict -the pin. Turns keep rejecting because the live service rejects. Vault and Credential -negative limits keep clamping because their pinned description says values are -"clamped between 1 and 100"; the official 400 for `-1` (VA-06) remains a recorded -pin conflict. - -No pinned single-resource operation sends a typed query parameter, including -Environment retrieval, so no single-resource route keeps a query rejection. - -Unchanged: every default (20, Files 10000), the maximum page capacity (100, Files -10000), `order` handling, cursor lookup order, `after` trimming, authentication, -tenant scoping and missing/foreign masking. A `tenant_id` query key is an ignored -unknown key; only authentication selects the tenant. Shared-parser list rejections -happen before any resource lookup, so tenant B receives the same error. The -Environment Files list validates its query only after the Environment lookup, so a -foreign or missing Environment returns 404 first. No rejected request writes. - -Families are still selected by path, as for order errors. Local choices for -unsampled inputs: - -- Artifact lists keep rejecting 0 and values above 100, now with the Turn error - fields. Subagent lists were a local choice here too; campaign scan 3 later - sampled them, and the Subagent visibility batch applies rows C1–C3 above. -- Beta limits above the signed 64-bit range still reject, with code - `invalid_request_error` and `...limit: number too large to fit in target type`. - A leading sign or any other non-digit input, including an empty value, uses the - C4 message. -- A non-integer Skills limit keeps the local `invalid_request` code; its message now - names the 0–100 range. A non-integer Files limit keeps `invalid_request`. The - hosted Files schema message and `detail` member are not copied. -- Duplicate keys are reported before value errors, and limit errors before order - errors. Vault status validation still precedes limit and order. Precedence - between simultaneous errors was not sampled. - -### Deferred - -These remain registered differences and are not changed here: repeated Files -`purpose` values (SFT-18); unsampled overflowing limits; Skill sole-version -deletion and number reuse (SFT-01/02, since resolved or recorded in -[file resource semantics](file-resource-semantics.md#sole-version-deletion--september-23-2026)); -Session deletion lifecycle (SES-29/30, since addressed by the -[deletion batch](official-semantics-alignment.md#session-deletion-lifecycle--september-23)); -whitespace input (SES-01..04, since addressed by the -[whitespace batch](official-semantics-alignment.md#whitespace-only-message-text--september-23)); Template network forms (SFT-21/22); and response -defaults (VA-11, SES-23/25). Malformed path IDs (SES-28), metadata and name error -fields (VA-07/08/09), U+0000 (VA-10) and Template network codes (SFT-20) are -addressed by the [validation error batch](official-semantics-alignment.md#validation-error-fields--september-23). -Unknown and repeated keys on the Environment Files list (HE-34/35) follow rows A1 -and B1 through the [Environment Files wire batch](environment-files.md#wire-alignment--september-23-2026); -that list still rejects malformed query encoding (such as `?foo=%GG` or `;` -separators), which the shared lists drop. - -### Acceptance boundary - -Go handler and parser tests cover every row, including tenant isolation, and the -Skills store test covers the zero page against PostgreSQL. `official_list_query.py` -replays rows A1–D2 through raw HTTP and the pinned SDK against Core and a dedicated -PostgreSQL database as tenant A and tenant B, and confirms that rejected requests -change no resource. It also checks that single-resource reads, event admission, -streams, updates and deletions ignore unknown keys. This is resource and query -acceptance with no model execution: query parsing does not affect execution, so -live model acceptance does not apply. The independent batch acceptance, server gate -and review are recorded separately when complete. - -## List cursor errors — September 23, 2026 - -The pin is unchanged: SDK 3.13.0, commit `d7c41ef`, `agents=v1`. This batch starts -from main `c5cb6b56` and aligns the response to an `after` cursor that does not -resolve within its list (ERROR-PROTOCOL-001). Findings with request IDs are retained -in `~/.parsar/remediation/20260923/campaign-scan-4/errors/findings.json` (ERR-01..06, -raw records labelled `cur-*` in `official/results.json`), with SAT-04 from campaign -scan 3 and HE-57 from campaign scan 2; the batch plan is -`~/.parsar/remediation/20260923/cursor-errors/PLAN.md`. "Unresolved" covers a random -well-formed ID, a malformed value, an ID of another resource type, a resource of -another parent, a deleted resource and a resource of another tenant. - -| Row | Lists | Core behavior for an unresolved cursor | Evidence (finding: request ID) | -| --- | --- | --- | --- | -| C1 | Agents, Sessions, Turns, Templates, Vaults, Credentials | 404, type and code `not_found_error`, param null, `Resource not found.` A malformed cursor takes the missing-cursor path instead of the former 400 `invalid_request`, so every unresolved cursor, including a foreign one, gives the same bytes | ERR-01: `req_68e57640f20d451e879f11e8882e9542`, `req_3aee69064ee544eeafaf1fead240e3cb`, `req_a6c9a7b7d84d4db49117a57be9b1e16d`; ERR-13 (random, other type, other parent): `req_a740ac9e3a55476ea9bf0ebf7e53272c`, `req_9df334089c33493584d3ceb5d8c86041`, `req_90aa82ac18ab4a82bf49f940c25a467e`, `req_44a2293ebaf049df8674ef5ea0379abd`, `req_ff3df24a0d4a4e21a2464b5c1ee38179` | -| C2 | Session Items, Subagent Items, Subagent Turn Items | 400, type and code `invalid_request_error`, param null, ``Invalid session item ID in `after` `` | ERR-02: `req_bab2c1aee7dd4415a814bcb94d0dac24`, `req_8d981f935c144b8ebe1a6c2866edb2ea`, `req_aa6340e02f1345758e7b28e83e8f1095`, `req_c71cb02b9363422d92222bf0a5b3df80`; ERR-03: `req_bf7fb4a67c004405b549e848d9dc4be8`, `req_0486db0d18a8415a96eb9410fc62d369`, `req_b21069abb6ad42608e8d3d1ca6f7fa1c`, `req_7531a7851063434098949ed171432af9` | -| C3 | Subagents, Subagent Turns | 400, type and code `invalid_request_error`, param null, ``Invalid resource ID in `after` `` | ERR-04: `req_6d21661de5a74df4bc35c27ad1d1dca8`, `req_6368d24f88dd4eb2acd230745e704621`, `req_5803615424dd424bbf8f007291f37cf9`, `req_f7f5232125bc49278ebde73738d00f7b`; SAT-04: `req_23bdf5f0b8f24813b191634eb6e69256`, `req_e76ce5d4c6494b06a788d3c342f76ad2`, `req_83b9d26129fd43eea67b795b684bbb55` | -| C4 | Session Artifacts | 400, type and code `invalid_request_error`, param null, `after is not a valid artifact ID` | ERR-05: `req_9c4aa6bfabfe4e6599cb317c416c04f4`, `req_333f2ca75b274ad7aedb1e854b28eb57`, `req_5ef5b49127c443f19f6a18d324d2e8e5`; HE-57: `req_22bb488390324d9ebd7779e0be6b7f6b` | -| C5 | Skill versions | A value that does not begin with `skillver`: 400, type `invalid_request_error`, code `invalid_value`, param `after`, ``Invalid 'after': ''. Expected an ID that begins with 'skillver'.`` (``Invalid 'after'. Expected an ID that begins with 'skillver'.`` when the value is not echoed). A version of another Skill in the tenant: the same fields with `Skill version cursor does not match this skill.` A missing, deleted or foreign version, or a `skillver` value with a malformed tail: unchanged 404, type `invalid_request_error`, null code and param | ERR-06: `req_0e7ac00f5359403b952b66d88afb8011`, `req_41d7bcc6522c4230a785b0f23e7043ae`, `req_9b3cd31591594c9797252408cc871512`, `req_7e78eb04a9d04c8497b93d9206a206be`, `req_5d037c1b64cb4475ae11b5b5ab572c84` | -| K1 | Files, Skills, Environment Files `page` | Unchanged: Files 404 with param `after`, Skills 404 with null code and param, Environment Files keeps its page token error | ERR-08: `req_e8a09b54bc804eaa9344270a69943252`; ERR-09: `req_a0b0414f74474cb2a1f1b61dac54edb7`; ERR-17: `req_5c2494e082714500bccbadbf60c6195d` | -| K2 | Valid cursors | Unchanged ordering, paging, `has_more`, first and last IDs and limits | — | -| K3 | Missing or foreign parent: Session, Vault, Subagent, child Turn, Skill | Still 404 before the cursor is read, including a Skill version cursor that does not begin with `skillver` | — | - -### Decisions - -- One typed store error carries the family message. The API chooses the fields by - path family, as for order errors: Skill versions use `invalid_value` with param - `after`, the Beta lists `invalid_request_error` with a null param. -- C1 lists reuse the malformed path-ID approach: a cursor that cannot name a - resource resolves to the never-assigned maximum UUID and the normal lookup runs, - so storage failures and missing rows behave exactly as for a well-formed cursor. - The Core Runtime observation list pages by Session ID and follows the same rule. -- Every cursor is resolved only inside its already resolved parent and tenant. - Other-parent, other-type and foreign cursors therefore take the same path as a - missing one and cannot reveal another tenant's resources. -- The Skill version cursor lookup is now tenant-wide, without a schema change, so - another Skill's version can be told apart from a missing one; another tenant's - version is still missing. The `skillver` prefix check is case-sensitive and uses - the observed prefix. -- Like the official message, the prefix error repeats the caller's value, but - only when it is at most 256 bytes of valid, printable UTF-8, the rule already - used for echoed field names. A longer, unprintable or invalid UTF-8 value is not - echoed, so the error body stays bounded; the Skills `order` error follows the - same rule (`Invalid value. Supported values are: 'asc' and 'desc'.`). -- A parent lookup and its cursor lookup that run as separate statements - (Artifacts and Skill versions) re-check the parent before reporting a 400, so a - parent deleted in between still gives its 404. Deleted Sessions and Skills never - reappear, so a parent found by the re-check also existed when the cursor was - read. Item and Subagent lists read both inside one locked Session transaction. -- Observed upstream failures are not copied: a Turn cursor from another Session - (ERR-11) and a Credential cursor equal to its Vault ID (ERR-12) stay 404, and a - Skills cursor that is not a Skill ID (ERR-10) stays the Skills 404. - -Unobserved and inferred: a Subagent Turn Items cursor naming another Turn of the -same child, another Session's Artifact, and every foreign-tenant cursor follow their -family's rule without an official sample. - -Deferred: deleted Agent and Session cursors, which still anchor pages officially -(ERR-15), would need tombstones; the official 404 message text naming the resource -(ERR-16) is not copied; 409 fields belong to a separate batch. - -### Acceptance boundary - -Go store tests cover each changed list, and `list_cursor_public_test.go` replays -rows C1–C5 and K1–K3 over real HTTP and PostgreSQL for tenant A and tenant B, with -seeded Subagent and Artifact history and exact response bytes, including long, -control-character and invalid UTF-8 Skill version cursors. For K2 it pages every -changed list one resource at a time in both orders and checks the page contents, -`has_more` and the first and last IDs. The path-ID and -Artifact filter replays and the API error mapping test were updated. -`official_list_query.py` checks rows C1, C2, C5, K1 and K3 through raw HTTP and the -pinned SDK for both tenants, and the Subagent acceptance script asserts the C2 and -C3 errors on child lists. This is resource and query acceptance with no model -execution. The independent replay, server gate and review are recorded separately -when complete. diff --git a/contracts/agents-api/message-content.md b/contracts/agents-api/message-content.md new file mode 100644 index 000000000..bfd45bf65 --- /dev/null +++ b/contracts/agents-api/message-content.md @@ -0,0 +1,70 @@ +# Message content + +User messages and function results share one content model: an ordered list of `input_text` and `input_image` parts. Core stores message boundaries, part order and image references exactly as sent and returns them unchanged in user Items. It never downloads, transcodes or repairs media. Session creation `input` and `events.create` messages share validation and admission; [Sessions, events and history](sessions-events.md#send-input) covers admission, request limits and errors. + +## Messages + +A message has `role: "user"`, an optional `type: "message"` and a non-empty `content` array. Event `input` is an array of messages; Session creation `input` may also be a string, which becomes one text message. An explicit null or empty message `type`, a string `content` or a string event `input` is invalid. + +A message is valid when it has an image or at least one non-empty text part. Core never trims text. These requests return 400 `invalid_request` and write nothing: + +- an empty `input` string, `input` array or `content` array; +- a message whose text parts are all empty and that has no image. + +An empty text part beside other content, such as `["", "text"]`, is accepted and stored as sent. + +## Images + +An `input_image` part carries `image_url` as an inline data URI: `data:image/png;base64,…` or `data:image/jpeg;base64,…`. The base64 must be canonical and the decoded image must match the declared type. Core accepts no remote URL, `file_id` or `detail`. + +| Harness | Message images | +| --- | --- | +| Codex | Accepted on every placement the harness supports: `none`, `openai_hosted` and `self_hosted` | +| Claude Code | Accepted on `none`, `openai_hosted` and `self_hosted` | +| MiniMax Code | Rejected | + +Admission checks the harness before anything is written; an image the harness cannot take returns 400. The Runtime must also report message image support: Core binds a Session whose input carries images only to such a Runtime, and delivery to a Runtime without it fails. Use a model that accepts images. + +## Whitespace-only text + +Whitespace-only text such as `" "` or `"\n\t"` is valid content and is stored and returned verbatim. Whether a harness can run it is declared in its engine profile: + +| Harness | A message with no image and no non-whitespace text | +| --- | --- | +| Codex | Admitted and delivered unchanged | +| Claude Code | 400 `unsupported_or_invalid_configuration` | +| MiniMax Code | 400 `unsupported_or_invalid_configuration` | + +The rejection applies to Session creation (including streaming and `self_hosted` creation) and to `events.create`, before any write, reservation or Turn, so a running Turn is never disturbed. Whitespace beside non-whitespace text in the same message is admitted for every harness. Whitespace is the union of Go `unicode.IsSpace` and ECMAScript `String.prototype.trim`, for example U+0085 and U+FEFF; Core admission and the Claude bridge use the same set. + +## Function results + +An `agent.session.input.tool_result` event carries `success`, an optional nullable `error` string and an optional nullable `output`: a string or an ordered array of `input_text` and `input_image` parts. + +- Core stores the result as submitted, including which of `output` and `error` were present, and uses it for retry identity. Public Items always carry both fields ([Item rules](sessions-events.md#turns-and-items)). +- The Runtime receives one ordered content list: the `output` parts, then the `error` text as a final text part. This conversion never changes the stored result. +- A result with images needs a Runtime that reports function-result image support; only image-bearing results check it. When the Runtime refuses a result after admission, the Turn fails without a confirmed application, and the stored result stays readable. + +| Harness | Function results | +| --- | --- | +| Codex | Text and ordered text/image output. Core checks only that each part is well formed and passes image references to the harness unchanged | +| Claude Code | Text output. Images only in successful results and only as inline PNG or JPEG; an image in a failed result or a remote reference returns 400 before anything is stored, and the pending call stays open. The harness may resize or re-encode images in its own history; public Items keep the submitted bytes | +| MiniMax Code | No public functions | + +### Application receipts + +Admission (202) does not mean the harness used the result. The pending call clears when the adapter confirms native application: + +- **Claude Code** confirms with a live root native tool result that matches the Session, call ID, success flag, exact text, block count and order, with a native image at every image position. Replayed, synthetic and Subagent records do not confirm it. +- **Codex** confirms with the live root dynamic-tool `item/completed` observation that matches thread, Turn, call, function name, status, success flag and the exact ordered content. A write to the harness alone does not confirm it. + +Both adapters wait at most 10 seconds for a receipt. A timeout or native release without confirmation leaves application uncertain. Confirmation means the harness recorded the result, not that the model provider consumed it. Core never replays a result automatically; the submission stays stored for recovery reads. + +## Runtime boundary + +The Core–Runtime wire carries messages as `MessageInput` for initial input, prepared start and steering, with the same ordered `InputContent` parts that function results use ([Core–Runtime protocol](../../docs/runtime-protocol.md)). Admission checks the harness's declared profile; binding and delivery check what the Runtime reports. Adapters own native encoding and application receipts. A text-only adapter rejects image parts instead of dropping them. + +- **Codex** flattens a batch into its native input list with a blank-line separator between public messages. Public message boundaries stay in Core's storage; the native history does not keep them. +- **Claude Code** sends native image blocks and a UUID per native user message. One public input is applied only after every message in its batch is consumed. Within one native Turn the bridge accepts at most 64 user messages, including the opening prompt; it rejects a steering batch that would exceed the bound before submitting any part of it, which ends the running Turn. The daemon requires bridge protocol 3. + +Native message and function-result image checks are listed under [qualify the adapter](harness-onboarding.md#qualify-the-adapter). diff --git a/contracts/agents-api/message-input.md b/contracts/agents-api/message-input.md deleted file mode 100644 index 710d417b4..000000000 --- a/contracts/agents-api/message-input.md +++ /dev/null @@ -1,185 +0,0 @@ -# Ordered message input - -The pinned SDK defines user messages as ordered `input_text` and `input_image` -parts. Core retains message boundaries, content order and the supplied image -reference in input persistence and user Items. Session creation and subsequent -events share validation and atomic admission. - -## Supported profile - -Codex and Claude SDK support inline PNG/JPEG data URIs on `environment:none` and -Core-managed Docker `openai_hosted`, for initial, prepared and active input. Use a real vision-capable model. -The existing 1 MiB HTTP and 512 KiB durable input limits still apply. A successful -events response acknowledges persistence, not native consumption. Active input -advances its durable receipt only after the adapter confirms application. - -```python -import base64 -from pathlib import Path - -image_url = "data:image/png;base64," + base64.b64encode( - Path("example.png").read_bytes() -).decode() -session = client.beta.agents.sessions.create( - agent={"model": model}, - environment={"type": "none"}, - input=[{"role": "user", "content": [ - {"type": "input_text", "text": "Describe this image."}, - {"type": "input_image", "image_url": image_url}, - ]}], -) -``` - -Query Session/Turn/Items to recover results. SSE remains live-only; reconnecting -does not replay inputs or recreate completed Turns. The request contains no -image `detail` or file-ID extension. - -## Text content - -Text is never trimmed. A message is valid when it contains an image or at least -one non-empty `input_text` part, so whitespace-only text such as `" "` or -`"\n\t"` is admitted at Session creation (string or message input) and by -`events.create`, then stored and projected in Items verbatim, as the official -service does (SES-01..04). The empty string, an empty `content` array and an -empty `input` array keep the existing 400 `invalid_request` response. A message -whose parts are `["", "text"]` is still accepted; its official behavior is -unobserved (SES-08). String `content` and string event `input` remain type errors, -as observed officially. The shared daemon validator, daemon dispatch and the -TypeScript client apply the same rule. Core Web keeps local UI rules: its -composer trims leading and trailing whitespace from every message it sends and -does not send blank text, and its Start Session form omits whitespace-only simple -text ([console API usage](../../docs/web/console-api-usage.md#not-consumed)). Other clients' text is -never trimmed. - -Harness profiles declare whether whitespace-only text is qualified, through the -same engine profile that declares image support. Only Codex is qualified; it -delivers such text unchanged and completed whitespace-only Turns in live -acceptance. Claude SDK is not qualified: the bridge and Anthropic-compatible -providers reject text without a non-whitespace character. MiniMax Code is not -qualified: its native runtime refused a whitespace-only prompt with "Local message -content or attachments are required.", failing the Turn with `engine_failed` -(public `internal_error`). Whitespace here is one explicit set, the union of Go -`unicode.IsSpace` and ECMAScript `String.prototype.trim` (for example U+0085 and -U+FEFF), shared by Core admission and the Claude bridge. For Claude SDK and -MiniMax Code Sessions, a message without an image or any non-whitespace text is -rejected at Session creation and -`events.create` with 400 `unsupported_or_invalid_configuration`, before any write, -input reservation or promotion, so no Turn starts and a running Turn is never -disturbed. Whitespace beside non-whitespace text in the same message is admitted -unchanged. Core never trims or pads model input to fit a harness. Admission makes -the Claude bridge's own check unreachable. If such a message still reached the bridge, -initial and prepared input would end with `invalid_request`, and a steering -message would be reported as `input_rejected`, which ends the running Turn as -before. - -## Runtime boundary - -The current private wire uses `MessageInput` for initial requests, prepared start and -active steering. Each message contains ordered `InputContent` parts, also reused -by function-result content. Core does not download, transcode or repair media. -The separate Claude bridge reports protocol 2; daemon readiness rejects protocol 1. -Adapters own native encoding and native application-receipt mapping. Existing -text-only adapters reject image parts instead of dropping them. - -Codex converts content to its native flat input list and inserts blank-line -separators between public messages. Public message boundaries remain durable; -independent native message boundaries are not claimed. Claude uses native image -blocks and a UUID for each native user message; one public active input is applied -only after every message in its batch is consumed. Its native 64-message bound is -checked before any part of a batch enters the iterator. - -Public profile qualification and Runtime advertisement are separate. Admission -checks the registered operation-specific image profile; device selection and -delivery check actual image support. Text-only operations retain existing offline -admission. Adding an adapter must implement the shared contract and qualify the -public operation, without adding engine-name branches to Core. - -## Validation and remaining gaps - -`TestNativeMessageImagePublicExecution` runs the pinned SDK and raw HTTP against -Core, a dedicated PostgreSQL database, the real daemon and a native harness. -Set `OAC_TEST_MESSAGE_IMAGE_ENGINE` to `codex` or `claude_sdk`, provide private real -provider options via `OAC_TEST_MESSAGE_IMAGE_REAL_OPTIONS`, and use the existing -`OAC_TEST_NATIVE_DAEMON_BIN`, `OAC_TEST_NATIVE_PROOF_DIR` and -`OAC_TEST_OFFICIAL_SDK_PYTHON` fixture settings. The test never supplies model responses. - -Real Kimi K3 acceptance on 2026-09-22 used randomized four-color PNGs whose answers -were absent from the input text. Both adapters passed initial ordered input, -active replacement with a different image, retry deduplication, exact retained -user Items, native receipts, cold daemon history continuation, cancellation and -ordinary text continuation, malformed-batch atomic rejection and tenant isolation. -Evidence is under `zju_a100_2:~/.parsar/remediation/20260922/message-image-input/`: -`public-codex-final.log` / `message-image-public-2718944397/public.json` and -`public-claude_sdk-final.log` / `message-image-public-700202073/public.json`. -After final lifecycle and bridge-version changes, both complete public chains -passed again: `public-codex-reviewed.log` (59.416s) and -`public-claude-reviewed.log` (79.525s). MiniMax ordinary text, active steering, -receipt, cancellation and cold continuation passed in -`public-mcode-network-fixed.log` (48.963s). The first MiniMax attempt hit a stale -hosted-fixture assertion; a subsequent attempt failed because its test network -relay was absent. Both failures are retained, and no production model behavior -was changed to obtain the passing result. - -The final required gate was split by host: server `make -o check-web check` -passed, and local Node22/pnpm10.30.3 `make check-web` passed all 73 browser tests. -An earlier complete server `make check` stopped at its missing Chrome executable. -The dedicated PostgreSQL cancellation/deletion interleaving regression passed -five repetitions, and shared input/dispatch race checks passed. These checks do -not qualify workspace images or additional native/provider combinations. -Native-only probes are feasibility evidence, not public qualification. - -## Docker workspace acceptance - -`tests/official_workspace_images_native.py` exposes `verify_workspace_images` -for an operator-owned standalone deployment. Supply the fixed SDK clients for -two tenants, their raw HTTP transport, a real vision model, the selected harness, -a cold Core/Runtime restart callback and a private evidence path. It creates and -deletes its own hosted Sessions; it never supplies model responses or credentials. - -The common workflow covers initial text/PNG/text input, prepared image-only JPEG -followed by a separate message, active PNG input while a function waits, and a -6000x2100 PNG function result. The model must read the band order from image pixels -and write the corresponding bytes with native tools. Files listing and immutable -Artifact downloads verify those bytes. SDK and HTTP Items must retain the original -ordered input/results. The same workflow checks retry/conflict, unsupported-input -non-mutation, tenant and same-tenant Session isolation, cold history continuation, -pending cancellation and ordinary text after cancellation. - -Real Kimi K3 acceptance on 2026-09-22 passed this workflow for Codex 0.153.4 and -Claude SDK 0.3.269/native 2.1.269. Evidence is retained under -`zju_a100_2:~/.parsar/remediation/20260922/workspace-images/`: -`codex/public-run-_lr5uiy_/` (152.78s) and -`claude_sdk/public-run-sb65p9e9/` (197.84s). Each run completed seven public Turns, -including one cancelled Turn, and cleaned up both owned hosted Sessions. Initial -Codex attempts exposed two acceptance-script errors: treating an earlier idle -event as the submitted Turn's completion and sending a scalar message to the -array-only events endpoint in a conflict probe. Both failures are retained; -production lifecycle behavior was not changed to obtain passing results. - -Here `openai_hosted` means the Core-managed Docker deployment, with daemon, native -harness, tools and workspace in one sandbox. Image admission uses Environment type -and the common operation-specific Runtime support; it adds no provider-name branch, -media downloader, file permission or preparation lifecycle. Docker evidence does -not qualify other providers or user-managed deployment. Both adapters now require -native function-result confirmation as described in the -[receipt coverage](function-result-images.md). This does not imply crash-safe -exactly-once tool effects; the original workspace acceptance predates Codex's -native receipt qualification. - -## Remaining gaps - -MiniMax Code image input, remote HTTP(S) image URLs, -other media types and full upstream error/default semantics remain unqualified. -MiniMax's fixed ACP advertises `image:false`; its adapter rejects images. These are -implementation gaps, not changes to the official protocol. JPEG parsing/conversion -has deterministic coverage; the original `none` message fixtures are PNG, and -the hosted workflow also exercises JPEG. Empty -parts beside text (SES-08), local payload limits and native batch-size parity need -upstream evidence. -Function-result image support has its own [coverage record](function-result-images.md). No full protocol -compatibility or support for arbitrary vision-model/provider combinations is claimed. - -User-managed Linux image execution follows the same native workspace path. See -[current qualification](harness-capabilities.md) for the tested -Harnesses, formats, continuation and remaining boundaries. Historical evidence -above retains its original deployment scope. diff --git a/contracts/agents-api/official-semantics-alignment.md b/contracts/agents-api/official-semantics-alignment.md deleted file mode 100644 index 3bd005253..000000000 --- a/contracts/agents-api/official-semantics-alignment.md +++ /dev/null @@ -1,973 +0,0 @@ -# Official wire semantics: September 22, 2026 - -This batch compares owned-resource requests to the official Agents API with Core -main `692e32daafb19521e0919915c7685eef78b813fa`. The protocol remains Python SDK -3.13.0, upstream `d7c41efee1b0802b79f3f88a678ef2052b06e9ce`, `agents=v1`. -The current [Session reference](https://developers.openai.com/api/reference/python/resources/beta/subresources/agents/subresources/sessions) -and [overview](https://developers.openai.com/api/docs/guides/agents-api/overview) -were consulted alongside the pinned source. Current documentation is not a -replacement baseline. A successful SDK parse alone is not conformance evidence. - -## Selected common behavior - -| Operation | Official observation and Core behavior in this batch | -| --- | --- | -| Create Agent, Vault, Credential, EnvironmentTemplate or Session | HTTP 201. Session JSON and live SSE creation use the same status. Non-creation successes retain their operation-specific status. | -| Submit Session events | HTTP 202 with an empty response body. An empty array is an authenticated no-op; null is invalid. No-op requests do not create a Turn, Item or execution retry reservation. | -| Session, Turn and Item lists | `object: list`, `data`, `has_more`, `first_id` and `last_id`. Empty pages contain null first/last IDs. | -| Empty Agent/Template update | Advance `updated_at` through the existing atomic update, preserving IDs, content, ownership and frozen Session snapshots. Timestamp precision is seconds; immediate updates may have the same serialized timestamp. | -| Static bearer create/replacement | Reject an explicitly empty token before mutation. Preserve valid opaque token bytes without trimming. | -| OAuth grant create/replacement | Reject an explicitly empty access token and a replacement without mutable grant fields. Preserve previously qualified refresh/expiry/null handling. | -| Input message discriminator | Omission remains valid; a supplied `type` must be `message`. Explicit null or empty strings reject through the shared initial/event decoder, matching the pinned literal type. | -| Missing beta resource | HTTP 404 with `type` and `code` equal to `not_found_error`. Missing and foreign resources remain indistinguishable. | -| Missing required Beta header | HTTP 400 with `type` and `code` equal to `invalid_beta`, before authentication (see [HTTP routing and response headers](#http-routing-and-response-headers--september-23)). | -| Missing non-beta File or Skill | HTTP 404 with `type: invalid_request_error`, `code: null`. Exact message, File `param` and additional detail payload remain outside this batch. | - -The resource comparison made 40 raw requests over six newly owned resources -(one Agent, two Vaults, two Credentials and one Template), including actual -rejected replacements and post-delete reads. All six resources were deleted. -Session probes used real `gpt-6-astra` executions and compared event admission, -query envelopes and accumulated usage; their owned Sessions and Agent were also -deleted. Separate missing File/Skill probes used randomly generated IDs. -Private request/status/body evidence and cleanup results are retained under -`~/.parsar/remediation/20260922/official-semantics-alignment/`; no API keys or -credential values belong in the repository. - -## Explicit remaining differences - -- Current documentation supports Session Agent configuration updates; the fixed - `SessionUpdateParams` exposes only metadata. New fields and newer Environment - status/configuration shapes are a queued baseline upgrade, as approved by the user. -- The Session admission batch rejects missing/null input for `none` and for - streaming creation outside `self_hosted`. September 23 official probes confirmed - these conditions and idle self-hosted creation. Official whitespace-only string - input returned 201; the - [whitespace batch](#whitespace-only-message-text--september-23) admits it. -- Two otherwise identical official creates with the same `Idempotency-Key` - returned 201 and distinct Session IDs. Core retains its durable creation retry - guarantee. This is a local behavior, not evidence of official idempotency parity. -- Empty Session update now returns the observed 400 error; explicit metadata null - and empty-object clearing remain supported. Generic validation codes/field `param`, - malformed queries, page limits and overlapping mutation behavior need qualification. - The error mapping above must not be extrapolated to every status or resource. -- Template references with inline installation overrides, optional Skill version - semantics and the other open items remain outstanding. -- Files.create cannot tell a file written by an earlier Files.create from any other - existing file, so both report the untracked-file message; see the - [write semantics](environment-files.md#write-semantics--september-23-2026). -- Native model defaults, tool combinations and unavailable usage counters retain - their documented multi-harness differences. Core does not reconstruct model - output, guess counters or introduce a second tool loop to manufacture equality. - -These observations establish a bounded comparison, not complete official protocol -compatibility. Core regression and real-model validation are recorded with the -implementation acceptance before merge. - -## Core acceptance - -The deployed server/runtime at `7c80d604` passed seven real-PostgreSQL resource and -safety groups, including unchanged credential-row hashes after rejected writes -and restart. Codex and Claude Code used Kimi K3; MiniMax Code used MiniMax M2.7. -Each completed three real Turns covering JSON creation, live SSE creation and -event continuation, with history paging, empty no-op requests and tenant isolation. -Four earlier attempts were interrupted by failed test-network relays and retained -as unsuccessful evidence. After end-to-end TLS checks, the controlled rerun passed. - -The server `make check` gate passed with Web checks run separately: type checks, -production build, 287 client tests, 573 Web tests and 74 fixture browser cases -(the corrected Beta-error fixture was rerun separately). Full standalone fixed -Python SDK and Go-client acceptance passed. A fresh Astra high full-diff review -found no blockers and independently ran API/contract tests. Rebase onto main -`c96ea82` preserved every batch patch; the combined tree passed API/execution and -three PostgreSQL scheduling regressions. Test resources were scoped to this batch. -E2B, OAuth provider refresh and new native capability combinations were not requalified. - -## Session admission batch — September 23 - -The conditional input requirements are checked before creation lookup, credential -binding or execution. No idle-none legacy creation exception is retained; existing -Session GET and events remain available. Valid creation requests keep the local -same-key guarantee. An empty metadata update is rejected after authentication and -before resource lookup; supplied metadata still uses the existing tenant-scoped -update path. - -Core Web requires initial input for conversation-only creation. Hosted creation -without input uses JSON; input-bearing creation retains SSE. If a creation stream -fails before revealing the Session ID, the next user-initiated retry sends the -same draft and key as JSON to recover that creation. It does not replay an input -or introduce an automatic retry loop. - -The official probe made nine bounded create requests and created two owned -Sessions (self-hosted without input and none with whitespace). Both were deleted -successfully; no hosted environment was created. Metadata observations are reused -from September 22. Evidence: `~/.parsar/remediation/20260923/session-admission-alignment/official/`. -See [operation evidence](operation-evidence.md) for the wider 58-operation audit. -Real Core/daemon/native-model acceptance ran on production source `7ccc636`: -Codex and Claude used Kimi K3; MiniMax Code used MiniMax M2.7. Each completed one -JSON string-input Turn and one SSE ordered-message Turn, six total with no failed -attempt or rerun. Same-key JSON recovery retained each Session and exactly one -Turn. Twenty-one initial invalid creates plus three missing-input retries rejected -without adding rows to the eight checked execution tables. Empty updates preserved -metadata; null/empty clearing and foreign-tenant reads/valid updates were checked. -Hosted JSON no-input admission was retained for all three configured profiles; -self-hosted idle admission was checked on Codex only. Neither check claims a new -hosted or self-hosted execution capability. - -Evidence is retained under -`~/.parsar/remediation/20260923/session-admission-alignment/live/`, including raw -HTTP, fixed-SDK responses, SSE, history, native outputs, exact source/image hashes -and cleanup. Six assistant results matched the requested markers. Claude emitted -two assistant Items for its two-message input within one Turn; native Item counts -were preserved. All 39 evidence hashes and known-secret scans passed. Owned -Runtime/Core processes, three databases, three derived images and the dedicated -network forward were removed, and temporary devices were revoked. - -The integrated Web gate passed 287 client and 583 Web unit tests, type checks, -builds and all 76 fixture browser cases. These controlled UI checks include -empty-input prevention and same-key JSON recovery after a lost creation response; -they are separate from the real-model evidence. The server `make -o check-web check` -passed at `01f9356` with dedicated PostgreSQL, generated-query checks, Go service -and adapter tests/builds, and Rust tests/format/Clippy. The optional packaged -MiniMax native-tools probe was skipped; the separate real-model evidence above -qualifies this batch, not every native capability. Fixed Python SDK and Go-client -service acceptance also passed. The subsequent tool-policy evidence-count test -change compiled; its opt-in native run was not repeated. - -A fresh independent Astra high reviewer inspected all 66 changed files and found -no grounded in-scope blockers; API/contract tests, 39 focused Web tests and whitespace -checks passed independently. Rebase onto main `6a3131e` preserved every batch patch. -The combined tree at `4981580` passed API, execution, contract and dedicated-PostgreSQL -Environment scheduling/initial-input/creation-stream regressions. No new E2B, -OAuth provider or native capability combination was qualified. - -## Validation error fields — September 23 - -This batch aligns validation failures that Core already rejected with the -official `code` and `param` fields. It does not change any limit. Evidence comes -from the campaign scan at main `284cbcf`, recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-1/{vaults-agents,sessions,skills-files-templates}/findings.json` -(VA-07, VA-08, VA-09, VA-10, SES-28 and SFT-20), plus the September 22 Session -observation that `{"metadata":{"a":null}}` returns param `metadata.a`. - -| Row | Case | Core behavior | -| --- | --- | --- | -| M1–M3 | More than 16 metadata pairs, a key over 64 characters, a value over 512 characters (Agent create/update, Session create/update) | 400 with type and code `invalid_request_error`, param `metadata` or `metadata.`, and the observed official message with the actual count or length. Pairs are checked before keys and values, and keys in sorted order. | -| M4 | A non-string metadata value: integer, number, boolean, object, array or null (Agent create/update, Session create/update and Vault create; neither Core nor the pinned SDK has a Vault update) | 400 `invalid_request_error`, param `metadata.`, message `Invalid type for 'metadata.': expected a string, but got instead.` The first such value in document order is reported before the generic whole-body error. Templates accept no metadata. | -| M5 | Vault metadata size | Unchanged: no pair or length limits, only the local 64 KiB storage bound. | -| N1 | Agent `name` over 128 characters | 400 `invalid_request_error`, param `name`, observed message. Empty and untrimmed names stay accepted. | -| U1 | U+0000 in a stored string | Never 500 and nothing is written. Metadata keys and values report `metadata.`; other strings return 400 `invalid_request_error` with a null param. This is a local limit: PostgreSQL text and jsonb cannot store U+0000, while the official service accepts and echoes it. | -| I1/I2 | A malformed path identifier on any Beta resource route, and on Files, Skills and Skill versions | Byte-for-byte the response of a well-formed missing identifier on that route, including invalid bodies and queries, and a deployment without credential encryption. Foreign, missing and malformed identifiers stay indistinguishable. | -| T1 | Template network rejections (wildcard, port, scheme, IPv6, empty host, empty/null/omitted list with `restricted`, more than 100 domains, domains with another access) and the shared inline Session network | 400 `invalid_request_error` with a null param. Accepted hostname forms are unchanged; other unsupported installation fields keep `unsupported_or_invalid_configuration`. | - -Decisions: - -- A typed field error carries the param and message through the existing error - writer. Metadata type errors are found by reading the metadata object in - document order before generic decoding; limit checks keep their previous - position, so validation order relative to lookups (SES-33) is unchanged. -- U+0000 is checked explicitly in metadata, so the param is exact. All other - stored strings rely on mapping PostgreSQL `22021` (U+0000 or invalid UTF-8 in - a text parameter) and `22P05` (`\u0000` in jsonb) to 400 with the generic - message "Request text contains characters this service cannot store or compare, - such as U+0000 or invalid UTF-8." The same mapping covers query filters, for - example `agent_id=%ff` on the Session list. The persisted string fields are too - many to check one by one, and the database is the single place that knows which - strings are stored. The failing statement aborts its transaction; real-PostgreSQL - tests compare every public table before and after the rejected requests. -- A malformed path identifier resolves to the maximum UUID, which Core never - assigns because it only generates version 4 and 5 UUIDs. The request then - follows exactly the missing-identifier path, including body, query and storage - checks. Routes whose lookup is the next check keep their direct not-found - response. Request-body references are unchanged. Malformed list cursors were - later aligned by the [list cursor error batch](list-query-semantics.md#list-cursor-errors--september-23-2026): - Agent, Session, Turn, Template, Vault and Credential cursors take the same - missing-cursor path, and Item, Subagent, Artifact and Skill version cursors - return their list's official cursor error. -- Network messages are Core wording; the official prose is not copied. -- Documented message difference for M2: the official message abbreviated a - 65-character key as `'KKK...KKK'`. That single sample of identical characters - cannot reveal the abbreviation rule, so Core quotes the full key. Status, type, - code and param match. - -Deferred and unchanged: accepting and storing U+0000; hostname forms accepted -officially (SFT-21) and `disabled` with domains, which the official service -accepts (SFT-22); non-canonical UUID spellings such as uppercase, braces or -`urn:uuid:` still resolve to the same resource; Skill sole-version deletion and -number reuse (since resolved or recorded in -[file resource semantics](file-resource-semantics.md#sole-version-deletion--september-23-2026)); -Session deletion lifecycle; whitespace input (since addressed by the -[whitespace batch](#whitespace-only-message-text--september-23)); response defaults; -and the Files `limit=abc` code. The Environment Files list query parser is aligned -for unknown and repeated keys by the [Environment Files wire batch](environment-files.md#wire-alignment--september-23-2026); -it still rejects malformed query encoding locally. - -Go handler tests cover every row. Real-PostgreSQL tests replay every path-ID -route for malformed, missing and foreign identifiers (tenant B), with valid and -invalid bodies and queries, and replay U+0000 on every create/update family with -a database digest proving no writes. The pinned-SDK acceptance scripts assert the -new codes, params and messages. Independent real-Core acceptance is recorded -separately by the coordinator. - -## Artifact capture and listing — September 23 - -This batch aligns Session Artifact capture and listing with the first official -Artifact observations. Evidence comes from the hosted-environment campaign scan -recorded privately in `~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/` -(`findings.json` HE-50..62, raw records under `official/` and `run1/`). The probe -used three owned Sessions and two tiny `gpt-6-astra` Turns; all three Sessions -were deleted. Official Turn 1 created regular, nested and empty outputs plus -`outputs/link.txt -> a.txt`; Turn 2 only wrote `outputs/c.txt` after one Artifact -was deleted. - -| Row | Case | Core behavior | -| --- | --- | --- | -| A1 | A symlink below `outputs/` at Turn completion: to a file or directory, dangling, or pointing outside the workspace (HE-51) | Skipped by its `lstat` type: never followed, opened or resolved, and no Artifact. Every regular file is still captured and the Turn completes. | -| A2 | Later Turns in the same Session (HE-52) | A path is published again only when it has no remaining published Artifact in the Session, or its bytes (sha256) differ from the newest remaining one. Unchanged paths keep their existing Artifact IDs. The first Turn is unchanged. | -| A3 | List envelope (HE-53) | `object: list`, `data`, `first_id`, `last_id`, `has_more`, with null first/last IDs on an empty page, like the Session, Turn and Item lists. Paging and cursors are unchanged. | -| A4 | Malformed `environment_id` filter (HE-56) | 200 with an empty page, as for another existing Environment. Session lookup still runs first, so foreign and missing Sessions remain 404; cursor and limit errors are unchanged. | - -Decisions: - -- The Rust export helper handles every link kind the same way. Official evidence - shows one relative link to a file; telling the other kinds apart would require - resolving the link, which the confinement rules forbid. A link still counts as - a directory entry, so creating or removing one during export is a concurrent - change. -- The republication decision runs in the Turn's terminal transaction, not in the - private capture transaction. Capture commits and releases the Session lock - before the Turn completes, so an Artifact deletion can commit in between. The - terminal transaction holds the Session lock that also orders Artifact deletion, - and only one Turn per Session can be active, so the decision sees exactly the - Artifacts that remain at completion. Unchanged staged rows are deleted and their - private large objects unlinked in that transaction; published rows are never - modified. -- "Newest" follows the producing Turn's database creation time, then its ID. - Publication time can come from the Runtime's reported completion and is not a - reliable order between Turns. -- Known difference from the batch plan's wording, accepted as a local decision: - the plan republishes a path whose newest Artifact was deleted, but Core compares - against the newest *remaining* published Artifact. Deletion is physical and - leaves no record, and adding one would need a schema change outside this batch. - Example: Turn 1 publishes `b.txt` as `bravo`, Turn 2 publishes `bravo-v2`, and - the Turn 2 Artifact is then deleted. A later Turn whose `b.txt` is `bravo-v2` - republishes it, because the remaining Turn 1 version differs. A later Turn whose - `b.txt` is `bravo` publishes nothing, because the remaining Turn 1 Artifact - already has those bytes. The official behavior for this case is unobserved. -- A malformed filter resolves to the never-assigned maximum UUID, as for - malformed path identifiers, so it matches nothing without a database text - comparison. An empty `environment_id=` still means no filter. - -Deferred and unchanged: a linked `outputs` root, hard links, FIFOs, sockets, -devices and device crossings still reject the whole capture and fail the Turn -with `artifact_capture_failed`; there is no official evidence for them yet. -Republication after changed bytes is inferred rather than observed, and the -deleted-newest case above is unobserved. The unknown `after` cursor (HE-57) -was later aligned by the -[list cursor error batch](list-query-semantics.md#list-cursor-errors--september-23-2026). -Subagent lists keep their `data`/`has_more` envelope until there is official -Subagent evidence. Artifact IDs keep the Core UUID format. Paths removed from -the workspace keep their Artifacts. - -Rust tests cover every link kind, including absolute links to a secret outside -the workspace and a relative link to a workspace file outside `outputs/`; an -inotify watch proves no target is opened or read, with a positive control. They -also keep the hard-link, socket, FIFO, linked-root and concurrent-change -rejections. Real-PostgreSQL store tests cover new, unchanged, changed, -changed-back, deleted-then-unchanged and deleted-during-capture paths, a deletion -that holds the Session lock while Turn completion waits, Turn-ordered newest -versions with inverted publication times, Session scoping and private object -accounting. Handler and real-PostgreSQL HTTP tests -cover the envelope, other, foreign and malformed filters, and foreign or missing -Sessions. The pinned-SDK and raw HTTP verifier used by live acceptance runs -against PostgreSQL across three Turns. Real Core, daemon and model acceptance is -recorded separately by the coordinator. - -## Session deletion lifecycle — September 23 - -This batch aligns Session deletion with the observed official lifecycle rules. -Evidence comes from the campaign scan at main `beb18fd`, recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-1/sessions/findings.json` (SES-29 -and SES-30) with raw records under `official/`: `q5-delete-repeat.json`, -`q5b-delete-while-in-progress.json`, `q5-delete-never-existed.json` and -`q5-delete-while-running.json`, plus the September 22 retry-session cleanup that -first returned 409. - -| Row | Case | Core behavior | -| --- | --- | --- | -| D1 | DELETE of the caller's own Session that is already publicly deleted (SES-29) | 200 `{id, object: "agent.session.deleted", deleted: true}`, identical to the first confirmation, with no database write. GET, update, events, Turns and Items stay 404. | -| D2 | DELETE of a never-existing, malformed or foreign Session, including a foreign deleted one | Unchanged: the byte-identical 404 `not_found_error` of a missing Session. | -| D3 | DELETE while a root Turn is queued, in progress (including a requested cancellation) or waiting on required actions or function results, or while an input reservation is pending: a queued later input, self-hosted input awaiting a connection, or hosted initial input while provisioning (SES-30). Subagent child Turns and pending Environment file writes are not checked (see follow-ups) | 409 with type and code `conflict_error`, param null and message "session must be durably idle or failed without required actions before deletion". Nothing changes: no cancellation, marker, event, Artifact removal or Runtime cleanup. | -| D4 | DELETE of an idle Session, including an idle hosted Session still provisioning without input, and of a failed Session without required actions, including expired initial input | 200 with the existing public deletion and managed Runtime cleanup. | -| D5 | Callers that need to delete running work | Cancel first with `agent.session.input.cancel`, wait until the Session is idle, then delete. The Core Web offers this as an explicit action after a 409. | - -Decisions: - -- The rule is the one the creation stream already uses to settle: the Session is - idle or failed, no root Turn is queued, running or waiting, and the latest input - reservation is not pending. Deletion reuses the Store's active-Turn query and - reservation state, so a pending reservation blocks deletion even while the - public status projects idle. -- The decision and the marker commit in one transaction under the tenant Session - row lock that also orders Turn and input admission. Either admission commits - first and deletion returns 409 without mutation, or deletion commits first and - admission returns 404. A rejected deletion rolls its transaction back. -- A repeated deletion locks the owner's deleted row and returns the confirmation - without a write. Foreign and missing rows are never locked, so they stay - indistinguishable. Physical purge, when implemented, may end this idempotency. -- Documented stricter local behavior: official DELETE immediately after an - `events.create` 202 on an idle Session returned 200 (`q5-delete-while-running`); - its Turn was apparently not yet durably in progress. Core admits the Turn - synchronously in the 202 transaction, so Core returns 409 in that window. -- A self-hosted or hosted Session whose reserved input waits for its Environment - cannot be cancelled publicly (the pending reservation rejects new batches), so - it stays undeletable until the input starts, its five-minute deadline expires - or its Environment fails. The official behavior of that window is unobserved. -- Capacity change: previously, deleting a provisioning hosted Session with - reserved input released its sandbox node placement immediately. Now the - deletion returns 409, and the placement keeps counting toward the node's - retained and reserved capacity until the input is admitted or its five-minute - deadline expires. A later allowed deletion releases an unallocated placement. -- Earlier releases deleted busy Sessions after requesting cancellation. Their - markers can remain in upgraded databases; hidden-work settlement, restart - reconciliation and Runtime cleanup keep handling them unchanged. -- The Core Web keeps the plain delete action. When Core returns the busy 409, the - dialog reads the Session once. If a Turn is still busy it replaces the action - with Cancel work and delete, which sends one cancellation, reads the Session - until it is idle or failed without required actions (a 30-second bound checked - between reads) and sends one deletion. A rejected or uncertain cancellation, a - timeout, a connection change or another 409 stops without retrying. If the - Session reads idle or failed, or only awaits its Environment connection, only - pending input blocks deletion; Core rejects its cancellation, so the dialog - explains that the input must start, expire or fail first and offers no - cancellation. - -Follow-up: deletion checks only root Turns and input reservations. A subagent child -Turn that is still running and a pending Environment file write do not block it, -which matches the permissive behavior before this batch. The official behavior for -both is unobserved; decide whether they should return 409 once it is sampled. - -Unchanged: physical retention and purge (SESSION-CLEANUP-001 remainder), 404 for -reads of deleted Sessions, Artifact retention rules after deletion, managed Runtime -cleanup once deletion is allowed, and caller-owned self-hosted compute, which is -never reclaimed. No schema change. - -Real-PostgreSQL HTTP tests replay D1–D4 across every busy and settled state with -exact bodies, tenant B requests and a whole-database digest proving that a 409 -and a repeated deletion write nothing. Store tests race deletion against Turn and -input admission on one real row lock in both commit orders and concurrently, and -a Worker test cancels a waiting Turn through the daemon protocol before deleting. -Handler, pinned-SDK, TypeScript client and Web unit tests cover the error fields -and the cancel-then-delete flow. Real Core, daemon and model acceptance is -recorded separately by the coordinator. - -## Agent configuration validation — September 23 - -This batch moves protocol validation of Agent configuration into Core with the -official error fields: saved Agent create and update bodies and the inline -`agent` on Session create. Evidence comes from the campaign scan at main -`beb18fd`, recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` -(TV-01..07) with raw official records in `official/validation-{B1,B2A,B2B,B3}.json` -and the Core replay under `core/`. The official probe used one owned Agent and 44 -requests without a model: Agent updates, two Agent creates and nine Session -creates without input on `none`, so no Session or Turn could start. The Agent was -deleted and a read confirmed 404. - -| Row | Case | Core behavior | -| --- | --- | --- | -| C1 | A missing required member, an unknown member, a wrong JSON type or an unsupported enum value in `tools[]`, `text`, `reasoning`, `service_tier`, `multi_agent`, `model`, `name` or `instructions`, including the unpinned `tool_choice` (TV-01) | 400 with type and code `invalid_request_error`, param set to the JSON path (`tools[0].parameters`; `agent.tools[0].parameters` on Session create) and the observed messages: `Missing required parameter: ''.`, `Unknown parameter: ''.`, `Invalid type for '': expected , but got instead.`, `Invalid value: ''. Supported values are: ...` with the pinned literals, and `Invalid '': integer below minimum value. Expected a value >= 1, but got instead.` | -| C2 | Repeated function name, more than one `web_search` or more than one `tool_search` (TV-02) | 400 `invalid_request_error`, param null: `duplicate function tool name: `, `duplicate web_search tool`, `duplicate tool_search tool`. | -| C3 | Function `parameters` or `text.format` json_schema with an explicit string root `type` other than `object` (TV-03) | 400 `invalid_request_error`, param null: `Invalid schema for function '': schema must be a JSON Schema of 'type: "object"', got 'type: ""'.` and `agent.text.format.schema must have top-level type "object"; got ""`, for every harness and before harness admission. | -| C4 | Session create on `none` without input and with an invalid inline agent (TV-04) | The configuration error first. Valid configurations, including enabled `web_search` or programmatic tool calling, still receive the input requirement. | -| K1 | Function names with any characters or over 64 characters, programmatic tool calling enabled on a saved Agent, reasoning effort `max`, service tier `flex` (TV-07) | Unchanged: saved and echoed. | -| K2 | Harness and execution admission limits: enabled `web_search` or programmatic tool calling, structured output on an unqualified harness, explicit reasoning or a non-`auto` service tier on Session create (TV-06) | Unchanged: `unsupported_or_invalid_configuration` with the existing messages, after protocol validation. | -| K3 | Saved `web_search` with mode `live`, `cached`, null or omitted (TV-05) | Resolved by [Saved web_search modes](#saved-web_search-modes--september-23): saved with the official projection; Session admission keeps the K2 rejection. | - -Decisions: - -- A compact validator walks the raw JSON along the pinned shapes - (`PersistedAgentToolParam`/`AgentToolParam`, `AgentTextParam`, - `AgentReasoningParam`, `MultiAgentConfigParam` and the `service_tier` literal) - and reports the first violation through the typed field error from the - validation error batch. It runs before the existing parsers, which keep Core's - local limits and codes, and before harness admission. It is not a JSON Schema - engine: function and output schemas, `request_metadata` values, MCP `transport` - members, `metadata` and `x_agents_core` stay with their existing parsers. -- In each object, a union's `type` is checked first. Unknown members are then - reported in document order, followed by member values in document order and - missing required members in the pinned order. The whole object is checked - before the C2/C3 conflicts, and tools before `text`. The official order - between several errors in one body was not observed. -- Member names match exactly, so a name that differs from a member only by case, - such as `reasoning.Effort`, is an unknown parameter. This is needed because - encoding/json matches names case-insensitively. It also merges repeated - objects into the decoded structs; a key repeated anywhere in the body is now - rejected earlier by the shared body gate with the official message (see - [Request body parsing](#request-body-parsing--september-23)), which replaces - this batch's local `Duplicate parameter: ''.` error. Members left to their - parsers are checked on the decoded values, so they cannot differ from what is - stored. -- Observed expected-kind phrases are `an object`, `a boolean` and - `an object with string keys and unknown value values`. At unsampled positions - Core uses `a string` (also for enum members), `an integer` and `an array`, and - reports a missing Agent create `model` and a non-object Session `agent` in the - same forms. -- Caller-supplied member names, enum values, function names and schema root types - are repeated only when they are at most 256 bytes of printable UTF-8, the - Environment Files rule. Otherwise the error keeps its code and path param, or - a null param for an unknown member, and drops the value: `Unknown parameter.`, - `Invalid value. Supported values are: ...`, `duplicate function tool name`, or - the schema message without the name or `got` clause. -- C3 rejects only an explicit string root type. Schemas without a root type, or - with a non-string `type` such as an array, are unchanged; neither was sampled. - The output schema message names `agent.text.format.schema` on Agent requests as - well, as observed on Agent update; Agent create was not sampled. -- Session admission also applies C2 and C3 to the resolved saved configuration, so - Agents saved before this batch cannot execute with such tools or schemas; - replacing the field in the Session override admits them. An invalid inline - override is reported before the saved-Agent lookup, so owned, foreign and missing - Agents give the same response. Otherwise the lookup order (SES-33) is unchanged. -- Validation of the update body precedes the Agent lookup, so owned, foreign, - missing and malformed Agent IDs give the same response. -- Inline agent validation runs before same-key creation recovery. A same-key - retry of an inline Session created before this batch therefore returns the new - 400 if its original configuration is now invalid, instead of the original - Session. Saved-Agent retries still recover, because the resolved-configuration - checks run after recovery. Core is pre-release, so the order is not changed for - such retries. - -TV-05, saving `web_search` with mode `live`, `cached` or omitted, was deferred -here and is resolved by [Saved web_search modes](#saved-web_search-modes--september-23). -Deferred and unchanged: duplicate `programmatic_tool_calling` -declarations and MCP server labels were not sampled: saved Agents accept them and -Session admission keeps "Execution requires distinct tool controls." and -"Execution requires distinct MCP server labels.". A missing model without -`agent_id`, unknown top-level Session members, `max_concurrent_subagents` above -4294967295, nonblank and 512-byte function names and the 64-function Session bound -keep their local codes. - -Go handler tests cover every C and K row on Agent create and update and on Session -create with and without input, the echo bounds and saved records from before this -batch. A real-PostgreSQL test replays the TV-01..03 rows on Agent create by two -tenants, on updates of owned, foreign, missing and malformed Agents, and on inline -and saved-override Session creates, with a database digest proving no writes; it -then saves and reads back the K1 values and checks tenant isolation. The -pinned-SDK acceptance scripts assert the new codes, params and messages. Real -Core, daemon and model acceptance is recorded separately by the coordinator. - -## Whitespace-only message text — September 23 - -Core admitted user text only when it had a non-whitespace character; the official -service admits any non-empty text and stores it unchanged. Evidence comes from the -campaign scan recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-1/sessions/findings.json` -(SES-01..08) with raw official records under `official/`: `s1-create-string-spaces`, -`s4-create-string-newline-tab`, `s2-create-message-part-newline-tab`, -`e1-events-two-whitespace-messages`, `q2-items-l100` (user Item text `" "`), -`p1a-create-missingagent-empty-string`, `e2-events-empty-content`, -`e3-events-text-emptystring` and `e4-events-empty-input`. The probed Sessions were -deleted. - -| Row | Case | Core behavior | -| --- | --- | --- | -| W1 | Session create with string input `" "` or `"\n\t"` (SES-01/02) | 201; the user Item keeps the exact text. | -| W2 | Session create with a message whose only `input_text` part is whitespace-only (SES-03) | 201; stored verbatim. | -| W3 | `events.create` message with whitespace-only `input_text` parts, including two such messages in one event (SES-04) | 202; one Turn, Items verbatim. String `content` or string `input` stay type errors, as observed officially. | -| W4 | Empty string, empty `content`, empty `input`, or a message whose text parts are all empty (SES-05..07) | Unchanged 400 `invalid_request` with the generic message and null param, without writes. The official responses use code `invalid_request_error`, specific messages and, for the empty create string, param `input`; aligning them is outside this batch. | -| W5 | A message with parts `["", "real text"]` (SES-08) | Unchanged: accepted and stored with the empty part. Official behavior is unobserved; its per-part error message suggests it may reject. | -| W6 | Native execution of a whitespace-only Turn on Codex, Claude SDK and MiniMax Code | Declared per harness through the engine profile. Codex admits and delivers the text unchanged; live acceptance at `898b197a` completed its whitespace-only Turns. Claude SDK and MiniMax Code are not qualified: a message without an image or non-whitespace text returns 400 `unsupported_or_invalid_configuration` at Session creation (including streaming and self-hosted creation) and `events.create`, before any write, reservation or promotion. Live evidence for MiniMax Code: after admission its native runtime refused the prompt with "Local message content or attachments are required." and the Turn failed with `engine_failed` (public `internal_error`). The Claude SDK admission rejection was confirmed live at `d88ffba6`. | - -Decisions: - -- `MessageInput.Validate` treats any non-empty text part as content and no longer - trims. Image reference checks are unchanged. The rule applies wherever the - validator runs: Core admission for create and events, Worker delivery, daemon - steering and prepared start, and the Codex and MiniMax adapters. -- W6 reuses the engine profile that declares image placements: a - `WhitespaceOnlyText` qualification checked with the other input profile rules - during Worker admission, without engine-name branches in handlers. The Claude - bridge and Anthropic-compatible providers reject text blocks without - non-whitespace characters, and the MiniMax Code native runtime refuses such a - prompt, so Core declares both combinations instead of failing the Turn or - rewriting input. Whitespace beside non-whitespace text in one - message stays admitted for every harness. Whitespace is one explicit set, the - union of Go `unicode.IsSpace` and ECMAScript `String.prototype.trim`, used by - both Core admission and the Claude bridge; a shared table test keeps them - equal. The MiniMax native check is covered only by live evidence. - Admission makes the bridge's own check unreachable. If such a steering message - still reached the bridge it would report `input_rejected`, and Core would end - the running Turn as before; the delivery lifetime is unchanged. -- The TypeScript client mirrored the old rule for event batches; it now rejects - only messages whose text is empty. Core Web keeps its local nonblank composer - and Start Session rules; they are a UI choice, not protocol validation. -- No schema, model output or image rule changes. - -Follow-ups: - -- Codex omits the `text` field of an empty text part (`omitempty` on its native - input), so a W5 message `["", "text"]` may be rejected natively. It is recorded - rather than changed here, because removing the tag would also add empty text to - image parts. -- On Claude SDK, a mixed message such as `[" ", "text"]` is admitted and sends - a whitespace-only native text block, and `["", "text"]` sends an empty block; - the provider's behavior for such blocks is unverified. -- The Core Web composer trims leading and trailing whitespace from all sent text, - not only blank sends. This is a UI choice; other clients' text is unchanged. - -Go proto, dispatch, Codex and API handler tests cover W1–W5. Profile, error -mapping and real-PostgreSQL Worker tests cover W6 admission: Codex admits and -stores the text, and Claude SDK and MiniMax Code reject at none, streaming and -self-hosted creation and at events.create without writes. A real-PostgreSQL -test creates Sessions and submits events over HTTP, reads back the exact user -Item text, and proves the W4 rejections write nothing; the pinned-SDK initial -input script asserts the same. Real Core, daemon and model acceptance is recorded -separately by the coordinator. - -## Session input conflicts and result targets — September 23 - -This batch gives every 409 the official conflict type and aligns the conflict and -tool result target errors of `events.create`, from Core main `0035435a`. Evidence -comes from the campaign scans recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-4/errors/findings.json` (ERR-22 and -ERR-27, raw records in `official/results.json`: `sessB-message-while-running`, -`sessA-delete-while-waiting`) and -`~/.parsar/remediation/20260923/campaign-scan-2/events-tools/findings.json` -(EVT-11, EVT-12 and EVT-14, raw records in `official/calls-s2.json`: -`s2-result-unknown-call`, `s2-result-unknown-turn`, -`s2-result-duplicate-after-terminal`, `s2-result-changed-after-terminal` and -`s2-result-after-cancel`). Every observed official 409 has type and code -`conflict_error` and a null param. - -| Row | Case | Core behavior | -| --- | --- | --- | -| CF1 | Every 409 response (ERR-27) | Type `conflict_error`. The code stays specific to the case (CF2–CF5). | -| CF2 | `events.create` input that the Session cannot accept in its current state: a tool result after its Turn was cancelled, or ended without a saved result (EVT-12), and any other Turn conflict on this route; a batch while earlier input still waits for admission, such as the reserved initial input of a provisioning hosted Session or of a self-hosted Session awaiting its connection (ERR-22) | 409 with code `conflict_error` and a null param. Core keeps its message "The Turn cannot accept this input in its current state."; pending input reports "Earlier input to this Session is still pending." (official: "session initial input is still pending"). | -| CF3 | A tool result that differs from the call's saved result, before or after its Turn ends (EVT-12) | 409 `conflict_error`, "The tool call already has a different result." | -| CF4 | Idempotency-Key reuse with a different body on Session creation or `events.create` | Unchanged local code `idempotency_conflict` and message, with type `conflict_error`. Request idempotency is a documented Core extension of these operations. | -| CF5 | Other Core-only conflicts: `sandbox_deployment_conflict`, `runtime_node_in_use`, `runtime_local_node_configured`, `environment_unavailable`, `environment_input_expired`, `environment_input_cancelled`, `runtime_history_unsupported`, and `turn_conflict` from an Environment file write while Session work or input is active | Codes and messages unchanged, with type `conflict_error`. | -| CF6 | A tool result whose `call_id` names no function call of the caller's own Session, with any `turn_id` (EVT-11) | 400 with type and code `invalid_request_error`, param null, "Unknown pending tool call." Nothing is written and the pending action is unchanged. | -| CF7 | A tool result for a call of the Session whose `turn_id` names another Turn, an unknown UUID or no UUID at all (EVT-11) | 400 `invalid_request_error`, param null, "The tool call belongs to a different Turn." Nothing is written. | -| CF8 | Any input to a missing, malformed or foreign Session | Unchanged: the byte-identical 404 `not_found_error`, whatever the result target. | -| CF9 | An identical tool result repeated before or after its Turn ends (EVT-14) | Unchanged: 202 without another application or event. The official repeated `item.added` is not copied. | - -Decisions: - -- The error writer selects type `conflict_error` from the 409 status, so later - conflicts cannot drift. The Session input writer maps Turn conflicts to code - `conflict_error`; other routes keep `turn_conflict`, because official conflicts - there, such as an Environment file write during work, are unsampled. -- Result targets are resolved under the tenant Session lock after the Session - lookup. A well-formed Turn ID is looked up in that Session and must own the call; - otherwise the Session's own calls decide between CF6 and CF7. The `turn_id` is - therefore no longer rejected as a malformed UUID before the Session lookup, and - malformed, missing and foreign Sessions keep one 404. The decision reads only - the caller's Session, so it reveals nothing about other Sessions or tenants. -- The official messages name the call or internal executor IDs ("Unknown pending - tool call: ", "function call exec-... belongs to a different managed - agent turn"). Core's messages are fixed and repeat neither caller input nor - internal identifiers. -- Checks keep their order: request validation, the Session lookup, the retry - lookup (CF4), for batches with a message the Environment file-write gate (a - Turn conflict, also CF2 with the Turn message), the pending input gate (CF2), - then each event in batch order. A - batch sent while input is pending therefore returns the CF2 409 even when its - result target is unknown; the official order between these errors is - unobserved. An empty `turn_id` or a - blank `call_id` remains the generic 400 `invalid_request`. -- The official pending-input sample is the asynchronous admission window of - `none` initial input (ERR-22). Core admits `none` input synchronously and does - not emulate that window; the same fields apply to Core's reserved hosted and - self-hosted input, initial or later. -- The TypeScript client and Core Web did not branch on the old codes. The client - documents that `isSessionDeletionConflict` classifies only a `deleteSession` - failure, since input conflicts now share its code. Cores before this batch - returned 409 `turn_conflict` or `idempotency_conflict` with type - `invalid_request_error`, and 404 for unknown result targets; clients that span - both should treat any 409 as a conflict, and a 400 on new Cores or a 404 on - older Cores as an unknown result target. - -Unchanged: the ERR-22 asynchronous admission window (an architectural difference), -Idempotency-Key semantics, Session-level 404 isolation, admission timing and the -schema. The pinned SDK still retries a 409 by default; the status did not change. - -Go API tests pin the exact CF2–CF9 bodies and the conflict type of every Core-only -409 code. A real-PostgreSQL HTTP test replays CF2–CF4 and CF6–CF9 with tenant B -requests, missing and malformed Sessions, a rolled-back mixed batch, a -whole-database digest and Session reads proving that every rejection writes -nothing and keeps the pending action. Store tests cover target classification and -rollback. The pinned-SDK scripts `official_function_inputs.py`, -`official_pending_actions_native.py` and `official_session_creators.py` assert the -new fields, and TypeScript client and Core Web unit tests cover them. Real Core, -daemon and model acceptance is recorded separately by the coordinator. - -## Saved web_search modes — September 23 - -Saved Agents now keep every pinned `web_search` mode, as the official service -does, from Core main `1eb60c27`. Evidence is TV-05 (official W01/W02) in the -campaign scan recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json`, -and owned probes under `~/.parsar/remediation/20260923/saved-web-search/official/` -(`results.json`, `ledger.jsonl`): four owned Agents with the create records -`type-only`, `mode-null`, `mode-cached` and `mode-cached-full`, the update records -`update-disabled`, `update-omitted-low` and `update-live-domains-empty`, and a -`retrieve`, without a Session or model. All four Agents were deleted. A second -probe created two more Agents, both deleted and confirmed 404 afterwards: -`location-partial-omitted` (`{"city":"Paris","country":"FR"}` saved with null -`region` and `timezone`, `req_db41d2f6261b4abfb69465eafe719ab5`) and `location-empty` -(`{}` saved with all four keys null, `req_165d53b88445490b9146d8272c54134d`). - -| Row | Case | Core behavior | -| --- | --- | --- | -| W1 | Agent create with `web_search` mode `live`, `cached`, null or omitted | 201. Omitted or null mode is saved as `live`, and omitted or null `context_size` as `medium`. `allowed_domains` keeps null versus `[]`. `location` stays null, or else includes `city`, `country`, `region` and `timezone`, with null for omitted keys, also for `{}`. | -| W2 | Agent update replacing tools with these forms, including a return to `disabled` | 200 with the same projection. Retrieve and list return the saved form. | -| W3 | Protocol errors in a `web_search` declaration | Unchanged: the C1 and C2 fields. | -| W4 | Session creation from a saved Agent with enabled search, without a Session `tools` replacement | Unchanged K2 rejection: 400 `unsupported_or_invalid_configuration`, "Only disabled web_search is qualified for execution.", writing no Session, Turn, Environment or reservation, for plain, streamed, self-hosted and hosted creation. Protocol errors and the C4 input requirement still come first. | -| W5 | The same creation with a per-Session `tools` replacement | Admitted as before; the saved search is not used. | -| W6 | Inline Session agent with enabled or omitted-mode search | Unchanged K2 rejection. | -| W7 | Explicit `disabled` search, saved or inline, including records saved before this batch | Unchanged, including the frozen Runtime control on all three harnesses. | -| W8 | Another tenant's `agent_id` | Unchanged: the same 404 as a missing Agent. | - -Decisions: - -- Saved Agents use a separate saved-form parser. Session admission re-resolves the - effective tools with the unchanged execution parser, which admits only disabled - search. This is the path saved enabled `programmatic_tool_calling` already takes - (K1, K2). All creation modes share that admission, and Worker device selection - and the final preclaim still refuse a non-disabled search control, so enabled - search cannot reach dispatch. -- Same-key creation retries keep their rules. A retry recovers the earlier Session - with its frozen disabled control only when that Session recorded its creation - request hash, as current Sessions do. An older Session without that record falls - through to admission and, like a new key, receives the 400. -- An inline agent passes the saved-form parser before execution admission, so its - enabled search is now reported by the execution check, with the same message. - When one configuration hits several execution limits, another limit, such as - explicit reasoning, can be reported first, as for saved Agents. -- Saved tools keep the stored key order of other saved configuration; Session - snapshots keep their own. Clients compare decoded values. -- The TypeScript client types the saved `web_search` declaration with its mode - union; inline execution types are unchanged. Core Web shows saved search as a - read-only tool. Its Session admission check mirrors Core: it starts Sessions - from Agents with one well-formed disabled search and blocks enabled or - omitted-mode search, which Core does not run. - -Unchanged: execution qualification, Runtime controls, the schema and the -configuration validation errors. Enabling search execution needs its own -qualification. - -Go handler tests cover the projection table and W4–W7 on every creation mode. A -real-PostgreSQL HTTP test creates, updates, retrieves and lists Agents with exact -tool bytes, rejects W4 on plain, streamed, self-hosted and hosted creation under a -whole-database digest, checks a same-key retry, admits W5 and checks tenant B. -The pinned-SDK scripts `official_agents.py` and `official_agent_update.py` assert -the saved projections and the admission rejection, and `official_tool_policy.py` -adds saved enabled search to its live rejection cases. Real Core acceptance is -recorded separately by the coordinator. - -## MCP origin and credential selection — September 23 - -The minimal pinned-SDK MCP tool `{type, server_label, transport}` now works on -Core, and Session MCP credential selection projects and reports errors as the -official service does. Evidence is MV-01..03 in the campaign scan recorded -privately in `~/.parsar/remediation/20260923/campaign-scan-6/mcp-vaults/` -(`findings.json`, `REPORT.txt`, `official-ledger.jsonl`): owned Agents, three -owned Sessions and two Vaults with four static-bearer Credentials, all deleted -and read back 404. The error records are `ERR-UNATTACHED` -(`req_b687ac760c03451caa5973be8d65a3ae`), `ERR-URL-MISMATCH` -(`req_90009e0ba2e548ad88a30516faeec852`), `ERR-AMBIGUOUS` -(`req_18b4777d35844be9a5549c8e58fc8747`), `ERR-CREDENTIAL-BOGUS` -(`req_5ba08377d4a841c98849cd4649e6fcaa`) and `ERR-VAULT-BOGUS` -(`req_0274192656b549739951de213e16514a`); the origin default is -`req_4a99a59eba2445c4b9ae22e74667f2b8` and `req_108ecc7c779240528efccd5ad55eebed`. - -| Row | Case | Core behavior | -| --- | --- | --- | -| M1 | HTTP MCP tool with omitted or null `connection_origin`, on a saved Agent, an inline Session agent or a per-Session replacement | Saved and projected as `"service"`. The stored and frozen configuration equals an explicit `service` declaration, so execution is unchanged. Explicit `"environment"` follows the [qualified Environment MCP contract](environments.md#public-mcp-connection-origin); other transports remain rejected. | -| M2 | Session tool without an explicit `credential_id` whose attached credential was selected | Retrieve, list and the created, in-progress and idle event snapshots show the selected credential ID, also after that credential is deleted. Anonymous and unmatched tools stay null; explicit IDs are echoed as sent. | -| M3 | `credential_id` with omitted, null or empty `vault_ids` | 400 `invalid_request_error`, null param: "MCP credential_id requires an attached vault". | -| M4 | `credential_id` not in an attached Vault: missing, foreign tenant, another Vault of the tenant, or malformed | 400 `invalid_request_error`, null param: "MCP credential_id `` was not found in an attached vault". Byte-identical for one ID across the missing, foreign and unattached cases. | -| M5 | Credential in an attached Vault for another URL | 400 `invalid_request_error`, null param: "MCP credential_id `` does not match server_url ``". | -| M6 | Several attached credentials match implicitly | 409 `conflict_error`, null param: "multiple attached vault credentials match MCP server_url ``; specify credential_id". | -| M7 | Unknown or foreign Vault in `vault_ids` | Unchanged 404 `not_found_error`, "Resource not found." (the official message names the ID). | -| M8 | Order and writes | Inline agent protocol errors and the input requirement come first; selection precedes any write, and a rejection writes nothing. | -| M9 | Dispatch | Unchanged: frozen bindings, scoped recheck before decryption, fail-closed on missing keys or decryption, no anonymous fallback. | - -Decisions: - -- `` and `` are the request's values, repeated only under the shared - bounded-echo rule (`internal/echotext`: at most 256 bytes of printable UTF-8); - otherwise the message leaves the value out. -- Selection searches only attached Vaults, which must all belong to the caller. - An explicit ID is found there by ID alone, so a missing, foreign or unattached - ID yields one response, and only a credential of an attached Vault can report - a server_url mismatch. The mismatch message repeats the tool's URL, not the - credential's. -- The projection reads the frozen private binding of the tool's label and URL, - and shows it only while the binding's Vault is among the Session's - attachments. It exposes a credential ID only, never tokens or ciphertext. - Stored configuration keeps the caller's null, so creation retries, recorded - caller intent and dispatch are unchanged. Retries recover the original - projection, also after deletion; a new creation can no longer select a deleted - credential. -- A same-key retry that omits the origin recovers a Session created with the - explicit form when the request has no recorded caller intent; with recorded - intent (attached Vaults or credential references) it remains the local - `idempotency_conflict`, as for any changed request. -- A deleted, previously selected credential is still admitted at later input and - fails at dispatch (MV-04); that remains a separate batch. - -Go tests cover the origin default, the projection and its private-binding -checks, and the typed store errors. A real-PostgreSQL HTTP test with tenants A -and B covers M1–M8, the byte-identical M4 responses with headers, a -whole-database digest over every rejection, M2 across creation, retrieve, list, -creation-stream and live events, deletion and retries. The pinned-SDK scripts -`official_mcp.py` and `official_mcp_credentials.py` assert the omitted origin, -the projection and the error fields. Real Core acceptance is recorded separately -by the coordinator. - -## Hosted initialization failure — September 23 - -Hosted Environments that fail to provision now surface the failure as the -official service does, from Core main `e1970fd7`. Evidence is HI-01..04 in the -campaign scan recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-6/hosted-init/findings.json`, with -raw official records under `official/`: `006-S2-create-setup-exit3`, -`007-S2-events`, `009-S3-events`, `021-S2-env-after-failed`, -`024-S2-session-after-failed`, `025-S3-session-after-failed`, -`026-S2-input-after-failure`, `027-S3-input-after-failure`, `038-S2-delete` and -`041-S3-delete`. Two owned `openai_hosted` Sessions without a Turn failed, one on a -setup command that echoed a value and exited 3, one on a nonexistent Python -package; both were deleted. The batch plan is -`~/.parsar/remediation/20260924/hosted-init-failure/PLAN.md`. - -| Row | Case | Core behavior | -| --- | --- | --- | -| H1 | Hosted initialization fails in any step | One transaction records the Environment failure, `agent.session.environment.failed`, `error` and one `agent.session.failed`. The Session reads `status: failed`, the stored safe reason as `error`, `required_actions: []` and the failure time as `last_active_at`; retrieve, list and the event snapshot agree. Pending input reserved for the Environment settles as failed exactly as before, captured in the same snapshot. | -| H2 | `environment.failed` payload | `error` is `{type: environment_error, code: environment_connection_failed, message: "The environment failed to connect."}`. | -| H3 | `error` event | `{type: environment_error, code: sandbox_error, message: , param: null}`. | -| H4 | Reason | `Failed to provision environment: script "setup_commands[i]" failed with exit code N`, and `script "Python package installation"` for Python packages; the official Python reason also appends raw pip output, which Core never copies. npm, system package, initial file and Skill labels are unverified. Other failures use `Failed to provision environment: initialization did not complete`; see the [initialization lifecycle](environments.md#initialization-state-and-failure). | -| H5 | Live SSE | GET and creation streams end right after that `agent.session.failed`. | -| H6 | Later `events.create` | 409 `conflict_error`/`conflict_error` "the hosted environment failed to provision", param null. Expired Environments, and input already waiting when the Environment failed, keep 409 `environment_unavailable`. | -| H7 | Delete | 200 `agent.session.deleted`, as officially, then 404. Deletion while provisioning is unchanged (HI-05 awaits a decision). | -| H8 | Unchanged | `self_hosted` and `none` Environments, successful initialization and its timing, the two-minute step limit (HI-06), expiry and tenant isolation. | - -Historical implementation decisions at the September 23 revision follow. The -current daemon uses shared Go preparation and process settlement, with no bwrap, -Python receipt decoder or old-image compatibility. System packages now reject -before initialization. Safe fixed labels and bounded exit statuses remain the -current public failure policy; see the [current initialization contract](environments.md#runtime-capability-preparation). - -Recorded decisions: - -- **No output.** The shared Runtime initializer adds only an integer `exit_code` - to its failed receipt, and only for a step run inside its bwrap isolation; - Runtime helpers, signals and other errors keep the generic receipt, and the - exception is never serialized. Core confirms a failed step only when the - process exits 1 with empty stderr and a version-1 `failed` receipt. The - decoder is deliberately lenient for older images and ignores other fields; the - only value ever taken from the receipt is `exit_code`, and only as an integer - from 1 to 255. The Store composes the reason from a fixed step label and - integers, so commands, env values, package names, paths and process output - cannot reach the reason, events, logs or responses. -- **Storage.** The additive migration `000062_environment_failure.sql` adds the - nullable `environments.failure_reason` and `failed_at`; a check ties them to - `status = failed` and bounds the reason to 256 characters. Environments that - failed earlier keep NULL and their previous projection and events; new input - on them gets the H6 409. -- **Runtime images.** A Runtime image built before this change reports no - `exit_code`; its failures use the generic reason and otherwise follow H1–H7. - The Codex, Claude and MiniMax Code images must be rebuilt for exit statuses. -- **Scope of the terminal state.** Only a recorded hosted provisioning failure - makes the Session terminal. A Turn failure still leaves GET streams open, and - a GET stream opened after the failure stays open; that case was not observed. -- **Clients.** Official-shaped `error` events carry `param: null`; Core's own - `stream_interrupted` frame keeps its three-field error without `param`, so - released clients still report it as an interruption. The TypeScript client - delivers error events other than Core's `stream_interrupted` to `onEvent` - before the failed snapshot, instead of raising them. Core Web already renders - the failed Session, its error and the blocked input. - -Go store tests on a dedicated PostgreSQL database drive the managed Worker with a -controlled Provider through setup exit statuses (first and later command), -Python packages, a receipt without `exit_code`, an unknown effect, raw output -instead of a receipt and a failed initial file write. They check the Session -read, list, the exact three events and snapshot, the H6 rejection, pending-input -settlement, tenant B and the absence of a canary. A real-PostgreSQL HTTP test -checks retrieve, list, the live GET stream and its end, the exact 409, tenant B -404s, delete and the canary in every body. Go API and contract tests pin the -projection, stream lifetime, wire shapes and error mapping; Python tests pin the -initializer receipt, and the TypeScript client and Web unit tests pass. Live -Docker acceptance is recorded separately by the coordinator. - -## HTTP routing and response headers — September 23 - -This batch aligns path handling, the Beta check, 401 envelopes and response -headers with campaign scan 6 at Core main `1eb60c27`, recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-6/http-protocol/` (`findings.json` -HP-02..24, raw `official-ledger.jsonl` and `REPORT.txt`): 90 official requests on -one owned Agent, deleted afterwards, without a Session or model. Labels below are -ledger records. - -| Row | Case | Core behavior | -| --- | --- | --- | -| RH1 | `//`, `.` or `..` path segments (HP-17: `R10`, `R11`, `R17`, `R18`) | Served on the canonical path, never redirected. Empty and dot segments resolve with ServeMux semantics and a trailing slash is kept, so `/v1/agents/x/../` still reaches the trailing-slash 404. The former 301 made the pinned SDK resend an update as a GET and drop it. | -| RH2 | A percent-encoded unreserved character in the path (HP-18: `R12`) | Decoded before routing, including `%2E` dot segments. The canonical path is built from the request's own path spelling; bytes that are invalid in an escaped path (such as `{`, `"`, a backslash or non-ASCII) are percent-encoded first. Other escapes, such as `%2F`, `%2f`, `%5C` and double encodings, stay encoded and never separate segments. Malformed, missing and foreign IDs keep the single 404. | -| RH3 | HEAD on a GET route (HP-19: `R13`, `R16`) | The GET route runs after the same Beta and authentication checks; 200 with its headers and no body. `Content-Length` is present when Go buffers the whole body (about 2 KiB) and omitted for larger responses. The events stream, the File, Skill, Skill version and Artifact content downloads, and the live Environment Files directory list answer HEAD with Core's 405 instead, so HEAD never holds a stream open or reads content. That exclusion is a documented Core difference; official HEAD on those routes is unobserved. | -| RH4 | Unsupported method (HP-20: `R04`–`R06`) | Unchanged 405 JSON `unsupported_operation`, now with `Allow` listing the route's methods in the observed order, such as `GET,HEAD,POST,DELETE`. Routes outside the Beta group, such as `/healthz` and executor credentials, and methods chi does not know, such as `FOO`, now use the same JSON 405 instead of chi's empty one; an unknown method is answered before the Beta and authentication checks, as before. | -| RH5 | `X-Request-Id` (HP-23) | Every response of the Agents API handler, including 400, 401, 404, 405 and SSE streams, carries a fresh random `req_` plus 32 lowercase hex characters. The ID is attached to the request log context as `request_id` next to the trace carrier. | -| RH6 | No or invalid credentials without OpenAI-Beta on a Beta route (HP-05: `A06`, `A11`) | 400 `invalid_beta`: the constant Beta check now precedes authentication. Files, Skills and Core project extensions still ignore the header. | -| RH7 | Repeated OpenAI-Beta header lines (HP-03: `B08`) | 400 `invalid_beta` unless there is exactly one field value, equal to `agents=v1`. | -| RH8 | 401 (HP-07: `A01`–`A04`, `A07`–`A10`) | Type `invalid_request_error`. Beta routes report a null code for every failure. Files, Skills and Core project extensions report a null code without a Bearer credential (missing, other scheme, empty or repeated header) and `invalid_api_key` for a rejected one, including mismatched scope headers. Core's message and `WWW-Authenticate: Bearer` are kept. | -| RH9 | `invalid_beta` message (HP-02: `B01`) | "To access the Agents API, set the 'OpenAI-Beta' header to 'agents=v1'." | -| RH10 | Optional headers (HP-24) | `OpenAI-Version: 2020-10-01`, `OpenAI-Processing-Ms` and `X-Content-Type-Options: nosniff`. Organization and project headers are not reported: Core's project scope is configured, not account-derived. | -| RH11 | Trailing slashes and unknown sub-routes (HP-21), OPTIONS and CORS (HP-22), `Cache-Control` and `traceparent` (HP-26), `agents=v0` (HP-04) | Unchanged: 404 JSON, no CORS handling, Core's extension headers kept, `agents=v0` still rejected (an upstream anomaly, not copied). | - -Decisions: - -- One canonicalizing handler wraps the complete server handler in both server - configurations: around the ServeMux that also serves daemon, enrollment and node - transport, and around the API router when it is served alone. The ServeMux, the - router, every middleware, authentication check and handler see only the - rewritten path. A dirty or encoded path therefore reaches exactly the route - group and authentication of its canonical path written literally; internal - daemon, node and sandbox routes keep their own authentication. The input is the - request's own path spelling (`RawPath` when Go keeps one), never a path - re-escaped from its decoded form, which would turn `%2F` into a separator when - the spelling holds a byte Go considers invalid. `Path` and `RawPath` are then - set consistently, so chi, which prefers `RawPath`, and the ServeMux, which uses - `EscapedPath`, route on the same string. Decoding only unreserved characters is - RFC 3986 normalization, so a proxy that normalizes URIs the same way sees the - same route. The ServeMux still redirects the exact daemon prefix - `/api/v1/agent-daemon` to `/api/v1/agent-daemon/`; that is daemon transport, not - an Agents API path. -- The Beta check reads only a constant header and returns no tenant or resource - data. Moving it first changes only responses that were rejected either way: - every request that passes it is authenticated before the router reaches any - Beta handler, 404 or 405. -- Wrong methods and unknown sub-routes below `/v1/files` and `/v1/skills` still - reach the Beta group's 404 and 405 after its checks, as before; without the Beta - header they now report `invalid_beta` instead of 401. Official behavior there is - unobserved. -- Every 401 of the Agents API handler has type `invalid_request_error`, including - the deployment administrator (`invalid_admin_key`) and sandbox node - (`invalid_node_credential`) extensions, whose codes are unchanged. Project API - key management keeps its deployment administrator authentication behind the - same canonical path; derived project keys authenticate exactly as their static - parent binding, under the Beta and Files rules above. Daemon, - enrollment and node transport served beside it keep their own formats. -- The response headers belong to the Agents API handler. Daemon, enrollment and - node transport routes do not carry them, and the shared log middleware is - unchanged. A caller-supplied request ID is not echoed; that header is not pinned. -- Core Web recognizes both the current 401 envelope and the older - `invalid_api_key` code in its connection probe. The TypeScript client already - exposes status, type and code without branching on them. The Parsar product - repository has no code branching on these 401 fields, and its Go client refuses - redirects. - -Go handler tests cover RH1–RH11, including a walk over every registered route: -unauthenticated requests, with and without the Beta header and with foreign -credentials, are rejected before any handler, and ten dirty and encoded spellings -of each path, including traversal from the daemon and sandbox prefixes, give the -clean path's exact response. Raw request-line tests over a real listener cover -invalid bytes, non-ASCII, `%2F`, `%2f`, `%5C`, double encoding and absolute-form -URIs in both server configurations, and two fuzz targets assert that any request -path reaches the same handler, route group and response as its canonical form, -with chi, the ServeMux and `Path` agreeing on it. A server test replays the -daemon-enabled composition and checks that no Agents API request is redirected. The pinned-SDK script -`official_http_routing.py` updates an Agent through base URL `/v1//`, checks -`_request_id` and the error `request_id`, and the raw checks in the other official -scripts now expect the Beta check first. - -## Request body parsing — September 23 - -Every Agents API JSON request body now passes one shared gate, before any -route-specific decoding, validation or lookup, with the official parse semantics. -Evidence is HP-09..HP-15 of the HTTP protocol campaign scan, recorded privately in -`~/.parsar/remediation/20260923/campaign-scan-6/http-protocol/` (`findings.json`, -`REPORT.txt`, raw requests in `official-ledger.jsonl`, labels `C01`–`C20`, -`U01`–`U11`, `S1`): one owned Agent, created, updated and deleted (404 confirmed), -without a Session or model. The official records cover Agent create and update; -the other routes are assumed to share the official parser. Two later owned probes -in `~/.parsar/remediation/20260924/http-json-body/official/results.json` created -nothing: a lone high surrogate escape (`req_1a9b7680d615454ca97c816b25e2f401`) and -Agent create with `metadata` and `Metadata` but no `model` -(`req_6ba2a50c71a4410f87a1baac855e82df`). - -| Row | Case | Core behavior | -| --- | --- | --- | -| B1 | Malformed JSON, trailing data, two concatenated values, a UTF-8 byte order mark, a whitespace-only body (`C01`, `C06`, `C07`, `C12`, `C20`, `U01`, `U05`, `U06`), or a string escape that forms a lone or mis-paired UTF-16 surrogate, such as `"\ud800"`, in a key or value (`req_1a9b7680d615454ca97c816b25e2f401`) | 400, type and code `invalid_request_error`, param null: "Invalid body: failed to parse JSON value. Please check the value to ensure it is valid JSON. (Common errors include trailing commas, missing closing brackets, missing quotation marks, etc.)". Agent update no longer returns `unsupported_or_invalid_configuration`. Valid surrogate pairs are accepted; lone surrogates were previously stored as U+FFFD. | -| B2 | Invalid UTF-8 anywhere in the body (`C13`) | 400 with the same fields: "Invalid body: encountered a unicode decode error when parsing this JSON value. Please check the value to ensure it is valid unicode." Previously the bytes were stored as U+FFFD. | -| B3 | A repeated object key at any depth (`C08`, `C15`, `C16`, `U07`) | 400 with the same fields: "Invalid body: duplicate JSON key '' at ''. Duplicate JSON keys are not supported." The path joins object keys with `.` and omits array indices: `name`, `metadata.k`, `tools.type`. Keys compare after unescaping and case-sensitively: `metadata` and `Metadata` are distinct keys (`req_6ba2a50c71a4410f87a1baac855e82df`). The first repeat in document order is reported. Previously metadata, Vaults and Templates kept the last value, and Agent configuration returned the local "Duplicate parameter". | -| B4 | A valid root that is not an object (`C05`) | 400 with the same fields: "Invalid type: expected an object, but got instead." with the existing kind phrases (`a string`, `an integer`, `a number`, `a boolean`). | -| B5 | A zero-length body or `null` (`C02`, `C03`, `U02`, `U03`) | Treated as `{}`: Agent create reports the missing `model`, Agent update is the documented empty update, Session update keeps its "At least one update field is required" rejection, and Vault create creates an unnamed Vault. A whitespace-only body stays B1. | -| B6 | Content-Type missing, `text/plain` or form-encoded, including a bodyless POST without Content-Type (`C09`–`C11`, `C19`, `U08`–`U10`) | 400 with the same fields, "expected request with Content-Type: application/json", checked before the body is read. `application/json` and `application/*+json` are accepted case-insensitively, with parameters (`S1`, `U11`, `C17`, `C18`); a malformed media type, such as `application/foo bar+json` or a conflicting repeated parameter, is rejected the same way. Previously Core ignored the header and applied the update. | -| B7 | Valid bodies | Unchanged, including unknown-member errors, configuration validation, route body limits (413) and Core extensions such as `x_agents_core`. | - -Order: authentication and Beta handling as before, then B6, then the route's body -limit, then B2, B1, B3 and B4/B5, then route validation. The gate covers Agent -create and update, Vault create, Credential create and update, Template create -and update, Environment Files create, Session create and update, and Session -events. Environment Files create now reads its body before the Environment -lookup; a missing Environment with a valid body still returns 404. - -Decisions: - -- An array root is rejected with the B4 message. The official service treats `[]` - as `{}` (HP-14); that upstream anomaly is not copied. -- The duplicate key and path are repeated only when each is at most 256 bytes of - printable UTF-8, the shared `echotext` rule; otherwise the message is "Invalid - body: duplicate JSON key. Duplicate JSON keys are not supported." -- The gate scans each body once in linear time and keeps key positions in the - body, not copies: objects with more than 16 keys use an open-addressing set of - 8-byte slots, a position and 32 hash bits that skip comparing unequal keys. A - body of many short keys allocates about twice its size in the gate. Bodies are - read into a doubling buffer, which allocates two to four times the body in - total (four near the route limit), against 4.4 to 6.1 times for `io.ReadAll`. -- Member names match exactly on every gated route. encoding/json would match a - case variant such as `Metadata`, `Input` or a nested `Role` to the field and - let the last copy win; such a key is now an unknown member at any depth, - rejected with the route's existing unknown-member error before any write. - Agent configuration already did this. On Agent create, a body with `Metadata` - and no `model` still reports the unknown member first, while the official - service reported the missing `model`; the official order between several - errors in one body remains unaligned. -- A walk over the body bytes checks member names before a decoder runs and stops - at the first unknown or case-variant key, and the Session metadata check reads - only the `metadata` member. For unknown top-level keys this makes rejection - cheap: a whole 16 MiB Session create allocates about 96 MiB and events about - six times a 1 MiB body, instead of 482 MiB and 9 MiB before this batch, when - the decoder formatted an error for every unknown key. Unknown keys nested under - `agent` or `environment` still pass the existing object decoding of - `decodeInputObject` and stay linear, at or below main: about 779 and 871 MiB - for 16 MiB bodies, against 850 and 889 MiB on main. -- Invalid UTF-8 is checked before JSON syntax; the official order for a body with - both faults was not observed. -- Not gated: DELETE routes, which keep their empty-body rule, the multipart Files - and Skills uploads, Skills update (a non-Beta API with its own observed error - fields), the Core extension `/core/v1/*` routes, including executor credential - issuance, and the internal daemon, sandbox and node routes. - -Unchanged: the schema and the route validation and error codes of valid bodies. - -Caller check: the TypeScript client sends `application/json` with every JSON body -(an empty Agent update sends `{}`), Core Web uses that client, the Go client uses -the pinned openai-go SDK, the Python acceptance tools send JSON through the pinned -SDK or `json=`, and the documentation has no JSON POST examples. The Parsar -product repository does not call these routes. - -Go tests cover the gate on its own (B1–B5, surrogate escapes, the echo bound, -media type parsing, deep bodies, a differential check of the duplicate-key scan -against an encoding/json token walk on small and large objects, and an allocation -bound on many short keys), case-variant members, route-level allocation bounds for -unknown keys on Session create and events, and all eleven route families -(B1–B4, B6, the order against authentication, Beta and the body limit, B5/B7). A -real-PostgreSQL test sends B1–B4, B6 and case-variant members to every route -family as the owner and as tenant B under a whole-database digest, then checks -B5/B7 writes; another exercises the excluded DELETE, Files and Skills upload, -Skills update and executor credential routes with real storage. The pinned-SDK -scripts `official_agents.py`, `official_agent_update.py`, `official_vaults.py`, -`official_credentials.py`, `official_credential_rotation.py` and -`official_session_metadata.py` assert the official messages and recount or reread -the resources to show no writes; `official_session_requests.py` rejects -case-variant members without creating a Session. -Independent acceptance is recorded separately by the coordinator. diff --git a/contracts/agents-api/operation-evidence.md b/contracts/agents-api/operation-evidence.md deleted file mode 100644 index bd59162b9..000000000 --- a/contracts/agents-api/operation-evidence.md +++ /dev/null @@ -1,202 +0,0 @@ -# Pinned operation evidence inventory — 2026-09-23 - -Baseline inventory of main `b5715912f09333e2b4449ec6f0eecaabce44c9b7`. The Session admission batch below updates creation and metadata validation, the list query tolerance batch (L) updates list and resource query handling, the validation error batch (X) updates field error codes/params, malformed path IDs and U+0000 handling, the Environment Files wire batch (G) updates Files.create/list status, envelope, query, path and empty-page behavior, and the creation stream settlement batch (J) updates creation SSE lifetime/snapshot, terminal Turn usage and Turn start order, and the Session deletion batch (Z) updates the deletion lifecycle, and the Agent configuration validation batch (M) updates saved and inline Agent configuration errors, the whitespace input batch (P) admits whitespace-only message text, the list cursor error batch (CE) updates unresolved `after` cursor errors on every list, the input conflict batch (CF) gives every 409 type `conflict_error` and aligns Session input conflicts and tool result target errors, and the saved web_search batch (SW) saves every pinned `web_search` mode while Session admission keeps rejecting enabled search, and the workspace file write batch (FW) aligns Files.create parent creation, no-replacement and the inline size bound, and the item serialization batch (SR) aligns Item/event null fields, assistant message event framing, reasoning keys and the Session usage rule, and the MCP credential selection batch (MV) saves an omitted HTTP MCP origin as `service`, projects the implicitly selected Session credential and aligns selection errors, and the hosted initialization failure batch (HF) aligns failed hosted provisioning: Session status/error, failure events, stream end and later-input 409, and the HTTP routing and header batch (RH) serves canonical paths without redirects, checks OpenAI-Beta before authentication and aligns 401 envelopes, HEAD, `Allow` and request ID headers on every operation, and the JSON body batch (JB) checks every Agents API JSON request body in one shared gate; historical evidence retains its original revision and scope. This inventory guides repeated qualification and does not assert complete compatibility. - -Baseline: `contracts/agents-api/upstream.json`, SDK **3.13.0**, upstream commit **d7c41efee1b0802b79f3f88a678ef2052b06e9ce**, `OpenAI-Beta: agents=v1`. AGENTS.md and relevant CONTRIBUTING.md compatibility, ownership and evidence rules govern this inventory. - -## Scope and evidence interpretation - -There are **58 distinct HTTP operations: 42 beta/agents + 5 general Files + 11 Skills/version/content**. Every one has a registered Core route; every family retains partial or unverified semantics. This is an exhaustive operation inventory, not an exhaustive compatibility claim. Methods/paths were enumerated from the complete installed fixed SDK, not from Core OpenAPI; the appendix gives exact source lines. The installed beta subtree matches the cached pinned commit byte-for-byte. The unrelated experiments virtualenv contains SDK 3.14.0 and was excluded. - -- **Implemented** means current source supports the stated bounded behavior. A route or SDK parse does not establish matching server behavior. -- **Official wire** means retained raw official-service observations, confined to their cases. `None located` means no positive wire observation was found in the specified evidence set, not that no such evidence exists anywhere. -- **Core validation** distinguishes **DB** (actual HTTP/PostgreSQL resource checks, no model execution), **Live** (service → daemon → native harness → real model), and **Recorded** (historical acceptance documented in the repository; underlying remote artifacts were not independently replayed/read in this inventory). -- **Gap** distinguishes known local restrictions/differences from unknown upstream semantics. A native limitation is not a redefinition of the fixed protocol. Historical runs are not fresh validation of b571591. - -## Evidence register - -Repository paths below are relative to the inspected worktree; private evidence paths are explicit. These locators retain scope and original revisions rather than claiming all older tests were rerun. - -| ID | Exact source and evidentiary boundary | -| --- | --- | -| R | `~/.parsar/remediation/20260922/official-semantics-alignment/resources/raw-evidence.json` and sibling `REPORT.md`, `request-index.md`: 40 official raw requests on six owned Agent/Vault/Credential/Template resources. Row labels below identify records; repeated cleanup labels additionally require the resource path. No model/hosted-environment execution. Historical source-derived mismatches in that report must be reconciled with the merged fixes, not copied as current bugs. | -| S | `~/.parsar/remediation/20260922/official-semantics-alignment/sessions/REPORT.md` plus named sibling JSON files: official `none` text, metadata, reads, pages, errors and cleanup. Final follow-up totals are four owned Sessions/five tiny gpt-6-astra Turns, including duplicate creation on retry; the earlier report's two-Session/three-Turn scope was expanded explicitly. | -| H | `contracts/agents-api/history-events-usage.md:85`; raw directory `~/.parsar/remediation/20260922/history-events-usage/official/`, including `create-http.json`, `create-sse-frames.json`, `reconnect-events.json`, `turns-asc-pagination.json`, `items-asc-pagination.json`, `analysis.json`. Official one-owned-Session/two-Turn text observation; no tools/Subagents/hosted execution, no full timing or accounting proof. | -| E | `~/.parsar/remediation/20260922/official-semantics-alignment/error-surface-probe.json`: missing non-beta File/Skill observations only. Not successful-resource or full error-surface coverage. | -| C | `~/.parsar/remediation/20260922/official-semantics-alignment/live/acceptance-summary.json`, source `7c80d604ba47c578084ebc9fa99cff332bd5d5f2`. Seven resource/safety groups (DB), plus three real Turns per harness (Live): Codex/Claude Kimi K3; MiniMax M2.7. Successful runs: `live/resources/run-1790091139881592673`, `live/codex/run-1790091781322707360`, `live/claude_sdk/run-1790091781322554888`, `live/mcode/run-1790091781324544033`; each has `result.json` and `raw-evidence.json`, model runs also history/SSE evidence. Four network-interrupted attempts remain failures. No new native capability/OAuth refresh/E2B qualification. | -| N | `contracts/agents-api/official-semantics-alignment.md`, September 23 admission section; private `~/.parsar/remediation/20260923/session-admission-alignment/{official,live}/`: conditional create/update official probes plus exact-source `7ccc636` Core/daemon real none execution (six Turns), write rejection, local retry, metadata and isolation. Qualified idle hosted admission and Codex self-hosted admission are not new execution profiles. | -| Q | [List query semantics](list-query-semantics.md); private `~/.parsar/remediation/20260923/error-query-survey/REPORT.md` and `request-index.md`: 35 read-only official GETs, nine empty-order rejections and family-specific error fields. Missing-parent/cursor cases do not qualify later numeric bounds. No model execution or resource creation. Core acceptance is recorded in the linked batch document when complete. | -| A | `contracts/agents-api/official-semantics-alignment.md`: merged status/envelope/no-op/error/rejection changes and explicit remaining differences. `contracts/agents-api/README.md:63` supplies the current resource ledger; it is a coverage summary, not raw evidence. | -| T | `contracts/agents-api/execution-tools.md`: per-operation input, function, required-action, structured-output, discovery and policy matrix; evidence register M1/M2/F1/F2/F3/S1/S2/D1/P1/E1/E2 gives exact private run paths. These are profile-specific real acceptances. The older blanket OAuth rejection in this document is superseded by O. | -| D | `contracts/agents-api/README.md:115`, `contracts/agents-api/user-managed-runtime-v1.md`: recorded three-harness Docker MVP and separate user-managed Docker/E2B qualification. Historical `~/.parsar/remediation/20260920/three-harness-mvp/REPORT.md`, `acceptance-results.json` on zju_a100_2; prior Core-managed E2B qualification is retired-route evidence, not current enrollment qualification. | -| B | `contracts/agents-api/subagents.md:109`: recorded six-GET real Docker matrix, two children, identity/pagination/isolation, cancellation and cold same-child continuation. Remote `~/.parsar/remediation/20260922/subagent-contract/{public-codex-kimi1,public-claude_sdk-2,public-mcode-1}` on zju_a100_2. Codex Kimi K3, Claude Kimi K3, MiniMax M2.7. `resources_passed`/`requested_phase_passed` qualify common reads; aggregate `passed` additionally requires optional Codex close/reopen. | -| F | `contracts/agents-api/source-files.md:11,65`: implemented general Files workflow and recorded PostgreSQL/SDK/raw-list checks; `contracts/agents-api/environment-files.md:1,30,71`: bounded local workspace list/copy and real deployment references. D and K add real consumption/copy/retention workflows. | -| K | `contracts/agents-api/environment-templates.md:94,764`: recorded fixed-SDK/raw/PostgreSQL Skill resource checks plus all-three-harness Docker reference workflows (Codex/Claude Kimi K3, MiniMax M2.7). Remote `~/.parsar/remediation/20260921/template-skill-references/`; local index `~/.parsar/remediation/20260921/template-capabilities-design/`. Covers upload, frozen template/Session resolution, native supporting files, source/template deletion, retry and cold continuation. Individual resource mutation cases not explicitly claimed by the summary remain DB-only or unspecified. | -| I | `contracts/agents-api/environment-templates.md`: separate recorded initial-file (:589), env/setup/npm/Python (:621), inline-Skill (:657), system-package (:692), capability-directory, Plugin MCP (:241), composition (:327) and restricted-network (:523) qualification. Exact historical run roots are in each section. No combinatorial/full hosted parity inference. | -| O | `services/core/oauth-credentials.md:144`: recorded genuine Keycloak 26.7.4 PKCE grants, real TLS MCP and Kimi execution. Codex Basic client-auth initial/refresh/restart/replacement/revocation/delete; Claude POST initial/refresh. Trusted `none` profiles only; not MiniMax, hosted, arbitrary-provider or complete error equivalence. | -| V | [Resource selector and error qualification](resource-selector-semantics.md); private September23 `skill-version-alignment/official/` and `files-error-alignment/official-probe.json`: owned Skill/Template/Session selector observations and seven missing File/cursor requests. Preserve each probe phase and distinguish actual hosted execution from metadata-only reads. | -| W | `~/.parsar/remediation/20260923/file-resource-semantics/official-files/` and `official-skills/`: 56 owned resource requests, no Sessions/models; [qualified observations and limitations](file-resource-semantics.md). | -| X | [Validation error fields](official-semantics-alignment.md#validation-error-fields--september-23); private `~/.parsar/remediation/20260923/campaign-scan-1/{vaults-agents,sessions,skills-files-templates}/findings.json` VA-07/08/09/10, SES-28, SFT-20 and the September 22 Session `metadata.a` null observation. Metadata/name field errors, U+0000 local limit, malformed path IDs and Template network codes; Go handler and real-PostgreSQL route/no-write tests, no model execution. | -| J | [Creation stream settlement](history-events-usage.md#creation-stream-settlement-2026-09-23); private `~/.parsar/remediation/20260923/campaign-scan-2/events-tools/findings.json` EVT-01..04 with raw frames under `official/`: four owned official `none` Sessions (text, structured output, function result, cancellation), 1091 events validated against the pinned types. Core-side changes carry API/store/contract DB tests; live acceptance is recorded with the batch. EVT-05..13 and EVT-19 stay deferred or record-only. | -| L | [List query tolerance](list-query-semantics.md#list-query-tolerance--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-1/{vaults-agents,sessions,skills-files-templates}/findings.json`: owned-collection unknown/repeated keys, limit bounds, Vault status union and Files empty purpose, plus unknown keys on a deleted Vault read and Agent delete. Rows A1–D2 of that section; no model execution. | -| G | [Environment Files wire alignment](environment-files.md#wire-alignment--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/findings.json` HE-10, 16, 18, 32, 34–39 with raw records under `official/` (labels `fc01`–`fc17`, `fl01`–`fl16`, `files-*-pending`) and `run1/`: three owned hosted Sessions, all deleted; first official Environment Files observations. Rows F1–F9 of that section. Go handler, real-PostgreSQL Worker, gateway/daemon and Rust helper tests without a model; live acceptance is recorded with the batch. | -| M | [Agent configuration validation](official-semantics-alignment.md#agent-configuration-validation--september-23); private `~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` TV-01..07 with raw records under `official/validation-{B1,B2A,B2B,B3}.json`: one owned Agent (deleted) and 44 requests without a model, Session creates without input on `none`. Rows C1–C4 and K1–K3 of that section. Go handler, real-PostgreSQL no-write/tenant replay and pinned-SDK tests without a model; live acceptance is recorded with the batch. | -| U | [Subagent visibility](subagents.md#subagent-visibility); private `~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` SAT-01, 02, 07, 08, 09 (SAT-03 for the kept limit rejections) with raw records under `official/` (labels `R00`–`R33`, `C01`–`C06`, S1/S2 stream frames): two owned official Sessions with two child Turns, all deleted; the first official Subagent observations. Rows A1–A5 of that section; SAT-03/05/13/14 match, SAT-04/06/10 and SAT-12 stay deferred or unknown. Go store/API and real-PostgreSQL HTTP tests, TypeScript client and Core Web tests; live acceptance is recorded with the batch. | -| P | [Whitespace-only message text](official-semantics-alignment.md#whitespace-only-message-text--september-23); private `~/.parsar/remediation/20260923/campaign-scan-1/sessions/findings.json` SES-01..08 with raw records under `official/` (`s1-create-string-spaces`, `s4-create-string-newline-tab`, `s2-create-message-part-newline-tab`, `e1-events-two-whitespace-messages`, `q2-items-l100`, `p1a`, `e2`–`e4`). Rows W1–W6 of that section. Proto, dispatch, Codex, API handler, TypeScript client and real-PostgreSQL HTTP/Worker tests without a model, including Claude SDK and MiniMax Code admission rejection; live acceptance completed Codex whitespace-only Turns, and the MiniMax native refusal is recorded in the linked section. | -| Y | [Artifact capture and listing](official-semantics-alignment.md#artifact-capture-and-listing--september-23); private `~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/findings.json` HE-50..62 with raw records under `official/` (labels `al01`–`al09`, `ar01`–`ar04`, `ac01`–`ac05`, `ad01`/`ad02`): three owned Sessions and two tiny Turns, all deleted; first official Artifact observations. Rows A1–A4 of that section: symlink skip, republication, list envelope and malformed filter. Rust link tests, real-PostgreSQL store/HTTP and pinned-SDK tests without a model; live acceptance is recorded with the batch. | -| CE | [List cursor errors](list-query-semantics.md#list-cursor-errors--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-4/errors/findings.json` ERR-01..06 (ERR-07..13 and ERR-17 match or record upstream failures) with raw records labelled `cur-*` in `official/results.json`, plus SAT-04 (campaign scan 3) and HE-57 (campaign scan 2): owned Agent, Session, Turn, Item, Subagent, Artifact, Template, Vault, Credential, File, Skill and Skill version cursors without a model. Rows C1–C5 and K1–K3 of that section. Go store, real-PostgreSQL HTTP tenant A/B and pinned-SDK tests; independent acceptance is recorded with the batch. | -| Z | [Session deletion lifecycle](official-semantics-alignment.md#session-deletion-lifecycle--september-23); private `~/.parsar/remediation/20260923/campaign-scan-1/sessions/findings.json` SES-29/30 with raw records under `official/` (`q5-delete-repeat`, `q5b-delete-while-in-progress`, `q5-delete-never-existed`, `q5-delete-while-running`, `q5-get-after-delete`). Rows D1–D5 of that section. Handler, real-PostgreSQL HTTP/store/lock-race, Worker, pinned-SDK, TypeScript client and Web tests without a model; live acceptance is recorded with the batch. | -| FW | [Environment Files write semantics](environment-files.md#write-semantics--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/findings.json` HE-11, 12, 13 and 15 with raw records under `official/` (`fc02-nested-missing-parents`, `fc03-overwrite`, `fc06-5mib-plus-1`, `fc15-onto-directory`, `fc20-overwrite-untracked`, `fc21-through-symlink-dir`, `fc22-symlink-outside`) and the saved guide file limits. Rows FW1–FW11 of that section. Rust installer and race tests, daemon writer tests (the native-helper case is opt-in and was run by hand and by live Docker acceptance), gateway/dispatch, Go handler and real-PostgreSQL HTTP/Worker tests with tenant B, TypeScript client and Web tests, without a model; live acceptance is recorded with the batch. | -| CF | [Session input conflicts and result targets](official-semantics-alignment.md#session-input-conflicts-and-result-targets--september-23); private `~/.parsar/remediation/20260923/campaign-scan-4/errors/findings.json` ERR-22/27 (raw `official/results.json`: `sessB-message-while-running`, `sessA-delete-while-waiting`) and `~/.parsar/remediation/20260923/campaign-scan-2/events-tools/findings.json` EVT-11/12/14 (raw `official/calls-s2.json`: `s2-result-unknown-call`, `s2-result-unknown-turn`, `s2-result-duplicate-after-terminal`, `s2-result-changed-after-terminal`, `s2-result-after-cancel`). Rows CF1–CF9 of that section. Go API bodies, real-PostgreSQL HTTP replay with tenant B and a no-write digest, pinned-SDK scripts; live acceptance is recorded with the batch. | -| SW | [Saved web_search modes](official-semantics-alignment.md#saved-web_search-modes--september-23); private `~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` TV-05 (W01/W02) and `~/.parsar/remediation/20260923/saved-web-search/official/{results.json,ledger.jsonl}`: four owned Agents (all deleted) with create records `type-only`, `mode-null`, `mode-cached`, `mode-cached-full`, update records `update-disabled`, `update-omitted-low`, `update-live-domains-empty` and `retrieve`, plus a second probe of two Agents (deleted, 404 confirmed) with `location-partial-omitted` and `location-empty`, without a Session or model. Rows W1–W8 of that section. Go handler, real-PostgreSQL HTTP exact-bytes/no-write/tenant B, pinned-SDK, TypeScript client and Web tests; independent acceptance is recorded with the batch. | -| SR | [Item serialization](history-events-usage.md#item-serialization-2026-09-23); private `~/.parsar/remediation/20260923/campaign-scan-2/events-tools/findings.json` EVT-09, EVT-10, EVT-13 with raw frames `official/streams-s1..s4.json` and Items pages `official/calls-s2.json` `s2-items-after-t1`, `calls-s4.json` `s4-items`; `campaign-scan-1/sessions/findings.json` SES-23/25 and `campaign-scan-1/vaults-agents/findings.json` VA-11; `campaign-scan-5/sessions-turns/findings.json` ST-03 with raw `official/calls.json`; live-kit EVT-24 `creation-stream-settlement/acceptance/candidate-evidence/codex-kimi/attempt-1/r1-*.json`. Rows S1–S8 of that section. Go contract, API and real-PostgreSQL store tests, Codex adapter usage tests, the pinned-SDK official client suite and TypeScript client/Web tests; live acceptance is recorded with the batch. | -| MV | [MCP origin and credential selection](official-semantics-alignment.md#mcp-origin-and-credential-selection--september-23); private `~/.parsar/remediation/20260923/campaign-scan-6/mcp-vaults/findings.json` MV-01..03 with raw records in `official-ledger.jsonl` (`AG1-minimal-and-nulls`, `S1-T1-create-stream`, `S1-after-c1-delete-session-get`, `S1-final-session-get`, `ERR-UNATTACHED`, `ERR-URL-MISMATCH`, `ERR-AMBIGUOUS`, `ERR-CREDENTIAL-BOGUS`, `ERR-VAULT-BOGUS`): owned Agents, three owned Sessions, two Vaults and four Credentials, all deleted. Rows M1–M9 of that section. Go API/store and real-PostgreSQL tenant A/B HTTP tests with a no-write digest, pinned-SDK official client scripts; live acceptance is recorded with the batch. | -| HF | [Hosted initialization failure](official-semantics-alignment.md#hosted-initialization-failure--september-23); private `~/.parsar/remediation/20260923/campaign-scan-6/hosted-init/findings.json` HI-01..04 with raw records under `official/` (`006-S2-create-setup-exit3`, `007-S2-events`, `009-S3-events`, `021-S2-env-after-failed`, `024-S2-session-after-failed`, `025-S3-session-after-failed`, `026-S2-input-after-failure`, `027-S3-input-after-failure`, `038-S2-delete`, `041-S3-delete`): two owned `openai_hosted` Sessions without a Turn (setup exit 3, nonexistent Python package), both deleted. Rows H1–H8 of that section; HI-05/06 stay deferred. Python initializer receipt tests, Go contract/API tests, real-PostgreSQL managed-Worker store tests with a controlled Provider and an HTTP test with tenant B and a canary, TypeScript client and Web tests; live Docker acceptance is recorded with the batch. | -| RH | [HTTP routing and response headers](official-semantics-alignment.md#http-routing-and-response-headers--september-23); private `~/.parsar/remediation/20260923/campaign-scan-6/http-protocol/findings.json` HP-02, 03, 05, 07, 17–20, 23, 24 (HP-04, 21, 22, 26 kept) with raw records in `official-ledger.jsonl` (labels `A01`–`A11`, `B01`, `B08`, `R04`–`R18`, `S2-retrieve-baseline`, `C01-malformed`, `X1-delete-owned`) and `REPORT.txt`: one owned Agent (deleted), 90 requests without a Session or model. Rows RH1–RH11 apply to every operation's HTTP layer; HEAD on the events stream and content downloads is a documented local 405. | -| JB | [Request body parsing](official-semantics-alignment.md#request-body-parsing--september-23); private `~/.parsar/remediation/20260923/campaign-scan-6/http-protocol/findings.json` HP-09..15 with raw records in `official-ledger.jsonl` (labels `C01`–`C20`, `U01`–`U11`, `S1`) and `REPORT.txt`: one owned Agent (deleted, 404 confirmed), no Session or model; plus `~/.parsar/remediation/20260924/http-json-body/official/results.json` (`lone-high-surrogate`, `case-variant-metadata`), two requests that created nothing. Rows B1–B7 of that section; HP-14 (`[]` as `{}`) is a recorded upstream anomaly, not copied. Official records cover Agent create/update; rows 1, 3, 6, 8, 11, 27, 29, 31, 34, 38 and 40 share the gate. Go gate and all-route handler tests, a real-PostgreSQL owner/tenant B no-write digest across every route family and pinned-SDK scripts; independent acceptance is recorded with the batch. | - -## Per-operation evidence matrix - -Paths in the appendix include `/v1`. SDK names here omit `client.`. `P` means partial implementation. General unresolved status/error/null/default/header/pagination/race semantics apply even where a row lists a narrower gap. - -| # | SDK operation | Implemented behavior | Official wire observation | Core validation | Known difference / unverified semantics | -| --- | --- | --- | --- | --- | --- | -| 1 | beta.agents.create | P: saved configuration, 201; metadata/name errors use official code and param; configuration protocol errors use official code, JSON-path param and message; repeated tools and non-object schema roots reject; every pinned `web_search` mode is saved, omitted/null mode as `live`; omitted/null HTTP MCP `connection_origin` is saved as `service`; the shared body gate rejects malformed, invalid UTF-8, repeated-key and non-object bodies and a non-JSON Content-Type with the official messages, and an empty or null body is `{}` | R `agent-create-supported`; initial unsupported model case 400; X VA-07/08/09; M TV-01/02 `AC01`/`AC02`; SW TV-05 `type-only`, `mode-null`, `mode-cached`, `mode-cached-full`; MV MV-01 `AG1-minimal-and-nulls`; JB HP-09..13/15 `C01`–`C20`, `S1` | C DB create; T saved-Agent execution references; X DB field errors and U+0000 no-write; M DB configuration errors, no-write and tenant replay; SW DB exact tool bytes on create/retrieve/list, admission rejection without writes and tenant B; MV DB omitted/null origin equals explicit; JB DB body gate no-write digest with tenant B | Model-derived reasoning defaults and unsupported configurations; enabled `web_search` is saved but rejects at Session admission; multi-error order, unsampled kind phrases and complete default/null/errors unknown | -| 2 | beta.agents.retrieve | P: tenant-owned saved read | R `agent-read`, `agent-read-deleted` | C DB own/foreign/deleted read | Full field defaults and inline-vs-saved lifetime | -| 3 | beta.agents.update | P: atomic replacements; empty body touches timestamp; metadata/name errors use official code and param; configuration errors as for create, before the Agent lookup; `web_search` and MCP origin projections as for create; body gate as for create, before the Agent lookup, so a zero-length or null body is the empty update | R `agent-patch-metadata`, `agent-null-fields`, `agent-noop`, `agent-nested-reasoning`, rejection labels; X VA-07/08/09/10; M TV-01..03 and TV-07 update labels (`F01`–`X01`, `W04`–`M02`); SW `update-disabled`, `update-omitted-low`, `update-live-domains-empty`, `retrieve`; JB HP-09/11/13/15 `U01`–`U11` | C DB no-op/unchanged snapshot; resource implementation-validation.md actual PostgreSQL SDK update tests; M DB owned/foreign/missing/malformed-ID replay, K1 values and isolation; SW DB update projections, retrieve/list bytes and same-key Session retry; MV DB null-origin update; JB DB body gate no-write digest with tenant B and pinned-SDK raw bytes | Model-dependent default recomputation; uncommon nested/null/error variants | -| 4 | beta.agents.list | P: scoped cursor list; unknown keys ignored, limit 0/above 100 clamp; any unresolved cursor, malformed included, is the missing 404 | R `agent-list-empty-scoped`, `agent-list-limit101`; L VA-01/02/03/04/18; CE ERR-13 `cur-agents-random`, `cur-agents-othertype-session` | L DB `official_list_query.py` tenant A/B; CE DB cursor matrix tenant A/B | Core page capacity 100; no inferred official cap. Overflowing limits unsampled | -| 5 | beta.agents.delete | P: resource deletion | R `cleanup-agent`, subsequent 404 | C DB delete/post-delete | Referenced/in-flight/repeated-delete exact parity | -| 6 | beta.agents.sessions.create | P: JSON/live SSE 201, saved/inline frozen config, initial messages, native profiles; inline agent configuration errors use official fields with `agent.` params before the input requirement, and saved records with repeated tools or non-object schema roots reject admission; fresh creation SSE sends the JSON 201 projection, then ends right after the first idle recorded when a Turn ends or an input reservation stops being pending, or any failed, never sending later events; nothing admitted ends after `created`; a silent settlement ends after events up to the cursor read with a settled projection in one snapshot. A same-key stream retry returns 201, sends no events and ends at once; whitespace-only text is admitted and stored verbatim, while empty text/content/input keep the local 400; omitted/null HTTP MCP origin is `service`; MCP credential selection errors use the official status, code, null param and message after the input requirement, with one message for missing, foreign and unattached references, and write nothing | S `create-1/2.json`, `omitted-input.json`, `null-input.json`, `empty-array-input.json`, `retry-status-original/repeat.json`; H stream; J creation streams closed after idle, open through requires_action; M TV-01..04 `SC01`–`SC11`; P SES-01..03 whitespace-only string and message input 201, verbatim Item; MV MV-01 `S1-T1-create-stream`, MV-03 `ERR-UNATTACHED`, `ERR-URL-MISMATCH`, `ERR-AMBIGUOUS`, `ERR-CREDENTIAL-BOGUS`, `ERR-VAULT-BOGUS` | N Live none admission and retry; C Live three hosted profiles; D/T/K/I recorded additional workflows; J DB creation-stream lifetime/snapshot/retry; M DB inline and saved-override configuration errors without writes; P DB verbatim whitespace Items and unchanged empty-input 400 without writes; MV DB M1–M8 tenant A/B, byte-identical not-found bodies and no-write digest | Session admission batch removes idle `none` creation; local idempotent create still differs from two official IDs. Stream retry, self-hosted, hosted and no-input stream lifetimes are local; work drained before a silent-settlement read can still be sent. Many input/tool/environment combinations restricted; harness admission limits keep the local code (TV-06). Empty-input 400s keep the local code and message; `["", text]` parts are accepted but unobserved officially (SES-08); Claude SDK and MiniMax Code reject whitespace-only messages at admission as a declared native limitation (W6); the unknown-Vault message stays `Resource not found.`; a deleted selected credential still fails at dispatch rather than input (MV-04) | -| 7 | beta.agents.sessions.retrieve | P: persisted state, required actions, usage; malformed ID equals missing; `agent.reasoning` carries both keys, null when unset; usage is the recorded root Turn sum only when every root Turn has ended with known usage, otherwise null; an MCP tool without explicit `credential_id` shows the implicitly selected credential ID, also after deletion; a hosted Environment that failed to provision reads `failed` with its safe step/exit-status reason and failure time | S `retrieve-1.json`, `session-after-1.json`; H recovered state; X SES-28; SR SES-23, EVT-13, ST-03; MV MV-02 `S1-after-c1-delete-session-get`, `S1-final-session-get`; HF HI-01/04 `024-S2-session-after-failed`, `025-S3-session-after-failed` | C Live history; T pending actions; H Core acceptance recorded; SR API/DB reasoning and usage rule; MV API/DB selected-credential projection before and after deletion; HF DB/HTTP failure projection, tenant B | Complete statuses/actions/lifecycle timing; Claude/MiniMax public usage remains null; model-derived default effort not resolved (VA-11); queued-Turn usage unobserved (treated as not ended) | -| 8 | beta.agents.sessions.update | P: metadata-only replacement/clear; metadata errors use official code and `metadata`/`metadata.` param | S `metadata-replace/null/empty/omit/invalid-value.json`; `update-agent.json` uses newer unpinned field | N Live completed Session metadata rejection/clear/isolation; recorded active controlled metadata coverage | Session admission batch changes empty update to observed official 400; Session agent update belongs to baseline upgrade, not fixed-pin operation gap | -| 9 | beta.agents.sessions.list | P: Agent filter, full envelope, cursor paging; unknown keys ignored, limit 0/above 100 clamp; any unresolved cursor, malformed included, is the missing 404; Sessions project the selected MCP credential as retrieve does; failed hosted provisioning projected as on retrieve | S `list-filter.json`, `list-empty-after.json`, `list-owned-cross-filter-cursor.json`, limit/order/unknown-query samples; L SES-10/11/15/16/17; CE ERR-01 `cur-sessions-malformed`, ERR-13/14 | C Live order/cursors/empty/tenant checks; L DB tenant A/B; CE DB cursor matrix tenant A/B; MV DB selected-credential projection; HF DB/HTTP failed projection, tenant B | Core page capacity 100; official cap unknown. Eventual visibility sample is not a required delay | -| 10 | beta.agents.sessions.delete | P: deletion only of a durably idle or failed Session without required actions or pending input; a busy root Turn or pending reservation gives 409 `conflict_error` with no change (subagent child Turns and pending Environment file writes are not checked); owner repeat returns the same 200; owned managed cleanup, user compute retained | S cleanup files 200/deleted; retry-session active cleanup initially 409; Z SES-29 `q5-delete-repeat` 200, SES-30 `q5b-delete-while-in-progress` 409, `q5-delete-never-existed` 404, `q5-delete-while-running` 200 right after an `events.create` 202; HF `038-S2-delete`/`041-S3-delete` 200 after a hosted initialization failure | D recorded real cleanup; C cleanup separately recorded; Z DB matrix D1–D4 with no-write digest, admission lock race and pinned-SDK script; HF HTTP delete of a failed hosted Session | Core returns 409 right after an `events.create` 202 because it admits Turns synchronously (official 200); input awaiting its Environment cannot be cancelled publicly and is unobserved officially; physical purge/retention may end repeat idempotency; caller compute ownership preserved | -| 11 | beta.agents.sessions.events.create | P: 202/empty body, empty-array authenticated no-op, text/cancel/function admission; whitespace-only text is admitted and stored verbatim, empty text/content/input keep the local 400; input the Session cannot accept (result after cancellation, batch while input is pending) and changed results give 409 `conflict_error`/`conflict_error`; in an owned Session an unknown call or a call of another Turn gives 400 `invalid_request_error` without writes; missing/foreign Sessions stay 404; key reuse keeps local 409 `idempotency_conflict` with type `conflict_error`; a Session whose hosted Environment failed to provision gives 409 `conflict_error`/`conflict_error` "the hosted environment failed to provision" | S `second-turn-create.json`, `events-empty/null.json`; H second-input; P SES-04 two whitespace-only messages 202, SES-06/07 empty text/content/input 400; CF EVT-11 `s2-result-unknown-call`/`s2-result-unknown-turn` 400, EVT-12 `s2-result-changed-after-terminal`/`s2-result-after-cancel` 409, EVT-14 duplicate 202, ERR-22 `sessB-message-while-running` 409; HF HI-03 `026-S2-input-after-failure`, `027-S3-input-after-failure` 409 | C Live real continuation/no-op; T qualified message/result/cancel workflows; P DB verbatim whitespace Items and no-write empty-input rejection; CF DB rows CF2–CF4 and CF6–CF9 with tenant B, no-write digest and unchanged pending action, pinned-SDK result script; HF DB/HTTP exact 409, tenant B | Mixed prepared-environment batches, native receipt vs durable acceptance, cancel-before-result-publication timing; unqualified content/tools; empty-input error code/message differ; `["", text]` unobserved officially (SES-08); Claude SDK and MiniMax Code reject whitespace-only messages at admission as a declared native limitation (W6); the official asynchronous pending-input window (ERR-22) is not emulated; Core messages omit call/executor IDs; official order between pending-input and target errors unobserved | -| 12 | beta.agents.sessions.events.stream | P: live-only SSE, typed persisted projections; terminal Turn events carry top-level Turn usage (null when unknown); new Turns start `turn.created`, user `item.added`, `session.in_progress`, `turn.in_progress`; Item events carry `output_index`, null for input Items; assistant text is added in progress with empty content, then an empty part, deltas (a non-streamed final in one delta) and the done events; Session snapshots project the selected MCP credential as retrieve does; a hosted provisioning failure sends `environment.failed` (`environment_error`/`environment_connection_failed`), `error` (`environment_error`/`sandbox_error`, `param` null) and `agent.session.failed`, then GET and creation streams end | H create/reconnect frames; no historical frames in sampled idle interval; J GET streams never ended, terminal `usage` present; SR EVT-09/10 `streams-s1..s4`; MV MV-02 `S1-T1-create-stream` snapshots; HF HI-01/02 `007-S2-events`, `009-S3-events` (stream closed) | C Live; H recorded three-harness disconnect/recovery; T pending actions; J DB order/usage; SR DB sequences and API rendering; MV DB created, in-progress and idle snapshots; HF DB order and HTTP GET stream end | Full SSE/Item variants/order (EVT-05..08 deferred); child Turns and Items are not streamed on the Session (U, SAT-09) and are read through the Subagent routes; no replay guarantee or observer-disconnect proof for every state | -| 13 | beta.agents.sessions.turns.retrieve | P: persisted root Turn identity; malformed and child Turn IDs equal missing | H/S contain Turn list payloads; no isolated positive retrieve raw request identified in this set; X SES-28 official `turn_` 404; U SAT-07 `C04` child Turn ID 404 | Recorded H/B scoped Turn recovery (B under the earlier mixed root/child contract); C history uses list; U DB child-ID 404 equal to missing, tenant B | Distinguish list-shape evidence from retrieve wire qualification; full lifecycle/usage; official 404 message text differs | -| 14 | beta.agents.sessions.turns.list | P: full envelope, ordered root Turns only; a child Turn, malformed or other unresolved cursor equals a missing one; limit outside 1–100 rejects with the Beta code | S `turns-final/empty-page/limit-high/order-empty.json`; H both directions; L SES-14/15/16/17; X SES-28; U SAT-07 `R06` root-only list beside a child Turn; CE ERR-01 `cur-turns-malformed`, ERR-13 | C Live paging; H/B real child/root identity recorded under the earlier mixed contract; L DB tenant A/B; U DB root-only list and cursor; CE DB cursor matrix tenant A/B | All interleavings, same-timestamp paging, interim/failed usage | -| 15 | beta.agents.sessions.items.list | P: scoped root Items, full envelope; limit 0/above 100 clamp within the pinned 1–100 page; a cursor that is not an Item of the Session is 400 ``Invalid session item ID in `after` ``; messages carry `phase`, null for user messages and when no native phase is reported; function results carry `output` and `error`, null when not submitted | S `items-final/empty-page/limit-high.json`; H both directions; L SES-12/13/15/16/17; CE ERR-02 `cur-items-*`; SR SES-25, EVT-09 `s2-items-after-t1`, `s4-items` | C Live; H/T recorded content/coordination/result variants; L DB tenant A/B; CE DB cursor matrix tenant A/B; SR DB/SDK null fields | Full Item union. Newer turn_id filter excluded from pin. Official pages showed a failed result's submitted array output as null; Core keeps the submitted content | -| 16 | beta.agents.sessions.artifacts.retrieve | P: immutable captured metadata; unchanged outputs keep their Artifact ID across Turns | Y HE-58/59 `ar01-retrieve`, missing and `ar03-cross-session` 404 (fields match; official `artifact_` IDs vs Core UUIDs) | D recorded live Docker/user-managed workspace output; Y DB unchanged-path ID and metadata stability | Error messages differ; hard-link/special-file capture edges unobserved | -| 17 | beta.agents.sessions.artifacts.list | P: scoped stored list, common list envelope; later Turns publish only new, changed or no-remaining-Artifact paths; symlinks skipped; malformed `environment_id` gives an empty page; a cursor that is not an Artifact of the Session is 400 `after is not a valid artifact ID` | Y HE-50..56 `al01`–`al09`: Turn 1 capture with a skipped link, Turn 2 republication, envelope, order/cursor/limits, other and malformed filters; CE HE-57, ERR-05 `cur-artifacts-*` | D recorded live output enumeration; Y Rust link tests, DB republication, HTTP envelope/filter/tenant and pinned-SDK checks; CE DB cursor matrix tenant A/B | Changed-bytes republication inferred; deleted-newest edge and paging during capture/delete unobserved; another Session's Artifact as a cursor is inferred | -| 18 | beta.agents.sessions.artifacts.delete | P: stored deletion; the next Turn republishes a deleted path | Y HE-61 `ad01-delete`, `ad02-delete-repeat` 404, reads after delete 404; HE-52 deleted path republished by Turn 2 | D recorded workflow summary; Y DB deleted-then-unchanged and deleted-during-capture republication | In-flight deletion and physical retention parity | -| 19 | beta.agents.sessions.artifacts.content | P: immutable download after Runtime loss | Y HE-60/62 `ac01-content`, `ac02-content-range` (Range ignored, full 200), `ac05` 404 after Session deletion | D recorded live retained download; Y DB earlier versions keep their bytes | Core's extra Content-Disposition; cancellation-edge capture, partial transfer and content after Environment expiry unobserved officially | -| 20 | beta.agents.sessions.subagents.retrieve | P: owned durable child identity/lifecycle | U `R01`, `R22`–`R26` (SAT-05/13 match; instructions hidden, SAT-11 ARCH) | B recorded Live six-read matrix; U DB tenant B 404 | Full lifecycle/multi-agent parity; native close differences; 404 message text (SAT-06) | -| 21 | beta.agents.sessions.subagents.list | P: direct/nested/closed child records; `object`/`first_id`/`last_id` envelope with null IDs on an empty page; limit outside 1–100 rejects; a cursor that is not a Subagent of the Session is 400 ``Invalid resource ID in `after` `` | U SAT-01 `R00`, `R11`; SAT-03 `R12`/`R13` rejections; CE SAT-04 `R32`/`R33`, ERR-04 `cur-subagents-*` | B recorded Live scopes/pages/continuation; U DB envelope, empty page, rejection, tenant B; CE DB cursor matrix tenant A/B | Publication timing/parent propagation and unsupported native nesting | -| 22 | beta.agents.sessions.subagents.items.list | P: child-owned history only; limit 0/above 100 clamp; a cursor outside the child's Items is 400 ``Invalid session item ID in `after` `` | U SAT-02 `R14`/`R15`; SAT-12 `R02` (contents unknown); CE ERR-03 `cur-subitems-*` | B recorded Live child separation/recovery; U DB clamp, tenant B; CE DB cursor matrix tenant A/B | Full child Item union; child input Item retained (SAT-12); continuous child progress is not streamed on the Session | -| 23 | beta.agents.sessions.subagents.turns.retrieve | P: scoped child Turn; `agent_id` is the Session Agent ID, `subagent_id` the child | U SAT-08 `R04`; `R27`/`R28` 404 | B recorded Live six-read matrix; U DB direct and nested child identity, tenant B | Nested-child `agent_id` unobserved; full status/usage/timestamp official semantics | -| 24 | beta.agents.sessions.subagents.turns.list | P: persisted child history with the Session Agent ID; limit outside 1–100 rejects; a cursor that is not a Turn of the child is 400 ``Invalid resource ID in `after` `` | U SAT-08 `R03`; SAT-03 `R16`/`R17` rejections; CE SAT-04 `C06`, ERR-04 `cur-subturns-*` | B recorded Live paging/cold continuation; U DB identity and rejection; CE DB cursor matrix tenant A/B | Full concurrent/cancel ordering and accounting | -| 25 | beta.agents.sessions.subagents.turns.items.list | P: exact child-and-Turn Items; limit 0/above 100 clamp; a cursor outside the child Turn's Items is 400 ``Invalid session item ID in `after` `` | U SAT-02 `C05`; `R05`; CE ERR-03 `cur-subturnitems-*` | B recorded Live six-read matrix; U DB clamp, tenant B; CE DB cursor matrix tenant A/B | Complete union and live ordering, overlapping mutation/cursors | -| 26 | beta.agents.environments.retrieve | P: durable status, safe configured installation metadata | None located | D/I/K recorded native readiness and metadata | All lifecycle timing and installation inventory; configured metadata is not arbitrary workspace discovery | -| 27 | beta.agents.environments.files.create | P: inline/source-file copy to qualified workspace; 201; observed path, unknown-field and provisioning errors; missing parents created, existing paths never replaced, 5 MiB decoded inline bound with official conflict, symlink/overwrite and size errors | G HE-10 success/nested/empty, HE-16 path forms and unknown field, HE-18 pending; FW HE-11 nested, HE-12 overwrite/untracked/directory, HE-13 symlinks, HE-15 5 MiB + 1 | F/D/K recorded real copy, hashes/consumption/retention; G handler tests; FW Rust, daemon, handler and real-PostgreSQL tests | A file an earlier Files.create wrote gets the untracked-file message (no path ledger); unsampled non-directory parents and non-regular destinations; other body errors and unknown-write semantics | -| 28 | beta.agents.environments.files.list | P: direct regular-file directory, opaque cursor; `page` envelope; unknown keys ignored (malformed query encoding still rejected locally), repeated keys rejected; missing, file and symlink paths list empty without following on local workspace readers (the Claude SDK adapter reader keeps 404/503); cleaned-path, token and provisioning errors | G HE-32 envelope, HE-34/35 query keys, HE-36/37 empty pages, HE-38/39 path and token errors, HE-18 pending | F/D/K recorded live workspace listing; G handler, Worker and native helper tests | 1,024-entry prefilter bound; no recursion (HE-30); exact defaults, unsampled path forms and mutation invalidation unknown | -| 29 | beta.agents.environments.templates.create | P: reusable network/files/env/setup/packages/Skills/Plugins/capability config, 201; network rejections use `invalid_request_error` | R `template-create`; V nullable Skill selector projection; X SFT-20 | C DB; I/K recorded real frozen-reference initialization | Restricted forms and unqualified combinations; complete hosted initialization semantics | -| 30 | beta.agents.environments.templates.retrieve | P: safe resource read | R `template-read`, deleted owned read; V nullable Skill selector projection | C DB; I/K recorded reference workflow | Full field/default/redaction parity; no live-secret projection inference | -| 31 | beta.agents.environments.templates.update | P: field replacement/null clearing, empty timestamp touch; network rejections use `invalid_request_error` | R `template-patch`, `template-null`, `template-noop`; V nullable Skill selector projection; X SFT-20 | C DB no-op; I/K recorded frozen Session behavior | Template update and referencing Session selection are distinct; composition and null inheritance are qualified separately in environment-templates.md/template-null-selection.md; uncommon fields/errors remain unverified | -| 32 | beta.agents.environments.templates.list | P: scoped cursor list; unknown keys ignored, limit 0/above 100 clamp; any unresolved cursor is the missing 404 | R `template-list-empty-scoped`; L SFT-11/23/24; CE ERR-07 `cur-templates-*` (match) | Recorded resource DB checks; L DB tenant A/B; CE DB cursor matrix tenant A/B | Multipage mutation and default parity | -| 33 | beta.agents.environments.templates.delete | P: delete resource, preserve committed Session snapshot | R `cleanup` at Template path, post-delete read | C DB; K recorded live deletion then continuation | Concurrent references/delete and exact errors | -| 34 | beta.agents.vaults.create | P: tenant resource, 201; non-string metadata reports `metadata.`; no pair/length limits | R `vault-create`, `vault-empty-token-fixture`; X VA-08/10/18 | C DB; O recorded MCP attachment workflow; X DB U+0000 no-write | Archive lifecycle, full defaults and selection parity | -| 35 | beta.agents.vaults.retrieve | P: safe metadata/status | R `vault-read`, post-cleanup read | C DB deleted/error checks; O recorded lifecycle | Full archived-state and visibility semantics | -| 36 | beta.agents.vaults.list | P: stored status filter (scalar and `status[]` union), separate clamping pagination; any unresolved cursor, malformed included, is the missing 404 | R `vault-list-empty-scoped`; L VA-02/03/04/05/06/18; CE ERR-01, ERR-13 `cur-vaults-*` | Recorded resource DB coverage; L DB tenant A/B; CE DB cursor matrix tenant A/B | Real archive transitions; official negative-limit 400 is a recorded pin conflict | -| 37 | beta.agents.vaults.delete | P: atomic credential cascade, frozen attachment boundaries | R `cleanup` at Vault path | C DB deletion; O recorded grant lifecycle | Archive vs delete, already-delivered tokens and active effects | -| 38 | beta.agents.vaults.credentials.create | P: encrypted static/OAuth, nonempty tokens, 201 | R static/OAuth create + empty token/access-token rejection labels | C DB exact-row preservation/restart; O recorded genuine grants/real MCP | Full provider grants/default/error/selection behavior; qualified none profiles only | -| 39 | beta.agents.vaults.credentials.retrieve | P: safe metadata, secret-free read | R `credential-read`, post-delete read | C DB own/foreign/restart; O recorded lifecycle | Full response metadata and archived-state parity | -| 40 | beta.agents.vaults.credentials.update | P: static replacement/OAuth grant patch, reject empty effective update | R `credential-rotate`, `credential-empty-token`, `credential-oauth-empty-patch` | C DB rejected-write preservation; O recorded real Codex replacement/refresh | Official empty-access-token update not separately probed; concurrent refresh/replacement/provider errors and in-flight withdrawal | -| 41 | beta.agents.vaults.credentials.list | P: safe scoped metadata/status list (scalar and `status[]` union); any unresolved cursor, malformed or the Vault's own ID included, is the missing 404 | R `credential-list`; L VA-03/04/05/06/18; CE ERR-01 `cur-creds-malformed`, ERR-12/13 | C DB mixed list/rejected-write/restart; O recorded resource checks; L DB tenant A/B; CE DB cursor matrix tenant A/B | Archive behavior; official negative-limit 400 is a recorded pin conflict | -| 42 | beta.agents.vaults.credentials.delete | P: encrypted resource deletion and dispatch denial | R `cleanup` at Credential paths | C DB; O recorded live Codex refusal after deletion | Cannot recall already-delivered token; exact hosted withdrawal/error timing | -| 43 | files.create | P: immutable multipart purpose=user_data | W three owned200 uploads (two user_data, assistants control) | F recorded DB upload; D/K recorded native source consumption | Other Core upload purposes/expires_after unsupported; purpose acceptance does not imply platform processing | -| 44 | files.retrieve | P: project-owned metadata | E/V missing-ID404 param=id; W positive metadata/status/nulls | F recorded DB workflow; V actual HTTP/PG foreign==missing | Other purposes/errors/defaults beyond the sampled profile | -| 45 | files.list | P: validated purpose filter, default/max10000, cursor list; unknown keys and empty purpose ignored, range errors null code | Q order error; V missing-after404; W nine purpose values reach cursor lookup, unknown/case variant400purpose; L SFT-11/14/17/18 | F recorded actual PostgreSQL SDK/raw pages/restart/isolation; L DB tenant A/B | Successful official filtering, repeated purpose, ordering/concurrent mutation remain unqualified | -| 46 | files.delete | P: transactional metadata/body unlink | V random nonexistent ID404 param=id | F recorded DB deletion; V actualHTTP/PG foreign delete denial preserves bytes; D/K retained workspace copies | Admitted immutable read may complete after deletion; hosted concurrency/retention unknown | -| 47 | files.content | P: tenant lookup then user_data400 denial; internal source consumption retained | V missing-ID404param=id; W user_data/assistants400 nullcode/nullparam | Source store/copy regressions; joint acceptance in file-resource-semantics.md | Other-purpose content policy and full error details remain unqualified | -| 48 | skills.create | P: encrypted bounded directory/ZIP bundle | V successful owned directory uploads | K recorded DB resources and Live three-harness upload/reference | Fixed SDK single-tuple multipart issue; local upload limits remain partial | -| 49 | skills.retrieve | P: safe metadata follows default version | E missing-ID404; W distinct names/descriptions follow default1→2→1 | Skill store and joint resource acceptance | Other error/visibility behavior remains unqualified | -| 50 | skills.update | P: atomic default pointer and descriptive metadata update | V default mutation/frozen Sessions; W default1→2→1 name/description changes | Skill store and joint resource acceptance; retained frozen Session regression | Complete errors and concurrent official behavior | -| 51 | skills.list | P: scoped resource list; limit 0 empty page with has_more, 0–100 with observed range and duplicate codes | L SFT-08/09/10/11 | K recorded DB resources; L DB tenant A/B | Non-integer and overflowing limits unsampled | -| 52 | skills.delete | P: remove owned source and every encrypted version, retain committed Session content; sole-version deletion uses the same cascade | [sole-version deletion](file-resource-semantics.md#sole-version-deletion--september-23-2026) SFT-01 sole-version deletion then Skill 404; SFT-06 repeat delete 404; SFT-07 retrieve/version reads 404 | K recorded DB and Live source deletion/continuation; sole-version store, HTTP tenant A/B and pinned-SDK DB tests | Official 404 message names the Skill, Core keeps a generic message; physical erasure and concurrent hosted deletion unobserved | -| 53 | skills.content.retrieve | P: unversioned content selects default | W distinct default1/latest2 bytes and default2 transition | K recorded DB plus joint resource acceptance | Headers/errors and source deletion/read races | -| 54 | skills.versions.create | P: immutable increasing version; optional default change | V second version default=false preserves default1/latest2 | K recorded DB resources | Broader numbering/default/top-level metadata/error/null semantics; upload limits | -| 55 | skills.versions.retrieve | P: owned immutable version metadata; malformed version path equals missing | V immediate version1 read404, bounded delayed read200 | K recorded DB resources and Live concrete Session freeze | Visibility timing is observational; full selector/metadata/error parity unqualified | -| 56 | skills.versions.list | P: scoped version cursor list; limit 0 empty page with has_more; a cursor not beginning with `skillver` or naming another Skill's version is 400 `invalid_value` on `after`, a missing version 404 | V delayed owned list contains created versions; L SFT-08/09; CE ERR-06 `cur-skillvers-*` | K recorded DB resource checks; L DB zero page; CE DB cursor matrix tenant A/B and pinned SDK | Exact ordering/cursors/default and concurrent version mutation | -| 57 | skills.versions.delete | P: nondefault deletion; default rejects invalid_value/version while other versions remain; deleting the only version deletes the Skill in the same locked transaction | W two-version default400; latest200 with parent pointer fallback; [sole-version deletion](file-resource-semantics.md#sole-version-deletion--september-23-2026) SFT-01 sole200 then Skill 404, SFT-04 default400 with visible v2 | Skill store (atomic removal, frozen Session, upload lock order), HTTP tenant A/B and joint resource acceptance | SFT-02 number reuse is an intentional difference (numbers stay immutable); SFT-03 whole-Skill removal and stale official version reads are not emulated | -| 58 | skills.versions.content.retrieve | P: decrypt/read immutable concrete bundle | W v1/v2 ZIP members and markers | K recorded DB/live consumption; joint resource acceptance | Content headers/errors and source deletion/read races | - -## Remaining gaps without task ordering - -1. **Public generic semantics:** sampled create/event/envelope/error/no-op corrections are merged. L aligns unknown/repeated list keys, sampled limit bounds and single-resource unknown keys. X gives malformed path IDs on every Beta, Files and Skills route the exact missing-resource response, reports metadata/name field errors with official code and param, maps Template network rejections to `invalid_request_error`, and rejects U+0000 in stored strings as a documented local limit (the official service stores it). G aligns Environment Files list query tolerance and the sampled Files.create/list errors; FW aligns Files.create parent creation, no-replacement and its conflict and size errors. CE gives every Beta and Skill version list the observed unresolved-cursor error; Files, Skills and Environment Files cursors are unchanged. Other resource-by-resource omissions/null/default/error params, overflowing limits, list caps, concurrent mutation and deletion require separate evidence. -2. **Session differences:** The Session admission batch removes idle `none` creation and empty metadata update. Local durable creation idempotency remains an explicit difference. P admits whitespace-only input verbatim, as observed officially; Claude SDK and MiniMax Code declare whitespace-only messages unsupported at admission. Session agent updates, newer Environment shapes and root Item turn_id are baseline-upgrade questions. -3. **Template/Skill composition:** shared env/files/setup/packages selection is covered by the composition batch; template-reference null network/capability lists are covered by the null-selection batch. Official derived capability-directory projection remains different. Skill content/default metadata and sole-version deletion are covered by file-resource-semantics.md; number reuse is an intentional difference; visibility and broader error behavior remain unverified. -4. **Execution coverage:** use T's qualified matrix, not a blanket missing-image/structured-output claim. MiniMax functions/service MCP, optional tool combinations, unsupported images/placements and broader native lifecycle are explicit restrictions. PTC omission retains approved native behavior; Claude/MiniMax public Usage remains null; child settlement cadence/native close limits remain visible. No second executor/model loop or guessed counters are justified. -5. **Workspace and resources:** live Files recursion (G aligns the sampled envelope, empty pages and path errors; FW aligns parent creation, no-replacement and the inline bound), artifact capture edges for hard links/special files and cancellation (Y aligns output symlinks, republication and the list envelope), full Environment metadata/lifecycle and Vault archive/in-flight-token semantics remain partial or unknown. Retired Core-managed E2B acceptance cannot qualify current user enrollment. - -## Enumeration and wording mismatches - -- The README's **42** is correct only for beta/agents. A campaign that includes referenced Files and Skills needs the explicit **58** inventory. Its combined Skills/Versions row abbreviates six operation names, while the fixed SDK actually exposes eleven distinct HTTP operations across resource, version and content classes. -- SDK `files.retrieve_content` is deprecated but makes the same `GET /files/{file_id}/content` request as `files.content`; count it as an alias, not another HTTP operation. Overloads, async mirrors, raw/streaming response wrappers and polling helpers add no distinct paths/methods. Likewise Session create's stream overload shares POST create; events.stream is its own GET. -- Vault paths are `/v1/vaults`, not `/v1/agents/vaults`. Files/Skills are non-beta project-authenticated routes. No public Environment create/delete, Plugin CRUD or extra Subagent mutation operation exists in this fixed resource inventory. -- Existing stale prose must not drive new backlog: `execution-tools.md` blanket OAuth rejection is superseded by O; README broad non-text-input and safe-empty Environment metadata summaries need qualification against T/I/K; environment-templates.md older no-op timestamp uncertainty is superseded by A. Older API comments also mention retired registry transport or unsupported features already delivered. Current code plus latest bounded evidence wins over historical narratives. -- **58 registered routes ≠ 58 compatible operations.** The positive official sample is concentrated in Agents/Sessions/history/Template/Vault/Credentials. No positive official artifact, Environment Files, Subagent or Skill resource workflow is established by the inspected September 22 comparison set. - -## Fixed SDK and current route appendix - -The local fixed-SDK source links below identify the method implementation (overloads excluded). All paths have the base `/v1` added. Source version is verified in [_version.py](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/_version.py#L2). Core registration sources: [handler.go](../../services/core/internal/api/handler.go), [subagents.go](../../services/core/internal/api/subagents.go), [skills.go](../../services/core/internal/api/skills.go). - -| Fixed resource source / method | HTTP operation | -| --- | --- | -| [resources/beta/agents/agents.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/agents.py#L85) | `POST /v1/agents` | -| [resources/beta/agents/agents.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/agents.py#L171) | `GET /v1/agents/{agent_id}` | -| [resources/beta/agents/agents.py:update](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/agents.py#L211) | `POST /v1/agents/{agent_id}` | -| [resources/beta/agents/agents.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/agents.py#L301) | `GET /v1/agents` | -| [resources/beta/agents/agents.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/agents.py#L359) | `DELETE /v1/agents/{agent_id}` | -| [resources/beta/agents/environments/environments.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/environments.py#L63) | `GET /v1/agents/environments/{environment_id}` | -| [resources/beta/agents/environments/files.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/files.py#L119) | `POST /v1/agents/environments/{environment_id}/files` | -| [resources/beta/agents/environments/files.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/files.py#L158) | `GET /v1/agents/environments/{environment_id}/files` | -| [resources/beta/agents/environments/templates.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py#L49) | `POST /v1/agents/environments/templates` | -| [resources/beta/agents/environments/templates.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py#L129) | `GET /v1/agents/environments/templates/{environment_template_id}` | -| [resources/beta/agents/environments/templates.py:update](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py#L174) | `POST /v1/agents/environments/templates/{environment_template_id}` | -| [resources/beta/agents/environments/templates.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py#L260) | `GET /v1/agents/environments/templates` | -| [resources/beta/agents/environments/templates.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py#L318) | `DELETE /v1/agents/environments/templates/{environment_template_id}` | -| [resources/beta/agents/sessions/artifacts.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/artifacts.py#L52) | `GET /v1/agents/sessions/{session_id}/artifacts/{artifact_id}` | -| [resources/beta/agents/sessions/artifacts.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/artifacts.py#L97) | `GET /v1/agents/sessions/{session_id}/artifacts` | -| [resources/beta/agents/sessions/artifacts.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/artifacts.py#L162) | `DELETE /v1/agents/sessions/{session_id}/artifacts/{artifact_id}` | -| [resources/beta/agents/sessions/artifacts.py:content](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/artifacts.py#L207) | `GET /v1/agents/sessions/{session_id}/artifacts/{artifact_id}/content` | -| [resources/beta/agents/sessions/events.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/events.py#L44) | `POST /v1/agents/sessions/{session_id}/events` | -| [resources/beta/agents/sessions/events.py:stream](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/events.py#L91) | `GET /v1/agents/sessions/{session_id}/events` | -| [resources/beta/agents/sessions/items.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/items.py#L44) | `GET /v1/agents/sessions/{session_id}/items` | -| [resources/beta/agents/sessions/sessions.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/sessions.py#L290) | `POST /v1/agents/sessions` | -| [resources/beta/agents/sessions/sessions.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/sessions.py#L336) | `GET /v1/agents/sessions/{session_id}` | -| [resources/beta/agents/sessions/sessions.py:update](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/sessions.py#L376) | `POST /v1/agents/sessions/{session_id}` | -| [resources/beta/agents/sessions/sessions.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/sessions.py#L422) | `GET /v1/agents/sessions` | -| [resources/beta/agents/sessions/sessions.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/sessions.py#L486) | `DELETE /v1/agents/sessions/{session_id}` | -| [resources/beta/agents/sessions/subagents/items.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/subagents/items.py#L44) | `GET /v1/agents/sessions/{session_id}/subagents/{subagent_id}/items` | -| [resources/beta/agents/sessions/subagents/subagents.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/subagents/subagents.py#L67) | `GET /v1/agents/sessions/{session_id}/subagents/{subagent_id}` | -| [resources/beta/agents/sessions/subagents/subagents.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/subagents/subagents.py#L112) | `GET /v1/agents/sessions/{session_id}/subagents` | -| [resources/beta/agents/sessions/subagents/turns/items.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/subagents/turns/items.py#L44) | `GET /v1/agents/sessions/{session_id}/subagents/{subagent_id}/turns/{turn_id}/items` | -| [resources/beta/agents/sessions/subagents/turns/turns.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/subagents/turns/turns.py#L55) | `GET /v1/agents/sessions/{session_id}/subagents/{subagent_id}/turns/{turn_id}` | -| [resources/beta/agents/sessions/subagents/turns/turns.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/subagents/turns/turns.py#L106) | `GET /v1/agents/sessions/{session_id}/subagents/{subagent_id}/turns` | -| [resources/beta/agents/sessions/turns.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/turns.py#L43) | `GET /v1/agents/sessions/{session_id}/turns/{turn_id}` | -| [resources/beta/agents/sessions/turns.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/sessions/turns.py#L87) | `GET /v1/agents/sessions/{session_id}/turns` | -| [resources/beta/agents/vaults/credentials.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/credentials.py#L52) | `POST /v1/vaults/{vault_id}/credentials` | -| [resources/beta/agents/vaults/credentials.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/credentials.py#L107) | `GET /v1/vaults/{vault_id}/credentials/{credential_id}` | -| [resources/beta/agents/vaults/credentials.py:update](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/credentials.py#L152) | `POST /v1/vaults/{vault_id}/credentials/{credential_id}` | -| [resources/beta/agents/vaults/credentials.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/credentials.py#L201) | `GET /v1/vaults/{vault_id}/credentials` | -| [resources/beta/agents/vaults/credentials.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/credentials.py#L269) | `DELETE /v1/vaults/{vault_id}/credentials/{credential_id}` | -| [resources/beta/agents/vaults/vaults.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/vaults.py#L58) | `POST /v1/vaults` | -| [resources/beta/agents/vaults/vaults.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/vaults.py#L110) | `GET /v1/vaults/{vault_id}` | -| [resources/beta/agents/vaults/vaults.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/vaults.py#L150) | `GET /v1/vaults` | -| [resources/beta/agents/vaults/vaults.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/vaults/vaults.py#L215) | `DELETE /v1/vaults/{vault_id}` | -| [resources/skills/content.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/content.py#L43) | `GET /v1/skills/{skill_id}/content` | -| [resources/skills/skills.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/skills.py#L80) | `POST /v1/skills` | -| [resources/skills/skills.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/skills.py#L126) | `GET /v1/skills/{skill_id}` | -| [resources/skills/skills.py:update](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/skills.py#L163) | `POST /v1/skills/{skill_id}` | -| [resources/skills/skills.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/skills.py#L204) | `GET /v1/skills` | -| [resources/skills/skills.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/skills.py#L257) | `DELETE /v1/skills/{skill_id}` | -| [resources/skills/versions/content.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/versions/content.py#L43) | `GET /v1/skills/{skill_id}/versions/{version}/content` | -| [resources/skills/versions/versions.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/versions/versions.py#L68) | `POST /v1/skills/{skill_id}/versions` | -| [resources/skills/versions/versions.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/versions/versions.py#L126) | `GET /v1/skills/{skill_id}/versions/{version}` | -| [resources/skills/versions/versions.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/versions/versions.py#L168) | `GET /v1/skills/{skill_id}/versions` | -| [resources/skills/versions/versions.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills/versions/versions.py#L223) | `DELETE /v1/skills/{skill_id}/versions/{version}` | -| [resources/files.py:create](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/files.py#L63) | `POST /v1/files` | -| [resources/files.py:retrieve](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/files.py#L157) | `GET /v1/files/{file_id}` | -| [resources/files.py:list](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/files.py#L194) | `GET /v1/files` | -| [resources/files.py:delete](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/files.py#L256) | `DELETE /v1/files/{file_id}` | -| [resources/files.py:content](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/files.py#L293) | `GET /v1/files/{file_id}/content` | diff --git a/contracts/agents-api/resource-selector-semantics.md b/contracts/agents-api/resource-selector-semantics.md deleted file mode 100644 index 4d6f6121d..000000000 --- a/contracts/agents-api/resource-selector-semantics.md +++ /dev/null @@ -1,82 +0,0 @@ -# Resource selector and error semantics - -This bounded September 23 alignment uses openai-python **3.13.0**, upstream -`d7c41efee1b0802b79f3f88a678ef2052b06e9ce`, and `agents=v1` for beta resources. -It does not qualify complete Skills, Templates, Sessions or Files compatibility. - -## Skill reference versions - -The pinned Template create/update and Session environment request types all import -`HostedSkillParam`, whose reference version is `Optional[str]`. The unrelated -`BetaSkillReferenceParam` is not their request type. An omitted or null selector -uses the Skill default; `latest` selects the latest version and a concrete string -selects that version. A Template retains unresolved intent and emits `version: -null` for the default selector. A Session exposes the concrete installed version. - -The existing tenant-scoped creation transaction freezes metadata and bundle bytes -together. Later source/default/Template changes do not modify the installation or -cause a creation retry to resolve mutable sources again. This change does not -redefine omitted-vs-null creation idempotency equivalence. Null Skill-list overrides, -version deletion rules, unversioned content selection, query bounds, and Template -plus inline initialization composition remain outside this batch. - -Official qualification uses owned resources. The first immediate version read -returned 404 despite successful creation; a separate bounded follow-up observed -that version after approximately 31 seconds. This is a recorded transient -visibility observation, not a guaranteed consistency interval or a behavior Core -should imitate. Template omission/null admission and projection, latest and exact -selectors are separately recorded. A further hosted Session probe made 36 calls: with default version 1 and latest -version 2, omitted/null resolved to 1, latest to 2, and exact "1" to 1. All four -retained their versions after the default changed to 2. Two actual gpt-6-astra -Turns read installed files and returned distinct private markers, proving frozen -null/default version 1 and latest version 2 content. All four Sessions and the -Skill were deleted; asynchronous physical sandbox destruction was not observed. Private evidence is under -`~/.parsar/remediation/20260923/skill-version-alignment/official/`. - -## Source File errors - -For general `/v1/files`, missing retrieve/content/delete resources use HTTP 404, -`type: invalid_request_error`, `code: null`, and `param: id`. A missing list cursor -uses the same envelope with `param: after`. Foreign resources retain the same -missing-resource response. These parameter hints pass through the existing error -serializer only for a not-found store error; other failures and Skills responses -retain their own mappings. - -Seven official requests qualify these cases, including raw HTTP, fixed SDK -exceptions and one DELETE of a random nonexistent identifier. No account files -were read or deleted. Evidence: -`~/.parsar/remediation/20260923/files-error-alignment/official-probe.json`. -Exact error prose, additional parser detail, purpose filtering, bounds and lookup -order are not changed or claimed as aligned. - -## Core acceptance - -`TestSkillSelectorsOfficialClientPostgres` and -`TestSourceFileErrorsOfficialClientPostgres` exercise real HTTP/PostgreSQL with -fixed SDK and raw requests. They cover exact Template reference projection, -independent handler reads, missing and foreign resources, safe errors and retained -owned data. They do not represent native model execution. - -The separate Codex/Docker run at `2ccc729d3acfa8d7109f671d480753d74d568279` -passed on its first attempt with two real Kimi K3 Turns. Null through a Template -froze version 1; direct latest froze version 2. Changing source default and Template -selectors left same-key creation retries unchanged and created no Turn. After -both sources were deleted and Core restarted, each Session still exposed its -concrete version and returned only its own previously undisclosed Skill marker. -Foreign Session reads failed. Source/binary/image hashes, request/event/history -records, secret scans and verified owned-resource cleanup are retained under -`~/.parsar/remediation/20260923/skill-version-alignment/live/`. - -The integrated Files error change does not alter Skill selection, initialization, -Runtime or adapter code. This live run does not qualify another harness/Provider, -upstream physical retention, or omitted/null cross-form retry equivalence. - -At integrated source `9d9639bd90bc74e7c27832b28dca27aac6c18781`, server -`make -o check-web check` passed, including the real PostgreSQL suite, fixed SDK -resource checks, byte-for-byte sqlc generation, native package checks and builds. -`make openapi` produced no schema change. A fresh local `make check-web` at -`756654b` passed typechecks, Core doctor tests, 287 client tests, 583 Web tests, -build and all 76 browser cases on isolated ports. These split runs cover every -`make check` target; the commits between them change only evidence documentation. -Optional live adapter profiles and the 512 MiB storage stress case are not newly -qualified. Subsequent changes only record evidence in documentation. diff --git a/contracts/agents-api/runtime-history-api.md b/contracts/agents-api/runtime-history-api.md index 3aa6a2349..b44716154 100644 --- a/contracts/agents-api/runtime-history-api.md +++ b/contracts/agents-api/runtime-history-api.md @@ -97,7 +97,7 @@ counter regression produces a gap; it is never filled with zero. The sampled value is Core's measured Session usage, a Core extension that sums every recorded root Turn snapshot, active Turns included. It differs by design from public Session usage, which is null while a root Turn runs or after one ends -unmeasured ([item serialization](history-events-usage.md#item-serialization-2026-09-23)). These counters +unmeasured ([usage](sessions-events.md#usage)). These counters are measured model usage, not price, cost, or billing records. Core Web queries each current managed Session through the administrator Session diff --git a/contracts/agents-api/runtime-observability-api.md b/contracts/agents-api/runtime-observability-api.md index 9ced075e9..507bc07ab 100644 --- a/contracts/agents-api/runtime-observability-api.md +++ b/contracts/agents-api/runtime-observability-api.md @@ -190,7 +190,7 @@ Use the existing Agents API error envelope. | 400 | `invalid_request_error` / `invalid_request_error` | List: a repeated supported query key, or an empty or invalid limit or order, with the shared Beta list messages. Unknown list query keys are ignored. | | 400 | `invalid_request_error` / `unsupported_parameter` | Single-Session retrieval with any query parameter. | | 401 | `invalid_request_error` / `invalid_admin_key` | Missing or invalid Core key. | -| 404 | `not_found_error` / `not_found_error` | Missing, malformed or foreign Session/cursor, indistinguishably, as for the [Session list cursor](list-query-semantics.md#list-cursor-errors--september-23-2026). | +| 404 | `not_found_error` / `not_found_error` | Missing, malformed or foreign Session/cursor, indistinguishably, as for the [Session list cursor](wire-semantics.md#cursors). | | 500 | `server_error` / `internal_error` | Integrity, ownership, or invalid provider evidence. | | 503 | `server_error` / `execution_unavailable` | Required Runtime observation service is not configured, or list collection exceeded its request budget. | diff --git a/contracts/agents-api/runtime-observability-design.md b/contracts/agents-api/runtime-observability-design.md index d968ab710..dec05fa85 100644 --- a/contracts/agents-api/runtime-observability-design.md +++ b/contracts/agents-api/runtime-observability-design.md @@ -471,7 +471,7 @@ history sweep, the Core resolver reads the cumulative measured Session usage (`MeasuredSessionUsage`: every recorded root Turn snapshot, active Turns included) from the execution store alongside Runtime identity. Public Session usage follows the stricter official rule and can be null meanwhile -([item serialization](history-events-usage.md#item-serialization-2026-09-23)). The exporter +([usage](sessions-events.md#usage)). The exporter emits Session-scoped input/output token gauges with the same Session and sampling time, independently of Docker, microsandbox, Kubernetes, or another provider. PostgreSQL retains those cumulative snapshots alongside the sample; query results diff --git a/contracts/agents-api/sessions-events.md b/contracts/agents-api/sessions-events.md new file mode 100644 index 000000000..cb30d840f --- /dev/null +++ b/contracts/agents-api/sessions-events.md @@ -0,0 +1,152 @@ +# Sessions, events and history + +This contract covers what happens inside a Session: sending input, the live event stream, and reading the durable history of Turns, Items and usage. The Session resource itself (creation configuration, retry identity, update, list and deletion) is in [Core wire behavior](wire-semantics.md). Message and function-result content is in [message content](message-content.md). The [Agents API guide](../../docs/api/public-agent-api.md) shows the calls with the SDK and HTTP. + +## Recovery model + +The event stream is live only; Turns and Items are durable. A client that needs every result: + +1. Opens `GET /v1/agents/sessions/{session_id}/events` before sending input. +2. After a disconnect, subscribes again and buffers new events. +3. Reads the Session, its Turns and its Items, deduplicates Items by ID and keeps finalized Items when it applies the buffered updates. + +A stream is an observer. Closing it never cancels admitted work, and a stream never replays events it did not send. After a lost input response, retry with the same `Idempotency-Key`, then read the Session, Turns and Items. + +## Session status + +A Session's `status` and `last_active_at` derive from its latest root Turn and from input still waiting for its Environment: + +| `status` | When | +| --- | --- | +| `idle` | No Turn yet, or the latest Turn completed or was cancelled. Input reserved for a hosted Environment that is still provisioning also reads `idle` | +| `in_progress` | The latest Turn is queued, running or waiting | +| `requires_action` | The latest Turn waits for a function result and no cancellation was requested, or input waits for a `self_hosted` machine to connect. `required_actions` lists `function_call` or `environment_connection` entries | +| `failed` | The latest Turn failed (`error` is "The execution could not complete."), input reserved for the Environment failed, initial input expired before admission, or the Environment failed to initialize (see [Environment initialization failure](#environment-initialization-failure)) | + +A Session stays usable after a Turn fails: new input starts a new Turn. Later reserved input that expires leaves the Session idle. An Environment initialization failure or expiry prevents new work. + +## Send input + +`POST /v1/agents/sessions/{session_id}/events` takes an ordered array of 1 to 64 events: `agent.session.input.message`, `agent.session.input.cancel` and `agent.session.input.tool_result`. The whole batch is admitted atomically, and the response is 202 with no body once the batch is stored, before any harness reads it. Admission never confirms native application. + +- **Limits.** The request body is at most 1 MiB, and the stored input of one request at most 512 KiB. +- **Null batch.** An explicit `null` for `events` is invalid. +- **Empty batch.** `{"events": []}` checks that the Session exists and returns 202. It creates no Turn, Item or retry identity. +- **Retries.** An `Idempotency-Key` of up to 128 bytes identifies the whole ordered batch. The same key and batch return 202 again without admitting anything twice; the same key with another batch returns 409 `idempotency_conflict`. A request without a key is always new. +- **Messages.** On an idle Session a message batch starts a queued Turn. While a Turn runs, messages join it (steering); they never start a parallel Turn. Each message stays its own user Item, even when the harness receives several as one prompt. +- **Cancellation.** A queued Turn is cancelled without a live Runtime. A running Turn is cancelled when the Runtime confirms it; completion can win that race. The Turn has stopped when it reads `cancelled`, not when the request returns. A cancellation on an idle Session with no pending input is accepted and has no effect; while an input reservation is pending, it returns 409. +- **Function results.** `turn_id`, `call_id` and `success` are required; `output` and `error` are optional and nullable ([content rules](message-content.md#function-results)). An identical repeated result returns 202 without another application or event. The result Item appears when the harness applies the result; a result that cancellation prevents from being applied stays stored but produces no Item. +- **Queueing.** A queued Turn starts when a Runtime that supports the Session's harness and configuration is connected and one of Core's [`core.execution_concurrency`](../../docs/configuration.md#settings) work slots is free. A Session stays bound to the Runtime that first ran it. +- **Execution availability.** A service without execution returns 503 `execution_unavailable`, and a Worker that loses execution ownership returns 503. A Session created without a model provider rejects new messages with 400 `model_provider_required` ([model execution](model-execution.md)). + +### Sessions with an Environment + +On `openai_hosted` and `self_hosted` Sessions, messages sent while a Turn runs join it at once. Messages sent to an idle Session reserve the batch for the Environment: the request waits until a Turn starts, for at most five minutes from the reservation. While the reservation waits for a `self_hosted` machine, the Session reads `requires_action` with an `environment_connection` action. A batch that carries messages on these placements may contain only messages. Cancellation-only and result-only batches are admitted at once and create no Turn. + +The waiting request ends with 202 when the Turn starts, or with 409 `environment_input_expired` when the deadline passes, 409 `environment_input_cancelled` when an administrator archives the Session or resets its deployment and cancels the reservation, or 409 `environment_unavailable` when the Environment fails or expires. Disconnecting the waiting request does not cancel the reservation or restart its deadline. + +### Input errors + +Checks run in this order: request validation, Session lookup, retry lookup, the Environment file-write gate for batches with a message, the pending-input gate, then each event in batch order. A rejected batch writes nothing and leaves any pending action unchanged. Every 409 has type `conflict_error` ([error envelope](wire-semantics.md)). + +| Case | Status and code | Message | +| --- | --- | --- | +| Earlier input to the Session still waits for admission, such as the reserved initial input of a provisioning hosted Session or of an offline `self_hosted` Session | 409 `conflict_error` | "Earlier input to this Session is still pending." | +| Input the Turn cannot accept in its state, such as a result after cancellation or after its Turn ended without a stored result | 409 `conflict_error` | "The Turn cannot accept this input in its current state." | +| A result that differs from the call's stored result, before or after its Turn ends | 409 `conflict_error` | "The tool call already has a different result." | +| The same `Idempotency-Key` with a different batch | 409 `idempotency_conflict` | "This idempotency key was used with different input." | +| A result whose `call_id` names no function call of this Session | 400 `invalid_request_error`, param null | "Unknown pending tool call." | +| A result for a call of this Session whose `turn_id` names another Turn, an unknown Turn ID or a value that is not a Turn ID | 400 `invalid_request_error`, param null | "The tool call belongs to a different Turn." | +| A missing, malformed or foreign Session | 404 `not_found_error` | "Resource not found." | +| New input after an `openai_hosted` Environment failed to provision | 409 `conflict_error` | "the hosted environment failed to provision" | +| New input after a `self_hosted` Environment failed, input already waiting when the Environment failed, or an expired Environment | 409 `environment_unavailable` | "The environment is no longer available for new input." | +| A message the Session's harness cannot take, such as whitespace-only text on Claude Code | 400 `unsupported_or_invalid_configuration` | See [whitespace-only text](message-content.md#whitespace-only-text) | + +An empty `turn_id` or blank `call_id` is the generic 400 `invalid_request`. Error messages never repeat caller input or internal identifiers. + +## Initial input at Session creation + +`POST /v1/agents/sessions` accepts `input` as a string (one text message) or an array of user messages, with the same validation and admission as the events endpoint. + +- Initial input is required on `none` (400 `invalid_request_error`, "conversation-only sessions currently require initial input") and for `stream: true` on every placement except `self_hosted` (400, "streaming session creation requires initial input"). These checks run before the creation retry lookup. +- The Session and its initial work commit in one transaction. On `none` that includes the first Turn and the input Items. On `openai_hosted` the input is reserved while the Environment provisions. On `self_hosted` it is reserved with an `environment_connection` action, and creation returns while the machine is offline. +- Reserved initial input has the same five-minute deadline as later input. When it passes before a Turn starts, the Session reads `failed` without a Turn. +- A creation retry returns the original Session and never admits its input again, including after later Turns ([creation retries](wire-semantics.md)). + +## Creation streaming + +`stream: true` on Session creation returns 201 with an event stream instead of JSON. + +1. The first event is `agent.session.created` with the same committed Session as the JSON 201 body, read after the commit. With initial input on `none` it already reads `in_progress`; on `self_hosted` it already shows the `environment_connection` action. +2. The stream then sends every committed event of the creation exactly once, starting at the creation's own position, so fast execution cannot skip its first events. +3. It ends right after the first `agent.session.idle` recorded when a Turn ends or an input reservation stops waiting (expired, cancelled or failed), or after any `agent.session.failed`, and never sends later events. `requires_action`, function results, resumed work and a `self_hosted` connection keep it open. A reservation keeps it open until a Turn settles or the reservation ends. +4. A creation that admitted nothing ends right after `agent.session.created`. When a settlement records no event, the stream ends after the events committed up to the settled state; another client's work committed before that point can still be sent. + +Input reserved while the ending Turn captures Artifacts can start a later Turn that the creation stream does not follow. + +A retry with the same `Idempotency-Key` and `stream: true` returns 201 with only the connection comment and ends at once: it admits nothing and follows no work. To recover a lost Session ID, repeat the request with the same key and `stream: false`, then read the Session, Turns and Items. Disconnecting stops only the observer. Observe later Turns with the GET stream. + +## Live event stream + +`GET /v1/agents/sessions/{session_id}/events` starts at the latest committed event and sends only events committed after that. `Last-Event-ID` is ignored. Events publish after their transaction commits. + +- **Lifetime.** The stream stays open across Turns and after a Turn fails. It ends when the Session is deleted or after the terminal `agent.session.failed` of an [Environment initialization failure](#environment-initialization-failure); a stream opened after that failure stays open. A keepalive comment is sent every 15 seconds. +- **Buffer.** Core keeps at most 256 events and 64 MiB of events per Session, plus one larger event when needed. A reader that falls behind the buffer receives an `error` event with type `server_error` and code `stream_interrupted`, and the stream closes. A socket write that blocks for five seconds also closes it. Execution never waits for a reader. +- **Key recheck.** An open stream checks the original Project key at most once per second while idle and before sending output. Revoking the key or archiving the Project closes the stream, and so does an authentication failure. A recheck uses the normal five-second authentication timeout and sends no Session data while it waits. Bytes already sent cannot be recalled. +- **Root work only.** Child Turns and child Items publish no Session events; `agent.session.subagent.*` events and root coordination Items do. Read child work through the [Subagent resources](subagents.md). + +### Event rules + +- Session events carry `event_id`, `type` and the `session` snapshot at that transition. Turn events carry `session_id` and `turn_id`. There is no Turn `waiting` event. +- A new Turn publishes `agent.session.turn.created`, the user `item.added`, `agent.session.in_progress`, then `agent.session.turn.in_progress`, all from one transaction. When a batch holds several messages, the later messages follow the Session activity. +- Terminal Turn events (`completed`, `failed`, `cancelled`) carry top-level `usage` copied from the Turn snapshot at that moment, null when unknown. Other events omit it. +- `item.added` and `item.done` always carry `output_index`, null for input Items. Function results emit `item.added` only; `item.done` is for agent output. +- An assistant message follows one sequence: `item.added` in progress with empty `content`, `content_part.added` with empty text, `output_text.delta` events, `output_text.done`, `content_part.done` and `item.done`. A message first observed complete, such as structured output, sends its whole text in one `output_text.delta`, byte for byte. The completed text replaces the accumulated deltas. +- Codex command output streams as `agent.output.command_execution_output.delta` with the command's Item ID and output index. Native output quotas and text conversion apply, so the deltas are not a byte-exact capture; the completed Item is authoritative. +- A cancelled or failed Turn marks its unfinished Items `incomplete` and keeps their partial content. +- A function call stays in `required_actions` until the harness applies its result, or cancellation or the Turn's end removes it. A repeated notification emits no new state. + +## Turns and Items + +Turn and Item lists take `after`, `limit` and `order` ([list rules](wire-semantics.md)). Cursors are IDs within the same Session. + +**Turns.** Session Turn routes hold root Turns only, ordered by creation time then ID; a child Turn ID returns 404 there. A failed Turn has `error: {code: "internal_error", message: "The execution could not complete."}` and never raw engine diagnostics. Administrators read the failure category through [Session diagnostics](session-diagnostics.md). + +**Items.** Items are ordered by the time they were first observed, then by their position in the Session, then by ID. Updates and retries never move an Item or change its `output_index`, a zero-based position among the Turn's output Items; input Items have none. Reads use the stored history index and never rebuild it from native journals. The Items list includes Items that are still in progress or incomplete. + +| Item `type` | Content | +| --- | --- | +| `message` | User or assistant content parts. `phase` is the harness's phase when it reports one (`commentary`, `final_answer`), otherwise null | +| `command_execution` | Command, reported output, exit code, duration and working directory | +| `mcp_call` | Server and tool identity, arguments, structured result or error | +| `function_call`, `function_call_output` | A linked call and its result. The result always carries `output` and `error`, null when the submission omitted them; stored results keep the submitted field presence. Native file changes appear as an `apply_patch` function call with the changes as arguments and no invented result | +| `web_search_call` | The supported action fields (`search`, `open_page`, `find_in_page`, `other`) | +| `reasoning`, `agent_message` and the Subagent coordination calls | See [Subagents](subagents.md). `agent_message` has no status; reasoning status can be absent or null | + +A failed tool does not fail its Turn. Tool output is readable by the Session's Project and can contain the tool's own diagnostic text. + +## Usage + +Turn and Session `usage` uses the pinned `TokenUsage` fields: input, cached input, output, reasoning output and total tokens. Null means unknown, never zero. + +- **Turn usage** is the latest complete snapshot the harness reported for that Turn. A new snapshot replaces the previous one; repeats never add. Stored snapshots survive cancellation and Worker restarts. +- **Session usage** is the sum of root Turn usage when every root Turn has ended (completed, failed or cancelled) with known usage. It is null while any root Turn is queued, running or waiting, and stays null once a root Turn ends without usage. Subagent Turns do not count. +- **By harness.** Codex reports measured snapshots while a Turn runs and at its end; a Turn interrupted before any usage report stays null. Claude Code and MiniMax Code report no complete public breakdown, so their usage is null. + +Usage is best-effort accounting of reported measurements. It is not an invoice, and Core never estimates missing usage. + +## Environment initialization failure + +When an Environment fails to initialize, whether an `openai_hosted` sandbox or a `self_hosted` machine, one transaction records the failure and three events, in this order: + +| Event | Payload | +| --- | --- | +| `agent.session.environment.failed` | `error: {type: "environment_error", code: "environment_connection_failed", message: "The environment failed to connect."}` | +| `error` | `error: {type: "environment_error", code: "sandbox_error", message: , param: null}` | +| `agent.session.failed` | The Session: `status: "failed"`, the reason as `error`, `required_actions: []` and the failure time as `last_active_at` | + +Session reads and lists return the same Session, and input that was waiting for the Environment settles as failed in the same snapshot. The GET stream and the creation stream end after `agent.session.failed`. New input returns 409 ([input errors](#input-errors)); the Session can be deleted. + +The [Environment initialization contract](environments.md#initialization-state-and-failure) defines the fixed failure reasons. + +Core's own `stream_interrupted` error event carries `type`, `code` and `message` without `param`; `error` events shaped like the official ones carry `param: null`. diff --git a/contracts/agents-api/source-files.md b/contracts/agents-api/source-files.md index 4419bb486..666f1b936 100644 --- a/contracts/agents-api/source-files.md +++ b/contracts/agents-api/source-files.md @@ -1,102 +1,110 @@ -# Referenced source Files +# Files and Skills -Environment `file_id` refers to a general Files API upload, not a local path or -an external provider's file. The contract uses the same SDK/source pin as -[upstream.json](upstream.json): [Files resource](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/files.py), -[create parameters](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/file_create_params.py) -and [FileObject](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/file_object.py). +Files (`/v1/files`) and Skills (`/v1/skills`) are Project resources with their own lifecycle, independent of Sessions. A File holds uploaded bytes that Environments copy by ID. A Skill holds immutable, versioned bundles that Templates and Sessions reference. Every API key of a Project shares them. -## Implemented source workflow +These routes follow the SDK pinned in [upstream.json](upstream.json): the [Files resource](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/files.py), [create parameters](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/file_create_params.py), [FileObject](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/file_object.py) and the [Skills resource](https://github.com/openai/openai-python/tree/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/skills). They need a Project API key and no `OpenAI-Beta` header. A missing ID and another Project's ID return the same 404. -| Operation | Current behavior | +## Files + +| Operation | Behavior | +| --- | --- | +| `POST /files` | Multipart upload with one `file` part and `purpose=user_data`, in either order. Returns 200 with the File | +| `GET /files` | Lists the Project's Files without reading their bytes | +| `GET /files/{file_id}` | Returns the File | +| `GET /files/{file_id}/content` | 400 `Not allowed to download files of purpose: user_data`, with a null code and param. The ID is checked first, so a missing File returns 404 | +| `DELETE /files/{file_id}` | Deletes the File and its bytes; returns `{"id": …, "object": "file", "deleted": true}` | + +Use a File by passing its ID to [Environment files](environment-files.md#create-a-file) or to a Template's or Session's initial `files` ([Environments](environments.md)). Those copies read the bytes internally; the public download stays refused. + +### Upload + +- Only `purpose=user_data` is accepted. Other purposes, `expires_after` and the Uploads API are not supported. +- The file may be empty and holds up to 512 MiB; the whole multipart body may exceed that by 64 KiB. A larger upload returns 413 `request_too_large`. The transfer must finish within five minutes. +- A missing, repeated or unknown part, a `Content-Encoding` or `Content-Transfer-Encoding` header, or a filename that is empty, longer than 1,024 bytes, not UTF-8 or contains NUL returns 400. Core stores nothing until the whole request validates. +- Core does not deduplicate uploads. After a lost response, list Files before uploading again. + +### File object + +| Field | Value | +| --- | --- | +| `id`, `object` | File ID; `file` | +| `bytes` | Size in bytes | +| `created_at` | Unix seconds | +| `filename` | The uploaded name. It is metadata only and never becomes a filesystem path | +| `purpose` | `user_data` | +| `status` | `processed`, meaning the bytes are stored. Core does not parse, index or scan them | +| `expires_at`, `status_details` | `null` | + +### List Files + +| Parameter | Rule | +| --- | --- | +| `order` | `desc` (default) or `asc`, by creation time, then ID | +| `after` | ID of a File this Project can see | +| `purpose` | One of `user_data`, `assistants`, `batch`, `fine-tune`, `vision`, `evals`, `assistants_output`, `batch_output`, `fine-tune-results`. Any other value, including a different case, returns 400 with `param: "purpose"` before the cursor is resolved. Values other than `user_data` return an empty page. An empty value means no filter | + +The response is `{"object": "list", "data": [...], "first_id", "last_id", "has_more"}`; an empty page has null IDs. `limit`, query parsing and their errors follow the shared [list rules](wire-semantics.md#lists). + +### Errors + +A missing or foreign File returns 404 with type `invalid_request_error`, a null code and `param: "id"` for retrieve, content and delete. + +### Storage and deletion + +Core stores File bytes as PostgreSQL large objects in its own database. An upload and a deletion each commit in one transaction, so a failure leaves neither partial bytes nor metadata. [Back up](../../docs/getting-started/operations.md#back-up) the database with its large objects; deleting a File does not remove it from write-ahead logs or earlier backups. + +The source Files schema refuses a downgrade while File rows remain. Delete Files through the API first so their large objects are removed. + +A copy into a workspace reads a consistent snapshot of the File and can finish after the File is deleted; later lookups fail. Deleting a File never changes a workspace copy. + +## Skills + +| Operation | Behavior | | --- | --- | -| `POST /files` | Multipart `file` and `purpose=user_data`; either part order; immutable bytes and metadata commit after full validation | -| `GET /files` | Project-scoped metadata listing with `after`, `limit`, `order` and `purpose`; deterministic creation-time/ID keysets | -| `GET /files/{id}` | Project-owned metadata, without reading the body | -| `GET /files/{id}/content` | Tenant-scoped lookup, then 400 for the supported `user_data` purpose; internal initialization/copy reads remain available | -| `DELETE /files/{id}` | Atomic metadata removal and body unlink; `id`, `object: file`, `deleted: true` | - -The configured SDK base URL includes `/v1`. These routes reuse bearer and optional -organization/project header validation but do not require `OpenAI-Beta`. Existing -Agents/Vault routes retain their Beta check. API keys in the same database-owned Project share the source resource. Every read/delete/copy lookup -uses that project partition; missing and foreign IDs return the same safe 404. - -Listing defaults to 10,000 resources and rejects limits outside the pinned -1–10,000 range. Omitted order uses descending creation order; `asc` and `desc` -use the stored timestamp plus ID as a deterministic keyset. The response includes -`object`, `data`, `first_id`, `last_id` and `has_more`; empty pages use null IDs. -The cursor must name a currently visible File in the same project. The optional -purpose filter accepts the nine values qualified by validation probes (including -`evals` and output-purpose names); an unknown or case-variant value returns 400 -with `param: purpose` before cursor resolution. Valid other-purpose filters return -an empty page because storage currently accepts only `user_data`. An explicit empty -purpose is treated as omitted and unknown query keys are ignored, as observed in the -[list query tolerance](list-query-semantics.md#list-query-tolerance--september-23-2026) batch. Successful official filtering/order, deleted-cursor behavior and -pagination during concurrent mutation remain unqualified. - -Metadata includes `id`, `object: file`, `bytes`, Unix-second `created_at`, -`filename`, `purpose: user_data`, deprecated `status: processed`, and nullable -`expires_at`/`status_details`. Here processed means stored bytes are available, -not parsed, indexed or scanned. Filename is metadata only and never a filesystem -path. The current service bounds it to 1–1024 UTF-8 bytes without NUL. - -The pinned general upload documentation states 512 MB. This implementation uses -512 MiB with a separate 64 KiB multipart-envelope allowance; exact hosted size-unit -and overhead/error parity are unverified. Streams use bounded chunks and a -five-minute transfer deadline. Complete multipart validation rejects missing, -duplicate, unknown or unsupported parts, invalid purpose and incomplete bodies. -Empty file bytes are valid. No partially validated upload is published. - -## Persistence and deletion - -The execution database owns source metadata and PostgreSQL large objects through -the already-pinned pgx driver. Upload and deletion are single transactions; a -rollback does not orphan a body or publish partial metadata. Object OIDs are private -and cannot be supplied through the API. No product tables, temporary local upload -directory, external object-storage service or model invocation are required. -Backups must include PostgreSQL large objects. Physical deletion from the live -database does not erase historical WAL/backups; database maintenance governs -reclamation. Downgrade refuses to drop a populated source table. - -An admitted internal content read uses an immutable database snapshot and may finish after -deletion. Later source lookups reject. Environment copy resolves up to its existing -50 MiB destination limit before invoking the same durable writer used by inline -uploads. Source deletion does not undo an admitted or completed workspace copy. -These concurrency/error choices are local policies, not verified hosted parity. -Ambiguous upload commits are not automatically retried; clients may need to retain -their source request evidence. Destination unknown-write handling remains unchanged. - -Missing source Files and missing cursors expose the measured `id` and `after` -error parameters without revealing foreign resource existence. See the bounded -[resource error qualification](resource-selector-semantics.md#source-file-errors). - -## Qualified resource semantics - -See [file resource qualification](file-resource-semantics.md) for owned official -metadata/content/purpose probes and the corresponding Core acceptance. Public -`user_data` content rejects with `invalid_request_error`, null code and null param; -missing or foreign IDs return the existing safe 404 before purpose is considered. -This does not restrict internal source consumption by Environment initialization -or workspace copies. Skill and Artifact downloads retain their separate rules. - -## Remaining scope and verification - -Other upload purposes, `expires_after`, resumable Uploads, quotas, -rate-limit parity, Artifacts and full status/error/header compatibility remain -unimplemented or unverified. Unsupported purposes/expiration are rejected. The -pinned request accepts `evals` while FileObject's purpose union omits it; this -discrepancy is recorded, not resolved by inventing a new contract. Current online -Agents limits require separate version qualification before changing the pinned -baseline. This workflow does not enable public hosted Session provisioning. - -Store tests use actual PostgreSQL for rollback, project isolation, independent -reads and concurrent deletion. The opt-in 512 MiB test exercises streaming storage. -API tests cover multipart ordering/validation and response/authentication behavior. -The fixed SDK and raw HTTP list regression covers default and bounded pages, -automatic continuation, purpose filtering, same-project sharing, foreign-project -isolation, restart and exact list envelopes against real PostgreSQL. -`official_environment_files_create.py` includes the fixed SDK/raw HTTP source -workflow through `official_source_files.py`; its invoking native fixture must -verify copied hashes, retained copies after source deletion, absence of leaked -database objects and real-model consumption. Controlled tests or SDK parsing alone -do not establish that execution acceptance or complete protocol compatibility. +| `POST /skills` | Uploads a new Skill. Its first version is both default and latest | +| `POST /skills/{skill_id}/versions` | Uploads a new version. A `default` form field of `true` makes it the default; `false` or omitted leaves the default unchanged | +| `GET /skills`, `GET /skills/{skill_id}` | Skill metadata, without decrypting any bundle | +| `POST /skills/{skill_id}` | `{"default_version": ""}` changes the default version | +| `DELETE /skills/{skill_id}` | Deletes the Skill and every version | +| `GET /skills/{skill_id}/content` | ZIP of the default version | +| `GET /skills/{skill_id}/versions`, `GET /skills/{skill_id}/versions/{version}` | Version metadata. The list orders by version number, and `after` is a version ID (`skillver_…`), not a number | +| `GET /skills/{skill_id}/versions/{version}/content` | ZIP of that version | +| `DELETE /skills/{skill_id}/versions/{version}` | See [Delete a version](#delete-a-version) | + +`limit`, cursors and query errors follow the shared [list rules](wire-semantics.md#lists). + +### Upload a bundle + +Send one ZIP as a `files` part, or a directory as repeated `files[]` parts whose filenames are relative paths such as `report/SKILL.md`. SDK 3.13.0 sends no part when `files` is a single file rather than a list, so upload a single ZIP with plain HTTP: + +```sh +curl "$OPENAI_BASE_URL/skills" -H "Authorization: Bearer $OPENAI_API_KEY" -F files=@report.zip +``` + +A bundle has one top-level folder containing `SKILL.md` and any supporting files: + +- `SKILL.md` is UTF-8, at most 256 KiB, and starts with YAML front matter. `name` is required: lowercase letters and digits, optionally separated by single `-` or `_`, at most 64 characters. `description` is required and non-empty. `license`, `compatibility` and a string-valued `metadata` map are optional; any other key is rejected. +- Entries are regular files or directories with clean relative paths. Links, special files, absolute paths, `..` components and duplicates are rejected. +- Limits: 5 MiB compressed, 20 MiB expanded, 500 files, and 1,000 ZIP entries including directories. + +Core encrypts each version's bundle bound to its Project, Skill and version. ZIP uploads keep executable bits; directory uploads store files with mode 0644. + +### Versions and metadata + +- Version numbers start at 1, increase by one per upload and are never reused, even after the latest version is deleted. Template and Session selectors name versions by number, so a reused number could point a stored selector at different bytes. +- The Skill's `name` and `description` are those of its default version. Changing the default, by `POST /skills/{skill_id}` or by uploading with `default=true`, updates the pointer and both fields together; `id` and `created_at` stay the same. +- `latest_version` is the highest remaining version. +- Uploads and deletions of one Skill run one at a time, so a deletion never removes a version whose upload was acknowledged. + +How Templates and Sessions select a version (default, `latest` or a number) and freeze its bytes is in [Environments](environments.md#skills-plugins-and-environment-mcp). + +### Delete a version + +| Version | Result | +| --- | --- | +| The default, and the only version | 200 `{"id": "skillver_…", "object": "skill.version.deleted", "deleted": true, "version": "1"}`. The Skill is deleted in the same transaction | +| The default, while other versions exist | 400, type `invalid_request_error`, code `invalid_value`, `param: "version"`, `Cannot delete the default skill version.` | +| Any other version | 200 with the same body. If it was the latest, `latest_version` falls back to the highest remaining | +| A missing or foreign Skill or version | 404 | + +Deleting a Skill or a version does not change Sessions that already installed it; Templates keep the reference they stored. diff --git a/contracts/agents-api/v1/upstream_contract_test.go b/contracts/agents-api/v1/upstream_contract_test.go index d5b7d6395..57b44bdd5 100644 --- a/contracts/agents-api/v1/upstream_contract_test.go +++ b/contracts/agents-api/v1/upstream_contract_test.go @@ -93,7 +93,7 @@ var fieldPlacementPending = map[string]string{ } // Official list pages also carry object, first_id and last_id (recorded in -// official-semantics-alignment.md), which the SDK's hand-written page classes +// wire-semantics.md#lists), which the SDK's hand-written page classes // do not declare. var listEnvelopeFields = []string{"object", "first_id", "last_id"} diff --git a/contracts/agents-api/vaults.md b/contracts/agents-api/vaults.md new file mode 100644 index 000000000..ca3ffe511 --- /dev/null +++ b/contracts/agents-api/vaults.md @@ -0,0 +1,146 @@ +# Vaults and Credentials + +A Vault is a Project-owned container of Credentials. A Credential holds the secret for one HTTPS MCP server: a `static_bearer` token or an `mcp_oauth` grant. A Session attaches Vaults in `vault_ids`; Core selects one Credential per HTTP MCP server when the Session is created and hands the decrypted token to the Runtime only when it dispatches work. Secrets are write-only: no read returns a token, refresh token, client secret or ciphertext. + +The application owns OAuth authorization and consent, provider revocation and any approval policy. Core has no authorization redirect, callback, revocation or public refresh endpoint, and creating or replacing a Credential never contacts the MCP server or the OAuth provider. + +## Store and use a credential + +```python +vault = client.beta.agents.vaults.create(name="internal") +credential = client.beta.agents.vaults.credentials.create( + vault.id, + name="Internal MCP", + auth={ + "type": "static_bearer", + "mcp_server_url": "https://mcp.example.com/endpoint", + "token": token_from_private_configuration, + }, +) +session = client.beta.agents.sessions.create( + environment={"type": "none"}, + vault_ids=[vault.id], + input="List the open incidents.", + agent={ + "model": "your-model-id", + "tools": [{ + "type": "mcp", + "server_label": "internal", + "transport": {"type": "http", "server_url": "https://mcp.example.com/endpoint"}, + }], + }, +) +``` + +The MCP tool may name the Credential in `credential_id`; without it, Core picks the attached Credential whose `mcp_server_url` equals the tool's `server_url` ([selection](#credential-selection-in-a-session)). Which Harness and placement can connect to the server depends on the tool's `connection_origin` ([MCP connection origin](environments.md#public-mcp-connection-origin)). + +## Routes + +All routes are under `/v1`, take a Project API key and require `OpenAI-Beta: agents=v1`. The Project's keys share its Vaults; another Project's Vault or Credential answers 404, the same as a missing one. + +| Operation | Route | Result | +| --- | --- | --- | +| Create a Vault | `POST /vaults` | 201 Vault | +| Retrieve a Vault | `GET /vaults/{vault_id}` | Vault | +| List Vaults | `GET /vaults` | List of Vaults | +| Delete a Vault | `DELETE /vaults/{vault_id}` | `{id, object: "vault.deleted", deleted: true}` | +| Create a Credential | `POST /vaults/{vault_id}/credentials` | 201 Credential | +| Retrieve a Credential | `GET /vaults/{vault_id}/credentials/{credential_id}` | Credential | +| List Credentials | `GET /vaults/{vault_id}/credentials` | List of Credentials | +| Replace secrets | `POST /vaults/{vault_id}/credentials/{credential_id}` | Credential | +| Delete a Credential | `DELETE /vaults/{vault_id}/credentials/{credential_id}` | `{id, object: "vault.credential.deleted", deleted: true}` | + +A malformed, missing or foreign ID returns 404 `not_found_error`, and so does a Credential addressed through a Vault that does not own it. Both lists order by creation time, then ID, and filter by `status` (`active`, `archived` or both); the [list rules](wire-semantics.md#lists) give the paging and filter details. Core stores the status privately, defaults it to `active` and has no operation that archives a Vault or Credential. A Credential's status is independent of its Vault's. + +## Vaults + +| Field | Rules | +| --- | --- | +| `name` | Optional. Omitted stays `null`; explicit `null` is rejected. A string is trimmed and must then hold 1–256 UTF-8 bytes | +| `metadata` | Omitted or `null` becomes `{}`. Values must be strings; another type returns 400 `invalid_request_error` with param `metadata.`. The encoded object is limited to 64 KiB, with no pair-count or length limits | + +A Vault reads as `id`, `object: "vault"`, `created_at`, `name` and `metadata`. There is no update route. + +Deleting a Vault removes it and all its Credentials in one transaction, whatever their status. It needs no storage key and sends no request to any provider. The effects on Sessions are under [Deletion](#deletion). + +## Credentials + +Creation takes a required `name` (trimmed, 1–256 UTF-8 bytes) and an `auth` object whose `type` is `static_bearer` or `mcp_oauth`. `mcp_server_url` must be an absolute HTTPS URL without userinfo or fragment; Core keeps it byte for byte, including any query, and performs no DNS or HTTP request. + +A Credential reads as `id`, `vault_id`, `name`, `object: "vault.credential"`, `created_at`, `updated_at` and `auth`. `auth` holds `type` and `mcp_server_url`; an OAuth Credential adds `expires_at` and `refresh` with `client_id`, `token_endpoint`, `token_endpoint_auth.type`, `resource` and `scope`. Reads and lists need no storage key. + +### Static bearer + +`auth` is `{type: "static_bearer", mcp_server_url, token}`. `token` is a required nonempty string. Core stores it as opaque bytes without trimming. To run, the token must be an RFC 6750 `b64token`; a Session that selects a stored token with other characters, such as whitespace, fails at dispatch. + +Replace the token with `POST /vaults/{vault_id}/credentials/{credential_id}` and exactly `{"auth": {"type": "static_bearer", "token": "…"}}`. A missing, null or empty token or any other field is rejected. Only the token and `updated_at` change; the ID, Vault, name, type, `mcp_server_url`, `created_at` and every Session binding stay. A failed replacement leaves the old token in place. + +### OAuth + +`auth` is `{type: "mcp_oauth", mcp_server_url, access_token, expires_at, refresh}`: + +| Field | Rules | +| --- | --- | +| `access_token` | Required, nonempty | +| `expires_at` | Optional, nullable RFC 3339 timestamp. Core accepts an expired value at creation | +| `refresh` | Optional, nullable: `client_id`, `refresh_token`, HTTPS `token_endpoint` and `token_endpoint_auth` are required; `resource` and `scope` are optional nullable strings | +| `refresh.token_endpoint_auth` | `{type: "none"}` with no `client_secret` member, or `client_secret_basic` / `client_secret_post` with a write-only `client_secret` | + +Replace grant material with the same update route and `auth.type: "mcp_oauth"`. The patch must change at least one of `access_token`, `expires_at`, `refresh.refresh_token`, `refresh.token_endpoint_auth.client_secret` or `refresh.scope`; an empty `access_token` is rejected. The ID, name, type, `mcp_server_url`, `client_id`, `token_endpoint`, `resource` and endpoint authentication method never change, and a `refresh` block cannot be added to a Credential created without one. `token_endpoint_auth`, when sent, must name the stored method, which must be `client_secret_basic` or `client_secret_post`. + +| Update field | Omitted | `null` | +| --- | --- | --- | +| `access_token` | Keep | Keep | +| `expires_at` | Keep; cleared when a new `access_token` is sent | Clear | +| `refresh` | Keep | Keep | +| `refresh.refresh_token` | Keep | Keep | +| `refresh.scope` | Keep | Stop sending a scope | +| `refresh.token_endpoint_auth` | Keep | Keep | +| `refresh.token_endpoint_auth.client_secret` | Keep | Keep | + +Changing a Credential's `auth.type` through an update returns 400. + +## Credential selection in a Session + +`vault_ids` on Session creation lists the Vaults whose Credentials the Session may use. Omitted, `null` and `[]` attach none; a `null` entry is invalid; every Vault must belong to the Project, or creation returns 404 `not_found_error`. Saving `credential_id` on an Agent authorizes nothing; only the Session's attachments do. + +For each HTTP MCP tool, Core selects a Credential among the attached Vaults when the Session is created: + +- With `credential_id`, that Credential must be in an attached Vault and its `mcp_server_url` must equal the tool's `server_url`. +- Without it (omitted or `null`), the one static or OAuth Credential whose `mcp_server_url` equals `server_url` is selected. With none, the tool connects anonymously. With several, creation fails; Core never prefers one type. + +Selection runs after the inline Agent's validation and the input requirement, and before anything is written. A rejected creation writes nothing. Failures have a null `param`: + +| Case | Response | +| --- | --- | +| `credential_id` with no attached Vault | 400 `invalid_request_error`: `MCP credential_id requires an attached vault` | +| `credential_id` missing, malformed, in another Project or in a Vault not attached | 400 `invalid_request_error`: `MCP credential_id was not found in an attached vault` | +| Credential in an attached Vault for another URL | 400 `invalid_request_error`: `MCP credential_id does not match server_url ` | +| Several Credentials match without `credential_id` | 409 `conflict_error`: `multiple attached vault credentials match MCP server_url ; specify credential_id` | +| Unknown or foreign Vault in `vault_ids` | 404 `not_found_error` | + +`` and `` repeat the request's values only when they are at most 256 bytes of printable UTF-8; otherwise the message leaves them out. The missing, foreign, unattached and malformed cases give byte-identical responses for the same ID, so a reference reveals nothing about Vaults the caller did not attach. The URL in a mismatch is the tool's. + +The Session freezes its attachments and each selection, anonymous ones included. Later changes to the Vaults never reselect: an identical creation retry returns the original Session and selection. Session reads, lists and event snapshots show an implicitly selected Credential's ID in the tool's `credential_id`, also after that Credential is deleted; anonymous tools show `null` and an explicit ID reads as sent. The stored request keeps the caller's value, so retries compare the original request. + +At each dispatch Core rechecks the Project, the attached Vault, the selected Credential, its type and the exact URL, then decrypts the token only into the Runtime's execution request. Reads of the Session, its configuration and its history never contain it. A server with a selected Credential runs only on a Runtime that advertises `mcp_http_bearer_auth`. A missing Credential, a failed decryption or a failed refresh fails the work; Core never falls back to another Credential or to an anonymous connection. + +## OAuth refresh + +At dispatch, an OAuth Credential whose `expires_at` has passed is refreshed before Core sends the work to the Runtime. A grant with a null `expires_at` is used as is and never refreshed proactively. An expired grant without `refresh` fails the work. + +The refresh is one `refresh_token` exchange with the stored endpoint authentication method, `scope` and `resource`, with no redirects, no method probing and no retry after a failure. The provider must return a `Bearer` token whose expiry, if any, is in the future. Core commits the new access token, the new expiry (or none) and any rotated refresh token before it uses the token; an omitted refresh token keeps the old one. A PostgreSQL row lock serializes refresh with manual replacement and with deletion, so a stale refresh never restores a deleted or replaced grant. A failed exchange or commit returns nothing and is not retried. When a provider rotates the grant and the local commit then fails, the application must reauthorize. Provider error bodies are neither returned nor logged, and Harnesses receive only the access token, never the refresh token or client secret. + +Token endpoints must be HTTPS. Core resolves the host, rejects loopback, private, link-local and shared (`100.64.0.0/10`) addresses, and dials the checked address so a second DNS answer cannot redirect it; TLS still verifies the host name. It ignores ambient HTTP proxies. For a private issuer, the operator lists its exact HTTPS origin in [`core.oauth_trusted_origins`](../../docs/configuration.md#settings); that allows private addresses for that host and port only, never plain HTTP, redirects or invalid certificates. The issuer's CA must be in Core's trust store. + +## Storage key + +Core seals every token, refresh token and client secret with AES-256-GCM under the installation's [`secrets/credential.key`](../../docs/configuration.md#installation-directory), bound to the Project, Vault, Credential, auth type and `mcp_server_url`. A wrong key, a modified row or a row moved to another binding fails to decrypt. Names are metadata outside the binding. The key and plaintext tokens exist in trusted service memory; encryption protects stored secrets and does not protect against a compromised service host. + +Without a configured key, Credential creation and replacement return 503 `credential_storage_unavailable` before writing; reads, lists, deletion and Vault operations still work. An unreadable or malformed key file stops Core at startup. Losing or replacing the key makes every stored secret unusable; Core supports one key, with no rotation or re-encryption. + +## Deletion + +Deleting a Credential, or its Vault, removes the credential rows. It does not erase secrets from PostgreSQL pages, write-ahead logs, backups or native history. Afterwards its reads, updates and repeated deletion return 404, lists omit it, new Sessions cannot select it and new Credentials cannot be created in a deleted Vault (a missing storage key still answers 503 first). + +Existing Sessions keep their frozen attachments and selections, and their history stays readable. The next dispatch that needs a deleted Credential fails; Core neither selects another Credential nor connects anonymously. Deletion does not cancel running work, withdraw a token already sent to a Runtime or revoke the grant at its provider: cancel the Session and revoke the grant yourself when needed. Revoked or invalid OAuth grants likewise fail until replaced. diff --git a/contracts/agents-api/wire-semantics.md b/contracts/agents-api/wire-semantics.md new file mode 100644 index 000000000..13bbbe92a --- /dev/null +++ b/contracts/agents-api/wire-semantics.md @@ -0,0 +1,247 @@ +# Core wire behavior + +The pinned OpenAI Python SDK ([upstream.json](upstream.json)) defines the `/v1` routes, fields and types. This page states what Core does where those types are silent, such as status codes, error fields, defaults and list bounds, and where Core behaves differently from the official service. The [coverage ledger](README.md) lists the differences and open gaps; [Sessions, events and history](sessions-events.md), [message content](message-content.md), [Vaults](vaults.md), [source Files and Skills](source-files.md) and [Environment files and Artifacts](environment-files.md) own the rules of their resources. + +"Beta routes" below are the routes under `/v1/agents` and `/v1/vaults`. "Files and Skills" are the routes under `/v1/files` and `/v1/skills`, which ignore `OpenAI-Beta`. + +## Requests + +### Paths and methods + +| Case | Core behavior | +| --- | --- | +| Empty, `.` or `..` path segments | Served on the canonical path, never redirected. Segments resolve with Go ServeMux semantics and a trailing slash is kept, so `/v1/agents/x/../` reaches the trailing-slash 404. | +| Percent-encoded unreserved characters (`A–Z`, `a–z`, `0–9`, `-`, `.`, `_`, `~`) | Decoded before routing, including `%2E` dot segments. Other escapes, such as `%2F`, `%5C` and double encodings, stay encoded and never separate segments. Every spelling reaches the route and authentication of its canonical path. | +| `HEAD` on a `GET` route | Runs the `GET` route after the same Beta and authentication checks and returns its headers without a body. | +| `HEAD` on the event stream, on File, Skill, Skill version and Artifact content, and on the Environment files list | 405, so `HEAD` never holds a stream open or reads content. | +| Unsupported method, including unknown methods such as `FOO` | 405, code `unsupported_operation`, message "This API method is not supported.", and an `Allow` header listing the route's methods in the order `GET,HEAD,POST,DELETE`. | +| Unknown sub-route under `/v1`, including a trailing slash | 404, code `unsupported_operation`, after the Beta and authentication checks. | +| `OPTIONS` and CORS | No CORS handling. | + +Creating an Agent, Vault, Credential, Environment Template, Environment file or Session returns 201, also for a streamed Session creation. An empty update body on an Agent or Environment Template advances `updated_at` and changes nothing else; timestamps have one-second precision. + +### Headers + +Beta routes require exactly one `OpenAI-Beta` header value, equal to `agents=v1`. A missing, different or repeated value returns 400 with type and code `invalid_beta` and the message "To access the Agents API, set the 'OpenAI-Beta' header to 'agents=v1'." This check runs before authentication. `agents=v0` is rejected. + +Every Agents API response, including errors and event streams, carries: + +| Header | Value | +| --- | --- | +| `X-Request-Id` | A fresh `req_` followed by 32 lowercase hex characters. Core also logs it as `request_id`. A caller-supplied value is not echoed. | +| `OpenAI-Version` | `2020-10-01` | +| `OpenAI-Processing-Ms` | Handling time when the headers are written | +| `X-Content-Type-Options` | `nosniff` | +| `Cache-Control` | `no-store` on JSON responses | + +### Authentication + +`/v1` accepts only a Project API key as `Authorization: Bearer `. All keys of one Project act as the same caller: they share its resources and its Session creation retries. Core resolves the key and its Project in the database on every request, with no credential cache and a five-second timeout. Revoking a key or archiving its Project takes effect on the next request. All keys of a Project act as subject `service_account/project:`. [Projects and keys](admin-api.md#projects-and-keys) describes key management. + +The optional `OpenAI-Organization` and `OpenAI-Project` headers must, when sent, appear once and equal `core` and `proj_`; any other value rejects the key. + +| Failure | Response | +| --- | --- | +| No key, another scheme, an empty or repeated `Authorization` header, an unknown or revoked key, a key of an archived Project, the Core key, or mismatched scope headers | 401, type `invalid_request_error`, message "A valid Agents API bearer key is required.", `WWW-Authenticate: Bearer`. The code is null on Beta routes. On Files and Skills it is `invalid_api_key` when exactly one Bearer credential was sent and rejected, and null otherwise. | +| The key lookup fails, for example because the database is unavailable | 503, type `server_error`, code `authentication_unavailable` | + +### Request bodies + +Every `/v1` JSON route passes one body gate before route decoding, validation or lookup: Agent create and update, Vault create, Credential create and update, Environment Template create and update, Environment file create, Session create and update, and Session events. DELETE routes, multipart Files and Skills uploads and Skill update keep their own readers. Except for the 413 size error below, gate errors are 400 with type and code `invalid_request_error` and a null param. + +| Order | Case | Response | +| --- | --- | --- | +| 1 | `Content-Type` missing, not JSON or malformed, including a bodyless `POST`. `application/json` and `application/*+json` are accepted case-insensitively, with parameters | "expected request with Content-Type: application/json" | +| 2 | Body over the route limit: 16 MiB for Session create and Environment Templates, the [Environment file](environment-files.md) bound for file create, 1 MiB elsewhere | 413, code `request_too_large` | +| 3 | Invalid UTF-8 | "Invalid body: encountered a unicode decode error when parsing this JSON value. Please check the value to ensure it is valid unicode." | +| 4 | Malformed JSON, trailing data, two values, a byte order mark, a whitespace-only body, or an escape that forms a lone UTF-16 surrogate | "Invalid body: failed to parse JSON value. Please check the value to ensure it is valid JSON. (Common errors include trailing commas, missing closing brackets, missing quotation marks, etc.)" | +| 5 | A key repeated in one object, at any depth | "Invalid body: duplicate JSON key '' at ''. Duplicate JSON keys are not supported." The path joins object keys with `.` and omits array indices, such as `metadata.k` or `tools.type`. Keys compare after unescaping and case-sensitively. The first repeat in document order is reported. | +| 6 | A root that is not an object, including an array | "Invalid type: expected an object, but got instead." | +| — | An empty body or `null` | Treated as `{}` | + +Member names match exactly. A case variant such as `Metadata` or a nested `Role` is an unknown member and gets the route's unknown-member error before any write. + +Errors guarded by `echotext.Allowed`, such as unknown-member, enum, schema-root and cursor errors, repeat caller values only when they are at most 256 bytes of printable UTF-8. An unknown member that cannot be repeated gets a generic message and a null param. Metadata errors use their own validation and may repeat longer keys. + +### Resource identifiers + +A malformed path identifier gets exactly the response of a well-formed missing one on that route, including when the body or query is also invalid. Missing, malformed and foreign resources are indistinguishable. Identifiers that are UUIDs also resolve when written in another spelling that Go's UUID parser accepts, such as uppercase, braces or `urn:uuid:`. + +## Errors + +### Error types + +| Status | `type` | `code` | +| --- | --- | --- | +| 400 for a missing or invalid `OpenAI-Beta` | `invalid_beta` | `invalid_beta` | +| 404 for a missing, malformed or foreign resource on Beta routes | `not_found_error` | `not_found_error`, message "Resource not found." | +| 404 for a missing File or Skill | `invalid_request_error` | null | +| 401 | `invalid_request_error` | See [Authentication](#authentication) | +| 409 | `conflict_error` | `conflict_error` for conflicts the official service reports; Core-only conflicts keep their own code, such as `idempotency_conflict` | +| 5xx | `server_error` | Core's code | + +Other 400 responses use type `invalid_request_error`. Validation failures with an official equivalent use code `invalid_request_error` and the observed param and message; other request errors keep Core's local codes, such as `invalid_request` or `unsupported_or_invalid_configuration`. + +### Validation errors + +| Case | Response | +| --- | --- | +| More than 16 metadata pairs, a key over 64 characters or a value over 512 characters on Agent or Session create or update | Param `metadata` or `metadata.` and the official message with the actual count or length. Pairs are checked before keys and values, keys in sorted order. | +| A non-string metadata value on Agent or Session create or update, or Vault create | Param `metadata.`, message "Invalid type for 'metadata.': expected a string, but got instead." The first such value in document order is reported first. | +| Agent `name` over 128 characters | Param `name`. Empty and untrimmed names are accepted. | +| U+0000 in a metadata key or value | Param `metadata.` | +| U+0000 or invalid UTF-8 in any other stored string or query filter | Null param, message "Request text contains characters this service cannot store or compare, such as U+0000 or invalid UTF-8." Nothing is written. PostgreSQL cannot store U+0000, which the official service accepts. | +| Environment Template or inline Session network rejections: wildcards, ports, schemes, IPv6, empty hosts, a `restricted` policy without domains, more than 100 domains, domains with another access mode | Null param, Core's message | +| Empty Session update body | "At least one update field is required" | + +## Lists + +Lists return `object: "list"`, `data`, `has_more`, `first_id` and `last_id`; an empty page has null first and last IDs. `order` defaults to `desc`. [Environment files](environment-files.md) page with their own `page` token and are not covered here. + +### Query parameters + +| Case | Beta lists | Files | Skills and Skill versions | +| --- | --- | --- | --- | +| Unknown key, including `tenant_id` | Ignored, on lists and on single-resource routes | Ignored | Ignored | +| Repeated supported key, including a scalar `status` | 400 `invalid_request_error`, null param, "Failed to deserialize query string: duplicate field ``" | 400 `unsupported_parameter` | 400 `duplicate_parameter`, param ``, the official message | +| `order` other than `asc` or `desc`, including an explicit empty `order=` | 400 `invalid_request_error`, null param, "Failed to deserialize query string: order: unknown variant ``, expected `asc` or `desc`" | 400, null code, "order must be asc or desc." | 400 `invalid_value`, param `order`, "Invalid value: ''. Supported values are: 'asc' and 'desc'." | + +The pinned Python SDK drops empty query values, so `list(order="")` sends no `order` and uses the default. `after` is trimmed of surrounding whitespace. Checks run in this order: repeated keys, then `limit`, then `order`; Vault and Credential `status` is checked first. These checks run before any resource lookup. + +Vault and Credential lists accept `status` as a scalar, as `status[]` entries, or both, and filter by their union. Both statuses are listed by default. Another value returns 400 `invalid_request_error` with a null param and "Failed to deserialize query string: status: data did not match any variant of untagged enum VaultStatusFilterParam". + +### Page size + +| Lists | Default | Accepted | Other values | +| --- | --- | --- | --- | +| Agents, Sessions, Items, Environment Templates, Subagent Items, Subagent Turn Items | 20 | 1–100 | 0 becomes 1; above 100 becomes 100 | +| Vaults, Credentials | 20 | 1–100 | 0, negative and larger integers, including overflowing ones, are clamped into 1–100 | +| Turns, Subagents, Subagent Turns, Artifacts | 20 | 1–100 | 400 `invalid_request_error`, "limit must be between 1 and 100" | +| Skills, Skill versions | 20 | 0–100 | 0 returns an empty page whose `has_more` reports whether a resource follows the cursor. Negative: 400 `integer_below_min_value`, param `limit`. Above 100: 400 `integer_above_max_value`, param `limit` | +| Files | 10000 | 1–10000 | 400 with a null code, "limit must be between 1 and 10000." | + +A `limit` that is not a decimal integer, including an empty value, returns 400 `invalid_request_error`, "Failed to deserialize query string: limit: invalid digit found in string" on Beta lists; outside Vault and Credential lists, a value above the signed 64-bit range returns "Failed to deserialize query string: limit: number too large to fit in target type". A leading `+` is accepted when encoded as `%2B`; a leading `-` returns the invalid-digit error on Beta lists except Vaults and Credentials. Skills return `invalid_request`, "limit must be an integer between 0 and 100."; Files return `invalid_request` with the Files range message. + +### Cursors + +`after` names a resource of the same list, inside its already resolved parent and tenant. The parent is resolved first: a missing or foreign parent returns its 404 before the cursor is read. A cursor that does not resolve, whether random, malformed, of another type, of another parent, deleted or of another tenant, returns: + +| Lists | Response | +| --- | --- | +| Agents, Sessions, Turns, Environment Templates, Vaults, Credentials | 404, type and code `not_found_error`, "Resource not found." | +| Session Items, Subagent Items, Subagent Turn Items | 400 `invalid_request_error`, null param, "Invalid session item ID in `after`" | +| Subagents, Subagent Turns | 400 `invalid_request_error`, null param, "Invalid resource ID in `after`" | +| Session Artifacts | 400 `invalid_request_error`, null param, "after is not a valid artifact ID" | +| Skill versions | A value that does not begin with `skillver`: 400 `invalid_value`, param `after`, "Invalid 'after': ''. Expected an ID that begins with 'skillver'." A version of another Skill: the same fields, "Skill version cursor does not match this skill." A malformed `skillver` suffix or a missing, deleted or foreign version: 404 with a null code and param | +| Skills | 404 with a null code and param | +| Files | 404, param `after` | + +## Agents + +### Saved configuration + +Agent create requires `model`. Core saves and returns these values for omitted fields: + +| Field | Saved value | +| --- | --- | +| `name`, `instructions` | null | +| `metadata` | `{}` | +| `tools` | `[]` | +| `text` | `{"format": {"type": "text"}, "verbosity": "medium"}` | +| `reasoning` | Saved as sent; an omitted effort stays unset rather than taking a model default. Agent and Session responses always carry `reasoning.effort` and `reasoning.summary`, null when unset | +| `service_tier` | `auto` | +| `multi_agent` | Disabled. When enabled without `max_concurrent_subagents`, 6 | +| Function `defer_loading` | `false` | +| `programmatic_tool_calling.enabled` | `true` | +| `web_search` | Every pinned mode is saved; see [tool policy](execution-tools.md#web-search-and-programmatic-tool-calling) | +| HTTP MCP transport | Saved with `headers: {}`; nonempty headers are rejected. Origin and allowlist defaults are in [public MCP connection origin](environments.md#public-mcp-connection-origin) | + +Saving a value does not make it executable. Session creation admits a smaller set; see [Session admission](#session-admission). + +### Configuration validation + +Agent create and update bodies and the inline `agent` of Session create are checked against the pinned shapes of `tools`, `text`, `reasoning`, `service_tier`, `multi_agent`, `model`, `name` and `instructions`, before their parsers and before Harness admission. Failures return 400 with type and code `invalid_request_error`: + +| Case | Param | Message | +| --- | --- | --- | +| Missing required member | JSON path, such as `tools[0].parameters`; on Session create `agent.tools[0].parameters` | `Missing required parameter: ''.` | +| Unknown member, including a case variant and the unpinned `tool_choice` | JSON path | `Unknown parameter: ''.` | +| Wrong JSON type | JSON path | `Invalid type for '': expected , but got instead.` | +| Unsupported enum value | JSON path | `Invalid value: ''. Supported values are: ...` with the pinned values | +| Integer below minimum | JSON path | `Invalid '': integer below minimum value. Expected a value >= 1, but got instead.` | +| Repeated function name, more than one `web_search`, more than one `tool_search` | null | `duplicate function tool name: `, `duplicate web_search tool`, `duplicate tool_search tool` | +| Function `parameters` with a string root `type` other than `object` | null | `Invalid schema for function '': schema must be a JSON Schema of 'type: "object"', got 'type: ""'.` | +| `text.format` JSON schema with a string root `type` other than `object` | null | `agent.text.format.schema must have top-level type "object"; got ""`, also on Agent requests | + +Within one object Core reports a union's `type` first, then unknown members, then member values in document order, then missing members; tools before `text`, and the whole object before the duplicate and schema-root checks. Schemas without a string root `type` are not checked. Function and output schemas, MCP `transport`, `request_metadata`, `metadata` and `x_agents_core` keep their own parsers. Update bodies and the inline Session agent are validated before the Agent lookup, so owned, foreign, missing and malformed Agent IDs give the same response. + +Core saves values the pinned shapes allow even when it cannot run them: function names of any length, enabled programmatic tool calling, reasoning effort `max` and service tier `flex`. + +### Session admission + +A Session's effective configuration must also pass execution admission, which applies to saved and inline configuration alike. Admission reports protocol errors from the table above first, including duplicate tools and schema roots in saved Agents, then these, all 400 `unsupported_or_invalid_configuration` before any write: + +| Configuration | Message | +| --- | --- | +| Explicit `reasoning.effort` or `reasoning.summary` | "Explicit reasoning execution options are not supported by this service yet." | +| `service_tier` other than `auto` | "Execution currently supports service_tier=auto only." | +| Enabled or omitted-mode `web_search`, enabled `programmatic_tool_calling` | See [tool policy](execution-tools.md#web-search-and-programmatic-tool-calling) | +| More than 64 functions, or a function name that is blank or longer than 512 bytes | "This service supports at most 64 function tools." or "Function names must be nonempty, unique and at most 512 bytes." | +| Two `programmatic_tool_calling` declarations, two MCP servers with one label | "Execution requires distinct tool controls.", "Execution requires distinct MCP server labels." | + +A per-Session `tools` replacement admits a Session whose saved tools would be rejected. Support for each tool and Harness is in [execution and tools](execution-tools.md). + +Omitted, null and explicit `medium` text verbosity give the same Session configuration. For a model whose native catalog declares no verbosity support, the Codex adapter drops a `medium` setting and uses the model's default, and rejects `low` or `high`. + +### Update, delete and list + +| Operation | Core behavior | +| --- | --- | +| `POST /agents/{agent_id}` | Replaces only the supplied fields. Nested objects replace the whole field; null `name` or `instructions` clears it; null or `{}` metadata clears all pairs, and an object replaces them. Existing Sessions keep their snapshots. | +| `DELETE /agents/{agent_id}` | Returns `{id, object: "agent.deleted", deleted: true}`. Sessions created from the Agent, their history and their creation retries are unaffected. A repeated or missing deletion returns 404, and new Sessions that name the Agent return 404. | +| `GET /agents` | Pages by creation time, then ID. | + +## Sessions + +### Configuration snapshot + +Session creation copies the effective Agent configuration into an immutable snapshot. With `agent_id`, the saved Agent is read once; fields in the inline `agent` replace the saved field whole, including arrays, and null `tools` clears the list. Omitted fields inherit; an inline `x_agents_core` that omits `harness` keeps the saved harness. Saved Agent metadata never becomes Session metadata. Later Agent updates or deletion affect only new Sessions. + +`stream` defaults to false. `stream` and `agent_id` cannot be null. Omitted or null `metadata` is `{}`. + +Creation validates the body and metadata types, the request fields and initial input, and the placement and streaming input requirements before looking up a creation retry. For new work, Core resolves the Template, saved Agent and model configuration, binds Vault Credentials, then validates the selected Harness and execution configuration before writing. A failed dependency lookup rechecks the retry identity so an already committed creation remains recoverable. + +### Creation retries + +Send an `Idempotency-Key` of 1–128 bytes that is not only whitespace; a longer or whitespace-only key returns 400 `invalid_request`. An empty header counts as no key. Without a key, every request creates a new Session. The official service creates a new Session for each request even with the same key; Core returns the original one. + +| Case | Response | +| --- | --- | +| Same key, same request, same Project | 201 with the Session's current state. No input is admitted again. A `stream=true` retry returns 201 with no events and closes. | +| Same key, different request | 409 `idempotency_conflict` | +| Same key after the Session was deleted | 409 `idempotency_conflict` | + +Keys are scoped to the Project; any key of the Project, including one issued after a rotation, can retry. A request that has an inline Agent without `model` or names a saved Agent, a template, initial files or preparation, `vault_ids` or credential references, `x_agents_core`, or an `openai_hosted` environment is compared as sent, before any of those sources is read: a matching retry returns the original Session even after the Agent, template, Credential or deployment default changes or is deleted. Other requests are compared by their resolved configuration. Model provider keys enter the comparison only as fingerprints. + +### Update and list + +`POST /agents/sessions/{session_id}` accepts only `metadata`, which is required: null or `{}` clears it and an object replaces all pairs. Execution state and the creation retry identity are unchanged. + +`GET /agents/sessions` accepts `agent_id`, which matches the Session's immutable root Agent ID, including inline Agent IDs and Agents that were since updated or deleted. The filter applies before pagination; an empty `agent_id` is a filter, not an omission. + +### Delete + +`DELETE /agents/sessions/{session_id}` deletes a Session that is idle or failed, has no queued, running or waiting root Turn and no pending input reservation. + +| Case | Response | +| --- | --- | +| Deletable | 200 `{id, object: "agent.session.deleted", deleted: true}`. Reads, updates, input, Turns and Items of the Session then return 404, and its open event streams end. Core releases the Session's Core-managed sandbox; a `self_hosted` machine and its files are left alone. | +| A root Turn is queued, in progress or waiting for required actions, or input is waiting for admission, a self-hosted connection or hosted provisioning | 409, type and code `conflict_error`, null param, "session must be durably idle or failed without required actions before deletion". Nothing changes. | +| The caller's own Session, already deleted | 200 with the same confirmation | +| Missing, malformed or foreign | 404 | + +Subagent child Turns and pending Environment file writes do not block deletion. To delete running work, send `agent.session.input.cancel`, wait until the Session is idle, then delete. Input waiting for its Environment cannot be cancelled; the Session becomes deletable when the input starts, its five-minute deadline passes or the Environment fails. Core admits a Turn in the same transaction that returns 202 for its input, so a deletion right after that 202 returns 409. + +### Response fields + +A Session's `agent.tools` omits `tool_search` declarations, which the pinned Session tool union does not include; the frozen configuration keeps them. On `self_hosted` Sessions, `environment.remote_url` is Core's daemon WebSocket URL, `/api/v1/agent-daemon/ws` under the public URL, which only OpenAgentCore's Runtime daemon speaks; requests cannot set it. Other Session fields follow the pinned types; `x_agents_core` is described in the [Agents API guide](../../docs/api/public-agent-api.md#core-extensions-x_agents_core). diff --git a/contracts/agents-api/write-audit.md b/contracts/agents-api/write-audit.md deleted file mode 100644 index 5f72cc843..000000000 --- a/contracts/agents-api/write-audit.md +++ /dev/null @@ -1,143 +0,0 @@ -# API-key write provenance - -This Core extension does not change the pinned public `/v1` protocol. It serves -administrator consoles; it is not an HTTP access log, execution event stream or -replacement for a Session's creator identity. - -## Authentication and ownership - -Authentication carries the issued key UUID and its Project's tenant/principal. -Keys in one Project share assets and permissions; provenance records which key -performed each write. All business keys live in PostgreSQL. Configuration contains -no business keys or Projects. The Core key authenticates only `/core/v1` and is -never a public API identity. Client headers cannot assert a -key identity, and display names do not change an ID. - -True creation stores ownership in the same transaction as the resource and operation. -Updates, retries and successful no-ops never replace it. Initial Skill versions and -new Session Environments share their parent creation operation's provenance. Runtime -produced Artifacts have no API-key creation anchor; deleting one is still recorded. -Missing historical or foreign ownership is `null`. No historical resources are backfilled. - -An issued key's current `revoked_at` comes from its retained key row. Revocation -prevents new authentication but does not invalidate already admitted work or erase -history. Existing audit metadata remains an immutable snapshot, including any -historical static-key and console-key records; retaining history does not enable static-key -authentication. Deleting a resource does not delete its -operation history or creation anchor. - -## Transaction and success boundary - -A successful write means a durable business commit, not successful HTTP delivery or -successful subsequent model execution. Each request has a server-generated -`request_id` for deduplication and the request's diagnostic `trace_id`. A shared -trace across requests is not an idempotency key. Business rollback also rolls back -the operation and ownership. Failure to persist the audit fails the transaction. -Reads, failed validation/authorization and internal maintenance writes are not logged. - -Session input is recorded when durably admitted, including prepared-Environment -reservations. Later execution failure or response disconnection does not undo that -accepted write. Explicit successful empty-input and creation/deletion retries get -one operation without changing ownership. Automatic promotion/refresh/cleanup does -not create another public operation. - -Environment file bytes are written by the existing Runtime, outside PostgreSQL. -Preparation durably stores only the safe key/request/trace origin. The confirmed -success receipt and audit commit together; uncertain/failed uploads are not reported -as successful operations. Recovery uses the saved origin. This is not a claim of -atomic transactions across the filesystem and PostgreSQL. - -Recorded fields are ID, timestamp, safe key metadata, action, resource type/ID, -parent ID, request ID and trace ID. Never store request/response bodies, bearer -secrets, model credentials, tokens, file paths or file contents in these tables. - -| Public write | Action | Resource / parent | -| --- | --- | --- | -| Agent create/update/delete | create/update/delete | agent | -| Session create/update/delete | create/update/delete | session | -| Session events | send_events | session | -| Session Artifact delete | delete | artifact / session | -| Environment file upload | upload_file | environment / session | -| Environment Template create/update/delete | create/update/delete | environment_template | -| Skill create/default/delete | create/update_default_version/delete | skill | -| Skill version upload/delete | upload_version/delete | skill_version / skill | -| Source File upload/delete | create/delete | file | -| Vault create/delete | create/delete | vault | -| Credential create/replace/delete | create/update/delete | credential / vault | - -Explicit public OAuth Credential writes are covered. Automatic OAuth refresh is not -a separate public write. Public resource IDs retain their existing formats. - -Public resource writes carry authenticated key provenance separately from the -execution principal. Persist their operation record and genuine creation ownership -in the same business transaction; no best-effort response middleware or async audit -queue. A failed audit must roll back the write. Internal lifecycle/refresh work does -not acquire public provenance. Retries never replace ownership. Environment uploads -persist safe request origin before dispatch and record success with the confirmed -receipt, not the native filesystem call. Never put payloads, paths or secrets in -audit metadata. Read models are Core-key `/core/v1` routes; keep `/v1` wire -contracts unchanged. See -[write-audit.md](write-audit.md) for coverage, retention and -console integration. Do not confuse key identity with Session creator identity. - -## Console queries - -Both endpoints require the Core key under `/core/v1/projects/{project_id}`. The path identifies the Project, including an -archived Project; it does not authenticate. API keys cannot call these -routes. - -### Batch ownership - -`GET /core/v1/projects/{project_id}/resource-owners?resource_type=agent&resource_ids=id1,id2` - -`resource_type` is one of `agent`, `session`, `environment`, -`environment_template`, `skill`, `skill_version`, `file`, `vault`, `credential`, -`artifact`. Pass 1–100 comma-separated public IDs. Results preserve input order: - -```json -{"data":[ - {"resource_id":"id1","api_key":{"id":"key-uuid","name":"SDK","prefix":"pc_example","kind":"issued","revoked_at":null},"source":"api_key","admin_audit_id":null}, - {"resource_id":"id2","api_key":null,"source":null,"admin_audit_id":null} -]} -``` - -Resources created by the removed administrator copy operation keep -`api_key:null`, `source:"admin_copy"` and a non-null `admin_audit_id`. They never -had a public API-key write. See [historical copy provenance](admin-api.md#historical-copy-provenance). - -### Operations - -`GET /core/v1/projects/{project_id}/write-operations?limit=50` - -Optional filters: `key_id`, `resource_type`, `resource_id`, `created_after` -(inclusive RFC3339 timestamp), `created_before` (exclusive RFC3339 timestamp). -`limit` is 1–100, default 50. Pass the previous `next_cursor` as `after`, keeping -filters unchanged. Results sort by descending `(created_at, id)` with keyset -pagination; a cursor is not an authorization token or a snapshot of future writes. -The response is `{ "data": [...], "has_more": false, "next_cursor": "" }`. -Each operation contains `id`, `created_at`, `api_key`, `action`, `resource_type`, -`resource_id`, `parent_id`, `request_id`, `trace_id`. An absent parent is an empty -string. Empty pages contain `data: []`. - -Malformed/duplicate/unknown query parameters return 400, missing Project -404, invalid deployment authentication 401. These rules belong only to the new -Core routes and do not alter public list parsing or errors. Creation records remain -queryable after resource deletion; expired non-creation records do not. - -## Retention - -`OAC_WRITE_AUDIT_RETENTION` accepts a Go duration of at least one hour; -default `2160h` (90 days). Every minute Core removes at most 1,000 expired -non-creation records in a bounded transaction. Creation records and anchors are -retained permanently, which includes the entire lifetime of a resource and its -history after deletion. Key revocation/resource deletion never triggers cleanup. -Retention removes only aged non-creation operations; it does not erase key metadata, -resource rows or ownership. A retention backlog can take multiple passes to drain. - -## Validation - -The batch's Store tests cover each mutation, induced audit-write failure with -business rollback, tenant isolation, ownership immutability, revocation, cursor -filters and retention. HTTP tests cover authenticated source propagation and -administrator-only query validation. Live evidence and remaining limitations are -recorded separately; passing these tests does not claim full Agents API compatibility. diff --git a/docs/api/README.md b/docs/api/README.md index 64f677f29..7bcef1d52 100644 --- a/docs/api/README.md +++ b/docs/api/README.md @@ -1,136 +1,27 @@ -# API documentation +# API namespaces and credentials -Core serves three namespaces. Each has one kind of caller and its own credential; -no credential works in another namespace. Versioned native installer downloads are -public release content. +Core serves three namespaces. Each has one kind of caller and its own credential, and a credential works only in its own namespace. -| Namespace | Caller | Credential | Contents | Reference | +| Namespace | Caller | Credential | Contents | Owner | | --- | --- | --- | --- | --- | -| `/v1` | Applications (business systems, SDKs) | Project API key | Exactly the pinned official Agents API routes. Core-only fields live only in `x_agents_core` (`harness`, `model_provider`, `harness_config`, Session `installation`) | [Agents API guide](public-agent-api.md) | -| `/core/v1` | Core Web's server and operator scripts | [Core key](../getting-started/operations.md#core-key) | Installation facts, Projects and keys, resource reads and deletion, Session archive, credential issuance, metrics, audit, sandbox deployment and nodes, deployment model providers | [Core API](#core-api), [Web API](web-management.md), [Core OpenAPI](../../contracts/agents-api/core.openapi.yaml) | -| `/api/v1` | Nodes, Runtime daemons, self-hosted executors | Machine credentials: short-lived Session installation grants, node enrollment tokens and executor credentials issued through `/core/v1` or claimed by installation, node credentials registered with an enrollment token, and daemon credentials Core issues for hosted sandboxes | Machine bootstrap and connections: `/api/v1/sandbox-node/*` and `/api/v1/agent-daemon/*`, including WebSockets; each credential works only on its own routes | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes), [executor credentials](../../contracts/agents-api/environment-executor-credentials.md), [machine OpenAPI](../../contracts/agents-api/runtime.openapi.yaml) | +| `/v1` | Applications: business systems and the official OpenAI SDK | Project API key | Exactly the 58 method and path pairs of the pinned official Agents API, listed in [upstream-routes.json](../../contracts/agents-api/upstream-routes.json). Core-only fields sit inside `x_agents_core`: `harness`, `model_provider`, `harness_config`, `environment`, and the read-only Session `installation` | [Agents API guide](public-agent-api.md) | +| `/core/v1` | Web's console server and operator scripts | [Core key](../getting-started/operations.md#core-key) | Installation facts, Projects and keys, resource reads and deletion, Session archive, executor credentials, default models, metrics, audit, sandbox deployment and nodes | [Core administration API](../../contracts/agents-api/admin-api.md) | +| `/api/v1` | Nodes, Runtime daemons, self-hosted executors and their installers | Machine credentials: node enrollment tokens and node credentials, installation grants, executor credentials, and daemon credentials. Each works only on its own routes | Machine bootstrap and connections under `/api/v1/sandbox-node/*` and `/api/v1/agent-daemon/*`, including WebSockets, and the public native installer downloads | [Machine connection API](#machine-connection-api) | -A credential used in another namespace gets 401: a Project API key on `/core/v1` or -`/api/v1`, the Core key on `/v1` or `/api/v1`. How Projects and keys behave is in -[Projects own assets](../design-principles.md#projects-own-assets). +A credential used in another namespace gets 401: a Project API key on `/core/v1` or `/api/v1`, the Core key on `/v1` or `/api/v1`. How Projects and keys behave is in [Projects own assets](../design-principles.md#projects-own-assets). -**Routing.** The reverse proxy sends `/v1` and `/api/v1` to Core and everything else -to Web ([proxy setup](../getting-started/install-options.md#https-and-the-reverse-proxy)). -Browsers reach `/core/v1` only through Web's server, which adds the Core key after -sign-in; Web returns 404 for `/v1` and `/api/v1`. Operator scripts call `/core/v1` -on Core's loopback port. Details: [Web and Core](web-management.md). - -## Public API - -Applications call `/v1` with a Project API key. The routes are exactly the 58 pairs -in [upstream-routes.json](../../contracts/agents-api/upstream-routes.json). The -[Agents API guide](public-agent-api.md) explains every resource with SDK and HTTP -examples. - -## Core API - -All routes are under `/core/v1`. Core Web's server and operator scripts call them -with the Core key. Most administrators use Web instead; call the API directly for -automation. - -### Common tasks - -Run these on the Core host, against Core's loopback port. The helper reads the key -from its file, keeping it off the command line: - -```sh -core() { # core METHOD PATH [JSON body] - curl -fsS -X "$1" "http://127.0.0.1:8091/core/v1$2" \ - -H @<(printf 'Authorization: Bearer %s\n' "$(cat ~/.oac/core/secrets/core.key)") \ - -H 'Content-Type: application/json' ${3:+-d "$3"} -} -``` - -| Task | Command | -| --- | --- | -| List Projects | `core GET /projects` | -| Create a Project | `core POST /projects '{"name": "billing-bot"}'` | -| Issue an API key (shown once, as `key`) | `core POST /projects/$PROJECT_ID/keys '{"name": "prod"}'` | -| Revoke a key | `core DELETE /projects/$PROJECT_ID/keys/$KEY_ID` | -| Archive a Project (revokes all keys) | `core POST /projects/$PROJECT_ID/archive` | -| See harnesses and their default models | `core GET /harnesses` | -| Set Codex's default model | `core PUT /harnesses/codex/model-configuration '{"model": "your-model-id", "model_provider": {"protocol": "responses", "base_url": "https://provider.example/v1", "api_key": "sk-..."}}'` | -| Issue an executor credential | See [executor credentials](../../contracts/agents-api/environment-executor-credentials.md#core-key-routes) | -| Installation facts, including the API base URL | `core GET /installation` | - -Errors use the [Core error envelope](../../contracts/agents-api/core-errors.md). - -### All routes - -| Routes | Contents | Contract | -| --- | --- | --- | -| `projects`, `projects/{project_id}[/archive]`, `projects/{project_id}/keys[/{key_id}]` | Projects and their API keys | [Administrator contract](../../contracts/agents-api/admin-api.md) | -| `projects/{project_id}/{agents,environment-templates,skills,files,vaults,sessions}/**` | Resource reads and deletion, Session history, artifacts and archive | [Administrator contract](../../contracts/agents-api/admin-api.md) | -| `projects/{project_id}/sessions/{session_id}/execution-configuration` | Committed harness and model selection | [Execution configuration](../../contracts/agents-api/execution-configuration.md) | -| `projects/{project_id}/sessions/{session_id}/diagnostics`, `projects/{project_id}/sessions/{session_id}/turns/{turn_id}/diagnostics` | Root failure categories and Item receipt timing | [Session diagnostics](../../contracts/agents-api/session-diagnostics.md) | -| `projects/{project_id}/sessions/{session_id}/runtime-observation`, `sandbox/runtime-observations` | Current Runtime observations | [Runtime observations](../../contracts/agents-api/runtime-observability-api.md) | -| `projects/{project_id}/sessions/{session_id}/runtime-history` | Stored Runtime history | [Runtime history](../../contracts/agents-api/runtime-history-api.md) | -| `projects/{project_id}/{resource-owners,write-operations}`, `audit-log`, `summary` | Provenance, write history, administrator audit and usage summary | [Write audit](../../contracts/agents-api/write-audit.md), [administrator contract](../../contracts/agents-api/admin-api.md) | -| `projects/{project_id}/environments/{environment_id}/executor-credentials[/{key_id}]` | Executor credentials for a self-hosted Environment | [Executor credentials](../../contracts/agents-api/environment-executor-credentials.md) | -| `installation` | Public URL, API base URL, source commit, the installer's process settings and what is bound to the public URL; available before any deployment | [Installation](../../contracts/agents-api/installation.md) | -| `metrics` | Core's own process metrics | [Core metrics](../../contracts/agents-api/core-metrics.md) | -| `sandbox/deployment[/reset]`, `sandbox/providers/{provider}/discovery`, `sandbox/enrollment-tokens`, `sandbox/nodes[/{node_id}[/allocations]]` | Sandbox deployment, read-only provider configuration discovery, node enrollment tokens and nodes | [Sandbox deployment](../../contracts/agents-api/sandbox-deployment.md), [nodes guide](../getting-started/nodes.md), [node host history](../../contracts/agents-api/node-host-history.md) | -| `harnesses`, `harnesses/{harness}/model-configuration` | Supported harnesses and each harness's deployment default model configuration (write-only provider key) | [Model execution](../../contracts/agents-api/model-execution.md#deployment-defaults) | +**Routing.** The reverse proxy sends `/v1` and `/api/v1` to Core and everything else to Web ([proxy setup](../getting-started/install-options.md#https-and-the-reverse-proxy)). Browsers reach `/core/v1` only through Web's console server, which adds the Core key after sign-in and answers 404 for `/v1` and `/api/v1` ([console server](../web/console-server.md)). Operator scripts call `/core/v1` on Core's loopback port ([script the Core API](../getting-started/operations.md#script-the-core-api)). ## Machine connection API -These routes are under `/api/v1`. Each accepts only the credential listed, never -the Core key or a Project API key. +These routes are under `/api/v1`. Each accepts only the credential listed, never the Core key or a Project API key. The generated [machine OpenAPI](../../contracts/agents-api/runtime.openapi.yaml) covers the annotated node HTTP and installation grant routes; the node WebSocket `connect` route is described in the node generation protocol. | Routes | Caller | Credential | Contract | | --- | --- | --- | --- | -| `POST sandbox-node/enroll` | Node installer | One-use enrollment token from `POST /core/v1/sandbox/enrollment-tokens` | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes), [machine OpenAPI](../../contracts/agents-api/runtime.openapi.yaml) | -| `GET sandbox-node/configuration` | Node installer and node | Enrollment token, or node credential with `X-OAC-Node-ID` | [Sandbox deployment](../../contracts/agents-api/sandbox-deployment.md), [machine OpenAPI](../../contracts/agents-api/runtime.openapi.yaml) | -| `GET sandbox-node/identity`, WebSocket `GET sandbox-node/connect` | Node | Node credential registered at enrollment | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes) | -| `POST agent-daemon/enroll`, `GET agent-daemon/connection` | Self-hosted executor and its installer | Executor credential from `/core/v1/projects/{project_id}/environments/{environment_id}/executor-credentials` | [Executor credentials](../../contracts/agents-api/environment-executor-credentials.md) | -| WebSocket `GET agent-daemon/ws`, `POST agent-daemon/bootstrap`, `GET agent-daemon/device-status` | Runtime daemons | Daemon credential: Core writes one into each hosted sandbox it prepares; a self-hosted executor uses its executor credential | [Runtime enrollment](../../services/core/README.md#user-managed-runtime-enrollment) | - -## Contract sources - -- [Pinned upstream baseline](../../contracts/agents-api/upstream.json): OpenAI - Python SDK 3.13.0, exact upstream commit and `agents=v1`. -- [Pinned routes](../../contracts/agents-api/upstream-routes.json) and - [fields](../../contracts/agents-api/upstream-fields.json): extracted from the - pinned SDK by `scripts/extract-agents-api-upstream.py`. Contract tests require the - public OpenAPI to have exactly these routes and to keep every other field inside - `x_agents_core`. -- [Public OpenAPI](../../contracts/agents-api/openapi.yaml): public schema snapshot; - combine it with the fixed SDK and [operation evidence](../../contracts/agents-api/operation-evidence.md). -- [Core OpenAPI](../../contracts/agents-api/core.openapi.yaml): generated `/core/v1` - routes, all authenticated by the Core key. -- [Machine connection OpenAPI](../../contracts/agents-api/runtime.openapi.yaml): - generated `/api/v1` sandbox node routes. The node WebSocket is described in the [generation protocol](../../contracts/agents-api/node-generation-protocol.md); the private - daemon transport is described in the Runtime credential guide. -- [Administrator contract](../../contracts/agents-api/admin-api.md): Project/key - lifecycle, resources, explicit hosted Session archive, summary, - errors/deletion preconditions, audit and historical copy provenance. -- [Web integration](web-management.md): browser, console and Core boundaries and frontend handoff. -- [Design rules](../design-principles.md) and [contributor guide](../../CONTRIBUTING.md): - ownership, security and change requirements. - -Generated schemas do not establish complete compatibility or real execution -support. The [coverage record](../../contracts/agents-api/README.md) identifies -qualified workflows, native differences and unresolved behavior. Update the -relevant contract and this index when adding or moving an API surface. - -## Self-hosted installation - -Creating or reading a `self_hosted` Session returns short-lived install commands in -`x_agents_core.installation`. Core Web reads the same commands at -`GET /core/v1/projects/{project_id}/environments/{environment_id}/installation`. -Machine installers use `POST /api/v1/agent-daemon/installation` and its `/claim` -subroute with the installation Bearer authorization. Qualified artifacts under -`/api/v1/agent-daemon/install/{version}/` are public, immutable release content. The -[installation grant](../../contracts/agents-api/environment-executor-credentials.md#installation-grant) -owns expiry, retry and credential ownership; the -[self-hosted guide](../getting-started/self-hosted.md#platforms) lists platforms. - -The console-local `GET`/`POST /console/installation/domain` surface uses the signed-in -browser session and same-origin checks. It delegates only domain setup to the -installer, with the server-held Core key over a private Unix socket; it is not part -of the Agents API or Core management API. See [console domain setup](../web/console-server.md#domain-setup). +| `POST sandbox-node/enroll` | Node installer | One-use enrollment token from `POST /core/v1/sandbox/enrollment-tokens` | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes) | +| `GET sandbox-node/configuration` | Node installer and node | Enrollment token, or node credential with `X-OAC-Node-ID` | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes) | +| `GET sandbox-node/identity`, WebSocket `GET sandbox-node/connect` | Node | Node credential registered at enrollment | [Node routes](../../contracts/agents-api/sandbox-deployment.md#authority-and-routes), [node generation protocol](../../contracts/agents-api/node-generation-protocol.md) | +| `GET agent-daemon/install/{version}/*` | Self-hosted installer | None: public, immutable release content | [Installation grant](../../contracts/agents-api/environment-executor-credentials.md#installation-grant) | +| `POST agent-daemon/installation`, `POST agent-daemon/installation/claim` | Self-hosted installer | Installation grant: the short-lived authorization in a `self_hosted` Session's `x_agents_core.installation` commands | [Installation grant](../../contracts/agents-api/environment-executor-credentials.md#installation-grant) | +| `POST agent-daemon/enroll`, `GET agent-daemon/connection` | Self-hosted executor and its installer | Executor credential | [Executor credentials](../../contracts/agents-api/environment-executor-credentials.md) | +| WebSocket `GET agent-daemon/ws`, `POST agent-daemon/bootstrap`, `GET agent-daemon/device-status` | Runtime daemons | Daemon credential: Core issues one to each hosted sandbox through the [bootstrap file](../runtime-bootstrap.md); a self-hosted executor uses its executor credential; a device for `none` Sessions uses the credential an operator provisions with `oac-core-device` | [Core–Runtime protocol](../runtime-protocol.md), [device provisioning](../../services/core/README.md#internal-execution-device-connection) | diff --git a/docs/api/public-agent-api.md b/docs/api/public-agent-api.md index e94221904..3e31a6067 100644 --- a/docs/api/public-agent-api.md +++ b/docs/api/public-agent-api.md @@ -1,15 +1,12 @@ # Agents API guide -Core serves the [OpenAI Agents API](https://platform.openai.com/docs/api-reference) at -`/v1`. Use the official OpenAI SDK or plain HTTP. This guide shows both for every -common operation, and notes where Core differs from OpenAI. +Core serves the [OpenAI Agents API](https://platform.openai.com/docs/api-reference) at `/v1`. Use the official OpenAI SDK or plain HTTP. This guide shows both for every common operation, and notes where Core differs from OpenAI. New to the API? Run the [quickstart](../getting-started/quickstart.md) first. ## Before you start -**Base URL and key.** Your administrator gives you the API base URL, such as -`https://core.example/v1`, and a Project API key. +**Base URL and key.** Your administrator gives you the API base URL, such as `https://core.example/v1`, and a Project API key. ```sh export OPENAI_BASE_URL=https://core.example/v1 @@ -28,9 +25,7 @@ from openai import OpenAI client = OpenAI() # reads OPENAI_BASE_URL and OPENAI_API_KEY ``` -**HTTP.** Every request needs a Bearer key. Routes under `/agents` and `/vaults` also -need `OpenAI-Beta: agents=v1`; `/files` and `/skills` do not. The SDK sets both. The -HTTP examples below use this shell helper: +**HTTP.** Every request needs a Bearer key. Routes under `/agents` and `/vaults` also need `OpenAI-Beta: agents=v1`; `/files` and `/skills` do not. The SDK sets both. The HTTP examples below use this shell helper: ```sh oac() { # oac PATH [curl options]: call /agents or /vaults with the required headers @@ -42,8 +37,7 @@ oac() { # oac PATH [curl options]: call /agents or /vaults with the required he oac /agents ``` -On a shared host, `-H @<(printf 'Authorization: Bearer %s\n' "$OPENAI_API_KEY")` -keeps the key out of the process list. +On a shared host, `-H @<(printf 'Authorization: Bearer %s\n' "$OPENAI_API_KEY")` keeps the key out of the process list. ## Resources at a glance @@ -60,9 +54,22 @@ keeps the key out of the process list. | [Vaults](#vaults) | `/vaults` | Write-only credentials for MCP servers | | Subagents | `/agents/sessions/{id}/subagents` | Read-only child work; see [subagents](../../contracts/agents-api/subagents.md) | -Core has exactly the routes of the pinned SDK, listed in -[upstream-routes.json](../../contracts/agents-api/upstream-routes.json). It adds no -route; its additions live in [`x_agents_core`](#core-extensions-x_agents_core). +Core has exactly the routes of the pinned SDK, listed in [upstream-routes.json](../../contracts/agents-api/upstream-routes.json). It adds no route; its additions live in [`x_agents_core`](#core-extensions-x_agents_core). + +## Common tasks + +| To | Use | +| --- | --- | +| Continue a conversation, or steer a running Turn | [Send a message](#send-a-message) to the same Session. For a different configuration or workspace, create a new Session | +| Watch output live | [Stream events](#stream-events) | +| Stop the current Turn, or recover after a lost response | [Cancel](#cancel), [idempotency](#idempotency) | +| Call your own code from the agent | [Function tools](#function-tools) | +| Give the agent Skills, packages, files and setup commands | [Skills](#skills), [Environment Templates](#environment-templates) | +| Connect an MCP server with credentials | [Vaults](#vaults) and [execution tools](../../contracts/agents-api/execution-tools.md) | +| Use Skill or Plugin directories on your own machine | [Local capability directories](../getting-started/self-hosted.md#local-capability-directories) | +| Put files in the workspace, or download what the agent wrote | [Files](#files) | +| Read the conversation and tool results | [Turns and Items](#turns-and-items) | +| Find out why a Session or Turn failed | [Diagnose a failure](#diagnose-a-failure) | ## Conventions @@ -91,14 +98,11 @@ for session in client.beta.agents.sessions.list(limit=100): oac "/agents/sessions?limit=100&after=$LAST_ID" ``` -Two lists differ: `GET /files` returns up to 10,000 files at once, and workspace -files use an opaque `page` token (see [Workspace files](#workspace-files)). +Two lists differ: `GET /files` returns up to 10,000 files at once, and workspace files use an opaque `page` token (see [Workspace files](#workspace-files)). ### Idempotency -Send an `Idempotency-Key` (up to 128 bytes) when creating a Session or sending input. -A retry with the same key and body returns the original result instead of doing the -work twice. The same key with a different body fails with 409 `idempotency_conflict`. +Send an `Idempotency-Key` (up to 128 bytes) when creating a Session or sending input. A retry with the same key and body returns the original result instead of doing the work twice. The same key with a different body fails with 409 `idempotency_conflict`. ```python import uuid @@ -108,12 +112,9 @@ client.beta.agents.sessions.events.create(session_id, events=[...], idempotency_ client.beta.agents.sessions.create(environment=..., extra_headers={"Idempotency-Key": key}) ``` -**After a lost response,** retry with the same key, then read the Session, Turns and -Items. Never resend without the key. Core keeps creation keys even after the Session -is deleted. +**After a lost response,** retry with the same key, then read the Session, Turns and Items. Never resend without the key. Core keeps creation keys even after the Session is deleted. -> Core difference: OpenAI creates a new Session for each create request even with the -> same key. Core returns the original. +See [creation retries](../../contracts/agents-api/wire-semantics.md#creation-retries) for comparison rules and the difference from OpenAI. ### Errors @@ -125,36 +126,41 @@ is deleted. | --- | --- | --- | | 400 | `invalid_request_error`, `invalid_beta`, `model_provider_required`, `unsupported_or_invalid_configuration` | Fix the request. `invalid_beta` means a missing or wrong `OpenAI-Beta` header | | 401 | `invalid_api_key` or none | Wrong or missing key, or a key from another namespace | -| 404 | | Missing, or belongs to another Project. The two look the same | +| 404 | `not_found_error` on Beta routes | Missing, or belongs to another Project. The two look the same | | 405 | `unsupported_operation` | Core doesn't support this operation | | 409 | `conflict_error`, `idempotency_conflict` | State conflict, for example deleting a busy Session | -| 413 | | Body too large | -| 503 | | Temporarily uncertain; read state before retrying | +| 413 | `request_too_large` | Body too large | +| 503 | `authentication_unavailable`, `execution_unavailable` | Temporarily uncertain; read state before retrying | Every response carries `X-Request-Id`; include it when reporting a problem. ### Core extensions: `x_agents_core` -Core runs several harnesses and accepts your own model access. Those settings are the -only additions to the OpenAI shapes, and they sit inside `x_agents_core`: +Core runs several harnesses and accepts your own model access. Those settings are the only additions to the OpenAI shapes, and they sit inside `x_agents_core`: | Field | Where | Value | | --- | --- | --- | | `harness` | Agent, or a Session's inline `agent` | `codex`, `claude_sdk` or `mcode` | -| `model_provider` | Agent, or Session creation (top level) | `protocol`, `base_url`, `api_key`, and for `mcode` also `context_window` and `max_output_tokens`. The protocol must be one the harness supports: Codex `responses`, Claude Code `anthropic`, MiniMax Code any of `anthropic`, `responses`, `chat_completions` | +| `model_provider` | Agent, or Session creation (top level) | `protocol`, `base_url`, `api_key`, and for `mcode` also `context_window` and `max_output_tokens`. The protocol must be one of the harness's native protocols | | `harness_config` | Agent, inline `agent`, or Session creation (top level, wins) | The harness's native model parameters, such as Codex's `model_reasoning_effort` | -| `environment` | Session creation (top level), any placement | Portable preparation: `environment_template_id`, `files`, `env`, `packages`, `setup_commands`, `skills`, `plugins`, `capability_directories`. A field may not also appear in `environment`; see [Environments](../../contracts/agents-api/environments.md#preparation-order) | +| `environment` | `openai_hosted` or `self_hosted` Session creation (top level) | Portable preparation: `environment_template_id`, `files`, `env`, `packages`, `setup_commands`, `skills`, `plugins`, `capability_directories`. A field may not also appear in `environment`; see [Environments](../../contracts/agents-api/environments.md#preparation-order) | | `installation` | Read-only, on `self_hosted` Sessions | Short-lived install commands for your machine; see [self-hosted execution](../getting-started/self-hosted.md) | -Any other member is rejected with 400. `api_key` is write-only: reads return -`api_key_configured`. `harness_config` replaces the whole object; `{}` clears it. With the SDK, pass these through `extra_body`. How to pick a -harness and model, and which provider a Session uses, is in the -[user guide](../user-guide.md#choose-a-harness-and-a-model). +Any other member is rejected with 400. `api_key` is write-only: reads return `api_key_configured`. `harness_config` replaces the whole object; `{}` clears it. With the SDK, pass these through `extra_body`. + +## Choose a harness and a model + +The harness is the agent program that runs a Session: Codex (`codex`), Claude Code (`claude_sdk`) or MiniMax Code (`mcode`). Set `x_agents_core.harness` on the Agent or the inline `agent`; without it, the installation's default harness applies ([`core.default_harness`](../configuration.md#settings), Codex unless the operator changed it). + +- **Model.** `model` is the provider's exact model ID. An inline Agent on an `openai_hosted` or `none` Session may omit it to use the default model configuration of its harness. A saved Agent always needs one. +- **Provider.** The harness calls your provider directly, with one of the harness's native protocols; there is no conversion, and a mismatch is rejected when the Session is created. [Model execution](../../contracts/agents-api/model-execution.md#saved-defaults-and-precedence) lists each harness's protocols and which provider a Session uses on each Environment type. A Session freezes its provider at creation. +- **Native parameters.** `harness_config` carries the harness's own model settings; see [native model parameters](../../contracts/agents-api/model-execution.md#native-model-parameters). + +Not every combination of harness, placement and operation is supported; the [Harness capabilities](../../contracts/agents-api/harness-capabilities.md) lists them. ## Agents -An Agent is saved configuration. Sessions copy it when they start, so editing an -Agent affects only new Sessions. +An Agent is saved configuration. Sessions copy it when they start, so editing an Agent affects only new Sessions. ```python agent = client.beta.agents.create( @@ -175,7 +181,7 @@ oac "/agents" -d '{ ``` ```json -{"id": "agent_...", "object": "agent", "model": "your-model-id", "name": "Reviewer", +{"id": "00000000-0000-4000-8000-000000000001", "object": "agent", "model": "your-model-id", "name": "Reviewer", "instructions": "...", "tools": [], "metadata": {}, "created_at": 1790000000, "x_agents_core": {"harness": "codex"}} ``` @@ -190,12 +196,9 @@ oac "/agents" -d '{ (`agents` is `client.beta.agents` throughout.) -- **Update** changes only the fields you send. `metadata` replaces all pairs; `null` - clears `name` or `instructions`. +- **Update** changes only the fields you send. `metadata` replaces all pairs; `null` clears `name` or `instructions`. - **Delete** keeps existing Sessions. -- **Tools** are functions, MCP servers, `tool_search` (Claude) and `web_search` with - `mode: "disabled"`. Support depends on harness and Environment; see - [execution tools](../../contracts/agents-api/execution-tools.md). +- **Tools** are functions, MCP servers, `tool_search` (Claude) and `web_search` with `mode: "disabled"`. Support depends on harness and Environment; see [execution tools](../../contracts/agents-api/execution-tools.md). ## Sessions @@ -215,7 +218,7 @@ session = client.beta.agents.sessions.create( ```sh oac "/agents/sessions" -H "Idempotency-Key: $(uuidgen)" -d '{ "environment": {"type": "openai_hosted"}, - "agent_id": "agent_...", + "agent_id": "00000000-0000-4000-8000-000000000001", "input": "Review the files in /workspace and summarize the risks.", "metadata": {"ticket": "T-123"} }' @@ -224,7 +227,7 @@ oac "/agents/sessions" -H "Idempotency-Key: $(uuidgen)" -d '{ Returns 201 with the Session: ```json -{"id": "sess_...", "object": "agent.session", "status": "in_progress", +{"id": "00000000-0000-4000-8000-000000000002", "object": "agent.session", "status": "idle", "agent": {"model": "...", ...}, "environment": {"id": "env_...", "type": "openai_hosted", ...}, "metadata": {"ticket": "T-123"}, "required_actions": [], "vault_ids": [], "created_at": 1790000000, "last_active_at": 1790000000} @@ -234,26 +237,23 @@ Returns 201 with the Session: | --- | --- | | `environment` | Required. Where the agent works; see the table below | | `agent_id` or `agent` | A saved Agent, or an inline Agent object (same fields as create). An inline Agent on `openai_hosted` or `none` may omit `model` to use the installation default | -| `input` | The first message: a string or a message array. Optional for `openai_hosted` and `self_hosted` | +| `input` | The first message: a string or a message array. Required on `none`, and with `stream: true` except on `self_hosted` ([initial input](../../contracts/agents-api/sessions-events.md#initial-input-at-session-creation)) | | `metadata` | Your own string key-value pairs | | `vault_ids` | [Vaults](#vaults) whose credentials MCP servers may use | | `stream` | `true` returns [server-sent events](#stream-events) instead of JSON | -| `x_agents_core.model_provider` | This Session's model access, if not from the Agent or the default | +| `x_agents_core.model_provider` | This Session's model access, if not from the Agent or the default. Rejected on `none` | | `environment.type` | Runs on | Notes | | --- | --- | --- | -| `openai_hosted` | A sandbox Core creates (node or E2B) | Optional `network`, `packages`, `files`, `skills`, `plugins`, `setup_commands`, or a template | -| `self_hosted` | Your machine | Requires an absolute `workspace_directory`. Skills, packages, files or a template go in `x_agents_core.environment`. The response carries install commands in `x_agents_core.installation`; see [self-hosted execution](../getting-started/self-hosted.md) | -| `none` | An existing device connection | `input` required; uses the default model only | +| `openai_hosted` | A sandbox Core creates on a node or E2B; the administrator provides the capacity | Optional `network`, `packages`, `files`, `skills`, `plugins`, `env`, `capability_directories`, `setup_commands`, or a template | +| `self_hosted` | Your own Linux, macOS or Windows machine | Requires an absolute `workspace_directory`. Skills, packages, files or a template go in `x_agents_core.environment`. The response carries install commands in `x_agents_core.installation`; see [self-hosted execution](../getting-started/self-hosted.md). The Session brings its own `model_provider` | +| `none` | A device connection an operator registered, with no workspace | `input` required. The model comes from the installation default, or from the device when no default is configured | + +A new `openai_hosted` Session reads `idle` while Core prepares its sandbox; its first Turn starts when the Environment is ready. The [Environment contract](../../contracts/agents-api/environments.md) owns placement, expiry and preparation. ### Session status -| `status` | Meaning | -| --- | --- | -| `in_progress` | A Turn is running | -| `idle` | Waiting for input | -| `requires_action` | Waiting for you: a [function result](#function-tools) or a self-hosted machine to connect | -| `failed` | The Environment could not be prepared. See `error` | +Read `status`, `error` and `required_actions` to decide whether to send input, return a [function result](#function-tools), connect a machine or diagnose a failure. [Session status](../../contracts/agents-api/sessions-events.md#session-status) defines every state and which failures allow new input. ### Update, list and delete @@ -273,8 +273,7 @@ client.beta.agents.sessions.delete(session.id) ## Send input -All input goes to one endpoint as a list of events. It returns 202 once the input is -stored, before the agent reads it. +All input goes to one endpoint as a list of events. It returns 202 after durable admission, before native application. On an idle `openai_hosted` or `self_hosted` Session, a message request can wait up to five minutes for its Turn to start and can end with a 409 expiry, cancellation or Environment error; allow that wait in client timeouts ([Environment input](../../contracts/agents-api/sessions-events.md#sessions-with-an-environment)). ### Send a message @@ -296,13 +295,9 @@ oac "/agents/sessions/$SESSION_ID/events" -H "Idempotency-Key: $KEY" -d '{ }' ``` -- **When idle,** a message starts a new Turn. **While a Turn runs,** it joins that - Turn (steering); it does not start a parallel task. -- **Content** is `input_text`, plus `input_image` as an inline PNG or JPEG data URI. - Images work with Codex and Claude Code wherever the Runtime supports them; - MiniMax Code rejects them. The - whole request is limited to 1 MiB. -- Full rules: [message input](../../contracts/agents-api/message-input.md). +- **When idle,** a message starts a new Turn. **While a Turn runs,** it joins that Turn (steering); it does not start a parallel task. +- **Content** is `input_text`, plus `input_image` as an inline PNG or JPEG data URI. Codex and Claude Code accept images; MiniMax Code rejects them. The whole request is limited to 1 MiB. +- Full rules: [message content](../../contracts/agents-api/message-content.md). ### Cancel @@ -314,19 +309,11 @@ client.beta.agents.sessions.events.create(session.id, events=[{"type": "agent.se oac "/agents/sessions/$SESSION_ID/events" -d '{"events": [{"type": "agent.session.input.cancel"}]}' ``` -The Turn is cancelled when it reaches `cancelled`, not when the request returns. A -cancel while idle does nothing. +The Turn is cancelled when it reaches `cancelled`, not when the request returns. A cancel while idle does nothing when no input is pending; a pending Environment input reservation returns 409. A self-hosted machine keeps its workspace and history when you restart the same installation; see [operate the installation](../getting-started/self-hosted.md#operate-the-installation). ## Stream events -`GET /agents/sessions/{id}/events` is a server-sent event stream. It is **live only**: -events sent while you were disconnected are not replayed. Open it before sending -input, and recover gaps from [Turns and Items](#turns-and-items). - -Open streams recheck the original Project key once per second, including before -output. Revocation or Project archival closes the stream; authentication failure -also closes it. Rechecks use the normal five-second authentication timeout and -send no Session data while waiting. Bytes already sent cannot be recalled. +`GET /agents/sessions/{id}/events` is a server-sent event stream. It is **live only**: events sent while you were disconnected are not replayed. Open it before sending input, and recover gaps from [Turns and Items](#turns-and-items). ```python with client.beta.agents.sessions.events.stream(session.id) as stream: @@ -356,21 +343,15 @@ data: {"type": "agent.session.turn.output_text.delta", "item_id": "item_...", "d | `agent.session.subagent.*` | Child work starts or ends | | `error` | A stream-level error | -The stream stays open across Turns. To stream one Turn and handle -[function calls](#function-tools) automatically, the SDK's `sessions.stream` helper -does both. +The stream stays open across Turns. To stream one Turn and handle [function calls](#function-tools) automatically, the SDK's `sessions.stream` helper does both. -**Stream the creation itself** with `stream=True` on `sessions.create`. You get -`agent.session.created` first, and the stream ends at the first `idle` or `failed`. +**Stream the creation itself** with `stream=True` on `sessions.create`. You get `agent.session.created` first, and the stream ends at the first `idle` or `failed`. -**Reconnecting:** resubscribe, then read Items and drop any you already have by ID. -Details: [history and events](../../contracts/agents-api/history-events-usage.md). +**Reconnecting:** resubscribe, then read Items and drop any you already have by ID. An open stream closes when its Project key is revoked or its Project archived. Details: [recovery model](../../contracts/agents-api/sessions-events.md#recovery-model). ## Turns and Items -A Turn is one piece of work started by input. Items are its recorded content: -messages, reasoning, tool calls and their results. Both are durable; read them to -check results or recover after a disconnect. +A Turn is one piece of work started by input. Items are its recorded content: messages, reasoning, tool calls and their results. Both are durable; read them to check results or recover after a disconnect. ```python turns = client.beta.agents.sessions.turns.list(session.id, order="desc") @@ -391,14 +372,12 @@ oac "/agents/sessions/$SESSION_ID/items?order=asc" | `waiting` | Waiting for a function result | | `completed`, `failed`, `cancelled` | Finished | -- **Usage** (`input_tokens`, `output_tokens`, `total_tokens`, …) may arrive after the - Turn ends. Treat `null` as "not yet known", not zero. +- **Usage** (`input_tokens`, `output_tokens`, `total_tokens`, …) is null when unknown, never zero. A Session's usage stays null while a Turn runs; Claude Code and MiniMax Code report none ([usage rules](../../contracts/agents-api/sessions-events.md#usage)). - Turn lists contain top-level Turns only. Read child work under `/subagents`. ## Function tools -Declare a function on the Agent. When the model calls it, the Session enters -`requires_action` and the Turn `waiting` until you return a result. +Declare a function on the Agent. When the model calls it, the Session enters `requires_action` and the Turn `waiting` until you return a result. ```python agent = client.beta.agents.create( @@ -414,6 +393,10 @@ agent = client.beta.agents.create( def get_weather(args): return f"Sunny in {args['city']}" +session = client.beta.agents.sessions.create( + environment={"type": "openai_hosted"}, agent_id=agent.id, +) + with client.beta.agents.sessions.stream( session.id, input="What's the weather in Paris?", tool_handlers={"get_weather": get_weather} ) as stream: @@ -430,8 +413,7 @@ oac "/agents/sessions/$SESSION_ID/events" -d '{ }' ``` -Resending the same result is safe. A different result for the same call, or one sent -after cancellation, fails with 409. MiniMax Code doesn't support public functions. +Resending the same result is safe. A different result for the same call, or one sent after cancellation, fails with 409. MiniMax Code doesn't support public functions. ## Files @@ -454,8 +436,7 @@ curl "$OPENAI_BASE_URL/files" -H "Authorization: Bearer $OPENAI_API_KEY" \ -F purpose=user_data -F file=@data.csv ``` -Only `purpose=user_data` is accepted, up to 512 MiB. No Beta header. Their content -cannot be downloaded again; use the ID in workspace files or templates. +Only `purpose=user_data` is accepted, up to 512 MiB. No Beta header. Their content cannot be downloaded again; use the ID in workspace files or templates. ### Workspace files @@ -469,18 +450,16 @@ page = client.beta.agents.environments.files.list(env_id, path="/workspace") ``` ```sh -oac "/agents/environments/$ENV_ID/files" -d '{"type": "file_id", "file_id": "file_...", "path": "/workspace/data.csv"}' +oac "/agents/environments/$ENV_ID/files" -d '{"type": "file_id", "file_id": "file-...", "path": "/workspace/data.csv"}' ``` - `inline` data is base64, up to 5 MiB decoded; `file_id` up to 50 MiB. - Parent directories are created. An existing file is never overwritten (400). -- The list shows regular files in one directory, not recursively. It pages with a - `page` token and returns `next`. +- The list shows regular files in one directory, not recursively. It pages with a `page` token and returns `next`. ### Artifacts -When a Turn completes, Core captures the regular files under the workspace's -`outputs/` directory. Artifacts stay readable after the Environment is gone. +When a Turn completes, Core captures the regular files under the workspace's `outputs/` directory. Artifacts stay readable after the Environment is gone. ```python for a in client.beta.agents.sessions.artifacts.list(session.id): @@ -493,13 +472,11 @@ oac "/agents/sessions/$SESSION_ID/artifacts" oac "/agents/sessions/$SESSION_ID/artifacts/$ARTIFACT_ID/content" -o report.md ``` -Each Artifact has `path`, `size_bytes`, `turn_id` and `environment_id`. Deleting one -leaves the workspace file alone. +Each Artifact has `path`, `size_bytes`, `turn_id` and `environment_id`. Deleting one leaves the workspace file alone. ## Skills -A Skill is a versioned bundle of instructions and files an agent can use. Upload a -directory or a ZIP; each upload is a version. +A Skill is a versioned bundle of instructions and files an agent can use. Upload a directory or a ZIP; each upload is a version. ```sh curl "$OPENAI_BASE_URL/skills" -H "Authorization: Bearer $OPENAI_API_KEY" -F files=@my-skill.zip @@ -510,17 +487,15 @@ skill = client.skills.create(files=[("my-skill/SKILL.md", open("my-skill/SKILL.m client.skills.versions.create(skill.id, files=[...], default=True) ``` -- No Beta header. Up to 50 Skills; 5 MiB compressed, 20 MiB expanded per archive. +- No Beta header. Select up to 50 Skills per Environment; each archive is at most 5 MiB compressed and 20 MiB expanded. - SDK 3.13.0 drops a single ZIP file from the upload; use HTTP for a ZIP. -- Attach Skills to a Session through its `environment.skills` or a - [template](#environment-templates). +- Attach Skills to a Session through its `environment.skills` or a [template](#environment-templates). -Details: [Source Files and Skills](../../contracts/agents-api/source-files.md). +Details: [Files and Skills](../../contracts/agents-api/source-files.md). A Session installs its Skills, Plugins and packages once, when it is prepared; editing the source later doesn't change a running Session. Preparation errors fail the Session before any work runs: fix the cause instead of retrying in a new Session. ## Environment Templates -A template saves workspace setup for reuse. `openai_hosted` Sessions reference it in -`environment`; `self_hosted` Sessions in `x_agents_core.environment`: +A template saves workspace setup for reuse. `openai_hosted` Sessions reference it in `environment`; `self_hosted` Sessions in `x_agents_core.environment`: ```python template = client.beta.agents.environments.templates.create( @@ -532,18 +507,16 @@ template = client.beta.agents.environments.templates.create( | Field | Meaning | | --- | --- | -| `network` | `access`: `enabled` (default), `disabled`, or `restricted` to 1–100 exact hosts in `allowed_domains`. A Session can only narrow it. Current Runtimes don't enforce `disabled` or `restricted`, so a Session that needs them is rejected; see [restricted network policy](../../contracts/agents-api/environments.md#restricted-network) | -| `packages` | `npm` and `python` packages. `system` packages are rejected: preinstall them in the image or on the machine | +| `network` | `access`: `enabled` (default), `disabled`, or `restricted` to 1–100 exact hosts in `allowed_domains`. A Session can only narrow it. See execution limits in [restricted network policy](../../contracts/agents-api/environments.md#restricted-network) | +| `packages` | Package setup; see [package admission](../../contracts/agents-api/environments.md#preparation-order) | | `setup_commands`, `env` | Run and set at preparation. Never returned by reads | | `files`, `skills`, `plugins` | Initial content. Up to 50 files, 10 MiB inline in total | -A Session freezes the template when it starts. Details: -[Environment Templates](../../contracts/agents-api/environments.md#templates). +A Session freezes the template when it starts. Details: [Environment Templates](../../contracts/agents-api/environments.md#templates). ## Vaults -Vaults hold credentials for HTTP MCP servers: a `static_bearer` token or an `mcp_oauth` -token with optional refresh. Tokens are write-only. +Vaults hold credentials for HTTP MCP servers: a `static_bearer` token or an `mcp_oauth` token with optional refresh. Tokens are write-only. ```python vault = client.beta.agents.vaults.create(name="github") @@ -554,31 +527,17 @@ client.beta.agents.vaults.credentials.create( session = client.beta.agents.sessions.create(environment={"type": "none"}, input="...", vault_ids=[vault.id], agent_id=agent.id) ``` -The HTTP path is `/vaults`, with the Beta header. A credential is used when an MCP -server's URL matches its `mcp_server_url` exactly, or when the tool names its -`credential_id`. +The HTTP path is `/vaults`, with the Beta header. A Session selects credentials from its `vault_ids`, optionally by `credential_id`; the MCP server's URL must match the selected credential's `mcp_server_url` exactly in either case. The [Vaults contract](../../contracts/agents-api/vaults.md) owns selection, errors, OAuth refresh and deletion. An MCP tool's `connection_origin` decides whether Core's side or the workspace connects to the server, and each harness supports a different set: see [MCP connection origin](../../contracts/agents-api/environments.md#public-mcp-connection-origin). -An MCP tool's `connection_origin` decides who connects: +## Diagnose a failure -| `connection_origin` | Connects from | Works with | -| --- | --- | --- | -| `service` (default) | Core's side, for `none` Sessions | Codex and Claude Code | -| `environment` | Inside the workspace (managed or your own machine) | Codex, Claude Code and MiniMax Code. MiniMax Code needs a null allowlist and `required: false` | +1. Read the Session's `status` and `error`, and the latest Turn's `error`. A failed Turn reports only a generic `internal_error`. +2. Check that the Environment is connected and its harness is available. +3. Check the harness, model and tool combination in [Harness capabilities](../../contracts/agents-api/harness-capabilities.md). +4. Ask the administrator for the Session's [diagnostics](../../contracts/agents-api/session-diagnostics.md), which name the failure category, and to check [troubleshooting](../getting-started/operations.md#troubleshooting) for service logs, credentials and node readiness. -Current rules: [public MCP connection origin](../../contracts/agents-api/environments.md#public-mcp-connection-origin). +A 401 usually means a key from another namespace; see [API namespaces and credentials](README.md). -## Limits and differences from OpenAI +## Differences from OpenAI -| Area | Core behavior | -| --- | --- | -| Routes | Exactly the pinned SDK's routes; no extra routes | -| Extensions | Only `x_agents_core.harness` and `x_agents_core.model_provider` | -| Session create idempotency | Same key returns the original Session | -| Event stream | Live only; no replay, `Last-Event-ID` ignored | -| Tools | No enabled `web_search`; no multi-agent with functions; MiniMax Code has no public functions | -| Images | Inline PNG/JPEG only; Codex and Claude Code; MiniMax Code rejects them | -| Packages | `packages.system` rejected | - -Per-operation status and harness differences are in the -[coverage record](../../contracts/agents-api/README.md). The exact schemas are in the -[public OpenAPI](../../contracts/agents-api/openapi.yaml). +Core differs from the OpenAI service in some behavior, such as Session creation idempotency and harness-specific tool support. The [coverage ledger](../../contracts/agents-api/README.md#differences-from-openai) lists every difference and the per-resource status; the [public OpenAPI](../../contracts/agents-api/openapi.yaml) has the exact schemas. diff --git a/docs/api/web-management.md b/docs/api/web-management.md deleted file mode 100644 index f942e1c2c..000000000 --- a/docs/api/web-management.md +++ /dev/null @@ -1,195 +0,0 @@ -# Web management API - -Core Web is an administrator console over `/core/v1`, authorized by the Core key: -resource inspection, Project and key management, credential issuance, audit, usage -and sandbox operations. It never calls `/v1` or `/api/v1`, and has no Agent -execution, copy or arbitrary asset editing operation. - -Core request failures use the [Core administration error envelope](../../contracts/agents-api/core-errors.md), -including optional typed safe details. Console sign-in endpoints retain their -separate error shape described below. - -## Browser to console - -The browser uses the console's own origin and signs in with the -[Core key](../getting-started/operations.md#core-key): - -| Method and route | Request | Result | -| --- | --- | --- | -| `GET /console/auth` | No body | `200 {"mode":"login"}` or `200 {"mode":"authenticated"}` | -| `POST /console/auth/login` | `Content-Type: application/json`, body `{"core_key":"…"}`; other members are rejected | `200 {"mode":"authenticated"}` and an HttpOnly, SameSite=Strict session cookie (Secure over HTTPS) | -| `POST /console/auth/logout` | No credential payload | `200 {"mode":"login"}`; clears the cookie and the server-side session | -| `GET /console/config` | Signed-in session | `node_installer`, `node_installer_sha256`, `node_artifacts`: the providers (`docker`, `microsandbox`) whose node assets this console holds | - -Sign-in errors use the console's `{"error": "…"}` envelope: 400 for a malformed -body, 401 for a wrong key, 415 for a non-JSON body, 429 with `Retry-After` when -failed attempts are limited or sign-in is busy, and 503 when the console cannot -start a session. The console compares the submitted key with its configured Core -key in constant time and never logs or returns it. Only failed attempts count -toward the limit; the correct key signs in even while failures are limited. The -console refuses to start with a Core key shorter than 32 characters. Sessions live only in the console's memory; a console restart -or Core key rotation requires signing in again. There are no console accounts, -usernames, passwords, account setup or Basic authentication. - -Use same-origin browser requests and cookies. Mutations require the same-origin -request checks; never put the Core key in JavaScript or browser storage. - -### Managed domain setup - -This console-local surface uses the signed-in session and the same-origin checks -above. It forwards to the installation controller, not Core. - -| Method and route | Request | Result | -| --- | --- | --- | -| `GET /console/installation/domain` | No body | Domain setup status | -| `POST /console/installation/domain` | `{"hostname":"core.example.com"}`; optional `confirm_public_url_change` equal to `https://core.example.com` | `202` and the current status; one asynchronous installer operation | - -Status contains `supported`, `state` (`unconfigured`, `checking`, `applying`, -`ready`, `failed`), nullable `public_url`, `target_url` and `message`. External -proxy installations report `supported: false`. A POST with existing address -bindings requires explicit confirmation and otherwise returns `409` with code -`public_url_confirmation_required`. Other rejections include invalid hostnames, -pending configuration edits and another installation operation holding the lock. -Action errors use `{"error":{"code":"…","message":"…"}}`; console authentication -failures retain the sign-in error shape above. - -Poll the same-origin GET while preparing. Applying the change restarts Web and -ends its sign-in sessions; provide a link to the submitted HTTPS origin for a -fresh login. A dropped request or cross-origin browser probe does not prove -success. The [installer contract](../../deploy/install/README.md#managed-https) owns -certificate verification, locking, retry and rollback. - -## Console to Core - -After sign-in and the same-origin checks, the console server forwards every -`/core/v1/*` request with its private Core key as the Bearer credential; Core -alone decides whether the route exists. It removes browser Authorization and -forwarding-sensitive headers, and sets `X-Core-Console-Actor: console`. Core -records the header as the audit `actor_label`. The label is caller-asserted and -display-only: direct Core key scripts normally send none, which records an empty -label, but could set any value. Never use it for authorization or as proof of -origin. - -Web calls Core only through the typed Core clients in `packages/agents-client`: -[AdminClient](../../packages/agents-client/src/admin-client.ts) (`/core/v1`), -`SandboxAdminClient` (`/core/v1/sandbox`) and `CoreMetricsClient` -(`/core/v1/metrics`). See [the complete administrator reference](../../contracts/agents-api/admin-api.md) -for methods, fields, filters, pagination and response shapes. -Routes below are relative to `/core/v1`: - -| Workflow | Routes | -| --- | --- | -| Projects | `GET/POST /projects`, `POST /projects/{id}`, `POST /projects/{id}/archive` | -| Project keys | `GET/POST /projects/{id}/keys`, `DELETE /projects/{id}/keys/{key_id}` | -| Resource lists/details/deletion | `/projects/{id}/agents`, `/sessions`, `/environment-templates`, `/skills`, `/files`, `/vaults`, including the documented nested reads | -| Asset ownership | `GET /projects/{id}/resource-owners` with batched resource IDs | -| Key operation history | `GET /projects/{id}/write-operations` with key/resource/time filters | -| Executor credentials | `GET/POST /projects/{id}/environments/{environment_id}/executor-credentials`, `DELETE …/executor-credentials/{key_id}` ([contract](../../contracts/agents-api/environment-executor-credentials.md)) | -| Deployment model configuration | `GET /harnesses`, `GET/PUT/DELETE /harnesses/{harness}/model-configuration`; the key is write-only ([contract](../../contracts/agents-api/model-execution.md#deployment-defaults)) | -| Usage and health | `GET /summary`, `/sandbox/runtime-observations`, `/metrics` | -| Administrator audit | `GET /audit-log`; deployment-wide entries have `project_id: null` | - -A Project UUID in a management path selects the target; it is not a credential. -API-key plaintext is returned only by successful issuance, so display it once and -never cache it. Reconcile uncertain issuance before explicitly issuing another key. -Deleted/revoked resources retain their audit records. Historical unknown ownership -stays null. Session usage grouped by key belongs to the Session's creation key; -it is operational attribution, not per-key billing. - -## Sandbox administration - -The [deployment configuration contract](../../contracts/agents-api/sandbox-deployment.md), -[nodes guide](../getting-started/nodes.md) -and [generated OpenAPI](../../contracts/agents-api/core.openapi.yaml) -define deployment and node operations: - -- `GET/POST/PUT /core/v1/sandbox/deployment` and - `POST/DELETE /core/v1/sandbox/deployment/reset`. -- `GET /core/v1/sandbox/nodes`, `PATCH/DELETE /core/v1/sandbox/nodes/{node_id}`, - and `GET /core/v1/sandbox/nodes/{node_id}/allocations`. -- `POST /core/v1/sandbox/enrollment-tokens` for a one-time node installation command. - Its non-secret `enrollment_id` reappears on the node that command registers. - -PostgreSQL owns one provider, per-sandbox resource specification and immutable -Runtime selection. POST initializes it and PUT updates the same provider; both -require the observed `expected_generation`, including zero at first setup. Requests carry `resources` and, for Docker/microsandbox, -`runtime`; safe responses return `specification` and `specification_digest`. -Response `resources.allocations` and `resources.pending` are cleanup counts, not -CPU, memory or disk settings. E2B accepts a write-only key and exact template build -instead of a node Runtime release, and provisions without a node installation. E2B -may omit `resources` to adopt the validated build's CPU and memory. An E2B-compatible -service may also supply paired `configuration.api_url` and `configuration.domain`; omitted selectors -use official E2B. Responses expose these addresses but never the key, and changing -them requires the same drained maintenance transition as changing the template. -Responses show -the build as read at selection time in `metadata.template_build`. Microsandbox responses -return its idle `suspension` policy; other providers return null. - -Same-team E2B updates apply online after verification. Existing sandboxes retain -their generation and use the committed credential for management; omitting the key -preserves it, while explicitly submitting even the same key verifies a replacement. -Node-provider updates still require zero retained/pending resources and no reset. -Changing backend or E2B team requires explicit durable reset before a new POST. -Auto archives idle/queued/suspended hosted Sessions, waits for started work and -file writes, and escalates at its persisted deadline; force requests cancellation -and verified cleanup. The deployment response supplies the authoritative -`reset.remaining` partition and offline-node subset; Web must not derive either -from independently loaded lists. The separate `rollout.state` describes target -preparation; old-generation resource counts alone do not imply active preparation. -Poll rapidly while reset is active or rollout is preparing. Node `ready_generation` -is a durable serving pin, not proof of current connectivity. Target unknown, failed -or update-required state does not by itself invalidate confirmed old-generation -service; consume Core's connection/provider facts separately. - -Administrators may [archive an individual hosted Session](../../contracts/agents-api/admin-api.md#administrative-session-archive) -at the current generation without reset. History and persisted Files/Artifacts -survive; unpersisted workspace is lost and the original Session cannot resume. -Cancel reset stops further archives, not cleanup already requested. After zero -resources Core clears the selection and advances generation; configure again using -that new generation. Never automatically replay an uncertain write. -Public Environment Templates, the `/v1` contract and -caller-owned `self_hosted` provisioning remain unchanged. - -`GET /api/v1/sandbox-node/configuration` uses an enrollment Bearer token, or a -retained node Bearer credential with `X-OAC-Node-ID`. This read does not consume -enrollment. Retained matching nodes can read their configuration during reset. -Installers must verify the returned generation, specification digest and Runtime -before registration; local files cannot override the saved limits. A mismatch -returns `sandbox_specification_mismatch` without replacing node state. - -Node configuration, enrollment, identity and connection routes live under -`/api/v1/sandbox-node`, beside the daemon's `/api/v1/agent-daemon`. The reverse -proxy sends `/api/v1` directly to Core; the console returns 404 for it and never -forwards machine traffic. These routes use their own credentials, do not inherit a -browser login and gain no management authority. E2B credentials are absent -from node configuration and safe deployment views. - -## Frontend handoff and errors - -Frontend screen implementation is owned by the separate frontend task. The -backend provides the `/core/v1` contract, the typed Core clients and the console -proxy; backend tests do not qualify the screens. - -The console returns 404 for `/v1`, `/api/v1` and old `/console/api-keys` routes, -even with an explicit application or machine Bearer credential. Invalid -Origin/Host requests are rejected; unavailable Core or rejected upstream redirects -return 502. Console authentication uses its own error envelope. Management resource -errors and deletion constraints are documented in the administrator reference and -generated schema; do not interpret every empty or failed read as an absent resource. - -## Core metrics - -`GET /core/v1/metrics?range=1h|6h|24h|7d` returns Core process, execution -queue/slots, PostgreSQL and background-job measurements. It requires the Core key, -rejects arbitrary query filters and never grants -Agent execution access. See the [exact measurement contract](../../contracts/agents-api/core-metrics.md) -for complete buckets, null values, units and process-local retention. Frontend -implementation is maintained separately; this backend change does not modify -Agent metrics or the public Agent API. - -## Node host history - -The Core key reads a node and its host history through -`GET /core/v1/sandbox/nodes/{node_id}?range=1h|6h|24h`. See the -[node host history contract](../../contracts/agents-api/node-host-history.md) -for nullable observations, freshness and aggregation. The node list is unchanged. diff --git a/docs/architecture.md b/docs/architecture.md index dcb19e699..fcd3566cb 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -89,5 +89,5 @@ The `none` profile shares the execution protocol without workspace preparation. - **Isolation belongs to the outer Environment.** The daemon is not a sandbox ([Runtime and outer isolation](design-principles.md#runtime-and-outer-isolation)). - **Execution and compute have separate lifetimes.** Closing an executor does not release its allocation, destroy its Environment or delete its workspace. Reclamation is an explicit Sandbox Provider operation. -- **Model keys stay with the compute that owns them.** A self-hosted Session brings its own model provider ([why](user-guide.md#which-model-provider-a-session-uses)). +- **Model keys stay with the compute that owns them.** A self-hosted Session brings its own model provider ([why](../contracts/agents-api/model-execution.md#saved-defaults-and-precedence)). - **Core Web is an administrator console.** It calls only `/core/v1` and cannot start Sessions or send input ([console API usage](web/console-api-usage.md#not-consumed)). diff --git a/docs/configuration.md b/docs/configuration.md index b258e0d05..4a20e1b14 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -96,7 +96,7 @@ Runtime settings live in Core's database. Change them in Web; scripts use the sa | Default model per harness | **System** → **Default model configuration**: **Set** | `/core/v1/harnesses/{harness}/model-configuration` | See [Default models](#default-models) | | Executor credentials of a self-hosted Session | **Session log**, then the **Session** page: **Executor credentials** | `/core/v1/projects/{project_id}/environments/{environment_id}/executor-credentials` | See [self-hosted executors](getting-started/self-hosted.md) | -Which harnesses are enabled, and the default one, are process settings (`core.harnesses`, `core.default_harness`); System shows them read-only. The [API index](api/README.md#core-api) lists every Core API route, and the [deployment contract](../contracts/agents-api/sandbox-deployment.md) defines the sandbox fields, limits and change rules. +Which harnesses are enabled, and the default one, are process settings (`core.harnesses`, `core.default_harness`); System shows them read-only. The [Core administration API](../contracts/agents-api/admin-api.md) lists every Core API route, and the [deployment contract](../contracts/agents-api/sandbox-deployment.md) defines the sandbox fields, limits and change rules. ### Node capacity diff --git a/docs/getting-started/README.md b/docs/getting-started/README.md index 1c1b4cf52..3648f1174 100644 --- a/docs/getting-started/README.md +++ b/docs/getting-started/README.md @@ -22,8 +22,7 @@ Application developers call the API with a Project API key, and can run Sessions | Guide | Covers | | --- | --- | | [Quickstart](quickstart.md) | From a Project API key to a finished Session | -| [User guide](../user-guide.md) | Common tasks: harness and model choice, follow-ups, capabilities, files, cancel | -| [Agents API guide](../api/public-agent-api.md) | Every resource, with SDK and HTTP examples | +| [Agents API guide](../api/public-agent-api.md) | Common tasks, harness and model choice, and every resource with SDK and HTTP examples | | [Self-hosted execution](self-hosted.md) | Running a Session on your own machine | | [Examples](../examples.md) | Complete applications built on the API | | [API index](../api/README.md) | All three namespaces and their credentials | diff --git a/docs/getting-started/nodes.md b/docs/getting-started/nodes.md index 07365e55d..cc09d9d49 100644 --- a/docs/getting-started/nodes.md +++ b/docs/getting-started/nodes.md @@ -103,7 +103,7 @@ Sandbox settings apply to the whole installation; nodes follow them. - **Size, Runtime release, E2B key or template build.** Open **System** → **Manage sandbox configuration** → **Change resources**, edit and save. Existing sandboxes keep their configuration. Each node prepares the new one while it keeps serving the old one, and **Nodes** shows its progress: **Preparing target**, **Ready for target** or **Preparation failed** with the [reason](#readiness-codes). Core places new Sessions on nodes ready for the new configuration first, and on nodes still serving an older one when those have no room. E2B changes apply at once. - **Backend.** On the same page, choose **Reset deployment**. Auto reset archives idle hosted Sessions at once and lets running work finish until the deadline you set; force cancels it now. Offline nodes must come back so Core can confirm their cleanup. When the reset completes, Core has retired every node and unused command: set up the new backend, then add nodes again. **Cancel reset** stops the remaining work; archived Sessions stay archived. -The [reset contract](../../contracts/agents-api/sandbox-deployment.md#generation-ownership-and-rollout) describes what reset archives and keeps. To archive a single hosted Session, use the [Core API](../../contracts/agents-api/admin-api.md#administrative-session-archive). +The [reset contract](../../contracts/agents-api/sandbox-deployment.md#generation-ownership-and-rollout) describes what reset archives and keeps. To archive a single hosted Session, use the [Core API](../../contracts/agents-api/admin-api.md#session-archive). ## Remove a node diff --git a/docs/getting-started/operations.md b/docs/getting-started/operations.md index 8eabf4e32..378614282 100644 --- a/docs/getting-started/operations.md +++ b/docs/getting-started/operations.md @@ -92,7 +92,7 @@ core() { # core METHOD PATH [JSON body] | Set Codex's default model | `core PUT /harnesses/codex/model-configuration '{"model": "your-model-id", "model_provider": {"protocol": "responses", "base_url": "https://provider.example/v1", "api_key": "sk-..."}}'` | | Installation facts, including the API base URL | `core GET /installation` | -The [API index](../api/README.md#core-api) lists every route; errors use the [Core error envelope](../../contracts/agents-api/core-errors.md). +The [Core administration API](../../contracts/agents-api/admin-api.md) lists every route; errors use the [Core error envelope](../../contracts/agents-api/core-errors.md). ### Rotate the Core key diff --git a/docs/getting-started/quickstart.md b/docs/getting-started/quickstart.md index 29e6ad915..963ec79d0 100644 --- a/docs/getting-started/quickstart.md +++ b/docs/getting-started/quickstart.md @@ -106,7 +106,7 @@ Success is a `completed` Turn whose Items describe the new file. A timeout neith | To | Read | | --- | --- | -| Stream output, send follow-up messages, upload files, add Skills or MCP, cancel | [User guide](../user-guide.md) | +| Stream output, send follow-up messages, upload files, add Skills or MCP, cancel | [Agents API guide](../api/public-agent-api.md#common-tasks) | | See every resource with request and response examples | [Agents API guide](../api/public-agent-api.md) | | Run the agent on your own machine | [Self-hosted execution](self-hosted.md) | | See a complete application | [Examples](../examples.md) | diff --git a/docs/getting-started/self-hosted.md b/docs/getting-started/self-hosted.md index 011c665f8..7462d3209 100644 --- a/docs/getting-started/self-hosted.md +++ b/docs/getting-started/self-hosted.md @@ -4,7 +4,7 @@ A `self_hosted` Session runs on a machine your application owns: a workstation, **The daemon is not a sandbox.** Tools run with the permissions of the account that starts it and can reach whatever that account can. Use a container or VM when you need isolation; see [Runtime and outer isolation](../design-principles.md#runtime-and-outer-isolation). The daemon does not restrict network access, so a Template that requires a network policy is rejected for a self-hosted Session. -The Session brings its own model provider; the installation default never applies ([why](../user-guide.md#which-model-provider-a-session-uses)). The machine gets an executor credential that works for this one Environment and nothing else. +The Session brings its own model provider; the installation default never applies ([why](../../contracts/agents-api/model-execution.md#saved-defaults-and-precedence)). The machine gets an executor credential that works for this one Environment and nothing else. ## Platforms diff --git a/docs/runtime-protocol.md b/docs/runtime-protocol.md index 2782464eb..96681656b 100644 --- a/docs/runtime-protocol.md +++ b/docs/runtime-protocol.md @@ -280,7 +280,7 @@ Core-managed Docker `openai_hosted` and user-managed `self_hosted`. MiniMax images and remote URLs remain explicit implementation gaps. Workspace images reuse the existing preparation, active-input and workspace authority; they do not add a downloader, a mount or a separate execution lifecycle. Core -does not fetch or transform media. See [message input coverage](../contracts/agents-api/message-input.md). +does not fetch or transform media. See [message input coverage](../contracts/agents-api/message-content.md#images). ## Preparation and execution order diff --git a/docs/user-guide.md b/docs/user-guide.md deleted file mode 100644 index 4548a1f1b..000000000 --- a/docs/user-guide.md +++ /dev/null @@ -1,156 +0,0 @@ -# User guide - -Common tasks for application developers. Start with the -[quickstart](getting-started/quickstart.md); it sets up `client` and `session`. Every -operation here is shown with SDK and HTTP examples in the -[Agents API guide](api/public-agent-api.md). - -## Choose where a task runs - -A Session's `environment.type` decides where its agent works: - -| Environment | Runs on | You prepare | -| --- | --- | --- | -| `openai_hosted` | A sandbox Core creates on a node or E2B | Nothing; the administrator provides capacity | -| `self_hosted` | Your own Linux, macOS or Windows machine | The machine, a workspace directory and the daemon. See [self-hosted execution](getting-started/self-hosted.md) | -| `none` | An existing device connection, no workspace | A connected device with its harness | - -Placement, expiry and supported combinations: -[Environment contract](../contracts/agents-api/environments.md). - -## Choose a harness and a model - -The harness is the agent program that runs your Session. Set it in -`x_agents_core.harness` on the Agent; if you don't, the installation's default -(Codex unless changed) is used. - -| Harness | `harness` | Provider `protocol` | Native model parameters (`harness_config`) | -| --- | --- | --- | --- | -| Codex | `codex` | `responses` | `model_reasoning_effort` | -| Claude Code | `claude_sdk` | `anthropic` | `effort`, `thinking` | -| MiniMax Code | `mcode` | `anthropic` (default), `responses` or `chat_completions` | None. The provider must set `context_window` and `max_output_tokens` | - -- **Model.** `agent.model` is the provider's exact model ID. An inline Agent on an - `openai_hosted` or `none` Session may omit it to use the installation default's - model. Saved Agents always need one. -- **Provider protocol.** The harness connects to your provider directly, so the - provider must speak one of the harness's protocols above. There is no conversion; - a mismatch is rejected when the Session is created. -- **Native parameters.** `x_agents_core.harness_config` passes the harness's own - settings; see [native model configuration](../contracts/agents-api/harness-onboarding.md#native-model-configuration). - -Selecting a harness doesn't make an unsupported model or operation work; see -[harness selection](../contracts/agents-api/model-execution.md#harness-selection). - -### Which model provider a Session uses - -A Session takes its model provider from the first of these that has one. They are -never merged: - -1. `x_agents_core.model_provider` in the Session creation request; -2. the saved Agent's `x_agents_core.model_provider`; -3. the installation's default model configuration for the harness, set by the - administrator. It also supplies the model and native parameters when the Session - leaves them out. - -| Environment | Request or Agent provider | Installation default | Neither | -| --- | --- | --- | --- | -| `openai_hosted` | Used | Used | 400 `model_provider_required` | -| `self_hosted` | Used | Never | 400 `model_provider_required` | -| `none` | Rejected (400) | Used | Allowed: the device supplies the model | - -Self-hosted Sessions never use the default, because it holds the operator's key and -your machine is outside the operator's control. A Session freezes its provider at -creation; later changes affect only new Sessions. Full rules: -[model execution](../contracts/agents-api/model-execution.md). - -## Follow up and steer - -Send another message to the same Session to continue with its history and workspace: - -```python -client.beta.agents.sessions.events.create(session.id, events=[{ - "type": "agent.session.input.message", - "input": [{"role": "user", "content": [{"type": "input_text", "text": "Summarize the file you created."}]}], -}]) -``` - -- **Idle Session:** the message starts a new Turn. -- **Running Turn:** the message joins it (steering). It never starts a parallel task. -- **Different configuration or workspace:** create a new Session. - -See [Send input](api/public-agent-api.md#send-input). - -## Watch the output live - -Open the event stream before sending input, and print text as it arrives: - -```python -with client.beta.agents.sessions.events.stream(session.id) as stream: - for event in stream: - if event.type == "agent.session.turn.output_text.delta": - print(event.delta, end="", flush=True) - if event.type in {"agent.session.turn.completed", "agent.session.turn.failed", "agent.session.turn.cancelled"}: - break -``` - -The stream is live only. After a disconnect, read Turns and Items to fill the gap. -See [Stream events](api/public-agent-api.md#stream-events). - -## Give the agent Skills, Plugins and MCP - -Choose capabilities when you create the Session. The Runtime installs them once, as a -snapshot; editing the source later doesn't change a running Session. - -| To | Use | -| --- | --- | -| Upload a Skill and pick a version | [Skills](api/public-agent-api.md#skills) | -| Reuse packages, files and setup commands | [Environment Templates](api/public-agent-api.md#environment-templates) | -| Call your own code from the agent | [Function tools](api/public-agent-api.md#function-tools) | -| Connect an MCP server, with credentials | [Execution tools](../contracts/agents-api/execution-tools.md) and [Vaults](api/public-agent-api.md#vaults) | -| Use Skill or Plugin directories on your machine | [Local capability directories](getting-started/self-hosted.md#local-capability-directories) | - -System packages must already be in the machine or image. Preparation errors fail the -Session before any work runs; fix the cause instead of retrying in a new Session. - -## Send files and get results - -| To | Use | -| --- | --- | -| Put a file in the workspace | [Workspace files](api/public-agent-api.md#workspace-files) | -| Upload once, reference by ID | [Source Files](api/public-agent-api.md#source-files) | -| Download what the agent wrote to `outputs/` | [Artifacts](api/public-agent-api.md#artifacts) | -| Read the conversation and tool results | [Turns and Items](api/public-agent-api.md#turns-and-items) | - -Deleting a Session doesn't erase files on your own machine. Reclaiming a managed -sandbox is a separate operation. - -## Cancel and recover - -Cancel the current Turn: - -```python -client.beta.agents.sessions.events.create(session.id, events=[{"type": "agent.session.input.cancel"}]) -``` - -The work has stopped when the Turn reads `cancelled`, not when the request returns. - -After a lost response or connection: - -1. Read the Session, its Turns and Items. -2. If the work isn't there, retry with the **same** `Idempotency-Key`. -3. Never resend without a key; you may run the work twice. - -On a self-hosted machine, restart the same installation to keep its workspace and -history; see [operating the installation](getting-started/self-hosted.md#operate-the-installation). - -## Diagnose a failure - -1. Read the Session's `status` and `error`, and the latest Turn's `error`. -2. Check that the Environment is connected and its harness is available. -3. Check the harness, model and tool combination in - [execution tools](../contracts/agents-api/execution-tools.md). -4. Ask the administrator to check [operations and troubleshooting](getting-started/operations.md#troubleshooting) - for service logs, credentials and node readiness. - -A 401 usually means a key from the wrong namespace; see the [API index](api/README.md). diff --git a/packages/agents-client/README.md b/packages/agents-client/README.md index e8564d06a..1e2854aad 100644 --- a/packages/agents-client/README.md +++ b/packages/agents-client/README.md @@ -87,7 +87,7 @@ const defaults = await admin.retrieveHarnessModelConfiguration("codex"); console.log(defaults.model, defaults.harness_config, defaults.model_provider.api_key_configured); ``` -[Execution configuration queries](../../contracts/agents-api/execution-configuration.md) and [deployment defaults](../../contracts/agents-api/model-execution.md#deployment-defaults) define these reads and writes. +[Execution configuration queries](../../contracts/agents-api/admin-api.md#execution-configuration) and [deployment defaults](../../contracts/agents-api/model-execution.md#deployment-defaults) define these reads and writes. ### Checks diff --git a/packages/claude-sdk-adapter/README.md b/packages/claude-sdk-adapter/README.md index 485a97560..2eed904d1 100644 --- a/packages/claude-sdk-adapter/README.md +++ b/packages/claude-sdk-adapter/README.md @@ -72,7 +72,7 @@ SDK state lives under `paths.ProfileDir(profile)/runtime/claude-sdk`, independen The SDK descriptor advertises the validated daemon subset, including durable Turns/input receipts, text observations, function tools, raw usage and restrictive execution controls. It does not advertise permissions, product authoring, raw tool Items, general web-search control or text-verbosity levels. Router admission for `environment:none` uses the available engine capability, not an engine name. -Core owns public admission for Claude: [harness selection](../../contracts/agents-api/model-execution.md#harness-selection) chooses the engine for each Session; one [engine policy](../../contracts/agents-api/harness-onboarding.md#add-the-engine-to-core) serves API admission, device selection and the final claim; the [qualified operations](../../contracts/agents-api/harness-capabilities.md) table records Claude's medium-only verbosity and object-root function schemas; [function result images](../../contracts/agents-api/function-result-images.md) defines which placements accept image results; and [deployment defaults](../../contracts/agents-api/model-execution.md#deployment-defaults) define the model provider Core freezes for a Session. The adapter receives that provider as the adapter-owned `model_provider`, never in public Session configuration. Without one, a `none` host uses the daemon's own provider environment. The adapter alone selects the provider environment and removes credentials from native tool environments. +Core owns public admission for Claude: [harness selection](../../contracts/agents-api/model-execution.md#harness-selection) chooses the engine for each Session; one [engine policy](../../contracts/agents-api/harness-onboarding.md#add-the-engine-to-core) serves API admission, device selection and the final claim; the [qualified operations](../../contracts/agents-api/harness-capabilities.md) table records Claude's medium-only verbosity and object-root function schemas; [function result images](../../contracts/agents-api/message-content.md#function-results) defines which placements accept image results; and [deployment defaults](../../contracts/agents-api/model-execution.md#deployment-defaults) define the model provider Core freezes for a Session. The adapter receives that provider as the adapter-owned `model_provider`, never in public Session configuration. Without one, a `none` host uses the daemon's own provider environment. The adapter alone selects the provider environment and removes credentials from native tool environments. The `none` profile accepts only text, explicit model/system instructions, managed state, exact native resume and declared functions with ordered text or successful inline PNG/JPEG results, and the HTTP MCP subset described above. It rejects unsupported request options and disables built-in tools and undeclared MCP discovery. `DisableExecutionEnvironment` and `DisableSubagents` are accepted assertions about the single-Agent restrictive profile. Omission does not enable built-in tools. Explicit Subagent observation enables only the native delegation tools described under [Subagents](#subagents). Single-Agent new and resumed queries use the SDK's empty built-in tool set, explicit function MCP configuration and allowlist, strict MCP configuration and empty user/project/local setting sources. Without HTTP MCP declarations, native initialization and real provider request inventories must contain only the declared host functions. Managed operator policy may further restrict execution; it must not widen the profile. The profile limits model tool access; it does not isolate native state files or filesystem access by an explicitly supplied host function. The private factory accepts typed execution controls only for disabled search and medium text verbosity. Search remains excluded by the native tool inventory; medium retains the SDK's default text generation, without adding instructions or changing caller input. The pinned SDK has no native verbosity-level option, so low and high verbosity and enabled search are unsupported. Missing/invalid fields in a supplied control block fail before native setup; omitting the block keeps the same restrictive profile. Use the SDK's history lookup before explicit resume; never fall back to a new Session. Native files remain device-affine under a caller-selected managed runtime directory. The launch configuration supplies trusted provider environment; request options cannot supply environment variables or business write authority. Omitted, null and empty `system_prompt` map to empty SDK instructions only at this adapter boundary; null model values and unsupported options remain rejected. diff --git a/scripts/core-distribution-manifest.py b/scripts/core-distribution-manifest.py index 2b7a98226..114af72c7 100644 --- a/scripts/core-distribution-manifest.py +++ b/scripts/core-distribution-manifest.py @@ -36,7 +36,7 @@ BUNDLED_DOCS = ( "README.md", "README.zh-CN.md", - "docs/user-guide.md", + "docs/api/public-agent-api.md", "docs/development.md", "docs/configuration.md", "docs/getting-started/README.md", diff --git a/scripts/name-allowlist.json b/scripts/name-allowlist.json index 8dccb5db3..9fa95fb93 100644 --- a/scripts/name-allowlist.json +++ b/scripts/name-allowlist.json @@ -234,16 +234,6 @@ "regex": "in Parsar\\.|apps/parsar/|Parsar is an ordinary client|Parsar integration|Parsar owns|`parsar` provider slug", "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." }, - { - "path": "contracts/agents-api/README.md", - "regex": "Parsar owns product|Parsar, not the HTTP contract|Parsar cutover|Parsar services|no Parsar dependency", - "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." - }, - { - "path": "contracts/agents-api/official-semantics-alignment.md", - "regex": "The Parsar\\n", - "reason": "This caller verification paragraph records the separate product client audit." - }, { "path": "docs/design-principles.md", "regex": "Parsar is an ordinary API-key holder", @@ -259,11 +249,6 @@ "regex": "Parsar's product|Parsar\\sworkspace|Parsar Skill/SP", "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." }, - { - "path": "services/core/credentials.md", - "regex": "Parsar approval", - "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." - }, { "path": "services/core/deploy/claude/README.md", "regex": "Parsar is not a dependency", @@ -354,11 +339,6 @@ "regex": "agents-api-", "reason": "The regression fixture uses the existing persisted Session state-key namespace required by Runtime binding validation." }, - { - "path": "contracts/agents-api/README.md", - "regex": "https://github\\.com/MiniMax-AI-Dev/parsar/pull/[0-9]+", - "reason": "Historical acceptance evidence links to pull requests in the separate upstream Parsar repository." - }, { "path": "services/core/migrations/000015_retire_item_backfill.sql", "regex": "services/agents-api/README\\.md", diff --git a/services/core/IMPLEMENTATION.md b/services/core/IMPLEMENTATION.md index 5a203a8ea..22dad4f2a 100644 --- a/services/core/IMPLEMENTATION.md +++ b/services/core/IMPLEMENTATION.md @@ -29,7 +29,7 @@ per-family limit bounds rather than applying one policy to every resource. Chang page bounds, cursor ownership or parent lookup order only with owned evidence for that family. Record uncertain range/lookup behavior separately; do not reproduce observed upstream server failures as compatibility behavior. See -`contracts/agents-api/list-query-semantics.md` for the bounded evidence. +[list rules](../../contracts/agents-api/wire-semantics.md#lists). Every Agents API JSON route reads its body through the shared gate (`readJSONObject`) before route decoding, validation or lookup. It requires a JSON @@ -40,7 +40,7 @@ Core extension and internal routes keep their own readers. Member names match exactly: decode request objects with `decodeInputObject`, or check `inexactMember` before another decoder, so that encoding/json never matches a case variant to a field. See -`contracts/agents-api/official-semantics-alignment.md#request-body-parsing--september-23`. +[request bodies](../../contracts/agents-api/wire-semantics.md#request-bodies). Report validation failures with official evidence through the typed field error, which emits `invalid_request_error` with the observed param and message; keep other local codes until their official fields are sampled. Every 409 has type @@ -58,10 +58,10 @@ their own errors. An `after` cursor that does not resolve inside its already resolved parent, malformed ones included, returns that list family's observed error: the missing-resource 404 on lookup lists, otherwise the typed store cursor error. Foreign and missing cursors stay identical; see -`contracts/agents-api/list-query-semantics.md`. Reject U+0000 in metadata +[list rules](../../contracts/agents-api/wire-semantics.md#lists). Reject U+0000 in metadata explicitly with its `metadata.` param; other stored strings rely on the PostgreSQL error mapping, so keep each request's writes in one transaction. See -`contracts/agents-api/official-semantics-alignment.md`. +[validation errors](../../contracts/agents-api/wire-semantics.md#validation-errors). Serve requests on their canonical path and never redirect. `api.CanonicalPaths` wraps the complete server handler in both configurations (the daemon ServeMux and @@ -216,7 +216,7 @@ from prior Docker or retired remote-executor evidence. ## Current implementation constraints The constraints below describe existing code, not requirements to preserve legacy -design. The [protocol assessment](../../contracts/agents-api/README.md#implementation-direction) +design. The [protocol assessment](../../contracts/agents-api/README.md#known-gaps) identifies replacements and gaps. Update these rules when their implementation is replaced; do not carry obsolete compatibility code forward to satisfy this section. @@ -317,7 +317,7 @@ replaced; do not carry obsolete compatibility code forward to satisfy this secti Never reuse product master-key conventions or daemon transport encryption for this storage boundary. Missing key configuration disables credential writes; malformed explicit configuration fails startup. See - [`services/core/credentials.md`](credentials.md) for + [Vaults and Credentials](../../contracts/agents-api/vaults.md#storage-key) for key persistence and current limits. Storage-key rotation remains separate work; resource creation never contacts the destination. - `GET /v1/vaults/{vault_id}/credentials` lists safe metadata only, with both @@ -363,7 +363,7 @@ replaced; do not carry obsolete compatibility code forward to satisfy this secti in `OAC_OAUTH_TRUSTED_ORIGINS`; tenants cannot relax that boundary and TLS verification remains mandatory. Keycloak is acceptance infrastructure only. Preserve the pinned update omission/null and immutable-field rules described in - [OAuth credentials](oauth-credentials.md); record unspecified + [OAuth credentials](../../contracts/agents-api/vaults.md#oauth); record unspecified hosted semantics. Native processes receive only access tokens. Provider revocation, withdrawal of already-dispatched tokens and Session cancellation remain distinct. - Credential `DELETE /v1/vaults/{vault_id}/credentials/{credential_id}` removes one @@ -548,7 +548,7 @@ replaced; do not carry obsolete compatibility code forward to satisfy this secti unknown sources and unavailable provider metadata, without backfill. Projection metadata does not alter retry identity; retries cannot replace it. Keep this administrator query separate from runtime observations and do not touch activity or wake sandboxes. The versioned - contract is `contracts/agents-api/execution-configuration.md`. + contract is [execution configuration](../../contracts/agents-api/admin-api.md#execution-configuration). - Model communication uses native direct connections only. Core sends one frozen confidential `model_provider` bundle, independent of engine and placement; adapters apply it through their native provider configuration. The shared @@ -1096,7 +1096,7 @@ replaced; do not carry obsolete compatibility code forward to satisfy this secti including after terminal or later Turns. Omitted/null input is permitted only for non-streaming hosted creation and self-hosted creation. Creation streaming uses the shared live path above. Image support requires the - qualification in [the message-input contract](../../contracts/agents-api/message-input.md). + qualification in [the message-input contract](../../contracts/agents-api/message-content.md#images). ### Worker ownership @@ -1602,3 +1602,46 @@ required history is missing. Daemon tools run with the launching account's full permissions. Managed isolation is provided by the outer Environment. A user who installs on a host does not receive a sandbox or protection from their own tools. Keep failed probes and unverified platform combinations explicit. + +## Public and administration API constraints + +These constraints implement the rules in the [Agents API contracts](../../contracts/agents-api/README.md) and the [Core administration API](../../contracts/agents-api/admin-api.md). + +### Wire validation mechanics + +- Stored strings other than metadata rely on PostgreSQL rejecting U+0000 and invalid UTF-8: map SQLSTATE `22021` (text parameter) and `22P05` (`\u0000` in jsonb) to the 400 unstorable-text error, including query filters such as `agent_id`. The failing statement aborts its transaction, so keep each request's writes in one transaction. +- An `after` cursor that cannot name a resource on a lookup list (Agents, Sessions, Turns, Templates, Vaults, Credentials, the Core Runtime observation list) resolves to the never-assigned maximum UUID and runs the normal lookup, so storage failures and missing rows behave as for a well-formed cursor. Resolve every cursor only inside its already resolved parent and tenant. +- Lists whose parent and cursor lookups are separate statements (Artifacts, Skill versions) re-check the parent before reporting a cursor 400, so a parent deleted in between still returns its 404. Item and Subagent lists read both inside one locked Session transaction. The Skill version cursor lookup is tenant-wide so another Skill's version can be told apart from a missing one; another tenant's version stays missing. +- Saved Agents parse tools with the saved-form parser, which keeps every pinned `web_search` mode. Session admission re-resolves the effective tools with the execution parser, which admits only disabled search; Worker device selection and the final preclaim also refuse a non-disabled search control. + +### Message text and result targets + +- Admission and the Claude bridge share one whitespace set, the union of Go `unicode.IsSpace` and ECMAScript `String.prototype.trim` (`blankTextRune` in `services/core/internal/execution/message_support.go`). A shared table test keeps both sides equal; change them together. +- Whitespace-only admission is a field of the engine profile (`WhitespaceOnlyText`) checked with the image profile during Worker admission. Never branch on the harness name in handlers. +- `MessageInput.Validate` treats any non-empty text part as content and never trims. It runs in Core admission, Worker delivery, daemon steering and prepared start, and in the Codex and MiniMax adapters. +- Resolve a tool result's target under the tenant Session lock, after the Session lookup: a well-formed `turn_id` is looked up in that Session and must own the call; otherwise the Session's own calls decide between "unknown call" and "different Turn". Never reject a malformed `turn_id` before the Session lookup, so missing, malformed and foreign Sessions keep one 404, and read only the caller's Session. +- The recorded Environment failure reason is composed only from a fixed step label and integers (setup command index, exit status 1-255), so commands, environment values, package names, paths and process output cannot reach the reason, events, logs or responses. + +### Environment files, Skills and Artifacts + +- Environment file list tokens bind a digest of the tenant, Environment, requested directory, effective order and limit, plus a fingerprint of the full sorted regular-file path and size list and an offset that is a multiple of the limit. Every page rereads the directory; there is no cursor registry, cache or snapshot. Reject every token mismatch with the single official token message. +- The directory helper checks each requested path component with `Root.Lstat` below `os.OpenRoot(workspace)` and opens the final directory with `O_NOFOLLOW`. A missing component, a regular file or a symbolic link maps to the distinct `not_directory` result, which the daemon and gateway carry only for directory reads; Core turns it into an empty page. `not_found` (Claude SDK adapter reader), permission, transport and uncertain results keep their errors. +- A Files.create write intent stores a digest of the path, size and content, not a path ledger, so Core cannot tell a file an earlier Files.create wrote from any other file; an existing regular file therefore gets the untracked-file message. Reserve the intent under the Session lock before dispatch. The daemon verifies the complete body's SHA-256 before calling the writer, and the writer creates parents with `Root.MkdirAll(0700)`, writes `.oac-write-` in the workspace root and publishes it with `Root.Link`, which never replaces an existing entry. Known refusals return `write_rejected` with `reason` `destination_directory` or `unsafe_destination`; Core settles the intent as `rejected`, which leaves no committed receipt and releases the mutation owner. Only an exact committed or rejected receipt settles an intent; nothing settles an unknown one automatically. +- Check the 5 MiB inline bound after path validation and Base64 decoding, and before the pending-hosted check, the source File lookup and execution. The JSON body limit still admits the Base64 form of 50 MiB so that oversized inline bodies up to that size get the official message. +- Serialize Skill version uploads, default changes and version deletion on the owning Skill row lock. Deleting the default version deletes the Skill only when no other version row exists, through the same cascade as Skill deletion, so every encrypted version row goes in the same commit. `next_version` only increases. +- Decide Artifact republication in the Turn's terminal transaction, not the capture transaction: capture commits and releases the Session lock before the Turn completes, and the terminal transaction holds the Session lock that also orders Artifact deletion. Drop staged rows whose SHA-256 equals the newest remaining published Artifact for the path, ordered by the producing Turn's creation time and then ID (publication time can come from the Runtime and does not order Turns), and unlink their large objects in the same transaction. Published rows are never modified. +- A malformed `environment_id` Artifact filter resolves to the never-assigned maximum UUID, so it matches nothing without a text comparison. + +### Credential encryption and OAuth refresh bounds + +- `credentialcrypto` ciphertext is a format version byte followed by the standard AEAD nonce, ciphertext and tag. The authenticated data holds a fixed domain and version plus the binding (tenant, Vault, Credential, auth type, exact destination). Keep the domain string unchanged: existing rows must still decrypt. +- Random-nonce GCM allows at most 2^32 encryptions per key. `secrets/credential.key` also seals model providers, the E2B key, Skills, initial files and environment setup, so every sealed write counts toward that bound; there is no rotation or re-encryption path. +- OAuth dispatch refresh holds the Credential row lock and the external exchange under one 20-second context (`store.oauthRefreshTimeout`). The refresh HTTP client has a 10-second overall timeout and 5-second TLS handshake and response-header timeouts, uses no proxy and treats any redirect as failure. + +### Core administration errors, metrics and write provenance + +- Core error details are scoped by the `/core/v1` router's writer mark, never by a request path test. Write Core errors with `writeCoreError` and typed `CoreErrorDetails` values (string, number, boolean, null and string-array constructors); invalid or empty details are omitted as a whole. The mark preserves error observation, flushing and `http.ResponseController` access. A shared handler or a Core-looking path alone never changes a public or machine error envelope. Core authentication runs before operation configuration checks, and unknown paths keep their status and admission rules. When adding a code with details, document its fixed keys in `contracts/agents-api/core-errors.md`, and pass only safe Core-owned facts: never submitted values, secrets, native text or provider bodies. +- Operation validators keep their original error text, sentinel identity and validation precedence. Package-owned typed errors carry fixed field metadata; only the marked Core error mapper translates it into operation codes and safe bound or catalog details. Keep Project and key rune limits separate from node byte limits. Sandbox validation metadata travels through its store wrapper without changing transaction or provider authority. Public Session provider validation stays byte-for-byte unchanged; cover it with handler-level golden responses. The Core clients ignore malformed optional details and never retry a write. +- Core metrics instrument the existing worker and job owners without changing scheduling, lease or retention behavior. Count `execution_unavailable` at the HTTP error writer, once per rejected response; never capture request or response bodies and never infer the count from other 503s or failed Turns. Process CPU, RSS and cgroup limits are sampled by the 30-second Core metrics loop into the same bounded in-memory ring; the first CPU interval and restart gaps stay null, and host usage never substitutes for process usage. Root Turn history is queried read-only from PostgreSQL with native timestamps. Builds inject the source commit with `-ldflags` into `main.buildRevision`. Keep the response shape aligned with `packages/agents-client/src/core-metrics.ts`. +- Public resource writes carry the authenticated key's provenance separately from the execution principal. Record the operation and any creation ownership in the business transaction, never in response middleware or an asynchronous queue; a failed record rolls back the write. Internal lifecycle and refresh work never acquires public provenance, and retries never replace ownership. Environment uploads persist the safe request origin before dispatch and record success with the confirmed Runtime receipt, not the native filesystem call. Never put payloads, paths or secrets in audit metadata. Do not confuse key identity with the Session creator identity used for retries. +- Administrator writes reuse the public resource deletion and serialization code and record their administrator audit entry in the same transaction. diff --git a/services/core/README.md b/services/core/README.md index 0c1eea029..1720819af 100644 --- a/services/core/README.md +++ b/services/core/README.md @@ -25,7 +25,7 @@ Parsar product execution and its eventual public-client cutover are separate. Static-bearer and OAuth Vault Credentials support creation, replacement, deletion and safe metadata retrieval/listing. Vault deletion atomically removes its Credentials. Configure their independent encryption key and authenticated Session -use through the [credential guide](credentials.md); see [OAuth credentials](oauth-credentials.md) +use through the [credential guide](../../contracts/agents-api/vaults.md); see [OAuth credentials](../../contracts/agents-api/vaults.md#oauth) for application authorization, dispatch-time refresh and revocation boundaries. The pinned Python client saves an Agent independently, then starts a Session with initial input: @@ -53,7 +53,7 @@ Source and Session metadata stay separate; execution never looks up the source a - List Sessions with `client.beta.agents.sessions.list(agent_id=agent.id)`. Filtering uses the immutable root ID, including inline Agents and history after source changes. -See the [configuration and retry limits](../../contracts/agents-api/README.md#public-semantics) +See the [configuration and retry limits](../../contracts/agents-api/wire-semantics.md#sessions) before relying on optional settings or hosted error/default equivalence. ## Build standalone binaries @@ -99,7 +99,7 @@ including across key rotation. Keys in the same Project share the creator princi so rotating a key preserves that retry identity. Records with a known creator but no recorded request intent retain resolved-snapshot behavior; records without a creator cannot be retried. These retry policies are not verified hosted semantics. See the -[retry boundary](../../contracts/agents-api/README.md#public-semantics). +[retry boundary](../../contracts/agents-api/wire-semantics.md#creation-retries). The Store uses internal creation keys and preserves immutable engine/configuration, native continuity and same-tenant device bindings. The public API applies schema @@ -176,7 +176,7 @@ node transport contracts are unchanged. `OAC_ADDR` defaults to `127.0.0.1:8091`; use a TLS reverse proxy for remote access. `OAC_DEFAULT_HARNESS` defaults to `codex`; use `claude_sdk` for Claude Code or `mcode` for MiniMax Code. Configure the corresponding qualified Runtime through -its [deployment guide](../../contracts/agents-api/README.md#public-engine-profiles). +its deployment guide: [Codex](deploy/codex/README.md), [Claude Code](deploy/claude/README.md) or [MiniMax Code](deploy/mcode/README.md). It selects new Sessions independently of the requested model. Existing Sessions retain their stored engine. Set `OAC_HARNESSES=codex,claude_sdk,mcode` to explicitly enable installed profiles @@ -202,12 +202,12 @@ resources); general Files routes do not. Supported operations include: deletion; see [source Files](../../contracts/agents-api/source-files.md). - Vault create/retrieve/list/delete, project-scoped pagination and stored status filtering; static-bearer and OAuth Credential create/retrieve/list/replacement/delete, - plus [dispatch-time OAuth refresh](oauth-credentials.md). Public archive semantics remain gaps. Already-delivered credentials + plus [dispatch-time OAuth refresh](../../contracts/agents-api/vaults.md#oauth-refresh). Public archive semantics remain gaps. Already-delivered credentials are not withdrawn by local deletion. Session attachments support - [authenticated HTTPS MCP](credentials.md#use-a-credential-in-a-session). + [authenticated HTTPS MCP](../../contracts/agents-api/vaults.md#credential-selection-in-a-session). Execution uses the selected -[engine profile](../../contracts/agents-api/README.md#public-engine-profiles), +[engine profile](../../contracts/agents-api/harness-capabilities.md), including `none` and the colocated self-hosted profile described below. Ordinary JSON requests have a 1 MiB body limit; file transfers use the separate bounds in the Files contracts. Session lists support `after`, `limit` (0 is treated @@ -249,7 +249,7 @@ it executable. Unsupported requests fail explicitly. `/healthz` reports liveness The basic `openai_hosted` profiles for Codex, Claude Code and MiniMax Code require explicit operator configuration. Select the qualified native image using the -[engine profile guides](../../contracts/agents-api/README.md#public-engine-profiles), +[Codex](deploy/codex/README.md), [Claude Code](deploy/claude/README.md) or [MiniMax Code](deploy/mcode/README.md) guides, then follow the [Docker setup](deploy/codex/README.md#standalone-operator-configuration). Core-managed hosting supports deployment-selected E2B, Docker or microsandbox; see the [nodes guide](../../docs/getting-started/nodes.md). For the separate @@ -428,10 +428,10 @@ created/live events. Open a GET event stream before submitting later work, or use the official `sessions.stream` helper for one Turn. Function handlers return results through the same public events endpoint. Recover missed output with Session/Turn/Items queries; reconnecting SSE does not replay history. -See [creation streaming](../../contracts/agents-api/README.md#session-creation-streaming) +See [creation streaming](../../contracts/agents-api/sessions-events.md#creation-streaming) for retry behavior and unverified hosted timing. -The [accepted workflows](../../contracts/agents-api/README.md#acceptance-evidence-and-remaining-scope) +The [accepted workflows](../../contracts/agents-api/harness-capabilities.md) include real MiniMax execution through built API/daemon/Codex and Claude SDK, function success/error, cancellation and native continuation. Controlled fixtures remain useful but do not replace real-provider acceptance for execution changes. @@ -583,7 +583,7 @@ selects a unique matching credential, or stays anonymous if none matches. Ambigu fails before Session creation with 409 `conflict_error`. Selection is frozen privately; Session reads and events show an implicitly selected credential ID in the public tool, while the stored request keeps the caller's field. See -[credential setup and limits](credentials.md). +[credential setup and limits](../../contracts/agents-api/vaults.md). Authenticated execution additionally requires `mcp_http_bearer_auth`; missing keys or failed authorization/decryption never fall back to anonymous execution. diff --git a/services/core/credentials.md b/services/core/credentials.md deleted file mode 100644 index b7044e402..000000000 --- a/services/core/credentials.md +++ /dev/null @@ -1,241 +0,0 @@ -# Vault credential storage - -The standalone service supports static-bearer and OAuth Credential creation, token replacement, -deletion and safe metadata retrieval through the pinned official SDK. It stores tokens as authenticated -ciphertext in its own PostgreSQL database. There is no product-service dependency, -public secret-read endpoint. Resource creation, replacement and retrieval do not contact the -configured MCP destination; an attached Session can use it during execution. - -## Configure the storage key - -`OAC_CREDENTIAL_KEY_FILE` points to a file containing one base64-encoded, -random 32-byte key. Generate it once in private service configuration; the following -command refuses to replace an existing file: - -```sh -( - umask 077 - set -C - mkdir -p "$HOME/.oac/core" - openssl rand -base64 32 > "$HOME/.oac/core/credential.key" -) -export OAC_CREDENTIAL_KEY_FILE="$HOME/.oac/core/credential.key" -``` - -Keep the same key across service restarts and retain a protected backup separately -from database backups. The service reads it at startup; it never generates a -replacement or falls back to `PARSAR_MASTER_KEY`. Invalid configured files fail -startup with a safe error. If the setting is absent, other resources and Credential -metadata reads continue working, but Credential creation and replacement return local -`503 credential_storage_unavailable` before writing. -Session attachment/selection also uses safe metadata. If a selected credential -cannot be decrypted at dispatch, execution fails without contacting its MCP server -or falling back to anonymous authentication. - -Losing or replacing the key prevents decryption of existing credentials. Metadata -reads do not decrypt tokens and therefore do not prove that a key can recover them. -This release supports one retained key; storage-key rotation and re-encryption are not -implemented. Go's random-nonce GCM requires no more than 2^32 encryptions per key; -stop new credential writes before that bound until a supported storage-key rotation process -is available. Never treat editing the key file as rotation. - -## Public resource contract - -```python -credential = client.beta.agents.vaults.credentials.create( - vault.id, - name="Internal MCP", - auth={ - "type": "static_bearer", - "mcp_server_url": "https://mcp.example.com/endpoint", - "token": token_from_private_configuration, - }, -) -metadata = client.beta.agents.vaults.credentials.retrieve( - credential.id, vault_id=vault.id, -) -for metadata in client.beta.agents.vaults.credentials.list(vault.id, order="asc"): - print(metadata.id, metadata.name, metadata.auth.mcp_server_url) -``` - -These operations use ordinary project authentication and `OpenAI-Beta: agents=v1`. -Users and service accounts in the same project share access; a foreign project or -wrong owning Vault cannot retrieve the Credential. Parsar approval and personal -credential policies belong in the product client. - -Listing supports `after`, creation order (default `desc`), a limit defaulting to -20 and clamped to 1–100, and scalar or array `status` filters (`active`/`archived`, -both by default). Credential classification is private and independent of Vault -classification. Listing requires no encryption key and reads only safe metadata; -the parent and cursor must belong to the requested project and Vault. Synthetic -archived fixtures verify filtering, not a public archive operation. Archive behavior -and exact hosted query/concurrent-page semantics remain separate gaps. - -Required name is trimmed to 1–256 UTF-8 bytes. Required `auth` accepts -`static_bearer` or the [OAuth variant](oauth-credentials.md). Static auth requires an HTTPS `mcp_server_url` and a string `token`. The token is -preserved as opaque, nonempty data; whitespace is not trimmed. An explicitly -empty token is rejected before storage or replacement. This does not -verify that it will authenticate to a destination. The local URL profile excludes -userinfo and fragments, preserves queries and performs no DNS or HTTP request. -Official empty-token create/update rejection was observed directly. Other hosted -URL normalization rules remain unverified. - -The response contains `id`, `vault_id`, `name`, `object: vault.credential`, -`created_at`, `updated_at` and `auth`. Static auth contains only `type` and -`mcp_server_url`. There is no token, ciphertext or key information in the response. -The existing 1 MiB body bound is a local implementation limit. Creation makes a -fresh resource; hosted retry/idempotency semantics remain unverified. - -## Encryption boundary and remaining work - -The implementation uses standard-library AES-256-GCM with random nonces, without -custom nonce generation or password-derived keys. Ciphertext is a format version -byte followed by the standard AEAD nonce/ciphertext/tag payload. Authenticated data -contains a fixed domain/version and the tenant, Vault, Credential, auth type and -exact destination. A wrong key, modified payload or substituted binding fails -authentication. Names are public mutable metadata and are not part of this binding. -Resource SQL reads select no secret ciphertext. The key and request token exist in -trusted service memory; this protects stored secrets, not a compromised service host. - -See [OAuth credentials](oauth-credentials.md) for grant storage, dispatch-time refresh, -replacement and provider-revocation boundaries. Credential archive behavior, -restricted-key scopes and storage-key rotation remain separate gaps. The foreign key preserves Vault ownership and -atomically removes all dependent Credentials when their Vault is deleted. The full -protocol target is unchanged. - -## Use a credential in a Session - -Attach the owning Vault and declare the same exact HTTPS destination: - -```python -session = client.beta.agents.sessions.create( - agent={ - "model": model, - "tools": [{ - "type": "mcp", - "server_label": "internal", - "transport": {"type": "http", "server_url": "https://mcp.example.com/endpoint"}, - "connection_origin": "service", - "credential_id": credential.id, - }], - }, - environment={"type": "none"}, - vault_ids=[vault.id], -) -``` - -Trusted service-side Codex and Claude SDK support `environment:none`; Codex also -supports the documented `self_hosted` combination. The daemon must advertise both -`mcp_http_tools` and `mcp_http_bearer_auth`. The usual -[MCP profile limits](README.md#http-mcp-execution) still apply. Without an explicit -`credential_id`, one exact-URL static or OAuth credential among attached Vaults is selected; -zero matches remains anonymous. `connection_origin` may be omitted or null; it is -saved as `"service"`. Saving a reference on an Agent does not authorize it for a -Session. Selection failures use the official messages, with a null `param`, after -the input requirement and before anything is written: - -| Case | Response | -| --- | --- | -| `credential_id` without `vault_ids` | 400 `invalid_request_error`: `MCP credential_id requires an attached vault` | -| Missing, foreign-tenant, unattached or malformed reference | 400 `invalid_request_error`: `MCP credential_id was not found in an attached vault` | -| Credential in an attached Vault for another URL | 400 `invalid_request_error`: `MCP credential_id does not match server_url ` | -| Several implicit matches | 409 `conflict_error`: `multiple attached vault credentials match MCP server_url ; specify credential_id` | -| Unknown or foreign Vault in `vault_ids` | 404 `not_found_error` | - -`` and `` repeat the request's values only when they are at most 256 bytes -of printable UTF-8; otherwise the message leaves them out. Missing, foreign-tenant -and unattached references return identical responses for the same ID, so a -reference reveals nothing about Vaults the caller has not attached. - -The Session freezes its attachment list and private selection, including anonymous -decisions. Session reads, lists and event snapshots show an implicitly selected -credential ID in a null or omitted `credential_id`, also after that credential is -deleted. Anonymous tools stay null and explicit values are echoed as sent. The -stored request keeps the caller's field, so creation retries compare the original intent. -Identical creation retries recover the accepted Session before selecting again; -adding another credential does not change an existing binding. Each dispatch -rechecks the complete scope before decryption. The token goes only through the -private daemon request and a fresh native child environment variable, never public -configuration, history, arguments or logs. Native execution requires nonempty RFC -6750 b64token bytes and rejects other opaque stored strings without trimming them. -Exact hosted matching, selection timing and redirect semantics remain unverified. - -## Replace a stored token - -Use the pinned auth-only update operation when the MCP server's bearer token changes: - -```python -updated = client.beta.agents.vaults.credentials.update( - credential.id, - vault_id=vault.id, - auth={"type": "static_bearer", "token": replacement_from_private_configuration}, -) -``` - -This calls `POST /v1/vaults/{vault_id}/credentials/{credential_id}`. Both `auth` and -its `type`/string `token` are required; null, missing fields and extra mutation -fields are rejected. Empty and whitespace tokens remain opaque stored values, -subject to the existing native execution syntax limit when used. The response is -the same safe Credential metadata. ID, Vault, name, auth type, exact destination -and creation time stay unchanged; only ciphertext and update time are replaced -atomically. Missing encryption configuration or a failed mutation leaves the old -row intact. Unknown, foreign, wrong-Vault and malformed references use local 404. - -Existing explicit and implicit Session bindings keep the same selected identity -and creation retry behavior. A later dispatch reads the replacement after commit; -an already-resolved or running request may still hold the previous token. Updating -this resource performs no MCP call, changes no server-side token independently, -and provides no in-flight revocation, hot reload or cancellation. Coordinate the -destination's token change operationally. Replacing this token does not rotate the -storage encryption key or reset its encryption budget. OAuth partial replacement -is described in [OAuth credentials](oauth-credentials.md). Exact hosted -overlapping-update, replay and timestamp semantics remain unverified. - -## Delete a stored credential - -```python -deleted = client.beta.agents.vaults.credentials.delete( - credential.id, - vault_id=vault.id, -) -``` - -The response confirms the ID, `deleted: true` and `object: vault.credential.deleted`. -This operation removes the owned database row and encrypted token without loading -the storage key. Local retrieval, update and repeated deletion return 404 afterward; -listings omit the row. The parent Vault and other credentials remain available. - -Existing Session snapshots and history keep their frozen credential identity. A -subsequent secret lookup fails without selecting another credential or switching -to anonymous MCP. A token already read before deletion may remain available to -dispatched work. Deletion does not stop running Sessions or revoke tokens at their -providers; use Session cancellation and provider management for those operations. - -The relationship between deletion and archived status, exact hosted metadata -visibility and duplicate-deletion errors remain unverified. This implementation -does not infer an archive transition. Row removal is not evidence of physical -erasure from PostgreSQL pages, WAL, backups or native history. - -## Delete a Vault and its credentials - -```python -deleted = client.beta.agents.vaults.delete(vault.id) -``` - -The response contains `id`, `deleted: true` and `object: vault.deleted`. Deletion -removes the project-owned Vault and every stored Credential in one database -transaction, including active and archived classifications. It needs no storage -key and sends no provider requests. Other Vaults and their Credentials are unchanged. - -Local retrieval, repeated deletion, child reads/updates/listing and new Session -attachments return 404 after removal. New child creation also returns 404 when -credential writes are configured; the existing missing-key 503 still takes -precedence when writes are disabled. Vault lists omit the deleted parent. -Existing Session snapshots and recorded creation retries keep the original Vault -and Credential IDs, and historical Items remain available. Later secret lookup -fails without selecting another Credential from an attached Vault or downgrading -to anonymous MCP. Already-read tokens may remain in dispatched work. - -This operation follows the same provider revocation, running-Session and physical -erasure limits as single-Credential deletion. Exact hosted archive relationships, -post-delete visibility and overlapping-mutation/error semantics remain unverified. diff --git a/services/core/oauth-credentials.md b/services/core/oauth-credentials.md deleted file mode 100644 index 8581b2024..000000000 --- a/services/core/oauth-credentials.md +++ /dev/null @@ -1,170 +0,0 @@ -# OAuth MCP credentials - -Core implements the pinned `mcp_oauth` Credential variant alongside -`static_bearer`. The application obtains authorization and consent from its OAuth -provider, then stores the resulting grant through the ordinary Vault API. Core -has no authorization redirect, callback, provider-revocation or public refresh -endpoint. These responsibilities follow the [official Vault guide](https://developers.openai.com/api/docs/guides/agents-api/tools/vaults). -The protocol pin in `contracts/agents-api/upstream.json` remains authoritative. - -## Store and use an authorized grant - -```python -credential = client.beta.agents.vaults.credentials.create( - vault.id, - name="Authorized MCP account", - auth={ - "type": "mcp_oauth", - "mcp_server_url": mcp_url, - "access_token": grant.access_token, - "expires_at": grant.expires_at, - "refresh": { - "client_id": oauth_client_id, - "refresh_token": grant.refresh_token, - "token_endpoint": issuer_token_endpoint, - "token_endpoint_auth": { - "type": "client_secret_basic", - "client_secret": oauth_client_secret, - }, - }, - }, -) -``` - -`token_endpoint_auth` accepts `none`, `client_secret_basic`, or -`client_secret_post`. The latter two require a write-only client secret. -`refresh.resource` and `refresh.scope` are optional nullable strings. Expiry is an -optional nullable RFC 3339 timestamp. Without refresh configuration Core can use a -known-valid or unknown-expiry access token, but rejects an expired token. - -Attach the Vault and reference the exact HTTPS MCP destination as described in -[credential storage](credentials.md#use-a-credential-in-a-session). Static and -OAuth credentials share tenant checks, frozen selection and the existing bearer -Runtime contract. A unique implicit selection searches both authentication types; -ambiguity fails rather than preferring either type. Only existing qualified MCP -harness/placement profiles are supported. OAuth is not a new harness capability: -adapters receive the access token, never refresh tokens or client secrets. - -Resource reads and lists select safe metadata without decryption. OAuth metadata -contains expiry and refresh client ID, endpoint, authentication method, resource -and scope. Access tokens, refresh tokens and client secrets never appear there. -Creation and manual replacement do not contact the OAuth provider or MCP server. - -## Refresh and replace - -At execution dispatch, Core rechecks the selected tenant, attached Vault, -Credential identity, auth type and exact MCP URL. A known-expired grant is refreshed -before native execution. Core uses the existing OAuth library with the declared -endpoint authentication method and stored scope/resource. Refresh is a single -bounded exchange, without authentication-method probing, redirects or automatic -401 retries. Provider error bodies are not returned or logged. - -A PostgreSQL row lock serializes refresh and manual replacement with Credential -and Vault deletion. The complete grant is authenticated before use. Refreshed -access token, expiry and an optional replacement refresh token are committed -atomically before an access token is returned. An omitted refresh token retains -the previous value. A failed exchange or commit never returns the new token or -falls back to anonymous access. A provider may rotate its grant before a local -commit fails; that uncertain outcome can require application reauthorization. -Core does not retry an uncertain external grant exchange to hide it. - -Use the pinned Credential update operation to replace grant material: - -```python -client.beta.agents.vaults.credentials.update( - credential.id, - vault_id=vault.id, - auth={"type": "mcp_oauth", "access_token": new_token}, -) -``` - -Explicitly empty access tokens are rejected at creation and replacement. A -replacement must include a mutable grant field; a type-only or otherwise empty -patch is rejected before reading or changing secret material. Existing optional -null semantics below remain qualified separately. - -Identity, name, auth type, destination, refresh endpoint/client ID/resource and -endpoint authentication method remain unchanged. A refresh configuration cannot -be added to a Credential that was created without one. - -| Update field | Omitted | Explicit null | -| --- | --- | --- | -| `access_token` | Keep | Keep (local interpretation) | -| `expires_at` | Keep, or clear when a new access token is supplied | Clear | -| `refresh` | Keep | Keep (local interpretation) | -| `refresh.refresh_token` | Keep | Keep | -| `refresh.scope` | Keep | Stop sending scope | -| `refresh.token_endpoint_auth` | Keep | Keep (local interpretation) | -| `refresh.token_endpoint_auth.client_secret` | Keep | Keep | - -Supplying `token_endpoint_auth` during replacement requires the existing -`client_secret_basic` or `client_secret_post` method. Read the Credential again to -observe updated safe metadata. Whole-object null rules marked above are not proven -hosted behavior. Exact refresh timing, retry/error equivalence and unspecified -field-edge behavior are not complete protocol qualification. - -## Provider network policy - -Refresh endpoints must be HTTPS. By default Core rejects nonpublic destinations, -including loopback, private, link-local and shared-address ranges. Resolution is -checked before dialing the resolved IP, retaining TLS hostname verification and -preventing a second DNS lookup from changing the destination. Ambient HTTP proxies -are not used for grant exchange. - -For an operator-controlled private issuer, configure exact HTTPS origins: - -```sh -export OAC_OAUTH_TRUSTED_ORIGINS='https://issuer.internal:8443' -``` - -The comma-separated list is server configuration, not a tenant parameter. It -allows private addresses only for those exact host/port origins; it does not -allow HTTP, redirects or invalid certificates. Install the issuer CA in the -service trust store (for example with Go's `SSL_CERT_FILE`). An invalid explicit -origin fails service startup. Trust only destinations authorized to receive tenant -grants; this setting is not an unrestricted network bypass. - -## Revocation and limits - -Deleting the stored Credential or Vault blocks later Core lookups while preserving -Session history and frozen selection. It does not revoke the provider grant, -withdraw an access token already delivered to a native process, or cancel running -Sessions. The application performs provider revocation and cancellation when -required. Revoked or invalid refresh grants fail closed until replaced; they do -not cause automatic resource deletion or credential reselection. - -This dispatch-time refresh does not promise mid-turn hot replacement, immediate -provider revocation detection, native 401 recovery or transparent replay. Unknown -expiry does not trigger proactive refresh. Storage-key rotation, archive lifecycle -and arbitrary provider/harness combinations remain separate work. - -Keycloak is used only by private real-acceptance infrastructure. No production -code depends on its realm, endpoints, token format or administrative APIs. - -## Acceptance scope - -The 2026-09-22 credential lifecycle batch uses a standalone Core database, genuine -Keycloak 26.7.4 authorization-code grants with S256 PKCE, a TLS MCP server checking -issuer/audience/expiry through provider introspection, and real Kimi model calls. -Keycloak's public, Basic and POST client authentication flows exercise consent, -refresh-token rotation, rejected reuse, revocation and code/PKCE rejection. All -three grant variants also exercise fixed-SDK/raw-HTTP resource operations. - -Codex with Basic client authentication exercises initial MCP access, dispatch -refresh, another refresh after Core restart, manual replacement, refusal after -refresh-grant revocation, reauthorization and refusal after Credential deletion. -Claude with POST client authentication exercises initial access and refresh -through the same Core path. Revocation/deletion failures are checked before new -native or MCP work. These are the existing trusted `environment:none` public -MCP profiles; no new hosted/user-managed/MiniMax profile is qualified. - -Refresh acceptance explicitly moves the stored declared expiry into the past; -it does not claim waiting for the provider JWT to expire. Reads, public histories -and owned logs are checked for known grant/model secrets. Controlled PostgreSQL -concurrency and failed-commit tests supplement these live checks. The browser -console reads mixed static/OAuth metadata and permits deletion; application-owned -OAuth authorization/replacement is performed through the public API, while its -existing static-token replacement UI remains static-only. - -This evidence does not qualify Google, GitHub or arbitrary OAuth services, -provider-independent error equivalence, or the complete Agents API protocol. diff --git a/services/core/tests/official_list_query.py b/services/core/tests/official_list_query.py index 21d79b18f..876023985 100644 --- a/services/core/tests/official_list_query.py +++ b/services/core/tests/official_list_query.py @@ -2,7 +2,7 @@ Owned fixtures use real Worker admission with dispatch paused. No native executor or model runs, and initial Turn/Item history remains visible throughout the test. -The tolerance and cursor checks follow contracts/agents-api/list-query-semantics.md. +The tolerance and cursor checks follow contracts/agents-api/wire-semantics.md#lists. """ import importlib.metadata