diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index c03c5baaa..f0cd6ead7 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -17,9 +17,9 @@ This guide owns how to work in the repository: documentation ownership, the repo | Environment ownership and capability preparation (Skills, Plugins, MCP, `packages.system`) | [Environments](contracts/agents-api/environments.md) | | Built-in Harness identifiers, configuration/profile bindings and display names | `internal/harnessconfig/builtin/catalog.json` and its [generated reference](contracts/agents-api/harness-catalog.md) | | Effective MCP bindings and credential authority | [Environment MCP](contracts/agents-api/environments.md#skills-plugins-and-environment-mcp) and `apps/daemon/internal/agent/mcp_binding.go` | -| Harness qualification and acceptance | [Harness integration](contracts/agents-api/harnesses.md) | +| Harness qualification and acceptance | [Harness capabilities](contracts/agents-api/harness-capabilities.md) and [Harness onboarding](contracts/agents-api/harness-onboarding.md#qualify-the-adapter) | | Harness service qualification declarations and registration | [Explicit service qualification](contracts/agents-api/harness-onboarding.md#explicit-service-qualification) and `services/core/internal/engine/profile.go` | -| Harness selection and Agent defaults | [Harness selection](contracts/agents-api/harness-selection.md) | +| Harness selection and Agent defaults | [Harness selection](contracts/agents-api/model-execution.md#harness-selection) | | Provider registration validation | [Sandbox Provider guide](docs/sandbox-provider.md#registration-validation) | | Provider selection, sandbox deployment and E2B setup | [Sandbox deployment](contracts/agents-api/sandbox-deployment.md) | | Hosted sandbox nodes | [Nodes guide](docs/getting-started/nodes.md) and [sandbox deployment contract](contracts/agents-api/sandbox-deployment.md) | diff --git a/README.md b/README.md index fef82815b..3ab763d42 100644 --- a/README.md +++ b/README.md @@ -72,7 +72,7 @@ the [architecture guide](docs/architecture.md). | Build an application on the API | [Quickstart](docs/getting-started/quickstart.md), then the [Agents API guide](docs/api/public-agent-api.md) | | See a complete application | [Examples](docs/examples.md) | | Run agents on my own machine | [Self-hosted execution](docs/getting-started/self-hosted.md) | -| Check verified Runtime capabilities and limits | [Capability qualification](contracts/agents-api/environment-capabilities-qualification.md) | +| Check verified Runtime capabilities and limits | [Harness capabilities](contracts/agents-api/harness-capabilities.md) | | Understand the design | [Architecture](docs/architecture.md) | | Add a sandbox, harness or other component | [Developer guide](docs/development.md) | diff --git a/README.zh-CN.md b/README.zh-CN.md index cc2313230..7474308ea 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -68,7 +68,7 @@ Core 对外提供两组 API: | 基于 API 开发应用 | [快速开始](docs/getting-started/quickstart.md),然后看 [Agents API 指南](docs/api/public-agent-api.md) | | 看一个完整的应用 | [示例](docs/examples.md) | | 在自己的机器上运行 Agent | [自托管执行](docs/getting-started/self-hosted.md) | -| 查看 Runtime 能力和验收范围 | [能力验收记录](contracts/agents-api/environment-capabilities-qualification.md) | +| 查看 Runtime 能力和验收范围 | [Harness 能力](contracts/agents-api/harness-capabilities.md) | | 了解设计 | [架构说明](docs/architecture.md) | | 接入新的沙箱、Harness 或其他组件 | [开发指南](docs/development.md) | diff --git a/contracts/agents-api/README.md b/contracts/agents-api/README.md index 7aed59bfc..e4a59ded1 100644 --- a/contracts/agents-api/README.md +++ b/contracts/agents-api/README.md @@ -33,7 +33,7 @@ with official observations separated from Core acceptance. Adapter and persistence design follows the [design rules](../../AGENTS.md#complexity-stays-in-the-adapter). -The [harness contract and parity baseline](harnesses.md) describes equal-engine +The [harness contract and parity baseline](harness-onboarding.md) describes equal-engine registration, qualification and shared acceptance. Verify configuration against actual execution: response defaults must not merely describe values the adapter never applied. @@ -65,7 +65,7 @@ still differ. See the [accepted scope and evidence](#accepted-milestone-and-evid The same three harnesses passed historical Core-managed E2B V1 qualification in PR #705. That original route is historical evidence; it does not qualify the later user-managed enrollment or current deployment-level E2B configuration. See the -[current hosted provider contract](sandbox-deployment.md). The [user-managed V1 qualification](user-managed-runtime-v1.md) +[current hosted provider contract](sandbox-deployment.md). The [user-managed qualification](harness-capabilities.md) records separate real deployment acceptance and its exact scope. Select further work only within current user authorization. Parsar cutover and business Team orchestration are separate from protocol coverage. @@ -96,19 +96,19 @@ paths start at `/vaults`, not `/agents/vaults`. | Resource | Upstream operations | Current coverage | | --- | --- | --- | | Root reusable Agents | create, retrieve, update, list, delete | Partial create/retrieve/update/list/delete and Session references; configuration/error gaps remain | -| Skills and Versions | create, retrieve, update default, list, delete, content | [Tenant-owned encrypted bundles and hosted references](environment-templates.md); [default metadata/content and deletion evidence](file-resource-semantics.md), qualified upload limits and unresolved semantics | +| Skills and Versions | create, retrieve, update default, list, delete, content | [Tenant-owned encrypted bundles and hosted references](environments.md#skills); [default metadata/content and deletion evidence](file-resource-semantics.md), qualified upload limits and unresolved semantics | | sessions | create, retrieve, update, list, delete | Create (ordinary/live), retrieve, list with root-Agent filter, metadata-only update, [idle-only public deletion](official-semantics-alignment.md#session-deletion-lifecycle--september-23) with idempotent owner repeat and owned Docker cleanup; user-managed compute stays caller-owned; general physical cleanup and exact hosted semantics remain open | | sessions.events | create, stream | Text/cancel/function-result admission and live events; function-action state snapshots supported | | sessions.turns | retrieve, list | Implemented reads; lifecycle conformance still partial | | sessions.items | list | Partial Item variants | -| sessions.artifacts | retrieve, list, delete, content | Shared output capture and immutable stored reads/deletion on accepted Docker profiles and [qualified user-managed workflows](user-managed-runtime-v1.md) (prior Core-managed E2B evidence remains historical), including retained downloads after Runtime loss. [Aligned](official-semantics-alignment.md#artifact-capture-and-listing--september-23) output symlink skipping, unchanged-path non-republication, the list envelope and malformed filters; exact upstream defaults/errors, hard-link/special-file capture and cancellation-edge parity remain unverified | +| sessions.artifacts | retrieve, list, delete, content | Shared output capture and immutable stored reads/deletion on accepted Docker profiles and [qualified user-managed workflows](harness-capabilities.md) (prior Core-managed E2B evidence remains historical), including retained downloads after Runtime loss. [Aligned](official-semantics-alignment.md#artifact-capture-and-listing--september-23) output symlink skipping, unchanged-path non-republication, the list envelope and malformed filters; exact upstream defaults/errors, hard-link/special-file capture and cancellation-edge parity remain unverified | | sessions.subagents | retrieve, list | [Three-harness Docker reads, native lifecycle limits and real evidence](subagents.md); full multi-agent semantics remain partial | -| sessions.subagents.items | list | Qualified own-child history reads; [limit clamping and the list envelope](subagents.md#subagent-visibility--september-23-2026) aligned; full Item variants remain partial. Child work is not streamed on the Session, as observed officially | +| sessions.subagents.items | list | Qualified own-child history reads; [limit clamping and the list envelope](subagents.md#subagent-visibility) aligned; full Item variants remain partial. Child work is not streamed on the Session, as observed officially | | sessions.subagents.turns | retrieve, list | Implemented; child Turns carry the Session's Agent ID and are not Session Turns | | sessions.subagents.turns.items | list | Implemented; scoped persisted reads | | environments | retrieve | Three-harness colocated self-hosted implementation and qualified Docker hosted profiles: durable status and safe initial-file metadata; other installation inventory and full lifecycle parity remain gaps | -| environments.files | create, list | [Bounded live listing and inline/source-file creation](environment-files.md) on qualified Docker workspaces; [user-managed enrollment](user-managed-runtime-v1.md) reuses the local implementation with separate real public acceptance. [Aligned](environment-files.md#wire-alignment--september-23-2026) the 201 status, page envelope, query keys, empty pages for non-directory paths on local workspace readers, sampled path/token errors and pending hosted rejection; [aligned](environment-files.md#write-semantics--september-23-2026) parent creation, no-replacement and the 5 MiB inline bound; recursion and other errors remain partial | -| environments.templates | create, retrieve, update, list, delete | [Reusable network, files, env/setup/packages, inline/referenced Skills and Session snapshots](environment-templates.md); other initialization and full semantics remain gaps | +| environments.files | create, list | [Bounded live listing and inline/source-file creation](environment-files.md) on qualified Docker workspaces; [user-managed enrollment](harness-capabilities.md) reuses the local implementation with separate real public acceptance. [Aligned](environment-files.md#wire-alignment--september-23-2026) the 201 status, page envelope, query keys, empty pages for non-directory paths on local workspace readers, sampled path/token errors and pending hosted rejection; [aligned](environment-files.md#write-semantics--september-23-2026) parent creation, no-replacement and the 5 MiB inline bound; recursion and other errors remain partial | +| environments.templates | create, retrieve, update, list, delete | [Reusable network, files, env/setup/packages, inline/referenced Skills and Session snapshots](environments.md#templates); other initialization and full semantics remain gaps | | vaults | create, retrieve, list, delete | Create/retrieve/list/delete with independent tenant persistence, stored status filtering, atomic Credential cascade and frozen Session attachments; archive semantics and full hosted lifecycle parity remain missing | | vaults.credentials | create, retrieve, update, list, delete | Static-bearer and OAuth create/retrieve/list/replacement/deletion with scoped encrypted storage and dispatch-time refresh; Session attachment and exact-URL HTTPS MCP binding; archive semantics and full hosted lifecycle parity remain missing | @@ -195,9 +195,9 @@ user-managed enrollment remain outside this qualification. | Area | Missing or unverified scope | | --- | --- | | Subagents / multi_agent | Six reads and same-child recovery have three-harness Docker evidence; optional native operations, live child progress, full lifecycle/interactions and tool combinations remain explicit gaps | -| Environment Templates | Unsupported restricted hostname forms and exact hosted errors remain gaps. Referenced null network and capability-list selection follow [qualified inheritance rules](template-null-selection.md). Template-reference env/files/commands/packages composition follows [qualified field rules](environment-templates.md#template-and-inline-configuration-composition). CRUD/list, files, env/setup/npm/Python, inline/referenced Skills, Plugins, workspace capability directories and Session references have recorded coverage. System-package and inner-isolation evidence is historical; current `packages.system` rejects and the daemon adds no sandbox. Environment Plugin MCP transport and placement limits are [listed separately](environment-templates.md#environment-origin-mcp-plugins) | -| Input and configuration | Non-text initial input, broader content/configuration unions and reasoning/verbosity combinations; [structured output](structured-output.md) has qualified Claude function profiles on none and Core-managed Docker openai_hosted, with other combinations remaining gaps | -| Tools and interactions | [Deferred discovery qualification](tool-search.md), other tool types, effective tool-set enforcement and result/cancel publication ordering; MiniMax public functions and service-origin MCP remain unsupported | +| Environment Templates | Unsupported restricted hostname forms and exact hosted errors remain gaps. Referenced null network and capability-list selection follow [qualified inheritance rules](environments.md#inheritance). Template-reference env/files/commands/packages composition follows [qualified field rules](environments.md#inheritance). CRUD/list, files, env/setup/npm/Python, inline/referenced Skills, Plugins, workspace capability directories and Session references have recorded coverage. System-package and inner-isolation evidence is historical; current `packages.system` rejects and the daemon adds no sandbox. Environment Plugin MCP transport and placement limits are [listed separately](environments.md#plugin-mcp) | +| Input and configuration | Non-text initial input, broader content/configuration unions and reasoning/verbosity combinations; [structured output](execution-tools.md#structured-output) has qualified Claude function profiles on none and Core-managed Docker openai_hosted, with other combinations remaining gaps | +| Tools and interactions | [Deferred discovery qualification](execution-tools.md#deferred-function-discovery), other tool types, effective tool-set enforcement and result/cancel publication ordering; MiniMax public functions and service-origin MCP remain unsupported | | Vault and Credentials | Archive semantics, in-flight token withdrawal and exact hosted selection/error behavior; static/OAuth CRUD, replacement and scoped dispatch-time refresh are implemented (see credential guide for qualification) | | Existing resources | Full Item/SSE/Usage variants, omitted/null/default/error semantics, pagination and overlapping lifecycle behavior beyond recorded cases | @@ -290,7 +290,7 @@ upgrade the protocol. multi-agent settings default to six concurrent subagents. Function defer-loading defaults to false and programmatic tool calling to true. Saving these values does not itself admit a native execution. Session references are admitted separately. -- [Explicit disabled tools](tool-policy.md) can be saved, used inline or resolved from saved Agents: +- [Explicit disabled tools](execution-tools.md#web-search-and-programmatic-tool-calling) can be saved, used inline or resolved from saved Agents: `web_search.mode=disabled` and `programmatic_tool_calling.enabled=false`. Search responses include `context_size=medium` for omitted/null size, nullable domains and location; an empty domain list stays empty. Saved Agents keep every pinned @@ -434,15 +434,13 @@ upgrade the protocol. Remaining unsupported installation configuration, full hosted lifecycle and exact hosted error semantics remain gaps. -[Environment Templates](environment-templates.md) provide tenant-owned CRUD/list +[Environment Templates](environments.md#templates) provide tenant-owned CRUD/list and immutable Session resolution through common Environment preparation. The `x_agents_core.environment` extension supplies that configuration to either placement; self-hosted machines never need a managed allocation. See -[shared preparation qualification](environment-preparation-qualification.md). They do not +[shared preparation qualification](harness-capabilities.md#environment-preparation). They do not select an E2B image or make unsupported initialization executable. -The supported Docker configuration has [composed real acceptance](environment-templates.md#composed-initialization-acceptance) -across the three harnesses, including frozen source deletion, cold continuation -and cancellation. This does not close the remaining protocol/transport gaps. +[Harness capabilities](harness-capabilities.md#environment-preparation) records composed preparation qualification by Harness and placement. ## Delivery and verification @@ -461,12 +459,12 @@ and cancellation. This does not close the remaining protocol/transport gaps. ### Public engine profiles `OAC_DEFAULT_HARNESS` supplies the default for new Sessions. The optional -[Core harness extension](harness-selection.md) explicitly selects an enabled engine; +[Core harness extension](model-execution.md#harness-selection) explicitly selects an enabled engine; existing Sessions retain their immutable choice. `OAC_HARNESSES` explicitly adds installed deployment profiles without requiring a managed Provider. Model identity is independent. All three profiles require implicit reasoning and service tier `auto`. Ordinary -text output is the baseline; [structured output](structured-output.md) has a +text output is the baseline; [structured output](execution-tools.md#structured-output) has a separately qualified Claude profile. Enabled `multi_agent` qualification is tracked separately in [Subagents](subagents.md); other profiles continue to reject unsupported execution. Optional tools/configuration are qualified per operation and placement; native support is not public admission by itself. @@ -478,14 +476,14 @@ operation and placement; native support is not public admission by itself. | `mcode` | Qualified `none` text and Docker `openai_hosted`; medium verbosity; Environment-origin HTTP MCP with null/omitted allowlist and optional initialization; public functions/service-origin MCP, image input and complete public usage breakdown remain unsupported | All three profiles implement user-managed `self_hosted` enrollment at `/workspace` -through our private daemon transport; [separate real acceptance](user-managed-runtime-v1.md) +through our private daemon transport; [separate real acceptance](harness-capabilities.md) records qualified deployments and limits. A `self_hosted` Session supplies its own model provider in the request or through a saved Agent; deployment defaults apply to `openai_hosted` and `none`, never to `self_hosted` ([model execution](model-execution.md)). Service-origin HTTP MCP is rejected on `self_hosted` and hosted local placements. Explicit Environment-origin HTTP declarations use the same Runtime binding path as Plugin MCP; see the [origin matrix](environments.md#public-mcp-connection-origin) -and [public qualification](public-mcp-qualification.md). -The [Docker lifecycle](environments.md#basic-public-docker-hosted-profile) retains +and [public qualification](harness-capabilities.md#tools). +The [Docker lifecycle](environments.md#hosted-openai_hosted) retains workspace Files/Artifacts, cancellation and recovery. Managed isolation belongs to the outer Environment; native tools use the starting account's permissions. Configure immutable Runtime images explicitly: [Codex](../../services/core/deploy/codex/README.md), @@ -496,16 +494,12 @@ with the official SDK; the user owns provisioning, renewal and destruction. The shared initialization path supports env/setup and user-directory npm/Python packages; `packages.system` rejects explicitly and system dependencies must be preinstalled. -See the [evidence and limits](environment-templates.md#verification). Remaining +See the [evidence and limits](harness-capabilities.md#environment-preparation). Remaining unsupported startup installations, unqualified restricted hostname forms and hosted service-origin HTTP MCP remain outside these accepted profiles. Environment-origin -MCP Plugins have a separate [Docker qualification and transport matrix](environment-templates.md#environment-origin-mcp-plugins): -stdio on all three harnesses, Codex HTTP with literal headers or HTTPS bearer, -and Claude/MiniMax anonymous HTTP or HTTPS bearer without literal headers. This batch -does not qualify those new Plugin paths on E2B. MiniMax's private workspace MCP -bridge remains internal transport, distinct from installed Environment MCP servers. +MCP Plugin qualification and transport limits are recorded in [Harness capabilities](harness-capabilities.md#environment-preparation). -The [self-hosted profile](environments.md#initial-public-self-hosted-profile) uses +The [self-hosted profile](environments.md#self-hosted-self_hosted) uses user-managed Runtime enrollment and remains distinct from Core-managed Docker. Product `claude_code` is likewise a separate integration from the API's `claude_sdk` engine key. Unsupported configurations fail before Session creation; unsupported results fail @@ -727,7 +721,7 @@ required name, description and JSON Schema parameter object. Missing `defer_loading` resolves to `false`; null and other types are rejected. Omitted, null and empty tool lists resolve to an empty list. The resolved tools are part of the immutable Session configuration and creation retry identity. Saved-Agent -inheritance uses the same resolved tools. The bounded [deferred discovery path](tool-search.md) +inheritance uses the same resolved tools. The bounded [deferred discovery path](execution-tools.md#deferred-function-discovery) adds type-only `tool_search` for its qualified profile. Other discovery combinations, other tool kinds, the native 64-definition cap and nonblank names of at most 512 bytes remain compatibility gaps; repeated names and explicit non-object root types reject as @@ -841,7 +835,7 @@ up to the cursor read with a settled projection in one snapshot; another client' work drained before that read can still be sent. Observe later Turns with the GET event stream, which does not end on settlement or a Turn failure; it ends only after the terminal `agent.session.failed` of a hosted provisioning failure, as -officially observed ([initialization failure](environment-templates.md#initialization-failure--september-23)), +officially observed ([initialization failure](environments.md#initialization-state-and-failure)), or when the Session is deleted. Terminal Turn events carry the Turn snapshot's `usage` at the top level, null when unknown. @@ -914,7 +908,7 @@ executor-specific prerequisite does not open public Environment admission or establish complete ownership, hosted key lifecycle or error compatibility. See the [standalone configuration](../../services/core/README.md#standalone-http-service). -Core documents its optional [harness selection extension](harness-selection.md) separately from the pinned upstream contract. +Core documents its optional [harness selection extension](model-execution.md#harness-selection) separately from the pinned upstream contract. The former Core startup configuration read is removed; the [installation read](installation.md) reports the public URL and installer process settings. diff --git a/contracts/agents-api/environment-capabilities-qualification.md b/contracts/agents-api/environment-capabilities-qualification.md deleted file mode 100644 index ae7209e4c..000000000 --- a/contracts/agents-api/environment-capabilities-qualification.md +++ /dev/null @@ -1,104 +0,0 @@ -# Environment capability qualification, September 30, 2026 - -This record covers the existing Linux Runtime path. It does not qualify every -Harness, model, operating system or infrastructure provider. Current rules remain -in [Environments](environments.md), [structured output](structured-output.md) and -[deferred function discovery](tool-search.md). - -## Verified public workflows - -Real requests used the pinned official Python SDK against a dedicated Core and -PostgreSQL, the official native daemon installer, and MiniMax-M2.7 or Kimi K3. No test model -or replacement model/tool loop was used. Remote evidence is under -`zju_a100_2:~/.oac/acceptance/environment-capabilities-20260930`. - -| Workflow | Evidence | Outcome | -| --- | --- | --- | -| Self-hosted Claude structured output | `claude-structured`, Session `2a70e3ca-b67c-4222-9c3a-2fa9b7db2b9d` | Native file read followed by schema-constrained proof; another completed Turn after a fresh daemon process resumed the same Session. | -| Self-hosted Claude deferred function discovery | `claude-discovery-qualified`, Session `7801d632-697f-419d-8531-848a60d9f1a8` | Native history records `ToolSearch` and the declared deferred callback. The callback used the random schema argument and the answer contained a fresh result supplied only by the client. Cold continuation completed another callback; cancellation during a third pending callback cancelled the Turn and settled the function call. | -| MiniMax installed HTTP MCP | `minimax-http-common` | Anonymous and explicit HTTPS bearer servers both returned fresh proofs, including after daemon restart. Cancelling an actual waiting HTTP tool call cancelled its Turn. The original scan covered only `config.yaml` and `mcp.json`; later review found an extra tool-environment copy in `workspace-profile.json`, so that scan did not establish secret-free native state. | - -The HTTP fixture implements MCP, not model responses. It verifies that the -anonymous server does not receive the other server's Authorization header and -that the authenticated server receives its selected token. It does not establish -cross-origin redirect qualification for arbitrary custom headers; those remain -rejected for MiniMax and Claude. - -The structured-output continuation and the first discovery result were recovered -through durable public reads. The initial test scripts made assumptions about an -initial SSE snapshot or optional assistant phase, so their failed script records -are retained separately. They are not presented as clean end-to-end SSE evidence. -The recovered results and native history verify the actual completed work; inputs -were not replayed to obtain those results. - -## MCP credential persistence follow-up - -Independent review found that MiniMax's workspace profile copied the complete -tool environment, including selected bearer values. The adapter now persists only -the validated Runtime-owned snapshot path; the tool launcher reads it when starting -tools. Missing or malformed snapshots fail explicitly without parser contents in -errors. No historical profile or credential migration is provided. - -A newly built daemon and MiniMax companion passed `minimax-http-private-env`: -anonymous and selected HTTPS bearer calls, cold continuation and cancellation of -an active HTTP call. After each Turn, the test scanned every native MiniMax state -file for the fresh token. A final check confirmed a real workspace profile with -no inline `toolEnv`, a snapshot reference outside native state, and no token in -any native state file. The earlier narrower scan above is not reused as evidence -for this guarantee. The Runtime's authorized private snapshot intentionally -retains the environment values required by tools. - -## Funded visual-model validation - -After restoring Kimi K3 credit, real pinned official-client requests used the -same Core/Runtime and native adapters on Linux user-managed machines: - -| Workflow | Evidence | Verified behavior | -| --- | --- | --- | -| Claude image messages | `claude-image-funded`, Session `be8c8251-5c07-4c39-b71c-3ac74139bbb7` | Randomized PNG colors, a second PNG after daemon restart, and JPEG on the rebuilt final Core. | -| Codex image messages | `codex-image-funded`, Session `eed03095-b3e9-4c97-8e23-3343287c797d` | Randomized PNG colors, a second PNG after daemon restart, and JPEG on the rebuilt final Core. | -| Claude function image results | `claude-function-image-funded`, Session `fc0f373d-d5a2-41bb-9061-7c0da3997387` | PNG supplied only in the callback; remote/failed image rejection leaves the call pending; identical result retry succeeds, conflicting retry rejects, and Items retain submitted content. Cold continuation recalls the prior image, pending-call cancellation settles, and final-Core JPEG succeeds. | - -The image follow-up changes operation admission, not native encoding, credential -selection, preparation or execution ownership. It removes the Environment-source -image gate for Codex and Claude; Runtime image support is still checked before -delivery. MiniMax image input remains rejected. This record adds no new managed -provider run, macOS/Windows live model run, or Codex function-result image -qualification. No production deployment was updated; only the owned acceptance -Core was restarted, preserving its database and Sessions. - -## Runtime boundaries - -All three adapters consume the same transient effective MCP binding resolver. -Unit and adapter tests retain nil versus empty tool allowlists, required flags, -service versus Environment origin, credential authority, reserved/duplicate -identities, fixed installed stdio launchers, secret-free generated configuration, -and existing failure/cancellation settlement. Public service-origin HTTP remains -restricted to its existing service execution host. This change does not add a -Core proxy or claim public `connection_origin=environment` support. - -The daemon CLI now waits up to ten seconds for confirmed shutdown, allowing the -native process owner's three-second grace period and subsequent pipe/owner -cleanup. A real MiniMax HTTP Session reproduced the old three-second CLI timeout; -its process exited shortly afterward. The new timeout passed cold restart and -cancellation. Genuine timeout still fails and retains process ownership records. - -## Limits and retained failures - -- The first Kimi attempt returned HTTP 429/quota and was cancelled. After the - account was funded, separate image workflows below used real visual requests; - the failed quota attempt remains separate evidence. -- An earlier ToolSearch bundle did not advertise workspace discovery. Its input - remained pending without a Turn; public cancellation returned 409. That Runtime - was stopped and the records retained. The corrected bundle advertises discovery - only with the explicit workspace/function features. This does not claim a fix - for the general pending-input cancellation gap. -- These model runs use Linux user-managed environments. The common managed - preparation flow has separate evidence in - [preparation qualification](environment-preparation-qualification.md); this - record does not add a new Docker/E2B/microsandbox model run. -- No Claude installation was performed on the Mac. Native platform CI, including - Windows, is required for the final PR. Windows manual acceptance remains absent. -- Public MiniMax MCP, Plugin tool allowlists/required flags, custom HTTP headers, - workspace discovery combined with Skills/Plugins/MCP, and MiniMax image - execution are not qualified by this work. diff --git a/contracts/agents-api/environment-files.md b/contracts/agents-api/environment-files.md index 105cb6996..13ff4c5c4 100644 --- a/contracts/agents-api/environment-files.md +++ b/contracts/agents-api/environment-files.md @@ -9,7 +9,7 @@ hosted provisioning and shared Artifacts are accepted within the [recorded Docker MVP scope](README.md#accepted-milestone-and-evidence) and separate historical Core-managed [E2B qualification](README.md#e2b-v1-qualification). The new user-managed enrollment chain reuses local Files with separate -[real public acceptance](user-managed-runtime-v1.md); complete Files/Environment +[real public acceptance](harness-capabilities.md); complete Files/Environment semantics and other providers are not implied. ## Pinned contract diff --git a/contracts/agents-api/environment-preparation-qualification.md b/contracts/agents-api/environment-preparation-qualification.md deleted file mode 100644 index f2a8dccfd..000000000 --- a/contracts/agents-api/environment-preparation-qualification.md +++ /dev/null @@ -1,59 +0,0 @@ -# Shared Environment preparation qualification - -This record covers the shared preparation change based on Core -`5cac06c13eccee954575a359df9b014d728f4d14`. The current behavioral contract is -[Environments](environments.md); this record does not broaden Harness capabilities. - -## Native execution - -Every passing row used the pinned official Python SDK, a real Core and PostgreSQL, -the daemon and the native Harness against MiniMax-M2.7. Model credentials were -supplied explicitly by each Session and excluded from retained public evidence. - -| Environment | Harness | Verified | -| --- | --- | --- | -| Linux user machine | Codex 0.153.4 | Uploaded Skill script, Plugin stdio MCP, initial file, setup, two Turns and daemon restart | -| Linux user machine | Claude SDK 0.3.269 | Same workload, two Turns and daemon restart | -| Linux user machine | MiniMax Code 0.4.12 | Same workload and declared workspace, two Turns and daemon restart | -| Docker managed Linux | Codex 0.153.4 | Same preparation input and workload, two Turns and container restart | -| Native macOS arm64 | Codex 0.153.4 | Same workload, physical workspace, two Turns and daemon restart | - -Skill execution wrote a marker loaded from the installed archive; a separate MCP -call returned another marker. Tests checked the resulting workspace files and -hashes of the installed capability tree independently of model text. Editing the -Template after preparation left existing Sessions unchanged; new Sessions read the -edited setup command. No allocation was fabricated for user-managed machines. - -Additional Linux Codex runs exercised an explicit operator tool-environment file: -local defaults, a Session override, preserving the source, editing the source and -reconnecting with the original values. Separate empty-configuration and initial-file-only -Sessions also retained their original tool variables after restart. - -## Contract and regression coverage - -Focused tests cover the shared extension and Template parser, encrypted frozen -resources and tenant access, authenticated preparation without an allocation, -missing Harnesses, failed or unknown steps, revoked or stale bindings, recovery -without replay, independent progress on healthy nodes and unchanged typed client -requests. Runtime tests cover local configuration freezing, missing snapshot -rejection, capability integrity and the existing Harness adapters. - -Required repository checks and native CI results are recorded on the change's PR. -The remote browser checks use cached Chromium instead of the local Google Chrome -channel; they run the same acceptance cases. No large build runs on the Mac. - -## Limits - -These probes do not establish complete upstream parity or qualify new image, -structured-output, tool-discovery or public HTTP MCP combinations. Those require -separate adapter changes and qualification. The shared setup/package mechanisms -retain their contract tests; these live probes do not download npm/Python packages. -Windows requires its native CI job; no Windows manual execution was performed. -No Claude installation was performed on the Mac. E2B and microsandbox were not -requalified by this change. - -An initial managed probe used a cached image with retired Codex requirements; -native shell execution failed. The passing managed row used a clean image without -those requirements. An initial MiniMax probe exposed a wrong native working -directory; the passing row includes the separately merged workspace fix. Neither -failed attempt is counted as acceptance. diff --git a/contracts/agents-api/environment-templates.md b/contracts/agents-api/environment-templates.md deleted file mode 100644 index c3bad53c5..000000000 --- a/contracts/agents-api/environment-templates.md +++ /dev/null @@ -1,996 +0,0 @@ -# Environment Templates and initialization - -Core owns reusable configuration through the five pinned -[Template operations](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py). -Templates do not contain a running workspace and do not select a provider image. -An E2B `templateID:build_UUID` remains private operator packaging configuration. -Each referencing Session obtains its own Environment through the same initialization -and five-operation SandboxProvider path as inline configuration. - -## Supported batch - -- Create, retrieve, update, delete and list under `/v1/agents/environments/templates`. - Every operation requires project authentication and `OpenAI-Beta: agents=v1`. - CRUD/list works without an execution deployment. -- Optional nullable name, preserved verbatim, with a local 1–256 Unicode character - bound. Network supports `enabled`, `disabled` and exact-host `restricted`; omitted/null create network - defaults to the pinned enabled policy. These are configuration values, not daemon - network enforcement; unsupported execution combinations reject. Update omission preserves; supplied name - or network replaces, with null clearing name or resetting network. -- Empty/null installation fields retain empty defaults. Responses contain safe - metadata and never `env`, `setup_commands` or inline file data. Initial files are - supported as described below, together with inline/referenced Skills, env, ordered setup and npm/Python packages; `packages.system` rejects explicitly and system dependencies must be preinstalled. -- Listing uses `after`, `limit` (default 20; 0 is treated as 1 and larger values as - 100), and `order` (default `desc`). - Creation timestamp plus ID supplies stable local ordering. Missing/foreign IDs - and cursors return the same not-found result. No compute is allocated by CRUD. -- Session `environment_template_id` resolves under the caller's tenant. Omitted - or null network inherits; enabled can narrow to restricted or disabled. - Restricted can narrow to an exact-host subset or disabled; disabled cannot widen. Effective - configuration is frozen without passing the template ID to execution. -- Updating/deleting a template does not change existing Sessions. Creation retries - recover recorded caller intent before template lookup, including after deletion; - changed intent conflicts. This is the existing local retry policy, not a claim of - complete upstream idempotency semantics. - -```python -from openai import OpenAI - -client = OpenAI(base_url="https://your-core.example/v1", api_key="your-project-key") -template = client.beta.agents.environments.templates.create( - name="Python workspace", network={"access": "enabled"}, - env={"APP_MODE": "analysis"}, packages={"python": ["packaging==26.0"]}, - setup_commands=[{"command": "mkdir -p /workspace/outputs"}] -) -session = client.beta.agents.sessions.create( - agent={"model": "your-configured-model"}, - environment={"type": "openai_hosted", "environment_template_id": template.id}, - input="Create /workspace/outputs/report.txt containing the result of 6 * 7.", -) -# Inspect Session/Turn/Items and retrieve published Artifacts after completion. -# Delete the Session to reclaim its Environment; template deletion is independent. -``` - -## Initial files - -Both inline hosted configuration and reusable templates accept `files` entries with -an absolute destination inside `/workspace`: `inline` with standard-base64 `data`, or -`file_id` referencing a project-owned Files API upload. The guide's limits are 50 -initial files, 5 MiB per inline file, 10 MiB total inline content, and 50 MiB per -referenced file. Session/Template JSON requests allow 16 MiB for the base64 envelope. -Paths must be canonical, distinct and stay within the logical workspace. Runtime -anchors the operation to its bound workspace; this is API path scope, not a -restriction on native tools running as the same user. A failed install never starts native execution. - -Configure `OAC_CREDENTIAL_KEY_FILE` with the existing execution-service -base64 32-byte encryption key. Template writes and Session resolution need it; -ordinary metadata reads do not. Template inline metadata contains type/path/size, -while references contain type/path/file_id. Sessions receive fresh file IDs and -sizes for both variants. Initialization keeps file data out of ordinary configuration, -resource responses, lifecycle events and command arguments. Templates keep references; each Session authorizes and -freezes its own encrypted source bytes. Later source deletion cannot change them. - -Template `files` omission preserves on update; null/[] clears. In a referencing -Session, omission and null inherit, while a supplied list replaces the entire file -set, including `[]` clearing it. Paths are not merged between the two sources. -The effective list retains the existing validation and tenant-owned source checks. - -Core initializes both paths through the authenticated daemon's typed `runtime_prepare` -file operation. The Runtime resolves the logical workspace path and owns the trusted -file installer. Daemon authentication remains available, but native preparation and live -Files wait for all writes. Each file gets a two-minute transfer budget; the batch has -a thirty-minute local budget in the common Environment preparation scheduler. -These are local operational limits, not verified upstream timing. Initial input -retains its existing five-minute admission deadline; large installations can use an -idle Session; connection status alone does not establish preparation readiness. - -Uncertain writes and Core restart during initialization fail the new Environment while retaining compute ownership and files; they do not replay partial installation. After completion, reconnect and -native-history recovery preserve user modifications instead of reinstalling files. -Managed and user-owned Runtime connections use this lifecycle. The Provider -API remains five operations; public Templates are never E2B image templates and -are also available to self-hosted Sessions through `x_agents_core.environment`. E2B now uses user-managed Runtime enrollment through the official -SDK. Historical Core-managed E2B evidence below retains its original scope and does -not qualify that new chain. - -## Skills and versioned references - -Both templates and standalone preparation configuration accept project-owned Skill -references and inline Skill ZIPs. Upload a directory through the pinned SDK, then -reference its default version from a template: - -```python -skill = client.skills.create(files=[ - ("report/SKILL.md", b"---\nname: report\ndescription: Create the report.\n---\nFollow the report procedure.", "text/markdown"), -]) -template = client.beta.agents.environments.templates.create( - skills=[{"type": "skill_reference", "skill_id": skill.id}] -) -session = client.beta.agents.sessions.create( - agent={"model": "your-configured-model"}, - environment={"type": "openai_hosted", "environment_template_id": template.id}, - input="Use the report Skill.", -) -``` - -Core exposes the pinned `/v1/skills` resource, version and content operations using -ordinary project bearer authentication. No Agents beta header is required on those -resource routes. Metadata reads do not decrypt or fetch bundles. Content is encrypted -with its tenant, Skill and immutable version identity. ZIP uploads use `files` and -directory uploads use repeated `files[]`; the fixed SDK directory form above works. -SDK 3.13.0 drops a single FileTypes tuple during multipart extraction before sending -it. Use raw HTTP for a single ZIP with this fixed client; Core does not synthesize -missing bytes or alter the pinned SDK. - -Templates preserve reference selectors: omission or null selects default at Session -creation, `"latest"` selects latest, and a positive version string selects that -version. A Session freezes tenant-authorized bytes and concrete version metadata -in its creation transaction. Later source deletion, default changes or template -updates cannot change that Session or its committed creation retry. A supplied -Session Skill list replaces the template list; omission/null inherit. Template -responses include `version: null` for an unresolved default selector; resolved -Session references retain a concrete version string. See -[resource selector qualification](resource-selector-semantics.md) and [null selection evidence](template-null-selection.md). -References return type/skill_id/version/name/description in Session metadata, -while template responses retain unresolved selectors. Confidential bundle content -never appears in these metadata responses. The common Runtime installation path -receives frozen files and descriptive metadata, without source or template IDs. - -Inline Skill ZIPs use the shared Runtime capability installer: - -```python -import base64 -from pathlib import Path - -skill = { - "type": "inline", "name": "report", "description": "Create the report.", - "source": {"type": "base64", "media_type": "application/zip", - "data": base64.b64encode(Path("report.zip").read_bytes()).decode()}, -} -template = client.beta.agents.environments.templates.create(skills=[skill]) -``` - -Each archive contains one top-level folder with `SKILL.md` and optional supporting -files. The manifest name/description must match the request. Portable descriptive -frontmatter supports `name`, `description`, `license`, `compatibility` and string -`metadata`; native hooks, permission controls and subagent directives reject. -Local limits are 50 Skills, 5 MiB compressed and 20 MiB expanded per archive, -10 MiB compressed and 50 MiB expanded in total, and 1,000 entries per archive. -Regular files only: path traversal, links, duplicate destinations, special files -and invalid manifests reject. Content is inert during installation; executable -files retain their executable bit. These operational limits are not claims about -upstream limits. - -Inline metadata contains only type/name/description. Archive content stays in encrypted, -resource-bound template and Session snapshots. Updates replace supplied `skills`; -omission preserves and null/[] clears. Existing Sessions retain their frozen -content after template update/deletion. - -Core transfers frozen Skill archives over the authenticated Runtime connection. -Runtime's shared parser installs them under its operator-bound capability root -(`skills/` below the configured capability directory) -before setup and native execution. -The same daemon handles initial files, configuration, packages and setup; Providers -only bootstrap it and manage compute resources. Setup and native tools run with -the launching user's permissions; snapshot file modes are an integrity hint, not -protection from that user. Completed recovery loads the recorded installation -without reinstalling. The common Runtime resolves installed metadata and paths before -invoking the native executor factory. -Codex registers native extra roots, Claude creates its own explicit Skill plugin -envelope, and MiniMax points its native user-global catalog at the shared root. -MiniMax retains disabled unrestricted built-in tools and uses its existing -workspace tool worker with the launching user's permissions. No Provider or model/tool loop is added. - -Codex nested `SKILL.md` discovery, `agents/openai.yaml` native dependency -configuration and Claude inline/fenced shell preprocessing are not qualified in this batch and explicitly fail adapter -preparation. Other files are not interpreted as a public plugin installation. -The initial skill-only Plugin and capability-directory batch below extends installation; -environment-origin MCP has a separate qualification boundary described below. Native built-in Skill visibility -is not evidence of exact public tool-set parity. Qualification probes alone do not -establish complete public support; record real service acceptance separately. - -## Skill-only Plugins and capability directories - -A hosted -`plugins` entry uses the pinned inline ZIP source with type/name/description and -`.codex-plugin/plugin.json` inside its single archive root. The manifest declares -Skill directories with `skills`; complete package-relative resources are retained. -The same parser, template resolution and encrypted Session snapshot serve inline -and template requests. Only type/name/description appears in public Plugin metadata. -Referencing Session Plugin lists inherit on omission/null and replace when a -non-null list is supplied, including empty-list clearing. Template resource updates -continue to accept null/empty clearing. - -Hosted Template and inline `capability_directories` accept clean absolute paths -within `/workspace`; self-hosted local selections have their -[separate public input](environments.md#runtime-capability-preparation). -Initial files and setup can populate hosted directories. Runtime snapshots them after -setup and writes the shared `installed.json`, bound to the Session, -Environment and source-selection digest. Recovery loads those installed bytes even -if a source directory changes or is removed; a new Session captures its own snapshot. -Partial or conflicting installation fails rather than silently recapturing sources. -This timing and the hosted path restriction are local implementation choices, not -claims about unspecified upstream behavior. A supplied Session directory list replaces -the template list; omitted/null lists inherit. Missing, overlapping duplicate Skill -names, unsupported manifests and nonregular files reject initialization without -publishing completion. Directory-discovered Skills do not become fabricated -inline entries in public `skills` or `plugins` metadata. - -The common Runtime resolves the installed manifest during admitted asynchronous -preparation, before creating a native executor; -read-only Files operations retain their existing minimal requirements. Native -adapters consume Runtime-owned Skill and package roots. Codex uses explicit roots, -MiniMax uses its native registry and Claude uses controlled envelopes with explicit -roots into a hardlinked content tree. Native component configuration is never -passed wholesale to a harness. Provider APIs and the native execution loops are -unchanged. No new installation/recovery lifecycle or framework is introduced. - -The original accepted batch supported skill-only packages and inert empty MCP maps. -Environment-origin MCP uses the separate transport path below. Archives retain the shared 5 MiB compressed/20 MiB expanded/ -1,000-entry limits. Plugin lists are limited to 50 entries and 10 MiB compressed; -combined installed capabilities are limited to 50 Skills and 50 MiB. Limits are -implementation bounds. Native shell preprocessing, dependency activation and -nested discovery retain the existing adapter restrictions. - -Historical standalone Docker acceptance on 2026-09-21 passed with the fixed official -SDK and raw HTTP against the then-current Core/daemon/adapters on qualified native images: -Codex/Kimi (258.08s), Claude/Kimi (187.81s) and MiniMax Code/MiniMax (250.23s). -Each exercised two Plugin Skills with package-relative resources and one directory -Skill generated during setup, public Files/Artifacts and tenant rejection, retained -native history after Core/Runtime restart and source/template deletion, no repeated -initialization, cancellation with stopped effects and stable duplicate cancellation. -Operators verified nonempty private canaries and actual daemon credentials before -native scripts proved the contents unreadable; canary hashes survived cold recovery. -All owned execution resources were cleaned. The complete Core-only `make check` -passed, including dedicated real PostgreSQL checks and native packaging. - -Evidence is under `~/.parsar/remediation/20260921/template-plugins/`: -`acceptance-summary.json`, `native-final-{codex,claude,mcode}.json`, -`runtime-builds.json` and `full-check-attempt2-result.json`. The shared reproducible -fixture is `services/core/tests/official_environment_plugins.py`; operator -runners reuse existing standalone acceptance and private model configuration. -A user-authorized reused-context GPT-6 Astra high independent review of all 60 -changed files found no material actionable findings. Those historical results do -not qualify Plugin MCP or the later shared Runtime transport and self-hosted directory -preparation. Portable root `plugin.json` applicability and the unconfirmed -semantics above remain gaps; -these results do not establish complete Environment Templates or protocol compatibility. - -## Environment-origin MCP Plugins - -MCP-only and combined Skill/MCP packages use the same encrypted template/Session -snapshot and installation as Skill-only Plugins. The fixed format remains -`.codex-plugin/plugin.json` with `mcpServers: "./.mcp.json"`, or an omitted path -using the root `.mcp.json`. The file contains `mcpServers` keyed by server name. -An exact capability-directory Plugin root activates its MCP declarations; selecting -a parent directory discovers Skills without activating every nested MCP server. - -The common parser accepts HTTP `url`, `bearer_token_env_var`, literal `http_headers`, -and stdio `command`, `args`, selected `env_vars`, package-relative `cwd`. It does not -resolve credentials. Runtime reloads frozen installed packages and resolves selected -values only from initialized caller env. A missing value fails instead of using a -model or daemon variable. Public `env_http_headers` remains unsupported; an adapter -may privately use native reference fields without changing public literal values. -Caller-owned env remains available to caller code under the normal env contract. - -Stdio executes through the common daemon helper, which resolves the frozen package -and explicitly selected env/cwd, then launches the declared command. Unix replaces -the helper process; Windows forwards stdio within the owned process tree. Native -MCP still owns the protocol. Process groups and Windows Jobs handle cancellation, -not isolation. There is no Python launcher or inner mount/network sandbox. System -dependencies must be preinstalled; npm/Python dependencies may use the ordinary -Runtime initialization flow. The implemented transport boundaries are: - -| Adapter | Environment MCP implementation | -| --- | --- | -| Codex | Stdio; HTTP with literal headers and HTTPS bearer references | -| Claude Code | Stdio; anonymous HTTP or HTTPS bearer, without literal custom headers | -| MiniMax Code | Stdio; HTTP explicitly rejected pending safe native qualification | - -All current Environment MCP execution requires enabled network. Restricted/disabled -HTTP and hosted stdio under a restricted/disabled policy are not qualified. Claude -literal headers are rejected because the pinned client expands them again and -forwards custom headers across origins. MiniMax ACP does not enable the native -custom-header redirect protection. Duplicate global server identities are rejected; -official duplicate namespace and required/optional connection-failure semantics -remain unconfirmed. These are implementation limits, not changes to the upstream -protocol or claims of equal optional feature sets. - -### Docker qualification (2026-09-21) - -The fixed OpenAI SDK 3.13.0 and raw HTTP passed against independent Core, -PostgreSQL and Docker Runtimes using real Kimi K3 (Codex/Claude) and MiniMax-M2.7 -(MiniMax Code). Each stdio run exercised MCP-only, combined Skill/MCP and exact -capability-directory packages, selected env/cwd, Files/Artifacts, tenant checks -and positive private credential/history canaries. Removing mutable sources and -the template, then restarting Core and Runtime, retained installed tools and native -history without replaying setup. Public cancellation stopped owned tool effects; -duplicate cancellation remained stable. SDK/raw Items and live event ordering -were compared without replacing native failed/incomplete tool observations. - -Separate real Kimi HTTPS runs covered inline/template configuration and cold -continuation: Codex literal headers plus bearer, and Claude anonymous plus bearer. -Codex preserved same-origin credentials and rejected cross-origin redirects. -Private CA trust and hostname validation remained enabled. These runs used the -same Core/daemon/adapters before the final stdio-only launcher correction and -Codex shell hook fix; affected paths were then qualified separately on final images. - -Evidence is under `~/.parsar/remediation/20260921/template-plugin-mcp/`: - -| Profile | Passed public run | -| --- | --- | -| Codex stdio | `docker/codex/plugin-d1anzd48/result.json` | -| Claude stdio | `docker/claude/plugin-l05txyye/result.json` | -| MiniMax stdio | `docker/mcode/plugin-ttnduvtn/result.json` | -| Codex HTTPS | `final-codex-http/run-8tldu8v2/result.json` | -| Claude HTTPS | `final-claude-http/run-67b8vj1o/result.json` | - -Core binary SHA-256: `4b0b3360bdc776e715d9239c3b9564cd7760dd4cef6a42d1a5b82db565f84912`. -Daemon: `d6ac4eeb7657884ab3704bcdc74456044e8e702f2f89b2d30486f29ea5fe083c`. -Per-run records retain exact Runtime image IDs and cleanup results. The final -`make-check-final.log` passed; OpenAPI regeneration and focused Go/Claude tests -also passed. Real Linux process checks cover idle creator-thread exit, native and -wrapper exit, active-call cleanup and the reproduced pre-exec orphan window. -`services/core/deploy/runtime/initialize_stdio_test.py` retains that OS -regression; it requires a disposable Linux packaged Runtime, not a model fixture. -The Codex environment hook additionally passed 30 actual sh/Bash cases after its -Bash-specific `eval --` failed under native `/bin/sh`. - -The reusable public fixture is -`services/core/tests/official_environment_plugin_mcp.py`. Mechanism probes -and failed attempts remain separate evidence. This batch does not qualify new -Plugin MCP paths on E2B, service-origin hosted MCP, OAuth, unlisted transports or -complete upstream protocol compatibility. - -### Composed initialization acceptance - -This section records the historical September 21 fixture and binaries. Its -system-package and inner-isolation checks are not current support requirements; -`packages.system` now rejects and tools run with the starting user's permissions. - -`services/core/tests/official_environment_composition.py` combines the -existing public fixtures in one enabled-network configuration: inline/referenced -files, caller env, system/npm/Python packages, ordered setup, an uploaded Skill -reference, Skill/MCP Plugins and exact capability-directory roots. Pass the same -configuration through a Template reference or inline hosted Session; Core resolves -both through the shared initialization flow. - -Use a real standalone service, fixed official SDK/raw HTTP and native model API. -The operator supplies package-network access and creates the nonempty private -canaries required by `official_environment_plugin_mcp.py`. The uploaded Skill and -each selected MCP server must actually use the installed dependencies; compare the -complete expected Files/Artifacts, public metadata, Items and live event sequence. -The fixture is not a synthetic model or a replacement execution loop. - -After the first successful Turn, change the Skill default and Template, delete -source resources and mutable capability directories, then retry the original -Session creation and restart Core/Runtime. Verify frozen bytes/configuration, -owned native history, retained user edits and a setup count of one. Finally cancel -an MCP call with observed ongoing effects and verify terminal Items, stopped -effects, stable duplicate cancellation and owned resource cleanup. A successful -initial Turn alone does not pass this composed workflow. - -The final Docker composition passed on 2026-09-21 with SDK 3.13.0/raw HTTP, -independent Core/PostgreSQL, Kimi K3 (Codex/Claude) and MiniMax-M2.7. Core and the -unchanged Codex/MiniMax images are from main `8eb3c089`; the Claude image includes -the historical native-prefix correction recorded by that batch. All four public -runs completed with zero cleanup errors: - -| Profile | Result under `~/.parsar/remediation/20260921/template-composed-acceptance/` | -| --- | --- | -| Codex template | `docker/codex/composition-template-f19l27kz/result.json` | -| Codex inline | `docker/codex/composition-inline-dhioacmy/result.json` | -| MiniMax template | `docker/mcode/composition-template-9txsyiut/result.json` | -| Claude template | `docker/claude/composition-template-fmros4h_/result.json` | - -`acceptance-summary.json` records exact images, binary/result hashes and checks. -Claude's final image is -`sha256:287e7634d932c20edc0895d2095ae79c59cc4e82d2e74f6d4996b4d35f091458`. -Its separate real native regression, `claude-prefix-regression/run-6bd24d52`, -verified cross-call cwd, quoting/exit status, installed jq, private credential/history -isolation and restricted-network HTTP/HTTPS allows, domain/subdomain/redirect denials -and direct-IP/proxy-free denial. This supplements the public runs; it does not -replace them. The final `make-check-candidate2.log` passed, including the prefix -regression and 115 Claude SDK tests. An unchanged opt-in packaged scratch/large-output -probe was skipped; live execution checks ran separately. - -Failed attempts remain evidence. A deterministic Claude MCP startup failure led -to the prefix correction. A later workspace permission rejection on the final -image remains unexplained because its original native input was not retained; -the successful repeat used explicit foreground/sandbox instructions and did not -change production authorization or acceptance assertions. It does not establish -a fix for that separate rejection. Diagnostic images and the superseded Bash-input -prototype are excluded from final qualification. These results close the supported -Templates workflow, not all upstream semantics, transports or Provider combinations. - -## Packaged Runtime initialization contract - -Template handlers and stores resolve public configuration without choosing a -harness, native path or compute backend. The common runner uses neutral -Environment/Session identity and an authenticated Runtime peer. It sends initial -files, configure/npm/python/setup operations, Skills, Plugins and finalization -through `runtime_prepare`, then uses the same daemon for execution. Providers own -placement, creation, bootstrap, inspection, renewal and reclamation; they do not -run Core initialization commands. Harness differences stay in adapters. - -One Go preparation implementation serves Linux, macOS and Windows. Initial files -and setup working directories use logical `/workspace` addresses; physical paths -and executable selection belong to Runtime. Initial files use the shared atomic -writer, including replacement, while public Files creation retains its separate -semantics. Transfer input and stored tool configuration are bounded. Process -ownership and I/O settle before a typed receipt; unknown effects are never replayed. - -Initialization and packages default under `OAC_RUNTIME_HOME`. Operators can select -`OAC_RUNTIME_INITIALIZATION_DIRECTORY` and `OAC_RUNTIME_PACKAGE_DIRECTORY`; -packaged Linux images set `/environment/initialization` and `/environment/packages`. -`OAC_RUNTIME_TOOL_ENV_FILE` selects explicit tool-variable JSON. Runtime does not -overwrite an existing configuration. No Python initialization wrapper, system-root -seed or managed shell hook is shipped. - -Setup uses Bash, with Git Bash on Windows; a missing dependency fails explicitly -rather than substituting a different shell. npm uses a local prefix and Python/pip -a local target under the Runtime package directory. Node/npm and Python/pip must -already be installed. No automatic system-package installation, sudo or daemon -privilege increase occurs. - -All commands use the launching user's permissions and host network. The daemon -provides no filesystem, permission or network sandbox, including on Linux. Managed -isolation belongs to the outer Environment. Configured user variables are applied -to the command, not automatically inherited from daemon credentials, but tools may -read any local state the same user can read. Output is discarded rather than -included in failure diagnostics. Only a confirmed command failure can report its -bounded integer exit status; cancellation or unknown effects retain a generic -reason. Files reads retain their separate authorization and readiness requirements. - -## Initialization failure — September 23 - -The September 23 observations below are historical acceptance evidence. Current -initialization uses the shared daemon protocol described above; these observations -do not qualify that new transport against a real deployment. - -When a hosted Environment fails to provision, Core now reports it the way the -official service does (evidence and rows H1–H8 in -[official semantics](official-semantics-alignment.md#hosted-initialization-failure--september-23)). -In the allocation's cleanup transaction, Core marks the Environment failed and -records, in order, `agent.session.environment.failed`, an `error` event and one -`agent.session.failed`. The Session reads `failed` with the safe reason as `error` -and the failure time as `last_active_at`; live streams end after the failed event. -New input on that Session returns 409 `conflict_error` "the hosted environment -failed to provision". Pending input settles as failed exactly as before. - -The reason names only the failed step and its exit status: - -| Step | Reason | -| --- | --- | -| Setup command `i` | `Failed to provision environment: script "setup_commands[i]" failed with exit code N` (observed) | -| Python packages | `... script "Python package installation" failed with exit code N` (observed label; the official reason appends raw pip output, Core never does) | -| npm packages | `... script "npm package installation" failed with exit code N` (unverified label; system packages now reject before initialization) | -| Initial file write or Runtime Skill preparation with a confirmed failure | `Failed to provision environment: initial file installation failed` / `Skill installation failed` (unverified) | -| Anything else | `Failed to provision environment: initialization did not complete` | - -"Anything else" covers timeouts, the thirty-minute budget, unknown effects, -missing or malformed receipts, Plugin installation and directory finalization, -bootstrap rejection and Core restart during initialization. All initialization -operations return typed Runtime `rejected`, `failed` or `unknown` outcomes. The -daemon confirms process exit and I/O settlement before reporting completion or -failure. Only `failed` can carry a bounded `exit_code`. For setup and package steps, Core projects a confirmed nonzero status, never -process output; zero or missing status retains the generic reason. Only a confirmed -Skill preparation failure gets the Skill label. Plugin/finalization failures and -uncertain effects retain the generic reason. The Store composes the reason from a fixed label and integers, so commands, env values, package names, -paths and any process output never reach the reason, events, logs or responses. -The failed step is -not retried and later steps do not run. - -## System packages - -`packages.system` is unsupported in both Templates and inline configuration. -Supplying the member, including null or an empty list, returns an explicit -validation error with param `packages.system`: preinstall system dependencies in -the sandbox image/template or on the host machine. The field is not silently -ignored, inherited into a privileged action or translated to apt. - -The daemon runs as its starting account. It does not install system packages, -request sudo or elevate permissions. Operators build managed images/templates -with required system dependencies; self-hosted users prepare them before execution. -A missing executable or library fails the operation that requires it. npm/Python -packages and setup retain the common user-directory initialization path. -Historical system-package qualification below applies only to its old binaries; -there is no system-root seed, package-root launcher or upgrade compatibility path. - -## Restricted network policy - -Template and inline configuration share Core validation, persistence and resolution. -`restricted` requires 1–100 exact ASCII hostnames; subdomains and redirect destinations -need their own entries. Unsupported host forms (wildcards, URL/port syntax, IP literals, -Unicode and trailing dots) reject explicitly. This is a qualified subset, not a -claim of complete upstream hostname normalization or TLS routing semantics. -Public reads preserve supplied spelling, order and duplicates. Effective native -comparison uses a separate lowercase, deduplicated copy; updates cannot change -existing Session snapshots or retry intent. - -The daemon does not enforce `disabled` or `restricted` network modes. Stored -configuration and successful Template CRUD do not establish execution support. -An execution combination must reject unless its outer Environment provides and -qualifies the requested restriction; it must not silently run unrestricted. -The normal native self-hosted combination uses the host's existing network. -Historical native-proxy or sandbox allowlist results below do not qualify current -outer enforcement. Read-only Files retains its separate minimal prerequisites. - -## Explicit gaps and evidence boundaries - -Templates and inline initialization share Plugin, capability-directory and Skill -reference installation. Unsupported native activation and unqualified protocol -semantics remain explicit gaps as described above. The separate live Files API remains -available after initialization. Unsupported requests reject without echoing payloads. - -The [hosted guide](https://developers.openai.com/api/docs/guides/agents-api/environments/openai-hosted) -clarifies that configured env values are readable by Agent code, files/packages -precede setup commands, nonzero setup prevents start, and runtime-reserved env names -must reject. OpenAgentCore reserves the complete `OAC_` prefix in Environment -`env` (including Runtime, Web, logging, developer and test names); it replaces the -former `PARSAR_` reservation. `PATH`, `OPENAI_API_KEY` and `CODEX_` remain reserved. -The shared initialization batch implements those fields with encrypted snapshots -and the existing readiness gate. Public reads show packages but omit env/commands. -Template updates replace each supplied field; omission preserves it and null clears -it. Referencing Sessions use the composition rules below; these differ from -Template.update replacement rules. - -Files and resolved Skills are installed first, followed by npm/Python packages and ordered commands; -the default cwd is `/workspace`. One command or package operation has the existing -two-minute local budget, within the thirty-minute initialization budget. No command -is retried after unknown effects. Completed setup never runs on reconnect. -Package dependencies are available to native tools across working directories. -System dependencies must already be installed; `packages.system` rejects. - -The [update Reference](https://developers.openai.com/api/reference/python/resources/beta/subresources/agents/subresources/environments/subresources/templates/methods/update) -defines runtime network as post-setup and packages as preceding that policy. -Runtime initialization uses the host's existing network. The daemon does not -provide the reference's later restricted-network enforcement. Unsupported execution -combinations must reject; the phase ordering is not an isolation guarantee. Env values are intentionally readable by Agent code; they must not -appear automatically in public metadata or initialization diagnostics. - -The [current Template reference](https://developers.openai.com/api/reference/python/resources/beta/subresources/agents/subresources/environments/subresources/templates) -mentions different GA/beta defaults; this service retains `agents=v1` and the -[fixed baseline](upstream.json), whose omitted network is enabled. Exact upstream -errors, concurrent pagination and broader network combinations remain unverified. -Referenced null network follows the [qualified inheritance rule](template-null-selection.md); -unsupported hostname forms still reject. This is not full protocol compatibility. - -## Template and inline configuration composition - -The pinned Session description applies the template before inline configuration. -Owned official API probes on 2026-09-23 established composition behavior. The -current implementation applies it only to supported fields; the table below -excludes the now-unsupported `packages.system` member: - -| Session field | Omitted or null | Non-null inline value | -|---|---|---| -| `env` | Inherit template keys | Overlay by key; inline value wins, `{}` preserves all keys | -| `setup_commands` | Inherit template sequence | Replace the sequence; `[]` clears it | -| `files` | Inherit template file set | Replace the complete set; `[]` clears it | -| `packages` | Inherit supported managers | Resolve Python/npm independently; `{}` inherits both; supplying `system` rejects | -| Individual package manager | Inherit its template list | Replace that list; `[]` clears it | - -Core performs this composition once, before the existing encrypted Session snapshot -transaction. Env values, command bodies and inline bytes remain absent from public -metadata. Effective package/file metadata reflects the selected inputs. Caller -intent still distinguishes omission, null and explicit fields for the local creation -retry policy; composition does not rewrite it. Later template/source changes do not -rewrite committed initialization. Existing per-source and effective-size validation, -network narrowing, Skill/Plugin selection and source authorization remain in place. -No harness or Provider participates in the merge. - -The official evidence used SDK 3.13.0, upstream `d7c41ef`, and `agents=v1`: -86 authenticated calls, nine owned Sessions, four lifetime Templates (including two -cleaned setup-script assertion failures), and six completed real `gpt-6-astra` Turns. -Four exit-zero command outputs confirmed env values and command replacement/order -for omitted, populated, empty and null inputs. Two further outputs confirmed full -file replacement and null inheritance, and an initialized empty listing confirmed -`files=[]`. The omitted-files case reached the bounded readiness limit, so its -inheritance proof is metadata-only. Package composition is also metadata evidence; -these official probes do not independently establish installed package versions. -All owned resources have successful public DELETE receipts; physical upstream -sandbox destruction was not independently observed. - -Private evidence: `~/.parsar/remediation/20260923/template-inline-composition/`, -subdirectories `official-env-setup` and `official-files-packages`. Reports retain -fixed-source snapshots, raw status/body/request IDs, command-output proofs, -accounting, cleanup and credential scans. This covers the observed fixtures rather -than all possible combinations. Null network and null Skill/Plugin/directory list -selection remained outside that composition batch; the -[subsequent qualification](template-null-selection.md) records those rules. -No new protocol version is introduced. - -### Core checks for composition (2026-09-23) - -`TestTemplateCompositionOfficialClientPostgres` uses the fixed strict SDK and raw -HTTP against real Core handlers and PostgreSQL. Six accepted cases and seven -rejected creation keys cover effective metadata, confidential frozen bytes and -ordered commands, tenant/source isolation, combined setup-size rejection without -partial records, and same-intent retries after template and source deletion. -A second Store/handler verifies reopened persistence; it is not an OS process -restart and does not run a model. API tests additionally cover raw caller intent, -non-mutating composition and retained field validation. - -The integrated server gate ran `make -o check-web check` at `5589df0`; this includes -real PostgreSQL tests, byte-for-byte sqlc generation, builds, native bridge checks -and Rust tests/format/Clippy. A fresh `make check-web` ran locally against the same -unchanged production diff: 287 client, 583 Web and 76 browser tests passed. Together -they cover every required `make check` target. `make openapi` regenerated the -Session description. The optional MiniMax packaged-native scratch/large-output -probe was skipped because its optional profile variables were unset; this batch -does not add MiniMax, Claude or E2B native qualification. - -### Real Codex Docker composition acceptance (2026-09-23) - -This is historical evidence for the recorded images, including their former -system-package and inner-sandbox implementation. It does not qualify current -Runtime permissions or permit `packages.system`. - -The exact `8019ac5` production build ran independently with Core, PostgreSQL and -Docker Runtime against the real Kimi API. Three completed native model Turns prove: - -- A populated-inline Session observed env key precedence, whole-file replacement, - ordered replacement commands, Python `packaging==26.0`, and inherited npm/system - tools. Its prompt named only the read operation, not the expected canary values. -- A separate inheritance Session observed null env/files/commands inheritance, - retained npm/system tools and absence of the explicitly cleared Python installation - directory. After template mutation/deletion and an actual Core process restart, - identical-key creation recovered its existing identity/configuration. A second - native Turn returned the exact same values and append trace: initialization did - not run again. The populated Session itself was not resumed in this run. - -Five owned Core Sessions were created over the acceptance attempts. The first two -failed before any model Turn because the reused private Skill-only runner omitted -Docker `nested_sandbox: true`, required by that historical initialization profile. Correcting that operator setting resolved the proc-mount failure without -production changes. In the corrected pair, the populated Session completed its -native command, but the private runner then required an optional assistant `phase` -and unwrapped JSON. The original message instead contained matching fenced JSON -without phase. An offline check preserved and verified the original native output -and message bytes; no model call was repeated. One new inheritance Session completed -the remaining two Turns. These failed assertions and operator diagnostics are -retained alongside successful evidence, not counted as extra successful runs. - -Private evidence is under the same batch's `live/` directory and records source, -binary/image hashes, operator-setting changes, exact native output and cleanup. -This is Codex/Docker qualification for the described composition paths, not renewed -qualification of other harnesses or Providers, every package manager combination, -or complete Template/Agents API semantics. - -## Verification - -The dated results in this section are historical evidence for their exact source, -images and inputs. They do not qualify current bypass execution, cross-platform -preparation, or outer network enforcement. In particular, former system-package -installation and private-file denial results are not current support promises. - -### Restricted-network Docker acceptance (2026-09-21) - -The fixed SDK 3.13.0 and raw HTTP passed restricted-policy CRUD, tenant isolation, -template narrowing and immutable creation-retry checks against the standalone Core. -Current Core/daemon builds with Codex 0.153.4, Claude Code and MiniMax Code passed -real Kimi/MiniMax execution, initialized system tools and Skills, Files/Artifacts, -credential isolation, cancellation with observed descendant cleanup, and retained -workspace/native history after separate Core and Runtime restarts. Codex exercised -both template and inline configuration; Claude and MiniMax exercised templates. - -Separate real-model network runs on all three Docker profiles verified HTTP and -certificate-checked HTTPS to allowed sites, rejection of an unlisted host and -subdomain, rejection after an allowed site's redirect, and failure of direct-IP or -proxy-free access. Allowed HTTPS, host rejection and direct-bypass checks repeated -after Core/Runtime restart with the frozen policy and retained conversation history. -The checks use actual tool effects, native command Items where available and exact -transport connection counts, not the model's assessment. All accepted runs cleaned -their owned containers, volumes and test transport. - -The test host lacked direct DNS/TCP egress. A task-only network-namespace route and -transparent sidecar carried unchanged HTTP/TLS bytes to real sites through the -existing outlet; host routes, Runtime capabilities, native policy and certificate -validation stayed unchanged. This qualifies native enforcement through that test -outlet, not production direct egress or DNS. An earlier Codex probe incorrectly -required a complete result file after native denial; the corrected probe requires -the corresponding native rejection when the tool is interrupted. One initial -MiniMax run failed before its first tool call; an independent rerun passed, without -a Core change or an established root cause for that failure. - -Focused policy/API/adapter/PostgreSQL tests, the real Codex managed-process lifecycle -test and the complete Core `make check` passed. Optional Docker fault fixtures were -not enabled; real public Docker runs cover the accepted paths. The first full-check -invocation lacked the server's OpenSSL development paths; the corrected invocation -passed with a fresh database. Evidence and failed attempts remain under -`~/.parsar/remediation/20260921/template-network-native/`. -Core SHA-256: `0dc40384192fc75c4be9896072dfe0089c30a02e44804faca7a14c7d8525efa3`. -E2B probes remain mechanism evidence only; this batch does not qualify official -E2B self-hosted onboarding or complete upstream network semantics. - -### Resource and initialization checks - -The E2B fixture and opt-in flags described below are historical evidence for the -retired Core-managed deployment, retained in Git history at `d03e1d25`. Current -user-managed E2B uses the shared daemon enrollment path and does not resolve hosted -Templates. These historical tests do not qualify the replacement deployment. - -`official_environment_templates.py` checks all five fixed-SDK operations plus raw -HTTP, exact safe response shapes, field replacement/defaults, pagination, tenant -isolation and rejected confidential canaries. `official_e2b_v1.py` opts in with -private `verify_environment_templates: true`; it creates its actual native/model -Sessions from public templates, verifies frozen snapshots and creation retries -after update/delete, then reuses the existing execution, Files/Artifacts, isolation, -cancellation and crash/history-recovery assertions. Its disabled-network Session -inherits that policy from another template. Runtime and provider packaging were -unchanged in the original metadata-only batch. Database integration tests cover persistence and concurrent field updates; -API tests cover parsing and caller-intent distinctions. - -`official_environment_initial_files.py` and the `verify_initial_files: true` option -together with `verify_environment_templates: true` in the real E2B runner add -both-source/template/inline metadata, source-deletion, -foreign-tenant and actual first-native-read checks. Existing Files/Artifacts, -cancel/crash/history checks then verify that initialization did not change the -execution loop or overwrite later user modifications. Controlled PostgreSQL lifecycle -tests separately exercise interrupted installation, readiness and maintenance fairness. -A test's presence is not a passing acceptance result; retain actual run evidence. - -### Accepted initial-file profiles (2026-09-20) - -The batch passed fixed SDK 3.13.0/raw HTTP acceptance with real models on Docker -and E2B for Codex, Claude Code and MiniMax Code. Both template and inline paths -verified initial native reads, Files/Artifacts, tenant and credential isolation, -source/template deletion followed by creation retry, cancellation, and preserved -workspace changes/native history after Core and Runtime restarts. E2B also verified -Core interruption during initialization: no native execution, no replay and owned -resource reclamation. All six completed runs reported clean resource cleanup. - -Separate real Provider checks covered Docker binary stdin/backpressure and E2B -50 MiB stdin. The real shared installer verified empty, binary, nested and 50 MiB -files, rejected symlink destinations, and preserved outside bytes. PostgreSQL/race -suites and `make check` passed. The optional native build probe skipped by the -default gate is not counted as real acceptance. Runtime images were the retained -qualified builds; Core was built from this batch. E2B runs preceded the final -readiness guard and store-interface cleanup, which received targeted regression; -The six-profile matrix preceded final creation-intent size and canonical-identity -corrections. Real HTTP/PostgreSQL regression accepted a 1 MiB file and two 5 MiB -inline files with retries, and verified canonical template/file encryption bindings. -A further rebuilt standalone Docker/Codex run passed a 5 MiB initial file with -real model reads, Artifacts, cancellation and retained history in 101.23 seconds. -The original Docker matrix used Core SHA-256 -`31973b17dd96106743e581c400555e3a4b036ad8cb3e68b51530a2b56023abe3`. - -Docker MiniMax Code passed with the real MiniMax API at its standard HTTPS origin -through the test network relay. Earlier Kimi/MiniMax connection timeouts remain -recorded with unknown cause, as does a Docker reconnect failure under a different -Core/Runtime restart order. They are not claimed as fixed. Sanitized run results, -checks, build hashes and failed attempts are retained under the private -`environment-template-files` acceptance directory and the linked task record. - -### Accepted env/setup and npm/Python batch - -`official_environment_setup.py` adds fixed-client/raw-response assertions for -confidential snapshots, safe package metadata, real registry dependencies, ordered -setup, native visibility across cwd and the post-setup network boundary. Runtime -mechanism tests cover private files/processes, immutable configuration, child -cleanup and failure receipts. The batch passed real-model template and inline acceptance on newly built Docker -Runtimes for Codex, Claude Code and MiniMax Code, plus Codex on a newly built E2B -template. Each verified actual npm/PyPI installs, ordered setup, native env and -dependency visibility across working directories, Files/Artifacts, cancellation, -post-setup disabled networking and recovery without repeating initialization. E2B -also verified daemon/history/process/envd isolation and separate Core/Runtime -crashes with exact native history and no automatic input replay. - -The shared initializer additionally passed actual Docker isolation probes for all -three profiles and E2B registry installation. Docker nonzero setup and missing cwd -failed before native Turns and reclaimed the Environment. PostgreSQL/race checks -cover encrypted owner/field-bound snapshots, readiness and uncertain-install cleanup. -A rebuilt standalone Core passed real fixed-SDK/raw-HTTP retries with changed, added -and removed inline env/setup under a saved Agent; unchanged retries still recover -after Agent deletion. Only this creation-identity regression required the final -Core rebuild; the completed model matrix preceded that isolated hash correction. - -One MiniMax inline post-restart model request reported an upstream timeout after -100 seconds. The affected inline rerun passed in 214.75 seconds, with no production -transport changes; this does not establish or fix the timeout cause. All completed -runs confirmed owned resource cleanup. The three-harness-by-two-Provider matrix -was not repeated: shared E2B initialization and the changed native adapter paths -were covered separately. System packages were outside that batch; a later -historical qualification is recorded below. Unconfirmed reference overrides -remained gaps in that evidence, and the retired native Codex hook's failure -limitation was not resolved by this run. -Private sanitized run/check/build evidence is retained under -`~/.parsar/remediation/20260920/environment-template-setup/`. -These results do not establish complete Template or Agents API compatibility. - - -### Accepted inline-Skill profiles (2026-09-20) - -`official_environment_skills.py` supplies fixed-client/raw-HTTP checks and a -native-discovered Skill whose helper produces an unpredictable Artifact, checks -private credentials/staging, and attempts to modify its own installed manifest. -Standalone Core, dedicated PostgreSQL and freshly packaged Docker Runtimes passed -with Codex 0.153.4/Kimi, Claude SDK 0.3.269 (native 2.1.269)/Kimi, and MiniMax Code -0.4.12/MiniMax-M3. Codex covered template and inline configuration; Claude and -MiniMax covered the template path through the same initializer. All verified safe -metadata, foreign-tenant rejection, frozen snapshots after template clear/delete, -creation retry, native Skill execution, Files/Artifacts, cancellation, and owned -history/workspace recovery after Core and Runtime restart. Cleanup completed. -The Core binary SHA-256 was -`85be1bc26ca6c03617ba74bf092485656dda9311bb636d507be6576a2f087f16`. - -The shared initializer also passed on a real Docker container and E2B VM, including -binary/executable content, read-only Skill access from setup, duplicate/path -rejection and private-state isolation. This batch did not repeat the E2B model -matrix: Provider code is unchanged, while its shared initialization boundary was -exercised in a real VM. PostgreSQL/API/archive tests, Claude SDK tests/build, -`make openapi`, `make sqlc-generate` and `make check` passed. The default gate's -optional native build probe remains skipped and is not counted as live acceptance. - -Initial integration failed safely because the proposed Skill parent was root-owned; -using the existing Runtime initialization directory resolved that packaging -boundary without broadening permissions. The earlier MiniMax/Kimi native timeout -and probe-only unrestricted-tool configuration failure remain recorded. The latter -passed after restoring the unchanged production tool settings. Docker builds reused -qualified base images after registry DNS failure; current daemon, adapter and -initializer artifacts were copied using the repository packaging recipe. Raw -receipts retain their inherited historical manifest fields; accompanying source, -Core and image hashes identify the actual candidates. Evidence is retained under -`~/.parsar/remediation/20260920/environment-template-skills`. This profile does not establish complete upstream Skill semantics. - -### System-package qualification (2026-09-20) - -The batch passed standalone Docker acceptance with current-source Core/daemon and -newly packaged Codex, Claude Code and MiniMax Code Runtimes. Fixed SDK 3.13.0 and -raw HTTP exercised public templates; Codex also exercised inline configuration. -Actual Kimi/MiniMax requests verified jq, compiler/libpq linkage, dependent -npm/Python packages, ordered setup, native visibility, read-only installed roots, -Skill/credential protection, Files/Artifacts, public cancellation and retained -workspace/native history after Core and Runtime restart. Cancellation checks -observed tool identities disappear before sandbox teardown. All three completed -runs reported clean resource cleanup. Template omission, replacement, null/empty -values and atomic invalid-input rejection received additional real HTTP checks. - -The Core SHA-256 was -`9466a8419fd0e4ad8cd4fb1786ef131c2cc504513bf77642fca43ce07b8114a6`. -The Codex template/inline run took 599.72 seconds; Claude and MiniMax template -runs took 271.69 and 357.34 seconds. Real initialization mechanism checks separately -covered isolated package scripts and compilation. Focused Go/SDK tests, OpenAPI -generation and `make check` passed. The optional native build probe skipped by -the default gate is not counted as real acceptance. - -E2B finalization rewrites `/usr/local` permissions. Its trusted bootstrap must -restore the common system-tool launcher's packaged `0555` mode before native -preparation; root ownership alone does not satisfy that Runtime receipt check. -The first qualified Codex E2B template passed real Kimi template and inline acceptance in -576.41 seconds, including final seed/launcher protection, actual package/setup -visibility, Files/Artifacts, credential/history/process/envd isolation, public -cancellation, separate Core/Runtime crashes, continued owned history without -input or initialization replay, preserved user files, and disabled native-tool -networking. Cleanup completed without fallback errors. This run uses the updated -daemon with the bounded discovery adjustment described below. The immutable build -is `1b60xhq0j13fnr5zipkg:7d11189a-b2bd-4690-bc3f-da0792439f91`. The full -three-harness E2B matrix was not repeated: shared initialization and the changed -native adapter paths were covered separately. - -Independent review then identified a missing native cwd alias: the installed tool -root exposed `/workspace`, while Codex retained `/environment/workspace`. Both -now mount the same authorized workspace. A rebuilt Docker Runtime passed actual -system/npm/Python initialization and entry from the default directory and its -subdirectory in 106.30 seconds, including private-state isolation and read-only -tools. The earlier model runs selected `/workspace` and do not prove this fix. -The rebuilt E2B template -`1b60xhq0j13fnr5zipkg:e6437586-927e-4683-99fe-632b51a974fd` then passed the -real Kimi template/inline loop in 457.91 seconds. Native command Items and actual -effects verified the default directory and subdirectory; the same run passed -Files/Artifacts, private-state isolation, cancellation, Core/Runtime recovery, -preserved history and user modifications, and disabled tool networking. Cleanup -reported no errors. The final mount-only correction received this actual regression -and Python source checks; the two full `make check` runs precede it. - -Failed attempts are retained: early admission incorrectly required the private -initialization receipt; execution preparation now owns that check. Test-only proxy -configuration and simultaneous package installation attempts failed before the -sequential accepted runs, without extending production budgets. E2B cold discovery -once killed `codex --version`; unchanged discovery subsequently passed, but a later cold deployment repeated -the failure with no observed OOM. The shared CLI availability probe now allows -15 seconds instead of five; no retry or Provider-specific startup path is added. -The precise initial paging/contention cause remains unconfirmed. Docker execution -results above precede this isolated startup-budget adjustment. The E2B launcher-mode mismatch failed preparation before any -native input was applied. The subsequent native isolation fixture assumed -`sudo` existed; the real tool transcript showed `FileNotFoundError`. The fixture -now records an absent privilege command explicitly while retaining all authority -and private-state checks. Interactive PTY behavior and packages needing additional -Unix identities or privileged services are not qualified by these results. - -Sanitized results, image/source hashes, full checks and failed evidence are retained -under `~/.parsar/remediation/20260920/environment-template-capabilities`. -Early Docker result manifests contain inherited installer archive fields; those -fields do not qualify a new installer archive. Current binary and image hashes -identify the tested deployment. These checks do not establish complete upstream -Template or Agents API compatibility. - -## Reference batch validation and limits - -Resource operations passed fixed SDK/raw HTTP and real PostgreSQL checks. The -shared reference path passed real Docker execution on all three qualified profiles: - -| Harness | Real model | Complete driver result | -| --- | --- | --- | -| Codex | Kimi K3 | Passed, 308.80 seconds | -| Claude Code | Kimi K3 | Passed, 211.68 seconds | -| MiniMax Code | MiniMax M2.7 | Passed, 162.71 seconds | - -Each run covers SDK upload, unresolved template intent, concrete Session metadata, -native Skill supporting files, public Files/Artifacts and tenant checks, source -and template deletion, committed retry, and cold Core/Runtime continuation without -reinstalling or replaying the original Turn. All report zero cleanup errors. The -current Core uses the previously qualified Runtime images and native adapters; -this batch does not change their execution architecture or qualify E2B references. - -MiniMax's initial attempts failed because the test host's SOCKS route exceeded the -native TLS connection deadline. A temporary operator SSH byte relay restored normal -TLS latency, retaining native certificate validation and the same model/credentials. -One subsequent response recalled the exact random history value but omitted its -fixed prefix. Clarifying the test prompt to request the complete literal token, -without supplying its random value again, passed the unchanged exact assertion and -remaining checks. Failed evidence is retained; no Core output repair or native -connection-timeout change was made. - -`make sqlc-generate`, `make openapi` and the standalone `make check` passed; the -full gate took 484.22 seconds. The optional MiniMax packaged-native scratch/large-output -probe was skipped because its profile/artifact variables were unset. Unchanged -native-profile qualification and the real reference workflows are separate evidence. -A user-authorized reused-context GPT-6 Astra high reviewer inspected the entire -52-file diff and original acceptance records, with no material actionable findings. -This was not a fresh-context blind review. These checks do not establish complete -upstream Skill, Template or Agents API compatibility. - -Sanitized results and operator drivers are retained on `zju_a100_2` under -`~/.parsar/remediation/20260921/template-skill-references/`, with a local evidence -index under `~/.parsar/remediation/20260921/template-capabilities-design/`. - -The current upload profile accepts at most 500 regular files, 5 MiB compressed -and 20 MiB expanded per bundle. These are qualified implementation limits, not -published protocol maxima. [File resource qualification](file-resource-semantics.md) -records default-selected unversioned content and descriptive metadata, default -deletion rejection with multiple versions, and nondefault latest pointer fallback. -Core updates the default pointer and top-level name/description atomically, while -concrete version bytes and previously frozen Sessions remain immutable. Deleting the -last version deletes the Skill; version numbers are intentionally never reused. -Complete errors and visibility timing remain gaps. - -## Implementation rules - -Public Environment Templates belong to Core and its execution database, independently -of provider image/build templates. Resolve a tenant-owned reference once at Session -creation, freeze the effective ordinary hosted configuration and reuse inline -initialization. Do not pass template IDs into Provider or Runtime. Omitted or null -network inherits the complete template policy; overrides may only narrow policy. -Preserve unresolved caller intent for creation retries and recover committed -results before reading mutable templates. -For template-reference Session initialization, omitted/null env, files, commands -and packages inherit. Overlay non-null env keys; replace non-null files and command -lists, including empty lists. Select each package manager independently: omitted/null -inherits, while a supplied list replaces that manager. Revalidate the effective -configuration through the existing validators and freeze it through the same -transaction as inline initialization. Keep caller intent separate from resolved -configuration; do not add a second installer or pass merge rules to adapters. -Updates and deletion cannot rewrite existing Session snapshots. Initial files use -one Core-owned installer for template and inline configurations. Keep confidential -bytes encrypted under the execution-service key and resource-bound AEAD, separately -from ordinary configuration and public metadata. Templates retain source references; -Session creation freezes tenant-authorized source bytes in the same commit, independent -of later source/template deletion. Public resource reads must not require decryption -or load encrypted file bodies. Record original creation intent before resolution. - -Template parsing, persistence and resolution must not select a harness or Provider, -or depend on native tool names and private harness paths. The shared initializer -uses the packaged Runtime contract for trusted commands, workspace/staging paths, -confidential input and completion receipts, following -[Runtime capability preparation](environments.md#runtime-capability-preparation). -Each adapter owns native tool configuration. A new harness or Provider must not require template -business-logic changes. Reuse qualified shared helpers even when their executable -names have historical engine prefixes; renaming is not a boundary fix. Select real -regressions by the changed shared, Provider and adapter boundaries, rather than -repeating every deployment combination for each configuration field. - -Name, initial files, inline/referenced Skills, Plugins, workspace capability directories and env/setup/npm/Python are -implemented independently of remaining installation fields. Reject unsupported -inputs rather than persisting them for silent -omission; expand inline and template initialization together in separately qualified -batches. Resource reads need only tenant authorization, not a live Runtime. diff --git a/contracts/agents-api/environments.md b/contracts/agents-api/environments.md index bb509ed80..a6a59b674 100644 --- a/contracts/agents-api/environments.md +++ b/contracts/agents-api/environments.md @@ -1,849 +1,387 @@ -# Environment contract and implementation path - -This assessment covers the fixed [Python SDK contract](upstream.json), with partial -coverage. Core-managed Docker runs the three qualified native harnesses. V1 -user-managed `self_hosted` enrollment uses the same colocated Runtime for Codex, -Claude SDK and MiniMax. `/workspace` names the packaged workspace; a self-hosted -Session may instead select its exact canonical physical directory. The -[recorded deployment qualification](user-managed-runtime-v1.md) retains its historical -source, paths and tested capability scope. -For hosted compute, the deployment selects E2B, Docker or microsandbox through -[sandbox deployment](sandbox-deployment.md). -For the separate caller-managed E2B path, the user owns allocation, renewal and -cleanup through the official SDK and [Runtime packaging](../../services/core/deploy/e2b/README.md). -[Templates](environment-templates.md) provide reusable preparation; the Core extension also applies their execution configuration to user-managed machines; -[Files](environment-files.md) reuse the exact authorized local workspace. - -The internal Store now owns a durable Environment association for newly created -`self_hosted` and `openai_hosted` snapshots, atomically with Session creation. -It derives configuration and tenant ownership from the Session; retries preserve -the existing identity. Scoped reads hide associations after Session deletion while -retaining the underlying record. The public text profile reuses this association -and the preparation/admission path below; additional provider profiles remain open. -Missing/`none` configurations and historical internal snapshots gain no backfill. - - -The private daemon gateway authenticates the enrolled Environment executor key and -exact dedicated device. Existing Worker connection observations retain generation -and revision fencing; registration and connectivity do not establish native -readiness. Rotation/revocation, Session deletion and ownership loss deny further -access without promising immediate cessation of native effects. Native history -remains local to the bound Runtime and cannot be replaced on retry. There is no -registry/Noise relay or transient service-side harness credential. -See the [enrollment guide](../../services/core/README.md#user-managed-runtime-enrollment). - -## Basic public Docker-hosted profile - -An explicitly configured default managed Docker provider enables `type=openai_hosted` -for the qualified Codex, Claude Code and MiniMax Code profiles. Each uses the same -Runtime lifecycle and workspace interfaces with its native adapter. The outer -Environment provides managed isolation; the daemon adds no inner sandbox. -See the [engine profile guides](README.md#public-engine-profiles) for setup and limits. -The standalone [operator configuration](../../services/core/deploy/codex/README.md#standalone-operator-configuration) -selects the qualified immutable Runtime image; advertised capabilities alone do -not enable admission. An idle or initial-text creation commits Session, Environment -and retry identity before the existing leased Worker provisions its allocation. -A committed creation interrupted before bootstrap is recovered without replaying -an existing allocation's Create. - -Omitted/null network defaults to enabled. The daemon does not enforce disabled -or restricted networking. Current execution profiles reject both before Session -creation and resource allocation. -Templates and inline configuration share initial files, env, packages, ordered setup -and inline or tenant-owned referenced Skills through ordered Environment initialization. -The same authenticated daemon handles initial files, tool configuration, packages, -setup, Skill/Plugin import and directory snapshot finalization through -`runtime_prepare`. Providers place, create, bootstrap, inspect, renew and reclaim -compute. npm/Python installation and setup use the host's existing network and -starting account. System dependencies must be preinstalled; `packages.system` -rejects explicitly, including null and empty lists, and never invokes apt, sudo -or elevated execution. Package responses include the official required -`system: []` field as empty metadata; it does not enable installation. -Unsupported hostname forms and installation combinations reject explicitly; see -the [Template coverage and limits](environment-templates.md). Empty/null installation -defaults produce safe empty metadata, not a live workspace inventory. Service-origin -hosted MCP remains unsupported; Environment Plugin MCP has its -own qualified transport matrix. - -Initial provisioning leaves a Session idle until a Turn starts, with no caller -connection action. The managed scan records authenticated, exactly bound daemon -connections through existing fenced generations. Native preparation remains -separate. Core restart preserves allocation/workspace/native identity; it does not -blindly replay uncertain work. Terminal cleanup revokes authority and settles -pending input atomically before external reclamation. Matching retries preserve -outcomes; new inputs reject terminal Environments. Expiry has no invented SSE -variant. Local failure codes and exact event ordering remain unverified upstream -semantics; this profile does not establish complete Environment compatibility. - -## User-managed E2B profile - -The application creates, renews and destroys its E2B sandbox through the official -SDK. It deploys the shared Runtime, then enrolls that Runtime into a `self_hosted` -Session. Core neither keeps an E2B allocation nor issues Provider renew/kill calls. -The [E2B guide](../../services/core/deploy/e2b/README.md) owns packaging and -user-side lifecycle instructions. Expiry or lost workspace/history must not trigger -transparent replacement or replay. The [new enrollment qualification](user-managed-runtime-v1.md) -records its own real deployment evidence. -The [prior E2B qualification](README.md#e2b-v1-qualification) records a historical -Core-managed route; it does not establish this caller-managed enrollment result. -Current deployment-managed E2B is a separate hosted configuration in the manager guide. - -## Initial public self-hosted profile - -Create a Session with `environment.type=self_hosted`, -a clean absolute `workspace_directory` and optional absolute local -`capability_directories`. `/workspace` maps to the Runtime-bound workspace. Creation -accepts initial text as a string or ordered user-message array. Omitted/null input -creates no Turn or connection action. Configured execution, an enabled harness and -the exact local profile are validated -before persistence. Supported optional functions remain engine-specific. - -Initial text commits a reservation and connection action, then returns the Session -and Environment connection target while offline. Streamed creation sends the same -committed projection as its `created` snapshot, already showing `requires_action` -and the connection action, then the committed `requires_action` event. Like -every fresh creation stream it ends right after the idle recorded when the -admitted Turn ends or the reservation stops being pending, or after a failure; -the connection alone clearing the action does not end it. A creation without -input ends right after its created snapshot, and a same-key stream retry ends at -once without events. The existing Worker prepares and admits the input; closing the stream -leaves committed work intact. -Initial expiry leaves a failed Session, safe error and empty actions without a -Turn or an Environment failure. Creation retries preserve the original identity, -deadline and input. Later live observers do not replay creation events. - -Later text-only batches recover their original reservation or direct receipt under -the Session lock. New active messages append to the current Turn through existing -ordered admission and native delivery; they create no Turn, preparation or reservation. -If no Turn is active, reserve and wait for the existing Worker to retain native -preparation, admit and claim. Return 204 only after that -transaction commits. Connection actions precede a Turn and clear on connection; -connection alone is not readiness. Retries preserve identity and the five-minute -database deadline. HTTP disconnect retains the reservation. Local expired/cancelled -outcomes return 409, lost ownership 503, and deletion 404; exact hosted error -status/body and pending-input crash recovery are unverified. See the -[canonical wait rules](#ownership-and-placement-decision). - -Cancellation-only batches use the existing locked admission and native delivery. -An idle cancellation creates no Turn, and retry identity preserves the original -target during later work. Pending pre-Turn reservations still block new cancellation. -HTTP 204 confirms admission, not native completion or OS quiescence. A cancellation -before native Session transfer can still lack a final Outcome and conservatively -fail; complete cancellation settlement remains open. -Homogeneous function-result batches reuse scoped locked admission and native -application receipts, without creating a Turn or preparation. Definitions use the -existing function parser and remain fixed through native preparation and continuation; -output/error field presence and ordered content keep their existing semantics. -Retries retain the original call, including during later work. New results cannot -bypass pending input. Function callbacks do not populate Environment installations. -Mixed events and non-text input remain unsupported. Local capability directories -use the same Runtime parser and installed snapshot as managed Skills and Plugins; -preparation must finish before native execution. Reconnect reuses installed bytes, -while new Sessions capture their own sources. Recursive references to the installed snapshot and escaping directory -entries are rejected. Platform support and isolation follow -[Platforms and isolation](#platforms-and-isolation). Enrollment must match either the `/workspace` -logical alias or the exact canonical directory bound by the Runtime; selecting an -arbitrary path does not grant access. `remote_url` is the configured daemon WebSocket URL, -returned unchanged. This is our private connection contract and does not claim -stock `exec-server` compatibility. Service-origin HTTP MCP is explicitly rejected -on `self_hosted`; `none` MCP and hosted Template Plugin MCP retain their own scope. - -Mechanism tests cover enrollment, credential checks and exact-device dispatch. They -do not establish real public qualification. Before claiming that qualification, -exercise fixed SDK/raw HTTP creation/input, actual native tools, Files/Artifacts, -second-Turn history, Core/Runtime restart, cancellation and credential rotation/ -revocation/deletion on each declared deployment. - -User-managed onboarding creates a `self_hosted` Session first, then passes its -Environment ID and unchanged `remote_url` to our Runtime with connect-only -authorization. This is our daemon connection contract, not stock OpenAI -`exec-server` transport compatibility. Validate the pinned public HTTP/SDK -resources, state transitions and lifecycle separately; do not infer complete -compatibility from a working connection. User-side tooling owns local Runtime or -E2B allocation, renewal and cleanup. Session deletion and credential revocation do -not transfer ownership of user compute to Core or prove process quiescence. - -Keep prerequisites specific to the public operation being implemented. Native -harnesses execute; adapters translate protocols and fill demonstrated capability -gaps; Core owns public semantics, authorization and resources. Before adding a -mechanism, identify the current operation it enables and why existing native -capabilities or interfaces do not suffice. Durable metadata queries need no live -runtime. Live file reads require an authorized, isolated view of the exact workspace -and bounded operation ownership, but not a complete file-write, environment -replacement or placement-retirement implementation. Apply mutation fencing and -retirement guarantees where an operation can write, replace or retire that owner. -Read-only access still requires tenant/resource checks, path isolation and safe -failure when the authorized workspace cannot be reached; it never grants public -admission merely because an adapter advertises a capability. - -Use distinct authorization for callers, devices and environment connections. A -managed outer Environment must exclude broader application credentials and other -tenants' secrets. The daemon does not hide its own state from same-user tools. -Directory bindings and process identities do not provide filesystem isolation. -Preserve or demonstrably restore native history across -compute replacement; never silently move a bound Session or replay unknown work. -Self-hosted compute/files remain caller-owned, with explicit cleanup separate from -Session deletion. The full Environment implementation remains pending; follow the -[pinned contract and acceptance sequence](environments.md) -and the [workspace placement map](workspace-placement.md). -For co-location, qualify the outer deployment boundary and the shared Runtime -lifecycle. Native tools use the starting account's permissions; do not claim a -daemon or harness sandbox. Host operators choose their own outer isolation. -Never enable an execution combination whose required outer behavior is unverified. - -The internal Store creates one Environment with an environment-bearing Session in -its creation transaction. The Session upsert selects the retry winner; retries -never create or repair associations. Environment identity/state live in their own -table. Tenant ownership and immutable configuration come from the owning Session, -without duplicated JSON, tenant columns or generated IDs in the creation hash. -Environment reads join that Session and exclude deleted Sessions; deletion retains -ownership for later settlement/cleanup. Existing `none` and legacy missing -configuration create no Environment, and historical internal snapshots are not -backfilled. Creation and recorded-intent retry snapshots load the Environment with -the Session row/cursor in the same transaction, without borrowing subsequent -activity or Turn state. Initial state is `pending`; authenticated connection observations follow -the lifecycle rules below. -Public creation supports `self_hosted` on the three enabled native profiles when -the daemon gateway is configured. The workspace is a clean absolute host path matching the bound Runtime root; -`/workspace` denotes that root through the existing logical mapping. Optional -capability directories are clean absolute host-local selections. Supported non-deferred functions keep their engine-specific -validation and native callback bridge. Enrollment binds the dedicated Runtime; -Session output uses the owned Environment association. Public Files reuse the exact -local workspace; populated self-hosted installation metadata remains unsupported. - -Environment retrieval uses the existing tenant-scoped join to a live owning Session -and its durable connection status, independently of a live Runtime or gateway setup. -It preserves project-shared read access and exposes only the pinned resource fields. -The current closed self-hosted configuration has no API-managed file, plugin or skill -installations, so those required arrays are empty. They are not a filesystem listing -or a claim about native discovery. Reuse the strict Environment configuration parser -and reject unsupported installation fields/capabilities or resource states instead -of treating unknown inventory as empty. No read initiates native work or changes -connection state; connection does not establish readiness or process quiescence. - - -The private Environment input reservation stores one canonical message batch before -Turn admission, with a five-minute deadline from the database clock. It requires -an Environment-bearing Session without active work. Reservation and direct input -paths share the Session lock and retry identity; pending or settled keys cannot -bypass the reservation through direct admission. A pending reservation blocks new -direct batches, including cancellation, while successful earlier retries remain -readable. Promotion commits the original inputs, history, reservation settlement -and execution claim (`queued` to `in_progress`) together; expiration and targeted -cancellation retain the terminal identity. Session deletion -is rejected while input is pending and changes nothing. A terminal reservation retry must not -affect a later reservation or Turn. Evaluate deadlines after acquiring the Session -lock, and return terminal storage outcomes without rolling their transaction back. - -A validated Runtime `failed` response with `preparation_failed` and no Run settles -the pending input immediately with `runtime_preparation_failed`. Before admission, -a rejection of the current `execution_prepare` request with no Run and code -`invalid_configuration`, `unsupported_configuration` or `unsupported_preparation` -settles through that same failure path. It records a -safe Session failure before any Turn exists and releases the input gate; fixing -the local cause allows new input. Transport loss, capacity rejection and -unconfirmed cleanup remain retryable within the original deadline. Core uses -these common control states, never Harness-specific error text. - -Initial messages for a newly created Environment-bearing Session use that same -reservation in the creation transaction, including its connection-action event. -The creation winner alone inserts it; the stream cursor still precedes that -event, and creation retries never re-insert it. A durable initial/later flag -defaults historical rows to later input without inferring origin. Initial expiry -projects a failed Session and safe error before any Turn exists; later expiry -retains idle semantics. Failure -events capture the settled activity and Usage atomically. Late connections and -creation retries cannot reset or replay expired input, and newer work supersedes -old activity without changing its event snapshots. The Environment itself is not -failed by an input deadline. This Store rule covers actual self-hosted and internal -hosted associations; it does not enable hosted providers. None/absent Environment initial input retains -immediate Turn admission. Cancellation/deletion keep their existing semantics. - -Ordinary and streamed public self-hosted creation accept initial text through this -transaction after configuration and new-work lease checks. They return the owned -Environment ID and executor URL while offline, without waiting for admission. -The creation stream's `created` snapshot is the committed JSON 201 projection, -including the connection action; the committed action event then follows from the -creation cursor. A disconnected observer leaves committed input intact; -only the existing Worker prepares, promotes and starts it. Saved-Agent retries with recorded intent -recover before fresh execution admission or source resolution; inline retries keep -their existing resolved-snapshot validation. -Later live subscribers observe only future events and recover history through queries. - -The public self-hosted input profile accepts message-only batches. Under the same -Session lock, recover the original reservation or direct receipt before choosing -current active input or idle reservation. Active messages use existing ordered input -receipts without a new Turn, preparation or reservation; idle messages retain the -readiness and promotion path, including already-connected environments. Pending -reservations keep their gate and deadline. An unlocked activity read or retry after -a conflict must never choose a different admission path. Mixed inputs remain gaps; -these restrictions do not narrow the pinned protocol target. Cancellation-only batches -use the existing direct admission after configuration and execution-ownership checks. -Only the Session-locked transaction chooses the active Turn or an idle receipt; -matching retries retain that target even during later work. Cancellation cannot -create a Turn or bypass preparation. A new cancellation still conflicts with a -pending reservation; it does not cancel pre-Turn input. Its 204 response confirms -durable admission, not native completion or process exit. Homogeneous function-result -batches also use direct admission after those same checks: their explicit Turn/call -identity selects an existing pending call, never new work. Reuse function validation, -Session-locked whole-batch receipts, preserved output/error fields and native -application acknowledgements. Matching retries remain bound to their original calls -after completion or during later work; new results cannot bypass a pending reservation. -Definitions remain fixed through preparation and cold native continuation. These -callbacks are not installed Environment metadata. Mixed result/cancel publication -and exact hosted action-removal timing remain separate gaps. -Promotion requires the current leased execution writer and the caller's -retained native preparation; never hold a database lock during external preparation. Only -the first successful non-replay receipts authorize Start on that same preparation. -An admitted retry returns the original receipts without reclaiming execution; a -read or uncertain commit never authorizes another Start. A crash after promotion -but before Start uses existing claimed-Turn reconciliation (`execution_interrupted`), -including unbound or deleted Sessions, rather than ordinary queued dispatch. Deletion -after claim is rejected like any active Turn. -The Worker expires at most 32 due reservations on each existing tick, after -checking ownership and before checking devices or execution slots. The sweep -requires the leased Store and uses its connection with the existing transaction -timeout; it never falls back to a pooled writer. A partial deadline index and -Session row locks with SKIP LOCKED let unrelated work proceed around contention. -The candidate cutoff is statement time; settlement rechecks the database clock -after acquiring the Session lock. This bounds mutations and transaction time, not -the number of examined locked rows. Restart resumes expiry on normal ticks without -a separate scheduler or backlog-draining loop. No failed Turn may stand in for a -pre-Turn connection failure. - -Session activity before a Turn is derived from the latest relevant reservation and -authenticated connection state. Offline input requests `environment_connection`; -connection arrival clears that action to `idle`, while the prepared Worker still -owns native readiness and admission. An idle offline Environment alone requests no -connection. Reservation/connection changes commit immutable Session activity and -usage snapshots in the same transaction; SSE must not substitute a later Turn or -action set. A newer or active Turn owns subsequent activity. Settled non-initial -reservations clear their action to `idle` until newer work exists, even after an -earlier failed Turn. This local settlement policy does not establish hosted expiry -errors or initial-input asynchronous failure semantics; those remain unverified. - -Session GET/list/metadata responses and live SSE share the safe `self_hosted` -output projection. Its `remote_url` comes only from the daemon gateway -configuration, never request headers or a daemon address. Include -the owned Environment ID, workspace and capability directories without exposing -private configuration. The standalone Environment resource remains separate. -Acceptance must pass that exact URL and ID to the -caller-started executor and observe real remote execution through the existing -Worker, daemon and harness, with fixed SDK and raw HTTP/SSE checks. - -A self-hosted input HTTP request returns 204 only after durable admission. Its -wait uses bounded pooled operations, outside transactions and execution lease -ownership; it cannot prepare or start native work. Only that route extends its -response write deadline to six minutes for the original five-minute database -admission deadline plus response grace. Request/observer disconnect stops waiting, -not the durable reservation or execution; retries keep the original identity and -deadline. The Worker remains the readiness, promotion and Start owner. Local failure -mapping uses 409 `environment_input_expired` / `environment_input_cancelled`, 503 -`execution_unavailable` for ownership loss, and existing 404 for deletion. New input -after a hosted provisioning failure returns the observed 409 `conflict_error`; other -exact hosted failure statuses/bodies and pending-input crash recovery remain unverified. -Principal acceptance must use public Session creation and input against the built -standalone service, including a wait exceeding its ordinary 30-second write timeout, -real remote commands/files and a second native-history Turn. Private provisioning -or injected API handlers cannot substitute for that workflow. +# Environments and Templates + +An Environment is the execution resource of a Session: the machine, workspace and prepared capabilities that a Harness runs in. A Session creates its Environment through its `environment` configuration; there is no standalone create call. An Environment Template is reusable preparation configuration that a Session resolves when it is created. This contract covers both resources, the two placements, input admission, capability preparation, Skills, Plugins and MCP connection origins. + +Related owners: + +- [Environment files](environment-files.md): the Files API on a live workspace. +- [Executor credentials](environment-executor-credentials.md): enrollment, the installation grant and connection status of a `self_hosted` machine. +- [Sandbox deployment](sandbox-deployment.md): which Sandbox Provider (E2B, Docker or microsandbox) hosts `openai_hosted` Environments. +- [Core–Runtime protocol](../../docs/runtime-protocol.md): the `runtime_prepare` transfer and every other wire message. +- [Runtime and outer isolation](../../docs/design-principles.md#runtime-and-outer-isolation): the daemon runs tools with its launching user's permissions; isolation comes from the outer Environment. + +## Resources and states + +Paths follow the SDK resource methods, before the service's `/v1` prefix. The linked pinned source owns the fields and unions. + +| Resource | Operations | Rules | +| --- | --- | --- | +| [Environment](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/environments.py) | `GET /agents/environments/{id}` | Created through Session configuration. Returns `id`, `type`, `status` and the `files`, `plugins` and `skills` recorded at Session creation. | +| [Template](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py) | `POST`, `GET /agents/environments/templates`; `GET`, `POST`, `DELETE /agents/environments/templates/{id}` | See [Templates](#templates). | +| [Files](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/files.py) | `POST`, `GET /agents/environments/{id}/files` | See [Environment files](environment-files.md). | + +An Environment read joins its live owning Session within the caller's Project and returns the durable connection status. It needs no live Runtime, starts no native work and changes no connection state. The installation arrays list API-managed files, Plugins and Skills for both placements: files as `{id, type, path, file_id, size_bytes}` without content, Skills as `{type, name, description, skill_id, version}` and Plugins as `{type, name, description}`. Capabilities discovered in local capability directories are not listed. + +[Session environment input](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/environment_param.py) and [output](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/environment.py) have different shapes: + +- `none` selects no Environment. +- `self_hosted` input requires `workspace_directory`; the nullable `capability_directories` defaults to an empty list. Output adds the Environment ID and the output-only `remote_url`. The output's `/workspace` default does not make the input field optional. +- `openai_hosted` can reference a Template and supply `capability_directories`, `network`, `packages`, `files`, `plugins`, `skills`, `env` and `setup_commands`. + +| Projection | States | +| --- | --- | +| [Environment resource](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agents/environment_info.py) | `pending`, `connected`, `disconnected`, `expired`, `failed` | +| [Session environment event](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agent_session_environment_state.py) | `pending`, `ready`, `connected`, `disconnected`, `failed`, with a nullable error | + +The two vocabularies are separate; never cast one to the other. Session environment events carry Session and Environment identity and an optional Turn identity, and never configuration, credentials, registration IDs or revisions. A Session's `required_actions` can contain `{type: "environment_connection", environment_id}`, separate from function calls. Environment connection, Session status and Turn status are independent. + +Core creates the Environment record in the Session creation transaction; the Session upsert picks the retry winner, and retries never create or repair an Environment. The Environment derives its Project and immutable configuration from the owning Session. Its first state is `pending`. After the Session is deleted, reads hide the Environment while Core keeps the record for settlement and cleanup. A Session with `none` has no Environment. + +A managed outer Environment must exclude broader application credentials and other tenants’ secrets. + +## Placements + +Both placements run the same Runtime: the daemon, the selected Harness, native tools and the workspace run together on one machine. They differ only in who owns that machine. + +| Object | Responsibility | +| --- | --- | +| Session and Environment | Durable ownership, configuration, pending interaction and connection observations (Core) | +| Provider allocation | Compute and filesystem lifetime: Core's Sandbox Provider for `openai_hosted`, the application for `self_hosted` | +| Device and daemon connection | Authenticated Runtime identity and the replaceable dispatch transport | +| Harness process and native session | The native model and tool loop, its execution state and native history | +| Enrollment | The exact Environment, device and executor-key binding of a `self_hosted` machine | + +### Hosted (`openai_hosted`) + +The deployment's configured Sandbox Provider (E2B, Docker or microsandbox, see [sandbox deployment](sandbox-deployment.md)) hosts the Environment. [Harness capabilities](harness-capabilities.md) lists which Harnesses run there. + +- Session creation, with or without initial input, commits the Session, Environment and retry identity before the Worker provisions compute. A creation interrupted before bootstrap is recovered without repeating the Provider's Create. +- Provisioning needs no caller action; the Session stays idle until a Turn starts. +- Omitted or null `network` means enabled; `disabled` and `restricted` are rejected before the Session is created ([restricted network](#restricted-network)). +- A Core restart keeps the allocation, workspace and native identity and never replays uncertain work. +- Terminal cleanup revokes authority and settles pending input in one transaction before the Provider reclaims compute. New input on a terminal Environment is rejected. +- Deleting the Session reclaims its Environment; deleting a Template does not. + +### Self-hosted (`self_hosted`) + +The application owns the machine. It creates the Session with a clean absolute `workspace_directory` and optional absolute local `capability_directories`. Core returns the Environment ID, the `remote_url` and an install command in `x_agents_core.installation`; running that command on the machine installs the daemon and enrolls it ([self-hosted guide](../../docs/getting-started/self-hosted.md), [executor credentials](environment-executor-credentials.md)). + +- `remote_url` is the daemon WebSocket URL derived from Core's public URL, never from request headers or a daemon address. It names Core's private daemon transport. +- Enrollment binds the exact Session, Environment, device and executor key. It creates no allocation and cannot move a Session to another device. +- The Session's workspace must equal the `/workspace` alias or the exact canonical directory the Runtime is bound to. Naming a path grants no access to it. +- The Session carries its own model provider; deployment defaults never apply ([model execution](model-execution.md#saved-defaults-and-precedence)). +- Session reads, lists and events return the `self_hosted` output with the Environment ID, workspace and capability directories, never private configuration. `capability_directories` lists the caller's selections; the Runtime's installation locations stay private. +- Compute, workspace and files stay the application's. Deleting the Session or revoking the credential denies further access but does not stop native processes; the machine owner stops and cleans up. +- The workspace and native history must survive a daemon restart. Losing them never authorizes silent replacement or replay. + +**Application-managed E2B.** An application can run the Runtime in an E2B sandbox it creates, renews and destroys with the E2B SDK, then enroll that Runtime as a `self_hosted` Environment ([E2B Runtime guide](../../services/core/deploy/e2b/README.md)). Core keeps no E2B allocation for it and never renews or kills it. + +### Ownership rules + +- Keep Environment identity, ownership, configuration and lifecycle in Core, separate from Provider compute, device identity, daemon sockets and native sessions. Keep mutable connection state out of immutable configuration; a replacement owner fences stale observations. +- Callers, devices and Environment connections use distinct credentials. The daemon gateway authenticates the enrolled executor key and the exact device. Connection observations keep generation and revision fencing. Registration and connection do not establish readiness. +- Rotation, revocation, Session deletion and loss of ownership deny further access; they do not promise that native effects stop at once. +- Native history stays on the bound Runtime. Preserve it, or demonstrably restore it, across compute replacement; never silently move a bound Session or replay unknown work. +- Durable metadata reads need no live Runtime. Live file reads need an authorized view of the exact workspace and bounded operation ownership. Operations that write, replace or retire a workspace owner also apply mutation fencing. +- When a Session has a started Turn but no recorded native session ID, the next Turn requires existing-history recovery through a verified Runtime capability. A supplied native ID stays authoritative. The Codex adapter recovers only a unique, non-archived root in the Session's private native home and expected working directory, through native listing and exact-ID resume. Missing, incomplete or ambiguous history fails without starting a new root, as does a recorded start that may precede native work. Recovery never replays interrupted input. Device identity and Environment scope belong to `ExecutionDevice`; native session identity and prior Turn state belong to `SessionExecutionBinding`. +- Runtime discovery gives each installed Harness a 15-second version probe; a missing binary fails at once. A version result is availability, not Environment readiness. + +## Input admission + +Session input goes through `POST /agents/sessions/{id}/events` in ordered batches of 1–64 events. On an Environment-bearing Session, Core reserves input that cannot start yet until the Environment is connected and prepared. + +### Initial input + +Creation accepts initial text as a string or an ordered array of user messages. Omitted or null input creates no Turn and no connection action. + +- The creation transaction stores the initial batch as a reservation; for `self_hosted` it also records a connection action. Only the creation winner inserts them; retries never insert them again. +- A `self_hosted` creation returns the Session and the Environment connection target at once, while the machine is still offline, without waiting for admission. A streamed `self_hosted` creation sends the committed JSON projection as its `created` snapshot, already showing `requires_action` and the connection action, then the committed `requires_action` event. +- A fresh creation stream ends right after the idle state recorded when the admitted Turn ends or the reservation stops being pending, or after a failure. The connection alone clearing the action does not end it. A creation without input ends right after its created snapshot; a same-key stream retry ends at once without events. Later subscribers see only future events and recover history through queries. +- Closing the stream leaves committed work intact; the Worker prepares and admits the input. +- If the initial reservation expires, the Session becomes `failed` with a safe error and no actions, and no Turn is created. The Environment itself is not failed. +- Creation retries keep the original identity, deadline and input. Saved-Agent retries with recorded intent recover before execution admission or source resolution. + +### Later input + +| Batch | Behavior | +| --- | --- | +| Messages, Turn active | Appended to the current Turn through ordered admission and native delivery. No new Turn, preparation or reservation. | +| Messages, Session idle | Reserved; the request waits for the Worker to prepare, admit and claim. | +| Cancellation only | Admitted through the ordinary durable cancellation path. An idle cancellation creates no Turn. A new cancellation conflicts while a reservation is pending; it never cancels pre-Turn input. | +| Function results only | Admitted to the named pending calls with the ordinary result receipts, without a Turn or preparation. Cannot bypass a pending reservation. Function results never become Environment installation metadata. | +| Mixed kinds | Rejected. | + +Core returns 202 only after the batch is durably admitted. For a cancellation, 202 confirms admission, not native completion or process exit. Under the Session lock, a retry recovers its original reservation or receipt before Core chooses between the active Turn and a new reservation, and keeps that target during later work. An unlocked read or a retry after a conflict never picks a different path. + +The waiting request uses bounded pooled operations, outside transactions and execution leases; it never prepares or starts native work. Only this route extends its response write deadline, to six minutes, covering the five-minute database deadline plus response time. Disconnecting stops the wait, not the reservation. Errors: + +| Condition | Response | +| --- | --- | +| Reservation expired | 409 `environment_input_expired` | +| Reservation cancelled | 409 `environment_input_cancelled` | +| New input on a failed `openai_hosted` Environment | 409 `conflict_error`, "the hosted environment failed to provision" | +| New input on a failed `self_hosted` or an expired Environment, or input that was waiting when the Environment failed | 409 `environment_unavailable` | +| Execution ownership lost | 503 `execution_unavailable` | +| Session deleted | 404 | + +### Reservations + +A reservation stores one canonical message batch before Turn admission, with a five-minute deadline from the database clock. + +- It requires an Environment-bearing Session without active work. Reservations and direct input share the Session lock and retry identity; a pending or settled key cannot bypass its reservation through direct admission. +- A pending reservation blocks new direct batches, including cancellation, while successful earlier retries stay readable. +- Promotion commits the original inputs, history, reservation settlement and the Turn claim (`queued` to `in_progress`) together. It requires the current leased execution writer and the retained native preparation; no database lock is held during preparation. Only the first successful, non-replay receipts authorize Start on that preparation. An admitted retry returns the original receipts without reclaiming execution; a read or an uncertain commit never authorizes another Start. +- A crash after promotion but before Start goes through claimed-Turn reconciliation (`execution_interrupted`), including for unbound or deleted Sessions. +- Session deletion is rejected while input is pending and changes nothing. Deletion after the claim is rejected like any active Turn. +- Deadlines are checked after taking the Session lock. Expiration and targeted cancellation keep their terminal identity, and a retry of a terminal reservation cannot affect a later reservation or Turn. +- Initial reservations fail the Session when they expire; later ones return the Session to idle. Failure events capture the settled activity and Usage atomically. A late connection or a creation retry cannot reset or replay expired input. + +A Runtime `failed` result with `preparation_failed` and no Run settles the pending input at once with `runtime_preparation_failed`. So does a rejection of the current `execution_prepare` with no Run and code `invalid_configuration`, `unsupported_configuration` or `unsupported_preparation`. Core records a safe Session failure before any Turn exists and releases the input gate; after the local cause is fixed, new input is accepted. Transport loss, capacity rejection and unconfirmed cleanup stay retryable within the original deadline. Core uses these common control states, never Harness-specific error text. + +### Activity and required actions + +Before a Turn exists, Session activity comes from the latest relevant reservation and the connection state. Pending input on an offline `self_hosted` Environment requests `environment_connection`; the connection arriving clears it to `idle`, while the Worker still owns native readiness and admission. An offline Environment without pending input requests nothing, and `openai_hosted` provisioning never requests a connection. A reservation or connection change commits immutable Session activity and usage snapshots in the same transaction; a newer or active Turn owns later activity. Settled non-initial reservations return the action to `idle` until newer work exists. + +### Expiry and scheduling + +The Worker expires at most 32 due reservations per tick, after checking its lease and before checking devices or execution slots. The sweep uses the leased Store connection with its transaction timeout, never a pooled writer. A partial deadline index and `SKIP LOCKED` Session locks let unrelated work proceed. The cutoff is statement time, and settlement rechecks the database clock after taking the Session lock. A restart resumes expiry on normal ticks; there is no separate scheduler. A failed Turn never stands in for a pre-Turn connection failure. + +When next-Turn input is pending on a completed managed allocation that is suspended or recovering, Core hints the managed Runtime maintenance loop. Initial input, cold creation, running or disabled compute, terminal receipts, cancellation and function results, history and file operations send no hint. Hints are best effort and coalesced without blocking; they add at most one scan per normal five-second cycle and never bypass capacity, a busy lifecycle gate or ownership checks. Persisted work and the normal ticker stay authoritative. ## Runtime capability preparation -This section is the single owner of capability preparation: initial files, tool -configuration, Skills, Plugins, Environment MCP, npm/Python packages, setup and -`packages.system`. Hosted and self-hosted Environments use one Runtime path. +Preparation installs what a Session selected: initial files, tool configuration, Skills, Plugins, Environment MCP, npm and Python packages and setup commands. Both placements use one Runtime path. ### Ownership and lifetimes -Resource management retains provider placement, capacity, allocation and -Environment create/renew/reclaim. The authenticated Runtime connection carries -initialization, capability preparation and executor operations; it never -implicitly allocates or destroys compute. +Resource management keeps Provider placement, capacity, allocation and Environment create, renew and reclaim. The authenticated Runtime connection carries initialization, capability preparation and executor operations; it never allocates or destroys compute. Providers never run Core initialization commands. -Executor close, Turn cancellation and transport loss preserve the installed -snapshot, workspace and allocation. Reclamation is an explicit operation -coordinated with active work; a disconnected socket is not proof that native -effects have stopped. See -[Executor and Turn lifetimes](../../docs/runtime-protocol.md#executor-and-turn-lifetimes). +Closing an Executor, cancelling a Turn or losing the transport keeps the installed snapshot, workspace and allocation. Reclamation is an explicit operation coordinated with active work; a disconnected socket does not prove that native effects stopped. [Harness onboarding](harness-onboarding.md#executor-and-turn-lifetimes) owns the Executor and Turn lifetimes. ### Preparation order -Core freezes resource versions, metadata and source selections. Environment -initialization runs in this order, every step over the common `runtime_prepare` -exchange with the same daemon: +Core freezes resource versions, metadata and source selections at Session creation. Initialization then runs in this order, each step over `runtime_prepare` with the same daemon: 1. initial files and tool configuration; -2. Skill/Plugin bundle import; -3. npm/Python packages and ordered setup commands; -4. capability directory snapshot finalization. - -Providers only place, create, bootstrap, inspect, renew and reclaim resources; -they never execute Core initialization commands. The common runner uses only -neutral Environment/Session identity and a Runtime peer, with no Provider, -deployment or OS branch. Harness differences belong to native adapters. - -A missing authenticated Runtime connection waits before initialization is claimed. -The Environment owns durable `pending`, `running`, `complete` or `failed` -initialization state. The leased Worker processes both locations with the same -bounded initializer, independently of Provider maintenance. A running operation -whose process-local owner is lost is failed as unconfirmed; setup is never replayed. -A confirmed failure records only the safe step and exit status. Both failure paths -settle pending input through the common Environment failure transaction. Failure -alone does not destroy compute or remove a workspace. - -The pinned self-hosted input retains `workspace_directory` and optional local -`capability_directories`. For portable preparation, applications may supply -`x_agents_core.environment` at Session creation in either location. This is a Core -extension, not an upstream self-hosted field. It accepts `environment_template_id`, -`files`, `env`, `packages`, `setup_commands`, `skills`, `plugins` and -`capability_directories`, with the same parsers and inheritance rules as hosted -input. A field supplied in both places rejects, including explicit null; there is -no hidden precedence. Machine location, sizing and network policy are not extension -preparation fields. A self-hosted selection cannot use a template requiring a -managed network restriction. +2. Skill and Plugin bundle import; +3. npm and Python packages, then setup commands in order; +4. snapshot of the capability directories. + +The runner uses only neutral Environment and Session identity and a Runtime peer, with no Provider, deployment or operating-system branch. Harness differences stay in the adapters. + +**Portable preparation input.** `self_hosted` input carries only `workspace_directory` and `capability_directories`. A Session with either workspace placement can also send `x_agents_core.environment`, a Core extension with `environment_template_id`, `files`, `env`, `packages`, `setup_commands`, `skills`, `plugins` and `capability_directories`. It uses the same parsers and inheritance rules as `openai_hosted` input. A field supplied both there and in `environment` is rejected, including an explicit null. A `self_hosted` Session names its Template only through this extension and cannot use a Template whose network is not `enabled` (400, param `x_agents_core.environment.environment_template_id`). Machine location, sizing and network policy are not extension fields. ```json { "agent_id": "agent_example", - "environment": { - "type": "self_hosted", - "workspace_directory": "/home/user/project" - }, - "x_agents_core": { - "environment": { - "environment_template_id": "env_template_example" - } - } + "environment": {"type": "self_hosted", "workspace_directory": "/home/user/project"}, + "x_agents_core": {"environment": {"environment_template_id": "env_template_example"}} } ``` -Switching `environment` to `{"type":"openai_hosted"}` reuses that same -preparation input. Resource resolution, Project authorization, concrete Skill -versions, encrypted file contents and confidential tool variables are frozen at -Session creation. Same-intent retries and reconnects reuse those snapshots; -new Sessions resolve new versions. User-managed Environments never require an -allocation record. Local directory discovery does not populate API-managed -installation arrays; managed Skills, Plugins and initial files do, in either -location. Session self-hosted responses retain their pinned connection shape; -the Environment resource exposes the safe installation metadata. - -Transport `connected` is still a connection observation, not readiness. Execution -and live file access wait for initialization, then the common native preparation -owner validates the installed snapshot and Harness before admitting a Turn. The -model-provider authority rule is unchanged: deployment credentials are not -implicitly sent to user-owned machines. - -### Transfer and paths - -The private `runtime_prepare`/`runtime_prepare_result` exchange carries typed -initial files, configure/npm/python/setup operations, inert Skill/Plugin -archives and finalization selections, with canonical Session and Environment -identities. - -- Files and setup working directories use logical `/workspace` addresses. - Setup commands are explicit typed input; executable selection and physical - destinations stay Runtime-owned. Core supplies no executable or host-platform field. -- Ordered 64 KiB chunks and SHA-256 receipts bound each file or archive to - 50 MiB, each control frame to 1 MiB and each connection to one transfer. - Initialization and finalization headers carry no file data. -- Begin, chunk and commit are never retried. Rejected, failed and unknown effects - stay distinct. Runtime owns transfer and filesystem work through settlement, - including disconnect or cancellation. -- Source selections accept portable absolute Unix, Windows drive and UNC paths. - Core never resolves them on its own host; the daemon applies its local path and - access checks. +Changing `environment` to `{"type":"openai_hosted"}` reuses the same preparation input. Resource resolution, Project authorization, concrete Skill versions, encrypted file contents and confidential tool variables freeze at Session creation. Retries and reconnects reuse those snapshots; new Sessions resolve new versions. A `self_hosted` Environment never needs an allocation record. + +**Readiness.** Transport `connected` is a connection observation, not readiness. Execution and live file access wait for initialization; then the native preparation owner validates the installed snapshot and Harness before admitting a Turn. File reads keep their own readiness and authorization and do not require capability or native readiness. Deployment model credentials are never sent to application-owned machines. + +**Transfer.** Initial files, configure, npm, Python and setup operations, inert Skill and Plugin archives and finalization selections travel as typed `runtime_prepare` operations with canonical Session and Environment identities; the [protocol](../../docs/runtime-protocol.md#preparation-and-execution-order) owns chunking and receipts. Files and setup working directories use logical `/workspace` addresses; the Runtime chooses executables and physical destinations, and Core supplies no executable or host-platform field. Source selections accept portable absolute Unix, Windows drive and UNC paths; Core never resolves them on its own host, and the daemon applies its local path and access checks. ### Installed snapshot -Both origins use the common Runtime parser and `installed.json` manifest on -Linux, macOS and Windows. The operator chooses the capability root (by default -`capabilities` under `OAC_RUNTIME_HOME`); Core and transfer requests cannot. - -- The manifest binds the Session and Environment to the ordered source-selection - digest. The admitted asynchronous preparation owner verifies or creates it - before native execution, including for an empty selection. -- Filesystem locking prevents concurrent installation. A private completion - record keeps only the operator installation root, so a deleted snapshot is - never mistaken for first preparation or recaptured. -- Missing, partial, conflicting or foreign snapshots fail without deleting data, - automatic repair or replay. -- Reconnection and replacement Executors load installed contents without - rereading sources. Source edits are seen only by a new Session with its own - Environment and fresh snapshot. -- Read-only snapshot modes are integrity hints, not protection from the - launching user. - -### Executor admission - -Synchronous Executor admission validates the frozen descriptor only. The -asynchronous preparation owner ensures capabilities are ready before invoking the -native factory. Adapters receive only the resolved Runtime-owned Skill paths and -MCP declarations; these remain Runtime-only fields. Reused Session Executors -keep their original configuration. - -Files reads keep their separate readiness and authorization and do not require -capability or native execution readiness. MCP environment variables come only -from explicit immutable Runtime tool configuration; missing variables never fall -back to daemon credentials or ambient process variables. +Both origins use the common Runtime parser and an `installed.json` manifest on Linux, macOS and Windows. The Runtime operator chooses the capability root ([installer options](../../docs/getting-started/self-hosted.md#options-for-automation)); Core and transfer requests cannot. + +- The manifest binds the Session and Environment to the ordered source-selection digest. The preparation owner verifies or creates it before native execution, also for an empty selection. +- After the setup commands, the initializer snapshots the declared workspace-contained capability directories into Runtime storage. Directory bytes are read after setup, not at Session creation. +- A filesystem lock prevents concurrent installation. A private completion record keeps only the operator's installation root, so a deleted snapshot is never mistaken for a first preparation or captured again. +- Missing, partial, conflicting or foreign snapshots fail without deleting data, repairing or replaying. +- Reconnecting and replacement Executors load the installed contents without rereading sources. Source edits reach only a new Session. +- Recursive references to the snapshot and directory entries that escape it are rejected. +- Read-only snapshot modes are integrity hints, not protection from the launching user. + +Executor admission validates only the frozen descriptor. The preparation owner makes the capabilities ready before calling the native factory; adapters receive only the resolved Runtime-owned Skill paths and MCP declarations. A reused Executor keeps its original configuration. ### System dependencies and Runtime directories -The daemon runs as its launching account and never uses sudo or elevates its -permissions. Only user-directory dependencies install during preparation. - -- System dependencies must be preinstalled in the managed image/template or by - the self-hosted user. -- `packages.system` is rejected explicitly, in Templates and inline - configuration, including a supplied null or empty list. It is never ignored or - translated into a privileged operation. -- Runtime runs no apt, sudo, unprivileged system-root installer or other - automatic system-package operation. No managed-only exception exists. Missing - dependencies fail the consuming operation. -- npm and Python use local prefix/target directories. Setup requires Bash; - Windows requires Git Bash, without a substitute shell. -- Package responses keep the official required `system: []` field as empty - response metadata; it is not stored as an initialization option. - -Initialization and package directories default under `OAC_RUNTIME_HOME` and may -be selected with `OAC_RUNTIME_INITIALIZATION_DIRECTORY` and -`OAC_RUNTIME_PACKAGE_DIRECTORY`. Packaged Linux images select their existing -`/environment` layout through those settings. `OAC_RUNTIME_TOOL_ENV_FILE` -selects explicit tool configuration. These are resource paths, never -Environment-source or OS switches in Core. - -Runtime process ownership waits for exit and I/O settlement. Confirmed failures -keep only a bounded exit status, never command output. - -### Platforms and isolation - -The protocol and Go preparation implementation are shared by all three native -platforms. Managed Providers remain Linux-only. Commands run directly with the -starting account's permissions and host network; the daemon does not sandbox -tools, files or network access. Outer Environments own managed isolation, and -unsupported network restrictions reject instead of silently running unrestricted. - -Native installation and validation limits are in the -[self-hosted guide](../../docs/getting-started/self-hosted.md#platforms). Historical acceptance evidence -stays limited to its recorded binaries and inputs. - -On Windows, npm package installation and stdio MCP commands named `npm` or `npx` -(including their `.cmd` shims) run through the resolved npm installation's -JavaScript entrypoint with Node, without an extra shell. - -### Environment initialization and compute wake - -Environment initialization has pending, running, complete and failed states. -Authentication and connection publication are independent of preparation. Execution -bindings and live Files wait for initialization. The leased Worker's common -scheduler scans 32 Environments at a time, wraps at EOF and bounds concurrent -preparations by execution concurrency. Missing sockets do not consume a pending -attempt; unavailable Harnesses fail before installation. Each operation rechecks -current authority and the original socket. Completion rechecks the exact binding. -Provider bootstrap and resource cleanup retain their own owners and settlement. - -After a next-Turn input is durably pending, a completed managed allocation in a -suspension/recovery phase may hint this loop. Initial inputs, cold creation, -running/disabled compute, terminal receipts, cancellation/tool-result events, -history and file operations do not use this hint. Eligibility lookup and delivery -are best effort; persisted work and the normal ticker remain authoritative. -Coalesce hints without blocking, and allow at most one extra scan per normal -five-second cycle. Keep the ticker independent of requests. A normal tick consumes -already queued hints before scanning; simultaneous tick/hint readiness is one -normal scan. Preserve hints arriving during a scan, the allocation cursor and all -ownership checks. Never close the hint channel while handlers may still send. -This bounds extra maintenance work but does not bypass capacity, a busy lifecycle -gate or multi-page scheduling, and does not guarantee a resume deadline. - -A recovered or uncertain running installation fails without replaying writes. -Preparation failure is terminal for its Session. It does not destroy compute or -user files; explicit resource cleanup retains its existing owner. The Environment -transaction stores a safe reason with the failed Environment and records -`environment.failed`, `error` and one `agent.session.failed`; Session reads derive -`failed`, that reason and the failure time from the same record, and live streams end -after the failed event. The initializer reports only the integer exit status of a -failed initialization command. Core composes the reason from a fixed step label and that status, -never from command, package-manager or file output; unknown effects, timeouts and -receipts without a status keep a generic reason. New input then gets the observed 409 -`conflict_error`; expiry and pending-input settlement keep their behavior. -Completed environments never reinstall initial files on reconnect or native recovery. -The typed Runtime preparation protocol carries bounded confidential input; the -daemon owns the common Go initialization implementation. Confidential env -and setup snapshots are encrypted independently of ordinary metadata. Adapters -apply explicit tool variables without automatically inheriting daemon credentials; -this does not prevent same-user tools from reading local Runtime state. -Reuse runtimefs atomic replacement and anchored logical-workspace paths for initial -files on every platform. Public Files creation keeps its own non-replacement rule. - -### Skills, Plugins and Environment MCP - -Skills and their immutable versions are Core-owned tenant resources, independent of -Sessions and native Skill installations. Serialize version allocation and pointer -mutations under the owning Skill row. Keep top-level name and description aligned -with the default version in the same transaction, including default-changing uploads; -nondefault uploads preserve that metadata. Preserve unique version identities across -concurrent uploads and deletion. Metadata reads never load or decrypt bundle bytes. -Encrypt bundle contents with a tenant, Skill and version binding using the existing -service cipher. Deleting a Skill reclaims its versions without affecting already -frozen Session initialization. Public reference metadata, unresolved template intent -and the resolved Runtime bundle are distinct; do not report a reference as inline -merely because it reuses the same installer. No compatibility reader, source -cache, extra lifecycle owner or per-harness resource implementation is required. -Resolve references inside the Session creation transaction, after the creation -upsert establishes ownership. Lock referenced resources in a stable order; freeze -the selected version, descriptive metadata and bytes together. Creation retries -recover the recorded intent before reading mutable templates or Skill sources. -Templates preserve default, latest and explicit version selectors. An omitted or -null reference version selects the default at Session creation and projects as -`version: null` in Template responses. Session responses contain concrete versions; -only validated installation metadata crosses the Runtime boundary. A supplied -Session Skill, Plugin or capability-directory list replaces its template list; -omission and null inherit, while an empty list clears that selection. This differs -from Template resource updates, where null clears lists and resets network to the -pinned enabled default. Preserve caller intent and frozen Session snapshots in -both cases. Public capability directories remain caller paths; adapter-owned -installation directories are not portable public paths. - -Inline and referenced Skill ZIPs use the same confidential initialization snapshot and installer. -Core validates portable manifests and bounded regular-file archives, returns only -safe Skill metadata, and freezes content before native preparation. The Runtime -owns `skills/` below its configured capability root; the manifest validates -the installed snapshot before reuse. This is state -consistency, not protection from the launching user. Public Plugin ZIPs preserve their complete package -layout and reuse the shared archive and portable Skill parsers. Core keeps safe -Plugin metadata separate from encrypted archives. Templates inherit or replace -Plugin and capability-directory lists through the same hosted resolver. - -After ordered setup, the existing initializer snapshots declared workspace-contained -capability directories into Runtime storage and writes one installed manifest. -This is an initialization artifact, not a second lifecycle owner or database ledger. -Directory bytes are observed after setup; they are not frozen at Session creation. -The common daemon resolves the manifest only for executable preparation and passes -validated Runtime-owned Skill/package roots to adapters. Files reads do not require -that artifact. Reconnection and recovery read installed bytes, never mutable source -directories. Missing or inconsistent installations fail preparation without replay. - -Adapters register only selected Skill roots without changing the execution loop. -Codex uses explicit extra roots; MiniMax projects its native catalog; Claude creates -one controlled envelope per package with real directories and immutable hardlinks -under content. Its explicit paths remain inside that envelope; original native -control files are not activated. Keep automatic native MCP discovery disabled. -Explicit Environment MCP declarations -follow the separately qualified transport path below; unsupported native activation -fails explicitly, without silent partial activation or a generic plugin framework. - -Environment-origin MCP declarations use the shared Plugin parser and frozen -installed packages. The installation manifest retains selected MCP package roots; -Runtime re-parses those installed packages without another configuration copy, -credential cache or lifecycle ledger. Native adapters must explicitly qualify and -project supported declarations before public admission. Parent-directory Skill -discovery does not activate nested Plugin MCP configuration. - -The common daemon stdio entry resolves the installed declaration and launches its -server under the same user permissions as the Harness. Unix replaces the helper -process; Windows forwards native stdio within the owned process tree. Explicit -initialized values override the selected declaration's variables. There is no -Python sandbox launcher, mount policy, shell prefix or separate credential sandbox. -Process groups and Windows Jobs own cancellation and descendant cleanup only. -Claude composes MCP identity and observation with the same workspace profile; -MiniMax uses the same effective bindings for stdio and HTTP. A public capability -advertisement must still match the installed engine's actual supported transport. - -Runtime resolves public HTTP declarations and installed Plugin MCP through -`agent.ResolveMCPBindings` before adapter projection. Each transient binding -retains its connection origin, transport, nullable tool allowlist, required flag, -credential authority and installed stdio identity. Bindings are never persisted -or logged. Duplicate identities and unavailable selected credentials reject. -Public HTTP MCP uses the same binding path as installed Plugin MCP. Its explicit -`connection_origin` is retained from saved configuration through the Session -snapshot and private Runtime request; the exact Runtime wire version is required. -A service-origin request cannot silently become an Environment-origin connection. +The daemon runs as its launching account and never uses sudo or raises its permissions. Only user-directory dependencies install during preparation. -### Public MCP connection origin +- System dependencies must be preinstalled in the managed image or by the owner of a self-hosted machine. A missing executable or library fails the operation that needs it. +- `packages.system` is rejected in Templates and inline configuration, including a null or empty list (400, param `packages.system`). Package responses still carry the official required `system: []`. +- npm installs into a local prefix and Python/pip into a local target under the Runtime package directory; Node/npm and Python/pip must already be installed. Their dependencies are visible to native tools in every working directory. +- Setup commands run with Bash; on Windows, Git Bash is required and no other shell substitutes. The default working directory is `/workspace`. +- On Windows, npm installation and stdio MCP commands named `npm` or `npx` (including their `.cmd` shims) run through npm's JavaScript entry point with Node, without an extra shell. + +Initialization and package directories default to `initialization` and `packages` under the Runtime home (`OAC_RUNTIME_HOME`) and can be set with `OAC_RUNTIME_INITIALIZATION_DIRECTORY` and `OAC_RUNTIME_PACKAGE_DIRECTORY`; packaged Linux images use `/environment/initialization` and `/environment/packages`. These are resource paths, never Environment-source or operating-system switches in Core. + +Every command uses the launching user's permissions and the host network. Process ownership waits for exit and I/O settlement. Command output is discarded; a confirmed failure keeps only a bounded integer exit status. + +### Explicit local tool environment + +The installer's `--tool-env-file` (`OAC_RUNTIME_TOOL_ENV_FILE`) supplies the Runtime operator's base tool variables. Preparation copies these values into its private initialization snapshot, and the Session's `env` keys override them. The Runtime never rewrites the source file or inherits unrelated ambient credentials. Setup, capability resolution and Harness execution read the same prepared snapshot. Reconnecting keeps that snapshot even if the operator edits the file; a new Session reads the current file. A Harness profile may reference the Runtime-owned file but must not persist copies of its values. A missing or invalid configured file fails preparation. + +### Initialization state and failure + +An Environment's initialization is `pending`, `running`, `complete` or `failed`, independent of any allocation and of authentication and connection publication. + +- The Worker's initialization scheduler scans 32 Environments at a time, wraps at the end and bounds concurrent preparations by execution concurrency, independently of Provider maintenance. +- A missing socket does not consume a pending attempt. An unavailable Harness fails before installation. Each operation rechecks current authority and the original socket; completion rechecks the exact binding. +- Each file transfer, configure, Skill, Plugin, package and setup step has a two-minute budget; the whole initialization has 30 minutes. Initial input keeps its five-minute admission deadline, so large installations should start from an idle Session. +- A running initialization whose owner is lost, including across a Core restart, fails as unconfirmed; nothing is replayed. A completed Environment never reinstalls on reconnect or native recovery, so later user changes survive. +- Failure is terminal for the Session but destroys neither compute nor files. + +On failure, one transaction marks the Environment failed and records `agent.session.environment.failed`, an `error` event and one `agent.session.failed`. Session reads return `failed`, the reason as `error` and the failure time as `last_active_at`; live streams end after the failed event. Pending input settles as failed. The reason names only the step and its exit status: + +| Failed step | Reason | +| --- | --- | +| Setup command `i`, confirmed exit 1–255 | `Failed to provision environment: script "setup_commands[i]" failed with exit code N` | +| Python packages, confirmed exit 1–255 | `Failed to provision environment: script "Python package installation" failed with exit code N` | +| npm packages, confirmed exit 1–255 | `Failed to provision environment: script "npm package installation" failed with exit code N` | +| Initial file installation, confirmed failure | `Failed to provision environment: initial file installation failed` | +| Skill preparation, confirmed failure | `Failed to provision environment: Skill installation failed` | +| Harness not installed on the Runtime | `Failed to prepare environment: the selected Harness is unavailable. Install the supported Harness version on the Runtime and create a new Session.` | +| Anything else: timeouts, unknown effects, missing or malformed receipts, Plugin installation, snapshot finalization, bootstrap rejection, Core restart | `Failed to provision environment: initialization did not complete` | + +Every initialization operation returns a typed `rejected`, `failed` or `unknown` outcome, and the daemon confirms process exit and I/O settlement first. Core composes the reason from a fixed label and integers, so commands, env values, package names, paths and process output never reach the reason, events, logs or responses. The failed step is not retried and later steps do not run. Confidential env and setup snapshots are encrypted separately from ordinary metadata. Initial files use the atomic replacing writer and anchored workspace paths on every platform; Files API creation keeps its own no-overwrite rule. + +## Templates + +A Template is Project-owned configuration for `openai_hosted` Sessions and for `x_agents_core.environment`. It holds no running workspace and is unrelated to Provider images such as E2B templates. Every Session that references it gets its own Environment through the same preparation as inline configuration. Template parsing, storage and resolution never select a Harness or Provider or depend on native tool names or private Harness paths, so a new Harness or Provider needs no Template change. + +```python +from openai import OpenAI + +client = OpenAI() # reads OPENAI_BASE_URL and OPENAI_API_KEY +template = client.beta.agents.environments.templates.create( + name="Python workspace", network={"access": "enabled"}, + env={"APP_MODE": "analysis"}, packages={"python": ["packaging==26.0"]}, + setup_commands=[{"command": "mkdir -p /workspace/outputs"}], +) +session = client.beta.agents.sessions.create( + agent={"model": "your-configured-model"}, + environment={"type": "openai_hosted", "environment_template_id": template.id}, + input="Create /workspace/outputs/report.txt containing the result of 6 * 7.", +) +``` + +### Operations -Core's Harness profile declares `MCPOrigins`; shared admission and dispatch check -the origin against the Environment and Runtime's advertised HTTP/bearer/required -capabilities. Runtime validates the same origin before invoking an adapter. No -Harness-name or Sandbox Provider branch selects a different connection path. - -| Harness | `service` origin | `environment` origin | Optional policy | -| --- | --- | --- | --- | -| Codex | Service execution host, `environment:none` | Managed or self-hosted workspace | Nullable tool allowlist and required initialization | -| Claude SDK | Service execution host, `environment:none` | Managed or self-hosted workspace; packaged `workspace_mcp_http` feature required | Nullable tool allowlist and required initialization | -| MiniMax Code | Unsupported | Managed or self-hosted workspace | `allowed_tools` must be null/omitted; `required` must be false | - -Both origins support the declared Harness's anonymous HTTP and selected HTTPS -bearer path. The existing attached-Vault selection freezes credential identity, -including a unique implicit URL match or an anonymous selection. Only that -Project-authorized credential may enter the transient Runtime request; Core -defaults and unrelated Vaults are not searched. Decryption failure or a missing -credential fails execution without an anonymous fallback. Public Environment -MCP retains `project_vault` authority; Plugin credentials retain -`environment_configuration` authority. Neither source overrides duplicate -server labels. Bearers never enter persisted native configuration or argv. - -Omitted/null origin still means `service`, including on self-hosted requests; -it does not select the local network automatically. Service-origin requests with -a workspace remain rejected because they require separate service-side connection -forwarding. This implementation adds no proxy. Environment origin requires an -initialized workspace with enabled network access and is invalid on `none`. - -Null/omitted `allowed_tools` permits all server tools; an empty list permits none. -MiniMax rejects every non-null allowlist, including an empty list, rather than -silently expanding it. Codex and Claude preserve native allowlists and initialize -required servers before releasing native input, including cold recovery. -Public MCP with native Subagents remains unqualified. Nonempty literal HTTP -headers, request metadata and public stdio declarations remain unsupported. -See [public MCP qualification](https://github.com/MiniMax-AI/OpenAgentCore/blob/e974a7f880a2eb799f0dd39e6ba0870462854a53/contracts/agents-api/public-mcp-qualification.md) for actual model, -platform and infrastructure coverage; admission support is not a claim of -complete cross-platform/provider qualification. - -Environment-origin literal HTTP headers remain rejected for the pinned Claude -and MiniMax clients because their cross-origin forwarding cannot preserve header -authority. MiniMax supports installed stdio and HTTP servers with anonymous or -explicit user-selected HTTPS bearer authentication. ACP HTTP declarations remain -Session-local native memory; tokens do not enter native configuration files or -process arguments. Required initialization and tool allowlists are not exposed -through the Plugin manifest. Public MiniMax HTTP uses the same transient ACP map -with the stricter admission limits above. -MiniMax reads the existing Session-private native runtime-name registry for exact -first-frame identities and cross-checks completed native results for both transports. -Reuse existing observation and cancellation settlement; never fabricate a delayed -start event, guess normalized identities or add a registry of our own. - -See [capability qualification](environment-capabilities-qualification.md) for -real model evidence and the remaining public-origin and image gaps. - -## Contract inventory - -Paths below follow the SDK resource methods, before the service's `/v1` prefix. -The authoritative fields and unions are linked to the pinned source; this table -is an inventory, not a replacement schema. - -| Resource | Operations | Contract distinctions | +- Every operation needs a Project API key and `OpenAI-Beta: agents=v1`. Template operations work without an execution deployment and allocate no compute. +- `name` is optional and nullable, kept verbatim, 1–256 Unicode characters. +- `network.access` is `enabled`, `disabled` or `restricted`; omitted or null means enabled. See [Restricted network](#restricted-network). +- Responses carry safe metadata and never `env`, `setup_commands` bodies or inline file data. +- List uses `after`, `limit` (default 20; 0 is treated as 1 and values above 100 as 100) and `order` (default `desc`), ordered by creation time and ID. Missing and foreign Template IDs and cursors return the same 404. +- Update: an omitted field keeps its value and a supplied field replaces it. Null clears `name` and every list and resets `network` to enabled. +- Writes and Session resolution that seal or open confidential content (files, env, setup commands, Skills, Plugins) need Core's [credential key](../../docs/configuration.md#installation-directory); metadata reads do not. +- A Session resolves `environment_template_id` within its Project once, at creation, freezes the effective configuration and never passes the Template ID to the Provider or Runtime. Updating or deleting a Template never changes an existing Session. Creation retries recover the recorded caller intent before reading the Template, even after it is deleted; a changed intent conflicts. + +### Inheritance + +The Session applies the Template first, then its own fields. Composition happens once, before the encrypted Session snapshot is written, and the result is revalidated with the ordinary validators. Caller intent (omitted, null or explicit) is kept separately for the retry policy. + +| Field | Omitted or null in the Session | Supplied in the Session | | --- | --- | --- | -| [Environment](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/environments.py) | `GET /agents/environments/{id}` | Created through Session configuration, with no standalone create/list/update/delete method in this resource. Safe metadata includes files, plugins, skills, type and status. | -| [Template](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/templates.py) | `POST`, `GET /agents/environments/templates`; `GET`, `POST`, `DELETE /agents/environments/templates/{id}` | Reusable hosted configuration, resolved for each Session. List uses `after`, `limit` and `order`. Supplied update fields replace their value; omitted fields stay unchanged. Deletion includes confidential inputs. | -| [Files](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/resources/beta/agents/environments/files.py) | `POST`, `GET /agents/environments/{id}/files` | Create accepts `file_id` or inline base64 data with an absolute path inside `/workspace`. List uses opaque `page`, not `after`, with stable path/order/limit across pages. | - -These are eight operations, separate from Session creation and live events. -Templates use an `after` cursor and limit default 20, clamped to 1–100; file listing has nullable -limit/path, non-null order/page when supplied, and case-sensitive path-component -ordering. Both default to descending order. Do not reuse cursor decoding merely -because both endpoints paginate. - -[Session environment input](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/environment_param.py) -and [output](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/environment.py) -have different shapes: - -- `none` selects no execution environment. -- `self_hosted` input requires `type` and `workspace_directory`; its optional - nullable `capability_directories` defaults to an empty list. `remote_url` is - output-only. The output also includes the Environment ID and capability paths. - The output description's `/workspace` default does not make the input field - optional. -- `openai_hosted` can reference a template and supply capability paths, network, - packages, files, plugins, skills, environment variables and setup commands. - Omitted template-backed values inherit; Session overrides cannot broaden the - template network policy. The service implementing this discriminator owns - provisioning; it does not rename the public discriminator for its provider. - -[Template responses](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agents/environments/environment_template.py) -expose safe metadata, retaining unresolved skill version selectors and file -references. They do not return inline file/archive contents, environment variables -or setup command bodies. Nullability and replacement behavior must be checked -against each request type, not inferred from these response models. - -| Projection | State vocabulary | +| `network` | Inherit the whole Template policy | Must narrow it: enabled can become restricted or disabled; restricted can become a subset of its hosts or disabled; disabled cannot widen | +| `env` | Inherit the Template keys | Overlay by key; the Session value wins; `{}` keeps all Template keys | +| `setup_commands` | Inherit the sequence | Replace it; `[]` clears | +| `files` | Inherit the file set | Replace the whole set; `[]` clears | +| `packages` | Inherit both managers | Resolve `python` and `npm` separately: an omitted or null manager inherits, a list replaces, `[]` clears; `system` is rejected | +| `skills`, `plugins`, `capability_directories` | Inherit the list | Replace it; `[]` clears | + +For comparison, an inline `openai_hosted` Session without a Template treats omitted and null `network` as enabled. + +### Restricted network + +`restricted` requires 1–100 exact ASCII hostnames in `allowed_domains`; subdomains and redirect targets need their own entries. Wildcards, URL or port syntax, IP literals, Unicode and trailing dots are rejected. Reads return the supplied spelling, order and duplicates; comparison uses a separate lowercase, deduplicated copy. + +The Runtime does not enforce `disabled` or `restricted`, and no Provider enforces them for it. Core therefore stores these policies in Templates but rejects any Session whose effective network is not `enabled`, before allocating compute. Initialization and tools use the host's existing network. + +### Env and setup commands + +Env values are readable by Agent code but never appear in public metadata or initialization diagnostics. Names must match `^[A-Za-z_][A-Za-z0-9_]*$` and values cannot contain NUL. Core reserves `PATH`, `OPENAI_API_KEY` and every name starting with `OAC_` or `CODEX_`. Files and packages are installed before setup commands; a nonzero setup command fails initialization; no command is retried after unknown effects; completed setup never runs again on reconnect. + +### Initial files + +`files` entries place a file at an absolute destination inside `/workspace`, from inline standard-base64 `data` or from a Project-owned uploaded `file_id`. + +| Limit | Value | | --- | --- | -| [Environment resource](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agents/environment_info.py) | `pending`, `connected`, `disconnected`, `expired`, `failed` | -| [Session environment event state](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agent_session_environment_state.py) | `pending`, `ready`, `connected`, `disconnected`, `failed`; nullable error | +| Files per configuration | 50 | +| Inline file | 5 MiB | +| All inline content | 10 MiB | +| Referenced file | 50 MiB | +| Session or Template request body | 16 MiB | -These projections cannot share an unchecked string cast. Session environment -notifications carry Session/Environment identity and optional Turn identity. -The Session's `required_actions` union includes `environment_connection` with an -Environment ID, separately from function calls. Environment readiness is distinct -from Session and Turn status. +Paths must be canonical, distinct and inside the logical workspace; the Runtime anchors each write to its bound workspace. This is API path scope, not a restriction on native tools running as the same user. Template metadata shows inline files as type, path and size and references as type, path and `file_id`; each Session gets fresh file IDs and sizes for both. File data stays out of ordinary configuration, responses, events and command arguments. A Template keeps references; each Session authorizes and freezes its own encrypted source bytes, so later source deletion cannot change them. -## Ownership and placement decision +### Skills -The canonical [architecture rules](#ownership-and-placement-decision) -keep Environment lifecycle common while leaving process placement and native -transport to adapters. Logical ownership does not require a machine per object. +Templates and inline configuration accept Project-owned Skill references and inline Skill ZIPs. Upload a directory through the pinned SDK, then reference its default version: -| Object | Responsibility | +```python +skill = client.skills.create(files=[ + ("report/SKILL.md", b"---\nname: report\ndescription: Create the report.\n---\nFollow the report procedure.", "text/markdown"), +]) +template = client.beta.agents.environments.templates.create( + skills=[{"type": "skill_reference", "skill_id": skill.id}], +) +``` + +The pinned `/v1/skills` resource, version and content routes use the Project API key without the Agents beta header. ZIP uploads use `files` and directory uploads repeated `files[]`. The pinned SDK 3.13.0 drops a single file tuple during multipart extraction, so upload a single ZIP with raw HTTP. An upload holds at most 500 regular files and exactly one `SKILL.md`, 5 MiB compressed and 20 MiB expanded. [File resource semantics](file-resource-semantics.md) owns default-version and deletion rules. + +A reference with an omitted or null version selects the default at Session creation, `"latest"` the latest version, and a positive version string that version. Template responses keep the unresolved selector (`version: null` for the default); Session metadata shows `{type, skill_id, version, name, description}` with a concrete version. A Session freezes the selected version's bytes and metadata in its creation transaction; later default changes, source deletion or Template updates cannot change it. + +An inline Skill carries `name`, `description` and a base64 ZIP `source` (`media_type` `application/zip`). The archive has one top-level folder with `SKILL.md` and optional supporting files; the manifest name and description must match the request. Frontmatter may contain `name`, `description`, `license`, `compatibility` and string `metadata`; native hooks, permission controls and subagent directives are rejected. Only regular files are allowed: path traversal, links, duplicate destinations, special files and invalid manifests are rejected. Content is inert during installation, and executable bits are kept. Inline metadata shows only type, name and description. + +### Plugins and capability directories + +A Plugin is an inline ZIP with type, name and description whose single archive root contains `.codex-plugin/plugin.json`. The manifest's `skills` names Skill directories; the whole package layout is kept. Public Plugin metadata shows only type, name and description. + +`openai_hosted` and Template `capability_directories` accept clean absolute paths inside `/workspace`, which initial files and setup can populate. `self_hosted` capability directories are absolute local paths on the machine. Directory-discovered Skills never appear as `skills` or `plugins` entries. Missing directories, duplicate Skill names, unsupported manifests and non-regular files fail initialization. + +| Archive and installation limit | Value | | --- | --- | -| API Session and Environment | Durable tenant ownership, configuration, pending interaction and connection observations. | -| Provider allocation | Compute and filesystem lifetime; caller-owned for `self_hosted`, service-owned for hosted provisioning. | -| Device and daemon connection | Authenticated engine-host identity and replaceable internal dispatch transport. | -| Harness process and native Session | Native model/tool loop, execution state and proven history/continuation path. | -| Runtime enrollment | Exact Environment/device/key binding for user-managed compute; no service-owned allocation. | - -Daemon, harness, tools and workspace are colocated in V1. Our daemon fills the -executor role; a separate native executor and service-side harness are retired. -The pinned public resources remain the target, while stock executor wire -interoperability is explicitly outside this implementation. - -Co-location needs a real credential and isolation design: generated code must not -gain the broader application credential or cross-tenant secrets through a shared -unrestricted process account. A directory binding alone is not isolation. Retain -native history independently of disposable compute, or prove native restoration; -never treat an Environment ID as a filesystem or history backup. Do not silently -move an existing Session away from its bound device. - -Environment identity, tenant/Session association, configuration and lifecycle belong -in Agents API, independently of provider compute, authenticated device identity, -daemon sockets and native harness Sessions. Create an Environment association in -the same transaction as its Session and creation identity when this resource is -implemented. Keep mutable connection/registration state out of immutable -configuration; replacement ownership must fence stale observations. - -The current hosted architecture is V1: Core runs independently; each Environment -sandbox contains its daemon, selected native harness, local tools and workspace. -Execution and Files use the same authorized workspace through the existing -Core/Runtime contract. Native tool calls stay local. -CLI discovery uses a bounded 15-second version probe per installed harness; -missing binaries fail immediately. A version result is availability, not Environment -readiness, and does not change initialization or connection ownership. Process placement and native -transport remain adapter responsibilities, without a second model/tool loop. -The former separated Runtime/harness and workspace executor topology is a distant -future V2 option, to revisit only after V1 is stable and concrete needs justify it. -Do not extend that topology for hosted delivery, maintain two current hosted routes, -or introduce dormant V2 compatibility scaffolding. - -When an execution Session has a previously started Turn but no recorded native -Session ID, Core requires existing-history recovery through a verified Runtime -capability. Read that condition before claiming the next Turn. A supplied native ID -remains authoritative. The Codex adapter may recover only a unique, nonarchived -root in the exact Session-private native home and expected working directory, using -native listing and exact-ID resume. Missing, incomplete or ambiguous history must -fail without starting a fresh root. A recorded start can precede native work; that -uncertain case also fails conservatively. Recovery does not replay interrupted -inputs, erase prior outcomes or promise transparent continuation of running tools. -Keep Device identity and Environment scope in `ExecutionDevice`; native Session -identity and prior API Turn state belong to `SessionExecutionBinding`. - -Platform-managed and user-managed deployment reuse this same Runtime. For platform -management, SandboxProvider creates and reclaims it. For user management, the user -starts the Runtime and its daemon authenticates and initiates the Core connection; -Core verifies principal ownership and the exact Environment binding. These are -management responsibilities, not separate execution architectures. - -In V1, our daemon fills the user-side executor role. Users deploy daemon, the -selected harness, local tools and workspace together. Do not require Codex -`exec-server`, a service-side harness, registry/Noise transport or remote tool -forwarding. The explicit daemon-executor decision supersedes the previous native -executor interoperability requirement. The superseded execution route is removed; -retain reusable filesystem helpers, -necessary regressions and historical evidence without a compatibility layer. +| Per archive | 5 MiB compressed, 20 MiB expanded, 1,000 entries | +| Inline Skills | 50, 10 MiB compressed and 50 MiB expanded in total | +| Plugins | 50, 10 MiB compressed and 50 MiB expanded in total | +| Installed snapshot | 50 Skills, 50 Plugins, 50 MiB | -### Explicit local tool environment +## Skills, Plugins and Environment MCP + +Skills and their immutable versions are Project resources, independent of Sessions and native installations. + +- Version allocation and pointer changes serialize on the owning Skill row. The top-level name and description follow the default version in the same transaction; non-default uploads keep them. Version identities stay unique across concurrent uploads and deletion. +- Bundles are encrypted with the service cipher, bound to Project, Skill and version. Metadata reads never load or decrypt bundles. Deleting a Skill reclaims its versions without affecting frozen Sessions. +- References resolve inside the Session creation transaction, after the upsert establishes ownership. Referenced resources are locked in a stable order, and the version, metadata and bytes freeze together. Public reference metadata, unresolved Template intent and the resolved Runtime bundle stay distinct; a reference is never reported as inline. + +Inline and referenced Skills use the same confidential snapshot and installer. The Runtime installs them under `skills/` in its capability root before setup and native execution. Adapters register only the selected Skill roots: Codex as explicit extra roots, MiniMax Code through its native catalog, and Claude as one controlled envelope per package with real directories and immutable hard links. Automatic native MCP discovery stays disabled. Codex nested `SKILL.md` discovery, `agents/openai.yaml` dependency configuration and Claude inline shell preprocessing fail adapter preparation. Native Harness configuration is never passed through wholesale. + +### Plugin MCP + +A Plugin declares MCP servers with `mcpServers: "./.mcp.json"` in `.codex-plugin/plugin.json`, or through a root `.mcp.json` when the path is omitted. The file holds `mcpServers` keyed by server name. Selecting a Plugin root as a capability directory activates its MCP declarations; selecting a parent directory discovers Skills without activating nested MCP servers. + +The shared parser accepts HTTP `url`, `bearer_token_env_var` and literal `http_headers`, and stdio `command`, `args`, selected `env_vars` and a package-relative `cwd`. Public `env_http_headers` is unsupported. The Runtime re-parses the frozen installed packages and resolves selected values only from the initialized env; a missing value fails instead of falling back to a model or daemon variable. + +A stdio server starts through the daemon's stdio helper, which resolves the installed declaration and launches the command with the Harness's permissions. On Unix the helper replaces itself with the server; on Windows it forwards stdio inside the owned process tree. Initialized values override the declaration's variables. Process groups and Windows Jobs own cancellation and descendant cleanup, not isolation. + +[Harness capabilities](harness-capabilities.md#environment-preparation) owns the supported Plugin transports and per-Harness limits. + +Environment MCP needs enabled network. Duplicate server identities are rejected. Claude rejects literal headers because the pinned client expands them again and forwards custom headers across origins. MiniMax ACP HTTP declarations stay in session-local native memory; tokens never enter native configuration files or process arguments. Required initialization and tool allowlists cannot be set through the Plugin manifest. + +### Effective bindings + +The Runtime resolves public HTTP declarations and installed Plugin MCP through `agent.ResolveMCPBindings` before adapter projection. Each transient binding keeps its connection origin, transport, nullable tool allowlist, required flag, credential authority and installed stdio identity. Bindings are never persisted or logged; duplicate identities and unavailable selected credentials are rejected. MiniMax reads its session-private native runtime-name registry for exact first-frame identities and cross-checks completed native results for both transports; adapters never fabricate a delayed start event or guess identities. + +### Public MCP connection origin + +An Agent's HTTP MCP tool ([declaration](execution-tools.md#http-mcp)) has a `connection_origin`. Omitted or null means `service`, including on `self_hosted` Sessions. The origin is kept from saved configuration through the Session snapshot to the Runtime request, which requires the exact wire version; a `service` request never becomes an `environment` connection. + +| Origin | Connects from | Allowed placements | +| --- | --- | --- | +| `service` | Core's service-side execution host | `none` only | +| `environment` | The Environment's workspace | `openai_hosted` and `self_hosted` with enabled network | + +Core's Harness profile declares `MCPOrigins`; admission and dispatch check the origin against the placement and the Runtime's advertised HTTP, bearer and required-initialization capabilities, and the Runtime validates the origin again before invoking an adapter. No Harness-name or Provider branch selects a different path. + +[Harness capabilities](harness-capabilities.md#tools) owns per-Harness origin support and policy limits. + +Both origins support anonymous HTTP and HTTPS bearer credentials. The attached-Vault selection freezes the credential identity, including a unique implicit URL match or an anonymous selection. Only that Project-authorized credential enters the transient Runtime request; Core defaults and unrelated Vaults are never searched. A decryption failure or missing credential fails execution without an anonymous fallback. Public Environment MCP keeps `project_vault` authority and Plugin credentials keep `environment_configuration` authority; neither overrides a duplicate server label. Bearers never enter persisted native configuration or process arguments. -`--tool-env-file` supplies the Runtime operator's base tool variables. Common -preparation copies these explicit values into its private initialization snapshot; -explicit Session `env` keys override the base. Runtime does not rewrite the source -file or inherit unrelated ambient credentials. Setup, capability resolution and -Harness execution read the same prepared snapshot. Reconnect preserves that -snapshot even if the operator edits the source; a new installation for a new -Session reads the current source. Harness profiles may reference the Runtime-owned -environment file, but must not persist copies of its values. A missing or invalid explicitly configured file -fails preparation. +Tool allowlists and required initialization follow the [HTTP MCP contract](execution-tools.md#http-mcp). diff --git a/contracts/agents-api/execution-tools.md b/contracts/agents-api/execution-tools.md index d4b722515..991bb02af 100644 --- a/contracts/agents-api/execution-tools.md +++ b/contracts/agents-api/execution-tools.md @@ -1,120 +1,91 @@ -# Execution and tools coverage - -Assessed against Core `206474a4c10901442b1faa094281bfb5559082b5` on 2026-09-22. -The contract remains [OpenAI Python 3.13.0 at `d7c41ef`](upstream.json), Beta -`agents=v1`. This matrix records the bounded execution/tools milestone. It does not -establish complete protocol compatibility; the remaining limits below stay open. - -**Implemented** means current source and controlled checks support the operation. -**Qualified** additionally identifies a real public Core/PostgreSQL/daemon/native -model workflow in the evidence register below. **Unsupported** describes a current -admission/native limitation, not a restriction in the official contract. -**Unverified** means evidence does not establish the stated semantics or profile. -Qualification never extends automatically to another model, placement or combination. - -## Operation matrix - -All function rows refer to application-defined functions. Native workspace tools -and Environment Plugin MCP have separate inventories and qualification. - -| Operation | Implemented behavior and real qualification | Unsupported or unverified boundary | -| --- | --- | --- | -| Initial message input, `sessions.create` | String input and ordered user-message arrays share atomic admission. Codex/Claude text and inline image execution: M1/M2; ordinary MiniMax text: M1/P1. [Input contract](message-input.md), [initial parser](../../services/core/internal/api/session_initial_input.go). | Whitespace-only text is admitted verbatim for Codex and rejected at admission for Claude SDK and MiniMax Code ([text content](message-input.md#text-content)). Empty parts beside text (SES-08) and local size-limit parity need upstream evidence. Inline PNG/JPEG is qualified on `none`, Core-managed Docker `openai_hosted` and `self_hosted`; MiniMax images and remote URLs remain gaps. | -| Prepared and active messages, `sessions.events.create` | Ordered `input_text`/`input_image` arrays retain original content and distinct public user Items. HTTP 202 confirms persistence; native receipts establish application. Initial/prepared/active Docker PNG/JPEG: M2; active PNG on `none`: M1. [Shared admission](../../services/core/internal/api/inputs.go). | Codex flattens native messages with blank-line separators. Claude can fold or queue native turns; it does not promise Codex's same-native-turn behavior. Public durability does not prove native consumption. Claude SDK and MiniMax Code reject messages without an image or non-whitespace text at admission (`unsupported_or_invalid_configuration`); Codex delivers them unchanged. | -| Structured output, `agent.text.format` | Save/inherit/freeze `{type:json_schema,schema:...}`. Claude SDK object-root, single Agent, medium verbosity, ordinary functions returning text: S1 (`none`) and S2 (Docker, including prepared/active input and Files/Artifacts), plus [self-hosted execution and cold continuation](environment-capabilities-qualification.md). Native final text is retained unchanged; its stream follows the official message sequence with the whole text in one `output_text.delta` ([item serialization](history-events-usage.md#item-serialization-2026-09-23)). [Contract](structured-output.md), [profile](../../services/core/internal/engine/claude.go). | An explicit non-object root type is a protocol error for every harness ([validation](official-semantics-alignment.md#agent-configuration-validation--september-23)). Codex/MiniMax, schemas without an object root, schema numbers changed by binary64, Skills/Plugins, MCP, Subagent and discovery combinations reject execution. No output repair, coercion or extra model loop. Arbitrary schema dialects are unverified. | -| Function configuration, saved/inline Agents | Required name/description/schema; `defer_loading` defaults false. Protocol errors, repeated names and explicit non-object root types reject with the official fields ([validation](official-semantics-alignment.md#agent-configuration-validation--september-23)). Saved references resolve into an immutable Session snapshot. Codex/Claude real calls: F1/F2/M2/S1/S2. [Parser](../../services/core/internal/api/function_configuration.go), [saved tools](../../services/core/internal/api/saved_tools.go). | MiniMax public functions reject. Claude requires an explicit object root. Local nonblank/512-byte name and 64-definition Session bounds are compatibility gaps. Saving configuration alone does not qualify execution. | -| Function-result admission, `events.create` | Required `turn_id`, `call_id`, `success`; optional nullable `error` and `output`. Output is string or ordered text/image content. Scoped atomic batches retain field presence, original content and retry identity in storage; public result Items and events always carry `output` and `error`, null when not submitted ([item serialization](history-events-usage.md#item-serialization-2026-09-23)). Same result retries are accepted; changed results and results after cancellation return 409 `conflict_error`, including after terminal state. In the caller's Session, an unknown call or a call of another Turn returns 400 `invalid_request_error` without changing the pending action; missing and foreign Sessions stay 404 ([error rows](official-semantics-alignment.md#session-input-conflicts-and-result-targets--september-23)). F1/F2/M2 plus [controlled SDK/raw checks](../../services/core/tests/official_function_inputs.py). [Parser](../../services/core/internal/api/function_inputs.go), [Store](../../services/core/internal/store/function_inputs.go). | Admission is separate from application and public Item publication. Invalid or unqualified content cannot consume a pending call. Core messages omit the call and executor IDs that official messages name; hosted defaults and publication timing remain unverified. | -| Function text results and native application | Codex waits for a matching live root dynamic-tool completion; Claude waits for a matching live root native tool result. Text/error results, retry/conflict, cancellation and cold continuation: F1/F2. [Receipt contract](function-result-images.md). | Transport writes alone do not confirm application. Confirmation is not provider consumption, crash recovery or exactly-once external effects. No automatic result replay. | -| Function image results | Successful ordered inline PNG/JPEG with text, large PNG and image-only JPEG: Codex/Claude `none` F1/F2; Docker M2. Public Items retain submitted bytes. Codex receipt regression: F2. | Claude rejects failed images and remote references before persistence; native resizing may change its image bytes. Self-hosted Claude image results use the same native path; see the current [qualification record](environment-capabilities-qualification.md). MiniMax functions remain unqualified. Other Codex image/error/reference combinations cannot be inferred from successful-inline evidence. | -| Deferred function discovery | Type-only `tool_search` plus mixed eager/deferred functions: Claude SDK 0.3.269/native 2.1.269, Kimi K3, single Agent, medium, `none`, text results; text and PNG input D1. Native provider observations establish lazy schema loading for D1. [Self-hosted workspace callback, continuation and cancellation](environment-capabilities-qualification.md) use the same native path; they do not add model-request observer evidence. [Contract](tool-search.md). | A repeated `tool_search` is a protocol error. Codex/MiniMax discovery, search-only/missing-search, workspace with Skills/Plugins, MCP, structured-output and Subagent combinations remain gaps. Opaque native policy changes lack a reliable pre-input deferral signal. Saved tools include `tool_search`; the pinned Session response union excludes it. Exact hosted projection is unverified. | -| Explicit disabled search/PTC | Saved and inline `web_search.mode=disabled` and `programmatic_tool_calling.enabled=false`; shared native controls on initial and cold execution. All three harnesses on `none`: P1, including native inventory/control evidence. [Contract](tool-policy.md), [parser](../../services/core/internal/api/disabled_tools.go). | Unsupported explicit enablement rejects at Session admission; saved Agents keep every pinned search mode as resource data (TV-05), and Sessions from such Agents reject unless they replace the tools. A repeated `web_search` is a protocol error. Omission retains approved native behavior, which does not establish official default-on PTC parity. Enabled search and default/error parity remain gaps; unrelated native utilities are not implicitly removed. | -| Agent service-origin MCP | Implemented Codex/Claude HTTP `none` profiles, anonymous/static bearer, scoped Vault selection and native `mcp_call` Items. [Configuration and qualification limits](../../services/core/README.md#http-mcp-execution), [profiles](../../services/core/internal/engine/profile.go). | This closure batch does not requalify MCP/provider combinations. MiniMax, self-hosted/hosted service-origin MCP, OAuth, nonempty inline headers/metadata and other transports reject. Environment origin follows the separate row below. Claude requires a static connected inventory; original MCP-envelope fidelity and continuing server health are unverified. | -| Agent Environment-origin MCP | Public HTTP declarations reuse installed MCP's effective Runtime bindings; attached Vault selection and native observations remain common. [Origin and Harness matrix](environments.md#public-mcp-connection-origin), [real qualification](https://github.com/MiniMax-AI/OpenAgentCore/blob/e974a7f880a2eb799f0dd39e6ba0870462854a53/contracts/agents-api/public-mcp-qualification.md). | Requires a workspace and enabled network. MiniMax rejects every non-null allowlist and required initialization. No service-origin relocation, automatic fallback or credential copy into native profiles. | -| Environment Plugin MCP/native tools | Separate [Docker Plugin transport matrix](environment-templates.md#environment-origin-mcp-plugins) and [V1 deployment evidence](user-managed-runtime-v1.md). Native tools stay within the existing colocated Runtime. | Plugin MCP is not Agent service-origin MCP or a public function action. No new optional cross-product is qualified here. Native utility inventories need not be identical. | -| Required action: `function_call` | Session GET/list and Session SSE expose persisted `{arguments,call_id,name,turn_id,type}`. Actions remain pending until native application or cancellation/terminal settlement. [Projection](../../services/core/internal/store/function_state.go), [confirmation](../../services/core/internal/store/function_results.go); real handling F1/F2/M2/S1/S2/D1; explicit client disconnect/query/reconnect: F3. | Function-call history is not pending-state authority. Client reconnect does not replay events or redo external application effects. | -| Required action: `environment_connection` | Offline waiting input exposes `{environment_id,type}` before Turn creation. Connect the exact Session Environment through our scoped daemon enrollment. Three-harness Docker/E2B execution: E1; explicit initial pending-input/query/connection/native completion: E2; expiry: [controlled initial-input test](../../services/core/tests/official_self_hosted_initial.py). | Idle offline Sessions without pending input request nothing. Registration alone is not connectivity or native readiness. Stock `exec-server`/Noise transport is outside the approved V1 route. Exact upstream registration/action-removal timing remains unverified. | -| Stream disconnect and execution loss | SSE is live-only. Reconnect, buffer events, retrieve Session/Turn/Items and deduplicate Item IDs. Worker recovery fails previously claimed work without replay; queued work can remain. [Recovery contract](README.md#public-execution-admission), [connection reconciliation](../../services/core/internal/store/environment_connection_recovery.go). | Client disconnect and daemon/Core loss are different cases. F3 qualifies client-stream loss and continued pending function handling; F2 qualifies native process loss with an unapplied saved result. Neither promises replay or recovery of unknown external effects. | - -## Required-action recovery contract - -The fixed [Session union](https://github.com/openai/openai-python/blob/d7c41efee1b0802b79f3f88a678ef2052b06e9ce/src/openai/types/beta/agent_session.py) -defines both variants. The official [Functions](https://developers.openai.com/api/docs/guides/agents-api/tools/functions) -and [Manage sessions](https://developers.openai.com/api/docs/guides/agents-api/sessions/manage) -guides, checked on 2026-09-22, direct clients to retrieve `required_actions` after -restart or stream loss. A historical `function_call` Item alone cannot establish -that a result is still pending. The fixed SDK governs field/type compatibility; -current prose does not silently change that pin. - -For a pending function, use the returned Session, Turn and call identity. If the -application already performed the function, keep and submit that saved result; -do not run an external effect again merely because an action remains visible. -Core's acknowledgement boundary is native application. For an environment action, -connect its exact Environment using the existing enrollment authorization. -[Connection observations](../../services/core/internal/store/environment_connections.go) -fence stale generations; transport connection can clear pre-Turn activity before -native preparation. Neither registration nor an action disappearing proves that -a model ran successfully. Query the matching Turn and Items for its outcome. - -Two timing questions remain open. Repeated `requires_action` notification order -and acknowledgement timing are local implementation choices, not proven upstream -semantics. Also, an admitted result that is cancelled before native observation -can remain internally saved without a public output Item or `item.added`. -[Item publication](../../services/core/internal/store/item_projection.go) and -[the existing coverage record](README.md#public-function-configuration) preserve -this gap. Successful retry cannot manufacture the missing observation or establish -that the model consumed the result. - -## Evidence register - -Real records below use fixed SDK 3.13.0 and raw HTTP through Core, dedicated -PostgreSQL, daemon and native harness. Evidence is private under `~/.parsar/remediation/` on `zju_a100_2` and the -development host. Paths containing `validated/` or `server-evidence/` identify -downloaded local mirrors; summary files may also live only on the development -host. Locators identify retained evidence without publishing credentials or -private launch configuration. Revisions identify the accepted change, -not a claim that every historical workflow was rerun at the current baseline. - -| ID | Exact evidence and accepted scope | +# Execution tools + +An Agent declares application functions, controls and MCP servers in `tools`, and an optional output schema in `text.format`. This contract states how Core validates each declaration, what crosses the Runtime boundary and how callers recover required actions. [Harness capabilities](harness-capabilities.md) lists which Harness supports each operation on which placement. Native workspace tools and Environment Plugin MCP are part of the [Environment](environments.md#skills-plugins-and-environment-mcp). + +## Admission + +- Saved Agents keep every pinned tool declaration as resource data. Saving never qualifies execution. +- Session creation resolves saved references and inline declarations with the execution parser into the immutable Session snapshot, then checks the combination against the selected Harness's profile in `services/core/internal/engine`. An unsupported combination returns 400 `unsupported_or_invalid_configuration` before anything is written. Protocol errors, such as a repeated `web_search` or `tool_search` or a non-object schema root, use the official error fields ([validation](official-semantics-alignment.md#agent-configuration-validation--september-23)). +- Before dispatch, the selected Runtime must also advertise the operation's capability. An advertisement alone never enables an operation. +- The native Harness runs the model and tool loop. Core adds no second loop, output repair, schema coercion or prompt wrapper, and selects no native tool names. + +## Functions + +A function declaration requires `name`, `description` and a JSON Schema in `parameters`; `defer_loading` defaults to false and cannot be null. Names are nonblank, unique and at most 512 bytes; a Session holds at most 64 function definitions. Saved references resolve into the Session snapshot, and the definitions stay fixed through native preparation and continuation. + +**Results.** A caller submits `agent.session.input.tool_result` events with `turn_id`, `call_id` and `success`, and optional nullable `error` and `output`. Output is a string or ordered text and image parts, subject to the Harness. Batches are atomic and keep field presence, original content and retry identity; public result Items and events always carry `output` and `error`, null when not submitted. + +| Case | Response | +| --- | --- | +| Identical retry, also after the Turn ends | Accepted | +| Different result for the same call | 409 `conflict_error` | +| First result after the Turn was cancelled | 409 `conflict_error` | +| Unknown call, or a call of another Turn, in the caller's Session | 400 `invalid_request_error`; the pending action is unchanged | +| Missing or foreign Session | 404 | + +[Session input conflicts](official-semantics-alignment.md#session-input-conflicts-and-result-targets--september-23) records the exact messages. Invalid or unsupported content cannot consume a pending call. Admission is separate from application: the adapter confirms a result only when the matching native tool result appears in the live root Turn ([receipt contract](function-result-images.md)). A transport write alone confirms nothing, and a confirmation says nothing about provider consumption or exactly-once external effects. Core never replays a result automatically. + +### Required actions and recovery + +A pending function appears in Session reads and Session SSE as a required action `{arguments, call_id, name, turn_id, type: "function_call"}` until native application, cancellation or terminal settlement. An offline `self_hosted` Environment with pending input shows `{environment_id, type: "environment_connection"}` ([Environments](environments.md#activity-and-required-actions)). + +SSE is live only. After a restart or a lost stream, read the Session's `required_actions`; a `function_call` Item in history does not prove the call is still pending. For a pending function, use the returned Session, Turn and call identity. If the application already ran the function, submit the saved result instead of running the external effect again. For an environment action, connect that exact Environment with its enrollment. Neither a registration nor a disappearing action proves that the model ran; read the Turn and its Items for the outcome. Reconnecting never replays events or external effects. + +The Worker fails previously claimed work without replay after execution loss; queued work can remain queued. A result admitted and then cancelled before native observation stays saved internally but can lack a public output Item and `item.added`. + +## Structured output + +`text.format` accepts `{type: "json_schema", schema: {...}}`, the Agents API form: it has no `name`, `strict` or other Responses API wrapper fields. The schema is saved, inherited through Agent and Session resolution and frozen in the Session snapshot. An explicit non-object root type is a protocol error on save and Session creation for every Harness. Claude requires an explicit `type: "object"` at the schema root. The Claude SDK reads JSON numbers as binary64, so Session admission rejects schemas whose numbers would change in that conversion; saved Agents keep them unchanged. + +Core carries the schema in `ExecutionControls.OutputFormat` and requires the profile's structured-output qualification plus the Runtime's `structured_output` and message-observation capabilities, only for requests that use the option. Frozen schemas reach preparation before input and apply to initial and resumed execution; Start cannot replace them. + +The Claude adapter passes `outputFormat` to the pinned SDK and allows its native `StructuredOutput` terminal tool, which is internal and never an extra caller function. A matching live root tool result and an attributed successful SDK result confirm the output. The adapter publishes the native `result.result` string unchanged as a completed `final_answer` message with the native tool-use ID; parent assistant prose keeps its own ID. Unvalidated retries and cancelled candidates never become the answer, and the adapter never serializes `structured_output` back to JSON. The stream follows the official message sequence with the whole text in one `output_text.delta`. The bridge advertises the operation only when it reports `structured_output`, and a workspace Runtime also needs `workspace_structured_output`. + +## Deferred function discovery + +A `tool_search` tool has only `type`; Responses-only execution fields are rejected. Function `defer_loading` marks which definitions load lazily. Discovery requires both: `tool_search` without a deferred function, or deferred functions without `tool_search`, is rejected. The saved-Agent tool union keeps `tool_search`; the pinned Session response union omits it, so Session and SSE resources project it out while the frozen configuration keeps it. The pinned Items union has no tool-search Item, and Core invents none. + +Core sends `PromptRequestPayload.ToolSearch` and each `FunctionTool.DeferLoading`, and requires the profile's qualification plus the Runtime's `tool_search` capability. Native search and lazy schema loading belong to the adapter. The Claude adapter's MCP server marks eager definitions `anthropic/alwaysLoad:true` and deferred ones false, and enables native ToolSearch; the function profile allows only declared callbacks and ToolSearch besides the selected workspace tools. A workspace Runtime derives `tool_search` from the bridge's `workspace_tool_search` feature. The native Harness owns model and provider policy; known conflicting modes and beta settings reject in the adapter, and the SDK gives no reliable pre-input signal that deferral took effect after an opaque policy change. + +## Web search and programmatic tool calling + +```json +[ + {"type": "web_search", "mode": "disabled"}, + {"type": "programmatic_tool_calling", "enabled": false} +] +``` + +Saved Agents keep every pinned `web_search` mode: omitted or null is saved as `live`, and `cached` and `live` as sent ([saved modes](official-semantics-alignment.md#saved-web_search-modes--september-23)). Search settings are resource data: omitted or null `context_size` resolves to `medium`; omitted domains and location resolve to null; an empty domain list stays empty; a supplied location, including `{}`, has `city`, `country`, `region` and `timezone`, null where omitted. + +Execution admits only `mode: "disabled"` and `enabled: false`. Enabled or omitted-mode search and enabled or omitted-`enabled` programmatic calling are rejected at Session admission unless the Session replaces the saved tools. Omitting programmatic configuration keeps each Harness's native behavior, which differs from the official default-on behavior. Unrelated native utility tools are not removed. + +`DisableProgrammaticToolCalling` carries the disabled intent on initial execution and cold continuation and requires the Runtime capability only when present. Search uses the existing disabled control. + +| Harness | Native enforcement | | --- | --- | -| M1 | `20260922/message-image-input/public-codex-reviewed.log` (59.41s), `public-claude-reviewed.log` (79.52s), `public-mcode-network-fixed.log` (48.95s); [PR #12](https://github.com/MiniMax-AI/OpenAgentCore/pull/12), merge `37a2923`. Codex/Claude Kimi K3 PNG initial/active inputs, receipts, retries, cancellation, cold continuation and isolation; MiniMax M2.7 ordinary text regression. | -| M2 | `20260922/workspace-images/validation-summary.json`; Codex `public-run-_lr5uiy_`, Claude `public-run-sb65p9e9`; candidate `2c295116d66ed9aafc808ec49a8cb6dcb5c674c2`, merge `1adb2489bc3415039440224e9a7905c53e68a557` ([#17](https://github.com/MiniMax-AI/OpenAgentCore/pull/17)). Kimi K3, Codex 0.153.4, Claude SDK 0.3.269/native 2.1.269; Docker seven-Turn image/workspace workflow. Codex's receipt conclusion is superseded by F2. | -| F1 | `20260922/function-result-images/validation-summary.json`, `validated/function-image-public-2567456666/public.json`; candidate `d3cdb225bc7628b968b76c7e1b400101894ce8cd`, merge `fd23f169e1ede1b2f1c39a1d9dcce71886e3c006` ([#16](https://github.com/MiniMax-AI/OpenAgentCore/pull/16)). Claude SDK 0.3.269/native 2.1.269 with Kimi K3 on `none`; success PNG/JPEG, failed text, retries, cancellation, history continuation. Earlier Codex transport-only evidence does not qualify current receipts. | -| F2 | `20260922/function-result-receipts/validation-summary.json`, `candidate-public-owner-sdk.log`; hosted `public-run-4fn70foy` (104.13s), `none` `function-image-public-945075091` (83.918s). Candidate `6dceb98e1c56469fcc70cd82025ce77fc870ed3d`, merge `206474a4c10901442b1faa094281bfb5559082b5` ([#20](https://github.com/MiniMax-AI/OpenAgentCore/pull/20)); real Kimi, Codex 0.153.4. Qualifies native receipts, not completed provider consumption. | -| F3 | `20260922/execution-tools-closure/pending-actions/validation-summary.json`; Codex `codex/public-run-8gexf5a8` (65.58s), Claude `claude_sdk/public-run-5f2tr7e_` (74.33s), production source `206474a`. The same [public verifier](../../services/core/tests/official_pending_actions_native.py) checks disconnected-client pending queries, success/error/cancel, original results, target isolation, retry/conflict and no implicit call replay with real Kimi on Docker. No Core restart or native-loss injection is claimed by these runs. | -| S1 | `20260922/structured-output/acceptance-summary.json`, `structured-public-1905281777/public.json`; real candidate `7f533d46c472f3cdf7c16ce9471c225a2ba7eded`, merge `8716e699dbc34904499c94b6b0b03857bbebfaad` ([#11](https://github.com/MiniMax-AI/OpenAgentCore/pull/11)). Claude/Kimi `none` schema/function execution and cold continuation. Schema-specific active steering was not separately qualified here. | -| S2 | `20260922/hosted-structured-output/validation-summary.json`, `validated/public-inline.log`, `public-run-qvq3csf1` (186.09s); candidate `f6d2c3f2a1dc2ff5f177213bd2b20d78496011d0`, merge `aa85d8ecc60e0a8b7f2494b73caa81216ee197e2` ([#18](https://github.com/MiniMax-AI/OpenAgentCore/pull/18)). Claude SDK 0.3.269/native 2.1.269, real Kimi; Docker saved/inline schema, prepared/active input, native files and four completed structured answers. | -| D1 | `20260922/deferred-tools/tool-search-public-2039155387/public.json`, `public-image-isolated-tests.log` (86.71s); merge `178507ae7a63d4068e1e82bef9cd256ba398ae00` ([#13](https://github.com/MiniMax-AI/OpenAgentCore/pull/13)). Claude SDK 0.3.269/native 2.1.269/Kimi K3 `none`; text results with initial/active PNGs, cancellation and cold continuation. `claude-native-1790052750` separately records provider schema inventories and native feasibility. | -| P1 | `20260922/tool-policy/acceptance.json`, `server-evidence/tool-policy-{codex-2495445746,claude_sdk-2539346823,mcode-3747575401}/public.json`; candidate `eeb432c7495265ae3a9309e32f5b4323611c3245`, merge `4b8754a6eda6696d9b18eb8d583953ac16c6800f` ([#14](https://github.com/MiniMax-AI/OpenAgentCore/pull/14)). Codex/Claude Kimi K3, MiniMax M2.7; four `none` Sessions per harness, native disable controls and cold continuation. | -| E1 | `20260921/self-hosted-onboarding/{source-candidate.json,source-verified.json,live,live-e2b}`; candidate `8f0cd2530d7b58cb7fb3ea124a1fc43dacecca36`. [Exact Docker/E2B run paths and limits](user-managed-runtime-v1.md#evidence-and-verification-boundaries): Codex Kimi K3 Responses, Claude Kimi K3 Anthropic-compatible API, MiniMax M2.7, immutable Linux amd64 Runtime builds. Native execution, restart/history, cancellation, Files/Artifacts and isolation; shared credential lifecycle evidence is not a separate live rotation run for every E2B profile. | -| E2 | `20260922/execution-tools-closure/environment-actions/result.json` contains Codex/Claude Kimi passes (13 checks each) and the retained MiniMax direct-network failure; `mcode-relay-retry/result.json` separately passes MiniMax M2.7 (13 checks) after correcting provider routing. Source `206474a`; same shared SDK/raw workflow, isolated Core/PostgreSQL and user-managed Docker. Covers initial creation SSE, client disconnect, exact pending Environment/no Turn or Items, scoped enrollment, one actual completed model Turn and creation retries preserving history. This is connection-action qualification, not another full lifecycle or E2B regression. | - -Controlled tests establish parser, projection, atomicity and race behavior; native -feasibility probes establish only feasibility. Neither replaces the real records. -Failed setup, provider and assertion runs remain failures with their original -scope. The unchanged gate requirements and independent review apply to this batch. - -## Milestone closure and remaining limits - -1. **Principal function recovery:** F3 directly verifies live client-stream loss - while a function is pending on Codex and Claude, authoritative queries, exact - targeting, rejection without mutation, original results, retry/conflict, - application, cancellation and isolation. F2 separately establishes truthful - native-loss settlement without replay. No active execution recovery is promised. -2. **Principal environment recovery:** E2 explicitly verifies pending - `environment_connection` recovery and single initial execution on all three - harnesses. E1 retains broader deployment/history/isolation qualification. - Registration, connection and preparation remain distinct; no exact upstream - timing or arbitrary-provider qualification is inferred. -3. **Recorded nonblocking timing:** an accepted result cancelled before native - application may lack a public result Item. The submitted data remains durable; - F2 confirms that native loss does not turn unknown delivery into Applied. The - fixed sources do not establish the publication point for this interleaving. - Keep it queued rather than inventing a public state or rewriting safe receipt - ownership. This limited timing question is not accepted protocol parity and - does not invalidate F3's principal pending-result workflow. - -Approved native capability differences (including default PTC), unsupported optional -combinations, arbitrary provider parity and low-frequency error/default edge cases -remain separate limitations. They do not authorize silently relabeling an unproved -principal workflow as deferred. OAuth, full Usage, release publication, new tool -kinds and new execution loops are outside this closure batch. +| Codex | Disables code-mode features; checks native managed requirements before thread start or resume and rejects a forced conflicting feature | +| Claude SDK | Keeps the restricted built-in inventory and verifies native initialization against it | +| MiniMax Code | Keeps the restricted native tool profile, an empty text-execution inventory and disabled web search | + +## HTTP MCP + +```json +{ + "type": "mcp", + "server_label": "tickets", + "transport": {"type": "http", "server_url": "https://mcp.example.com/mcp"}, + "connection_origin": "service", + "allowed_tools": ["lookup_ticket"], + "required": false +} +``` + +- `server_label` is nonempty and unique within the Session. Only the `http` transport is accepted; `server_url` is an absolute HTTP or HTTPS URL without credentials, query or fragment. Nonempty `headers` and `request_metadata` are rejected. +- [Public MCP connection origin](environments.md#public-mcp-connection-origin) owns origin defaults, placement and credential authority; [Harness capabilities](harness-capabilities.md#tools) owns per-Harness support. +- Omitted or null `allowed_tools` permits every server tool; `[]` permits none. +- `required: true` makes native thread creation and cold resume wait for the server to initialize; a failure stops execution without replacing retained history. It needs the Runtime's `mcp_http_required` capability. Public work can be accepted or queued during the wait. +- Bearer authentication uses an attached static or OAuth Vault credential. [Vault credentials](../../services/core/credentials.md) owns selection, and [MCP credential authority](environments.md#public-mcp-connection-origin) owns the frozen Runtime binding. Authenticated execution requires `mcp_http_bearer_auth`. +- The Runtime must advertise `mcp_http_tools`. The native Harness owns discovery, calls and results; public `mcp_call` Items use the original server and tool names and keep the observed native result. + +Codex verifies the exact effective MCP configuration before starting or resuming a thread, excludes undeclared servers, disables native apps and plugins, and rejects reserved native labels and stored native MCP credentials. Claude accepts labels of ASCII letters, digits, underscore and hyphen except `functions`, tool names that may also contain dots, and requires connected servers with static inventories. Anonymous Claude requests send a blank Authorization header to suppress native OAuth injection. Native OAuth login is not supported. diff --git a/contracts/agents-api/function-result-images.md b/contracts/agents-api/function-result-images.md index 699d10e75..a1c7af0b6 100644 --- a/contracts/agents-api/function-result-images.md +++ b/contracts/agents-api/function-result-images.md @@ -67,6 +67,6 @@ parity, arbitrary managed output rewrites, crash recovery or full Agents API compatibility. No downloader, image converter or second tool loop belongs in Core. User-managed Linux image execution follows the same native workspace path. See -[current qualification](environment-capabilities-qualification.md) for the tested +[current qualification](harness-capabilities.md) for the tested Harnesses, formats, continuation and remaining boundaries. Historical evidence above retains its original deployment scope. diff --git a/contracts/agents-api/harness-capabilities.md b/contracts/agents-api/harness-capabilities.md new file mode 100644 index 000000000..fd0350449 --- /dev/null +++ b/contracts/agents-api/harness-capabilities.md @@ -0,0 +1,59 @@ +# Harness capabilities + +This page lists what each Harness supports on each placement. Core decides admission from the Harness's engine profile in `services/core/internal/engine`, and the Runtime that runs the Session must also advertise the operation. The linked contracts define each operation; [Harness onboarding](harness-onboarding.md#qualify-the-adapter) describes how a Harness is qualified. + +| Status | Meaning | +| --- | --- | +| Verified | Core admits it, and real-model acceptance through the pinned official client passed on that placement | +| Admitted | Core admits it through the same Runtime path, but no real-model acceptance has run on that placement | +| Rejected | Core rejects the request before execution | + +Placements are `none` (no Environment), hosted (`openai_hosted`) and self-hosted (`self_hosted`). Hosted acceptance ran on Docker nodes; E2B and microsandbox run the same Runtime and adapters, and every Verified hosted cell counts as Admitted there. Self-hosted acceptance ran on Linux machines; the [self-hosted guide](../../docs/getting-started/self-hosted.md#platforms) lists the supported platforms. A verified operation is verified on its own, not in every combination with other options; combinations that a profile rejects are listed in the operation's contract. + +## Execution and input + +| Operation | Codex | Claude SDK | MiniMax Code | +| --- | --- | --- | --- | +| Text Turns, active input, cancellation, restart and continuation | Verified on all placements | Verified on all placements | Verified on all placements | +| [Files and Artifacts](environment-files.md) | Verified: hosted, self-hosted | Verified: hosted, self-hosted | Verified: hosted, self-hosted | +| [Whitespace-only message text](message-input.md) | Admitted; delivered unchanged | Rejected | Rejected | +| [Inline PNG and JPEG message images](message-input.md) | Verified on all placements | Verified on all placements | Rejected | +| Remote image URLs | Rejected | Rejected | Rejected | +| Explicit `reasoning`; `service_tier` other than `auto` | Rejected | Rejected | Rejected | +| `text.verbosity` other than `medium` | Admitted; the native model decides | Rejected | Rejected | +| [Public token usage](history-events-usage.md) | Measured counters | Null | Null | + +Native model parameters and provider protocols per Harness are in [model execution](model-execution.md). + +## Tools + +| Operation | Codex | Claude SDK | MiniMax Code | +| --- | --- | --- | --- | +| [Public functions](execution-tools.md#functions) with text results | Verified: `none`, hosted; admitted: self-hosted | Verified on all placements; object-root schemas only | Rejected | +| [Function results with images](function-result-images.md) | Verified: `none`, hosted; admitted: self-hosted | Verified on all placements; successful inline PNG or JPEG results only | Rejected | +| [Structured output](execution-tools.md#structured-output) | Rejected | Verified on all placements | Rejected | +| [Deferred function discovery](execution-tools.md#deferred-function-discovery) | Rejected | Verified: `none`, self-hosted; admitted: hosted | Rejected | +| [Disabled web search and programmatic tool calling](execution-tools.md#web-search-and-programmatic-tool-calling) | Verified: `none`; admitted: hosted, self-hosted | Verified: `none`; admitted: hosted, self-hosted | Verified: `none`; admitted: hosted, self-hosted | +| Enabled web search or programmatic tool calling | Rejected | Rejected | Rejected | +| [Service-origin HTTP MCP](environments.md#public-mcp-connection-origin) | Verified: `none`; rejected elsewhere | Verified: `none`; rejected elsewhere | Rejected | +| [Environment-origin HTTP MCP](environments.md#public-mcp-connection-origin) | Verified: hosted, self-hosted | Verified: hosted, self-hosted | Verified: hosted, self-hosted; `allowed_tools` null and `required` false only | +| [HTTP MCP bearer credentials](execution-tools.md#http-mcp) | Verified | Verified | Verified | +| Required MCP initialization | Verified | Verified | Rejected | +| [Subagents](subagents.md) | Verified: hosted; admitted: `none`, self-hosted | Verified: hosted; admitted: `none`, self-hosted | Verified: hosted; admitted: `none`, self-hosted | +| Subagents together with functions or HTTP MCP | Rejected | Rejected | Rejected | + +Claude structured output requires a single Agent, medium verbosity and no Skills, Plugins, capability directories, MCP or `tool_search`. Claude deferred discovery requires a single Agent, only function tools besides `tool_search` and the disabled controls, no Skills, Plugins or capability directories and no structured output. Claude on a workspace placement needs the packaged bridge features for each operation (`workspace_functions`, `workspace_structured_output`, `workspace_tool_search`, `workspace_mcp_http`). + +## Environment preparation + +These operations need a workspace, so they apply to hosted and self-hosted placements only. + +| Operation | Codex | Claude SDK | MiniMax Code | +| --- | --- | --- | --- | +| [Initial files, setup commands, Skills and Plugins](environments.md#runtime-capability-preparation) | Verified: hosted, self-hosted | Verified: self-hosted; admitted: hosted | Verified: self-hosted; admitted: hosted | +| Capability directories | Admitted | Admitted | Admitted | +| npm and Python packages | Admitted | Admitted | Admitted | +| `packages.system` | Rejected | Rejected | Rejected | +| Network `disabled` or `restricted` | Rejected | Rejected | Rejected | +| [Plugin MCP over stdio](environments.md#plugin-mcp) | Verified: hosted, self-hosted | Verified: self-hosted; admitted: hosted | Verified: self-hosted; admitted: hosted | +| Plugin MCP over HTTP | Admitted, with literal headers or HTTPS bearer | Admitted, anonymous or HTTPS bearer | Verified: self-hosted; admitted: hosted; anonymous or HTTPS bearer | diff --git a/contracts/agents-api/harness-catalog.md b/contracts/agents-api/harness-catalog.md index 5990b3c9e..bdb1a34ff 100644 --- a/contracts/agents-api/harness-catalog.md +++ b/contracts/agents-api/harness-catalog.md @@ -1,10 +1,7 @@ # Built-in Harness registrations -The authored registration list is -[`catalog.json`](../../internal/harnessconfig/builtin/catalog.json). -Change it and run `make generate-harness-catalog`; generated projections are checked -by `make check-harness-catalog`. +The authored registration list is [`catalog.json`](../../internal/harnessconfig/builtin/catalog.json). Change it and run `make generate-harness-catalog`; `make check-harness-catalog` checks the generated projections. | Identifier | Display name | Model configuration declaration | Core qualification constructor | | --- | --- | --- | --- | @@ -12,8 +9,4 @@ by `make check-harness-catalog`. | `codex` | Codex | `codex.Configuration` | `codexProfile` | | `mcode` | MiniMax Code | `mcode.Configuration` | `mcodeProfile` | -Configuration declarations live in `internal/harnessconfig/`; -qualification constructors live in `services/core/internal/engine`. -These registrations describe the build. Deployment enablement, qualified -operations and a connected Runtime's actual availability remain separate checks. -See [Harness onboarding](harness-onboarding.md) for adapter and packaging steps. +Configuration declarations live in `internal/harnessconfig/` and qualification constructors in `services/core/internal/engine`. These registrations describe the build. Which Harnesses a deployment enables is the `core.harnesses` [process setting](../../docs/configuration.md#settings), [Harness capabilities](harness-capabilities.md) lists what each Harness supports, and a connected Runtime reports its own availability. [Harness onboarding](harness-onboarding.md) describes the adapter and packaging steps. diff --git a/contracts/agents-api/harness-onboarding.md b/contracts/agents-api/harness-onboarding.md index 81bf61861..1ad908c0c 100644 --- a/contracts/agents-api/harness-onboarding.md +++ b/contracts/agents-api/harness-onboarding.md @@ -1,18 +1,11 @@ # Add a native Harness to OpenAgentCore -A **Harness** is a native agent engine (Codex, Claude Code, MiniMax Code) that -runs the model/tool loop. A **Harness adapter** translates the common -Core–Runtime execution contract into that engine's SDK or protocol. This guide is -the single numbered path for adding one. Qualification evidence and the -per-operation support table live in [Harness integration](harnesses.md). +A **Harness** is a native agent engine (Codex, Claude Code, MiniMax Code) that runs the model and tool loop. A **Harness adapter** translates the Runtime's Executor and Turn contract into that engine's SDK or protocol. This document is the Runtime–Harness protocol: the adapter interfaces and their lifecycle obligations, registration, Core qualification and acceptance. [Harness capabilities](harness-capabilities.md) records what each current Harness supports. Start from two entry points: -- [`internal/harnessconfig/harness.go`](../../internal/harnessconfig/harness.go): - the shared model configuration contract (declarations and preparation). -- [`agent/harness.go`](../../apps/daemon/internal/agent/harness.go): the - execution lifecycle, explicit extension contracts and registration methods. - +- [`internal/harnessconfig/harness.go`](../../internal/harnessconfig/harness.go): the shared model configuration contract (declarations and preparation). +- [`agent/harness.go`](../../apps/daemon/internal/agent/harness.go): the execution lifecycle, the extension contracts and the registration methods. ## Ownership @@ -31,293 +24,155 @@ Runtime: Executor preparation, reuse, idle expiry, recovery | Component | Responsibility | Location | | --- | --- | --- | | Core | Public API, authority, durable state, scheduling and configuration snapshots | `services/core` | -| Runtime | Authenticated connection, shared capability preparation and common Executor/Turn lifecycle | `apps/daemon/internal/dispatch` | +| Runtime | Authenticated connection, shared capability preparation and the common Executor and Turn lifecycle | `apps/daemon/internal/dispatch` | | Adapter | Native configuration, resources, API calls, event translation and restrictions | `apps/daemon/internal/agent/` | -| Harness | Native model/tool loop and history | Pinned SDK or executable | +| Harness | Native model and tool loop and history | Pinned SDK or executable | | Service profile | Pure validation of qualified operations and placements | `services/core/internal/engine` | | Registration | Installed factories and verified capability declarations | `apps/daemon/internal/cli` | -An Environment supplies execution resources. Managed Docker, E2B and user-managed -machines differ in provisioning and connection; their connected Runtime uses this -same contract. Operating-system support belongs in the implementation and its -qualification. The native daemon supports Linux, macOS and Windows; each adapter -declares its qualified platform scope. Managed Providers remain Linux-only. -A platform-neutral interface alone does not qualify a harness on another platform. -See [self-hosted platforms](../../docs/getting-started/self-hosted.md#platforms) for the current -acceptance limits. Runtime connection, installed capability snapshot, Session Executor and Turn -each have their own lifetime; see -[Executor and Turn lifetimes](../../docs/runtime-protocol.md#executor-and-turn-lifetimes). -Native factories receive -capabilities only after the common Runtime has loaded its bound installed snapshot; -see [capability preparation](environments.md#runtime-capability-preparation). -Model providers supply -model communication configuration, not Turn scheduling or native process ownership. +An Environment supplies execution resources. Managed E2B, Docker and microsandbox machines and application-owned machines differ in provisioning and connection; the connected Runtime uses this same contract. The daemon runs on Linux, macOS and Windows, managed Providers are Linux-only, and each adapter qualifies its own platforms ([self-hosted platforms](../../docs/getting-started/self-hosted.md#platforms)). Native factories receive capabilities only after the Runtime has loaded the bound installed snapshot ([capability preparation](environments.md#runtime-capability-preparation)). Model providers supply model communication settings, not Turn scheduling or native process ownership. ## Steps -1. **Pin the native source.** Record the upstream package version and source - revision and document the native entry point next to the adapter. -2. **Implement the adapter** in `apps/daemon/internal/agent/`: an - `ExecutorFactory`, an `Executor` and a `Turn`. See - [Required adapter interfaces](#required-adapter-interfaces). Reuse shared - process, credential/configuration and local workspace helpers. -3. **Register the kind in the Runtime** in `apps/daemon/internal/cli`. - See [Register the adapter](#register-the-adapter). -4. **Add the service profile and one catalog entry.** See - [Add the engine to Core](#add-the-engine-to-core). -5. **Package native prerequisites.** Add a Runtime image under - `services/core/deploy/` and, optionally, - [native installer participation](#native-installer-participation). -6. **Enable and select the engine** through - [engine selection](#engine-selection). -7. **Qualify it.** Run the synthetic integration test, then the real acceptance in - [Harness integration](harnesses.md#acceptance-checklist). Record results in - the [qualification table](harnesses.md#current-qualified-operations). - -Implement the mandatory text lifecycle and explicitly handle every extension. -Qualify supported extensions one at a time; an unqualified extension returns -`agent.ErrUnsupportedOperation` without native effects. Call the reusable `agent/contracttest.TextLifecycle` assertions with the -adapter's prepared Executor and deterministic native fixture. These assertions -cover healthy reuse, durable input and cancellation for adapters whose native -owner remains reusable. A native cancellation may instead require retirement: -`Reusable=false` carries a reason and the caller must confirm `Executor.Close`. -Do not force reuse to fit a test helper. Keep native fault and live acceptance -separate. Name the entry test `TestSharedTextLifecycle` so -`make check-runtime-contract` includes it. Do not copy an adapter's native -limitations into the shared Core protocol. +1. **Pin the native source.** Record the upstream package version and source revision and document the native entry point next to the adapter. +2. **Implement the adapter** in `apps/daemon/internal/agent/`: an `ExecutorFactory`, an `Executor` and a `Turn` ([required interfaces](#required-adapter-interfaces), [lifetimes](#executor-and-turn-lifetimes)). Reuse the shared process, credential, configuration and local workspace helpers. +3. **Register the kind in the Runtime** in `apps/daemon/internal/cli` ([register the adapter](#register-the-adapter)). +4. **Add the service profile and one catalog entry** ([add the engine to Core](#add-the-engine-to-core)). +5. **Package native prerequisites.** Add a Runtime image under `services/core/deploy/` and, optionally, [native installer participation](#native-installer-participation). +6. **Enable and select the engine** with the `core.harnesses` setting and [Harness selection](model-execution.md#harness-selection). +7. **Qualify it** ([qualify the adapter](#qualify-the-adapter)) and record the result in [Harness capabilities](harness-capabilities.md). + +Implement the mandatory text lifecycle and handle every extension explicitly. Qualify supported extensions one at a time; an unqualified extension returns `agent.ErrUnsupportedOperation` without native effects. A native cancellation may require retirement instead of reuse: `Reusable=false` carries a reason and the caller must confirm `Executor.Close`. Do not force reuse to fit a test helper, and do not copy an adapter's native limitations into the shared Core protocol. ## Architecture rules -- Codex, Claude Code and future Harnesses have equal architectural status. The - common Runtime wire protocol and Executor/Turn interfaces own lifecycle, input - receipts, cancellation, recovery and resource access. Each adapter keeps its - native implementation and model/tool loop. -- A new engine supplies an adapter, a qualified profile, registration and an - independently verified deployment. It adds no engine-name branches to API - handlers, persistence, dispatch, scheduling or Environment providers. -- Keep required lifecycle declarations, extension interfaces and registration - methods in `agent/harness.go`. Result types, errors and Registry storage may - stay in focused files. Keep this guide linked to that entry point. -- Use the existing `proto.SupportedAgentKind` and `AgentKindCapabilities` - schema. Do not add a second capability descriptor or a combined optional - interface. -- Onboarding does not require feature equality. Verify common lifecycle - obligations and use the same public assertions for each declared operation. - Native differences do not block onboarding, but an omitted declaration or - missing extension implementation does. +- Codex, Claude Code and future Harnesses have equal standing. The common Runtime wire protocol and Executor and Turn interfaces own lifecycle, input receipts, cancellation, recovery and resource access; each adapter keeps its native implementation and model and tool loop. +- A new engine supplies an adapter, a qualified profile, registration and an independently verified deployment. It adds no engine-name branches to API handlers, persistence, dispatch, scheduling or Environment providers, and no handler, store table, scheduler, event projector or model loop for capabilities the contract already represents. +- Keep required lifecycle declarations, extension interfaces and registration methods in `agent/harness.go`. Result types, errors and Registry storage may stay in focused files. +- Use the existing `proto.SupportedAgentKind` and `AgentKindCapabilities` schema. Do not add a second capability descriptor or a combined optional interface. +- Onboarding does not require feature equality. Harnesses need not match each other's optional features, and MCP, functions, images or verbosity control are not required to register. Verify the common lifecycle obligations and use the same public assertions for each declared operation. An omitted declaration or a missing extension implementation blocks onboarding; a native difference does not. +- The service profile catalog is the qualification boundary. Unknown profiles fail closed, and a Runtime heartbeat cannot authorize new public functionality. Schema validity, service qualification and the available Runtime are independent checks. - Never equate accepted parameters with applied native behavior. -## Native model configuration +## Required adapter interfaces -[`internal/harnessconfig/harness.go`](../../internal/harnessconfig/harness.go) owns -the shared configuration declaration and pure preparation contract. Each adapter -supplies one `Configuration` to Core's composition and Runtime's `RegisterKind`. -All three Runtime entry paths validate through that declaration before native -side effects: direct factory, preparation and Executor. Registry wrappers retain -the declaration alongside the factory. Lifecycle and cleanup ownership remain in -[`agent/harness.go`](../../apps/daemon/internal/agent/harness.go). -The shared wire object is `proto.HarnessConfig`. - -A supplied `model` must be a nonempty string, and an explicit `model_provider` -requires it. The native-owned connection path may omit both; explicit null is -invalid. An explicitly empty adapter declaration accepts no provider or nonempty -native parameters. It does not advertise provider support. Unknown protocol -formats and duplicate protocol declarations fail at registration. -Adapter declarations in -`internal/harnessconfig/` own these native fields: - -| Harness | Accepted native fields | Application | +[`agent/harness.go`](../../apps/daemon/internal/agent/harness.go) is the interface entry point. The required lifecycle is `ExecutorFactory`, `Executor`, `Turn` (including `DurableSteerer`) and `TurnSettlement`. Required methods perform their native obligations; returning Unsupported is not an implementation of cancellation, receipts, settlement or cleanup. Turn and workspace extension interfaces stay small and separate, but every public adapter implements each one explicitly. All use the neutral protocol types. + +For example, the Codex adapter keeps its app-server and thread, the Claude adapter one streaming Query, and the MiniMax adapter its ACP connection and native session. All expose the same Executor and Turn contract. Native callbacks and resources stay inside the adapter; the Runtime owns admission, idle expiry and replacement. Cancellation targets the exact Turn through `agent.Session`, and the adapter supplies native completion evidence to the Runtime. + +| Interface or contract | Required handling | Obligation | | --- | --- | --- | -| Codex | `model_reasoning_effort`: `none`, `minimal`, `low`, `medium`, `high`, `xhigh` | Native app-server `-c model_reasoning_effort=...` and each Turn's `collaborationMode.settings.reasoning_effort` | -| Claude SDK | `effort`: `low`, `medium`, `high`, `xhigh`, `max`; `thinking`: the SDK's `adaptive`, `enabled` or `disabled` object | SDK `Options.effort` and `Options.thinking` | -| MiniMax Code | Empty object only | Existing provider token limits remain required; additional native generation parameters are not qualified | - -Claude `thinking` accepts `display` (`summarized` or `omitted`) for adaptive or -enabled thinking, and an optional positive integer `budgetTokens` only for -`enabled`. The deprecated `maxThinkingTokens` alternative is rejected. -These are native settings, not a shared reasoning vocabulary; model availability -and provider support remain the selected harness's responsibility. - -Native protocol and parameter declarations also feed Core administration's small -configuration-support descriptor. Its ordered `protocols` list is the sole source -for accepted protocols and the default (the first entry). Core and Runtime reject -unsupported combinations through the same declaration. The current protocol -matrix belongs to [model execution](model-execution.md#saved-defaults-and-precedence). -Adapters connect directly through native configuration; they must not introduce -a model API proxy or protocol converter. Follow the per-input capability -requirements in the execution lifecycle contract; do not add a second model capability registry -or infer capabilities from model names. Native protocol acceptance and remote -model support remain separate facts. - -Claude's private bridge receives compiled native options and performs structural -checks only, not a second copy of the declaration's rules. Planned fields and -ownership are in the [unified model configuration design](model-configuration-design.md). +| `ExecutorFactory`, `Executor.StartTurn`, `Executor.Close` | Real implementation | Prepare without model input; keep ownership of failed or uncertain resources; confirm cleanup | +| `Turn`, `Session.Cancel`, `CancellationOutcome`, `AwaitSettlement` | Real implementation | Cancel the exact Turn, keep observed results and confirm settlement independently of cancellation requests | +| `DurableSteerer` | Real implementation on every Turn | Distinguish a complete write from the native application receipt; keep retry identity | +| `Steerer` | Explicit implementation or Unsupported | Additional non-durable active-Turn input | +| `FunctionResultSubmitter` | Explicit implementation or Unsupported | Match native call and result identity and acknowledge application | +| `PermissionResponder`, `UserChoiceResponder` | Explicit implementation or Unsupported | Respond to exact emitted identities; unknown or expired interactions stay distinct from Unsupported | +| `WorkspaceReader`, `WorkspaceDirectoryLister`, `WorkspaceWriter` | Explicit on Turn, Executor and Prepared owners | Use the authorized workspace, confirm access, commit or close, or return the operation's Unsupported error | +| `Prepared`, `PreparedCancellation` | Real implementation for an executable preparation | Keep resource and output ownership across Start, cancellation and unused cleanup | +| Neutral messages, images, MCP, structured output and Subagent observations | Explicit capability decisions | Keep each operation's protocol semantics; reject unsupported input before submission | +Each adapter's `contracts.go` holds an individual compile-time assertion for each small interface. Do not embed a default implementation that makes future interfaces appear implemented. Adding a contract also requires a classification in the common completeness check and an explicit assertion in every public adapter; the check follows the authored Harness catalog. -## Required adapter interfaces +For a design-level refusal, implement the method directly: + +```go +func (s *Session) SubmitFunctionResult(context.Context, proto.FunctionResultPayload) error { + return fmt.Errorf("%w: native public function tools are not qualified", agent.ErrUnsupportedOperation) +} +``` + +The reason is a fixed safe string, never submitted content, a credential or raw native diagnostics. Unsupported guarantees no native side effect and is not a successful empty operation. Installation unavailability, unknown interaction IDs, native failures and uncertain outcomes keep their own errors and ownership. A nil `Turn` still means that no input was submitted and the output stays with the caller; never use it as an Unsupported marker. -[`agent/harness.go`](../../apps/daemon/internal/agent/harness.go) is the -canonical interface entry point. Its required lifecycle is `ExecutorFactory`, -`Executor`, `Turn` (including `DurableSteerer`) and `TurnSettlement`. Required -methods must perform their native obligations; returning Unsupported is not an -implementation of cancellation, receipts, settlement or cleanup. Turn and workspace -extension interfaces remain small and separate, but every public adapter implements -each explicitly. Their result types and errors stay in focused operation files. All use the existing neutral protocol types. +Workspace capability describes the actual Runtime and resource-owner combination. The Codex and MiniMax resource objects reject native workspace access while the common authorized `localworkspace` owner provides it; Claude can expose native read and list access, and the common owner provides writes. Interface presence alone never selects a resource or advertises support. -The [Core–Runtime lifecycle contract](../../docs/runtime-protocol.md#executor-and-turn-lifetimes) -owns preparation failure, partial StartTurn results, output closure, settlement, -reuse and cleanup. Implement those obligations through the interfaces above. -Native callbacks and resources stay inside the adapter; Runtime owns admission, -idle expiry and replacement. +The service profile qualifies public combinations and the Runtime advertises the installed combination; neither replaces schema validation or Project authorization. Native behavior tests must agree with the declarations. An advertised operation that returns Unsupported is a contract violation, never success or grounds for replay. -For example, a Codex adapter retains its app-server and thread, a Claude adapter -retains one streaming Query, and a MiniMax adapter retains its ACP connection and -native Session. Their implementations expose the same Executor and Turn contract. +## Executor and Turn lifetimes -### Cancellation ownership +| Lifetime | Owner | Ends when | +| --- | --- | --- | +| Environment allocation | Sandbox Provider | Explicit reclamation, coordinated with Runtime execution | +| Runtime connection | Runtime transport | Disconnection or replacement by a newer connection | +| Installed capability snapshot | Runtime | Its Environment is reclaimed; never on Executor close | +| Session Executor | Runtime | `Executor.Close` on idle expiry, shutdown or confirmed invalidation | +| Turn | Adapter `Turn`, tracked by the Runtime | `AwaitSettlement` confirms settlement | + +A Session owns one reusable Executor in its connected Runtime; a Turn owns one input execution, its output stream and its cancellation. `agent.ExecutorFactory` prepares the fixed configuration without model input, and `Executor.StartTurn` creates a new `agent.Turn` without replacing healthy native resources. Normal completion settles only the Turn. `Executor.Close` releases native resources on idle expiry, Environment shutdown or confirmed invalidation; it releases neither the Environment allocation nor the workspace. Core keeps no second Executor cache. The same lifecycle applies to hosted, self-hosted and `none` placements. + +**Binding.** The Runtime binds its Executor record to the Session, Environment, connection and immutable execution configuration. Resume identity and prior-Turn recovery flags are continuity assertions, not configuration changes. A supplied native identity must match the retained owner, and when existing history is required, recovery never starts a new root. A configuration conflict is an error, not a hot switch. A lost connection retires its owners and handles; old timers, output and cancellation cannot affect their replacements. + +**Per-Turn state.** Each Turn gets a fresh wrapper, output channel and receipt state. Steering, function, permission and user-choice interfaces belong to that Turn. Native callbacks capture the originating Turn before asynchronous work, so a late event is never attributed to whichever Turn is active. Native processes, query or transport connections, fixed capability configuration and native session identity belong to the Executor. Do not reset completed `sync.Once` values or reuse an old Turn object. -Implement cancellation on the exact Turn through `agent.Session`. Follow the -[lifecycle and settlement rules](../../docs/runtime-protocol.md#executor-and-turn-lifetimes); -the adapter must supply native completion evidence to the shared Runtime. -Permission and user-choice responses use the explicit extension interfaces below. +**Start.** A nil Turn from `StartTurn` guarantees that no native input was submitted and the output channel was not retained; the Runtime then closes the channel. Once input may have been submitted, return a non-nil Turn even with an error: that Turn owns exactly-once output closure and stays tracked until settlement. Unknown input is never replayed. A definite `executor_unavailable` Start rejection allows one common recovery attempt, only after the previous Executor has been closed and no input was submitted; the Runtime rechecks the same physical peer and the current authorization. + +**Cancellation and settlement.** `Turn.Cancel` targets only that Turn and does not close a healthy Executor. `AwaitSettlement` applies after both natural completion and cancellation. Success means output can no longer be written and the Turn's native events, input, functions, interactions and child work have settled. Native completion or cancellation confirmation is independent of resource retirement: closing a transport cannot supply a missing native terminal or operation receipt. + +- `Reusable=true` also confirms that the native owner can accept the next Turn. `Reusable=false` requires a reason and a later confirmed Executor close. +- An error means settlement is unconfirmed and frees neither ownership nor capacity. Caller deadlines stop the wait, not the tracked cleanup. Retry the same cleanup target serially; a failed cleanup blocks replacement and keeps its resource slot. +- `Executor.Close` confirms resource retirement independently of the Turn outcome: an immutable Turn error must not prevent closing the native transport once its work and output have stopped. +- Include owned background work in settlement and keep the exact native cleanup target after a failure. Native termination belongs to the adapter; a bulk cleanup acknowledgement alone does not establish quiescence. +- The observed cancellation outcome keeps native identity, Usage and output without fabricating missing evidence. + +**What the Runtime does around a Turn.** One output consumer starts before native Start, drains the bounded 64-frame channel and keeps the terminal observation until Start publication, Turn settlement and admitted operation receipts finish. Natural completion never calls Cancel. Input, function and interaction admission close before settlement; operations already admitted hold their barrier through native receipts and outbound acknowledgement. The Runtime sends cancellation to the Turn before waiting on that barrier, because a written input may need a native interrupt to produce its receipt. It joins native settlement, any required confirmed Executor close, output drain and all admitted operations before an applied acknowledgement or reuse, and only then forwards Done or an applied cancellation receipt. A failed Close can report failure while keeping the same Run and outstanding operations for retry; a closed caller wait cannot manufacture an applied input receipt. The Runtime commits native continuity and releases the old Run's admission before publishing Done, since the receiver may start another Turn at once; a late terminal-send failure belongs to the old Run and cannot invalidate a successor that already owns the Executor. Connection shutdown owns transport-loss cleanup. The settlement wait is ten seconds and the receipt send budget five seconds; a timeout is not proof of quiescence. ## Events, inputs and optional capabilities -Use [`internal/agentdaemon/proto`](../../internal/agentdaemon/proto) for neutral -requests, events and receipts. Each Turn emits only its own events with its Run ID, -in order, and one terminal outcome. Native IDs and usage must be observed rather -than invented. Capture the originating Turn before asynchronous native callbacks; -never attribute a late result to whichever Turn is currently active. - -Initial input and steering use ordered `proto.MessageInput`. Preserve user-message -and content order. Text-only adapters reject images through `TextOnly()` instead -of dropping them. A successful transport write is distinct from confirmed native -application. User-choice answers use the emitted question ID and an array of -values; shared `PromptForUserChoiceDecisionPayload.AnswersFor` validates identity -before consuming a pending interaction. Do not map answers by header or position. -Resume only the exact history bound to the Session; missing or ambiguous required -history fails before new model input. +Use [`internal/agentdaemon/proto`](../../internal/agentdaemon/proto) for neutral requests, events and receipts. Each Turn emits only its own events with its Run ID, in order, and one terminal outcome. Native IDs and usage are observed, never invented; a missing measurement is unknown, not zero. -| Interface or contract | Required handling | Obligation | -| --- | --- | --- | -| `ExecutorFactory`, `Executor.StartTurn`, `Executor.Close` | Real implementation | Prepare without model input; keep failed or uncertain resource ownership; confirm cleanup | -| `Turn`, `Session.Cancel`, `CancellationOutcome`, `AwaitSettlement` | Real implementation | Cancel the exact Turn, preserve observed results and confirm settlement independently of cancellation requests | -| `DurableSteerer` | Real implementation on every Turn | Distinguish complete write from native application receipt; preserve retry identity | -| `Steerer` | Explicit implementation or Unsupported | Additional non-durable active-turn input | -| `FunctionResultSubmitter` | Explicit implementation or Unsupported | Match native call/result identity and acknowledge application | -| `PermissionResponder`, `UserChoiceResponder` | Explicit implementation or Unsupported | Respond to exact emitted identities; unknown/expired interactions remain distinct from Unsupported | -| `WorkspaceReader`, `WorkspaceDirectoryLister`, `WorkspaceWriter` | Explicit on Turn, Executor and Prepared owners | Use the authorized workspace, confirm access/commit/close, or return the operation's Unsupported error | -| `Prepared`, `PreparedCancellation` | Real implementation for an executable preparation | Preserve resource and output ownership across Start, cancellation and unused cleanup | -| Neutral messages, images, MCP, structured output and Subagent observations | Explicit capability decisions | Preserve each operation's protocol semantics; reject unsupported input before submission | - -Each adapter's `contracts.go` contains individual compile-time assertions for these -small interfaces. Do not embed a default implementation that makes future -interfaces appear implemented. Adding a contract also requires classification in -the common completeness check and an explicit assertion for every public adapter; -the check follows the authored Harness catalog. - -For a design-level refusal, implement the method directly, for example: +Initial input and steering use ordered `proto.MessageInput`. Keep user-message and content order. Text-only adapters reject images through `TextOnly()` instead of dropping them; image adapters translate each part natively and acknowledge an active batch only after all its messages are applied. A successful transport write is distinct from confirmed native application. User-choice answers use the emitted question ID and an array of values; the shared `PromptForUserChoiceDecisionPayload.AnswersFor` validates identity before consuming a pending interaction, so never map answers by header or position. Resume only the exact history bound to the Session; missing, ambiguous or foreign history fails before new model input. Device identity is not native session ownership. -```go -func (s *Session) SubmitFunctionResult(context.Context, proto.FunctionResultPayload) error { - return fmt.Errorf("%w: native public function tools are not qualified", agent.ErrUnsupportedOperation) -} -``` +### Required and extension operations + +The public text path requires durable Turns, applied input receipts, ordered observations, cancellation and enforcement of disabled execution controls; `execution.Policy.engineCapabilities` holds the exact requirements. An engine without native tools can guarantee their absence; an engine with tools must actually disable them when asked. Accepting a configuration is not proof of enforcement. -The reason is a fixed safe string, never submitted content, a credential or raw -native diagnostics. Unsupported guarantees no native side effect. It is not a -successful empty operation. Installation unavailability, unknown interaction IDs, -native failures and uncertain outcomes keep their existing errors and ownership. -A nil `Turn` still means no input was submitted and output remains with the caller; -it must not be repurposed as an Unsupported marker. - -Workspace capability describes the actual Runtime/resource-owner combination. -Codex and MiniMax resource objects explicitly reject native workspace access while -the common authorized `localworkspace` owner provides it. Claude can expose native -read/list access; writes are provided by the common owner. Interface presence alone -must never select a resource or advertise support. - -The service profile qualifies public combinations; the Runtime advertises the -installed combination. Neither replaces schema validation or tenant authorization. -Native behavior tests must agree with supported declarations. An advertised -operation returning Unsupported is a contract violation, never success or grounds -for automatic replay. +MCP, public functions, deferred function discovery, structured output, image input, verbosity controls and other optional operations need not match another engine. Reject an unqualified combination with Unsupported and record the gap; never advertise a capability to bypass selection. + +- Structured output: consume `ExecutionControls.OutputFormat` and publish confirmed native output through the Message contract ([execution tools](execution-tools.md#structured-output)). Register the public qualification separately from the Runtime capability. +- Images: register the Runtime's `MessageImages` and qualify the profile's `MessageImages` separately ([message input](message-input.md)). +- Workspace placements additionally need verified preparation, workspace reads and output export and the dedicated Runtime binding with the shared Files helpers. Enable a placement only after its lifecycle behavior is demonstrated. + +### MCP origin and native limits + +Declare supported public origins in the engine profile's `MCPOrigins` and bearer support in `MCPBearer`. The Runtime advertises its actual HTTP, bearer and required-initialization capabilities. Shared admission validates origin and placement; adapter validation keeps native label, allowlist and initialization limits. + +Consume `agent.ResolveMCPBindings` for public and installed declarations, and keep origin, credential authority, null versus empty allowlists and required startup. Do not copy tokens into native profiles or reinterpret a service request as an Environment request. Reject unsupported native policies instead of dropping them. Follow the [MCP origin contract](environments.md#public-mcp-connection-origin) and run public-client, failure, cancellation and cold-recovery qualification for each advertised combination. Model capability is separate from Harness transport support; never infer it from model names or silently degrade input. + +### Subagent observations + +A Harness that supports the Subagent reads implements the [neutral observation contract](subagents.md#adapter-contract). It reports verified child identity, lifecycle effects and owned Turn and Item history through the authenticated Run and qualifies those facts with real execution, without routes, storage branches or a Harness-specific scheduler. Report unsupported native facts explicitly; completing a child task is not closing its Subagent. Native background work stays owned through settlement and cancellation. ## Register the adapter -Registration is static and requires a build; dynamic plugins are outside this -contract. The methods live in `agent/harness.go`, and built-in adapters call them -from [`cli/agent_registration.go`](../../apps/daemon/internal/cli/agent_registration.go) -(Codex, MiniMax Code) and [`cli/claude_sdk.go`](../../apps/daemon/internal/cli/claude_sdk.go) -(Claude Code). +Registration is static and requires a build. The methods live in `agent/harness.go`, and the built-in adapters call them from [`cli/agent_registration.go`](../../apps/daemon/internal/cli/agent_registration.go) (Codex, MiniMax Code) and [`cli/claude_sdk.go`](../../apps/daemon/internal/cli/claude_sdk.go) (Claude Code). | Order | Method | Registers | | --- | --- | --- | -| 1 | `RegisterKind(proto.SupportedAgentKind, harnessconfig.Configuration, agent.Factory)` | Kind, availability, version, `AgentKindCapabilities`, the adapter's model configuration declaration and the direct-call factory. Resets the other registrations, so call it first. | -| 2 | `RegisterExecutor(kind, agent.ExecutorFactory)` | The shared Executor/Turn lifecycle used for execution. Derives the `Preparation` capability. | -| 3 | `RegisterPreparation(kind, workspaceRead, agent.PreparationFactory)` | Optional. Separate read-only workspace preparation when qualified workspace operations need it. | - -The direct-call `agent.Factory` should delegate to the same Executor -implementation. Every `proto.AgentKindCapabilities` field must be explicitly -`proto.CapabilitySupported` or `proto.CapabilityUnsupported`. -`proto.CapabilityUnspecified` is invalid: zero values and omitted fields do not mean -Unsupported. Installation probes may use `proto.CapabilityFromBool` for an -individual field; they must not populate all unmentioned or future fields. -Availability remains separate in `SupportedAgentKind.Available`. - -Registration and wire decoding validate the complete declaration. The wire carries -an explicit boolean for every field; omitted and null fields are invalid. A new -field requires a decision by every production declaration. Runtime consumers use -`IsSupported()` and reject unsupported requests before native operations; an -interface assertion only verifies implementation, never support. Every declaration -must match behavior verified for that installation. - -The admission mapping is explicit: `Steering` controls non-durable `Steerer` input; -`DurableInputReceipts` controls `DurableSteerer` input and also requires the Turn -settlement contract. Neither implies the other. Core's current public text profile -requires both advertised capabilities. `Permissions` qualifies permission and -user-choice responses together; a supported declaration requires both native -response paths. Workspace declarations describe the selected authorized resource -owner, including the common Runtime workspace implementation. - -Runtime registration does not grant Core qualification; -that belongs to the service profile. - -The runnable test-only example is -[`testdata/onboarding/main.go`](../../apps/daemon/testdata/onboarding/main.go). -It registers a text-only synthetic Harness, demonstrates a Session-owned -Executor, fresh Turns, durable steering, cancellation and history binding, and is -never shipped as a real engine. +| 1 | `RegisterKind(proto.SupportedAgentKind, harnessconfig.Configuration, agent.Factory)` | Kind, availability, version, `AgentKindCapabilities`, the model configuration declaration and the direct-call factory. It resets the other registrations, so call it first. | +| 2 | `RegisterExecutor(kind, agent.ExecutorFactory)` | The Executor and Turn lifecycle used for execution; derives the `Preparation` capability | +| 3 | `RegisterPreparation(kind, workspaceRead, agent.PreparationFactory)` | Optional: separate read-only workspace preparation for qualified workspace operations | + +The direct-call `agent.Factory` delegates to the same Executor implementation. + +Every `proto.AgentKindCapabilities` field must be explicitly `proto.CapabilitySupported` or `proto.CapabilityUnsupported`, even for an unavailable Harness. `proto.CapabilityUnspecified` is invalid: zero values and omitted fields never mean Unsupported. An installation probe may set an individual field with `proto.CapabilityFromBool`; it must not populate unmentioned or future fields. Availability stays separate in `SupportedAgentKind.Available`. Registration validates the complete declaration before changing the registry, and the wire carries an explicit boolean for every field, so omitted and null fields are invalid. A new field requires a decision in every production declaration. Runtime consumers use `IsSupported()` and reject unsupported requests before native operations; an interface assertion verifies implementation, never support. Every declaration must match the behavior verified for that installation; the [Core–Runtime protocol](../../docs/runtime-protocol.md#explicit-capability-declarations) owns how declarations travel and are frozen. + +The admission mapping is explicit. `Steering` controls non-durable `Steerer` input. `DurableInputReceipts` controls `DurableSteerer` input and also requires the Turn settlement contract; neither implies the other, and Core's public text profile requires both. `Permissions` qualifies permission and user-choice responses together and requires both native response paths. Workspace declarations describe the authorized resource owner, including the common Runtime workspace implementation. Runtime registration does not grant Core qualification; the service profile does. + +The runnable test-only example [`testdata/onboarding/main.go`](../../apps/daemon/testdata/onboarding/main.go) registers a text-only synthetic Harness. It shows a Session-owned Executor, fresh Turns, durable steering, cancellation and history binding, and is never shipped. ## Add the engine to Core -Core recognizes the [built-in Harness registrations](harness-catalog.md). -Add one entry to `internal/harnessconfig/builtin/catalog.json` with: +Core recognizes the [built-in Harness registrations](harness-catalog.md). Add one entry to `internal/harnessconfig/builtin/catalog.json` with: - the public `kind` and display `label`; - the model `configuration` package under `internal/harnessconfig`; - the `profile` constructor under `services/core/internal/engine`. -Implement the profile constructor, then run `make generate-harness-catalog`. -This generates the model configuration registry, Core profile catalog, client -identifiers/display names and registration reference. Public input validators read -the generated registry. `make openapi` derives Harness enums from the same authored -catalog; do not add handwritten enums to DTO tags or route annotations. -`make check-harness-catalog` rejects stale projections. - -Runtime registration uses `builtin.Configuration(kind)` for public Harnesses and -separately registers native factories, probes and installed capability evidence. -The catalog cannot declare a machine's availability. Adapter discovery and -packaging still require their own implementation and qualification; no dynamic -plugin loader is introduced. - -The profile is pure: it declares supported placements, public configuration and -result limits and required Runtime controls, using existing public/protocol -types. Profile callbacks cannot query business data, decrypt credentials or -control native processes. Public schema validation, qualified engine support and -actual Runtime capabilities remain separate checks; Runtime advertisements alone -never enable public operations. Shared dispatch checks capability combinations, -not a whitelist of engine names. +Implement the profile constructor, then run `make generate-harness-catalog`. It generates the model configuration registry, Core profile catalog, client identifiers and display names and the registration reference; public input validators read the generated registry. `make openapi` derives Harness enums from the same catalog, so do not add handwritten enums to DTO tags or route annotations. `make check-harness-catalog` rejects stale projections. + +The Runtime registers public Harnesses with `builtin.Configuration(kind)` and separately registers native factories, probes and installed capability evidence. The catalog cannot declare a machine's availability, and there is no dynamic plugin loader. + +The profile is pure: it declares supported placements, public configuration and result limits and required Runtime controls, using existing public and protocol types. Profile callbacks cannot query business data, decrypt credentials or control native processes. Shared dispatch checks capability combinations, not a whitelist of engine names. ### Explicit service qualification @@ -334,83 +189,41 @@ An omitted or unknown policy, missing required callback, or callback paired with Run the `engine` and `execution` tests for omission, policy, combination and error precedence coverage, and the public onboarding/store tests for admission and Runtime dispatch. Test fixtures use `engine/enginetest`, whose exhaustive literal also requires a decision when a field is added; it is not a production profile. -`execution.Policy` supplies immutable service qualification to HTTP admission, -Worker device selection and final dispatch. Custom composition gives the same -Policy to `api.WithExecutionPolicy` and `Dispatcher.Policy`. The zero value uses -built-in profiles; an explicitly empty catalog authorizes none. There is no -mutable global registration or compatibility fallback. - -## Engine selection - -The [Harness selection contract](harness-selection.md) owns the public selector, -deployment enablement, defaults and immutable Session binding. This guide adds no -second selector or fallback rule. - -## Required versus extension operations - -The current public text path requires durable turns, applied input receipts, -ordered observations, cancellation and enforcement of disabled execution controls. -Check `execution.Policy.engineCapabilities` for the exact current requirements. -An engine without native tools can guarantee their absence; an engine with tools -must actually disable them when requested. Configuration acceptance is not proof -of enforcement. - -MCP, public function calls, deferred function discovery, structured output, image inputs, verbosity controls and other optional -operations do not need to match another engine. Reject unqualified combinations -with Unsupported and record the gap. Never advertise a capability to bypass selection. - -Structured-output adapters consume `ExecutionControls.OutputFormat` and publish -confirmed native output through the existing Message contract. Register public -qualification separately from the Runtime capability; see the -[structured-output boundary](structured-output.md). No Core engine-name branch is required. - -Initial requests, `Executor.StartTurn` and steering consume the same ordered -`proto.MessageInput`. Text-only adapters use `TextOnly()` to reject images without -discarding content. Image adapters translate each part natively and acknowledge -an active batch only after all its messages are applied. Register -Runtime `MessageImages` and qualify the engine profile’s `MessageImages` separately; see the -[message-input contract and real acceptance](message-input.md). - -Hosted workspace execution additionally requires verified preparation, workspace -reads/output export, network behavior and credential/history isolation. Reuse the -same dedicated Runtime binding and shared Files helpers. A native Bash sandbox -alone does not establish isolation for other native file tools. Enable a placement -only after its required security and lifecycle behavior is demonstrated. +`execution.Policy` supplies immutable service qualification to HTTP admission, Worker device selection and final dispatch. Custom composition gives the same Policy to `api.WithExecutionPolicy` and the Core dispatcher's `Policy`. The zero value uses the built-in profiles; an explicitly empty catalog authorizes none. There is no mutable global registration. -### MCP origin and native limits +## Native model configuration -Declare supported public origins in the existing engine profile's `MCPOrigins` -and bearer support in `MCPBearer`. Runtime advertises actual HTTP, bearer and -required-initialization capabilities. Shared admission validates origin and -placement; adapter validation retains native label, allowlist and initialization -limits. These are separate checks, not a second MCP executor. - -Consume `agent.ResolveMCPBindings` for public and installed declarations; preserve -origin, credential authority, null versus empty allowlists and required startup. -Do not copy tokens into native profiles or reinterpret a service request as an -Environment request. Reject unsupported native policies instead of dropping them. -Follow [the MCP origin contract](environments.md#public-mcp-connection-origin) -and run public-client, failure, cancellation and cold-recovery qualification for -each advertised combination. Model capability remains separate from Harness -transport support; never infer it from model names or silently degrade input. - -## Optional Subagent observations - -A harness that supports the Subagent resource reads implements the existing -[neutral observation contract](subagents.md#common-adapter-contract). It reports -verified child identity, lifecycle effects and owned Turn/Item history through -the authenticated Run, then qualifies those facts with real execution. It does -not add routes, storage branches or a harness-specific Core scheduler. Report -unsupported native facts explicitly; completing a child task is not closing its -Subagent. Native background work must remain owned through settlement and cancel. - -## Contract verification - -Run `make check-runtime-contract`, the three adapter test packages and `make check`. -The common completeness gate covers capability omissions and interface assertions; -adapter tests must cover actual native semantics, not only method presence. - -| Boundary | Existing focused evidence | +[`internal/harnessconfig/harness.go`](../../internal/harnessconfig/harness.go) owns the shared configuration declaration and pure preparation contract. Each adapter supplies one `Configuration`, in `internal/harnessconfig/`, to Core's composition and to the Runtime's `RegisterKind`. The direct factory, preparation and Executor paths all validate through that declaration before native side effects, and Registry wrappers keep the declaration with the factory. The wire object is `proto.HarnessConfig`. [Model execution](model-execution.md#native-model-parameters) lists each Harness's accepted fields. + +A supplied `model` must be a nonempty string, and an explicit `model_provider` requires it. The native-owned connection path may omit both; explicit null is invalid. An explicitly empty declaration accepts no provider or nonempty native parameters and advertises no provider support. Unknown protocol formats and duplicate protocol declarations fail at registration. + +The declaration's ordered `protocols` list is the only source of accepted protocols and the default (the first entry); it also feeds Core's configuration-support descriptor, and Core and Runtime reject unsupported combinations through it. Adapters connect directly through native configuration; they never introduce a model API proxy or protocol converter, a second model capability registry, or capabilities inferred from model names. Claude's private bridge receives compiled native options and performs structural checks only, not a second copy of the declaration's rules. + +## Qualify the adapter + +Before starting, record the operation set, expected results, exclusions and stopping conditions. A qualification ends when its declared operations pass; it does not expand to match another Harness's feature list. + +1. **Contract tests.** Call `agent/contracttest.TextLifecycle` from a test named `TestSharedTextLifecycle` with the adapter's prepared Executor and a deterministic native fixture; `claudesdk/executor_test.go` is the reference. It checks independent Turn streams, native owner and history continuity, durable write and application receipts, stale cancellation and healthy continuation after cancellation. `make check-runtime-contract` runs it together with the shared wire, gateway, transport and dispatcher tests, the declaration completeness check and each adapter's `TestUnsupportedExtensionsHaveNoNativeEffects`. Adapter tests also cover two ordinary Turns sharing one native process or connection and history, cancellation followed by another Turn, stale cancellation and late events, native exit, cleanup failure, input write and application receipts, unknown outcomes and fresh per-Turn usage, function, input and child-observation state. State whether a fixture is controlled or a real provider. +2. **Shared integration.** `TestThirdHarnessPublicOnboarding` runs the synthetic Harness through public Session and input admission, Worker device selection, the real WebSocket gateway, the daemon Registry and Router, neutral events and durable terminal projection. It uses a custom immutable `engine.Catalog` in the same `execution.Policy` given to the API handler and the dispatcher, and checks applied input receipts, saved native identity, continuation, cancellation, unsupported optional requests and missing mandatory Runtime support. The fixture has no workspace, MCP, public functions, permissions or user-choice handlers, and its registration stays local to the test. It proves the integration path, not native execution. +3. **Real acceptance.** Use the pinned official Python SDK and raw HTTP against Core, a real provider API, the native Harness and a dedicated database. Verify initial execution, a warm follow-up, cancellation and restart with continuation; record native owner identity and same-condition cold and warm timing. For workspace placements also verify Files and Artifacts, workspace identity, that no credentials appear in public responses and that foreign history is rejected. `services/core/tests/official_hosted_functions_native.py` holds the shared function assertions: success and error, native file output and public Artifact bytes, same-history continuation after restart, foreign result rejection and pending-call cancellation. Synthetic or failed runs never count. The opt-in tests below run the pinned-SDK fixtures in `services/core/tests` against a real daemon and model; each runs when `OAC_TEST_OFFICIAL_SDK_PYTHON`, `OAC_TEST_NATIVE_DAEMON_BIN`, `OAC_TEST_NATIVE_PROOF_DIR` and its private options file are set. +4. **Regression.** Existing Harnesses keep working. Run targeted tests, then `make check`; run `make openapi` after API changes and `make sqlc-generate` after query changes. +5. **Review.** Follow the [blind review workflow](../../CONTRIBUTING.md#review). + +| Operation | Test in `services/core/internal/store` | Options file variable, and Harness variable where the test takes one | +| --- | --- | --- | +| Model provider protocols | `TestNativeModelProtocolPublicExecution` | [Model execution](model-execution.md#acceptance) | +| MiniMax Code text | `TestNativeMCodePublicExecution` | `OAC_TEST_MCODE_REAL_OPTIONS` | +| Message images | `TestNativeMessageImagePublicExecution` | `OAC_TEST_MESSAGE_IMAGE_REAL_OPTIONS`, `OAC_TEST_MESSAGE_IMAGE_ENGINE` | +| Function results with images | `TestNativeFunctionImagePublicExecution` | `OAC_TEST_FUNCTION_IMAGE_REAL_OPTIONS`, `OAC_TEST_FUNCTION_IMAGE_ENGINE` | +| Structured output | `TestNativeStructuredOutputPublicExecution` | `OAC_TEST_STRUCTURED_OUTPUT_REAL_OPTIONS` | +| Deferred function discovery | `TestNativeToolSearchPublicExecution` | `OAC_TEST_TOOL_SEARCH_REAL_OPTIONS` | +| Disabled web search and programmatic tool calling | `TestNativeToolPolicyPublicExecution` | `OAC_TEST_TOOL_POLICY_REAL_OPTIONS`, `OAC_TEST_TOOL_POLICY_ENGINE` | + +Environment acceptance uses `services/core/tests/official_environment_{templates,setup,skills,plugins,plugin_mcp,composition,initial_files,network,skill_references}.py`. For composed preparation, change the Skill default and Template, delete the sources, retry and restart; verify frozen bytes, one setup execution and MCP cancellation. `official_hosted_structured_native.py` covers hosted structured output. Record exact source revisions, native versions and commands with each acceptance result. + +Keep provider keys in private operator files, never in commits or logs. Existing focused tests, relative to `apps/daemon/internal/agent`: + +| Boundary | Tests | | --- | --- | | Codex reuse, cancellation and unconfirmed cleanup | `codex/executor_test.go`, `terminal_cleanup_test.go`, `prepared_cancel_test.go` | | Codex input receipts and strict recovery | `codex/function_write_receipt_test.go`, `function_receipt_test.go`, `resume_test.go`, `recovery_test.go` | @@ -419,55 +232,21 @@ adapter tests must cover actual native semantics, not only method presence. | MiniMax native history binding | `mcode/session_test.go` | | Explicit refusals without native effects or fabricated results | Each adapter's `unsupported_test.go` | -These paths are relative to `apps/daemon/internal/agent`. Controlled native -transport fixtures establish failure and ownership behavior; they are not live model -qualification. Preserve the separate native acceptance requirements in -[Harness integration](harnesses.md#acceptance-checklist). - ## Native installer participation -An adapter may supply `agent.Installation` from `installation.go` in its own -package: registered agent kind, pinned version, supported platforms, activation -environment and a bounded -readiness probe. Register it in `cli/native_harness.go` and add its pinned component -to the native distribution builder. This optional contract does not change -Executor/Turn semantics. Runtime owns checksums, copying, locks and additive -installation; adapters own native layout and probes. Validate installation and -actual execution on each advertised platform. Missing or incompatible native -content must fail, never install itself during a Turn. +An adapter may supply `agent.Installation` from `installation.go` in its own package: registered kind, pinned version, supported platforms, activation environment and a bounded readiness probe. Register it in `cli/native_harness.go` and add its pinned component to the native distribution builder. This optional contract does not change Executor and Turn semantics. The Runtime owns checksums, copying, locks and additive installation; adapters own native layout and probes. Validate installation and execution on each advertised platform. Missing or incompatible native content fails; it never installs itself during a Turn. ## Native process ownership -The shared daemon `clirunner` offers opt-in Unix process-group ownership for -adapters whose SDK launches a native child. Existing callers keep their current -process policy. Explicit cancellation and parent-context cancellation share a -TERM grace period and bounded KILL escalation. An internal reaper also cleans -remaining group members when the direct process exits, even if a descendant -still holds stdout open. During cancellation, surviving descendants keep the -remaining TERM grace after the leader exits. Unsupported hosts reject this mode before launch. - -Owned output pipes remain readable after the leader exits. Consumers must drain -stdout and stderr before calling `Wait`, which joins the cached process result -and closes the readers. `Done` reports leader reaping and group cleanup signals; -it is not a native execution receipt or proof of persisted history. SDK adapters -must settle each Turn and drain its observations before publishing completion. -Executor close additionally closes the query and awaits the native child. Process groups are lifecycle supervision, not OS isolation -or containment of descendants that deliberately leave the group. - -All Harness adapters use bypass execution. Do not restore named Codex permission -profiles, bubblewrap wrappers, native Claude sandbox settings or MiniMax -SandboxManager branches. There is one execution path for every Environment origin. -Resource paths are ordinary operator configuration, not a permission boundary. - -The daemon does not enforce disabled or restricted network policies. Such a -combination must be rejected unless its outer Environment implementation provides -and qualifies the requested behavior. Do not advertise daemon-level network -isolation or silently run a restricted request with unrestricted semantics. -The normal self-hosted combination uses the host's existing network access. +The daemon's `clirunner` offers opt-in Unix process-group ownership for adapters whose SDK launches a native child; unsupported hosts reject this mode before launch. Explicit and parent-context cancellation share a TERM grace period (three seconds by default) and a bounded KILL escalation. An internal reaper also cleans remaining group members when the direct process exits, even if a descendant still holds stdout open; during cancellation, surviving descendants keep the remaining grace after the leader exits. The daemon's `stop` command waits up to ten seconds for confirmed shutdown, which covers that grace period and the pipe and owner cleanup after it. -## Native references +Owned output pipes stay readable after the leader exits. Consumers drain stdout and stderr before calling `Wait`, which joins the cached process result and closes the readers. `Done` reports leader reaping and group cleanup signals; it is not a native execution receipt or proof of persisted history. SDK adapters settle each Turn and drain its observations before publishing completion, and Executor close also closes the query and awaits the native child. Process groups are lifecycle supervision, not isolation or containment of descendants that leave the group. -Use these adapters as implementation references after choosing a native API: +Adapters run native tools unattended with the launching user's permissions: Codex with approval policy `never` and full access, Claude through the adapter's tool callback in native `default` permission mode with the SDK sandbox disabled, and MiniMax with bypassed permissions and its sandbox disabled. Do not add permission profiles, bubblewrap wrappers or native sandbox settings; there is one execution path for every Environment origin. Resource paths are operator configuration, not a permission boundary. + +Network admission follows [Restricted network](environments.md#restricted-network). + +## Native references | Harness | Adapter | Native transport | Runtime guide | | --- | --- | --- | --- | diff --git a/contracts/agents-api/harness-selection.md b/contracts/agents-api/harness-selection.md deleted file mode 100644 index 395d77a57..000000000 --- a/contracts/agents-api/harness-selection.md +++ /dev/null @@ -1,61 +0,0 @@ -# Core harness selection extension - -The pinned official Agent contract has no harness selector. Core adds one optional -`x_agents_core` field to saved Agent create/update/read and Session inline/effective -Agent configuration. This is a Core extension, not an upstream field. - -```json -{"x_agents_core":{"harness":"claude_sdk"}} -``` - -Supported identifiers come from the [generated Harness catalog](harness-catalog.md). -Unknown identifiers, empty objects and unknown nested fields are -rejected. Omission inherits a saved Agent value, or uses `OAC_DEFAULT_HARNESS` for -an inline Agent. An explicit null clears the saved selection or replaces it for -one Session, restoring deployment-default selection. Agent updates preserve omitted -fields and replace the entire supplied extension. Saved resources do not start -execution and can retain protocol configuration beyond the selected engine's -current execution profile. - -Session creation resolves the extension after saved-Agent overrides, validates the -selected execution profile, and persists the resulting existing `Session.Engine`. -An explicitly selected harness must be enabled by the deployment; it never falls -back to a different engine. When the effective Agent includes the extension, Session -reads report its persisted engine. Sessions without the extension retain the official -Agent response shape, including historical Sessions. Reads never consult current -Agent defaults or the current deployment default. Explicit-selector creation retries -retain caller intent before mutable configuration resolution. Changing a selector -under an existing creation key conflicts. - -The environment remains the separate official Session `environment` parameter. -Environment templates select startup configuration, not engines, containers or -providers. Multiple Sessions using the same Agent/template have separate Environment -resources and native histories. Agent edits do not change accepted Sessions. - -## Operator configuration - -`OAC_DEFAULT_HARNESS` selects the default engine. `OAC_HARNESSES` explicitly -adds comma-separated deployment-supported engines, for example -`codex,claude_sdk,mcode`, without requiring a managed Provider. The default engine -remains enabled; unknown names fail startup. -This setting does not install a harness or qualify a native deployment. - -Harness selection is independent of Environment provisioning. Provider selection, -node generations and reset behavior are owned by -[Sandbox deployment](sandbox-deployment.md); node wire compatibility is owned by -[the node protocol](node-generation-protocol.md). Enabling a Harness does not -install or qualify it on an existing Runtime. - -Model/provider defaults, source precedence and credential handling are owned by -[Model execution](model-execution.md#saved-defaults-and-precedence). They do not -introduce another Harness selector. - -Current profile limits remain in the [engine coverage table](README.md#public-engine-profiles). - -## Implementation rules - -Core accepts the optional `agent.x_agents_core.harness` extension through the -saved Agent and inline Session configuration paths. Define the extension once in -`contracts/agents-api/v1`; never use metadata or a competing top-level selector. -Use the catalog-derived validators for identifiers and the selection semantics -above. Do not maintain another accepted-name list in handlers or schemas. diff --git a/contracts/agents-api/harnesses.md b/contracts/agents-api/harnesses.md deleted file mode 100644 index 0ebc47cb7..000000000 --- a/contracts/agents-api/harnesses.md +++ /dev/null @@ -1,164 +0,0 @@ -# Native harness contract and qualification - -Codex, Claude Code and future harnesses are equal execution engines. Core owns -public protocol, authority and durable state. Each adapter owns native -configuration, transport and process translation; -the native harness owns the model/tool loop. Supporting this contract -means implementing its observable semantics. Harnesses do not need identical -feature sets. Optional native limitations are separate capability work and do not -block completion of otherwise qualified onboarding. - -For implementation steps and interface obligations, see -[Add a native Harness](harness-onboarding.md). - -## Integration surface - -Follow the numbered steps in [Add a native Harness](harness-onboarding.md#steps) -to implement, register and select an adapter. This page owns what counts as -qualified. - -The service profile catalog is the explicit qualification boundary; a Runtime -heartbeat cannot authorize new public functionality. Unknown profiles fail -closed. Schema validity, qualified service support and the available Runtime -remain independent checks. Additional capability combinations require evidence, -not an engine-name exception. No new handler, store table, scheduler, event -projector or model loop is needed for capabilities the contract already represents. - -Do not require MCP, functions, images, verbosity control or another engine's -optional features simply to register a Harness. The default engine is a -deployment convenience, not a different contract or authority level; see -[engine selection](harness-onboarding.md#engine-selection). - -## Shared behavioral obligations - -- Preparation holds resources without consuming input. Normal Turn completion - retains a settled healthy Executor. Cancellation targets one Turn; idle expiry, - environment shutdown and invalidation close the Executor. Failed cleanup retains - ownership and capacity, and never authorizes replay of uncertain input. -- Confirm accepted/applied inputs separately. Preserve ordered public Items and - events, stable call identity and one terminal outcome. Never replay uncertain - work merely because a connection closed. -- Resume only the bound Session's native history. Missing, ambiguous or foreign - history fails closed. Device identity is not native Session ownership. -- Native tools and public Files use the same bound workspace. Public Files retains - tenant and path authorization. Native tools run with the starting account's - permissions; daemon/model credentials are not isolated from that same user. - Managed outer Environments must exclude other tenants' resources. -- Emit verified measurements; absence of native usage detail is not a zero value. - Explicit unsupported operations remain implementation gaps in protocol coverage. - -## Current qualified operations - -The baseline is actual supported behavior on main, not everything Codex accepts -syntactically or everything an upstream Harness can theoretically perform. The -MiniMax Code column summarizes the `mcode` row of the -[public engine profiles](README.md#public-engine-profiles) and the linked operation -contracts. - -| Operation | Codex | Claude Code | MiniMax Code | -| --- | --- | --- | --- | -| Docker hosted text execution, native local tools | Qualified | Qualified | Qualified (Docker `openai_hosted`) | -| Files, immutable Artifacts, cancellation, restart/history recovery | Qualified | Qualified | Qualified on the dedicated Docker profile; see [MiniMax Code Runtime](../../services/core/deploy/mcode/README.md) | -| Public functions in `none` | Qualified | Qualified; object-root schemas; text or successful inline PNG/JPEG results | Unsupported | -| Public functions alongside hosted workspace tools | Qualified | Qualified; object-root schemas and text or successful inline PNG/JPEG results | Unsupported | -| HTTP MCP and static-bearer Vault credentials in `none` | Qualified | Qualified subset | Unsupported | -| Required MCP initialization | Qualified | Qualified on `none`; native readiness before initial input | Unsupported | -| Hosted HTTP MCP | Gap | Gap | Gap | -| Function image results | Supported subset; early acknowledgement is transport-only | Successful inline PNG/JPEG on `none` and Docker `openai_hosted`; native resizing allowed, error images rejected | Unsupported | -| Non-default verbosity | Native/model-dependent support | No equivalent qualified; medium only | Medium only | -| Public detailed Usage | Supported native counters | Native raw usage retained; public breakdown gap | Public breakdown unsupported | -| V1 `self_hosted` daemon enrollment at `/workspace` | [Qualified deployment scope](user-managed-runtime-v1.md) | [Qualified deployment scope](user-managed-runtime-v1.md) | [Qualified deployment scope](user-managed-runtime-v1.md) | -| Deferred function discovery | Unqualified; explicit rejection | [Single-agent text/function profile](tool-search.md) | Gap; see [tool search](tool-search.md) | -| Structured output | Unqualified; explicit rejection | [Qualified single-agent function profile](structured-output.md) | Gap; see [structured output](structured-output.md) | -| Message images | Inline PNG/JPEG on `none`, Docker `openai_hosted` and `self_hosted` | Inline PNG/JPEG on `none`, Docker `openai_hosted` and `self_hosted` | Unsupported; see [message input](message-input.md) | -| Explicit reasoning | Shared service gap | Shared service gap | Shared service gap | -| Six Subagent reads | [Qualified scope](subagents.md) | [Qualified scope](subagents.md) | [Qualified scope](subagents.md) | - -This inventory records supported combinations, not a feature-equality checklist. -Do not silently drop options, fabricate measurements, weaken isolation or remove -working features. Unsupported operations stay explicit; implementing them is a -separately prioritized decision, not an onboarding prerequisite. MiniMax also implements the same colocated Runtime enrollment. This V1 decision -uses our daemon as executor and explicitly does not claim stock `exec-server` -interoperability. The old service-side harness/remote executor route is retired. -Service-origin HTTP MCP remains unsupported on `self_hosted`; `none` MCP and -qualified hosted Template Plugin MCP retain their separate scopes. - -## Acceptance checklist - -Before implementation, record the operation set, expected results, exclusions and -stopping conditions. A batch ends when its declared operations pass; it does not -expand to match another Harness's feature list. A small adapter does not remove -the need for native qualification. - -- Adapter tests: reuse `agent/contracttest.TextLifecycle` with a controlled native - fixture or real provider. Record which one was used. Two ordinary Turns share - one native process/connection and history; - cancellation followed by another Turn; stale cancellation and late events; native - exit, cleanup failure, input write/application receipts and unknown outcomes. - Verify fresh per-Turn usage, function, input and child-observation state. -- Shared integration: public admission, actual Worker/device selection, gateway, - daemon registration/dispatch and durable terminal projection. See - `TestThirdHarnessPublicOnboarding` for a synthetic example, not native evidence. -- Real acceptance: pinned official Python SDK and raw HTTP against our Agents API, - real provider API, native harness and independent execution database. Verify - initial execution, warm follow-up, cancellation and restart/continuation. - Record native owner identity and same-condition cold/warm timing; mock results - cannot establish native reuse or performance gains. - For hosted qualification also verify Files/Artifacts, workspace identity, - credential protection and foreign-history rejection. -- Regression: existing qualified engines keep working. Run targeted tests during - development, then `make check` and applicable real regressions. API changes - require `make openapi`; query changes require `make sqlc-generate`. -- Review: follow the repository - [blind review workflow](../../CONTRIBUTING.md#review). - -Record exact revisions, image/package versions, commands, results and limits. -Keep keys in private operator files; never commit them or include them in logs. -Failed or synthetic runs cannot be counted as real acceptance. Acceptance of one -engine is not complete public protocol compatibility. - -## Common contract acceptance - -The synthetic [third-harness fixture](../../apps/daemon/testdata/onboarding/main.go) -implements only the current text execution contract: cancellation, durable active -input receipts and strict bound-history continuation. It has no workspace, MCP, -public functions, permissions or user-choice handlers. Its registration is local -to the fixture; production builds never register it. - -`TestThirdHarnessPublicOnboarding` uses a custom immutable `engine.Catalog` in the -same `execution.Policy` supplied to both the API handler and Dispatcher. The zero -policy selects built-ins; an explicitly empty catalog authorizes no engines. -The test runs public Session/input admission, Worker device selection, the real -WebSocket gateway, daemon Registry/Router, neutral events and durable terminal -projection. It checks applied input receipts, saved native identity, continuation, -cancellation, unsupported optional requests and missing mandatory Runtime support. -This proves the integration path, not real native execution or sandbox security. -The existing internal registration functions suffice for this fixture; no dynamic -registry or global mutable test registration is required. - -The current text execution contract still requires durable input/Turn semantics, -ordered observations and enforcement of disabled execution controls. An adapter -without tools or subagents can guarantee their absence; it must not pretend to -apply unsupported requested behavior. Hosted qualification has additional workspace -and isolation obligations. These guarantees are independent of feature equality. - -## Native operation acceptance - -Use the pinned official Python SDK, raw HTTP and real model APIs. The common -`services/core/tests/official_hosted_functions_native.py` assertions exercise -function success/error, native file output and public artifact bytes, same-history -continuation after restart, foreign result rejection and pending-call cancellation. -The operator fixture supplies only deployment/restart and model configuration; -public assertions are shared by adapters that support this operation. Existing workflow, file, artifact -and interruption fixtures remain applicable. Native isolation canaries supplement -these tests; synthetic responses alone do not establish live qualification. - -The 2026-09-19 candidate passed the same hosted-function assertions with Codex -0.153.4 and Claude SDK 0.3.269/native 2.1.269 using real Kimi K3. Each run used -an independent Core, dedicated Agents API database and Docker Runtime. Cold -restart acceptance restores Core before restarting Runtime; daemon startup while -Core is unavailable is not qualified by this test. Evidence is retained under -`~/.parsar/remediation/20260919/harness-parity/` on the validation server. - -The broader protocol inventory remains in [README.md](README.md). Passing one -profile or these shared assertions does not establish complete compatibility. diff --git a/contracts/agents-api/history-events-usage.md b/contracts/agents-api/history-events-usage.md index 774c02299..79c36d6c5 100644 --- a/contracts/agents-api/history-events-usage.md +++ b/contracts/agents-api/history-events-usage.md @@ -30,7 +30,7 @@ Use response cursors to page history in the requested direction. Session Turn lists contain root Turns only; a child Turn ID on the Session Turn routes is not found. Root Items and child Items have separate query resources; use the Subagent resources for child Turns and history. See -[Subagent visibility](subagents.md#subagent-visibility--september-23-2026). +[Subagent visibility](subagents.md#subagent-visibility). Tenant ownership is enforced by Core for both queries and streams. ## Measurement boundary @@ -182,7 +182,7 @@ stream differences EVT-01..04; the plan is `cancelled`) now carry `usage` copied from the rendered Turn snapshot, with explicit null when unknown; other events omit it. This batch applied it to root and child Turn events; the - [Subagent visibility batch](subagents.md#subagent-visibility--september-23-2026) + [Subagent visibility batch](subagents.md#subagent-visibility) later stopped publishing child Turn events on the Session stream. Codex can therefore publish measured counters at settlement, while Claude and MiniMax stay null; no counter is derived or summed. The TypeScript client accepts the diff --git a/contracts/agents-api/list-query-semantics.md b/contracts/agents-api/list-query-semantics.md index 68a49dfaf..69363ac85 100644 --- a/contracts/agents-api/list-query-semantics.md +++ b/contracts/agents-api/list-query-semantics.md @@ -117,7 +117,7 @@ deferred below. | B1 | Repeated supported key on Beta lists, including scalar `status` | 400, code `invalid_request_error`, param null, ``Failed to deserialize query string: duplicate field `` `` | VA-03: `req_83ee26a0b9ba4e84af8996a9346f261b`, `req_de6a02d9a9c547e490e6e53aeb45544d`, `req_92b924b186cf48488cf625638130bb16`; SES-16: `req_86ab2f32451a426399e76b21debeadba`; SFT-24: `req_794a86f4e72946dfb688a2e6f23fff32` | | B2 | Repeated supported key on Skills and Skill versions | 400, code `duplicate_parameter`, param ``, with the observed message | SFT-10: `req_36e64628c2a44b1198d00949c4ed6cc8` | | B3 | Repeated key on Files | Unchanged local `unsupported_parameter` rejection | SFT-18: `req_9a38e0628b4e40bd8e7202ac0880209d` accepted one identical repeated `purpose`; a single sample does not define which differing value wins | -| C1, C2 | `limit=0` and `limit>100` on Agents, Sessions, Items, Templates; Subagent Items and Subagent Turn Items since the [Subagent visibility batch](subagents.md#subagent-visibility--september-23-2026) | Clamped to 1 and 100 | VA-01: `req_4954e90768b54c588167333e83f0a958`; VA-18: `req_d43d918872cb40b5b6e545582ea94205`, `req_2efee54ffd6f41179e294870b0e62e02`; SES-10: `req_4f1fb6cb479a4d1cb64486c72d250e9f`; SES-11: `req_5af4e60768b14c9eab284bce8638348a`; SES-12: `req_2f7cab87400348408ab8bb6e68789dc9`; SES-13: `req_b5ebab5691314bfba6ad59c84c42d04e`; SFT-23: `req_f63fef0e08c1462c8158c0ee8523a88d`; SAT-02: `req_6179ae6c1d1640d899ee4798e7f9fa57`, `req_7436104afbae4e73a0eb43b00ec9e660`, `req_32899313414b4031849a22cd2927f0ad` | +| C1, C2 | `limit=0` and `limit>100` on Agents, Sessions, Items, Templates; Subagent Items and Subagent Turn Items since the [Subagent visibility batch](subagents.md#subagent-visibility) | Clamped to 1 and 100 | VA-01: `req_4954e90768b54c588167333e83f0a958`; VA-18: `req_d43d918872cb40b5b6e545582ea94205`, `req_2efee54ffd6f41179e294870b0e62e02`; SES-10: `req_4f1fb6cb479a4d1cb64486c72d250e9f`; SES-11: `req_5af4e60768b14c9eab284bce8638348a`; SES-12: `req_2f7cab87400348408ab8bb6e68789dc9`; SES-13: `req_b5ebab5691314bfba6ad59c84c42d04e`; SFT-23: `req_f63fef0e08c1462c8158c0ee8523a88d`; SAT-02: `req_6179ae6c1d1640d899ee4798e7f9fa57`, `req_7436104afbae4e73a0eb43b00ec9e660`, `req_32899313414b4031849a22cd2927f0ad` | | C3 | `limit` 0 or above 100 on Turns, Subagents and Subagent Turns | 400, code `invalid_request_error`, param null, `limit must be between 1 and 100` | SES-14: `req_2793f57b6a454c399a031277b6a02e45`; SAT-03 (campaign scan 3): `req_7df58d9579be4ee3ab7fdab55286aa05`, `req_b4321de4480c4a8e96b9ea285ff63a46`, `req_0f437ad4713d47a8af1f61a88636bf79`, `req_e9d476dd2a69472694cffc0851d0574c` | | C4 | Negative or non-integer `limit` on Beta lists other than Vaults and Credentials | 400, code `invalid_request_error`, param null, `Failed to deserialize query string: limit: invalid digit found in string` | VA-04: `req_1362046e9d69497da9c23ca69517a026`, `req_c384455192dc4a99a032299a91416a74`; SES-15: `req_918128738f3a47b69203ca091f9cb9ca` | | C5 | Vault and Credential `limit` | 0, negative and above-100 values keep the pinned clamp; a non-integer uses the C4 error | VA-04: `req_f08cda4e1b9048808be9465b9a56c4a4`, `req_e53033a2e01e4b87aa0bb8d32c932a44`; VA-06 (pin conflict): `req_2ca22663c9414afa920b5509d3574812`, `req_1d088639a3614f6ba545cd36497ff35c`; VA-18: `req_bc47269acf8143f7869908567272e433`, `req_62eaf8fd15e8483a98a9a09ebc4d5e51` | diff --git a/contracts/agents-api/message-input.md b/contracts/agents-api/message-input.md index fd34a9f95..710d417b4 100644 --- a/contracts/agents-api/message-input.md +++ b/contracts/agents-api/message-input.md @@ -180,6 +180,6 @@ Function-result image support has its own [coverage record](function-result-imag compatibility or support for arbitrary vision-model/provider combinations is claimed. User-managed Linux image execution follows the same native workspace path. See -[current qualification](environment-capabilities-qualification.md) for the tested +[current qualification](harness-capabilities.md) for the tested Harnesses, formats, continuation and remaining boundaries. Historical evidence above retains its original deployment scope. diff --git a/contracts/agents-api/model-configuration-design.md b/contracts/agents-api/model-configuration-design.md deleted file mode 100644 index 87a441632..000000000 --- a/contracts/agents-api/model-configuration-design.md +++ /dev/null @@ -1,288 +0,0 @@ -# Unified model configuration protocol design - -Status: planned protocol. This document defines the target design, not newly -accepted HTTP input or current adapter support. [Model execution](model-execution.md) -owns the implemented API; [Harness onboarding](harness-onboarding.md#native-model-configuration) -lists current native parameters. This planned contract does not add runtime -behavior, public fields or adapter support. Existing fields below are reused owners, -not a claim that every value already executes on every Harness. - -This is the canonical owner of the planned public model field vocabulary, -resolution and application semantics. Future adapter support extends the existing shared contract in -[internal/harnessconfig/harness.go](../../internal/harnessconfig/harness.go). -Implementation must update the current API documentation and generated schemas -together before advertising new fields. - -## 1. Shared path - -Web deployment defaults and public Agent/Session configuration use the same value types, validation, resolution and planning functions. Their differences are authority, source and when the complete Agent/input is available. - - Deployment model defaults ─┐ - Saved Agent configuration ─┼─> resolve one immutable model configuration - Inline Session Agent ──────┘ | - validate requested behavior - | - compile native Harness plan - | - native preparation and inference - -A default is a reusable model/generation fragment. It contains no instructions, messages, tool definitions, MCP declarations, Skills, workspace settings or lifecycle policy. Saving defaults checks the shared fragment; Session admission supplies Agent tools and other context and completes the same validation. Unresolved contextual obligations are never permission to execute. - -There is no capability-certification database, model catalog requirement or second execution loop. - -## 2. One public owner per field - -Reuse the existing public Agents API fields instead of placing another copy inside generation. Public wire envelopes need not look identical across different resources; their model/generation fragments use the same types and resolver. - -| Semantic field | Public Agent / Session request owner | Deployment default owner | -| --- | --- | --- | -| Exact model ID | Saved Agent model; Session agent.model | model | -| Upstream connection | Saved Agent x_agents_core.model_provider; Session x_agents_core.model_provider override | model_provider | -| Model capability declarations and limits | Saved/inline Agent x_agents_core.model_capabilities | model_capabilities | -| Reasoning effort and summary | Agent reasoning.effort and reasoning.summary | reasoning, using the same type | -| Text format/schema and verbosity | Agent text.format and text.verbosity | text, using the same type | -| Service tier | Agent service_tier | service_tier | -| New generation options absent from the public contract | Saved/inline Agent x_agents_core.generation | generation | -| Tool definitions, native search declaration and search options | Existing Agent tools and its typed variants | Not copied into defaults | -| Actual message/file/image content | Existing input / MessageInput and tool-result content | Not copied into defaults | -| Optional vendor-only model options | Saved/inline Agent x_agents_core.harness_config | harness_config | - -There is no additional Session-level generation, model_capabilities or harness_config override. A Session overrides its saved Agent through agent.x_agents_core. The connection override remains in its current confidential Session envelope. This removes the current three-location native-parameter precedence rather than retaining aliases. - -Provider becomes connection-only: protocol, base_url and write-only api_key. Move its current context_window and max_output_tokens into model_capabilities.limits.context_window and model_capabilities.limits.max_output_tokens. Remove the old provider locations in the same change; do not accept both. - -The existing endpoint /core/v1/harnesses/{harness}/model-configuration remains the deployment-default resource. Its response includes safe configuration and a support description, never credentials. - -## 3. Common vocabulary - -The public vocabulary is not reduced to the switches available in today's three adapters. A recognized option can still be explicitly unsupported by a particular execution plan. - -| Group | Fields / concepts | Meaning | -| --- | --- | --- | -| Model limits | context_window, max_output_tokens | Declared model ceilings, not the requested generation budget | -| Modalities | input: text, image, audio, video, file; output: text, image, audio, video, file | Capability vocabulary; modalities need format/placement constraints, not just one boolean | -| Generation limits | generation.max_output_tokens, generation.stop | Per model-inference-call constraints; not a whole Agent Turn or Session budget | -| Sampling | generation.sampling.temperature, top_p, top_k, min_p, presence_penalty, frequency_penalty, repetition_penalty, seed | Distinct controls; no renaming one penalty into another | -| Reasoning | Existing effort and summary; generation.reasoning.mode and budget_tokens where the official fields do not express them | Relative effort, explicit budget and reasoning visibility are different controls | -| Structured output | Existing text.format and schema; planned json_object variant | Text and JSON Schema retain their single current owner; JSON-object mode requires an explicit contract extension and qualified output handling | -| Token probabilities | generation.logprobs, generation.top_logprobs | Token probability outputs need a qualified response/event representation | -| Tool selection | generation.tool_choice, generation.parallel_tool_calls | Selection/parallelism over the resolved Agent tool inventory; not new tool definitions | -| Tool capabilities | Function calling, parallel calling, native web search, tool search, image/tool results | Separate execution capabilities with separate requirements | -| Output modalities | generation.output.modalities and modality-specific settings such as audio format/voice | Typed extensions; message/event representation must also support the requested output | -| Delivery | Streaming, usage accounting, cancellation, continuation | Existing execution protocol capabilities; not duplicated model-generation settings | - -Keep the existing public reasoning effort values: none, minimal, low, medium, high, xhigh, max. Support declarations can expose a subset. Effort is relative to the selected model; it is not a promise of equal latency, price or exact reasoning-token count across models. - -Proposed reasoning.mode values are disabled, adaptive and budgeted. These describe different intentions. A budgeted request requires a positive budget_tokens value. Explicit mode disabled conflicts with an enabled effort or reasoning budget. Explicit effort none conflicts with adaptive/budgeted mode. Simultaneous relative effort and a token budget require a specifically supported combination; never assume one wins. - -Numeric values must be finite and validated using typed ranges. Temperature/top-k/penalties have no universal vendor-compatible range. top_k is an integer; counts and token budgets are positive integers; seed is an integer and does not promise deterministic results. Support declarations provide applicable ranges and combinations. They must not clamp values or substitute the nearest enum. - -max_output_tokens is a ceiling on tokens generated by one inference call, including reasoning tokens where the upstream counts them in completion output. If an endpoint can cap only visible text while hidden generation remains uncapped, it does not satisfy that total-output ceiling. An exact native mapping or an explicit rejection is required. It is not an instruction to truncate the returned answer. - -A declared context window informs Harness context management; it does not authorize Core to delete/truncate conversation history. Unknown limits are not zero. Request budgets above a known model ceiling reject rather than being silently reduced. - -Tools, images and audio/video are end-to-end behaviors. Merely forwarding or injecting a request member does not implement native tool handling, input consumption, output events or recovery. - -### Field types and constraints - -All entries below describe the target contract. Existing public owners retain their -current parsing/update rules until an explicitly qualified implementation changes -them. A recognized field is not automatically executable. Planned fields must not -appear as accepted HTTP input or advertised support before their implementation. - -Objects reject unknown/duplicate members. Numbers must be finite. Optional values -may be absent; absence is not zero. Counts use positive integers unless the table -explicitly permits zero. Native support descriptors can narrow ranges, never silently -clamp or rename a value. No arbitrary provider-specific key/value bag is added. - -| Field (relative to its owner above) | Type / basic range | Meaning and combinations | -| --- | --- | --- | -| model | Nonempty string | Exact upstream identity; no alias or name-based capability inference | -| model_provider.protocol | Enum: responses, anthropic, chat_completions | Upstream protocol, independently selected from Harness | -| model_provider.base_url / api_key | HTTPS URL / nonempty write-only string | Existing connection validation and secret-storage rules remain in [model execution](model-execution.md); replace as one bundle | -| model_capabilities.limits.context_window | Positive integer | Declared total context ceiling; absence means unknown | -| model_capabilities.limits.max_output_tokens | Positive integer, no greater than declared context_window | Model output ceiling, distinct from the requested per-call budget | -| model_capabilities.input_modalities / output_modalities | Optional unique arrays of text, image, audio, video, file | Absent list means unknown; supplied list is exhaustive, so an omitted modality is unsupported. These are declarations, not actual content | -| model_capabilities.features | Typed object of feature status enums | Keys: reasoning, function_calling, parallel_tool_calls, structured_output, web_search, tool_search, logprobs. Each is supported, unsupported or unknown; absent means unknown | -| reasoning.effort | Existing enum: none, minimal, low, medium, high, xhigh, max | Relative effort; conflicts with incompatible explicit mode/budget; supported subset depends on the plan | -| reasoning.summary | Existing enum: auto, concise, detailed | Requested summary visibility/detail; does not authorize revealing private reasoning | -| text.format | Discriminated object, type text or json_schema; json_object is planned | json_schema requires its schema object. text/json_object reject schema. json_object requests valid JSON without a schema; output representation and native behavior must be qualified | -| text.verbosity | Existing enum: low, medium, high | Output detail preference, not a token ceiling | -| service_tier | Existing enum: auto, default, flex, priority, fast | Preserve public vocabulary; qualification may accept a subset | -| generation.max_output_tokens | Positive integer, at most a known model output ceiling | Per-inference generation ceiling, including counted reasoning tokens | -| generation.stop | Array of distinct nonempty strings; [] means no configured stop sequences | Stop-generation sequences; provider-specific count/length limits come from native protocol support | -| generation.sampling.temperature | Finite number >= 0 | Randomness control; zero is explicit. Require a qualified combination if other samplers are also supplied | -| generation.sampling.top_p | Number in (0, 1] | Cumulative-probability sampling threshold | -| generation.sampling.top_k | Integer >= 1 | Candidate-count threshold; zero is invalid rather than a vendor-specific disable alias | -| generation.sampling.min_p | Number in [0, 1] | Minimum candidate probability relative to the highest-probability token; zero is explicit | -| generation.sampling.presence_penalty / frequency_penalty | Finite numbers; qualified native adapter supplies min/max | Presence-based / frequency-based penalties; independent meanings, not aliases | -| generation.sampling.repetition_penalty | Finite number > 0 | Multiplicative repetition penalty; do not map to additive presence/frequency penalties | -| generation.sampling.seed | Integer in [-(2^53-1), 2^53-1] | Random seed, limited to exact client JSON integers; not a determinism guarantee | -| generation.reasoning.mode | Enum: disabled, adaptive, budgeted | budgeted requires budget_tokens; other modes reject budget_tokens. disabled conflicts with enabled effort | -| generation.reasoning.budget_tokens | Positive integer | Reasoning budget within the total per-call budget when supplied; cannot exceed a known model output ceiling. Relative effort + budget requires a qualified combination | -| generation.tool_choice | Tagged object: type auto, none, required or function; function requires name | Function name must identify a callable tool in the resolved inventory; other types reject name. required/function need tools; none remains meaningful without tools | -| generation.parallel_tool_calls | Boolean | Explicit false constrains tool-call parallelism; true needs end-to-end consumption support | -| generation.logprobs | Boolean | Request token log probabilities; requires qualified response/event representation, not just request forwarding | -| generation.top_logprobs | Integer >= 0 | Number of alternative token probabilities; requires explicit logprobs: true; upper bound comes from native adapter support | -| generation.output.modalities | Nonempty unique array of text, image, audio, video, file | Requested output modalities; each needs native consumption, public output/events and native qualification | -| generation.output.audio | Typed object with nonempty format and optional nonempty voice strings | Requires audio in output.modalities; formats/voices are exact provider values restricted by qualified mappings, not interchangeable labels | -| harness_config | Optional bounded JSON object with adapter-owned vendor-only keys | Current native schema remains implemented until explicitly replaced. A future common-field mapping removes its native duplicate atomically | - -Images/files in messages and tool results retain the existing MessageInput and -placement types. Their supported forms, sizes and event contracts have existing -owners. No image payload, tool definition, instruction or message is moved into -model_capabilities or generation. Constraints involving format, placement or -feature combinations use typed support descriptions and pure validators; the -feature statuses alone do not establish those constraints. - -## 4. Capability facts and unknowns - -Keep three distinguishable facts: - -1. Model declaration: supplied by the operator/application or trustworthy provider metadata. Missing information is unknown; a declaration is not execution evidence. -2. Harness expression: the adapter and installed native version can implement the requested behavior through native configuration, SDK or Turn settings. -3. Native protocol support: the selected protocol belongs to the adapter's shared ordered `protocols` list, and the native path preserves the requested behavior. - -Use supported / unsupported / unknown, with field-specific ranges, accepted values and combinations. A protocol name or model-name pattern is not proof of behavior. - -Known model unsupported, or unqualified local Harness behavior, rejects before native submission. Unknown model support is allowed when the local native plan is qualified; the upstream then confirms acceptance or returns its normal error. Return the model status as unknown, not supported. Never force users to register every model in a catalog. - -Do not introduce a rules DSL. Shared typed descriptors describe ordinary ranges and supported choices; adapter-owned pure functions validate combinations. The same declarations serve Core admission, Runtime and Web. The implemented protocol matrix and descriptor contract remain in [model execution](model-execution.md); this design does not add another protocol list. - -## 5. Native application of common settings - -Each supported explicit generation field is applied once by its Harness adapter through native configuration, SDK options or Turn settings. Unified public vocabulary expresses intent; the native adapter maps only fields whose semantics it can preserve. - -A missing native CLI switch is not by itself decisive if a supported native SDK or Turn setting implements the field. If the native interface cannot express the requested semantics, reject the field explicitly as unsupported. Do not inject fields through a proxy, rewrite outgoing model requests, introduce a passthrough gateway or add cross-protocol conversion inside the Harness. - -Compile native application before execution. Do not silently drop unsupported members, retry with weaker controls or add prompt instructions to imitate a setting. Native defaults must not replace explicit frozen values. - -Fields affecting tool handling, context accounting, reasoning state/signatures or output consumption require adapter-specific qualification across the native lifecycle. A common field definition alone does not establish that support. - -## 6. Required Harness contract - -The existing shared contract in -[internal/harnessconfig/harness.go](../../internal/harnessconfig/harness.go) -remains the only model configuration contract. It already declares native -protocols and parameter validation, prepares the current model/provider/native -inputs, and is mandatory at Runtime registration. Core consumes that same -declaration. The separate Executor/Turn lifecycle remains unchanged. - -Implement future common fields by extending this contract's resolved input, -support description and pure preparation together. Do not introduce a second -adapter interface, registration hook or parallel configuration planner. Public -wire types retain their existing owners; this design does not add those fields -to current HTTP input or require placeholder types now. - -Mappings are ordinary tested pure functions within native adapters. Existing wire and adapter versions identify compatibility. Do not add mapping IDs, a mapping registry or a version/rules DSL. - -When common fields are implemented, shared preparation combines native adapter support with model declarations. No Harness-name branches in Core, source-specific adapter flows or duplicate validators. - -The compiled plan must account for every resolved explicit non-clear field and every requested behavioral requirement. Missing or duplicate application is an error. Identity, model limits and model declarations retain their own typed handling. - -Core freezes the resolved configuration, its sources and semantic/wire version in the existing Session record. It does not store an executable local plan. Runtime recompiles using the installed native adapter and rejects a plan it cannot preserve. - -Preparation starts native resources only after local planning succeeds. The factory receives the frozen direct provider connection and compiled native settings. Preparation failure, cleanup ownership and uncertain submission retain the existing contracts. - -Native settings must hold for every inference: first Turn, subsequent Turns, retries owned by the native Harness, steering, tool continuations and cold recovery. Qualify this behavior through native configuration and lifecycle evidence, without introducing request interception. New content requirements are rechecked before sending messages or tool results; a previously text-only Session does not authorize a later unsupported image. - -The scope is the configured Agent's inference calls. Existing subagent configuration/override rules remain authoritative; do not silently apply the parent's model-specific limits or native settings to a child with a different model binding. - -Errors identify safe field paths, reason and blocking layer. They never include keys, submitted secret values or raw upstream bodies. - -## 7. Presence, inheritance, replacement and conflicts - -Use presence-aware typed inputs. Do not let JSON omitempty or truthiness turn zero/false/disabled into absence. - -For the new generation extension, the object is the unit of replacement: - -| Input | Meaning | -| --- | --- | -| generation omitted | Inherit the entire object from the next applicable source | -| generation: null or {} | Clear the entire inherited generation object | -| Nonempty generation object | Replace the entire object; omitted fields and nested groups do not inherit | -| Explicit 0 or false inside that object | A requested value; validate it, never treat it as omission | -| Nested null | Invalid; clear by replacing the object without that field | -| A supplied list | Use that list as supplied; an empty stop list requests no stop sequences | - -There is no field-level or nested-group inheritance. A replacement such as -{"sampling":{"temperature":0}} discards an inherited top_p and output budget. -A cleared option has no applied-value obligation, while explicit false/disabled -does. Safe reads expose the resolved values and their source. - -Keep the existing documented public Agent update/override/null semantics for reasoning, text, tools and service_tier. The input envelope translates those operations into the same resolved representation; the new extension does not retroactively reinterpret official fields. - -Provider replacement is atomic: protocol, address and key come from one source. Never combine a new endpoint with an inherited credential. API-key rotation alone is not a change in model capabilities. - -Bind model capability declarations and native escape-hatch options to the selected model, provider endpoint/protocol and Harness as applicable. Changing that binding clears inherited declarations/native options unless replacements are explicitly supplied. Key-only rotation preserves the binding. - -Portability rule: unified generation values explicitly stored in the Agent represent portable application intent. Preserve and revalidate them on a model change rather than silently dropping them. Do not inherit model-specific tuning from a deployment default bound to a different explicitly selected model. Existing official reasoning/text/tools inheritance remains unchanged. This deliberately differs from retaining opaque native settings belonging to the former model. - -Native escape-hatch fields cannot duplicate any canonical common field, even with an equal value. For example, native model_reasoning_effort/effort must move to the common reasoning owner when that mapping lands. A conflict rejects; there is no "native wins" precedence or old alias. - -Validate combinations such as reasoning-disabled plus budget, incompatible samplers, forced tool choice missing from the resolved inventory, tool choice with no tools, parallel calls that the Harness cannot consume, and requested budget above declared ceilings. - -Saved/default edits affect only future Sessions. New versions must not silently reinterpret an existing Session's frozen native keys as common fields. Keep existing Session records/data untouched. If the new Runtime cannot execute an older snapshot contract, return an explicit unsupported-version error; do not translate native keys implicitly or delete data. There is no compatibility alias, automatic migration or fallback parser. - -## 8. Example of the shared fragment - -Deployment-default input: - -~~~json -{ - "model": "provider-model-id", - "model_provider": { - "protocol": "responses", - "base_url": "https://provider.example/v1", - "api_key": "" - }, - "model_capabilities": { - "limits": { "context_window": 128000, "max_output_tokens": 16000 } - }, - "reasoning": { "effort": "low" }, - "generation": { - "max_output_tokens": 4000 - } -} -~~~ - -Equivalent explicit Session configuration: - -~~~json -{ - "agent": { - "model": "provider-model-id", - "reasoning": { "effort": "low" }, - "x_agents_core": { - "model_capabilities": { - "limits": { "context_window": 128000, "max_output_tokens": 16000 } - }, - "generation": { "max_output_tokens": 4000 } - } - }, - "environment": { "type": "openai_hosted" }, - "x_agents_core": { - "model_provider": { - "protocol": "responses", - "base_url": "https://provider.example/v1", - "api_key": "" - } - } -} -~~~ - -These are planned shapes, not currently accepted requests. Both must reach identical model/generation validation and planning after source resolution. - -## 9. Delivery boundaries and acceptance - -1. Establish this field table, value semantics and ownership contract. Extend the existing shared harness.go contract and link one public field owner document. -2. Introduce shared typed fragments and resolver; separate provider connection from model limits. Remove superseded aliases/duplicate inputs together. Update Web and public clients in the same change. -3. The first implementation phase unifies existing reasoning, text, tools and images through the common planner, then qualifies per-call output budgets, temperature and top_p. Keep native adapter ownership and immutable Session behavior. Common fields are the normal input; native JSON remains only a vendor-specific escape hatch. -4. Qualify sampling, per-call output ceilings and other common controls through each adapter's native interfaces. Explicitly reject fields or combinations those interfaces cannot express; do not add a model proxy, transport mapper or protocol conversion to expand support. -5. Add audio/video/file output and other operations only with native input/output, event, cancellation and recovery coverage. Defining the vocabulary now does not claim these operations are implemented. - -Required evidence: equivalent default/public inputs produce the same resolved plan; presence/null/zero semantics; no secrets in safe reads; immutable snapshots and retries; new/reused/resumed native Turns; native application and lifecycle assertions for supported fields; no double application; explicit rejection before sending unsupported inputs; focused real-model acceptance for newly claimed behavior. - -Web should render ordinary unified controls from the shared support description, distinguish unknown model support from unavailable local implementation, and show the reason for an unavailable option. Do not build another handwritten per-Harness parameter schema in the frontend. - -Not in this draft: model catalog/autodiscovery service, capability-certification storage, cost optimization, Session-wide budgets, new tool or Agent loops, cross-platform installers, or replacing a native Harness with a direct-model execution engine. diff --git a/contracts/agents-api/model-execution.md b/contracts/agents-api/model-execution.md index 09394ddc1..669a47ee9 100644 --- a/contracts/agents-api/model-execution.md +++ b/contracts/agents-api/model-execution.md @@ -1,108 +1,65 @@ -# Agent defaults and Session model execution +# Model execution -This page describes the implemented API. The -[unified model configuration design](model-configuration-design.md) owns the planned -common parameter vocabulary and replacement rules; its planned fields are not yet -accepted by these endpoints. +Each Session runs one Harness with one model provider. Core selects them through three Core extensions that the pinned upstream protocol does not define: `x_agents_core.harness` chooses the Harness, `x_agents_core.model_provider` supplies the endpoint and key, and `x_agents_core.harness_config` carries native model parameters. Core has no provider catalog, model alias resolution or product permission model; besides Session and saved-Agent bundles, the only stored bundle is one [deployment default](#deployment-defaults) per Harness. This document is the Harness–model provider protocol: [`internal/modelprovider/config.go`](../../internal/modelprovider/config.go) validates the frozen provider connection, and each Harness declares its protocols and native parameters through [`internal/harnessconfig/harness.go`](../../internal/harnessconfig/harness.go). -Core accepts optional `x_agents_core.model_provider` on saved Agent creation and -update, and the same top-level bundle as an explicit Session creation override. -This is a Core extension, not part of the pinned upstream protocol. It supplies -execution input only: there is no provider catalog, model alias resolution or -product permission model in Core. Besides Session and saved Agent bundles, the only -stored bundle is one [deployment default](#deployment-defaults) per harness. +## Harness selection + +```json +{"x_agents_core": {"harness": "claude_sdk"}} +``` + +Saved Agents accept `x_agents_core.harness` on create, update and read, and Sessions accept it in the inline `agent.x_agents_core`. Identifiers come from the [Harness catalog](harness-catalog.md); unknown identifiers and unknown nested fields are rejected, as is an empty inline Session extension. Which Harnesses a deployment enables, and its default, are the process settings `core.harnesses` and `core.default_harness` ([configuration](../../docs/configuration.md#settings)). + +- Omitted: a Session inherits its saved Agent's Harness; an inline Agent uses the deployment default. +- Explicit null on the Session's inline extension: resets to the deployment default Harness while keeping inherited provider bundles. On a saved Agent, a null extension clears its Harness and provider. +- A selected Harness must be enabled; Core never falls back to another one. +- Session creation resolves the selection after saved-Agent overrides, validates the Harness profile and stores the result as the Session's engine. Session reads report it when the effective Agent includes the extension; other Sessions keep the official Agent shape. Reads never consult the current Agent or deployment default. +- Creation retries with an explicit selector keep the caller's intent; changing the selector under an existing Idempotency-Key conflicts. + +The Session's `environment` and Environment Templates select preparation, not Harnesses or providers. The extension is defined once in `contracts/agents-api/v1`; validators derive from the catalog, and no handler or schema keeps its own list of names. ## Saved defaults and precedence -A saved Agent is editable configuration, not a permanently bound runtime. Create -or update it with `model`, optional `x_agents_core.harness`, and an optional complete -`x_agents_core.model_provider`. Create/update/retrieve/list responses return safe -provider fields and the output-only `api_key_configured` flag. They never return -`api_key`, ciphertext or a reusable credential reference. Normal Agent JSON stores -only the safe view; the secret bundle has a separate encrypted row bound to the -tenant and Agent with a distinct encryption purpose. Agent writes commit safe -configuration and ciphertext together. A model-only edit does not require a key. - -Session creation resolves each explicit model/harness override before saved defaults; -when no harness is selected, the deployment harness applies. Saved Agent creation still requires a model. As a Core extension, an inline -`openai_hosted` or `none` Session may omit its model to use the deployment model -for the resolved harness. `self_hosted` never uses deployment model settings. -There is no model-name inference. Provider precedence is: complete Session bundle, complete -saved bundle, then the deployment default for the resolved harness. Never merge a -replacement endpoint with an inherited key. A model-only override reuses the entire -inherited bundle. Connections use only the selected Harness's native protocols: - -| Harness | Supported protocols, in default order | +A saved Agent is editable configuration, not a bound runtime. Create or update it with `model`, optional `x_agents_core.harness` and an optional complete `x_agents_core.model_provider`. Responses return the safe provider fields and the output-only `api_key_configured` flag, never `api_key`, ciphertext or a reusable credential reference. Agent JSON stores only the safe view; the secret bundle has its own encrypted row, bound to the Project and Agent with a distinct encryption purpose, and is written in the same transaction. A model-only edit needs no key. + +Session creation resolves each explicit model or Harness override before saved defaults; without a selected Harness, the deployment default applies. Saved Agents require a model. An inline `openai_hosted` or `none` Session may omit its model to use the deployment model of the resolved Harness; `self_hosted` never uses deployment model settings. Core never infers a model from its name. + +Provider precedence is: a complete Session bundle, then a complete saved bundle, then the deployment default for the resolved Harness. Core never merges a replacement endpoint with an inherited key; a model-only override reuses the whole inherited bundle. Each Harness connects only through its native protocols: + +| Harness | Supported protocols, default first | | --- | --- | | Codex | `responses` | -| Claude SDK / Claude Code | `anthropic` | +| Claude SDK | `anthropic` | | MiniMax Code | `anthropic`, `responses`, `chat_completions` | -The first protocol is the default. MiniMax Code requires positive context/output -limits. Validate the resolved combination before writing a Session. Core and -Runtime consume the same ordered `protocols` declaration in -`internal/harnessconfig`. There is no built-in model API proxy, passthrough -gateway or automatic cross-protocol conversion, including inside a Harness. -Unsupported saved configurations and Session snapshots fail explicitly when used; -they are never silently rewritten, aliased or migrated. +MiniMax Code requires positive context and output limits. Core validates the resolved combination before writing the Session. Core and Runtime read the same ordered `protocols` declaration in `internal/harnessconfig`. There is no model API proxy, passthrough gateway or cross-protocol conversion, including inside a Harness. Unsupported saved configurations and Session snapshots fail when used; they are never rewritten, aliased or migrated. -Where each source applies depends on who owns the compute that receives the key: +Which sources apply depends on who owns the compute that receives the key: -| Environment | Session or saved Agent bundle | Deployment default | No bundle resolved | +| Environment | Session or saved-Agent bundle | Deployment default | No bundle resolved | | --- | --- | --- | --- | | `openai_hosted` | Accepted | Applied | 400 `model_provider_required` | | `self_hosted` | Accepted | Never applied | 400 `model_provider_required` | | `none` | Rejected with 400 | Applied when configured | Accepted; the device's own environment supplies the model | -The deployment default holds the operator's key, so it stays on operator compute: -Core-managed sandboxes and operator-registered `none` devices. A `self_hosted` -executor belongs to the application; the caller supplies its own bundle. Hosted and -self-hosted Runtimes carry no model configuration of their own, so a Session there -without a bundle is rejected before any write, with `param` -`x_agents_core.model_provider` and a message that says what to configure, instead -of starting a harness that would fall back to a built-in endpoint. +The deployment default holds the operator's key, so it stays on operator compute: Core-managed sandboxes and operator-registered `none` devices. A `self_hosted` executor belongs to the application, which supplies its own bundle. Hosted and self-hosted Runtimes carry no model configuration of their own, so a Session there without a bundle is rejected before any write, with param `x_agents_core.model_provider` and a message saying what to configure. | Operation | Omitted | Explicit null | | --- | --- | --- | -| Agent update `x_agents_core` | Preserve both defaults | Clear harness and provider, including its secret | -| Agent update nested `model_provider` | Preserve the existing bundle | Clear the entire saved bundle | -| Agent update nested `harness` | Preserve existing harness | Reject; use extension null to reset | +| Agent update `x_agents_core` | Keep both defaults | Clear Harness and provider, including its secret | +| Agent update nested `model_provider` | Keep the bundle | Clear the whole saved bundle | +| Agent update nested `harness` | Keep the Harness | Rejected; use a null extension to reset | | Session top-level `x_agents_core` | Inherit provider defaults | Inherit provider defaults | | Session nested `model_provider` | Inherit provider defaults | Inherit provider defaults | -| Session inline `agent.x_agents_core` | Inherit saved harness | Reset to deployment harness | - -An empty Session execution extension remains invalid. An explicitly null provider -is a defined inheritance request; an empty/partial provider object is invalid. -Unknown, duplicate or output-only saved-provider input fields are rejected. Saved -Agent creation without a harness may save a valid bundle, with final harness -compatibility checked at Session admission. Provider-only Agent updates preserve -the saved harness and validate their merged compatibility under the row lock. -Session inline `agent.x_agents_core` accepts `harness` and `harness_config`; -the provider override belongs at the Session request's top level. - -Read the Agent configuration and encrypted bundle from one coherent database -snapshot. Explicit complete Session overrides do not need to decrypt a saved -bundle. Persist a new Session-owned encrypted snapshot atomically with Session and -environment creation. Existing Sessions never consult the Agent again: edits, key -replacement, deletion, suspend/resume and process restarts cannot change their -model, harness or provider. A missing/wrong encryption key fails closed. Retain the -same deployment credential-encryption key across restarts. V1 has no Turn override, -provider catalog, Session migration or new execution loop. - -All new hosted requests and requests that omit the inline model record caller intent before resolving mutable defaults, -including inline requests that use deployment defaults. Other inline requests, -such as `none`, keep the resolved-request retry rule; that hash leaves out a -deployment default, so setting, replacing or removing the default does not change -their retry identity. Matching creation retries -recover the committed Session before resolving the Agent or provider again and do -not enqueue another input. Streaming remains outside the retry identity. Existing -historical rows keep their documented retry limitations; this change does not -rewrite them. Omitted and explicit fields retain the existing local intent-hash -semantics rather than promising upstream equivalence. - -See the [TypeScript client example](../../packages/agents-client/README.md#saved-agent-and-deployment-defaults). - -## Session override example +| Session inline `agent.x_agents_core` | Inherit the saved Harness | Reset to the deployment Harness | + +An empty Session execution extension is invalid. An explicit null provider requests inheritance; an empty or partial provider object is invalid. Unknown, duplicate or output-only saved-provider fields are rejected. A saved Agent without a Harness may save a valid bundle; its Harness compatibility is checked at Session admission. A provider-only Agent update keeps the saved Harness and validates the merged combination under the row lock. The Session's inline `agent.x_agents_core` accepts `harness` and `harness_config`; the provider override belongs at the request's top level. + +Core reads the Agent configuration and encrypted bundle from one database snapshot; an explicit complete Session override needs no decryption of the saved bundle. The Session's own encrypted snapshot is written atomically with the Session and its Environment. Existing Sessions never consult the Agent again: edits, key replacement, deletion, suspension and restarts cannot change their model, Harness or provider. A missing or wrong encryption key fails closed; keep the same [credential key](../../docs/configuration.md#installation-directory) across restarts. There is no Turn-level override. + +New hosted requests, and requests that omit the inline model, record caller intent before resolving mutable defaults. Other inline requests, such as `none`, keep the resolved-request retry rule; that hash leaves out the deployment default, so changing the default does not change their retry identity. A matching creation retry recovers the committed Session before resolving the Agent or provider again and enqueues no further input. Streaming is outside the retry identity. The [TypeScript client](../../packages/agents-client/README.md#saved-agent-and-deployment-defaults) shows saved Agents and deployment defaults. + +## Session override ```json { @@ -120,145 +77,64 @@ See the [TypeScript client example](../../packages/agents-client/README.md#saved } ``` -`protocol` names the upstream API: `anthropic`, `responses` or `chat_completions`. -It does not select an engine. The selected Harness must support that protocol -natively, as listed above; a mismatch is rejected. -The endpoint must use HTTPS without embedded credentials, a query or a fragment. -Keys must be nonempty, at most 16 KiB, and contain no NUL/CR/LF. Unknown fields and -unsupported protocol/Harness/environment combinations are rejected before creating -a Session. The Session's `x_agents_core` accepts `model_provider` and `harness_config`; hosted node -placement is automatic, and the removed `sandbox_node_id` is rejected with 400 like -any other unknown member. Context/output limits are optional nonnegative integers, with output no -larger than context; both must be positive for MiniMax Code. Use the actual model's -limits. Native provider availability is checked during execution, not by a new probe. -`agent.model` retains its exact provider model identity; a supplied value always -replaces the deployment model. - -The entire resolved provider configuration is frozen and encrypted in the Session creation -transaction, with a distinct credential-crypto purpose and tenant/Session binding. -Creation retries include this intent in their request hash; changing the key or -endpoint under the same idempotency key conflicts. A key enters any stored hash -only as a fingerprint keyed by the deployment credential key, never directly. Recovery reads the committed -Session before mutable Agent/template resolution. No public Session, Agent, -Environment, event or ordinary configuration contains the key. The top-level -extension is write-only and has no update endpoint. - -At dispatch, Core delivers its encrypted snapshot as one common confidential -provider bundle. Runtime adapters apply native options and connect directly to the -configured provider. Core does not fall back to other credentials when a snapshot is missing or -cannot decrypt. The same snapshot path serves every environment: Core sends the -options only over the daemon connection bound to the Session. For `self_hosted`, -that is the executor enrolled for the Session's own Environment with a current -executor credential of the Session creator's principal; rotation or revocation -closes the socket before further dispatch. The executor host stores the bundle in -its native harness home, as hosted Runtimes do. Public Files remains scoped to the -bound workspace, but native tools use the starting account's permissions and can -access whatever that user can read. The daemon does not isolate its local -credentials from same-user tools. Revocation does not erase an already delivered -bundle. +- `protocol` names the upstream API (`anthropic`, `responses` or `chat_completions`), not an engine. The selected Harness must support it natively. +- `base_url` uses HTTPS with a valid host, without credentials, query or fragment. +- `api_key` is nonempty, at most 16 KiB and contains no NUL, CR or LF. +- `context_window` and `max_output_tokens` are optional nonnegative integers, with output no larger than context; both must be positive for MiniMax Code. Use the real model's limits. +- `agent.model` is the exact provider model ID; a supplied value always replaces the deployment model. +- The Session's `x_agents_core` accepts `model_provider`, `harness_config` and `environment` ([Environments](environments.md#preparation-order)); any other member, such as `sandbox_node_id`, is rejected with 400. Hosted node placement is automatic. + +Unsupported protocol, Harness or Environment combinations are rejected before the Session is created. Provider availability is checked during execution, not by a probe. + +The resolved provider configuration is frozen and encrypted in the Session creation transaction, with its own encryption purpose and Project and Session binding. Creation retries include it in their request hash, so a changed key or endpoint under the same Idempotency-Key conflicts; a key enters any stored hash only as a fingerprint keyed by the deployment credential key. No public Session, Agent, Environment, event or ordinary configuration contains the key. The top-level extension is write-only and cannot be updated. + +At dispatch, Core sends the snapshot as one confidential provider bundle over the daemon connection bound to the Session, and the adapter applies it natively and connects directly to the provider. Core never falls back to other credentials when a snapshot is missing or cannot be decrypted. For `self_hosted`, the receiving daemon is the executor enrolled for the Session's own Environment with a current executor credential of the Session creator's principal; rotation or revocation closes the socket before further dispatch. The executor host stores the bundle in its native Harness home, as hosted Runtimes do. Native tools run with the starting account's permissions and can read what that account can read, and revocation does not erase a bundle already delivered. ## Native model parameters -The optional `harness_config` object uses the selected harness's native model -parameters. It is accepted on saved Agent `x_agents_core`, inline -`agent.x_agents_core`, and the Session's top-level `x_agents_core`. The Session -extension takes precedence over the inline extension. See -[Harness onboarding](harness-onboarding.md#native-model-configuration) for the -single supported-field reference and Runtime preparation contract. - -An explicitly supplied object replaces the entire object; `{}` clears it, and -null is invalid. A Session that explicitly selects a model or provider without -supplying native parameters uses `{}` rather than inheriting another model's -parameters. Otherwise a saved Agent supplies its object. An inline Session using -the deployment model uses its deployment object. Changing the selected harness -also clears inherited parameters. Updating a saved Agent's model, provider or -harness follows the same rule unless that update supplies `harness_config`. -There is no deep merge. An inline extension that only sets native parameters -preserves the saved harness; a null inline extension resets harness selection. Neither this object nor a deployment default enables -public `reasoning` execution options that the service does not already support. - -Core validates the resolved configuration before persistence and freezes it in the -Session Agent configuration. The administrator execution-configuration read records -its value and source. Native protocol support does not qualify structured output, -tool discovery, native web search, verbosity or image forms by itself. Existing -operation and input capability checks still apply before native submission, -including newly introduced images. A supported native connection does not -establish that the remote model accepts a parameter. - -Model parameters are safe, non-confidential fields; provider -keys remain in the separate encrypted bundle. Reconnect uses the frozen logical -configuration and cannot change the selected provider, protocol or native -parameters. An incompatible frozen configuration fails explicitly. +`harness_config` holds the selected Harness's native model parameters. Saved Agents accept it in `x_agents_core`, Sessions in the inline `agent.x_agents_core` and in the top-level `x_agents_core`; the top-level value wins. + +| Harness | Accepted fields | Applied as | +| --- | --- | --- | +| Codex | `model_reasoning_effort`: `none`, `minimal`, `low`, `medium`, `high`, `xhigh` | App-server `-c model_reasoning_effort=...` and each Turn's `collaborationMode.settings.reasoning_effort` | +| Claude SDK | `effort`: `low`, `medium`, `high`, `xhigh`, `max`; `thinking`: the SDK's `adaptive`, `enabled` or `disabled` object | SDK `Options.effort` and `Options.thinking` | +| MiniMax Code | Empty object only | Provider token limits remain required | + +Claude `thinking` accepts `display` (`summarized` or `omitted`) for `adaptive` and `enabled`, and a positive integer `budgetTokens` only for `enabled`; `maxThinkingTokens` is rejected. These are native settings, not a shared reasoning vocabulary; model availability and provider support are the Harness's responsibility. + +A supplied object replaces the whole object; `{}` clears it and null is invalid. A Session that selects a model or provider without supplying native parameters uses `{}` instead of inheriting another model's parameters. Otherwise a saved Agent supplies its object, and an inline Session using the deployment model uses the deployment object. Changing the Harness clears inherited parameters, and so does an Agent update of model, provider or Harness that does not supply `harness_config`. There is no deep merge. An inline extension that sets only native parameters keeps the saved Harness. Neither this object nor a deployment default enables public `reasoning` options. + +Core validates the resolved configuration before writing the Session and freezes it in the Session's Agent configuration; the administrator execution-configuration read shows its value and source. Native parameters are not confidential; provider keys stay in the separate encrypted bundle. Reconnection uses the frozen configuration and cannot change the provider, protocol or parameters; an incompatible frozen configuration fails. Protocol support does not qualify structured output, tool discovery, web search, verbosity or image input by itself, and a supported connection does not prove that the remote model accepts a parameter. ## Deployment defaults -The deployment default is a runtime setting stored in Core, one complete model configuration per -harness, managed with the Core key through Web or `/core/v1`: +The deployment default is a runtime setting stored in Core: one complete model configuration per Harness, managed with the Core key through Web or `/core/v1`. | Method and route | Result | | --- | --- | -| `GET /core/v1/harnesses` | Every harness this build supports, with `enabled` and `default` from the process configuration and its `model_configuration` (safe view) or null | +| `GET /core/v1/harnesses` | Every Harness this build supports, with `enabled` and `default` from the process configuration, its `model_configuration` (safe view) or null, and `model_configuration_support` | | `GET /core/v1/harnesses/{harness}/model-configuration` | The safe view; 404 when none is set | | `PUT /core/v1/harnesses/{harness}/model-configuration` | Replace `{model_provider, model, harness_config}` using the shared provider and native-parameter validators | | `DELETE /core/v1/harnesses/{harness}/model-configuration` | Remove it; idempotent, 204 | -PUT requires `model` and a complete `model_provider` bundle; `harness_config` -defaults to `{}`. Reads return `model`, `harness_config`, the safe `model_provider` -view, observations and `updated_at`, never the key. The provider view owns -`protocol`, `base_url`, optional token limits and `api_key_configured`. The bundle is encrypted with its own -credential-encryption purpose, bound to the harness, and each write records an -administrator audit entry (`resource_type: deployment_model_provider`, the harness -as `resource_id`, action `set` or `delete`, `project_id` null) without the key. A -missing or wrong encryption key fails closed: writes and Session creation that -needs the default return 503 `credential_storage_unavailable`. - -Session creation decrypts the default for the resolved harness and freezes it in the -Session's encrypted snapshot, like any other bundle. Changing or removing the -default never reaches existing Sessions, so a Session's first and later Turns always -use the same provider. The execution-configuration read shows the frozen safe view -with source `deployment`. - -Configure deployment defaults through `PUT /core/v1/harnesses/{harness}/model-configuration`. -The operator options file and historical native-option snapshots are unsupported. -A missing or invalid provider snapshot fails closed; no upgrade reader, automatic -migration or fallback to another model/provider is provided. - -Parsar manages its own workspace catalog and encrypted keys, sends this extension -only on the first Core Session request, and retains a private encrypted snapshot -for uncertain creation retries. Catalog updates and deletion affect new Sessions; -existing Sessions retain their original model, endpoint and key. +PUT requires `model` and a complete `model_provider`; `harness_config` defaults to `{}`. Reads return `model`, `harness_config`, the safe `model_provider` view (`protocol`, `base_url`, optional token limits and `api_key_configured`), observations and `updated_at`, never the key. The bundle is encrypted with its own purpose and bound to the Harness. Each write records an administrator audit entry (`resource_type: deployment_model_provider`, the Harness as `resource_id`, action `set` or `delete`, `project_id` null) without the key. A missing or wrong encryption key fails closed: writes, and Session creation that needs the default, return 503 `credential_storage_unavailable`. + +Session creation decrypts the default of the resolved Harness and freezes it in the Session's encrypted snapshot like any other bundle, so changing or removing the default never reaches existing Sessions. The execution-configuration read shows the frozen safe view with source `deployment`. A missing or invalid provider snapshot fails closed, with no fallback to another model or provider. + +`model_configuration_support` derives from the adapter declaration that Core and Runtime share: `protocols` lists the selectable native protocols with the default first, `accepts_harness_config` says whether native parameters are accepted and `token_limits_required` whether provider token limits are required. It describes the build, not a live Runtime or a remote model. ### Deployment default observations -The Core-only harness/default-provider reads include nullable `last_used_at`, -`last_error_code` and `last_error_at`. PUT resets all three, even for the same bundle. -Completed root Turns contribute use observations only when their Session froze -that exact current default revision. Failed root Turns contribute only the fixed -native provider codes `authentication_error`, `connection_failed`, -`rate_limit_exceeded`, `usage_limit_exceeded`, `server_overloaded`, `server_error`, -`resource_not_found`, `request_timeout` and `invalid_request`. Input-policy, -Core/runtime, cancelled and waiting outcomes do not contribute. Successful use -retains the earlier error; comparing timestamps is only a display convention. - -These are best-effort Core receipt times after terminal commit, not provider health -or remote completion times. Private revision identity is independent of timestamps; -old, explicit-provider and historical Sessions cannot update a replacement default. -No public Session/Turn fields or retry identity change. The observation has a -one-second budget including pool/row-lock acquisition and cannot change the committed -Turn. A crash, failure or throttle can omit the last observation indefinitely. - -Errors throttle for 30 seconds regardless of code. Ordinary successful writes -throttle for 30 seconds; the first success after an accepted error records recovery -immediately. For an unchanged revision this allows at most three effective metadata -writes in any half-open 30-second interval of nondecreasing DB-clock time. Equal or -backwards clock readings are preserved and can make timestamp-based display -ambiguous; recovery does not impose a total order. No readiness probe, automatic -refresh, observation history or credential/raw-error read is provided. - -The harness list also returns `model_configuration_support`, derived from the -same adapter declaration used by Core and Runtime: `protocols` lists selectable -native upstream protocols in default order, `accepts_harness_config` reports -whether native parameters are accepted, and `token_limits_required` reports -required provider token limits. `protocols` is the sole protocol list; its first -entry is the configuration form's default. This descriptor describes the current -build, not a live Runtime or a remote model's availability. +The Harness and default-model reads include nullable `last_used_at`, `last_error_code` and `last_error_at`; PUT resets all three, even for the same bundle. + +- Completed root Turns record use only when their Session froze the exact current default revision. +- Failed root Turns record only the native provider codes `authentication_error`, `connection_failed`, `rate_limit_exceeded`, `usage_limit_exceeded`, `server_overloaded`, `server_error`, `resource_not_found`, `request_timeout` and `invalid_request`. Input-policy, Core or Runtime, cancelled and waiting outcomes do not count. +- A success keeps the earlier error; comparing the timestamps is a display convention. +- Errors are throttled for 30 seconds regardless of code, and ordinary successes for 30 seconds; the first success after an error records recovery at once. For an unchanged revision this allows at most three effective writes in any 30-second window of database time. +- Each observation has a one-second budget, including pool and row-lock acquisition, and cannot change the committed Turn. A crash, failure or throttle can omit the latest observation. + +These are best-effort Core receipt times after the terminal commit, not provider health or remote completion times. Old, explicit-provider and earlier Sessions cannot update a replacement default. Public Session and Turn fields and retry identity are unaffected. Core has no readiness probe, automatic refresh, observation history or read of credentials or raw errors. + +## Acceptance + +`TestNativeModelProtocolPublicExecution` (`services/core/internal/store/model_protocol_native_test.go`) with `services/core/tests/official_model_protocol_native.py` runs each Harness against real provider APIs through the pinned official client. It runs when `OAC_TEST_OFFICIAL_SDK_PYTHON`, `OAC_TEST_NATIVE_DAEMON_BIN`, `OAC_TEST_NATIVE_PROOF_DIR` and `OAC_TEST_MODEL_PROTOCOL_OPTIONS` are set; the last names a private model settings file. Never commit those settings or print their values. diff --git a/contracts/agents-api/model-protocol-benchmarks.md b/contracts/agents-api/model-protocol-benchmarks.md deleted file mode 100644 index 3444b17b4..000000000 --- a/contracts/agents-api/model-protocol-benchmarks.md +++ /dev/null @@ -1,106 +0,0 @@ -# Model transport benchmarks — historical, retired - -> Historical record of the retired built-in model proxy and converter. The -> implementation, dependency choices, commands and results below describe that -> earlier revision only; they are not current setup instructions, dependencies -> or supported protocol combinations. The current native-only contract is -> [model execution](model-execution.md#saved-defaults-and-precedence). No -> compatibility alias, automatic migration or restoration of this converter is supported. - -Measured on 2026-09-28 on Linux amd64, Intel Xeon Platinum 8358P @ 2.60 GHz, -Go 1.26.8, GOMAXPROCS=8, CLIProxyAPI v8.0.3. This measures the thin Exchange -integration, including the composed Responses-to-Messages route. The shared host -was not isolated from other workloads; figures are microbenchmarks, not latency SLOs. -No model or network is used by these benchmark workloads. - -Short input is one user message and a function declaration. Large input contains -128 alternating messages with 2 KiB each. Tool loop includes a streamed function -call, its result in a second request, and a text stream. Concurrent SSE uses 32 -workers, 64 deltas and 8 KiB of text per stream. Outputs are consumed and discarded. -HTTP framing, backpressure and native engine processing are excluded. - -The tables report the median of three testing.Benchmark allocation runs with a -200 ms target. Sequential latency uses 20 warmups and 500 samples; concurrent -latency uses 800 streams. p95/p99 include scheduling and GC. Concurrent mean is -aggregate throughput per stream, rather than an individual stream's latency. - -## Short request - -| Client → upstream | Mean (ms) | p95 (ms) | p99 (ms) | Bytes/op | Allocs/op | -| --- | ---: | ---: | ---: | ---: | ---: | -| anthropic → responses | 0.076 | 0.046 | 0.076 | 13,605 | 137 | -| anthropic → chat_completions | 0.059 | 0.065 | 0.099 | 7,614 | 84 | -| responses → anthropic | 0.205 | 0.276 | 0.552 | 19,807 | 239 | -| responses → chat_completions | 0.074 | 0.100 | 0.215 | 10,996 | 120 | -| chat_completions → anthropic | 0.129 | 0.189 | 0.316 | 10,777 | 132 | -| chat_completions → responses | 0.090 | 0.040 | 0.065 | 11,579 | 106 | - -## 256 KiB history - -| Client → upstream | Mean (ms) | p95 (ms) | p99 (ms) | Bytes/op | Allocs/op | -| --- | ---: | ---: | ---: | ---: | ---: | -| anthropic → responses | 15.484 | 18.166 | 34.967 | 4,021,333 | 2,568 | -| anthropic → chat_completions | 15.872 | 18.228 | 19.874 | 3,588,969 | 1,553 | -| responses → anthropic | 56.984 | 63.098 | 75.256 | 8,020,006 | 5,743 | -| responses → chat_completions | 26.077 | 27.209 | 33.847 | 3,631,486 | 1,603 | -| chat_completions → anthropic | 37.060 | 41.171 | 46.010 | 5,740,777 | 4,153 | -| chat_completions → responses | 17.402 | 30.225 | 43.999 | 5,161,877 | 3,370 | - -## Two-request tool loop - -| Client → upstream | Mean (ms) | p95 (ms) | p99 (ms) | Bytes/op | Allocs/op | -| --- | ---: | ---: | ---: | ---: | ---: | -| anthropic → responses | 0.676 | 1.056 | 1.599 | 133,175 | 923 | -| anthropic → chat_completions | 0.555 | 1.006 | 1.515 | 116,540 | 794 | -| responses → anthropic | 1.927 | 3.170 | 3.520 | 315,546 | 2,410 | -| responses → chat_completions | 1.180 | 1.874 | 2.153 | 234,483 | 1,541 | -| chat_completions → anthropic | 0.782 | 0.803 | 0.984 | 95,143 | 993 | -| chat_completions → responses | 0.605 | 0.965 | 1.525 | 83,962 | 808 | - -## Concurrent SSE - -| Client → upstream | Mean (ms) | p95 (ms) | p99 (ms) | Bytes/op | Allocs/op | -| --- | ---: | ---: | ---: | ---: | ---: | -| anthropic → responses | 0.318 | 40.509 | 64.359 | 728,156 | 2,814 | -| anthropic → chat_completions | 0.285 | 22.059 | 33.589 | 593,693 | 3,028 | -| responses → anthropic | 0.625 | 38.772 | 55.790 | 1,493,381 | 6,200 | -| responses → chat_completions | 0.513 | 28.864 | 39.006 | 1,233,740 | 4,003 | -| chat_completions → anthropic | 0.167 | 19.135 | 28.550 | 323,167 | 2,796 | -| chat_completions → responses | 0.211 | 24.962 | 33.010 | 461,607 | 2,814 | - -## Sampled heap - -During the concurrent run, runtime.MemStats is sampled every 1 ms after a baseline -GC. These are observed Go heap peaks, not exact maxima, OS RSS, per-Session retained -memory or a ceiling. Sampling can miss short spikes and perturbs timings. - -| Client → upstream | Baseline heap (B) | Sampled peak (B) | Increase (B) | -| --- | ---: | ---: | ---: | -| anthropic → responses | 2,256,344 | 8,569,920 | 6,313,576 | -| anthropic → chat_completions | 2,260,440 | 7,682,336 | 5,421,896 | -| responses → anthropic | 2,283,664 | 12,052,232 | 9,768,568 | -| responses → chat_completions | 2,310,952 | 10,051,656 | 7,740,704 | -| chat_completions → anthropic | 2,316,000 | 9,032,160 | 6,716,160 | -| chat_completions → responses | 2,437,928 | 7,721,664 | 5,283,736 | - -## Binary cost - -With Go 1.26.8 and `go build -trimpath -ldflags="-s -w"`, the Runtime executable -is 10,051,849 bytes at baseline `5b5c10fc4afeb20a1f5ac52207a9810ece85dbed` and -19,157,257 bytes with this integration, an increase of 9,105,408 bytes. -This is separate from container image size and process memory. Same-protocol -production connections bypass the converter and are covered by selection tests. - -## Reproduce - -```sh -GOMAXPROCS=8 go test ./internal/modeltransport -run '^$' \ - -bench '^BenchmarkProduct' -benchmem -benchtime=200ms -count=3 -OAC_MODELTRANSPORT_PROFILE=1 GOMAXPROCS=8 go test ./internal/modeltransport \ - -run '^TestProductBenchmarkProfile$' -count=1 -v -``` - -The profile is opt-in. The source fixtures live in `internal/modeltransport`. -Raw samples are retained as `thin-bench.log` and `thin-profile.log` in the Linux -qualification workspace. The independent library comparison uses Go 1.27 and -is not a toolchain-controlled comparison with these product numbers. diff --git a/contracts/agents-api/model-protocol-conversion.md b/contracts/agents-api/model-protocol-conversion.md deleted file mode 100644 index cf5a9139b..000000000 --- a/contracts/agents-api/model-protocol-conversion.md +++ /dev/null @@ -1,96 +0,0 @@ -# Runtime model protocol conversion — historical, retired - -> Historical record of the retired built-in model proxy and converter. The -> implementation, dependency choices, commands and results below describe that -> earlier revision only; they are not current setup instructions, dependencies -> or supported protocol combinations. The current native-only contract is -> [model execution](model-execution.md#saved-defaults-and-precedence). No -> compatibility alias, automatic migration or restoration of this converter is supported. - -The model provider bundle names the upstream protocol, endpoint, credential and -limits. The Agent model and execution harness are separate selections. Runtime -uses the adapter's native protocols to choose a direct connection or a local -conversion endpoint automatically; model names never select a protocol. - -| Engine | Anthropic Messages upstream | OpenAI Responses upstream | OpenAI Chat Completions upstream | -| --- | --- | --- | --- | -| Claude SDK / Claude Code | Native | Convert Messages to Responses and back | Convert Messages to Chat Completions and back | -| Codex | Convert Responses to Messages and back | Native | Convert Responses to Chat Completions and back | -| MiniMax Code | Native | Native | Native | - -These are model communication paths, not claims that every model implements every -engine capability. MiniMax Code uses its pinned native `custom_provider.api` -selection (`anthropic-messages`, `openai-responses`, `openai-completions`). Do not -force it through another protocol merely to exercise a converter. - -## Ownership and transport - -Core freezes one confidential provider bundle and sends it through the existing -execution contract. It does not emit engine-specific provider options. The same -Runtime implementation serves Core-created and self-hosted environments. Session -scheduling, environments, native execution, tools, Skills and MCP retain their -existing owners. No gateway service, accounts, routing catalog, billing subsystem -or additional Agent loop is installed. - -A conversion endpoint binds only loopback and requires a random per-endpoint -credential. Native configuration receives that credential; the upstream key stays -inside Runtime. Only the configured model endpoint is exposed. Upstream redirects -are refused, and upstream error bodies are never exposed as native diagnostics. -Native URL query flags are accepted on the configured operation but are not -forwarded to the provider. Streaming Chat requests enable `include_usage` so the -provider can send its final usage frame. Token conversion remains in the SDK. -SSE data is converted incrementally and flushed without waiting for the complete -answer. Chat completion requires either `[DONE]` after a valid finish reason, or -a separate empty-choice usage tail after that finish reason. The latter is used by compatible providers that omit `[DONE]`. A bare EOF without a protocol end cannot complete an exchange. -Closing the execution resource cancels requests and releases the endpoint; -individual Turn completion does not establish Session resource ownership. -Same-protocol execution keeps the original native connection and behavior. - -The upstream base URL has the native protocol's usual semantics: Responses and -Chat Completions append `/responses` or `/chat/completions`; include `/v1` in that -base when the provider requires it. Messages appends `/v1/messages`, avoiding a -second `/v1` when already present. Public provider admission requires HTTPS. -There is no endpoint probing or automatic model substitution. - -## Conversion limits - -The conversion dependency is pinned to CLIProxyAPI v8.0.3 (MIT). Only its -public translator SDK is embedded; gateway, account and management services are -not started. The SDK owns field, tool, image, reasoning and usage conversion. -Runtime supplies per-request SDK state and the HTTP/SSE connection. It does not -maintain field allowlists, parameter restoration, tool-name mappings, reasoning -signature interpretation or token accounting alongside the dependency. - -Cross-protocol conversion inherits the pinned library's behavior. Provider-only -tools, grammar enforcement, opaque reasoning state, detailed reasoning usage and -other native extensions are not universally portable. The SDK may omit or -normalize those fields. Exact preservation of arbitrary provider extensions is -not a supported cross-protocol guarantee. The SDK's Responses target applies -Codex defaults to stream/store/parallel-tool settings and can omit sampling and -token-limit controls. Custom grammar tools become string-argument functions; -the target does not enforce the original grammar. For Chat providers that send -empty or provisional usage on a finish chunk before definitive usage, the pinned -SDK's Messages output can retain the earlier counts. These inherited behaviors -are not repaired with local conversion rules. Same-protocol connections retain -the native API. Gemini conversion is not enabled in this first set. - -Dependency updates are explicit version changes, validated with request/response, -stream and real engine regression tests. Prefer an upstream release for conversion -repairs; any unavoidable local patch requires a separately justified narrow scope. - -The project is unpublished. Historical engine-specific provider option snapshots -and alternate legacy wire shapes are unsupported. No automatic migration, old -configuration reader or data deletion accompanies this change. - -## Validation - -Use the same engine/protocol/capability matrix for hosted and self-hosted model -communication. Validate multiple tool rounds, text deltas, supported image and -reasoning forms, cancellation, abnormal termination, errors and usage. A synthetic -HTTP fixture or one-line native answer alone is not real model qualification. -Pure conversion benchmarks and HTTP control tests establish separate evidence; -record real native-model acceptance and any unverified cells explicitly. - -See the [pinned library evaluation](model-protocol-library-evaluation.md) and -[product benchmarks](model-protocol-benchmarks.md) for dependency and measured -conversion costs. These measurements are separate from live model acceptance. diff --git a/contracts/agents-api/model-protocol-library-evaluation.md b/contracts/agents-api/model-protocol-library-evaluation.md deleted file mode 100644 index 1082c6e50..000000000 --- a/contracts/agents-api/model-protocol-library-evaluation.md +++ /dev/null @@ -1,123 +0,0 @@ -# Model protocol conversion library evaluation (2026-09-28) — historical, retired - -> Historical record of the retired built-in model proxy and converter. The -> implementation, dependency choices, commands and results below describe that -> earlier revision only; they are not current setup instructions, dependencies -> or supported protocol combinations. The current native-only contract is -> [model execution](model-execution.md#saved-defaults-and-precedence). No -> compatibility alias, automatic migration or restoration of this converter is supported. - -Scope: in-process conversion among OpenAI Chat Completions, OpenAI Responses and Anthropic Messages. No gateway processes, account pools, billing, agent loops, real model requests or credentials were used. - -## Decision - -The repository is MIT-licensed. Prefer native protocol operation when a harness already supports the configured provider protocol. For cross-protocol operation, use the maintained CLIProxyAPI public translator SDK at v8.0.3 behind the Runtime-owned Exchange interface. This keeps the upstream transformation implementation upgradeable. The dependency owns protocol conversion, including field selection, tools, reasoning and usage. The host only selects registered conversion routes, supplies per-request SDK state and handles HTTP/SSE transport. Pin the dependency and qualify upgrades; do not maintain a second converter through local field maps or parameter restoration. SDK behavior is not universally lossless; limitations need explicit documentation and engine-level regression evidence. - -RelayKit is the cleanest small protocol-only module, but is AGPL-3.0. Do not add it as the MIT repository's default dependency without resolving licensing compatibility. Bifrost is Apache-2.0, faster than CLIProxyAPI in this bounded sample, but its provider/schema surface is substantially broader than a converter, requires Go 1.27, and brings MCP and transport dependencies even when only the two provider packages are imported. - -## Fixed sources - -- RelayKit: new-api commit 789c970199ea527e6a26e071915f4a4cd2c64178; module github.com/QuantumNous/new-api/relaykit; Go 1.25.1. Published relaykit/v0.2.1 points to 0aec08fee811ec6136828fda790551b49e410301 (2026-09-21). Bench/tests used the newer pinned commit. -- CLIProxyAPI: v8.0.3, acdace936fa7df2905500c7f5e0a97d683138dea; github.com/router-for-me/CLIProxyAPI/v8; MIT; Go 1.26.0 minimum. -- Bifrost: 51172dbe37df018384d0f15f94940e01c40a7fff; github.com/maximhq/bifrost/core; Apache-2.0; Go 1.27.0. - -CLI v6.10.9 (2026-05-07), v6.9.0 and v6.8.55 already require Go 1.26.0. v6.8.9 uses Go 1.24.0, but predates the relevant recent translator repairs; no recent equivalent Go-1.25 version was established. Do not silently lower its go directive. - -## Import closure and reusable APIs - -| Candidate/import | Non-standard packages | Modules, including candidate | Relevant dependencies | -|---|---:|---:|---| -| RelayKit relayconvert | 35 | 8 | uuid, lo, gjson/sjson, match/pretty, x/text | -| CLI sdk/translator/builtin | 135 | 26 | Gin, Redis, logrus, protobuf, crypto, gjson/sjson | -| Bifrost providers/anthropic + providers/openai | 126 | 32 | sonic, fasthttp, websocket, MCP, compression, JSON-schema utilities | - -Counts are actual go list -deps results, not root go.mod requirement counts. Raw evidence remains in the Linux qualification workspace; it is not a repository artifact. - -RelayKit exposes ConvertRequest, ConvertResponse, NewResponseStreamState, ConvertStreamResponseChunk and FinalizeStreamResponse. It owns no HTTP or SSE reading/writing. Result objects include route, quality, usage and diagnostics. Images needing materialization require a host MediaResolver; the library does not download content. - -CLI exposes sdk/translator.Registry and sdk/translator/builtin.Registry(). Request translation returns bytes; response/stream translation returns bytes/chunks plus caller-owned per-exchange state through *any. Responses as an incoming client format is openai-response; the outgoing generic Responses-shaped implementation is registered as codex. The bare registry imports only 4 internal packages and logrus/gjson/sjson/yaml, but contains no registered concrete conversions. builtin imports 50 internal packages. A static closure of the six relevant conversion directions still traverses 23 internal packages and Gin/config/util dependencies. Internal conversion packages cannot be imported outside that module. A copied subset would require removing registration init files and extracting tool/schema/JSON helpers from internal/util plus common/signature/thinking helpers; this is a maintained fork, not a drop-in tiny module. - -Bifrost can be used without constructing its gateway. OpenAIChatRequest.ToBifrostChatRequest plus anthropic.ToAnthropicChatRequest is the tested conversion path. AnthropicMessageRequest.ToBifrostResponsesRequest, ToAnthropicResponsesRequest, response methods and AnthropicResponsesStreamState/ToBifrostResponsesStream/ToAnthropicResponsesStreamResponse cover the corresponding typed conversions. Chat/Responses interchange additionally uses core/schemas conversion methods. Importing those packages also compiles their provider HTTP code and shared schema/MCP definitions. - -## Evidence of maintenance and defects - -RelayKit commit 4eb3b9160566dc4c1f4f0e9340de254f7febf056 (2026-09-21) repairs tool_result image blocks previously sent as text and Responses reasoning segmentation after a mid-stream finish. It adds 292 lines of tests across request/response tests. Current tests include conversion matrix golden fixtures, tool-loss policies, terminal stream tails, usage and failed Responses events. Strict loss policy rejects only request conversion loss; response and stream loss still return successful results with diagnostics. This is a deliberate documented limitation, not a hard fail-closed library. - -CLI commit 0a45f253344089c2f47ae698c34abe8837ec6135 (2026-09-28) repairs missing blank-line SSE terminators and adds regression tests. 4ad5bba repairs false namespace-prefix matches (2026-09-28); 75b854e repairs Claude tool-name sanitation and parameterless schemas (2026-09-24). Tests cover parallel tools, cached usage, incomplete Responses, custom-tool replay, namespace collisions and tool-result adjacency. However: -- The public transforms have no error return; missing registrations return the original body. -- TestConvertOpenAIResponsesRequestToClaude_DropsApplyPatchCustomTool explicitly requires dropping the reserved apply_patch custom grammar tool. -- The codex target rewrites generic parameters, reasoning defaults, stream/store/include. -- Claude-to-Chat and Claude-to-Responses non-stream functions expect an aggregated SSE transcript, while the Codex non-stream response functions expect a response.completed/incomplete envelope. -- Some stream conversions emit client completion at a Chat finish_reason before upstream [DONE]. -- Claude-to-Responses estimates reasoning_tokens from visible text length. - -The integration uses only public registered transformers. Responses-to-Messages -can compose the library's Responses-to-Chat and Chat-to-Messages routes, retaining -separate SDK state and original/translated request bytes for each step. This avoids -the direct route's reserved custom-tool behavior without local tool-name or schema -mapping. Composition must pass the same tool-history and stream tests as direct -routes. HTTP framing handles stream termination; it does not calculate usage or -interpret reasoning signatures. Non-stream Messages conversion can request an -upstream stream and pass its bounded transcript to the library's aggregate entry. -No local synthetic Messages events or field restoration are retained. - -The dependency may normalize parameters, omit provider-specific extensions or -estimate a reasoning-token breakdown. Those behaviors belong to the pinned SDK; -this integration does not claim native fidelity across protocols. Prefer a native -connection when exact provider-specific semantics are needed. A library upgrade -must pass the stored contract and real engine checks before adoption. - -Bifrost commit a44106ad (2026-09-24) repairs Responses types and Anthropic tool-result is_error preservation with dedicated tests; e0ad6638 (2026-09-23) preserves Anthropic extra parameters through Bedrock conversion. Its current tests also deliberately drop forced tool choice on unsupported model capabilities. Typed conversion and model-aware normalization can therefore still change semantics; typed APIs do not imply universal field preservation. - -## First-batch boundary - -CLIProxyAPI registers Gemini conversions in the same public registry. Enabling -Gemini also requires a provider configuration contract, model/credential URL -handling and real engine qualification. Those are not established in this batch; -the admitted protocols remain the initial three. No parallel Gemini gateway or -new provider manager is introduced. - -The measured costs support using CLIProxyAPI despite its larger dependency -closure: the chosen license and tested tool coverage avoid a maintained fork. -The product qualification budget for this host is a short-request p99 below 1 ms, -a 256 KiB request p99 below 100 ms and an executable increase below 12 MiB. These -are review thresholds for dependency upgrades, not production SLOs. Concurrent -heap samples are retained for comparison and do not establish a memory ceiling. - -## Bounded performance sample - -All final samples ran on the same Linux host using Go 1.27.0, GOMAXPROCS=1, stripped binaries (-s -w), no HTTP or real model. Direction: Chat -> Claude, including JSON decode/encode for typed libraries. CLI's public API operates directly on bytes. Baseline: JSON decode and re-encode only. Inputs: 94-byte single text; 263,212-byte history with 32 alternating user/assistant messages (8 KiB text each); 1,496-byte function declaration/call/result/follow-up. testing.Benchmark supplies mean and allocation counts; a separate 1,000-operation sample after GC supplies percentiles. These are comparative microbenchmarks, not production SLOs. - -| Candidate | Input | Mean | p95 | p99 | Bytes/op | Allocs/op | -|---|---|---:|---:|---:|---:|---:| -| Baseline | small | 6.44 us | 6.29 us | 8.47 us | 1,176 | 32 | -| Baseline | large | 0.636 ms | 0.603 ms | 0.671 ms | 548,346 | 350 | -| Baseline | tools | 24.86 us | 34.50 us | 40.57 us | 8,273 | 132 | -| RelayKit | small | 6.77 us | 10.29 us | 15.91 us | 5,008 | 22 | -| RelayKit | large | 0.721 ms | 0.898 ms | 0.961 ms | 567,051 | 194 | -| RelayKit | tools | 31.12 us | 41.94 us | 89.09 us | 15,698 | 132 | -| CLIProxyAPI | small | 19.60 us | 21.79 us | 58.41 us | 3,536 | 59 | -| CLIProxyAPI | large | 7.788 ms | 7.705 ms | 8.402 ms | 3,252,011 | 1,040 | -| CLIProxyAPI | tools | 93.12 us | 129.67 us | 176.01 us | 27,677 | 212 | -| Bifrost | small | 9.00 us | 12.31 us | 41.72 us | 3,802 | 32 | -| Bifrost | large | 1.221 ms | 1.753 ms | 1.956 ms | 1,988,175 | 296 | -| Bifrost | tools | 46.38 us | 90.94 us | 144.14 us | 25,284 | 173 | - -Stripped binary sizes: baseline 2,834,592 B; RelayKit 6,549,767 B (+3,715,175 B); CLI 14,782,727 B (+11,948,135 B); Bifrost 9,924,871 B (+7,090,279 B). This is an independent harness delta, not a prediction of the exact final daemon delta. - -## Validation and reproducing - -RelayKit GOWORK=off go test ./... passed at the pinned commit. CLI sdk/translator and the selected Claude/OpenAI conversion test packages passed. Bifrost selected pure converter tests (TestToAnthropic, TestToBifrost, tool conversion/stream names) passed. None of these results establishes live provider compatibility. - -Each checkout contains benchprotocol/main.go; baseline/main.go is the identical fixture/measurement harness without conversion. *-bench.log, *-tests.log, and *-deps.json hold raw output. The root task's product worktree contains its separate Exchange contract tests and HTTP mock-provider tests; this report does not claim full repository make check. - -## Primary source links - -- [RelayKit API and license](https://github.com/QuantumNous/new-api/blob/789c970199ea527e6a26e071915f4a4cd2c64178/relaykit/README.md) -- [RelayKit actual dependencies](https://github.com/QuantumNous/new-api/blob/789c970199ea527e6a26e071915f4a4cd2c64178/relaykit/go.mod) -- [RelayKit media/stream repair](https://github.com/QuantumNous/new-api/commit/4eb3b9160566dc4c1f4f0e9340de254f7febf056) -- [CLI public registry](https://github.com/router-for-me/CLIProxyAPI/blob/acdace936fa7df2905500c7f5e0a97d683138dea/sdk/translator/registry.go) -- [CLI reserved custom-tool test](https://github.com/router-for-me/CLIProxyAPI/blob/acdace936fa7df2905500c7f5e0a97d683138dea/internal/translator/claude/openai/responses/claude_openai-responses_request_test.go) -- [CLI SSE repair](https://github.com/router-for-me/CLIProxyAPI/commit/0a45f253344089c2f47ae698c34abe8837ec6135) -- [Bifrost module](https://github.com/maximhq/bifrost/blob/51172dbe37df018384d0f15f94940e01c40a7fff/core/go.mod) -- [Bifrost Responses converters](https://github.com/maximhq/bifrost/blob/51172dbe37df018384d0f15f94940e01c40a7fff/core/providers/anthropic/responses.go) diff --git a/contracts/agents-api/model-protocol-qualification.md b/contracts/agents-api/model-protocol-qualification.md deleted file mode 100644 index 7ff0bfa9e..000000000 --- a/contracts/agents-api/model-protocol-qualification.md +++ /dev/null @@ -1,54 +0,0 @@ -# Model protocol qualification (2026-09-28) — historical, retired - -> Historical record of the retired built-in model proxy and converter. The -> implementation, dependency choices, commands and results below describe that -> earlier revision only; they are not current setup instructions, dependencies -> or supported protocol combinations. The current native-only contract is -> [model execution](model-execution.md#saved-defaults-and-precedence). No -> compatibility alias, automatic migration or restoration of this converter is supported. - -The pinned converter is CLIProxyAPI v8.0.3. Real execution used MiniMax-M2.7 -through the configured MiniMax Chat Completions, Responses and Messages APIs, -Codex 0.153.4, Claude Agent SDK 0.3.269 / Claude Code 2.1.269 and MiniMax Code -0.4.12. The Claude bridge uses Executor protocol 3. Each run creates a fresh -public Session through the pinned official Python client and freezes the provider -through normal deployment-default admission. No test Dispatcher.Options override -supplies model selection. - -| Engine | Messages | Responses | Chat Completions | -| --- | --- | --- | --- | -| Claude Code | Native: passed | Converted: passed | Converted: passed | -| Codex | Converted: passed | Native: passed | Converted: passed | -| MiniMax Code | Native: passed | Native: passed | Native: passed | - -Each Claude/Codex case covers two successful public function rounds, one failed -function result, incremental text, cancellation, continuation after cancellation -and a Runtime restart followed by exact native Session continuation. Persisted -function-result delivery is checked for loss or duplication. The engine's public -Session response is checked for provider-key exposure. Each MiniMax Code case -covers remembered context and exact continuation after a Runtime restart; public -function calls and cancellation are not claimed by those MiniMax cases. - -This is a model-communication acceptance suite using the existing environment:none -profile. Hosted and self-hosted use the same Runtime path in code; this batch does -not repeat a complete Docker/E2B/self-hosted deployment qualification. Native -model reasoning appeared in the real responses. Image payload preservation, -visible reasoning projection, usage, namespaced/custom tools and abnormal stream -termination are additionally checked by controlled converter fixtures. No real -vision-model qualification is claimed. Public token usage can remain null when an -adapter preserves only a partial native breakdown; this task does not add billing -or invent counters to fill that gap. - -The endpoint tests cover upstream credential isolation, redirect refusal, -authentication, first-text flushing before completion, cancellation, final cleanup, -non-stream SDK aggregation and interrupted streams. An Executor test verifies that -the endpoint survives completed and cancelled Turns and closes on final teardown. -The console's three affected browser scenarios passed, including protocol choice, -save/edit/cancel, token limits and write-only credentials. - -The real entrypoints are TestNativeModelProtocolPublicExecution and -services/core/tests/official_model_protocol_native.py. Private model settings -are supplied by OAC_TEST_MODEL_PROTOCOL_OPTIONS; do not commit them or print their -values. The suite records only controlled checks, IDs and event counts under -OAC_TEST_NATIVE_PROOF_DIR. Library and product performance evidence are separate: -see model-protocol-library-evaluation.md and model-protocol-benchmarks.md. diff --git a/contracts/agents-api/official-semantics-alignment.md b/contracts/agents-api/official-semantics-alignment.md index 22f1ff9a8..3bd005253 100644 --- a/contracts/agents-api/official-semantics-alignment.md +++ b/contracts/agents-api/official-semantics-alignment.md @@ -745,7 +745,7 @@ package; both were deleted. The batch plan is | H1 | Hosted initialization fails in any step | One transaction records the Environment failure, `agent.session.environment.failed`, `error` and one `agent.session.failed`. The Session reads `status: failed`, the stored safe reason as `error`, `required_actions: []` and the failure time as `last_active_at`; retrieve, list and the event snapshot agree. Pending input reserved for the Environment settles as failed exactly as before, captured in the same snapshot. | | H2 | `environment.failed` payload | `error` is `{type: environment_error, code: environment_connection_failed, message: "The environment failed to connect."}`. | | H3 | `error` event | `{type: environment_error, code: sandbox_error, message: , param: null}`. | -| H4 | Reason | `Failed to provision environment: script "setup_commands[i]" failed with exit code N`, and `script "Python package installation"` for Python packages; the official Python reason also appends raw pip output, which Core never copies. npm, system package, initial file and Skill labels are unverified. Other failures use `Failed to provision environment: initialization did not complete`; see the [initialization lifecycle](environment-templates.md#initialization-failure--september-23). | +| H4 | Reason | `Failed to provision environment: script "setup_commands[i]" failed with exit code N`, and `script "Python package installation"` for Python packages; the official Python reason also appends raw pip output, which Core never copies. npm, system package, initial file and Skill labels are unverified. Other failures use `Failed to provision environment: initialization did not complete`; see the [initialization lifecycle](environments.md#initialization-state-and-failure). | | H5 | Live SSE | GET and creation streams end right after that `agent.session.failed`. | | H6 | Later `events.create` | 409 `conflict_error`/`conflict_error` "the hosted environment failed to provision", param null. Expired Environments, and input already waiting when the Environment failed, keep 409 `environment_unavailable`. | | H7 | Delete | 200 `agent.session.deleted`, as officially, then 404. Deletion while provisioning is unchanged (HI-05 awaits a decision). | @@ -755,7 +755,7 @@ Historical implementation decisions at the September 23 revision follow. The current daemon uses shared Go preparation and process settlement, with no bwrap, Python receipt decoder or old-image compatibility. System packages now reject before initialization. Safe fixed labels and bounded exit statuses remain the -current public failure policy; see the [current initialization contract](environment-templates.md#packaged-runtime-initialization-contract). +current public failure policy; see the [current initialization contract](environments.md#runtime-capability-preparation). Recorded decisions: diff --git a/contracts/agents-api/operation-evidence.md b/contracts/agents-api/operation-evidence.md index 4ca748764..bd59162b9 100644 --- a/contracts/agents-api/operation-evidence.md +++ b/contracts/agents-api/operation-evidence.md @@ -41,7 +41,7 @@ Repository paths below are relative to the inspected worktree; private evidence | L | [List query tolerance](list-query-semantics.md#list-query-tolerance--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-1/{vaults-agents,sessions,skills-files-templates}/findings.json`: owned-collection unknown/repeated keys, limit bounds, Vault status union and Files empty purpose, plus unknown keys on a deleted Vault read and Agent delete. Rows A1–D2 of that section; no model execution. | | G | [Environment Files wire alignment](environment-files.md#wire-alignment--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/findings.json` HE-10, 16, 18, 32, 34–39 with raw records under `official/` (labels `fc01`–`fc17`, `fl01`–`fl16`, `files-*-pending`) and `run1/`: three owned hosted Sessions, all deleted; first official Environment Files observations. Rows F1–F9 of that section. Go handler, real-PostgreSQL Worker, gateway/daemon and Rust helper tests without a model; live acceptance is recorded with the batch. | | M | [Agent configuration validation](official-semantics-alignment.md#agent-configuration-validation--september-23); private `~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` TV-01..07 with raw records under `official/validation-{B1,B2A,B2B,B3}.json`: one owned Agent (deleted) and 44 requests without a model, Session creates without input on `none`. Rows C1–C4 and K1–K3 of that section. Go handler, real-PostgreSQL no-write/tenant replay and pinned-SDK tests without a model; live acceptance is recorded with the batch. | -| U | [Subagent visibility](subagents.md#subagent-visibility--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` SAT-01, 02, 07, 08, 09 (SAT-03 for the kept limit rejections) with raw records under `official/` (labels `R00`–`R33`, `C01`–`C06`, S1/S2 stream frames): two owned official Sessions with two child Turns, all deleted; the first official Subagent observations. Rows A1–A5 of that section; SAT-03/05/13/14 match, SAT-04/06/10 and SAT-12 stay deferred or unknown. Go store/API and real-PostgreSQL HTTP tests, TypeScript client and Core Web tests; live acceptance is recorded with the batch. | +| U | [Subagent visibility](subagents.md#subagent-visibility); private `~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` SAT-01, 02, 07, 08, 09 (SAT-03 for the kept limit rejections) with raw records under `official/` (labels `R00`–`R33`, `C01`–`C06`, S1/S2 stream frames): two owned official Sessions with two child Turns, all deleted; the first official Subagent observations. Rows A1–A5 of that section; SAT-03/05/13/14 match, SAT-04/06/10 and SAT-12 stay deferred or unknown. Go store/API and real-PostgreSQL HTTP tests, TypeScript client and Core Web tests; live acceptance is recorded with the batch. | | P | [Whitespace-only message text](official-semantics-alignment.md#whitespace-only-message-text--september-23); private `~/.parsar/remediation/20260923/campaign-scan-1/sessions/findings.json` SES-01..08 with raw records under `official/` (`s1-create-string-spaces`, `s4-create-string-newline-tab`, `s2-create-message-part-newline-tab`, `e1-events-two-whitespace-messages`, `q2-items-l100`, `p1a`, `e2`–`e4`). Rows W1–W6 of that section. Proto, dispatch, Codex, API handler, TypeScript client and real-PostgreSQL HTTP/Worker tests without a model, including Claude SDK and MiniMax Code admission rejection; live acceptance completed Codex whitespace-only Turns, and the MiniMax native refusal is recorded in the linked section. | | Y | [Artifact capture and listing](official-semantics-alignment.md#artifact-capture-and-listing--september-23); private `~/.parsar/remediation/20260923/campaign-scan-2/hosted-env/findings.json` HE-50..62 with raw records under `official/` (labels `al01`–`al09`, `ar01`–`ar04`, `ac01`–`ac05`, `ad01`/`ad02`): three owned Sessions and two tiny Turns, all deleted; first official Artifact observations. Rows A1–A4 of that section: symlink skip, republication, list envelope and malformed filter. Rust link tests, real-PostgreSQL store/HTTP and pinned-SDK tests without a model; live acceptance is recorded with the batch. | | CE | [List cursor errors](list-query-semantics.md#list-cursor-errors--september-23-2026); private `~/.parsar/remediation/20260923/campaign-scan-4/errors/findings.json` ERR-01..06 (ERR-07..13 and ERR-17 match or record upstream failures) with raw records labelled `cur-*` in `official/results.json`, plus SAT-04 (campaign scan 3) and HE-57 (campaign scan 2): owned Agent, Session, Turn, Item, Subagent, Artifact, Template, Vault, Credential, File, Skill and Skill version cursors without a model. Rows C1–C5 and K1–K3 of that section. Go store, real-PostgreSQL HTTP tenant A/B and pinned-SDK tests; independent acceptance is recorded with the batch. | diff --git a/contracts/agents-api/public-mcp-qualification.md b/contracts/agents-api/public-mcp-qualification.md deleted file mode 100644 index 83f27445b..000000000 --- a/contracts/agents-api/public-mcp-qualification.md +++ /dev/null @@ -1,91 +0,0 @@ -# Public Environment-origin MCP qualification - -This record covers explicit public HTTP MCP declarations through the pinned -official client, Core, Runtime and native Harness adapters. The -[Environment contract](environments.md#public-mcp-connection-origin) owns origin, -credential and policy semantics. Earlier Plugin-only evidence does not qualify -this public entrypoint. - -## Linux execution evidence (2026-09-30) - -Tests use real Kimi K3 through Codex 0.153.4 and Claude SDK 0.3.269, and real -MiniMax-M2.7 through MiniMax Code 0.4.12. Each successful invocation obtains a -fresh random proof from the actual MCP HTTPS server; the model must return both -anonymous and authenticated proofs. Tests use public Session, input, Turn and -Items APIs. Authentication uses an explicitly selected credential from an -attached Project Vault. The anonymous server rejects nonempty Authorization. - -| Harness | User-managed Linux | Core-managed Docker | -| --- | --- | --- | -| Codex | Anonymous/bearer, cold continuation, cancellation, tool failure | Same workflow passed | -| Claude SDK | Anonymous/bearer, cold continuation, cancellation, tool failure | Same workflow passed | -| MiniMax Code | Anonymous/bearer, cold continuation, cancellation, tool failure | Same workflow passed | - -Codex and Claude declare a named tool allowlist and required initialization. -The fixture also offers an excluded tool; native inventory/call authorization -must preserve the allowlist. Entry-point tests additionally cover an empty -allowlist, unavailable required servers and holding input until initialization. -MiniMax uses null allowlists and required=false; raw HTTP and adapter tests -reject both an empty allowlist and required=true. - -After stopping and restarting the daemon (or managed container), the same public -Session continues. Cancellation interrupts an in-flight native MCP call and -settles the Turn as cancelled. A controlled MCP isError response produces a -failed MCP Item while the model can finish its Turn. Codex exposes that failure -with output and a nullable error field; requiring error to be non-null was a -test assertion error, corrected by checking the durable public failed status. - -Raw HTTP and official SDK Session creation preserve environment origin. Negative -cases reject service-origin relocation, an unattached selected credential and -MiniMax's unsupported policies. Store tests also verify anonymous and unique -implicit credential selection without writes on rejection. Native state scans -check for the selected bearer after normal calls, cold continuation and -cancellation; it must not appear in native files. - -## Reproducible evidence - -Source revisions used during acceptance: - -- Core: 3448797d33026734d7f1f39d8b201d1af6cf7317, including the common - preparation-failure fix from main. -- Runtime executable: 04fee4e3b2a4c44f124438b03dddb1be836f38ad for Codex/Claude; - 04883f329ebebf17e35944e39b6ca854eedb061d for final MiniMax observation qualification. -- Claude bridge: cb5a42253cac16921e17e115669cdc87798a3294. -- Private Runtime protocol: 0.10.0 on both peers. - -These are development artifacts assembled from the recorded revisions, not a -published distribution. The acceptance image includes a private test CA and a -test model-network proxy. The final MiniMax fixture makes the test CA readable -by the native process; earlier root-only certificate copies did not qualify. -It is not a release image. Only owned acceptance -infrastructure was created or restarted; existing deployments and histories -were retained. - -| Evidence | Public Session ID | -| --- | --- | -| Codex self-hosted | b7300fae-6e33-4ce7-8345-d69824986c5c | -| Claude self-hosted | ce1e7e3e-37d4-439a-b017-688523fec37f | -| MiniMax self-hosted | 89a63c73-dec5-49ef-974c-1548bbb82cd1 | -| Codex managed | 23389dd4-2afb-495b-9b14-c403362f5100 | -| Claude managed | 14e5a213-cd9a-4989-89d2-1b814a653395 | -| MiniMax managed | 633a8bc6-1c8c-4674-bb8d-e47425d2abac | - -Private scripts, identities, durable Turn records, native-state checks and build -logs are retained under ~/.oac/acceptance/public-mcp-20260930 on the Linux -acceptance host. Credentials are excluded from repository evidence. - -## Limits and checks - -Real macOS/Windows model execution and E2B/microsandbox qualification are not -claimed by this batch. Service-origin workspace forwarding, public stdio, -literal headers/metadata, public functions and Subagent/MCP combinations are -outside scope. MiniMax's non-null allowlists and required initialization remain -explicitly unsupported. There is no service-network proxy or model-loop fallback. - -Focused Core/Runtime tests, 170 Claude adapter tests and the pinned official -client workflow pass. Full make check, native platform CI and independent review -results are recorded on [PR #251](https://github.com/MiniMax-AI/OpenAgentCore/pull/251). -The first local full run stopped because its new test database did not match the -required oac_*_tests naming rule. A later run passed Go/database/adapter checks -but reached an occupied browser fixture port; remaining checks use separate -ports. These interrupted attempts are not successful full gates. diff --git a/contracts/agents-api/structured-output.md b/contracts/agents-api/structured-output.md deleted file mode 100644 index b31a697fb..000000000 --- a/contracts/agents-api/structured-output.md +++ /dev/null @@ -1,125 +0,0 @@ -# Structured output - -The pinned Agents API accepts `text.format` with `type: "json_schema"` and a -`schema` object. This is not the Responses API wrapper: do not add `name`, `strict` -or an alternative public output field. The schema is saved and inherited through -the existing Agent/Session configuration resolution and immutable snapshot. - -## Qualified execution profile - -Claude SDK supports object-root schemas with `environment:none` or Core-managed -Docker `openai_hosted` or user-managed `self_hosted`, medium verbosity, `multi_agent.enabled=false` and optional -ordinary function tools returning text. Hosted execution uses the existing native -workspace tools, preparation, Files and Artifacts. The native SDK remains -responsible for its model/tool loop and schema validation. Codex and MiniMax -structured-output execution, Skills/Plugins/capability -directories (including inherited template contents), HTTP MCP, Subagent/tool-search -combinations and schemas without an explicit object root are unqualified and -explicitly rejected. E2B uses the common execution path but has not been -requalified in this change. These are implementation gaps, not a redefinition of the -official protocol. An explicit non-object root type is an official protocol error -on save and Session creation for every harness -([validation](official-semantics-alignment.md#agent-configuration-validation--september-23)). - -The Claude SDK consumes JSON numbers as binary64. Session admission rejects schema -numbers whose values cannot survive that conversion; saved Agent resources still -retain such schemas exactly. This does not claim that the upstream model/harness -preserves arbitrary numeric candidates internally. The adapter never reparses and -serializes a final answer to construct public text. - -## Core and Runtime boundary - -`ExecutionControls.OutputFormat` carries the schema through the shared execution -request. The selected service profile qualifies the operation; `structured_output` -and message observations are checked only for requests using the option. An -advertised capability alone never enables public support. Other harnesses can -implement this same request and existing Message observations without Core -engine-name branches or a second execution loop. - -The Claude adapter passes `outputFormat` to the fixed SDK. Its native -`StructuredOutput` tool is internal and is not an extra caller-defined function. -Successful live root calls are confirmed by their native tool-result receipt and -an attributed successful result containing `structured_output`. The adapter emits -a completed `final_answer` Message with the native tool-use ID and unchanged -`result.result` text. Parent assistant prose retains its own ID. Failed retries -and cancelled candidates cannot become a completed structured answer. No private -history read, output repair, schema coercion or prompt wrapper supplies the result. - -Workspace preparation additionally verifies the installed bridge's -`workspace_structured_output` feature and complete local Runtime contract. Native -inventory includes `StructuredOutput` only when output configuration requests it; -root tool identity, abort checks, filesystem and credential protections remain -unchanged. A bundle feature alone does not qualify another public combination. - -Recovery uses the existing Session/Turn/Items queries and native continuation. -SSE is still live-only. Frozen schemas apply to both initial and resumed execution; -ordinary text configuration retains its prior behavior. - -The public `text.format={type:"json_schema",schema:{...}}` is resolved with saved -Agent overrides and frozen in the existing Session configuration. Core transports -it in `ExecutionControls.OutputFormat`; it does not append prompt instructions, -validate/retry model answers, repair JSON or select native tool names. Public -admission requires the selected profile's structured-output qualification, and -only requests using this option require the Runtime's `structured_output` and -message-observation capabilities. A capability advertisement does not qualify a -new public combination. Claude advertises this operation only when the installed -SDK bridge reports its `structured_output` feature. A workspace Runtime also -requires the complete local Runtime contract and `workspace_structured_output`; -preparation checks that bundle before native launch. These remain adapter readiness -features, not new Core lifecycle or public protocol variants. - -The current qualified path is Claude SDK, `environment:none` or Core-managed -Docker `openai_hosted` or `self_hosted`, medium verbosity, single Agent, with optional ordinary -function tools and text results. The workspace uses its existing preparation and -bypass execution profile with the SDK's configured `StructuredOutput` tool added to -inventory and permission checks. Frozen schemas reach preparation before the -input handoff; Start cannot replace them. Skills, Plugins, capability directories, -HTTP MCP, Subagent/tool-discovery combinations and schemas without an explicit -object root remain unqualified; an explicit non-object root type is a protocol -error for every harness. Check resolved template contents as well as inline configuration; -ordinary text requests retain their existing qualifications. -The SDK uses binary64 JSON numbers: reject execution schemas whose numeric values -would change during that conversion, without narrowing saved Agent storage. -Codex and MiniMax structured output remain explicit execution gaps. - -The Claude adapter passes `outputFormat` to the maintained native SDK and allows -its native `StructuredOutput` terminal tool. A matching live root tool result and -an attributed successful SDK result confirm the final output. Publish the native -`result.result` string unchanged as a completed `final_answer` Message using the -native tool-use ID; the parent assistant ID can already own a prose Item. Do not -publish unvalidated retry candidates or serialize `structured_output` back to -JSON. Native retries remain harness-owned. Existing input receipts, usage, -cancellation, release and recovery rules apply unchanged. The implementation and -qualification limits are recorded in [the coverage note](structured-output.md). - -## Acceptance - -`TestNativeStructuredOutputPublicExecution` and -`services/core/tests/official_structured_output.py` exercise the pinned SDK, -raw HTTP, actual PostgreSQL/Worker/gateway/daemon and real model APIs. They require -explicit private operator options and never supply model responses. The workflow -covers a function-only random value, unchanged saved configuration, native result -application receipts, ordered terminal SSE, persisted final JSON, daemon restart -and same-history continuation, cancellation, text override and tenant isolation. - -`services/core/tests/official_hosted_structured_native.py` extends public -acceptance to an independently deployed Core, dedicated PostgreSQL and Docker -Runtime using the real Kimi API. It covers initial saved configuration and inline -prepared configuration, function-only random values, active input receipts, -native file writes consistent with final JSON, Files/Artifact reads, unchanged -terminal SSE, result retries/conflicts, cold Core/Runtime continuation, pending -cancellation without a fabricated final, ordinary text and tenant/Session isolation. -The accepted image retains the pinned SDK 0.3.269 and Claude Code 2.1.269. This -qualifies that native/provider combination, not every model or schema dialect. - -Focused tests cover native failure/retry projection, exact result bytes, schema -numeric admission, configuration transport and independent service qualification -for another harness. Running only these tests or importing the SDK is not a claim -of complete protocol compatibility. - -The retained Codex provider/tool-chain failure is not reopened by Claude's native -qualification. Its next attempt needs a concrete changed prerequisite; do not -weaken assertions, rewrite model output or repeatedly sample until one run passes. - -Self-hosted Claude structured output and cold continuation use the same workspace -adapter as managed execution. See [current qualification](environment-capabilities-qualification.md). diff --git a/contracts/agents-api/subagents.md b/contracts/agents-api/subagents.md index 7ff504ede..67bae316d 100644 --- a/contracts/agents-api/subagents.md +++ b/contracts/agents-api/subagents.md @@ -1,15 +1,10 @@ -# Subagent resources and Runtime observations +# Subagents -The target is the six read operations in the pinned SDK in `upstream.json`. -Implementation and qualification are separate: the public types and handlers do -not qualify a harness merely because it can deserialize them. The Docker V1 -workflows below have real execution evidence; they do not establish complete -multi-agent or protocol compatibility. +With `multi_agent.enabled`, a Harness may start native child agents. Core exposes them through the six Subagent read operations of the pinned SDK in [`upstream.json`](upstream.json) and records them from adapter observations. With `multi_agent.enabled=false`, the Runtime removes native child tools. [Harness capabilities](harness-capabilities.md) lists which Harness supports Subagents in which combinations. ## Public reads -All paths are under `/v1/agents/sessions/{session_id}` and require the same tenant -authentication and `OpenAI-Beta: agents=v1` as ordinary Session reads. +All paths are under `/v1/agents/sessions/{session_id}` and need the same Project authentication and `OpenAI-Beta: agents=v1` as ordinary Session reads. | Path | Result | | --- | --- | @@ -20,205 +15,45 @@ authentication and `OpenAI-Beta: agents=v1` as ordinary Session reads. | `/subagents/{subagent_id}/turns/{turn_id}` | One owned Turn | | `/subagents/{subagent_id}/turns/{turn_id}/items` | Only that child's Items in that Turn | -Lists use `after`, `limit` (default 20) and `order` (default `desc`) and return -`object: "list"`, `data`, `first_id`, `last_id` (null on an empty page) and -`has_more`. The Subagent and Subagent Turn lists reject a limit outside 1–100; -the two Item lists treat 0 as 1 and larger values as 100, like Session Items. -Cursors must belong to the requested tenant, Session, child and optional Turn. - -Child work appears only on these routes. Session Turn list and retrieve return -root Turns only; a child Turn ID there, including as a list cursor, gets the same -404 as a missing Turn. A child Turn's `agent_id` is the Session's Agent ID (a -direct child's `parent_agent_id` and the `create_subagent_call` `agent_id`), and -its `subagent_id` identifies the child; nested children use the same rule, which -is not observed officially. Root Turns have `subagent_id: null`. Session Items -remain root-owned; inherited native parent transcripts are not child work. The -Session event stream carries root work only: child Turns and child Items publish -no `agent.session.turn.*` events. `agent.session.subagent.*` events and root -coordination Items are unchanged. See [Subagent visibility](#subagent-visibility--september-23-2026). - -Active includes idle. Successful close records native time; successful reopen -preserves identity and `opened_at`, clears `closed_at`, and emits `active` once. -An already-active resume is a no-op. Turn completion, interruption and process -release do not close a Subagent. Unknown token measurements remain null. - -## Common adapter contract - -Use `internal/agentdaemon/proto/subagents.go` through the existing authenticated -Run and execution journal. No new transport, scheduler or model/tool loop exists. +Lists use `after`, `limit` (default 20) and `order` (default `desc`) and return `object: "list"`, `data`, `first_id`, `last_id` (null on an empty page) and `has_more`. The Subagent and Subagent Turn lists reject a limit outside 1–100 with `limit must be between 1 and 100`; the two Item lists treat 0 as 1 and larger values as 100, like Session Items. Cursors must belong to the requested Project, Session, child and optional Turn. + +## Subagent visibility + +Child work appears only on the Subagent routes. + +- Session `turns.list` and `turns.retrieve` return root Turns only. A child Turn ID, as a path or a list cursor, returns the same 404 as a missing Turn; Core's 404 message differs from the official one. +- The Session event stream and the creation stream carry root work only: child Turns and their Items publish no `agent.session.turn.*` or Item content events. `agent.session.subagent.*` events and root coordination Items stay. The creation stream still ends on the root's settled idle state. +- A child Turn's `agent_id` is the Session's Agent ID (a direct child's `parent_agent_id` and the `create_subagent_call` `agent_id`); `subagent_id` identifies the child. Nested children follow the same rule. Root Turns have `subagent_id: null`. +- Session Items stay root-owned; inherited native parent transcripts are not child work. Session usage sums root Turns only. + +Active includes idle. A successful close records the native time; a successful reopen keeps the identity and `opened_at`, clears `closed_at` and emits `active` once. Resuming an active Subagent does nothing. Turn completion, interruption and process release never close a Subagent. Unknown token measurements stay null. + +## Adapter contract + +Adapters report Subagent facts through `internal/agentdaemon/proto/subagents.go` over the existing authenticated Run and execution journal. There is no separate transport, scheduler or model and tool loop. | Fact | Adapter obligation | | --- | --- | -| Identity | Prove native ID, original parent and creation time; publish parents first | -| Lifecycle effect | Prove a successful close/reopen and its original time; stable effect identity across history reads | -| Child Turn | Supply native-owned ID, state and source timestamps, with known Usage only; a missing native cancellation timestamp requires a durable confirmed-effect receipt | -| Child Item | Supply an ordered complete message/tool snapshot using the existing neutral vocabulary | -| Coordination | Translate the native operation and actor/recipient identities without putting native tool names in Core | - -Core supplies public IDs and ownership from the authorized Session binding. -Identity, lifecycle, child history and live event projections commit atomically -under the existing Session lock and execution lease. Repeated observations retain -IDs and do not duplicate lifecycle events. Conflicting effects fail. A failed -coordination request to a nonexistent child preserves its opaque requested target; -it does not create a Subagent or imply that the target exists. - -Child Turns have a native writer, so their storage is separate from the Core work -queue. Session Turn reads query root Turns directly; the read-only SQL view from -migration 000051 that joins root and child Turns stays in the schema without a -public reader. Child work never becomes a second queued Core execution. Public GETs read durable -resources; they neither start native processes nor replay execution. - -An adapter freezes root output before child settlement, keeps the existing native -owner/reader alive while finite child work completes, and delivers child Items -before their terminal Turn snapshot and Run completion. Cancellation uses the -same owner and must settle child writes before release. A failed observation or -uncertain native effect cannot be converted to a successful empty history. -Codex cancellation continues under the same owner after a caller deadline; a -later call can confirm settlement without repeating the native interrupt. Actual -observation failures remain fail-closed, independently of local process cleanup. - -## Native evidence and remaining qualification - -Fixed Codex 0.153.4 real probes established original close/resume/no-op/failure -facts by correlating direct-tool output with the same call's persisted canonical -completion. Cold resume has no raw-result opt-in, so this requires native persisted -receipts. The proven profile excludes result-rewriting hooks, code-mode and -plugins. A packaged immutable Bash PreToolUse hook is allowed only with enforced -managed-only hook discovery and its exact Runtime command and matcher. Native Turn times have second precision; converting to milliseconds does -not create additional precision. A real root-first probe confirmed child file -work can finish in the same owner after the root finishes. - -Multi-agent execution with required ToolEnvironment initialization remains an -explicit combination gap until child hook-process failures are handled by the -same execution owner. The managed PreToolUse hook itself does not alter lifecycle -result receipts, but existing hook-failure handling only covers the root Turn. Ordinary single-agent environment/package -execution is unchanged. Public function and MCP tools with enabled multi-agent execution -remain unqualified. These limitations do not redefine the official protocol. - -Claude's fixed SDK uses native Agent and idle-child SendMessage calls. Original -private child records establish parentage, the first own input time and subsequent -own Turns; inherited parent context is excluded. The existing query owner admits -children before start and retains their history through settlement. Confirmed -cancellation uses a protected immutable effect receipt because native abort can -leave no terminal record. The receipt preserves the confirmed effect time across -reads without rewriting native history. This profile has no qualified close -operation, and completed or cancelled children remain active. See the -[adapter contract](../../packages/claude-sdk-adapter/README.md#subagents) for restrictions. - -MiniMax's fixed ACP supplies native delegation operations. Its Session-private -SQLite records supply original child identity, accepted inputs, terminal times -and own messages. The native task-create transaction enforces the requested -concurrent limit before start; native preparation must acknowledge that applied -limit and the protected tool profile before any model input. Discovery describes -adapter support, not proof that an arbitrary installed CLI applied these controls. -Neither adapter may substitute parent output, task completion, observation time -or an empty list for missing facts. - -## Qualified Docker workflows - -The three harnesses passed the same six GET checks with Python SDK 3.13.0 and -raw HTTP against the independent Core, dedicated PostgreSQL and colocated Runtime. -Checks include two real children with their own model output, ascending/descending -pagination, scoped cursors, root/child Item separation, Session/child Turn identity -(under the earlier contract that listed child Turns in Session Turn reads) and -cross-project denial. A new native process continued the same child without -changing old IDs, timestamps or history. Public cancellation stopped actual child -workspace writes, persisted cancelled Turns and left the Subagent active. Core -restart preserved all previously captured resources byte-for-byte after JSON -normalization. - -| Harness | Native execution | Verified optional behavior | Explicit limits | -| --- | --- | --- | --- | -| Codex 0.153.4 | Native app-server, Kimi K3 Responses | Nested children, successful close and reopen, same-child continuation | Required ToolEnvironment and public function/MCP combinations are not qualified | -| Claude Agent SDK 0.3.269 | Native Agent/SendMessage, Kimi K3 Anthropic endpoint | Foreground `oac_worker`, idle-child continuation, protected Bash | No qualified close; running-child messages, background work, alternate child profiles and per-call model overrides are rejected | -| MiniMax Code 0.4.12 | Fixed source `33b259bbbeb1c16433390869938191d09bdb0680` and recorded bounded patch, MiniMax M2.7 | Native task/task_append/task_stop, protected workspace tools | No qualified close/reopen; native workers do not delegate nested work; public function/MCP combinations remain unsupported | - -Native probes separately verified concurrency admission, child credential/history -protection and cancellation settlement under each supported profile. MiniMax's -initial child-tool isolation failure and Claude's initial input-projection failure -remain failed evidence; subsequent fixed executions supply the acceptance proof. -The MiniMax M3/Codex empty-tool-argument failure remains a separate model-profile -investigation; Core does not repair model output. - -Evidence root on `zju_a100_2`: -`~/.parsar/remediation/20260922/subagent-contract/`. Public proofs are in -`public-codex-kimi1`, `public-claude_sdk-2` and `public-mcode-1`; native mechanism -proofs are in `native-proof`, `claude-native` and `mcode-native`. The shared script -`scripts/core-subagents-acceptance.py --phase spawn-direct` validates common -reads; its `resources_passed` and `requested_phase_passed` fields qualify that -phase. Since the visibility batch, `resources_passed` also requires every named -check recorded under `visibility`. Its aggregate `passed` field additionally -requires the optional Codex close/reopen scenario. Do not require unsupported native close operations merely -to make that separate aggregate flag true. - -This batch does not rerun the E2B deployment matrix or establish live child-delta -timing equivalence. Claude and MiniMax publish verified child history at settlement; -the accepted read/recovery workflow must not be advertised as continuous native -child progress streaming. Existing single-Agent deployment evidence retains its -original scope. - -Still unconfirmed upstream semantics include root-completion child propagation, -complete child-delta ordering and Session Usage aggregation. Unlimited background -work across root Turns, complete multi-agent conformance and business Teams are -not established by these six resource reads. - -## Subagent visibility — September 23, 2026 - -The pin is unchanged: SDK 3.13.0, commit `d7c41ef`, `agents=v1`. This batch is -based on main `73ecc152`. Its plan is -`~/.parsar/remediation/20260923/subagent-visibility/PLAN.md`. The first owned -official Subagent evidence is -`~/.parsar/remediation/20260923/campaign-scan-3/subagents-tools/findings.json` -(SAT-01, 02, 07, 08, 09, with SAT-03 for the kept rejections), with raw records -under `official/`. It covers two owned -Sessions and two child Turns, all deleted. - -| Row | Case | Core behavior | Evidence (finding: request ID) | -| --- | --- | --- | --- | -| A1 | Session `turns.list` and `turns.retrieve` | Root Turns only. A child Turn ID, as a path or a list cursor, returns the same 404 as a missing Turn. Subagent Turn list/retrieve and their Item lists still serve child Turns | SAT-07: `req_4bb89ada3457444f994e7a90374d114e` (root-only list), `req_85e7eb58da8e402c8103379ff5bb11d2` (child list), `req_8a599dc455014b0398d884dfa5cc289c` (child ID 404) | -| A2 | GET events and the creation stream | No `agent.session.turn.*` event for a child Turn, including its Item and content events. `agent.session.subagent.*` events and root coordination Items stay. The creation stream still ends on the root's settled idle | SAT-09: `req_e0f7fb0ca13f4eb98b4d677be046e1da`, `req_7a68fa8c18e344cfa0ed202df92a875e` (S1 20 and S2 43 frames, no child Turn or Item event) | -| A3 | Child Turn `agent_id` | The Session's Agent ID; `subagent_id` unchanged. Nested children follow the same rule (not observed) | SAT-08: `req_85e7eb58da8e402c8103379ff5bb11d2`, `req_fc10f0d1a2e84bd086f006c01aa7ee54` | -| A4 | Subagent list envelope | `object`, `data`, `first_id`, `last_id`, `has_more`; null IDs on an empty page | SAT-01: `req_089f86e8088d441380a22de2723e6179`, `req_5f79af4eaea44cb7ab4e92920e0f88c8` | -| A5 | `limit` 0 or above 100 | Subagent Item and Subagent Turn Item lists clamp to 1 and 100. The Subagent and Subagent Turn lists keep rejecting with `limit must be between 1 and 100` | SAT-02: `req_6179ae6c1d1640d899ee4798e7f9fa57`, `req_7436104afbae4e73a0eb43b00ec9e660`, `req_32899313414b4031849a22cd2927f0ad`; SAT-03 rejections: `req_7df58d9579be4ee3ab7fdab55286aa05`, `req_b4321de4480c4a8e96b9ea285ff63a46`, `req_0f437ad4713d47a8af1f61a88636bf79`, `req_e9d476dd2a69472694cffc0851d0574c` | - -### Decisions - -- Session Turn reads use a new root-only query instead of changing the view, so no - migration is needed. Child data is not deleted or rewritten. -- The Session event log has one reader: the public GET and creation streams. - Creation-stream settlement reads the settled idle and the latest root Turn, and - Session usage sums root Turns, so neither used child Turn events. The Core Web - timeline had no Subagent view; it only showed child Turns as ordinary Turn rows - and now keeps its timeline root-only even against an earlier Core, hiding the - Items of Subagent Turns that Core listed or streamed. Recovery - continues through Session, Turn and Item reads plus the Subagent routes. No - internal signal had to be kept. -- The scan recorded child Item events as already absent. They were not: every - child Item recorded `turn.item.*` and content events with the child Turn ID. - They are removed with the child Turn events, since the official parent stream - carried neither. -- The official 404 message for a child Turn ID (`No managed agent resource found: - …`) and the Subagent 404 messages (SAT-06) differ from Core's local text. Only - the status, error fields and "same as missing" behavior are aligned here. -- Core may expose documented extensions beyond the official API. This batch - removes only the mixed Session Turn pages and child Session events, which were - not documented extensions. Root-only extensions such as - `agent.output.command_execution_output.delta` and the Web's handling of older - Core releases are unchanged or additive. - -Unchanged: Subagent retrieve fields and statuses, child history contents (SAT-12 -remains unknown; Core keeps the child input Item), the hidden task text, the -single `subagent.created` emission, cursor error semantics (SAT-04, HE-57, since -aligned by the [list cursor error batch](list-query-semantics.md#list-cursor-errors--september-23-2026)), native -history ownership, cancellation, cold continuation and tenant isolation. - -### Acceptance boundary - -Go store and API tests cover each row, including tenant isolation. A real-PostgreSQL -HTTP test also checks creation-stream settlement with Subagent facts. TypeScript -client and Core Web unit tests cover the official child Turn shape and the -root-only timeline. `scripts/core-subagents-acceptance.py` records A1–A5 -as named `visibility` checks per phase, including the observed stream. Its -inspect phase, run against a controlled local fixture without a model or stream, -reported all six read differences on baseline main and passed on this branch. Live model acceptance and the server gate are recorded with the -batch when complete. +| Identity | Prove the native ID, original parent and creation time; publish parents first | +| Lifecycle effect | Prove a successful close or reopen and its original time; keep the effect identity stable across history reads | +| Child Turn | Supply the native-owned ID, state and source timestamps, with known Usage only; a missing native cancellation timestamp requires a durable confirmed-effect receipt | +| Child Item | Supply an ordered, complete message or tool snapshot in the neutral vocabulary | +| Coordination | Translate the native operation and actor and recipient identities without putting native tool names in Core | + +Core assigns public IDs and ownership from the authorized Session binding. Identity, lifecycle, child history and live event projections commit atomically under the Session lock and execution lease. Repeated observations keep their IDs and never duplicate lifecycle events; conflicting effects fail. A failed coordination request to a nonexistent child keeps its opaque requested target and creates no Subagent. + +Child Turns have a native writer, so they are stored apart from Core's work queue and never become a second queued execution. Session Turn reads query root Turns directly. Public GETs read durable resources; they never start native processes or replay execution. + +An adapter freezes root output before child settlement, keeps its native owner and reader alive while finite child work completes, and delivers child Items before their terminal Turn snapshot and the Run's completion. Cancellation uses the same owner and settles child writes before release. A failed observation or uncertain native effect never becomes a successful empty history; parent output, task completion, observation time or an empty list never substitutes for a missing fact. + +## Native profiles + +[Harness capabilities](harness-capabilities.md#tools) lists rejected tool combinations. + +**Codex.** The adapter enables the native `multi_agent` feature with a nesting depth of 64 and maps the concurrency limit to `agents.max_threads`. It disables native hooks, plugins, code mode and `multi_agent_v2`, and refuses to start if the native hook list is not empty or managed requirements force a conflicting feature. Close and reopen facts come from direct tool output correlated with the same call's persisted completion, so they need native persisted receipts. Native Turn times have second precision. Child file work can finish under the same owner after the root Turn finishes. Cancellation continues under the same owner after a caller deadline; a later call can confirm settlement without repeating the native interrupt. + +**Claude SDK.** The pinned SDK's native Agent and SendMessage calls run the single child type `oac_worker`, which inherits the model and has workspace Bash, Agent and SendMessage; the bridge's `subagent_resources` feature gates it. Children use native Bash with the parent's launching-user permissions. Private child records establish parentage, the first own input time and later own Turns; inherited parent context is excluded. The query owner admits children before start and keeps their history through settlement. Confirmed cancellation writes an immutable effect receipt because a native abort can leave no terminal record. There is no close operation: completed or cancelled children stay active. Messages to running children, background work, other child profiles and per-call model overrides are rejected. The [Claude SDK adapter](../../packages/claude-sdk-adapter/README.md#subagents) documents the details. + +**MiniMax Code.** Native ACP delegation (`task`, `task_append`, `task_stop`) creates children. Session-private SQLite records supply child identity, accepted inputs, terminal times and own messages. The native task-create transaction enforces the concurrency limit before start, and native preparation must acknowledge that limit and the restricted tool profile before any model input. There is no close or reopen, and native workers do not delegate nested work. + +Claude and MiniMax publish verified child history at settlement, not as continuous child progress. Root-completion propagation to children, complete child-delta ordering and unbounded background work across root Turns are not established. diff --git a/contracts/agents-api/template-null-selection.md b/contracts/agents-api/template-null-selection.md deleted file mode 100644 index 39d79de99..000000000 --- a/contracts/agents-api/template-null-selection.md +++ /dev/null @@ -1,112 +0,0 @@ -# Template-reference null selection - -This batch retains SDK 3.13.0, upstream `d7c41efee1b0802b79f3f88a678ef2052b06e9ce` and `agents=v1`. It changes shared Core configuration selection only; Provider and Runtime receive the existing frozen effective configuration. - -## Qualified selection rules - -| Operation / field | Omitted | Null | Non-null | -| --- | --- | --- | --- | -| Referencing Session network | Inherit complete policy | Inherit complete policy | Must narrow template authority | -| Referencing Session Skills, Plugins, capability directories | Inherit list | Inherit list | Replace list; empty clears | -| Inline hosted Session network | Enabled default | Enabled default | Existing policy validation | -| Template resource update network | Preserve | Reset to enabled | Replace saved policy | -| Template resource update capability lists | Preserve | Clear | Replace saved list | - -Resolve only the selected tenant-owned sources, using the existing encrypted initialization transaction. Keep unresolved caller intent distinct from the frozen snapshot. Template mutation/deletion cannot alter existing Sessions or same-intent creation retries. No new installer, native path, harness branch or compatibility reader is introduced. - -Official Session capability-directory responses can include automatically derived Skill/Plugin installation directories in addition to caller-selected paths. Core continues to return caller paths and keeps Runtime-owned installation locations private. The selection matrix does not qualify complete public directory projection parity; copying official internal paths would not establish usable Core paths. - -## Evidence and limits - -Owned official probes and Core acceptance records are under -`~/.parsar/remediation/20260923/template-null-selection/`. Together they made -82 HTTP requests and created six Templates and eleven Sessions. All seventeen -owned resources have successful public DELETE receipts; physical upstream -destruction was not independently observed. Credential scans passed. - -- `official-network/REPORT.md`: 46 requests, four Templates, six Sessions, zero - Turns. Disabled and populated restricted policies distinguish inheritance from - an enabled default. Narrowing succeeds; broadening rejects. Inline null and - Template create/update null produce enabled policy. Existing Sessions preserve - their policies after Template reset. Four provisioning-time DELETE conflicts - each succeeded on one later bounded cleanup attempt. This does not requalify - native network enforcement. -- `official-capabilities/REPORT.md`: 36 requests, two Templates, five Sessions, - one real official `gpt-6-astra` Turn. Four distinct request cases establish - list selection. The null case additionally has three completed, exit-zero - native commands reading unknown markers from inherited Skill, Plugin Skill - and caller directory contents. Empty/replacement native bytes were not newly - qualified against the official service. - -The first capability attempt stopped before any model call because its test -incorrectly equated caller directories with the entire official Session projection. -Its original script, raw results and cleanup remain. The bounded continuation -separated caller-selected paths from observed derived paths; it did not change -the requests or hide a model failure. Official Environment reads contain Skill -and Plugin metadata but no capability-directory field; do not equate them with -the richer Session environment projection. Mixed individual-field official null -cases and all native/provider combinations remain unqualified. - -## Core verification - -`TestTemplateNullSelectionOfficialClientPostgres` and -`official_template_null_selection.py` exercise the actual HTTP handler, isolated -PostgreSQL, pinned strict SDK and raw HTTP. Eleven successful selections cover -combined and mixed fields, disabled/restricted inheritance, narrowing, inline null, -Template resets and exclusion of an unselected foreign Skill. Eight rejected -creations leave no Session, Environment or initialization residue. Foreign and -missing template lookups remain indistinguishable; direct foreign snapshot reads -reject. Encrypted archive digests and unrelated initial file/env contents remain -frozen after source/default/template mutation and deletion and reopening the Store. -Changed retry intent conflicts. This test does not execute a native model. - -The first acceptance assertion incorrectly expected Session-only directory/network -fields on the Environment resource. The one-line test correction uses its pinned -projection; the original failure remains. Independent PostgreSQL acceptance passed -in 1.74 seconds. The complete server gate then passed at `65f9898`, including -228.891 seconds of Store tests, sqlc byte comparison, Go/build/native package and -Rust checks. No database query or migration changed. - -Fresh `make check-web` passed with Node22 and pinned pnpm10.30.3: 63 doctor, -287 client, 583 Web unit and 76 browser cases. Its initial dependency setup used -the app's pnpm11 fallback and stopped before tests; the tool-generated workspace -placeholder was removed and the unchanged repository was tested with its pinned -toolchain. No dependency or build-policy change was made. `make openapi` and -`git diff --check` passed. Optional 512 MiB/source streaming and packaged MiniMax -scratch/large-output live profiles were not enabled. - -The first Codex/Docker attempt initialized three Sessions. Its Kimi null-selection -Turn read the three inherited capability markers through native commands. The next -empty-selection Turn failed on provider HTTP 429 before a native command or reported -usage; replacement and restart checks were not submitted. These records are retained. - -An OpenAI-provider continuation was stopped when the user clarified that its key -was restricted to official Agents API probes. One submitted Core Turn was cancelled; -its stored usage was 10,576 input and 61 output tokens. This attempt is not acceptance. -Its owned resources and temporary credential copies were removed. Subsequent Core -model verification uses only the authorized Kimi or MiniMax providers. - -MiniMax-M3 completed the empty and replacement Turns. Each produced one completed, -exit-zero native command whose unmodified JSON output exactly matched the selected -capabilities, excluded markers, inherited confidential env and single setup trace. -The original runner still failed because the replacement assistant answer omitted -one trailing newline from that trace. Both native proofs and the mismatching answer -remain preserved; Core and model output were not changed. Initialization acceptance -uses the actual command output, with the assistant answer retained as an observation. - -The separate one-Session null continuation passed two MiniMax-M3 Turns. After the -first native proof, the test changed the source Skill version and Template, deleted -both resources, restarted Core, and retried the original creation key. The retry -returned the same Session; its second native command read the original markers, -confidential env and unchanged one-line setup trace. Initialization was not replayed. - -All Core native execution used production source `f86223f`; the final server gate -adds only the corrected resource-test assertion and documentation. The null -continuation reused the verified Core/daemon binaries and base image. Rebuilding -the deleted wrapper image changed its image ID; base identity, in-image daemon -hash, Dockerfile and runtime configuration established the same content provenance. -No native feature, model response or production behavior was changed for acceptance. -Reports in `live/attempt-minimax/` and `live/attempt-minimax-null/` retain the original -failures, command proofs, source hashes and owned-resource cleanup records. - -This batch does not widen hostname syntax, change native capability support, requalify every harness/Provider combination, or claim full protocol compatibility. Provisioning-delete retries and private test assertion failures remain in the evidence rather than being counted as successful first attempts. diff --git a/contracts/agents-api/tool-policy.md b/contracts/agents-api/tool-policy.md deleted file mode 100644 index 8d0ce8ae2..000000000 --- a/contracts/agents-api/tool-policy.md +++ /dev/null @@ -1,81 +0,0 @@ -# Explicit disabled tools - -Protocol baseline: `upstream.json` (OpenAI Python SDK 3.13.0). This covers a -bounded execution profile, not complete tool or Agents API compatibility. - -## Public behavior - -Saved Agents and inline Session configuration accept these declarations: - -```json -[ - {"type": "web_search", "mode": "disabled"}, - {"type": "programmatic_tool_calling", "enabled": false} -] -``` - -Saved Agents also keep every other pinned `web_search` mode, as the official -service does: omitted/null mode is saved as `live`, and `cached` and `live` are -saved as sent ([TV-05](official-semantics-alignment.md#saved-web_search-modes--september-23)). -Search settings are resource data in every mode: omitted/null context size -resolves to `medium`; domain and location omission resolves to null; an empty -domain list remains empty; a supplied location, including `{}`, includes `city`, -`country`, `region` and `timezone`, with null for omitted keys (observed -officially: `req_db41d2f6261b4abfb69465eafe719ab5` and `req_165d53b88445490b9146d8272c54134d`). These settings cannot -enable execution. Session admission resolves saved references and inline -declarations with the execution parser into the immutable Session snapshot. -Only explicit `disabled` search is qualified: enabled or omitted-mode search, -inline or saved, rejects with `unsupported_or_invalid_configuration` before any -write unless the Session replaces the saved tools. Protocol errors in these -declarations, such as an unsupported `mode` or `context_size`, a non-boolean -`enabled` or a repeated `web_search`, use the official fields -([validation](official-semantics-alignment.md#agent-configuration-validation--september-23)). -Explicit enabled programmatic execution likewise saves and rejects at Session -admission; saving intent remains separate from execution qualification. - -Omitting programmatic configuration preserves each harness's native behavior. -The user approved this difference from the official default-on behavior. Native -feature differences stay in adapters; Core does not supply another executor or -model loop. Unrelated native utility tools are not implicitly removed. - -## Runtime boundary - -The common `DisableProgrammaticToolCalling` control carries explicit disabled -intent on initial execution and cold continuation. Public qualification and the -operation-specific Runtime capability must both permit this request. Omission -does not require the new capability. Search uses the existing disabled control. - -| Adapter | Native enforcement | -| --- | --- | -| Codex | Disable code-mode features; check native managed requirements before thread start/resume and reject a forced conflicting feature. | -| Claude Code | Retain the restricted built-in inventory and verify native initialization against that inventory. | -| MiniMax Code | Retain the protected native tool profile, empty text-execution inventory and disabled web-search feature. | - -## Acceptance - -The opt-in `TestNativeToolPolicyPublicExecution` fixture and -`services/core/tests/official_tool_policy.py` exercise a real PostgreSQL -database, independent API, Docker daemon and native harness using the pinned -official SDK with strict response validation and raw HTTP. Supply the existing -native-test environment variables plus `OAC_TEST_TOOL_POLICY_ENGINE` and a private -`OAC_TEST_TOOL_POLICY_REAL_OPTIONS` file. Each concurrent execution worker requires -its own dedicated test database. - -On 2026-09-22, Codex and Claude with Kimi K3, and MiniMax Code with MiniMax-M2.7, -passed SDK/raw HTTP through both saved and inline configurations. Each of four -Sessions completed a first Turn and continued its random marker after a full -daemon restart with the same native Session. Checks include public configuration, -SSE order, Items, native input receipts, unsupported enablement without Session -persistence, omitted configuration admission and tenant isolation. - -Native evidence separately records Codex's disabled feature arguments and search -setting, Claude's empty built-in tool argument and initialization validation, and -MiniMax's restricted native configuration. Successful model text alone does not -prove disabled tools. Controlled tests cover managed requirement conflicts, -invalid fields and the common qualification contract. - -Evidence is retained privately under `~/.parsar/remediation/20260922/tool-policy` -on the development host and test server. Live qualification in this batch uses -`environment:none`; workspace provisioning, provider matrices, enabled tools, -native question extensions and complete hosted default/error semantics were not -requalified. Existing workspace mechanisms are unchanged. diff --git a/contracts/agents-api/tool-search.md b/contracts/agents-api/tool-search.md deleted file mode 100644 index 2d1a8c8b0..000000000 --- a/contracts/agents-api/tool-search.md +++ /dev/null @@ -1,107 +0,0 @@ -# Deferred function discovery - -Baseline: Python SDK 3.13.0, upstream -`d7c41efee1b0802b79f3f88a678ef2052b06e9ce`, Beta `agents=v1`. -This is bounded execution coverage, not complete Agents API compatibility. - -## Contract and boundary - -Saved and inline tools preserve the full function definition and `defer_loading` -(default false). The pinned Agents `tool_search` configuration has only `type`; -Responses-specific execution/parameter fields are not accepted here. The saved-Agent `PersistedAgentTool` union retains `tool_search`, while the -Session `AgentTool` response union omits it. Session/SSE resource projection follows -that pinned distinction; the full frozen configuration still includes it. Exact -hosted response behavior is unverified. The shared -Runtime carries search intent and each deferral flag. Native search and lazy schema -loading belong to the harness adapter. Core retains the frozen definitions and -existing public function calls, results and application receipts. The pinned Items -union contains no tool-search Item; do not invent one. - -The implementation is Claude SDK 0.3.269 / native 2.1.269, single Agent, -`environment:none` or a managed/user-owned workspace, medium verbosity, object-root -function schemas and text results. Workspace tool discovery excludes Skills, -Plugins and local capability directories; its ordinary native workspace tools remain available. -It supports a mixture of eager and deferred application functions with text or -previously qualified inline PNG/JPEG message inputs. The native MCP -server marks eager definitions `anthropic/alwaysLoad:true`; deferred definitions use -false and the adapter explicitly enables native ToolSearch. The function profile allows only declared callbacks and ToolSearch in addition -to the selected workspace tools. No search index, callback protocol, provider proxy or -model loop is added to production. - -Search-only, missing-search, HTTP MCP, structured-output and Subagent -combinations remain unqualified. A repeated `tool_search` is an official protocol -error on saved and inline configuration. They are implementation/verification -gaps, not claimed upstream restrictions. Codex and MiniMax discovery remain gaps. -Unknown public combinations reject before execution; an actual Runtime must also -advertise the operation. An advertisement alone cannot qualify a public profile. - -Public `tool_search` and function `defer_loading` are shared Runtime intent. -Core preserves complete immutable definitions, sends `PromptRequestPayload.ToolSearch` -and each `FunctionTool.DeferLoading`, and requires the operation's existing profile -qualification plus the Runtime `tool_search` capability. It never performs native -search, selects native names, interprets provider policies or implements another -model/tool loop. Additional harnesses implement the same intent in their adapters. -The pinned Session AgentTool response union excludes the tool_search input member; -project it out of Session/SSE resources while preserving saved and frozen input. - -Workspace discovery uses the same native ToolSearch and callback path. The -installed bridge must additionally advertise `workspace_tool_search` and the -complete local Runtime/function contract. The daemon derives its ordinary -`tool_search` capability from these verified bridge features; Core does not branch -on environment ownership. Existing callback receipts, cancellation and native cold -continuation retain their semantics. See [current workspace qualification](environment-capabilities-qualification.md). - -## Evidence and limitations - -Native feasibility passed with Kimi `kimi-k3`: initial provider requests excluded -the deferred target schema; after a real native ToolSearch call that schema became -available, while an unrelated deferred schema stayed unloaded. The model used a -random required argument available only in that schema and returned the exact fresh -callback result. A new native process resumed the same native Session and called -it again with a new callback result. Existing discovered definitions remained in -that native history. Evidence: `~/.parsar/remediation/20260922/deferred-tools/claude-native-1790052750`. - -A first test proxy omitted response Content-Encoding and failed after discovery; -the failed evidence is retained. The corrected test forwards unchanged real provider -responses. The proxy is test observation only and is not part of Runtime. - -Codex source exposes deferred dynamic functions, but the available Kimi Responses -probe rejected native `tool_search`; the MiniMax sample did not produce discovery. -Neither establishes Codex qualification. These observations do not prove that all -models from either provider lack the feature. - -The native harness owns dynamic model/provider policy. Known conflicting modes and -beta/search settings reject in the adapter. Its SDK has no reliable pre-input signal -proving actual deferral after opaque policy changes; init tool inventory is insufficient. -That detection gap and other providers/models remain unverified. No endpoint/model -allowlist or copied native policy evaluator is added to imply a stronger guarantee. - -Public-chain real acceptance passed on 2026-09-22 with Kimi `kimi-k3`: -`tool-search-public-2315339108`, 97.83 seconds. The initial strict-SDK attempt -failed on the Session response union before any model request; that failure is -retained as `public-first.log`. Resource projection follows the fixed schema, -without relaxing SDK validation. Run `TestNativeToolSearchPublicExecution` -with the fixed official SDK, a real private model configuration, a real daemon and -PostgreSQL. It checks saved/inline configuration, mixed functions, result retries, -SSE/Items, cancellation, cold daemon continuation and tenant isolation. Independent review is required before release qualification. - -Shared Go race checks and 132 Claude SDK checks passed. The required browser gate -passed 73 cases locally with Node22 and real Chrome. The server has no browser; -the remaining `make check` targets passed on zju against a dedicated PostgreSQL. -No database query/schema changed, so explicit sqlc generation is not applicable; -the full gate still verifies generated queries. - -The image/discovery combination also passed through the same public chain in -`tool-search-public-2039155387` (`public-image-isolated-tests.log`). A random PNG -was submitted with the opening prompt and a different random PNG was submitted -while the deferred function waited. The answer identified the latest image's band -order and the fresh callback result; cancellation and cold continuation passed in -the same run. The existing PNG generator is shared with the image-input regression. -No new image admission restrictions or native lifecycle were needed. - -The full server gate is `make -o check-web check`, with `make check-web` completed -locally on the same changes. Packaging initially caught an outdated expected Runtime -feature list; it was updated and the full server gate rerun successfully. Image -acceptance uses a separate dedicated database from the full gate. Failed setup -attempts (execution-owner lock and test-database naming guard) are retained; neither -reached model execution. diff --git a/contracts/agents-api/user-managed-runtime-v1.md b/contracts/agents-api/user-managed-runtime-v1.md deleted file mode 100644 index 3e8841ab1..000000000 --- a/contracts/agents-api/user-managed-runtime-v1.md +++ /dev/null @@ -1,122 +0,0 @@ -# User-managed V1 Runtime qualification - -The V1 executor is the OpenAgentCore daemon, colocated with the selected native harness, -local tools and workspace. Core runs separately with its own PostgreSQL database. -A caller creates a public `self_hosted` Session and starts Runtime with its exact -Environment ID, returned `remote_url` and scoped executor credential. This private -transport does not interoperate with stock Codex `exec-server` or Noise. - -For `openai_hosted`, Core uses the deployment-selected -[E2B, Docker or microsandbox provider](sandbox-deployment.md). -This document qualifies the separate caller-managed path: Docker and E2B use the -same Runtime contract, but the application owns their compute. In that path, E2B -create, information, renewal and deletion use the official E2B SDK outside Core. Public execution and -file operations continue through Core and daemon, not E2B commands or files. - -## Accepted scope - -Historical qualification recorded on 2026-09-22 against source candidate `8f0cd2530d7b58cb7fb3ea124a1fc43dacecca36`. -These records qualify only those binaries and tested inputs. Later implementation -changes, including shared Runtime capability preparation and native-platform -support, require separate acceptance. The current daemon has no inner sandbox; -private-file denial and namespace probes below describe the former tested -implementation, not current tool permissions or an installation requirement. - -| Deployment | Codex | Claude Code | MiniMax Code | -| --- | --- | --- | --- | -| User-managed Linux amd64 Docker | Accepted | Accepted | Accepted | -| User-managed E2B, immutable Runtime template | Accepted | Accepted | Accepted | -| Real model | Kimi K3, Responses | Kimi K3, Anthropic-compatible API | MiniMax M2.7, Anthropic-compatible API | - -Each complete deployment run uses official OpenAI Python SDK 3.13.0 at the pinned -upstream commit plus raw HTTP and live SSE. The common four-Turn acceptance checks: - -- Public inline Files.create bytes, native command execution and two actual output - files; Files.list path/order/pagination and immutable Artifact downloads through - SDK/raw HTTP, including foreign-tenant denial. -- Daemon and Core restart with unchanged committed Items, preserved conversation - history and outputs, and no repeated publication effects. -- Cancellation of a foreground native process, stopped file effects, idempotent - repeat cancellation and continued execution without restarting cancelled work. -- Native tools cannot read executor credentials/private staging witnesses or - inherit known private credential values. Public responses contain no known - private credentials. Private witnesses remain unchanged across restart. - -E2B additionally executes the native isolation probe through a fifth real model -Turn. It checks a witness in the actual native-history directory, the executor -key, staging and protected startup receipt, PID-namespace separation, inaccessible -outer process secrets, and denied envd/sudo/privileged-account access. Official -E2B SDK 2.51.0 receipts record creation, metadata-bound inspection, lease renewal -and explicit kill of each owned sandbox. No replacement Runtime is called recovery. - -The separate real Docker credential lifecycle check covers caller/executor role -separation, principal/tenant/Environment binding, rejection of history rebinding, -key rotation fencing the old socket while retaining device identity, revocation, -and Session deletion without reclaiming caller-owned compute. These shared Core -checks are not claimed as a separate live rotation/revocation run on every E2B -profile. User-managed Sessions create no managed Runtime allocation. - -## Evidence and verification boundaries - -Evidence is retained on `zju_a100_2` under -`~/.parsar/remediation/20260921/self-hosted-onboarding/`: - -- `source-candidate.json` and `source-verified.json`: exact source and binary hashes. -- `live/{codex,claude,mcode}/accepted.json`: complete Docker four-Turn runs before - removal of unused private wire fields. `live/codex/security.json` records the shared - real credential lifecycle checks. -- `live/final-{codex,claude,mcode}/final-smoke.json`: final private wire 0.3.0 - binaries, real native command and Files/Artifacts/tenant readback. This is a - focused regression, not another complete Docker lifecycle run. -- `live-e2b/{codex-attempt3,claude-attempt1,mcode-attempt1}/`: final five-Turn - E2B acceptance and cleanup (208.562s, 147.053s and 125.647s respectively). -- `e2b-builds/final-daemon/`: immutable template receipts and verified daemon hash. -- `make-check-final.log`: full gate at `ce11501`, including dedicated real - PostgreSQL, fixed official client, all service/daemon packages, sqlc regeneration, - independent builds, Claude/MiniMax packaging and retained Rust helper checks. - The subsequent connection-cache cleanup fix at `8f0cd25` passed the full - execution-package regression suite and independent Core build. -- `review-final.json`: independent Astra high review of the complete diff, - performed with reused context under explicit user authorization. It is not a - fresh-context blind review. - -The final Docker smoke runner initially called a test helper with the wrong -signature after native execution. A read-only follow-up completed the assertions -against those exact Turns without model replay. Original failed logs are retained. -The first E2B Codex launch had an uncertain network result and was reconciled by -metadata without retrying that Create. The second run stopped on an operator -readiness-wrapper error after restart. Its owned VM was removed before the fresh -third run; the failed records remain failures, not accepted execution evidence. - -## Limits - -Qualification is bounded to these immutable Linux amd64 Runtime builds and their -recorded real models. It does not establish arbitrary-host isolation, high -availability, full Agents API conformance, Anthropic-model acceptance for Claude, -or identical optional capabilities across harnesses. - -These runs received their model through the operator options file, which is now -retired. A self-hosted Session now carries its own model provider, from the request -or a saved Agent, frozen and delivered over the same daemon transport; see -[model execution](model-execution.md). Deployment default model providers never -reach self-hosted executors. - -The historical runs above used `/workspace` and empty capability directories. -Current self-hosted input accepts a canonical absolute workspace and optional local -capability directories, prepared through the -[shared Runtime snapshot](environments.md#runtime-capability-preparation). That added -scope is not qualified by these earlier runs. Service-origin HTTP MCP on self-hosted -remains explicitly rejected; -none-environment HTTP MCP and separately qualified hosted Plugin MCP keep their own -scope. Environment Templates remain hosted-only. Unspecified upstream defaults, -errors, lifecycle edge cases and broader resource semantics remain in the protocol -coverage ledger. - -Runtime workspace and native history must survive restart. Expiry or loss of that -state cannot authorize silent replacement or replay. After a permanent executor -rejection the daemon parks instead of exiting, so the Docker Runtime keeps its -`unless-stopped` restart after reboot without a restart loop; rerunning the -installer with the rotated credential resumes the same container and history -([executor credentials](environment-executor-credentials.md#revoked-or-rotated-credential)). A successful disconnect, -revocation or Session deletion is not a guarantee that every native effect has -stopped; the compute owner remains responsible for termination and cleanup. diff --git a/contracts/agents-api/workspace-placement.md b/contracts/agents-api/workspace-placement.md deleted file mode 100644 index ed40579bf..000000000 --- a/contracts/agents-api/workspace-placement.md +++ /dev/null @@ -1,40 +0,0 @@ -# Workspace placement - -The daemon, selected harness, native tools and workspace are colocated in one Runtime. -Our daemon is the user-side executor for `self_hosted`. Enrollment freezes the exact -Session/Environment/device/key binding; it creates no managed allocation and cannot -move a Session to another device. Core-managed hosting uses the deployment-selected E2B, Docker or microsandbox -provider. Users manage -local or E2B Runtime creation, renewal and destruction through the official SDK. - -All three harnesses reuse typed `LocalEnvironment`, existing preparation/start/ -cancel ownership and authorized local Files/Artifacts. Native tools use the -starting account's permissions and may access that account's Runtime credentials -and history. The daemon adds no inner sandbox; managed isolation belongs to the -outer Environment. Strict resume requires retained history; connectivity alone -establishes neither readiness nor isolation. Current -implementation is recorded in the [Environment profile](environments.md#initial-public-self-hosted-profile); -[real qualification](user-managed-runtime-v1.md) identifies accepted deployments and limits. - -The pinned public Environment resources and `remote_url` remain unchanged, while -that URL names our private daemon transport, not stock `exec-server`. -Service-origin HTTP MCP is rejected on `self_hosted`; qualified `none` MCP and -hosted Template Plugin MCP retain their separate boundaries. - -## Contract owners - -- [Environments](environments.md#ownership-and-placement-decision) owns placement, - preparation and authorization. -- [Core–Runtime protocol](../../docs/runtime-protocol.md) owns workspace operation - messages, bounds and operation lifetimes. -- [Environment Files](environment-files.md) records the public file behavior and - its qualification evidence. -- [Harness onboarding](harness-onboarding.md) owns native adapter integration. - -## Recorded output limitation - -The earlier Codex placement assessment recorded `NATIVE-COMMAND-OUTPUT-001`, -a retained-command-output gap. That failed run is historical evidence, not a -current installation prerequisite or a claim that subsequent implementations -fixed output fidelity. Current behavior must be qualified through the shared -Runtime path; do not fabricate missing output. diff --git a/docs/api/public-agent-api.md b/docs/api/public-agent-api.md index df311a662..e94221904 100644 --- a/docs/api/public-agent-api.md +++ b/docs/api/public-agent-api.md @@ -532,13 +532,13 @@ template = client.beta.agents.environments.templates.create( | Field | Meaning | | --- | --- | -| `network` | `access`: `enabled` (default), `disabled`, or `restricted` to 1–100 exact hosts in `allowed_domains`. A Session can only narrow it. Current Runtimes don't enforce `disabled` or `restricted`, so a Session that needs them is rejected; see [restricted network policy](../../contracts/agents-api/environment-templates.md#restricted-network-policy) | +| `network` | `access`: `enabled` (default), `disabled`, or `restricted` to 1–100 exact hosts in `allowed_domains`. A Session can only narrow it. Current Runtimes don't enforce `disabled` or `restricted`, so a Session that needs them is rejected; see [restricted network policy](../../contracts/agents-api/environments.md#restricted-network) | | `packages` | `npm` and `python` packages. `system` packages are rejected: preinstall them in the image or on the machine | | `setup_commands`, `env` | Run and set at preparation. Never returned by reads | | `files`, `skills`, `plugins` | Initial content. Up to 50 files, 10 MiB inline in total | A Session freezes the template when it starts. Details: -[Environment Templates](../../contracts/agents-api/environment-templates.md). +[Environment Templates](../../contracts/agents-api/environments.md#templates). ## Vaults diff --git a/docs/architecture.md b/docs/architecture.md index 83355612d..dcb19e699 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -64,7 +64,7 @@ The numbers below match the overview. Each protocol defines behavior, ownership, The [bootstrap contract](runtime-bootstrap.md) carries the Runtime's startup input across the provisioning boundary. After connection, capability preparation belongs to Runtime; the Provider does not become a second execution path. -Not every combination of Harness, model and Environment works. The supported ones are recorded in [Harness selection](../contracts/agents-api/harness-selection.md) and the [coverage record](../contracts/agents-api/README.md). +Not every combination of Harness, model and Environment works. The supported ones are recorded in [Harness capabilities](../contracts/agents-api/harness-capabilities.md) and the [coverage record](../contracts/agents-api/README.md). ## A Session, end to end diff --git a/docs/development.md b/docs/development.md index f50b9c75f..1fb0ba1a7 100644 --- a/docs/development.md +++ b/docs/development.md @@ -121,7 +121,7 @@ Implement the shared `ExecutorFactory`, `Executor` and `Turn` interfaces in [`agent/harness.go`](../apps/daemon/internal/agent/harness.go), register the adapter and add its profile/configuration entry to the shared catalog. Follow the numbered steps in [Harness onboarding](../contracts/agents-api/harness-onboarding.md); qualification -evidence belongs in [Harness integration](../contracts/agents-api/harnesses.md). +evidence belongs in [Harness capabilities](../contracts/agents-api/harness-capabilities.md). ### Add a Sandbox Provider diff --git a/docs/user-guide.md b/docs/user-guide.md index 0f3351213..4548a1f1b 100644 --- a/docs/user-guide.md +++ b/docs/user-guide.md @@ -40,7 +40,7 @@ The harness is the agent program that runs your Session. Set it in settings; see [native model configuration](../contracts/agents-api/harness-onboarding.md#native-model-configuration). Selecting a harness doesn't make an unsupported model or operation work; see -[harness selection](../contracts/agents-api/harness-selection.md). +[harness selection](../contracts/agents-api/model-execution.md#harness-selection). ### Which model provider a Session uses diff --git a/packages/agents-client/README.md b/packages/agents-client/README.md index 4f81629af..e8564d06a 100644 --- a/packages/agents-client/README.md +++ b/packages/agents-client/README.md @@ -62,7 +62,7 @@ await client.updateAgent(agent.id, { model: "another-model" }); await client.updateAgent(agent.id, { x_agents_core: { model_provider: null } }); ``` -Reads return `ModelProviderView`, which has `api_key_configured` and never `api_key`; writes take `ModelProviderInput`, so a read cannot be resubmitted as an update. [Model execution](../../contracts/agents-api/model-execution.md#saved-defaults-and-precedence) defines what omission and `null` mean on each field, which provider a Session uses and which protocols each harness accepts. [Harness selection](../../contracts/agents-api/harness-selection.md) defines the `harness` field. +Reads return `ModelProviderView`, which has `api_key_configured` and never `api_key`; writes take `ModelProviderInput`, so a read cannot be resubmitted as an update. [Model execution](../../contracts/agents-api/model-execution.md#saved-defaults-and-precedence) defines what omission and `null` mean on each field, which provider a Session uses and which protocols each harness accepts. [Harness selection](../../contracts/agents-api/model-execution.md#harness-selection) defines the `harness` field. With the Core key, `AdminClient` reads the configuration a Session froze at creation and sets each harness's deployment default: diff --git a/packages/claude-sdk-adapter/README.md b/packages/claude-sdk-adapter/README.md index 693935775..485a97560 100644 --- a/packages/claude-sdk-adapter/README.md +++ b/packages/claude-sdk-adapter/README.md @@ -2,7 +2,7 @@ This package translates the pinned native Claude Agent SDK into OpenAgentCore's common Executor and Turn lifecycle. It owns the private TypeScript bridge and the native SDK configuration. The [Go adapter](../../apps/daemon/internal/agent/claudesdk) owns the bridge subprocess and translates bridge frames into the shared Runtime protocol. -[Harness onboarding](../../contracts/agents-api/harness-onboarding.md) defines the shared interfaces, registration and acceptance. Public qualification belongs to [the harness contract](../../contracts/agents-api/harnesses.md) and its linked operation contracts; local readiness cannot expand it. +[Harness onboarding](../../contracts/agents-api/harness-onboarding.md) defines the shared interfaces, registration and acceptance. Public qualification belongs to [Harness capabilities](../../contracts/agents-api/harness-capabilities.md) and its linked operation contracts; local readiness cannot expand it. ## Develop and verify @@ -60,7 +60,7 @@ Root assistant tool calls and live root user results produce the neutral MCP obs ### Deferred function discovery -Workspace deferred-function discovery uses native ToolSearch alongside the normal workspace tool profile. Its readiness feature is `workspace_tool_search`, in addition to `tool_search` and the workspace/function features. Qualification, combination limits and model-policy limitations are owned by [Deferred function discovery](../../contracts/agents-api/tool-search.md). +Workspace deferred-function discovery uses native ToolSearch alongside the normal workspace tool profile. Its readiness feature is `workspace_tool_search`, in addition to `tool_search` and the workspace/function features. Qualification, combination limits and model-policy limitations are owned by [Deferred function discovery](../../contracts/agents-api/execution-tools.md#deferred-function-discovery). ### Registration and state @@ -72,7 +72,7 @@ SDK state lives under `paths.ProfileDir(profile)/runtime/claude-sdk`, independen The SDK descriptor advertises the validated daemon subset, including durable Turns/input receipts, text observations, function tools, raw usage and restrictive execution controls. It does not advertise permissions, product authoring, raw tool Items, general web-search control or text-verbosity levels. Router admission for `environment:none` uses the available engine capability, not an engine name. -Core owns public admission for Claude: [harness selection](../../contracts/agents-api/harness-selection.md) chooses the engine for each Session; one [engine policy](../../contracts/agents-api/harness-onboarding.md#add-the-engine-to-core) serves API admission, device selection and the final claim; the [qualified operations](../../contracts/agents-api/harnesses.md#current-qualified-operations) table records Claude's medium-only verbosity and object-root function schemas; [function result images](../../contracts/agents-api/function-result-images.md) defines which placements accept image results; and [deployment defaults](../../contracts/agents-api/model-execution.md#deployment-defaults) define the model provider Core freezes for a Session. The adapter receives that provider as the adapter-owned `model_provider`, never in public Session configuration. Without one, a `none` host uses the daemon's own provider environment. The adapter alone selects the provider environment and removes credentials from native tool environments. +Core owns public admission for Claude: [harness selection](../../contracts/agents-api/model-execution.md#harness-selection) chooses the engine for each Session; one [engine policy](../../contracts/agents-api/harness-onboarding.md#add-the-engine-to-core) serves API admission, device selection and the final claim; the [qualified operations](../../contracts/agents-api/harness-capabilities.md) table records Claude's medium-only verbosity and object-root function schemas; [function result images](../../contracts/agents-api/function-result-images.md) defines which placements accept image results; and [deployment defaults](../../contracts/agents-api/model-execution.md#deployment-defaults) define the model provider Core freezes for a Session. The adapter receives that provider as the adapter-owned `model_provider`, never in public Session configuration. Without one, a `none` host uses the daemon's own provider environment. The adapter alone selects the provider environment and removes credentials from native tool environments. The `none` profile accepts only text, explicit model/system instructions, managed state, exact native resume and declared functions with ordered text or successful inline PNG/JPEG results, and the HTTP MCP subset described above. It rejects unsupported request options and disables built-in tools and undeclared MCP discovery. `DisableExecutionEnvironment` and `DisableSubagents` are accepted assertions about the single-Agent restrictive profile. Omission does not enable built-in tools. Explicit Subagent observation enables only the native delegation tools described under [Subagents](#subagents). Single-Agent new and resumed queries use the SDK's empty built-in tool set, explicit function MCP configuration and allowlist, strict MCP configuration and empty user/project/local setting sources. Without HTTP MCP declarations, native initialization and real provider request inventories must contain only the declared host functions. Managed operator policy may further restrict execution; it must not widen the profile. The profile limits model tool access; it does not isolate native state files or filesystem access by an explicitly supplied host function. The private factory accepts typed execution controls only for disabled search and medium text verbosity. Search remains excluded by the native tool inventory; medium retains the SDK's default text generation, without adding instructions or changing caller input. The pinned SDK has no native verbosity-level option, so low and high verbosity and enabled search are unsupported. Missing/invalid fields in a supplied control block fail before native setup; omitting the block keeps the same restrictive profile. Use the SDK's history lookup before explicit resume; never fall back to a new Session. Native files remain device-affine under a caller-selected managed runtime directory. The launch configuration supplies trusted provider environment; request options cannot supply environment variables or business write authority. Omitted, null and empty `system_prompt` map to empty SDK instructions only at this adapter boundary; null model values and unsupported options remain rejected. @@ -90,7 +90,7 @@ Each SDK result supplies one native usage snapshot, including reported failures. Active text uses the SDK's `AsyncIterable` input, with a fresh native UUID mapped to each daemon input ID. A native query may fold text into its current native turn or queue another; one daemon Run can therefore contain several native turns. Never promise Codex's same-native-turn semantics. Writes, queued notifications and user-message echoes do not confirm consumption. Only matching root assistant/partial/result `user_message_uuids` (or the singular fallback) confirm applied input. Typed mid-turn folds may appear only on the native result. Preserve that receipt even when the result reports failure. Check pending functions after the query drains: the SDK may dispatch later-turn callbacks before the earlier result handler finishes. Keep Turn input admission open until every submitted input has a consuming result, even when an earlier result reports an empty native queue. Close admission before releasing final receipt waiters. The outer SDK iterator remains open across Turns. Cancellation resolves unconfirmed pending receipts as unknown and interrupts the exact native Turn. Successful Cancel requires confirmed Turn settlement; process exit alone cannot establish a successful cancellation. Reuse also requires an empty confirmed native queue and settled child work. Otherwise close the Executor. The settled CancellationOutcome retains verified native identity, partial text and observed Usage; a requested resume identity alone is not evidence. Caller wait expiry reports failure/unknown while cleanup retains ownership. Output backpressure cannot turn missing native facts into confirmed settlement. Closed Turn output precedes successful AwaitSettlement; the settled outcome remains readable when connection loss prevents publication. A receipt timeout after a full write preserves the process and pending identity without redelivery; a blocked write is cancelled and released. The private adapter permits one input awaiting consumption and at most 63 extra inputs per Run, preserving the native 64-UUID receipt bound. Durable receipt opt-in separates bounded writes from native consumption waits; calls without it keep the router's ten-second deadline. Larger input capacity and recovery of interrupted input are not supported. Daemon registration alone does not establish public acceptance. -The [harness contract](../../contracts/agents-api/harnesses.md) and its linked operation contracts record current public qualification. Keep native execution evidence separate from local registration and packaging checks. Live adapter acceptance uses a real provider with private credentials; fixture tests cannot substitute for it. +[Harness capabilities](../../contracts/agents-api/harness-capabilities.md) and its linked operation contracts record current public qualification. Keep native execution evidence separate from local registration and packaging checks. Live adapter acceptance uses a real provider with private credentials; fixture tests cannot substitute for it. ## Subagents diff --git a/packages/mcode-harness/README.md b/packages/mcode-harness/README.md index 5b468ce7c..d0893d1c8 100644 --- a/packages/mcode-harness/README.md +++ b/packages/mcode-harness/README.md @@ -32,7 +32,7 @@ Native tool schemas are retained. Text and image results use standard MCP conten ## Tests -`make check-mcode-harness` runs the package's Node tests and syntax checks. Qualify changes with the [Harness acceptance checklist](../../contracts/agents-api/harnesses.md#acceptance-checklist); synthetic probes and native model runs do not complete public Files/Artifacts or independent Core acceptance. +`make check-mcode-harness` runs the package's Node tests and syntax checks. Qualify changes with the [Harness acceptance checklist](../../contracts/agents-api/harness-onboarding.md#qualify-the-adapter); synthetic probes and native model runs do not complete public Files/Artifacts or independent Core acceptance. For the packaged Linux regression, provide an operator-owned private profile and artifact directory, then run `native.test.mjs` inside the qualified Docker Runtime: diff --git a/scripts/check-names.test.py b/scripts/check-names.test.py index 1410b700e..0dcad8c35 100644 --- a/scripts/check-names.test.py +++ b/scripts/check-names.test.py @@ -93,8 +93,8 @@ def test_persisted_domain_exception_does_not_allow_other_settings_or_paths(self) self.assertEqual([item[2] for item in found], ["PARSAR"]) self.assertTrue(names.violations("README.md", content, rules)) - def test_retirement_table_exception_does_not_hide_runtime_setting(self): - rules = names.load_rules(Path(__file__).with_name("name-allowlist.json")) + def test_scoped_exception_does_not_hide_runtime_setting(self): + rules = [self.rule(r"AGENTS_API_ADDR", path="services/core/cmd/server/process_configuration.go")] content = '{"AGENTS_API_ADDR", "OAC_ADDR"}; os.Getenv("PARSAR_HOME")' found = names.violations("services/core/cmd/server/process_configuration.go", content, rules) self.assertEqual([item[2] for item in found], ["PARSAR"]) diff --git a/scripts/generate-harness-catalog.py b/scripts/generate-harness-catalog.py index c9b8d6eb8..eedac3df3 100644 --- a/scripts/generate-harness-catalog.py +++ b/scripts/generate-harness-catalog.py @@ -99,20 +99,13 @@ def render(entries): Path("contracts/agents-api/harness-catalog.md"): f''' # Built-in Harness registrations -The authored registration list is -[{tick}catalog.json{tick}](../../internal/harnessconfig/builtin/catalog.json). -Change it and run {tick}make generate-harness-catalog{tick}; generated projections are checked -by {tick}make check-harness-catalog{tick}. +The authored registration list is [{tick}catalog.json{tick}](../../internal/harnessconfig/builtin/catalog.json). Change it and run {tick}make generate-harness-catalog{tick}; {tick}make check-harness-catalog{tick} checks the generated projections. | Identifier | Display name | Model configuration declaration | Core qualification constructor | | --- | --- | --- | --- | {rows} -Configuration declarations live in {tick}internal/harnessconfig/{tick}; -qualification constructors live in {tick}services/core/internal/engine{tick}. -These registrations describe the build. Deployment enablement, qualified -operations and a connected Runtime's actual availability remain separate checks. -See [Harness onboarding](harness-onboarding.md) for adapter and packaging steps. +Configuration declarations live in {tick}internal/harnessconfig/{tick} and qualification constructors in {tick}services/core/internal/engine{tick}. These registrations describe the build. Which Harnesses a deployment enables is the {tick}core.harnesses{tick} [process setting](../../docs/configuration.md#settings), [Harness capabilities](harness-capabilities.md) lists what each Harness supports, and a connected Runtime reports its own availability. [Harness onboarding](harness-onboarding.md) describes the adapter and packaging steps. ''', } diff --git a/scripts/name-allowlist.json b/scripts/name-allowlist.json index d1967cda5..721dd2745 100644 --- a/scripts/name-allowlist.json +++ b/scripts/name-allowlist.json @@ -309,11 +309,6 @@ "regex": "Parsar owns product|Parsar, not the HTTP contract|Parsar cutover|Parsar services|no Parsar dependency", "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." }, - { - "path": "contracts/agents-api/model-execution.md", - "regex": "Parsar manages", - "reason": "These exact phrases refer to the separate Parsar product, its ownership or historical source, not the OpenAgentCore brand." - }, { "path": "contracts/agents-api/official-semantics-alignment.md", "regex": "The Parsar\\n", @@ -349,21 +344,6 @@ "regex": "(?:AGENTS_API_|CORE_CONSOLE_|PARSAR_)[A-Z0-9_]*\\*?|X-Parsar-Node-ID", "reason": "These names appear in retirement documentation or instructions for an explicitly pinned historical release; current configuration uses OAC settings." }, - { - "path": "contracts/agents-api/environment-templates.md", - "regex": "(?:AGENTS_API_|CORE_CONSOLE_|PARSAR_)[A-Z0-9_]*\\*?|X-Parsar-Node-ID", - "reason": "These names appear in retirement documentation or instructions for an explicitly pinned historical release; current configuration uses OAC settings." - }, - { - "path": "contracts/agents-api/harness-selection.md", - "regex": "(?:AGENTS_API_|CORE_CONSOLE_|PARSAR_)[A-Z0-9_]*\\*?|X-Parsar-Node-ID", - "reason": "These names appear in retirement documentation or instructions for an explicitly pinned historical release; current configuration uses OAC settings." - }, - { - "path": "contracts/agents-api/model-execution.md", - "regex": "(?:AGENTS_API_|CORE_CONSOLE_|PARSAR_)[A-Z0-9_]*\\*?|X-Parsar-Node-ID", - "reason": "These names appear in retirement documentation or instructions for an explicitly pinned historical release; current configuration uses OAC settings." - }, { "path": "docs/getting-started/operations.md", "regex": "~/\\.parsar/core", diff --git a/services/core/IMPLEMENTATION.md b/services/core/IMPLEMENTATION.md index 34a9a14a0..5a203a8ea 100644 --- a/services/core/IMPLEMENTATION.md +++ b/services/core/IMPLEMENTATION.md @@ -832,7 +832,7 @@ replaced; do not carry obsolete compatibility code forward to satisfy this secti unsupported non-default levels remain an explicit implementation gap. Product requests that omit the native option retain their existing defaults. Structured output has a separately qualified profile described in - [Structured output execution](../../contracts/agents-api/structured-output.md#core-and-runtime-boundary). + [Structured output execution](../../contracts/agents-api/execution-tools.md#structured-output). - `subagent_control` advertises native subagent tool control. Agents API requires it when resolved `multi_agent.enabled` is false and sends the typed internal `disable_subagents` policy on both new and resumed Turns. Native translation diff --git a/services/core/README.md b/services/core/README.md index a15a98fb5..0ca8f7b67 100644 --- a/services/core/README.md +++ b/services/core/README.md @@ -7,7 +7,7 @@ It owns reusable Agents, durable Sessions/Turns/Items, live events, function act and a daemon execution worker. Public execution supports qualified Codex, Claude Code (`claude_sdk`) and MiniMax Code (`mcode`) profiles through the shared Runtime contract. The three-harness Linux amd64 Docker V1 MVP has accepted evidence. V1 user-managed -Runtime enrollment has [recorded real acceptance](../../contracts/agents-api/user-managed-runtime-v1.md) +Runtime enrollment has [recorded real acceptance](../../contracts/agents-api/harness-capabilities.md) with explicit deployment coverage. [E2B](deploy/e2b/README.md) can back Core-managed hosted sandboxes, selected in Web's sandbox setup, or application-managed `self_hosted` Environments. @@ -263,12 +263,12 @@ text Sessions share the existing preparation, execution, Files and recovery path Networking defaults to enabled; disabled and exact-domain restricted policy are supported after setup. System/npm/Python packages use the shared initializer; remaining unsupported combinations are explicit gaps. Initial inline/file_id files, confidential env, npm/Python packages, -ordered setup and [public Environment Templates](../../contracts/agents-api/environment-templates.md) +ordered setup and [public Environment Templates](../../contracts/agents-api/environments.md#templates) resolve to the same immutable hosted configuration, independently of provider templates. Additional harnesses require separate integration and qualification. Connected describes the authenticated Runtime connection, not native readiness. A provisioning failure fails the Session with a safe step and exit-status reason -([initialization failure](../../contracts/agents-api/environment-templates.md#initialization-failure--september-23)); +([initialization failure](../../contracts/agents-api/environments.md#initialization-state-and-failure)); other exact hosted failure and expiry semantics remain unverified. The independent Docker Provider consumes an immutable Runtime image and retains @@ -325,7 +325,7 @@ connections alone do not start a Turn. Submit text, cancellation or function res through the official Session events endpoint; the worker assigns a same-tenant host and preserves that binding. Managed Docker has three-harness evidence. This generic device provisioning path is for `none`; self-hosted Sessions require the dedicated enrollment below. -User-managed enrollment has a separate [qualification record](../../contracts/agents-api/user-managed-runtime-v1.md); complete protocol semantics remain partial. See the [ownership rules](../../AGENTS.md#public-api). +User-managed enrollment has a separate [qualification record](../../contracts/agents-api/harness-capabilities.md); complete protocol semantics remain partial. See the [ownership rules](../../AGENTS.md#public-api). ### Enable Claude SDK execution @@ -536,7 +536,7 @@ and reads. Exact upstream failure/error timing remains unverified. The Runtime's shared Go workspace implementation owns directory reads, writes and output export; Harness adapters use the same authorized workspace binding. -The [qualification record](../../contracts/agents-api/user-managed-runtime-v1.md) +The [qualification record](../../contracts/agents-api/harness-capabilities.md) identifies fixed-SDK/raw HTTP, real-model, Files/Artifacts, cancellation, restart/history and credential lifecycle evidence, with its recorded revision limits. diff --git a/services/core/deploy/claude/README.md b/services/core/deploy/claude/README.md index 587ca9c1b..2cbd3d301 100644 --- a/services/core/deploy/claude/README.md +++ b/services/core/deploy/claude/README.md @@ -89,4 +89,4 @@ it implements the shared `agent.ExecutorFactory`, `Executor` and `Turn` interfaces and registers them in [`cli/claude_sdk.go`](../../../../apps/daemon/internal/cli/claude_sdk.go). Acceptance requirements are in -[Harness integration](../../../../contracts/agents-api/harnesses.md#acceptance-checklist). +[Harness onboarding](../../../../contracts/agents-api/harness-onboarding.md#qualify-the-adapter). diff --git a/services/core/deploy/codex/README.md b/services/core/deploy/codex/README.md index cb75becfe..6917d71c3 100644 --- a/services/core/deploy/codex/README.md +++ b/services/core/deploy/codex/README.md @@ -124,7 +124,7 @@ deletion revokes authority before owned container/volume cleanup. The daemon doe not enforce `disabled` or `restricted` networking; combinations without matching outer enforcement are unsupported. Templates and inline configuration share initial files, env, npm/Python packages, ordered setup and capabilities; see the -[supported fields and limits](../../../../contracts/agents-api/environment-templates.md). +[supported fields and limits](../../../../contracts/agents-api/environments.md#templates). ### Environment initialization @@ -142,4 +142,4 @@ daemon never runs apt, sudo or another privilege escalation, and operation requiring it. The image no longer contains a system-root seed or system-package launcher. Native self-hosted users prepare their own dependencies before starting the daemon. See -[initialization contract and limits](../../../../contracts/agents-api/environment-templates.md). +[initialization contract and limits](../../../../contracts/agents-api/environments.md#runtime-capability-preparation). diff --git a/services/core/deploy/e2b/README.md b/services/core/deploy/e2b/README.md index 904d83a74..27842da91 100644 --- a/services/core/deploy/e2b/README.md +++ b/services/core/deploy/e2b/README.md @@ -185,7 +185,7 @@ python -m unittest discover -s services/core/deploy/e2b -p '*_test.py' -v These are controlled startup-contract tests: input binding, protected key output, unchanged URL, no secret in argv/environment/record, one-shot claim, and retained provider ID on unknown outcomes. They do not create billable resources or qualify -E2B security, enrollment or model execution. The [qualification record](../../../../contracts/agents-api/user-managed-runtime-v1.md) +E2B security, enrollment or model execution. The [qualification record](../../../../contracts/agents-api/harness-capabilities.md) records historical three-harness deployment results and their verification limits, including shared Core credential lifecycle checks and explicit application-owned cleanup. `tests/official_user_runtime.py` supplies the shared diff --git a/services/core/deploy/mcode/README.md b/services/core/deploy/mcode/README.md index 9ef896140..0539f74b7 100644 --- a/services/core/deploy/mcode/README.md +++ b/services/core/deploy/mcode/README.md @@ -159,7 +159,7 @@ Multi-agent workspace execution installs only the existing authorized workspace MCP entry in that private native configuration so children inherit the same tools; public MCP and Environment-origin MCP combinations remain separately qualified. Complete the hosted -[acceptance checklist](../../../../contracts/agents-api/harnesses.md#acceptance-checklist) +[acceptance checklist](../../../../contracts/agents-api/harness-onboarding.md#qualify-the-adapter) before enabling hosted execution. The standalone companion uses its own npm lock; `make check` runs its lifecycle tests and script checks, while its exact-source Linux build and native qualification for the actual supported