Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 32 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -1518,7 +1518,8 @@ replaced; do not carry obsolete compatibility code forward to satisfy this secti
Unsupported `low`/`high` and unreadable catalogs fail before model execution;
unsupported non-default levels remain an explicit implementation gap.
Product requests that omit the native option retain their existing defaults.
Structured output formats remain a separate protocol gap.
Structured output has a separately qualified profile described in
[Structured output execution](#structured-output-execution).
- `subagent_control` advertises native subagent tool control. Agents API requires
it when resolved `multi_agent.enabled` is false and sends the typed internal
`disable_subagents` policy on both new and resumed Turns. Native translation
Expand Down Expand Up @@ -1907,6 +1908,36 @@ foreign, ambiguous or metadata-only history rejects before model input. The Runt
volume and shared Environment/Session binding establish ownership; this lookup
cannot select another Session's home or infer ownership from a model response.

### Structured output execution

The public `text.format={type:"json_schema",schema:{...}}` is resolved with saved
Agent overrides and frozen in the existing Session configuration. Core transports
it in `ExecutionControls.OutputFormat`; it does not append prompt instructions,
validate/retry model answers, repair JSON or select native tool names. Public
admission requires the selected profile's structured-output qualification, and
only requests using this option require the Runtime's `structured_output` and
message-observation capabilities. A capability advertisement does not qualify a
new public combination. Claude advertises this operation only when the installed
SDK bridge reports its `structured_output` feature and the selected Runtime is
not a workspace profile.

The current qualified path is Claude SDK, `environment:none`, medium verbosity,
single Agent, with optional ordinary function tools and text results. Workspace,
HTTP MCP, Subagent combinations and non-object root schemas remain unqualified.
The SDK uses binary64 JSON numbers: reject execution schemas whose numeric values
would change during that conversion, without narrowing saved Agent storage.
Codex and MiniMax structured output remain explicit execution gaps.

The Claude adapter passes `outputFormat` to the maintained native SDK and allows
its native `StructuredOutput` terminal tool. A matching live root tool result and
an attributed successful SDK result confirm the final output. Publish the native
`result.result` string unchanged as a completed `final_answer` Message using the
native tool-use ID; the parent assistant ID can already own a prose Item. Do not
publish unvalidated retry candidates or serialize `structured_output` back to
JSON. Native retries remain harness-owned. Existing input receipts, usage,
cancellation, release and recovery rules apply unchanged. The implementation and
qualification limits are recorded in [the coverage note](contracts/agents-api/structured-output.md).

### Claude SDK adapter foundation

`packages/claude-sdk-adapter` privately owns the pinned official TypeScript SDK
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -75,3 +75,27 @@ func TestMCPWithoutEnvironmentNoneRejectedBeforeSetup(t *testing.T) {
t.Fatal("MCP reached native setup", err)
}
}

func TestStructuredOutputConfigurationReachesNativeUnchanged(t *testing.T) {
root := t.TempDir()
t.Setenv("PARSAR_HOME", root)
config := Config{Entrypoint: filepath.Join(root, "worker"), StateDir: filepath.Join(root, "state")}
schema := json.RawMessage(`{"type":"object","properties":{"n":{"const":9007199254740992}}}`)
request := proto.PromptRequestPayload{RunID: "run", Prompt: "Original input.", ObserveMessages: true, DisableSubagents: true, AgentOptions: map[string]any{"model": "model", "system_prompt": "Original instructions."}, ExecutionControls: &proto.ExecutionControls{WebSearch: "disabled", TextVerbosity: "medium", OutputFormat: &proto.OutputFormat{Type: "json_schema", Schema: schema}}}
start, _, err := prepare(config, request)
if err != nil {
t.Fatal(err)
}
if start.OutputFormat == nil || string(start.OutputFormat.Schema) != string(schema) || start.Prompt != request.Prompt || start.SystemPrompt != "Original instructions." {
t.Fatal("native configuration changed")
}
request.ExecutionControls.OutputFormat.Schema = json.RawMessage(`{"type":"object","const":9007199254740993}`)
if _, _, err := prepare(config, request); err == nil {
t.Fatal("lossy schema accepted")
}
request.ExecutionControls.OutputFormat.Schema = schema
request.DisableSubagents = false
if _, _, err := prepare(config, request); err == nil {
t.Fatal("unqualified subagent combination accepted")
}
}
11 changes: 11 additions & 0 deletions apps/parsar-daemon/internal/agent/claudesdk/options.go
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@ type subagentOptions struct {

type startRequest struct {
Subagents *subagentOptions `json:"subagents,omitempty"`
OutputFormat *proto.OutputFormat `json:"output_format,omitempty"`
Type string `json:"type"`
Prompt string `json:"prompt,omitempty"`
Model string `json:"model"`
Expand Down Expand Up @@ -66,6 +67,16 @@ func prepareConfiguration(config Config, req proto.PromptRequestPayload) (startR
if controls := req.ExecutionControls; controls != nil && (controls.WebSearch != "disabled" || controls.TextVerbosity != "medium") {
return fail("execution controls require disabled web search and medium text verbosity")
}
if req.ExecutionControls != nil && req.ExecutionControls.OutputFormat != nil {
format := req.ExecutionControls.OutputFormat
if format.Type != "json_schema" || !req.ObserveMessages || !req.DisableSubagents || config.Workspace != nil || req.MCPHTTPServers != nil {
return fail("structured output requires the qualified message-observing single-agent function profile")
}
if err := proto.ValidateBinary64Schema(format.Schema); err != nil {
return startRequest{}, nil, err
}
start.OutputFormat = format
}
if req.ObserveSubagentIdentities {
if req.DisableSubagents || len(req.FunctionTools) != 0 || req.MCPHTTPServers != nil || (req.LocalEnvironment != nil && len(req.LocalEnvironment.MCP) != 0) {
return fail("subagent execution does not support this tool combination")
Expand Down
4 changes: 4 additions & 0 deletions apps/parsar-daemon/internal/agent/claudesdk/readiness.go
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,10 @@ type RuntimeInfo struct {
Features []string `json:"features"`
}

func (info RuntimeInfo) SupportsStructuredOutput() bool {
return slices.Contains(info.Features, "structured_output")
}

func (info RuntimeInfo) SupportsSubagents() bool {
return slices.Contains(info.Features, "subagent_resources")
}
Expand Down
3 changes: 3 additions & 0 deletions apps/parsar-daemon/internal/agent/codex/preparation.go
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,9 @@ func newSession(parent context.Context, req proto.PromptRequestPayload, out chan
}

func newPreparation(parent context.Context, req proto.PromptRequestPayload, cfg sessionConfig) (*Prepared, error) {
if req.ExecutionControls != nil && req.ExecutionControls.OutputFormat != nil {
return nil, errors.New("codex: structured output is not qualified")
}
if req.WorkspaceReadOnly {
return nil, errors.New("codex: workspace reads use the local Runtime interface")
}
Expand Down
2 changes: 1 addition & 1 deletion apps/parsar-daemon/internal/agent/mcode/execution.go
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ func validateExecutionRequest(req proto.PromptRequestPayload) error {
if !req.ReleaseOnCompletion || !req.DisableExecutionEnvironment || req.WorkDir != "" || req.AgentStateKey == "" || req.LocalEnvironment != nil || req.RequireExistingNativeSession || len(req.FunctionTools) != 0 || (req.MCPHTTPServers != nil && len(*req.MCPHTTPServers) != 0) {
return fmt.Errorf("mcode: unsupported execution configuration")
}
if req.ExecutionControls == nil || req.ExecutionControls.WebSearch != "disabled" || (req.ExecutionControls.TextVerbosity != "" && req.ExecutionControls.TextVerbosity != "medium") {
if req.ExecutionControls == nil || req.ExecutionControls.OutputFormat != nil || req.ExecutionControls.WebSearch != "disabled" || (req.ExecutionControls.TextVerbosity != "" && req.ExecutionControls.TextVerbosity != "medium") {
return fmt.Errorf("mcode: unsupported execution controls")
}
if !req.DisableSubagents {
Expand Down
1 change: 1 addition & 0 deletions apps/parsar-daemon/internal/cli/claude_sdk.go
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,7 @@ func discoverClaudeSDK(rc *runContext, profile string, check func(context.Contex
caps.WorkspaceReadPreparation, caps.NativeSessionRecovery = true, true
}
out.Info.Available, out.Info.Version = true, info.SDK
out.Info.Capabilities.StructuredOutput = out.Config.Workspace == nil && info.SupportsStructuredOutput()
out.Info.Capabilities.SubagentObservations = info.SupportsSubagents()
out.Info.Capabilities.MCPHTTPTools = info.SupportsHTTPMCP()
out.Info.Capabilities.MCPHTTPBearerAuth = info.SupportsHTTPMCPBearer()
Expand Down
5 changes: 4 additions & 1 deletion apps/parsar-daemon/internal/cli/claude_sdk_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -133,7 +133,7 @@ func TestClaudeSDKFeatureDiscovery(t *testing.T) {
t.Fatal(err)
}
t.Setenv(claudeSDKNodeEnv, node)
for _, features := range [][]string{nil, {"mcp_http_tools"}, {"mcp_http_bearer_auth"}, {"mcp_http_tools", "mcp_http_bearer_auth"}, {"mcp_http_required"}, {"mcp_http_tools", "mcp_http_required"}, {"subagent_resources"}} {
for _, features := range [][]string{nil, {"mcp_http_tools"}, {"mcp_http_bearer_auth"}, {"mcp_http_tools", "mcp_http_bearer_auth"}, {"mcp_http_required"}, {"mcp_http_tools", "mcp_http_required"}, {"subagent_resources"}, {"structured_output"}} {
out := discoverClaudeSDK(&runContext{stdout: &strings.Builder{}, stderr: &strings.Builder{}}, "default", func(context.Context, claudesdk.Config) (claudesdk.RuntimeInfo, error) {
info := claudesdk.RuntimeInfo{SDK: "0.3.269", Native: "2.1.269 (Claude Code)", Features: features}
return info, nil
Expand All @@ -142,6 +142,9 @@ func TestClaudeSDKFeatureDiscovery(t *testing.T) {
if out == nil || !out.Info.Available || out.Info.Capabilities.MCPHTTPTools != supported || out.Info.Capabilities.MCPHTTPBearerAuth != (supported && slices.Contains(features, "mcp_http_bearer_auth")) || out.Info.Capabilities.MCPHTTPRequired != (supported && slices.Contains(features, "mcp_http_required")) {
t.Fatal("MCP feature discovery widened the runtime profile")
}
if out.Info.Capabilities.StructuredOutput != slices.Contains(features, "structured_output") {
t.Fatal("structured output feature does not match the installed runtime")
}
if out.Info.Capabilities.SubagentObservations != slices.Contains(features, "subagent_resources") {
t.Fatal("Subagent feature discovery does not match the runtime contract")
}
Expand Down
4 changes: 2 additions & 2 deletions contracts/agents-api/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -164,7 +164,7 @@ user-managed enrollment remain outside this qualification.
| --- | --- |
| Subagents / multi_agent | Six reads and same-child recovery have three-harness Docker evidence; optional native operations, live child progress, full lifecycle/interactions and tool combinations remain explicit gaps |
| Environment Templates | Unsupported restricted hostname forms, unqualified installation overrides/null network and exact hosted errors remain gaps. CRUD/list, files, env/setup/system/npm/Python, inline/referenced Skills, Plugins, workspace capability directories and Session references have recorded coverage. Environment Plugin MCP transport and placement limits are [listed separately](environment-templates.md#environment-origin-mcp-plugins) |
| Input and configuration | Non-text initial input, broader content/configuration unions, structured output and reasoning/verbosity combinations |
| Input and configuration | Non-text initial input, broader content/configuration unions and reasoning/verbosity combinations; [structured output](structured-output.md) has a qualified Claude function profile, with other combinations remaining gaps |
| Tools and interactions | Deferred functions, other tool types, effective tool-set enforcement and result/cancel publication ordering; MiniMax public functions and service-origin MCP remain unsupported |
| Vault and Credentials | OAuth/refresh, archive semantics, revocation/concurrent mutation and exact hosted selection/error behavior; static bearer CRUD/token replacement is already present |
| Existing resources | Full Item/SSE/Usage variants, omitted/null/default/error semantics, pagination and overlapping lifecycle behavior beyond recorded cases |
Expand Down Expand Up @@ -453,7 +453,7 @@ extension separately from upstream fields and document it here when implemented.

`openapi.yaml` is our generated supported surface; it is not the full upstream
specification. The shared Go wire types are in `v1`. Physical Session cleanup, non-text
message input, structured output execution, broader options/tools, remaining Vault lifecycle,
message input, broader structured-output combinations, broader options/tools, remaining Vault lifecycle,
Subagents and environment/provider resources remain incomplete. Reject unsupported
requests explicitly; persisted saved configuration is not execution admission.

Expand Down
7 changes: 6 additions & 1 deletion contracts/agents-api/harness-onboarding.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,10 +81,15 @@ An engine without native tools can guarantee their absence; an engine with tools
must actually disable them when requested. Configuration acceptance is not proof
of enforcement.

MCP, public function calls, image inputs, verbosity controls and other optional
MCP, public function calls, structured output, image inputs, verbosity controls and other optional
operations do not need to match another engine. Reject unqualified combinations
explicitly and record the gap. Never advertise a capability to bypass selection.

Structured-output adapters consume `ExecutionControls.OutputFormat` and publish
confirmed native output through the existing Message contract. Register public
qualification separately from the Runtime capability; see the
[structured-output boundary](structured-output.md). No Core engine-name branch is required.

Hosted workspace execution additionally requires verified preparation, workspace
reads/output export, network behavior and credential/history isolation. Reuse the
same dedicated Runtime binding and shared Files helpers. A native Bash sandbox
Expand Down
4 changes: 3 additions & 1 deletion contracts/agents-api/harnesses.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,7 +80,9 @@ syntactically or everything either upstream harness can theoretically perform.
| Non-default verbosity | Native/model-dependent support | No equivalent qualified; medium only |
| Public detailed Usage | Supported native counters | Native raw usage retained; public breakdown gap |
| V1 `self_hosted` daemon enrollment at `/workspace` | [Qualified deployment scope](user-managed-runtime-v1.md) | [Qualified deployment scope](user-managed-runtime-v1.md) |
| Explicit reasoning, structured output, enabled `multi_agent`, message images | Shared service gaps | Shared service gaps |
| Structured output | Unqualified; explicit rejection | [Qualified single-agent function profile](structured-output.md) |
| Explicit reasoning, message images | Shared service gaps | Shared service gaps |
| Six Subagent reads | [Qualified scope](subagents.md) | [Qualified scope](subagents.md) |

This inventory records supported combinations, not a feature-equality checklist.
Do not silently drop options, fabricate measurements, weaken isolation or remove
Expand Down
15 changes: 10 additions & 5 deletions contracts/agents-api/openapi.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1598,9 +1598,12 @@ definitions:
type: object
v1.TextFormat:
properties:
schema:
type: object
type:
enum:
- text
- json_schema
type: string
required:
- type
Expand Down Expand Up @@ -2689,11 +2692,13 @@ paths:
Unknown historical creators reject retries; known creators without recorded
intent retain resolved-snapshot retry rules. These conflict policies are local
and not verified hosted parity. Creation retries observe future events without
replay; retry with stream=false to retrieve the Session. Non-text initial
input remains unsupported. Basic Codex and Claude SDK openai_hosted creation
requires an explicitly configured managed provider. The Claude workspace profile
supports non-deferred function tools with text results alongside native workspace
tools; HTTP MCP remains unsupported. Idle Sessions provision automatically;
replay; retry with stream=false to retrieve the Session. Claude SDK environment:none
supports qualified object-root json_schema output with medium verbosity, single-Agent
execution and ordinary functions; other combinations remain unsupported. Non-text
initial input remains unsupported. Basic Codex and Claude SDK openai_hosted
creation requires an explicitly configured managed provider. The Claude workspace
profile supports non-deferred function tools with text results alongside native
workspace tools; HTTP MCP remains unsupported. Idle Sessions provision automatically;
initial provisioning has no caller connection action. Network defaults to
enabled; disabled and restricted exact ASCII hostnames are supported. Restricted
policy requires 1–100 allowed domains. Unsupported hostname forms and startup
Expand Down
62 changes: 62 additions & 0 deletions contracts/agents-api/structured-output.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# Structured output

The pinned Agents API accepts `text.format` with `type: "json_schema"` and a
`schema` object. This is not the Responses API wrapper: do not add `name`, `strict`
or an alternative public output field. The schema is saved and inherited through
the existing Agent/Session configuration resolution and immutable snapshot.

## Qualified execution profile

Claude SDK supports object-root schemas with `environment:none`, medium verbosity,
`multi_agent.enabled=false` and optional ordinary function tools returning text.
The native SDK remains responsible for its model/tool loop and schema validation.
Codex and MiniMax structured-output execution, workspace/HTTP MCP/Subagent
combinations and other root types are unqualified and explicitly rejected. These
are implementation gaps, not a redefinition of the official protocol.

The Claude SDK consumes JSON numbers as binary64. Session admission rejects schema
numbers whose values cannot survive that conversion; saved Agent resources still
retain such schemas exactly. This does not claim that the upstream model/harness
preserves arbitrary numeric candidates internally. The adapter never reparses and
serializes a final answer to construct public text.

## Core and Runtime boundary

`ExecutionControls.OutputFormat` carries the schema through the shared execution
request. The selected service profile qualifies the operation; `structured_output`
and message observations are checked only for requests using the option. An
advertised capability alone never enables public support. Other harnesses can
implement this same request and existing Message observations without Core
engine-name branches or a second execution loop.

The Claude adapter passes `outputFormat` to the fixed SDK. Its native
`StructuredOutput` tool is internal and is not an extra caller-defined function.
Successful live root calls are confirmed by their native tool-result receipt and
an attributed successful result containing `structured_output`. The adapter emits
a completed `final_answer` Message with the native tool-use ID and unchanged
`result.result` text. Parent assistant prose retains its own ID. Failed retries
and cancelled candidates cannot become a completed structured answer. No private
history read, output repair, schema coercion or prompt wrapper supplies the result.

Recovery uses the existing Session/Turn/Items queries and native continuation.
SSE is still live-only. Frozen schemas apply to both initial and resumed execution;
ordinary text configuration retains its prior behavior.

## Acceptance

`TestNativeStructuredOutputPublicExecution` and
`services/agents-api/tests/official_structured_output.py` exercise the pinned SDK,
raw HTTP, actual PostgreSQL/Worker/gateway/daemon and real model APIs. They require
explicit private operator options and never supply model responses. The workflow
covers a function-only random value, unchanged saved configuration, native result
application receipts, ordered terminal SSE, persisted final JSON, daemon restart
and same-history continuation, cancellation, text override and tenant isolation.

Focused tests cover native failure/retry projection, exact result bytes, schema
numeric admission, configuration transport and independent service qualification
for another harness. Running only these tests or importing the SDK is not a claim
of complete protocol compatibility.

The retained Codex provider/tool-chain failure is not reopened by Claude's native
qualification. Its next attempt needs a concrete changed prerequisite; do not
weaken assertions, rewrite model output or repeatedly sample until one run passes.
Loading
Loading