Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,17 @@

All notable changes to GopherAgent are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/); versions follow [Semantic Versioning](https://semver.org/) — pre-1.0, breaking API changes only require a minor bump.

## [v0.42.0] — 2026-08-15

### Added

- **`openai.JSONMode` — a compatible endpoint now states which JSON mode it implements instead of being assumed into OpenAI's.** "OpenAI-compatible" describes the endpoint, the request and response shape, and the SDK; it does not promise every extension built on top of them. Structured output is where that gap bites: `api.openai.com` takes `response_format {type:"json_schema"}` with the schema inline and enforces it server-side, while several compatible endpoints publish only the older `{type:"json_object"}`, which guarantees JSON syntax and carries no schema field at all. The adapter sent `json_schema` unconditionally, so on such an endpoint every schema-constrained call was a 400 while plain chat and tool calling kept working — the failure landed exactly on the planner, extractor, and judge stages that depend on structured output, and nowhere else. `WithJSONMode` takes `JSONModeSchema`, `JSONModeObject`, or `JSONModeNone`; `WithImageInput` declares the other feature a gateway may or may not have. Under `JSONModeObject` the schema moves into the prompt as a trailing system message, appended last rather than merged into an existing system prompt because the schema is an instruction and the final message is the one models follow most closely. The rendered text always contains the literal word JSON, which is load-bearing rather than stylistic: endpoints implementing this format commonly reject a request whose messages never mention it, to stop callers switching on JSON mode and then asking for prose. What `JSONModeObject` cannot do is enforce anything — the endpoint guarantees only that the reply parses, so schema conformance degrades from a server-side constraint to a model instruction and `Strict` becomes a request rather than a rule. Callers on that path must validate the result and be ready to retry; the adapter does not retry on their behalf, because how many attempts a malformed reply is worth is policy the caller owns. (`pkg/llm/openai/json_mode.go`)

### Changed (breaking)

- **`openai.NewCompat` no longer claims image input or structured output on the endpoint's behalf.** `Capabilities()` was a method on the shared `*Provider` type returning a hardcoded `{ImageInput: true, StructuredOutput: true}`, and `NewCompat` returns that same type — so a provider pointed at any gateway in the ecosystem reported full support regardless of what the gateway actually implemented. That inverts the entire point of the signal. `CapabilityProvider` exists so a consumer can reject an unsuitable provider at construction instead of discovering the gap from a confident, wrong answer, and for compatible endpoints it produced precisely that: a pre-flight check that passed, followed by a failure mid-run. The report now derives from configuration — `StructuredOutput` is true whenever a JSON mode is declared — so a single source of truth governs both what is claimed and what goes on the wire, and the two cannot drift. `New` is unchanged and still reports both, since it does speak for `api.openai.com`. `NewCompat` starts from `JSONModeNone` and no image claim, in the same spirit as its already requiring an explicit model: a compatible endpoint has no sensible default, and the adapter declares nothing it cannot know. Callers using structured output through `NewCompat` must add `WithJSONMode(...)`; the alternative was to keep defaulting to `json_schema`, which preserves both the false claim and the 400.
- **A structured-output call against `JSONModeNone` fails before the request is sent.** The error names the option that fixes it. Silently dropping the constraint was the worse option in both directions: the model returns prose, and the caller's unmarshal fails somewhere unrelated to the cause — while forwarding a `response_format` the endpoint never published produces a 400 that names neither. (`pkg/llm/openai/openai.go`, `pkg/llm/openai/openai_compat.go`)

## [v0.41.0] — 2026-08-10

### Added
Expand Down Expand Up @@ -635,6 +646,7 @@ Multi-user, long-running, audit-friendly chat surface — the foundation for sid
- README section on the permission flow — documents `RequiresConfirmation` × `ConfirmHITL` × `Permissions` interaction.
- Enum struct tag support in `tools.SchemaFor[T]()` — emit values into JSON-Schema's `enum` array so providers reject invalid values upstream.

[v0.42.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.42.0
[v0.41.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.41.0
[v0.40.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.40.0
[v0.39.0]: https://github.com/hung12ct/gopheragent/releases/tag/v0.39.0
Expand Down
44 changes: 41 additions & 3 deletions docs/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,43 @@ Compatible base URLs must be absolute HTTP(S) URLs without embedded
credentials, query parameters, or fragments. Use HTTPS for remote gateways;
plain HTTP remains available for local Ollama/vLLM development.

### Declare what the endpoint implements

"OpenAI-compatible" covers the endpoint, the request/response shape, and the
SDK — not every extension built on top of them. `NewCompat` therefore claims
nothing beyond plain chat and tool calling, in the same spirit as requiring an
explicit model: the adapter cannot infer the rest from a base URL.

Two declarations matter:

| Option | Effect |
|---|---|
| `WithJSONMode(JSONModeSchema)` | `response_format {type:"json_schema"}` with the schema inline, enforced server-side. What `api.openai.com` implements. |
| `WithJSONMode(JSONModeObject)` | `response_format {type:"json_object"}` plus the schema appended to the prompt, for endpoints publishing only the older format. |
| `WithImageInput(true)` | Declares that the endpoint accepts image parts on user messages. |

```go
// Endpoint publishing OpenAI's json_schema:
p, _ := openai.NewCompat(key, "openai/gpt-4o", "https://openrouter.ai/api/v1",
openai.WithJSONMode(openai.JSONModeSchema), openai.WithImageInput(true))

// Endpoint publishing only json_object:
p, _ := openai.NewCompat(key, "deepseek-v4-flash", "https://api.deepseek.com/v1",
openai.WithJSONMode(openai.JSONModeObject))
```

Without a declaration the JSON mode is `JSONModeNone`, and a structured-output
call fails with an error naming the option instead of sending a
`response_format` the endpoint may reject. Check the endpoint's own docs for
which formats it publishes — the two are not interchangeable, and sending the
wrong one is a 400 on every structured call.

`JSONModeObject` guarantees only that the reply parses as JSON. Schema
conformance degrades from a server-side constraint to a prompt instruction, so
validate the result and be ready to retry; `Strict` becomes a request rather
than a rule. This is the same trade the Anthropic adapter avoids by
synthesizing a tool — an option `json_object` endpoints do not offer.

### Point the non-chat clients at the same endpoint

`NewCompat` configures the **chat provider only**. The embedder, vision
Expand All @@ -52,9 +89,10 @@ silently does nothing.

`agent.CapabilityProvider` lets a consumer that requires image input or
structured output reject an unsuitable provider at construction, instead of
discovering the gap from a confident, wrong answer. The OpenAI, Anthropic, and
Gemini adapters report both; `llmfake.ScriptedProvider` reports neither, since
it replays a script without ever reading a message's media parts.
discovering the gap from a confident, wrong answer. `openai.New`, Anthropic,
and Gemini report both; `openai.NewCompat` reports what the caller declared
(see above); `llmfake.ScriptedProvider` reports neither, since it replays a
script without ever reading a message's media parts.

```go
if c, ok := provider.(agent.CapabilityProvider); ok && !c.Capabilities().ImageInput {
Expand Down
4 changes: 4 additions & 0 deletions pkg/agent/structured_output.go
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,10 @@ import (
//
// - OpenAI (gpt-4o-*, o-series): response_format: {type:"json_schema",
// json_schema:{name, description, schema, strict}}. Native and strict.
// Compatible endpoints publishing only the older {type:"json_object"}
// declare it with openai.WithJSONMode, which moves the schema into the
// prompt: still JSON, but conformance stops being enforced server-side,
// so validate the result.
// - Gemini (1.5+): response_mime_type:"application/json" +
// response_schema. Native; Strict is ignored (Gemini always enforces).
// - Anthropic Claude: no native JSON mode. Providers synthesize a single
Expand Down
154 changes: 154 additions & 0 deletions pkg/llm/openai/json_mode.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,154 @@
package openai

import (
"context"
"fmt"
"strings"

"github.com/hung12ct/gopheragent/pkg/agent"
"github.com/sashabaranov/go-openai"
)

// JSONMode selects the wire encoding an endpoint accepts for a
// schema-constrained response.
//
// api.openai.com takes response_format {type:"json_schema"} with the schema
// inline and enforces it server-side. Compatible endpoints implement an
// unknown subset of that: several publish only the older
// {type:"json_object"}, which guarantees JSON syntax but carries no schema
// field at all, and some publish neither. Sending json_schema to an endpoint
// that does not implement it is a 400 on every structured call, so the mode
// is configuration — the adapter cannot infer it from a base URL.
type JSONMode int

const (
// JSONModeSchema sends response_format {type:"json_schema", ...} with the
// schema inline, and the endpoint enforces it. Zero value because it is
// what this package is named for; NewCompat overrides it.
JSONModeSchema JSONMode = iota
// JSONModeObject sends response_format {type:"json_object"} and appends
// the schema to the prompt, because that wire format has nowhere to put
// it. The endpoint guarantees only that the reply parses as JSON —
// conformance to the schema degrades to a model instruction, so callers
// should validate the result and be ready to retry.
JSONModeObject
// JSONModeNone declares no JSON mode. A structured-output request against
// such an endpoint fails at the call, rather than silently downgrading to
// free-form prose that the caller would then try to unmarshal.
JSONModeNone
)

// String implements fmt.Stringer so the mode reads as its constant name in
// error messages instead of an integer.
func (m JSONMode) String() string {
switch m {
case JSONModeSchema:
return "JSONModeSchema"
case JSONModeObject:
return "JSONModeObject"
case JSONModeNone:
return "JSONModeNone"
default:
return fmt.Sprintf("JSONMode(%d)", int(m))
}
}

// WithJSONMode declares which JSON mode this endpoint implements. It is
// meaningful only on a compatible endpoint: New already reports the mode
// api.openai.com implements, and NewCompat claims none until this says
// otherwise.
//
// It also drives Capabilities().StructuredOutput, so declaring the mode is
// what lets a consumer that depends on schema enforcement accept the
// provider at construction.
func WithJSONMode(m JSONMode) Option {
return providerOptionFunc(func(p *Provider) { p.jsonMode = m })
}

// WithImageInput declares whether this endpoint accepts image parts on user
// messages. Same reasoning as WithJSONMode: New reports OpenAI's own
// support, NewCompat claims none until the caller declares it.
//
// It sets Capabilities().ImageInput only. Media parts are rendered on the
// wire either way, so this changes what the provider promises, not what it
// sends — a caller that skips the capability check still reaches the
// endpoint's own rejection.
func WithImageInput(ok bool) Option {
return providerOptionFunc(func(p *Provider) { p.imageInput = ok })
}

// structuredOutputFor translates the provider-neutral structured-output
// request on ctx into this endpoint's JSON mode. It returns the
// response_format to send plus the messages to send with it, since
// JSONModeObject carries the schema as an extra message. Both are returned
// unchanged when no structured output was requested.
func (p *Provider) structuredOutputFor(
ctx context.Context,
msgs []openai.ChatCompletionMessage,
) (*openai.ChatCompletionResponseFormat, []openai.ChatCompletionMessage, error) {
so := agent.StructuredOutputFromContext(ctx)
if so == nil {
return nil, msgs, nil
}
switch p.jsonMode {
case JSONModeSchema:
name := so.Name
if name == "" {
name = "response"
}
return &openai.ChatCompletionResponseFormat{
Type: openai.ChatCompletionResponseFormatTypeJSONSchema,
JSONSchema: &openai.ChatCompletionResponseFormatJSONSchema{
Name: name,
Description: so.Description,
Schema: jsonSchemaMarshaler(so.Schema),
Strict: so.Strict,
},
}, msgs, nil
case JSONModeObject:
instruction, err := jsonObjectInstruction(so)
if err != nil {
return nil, nil, err
}
// Appended last rather than merged into an existing system message:
// the schema is an instruction, and the final message is the one
// models follow most reliably.
return &openai.ChatCompletionResponseFormat{
Type: openai.ChatCompletionResponseFormatTypeJSONObject,
}, append(msgs, openai.ChatCompletionMessage{
Role: openai.ChatMessageRoleSystem,
Content: instruction,
}), nil
default:
return nil, nil, fmt.Errorf(
"openai: structured output requested but this endpoint declares %s: pass WithJSONMode(JSONModeSchema) or WithJSONMode(JSONModeObject) to the constructor",
p.jsonMode)
}
}

// jsonObjectInstruction renders the schema as a prompt fragment, which is
// where it has to live under JSONModeObject.
//
// The literal word JSON is load-bearing, not decoration: endpoints
// implementing this format commonly reject a request whose messages never
// mention it, to stop callers from switching on JSON mode and then asking
// for prose.
func jsonObjectInstruction(so *agent.StructuredOutput) (string, error) {
schema, err := so.MarshalSchema()
if err != nil {
return "", fmt.Errorf("openai: encoding structured-output schema: %w", err)
}
var b strings.Builder
b.Grow(len(schema) + 256)
b.WriteString("Reply with a single JSON object and nothing else: no prose, no code fence.\n")
b.WriteString("It must validate against this JSON Schema:\n")
b.Write(schema)
if so.Description != "" {
b.WriteString("\nThe object represents: ")
b.WriteString(so.Description)
}
if so.Strict {
b.WriteString("\nDo not emit any property the schema does not declare.")
}
return b.String(), nil
}
Loading
Loading