Skip to content

feat(binding-llm): llm(proxy) — route on dialect/model - #2598

Merged
jfallows merged 97 commits into
developfrom
claude/serene-ride-hlw83v
Sep 30, 2026
Merged

jfallows merged 97 commits into
developfrom
claude/serene-ride-hlw83v

Conversation

@jfallows

Copy link
Copy Markdown
Contributor

Description

Adds kind: proxy to type: llm: a same-protocol proxy that routes purely
on LlmBeginEx.dialect/model — both already resolved and populated by the
upstream llm(server)/llm(proxy) hop that produced the stream's own begin
extension. llm(proxy) never touches request/response content: it reads the
two fields off the flyweight once, resolves a route synchronously in
newStream() (same pattern TlsProxyFactory/McpProxyFactory use for
BEGIN-time routing), and relays DATA/END/ABORT/RESET/FLUSH/WINDOW/CHALLENGE
verbatim between the accepted network stream and the resolved exit. A
routing miss returns null from newStream(), and the engine resets the
stream automatically.

Also fixes a real flow-control bug this feature's own test coverage caught:
LlmProxyFactory's window-relay math computed the peer's new acknowledge
position as mySeq - peerWindow, which only advances correctly when the
peer's advertised window maximum grows monotonically. Once a relayed peer
instead holds its maximum steady while advancing acknowledge (a valid,
commonly-used convention), the relayed acknowledge never moved and the
stream stalled as soon as content exceeded a single window. The formula now
derives the relayed acknowledge from the peer's in-flight delta
(peerMax - peerWindow), which is correct under both conventions.

LlmProxyIT gained flow-control coverage (10k/100k, openai/anthropic,
routed to both configured exits) exercising payloads that span multiple
windows through the proxy — the existing small-payload routing tests never
exceeded a single window, so this gap was previously uncovered.
examples/llm.proxy now demonstrates llm(proxy) routing a subset of
models to a second same-dialect deployment (no translation) alongside the
existing cross-dialect translation routes.

Fixes #2502

Test plan

  • ./mvnw checkstyle:check -pl incubator/binding-llm.conf,incubator/binding-llm.spec,incubator/binding-llm
  • ./mvnw clean verify -pl incubator/binding-llm.spec,incubator/binding-llm — NetworkIT, ApplicationIT, LlmServerIT, LlmClientIT, LlmProxyIT all green (100 tests, 0 failures)
  • ./mvnw clean install -DskipITs across the full reactor
  • Manually exercised examples/llm.proxy's model-based routing scenarios via etc/test/verify.sh

🤖 Generated with Claude Code

https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG


Generated by Claude Code

Scaffold incubator/binding-llm.spec per AGENTS.md conventions and define
LlmBeginEx, LlmDataEx, and the LlmFlushEx union, modelled on
binding-mcp.spec's idl.

LlmBeginEx carries dialect only; model routing is deferred. LlmDataEx has
no fields: content flows through the DATA frame's own payload octets and
INIT/FIN through its existing flags, so nothing survives in the extension
once block identity moves to the FLUSH plane. LlmFlushEx is a 7-case union
covering message start, block start/end, finish, usage, keepalive, and an
opaque native/raw case for re-encoding events a same-dialect route doesn't
recognize.

Fixes #2476

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AuZoMsETwEczJx3cbkb8EJ
Scaffolds incubator/binding-llm and incubator/binding-llm.conf, modelled
on binding-mcp's SERVER/CLIENT BindingContext structure. LlmBindingInfo
is annotated @Incubating so type: llm config loading is gated behind
ZILLA_INCUBATOR_ENABLED via FeatureFilter, matching the AmqpBindingInfo/
PgsqlBindingInfo/RisingwaveBindingInfo precedent.

Fixes #2477

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0142buJWS7C89AKr9uDJtSy4
…t-type

Registers by content-type and hands back a per-stream LlmContentDecoder;
stays in an internal, unexported package for now with no concrete
implementation registered yet.

Fixes #2478

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mw32oxEw24fLt5Ypakj6pH
Closes the non-streaming half of #2478's own scope: "Non-streaming
application/json goes through the same abstraction as one event, rather
than a special-cased branch." Only text/event-stream had an
LlmContentDecoderSpi implementation; application/json requests
(non-streaming dialect responses) had no decoder to dispatch to.

LlmJsonContentDecoder treats the entire buffered document as a single
event (one data + one flush call, no framing loop), mirroring
LlmSseContentDecoder's structure and unit-test conventions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mw32oxEw24fLt5Ypakj6pH
…odecSpi

LlmContentDecoderSpi and LlmContentEncoderSpi (the latter previously
living downstream in #2571) let a content-type register a decoder with
no matching encoder, or vice versa, since each was ServiceLoader-discovered
and dispatched independently. LlmContentCodecSpi makes "this content-type
is fully supported" one type-enforced fact: a single contentType() key
with both supplyDecoder() and supplyEncoder(), one META-INF/services
registration per content-type, one LlmContentCodecFactory dispatching
both directions.

Pulls the internal/encode/ base package and its text/event-stream
implementation (LlmContentEncoder, LlmSseContentEncoder) forward from
#2571 so the collapse can happen where decode already lives, rather than
forking that package ahead of its own introduction there; #2571 will
need to rebase on top of this and drop its now-duplicate copies.

LlmSseContentDecoder, LlmSseContentEncoder, LlmJsonContentDecoder widen
from package-private to public (unchanged otherwise) since their new
LlmSseContentCodecSpi/LlmJsonContentCodecSpi providers construct them
from the sibling internal.codec package. Adds LlmJsonContentEncoder
(new): the application/json inverse of LlmJsonContentDecoder, copying
content bytes through unchanged with no framing on flush.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mw32oxEw24fLt5Ypakj6pH
Defines the pluggable-dialect contract for binding-llm, exported from the
start (unlike LlmContentDecoderSpi, which stays internal): LlmDialect
exposes name()/detect()/contentType() plus supplyDecoder(Kind)/
supplyEncoder(Kind) returning common-json JsonTransform stages, and
LlmDialectFactorySpi is the ServiceLoader-registered entry point.
HttpHeaders is a minimal read-only accessor for detect(path, headers),
since no HTTP header abstraction previously existed in this codebase and
pulling in jakarta.ws.rs would add a dependency never otherwise used here.
Kind is nested on LlmDialect, distinguishing request/response schemas.

No concrete dialect implementations yet (OpenAI/Anthropic land later) --
module-info.java exports the dialect package and declares uses without a
corresponding provides. Unit-tested via a stub LlmTestDialect/
LlmTestDialectFactorySpi registered under test-scope META-INF/services,
mirroring this module's existing LlmContentDecoderSpi/
LlmTestContentDecoderFactorySpi pattern.

Fixes #2480

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NcAVwRPN1Pwjzpobr75w6
…er instance

contentType() previously took no parameters, so a dialect could only
report one fixed content-type for its lifetime -- insufficient for an
API whose response framing (event-stream vs. a single JSON document)
depends on a flag in the request body, since neither contentType() nor
detect(String, HttpHeaders) offered any way to inspect it.

Adds HttpRequestBody, a minimal read-only scalar-member accessor
mirroring HttpHeaders, and changes contentType() to
contentType(Kind, HttpHeaders, HttpRequestBody): Kind lets request and
response resolve independently (a dialect's request body content-type
can be fixed while its response varies), and the headers/body context
lets that resolution depend on the actual request rather than being
fixed at dialect-instance-creation time. Both parameters are nullable
for callers without that context available.

LlmTestDialect now resolves text/test-event-stream for a streaming
response and application/test+json otherwise, exercising the new
per-Kind, per-request resolution the stub previously couldn't express.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NcAVwRPN1Pwjzpobr75w6
…velope, not whole-body buffering

Replaces LlmDialect's HttpHeaders/HttpRequestBody/contentType() with the
engine's existing ModelEnvelope/ModelTransform (runtime/engine/.../model/),
reusing machinery this codebase already has instead of inventing
LLM-specific buffering to resolve text/event-stream vs. application/json,
or a model name, by peeking one field of an otherwise-unbuffered body.

detect(ModelEnvelope) folds :path/:method into the same envelope as
ordinary headers -- no separate path parameter. supplyDecoder/supplyEncoder
now take (Kind, ModelEnvelope) and return ModelTransform: a per-field
stage that can extract a field (e.g. a model name) into the envelope while
the body still flows through unchanged, mirroring
KafkaExtractTransform (runtime/binding-kafka/.../cache/) -- so a caller
reads that signal back off the envelope as decoding proceeds rather than
buffering the whole body first to inspect it. contentType() is removed
entirely: nothing in this shape needs it once the streaming/non-streaming
signal is just another envelope entry a caller reads after extraction.

common-json/JsonTransform is no longer used anywhere in this module now
that the dialect SPI itself doesn't need it, so the dependency comes back
out of module-info.java and pom.xml along with it.

LlmTestDialect/LlmTestDialectFactorySpi become LlmTestConditionalDialect/
LlmTestConditionalDialectFactorySpi, since what they now demonstrate is
exactly this: request detection and model-name extraction conditional on
envelope contents, not a fixed per-instance answer.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016NcAVwRPN1Pwjzpobr75w6
…ken detail

Extend LlmFlushEx's block-lifecycle skeleton (from #2476) with the two
places OpenAI's parallel completions and per-token detail need to survive
in the vocabulary, per the issue's scope:

- choiceIndex (default 0) on messageStart, blockStart, blockEnd, finish,
  and native/raw: the (choice, block) compound index's outer half.
  Anthropic is always choice 0, so every existing dialect mapping is
  unaffected; OpenAI's n > 1 becomes one messageStart per parallel
  completion, distinguished by choiceIndex. usage stays choiceIndex-free
  since every dialect reports it aggregated across choices, never per
  choice.
- logProbability (nullable) on LlmDataEx: the one per-delta detail the
  vocabulary carries directly, for dialects exposing per-token detail
  (OpenAI logprobs) without reopening the DATA/FLUSH split from #2476 or
  growing the vocabulary for the full log-probability structure — richer
  detail than one value per token stays behind LlmNativeFlushEx.

Documents the three lossiness cases as doc comments alongside the fields
they concern: choiceIndex and logProbability both drop out on any
cross-dialect route to Anthropic (structurally exactly one choice, no
per-token detail); message-start input token counts are resolved by
decoupling inputTokens into its own deferred usage event rather than
emitting a placeholder on messageStart and correcting it later, so a
source that discloses tokens late (OpenAI) just emits usage late instead
of needing a correction event.

Fixes #2481

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AYGZDrytLjmG2AQqUN1hsu
…lects

Translates between each dialect's native streaming event sequence and the
canonical vocabulary from #2557, in both directions:

- LlmAnthropicEventMapper: holds input_tokens from message_start until the
  paired usage event at message_delta; tracks the currently open block's
  type to know when content_block_stop needs a canonical blockEnd (tool
  calls only) and to route content_block_delta payloads (text_delta vs
  input_json_delta) on encode.
- LlmOpenAiEventMapper: translates OpenAI's tool-call-only index space into
  the canonical (Anthropic-shaped) block index via a per-stream map, and
  synthesizes blockEnd lazily -- deferred until the next tool call starts
  or the stream finishes, since OpenAI has no explicit block-close event.

Unit-tested against both worked-example tables from the issue (message
role/content/tool-call cardinality changes in each direction), plus the
held-usage/already-consumed and lazy-blockEnd edge cases.

Fixes #2482

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019BHBkLjV2tcbNxMEgxkpSw
LlmDialectResolver dispatches path/header detection across every LlmDialect
registered via LlmDialectFactorySpi, without hardcoding any dialect's
signals. A configured fixed dialect name bypasses detection entirely,
including when it matches no registered dialect. When detection matches
more than one dialect, or none, resolution is ambiguous and returns null
so the caller rejects the request rather than guessing.

LlmOptionsConfig adds the optional server-kind `dialect` option (schema,
config, and adapter) used to pin a fixed dialect.

Fixes #2483

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AtkteS7C66qiVLGeEDJ2FP
The sys: namespace's patch contract (Binding.system()/Exporter.system())
lets a component contribute a shared binding (e.g. http_client) but has
no pre-seeded slot for a shared catalog, so a patch adding one has
nothing to append into. Pre-seed catalogs: {} alongside the existing
bindings: {} in the base sys namespace skeleton, the same "pre-seed the
extension point" convention already used for the JSON-schema *-ext
scaffolds, so any component can share a catalog-backed resource the way
bindings are already shared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
…ect schemas

Adds LlmDialectFactorySpi.schema(Kind), letting a dialect contribute a URL
to its own JSON schema for the request or response direction without
knowing anything about catalogs or patches. LlmBinding.system() enumerates
every registered dialect, reads each contributed schema, and generates a
sys: namespace patch adding one shared inline catalog (llm_dialects) with
a <dialect>.request/<dialect>.response subject per schema a dialect
contributes -- built once, at engine startup, since the dialect set is
ServiceLoader-discovered off the classpath and therefore fixed for the
JVM's life, the same way sys: already shares a binding (e.g. http_client)
across every binding that references it.

LlmDataUrlStreamHandler decodes a base64 data: URL (RFC 2397) in memory,
scoped to a single URL via the URL.of(URI, URLStreamHandler) factory --
no temporary file, no globally-registered protocol handler -- used to
hand the generated patch to Binding.system()'s URL-returning contract
without writing it to disk.

LlmBeginEx gains contentType and model fields alongside dialect, for a
server-kind stream factory to stamp once it has resolved the dialect and
read the request's content-type.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
… LlmBeginEx.dialect

Implements request/response framing decode-and-forward for the LLM server
binding (#2484): LlmServerFactory drives the request body
through the resolved dialect's ModelPipeline as bytes arrive off the wire,
forwards transformed content to app0 incrementally rather than buffering
the whole body, and threads the resolved dialect name onto LlmBeginEx so
app0 can see which wire dialect produced the request.

Flow control between the client, this binding, and app0 is enforced with
dedicated decodeSlot/encodeSlot buffers on each side of the exchange
(LlmServer for the network-facing leg, LlmStream for the app-facing leg),
each granting credit strictly from its own local slot occupancy rather than
copying a peer's sequence numbers across independent byte domains. LlmState
tracks per-direction open/closing/closed transitions and end-of-stream
deferral while a buffer is still draining.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
…fferPool handles

DefaultBufferPool.buffer(slot) rewraps and returns a single shared mutable
field per pool instance, so holding a buffer reference across a nested call
that fetches a different slot from the same pool silently repoints it.
LlmServer.decodeNetwork() directly, synchronously calls LlmStream's
request-relay staging method as a plain nested Java method call, so a
single pool instance backing both the network-decode slot and the app0
request-relay slot would alias between them.

Give LlmServerFactory two BufferPool handles instead of one: decodePool
(network decode) and encodePool (app0 relay, both directions), obtained via
context.bufferPool() and .duplicate() -- matching McpServerFactory's
decodePool/encodePool precedent. The two relay directions sharing
encodePool (the reply-direction slot on LlmServer and the request-direction
slot on LlmStream) never appear in the same call stack: cross-binding
accept() is ring-buffer-mediated and dispatched on a later engine tick, not
a nested call, so one encodePool instance can't alias between them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
…nline.conf

LlmSystemNamespaceGenerator emits a real "type": "inline" catalog config
into the generated sys: namespace patch, serviced at actual runtime by
catalog-inline's CatalogFactorySpi via ServiceLoader -- not just exercised
by this module's own tests. test scope kept it off the runtime classpath
entirely; provided scope wouldn't fit either, since nothing in main source
compiles against catalog-inline's Java API (the reference is a plain
string), so there's nothing to satisfy at compile time.

Match binding-asyncapi/binding-openapi's precedent for this same
generated-"type: inline"-config pattern: catalog-inline at runtime scope,
plus the companion catalog-inline.conf (config-schema side) at default
scope alongside the existing model-json.conf dependency.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
LlmBindingConfig.newModelConfig() constructs a real JsonModelConfig in main
source, resolved to a working ModelHandler by context.supplyModel() via a
ModelFactorySpi lookup at actual runtime -- not just exercised by this
module's own tests. Matches binding-asyncapi/binding-mcp-openapi/
binding-openapi, each of which also builds JsonModelConfig directly in main
source and declares model-json at runtime scope.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
Stale relative to the catalog-inline/model-json scope fixes on binding-llm
(catalog-inline.conf and model-json.conf now roll up into this aggregate),
plus a pre-existing gap for binding-http.spec's license entry. Regenerated
via ./mvnw notice:generate -pl incubator -amd, never hand-edited, per
AGENTS.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
Add LlmServerConfig/LlmServerConfigBuilder (host/port, nested builder)
following the existing *OptionsConfigAdapter pattern used by
binding-kafka's options.servers, wired into LlmOptionsConfig via a new
`server` field so `llm client` bindings can configure their upstream
endpoint. Config adapter unit tests cover parsing and serializing
options.server.

Fixes #2485

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FQ4XvYNCp1NzQC4Ne4eWbK
Exercise the inject() path on LlmServerConfigBuilder so the nested
server builder reaches the module's required 100% instruction
coverage, mirroring the existing shouldInjectBuilder test for
LlmOptionsConfigBuilder.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FQ4XvYNCp1NzQC4Ne4eWbK
…→encoder chaining

Adds LlmClientFactory (kind: client), the first real integration point for
LlmDialect.supplyDecoder/supplyEncoder: it compares the inbound app-declared
dialect (LlmBeginEx.dialect) against the binding's own configured dialect and
only chains a JsonPipeline payload transform when they differ, forwarding
framing-only re-encoded content when they match.

- internal/encode: LlmContentEncoder/Spi/Factory + LlmSseContentEncoder, the
  SSE-framing encode counterpart to internal/decode's existing SSE decoder
- LlmDialectResolver.dialectNamed(String): by-name lookup for resolving the
  inbound dialect when it differs from the client's configured one
- Schema patch: kind: client options (dialect, server); also fixes kind:
  server to accept options.server (previously unreachable under its
  additionalProperties: false, despite LlmOptionsConfig already supporting it
  since PR 2570)
- Spec scripts (same.dialect, cross.dialect, client.opaque.fallback,
  client.abort) plus a new LlmClientIT, written first per this repo's
  test-first discipline, confirmed failing before LlmClientFactory existed

Fixes #2486

Real dialect implementations (openai, anthropic), the mock backend, and the
full round-trip identity test are separate, sibling issues (#2487, #2490,

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
…aders, body)

Rebased onto #2570's latest, which changed LlmDialect.contentType() to take
(Kind, headers, body) parameters, resolved per request rather than fixed
once per dialect instance. The client resolves content-type separately for
each direction now (REQUEST for its outbound encoder, RESPONSE for its
inbound decoder) rather than a single shared call, matching the interface's
own point: request and response can have different native content-types.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
LlmClientFactory.newStream resolved the RESPONSE content decoder eagerly at
BEGIN time with body=null, so LlmDialect.contentType(Kind.RESPONSE, ...)
could never actually see request-body content when deciding between a
streaming and non-streaming response, defeating the per-request capability
added for that method. REQUEST-side encoder resolution is unaffected since
it does not depend on body content.

Accumulate the request body (post cross-dialect transform, pre-framing) into
a buffer-pool-backed slot as DATA arrives, and resolve the decoder at
onAppEnd, once the full request is available, relying on this binding's
half-duplex transmission convention to guarantee no response bytes arrive
before then. Add LlmJsonRequestBody, a pure-Java HttpRequestBody view driven
by common-json's one-shot parser API, plus a unit test. Add a new
test-conditional dialect and two client IT/spec scenarios proving the
decoder now differs based on a stream field in the request body.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
…line SPI

LlmDialect no longer exposes contentType(Kind, HttpHeaders, HttpRequestBody);
dialects now detect via ModelEnvelope and supply ModelTransform-based
decoders/encoders driven through the engine's ModelHandler/ModelPipeline SPI.
Content-type is read directly off real wire headers instead of being computed
per-dialect, so the request-body buffering added to defer response-decoder
selection (LlmJsonRequestBody) is no longer needed and is removed.

LlmClientFactory is rebuilt against the new SPI: dialect resolution via
LlmBindingConfig.resolveDialect/dialectNamed, a per-stream ModelEnvelope, and
a ModelPipeline (ModelTransform.NONE for same-dialect) run unconditionally in
both directions so downstream code always sees validated, well-formed
payloads. Test dialects are renamed (test-client, test-client-sse,
test-client-sse-alt) to avoid colliding with the real upstream test fixtures,
and gain request/response JSON schemas so the pipeline can actually validate
their payloads instead of rejecting everything for lack of a schema.

Adds the missing llm:flushEx()/llm:matchFlushEx() k3po functions that the
client k3po scripts already relied on.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
…sponse transforms

Implements LlmDialect for OpenAI Chat Completions, registered via
LlmDialectFactorySpi: detect() matches POST /v1/chat/completions and
POST /v1/completions; contentType() is text/event-stream, the only
LlmContentDecoderSpi/LlmContentEncoderSpi framing codec this module
has today (a non-streaming application/json pair is a content-decoder-
layer gap, not a dialect one, and is left as follow-up).

supplyDecoder/supplyEncoder rename OpenAI-native request/response JSON
members to the canonical vocabulary this dialect defines a synonym
for -- max_tokens/maxOutputTokens, top_p/topP, n/choiceCount and
friends on requests; index/choiceIndex, finish_reason/finishReason
(remapping tool_calls/tool_call), logprobs/logProbability, and the
usage token counts on responses -- via a depth-tracked JsonTransform
that renames JSON events without DOM parsing. Everything without an
established canonical synonym (id, model, messages, tools, the whole
delta/tool_calls structure including streamed function.arguments
fragments) forwards unchanged, at any depth, so decode -> encode
round-trips with no loss.

Tests cover dialect detection/registration and both Kinds of
transform, including full round-trips through the canonical form.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
… rename

model-json's ModelTransform integration was observation-only (ModelFieldBridge),
discarding REPLACED/DECLINED answers instead of writing them. JsonModelFieldTransform
drives a wired ModelTransform inline as JSON streams through, computing a JSON-pointer
path per scalar field (with array-index segments) at any nesting depth, and writes
FIELD/REPLACED/DECLINED answers straight to the destination -- including a REPLACED
substitute redirecting a field to a sibling key of the same enclosing object.

Adds the missing key-write for container-valued members entering a named object
member, fixes a resumed key/value write re-offering the whole text instead of the
remainder (TextSource now tracks its own consumed() progress), and reuses the
already-decoded scalar text/key for an unchanged FIELD answer instead of a fresh
per-field allocation.

Also fixes JsonModelHandlerImpl.supplyEncoder silently dropping its transform
parameter.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
…e SPI

Rebases LlmOpenAiDialect and its request/response transforms onto the
redesigned LlmDialect SPI: detect(ModelEnvelope), no more contentType()
(content-type is now resolved from real upstream Content-Type headers),
supplyDecoder/supplyEncoder(Kind, ModelEnvelope) returning a ModelTransform.

The request/response transforms are rewritten as plain ModelTransform
implementations matching each field's own full path (e.g. $.choices[0].index,
$.usage.prompt_tokens) rather than tracking JSON structural depth, since the
model-json adapter now computes paths itself. This drops the old
JsonEvent-token depth-tracking machinery and the JsonSource/JsonController
wrappers (LlmOpenAiStructuredController is no longer needed); the renamed
LlmOpenAiSubstitutedSource is now a plain ModelSource.

The choices[]/usage direct-member path checks defer their substring() until
after confirming the path is actually a direct member, so the many deeply
nested per-chunk fields that share the prefix (delta.tool_calls[].index and
the like) cost no allocation.

The one dropped rename from the prior implementation is logprobs/logProbability:
it names a container-valued field, and this dialect only renames scalar leaves.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
…openai dialect

LlmOpenAiRequestTransform renamed a fixed set of top-level fields but never
observed model, so LlmServerFactory.doAppBegin's server.envelope.get("model", 0)
read right after running the request through this transform came back empty and
LlmBeginEx.model was never stamped for real dialect: openai traffic.

Mirrors LlmTestPermissiveDialect's inline ModelExtractTransform (and
KafkaExtractTransform's own pattern): on a FIELD event at $.model, copy the
value into the envelope alongside the existing rename-or-forward decision,
without touching the RENAMES table. LlmOpenAiResponseTransform needs no
equivalent change -- nothing reads model back off a response.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
…requests

CI's LlmServerIT.shouldRejectRequestWithUnresolvedDialect failed: detect()
matched on :method/:path alone, so it ambiguously co-matched every existing
k3po server fixture -- all of which hit the same /v1/chat/completions path
with a test-specific content-type (application/vnd.zilla.test-permissive+json,
application/vnd.zilla.test-strict+json) to select a *different* dialect
unambiguously. Only one failure surfaced in CI because failsafe stops after
the first failure, but the same ambiguity affected every other fixture in
that class too -- confirmed by LlmServerIT going from 6 run/1 failed/1
skipped to 8/8 passing with this fix, and LlmClientIT unaffected at 4/4.

detect() now also requires content-type: application/json, the only
content-type a genuine OpenAI request ever carries, so a request to the same
path with a different dialect's own content-type no longer collides.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
The llm.schema.patch.json allowed options.server on kind:server bindings,
but LlmServerFactory never reads binding.options.server — only
LlmClientFactory dials it as the upstream endpoint. Restrict the
kind:server options schema to dialect (fixed-dialect mode), matching
actual runtime usage and the milestone's example configs.

Add LlmSchemaValidationTest exercising the full EngineConfigReader
pipeline against real zilla.yaml text for both kind:server and
kind:client, covering acceptance (bare server, fixed-dialect server,
full client) and rejection (server option on kind:server, missing
required client fields, malformed server pattern, unknown kind).

Fixes #2488

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aawd4ZFaUE3FgQnzkABj9F
…d client ITs

The two new client-side forwarded-credentials scenarios
(openai.request.guarded / anthropic.request.guarded) drive the caller's
already-authorized session by connecting to app0 with a nonzero
zilla:authorization option. LlmClientFactory forwards that same
authorization value unchanged onto the outbound stream to net0. Since
net0 is k3po's own "external" accept (not a real engine binding), its
zilla:// transport rejects the connection with a RESET when the connect
and accept sides carry mismatched authorization values -- so the
network server script's accept needs the same option as the app-side
connect for the scenario to reach the point where it can actually
assert the injected credentials header.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BG4hRScwRVV93odJitLfNs
…skeletons

Sets up the directory layout, docker-compose stack, and zilla.yaml wiring
chained llm(server)+llm(client) bindings both directions (openai-facing
frontend to an anthropic backend, and anthropic-facing frontend to an
openai backend), with inline-guard credential pass-through. Mock backend
server implementations and CI verify.sh assertions land in a follow-up
commit.
…inding-llm

Adds the mock OpenAI and Anthropic backend servers, the etc/test/verify.sh
CI assertions (both translation directions, tool-call round trips, and
both credential pass-through directions), and the example README.

Also adds binding-llm to cloud/docker-image's dependencies and
zpm.json.template -- it was missing from the docker image build, so the
example's zilla.yaml (type: llm) would fail schema validation against the
published image.
… mock responses

llm(server)/llm(client) sit on top of http, the same way mcp does in
examples/mcp.proxy -- they don't decode/encode HTTP/1.1 framing
themselves. Insert http(server) between each tcp(server) and llm(server),
and http(client) between each llm(client) and its tcp(client).

Also fixes the mock backends: the llm(client)'s outbound request always
targets "/" (PATH_DEFAULT in LlmClientFactory -- no per-dialect path is
modeled yet), and Express's res.json() sends a "charset=utf-8" parameter
that breaks the llm binding's exact-match response content-type lookup.
Both mocks now listen on "/" and set the response Content-Type header
explicitly without a charset.
onRequestBegin granted the request window unconditionally, before the
network transport had actually connected. A request whose first DATA
frame arrives before the connect completes (e.g. because the exit
binding streams it eagerly right after BEGIN) could reach the network
write path while the connection was still pending, which collapsed
into the abort-on-pending-connect path in doNetShutdownOutput and
silently dropped the request.

Grant the window when the transport is already open, and otherwise
defer it to flushNext(), which runs once onNetworkBegin confirms the
connection is ready. This uses the existing flow-control mechanism to
prevent the premature send rather than buffering the request data.
LlmClient.onAppBegin granted the application window as soon as the
stream opened, without waiting for the network transport (delegated to
an http client) to actually be ready. Mirrors the binding-http fix:
the window is now granted once LlmHttpClient.onNetWindow reports the
transport has become ready, using flow control to prevent a request
from arriving before the connection can accept it.
LlmContentCodecFactory looked up decoders/encoders by exact
content-type string match, which fails against a real-world
Content-Type header carrying parameters (e.g. "application/json;
charset=utf-8"). A missed lookup silently falls back to raw
passthrough, skipping dialect translation entirely. Strip any
";"-separated parameters and match on the bare media type.
LlmServer's initialAuthorization was overwritten on every DATA frame
with the raw per-frame authorization instead of keeping the
guard-authorized session value captured at construction, so the exit
stream never carried a valid guard session. This breaks the
authorization-forwarding contract llm(client)'s
options.authorization pass-through depends on: it resolves
guard.credentials(authorization) only when the inbound stream's own
authorization is itself a valid guard session.

Threads authResult.authorization() through the LlmServer constructor
as a final field instead.

Add paired k3po scripts (server.openai.request.guarded,
server.anthropic.request.guarded) for the exit-side stream carrying a
non-zero authorization, since the existing openai.request.guarded /
anthropic.request.guarded scripts are owned by the client-side
scenario and expect a different request shape. Point
LlmServerIT's guarded tests at the new scripts -- they previously
referenced the unguarded server script, which only happened to work
because the authorization value being forwarded was always 0.
llm server routes purely on authorization today -- LlmRouteConfig's own
comment noted there was no routes[].when condition schema. LlmBeginEx
already carries both dialect (resolved from headers at connect time) and
model (extracted from the request body via the existing bounded
LlmModelExtractTransform), so route selection just needed a condition
matcher over those two already-populated fields, mirroring mcp(proxy)'s
toolkit/tool route matching.

Adds LlmConditionConfig (dialect exact-match, model glob allow-set) and
its adapter/schema patch for kind: server routes, plus LlmConditionMatcher
and a dialect+model overload of LlmBindingConfig.resolve() on the runtime
side. Since model is only known once request body bytes have started
decoding, the exit route is now resolved at doAppBegin() time (after the
model has had a chance to be extracted) rather than eagerly at connect
time; doAppBegin returns false and the request is reset when no route
matches.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…tchers

Routing on dialect/model belongs to the llm(proxy) hop described in #2502,
not to llm(server) itself. llm(server) extracts model from raw request
bytes as it decodes, so giving it its own routes[].when meant deferring
route resolution until mid-decode (doAppBegin()) -- which also introduced
a stream/app-lifecycle bug (onNetAbort/onNetReset assumed stream != null
implied the app stream was open). llm(proxy) doesn't have that problem:
it receives an already-populated LlmBeginEx from an upstream llm(server)
and can resolve dialect+model synchronously at newStream() time, same as
any other condition-matched proxy binding.

Reverts LlmServerFactory/LlmServerIT/the routes[].when schema patch under
kind: server, and the k3po scripts that exercised server-level routing.
Keeps LlmConditionConfig/LlmConditionMatcher/LlmBindingConfig.resolve(
authorization, dialect, model) -- kind-agnostic and needed by the actual
llm(proxy) implementation.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
llm(proxy) is a third kind alongside server/client: it accepts a BEGIN
whose extension is already a populated LlmBeginEx -- produced upstream by
an llm(server) (or another llm(proxy)) that already did the model
extraction -- and resolves its exit route synchronously in newStream()
from LlmBeginEx.dialect()/model(), mirroring how mcp(proxy) and
TlsProxyFactory resolve routes from BEGIN-time data. No content is ever
decoded or re-encoded: DATA/FLUSH/END/ABORT/RESET/WINDOW/CHALLENGE relay
verbatim (including extensions) between the accepted network stream and
the resolved exit once picked, same-protocol-proxy shape (LlmServer +
LlmClient inner classes) per runtime/AGENTS.md.

Reuses the condition-matching infrastructure already in place
(LlmConditionConfig/LlmConditionMatcher/LlmBindingConfig.resolve) --
unchanged from the earlier attempt at llm(server)-level routing, which
was the wrong layer and has been reverted. Adds "proxy" to the kind enum
and a routes[].when.{dialect,model} schema block scoped to kind: proxy;
wires LlmProxyFactory into LlmBindingContext's kind->factory map.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
Both frontends' llm(server) bindings now exit into a shared
north_llm_proxy, which routes purely on LlmBeginEx.dialect/model. Two
new routes send gpt-4o/claude-3-5-haiku-20241022 to a second
same-dialect deployment instead of the default cross-dialect
translation, demonstrating the proxy's intra-dialect model-based
routing alongside the existing translation demo. Adds
mock-openai-secondary/mock-anthropic-secondary compose services for
CI, verify.sh assertions for both new routes, and a README section
showing how to swap any llm(client) leg for the real OpenAI/Anthropic
API with a caller's own key, following the same mock-for-CI pattern as
mcp.proxy.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
Full ./mvnw install surfaced a jacoco coverage regression: this module
requires 100% instruction coverage, and LlmConditionConfig's
Function-mapper builder overload plus LlmConditionConfigBuilder's
thisType() were both added but never exercised by a test. Cover them
the same way sibling *ConditionConfigAdapterTest classes do -- inject(
identity()) to hit thisType(), and a direct call to the mapper-based
builder() overload.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…end until decode flushes

Live end-to-end testing of the llm.proxy example (a real docker-built
zilla image, not just k3po) surfaced two bugs neither the config-level
tests nor the existing IT payloads were large/slow enough to expose:

- LlmProxyFactory's window relay plugged the wrong side's sequence
  counter into the ack computation on both directions -- numerically
  equivalent at idle, but wrong in general once real traffic diverges
  the two independently-numbered streams. Fixed to match
  ProxyServerFactory's pattern: derive an intermediate window-size
  from the granter's own seq/ack/max, then re-anchor the new ack
  against the receiver's own sequence.

- LlmClientFactory.onNetEnd unconditionally discarded any
  backend-response bytes still sitting in decodeSlot (buffered because
  they arrived before reply-window credit did) and closed the reply
  immediately -- silently dropping a fully-received response. Defer
  the close instead, via the same LlmState.deferReplyEnd/
  replyEndDeferred bits LlmServerFactory already uses symmetrically on
  its own encode side, and complete it once decodeNet actually drains
  the slot.

Verified against a locally built docker image (not just k3po): all
four llm(proxy) routing scenarios, both cross-dialect translations,
both tool-call round trips, and both credential pass-through
directions now return full bodies end-to-end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…/closed

LlmClientFactory's decode-side deferred-end tracked a single object's
own reply direction (this class's own state), unlike LlmServerFactory's
encode-side deferReplyEnd/replyEndDeferred, which tracks a separate
object's (the network-side encode buffer's) drain state independent of
the app-stream's own lifecycle -- a genuinely distinct condition that
needs its own bit. Here, "end arrived but not yet flushed" is exactly
what closing-without-closed already means on the same state variable,
so add LlmState.replyClosing() to query the existing bit instead of
introducing a second, overlapping one.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
Building blocks reusable across LlmServerIT and LlmClientIT: for each
real dialect (openai, anthropic), a network/ pair (http:beginEx, the
real API's own path/content-type) and an application/ pair (llm:beginEx
with that dialect), each carrying a 10k blob in both the request and
response body. LlmServerIT pairs ${net}/client + ${app}/server per
dialect (both pass, same shape as its existing test-permissive 10k/100k
coverage). LlmClientIT recombines ${app}/client + ${net}/server across
dialects -- same-dialect (openai+openai, anthropic+anthropic) and
cross-dialect (anthropic app + openai net, and the reverse) -- without
needing dedicated "X.to.Y"-named scripts, since routing which script
pairs with which is entirely a Java-side @Specification choice.

The four new LlmClientIT tests currently fail: tracing shows
LlmClientFactory's request-encode path (onAppEnd/onAppFlush) sends the
entire encoded request in one doNetData call sized to the full content
length, regardless of the actually granted initialMax window --
confirmed via instrumentation as length=13410 against initialMax=8192.
No existing test sends a request over 8192 bytes through llm(client),
so this window violation was previously unreachable. LlmServerFactory's
response-encode path already solves the mirror-image problem via its
encodeSlot/SlotChunk queue + flushEncodeSlot, re-sliced against
replyMax on every write and every window grant; LlmClientFactory's
request-encode path has no equivalent for its own initialMax. Fix
tracked as a follow-up commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…entFactory's pattern

onNetEnd now attempts an immediate decode-slot drain (renaming
resumeDecode to a decodeNet(traceId, authorization) overload, matching
the two-decodeNet-overload + onNetEnd-calls-decodeNet-directly shape
already used by TlsClientFactory/McpClientFactory) instead of only
waiting on a later onAppWindow-triggered resume. Both the immediate
attempt and the later one (if decode is still pending when onNetEnd
runs) now route through a single doAppEnd(traceId, authorization)
helper guarded by !replyClosed(state), so whichever path actually
drains the slot is the only one that fires the app-facing end.

No behavior change for any passing scenario -- LlmServerIT still
21/21, and LlmClientIT's pre-existing 21 tests are unaffected. The 4
new 10k tests still fail identically to before this change, against
the separate, already-diagnosed request-encode window violation (not
yet fixed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
LlmClientFactory.LlmHttpClient.doNetData() sent the entire encoded
request body in a single doNetData() call regardless of the net
stream's currently granted initialMax, violating the window contract
whenever encoded content exceeded one window's worth of bytes (e.g.
10k+ request bodies against an 8192-byte window) -- confirmed via
HttpClientFactory's own onRequestData reset/abort on window violation.

Adds an encodeSlot on LlmHttpClient (net-facing only, mirroring
McpClientFactory's HttpStream encode pattern): doNetData() appends
onto the slot's tail when one is already draining, otherwise encodes
straight from the caller's buffer; encodeNet() writes only as much as
initialWindow currently allows and buffers any remainder, resumed by
onNetWindow on every subsequent grant (not just the first). doNetEnd()
now defers the actual net END until the slot fully drains, guarded by
the same closingInitial/closeInitial state bits (no new deferred-flag
pair, since this is single-object state). A dedicated encodePool
(duplicate of the shared bufferPool) avoids aliasing the encodeSlot's
buffer against LlmClient's own decodeSlot buffer live in the same
call stack, matching LlmServerFactory's existing decodePool/encodePool
precedent.

pendingContent (LlmClient) is unchanged and still required: it
accumulates content between SSE event boundaries for the request
direction's raw/native flush framing (same.dialect/cross.dialect
scenarios), which is a distinct concern from net-side window credit.

Verified: LlmClientIT's new openai.10k/anthropic.10k request bodies
now reach the backend intact (previously hung on the window
violation). This surfaces a separate, pre-existing bug in the same
class's response-decode direction: LlmJsonContentDecoder.decode()
documents that it expects one complete buffered document per call,
but LlmHttpClient.decodeContent() truncates each call to the current
reply window, so a response spanning multiple network reads (10k+)
is decoded as several bogus "complete documents" -- not yet fixed,
tracked as the next item.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…ontrol

LlmProxyFactory's doNetWindow/doAppWindow computed the relayed ack as
initialSeq - minInitialWin (mirroring binding-proxy's ProxyServerFactory),
which only advances correctly when the peer's advertised maximum grows
monotonically. Once content exceeds a single window and the peer instead
holds maximum steady while advancing acknowledge (as k3po's own accept/
connect roles do), the relayed ack never moves and the stream stalls after
the first window's worth of data.

Fix the formula to derive the relayed ack from the peer's in-flight delta
(minInitialMax - minInitialWin) rather than the raw window value, which
tracks correctly under both conventions.

Add LlmProxyIT flow-control coverage (10k/100k, openai/anthropic, routed
to both app0 and app1) that exercises payloads spanning multiple windows
through llm(proxy) and caught this bug; the existing small-payload proxy
routing tests never exceeded a single window. Reuses the application
openai/anthropic 10k/100k server scripts via a new serverAddress k3po
property override, mirroring mcp(proxy)'s pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
@jfallows
jfallows force-pushed the claude/serene-ride-hlw83v branch from c44c62e to e2fc451 Compare September 18, 2026 23:09
…lm binding

The llm binding (client/server/proxy kinds, dialect enum, proxy routes[].when)
is now part of the merged zilla.yaml JSON Schema, but the checked-in golden
file predates its addition. Regenerated via the documented process: built
the develop-SNAPSHOT Docker image from source and ran
`docker compose exec zilla zilla inspect schema` against it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…d message_start

LlmAnthropicEventMapper.encodeMessageStart built the outbound Anthropic
message_start event with only id/type/role/model, omitting content,
stop_reason, and stop_sequence -- fields the real Anthropic API always
includes and that the official anthropic-sdk-python's streaming
accumulator depends on to initialize its per-message state. Without
content: [], the SDK crashed on the first content_block_start/delta with
'NoneType' object has no attribute 'append', surfaced while exercising
examples/llm.proxy's cross-dialect streaming translation through the real
SDK instead of hand-built assertions.

usage stays intentionally omitted here, per the existing comment on
LlmUsageFlushEx: input token counts are decoupled from message_start so a
dialect that only discloses them later (OpenAI) doesn't need a
placeholder zero corrected afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…eal SDKs

Replace etc/test/verify.sh's hand-built curl+grep assertions with
etc/test/verify.py, driven through the official openai and anthropic
Python SDKs against the example's openai-facing and anthropic-facing
frontends respectively. Proves the dialects are wire-compatible enough
for an off-the-shelf client to succeed unmodified, not just compatible
with our own request/response shapes -- and, in doing so, caught the
message_start gap fixed in the preceding commit that curl+grep never
would have (an official SDK enforces the real wire contract; a substring
match on a curl response does not).

Covers both dialects' default (cross-dialect translation) and secondary
(intra-dialect, model-routed) legs, non-streaming and streaming, plus
tool calls and credential pass-through, mirroring the prior script's
scenario coverage.

All four mock backends (mock-openai, mock-openai-secondary,
mock-anthropic, mock-anthropic-secondary) gained SSE streaming support --
previously they only ever returned static non-streaming JSON regardless
of the request's `stream` field -- since the SDKs' streaming iterators
need a real event stream to consume through the proxy/translation layer.

The verify compose service switches from node:20-alpine (curl) to
python:3.13-alpine (openai + anthropic pip packages), since the checks no
longer need Node or curl at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
The prior commit added content/stop_reason/stop_sequence to
LlmAnthropicEventMapper's encoded message_start but left usage omitted
when input tokens aren't known yet (still true for an OpenAI-sourced
stream, which never discloses them before its terminal chunk). That
omission is a different problem from the deferred-disclosure one
LlmUsageFlushEx's own comment addresses: anthropic-sdk-python's streaming
accumulator initializes its per-message snapshot from message_start and
then patches usage.output_tokens on it in place when message_delta
arrives -- with no usage object to patch, that crashed with 'NoneType'
object has no attribute 'output_tokens' (confirmed streaming an
openai-to-anthropic cross-dialect response through the real SDK via
examples/llm.proxy).

message_start now always carries a usage object; input_tokens is a
provisional 0 for a source dialect that hasn't disclosed it yet,
corrected by nothing further since Anthropic's own message_delta usage
never carries input_tokens either -- only output_tokens, which
encodeFinish already corrects from whatever the source discloses.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
Replace the curl commands with the openai/anthropic Python SDKs for all
four demo scenarios, folding the separate SDK section into Try it so
there's one walkthrough instead of two overlapping ones.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
jfallows pushed a commit that referenced this pull request Sep 23, 2026
…izfewx

Resolves overlapping changes with PR #2598 (llm(proxy) routing):

- Mock server route paths: kept this branch's fix (register on the
  dialect's real canonical path /v1/messages or /v1/chat/completions)
  layered on top of PR #2598's added SSE streaming support for each mock,
  since both changes touch the same route-registration line but are
  otherwise independent.
- README.md: took PR #2598's version entirely (SDK-based walkthrough
  rewrite); this branch never touched the file.
- LlmAnthropicEventMapper.java/Test.java and the openai.to.anthropic.streaming
  k3po fixture: kept deleted. These predate this branch's LlmFlushEx ->
  LlmDataEx wire-model rewrite and JsonPipeline consolidation, and the
  fixture still asserts the retired zilla:flush/matchFlushEx() shape.
  Ported the real bug PR #2598 fixed in the now-deleted mapper --
  Anthropic's message_start event needs content/stop_reason/stop_sequence
  and a structural usage object, or anthropic-sdk-python's streaming
  accumulator crashes -- into this branch's replacement,
  LlmAnthropicEncodeSink.messageStartSteps(), and updated this branch's
  own equivalent fixture (application/anthropic.streaming.transformed)
  to match. Full binding-llm.spec + binding-llm verify suite green after
  the port.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AKZMN6ckSpx1QpHvaE4DwM
Reconciles binding-llm's independent evolution on develop (an earlier
snapshot merged via a different route: LlmContentDecoderSpi, dialect
detection/resolver, event mapper, etc., without kind:proxy) with this
branch's later state (adds kind:proxy routing for #2502, plus two real
message_start encoding bug fixes found via SDK testing). Every conflicted
file was a clean superset in this branch's favor; resolved by keeping this
branch's content throughout.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
…ct resolution

The kind:proxy enum entry and its closing bracket got glued onto one line
during the develop merge's conflict resolution, so the checked-in fixture no
longer matched zilla inspect schema's actual output byte-for-byte. Verified
against a real docker compose run of the example.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01C5EdovtZRhxkqY5DF6MsoG
@jfallows
jfallows merged commit 94a6598 into develop Sep 30, 2026
45 checks passed
jfallows pushed a commit that referenced this pull request Sep 30, 2026
develop's #2598 (now merged) landed an independently-evolved, older
implementation of the same llm binding feature under incubator/binding-llm,
incubator/binding-llm.conf, and incubator/binding-llm.spec — built on the
pre-JsonPipeline LlmAnthropicEventMapper/LlmOpenaiEventMapper architecture,
with its own synthetic test dialects and k3po fixture naming. It shares only
a distant common ancestor with this branch's JsonPipeline-based rewrite
(LlmAnthropicEncodeSink/LlmOpenaiEncodeSink/LlmCanonicalEncodeSink etc.).

Per explicit direction, this branch's implementation wins wholesale: every
develop-only file under the three llm modules is removed, and every add/add
conflict resolves to this branch's version. examples/llm.proxy's four mock
server files merge develop's new SSE-streaming helpers with this branch's
route-path fix; its README takes develop's SDK-based rewrite since this
branch never touched it. The one non-llm conflict, the engine config JSON
schema fixture, keeps this branch's llm(client) options.server description.

incubator/NOTICE regenerated via notice:generate to pick up two Aklivity
Community License entries (catalog-inline.conf, model-json.conf) that
landed via develop in the interim.

Full binding-llm module suite verified green post-merge: 136 unit tests,
57 integration tests (LlmClientIT, LlmServerIT, LlmProxyIT), checkstyle,
license check, and coverage all passing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AKZMN6ckSpx1QpHvaE4DwM
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

llm proxy: route on LlmBeginEx.model, extracted at llm server via bounded JsonPath (design)

2 participants