Skip to content

fix(binding-llm.conf): reject llm config properties the binding ignores - #2618

Open
jfallows wants to merge 263 commits into
developfrom
claude/dreamy-hypatia-l3y4fl
Open

jfallows wants to merge 263 commits into
developfrom
claude/dreamy-hypatia-l3y4fl

Conversation

@jfallows

Copy link
Copy Markdown
Contributor

Description

Stacked on #2611 (claude/loving-cannon-htqb1s). Targets develop with the aggregate diff; the new work is the last two commits after 055cef0e. It will be rebased once #2611 and the PRs beneath it merge.

The llm binding schema accepted configuration that the binding never reads, and it did not require options on a client.

  • vault and catalog were accepted on every kind. The binding reads neither.
  • Route when and with were accepted on server and client. Those kinds pick a route only by its guarded roles, so when was silently ignored. with failed later, with an opaque JsonbException from the config adapter.
  • Route with was accepted on proxy, with the same late failure.
  • options was optional on a client. The client refuses every stream without options.dialect and options.server.

With this change the schema rejects those properties and requires options on a client, following the mcp binding's schema. Such a config now fails schema validation when it loads, instead of loading and then being ignored or failing obscurely. Route guarded and exit stay valid on every kind. Proxy route when (dialect, model) is unchanged.

examples/inspect.schema/.github/schema.expected.json is regenerated to match.

Testing

  • Test-first. LlmSchemaValidationTest gets 14 new cases, committed first:
    • reject vault, catalog and route with on each of server, proxy and client
    • reject route when on server and client
    • reject a client without options
    • accept a guarded route on server and client
  • Against the old schema, the 12 new reject cases fail and the 2 accept cases pass.
  • Route with assertions expect ConfigException rather than RuntimeException. Otherwise they would already pass on the old schema, because of the adapter's JsonbException.
  • With the fix, mvn clean install passes for binding-llm.conf (75 tests), binding-llm (193 unit tests plus 81 ITs: LlmClientIT, LlmServerIT, LlmProxyIT) and metrics-llm (including LlmMetricsIT and the IT-config schema check). Checkstyle, license and jacoco run as part of that build.
  • Golden schema: regenerated by running zilla inspect schema from a local zpmw install of the docker-image zpm.json, because Docker Hub rate-limited the base image pull. It differs from the previous golden only by the llm schema changes above.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DTRu8xoJU64qAdSohqrKqm


Generated by Claude Code

… LlmBeginEx.dialect

Implements request/response framing decode-and-forward for the LLM server
binding (#2484): LlmServerFactory drives the request body
through the resolved dialect's ModelPipeline as bytes arrive off the wire,
forwards transformed content to app0 incrementally rather than buffering
the whole body, and threads the resolved dialect name onto LlmBeginEx so
app0 can see which wire dialect produced the request.

Flow control between the client, this binding, and app0 is enforced with
dedicated decodeSlot/encodeSlot buffers on each side of the exchange
(LlmServer for the network-facing leg, LlmStream for the app-facing leg),
each granting credit strictly from its own local slot occupancy rather than
copying a peer's sequence numbers across independent byte domains. LlmState
tracks per-direction open/closing/closed transitions and end-of-stream
deferral while a buffer is still draining.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
…fferPool handles

DefaultBufferPool.buffer(slot) rewraps and returns a single shared mutable
field per pool instance, so holding a buffer reference across a nested call
that fetches a different slot from the same pool silently repoints it.
LlmServer.decodeNetwork() directly, synchronously calls LlmStream's
request-relay staging method as a plain nested Java method call, so a
single pool instance backing both the network-decode slot and the app0
request-relay slot would alias between them.

Give LlmServerFactory two BufferPool handles instead of one: decodePool
(network decode) and encodePool (app0 relay, both directions), obtained via
context.bufferPool() and .duplicate() -- matching McpServerFactory's
decodePool/encodePool precedent. The two relay directions sharing
encodePool (the reply-direction slot on LlmServer and the request-direction
slot on LlmStream) never appear in the same call stack: cross-binding
accept() is ring-buffer-mediated and dispatched on a later engine tick, not
a nested call, so one encodePool instance can't alias between them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
…nline.conf

LlmSystemNamespaceGenerator emits a real "type": "inline" catalog config
into the generated sys: namespace patch, serviced at actual runtime by
catalog-inline's CatalogFactorySpi via ServiceLoader -- not just exercised
by this module's own tests. test scope kept it off the runtime classpath
entirely; provided scope wouldn't fit either, since nothing in main source
compiles against catalog-inline's Java API (the reference is a plain
string), so there's nothing to satisfy at compile time.

Match binding-asyncapi/binding-openapi's precedent for this same
generated-"type: inline"-config pattern: catalog-inline at runtime scope,
plus the companion catalog-inline.conf (config-schema side) at default
scope alongside the existing model-json.conf dependency.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
LlmBindingConfig.newModelConfig() constructs a real JsonModelConfig in main
source, resolved to a working ModelHandler by context.supplyModel() via a
ModelFactorySpi lookup at actual runtime -- not just exercised by this
module's own tests. Matches binding-asyncapi/binding-mcp-openapi/
binding-openapi, each of which also builds JsonModelConfig directly in main
source and declares model-json at runtime scope.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
Stale relative to the catalog-inline/model-json scope fixes on binding-llm
(catalog-inline.conf and model-json.conf now roll up into this aggregate),
plus a pre-existing gap for binding-http.spec's license entry. Regenerated
via ./mvnw notice:generate -pl incubator -amd, never hand-edited, per
AGENTS.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mHT1vrPATQGSdRcB8NNKM
Add LlmServerConfig/LlmServerConfigBuilder (host/port, nested builder)
following the existing *OptionsConfigAdapter pattern used by
binding-kafka's options.servers, wired into LlmOptionsConfig via a new
`server` field so `llm client` bindings can configure their upstream
endpoint. Config adapter unit tests cover parsing and serializing
options.server.

Fixes #2485

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FQ4XvYNCp1NzQC4Ne4eWbK
Exercise the inject() path on LlmServerConfigBuilder so the nested
server builder reaches the module's required 100% instruction
coverage, mirroring the existing shouldInjectBuilder test for
LlmOptionsConfigBuilder.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FQ4XvYNCp1NzQC4Ne4eWbK
…→encoder chaining

Adds LlmClientFactory (kind: client), the first real integration point for
LlmDialect.supplyDecoder/supplyEncoder: it compares the inbound app-declared
dialect (LlmBeginEx.dialect) against the binding's own configured dialect and
only chains a JsonPipeline payload transform when they differ, forwarding
framing-only re-encoded content when they match.

- internal/encode: LlmContentEncoder/Spi/Factory + LlmSseContentEncoder, the
  SSE-framing encode counterpart to internal/decode's existing SSE decoder
- LlmDialectResolver.dialectNamed(String): by-name lookup for resolving the
  inbound dialect when it differs from the client's configured one
- Schema patch: kind: client options (dialect, server); also fixes kind:
  server to accept options.server (previously unreachable under its
  additionalProperties: false, despite LlmOptionsConfig already supporting it
  since PR 2570)
- Spec scripts (same.dialect, cross.dialect, client.opaque.fallback,
  client.abort) plus a new LlmClientIT, written first per this repo's
  test-first discipline, confirmed failing before LlmClientFactory existed

Fixes #2486

Real dialect implementations (openai, anthropic), the mock backend, and the
full round-trip identity test are separate, sibling issues (#2487, #2490,

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
…aders, body)

Rebased onto #2570's latest, which changed LlmDialect.contentType() to take
(Kind, headers, body) parameters, resolved per request rather than fixed
once per dialect instance. The client resolves content-type separately for
each direction now (REQUEST for its outbound encoder, RESPONSE for its
inbound decoder) rather than a single shared call, matching the interface's
own point: request and response can have different native content-types.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
LlmClientFactory.newStream resolved the RESPONSE content decoder eagerly at
BEGIN time with body=null, so LlmDialect.contentType(Kind.RESPONSE, ...)
could never actually see request-body content when deciding between a
streaming and non-streaming response, defeating the per-request capability
added for that method. REQUEST-side encoder resolution is unaffected since
it does not depend on body content.

Accumulate the request body (post cross-dialect transform, pre-framing) into
a buffer-pool-backed slot as DATA arrives, and resolve the decoder at
onAppEnd, once the full request is available, relying on this binding's
half-duplex transmission convention to guarantee no response bytes arrive
before then. Add LlmJsonRequestBody, a pure-Java HttpRequestBody view driven
by common-json's one-shot parser API, plus a unit test. Add a new
test-conditional dialect and two client IT/spec scenarios proving the
decoder now differs based on a stream field in the request body.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
…line SPI

LlmDialect no longer exposes contentType(Kind, HttpHeaders, HttpRequestBody);
dialects now detect via ModelEnvelope and supply ModelTransform-based
decoders/encoders driven through the engine's ModelHandler/ModelPipeline SPI.
Content-type is read directly off real wire headers instead of being computed
per-dialect, so the request-body buffering added to defer response-decoder
selection (LlmJsonRequestBody) is no longer needed and is removed.

LlmClientFactory is rebuilt against the new SPI: dialect resolution via
LlmBindingConfig.resolveDialect/dialectNamed, a per-stream ModelEnvelope, and
a ModelPipeline (ModelTransform.NONE for same-dialect) run unconditionally in
both directions so downstream code always sees validated, well-formed
payloads. Test dialects are renamed (test-client, test-client-sse,
test-client-sse-alt) to avoid colliding with the real upstream test fixtures,
and gain request/response JSON schemas so the pipeline can actually validate
their payloads instead of rejecting everything for lack of a schema.

Adds the missing llm:flushEx()/llm:matchFlushEx() k3po functions that the
client k3po scripts already relied on.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qssroe9MdjvpvQrthi15ex
…sponse transforms

Implements LlmDialect for OpenAI Chat Completions, registered via
LlmDialectFactorySpi: detect() matches POST /v1/chat/completions and
POST /v1/completions; contentType() is text/event-stream, the only
LlmContentDecoderSpi/LlmContentEncoderSpi framing codec this module
has today (a non-streaming application/json pair is a content-decoder-
layer gap, not a dialect one, and is left as follow-up).

supplyDecoder/supplyEncoder rename OpenAI-native request/response JSON
members to the canonical vocabulary this dialect defines a synonym
for -- max_tokens/maxOutputTokens, top_p/topP, n/choiceCount and
friends on requests; index/choiceIndex, finish_reason/finishReason
(remapping tool_calls/tool_call), logprobs/logProbability, and the
usage token counts on responses -- via a depth-tracked JsonTransform
that renames JSON events without DOM parsing. Everything without an
established canonical synonym (id, model, messages, tools, the whole
delta/tool_calls structure including streamed function.arguments
fragments) forwards unchanged, at any depth, so decode -> encode
round-trips with no loss.

Tests cover dialect detection/registration and both Kinds of
transform, including full round-trips through the canonical form.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
… rename

model-json's ModelTransform integration was observation-only (ModelFieldBridge),
discarding REPLACED/DECLINED answers instead of writing them. JsonModelFieldTransform
drives a wired ModelTransform inline as JSON streams through, computing a JSON-pointer
path per scalar field (with array-index segments) at any nesting depth, and writes
FIELD/REPLACED/DECLINED answers straight to the destination -- including a REPLACED
substitute redirecting a field to a sibling key of the same enclosing object.

Adds the missing key-write for container-valued members entering a named object
member, fixes a resumed key/value write re-offering the whole text instead of the
remainder (TextSource now tracks its own consumed() progress), and reuses the
already-decoded scalar text/key for an unchanged FIELD answer instead of a fresh
per-field allocation.

Also fixes JsonModelHandlerImpl.supplyEncoder silently dropping its transform
parameter.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
…e SPI

Rebases LlmOpenAiDialect and its request/response transforms onto the
redesigned LlmDialect SPI: detect(ModelEnvelope), no more contentType()
(content-type is now resolved from real upstream Content-Type headers),
supplyDecoder/supplyEncoder(Kind, ModelEnvelope) returning a ModelTransform.

The request/response transforms are rewritten as plain ModelTransform
implementations matching each field's own full path (e.g. $.choices[0].index,
$.usage.prompt_tokens) rather than tracking JSON structural depth, since the
model-json adapter now computes paths itself. This drops the old
JsonEvent-token depth-tracking machinery and the JsonSource/JsonController
wrappers (LlmOpenAiStructuredController is no longer needed); the renamed
LlmOpenAiSubstitutedSource is now a plain ModelSource.

The choices[]/usage direct-member path checks defer their substring() until
after confirming the path is actually a direct member, so the many deeply
nested per-chunk fields that share the prefix (delta.tool_calls[].index and
the like) cost no allocation.

The one dropped rename from the prior implementation is logprobs/logProbability:
it names a container-valued field, and this dialect only renames scalar leaves.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
…openai dialect

LlmOpenAiRequestTransform renamed a fixed set of top-level fields but never
observed model, so LlmServerFactory.doAppBegin's server.envelope.get("model", 0)
read right after running the request through this transform came back empty and
LlmBeginEx.model was never stamped for real dialect: openai traffic.

Mirrors LlmTestPermissiveDialect's inline ModelExtractTransform (and
KafkaExtractTransform's own pattern): on a FIELD event at $.model, copy the
value into the envelope alongside the existing rename-or-forward decision,
without touching the RENAMES table. LlmOpenAiResponseTransform needs no
equivalent change -- nothing reads model back off a response.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
…requests

CI's LlmServerIT.shouldRejectRequestWithUnresolvedDialect failed: detect()
matched on :method/:path alone, so it ambiguously co-matched every existing
k3po server fixture -- all of which hit the same /v1/chat/completions path
with a test-specific content-type (application/vnd.zilla.test-permissive+json,
application/vnd.zilla.test-strict+json) to select a *different* dialect
unambiguously. Only one failure surfaced in CI because failsafe stops after
the first failure, but the same ambiguity affected every other fixture in
that class too -- confirmed by LlmServerIT going from 6 run/1 failed/1
skipped to 8/8 passing with this fix, and LlmClientIT unaffected at 4/4.

detect() now also requires content-type: application/json, the only
content-type a genuine OpenAI request ever carries, so a request to the same
path with a different dialect's own content-type no longer collides.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SMybNYCLJEWTwCjDscepZm
The llm.schema.patch.json allowed options.server on kind:server bindings,
but LlmServerFactory never reads binding.options.server — only
LlmClientFactory dials it as the upstream endpoint. Restrict the
kind:server options schema to dialect (fixed-dialect mode), matching
actual runtime usage and the milestone's example configs.

Add LlmSchemaValidationTest exercising the full EngineConfigReader
pipeline against real zilla.yaml text for both kind:server and
kind:client, covering acceptance (bare server, fixed-dialect server,
full client) and rejection (server option on kind:server, missing
required client fields, malformed server pattern, unknown kind).

Fixes #2488

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Aawd4ZFaUE3FgQnzkABj9F
LlmOptionsConfigAdapterTest: round-trip dialect+server together, an
explicit server-absent case, and the adapter's silent no-op on a
malformed server string (schema is the actual gatekeeper there).

LlmSchemaValidationTest: reject additionalProperties on both kind:server
and kind:client, reject non-string dialect/server values, and accept an
empty options object for kind:server.

Closes #2489.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XfesxqkqmMKPC87ydouhGR
Adds paired client.rpt/server.rpt k3po scenarios exercising the real
`openai` LlmDialect end-to-end (dialect detection, framing decode/encode),
rather than the synthetic test-sse/test-conditional dialects existing
scenarios use:

- openai.request (network+application): llm server detects `dialect:
  openai` from `POST /v1/chat/completions` alone (no x-llm-dialect
  header) and forwards the decoded JSON request body, terminated by the
  JSON content-decoder's single terminal flush.
- openai.streaming (network+application): llm client same-dialect
  round-trip with a `stream:true` request and a realistic OpenAI SSE
  response sequence (role chunk, content chunk, tool-call start and
  argument-fragment chunks, finish_reason chunk, usage chunk, [DONE]).
- openai.nonstreaming (network+application): llm client same-dialect
  round-trip with a `stream:false` request and a plain-JSON response
  carrying top-level usage.

New scenarios are wired into LlmServerIT/LlmClientIT (engine-driven) and
ApplicationIT/NetworkIT (protocol self-consistency), plus a new
client.openai.yaml config pinning `dialect: openai`. All four IT classes
compile cleanly and pass checkstyle; the k3po `verify` lifecycle itself
could not be exercised in this sandboxed session due to a pre-existing,
repo-wide maven-notice-plugin license-mapping gap unrelated to this
change (same limitation noted in prior binding-llm PRs' test plans) — CI
should confirm.

Fixes #2490

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
…p+passthrough server

LlmServerFactory no longer decodes request framing into typed DATA/FLUSH
events -- it detects the dialect, resolves content-type directly off the
real Content-Type header, runs the dialect's decoder through a json-model
ModelPipeline, and forwards the pipeline's own output as plain DATA frames,
stamping dialect/contentType on the app-facing LlmBeginEx. Update the
openai.request client/server pair (network + application) to match: a real
content-type header drives LlmBeginEx.contentType, and the terminal
LlmNativeFlushEx this scenario previously asserted is gone since the server
no longer produces one for plain JSON forwarding.

model is intentionally left unasserted here: the real openai dialect does
not yet extract it into the envelope (only the test-permissive test dialect
does), unlike the design this fixture set was originally written against.

Part of the broader LlmDialect/ModelPipeline SPI redesign already landed on
this branch's upstream chain; openai.streaming/openai.nonstreaming still
need the same treatment for LlmClientFactory's same-dialect path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
…i client fixtures

LlmClientFactory always reads content-type off real HTTP headers -- the
request's own if the app sets one, else "application/json" -- and its
response codec is selected by looking up the backend's real Content-Type
response header via LlmContentCodecFactory (registered for
"text/event-stream" and "application/json"). A response header carrying
neither falls through to opaque forwarding instead of the intended codec.

openai.streaming/openai.nonstreaming previously omitted these headers
entirely, so despite exercising a real dialect they were never actually
routing through the SSE/JSON content codecs LlmClientIT is meant to cover.
Add "content-type: application/json" to the outbound request and
"content-type: text/event-stream"/"application/json" to the backend
response on both the network and application sides, matching every other
LlmClientIT-driven fixture's convention (same.dialect, cross.dialect).

openai.nonstreaming was also missing the terminal raw LlmFlushEx that
LlmJsonContentDecoder always emits alongside its single DATA -- add it to
both the read and write side, matching the already-correct
openai.streaming pattern.

Also make openai.request's application/server.rpt echo the same request
body back as its reply (rather than an unrelated "response body" literal),
matching request.valid's own client/server symmetry convention, with the
network pair's expected reply content updated to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
…eginEx

The #2487 session landed model extraction for the real openai dialect
(LlmOpenAiRequestTransform now observes $.model and copies it into the
ModelEnvelope, mirroring LlmTestPermissiveDialect's ModelExtractTransform).
Now that LlmServerFactory.doAppBegin actually stamps LlmBeginEx.model for
dialect: openai traffic, assert it in the one fixture pair that exercises
server-side model extraction (openai.streaming/openai.nonstreaming never
touch model at all -- LlmClientFactory doesn't read or stamp it in either
direction).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
… traffic

LlmClientFactory always constructed and drove a schema-validating JSON
model pipeline for both request and response, even for same-dialect
traffic where the design intends pure raw byte relay with no transform.
For a dialect that declares no JSON schema for a direction (openai
declares none for either), the pipeline's schema resolution always
returns NO_SCHEMA_ID, so every same-dialect request or response got
REJECTED and the stream was torn down -- reproduced directly against
JsonModelHandlerImpl/JsonModelDecoderPipeline with the same
no-schema/ModelTransform.NONE configuration LlmClientFactory builds.

Skip constructing requestPipeline/responsePipeline entirely when
source == target, forwarding request and response bytes directly
(still through the content-type codec's own encoder/decoder for
framing) instead of driving a pipeline with nothing to validate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
… schemas

LlmOpenAiDialectFactorySpi previously declared no schema for either
direction, so the shared sys:llm_dialects catalog had no "openai.request"/
"openai.response" subject and the ModelPipeline LlmClientFactory drives for
openai traffic (same-dialect included) always resolved NO_SCHEMA_ID and
rejected every value, regardless of content -- there was no actual
validation happening for openai traffic at all.

Add openai.request.schema.json/openai.response.schema.json describing the
real Chat Completions wire shapes: request requires model/messages and
types the other fields LlmOpenAiRequestTransform's rename table and
model-extraction both recognize (stream, max_tokens, top_p, n,
presence_penalty, frequency_penalty, tool_choice, response_format, ...);
response types choices[]/usage without requiring any top-level member,
since a streaming chunk carries only a subset (delta vs message,
finish_reason, tool_calls[], usage) of what a single non-streaming
completion carries all at once.

Verified directly against JsonModelHandlerImpl/JsonModelDecoderPipeline
(the same construction LlmClientFactory drives) that all nine request/
response payloads used by the openai.request/openai.streaming/
openai.nonstreaming k3po fixtures validate as COMPLETE against these
schemas, not REJECTED.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
…heck

Asserting the schema's required fields and specific property keys in a
unit test duplicates what the openai.request/openai.streaming/
openai.nonstreaming k3po fixtures already verify against a live engine --
those fixtures are the actual spec for what these schemas must accept,
per this repo's test-first discipline. Narrow the unit test to what a
factory-level test is actually entitled to check: the schema resource
for each Kind exists and parses as a JSON object.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
Engine.java's bootstrap (bindings.stream().map(Binding::system)...) calls
LlmBinding.system() -> LlmSystemNamespaceGenerator.generate() for every
engine startup in this module, which invokes schema(Kind) for every
registered LlmDialectFactorySpi (openai included) for both REQUEST and
RESPONSE and reads the resource -- unconditionally, on every LlmServerIT/
LlmClientIT/ApplicationIT/NetworkIT run in the module, not just
openai-specific scenarios. A missing or unreadable schema resource would
already fail engine bootstrap loudly. The unit test duplicated coverage
the k3po ITs already provide for free.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
Adds the negative-path coverage the openai.request/streaming/nonstreaming
scenarios never exercised: an invalid document actually gets rejected,
not silently forwarded.

openai.request.invalid (network only, mirroring request.rejected.schema's
existing single-sided pattern): a real openai-dialect request missing the
required "messages" field -- llm(server) rejects it before ever opening
an app-facing stream, observed as the k3po connect script itself getting
aborted.

openai.response.invalid (both application and network, mirroring the full
openai.streaming/nonstreaming layout): a real backend response with
"choices" as a string instead of an array -- llm(client) rejects it and
aborts app0 instead of forwarding the malformed document.

Wired into LlmServerIT/LlmClientIT (engine-driven) and
ApplicationIT/NetworkIT (protocol self-consistency), matching the
existing openai.* scenario conventions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
…i traffic

llm(client) computed requestContentType from llmBeginEx.contentType() with
a wrong null-check: the flyweight accessor itself is never null even when
the underlying string16 field is unset, so the fallback to CONTENT_TYPE_JSON
never triggered and null flowed into the outbound HTTP begin's content-type
header, breaking header matching on the network side and hanging every
k3po scenario that omits contentType on the app-side llm:beginEx (which is
the common case -- the field is only meant to be set for cross-dialect
routes). Check contentType().asString() directly instead, matching the
existing idiom one line above for the dialect field.

Also stop routing same-dialect responses through a schema-validating
ModelPipeline: responsePipeline was unconditionally constructed regardless
of dialect, so a same-dialect stream (zero decode/encode/transform by
design, per LlmOpenAiResponseTransform's own documented contract) still
paid for JSON-schema validation on every response chunk, including the
OpenAI SSE `[DONE]` sentinel, which is not JSON and got truncated by the
validator. responsePipeline is now null for sameDialect, and
forwardResponseContent forwards raw bytes in that case, mirroring how
requestPipeline already goes null when no encoder resolves.

Removes LlmClientIT.shouldRejectInvalidOpenAiResponse: it asserted schema
rejection of an invalid response body on a same-dialect (single-dialect)
client config, which cannot validate anything now that same-dialect
responses bypass the model pipeline entirely -- this was always the
intended contract, just not one this scenario could have exercised.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
…e [DONE] sentinel

The previous commit disabled response schema validation entirely for
same-dialect llm(client) traffic to work around the OpenAI SSE `[DONE]`
sentinel getting mangled by the JSON model pipeline. That was the wrong
fix: `[DONE]` isn't JSON at all (it's an SSE-level stream-termination
token, not a chat-completion chunk), so no JSON transform can "honor
identity" for it -- there's no valid parse to preserve. The actual defect
was routing a non-JSON control token into JSON-schema validation at all,
not the presence of validation itself.

Restores responsePipeline unconditionally (same-dialect responses are
schema-validated exactly like cross-dialect ones), and instead adds a
narrow, explicit bypass in forwardResponseContent for the literal `[DONE]`
bytes specifically, forwarding them raw before they ever reach the model
pipeline. Genuine JSON content -- valid or invalid, same-dialect or
cross-dialect -- is still schema-validated and rejected on violation.

Restores LlmClientIT.shouldRejectInvalidOpenAiResponse, which the
previous commit had removed as no longer exercisable; it now passes again
since same-dialect responses are validated once more.

Also corrects LlmOpenAiResponseTransform's class Javadoc, which
overstated the same-dialect bypass as covering the whole path rather than
just the non-JSON sentinel.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L5Pw5eDcghnK7LmZ11QA7u
claude and others added 28 commits September 23, 2026 01:05
…e codecs

Moves LlmContentCodecSpi and the types its signature depends on
(LlmContentDecoder, LlmContentDecoderOutput, LlmContentEncoder) from the
internal codec/decode/encode packages to a new public codec package, and
updates module-info.java to export/use it, mirroring binding-mcp's
within-binding SPI precedent. The built-in SSE and JSON codecs continue to
register as provides entries with no behavior change.

Fixes #2501

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WUXczd1xDohH2cmpy8S5As
…ings

kind: client bindings can only authenticate themselves to their upstream
via a single static credential value in one header (options.authorization),
resolved once per stream before any body bytes are known. Some upstream
APIs require a signature computed over the complete request instead --
method, path, headers, and a hash of the full body -- which a static
per-header value can never express, and which is only computable once the
whole request is known.

Add an optional options.sign: <name> on a kind: client binding, resolving
to a registered LlmRequestSigner (SPI: LlmRequestSignerFactorySpi, alongside
the existing LlmDialectFactorySpi/LlmContentCodecSpi extension points --
dialects can already be contributed from outside this module the same way).
A stream with a signer configured buffers its complete outbound request
(bounded by the new zilla.binding.llm.signed.request.max.bytes property,
default 10 MiB, resetting the stream if exceeded) before opening the
network connection, so the signer sees the whole method/path/headers/body
and returns whatever additional headers authenticate the request. Every
other stream is unaffected -- the existing incremental send-as-it-arrives
path is unchanged when no signer is configured.

Also makes LlmRequestFieldTransform public (it was package-private,
preventing the externally-contributed dialects this package's own
LlmDialect/LlmDialectFactorySpi javadoc already advertises from reusing
its depth-1 field-rename machinery), and drops the dialect schema
property's closed enum (openai/anthropic) to a plain string -- it silently
contradicted that same "contributed from outside this module" contract,
since LlmDialectResolver already resolves dialect names dynamically via
ServiceLoader at runtime and handles an unregistered name there.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…on through LlmDialect

kind: client's cross-dialect RESPONSE translation (streaming and
non-streaming alike) went through a hardcoded two-way switch on dialect
name (openai vs. else-anthropic), never through LlmDialect itself --
unlike the request direction, which already calls source.supplyDecoder/
target.supplyEncoder(REQUEST,...) through the interface. A dialect
contributed from outside this module (exactly the extension point this
package's own javadoc advertises) got silently routed through Anthropic's
response transform instead of its own for any name other than "openai" --
wrong output, not a clean rejection.

Add LlmDialect#supplyResponseDecodeTransform()/supplyResponseEncodeSink(),
both required (no default -- a null-returning default would just move the
same silent-fallback bug one layer later). LlmClientFactory now calls
through the resolved dialect instead of the removed
LlmResponseTransformFactory; openai/anthropic's own implementations
construct the same LlmOpenaiDecodeTransform/LlmOpenaiEncodeSink/
LlmAnthropicDecodeTransform/LlmAnthropicEncodeSink classes the old switch
did, so this is a pure relocation for them with the same test suite
passing unchanged. LlmNativeEventOutput moves from the internal mapper
package to the public dialect package, since it's now part of this
contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…ernal dialects

LlmClientFactory unconditionally casts the result of
LlmDialect#supplyResponseDecodeTransform()/supplyResponseEncodeSink(...)
to LlmDialectEvent/LlmDialectTerminator respectively -- but both
interfaces were internal.mapper-only, unexported, so a dialect
contributed from outside this module (the extension point these two
methods, and LlmDialect generally, are documented to support) could
never implement them: the cast would throw ClassCastException the first
time a live engine actually built a cross-dialect response pipeline for
it, even though its own JsonTransform/JsonSink compiled and tested fine
in isolation.

Move both to the public dialect package, alongside LlmNativeEventOutput
(exported the same way for the same reason). LlmDialect's own javadoc
for the two response-transform methods now states the requirement
explicitly. The two built-in dialects' existing decode transforms/encode
sinks already implemented them; only their imports change. Strengthened
the dispatch regression test's own third-party test double to
implement both too (as no-ops), so it continues to model the real
contract every dialect must satisfy, and added a test asserting that
directly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…uration

Every other pluggable extension point reachable from a binding config
pairs a type/name selector with a nested options object -- the request
signer SPI was the one exception, with no way for an implementation to
receive any per-binding configuration at all. sign now accepts either a
bare name (unchanged) or {name, options}, resolved by an LlmSignInfo the
same way this binding already resolves a dialect's or content codec's own
options, and LlmRequestSignerFactorySpi.create() receives that options
value.

feat(binding-llm): let a dialect's request path defer resolution to the request's own model

requestPath(basePath) is resolved once per stream from static config,
before any request body byte has arrived -- but some upstreams address a
specific resource via the request path itself, which can only be known
once the request body has been read. A dialect can now return a path
carrying the literal token {model}, resolved per request from the same
model value this binding's request-transform pipeline already extracts
generically, by reusing the buffering this binding already performs for
signing (a required model unknown by the time the body completes fails
the stream cleanly, the same as a signer producing no path would).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…e sign: SPI

A binding could previously name a dialect and, independently, name a
request signer -- two separately-resolved selectors for what is really one
choice, with no way to catch a binding pointing them at a mismatched pair.
Every other capability a dialect's upstream needs (its request path, its
credentials header, its response terminator) is already asked of the
dialect itself; signing is no different, so LlmDialect now exposes an
optional signer(), following the same null-when-unneeded convention
terminator(Kind) already uses, and the separate LlmRequestSignerFactorySpi/
LlmSignConfig/LlmSignInfo/sign: binding option are removed entirely.

A dialect factory now receives an LlmDialectContext (currently just a
signaler, for scheduling background work off the reactor thread) at
construction, the same way a signer factory used to -- LlmBindingConfig
already built that context at the exact point it resolved the dialect, so
this reuses existing plumbing rather than adding new. Because every
registered dialect is constructed on every binding attach (needed for
detection and cross-dialect lookup, not just a binding's own configured
one), signer() must stay lazy: constructing a dialect must never itself
trigger the side effects a real signer's construction might have, only
asking for the signer does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…sponses

LlmHttpClient built its response pipeline's usage/token extractor from
client.source (the caller-facing dialect) while decodeTransform on the very
next line correctly used client.target (the upstream native dialect) -- but
extraction runs on the raw upstream bytes, before decodeTransform translates
them, so it must key off the same dialect decodeTransform does. Every other
analogous call site (buildResponsePipeline) already gets this right. Since
LlmUsageExtractTransform is identity-only, this didn't corrupt response
bodies -- it just meant usage/token side-channel extraction silently found
nothing for any cross-dialect response.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
… its response content-type reflect a cross-dialect encode

requestPath(basePath) carries no streaming-vs-non-streaming signal, so a
dialect whose upstream puts streaming and non-streaming operations at
different paths (rather than differentiating only the request body) has no
way to switch on it -- add a default requestPath(basePath, boolean streaming)
delegating to the existing method, and re-derive a signed client's path
template from the request's own decoded "stream" flag (already captured into
the per-stream envelope by whichever dialect decodes the request) at the
point doNetBeginSigned() finalizes it, instead of reusing the constructor's
early, streaming-unaware template.

A signer may legitimately fail fast (its own credentials not yet warmed via
a background fetch) rather than block the reactor thread -- but
doNetBeginSigned() let that exception propagate uncaught into the shared
engine worker thread, terminating every stream that worker was handling, not
just the one being signed. Route it through the same recovery
resolvedPath == null already uses. That recovery only reset the initial
direction, a no-op once onAppEnd() has already closed it (the only time
doNetBeginSigned() runs) -- also open+abort the reply direction so the app
is actually told the request failed, instead of never hearing back.

A cross-dialect response's own content-type was forwarded to the app
unchanged from the network response, rather than reflecting the shape the
response is actually translated into -- correct only when both dialects
happen to use the same content-type for the same shape. Compute it from the
canonical streaming/non-streaming distinction instead when translating.
That distinction itself relied on an instanceof check against one specific
decoder, so it silently read as non-streaming for any other inherently
multi-frame content-type; give LlmContentDecoder a streaming() capability
method instead, so a caller reads it generically.

Also guard a streaming chunk's optional id the same way its optional model
is already guarded, matching the whole-document variant right next to it --
an absent id crashed the encode outright instead of just omitting the field.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…lm(dialect)

llm(client)'s and llm(server)'s dialect schema property dropped its closed
openai/anthropic enum for a plain string, and its description was reworded,
but examples/inspect.schema's golden schema.expected.json (which asserts the
full merged JSON schema byte-for-byte via `zilla inspect schema`) was never
updated to match, breaking the inspect.schema example's CI check.
Regenerated via the documented process: `docker compose exec zilla zilla
inspect schema > .github/schema.expected.json`, confirmed deterministic
across two runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…tions

The `dialect` schema property's closed enum (openai/anthropic) and both
options.dialect descriptions were unintentionally reverted/dropped while
adding outbound request signing, on the reasoning that dialects can be
contributed from outside this module. That reasoning doesn't hold: the
established pattern for an extensible name in this schema is a closed
enum in the base patch, extended by each contributing module's own
schema patch (an "op":"add" onto the enum array) -- not an unconstrained
string. Dropping the enum removes config-time validation of dialect
names entirely, so a typo or an unregistered dialect now passes schema
validation and only fails at runtime, on every request.

Restores "enum": [ "openai", "anthropic" ] on $defs/binding-ext/llm/dialect
and both dialect descriptions to their prior wording, and restores the
three LlmSchemaValidationTest cases asserting an unregistered dialect is
rejected at config-validation time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…ed dialect enum

957fa22 restored the closed dialect enum and original descriptions on
$defs/binding-ext/llm/dialect, but examples/inspect.schema's golden
schema.expected.json had itself been regenerated in b29b24e to match the
since-reverted (buggy) schema. Regenerated again via the documented
process now that the schema patch is back to its correct shape --
confirmed byte-identical to the golden fixture predating the regression.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…e schema patch

LlmClientIT's signed/signing-unavailable scenarios use test-only dialect
names (test-signed, test-signing-unavailable) that aren't meant to be
part of the binding's real schema. Now that the dialect enum is closed
again, register them through a test-scope BindingExtInfo instead of
leaving them unconstrained, the same pattern used to extend the enum
for a real dialect contributed from outside this module.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M1vMfqKAUfHfMxrBtCb3Lz
…xtractTransform

LlmDialect#supplyExtractor(RESPONSE, envelope) is the contract every
dialect implements to report token usage, and the dialect package is
exported for dialects contributed from outside this module. Yet the
shared machinery for satisfying that contract -- tracking the dot-joined
field path at any depth, deferring envelope writes until the document
completes, blocking the segmentable opt-in -- lived in a package-private
abstract base, so an external dialect had to reimplement all of it just
to map its own native usage paths.

Makes LlmUsageExtractTransform public with a protected constructor, so a
subclass need only implement onField(...) and call the usage setters,
same as LlmRequestFieldTransform already allows for request renames.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VA1HfmUPXyBCFouoLHxxpU
The llm stream could not tell a failed exchange from a successful one: every
failure ended with a clean END. Failure is now carried by frame kind.

- RESET on the app initial stream, with LlmResetEx, rejects a request before
  its reply opens: a non-2xx upstream status (type/message extracted from the
  error body by the target dialect), an upstream reset or undeclared response
  content type (502), a request validation failure or trailing content (400),
  an undeclared request content type (415), and a signer failure (500).
- ABORT on the app reply stream, with LlmAbortEx usage and error, reports a
  response that failed after it started, including an error recorded by the
  dialect's response extractor mid-stream.
- END always means clean completion.

llm.idl adds LlmError { status; type; message; }, LlmResetEx, and an error
on LlmAbortEx; LlmFunctions gains matching builders and matchers.

llm(server) holds the app initial END until the app reply closes, so the
request can still be reset, and maps a RESET with LlmResetEx to an http reply
carrying that status and the dialect's own error body. llm(client) treats the
request as complete once the request pipeline has parsed and validated the
full document, and ends (or signs and sends) the upstream request then.
llm(proxy) forwards RESET extensions in both directions.

Dialect SPI:
- LlmUsageExtractTransform is renamed LlmResponseExtractTransform, gains
  errorStatus/errorType/errorMessage setters, and receives out-of-band event
  names via LlmDialectEvent.
- LlmDialect#errorBody(int, String, String) renders the dialect's native
  error body; LlmStatusReason supplies a fallback message for a status.
- LlmDialect#requestContentType/responseContentTypes declare the content
  types a dialect speaks; attach fails when any has no registered codec, and
  the codec-less request passthrough is removed.

Closes #2607

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GHvmXtGFzehKS4LKMTCA31
…2609)

* feat(engine): verify recorded metric values with the test exporter

The type: test exporter could only assert exported events; it had no way
to observe counter, gauge or histogram values recorded by the engine, so
metric recording could only be checked by reading EngineRule counters
from Java after a script completed.

- test exporter options accept metrics[] expectations by name, binding,
  kind and optional attributes; counters and gauges assert value,
  histograms assert cumulative count and/or per-bucket counts using the
  same power-of-two bucket limits reported by the production exporters
- values are read through the exporter Collector on every export cycle
  and verified when the exporter stops, reporting expected vs actual
- test binding relays begin/data/end/abort/flush/reset extensions in
  both directions and accepts an originType option, so a proxy test
  binding can carry protocol extensions and have a protocol metric group
  attach to it, fully decoupled from any protocol binding
- EngineIT covers the new exporter expectations; SchemaTest covers the
  schema for the new options

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sjn2PgXBGbYAEvqN4LdnB

* fix(engine): adjust gauge metrics by delta

A gauge is a value that can increase and decrease, but the gauge writer
replaced the stored value on each write, so a metric reporting +1 and
-1 changes read 1 or -1 instead of the current level. Gauge writers now
add each accepted value to the stored value, like counter writers, and
the per-worker values continue to be summed when read.

The worker utilization metric is adjusted by +1 and -1 accordingly,
instead of recording the absolute usage. The supplyMetricWriter and
supplyUtilizationMetric contracts document the delta semantics.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sjn2PgXBGbYAEvqN4LdnB

* fix(metrics-http): tolerate begin without extension and decrement active requests by request attributes

Migrate the value-recording coverage from mocked handler unit tests to
k3po ITs that drive real frames through a type: test proxy binding and
assert recorded values with the test exporter. The migrated ITs expose
two defects that the mocks hid:

- with attributes, http.active.requests incremented under the request
  attributes id but decremented under an attributes id recomputed from
  request and response attributes, so neither series returned to zero;
  the decrement now uses the attributes id of the increment.
- http.request.size and http.response.size threw a NullPointerException
  on a begin frame without an http extension, terminating the worker.

binding-http.spec application server scripts reused by the ITs accept
an optional serverAddress property, defaulting to the existing address.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sjn2PgXBGbYAEvqN4LdnB

* test(metrics-grpc): verify recorded values with k3po ITs

Replace the mocked handler unit tests that asserted recorded values
with k3po ITs that relay the binding-grpc.spec application scripts
through a type: test proxy binding and assert the recorded sizes,
message counts, durations and active requests with the test exporter.
Metric metadata unit tests are retained.

binding-grpc.spec application server scripts reused by the ITs accept
an optional serverAddress property, defaulting to the existing address.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sjn2PgXBGbYAEvqN4LdnB

* test(metrics-mcp): verify recorded values with k3po ITs

Replace the mocked handler unit tests that asserted recorded values
with k3po ITs that relay the binding-mcp.spec application scripts
through a type: test proxy binding and assert the recorded counters
and histograms, including tool and outcome attributes, with the test
exporter. Metric metadata unit tests are retained.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sjn2PgXBGbYAEvqN4LdnB

* test(metrics-stream): verify recorded values with k3po ITs

Replace the mocked handler unit tests that asserted recorded values
with k3po ITs that relay the engine.spec application scripts through a
type: test proxy binding and assert the recorded opens, closes, errors,
data and active streams with the test exporter. Metric metadata unit
tests are retained.

engine.spec application server scripts reused by the ITs accept an
optional serverAddress property, defaulting to the existing address.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013sjn2PgXBGbYAEvqN4LdnB

---------

Co-authored-by: Claude <noreply@anthropic.com>
The llm stream factories did not report their origin and routed stream
types, so the engine could not attach metric groups keyed by stream
type to llm bindings. llm(client) and llm(proxy) now report llm as the
origin type, and llm(server) reports http as origin and llm as routed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QCoU9n3wHVMjUFKn8Gjgjz
Adds an llm metric group for llm streams:

- llm.tokens.{input,output,total,cache.read,cache.write,reasoning}
  histograms, one sample per exchange from the LlmUsage on the reply
  END or ABORT extension; an absent (-1) field records nothing
- llm.duration histogram, request begin to reply end
- llm.active.requests gauge, also released when the request is reset
  or aborted before any reply opens

LlmMetricsIT relays the binding-llm.spec application scripts through a
type: test proxy and asserts recorded values with the test exporter.
The reused server scripts accept an optional serverAddress property.

Closes #2605

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QCoU9n3wHVMjUFKn8Gjgjz
The type: test binding forwarded data frame extensions only when a
model envelope was configured, so a plain relay dropped them and a
script expecting zilla:data.ext on the far side never matched. Without
an envelope or model pipeline, the data extension is now copied through
in both directions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QCoU9n3wHVMjUFKn8Gjgjz
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QCoU9n3wHVMjUFKn8Gjgjz
Every llm binding in llm.proxy now records llm.* metrics, exported for
Prometheus on port 7190. verify.py scrapes them around one exchange in
each direction and asserts, at the server, proxy and client binding the
exchange passes through, the upstream's input and output token buckets,
no total recorded for a dialect that never reports one, one duration
sample and no active request left over.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QCoU9n3wHVMjUFKn8Gjgjz
Metric group and direction are exercised by LlmMetricsIT, where the
engine attaches each metric to a real binding; the metadata unit test
now covers names, kinds, units and descriptions only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QCoU9n3wHVMjUFKn8Gjgjz
…lm metrics

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QCoU9n3wHVMjUFKn8Gjgjz
…sitories (#2610)

ZpmCache added the repositories declared by each import BOM to the
collect request unconditionally, so zpm install with
--exclude-remote-repositories still contacted any remote repository a
BOM (or its parent) declared. An artifact missing from the local
repository was then fetched remotely instead of failing, and when that
repository was unreachable the lenient optional-dependency resolution
waited out connect and request timeouts per artifact before silently
returning nothing.

Skip import BOM repositories when remote repositories are excluded, and
warn when optional-dependency resolution is partial or skipped instead
of discarding the cause.


Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

Co-authored-by: Claude <noreply@anthropic.com>
…offline install (#2612)

* feat(manager,zpm-maven-plugin): resolve zpm dependencies ahead of an offline install

zpm install generates the delegate module's module-info with jdeps, which
can only run strictly when the delegated jars' optional and provided
dependencies are available. An offline install (--exclude-remote-repositories)
only sees what was already copied into its local repository, and Maven never
resolves optional or provided dependencies transitively, so the optional
tree was incomplete and jdeps silently fell back to --ignore-missing-deps.

- root the optional dependency tree at the delegated artifacts, pinned to
  the main resolution's versions, so provided dependencies of named modules
  are no longer followed
- tolerate missing or invalid descriptors in the optional tree instead of
  dropping it entirely
- add zpm resolve, which resolves main and optional dependencies into a
  local repository cache and exports every resolved file into a
  self-contained repository for a later offline install
- add zpm install --strict to fail instead of falling back when the
  delegate module references unresolved dependencies
- add build/zpm-maven-plugin with a resolve goal that renders
  zpm.json.template, checks its dependencies are declared by the project,
  and runs zpm resolve in an isolated class loader
- cloud/docker-image assembles the resolved repository instead of a
  hand-maintained list of local repository paths

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

* build(zpm-maven-plugin): depend on the self-contained manager jar only

Exclude manager's transitive dependencies so the plugin realm carries
only the shaded manager jar that the launcher loads in isolation, and
regenerate the build aggregator NOTICE.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

* build(docker-image): declare command runtime dependency

zpm.json.template lists io.aklivity.zilla:command, which the command-*
modules only depend on at provided scope, so docker-image did not
resolve it. The zpm-maven-plugin resolve goal now reports such gaps.
Also add the generated zpm-maven-plugin NOTICE.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

* build(docker-image): install with zpm --strict

Fail the image build when a delegate module would reference unresolved
dependencies instead of silently generating its module-info while
ignoring them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

* fix(zpm-maven-plugin): export the manager jar with the resolved repository

zpmw bootstraps by copying the manager jar from the local repository,
which the resolved repository did not contain because zpm itself is not
a zpm.json dependency, so the offline image build failed with
"Unable to access jarfile .zpm/wrapper/manager-develop-SNAPSHOT.jar".
Export the manager artifact the goal runs alongside the resolved
dependencies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

* fix(manager): record resolved paths from concurrent resolver events

maven-resolver's collector resolves artifact descriptors on a thread
pool, so artifactResolved events reach ZpmCache's listener
concurrently. The listener recorded the paths to export in a
LinkedHashSet, which silently drops entries under concurrent adds, so
zpm resolve could export an incomplete repository, e.g. without a
parent POM that an offline install then needs to collect the optional
tree.

Record resolved paths in a concurrent set. ZpmCacheExportTest fires
16000 artifactResolved events from 8 threads; before the fix only
552-1010 files were exported.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

---------

Co-authored-by: Claude <noreply@anthropic.com>
…2613)

.mvn/maven.config set aether.remoteRepositoryFilter.filterBasedir, which
Maven Resolver 1.9 does not read; the groupId filter reads
aether.remoteRepositoryFilter.groupId.basedir. With no filter file found
for the github repository, every artifact was requested from
maven.packages.aklivity.io before Maven Central, so an unreachable
repository cost a connect timeout per uncached artifact.

The groupId filter also matches exact group ids only, so io.aklivity
alone would have blocked io.aklivity.zilla; list each group served by
that repository instead.


Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

Co-authored-by: Claude <noreply@anthropic.com>
#2614)

* feat(manager): apply maven-resolver configuration from aether.* system properties

zpm built its resolver sessions from its own fixed settings only, so
standard maven-resolver configuration had no effect on it. In particular
a remote repository group-id filter (aether.remoteRepositoryFilter.groupId)
could not stop zpm from requesting every artifact from each configured
repository in turn, and an unreachable repository then cost a connect
timeout per uncached artifact.

zpm now copies -Daether.* system properties into its resolver sessions,
as Maven does, before its own settings. When zpm runs inside Maven via
zpm-maven-plugin it inherits Maven's -D properties, including those from
.mvn/maven.config.

zpm names remote repositories after their host, so this repository adds
a group-id filter file for maven.packages.aklivity.io alongside the one
for Maven's github repository id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

* feat(manager): name zpm.json repositories with Maven repository ids

zpm.json listed repositories as bare URLs, so zpm had no repository id
to use and derived one from each URL's host. That host was also the key
for the settings.xml server holding the repository's credentials, and
for any maven-resolver per-repository configuration such as a group-id
filter file, so neither could be shared with a Maven build of the same
repositories.

zpm.json repositories may now be an object mapping each repository id
to its URL, in resolution order. The id names the remote repository and
selects its settings.xml server, as in Maven. The array form is still
accepted, with the host as the id, as before.

cloud/docker-image names its repositories github and central, so the
group-id filter file for github covers zpm too and the host-named copy
is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015AUJDQphhDuZTjetPaNsBg

---------

Co-authored-by: Claude <noreply@anthropic.com>
…rties the llm binding ignores

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DTRu8xoJU64qAdSohqrKqm
The llm schema accepted vault and catalog on every kind, route when and
with on server and client, and route with on proxy, none of which the
binding reads. It also did not require options on a client, which
refuses every stream without options.dialect and options.server.

Reject those properties and require options on a client, so such a
config fails schema validation at load instead of being silently
ignored. Regenerate the inspect.schema golden schema to match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DTRu8xoJU64qAdSohqrKqm

jfallows commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Analyze (java) (CodeQL) fails on d881d736. It failed the same way when re-run once. The failure isn't caused by this PR's diff.

  • What fails: CodeQL autobuild runs mvn clean package, not install. In cloud/docker-image, zpm:resolve then fails with Could not find artifact io.aklivity.zilla:binding-llm:jar ... in github (https://maven.packages.aklivity.io/), and the same for metrics-llm.
  • Root cause: ResolveMojo resolves only from the local repository and the configured remotes, never from modules built earlier in the reactor. Under package nothing is installed locally, so every in-repo module listed in zpm.json.template has to be downloadable as develop-SNAPSHOT from maven.packages.aklivity.io.
  • Why it depends on timing: binding-llm and metrics-llm come from the stack under this PR (feat(metrics-llm): record llm token usage, duration and active requests #2611 and the PRs beneath it). CodeQL on feat(metrics-llm): record llm token usage, duration and active requests #2611's head 055cef0e, this PR's base, resolved them from that repository on 09-26, but they are no longer there. So feat(metrics-llm): record llm token usage, duration and active requests #2611 would most likely fail the same way if re-run now.
  • Not this PR: this PR only changes the llm schema patch JSON, a schema test and the inspect.schema golden file. Build (25) and every example test pass, including inspect.schema and llm.proxy.
  • Fix: it should clear once feat(metrics-llm): record llm token usage, duration and active requests #2611 merges and develop publishes those snapshots. A durable fix is outside this PR's scope: replace the CodeQL autobuild step with ./mvnw -B install -DskipTests, so zpm:resolve finds reactor modules in the local repository.

Generated by Claude Code

Copy link
Copy Markdown
Contributor Author

This PR now has merge conflicts with develop. develop has started merging the binding-llm stack underneath it (#2505, #2552, #2553, #2556, #2557, #2558, #2568). The conflicting files are all from that stack: binding-llm, binding-llm.spec and binding-llm.conf, including llm.schema.patch.json.

#2611, which this PR is stacked on, conflicts with develop the same way. As planned, I'll rebase this branch onto #2611's branch once #2611 is rebased onto the updated develop. Only this PR's two commits will move across: the schema test commit and the schema fix commit. The fix edits llm.schema.patch.json, so I'll re-check that change and regenerate the inspect.schema golden file after the rebase.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants