Repository navigation
feat(server): agent runtime runner and Anthropic Messages adapter with typed failures and configurable timeout (part 1 of #20) - #108
XxHugheadxX wants to merge 1 commit into
Conversation
AgentRunner runs one prompt-only task (ADR-0004) on any ModelRuntime with the rules shared by every provider: the manifest's input and output max_chars, and a configurable runtime timeout. Failures are typed RuntimeFailure values whose code is safe to store as the hire's failureReason (timeout, provider_error:<type>, refused, output_truncated, input_too_long, output_too_long, empty_output, unsupported_provider). AnthropicRuntime calls the Claude Messages API over HTTP (no official Dart SDK), with model, system prompt and input taken from the task, server-side fallbacks on the models that support them, and status and error type kept for retry decisions. The API key is sent only in x-api-key and never appears in a failure or toString. Nothing outside lib/src/runtime/adapters/ knows the provider API. RuntimeConfig reads PULS3_ANTHROPIC_API_KEY (an unset key disables the runtime) and PULS3_RUNTIME_TIMEOUT_SECONDS (default 120 s, fails fast on an invalid value). Refs #20
Deploying puls3 with
|
| Latest commit: |
c39033c
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://398c3bb4.puls3-4lw.pages.dev |
| Branch Preview URL: | https://feat-20-agent-runtime.puls3-4lw.pages.dev |
TOMOKI977
left a comment
There was a problem hiding this comment.
Great foundation, thanks. The Claude API usage checks out against the docs: model IDs, anthropic-version: 2023-06-01, fallbacks: "default" with server-side-fallback-2026-07-01 only on the four models that support it, thinking blocks ignored, refusal and max_tokens handled. The key never leaks, consumer input stays in the user turn, and the tests pin the exact wire request. Requesting changes for a small set before this lands:
- The timeout doesn't cancel the call (
agent_runner.dart:29).Future.timeoutonly stops waiting: thehttp.postkeeps running and is billed (up to 16000 output tokens), and a retry starts a second paid call while the first is still in flight.package:http1.6 (already our version) supports aborting requests: send anAbortableRequestwith anabortTrigger(seeio_client.dart), and have the runner pass the deadline down to the runtime so it can abort. - The API key name breaks the
PULS3_*convention. In this repoPULS3_*holds public values only (StellarConfigdoc,puls3_server/README.md), and #106'sdocs/infra/secrets.mdplans the LLM key as a Serverpod password (SERVERPOD_PASSWORD_<name>, read throughpasswords.yaml/ Serverpod Cloud secrets). Reading it that way also keeps it out of the shared root.envthat the Flutter app loads with--dart-define-from-file. max_tokensis fixed at 16000 (anthropic_runtime.dart:29) regardless of the manifest'soutput.max_chars. A small output cap still pays for up to 16000 tokens and then fails asoutput_too_long. Deriving it frommaxOutputChars(with a ceiling) caps cost per hire.
Fine as follow-ups for the wiring PR:
- No model allow-list yet: any manifest
modelIdreaches the paid API. I know it's planned as manifest validation in #34. - The 120 s default is short for non-streaming long turns, and nothing ties the timeout to the job's
expired_at(the runner takes a fixed duration). String.lengthcounts UTF-16 code units, so emoji count twice againstmax_chars. Untested for non-ASCII output.pause_turnfalls intounexpected_stop_reason. 408/409 aren't retryable, andretry-afteron 429 is dropped.- No
request-idorusagecaptured, so failures can't be traced and spend per hire isn't measurable. ModelRuntimesits next to the domain'sAgentRuntimePort(ADR-0001, "implemented by the agent runtime (#20)"). A line on howAgentRunnerrelates to it would help.RuntimeConfigthrowsFormatExceptionwhileTrackerLoopConfigthrowsArgumentErrorfor the same kind of bad input.
|
@XxHugheadxX Heads-up: CI changes when #114 (Serverpod 4.0.4) merges. Once it lands on
Specific to this PR: it reads |
|
@TOMOKI977 thanks for the review. I agree with the three blocking points, and #2 is a real flaw in this PR. Here is what changes. Nothing is pushed yet. 1. The timeout does not cancel the call. Confirmed:
2. The API key in
3. Branch up to date with
The push goes out with fixes 1–3. Follow-ups for the wiring PR (I'll track them in #20):
|
Refs #20
This is the first part of #20: the provider-neutral runner and the Anthropic adapter. Wiring the runner to hires, the demo manifests and the testnet end-to-end run come in follow-up PRs. Their prerequisites are in this comment on #20.
Summary
AgentRunner(lib/src/runtime/agent_runner.dart) runs one prompt-only task (ADR-0004) on anyModelRuntime. It applies the rules that hold for every provider:input.max_chars: input over the limit fails before the model is called;output.max_chars;expired_at(ADR-0005 D3/D4).RuntimeFailure(lib/src/runtime/runtime_task.dart) is a sealed set of typed failures. Each has acodethat is safe to store as the hire'sfailureReasonand to show to the client:timeout,provider_error:<type>,refused,output_truncated,input_too_long,output_too_long,empty_output,unsupported_provider. Codes never include provider messages, prompts, inputs or keys.RuntimeProviderFailed.retryableis true only when no response arrived, or the status is 429 or 5xx.AnthropicRuntime(lib/src/runtime/adapters/anthropic_runtime.dart) makes one non-streamingPOST /v1/messagesover HTTP with the existinghttppackage. There is no official Dart SDK, so it adds no new dependency.model,systemand the user input come from the task, so the manifest drives behavior. It setsmax_tokens: 16000and no sampling or thinking parameters.claude-opus-5-5,claude-sonnet-5-5,claude-opus-5andclaude-fable-5-1, it asks for server-side fallbacks on a refusal (fallbacks: "default"plus theserver-side-fallback-2026-07-01beta header).refusalstop reason becomesRuntimeRefusedwith its category, andmax_tokensbecomesRuntimeOutputTruncated.x-api-key. It never appears in a failure or intoString().RuntimeConfig(lib/src/runtime/runtime_config.dart) reads two variables:PULS3_ANTHROPIC_API_KEY: an unset key disables the runtime.PULS3_RUNTIME_TIMEOUT_SECONDS: defaults to 120 s. A value that is not a positive whole number fails fast.Acceptance criteria (#20)
Delivered(nowsubmitted, ADR-0005 D6) with a real result. Follow-up PR.a different manifest prompt changes the request, with no code changecovers it at the adapter level; the manifest-to-task mapping lands with feat: AgentManifest domain model and validation #34.failed(nowruntimeStatus = failedviaHire.failRun, feat: align Hire lifecycle with ERC-8183 #73) comes with the wiring PR.PULS3_RUNTIME_TIMEOUT_SECONDS)..env.example, which waits for the env inventory in chore(infra): testnet configuration and secrets management #106 to avoid a conflict.Verification evidence
I broke each guard one at a time to check that the tests catch it. All 6 changes made the suite fail:
Notes for reviewers
AgentRunner, not in the adapter, so every provider gets the same rule. The 120 s default is a placeholder until ADR-0005 D3 sets the real values.output.max_charsfails the run instead of being truncated, so a client never receives a cut-off result as if it were complete.retryableis exposed so the hire wiring can decide whether to retry within the timeout.