Skip to content

Add Bedrock Mantle through the existing Responses client - #33

Closed
edwardsb wants to merge 3 commits into
pgrundev:mainfrom
cavenine:bedrock-mantle
Closed

Add Bedrock Mantle through the existing Responses client#33
edwardsb wants to merge 3 commits into
pgrundev:mainfrom
cavenine:bedrock-mantle

Conversation

@edwardsb

@edwardsb edwardsb commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Lets pgbot use OpenAI GPT models on AWS Bedrock Mantle with an existing bearer token.

  • New explicit provider PGBOT_AI_PROVIDER=bedrock (alias mantle) reusing ResponsesProvider; auth via AWS_BEARER_TOKEN_BEDROCK (overridable with PGBOT_AI_API_KEY).
  • Defaults to openai.gpt-5.6-terra; region precedence AWS_REGION > AWS_DEFAULT_REGION > us-east-1; default base URL https://bedrock-mantle.<region>.api.aws/openai/v1.
  • Reasoning detection handles Bedrock's openai. prefix (incl. gpt-6): omits temperature, floors output at 32,000 tokens, sends store=false. No new dependencies, signing, token caching, or IAM resolution — AWS_PROFILE alone does not authenticate.

Usage:

export PGBOT_AI_PROVIDER=bedrock
export AWS_REGION=us-east-1
export PGBOT_AI_MODEL=openai.gpt-5.6-terra
pgbot ask "What needs attention?" --url "$DATABASE_URL"

Validation:

  • Unit tests added: TestBedrockResponses (endpoint/auth, request shape incl. omitted temperature and 32k floor, for gpt-5.6-terra and gpt-6-astra) and TestBedrockRequiresToken; existing unknown-provider test updated.
  • Live (reporter-validated, us-east-1): pgbot ask succeeded against a local PostgreSQL DB via openai.gpt-5.6-terra through Mantle. openai.gpt-6-astra returned HTTP 404 (model does not exist for that account); /v1/models discovery returned 403 — set PGBOT_AI_MODEL to an ID your account/region serves.
  • Also validated: Anthropic Sonnet 5 through Mantle needs no code changes — PGBOT_AI_PROVIDER=anthropic, PGBOT_AI_MODEL=anthropic.claude-sonnet-5, PGBOT_AI_BASE_URL=https://bedrock-mantle.us-east-1.api.aws/anthropic with a privately supplied token via PGBOT_AI_API_KEY; direct request HTTP 200 and pgbot ask returned findings explanations.

edwardsb and others added 3 commits September 1, 2026 23:27
The AI layer only spoke Gemini: NewFromEnv demanded GEMINI_API_KEY, Generate
wrote the generateContent wire format inline, and both `explain` and `ask`
hardcoded "this sends data to Google" in their consent prompts. Anyone holding
an Anthropic, OpenAI, OpenRouter or xAI key — or running a model locally —
couldn't use it at all.

internal/ai now has a provider seam modelled on charmbracelet/fantasy's
Provider/LanguageModel pair, implemented over net/http rather than depending on
fantasy: it wraps the real vendor SDKs (anthropic-sdk-go, openai-go, genai,
aws-sdk-go-v2) and measured at 23 MB → ~65 MB of binary and 61 → ~170 modules,
for one non-streaming POST in an optional feature. Four providers, two wire
formats, zero new dependencies:

- gemini      generateContent — unchanged behavior, default gemini-flash-latest
              (a moving alias, so it tracks 3.7 Flash without a re-pin)
- anthropic   /v1/messages, default claude-opus-5
- openai      /chat/completions, default gpt-5.6-terra — also the compatibility
              path for OpenRouter, Groq, Together, DeepSeek, xAI, Mistral and
              every local runtime (Ollama, vLLM, LM Studio)
- xai         /responses, default grok-4.6

Provider comes from PGBOT_AI_PROVIDER, else auto-detects from whichever key is
present (Gemini first, so existing setups are untouched). PGBOT_AI_MODEL,
PGBOT_AI_BASE_URL, PGBOT_AI_API_KEY and PGBOT_AI_REASONING_EFFORT override;
PGBOT_GEMINI_MODEL/URL still work. Keys are still read only from the
environment, never a flag — now enforced in one place for every provider.

Wire-format details that are load-bearing, each found by probing the live APIs:

- Anthropic rejects `temperature` on current models (400), so the provider never
  sends it — Call.Temperature is a hint, not a contract. A refusal is an HTTP 200
  with empty content, so stop_reason is checked before reading blocks.
- OpenAI reasoning models reject `max_tokens` ("use max_completion_tokens") and
  take reasoning_effort; their cap covers hidden reasoning as well as the answer,
  so it is floored at 32k — 8192 returns empty text with finish_reason "length".
  Dispatch is by model id, so local runtimes keep the plain shape.
- The Responses API defaults `store` to true, retaining the findings server-side
  after the call. pgbot always sends store:false — that retention isn't the
  disclosure the user consented to.
- xAI returns errors as {"code":…,"error":"<string>"} where OpenAI nests an
  object. Both providers now decode either shape, or the real message is lost.

`explain` and `ask` name the actual destination before sending ("…to xai at
api.x.ai (model grok-4.6)"), and with a local endpoint say nothing leaves the
machine and skip the prompt entirely. The model call also gets its own deadline
instead of the leftovers of the 45s collection budget, which a CPU-bound local
model always blew.

Verified end-to-end against live endpoints: Grok 4.6 via /responses and a local
Ollama via /chat/completions both produced real explanations for `explain` and
`ask`; Anthropic reached auth and surfaced its error; bad-key paths degrade to
the deterministic report with the vendor's message intact. Unit tests cover each
provider's request shape, error handling and env precedence. Binary unchanged at
16 MB, still 61 modules.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
BYOK: pluggable model providers (Gemini, Anthropic, OpenAI, xAI)
PGBOT_AI_PROVIDER=bedrock (alias mantle) reuses ResponsesProvider with
AWS_BEARER_TOKEN_BEDROCK, defaulting to openai.gpt-5.6-terra in
AWS_REGION > AWS_DEFAULT_REGION > us-east-1. Reasoning detection handles
the openai. prefix and gpt-6, omits temperature, and floors output at
32k tokens with store=false.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant