BYOK: pluggable model providers (Gemini, Anthropic, OpenAI, xAI) - #29
Conversation
The AI layer only spoke Gemini: NewFromEnv demanded GEMINI_API_KEY, Generate
wrote the generateContent wire format inline, and both `explain` and `ask`
hardcoded "this sends data to Google" in their consent prompts. Anyone holding
an Anthropic, OpenAI, OpenRouter or xAI key — or running a model locally —
couldn't use it at all.
internal/ai now has a provider seam modelled on charmbracelet/fantasy's
Provider/LanguageModel pair, implemented over net/http rather than depending on
fantasy: it wraps the real vendor SDKs (anthropic-sdk-go, openai-go, genai,
aws-sdk-go-v2) and measured at 23 MB → ~65 MB of binary and 61 → ~170 modules,
for one non-streaming POST in an optional feature. Four providers, two wire
formats, zero new dependencies:
- gemini generateContent — unchanged behavior, default gemini-flash-latest
(a moving alias, so it tracks 3.7 Flash without a re-pin)
- anthropic /v1/messages, default claude-opus-5
- openai /chat/completions, default gpt-5.6-terra — also the compatibility
path for OpenRouter, Groq, Together, DeepSeek, xAI, Mistral and
every local runtime (Ollama, vLLM, LM Studio)
- xai /responses, default grok-4.6
Provider comes from PGBOT_AI_PROVIDER, else auto-detects from whichever key is
present (Gemini first, so existing setups are untouched). PGBOT_AI_MODEL,
PGBOT_AI_BASE_URL, PGBOT_AI_API_KEY and PGBOT_AI_REASONING_EFFORT override;
PGBOT_GEMINI_MODEL/URL still work. Keys are still read only from the
environment, never a flag — now enforced in one place for every provider.
Wire-format details that are load-bearing, each found by probing the live APIs:
- Anthropic rejects `temperature` on current models (400), so the provider never
sends it — Call.Temperature is a hint, not a contract. A refusal is an HTTP 200
with empty content, so stop_reason is checked before reading blocks.
- OpenAI reasoning models reject `max_tokens` ("use max_completion_tokens") and
take reasoning_effort; their cap covers hidden reasoning as well as the answer,
so it is floored at 32k — 8192 returns empty text with finish_reason "length".
Dispatch is by model id, so local runtimes keep the plain shape.
- The Responses API defaults `store` to true, retaining the findings server-side
after the call. pgbot always sends store:false — that retention isn't the
disclosure the user consented to.
- xAI returns errors as {"code":…,"error":"<string>"} where OpenAI nests an
object. Both providers now decode either shape, or the real message is lost.
`explain` and `ask` name the actual destination before sending ("…to xai at
api.x.ai (model grok-4.6)"), and with a local endpoint say nothing leaves the
machine and skip the prompt entirely. The model call also gets its own deadline
instead of the leftovers of the 45s collection budget, which a CPU-bound local
model always blew.
Verified end-to-end against live endpoints: Grok 4.6 via /responses and a local
Ollama via /chat/completions both produced real explanations for `explain` and
`ask`; Anthropic reached auth and surfaced its error; bad-key paths degrade to
the deterministic report with the vendor's message intact. Unit tests cover each
provider's request shape, error handling and env precedence. Binary unchanged at
16 MB, still 61 modules.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
@alexshapalov hey Alex, not sure if you're accepting PRs, but I thought it would be cool to give folks more options on inference providers. I tested the ones listed in the description. I also recently tested Muse 1.3 from Meta using the OpenAI configuration (muse is openai compatible). If this isn't something you'd like to maintain, feel free to close the PR, nbd |
Code reviewFound 5 issues, all fixed in 2c48f04 (pushed to this branch with maintainer edits, on top of a merge of main):
Lines 71 to 83 in 0a8ade2
Lines 101 to 109 in 0a8ade2
Lines 28 to 35 in 0a8ade2
pgbot/internal/ai/anthropic.go Lines 84 to 88 in 0a8ade2
Lines 62 to 68 in 0a8ade2 Also added a cleartext warning on the consent line for a plain-http remote endpoint, and CHANGELOG entries. One thing for the maintainer rather than a defect: the OpenAI default moves from Checked for bugs, consent and key handling, git history, prior PRs and issues (#32 is satisfied by this design without a vendor-specific case), and code-comment guidance; this repo has no CLAUDE.md. Verified the local-endpoint check fails safe against spoofed hostnames, keys travel only in headers, and response bodies are bounded. 🤖 Generated with Claude Code - If this code review was useful, please react with 👍. Otherwise, react with 👎. |
BYOK: pluggable model providers (Gemini, Anthropic, OpenAI, xAI) — #29 plus review fixes
Summary
Expand the existing OpenAI and Gemini integration to support Anthropic, xAI, and OpenAI-compatible services. The implementation uses
net/httpand adds no dependencies.Providers
gemini-flash-latestgenerateContentclaude-opus-5/v1/messagesgpt-5.6-terra/chat/completionsgrok-4.6/responsesThe OpenAI-compatible provider also supports OpenRouter, Groq, Together, DeepSeek, Mistral, Ollama, vLLM, and LM Studio.
Configuration
PGBOT_AI_PROVIDERselects a provider. If unset, pgbot detects one from the available API keys. OpenAI remains first in the detection order to preserve current behavior.General overrides:
PGBOT_AI_MODELPGBOT_AI_BASE_URLPGBOT_AI_API_KEYPGBOT_AI_REASONING_EFFORTExisting
PGBOT_OPENAI_MODEL,PGBOT_OPENAI_URL,PGBOT_GEMINI_MODEL, andPGBOT_GEMINI_URLsettings remain supported. Keys are read only from environment variables.Behavior
store: false.temperature, which current models reject.max_completion_tokensandreasoning_effort./chat/completions.The provider interface follows the
ProviderandLanguageModelstructure fromcharmbracelet/fantasy, implemented directly overnet/http.Verification
Live end-to-end tests against PostgreSQL passed after rebasing onto current upstream
main:gemini-flash-latestclaude-opus-5gpt-5.6-terragrok-4.6via/responsesOllama was also verified through
/chat/completions, including local-endpoint consent behavior.The repository gate passes: build, vet, golangci-lint, all Go tests, and Linux/macOS builds for amd64 and arm64.