Provider audit, SDK major upgrades (openai 3 / anthropic 1 / genai 2), and media + Colab specs - #18
Merged
Merged
Conversation
…ainer) Implement FriendliProvider covering all three FriendliAI inference surfaces from one [providers.friendli] section, selected with endpoint_type: "serverless" (Model APIs, the hosted pay-per-token catalog), "dedicated" (the model field is the endpoint ID, or ID:ADAPTER_ROUTE for Multi-LoRA), and "container" (self-hosted Friendli Engine; base_url required, API key optional). The chat endpoint is OpenAI-compatible, plus the Friendli extensions: reasoning controls (reasoning_effort incl. the Friendli-only "ultracode" tier, reasoning_budget, parse_reasoning, include_reasoning), the chat-template switches enable_thinking / clear_thinking folded into chat_template_kwargs, Friendli Engine sampling (top_k, min_p, min_tokens, repetition_penalty, eos_token, XTC), regex-constrained structured output, cache-aware usage, and the exact /tokenize endpoint. Mutually exclusive body fields (tools vs min_tokens/response_format) are dropped with a warning instead of 422-ing. Transport is selectable via `backend` and auto-resolves openai -> httpx -> sdk. The vendor `friendli` SDK is supported but ranked last on purpose: its generated response models ignore unknown fields, so reasoning_content and reasoning are silently dropped, and it offers no extra_body escape hatch. The provider warns at startup when backend="sdk" meets parse_reasoning, and filters kwargs the SDK cannot type rather than surfacing a TypeError from inside the vendor package. Beyond chat: rich catalog discovery (context, pricing, modalities, reasoning options) cached and primed by warm_up(), tokenize/detokenize/ render_chat, text_completion, transcribe_audio, and — gated to dedicated/container — create_embeddings and generate_image. get_team_cost() and get_team_usage() read the Friendli Suite billing APIs for the configured team, which is also sent as X-Friendli-Team on every request. Token counting stays local (tiktoken) by default: llmcore counts tokens every turn and Model APIs rate limits are tier-based, so exact /tokenize counts are opt-in via native_token_count. Register the provider in ProviderManager with friendliai / friendli_ai aliases, add the llmcore[friendli] extra, the [providers.friendli] config section, and the provider_friendli confy schema section. Validated live against api.friendli.ai: catalog discovery, exact tokenization, chat on all three backends, SSE streaming with reasoning deltas, tool-call round trip, team cost/usage, and 401/429 mapping. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Friendli's GET /serverless/v1/models is a rich catalog rather than the minimal OpenAI /models shape: context length, max completion tokens, per-token pricing (input/output/cache-read/cache-write/audio-minute), a functionality capability block, input/output modalities, reasoning support with the available reasoning_options, the canonical models.dev base_model, the serving mode, and the deprecation date. FriendliAdapter derives nearly every card field from that live data, so the friendli.toml enrichment overlay only carries what the API cannot know: architecture family/type for the open-weight checkpoints Friendli hosts, short display names, and aliases. Pricing is deliberately not pinned in the overlay — the live catalog is authoritative and Friendli adjusts rates. Only the hosted Model APIs catalog is discoverable; Dedicated Endpoints and Container serve a single deployment each and expose no listing endpoint. The adapter also accepts every documented key spelling (FRIENDLI_TOKEN, FRIENDLIAI_API_KEY, FRIENDLI_API_KEY), matching the provider. Registered as "friendli" with a "friendliai" alias. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Seven cards generated from the live Friendli Model APIs catalog with `python -m tools.cardctl generate friendli` (2026-09-20): zai-org/GLM-5.3 1048576 ctx $1.26 / $3.96 per 1M zai-org/GLM-5.3-Flash 1048576 ctx $0.15 / $0.50 (text+image+video) zai-org/GLM-5.2 1048576 ctx $1.40 / $4.40 zai-org/GLM-5.1 202752 ctx $1.40 / $4.40 google/gemma-4-31B-it 262144 ctx $0.14 / $0.40 (text+image) deepseek-ai/DeepSeek-V3.2 163840 ctx $0.50 / $1.50 MiniMaxAI/MiniMax-M2.5 196608 ctx $0.30 / $1.20 Context, pricing, capabilities, modalities and per-model reasoning options come straight from the API. `cardctl diff friendli` reports no differences and all seven validate. Data only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ContextLengthError takes (model_name, limit, actual, message), but the
OpenAI, DeepSeek and Z.ai providers each constructed it with a keyword set
it has never accepted (provider_name / model / max_tokens /
requested_tokens). The raise statement therefore blew up inside __init__:
TypeError: ContextLengthError.__init__() got an unexpected keyword
argument 'provider_name'
Every context-overflow response produced an opaque TypeError carrying no
model and no limit, and nothing catching ContextLengthError — including
llmcore's own context-management and agent retry paths — ever saw it.
Fixing OpenAIProvider also fixes its subclasses (DeepInfra, vLLM, Poe,
OpenRouter). Anthropic, Mistral, Gemini, Kimi and Friendli already used the
documented signature. actual=0 is the faithful translation of the
requested_tokens=None all three were passing.
The defect survived because no test touched those branches, so add
tests/providers/test_context_length_error_mapping.py with two independent
guards:
- A static AST check over src/llmcore asserting that every
ContextLengthError(...) call site uses keywords the constructor accepts.
It is import-free, so it covers providers with no error-path tests and
any added later — this is the guard that would have caught the bug.
- Behavioural tests driving the real chat_completion() failure path of each
fixed provider, asserting the mapped exception carries the model name and
the model's context limit, plus negative cases (a plain 400, a 401) that
must not become ContextLengthError.
Both guards were verified to fail against the pre-fix code.
The behavioural tests skip with an explicit reason when
tests/providers/test_openai_provider.py has already replaced the openai
package in sys.modules with MagicMock placeholders — in that case the
provider binds a mock exception class no except clause can match. The
static check still runs unconditionally.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add docs/Friendli_provider_usage.md covering the three endpoint types, the transport backends (including the measured reason the vendor SDK is not the default — its response models drop reasoning_content), reasoning controls, tool calling and regex structured output, multimodal input, catalog and cardctl workflow, the auxiliary endpoints, token-counting trade-offs, and error/rate-limit mapping. Record the Friendli-only "ultracode" reasoning tier in docs/model_cards.md alongside the canonical vocabulary, add Friendli to the per-provider wire mapping table, and note that it has no "none" tier (reasoning is turned off through chat_template_kwargs.enable_thinking instead). Also note two verified Friendli behaviours callers will hit: /detokenize and /chat/render are documented but currently 404 on Model APIs (they work on Dedicated Endpoints and Container), and tier-0 rate limits are adaptive and in practice allow only a couple of requests per minute — which is why native token counting is opt-in and the example paces its calls. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Audit every curated provider against its upstream SDK/API and record the result in two tracking documents. docs/PROVIDER_SUPPORT_MATRIX.md is the ongoing tracker: per provider, the vendor SDK clone in /av/avalon/xrepos with its tag, commit and date, our pyproject pin, the installed version, the transport shape, and a capability matrix (chat/stream/tools/structured/reasoning/vision/audio/image/video/ embeddings/OCR/search/tokenizer) extracted from the provider classes rather than assumed. Section 6 is a runnable refresh procedure so the document can be regenerated per release. docs/PROVIDER_MODERNIZATION_PLAN.md turns the gaps into a phased program, starting from the dual-transport and one-contract principles. Findings that drove the plan: - openai (2.31 pin vs 3.22.1), anthropic (0.94 vs 1.9.0) and google-genai (1.72 vs 2.25.0) are each a MAJOR version behind. - openai 3.x and anthropic 1.x moved to httpx2 and no longer install httpx. Verified we never hand httpx objects to those clients and that no respx test routes traffic through a vendor SDK, so the port is packaging-only — but six providers (mistral, kimi, poe, openrouter, vllm, huggingface) import httpx with no extra of their own and would fail at import. - httpx2 verifies against the OS trust store, not certifi: a deployment risk worth documenting. - Only 4 of 16 providers implement the full extractor contract; ollama and gemini surface reasoning under provider-specific names that callers cannot use polymorphically. - anthropic's thinking_budget_tokens config key is now rejected with a 400 on every current Claude model; adaptive thinking + output_config.effort is the current API. Model defaults are several generations stale (openai gpt-4o, anthropic claude-sonnet-4-6, ollama llama3). - Gemini's media surface (Imagen, Veo, native TTS, Live API, embeddings) is entirely unexposed; xai/groq/together now ship native SDKs we do not use; mistralai v3.0.0 sits unused while the provider is httpx-only. - OpenAI deprecated the Sora video APIs in 3.1 — recorded so we don't add them. Vendor SDK clones under /av/avalon/xrepos were fast-forwarded as part of this audit, and xai-sdk-python, groq-python and together-python were cloned (they were missing). No llmcore code changes yet — the phases are the follow-up. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Phase 0 of docs/PROVIDER_MODERNIZATION_PLAN.md: unblock the SDK upgrade so the later capability phases have something current to build on. BREAKING: minimum SDK versions move across a major boundary. openai >=2.31.0 -> >=3.0.0,<4 anthropic >=0.94.0 -> >=1,<2 google-genai >=1.72.0 -> >=2,<3 ollama >=0.6.0 -> >=0.6.3 deepgram-sdk >=7.0.0 -> >=7.11.0 zai-sdk >=0.2.0 -> >=0.2.3 The three majors land together because openai 3.x and anthropic 1.x share one breaking change: their HTTP layer moved from httpx to httpx2 (Pydantic's maintained fork), which is now installed in place of httpx and certifi. Two consequences, both handled here: 1. Six providers (mistral, kimi, poe, openrouter, vllm, huggingface) import httpx but had no extra of their own — they worked only because openai installed httpx transitively. Under openai>=3 they fail at import. Each now has an extra declaring what it actually needs, and all six are in [all]. 2. httpx2 verifies TLS against the OS trust store rather than certifi, which can break minimal containers and TLS-inspecting proxies. Documented in CONFIG_REFERENCE.md with the SSL_CERT_FILE / SSL_CERT_DIR escape hatches. No provider code needed porting: llmcore only ever passes numeric timeouts to the vendor clients (never httpx objects), and no respx test routes traffic through a vendor SDK. Both were verified before bumping rather than assumed. Also fixes a latent test-isolation bug this upgrade exposed. Installing zai-sdk flipped ZaiProvider's backend auto-resolution from "openai" to "sdk", bypassing the AsyncOpenAI mocks in 21 tests — the exact hazard the old CI comment described when it deliberately left zai-sdk uninstalled. The tests now pin `backend` explicitly and patch the availability flags for resolution assertions, so they no longer depend on what happens to be installed. That removes the carve-out, and Z.ai's preferred SDK transport is exercised in CI and validated live for the first time. CI now installs .[dev,all] rather than a hand-maintained extras subset, so a new extra is covered the moment it is added to pyproject. The friendli pin stays at >=0.15.1: the vendor repo's pyproject reads 0.15.2 but that version is not published on PyPI. Verified: all 17 provider modules import; full unit suite green (5164 passed, 29 skipped); live calls through OpenAI 3.22.1, Google Gemini 2.25.0 (47 models discovered), Z.ai on the native SDK backend, and DeepSeek. NOT verified live: anthropic 1.9.0 — import- and test-clean, but no ANTHROPIC_API_KEY is available in this environment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two design+specification documents for the capability programs that add subsystems rather than extending providers. Specification only — no code. docs/MEDIA_SUBSYSTEM_SPEC.md turns the provider survey in /av/data/repos/docs/llmcore/researches into an llmcore-side design: a first-class llmcore.media subsystem with MediaArtifact / MediaUsage / MediaJob, capability Protocols per modality, and — the core distinction — three execution classes rather than one, because image generation is request/response, TTS is a byte stream and video is a long-running job. Capability metadata lives in model cards (extended with media/sourcing/policy blocks) instead of provider-specific tables in code; aggregators get the four-way split between who we call, whose weights, whose licence and whose AUP, and "uncensored" is represented honestly as supports_custom_weights plus provider_policy_applies rather than a boolean no vendor actually offers. It also records what the research could not see: - OpenAI's Sora video APIs were DEPRECATED in openai 3.1.0 (confirmed in the vendor CHANGELOG), so the survey's P0 "add Sora" item is dropped; frontier video comes from Veo and fal-hosted models. - llmcore already returns SpeechResult / ImageGenerationResult / OCRResult from seven providers, so models_multimodal types become views over MediaArtifact rather than being replaced. - Deepgram's surface is 12 public methods including a bidirectional voice agent, so the "refactor behind protocols" step is bigger than it looks — and it is the right first migration precisely because it exercises batch, realtime WebSocket and voice agent. - fal's own env var is FAL_KEY while the configured key is FAL_API_KEY; accept both, as the Friendli provider does for its three spellings. - Artifacts carry expires_at and checksum_sha256 from day one: every aggregator returns short-lived URLs, so storing a URI instead of bytes yields dead links. - Long media jobs get idempotency keys so a retried submit cannot double-bill. docs/COLAB_RUNTIME_SPEC.md designs llmcore.runtimes, a remote-compute abstraction with Colab as the first backend, after studying agent-lens's implemented design (391-line spec + ~3,750 lines across 13 modules) and the official google-colab-cli. The factoring argument: agent-lens already ends its bootstrap by registering the endpoint as an llmcore provider, which means the capability is being built on top of llmcore by a consumer and every other consumer must rebuild it. llmcore should own provisioning/bootstrap/tunnel/ lifecycle; agent-lens keeps its CLI and heuristics and deletes the duplication. Because the endpoint vLLM exposes is OpenAI-compatible, no new provider class is required — only dynamic instance registration in ProviderManager, which is also the one capability the media program needs. The safety model is the part that differs from every other provider: a Colab runtime bills per minute from assignment, not per request. Hence explicit-action-only provisioning (LLMCore.create() must never boot a VM), fail-closed bootstrap that releases the VM on any error, orphan detection so an unmonitored VM is visible, and a new max_lifetime_minutes hard cap on top of the reference idle reaper — an idle reaper does not protect against a runtime that is busy in a loop. Also records this session's live validation in the support matrix, including the two results that are not green: Anthropic 1.9.0 authenticates and maps errors correctly through the new major but every request returns "credit balance is too low", so no completion was validated; and Mistral's refreshed key works (46 models, open-mistral-nemo verified) but the configured default mistral-large-latest returns 403 — not in the account's tier. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Partial installations can fail at import time, and streaming error handling plus specification inconsistencies remain unresolved.
Review effort: Balanced
Findings: 1
Open (5)
What changed in this PR
Adds FriendliAI support, upgrades major provider SDK dependencies, fixes context-overflow mapping, and documents future provider, media, and runtime modernization.
Changes:
- Adds FriendliAI transports, configuration, model cards, tests, documentation, and examples.
- Upgrades provider SDK dependencies and CI installation coverage.
- Adds provider-audit, media-subsystem, and Colab-runtime specifications.
| File | Description |
|---|---|
.github/workflows/ci.yml |
Installs all optional dependencies in CI. |
CHANGELOG.md |
Records provider, SDK, and specification changes. |
README.md |
Documents FriendliAI support. |
docs/COLAB_RUNTIME_SPEC.md |
Specifies remote GPU runtimes. |
docs/CONFIG_REFERENCE.md |
Documents TLS and Friendli configuration. |
docs/Friendli_provider_usage.md |
Adds Friendli usage guidance. |
docs/MEDIA_SUBSYSTEM_SPEC.md |
Specifies the media subsystem. |
docs/PROVIDER_MODERNIZATION_PLAN.md |
Defines provider modernization phases. |
docs/PROVIDER_SUPPORT_MATRIX.md |
Tracks provider versions and capabilities. |
docs/model_cards.md |
Documents Friendli cards and reasoning tiers. |
examples/README.md |
Lists the Friendli example. |
examples/friendli_example.py |
Demonstrates Friendli capabilities. |
pyproject.toml |
Upgrades SDKs and adds provider extras. |
src/llmcore/config/default_config.toml |
Adds default Friendli configuration. |
src/llmcore/model_cards/default_cards/friendli/__init__.py |
Initializes Friendli cards. |
src/llmcore/model_cards/default_cards/friendli/MiniMaxAI--MiniMax-M2.5.json |
Adds MiniMax card. |
src/llmcore/model_cards/default_cards/friendli/deepseek-ai--DeepSeek-V3.2.json |
Adds DeepSeek card. |
src/llmcore/model_cards/default_cards/friendli/google--gemma-4-31B-it.json |
Adds Gemma card. |
src/llmcore/model_cards/default_cards/friendli/zai-org--GLM-5.1.json |
Adds GLM-5.1 card. |
src/llmcore/model_cards/default_cards/friendli/zai-org--GLM-5.2.json |
Adds GLM-5.2 card. |
src/llmcore/model_cards/default_cards/friendli/zai-org--GLM-5.3-Flash.json |
Adds GLM-5.3 Flash card. |
src/llmcore/model_cards/default_cards/friendli/zai-org--GLM-5.3.json |
Adds GLM-5.3 card. |
src/llmcore/providers/deepseek_provider.py |
Fixes context-overflow exceptions. |
src/llmcore/providers/friendli_provider.py |
Implements Friendli provider behavior. |
src/llmcore/providers/manager.py |
Registers Friendli and aliases. |
src/llmcore/providers/openai_provider.py |
Fixes context-overflow exceptions. |
src/llmcore/providers/zai_provider.py |
Fixes context-overflow exceptions. |
tests/providers/test_context_length_error_mapping.py |
Adds regression coverage. |
tests/providers/test_friendli_provider.py |
Tests Friendli functionality. |
tests/providers/test_zai_provider.py |
Makes backend tests deterministic. |
tools/cardctl/adapters/__init__.py |
Registers the Friendli adapter. |
tools/cardctl/adapters/friendli_adapter.py |
Maps Friendli catalog metadata. |
tools/cardctl/enrichments/friendli.toml |
Adds curated Friendli metadata. |
tools/llmcore.confy-schema.json |
Adds Friendli configuration fields. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+59
to
+65
| # NOTE on httpx: `openai` 3.x and `anthropic` 1.x moved their HTTP layer to | ||
| # httpx2 (https://httpx2.pydantic.dev/) and NO LONGER install `httpx` or | ||
| # `certifi` transitively. Every extra whose provider imports `httpx` directly | ||
| # must therefore declare it explicitly — previously they got it for free from | ||
| # `openai`. httpx2 also verifies TLS against the OS trust store rather than | ||
| # certifi; see docs/CONFIG_REFERENCE.md if you deploy into minimal containers | ||
| # or behind a TLS-inspecting proxy. |
Comment on lines
+157
to
+161
| class StreamingTTSProvider(Protocol): | ||
| async def stream_tts( | ||
| self, text: str, *, model: str | None = None, voice: str | None = None, | ||
| sample_rate_hz: int | None = None, **kwargs: Any, | ||
| ) -> AsyncIterator[bytes]: ... |
Comment on lines
+1088
to
+1091
| except (ProviderError, ContextLengthError, ValueError): | ||
| raise | ||
| except Exception as e: | ||
| self._raise_error(e, model_name) |
Comment on lines
+63
to
+66
| Live-validated after the upgrade: OpenAI (3.22.1), Google Gemini (2.25.0, 47 | ||
| models discovered), Z.ai (SDK backend), DeepSeek. Full unit suite green (5164 | ||
| passed). **Anthropic 1.9.0 is import- and test-verified but not live-validated — | ||
| no `ANTHROPIC_API_KEY` is available in this environment.** |
|
|
||
| - **Written:** 2026-09-29, from the audit of the same date | ||
| - **Scope:** all 16 providers + the 3 OpenAI-compatible aliases (xai, groq, together) | ||
| - **Status:** plan only — no phase has landed yet |
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Audits every curated provider against its upstream SDK/API, adopts three major SDK versions, and specifies the two capability programs that follow.
1. Audit → two tracking documents
docs/PROVIDER_SUPPORT_MATRIX.md— the ongoing tracker. Per provider: the vendor SDK clone in/av/avalon/xreposwith tag + commit + date, our pyproject pin, the installed version, the transport shape, and a capability matrix extracted from the provider classes rather than assumed. §6 is a runnable refresh procedure; §7 is a live-validation log.docs/PROVIDER_MODERNIZATION_PLAN.md— the phased program, built on two principles: dual transport always and one extractor contract everywhere.SDK clones were fast-forwarded as part of the audit;
xai-sdk-python,groq-pythonandtogether-pythonwere missing and are now cloned.What the audit found
openai2.31 vs 3.22.1,anthropic0.94 vs 1.9.0,google-genai1.72 vs 2.25.0httpx, which is no longer installed — nor iscertifimistral,kimi,poe,openrouter,vllm,huggingfaceimporthttpxwith no extra of their ownollama/geminisurface reasoning under names callers can't use polymorphicallythinking_budget_tokensis rejected with a 400 on every current Claude model — adaptive thinking +output_config.effortreplaced itopenai→gpt-4o,anthropic→claude-sonnet-4-6,ollama→llama3xai-sdk,groq,togetherall ship one; all three are bareOpenAIProvider+base_url2. Phase 0 — adopt the majors (BREAKING)
The three majors land together because openai 3.x and anthropic 1.x share one breaking change (httpx2).
[all].httpxobjects) and that norespxtest routes traffic through a vendor SDK.certifi.CONFIG_REFERENCE.mdgained an "HTTP transport and TLS" section withSSL_CERT_FILE/SSL_CERT_DIR..[dev,all]instead of a hand-maintained subset, so a new extra is exercised the moment it's added.A latent test bug this exposed
Installing
zai-sdkflipped Z.ai's backend auto-resolution fromopenaitosdkand silently bypassed the mocks in 21 tests — exactly the hazard the old CI comment described when it deliberately left the SDK uninstalled. Those tests now pinbackendexplicitly and patch the availability flags for resolution assertions, so they no longer depend on what happens to be installed. The carve-out is gone, and Z.ai's preferred transport is exercised in CI and validated live for the first time.Validation
Full unit suite green (5164 passed). Live: OpenAI 3.22.1 ✅ · Gemini 2.25.0 ✅ (47 models) · Z.ai on the native SDK backend ✅ · DeepSeek ✅ · Mistral ✅ (46 models).
Two results are not green, both account-level rather than code:
invalid_request_error: credit balance is too low. No completion validated. Recorded 🟡 in the matrix, not ✅.mistral-large-latest(llmcore's configured default) returns 403, not in the account's tier;open-mistral-nemoverified working.3. Two specifications (no code)
docs/MEDIA_SUBSYSTEM_SPEC.md— a first-classllmcore.mediasubsystem.MediaArtifact/MediaUsage/MediaJob, capabilityProtocols per modality, and the core distinction: three execution classes, because image generation is request/response, TTS is a byte stream and video is a long-running job. Capability metadata lives in model cards (newmedia/sourcing/policyblocks) instead of provider tables in code; aggregators get the four-way split between who we call, whose weights, whose licence and whose AUP; "uncensored" is represented honestly assupports_custom_weights+provider_policy_appliesrather than a boolean no vendor offers. Nine-phase rollout: Deepgram refactor → OpenAI → Google/Veo → fal → ElevenLabs → Replicate → HF Endpoints → direct specialists. Backward compatibility keepsBaseProvider's five media methods and themodels_multimodaltypes working.Corrections to the source research, which couldn't see the code or the newest changelogs:
openai3.1 — confirmed in the vendor CHANGELOG. The survey's P0 "add Sora" is dropped; frontier video comes from Veo and fal-hosted models.SpeechResult/ImageGenerationResult/OCRResultfrom seven providers → those become views overMediaArtifact, not replacements.expires_atandchecksum_sha256from day one: every aggregator returns short-lived URLs, so storing a URI instead of bytes yields dead links.docs/COLAB_RUNTIME_SPEC.md— allmcore.runtimessubsystem: provision and control remote GPU runtimes (Colab first) and attach the resulting OpenAI-compatible endpoint as a provider instance, so a remotely served model is reachable through the normalllm.chat(provider_name=...).The factoring argument:
agent-lensalready ends its bootstrap by registering the endpoint as an llmcore provider — so the capability is being built on top of llmcore by a consumer, and every other consumer has to rebuild it. llmcore should own provisioning/bootstrap/tunnel/lifecycle; agent-lens keeps its CLI and heuristics and deletes the duplication (~3,750 lines studied). Because vLLM's endpoint is OpenAI-compatible, no new provider class is needed — only dynamic instance registration inProviderManager, which is also the one capability the media program needs.The safety model is what differs from every other provider: a Colab runtime bills per minute from assignment, not per request. Hence explicit-action-only provisioning (
LLMCore.create()must never boot a VM), fail-closed bootstrap that releases the VM on any error, orphan detection so an unmonitored VM is visible, and a newmax_lifetime_minuteshard cap on top of the reference idle reaper — an idle reaper doesn't protect against a runtime that's busy in a loop.4. Follow-ups
mistral-large-latestdefault should change, or the tier upgraded.REPLICATE_API_TOKEN,BFL_API_KEY,LUMA_API_KEY,XAI_API_KEY,GROQ_API_KEY,TOGETHER_API_KEY.colabCLI isn't installed on this machine —uv tool install google-colab-clibefore spec phase R3.🤖 Generated with Claude Code