Skip to content

feat(media): media subsystem core (M1) + dynamic provider registration - #19

Merged
araray merged 9 commits into
mainfrom
av/media_core
Sep 30, 2026
Merged

araray merged 9 commits into
mainfrom
av/media_core

Conversation

@araray

@araray araray commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Implements phase M1 of docs/MEDIA_SUBSYSTEM_SPEC.md — the llmcore.media subsystem — plus the one ProviderManager capability both that program and the remote-runtime program need.

Stacked on av/provider_modernization (#18), which is stacked on av/friendli_provider (#17). New work here is the last commit, c890cde.

No vendor adapters. That's the gate the spec requires: land the primitives first, or each vendor implementation invents its own incompatible conventions.

The design choice that matters

Three execution classes, not one. Image generation is request/response, TTS is a byte stream, video is a long-running job:

result = await llm.media.images.generate("an orange tabby", n=2)      # MediaResult
async for chunk in llm.media.audio.stream_tts("hello"): ...           # AsyncIterator[bytes]
job = await llm.media.video.generate("a drone shot over dunes")       # MediaJob
result = await llm.media.wait(job)

Deepgram's 12 provider-private public methods are what happens without that distinction — it had nowhere else to put realtime audio.

What landed

Module Role
models.py MediaKind, 19 MediaCapability values, MediaExecution, MediaJobStatus, MediaRef, MediaArtifact, MediaProvenance, MediaUsage, MediaResult, MediaJob
protocols.py Runtime-checkable Protocol per capability group + CAPABILITY_PROTOCOLS table
manager.py MediaManager, routers, capability discovery, resolution policy
jobs.py MediaJobManager — polling, capped backoff, timeouts, cancellation
artifacts.py Content-addressed ArtifactStore with materialization policies
testing.py FakeMediaProvider implementing every protocol

Design details worth reviewing:

  • Adapters are the chat providers. Any [providers.*] instance implementing the protocols becomes a media adapter — one credential per vendor, no parallel [media.providers.*] tree to keep in sync. [media.routing] expresses preference order separately.
  • MediaRef accepts url/path/bytes/artifact, so callers never hand-roll base64 and adapters decide what their API wants.
  • MediaArtifact carries expires_at + checksum_sha256. Every aggregator returns short-lived URLs; a caller who stores the URI gets a dead link hours later. The store uses both to decide what to materialize and to dedupe.
  • MediaUsage keeps the vendor's native units (images, megapixels, seconds, audio-minutes, characters, compute-seconds) with a stamped estimate, rather than synthesizing a token count for a video.
  • A timeout does not cancel the job. wait() raises and leaves the handle valid so it can be waited on again — an expensive generation must not be discarded over a client-side deadline. That's also why it's an explicit parameter rather than asyncio.timeout, which would cancel the task.
  • A declared-but-unbacked capability is dropped with a warning, not trusted — better a missing capability than a confident AttributeError at call time.
  • FakeMediaProvider ships inside the package, not under tests/, so downstream projects building adapters can use it. All 134 new tests run against it: no network, no vendor account.

Dynamic provider registration

ProviderManager.register_instance() / unregister_instance() + is_ephemeral(). Providers were only constructible during __init__, but a subsystem that creates an endpoint must add one afterwards — a booted Colab VM's OpenAI-compatible endpoint registers as a vllm instance (COLAB_RUNTIME_SPEC.md §3.2).

Guards: a name collision raises unless replace=True (never silently swap a live provider out from under its callers), the configured default can't be unregistered, construction failures surface as ConfigError, and a failing close() is logged rather than blocking teardown.

A real bug the tests caught

The polling backoff computed 2 ** attempt, which stops converting to float past ~1024 polls and killed the wait loop with OverflowError. A multi-hour video job polled every few seconds actually reaches that. The exponent is now capped, regression tested at 10,000 polls.

Backward compatibility

BaseProvider's five media methods and the models_multimodal result types are untouched, so the seven providers using them keep working. Until M2 migrates them, llm.media.adapter_names is empty and reports so honestly rather than pretending to route.

Verification

134 new tests · full unit suite 5298 passed, 29 skipped · ruff CI gate clean · llm.media exercised end-to-end through LLMCore.create().

Remaining lint in the new code is 4 findings across RUF022/ASYNC109/ASYNC240, all matching substantial pre-existing baselines (49/25/17 repo-wide).

Next

M2 — refactor Deepgram behind the audio protocols. It's the right first migration precisely because it already exercises batch + realtime WebSocket + voice agent, so it validates the hard parts before any new vendor lands.

🤖 Generated with Claude Code

araray and others added 9 commits September 20, 2026 23:00
…ainer)

Implement FriendliProvider covering all three FriendliAI inference surfaces
from one [providers.friendli] section, selected with endpoint_type:
"serverless" (Model APIs, the hosted pay-per-token catalog), "dedicated"
(the model field is the endpoint ID, or ID:ADAPTER_ROUTE for Multi-LoRA),
and "container" (self-hosted Friendli Engine; base_url required, API key
optional).

The chat endpoint is OpenAI-compatible, plus the Friendli extensions:
reasoning controls (reasoning_effort incl. the Friendli-only "ultracode"
tier, reasoning_budget, parse_reasoning, include_reasoning), the
chat-template switches enable_thinking / clear_thinking folded into
chat_template_kwargs, Friendli Engine sampling (top_k, min_p, min_tokens,
repetition_penalty, eos_token, XTC), regex-constrained structured output,
cache-aware usage, and the exact /tokenize endpoint. Mutually exclusive
body fields (tools vs min_tokens/response_format) are dropped with a
warning instead of 422-ing.

Transport is selectable via `backend` and auto-resolves openai -> httpx ->
sdk. The vendor `friendli` SDK is supported but ranked last on purpose: its
generated response models ignore unknown fields, so reasoning_content and
reasoning are silently dropped, and it offers no extra_body escape hatch.
The provider warns at startup when backend="sdk" meets parse_reasoning, and
filters kwargs the SDK cannot type rather than surfacing a TypeError from
inside the vendor package.

Beyond chat: rich catalog discovery (context, pricing, modalities,
reasoning options) cached and primed by warm_up(), tokenize/detokenize/
render_chat, text_completion, transcribe_audio, and — gated to
dedicated/container — create_embeddings and generate_image. get_team_cost()
and get_team_usage() read the Friendli Suite billing APIs for the
configured team, which is also sent as X-Friendli-Team on every request.

Token counting stays local (tiktoken) by default: llmcore counts tokens
every turn and Model APIs rate limits are tier-based, so exact /tokenize
counts are opt-in via native_token_count.

Register the provider in ProviderManager with friendliai / friendli_ai
aliases, add the llmcore[friendli] extra, the [providers.friendli] config
section, and the provider_friendli confy schema section.

Validated live against api.friendli.ai: catalog discovery, exact
tokenization, chat on all three backends, SSE streaming with reasoning
deltas, tool-call round trip, team cost/usage, and 401/429 mapping.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Friendli's GET /serverless/v1/models is a rich catalog rather than the
minimal OpenAI /models shape: context length, max completion tokens,
per-token pricing (input/output/cache-read/cache-write/audio-minute), a
functionality capability block, input/output modalities, reasoning support
with the available reasoning_options, the canonical models.dev base_model,
the serving mode, and the deprecation date.

FriendliAdapter derives nearly every card field from that live data, so the
friendli.toml enrichment overlay only carries what the API cannot know:
architecture family/type for the open-weight checkpoints Friendli hosts,
short display names, and aliases. Pricing is deliberately not pinned in the
overlay — the live catalog is authoritative and Friendli adjusts rates.

Only the hosted Model APIs catalog is discoverable; Dedicated Endpoints and
Container serve a single deployment each and expose no listing endpoint.

The adapter also accepts every documented key spelling (FRIENDLI_TOKEN,
FRIENDLIAI_API_KEY, FRIENDLI_API_KEY), matching the provider.

Registered as "friendli" with a "friendliai" alias.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Seven cards generated from the live Friendli Model APIs catalog with
`python -m tools.cardctl generate friendli` (2026-09-20):

  zai-org/GLM-5.3            1048576 ctx   $1.26 / $3.96 per 1M
  zai-org/GLM-5.3-Flash      1048576 ctx   $0.15 / $0.50   (text+image+video)
  zai-org/GLM-5.2            1048576 ctx   $1.40 / $4.40
  zai-org/GLM-5.1             202752 ctx   $1.40 / $4.40
  google/gemma-4-31B-it       262144 ctx   $0.14 / $0.40   (text+image)
  deepseek-ai/DeepSeek-V3.2   163840 ctx   $0.50 / $1.50
  MiniMaxAI/MiniMax-M2.5      196608 ctx   $0.30 / $1.20

Context, pricing, capabilities, modalities and per-model reasoning options
come straight from the API. `cardctl diff friendli` reports no differences
and all seven validate.

Data only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ContextLengthError takes (model_name, limit, actual, message), but the
OpenAI, DeepSeek and Z.ai providers each constructed it with a keyword set
it has never accepted (provider_name / model / max_tokens /
requested_tokens). The raise statement therefore blew up inside __init__:

    TypeError: ContextLengthError.__init__() got an unexpected keyword
    argument 'provider_name'

Every context-overflow response produced an opaque TypeError carrying no
model and no limit, and nothing catching ContextLengthError — including
llmcore's own context-management and agent retry paths — ever saw it.
Fixing OpenAIProvider also fixes its subclasses (DeepInfra, vLLM, Poe,
OpenRouter). Anthropic, Mistral, Gemini, Kimi and Friendli already used the
documented signature. actual=0 is the faithful translation of the
requested_tokens=None all three were passing.

The defect survived because no test touched those branches, so add
tests/providers/test_context_length_error_mapping.py with two independent
guards:

- A static AST check over src/llmcore asserting that every
  ContextLengthError(...) call site uses keywords the constructor accepts.
  It is import-free, so it covers providers with no error-path tests and
  any added later — this is the guard that would have caught the bug.
- Behavioural tests driving the real chat_completion() failure path of each
  fixed provider, asserting the mapped exception carries the model name and
  the model's context limit, plus negative cases (a plain 400, a 401) that
  must not become ContextLengthError.

Both guards were verified to fail against the pre-fix code.

The behavioural tests skip with an explicit reason when
tests/providers/test_openai_provider.py has already replaced the openai
package in sys.modules with MagicMock placeholders — in that case the
provider binds a mock exception class no except clause can match. The
static check still runs unconditionally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add docs/Friendli_provider_usage.md covering the three endpoint types, the
transport backends (including the measured reason the vendor SDK is not the
default — its response models drop reasoning_content), reasoning controls,
tool calling and regex structured output, multimodal input, catalog and
cardctl workflow, the auxiliary endpoints, token-counting trade-offs, and
error/rate-limit mapping.

Record the Friendli-only "ultracode" reasoning tier in docs/model_cards.md
alongside the canonical vocabulary, add Friendli to the per-provider wire
mapping table, and note that it has no "none" tier (reasoning is turned off
through chat_template_kwargs.enable_thinking instead).

Also note two verified Friendli behaviours callers will hit: /detokenize
and /chat/render are documented but currently 404 on Model APIs (they work
on Dedicated Endpoints and Container), and tier-0 rate limits are adaptive
and in practice allow only a couple of requests per minute — which is why
native token counting is opt-in and the example paces its calls.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Audit every curated provider against its upstream SDK/API and record the
result in two tracking documents.

docs/PROVIDER_SUPPORT_MATRIX.md is the ongoing tracker: per provider, the
vendor SDK clone in /av/avalon/xrepos with its tag, commit and date, our
pyproject pin, the installed version, the transport shape, and a capability
matrix (chat/stream/tools/structured/reasoning/vision/audio/image/video/
embeddings/OCR/search/tokenizer) extracted from the provider classes rather
than assumed. Section 6 is a runnable refresh procedure so the document can
be regenerated per release.

docs/PROVIDER_MODERNIZATION_PLAN.md turns the gaps into a phased program,
starting from the dual-transport and one-contract principles.

Findings that drove the plan:

- openai (2.31 pin vs 3.22.1), anthropic (0.94 vs 1.9.0) and google-genai
  (1.72 vs 2.25.0) are each a MAJOR version behind.
- openai 3.x and anthropic 1.x moved to httpx2 and no longer install httpx.
  Verified we never hand httpx objects to those clients and that no respx
  test routes traffic through a vendor SDK, so the port is packaging-only —
  but six providers (mistral, kimi, poe, openrouter, vllm, huggingface)
  import httpx with no extra of their own and would fail at import.
- httpx2 verifies against the OS trust store, not certifi: a deployment risk
  worth documenting.
- Only 4 of 16 providers implement the full extractor contract; ollama and
  gemini surface reasoning under provider-specific names that callers cannot
  use polymorphically.
- anthropic's thinking_budget_tokens config key is now rejected with a 400 on
  every current Claude model; adaptive thinking + output_config.effort is the
  current API. Model defaults are several generations stale (openai gpt-4o,
  anthropic claude-sonnet-4-6, ollama llama3).
- Gemini's media surface (Imagen, Veo, native TTS, Live API, embeddings) is
  entirely unexposed; xai/groq/together now ship native SDKs we do not use;
  mistralai v3.0.0 sits unused while the provider is httpx-only.
- OpenAI deprecated the Sora video APIs in 3.1 — recorded so we don't add them.

Vendor SDK clones under /av/avalon/xrepos were fast-forwarded as part of this
audit, and xai-sdk-python, groq-python and together-python were cloned (they
were missing). No llmcore code changes yet — the phases are the follow-up.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Phase 0 of docs/PROVIDER_MODERNIZATION_PLAN.md: unblock the SDK upgrade so the
later capability phases have something current to build on.

BREAKING: minimum SDK versions move across a major boundary.

  openai        >=2.31.0  -> >=3.0.0,<4
  anthropic     >=0.94.0  -> >=1,<2
  google-genai  >=1.72.0  -> >=2,<3
  ollama        >=0.6.0   -> >=0.6.3
  deepgram-sdk  >=7.0.0   -> >=7.11.0
  zai-sdk       >=0.2.0   -> >=0.2.3

The three majors land together because openai 3.x and anthropic 1.x share one
breaking change: their HTTP layer moved from httpx to httpx2 (Pydantic's
maintained fork), which is now installed in place of httpx and certifi.

Two consequences, both handled here:

1. Six providers (mistral, kimi, poe, openrouter, vllm, huggingface) import
   httpx but had no extra of their own — they worked only because openai
   installed httpx transitively. Under openai>=3 they fail at import. Each now
   has an extra declaring what it actually needs, and all six are in [all].

2. httpx2 verifies TLS against the OS trust store rather than certifi, which
   can break minimal containers and TLS-inspecting proxies. Documented in
   CONFIG_REFERENCE.md with the SSL_CERT_FILE / SSL_CERT_DIR escape hatches.

No provider code needed porting: llmcore only ever passes numeric timeouts to
the vendor clients (never httpx objects), and no respx test routes traffic
through a vendor SDK. Both were verified before bumping rather than assumed.

Also fixes a latent test-isolation bug this upgrade exposed. Installing zai-sdk
flipped ZaiProvider's backend auto-resolution from "openai" to "sdk", bypassing
the AsyncOpenAI mocks in 21 tests — the exact hazard the old CI comment
described when it deliberately left zai-sdk uninstalled. The tests now pin
`backend` explicitly and patch the availability flags for resolution
assertions, so they no longer depend on what happens to be installed. That
removes the carve-out, and Z.ai's preferred SDK transport is exercised in CI
and validated live for the first time.

CI now installs .[dev,all] rather than a hand-maintained extras subset, so a
new extra is covered the moment it is added to pyproject.

The friendli pin stays at >=0.15.1: the vendor repo's pyproject reads 0.15.2
but that version is not published on PyPI.

Verified: all 17 provider modules import; full unit suite green (5164 passed,
29 skipped); live calls through OpenAI 3.22.1, Google Gemini 2.25.0 (47 models
discovered), Z.ai on the native SDK backend, and DeepSeek.

NOT verified live: anthropic 1.9.0 — import- and test-clean, but no
ANTHROPIC_API_KEY is available in this environment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two design+specification documents for the capability programs that add
subsystems rather than extending providers. Specification only — no code.

docs/MEDIA_SUBSYSTEM_SPEC.md turns the provider survey in
/av/data/repos/docs/llmcore/researches into an llmcore-side design: a
first-class llmcore.media subsystem with MediaArtifact / MediaUsage /
MediaJob, capability Protocols per modality, and — the core distinction — three
execution classes rather than one, because image generation is
request/response, TTS is a byte stream and video is a long-running job.
Capability metadata lives in model cards (extended with media/sourcing/policy
blocks) instead of provider-specific tables in code; aggregators get the
four-way split between who we call, whose weights, whose licence and whose AUP,
and "uncensored" is represented honestly as supports_custom_weights plus
provider_policy_applies rather than a boolean no vendor actually offers.

It also records what the research could not see:

- OpenAI's Sora video APIs were DEPRECATED in openai 3.1.0 (confirmed in the
  vendor CHANGELOG), so the survey's P0 "add Sora" item is dropped; frontier
  video comes from Veo and fal-hosted models.
- llmcore already returns SpeechResult / ImageGenerationResult / OCRResult from
  seven providers, so models_multimodal types become views over MediaArtifact
  rather than being replaced.
- Deepgram's surface is 12 public methods including a bidirectional voice
  agent, so the "refactor behind protocols" step is bigger than it looks — and
  it is the right first migration precisely because it exercises batch,
  realtime WebSocket and voice agent.
- fal's own env var is FAL_KEY while the configured key is FAL_API_KEY; accept
  both, as the Friendli provider does for its three spellings.
- Artifacts carry expires_at and checksum_sha256 from day one: every aggregator
  returns short-lived URLs, so storing a URI instead of bytes yields dead links.
- Long media jobs get idempotency keys so a retried submit cannot double-bill.

docs/COLAB_RUNTIME_SPEC.md designs llmcore.runtimes, a remote-compute
abstraction with Colab as the first backend, after studying agent-lens's
implemented design (391-line spec + ~3,750 lines across 13 modules) and the
official google-colab-cli. The factoring argument: agent-lens already ends its
bootstrap by registering the endpoint as an llmcore provider, which means the
capability is being built on top of llmcore by a consumer and every other
consumer must rebuild it. llmcore should own provisioning/bootstrap/tunnel/
lifecycle; agent-lens keeps its CLI and heuristics and deletes the duplication.

Because the endpoint vLLM exposes is OpenAI-compatible, no new provider class
is required — only dynamic instance registration in ProviderManager, which is
also the one capability the media program needs.

The safety model is the part that differs from every other provider: a Colab
runtime bills per minute from assignment, not per request. Hence
explicit-action-only provisioning (LLMCore.create() must never boot a VM),
fail-closed bootstrap that releases the VM on any error, orphan detection so an
unmonitored VM is visible, and a new max_lifetime_minutes hard cap on top of
the reference idle reaper — an idle reaper does not protect against a runtime
that is busy in a loop.

Also records this session's live validation in the support matrix, including
the two results that are not green: Anthropic 1.9.0 authenticates and maps
errors correctly through the new major but every request returns "credit
balance is too low", so no completion was validated; and Mistral's refreshed
key works (46 models, open-mistral-nemo verified) but the configured default
mistral-large-latest returns 403 — not in the account's tier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Phase M1 of docs/MEDIA_SUBSYSTEM_SPEC.md: the llmcore.media subsystem, reached
through llm.media, plus the one ProviderManager capability that both the media
and remote-runtime programs need. No vendor adapters — that is the gate the
spec requires before any provider work lands, so that each implementation does
not establish its own incompatible conventions.

The central design choice is three execution classes rather than one.
MediaResult covers request/response (image generation, batch ASR),
AsyncIterator[bytes] covers byte streams (TTS, realtime ASR), and MediaJob
covers long-running work (video, queue-based vendors). Deepgram's existing
12-method provider-private surface is what happens without that distinction.

Core pieces:

- models.py — MediaKind, 19 MediaCapability values, MediaExecution,
  MediaJobStatus, MediaRef, MediaArtifact, MediaProvenance, MediaUsage,
  MediaResult, MediaJob. MediaRef accepts url/path/bytes/artifact so callers
  never hand-roll base64. MediaArtifact carries expires_at and
  checksum_sha256 because every aggregator returns short-lived URLs and a
  caller who stores the URI gets a dead link hours later. MediaUsage keeps the
  vendor's own billing units (images, megapixels, seconds, audio minutes,
  characters, compute seconds) with a stamped estimate, rather than
  synthesizing a token count for a video.
- protocols.py — runtime-checkable Protocols per capability group, plus a
  CAPABILITY_PROTOCOLS table so adding a capability is one entry rather than
  an edit to routing code. A capability an adapter declares but does not back
  is dropped with a warning: better a missing capability than a confident
  AttributeError at call time.
- manager.py — MediaManager with per-modality routers, capability discovery,
  and the documented resolution order (explicit provider, explicit model,
  [media.routing], built-in defaults, any capable adapter). Adapters ARE the
  chat providers: any [providers.*] instance implementing the protocols
  becomes one, so there is one credential per vendor and no parallel
  [media.providers.*] tree to keep in sync.
- jobs.py — MediaJobManager owns polling, capped backoff with jitter, timeouts
  and cancellation so no adapter writes its own loop. A timeout raises WITHOUT
  cancelling the job: the handle stays valid and can be waited on again,
  because an expensive generation must not be thrown away over a client-side
  deadline. That is also why the deadline is an explicit parameter rather than
  asyncio.timeout, which would cancel the task.
- artifacts.py — content-addressed store, sharded by SHA-256, atomic publish
  via a .part rename, with always/on_expiry/never policies. The byte fetcher is
  injected, so the store carries no hard httpx dependency.
- testing.py — FakeMediaProvider implements every protocol and ships inside the
  package, so downstream projects building adapters can use it too. All 134 new
  tests run against it: no network, no vendor account.

ProviderManager gains register_instance() / unregister_instance() with
ephemeral tracking. Providers were previously only constructible during
__init__, but a subsystem that creates an endpoint has to add one afterwards —
a booted Colab VM's OpenAI-compatible endpoint registers as a vllm instance
(docs/COLAB_RUNTIME_SPEC.md §3.2). Guards: a name collision raises unless
replace=True so a live provider is never silently swapped out from under its
callers; the configured default cannot be unregistered; construction failures
surface as ConfigError; a failing close() is logged rather than blocking
teardown.

Also fixes a real bug the new tests caught: the polling backoff computed
2 ** attempt, which stops converting to float past ~1024 polls and killed the
wait loop with OverflowError. A multi-hour video job polled every few seconds
reaches that. The exponent is now capped and regression tested at 10,000 polls.

Backward compatible throughout: BaseProvider's five media methods and the
models_multimodal result types are untouched, so the seven providers using them
keep working. Until M2 migrates them, llm.media.adapter_names is empty and says
so honestly rather than pretending to route.

Verified: 134 new tests; full unit suite 5298 passed, 29 skipped; ruff CI gate
clean; llm.media exercised end-to-end through LLMCore.create().

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings September 30, 2026 05:00

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Capability routing, provider lifecycle synchronization, timeout enforcement, and artifact integrity contain unresolved correctness issues.

Review effort: Balanced
Findings: 5 Medium severity

Open (5)
What changed in this PR

Adds the M1 media subsystem and runtime provider registration, atop stacked FriendliAI and provider-modernization changes.

Changes:

  • Adds media models, protocols, routing, jobs, artifacts, and facade integration.
  • Adds dynamic provider registration and fixes context-length exception mapping.
  • Includes FriendliAI configuration, model cards, documentation, and dependency upgrades.
File Description
tools/​llmcore.confy-schema.json Adds Friendli configuration schema.
tools/​cardctl/​enrichments/​friendli.toml Adds Friendli card enrichments.
tools/​cardctl/​adapters/​friendli_adapter.py Maps Friendli catalog metadata.
tools/​cardctl/​adapters/​__init__.py Registers Friendli adapter aliases.
tests/​providers/​test_zai_provider.py Stabilizes backend-selection tests.
tests/​providers/​test_context_length_error_mapping.py Covers context-error mapping.
tests/​media/​test_media_models.py Tests media data models.
tests/​media/​test_media_integration.py Tests facade and provider integration.
tests/​media/​__init__.py Defines media test package.
src/​llmcore/​providers/​zai_provider.py Fixes context exception construction.
src/​llmcore/​providers/​openai_provider.py Fixes context exception construction.
src/​llmcore/​providers/​manager.py Adds Friendli and dynamic registration.
src/​llmcore/​providers/​deepseek_provider.py Fixes context exception construction.
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.3.json Adds GLM-5.3 card.
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.3-Flash.json Adds GLM-5.3 Flash card.
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.2.json Adds GLM-5.2 card.
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.1.json Adds GLM-5.1 card.
src/​llmcore/​model_cards/​default_cards/​friendli/​MiniMaxAI--MiniMax-M2.5.json Adds MiniMax card.
src/​llmcore/​model_cards/​default_cards/​friendli/​google--gemma-4-31B-it.json Adds Gemma card.
src/​llmcore/​model_cards/​default_cards/​friendli/​deepseek-ai--DeepSeek-V3.2.json Adds DeepSeek card.
src/​llmcore/​model_cards/​default_cards/​friendli/​__init__.py Defines Friendli card package.
src/​llmcore/​media/​testing.py Adds reusable fake adapter.
src/​llmcore/​media/​routers.py Adds modality routers.
src/​llmcore/​media/​protocols.py Defines capability protocols.
src/​llmcore/​media/​manager.py Adds media discovery and routing.
src/​llmcore/​media/​jobs.py Adds job polling lifecycle.
src/​llmcore/​media/​artifacts.py Adds artifact persistence.
src/​llmcore/​media/​__init__.py Exports media public API.
src/​llmcore/​exceptions.py Adds media exceptions.
src/​llmcore/​config/​default_config.toml Adds Friendli and media defaults.
src/​llmcore/​api.py Wires media into LLMCore.
README.md Documents Friendli support.
pyproject.toml Updates SDKs and extras.
examples/​README.md Lists Friendli example.
examples/​friendli_example.py Demonstrates Friendli usage.
docs/​model_cards.md Documents Friendli cards.
docs/​Friendli_provider_usage.md Adds Friendli usage guide.
docs/​CONFIG_REFERENCE.md Documents TLS and Friendli config.
docs/​COLAB_RUNTIME_SPEC.md Specifies remote runtimes.
CHANGELOG.md Records stacked changes.
.github/​workflows/​ci.yml Installs all extras in CI.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +164 to +166
suffix = _suffix_for(artifact)
checksum, path = self.put(data, suffix=suffix)
stored = artifact.with_data(data)
Comment thread src/llmcore/media/jobs.py
Comment on lines +236 to +238
delay = min(delay, remaining)
await asyncio.sleep(delay)
job = await self.poll(job)
Comment on lines +117 to +125
adapters: dict[str, MediaCapableProvider] = {}
for name in provider_manager.get_available_providers():
try:
provider = provider_manager.get_provider(name)
except Exception as e:
logger.debug("Skipping provider '%s' during media discovery: %s", name, e)
continue
if isinstance(provider, MediaCapableProvider):
adapters[name] = provider
MediaCapability.IMAGE_GENERATE: ImageGenerationProvider,
MediaCapability.IMAGE_EDIT: ImageEditProvider,
MediaCapability.IMAGE_UPSCALE: ImageUpscaleProvider,
MediaCapability.IMAGE_VARIATE: ImageGenerationProvider,
Comment on lines +782 to +783
previous = self._providers.get(key)
self._providers[key] = provider
@araray araray self-assigned this Sep 30, 2026
@araray
araray merged commit 955d4c3 into main Sep 30, 2026
4 checks passed
@araray
araray deleted the av/media_core branch September 30, 2026 05:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants