Skip to content

feat(providers): add FriendliAI provider (Model APIs, Dedicated Endpoints, Container) - #17

Merged
araray merged 5 commits into
mainfrom
av/friendli_provider
Sep 30, 2026
Merged

araray merged 5 commits into
mainfrom
av/friendli_provider

Conversation

@araray

@araray araray commented Sep 21, 2026

Copy link
Copy Markdown
Owner

Adds a first-class FriendliAI provider, its cardctl adapter and model cards, and fixes a latent ContextLengthError bug found in three existing providers along the way.

FriendliAI provider

One [providers.friendli] section covers all three Friendli inference surfaces, selected with endpoint_type:

endpoint_type Surface model means
serverless (default) Model APIs — hosted pay-per-token catalog catalog model ID
dedicated Dedicated Endpoints — your GPU deployments endpoint ID (ID:ADAPTER_ROUTE for Multi-LoRA)
container self-hosted Friendli Engine whatever the container serves

The chat endpoint is OpenAI-compatible, so streaming, tool calling, sessions, RAG and agents work unchanged. On top of that:

  • Reasoning controls — reasoning_effort (incl. the Friendli-only ultracode tier), reasoning_budget, parse_reasoning, include_reasoning, plus enable_thinking / clear_thinking folded into chat_template_kwargs. Parsed chains of thought surface through extract_reasoning_content() / extract_delta_reasoning_content() in both streaming and non-streaming mode.
  • Friendli Engine sampling — top_k, min_p, min_tokens, repetition_penalty, eos_token, XTC. Mutually exclusive body fields (tools vs min_tokens/response_format) are dropped with a warning instead of 422-ing.
  • Structured output including Friendli's regex response_format; tool calling honours first-class Message.tool_calls (R-2); multimodal input via inline_images / inline_audio / inline_videos / content_parts.
  • Beyond chat — cached catalog discovery primed by warm_up(), tokenize/detokenize/render_chat, text_completion, transcribe_audio, and (dedicated/container only, gated with actionable errors) create_embeddings / generate_image. get_team_cost() / get_team_usage() read the Suite billing APIs; the team ID is sent as X-Friendli-Team on every request.

Transport: direct first, vendor SDK last

backend = "openai" | "httpx" | "sdk", auto-resolving openai → httpx → sdk. The vendor SDK is fully supported but ranked last deliberately — its generated response models ignore unknown fields, so reasoning_content is silently dropped and there is no extra_body escape hatch:

ServerlessChatCompleteSuccess.model_validate({... "reasoning_content": "THINK" ...}).model_dump()
# -> message == {"role": "assistant", "content": "hi"}   # gone

The provider warns at startup when backend="sdk" meets parse_reasoning, and filters kwargs the SDK cannot type rather than surfacing a TypeError from inside the vendor package.

Token counting is local by default

llmcore counts tokens every turn for context budgeting and Model APIs rate limits are tier-based (tier 0 is adaptive — in practice a couple of requests/minute), so exact /tokenize counts are opt-in via native_token_count. tokenize() / detokenize() stay available either way, and a failed native count falls back locally.

Environment variables

Accepts every spelling in circulation — FRIENDLI_TOKEN (the SDK's), FRIENDLIAI_API_KEY (used in friendli.ai's own docs), FRIENDLI_API_KEY; likewise FRIENDLI_TEAM_ID / FRIENDLIAI_TEAM_ID. No renaming needed for existing setups.

Model cards

FriendliAdapter derives nearly every field from Friendli's unusually rich /models catalog — context, max completion tokens, per-token pricing, capability block, modalities, reasoning options, base_model, serving mode, deprecation date. The friendli.toml overlay carries only what the API cannot know (architecture, display names, aliases); pricing is left to the live catalog.

Seven cards generated, cardctl diff friendli clean, all validating:

zai-org/GLM-5.3            1048576 ctx   $1.26 / $3.96 per 1M
zai-org/GLM-5.3-Flash      1048576 ctx   $0.15 / $0.50   (text+image+video)
zai-org/GLM-5.2            1048576 ctx   $1.40 / $4.40
zai-org/GLM-5.1             202752 ctx   $1.40 / $4.40
google/gemma-4-31B-it       262144 ctx   $0.14 / $0.40   (text+image)
deepseek-ai/DeepSeek-V3.2   163840 ctx   $0.50 / $1.50
MiniMaxAI/MiniMax-M2.5      196608 ctx   $0.30 / $1.20

Fix: ContextLengthError in OpenAI / DeepSeek / Z.ai

All three constructed the exception with a keyword set it has never accepted (provider_name / model / max_tokens / requested_tokens) instead of its real (model_name, limit, actual, message), so the raise blew up inside __init__:

TypeError: ContextLengthError.__init__() got an unexpected keyword argument 'provider_name'

Every context overflow produced an opaque TypeError with no model or limit, and nothing catching ContextLengthError — including llmcore's own context-management and agent retry paths — ever saw it. Fixing OpenAIProvider also fixes DeepInfra, vLLM, Poe and OpenRouter. Anthropic, Mistral, Gemini, Kimi and Friendli were already correct.

It survived because no test touched those branches, so tests/providers/test_context_length_error_mapping.py adds two independent guards: an import-free static AST check that every ContextLengthError(...) call site in src/llmcore uses accepted keywords (covering untested and future providers), and behavioural tests driving each fixed provider's real failure path. Both were verified to fail against the pre-fix code.

Validation

Live against api.friendli.ai with real credentials: catalog discovery, exact tokenization, non-streaming chat, SSE streaming with reasoning deltas and usage, all three backends, enable_thinking=False on controllable GLM-5.2, tool call + round trip, team cost/usage, 401/429 mapping, and end-to-end llm.chat() / chat_with_usage() through the facade.

Offline: 128 provider tests + 9 regression tests. Full suite 5162 passed, 29 skipped, 0 failures; ruff blocking gate clean.

Two things to know

  • Friendli's documented /detokenize and /chat/render return 404 on Model APIs today (verified 2026-09-20); they work on Dedicated Endpoints / Container, and every caller degrades gracefully.
  • tests/providers/test_openai_provider.py replaces the openai package in sys.modules with MagicMock at import time. The new behavioural tests detect that and skip with an explicit reason rather than fail on another module's side effect; the static check always runs. Worth cleaning up separately.

Not included

Version left at 0.53.0 with the changelog entry under ## Unreleased — bump to 0.54.0 if you want this to cut a release.

🤖 Generated with Claude Code

araray and others added 5 commits September 20, 2026 23:00
…ainer)

Implement FriendliProvider covering all three FriendliAI inference surfaces
from one [providers.friendli] section, selected with endpoint_type:
"serverless" (Model APIs, the hosted pay-per-token catalog), "dedicated"
(the model field is the endpoint ID, or ID:ADAPTER_ROUTE for Multi-LoRA),
and "container" (self-hosted Friendli Engine; base_url required, API key
optional).

The chat endpoint is OpenAI-compatible, plus the Friendli extensions:
reasoning controls (reasoning_effort incl. the Friendli-only "ultracode"
tier, reasoning_budget, parse_reasoning, include_reasoning), the
chat-template switches enable_thinking / clear_thinking folded into
chat_template_kwargs, Friendli Engine sampling (top_k, min_p, min_tokens,
repetition_penalty, eos_token, XTC), regex-constrained structured output,
cache-aware usage, and the exact /tokenize endpoint. Mutually exclusive
body fields (tools vs min_tokens/response_format) are dropped with a
warning instead of 422-ing.

Transport is selectable via `backend` and auto-resolves openai -> httpx ->
sdk. The vendor `friendli` SDK is supported but ranked last on purpose: its
generated response models ignore unknown fields, so reasoning_content and
reasoning are silently dropped, and it offers no extra_body escape hatch.
The provider warns at startup when backend="sdk" meets parse_reasoning, and
filters kwargs the SDK cannot type rather than surfacing a TypeError from
inside the vendor package.

Beyond chat: rich catalog discovery (context, pricing, modalities,
reasoning options) cached and primed by warm_up(), tokenize/detokenize/
render_chat, text_completion, transcribe_audio, and — gated to
dedicated/container — create_embeddings and generate_image. get_team_cost()
and get_team_usage() read the Friendli Suite billing APIs for the
configured team, which is also sent as X-Friendli-Team on every request.

Token counting stays local (tiktoken) by default: llmcore counts tokens
every turn and Model APIs rate limits are tier-based, so exact /tokenize
counts are opt-in via native_token_count.

Register the provider in ProviderManager with friendliai / friendli_ai
aliases, add the llmcore[friendli] extra, the [providers.friendli] config
section, and the provider_friendli confy schema section.

Validated live against api.friendli.ai: catalog discovery, exact
tokenization, chat on all three backends, SSE streaming with reasoning
deltas, tool-call round trip, team cost/usage, and 401/429 mapping.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Friendli's GET /serverless/v1/models is a rich catalog rather than the
minimal OpenAI /models shape: context length, max completion tokens,
per-token pricing (input/output/cache-read/cache-write/audio-minute), a
functionality capability block, input/output modalities, reasoning support
with the available reasoning_options, the canonical models.dev base_model,
the serving mode, and the deprecation date.

FriendliAdapter derives nearly every card field from that live data, so the
friendli.toml enrichment overlay only carries what the API cannot know:
architecture family/type for the open-weight checkpoints Friendli hosts,
short display names, and aliases. Pricing is deliberately not pinned in the
overlay — the live catalog is authoritative and Friendli adjusts rates.

Only the hosted Model APIs catalog is discoverable; Dedicated Endpoints and
Container serve a single deployment each and expose no listing endpoint.

The adapter also accepts every documented key spelling (FRIENDLI_TOKEN,
FRIENDLIAI_API_KEY, FRIENDLI_API_KEY), matching the provider.

Registered as "friendli" with a "friendliai" alias.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Seven cards generated from the live Friendli Model APIs catalog with
`python -m tools.cardctl generate friendli` (2026-09-20):

  zai-org/GLM-5.3            1048576 ctx   $1.26 / $3.96 per 1M
  zai-org/GLM-5.3-Flash      1048576 ctx   $0.15 / $0.50   (text+image+video)
  zai-org/GLM-5.2            1048576 ctx   $1.40 / $4.40
  zai-org/GLM-5.1             202752 ctx   $1.40 / $4.40
  google/gemma-4-31B-it       262144 ctx   $0.14 / $0.40   (text+image)
  deepseek-ai/DeepSeek-V3.2   163840 ctx   $0.50 / $1.50
  MiniMaxAI/MiniMax-M2.5      196608 ctx   $0.30 / $1.20

Context, pricing, capabilities, modalities and per-model reasoning options
come straight from the API. `cardctl diff friendli` reports no differences
and all seven validate.

Data only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ContextLengthError takes (model_name, limit, actual, message), but the
OpenAI, DeepSeek and Z.ai providers each constructed it with a keyword set
it has never accepted (provider_name / model / max_tokens /
requested_tokens). The raise statement therefore blew up inside __init__:

    TypeError: ContextLengthError.__init__() got an unexpected keyword
    argument 'provider_name'

Every context-overflow response produced an opaque TypeError carrying no
model and no limit, and nothing catching ContextLengthError — including
llmcore's own context-management and agent retry paths — ever saw it.
Fixing OpenAIProvider also fixes its subclasses (DeepInfra, vLLM, Poe,
OpenRouter). Anthropic, Mistral, Gemini, Kimi and Friendli already used the
documented signature. actual=0 is the faithful translation of the
requested_tokens=None all three were passing.

The defect survived because no test touched those branches, so add
tests/providers/test_context_length_error_mapping.py with two independent
guards:

- A static AST check over src/llmcore asserting that every
  ContextLengthError(...) call site uses keywords the constructor accepts.
  It is import-free, so it covers providers with no error-path tests and
  any added later — this is the guard that would have caught the bug.
- Behavioural tests driving the real chat_completion() failure path of each
  fixed provider, asserting the mapped exception carries the model name and
  the model's context limit, plus negative cases (a plain 400, a 401) that
  must not become ContextLengthError.

Both guards were verified to fail against the pre-fix code.

The behavioural tests skip with an explicit reason when
tests/providers/test_openai_provider.py has already replaced the openai
package in sys.modules with MagicMock placeholders — in that case the
provider binds a mock exception class no except clause can match. The
static check still runs unconditionally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add docs/Friendli_provider_usage.md covering the three endpoint types, the
transport backends (including the measured reason the vendor SDK is not the
default — its response models drop reasoning_content), reasoning controls,
tool calling and regex structured output, multimodal input, catalog and
cardctl workflow, the auxiliary endpoints, token-counting trade-offs, and
error/rate-limit mapping.

Record the Friendli-only "ultracode" reasoning tier in docs/model_cards.md
alongside the canonical vocabulary, add Friendli to the per-provider wire
mapping table, and note that it has no "none" tier (reasoning is turned off
through chat_template_kwargs.enable_thinking instead).

Also note two verified Friendli behaviours callers will hit: /detokenize
and /chat/render are documented but currently 404 on Model APIs (they work
on Dedicated Endpoints and Container), and tier-0 rate limits are adaptive
and in practice allow only a couple of requests per minute — which is why
native token counting is opt-in and the example paces its calls.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 21, 2026 02:03

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved provider issues affect context limits, streaming error mapping, retry behavior, and SDK transport coverage.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 3 Medium severity

Open (3)
What changed in this PR

Adds first-class FriendliAI support across multiple deployment surfaces, catalog-backed model cards, and fixes for ContextLengthError construction.

Changes:

  • Adds Friendli provider functionality, configuration, registration, and tests.
  • Adds model-card tooling, seven cards, documentation, examples, and packaging updates.
  • Adds context-error regression coverage and fixes.
File Reviewed change
tools/​llmcore.confy-schema.json Friendli configuration schema
tools/​cardctl/​enrichments/​friendli.toml Friendli card metadata
tools/​cardctl/​adapters/​friendli_adapter.py Catalog adapter
tools/​cardctl/​adapters/​__init__.py Adapter registration
tests/​providers/​test_friendli_provider.py Friendli provider tests
tests/​providers/​test_context_length_error_mapping.py Context-error regression tests
src/​llmcore/​providers/​zai_provider.py Context-error fix
src/​llmcore/​providers/​openai_provider.py Context-error fix
src/​llmcore/​providers/​manager.py Provider registration and aliases
src/​llmcore/​providers/​friendli_provider.py Friendli provider implementation
src/​llmcore/​providers/​deepseek_provider.py Context-error fix
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.3.json Friendli model card
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.3-Flash.json Friendli model card
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.2.json Friendli model card
src/​llmcore/​model_cards/​default_cards/​friendli/​zai-org--GLM-5.1.json Friendli model card
src/​llmcore/​model_cards/​default_cards/​friendli/​MiniMaxAI--MiniMax-M2.5.json Friendli model card
src/​llmcore/​model_cards/​default_cards/​friendli/​google--gemma-4-31B-it.json Friendli model card
src/​llmcore/​model_cards/​default_cards/​friendli/​deepseek-ai--DeepSeek-V3.2.json Friendli model card
src/​llmcore/​model_cards/​default_cards/​friendli/​__init__.py Card package marker
src/​llmcore/​config/​default_config.toml Default Friendli configuration
README.md Provider overview and setup
pyproject.toml Dependencies and extras
examples/​README.md Example index
examples/​friendli_example.py Friendli usage example
docs/​model_cards.md Model-card documentation
docs/​Friendli_provider_usage.md Friendli usage guide
docs/​CONFIG_REFERENCE.md Configuration reference
CHANGELOG.md Release notes
.github/​workflows/​ci.yml CI dependency setup

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.


try:
registry = get_model_card_registry()
card = registry.get(self.get_name(), model_name)
Comment on lines +1119 to +1126
async def stream_wrapper() -> AsyncGenerator[dict[str, Any], None]:
async for chunk in resp: # type: ignore[union-attr]
chunk_dict = self._normalize_obj(chunk)
if self.log_raw_payloads_enabled and logger.isEnabledFor(logging.DEBUG):
logger.debug(
"RAW FRIENDLI STREAM CHUNK: %s", json.dumps(chunk_dict, default=str)
)
yield chunk_dict
Comment on lines +1335 to +1339
raise ProviderError(
self.get_name(),
f"Friendli rate limit exceeded. Model APIs limits scale with your "
f"usage tier (https://friendli.ai/docs/guides/model-apis/rate-limits). "
f"Error: {body}",
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants