Skip to content

modelcards update: refresh generated cards for 9 providers (+353 / ~140) - #12

Merged
araray merged 1 commit into
mainfrom
av/modelcards_refresh_2026-09
Sep 19, 2026
Merged

araray merged 1 commit into
mainfrom
av/modelcards_refresh_2026-09

Conversation

@araray

@araray araray commented Sep 19, 2026

Copy link
Copy Markdown
Owner

Summary

Data-only refresh of the bundled model cards via the provider management tool (python -m tools.cardctl generate <provider>, no --force, so hand-written source: builtin cards are untouched and only source: generated cards are added or updated).

Stacked on #11 (TypeSafe provider) → #10 (av/26-july_improvements). Once those merge, this PR is purely src/llmcore/model_cards/default_cards/**.

Scope

Provider New cards Updated cards Notes
openai 11 0 e.g. gpt-5.6-*, gpt-6-astra, gpt-image-2.5-flare
google 7 0 e.g. gemini-3.5-flash-lite, gemini-3.6/3.7/3.8-flash, gemini-omni-1.1-flash
zai 3 0 glm-5.3, glm-5.3-flash, glm-5.3-flashx
deepseek 1 0 deepseek-flash
kimi 1 2 kimi-k3; two release_date bumps from the Moonshot listing
poe 69 10 descriptions / context / family updates
openrouter 70 43 pricing (input/output), max_output_tokens, display names
deepinfra 122 80 capabilities (json_mode / tool_use / structured_output / reasoning), pricing, context; 126 deprecated models skipped
huggingface 69 5 parameter_count / family / vision flags
typesafe 0 0 already in sync (cardctl diff typesafe: no differences)
total 353 140

Not refreshed: mistral (the MISTRAL_API_KEY in the env file returns 401 Unauthorized on GET /v1/models — needs a new key), anthropic / xai / qwen (no ANTHROPIC_API_KEY / XAI_API_KEY / DASHSCOPE_API_KEY available), ollama (local models are machine-specific), deepgram (no cardctl adapter).

Verification

  • cardctl validate <provider> on every regenerated directory: all pass (openai 130, google 42, zai 13, deepseek 7, kimi 16, poe 489, openrouter 586, deepinfra 221, huggingface 297 — 0 failures).
  • tests/model_cards + tests/providers re-run on the branch (see the last commit message for counts).

🤖 Generated with Claude Code

Copilot AI lite review requested due to automatic review settings September 19, 2026 04:07

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Provider tool-call precedence currently uses msg.tool_calls or metadata.get("tool_calls"), which can incorrectly resurrect legacy tool calls when msg.tool_calls is an empty list.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 Medium severity

Open (2)
What changed in this PR

This PR primarily refreshes bundled, generated model cards across multiple providers (OpenAI, Google, Poe, OpenRouter, DeepInfra, HuggingFace, etc.). Because it’s stacked on earlier PRs, it also includes supporting runtime changes around prompt registry wiring, tool-calling (R-2) message serialization, and small provider/storage updates.

Changes:

  • Regenerates/updates src/llmcore/model_cards/default_cards/** model card JSONs for many providers.
  • Introduces first-class Message.tool_calls and updates several providers/storage backends to round-trip tool-calling data correctly.
  • Adds/adjusts test fixtures and agent/prompt registry wiring for the grimoire-backed control plane.
File Description
tools/​cardctl/​adapters/​__init__.py Registers TypeSafe adapter and jev alias for card generation/diffing.
tests/​providers/​test_kimi_provider.py Updates/extends Kimi provider tests for non-empty assistant content constraints.
tests/​conftest.py Adds a session-scoped bundled prompt registry fixture for tests.
tests/​agents/​test_persona_and_single_agent.py Injects prompt registry into SingleAgentMode tests; asserts it’s required.
tests/​agents/​test_circuit_breaker_integration.py Wires prompt registry fixture into cognitive cycle construction.
tests/​agents/​observability/​conftest.py Ensures llmcore.shared_events can be imported in isolated observability tests.
src/​llmcore/​models.py Adds Message.tool_calls as a first-class field.
src/​llmcore/​storage/​sqlite_session.py Persists tool-protocol fields via metadata so they round-trip without schema changes.
src/​llmcore/​storage/​constants.py Centralizes :memory: token handling to prevent accidental directory creation.
src/​llmcore/​storage/​chromadb_vector.py Treats :memory: as in-memory storage rather than a filesystem path.
src/​llmcore/​providers/​kimi_provider.py Prefers first-class tool calls and enforces non-empty assistant content placeholder for Moonshot/Kimi.
src/​llmcore/​providers/​mistral_provider.py Prefers first-class tool calls over legacy metadata for assistant messages.
src/​llmcore/​providers/​zai_provider.py Prefers first-class tool calls over legacy metadata for assistant messages.
src/​llmcore/​providers/​anthropic_provider.py Accepts tool calls from first-class field or legacy metadata when building tool-use blocks.
src/​llmcore/​providers/​poe_provider.py Flattens tool-role messages for Poe’s text-only protocol when using native path.
src/​llmcore/​providers/​manager.py Registers TypeSafe provider and jev alias for provider lookup/config.
src/​llmcore/​tracing.py Imports OpenTelemetry trace inside the exception-recording helper.
src/​llmcore/​observability/​events.py Forwards emitted events into the shared-events spine when sinks exist.
src/​llmcore/​agents/​observability/​logger.py Forwards agent events into the shared-events spine when sinks exist.
src/​llmcore/​model_cards/​__init__.py Re-exports model_supports_native_search helper.
src/​llmcore/​memory/​__init__.py Re-exports create_semantiscan_memory_backend.
src/​llmcore/​grimoire_pack/​** Adds bundled pack metadata, spells, runes, and promptlets for control-plane prompting.
src/​llmcore/​model_cards/​default_cards/​** Large generated refresh of bundled model card JSONs across providers.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +535 to +538
# First-class Message.tool_calls (R-2) takes precedence over the
# legacy metadata["tool_calls"] channel.
tool_calls = getattr(msg, "tool_calls", None) or metadata.get("tool_calls")

Comment on lines +666 to +670
# metadata channel so native tool-role results pair correctly.
if role_str == "assistant":
if "tool_calls" in metadata:
msg_dict["tool_calls"] = metadata["tool_calls"]
if not msg.content:
msg_dict["content"] = None
tool_calls = getattr(msg, "tool_calls", None) or metadata.get("tool_calls")
if tool_calls:
msg_dict["tool_calls"] = tool_calls
@araray
araray force-pushed the av/modelcards_refresh_2026-09 branch from cf16a56 to 19c3bee Compare September 19, 2026 04:18
…er, deepinfra, huggingface

Data-only refresh via `python -m tools.cardctl generate <provider>` (no
--force: hand-written `source: builtin` cards untouched; only generated
cards added/updated). 353 new cards, 140 updated:

  openai      +11        google   +7         zai        +3
  deepseek    +1         kimi     +1 / 2     poe        +69 / 10
  openrouter  +70 / 43   deepinfra +122 / 80 (126 deprecated skipped)
  huggingface +69 / 5

Updates carry upstream pricing, context, capability and description
changes (openrouter/deepinfra/poe) plus Moonshot release_date bumps.
Not refreshed: mistral (MISTRAL_API_KEY returns 401), anthropic/xai/qwen
(no keys), ollama (local), deepgram (no adapter); typesafe already in sync.

Verification: `cardctl validate` passes for every regenerated provider
(openai 130, google 42, zai 13, deepseek 7, kimi 16, poe 489, openrouter
586, deepinfra 221, huggingface 297 — 0 failures); the registry loads all
1983 builtin cards; tests/model_cards: 60 passed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@araray
araray force-pushed the av/modelcards_refresh_2026-09 branch from 19c3bee to 925aac2 Compare September 19, 2026 04:22
@araray
araray merged commit 69dccfa into main Sep 19, 2026
1 of 3 checks passed
@araray
araray deleted the av/modelcards_refresh_2026-09 branch September 19, 2026 04:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants