From 6d83fe5499755d112c35f5e6116659b2b1c08b38 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 2 Oct 2026 15:26:20 +0000 Subject: [PATCH 1/2] docs: add frozen LLM compute-cost research notes (2026-10-02), verbatim Co-authored-by: RaapTechllc --- ...-compute-cost-research-notes-2026-10-02.md | 382 ++++++++++++++++++ 1 file changed, 382 insertions(+) create mode 100644 docs/research/llm-compute-cost-research-notes-2026-10-02.md diff --git a/docs/research/llm-compute-cost-research-notes-2026-10-02.md b/docs/research/llm-compute-cost-research-notes-2026-10-02.md new file mode 100644 index 0000000..11f9db3 --- /dev/null +++ b/docs/research/llm-compute-cost-research-notes-2026-10-02.md @@ -0,0 +1,382 @@ +# LLM cost-efficiency benchmark — frozen research notes + +- Status: FROZEN 2026-10-02 (America/Chicago). Research only. No code, no repo changes, no PRs. +- Idea owner: Kyle Raap (RaapTech). Proposed home: `RaapTechllc/minibench`. +- Labels: **FACT** (URL + access date), **ASSUMPTION** (my reasoning, not sourced), **UNKNOWN** (not found or not verified). +- Every URL below was accessed 2026-10-02 unless noted. Prices are cited examples tied to their source date. They are not current quotes. +- **Addenda (folded in 10:18 CT, same day):** (1) output-weighted cost, §8. (2) **Reframe:** the project is now mainly an *investigation of what labs actually pay for inference compute*. The size-over-price ratio and the Pareto chart become estimation tools, not the product. See §9–§12. +- Model names such as "GPT-6 Astra", "Claude Fable 5.1" and "GPT-5.6 Luna" are reproduced exactly as the cited pages print them. I did not check them against vendor pages. + +--- + +## §0 TL;DR + +1. **FACT:** Capability-per-dollar Pareto charts already exist at Artificial Analysis (intelligence vs cost per task), LMArena (Arena score vs blended $/1M), Epoch AI (cost frontier per benchmark) and MIT FutureTech (Gundlach et al.). Plotting price against capability is not a new idea. Dividing parameters by price is uncommon. I found no public board that does it (**UNKNOWN** whether one exists). +2. **ASSUMPTION, supported by the FACTs in §3:** "params ÷ price" is a poor capability proxy. MoE total vs active counts alone can differ by about 18–23× (DeepSeek-V3, gpt-oss-120b). Distillation, quantization, provider margins, and reasoning-token volume also break the link between size and either capability or cost. +3. **FACT:** Closed labs (OpenAI, Google, xAI, Amazon, Mistral) score zero on FMTI's model-information indicators, which include model size. Anthropic does not publish counts either (§2). Published estimators of hidden size are wide: Epoch's inference-economics estimates are "off by a factor of 2", IKP's 90% prediction interval is about 3.2×, and NightVision's parameter error is about 53%. +4. **FACT (repo inspection):** minibench already ships a cost-vs-pass-rate Pareto frontier (`_pareto_frontier` in `backend/app/agents_router.py`, "Capability vs. cost" chart in `frontend/src/pages/Models.tsx`) and a 1:3 in:out blended price. The brief says minibench has "no cost or Pareto work yet". That is wrong. What it lacks is parameter counts. +5. **Verdict:** minibench is a good home for a score-per-dollar view with size shown as a secondary channel. The price-to-size inverse estimator belongs in a separate research repo or notebook (§4.5). +6. **Addendum 1 (FACT §8):** Kyle's premise is that agentic work is output-heavy. By *token count* the opposite holds: OpenRouter averages about 6K prompt vs 400 completion tokens per request, and code requests "routinely exceed 20K input tokens". By *dollars*, output and reasoning can still dominate: published output prices run 1–5× input, and in AA's own example payload reasoning cost exceeds input cost. AA itself moved from a 3:1 blend to a 7:2:1 cache-hit:input:output blend. Use a configurable workload profile, not output-only. +7. **Addendum 2 (FACT §9–§11):** API price ≠ lab cost, and the gap differs by lab. Google owns its silicon (TPUs). xAI owns Colossus and now *leases* it to Anthropic for $1.25B/month (SpaceX S-1, via TechCrunch). Anthropic rents (AWS Trainium, Google TPUs, xAI Colossus). OpenAI rents across Azure, Oracle, AWS and Cerebras, with its own Broadcom chip arriving from late 2026. Hardware ownership is a first-order methodology risk for any price-to-size inversion. +8. **Kyle's hypotheses (§11):** (a) OpenAI uses Cerebras = FACT. "Keeps most of it for itself" is UNKNOWN; the only evidence is a Pro-only Cerebras model. "Cerebras released a $500 plan" is **not supported**: the $500 plan is *OpenAI's* ChatGPT Pro 500 (2026-09-29), and OpenAI does not name the hardware behind it. Cerebras Code plans cost $50 and $200. (b) FACT for both parts: xAI owns its data centres and sells compute to Anthropic. (c) The general principle is supported in direction. Its magnitude is UNKNOWN. + +--- + +## §1 Existing public work on price-performance + +| # | Source | What it measures | Gap vs Kyle's idea | +|---|---|---|---| +| 1.1 | Artificial Analysis (AA) | Intelligence Index plus **cost per task** and cost to run the index, built from input, cache-hit, cache-write, reasoning and answer token prices. Pareto chart of intelligence vs cost per task. | Measured cost × score, not size. Shows open-weights intelligence vs total and active params as a separate chart, never params ÷ price. Closed models have no size. | +| 1.2 | LMArena / Arena | Pareto frontier of Arena score vs **blended $/1M at a 3:1 ratio** (added April 2026). Agent Arena adds cost per task at p25/p50/p95. | Preference Elo, not size. A public comment on the launch post asks for token-normalized cost. | +| 1.3 | Epoch AI | Cost frontier per benchmark over time ("price of thought"), plus parameter and compute database with confidence tiers. | Research outputs, not a live leaderboard. Size is often "Unknown" for closed models. | +| 1.4 | MIT FutureTech (Gundlach et al.) | Benchmark-level price trends (5–10×/yr). Open vs closed and MoE vs dense frontiers. | Academic, snapshot-based. Splits by architecture but does not normalize by size. | +| 1.5 | HF Open LLM Leaderboard | Open models only. Recorded `#Params (B)`, an MoE flag, and an evaluation CO₂ cost. **Retired** in 2025. | No API price. Archived. | +| 1.6 | OpenRouter `/api/v1/models` | Per-token prices (prompt, completion, cache, reasoning), sort by price, throughput, latency, AA intelligence index. | **No parameter field** in the documented schema. | +| 1.7 | the-frontier.app | LMArena Elo vs OpenRouter $/1M, adjustable token ratio and caching. | Third-party mashup. No size. | +| 1.8 | Densing Law (Huang et al. 2025) | Capability per parameter ("capability density") over time. | Per parameter, not per dollar. Contested by IKP for factual knowledge (§2.4). | +| 1.9 | Cost-of-pass (Erol et al. 2025) | Expected $ to get one correct answer. | Closest to a principled score-per-$ metric. Not a live board. | + +Evidence: +- 1.1 **FACT:** AA "Cost per Intelligence Index Task" is a "weighted average cost … calculated from input, cache hit, cache write, reasoning, and answer token prices" — https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3 (accessed 2026-10-02). AA reports a 4-lab Pareto frontier and gives an example: "GPT-6 Astra (max) and Claude Fable 5.1 … both score 53, but … $3.26 and $7.63" per task (same URL). +- 1.1 **FACT:** AA has open-weights charts of intelligence vs total parameters and vs active parameters — https://artificialanalysis.ai/models/open-source (accessed 2026-10-02, via search summary. I did not render the chart myself). +- 1.1 **FACT:** AA v4.1 added cost, time and tokens per task, plus cached-input reporting — https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1 (accessed 2026-10-02). +- 1.2 **FACT:** "Arena score vs. a blended price per 1M tokens (3:1 Ratio)" Pareto, posted 2026-04-01 — https://www.linkedin.com/posts/arenaai_weve-added-pareto-frontier-charts-to-the-activity-7445144386054828033-jYa1 (accessed 2026-10-02). Agent cost per task with p25/p50/p95 — https://arena.ai/blog/agent-categories-and-cost (accessed 2026-10-02). +- 1.2 **UNKNOWN:** whether Arena's "3:1" means input:output or the reverse. The post does not say. +- 1.3 **FACT:** Epoch, "The plunging price of thought" (2026-09-22): the cost of a fixed performance level fell about 47%/quarter (13×/yr). It uses the CAISI truncation method to trace cost-performance curves — https://epoch.ai/publications/the-plunging-price-of-thought (accessed 2026-10-02). +- 1.3 **FACT:** Epoch's database confidence tiers: "Confident" within 3×, "Likely" within 10×, "Speculative" within 30× — https://epoch.ai/data/ai-models (accessed 2026-10-02). +- 1.4 **FACT:** Gundlach et al. find 5–10×/yr price declines at fixed performance. Frontier-run cost rises 3–18×/yr. Open models dominate the low and mid-cost GPQA frontier — https://arxiv.org/abs/2511.23455 (accessed 2026-10-02). +- 1.5 **FACT:** Retirement announcement — https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135. Dataset columns include `#Params (B)`, `MoE`, `CO₂ cost (kg)` — https://huggingface.co/datasets/open-llm-leaderboard/contents. CO₂ is evaluation-time only (8×H100 at 5.6 kW, Virginia grid) — https://huggingface.co/docs/leaderboards/open_llm_leaderboard/emissions (all accessed 2026-10-02). +- 1.6 **FACT:** The OpenRouter models schema shows `pricing`, `context_length`, `architecture` (modality, tokenizer), sort by price, throughput, latency or intelligence. It has no parameter-count field — https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties (accessed 2026-10-02). +- 1.7 **FACT:** https://the-frontier.app/ (accessed 2026-10-02). +- 1.8/1.9 **FACT (secondary citation):** Densing Law, Huang et al. 2025, *Nature Machine Intelligence*, as cited in arXiv 2604.24827. Cost-of-pass, Erol et al., arXiv 2504.13359, as cited in arXiv 2511.23455. I did not open either primary paper (**UNKNOWN** details). +- **UNKNOWN:** a public board that plots parameters ÷ $/token. One targeted search found none: https://llmprice.gitlab.io/ and similar sites are price-only. + +**Gap summary (ASSUMPTION):** Nobody shows on one chart (a) measured score per $, (b) disclosed size for open weights, and (c) *banded* size estimates for closed models with provenance labels. Charts that mix open and closed models either drop size altogether (AA main chart, Arena) or restrict themselves to open weights (AA size charts, HF). + +--- + +## §2 Parameter disclosure and hidden-size estimation + +### 2.1 How open models publish size +- **FACT:** DeepSeek-V3 has 671B total and 37B activated per token. The HF checkpoint is 685B because it adds a 14B MTP module — https://github.com/deepseek-ai/DeepSeek-V3/blob/main/README.md (accessed 2026-10-02). +- **FACT:** gpt-oss-120b has 116.83B total (≈117B) and 5.13B active per token. Unembedding counts as active, embeddings do not — https://arxiv.org/html/2508.10925 (accessed 2026-10-02). +- **FACT:** Grok-1 (xAI, open-weight 2024-03-17) has 314B MoE parameters, 8 experts with 2 active per token — https://x.ai/news/grok-os and https://github.com/xai-org/grok-1 (accessed 2026-10-02). +- **FACT:** NightVision Appendix C tabulates config-derived total and active counts (e.g. Qwen3-30B-A3B is 30.5B / 3.4B) — https://arxiv.org/abs/2607.01313 (accessed 2026-10-02). +- **ASSUMPTION:** Conventions differ between vendors: embeddings in or out, MTP or vision towers in or out, and "active" defined differently. The DeepSeek 671 vs 685 and gpt-oss embedding rule above are examples. A benchmark must store `total_params`, `active_params`, `count_convention`, and `source_url`. + +### 2.2 How closed labs withhold size +- **FACT:** The GPT-4 Technical Report says it gives "no further details about the architecture (including model size)" — https://arxiv.org/abs/2303.08774 (accessed 2026-10-02). +- **FACT:** FMTI 2025: "Amazon, Google, Midjourney, Mistral, OpenAI and xAI do not score any indicators in the model information subdomain … (which includes … model size …)" — https://crfm.stanford.edu/fmti/December-2025/index.html (accessed 2026-10-02). +- **FACT:** Epoch lists Claude 3.5 Sonnet parameters as "Unknown", training compute 2.7e25 FLOP, and price $3 / $15 per 1M in/out — https://epoch.ai/models/claude-3-5-sonnet (accessed 2026-10-02). This is a cited example only. +- **FACT:** xAI open-sourced Grok-1 (above) but scores 14 on FMTI 2025, tied for lowest (FMTI URL above). Opening an old model does not mean disclosing current ones. +- **UNKNOWN:** whether Anthropic or Google have disclosed a parameter count for any current frontier API model. I found none. + +### 2.3 Published methods for inferring size + +| Method | Signal | Reported accuracy | Source | +|---|---|---|---| +| Inference economics (Epoch) | price + tokens/s + assumed GPU markup | "could easily be off by a factor of 2" | Epoch Gradient Update (2024-12) | +| Logit-bias / softmax-bottleneck attack (Carlini et al., ICML 2024) | full or top-k logprobs + logit bias | exact hidden dim for ada/babbage (1024/2048), gpt-3.5-turbo hidden dim recovered | PMLR v235 | +| NightVision (CMU, arXiv 2607.01313) | single-token logprob + TTFT timing | hidden dim ~23% mean rel. err.; params ~53% for ≥3B models; needs >2×10¹¹ tokens | arXiv | +| IKP — Incompressible Knowledge Probes (arXiv 2604.24827) | accuracy on 1,311 obscure facts | LOO median 1.48×; 72% within 2×; 90% PI ≈3.2× | arXiv | +| Press/"mined" estimates (e.g., MEDEC paper) | secondary reporting | none; authors say unverified | arXiv 2412.19260 | +| LeakyLMs per-token timing | timing side channel | recovers layers/dims, not total params | arXiv 2607.20723 | + +Evidence and critique: +- **FACT:** Epoch estimated GPT-4o at about 200B total parameters and Claude 3.5 Sonnet at about 400B. The inputs were serving speed (GPT-4o at 100–150 tok/s vs GPT-4 Turbo at ~55) and price (GPT-4o at $10/1M output in Nov 2024). The authors call it "rough" and possibly "off by a factor of 2" — https://epoch.ai/gradient-updates/frontier-language-models-have-become-much-smaller (accessed 2026-10-02). +- **FACT:** Carlini et al.: "For under $20 … extracts the entire projection matrix of … ada and babbage … hidden dimension of 1024 and 2048." After disclosure, providers added defenses — https://proceedings.mlr.press/v235/carlini24a.html and https://arxiv.org/pdf/2403.06634 (accessed 2026-10-02). +- **FACT:** NightVision notes that providers now expose only single-token logprobs with no logit bias. The paper estimates params using a LLaMA-style formula and was evaluated on models from 135M to 30B. Error exceeds 100% below 3B. MoE serving deployments are untested. The authors put the cost at 12–190× Carlini's — https://arxiv.org/abs/2607.01313 (accessed 2026-10-02). +- **FACT:** IKP calibrates on 93 open models (135M–1.6T, R²=0.910). Total parameters predict MoE knowledge better than active (R² 0.67 vs 0.41). Refusal-heavy models read as lower bounds. Above 1T there are only 2 calibration anchors. It gives banded estimates for 97 proprietary models, e.g. "GPT-5.5 ~4.7T [1.5T–15.1T]" — https://arxiv.org/abs/2604.24827 (accessed 2026-10-02). Code: https://github.com/19PINE-AI/ikp (not opened, **UNKNOWN** license). +- **FACT:** The IKP paper itself describes inference economics as carrying "2×+ uncertainty from hardware, batching, and serving-stack assumptions" (same URL). +- **FACT:** MEDEC lists GPT-4o-mini at about 8B and Claude 3.5 Sonnet at about 175B, mined from public articles and described as unverified — https://arxiv.org/abs/2412.19260 (accessed 2026-10-02, via search summary). Note that this contradicts Epoch's ~400B Sonnet guess, which shows how widely estimates spread. +- **FACT (secondary):** "Inference economics of language models", Erdil 2025, arXiv 2506.04645, as cited in arXiv 2511.23455. I did not open it. + +### 2.4 Reliability critique +- **ASSUMPTION:** Every method above produces a *band*, not a number. At 2–3× error, a closed model's "params ÷ price" value can move by up to about 10× end to end, enough to move it from one end of the Pareto to the other. +- **ASSUMPTION:** Estimating size *from price*, then dividing size *by price*, is circular. The output becomes mostly a function of price plus the assumed markup. This is the main methodological trap in Kyle's "second pass". +- **FACT-based caution:** IKP (knowledge capacity) and NightVision or Epoch (serving cost) measure different things. IKP effective capacity could exceed the compute actually served per token: total vs active, R² 0.67 vs 0.41. +- **ASSUMPTION:** The IKP and NightVision papers are 2026 preprints (April and July). Neither has been independently replicated as far as I found. IKP's own acknowledgements cite an external sanity check that forced a scoring revision. + +--- + +## §3 Metric validity: problems with "params ÷ price" + +Define Kyle's raw metric as `V = P / price` for one value per model. Problems: + +| # | Problem | Evidence | Effect on V | +|---|---|---|---| +| 3.1 | MoE: which P? | DeepSeek-V3 671B/37B (≈18× ratio); gpt-oss-120b 117B/5.1B (≈23× ratio) — FACT counts in §2.1; ratios are my arithmetic | V swings ~20× on a definitional choice | +| 3.2 | Capability per param is not constant over time | Densing Law (reasoning density rising) vs IKP (factual capacity flat, +0.0013/mo, p=0.19) — FACT, arXiv 2604.24827 | Newer small models look "worse value" under V despite being better | +| 3.3 | Distillation / post-training | IKP: refusal tuning cuts scores by tens of pp (Sonnet 4 vs 3.7 Sonnet: 21.5 pp) — FACT, same URL | Same P, very different capability | +| 3.4 | Quantization | NightVision notes production bf16 / reduced precision; Cai et al. 2025 (via IKP) say text tests detect quantization only ~50% of the time — FACT (secondary) | Provider can serve P at lower cost and quality; V can't see it | +| 3.5 | Provider margins / subsidies | DeepSeek reported a *theoretical* 545% daily margin (2025-02-27/28) — FACT, https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md (accessed 2026-10-02); Gundlach et al. treat closed vs open price gap as partly competitive/markup — FACT | Price ≠ cost; V measures pricing strategy | +| 3.6 | Price falls fast | ~13×/yr at fixed performance (Epoch, 2026-09-22) — FACT | V for a model drifts with time; must be timestamped | +| 3.7 | Reasoning tokens / verbosity | Grok 4.6 cost per AA task $0.48 (low) → $2.32 (xhigh) for one model; avg output 10.3k vs 37.6k tokens — FACT, https://aiagentstore.ai/data/model-effort/grok-4-6/aa-intelligence-v4-3-2.csv (third-party mirror of AA data, accessed 2026-10-02) | Same $/token, ~5× cost/task spread across efforts | +| 3.8 | Cached / batch pricing | AA prices cache hit and cache write separately — FACT (§1.1). AA comparison example cache-hit $0.25 vs $1.00 for two models at the same $10 input — FACT, https://artificialanalysis.ai/models/comparisons (accessed 2026-10-02) | Workload-dependent; one "price" is a fiction | +| 3.9 | Input/output blend | Arena uses "3:1"; minibench Usage Board uses **1:3 in:out** — FACT (§1.2, §4) | Ranking order flips with the ratio (worked example §8.3). AA now uses 7:2:1 cache:in:out (FACT, §8.2) | +| 3.10 | Same model, many prices | OpenRouter `/models/{author}/{slug}/endpoints` returns per-provider price/latency/throughput — FACT, OpenRouter docs (§1.6) | Which provider's price? Cheapest vs first-party differ | +| 3.11 | Bigger isn't the goal | ASSUMPTION: maximising P/$ rewards buying the *most weights* per dollar, not the most work done | Wrong objective for the "vibe coder" audience (minibench CONTEXT.md) | + +**Better formulations (ranked by validity, all ASSUMPTION unless cited):** +1. **Score per dollar on a fixed task set:** `pass_rate` (or AA, Arena or minibench score) on the y axis against **measured $ per task** on the x axis, then the Pareto frontier. This is what AA, Arena-Agent and minibench `/models` already do. Matches cost-of-pass (Erol et al.). +2. **Cost to reach a threshold:** cheapest $ at which a model reaches X% (Epoch's cost frontier `C(t,a)`, FACT, §1.3). Stable across blend ratios because it uses actual token counts. +3. **Effort-curve frontier:** use one run per reasoning effort, or the CAISI truncation method, to trace each model's whole cost-score curve rather than a single point (Epoch, FACT §1.3). +4. **Parameters as a secondary channel only:** marker size or colour = `active_params` (open models, FACT tier) or an IKP or NightVision band (closed models, ESTIMATE tier, drawn as an interval). Never on an axis. +5. **Where params ÷ price *is* valid:** for **open-weight models with known P**, `$ per 1M tokens ÷ active_params` across *hosting providers* measures **provider markup and serving efficiency**, not model capability. It answers "who hosts model X cheapest per unit of compute", and it is a defensible use of Kyle's ratio. +6. Always store `as_of`, `blend_ratio`, `price_source`, `cache_assumption`, and `provider`. + +--- + +## §4 RaapTech fit — `RaapTechllc/minibench` (read-only inspection) + +Inspected main @ `f9e2b313d1c103743faa57a5a75adcfc511b683a` (commit time 2026-09-14 10:46 CT), using the GitHub MCP tools and a throwaway read-only shallow clone in `/tmp/mb` on the box. Nothing was pushed. Repo: https://github.com/RaapTechllc/minibench (accessed 2026-10-02). + +### 4.1 Structure (FACT) +- Monorepo: `backend/` (FastAPI + async SQLAlchemy + Postgres, port 3070), `frontend/` (React 19, Vite, TS, Tailwind v4, **Recharts**, port 3071), `cli/` (legacy hardware bench via Ollama), `agentbench/` (Python evaluation runners, graders, MoA, OpenRouter poller), `docs/` (PRDs, ADRs 0001–0004), Docker Compose, 4-job CI. Source: README.md. +- The project describes itself as measuring "whether AI models and agent configurations finish useful work, with completion, cost, latency, and reproducibility receipts". The active focus is the **Real-Work Agent Cabinet** (docs/PROJECT-STATUS.md, reviewed 2026-09-04). +- Surfaces: Solo Cabinet `/models`, Multiplayer `/agents` (cost per task, "$/quarter"), Agent Cabinet, and the OpenRouter **Usage Board** `/usage/cost|task|latency`. The README says they share no composite score by design. + +### 4.2 Existing cost / Pareto work (FACT, contradicts the brief) +- `backend/app/agents_router.py::_pareto_frontier()` computes the "accuracy-vs-cost Pareto frontier", flags each row with `on_pareto_frontier`, and is used on `/api/v1/agents/leaderboard` and `/api/v1/agents/models/leaderboard`. +- `frontend/src/pages/Models.tsx` renders a "Capability vs. cost" scatter (cost per task × pass rate) with "Gold = Pareto-optimal". +- `agentbench/board.py::blended_per_million()`: "USD per 1M tokens at a 1:3 in:out mix". CONTEXT.md defines "Blended price" the same way. +- `agentbench/cost_check.py`: a CI budget guard that caps a suite's worst-case cost on the priciest catalog model (default $15). +- `backend/app/data/known_models_seed.json`: 42 OpenRouter models. Fields: `provider, model_id, display_name, family, license, context_length, prompt_price, completion_price, snapshot_date`. **No parameter fields.** Generated 2026-07-06 from `openrouter.ai/api/v1/models`. +- `agentbench/results/`: 26 committed result artifacts, e.g. `minibench-core-v1-anthropic-claude-haiku-4.5.json`. +- docs/PROJECT-STATUS.md mentions an unfinished July "Value Terminal" proposal on a local branch. It was not accepted. **UNKNOWN** whether its content overlaps this idea. +- Proposed **Benchmark Lens** (ADR 0004, not implemented): "republishes cited external agentic benchmark claims with provenance tiers … adds no composite." + +### 4.3 The three open PRs (FACT, all opened 2026-09-14, all UTC-stamped and converted to CT) +| PR | Title | State | Summary | Interaction with this idea | +|---|---|---|---|---| +| #68 | E2.1: cut legacy hardware/dashboard frontend, add landing page | open, ready (not draft), opened 11:06 CT | Deletes Dashboard, Hardware, Submit, Compare, Leaderboard, and MoaCalculator pages. Adds `Home.tsx`. Nav becomes Home / Agent Cabinet / Models / Agents / Usage / Methodology. | Removes the legacy HEI = (tok/s × quality)/price "value per dollar" chart. Any new value chart should live on `/models` or `/usage`. | +| #69 | E2.2: cut legacy hardware API, make backend run on SQLite | draft, opened 12:13 CT | Deletes `Benchmark`, `HardwareSpec`, `ModelQuality` models and `/api/v1/{leaderboard,hardware,models,…}` routes. Adds SQLite dev support. | Removes legacy `/api/v1/models` (quality table). A parameter catalog would extend `KnownModel` and the seed JSON instead. Merge conflicts are likely if built before #69 lands (ASSUMPTION). | +| #70 | Propose oddball task categories (data security and constrained creative writing) | draft, docs only, opened 12:15 CT | New categories, a `constraint_pack` grader, and a separate "Taste" human-vote evidence class. | No direct overlap. Its rule that each evidence class stays separate and never enters a score supports keeping size estimates in a separate tier. | + +### 4.4 How it could slot in (ASSUMPTION — design sketch, not a change) +- **Data:** add optional `total_params_b`, `active_params_b`, `params_source_url`, `params_tier` (`disclosed` | `estimated:` | `unknown`), and `params_band_low/high` to the `KnownModel` seed. Open models get `disclosed` with a model-card URL. Closed models get `unknown` by default, or an `estimated:ikp` band with citation. +- **Surface A, low risk:** on `/models` "Capability vs. cost", encode marker size or opacity by `active_params` when the tier is `disclosed`. Closed or unknown models get a hollow marker. Axes and Pareto logic stay as they are. +- **Surface B:** in the Usage Board, add a "best-by-$-per-active-param" compare route for open weights only. This is the provider-markup view from §3 item 5. It fits Mode A because it needs only `/models` plus a static param table. +- **Surface C:** the Benchmark Lens (ADR 0004) is the natural place to republish external size *estimates* (Epoch, IKP) with provenance tiers. Its stated design is cited, tiered and never merged into a composite. +- **Sequencing:** wait for #68 and #69 to land (or be closed) before touching the frontend or the backend models. + +### 4.5 Fit verdict +- **Good fit** for score-per-dollar with size as an annotation (sparks 1 and 2). The Pareto code, blended price, OpenRouter catalog and provenance culture (citations, `as_of`, fixture-vs-live labels) already exist. +- **Poor fit** for "infer hidden params from price" as a headline metric. It conflicts with the repo's "no composite" and provenance rules, and §2.4 shows it is circular. Alternatives: a standalone `RaapTechllc/price-implied-size` research repo or notebook, or a blog post. Its outputs can then be fed into minibench's Benchmark Lens as an `estimated` tier. + +--- + +## §5 Idea sparks (ranked; see §12 for the re-rank under the reframe) + +**#1 — "Value Frontier with Size Lens" inside minibench (recommended)** +- Score per $ (measured cost per task) Pareto. Marker encodes `active_params` for disclosed open weights and a hollow or banded marker for closed models. Toggles for blend ratio and cache assumption. +- Why first: reuses `_pareto_frontier`, Recharts and the seed catalog. It is valid under §3 and visibly different from AA and Arena because size is shown with provenance. +- Risk: limited novelty (ASSUMPTION). The size channel is the differentiator. + +**#2 — Open-Weight Provider Markup Index** +- For each open model with disclosed P, compute `$ per 1M tokens / active_params` for each OpenRouter provider endpoint. Rank providers and track over time. +- Why: Kyle's ratio on the one population where it is valid (P is a FACT). It answers a real buyer question: who is overcharging for the same weights? +- Gaps: the endpoints API needs live polling (Mode A allows only the listed Data API paths, so `/models/{id}/endpoints` would need an ADR change, **UNKNOWN** whether permitted). Quantization differences between providers confound results (§3.4). + +**#3 — Price-Implied Size estimator, published as research (not a leaderboard)** +- Fit `log(price)` against `log(active_params)` on open-weight models using cheapest-provider prices. Invert the fit for closed models and report 90% intervals. Cross-check against IKP bands and Epoch's guesses, and publish the disagreement. +- Why: a novel, shareable artifact for RaapTech. It is honest only if presented as "what the price implies under assumption X". +- Risk: circular (§2.4). Closed-lab markup is unobservable. Expect at least 2–3× error. Keep it separate from minibench scores. + +--- + +## §6 Known gaps and unverified items + +- **UNKNOWN:** whether any public site plots params ÷ price directly. One search found none. A broader search, including GitHub, was not done. +- **RESOLVED (addendum work):** AA's blended price is now 7:2:1 cache-hit:input:output (§8.2). That explains the $7.175 figure. **Still UNKNOWN:** whether Arena's "3:1" means input:output. It probably mirrors AA's legacy 3:1 input:output (ASSUMPTION). +- **Unverified:** AA open-weights parameter charts (search summary only, not rendered). AA's per-task figures come partly from a third-party CSV mirror (aiagentstore.ai), not AA's own export. +- **Unverified:** MEDEC estimate details (search summary only). The Densing Law and Cost-of-pass primary papers were not opened (secondary citations). +- **Unverified:** IKP code license and probe availability. The paper says probes are released, but the Limitations section says they "should remain private", which is internally inconsistent. +- **UNKNOWN:** current disclosed parameter counts for any Anthropic, Google, OpenAI or xAI frontier API model as of 2026-10. None found. +- **UNKNOWN:** whether OpenRouter's per-provider endpoint data records quantization, and whether minibench's Mode A policy (ADR 0002/0003) can be extended to `/endpoints`. +- **UNKNOWN:** the status and contents of PR #67 (the "era 2 cut" plan referenced by #68 and #69). I did not open it. Also the content of the unaccepted "Value Terminal" branch, which exists only locally per PROJECT-STATUS.md. +- **Not done:** I did not use the `grok` CLI. WebSearch and WebFetch were enough. I fetched no vendor pricing pages, so no current per-token prices are asserted here beyond the cited, dated examples. +- **Addenda gaps:** Many deal figures come from search summaries or secondary press (Reuters/The Information, TechCrunch reading the S-1). I did not read the S-1 itself. + - **UNKNOWN:** per-token costs of TPU, Trainium, Cerebras or Broadcom silicon. + - **UNKNOWN:** OpenAI's internal routing of Cerebras capacity. + - **UNKNOWN:** any Microsoft/Azure–OpenAI or Anthropic–Azure terms (not researched). + - **UNKNOWN:** a per-category in:out token ratio for agentic loops. + - The §10.1 GPU-hour figure is illustrative arithmetic. +- **Caveat:** many 2026 model names and figures come from single sources (AA, Epoch, the IKP preprint) and were not cross-checked against vendor announcements. + +## §8 Addendum 1: output-weighted cost + +### 8.1 Published input:output price ratios (cited examples, each dated) +| Model / offer | In $/1M | Out $/1M | Out÷In | Source & date | +|---|---|---|---|---| +| Claude 3.5 Sonnet | 3 | 15 | 5× | Epoch model page, accessed 2026-10-02 — https://epoch.ai/models/claude-3-5-sonnet | +| GPT-4 (OpenRouter docs sample) | 30 | 60 | 2× | OpenRouter API docs example payload, accessed 2026-10-02 (§1.6 URL) | +| "Claude Fable 5.1" / "GPT-6 Astra" (AA comparison) | 10 | 50 | 5× | https://artificialanalysis.ai/models/comparisons, accessed 2026-10-02 | +| Third model on same AA page (name not captured) | 1.25 | 4.25 | 3.4× | same | +| AA Data API example payload (model unnamed) | 0.06 | 0.20 | 3.3× | https://artificialanalysis.ai/data-api/docs, accessed 2026-10-02 | +| Qwen3.7 Max (OpenRouter snapshot) | 1.25 | 3.75 | 3× | minibench `known_models_seed.json`, snapshot_date 2026-05-21 | +| Cerebras Qwen3-Coder at Cerebras Code launch | 2 | 2 | 1× | secondary: https://tomrochette.com/agents/model-access/cerebras-code/ (says launch price 2025-08; current Cerebras rate card not extractable, **unverified**) | +- **FACT:** Cache-hit prices are far below input prices: $0.25 vs $1.00 for two models that both charge $10 input (AA comparisons). AA Data API example: cache hit $0.015, input $0.06, cache write $0.075. +- **ASSUMPTION:** Across these examples output costs 1–5× input. Output-weighting therefore matters most for 5× vendors and least for flat-priced fast-silicon hosts. + +### 8.2 Token-mix evidence +- **FACT:** AA methodology: "we calculate a blended price assuming a 7:2:1 ratio of cache hit, input, and output tokens" — https://artificialanalysis.ai/methodology (accessed 2026-10-02). The Data API still exposes `price_1m_blended_3_to_1` alongside `price_1m_blended_7_to_2_to_1` — https://artificialanalysis.ai/data-api/docs (accessed 2026-10-02). Check: (7×0.25 + 2×10 + 50)/10 = 7.175, which matches the AA comparisons figure in §1. +- **FACT:** OpenRouter/a16z "State of AI" (100T tokens): average prompt tokens per request rose from about 1.5K to over 6K, and completions from about 150 to 400. Programming "routinely exceed[s] 20K input tokens". Completion growth comes "mostly due to reasoning tokens". OpenRouter counts reasoning tokens within completion tokens — https://arxiv.org/abs/2601.10088 (PDF read 2026-10-02). +- **ASSUMPTION (arithmetic on the FACT above):** 6K:400 is about 15:1 input:output *by tokens*. Agentic and code traffic is input-heavy in volume, contrary to the addendum's premise. +- **FACT:** AA Data API example payload (one unnamed model, full Intelligence Index): input 140.3M tokens and output 61.3M, of which reasoning is 58.3M and answer 3.0M. Cost fields in that payload show reasoning about $11.66 and input about $8.42 of a $20.69 total. The cost field names were partly truncated in the fetch, so treat the breakdown as **unverified** — https://artificialanalysis.ai/data-api/docs (accessed 2026-10-02). +- **ASSUMPTION (arithmetic):** In that example, reasoning is about 95% of output tokens. Input:output is about 2.3:1 by tokens, yet output still makes up about 56% of cost. So Kyle's intuition holds *in dollars* for reasoning-heavy evaluations even though tokens are input-heavy. +- **FACT:** Effort setting alone moves output volume about 3.7× for one model (Grok 4.6: about 10.3K output tokens per task at low, 37.6K at xhigh; cost per task $0.48 to $2.32). Source: third-party mirror of AA data (§3.7 URL). +- **UNKNOWN:** a public token-mix breakdown specifically for *agentic tool-call loops*, meaning per-turn input re-reads vs output. OpenRouter's study reports a rising tool-call share but no per-category in:out ratio that I could extract. + +### 8.3 Candidate weightings and ranking effects +Define `cost = w_cache·p_cache + w_in·p_in + w_out·p_out` (weights sum to 1), or use measured tokens per task. + +| Weighting | Formula | Who it favours (ASSUMPTION) | +|---|---|---| +| Output-only | `p_out` | Flat-priced or fast-silicon hosts (1×). Penalises 5× vendors. | +| Legacy 3:1 in:out (AA legacy; Arena "3:1") | `(3p_in+p_out)/4` | Cheap-input vendors | +| minibench current 1:3 in:out | `(p_in+3p_out)/4` | Between the two above. Already in `agentbench/board.py` | +| AA 7:2:1 cache:in:out | `(7p_c+2p_in+p_out)/10` | Vendors with deep cache discounts | +| OpenRouter observed ~15:1 in:out | `(15p_in+p_out)/16` | Cheap-input vendors (strongest) | +| **Measured** tokens/task × prices (recommended) | Σ tokens_type × price_type | Reflects verbosity and reasoning. Already how AA and minibench cost per task work | + +- **Worked flip (ASSUMPTION, arithmetic on dated cited prices from different dates, illustration only):** Cerebras-launch Qwen3-Coder ($2/$2) vs Qwen3.7 Max ($1.25/$3.75): + - 3:1 → 2.00 vs 1.875 (Qwen Max cheaper) + - 15:1 → 2.00 vs 1.41 (Qwen Max cheaper) + - 1:3 → 2.00 vs 3.125 (Cerebras cheaper) + - output-only → 2.00 vs 3.75 (Cerebras cheaper) + - **The ranking flips with the weighting.** +- **Recommendation (ASSUMPTION):** make a **workload profile** a first-class, labelled parameter with presets: `chat-3:1`, `aa-7:2:1`, `agentic-measured`, `output-only`. Default the headline to *measured* cost per task. Show rank stability across profiles, e.g. "rank changes under N of 4 profiles". Output-only is a useful stress test but should not be the default, given the token-volume evidence in §8.2. + +--- + +## §9 Addendum 2: reframe as an investigation of what labs actually pay + +**New question (Kyle):** what does each lab actually pay per token or per parameter for inference? Estimate it from disclosed sizes (open weights), API prices, and evidence about hardware and deals. + +**Method sketch (ASSUMPTION):** +1. **Calibrate on open weights served by competitive third-party hosts.** P is known (FACT tier) and markups are competitive. Gundlach et al. assume open models are "not priced at a significant markup relative to the necessary GPU resources" (FACT, arXiv 2511.23455). Epoch reports direct rented-hardware costs within 30% of API prices for 5 open models (FACT, §1.3 URL). Fit `$/1M out` against `active_params` and hardware class to get a **cost curve**. +2. **Anchor dollar-per-hardware-hour from disclosed deals** (§10). Examples: DeepSeek's H800 rental assumption and the Anthropic→xAI lease. +3. **For each closed lab,** compare API price against the calibrated cost curve evaluated at an *external* size band (IKP, Epoch). Never use a size inferred from the same price, which is circular (§2.4). The residual is markup ± hardware advantage ± subsidy. Report it as a **range** (§12). +4. **Throughput cross-check:** tokens/s per user constrains size *given* a hardware class. On Cerebras or Groq-class silicon, the same size runs much faster, so hardware must be identified first (§10). + +--- + +## §10 Compute deals, ownership and silicon: public evidence + +### 10.1 By lab +**OpenAI: mostly rented, diversifying into custom and alternative silicon** +- **FACT:** Cerebras. 750 MW over three years: 250 MW by end-2026, 500 MW by end-2027, 750 MW by end-2028. Cerebras-hosted, "dedicated low-latency inference". Initial value over $10B (Reuters). The SEC-filed master agreement gives OpenAI an option on more capacity on equal or better terms. Later reporting (The Information, via Reuters, 2026-04-17) says over $20B plus warrants for up to about 10% equity. Sources: + - https://www.cerebras.ai/blog/openai-partners-with-cerebras-to-bring-high-speed-inference-to-the-mainstream + - https://www.reuters.com/technology/openai-buy-compute-capacity-startup-cerebras-around-10-billion-wsj-reports-2026-01-14/ + - https://www.sec.gov/Archives/edgar/data/2021728/000162828026029503/exhibit1011-sx1a.htm + - https://www.reuters.com/technology/openai-spend-more-than-20-billion-cerebras-chips-receive-equity-stake-2026-04-17/ + - All accessed 2026-10-02. +- **FACT:** GPT-5.3-Codex-Spark runs on Cerebras at over 1,000 tok/s. It is a research preview for ChatGPT Pro users. API access is limited to "select design partners" — https://openai.com/index/introducing-gpt-5-3-codex-spark/ (accessed 2026-10-02, via search summary). +- **FACT:** Other deals (all via search summaries of first-party pages, accessed 2026-10-02): + - Oracle/Stargate, up to 4.5 GW: https://openai.com/index/five-new-stargate-sites/ + - AMD, 6 GW, first 1 GW of MI450 in late 2026: https://newsroom.amd.com/news/amd-and-openai-announce-strategic-partnership-to-d/ + - Broadcom, 10 GW of *OpenAI-designed* accelerators, late 2026 to 2029: https://openai.com/index/openai-and-broadcom-announce-strategic-collaboration/ + - AWS, $38B over 7 years: https://openai.com/index/aws-and-openai-partnership/ + - Nvidia, at least 10 GW +- **FACT:** Paid priority. OpenAI sells a Scale Tier (committed TPM, minimum 30 days, 99.9% uptime) and "Fast mode" / `service_tier="priority"` at premium per-token prices — https://openai.com/api-scale-tier/ and https://developers.openai.com/api/docs/guides/fast-mode (accessed 2026-10-02, via search summary). Consumer side: ChatGPT **Pro 500** ($500/month, launched 2026-09-29) includes "Astra Ultrafast", up to 8× faster in Codex — https://www.businessinsider.com/chatgpt-new-plan-pro-500-cost-compute-allowance-2026-9 (accessed 2026-10-02). + +**Anthropic: renter on multiple silicon types, no owned fleet found** +- **FACT:** Google Cloud TPUs, "up to one million", tens of billions of dollars, well over 1 GW in 2026 — https://www.anthropic.com/news/expanding-our-use-of-google-cloud-tpus-and-services. AWS Trainium2 via Project Rainier — https://www.aboutamazon.com/news/aws/aws-project-rainier-ai-trainium-chips-compute-cluster. Expanded Google+Broadcom and Amazon (up to 5 GW) deals — https://www.anthropic.com/news/google-broadcom-partnership-compute and https://www.anthropic.com/news/anthropic-amazon-compute (all accessed 2026-10-02; headlines and search summaries only). +- **FACT:** xAI lease. Access to Colossus 1 (over 220k Nvidia GPUs) "to directly improve capacity for Claude Pro and Claude Max" (xAI, 2026-05-06). The SpaceX S-1 discloses $1.25B/month through May 2029 for roughly 300 MW, discounted for the first two months, terminable on 90 days' notice by either side. The S-1 calls it a way to "monetize unused compute capacity" — https://x.ai/news/anthropic-compute-partnership and https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/ (accessed 2026-10-02). +- **ASSUMPTION (arithmetic, illustrative):** $1.25B ÷ about 220k GPUs ≈ $5.7K per GPU-month ≈ $7.8 per GPU-hour, blended across H100/H200/GB200, power and facility included. Caveats: unknown share of the 220k allocated to Anthropic, ramp discount, and Colossus II inclusion per Network World. This is the first public *renter-price* anchor for a frontier lab. +- **FACT:** Anthropic's "Priority Tier" reserved capacity is "no longer available for purchase"; existing commitments run to contract end — https://docs.anthropic.com/en/api/service-tiers (accessed 2026-10-02, via search summary). + +**Google: owns its silicon and data centres (vertically integrated)** +- **FACT:** Gemini is served on Google's own TPUs, including Ironwood, described as "the first Google TPU for the age of inference" — https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/ironwood-tpu-age-of-inference/ (accessed 2026-10-02). Google also rents TPUs to Anthropic (above) and reportedly to Meta (secondary: https://winbuzzer.com/2026/03/03/meta-signs-multibillion-dollar-deal-rent-google-tpus-xcxwbn/, **unverified**). + +**xAI: owns data centres, now also a seller** +- **FACT:** Colossus 1 has over 220k Nvidia GPUs and is xAI-built. xAI sells capacity to Anthropic (above). The S-1 says it "expect[s] to enter into additional similar services contracts" (TechCrunch, above). +- **Reported, not FACT:** TechCrunch writes that Grok usage "has dropped significantly", freeing servers. This is the outlet's characterisation, **unverified**. + +**DeepSeek: renter accounting (self-disclosed)** +- **FACT:** One day (2025-02-27/28): $87,072 GPU rental for an average of 226.75 H800 nodes against $562,027 theoretical revenue at R1 prices, a "545%" theoretical margin. DeepSeek notes actual revenue is lower — §3.5 URL (accessed 2026-10-02). + +**Alternative silicon vendors (context)** +- **FACT:** Nvidia paid about $20B for a *non-exclusive license* to Groq's inference technology and hired executives. Groq remains independent — https://groq.com/newsroom/groq-and-nvidia-enter-non-exclusive-inference-technology-licensing-agreement-to-accelerate-ai-inference-at-global-scale and https://www.reuters.com/business/nvidia-buy-ai-chip-startup-groq-about-20-billion-cnbc-reports-2025-12-24/ (accessed 2026-10-02). +- **Search summary only, unverified:** SambaNova positioning SN50 RDUs for decode. +- **FACT:** Cerebras resells API access through OpenRouter, Hugging Face and AWS Marketplace — https://www.cerebras.ai/pricing (accessed 2026-10-02). +- **ASSUMPTION:** Custom ASICs (Google TPU, AWS Trainium, OpenAI–Broadcom) shift cost away from Nvidia margins. Magnitude **UNKNOWN**: no public per-token cost for any of them. + +### 10.2 Per-lab distortion table (how each bias moves a price-to-size estimate) +Convention: estimating size from price under an *assumed* markup over a *GPU-rental* cost curve. If true cost is lower than the curve (owned or cheaper silicon, subsidy) or the markup is smaller than assumed, the estimate **under**-states size. If the lab pays a renter premium or charges a speed or priority premium, it **over**-states size. The throughput method has its own row. All entries are ASSUMPTIONs built on the §10.1 FACTs. + +| Lab | Ownership / deals (FACT basis) | Price-to-size bias | Throughput-to-size bias | Confidence | +|---|---|---|---|---| +| Google | Owns TPUs and data centres; rents TPUs out | **Under**-estimates (no renter margin; can price near marginal cost; free and promo tiers) | Moderate (TPU speed differs from GPU curve) | Medium | +| xAI | Owns Colossus; sells surplus to Anthropic | **Under**-estimates (owned silicon; incentive to fill idle capacity cheaply) | Low | Medium | +| Anthropic | Rents Trainium, TPU, Nvidia (incl. xAI at ~$1.25B/mo); priority tier closed to new buyers | **Over**-estimates (renter premium stacked on margin); partly offset if Trainium/TPU are cheaper per token than GPUs (UNKNOWN) | Mixed: three silicon types, unknown routing | Low | +| OpenAI | Rents Azure, Oracle, AWS, CoreWeave; Cerebras capacity-as-service; own Broadcom chip from late 2026 | **Over**-estimates on standard tiers (renter + margin); strongly **over** on Fast/priority/Ultrafast tiers (speed premium, not size) | **Under**-estimates on Cerebras-served models (wafer-scale speed looks like a small model) | Low | +| DeepSeek | Rents H800; self-reported theoretical 545% margin | **Over**-estimates if a thin margin is assumed | Low | Medium (self-reported) | +| Third-party open-weight hosts | Competitive renters | ≈ calibration baseline; quantization can bias toward **under** (§3.4) | Hardware-dependent (Groq/Cerebras/SambaNova hosts are fast) | Medium-High | + +--- + +## §11 Kyle's hypotheses: verdicts +| # | Claim | Verdict | Evidence | +|---|---|---|---| +| a1 | OpenAI uses Cerebras | **FACT** | 750 MW deal; Codex-Spark on Cerebras (§10.1) | +| a2 | …but keeps most of that capacity for itself | **UNKNOWN / partially supported** | Only public signal: Codex-Spark is Pro-only, API limited to "select design partners". No disclosure of capacity split. The deal is Cerebras-hosted capacity for "OpenAI customers" (Cerebras blog) | +| a3 | Cerebras just released a $500 plan | **Not supported as stated** | The $500/month plan is **OpenAI's ChatGPT Pro 500** (DevDay 2026-09-29, "Astra Ultrafast"). Pre-launch leaks tied it to Cerebras, but OpenAI pages do not name the hardware (Business Insider; https://pasqualepillitteri.it/en/news/19330/openai-launches-pro-500-500-dollars-month, accessed 2026-10-02). Cerebras Code plans are $50 and $200, shown as sold out — https://www.cerebras.ai/code, https://support.cerebras.net/articles/9996007307-cerebras-code-faq (accessed 2026-10-02). Ultrafast on Cerebras: **UNKNOWN** | +| b1 | xAI owns its own data centres | **FACT** | Colossus 1, "built from the ground up" (xAI), Memphis | +| b2 | xAI sells compute to others, e.g. Anthropic | **FACT** | xAI 2026-05-06 announcement; S-1 terms $1.25B/month to May 2029 (TechCrunch) | +| c1 | Owned or reserved silicon and priority deals mean API price ≠ lab cost | **Supported in direction** | Priority/Scale tiers price capacity, not size; DeepSeek theoretical margin; xAI monetising "unused" capacity. Magnitude per lab **UNKNOWN** | +| c2 | Owners price near marginal cost; renters carry a premium | **ASSUMPTION (plausible, unproven)** | No lab discloses marginal cost. Gundlach et al. attribute part of the closed-vs-open price gap to competition and markup, not ownership. Owners may also price *above* cost to recover capex | + +--- + +## §12 Hardware ownership as a methodology risk, and open decisions for Kyle + +**Risk statement (ASSUMPTION):** Price-to-size and size-to-cost inversions assume one shared cost curve. Ownership (Google, xAI), alternative silicon (Cerebras, Groq, TPU, Trainium, custom ASICs), lab-to-lab leases (Anthropic→xAI), and priced priority tiers (OpenAI Fast/Scale, Pro 500) each move a lab off that curve, by unknown amounts and in different directions (§10.2). On top of the about 2–3× size uncertainty (§2.3), any "true cost" point estimate is unsupported. + +**Open decisions:** +1. **Present "estimated true cost" as a range?** Recommendation (ASSUMPTION): yes, always. Show low/high bands that combine the size band (IKP/Epoch) × hardware-class cost range × markup range. Never show a point. Label the tier (`disclosed` / `estimated:`). +2. **Back out hardware ownership as a variable?** Options: + - (i) Treat it as a **categorical covariate** (owned / rented-GPU / rented-alt-silicon / mixed) from §10 evidence and report results stratified by category. Lowest risk. + - (ii) Fit it as a **latent parameter**. It is not identifiable from price alone: markup and hardware advantage are confounded. + - (iii) Leave it out and publish the residual as "unexplained price gap". + - Recommendation (ASSUMPTION): (i) plus (iii). Avoid (ii) unless an independent per-token cost anchor appears, such as a disclosed lease rate × measured throughput. +3. **Which price is "the" price per lab?** First-party standard tier, cheapest third-party host, or batch. This affects every row of §10.2. +4. **Do priority/fast tiers count as separate data points?** They are useful because they isolate the speed premium. +5. **Default workload profile** (§8.3): measured, 7:2:1, or output-only stress test. +6. **Home:** under the reframe, this is research rather than a leaderboard. That strengthens the §4.5 verdict: run the investigation in a separate research repo or notebook, and republish only cited, tiered outputs into minibench (Benchmark Lens / Usage Board). + +**Sparks re-ranked under the reframe (ASSUMPTION):** +1. Open-Weight Provider Markup Index (was #2). It now becomes the **calibration layer** for the cost curve and is valid on FACT-tier sizes. +2. Lab True-Cost Range Report (evolves #3). The cost curve plus external size bands plus the §10.2 ownership strata give per-lab *ranges* and an "unexplained gap", published as research. +3. Value Frontier with Size Lens in minibench (was #1). Still the easiest build, but now a by-product rather than the goal. + +--- + +## §7 Source index (all accessed 2026-10-02; placed last, after the addenda §8–§12) +- AA v4.3: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3 +- AA v4.1: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1 +- AA comparisons: https://artificialanalysis.ai/models/comparisons · open-source: https://artificialanalysis.ai/models/open-source +- Arena Pareto: https://www.linkedin.com/posts/arenaai_weve-added-pareto-frontier-charts-to-the-activity-7445144386054828033-jYa1 · https://arena.ai/blog/agent-categories-and-cost +- Epoch price of thought: https://epoch.ai/publications/the-plunging-price-of-thought +- Epoch smaller models: https://epoch.ai/gradient-updates/frontier-language-models-have-become-much-smaller +- Epoch data: https://epoch.ai/data/ai-models · https://epoch.ai/models/claude-3-5-sonnet +- Gundlach et al.: https://arxiv.org/abs/2511.23455 +- IKP: https://arxiv.org/abs/2604.24827 · NightVision: https://arxiv.org/abs/2607.01313 +- Carlini et al.: https://proceedings.mlr.press/v235/carlini24a.html +- MEDEC: https://arxiv.org/abs/2412.19260 · GPT-4 TR: https://arxiv.org/abs/2303.08774 +- FMTI 2025: https://crfm.stanford.edu/fmti/December-2025/index.html +- DeepSeek-V3: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/README.md · DeepSeek margin: https://github.com/deepseek-ai/open-infra-index +- gpt-oss: https://arxiv.org/html/2508.10925 · Grok-1: https://x.ai/news/grok-os +- HF OLL: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard/discussions/1135 +- OpenRouter models API: https://openrouter.ai/docs/api/api-reference/models/list-all-models-and-their-properties +- the-frontier.app: https://the-frontier.app/ +- minibench: https://github.com/RaapTechllc/minibench (main @ f9e2b31; PRs #68, #69, #70) +- Addenda: AA methodology https://artificialanalysis.ai/methodology · AA Data API https://artificialanalysis.ai/data-api/docs · OpenRouter State of AI https://arxiv.org/abs/2601.10088 +- Cerebras/OpenAI: https://www.cerebras.ai/blog/openai-partners-with-cerebras-to-bring-high-speed-inference-to-the-mainstream · SEC exhibit https://www.sec.gov/Archives/edgar/data/2021728/000162828026029503/exhibit1011-sx1a.htm · Reuters 2026-01-14 & 2026-04-17 · https://openai.com/index/introducing-gpt-5-3-codex-spark/ +- Cerebras Code: https://www.cerebras.ai/code · https://support.cerebras.net/articles/9996007307-cerebras-code-faq · https://www.cerebras.ai/pricing +- OpenAI Pro 500: https://www.businessinsider.com/chatgpt-new-plan-pro-500-cost-compute-allowance-2026-9 +- xAI/Anthropic: https://x.ai/news/anthropic-compute-partnership · https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/ +- Anthropic compute: https://www.anthropic.com/news/expanding-our-use-of-google-cloud-tpus-and-services · https://www.aboutamazon.com/news/aws/aws-project-rainier-ai-trainium-chips-compute-cluster · https://docs.anthropic.com/en/api/service-tiers +- OpenAI compute: https://openai.com/index/five-new-stargate-sites/ · https://newsroom.amd.com/news/amd-and-openai-announce-strategic-partnership-to-d/ · https://openai.com/index/openai-and-broadcom-announce-strategic-collaboration/ · https://openai.com/index/aws-and-openai-partnership/ · https://openai.com/api-scale-tier/ +- Google TPU: https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/ironwood-tpu-age-of-inference/ · Groq/Nvidia: https://groq.com/newsroom/groq-and-nvidia-enter-non-exclusive-inference-technology-licensing-agreement-to-accelerate-ai-inference-at-global-scale From 2d4e7bcc19dc6b5f1886d7deb06cabcd68d3da42 Mon Sep 17 00:00:00 2001 From: Cursor Agent Date: Fri, 2 Oct 2026 15:31:39 +0000 Subject: [PATCH 2/2] docs: add PRD for LLM true-compute-cost investigation (spec only) Co-authored-by: RaapTechllc --- docs/PRD-llm-compute-cost.md | 366 +++++++++++++++++++++++++++++++++++ 1 file changed, 366 insertions(+) create mode 100644 docs/PRD-llm-compute-cost.md diff --git a/docs/PRD-llm-compute-cost.md b/docs/PRD-llm-compute-cost.md new file mode 100644 index 0000000..971808d --- /dev/null +++ b/docs/PRD-llm-compute-cost.md @@ -0,0 +1,366 @@ +# PRD: LLM true-compute-cost investigation (spec only) + +- **Status:** PROPOSED. Awaiting acceptance by Kyle Raap. No build until §11 GO gate clears. +- **Date:** 2026-10-02. +- **Idea owner:** Kyle Raap (RaapTech). Proposed home: `RaapTechllc/minibench` (notes header). +- **Single source of market facts:** [`docs/research/llm-compute-cost-research-notes-2026-10-02.md`](research/llm-compute-cost-research-notes-2026-10-02.md), frozen 2026-10-02. Cited below as `notes §n`. No price, parameter count, deal figure, or statistic in this document comes from anywhere else. +- **Repo baseline:** `main` at `f9e2b31`, the same commit the notes inspected (notes §4). Open PRs #68, #69, #70 are not touched by this document (notes §4.3). +- **Glossary and decision records:** `CONTEXT.md`; ADR 0002 (Usage Board, Mode A), ADR 0003 (board path), ADR 0004 (Benchmark Lens, proposed). Where this PRD would require an ADR change it says so explicitly, per `docs/agents/domain.md`. + +## 0. How to read the labels + +| Label | Meaning | +|---|---| +| **FACT** | Carried over from the notes with their URL and access date. Cited as `notes §n`. | +| **ASSUMPTION** | The notes' own reasoning, not sourced. Carried over unchanged. | +| **UNKNOWN** | Not found or not verified in the notes. Never upgraded here. | +| **PROPOSED** | A design choice introduced by this PRD. Not in the notes. Kyle may reject it. | + +The project has been reframed (notes header, addendum 2). It started as a size-over-cost ratio (parameter count ÷ per-million-token price) on a Pareto chart. It is now an *investigation of what labs actually pay for inference compute*. The ratio and the Pareto chart are estimation tools, not the product (notes §9–§12). + +## 1. Problem + +1. **The question.** What does each lab actually pay per token or per parameter for inference? API price is not lab cost, and the gap differs by lab (notes §0.7, §9). The direction of this claim is supported; the magnitude per lab is **UNKNOWN** (notes §11 c1). +2. **Nobody shows the pieces together.** Capability-per-dollar Pareto charts already exist at Artificial Analysis, LMArena, Epoch AI and MIT FutureTech (FACT, notes §0.1). Dividing parameters by price is uncommon and no public board doing it was found (**UNKNOWN** whether one exists, notes §0.1, §6). Charts that mix open and closed models either drop size altogether or restrict themselves to open weights (ASSUMPTION, notes §1 gap summary). +3. **Closed labs withhold size.** OpenAI, Google, xAI, Amazon and Mistral score zero on FMTI's model-information indicators, which include model size; Anthropic does not publish counts either (FACT, notes §0.3, §2.2). Published hidden-size estimators are wide: Epoch's inference-economics estimates are "off by a factor of 2", IKP's 90% prediction interval is about 3.2×, NightVision's parameter error is about 53% (FACT, notes §0.3, §2.3). +4. **The raw ratio is a poor capability proxy.** `V = P / price` swings about 20× on the MoE total-vs-active definitional choice alone, and is broken by distillation, quantization, provider margins, reasoning-token volume, price decline, caching, blend ratio and multi-provider pricing (notes §3, table rows 3.1–3.11; ASSUMPTION, supported by FACTs in §3). The ratio is defensible on one population only: open-weight models with known `P`, across hosting providers, where it measures provider markup and serving efficiency rather than capability (ASSUMPTION, notes §3 item 5). +5. **Hardware ownership breaks a shared cost curve.** Google owns TPUs; xAI owns Colossus and leases it to Anthropic for $1.25B/month; Anthropic rents Trainium, TPUs and Colossus; OpenAI rents across Azure, Oracle, AWS and Cerebras with its own Broadcom chip arriving from late 2026 (FACT, notes §0.7, §10.1). Any price-to-size or size-to-cost inversion that assumes one cost curve is therefore biased per lab, in different directions and by unknown amounts (ASSUMPTION, notes §10.2, §12). +6. **minibench is further along than the brief said.** The brief claimed minibench had "no cost or Pareto work yet". That is wrong: `_pareto_frontier` in `backend/app/agents_router.py`, the "Capability vs. cost" chart in `frontend/src/pages/Models.tsx`, and a 1:3 in:out blended price already exist. What minibench lacks is parameter counts (FACT, repo inspection, notes §0.4, §4.2). +7. **Output weighting.** Kyle wants cost weighted toward output tokens. By token count agentic traffic is input-heavy (OpenRouter averages about 6K prompt vs 400 completion tokens per request; code requests "routinely exceed 20K input tokens"), yet by dollars output and reasoning can still dominate (published output prices run 1–5× input; in AA's example payload reasoning cost exceeds input cost) (FACT, notes §0.6, §8.2). The weighting must therefore be a configurable workload profile, with output-only as a stress test, not the default (ASSUMPTION, notes §8.3). + +## 2. Buyer and user: agent vs human + +### 2.1 Buyer + +Kyle Raap, RaapTech (notes header). Kyle decides GO (§11), resolves the open decisions (§10), and authorizes any publication. The MVP spends nothing on model calls (§5, `AGENTS.md` hard invariants). + +### 2.2 Users, by deliverable + +| Deliverable (§4) | Primary consumer | Secondary consumer | Is the consumer an agent? | +|---|---|---|---| +| D1 Open-Weight Provider Markup Index | **Vibe coder / normal dev** (`CONTEXT.md` audience) asking "who hosts model X cheapest per unit of compute" (ASSUMPTION, notes §3 item 5, §5 #2) | **Agents** picking a host for a known open model. The Usage Board already exposes an MCP `recommend` tool over a cached board (`docs/PRD-openrouter-usage-board.md`, story 6); D1 output is shaped to be consumable the same way (PROPOSED) | Both. Machine-readable JSON is a first-class output. | +| D2 Lab True-Cost Range Report | **Humans**: Kyle, researchers, readers of a RaapTech research repo, notebook or blog post (ASSUMPTION, notes §4.5, §12 d6) | None by design. Ranges plus ownership strata need interpretation; the report must not be fed to an agent as a cost input (PROPOSED) | No. Human-read research. | +| D3 Value Frontier with Size Lens | **Vibe coder** on `/models` (notes §4.4 Surface A) | **Agents** reading `/api/v1/agents/models/leaderboard`, which already returns `on_pareto_frontier` (FACT, notes §4.2) | Both. Size is an annotation on an existing agent-readable payload. | + +### 2.3 Operator + +PROPOSED: the computation pipeline (ledger validation, cost-curve fit, range generation, report rendering) is designed to be run end-to-end by a coding agent under the repo's hard invariants, with Kyle as the human gate on publication. Ledger *curation* (entering a parameter count, a price, a deal figure) follows the Benchmark Lens rule: never type a number without the URL and date it came from, and leave untraceable values `null` (`docs/PLAN-benchmark-lens.md`, guardrails). An agent may draft ledger rows; a human verifies URL and date before a row leaves fixture status. + +## 3. Agent-First score intent + +The notes contain no Agent-First rubric and the repo has none. The six dimensions below are **PROPOSED**. Intent is a 1–5 target per dimension, scored per deliverable, with the evidence that would justify the score. + +| # | Dimension (PROPOSED) | D1 Markup Index | D2 Range Report | D3 Value Frontier | Evidence needed to claim the score | +|---|---|---|---|---|---| +| A | **Machine-readable output.** Every artifact has a JSON form with a published schema. | 5 | 3 | 5 | Schema file in `agentbench/data/`; CLI `--out` writes JSON; D2's markdown report is generated from the same JSON (hence 3, not 5, because the narrative is the deliverable). | +| B | **Provenance per number.** Every number carries `source_url`, `as_of`, and a tier. | 5 | 5 | 5 | Loader rejects rows missing URL or `as_of` (ADR 0004 item 1 pattern); tier badge visible wherever a number renders. | +| C | **Offline reproducibility.** Rebuildable from committed fixtures with no key and no network. | 5 | 4 | 5 | CI job without `OPENROUTER_API_KEY` is green; D2 is 4 because the cost-curve fit depends on curated ledgers that a human verified. | +| D | **Agent operability.** An agent can run the pipeline and act on exit codes. | 4 | 3 | 4 | `python -m agentbench. --check` exits non-zero on loader violations; dry-run mode; no interactive prompts. | +| E | **Consumable without human interpretation.** | 4 | 2 | 3 | D1 is a ranked table with one meaning (provider markup). D2 is intervals plus strata and an unexplained gap; it is written for a human reader (notes §12 d6). D3 adds a size channel that still needs the tier legend. | +| F | **Safety and gating.** No paid calls, no publishing, no external communication without owner authorization. | 5 | 5 | 5 | Allowlist tests on the OpenRouter client; `live=false` on fixtures; GO gate (§11) and `AGENTS.md` hard invariants. | + +## 4. MVP scope + +Scope follows the notes' re-ranked sparks under the reframe (ASSUMPTION, notes §12 "Sparks re-ranked"): the markup index is the baseline and calibration layer, the range report is the investigation, and the minibench view is a by-product. + +### 4.1 D1 — Open-Weight Provider Markup Index (baseline; calibration layer) + +- For each open-weight model with a disclosed parameter count, compute `$ per 1M tokens ÷ active_params` for each hosting-provider endpoint, rank providers, and track over time (notes §5 #2). +- Why first: it is Kyle's ratio on the one population where it is valid, because `P` is a FACT, and it answers a real buyer question: who is overcharging for the same weights? (ASSUMPTION, notes §5 #2.) Under the reframe it becomes the calibration layer for the cost curve (notes §12). +- Inputs: a **parameter ledger** (open weights, `disclosed` tier) and **per-provider endpoint prices**. The endpoint data path is constrained; see §7 and open decision D7. +- Known confounder: quantization differences between providers (notes §3.4, §5 #2). Whether OpenRouter endpoint data records quantization is **UNKNOWN** (notes §6). +- Fit in minibench: Usage Board "best-by-$-per-active-param" compare route for open weights only (ASSUMPTION, notes §4.4 Surface B). + +### 4.2 D2 — Lab True-Cost Range Report (the investigation) + +- The cost curve from D1, plus external size bands, plus the ownership strata of notes §10.2, give per-lab cost *ranges* and an "unexplained gap", published as research (notes §12 spark 2). +- Method follows notes §9 (ASSUMPTION): calibrate on open weights served by competitive hosts; anchor dollar-per-hardware-hour from disclosed deals; for each closed lab compare API price against the calibrated curve evaluated at an *external* size band, never one inferred from the same price; cross-check with throughput once hardware is identified. +- Output is always a range, never a point (notes §12 d1, recommendation ASSUMPTION; confirmed as open decision D1). +- Home: the notes' verdict is that this is research, not a leaderboard, and belongs in a separate research repo or notebook with only cited, tiered outputs republished into minibench via the Benchmark Lens or Usage Board (notes §4.5, §12 d6). This is open decision D6, not a settled fact. + +### 4.3 D3 — Value Frontier with Size Lens (minibench; by-product) + +- On `/models` "Capability vs. cost", encode marker size or opacity by `active_params` when the tier is `disclosed`; closed or unknown models get a hollow marker. Axes and Pareto logic stay as they are (ASSUMPTION, notes §4.4 Surface A, §5 #1). +- Add toggles for workload profile and cache assumption (notes §5 #1). +- Size estimates for closed models, if ever shown, are republished through the Benchmark Lens as an `estimated:` tier with citation, never on an axis (notes §3 item 4, §4.4 Surface C). +- Sequencing: wait for PRs #68 and #69 to land or close before touching the frontend or backend models (ASSUMPTION, notes §4.4). PR #69 removes legacy `/api/v1/models`; a parameter catalog would extend `KnownModel` and the seed JSON instead (notes §4.3). + +### 4.4 Deliverable artifacts (PROPOSED) + +| Artifact | Form | Lives in | +|---|---|---| +| Parameter ledger | JSON, one row per open-weight model: `model_id`, `total_params_b`, `active_params_b`, `count_convention`, `params_source_url`, `params_tier`, `as_of` (fields per notes §2.1 ASSUMPTION and §4.4) | `agentbench/data/` if D6 picks minibench; otherwise the research repo, mirrored into minibench as a fixture | +| Endpoint price ledger | JSON, one row per (model, provider, service tier): prices by token type, `quantization` or `unknown`, `price_source`, `as_of` | Same | +| Size-band ledger (closed models) | JSON, one row per model: `band_low_b`, `band_high_b`, `size_band_method` (`ikp`, `epoch`), `source_url`, `as_of` | Same; republished via Benchmark Lens tier `estimated:` | +| Deal / ownership ledger | JSON: lab, stratum, deal figure, unit, `source_kind`, `source_url`, `as_of`; derived per-hour arithmetic in a separate ASSUMPTION-labelled field | Same | +| Markup index output | JSON plus table; CLI `--out` | D1 | +| Cost curve and ranges | JSON plus a generated markdown or notebook report | D2 | +| Size-lens rendering | Marker encoding on the existing `/models` scatter | D3 | + +## 5. Non-goals + +1. **`params ÷ price` as a headline metric or chart axis** for any population. It rewards buying the most weights per dollar, not the most work done (ASSUMPTION, notes §3.11), and conflicts with the repo's no-composite and provenance rules (notes §4.5). +2. **Inferring a closed model's size from its price, then dividing by price.** Circular (ASSUMPTION, notes §2.4). The "second pass" in the original brief is dropped. +3. **Point estimates of any lab's true cost, margin, or profit.** Ranges only (notes §12 d1). The word "margin" appears only when quoting DeepSeek's self-reported theoretical figure (FACT, notes §3.5, §10.1). +4. **Fitting hardware ownership as a latent parameter.** Not identifiable from price alone; markup and hardware advantage are confounded (notes §12 d2 ii). +5. **Active size-extraction attacks** (logit-bias, softmax-bottleneck, timing side channels) against any provider (methods in notes §2.3). This PRD only republishes third-party published bands. +6. **Any composite score.** Extends the existing no-composite rule (ADR 0004 item 5; `CONTEXT.md`). +7. **Live model calls, `/chat/completions`, `/analytics`, scraping, Mode B, new packages, new secrets in the repo, paid campaigns** (ADR 0002, including its "adding a package is a STOP" rule; `CONTEXT.md` Mode A/Mode B; `AGENTS.md` hard invariants). +8. **Replacing or re-scoring the Solo, Multiplayer or Agent Cabinet boards.** D3 annotates an existing chart; nothing from D1 or D2 joins a measured cabinet score (notes §4.3 PR #70 rule; ADR 0004 item 7). +9. **Current price quotes.** The notes' prices are cited examples tied to their source date, not current quotes (notes header). This PRD asserts none. +10. **Resolving the hypotheses the notes leave UNKNOWN** (notes §11 a2, c2; §6). They stay at their verdicts (Appendix A). + +## 6. Methodology + +### 6.1 Metric definitions + +| ID | Metric | Definition | Population | What it measures | Source | +|---|---|---|---|---|---| +| M0 | Raw ratio `V` | `V = P / price`, one value per model | Any | Rejected as a capability proxy (§5 item 1) | notes §3 | +| M1 | **Provider markup index** | For open-weight model `m` with disclosed `active_params_b(m)` and provider endpoint `e`: `markup_index(m, e, profile) = blended_price(e, profile) [$ per 1M tokens] ÷ active_params_b(m)` | Open weights, `params_tier = disclosed`, per provider endpoint | Provider markup and serving efficiency for the same weights. Not capability. | ASSUMPTION, notes §3 item 5, §5 #2 | +| M2 | **Cost curve** | Fit `log(blended_price(profile))` against `log(active_params_b)` and a hardware-class covariate, on open-weight models served by competitive third-party hosts. Report fit statistics and the calibration table. | Open weights only | The competitive rental cost of serving a given active size on a given hardware class | ASSUMPTION, notes §9 step 1. Calibration justification: Gundlach et al. assume open models are "not priced at a significant markup relative to the necessary GPU resources"; Epoch reports direct rented-hardware costs within 30% of API prices for 5 open models (FACT, notes §9) | +| M3 | **Hardware-hour anchors** | Dollar per accelerator-hour derived from disclosed deals. Deal figures are FACT; the per-hour arithmetic is ASSUMPTION and stored in a separate field. | Labs with disclosed rental or lease terms | A renter-price anchor for the cost curve | notes §9 step 2, §10.1. The notes' two anchors: DeepSeek's $87,072 one-day GPU rental for an average of 226.75 H800 nodes (FACT); Anthropic→xAI $1.25B/month for roughly 300 MW and over 220k GPUs (FACT), illustratively ≈ $5.7K per GPU-month ≈ $7.8 per GPU-hour (ASSUMPTION, with the notes' caveats) | +| M4 | **Implied cost range** | For closed model `m` of lab `L`: evaluate M2 at `[band_low, band_high]` from an **external** size band, over the hardware classes in `L`'s stratum. `implied_cost(m) = [low, high]`. | Closed models with an external band | What the calibrated competitive curve says serving `m` would cost | ASSUMPTION, notes §9 step 3, §12 d1 | +| M5 | **Unexplained gap** | `unexplained_gap(m) = api_price(m, price_basis) − implied_cost(m)`, as an interval. Sign and expected direction annotated from the §10.2 table. | Same as M4 | Markup ± hardware advantage ± subsidy, which this method cannot separate | ASSUMPTION, notes §9 step 3, §12 d2 iii | +| M6 | **Throughput cross-check** | Tokens/s per user constrains size *given* a hardware class. Applied only after hardware is identified. | Any model with published throughput | A consistency check on the size band, not an estimator | ASSUMPTION, notes §9 step 4; bias rows in §10.2 | +| M7 | **Value frontier** | y = cabinet pass rate; x = measured $ per task; Pareto via existing `_pareto_frontier`. Marker = `active_params` for `disclosed`; hollow or banded otherwise. Size never on an axis. | minibench models with cabinet results | Score per dollar with size as a secondary channel | FACT for existing code (notes §4.2); ASSUMPTION for the size channel (notes §3 item 4, §4.4) | +| M8 | **Rank stability** | For any ranked output, the count of profiles under which the top-k order changes: "rank changes under N of K profiles" | Any ranking | How much the ranking is an artifact of the weighting | ASSUMPTION, notes §8.3 | + +Which `P`: the default divisor is `active_params`. `total_params` and `count_convention` are stored alongside, because conventions differ between vendors (embeddings in or out, MTP or vision towers in or out) and the MoE total:active ratio alone is about 18× for DeepSeek-V3 (671B/37B) and about 23× for gpt-oss-120b (117B/5.1B) (FACT counts, notes §2.1; ratios are the notes' arithmetic, §3.1). Mapping a *total*-parameter external band (IKP predicts MoE knowledge better from total than active, R² 0.67 vs 0.41, FACT notes §2.3) onto an *active*-parameter cost curve is unresolved; see open decision D8. + +Every stored value carries `as_of`, `profile_id` (blend ratio), `price_source`, `cache_assumption`, and `provider` (ASSUMPTION, notes §3 item 6). + +### 6.2 Workload profiles + +Cost under a profile is `cost = w_cache·p_cache + w_in·p_in + w_out·p_out` with weights summing to 1, or measured tokens per task times prices (notes §8.3). + +| Profile id | Formula | Origin | Who it favours (ASSUMPTION, notes §8.3) | Role in this PRD | +|---|---|---|---|---| +| `agentic-measured` | Σ tokens_type × price_type from actual runs | AA and minibench cost per task already work this way (FACT, notes §8.3) | Reflects verbosity and reasoning | **Default headline** where measured tokens exist (notes §8.3 recommendation) | +| `aa-7:2:1` | `(7p_cache + 2p_in + p_out)/10` | AA methodology (FACT, notes §8.2) | Vendors with deep cache discounts | Preset | +| `chat-3:1` | `(3p_in + p_out)/4` | AA legacy; Arena "3:1" (FACT, notes §1.2, §8.3). Whether Arena's 3:1 is input:output is **UNKNOWN** (notes §1.2, §6) | Cheap-input vendors | Preset | +| `output-only` | `p_out` | Kyle's premise | Flat-priced or fast-silicon hosts; penalises 5× vendors | **Stress test only**, never default (notes §8.3) | +| `minibench-1:3` | `(p_in + 3p_out)/4` | Existing `agentbench/board.py::blended_per_million` (FACT, notes §4.2) | Between 3:1 and output-only | PROPOSED additional preset, so the existing Usage Board blend is reproducible | +| `openrouter-15:1` | `(15p_in + p_out)/16` | OpenRouter observed ~6K:400 tokens (FACT, notes §8.2; 15:1 is the notes' arithmetic, ASSUMPTION) | Cheap-input vendors (strongest) | PROPOSED additional preset, as the token-volume reality check | + +Rules: +- The profile is a first-class, labelled parameter (notes §8.3). Every output states its `profile_id`. +- Where no measured tokens exist (D1 has no runs), the default profile is open decision D5. +- Rankings are shown with M8 rank stability. The notes' worked flip is the fixture test: Cerebras-launch Qwen3-Coder ($2/$2) vs Qwen3.7 Max ($1.25/$3.75) gives 2.00 vs 1.875 under 3:1, 2.00 vs 1.41 under 15:1, 2.00 vs 3.125 under 1:3, 2.00 vs 3.75 under output-only; the ranking flips with the weighting (ASSUMPTION, arithmetic on dated prices from different dates, illustration only, notes §8.3). +- Why output-only is not the default: token mix is input-heavy (about 15:1 by tokens, ASSUMPTION on FACT, notes §8.2), yet reasoning can be about 95% of output tokens and output about 56% of cost in AA's example payload, whose cost breakdown is **unverified** (notes §8.2). Effort setting alone moves output volume about 3.7× for one model (FACT, notes §8.2). A per-category in:out ratio for agentic tool-call loops is **UNKNOWN** (notes §6, §8.2). + +### 6.3 How ranges are produced + +PROPOSED procedure, implementing notes §9 and §12 d1: + +1. **Calibration set.** Open-weight models with `params_tier = disclosed`, served by competitive third-party hosts. Each row: model, provider, hardware class, prices by token type, `quantization` or `unknown`, `as_of`. No closed-model row is ever in the calibration set. +2. **Fit M2** per hardware class. Report fit statistics and the residual spread of the calibration set. The residual spread is the first component of the range. +3. **Size band.** Take `[band_low, band_high]` from an external source (`ikp`, `epoch`) with URL and `as_of`. Example of the width to expect: IKP reports "GPT-5.5 ~4.7T [1.5T–15.1T]" (FACT, notes §2.3); Epoch says its GPT-4o (~200B) and Claude 3.5 Sonnet (~400B) estimates could be "off by a factor of 2" (FACT, notes §2.3); MEDEC's ~175B for Sonnet contradicts Epoch's ~400B and is described by its authors as unverified (FACT, notes §2.3). Where no external band exists the row is `unknown` and produces no range. +4. **Hardware-class range.** From the lab's ownership stratum (§6.4), the set of hardware classes the lab is known to serve on. Per-token costs of TPU, Trainium, Cerebras or Broadcom silicon are **UNKNOWN** (notes §6), so alternative-silicon strata widen the range rather than shift it. +5. **Interval arithmetic.** `implied_cost = [min, max]` over band endpoints × hardware-class endpoints × calibration residual. Endpoints, not distributions: the inputs are bands with stated tiers, not calibrated distributions (PROPOSED). +6. **Gap.** `unexplained_gap = api_price − implied_cost`, as an interval, with the §10.2 expected bias direction annotated. Never labelled margin or profit. +7. **Throughput check (M6).** If published tokens/s is inconsistent with the band on the identified hardware class, flag the row; do not adjust the band. +8. **Report.** One row per (lab, model, price basis, service tier, profile), stratified by ownership stratum, with every input's tier visible. A point value is never rendered. + +The notes are explicit that at 2–3× size error a closed model's ratio can move about 10× end to end (ASSUMPTION, notes §2.4) and that on top of that uncertainty "any 'true cost' point estimate is unsupported" (ASSUMPTION, notes §12). The report's copy must say so. + +### 6.4 Per-lab distortion handling + +Convention (notes §10.2): estimating size from price under an assumed markup over a GPU-rental cost curve. If true cost is below the curve (owned or cheaper silicon, subsidy) or the markup is smaller than assumed, the estimate **under**-states size. If the lab pays a renter premium or charges a speed or priority premium, it **over**-states size. All entries are ASSUMPTIONs built on the §10.1 FACTs. + +| Lab | Ownership / deals (FACT basis, notes §10.1) | Stratum (PROPOSED enum, from notes §12 d2 i) | Price-to-size bias | Throughput-to-size bias | Confidence (notes) | +|---|---|---|---|---|---| +| Google | Owns TPUs and data centres; rents TPUs out | `owned` | **Under** | Moderate | Medium | +| xAI | Owns Colossus; sells surplus to Anthropic | `owned` | **Under** | Low | Medium | +| Anthropic | Rents Trainium, TPU, Nvidia (incl. xAI at ~$1.25B/mo); priority tier closed to new buyers | `mixed` | **Over**; partly offset if Trainium/TPU are cheaper per token than GPUs (UNKNOWN) | Mixed: three silicon types, unknown routing | Low | +| OpenAI | Rents Azure, Oracle, AWS, CoreWeave; Cerebras capacity-as-service; own Broadcom chip from late 2026 | `mixed` | **Over** on standard tiers; strongly **over** on Fast/priority/Ultrafast tiers | **Under** on Cerebras-served models | Low | +| DeepSeek | Rents H800; self-reported theoretical 545% margin | `rented-gpu` | **Over** if a thin margin is assumed | Low | Medium (self-reported) | +| Third-party open-weight hosts | Competitive renters | `rented-gpu` or `rented-alt-silicon` | ≈ calibration baseline; quantization can bias toward **under** | Hardware-dependent (Groq/Cerebras/SambaNova hosts are fast) | Medium-High | + +Handling (notes §12 d2 recommendation, ASSUMPTION; pending open decision D2): +- Option (i): treat ownership as a **categorical covariate** (`owned` / `rented-gpu` / `rented-alt-silicon` / `mixed`) and report results stratified by category. Lowest risk. +- Option (iii): publish the residual as an **unexplained price gap** (M5). +- Option (ii), a latent parameter, is avoided unless an independent per-token cost anchor appears, such as a disclosed lease rate × measured throughput. +- Priority, fast and Scale tiers are kept as separate rows because they isolate the speed premium (notes §12 d4; open decision D4). Examples in the notes: OpenAI Scale Tier and `service_tier="priority"`, ChatGPT Pro 500 with "Astra Ultrafast"; Anthropic's Priority Tier "no longer available for purchase" (FACT, notes §10.1). + +## 7. Data sources and refresh + +| Source | What it supplies | Tier | Access path and constraint | Refresh | +|---|---|---|---|---| +| Open-weight model cards and papers: DeepSeek-V3 README, gpt-oss paper, Grok-1 release, NightVision Appendix C | `total_params`, `active_params`, `count_convention` | FACT (notes §2.1) | Hand-curated ledger with URL and access date (ADR 0004 item 6 pattern: curated ledger, no scraping) | On each new open-weight release; PROPOSED review at least monthly | +| OpenRouter `/models` | Per-token prices (prompt, completion, cache, reasoning), `context_length`, sort by price, throughput, latency, AA index. **No parameter field** | FACT (notes §1.6) | Already allowlisted Mode A GET path, polled daily (ADR 0002; `CONTEXT.md`) | Existing daily poll; fixture when no key | +| OpenRouter `/models/{author}/{slug}/endpoints` | Per-provider price, latency, throughput | FACT that it exists (notes §3.10) | **Not** in the Mode A allowlist. Would need an ADR change; **UNKNOWN** whether permitted (notes §5 #2, §6). Whether it records quantization is **UNKNOWN** (notes §6). See D7 | Only if D7 allows; otherwise a curated endpoint ledger | +| minibench `backend/app/data/known_models_seed.json` | 42 OpenRouter models: `provider, model_id, display_name, family, license, context_length, prompt_price, completion_price, snapshot_date`. No parameter fields. Generated 2026-07-06 | FACT (notes §4.2) | Repo file | Already stale relative to price decline (§9.3); regenerate from the daily poll (PROPOSED) | +| IKP (arXiv 2604.24827) | Banded size estimates for 97 proprietary models; calibrated on 93 open models | FACT for the paper's claims (notes §2.3). Preprint, not independently replicated (ASSUMPTION, notes §2.4). Code license **UNKNOWN**; probe-release statement internally inconsistent (notes §6) | Curated size-band ledger citing the paper; tier `estimated:ikp` | On paper revision; see D9 | +| Epoch AI data and gradient updates | Parameter and compute database with confidence tiers ("Confident" within 3×, "Likely" within 10×, "Speculative" within 30×); size guesses for GPT-4o and Claude 3.5 Sonnet | FACT (notes §1.3, §2.3) | Curated ledger; tier `estimated:epoch`, carrying Epoch's own confidence tier | On Epoch update; see D9 | +| MEDEC (arXiv 2412.19260) | Press-mined sizes, described as unverified | FACT that the paper lists them; values unverified (notes §2.3) | PROPOSED: excluded from ranges; may be shown as `imported` for the disagreement note only | n/a | +| Deal and ownership evidence (notes §10.1 URLs: Cerebras/OpenAI blog, SEC exhibit, Reuters, xAI announcement, TechCrunch S-1 reading, Anthropic TPU/Trainium posts, Google Ironwood, DeepSeek inference overview) | Ownership stratum, lease and rental figures, capacity commitments | Mixed: FACT where first-party; many figures via search summaries or secondary press, S-1 not read directly (notes §6) | Curated deal ledger with `source_kind` per row (ADR 0004 enum) | On new filings or announcements; PROPOSED review monthly | +| AA methodology and Data API docs | 7:2:1 blend definition; example payload token and cost split (cost breakdown **unverified**) | FACT / unverified (notes §8.2) | Profile definitions only; no AA data fetched | On AA methodology change | +| Third-party mirror of AA data (aiagentstore.ai) | Effort-level cost and token spread for one model | Third-party mirror, not AA's own export (notes §3.7, §6) | Not used as a data source; cited only as evidence for the effort-curve risk | n/a | +| OpenRouter / a16z "State of AI" (arXiv 2601.10088) | Token-mix evidence | FACT (notes §8.2) | Profile justification only | n/a | + +Refresh and staleness rules: +- The cost of a fixed performance level fell about 47%/quarter, about 13×/yr (FACT, notes §1.3); 5–10×/yr per Gundlach et al. (FACT, notes §1.4). Any price without `as_of` is meaningless; `V` for a model drifts with time and must be timestamped (notes §3.6). +- Every row and every rendered number carries `as_of` (existing repo rule: `CONTEXT.md` as_of; ADR 0004 item 1). +- PROPOSED: a staleness badge when a row's `as_of` is older than 90 days, and exclusion from default rankings when older than 180 days. Kyle sets the thresholds (D11). +- Fixture data is never labelled live (ADR 0003 reader rule). + +## 8. Acceptance criteria + +Each criterion is numbered and has a check that can be run or inspected. "Fixture" means a committed, `live=false` JSON file. + +### Spec-level (this document) + +- **AC-1** Every price, parameter count, deal figure, and statistic in this PRD cites `notes §n` in the same sentence or table row. Check: search this file for `$`, `B total`, `×`, `%` and confirm an adjacent `notes §` cite. +- **AC-2** No UNKNOWN or hypothesis from the notes is restated as FACT. Check: Appendix A matches notes §11 verbatim; §1 items 1–3 keep their UNKNOWN markers. +- **AC-3** The research notes file is a verbatim copy. Check: `sha256sum` of `docs/research/llm-compute-cost-research-notes-2026-10-02.md` equals that of the frozen notes supplied on 2026-10-02 (`41783be1da754c0c14ec7174dfeda44e60c55f4ca6e14dd6e1fa8da21edfe575`). + +### D1 — Open-Weight Provider Markup Index + +- **AC-4** A parameter ledger exists with the fields `model_id`, `total_params_b`, `active_params_b`, `count_convention`, `params_source_url`, `params_tier`, `as_of`; the loader rejects a row missing `params_source_url` or `as_of`. Check: `--check` exits non-zero on a fixture row with the URL removed. +- **AC-5** The ledger's seed rows reproduce the FACT-tier counts in notes §2.1 with their URLs: DeepSeek-V3 671 total / 37 active (HF checkpoint 685 incl. 14B MTP noted in `count_convention`), gpt-oss-120b 116.83 total / 5.13 active, Grok-1 314 total, Qwen3-30B-A3B 30.5 / 3.4. Check: values equal notes §2.1 at the stated precision. +- **AC-6** For each (model, provider endpoint, profile) the index equals blended price ÷ `active_params_b`, and the output row records `provider`, `price_source`, `as_of`, `profile_id`, `cache_assumption`. Check: a hand-computed fixture value matches to four significant figures. +- **AC-7** The index is computed only for `params_tier = disclosed`; an `estimated` or `unknown` row is refused with a named reason. Check: fixture test. +- **AC-8** No UI copy, API field description, or report text describes the index as capability, intelligence, or value. Copy says provider markup or serving efficiency. Check: string test over the methodology copy and field descriptions. +- **AC-9** Each endpoint row carries `quantization` with a value or `unknown`; rows with `unknown` are flagged in the output rather than silently compared. Check: fixture with one `unknown` row produces a flag. +- **AC-10** Endpoint prices come only from an allowlisted documented GET path or from a curated fixture. If `/models/{author}/{slug}/endpoints` has not been allowlisted by an accepted ADR, the client refuses it. Check: `agentbench/tests/test_openrouter_data.py::test_allowlist_is_exactly_the_four_gate1_paths` still passes, and a request for the endpoints path raises `DataAPIError` "not allowlisted". + +### D2 — Lab True-Cost Range Report + +- **AC-11** The cost curve is fitted on open-weight rows only; the report prints the calibration table (model, provider, hardware class, price, `as_of`) and a goodness-of-fit statistic. Check: no row in the calibration table has `params_tier ≠ disclosed`. +- **AC-12** Every closed-model row uses an external size band with `size_band_method ∈ {ikp, epoch}` and a URL; a closed model without a band is `unknown` and yields no numeric range. Check: fixture closed model with no band renders "no range" and no number. +- **AC-13** No size band is derived from the same lab's API price. Check: `size_band_method` never takes a value of `price-implied`; a test asserts the enum. +- **AC-14** Every true-cost output is a `[low, high]` interval; the output schema has no scalar point-estimate field. Check: JSON schema validation. +- **AC-15** Each lab row carries an ownership stratum from `{owned, rented-gpu, rented-alt-silicon, mixed}`, a bias direction, and a confidence consistent with §6.4; the report is stratified by stratum. Check: fixture rows for Google, xAI, Anthropic, OpenAI, DeepSeek match the §6.4 table. +- **AC-16** The residual is named `unexplained_gap` and is an interval; the strings "margin" and "profit" appear only inside a quotation of DeepSeek's self-reported figure. Check: string test. +- **AC-17** Priority, fast, Scale and similar tiers are separate rows with `service_tier` set and are never averaged into a standard-tier row. Check: fixture with a standard and a priority row for one model yields two output rows. +- **AC-18** The deal ledger stores the FACT deal figure and any derived per-hour value in separate, separately labelled fields. Check: the Anthropic→xAI row has `deal_value = $1.25B/month` labelled FACT and `≈ $7.8 per GPU-hour` labelled ASSUMPTION, with the notes' caveats attached. +- **AC-19** Every ranking carries M8 rank stability as "rank changes under N of K profiles". Check: the notes §8.3 worked flip fixture reproduces Qwen3.7 Max cheaper under `chat-3:1` and `openrouter-15:1`, Cerebras-launch Qwen3-Coder cheaper under `minibench-1:3` and `output-only`. +- **AC-20** The report's copy states that point estimates are unsupported and that size bands carry about 2–3× error. Check: the two sentences are present and cite the notes. +- **AC-21** Nothing from D2 joins a Solo, Multiplayer or Agent Cabinet score, and no route returns a field named `composite`, `overall`, or `index_score` computed by minibench. Check: the no-composite guard test from `docs/PLAN-benchmark-lens.md` L2.2 passes. + +### D3 — Value Frontier with Size Lens + +- **AC-22** `/models` "Capability vs. cost" axes and `_pareto_frontier` output are unchanged; size appears only as marker size or opacity for `disclosed` rows and a hollow marker otherwise. Check: existing Pareto tests pass without modification; a component test asserts marker encoding by tier. +- **AC-23** A workload-profile selector offers at least `agentic-measured`, `aa-7:2:1`, `chat-3:1`, `output-only`; the default is `agentic-measured`; `output-only` is labelled "stress test". Check: helper test on the default and labels. +- **AC-24** Every rendered number shows `as_of`, `price_source`, `profile_id` and a tier badge. Check: component test. +- **AC-25** D3 work does not start until PRs #68 and #69 have landed or been closed. Check: the implementing issue records the dependency and its resolution date. + +### Cross-cutting + +- **AC-26** No live model calls, no `/chat/completions`, no `/analytics`, no scraping, and no new secrets; the CI-equivalent checks in `AGENTS.md` pass with no `OPENROUTER_API_KEY` set. Check: CI run log shows fixture path only. +- **AC-27** Fixture data is never labelled `live=true`. Check: loader test per ADR 0003. +- **AC-28** Every output artifact (D1 table, D2 ranges, D3 payload) is available as JSON conforming to a committed schema, so an agent can consume it without parsing a chart or prose. Check: schema files exist and the CLI `--out` output validates. +- **AC-29** The methodology page and any README copy state Kyle's hypotheses at exactly the verdicts in notes §11. Check: text match against Appendix A. + +**Total: 29 acceptance criteria.** + +## 9. Risks + +### 9.1 Methodology + +| Risk | Evidence | Mitigation in this PRD | +|---|---|---| +| **Circularity.** Estimating size from price then dividing size by price makes the output a function of price plus assumed markup | ASSUMPTION, notes §2.4 | Non-goal §5 item 2; AC-13; size bands must be external (M4) | +| **Size-band width.** 2–3× error moves a closed model's ratio about 10× end to end; IKP 90% PI ≈ 3.2×; Epoch "off by a factor of 2"; NightVision ≈ 53% params error; above 1T IKP has only 2 calibration anchors | FACT/ASSUMPTION, notes §0.3, §2.3, §2.4 | Ranges only (M4, AC-14); report copy (AC-20); `unknown` when no band (AC-12) | +| **Which `P`.** MoE total vs active differ about 18–23× | FACT counts, notes §2.1; §3.1 | `active_params` default with `total_params` and `count_convention` stored (§6.1); D8 | +| **Total-band vs active-curve mismatch.** IKP predicts from total (R² 0.67 vs 0.41) while the cost curve is in active params | FACT, notes §2.3, §2.4 | Unresolved; D8 | +| **One shared cost curve does not exist.** Ownership, alternative silicon, lab-to-lab leases, and priced priority tiers each move a lab off the curve by unknown amounts and in different directions | ASSUMPTION on §10.1 FACTs, notes §10.2, §12 | Strata (§6.4); unexplained gap (M5); no latent fit (§5 item 4); D2 | +| **Alt-silicon per-token cost unknown.** TPU, Trainium, Cerebras, Broadcom | UNKNOWN, notes §6 | Widens the range rather than shifting it (§6.3 step 4) | +| **Quantization.** Providers can serve `P` at lower cost and quality; text tests detect quantization only ~50% of the time; whether endpoint data records it is UNKNOWN | FACT (secondary), notes §3.4; §6 | `quantization` field, flagged when `unknown` (AC-9) | +| **Effort and verbosity.** One model's cost per task spans $0.48 to $2.32 across effort settings, about 5× at the same $/token | FACT via third-party mirror, notes §3.7 | Measured profile as default; effort recorded per row; rank stability (M8) | +| **Blend-ratio artifacts.** Ranking order flips with the weighting | ASSUMPTION, notes §3.9, §8.3 | Profiles first-class; M8; output-only never default | +| **Unreplicated preprints.** IKP and NightVision are 2026 preprints; IKP's own acknowledgements record an external sanity check that forced a scoring revision | ASSUMPTION, notes §2.4 | Tier `estimated:` always visible; D9 | +| **Open-weight calibration assumption.** "Not priced at a significant markup" is an assumption of Gundlach et al.; DeepSeek reported a *theoretical* 545% daily margin on rented H800s | FACT, notes §9, §3.5 | Calibration residual enters the range (§6.3 step 2); DeepSeek stratified as `rented-gpu` | + +### 9.2 Legal and reputational risk of publishing cost estimates + +| Risk | Evidence | Mitigation | +|---|---|---| +| **Named-lab cost or margin claims.** A range labelled "true cost" for OpenAI or Anthropic is a claim about a private company's economics that no lab discloses | No lab discloses marginal cost (notes §11 c2) | Ranges only; `unexplained_gap` never called margin or profit (AC-16); research framing outside the leaderboard (D6); publication requires explicit owner authorization (`AGENTS.md`); D10 | +| **Secondary and unverified sources.** Deal figures from search summaries and secondary press; the S-1 was not read directly; AA figures partly from a third-party mirror; 2026 model names reproduced as printed and not checked against vendor pages | notes §6, notes header | `source_kind` per row (ADR 0004 enum); no third-party mirror as a data source (§7) | +| **Third-party IP.** IKP code license UNKNOWN; probe-release statement inconsistent | notes §6 | Cite the paper's published bands only; do not vendor code or probes (D9) | +| **OpenRouter terms.** Existing citation and GET-only rules | ADR 0002; `CONTEXT.md` | Unchanged; `/endpoints` only via an accepted ADR (D7, AC-10) | +| **Disputing a lab's hypothesis in public.** For example "Cerebras released a $500 plan" is not supported as stated | FACT, notes §11 a3 | Appendix A keeps the notes' exact verdicts; nothing is upgraded or softened (AC-2, AC-29) | +| **Being read as a leaderboard.** A ranked table of labs by "true cost" would be read as a verdict | ADR 0004 item 4 ("flags, not verdicts") | D2 stratified, interval-only, human-read; D6 | + +### 9.3 Data staleness + +| Risk | Evidence | Mitigation | +|---|---|---| +| **Prices fall fast.** About 13×/yr at fixed performance; 5–10×/yr per Gundlach et al. | FACT, notes §1.3, §1.4, §3.6 | `as_of` on everything; staleness badge and exclusion thresholds (D11) | +| **The seed catalog is already dated.** `known_models_seed.json` was generated 2026-07-06 | FACT, notes §4.2 | Regenerate from the daily poll (PROPOSED, §7) | +| **The notes' prices are examples, not quotes.** | notes header | Non-goal §5 item 9; fixtures carry the notes' dates | +| **Deals change.** The Anthropic→xAI lease is terminable on 90 days' notice by either side; Anthropic's Priority Tier is closed to new buyers; OpenAI's Broadcom chip arrives from late 2026 | FACT, notes §10.1 | Deal ledger `as_of`; stratum re-review on each filing (§7) | +| **Fixture-only operation.** The Usage Board's scheduled run has used fixtures when no key was configured | `docs/PROJECT-STATUS.md` | `live=false` labelling (AC-27); no fixture presented as current | + +## 10. Open decisions for Kyle + +Decisions D1–D6 are the notes' §12 list. D7–D12 are raised by this PRD. Where the notes give a recommendation it is shown as an ASSUMPTION; where this PRD proposes a default it is marked PROPOSED. + +| # | Decision | Options | Notes' recommendation / PRD default | +|---|---|---|---| +| **D1** | **Present "estimated true cost" as a range?** | Range always; range with a central tendency; point | Notes: yes, always. Show low/high bands combining size band × hardware-class cost range × markup range. Never show a point. Label the tier (ASSUMPTION, notes §12 d1). AC-14 assumes this. | +| **D2** | **Back out hardware ownership as a variable?** | (i) categorical covariate, stratified results; (ii) latent parameter; (iii) leave out, publish the residual as "unexplained price gap" | Notes: (i) plus (iii). Avoid (ii) unless an independent per-token cost anchor appears (ASSUMPTION, notes §12 d2). §6.4 assumes this. | +| **D3** | **Which price is "the" price per lab?** | First-party standard tier; cheapest third-party host; batch | Notes: affects every row of §10.2 (notes §12 d3). PROPOSED default: first-party standard tier as the headline, with cheapest-host and batch as additional rows under `price_basis`. | +| **D4** | **Do priority/fast tiers count as separate data points?** | Separate rows; excluded; merged | Notes: useful because they isolate the speed premium (notes §12 d4). PROPOSED default: separate rows (AC-17). | +| **D5** | **Default workload profile** when measured tokens do not exist (D1) | `agentic-measured` (n/a without runs); `aa-7:2:1`; `chat-3:1`; `output-only` stress test; `minibench-1:3` | Notes: default the headline to measured cost per task; output-only is a stress test, not the default (ASSUMPTION, notes §8.3, §12 d5). PROPOSED fallback where no runs exist: `aa-7:2:1`, with M8 stability always shown. | +| **D6** | **Where does this live?** | (a) all in minibench; (b) D2 in a separate research repo or notebook (for example `RaapTechllc/price-implied-size`) with only cited, tiered outputs republished into minibench via Benchmark Lens or Usage Board; (c) blog post only | Notes: under the reframe this is research rather than a leaderboard, which strengthens the §4.5 verdict for (b) (notes §4.5, §12 d6). This PRD is written so that D1 and D3 fit minibench and D2 can be either. | +| **D7** | **Extend the Mode A allowlist to `/models/{author}/{slug}/endpoints`?** | New ADR allowing one more GET path; curated endpoint ledger only; defer D1 live tracking | Notes: needs an ADR change; UNKNOWN whether permitted (notes §5 #2, §6). PROPOSED default: curated ledger first (AC-10), ADR later if D1 proves useful. | +| **D8** | **Which `P`, and how to map total-parameter external bands onto the active-parameter cost curve?** | Active only (drop total-only bands); carry both and show the band in total with a sensitivity note; attempt a conversion | Notes: store `total_params`, `active_params`, `count_convention` (notes §2.1); IKP predicts better from total (notes §2.3). PROPOSED default: active for the curve; closed-model bands shown as reported with their basis labelled; no conversion. | +| **D9** | **Which external size-band sources are admissible?** | IKP; Epoch; both; MEDEC excluded or `imported` only | Notes: IKP and NightVision unreplicated (notes §2.4); IKP license UNKNOWN (notes §6); MEDEC unverified (notes §2.3). PROPOSED default: IKP and Epoch bands as `estimated:`; MEDEC excluded from ranges. | +| **D10** | **Publication posture for named-lab ranges** | Publish with lab names; publish with strata only and no lab names; internal only until reviewed | PROPOSED default: internal draft first; publication is a separately authorized step (`AGENTS.md`). | +| **D11** | **Refresh cadence and staleness thresholds** | Badge at 90 days / exclude at 180 (PROPOSED); other values; no exclusion | Notes: ~13×/yr price decline makes timestamps essential (notes §1.3, §3.6). | +| **D12** | **Source for the throughput cross-check (M6)** | OpenRouter `/models` throughput sort (allowlisted, notes §1.6); none; defer | PROPOSED default: allowlisted catalog throughput only; no measurement runs. | + +**Total: 12 open decisions** (6 from notes §12, 6 raised by this PRD). + +## 11. GO gate + +**No build until Kyle accepts this PRD.** + +Acceptance means all of the following, recorded in this file: + +1. The Status line at the top changes from `PROPOSED` to `ACCEPTED` with a date and Kyle's name. +2. Decisions D1, D2, D3, D5 and D6 have a recorded choice. The remaining decisions may fall to their PROPOSED defaults if Kyle writes "default" against them. +3. If D6 chooses a separate research repo or notebook for D2, the minibench scope reduces to D1 and D3 plus a republishing path, and that is written into §4. + +Until the gate clears: + +- No product code, schema change, fixture, ledger, ADR status flip, or frontend change is made for this PRD. +- Open PRs #68, #69 and #70 are not touched. +- The research notes stay frozen; corrections go in a new dated notes file, not edits to the frozen one. + +After the gate clears, the next artifacts are a handoff plan in the shape of `docs/PLAN-benchmark-lens.md` and, if D7 is answered yes, a new ADR for the allowlist change. Paid live model calls, publishing, deployment and external communication remain separately authorized under `AGENTS.md` regardless of this gate. + +--- + +## Appendix A. Kyle's hypotheses: verdicts (reproduced exactly from notes §11) + +| # | Claim | Verdict | Evidence | +|---|---|---|---| +| a1 | OpenAI uses Cerebras | **FACT** | 750 MW deal; Codex-Spark on Cerebras (§10.1) | +| a2 | …but keeps most of that capacity for itself | **UNKNOWN / partially supported** | Only public signal: Codex-Spark is Pro-only, API limited to "select design partners". No disclosure of capacity split. The deal is Cerebras-hosted capacity for "OpenAI customers" (Cerebras blog) | +| a3 | Cerebras just released a $500 plan | **Not supported as stated** | The $500/month plan is **OpenAI's ChatGPT Pro 500** (DevDay 2026-09-29, "Astra Ultrafast"). Pre-launch leaks tied it to Cerebras, but OpenAI pages do not name the hardware (Business Insider; https://pasqualepillitteri.it/en/news/19330/openai-launches-pro-500-500-dollars-month, accessed 2026-10-02). Cerebras Code plans are $50 and $200, shown as sold out — https://www.cerebras.ai/code, https://support.cerebras.net/articles/9996007307-cerebras-code-faq (accessed 2026-10-02). Ultrafast on Cerebras: **UNKNOWN** | +| b1 | xAI owns its own data centres | **FACT** | Colossus 1, "built from the ground up" (xAI), Memphis | +| b2 | xAI sells compute to others, e.g. Anthropic | **FACT** | xAI 2026-05-06 announcement; S-1 terms $1.25B/month to May 2029 (TechCrunch) | +| c1 | Owned or reserved silicon and priority deals mean API price ≠ lab cost | **Supported in direction** | Priority/Scale tiers price capacity, not size; DeepSeek theoretical margin; xAI monetising "unused" capacity. Magnitude per lab **UNKNOWN** | +| c2 | Owners price near marginal cost; renters carry a premium | **ASSUMPTION (plausible, unproven)** | No lab discloses marginal cost. Gundlach et al. attribute part of the closed-vs-open price gap to competition and markup, not ownership. Owners may also price *above* cost to recover capex | + +Section references inside this table (§10.1) point into the notes, not this PRD. + +## Appendix B. Repo fit map + +Where each piece would live if D6 keeps it in minibench. All rows are ASSUMPTION or PROPOSED design sketches, not changes. + +| Piece | Location | Basis | +|---|---|---| +| Parameter fields on the catalog | Extend `KnownModel` and `backend/app/data/known_models_seed.json` with `total_params_b`, `active_params_b`, `params_source_url`, `params_tier`, `params_band_low/high`; not the legacy `/api/v1/models` quality table that PR #69 removes | notes §4.3, §4.4 | +| D1 compare route | Usage Board, "best-by-$-per-active-param", open weights only | notes §4.4 Surface B | +| D3 size channel | `frontend/src/pages/Models.tsx` "Capability vs. cost" marker encoding; `_pareto_frontier` unchanged | notes §4.2, §4.4 Surface A | +| Closed-model size bands | Benchmark Lens (ADR 0004, proposed, not implemented) as `estimated:` rows with provenance tier | notes §4.4 Surface C; ADR 0004 items 1–2, 6 | +| Ledgers and loaders | `agentbench/data/*.json` with a runtime override path following the ADR 0003 pattern; pure-function loaders in `agentbench/` | `docs/PLAN-benchmark-lens.md` concept map | +| Glossary | `CONTEXT.md` gains: Workload profile, Provider markup index, Size tier (`disclosed` / `estimated:` / `unknown`), Ownership stratum, Unexplained gap | `docs/agents/domain.md` | +| Sequencing | After #68 and #69 land or close | notes §4.4 | + +ADR interactions flagged per `docs/agents/domain.md`: this PRD does not change ADR 0002, 0003 or 0004. D7 would require a new ADR because ADR 0002 fixes the Mode A path allowlist and ADR 0004 item 6 says any automated fetch from a new host or path is a separate ADR.