Stop overpaying to run your agents. Kalibr routes every request to lower-cost model and tool paths without degrading performance.
-
Updated
Jun 3, 2026 - Python
Stop overpaying to run your agents. Kalibr routes every request to lower-cost model and tool paths without degrading performance.
Measures dollars per correct outcome on LLMs. Contamination-resistant: benchmark questions are generated fresh at runtime.
What LLM inference actually costs. 874 verified price rows across 24 providers, 378 GPU rental rates, 395 cited throughput datapoints, and a self-hosting break-even model. Every number carries a source URL, a retrieval date, and passed a mechanical evidence gate. Dated immutable snapshots. CC BY 4.0.
Caveman prompting, measured. A two-channel evaluation protocol scoring what input and output compression actually cost LLMs in dollars, accuracy, and surface-text fidelity across seven models and five benchmarks.
Measures and protects the server-side KV cache your coding agent keeps invalidating
Drop-in OpenAI-compatible proxy that routes each request to the best model by cost, quality, or latency — with a pluggable scoring engine and a built-in eval harness. Self-hosted.
Tamper-evident, stranger-verifiable receipts for LLM cost-savings — anyone can recompute your caching/routing savings math offline, no trust in your dashboard required. Pure stdlib, zero-dependency.
Budget-aware degradation for autonomous AI agents — map a spend balance to survival tiers, enforce per-call, hourly and daily LLM inference budgets, and automatically downgrade the model when funds run low.
Inference cost allocation for autonomous AI agent collaborations — Shapley-fair splitting, congestion pricing, token metering. Part of the Agent Trust Stack.
Cost-aware LLM gateway: semantic cache + difficulty-based model routing that cuts spend (~66%) and latency vs always using the frontier model, measured against a baseline and gated in CI. Fully offline.
Locational Cost of Intelligence: a location-adjusted cost function with QoS chance constraints, with code and data schemas for an Intelligence Price Deflator. Headline empirical result withdrawn pending reconstruction, see STATUS.md.
AI Infrastructure Portfolio: agent orchestration, model routing, inference cost analysis — working code, real measurements.
NeoSmith Maestro — frontier-level coding accuracy from a bouquet of small models at ~1/30th the cost. Technical report + reproducible evidence (LiveCodeBench v6 92.2%, SWE-bench Pro 74.2/88.9).
Inference cost allocation for autonomous AI agent collaborations — Shapley-fair splitting, congestion pricing, token metering. Part of the Agent Trust Stack.
Add a description, image, and links to the inference-cost topic page so that developers can more easily learn about it.
To associate your repository with the inference-cost topic, visit your repo's landing page and select "manage topics."