Repository navigation
feat(ai): warn when a day's AI spend crosses the line, and never stop on it - #143
Merged
Merged
Conversation
… on it The provider will not tell the application what it is spending. Anthropic's cost and usage reports need an Admin key; a normal API key gets a 401 on them, so the app doing the spending cannot read its own running total or the account balance. Every response does carry its token counts and the per-model rates are published, so the cost is computable at the point of the call — which is also the only place that knows which feature caused it. Spend is recorded at the one function every generation goes through, and recorded before the parse check: the tokens are gone whether or not the output turns out usable, and a failed parse is exactly the waste worth seeing in a total. Writing the row cannot fail the request that already succeeded. Amounts are micro-dollars. A Haiku draft costs about a fifth of a cent, so in whole cents a day of real spending reports as zero. The rates used are stored alongside each row so a later price change does not silently rewrite history, and an unknown model records a null rate rather than a confident zero. Cache reads and writes are priced at their own multipliers rather than ignored. Counting only fresh input would understate a cached workload by most of its bill. It warns and nothing else. No throttle, no pause, no circuit breaker — a budget alarm that turns the product off is worse than the bill it was meant to prevent. De-duplicated on (day, threshold) by a unique index and an insert-first claim, so a day that crosses the line mails once instead of on every subsequent run, even if two runs overlap. The mail leads with the number and the per-feature breakdown, because "you spent $18" prompts "on what" and the answer should not require opening a dashboard. It says plainly that nothing was stopped, so it is not read as an outage. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vu1nz Security Review0 finding(s) in PR #? No security issues found. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Emails
anthony@profullstack.comwhen a day's AI spend passes $15. It warns only — nothing throttles, pauses, or blocks.Why spend is computed here rather than read from Anthropic
Anthropic's cost/usage endpoints need an Admin key (
sk-ant-admin01-…). Verified against your key:So the app that's spending the money can't read its own total, or the account balance. What every response does carry is token counts, and per-model rates are published — so cost is computable at the call site, which is also the only place that knows which feature spent it.
Balance is therefore not covered. If you want the actual account balance, that needs an Admin key — say the word and I'll add it as a second source.
Recorded at the chokepoint
generateStructuredOutputis the single function all generation passes through. Recorded before the parse check: tokens are spent whether or not the output parses, and a failed parse is exactly the waste worth seeing. Writing the row can't fail a request that already succeeded.Micro-dollars, not cents
A Haiku draft costs ~$0.00185. In whole cents, a full day of real spending reports as
$0.00. Rates used are stored per row so a later price change doesn't rewrite history, and an unknown model records a null rate rather than a confident zero.Cache reads (~0.1×) and writes (~1.25×) are priced at their own multipliers — counting only fresh input would understate a cached workload by most of its bill.
Warns once, never stops
De-duplicated on
(day, threshold)via a unique index plus an insert-first claim, so a day that crosses the line mails once rather than on every subsequent run — and correctly even if two cron runs overlap.The email leads with the number and a per-feature breakdown, because "you spent $18" prompts "on what". It states plainly that nothing was stopped.
Setup
GET /api/cron/ai-spendhourly withx-cron-secretAI_SPEND_DAILY_ALERT_USD(default15),AI_SPEND_ALERT_EMAIL(defaultanthony@profullstack.com)Checks
tsc --noEmitcleanclaude-haiku-4-5-20251001) still price correctly, and that sub-cent calls don't round to zero🤖 Generated with Claude Code