Skip to content

feat(ai): warn when a day's AI spend crosses the line, and never stop on it - #143

Merged
ralyodio merged 1 commit into
masterfrom
feat/ai-spend-alert
Jul 28, 2026
Merged

ralyodio merged 1 commit into
masterfrom
feat/ai-spend-alert

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Emails anthony@profullstack.com when a day's AI spend passes $15. It warns only — nothing throttles, pauses, or blocks.

Why spend is computed here rather than read from Anthropic

Anthropic's cost/usage endpoints need an Admin key (sk-ant-admin01-…). Verified against your key:

v1/organizations/cost_report            HTTP=401
v1/organizations/usage_report/messages  HTTP=401

So the app that's spending the money can't read its own total, or the account balance. What every response does carry is token counts, and per-model rates are published — so cost is computable at the call site, which is also the only place that knows which feature spent it.

Balance is therefore not covered. If you want the actual account balance, that needs an Admin key — say the word and I'll add it as a second source.

Recorded at the chokepoint

generateStructuredOutput is the single function all generation passes through. Recorded before the parse check: tokens are spent whether or not the output parses, and a failed parse is exactly the waste worth seeing. Writing the row can't fail a request that already succeeded.

Micro-dollars, not cents

A Haiku draft costs ~$0.00185. In whole cents, a full day of real spending reports as $0.00. Rates used are stored per row so a later price change doesn't rewrite history, and an unknown model records a null rate rather than a confident zero.

Cache reads (~0.1×) and writes (~1.25×) are priced at their own multipliers — counting only fresh input would understate a cached workload by most of its bill.

Warns once, never stops

De-duplicated on (day, threshold) via a unique index plus an insert-first claim, so a day that crosses the line mails once rather than on every subsequent run — and correctly even if two cron runs overlap.

The email leads with the number and a per-feature breakdown, because "you spent $18" prompts "on what". It states plainly that nothing was stopped.

Setup

  • Schedule GET /api/cron/ai-spend hourly with x-cron-secret
  • AI_SPEND_DAILY_ALERT_USD (default 15), AI_SPEND_ALERT_EMAIL (default anthony@profullstack.com)

Checks

  • tsc --noEmit clean
  • 863/863 tests pass, 11 new — including that dated model ids (claude-haiku-4-5-20251001) still price correctly, and that sub-cent calls don't round to zero
  • production build compiles; route registered
  • migration applied to prod

🤖 Generated with Claude Code

… on it

The provider will not tell the application what it is spending. Anthropic's
cost and usage reports need an Admin key; a normal API key gets a 401 on
them, so the app doing the spending cannot read its own running total or
the account balance. Every response does carry its token counts and the
per-model rates are published, so the cost is computable at the point of
the call — which is also the only place that knows which feature caused it.

Spend is recorded at the one function every generation goes through, and
recorded before the parse check: the tokens are gone whether or not the
output turns out usable, and a failed parse is exactly the waste worth
seeing in a total. Writing the row cannot fail the request that already
succeeded.

Amounts are micro-dollars. A Haiku draft costs about a fifth of a cent, so
in whole cents a day of real spending reports as zero. The rates used are
stored alongside each row so a later price change does not silently rewrite
history, and an unknown model records a null rate rather than a confident
zero.

Cache reads and writes are priced at their own multipliers rather than
ignored. Counting only fresh input would understate a cached workload by
most of its bill.

It warns and nothing else. No throttle, no pause, no circuit breaker — a
budget alarm that turns the product off is worse than the bill it was
meant to prevent. De-duplicated on (day, threshold) by a unique index and
an insert-first claim, so a day that crosses the line mails once instead
of on every subsequent run, even if two runs overlap.

The mail leads with the number and the per-feature breakdown, because
"you spent $18" prompts "on what" and the answer should not require
opening a dashboard. It says plainly that nothing was stopped, so it is
not read as an outage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

@ralyodio
ralyodio merged commit c0959d8 into master Jul 28, 2026
8 checks passed
@ralyodio
ralyodio deleted the feat/ai-spend-alert branch July 28, 2026 07:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant