Skip to content

feat(engine): record agent spend on every phase exit (U1) - #25

Merged
Steel-tech merged 3 commits into
mainfrom
feat/spend-every-exit
Sep 28, 2026
Merged

Steel-tech merged 3 commits into
mainfrom
feat/spend-every-exit

Conversation

@Steel-tech

Copy link
Copy Markdown
Contributor

Implements U1 of the factory-quality follow-ups plan (PR #21): R1, KTD2.

What changed

  • agent_end on every agent phase exit. Before this, it was emitted only when an entry succeeded, so every failed, dead, or attempt-ending entry dropped out of any spend figure. Now every entry that emits agent_start closes it with exactly one agent_end. The event comes before whatever closes the phase or the attempt (phase_end, phase_death, attempt_send_budget_exhausted, attempt_ceiling_exceeded). The payload adds:
    • outcome, one of passed | failed | died | aborted | cancelled | ceiling | send_budget. These are pinned as protocol.Agent* constants for the report (U2).
    • unmetered_sends
  • Runtime-error sends are metered. Their reported usage is added before send returns (KTD2).
  • Killed sends count as unmetered, not free. This covers watchdog, phase clock, ceiling, and cancel, plus a send that ends without a usable result (crash). Invariant: sends = metered + unmetered per entry.
  • Claude Code cost is now per send. total_cost_usd turned out to be cumulative (see below). The adapter records the session's reported total on a new runtime.Session.ReportedCostUSD and returns each send's own share as Usage.CostUSD. The engine's += is unchanged and is now correct. Sessions are only ever built fresh in memory (deaths and CI-repair rounds start new ones), so the running total never resets mid-conversation.

total_cost_usd semantics: cumulative across --resume

Two real calls, claude 2.1.283, --model haiku, -p "say ok":

call usage (per invocation) total_cost_usd modelUsage tokens
1 (new session) in 10 / out 35 / cache-write 19,483 / cache-read 13,689 0.0405199 same as usage
2 (--resume same id) in 10 / out 36 / cache-write 7,237 / cache-read 33,172 0.0585011 in 20 / out 71 / cache-write 26,720 / cache-read 46,861 (sums of both)

Pricing call 2's own usage at Haiku 4.5 rates gives $0.0179812, and 0.0405199 + 0.0179812 = 0.0585011 exactly. So total_cost_usd is the session total, while usage is per invocation. Before this fix, every resumed correction send was double-counting all earlier sends in its session.

Verification

  • just check passes (format, vet, boundary, definitions, tests, build).
  • go test -race ./internal/engine/... ./internal/runtime/... passes.
  • New engine tests cover each exit path:
    • pass, status: fail, runtime error, death with re-entry, send-budget mid-phase, send budget refusing an entry's first send, ceiling, cancel, and boundary abort;
    • the unit's verification check: summed agent_end.cost equals the usage of the sends that returned a result, and summed unmetered_sends equals the killed and crashed sends.
  • Adapter test: a resumed send reporting a running total of 1.25 after 0.5 yields a per-send cost of 0.75.
  • Exact-sequence tests. The existing three-phase test has only passing entries, so its sequence doesn't change; it now also pins each agent_end outcome. A new exact-sequence test pins the new exits in order: crash → agent_end(died) → phase_death → re-entry → pass → next phase refused by the send budget → agent_end(send_budget) → attempt_send_budget_exhausted.

Mutation testing

18 mutants were tried, and each one compiled (go vet), so no compile error was counted as a catch. The first run killed 16. The two survivors were the agent_end calls on the breach (abort) branches inside the failed-entry and attempt-terminal exits. I added agent_end assertions to the two existing boundary tests that drive those branches, and both mutants are now killed: 18/18.

The mutants also covered: runtime-error metering, the unmetered count on kill and on crash, the agent_end → phase_death order, every outcome label, the payload field, and both adapter delta lines.

Known gaps / untested

  • If the CLI exits nonzero after emitting a result, the adapter returns an error and the engine treats the send as a death. That send is counted unmetered even though a cost was reported. The accounting is conservative, and the death handling is unchanged.
  • No live-CLI test covers the per-send delta, only the stub. The real-CLI numbers above are the evidence. The codex adapter reports no cost (ReportsCost=false), so it's unaffected.
  • The early failInfra exits (snapshot, seed, prompt read) happen before agent_start, so they correctly emit no agent_end.

🤖 Generated with Claude Code

Steel-tech and others added 3 commits September 28, 2026 00:27
agent_end was emitted only when an agent phase entry succeeded, so any
cost figure undercounted every failed, dead, or attempt-ending entry.
Every entry that emits agent_start now closes it with agent_end carrying
an outcome (passed, failed, died, aborted, cancelled, ceiling,
send_budget) and an unmetered_sends count, before the event that closes
the phase or the attempt.

A runtime-error send's reported usage is now added to spend. A send
killed in flight (watchdog, phase clock, ceiling, cancel) or ending
without a usable result is counted as unmetered, not as free.

Claude Code's total_cost_usd is the session's running total across
--resume, not the invocation's cost: a real haiku resume reported
0.0585011 after a first send of 0.0405199. The adapter now records the
session's reported total on runtime.Session and returns each send's own
share, so the engine's per-send sum is correct.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A discontinuity is a send death and the engine drops that session, so
the recorded total is never read again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l entries

Mutation testing showed the abort paths inside fail and terminalExit
could drop agent_end without a test noticing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 46 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: ac282891-0c6e-43af-be32-6d5ad982d40a

📥 Commits

Reviewing files that changed from the base of the PR and between 8c542bb and 3589b38.

📒 Files selected for processing (6)
  • internal/engine/phase.go
  • internal/engine/phase_test.go
  • internal/protocol/types.go
  • internal/runtime/claudecode/adapter.go
  • internal/runtime/claudecode/adapter_test.go
  • internal/runtime/runtime.go

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Steel-tech
Steel-tech merged commit dcdcc28 into main Sep 28, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant