Skip to content

fix(agent-activity): cache a context id only after its line is written (#1303) - #1414

Merged
Juliusolsson05 merged 7 commits into
integration/batch-2026-09-27-rfrom
fix/activity-context-id-after-write
Sep 27, 2026
Merged

Juliusolsson05 merged 7 commits into
integration/batch-2026-09-27-rfrom
fix/activity-context-id-after-write

Conversation

@Juliusolsson05

@Juliusolsson05 Juliusolsson05 commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Fixes #1303.

Problem

AgentActivityStore.appendInterval cached a new context id (ids.set(key, id)) BEFORE the append that writes that context line. If the append failed (ENOSPC, EIO; the store's queue only console.warns), every later interval for that agent in the same month referenced a context line that never reached disk, and readIntervals silently dropped each one. One failed write cost the agent's hours for the rest of the month.

What merges

  • The id is cached only after appendLines succeeds (Agent Analytics silently loses an agent's hours after one failed context write #1303).
  • Ids come from a per-month counter and are never reissued (reviews a+b). A failed or partial write burns its id. After a restart the counter continues from the HIGHEST id on disk; it is not recomputed from the number of contexts, which would reuse an id after a gap. The ids.size + 1 fallback is removed.
  • "Unknown is never empty" (q115):
    • an unreadable month file refuses the append instead of restarting ids;
    • an unreadable or corrupt open.json is set aside to open.json.unrecovered-<time>, never overwritten. Later starts recover each readable copy. A recovery that fails partway sets aside only the unrecovered remainder, so nothing is counted twice;
    • an unreadable aliases.jsonl refuses the append (its tail cannot be checked) and is not cached as empty (B6 check).
  • The interval whose write failed is still lost: there is nowhere to put it.

Tests (real files, AgentActivityStore.test.ts)

  • a failed context write does not orphan later intervals;
  • a partially written context never lets another agent reuse its id;
  • an id gap left by a failed write is never filled after a restart;
  • an unreadable month file refuses the append, and ids continue once it is readable;
  • an unreadable open.json is set aside, not overwritten, and recovered once readable;
  • a partial recovery sets aside only the remainder; an unreadable set-aside copy is kept;
  • an unreadable aliases file refuses the append, and works once readable (pins the tail check);
  • the same store reading while aliases.jsonl is unreadable throws, then groups A and B once it is readable (pins loadAliases' ENOENT-only catch; red with a catch-all).

Each is red on the code it fixes, and the named mutations are killed. The aliases case has two separate guards, each with its own test. The read path depends on loadAliases alone. The append path is also refused by the tail check.

Residual: the ENOENT-only guard in contextsFor is pinned for EACCES only; EISDIR fails the append itself, and EIO is not reproducible on a real filesystem.

Gate: npx tsc -b is clean; src/main/agentActivity + src/renderer/src/features/agent-activity pass, 6 files / 78 tests at 4241aa25.

🤖 Generated with Claude Code

#1303)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Juliusolsson05 Juliusolsson05 added the type:bug Something works wrong label Sep 27, 2026
Juliusolsson05 and others added 6 commits September 27, 2026 05:44
…tial write (#1414 review a+b)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ead of restarting ids (#1303, #1414 review a round 2)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e and retried, never overwritten (#1414 review a round 3)

Recovery treated any failure as 'no open file' and overwrote open.json with
an empty snapshot, losing the pending interval. Only ENOENT is empty now;
anything else is moved to open.json.unrecovered-<time> (a rename needs no
read permission) and later starts recover each copy they can read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…recovered; no reissue-by-size fallback (#1414 review c)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d is not cached as empty (#1414, B6 check)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The same store reads while aliases.jsonl is unreadable, then must group A
and B once readable. Red with a catch-all; the append test could not pin
it because the tail check refuses that append on its own (B6 check 2110).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Juliusolsson05

Copy link
Copy Markdown
Owner Author

Disposition at 4241aa2 (a, b and c rounds; review cap reached; B6 manual inspection, then manager-verify a/b/c MERGE-READY):

  • a + b r1 (major): a partial write let another agent reuse a context id, so A's time could be attributed to B after a restart. Fixed: a per-month id counter that never reissues an id; after a restart it continues from the highest id on disk (8755ce2).
  • a r2 (q115): an unreadable month file restarted ids at 1. Fixed: only ENOENT means no file; anything else refuses the append (4c64d3d). b r2's restart-after-an-id-gap test was added.
  • a r3 (q115): recovery overwrote an unreadable open.json with an empty snapshot. Fixed: it is set aside to open.json.unrecovered-<time> and recovered on a later start.
  • c r1 (major): a partial recovery set aside the whole snapshot, so a retry double-counted it. Fixed: only the unrecovered remainder is set aside (7ccbdd6). The keep-an-unreadable-copy rule is pinned, and the dead ids.size + 1 fallback is removed.
  • B6 check (1950 + 2110): an unreadable aliases.jsonl was read as empty. Fixed: loadAliases and the append tail both treat only ENOENT as absent. Each has its own fail-first test (12960b6, 4241aa2).
  • b r1 minor, a lost interval is absent from the display: not fixed. The interval whose write failed has nowhere to go and is lost. The body says so.
  • Residuals:
    • the ENOENT-only guard in contextsFor is pinned for EACCES only; EISDIR fails the append itself, and EIO cannot be reproduced on a real filesystem;
    • c r2 S16: if the remainder write AND the set-aside write both fail, the whole snapshot is renamed aside, and the next start may re-append the prefix that already landed. That takes a double disk failure.

Gates: tsc -b clean; 78 tests at 4241aa2; CI green.

🤖 Generated with Claude Code

@Juliusolsson05
Juliusolsson05 changed the base branch from main to integration/batch-2026-09-27-r September 27, 2026 21:32
@Juliusolsson05
Juliusolsson05 merged commit 72a753a into integration/batch-2026-09-27-r Sep 27, 2026
2 checks passed
@Juliusolsson05
Juliusolsson05 deleted the fix/activity-context-id-after-write branch September 27, 2026 21:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something works wrong

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Agent Analytics silently loses an agent's hours after one failed context write

1 participant