Skip to content

fix(capture): resume from the last fire instead of skipping past 400 lines - #173

Merged
AkashGoenka merged 2 commits into
mainfrom
fix/capture-resume-from-last-fire
Sep 19, 2026
Merged

AkashGoenka merged 2 commits into
mainfrom
fix/capture-resume-from-last-fire

Conversation

@AkashGoenka

@AkashGoenka AkashGoenka commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Fixes #172.

What was wrong

A single prompt that runs the agent through a lot of turns ended up captured as nothing. The task finishes, /capture-notes says it was never asked to write anything up.

The capture hook keeps a marker in the OS temp directory holding how far it has read into the transcript. Temp gets swept every few days, the transcript does not, so a session resumed across days keeps losing the offset while the full history stays on disk. The guard for that read "transcript over 400 lines" as "already accounted for", snapped the offset to the end of the file and recorded nothing.

Line count cannot answer that question. Every tool call writes a couple of transcript lines, so 62 of 121 transcripts in this repo pass 400 in a single sitting, and ordinary long tasks got mistaken for stale history. Measured from capture.jsonl: 24 discards across 8 sessions, all spanning 5 to 11 days, one of them hit six times.

What this does

Asks the durable record instead of guessing from size. capture.jsonl lives in the repo, is append only, and stamps every fire with a session id and timestamp, so it survives the sweep that eats the marker.

When the marker is missing, find the newest fire for this session, find the first transcript line stamped after it, and resume there. Everything before that point was already put in front of the agent, everything after it never was.

Fire events only. A stop that merely processed evidence banked it in the marker that just got swept, and the old baseline events mark discarded history rather than offered history, so counting either would skip work nobody was ever asked about.

Both unknowns fail towards replay rather than skipping. No fire on record means the session was never asked for anything, so replay all of it however big. No timestamps to place the boundary means replay too.

New in hooks/elicit-core.mjs:

  • lastFireAt(root, sid) reads capture.jsonl for the newest fire for this session
  • lineIndexAfter(lines, isoTs) finds the first transcript line stamped after it

Both sides are new Date().toISOString(), fixed width UTC, so lexicographic compare is chronological and no date parsing is needed.

Verified

Ran the real hook three times against an identical transcript, changing only what was on record:

on record result captured
fire between the two reads REATTACH resumeAt=2 only the read after the fire
nothing (control) REPLAY both reads
only a baseline REPLAY both reads

The third case matters: the discards already sitting in history do not count as offered, so that work gets re-offered rather than staying lost.

While checking this I had a probe that showed cases 1 and 2 behaving identically, which looked like the fix was dead. It turned out the probe emitted "timestamp": "..." with a space while real transcripts are compact. Harness bug rather than a code bug, but the pattern is whitespace tolerant now, since a miss there silently degrades to a full replay.

Cost

Largest transcript in this repo, 39k lines and 126MB:

  • lastFireAt 0.6ms
  • lineIndexAfter 25ms
  • about 25k already-offered lines skipped, so reattaching does less work than replaying

The rule it replaces was trading correctness for roughly half a second.

Windows

CI already matrixes ubuntu and windows, so the new tests run on both. Checked specifically:

Rejected alternatives

Move the marker to durable storage. Looks like the obvious fix and breaks something else. soleMarkerUnderRoot resolves which session it is in by finding exactly one marker lying around, which only holds because temp keeps getting swept (there was exactly one present across 121 sessions when checked). Permanent markers mean every past session leaves one behind, the count is never one again, and manual /capture-notes on Cursor and Codex stops working. It also brings its own garbage collection problem.

Timestamp gap detection. A gap does not mean the work was captured, it means somebody went to lunch. hooks/trigger.mjs rules this class of signal out by design.

Known gap

Files a fire ranked past MAX_CAPTURE_FILES were read before that fire, so a sweep still forgets them. Much smaller than dropping the whole span, worth revisiting only with evidence it bites.

Test plan

  • npx vitest run, 861 tests across 57 files pass
  • npx tsc --noEmit clean
  • old test pinning the discard behaviour replaced by three pinning the new contract (full replay with no fire, reattach from last fire, replay when the boundary cannot be placed)
  • real hook exercised out of band for all three paths
  • CI green on windows-latest

🤖 Generated with Claude Code

AkashGoenka and others added 2 commits September 19, 2026 23:48
…lines

A single prompt that runs the agent through many turns was captured as
nothing at all: the task finishes and /capture-notes reports it was never
asked to write anything up.

The trigger marker holds the transcript read offset and lives in the OS temp
dir, which is swept every few days while the transcript survives. The guard
for that read "transcript > 400 lines" as "already accounted for" and snapped
the offset to the end of the file, discarding every read and edit since the
sweep. Line count cannot answer that question. Every tool call writes a couple
of transcript lines, so 62 of 121 transcripts in this repo pass 400 in one
sitting, and ordinary long tasks were mistaken for stale history. Measured
footprint: 24 discards across 8 sessions, all spanning 5 to 11 days, one hit
six times.

Ask the durable record instead of a proxy. capture.jsonl survives the sweep
and stamps every fire with session + ts, so the last fire marks the point up
to which this session was already asked for notes; resume there. Fire events
only: a stop that merely processed evidence banked it in the swept marker, and
baseline events mark discarded rather than offered history, so counting either
would skip work nobody was ever asked about. Both unknowns fail towards replay,
no fire on record and no placeable boundary each replay in full, because losing
unasked work is the failure that matters.

Costs 0.6ms + 25ms on the largest transcript here (39k lines) and skips ~25k
already-offered lines, so it does less work than the replay it replaces.

Rejected: relocating the marker to durable storage, which breaks
soleMarkerUnderRoot (it resolves the session by finding exactly one marker,
which only holds because temp is swept) and brings its own GC problem; and
timestamp-gap detection, which trigger.mjs rules out by design.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@AkashGoenka
AkashGoenka merged commit 4cf2d46 into main Sep 19, 2026
8 checks passed
@AkashGoenka
AkashGoenka deleted the fix/capture-resume-from-last-fire branch September 19, 2026 18:22
@AkashGoenka AkashGoenka mentioned this pull request Sep 19, 2026
AkashGoenka added a commit that referenced this pull request Sep 19, 2026
Bumps package.json, package-lock.json and server.json to 2.3.3 (all three must
match the tag; server.json is validated against the published tarball) and adds
the CHANGELOG entry for #173.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Capture silently drops a session's work once the transcript passes 400 lines

1 participant