fix(capture): resume from the last fire instead of skipping past 400 lines - #173
Merged
Merged
Conversation
…lines A single prompt that runs the agent through many turns was captured as nothing at all: the task finishes and /capture-notes reports it was never asked to write anything up. The trigger marker holds the transcript read offset and lives in the OS temp dir, which is swept every few days while the transcript survives. The guard for that read "transcript > 400 lines" as "already accounted for" and snapped the offset to the end of the file, discarding every read and edit since the sweep. Line count cannot answer that question. Every tool call writes a couple of transcript lines, so 62 of 121 transcripts in this repo pass 400 in one sitting, and ordinary long tasks were mistaken for stale history. Measured footprint: 24 discards across 8 sessions, all spanning 5 to 11 days, one hit six times. Ask the durable record instead of a proxy. capture.jsonl survives the sweep and stamps every fire with session + ts, so the last fire marks the point up to which this session was already asked for notes; resume there. Fire events only: a stop that merely processed evidence banked it in the swept marker, and baseline events mark discarded rather than offered history, so counting either would skip work nobody was ever asked about. Both unknowns fail towards replay, no fire on record and no placeable boundary each replay in full, because losing unasked work is the failure that matters. Costs 0.6ms + 25ms on the largest transcript here (39k lines) and skips ~25k already-offered lines, so it does less work than the replay it replaces. Rejected: relocating the marker to durable storage, which breaks soleMarkerUnderRoot (it resolves the session by finding exactly one marker, which only holds because temp is swept) and brings its own GC problem; and timestamp-gap detection, which trigger.mjs rules out by design. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Merged
AkashGoenka
added a commit
that referenced
this pull request
Sep 19, 2026
Bumps package.json, package-lock.json and server.json to 2.3.3 (all three must match the tag; server.json is validated against the published tarball) and adds the CHANGELOG entry for #173. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #172.
What was wrong
A single prompt that runs the agent through a lot of turns ended up captured as nothing. The task finishes,
/capture-notessays it was never asked to write anything up.The capture hook keeps a marker in the OS temp directory holding how far it has read into the transcript. Temp gets swept every few days, the transcript does not, so a session resumed across days keeps losing the offset while the full history stays on disk. The guard for that read "transcript over 400 lines" as "already accounted for", snapped the offset to the end of the file and recorded nothing.
Line count cannot answer that question. Every tool call writes a couple of transcript lines, so 62 of 121 transcripts in this repo pass 400 in a single sitting, and ordinary long tasks got mistaken for stale history. Measured from
capture.jsonl: 24 discards across 8 sessions, all spanning 5 to 11 days, one of them hit six times.What this does
Asks the durable record instead of guessing from size.
capture.jsonllives in the repo, is append only, and stamps every fire with a session id and timestamp, so it survives the sweep that eats the marker.When the marker is missing, find the newest fire for this session, find the first transcript line stamped after it, and resume there. Everything before that point was already put in front of the agent, everything after it never was.
Fire events only. A stop that merely processed evidence banked it in the marker that just got swept, and the old
baselineevents mark discarded history rather than offered history, so counting either would skip work nobody was ever asked about.Both unknowns fail towards replay rather than skipping. No fire on record means the session was never asked for anything, so replay all of it however big. No timestamps to place the boundary means replay too.
New in
hooks/elicit-core.mjs:lastFireAt(root, sid)readscapture.jsonlfor the newest fire for this sessionlineIndexAfter(lines, isoTs)finds the first transcript line stamped after itBoth sides are
new Date().toISOString(), fixed width UTC, so lexicographic compare is chronological and no date parsing is needed.Verified
Ran the real hook three times against an identical transcript, changing only what was on record:
REATTACH resumeAt=2REPLAYbaselineREPLAYThe third case matters: the discards already sitting in history do not count as offered, so that work gets re-offered rather than staying lost.
While checking this I had a probe that showed cases 1 and 2 behaving identically, which looked like the fix was dead. It turned out the probe emitted
"timestamp": "..."with a space while real transcripts are compact. Harness bug rather than a code bug, but the pattern is whitespace tolerant now, since a miss there silently degrades to a full replay.Cost
Largest transcript in this repo, 39k lines and 126MB:
lastFireAt0.6mslineIndexAfter25msThe rule it replaces was trading correctness for roughly half a second.
Windows
CI already matrixes ubuntu and windows, so the new tests run on both. Checked specifically:
capture.jsonlis gitignored via.metrics/, so it is only ever written locally with"\n"and git never checks it out, meaning autocrlf cannot convert itJSON.parsetreats a trailing\ras whitespace and the timestamp regex matches mid linepath.join, and nochild_processspawn is added, so the Windows: every child_process spawn missing windowsHide — daemon flashes a visible console/terminal window constantly #149windowsHideclass does not applyRejected alternatives
Move the marker to durable storage. Looks like the obvious fix and breaks something else.
soleMarkerUnderRootresolves which session it is in by finding exactly one marker lying around, which only holds because temp keeps getting swept (there was exactly one present across 121 sessions when checked). Permanent markers mean every past session leaves one behind, the count is never one again, and manual/capture-noteson Cursor and Codex stops working. It also brings its own garbage collection problem.Timestamp gap detection. A gap does not mean the work was captured, it means somebody went to lunch.
hooks/trigger.mjsrules this class of signal out by design.Known gap
Files a fire ranked past
MAX_CAPTURE_FILESwere read before that fire, so a sweep still forgets them. Much smaller than dropping the whole span, worth revisiting only with evidence it bites.Test plan
npx vitest run, 861 tests across 57 files passnpx tsc --noEmitclean🤖 Generated with Claude Code