Skip to content

feat(cli): add directa monitor for agent streaming tools - #37

Merged
quantizor merged 205 commits into
mainfrom
feat/survive-daemon-jetsam-restart
Sep 30, 2026
Merged

quantizor merged 205 commits into
mainfrom
feat/survive-daemon-jetsam-restart

Conversation

@quantizor

@quantizor quantizor commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

What changes for you

An agent can watch one server without filling its context

directa monitor <name> streams a server's output shaped for an agent's own streaming tool (Claude Code's Monitor tool, or Grok Build's monitor tool). It attaches with a start marker (checkout path, phase, pid, last exit), then polls for new lines. Repeats collapse. Stdout and stderr each have their own budget, so a flooding server cannot crowd lifecycle lines or errors out of the stream, and every skip names the exact directa logs command that reads what was held back.

The default budget is 120 stdout lines and 30 error lines per minute (--lines-per-minute, --errors-per-minute, and the per-run --lines-per-arm / --errors-per-arm totals). Re-run the command to reset the per-run totals. The stream keeps going across a daemon restart and never starts the daemon itself. It ends on its own after 29 minutes (under the usual 30-minute tool limit), when the server is unregistered, after a sustained daemon outage, or when the daemon refuses it, and the last line says what happened and how to keep watching.

Claude Code and Grok Build sessions now see a line at session start naming this command for the project's servers. Cursor, Antigravity, and a plain directa context are unchanged.

Logs stay small unless you ask for the history

directa logs <name> now shows the last 200 lines by default. Pass --all for the old full history, or --tail, --since, --since-mark, or --follow as before.

--tail (with no --since or --grep) reads from the end of the log instead of parsing the whole file. --follow streams through a pipe: lines arrive as they are written, lines stamped in the same millisecond are not dropped or repeated, and a follow that starts before the server has printed anything no longer re-reads the whole log on every poll. --head N reads the oldest lines from a starting point, for catching up on a stretch in order. A very large daemon reply (a full history of a long-running server) now finishes in seconds instead of minutes.

A memory-pressure kill no longer bounces work that was still running

When macOS kills the background daemon to free memory (jetsam: the system reclaiming memory from the largest processes), launchd brings the daemon back. Recovery used to stop every still-running dev server and start it again. Servers that are their own background jobs and are still healthy are now re-attached: the daemon watches their exit, checks health, and keeps tailing the existing log. A process id that no longer names the same process, or a server outside that mode, is still stopped and replaced as before.

directa doctor now includes how many daemon restarts landed in the last 24 hours in that jetsam finding, counted once per restart rather than once per server.

A still-running server is re-attached only when directa can actually watch it exit; otherwise it is stopped and started fresh, so nothing is left running unsupervised. After a restart of the Mac, a process id that now belongs to an unrelated app is never signaled: recovery checks the process start time first.

The menu bar app stays resident the same way once Start at login is on. Idle-culling or a memory-pressure kill used to leave it gone for the rest of the session. It now comes back, and a deliberate Quit still stays quit. Start at login stays your choice: upgrading does not turn it on, someone who had it on keeps it (the older login item is only retired once the new one is enabled, so it keeps working while macOS waits for your approval), and turning it off turns both off.

Watching server exits no longer costs one thread per server, so a machine with many servers no longer risks the daemon running out of threads. A start command that runs with elevated privileges (for example sudo caddy run) starts again instead of being torn down as failed. Restarting a server that prints a lot of output no longer grows the daemon until macOS kills it, and a stop that cannot finish gives up after a bounded wait and notes that in the server's log.

When the app is running the daemon, a server that exits immediately now reports a real exit, with the exit code or signal macOS recorded, instead of "never became a session leader", and its own error output reaches directa logs and directa why. A command that cannot run at all (a typo'd path) exits 127 and says directa: cannot run <command>: <reason>, where before it exited 0 silently. A command that exits before directa can watch it is not started again on every later daemon launch. The recorded exit reason is the real one, not a clean code=0 for every death. A server killed by something other than directa (an IDE stop, a forwarded Ctrl-C) shows as stopped when the signal was a graceful shutdown, and directa why says the signal came from outside. A server's log and event history say why directa itself stopped it (restart, a watch naming the file that changed, a resource lock, a down, or the daemon shutting down). directa status and the menu bar stay quick for a crashed server after a daemon restart, instead of rereading that server's whole log on every poll.

Projects, doctor, and unregister

A checkout that disappears for a moment (a drive unmounting, a folder mid-move) is no longer forgotten on the first check. The path has to stay missing for a full sweep interval, the same wait the timer sweep already used.

Forgetting a project now removes its log directory. directa doctor flags a directory left over from before this, naming its size, and directa doctor --fix removes it. That removal re-checks with the daemon right before deleting and only touches a real folder directa itself created under its own logs folder, never a link, another app's folder, or one a project uses. Doctor never prints a shell deletion command. It no longer recommends deleting a live project's logs while that project's config is briefly unreadable. directa doctor --fix actually forgets a leftover entry whose checkout is gone, and it says plainly when the running daemon is too old to do that.

directa unregister stops a server that is still running before removing it, and it no longer drops the project's approval as a side effect of removing the last ad hoc server. Unregistering a name that was never registered ad hoc, including one that exists only in devservers.json, now fails with a clear not-found error. Forgetting a project that still had a server running no longer writes a second "stopped" event. Unregistering or forgetting a server whose stop hangs, or one caught mid-restart, no longer lets it come back on the next daemon launch, and a hung process is cleaned up then. A stop sent before a server's process has fully launched now stops it once it appears, and a server marked failed for a port problem is really stopped by directa stop instead of left running.

Saved daemon state is written so a crash in the middle of a save does not silently drop the update, and a failed save no longer leaves a stray temporary file. An extremely long DIRECTA_SOCKET path now refuses to start with a clear error instead of printing that it is listening and then never opening a socket. A request that streams past 1 MiB without a newline is refused. That does not affect the directa command or the menu bar app.

Reading git information from a checkout that cannot run a process no longer leaks a pair of threads on every failure, and a checkout that prints a long warning before answering no longer hangs sibling-port rebind or worktree detection. The setup panel no longer hangs when a command it runs prints more than one pipe buffer of output.

Everyday commands

directa hook install with no --harness installs into every supported agent harness detected on the machine (Claude Code, Cursor, Grok Build, OpenCode, Antigravity), one line each for installed and for skipped. It fails with a clear error, naming what it checked, if none is detected. --harness still installs exactly that one.

directa up web is shorthand for directa up --only web. directa down api stops only that server. Omitting the name still targets the whole project. Passing a name and --only together is a usage error.

Every --timeout and --acquire-timeout rejects a non-finite value (inf, nan) or one outside 0 to 86400 seconds, instead of accepting it and letting the daemon clamp it quietly.

directa events and the dashboard timeline no longer drop the whole list when the daemon reports an event kind this build does not know yet. directa why and the jetsam finding no longer treat a file path that happens to contain the text "daemon-restart" or "(external)" as one of those events.

Messages that used to lead with ddirecta now say "the daemon", so they no longer read as "directa: ddirecta …". The warning directa lock prints for a resource with no path names the server and the devservers.json object to add. directa lock <resource> -- <command> runs <command> in the directory you ran it from, including when you pass --project.

A steadier daemon under heavy development

The daemon no longer ties up its own worker threads waiting on git, launchctl, or lsof. A slow or hung repository used to be able to pin enough threads to push the daemon toward macOS's per-process thread limit; these calls now wait on a small fixed pool, and a stuck git gives up after 10 seconds. The daemon also keeps a bounded diagnostic record of its own health under ~/Library/Logs/directa/daemon (thread and memory use, what it was waiting on, and why the previous run ended), about 70 MB and never more than about 100 MB, never sent anywhere, so the remaining unexplained daemon restarts can be diagnosed. scripts/daemon-deaths.sh prints one summary per restart.

A dev server that floods its output can no longer fill the disk: directa releases the raw output it has already recorded and keeps only the newest 1 MB, and says so in the server's log when a flood outruns it. Reading a server's logs stays small and fast however large the log grows, including after the system clock jumps backward.

When a dev server exits, a helper process it started in its own session a moment earlier is now stopped with it instead of being left running on its port. A server marked failed because of a port problem is treated as still running by every command that asks, and a sibling worktree's rebound port range no longer overlaps the main checkout's.

directa restart finishes cleanly across a daemon restart instead of failing with "daemon closed the connection" and inviting a retry that restarts the server twice. directa doctor reads the data folder of the daemon it is talking to, and doctor --fix removes a leftover log folder through the daemon, so a project that starts at that moment keeps its logs. Cursor and Antigravity sessions get the directa context block once instead of on every hook firing.

Faster status, firmer timeouts, and safer cleanup

directa status across many projects stays fast when servers share ports: a checkout's git layout is read from its files instead of by running git for every flagged conflict. The helper commands the daemon runs (launchctl, lsof, ps, and git) now each have a deadline and are stopped with everything they started when they pass it, so one hung call can no longer stall the daemon. When launchd does not answer at boot, the daemon leaves still-running servers alone instead of restarting them. directa restart against a daemon that stopped answering gives up within its stated time. A negative --tail or --head, or a negative tail sent straight to the daemon, is refused instead of crashing it.

A sibling worktree moved to another port no longer counts as holding its original port, so the main checkout keeps its own port, and directa doctor no longer calls a server's own still-running process an unmanaged listener. Forgetting a project works for any spelling of its path, even after its folder is deleted, and a second forget during an automatic one runs only one teardown. directa uninstall --purge and the background-agent commands refuse to run while a DIRECTA_* location override is set, so they never delete or manage the real install on behalf of a relocated test layout. directa lock hints now start with run: like every other hint.

Clearer next steps, and servers that come back after a lock

directa status in a folder with no servers points to directa status --all when other projects have some, directa wait that ends on a crashed or stopped server names directa ensure <name>, and directa restart prints why a server fell short the way up and switch do. Error hints from the daemon are now always a command you can paste.

A dev server that was running when the daemon restarted now comes back once a directa lock it depends on is released, instead of staying marked crashed. A server the daemon stopped on its way down reads stopped, not crashed, when it cannot be brought back right away. A slow or hung launchd no longer holds up the daemon for minutes while it reads a finished job's exit, and a helper command that floods its output is stopped at a fixed size instead of growing the daemon's memory.

A readable menu bar popover

The menu bar popover now has a solid background, so server rows stay readable whatever sits behind it, and the filter and sort controls no longer show rows scrolling beneath them. The popover follows your system Light or Dark setting instead of the menu bar's wallpaper-driven tint, so Dark Mode on a bright wallpaper gets a dark popover.

The menu bar app no longer keeps using a steady share of a CPU core after a server starts. The breathing dot a starting server shows kept redrawing after the server came up, even with the popover closed; it now stops the moment the server leaves the starting state.

For contributors: make build, make test, and an opt-in pre-commit hook sweep the empty temp folders the Swift toolchain leaves behind (see CONTRIBUTING.md). make dead-code is a new unused-code check built on Periphery; it runs entirely on your machine (the script blocks the tool's report of the project's line count to its vendor) and needs brew install periphery.

Migration

  • The app, CLI, and cask now declare macOS 26 as the minimum, the oldest macOS directa is tested on (continuous integration runs macOS 26, and day-to-day use is on macOS 27).
  • directa logs <name> with no bound now returns the last 200 lines. Use --all for the full history.
  • directa hook install with no --harness installs every detected harness, not only Claude Code.
  • A git worktree nested inside its main checkout (the layout isolated agents use) now acts on its own servers. A worktree with no devservers.json of its own gets a not-found message naming both fixes and the command that lists the main checkout's servers. A bare repository with one devservers.json in the parent folder and a worktree per branch now needs --project or a config in each worktree.
  • directa unregister of a name that was never added with directa register now fails instead of succeeding silently. A running server is stopped before it is unregistered.
  • A graceful signal from outside directa (SIGTERM, SIGINT, or SIGHUP) is stopped, not crashed. A kill or a nonzero exit of the server's own is still crashed.
  • directa up <name> and directa down <name> are new. directa up <name> --only … together is a usage error.
  • --timeout and --acquire-timeout must be finite and between 0 and 86400 seconds.
  • Start at login for the menu bar app stays off unless you had it on. Someone who had the older login item on moves to the new one once it is enabled. Quitting the app on purpose still leaves it quit.
  • directa doctor's leftover log directory finding now says to run directa doctor --fix instead of printing an rm -rf command.

Verification

Changesets are included for the user-facing changes above. On the current tree: make build passed, make test passed (three consecutive clean runs during this pass), and make dead-code reports no unused code. scripts/smoke.sh printed SMOKE PASS on the final tree, and scripts/test-narrow-pool.sh passes. GitHub Actions on this pull request runs swift build and make test, which caps how many tests run at once so the suites finish on a small three-core runner, and fails early if the runner's Swift is older than 6.3, which is needed for that cap.

Other changes in this pull request: fix(app) popover readability and idle CPU, refactor unused code removed, build the dead-code check and the macOS 26 minimum, test shared polling helpers, per-run test port blocks so two checkouts can run the suites at once, test processes that answer SIGTERM again, and test gates that always open so a failed check reports a failure instead of hanging the run, flood test servers that exit when the test run that started them dies, a crash fix and self-test for the thread-pool checker, and ci running the suites through make test.

…start

When macOS memory pressure jetsam-kills the daemon, launchd relaunches it and
boot-restore used to SIGTERM every still-running dev server and respawn it,
bouncing work that never actually stopped. In agent mode the child servers are
independent launchd jobs that outlive the daemon's own death, so recovery now
re-attaches to a survivor whose launchd job pid still matches a persisted
server instead of killing it: it arms a fresh NOTE_EXIT watch feeding the same
exit path a spawn uses, re-monitors health, and resumes log tailing at the end
of the existing spool. A pid whose kernel start time contradicts the recorded
one, a pid with no matching launchd job, or any orphan outside agent mode is
group-killed as before.
The jetsam finding already tells the user memory pressure killed the daemon
rather than a crash. It now also reports how many daemon restarts landed in the
last 24 hours, read from the on-disk event log and collapsed so one restart
that touched several servers counts once, so a burst of pressure reads as the
recurring problem it is.
The menu bar app was a login item with no KeepAlive at the background-app
jetsam band, so when macOS idle-culled the windowless menu extra or a memory
pressure pass killed it, the app vanished and stayed gone for the rest of the
session. It now ships its own KeepAlive LaunchAgent, the same mechanism the
daemon uses, so a kill relaunches it within a second and it runs in the lower
jetsam band where the system rarely chooses it. A deliberate Quit still stays
quit. Start at login is repointed to this agent and migrates existing installs
off the old login item.
RecoverAtStartupTests set runningAsAgent to true to exercise adoption,
which also armed the leftover-job reap. That reap called
LaunchdJobs.reapStaleChildJobs directly instead of going through the
injected childJobsProvider fake, so it shelled out to the real
launchctl list and booted out every real dev.quantizor.directa.job.*
label a test did not itself supervise, including the user's own
running dev servers.

Fold the job listing and the bootout action into one AgentJobs value
alongside the agent-mode flag, so a Router can never enable one
without the other. nil means "not the agent"; production builds the
live seam only when LaunchdJobLauncher.runningAsAgent is true.
Recovery's adoption match and leftover-job reap both read through the
same value now, and every RecoverAtStartupTests case that exercises
agent mode injects a fake that records bootouts instead of shelling
out.
armExitWatch armed EVFILT_PROC with NOTE_EXIT alone, so the kernel left
kevent.data at 0 on every exit; waitForExit then decoded every
launchd-run server's death as exited(code: 0), whatever the process
actually did (every event in the daemon's own live event log reads
code=0, though server logs show real exit statuses and signals).

Arm NOTE_EXIT | NOTE_EXITSTATUS instead, so data carries the real
wait(2) status. The kqueue man page calls NOTE_EXITSTATUS valid only
on child processes, but a launchd job is never this process's child;
LaunchdJobLauncherTests now pins the real status arriving anyway on
this Darwin version, for both a signaled and a nonzero-exit job.
NDJSONBuffer.feed rescanned the whole accumulated buffer for a
newline on every call and shifted the remainder once per line found,
both O(n) per call. An unbounded single-line response (a `directa
logs` query with no tail bound) arrives over the wire in ~8 KB reads,
so framing it cost O(n^2) in the line's length: a 36 MB line took
about 11.5 minutes to frame versus 6.85 seconds for the daemon to
produce it.

Track how many leading bytes have already been scanned with no
newline found, and resume from there on the next feed instead of
rescanning from the start; drop every consumed line with one
removeSubrange instead of one per line. A 36 MB single line now
frames in well under a second, and the 100k-small-lines case shows no
regression.
Watching a launchd-run dev server's exit parked a dedicated thread of
Swift's cooperative pool for the server's whole life; with roughly
hw.ncpu or more servers watched at once, the pool has nothing left to
run other work, and the daemon's own launchd jetsam thread limit (32)
made a thread-per-server design a dead end regardless. A single
process-wide kqueue now services every watched pid through one
dedicated thread, so thread count stays constant no matter how many
servers are watched.

NOTE_EXITSTATUS also turned out to be gated on whether this process
may signal the target (same user or root), not "child only" as the
kqueue man page claims: arming it against a process owned by another
user (a dev server whose start command runs via sudo) failed with
EACCES and tore the server down as spawnFailed. The watcher now falls
back to NOTE_EXIT alone on EACCES and reports that exit as
status-unknown instead of losing the watch.
docs/macos-lifecycle.md and several comments still described exit
tracking as a per-launch kqueue NOTE_EXIT watch and called exit
forensics for a non-child unknowable. Both describe ExitWatcher now:
NOTE_EXIT | NOTE_EXITSTATUS through one shared, process-wide kqueue,
which delivers the real wait(2) status for any process this daemon
may signal and falls back to status-unknown on EACCES.

AGENTS.md's logging note said log show cannot see past behavior;
error-level entries do persist, only info and debug do not. Also notes
scripts/smoke.sh is a zsh script to run directly, and maps
Supervisor/ExitWatcher.swift.

Several test comments narrated what a prior implementation used to do
instead of describing the current one; reworded to state the present
behavior and its justification directly.
LogQuery.run read and parsed every file in a server's log family (current.log
plus rotations, up to 10 MB each) before trimming to the requested tail, so
`directa logs <name> --tail 50` cost as much as an unbounded query. A tail
with no grep and no since now reads backward in doubling windows from the
newest file, falling back to older files only once the newest one runs dry,
which reproduces the exact same records without materializing bytes the
caller never asked to see.
… exits

readCLIVersion and runCLI waited for the CLI to exit before reading its
output pipe, and clearQuarantine never read its pipes at all. A child that
writes more than one pipe buffer before exiting blocks on that write since
nothing is draining it, which blocks the app waiting for an exit the child
can now never reach. All three now route through LaunchdAdmin.shell, which
already drains its pipe on a background thread while waiting.
CheckoutIdentity's git() helper waited for git to exit before reading its
stdout and stderr pipes. A git that fills either buffer before exiting, a
long "detached HEAD" or "unsafe repository" warning on stderr counts, blocks
on that write, which then blocks worktree and sibling-port detection waiting
for an exit that write can no longer reach. Both pipes now drain on
background threads started before the wait, matching the fix already in
LaunchdAdmin.shell.
… the message

A hint is the literal remediation command, prefixed run:. Three daemon
refusals carried prose in the hint instead.

- writeConfig stale baseline: no command fixes it, so the hint is omitted and
  the message ends "; reload it and re-apply your edit".
- acquire on a held resource: hint is now "run: ps -p <pid>"; the message ends
  "; wait for it to finish".
- start of a server whose resource is locked: hint is now "run: ps -p <pid>";
  the message ends "; it is released when the holder finishes".

No golden asserted these strings (the schema goldens cover shape, not text).
New assertions pin each hint and message in ResourceLockTests and
TrustAndInputValidationTests.
…ead of a bare -1

LaunchdAdmin.shell mapped a timeout to status -1 with empty output, so a hung
launchctl bootstrap read "launchctl bootstrap failed (-1): ". The output now
reads "timed out after N seconds", followed by "; output so far: ..." when the
child wrote anything first.

capturedPath relied on the empty output to fall back to the PATH floor, so it
now reads shellOutcome and treats a timeout as no answer. Every other caller
checks the status before reading the output (the already-bootstrapped check
runs on a nonzero status only and never matched an empty string), so their
behavior is unchanged.
A helper command (launchctl, lsof, ps, git, log show) that wrote without end grew the daemon's memory until its deadline, and a writer fast enough to keep the pipe full never reached the deadline check at all. HelperCommand now stops reading at 4 MiB, kills the command's whole process group exactly as at a deadline, and answers a new outputLimitExceeded outcome carrying the first 4 MiB.

Each place that reads a helper's outcome treats the new case like a timeout, with two differences: the (status, output) form says "output exceeded N bytes" without the partial output, which is the whole cap and too large for an error message, and the launchctl list reader reports the listing as unavailable.
…rint times out

When launchd's own exit record was needed for a finished job, the supervisor asked launchctl print up to ten times, and a print that hit its deadline read the same as "no record yet" and was retried, so one exit could hold a shared launchctl lane thread for about 100 seconds. A print that timed out or wrote past the output cap now ends the retries at once and the exit stays status-unknown, as it already did when every attempt failed. A print that answered without the exit record yet is still retried.

LaunchdJobs gains printOutcome, which tells a print that never answered from one that answered with no job.
isAgentLoaded treated any nonzero status as "not loaded", and a print killed at its deadline reads as status -1, so a slow launchd made a present agent look absent. Its callers act on "not loaded" by registering the agent again (the app's launch-time recovery) or by replacing the helper (the DMG replace and re-register paths), which is wrong while the job still exists. A print that timed out or wrote past the output cap now reads as loaded; a nonzero exit or a launchctl that cannot start still reads as not loaded.

The waits that poll for the agent to unload keep waiting and then report it still loaded, which stops the helper replacement: the cautious outcome.
A server the daemon drained on shutdown used to stay down after the next boot when a live `directa lock` holder owned a resource it declares: the restore was refused with resource-locked and nothing brought it back. It now joins that holder's paused set, reads stopped until the release, and starts then. A server declaring several held resources waits on each in turn, and the same check runs before a resume after another lock's release.

A drained row restored only through its boot intent keeps reading stopped when its restore is refused for any other reason (trust, config). Only a row whose run was live when the daemon died is marked crashed.
CI linked every suite into one process and Swift Testing started all of them at the same moment, so on a three-core runner ready work waited seconds for a thread and wall-clock checks failed. make test now caps the width (TEST_PARALLEL_WIDTH), CI runs make test and fails on a toolchain too old to honor the cap, and the narrow-pool sensor runs under the same width. Under an emulated three-core load the uncapped suite reproduced CI's failures while a width of 16 passed every run with no change in run time.
@quantizor
quantizor merged commit e4d22ea into main Sep 30, 2026
2 checks passed
@quantizor
quantizor deleted the feat/survive-daemon-jetsam-restart branch September 30, 2026 22:33
@github-actions github-actions Bot mentioned this pull request Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants