feat(cli): add directa monitor for agent streaming tools - #37
Merged
Merged
Conversation
…start When macOS memory pressure jetsam-kills the daemon, launchd relaunches it and boot-restore used to SIGTERM every still-running dev server and respawn it, bouncing work that never actually stopped. In agent mode the child servers are independent launchd jobs that outlive the daemon's own death, so recovery now re-attaches to a survivor whose launchd job pid still matches a persisted server instead of killing it: it arms a fresh NOTE_EXIT watch feeding the same exit path a spawn uses, re-monitors health, and resumes log tailing at the end of the existing spool. A pid whose kernel start time contradicts the recorded one, a pid with no matching launchd job, or any orphan outside agent mode is group-killed as before.
The jetsam finding already tells the user memory pressure killed the daemon rather than a crash. It now also reports how many daemon restarts landed in the last 24 hours, read from the on-disk event log and collapsed so one restart that touched several servers counts once, so a burst of pressure reads as the recurring problem it is.
The menu bar app was a login item with no KeepAlive at the background-app jetsam band, so when macOS idle-culled the windowless menu extra or a memory pressure pass killed it, the app vanished and stayed gone for the rest of the session. It now ships its own KeepAlive LaunchAgent, the same mechanism the daemon uses, so a kill relaunches it within a second and it runs in the lower jetsam band where the system rarely chooses it. A deliberate Quit still stays quit. Start at login is repointed to this agent and migrates existing installs off the old login item.
RecoverAtStartupTests set runningAsAgent to true to exercise adoption, which also armed the leftover-job reap. That reap called LaunchdJobs.reapStaleChildJobs directly instead of going through the injected childJobsProvider fake, so it shelled out to the real launchctl list and booted out every real dev.quantizor.directa.job.* label a test did not itself supervise, including the user's own running dev servers. Fold the job listing and the bootout action into one AgentJobs value alongside the agent-mode flag, so a Router can never enable one without the other. nil means "not the agent"; production builds the live seam only when LaunchdJobLauncher.runningAsAgent is true. Recovery's adoption match and leftover-job reap both read through the same value now, and every RecoverAtStartupTests case that exercises agent mode injects a fake that records bootouts instead of shelling out.
armExitWatch armed EVFILT_PROC with NOTE_EXIT alone, so the kernel left kevent.data at 0 on every exit; waitForExit then decoded every launchd-run server's death as exited(code: 0), whatever the process actually did (every event in the daemon's own live event log reads code=0, though server logs show real exit statuses and signals). Arm NOTE_EXIT | NOTE_EXITSTATUS instead, so data carries the real wait(2) status. The kqueue man page calls NOTE_EXITSTATUS valid only on child processes, but a launchd job is never this process's child; LaunchdJobLauncherTests now pins the real status arriving anyway on this Darwin version, for both a signaled and a nonzero-exit job.
NDJSONBuffer.feed rescanned the whole accumulated buffer for a newline on every call and shifted the remainder once per line found, both O(n) per call. An unbounded single-line response (a `directa logs` query with no tail bound) arrives over the wire in ~8 KB reads, so framing it cost O(n^2) in the line's length: a 36 MB line took about 11.5 minutes to frame versus 6.85 seconds for the daemon to produce it. Track how many leading bytes have already been scanned with no newline found, and resume from there on the next feed instead of rescanning from the start; drop every consumed line with one removeSubrange instead of one per line. A 36 MB single line now frames in well under a second, and the 100k-small-lines case shows no regression.
Watching a launchd-run dev server's exit parked a dedicated thread of Swift's cooperative pool for the server's whole life; with roughly hw.ncpu or more servers watched at once, the pool has nothing left to run other work, and the daemon's own launchd jetsam thread limit (32) made a thread-per-server design a dead end regardless. A single process-wide kqueue now services every watched pid through one dedicated thread, so thread count stays constant no matter how many servers are watched. NOTE_EXITSTATUS also turned out to be gated on whether this process may signal the target (same user or root), not "child only" as the kqueue man page claims: arming it against a process owned by another user (a dev server whose start command runs via sudo) failed with EACCES and tore the server down as spawnFailed. The watcher now falls back to NOTE_EXIT alone on EACCES and reports that exit as status-unknown instead of losing the watch.
docs/macos-lifecycle.md and several comments still described exit tracking as a per-launch kqueue NOTE_EXIT watch and called exit forensics for a non-child unknowable. Both describe ExitWatcher now: NOTE_EXIT | NOTE_EXITSTATUS through one shared, process-wide kqueue, which delivers the real wait(2) status for any process this daemon may signal and falls back to status-unknown on EACCES. AGENTS.md's logging note said log show cannot see past behavior; error-level entries do persist, only info and debug do not. Also notes scripts/smoke.sh is a zsh script to run directly, and maps Supervisor/ExitWatcher.swift. Several test comments narrated what a prior implementation used to do instead of describing the current one; reworded to state the present behavior and its justification directly.
LogQuery.run read and parsed every file in a server's log family (current.log plus rotations, up to 10 MB each) before trimming to the requested tail, so `directa logs <name> --tail 50` cost as much as an unbounded query. A tail with no grep and no since now reads backward in doubling windows from the newest file, falling back to older files only once the newest one runs dry, which reproduces the exact same records without materializing bytes the caller never asked to see.
… exits readCLIVersion and runCLI waited for the CLI to exit before reading its output pipe, and clearQuarantine never read its pipes at all. A child that writes more than one pipe buffer before exiting blocks on that write since nothing is draining it, which blocks the app waiting for an exit the child can now never reach. All three now route through LaunchdAdmin.shell, which already drains its pipe on a background thread while waiting.
CheckoutIdentity's git() helper waited for git to exit before reading its stdout and stderr pipes. A git that fills either buffer before exiting, a long "detached HEAD" or "unsafe repository" warning on stderr counts, blocks on that write, which then blocks worktree and sibling-port detection waiting for an exit that write can no longer reach. Both pipes now drain on background threads started before the wait, matching the fix already in LaunchdAdmin.shell.
… the message A hint is the literal remediation command, prefixed run:. Three daemon refusals carried prose in the hint instead. - writeConfig stale baseline: no command fixes it, so the hint is omitted and the message ends "; reload it and re-apply your edit". - acquire on a held resource: hint is now "run: ps -p <pid>"; the message ends "; wait for it to finish". - start of a server whose resource is locked: hint is now "run: ps -p <pid>"; the message ends "; it is released when the holder finishes". No golden asserted these strings (the schema goldens cover shape, not text). New assertions pin each hint and message in ResourceLockTests and TrustAndInputValidationTests.
…ead of a bare -1 LaunchdAdmin.shell mapped a timeout to status -1 with empty output, so a hung launchctl bootstrap read "launchctl bootstrap failed (-1): ". The output now reads "timed out after N seconds", followed by "; output so far: ..." when the child wrote anything first. capturedPath relied on the empty output to fall back to the PATH floor, so it now reads shellOutcome and treats a timeout as no answer. Every other caller checks the status before reading the output (the already-bootstrapped check runs on a nonzero status only and never matched an empty string), so their behavior is unchanged.
A helper command (launchctl, lsof, ps, git, log show) that wrote without end grew the daemon's memory until its deadline, and a writer fast enough to keep the pipe full never reached the deadline check at all. HelperCommand now stops reading at 4 MiB, kills the command's whole process group exactly as at a deadline, and answers a new outputLimitExceeded outcome carrying the first 4 MiB. Each place that reads a helper's outcome treats the new case like a timeout, with two differences: the (status, output) form says "output exceeded N bytes" without the partial output, which is the whole cap and too large for an error message, and the launchctl list reader reports the listing as unavailable.
…rint times out When launchd's own exit record was needed for a finished job, the supervisor asked launchctl print up to ten times, and a print that hit its deadline read the same as "no record yet" and was retried, so one exit could hold a shared launchctl lane thread for about 100 seconds. A print that timed out or wrote past the output cap now ends the retries at once and the exit stays status-unknown, as it already did when every attempt failed. A print that answered without the exit record yet is still retried. LaunchdJobs gains printOutcome, which tells a print that never answered from one that answered with no job.
isAgentLoaded treated any nonzero status as "not loaded", and a print killed at its deadline reads as status -1, so a slow launchd made a present agent look absent. Its callers act on "not loaded" by registering the agent again (the app's launch-time recovery) or by replacing the helper (the DMG replace and re-register paths), which is wrong while the job still exists. A print that timed out or wrote past the output cap now reads as loaded; a nonzero exit or a launchctl that cannot start still reads as not loaded. The waits that poll for the agent to unload keep waiting and then report it still loaded, which stops the helper replacement: the cautious outcome.
A server the daemon drained on shutdown used to stay down after the next boot when a live `directa lock` holder owned a resource it declares: the restore was refused with resource-locked and nothing brought it back. It now joins that holder's paused set, reads stopped until the release, and starts then. A server declaring several held resources waits on each in turn, and the same check runs before a resume after another lock's release. A drained row restored only through its boot intent keeps reading stopped when its restore is refused for any other reason (trust, config). Only a row whose run was live when the daemon died is marked crashed.
CI linked every suite into one process and Swift Testing started all of them at the same moment, so on a three-core runner ready work waited seconds for a thread and wall-clock checks failed. make test now caps the width (TEST_PARALLEL_WIDTH), CI runs make test and fails on a toolchain too old to honor the cap, and the narrow-pool sensor runs under the same width. Under an emulated three-core load the uncapped suite reproduced CI's failures while a width of 16 passed every run with no change in run time.
Open
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes for you
An agent can watch one server without filling its context
directa monitor <name>streams a server's output shaped for an agent's own streaming tool (Claude Code's Monitor tool, or Grok Build's monitor tool). It attaches with a start marker (checkout path, phase, pid, last exit), then polls for new lines. Repeats collapse. Stdout and stderr each have their own budget, so a flooding server cannot crowd lifecycle lines or errors out of the stream, and every skip names the exactdirecta logscommand that reads what was held back.The default budget is 120 stdout lines and 30 error lines per minute (
--lines-per-minute,--errors-per-minute, and the per-run--lines-per-arm/--errors-per-armtotals). Re-run the command to reset the per-run totals. The stream keeps going across a daemon restart and never starts the daemon itself. It ends on its own after 29 minutes (under the usual 30-minute tool limit), when the server is unregistered, after a sustained daemon outage, or when the daemon refuses it, and the last line says what happened and how to keep watching.Claude Code and Grok Build sessions now see a line at session start naming this command for the project's servers. Cursor, Antigravity, and a plain
directa contextare unchanged.Logs stay small unless you ask for the history
directa logs <name>now shows the last 200 lines by default. Pass--allfor the old full history, or--tail,--since,--since-mark, or--followas before.--tail(with no--sinceor--grep) reads from the end of the log instead of parsing the whole file.--followstreams through a pipe: lines arrive as they are written, lines stamped in the same millisecond are not dropped or repeated, and a follow that starts before the server has printed anything no longer re-reads the whole log on every poll.--head Nreads the oldest lines from a starting point, for catching up on a stretch in order. A very large daemon reply (a full history of a long-running server) now finishes in seconds instead of minutes.A memory-pressure kill no longer bounces work that was still running
When macOS kills the background daemon to free memory (jetsam: the system reclaiming memory from the largest processes), launchd brings the daemon back. Recovery used to stop every still-running dev server and start it again. Servers that are their own background jobs and are still healthy are now re-attached: the daemon watches their exit, checks health, and keeps tailing the existing log. A process id that no longer names the same process, or a server outside that mode, is still stopped and replaced as before.
directa doctornow includes how many daemon restarts landed in the last 24 hours in that jetsam finding, counted once per restart rather than once per server.A still-running server is re-attached only when directa can actually watch it exit; otherwise it is stopped and started fresh, so nothing is left running unsupervised. After a restart of the Mac, a process id that now belongs to an unrelated app is never signaled: recovery checks the process start time first.
The menu bar app stays resident the same way once Start at login is on. Idle-culling or a memory-pressure kill used to leave it gone for the rest of the session. It now comes back, and a deliberate Quit still stays quit. Start at login stays your choice: upgrading does not turn it on, someone who had it on keeps it (the older login item is only retired once the new one is enabled, so it keeps working while macOS waits for your approval), and turning it off turns both off.
Watching server exits no longer costs one thread per server, so a machine with many servers no longer risks the daemon running out of threads. A start command that runs with elevated privileges (for example
sudo caddy run) starts again instead of being torn down as failed. Restarting a server that prints a lot of output no longer grows the daemon until macOS kills it, and a stop that cannot finish gives up after a bounded wait and notes that in the server's log.When the app is running the daemon, a server that exits immediately now reports a real exit, with the exit code or signal macOS recorded, instead of "never became a session leader", and its own error output reaches
directa logsanddirecta why. A command that cannot run at all (a typo'd path) exits 127 and saysdirecta: cannot run <command>: <reason>, where before it exited 0 silently. A command that exits before directa can watch it is not started again on every later daemon launch. The recorded exit reason is the real one, not a cleancode=0for every death. A server killed by something other than directa (an IDE stop, a forwarded Ctrl-C) shows as stopped when the signal was a graceful shutdown, anddirecta whysays the signal came from outside. A server's log and event history say why directa itself stopped it (restart, a watch naming the file that changed, a resource lock, adown, or the daemon shutting down).directa statusand the menu bar stay quick for a crashed server after a daemon restart, instead of rereading that server's whole log on every poll.Projects, doctor, and unregister
A checkout that disappears for a moment (a drive unmounting, a folder mid-move) is no longer forgotten on the first check. The path has to stay missing for a full sweep interval, the same wait the timer sweep already used.
Forgetting a project now removes its log directory.
directa doctorflags a directory left over from before this, naming its size, anddirecta doctor --fixremoves it. That removal re-checks with the daemon right before deleting and only touches a real folder directa itself created under its own logs folder, never a link, another app's folder, or one a project uses. Doctor never prints a shell deletion command. It no longer recommends deleting a live project's logs while that project's config is briefly unreadable.directa doctor --fixactually forgets a leftover entry whose checkout is gone, and it says plainly when the running daemon is too old to do that.directa unregisterstops a server that is still running before removing it, and it no longer drops the project's approval as a side effect of removing the last ad hoc server. Unregistering a name that was never registered ad hoc, including one that exists only indevservers.json, now fails with a clear not-found error. Forgetting a project that still had a server running no longer writes a second "stopped" event. Unregistering or forgetting a server whose stop hangs, or one caught mid-restart, no longer lets it come back on the next daemon launch, and a hung process is cleaned up then. A stop sent before a server's process has fully launched now stops it once it appears, and a server marked failed for a port problem is really stopped bydirecta stopinstead of left running.Saved daemon state is written so a crash in the middle of a save does not silently drop the update, and a failed save no longer leaves a stray temporary file. An extremely long
DIRECTA_SOCKETpath now refuses to start with a clear error instead of printing that it is listening and then never opening a socket. A request that streams past 1 MiB without a newline is refused. That does not affect thedirectacommand or the menu bar app.Reading git information from a checkout that cannot run a process no longer leaks a pair of threads on every failure, and a checkout that prints a long warning before answering no longer hangs sibling-port rebind or worktree detection. The setup panel no longer hangs when a command it runs prints more than one pipe buffer of output.
Everyday commands
directa hook installwith no--harnessinstalls into every supported agent harness detected on the machine (Claude Code, Cursor, Grok Build, OpenCode, Antigravity), one line each for installed and for skipped. It fails with a clear error, naming what it checked, if none is detected.--harnessstill installs exactly that one.directa up webis shorthand fordirecta up --only web.directa down apistops only that server. Omitting the name still targets the whole project. Passing a name and--onlytogether is a usage error.Every
--timeoutand--acquire-timeoutrejects a non-finite value (inf,nan) or one outside 0 to 86400 seconds, instead of accepting it and letting the daemon clamp it quietly.directa eventsand the dashboard timeline no longer drop the whole list when the daemon reports an event kind this build does not know yet.directa whyand the jetsam finding no longer treat a file path that happens to contain the text "daemon-restart" or "(external)" as one of those events.Messages that used to lead with
ddirectanow say "the daemon", so they no longer read as "directa: ddirecta …". The warningdirecta lockprints for a resource with nopathnames the server and thedevservers.jsonobject to add.directa lock <resource> -- <command>runs<command>in the directory you ran it from, including when you pass--project.A steadier daemon under heavy development
The daemon no longer ties up its own worker threads waiting on
git,launchctl, orlsof. A slow or hung repository used to be able to pin enough threads to push the daemon toward macOS's per-process thread limit; these calls now wait on a small fixed pool, and a stuckgitgives up after 10 seconds. The daemon also keeps a bounded diagnostic record of its own health under~/Library/Logs/directa/daemon(thread and memory use, what it was waiting on, and why the previous run ended), about 70 MB and never more than about 100 MB, never sent anywhere, so the remaining unexplained daemon restarts can be diagnosed.scripts/daemon-deaths.shprints one summary per restart.A dev server that floods its output can no longer fill the disk: directa releases the raw output it has already recorded and keeps only the newest 1 MB, and says so in the server's log when a flood outruns it. Reading a server's logs stays small and fast however large the log grows, including after the system clock jumps backward.
When a dev server exits, a helper process it started in its own session a moment earlier is now stopped with it instead of being left running on its port. A server marked failed because of a port problem is treated as still running by every command that asks, and a sibling worktree's rebound port range no longer overlaps the main checkout's.
directa restartfinishes cleanly across a daemon restart instead of failing with "daemon closed the connection" and inviting a retry that restarts the server twice.directa doctorreads the data folder of the daemon it is talking to, anddoctor --fixremoves a leftover log folder through the daemon, so a project that starts at that moment keeps its logs. Cursor and Antigravity sessions get the directa context block once instead of on every hook firing.Faster status, firmer timeouts, and safer cleanup
directa statusacross many projects stays fast when servers share ports: a checkout's git layout is read from its files instead of by runninggitfor every flagged conflict. The helper commands the daemon runs (launchctl,lsof,ps, andgit) now each have a deadline and are stopped with everything they started when they pass it, so one hung call can no longer stall the daemon. When launchd does not answer at boot, the daemon leaves still-running servers alone instead of restarting them.directa restartagainst a daemon that stopped answering gives up within its stated time. A negative--tailor--head, or a negativetailsent straight to the daemon, is refused instead of crashing it.A sibling worktree moved to another port no longer counts as holding its original port, so the main checkout keeps its own port, and
directa doctorno longer calls a server's own still-running process an unmanaged listener. Forgetting a project works for any spelling of its path, even after its folder is deleted, and a second forget during an automatic one runs only one teardown.directa uninstall --purgeand the background-agent commands refuse to run while aDIRECTA_*location override is set, so they never delete or manage the real install on behalf of a relocated test layout.directa lockhints now start withrun:like every other hint.Clearer next steps, and servers that come back after a lock
directa statusin a folder with no servers points todirecta status --allwhen other projects have some,directa waitthat ends on a crashed or stopped server namesdirecta ensure <name>, anddirecta restartprints why a server fell short the wayupandswitchdo. Error hints from the daemon are now always a command you can paste.A dev server that was running when the daemon restarted now comes back once a
directa lockit depends on is released, instead of staying marked crashed. A server the daemon stopped on its way down reads stopped, not crashed, when it cannot be brought back right away. A slow or hung launchd no longer holds up the daemon for minutes while it reads a finished job's exit, and a helper command that floods its output is stopped at a fixed size instead of growing the daemon's memory.A readable menu bar popover
The menu bar popover now has a solid background, so server rows stay readable whatever sits behind it, and the filter and sort controls no longer show rows scrolling beneath them. The popover follows your system Light or Dark setting instead of the menu bar's wallpaper-driven tint, so Dark Mode on a bright wallpaper gets a dark popover.
The menu bar app no longer keeps using a steady share of a CPU core after a server starts. The breathing dot a starting server shows kept redrawing after the server came up, even with the popover closed; it now stops the moment the server leaves the starting state.
For contributors:
make build,make test, and an opt-in pre-commit hook sweep the empty temp folders the Swift toolchain leaves behind (see CONTRIBUTING.md).make dead-codeis a new unused-code check built on Periphery; it runs entirely on your machine (the script blocks the tool's report of the project's line count to its vendor) and needsbrew install periphery.Migration
directa logs <name>with no bound now returns the last 200 lines. Use--allfor the full history.directa hook installwith no--harnessinstalls every detected harness, not only Claude Code.devservers.jsonof its own gets a not-found message naming both fixes and the command that lists the main checkout's servers. A bare repository with onedevservers.jsonin the parent folder and a worktree per branch now needs--projector a config in each worktree.directa unregisterof a name that was never added withdirecta registernow fails instead of succeeding silently. A running server is stopped before it is unregistered.directa up <name>anddirecta down <name>are new.directa up <name> --only …together is a usage error.--timeoutand--acquire-timeoutmust be finite and between 0 and 86400 seconds.directa doctor's leftover log directory finding now says to rundirecta doctor --fixinstead of printing anrm -rfcommand.Verification
Changesets are included for the user-facing changes above. On the current tree:
make buildpassed,make testpassed (three consecutive clean runs during this pass), andmake dead-codereports no unused code.scripts/smoke.shprintedSMOKE PASSon the final tree, andscripts/test-narrow-pool.shpasses. GitHub Actions on this pull request runsswift buildandmake test, which caps how many tests run at once so the suites finish on a small three-core runner, and fails early if the runner's Swift is older than 6.3, which is needed for that cap.Other changes in this pull request:
fix(app)popover readability and idle CPU,refactorunused code removed,buildthe dead-code check and the macOS 26 minimum,testshared polling helpers, per-run test port blocks so two checkouts can run the suites at once, test processes that answer SIGTERM again, and test gates that always open so a failed check reports a failure instead of hanging the run, flood test servers that exit when the test run that started them dies, a crash fix and self-test for the thread-pool checker, andcirunning the suites throughmake test.