Unified scheduler + dispatcher for agent routines (Claude Code / Codex), with
global difficulty routing. One scheduler owns dispatch; non-pinned routines
declare only difficulty = "fast" | "normal" | "hard", while explicit provider
smokes use pin = true with harness and model. Spec of record:
fbrain get design-routines-orchestrator.
Today routine configs are split across three registries with two live
schedulers (~/.last-stack/routines/*.md prompts, ~/.codex/automations/*
Codex cron, ~/.claude/scheduled-tasks/). routines unifies dispatch into a
single launchd-supervised daemon reading one on-disk registry.
bun install
bun run install-shim # symlinks `routines` into ~/.local/binRun without the shim via bun run src/cli.ts <command> or bun run routines.
One TOML file per routine at $ROUTINES_HOME/registry/<id>.toml (default
~/.routines). The registry lives on disk, not in LastDB, on purpose: the
scheduler must keep firing β or fail loudly β during a brain outage. Run history
and heartbeats still flow to fbrain.
The versioned 3Γ3 provider/model matrix has one optional fleet owner at
$ROUTINES_HOME/routing-matrix.json. If it is absent, the checked-in version 1
bootstrap matrix is used. providerOrder selects the first available cell.
At each routinesd dispatch pass, an active Situation named
harness-outage-<provider> makes that provider unavailable to matrix-routed
routines; clearing the Situation restores the configured order. This is an
ephemeral choice: registry TOML is never rewritten, and meta.json retains
resolvedBy = "matrix" with the selected cell in matrixResolution. Explicit
and legacy pins do not rotate, so provider smokes keep exercising their named
harness. Migrate a routine by replacing its harness/model lines with one
difficulty line. Legacy route pairs remain pinned during migration. Provider
smokes should make that intent explicit with pin = true. Provider recovery
therefore uses the outage Situation rather than a routine-by-routine TOML sweep.
# ~/.routines/registry/disk-reclaim.toml (filename stem = id)
difficulty = "normal" # fast | normal | hard
tier = "worker" # optional: spine | worker | opportunistic
rrule = "FREQ=HOURLY;INTERVAL=2" # RFC 5545, same dialect as Codex automations
prompt_path = "/Users/you/.last-stack/routines/disk-reclaim.md" # or inline `prompt = "..."`
cwd = "/Users/you/code/edgevector"
status = "active" # active | paused
timeout_min = 30
error_priority = "P0" # optional; new error cards default to P3
heartbeat_slug = "routine-heartbeats" # optional; runs append the fleet heartbeat line
group = "ops" # optional dashboard group override
# fallback = "claude:sonnet,grok:grok-4.5" # optional; overrides fleet default tail
# gate_command = "routines-north-star-rollup-gate" # optional zero-LLM pre-dispatchAn intentional provider smoke is pinned explicitly:
pin = true
harness = "grok"
model = "grok-4.5"Optional gate_command runs before the LLM harness (zero-LLM):
| exit | meaning |
|---|---|
| 10 | proceed to the configured harness |
| 0 | skip harness; honor ROUTINE_RESULT outcome=ok|noop (missing trailer β noop) |
| other | fail the run |
Gate wall-clock is timeout_min (capped at 15 minutes). For
last-stack-north-star-rollup, set
gate_command = "routines-north-star-rollup-gate" after scripts/install-shim.sh
so the dashboard regenerates without an LLM (avoids false
dashboard-script-crash-prior-snapshot-retained from first-yield shells). The
runner also injects LAST_STACK_NORTH_STAR_DASHBOARD_CMD_TIMEOUT=120 for that
routine id when unset (dashboard default is 30s per subprocess).
last-stack-worktree-cleanup also receives a narrow host-level liveness
snapshot before the harness sandbox starts. Routinesd enumerates cwd and
command paths with both lsof and ps, filters them to managed worktree roots,
and passes only those paths through the reclaim helper's existing fixture
variables. Both probes must succeed; otherwise routinesd injects nothing and
the helper retains its conservative soft-degrade behavior. The LLM harness is
never switched to unrestricted sandbox mode.
For cloud-sync-health-fix, set
gate_command = "routines-cloud-sync-health-fix-gate" (see
examples/cloud-sync-health-fix.toml) so the default path is a bounded
zero-LLM observe that always heartbeats staging / upload_queue /
degraded_reasons / backup durability. That stops overnight Γ3 exit-124
(45m harness kill with no metrics) when the agent digresses under load.
The gate only exits 10 (LLM fix lane) on hard signals (auth/quota, staging
β₯50%, last_success=never). Probe timeouts flush
outcome=timeout_partial and exit 0.
For lastdb-local-smoke-test, set
gate_command = "routines-lastdb-local-smoke-gate". The gate resolves only a
prebuilt lastdbd, mirrors the live LaunchAgent's non-home LASTDB_* settings,
and runs the existing real-data CoW smoke under a hard timeout. It never invokes
Cargo; lastdb-canary-build-main owns cold builds. A missing staged candidate is
a classified noop, GREEN is ok, and RED/timeout is error, each with an
authoritative outcome.txt verdict.
tier is the shared capacity-control priority. spine is the essential
shipping loop, worker is normal product work, and opportunistic is the
first class a capacity controller may skip. Omitting it preserves legacy
behavior; controllers must treat an unset tier explicitly rather than guessing
from the routine id.
An optional ${ROUTINES_HOME}/capacity-controller.json policy makes routinesd
admission quota-aware. Each harness snapshot supplies usedPercent, live
resetAt, observedAt, and measured percentPerFire; routinesd computes
(100 - usedPercent) / hoursUntilReset and admits due work in
spine -> worker -> opportunistic order. Spine always runs. Missing, stale,
expired, or malformed quota data fails closed for every non-spine tier. The
unsetTier policy is mandatory and explicit (shed is the safe migration
default). See docs/capacity-controller.example.json.
routines capacity-controller --dry-run --json runs the companion ready-count
idle ladder without mutating its hysteresis/cooldown state. The ladder reads
kanban pickup status --json, leaves state untouched when the board is
unreadable, and tries one rung per tick in this order: unblock, promote, derive,
invent, harvest, entropy. Every tick is appended to
${ROUTINES_HOME}/logs/capacity-controller.jsonl.
Use scripts/migrate-capacity-controller.sh to prove the replacement in dry-run
mode. Its explicit --apply path installs the one routines-owned controller
and only then retires the vacation pilot and standalone idle-ladder launchd
jobs.
When a run fails because the harness itself is out of service (usage limit / credits / capacity / auth β same classifier as harness-outage), routinesd does not stop the fleet. In the same fire it retries the next agent:
- the routine's configured primary (
harness/model) - Claude Sonnet (default)
- Grok
grok-4.5(default)
The registry TOML is not rewritten (ephemeral). Outage state records an
expiry so later fires skip the dead primary until it clears. Ordinary agent
bugs still escalate as before (no chain hop). Disable with ROUTINES_FALLBACK=0;
override the default tail with ROUTINES_FALLBACK_CHAIN or per-routine
fallback = "claude:sonnet,grok:grok-4.5".
Supported rrule keys: FREQ (SECONDLY..YEARLY), INTERVAL, BYDAY, BYHOUR,
BYMINUTE, BYSECOND, BYMONTHDAY, and an optional DTSTART anchor. An
example lives in examples/.
routines list # registered routines
routines status # last run / next fire / harness / model β the single-pane view
routines run <id> # run a routine now (foreground)
routines pause|resume <id> # toggle status
routines route <id> --harness codex --model gpt-5.5
routines route <id> --harness grok --model grok-4.5
routines logs <id> # recent runs (--path, --tail, --json)
routines publish-status # publish slim fleet status + recent run summaries to LastDB
routines deliver-status # publish + stage/admin-approve a fleet-status delivery
routines hygiene # mechanical cleanup (prune runs/memory, daemon check, publish)
routines hygiene --dry-run # report only
routines import # import the legacy schedulers into the registry (dry-run)
routines web # serve the local dashboard (localhost); --port, --host
routines doctor # validate registry + environment (+ configurations project-config)
routines daemon # the scheduler loop (launchd entrypoint); --once, --catchup <s>
routines install-daemon # install + load the launchd user agent
routines install-hygiene # install + load hourly mechanical hygiene launchd agent
routines print-plist # preview the launchd plist
Two complementary layers:
routines hygiene(mechanical, no LLM) β prunes old run dirs under~/.routines/runs(keep last 20 per id or last 7 days), truncatesmemory.mdfiles to the last 100 lines, drops staleerror-escalate/*.json, checks thatcom.edgevector.routinesdis loaded, and runspublish-status. The installed hourly agent also runs--ff-install, so a clean CLI checkout fast-forwards to LastGitmainand kickstartsroutinesdafter updates. Install hourly viaroutines install-hygiene(labelcom.edgevector.routines-hygiene), which writes a stable launcher at~/.routines/daemon/run-hygiene.sh. Shell wrapper for manual runs:scripts/routines-hygiene.sh.routine-fleet-health(agent, hourly) β closes healedroutine-error-*cards, safe registry timeout bumps for chronic 124s, dedupes against error-escalate, files pickup cards only when needed. Canonical prompt:prompts/routine-fleet-health.md(copy into~/.routines/prompts/and/or last-stack as needed).
routines web serves the "coordinate them all from one place" view β the
single-pane routines status table rendered in the browser, with per-routine
actions. It is a single self-contained page (inline CSS + JS, no external
assets) plus a small JSON API on 127.0.0.1 only β localhost, no auth.
routines web # http://127.0.0.1:4778 (ROUTINES_WEB_PORT to override)
routines web --port 8080 # pick a portThe dashboard shows every routine's harness, model, schedule, status, next fire,
and last-run outcome, and wires each action to the same code path as the CLI
(src/actions.ts): run-now (shares the daemon's per-routine single-flight lock),
pause/resume (rewrites status in the registry TOML), and re-route harness/model
(rewrites harness/model in place). Expand a routine to see its recent runs
with exit status and a tail of the captured log. Serve it alongside the daemon β
it reads the same on-disk registry and run logs, so it reflects live scheduler
state.
The JSON API (for scripting): GET /api/routines, GET /api/routines/<id>/runs,
GET /api/routines/<id>/runs/<stamp|latest>, and POST /api/routines/<id>/{run,pause,resume,route}.
routines publish-status writes a slim, admin-deliverable fleet snapshot to the
local LastDB Mini socket. It reuses collectStatus() and listRuns(), declares
the app-owned schemas on first run, and upserts:
routines/RoutineFleetSnapshotkeyfleet-latestroutines/RoutineStatuskey<routine id>routines/RoutineRunSummarykey<routine id>/<run stamp>
The publisher intentionally excludes prompts and full logs. Recent run evidence
is capped (--tail-bytes, default 2048) and common secret-looking assignments
are redacted before write.
routines publish-status --json
routines publish-status --runs 5 --tail-bytes 2048
routines publish-status --dry-run --jsonroutines deliver-status dogfoods LastDB Mini deliver for the routines fleet
slice. It first runs the same publisher as routines publish-status, then
stages a lastdb.slice.v1 delivery with two legs:
routines/RoutineFleetSnapshotkeyfleet-latest- a capped
routines/RoutineStatussample (--max-records, default 20)
Recipient keys are operational inputs, not repository config. Pass them as flags or environment variables; do not commit them:
export ROUTINES_ADMIN_RECIPIENT_PUBKEY=...
export ROUTINES_ADMIN_MESSAGING_PUBLIC_KEY=...
export ROUTINES_ADMIN_MESSAGING_PSEUDONYM=...
routines deliver-status --dry-run --json
routines deliver-status --max-records 20
routines deliver-status --approve --max-records 20Without --approve, the command stages only and prints the pending
delivery_id; with --approve, Mini seals and sends a delivery_slice through
Exemem messaging and prints non-secret evidence (delivery_id, shared count,
message type, schema hashes). Mailbox polling/decryption is intentionally left
to the receiving admin consumer tooling because the send path and read path are
owned by different apps.
routinesd is a launchd user agent (com.edgevector.routinesd, KeepAlive) so
it survives session exit. Each tick it:
- loads the registry (per-file parse errors are reported, healthy routines still schedule),
- computes which active routines are due from their rrule + last-fire state,
- dispatches them as a free-slot pool: when a run finishes, the next due
routine is admitted immediately (the scheduler does not wait for a whole
batch to finish). Constraints are per-routine single-flight (lock file),
an optional global concurrency cap (
--concurrency N; default unlimited /0), and a per-run timeout kill, - enforces the dispatch-time Situation fence: a run whose id matches an
active
fsituationsSituation'sscope_routinesglob is skipped and logged.
Per-run evidence lands at $ROUTINES_HOME/runs/<id>/<ts>/
(meta.json, prompt.txt, stdout.log, stderr.log).
Every dispatched routine prompt requires one final, single-line verdict in
$ROUTINES_RUN_DIR/outcome.txt: ok <detail>, noop <detail>, or
error <detail>. This sink is authoritative over transcript heuristics and is
reported by routines status --json as outcomeSource="sink". Routines still
print their ROUTINE_RESULT outcome=... trailer as a compatibility fallback
for older runners and prompts that have not yet adopted the sink.
routines install-daemon # bootstrap under launchd
routines daemon --once --catchup 60 # single evaluation pass (testing / e2e)routines install-daemon also enables Sentry for the launchd-managed daemon by
setting OBS_SENTRY_DSN=lastsecrets://obs-sentry-dsn-routines,
OBS_SENTRY_ENVIRONMENT=production, and an OBS_SENTRY_RELEASE tag in the
plist. At process startup, routines resolves the locator with lastsecrets get
and initializes the shared Last Stack Bun/TypeScript Sentry helper from
$LAST_STACK_ROOT/lib/observability/sentry.ts (default ~/.last-stack). Harness
and triage children do not inherit that locator: unresolved
lastsecrets:// OBS_SENTRY_DSN / SENTRY_DSN values are omitted from
spawned agent/CLI environments. If the
locator, helper, or @sentry/node dependency is unavailable, Sentry stays off
and the daemon/web process continues. Reported events include uncaught process
errors, daemon tick/dispatch exceptions, non-zero routine run exits tagged by
routine id/harness/model, and dashboard handler exceptions; prompt text and log
bodies are not sent.
This repo merges through LastGit-native change requests, not GitHub PRs
(GitHub is a read-only mirror). Venue: .last-stack/pr-venue; CI gate:
.lastgit/ci.sh (ci-required). LastGit is homed at lastdb:///routines on
the canonical LastDB socket; see fbrain sop-lastgit-native-forge-workflow.
GitHub stays public for clone/browse only. It is not a review or CI venue:
repository Actions are disabled, this checkout contains no GitHub workflows,
and LastGit-to-GitHub mirror sync keeps origin/main aligned after CR merges.
Mirror sync proof: LastGit CRs are expected to appear on the GitHub mirror within the configured sync interval (validated 2026-07-12T23:11:25Z).
bun test # unit + daemon + dashboard integration (47 tests)
bun run typecheck # tsc --noEmit
bun run e2e # full both-adapter dispatch e2e on a throwaway ROUTINES_HOMEThe e2e stubs the leaf claude/codex/grok binaries by setting
ROUTINES_ALLOW_HARNESS_BIN_OVERRIDES=1 plus the relevant ROUTINES_*_BIN
values. It also stubs fsituations/fbrain via ROUTINES_FSITUATIONS_BIN and
ROUTINES_FBRAIN_BIN, so it is hermetic and spends no API credits while
exercising the full dispatch β spawn β log β heartbeat path that routines owns.
routines import reads the two legacy schedulers and generates paused
registry entries, preserving each routine's prompt / rrule / model / cwd /
harness. Importing configuration never activates it: inspect the result, then
use routines resume <id> for each routine that should run. Force re-imports
preserve an existing entry's live active/paused posture (including with
--replace-routing).
~/.codex/automations/*/automation.tomlβ only the ACTIVE crons (PAUSED ones are already off). A strayRRULE:value prefix is stripped; the huge inline prompt is preserved verbatim.- the Claude scheduler's
scheduled-tasks.json(auto-discovered) β only enabled tasks with a cron; one-shotfireAtreminders and disabled tasks are skipped. 5-field cron is converted to the same RRULE dialect. Each task'sSKILL.mdis referenced viaprompt_path. Claude tasks carry no per-task model, so they import at--claude-model(defaultsonnet); re-route withroutines route.
routines import # DRY-RUN: print the diff table, write nothing
routines import --write # generate ~/.routines/registry/<id>.toml files
routines import --json # machine-readable plan (incl. pauseTargets)Codex automation IDs that used last-stack-fkanban-pickup, last-stack-fkanban-watch,
or last-stack-fkanban-validate import as the canonical
last-stack-kanban-pickup, last-stack-kanban-watch, and
last-stack-kanban-validate registry IDs. Existing installs can run the
idempotent one-time filesystem migration before cutover:
routines migrate-kanban-ids # DRY-RUN: registry/state/memory/lock/run moves
routines migrate-kanban-ids --write # apply the moves under $ROUTINES_HOMEWhen both old and new paths exist, the migration keeps the new registry entry,
merges state/run/memory directories where possible, and archives the old path so
only the canonical last-stack-kanban-* registry files can fire.
Dual-scheduler dedup. Many routines are scheduled in both legacy
schedulers under different ids (e.g. Claude program-driver and Codex
last-stack-program-driver are one loop). Importing both would make routines
itself double-fire. import detects these by a normalized name and keeps one per
group (Codex wins by default β flip with --prefer claude, or import both with
--keep-duplicates). Every collapsed group is shown under CROSS-SCHEDULER
DUPLICATES so a human resolves routing before cutover. This is the
papercut-phantom-program-rollup-churn hazard, made visible.
Once the registry looks right, scripts/cutover.sh pauses the legacy schedulers
so routines becomes the sole scheduler:
scripts/cutover.sh # DRY-RUN: print the plan + write the rollback manifest
scripts/cutover.sh --apply # pause Codex ACTIVE->PAUSED + disable Claude tasks
scripts/cutover.sh --restore <manifest.json> # reverse a cutover
β οΈ --applyis a prod cutover of shared scheduling infrastructure (it pauses the kanban-pickup / fleet routines). Run it attended, quit the Claude app first (its scheduler rewritesscheduled-tasks.jsonon every fire, so--applyrefuses while a Claude process is running), and keep the rollback manifest. The manifest (every entry + its prior status) is written even in dry-run, so the rollback list exists before any change.
Core = the scheduler daemon + CLI + both adapters + the local routines web
dashboard (single-pane coordination) + the one-time import + cutover tooling.
Out of scope (separate cards): remote/authenticated access and historical
analytics for the dashboard, and any new routine content.
When a run ends with non-zero exit, timeout, or outcome=error,
routinesd resolves the priority once, preserving a human-set priority on an
existing card ahead of registry/default policy, then uses exactly one route:
- P0: upsert
routine-error-<id>on Kanban and dispatch a one-shot triage agent (same harness/model; 30m cooldown per id). Cards are attributed toroutine:routinesd-error-escalate. - P1/P2/P3: append the failure evidence and a stable cross-routine
signature to the single Brain reference
papercut-routine-non-p0-failures. No Kanban card is created and no immediate triage agent is dispatched, so Brain grooming can consolidate overlapping papercuts into systemic fixes.
New failures default to P3. A truly critical routine must explicitly set
error_priority = "P0" in its registry entry.
Disable with ROUTINES_ERROR_ESCALATE=0. The triage runner id
routine-error-triage is never re-escalated (no loops).