SAAGE is a deterministic, composable agentic workflow engine. Control flow (loops, retries, polling, exit conditions) is owned by code, not by an LLM's judgment, while individual steps still use LLMs to do the work. It's a graph engine: workflows are hydrated into a graph of nodes over a shared store.
Built on PocketFlow (graph + shared-store), plus a lightweight first-class harness (file CRUD + exec + git tools) that LLM steps drive through a native, provider-agnostic agent loop. Workflows are authored in YAML and hydrated into runnable flows. Skills are Claude-style markdown directories, imported as-is.
Composing skills into agent files inside existing harnesses (Claude Code, Gemini, Codex, Copilot, Windsurf) works but is non-deterministic: the harness, not your spec, decides control flow — and it decides badly (e.g. a poll step launched in the background that never returns). The same workflow run twice gives different results; swapping models changes behavior entirely. This engine makes the LLM choose only content, never control flow.
Requires Python ≥ 3.10. Clone, run the setup script, done:
git clone <this-repo> && cd saage
./setup.sh # Linux / macOS / WSL2
source .venv/bin/activategit clone <this-repo>; cd saage
setup_windows.bat # native Windows
.venv\Scripts\activateThe script creates .venv/, installs everything (CLI, web UI, test tools;
editable, so source edits take effect immediately), and finishes with a
saage doctor environment check. It's idempotent — re-run it after a git pull. Prefer doing it by hand? It's just python -m venv .venv +
pip install -e ".[dev,server,mcp]" (or uv venv + uv pip install, which
setup.sh uses automatically when uv is installed).
Platforms: Linux, macOS, and Windows — both WSL2 and native. On native
Windows you also need Git for Windows:
flow commands are POSIX sh everywhere, and its bundled bash.exe is what
runs them (found automatically next to git.exe, never System32\bash.exe;
set SAAGE_SHELL to override). Don't install an rsync port — saage
deliberately never uses rsync on Windows; remote handoffs go over
tar-into-ssh and work natively.
saage setup # one-time: pick a default provider + model, paste an API key
saage serve # web UI at http://127.0.0.1:8321 — browse and launch flowssaage setup is interactive (aws-configure style): it validates the key with a
cheap live call, saves the defaults to ~/.saage/config.yaml, and the key to
~/.saage/credentials.toml (chmod 600). Everything — CLI runs, the web UI, its
natural-language launcher, the MCP server — uses those defaults from then on.
The wizard ends by offering to wire up your coding agents (see below).
Or drive it from the command line:
saage run flows/story_writer/flow.yaml # a live run with your defaults
saage run flows/guessing_game/flow.yaml --set target=0.3 # override a flow knob
saage run flows/story_writer/flow.yaml --provider anthropic --model claude-opus-4-8
# different provider for one run
pytest -q # full test suite: offline, no API key neededThe example flows don't pin a provider, so they all run with whatever you chose
in setup. export OPENROUTER_API_KEY=... (etc.) always wins over the saved
key, and --provider/--model beats everything for a single run. A flow can
pin its own provider: { type: ..., model: ... } block (e.g. one that needs a
specific strong model) — a pin beats your defaults, and the model id must then
match that provider.
While a flow runs, the engine logs each step to stderr as it happens — flow
loading, skills loaded, every node entering/finishing, model calls, tool calls,
and loop iterations — so you see progress instead of a silent wait. At the end it
prints a run summary (steps run, loop outcomes, and which files were written).
Use -v for tool-output detail and the full per-node results, -q to quiet it:
12:00:01 loading flow: flows/story_writer/flow.yaml
12:00:01 provider: openrouter / anthropic/claude-3.5-sonnet
12:00:01 loaded 3 skill(s): add_twist, review, write_scene
12:00:01 workflow ready: 2 top-level step(s)
12:00:01 ▶ scene [agent: write_scene]
⠹ cogitating… (spinner shown during each model call)
12:00:03 ⚙ write_file story.md
12:00:03 ✓ scene → default
...
12:00:09 ↻ draft: iteration 1/3 done — continuing
...
12:00:30 ✓ draft: reached max_iterations (3) — exiting loop
12:00:31 run complete
── run summary ─────────────────────────────────
steps: scene ×3, twist ×3, critique
loop: draft → 3 iteration(s) (max_iterations)
files: review.md, story.md
────────────────────────────────────────────────
(Logging is configured by the CLI. As a library, saage never installs log
handlers — your app controls logging via the standard logging module.)
Run the flow job manager and web UI locally, with a natural-language launcher and live job monitoring. The server requires a POSIX OS (Linux/macOS): job control uses process groups and POSIX signals, which are unavailable on Windows.
saage serve # from the repo root: ./flows is picked up automatically
# Open http://127.0.0.1:8321(setup.sh installs the server; on a hand-rolled install without the
[server] extra, add it with pip install -e ".[server]".)
The provider and API key come from saage setup — jobs and the
natural-language launcher use your saved defaults; there is nothing
LLM-related to configure on the server. With no config file, saage serve
auto-discovers a flows/ directory under the current directory, and
--flow-path DIR (repeatable) adds any directory without a config. To pin
flow directories or the bind address persistently, write ~/.saage/server.yaml:
cat > ~/.saage/server.yaml << 'EOF'
flow_paths:
- ./flows # search these dirs for */flow.yaml
host: 127.0.0.1
port: 8321
EOF
saage serve
# Open http://127.0.0.1:8321Home page: Lists all flows in flow_paths and provides two ways to launch:
- Knob form — drop-down to select a flow, form fields for numeric/text parameters.
- Natural language — "Run the story writer with 5 iterations" → the LLM parses it into flow name + knob values; you confirm before launching.
Job detail page: Live DAG visualization (updated as the run progresses) + streaming logs + a cancel button.
History page: All past runs, newest first.
Everything the UI does is a plain JSON API — curl examples for every endpoint in docs/server_api.md.
saage doubles as an automation layer for your coding agent: the agent designs and authors flows in conversation with you, then launches and monitors them as native tool calls — and the engine guarantees the resulting automation runs deterministically, on a schedule or unattended, long after the chat ends.
One command wires everything — run the setup wizard and say y at the
"wire up coding agents?" step:
$ saage setup
...
wire up coding agents (flow skills + the `saage mcp` server)? (y/N) y
1) Claude Code detected skills + MCP (claude CLI / ~/.claude.json)
2) Cursor detected ~/.cursor/mcp.json
3) Codex - ~/.codex/config.toml
4) Windsurf - ~/.codeium/windsurf/mcp_config.json
5) Gemini CLI - ~/.gemini/settings.json
configure which? (numbers/names, 'all', 'none') [detected: claude-code, cursor]:
That installs two surfaces:
saage mcp— an MCP server (stdio) over the same job manager as the web UI:list_flows,launch_flow,wait_for_job,job_status,job_logs,cancel_job,validate_flow. The agent launches and monitors flows as native tool calls — and the server steers it away from token-burning poll loops (one blockingwait_for_job, only after asking you).- Two skills, installed to
~/.claude/skills/so they work in every project:designing-saage-flows(a guided interview that turns "help me automate X" into a concrete flow design) andbuilding-saage-flows(the author → validate → offline-test → run loop). Non-Claude agents get the same content viaAGENTS.md.
Already-open agent sessions must restart to pick up new skills/servers.
The intended loop, in your agent's chat:
- Describe the automation — "I want a weekly digest of the subreddits I
care about" or "walk me through automating my model-training retries".
The vague form triggers
designing-saage-flows: the agent interviews you (goal, loop shape, what's deterministic, knobs, bounds) and plays back a design. - Let it build —
building-saage-flowstakes over: it writesflows/<name>/flow.yaml+ skill directories, hydrate-checks them withvalidate_flow(free, no tokens), and writes an offline integration test before anything runs live. Ask for it directly with "build me a flow that…" when you already know what you want. - Run it as tool calls — "launch it" →
launch_flowreturns a job id immediately; the agent asks whether you want to wait (one blockingwait_for_job— zero tokens while blocked) or check back later withjob_status/job_logs. - Keep it — the flow is a directory in your repo: versioned, testable
offline, runnable without any agent via
saage runor cron, and visible in the web UI. The agent built the automation; the engine owns its determinism from then on.
Details, tool reference, and manual client configs: docs/agents.md.
flow_paths(list) — directories whose immediate subdirectories are scanned for<flow_name>/flow.yaml(one level deep, not recursive). Paths are relative to cwd or absolute.parser_provider(dict, optional) — advanced override: use a different LLM for natural-language parsing than yoursaage setupdefaults (same shape asprovider:in flow.yaml, e.g.{ type: openrouter, model: "openai/gpt-4o-mini" }for a cheaper parser model). Normally leave it unset — the parser uses the setup defaults; with neither, the NL launcher is disabled (503).host(str, default127.0.0.1) — bind address.port(int, default8321) — bind port.
Jobs are run as detached subprocesses of saage run, so they do not block the server.
Each job has its own checkpoint and ledger at ~/.saage/runs/<job_id>/, mirroring the
CLI's local run directories. The UI polls ledger events to render live DAG state and
streams logs via Server-Sent Events (SSE). Job cancellation sends SIGTERM to the subprocess.
Every saage run records a checkpoint under ~/.saage/runs/<run_id>/ after each
step (and each loop iteration). If the run is killed — Ctrl-C, a dead battery, an
ssh drop — pick it up where it left off:
saage runs # list runs: id, status, position, flow
saage resume # resume the most recent unfinished run
saage resume <id|prefix> # resume a specific run
saage resume --force <id> # resume even if the flow.yaml/skills changedsaage run always starts a fresh run. Resume granularity is one iteration of the
outermost loop: a 12-iteration hill-climb killed during iteration 10 resumes at
iteration 10, keeping 1–9. The killed iteration is redone from its start, so a
flow's loop body should be safe to re-run (e.g. clean a checkpoint dir, then
train) — the example ML flows already follow this pattern.
A loop nested inside another loop isn't resumed independently: a crash redoes the entire in-progress outer iteration, re-running the inner loop from scratch. The result stays correct, but keep inner loops cheap (or prefer a single loop level) if resumability matters.
The native agent loop is provider-agnostic. The provider for a run is resolved as:
--provider/--model/--base-urlCLI flags (single-run override), else- the flow's own
provider:block (a pin — for flows that need a specific model), else - your
saage setupdefaults (~/.saage/config.yaml).
The API key comes from the provider's env var when set, else from
~/.saage/credentials.toml [keys] (written by saage setup, chmod 600):
provider.type |
backend | env var |
|---|---|---|
anthropic |
Anthropic Messages | ANTHROPIC_API_KEY |
openai |
api.openai.com | OPENAI_API_KEY |
openrouter |
openrouter.ai/api/v1 | OPENROUTER_API_KEY |
nvidia |
integrate.api.nvidia.com/v1 (NIM) | NVIDIA_API_KEY |
local |
any OpenAI-compatible server (Ollama/vLLM/LM Studio/llama.cpp) | none |
provider: { type: anthropic, model: claude-opus-4-8 }
provider: { type: openrouter, model: "anthropic/claude-3.5-sonnet" }
provider: { type: nvidia, model: "nvidia/nemotron-3-ultra-550b-a55b" }
provider: { type: local, model: "llama3.1:8b", base_url: "http://localhost:11434/v1" }Every real provider call is wrapped in bounded exponential backoff with jitter, so a
transient API failure (network blip, 429 rate limit, 5xx) is retried instead of
aborting the whole run. Permanent errors (400 bad request, 401 auth) are not retried —
they propagate immediately. Defaults: 5 attempts, 0.5s base delay doubling up to 30s. Tune
per flow with an optional retry: sub-block:
provider: { type: anthropic, model: claude-opus-4-8, retry: { max_attempts: 8, base_delay: 1.0 } }You can override the flow's provider block without editing the YAML using
--provider, --model, and --base-url. For OpenRouter:
saage run flows/story_writer/flow.yaml \
--provider openrouter \
--model "anthropic/claude-3.5-sonnet" # any model id from openrouter.ai/models(The key comes from saage setup / the env var as usual — export
OPENROUTER_API_KEY=... if you haven't saved one for that provider.)
Same idea for a local model (no key needed):
saage run flows/story_writer/flow.yaml \
--provider local --model "llama3.1:8b" --base-url http://localhost:11434/v1The model id is whatever the backend expects — e.g. gpt-4o for openai,
openai/gpt-4o-mini or meta-llama/llama-3.1-70b-instruct for openrouter,
claude-opus-4-8 for anthropic.
Building a flow yourself (or pointing a coding agent at this repo)? See
AGENTS.mdfor a complete, self-contained guide to the flow/skill schema, step types, the shared store, and conventions.
A flow is a directory containing flow.yaml plus one sub-directory per skill
(skill.md = Claude-style frontmatter + instructions, with optional .py files the agent
runs via run_command). The YAML composes steps with three loop primitives:
retry_loop—action → check; onfailloop back (with the checker's feedback fed in) untilpassormax_iterations. (e.g. implement → run tests)polling_loop—poll → classify; onrunningwait and poll again untilcomplete/failed, with a hardmax_wait_secondscap so it can never hang. (e.g. submit to Slurm, pollsqueue)counting_loop— run a body of steps, looping untilmax_iterationsor anexit_whenpredicate over the shared store. (e.g. optimize untilaccuracy >= target_accuracy)
Plain steps are agent (an LLM skill with the harness tools) and command (a deterministic
shell step). set: { key: regex } captures values from a step's output into the shared store
so exit_when and {{ templates }} can use them.
{{ var }} placeholders are filled from the shared store (deterministically, by the engine —
the model only ever sees finished text) in every step's text: a command: run string and an
agent skill's description and body. So a skill can say Answer this question: {{ question }}
in its instructions. An undefined name renders to "" and logs a warning; wrap a literal brace
in {% raw %}…{% endraw %}.
read_file, write_file, edit_file, delete_file, run_command, and git: git_status,
git_diff, git_add, git_commit, git_branch, git_checkout, git_log.
Security note. The file tools are path-confined to the flow/workspace directory (
..and absolute escapes are rejected).run_commandand the git tools, however, run arbitrary shell with the engine's own privileges andcwdset to the workspace — they are not sandboxed and can read or modify anything the process can (e.g.run_commandcancat ../../etc/passwd). Run untrusted flows inside a container or VM.
As a first line of defense, run_command refuses an obviously destructive command
before running it — recursive force deletes (rm -rf), privilege escalation (sudo),
raw-device writes (dd of=/dev/…, mkfs), fork bombs, pipe-to-shell installs
(curl … | sh), reads of credential files (/etc/shadow, ~/.ssh/…), and more. A
refused command is returned to the agent as an ERROR: (non-fatal — it just can't do
that). The full built-in denylist is DEFAULT_DENY in saage/config.py.
The rules are configurable via an engine config YAML (--config engine.yaml):
command_policy:
use_defaults: true # keep the built-in denylist (default); false = start empty
deny: # extra regex patterns to refuse
- '\bkubectl\s+delete\b'
allow: # whole-command carve-outs (must match the FULL command)
- 'rm -rf \./build'saage run flows/story_writer/flow.yaml --config engine.yamlAn allow is a whole-command carve-out — it overrides a deny only when it matches the
entire command, so it can't wave through a chained extra (rm -rf ./build && rm -rf /
stays blocked). The policy guards the agent's run_command tool, where the LLM picks the
command; deterministic command: steps are author-written and run unfiltered.
See engine.example.yaml. This is defense in depth, not a
sandbox: a denylist over shell=True can always be evaded — the real isolation
boundary is still a container/VM (above).
Each is a runnable demo and a deterministic integration test:
| flow | demonstrates |
|---|---|
story_writer |
counting_loop with a multi-step body, then a terminal review |
fix_failing_test |
retry_loop driving real pytest, with feedback re-injection |
poll_job |
command capture + polling_loop + wall-clock timeout cap |
guessing_game |
multi-agent feedback loop: guesser + judge (higher/lower) homing in on a hidden target via counting_loop + exit_when |
greenfield_ml |
full ML auto-research: baseline classifier + hill-climb on MNIST |
Heavier, application-specific flows live in contrib/ — currently the
le-wm world-model hill-climbs (lewm_hillclimb, lewm_hillclimb_guided).
Develop a flow locally, then hand the entire run off to a remote GPU box over
ssh — package, push, start, disconnect. Observe with saage remote status/logs,
pull artifacts back with fetch, resume killed runs (even onto a fresh box),
and provision Lambda Cloud instances with spawn. Full guide:
docs/remote_handoff.md.
pytest -q # unit + integration, offline & reproducible
SAAGE_SSH_TESTS=1 pytest tests/remote/ # + live ssh handoffs to localhostIntegration tests run the real engine + real local tools/commands/files; only the LLM turns are scripted, so the suite is free, offline, and bit-reproducible. For a real end-to-end smoke test, run a flow live against a provider:
saage setup # once: default provider/model + key
saage run flows/story_writer/flow.yaml(or point it at another provider with --provider/--model, above).
(A live pytest marker is reserved in pyproject.toml for future provider-hitting tests.)
Working and in active use. See docs/plan.md for the original design.
Licensed under the Apache License 2.0.