A cross-vendor adversarial review pipeline for AI coding CLIs — with deterministic command gates.
One model writes, a different model reviews, and the loop repeats until the reviewer is clean. AgentPipe declares that flow as a small YAML manifest, runs your existing claude and codex CLIs as black boxes, and gates progress on real exit codes — not vibes.
中文说明见 README.zh-CN.md
Status: alpha / personal tool, made public as-is. Built for one person's workflow (Claude Code + Codex, macOS-first). It is small (~8.6k LOC), tested (100+ tests), and the engine is cross-platform-aware, but it has not been hardened for general use. Issues and PRs welcome — see CONTRIBUTING.md.
The 2026 multi-agent landscape is mostly parallel swarms and kanban boards: run N agents at once, review the diffs by hand. AgentPipe is the other shape — a serial, declarative quality pipeline with three opinions:
- Cross-vendor adversarial review. The author and the reviewer are different vendors on purpose (Claude writes, Codex reviews). Different models have different blind spots, so the reviewer catches what the author can't see in its own work — then the author fixes and the reviewer re-checks, looping until clean.
- Deterministic command gates. A verify gate can be
by: command— your own test/build/lint command, where exit code 0 = goal met. No LLM judging whether the tests "probably pass." Run the tests. - Fail-closed convergence. A parse failure, a missing verdict, or hitting the loop's
maxdoes not silently pass. It stops and asks a human. The conservative branch is always the default.
If you want a parallel agent swarm, use Vibe Kanban or Conductor. If you want "Codex keeps tearing apart Claude's work until it's actually clean, gated on my test suite", that's this.
cargo build --release # builds the `agentpipe` CLI
cargo test # ~100 testsTry the full flow with no API calls using the bundled stub binaries — this is exactly what the GIF above records:
# the empty commit gives the repo a resolvable `main` ref — review-mr's base-ref
# check (fail-loud by design) needs at least one commit to diff against.
# `-b main` pins the branch name: demo-task.yaml reviews against `main`, and on
# machines where `init.defaultBranch` is unset git would create `master` instead.
mkdir -p /tmp/ap-demo/repo && (cd /tmp/ap-demo/repo && git init -q -b main && git -c user.name=demo -c user.email=demo@example.com commit --allow-empty -q -m init)
AGENTPIPE_CLAUDE_BIN=$PWD/demo/stub-claude.sh \
AGENTPIPE_CODEX_BIN=$PWD/demo/stub-codex.sh \
AGENTPIPE_HOME=/tmp/ap-demo \
./target/release/agentpipe run demo/demo-task.yamlFor real runs, have the claude and codex CLIs on your PATH (or point AGENTPIPE_CLAUDE_BIN / AGENTPIPE_CODEX_BIN at them) and:
agentpipe run task.yaml # run a pipeline
agentpipe run task.yaml --dry-run # parse + validate + print the plan, spawn nothing
agentpipe validate task.yaml # parse + validate only
agentpipe runs # list past runs
agentpipe view <run-id> # replay a run's events
agentpipe cost <run-id> # per-step cost / turns / duration
agentpipe diff <run-a> <run-b> # diff two runstemplates/mr-review-loop.yaml is the flow in the GIF: paste an MR/PR link, Claude checks out the branch, then Codex reviews the diff and Claude fixes the findings — looping until Codex reports clean (or stopping for a human at max).
version: 1
name: "MR review-fix loop"
target: /abs/path/to/repo
mode: auto
steps:
- id: mr
kind: human
instruction: "Paste the MR/PR link"
expects: "MR link"
- id: checkout
kind: claude
prompt: "Check out the branch for {{mr.artifact}} ..."
- id: review-fix
kind: loop
until: codex-clean # converge when Codex's verdict is `clean`
max: 10 # hit the ceiling -> pause for a human, never pass silently
body:
- id: review
kind: codex
action: review-mr
base: main
- id: fix
kind: claude
prompt: "Fix Codex's findings on {{mr.artifact}} and push.\n\n{{review.findings}}"Data flows between steps by explicit interpolation — {{mr.artifact}}, {{review.findings}}, {{review.verdict}}, {{review.history}} (that step's findings from prior rounds, so the fixer doesn't ping-pong on feedback it already addressed).
Convergence is checked right after the loop's last codex step finishes each round, not at the end of the body — once Codex reports clean, any remaining steps (like fix) are skipped instead of running a pointless empty-findings pass.
| kind | what it does |
|---|---|
| claude | Run Claude once on a prompt (optionally referencing a skill). Runs at the CLI's highest permission (bypassPermissions). |
| codex | Codex as a reviewer: review-doc / review-mr (structured verdict + findings) / ask. Optional vet: true adds a read-only self-challenge pass — findings Codex can't back up with code evidence get dropped before they reach the fixer (verification itself failing falls back to the original findings). |
| acp | Run an external ACP agent by name. The command can be inline or resolved from agents.toml; supports verify gates and on_permission reject or ask. |
| human | A human does something (often in their own Claude Code session); the engine waits for approval + an artifact. Can be pre-seeded with value to run headless. |
| loop | Wrap a body of sub-steps; until: codex-clean converges. Exhausting max pauses at a decision gate (retry / skip / abort) — it never passes silently. Optional allow_residual: minor|major|nit also converges once every residual finding is at or below that severity. |
Any claude / codex step can carry a verify gate (verify: { by: codex | claude | command, ... }) that re-checks whether the step met its goal and retries with feedback if not.
cargo tauri dev # auto-starts the Vite UI + a Tauri windowThree panes: projects (runs grouped by target repo, persisted under ~/.agentpipe/runs/ with cost, pairwise diff), console (live execution / read-only replay / a quick-run prompt bar), and composer (build a task.yaml visually, save it as a template, or fill in the launch conditions and run). The same stub binaries drive the GUI demo.
Note: the GUI's labels are currently Chinese; the CLI is English. English GUI strings are a known gap.
Every run is appended as NDJSON to ~/.agentpipe/runs/<run-id>.ndjson (run-id is allowlisted to [A-Za-z0-9_-] to prevent path traversal). That's the basis for view / cost / diff and the GUI's read-only replay.
| Platform | CLI engine | Desktop GUI |
|---|---|---|
| macOS | primary, tested | primary, tested |
| Linux | should work (unix paths; same code paths as macOS) | should work (untested) |
| Windows | compiles (unix-only libc/process-group code is cfg-gated, with fallbacks); untested |
Tauri is cross-platform; untested |
Verified locally on macOS via cargo test --workspace + cargo clippy --workspace --all-targets -- -D warnings. Reports of breakage on Linux/Windows are welcome.
AgentPipe wires to claude and codex today. The injection points, in order of effort:
- Swap the binaries (no code):
AGENTPIPE_CLAUDE_BIN/AGENTPIPE_CODEX_BINpoint the runners at any compatible executable (this is how the stub demo works). - Add a new agent runner:
crates/engine/src/runner/holdsclaude.rs/codex.rsbehindRunnerBins; add a sibling and aStepKindvariant incrates/engine/src/manifest.rs. See CONTRIBUTING.md.
A general plugin/provider abstraction is intentionally not built yet (YAGNI for a two-vendor tool).
The repo carries its specs and plans — see docs/specs/ and docs/plans/. docs/research/ compares AgentPipe to ccswarm and Conductor.
MIT — see LICENSE.
