Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 22 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,28 @@ as `test-driven-development`, `spec-driven-development`, `pr-gate`, and
does not wrap or duplicate those surfaces.

Provider infrastructure remains below Pi model selection. ForgeFlow logical roles
may resolve to physical models, while endpoint selection, credentials, channel
health, weights, quotas, and transport belong to the provider layer.
choose model capability and thinking effort; policy v3 can also order equivalent quota
sources (for example Business Team before a commercial relay). Endpoint selection,
credentials, channel health, and transport remain provider-layer concerns. LiteLLM still
owns commercial-channel selection after ForgeFlow has selected the commercial physical
model.

ForgeFlow records credential-free local model/supply usage in `~/.pi/forgeflow-usage.jsonl`;
`forgeflow-usage` summarizes it. LiteLLM SpendLogs remain authoritative for relay-channel spend.

## Antigravity delegation

When the installed `pi-subagents` owner is present, ForgeFlow also registers two
external agents backed by the locally authenticated Antigravity CLI (`agy`):

- `antigravity` (alias `agy`) runs Antigravity in read-only `plan` mode;
- `antigravity-writer` (alias `agy-writer`) runs in `accept-edits` mode.

`pi-subagents` still owns child lifecycle, status, timeout, and stop. ForgeFlow only
bridges its stdin handoff to `agy --print`; Antigravity keeps its own authentication,
model selection, quota, and execution runtime. The bridge is local-only and requires
`agy` on `PATH`. It does not turn Antigravity subscription quota into a Pi model
provider.

See [`docs/model-policy.md`](./docs/model-policy.md) for logical-model policy and
[`docs/architecture.md`](./docs/architecture.md) for the ownership boundary.
Expand Down
64 changes: 64 additions & 0 deletions bin/agy-subagent.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
#!/usr/bin/env node

import { spawn } from "node:child_process";

const VALID_MODES = new Set(["plan", "accept-edits"]);
const mode = process.argv[2];

if (!VALID_MODES.has(mode)) {
process.stderr.write("Usage: agy-subagent.mjs <plan|accept-edits>\n");
process.exit(64);
}

let prompt = "";
process.stdin.setEncoding("utf8");
for await (const chunk of process.stdin) {
prompt += chunk;
}

if (!prompt.trim()) {
process.stderr.write("Antigravity subagent prompt must not be empty.\n");
process.exit(64);
}

const args = [
"--mode",
mode,
"--output-format",
"text",
"--dangerously-skip-permissions",
"--print-timeout",
"0s",
`--print=${prompt}`
];

const child = spawn("agy", args, {
env: process.env,
stdio: ["ignore", "pipe", "pipe"]
});

child.stdout.pipe(process.stdout);
child.stderr.pipe(process.stderr);

for (const signal of ["SIGINT", "SIGTERM"]) {
process.on(signal, () => {
if (!child.killed) child.kill(signal);
});
}

child.on("error", (error) => {
process.stderr.write(
`Failed to launch Antigravity CLI (agy): ${error.message}\n`
);
process.exitCode = 127;
});

child.on("exit", (code, signal) => {
if (signal) {
process.stderr.write(`Antigravity CLI terminated by ${signal}.\n`);
process.exitCode = 1;
return;
}

process.exitCode = code ?? 1;
});
70 changes: 70 additions & 0 deletions bin/forgeflow-usage.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
#!/usr/bin/env node
import { existsSync } from "node:fs";
import { readUsageEvents, resolveUsageLogPath } from "../extension/usage.js";

const args = new Set(process.argv.slice(2));
const json = args.has("--json");
const path = resolveUsageLogPath();
if (!existsSync(path)) {
console.error(`ForgeFlow usage log not found: ${path}`);
process.exitCode = 1;
} else {
const events = readUsageEvents(path);
const usageEvents = events.filter((event) => event.event === "model_usage");
const decisions = events.filter((event) => event.event === "route_decision");
const groups = new Map();

for (const event of usageEvents) {
const supply = event.supplyId ?? event.provider ?? "unknown";
const key = `${supply}\u0000${event.model ?? "unknown"}`;
const row = groups.get(key) ?? {
supply,
model: event.model ?? "unknown",
requests: 0,
failures: 0,
tokens: 0,
input: 0,
output: 0,
cacheRead: 0,
reasoning: 0,
cost: 0
};
row.requests += 1;
if (event.status !== "success") row.failures += 1;
row.tokens += Number(event.usage?.totalTokens ?? 0);
row.input += Number(event.usage?.input ?? 0);
row.output += Number(event.usage?.output ?? 0);
row.cacheRead += Number(event.usage?.cacheRead ?? 0);
row.reasoning += Number(event.usage?.reasoning ?? 0);
row.cost += Number(event.usage?.cost?.total ?? 0);
groups.set(key, row);
}

const failovers = new Map();
for (const event of decisions) {
if (event.basis !== "supply-failover") continue;
const reason = event.failoverReason ?? "unknown";
failovers.set(reason, (failovers.get(reason) ?? 0) + 1);
}

const result = {
path,
usage: [...groups.values()].sort((a, b) => b.tokens - a.tokens || b.requests - a.requests),
failovers: [...failovers.entries()].map(([reason, count]) => ({ reason, count })).sort((a, b) => b.count - a.count)
};

if (json) {
console.log(JSON.stringify(result, null, 2));
} else {
console.log(`ForgeFlow usage: ${path}`);
console.log("SUPPLY\tMODEL\tREQUESTS\tFAIL\tTOKENS\tCOST");
for (const row of result.usage) {
console.log(`${row.supply}\t${row.model}\t${row.requests}\t${row.failures}\t${row.tokens}\t${row.cost.toFixed(6)}`);
}
if (result.failovers.length > 0) {
console.log("\nFAILOVER_REASON\tCOUNT");
for (const row of result.failovers) console.log(`${row.reason}\t${row.count}`);
}
}
}

36 changes: 32 additions & 4 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,20 @@ workflow state, provider/channel routing, worktrees, or resume, it belongs in Pi
an existing plugin. If it is an engineering method or delivery procedure, it belongs
in a Skill. Provider/channel routing remains in the provider layer.

### External coding agents

ForgeFlow may register a thin transport adapter when an installed coding agent cannot
consume the `pi-subagents` stdin contract directly. The Antigravity integration is
one example: `pi-subagents` remains the orchestration owner, while a small bridge
converts the assembled stdin prompt into one `agy --print=<prompt>` argument.

The bridge owns no sessions, retries, workflow state, model routing, credentials, or
quota. `agy` remains the authority for Antigravity authentication, model entitlement,
and subscription usage. Two runtime agents are exposed when the `pi-subagents`
registration owner is present: `antigravity`/`agy` for plan-mode analysis and
`antigravity-writer`/`agy-writer` for explicit workspace mutation. Generic external
CLI runners are local-only under the current `pi-subagents` contract.

## Model policy boundary

ForgeFlow registers stable logical roles as Pi virtual models. Their physical model
Expand All @@ -42,9 +56,23 @@ ambient-extension discovery, so this keeps the `forgeflow/*` roles available in
foreground, detached, nested, and recovery child sessions without hard-coding an
installation path in operator profile settings.

New user/direct requests resolve the current role mapping. Continuation/retry
requests stay on the physical model already handling the turn to preserve cache and
reasoning-signature continuity.
New user/direct requests resolve the current role mapping. Policy v3 separates model
capability from supply: a role chooses a logical model/effort, then an ordered supply
group chooses the physical Pi model that pays for it. The current GPT-6.1 Sol supply
order is Business Team (`openai-codex`) before the commercial relay (`litellm`).
LiteLLM then chooses only among channels for the already-selected commercial physical
model.

Continuations remain sticky. Retries remain sticky unless the failed response is
classified as a narrow supply failure (quota/rate/capacity/upstream/transport/auth/model
availability); only then may v3 move to the next source in the same supply group. Context,
request, policy, and tool/schema failures do not consume the next paid source.

The router records its decision in Pi's native virtual-model state. ForgeFlow additionally
maintains a credential-free local usage JSONL projection for subscription/native-provider
traffic, while LiteLLM SpendLogs remain authoritative for commercial relay/channel spend.
Explicit task classes use `[[forgeflow:task=<class>]]`; ForgeFlow deliberately does not add
a hidden LLM classifier or per-turn semantic router.

Project-local `.pi/forgeflow-models.json` is considered only when Pi reports the
project trusted. User-level policy under `~/.pi/forgeflow-models.json` remains
Expand All @@ -55,7 +83,7 @@ Provider gateways such as LiteLLM sit below this boundary: ForgeFlow chooses a
logical role and Pi resolves its physical model; the provider plane chooses the
endpoint/channel/key for that already-selected physical model.

See `docs/model-policy.md` for the v1 schema and role mapping.
See `docs/model-policy.md` for policy schemas, supply priority, failure classification, and usage projection.

## Invariant preflight

Expand Down
Loading
Loading