Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,12 +2,17 @@

## [Unreleased]

## 0.24.4 — 2026-09-29

### Added

- **A review of content this deployment already reviewed is replayed from a durable store instead of paying for the reviewer again.** `src/automation/reviewVerdictStore.ts` keys a verdict on `(kind, base, sha256(content digest))`, the digest being every changed path paired with its content hash, and replays the stored outcome on a hit — from both the plain `openswarm review` path (`reviewCommand.ts`) and the publication hook (`publicationReviewHook.ts`). Measured over the recorded history (530 records, 512 carrying a verdict): 68 byte-identical pairs, 15 of them same-mode (8 direct→direct, 7 pr→pr) and 10 of those repeating the same verdict — on the order of 10 of 512 reviews, at a measured ~$0.26 and 93s p50 per reviewer call. Proven across two separate processes: a cold 27.0s / 6-model-call review became a 1.58s replay with **zero** model calls, same verdict. The key carries the base ref and the review mode deliberately, because 15 of the 53 cross-mode repeats in that history flipped their verdict — a CLI review of a tree and a publication review of the same tree are different reviews. What is skipped is the reviewer call; the advisor pass still runs over the reused verdict, so a replay can tighten (`approve`→`revise`/`reject`) exactly as a live run would and the cache cannot fail open past an escalation.
- **The reuse survives a daemon restart.** The verdict lives in its own `publication_reviews` table in the automation database rather than in the in-process `reviewedPublications` Set it replaces, which was empty after every restart — the failure census records 141 `owner_process_exited` and 157 `shutdown_cancelled`, each of which threw away a verdict already paid for. A store that cannot be opened does not suppress a review; the hook reviews instead.
- **`openswarm review --max --harness-only` runs a deterministic quality harness with no LLM cost.** The new harness (`src/verify/qualityHarness.ts`) enumerates every **tracked** source file via the git index and scans each one, then runs the discovered typecheck/lint/test/build commands inside the existing isolated verify sandbox. It is fail-closed by construction: a file over the 512 KiB ceiling, one containing NUL bytes, an unreadable path, a symlinked source, or a path that escapes the repository root each become an explicit error finding rather than a silent skip, and a listing that scanned nothing is itself an error — so the gate cannot report "passed" over a subset it never read. Findings are folded into the audit run as a synthetic `.openswarm/quality-harness` area, which means the markdown report and the exit-code contract carry the evidence on every `--max` run, not only when an LLM area happened to notice something.

### Fixed

- **`openswarm review --max` no longer changes its verdict with `--concurrency`.** The audit re-partitioned the source into reviewer areas until the fan-out saturated the pool (INT-2249), so the same files and the same reviewer produced 2 areas at concurrency 1 and 10 at concurrency 8 — and because `aggregateAuditResults` is worst-wins, the finer split could only turn an `approve` into a `revise`/`reject`. The audit's units of judgement now come from `planAuditAreas`, which is derived from `--max-files-per-area` alone (`--concurrency` is purely how many reviewers run at once), so a given file set gets one verdict. The finer, more parallel fan-out is still available, deterministically, by lowering `--max-files-per-area`; the `--fix` path keeps the pool-filling `balanceAreasToConcurrency`, where more areas only mean more parallel fix workers. The cost gate now prints the real agent-run count (the area count) and points at the granularity flag when the pool is under-filled.
- **A workflow execution can no longer be persisted in a state its DAG forbids.** `saveExecution` now rejects a step marked `completed` without a `completedAt`, `failed` without an `error`, or advanced past a dependency that is still pending, running or failed — states that previously reached disk and were then read back as if the pipeline had progressed. It also gains a `definitionStamp` fence: a definition replaced underneath a live execution makes the next save refuse rather than record a snapshot that never ran against it.
- **The local issue store no longer serves a stale snapshot after a same-size replacement.** The cache stamp was `mtimeMs:size`, which is unchanged when another process replaces the file by atomic rename within the same mtime tick and writes the same number of bytes — the process then kept serving the old contents indefinitely. The stamp now includes the inode.
- **A status transition's event log records the status the write actually saw.** `changeStatus` and `updateIssue` read the current status inside the write transaction instead of from a pre-transaction read, so a concurrent transition can no longer stamp a stale `oldValue` into the audit trail.
Expand All @@ -22,6 +27,8 @@
- **The `dev.ts` close handler always releases its task.** Reporting ran before `onComplete` and `activeTasks.delete`, so a throw while formatting cost/output left the task registered forever; the reporting is now contained and the cleanup runs in `finally`.
- **`memoryBridge` emits one `memory_linked` event per link**, not two (`linkMemory` already emits one).
- **A failed Linear SDK load is retryable.** The rejected init promise stayed cached, so every later `initLinearBridge` call awaited the same rejection and the bridge never recovered within the process.
- **`advisor` role: a second, independent review of the same diff.** Ported from the harness agent patterns. A separate, independently-prompted model is asked one narrow question — *what concrete defects did the reviewer miss?* — and its only permitted effect is to make the gate **more** cautious: it may append concrete findings the reviewer did not report and raise severity, but the merged decision is the max rank of the two (`approve < revise < reject`), so an advisor `approve` can never soften a reviewer `revise`/`reject`, and a raised severity with no concrete finding is discarded. Any error, timeout, empty, or unparseable output fails **open** — `ran: false`, the reviewer's result untouched, no exit code changed. `review`, `review --max` (per area), and the `--fix` re-review loop all run it before dedupe, so its findings are deduped against history like any other. Disabled by default (a second paid call per review), configured under `autonomous.defaultRoles.advisor`; its model must come from a different family than the reviewer's or it is a second identical opinion.
- **Declarative per-role subagent settings: `tools` and `effort`.** Each role (`worker`/`reviewer`/`advisor`/…) may declare `tools.allow` / `tools.deny` and `effort`. An allow-list can only **narrow** the role's default tool set — it can never grant a tool the role would not otherwise have, so a misconfiguration cannot hand a read-only reviewer `bash`. `deny` is applied after `allow` (deny wins) and supports a trailing `*` (`scratch_*`). `effort` selects the native-loop reasoning level for the stage.

## 0.24.3 — 2026-09-28

Expand Down
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,7 +95,8 @@ openswarm review --max --fix # after the audit, dependency-related findings
# a PR is published only after every re-review and trusted
# deterministic repository check passes
# add --in-place to edit the current working tree instead
openswarm review --max --concurrency 8 # widen the fan-out — areas auto-split to fill the pool
openswarm review --max --max-files-per-area 6 # finer fan-out — the partition is this knob
# alone; --concurrency only sets how many run at once
# more --max flags: --no-linear (report only) · --issues-per-area
# (legacy spray) · --issues <id> (set parent) · --fallback
# <adapter> · --out <file> · --dry-run (print the plan)
Expand Down Expand Up @@ -754,7 +755,7 @@ Agent-oriented module map, admission gates, and landmines: **[ARCHITECTURE.md](A
- **BS Detector** — Built-in static analysis engine that detects bad code patterns (empty catch, hardcoded secrets, `as any`, etc.) with pipeline guard integration
- **Autonomous Pipeline** — Cron-driven heartbeat fetches Linear issues, runs Worker/Reviewer pair loops, and updates issue state automatically
- **Worker/Reviewer Pairs** — Multi-iteration code generation with automated review, testing, and documentation stages
- **Codebase Audit (`review --max`)** — fans reviewer subagents out over directory-shaped areas (auto-split to fill `--concurrency`), aggregates a deduped verdict into a markdown report, and synthesizes ≤10 cohesive Linear issues via a PM agent. Both `review` modes consult repository-local prior review logs; resolved/stale findings are not repeated, and byte-identical duplicate follow-ups are suppressed while unresolved issues remain visible. `--fix` groups findings by repository dependency closure, injects the package manager/manifests/verification contract and repo knowledge, runs only independent fix units concurrently in isolated sandboxes, and promotes disjoint in-scope diffs into an audit worktree. It publishes the PR only when every area re-approves and trusted deterministic verification passes; unavailable dependencies/checks fail closed. `--in-place` keeps edits in the current working tree but uses the same gates. Language-agnostic; codex usage-limit aware with automatic `claude` fallback
- **Codebase Audit (`review --max`)** — fans reviewer subagents out over directory-shaped areas (partitioned by `--max-files-per-area`; `--concurrency` only sets how many run at once, so one file set gets one verdict), aggregates a deduped verdict into a markdown report, and synthesizes ≤10 cohesive Linear issues via a PM agent. Both `review` modes consult repository-local prior review logs; resolved/stale findings are not repeated, and byte-identical duplicate follow-ups are suppressed while unresolved issues remain visible. A review of content this deployment already reviewed is replayed from a durable `(kind, base, content digest)` store instead of re-paying for the reviewer, and that store survives a daemon restart. `--fix` groups findings by repository dependency closure, injects the package manager/manifests/verification contract and repo knowledge, runs only independent fix units concurrently in isolated sandboxes, and promotes disjoint in-scope diffs into an audit worktree. It publishes the PR only when every area re-approves and trusted deterministic verification passes; unavailable dependencies/checks fail closed. `--in-place` keeps edits in the current working tree but uses the same gates. Language-agnostic; codex usage-limit aware with automatic `claude` fallback
- **CI / test gate auto-fix (`openswarm fix`)** — runs the project's objective checks (lint / typecheck / build / test), groups the failures by file into areas, fans a fix-worker out over each, then **re-runs the checks and repeats until green** (or the round budget). Deterministic convergence — unlike the review fix pass, it verifies its own work. Multi-language: auto-detects npm scripts, `Cargo.toml` (`cargo check`/`test`, clippy on request), and Python tooling (`ruff`/`mypy`/`pytest`, gated on the repo's config); any other toolchain via a `"checks"` map in `openswarm.json`
- **PR autopilot (`openswarm pr`)** — on-demand surface over the daemon's PRProcessor + `commitAndCreatePR`. `status` reports conflicts / review feedback / CI; `fix` runs one autopilot pass (conflict → comments → CI); `review` re-applies reviewer feedback only, skipping conflict/CI work — recognizes Claude, Codex, and any formal `CHANGES_REQUESTED` review; `review --fresh` instead runs a brand-new code review of the PR's current diff (the same reviewer `openswarm review` uses) and posts the verdict as a PR comment, independent of any existing feedback; `review --all` reviews every open PR in the repo sequentially (combine with `--fresh`) instead of just the current branch's PR or `--number`; `watch` loops until merge-ready; `create` publishes the current feature branch (local fix → commit → push → `gh pr create`). Never merges or enables auto-merge.
- **Decision Engine** — Scope validation, priority-based task selection, and workflow mapping
Expand Down
29 changes: 29 additions & 0 deletions config.example.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -183,6 +183,18 @@ autonomous:
# run ends on completion, the repeated-tool-call guard, or timeoutMs — a turn
# count is not a property of the task. Set a number only to cap cost hard.
# maxTurns: 0
#
# Declarative tool scoping per role. `allow` can only NARROW the role's
# default tools — it can never grant one the role would not otherwise have,
# so an allow-list naming `bash` on a read-only role stays withheld. `deny`
# is applied after `allow` (deny wins), and supports a trailing `*`
# (`scratch_*`). Use this to keep a stage's blast radius explicit instead
# of relying on the stage's defaults.
# tools:
# allow: [read_file, write_file, edit_file, bash]
# deny: [web_fetch, web_search]
# reasoning effort for this role's native-loop adapter (low|medium|high)
# effort: medium
reviewer:
enabled: true
model: gpt-5.6-sol # Correctness gate; light profile lowers this to Terra
Expand All @@ -192,6 +204,23 @@ autonomous:
# diff-scaled default (300s base); a slow model needs more — measured
# 2026-09-17: half the reviews died at 300s. maxTurns 0 = no ceiling.
timeoutMs: 600000
# Second, independently-prompted review of the SAME diff. Its only permitted
# effect is to add findings the reviewer missed and to raise severity — it
# can never soften the reviewer's verdict, and any raised severity without a
# concrete finding is discarded. DISABLED by default: it is a second paid
# call on every review.
#
# Its model must come from a DIFFERENT family than the reviewer's. On the
# reviewer's own model the advisor is a second identical opinion — the same
# weights re-deriving the same blind spots (the trap modelCompat.ts
# documents for `escalate`). The shipped reviewer default is
# deepseek/deepseek-v4-flash; z-ai/glm-5.2 is the family-independent
# alternative measured at 100% detect / 6% false-reject (6s avg) on the
# planted-defect fixtures where the reviewer scored 0% false-reject at 36s.
advisor:
enabled: false
model: z-ai/glm-5.2
timeoutMs: 45000 # 45s single-turn ceiling, as the guard arbiter uses
tester:
enabled: false
model: gpt-5.6-terra # Used only if deterministic verify cannot run
Expand Down
4 changes: 2 additions & 2 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@intrect/openswarm",
"version": "0.24.3",
"version": "0.24.4",
"description": "Autonomous AI agent orchestrator — Claude, GPT, Codex, and local models (Ollama/LMStudio/llama.cpp)",
"license": "MIT",
"type": "module",
Expand Down
129 changes: 129 additions & 0 deletions src/adapters/agenticLoop.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -534,6 +534,135 @@ describe('runAgenticLoop tool exposure options', () => {
expect(toolNames).not.toContain('search_memory');
});

it('exposes only the allow-listed tools, and a deny entry removes what allow kept', async () => {
let toolNames: string[] = [];

await runAgenticLoop({
prompt: 'x',
cwd: process.cwd(),
model: 'test',
webTools: false,
memoryTools: false,
// `search_files` is allow-listed and then denied: allow can never add a
// tool back that deny removed, not even a member of `allow` itself.
toolAllow: ['read_file', 'search_files', 'write_file'],
toolDeny: ['search_files'],
maxTurns: 1,
callApi: async (_messages, tools) => {
toolNames = tools.map((tool) => tool.function.name);
return finalResp('done');
},
});

expect(toolNames).toEqual(['read_file', 'write_file']);
});

it('a deny wildcard withholds the whole scratch family', async () => {
let toolNames: string[] = [];

await runAgenticLoop({
prompt: 'x',
cwd: process.cwd(),
model: 'test',
webTools: false,
memoryTools: false,
scratchpadRunId: 'AGT-0000',
toolDeny: ['scratch_*'],
maxTurns: 1,
callApi: async (_messages, tools) => {
toolNames = tools.map((tool) => tool.function.name);
return finalResp('done');
},
});

expect(toolNames).not.toContain('scratch_write');
expect(toolNames).not.toContain('scratch_read');
expect(toolNames).toContain('read_file');
});

it('an allow-list cannot resurrect bash on a read-only run', async () => {
// The narrowing-only invariant: `allow` intersects with the composition, so
// naming a withheld tool does not expose it. A read-only run (the reviewer's
// shape) keeps bash hidden no matter what the role declared.
let toolNames: string[] = [];

await runAgenticLoop({
prompt: 'x',
cwd: process.cwd(),
model: 'test',
readOnly: true,
webTools: false,
memoryTools: false,
toolAllow: ['bash', 'read_file', 'write_file'],
maxTurns: 1,
callApi: async (_messages, tools) => {
toolNames = tools.map((tool) => tool.function.name);
return finalResp('done');
},
});

expect(toolNames).toEqual(['read_file']);
});

it('refuses a tool the role scope narrowed away, at dispatch as well as in the schema', async () => {
// The schema above and the dispatch set below are the same narrowed array,
// so a provider that emits a withheld name anyway is answered, not obeyed.
let turn = 0;
let deniedResult = '';

await runAgenticLoop({
prompt: 'x',
cwd: process.cwd(),
model: 'test',
webTools: false,
memoryTools: false,
toolAllow: ['read_file'],
maxTurns: 2,
callApi: async (messages, tools) => {
if (turn++ === 0) {
expect(tools.map((tool) => tool.function.name)).toEqual(['read_file']);
return toolCallResp('hidden-write', 'write_file', { path: 'out.txt', content: 'x' });
}
deniedResult = messages.at(-1)?.content ?? '';
return finalResp('done');
},
});

expect(deniedResult).toContain('TOOL_NOT_ALLOWED');
expect(deniedResult).toContain('write_file');
});

it('leaves MCP and coordination tools to their own flags, not the role list', async () => {
// RoleConfig.tools names built-ins (see its doc): a role narrowing to the
// file tools must not silently lose its board access, which the MCP and
// coordination flags govern.
let toolNames: string[] = [];

await runAgenticLoop({
prompt: 'x',
cwd: process.cwd(),
model: 'test',
webTools: false,
memoryTools: false,
toolAllow: ['read_file'],
maxTurns: 1,
mcpTools: [{
type: 'function',
function: { name: 'linear__get_issue', description: '', parameters: { type: 'object' } },
}],
coordinationContext: { repository: '/repo', taskId: 'supervisor', actor: 'orchestrator' },
callApi: async (_messages, tools) => {
toolNames = tools.map((tool) => tool.function.name);
return finalResp('done');
},
});

expect(toolNames).toContain('read_file');
expect(toolNames).toContain('linear__get_issue');
expect(toolNames).toContain('coordination_read');
expect(toolNames).not.toContain('write_file');
});

it('withholds the scratch tools when the run has no scratchpad', async () => {
let toolNames: string[] = [];
await runAgenticLoop({
Expand Down
Loading
Loading