Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,27 @@

## [Unreleased]

### Added

- **`openswarm review --max --harness-only` runs a deterministic quality harness with no LLM cost.** The new harness (`src/verify/qualityHarness.ts`) enumerates every **tracked** source file via the git index and scans each one, then runs the discovered typecheck/lint/test/build commands inside the existing isolated verify sandbox. It is fail-closed by construction: a file over the 512 KiB ceiling, one containing NUL bytes, an unreadable path, a symlinked source, or a path that escapes the repository root each become an explicit error finding rather than a silent skip, and a listing that scanned nothing is itself an error — so the gate cannot report "passed" over a subset it never read. Findings are folded into the audit run as a synthetic `.openswarm/quality-harness` area, which means the markdown report and the exit-code contract carry the evidence on every `--max` run, not only when an LLM area happened to notice something.

### Fixed

- **A workflow execution can no longer be persisted in a state its DAG forbids.** `saveExecution` now rejects a step marked `completed` without a `completedAt`, `failed` without an `error`, or advanced past a dependency that is still pending, running or failed — states that previously reached disk and were then read back as if the pipeline had progressed. It also gains a `definitionStamp` fence: a definition replaced underneath a live execution makes the next save refuse rather than record a snapshot that never ran against it.
- **The local issue store no longer serves a stale snapshot after a same-size replacement.** The cache stamp was `mtimeMs:size`, which is unchanged when another process replaces the file by atomic rename within the same mtime tick and writes the same number of bytes — the process then kept serving the old contents indefinitely. The stamp now includes the inode.
- **A status transition's event log records the status the write actually saw.** `changeStatus` and `updateIssue` read the current status inside the write transaction instead of from a pre-transaction read, so a concurrent transition can no longer stamp a stale `oldValue` into the audit trail.
- **The daily reporter retries a failed project once, and only advances its watermark when every project succeeded (after that retry).** A transient Linear error previously left a project unreported until the next day, while a permanently failing project could hold the window open; the watermark now reflects post-retry reality.
- **A Linear project description keeps its automation summary.** The compact `[Done:n InProgress:n Todo:n]` line was appended before truncation, so a long base description pushed it past Linear's 255-character limit and it was silently cut. The builder now reserves room for the summary first.
- **Repository-registry corruption that could not be quarantined says so.** `loadRepos` reported a failed `renameSync` as if the corrupt file had been preserved, naming a recovery path that does not exist; the quarantine failure is now surfaced as its own error.
- **A failed local Linear mapping can no longer orphan-recreate the remote issue.** If `updateIssue`/`addEvent` throws after Linear's `createIssue` succeeded, the new `linearId` is remembered in-process and re-persisted on the next call (with an `idempotencyKey` on the `linked` event), instead of the retry creating a second Linear issue.
- **Backlog grooming refuses to act on an issue id it was not given.** `validIssueIds` is now mandatory at both parse time and apply time, so a hallucinated or out-of-scope id from the planner is dropped before any mutation path sees it, and an omitted scope means "nothing is in scope" rather than "anything goes".
- **Shell/SQL source under data directories still requires validation evidence.** `isValidationRelevantFile` exempts locale/fixtures/snapshots trees as non-code, but its idea of "code" excluded `sh`, `bash`, `sql`, `swift` and others that a worker can meaningfully break; those now count.
- **Linear project-overview pagination surfaces a broken page walk.** A response with no issues connection, or a missing/repeated `endCursor`, now throws instead of silently returning a truncated issue list that the overview reports as complete.
- **An idempotent `createIssue` rejects a materially different artifact under the same id.** A caller-supplied `id` that already exists is treated as a retry only when title, description, parent and project agree; otherwise the store throws an explicit collision error instead of returning the stored row (which masked a re-planned decomposition) or overwriting it (which would destroy the first plan's artifact).
- **The `dev.ts` close handler always releases its task.** Reporting ran before `onComplete` and `activeTasks.delete`, so a throw while formatting cost/output left the task registered forever; the reporting is now contained and the cleanup runs in `finally`.
- **`memoryBridge` emits one `memory_linked` event per link**, not two (`linkMemory` already emits one).
- **A failed Linear SDK load is retryable.** The rejected init promise stayed cached, so every later `initLinearBridge` call awaited the same rejection and the bridge never recovered within the process.

## 0.24.3 — 2026-09-28

### Added
Expand Down
13 changes: 13 additions & 0 deletions src/__tests__/issueStore.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -129,6 +129,19 @@ describe('SqliteIssueStore', () => {

expect(done?.closedAt).toBeDefined();
});

it('event oldValue reflects the status read inside the write transaction', () => {
const issue = store.createIssue({ projectId: 'p1', title: 'task' });
store.changeStatus(issue.id, 'todo');
store.changeStatus(issue.id, 'in_progress');

// getEvents returns newest-first (created_at DESC, rowid DESC).
const events = store.getEvents(issue.id).filter((e) => e.type === 'status_changed');
expect(events.map((e) => [e.oldValue, e.newValue])).toEqual([
['todo', 'in_progress'],
['backlog', 'todo'],
]);
});
});

describe('listIssues', () => {
Expand Down
17 changes: 17 additions & 0 deletions src/agents/workerValidationEvidence.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,23 @@ describe('missingWorkerValidationIssues', () => {
})).length).toBeGreaterThan(0);
});

it('requires validation for executable formats under locale/i18n dirs', () => {
// sh/swift/sql (and other VALIDATION_RELEVANT executables) must not inherit
// the data-only exemption that applies to json locale strings.
expect(missingWorkerValidationIssues(worker({
filesChanged: ['src/locales/format.sh'],
commands: [],
})).length).toBeGreaterThan(0);
expect(missingWorkerValidationIssues(worker({
filesChanged: ['src/i18n/Localizable.swift'],
commands: [],
})).length).toBeGreaterThan(0);
expect(missingWorkerValidationIssues(worker({
filesChanged: ['src/locales/seed.sql'],
commands: [],
})).length).toBeGreaterThan(0);
});

it('treats a source module named readme.ts as code, not docs', () => {
// README.md is docs; readme.ts is a real module and must hit the gate.
expect(missingWorkerValidationIssues(worker({
Expand Down
9 changes: 7 additions & 2 deletions src/agents/workerValidationEvidence.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,10 @@ const DOC_ONLY_FILE_RE = /(^|\/)(README|CHANGELOG|LICENSE|NOTICE)(\.(md|mdx|txt|
// nothing to build or test on their own; exempt them so a data-only edit does
// not get bounced for "no validation command".
const DATA_ONLY_DIR_RE = /(^|\/)(locales?|i18n|fixtures?|__fixtures__|__snapshots__|snapshots?|__mocks__|mocks?|testdata|test-data)\//i;
// Executable / source formats under data dirs still require validation evidence.
// Broader than TESTER_CODE_FILE_RE: includes sh/swift/sql/etc. that VALIDATION_RELEVANT
// already tracks, but excludes pure data formats (json/yaml/toml).
const EXECUTABLE_SOURCE_FILE_RE = /\.(ts|tsx|mts|cts|js|jsx|mjs|cjs|py|rs|go|java|rb|c|cc|cpp|h|hpp|swift|kt|kts|scala|cs|php|sh|bash|zsh|sql)$/i;
const VALIDATION_COMMAND_RE = /\b(npm\s+(?:test|run\s+(?:test|build|lint|typecheck|check|ci|verify|validate|smoke))|pnpm\s+(?:test|run\s+(?:test|build|lint|typecheck|check|ci|verify|validate|smoke))|yarn\s+(?:test|run\s+(?:test|build|lint|typecheck|check|ci|verify|validate|smoke))|bun\s+(?:test|run\s+(?:test|build|lint|typecheck|check|ci|verify|validate|smoke))|vitest|jest|mocha|pytest|ruff|mypy|pyright|tsc|eslint|oxlint|cargo\s+(?:check|test|clippy|build)|go\s+(?:test|vet|build)|swift\s+test|gradle\s+(?:test|build|check)|mvn\s+(?:test|verify)|make\b|cmake\b|py_compile|compileall|clippy|fmt\s+--check)\b/i;
// Anchored at each segment start: a leading inspection verb means that segment
// ran no validation (e.g. `rg "npm test"` searches for the string, it does not
Expand All @@ -23,8 +27,9 @@ export function isValidationRelevantFile(file: string): boolean {
if (/(^|\/)docs?\//i.test(file)) return false;
if (VALIDATION_RELEVANT_BASENAME_RE.test(file)) return true;
// Data/asset trees (locale, fixtures, snapshots, mocks) are exempt ONLY for
// non-code assets. A real source module under such a dir still needs a check.
if (DATA_ONLY_DIR_RE.test(file) && !TESTER_CODE_FILE_RE.test(file)) return false;
// non-executable assets. Source/executable modules under such a dir (including
// sh/swift/sql and locale i18n modules) still need a validation check.
if (DATA_ONLY_DIR_RE.test(file) && !EXECUTABLE_SOURCE_FILE_RE.test(file)) return false;
return VALIDATION_RELEVANT_FILE_RE.test(file) && !DOC_ONLY_FILE_RE.test(file);
}

Expand Down
30 changes: 25 additions & 5 deletions src/automation/backlogGrooming.coverage.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -91,21 +91,21 @@ describe('parseBacklogGroomingOutput edge branches', () => {
it('drops non-object decision entries', () => {
const result = parseBacklogGroomingOutput(`\`\`\`json
{"decisions": [null, "not-an-object", 42, {"issueId":"id-1","status":"active","reason":"ok"}]}
\`\`\``);
\`\`\``, new Set(['id-1']));
expect(result.success).toBe(true);
expect(result.decisions.map(d => d.issueId)).toEqual(['id-1']);
});

it('drops decision entries missing required fields', () => {
const result = parseBacklogGroomingOutput(`\`\`\`json
{"decisions": [{"identifier":"INT-9"}, {"issueId":"id-1","status":"active"}, {"issueId":"id-2","reason":"no status"}]}
\`\`\``);
\`\`\``, new Set(['id-1', 'id-2']));
expect(result.success).toBe(true);
expect(result.decisions).toEqual([]);
});

it('returns a failure result when the output cannot be parsed as JSON', () => {
const result = parseBacklogGroomingOutput('not json at all, no brace here');
const result = parseBacklogGroomingOutput('not json at all, no brace here', new Set());
expect(result.success).toBe(false);
expect(result.decisions).toEqual([]);
expect(result.error).toBeTruthy();
Expand All @@ -131,7 +131,10 @@ describe('runBacklogGroomingPlanner', () => {
stdout: '```json\n{"decisions":[{"issueId":"id-1","status":"active","reason":"fine"}]}\n```',
stderr: '',
});
const result = await runBacklogGroomingPlanner({ tasks: [baseTask()], projectPath: '/repo' });
const result = await runBacklogGroomingPlanner({
tasks: [baseTask({ issueId: 'id-1' })],
projectPath: '/repo',
});
expect(result.success).toBe(true);
expect(result.decisions.map(d => d.issueId)).toEqual(['id-1']);
});
Expand All @@ -142,11 +145,28 @@ describe('runBacklogGroomingPlanner', () => {
stdout: '```json\n{"decisions":[{"issueId":"id-1","status":"active","reason":"fine"}]}\n```',
stderr: 'warning: partial output',
});
const result = await runBacklogGroomingPlanner({ tasks: [baseTask()], projectPath: '/repo' });
const result = await runBacklogGroomingPlanner({
tasks: [baseTask({ issueId: 'id-1' })],
projectPath: '/repo',
});
expect(result.success).toBe(true);
expect(result.decisions).toHaveLength(1);
});

it('drops out-of-scope decision ids from planner output', async () => {
spawnCli.mockResolvedValueOnce({
exitCode: 0,
stdout: '```json\n{"decisions":[{"issueId":"other","status":"stale","reason":"nope","closeState":"Done"}]}\n```',
stderr: '',
});
const result = await runBacklogGroomingPlanner({
tasks: [baseTask({ issueId: 'id-1' })],
projectPath: '/repo',
});
expect(result.success).toBe(true);
expect(result.decisions).toEqual([]);
});

it('reports stderr as the error when exit code is non-zero and stdout is empty', async () => {
spawnCli.mockResolvedValueOnce({ exitCode: 1, stdout: ' ', stderr: 'adapter blew up' });
const result = await runBacklogGroomingPlanner({ tasks: [baseTask()], projectPath: '/repo' });
Expand Down
20 changes: 17 additions & 3 deletions src/automation/backlogGrooming.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -28,10 +28,11 @@ describe('backlogGrooming (INT-1609)', () => {
"decisions": [
{"issueId":"id-1","identifier":"INT-1","status":"stale","reason":"implemented","evidence":["src/a.ts:10"],"closeState":"Done"},
{"issueId":"id-2","status":"bogus","reason":"bad"},
{"issueId":"id-3","status":"needs_update","reason":"drifted","updatedDescription":"new body"}
{"issueId":"id-3","status":"needs_update","reason":"drifted","updatedDescription":"new body"},
{"issueId":"hallucinated","status":"stale","reason":"out of scope","closeState":"Done"}
]
}
\`\`\``);
\`\`\``, new Set(['id-1', 'id-2', 'id-3']));
expect(result.success).toBe(true);
expect(result.decisions.map(d => d.issueId)).toEqual(['id-1', 'id-3']);
expect(result.decisions[0].closeState).toBe('Done');
Expand Down Expand Up @@ -63,12 +64,25 @@ describe('backlogGrooming (INT-1609)', () => {
{ issueId: 'id-2', status: 'needs_update', reason: 'drifted', evidence: ['src/a.ts:1'], updatedDescription: 'new body' },
{ issueId: 'id-3', status: 'stale', reason: 'implemented', evidence: ['src/b.ts:2'], closeState: 'Done' },
],
}, 'apply');
}, 'apply', new Set(['id-1', 'id-2', 'id-3']));
expect(applied).toEqual({ commented: 2, failedComments: 0, updatedDescriptions: 1, moved: 1, movedIssueIds: ['id-3'], skippedUnknown: 0 });
expect(src.updateDescription).toHaveBeenCalledWith('id-2', 'new body');
expect(src.updateState).toHaveBeenCalledWith('id-3', 'Done');
});

it('refuses apply mutations when no scope Set is supplied', async () => {
const src = source();
const applied = await applyBacklogGrooming(src, {
success: true,
decisions: [
{ issueId: 'id-1', status: 'stale', reason: 'implemented', evidence: ['src/a.ts:1'], closeState: 'Done' },
],
}, 'apply');
expect(applied.moved).toBe(0);
expect(applied.skippedUnknown).toBe(1);
expect(src.updateState).not.toHaveBeenCalled();
});

it('does not count a stale issue as moved when state transition fails', async () => {
const src = source();
src.updateState.mockResolvedValueOnce(false);
Expand Down
20 changes: 15 additions & 5 deletions src/automation/backlogGrooming.ts
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,10 @@ Rules:
- Keep updatedDescription concise and implementation-ready.`;
}

export function parseBacklogGroomingOutput(output: string): BacklogGroomingResult {
export function parseBacklogGroomingOutput(
output: string,
validIssueIds: Set<string>,
): BacklogGroomingResult {
try {
const fence = output.match(/```json\s*([\s\S]*?)```/i);
const jsonText = fence?.[1] ?? output.slice(output.indexOf('{'));
Expand All @@ -142,8 +145,11 @@ export function parseBacklogGroomingOutput(output: string): BacklogGroomingResul
const d = item as Partial<GroomingDecision>;
if (!d.issueId || !d.status || !d.reason) return [];
if (!['active', 'needs_update', 'stale'].includes(d.status)) return [];
const issueId = String(d.issueId);
// Drop hallucinated IDs before any downstream mutation path can see them.
if (!validIssueIds.has(issueId)) return [];
return [{
issueId: String(d.issueId),
issueId,
identifier: d.identifier ? String(d.identifier) : undefined,
status: d.status,
reason: String(d.reason),
Expand Down Expand Up @@ -177,7 +183,10 @@ export async function runBacklogGroomingPlanner(options: RunBacklogGroomingOptio
if (raw.exitCode !== 0 && !raw.stdout.trim()) {
return { success: false, decisions: [], error: raw.stderr.slice(0, 500) || `Planner adapter exited with code ${raw.exitCode}` };
}
return parseBacklogGroomingOutput(raw.stdout);
const validIssueIds = new Set(
options.tasks.map(task => task.issueId || task.id).filter(Boolean),
);
return parseBacklogGroomingOutput(raw.stdout, validIssueIds);
} catch (error) {
return { success: false, decisions: [], error: error instanceof Error ? error.message : String(error) };
}
Expand All @@ -198,7 +207,7 @@ export async function applyBacklogGrooming(
source: ITaskSource,
result: BacklogGroomingResult,
mode: BacklogGroomingMode = 'comment',
validIssueIds?: Set<string>,
validIssueIds: Set<string> = new Set(),
): Promise<ApplyBacklogGroomingResult> {
const applied: ApplyBacklogGroomingResult = {
commented: 0,
Expand All @@ -209,8 +218,9 @@ export async function applyBacklogGrooming(
skippedUnknown: 0,
};
if (!result.success) return applied;
// Scope is mandatory: an empty/missing Set must not mutate arbitrary IDs.
for (const decision of result.decisions) {
if (validIssueIds && !validIssueIds.has(decision.issueId)) {
if (!validIssueIds.has(decision.issueId)) {
applied.skippedUnknown++;
continue;
}
Expand Down
65 changes: 65 additions & 0 deletions src/automation/dailyReporter.retry.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { rmSync } from 'node:fs';

// Isolate the watermark file: generateDailyReports skips a window whose
// watermark already exists, and the real file lives in the developer's home.
const homeState = vi.hoisted(() => ({
home: `/tmp/openswarm-daily-retry-${process.pid}`,
}));

vi.mock('node:os', async (importOriginal) => ({
...(await importOriginal<typeof import('node:os')>()),
homedir: () => homeState.home,
}));

const postStatusUpdateMock = vi.fn();
vi.mock('../linear/index.js', () => ({
postStatusUpdate: (...args: unknown[]) => postStatusUpdateMock(...args),
}));

import {
generateDailyReports,
setLinearClient,
setTeamId,
} from './dailyReporter.js';

describe('generateDailyReports retry targeting', () => {
beforeEach(() => {
rmSync(homeState.home, { recursive: true, force: true });
postStatusUpdateMock.mockReset();
setTeamId('team-1');
});

it('retries only projects whose first publication update failed', async () => {
const projects = [
{ id: 'p-ok', name: 'Alpha', state: 'started' },
{ id: 'p-fail', name: 'Beta', state: 'started' },
{ id: 'p-ok2', name: 'Gamma', state: 'started' },
];

setLinearClient({
team: async () => ({
projects: async () => ({
nodes: projects,
pageInfo: { hasNextPage: false, endCursor: null },
}),
}),
} as never);

postStatusUpdateMock.mockImplementation(async (id: string) => {
if (id === 'p-fail') {
// Fail once, then succeed on retry.
if (postStatusUpdateMock.mock.calls.filter((c) => c[0] === 'p-fail').length === 1) {
throw new Error('transient Linear error');
}
}
});

await generateDailyReports();

const callsByProject = postStatusUpdateMock.mock.calls.map((c) => c[0] as string);
expect(callsByProject.filter((id) => id === 'p-ok')).toHaveLength(1);
expect(callsByProject.filter((id) => id === 'p-ok2')).toHaveLength(1);
expect(callsByProject.filter((id) => id === 'p-fail')).toHaveLength(2);
});
});
Loading
Loading