Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
52 changes: 52 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,58 @@ All notable changes follow Keep a Changelog. Versions follow Semantic Versioning

## [Unreleased]

## [0.15.0-alpha.0] - 2026-07-02

### Added

- Immutable host-owned `SubagentProfile` contracts with exact child Tools, independent Agent/
batch limits, no recursion, and `TrustSource.SUBAGENT`.
- `SubagentSupervisor` with fresh child contexts, preflight composition, structured
`asyncio.TaskGroup` concurrency, per-child and outer deadlines, ordered aggregation, and
cancellation propagation.
- Bounded child/batch result models, canonical SHA-256 projections, ToolResult evidence hashes,
static failures, and metadata-only lifecycle events.
- One dynamic governed parent analysis Tool per profile, with profile-specific JSON Schema,
medium-risk preview, canonical ASCII result JSON, and UTF-8 byte budget.
- Real parent/child Agent integration using governed `ReadFileTool`/`SearchTextTool`, including
Policy deny, non-recursion, sibling timeout isolation, event omission, byte-identical Workspace,
and parent cancellation.
- M6a architecture, threat-model, ADR, learning, and resume documentation.

### Changed

- Added `TrustSource.SUBAGENT` so child Tool Policy is independently addressable from parent model
and extension calls.
- The stable installed-package smoke imports the public Subagent profile, supervisor, Tool, and
builder API.
- Subagent unit tests use a package namespace so Pytest's default import mode can collect the full
suite alongside existing same-named test modules.

### Security

- Every child ID, Provider, and governed Tool executor is validated before child Provider I/O;
malformed/duplicate IDs, reused objects, capability drift, non-read-only definitions, and
non-SUBAGENT provenance fail the complete batch.
- Duplicate, empty, NUL-containing, oversized, or excessive tasks fail before child composition.
- Child sessions are non-interactive, cannot receive a delegation Tool, and cannot convert `ASK`
into authority.
- Child/batch timeout is isolated, external cancellation is re-raised, and no detached asyncio
task survives the parent ToolCall.
- Events exclude task/prompt/message/summary/argument/result/exception content; evidence retains
only bounded metadata and SHA-256.
- In-process children are not an OS sandbox. M6a does not claim Tool-implementation isolation,
durable parent-child audit, semantic proof from hashes, rollback, or exactly-once behavior.

### Verification

- Python 3.12.13 passed 1060 tests with 10 Windows symlink-privilege skips and 91.08% branch
coverage before the final two hardening regressions; the complete focused Subagent/integration
suite then passed 100 tests.
- Final Python 3.13.14 passed 1062 tests with the same 10 platform skips and 91.09% branch
coverage. Ruff format/check, strict Pyright, Bandit, and locked runtime pip-audit passed.
- Remote CI, reproducible artifact, installed wheel/sdist, tag, and GitHub prerelease evidence is
pending the release task and is not claimed here.

## [0.14.0-alpha.0] - 2026-07-01

### Added
Expand Down
31 changes: 27 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,15 +2,15 @@

A framework-light, provider-neutral coding agent built from first principles.

> Status: pre-alpha. M5b provides a provider-neutral Agent Core, Anthropic/OpenAI-compatible
> Status: pre-alpha. M6a provides a provider-neutral Agent Core, Anthropic/OpenAI-compatible
> adapters, a schema-validating Tool Registry, a cross-platform Workspace boundary, bounded
> Read/Search, conflict-aware Write/Edit, policy-governed argv command execution, and deterministic
> context admission, hardened read-only Git evidence, governed Pytest diagnostics, versioned SQLite
> Session/Trace persistence, fail-closed Checkpoint/Resume, and a host-controlled bounded Repair
> loop, provenance-aware lazy Skills, deterministic host-registered Tool Hooks, and host-pinned
> local MCP stdio Tools. OS sandboxing, shell-string execution, project-provided executable Hooks,
> automatic Repair resume, remote HTTP/OAuth MCP, Subagents/Worktrees, and live-provider CI are not
> implemented.
> local MCP stdio Tools, and bounded host-profiled read-only analysis Subagents. OS sandboxing,
> shell-string execution, project-provided executable Hooks, automatic Repair resume, remote
> HTTP/OAuth MCP, write-capable Subagents/Worktrees, and live-provider CI are not implemented.

## Requirements

Expand Down Expand Up @@ -236,6 +236,27 @@ approval are not OS sandboxing. Remote HTTP/OAuth, Resources, Prompts, Roots, Sa
Elicitation, Tasks, dynamic Tool lists, and package installation are not supported. See
`docs/architecture/governed-mcp.md`.

## Governed Analysis Subagents

M6a exposes one governed parent Tool per immutable host profile. The model supplies one to four
unique bounded tasks; the host fixes the child system prompt, exact read-only Tool names, Agent
limits, concurrency, deadlines, and result budgets.

Before any child Provider request, `SubagentSupervisor` requires distinct Providers/executors,
exact `READ_ONLY` definitions, `governance_enforced is True`, and
`TrustSource.SUBAGENT` for every child Tool. Each child receives one fresh task message, not the
parent or sibling transcript. Delegation Tools are structurally unavailable to children.

All children belong to one `asyncio.TaskGroup`. A semaphore bounds concurrency, individual and
batch timeouts have typed outcomes, input order is preserved, and external cancellation cancels
and joins every child before being re-raised. Parent results contain bounded untrusted summaries
and ToolResult metadata/SHA-256 evidence, not raw child transcripts or Tool content. Events omit
tasks, prompts, summaries, arguments, results, repository content, and exception text.

In-process context isolation is not an OS sandbox. M6a cannot write, run commands, call network
Tools, open nested approval prompts, persist durable child traces, create Worktrees, or merge
changes. See `docs/architecture/governed-subagents.md`.

## Documentation

- Product design: `docs/superpowers/specs/2026-06-29-mini-code-agent-design.md`
Expand All @@ -255,6 +276,7 @@ Elicitation, Tasks, dynamic Tool lists, and package installation are not support
- Bounded Repair loop: `docs/architecture/bounded-repair-loop.md`
- Governed Skills and Hooks: `docs/architecture/governed-extensions.md`
- Governed MCP stdio: `docs/architecture/governed-mcp.md`
- Governed analysis Subagents: `docs/architecture/governed-subagents.md`
- Threat model: `docs/architecture/threat-model.md`
- Provider protocol ADR: `docs/adr/0002-provider-wire-protocols.md`
- Workspace boundary ADR: `docs/adr/0003-workspace-boundary.md`
Expand All @@ -268,6 +290,7 @@ Elicitation, Tasks, dynamic Tool lists, and package installation are not support
- Host-controlled bounded Repair ADR: `docs/adr/0011-host-controlled-bounded-repair.md`
- Inert Skills and host Hooks ADR: `docs/adr/0012-inert-skills-host-hooks.md`
- Host-pinned stdio MCP ADR: `docs/adr/0013-host-pinned-stdio-mcp.md`
- Bounded host-profiled Subagents ADR: `docs/adr/0014-bounded-host-profiled-subagents.md`

## License

Expand Down
24 changes: 21 additions & 3 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,9 @@ after the repository is published. Until then, contact the repository owner priv

## Current Boundary

Model output, repository content, project Skills, Tool arguments, test reports, and MCP servers
are untrusted inputs. File, command, Git, test, Repair, and MCP Tool actions pass typed validation,
Policy, and approval where applicable.
Model output, repository content, project Skills, Tool arguments, test reports, MCP servers, child
Agent output, and child summaries are untrusted inputs. File, command, Git, test, Repair, MCP, and
Subagent Tool actions pass typed validation, Policy, and approval where applicable.

M5a Skills are inert Markdown data. Discovery rejects links/reparse points, unsafe YAML, invalid
metadata, conflicts, drift, and resource-limit violations. Parsing or hashing a Skill does not
Expand Down Expand Up @@ -44,6 +44,24 @@ Timeout/cancellation cannot prove that a remote side effect did not complete. Th
support remote HTTP/OAuth MCP, package installation, executable signatures, dynamic Tool lists,
Resources, Prompts, Roots, Sampling, Elicitation, or Tasks.

M6a analysis Subagents use immutable host profiles, fresh child contexts, exact read-only Tool
sets, independent Agent/result budgets, and `TrustSource.SUBAGENT`. Every child Tool executor must
prove governance and SUBAGENT provenance before any child Provider request. Child sessions are
non-interactive, so `ASK` fails closed instead of opening nested approval. Parent Policy deny
prevents child factories and Provider I/O.

All child tasks belong to one `asyncio.TaskGroup`; child and batch deadlines are separate, and
external cancellation is re-raised after children are cancelled and joined. Results expose only
bounded untrusted summaries and ToolResult metadata/SHA-256 evidence. Subagent events exclude
tasks, prompts, messages, summaries, Tool arguments/results, repository content, and exception
text.

M6a children are in-process and are not an OS, process, memory, credential, or network sandbox.
Read-only Tool admission does not constrain malicious host-supplied Provider or Tool code.
Evidence hashes are not signatures, semantic validation, confidentiality, or durable audit.
M6a does not support child writes, command/network Tools, recursive delegation, Worktrees,
candidate adoption, merge, rollback, or exactly-once execution.

The project does not claim OS-level sandboxing unless an explicit sandbox backend is enabled and
documented. It also does not claim that Hook timeout stops work delegated to another thread or
process, or that SHA-256 establishes extension authorship.
109 changes: 109 additions & 0 deletions docs/adr/0014-bounded-host-profiled-subagents.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
# ADR 0014: Use Bounded Host-Profiled Analysis Subagents

- Status: Accepted
- Date: 2026-07-02

## Context

Independent code-reading tasks can be parallelized, and a child Agent can keep exploratory
messages out of the parent transcript. A naive Subagent feature, however, can duplicate parent
authority, inherit unrelated context, recursively create more Agents, start background approval
prompts, swallow cancellation, leak raw Tool content into aggregation, or leave orphan tasks.

The project already has a typed `AgentRuntime`, Tool Registry, Workspace boundary, Policy,
provenance, and deterministic limits. A Subagent design should reuse these contracts rather than
introduce a second execution and authorization system.

Write-capable children introduce a different problem: concurrent mutation, repository identity,
candidate persistence, merge/adoption authority, and cleanup uncertainty. Combining that problem
with initial analysis delegation would make the first boundary too broad.

## Decision

M6a implements host-profiled, in-process, non-recursive analysis children.

The trusted host creates one immutable `SubagentProfile` per parent Tool. It fixes the child
system prompt, exact ordered Tool names, Agent limits, concurrency, deadlines, evidence, summary,
and result budgets. The model supplies only one to four unique bounded tasks plus a reason.

Before any child Provider request, `SubagentSupervisor` creates and validates every child:

- unique host child ID;
- distinct Provider and governed Tool executor;
- exact `READ_ONLY` definitions;
- `governance_enforced is True`;
- `TrustSource.SUBAGENT` for every child Tool;
- no delegation Tool.

Every child gets a fresh one-message context and an independent `AgentRuntime`. All child tasks
belong to one `asyncio.TaskGroup`; a semaphore bounds concurrency, child and batch timeouts are
separate, results are stored by input ordinal, and external `CancelledError` is re-raised.

Children run in `SessionMode.NON_INTERACTIVE`, so a child Policy `ASK` cannot open a nested prompt
and fails closed.

The parent receives only bounded typed projections: untrusted final summaries, usage/counts,
static failure metadata, and SHA-256 evidence for correlated ToolResult content. Lifecycle events
contain metadata and hashes but never task text, prompts, messages, summaries, Tool arguments,
ToolResult content, repository content, or exception text.

M6a is read-only. Worktree-backed implementation children and candidate adoption are deferred to
M6b.

## Consequences

Positive:

- child authority is an exact host capability profile, not copied parent authority;
- parent and sibling transcripts are not implicitly shared;
- Policy can distinguish parent-model calls from delegated calls;
- recursive delegation is structurally unavailable;
- one child timeout/failure does not erase sibling results;
- TaskGroup gives one lexical owner for cancellation and joining;
- input order remains stable under out-of-order completion;
- aggregation does not copy raw repository or ToolResult content;
- parent Policy deny prevents child factories and Provider I/O;
- the existing Agent/Tool/Workspace contracts remain the execution path.

Negative:

- each task adds a Provider session and may increase cost, latency, and rate-limit pressure;
- fresh children repeat context that a forked child might otherwise inherit;
- in-process Provider and Tool implementations share memory and OS authority;
- read-only admission cannot sandbox malicious host code;
- evidence hashes do not validate semantic correctness;
- child events are best-effort and lack durable parent run/turn linkage;
- `NON_INTERACTIVE` means child work requiring approval cannot proceed;
- M6a cannot implement or merge code changes.

## Alternatives Rejected

- **Give children the parent transcript:** leaks unrelated context, weakens attribution, and makes
context growth implicit.
- **Let the model define child prompts and Tools:** allows model output to create authority.
- **Copy the complete parent Tool Registry:** can expose writes, commands, network access, MCP, or
delegation without a separate decision.
- **Recursive delegation with a depth counter:** a depth limit does not solve capability
amplification, cost fan-out, or audit complexity; M6a admits no delegation Tool.
- **Detached `asyncio.create_task`:** permits orphan work and ambiguous cancellation ownership.
- **`asyncio.gather(return_exceptions=True)`:** makes cancellation and task-lifetime invariants
less explicit than TaskGroup plus typed child outcomes.
- **One timeout only:** cannot distinguish a slow child from a whole-batch deadline.
- **Interactive child approval:** background children must not compete for user prompts or reuse
parent approval.
- **Return complete child transcripts:** consumes parent context and leaks Tool arguments/results
rather than a bounded projection.
- **Run every child in a subprocess now:** stronger interpreter isolation also requires
credential transport, authenticated IPC, Provider lifecycle, process cleanup, and durable
result protocols; deferred until the in-process contract is stable.
- **Enable writes in the parent checkout:** concurrent children can collide with user work and
one another. M6b requires host-managed Worktrees and explicit candidate adoption.
- **Adopt a multi-agent framework:** would duplicate or obscure the project's existing
Provider/Tool/Policy semantics and cancellation evidence.

## Follow-up

M6b may add one write-capable implementation profile only through host-created, locked,
no-checkout Git Worktrees. A child result will produce a bounded candidate snapshot; a separate
parent-side approval will control adoption. No M6b change may weaken M6a's exact profiles,
SUBAGENT provenance, cancellation propagation, or no-recursion rule.
Loading
Loading