Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,9 @@ jobs:
- run: uv run --no-sync ruff check .
- run: uv run --no-sync pyright
- run: uv build --build-constraint build-constraints.txt --require-hashes
env:
SOURCE_DATE_EPOCH: "1580601600"
- run: uv run --no-sync python tests/artifact_test.py
- run: uv run --python 3.13 --isolated --no-project --with dist/*.whl tests/smoke_test.py
- run: uv run --python 3.13 --isolated --no-project --with dist/*.tar.gz tests/smoke_test.py

Expand Down
44 changes: 44 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,50 @@ All notable changes follow Keep a Changelog. Versions follow Semantic Versioning

## [Unreleased]

## [0.14.0-alpha.0] - 2026-07-01

### Added

- Host-pinned local MCP stdio profiles with absolute executable/cwd validation, SecretStr
environment values, independent connection approval, fixed identity, exact Tool grants, and
lifecycle/content budgets.
- Official stable MCP Python SDK v1 adapter for protocol `2025-11-25`, with dedicated owner-worker
lifecycle management, bounded snapshots, process-tree shutdown, and discarded stderr.
- Canonical JSON Schema SHA-256 verification for complete non-paginated Tool sets; host-owned
local aliases, descriptions, side-effect classes, and risk levels.
- MCP `RegisteredTool` adapters with governed ActionPreview, deterministic text/structured JSON
results, output-schema validation, and static public errors.
- Per-Tool trust provenance in `GovernedToolExecutor`, allowing MCP aliases to reach Hooks and
Policy as `TrustSource.EXTENSION` while preserving the constructor default for native Tools.
- Real official-SDK stdio/Agent integration proving handshake, call, structured output, shutdown,
extension deny, independent Tool approval, schema drift rejection, and cross-task close.

### Changed

- Added `mcp>=1.28.1,<2` as a bounded runtime dependency. SDK v2 remains pre-release and is not
selected.
- `GovernedToolExecutor` accepts an optional copied mapping from registered Tool names to
`TrustSource`; unknown names and invalid values fail construction.
- MCP unit test filenames are globally unique so Pytest's default import mode can collect the
complete suite with existing command/skill tests.

### Security

- MCP commands must be absolute existing executable regular files; command and cwd reject
symlink/reparse paths and are revalidated immediately before process launch.
- Process approval shows complete argv/cwd and environment names before any server code runs.
Environment values remain secret, and connection approval never replaces per-Tool Policy or
approval.
- Protocol/server identity, Tools capability, static Tool list, exact grant set, schema hashes,
and task mode all fail closed before any alias is published.
- Server instructions, descriptions, titles, annotations, icons, `_meta`, and stderr do not enter
model-facing definitions or results.
- Results reject image/audio/resource content, non-finite or excessive JSON, oversized text/bytes,
and successful output-schema mismatches without returning partial success.
- Local stdio processes retain the Agent user's OS authority. M5b does not claim sandboxing,
executable provenance, package safety, remote MCP/OAuth security, rollback, or exactly-once
side effects.

## [0.13.0-alpha.0] - 2026-07-01

### Added
Expand Down
29 changes: 25 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,14 +2,15 @@

A framework-light, provider-neutral coding agent built from first principles.

> Status: pre-alpha. M5a provides a provider-neutral Agent Core, Anthropic/OpenAI-compatible
> Status: pre-alpha. M5b provides a provider-neutral Agent Core, Anthropic/OpenAI-compatible
> adapters, a schema-validating Tool Registry, a cross-platform Workspace boundary, bounded
> Read/Search, conflict-aware Write/Edit, policy-governed argv command execution, and deterministic
> context admission, hardened read-only Git evidence, governed Pytest diagnostics, versioned SQLite
> Session/Trace persistence, fail-closed Checkpoint/Resume, and a host-controlled bounded Repair
> loop, provenance-aware lazy Skills, and deterministic host-registered Tool Hooks. OS sandboxing,
> shell-string execution, project-provided executable Hooks, automatic Repair resume, MCP, and
> live-provider CI are not implemented.
> loop, provenance-aware lazy Skills, deterministic host-registered Tool Hooks, and host-pinned
> local MCP stdio Tools. OS sandboxing, shell-string execution, project-provided executable Hooks,
> automatic Repair resume, remote HTTP/OAuth MCP, Subagents/Worktrees, and live-provider CI are not
> implemented.

## Requirements

Expand Down Expand Up @@ -217,6 +218,24 @@ actual result; timeout, exception, or invalid return cannot replace it. Reposito
prompt Hooks and dynamic Python imports are not supported. In-process Hooks have the Agent
process authority and are not sandboxed. See `docs/architecture/governed-extensions.md`.

## Governed MCP Stdio

Local MCP Tools use the official stable Python SDK v1 over direct stdio. A trusted host profile
pins an absolute executable/argv/cwd, server identity, exact Tool grant set, host-owned
description/side-effect/risk, canonical input/output-schema hashes, and hard lifecycle/content
limits.

Starting the process requires dedicated connection approval. Verified local aliases still pass
through the ordinary Tool Registry, Hooks, Policy, and optional Tool approval with
`TrustSource.EXTENSION`. Server instructions, descriptions, annotations, icons, and `_meta` are
not authority and are not copied into model-facing definitions.

Only bounded text and object-shaped structured JSON results are accepted. Calls are serialized
and never retried; a timed-out side-effecting Tool reports uncertain completion. Stdio and user
approval are not OS sandboxing. Remote HTTP/OAuth, Resources, Prompts, Roots, Sampling,
Elicitation, Tasks, dynamic Tool lists, and package installation are not supported. See
`docs/architecture/governed-mcp.md`.

## Documentation

- Product design: `docs/superpowers/specs/2026-06-29-mini-code-agent-design.md`
Expand All @@ -235,6 +254,7 @@ process authority and are not sandboxed. See `docs/architecture/governed-extensi
- Governed test execution: `docs/architecture/governed-test-execution.md`
- Bounded Repair loop: `docs/architecture/bounded-repair-loop.md`
- Governed Skills and Hooks: `docs/architecture/governed-extensions.md`
- Governed MCP stdio: `docs/architecture/governed-mcp.md`
- Threat model: `docs/architecture/threat-model.md`
- Provider protocol ADR: `docs/adr/0002-provider-wire-protocols.md`
- Workspace boundary ADR: `docs/adr/0003-workspace-boundary.md`
Expand All @@ -247,6 +267,7 @@ process authority and are not sandboxed. See `docs/architecture/governed-extensi
- Fixed Pytest/JUnit boundary ADR: `docs/adr/0010-fixed-pytest-junit-boundary.md`
- Host-controlled bounded Repair ADR: `docs/adr/0011-host-controlled-bounded-repair.md`
- Inert Skills and host Hooks ADR: `docs/adr/0012-inert-skills-host-hooks.md`
- Host-pinned stdio MCP ADR: `docs/adr/0013-host-pinned-stdio-mcp.md`

## License

Expand Down
25 changes: 22 additions & 3 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,9 @@ after the repository is published. Until then, contact the repository owner priv

## Current Boundary

Model output, repository content, project Skills, Tool arguments, test reports, and future MCP
servers are untrusted inputs. File, command, Git, test, and Repair actions pass typed validation,
Workspace boundaries, Policy, and approval where applicable.
Model output, repository content, project Skills, Tool arguments, test reports, and MCP servers
are untrusted inputs. File, command, Git, test, Repair, and MCP Tool actions pass typed validation,
Policy, and approval where applicable.

M5a Skills are inert Markdown data. Discovery rejects links/reparse points, unsafe YAML, invalid
metadata, conflicts, drift, and resource-limit violations. Parsing or hashing a Skill does not
Expand All @@ -25,6 +25,25 @@ output, or environment-selected modules. Pre-Hooks can deny but cannot grant Pol
post-Hook failures cannot rewrite Tool results. A malicious host-registered Hook still has the
Agent process authority.

M5b MCP supports host-configured local stdio Tools only. Before launch, a dedicated approver sees
the exact absolute executable, argv, cwd, and environment variable names; values remain secret.
The executable and cwd reject links/reparse points and are revalidated before process creation.
Initialization pins protocol and server identity; the complete Tool set and canonical input/output
schema hashes must exactly match host grants. Server descriptions, instructions, annotations,
icons, and metadata do not grant authority.

Verified MCP aliases use `TrustSource.EXTENSION` and still pass the ordinary Tool Policy and
optional per-call approval. Connection approval does not approve future Tool calls. Results accept
only bounded text and object-shaped structured JSON; unsupported or oversized content fails
without partial output.

Local MCP processes run with the Agent user's OS privileges and may act during startup before a
Tool call. Stdio limits protocol access but is not a filesystem, network, process, or credential
sandbox. Schema hashes detect reviewed-contract drift, not executable provenance or behavior.
Timeout/cancellation cannot prove that a remote side effect did not complete. The project does not
support remote HTTP/OAuth MCP, package installation, executable signatures, dynamic Tool lists,
Resources, Prompts, Roots, Sampling, Elicitation, or Tasks.

The project does not claim OS-level sandboxing unless an explicit sandbox backend is enabled and
documented. It also does not claim that Hook timeout stops work delegated to another thread or
process, or that SHA-256 establishes extension authorship.
80 changes: 80 additions & 0 deletions docs/adr/0013-host-pinned-stdio-mcp.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# ADR 0013: Use Host-Pinned Stdio MCP Grants

- Status: Accepted
- Date: 2026-07-01

## Context

MCP can expose external Tools through a common protocol, but protocol compatibility does not
establish trust. A local MCP command runs with the Agent user's privileges and can perform work
during startup, before local Tool Policy evaluates a call. A server can also change Tool names,
schemas, descriptions, annotations, and behavior between versions.

The public MCP surface includes local stdio, remote HTTP, OAuth, Resources, Prompts, Roots,
Sampling, Elicitation, Tasks, notifications, pagination, and dynamic Tool lists. Adding all of
these at once would combine process execution, network authorization, prompt injection, delegated
model access, credential handling, and changing capability sets in one boundary.

The official Python SDK v1 is the stable production line. SDK v2 is pre-release and has different
architecture and future protocol targets.

## Decision

M5b supports only direct local stdio Tools through `mcp>=1.28.1,<2`.

The trusted host supplies an immutable profile with:

- an absolute existing executable and exact argv/cwd/environment names;
- explicit connection approval before process creation;
- exact protocol and server identity;
- an exact Tool grant set;
- host descriptions, side-effect classes, and risk levels;
- canonical input/output-schema hashes;
- hard lifecycle and content limits.

The complete observed Tool set must equal the grants. Dynamic lists and pagination are rejected.
Server instructions, annotations, titles, descriptions, and metadata are ignored. Verified MCP
Tools use local aliases and flow through the ordinary Registry, ActionPreview, Hooks, Policy, Tool
approval, and result bounds with `TrustSource.EXTENSION`.

The production adapter owns SDK context managers in a dedicated task so context exit and process
cleanup obey AnyIO task affinity while callers may use or close the proxy from another task.

Remote transports, OAuth, other server features, automatic retries, package installation, and OS
sandbox claims are deferred.

## Consequences

Positive:

- MCP discovery cannot silently add Agent authority;
- package/server/schema drift fails before Tool publication;
- server prompt-like metadata cannot rewrite model-facing Tool definitions;
- process approval and per-call approval remain distinct and understandable;
- the existing Policy/Hook/Trace Tool path remains the single call authority;
- official SDK lifecycle and process-tree behavior are reused instead of hand-rolled JSON-RPC;
- SDK task-affinity details remain inside the adapter.

Negative:

- each approved server upgrade requires reviewing identity and schema hashes;
- local commands must be resolved to absolute executable paths;
- dynamic and paginated Tool servers are unsupported;
- stderr is discarded, reducing production diagnostics;
- calls are serialized and not retried;
- a local process still has the user's OS permissions;
- executable signatures, package provenance, and OS sandboxing are not provided.

## Alternatives Rejected

- **Hand-written JSON-RPC:** duplicates version negotiation, cancellation, protocol types, and
process shutdown without improving the product boundary.
- **Trust every discovered Tool:** lets server/package replacement create unreviewed authority.
- **Trust server annotations for side effects:** annotations are untrusted hints, not Policy.
- **One approval for connection and all calls:** hides the difference between starting code and
authorizing a represented action.
- **Repository-defined MCP commands:** lets inspected content execute before Tool Policy.
- **Shell command strings:** introduce expansion and injection; exact argv is required.
- **Remote HTTP in M5b:** requires an independent OAuth, SSRF, redirect, token, and endpoint
identity design.
- **SDK v2 pre-release:** inappropriate for the project's stable production dependency boundary.
Loading