Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,12 @@ All notable changes follow Keep a Changelog. Versions follow Semantic Versioning

### Added

- Loopback-only `mini-code-agent web` command with a responsive three-pane local workbench,
real-time SSE lifecycle activity, task cancellation, and browser approval for governed actions.
- Bounded in-memory Web run manager with one active run, monotonic replayable events,
Future-based single-use approvals, and deterministic cancellation cleanup.
- Learning and resume documentation for the Web adapter, asyncio/SSE flow, browser trust boundary,
and Java backend concept mapping.
- Provider-backed `run` and `chat` CLI commands that compose the existing Agent runtime,
OpenAI-compatible or Anthropic adapters, bounded Workspace, and governed built-in tools.
- SiliconFlow configuration through `provider`, `model`, `base_url`, and the existing
Expand All @@ -15,6 +21,10 @@ All notable changes follow Keep a Changelog. Versions follow Semantic Versioning

### Security

- Web requests that mutate state require a process-random token; browser Origins must be
loopback, CORS is not enabled, and the CLI rejects remote binding.
- Workspace selection and Provider credentials stay server-side. Browser-rendered model text,
paths, commands, and diffs use text nodes rather than dynamic HTML.
- Read-only tools remain allowed by default. Writes and CLI-enabled command execution require an
interactive approval; non-interactive mode denies both without prompting.
- CLI output uses normalized public errors and never renders API key values. Live provider calls
Expand All @@ -24,6 +34,9 @@ All notable changes follow Keep a Changelog. Versions follow Semantic Versioning

### Verification

- M8 local Windows verification passed 1218 tests with 13 privilege/platform skips and 88.52%
branch-aware package coverage. Ruff format/check, strict Pyright, and browser layout checks at
1440x1024, 1024x768, and 390x844 passed.
- Local uv-managed Python 3.13.14 passed 1201 tests with 13 Windows privilege/platform skips and
88.56% branch-aware package coverage. Ruff format/check and strict Pyright passed.
- MockTransport verified the SiliconFlow-compatible
Expand Down
20 changes: 18 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,15 +2,16 @@

A framework-light, provider-neutral coding agent built from first principles.

> Status: pre-alpha. M7 provides a provider-neutral Agent Core, Anthropic/OpenAI-compatible
> Status: pre-alpha. M8 provides a provider-neutral Agent Core, Anthropic/OpenAI-compatible
> adapters, a schema-validating Tool Registry, a cross-platform Workspace boundary, bounded
> Read/Search, conflict-aware Write/Edit, policy-governed argv command execution, and deterministic
> context admission, hardened read-only Git evidence, governed Pytest diagnostics, versioned SQLite
> Session/Trace persistence, fail-closed Checkpoint/Resume, and a host-controlled bounded Repair
> loop, provenance-aware lazy Skills, deterministic host-registered Tool Hooks, and host-pinned
> local MCP stdio Tools, bounded host-profiled read-only analysis Subagents, and host-managed
> Worktree implementation candidates with separately approved adoption, plus provider-backed
> `run` and `chat` terminal commands with governed action previews. OS sandboxing,
> `run` and `chat` terminal commands plus a loopback-only Web console with live activity,
> governed action previews, approval, and cancellation. OS sandboxing,
> shell-string execution, project-provided executable Hooks, automatic Repair resume, remote
> HTTP/OAuth MCP, automatic commit/merge/push, and live-provider CI are not implemented.

Expand Down Expand Up @@ -74,6 +75,21 @@ Start an interactive task loop:
mini-code-agent chat --config .\config.toml --workspace .
```

Start the local Web console:

```powershell
mini-code-agent web --config .\config.toml --workspace .
```

The browser opens `http://127.0.0.1:8765` by default. Use `--no-open` to start only the server or
`--port` to choose another local port. The command rejects non-loopback hosts. The Workspace is
fixed when the process starts; it cannot be changed from the browser.

The Web console shows Agent lifecycle events, Tool activity, token usage, bounded action previews,
and file diffs. Write and command actions pause until they are approved or rejected in the
inspector. API keys stay in server-side settings; the browser receives only a configured/not
configured flag and a process-random request token.

Each `chat` prompt starts an independent bounded Agent run against the same workspace; durable
conversation memory is not implied. Read-only tools run automatically. File writes and local argv
commands display an action preview and require explicit confirmation. Use `--non-interactive` with
Expand Down
3 changes: 3 additions & 0 deletions config.example.toml
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,6 @@ base_url = "https://api.siliconflow.cn/v1"

# Keep API keys out of this file. For SiliconFlow, set:
# MINI_CODE_AGENT_OPENAI_API_KEY
#
# Then start the local UI with:
# mini-code-agent web --config .\config.example.toml --workspace .
115 changes: 115 additions & 0 deletions docs/learning/m8-web-console.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
# M8 学习笔记:把 Agent Runtime 变成可交互的本地 Web 产品

## 1. 本阶段解决的问题

M7 已经能通过命令行调用真实模型,但运行过程、工具活动、审批和 diff 都依赖终端展示。
M8 增加一个本地 Web 适配层,不重写 Agent Runtime:

```text
浏览器
-> FastAPI REST(启动、审批、取消)
-> WebRunManager
-> run_task()
-> AgentRuntime -> Provider / Tool / Policy
-> EventSink / ApprovalHandler
-> SSE -> 浏览器活动面板
```

产品目标是让用户看见 Agent 正在做什么,并在有副作用的动作执行前做决定。Web 层是
交互适配器,不是新的 Agent 框架。

## 2. 前置知识与本项目知识点

| 学习主题 | 先掌握什么 | 在项目中的落点 |
|---|---|---|
| Python 异步 | coroutine、Task、Future、取消传播 | 后台 Agent run、待审批 Future、取消与清理 |
| FastAPI | 路由、依赖、Middleware、StreamingResponse | REST、CSRF 校验、Origin 校验、SSE |
| Pydantic | frozen model、字段上下界、序列化 | Web 请求、运行快照、事件信封 |
| 浏览器基础 | fetch、EventSource、DOM API、响应式 CSS | 启动任务、消费事件、安全渲染、移动端抽屉 |
| Agent Harness | EventSink、ApprovalHandler、Policy | 把已有运行时能力映射到 Web,而不绕过治理 |
| Web 安全 | secret boundary、CSRF、Origin、CSP、XSS | Key 留在服务端、随机令牌、文本节点渲染 |

建议边做边补,不需要先系统学完前端或 FastAPI。

## 3. 核心实现

### 3.1 WebRunManager

`WebRunManager` 是 Web 层的运行状态机:

- 一次只允许一个活跃任务,避免多个浏览器操作争用同一工作区;
- 使用递增 `sequence` 给事件排序,并在有界 `deque` 中保留最近事件;
- 浏览器断线后通过 `after` 序号重放事件;
- 完成、失败和取消都产生明确终态;
- 生命周期事件不包含用户 prompt、Tool 参数/结果或 API Key。

这与 Flink JobManager 的相似点是都管理任务生命周期和状态;差异是这里是单进程、
单活跃任务,没有分布式容错和 Checkpoint 语义。

### 3.2 Future 驱动的人工审批

当 `GovernedToolExecutor` 请求审批时,`WebApprovalHandler`:

1. 创建 `asyncio.Future[bool]`;
2. 发布有界 `approval_required` 事件;
3. Agent 协程等待 Future,不执行工具;
4. 浏览器通过 REST 提交允许或拒绝;
5. Manager 先移除 Future,再设置结果,保证决定只能使用一次。

这类似 Java 的 `CompletableFuture<Boolean>`:生产者暂停等待外部决策,另一个请求处理器
完成 Future。取消任务时所有待审批 Future 都会被拒绝并清理。

### 3.3 SSE 为什么适合这里

浏览器到服务端只有启动、审批和取消三个低频命令,服务端到浏览器则持续发送运行事件。
SSE 提供单向事件流、浏览器原生 `EventSource` 和自动重连,不需要为双向 WebSocket
协议增加额外状态。事件带序号,重连时可以从最后位置继续。

当前不是逐 Token 流式输出。Provider 完成后才展示最终文本,SSE 传输的是 Agent
生命周期和工具活动。

### 3.4 浏览器和服务端的信任边界

- CLI 启动时固定 Workspace,浏览器不能传入任意路径;
- CLI 只允许回环 Host,不能使用 `0.0.0.0` 暴露到局域网;
- API Key 只由服务端配置读取,bootstrap 只返回布尔状态;
- 修改请求必须携带进程随机令牌,并通过 loopback Origin 检查;
- 模型文本、命令、路径和 diff 通过 `textContent` 显示;
- CSP 禁止第三方脚本和页面嵌入。

这些措施降低本地浏览器攻击面,但 Workspace/Policy 仍不是 OS Sandbox。

## 4. Java 后端经验映射

| Java / 数据开发概念 | Python / M8 对应 |
|---|---|
| Spring MVC Controller | FastAPI 路由函数 |
| HandlerInterceptor / Filter | FastAPI Middleware 和依赖 |
| CompletableFuture | `asyncio.Future` |
| ExecutorService Future.cancel | `asyncio.Task.cancel()` 和取消传播 |
| WebFlux ServerSentEvent | `StreamingResponse` + `text/event-stream` |
| DTO + Bean Validation | frozen Pydantic model + Field 约束 |
| ConcurrentHashMap 中的运行状态 | event-loop 内的 run/pending 字典 |
| Flink event-time sequence / offset | WebEvent sequence 与断线重放 |
| SQL 权限审批 | Tool Policy + 一次性 ActionPreview 审批 |

Python 的关键差异是:同一事件循环中的共享状态通常不需要线程锁,但不能在协程中执行
阻塞 I/O;取消是协作式异常传播,必须在 `finally` 中清理资源。

## 5. 建议学习练习

1. 从 `POST /api/runs` 跟到 `AgentRuntime.run()`,画出对象创建和调用顺序。
2. 跟踪一次 `run_command`:Policy ASK -> Future -> 浏览器允许 -> Tool 执行。
3. 删除 CSRF Header 或改成外部 Origin,观察接口为什么返回 403。
4. 运行中刷新页面,解释 bootstrap 的 `active_run` 与 SSE `after` 如何恢复界面。
5. 为 Manager 增加事件保留边界测试,说明慢客户端不会让内存无限增长。
6. 对比 SSE 和 WebSocket,说明本项目为什么暂时不需要双向长连接。

## 6. 当前边界

- 仅支持本地单用户、单工作区、单活跃任务;
- 运行和会话只保存在内存中,进程退出后不恢复 Web 状态;
- 最终回答不是逐 Token 流式输出;
- 当前未把 Skills、MCP、Subagent、Worktree candidate 组合进 Web composition root;
- 图像生成 API 尚未接入,后续应作为受治理 Tool,而不是让浏览器直接持有 Key;
- 自动化测试使用 Mock/Scripted Provider,不声明已完成真实 SiliconFlow smoke。
1 change: 1 addition & 0 deletions docs/learning/progress.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@
| L10 MCP | Complete and released | Governed stdio, exact grants, real SDK integration; v0.14 evidence |
| L11 Subagent and Worktree | Complete and released | Host-profiled analysis plus governed Worktree candidates/adoption; v0.16 evidence |
| L12 CI, benchmark and release | In progress | v0.16 prerelease and cross-platform evidence complete; benchmark remains separate |
| L13 Local Web console | Complete locally | FastAPI adapter, SSE event replay, Future approval, responsive browser workbench |

## L0 Notes

Expand Down
109 changes: 109 additions & 0 deletions docs/resume/m8-web-console-profile.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
# Mini CodeAgent M8 简历与面试说明

## 项目介绍

Mini CodeAgent 是一个使用 Python 从零实现的 Coding Agent Harness。项目实现了
Model -> ToolCall -> ToolResult -> Model 的有界循环,并把 Workspace、文件/Git/命令工具、
Policy、人工审批、上下文预算和 Provider Adapter 拆成可测试模块。M8 在不改动核心运行时
的前提下增加本地 Web 工作台,用于提交项目任务、观察运行活动、审批副作用操作和取消任务。

## 技术栈

- Python 3.12/3.13、asyncio、强类型 Protocol
- FastAPI、Uvicorn、Pydantic
- Server-Sent Events、REST、HTML/CSS、原生 JavaScript
- OpenAI-compatible Chat Completions、SiliconFlow 配置
- Typer、httpx、Pytest、pytest-asyncio
- Ruff、Pyright、GitHub Actions

## 30 秒面试介绍

“我实现了一个 Python Mini Coding Agent,不只是调用一次大模型,而是完整实现有界
Agent Loop、Tool Calling、工作区边界和有副作用工具审批。为了让运行过程可观察,我又做了
一个只绑定本机回环地址的 Web 工作台。FastAPI 负责启动、审批和取消接口,SSE 推送模型与
工具生命周期事件;命令或写文件时,Agent 通过 asyncio Future 暂停,用户查看资源、argv
和 diff 后只能批准一次。API Key 和 Workspace 都留在服务端,前端只接收必要状态并用文本
节点渲染模型内容。测试用 Scripted Provider 和 ASGITransport,不依赖真实额度。”

## 项目亮点

### 1. 核心运行时与 Web 交互解耦

**为什么使用:** CLI 和 Web 的输入输出方式不同,但 Provider、Tool、Policy 和 Agent Loop
不应该复制两套。

**技术实现:** `run_task()` 作为 Composition Root;Web 层只实现 `EventSink` 和
`ApprovalHandler` 两个协议,再通过依赖注入调用已有 Runtime。

**实现功能:** 同一套 Agent 能从终端或浏览器运行,并共享相同工具治理语义。

**解决问题:** 避免界面层绕过 Policy,也降低新增交互入口时的重复代码和行为漂移。

**代码证据:** `src/mini_code_agent/application.py`、
`src/mini_code_agent/web/manager.py`、`src/mini_code_agent/web/app.py`。

### 2. SSE 可观察运行与有界事件重放

**为什么使用:** 运行事件主要是服务端单向推送,WebSocket 的双向协议复杂度在这里没有收益。

**技术实现:** Manager 为事件分配单调递增 sequence,使用有界 deque 保留事件;
FastAPI `StreamingResponse` 输出具名 SSE,浏览器用 `EventSource` 消费并按序号重连。

**实现功能:** 展示模型调用、Tool 开始/结束、Token 用量、完成/失败/取消状态。

**解决问题:** 终端黑盒运行变成可追踪时间线;短暂断线后可恢复最近事件,同时限制内存增长。

**代码证据:** `WebRunManager.subscribe()`、`run_events()`、`static/app.js`。

### 3. Future 驱动的一次性人工审批

**为什么使用:** 模型提出命令或写入动作不等于用户授权,Web 请求和 Agent 协程又是两个独立
控制流。

**技术实现:** 每个 ToolCall 创建一个 `asyncio.Future[bool]`;审批事件包含有界
ActionPreview;REST 决策先从 pending map 移除 Future 再完成它,重复和过期决定返回冲突。

**实现功能:** Agent 在副作用执行前暂停,用户查看工具、风险、原因、资源、argv 和 diff 后
允许一次或拒绝;取消任务会拒绝并清理全部待审批。

**解决问题:** 防止模型自授权、重复点击和过期审批触发工具,保证取消不会遗留悬挂协程。

**代码证据:** `_WebApprovalHandler`、`decide_approval()`、`cancel()` 及 Manager 单元测试。

### 4. 本地 Web 的服务端信任边界

**为什么使用:** Coding Agent 拥有本地文件和进程能力,普通“localhost 页面”仍需防止
远程绑定、跨站请求和模型内容注入。

**技术实现:** CLI 拒绝非 loopback Host;Workspace 启动时固定;修改接口校验随机请求令牌
和 loopback Origin;Key 使用服务端 `SecretStr`/环境变量;设置 CSP;动态内容只写
`textContent`。

**实现功能:** 浏览器可以操作 Agent,但不能选择任意服务器路径、读取 Key 或从外部站点静默
发起批准请求。

**解决问题:** 缩小本地管理界面的攻击面,并明确 Web 治理与 OS Sandbox 的边界。

**代码证据:** `create_web_app()` Middleware/依赖、`web()` Host 校验、静态资源契约测试。

### 5. 确定性的无凭证测试

**为什么使用:** 真实模型存在费用、网络波动和非确定性,不适合作为默认 CI 前提。

**技术实现:** Manager 注入 async runner;API 使用 `httpx.ASGITransport`;核心 Agent 使用
Scripted Provider/MockTransport;前端用静态契约和浏览器响应式检查。

**实现功能:** 覆盖运行冲突、事件顺序、密钥脱敏、审批、取消、CSRF、Origin、SSE 和 CLI。

**解决问题:** 在不消耗 Token、不上传凭证的情况下验证控制流和协议边界。

**代码证据:** `tests/unit/web/`、`tests/cli/test_cli.py`。

## 诚实边界

- 不描述为 Claude Code 的完整替代品;
- 不声称是多用户或可公网部署的 Agent 平台;
- 不把 loopback、Workspace 或 Policy 描述为 OS Sandbox;
- 不声称默认 CI 验证了真实 SiliconFlow 账户;
- 不编造效率、准确率或成本下降百分比;
- 图像生成尚未接入当前 Agent 工具链。
Loading
Loading