Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -7,12 +7,12 @@

# Anthropic
# ANTHROPIC_API_KEY=sk-ant-xxxxx
# ANTHROPIC_MODEL=claude-sonnet-4-20250514
# ANTHROPIC_MODEL=claude-opus-5-5

# OpenAI / 兼容服务
# OPENAI_API_KEY=sk-xxxxx
# OPENAI_BASE_URL=https://api.openai.com/v1
# OPENAI_MODEL=gpt-4o
# OPENAI_MODEL=gpt-6-astra

# 百炼
# DASHSCOPE_API_KEY=sk-xxxxx
Expand Down
97 changes: 97 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,103 @@

## [Unreleased]

## [0.2.130] - 2026-10-01

### Fixed (P131 — current models, and what P129/P130 broke)

**Every current Claude model was unusable.** The Anthropic adapter sent `temperature: 0.3`
on every request. Fable 5 / 5.1, Opus 4.7+ and Sonnet 5+ reject sampling parameters with a
400, and the retry only fired for errors that mentioned `thinking`. So every agent turn
failed on two of the three Claude presets. Reasoning levels were sent as `budget_tokens`,
which those models also reject (`off` sent `thinking.disabled`, a 400 on Opus 5.5).
- **Per-model request surface:** `thinking.ts` gains `claudeCaps()` /
`anthropicReasoningParams()`. These models now get no sampling fields, and reasoning
level maps to `thinking: {type:'adaptive', display:'summarized'}` + `output_config.effort`.
`off` on an always-thinking model becomes `effort: low`.
- **Broader retry:** a 400 naming temperature / thinking / effort is retried once without
those fields, so a model id we don't know yet degrades instead of dying.
- **Refusals:** `stop_reason: "refusal"` now shows a notice instead of an empty reply.

**GPT-6 Astra was unusable on api.openai.com.**
- The reasoning-model check only matched `gpt-5`, so `gpt-6-astra` got `max_tokens` +
`temperature`, both 400s.
- On `/chat/completions` it also rejects `reasoning_effort` together with tools, and
rejects the `minimal` effort. Agent turns now omit the effort (model default), and plain
chat maps `off` → `low`.
- The fallback detector now recognises "… are not supported" errors.
- "Test connection" sent `max_tokens` to every OpenAI reasoning model, so it failed for
GPT-5.x too. It now uses `max_completion_tokens`.
- The vision path had the same `max_tokens` / `temperature` problem.

**Sidecar stalls (since P129).** The stdin watchdog called `peek()` on stdin from a helper
thread. After ~2 s idle, the request that followed the next one was not read until more
input arrived, so back-to-back tool calls stalled into 60 s timeouts. That is likely the
real source of the "timed out but it finished" reports P130 worked around. The watchdog now
checks the parent PID and never touches stdin.

**P130 timeout ≠ failure, finished properly.**
- **Absolute cap:** heartbeats prove Python is alive, not that SolidWorks progresses. A COM
call stuck on a modal dialog heartbeated forever, and Stop could not break it (every new
message got AGENT_BUSY). Each call now has an absolute cap of 3× its budget (`sw_status`
has none), and `call()` takes the abort signal (code `CANCELLED`).
- **One request at a time:** a request queued behind a slow call burned its budget unread,
timed out, and its same-op_id retry re-ran a tool whose first run had failed (only
successes were cached). Requests now go out one at a time, budgets start when sent, and
failures are cached under their op_id too.
- **Namespaced op_ids:** op_ids are scoped per agent run. Providers that restart tool-call
ids per conversation (Kimi's `functions.<name>:<n>`) could otherwise receive an older
session's cached result.
- **Python 3.9 / 3.10:** `concurrent.futures.TimeoutError` is not the builtin before 3.11,
so the first heartbeat tick escaped as an empty TOOL_FAILED.
- **"Still running" note:** `onStillRunning` only fired after a timeout that heartbeats
prevented, so the note never appeared. It now fires on heartbeats, once a minute in chat.
- **Durations:** tool durations were dropped before reaching the step. They now show on the
tool row and in exports.
- **Amber busy dot:** the busy state never reached `StatusDot`. It is wired through now and
serves `connected:true` while a tool runs; concurrent status probes share one request.
- **`looksLikeQuestion`:** it now judges the reply's ending. Plans mentioning
which / 几个 / 哪个 / 多少 no longer read as questions, and
「我还缺少以下参数…」 / "I need a few values" no longer get auto-nudged.

**Settings and config**
- **Protocol switch:** switching to the OpenAI protocol paired `api.openai.com` with
`deepseek-v4-pro`. It now uses that endpoint's suggested model.
- **Quick-fill buttons:** they now apply the provider's model, context window and max
output. Those defaults were declared in `presets.ts` but never used.
- **Number fields:** they clamp on blur, not per keystroke. Typing "2" used to snap to 4096,
so values like 200000 could not be entered.
- **Env-fallback config:** an env-sourced config no longer replaces the saved preferences,
and the env key is never persisted to the store.
- **Config load:** saving before the stored config loaded overwrote it with the defaults
and wiped the API key.
- **Non-streaming timeout:** the 20 s connect-stage timeout no longer acts as a total timeout
for non-streaming requests (vision captions), and no longer re-sends finished-but-slow
requests up to 3×.
- **Truncation:** agent-loop truncation now counts the system prompt and tool schemas, and
never returns a lone tool result. The max-rounds summary turn passes the tool list, since
Anthropic 400s on tool blocks without tools.
- **IPC channels:** theme / locale channels moved into `ipc-channels.ts`.
- **CI:** `sidecar/tests/test_reliability.py` failed ruff E702, so the next CI run would
have been red.

### Changed
- **Recommended models:** GPT-6 Astra (`gpt-6-astra`, 1.05M context) for OpenAI;
Claude Opus 5.5 (`claude-opus-5-5`) and Claude Fable 5.1 (`claude-fable-5-1`) for
Anthropic.
- **Where they changed:** presets, provider defaults, README / README.zh-CN, USER-GUIDE,
ARCHITECTURE, `.env.example`. Saved model ids that left the preset list still load as
"Custom model".

### Tests
- **`tests/llm-request.test.mjs`:** the request bodies each adapter sends, per model, plus
truncation and `llmFetch`.
- **`tests/sidecar-client.test.mjs`:** the Node client against the real sidecar server with
fake tools: queueing, failure caching, cap + progress, cancel, and the idle-gap stall.
It fails 5/5 on 0.2.129.
- **Python:** failure caching, the futures timeout on <3.11, and the PID watchdog.
- **`looksLikeQuestion`:** plan and question cases.
- **Totals:** 191 JS + 58 Python tests.

## [0.2.129] - 2026-09-12

### Changed (P130 — timeout≠failure)
Expand Down
16 changes: 9 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,12 +28,12 @@
</p>

<p align="center">
<img src="https://img.shields.io/badge/version-0.2.37-blue" alt="version" />
<img src="https://img.shields.io/badge/version-0.2.130-blue" alt="version" />
<img src="https://img.shields.io/badge/electron-28-47848F?logo=electron" alt="electron" />
<img src="https://img.shields.io/badge/react-18-61DAFB?logo=react" alt="react" />
<img src="https://img.shields.io/badge/typescript-5.3-3178C6?logo=typescript" alt="typescript" />
<img src="https://img.shields.io/badge/python-3.9%2B-3776AB?logo=python&logoColor=white" alt="python" />
<img src="https://img.shields.io/badge/tests-167_JS_%2B_13_Python-brightgreen" alt="tests" />
<img src="https://img.shields.io/badge/tests-191_JS_%2B_58_Python-brightgreen" alt="tests" />
<img src="https://img.shields.io/badge/license-Apache_2.0-orange" alt="license" />
</p>

Expand Down Expand Up @@ -86,7 +86,7 @@ Millwright:
- **Agentic tool loop.** Observe → reason → act. The model chains multiple tool calls, reads structured JSON back from each one, and recovers from errors instead of failing silently.
- **Visual understanding.** Reorient, rotate, screenshot, and analyze the model — via a multimodal main model or a dedicated vision model.
- **Resident execution engine.** A persistent Python sidecar holds one COM connection open across an entire multi-step task.
- **Developer-friendly.** 167 TypeScript/Node tests plus a Python suite (`pytest sidecar/tests`) for the sidecar, a typed IPC boundary, and a `SKIP_SW_CONNECT` mode for UI-only development without SolidWorks installed.
- **Developer-friendly.** 191 TypeScript/Node tests plus a Python suite (`pytest sidecar/tests`) for the sidecar, a typed IPC boundary, and a `SKIP_SW_CONNECT` mode for UI-only development without SolidWorks installed.

## Cross-version compatibility

Expand Down Expand Up @@ -160,17 +160,19 @@ A `Millwright-*-x64.zip` is also published alongside the Setup installer for use

| Provider | Protocol | Base URL | Suggested model |
|---|---|---|---|
| OpenAI | OpenAI | `https://api.openai.com/v1` | `gpt-6-astra` (GPT-6 Astra) |
| Anthropic | Anthropic | `https://api.anthropic.com` | `claude-opus-5-5` (Opus 5.5, default) / `claude-fable-5-1` (Fable 5.1, most capable) |
| DeepSeek | OpenAI-compatible | `https://api.deepseek.com` | `deepseek-v4-pro` |
| Kimi / Moonshot | OpenAI-compatible | `https://api.moonshot.cn/v1` | `kimi-k3` |
| MiniMax | OpenAI-compatible | `https://api.minimaxi.com/v1` | `minimax-m3` |
| Anthropic | Anthropic | `https://api.anthropic.com` | `claude-sonnet-5` / `claude-opus-5` |
| OpenAI | OpenAI | `https://api.openai.com/v1` | `gpt-5.6` |
| Alibaba Bailian (Qwen) | OpenAI-compatible | `https://dashscope.aliyuncs.com/compatible-mode/v1` | `qwen-3.8max` |
| Zhipu (GLM) | OpenAI-compatible | `https://open.bigmodel.cn/api/paas/v4` | `glm-4.6` |
| SiliconFlow | OpenAI-compatible | `https://api.siliconflow.cn/v1` | — |
| Ollama (local) | OpenAI-compatible | `http://localhost:11434/v1` | — |

> Model IDs move fast — check your provider's docs for the current lineup. Agentic tool calling requires a model that supports function calling; DeepSeek V4, Kimi K3, MiniMax M3, and GLM-4.6 are first-class targets.
> Model IDs move fast — check your provider's docs for the current lineup. Agentic tool calling requires a model that supports function calling; GPT-6 Astra, Claude Opus 5.5 / Fable 5.1, DeepSeek V4, Kimi K3, MiniMax M3, and GLM-4.6 are first-class targets.
>
> The newest models reject request fields that older ones accepted: GPT-6 Astra needs `max_completion_tokens`, takes no `temperature`, and refuses `reasoning_effort` alongside tools on `/chat/completions`; Claude Opus 5.5 and Fable 5.1 reject `temperature` and `budget_tokens` and always think (depth is set with `effort`). Millwright detects these models by ID and sends the right fields — the **Reasoning depth** setting maps onto each model's own controls.

## Examples

Expand Down Expand Up @@ -244,7 +246,7 @@ Contributions welcome — see [CONTRIBUTING.md](docs/CONTRIBUTING.md). We especi

- [x] **v0.1** — MVP: Electron shell, LLM adapters, COM bridge, first tool set
- [x] **v0.2** — Python sidecar, agentic tool loop, dual-engine fallback, vision feedback, confirmation cards, Apache-2.0 open source
- [x] **v0.2.4 → v0.2.37** — Extensive hardening against real SolidWorks installs ← *current*: the sketch → feature → cut → visual-verification loop now runs end to end on real hardware
- [x] **v0.2.4 → v0.2.130** — Extensive hardening against real SolidWorks installs ← *current*: the sketch → feature → cut → visual-verification loop now runs end to end on real hardware
- [ ] **v0.3** — Streaming tool calls, sketching on model faces (not just reference planes), hole wizard, sheet metal, drawing annotations, remaining `#VERIFY` parameters confirmed
- [ ] **v1.0** — MCP server, multi-CAD support

Expand Down
16 changes: 9 additions & 7 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,12 +27,12 @@
</p>

<p align="center">
<img src="https://img.shields.io/badge/version-0.2.37-blue" alt="version" />
<img src="https://img.shields.io/badge/version-0.2.130-blue" alt="version" />
<img src="https://img.shields.io/badge/electron-28-47848F?logo=electron" alt="electron" />
<img src="https://img.shields.io/badge/react-18-61DAFB?logo=react" alt="react" />
<img src="https://img.shields.io/badge/typescript-5.3-3178C6?logo=typescript" alt="typescript" />
<img src="https://img.shields.io/badge/python-3.9%2B-3776AB?logo=python&logoColor=white" alt="python" />
<img src="https://img.shields.io/badge/tests-167_JS_%2B_13_Python-brightgreen" alt="tests" />
<img src="https://img.shields.io/badge/tests-191_JS_%2B_58_Python-brightgreen" alt="tests" />
<img src="https://img.shields.io/badge/license-Apache_2.0-orange" alt="license" />
</p>

Expand Down Expand Up @@ -85,7 +85,7 @@ Millwright:
- **Agent 工具循环。** 观察 → 推理 → 执行。模型串联多次工具调用,读取每次返回的结构化 JSON,出错能自愈而不是静默失败。
- **视觉理解。** 可翻转、旋转、截屏,再做分析——既支持多模态主模型,也支持独立视觉模型。
- **常驻执行引擎。** 常驻 Python 边车在一整个多步任务中复用同一条 COM 连接。
- **开发者友好。** 167 个 TS/Node 单元测试,另有独立的 Python 测试套件(`pytest sidecar/tests`),类型化 IPC 边界,`SKIP_SW_CONNECT` 纯 UI 开发模式(无需 SolidWorks)。
- **开发者友好。** 191 个 TS/Node 单元测试,另有独立的 Python 测试套件(`pytest sidecar/tests`),类型化 IPC 边界,`SKIP_SW_CONNECT` 纯 UI 开发模式(无需 SolidWorks)。

## 跨版本兼容

Expand Down Expand Up @@ -159,8 +159,8 @@ npm run dev

| 服务商 | 协议 | Base URL | 推荐模型 |
| ----------- | --------- | --------------------------------------------------- | ----------------------------------- |
| OpenAI | OpenAI | `https://api.openai.com/v1` | `gpt-5.6-sol` |
| Anthropic | Anthropic | `https://api.anthropic.com` | `claude-fable-5`(旗舰)/ `claude-opus-4-8`(更稳) |
| OpenAI | OpenAI | `https://api.openai.com/v1` | `gpt-6-astra`(GPT-6 Astra) |
| Anthropic | Anthropic | `https://api.anthropic.com` | `claude-opus-5-5`(Opus 5.5,默认)/ `claude-fable-5-1`(Fable 5.1,最强) |
| DeepSeek | OpenAI 兼容 | `https://api.deepseek.com` | `deepseek-v4-pro`(强) / `deepseek-v4-flash`(快) |
| Kimi / 月之暗面 | OpenAI 兼容 | `https://api.moonshot.cn/v1` | `kimi-k2.5` |
| MiniMax | OpenAI 兼容 | `https://api.minimaxi.com/v1` | `minimax-m3`(512K 上下文) |
Expand All @@ -169,7 +169,9 @@ npm run dev
| 硅基流动 | OpenAI 兼容 | `https://api.siliconflow.cn/v1` | —(用户自填) |
| Ollama(本地) | OpenAI 兼容 | `http://localhost:11434/v1` | —(用户自填) |

> 各家型号更新很快,请以服务商官方文档为准。Agent 工具调用需要模型支持 function calling;DeepSeek V4、Kimi K2、MiniMax M3、GLM-4.6 是一等公民。OpenAI `gpt-5.x` / o 系列和 Anthropic `claude-fable-5` 需要特定字段名,Millwright 已自动识别。
> 各家型号更新很快,请以服务商官方文档为准。Agent 工具调用需要模型支持 function calling;GPT-6 Astra、Claude Opus 5.5 / Fable 5.1、DeepSeek V4、Kimi K2、MiniMax M3、GLM-4.6 是一等公民。
>
> 新一代模型会拒绝老模型能接受的请求字段:GPT-6 Astra 要求 `max_completion_tokens`、不接受 `temperature`,并且在 `/chat/completions` 上不允许 `reasoning_effort` 与工具同时出现;Claude Opus 5.5 和 Fable 5.1 拒绝 `temperature` 和 `budget_tokens`,并且始终开启思考(深度用 `effort` 控制)。Millwright 按模型 ID 自动识别并发送正确的字段——设置里的「推理深度」会映射到各模型自己的控制参数上。OpenAI `gpt-5.x` / o 系列同样已自动识别。

## 使用示例

Expand Down Expand Up @@ -243,7 +245,7 @@ SolidWorks

- [x] **v0.1** — MVP:Electron 骨架、LLM 适配器、COM 桥接、首批工具
- [x] **v0.2** — Python 边车、agent 工具循环、双引擎降级、视觉反馈、确认卡片,Apache-2.0 开源
- [x] **v0.2.4 → v0.2.37** — 大量真机加固 ← *当前*:草图 → 特征 → 切除 → 视觉核验的完整闭环已在真机上端到端跑通
- [x] **v0.2.4 → v0.2.130** — 大量真机加固 ← *当前*:草图 → 特征 → 切除 → 视觉核验的完整闭环已在真机上端到端跑通
- [ ] **v0.3** — 流式工具调用、在模型面上画草图(而非仅基准面)、孔向导、钣金、工程图标注、剩余 `# VERIFY` 参数完成核验
- [ ] **v1.0** — MCP server、多 CAD 支持

Expand Down
4 changes: 2 additions & 2 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -213,8 +213,8 @@ const DEFAULT_SYSTEM_PROMPT = `You are a SolidWorks automation specialist.

| Provider | Protocol | Base URL | Example model |
|---------|----------|----------|----------|
| Anthropic | anthropic | https://api.anthropic.com | claude-sonnet-4-20250514 |
| OpenAI | openai | https://api.openai.com/v1 | gpt-4o |
| Anthropic | anthropic | https://api.anthropic.com | claude-opus-5-5 |
| OpenAI | openai | https://api.openai.com/v1 | gpt-6-astra |
| Bailian | openai | https://dashscope.aliyuncs.com/compatible-mode/v1 | qwen-coder-plus |
| MiniMax | openai | https://api.minimax.chat/v1 | MiniMax-Text-01 |
| DeepSeek | openai | https://api.deepseek.com | deepseek-chat |
Expand Down
4 changes: 2 additions & 2 deletions docs/ARCHITECTURE.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -213,8 +213,8 @@ const DEFAULT_SYSTEM_PROMPT = `你是一个 SolidWorks 自动化专家助手。

| 服务商 | 协议 | Base URL | 模型示例 |
|--------|------|----------|----------|
| Anthropic | anthropic | https://api.anthropic.com | claude-sonnet-4-20250514 |
| OpenAI | openai | https://api.openai.com/v1 | gpt-4o |
| Anthropic | anthropic | https://api.anthropic.com | claude-opus-5-5 |
| OpenAI | openai | https://api.openai.com/v1 | gpt-6-astra |
| 百炼 | openai | https://dashscope.aliyuncs.com/compatible-mode/v1 | qwen-coder-plus |
| MiniMax | openai | https://api.minimax.chat/v1 | MiniMax-Text-01 |
| DeepSeek | openai | https://api.deepseek.com | deepseek-chat |
Expand Down
4 changes: 2 additions & 2 deletions docs/USER-GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ If you have an Anthropic API key:
| API Protocol | Anthropic |
| Base URL | https://api.anthropic.com |
| API Key | Your `sk-ant-...` key |
| Model | `claude-sonnet-4-20250514` (recommended) |
| Model | `claude-opus-5-5` (recommended) or `claude-fable-5-1` (most capable) |

Get your API key by signing up at console.anthropic.com and creating a key.

Expand All @@ -53,7 +53,7 @@ Get your API key by signing up at console.anthropic.com and creating a key.
| API Protocol | OpenAI-compatible |
| Base URL | https://api.openai.com/v1 |
| API Key | Your `sk-...` key |
| Model | `gpt-4o` |
| Model | `gpt-6-astra` (recommended) |

### Option 3: Bailian (Alibaba Cloud)

Expand Down
4 changes: 2 additions & 2 deletions docs/USER-GUIDE.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@
| API 协议 | Anthropic |
| Base URL | https://api.anthropic.com |
| API Key | 你的 sk-ant-... 密钥 |
| 模型 | claude-sonnet-4-20250514(推荐) |
| 模型 | `claude-opus-5-5`(推荐)或 `claude-fable-5-1`(最强) |

获取 API Key:前往 console.anthropic.com 注册并创建密钥。

Expand All @@ -53,7 +53,7 @@
| API 协议 | OpenAI 兼容 |
| Base URL | https://api.openai.com/v1 |
| API Key | 你的 sk-... 密钥 |
| 模型 | gpt-4o |
| 模型 | `gpt-6-astra`(推荐) |

### 方式三:使用百炼(阿里云)

Expand Down
Loading
Loading