Skip to content

fix: P131 current-model request bodies, sidecar stall + P130 follow-ups; recommend GPT-6 Astra / Claude Opus 5.5 / Fable 5.1 - #3

Merged
raylanlin merged 1 commit into
masterfrom
claude/gracious-davinci-j7qso7
Oct 1, 2026
Merged

raylanlin merged 1 commit into
masterfrom
claude/gracious-davinci-j7qso7

Conversation

@raylanlin

Copy link
Copy Markdown
Owner

What Changed

v0.2.130 (P131) fixes three problems: current Claude models and GPT-6 Astra were unusable, the sidecar stalled after idle gaps, and several P130 features never took effect. It also updates the recommended models.

Model support

  • Anthropic:
    • temperature was sent on every request, a 400 on Fable 5/5.1, Opus 4.7+ and Sonnet 5+ (every Claude preset except Sonnet 4.6). It is no longer sent to those models.
    • Reasoning level now maps to thinking: {type:'adaptive'} + output_config.effort. The removed budget_tokens form is no longer sent.
    • Always-thinking models never get thinking.disabled.
    • A 400 that names sampling, thinking or effort fields is retried once without them.
    • stop_reason: "refusal" now shows a notice instead of an empty reply.
  • OpenAI:
    • gpt-6* is now treated as a reasoning model: max_completion_tokens, no temperature.
    • On GPT-6 /chat/completions, reasoning_effort is omitted whenever tools are present.
    • The reasoning-param fallback now recognises "… are not supported" errors.
    • Test connection and the vision path send max_completion_tokens to reasoning models.

Sidecar

  • Stall after idle: the P129 stdin watchdog peeked stdin from a helper thread. After about 2 s idle, back-to-back requests stalled until they timed out. It now checks the parent PID instead.
  • Request handling:
    • Requests go out one at a time, and a budget starts only when its request is sent.
    • Heartbeats can stretch a call to at most 3× its budget.
    • Stop now cancels the wait (code CANCELLED).
  • No double runs: failed results are cached under their op_id, so the TIMEOUT follow-up never re-runs a half-built generator. op_ids are namespaced per agent run.
  • Python < 3.11: the heartbeat loop now catches concurrent.futures.TimeoutError.
  • P130 features now wired up:
    • The "still running" note appears.
    • Tool durations reach the tool row and the exports.
    • The amber busy dot shows while a tool runs.
  • looksLikeQuestion: it now judges the end of the reply.

Settings / config

  • Protocol switch and quick-fill: they apply a matching model and the provider's defaults.
  • Number fields: they clamp on blur instead of on every keystroke.
  • Env fallback: it keeps the saved preferences, and the env key is never persisted.
  • Config load: nothing is saved before the stored config has loaded.
  • Non-streaming requests: they no longer hit the 20 s connect timeout, and finished requests are not re-sent.
  • Truncation: it now counts the system prompt and tool schemas.
  • Max-rounds summary: it passes the tool list.
  • IPC channels: theme and locale channels move into ipc-channels.ts.
  • Lint: fixes the ruff E702 error in test_reliability.py.

Docs

  • Presets, README / README.zh-CN, USER-GUIDE, ARCHITECTURE and .env.example now recommend gpt-6-astra for OpenAI and claude-opus-5-5 / claude-fable-5-1 for Anthropic.
  • Version bumped to 0.2.130, with a CHANGELOG entry.

Related Issue

N/A. The issues came from review of v0.2.129 and from the user's request to update the recommended models.

Type of Change

  • 🐛 Bug fix
  • ✨ New feature
  • 📝 Documentation
  • ♻️ Refactor
  • 🧪 Tests
  • 🔧 Build / CI

Testing

  • npm test passes: 191 tests, including the new tests/llm-request.test.mjs and tests/sidecar-client.test.mjs.
  • npm run lint passes. npm run typecheck, ruff check sidecar/, pytest sidecar/tests (58 passed), compileall, the renderer build and scripts/precommit-check.sh 0.2.130 also pass.
  • Manual test, Linux only: no Windows or SolidWorks was available.
    • tests/sidecar-client.test.mjs runs the Node client against the real sidecar server with fake tools. It fails 5/5 on 0.2.129 and passes 5/5 here.
    • Stubbed fetch captured the request bodies for Fable 5.1, Opus 5.5 and GPT-6 Astra.
    • The heartbeat fix was checked under Python 3.10.
    • The PID watchdog exits a busy sidecar 7 s after its parent dies.
  • Not verified: the Windows branch of the PID watchdog, and real SolidWorks sessions. These need a test of the packaged build.

Checklist

  • Code follows project conventions
  • Self-reviewed
  • Docs updated (if needed)
  • CHANGELOG.md updated (if needed)

🤖 Generated with Claude Code

https://claude.ai/code/session_015nT6r6mw2MoBrJhEfWJbQD


Generated by Claude Code

…ps; recommend GPT-6 Astra / Claude Opus 5.5 / Fable 5.1

Model support (every current Claude model and GPT-6 Astra were unusable):
- Anthropic: stop sending temperature to Fable 5/5.1, Opus 4.7+, Sonnet 5+ (400 on
  every request); map reasoning level to adaptive thinking + output_config.effort
  instead of budget_tokens; never send thinking.disabled to always-thinking models;
  retry once without sampling/thinking/effort on a 400 naming them; surface
  stop_reason "refusal" instead of an empty reply.
- OpenAI: treat gpt-6* as a reasoning model (max_completion_tokens, no temperature);
  omit reasoning_effort with tools on GPT-6 chat/completions and map 'minimal' to
  'low'; recognise "... are not supported" as a reasoning-param error; Test
  connection and the vision path use max_completion_tokens for reasoning models.

Sidecar:
- The P129 stdin watchdog peeked stdin from a helper thread; after ~2 s idle,
  back-to-back requests stalled into timeouts. It now watches the parent PID.
- Requests go out one at a time (budgets start when sent); heartbeats can extend a
  call to at most 3x its budget; onStillRunning fires on heartbeats; Stop cancels
  the wait (CANCELLED); failed results are cached under op_id so the TIMEOUT
  follow-up never re-runs a half-built generator; op_ids are namespaced per run.
- Python < 3.11: catch concurrent.futures.TimeoutError in the heartbeat loop.
- Durations reach the tool row and exports; the amber busy dot is wired up.
- looksLikeQuestion judges the reply's ending (plans proceed, parameter requests stop).

Settings / config:
- Protocol switch and quick-fill apply a matching model and provider defaults.
- Number fields clamp on blur; env fallback keeps saved preferences and is never
  persisted; no config save before the stored config has loaded.
- Non-streaming requests no longer hit the 20 s connect timeout (and are not re-sent).
- Truncation counts the system prompt + tool schemas; the max-rounds summary passes
  tools (Anthropic 400s on tool blocks without tools); theme/locale IPC channels
  moved into ipc-channels.ts; ruff E702 in test_reliability.py.

Docs: presets, README / README.zh-CN, USER-GUIDE, ARCHITECTURE and .env.example
recommend gpt-6-astra (OpenAI) and claude-opus-5-5 / claude-fable-5-1 (Anthropic).

Tests: tests/llm-request.test.mjs (request bodies per model, truncation, llmFetch),
tests/sidecar-client.test.mjs (Node client vs the real sidecar server; 5/5 fail on
0.2.129), Python tests for failure caching, futures timeout, PID watchdog.
191 JS + 58 Python tests pass; typecheck, lint, ruff, compileall, renderer build OK.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015nT6r6mw2MoBrJhEfWJbQD
@raylanlin
raylanlin merged commit c332442 into master Oct 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant