Skip to content

refactor(runtime): retire unowned recovery and judgment protocols - #942

Merged
XuPeng-SH merged 1 commit into
matrixorigin:mainfrom
XuPeng-SH:refactor/runtime-responsibility-cleanup
Oct 4, 2026
Merged

XuPeng-SH merged 1 commit into
matrixorigin:mainfrom
XuPeng-SH:refactor/runtime-responsibility-cleanup

Conversation

@XuPeng-SH

@XuPeng-SH XuPeng-SH commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Retire three protocols that no longer own production execution: local recovery retry/replay bookkeeping, the unwired RuntimeEnvironment/FakeRuntime execution API, and test-only host judgment overrides. Recovery now uses the existing owned checkpoint validator; Work and skill judgment tests exercise the actual provider response boundary. Live harness follow-up also keeps command budgets out of model arguments and reconciles tool evidence by its owning run.

Related issue

Responsibility cleanup of existing runtime contracts; no new subsystem or database schema.

Change type

  • Refactor or performance improvement
  • Test
  • Documentation

User and compatibility impact

Crash recovery offers continuation after acknowledging unknown effects, or abort. It no longer claims it will replay or skip tools: the retired choices did not execute tools. Owner-scoped checkpoints, canonical journal recovery and policy/capability checks remain authoritative. Removed unwired execution APIs have no in-repository production callers. No compatibility layer or migration is added.

Bash callback requests carry a required command_timeout_cap_ms separate from the delivery deadline. CLI execution preserves original model arguments for approval and journal evidence. Missing or invalid command authority fails closed. Server terminal tool summaries remain authoritative when CLI records cover only local executions.

Architecture and complexity delta

  • Canonical owners: shared checkpoint restoration and the existing typed Work/skill provider contracts.
  • Callers and implementations were traced across CLI, runtime, services and transport before retirement.
  • Removed recovery manager/hash/retry state, FakeRuntime-only execution protocol, host judge injection, generic unused turn-intent wire protocol and their self-only tests.
  • Ordinary recovery retains bounded/degraded audit reads; crash recovery requires its complete journal window. These policies are intentionally distinct.
  • Removed argument-injected Bash controls and the host's read-only transfer flag; the existing complete tool request carries authority into the existing executor metadata.
  • Harness journal projections share run identity extraction and retain strict parameter/result conflict checks. PTY journeys wait for selectable navigation rows; the admission test holds actual terminal storage rather than racing a sleep.
  • Physical line accounting against main d5cf67352: implementation/helpers −1,378; tests −1,612; docs +12; manifests/lockfile −2. Total −2,980. File moves are not counted as reduction.

Verification

  • GPT-6-Astra xhigh independently reviewed the stages, combined diff and rebase impact; all identified issues were resolved.
  • Public entrypoints: recovery from real owner-scoped checkpoint/journal files; Server-produced checkpoint restoration; actual Work/skill judgment request construction and response parsing; authenticated SSE parsing and real CLI batch execution/callbacks.
  • Unhappy paths: wrong owner/cursor/root/control, quarantine, journal gaps/torn receipts, unknown side effects without replay, malformed judgments, disabled policy making zero provider calls, and Work scope-to-workspace requirements; invalid command budgets, actual process timeout with unchanged arguments, and conflicting/cross-run journal evidence.
  • Database verification: no schema, query or transaction changes.
  • make check passed; strict workspace nextest: 21,204 passed, 800 skipped; workspace doctests and CLI/Server/harness production builds passed. Both child transcript journeys, the terminal-storage admission barrier and executable identity regression passed.
  • CLI, Server and harness report clean build SHA 164f0b413; the live Server health check confirmed that same SHA with database and Memoria connected.
  • Live harness: 4/4 passed, with no warnings. Unchanged cases: read_only_source_review_compliance, session_tool_result_artifact_recovery, skill_discovery_uses_skill_tool, and work_auto_decomposition_journey. Execution model: deepseek-flash. All 60 deterministic criteria passed. qwen3.8-max was configured as the judge model, but these cases contain no LLM-judge criterion.
  • Initial harness was 1/4; evidence identified duplicate nested tool counts, lost Server tool names and argument-injected runtime controls. The rerun preserves the original scenario thresholds and strict evidence checks. Artifact recovery now records both bash and introspect; Skill invocation is counted once.

Final checklist

  • Tests exercise the owning entrypoints and unhappy paths.
  • Runtime lifecycle and model-routing documentation updated.
  • Diff checked for credentials, private endpoints, generated files and sensitive data.
  • Conventional Commit-style title.

@XuPeng-SH
XuPeng-SH force-pushed the refactor/runtime-responsibility-cleanup branch from abdba93 to 164f0b4 Compare October 4, 2026 16:14
@XuPeng-SH
XuPeng-SH merged commit fd40b74 into matrixorigin:main Oct 4, 2026
37 of 38 checks passed
@XuPeng-SH
XuPeng-SH deleted the refactor/runtime-responsibility-cleanup branch October 4, 2026 16:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant