Skip to content

fix(agent): recover delegation judgments and preserve execution ownership - #943

Merged
XuPeng-SH merged 6 commits into
matrixorigin:mainfrom
XuPeng-SH:fix/delegation-admission-recovery
Oct 4, 2026
Merged

XuPeng-SH merged 6 commits into
matrixorigin:mainfrom
XuPeng-SH:fix/delegation-admission-recovery

Conversation

@XuPeng-SH

@XuPeng-SH XuPeng-SH commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fix named-model delegation recovery and observation ownership without adding a second execution path.

  • Align candidate-assessment prompts and parsers: supplied tasks require explicit slot arrays; remove implicit universal-scope normalization.
  • Validate judgment responses before using the existing second-call allowance. Invalid responses can receive one corrective attempt within the existing deadline/fallback budget, never a third logical judgment call. Distinguish invalid internal evidence from valid semantic ambiguity; fail closed without inviting repeated spawn attempts.
  • Include typed pre-execution rejection facts in Introspect. Bind request identity before surface rejection and keep child callbacks out of foreground incremental execution snapshots.
  • Make requested output formats override persona/progress/summary defaults. Share child-outcome guidance across launch receipts, retaining result inspection, pagination and recovery without encouraging re-fetching sufficient automatically delivered results.
  • Add live strict-integer/independent-parent and compact-JSON cases; measure fanout primary-round overhead.

Related issue

No linked issue. Follow-up to observed named-model delegation failures.

Change type

  • Bug fix
  • Documentation
  • Refactor or performance improvement
  • Test
  • Feature
  • Build, CI, or maintenance

User and compatibility impact

Natural named-model delegation continues through the existing authenticated candidate selector. Invalid selector responses no longer masquerade as user ambiguity. Child events no longer contaminate the foreground callback snapshot. No compatibility shim, migration, configuration change, model-name alias, task-specific matcher or output-rewriting filter is added.

Architecture and complexity delta

  • Canonical owners extended: existing delegation assessment/parser, shared SSE admission hook, CLI foreground ownership predicate, Introspect projection and child-result wire guidance.
  • Existing callers checked: server judgment/fallback calls, serial and parallel CLI callback paths, spawn/fanout receipts, resident/deferred tool schemas and shared prompt assembly.
  • Removed implicit universal-slot normalization and duplicated private child-outcome guidance; replaced unconditional explanation/summary instructions.
  • No new database schema/query, registry, state machine or execution path. Fixed prompt and resident-schema budgets are unchanged. Discovery and bounded result reads remain available when genuinely needed.
  • Branch diff against rebased main (fd40b7402): 23 files, 705 insertions / 103 deletions. This is a focused correctness repair, not a claimed large-scale cleanup milestone.

Verification

CI follow-up

Current head: 4af170d1e. Fresh remote CI for this head must still finish; local gates below do not claim all remote checks are green.

Rebased onto main commit fd40b7402 (#942), preserving its whole-request SSE API and command budget; no retired recovery/judgment path is restored. Updated the CLI budget fixture to admit the actual typed request. Post-rebase CLI callback, ownership and budget gates: 6/6 PASS.

The subsequent remote run on ddaef6055 passed all jobs except one bridge E2E fixture. That fixture now supplies both invalid assessment responses permitted by bounded repair and checks the actual delegation_model_assessment_unavailable error. It retains zero children, executed: false, exact total request count, and exactly two shared assessment calls for the batch. The discovery-only fixture now explicitly denies skills, independently of the machine's local skill catalog. Post-rebase full Web agent E2E suite: 54/54 PASS (two existing ignored tests); these are deterministic provider-boundary tests, not a fresh paid-provider cohort.

Adopted review comment 4178480735: a full admission-rejection list previously hid execution failures in Introspect. The shared snapshot now fairly merges both existing error sources within the same 10-entry cap, preserves each source's recency and redaction, and avoids a duplicate health scan. Regression coverage exercises a fresh timeout after ten rejections, both categories saturated, and each category alone. Runtime Introspect gate: 19/19 PASS. Runtime lib/tests Clippy (-D warnings): PASS, 3m56s; format/diff checks pass. No additional storage or database I/O. Independent gpt-6-astra / xhigh review approved this delta.

The initial CI found three deterministic gaps across two jobs:

Follow-up commit: ddaef6055. Independent gpt-6-astra / xhigh review approved the repair. Formatting and diff checks pass. cargo clippy -p astra-tools -p astra-test-harness --lib --tests -- -D warnings: PASS (1m29s).

  • Core crates: an agent discovery summary exceeded its existing display budget, truncating the shell-sleep prohibition. All three affected source summaries now fit 177 characters; the 180-character budget and both surfaces' visible-cue checks are unchanged.
  • Core crates: the shipped fanout case's fixture omitted round count after adding its 2–4-round criterion. The success fixture now carries four rounds; zero and five explicitly fail. No criterion is skipped or relaxed.
  • Runtime: the real correlated-question/answer test still expected the old final-answer phrasing. It now checks explicit agent(wait) and runtime completion waiting while retaining queued-not-applied and exact reply-correlation assertions.

cargo nextest run -p astra-tools -p astra-test-harness -p astra-runtime --lib -E 'test(schemas::tests::) | test(criteria::tests::) | test(messaging::e2e_loop_tests::) | test(tool_registry::surface)' --test-threads 2 --status-level fail --final-status-level fail: 277/277 PASS. Fresh CI after the follow-up push remains pending; this is not a claim that all remote checks are green.

Prior live cohort and owning-boundary gates

Live-cohort candidate: 8d472a035df38ed93416149d7d1b5220bc80998b, based on main (d5cf673523cff0b0fcf871d4a9303c39e56d0c92). Follow-up CI fixes shorten discovery summaries and update deterministic fixtures/assertions, without changing execution, selectors or output-format policy.

  • cargo fmt --all -- --check and git diff --check: PASS.
  • cargo nextest run -p astra-runtime --lib -E 'test(prompts::) | test(agent_fanout) | test(fanout_) | test(wait_) | test(parent_coordination) | test(tool_registry::surface)' --test-threads 1 --status-level fail --final-status-level fail: 343/343 PASS.
  • cargo clippy -p astra-runtime -p astra-tools -p astra-turn-core --lib --tests -- -D warnings: PASS.
  • Earlier unchanged-boundary gates: selected runtime 194/194, delegation-service 27/27, Introspect 143/143, and CLI callback/snapshot 6/6; affected lint checks passed.
  • Independent design and pre-commit review: configured gpt-6-astra / xhigh; no remaining blockers. Post-cohort evidence independently checked with the same configuration.
  • Clean CLI/Server/harness build from the candidate, matching healthy server revision, isolated test workspace, serial runs and bounded build memory. Actual DeepSeek Flash parent and GLM 5.2 children; no synthetic output substitution or external judging model.

First fixed-cohort outcomes, all criteria passed with no warnings:

Live case Elapsed Primary rounds Evidence
flash_delegate_simple_glm 12.356s 4 Actual GLM child, exact integer adoption and answer
flash_scoped_child_and_parent 7.788s 3 Exact child integer, independent parent result; no file/network/discovery calls
flash_child_compact_json 9.043s 3 Exact child JSON adoption and parent JSON; no file/network/discovery calls
flash_semantic_model_reference_glm 13.768s 4 Authorized GLM binding from natural-language reference, real child inference
flash_fanout_model_default 9.503s 4 Both child identities/results adopted; discovery and start only, no get_results
flash_spawn_prohibited_model_fail_closed 5.799s 2 Rejected delegation, zero children
  • Public entrypoint: real CLI → Server admission → child execution → parent result adoption; Introspect and CLI callback paths also covered at their owning boundaries.
  • Unhappy paths: invalid-response repair/exhaustion, semantic refusal without repair, prohibited/unavailable model, callback ownership and pre-execution rejection. Existing cancellation, wait, partial-result and pagination coverage remains intact.
  • Database verification: live tests used a real MatrixOne-backed test server through existing APIs. No database changes or added reads/writes; no claim of a new scale benchmark.

Evidence limits

The earlier strict-integer failure is retained; a later pass does not prove guaranteed LLM compliance. Two cohort cases still performed avoidable discovery despite resident agent availability. Fanout improved from a prior single sample of 5 rounds / 12.928s to 4 rounds / 9.503s, but this is not a controlled latency distribution or general cost-saving claim. Logical usage coverage is complete in these reports; complete child cache coverage and historical billing reconciliation are not established. No full-workspace CI pass or configured-Jev live cohort is claimed. Reproduction instructions and case definitions are checked in; private traces, credentials and local artifacts are not.

Final checklist

  • Tests updated at the owning layers and real public entrypoints exercised.
  • Design and harness documentation updated.
  • Diff checked for credentials, private URLs, customer data and generated files; none included.
  • Conventional Commit-style PR title.

@XuPeng-SH XuPeng-SH left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deep review of d5cf673523cff0b0fcf871d4a9303c39e56d0c92..ddaef6055ce07096421a7e1c3e321038aa806003, covering all four commits and the combined diff (22 files, +606/-96).

Recommendation: fix the P2 below before merging.

One actionable finding: the new Introspect error projection can completely hide recent execution failures after ten admission rejections have accumulated. The inline comment identifies the truncation point, trigger and regression coverage needed. This is a regression in the diagnostic view, not a claim that execution failures disappear from the underlying health tracker.

Other review conclusions:

  • Invalid-response correction, configured fallback and primary deadline retry share the existing second-call allowance. The added branches do not introduce a third call within one judgment invocation, and cancellation, lease-loss and remaining-budget checks remain in place.
  • Exhausted invalid initial candidate assessments become internal unavailability rather than user ambiguity. Repeated spawn attempts under the same authenticated intent reuse that failed assessment.
  • Requiring explicit slot arrays aligns the prompt examples with the parser and removes implicit universal-scope expansion.
  • The CLI foreground filter preserves child callback settlement. The updated serial/parallel tests check both foreground ownership and the run identities of the actual result callbacks.
  • Output-format guidance remains a prompt contract, without a task-specific answer rewriter. The new live cases require actual child model inference and adoption of the child result.
  • The implementation extends existing owners without a new execution path, registry, database query or transaction. The net +510 lines have concrete correctness and verification purposes. A malformed judgment can now incur one additional auxiliary inference; that cost is bounded.

Verification: git diff --check passed. At review time, CI lint/check and turn-core/services/plan had passed; runtime, CLI and other lanes were still running. No local Rust tests were executed because this environment has no Rust toolchain. Author-reported local and live-harness results were not independently rerun. The P2 follows directly from the iterator ordering and truncation in the production snapshot builder; the existing new test covers only one rejection with an empty execution-error history.

Comment on lines +1559 to +1560
.chain(state.turn_guard.health.recent_errors(10))
.take(10)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve execution errors when the rejection list reaches its cap

This appends execution failures after up to ten admission rejections and then truncates the combined list to ten. Once the current turn has accumulated ten Rejected records, even a later Bash timeout or file-operation failure is omitted from introspect(facet="errors"): health.recent_errors(10) contains the fresh failure, but none of its entries can survive this take(10). Subsequent successful calls do not remove the earlier rejection records, so the diagnostic view can keep showing stale admission problems while hiding the failure the agent needs to recover from.

Use a bounded merge that cannot let one category completely starve the other, and add a snapshot-level regression with ten earlier rejections followed by a new execution failure. Assert that the new failure remains visible while the projection stays bounded.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adopted in 4af170d. The shared Introspect snapshot alternates existing execution and admission error sources within the unchanged ten-entry cap, preserving source-local recency, redaction, and unknown admission timestamps. It reuses the health errors collected for alerts rather than scanning twice; no storage or database I/O added. Snapshot regression covers ten old rejections plus a fresh timeout, both categories saturated, and each category alone. Introspect 19/19 and runtime lib/tests Clippy pass; independent gpt-6-astra / xhigh review approved. Rebased onto fd40b74 and updated this PR.

@XuPeng-SH
XuPeng-SH force-pushed the fix/delegation-admission-recovery branch from ddaef60 to 4af170d Compare October 4, 2026 17:00
@XuPeng-SH
XuPeng-SH merged commit 2e36e7c into matrixorigin:main Oct 4, 2026
38 checks passed
@XuPeng-SH
XuPeng-SH deleted the fix/delegation-admission-recovery branch October 4, 2026 17:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant