From 57f3322bc07693d4c5cd6ef8878edb9d2834187e Mon Sep 17 00:00:00 2001 From: LBX154 <145820328+lbx154@users.noreply.github.com> Date: Wed, 30 Sep 2026 04:39:43 -0700 Subject: [PATCH 1/3] Consolidate the skill library's duplicated clusters into one file each The usage audit behind the preceding prune found that the largest problem in the built-in library was repetition, not gaps: the figure guidance, the diagnosis ladder and the method-card contract were each written three to five times across sibling files, some with verbatim-identical closing paragraphs and palettes, and every role banner restated its method file. Agents that opened one file of a cluster rarely opened the others, and the required playbooks repeated the clusters again. This commit merges each cluster into the file the audit named as its home and removes the absorbed originals: 40 files (2,076 lines) become sections of 21 rewritten or new files, and the library goes from 98 skill files and 7,488 lines to 61 files and 6,841 lines. Where the absorbed text asserted a number instead of explaining a judgement (three figures, three diagnosis attempts, a 45-60 minute time box, eight mandatory environment sections, one visual fix, two tasks for a team, missing anchors meaning NOT_IMPLEMENTED), the merged text says what the reader should weigh. Three files the audit listed for merging stay separate because runtime code injects them rather than agents reading them: the RL collapse diagnosis (injected into RL training supervision), the curator contract (the curator's prompt) and, in the other direction, the math Manager and Scientist banners now read the planning file, whose closing sections address them. Loader path lists, the math banner mapping, the figure code paths, docs links and the tests that named the absorbed files follow the new layout; seeded copies of the absorbed files retire through the existing table. Co-Authored-By: Claude Fable 5.1 --- argus/builtin_skills/agent-team-lead.md | 37 +- .../curator/argus-curator-role.md | 2 +- .../engineer/environment-readiness.md | 150 ---- .../engineer/presentation-master.md | 57 -- .../engineer/skill-authoring-guide.md | 145 ++- .../engineer/stale-world-model.md | 42 - .../engineer/study-and-curate.md | 34 - .../builtin_skills/engineer/wiki-collector.md | 49 - .../manager/argus-manager-role.md | 32 - .../manager/evidence-based-stage-decision.md | 51 +- .../planner/argus-planner-role.md | 27 - .../dependency-aware-task-decomposition.md | 57 +- .../project-venv-package-management.md | 35 +- .../reviewer/argus-reviewer-role.md | 66 +- .../reviewer/guiding-the-engineer.md | 69 -- argus/builtin_skills/reviewer/wiki-curator.md | 45 - .../stale-blocker-verification-probe.md | 112 ++- argus/life/supervisor/_helpers.py | 2 +- argus/skills/builtins.py | 135 ++- .../kernel-benchmark-measurement-integrity.md | 357 ++++++-- .../kernel-environment-first-engineering.md | 25 - .../engineer/kernel-optimization-knowledge.md | 129 --- .../reviewer/kernel-engineering-review.md | 25 - .../skills/engineer/learning-curation.md | 18 - .../learning/skills/learning-curation.md | 81 ++ .../skills/reviewer/curation-review.md | 53 -- .../engineer/math-research-execution.md | 56 +- .../skills/manager/math-research-manager.md | 13 - .../skills/planner/math-research-planning.md | 35 +- .../scientist/math-research-adaptation.md | 12 - .../scientist/math-research-distillation.md | 12 - argus/verticals/math/stages.py | 8 +- argus/verticals/research/figure_tool.py | 8 +- argus/verticals/research/pipeline_figure.py | 3 +- argus/verticals/research/prompt_policy.py | 14 +- .../skills/engineer/citation-check.md | 24 - .../engineer/claims-against-evidence.md | 51 -- .../skills/engineer/delta-on-reference.md | 54 -- .../skills/engineer/executable-spec.md | 84 -- .../engineer/framework-stand-up-pilot.md | 58 -- .../hypothesis-implementation-contract.md | 39 - .../skills/engineer/implementation-brief.md | 63 -- .../infrastructure-landscape-survey.md | 73 -- .../research/skills/engineer/method-card.md | 418 +++++++-- .../skills/engineer/method_card_template.md | 43 - .../skills/engineer/paper-chart-styling.md | 61 +- .../engineer/paper-exemplar-pdf-learning.md | 25 - .../engineer/paper-framework-figure-studio.md | 834 +++++++++--------- .../engineer/paper-illustration-image2.md | 40 - .../engineer/paper-infrastructure-review.md | 22 - .../skills/engineer/recipe-anchored-tuning.md | 58 -- .../skills/engineer/research-grind.md | 211 ++++- .../research-results-analysis-and-figures.md | 70 -- .../engineer/research-visualization-router.md | 117 --- .../skills/engineer/result-to-claim.md | 33 - .../engineer/semantic-scholar-search.md | 52 -- .../skills/engineer/sources-and-citations.md | 120 +++ .../skills/engineer/suspect-the-setup.md | 151 ---- .../engineer/training-infrastructure-guide.md | 42 - .../engineer/training-infrastructure.md | 186 ++++ .../skills/engineer/venue-format-preflight.md | 76 +- .../skills/engineer/venue-paper-drafting.md | 43 +- .../skills/engineer/write-for-review.md | 47 - .../skills/research-experiment-playbook.md | 66 +- .../skills/research-paper-playbook.md | 21 +- .../skills/research-review-playbook.md | 7 +- .../reviewer/experiment-results-review.md | 130 ++- .../reviewer/infrastructure-choice-review.md | 46 - .../skills/reviewer/reading-the-evidence.md | 56 -- .../software-change-implementation.md | 21 +- .../planner/software-project-grounding.md | 36 - docs/CORE_CONCEPTS.md | 4 +- docs/WHAT_ARGUS_GREW.md | 4 +- ...exploration-without-local-hill-climbing.md | 3 +- ...ation-without-local-hill-climbing.zh-CN.md | 3 +- docs/failure-modes-and-fixes.md | 6 +- docs/failure-modes-and-fixes.zh-CN.md | 6 +- frontend/web/src/lib/bundledSkillChinese.json | 441 ++------- tests/apps/test_cli_parser.py | 12 +- tests/life/test_planner_task_sanitization.py | 2 +- tests/skills/test_builtins_seeding.py | 25 +- tests/skills/test_fixed_claim_policy.py | 42 +- .../test_infrastructure_procedure_skills.py | 65 +- .../test_kernel_engineering_vertical.py | 12 +- tests/skills/test_math_vertical.py | 3 - tests/skills/test_method_card_policy.py | 20 +- tests/skills/test_paper_chart_style.py | 28 +- tests/skills/test_research_learning_prompt.py | 2 +- tests/skills/test_research_svg_pipeline.py | 1 - .../test_research_visualization_router.py | 113 ++- .../skills/test_visual_authoring_builtins.py | 52 +- tests/test_aris_adapted_skills.py | 13 +- tests/test_manager_skill_wiring.py | 12 +- tests/test_stage_authority_prompts.py | 2 +- tests/test_stale_world_model_skill.py | 2 +- 95 files changed, 2749 insertions(+), 3598 deletions(-) delete mode 100644 argus/builtin_skills/engineer/environment-readiness.md delete mode 100644 argus/builtin_skills/engineer/presentation-master.md delete mode 100644 argus/builtin_skills/engineer/stale-world-model.md delete mode 100644 argus/builtin_skills/engineer/study-and-curate.md delete mode 100644 argus/builtin_skills/engineer/wiki-collector.md delete mode 100644 argus/builtin_skills/manager/argus-manager-role.md delete mode 100644 argus/builtin_skills/planner/argus-planner-role.md delete mode 100644 argus/builtin_skills/reviewer/guiding-the-engineer.md delete mode 100644 argus/builtin_skills/reviewer/wiki-curator.md delete mode 100644 argus/verticals/kernel_engineering/skills/engineer/kernel-environment-first-engineering.md delete mode 100644 argus/verticals/kernel_engineering/skills/engineer/kernel-optimization-knowledge.md delete mode 100644 argus/verticals/kernel_engineering/skills/reviewer/kernel-engineering-review.md delete mode 100644 argus/verticals/learning/skills/engineer/learning-curation.md create mode 100644 argus/verticals/learning/skills/learning-curation.md delete mode 100644 argus/verticals/learning/skills/reviewer/curation-review.md delete mode 100644 argus/verticals/math/skills/manager/math-research-manager.md delete mode 100644 argus/verticals/math/skills/scientist/math-research-adaptation.md delete mode 100644 argus/verticals/math/skills/scientist/math-research-distillation.md delete mode 100644 argus/verticals/research/skills/engineer/citation-check.md delete mode 100644 argus/verticals/research/skills/engineer/claims-against-evidence.md delete mode 100644 argus/verticals/research/skills/engineer/delta-on-reference.md delete mode 100644 argus/verticals/research/skills/engineer/executable-spec.md delete mode 100644 argus/verticals/research/skills/engineer/framework-stand-up-pilot.md delete mode 100644 argus/verticals/research/skills/engineer/hypothesis-implementation-contract.md delete mode 100644 argus/verticals/research/skills/engineer/implementation-brief.md delete mode 100644 argus/verticals/research/skills/engineer/infrastructure-landscape-survey.md delete mode 100644 argus/verticals/research/skills/engineer/method_card_template.md delete mode 100644 argus/verticals/research/skills/engineer/paper-exemplar-pdf-learning.md delete mode 100644 argus/verticals/research/skills/engineer/paper-illustration-image2.md delete mode 100644 argus/verticals/research/skills/engineer/paper-infrastructure-review.md delete mode 100644 argus/verticals/research/skills/engineer/recipe-anchored-tuning.md delete mode 100644 argus/verticals/research/skills/engineer/research-results-analysis-and-figures.md delete mode 100644 argus/verticals/research/skills/engineer/research-visualization-router.md delete mode 100644 argus/verticals/research/skills/engineer/result-to-claim.md delete mode 100644 argus/verticals/research/skills/engineer/semantic-scholar-search.md create mode 100644 argus/verticals/research/skills/engineer/sources-and-citations.md delete mode 100644 argus/verticals/research/skills/engineer/suspect-the-setup.md delete mode 100644 argus/verticals/research/skills/engineer/training-infrastructure-guide.md create mode 100644 argus/verticals/research/skills/engineer/training-infrastructure.md delete mode 100644 argus/verticals/research/skills/engineer/write-for-review.md delete mode 100644 argus/verticals/research/skills/reviewer/infrastructure-choice-review.md delete mode 100644 argus/verticals/research/skills/reviewer/reading-the-evidence.md delete mode 100644 argus/verticals/software/skills/planner/software-project-grounding.md diff --git a/argus/builtin_skills/agent-team-lead.md b/argus/builtin_skills/agent-team-lead.md index 7e89add78..1057ecff1 100644 --- a/argus/builtin_skills/agent-team-lead.md +++ b/argus/builtin_skills/agent-team-lead.md @@ -13,7 +13,7 @@ Use a team only to parallelize several genuinely independent tasks. The lead wri Every role may discover this Skill, but it does not erase role boundaries: -- Manager recognizes a Team request and preserves it in the mission handoff. +- Manager recognizes a Team request and preserves it in the mission it passes on. - Planner delegates Team formation unchanged; Planner does not infer availability from its own role-specific Skill directory. - Engineer or an explicitly assigned lead forms and operates the Team. @@ -29,16 +29,18 @@ further decomposition to the parent lead. Do not create another Team, change the nesting switch, or bypass the runtime admission check. Explicitly authorized nested workflows must be configured by the host before execution, not by a child. -Form a team only when all of these hold: +A team pays for itself when the work splits into pieces whose files and +responsibilities do not overlap, each piece carries its own completion +evidence, and each piece is large enough that coordinating it — writing its +objective, reviewing its result, merging it — costs less than doing it in +sequence. Two such pieces are enough; twenty pieces that share a file or wait +on one another are not a team but a queue with extra bookkeeping. Small, +sequential, tightly coupled or same-file work stays solo. Provider, compute and +hardware capacity must be able to serve the width you ask for, and the width, +timeout and total spend must fit the current operator budget; widening a pool +is not a way around the shared budget or a failed admission. -- At least two tasks can make useful progress concurrently. -- Their writable paths do not overlap. -- Each task has its own completion evidence. -- Provider, compute, and hardware capacity can support the requested width. -- The width, timeout and total spend fit the current operator budget; widening - a pool is not a way around the shared budget or a failed admission. - -Stay solo for small, sequential, tightly coupled, or same-file work. `owns_paths` records the lead's partition for review and prior-work inheritance; it is not a filesystem sandbox, so do not form a team when prompt-level ownership is insufficient. +`owns_paths` records the lead's partition for review and prior-work inheritance; it is not a filesystem sandbox, so do not form a team when prompt-level ownership is insufficient. ## Form the rolling backlog Use `python -m argus.tools.team`. @@ -52,14 +54,15 @@ Use `python -m argus.tools.team`. `form --root --team-id --cwd --mission "" --tasks tasks.jsonl`. 3. Set deliberate capacity with: `pool-set --root --width --state running`. - The width is clamped to what the host can serve: with a provider concurrency - limit, one slot always stays with the lead, so asking for more than the - ceiling grants the ceiling. + The host clamps the width to what it can serve. Under a provider concurrency + limit the lead keeps one slot for its own reading, synthesis and review, so + the pool can never take the whole limit; asking for more than the remainder + grants the remainder. 4. Inspect progress with `status --root ` and read landed `shards/*.jsonl` plus `leaderboard.json`. 5. A task waiting on a real operator-owned decision is `blocked`, retains its owner and question, and is not retried. After the operator answers, run `resume --root --task-id --answer ""` to requeue it with that answer. 6. Refresh or extend the backlog with `form`. Re-forming claimed, running, or blocked work preserves its lifecycle state; re-forming a done or failed task deliberately reopens it. 7. Once the required results are ready, set `pool-set --state draining`, read - their final shards, and synthesize and verify the canonical artifact. For + their final shards, and synthesize and verify the canonical result. For alternative candidates, the lead may use a completed, independently reviewed result without waiting for every optional alternative. Remaining workers stay confined to their assigned private outputs and may not alter the chosen @@ -67,14 +70,14 @@ Use `python -m argus.tools.team`. Reviewer while the Curator reaps them. After all teammates settle, run `dissolve --root ` at a normal status check; optional candidate cleanup does not block review. Reviewer validates the durable project - artifacts and does not need a live Team runtime. + files and does not need a live Team runtime. The lead never manually spawns, claims, waits for, reassigns, or kills teammates. Those are Curator responsibilities. ## Task-objective contract Every task must state: -- the objective, and the separately checkable done condition as the task's `acceptance_check`, naming exactly once the single subject the task must move (the claim, kernel, or artifact id): a vertical's per-mission context block is resolved from the first task field that names exactly one, and a field naming two resolves to none; +- the objective, and the separately checkable done condition as the task's `acceptance_check`, naming exactly once the single subject the task must move (the claim, the kernel, or the id of the result it produces): a vertical's per-mission context block is resolved from the first task field that names exactly one, and a field naming two resolves to none; - the only paths it may modify; - the required result shard or output file; - the real measurement or verification command; @@ -84,7 +87,7 @@ A teammate runs one normal Engineer→Reviewer mission and exits. The Curator th ## Result and synthesis rules -- Teammates emit task-local artifacts and one shard; they never write the shared leaderboard or the lead's canonical merged artifact. +- Teammates emit task-local outputs and one shard; they never write the shared leaderboard or the lead's canonical merged result. - The Curator is the single writer for pool lifecycle and deterministic leaderboard folding. - The lead accepts measured, task-valid results only and is the single writer of the canonical synthesis. - Every teammate result passes its own Reviewer; the final synthesis still passes the mission Reviewer. diff --git a/argus/builtin_skills/curator/argus-curator-role.md b/argus/builtin_skills/curator/argus-curator-role.md index 3269aa5ea..b27569480 100644 --- a/argus/builtin_skills/curator/argus-curator-role.md +++ b/argus/builtin_skills/curator/argus-curator-role.md @@ -7,7 +7,7 @@ description: "The daemon-resident agent that maintains an agent team's pool and Team Curator ## Description -You are the **Curator** of an Argus agent team — the persistent, daemon-resident agent that maintains the teammate pool and the **leaderboard**, and distills a short forward **strategy** the next teammates inherit. You are NOT an engineer: you never write or optimize the artifact yourself. (Distinct from the `wiki-curator` reviewer skill.) +You are the **Curator** of an Argus agent team — the persistent, daemon-resident agent that maintains the teammate pool and the **leaderboard**, and distills a short forward **strategy** the next teammates inherit. You are NOT an engineer: you never write or optimize the artifact yourself. Two cadences run your work: - **Mechanical (high-frequency, no LLM):** keep N teammates in flight, reap finished/wedged ones, and re-fold the leaderboard from result shards. This is deterministic code — not your judgment. diff --git a/argus/builtin_skills/engineer/environment-readiness.md b/argus/builtin_skills/engineer/environment-readiness.md deleted file mode 100644 index 68f956076..000000000 --- a/argus/builtin_skills/engineer/environment-readiness.md +++ /dev/null @@ -1,150 +0,0 @@ ---- -name: "Environment Readiness Check" -description: "Verify the project environment, public data/evaluator, dependencies, storage, and only the compute/API resources an experiment actually uses before producing evidence." ---- - -# Environment Readiness Check - -## Purpose - -Prevent invalid or wasted runs without assuming every AI research project uses -CUDA, Hugging Face models, an LLM API, or a training framework. - -Run these checks before the first real benchmark/evidence call and before each -substantively different pilot, full, or ablation launch. - -## Applicability rule - -Verify only resources the experiment actually uses. Skip irrelevant sections; -state `NOT_APPLICABLE` only when a supplied contract asks for it. Do not fabricate -a CUDA, model-weight, or API dependency to satisfy the checklist. Follow Project -Environment and Dependencies for the existing interpreter, lockfile and cache policy. - -## Required checks - -### 1. Project environment - -- Record the interpreter/runtime/compiler executable and version. -- Confirm dependencies import or execute from the project environment rather - than the Argus framework environment. -- For Python projects with a venv, prefer `./.venv/bin/python` on POSIX or - `.\.venv\Scripts\python.exe` on Windows. -- Record package-lock and configuration versions when they are claim-relevant. - -POSIX-shell example: - -```bash -pwd -command -v python || true -python -V || true -test -x .venv/bin/python && .venv/bin/python -V || true -``` - -Windows PowerShell example (compatible with Windows PowerShell 5.1): - -```powershell -Get-Location -Get-Command python -ErrorAction SilentlyContinue -python -V -if (Test-Path -LiteralPath '.\.venv\Scripts\python.exe') { - & '.\.venv\Scripts\python.exe' -V -} -``` - -### 2. Public evidence source - -- Verify the public benchmark/dataset/task suite or official evaluation release - can be retrieved or is present locally. -- Record official URL/repository, version/commit, split/cohort, license/access - condition, checksum when practical, and any filtering/conversion script. -- Confirm synthetic/generated diagnostics are labeled separately from public - evidence. - -### 3. Evaluator or analysis path - -- Execute the official evaluator, metric implementation, statistical analysis, - theorem checker, profiler, simulator, or domain-native verifier on a tiny - known input. -- Confirm outputs are non-empty and semantically plausible. -- For custom evaluators, compare against an official/reference implementation - on at least one shared example. - -### 4. Compute backend - -Choose the applicable branch: - -**CPU / compiler / systems** - -- Record CPU/runtime/compiler/OS details relevant to the measurement. -- Verify required binaries, permissions, clocks/affinity policy, and timing - method where applicable. - -**GPU / accelerator** - -- Confirm the allocated devices match the operator/runtime allocation. -- Verify the framework sees the expected devices and has enough memory for the - planned smoke run. -- Record driver/runtime/framework versions when they affect correctness or - performance. - -Example for a GPU experiment: - -```bash -nvidia-smi --query-gpu=index,name,memory.free --format=csv,noheader -./.venv/bin/python - <<'PY' -import torch -print(torch.__version__, torch.cuda.is_available(), torch.cuda.device_count()) -PY -``` - -On Windows PowerShell, use the project interpreter and `-c` instead of a -POSIX heredoc: - -```powershell -nvidia-smi --query-gpu=index,name,memory.free --format=csv,noheader -& '.\.venv\Scripts\python.exe' -c 'import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.device_count())' -``` - -GPU checks are `NOT_APPLICABLE` for CPU-only, API-only, theoretical, or -non-accelerator work. - -### 5. Models and frameworks - -- Import only the frameworks selected by the plan. -- Verify required checkpoints/assets are present and match the declared - revision. -- HF/Torch cache checks apply only when the run uses those ecosystems. -- A custom runtime/trainer/evaluator is allowed when required by the research; - verify it against a trusted reference rather than rejecting it by category. - -### 6. External APIs - -- Test only routes the experiment will call. -- Confirm one minimal non-empty response without printing credentials, private - endpoints, or raw capability-vault contents. -- API checks are `NOT_APPLICABLE` when the experiment is fully local. - -### 7. Storage and outputs - -- Confirm enough disk space for expected outputs/checkpoints. -- Create the run directory and verify it is writable. - -### 8. Cancellation and observability - -For long-running work: - -- verify status/progress/log paths; -- verify the cancellation mechanism or scheduler stop path; -- confirm the worker reports an initial heartbeat. - -Short deterministic commands may simply capture stdout/stderr. - -Report only a concrete blocker and the decisive command or output through the -normal Engineer response. Do not create a separate preflight file. - -## Reviewer hook - -The Reviewer should keep the stage open when an applicable readiness check is -missing, stale, contradictory, or failed. The Reviewer must not require CUDA, -HF caches, base-model weights, or API calls for an experiment that does not use -them. diff --git a/argus/builtin_skills/engineer/presentation-master.md b/argus/builtin_skills/engineer/presentation-master.md deleted file mode 100644 index eab043a4f..000000000 --- a/argus/builtin_skills/engineer/presentation-master.md +++ /dev/null @@ -1,57 +0,0 @@ ---- -name: "Editable Presentations with PPT Master" -description: "创建或编辑可编辑的演示文稿和模板。 Use the installed PPT Master toolkit for editable PPTX delivery, honoring the requested format and existing authorization; research figure policy comes from the active research vertical." ---- - -# Editable Presentations with PPT Master - -Use when the user needs editable PowerPoint slides or templates. Ordinary measured -charts and Markdown diagrams can use simpler tools when that satisfies the request. -For paper-facing figures, first read the active research vertical's figure routing -Skill; this global adapter does not decide a Method D/Method B research policy. - -## Locate the supported toolkit - -```bash -PPT_MASTER_ROOT="${ARGUS_SKILL_HOME:-$HOME/.argus-skill}/tools/ppt-master" -SKILL_DIR="$PPT_MASTER_ROOT/skills/ppt-master" -"${ARGUS_SKILL_BIN:-argus}" --ppt-master-status -``` - -Use the revision managed by Argus. A failed status check may indicate a missing -installation, a dependency problem or a modified checkout: report the actual cause. -Use `argus --install-ppt-master` only within existing operator authorization; -do not silently clone, update or replace the shared toolkit to complete a slide. - -Read `$SKILL_DIR/SKILL.md`, its `workflows/routing.md`, then only the selected route's -required references. Reuse the installed Generate PPTX, Create Template, Fill Native -PPTX or Enhance Native PPTX workflow instead of inventing a second toolkit. - -Run toolkit scripts through the supplied framework interpreter: - -```bash -"${ARGUS_SKILL_PYTHON:-python3}" "$SKILL_DIR/scripts/