From 1c66f9ddf9d26a39ac2e8d0526f9ccc6db518c8f Mon Sep 17 00:00:00 2001 From: LBX154 <145820328+lbx154@users.noreply.github.com> Date: Wed, 30 Sep 2026 03:32:03 -0700 Subject: [PATCH 1/2] Add the getting-started, best-practices and vertical-building guides Three guides under docs/, linked from a new Tutorials entry in the README: - getting-started.md: from an empty machine to a finished first task by the command line, the web UI and the desktop app. Every command was run against this checkout (argus 0.1.8); the example runs are real projects from 2026-09-30 with their times and costs taken from their logs. - best-practices.md: how to write an objective, how to change direction while a project runs, how the backend and model are resolved per role, how spend is bounded, and which tasks suit Argus. Each point carries an example from one of three real projects (the GPU roofline campaign, the damped-oscillator task, the fresh-install trial); numbers are measurements, not rules. The spend section notes that the default unpriced-cost policy no longer stalls a fresh subscription install since #179 and #183. - building-a-vertical.md: a small real vertical, lab_notebook, from the contract's names and two built-ins to learn from, through stages, checklists and skills, to packaging, a local catalog, installing it and publishing to the store. The finished example lives under examples/verticals/ with a helper that builds the zip and catalog; the guide's contract check and install steps were run. Co-Authored-By: Claude Fable 5.1 --- README.md | 6 + docs/best-practices.md | 328 +++++++++++++ docs/building-a-vertical.md | 431 ++++++++++++++++++ docs/getting-started.md | 419 +++++++++++++++++ .../argus_verticals/lab_notebook/__init__.py | 5 + .../skills/engineer/measurement-record.md | 20 + .../skills/reviewer/measurement-review.md | 16 + .../argus_verticals/lab_notebook/stages.py | 81 ++++ examples/verticals/build_local_catalog.py | 97 ++++ 9 files changed, 1403 insertions(+) create mode 100644 docs/best-practices.md create mode 100644 docs/building-a-vertical.md create mode 100644 docs/getting-started.md create mode 100644 examples/verticals/argus_verticals/lab_notebook/__init__.py create mode 100644 examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md create mode 100644 examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md create mode 100644 examples/verticals/argus_verticals/lab_notebook/stages.py create mode 100644 examples/verticals/build_local_catalog.py diff --git a/README.md b/README.md index 37e397784..e993d7284 100644 --- a/README.md +++ b/README.md @@ -107,6 +107,12 @@ maintainers for the latest code.

Community Group 1 is full. Please join Group 2.

+## Tutorials + +- **[Getting started](docs/getting-started.md)** — install, first task, three ways in: command line, web UI, desktop app. +- **[Best practices](docs/best-practices.md)** — writing an objective, changing direction mid-run, choosing the model and backend, controlling spend, what suits Argus. +- **[Building a vertical](docs/building-a-vertical.md)** — a small real vertical from stages and skills to the Vertical Store. + ## Quick Install Choose the section for your operating system. Do not mix commands between diff --git a/docs/best-practices.md b/docs/best-practices.md new file mode 100644 index 000000000..ce1ee008f --- /dev/null +++ b/docs/best-practices.md @@ -0,0 +1,328 @@ +# Argus best practices + +This page explains how to get good work out of Argus and why each +recommendation holds. The reasons come from the code paths that read your +input; the examples come from three real projects run on one machine on +2026-09-30 (a GPU roofline research campaign, a damped-oscillator task, and a +fresh-install trial), with times and costs taken from their logs. Nothing here +is a quota; where a number appears it is a measurement, not a rule. + +Read [getting started](getting-started.md) first if you have not run a task +yet. + +## Writing an objective + +The objective is read by four different consumers, and writing for all of +them is what makes an objective good: + +1. **The Manager** decides from your text alone whether this is a + conversation, a small local job it can do itself, or team work for the + Planner, Engineer and Reviewer. It also decides the lifetime: finite, + casually worded work defaults to *bounded* (stop when done); only text that + expresses ongoing intent becomes a *standing* campaign that keeps generating + work. +2. **The Manager again** picks the vertical (research, software, math, ...) + from the objective; there is no keyword classifier and no flag to force it, + so name the kind of work in plain words. +3. **The Planner** decomposes it into tasks, each with an acceptance check. + What you leave unsaid, it will decide for you. +4. **The Reviewer** judges every round against it and against the vertical's + stage checklist. The Reviewer runs read-only and cannot ask the Engineer; + it can only check what the objective and the work directory let it check. + +So an objective should say what must be true at the end, what evidence you +will accept, what the scope and constraints are, and what form the result +takes. Compare the two objectives that ran on this machine. + +The campaign objective (`objective-roofline.txt`, project `s-02d3c282`): + +> 做一项小规模但真实的研究:在本机 4 块 NVIDIA RTX A6000 上,PyTorch 方阵乘法的实际吞吐(TFLOPS)随矩阵规模(512 到 16384)和精度(FP32、TF32、BF16、FP16)如何变化,与理论峰值的差距能否用一个简单的 roofline 式模型(计算强度与显存带宽)解释。要求:真实实验、每种配置多次重复取中位数并报告离散度,拟合模型并给出误差,画图,最后写成一篇 4 页以内的短论文(含方法、结果、讨论、复现说明),并通过内部评审。所有数字必须来自本机真实运行,不允许估算或编造。 + +Every clause did work. "小规模但真实的研究" and "短论文" put it in the `research` +vertical (`pipeline : vertical=research` in the status). The hardware and the +ranges fixed the scope, so the Planner did not have to guess a sweep. "多次重复取中位数并报告离散度" +and "拟合模型并给出误差" became things the Reviewer could check: the draft's +`results.tex` reports 17,280 samples, medians, and the fit's error on a held-out +confirmation set. "通过内部评审" is why a `paper` round came back +`replan_requested` instead of being accepted on the Engineer's word. And +"所有数字必须来自本机真实运行,不允许估算或编造" is the sentence that lets the Reviewer +refuse a plausible-looking number that has no raw file behind it. + +The small objective (`objective-oscillator.txt`, project `s-4a34a674`): + +> 用 numpy 数值求解一维阻尼谐振子 x'' + 2γx' + ω0² x = 0(取 ω0=2π, 分别取欠阻尼 γ=0.5、临界 γ=2π、过阻尼 γ=10),与解析解逐点比较给出最大误差,画出三条曲线,把方法、数值结果表和图写进本目录的 README.md。所有数字必须来自真实运行。 + +It names the equation, the three parameter values, the comparison ("与解析解逐点比较给出最大误差"), +the deliverable and its location. Because the acceptance check was in the +text, the task ran once, was sent back once, and was done after 12 minutes and +$0.91. + +The browser message from the trial (project `s-67fb6d62`) shows the casual end +of the scale: + +> 帮我在这台机器的 GPU 上测一下 PyTorch 矩阵乘在 fp16 和 fp32 下的 TFLOPS,写个脚本跑出真实数字,整理成表格。 + +"写个脚本跑出真实数字" and "整理成表格" were enough for a bounded, self-contained job +that replied with a per-GPU table three minutes later. What it did not say +(matrix sizes, repeats) the Engineer chose: 4096/8192/16384 and seven timed +batches. If those choices matter to you, say them. + +Things to avoid in an objective, and why: + +- **Work that needs your credentials, your money, or an irreversible or + public action.** Those are authority boundaries: whatever the autonomy + setting, a task that reaches one stops and asks. Put the credential in + place first (the cockpit accepts credentials and stores them in the + capability vault) or keep the objective short of the boundary. +- **A goal with no checkable end.** "Make it better" leaves the Reviewer + nothing to refuse, so rounds continue until a round cap or you intervene. +- **Numbers you already know the answer to.** The roles will try to + reproduce them; that is the point, and it costs rounds. + +On the command line the objective must travel with `--continuous` +(`--objective` alone is refused), and `--objective-file` keeps long text out of +shell quoting. Add `--bounded` when the objective is finite; without it the +worker keeps proposing follow-up work after the project is declared done. + +## Changing direction while it runs + +Argus gives you three ways to speak to a running project. They are not +interchangeable, because they enter at different points of the loop. + +| You want to | Command line | Cockpit | Web API | What happens | +|---|---|---|---|---| +| Add guidance the next round should read | `argus --notify "text"` | `/nudge text` | `POST /api/projects/{sid}/nudge` | the text is queued in the project's durable inbox and spliced into the next Engineer round's prompt as operator guidance | +| Hold guidance until a later stage | `argus --notify "text" --notify-stage paper` | | | same, delivered only when the project reaches that stage (aliases such as `writeup` are canonicalized; unknown stages are refused) | +| Change the direction of the current task, pause or abort | (type it) | type it, or `/abort` | `POST /api/projects/{sid}/message` | the Manager classifies the message: `STEER` records a new direction for the active task, `PAUSE` stops the campaign, `ABORT` ends the task, `NO_DISPATCH` stops new work | +| Set a rule for the whole project | (type it) | type it | `POST .../message` | text that sounds standing ("always", "never", "from now on", "do not ask") is stored as a project-wide directive; other amendments are scoped to the current task | +| Answer a question the task stopped on | `argus --answer "text"` (`--answer-item ID` if several wait) | answer in the pending-question prompt | `POST /api/projects/{sid}/backlog/{item_id}/answer` | the paused task is replaced by a continuation whose objective carries your answer as authority, and the worker is restarted if needed | +| Ask something without touching the run | `argus --ask "question"` | `/ask question` | `route_override: "chat"` on `/message` | the Manager answers from project state; nothing is queued | + +Why the nudge is deliberately weak: it is guidance the next round *reads*, not +an order the runtime *enforces*. The inbox is durable (a queue under the +project state that survives restarts, acknowledges only after delivery, and +answers HTTP 429 when full), but the Engineer decides what to do with the +text. That is the right tool for "prefer torch.compile off" or "the paper must +cite X"; it is the wrong tool for "stop". + +Why `--answer` is not a nudge: a task that stopped to ask a question is +`paused` until the question is answered. Clearing the question and resuming +the same task looks like it works but does not: the task re-reads the +objective that made it ask and asks again (the code comment on `--answer` +records a campaign that burned five attempts that way). The answer path +therefore enqueues a continuation whose objective states your answer, so the +next round reads what it was told. + +Why a typed message is classified rather than obeyed literally: the same box +carries greetings, questions, new tasks, settings changes and controls. The +Manager first decides whether the message is conversation or work, and only +when a task is active and the text clearly redirects it does the second check +turn it into a `STEER`. Questions, criticism and suggestions are not steering; +if you mean "change course", say so in the imperative. + +When Argus stops on its own to ask you is governed by one setting: + +```bash +export ARGUS_SKILL_AUTONOMY_MODE=pragmatic # default +``` + +`pragmatic` recovers technical problems (a failed test, a timeout, a route +that did not work) without asking and stops only at authority boundaries: +credentials, payment or a bigger budget, deleting or force-pushing shared +state, publishing or sending outward, and changes to an acceptance contract +you own. `cautious` asks on every explicit question the Reviewer raises. +`autonomous` still stops at those same boundaries; it only removes the +remaining discretionary pauses. Pick `cautious` for the first run of a new +vertical, when you want to see what it would have asked. + +Two real cases show what interjection looks like in practice. + +*The follow-up.* In the trial's browser project, after the first table +arrived, the second message at 07:03:44Z was + +> 现在项目里都有什么文件?把 bf16 也补测一下,更新表格。 + +Three minutes later the reply listed the project's files and gave the updated +table with a BF16 column. That message is a new bounded task in the same +project, not a steer of a running one; the earlier task had already finished. +Most "changes of direction" on small tasks are like this, and the message box +is the right place for them. + +*The pause that was not a question.* The roofline campaign's log shows, 43 +seconds after launch, a round ending with "The work was paused before the +Engineer finished this round because provider charges are awaiting +reconciliation or explicit risk approval, so this round was not judged". +Nothing was asked of the operator; the cost check had refused an unpriced +call. The operator's response was two CLI commands recorded in +`daemon.commands.jsonl`: a `drain` at 08:16:12Z and a `start` at 08:21:19Z +after changing the cost policy. Over the following two hours the campaign ran +without a single operator message: all twenty entries in its inbox were +background-job reports the runtime queued for itself. A precise objective is +what made that possible. + +## Choosing the backend and model + +Argus resolves the backend and model for each role separately, and every +resolution follows the same order: + +1. an environment variable for that role (`ARGUS_SKILL_ENGINEER_BACKEND`, + `ARGUS_SKILL_ENGINEER_MODEL`, ...), +2. the shared variable (`ARGUS_SKILL_RUNNER_BACKEND`, `ARGUS_SKILL_MODEL`), +3. the same names in the persisted settings file `~/.argus-skill/config.json`, + which is what `argus --setup`, `/backend`, `/config key=value` and a + natural-language "把模型换成 X" write, +4. the default (`codex` for the backend; `auto` for the model, meaning the + backend's own default). + +A `--backend` typed on a `--daemon` launch sits above all of these for that +daemon and is exported to its roles at boot, so a persisted choice can never +silently override what you typed. `argus --config-help` prints every setting +with its default, its current value and where the value came from +(`env`, `persisted`, `default`); `argus --config-snapshot` writes the same to +a file you can attach to a report. + +| Setting | Default | Notes | +|---|---|---| +| `ARGUS_SKILL_RUNNER_BACKEND` | `codex` | `codex`, `claude`, `copilot`, `cursor`, `opencode`, `pi`, `grok`, `qoder`, `dsh` | +| `ARGUS_SKILL__BACKEND` | inherits | `ENGINEER`, `REVIEWER`, `PLANNER`, `MANAGER`, `SUPERVISOR` | +| `ARGUS_SKILL_MODEL` | `auto` | a bare model id (`gpt-5.6-sol`, `copilot/opus-5`); free text reaches the CLI verbatim and every call then fails | +| `ARGUS_SKILL__MODEL` | `auto` | `ENGINEER`, `REVIEWER`, `PLAN`, `MANAGER`, `SUPERVISOR` | +| `ARGUS_SKILL_FRONTDOOR_MODEL` | `auto` | the cheap classifier that reads every message (`gpt-5.4-mini` on Copilot) | +| `ARGUS_SKILL__REASONING_EFFORT` | Engineer `xhigh`, others `high`, Supervisor `low` | `low`, `medium`, `high`, `xhigh`, `max` | +| `ARGUS_SKILL_PI_PROVIDER` / `ARGUS_SKILL_OPENCODE_PROVIDER` | unset | provider prefix for bare model ids on those two backends; on OpenCode the model is dropped without it | + +Why the roles are separable: they do different jobs. The Engineer runs long +tool-using turns and benefits from the strongest model at high effort; the +Reviewer reads and judges; the front door classifies one short message per +turn and should be cheap. The roofline campaign ran on Copilot with +`ARGUS_SKILL_MODEL=gpt-5.6-sol`, and its `usage.jsonl` splits as: $29.22 on +`gpt-5.6-sol` across engineer, planner, reviewer, manager and supervision +calls; $0.16 on `gpt-5.4-mini` for classification, reflection and fact +extraction. Changing the front-door model would have saved nothing; changing +the Engineer's would have changed everything. + +The same ledger explains where the money goes: 33.9 million input tokens +against 0.3 million output tokens after two hours. Long tool-using rounds +re-read context, which is why Argus by default sends the full task text only +on the first round of a session (`ARGUS_SKILL_COMPACT_CONTINUATION_PROMPTS`) +and why a precise objective with the evidence already in the work directory +costs less than a vague one that makes the Engineer explore. + +## Controlling spend + +Spend is controlled at three levels, and the one that surprises new users is +the third. + +**A daily cap across every project on the host.** +`ARGUS_SKILL_GLOBAL_DAILY_CAP_USD` (default `1000.0`) and, if you prefer to +count tokens, `ARGUS_SKILL_GLOBAL_DAILY_TOKEN_CAP` (default `0`, off). +`argus --status` shows the running total: `budget : global daily $1000.00 +(spent $30.71) · remaining $969.29` was the line two hours into the roofline +campaign, with `cost : $29.79 cumulative` for that project alone. The +per-call record is `~/.argus-skill/projects//usage.jsonl` (model, tokens, +`cost_usd`), and the web UI reads `GET /api/projects/costs`. + +**Provider-level circuit breakers.** `ARGUS_SKILL_CODEX_DAILY_CALL_CAP` +(default 300 calls per day), `ARGUS_SKILL_COPILOT_DAILY_CALL_CAP`, +`ARGUS_SKILL_COPILOT_DAILY_PREMIUM_CAP` and `ARGUS_SKILL_COPILOT_HOURLY_CALL_CAP` +(default 10000 each), `ARGUS_SKILL_PROVIDER_MAX_CONCURRENCY` (default 0, off) +and `ARGUS_SKILL_MAX_ACTIVE_DAEMONS` (default 64). These exist because a +subscription CLI is shared by every project on the host, and one runaway +campaign should not exhaust it for the others. `--mission-width` (default 2) +is the per-project counterpart; the roofline campaign used `1`, which is the +right choice when the tasks share four GPUs. + +**What to do with a call whose price is not known yet.** +`ARGUS_SKILL_UNPRICED_COST_POLICY` is `block` by default: a call whose cost the +provider has not settled is refused before it starts, because a cap that +cannot see the cost cannot enforce anything. With Copilot's subscription +billing the first calls on a fresh install report `pricing_status: partial` +with no `cost_usd`, and that is exactly what happened in the trial: + +- CLI project `s-600e27bf`: the first classify call settled unpriced (its + event has `"pricing_status":"partial","cost_usd":null`), the worker went + to `paused_cost`, and `cost-control.json` listed the call as unresolved with + the reason "Copilot token billing is awaiting local CLI usage + reconciliation". +- Browser project `s-67fb6d62`: the first message came back `[not dispatched] + Manager could not classify this message (refused before start: unresolved + provider cost: 1 call(s) awaiting usage reconciliation ...)`. +- The roofline campaign, on a second install a little later, paused 43 + seconds in for the same reason. + +The operator's fix each time was one line in `~/.argus-skill/config.json`, +`"ARGUS_SKILL_UNPRICED_COST_POLICY": "allow"`, followed by a drain and restart +(`/config ARGUS_SKILL_UNPRICED_COST_POLICY=allow` in the cockpit writes the +same key, and the web configuration view exposes it). `allow` means: run the +call now and let the ledger reconcile later. The trade-off is real. With a +metered API key, keep `block`; the price of every call is known and the cap +is exact. With a subscription CLI, `allow` is usually right, because the +"cost" the ledger reconciles is your subscription's usage, and refusing to run +does not save money you have already paid. The daily cap still applies to +everything that does get priced. + +Two changes merged on 2026-09-30 (#179 and #183) move this default: Copilot +calls made through the warm session are now priced from the CLI's own session +log, and a call Argus itself interrupts before the CLI recorded anything is +settled as one premium request. A fresh install on a subscription CLI no +longer pauses under `block`; the three examples above ran before that change. +`allow` is still the choice when you would rather run than wait for any late +reconciliation at all. + +A last practical point: `ARGUS_SKILL_COST_CONTROL` (default `on`) is the +switch for the whole admission-and-reconciliation layer. Leave it on; turning +it off also removes the `paused_cost` state that tells you something is wrong +with billing. + +## Which tasks suit Argus + +Argus is built for work that can be checked. The four-role loop only pays for +itself when there is evidence for the Reviewer to read and refuse. + +It fits well when: + +- **The result is measurable on the machine it runs on.** Every task in the + three example projects was of this kind: a TFLOPS sweep, a numerical + solution compared point-by-point with an analytic one, a script whose + output table can be rerun. The Reviewer can open the raw files. +- **The work has stages and takes hours.** The roofline campaign moved from + idea to experiment to a paper draft, launched background GPU jobs and + waited for them, and had a draft sent back for rework, all without an + operator message. That is what the daemon, the stage checklists and the + independent review exist for. +- **There is a vertical for the field.** Seven ship with Argus (`research`, + `software`, `math`, ...) and the store carries seventeen more. A vertical is + what turns "done" into a checklist the Reviewer can apply; without one the + Manager still routes the task, but the acceptance standard is generic. See + [building a vertical](building-a-vertical.md). +- **You want it to keep going.** A standing objective with `--continuous` and + no `--bounded` keeps proposing follow-up work in the same direction. + +It fits poorly when: + +- **You want an answer, not work.** `argus --ask` and `/ask` answer inline for + a fraction of a cent and queue nothing. Sending a question through the task + path costs a Manager turn plus, if it is misread as work, a Planner and + Engineer round. +- **The task is smaller than the loop.** The oscillator task, a script most + people would write in ten minutes, took 12 minutes and $0.91 because it went + through Manager, Planner, an Engineer round, a Reviewer `continue`, a second + round and a `done`. That overhead is worth it when you would otherwise not + check the work; it is not worth it for a throwaway. +- **Finishing requires something only you can do.** Logging in somewhere, + paying, publishing, deleting shared state, or approving a change to an + acceptance contract stops the task by design in every autonomy mode. Plan + the objective to end before that step, or do the step first. +- **"Done" cannot be written down.** If you cannot state what the Reviewer + should check, it will accept or refuse on its own reading, and rounds will + continue until a cap. +- **The evidence lives somewhere the roles cannot reach.** The Reviewer is + read-only and works from the work directory and the paths you allow + (`ARGUS_SKILL_REVIEWER_READ_DIRS`); a result that only exists in a service it + cannot query cannot be certified. + +A useful test before you start: write the sentence the Reviewer should be +able to say at the end ("the README reports the maximum error against the +analytic solution for all three damping cases, and each number matches the +saved run"). If you can write it, put it in the objective. If you cannot, +Argus will not be able to tell when it is finished either. diff --git a/docs/building-a-vertical.md b/docs/building-a-vertical.md new file mode 100644 index 000000000..bd794d9b0 --- /dev/null +++ b/docs/building-a-vertical.md @@ -0,0 +1,431 @@ +# Building a vertical + +A vertical teaches Argus what "done" means in one field: the stages work moves +through, the checklist the Reviewer applies at each stage, and the skills the +roles read before they start. This page builds a small real one, +`lab_notebook`, from nothing to an installed entry in the Vertical Store. The +finished files are under [`examples/verticals/`](../examples/verticals/) and +every command below was run against this checkout. + +Background reading: [the Vertical Store](vertical-store.md) for the store's +internals and hosted mode, and +[single-agent verticals](single-agent-vertical-20260915.md) for the design +discussion behind the contract. + +## What a vertical is, in code + +A vertical is a Python package whose `stages.py` declares a handful of +module-level names. There is no base class and no registration call: the +framework imports the module and validates the names through +`argus.core.vertical_contract.vertical_contract`, and (for anything not built +in) `argus.verticals._registry._validated_plugin`. The built-in seven are +listed by name in `argus/skills/vertical_select.py`; everything else arrives +through the store or a Python entry point in the group `argus.verticals`. + +The names the contract reads: + +| Name | Required | Meaning | +|---|---|---| +| `ARGUS_VERTICAL_API_VERSION` | yes, for non-built-ins | must be `1`; any other value hides the vertical | +| `VERTICAL_PURPOSE` | yes, for non-built-ins | one sentence; it is the line the Manager sees when choosing a vertical, so write it for routing, not as an abstract | +| `CHECKLIST_STAGE_ORDER` | yes | tuple of stage names, in order; no duplicates | +| `CHECKLIST_ITEMS` | yes | dict `stage -> tuple[ChecklistItem, ...]`; every non-optional stage needs a non-empty checklist | +| `completion_gate` | yes | `"none"`, `"metric"` or `"certified"` | +| `CHECKLIST_OPTIONAL_STAGES` | no | stages that may have no checklist | +| `STAGE_ALIASES` | no | `{"experiment": "measure"}`: other names the Manager may use for a stage | +| `WORKFLOW_MODE` | no | `"staged"`, `"direct"` or `"proportional"` | +| `MISSION_KIND` | no | `"custom"`, `"optimize"`, `"research"` or `"software"` | +| `REQUIRE_INDEPENDENT_REVIEW` | no | defaults to `True` | +| `role_banner(role)` | no | a paragraph each role reads before its task | +| `render_role_prompt_fragment`, `render_role_prompt_context`, `stage_completion_issues`, `prepare_mission`, `LIBRARY_PREPARER`, `EVIDENCE_SCHEMA`, `WORKFLOW_PROFILES` | no | hooks the larger built-ins use; not needed for a first vertical | +| `VERTICAL_SKILLS`, `VERTICAL_SKILL_PARENTS`, `VERTICAL_ROUTING_PATH` | no | an explicit skills root (a `skills/` directory next to `stages.py` is found automatically), verticals whose skills are seeded before yours, and a browsing category for the store | + +A `ChecklistItem` (from `argus/skills/stage_machine.py`) has three fields: + +```python +@dataclass(frozen=True) +class ChecklistItem: + id: str + statement: str + evidence_hint: str +``` + +The `statement` is what must be true; the `evidence_hint` tells the Engineer +what to show and the Reviewer what to look for. This is the whole mechanism: +the Reviewer is read-only, so the quality of your checklist is the quality of +your review. + +## Two built-ins to learn from + +**`software`** (`argus/verticals/software/`) is the smallest useful vertical: +one stage, `delivery`, with three checklist items, a `role_banner`, and four +skill files. Its `stages.py` is worth reading in full; the module docstring +explains why the checklist exists ("code that does not build" was the failure +it was written against), and the first two item ids are marked protected so a +Planner may add items but never weaken them: + +```python +STAGE_ORDER = ["delivery"] +CHECKLIST_STAGE_ORDER = tuple(STAGE_ORDER) +CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = () +completion_gate = "none" +MISSION_KIND = "software" +WORKFLOW_MODE = "staged" + +PROTECTED_ITEM_IDS = frozenset({"delivery.builds", "delivery.tests-executed"}) +``` + +Its skills sit at `skills//.md`: + +``` +software/skills/engineer/software-change-implementation.md +software/skills/manager/software-project-grounding.md +software/skills/planner/software-project-grounding.md +software/skills/reviewer/software-change-review.md +``` + +**`research`** (`argus/verticals/research/`) is the largest: four stages, + +```python +CANONICAL_STAGE_ORDER: tuple[str, ...] = ("idea", "experiment", "paper", "review") +STAGE_ALIASES = {"research": "idea", "plan": "experiment", ..., "submission": "review"} +``` + +with `completion_gate = "certified"`, `WORKFLOW_MODE = "proportional"`, role +prompts in `prompt_policy.py`, a `prepare_mission` hook in `mission_brief.py`, +and some forty engineer skills plus seven reviewer skills. It shows where a +vertical can grow; it is not where one should start. + +The directory shape both share, and the store expects: + +``` +/ + __init__.py docstring only + stages.py the contract + skills/ + engineer/*.md + reviewer/*.md + manager/*.md (optional) + planner/*.md (optional) + engineer/references/ supporting files, not skills (optional) + engineer/*_scripts/ .py/.json/.sh copied verbatim (optional) +``` + +## Step 1: decide the stages and the checklist + +`lab_notebook` is for small measurement tasks: run something on this machine, +then write it up so someone else can rerun it. Two stages follow from that, +`measure` and `report`, and each gets the two or three things a reviewer +would actually check. + +Write the checklist before anything else, and write it as statements a +read-only reviewer can verify from files. "The measurement was repeated" is +checkable (there are N raw files); "the measurement is good" is not. + +`examples/verticals/argus_verticals/lab_notebook/stages.py`: + +```python +from argus.skills.stage_machine import ChecklistItem + +ARGUS_VERTICAL_API_VERSION = 1 +VERTICAL_PURPOSE = ( + "small measurement tasks on this machine: run the measurement, record how " + "it was run, and write a notebook entry another person can reproduce" +) + +CHECKLIST_STAGE_ORDER: tuple[str, ...] = ("measure", "report") +CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = () +STAGE_ALIASES = {"experiment": "measure", "writeup": "report"} + +completion_gate = "none" +WORKFLOW_MODE = "staged" +MISSION_KIND = "custom" +REQUIRE_INDEPENDENT_REVIEW = True + +CHECKLIST_ITEMS: dict[str, tuple[ChecklistItem, ...]] = { + "measure": ( + ChecklistItem( + id="measure.ran-here", + statement=("The measurement was executed on this machine in this project, " + "and its raw output is saved in the work directory."), + evidence_hint="the command that was run and the path of its raw output", + ), + ChecklistItem( + id="measure.repeated", + statement=("The measurement was repeated, and the reported number is a " + "median or mean with its spread; a single run is not a result."), + evidence_hint="number of repeats, the aggregate and the spread", + ), + ), + "report": ( + ChecklistItem( + id="report.notebook-entry", + statement=("NOTEBOOK.md in the work directory states what was measured, " + "how, the result with its spread, and the exact command to rerun it."), + evidence_hint="the NOTEBOOK.md section and the rerun command", + ), + ChecklistItem( + id="report.numbers-traceable", + statement="Every number in NOTEBOOK.md can be traced to a saved raw output file.", + evidence_hint="raw file paths next to each number", + ), + ), +} + + +def role_banner(role: str) -> str: + return ( + "LAB NOTEBOOK VERTICAL: measure first, then write. A number without a " + "saved raw output and a rerun command is not a result. The Reviewer " + "checks the notebook entry against the raw files, not against the " + "Engineer's summary." + ) +``` + +Why these particular choices: + +- `completion_gate = "none"` means the project is complete once the last + stage's checklist is met. `"metric"` is for verticals whose completion is a + number reaching a target; `"certified"` adds an explicit certification step, + which the research vertical uses for papers. A notebook entry needs neither. +- `WORKFLOW_MODE = "staged"` makes the Manager move through the stages in + order. `"direct"` skips staging for one-shot work; `"proportional"` lets the + Manager scale the process to the task. Start with `staged`; the aliases let + a Manager that says "experiment" land on `measure`. +- `MISSION_KIND = "custom"` because the other three values switch on + behaviour written for optimization loops, research campaigns and repository + changes. +- The `role_banner` says one thing, in the field's own terms. It is read by + every role, so it is the place for the rule that ties the stages together. + +The contract check refuses, with a message naming the problem, a stage with no +checklist, a checklist for a stage that is not in the order, a duplicate stage, +an item without an `id` or `statement`, a repeated item id within a stage, and +an unknown `completion_gate` value. Run it directly while you iterate: + +```bash +python -c " +import importlib.util, pathlib +from argus.core.vertical_contract import vertical_contract +p = pathlib.Path('examples/verticals/argus_verticals/lab_notebook/stages.py') +spec = importlib.util.spec_from_file_location('lab_notebook_stages', p) +m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m) +c = vertical_contract('lab_notebook', m) +print(c.stage_order, c.completion_gate, c.workflow_mode) +" +``` + +## Step 2: write the skills + +A skill is a Markdown file with exactly two front-matter fields, `name` and +`description`, both quoted. The runtime does not parse skill bodies: the +roles receive the library paths and read the files themselves, so the +description is a routing line (it is what an agent reads to decide whether to +open the file; the repository's own test caps it at 1200 characters) and the +body is written for a reader who will act on it. + +`skills/engineer/measurement-record.md`: + +```markdown +--- +name: "Measurement Record" +description: "How to run a small measurement so the result can be checked and rerun: save raw output, repeat, report the spread, and record the exact command." +--- + +# Measurement record + +Run the measurement from a script saved in the work directory, never from an +interactive shell you cannot show later. Write the raw output to a file under +`results/` before computing any aggregate. + +Repeat the run. Report the median (or mean) together with the spread (min/max or +standard deviation) and the number of repeats. A single run is not a result. + +Record in `NOTEBOOK.md`: + +- what was measured and on which hardware +- the exact command to rerun it +- the aggregate, the spread and the number of repeats +- the path of the raw output every number came from +``` + +`skills/reviewer/measurement-review.md` tells the Reviewer to open the raw +files and recompute one aggregate before accepting, and to return `continue` +for any number without a file behind it. The two skills and the two checklists +say the same thing from three sides; that redundancy is deliberate, because +each role reads only its own. + +Rules the loader applies: skills live at `skills//` where the role is +`engineer`, `reviewer`, `planner` or `manager` (anything else is filed as +general); a `references/` directory is copied as supporting material rather +than as skills; `*_scripts/` directories are copied verbatim; names starting +with `_` or `.` are skipped. When the roles start a task with this vertical, +those files are seeded into the project's skill library in the order project, +vertical, global, so a project's own edits win. + +## Step 3: package it for the store + +The store installs a zip whose members sit under `argus_verticals//`, +described by a `catalog.json`. The catalog entry the store validates +(`argus/verticals/store.py`, `_validate_entry`) needs: + +| Field | Constraint | +|---|---| +| `name` | equals the key; `^[a-z][a-z0-9_]{0,47}$` | +| `version` | non-empty, no slashes | +| `module` | `argus_verticals..stages`, and it must live in `paths[0]` | +| `paths` | at least `["argus_verticals/"]` | +| `purpose` | non-empty | +| `archive.file`, `archive.url`, `archive.sha256`, `archive.size` | the zip's name, location, digest and exact byte size; a non-`https` URL is only accepted when the catalog itself is local | + +Optional: `purpose_zh`, `requires` (other verticals to install first), +`shared` (helper trees outside the vertical's own directory), +`python_requirements` (shown, never installed), `tags`, `skill_parents`, +`has_skills`, `routing_path`, `min_argus`, `maintainers`. + +`examples/verticals/build_local_catalog.py` produces both files for a directory +under `examples/verticals/argus_verticals/`: + +```bash +python examples/verticals/build_local_catalog.py lab_notebook --version 0.1.0 --out /tmp/lab-store +``` + +``` +archive : /tmp/lab-store/lab_notebook-0.1.0.zip (3099 bytes) +catalog : /tmp/lab-store/catalog.json +install : ARGUS_VERTICAL_CATALOG=/tmp/lab-store/catalog.json argus verticals install lab_notebook +``` + +The catalog it wrote: + +```json +{ + "schema": 1, + "verticals": { + "lab_notebook": { + "name": "lab_notebook", + "version": "0.1.0", + "module": "argus_verticals.lab_notebook.stages", + "paths": ["argus_verticals/lab_notebook"], + "purpose": "small measurement tasks on this machine: run the measurement, record how it was run, and write a notebook entry another person can reproduce", + "has_skills": true, + "archive": { + "file": "lab_notebook-0.1.0.zip", + "url": "file:///tmp/lab-store/lab_notebook-0.1.0.zip", + "sha256": "5b36…b8e0", + "size": 3099 + } + } + } +} +``` + +The `sha256` and `size` are checked byte-for-byte at install; regenerate the +catalog whenever the zip changes. + +## Step 4: install it and confirm Argus sees it + +Point the store at the local catalog. A local catalog turns its directory into +an offline mirror: an archive named in the catalog and sitting next to it is +used without any download. + +```bash +export ARGUS_VERTICAL_CATALOG=/tmp/lab-store/catalog.json +argus verticals install lab_notebook +argus verticals list +argus verticals info lab_notebook +``` + +What this checkout printed (the install was done into a throwaway +`ARGUS_SKILL_HOME` so the machine's real store stayed untouched): + +``` +lab_notebook: install started + [ 0%] lab_notebook: starting + [100%] lab_notebook: install finished +lab_notebook: install finished +``` + +``` +name kind version installed enabled update purpose +... +lab_notebook installed 0.1.0 0.1.0 yes - small measurement tasks on this machine: run the measurement +``` + +``` +name lab_notebook +kind installed +purpose small measurement tasks on this machine: run the measurement, record how it was run, and write a notebook entry another person can reproduce +version 0.1.0 +installed_version 0.1.0 +enabled True +actions disable, uninstall +operation install done: install finished +``` + +The files landed at `/verticals/argus_verticals/lab_notebook/` +next to the store's `registry.json`, and the registry advertises the vertical +with origin `store`, its two skills found automatically next to `stages.py`: + +``` +advertised: True +origin: store | skills_root: .../verticals/argus_verticals/lab_notebook/skills +skills: ['engineer/measurement-record.md', 'reviewer/measurement-review.md'] +stages: ('measure', 'report') | gate: none | mode: staged +``` + +From here the Manager can choose `lab_notebook` for a task whose text matches +its purpose line; there is no flag to force a vertical (`ARGUS_SKILL_VERTICAL` +is a legacy name with no authority), and the choice is saved per project in +`.argus/PIPELINE_STATE.json`. To try it, start a project with an objective in +the vertical's own words, for example "measure how long `python -c 'import +torch'` takes on this machine, repeat it, and write a notebook entry with the +rerun command", and check `pipeline : vertical=lab_notebook` in +`argus --status`. + +`argus verticals disable lab_notebook` hides it again without removing the +files; `argus verticals remove lab_notebook` deletes them (it refuses while a +local project's `PIPELINE_STATE.json` still names the vertical, unless you +pass `--force`). + +## Publishing to the shared store + +The default catalog is the `catalog.json` attached to the latest release of +[Argus-AiTeam/argus-verticals](https://github.com/Argus-AiTeam/argus-verticals); +each vertical is one directory in that repository, and its release workflow +builds the zips and the catalog (`scripts/build_catalog.py` there). Publishing +therefore means opening a pull request that adds +`argus_verticals//` to that repository, with the same layout as here. The +store on every user's machine only accepts archives from an allow-listed host +(`github.com` and its release asset hosts), so a catalog you host elsewhere +must be reached through `ARGUS_VERTICAL_CATALOG` as a local file, which is what +the steps above do. + + + +Two other ways exist to run a vertical without the store, useful during +development of a bigger one: `pip install` a package that declares the entry +point `argus.verticals` (`name = "argus_verticals..stages"`), which wins +over a store copy of the same name; or install it as a managed workbench +plugin. A name that collides with a built-in is ignored from every source. + +## Checks before you share it + +- The contract check in step 1 passes. +- Every skill header has quoted `name` and `description` values. The + repository's own skill writer (`argus/skills/store.py`) emits exactly that + shape, JSON-quoted, because an unquoted description containing a colon does + not survive a YAML parse; `tests/skills/test_skill_frontmatter_integrity.py` + checks the in-tree verticals for the same reason. +- `argus verticals install` from a local catalog succeeds and + `argus verticals list` shows the row as `installed` and enabled. +- One real task ran through it and the Reviewer's verdicts referred to your + checklist ids (they appear in the round events in `events.jsonl`). + +The example under `examples/verticals/` is not scanned by +`tests/skills/test_vertical_plugins.py` (that file builds synthetic plugins in +memory), so adding a vertical there cannot change the test's result. On this +machine one test in that file fails before and after these changes for a +local reason: a pip-installed `argus_verticals` package in the venv makes +`test_store_verticals_are_discovered_with_origin_store` see two package paths +instead of one. diff --git a/docs/getting-started.md b/docs/getting-started.md new file mode 100644 index 000000000..a2a3de193 --- /dev/null +++ b/docs/getting-started.md @@ -0,0 +1,419 @@ +# Getting started with Argus + +This guide takes you from an empty machine to a finished first task. It covers +three ways to work: the command line, the web UI, and the desktop app. Every +command below was checked against the `argus` CLI in this checkout +(`argus 0.1.8`); the example runs are real projects, with their times and costs +taken from their own logs. + +Companion guides: [best practices](best-practices.md) (objectives, changing +direction mid-run, models and spend) and +[building a vertical](building-a-vertical.md). + +## What you are installing + +Argus is a Python package plus a bundled terminal cockpit (Node.js) and a web +UI. It does not call a model API itself; it drives one of the coding-agent CLIs +you already use (Copilot, Codex, Claude Code, Cursor, Pi, OpenCode, Grok, Qoder, +DeepSeek Harness) and lets four roles share it: a Manager that routes your +message, a Planner that decides the next task, an Engineer that does the work, +and a read-only Reviewer that decides whether the work is done. + +Two things follow from that design and explain most of what you will see: + +- **A background worker does the work.** Your cockpit, browser tab or terminal + can close; the project's daemon keeps running and keeps its state under + `~/.argus-skill/projects//`. +- **Nothing is "done" until the Reviewer says so.** A task usually takes + several Engineer rounds. Expect the first small task to take minutes, not + seconds. + +## Prerequisites + +| Requirement | Why | +|---|---| +| Python 3.11+ | the `argus` package (`pyproject.toml` requires `>=3.11`) | +| Node.js 22.12+ | the terminal cockpit is an Ink app; `argus` refuses to start it on older Node (`argus: Ink TUI requires Node.js 22.12 or newer`) | +| One authenticated agent CLI | Argus reuses its login; there is no separate Argus account | + +Install and log into one CLI first. The README's +[Quick Install](../README.md#quick-install) table lists the install and login +command for each backend; for example GitHub Copilot CLI is +`npm install -g @github/copilot` then `copilot login`. + +## Install + +The commands are the README's, repeated here so this page stands alone. + +Linux (isolated venv; use this on servers): + +```bash +git clone https://github.com/microsoft/ArgusAgent.git "$HOME/Argus" +cd "$HOME/Argus" +python3 -m venv .venv +.venv/bin/python -m pip install --upgrade pip +.venv/bin/python -m pip install -e . +ARGUS_BIN="$HOME/Argus/.venv/bin/argus" +"$ARGUS_BIN" --version +``` + +Windows (PowerShell, no venv): + +```powershell +py -m pip install --upgrade pip +py -m pip install --upgrade --force-reinstall "argus @ https://github.com/microsoft/ArgusAgent/archive/refs/heads/main.zip" +``` + +macOS (managed command): + +```bash +uv tool install --force --python 3.12 "argus @ https://github.com/microsoft/ArgusAgent/archive/refs/heads/main.zip" +``` + +On Linux, `argus` below means `$HOME/Argus/.venv/bin/argus` unless the venv is +active. Everywhere, `python -m argus ...` runs the same CLI without ever +starting the Node cockpit; it is what you want in scripts and over SSH without +a terminal. + +## First-time setup and a health check + +```bash +argus --setup +argus doctor +``` + +`argus --setup` is an interactive wizard: it asks which backend to use, checks +that the CLI is installed and logged in, runs one real model turn through it, +and only then saves the profile. It ends with `Setup complete. Run `argus`.` +If you would rather not answer prompts: + +```bash +argus --setup --backend copilot --non-interactive +``` + +Two variants worth knowing: + +| Situation | Command | +|---|---| +| Keep Argus on its own Copilot account, separate from your interactive login | `argus --setup --backend copilot --copilot-home "$HOME/.copilot-argus" --copilot-login` | +| Use an OpenAI-compatible endpoint (served through Pi) | `argus --setup --api-url URL --api-model MODEL` with the key in `ARGUS_SETUP_API_KEY` rather than `--api-key`, so it stays out of shell history | + +`argus doctor` is read-only. This is what it printed on the machine this guide +was written on (a source checkout, Pi backend): + +``` +argus doctor — cross-platform diagnostics + +✓ ARGUS-HOST-001 [host/supported_host] Linux 6.8.0-139-generic x86_64 +✓ ARGUS-INSTALL-001 [install/source_checkout] /data/.../Argus +✓ ARGUS-ASSET-001 [install/assets_ready] release manifest and Web/TUI assets are present +✓ ARGUS-PYTHON-001 [cli/python_ready] 3.12.3 at /data/.../.venv/bin/python +✓ ARGUS-NODE-001 [cli/node_ready] v22.23.2 +✓ ARGUS-WEB-001 [web/compatible] 127.0.0.1:8799 is compatible +! ARGUS-DESKTOP-001 [desktop/tauri_dependencies_missing] Tauri Desktop sources exist but npm/Rust build dependencies are incomplete + fix: run `npm --prefix desktop-tauri ci` and install the Rust Windows toolchain +✓ ARGUS-DAEMON-001 [daemon/stopped] no daemon is running +✓ ARGUS-BACKEND-001 [backend/ready] pi 0.85.1 runnable at /home/.../argus-pi (subscription_cli; authentication checked; ...) + +all blocking checks passed +``` + +A `!` line is advice, not a failure; the desktop warning above only matters if +you intend to build the desktop app from this checkout. When something is +wrong, `argus doctor --fix-safe` applies the repairs the doctor itself marked +safe and reruns; `argus doctor --advisor auto` lets one of the installed agent +CLIs inspect and repair. Without one of those two options `doctor` changes +nothing. + +A note on help: `argus --help` (and `argus doctor --help`, +`argus verticals --help`) print the same short overview. The complete flag +reference is behind an environment variable: + +```bash +ARGUS_SKILL_DEBUG_HELP=1 argus --help +``` + +## Path 1: the command line + +There are two ways to use the CLI. A bare `argus` opens the terminal cockpit, +where you type in natural language. Everything else (`--daemon`, `--status`, +`--notify`, ...) is a one-shot command that talks to the project's worker and +exits; those work over SSH, in cron, and without Node. + +### Which project a command means + +Argus keys a project's state on the directory you run it from. Management +commands attach to the newest session for the current directory, or to the +one you name: + +| Flag | Meaning | +|---|---| +| (none) | the newest session for the current working directory | +| `--project-root DIR` | the newest session for `DIR` | +| `--resume ID` | a specific session id (with no id, a picker of recent sessions) | +| `--continue` | the most recently active session | +| `--life-dir DIR` | use `DIR` instead of `~/.argus-skill` as the state root (`ARGUS_SKILL_HOME` does the same) | + +### Start a task without opening the cockpit + +Write the objective in a file, `cd` into the directory the work should happen +in, and start a background worker with a campaign objective: + +```bash +cd ~/work/matmul-roofline +argus --daemon --continuous --objective-file objective.txt --bounded --backend copilot +``` + +The pieces: + +| Flag | What it does and why you would set it | +|---|---| +| `--daemon` | start a detached worker (`--daemon-fg` keeps it in the foreground, for systemd or debugging) | +| `--continuous` | give the worker an objective and let the Planner keep generating tasks toward it; refused without an objective (`--continuous requires a non-empty --objective`) | +| `--objective TEXT` / `--objective-file PATH` | the objective; the file form keeps long text and quotes out of the shell (`--objective` without `--continuous` is refused too) | +| `--bounded` | stop when the Planner certifies the project done; without it the worker keeps generating work for the same objective | +| `--backend NAME` | the agent CLI for this daemon; it outranks the persisted setup for this launch and is exported to every role | +| `--mission-width N` | how many tasks may run in parallel (default 2; `1` is serial) | + +Before the worker starts, Argus re-checks that the backend is installed and +logged in; if not, it prints the readiness report and exits with code 3 instead +of starting a worker that cannot call a model. + +### A real first campaign + +The objective below started the campaign that project `s-02d3c282` ran on this +machine on 2026-09-30 (`objective-roofline.txt`, quoted in full because its +shape is what makes the rest work): + +> 做一项小规模但真实的研究:在本机 4 块 NVIDIA RTX A6000 上,PyTorch 方阵乘法的实际吞吐(TFLOPS)随矩阵规模(512 到 16384)和精度(FP32、TF32、BF16、FP16)如何变化,与理论峰值的差距能否用一个简单的 roofline 式模型(计算强度与显存带宽)解释。要求:真实实验、每种配置多次重复取中位数并报告离散度,拟合模型并给出误差,画图,最后写成一篇 4 页以内的短论文(含方法、结果、讨论、复现说明),并通过内部评审。所有数字必须来自本机真实运行,不允许估算或编造。 + +It was launched with `--backend copilot --mission-width 1` (the live process +still shows those flags). About two hours in, `argus --status` reported: + +``` + project : .../projects/s-02d3c282 + daemon : alive (pid 1430124, up 1h 48m, backend live — see /roles, width 1) + budget : global daily $1000.00 (spent $30.71) · remaining $969.29 + active : 0 pending · 1 running · 0 paused + current : + title : 重绘正式数据图并完成论文稿 + history : 2 done + cost : $29.79 cumulative + continuous: on + pipeline : vertical=research + lifecycle: + state : writing +``` + +By then the work directory held `paper/main.tex`, `paper/results.tex`, +`results/attempt-01`, `figures/` and `src/`; `paper/results.tex` reports, at +n=16384, 22.81 / 57.01 / 121.11 / 115.96 TFLOP/s for FP32 / TF32 / BF16 / FP16 +(58.95 / 73.66 / 78.24 / 74.91 % of the spec sheet) over 17,280 samples, and a +fitted roofline whose calibrated form has a 53.92 % mean absolute error on the +confirmation set. The campaign was still running when this page was written +(its log had a `paper` round returned as `replan_requested`, which is normal: +the Reviewer sent the draft back), so treat those as a snapshot, not a result. + +For a small first task, a bounded objective finishes in minutes. Project +`s-4a34a674` on the same machine was started the same way with this objective: + +> 用 numpy 数值求解一维阻尼谐振子 x'' + 2γx' + ω0² x = 0(取 ω0=2π, 分别取欠阻尼 γ=0.5、临界 γ=2π、过阻尼 γ=10),与解析解逐点比较给出最大误差,画出三条曲线,把方法、数值结果表和图写进本目录的 README.md。所有数字必须来自真实运行。 + +Its log runs from 08:11:39Z to 08:23:33Z (12 minutes), one task in two +attempts, `continuous.json` ends with `"done_reason": "planner declared project +done"`, and it cost $0.91 (12 model calls in `usage.jsonl`). The README it +wrote gives maximum errors of 6.0e-07, 1.6e-07 and 5.1e-07 for the three +damping cases against the analytic solution. + +### Watch, steer, stop + +| Command | What it does | +|---|---| +| `argus --status` | one screen: daemon, budget, current task, stage, last events | +| `argus --follow` | stream the event log to the terminal (`tail -f` style, Ctrl-C to stop) | +| `argus --watch` | the read-only live cockpit | +| `argus --notify "text"` | queue guidance for the next Engineer round (`--notify-stage STAGE` holds it until that stage) | +| `argus --answer "text"` | answer the question a paused task is waiting on (`--answer-item ID` when several wait) | +| `argus --ask "question"` | ask the Manager something and exit; nothing is queued and no daemon is needed | +| `argus --daemon-stop --drain` | let the current task finish at a clean boundary, then exit; the safe way to stop before upgrading | +| `argus --daemon-stop --force` | SIGKILL if it does not exit in time (interrupts running work) | +| `argus --daemon-runbook` | print the restart playbook for the current project | + +The difference between `--notify` and `--answer` matters: a nudge is read at +the next round and does not wake a task that has stopped to ask you something; +`--answer` does, and it also rewrites the task so the next round reads your +answer instead of re-asking. [Best practices](best-practices.md#changing-direction-while-it-runs) +goes into this. + +### The terminal cockpit + +```bash +argus +``` + +The cockpit needs a real terminal (piped or cron use gets a message pointing +you to `--web`, `--watch`, `--status` or `--daemon`). Type what you want in +plain language. The Manager reads every message and decides whether to answer +it itself (a question, a status request, a small local check) or to turn it +into a task for the Planner, Engineer and Reviewer (multi-step work, anything +that needs an independent review). You do not choose a vertical, a model or a +backend for the first task; the Manager picks the vertical from your text and +the backend comes from setup. + +`/new` (optionally `/new `) opens a two-field form, Name and +Objective (Tab or arrows switch fields, Enter creates), and switches the +cockpit to the new session; its worker starts when the first task arrives. + +Commands you will use in the first hour (`/help` lists them all): + +| Command | Purpose | +|---|---| +| `/task ` | queue work directly, skipping the "is this a chat?" decision | +| `/ask ` | answer inline; no task is queued | +| `/plan ` | preview the Planner's execution plan before committing | +| `/status`, `/roles`, `/backlog`, `/journal` | what is running, on which backend/model, what is queued, what happened | +| `/nudge ` | inject guidance into the running task | +| `/abort` | stop the running task now | +| `/backend [name]`, `/config [key=value]` | view or change the runner backend and runtime settings (persisted) | +| `/resume`, `/daemons` | switch to another project | +| `/quit` | leave; background work keeps running | + +## Path 2: the web UI + +```bash +argus --web +``` + +From the `argus` command this goes through the cockpit launcher, which picks +the first free port from 8799 and opens your browser at +`http://127.0.0.1:8799`. With an explicit host or port, or through +`python -m argus --web`, the Python server runs directly and does not open a +browser: + +```bash +argus --web --web-port 8800 +python -m argus --web # over SSH: then `ssh -L 8799:127.0.0.1:8799 user@server` +``` + +Binding to a LAN address (`--web-host 0.0.0.0`) always requires a bearer +token: `ARGUS_SKILL_WEB_TOKEN` if set, otherwise one is minted for the run and +printed with a QR code. The web assets are checked into the repository and +packaged with the wheel, so nothing needs building first. + +### The first screen + +With no sessions yet the page shows **Start a project**, the line *No sessions +yet. Create one to begin.*, and a **New project** button. The button creates an +idle session (no form; the toast reads *New session ready. Say what it is for +and the work begins.*). Then type the objective into the message box. The +message goes to `POST /api/projects/{sid}/message`, the Manager classifies it, +and if it is work the session's daemon is started on demand and the task is +queued. The web UI follows your browser language (English or Simplified +Chinese); the sidebar has a language button. + +### A real first task from the browser + +Project `s-67fb6d62` in the fresh-install trial on this machine was created +from the web UI (`session.json` records `origin: web`). Its first message, at +06:57:51Z: + +> 帮我在这台机器的 GPU 上测一下 PyTorch 矩阵乘在 fp16 和 fp32 下的 TFLOPS,写个脚本跑出真实数字,整理成表格。 + +The reply, three minutes later at 07:02:21Z, reported an average of 121.31 +TFLOPS in FP16 and 23.96 in strict FP32 at 16384×16384 across the four GPUs, +with a per-GPU table and the script name (`benchmark_torch_matmul.py`, seven +timed batches, medians). A follow-up at 07:03:44Z, "现在项目里都有什么文件?把 bf16 +也补测一下,更新表格。", produced an updated table at 07:07:00Z (FP16 122.08, +BF16 125.00, FP32 23.87 on average). The project's usage ledger shows $0.80 for +the whole exchange, plus two early calls that were never priced (see below). + +That first message was actually sent twice. The first attempt (also 06:57:51Z) +came back as `[not dispatched] Manager could not classify this message +(refused before start: unresolved provider cost ...)`: on a fresh Copilot +install the first call's price was not yet known, and the default policy +refuses to spend money it cannot price. [Best practices](best-practices.md#controlling-spend) +explains the setting (`ARGUS_SKILL_UNPRICED_COST_POLICY`) and the trade-off. + +### The endpoints behind the UI + +If you script against the server, these are the calls the UI itself makes +(all need the bearer token when one is configured): + +| Endpoint | Body | Purpose | +|---|---|---| +| `POST /api/daemons` | `objective`, `name`, `workdir`, `launch_cwd` | create a session (an objective starts its worker immediately) | +| `POST /api/projects/{sid}/message` | `text`, optional `route_override` (`auto`/`chat`/`task`), `attachments` | send a message through the Manager; `/message/stream` is the SSE twin | +| `POST /api/projects/{sid}/tasks` | `text`, `autostart_daemon` (default true) | queue work directly | +| `POST /api/projects/{sid}/nudge` | `text` | guidance for the next round (429 when the queue is full) | +| `POST /api/projects/{sid}/backlog/{item_id}/answer` | `text` | answer a paused question and resume | +| `POST /api/projects/{sid}/continuous` | `enabled`, `objective` | turn a campaign objective on or off | +| `GET /api/projects/costs` | | spend per project | + +## Path 3: the desktop app + +The desktop app is a Tauri host around a frozen copy of the same Python +backend; it opens the same web cockpit. It is a packaged preview channel, +separate from source updates. + +- **Download:** the Releases page of + [lbx154/Argus](https://github.com/lbx154/Argus/releases). Windows ships as + an NSIS installer named `Argus--setup.exe`. The release workflow can + also build macOS (`.dmg`, arm64 and x86_64) and Linux (`.AppImage`, `.deb`) + packages, but the workflow file notes that v0.1.7 and v0.1.8 were published + for Windows only, so check the assets of the release you pick. + `microsoft/ArgusAgent` Releases is a separate channel; do not mix installers + or updates between the two. +- **What it needs:** nothing else on the machine. The installer bundles the + backend (`argus-backend`), so no Python, Node.js or venv is required. +- **First launch:** a setup wizard asks you to pick an installed, logged-in + agent CLI and confirm it; the local backend does not start until you do. If + you have an internal trial key instead, paste it and click **开始试用** ("start + trial"): the app downloads the official standalone Copilot program, checks + its SHA-256, runs one real model reply, and opens the workbench. That path + needs no CLI, Python, Node or coding-agent account of your own. Either + choice is remembered; **File → Settings** changes the CLI, executable or + port later. +- **After that** the app is the web UI from Path 2: create a project, type the + objective, watch the feed. + +To build it from source instead (Python 3.11+, Node 22.12+, Rust stable, plus +MSVC Build Tools on Windows or Xcode command-line tools on macOS), the +documented sequence in [docs/desktop-trial.md](desktop-trial.md) is: + +```bash +python -m pip install -e ".[trial]" "pyinstaller>=6.11,<7" tzdata +npm --prefix frontend/web ci +npm --prefix frontend/tui ci +npm --prefix desktop-tauri ci +npm --prefix frontend/web run build && npm --prefix frontend/tui run build +npm --prefix desktop-tauri run build:backend +npm --prefix desktop-tauri run build:unsigned +``` + +The unsigned bundle lands under `desktop-tauri/src-tauri/target/release/bundle/`. +[docs/windows-desktop.md](windows-desktop.md) covers signing and the update +channel. + +## What "finished" looks like + +Whichever path you used, the task ends the same way. The Reviewer returns +`done` for the last task; with `--bounded` (or a bounded objective in the +cockpit) the Planner then declares the project done and the worker stops, and +`argus --status` shows `history : N done` with no running task. The evidence is +in the work directory (the oscillator task's `README.md`; the roofline +campaign's `paper/` and `results/`), and the per-call record of what it cost is +in `~/.argus-skill/projects//usage.jsonl`. + +If instead the status shows a task `paused`, read the question with +`argus --status` or in the cockpit and answer it with `argus --answer`; if it +shows `paused_cost`, see [controlling spend](best-practices.md#controlling-spend). + +## Where things live + +| Path | Contents | +|---|---| +| `~/.argus-skill/` | the state root (`ARGUS_SKILL_HOME` or `--life-dir` overrides it) | +| `~/.argus-skill/config.json` | persisted settings from `--setup`, `/config`, `/backend` | +| `~/.argus-skill/projects//` | one project's state: `events.jsonl`, `usage.jsonl`, `backlog.jsonl`, `continuous.json`, `daemon.log` | +| `~/.argus-skill/verticals/` | verticals installed from the store | +| your work directory | where the roles read and write files; never inside the state root | diff --git a/examples/verticals/argus_verticals/lab_notebook/__init__.py b/examples/verticals/argus_verticals/lab_notebook/__init__.py new file mode 100644 index 000000000..f301d5db4 --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/__init__.py @@ -0,0 +1,5 @@ +"""``lab_notebook``: the worked example from docs/building-a-vertical.md. + +A two-stage vertical for small measurement tasks: measure something on this +machine, then write the result up so another person can rerun it. +""" diff --git a/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md b/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md new file mode 100644 index 000000000..a60c336b6 --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md @@ -0,0 +1,20 @@ +--- +name: "Measurement Record" +description: "How to run a small measurement so the result can be checked and rerun: save raw output, repeat, report the spread, and record the exact command." +--- + +# Measurement record + +Run the measurement from a script saved in the work directory, never from an +interactive shell you cannot show later. Write the raw output to a file under +`results/` before computing any aggregate. + +Repeat the run. Report the median (or mean) together with the spread (min/max or +standard deviation) and the number of repeats. A single run is not a result. + +Record in `NOTEBOOK.md`: + +- what was measured and on which hardware +- the exact command to rerun it +- the aggregate, the spread and the number of repeats +- the path of the raw output every number came from diff --git a/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md b/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md new file mode 100644 index 000000000..d9768ad3d --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md @@ -0,0 +1,16 @@ +--- +name: "Measurement Review" +description: "How to review a measurement task: open the raw output files, recompute one aggregate, and confirm the rerun command is complete before accepting." +--- + +# Measurement review + +Do not accept the Engineer's summary as evidence. Open the raw output files +named in `NOTEBOOK.md` and recompute at least one aggregate yourself. + +Confirm that the rerun command in `NOTEBOOK.md` names the script, its arguments +and the working directory. A command that cannot be pasted and run is +incomplete. + +Return `continue` when a number has no raw file behind it, and `done` only when +every checklist item of the current stage has evidence you inspected. diff --git a/examples/verticals/argus_verticals/lab_notebook/stages.py b/examples/verticals/argus_verticals/lab_notebook/stages.py new file mode 100644 index 000000000..491334251 --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/stages.py @@ -0,0 +1,81 @@ +"""Stages and checklists of the ``lab_notebook`` example vertical. + +Everything Argus needs to know about a vertical is declared at module level in +this file; there is no base class to inherit from. The framework reads the +attributes below through ``argus.core.vertical_contract.vertical_contract`` and +``argus.verticals._registry._validated_plugin``. +""" + +from __future__ import annotations + +from argus.skills.stage_machine import ChecklistItem + +# Required for every vertical that is not built in. Argus refuses any other +# API version, and an empty purpose hides the vertical from the Manager. +ARGUS_VERTICAL_API_VERSION = 1 +VERTICAL_PURPOSE = ( + "small measurement tasks on this machine: run the measurement, record how " + "it was run, and write a notebook entry another person can reproduce" +) + +# The stage order is the tuple order. Every stage that is not optional needs a +# non-empty checklist below. +CHECKLIST_STAGE_ORDER: tuple[str, ...] = ("measure", "report") +CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = () +# Names the Manager may use for a stage; they are canonicalized to the real one. +STAGE_ALIASES = {"experiment": "measure", "writeup": "report"} + +# "none": the project is complete when the last stage's checklist is met. +# "metric" and "certified" are the other two values the contract accepts. +completion_gate = "none" +WORKFLOW_MODE = "staged" # staged | direct | proportional +MISSION_KIND = "custom" # custom | optimize | research | software +REQUIRE_INDEPENDENT_REVIEW = True + +CHECKLIST_ITEMS: dict[str, tuple[ChecklistItem, ...]] = { + "measure": ( + ChecklistItem( + id="measure.ran-here", + statement=( + "The measurement was executed on this machine in this project, " + "and its raw output is saved in the work directory." + ), + evidence_hint="the command that was run and the path of its raw output", + ), + ChecklistItem( + id="measure.repeated", + statement=( + "The measurement was repeated, and the reported number is a " + "median or mean with its spread; a single run is not a result." + ), + evidence_hint="number of repeats, the aggregate and the spread", + ), + ), + "report": ( + ChecklistItem( + id="report.notebook-entry", + statement=( + "NOTEBOOK.md in the work directory states what was measured, " + "how, the result with its spread, and the exact command to rerun it." + ), + evidence_hint="the NOTEBOOK.md section and the rerun command", + ), + ChecklistItem( + id="report.numbers-traceable", + statement=( + "Every number in NOTEBOOK.md can be traced to a saved raw output file." + ), + evidence_hint="raw file paths next to each number", + ), + ), +} + + +def role_banner(role: str) -> str: + """One paragraph every role reads before its task; keep it about the field.""" + return ( + "LAB NOTEBOOK VERTICAL: measure first, then write. A number without a " + "saved raw output and a rerun command is not a result. The Reviewer " + "checks the notebook entry against the raw files, not against the " + "Engineer's summary." + ) diff --git a/examples/verticals/build_local_catalog.py b/examples/verticals/build_local_catalog.py new file mode 100644 index 000000000..f8dfee9a4 --- /dev/null +++ b/examples/verticals/build_local_catalog.py @@ -0,0 +1,97 @@ +"""Package one vertical directory as a Vertical Store archive plus a local catalog. + +Usage: + + python examples/verticals/build_local_catalog.py NAME --version 0.1.0 --out DIR + +Reads ``examples/verticals/argus_verticals/NAME`` (or ``--source``), writes +``DIR/NAME-VERSION.zip`` with members under ``argus_verticals/NAME/`` (the layout +``argus.verticals.store._verify_tree`` expects), and writes ``DIR/catalog.json`` +in the shape ``argus.verticals.store._validate_entry`` accepts. Point +``ARGUS_VERTICAL_CATALOG`` at that file and ``argus verticals install NAME`` uses +the archive next to it instead of downloading anything. + +The catalog is minimal on purpose: the fields the store requires, nothing else. +See docs/building-a-vertical.md for the walk-through. +""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import re +import sys +import zipfile +from pathlib import Path + +_NAME = re.compile(r"^[a-z][a-z0-9_]{0,47}$") + + +def _purpose(stages: Path) -> str: + """Read VERTICAL_PURPOSE from stages.py without importing it.""" + text = stages.read_text(encoding="utf-8") + match = re.search(r"VERTICAL_PURPOSE\s*=\s*\(?\s*((?:\"[^\"]*\"\s*)+)\)?", text) + if match is None: + raise SystemExit(f"{stages}: VERTICAL_PURPOSE not found") + return " ".join("".join(re.findall(r"\"([^\"]*)\"", match.group(1))).split()) + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0]) + parser.add_argument("name", help="vertical name, e.g. lab_notebook") + parser.add_argument("--version", default="0.1.0") + parser.add_argument("--source", type=Path, default=None, + help="vertical directory (default: examples/verticals/argus_verticals/NAME)") + parser.add_argument("--out", type=Path, required=True, help="directory for the zip and catalog.json") + args = parser.parse_args(argv) + + name = args.name + if not _NAME.fullmatch(name): + raise SystemExit(f"{name!r} is not a valid vertical name (^[a-z][a-z0-9_]{{0,47}}$)") + source = args.source or (Path(__file__).resolve().parent / "argus_verticals" / name) + stages = source / "stages.py" + if not stages.is_file(): + raise SystemExit(f"{source} has no stages.py") + + out = args.out.resolve() + out.mkdir(parents=True, exist_ok=True) + archive_name = f"{name}-{args.version}.zip" + archive = out / archive_name + tree = f"argus_verticals/{name}" + with zipfile.ZipFile(archive, "w", compression=zipfile.ZIP_DEFLATED) as zf: + for path in sorted(p for p in source.rglob("*") if p.is_file()): + if "__pycache__" in path.parts: + continue + zf.write(path, f"{tree}/{path.relative_to(source).as_posix()}") + + data = archive.read_bytes() + catalog = { + "schema": 1, + "verticals": { + name: { + "name": name, + "version": args.version, + "module": f"argus_verticals.{name}.stages", + "paths": [tree], + "purpose": _purpose(stages), + "has_skills": (source / "skills").is_dir(), + "archive": { + "file": archive_name, + "url": archive.as_uri(), + "sha256": hashlib.sha256(data).hexdigest(), + "size": len(data), + }, + } + }, + } + catalog_path = out / "catalog.json" + catalog_path.write_text(json.dumps(catalog, indent=2, ensure_ascii=False) + "\n", encoding="utf-8") + print(f"archive : {archive} ({len(data)} bytes)") + print(f"catalog : {catalog_path}") + print(f"install : ARGUS_VERTICAL_CATALOG={catalog_path} argus verticals install {name}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) From b33d793b9307294863dfe4d7ae082ff98b393018 Mon Sep 17 00:00:00 2001 From: LBX154 <145820328+lbx154@users.noreply.github.com> Date: Wed, 30 Sep 2026 03:45:25 -0700 Subject: [PATCH 2/2] List examples/ in the repository layout map The layout map must name every top-level directory; the guides' worked example directory was missing from it. Co-Authored-By: Claude Fable 5.1 --- docs/LAYOUT.md | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/LAYOUT.md b/docs/LAYOUT.md index d0ec9aa14..fb524d931 100644 --- a/docs/LAYOUT.md +++ b/docs/LAYOUT.md @@ -128,6 +128,7 @@ shipped" means no CI job, no test, and no wheel content comes from the directory - `deploy/` - systemd units and Dockerfiles for the hosted trial (`deploy/trial/`). - `desktop-tauri/` - the Tauri desktop shell and the PyInstaller spec (`argus_backend.spec`) for the frozen `argus-backend` binary (the spec's `name=`). - `docs/` - operator and developer documentation; `docs/audits/` holds dated audit reports and their data attachments. +- `examples/` - worked examples the guides refer to: `verticals/` holds the small `lab_notebook` vertical built in `docs/building-a-vertical.md` and the helper that packages a vertical into a local store catalog. Not built, not tested, not shipped. - `experiments/` - historical PR regression-study forwarding entry points and documentation (`pr_regression_50/`). The maintained implementation and offline tests live in `argus/release_tools/pr_gate/regression/` and `tests/tools/`; model/Docker studies are opt-in, not CI runs. This directory is not built or shipped in the wheel or sdist. - `frontend/` - `core` (shared TypeScript), `tui` (Ink terminal cockpit), `web` (React web cockpit). `frontend/web/dist` is committed on purpose and force-included into the wheel. - `integrations/` - the `agent-skills` package for external agent hosts (`SKILL.md` plus per-host adapters). Not the Python package `argus/integrations/`.