diff --git a/README.md b/README.md
index 37e397784..e993d7284 100644
--- a/README.md
+++ b/README.md
@@ -107,6 +107,12 @@ maintainers for the latest code.
Community Group 1 is full. Please join Group 2.
+## Tutorials
+
+- **[Getting started](docs/getting-started.md)** — install, first task, three ways in: command line, web UI, desktop app.
+- **[Best practices](docs/best-practices.md)** — writing an objective, changing direction mid-run, choosing the model and backend, controlling spend, what suits Argus.
+- **[Building a vertical](docs/building-a-vertical.md)** — a small real vertical from stages and skills to the Vertical Store.
+
## Quick Install
Choose the section for your operating system. Do not mix commands between
diff --git a/docs/LAYOUT.md b/docs/LAYOUT.md
index d0ec9aa14..fb524d931 100644
--- a/docs/LAYOUT.md
+++ b/docs/LAYOUT.md
@@ -128,6 +128,7 @@ shipped" means no CI job, no test, and no wheel content comes from the directory
- `deploy/` - systemd units and Dockerfiles for the hosted trial (`deploy/trial/`).
- `desktop-tauri/` - the Tauri desktop shell and the PyInstaller spec (`argus_backend.spec`) for the frozen `argus-backend` binary (the spec's `name=`).
- `docs/` - operator and developer documentation; `docs/audits/` holds dated audit reports and their data attachments.
+- `examples/` - worked examples the guides refer to: `verticals/` holds the small `lab_notebook` vertical built in `docs/building-a-vertical.md` and the helper that packages a vertical into a local store catalog. Not built, not tested, not shipped.
- `experiments/` - historical PR regression-study forwarding entry points and documentation (`pr_regression_50/`). The maintained implementation and offline tests live in `argus/release_tools/pr_gate/regression/` and `tests/tools/`; model/Docker studies are opt-in, not CI runs. This directory is not built or shipped in the wheel or sdist.
- `frontend/` - `core` (shared TypeScript), `tui` (Ink terminal cockpit), `web` (React web cockpit). `frontend/web/dist` is committed on purpose and force-included into the wheel.
- `integrations/` - the `agent-skills` package for external agent hosts (`SKILL.md` plus per-host adapters). Not the Python package `argus/integrations/`.
diff --git a/docs/best-practices.md b/docs/best-practices.md
new file mode 100644
index 000000000..ce1ee008f
--- /dev/null
+++ b/docs/best-practices.md
@@ -0,0 +1,328 @@
+# Argus best practices
+
+This page explains how to get good work out of Argus and why each
+recommendation holds. The reasons come from the code paths that read your
+input; the examples come from three real projects run on one machine on
+2026-09-30 (a GPU roofline research campaign, a damped-oscillator task, and a
+fresh-install trial), with times and costs taken from their logs. Nothing here
+is a quota; where a number appears it is a measurement, not a rule.
+
+Read [getting started](getting-started.md) first if you have not run a task
+yet.
+
+## Writing an objective
+
+The objective is read by four different consumers, and writing for all of
+them is what makes an objective good:
+
+1. **The Manager** decides from your text alone whether this is a
+ conversation, a small local job it can do itself, or team work for the
+ Planner, Engineer and Reviewer. It also decides the lifetime: finite,
+ casually worded work defaults to *bounded* (stop when done); only text that
+ expresses ongoing intent becomes a *standing* campaign that keeps generating
+ work.
+2. **The Manager again** picks the vertical (research, software, math, ...)
+ from the objective; there is no keyword classifier and no flag to force it,
+ so name the kind of work in plain words.
+3. **The Planner** decomposes it into tasks, each with an acceptance check.
+ What you leave unsaid, it will decide for you.
+4. **The Reviewer** judges every round against it and against the vertical's
+ stage checklist. The Reviewer runs read-only and cannot ask the Engineer;
+ it can only check what the objective and the work directory let it check.
+
+So an objective should say what must be true at the end, what evidence you
+will accept, what the scope and constraints are, and what form the result
+takes. Compare the two objectives that ran on this machine.
+
+The campaign objective (`objective-roofline.txt`, project `s-02d3c282`):
+
+> 做一项小规模但真实的研究:在本机 4 块 NVIDIA RTX A6000 上,PyTorch 方阵乘法的实际吞吐(TFLOPS)随矩阵规模(512 到 16384)和精度(FP32、TF32、BF16、FP16)如何变化,与理论峰值的差距能否用一个简单的 roofline 式模型(计算强度与显存带宽)解释。要求:真实实验、每种配置多次重复取中位数并报告离散度,拟合模型并给出误差,画图,最后写成一篇 4 页以内的短论文(含方法、结果、讨论、复现说明),并通过内部评审。所有数字必须来自本机真实运行,不允许估算或编造。
+
+Every clause did work. "小规模但真实的研究" and "短论文" put it in the `research`
+vertical (`pipeline : vertical=research` in the status). The hardware and the
+ranges fixed the scope, so the Planner did not have to guess a sweep. "多次重复取中位数并报告离散度"
+and "拟合模型并给出误差" became things the Reviewer could check: the draft's
+`results.tex` reports 17,280 samples, medians, and the fit's error on a held-out
+confirmation set. "通过内部评审" is why a `paper` round came back
+`replan_requested` instead of being accepted on the Engineer's word. And
+"所有数字必须来自本机真实运行,不允许估算或编造" is the sentence that lets the Reviewer
+refuse a plausible-looking number that has no raw file behind it.
+
+The small objective (`objective-oscillator.txt`, project `s-4a34a674`):
+
+> 用 numpy 数值求解一维阻尼谐振子 x'' + 2γx' + ω0² x = 0(取 ω0=2π, 分别取欠阻尼 γ=0.5、临界 γ=2π、过阻尼 γ=10),与解析解逐点比较给出最大误差,画出三条曲线,把方法、数值结果表和图写进本目录的 README.md。所有数字必须来自真实运行。
+
+It names the equation, the three parameter values, the comparison ("与解析解逐点比较给出最大误差"),
+the deliverable and its location. Because the acceptance check was in the
+text, the task ran once, was sent back once, and was done after 12 minutes and
+$0.91.
+
+The browser message from the trial (project `s-67fb6d62`) shows the casual end
+of the scale:
+
+> 帮我在这台机器的 GPU 上测一下 PyTorch 矩阵乘在 fp16 和 fp32 下的 TFLOPS,写个脚本跑出真实数字,整理成表格。
+
+"写个脚本跑出真实数字" and "整理成表格" were enough for a bounded, self-contained job
+that replied with a per-GPU table three minutes later. What it did not say
+(matrix sizes, repeats) the Engineer chose: 4096/8192/16384 and seven timed
+batches. If those choices matter to you, say them.
+
+Things to avoid in an objective, and why:
+
+- **Work that needs your credentials, your money, or an irreversible or
+ public action.** Those are authority boundaries: whatever the autonomy
+ setting, a task that reaches one stops and asks. Put the credential in
+ place first (the cockpit accepts credentials and stores them in the
+ capability vault) or keep the objective short of the boundary.
+- **A goal with no checkable end.** "Make it better" leaves the Reviewer
+ nothing to refuse, so rounds continue until a round cap or you intervene.
+- **Numbers you already know the answer to.** The roles will try to
+ reproduce them; that is the point, and it costs rounds.
+
+On the command line the objective must travel with `--continuous`
+(`--objective` alone is refused), and `--objective-file` keeps long text out of
+shell quoting. Add `--bounded` when the objective is finite; without it the
+worker keeps proposing follow-up work after the project is declared done.
+
+## Changing direction while it runs
+
+Argus gives you three ways to speak to a running project. They are not
+interchangeable, because they enter at different points of the loop.
+
+| You want to | Command line | Cockpit | Web API | What happens |
+|---|---|---|---|---|
+| Add guidance the next round should read | `argus --notify "text"` | `/nudge text` | `POST /api/projects/{sid}/nudge` | the text is queued in the project's durable inbox and spliced into the next Engineer round's prompt as operator guidance |
+| Hold guidance until a later stage | `argus --notify "text" --notify-stage paper` | | | same, delivered only when the project reaches that stage (aliases such as `writeup` are canonicalized; unknown stages are refused) |
+| Change the direction of the current task, pause or abort | (type it) | type it, or `/abort` | `POST /api/projects/{sid}/message` | the Manager classifies the message: `STEER` records a new direction for the active task, `PAUSE` stops the campaign, `ABORT` ends the task, `NO_DISPATCH` stops new work |
+| Set a rule for the whole project | (type it) | type it | `POST .../message` | text that sounds standing ("always", "never", "from now on", "do not ask") is stored as a project-wide directive; other amendments are scoped to the current task |
+| Answer a question the task stopped on | `argus --answer "text"` (`--answer-item ID` if several wait) | answer in the pending-question prompt | `POST /api/projects/{sid}/backlog/{item_id}/answer` | the paused task is replaced by a continuation whose objective carries your answer as authority, and the worker is restarted if needed |
+| Ask something without touching the run | `argus --ask "question"` | `/ask question` | `route_override: "chat"` on `/message` | the Manager answers from project state; nothing is queued |
+
+Why the nudge is deliberately weak: it is guidance the next round *reads*, not
+an order the runtime *enforces*. The inbox is durable (a queue under the
+project state that survives restarts, acknowledges only after delivery, and
+answers HTTP 429 when full), but the Engineer decides what to do with the
+text. That is the right tool for "prefer torch.compile off" or "the paper must
+cite X"; it is the wrong tool for "stop".
+
+Why `--answer` is not a nudge: a task that stopped to ask a question is
+`paused` until the question is answered. Clearing the question and resuming
+the same task looks like it works but does not: the task re-reads the
+objective that made it ask and asks again (the code comment on `--answer`
+records a campaign that burned five attempts that way). The answer path
+therefore enqueues a continuation whose objective states your answer, so the
+next round reads what it was told.
+
+Why a typed message is classified rather than obeyed literally: the same box
+carries greetings, questions, new tasks, settings changes and controls. The
+Manager first decides whether the message is conversation or work, and only
+when a task is active and the text clearly redirects it does the second check
+turn it into a `STEER`. Questions, criticism and suggestions are not steering;
+if you mean "change course", say so in the imperative.
+
+When Argus stops on its own to ask you is governed by one setting:
+
+```bash
+export ARGUS_SKILL_AUTONOMY_MODE=pragmatic # default
+```
+
+`pragmatic` recovers technical problems (a failed test, a timeout, a route
+that did not work) without asking and stops only at authority boundaries:
+credentials, payment or a bigger budget, deleting or force-pushing shared
+state, publishing or sending outward, and changes to an acceptance contract
+you own. `cautious` asks on every explicit question the Reviewer raises.
+`autonomous` still stops at those same boundaries; it only removes the
+remaining discretionary pauses. Pick `cautious` for the first run of a new
+vertical, when you want to see what it would have asked.
+
+Two real cases show what interjection looks like in practice.
+
+*The follow-up.* In the trial's browser project, after the first table
+arrived, the second message at 07:03:44Z was
+
+> 现在项目里都有什么文件?把 bf16 也补测一下,更新表格。
+
+Three minutes later the reply listed the project's files and gave the updated
+table with a BF16 column. That message is a new bounded task in the same
+project, not a steer of a running one; the earlier task had already finished.
+Most "changes of direction" on small tasks are like this, and the message box
+is the right place for them.
+
+*The pause that was not a question.* The roofline campaign's log shows, 43
+seconds after launch, a round ending with "The work was paused before the
+Engineer finished this round because provider charges are awaiting
+reconciliation or explicit risk approval, so this round was not judged".
+Nothing was asked of the operator; the cost check had refused an unpriced
+call. The operator's response was two CLI commands recorded in
+`daemon.commands.jsonl`: a `drain` at 08:16:12Z and a `start` at 08:21:19Z
+after changing the cost policy. Over the following two hours the campaign ran
+without a single operator message: all twenty entries in its inbox were
+background-job reports the runtime queued for itself. A precise objective is
+what made that possible.
+
+## Choosing the backend and model
+
+Argus resolves the backend and model for each role separately, and every
+resolution follows the same order:
+
+1. an environment variable for that role (`ARGUS_SKILL_ENGINEER_BACKEND`,
+ `ARGUS_SKILL_ENGINEER_MODEL`, ...),
+2. the shared variable (`ARGUS_SKILL_RUNNER_BACKEND`, `ARGUS_SKILL_MODEL`),
+3. the same names in the persisted settings file `~/.argus-skill/config.json`,
+ which is what `argus --setup`, `/backend`, `/config key=value` and a
+ natural-language "把模型换成 X" write,
+4. the default (`codex` for the backend; `auto` for the model, meaning the
+ backend's own default).
+
+A `--backend` typed on a `--daemon` launch sits above all of these for that
+daemon and is exported to its roles at boot, so a persisted choice can never
+silently override what you typed. `argus --config-help` prints every setting
+with its default, its current value and where the value came from
+(`env`, `persisted`, `default`); `argus --config-snapshot` writes the same to
+a file you can attach to a report.
+
+| Setting | Default | Notes |
+|---|---|---|
+| `ARGUS_SKILL_RUNNER_BACKEND` | `codex` | `codex`, `claude`, `copilot`, `cursor`, `opencode`, `pi`, `grok`, `qoder`, `dsh` |
+| `ARGUS_SKILL__BACKEND` | inherits | `ENGINEER`, `REVIEWER`, `PLANNER`, `MANAGER`, `SUPERVISOR` |
+| `ARGUS_SKILL_MODEL` | `auto` | a bare model id (`gpt-5.6-sol`, `copilot/opus-5`); free text reaches the CLI verbatim and every call then fails |
+| `ARGUS_SKILL__MODEL` | `auto` | `ENGINEER`, `REVIEWER`, `PLAN`, `MANAGER`, `SUPERVISOR` |
+| `ARGUS_SKILL_FRONTDOOR_MODEL` | `auto` | the cheap classifier that reads every message (`gpt-5.4-mini` on Copilot) |
+| `ARGUS_SKILL__REASONING_EFFORT` | Engineer `xhigh`, others `high`, Supervisor `low` | `low`, `medium`, `high`, `xhigh`, `max` |
+| `ARGUS_SKILL_PI_PROVIDER` / `ARGUS_SKILL_OPENCODE_PROVIDER` | unset | provider prefix for bare model ids on those two backends; on OpenCode the model is dropped without it |
+
+Why the roles are separable: they do different jobs. The Engineer runs long
+tool-using turns and benefits from the strongest model at high effort; the
+Reviewer reads and judges; the front door classifies one short message per
+turn and should be cheap. The roofline campaign ran on Copilot with
+`ARGUS_SKILL_MODEL=gpt-5.6-sol`, and its `usage.jsonl` splits as: $29.22 on
+`gpt-5.6-sol` across engineer, planner, reviewer, manager and supervision
+calls; $0.16 on `gpt-5.4-mini` for classification, reflection and fact
+extraction. Changing the front-door model would have saved nothing; changing
+the Engineer's would have changed everything.
+
+The same ledger explains where the money goes: 33.9 million input tokens
+against 0.3 million output tokens after two hours. Long tool-using rounds
+re-read context, which is why Argus by default sends the full task text only
+on the first round of a session (`ARGUS_SKILL_COMPACT_CONTINUATION_PROMPTS`)
+and why a precise objective with the evidence already in the work directory
+costs less than a vague one that makes the Engineer explore.
+
+## Controlling spend
+
+Spend is controlled at three levels, and the one that surprises new users is
+the third.
+
+**A daily cap across every project on the host.**
+`ARGUS_SKILL_GLOBAL_DAILY_CAP_USD` (default `1000.0`) and, if you prefer to
+count tokens, `ARGUS_SKILL_GLOBAL_DAILY_TOKEN_CAP` (default `0`, off).
+`argus --status` shows the running total: `budget : global daily $1000.00
+(spent $30.71) · remaining $969.29` was the line two hours into the roofline
+campaign, with `cost : $29.79 cumulative` for that project alone. The
+per-call record is `~/.argus-skill/projects//usage.jsonl` (model, tokens,
+`cost_usd`), and the web UI reads `GET /api/projects/costs`.
+
+**Provider-level circuit breakers.** `ARGUS_SKILL_CODEX_DAILY_CALL_CAP`
+(default 300 calls per day), `ARGUS_SKILL_COPILOT_DAILY_CALL_CAP`,
+`ARGUS_SKILL_COPILOT_DAILY_PREMIUM_CAP` and `ARGUS_SKILL_COPILOT_HOURLY_CALL_CAP`
+(default 10000 each), `ARGUS_SKILL_PROVIDER_MAX_CONCURRENCY` (default 0, off)
+and `ARGUS_SKILL_MAX_ACTIVE_DAEMONS` (default 64). These exist because a
+subscription CLI is shared by every project on the host, and one runaway
+campaign should not exhaust it for the others. `--mission-width` (default 2)
+is the per-project counterpart; the roofline campaign used `1`, which is the
+right choice when the tasks share four GPUs.
+
+**What to do with a call whose price is not known yet.**
+`ARGUS_SKILL_UNPRICED_COST_POLICY` is `block` by default: a call whose cost the
+provider has not settled is refused before it starts, because a cap that
+cannot see the cost cannot enforce anything. With Copilot's subscription
+billing the first calls on a fresh install report `pricing_status: partial`
+with no `cost_usd`, and that is exactly what happened in the trial:
+
+- CLI project `s-600e27bf`: the first classify call settled unpriced (its
+ event has `"pricing_status":"partial","cost_usd":null`), the worker went
+ to `paused_cost`, and `cost-control.json` listed the call as unresolved with
+ the reason "Copilot token billing is awaiting local CLI usage
+ reconciliation".
+- Browser project `s-67fb6d62`: the first message came back `[not dispatched]
+ Manager could not classify this message (refused before start: unresolved
+ provider cost: 1 call(s) awaiting usage reconciliation ...)`.
+- The roofline campaign, on a second install a little later, paused 43
+ seconds in for the same reason.
+
+The operator's fix each time was one line in `~/.argus-skill/config.json`,
+`"ARGUS_SKILL_UNPRICED_COST_POLICY": "allow"`, followed by a drain and restart
+(`/config ARGUS_SKILL_UNPRICED_COST_POLICY=allow` in the cockpit writes the
+same key, and the web configuration view exposes it). `allow` means: run the
+call now and let the ledger reconcile later. The trade-off is real. With a
+metered API key, keep `block`; the price of every call is known and the cap
+is exact. With a subscription CLI, `allow` is usually right, because the
+"cost" the ledger reconciles is your subscription's usage, and refusing to run
+does not save money you have already paid. The daily cap still applies to
+everything that does get priced.
+
+Two changes merged on 2026-09-30 (#179 and #183) move this default: Copilot
+calls made through the warm session are now priced from the CLI's own session
+log, and a call Argus itself interrupts before the CLI recorded anything is
+settled as one premium request. A fresh install on a subscription CLI no
+longer pauses under `block`; the three examples above ran before that change.
+`allow` is still the choice when you would rather run than wait for any late
+reconciliation at all.
+
+A last practical point: `ARGUS_SKILL_COST_CONTROL` (default `on`) is the
+switch for the whole admission-and-reconciliation layer. Leave it on; turning
+it off also removes the `paused_cost` state that tells you something is wrong
+with billing.
+
+## Which tasks suit Argus
+
+Argus is built for work that can be checked. The four-role loop only pays for
+itself when there is evidence for the Reviewer to read and refuse.
+
+It fits well when:
+
+- **The result is measurable on the machine it runs on.** Every task in the
+ three example projects was of this kind: a TFLOPS sweep, a numerical
+ solution compared point-by-point with an analytic one, a script whose
+ output table can be rerun. The Reviewer can open the raw files.
+- **The work has stages and takes hours.** The roofline campaign moved from
+ idea to experiment to a paper draft, launched background GPU jobs and
+ waited for them, and had a draft sent back for rework, all without an
+ operator message. That is what the daemon, the stage checklists and the
+ independent review exist for.
+- **There is a vertical for the field.** Seven ship with Argus (`research`,
+ `software`, `math`, ...) and the store carries seventeen more. A vertical is
+ what turns "done" into a checklist the Reviewer can apply; without one the
+ Manager still routes the task, but the acceptance standard is generic. See
+ [building a vertical](building-a-vertical.md).
+- **You want it to keep going.** A standing objective with `--continuous` and
+ no `--bounded` keeps proposing follow-up work in the same direction.
+
+It fits poorly when:
+
+- **You want an answer, not work.** `argus --ask` and `/ask` answer inline for
+ a fraction of a cent and queue nothing. Sending a question through the task
+ path costs a Manager turn plus, if it is misread as work, a Planner and
+ Engineer round.
+- **The task is smaller than the loop.** The oscillator task, a script most
+ people would write in ten minutes, took 12 minutes and $0.91 because it went
+ through Manager, Planner, an Engineer round, a Reviewer `continue`, a second
+ round and a `done`. That overhead is worth it when you would otherwise not
+ check the work; it is not worth it for a throwaway.
+- **Finishing requires something only you can do.** Logging in somewhere,
+ paying, publishing, deleting shared state, or approving a change to an
+ acceptance contract stops the task by design in every autonomy mode. Plan
+ the objective to end before that step, or do the step first.
+- **"Done" cannot be written down.** If you cannot state what the Reviewer
+ should check, it will accept or refuse on its own reading, and rounds will
+ continue until a cap.
+- **The evidence lives somewhere the roles cannot reach.** The Reviewer is
+ read-only and works from the work directory and the paths you allow
+ (`ARGUS_SKILL_REVIEWER_READ_DIRS`); a result that only exists in a service it
+ cannot query cannot be certified.
+
+A useful test before you start: write the sentence the Reviewer should be
+able to say at the end ("the README reports the maximum error against the
+analytic solution for all three damping cases, and each number matches the
+saved run"). If you can write it, put it in the objective. If you cannot,
+Argus will not be able to tell when it is finished either.
diff --git a/docs/building-a-vertical.md b/docs/building-a-vertical.md
new file mode 100644
index 000000000..bd794d9b0
--- /dev/null
+++ b/docs/building-a-vertical.md
@@ -0,0 +1,431 @@
+# Building a vertical
+
+A vertical teaches Argus what "done" means in one field: the stages work moves
+through, the checklist the Reviewer applies at each stage, and the skills the
+roles read before they start. This page builds a small real one,
+`lab_notebook`, from nothing to an installed entry in the Vertical Store. The
+finished files are under [`examples/verticals/`](../examples/verticals/) and
+every command below was run against this checkout.
+
+Background reading: [the Vertical Store](vertical-store.md) for the store's
+internals and hosted mode, and
+[single-agent verticals](single-agent-vertical-20260915.md) for the design
+discussion behind the contract.
+
+## What a vertical is, in code
+
+A vertical is a Python package whose `stages.py` declares a handful of
+module-level names. There is no base class and no registration call: the
+framework imports the module and validates the names through
+`argus.core.vertical_contract.vertical_contract`, and (for anything not built
+in) `argus.verticals._registry._validated_plugin`. The built-in seven are
+listed by name in `argus/skills/vertical_select.py`; everything else arrives
+through the store or a Python entry point in the group `argus.verticals`.
+
+The names the contract reads:
+
+| Name | Required | Meaning |
+|---|---|---|
+| `ARGUS_VERTICAL_API_VERSION` | yes, for non-built-ins | must be `1`; any other value hides the vertical |
+| `VERTICAL_PURPOSE` | yes, for non-built-ins | one sentence; it is the line the Manager sees when choosing a vertical, so write it for routing, not as an abstract |
+| `CHECKLIST_STAGE_ORDER` | yes | tuple of stage names, in order; no duplicates |
+| `CHECKLIST_ITEMS` | yes | dict `stage -> tuple[ChecklistItem, ...]`; every non-optional stage needs a non-empty checklist |
+| `completion_gate` | yes | `"none"`, `"metric"` or `"certified"` |
+| `CHECKLIST_OPTIONAL_STAGES` | no | stages that may have no checklist |
+| `STAGE_ALIASES` | no | `{"experiment": "measure"}`: other names the Manager may use for a stage |
+| `WORKFLOW_MODE` | no | `"staged"`, `"direct"` or `"proportional"` |
+| `MISSION_KIND` | no | `"custom"`, `"optimize"`, `"research"` or `"software"` |
+| `REQUIRE_INDEPENDENT_REVIEW` | no | defaults to `True` |
+| `role_banner(role)` | no | a paragraph each role reads before its task |
+| `render_role_prompt_fragment`, `render_role_prompt_context`, `stage_completion_issues`, `prepare_mission`, `LIBRARY_PREPARER`, `EVIDENCE_SCHEMA`, `WORKFLOW_PROFILES` | no | hooks the larger built-ins use; not needed for a first vertical |
+| `VERTICAL_SKILLS`, `VERTICAL_SKILL_PARENTS`, `VERTICAL_ROUTING_PATH` | no | an explicit skills root (a `skills/` directory next to `stages.py` is found automatically), verticals whose skills are seeded before yours, and a browsing category for the store |
+
+A `ChecklistItem` (from `argus/skills/stage_machine.py`) has three fields:
+
+```python
+@dataclass(frozen=True)
+class ChecklistItem:
+ id: str
+ statement: str
+ evidence_hint: str
+```
+
+The `statement` is what must be true; the `evidence_hint` tells the Engineer
+what to show and the Reviewer what to look for. This is the whole mechanism:
+the Reviewer is read-only, so the quality of your checklist is the quality of
+your review.
+
+## Two built-ins to learn from
+
+**`software`** (`argus/verticals/software/`) is the smallest useful vertical:
+one stage, `delivery`, with three checklist items, a `role_banner`, and four
+skill files. Its `stages.py` is worth reading in full; the module docstring
+explains why the checklist exists ("code that does not build" was the failure
+it was written against), and the first two item ids are marked protected so a
+Planner may add items but never weaken them:
+
+```python
+STAGE_ORDER = ["delivery"]
+CHECKLIST_STAGE_ORDER = tuple(STAGE_ORDER)
+CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = ()
+completion_gate = "none"
+MISSION_KIND = "software"
+WORKFLOW_MODE = "staged"
+
+PROTECTED_ITEM_IDS = frozenset({"delivery.builds", "delivery.tests-executed"})
+```
+
+Its skills sit at `skills//.md`:
+
+```
+software/skills/engineer/software-change-implementation.md
+software/skills/manager/software-project-grounding.md
+software/skills/planner/software-project-grounding.md
+software/skills/reviewer/software-change-review.md
+```
+
+**`research`** (`argus/verticals/research/`) is the largest: four stages,
+
+```python
+CANONICAL_STAGE_ORDER: tuple[str, ...] = ("idea", "experiment", "paper", "review")
+STAGE_ALIASES = {"research": "idea", "plan": "experiment", ..., "submission": "review"}
+```
+
+with `completion_gate = "certified"`, `WORKFLOW_MODE = "proportional"`, role
+prompts in `prompt_policy.py`, a `prepare_mission` hook in `mission_brief.py`,
+and some forty engineer skills plus seven reviewer skills. It shows where a
+vertical can grow; it is not where one should start.
+
+The directory shape both share, and the store expects:
+
+```
+/
+ __init__.py docstring only
+ stages.py the contract
+ skills/
+ engineer/*.md
+ reviewer/*.md
+ manager/*.md (optional)
+ planner/*.md (optional)
+ engineer/references/ supporting files, not skills (optional)
+ engineer/*_scripts/ .py/.json/.sh copied verbatim (optional)
+```
+
+## Step 1: decide the stages and the checklist
+
+`lab_notebook` is for small measurement tasks: run something on this machine,
+then write it up so someone else can rerun it. Two stages follow from that,
+`measure` and `report`, and each gets the two or three things a reviewer
+would actually check.
+
+Write the checklist before anything else, and write it as statements a
+read-only reviewer can verify from files. "The measurement was repeated" is
+checkable (there are N raw files); "the measurement is good" is not.
+
+`examples/verticals/argus_verticals/lab_notebook/stages.py`:
+
+```python
+from argus.skills.stage_machine import ChecklistItem
+
+ARGUS_VERTICAL_API_VERSION = 1
+VERTICAL_PURPOSE = (
+ "small measurement tasks on this machine: run the measurement, record how "
+ "it was run, and write a notebook entry another person can reproduce"
+)
+
+CHECKLIST_STAGE_ORDER: tuple[str, ...] = ("measure", "report")
+CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = ()
+STAGE_ALIASES = {"experiment": "measure", "writeup": "report"}
+
+completion_gate = "none"
+WORKFLOW_MODE = "staged"
+MISSION_KIND = "custom"
+REQUIRE_INDEPENDENT_REVIEW = True
+
+CHECKLIST_ITEMS: dict[str, tuple[ChecklistItem, ...]] = {
+ "measure": (
+ ChecklistItem(
+ id="measure.ran-here",
+ statement=("The measurement was executed on this machine in this project, "
+ "and its raw output is saved in the work directory."),
+ evidence_hint="the command that was run and the path of its raw output",
+ ),
+ ChecklistItem(
+ id="measure.repeated",
+ statement=("The measurement was repeated, and the reported number is a "
+ "median or mean with its spread; a single run is not a result."),
+ evidence_hint="number of repeats, the aggregate and the spread",
+ ),
+ ),
+ "report": (
+ ChecklistItem(
+ id="report.notebook-entry",
+ statement=("NOTEBOOK.md in the work directory states what was measured, "
+ "how, the result with its spread, and the exact command to rerun it."),
+ evidence_hint="the NOTEBOOK.md section and the rerun command",
+ ),
+ ChecklistItem(
+ id="report.numbers-traceable",
+ statement="Every number in NOTEBOOK.md can be traced to a saved raw output file.",
+ evidence_hint="raw file paths next to each number",
+ ),
+ ),
+}
+
+
+def role_banner(role: str) -> str:
+ return (
+ "LAB NOTEBOOK VERTICAL: measure first, then write. A number without a "
+ "saved raw output and a rerun command is not a result. The Reviewer "
+ "checks the notebook entry against the raw files, not against the "
+ "Engineer's summary."
+ )
+```
+
+Why these particular choices:
+
+- `completion_gate = "none"` means the project is complete once the last
+ stage's checklist is met. `"metric"` is for verticals whose completion is a
+ number reaching a target; `"certified"` adds an explicit certification step,
+ which the research vertical uses for papers. A notebook entry needs neither.
+- `WORKFLOW_MODE = "staged"` makes the Manager move through the stages in
+ order. `"direct"` skips staging for one-shot work; `"proportional"` lets the
+ Manager scale the process to the task. Start with `staged`; the aliases let
+ a Manager that says "experiment" land on `measure`.
+- `MISSION_KIND = "custom"` because the other three values switch on
+ behaviour written for optimization loops, research campaigns and repository
+ changes.
+- The `role_banner` says one thing, in the field's own terms. It is read by
+ every role, so it is the place for the rule that ties the stages together.
+
+The contract check refuses, with a message naming the problem, a stage with no
+checklist, a checklist for a stage that is not in the order, a duplicate stage,
+an item without an `id` or `statement`, a repeated item id within a stage, and
+an unknown `completion_gate` value. Run it directly while you iterate:
+
+```bash
+python -c "
+import importlib.util, pathlib
+from argus.core.vertical_contract import vertical_contract
+p = pathlib.Path('examples/verticals/argus_verticals/lab_notebook/stages.py')
+spec = importlib.util.spec_from_file_location('lab_notebook_stages', p)
+m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
+c = vertical_contract('lab_notebook', m)
+print(c.stage_order, c.completion_gate, c.workflow_mode)
+"
+```
+
+## Step 2: write the skills
+
+A skill is a Markdown file with exactly two front-matter fields, `name` and
+`description`, both quoted. The runtime does not parse skill bodies: the
+roles receive the library paths and read the files themselves, so the
+description is a routing line (it is what an agent reads to decide whether to
+open the file; the repository's own test caps it at 1200 characters) and the
+body is written for a reader who will act on it.
+
+`skills/engineer/measurement-record.md`:
+
+```markdown
+---
+name: "Measurement Record"
+description: "How to run a small measurement so the result can be checked and rerun: save raw output, repeat, report the spread, and record the exact command."
+---
+
+# Measurement record
+
+Run the measurement from a script saved in the work directory, never from an
+interactive shell you cannot show later. Write the raw output to a file under
+`results/` before computing any aggregate.
+
+Repeat the run. Report the median (or mean) together with the spread (min/max or
+standard deviation) and the number of repeats. A single run is not a result.
+
+Record in `NOTEBOOK.md`:
+
+- what was measured and on which hardware
+- the exact command to rerun it
+- the aggregate, the spread and the number of repeats
+- the path of the raw output every number came from
+```
+
+`skills/reviewer/measurement-review.md` tells the Reviewer to open the raw
+files and recompute one aggregate before accepting, and to return `continue`
+for any number without a file behind it. The two skills and the two checklists
+say the same thing from three sides; that redundancy is deliberate, because
+each role reads only its own.
+
+Rules the loader applies: skills live at `skills//` where the role is
+`engineer`, `reviewer`, `planner` or `manager` (anything else is filed as
+general); a `references/` directory is copied as supporting material rather
+than as skills; `*_scripts/` directories are copied verbatim; names starting
+with `_` or `.` are skipped. When the roles start a task with this vertical,
+those files are seeded into the project's skill library in the order project,
+vertical, global, so a project's own edits win.
+
+## Step 3: package it for the store
+
+The store installs a zip whose members sit under `argus_verticals//`,
+described by a `catalog.json`. The catalog entry the store validates
+(`argus/verticals/store.py`, `_validate_entry`) needs:
+
+| Field | Constraint |
+|---|---|
+| `name` | equals the key; `^[a-z][a-z0-9_]{0,47}$` |
+| `version` | non-empty, no slashes |
+| `module` | `argus_verticals..stages`, and it must live in `paths[0]` |
+| `paths` | at least `["argus_verticals/"]` |
+| `purpose` | non-empty |
+| `archive.file`, `archive.url`, `archive.sha256`, `archive.size` | the zip's name, location, digest and exact byte size; a non-`https` URL is only accepted when the catalog itself is local |
+
+Optional: `purpose_zh`, `requires` (other verticals to install first),
+`shared` (helper trees outside the vertical's own directory),
+`python_requirements` (shown, never installed), `tags`, `skill_parents`,
+`has_skills`, `routing_path`, `min_argus`, `maintainers`.
+
+`examples/verticals/build_local_catalog.py` produces both files for a directory
+under `examples/verticals/argus_verticals/`:
+
+```bash
+python examples/verticals/build_local_catalog.py lab_notebook --version 0.1.0 --out /tmp/lab-store
+```
+
+```
+archive : /tmp/lab-store/lab_notebook-0.1.0.zip (3099 bytes)
+catalog : /tmp/lab-store/catalog.json
+install : ARGUS_VERTICAL_CATALOG=/tmp/lab-store/catalog.json argus verticals install lab_notebook
+```
+
+The catalog it wrote:
+
+```json
+{
+ "schema": 1,
+ "verticals": {
+ "lab_notebook": {
+ "name": "lab_notebook",
+ "version": "0.1.0",
+ "module": "argus_verticals.lab_notebook.stages",
+ "paths": ["argus_verticals/lab_notebook"],
+ "purpose": "small measurement tasks on this machine: run the measurement, record how it was run, and write a notebook entry another person can reproduce",
+ "has_skills": true,
+ "archive": {
+ "file": "lab_notebook-0.1.0.zip",
+ "url": "file:///tmp/lab-store/lab_notebook-0.1.0.zip",
+ "sha256": "5b36…b8e0",
+ "size": 3099
+ }
+ }
+ }
+}
+```
+
+The `sha256` and `size` are checked byte-for-byte at install; regenerate the
+catalog whenever the zip changes.
+
+## Step 4: install it and confirm Argus sees it
+
+Point the store at the local catalog. A local catalog turns its directory into
+an offline mirror: an archive named in the catalog and sitting next to it is
+used without any download.
+
+```bash
+export ARGUS_VERTICAL_CATALOG=/tmp/lab-store/catalog.json
+argus verticals install lab_notebook
+argus verticals list
+argus verticals info lab_notebook
+```
+
+What this checkout printed (the install was done into a throwaway
+`ARGUS_SKILL_HOME` so the machine's real store stayed untouched):
+
+```
+lab_notebook: install started
+ [ 0%] lab_notebook: starting
+ [100%] lab_notebook: install finished
+lab_notebook: install finished
+```
+
+```
+name kind version installed enabled update purpose
+...
+lab_notebook installed 0.1.0 0.1.0 yes - small measurement tasks on this machine: run the measurement
+```
+
+```
+name lab_notebook
+kind installed
+purpose small measurement tasks on this machine: run the measurement, record how it was run, and write a notebook entry another person can reproduce
+version 0.1.0
+installed_version 0.1.0
+enabled True
+actions disable, uninstall
+operation install done: install finished
+```
+
+The files landed at `/verticals/argus_verticals/lab_notebook/`
+next to the store's `registry.json`, and the registry advertises the vertical
+with origin `store`, its two skills found automatically next to `stages.py`:
+
+```
+advertised: True
+origin: store | skills_root: .../verticals/argus_verticals/lab_notebook/skills
+skills: ['engineer/measurement-record.md', 'reviewer/measurement-review.md']
+stages: ('measure', 'report') | gate: none | mode: staged
+```
+
+From here the Manager can choose `lab_notebook` for a task whose text matches
+its purpose line; there is no flag to force a vertical (`ARGUS_SKILL_VERTICAL`
+is a legacy name with no authority), and the choice is saved per project in
+`.argus/PIPELINE_STATE.json`. To try it, start a project with an objective in
+the vertical's own words, for example "measure how long `python -c 'import
+torch'` takes on this machine, repeat it, and write a notebook entry with the
+rerun command", and check `pipeline : vertical=lab_notebook` in
+`argus --status`.
+
+`argus verticals disable lab_notebook` hides it again without removing the
+files; `argus verticals remove lab_notebook` deletes them (it refuses while a
+local project's `PIPELINE_STATE.json` still names the vertical, unless you
+pass `--force`).
+
+## Publishing to the shared store
+
+The default catalog is the `catalog.json` attached to the latest release of
+[Argus-AiTeam/argus-verticals](https://github.com/Argus-AiTeam/argus-verticals);
+each vertical is one directory in that repository, and its release workflow
+builds the zips and the catalog (`scripts/build_catalog.py` there). Publishing
+therefore means opening a pull request that adds
+`argus_verticals//` to that repository, with the same layout as here. The
+store on every user's machine only accepts archives from an allow-listed host
+(`github.com` and its release asset hosts), so a catalog you host elsewhere
+must be reached through `ARGUS_VERTICAL_CATALOG` as a local file, which is what
+the steps above do.
+
+
+
+Two other ways exist to run a vertical without the store, useful during
+development of a bigger one: `pip install` a package that declares the entry
+point `argus.verticals` (`name = "argus_verticals..stages"`), which wins
+over a store copy of the same name; or install it as a managed workbench
+plugin. A name that collides with a built-in is ignored from every source.
+
+## Checks before you share it
+
+- The contract check in step 1 passes.
+- Every skill header has quoted `name` and `description` values. The
+ repository's own skill writer (`argus/skills/store.py`) emits exactly that
+ shape, JSON-quoted, because an unquoted description containing a colon does
+ not survive a YAML parse; `tests/skills/test_skill_frontmatter_integrity.py`
+ checks the in-tree verticals for the same reason.
+- `argus verticals install` from a local catalog succeeds and
+ `argus verticals list` shows the row as `installed` and enabled.
+- One real task ran through it and the Reviewer's verdicts referred to your
+ checklist ids (they appear in the round events in `events.jsonl`).
+
+The example under `examples/verticals/` is not scanned by
+`tests/skills/test_vertical_plugins.py` (that file builds synthetic plugins in
+memory), so adding a vertical there cannot change the test's result. On this
+machine one test in that file fails before and after these changes for a
+local reason: a pip-installed `argus_verticals` package in the venv makes
+`test_store_verticals_are_discovered_with_origin_store` see two package paths
+instead of one.
diff --git a/docs/getting-started.md b/docs/getting-started.md
new file mode 100644
index 000000000..a2a3de193
--- /dev/null
+++ b/docs/getting-started.md
@@ -0,0 +1,419 @@
+# Getting started with Argus
+
+This guide takes you from an empty machine to a finished first task. It covers
+three ways to work: the command line, the web UI, and the desktop app. Every
+command below was checked against the `argus` CLI in this checkout
+(`argus 0.1.8`); the example runs are real projects, with their times and costs
+taken from their own logs.
+
+Companion guides: [best practices](best-practices.md) (objectives, changing
+direction mid-run, models and spend) and
+[building a vertical](building-a-vertical.md).
+
+## What you are installing
+
+Argus is a Python package plus a bundled terminal cockpit (Node.js) and a web
+UI. It does not call a model API itself; it drives one of the coding-agent CLIs
+you already use (Copilot, Codex, Claude Code, Cursor, Pi, OpenCode, Grok, Qoder,
+DeepSeek Harness) and lets four roles share it: a Manager that routes your
+message, a Planner that decides the next task, an Engineer that does the work,
+and a read-only Reviewer that decides whether the work is done.
+
+Two things follow from that design and explain most of what you will see:
+
+- **A background worker does the work.** Your cockpit, browser tab or terminal
+ can close; the project's daemon keeps running and keeps its state under
+ `~/.argus-skill/projects//`.
+- **Nothing is "done" until the Reviewer says so.** A task usually takes
+ several Engineer rounds. Expect the first small task to take minutes, not
+ seconds.
+
+## Prerequisites
+
+| Requirement | Why |
+|---|---|
+| Python 3.11+ | the `argus` package (`pyproject.toml` requires `>=3.11`) |
+| Node.js 22.12+ | the terminal cockpit is an Ink app; `argus` refuses to start it on older Node (`argus: Ink TUI requires Node.js 22.12 or newer`) |
+| One authenticated agent CLI | Argus reuses its login; there is no separate Argus account |
+
+Install and log into one CLI first. The README's
+[Quick Install](../README.md#quick-install) table lists the install and login
+command for each backend; for example GitHub Copilot CLI is
+`npm install -g @github/copilot` then `copilot login`.
+
+## Install
+
+The commands are the README's, repeated here so this page stands alone.
+
+Linux (isolated venv; use this on servers):
+
+```bash
+git clone https://github.com/microsoft/ArgusAgent.git "$HOME/Argus"
+cd "$HOME/Argus"
+python3 -m venv .venv
+.venv/bin/python -m pip install --upgrade pip
+.venv/bin/python -m pip install -e .
+ARGUS_BIN="$HOME/Argus/.venv/bin/argus"
+"$ARGUS_BIN" --version
+```
+
+Windows (PowerShell, no venv):
+
+```powershell
+py -m pip install --upgrade pip
+py -m pip install --upgrade --force-reinstall "argus @ https://github.com/microsoft/ArgusAgent/archive/refs/heads/main.zip"
+```
+
+macOS (managed command):
+
+```bash
+uv tool install --force --python 3.12 "argus @ https://github.com/microsoft/ArgusAgent/archive/refs/heads/main.zip"
+```
+
+On Linux, `argus` below means `$HOME/Argus/.venv/bin/argus` unless the venv is
+active. Everywhere, `python -m argus ...` runs the same CLI without ever
+starting the Node cockpit; it is what you want in scripts and over SSH without
+a terminal.
+
+## First-time setup and a health check
+
+```bash
+argus --setup
+argus doctor
+```
+
+`argus --setup` is an interactive wizard: it asks which backend to use, checks
+that the CLI is installed and logged in, runs one real model turn through it,
+and only then saves the profile. It ends with `Setup complete. Run `argus`.`
+If you would rather not answer prompts:
+
+```bash
+argus --setup --backend copilot --non-interactive
+```
+
+Two variants worth knowing:
+
+| Situation | Command |
+|---|---|
+| Keep Argus on its own Copilot account, separate from your interactive login | `argus --setup --backend copilot --copilot-home "$HOME/.copilot-argus" --copilot-login` |
+| Use an OpenAI-compatible endpoint (served through Pi) | `argus --setup --api-url URL --api-model MODEL` with the key in `ARGUS_SETUP_API_KEY` rather than `--api-key`, so it stays out of shell history |
+
+`argus doctor` is read-only. This is what it printed on the machine this guide
+was written on (a source checkout, Pi backend):
+
+```
+argus doctor — cross-platform diagnostics
+
+✓ ARGUS-HOST-001 [host/supported_host] Linux 6.8.0-139-generic x86_64
+✓ ARGUS-INSTALL-001 [install/source_checkout] /data/.../Argus
+✓ ARGUS-ASSET-001 [install/assets_ready] release manifest and Web/TUI assets are present
+✓ ARGUS-PYTHON-001 [cli/python_ready] 3.12.3 at /data/.../.venv/bin/python
+✓ ARGUS-NODE-001 [cli/node_ready] v22.23.2
+✓ ARGUS-WEB-001 [web/compatible] 127.0.0.1:8799 is compatible
+! ARGUS-DESKTOP-001 [desktop/tauri_dependencies_missing] Tauri Desktop sources exist but npm/Rust build dependencies are incomplete
+ fix: run `npm --prefix desktop-tauri ci` and install the Rust Windows toolchain
+✓ ARGUS-DAEMON-001 [daemon/stopped] no daemon is running
+✓ ARGUS-BACKEND-001 [backend/ready] pi 0.85.1 runnable at /home/.../argus-pi (subscription_cli; authentication checked; ...)
+
+all blocking checks passed
+```
+
+A `!` line is advice, not a failure; the desktop warning above only matters if
+you intend to build the desktop app from this checkout. When something is
+wrong, `argus doctor --fix-safe` applies the repairs the doctor itself marked
+safe and reruns; `argus doctor --advisor auto` lets one of the installed agent
+CLIs inspect and repair. Without one of those two options `doctor` changes
+nothing.
+
+A note on help: `argus --help` (and `argus doctor --help`,
+`argus verticals --help`) print the same short overview. The complete flag
+reference is behind an environment variable:
+
+```bash
+ARGUS_SKILL_DEBUG_HELP=1 argus --help
+```
+
+## Path 1: the command line
+
+There are two ways to use the CLI. A bare `argus` opens the terminal cockpit,
+where you type in natural language. Everything else (`--daemon`, `--status`,
+`--notify`, ...) is a one-shot command that talks to the project's worker and
+exits; those work over SSH, in cron, and without Node.
+
+### Which project a command means
+
+Argus keys a project's state on the directory you run it from. Management
+commands attach to the newest session for the current directory, or to the
+one you name:
+
+| Flag | Meaning |
+|---|---|
+| (none) | the newest session for the current working directory |
+| `--project-root DIR` | the newest session for `DIR` |
+| `--resume ID` | a specific session id (with no id, a picker of recent sessions) |
+| `--continue` | the most recently active session |
+| `--life-dir DIR` | use `DIR` instead of `~/.argus-skill` as the state root (`ARGUS_SKILL_HOME` does the same) |
+
+### Start a task without opening the cockpit
+
+Write the objective in a file, `cd` into the directory the work should happen
+in, and start a background worker with a campaign objective:
+
+```bash
+cd ~/work/matmul-roofline
+argus --daemon --continuous --objective-file objective.txt --bounded --backend copilot
+```
+
+The pieces:
+
+| Flag | What it does and why you would set it |
+|---|---|
+| `--daemon` | start a detached worker (`--daemon-fg` keeps it in the foreground, for systemd or debugging) |
+| `--continuous` | give the worker an objective and let the Planner keep generating tasks toward it; refused without an objective (`--continuous requires a non-empty --objective`) |
+| `--objective TEXT` / `--objective-file PATH` | the objective; the file form keeps long text and quotes out of the shell (`--objective` without `--continuous` is refused too) |
+| `--bounded` | stop when the Planner certifies the project done; without it the worker keeps generating work for the same objective |
+| `--backend NAME` | the agent CLI for this daemon; it outranks the persisted setup for this launch and is exported to every role |
+| `--mission-width N` | how many tasks may run in parallel (default 2; `1` is serial) |
+
+Before the worker starts, Argus re-checks that the backend is installed and
+logged in; if not, it prints the readiness report and exits with code 3 instead
+of starting a worker that cannot call a model.
+
+### A real first campaign
+
+The objective below started the campaign that project `s-02d3c282` ran on this
+machine on 2026-09-30 (`objective-roofline.txt`, quoted in full because its
+shape is what makes the rest work):
+
+> 做一项小规模但真实的研究:在本机 4 块 NVIDIA RTX A6000 上,PyTorch 方阵乘法的实际吞吐(TFLOPS)随矩阵规模(512 到 16384)和精度(FP32、TF32、BF16、FP16)如何变化,与理论峰值的差距能否用一个简单的 roofline 式模型(计算强度与显存带宽)解释。要求:真实实验、每种配置多次重复取中位数并报告离散度,拟合模型并给出误差,画图,最后写成一篇 4 页以内的短论文(含方法、结果、讨论、复现说明),并通过内部评审。所有数字必须来自本机真实运行,不允许估算或编造。
+
+It was launched with `--backend copilot --mission-width 1` (the live process
+still shows those flags). About two hours in, `argus --status` reported:
+
+```
+ project : .../projects/s-02d3c282
+ daemon : alive (pid 1430124, up 1h 48m, backend live — see /roles, width 1)
+ budget : global daily $1000.00 (spent $30.71) · remaining $969.29
+ active : 0 pending · 1 running · 0 paused
+ current :
+ title : 重绘正式数据图并完成论文稿
+ history : 2 done
+ cost : $29.79 cumulative
+ continuous: on
+ pipeline : vertical=research
+ lifecycle:
+ state : writing
+```
+
+By then the work directory held `paper/main.tex`, `paper/results.tex`,
+`results/attempt-01`, `figures/` and `src/`; `paper/results.tex` reports, at
+n=16384, 22.81 / 57.01 / 121.11 / 115.96 TFLOP/s for FP32 / TF32 / BF16 / FP16
+(58.95 / 73.66 / 78.24 / 74.91 % of the spec sheet) over 17,280 samples, and a
+fitted roofline whose calibrated form has a 53.92 % mean absolute error on the
+confirmation set. The campaign was still running when this page was written
+(its log had a `paper` round returned as `replan_requested`, which is normal:
+the Reviewer sent the draft back), so treat those as a snapshot, not a result.
+
+For a small first task, a bounded objective finishes in minutes. Project
+`s-4a34a674` on the same machine was started the same way with this objective:
+
+> 用 numpy 数值求解一维阻尼谐振子 x'' + 2γx' + ω0² x = 0(取 ω0=2π, 分别取欠阻尼 γ=0.5、临界 γ=2π、过阻尼 γ=10),与解析解逐点比较给出最大误差,画出三条曲线,把方法、数值结果表和图写进本目录的 README.md。所有数字必须来自真实运行。
+
+Its log runs from 08:11:39Z to 08:23:33Z (12 minutes), one task in two
+attempts, `continuous.json` ends with `"done_reason": "planner declared project
+done"`, and it cost $0.91 (12 model calls in `usage.jsonl`). The README it
+wrote gives maximum errors of 6.0e-07, 1.6e-07 and 5.1e-07 for the three
+damping cases against the analytic solution.
+
+### Watch, steer, stop
+
+| Command | What it does |
+|---|---|
+| `argus --status` | one screen: daemon, budget, current task, stage, last events |
+| `argus --follow` | stream the event log to the terminal (`tail -f` style, Ctrl-C to stop) |
+| `argus --watch` | the read-only live cockpit |
+| `argus --notify "text"` | queue guidance for the next Engineer round (`--notify-stage STAGE` holds it until that stage) |
+| `argus --answer "text"` | answer the question a paused task is waiting on (`--answer-item ID` when several wait) |
+| `argus --ask "question"` | ask the Manager something and exit; nothing is queued and no daemon is needed |
+| `argus --daemon-stop --drain` | let the current task finish at a clean boundary, then exit; the safe way to stop before upgrading |
+| `argus --daemon-stop --force` | SIGKILL if it does not exit in time (interrupts running work) |
+| `argus --daemon-runbook` | print the restart playbook for the current project |
+
+The difference between `--notify` and `--answer` matters: a nudge is read at
+the next round and does not wake a task that has stopped to ask you something;
+`--answer` does, and it also rewrites the task so the next round reads your
+answer instead of re-asking. [Best practices](best-practices.md#changing-direction-while-it-runs)
+goes into this.
+
+### The terminal cockpit
+
+```bash
+argus
+```
+
+The cockpit needs a real terminal (piped or cron use gets a message pointing
+you to `--web`, `--watch`, `--status` or `--daemon`). Type what you want in
+plain language. The Manager reads every message and decides whether to answer
+it itself (a question, a status request, a small local check) or to turn it
+into a task for the Planner, Engineer and Reviewer (multi-step work, anything
+that needs an independent review). You do not choose a vertical, a model or a
+backend for the first task; the Manager picks the vertical from your text and
+the backend comes from setup.
+
+`/new` (optionally `/new `) opens a two-field form, Name and
+Objective (Tab or arrows switch fields, Enter creates), and switches the
+cockpit to the new session; its worker starts when the first task arrives.
+
+Commands you will use in the first hour (`/help` lists them all):
+
+| Command | Purpose |
+|---|---|
+| `/task ` | queue work directly, skipping the "is this a chat?" decision |
+| `/ask ` | answer inline; no task is queued |
+| `/plan ` | preview the Planner's execution plan before committing |
+| `/status`, `/roles`, `/backlog`, `/journal` | what is running, on which backend/model, what is queued, what happened |
+| `/nudge ` | inject guidance into the running task |
+| `/abort` | stop the running task now |
+| `/backend [name]`, `/config [key=value]` | view or change the runner backend and runtime settings (persisted) |
+| `/resume`, `/daemons` | switch to another project |
+| `/quit` | leave; background work keeps running |
+
+## Path 2: the web UI
+
+```bash
+argus --web
+```
+
+From the `argus` command this goes through the cockpit launcher, which picks
+the first free port from 8799 and opens your browser at
+`http://127.0.0.1:8799`. With an explicit host or port, or through
+`python -m argus --web`, the Python server runs directly and does not open a
+browser:
+
+```bash
+argus --web --web-port 8800
+python -m argus --web # over SSH: then `ssh -L 8799:127.0.0.1:8799 user@server`
+```
+
+Binding to a LAN address (`--web-host 0.0.0.0`) always requires a bearer
+token: `ARGUS_SKILL_WEB_TOKEN` if set, otherwise one is minted for the run and
+printed with a QR code. The web assets are checked into the repository and
+packaged with the wheel, so nothing needs building first.
+
+### The first screen
+
+With no sessions yet the page shows **Start a project**, the line *No sessions
+yet. Create one to begin.*, and a **New project** button. The button creates an
+idle session (no form; the toast reads *New session ready. Say what it is for
+and the work begins.*). Then type the objective into the message box. The
+message goes to `POST /api/projects/{sid}/message`, the Manager classifies it,
+and if it is work the session's daemon is started on demand and the task is
+queued. The web UI follows your browser language (English or Simplified
+Chinese); the sidebar has a language button.
+
+### A real first task from the browser
+
+Project `s-67fb6d62` in the fresh-install trial on this machine was created
+from the web UI (`session.json` records `origin: web`). Its first message, at
+06:57:51Z:
+
+> 帮我在这台机器的 GPU 上测一下 PyTorch 矩阵乘在 fp16 和 fp32 下的 TFLOPS,写个脚本跑出真实数字,整理成表格。
+
+The reply, three minutes later at 07:02:21Z, reported an average of 121.31
+TFLOPS in FP16 and 23.96 in strict FP32 at 16384×16384 across the four GPUs,
+with a per-GPU table and the script name (`benchmark_torch_matmul.py`, seven
+timed batches, medians). A follow-up at 07:03:44Z, "现在项目里都有什么文件?把 bf16
+也补测一下,更新表格。", produced an updated table at 07:07:00Z (FP16 122.08,
+BF16 125.00, FP32 23.87 on average). The project's usage ledger shows $0.80 for
+the whole exchange, plus two early calls that were never priced (see below).
+
+That first message was actually sent twice. The first attempt (also 06:57:51Z)
+came back as `[not dispatched] Manager could not classify this message
+(refused before start: unresolved provider cost ...)`: on a fresh Copilot
+install the first call's price was not yet known, and the default policy
+refuses to spend money it cannot price. [Best practices](best-practices.md#controlling-spend)
+explains the setting (`ARGUS_SKILL_UNPRICED_COST_POLICY`) and the trade-off.
+
+### The endpoints behind the UI
+
+If you script against the server, these are the calls the UI itself makes
+(all need the bearer token when one is configured):
+
+| Endpoint | Body | Purpose |
+|---|---|---|
+| `POST /api/daemons` | `objective`, `name`, `workdir`, `launch_cwd` | create a session (an objective starts its worker immediately) |
+| `POST /api/projects/{sid}/message` | `text`, optional `route_override` (`auto`/`chat`/`task`), `attachments` | send a message through the Manager; `/message/stream` is the SSE twin |
+| `POST /api/projects/{sid}/tasks` | `text`, `autostart_daemon` (default true) | queue work directly |
+| `POST /api/projects/{sid}/nudge` | `text` | guidance for the next round (429 when the queue is full) |
+| `POST /api/projects/{sid}/backlog/{item_id}/answer` | `text` | answer a paused question and resume |
+| `POST /api/projects/{sid}/continuous` | `enabled`, `objective` | turn a campaign objective on or off |
+| `GET /api/projects/costs` | | spend per project |
+
+## Path 3: the desktop app
+
+The desktop app is a Tauri host around a frozen copy of the same Python
+backend; it opens the same web cockpit. It is a packaged preview channel,
+separate from source updates.
+
+- **Download:** the Releases page of
+ [lbx154/Argus](https://github.com/lbx154/Argus/releases). Windows ships as
+ an NSIS installer named `Argus--setup.exe`. The release workflow can
+ also build macOS (`.dmg`, arm64 and x86_64) and Linux (`.AppImage`, `.deb`)
+ packages, but the workflow file notes that v0.1.7 and v0.1.8 were published
+ for Windows only, so check the assets of the release you pick.
+ `microsoft/ArgusAgent` Releases is a separate channel; do not mix installers
+ or updates between the two.
+- **What it needs:** nothing else on the machine. The installer bundles the
+ backend (`argus-backend`), so no Python, Node.js or venv is required.
+- **First launch:** a setup wizard asks you to pick an installed, logged-in
+ agent CLI and confirm it; the local backend does not start until you do. If
+ you have an internal trial key instead, paste it and click **开始试用** ("start
+ trial"): the app downloads the official standalone Copilot program, checks
+ its SHA-256, runs one real model reply, and opens the workbench. That path
+ needs no CLI, Python, Node or coding-agent account of your own. Either
+ choice is remembered; **File → Settings** changes the CLI, executable or
+ port later.
+- **After that** the app is the web UI from Path 2: create a project, type the
+ objective, watch the feed.
+
+To build it from source instead (Python 3.11+, Node 22.12+, Rust stable, plus
+MSVC Build Tools on Windows or Xcode command-line tools on macOS), the
+documented sequence in [docs/desktop-trial.md](desktop-trial.md) is:
+
+```bash
+python -m pip install -e ".[trial]" "pyinstaller>=6.11,<7" tzdata
+npm --prefix frontend/web ci
+npm --prefix frontend/tui ci
+npm --prefix desktop-tauri ci
+npm --prefix frontend/web run build && npm --prefix frontend/tui run build
+npm --prefix desktop-tauri run build:backend
+npm --prefix desktop-tauri run build:unsigned
+```
+
+The unsigned bundle lands under `desktop-tauri/src-tauri/target/release/bundle/`.
+[docs/windows-desktop.md](windows-desktop.md) covers signing and the update
+channel.
+
+## What "finished" looks like
+
+Whichever path you used, the task ends the same way. The Reviewer returns
+`done` for the last task; with `--bounded` (or a bounded objective in the
+cockpit) the Planner then declares the project done and the worker stops, and
+`argus --status` shows `history : N done` with no running task. The evidence is
+in the work directory (the oscillator task's `README.md`; the roofline
+campaign's `paper/` and `results/`), and the per-call record of what it cost is
+in `~/.argus-skill/projects//usage.jsonl`.
+
+If instead the status shows a task `paused`, read the question with
+`argus --status` or in the cockpit and answer it with `argus --answer`; if it
+shows `paused_cost`, see [controlling spend](best-practices.md#controlling-spend).
+
+## Where things live
+
+| Path | Contents |
+|---|---|
+| `~/.argus-skill/` | the state root (`ARGUS_SKILL_HOME` or `--life-dir` overrides it) |
+| `~/.argus-skill/config.json` | persisted settings from `--setup`, `/config`, `/backend` |
+| `~/.argus-skill/projects//` | one project's state: `events.jsonl`, `usage.jsonl`, `backlog.jsonl`, `continuous.json`, `daemon.log` |
+| `~/.argus-skill/verticals/` | verticals installed from the store |
+| your work directory | where the roles read and write files; never inside the state root |
diff --git a/examples/verticals/argus_verticals/lab_notebook/__init__.py b/examples/verticals/argus_verticals/lab_notebook/__init__.py
new file mode 100644
index 000000000..f301d5db4
--- /dev/null
+++ b/examples/verticals/argus_verticals/lab_notebook/__init__.py
@@ -0,0 +1,5 @@
+"""``lab_notebook``: the worked example from docs/building-a-vertical.md.
+
+A two-stage vertical for small measurement tasks: measure something on this
+machine, then write the result up so another person can rerun it.
+"""
diff --git a/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md b/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md
new file mode 100644
index 000000000..a60c336b6
--- /dev/null
+++ b/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md
@@ -0,0 +1,20 @@
+---
+name: "Measurement Record"
+description: "How to run a small measurement so the result can be checked and rerun: save raw output, repeat, report the spread, and record the exact command."
+---
+
+# Measurement record
+
+Run the measurement from a script saved in the work directory, never from an
+interactive shell you cannot show later. Write the raw output to a file under
+`results/` before computing any aggregate.
+
+Repeat the run. Report the median (or mean) together with the spread (min/max or
+standard deviation) and the number of repeats. A single run is not a result.
+
+Record in `NOTEBOOK.md`:
+
+- what was measured and on which hardware
+- the exact command to rerun it
+- the aggregate, the spread and the number of repeats
+- the path of the raw output every number came from
diff --git a/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md b/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md
new file mode 100644
index 000000000..d9768ad3d
--- /dev/null
+++ b/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md
@@ -0,0 +1,16 @@
+---
+name: "Measurement Review"
+description: "How to review a measurement task: open the raw output files, recompute one aggregate, and confirm the rerun command is complete before accepting."
+---
+
+# Measurement review
+
+Do not accept the Engineer's summary as evidence. Open the raw output files
+named in `NOTEBOOK.md` and recompute at least one aggregate yourself.
+
+Confirm that the rerun command in `NOTEBOOK.md` names the script, its arguments
+and the working directory. A command that cannot be pasted and run is
+incomplete.
+
+Return `continue` when a number has no raw file behind it, and `done` only when
+every checklist item of the current stage has evidence you inspected.
diff --git a/examples/verticals/argus_verticals/lab_notebook/stages.py b/examples/verticals/argus_verticals/lab_notebook/stages.py
new file mode 100644
index 000000000..491334251
--- /dev/null
+++ b/examples/verticals/argus_verticals/lab_notebook/stages.py
@@ -0,0 +1,81 @@
+"""Stages and checklists of the ``lab_notebook`` example vertical.
+
+Everything Argus needs to know about a vertical is declared at module level in
+this file; there is no base class to inherit from. The framework reads the
+attributes below through ``argus.core.vertical_contract.vertical_contract`` and
+``argus.verticals._registry._validated_plugin``.
+"""
+
+from __future__ import annotations
+
+from argus.skills.stage_machine import ChecklistItem
+
+# Required for every vertical that is not built in. Argus refuses any other
+# API version, and an empty purpose hides the vertical from the Manager.
+ARGUS_VERTICAL_API_VERSION = 1
+VERTICAL_PURPOSE = (
+ "small measurement tasks on this machine: run the measurement, record how "
+ "it was run, and write a notebook entry another person can reproduce"
+)
+
+# The stage order is the tuple order. Every stage that is not optional needs a
+# non-empty checklist below.
+CHECKLIST_STAGE_ORDER: tuple[str, ...] = ("measure", "report")
+CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = ()
+# Names the Manager may use for a stage; they are canonicalized to the real one.
+STAGE_ALIASES = {"experiment": "measure", "writeup": "report"}
+
+# "none": the project is complete when the last stage's checklist is met.
+# "metric" and "certified" are the other two values the contract accepts.
+completion_gate = "none"
+WORKFLOW_MODE = "staged" # staged | direct | proportional
+MISSION_KIND = "custom" # custom | optimize | research | software
+REQUIRE_INDEPENDENT_REVIEW = True
+
+CHECKLIST_ITEMS: dict[str, tuple[ChecklistItem, ...]] = {
+ "measure": (
+ ChecklistItem(
+ id="measure.ran-here",
+ statement=(
+ "The measurement was executed on this machine in this project, "
+ "and its raw output is saved in the work directory."
+ ),
+ evidence_hint="the command that was run and the path of its raw output",
+ ),
+ ChecklistItem(
+ id="measure.repeated",
+ statement=(
+ "The measurement was repeated, and the reported number is a "
+ "median or mean with its spread; a single run is not a result."
+ ),
+ evidence_hint="number of repeats, the aggregate and the spread",
+ ),
+ ),
+ "report": (
+ ChecklistItem(
+ id="report.notebook-entry",
+ statement=(
+ "NOTEBOOK.md in the work directory states what was measured, "
+ "how, the result with its spread, and the exact command to rerun it."
+ ),
+ evidence_hint="the NOTEBOOK.md section and the rerun command",
+ ),
+ ChecklistItem(
+ id="report.numbers-traceable",
+ statement=(
+ "Every number in NOTEBOOK.md can be traced to a saved raw output file."
+ ),
+ evidence_hint="raw file paths next to each number",
+ ),
+ ),
+}
+
+
+def role_banner(role: str) -> str:
+ """One paragraph every role reads before its task; keep it about the field."""
+ return (
+ "LAB NOTEBOOK VERTICAL: measure first, then write. A number without a "
+ "saved raw output and a rerun command is not a result. The Reviewer "
+ "checks the notebook entry against the raw files, not against the "
+ "Engineer's summary."
+ )
diff --git a/examples/verticals/build_local_catalog.py b/examples/verticals/build_local_catalog.py
new file mode 100644
index 000000000..f8dfee9a4
--- /dev/null
+++ b/examples/verticals/build_local_catalog.py
@@ -0,0 +1,97 @@
+"""Package one vertical directory as a Vertical Store archive plus a local catalog.
+
+Usage:
+
+ python examples/verticals/build_local_catalog.py NAME --version 0.1.0 --out DIR
+
+Reads ``examples/verticals/argus_verticals/NAME`` (or ``--source``), writes
+``DIR/NAME-VERSION.zip`` with members under ``argus_verticals/NAME/`` (the layout
+``argus.verticals.store._verify_tree`` expects), and writes ``DIR/catalog.json``
+in the shape ``argus.verticals.store._validate_entry`` accepts. Point
+``ARGUS_VERTICAL_CATALOG`` at that file and ``argus verticals install NAME`` uses
+the archive next to it instead of downloading anything.
+
+The catalog is minimal on purpose: the fields the store requires, nothing else.
+See docs/building-a-vertical.md for the walk-through.
+"""
+
+from __future__ import annotations
+
+import argparse
+import hashlib
+import json
+import re
+import sys
+import zipfile
+from pathlib import Path
+
+_NAME = re.compile(r"^[a-z][a-z0-9_]{0,47}$")
+
+
+def _purpose(stages: Path) -> str:
+ """Read VERTICAL_PURPOSE from stages.py without importing it."""
+ text = stages.read_text(encoding="utf-8")
+ match = re.search(r"VERTICAL_PURPOSE\s*=\s*\(?\s*((?:\"[^\"]*\"\s*)+)\)?", text)
+ if match is None:
+ raise SystemExit(f"{stages}: VERTICAL_PURPOSE not found")
+ return " ".join("".join(re.findall(r"\"([^\"]*)\"", match.group(1))).split())
+
+
+def main(argv: list[str] | None = None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
+ parser.add_argument("name", help="vertical name, e.g. lab_notebook")
+ parser.add_argument("--version", default="0.1.0")
+ parser.add_argument("--source", type=Path, default=None,
+ help="vertical directory (default: examples/verticals/argus_verticals/NAME)")
+ parser.add_argument("--out", type=Path, required=True, help="directory for the zip and catalog.json")
+ args = parser.parse_args(argv)
+
+ name = args.name
+ if not _NAME.fullmatch(name):
+ raise SystemExit(f"{name!r} is not a valid vertical name (^[a-z][a-z0-9_]{{0,47}}$)")
+ source = args.source or (Path(__file__).resolve().parent / "argus_verticals" / name)
+ stages = source / "stages.py"
+ if not stages.is_file():
+ raise SystemExit(f"{source} has no stages.py")
+
+ out = args.out.resolve()
+ out.mkdir(parents=True, exist_ok=True)
+ archive_name = f"{name}-{args.version}.zip"
+ archive = out / archive_name
+ tree = f"argus_verticals/{name}"
+ with zipfile.ZipFile(archive, "w", compression=zipfile.ZIP_DEFLATED) as zf:
+ for path in sorted(p for p in source.rglob("*") if p.is_file()):
+ if "__pycache__" in path.parts:
+ continue
+ zf.write(path, f"{tree}/{path.relative_to(source).as_posix()}")
+
+ data = archive.read_bytes()
+ catalog = {
+ "schema": 1,
+ "verticals": {
+ name: {
+ "name": name,
+ "version": args.version,
+ "module": f"argus_verticals.{name}.stages",
+ "paths": [tree],
+ "purpose": _purpose(stages),
+ "has_skills": (source / "skills").is_dir(),
+ "archive": {
+ "file": archive_name,
+ "url": archive.as_uri(),
+ "sha256": hashlib.sha256(data).hexdigest(),
+ "size": len(data),
+ },
+ }
+ },
+ }
+ catalog_path = out / "catalog.json"
+ catalog_path.write_text(json.dumps(catalog, indent=2, ensure_ascii=False) + "\n", encoding="utf-8")
+ print(f"archive : {archive} ({len(data)} bytes)")
+ print(f"catalog : {catalog_path}")
+ print(f"install : ARGUS_VERTICAL_CATALOG={catalog_path} argus verticals install {name}")
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())