diff --git a/README.md b/README.md index 37e397784..e993d7284 100644 --- a/README.md +++ b/README.md @@ -107,6 +107,12 @@ maintainers for the latest code.

Community Group 1 is full. Please join Group 2.

+## Tutorials + +- **[Getting started](docs/getting-started.md)** — install, first task, three ways in: command line, web UI, desktop app. +- **[Best practices](docs/best-practices.md)** — writing an objective, changing direction mid-run, choosing the model and backend, controlling spend, what suits Argus. +- **[Building a vertical](docs/building-a-vertical.md)** — a small real vertical from stages and skills to the Vertical Store. + ## Quick Install Choose the section for your operating system. Do not mix commands between diff --git a/docs/LAYOUT.md b/docs/LAYOUT.md index d0ec9aa14..fb524d931 100644 --- a/docs/LAYOUT.md +++ b/docs/LAYOUT.md @@ -128,6 +128,7 @@ shipped" means no CI job, no test, and no wheel content comes from the directory - `deploy/` - systemd units and Dockerfiles for the hosted trial (`deploy/trial/`). - `desktop-tauri/` - the Tauri desktop shell and the PyInstaller spec (`argus_backend.spec`) for the frozen `argus-backend` binary (the spec's `name=`). - `docs/` - operator and developer documentation; `docs/audits/` holds dated audit reports and their data attachments. +- `examples/` - worked examples the guides refer to: `verticals/` holds the small `lab_notebook` vertical built in `docs/building-a-vertical.md` and the helper that packages a vertical into a local store catalog. Not built, not tested, not shipped. - `experiments/` - historical PR regression-study forwarding entry points and documentation (`pr_regression_50/`). The maintained implementation and offline tests live in `argus/release_tools/pr_gate/regression/` and `tests/tools/`; model/Docker studies are opt-in, not CI runs. This directory is not built or shipped in the wheel or sdist. - `frontend/` - `core` (shared TypeScript), `tui` (Ink terminal cockpit), `web` (React web cockpit). `frontend/web/dist` is committed on purpose and force-included into the wheel. - `integrations/` - the `agent-skills` package for external agent hosts (`SKILL.md` plus per-host adapters). Not the Python package `argus/integrations/`. diff --git a/docs/best-practices.md b/docs/best-practices.md new file mode 100644 index 000000000..ce1ee008f --- /dev/null +++ b/docs/best-practices.md @@ -0,0 +1,328 @@ +# Argus best practices + +This page explains how to get good work out of Argus and why each +recommendation holds. The reasons come from the code paths that read your +input; the examples come from three real projects run on one machine on +2026-09-30 (a GPU roofline research campaign, a damped-oscillator task, and a +fresh-install trial), with times and costs taken from their logs. Nothing here +is a quota; where a number appears it is a measurement, not a rule. + +Read [getting started](getting-started.md) first if you have not run a task +yet. + +## Writing an objective + +The objective is read by four different consumers, and writing for all of +them is what makes an objective good: + +1. **The Manager** decides from your text alone whether this is a + conversation, a small local job it can do itself, or team work for the + Planner, Engineer and Reviewer. It also decides the lifetime: finite, + casually worded work defaults to *bounded* (stop when done); only text that + expresses ongoing intent becomes a *standing* campaign that keeps generating + work. +2. **The Manager again** picks the vertical (research, software, math, ...) + from the objective; there is no keyword classifier and no flag to force it, + so name the kind of work in plain words. +3. **The Planner** decomposes it into tasks, each with an acceptance check. + What you leave unsaid, it will decide for you. +4. **The Reviewer** judges every round against it and against the vertical's + stage checklist. The Reviewer runs read-only and cannot ask the Engineer; + it can only check what the objective and the work directory let it check. + +So an objective should say what must be true at the end, what evidence you +will accept, what the scope and constraints are, and what form the result +takes. Compare the two objectives that ran on this machine. + +The campaign objective (`objective-roofline.txt`, project `s-02d3c282`): + +> 做一项小规模但真实的研究:在本机 4 块 NVIDIA RTX A6000 上,PyTorch 方阵乘法的实际吞吐(TFLOPS)随矩阵规模(512 到 16384)和精度(FP32、TF32、BF16、FP16)如何变化,与理论峰值的差距能否用一个简单的 roofline 式模型(计算强度与显存带宽)解释。要求:真实实验、每种配置多次重复取中位数并报告离散度,拟合模型并给出误差,画图,最后写成一篇 4 页以内的短论文(含方法、结果、讨论、复现说明),并通过内部评审。所有数字必须来自本机真实运行,不允许估算或编造。 + +Every clause did work. "小规模但真实的研究" and "短论文" put it in the `research` +vertical (`pipeline : vertical=research` in the status). The hardware and the +ranges fixed the scope, so the Planner did not have to guess a sweep. "多次重复取中位数并报告离散度" +and "拟合模型并给出误差" became things the Reviewer could check: the draft's +`results.tex` reports 17,280 samples, medians, and the fit's error on a held-out +confirmation set. "通过内部评审" is why a `paper` round came back +`replan_requested` instead of being accepted on the Engineer's word. And +"所有数字必须来自本机真实运行,不允许估算或编造" is the sentence that lets the Reviewer +refuse a plausible-looking number that has no raw file behind it. + +The small objective (`objective-oscillator.txt`, project `s-4a34a674`): + +> 用 numpy 数值求解一维阻尼谐振子 x'' + 2γx' + ω0² x = 0(取 ω0=2π, 分别取欠阻尼 γ=0.5、临界 γ=2π、过阻尼 γ=10),与解析解逐点比较给出最大误差,画出三条曲线,把方法、数值结果表和图写进本目录的 README.md。所有数字必须来自真实运行。 + +It names the equation, the three parameter values, the comparison ("与解析解逐点比较给出最大误差"), +the deliverable and its location. Because the acceptance check was in the +text, the task ran once, was sent back once, and was done after 12 minutes and +$0.91. + +The browser message from the trial (project `s-67fb6d62`) shows the casual end +of the scale: + +> 帮我在这台机器的 GPU 上测一下 PyTorch 矩阵乘在 fp16 和 fp32 下的 TFLOPS,写个脚本跑出真实数字,整理成表格。 + +"写个脚本跑出真实数字" and "整理成表格" were enough for a bounded, self-contained job +that replied with a per-GPU table three minutes later. What it did not say +(matrix sizes, repeats) the Engineer chose: 4096/8192/16384 and seven timed +batches. If those choices matter to you, say them. + +Things to avoid in an objective, and why: + +- **Work that needs your credentials, your money, or an irreversible or + public action.** Those are authority boundaries: whatever the autonomy + setting, a task that reaches one stops and asks. Put the credential in + place first (the cockpit accepts credentials and stores them in the + capability vault) or keep the objective short of the boundary. +- **A goal with no checkable end.** "Make it better" leaves the Reviewer + nothing to refuse, so rounds continue until a round cap or you intervene. +- **Numbers you already know the answer to.** The roles will try to + reproduce them; that is the point, and it costs rounds. + +On the command line the objective must travel with `--continuous` +(`--objective` alone is refused), and `--objective-file` keeps long text out of +shell quoting. Add `--bounded` when the objective is finite; without it the +worker keeps proposing follow-up work after the project is declared done. + +## Changing direction while it runs + +Argus gives you three ways to speak to a running project. They are not +interchangeable, because they enter at different points of the loop. + +| You want to | Command line | Cockpit | Web API | What happens | +|---|---|---|---|---| +| Add guidance the next round should read | `argus --notify "text"` | `/nudge text` | `POST /api/projects/{sid}/nudge` | the text is queued in the project's durable inbox and spliced into the next Engineer round's prompt as operator guidance | +| Hold guidance until a later stage | `argus --notify "text" --notify-stage paper` | | | same, delivered only when the project reaches that stage (aliases such as `writeup` are canonicalized; unknown stages are refused) | +| Change the direction of the current task, pause or abort | (type it) | type it, or `/abort` | `POST /api/projects/{sid}/message` | the Manager classifies the message: `STEER` records a new direction for the active task, `PAUSE` stops the campaign, `ABORT` ends the task, `NO_DISPATCH` stops new work | +| Set a rule for the whole project | (type it) | type it | `POST .../message` | text that sounds standing ("always", "never", "from now on", "do not ask") is stored as a project-wide directive; other amendments are scoped to the current task | +| Answer a question the task stopped on | `argus --answer "text"` (`--answer-item ID` if several wait) | answer in the pending-question prompt | `POST /api/projects/{sid}/backlog/{item_id}/answer` | the paused task is replaced by a continuation whose objective carries your answer as authority, and the worker is restarted if needed | +| Ask something without touching the run | `argus --ask "question"` | `/ask question` | `route_override: "chat"` on `/message` | the Manager answers from project state; nothing is queued | + +Why the nudge is deliberately weak: it is guidance the next round *reads*, not +an order the runtime *enforces*. The inbox is durable (a queue under the +project state that survives restarts, acknowledges only after delivery, and +answers HTTP 429 when full), but the Engineer decides what to do with the +text. That is the right tool for "prefer torch.compile off" or "the paper must +cite X"; it is the wrong tool for "stop". + +Why `--answer` is not a nudge: a task that stopped to ask a question is +`paused` until the question is answered. Clearing the question and resuming +the same task looks like it works but does not: the task re-reads the +objective that made it ask and asks again (the code comment on `--answer` +records a campaign that burned five attempts that way). The answer path +therefore enqueues a continuation whose objective states your answer, so the +next round reads what it was told. + +Why a typed message is classified rather than obeyed literally: the same box +carries greetings, questions, new tasks, settings changes and controls. The +Manager first decides whether the message is conversation or work, and only +when a task is active and the text clearly redirects it does the second check +turn it into a `STEER`. Questions, criticism and suggestions are not steering; +if you mean "change course", say so in the imperative. + +When Argus stops on its own to ask you is governed by one setting: + +```bash +export ARGUS_SKILL_AUTONOMY_MODE=pragmatic # default +``` + +`pragmatic` recovers technical problems (a failed test, a timeout, a route +that did not work) without asking and stops only at authority boundaries: +credentials, payment or a bigger budget, deleting or force-pushing shared +state, publishing or sending outward, and changes to an acceptance contract +you own. `cautious` asks on every explicit question the Reviewer raises. +`autonomous` still stops at those same boundaries; it only removes the +remaining discretionary pauses. Pick `cautious` for the first run of a new +vertical, when you want to see what it would have asked. + +Two real cases show what interjection looks like in practice. + +*The follow-up.* In the trial's browser project, after the first table +arrived, the second message at 07:03:44Z was + +> 现在项目里都有什么文件?把 bf16 也补测一下,更新表格。 + +Three minutes later the reply listed the project's files and gave the updated +table with a BF16 column. That message is a new bounded task in the same +project, not a steer of a running one; the earlier task had already finished. +Most "changes of direction" on small tasks are like this, and the message box +is the right place for them. + +*The pause that was not a question.* The roofline campaign's log shows, 43 +seconds after launch, a round ending with "The work was paused before the +Engineer finished this round because provider charges are awaiting +reconciliation or explicit risk approval, so this round was not judged". +Nothing was asked of the operator; the cost check had refused an unpriced +call. The operator's response was two CLI commands recorded in +`daemon.commands.jsonl`: a `drain` at 08:16:12Z and a `start` at 08:21:19Z +after changing the cost policy. Over the following two hours the campaign ran +without a single operator message: all twenty entries in its inbox were +background-job reports the runtime queued for itself. A precise objective is +what made that possible. + +## Choosing the backend and model + +Argus resolves the backend and model for each role separately, and every +resolution follows the same order: + +1. an environment variable for that role (`ARGUS_SKILL_ENGINEER_BACKEND`, + `ARGUS_SKILL_ENGINEER_MODEL`, ...), +2. the shared variable (`ARGUS_SKILL_RUNNER_BACKEND`, `ARGUS_SKILL_MODEL`), +3. the same names in the persisted settings file `~/.argus-skill/config.json`, + which is what `argus --setup`, `/backend`, `/config key=value` and a + natural-language "把模型换成 X" write, +4. the default (`codex` for the backend; `auto` for the model, meaning the + backend's own default). + +A `--backend` typed on a `--daemon` launch sits above all of these for that +daemon and is exported to its roles at boot, so a persisted choice can never +silently override what you typed. `argus --config-help` prints every setting +with its default, its current value and where the value came from +(`env`, `persisted`, `default`); `argus --config-snapshot` writes the same to +a file you can attach to a report. + +| Setting | Default | Notes | +|---|---|---| +| `ARGUS_SKILL_RUNNER_BACKEND` | `codex` | `codex`, `claude`, `copilot`, `cursor`, `opencode`, `pi`, `grok`, `qoder`, `dsh` | +| `ARGUS_SKILL__BACKEND` | inherits | `ENGINEER`, `REVIEWER`, `PLANNER`, `MANAGER`, `SUPERVISOR` | +| `ARGUS_SKILL_MODEL` | `auto` | a bare model id (`gpt-5.6-sol`, `copilot/opus-5`); free text reaches the CLI verbatim and every call then fails | +| `ARGUS_SKILL__MODEL` | `auto` | `ENGINEER`, `REVIEWER`, `PLAN`, `MANAGER`, `SUPERVISOR` | +| `ARGUS_SKILL_FRONTDOOR_MODEL` | `auto` | the cheap classifier that reads every message (`gpt-5.4-mini` on Copilot) | +| `ARGUS_SKILL__REASONING_EFFORT` | Engineer `xhigh`, others `high`, Supervisor `low` | `low`, `medium`, `high`, `xhigh`, `max` | +| `ARGUS_SKILL_PI_PROVIDER` / `ARGUS_SKILL_OPENCODE_PROVIDER` | unset | provider prefix for bare model ids on those two backends; on OpenCode the model is dropped without it | + +Why the roles are separable: they do different jobs. The Engineer runs long +tool-using turns and benefits from the strongest model at high effort; the +Reviewer reads and judges; the front door classifies one short message per +turn and should be cheap. The roofline campaign ran on Copilot with +`ARGUS_SKILL_MODEL=gpt-5.6-sol`, and its `usage.jsonl` splits as: $29.22 on +`gpt-5.6-sol` across engineer, planner, reviewer, manager and supervision +calls; $0.16 on `gpt-5.4-mini` for classification, reflection and fact +extraction. Changing the front-door model would have saved nothing; changing +the Engineer's would have changed everything. + +The same ledger explains where the money goes: 33.9 million input tokens +against 0.3 million output tokens after two hours. Long tool-using rounds +re-read context, which is why Argus by default sends the full task text only +on the first round of a session (`ARGUS_SKILL_COMPACT_CONTINUATION_PROMPTS`) +and why a precise objective with the evidence already in the work directory +costs less than a vague one that makes the Engineer explore. + +## Controlling spend + +Spend is controlled at three levels, and the one that surprises new users is +the third. + +**A daily cap across every project on the host.** +`ARGUS_SKILL_GLOBAL_DAILY_CAP_USD` (default `1000.0`) and, if you prefer to +count tokens, `ARGUS_SKILL_GLOBAL_DAILY_TOKEN_CAP` (default `0`, off). +`argus --status` shows the running total: `budget : global daily $1000.00 +(spent $30.71) · remaining $969.29` was the line two hours into the roofline +campaign, with `cost : $29.79 cumulative` for that project alone. The +per-call record is `~/.argus-skill/projects//usage.jsonl` (model, tokens, +`cost_usd`), and the web UI reads `GET /api/projects/costs`. + +**Provider-level circuit breakers.** `ARGUS_SKILL_CODEX_DAILY_CALL_CAP` +(default 300 calls per day), `ARGUS_SKILL_COPILOT_DAILY_CALL_CAP`, +`ARGUS_SKILL_COPILOT_DAILY_PREMIUM_CAP` and `ARGUS_SKILL_COPILOT_HOURLY_CALL_CAP` +(default 10000 each), `ARGUS_SKILL_PROVIDER_MAX_CONCURRENCY` (default 0, off) +and `ARGUS_SKILL_MAX_ACTIVE_DAEMONS` (default 64). These exist because a +subscription CLI is shared by every project on the host, and one runaway +campaign should not exhaust it for the others. `--mission-width` (default 2) +is the per-project counterpart; the roofline campaign used `1`, which is the +right choice when the tasks share four GPUs. + +**What to do with a call whose price is not known yet.** +`ARGUS_SKILL_UNPRICED_COST_POLICY` is `block` by default: a call whose cost the +provider has not settled is refused before it starts, because a cap that +cannot see the cost cannot enforce anything. With Copilot's subscription +billing the first calls on a fresh install report `pricing_status: partial` +with no `cost_usd`, and that is exactly what happened in the trial: + +- CLI project `s-600e27bf`: the first classify call settled unpriced (its + event has `"pricing_status":"partial","cost_usd":null`), the worker went + to `paused_cost`, and `cost-control.json` listed the call as unresolved with + the reason "Copilot token billing is awaiting local CLI usage + reconciliation". +- Browser project `s-67fb6d62`: the first message came back `[not dispatched] + Manager could not classify this message (refused before start: unresolved + provider cost: 1 call(s) awaiting usage reconciliation ...)`. +- The roofline campaign, on a second install a little later, paused 43 + seconds in for the same reason. + +The operator's fix each time was one line in `~/.argus-skill/config.json`, +`"ARGUS_SKILL_UNPRICED_COST_POLICY": "allow"`, followed by a drain and restart +(`/config ARGUS_SKILL_UNPRICED_COST_POLICY=allow` in the cockpit writes the +same key, and the web configuration view exposes it). `allow` means: run the +call now and let the ledger reconcile later. The trade-off is real. With a +metered API key, keep `block`; the price of every call is known and the cap +is exact. With a subscription CLI, `allow` is usually right, because the +"cost" the ledger reconciles is your subscription's usage, and refusing to run +does not save money you have already paid. The daily cap still applies to +everything that does get priced. + +Two changes merged on 2026-09-30 (#179 and #183) move this default: Copilot +calls made through the warm session are now priced from the CLI's own session +log, and a call Argus itself interrupts before the CLI recorded anything is +settled as one premium request. A fresh install on a subscription CLI no +longer pauses under `block`; the three examples above ran before that change. +`allow` is still the choice when you would rather run than wait for any late +reconciliation at all. + +A last practical point: `ARGUS_SKILL_COST_CONTROL` (default `on`) is the +switch for the whole admission-and-reconciliation layer. Leave it on; turning +it off also removes the `paused_cost` state that tells you something is wrong +with billing. + +## Which tasks suit Argus + +Argus is built for work that can be checked. The four-role loop only pays for +itself when there is evidence for the Reviewer to read and refuse. + +It fits well when: + +- **The result is measurable on the machine it runs on.** Every task in the + three example projects was of this kind: a TFLOPS sweep, a numerical + solution compared point-by-point with an analytic one, a script whose + output table can be rerun. The Reviewer can open the raw files. +- **The work has stages and takes hours.** The roofline campaign moved from + idea to experiment to a paper draft, launched background GPU jobs and + waited for them, and had a draft sent back for rework, all without an + operator message. That is what the daemon, the stage checklists and the + independent review exist for. +- **There is a vertical for the field.** Seven ship with Argus (`research`, + `software`, `math`, ...) and the store carries seventeen more. A vertical is + what turns "done" into a checklist the Reviewer can apply; without one the + Manager still routes the task, but the acceptance standard is generic. See + [building a vertical](building-a-vertical.md). +- **You want it to keep going.** A standing objective with `--continuous` and + no `--bounded` keeps proposing follow-up work in the same direction. + +It fits poorly when: + +- **You want an answer, not work.** `argus --ask` and `/ask` answer inline for + a fraction of a cent and queue nothing. Sending a question through the task + path costs a Manager turn plus, if it is misread as work, a Planner and + Engineer round. +- **The task is smaller than the loop.** The oscillator task, a script most + people would write in ten minutes, took 12 minutes and $0.91 because it went + through Manager, Planner, an Engineer round, a Reviewer `continue`, a second + round and a `done`. That overhead is worth it when you would otherwise not + check the work; it is not worth it for a throwaway. +- **Finishing requires something only you can do.** Logging in somewhere, + paying, publishing, deleting shared state, or approving a change to an + acceptance contract stops the task by design in every autonomy mode. Plan + the objective to end before that step, or do the step first. +- **"Done" cannot be written down.** If you cannot state what the Reviewer + should check, it will accept or refuse on its own reading, and rounds will + continue until a cap. +- **The evidence lives somewhere the roles cannot reach.** The Reviewer is + read-only and works from the work directory and the paths you allow + (`ARGUS_SKILL_REVIEWER_READ_DIRS`); a result that only exists in a service it + cannot query cannot be certified. + +A useful test before you start: write the sentence the Reviewer should be +able to say at the end ("the README reports the maximum error against the +analytic solution for all three damping cases, and each number matches the +saved run"). If you can write it, put it in the objective. If you cannot, +Argus will not be able to tell when it is finished either. diff --git a/docs/building-a-vertical.md b/docs/building-a-vertical.md new file mode 100644 index 000000000..bd794d9b0 --- /dev/null +++ b/docs/building-a-vertical.md @@ -0,0 +1,431 @@ +# Building a vertical + +A vertical teaches Argus what "done" means in one field: the stages work moves +through, the checklist the Reviewer applies at each stage, and the skills the +roles read before they start. This page builds a small real one, +`lab_notebook`, from nothing to an installed entry in the Vertical Store. The +finished files are under [`examples/verticals/`](../examples/verticals/) and +every command below was run against this checkout. + +Background reading: [the Vertical Store](vertical-store.md) for the store's +internals and hosted mode, and +[single-agent verticals](single-agent-vertical-20260915.md) for the design +discussion behind the contract. + +## What a vertical is, in code + +A vertical is a Python package whose `stages.py` declares a handful of +module-level names. There is no base class and no registration call: the +framework imports the module and validates the names through +`argus.core.vertical_contract.vertical_contract`, and (for anything not built +in) `argus.verticals._registry._validated_plugin`. The built-in seven are +listed by name in `argus/skills/vertical_select.py`; everything else arrives +through the store or a Python entry point in the group `argus.verticals`. + +The names the contract reads: + +| Name | Required | Meaning | +|---|---|---| +| `ARGUS_VERTICAL_API_VERSION` | yes, for non-built-ins | must be `1`; any other value hides the vertical | +| `VERTICAL_PURPOSE` | yes, for non-built-ins | one sentence; it is the line the Manager sees when choosing a vertical, so write it for routing, not as an abstract | +| `CHECKLIST_STAGE_ORDER` | yes | tuple of stage names, in order; no duplicates | +| `CHECKLIST_ITEMS` | yes | dict `stage -> tuple[ChecklistItem, ...]`; every non-optional stage needs a non-empty checklist | +| `completion_gate` | yes | `"none"`, `"metric"` or `"certified"` | +| `CHECKLIST_OPTIONAL_STAGES` | no | stages that may have no checklist | +| `STAGE_ALIASES` | no | `{"experiment": "measure"}`: other names the Manager may use for a stage | +| `WORKFLOW_MODE` | no | `"staged"`, `"direct"` or `"proportional"` | +| `MISSION_KIND` | no | `"custom"`, `"optimize"`, `"research"` or `"software"` | +| `REQUIRE_INDEPENDENT_REVIEW` | no | defaults to `True` | +| `role_banner(role)` | no | a paragraph each role reads before its task | +| `render_role_prompt_fragment`, `render_role_prompt_context`, `stage_completion_issues`, `prepare_mission`, `LIBRARY_PREPARER`, `EVIDENCE_SCHEMA`, `WORKFLOW_PROFILES` | no | hooks the larger built-ins use; not needed for a first vertical | +| `VERTICAL_SKILLS`, `VERTICAL_SKILL_PARENTS`, `VERTICAL_ROUTING_PATH` | no | an explicit skills root (a `skills/` directory next to `stages.py` is found automatically), verticals whose skills are seeded before yours, and a browsing category for the store | + +A `ChecklistItem` (from `argus/skills/stage_machine.py`) has three fields: + +```python +@dataclass(frozen=True) +class ChecklistItem: + id: str + statement: str + evidence_hint: str +``` + +The `statement` is what must be true; the `evidence_hint` tells the Engineer +what to show and the Reviewer what to look for. This is the whole mechanism: +the Reviewer is read-only, so the quality of your checklist is the quality of +your review. + +## Two built-ins to learn from + +**`software`** (`argus/verticals/software/`) is the smallest useful vertical: +one stage, `delivery`, with three checklist items, a `role_banner`, and four +skill files. Its `stages.py` is worth reading in full; the module docstring +explains why the checklist exists ("code that does not build" was the failure +it was written against), and the first two item ids are marked protected so a +Planner may add items but never weaken them: + +```python +STAGE_ORDER = ["delivery"] +CHECKLIST_STAGE_ORDER = tuple(STAGE_ORDER) +CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = () +completion_gate = "none" +MISSION_KIND = "software" +WORKFLOW_MODE = "staged" + +PROTECTED_ITEM_IDS = frozenset({"delivery.builds", "delivery.tests-executed"}) +``` + +Its skills sit at `skills//.md`: + +``` +software/skills/engineer/software-change-implementation.md +software/skills/manager/software-project-grounding.md +software/skills/planner/software-project-grounding.md +software/skills/reviewer/software-change-review.md +``` + +**`research`** (`argus/verticals/research/`) is the largest: four stages, + +```python +CANONICAL_STAGE_ORDER: tuple[str, ...] = ("idea", "experiment", "paper", "review") +STAGE_ALIASES = {"research": "idea", "plan": "experiment", ..., "submission": "review"} +``` + +with `completion_gate = "certified"`, `WORKFLOW_MODE = "proportional"`, role +prompts in `prompt_policy.py`, a `prepare_mission` hook in `mission_brief.py`, +and some forty engineer skills plus seven reviewer skills. It shows where a +vertical can grow; it is not where one should start. + +The directory shape both share, and the store expects: + +``` +/ + __init__.py docstring only + stages.py the contract + skills/ + engineer/*.md + reviewer/*.md + manager/*.md (optional) + planner/*.md (optional) + engineer/references/ supporting files, not skills (optional) + engineer/*_scripts/ .py/.json/.sh copied verbatim (optional) +``` + +## Step 1: decide the stages and the checklist + +`lab_notebook` is for small measurement tasks: run something on this machine, +then write it up so someone else can rerun it. Two stages follow from that, +`measure` and `report`, and each gets the two or three things a reviewer +would actually check. + +Write the checklist before anything else, and write it as statements a +read-only reviewer can verify from files. "The measurement was repeated" is +checkable (there are N raw files); "the measurement is good" is not. + +`examples/verticals/argus_verticals/lab_notebook/stages.py`: + +```python +from argus.skills.stage_machine import ChecklistItem + +ARGUS_VERTICAL_API_VERSION = 1 +VERTICAL_PURPOSE = ( + "small measurement tasks on this machine: run the measurement, record how " + "it was run, and write a notebook entry another person can reproduce" +) + +CHECKLIST_STAGE_ORDER: tuple[str, ...] = ("measure", "report") +CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = () +STAGE_ALIASES = {"experiment": "measure", "writeup": "report"} + +completion_gate = "none" +WORKFLOW_MODE = "staged" +MISSION_KIND = "custom" +REQUIRE_INDEPENDENT_REVIEW = True + +CHECKLIST_ITEMS: dict[str, tuple[ChecklistItem, ...]] = { + "measure": ( + ChecklistItem( + id="measure.ran-here", + statement=("The measurement was executed on this machine in this project, " + "and its raw output is saved in the work directory."), + evidence_hint="the command that was run and the path of its raw output", + ), + ChecklistItem( + id="measure.repeated", + statement=("The measurement was repeated, and the reported number is a " + "median or mean with its spread; a single run is not a result."), + evidence_hint="number of repeats, the aggregate and the spread", + ), + ), + "report": ( + ChecklistItem( + id="report.notebook-entry", + statement=("NOTEBOOK.md in the work directory states what was measured, " + "how, the result with its spread, and the exact command to rerun it."), + evidence_hint="the NOTEBOOK.md section and the rerun command", + ), + ChecklistItem( + id="report.numbers-traceable", + statement="Every number in NOTEBOOK.md can be traced to a saved raw output file.", + evidence_hint="raw file paths next to each number", + ), + ), +} + + +def role_banner(role: str) -> str: + return ( + "LAB NOTEBOOK VERTICAL: measure first, then write. A number without a " + "saved raw output and a rerun command is not a result. The Reviewer " + "checks the notebook entry against the raw files, not against the " + "Engineer's summary." + ) +``` + +Why these particular choices: + +- `completion_gate = "none"` means the project is complete once the last + stage's checklist is met. `"metric"` is for verticals whose completion is a + number reaching a target; `"certified"` adds an explicit certification step, + which the research vertical uses for papers. A notebook entry needs neither. +- `WORKFLOW_MODE = "staged"` makes the Manager move through the stages in + order. `"direct"` skips staging for one-shot work; `"proportional"` lets the + Manager scale the process to the task. Start with `staged`; the aliases let + a Manager that says "experiment" land on `measure`. +- `MISSION_KIND = "custom"` because the other three values switch on + behaviour written for optimization loops, research campaigns and repository + changes. +- The `role_banner` says one thing, in the field's own terms. It is read by + every role, so it is the place for the rule that ties the stages together. + +The contract check refuses, with a message naming the problem, a stage with no +checklist, a checklist for a stage that is not in the order, a duplicate stage, +an item without an `id` or `statement`, a repeated item id within a stage, and +an unknown `completion_gate` value. Run it directly while you iterate: + +```bash +python -c " +import importlib.util, pathlib +from argus.core.vertical_contract import vertical_contract +p = pathlib.Path('examples/verticals/argus_verticals/lab_notebook/stages.py') +spec = importlib.util.spec_from_file_location('lab_notebook_stages', p) +m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m) +c = vertical_contract('lab_notebook', m) +print(c.stage_order, c.completion_gate, c.workflow_mode) +" +``` + +## Step 2: write the skills + +A skill is a Markdown file with exactly two front-matter fields, `name` and +`description`, both quoted. The runtime does not parse skill bodies: the +roles receive the library paths and read the files themselves, so the +description is a routing line (it is what an agent reads to decide whether to +open the file; the repository's own test caps it at 1200 characters) and the +body is written for a reader who will act on it. + +`skills/engineer/measurement-record.md`: + +```markdown +--- +name: "Measurement Record" +description: "How to run a small measurement so the result can be checked and rerun: save raw output, repeat, report the spread, and record the exact command." +--- + +# Measurement record + +Run the measurement from a script saved in the work directory, never from an +interactive shell you cannot show later. Write the raw output to a file under +`results/` before computing any aggregate. + +Repeat the run. Report the median (or mean) together with the spread (min/max or +standard deviation) and the number of repeats. A single run is not a result. + +Record in `NOTEBOOK.md`: + +- what was measured and on which hardware +- the exact command to rerun it +- the aggregate, the spread and the number of repeats +- the path of the raw output every number came from +``` + +`skills/reviewer/measurement-review.md` tells the Reviewer to open the raw +files and recompute one aggregate before accepting, and to return `continue` +for any number without a file behind it. The two skills and the two checklists +say the same thing from three sides; that redundancy is deliberate, because +each role reads only its own. + +Rules the loader applies: skills live at `skills//` where the role is +`engineer`, `reviewer`, `planner` or `manager` (anything else is filed as +general); a `references/` directory is copied as supporting material rather +than as skills; `*_scripts/` directories are copied verbatim; names starting +with `_` or `.` are skipped. When the roles start a task with this vertical, +those files are seeded into the project's skill library in the order project, +vertical, global, so a project's own edits win. + +## Step 3: package it for the store + +The store installs a zip whose members sit under `argus_verticals//`, +described by a `catalog.json`. The catalog entry the store validates +(`argus/verticals/store.py`, `_validate_entry`) needs: + +| Field | Constraint | +|---|---| +| `name` | equals the key; `^[a-z][a-z0-9_]{0,47}$` | +| `version` | non-empty, no slashes | +| `module` | `argus_verticals..stages`, and it must live in `paths[0]` | +| `paths` | at least `["argus_verticals/"]` | +| `purpose` | non-empty | +| `archive.file`, `archive.url`, `archive.sha256`, `archive.size` | the zip's name, location, digest and exact byte size; a non-`https` URL is only accepted when the catalog itself is local | + +Optional: `purpose_zh`, `requires` (other verticals to install first), +`shared` (helper trees outside the vertical's own directory), +`python_requirements` (shown, never installed), `tags`, `skill_parents`, +`has_skills`, `routing_path`, `min_argus`, `maintainers`. + +`examples/verticals/build_local_catalog.py` produces both files for a directory +under `examples/verticals/argus_verticals/`: + +```bash +python examples/verticals/build_local_catalog.py lab_notebook --version 0.1.0 --out /tmp/lab-store +``` + +``` +archive : /tmp/lab-store/lab_notebook-0.1.0.zip (3099 bytes) +catalog : /tmp/lab-store/catalog.json +install : ARGUS_VERTICAL_CATALOG=/tmp/lab-store/catalog.json argus verticals install lab_notebook +``` + +The catalog it wrote: + +```json +{ + "schema": 1, + "verticals": { + "lab_notebook": { + "name": "lab_notebook", + "version": "0.1.0", + "module": "argus_verticals.lab_notebook.stages", + "paths": ["argus_verticals/lab_notebook"], + "purpose": "small measurement tasks on this machine: run the measurement, record how it was run, and write a notebook entry another person can reproduce", + "has_skills": true, + "archive": { + "file": "lab_notebook-0.1.0.zip", + "url": "file:///tmp/lab-store/lab_notebook-0.1.0.zip", + "sha256": "5b36…b8e0", + "size": 3099 + } + } + } +} +``` + +The `sha256` and `size` are checked byte-for-byte at install; regenerate the +catalog whenever the zip changes. + +## Step 4: install it and confirm Argus sees it + +Point the store at the local catalog. A local catalog turns its directory into +an offline mirror: an archive named in the catalog and sitting next to it is +used without any download. + +```bash +export ARGUS_VERTICAL_CATALOG=/tmp/lab-store/catalog.json +argus verticals install lab_notebook +argus verticals list +argus verticals info lab_notebook +``` + +What this checkout printed (the install was done into a throwaway +`ARGUS_SKILL_HOME` so the machine's real store stayed untouched): + +``` +lab_notebook: install started + [ 0%] lab_notebook: starting + [100%] lab_notebook: install finished +lab_notebook: install finished +``` + +``` +name kind version installed enabled update purpose +... +lab_notebook installed 0.1.0 0.1.0 yes - small measurement tasks on this machine: run the measurement +``` + +``` +name lab_notebook +kind installed +purpose small measurement tasks on this machine: run the measurement, record how it was run, and write a notebook entry another person can reproduce +version 0.1.0 +installed_version 0.1.0 +enabled True +actions disable, uninstall +operation install done: install finished +``` + +The files landed at `/verticals/argus_verticals/lab_notebook/` +next to the store's `registry.json`, and the registry advertises the vertical +with origin `store`, its two skills found automatically next to `stages.py`: + +``` +advertised: True +origin: store | skills_root: .../verticals/argus_verticals/lab_notebook/skills +skills: ['engineer/measurement-record.md', 'reviewer/measurement-review.md'] +stages: ('measure', 'report') | gate: none | mode: staged +``` + +From here the Manager can choose `lab_notebook` for a task whose text matches +its purpose line; there is no flag to force a vertical (`ARGUS_SKILL_VERTICAL` +is a legacy name with no authority), and the choice is saved per project in +`.argus/PIPELINE_STATE.json`. To try it, start a project with an objective in +the vertical's own words, for example "measure how long `python -c 'import +torch'` takes on this machine, repeat it, and write a notebook entry with the +rerun command", and check `pipeline : vertical=lab_notebook` in +`argus --status`. + +`argus verticals disable lab_notebook` hides it again without removing the +files; `argus verticals remove lab_notebook` deletes them (it refuses while a +local project's `PIPELINE_STATE.json` still names the vertical, unless you +pass `--force`). + +## Publishing to the shared store + +The default catalog is the `catalog.json` attached to the latest release of +[Argus-AiTeam/argus-verticals](https://github.com/Argus-AiTeam/argus-verticals); +each vertical is one directory in that repository, and its release workflow +builds the zips and the catalog (`scripts/build_catalog.py` there). Publishing +therefore means opening a pull request that adds +`argus_verticals//` to that repository, with the same layout as here. The +store on every user's machine only accepts archives from an allow-listed host +(`github.com` and its release asset hosts), so a catalog you host elsewhere +must be reached through `ARGUS_VERTICAL_CATALOG` as a local file, which is what +the steps above do. + + + +Two other ways exist to run a vertical without the store, useful during +development of a bigger one: `pip install` a package that declares the entry +point `argus.verticals` (`name = "argus_verticals..stages"`), which wins +over a store copy of the same name; or install it as a managed workbench +plugin. A name that collides with a built-in is ignored from every source. + +## Checks before you share it + +- The contract check in step 1 passes. +- Every skill header has quoted `name` and `description` values. The + repository's own skill writer (`argus/skills/store.py`) emits exactly that + shape, JSON-quoted, because an unquoted description containing a colon does + not survive a YAML parse; `tests/skills/test_skill_frontmatter_integrity.py` + checks the in-tree verticals for the same reason. +- `argus verticals install` from a local catalog succeeds and + `argus verticals list` shows the row as `installed` and enabled. +- One real task ran through it and the Reviewer's verdicts referred to your + checklist ids (they appear in the round events in `events.jsonl`). + +The example under `examples/verticals/` is not scanned by +`tests/skills/test_vertical_plugins.py` (that file builds synthetic plugins in +memory), so adding a vertical there cannot change the test's result. On this +machine one test in that file fails before and after these changes for a +local reason: a pip-installed `argus_verticals` package in the venv makes +`test_store_verticals_are_discovered_with_origin_store` see two package paths +instead of one. diff --git a/docs/getting-started.md b/docs/getting-started.md new file mode 100644 index 000000000..a2a3de193 --- /dev/null +++ b/docs/getting-started.md @@ -0,0 +1,419 @@ +# Getting started with Argus + +This guide takes you from an empty machine to a finished first task. It covers +three ways to work: the command line, the web UI, and the desktop app. Every +command below was checked against the `argus` CLI in this checkout +(`argus 0.1.8`); the example runs are real projects, with their times and costs +taken from their own logs. + +Companion guides: [best practices](best-practices.md) (objectives, changing +direction mid-run, models and spend) and +[building a vertical](building-a-vertical.md). + +## What you are installing + +Argus is a Python package plus a bundled terminal cockpit (Node.js) and a web +UI. It does not call a model API itself; it drives one of the coding-agent CLIs +you already use (Copilot, Codex, Claude Code, Cursor, Pi, OpenCode, Grok, Qoder, +DeepSeek Harness) and lets four roles share it: a Manager that routes your +message, a Planner that decides the next task, an Engineer that does the work, +and a read-only Reviewer that decides whether the work is done. + +Two things follow from that design and explain most of what you will see: + +- **A background worker does the work.** Your cockpit, browser tab or terminal + can close; the project's daemon keeps running and keeps its state under + `~/.argus-skill/projects//`. +- **Nothing is "done" until the Reviewer says so.** A task usually takes + several Engineer rounds. Expect the first small task to take minutes, not + seconds. + +## Prerequisites + +| Requirement | Why | +|---|---| +| Python 3.11+ | the `argus` package (`pyproject.toml` requires `>=3.11`) | +| Node.js 22.12+ | the terminal cockpit is an Ink app; `argus` refuses to start it on older Node (`argus: Ink TUI requires Node.js 22.12 or newer`) | +| One authenticated agent CLI | Argus reuses its login; there is no separate Argus account | + +Install and log into one CLI first. The README's +[Quick Install](../README.md#quick-install) table lists the install and login +command for each backend; for example GitHub Copilot CLI is +`npm install -g @github/copilot` then `copilot login`. + +## Install + +The commands are the README's, repeated here so this page stands alone. + +Linux (isolated venv; use this on servers): + +```bash +git clone https://github.com/microsoft/ArgusAgent.git "$HOME/Argus" +cd "$HOME/Argus" +python3 -m venv .venv +.venv/bin/python -m pip install --upgrade pip +.venv/bin/python -m pip install -e . +ARGUS_BIN="$HOME/Argus/.venv/bin/argus" +"$ARGUS_BIN" --version +``` + +Windows (PowerShell, no venv): + +```powershell +py -m pip install --upgrade pip +py -m pip install --upgrade --force-reinstall "argus @ https://github.com/microsoft/ArgusAgent/archive/refs/heads/main.zip" +``` + +macOS (managed command): + +```bash +uv tool install --force --python 3.12 "argus @ https://github.com/microsoft/ArgusAgent/archive/refs/heads/main.zip" +``` + +On Linux, `argus` below means `$HOME/Argus/.venv/bin/argus` unless the venv is +active. Everywhere, `python -m argus ...` runs the same CLI without ever +starting the Node cockpit; it is what you want in scripts and over SSH without +a terminal. + +## First-time setup and a health check + +```bash +argus --setup +argus doctor +``` + +`argus --setup` is an interactive wizard: it asks which backend to use, checks +that the CLI is installed and logged in, runs one real model turn through it, +and only then saves the profile. It ends with `Setup complete. Run `argus`.` +If you would rather not answer prompts: + +```bash +argus --setup --backend copilot --non-interactive +``` + +Two variants worth knowing: + +| Situation | Command | +|---|---| +| Keep Argus on its own Copilot account, separate from your interactive login | `argus --setup --backend copilot --copilot-home "$HOME/.copilot-argus" --copilot-login` | +| Use an OpenAI-compatible endpoint (served through Pi) | `argus --setup --api-url URL --api-model MODEL` with the key in `ARGUS_SETUP_API_KEY` rather than `--api-key`, so it stays out of shell history | + +`argus doctor` is read-only. This is what it printed on the machine this guide +was written on (a source checkout, Pi backend): + +``` +argus doctor — cross-platform diagnostics + +✓ ARGUS-HOST-001 [host/supported_host] Linux 6.8.0-139-generic x86_64 +✓ ARGUS-INSTALL-001 [install/source_checkout] /data/.../Argus +✓ ARGUS-ASSET-001 [install/assets_ready] release manifest and Web/TUI assets are present +✓ ARGUS-PYTHON-001 [cli/python_ready] 3.12.3 at /data/.../.venv/bin/python +✓ ARGUS-NODE-001 [cli/node_ready] v22.23.2 +✓ ARGUS-WEB-001 [web/compatible] 127.0.0.1:8799 is compatible +! ARGUS-DESKTOP-001 [desktop/tauri_dependencies_missing] Tauri Desktop sources exist but npm/Rust build dependencies are incomplete + fix: run `npm --prefix desktop-tauri ci` and install the Rust Windows toolchain +✓ ARGUS-DAEMON-001 [daemon/stopped] no daemon is running +✓ ARGUS-BACKEND-001 [backend/ready] pi 0.85.1 runnable at /home/.../argus-pi (subscription_cli; authentication checked; ...) + +all blocking checks passed +``` + +A `!` line is advice, not a failure; the desktop warning above only matters if +you intend to build the desktop app from this checkout. When something is +wrong, `argus doctor --fix-safe` applies the repairs the doctor itself marked +safe and reruns; `argus doctor --advisor auto` lets one of the installed agent +CLIs inspect and repair. Without one of those two options `doctor` changes +nothing. + +A note on help: `argus --help` (and `argus doctor --help`, +`argus verticals --help`) print the same short overview. The complete flag +reference is behind an environment variable: + +```bash +ARGUS_SKILL_DEBUG_HELP=1 argus --help +``` + +## Path 1: the command line + +There are two ways to use the CLI. A bare `argus` opens the terminal cockpit, +where you type in natural language. Everything else (`--daemon`, `--status`, +`--notify`, ...) is a one-shot command that talks to the project's worker and +exits; those work over SSH, in cron, and without Node. + +### Which project a command means + +Argus keys a project's state on the directory you run it from. Management +commands attach to the newest session for the current directory, or to the +one you name: + +| Flag | Meaning | +|---|---| +| (none) | the newest session for the current working directory | +| `--project-root DIR` | the newest session for `DIR` | +| `--resume ID` | a specific session id (with no id, a picker of recent sessions) | +| `--continue` | the most recently active session | +| `--life-dir DIR` | use `DIR` instead of `~/.argus-skill` as the state root (`ARGUS_SKILL_HOME` does the same) | + +### Start a task without opening the cockpit + +Write the objective in a file, `cd` into the directory the work should happen +in, and start a background worker with a campaign objective: + +```bash +cd ~/work/matmul-roofline +argus --daemon --continuous --objective-file objective.txt --bounded --backend copilot +``` + +The pieces: + +| Flag | What it does and why you would set it | +|---|---| +| `--daemon` | start a detached worker (`--daemon-fg` keeps it in the foreground, for systemd or debugging) | +| `--continuous` | give the worker an objective and let the Planner keep generating tasks toward it; refused without an objective (`--continuous requires a non-empty --objective`) | +| `--objective TEXT` / `--objective-file PATH` | the objective; the file form keeps long text and quotes out of the shell (`--objective` without `--continuous` is refused too) | +| `--bounded` | stop when the Planner certifies the project done; without it the worker keeps generating work for the same objective | +| `--backend NAME` | the agent CLI for this daemon; it outranks the persisted setup for this launch and is exported to every role | +| `--mission-width N` | how many tasks may run in parallel (default 2; `1` is serial) | + +Before the worker starts, Argus re-checks that the backend is installed and +logged in; if not, it prints the readiness report and exits with code 3 instead +of starting a worker that cannot call a model. + +### A real first campaign + +The objective below started the campaign that project `s-02d3c282` ran on this +machine on 2026-09-30 (`objective-roofline.txt`, quoted in full because its +shape is what makes the rest work): + +> 做一项小规模但真实的研究:在本机 4 块 NVIDIA RTX A6000 上,PyTorch 方阵乘法的实际吞吐(TFLOPS)随矩阵规模(512 到 16384)和精度(FP32、TF32、BF16、FP16)如何变化,与理论峰值的差距能否用一个简单的 roofline 式模型(计算强度与显存带宽)解释。要求:真实实验、每种配置多次重复取中位数并报告离散度,拟合模型并给出误差,画图,最后写成一篇 4 页以内的短论文(含方法、结果、讨论、复现说明),并通过内部评审。所有数字必须来自本机真实运行,不允许估算或编造。 + +It was launched with `--backend copilot --mission-width 1` (the live process +still shows those flags). About two hours in, `argus --status` reported: + +``` + project : .../projects/s-02d3c282 + daemon : alive (pid 1430124, up 1h 48m, backend live — see /roles, width 1) + budget : global daily $1000.00 (spent $30.71) · remaining $969.29 + active : 0 pending · 1 running · 0 paused + current : + title : 重绘正式数据图并完成论文稿 + history : 2 done + cost : $29.79 cumulative + continuous: on + pipeline : vertical=research + lifecycle: + state : writing +``` + +By then the work directory held `paper/main.tex`, `paper/results.tex`, +`results/attempt-01`, `figures/` and `src/`; `paper/results.tex` reports, at +n=16384, 22.81 / 57.01 / 121.11 / 115.96 TFLOP/s for FP32 / TF32 / BF16 / FP16 +(58.95 / 73.66 / 78.24 / 74.91 % of the spec sheet) over 17,280 samples, and a +fitted roofline whose calibrated form has a 53.92 % mean absolute error on the +confirmation set. The campaign was still running when this page was written +(its log had a `paper` round returned as `replan_requested`, which is normal: +the Reviewer sent the draft back), so treat those as a snapshot, not a result. + +For a small first task, a bounded objective finishes in minutes. Project +`s-4a34a674` on the same machine was started the same way with this objective: + +> 用 numpy 数值求解一维阻尼谐振子 x'' + 2γx' + ω0² x = 0(取 ω0=2π, 分别取欠阻尼 γ=0.5、临界 γ=2π、过阻尼 γ=10),与解析解逐点比较给出最大误差,画出三条曲线,把方法、数值结果表和图写进本目录的 README.md。所有数字必须来自真实运行。 + +Its log runs from 08:11:39Z to 08:23:33Z (12 minutes), one task in two +attempts, `continuous.json` ends with `"done_reason": "planner declared project +done"`, and it cost $0.91 (12 model calls in `usage.jsonl`). The README it +wrote gives maximum errors of 6.0e-07, 1.6e-07 and 5.1e-07 for the three +damping cases against the analytic solution. + +### Watch, steer, stop + +| Command | What it does | +|---|---| +| `argus --status` | one screen: daemon, budget, current task, stage, last events | +| `argus --follow` | stream the event log to the terminal (`tail -f` style, Ctrl-C to stop) | +| `argus --watch` | the read-only live cockpit | +| `argus --notify "text"` | queue guidance for the next Engineer round (`--notify-stage STAGE` holds it until that stage) | +| `argus --answer "text"` | answer the question a paused task is waiting on (`--answer-item ID` when several wait) | +| `argus --ask "question"` | ask the Manager something and exit; nothing is queued and no daemon is needed | +| `argus --daemon-stop --drain` | let the current task finish at a clean boundary, then exit; the safe way to stop before upgrading | +| `argus --daemon-stop --force` | SIGKILL if it does not exit in time (interrupts running work) | +| `argus --daemon-runbook` | print the restart playbook for the current project | + +The difference between `--notify` and `--answer` matters: a nudge is read at +the next round and does not wake a task that has stopped to ask you something; +`--answer` does, and it also rewrites the task so the next round reads your +answer instead of re-asking. [Best practices](best-practices.md#changing-direction-while-it-runs) +goes into this. + +### The terminal cockpit + +```bash +argus +``` + +The cockpit needs a real terminal (piped or cron use gets a message pointing +you to `--web`, `--watch`, `--status` or `--daemon`). Type what you want in +plain language. The Manager reads every message and decides whether to answer +it itself (a question, a status request, a small local check) or to turn it +into a task for the Planner, Engineer and Reviewer (multi-step work, anything +that needs an independent review). You do not choose a vertical, a model or a +backend for the first task; the Manager picks the vertical from your text and +the backend comes from setup. + +`/new` (optionally `/new `) opens a two-field form, Name and +Objective (Tab or arrows switch fields, Enter creates), and switches the +cockpit to the new session; its worker starts when the first task arrives. + +Commands you will use in the first hour (`/help` lists them all): + +| Command | Purpose | +|---|---| +| `/task ` | queue work directly, skipping the "is this a chat?" decision | +| `/ask ` | answer inline; no task is queued | +| `/plan ` | preview the Planner's execution plan before committing | +| `/status`, `/roles`, `/backlog`, `/journal` | what is running, on which backend/model, what is queued, what happened | +| `/nudge ` | inject guidance into the running task | +| `/abort` | stop the running task now | +| `/backend [name]`, `/config [key=value]` | view or change the runner backend and runtime settings (persisted) | +| `/resume`, `/daemons` | switch to another project | +| `/quit` | leave; background work keeps running | + +## Path 2: the web UI + +```bash +argus --web +``` + +From the `argus` command this goes through the cockpit launcher, which picks +the first free port from 8799 and opens your browser at +`http://127.0.0.1:8799`. With an explicit host or port, or through +`python -m argus --web`, the Python server runs directly and does not open a +browser: + +```bash +argus --web --web-port 8800 +python -m argus --web # over SSH: then `ssh -L 8799:127.0.0.1:8799 user@server` +``` + +Binding to a LAN address (`--web-host 0.0.0.0`) always requires a bearer +token: `ARGUS_SKILL_WEB_TOKEN` if set, otherwise one is minted for the run and +printed with a QR code. The web assets are checked into the repository and +packaged with the wheel, so nothing needs building first. + +### The first screen + +With no sessions yet the page shows **Start a project**, the line *No sessions +yet. Create one to begin.*, and a **New project** button. The button creates an +idle session (no form; the toast reads *New session ready. Say what it is for +and the work begins.*). Then type the objective into the message box. The +message goes to `POST /api/projects/{sid}/message`, the Manager classifies it, +and if it is work the session's daemon is started on demand and the task is +queued. The web UI follows your browser language (English or Simplified +Chinese); the sidebar has a language button. + +### A real first task from the browser + +Project `s-67fb6d62` in the fresh-install trial on this machine was created +from the web UI (`session.json` records `origin: web`). Its first message, at +06:57:51Z: + +> 帮我在这台机器的 GPU 上测一下 PyTorch 矩阵乘在 fp16 和 fp32 下的 TFLOPS,写个脚本跑出真实数字,整理成表格。 + +The reply, three minutes later at 07:02:21Z, reported an average of 121.31 +TFLOPS in FP16 and 23.96 in strict FP32 at 16384×16384 across the four GPUs, +with a per-GPU table and the script name (`benchmark_torch_matmul.py`, seven +timed batches, medians). A follow-up at 07:03:44Z, "现在项目里都有什么文件?把 bf16 +也补测一下,更新表格。", produced an updated table at 07:07:00Z (FP16 122.08, +BF16 125.00, FP32 23.87 on average). The project's usage ledger shows $0.80 for +the whole exchange, plus two early calls that were never priced (see below). + +That first message was actually sent twice. The first attempt (also 06:57:51Z) +came back as `[not dispatched] Manager could not classify this message +(refused before start: unresolved provider cost ...)`: on a fresh Copilot +install the first call's price was not yet known, and the default policy +refuses to spend money it cannot price. [Best practices](best-practices.md#controlling-spend) +explains the setting (`ARGUS_SKILL_UNPRICED_COST_POLICY`) and the trade-off. + +### The endpoints behind the UI + +If you script against the server, these are the calls the UI itself makes +(all need the bearer token when one is configured): + +| Endpoint | Body | Purpose | +|---|---|---| +| `POST /api/daemons` | `objective`, `name`, `workdir`, `launch_cwd` | create a session (an objective starts its worker immediately) | +| `POST /api/projects/{sid}/message` | `text`, optional `route_override` (`auto`/`chat`/`task`), `attachments` | send a message through the Manager; `/message/stream` is the SSE twin | +| `POST /api/projects/{sid}/tasks` | `text`, `autostart_daemon` (default true) | queue work directly | +| `POST /api/projects/{sid}/nudge` | `text` | guidance for the next round (429 when the queue is full) | +| `POST /api/projects/{sid}/backlog/{item_id}/answer` | `text` | answer a paused question and resume | +| `POST /api/projects/{sid}/continuous` | `enabled`, `objective` | turn a campaign objective on or off | +| `GET /api/projects/costs` | | spend per project | + +## Path 3: the desktop app + +The desktop app is a Tauri host around a frozen copy of the same Python +backend; it opens the same web cockpit. It is a packaged preview channel, +separate from source updates. + +- **Download:** the Releases page of + [lbx154/Argus](https://github.com/lbx154/Argus/releases). Windows ships as + an NSIS installer named `Argus--setup.exe`. The release workflow can + also build macOS (`.dmg`, arm64 and x86_64) and Linux (`.AppImage`, `.deb`) + packages, but the workflow file notes that v0.1.7 and v0.1.8 were published + for Windows only, so check the assets of the release you pick. + `microsoft/ArgusAgent` Releases is a separate channel; do not mix installers + or updates between the two. +- **What it needs:** nothing else on the machine. The installer bundles the + backend (`argus-backend`), so no Python, Node.js or venv is required. +- **First launch:** a setup wizard asks you to pick an installed, logged-in + agent CLI and confirm it; the local backend does not start until you do. If + you have an internal trial key instead, paste it and click **开始试用** ("start + trial"): the app downloads the official standalone Copilot program, checks + its SHA-256, runs one real model reply, and opens the workbench. That path + needs no CLI, Python, Node or coding-agent account of your own. Either + choice is remembered; **File → Settings** changes the CLI, executable or + port later. +- **After that** the app is the web UI from Path 2: create a project, type the + objective, watch the feed. + +To build it from source instead (Python 3.11+, Node 22.12+, Rust stable, plus +MSVC Build Tools on Windows or Xcode command-line tools on macOS), the +documented sequence in [docs/desktop-trial.md](desktop-trial.md) is: + +```bash +python -m pip install -e ".[trial]" "pyinstaller>=6.11,<7" tzdata +npm --prefix frontend/web ci +npm --prefix frontend/tui ci +npm --prefix desktop-tauri ci +npm --prefix frontend/web run build && npm --prefix frontend/tui run build +npm --prefix desktop-tauri run build:backend +npm --prefix desktop-tauri run build:unsigned +``` + +The unsigned bundle lands under `desktop-tauri/src-tauri/target/release/bundle/`. +[docs/windows-desktop.md](windows-desktop.md) covers signing and the update +channel. + +## What "finished" looks like + +Whichever path you used, the task ends the same way. The Reviewer returns +`done` for the last task; with `--bounded` (or a bounded objective in the +cockpit) the Planner then declares the project done and the worker stops, and +`argus --status` shows `history : N done` with no running task. The evidence is +in the work directory (the oscillator task's `README.md`; the roofline +campaign's `paper/` and `results/`), and the per-call record of what it cost is +in `~/.argus-skill/projects//usage.jsonl`. + +If instead the status shows a task `paused`, read the question with +`argus --status` or in the cockpit and answer it with `argus --answer`; if it +shows `paused_cost`, see [controlling spend](best-practices.md#controlling-spend). + +## Where things live + +| Path | Contents | +|---|---| +| `~/.argus-skill/` | the state root (`ARGUS_SKILL_HOME` or `--life-dir` overrides it) | +| `~/.argus-skill/config.json` | persisted settings from `--setup`, `/config`, `/backend` | +| `~/.argus-skill/projects//` | one project's state: `events.jsonl`, `usage.jsonl`, `backlog.jsonl`, `continuous.json`, `daemon.log` | +| `~/.argus-skill/verticals/` | verticals installed from the store | +| your work directory | where the roles read and write files; never inside the state root | diff --git a/examples/verticals/argus_verticals/lab_notebook/__init__.py b/examples/verticals/argus_verticals/lab_notebook/__init__.py new file mode 100644 index 000000000..f301d5db4 --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/__init__.py @@ -0,0 +1,5 @@ +"""``lab_notebook``: the worked example from docs/building-a-vertical.md. + +A two-stage vertical for small measurement tasks: measure something on this +machine, then write the result up so another person can rerun it. +""" diff --git a/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md b/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md new file mode 100644 index 000000000..a60c336b6 --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/skills/engineer/measurement-record.md @@ -0,0 +1,20 @@ +--- +name: "Measurement Record" +description: "How to run a small measurement so the result can be checked and rerun: save raw output, repeat, report the spread, and record the exact command." +--- + +# Measurement record + +Run the measurement from a script saved in the work directory, never from an +interactive shell you cannot show later. Write the raw output to a file under +`results/` before computing any aggregate. + +Repeat the run. Report the median (or mean) together with the spread (min/max or +standard deviation) and the number of repeats. A single run is not a result. + +Record in `NOTEBOOK.md`: + +- what was measured and on which hardware +- the exact command to rerun it +- the aggregate, the spread and the number of repeats +- the path of the raw output every number came from diff --git a/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md b/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md new file mode 100644 index 000000000..d9768ad3d --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/skills/reviewer/measurement-review.md @@ -0,0 +1,16 @@ +--- +name: "Measurement Review" +description: "How to review a measurement task: open the raw output files, recompute one aggregate, and confirm the rerun command is complete before accepting." +--- + +# Measurement review + +Do not accept the Engineer's summary as evidence. Open the raw output files +named in `NOTEBOOK.md` and recompute at least one aggregate yourself. + +Confirm that the rerun command in `NOTEBOOK.md` names the script, its arguments +and the working directory. A command that cannot be pasted and run is +incomplete. + +Return `continue` when a number has no raw file behind it, and `done` only when +every checklist item of the current stage has evidence you inspected. diff --git a/examples/verticals/argus_verticals/lab_notebook/stages.py b/examples/verticals/argus_verticals/lab_notebook/stages.py new file mode 100644 index 000000000..491334251 --- /dev/null +++ b/examples/verticals/argus_verticals/lab_notebook/stages.py @@ -0,0 +1,81 @@ +"""Stages and checklists of the ``lab_notebook`` example vertical. + +Everything Argus needs to know about a vertical is declared at module level in +this file; there is no base class to inherit from. The framework reads the +attributes below through ``argus.core.vertical_contract.vertical_contract`` and +``argus.verticals._registry._validated_plugin``. +""" + +from __future__ import annotations + +from argus.skills.stage_machine import ChecklistItem + +# Required for every vertical that is not built in. Argus refuses any other +# API version, and an empty purpose hides the vertical from the Manager. +ARGUS_VERTICAL_API_VERSION = 1 +VERTICAL_PURPOSE = ( + "small measurement tasks on this machine: run the measurement, record how " + "it was run, and write a notebook entry another person can reproduce" +) + +# The stage order is the tuple order. Every stage that is not optional needs a +# non-empty checklist below. +CHECKLIST_STAGE_ORDER: tuple[str, ...] = ("measure", "report") +CHECKLIST_OPTIONAL_STAGES: tuple[str, ...] = () +# Names the Manager may use for a stage; they are canonicalized to the real one. +STAGE_ALIASES = {"experiment": "measure", "writeup": "report"} + +# "none": the project is complete when the last stage's checklist is met. +# "metric" and "certified" are the other two values the contract accepts. +completion_gate = "none" +WORKFLOW_MODE = "staged" # staged | direct | proportional +MISSION_KIND = "custom" # custom | optimize | research | software +REQUIRE_INDEPENDENT_REVIEW = True + +CHECKLIST_ITEMS: dict[str, tuple[ChecklistItem, ...]] = { + "measure": ( + ChecklistItem( + id="measure.ran-here", + statement=( + "The measurement was executed on this machine in this project, " + "and its raw output is saved in the work directory." + ), + evidence_hint="the command that was run and the path of its raw output", + ), + ChecklistItem( + id="measure.repeated", + statement=( + "The measurement was repeated, and the reported number is a " + "median or mean with its spread; a single run is not a result." + ), + evidence_hint="number of repeats, the aggregate and the spread", + ), + ), + "report": ( + ChecklistItem( + id="report.notebook-entry", + statement=( + "NOTEBOOK.md in the work directory states what was measured, " + "how, the result with its spread, and the exact command to rerun it." + ), + evidence_hint="the NOTEBOOK.md section and the rerun command", + ), + ChecklistItem( + id="report.numbers-traceable", + statement=( + "Every number in NOTEBOOK.md can be traced to a saved raw output file." + ), + evidence_hint="raw file paths next to each number", + ), + ), +} + + +def role_banner(role: str) -> str: + """One paragraph every role reads before its task; keep it about the field.""" + return ( + "LAB NOTEBOOK VERTICAL: measure first, then write. A number without a " + "saved raw output and a rerun command is not a result. The Reviewer " + "checks the notebook entry against the raw files, not against the " + "Engineer's summary." + ) diff --git a/examples/verticals/build_local_catalog.py b/examples/verticals/build_local_catalog.py new file mode 100644 index 000000000..f8dfee9a4 --- /dev/null +++ b/examples/verticals/build_local_catalog.py @@ -0,0 +1,97 @@ +"""Package one vertical directory as a Vertical Store archive plus a local catalog. + +Usage: + + python examples/verticals/build_local_catalog.py NAME --version 0.1.0 --out DIR + +Reads ``examples/verticals/argus_verticals/NAME`` (or ``--source``), writes +``DIR/NAME-VERSION.zip`` with members under ``argus_verticals/NAME/`` (the layout +``argus.verticals.store._verify_tree`` expects), and writes ``DIR/catalog.json`` +in the shape ``argus.verticals.store._validate_entry`` accepts. Point +``ARGUS_VERTICAL_CATALOG`` at that file and ``argus verticals install NAME`` uses +the archive next to it instead of downloading anything. + +The catalog is minimal on purpose: the fields the store requires, nothing else. +See docs/building-a-vertical.md for the walk-through. +""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import re +import sys +import zipfile +from pathlib import Path + +_NAME = re.compile(r"^[a-z][a-z0-9_]{0,47}$") + + +def _purpose(stages: Path) -> str: + """Read VERTICAL_PURPOSE from stages.py without importing it.""" + text = stages.read_text(encoding="utf-8") + match = re.search(r"VERTICAL_PURPOSE\s*=\s*\(?\s*((?:\"[^\"]*\"\s*)+)\)?", text) + if match is None: + raise SystemExit(f"{stages}: VERTICAL_PURPOSE not found") + return " ".join("".join(re.findall(r"\"([^\"]*)\"", match.group(1))).split()) + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0]) + parser.add_argument("name", help="vertical name, e.g. lab_notebook") + parser.add_argument("--version", default="0.1.0") + parser.add_argument("--source", type=Path, default=None, + help="vertical directory (default: examples/verticals/argus_verticals/NAME)") + parser.add_argument("--out", type=Path, required=True, help="directory for the zip and catalog.json") + args = parser.parse_args(argv) + + name = args.name + if not _NAME.fullmatch(name): + raise SystemExit(f"{name!r} is not a valid vertical name (^[a-z][a-z0-9_]{{0,47}}$)") + source = args.source or (Path(__file__).resolve().parent / "argus_verticals" / name) + stages = source / "stages.py" + if not stages.is_file(): + raise SystemExit(f"{source} has no stages.py") + + out = args.out.resolve() + out.mkdir(parents=True, exist_ok=True) + archive_name = f"{name}-{args.version}.zip" + archive = out / archive_name + tree = f"argus_verticals/{name}" + with zipfile.ZipFile(archive, "w", compression=zipfile.ZIP_DEFLATED) as zf: + for path in sorted(p for p in source.rglob("*") if p.is_file()): + if "__pycache__" in path.parts: + continue + zf.write(path, f"{tree}/{path.relative_to(source).as_posix()}") + + data = archive.read_bytes() + catalog = { + "schema": 1, + "verticals": { + name: { + "name": name, + "version": args.version, + "module": f"argus_verticals.{name}.stages", + "paths": [tree], + "purpose": _purpose(stages), + "has_skills": (source / "skills").is_dir(), + "archive": { + "file": archive_name, + "url": archive.as_uri(), + "sha256": hashlib.sha256(data).hexdigest(), + "size": len(data), + }, + } + }, + } + catalog_path = out / "catalog.json" + catalog_path.write_text(json.dumps(catalog, indent=2, ensure_ascii=False) + "\n", encoding="utf-8") + print(f"archive : {archive} ({len(data)} bytes)") + print(f"catalog : {catalog_path}") + print(f"install : ARGUS_VERTICAL_CATALOG={catalog_path} argus verticals install {name}") + return 0 + + +if __name__ == "__main__": + sys.exit(main())