diff --git a/.gitignore b/.gitignore index 40c24b8..43c8a9e 100644 --- a/.gitignore +++ b/.gitignore @@ -14,8 +14,9 @@ TestResults/ # Local tool output .playwright-mcp/ -# axm's local state: the task scope and caches +# axm's local state: the task scope and caches, except the scopes eval cases start with .axm/ +!*/*/evals/*/*/repo/.axm/ # OS files .DS_Store diff --git a/AGENTS.md b/AGENTS.md index f6e6870..13b62a5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -27,12 +27,12 @@ The tool is only worth trusting if these hold. Never break them, not even in deb - **Local first.** No accounts, telemetry, hosted services, update checks or uploads. The only network traffic is model calls made by `axm eval`, `axm triggers`, `axm fossil` and `axm conflicts --judge`, and only when the user runs them. - **A deterministic core.** `Axiomarium.Core` makes no model or network calls. A test fails if it references `System.Console`, Spectre.Console or `System.Net.Http`. -- **Writes only on request.** `list`, `validate`, `doctor`, `hook`, `detect`, `explain`, `conflicts` and `triggers` (overlap, `test` and `export`) are read-only. `init`, `sync`, `incident new` and `triggers generate` show what they will write and wait for approval in the same run, and refuse before any model call when there's no way to ask. `fossil` and `distill` recommend changes and never delete or rewrite instructions. +- **Writes only on request.** `list`, `validate`, `doctor`, `hook`, `detect`, `explain`, `conflicts` and `triggers` (overlap, `test` and `export`) are read-only. `init`, `sync`, `incident new` and `triggers generate` show what they will write and wait for approval in the same run, and refuse before any model call when there's no way to ask. `eval run` and `eval compare` write only their history, in `.axm/evals/`, and run sessions in a sealed home, never in the user's own. `fossil` and `distill` recommend changes and never delete or rewrite instructions. - **Label model output.** Anything a model produced, such as a judged contradiction or a distilled root cause, says so in the output. - **Every explain entry cites its rule.** Each loaded or dropped file carries the loading rule that produced it. If you can't name the rule, don't emit the entry. - **Harness models follow the docs, then the real harness.** When you change a model, cite the doc section and update the docs date and harness version recorded in the model. If the docs and the real harness disagree, the real harness wins, and the disagreement goes into `ROADMAP.md` as a decision. - **Don't guess what the model will read.** Skills are "available", never "loaded". When a harness leaves loading to the model, such as a Codex AGENTS.md below the launch directory, report it as "not loaded by the harness". -- **Tests never read the real machine.** Home, `CODEX_HOME`, Codex's system folder, managed-policy and settings locations are injected. Fixtures and docs use placeholder paths such as `/home/dev` and `C:\Users\dev`, never real ones. `CliRun` refuses `doctor`, `explain` and `triggers` without an injected machine, and `triggers generate` and `triggers test` without an injected runner, so no test runs a model. The native binary tests only run commands that read no harness files. +- **Tests never read the real machine.** Home, `CODEX_HOME`, Codex's system folder, managed-policy and settings locations are injected, and so is the operating system `axm eval` plans for, which `CliRun` pins to Linux. Fixtures and docs use placeholder paths such as `/home/dev` and `C:\Users\dev`, never real ones. `CliRun` refuses `doctor`, `explain`, `triggers`, `eval run`, `eval compare` and `conflicts` without an injected machine, and `triggers generate`, `triggers test`, `eval run`, `eval compare` and `conflicts` without an injected runner, so no test runs a model. The native binary tests only run commands that read no harness files. - **Scenarios are ground truth.** Each one in `scenarios//` is a tiny repo, a fake home, `scenario.yaml` and the recording in `expected.json`. Every Markdown file in it starts with its `MARKER ` line (after any frontmatter), so a recording can name the file even when a harness cuts it short. A skill's description starts with the same marker, because a listing shows descriptions, and a hook's command is `echo MARKER