From 812b7390be7c800ca8a37e7eeb3defc208061182 Mon Sep 17 00:00:00 2001 From: Hazel Granados Date: Tue, 29 Sep 2026 14:30:11 -0500 Subject: [PATCH 01/43] docs(roadmap): plan M5, evals --- ROADMAP.md | 37 +++++++++++++++++++++++++++++-------- 1 file changed, 29 insertions(+), 8 deletions(-) diff --git a/ROADMAP.md b/ROADMAP.md index 7f6c422..f8a14d9 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -32,10 +32,10 @@ Axiomarium is my lab for building, testing and debugging AI coding environments - **YAML follows the 1.2 core schema:** `yes` and `on` stay strings, and a number JSON can write as is keeps its source text, so an error can say `Quote it: "1.0"`. Other number forms, such as hex, octal, `+1` or `.5`, become their value. - **The doctor checks each asset in isolation.** YamlDotNet sometimes throws exceptions that aren't `YamlException`, and reports some errors (tab indentation) at line 1. So reading YAML never throws, a wrong line is replaced by the real one or by none, and anything unexpected while checking one manifest becomes an error on that file while the rest of the vault is still checked. - **Versions:** between releases, `main` carries the next version with a `-dev` suffix, such as `0.2.0-dev`. A release commit drops the suffix and dates the version's section in `CHANGELOG.md`, and its tag must match. -- **Projects:** `Axiomarium.Core` (assets, schemas, registry, instruction resolution) and `Axiomarium.Cli`. `Axiomarium.Eval` and `Axiomarium.Adapters` arrive with their milestones. +- **Projects:** `Axiomarium.Core` (assets, schemas, registry, instruction resolution) and `Axiomarium.Cli`. `Axiomarium.Adapters` arrives with its milestone. Evals get no project of their own (2026-09-29): as with triggers, their pure parts live in the Core and the processes in the CLI. - **The command is `axm`.** `axiom` would collide with the CLI of Axiom (axiom.co), and their npm package `axiom` is an AI evals SDK in the same space. - **Versioned releases.** Each milestone from M1 on is a release (v0.1, v0.2 …), tagged `vX.Y.Z`, with NativeAOT binaries for `win-x64`, `linux-x64` and `osx-arm64` on GitHub Releases, `SHA256SUMS`, the `Axiomarium` dotnet tool on NuGet, and a `CHANGELOG.md` entry. Each asset also carries its own SemVer `version` in its manifest. -- **Storage:** configuration lives in git (`axiomarium.yaml`, `registry/`, `evals/`, `incidents/`). SQLite is only for disposable local state in `.axm/cache.db`: eval history, hashes, evidence and cached scans. The repo never needs the database to be understood. +- **Storage:** configuration lives in git (`axiomarium.yaml`, `registry/`, `evals/`, `incidents/`). Disposable local state lives in `.axm/`, which is never committed, and the repo never needs it to be understood. Eval history is JSON run files, not SQLite (changed 2026-09-29): under NativeAOT, SQLite ships a native library with every binary and tool package, and history that is saved `--json` output keeps one contract instead of two. Evidence and cached scans choose their format in their milestones. ### Assets @@ -49,7 +49,7 @@ Axiomarium is my lab for building, testing and debugging AI coding environments - **Each asset has one content file named after its kind:** `agent.md`, `skill.md`, `hook.md`, `policy.md`, `workflow.md` or `experiment.md`. Adapters generate harness files such as `SKILL.md` from it. - **Skills, hooks and policies carry a block named after their kind** in `asset.yaml`, validated by `schemas/skill.schema.json`, `hook.schema.json` and `policy.schema.json`. A skill says when it should activate (`use_when`), a hook says whether it blocks, warns or only gathers evidence (`response`), and each policy rule names what enforces it (`enforced_by`). Agents need nothing beyond `permissions` yet, so there is no agent schema. - **References are checked:** the content file exists and isn't empty, each `enforced_by` names an asset that exists, and each `evals.: true` has at least one file in the asset's `evals//`. -- **Maturity evidence is data** in `registry/maturity.yaml`: each level's promise and what it requires. The binary embeds it, like the schemas, so every vault is judged by the same promises. Usage evidence is a dated entry in `docs/dogfooding.md`: a heading such as `## 2026-10-02 · carmine-workbench` whose section links to the asset's folder. Incubating needs 1 entry. Tested adds behavioral and regression evals, plus trigger evals for skills. Stable needs 3 entries, and battle-tested 5 entries across 3 repos. A maturity claim without its evidence is an error, because the manifest says something untrue. Until v0.5, eval evidence means the eval files exist, not that they pass. +- **Maturity evidence is data** in `registry/maturity.yaml`: each level's promise and what it requires. The binary embeds it, like the schemas, so every vault is judged by the same promises. Usage evidence is a dated entry in `docs/dogfooding.md`: a heading such as `## 2026-10-02 · carmine-workbench` whose section links to the asset's folder. Incubating needs 1 entry. Tested adds behavioral and regression evals, plus trigger evals for skills. Stable needs 3 entries, and battle-tested 5 entries across 3 repos. A maturity claim without its evidence is an error, because the manifest says something untrue. Until v0.6, eval evidence means the eval files exist, not that they pass: eval results become evidence with the M6 evidence store, which already tells fresh from stale (moved from v0.5 on 2026-09-29). - **`validate` and `doctor` render the same result.** `validate` prints only the problems, for CI and hooks. `doctor` adds the inventory, and from v0.2 the instruction findings. `list` shows the inventory and judges nothing. - **No `--json` in v0.1.** Agents and hooks read the plain output. JSON becomes a contract with `explain` in v0.2. - **`inspect` and `init` move to v0.10.** `inspect` would duplicate `detect`, and `init` has nothing to write until it can detect a stack and install assets. Until then a vault is recognized by its kind folders. @@ -169,6 +169,19 @@ Axiomarium is my lab for building, testing and debugging AI coding environments - **`promptfoo validate` checks structure, not names (2026-09-29).** promptfoo 0.123.1's `validate` rejects an `assert` that isn't a list, but accepts unknown keys and assertion types. So CI validates the export's structure, with a broken export as the positive control, and unit tests pin the names from promptfoo's docs: `skill-used`, `not-skill-used`, `anthropic:claude-agent-sdk` and `openai:codex-sdk`. The export says that `skill-used` passes only once the skill is installed where the harness finds skills, which is sync's job from v0.9. - **Docs stay in `docs/` (2026-09-28).** A synced GitHub wiki was considered and dropped: it's a separate git repo without PR review or CI, and links into `findings/` and `scenarios/` break there. `docs/README.md` is the home and the reading order. Reference pages sit next to the write-ups and link to what exists instead of repeating it, and tests check that links resolve and that every command and finding id is documented. The repo's GitHub Wiki tab goes off, so there's one place to look. A generated site can come from the same files near v1.0. +### Evals (v0.5) + +- **Evals run on `axm`'s own runner (2026-09-29).** It builds on M4's throwaway copy, runner interface and stream parsers. promptfoo can't seal the home or take a baseline from a git ref, and it would add Node, so an export for it waits in Later. The harnesses are the providers: Claude Code for Anthropic and Codex for OpenAI, behind the one runner interface, so no eval depends on one vendor, and there's still no API key and no HTTP client. +- **An eval case is a folder.** `/evals/behavioral//` or `/evals/regression//` holds `eval.yaml`, validated by `schemas/eval.schema.json`, and `repo/`, a tiny repo copied for each run with an empty `.git`. `eval.yaml` has the `prompt`, the `checks`, an optional `judge.rubric` and `allow`, the commands the session may run. A regression case names the failure it guards in `guards:`, and exists only for a failure that happened. Setting `evals.behavioral: true` or `evals.regression: true` stays the owner's claim. +- **Checks carry the verdict, and a judge is optional and labeled.** There are five checks, all deterministic: `file` (exists or not, contains a text), `run` (a command run in the copy after the session and its exit code, where `axm` means the running binary), `loaded` (skills loaded or not), `ran` (commands the session ran or not) and `reply` (the final message contains a text or not). A run passes when every check passes. A `judge.rubric` is graded by a model, shown on its own line as model judgment, and never changes what the checks decided. +- **The asset is installed in the copy the way sync will install it in v0.9.** A skill becomes each harness's skill file, as in M4. A hook becomes a hooks entry in the copy's `.claude/settings.json` that runs the current `axm` binary. An agent becomes `.claude/agents/.md`, and runs on Codex only if the spike finds how Codex loads one. Policies, workflows and experiments get no evals in M5: a policy's rules are tested through the assets that enforce them. +- **Sessions run in a sealed home by default.** A fake home sits next to the copy (`C:\axm-evals\` on Windows, `/tmp/axm-evals/` elsewhere): `CLAUDE_CONFIG_DIR` with the Claude Code login borrowed, as the ground-truth recorder does, and `CODEX_HOME` with Codex's `auth.json`. Only the model the user chose comes along (Claude Code's `model` setting, Codex's `model` and reasoning effort), so results depend on the asset and not on the user's plugins, hooks, MCP servers or global instructions, and history stays comparable when that setup changes. `--model` overrides it, and `--real-home` runs in the user's real setup, with a warning that their hooks and MCP servers run too. The copy, the fake home and the borrowed login are deleted afterwards, and a run removes what a killed run left behind and says so, because the login is a credential. Each sealed Claude Code run compares the skills in its `init` event with the ones it installed, and reports any other as "the home leaked". On macOS, Claude Code keeps its login in the Keychain, so a sealed run there refuses with a hint to use `--real-home` until it's confirmed to work. +- **Sessions write only in the copy and run only what the case allows.** Claude Code runs with `--permission-mode acceptEdits`, `--allowedTools` built from the case's `allow` (such as `Bash(axm validate:*)`) and `--strict-mcp-config` with no servers; nothing else can be approved in `-p`, so everything else is denied. Codex runs with `-s workspace-write`, which keeps writes in the copy and the network off; if its rules can't hold a session to `allow`, a command outside it shows in the `ran` checks instead of being blocked. Sessions run 4 at a time, each with a timeout, and Claude Code's with a turn cap, both set by the spike. +- **`axm eval run` reports counts, not verdicts.** Each case runs 3 times on each harness the asset supports, and `--harness`, `--runs` and `--model` override that. For each case it shows the runs passed, each failed check with its run, the judge's line, and the median and range of tokens (input, cached and output), wall time and turns, plus cost where the harness reports it, which is Claude Code only. It says how many sessions it will run before it starts. It exits 0 whenever it ran and 2 when nothing could run; failing on a result waits for `axm regress` in M7. +- **`axm eval compare` changes only the asset.** The baseline is the asset's files at a git ref (`--baseline `, HEAD by default, read with `git show`) or no asset at all (`--baseline none`), and the candidate is the working tree. It stops when the asset is unchanged. Both versions run the working tree's cases, and the output says so when the cases changed too. Baseline and candidate sessions alternate, so a rate limit or a slow hour hits both. The report sets the two side by side: runs passed, and tokens and time as median and range with the difference, with no verdict and no claim of significance. +- **History is saved output.** Every `run` and `compare` saves its `--json` output as `.axm/evals//