Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
812b739
docs(roadmap): plan M5, evals
hazeliscoding Sep 29, 2026
b358914
test(evals): add the M5 spike's harness streams as fixtures
hazeliscoding Sep 29, 2026
bf0991c
docs(roadmap): record what the M5 spike settled, and seal Codex by di…
hazeliscoding Sep 29, 2026
7549a0f
feat(schemas): add eval cases, checked by the doctor
hazeliscoding Sep 29, 2026
431f1ee
feat(evals): decide an eval case's checks from what the harness recorded
hazeliscoding Sep 29, 2026
320d870
feat(evals): read a whole session from each harness, subagents included
hazeliscoding Sep 29, 2026
4231a42
docs(roadmap): run M5's evals sealed only, and move --real-home to Later
hazeliscoding Sep 29, 2026
74b1170
feat(evals): install an asset for each harness the way sync will
hazeliscoding Sep 29, 2026
a215ad8
feat(evals): plan a sealed home that borrows only the logins and mode…
hazeliscoding Sep 29, 2026
835f6af
feat(evals): plan an eval run from the vault's cases
hazeliscoding Sep 29, 2026
7b56014
feat(evals): count each case's runs, with the median and range of wha…
hazeliscoding Sep 30, 2026
49709ec
feat(cli): give harness sessions their own environment, and remove th…
hazeliscoding Sep 30, 2026
0371894
feat(evals): run eval sessions in a sealed home, each in its own copy…
hazeliscoding Sep 30, 2026
7da1658
feat(evals): add axm eval run, which counts each case's passed runs a…
hazeliscoding Sep 30, 2026
4b4167f
fix(evals): delete copies and leftovers whose git objects are read-only
hazeliscoding Sep 30, 2026
37351bd
fix(evals): read a rollout that a stopped Codex still holds open
hazeliscoding Sep 30, 2026
9c64c72
fix(evals): keep why a session stopped, and its time, when its checks…
hazeliscoding Sep 30, 2026
29bd986
fix(evals): warm Codex's sandbox in a git repo shaped like a session'…
hazeliscoding Sep 30, 2026
69a23e0
feat(evals): time each Codex command, and name the paths a run read o…
hazeliscoding Sep 30, 2026
b7be25a
docs(roadmap): record that Codex on Windows reads the whole disk, and…
hazeliscoding Sep 30, 2026
b966782
feat(judging): read a judge's contradictions, and keep only those it …
hazeliscoding Sep 30, 2026
51be655
feat(conflicts): add axm conflicts --judge, which keeps only contradi…
hazeliscoding Sep 30, 2026
4c192fb
feat(evals): grade each run against its case's rubric, as model judgm…
hazeliscoding Sep 30, 2026
973875f
docs(roadmap): tick axm conflicts --judge and the eval judge
hazeliscoding Sep 30, 2026
07ae84e
feat(evals): run a session with one version of its asset, or none
hazeliscoding Sep 30, 2026
68cc024
feat(evals): read saved runs back, and find a baseline a compare can …
hazeliscoding Sep 30, 2026
ebdb9f2
test(cli): stop the fixed clock's timestamps, so a measured time can'…
hazeliscoding Sep 30, 2026
7adfdc9
feat(evals): add axm eval compare, which shows a baseline and the wor…
hazeliscoding Sep 30, 2026
690685d
docs(roadmap): record how compare reuses history, and tick axm eval c…
hazeliscoding Sep 30, 2026
13f8b7f
fix(evals): leave commands the harness denied out of what a session ran
hazeliscoding Sep 30, 2026
c62088b
fix(evals): read a Git Bash path such as /c/x as the Windows path it …
hazeliscoding Sep 30, 2026
ac5113c
fix(evals): keep a long case name apart from its count, and say a run…
hazeliscoding Sep 30, 2026
453abcc
fix(evals): don't count Codex reading its own skills in the sealed ho…
hazeliscoding Sep 30, 2026
965dd7e
test(assets): add behavioral evals for the four starter assets
hazeliscoding Sep 30, 2026
f28d823
docs(roadmap): tick the starter assets' behavioral evals
hazeliscoding Sep 30, 2026
ed1250f
fix(evals): fit a compare's columns to a median and its range
hazeliscoding Sep 30, 2026
e0ccb5b
fix(evals): don't read a glob's exclusion such as '!/.git/**' as a path
hazeliscoding Sep 30, 2026
28b3ee0
feat(evals): name each compared version's slowest call and the paths …
hazeliscoding Sep 30, 2026
afbdff4
feat(assets): have agent-asset-authoring name the tool that should en…
hazeliscoding Sep 30, 2026
5117299
docs: write "Did the change make the skill better?" on the first real…
hazeliscoding Sep 30, 2026
8602f79
docs(roadmap): record the stalls after Codex's warm-up, and tick the …
hazeliscoding Sep 30, 2026
6a84856
chore(release): prepare v0.5.0
hazeliscoding Sep 30, 2026
178f4c7
fix(evals): take the operating system from the caller, so eval tests …
hazeliscoding Sep 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -14,8 +14,9 @@ TestResults/
# Local tool output
.playwright-mcp/

# axm's local state: the task scope and caches
# axm's local state: the task scope and caches, except the scopes eval cases start with
.axm/
!*/*/evals/*/*/repo/.axm/

# OS files
.DS_Store
Expand Down
4 changes: 2 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,12 +27,12 @@ The tool is only worth trusting if these hold. Never break them, not even in deb

- **Local first.** No accounts, telemetry, hosted services, update checks or uploads. The only network traffic is model calls made by `axm eval`, `axm triggers`, `axm fossil` and `axm conflicts --judge`, and only when the user runs them.
- **A deterministic core.** `Axiomarium.Core` makes no model or network calls. A test fails if it references `System.Console`, Spectre.Console or `System.Net.Http`.
- **Writes only on request.** `list`, `validate`, `doctor`, `hook`, `detect`, `explain`, `conflicts` and `triggers` (overlap, `test` and `export`) are read-only. `init`, `sync`, `incident new` and `triggers generate` show what they will write and wait for approval in the same run, and refuse before any model call when there's no way to ask. `fossil` and `distill` recommend changes and never delete or rewrite instructions.
- **Writes only on request.** `list`, `validate`, `doctor`, `hook`, `detect`, `explain`, `conflicts` and `triggers` (overlap, `test` and `export`) are read-only. `init`, `sync`, `incident new` and `triggers generate` show what they will write and wait for approval in the same run, and refuse before any model call when there's no way to ask. `eval run` and `eval compare` write only their history, in `.axm/evals/`, and run sessions in a sealed home, never in the user's own. `fossil` and `distill` recommend changes and never delete or rewrite instructions.
- **Label model output.** Anything a model produced, such as a judged contradiction or a distilled root cause, says so in the output.
- **Every explain entry cites its rule.** Each loaded or dropped file carries the loading rule that produced it. If you can't name the rule, don't emit the entry.
- **Harness models follow the docs, then the real harness.** When you change a model, cite the doc section and update the docs date and harness version recorded in the model. If the docs and the real harness disagree, the real harness wins, and the disagreement goes into `ROADMAP.md` as a decision.
- **Don't guess what the model will read.** Skills are "available", never "loaded". When a harness leaves loading to the model, such as a Codex AGENTS.md below the launch directory, report it as "not loaded by the harness".
- **Tests never read the real machine.** Home, `CODEX_HOME`, Codex's system folder, managed-policy and settings locations are injected. Fixtures and docs use placeholder paths such as `/home/dev` and `C:\Users\dev`, never real ones. `CliRun` refuses `doctor`, `explain` and `triggers` without an injected machine, and `triggers generate` and `triggers test` without an injected runner, so no test runs a model. The native binary tests only run commands that read no harness files.
- **Tests never read the real machine.** Home, `CODEX_HOME`, Codex's system folder, managed-policy and settings locations are injected, and so is the operating system `axm eval` plans for, which `CliRun` pins to Linux. Fixtures and docs use placeholder paths such as `/home/dev` and `C:\Users\dev`, never real ones. `CliRun` refuses `doctor`, `explain`, `triggers`, `eval run`, `eval compare` and `conflicts` without an injected machine, and `triggers generate`, `triggers test`, `eval run`, `eval compare` and `conflicts` without an injected runner, so no test runs a model. The native binary tests only run commands that read no harness files.
- **Scenarios are ground truth.** Each one in `scenarios/<name>/` is a tiny repo, a fake home, `scenario.yaml` and the recording in `expected.json`. Every Markdown file in it starts with its `MARKER <path>` line (after any frontmatter), so a recording can name the file even when a harness cuts it short. A skill's description starts with the same marker, because a listing shows descriptions, and a hook's command is `echo MARKER <path> <label>`. Only the recorder writes `expected.json`, never a person.

## Vault assets
Expand Down
22 changes: 21 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,25 @@ This file records the user-visible changes to `axm` and the vault's schemas. The

## [Unreleased]

## [0.5.0] - 2026-09-30

Evals: what an asset does once it's active, measured on the real Claude Code and Codex in a sealed home, and whether a change made it better.

### Added

- `schemas/eval.schema.json`: an asset's eval cases, each a folder under `evals/behavioral/` or `evals/regression/` with an `eval.yaml` (the prompt, the commands the session may run, the checks and an optional judge's rubric) and a `repo/` for the files the session starts with. A check is one of `file`, `run`, `loaded`, `ran` or `reply`, and `not: true` turns it around. A regression case names the failure it guards in `guards`.
- `axm eval run [<asset>...]` runs each asset's eval cases on Claude Code and Codex, each session in a sealed home with only your logins and chosen model, and in its own git copy of the case's repo with the asset installed. It reports how many runs passed, each failed check, and the median and range of tokens, time and tool calls, plus turns and cost on Claude Code. `--json` prints it as JSON, shape 1, and each asset's part is saved in `.axm/evals/`. On Windows, one unscored Codex session starts Codex's sandbox first. Each Codex command is timed from Codex's own records, with each case's slowest named, and every path outside the copy that a run's commands touched is named, since Codex on Windows can read the whole disk.
- `axm eval compare <asset>` runs an asset's cases on a baseline, its files at a git ref or `--baseline none`, and on the working tree, with sessions alternating, and shows each case's passed runs and measures side by side with the change. It reuses a saved baseline that still matches, unless `--fresh`, and saves each version's runs in `.axm/evals/`. `--json` prints it as JSON, shape 1.
- An eval case's `judge.rubric` is graded by a model for each finished run, from the prompt, the commands, the copy's git diff and the final message, in the sealed home and through Claude Code unless `axm eval run --judge-with codex` says otherwise. Its verdicts are shown as model judgment, apart from the checks, which alone decide a run.
- `axm conflicts <path> --judge` has a model find instructions a file gets that contradict each other, through Claude Code or, with `--judge-with codex`, Codex. It keeps only the contradictions whose two passages the model quotes from the files, each with its `file:line`, drops and counts the rest, and labels everything as model judgment.
- The `agent-asset-authoring` skill, 0.2.0, tells the user which tool should enforce a rule, such as a pre-commit hook or a CI check, before writing anything. On Codex with gpt-6-sol, the judge passed its `prefers-tooling` case in 2 of 3 runs, where 0.1.0 passed in none.
- The four starter assets have behavioral evals, and set `evals.behavioral: true`: three cases for the `agent-asset-authoring` skill (a new hook, a new skill, and a rule tooling should enforce), one each for the `scope-sheriff` and `session-doctor` hooks, and one for the `determinism-auditor` agent.
- `axm validate` and `axm doctor` check every eval case: its schema, that each check is exactly one kind and takes only its own fields, that each `file` glob is valid, and that the case is a kebab-case folder with its `eval.yaml`.

### Changed

- Files directly in an asset's `evals/behavioral/` or `evals/regression/` are errors: each case is now a folder with an `eval.yaml`.

## [0.4.0] - 2026-09-29

Trigger testing: whether the agent picks the right skill for a prompt, measured on the real Claude Code and Codex.
Expand Down Expand Up @@ -74,7 +93,8 @@ The asset model: agent configuration can be inspected and validated like softwar
- the `deterministic-boundaries` policy: nine rules for which decisions belong to code instead of the model, eight of them enforced by the determinism auditor;
- the `prompt-fossil` experiment: the method for finding instructions that cost tokens but no longer change behavior. Designed, not run yet.

[Unreleased]: https://github.com/hazeliscoding/axiomarium/compare/v0.4.0...HEAD
[Unreleased]: https://github.com/hazeliscoding/axiomarium/compare/v0.5.0...HEAD
[0.5.0]: https://github.com/hazeliscoding/axiomarium/releases/tag/v0.5.0
[0.4.0]: https://github.com/hazeliscoding/axiomarium/releases/tag/v0.4.0
[0.3.0]: https://github.com/hazeliscoding/axiomarium/releases/tag/v0.3.0
[0.2.0]: https://github.com/hazeliscoding/axiomarium/releases/tag/v0.2.0
Expand Down
2 changes: 1 addition & 1 deletion Directory.Build.props
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
<ImplicitUsings>enable</ImplicitUsings>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
<InvariantGlobalization>true</InvariantGlobalization>
<Version>0.5.0-dev</Version>
<Version>0.5.0</Version>
<IncludeSourceRevisionInInformationalVersion>false</IncludeSourceRevisionInInformationalVersion>
</PropertyGroup>
</Project>
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@

Axiomarium is my lab for engineering reliable AI coding environments. It holds the agents, skills, hooks, policies, workflows and evals I use, and `axm`, a local CLI that inspects, validates, tests and debugs them. It treats agent configuration as real software infrastructure: kept in git, inspectable, testable, portable, and able to learn from its failures. It's built for my own setup first, and it's public in case it's useful to you too.

> **Status:** early development. `axm explain` shows what Claude Code and Codex actually load for a file, and since v0.3 which skills each lists for the model and which hooks run. `axm doctor` checks the instruction files, skills and hooks of any repo, and `axm triggers` finds skills whose descriptions overlap. Since v0.4, `axm triggers test` runs a vault skill's trigger prompts on both harnesses and reports whether the agent picks it. v0.1's vault checks stay: `axm list`, `axm validate` and `axm doctor` in a vault. Evals follow in v0.5. See [ROADMAP.md](ROADMAP.md).
> **Status:** early development. `axm explain` shows what Claude Code and Codex actually load for a file, and since v0.3 which skills each lists for the model and which hooks run. `axm doctor` checks the instruction files, skills and hooks of any repo, and `axm triggers` finds skills whose descriptions overlap. Since v0.4, `axm triggers test` runs a vault skill's trigger prompts on both harnesses and reports whether the agent picks it. Since v0.5, `axm eval run` runs an asset's eval cases in a sealed home, `axm eval compare` shows a change to the asset side by side with its last commit, and `axm conflicts --judge` has a model find instructions that contradict each other. v0.1's vault checks stay: `axm list`, `axm validate` and `axm doctor` in a vault. Evidence freshness follows in v0.6. See [ROADMAP.md](ROADMAP.md).

## The problem

Expand Down Expand Up @@ -156,7 +156,7 @@ axm explain src/app/main.cs what Claude Code and Codex load for that file, wh
axm triggers skills whose descriptions overlap, and the words they share
```

In a vault such as this repo, `axm list` and `axm validate` check the assets too, and `axm doctor` adds them to its report. `axm triggers generate <skill>` writes a skill's trigger prompts after you approve them, and `axm triggers test` runs them on Claude Code and Codex, on your own login. The [docs](docs/README.md) have the full command reference and the write-ups.
In a vault such as this repo, `axm list` and `axm validate` check the assets too, and `axm doctor` adds them to its report. `axm triggers generate <skill>` writes a skill's trigger prompts after you approve them, and `axm triggers test` runs them on Claude Code and Codex, on your own login. `axm eval run <asset>` runs an asset's eval cases and `axm eval compare <asset>` compares it with its last commit, each session in a sealed home that borrows only your logins and model. The [docs](docs/README.md) have the full command reference and the write-ups.

A few platform notes:

Expand Down
Loading
Loading