Skip to content

M5: v0.5 evals - #13

Merged
hazeliscoding merged 43 commits into
mainfrom
m5-evals
Sep 30, 2026
Merged

hazeliscoding merged 43 commits into
mainfrom
m5-evals

Conversation

@hazeliscoding

Copy link
Copy Markdown
Owner

What an asset does once it's active, measured on the real Claude Code and Codex in a sealed home, and whether a change made it better. The plan and its decisions are in ROADMAP.md under M5 and "Evals (v0.5)".

  • axm eval run: each asset's eval cases, each session in a sealed home with only the user's logins and chosen model, in its own git copy of the case's repo with the asset installed. Checks decide each run; an optional rubric is graded by a model as model judgment, apart from the checks.
  • axm eval compare: a baseline at a git ref, or no asset, against the working tree, with sessions alternating, a side-by-side report, and a saved baseline reused while it still matches.
  • axm conflicts <path> --judge: contradictions between the instructions a file gets, kept only when the model quotes both passages from the files.
  • Behavioral evals for the four starter assets, and agent-asset-authoring 0.2.0, whose change the first compare measured.
  • Codex on Windows reads the whole disk and starts its sandbox slowly; the report names paths a run touched outside its copy and each case's slowest call, and docs/cli.md states the limits.
  • Write-up: docs/did-the-change-make-the-skill-better.md.

@hazeliscoding
hazeliscoding merged commit 8def58d into main Sep 30, 2026
13 checks passed
@hazeliscoding
hazeliscoding deleted the m5-evals branch September 30, 2026 05:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant