Skip to content

Repository files navigation

Metis

Metis is a session-audit companion for Claude Code. It reads Claude Code's own run transcripts and reports whether your agent setup actually used the tools and workflows it was told to, then proposes evidence-based optimizations. Named for the Titaness of wise counsel, it advises; you decide.

It is not a telemetry pipeline or a dashboard. Think of it as a counsel you re-run. Each time you invoke /metis, it takes a fresh census of what you have installed, audits the transcripts that have accumulated since last time, and, on a Phanes-managed project, scores its own past suggestions before offering any new ones. The ground truth is always the transcript, never the console, which hides subagent internals and collapses tool-call detail to a single line.

No dependencies, pure Node. Everything deterministic is plain Node (18 or newer), streaming, with Windows-safe paths and no npm install. The only model-token cost is the short final report; the parsing and diffing are free.

Standalone and Phanes-aware. Metis works on its own in any Claude Code project, with no Phanes install at all. When it detects a .phanes/ directory it switches into an integrated mode with an optimization ledger and bounded autonomy, and Phanes in turn detects the /metis command and calls Metis during its update runs. Neither tool depends on the other; each degrades gracefully to solo operation.

Which mode are you in?

Your situation What Metis does
Any Claude Code project, no Phanes Standalone. Nothing is assumed. It builds the capability list purely from detection, audits your transcripts, and proposes optimizations as a per-item checklist. Every change is applied only on your yes.
A Phanes-managed project (.phanes/ present) Phanes-integrated. The Phanes standard tools are recognized, an optimization ledger records every change, and each run verifies its past suggestions before proposing new ones. A narrow autonomous whitelist may act and log; anything structural still asks first.

Contents


What it does

Metis installs like Phanes: you copy one file, metis.md, as your /metis command, and it bootstraps its own engine on the first run. It keeps a per-project run counter, so the first run installs the engine and runs an optimization pass, and every later run does an update check and another optimization pass. The engine itself is deterministic Node; you interpret its output.

Every optimization pass walks the same path:

1. Detect and gate. Metis works out whether the project is standalone or Phanes-managed, whether a transcript directory exists, and whether there is anything to optimize at all. A setup with no MCP servers, no plugins, no skills, and no agents has nothing to build policy around, so Metis stops and says so rather than emitting an empty report. It also self-checks this repository for a newer release; when one has shipped it asks whether to upgrade, and on yes runs /metisupgrade, which completely replaces the command and engine with the latest while preserving your audit state.

2. Census and consent. It enumerates every installed capability, the MCP servers, plugins, skills, slash commands, and foreign agents, and probes each MCP server to see whether it is actually reachable and authenticated rather than merely configured. It then proposes a per-item selection. On a Phanes project the standard set (context7, deepwiki, serena, semble, frontend-design) is pre-selected and marked recommended; on a standalone project nothing is assumed and every item is listed unchecked by its detected name. Your selection persists, so the next run stays silent unless the set actually changes, and then it asks only about what changed. A server that is switched off or signed out can never be mandated, even if selected.

MCP servers are registered across six surfaces, and no single one of them is complete: the project .mcp.json, ~/.claude/.mcp.json, ~/.claude.json at user scope and again per project (whose key is written with forward slashes on Windows while the project is addressed with backslashes), the live claude mcp list, and the claude.ai connection history. Plugin-provided and claude.ai servers appear in no config file at all and are visible only to the live probe, which is also cwd-sensitive, so Metis runs it from the audited project. Reading only some of these is how a v0.4 census found 1 server on a machine registering 18, and reported it as though the environment were bare.

3. Audit the transcripts. It streams the main-session JSONL, the durable subagent transcripts under <session>/subagents/ (the current Claude Code layout), and the volatile subagent task store (harvesting it first, because the harness discards it). It aggregates, per session and per agent: tool use by name (MCP split into server and tool), models, tokens, spawns, and errors. Then it diffs actual usage against the policy: a tool that was mandated but never called, a server that was configured but never used, an agent that was never spawned. It runs a redacted secret scan over commands and tool results, and summarizes where the tokens went. A subagent's internal MCP use is invisible in the parent session; recovering it from the subagent transcripts is the whole point.

--last N scopes both halves of the corpus, and the report states the scope. Scoping only the main sessions, as v0.4 did, let a 12-session window quietly roll up 606 subagent transcripts spanning two roster generations, which makes every per-role figure uninterpretable as evidence about the window that was asked for. --all-subagents widens it deliberately and says so in the header. Roles are resolved from the subagent transcripts as well as the main ones, because an engaged orchestrator dispatches its workers from inside its own transcript: reading only main sessions meant the better a project followed the delegation rule, the more of its behaviour went missing from the tables.

Beyond that binary check, Metis runs condition-aware checks whose absence is only a finding when the precondition was present, so a tool that was correctly not needed stays silent. On a Phanes project it reads the roster and thresholds to ask whether the orchestrator was engaged for a plan at or above its step threshold, and, on a pre-3.4 project, whether the effort bridge delivered an above-baseline archetype's rubric. What ran is read from the transcript; what should have run is read from the Phanes artifacts, at no model-token cost. Every such finding is advisory and carries its precondition and evidence.

Two rules govern what these checks are allowed to say, and both exist because v0.4 broke them:

  • An unreadable surface is not an empty one. Any capability count whose surface could not be read is reported as null, never 0, with the surface named. A broken sensor must never be indistinguishable from a bare environment.
  • A check whose evidence is unavailable stays silent. When a plan is correctly delegated, the primary session's observable scale is deliberately compressed and its worker spawns move into the orchestrator's transcript by design. If those transcripts were not ingested, Metis reports the scale as not measurable rather than inferring smallness from its own blindness.

Metis also cross-checks itself: adherence.selfCheck re-derives the never-called list against the observed-servers table and reports FAIL if the two disagree. That contradiction is what exposed the v0.4 defect class, so it is a deliberate check rather than something a fix tidied away.

4. Verify, then propose (Phanes mode). Before proposing anything new, Metis scores its open ledger entries against the new sessions: delivered, not-yet-measurable, or regressed. A regression produces a rollback proposal. Only then does it propose new changes, and only from strong signals, with a cooldown so it never oscillates on the same knob. Each proposal is gated. Trigger lines, annotations, and flags may be applied and logged; merging or removing agents, removing mandates, and single-writer reassignments always ask first.

Because that whitelist is the one path allowed to act without asking, it re-derives its own evidence before emitting anything. Proposals are keyed by logical capability, so a server registered several ways yields at most one, and a capability with observed calls under any of its names can never produce a zero-uses proposal. Every refusal is reported with the calls and the names it matched them under. The whitelist is safe because it is narrow and evidence-backed; it stops being safe the moment the evidence is wrong, which is exactly what v0.4 demonstrated when it offered to weaken the trigger lines of the three most-used capabilities in a roster.

5. The checklist. Findings and proposals are presented as a per-item checklist, most-evidenced first, each with its evidence, its proposed change, and its gate. Nothing outside the autonomous whitelist is applied without your yes. Quality signals such as error and retry rates are proxies and are marked as proxies; token spend is the only thing measured directly.

How to use

Most users only ever type /metis. That command is the front end: on the first run it asks the install scope and fetches its own engine, and on every run it asks the consent question through the harness, runs the audit, and presents the checklist. It never runs unbidden.

For direct or scripted use, call the engine dispatcher at <engineDir>/metis.mjs, where <engineDir> is the folder the first run created (~/.claude/metis/ for a global install, <project>/metis/ for a per-project one).

Command What it does
metis.mjs detect --project . Mode, companion presence, the update check, and whether there is anything to optimize (an early stop when not).
metis.mjs census --project . Detect capabilities, propose a per-item selection, diff against the prior manifest. --set-selection a,b persists the consented selection.
metis.mjs audit --project . --harvest <dir> --out <dir> --last N Run the transcript audit. --dir is derived from --project; --harvest preserves the volatile task store first. --last scopes main sessions and subagents alike; --all-subagents widens the subagent side and says so in the report.
metis.mjs ledger --project . --audit <report.json> --verify Verification-first scoring of the ledger (Phanes mode). --propose lists strong-signal proposals.
metis.mjs version --check Print the version and check the repository for a newer release.

Each piece also runs on its own from the source tree (node src/census.mjs, node src/session-audit.mjs, node src/ledger.mjs, node src/policy.mjs, node src/phanes-context.mjs, node src/version.mjs), and node --test runs the suite. src/names.mjs is the shared canonicalization layer: one MCP server is written three ways on one machine (plugin:neon:neon in the registration, plugin_neon_neon in the transcript, neon in a grant), and every comparison across those forms goes through it.

On a Phanes project the capabilities.selection[] array belongs to Phanes. Metis reads it, matches against it, and writes back only what Phanes' own census-diff can see, preserving every field Phanes put there. Registrations Metis finds on surfaces Phanes does not read are reported as a blind spot, never written. See DECISIONS.md.

Policy derivation

Metis never hardcodes an external tool name. It ships knowledge of exactly one named set, the Phanes standard tools, and recognizes it only on a Phanes project. Every other capability is discovered and reasoned about generically from three evidence sources: its own tool names and descriptions, transcript evidence of actual use, and domain matching against the agent roster. On a Phanes project the policy is the consented, reachable selection; standalone, every reachable detected server is treated as configured, and a never-called one is an advisory finding rather than a hard mandate.

Where the transcripts live

Claude Code stores main transcripts under ~/.claude/projects/<encoded-project-path>/ and volatile subagent transcripts under the Temp task store. The path is encoded by replacing : and the path separators with -, so C:\Projects\YourProject becomes C--Projects-YourProject. Metis derives this from --project for you. Reports contain aggregates only; raw transcript content is never emitted, and secret findings are always redacted.

Core principles

  • The transcript is ground truth. Not the console, which hides subagent internals and collapses tool-call detail. Metis reads the JSONL the harness actually wrote.
  • Advisory-first. Metis proposes; you dispose. The only self-applied changes are a narrow autonomous whitelist inside a Phanes update run, and each one is logged to the ledger.
  • One named set, everything else discovered. The code knows exactly one profile, the Phanes standard tools, and only on a Phanes project. No external tool ever gets a hardcoded playbook; usage rules are derived generically.
  • Verify before proposing. In Phanes mode, every run scores its past suggestions against the new sessions before offering new ones. No stacking new optimizations on unverified ones.
  • Strong signals only, with a cooldown. Metis acts on a capability that is reachable but unused across several sessions, and it never touches the same knob twice inside a window, so it cannot oscillate.
  • Consent once per project. The user picks, per item, which capabilities Metis may build policy around. The choice persists and is only revisited when the environment actually changes.
  • Redacted always, aggregates only. Secret findings show location, pattern type, and a masked value; reports never contain raw transcript content.
  • Quality is proxied, never claimed. Token spend is measured. Output quality is only approximated, through error and retry rates, and those are labeled as proxies.
  • An unreadable surface is not an empty one. A count Metis could not take is reported as null with the surface named, never as 0, and never triggers the early stop. A broken sensor must not be indistinguishable from a bare environment, because the two call for opposite responses.
  • A check with no evidence stays silent. Where a check cannot observe what it is judging, it says the answer is not measurable instead of inferring one. Reporting a gap is a useful output; guessing into it is not.
  • The report cross-checks itself. Two sensors that read the same transcripts must agree. Where they do not, the report says so, names it as a Metis defect rather than a finding about your project, and tells you not to act on the affected list.
  • Never damage what another tool owns. On a Phanes project the capability array belongs to Phanes. Metis merges rather than rebuilds, preserves every field it did not write, and reports what it cannot safely record instead of forcing it in.

How to install

Metis is one file, exactly like Phanes. You install the metis.md command, run /metis once, and it fetches its own engine on the first run. It needs Claude Code and Node 18 or newer; there is nothing to clone and no dependencies to install.

Install the command

For all projects (global), put it in your user commands folder:

Linux / macOS:

mkdir -p ~/.claude/commands
curl -L https://raw.githubusercontent.com/Aloim/metis/main/metis.md \
  -o ~/.claude/commands/metis.md

Windows (PowerShell):

New-Item -ItemType Directory -Force "$env:USERPROFILE\.claude\commands" | Out-Null
Invoke-WebRequest `
  -Uri https://raw.githubusercontent.com/Aloim/metis/main/metis.md `
  -OutFile "$env:USERPROFILE\.claude\commands\metis.md"

For a single project, put it in that project's commands folder instead (.claude/commands/metis.md under the project root).

Run it

Open a project in Claude Code and type:

/metis

The first run asks whether to install for all projects or just this one, fetches the engine into its folder, installs the sibling /metisupgrade command alongside, audits, and presents the checklist. Every later run does a version check and audits again; when a newer version is published it offers to upgrade in place with /metisupgrade (a complete command-and-engine replacement that preserves your audit state). Anything after the command is treated as a directive, for example /metis reinstall or /metis verify only.

What the first run creates

  • The engine: ~/.claude/metis/ for a global install, or <project>/metis/ for a per-project one.
  • The per-project state: <project>/.metis/ for a global install (so projects stay separated), or alongside the engine for a per-project install. It holds the run counter, the standalone manifest, and the reports.
  • On a Phanes project, the capability manifest and the optimization ledger live in <project>/.phanes/, per the Phanes contract, whichever install scope you chose.

The command calls the engine as node <metis-dir>/metis.mjs; point it at wherever you cloned the repository. Installing the command is also what lets a Phanes-managed project detect Metis and call it during update runs.


Consent and sequencing

Consent is asked once per project, and never twice:

  • During a Phanes run, Phanes owns the consent gate (its own pre-flight). When a Phanes update run invokes Metis it calls the CLI directly and reads the selection Phanes already wrote; Metis does not re-ask, and it owns only the ledger.
  • The /metis slash command is always user-initiated. On a standalone project, or a Phanes project with no selection yet, it asks the full per-item question. On a Phanes project that already has a selection it reads that and asks only about a delta.

First run. Metis does not audit during a Phanes first setup run: there are no steady-state sessions yet, and the ledger is empty. Phanes records that the companion is present and defers the audit to the first update run, when real sessions exist. On any project whose only sessions are a setup run, Metis says the sample is not steady-state work and defers strong conclusions.


Relationship to Phanes

Metis is a Phanes companion tool, in the same family as Charon. Like every companion, it is a full standalone tool that needs no Phanes install, and it also snaps into the structures Phanes builds the moment it lands in a Phanes-managed project. Phanes detects the /metis command during its capability census; on an update run it has Metis harvest the transcripts, verify its past suggestions against the new sessions, and file an adherence report the run then acts on. Absent Phanes, Metis is simply a standalone auditor.

Where the two tools overlap, Phanes owns the shared record. They both read capabilities.selection[] in .phanes/config.json, and they populate it from different sensors producing different name forms, which is enough for each tool's run to undo the other's. The resolution is that Phanes' display-name taxonomy is authoritative: Metis derives its own canonical keys from it, groups duplicate registrations of one logical server by Phanes' alias field, preserves every field Phanes wrote, and writes back only what Phanes' own census-diff can see. A registration Metis finds on a surface Phanes does not read is reported as a blind spot between the two tools, never written, because writing it would make Phanes' next diff report it as removed from a surface that genuinely lacks it. The reasoning, and the alternatives that were rejected, are in DECISIONS.md.

Metis also reads the Phanes spec version and adapts to it, rather than running checks that can no longer fire. Phanes v3.4 retires the per-agent effort bridge (effort is a single session-wide level, effort_class is gone from every agent file, and the CLI spawn bridge is deleted), so on a v3.4+ project that check reports not-applicable and says why, while an observed bridge spawn there becomes a finding in its own right. Metis recognizes the v3.4 script library, reads v3.4 agent frontmatter, and counts agent-persistence resumes so resume-before-respawn is not misread as fan-out.


Version

Current: v0.5 (2026-08-06), a sensor-correctness release built from a field bug report against a live v0.4 run. Six defects, three of them severity 1, and all three of those reduced to a single root cause: one MCP server is written three different ways on one machine, and v0.4 compared them as raw strings. The registration form (plugin:neon:neon), the transcript form (plugin_neon_neon) and the grant alias (neon) never matched, so servers called hundreds of times measured as zero, a census found 1 server where 18 were registered, and that inverted evidence fed the one proposal path allowed to act without asking. Measured on the reporting machine: MCP detection 1 to 20, proposals 12 (nine false or duplicate) to 5 (all correct), the shared-array diff 21 added / 24 removed to 0 / 0 / 0. names.mjs is the single canonicalization layer that closes the class, and the report now cross-checks itself so a recurrence is announced rather than absorbed.

It builds on v0.4 (condition-aware adherence), v0.3 (one-file self-bootstrapping install), and v0.2 (the first public release of the audit engine, the census and consent contract, genericized policy derivation, and the verification-first ledger). The full history is in Changelog.md, and the reasoning behind the structural choices is in DECISIONS.md. Metis also checks this repository on invocation and tells you when a newer release has shipped. /metisupgrade tracks the same version number from v0.5 on.

Phanes compatibility. Metis reads the Phanes spec version and adapts: it recognizes the v3.4 script library, reads v3.4 agent frontmatter, and gates retired checks on the detected version rather than running them as dead code. Phanes v3.4 retires the per-agent effort bridge, so that check reports not-applicable there and stays live on pre-3.4 projects. None of this is required: Metis runs standalone on any Claude Code project.


License

Metis is released under the Creative Commons Attribution-NonCommercial 4.0 International license (see LICENSE).

You are free to use, share, and adapt Metis for any non-commercial purpose with attribution. Commercial use is not granted by this license. For commercial licensing terms, contact the author directly.


Contributing

Issues and pull requests are welcome. Because Metis is an advisory tool whose value is trust, a substantive change should explain which failure mode it closes and carry a node --test case that proves it.

About

Session-audit companion for Claude Code: reads run transcripts and reports whether your agent setup actually used the tools it was told to, then proposes evidence-based optimizations. Standalone and Phanes-aware. Companion to Phanes.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages