docs: re-baseline the MCP context claims against tool search; scope /infinite - #98
Conversation
|
VERDICT: PASS Clean docs-only PR. The CHANGELOG |
62db375 to
6de8825
Compare
c0e614a to
d0a2ae8
Compare
|
VERDICT: PASS Clean docs correction. All changed files are documentation, template content, or template comments — no code logic, no fixture tests, no persona snapshots. The stale MCP token-cost claim is removed consistently across all affected locations (README, docs/04, docs/06, check-context SKILL.md in both the template and the example, claude-ctx.sh, mcp.minimal.json, servers-cookbook.md), and the addendum pattern used in the experiments-memory file is the correct approach for an immutable experiment log. The /infinite decision table is a sensible scope guard. CHANGELOG ## Unreleased entry is present. No third-party code, no AGPL issue, templates/discipline-skills/ untouched, no schema additions. |
6de8825 to
ada0e68
Compare
…infinite
Two headline claims had been overtaken by Claude Code and were overstating what
the modules buy.
MCP. README claimed per-task profiles "drop a bloated 4-MCP baseline from ~49%
context to under 5%", and docs/04 asserted that "every MCP tool is a chunk of
JSON schema loaded at session start". Tool search defers MCP schemas by default
(alwaysLoad: true is the opt-OUT), so the premise no longer holds.
Measured rather than re-guessed. Four local stdio servers advertising twelve
tools each (48 total), against an otherwise identical one-turn session on CC
2.1.245:
no MCP servers 26,665 prompt tokens
4 servers, deferred 27,361 (+696, ~14 tokens/tool)
4 servers, alwaysLoad 40,993 (+14,328, ~298 tokens/tool)
Deferral removes ~95% of the schema cost, and the numbers reproduced exactly
across runs.
The stale claim was also embedded in four SHIPPED templates, which is worse than
in the docs because it lands in every user's project: check-context/SKILL.md
(its budget guardrails and the "MCP > 10%" flag), claude-ctx.sh's rationale
comment, servers-cookbook.md (which already explained deferral correctly a few
sections earlier, so it contradicted itself), and the mcp.minimal.json profile
comment. All corrected.
Profiles are now documented for what they still genuinely buy -- which servers
connect: startup time, auth prompts, cold start, and the blast radius
--strict-mcp-config enforces. docs/06 picks up the same correction. The dated
experiments-memory example keeps its original result with a superseding
addendum rather than a rewrite, because an experiment log records what was true
when it ran; that is also a better demonstration of the format.
/infinite. Dynamic workflows now do staged, resumable, budgeted fan-out with
structured output between stages, and hand-rolled wave batching is the weaker
instrument for that job. The skill opens with a decision table routing staged,
merge-heavy or resumable work to a workflow, and keeps the one case it is
genuinely good at: N variants of a single spec into disjoint slots with no
cross-iteration coordination. README's module row and a new docs/04 section say
the same.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JVndNviHZSnbKJnWP7jFbV
d0a2ae8 to
a2b6c08
Compare
|
VERDICT: PASS Documentation-only corrections: MCP context numbers re-measured against CC 2.1.245 (the old "~49% of 100k window" claim predated tool-search deferral), and |
|
VERDICT: COMMENT-ONLY BlockingNone. AdvisoryScope (borderline): The PR bundles two distinct corrections — the MCP token-cost re-baselining and the
Both are advisory only; the docs corrections are accurate, the CHANGELOG entry is present and detailed, no code paths or persona snapshots changed, no AGPL or MIT subtree concerns. |
What & why
Split out of #94 at the reviewer's request. Two of the project's own claims had been overtaken by Claude Code and were overstating what the modules buy. Overclaiming is the real risk here: a user who measures the headline number and finds it wrong stops trusting the rest of the docs.
1. The MCP context claim was factually wrong. README claimed per-task profiles "drop a bloated 4-MCP baseline from ~49% context to under 5%", and
docs/04asserted "every MCP tool is a chunk of JSON schema loaded at session start". Tool search defers MCP schemas by default (alwaysLoad: trueis the opt-out), so the premise no longer holds.Measured rather than re-guessed — four local stdio servers advertising twelve tools each (48 total), identical one-turn session on CC 2.1.245:
alwaysLoad: trueDeferral removes ~95% of the schema cost; the numbers reproduced exactly across runs. The stale claim was also embedded in four shipped templates — worse than in the docs, because it lands in every user's project:
check-context/SKILL.md(its budget guardrails actively taught the model a wrong premise),claude-ctx.sh's rationale comment,servers-cookbook.md(which already explained deferral correctly a few sections earlier, so it contradicted itself), and themcp.minimal.jsoncomment. All corrected. Profiles are now documented for what they genuinely buy: which servers connect — startup time, auth prompts, cold start, and the blast radius--strict-mcp-configenforces.The dated
experiments-memoryexample keeps its original result with a superseding addendum rather than a rewrite — an experiment log records what was true when it ran, and that's also a better demonstration of the format.2.
/infinitescoped against dynamic workflows. Workflows now do staged, resumable, budgeted fan-out with structured output between stages. The skill opens with a decision table routing that work to a workflow and keeps the one case it's genuinely good at: N variants of one spec into disjoint slots.Type of change
Scope
mcp,multi-agent,token-efficiency,experiments-memoryTests
python3 configure.py --checkpasses locally — verified on this commit alone.CHANGELOG
## Unreleased.License & NOTICE
templates/discipline-skills/untouched by this PR.LICENSE/NOTICEuntouched.Signing
I understand
CONTRIBUTING.md.Fourth of seven stacked PRs — based on
fix/currency-corrections(#94).