Repeatable regression checks for stateful MCP tools. A tool can return valid JSON and still apply a change twice, overwrite a newer revision, or mutate state before approval. MCP Tool Check runs explicit fixtures over the real Streamable HTTP protocol and checks those behaviors before you connect an assistant.
This is an original TypeScript CLI and library built on the official @modelcontextprotocol/sdk. It initializes the connection, negotiates the protocol version, discovers tools, invokes only your allowlisted fixtures, and exits unsuccessfully if expectations fail. It requires no model, cloud account, payment information, or API key for the included local demo.
Use Node.js 22 or later. From this repository:
npm ci
npm run check
npm run demo:serverIn another terminal, in the same repository:
npm run demo:check
node dist/cli.js --config examples/demo.json --json report.jsonThe demo server listens only on 127.0.0.1:5191. All data is synthetic and held in memory. Every MCP session starts with a fresh household task list. Stop it with Ctrl-C.
The eight-step fixture verifies that:
- The initial revision is zero with no tasks.
- Creating a proposal leaves state unchanged.
- Committing the proposal advances the revision once.
- Repeating the same idempotency key does not add a duplicate task.
- A proposal based on an old revision is rejected with
isError: true. - The final task list contains exactly the approved task.
Observed output on the local demo:
PASS Household revision and idempotency regression
Endpoint: http://127.0.0.1:5191
Protocol: 2025-11-25 | Tools discovered: 3
...
8 passed, 0 failed, 0 skipped
examples/relay.json is a second, application-level fixture for the Relay Household planning assistant. Start Relay's MCP server using that project's instructions, then run from this repository:
node dist/cli.js --config examples/relay.json --json report.jsonThe fixture targets http://127.0.0.1:5190/mcp and allowlists only get_household, propose_change, commit_plan, and undo_last_plan. It uses Relay's synthetic household session to exercise a feasible replan, an unchanged revision after proposal, rejection without approval, an approved commit, duplicate-key replay, rejection of a stale proposal, an infeasible proposal and rejected commit, and undo to a new revision. Each new session starts from Relay's demo state; no household account or real calendar is involved.
The recorded run passed 12 steps, with 0 failures and 0 skipped steps, in 320 ms, negotiating protocol 2025-11-25 with relay-household version 0.1.0. See the machine-readable result and verification notes. Assertions are explicit in the fixture: the rejection steps check isError, while the final check verifies revision 2; the fixture does not claim to compare every restored household field.
Relay supplies a practical integration example. MCP Tool Check also runs independently with its own included demo and accepts other explicitly configured MCP servers without importing Relay code.
There is no default endpoint, endpoint discovery, or default tool invocation. Supply one exact MCP endpoint and every permitted tool by name:
{
"name": "Approval applies once",
"endpoint": "http://127.0.0.1:5191/mcp",
"allowTools": ["propose_task", "commit_task"],
"timeoutMs": 5000,
"concurrency": 1,
"scenarios": [{
"name": "One approved task",
"steps": [
{
"id": "proposal",
"tool": "propose_task",
"arguments": {"text": "Collect the library books"},
"expect": {"structuredIncludes": {"revision": 0}}
},
{
"id": "commit",
"tool": "commit_task",
"arguments": {
"proposalId": {"$ref": "proposal#/structuredContent/proposalId"},
"expectedRevision": 0,
"idempotencyKey": "fixture-approval-1"
},
"expect": {"structuredEquals": {"revision": 1, "replayed": false}}
}
]
}]
}Steps within a scenario always execute in order. Scenarios share one MCP session; use concurrency greater than one only for scenarios whose state changes are independent. The default is one and the maximum is 16. A failure skips the remaining steps in that scenario; other scenarios continue.
All fixture tools must appear in allowTools and in tools/list before the first invocation. A missing tool fails the suite. Tool discovery follows pagination, rejecting repeated cursors and limiting discovery to 100 pages.
| Field | Meaning |
|---|---|
expect.isError |
Defaults to false. Set true when intentionally testing a tool error. |
expect.structuredEquals |
Deep equality against structuredContent, including exact object keys and array order/length. |
expect.structuredIncludes |
Recursive object subset against structuredContent. Extra object keys are allowed. Arrays still require equal length/order. |
{"$ref":"stepId#/structuredContent/id"} |
Copy a value from an earlier successful step in the same scenario into tool arguments. |
References use JSON Pointer escaping (~1 for /, ~0 for ~) and access only own properties. They cannot evaluate code, access environment variables, or reference a future/different scenario. A marker is recognized only when $ref is its sole property. Stored values remain in memory and are omitted from reports.
Assertions deliberately operate on structured MCP output. Text that happens to contain JSON is not parsed implicitly. If your tool returns only text, an isError assertion still works; structured assertions require structuredContent.
# Human report on stdout; JSON report to a file
node dist/cli.js --config fixtures.json --json report.json
# JSON alone on stdout; human report on stderr
node dist/cli.js --config fixtures.json --json -| Exit code | Meaning |
|---|---|
0 |
All configured checks passed. |
1 |
A connection, discovery, tool call, or assertion failed. |
2 |
Invalid CLI options, configuration, missing credentials, or a file error. |
JSON reports have schemaVersion: 1, negotiated protocol/server information, discovered tool names, per-step statuses and durations, assertion paths, and passed/failed/skipped counts. They omit tool arguments, result values, headers, URL paths, and raw server errors. Known supplied credential values are additionally redacted. Human reports omit control sequences. Treat scenario/tool labels and object field names as report-visible metadata.
The included GitHub Actions workflow runs the checks on Node.js 22 and 24. JSON reports can be archived as CI artifacts by your own pipeline.
Keep credentials outside fixtures. Map header names to existing environment-variable names:
{
"headersFromEnv": {
"Authorization": "MCP_AUTHORIZATION",
"X-API-Key": "MCP_API_KEY"
}
}Add this field to your full configuration. MCP_AUTHORIZATION should contain the complete header value, such as a bearer authorization supplied by your secret manager. Neither the names nor values are printed. Missing/invalid variables fail before any network request.
Remote endpoints require HTTPS. Plain HTTP is accepted only for the explicit loopback names localhost, 127.0.0.1, and [::1]. URL credentials, queries, fragments, and redirects are rejected. Protocol-owned headers cannot be overridden. This version supports credentials supplied as headers; it does not perform OAuth registration or interactive login.
Use a disposable test account or synthetic server for mutating fixtures. The allowlist controls which tools this client invokes; it cannot undo a server-side effect. Calls are never retried. A timeout may occur after a server has applied a change, so later dependent steps are skipped. timeoutMs bounds each protocol request and fetch, not the entire suite. At completion the client attempts to terminate its own MCP session and closes its transport.
After building this checkout:
import { readFile } from 'node:fs/promises';
import { runSuite, formatReport } from './dist/index.js';
const config = JSON.parse(await readFile('fixtures.json', 'utf8'));
const report = await runSuite(config);
console.log(formatReport(report));
process.exitCode = report.passed ? 0 : 1;Exports include runSuite, formatReport, parseConfig, configSchema, compareJson, and their TypeScript types. runSuite throws for invalid configuration or unavailable required environment variables; connection and check failures are represented in the returned report. This repository is usable from source; no npm publication is claimed.
This checks your declared behaviors. It is not a complete MCP conformance suite, security audit, property-based generator, or proof that an assistant will choose the right tool. It does not crawl hosts, synthesize tool calls, use an LLM, execute code from tool responses, or run cleanup mutations beyond session termination.
See VERIFICATION.md for actual local results, CONTRIBUTING.md for development, and the standalone project pitch. SDK APIs were checked against the official v1 client guide and the installed, locked SDK package. Version 1.30.0 negotiated 2025-11-25 in the included test; this is an observed result, not a hardcoded protocol assumption.
MIT licensed. Copyright codeforcode111.