Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 20 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# AgentXRay

**See what your coding agent ran—and what its logs actually verify.**
**Trace your coding agent's execution back to the evidence.**

Read existing session logs locally. Trace failures, background exits and checks after edits back to their source, without an SDK, model call or mandatory human labeling.
Read existing session logs locally. Connect tool calls, background exits and checks after edits to their original records, without instrumentation, model calls or mandatory human labeling.

<p align="center">
<img src="assets/readme/hero.svg" width="100%" alt="AgentXRay execution evidence: a check passes, an edit follows, and the next check is unknown. Conceptual timeline, not a task-success verdict.">
Expand Down Expand Up @@ -49,11 +49,27 @@ Replace the path with your log; `codex` and `claude-code` are also accepted. `np

**Exit 0 means a report was generated, not that the task passed.** [JSON contract, coverage checks and exit policies →](docs/offline-inspect.md)

**Layered CLI:** [summary-first inspection, hash-checked evidence expansion and structured JSON errors](docs/offline-inspect.md#layered-cli-summary-first-evidence-on-demand). The full successful report remains compatible; explicit raw evidence may contain sensitive data. JSON-mode failures now return an error object instead of empty stdout; check `kind`, `error.code` and the exit status.

## Real-session check: background work in a 4.33 MB log

On one previously studied, frozen Codex session (**771 records**), AgentXRay and an independent raw-record parser agreed on **all 36 background-process chains**: **28 recorded successful exits** and **8 last-recorded running states**. The three questions chosen before this investigation produced matching answers, with launch, poll and terminal source lines where available; no uniquely associated nonzero terminal process was found.

This validates associations on **one real snapshot**, not general accuracy, current process status or time saved. The private log is not published, so the case is not publicly reproducible from this repository. [Questions, evidence and retrieval costs →](docs/session-forensics.md)

### Choose the path for your question

- **Get an overview:** `inspect --summary --json` returns counts and sampled references; it is not a complete list of evidence.
- **Find the latest launch, all exits or no matching result:** use `inspect --json`, then examine the full `processes.entries`. Do not infer absence from a short summary.
- **Verify a conclusion at its source:** use `evidence` with the report's source hash and physical line; follow `nextOffset` if the record is paginated.

These commands are available through the installed `agentxray` launcher. [Commands, required arguments and boundaries →](docs/session-forensics.md#existing-cli-workflow)

## See the evidence

![Actual AgentXRay UI on a synthetic session: the modification/check panel shows a successful earlier test, a later edit and no recognized post-edit check.](screenshots/verification-chronology.png)

*Real interface, synthetic data. The expanded panel separates an earlier test from an overlapping check and a later modification. [Open the full-size screenshot](screenshots/verification-chronology.png) or [run the interactive chronology demo](docs/diagnostics.md#modification-and-verification-chronology).*
*Real interface, synthetic demonstration data—not the private Codex case above. The expanded panel separates an earlier test from an overlapping check and a later modification. [Open the full-size screenshot](screenshots/verification-chronology.png) or [run the interactive chronology demo](docs/diagnostics.md#modification-and-verification-chronology).*

### A reproducible report

Expand Down Expand Up @@ -102,7 +118,7 @@ There are **8 historical failures**; **7 remain pending in 2 events**, and **1 h

**Implemented and tested:** local browsing, deterministic execution-evidence rules and the offline report. The UI and CLI share their diagnostic source; tests check source references, conservative matching and temporal counterexamples. [CLI validation](test/inspect.test.js) · [UI and fixture verification](docs/diagnostics-verification.md) · [Recompute published claims](claims.json)

**Not established:** improved real-world agent completion, lower costs or less developer time. In the [initial synthetic pilot](experiments/effectiveness-pilot/RESULTS.md), all three arms passed **12/12** tasks; AgentXRay did not demonstrate an advantage over mechanical context. These experiments live under `experiments/`, are not included in the npm package, and are not default product behavior.
**Real-log evidence:** process associations and source links agree with independent parsing on the single snapshot above. **Not established:** improved agent completion, lower model costs or faster human investigation. Controlled synthetic evaluations have not demonstrated an overall agent-performance advantage. [Initial pilot](experiments/effectiveness-pilot/RESULTS.md) · [Full versus layered CLI](experiments/layered-comparison/RESULTS.md) · [No-tool versus invocation policies](experiments/invocation-policy/RESULTS.md). Experiments are repository-only, excluded from npm and never run by default.

**Not a completion judge:** no root-cause inference, automatic repair, test-coverage proof or live process monitoring. A missing record is missing evidence, not proof that an operation did not happen. If you need instrumented production tracing or a hosted team service, this local log reader is not that product.

Expand Down
24 changes: 20 additions & 4 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# AgentXRay

**看清 coding agent 执行了什么,以及日志究竟验证了什么。**
**看清 coding agent 的执行经过,找到支持结论的原始证据。**

直接读取本机会话日志,把失败、后台退出和修改后的检查追溯到原始证据。无需接入 SDK、调用模型或人工标注。
直接读取本机已有会话日志,把工具调用、后台退出和修改后的检查关联到原始记录。无需埋点、调用模型或人工标注。

<p align="center">
<img src="assets/readme/hero.svg" width="100%" alt="AgentXRay 执行证据:检查通过后又发生修改,下一次检查仍未知。概念时间线,不是任务通过的判定。">
Expand Down Expand Up @@ -49,11 +49,27 @@ npx @alloevil/agent-xray inspect --platform omp /path/to/session.jsonl --json

**退出码 0 表示报告生成成功,不代表任务通过。**[JSON 契约、覆盖检查与退出策略 →](docs/offline-inspect.md)

**分层 CLI:**[先摘要、按哈希展开证据、结构化 JSON 错误](docs/offline-inspect.md#layered-cli-summary-first-evidence-on-demand)。完整成功报告保持兼容;显式展开的原文可能包含敏感信息。JSON 模式失败时现在返回错误对象,不再是空 stdout;请检查 `kind`、`error.code` 和退出码。

## 真实会话核查:4.33 MB 日志中的后台任务

对一份此前研究过、已冻结的 Codex 会话(**771 条记录**),AgentXRay 与独立原始记录解析在 **36 条后台进程链**上得到一致结果:**28 条记录为成功退出**,**8 条最后记录为运行中**。本次调查前选定的三个问题均得到一致答案,有对应记录时可追溯启动、轮询和终止行号;未找到可唯一关联的非零终止进程。

这验证的是**单个真实快照的关联一致性**,不是普遍准确率、进程实时状态或节省时间的证明。私人日志不公开,读者无法仅凭仓库复现该样本。[问题、证据与取证成本 →](docs/session-forensics.md)

### 根据问题选择入口

- **概览会话:**用 `inspect --summary --json` 查看计数和部分证据引用;它不是完整证据列表。
- **查最后一次启动、全部退出或是否存在某类结果:**用 `inspect --json` 查看完整 `processes.entries`,不要从短摘要推断“没有发生”。
- **核实一个结论:**用 `evidence`,传入报告中的原始日志哈希和物理行号;记录被分页时,按 `nextOffset` 继续读取。

这些命令可通过安装后的 `agentxray` 使用。[完整命令、必填参数与边界 →](docs/session-forensics.md#existing-cli-workflow)

## 直接看证据

![AgentXRay 真实界面中的合成会话:修改—检查面板展示先前成功的测试、之后发生的修改,以及缺少可识别后续检查。](screenshots/verification-chronology.png)

*真实界面,合成数据。展开的面板区分先前测试、与修改重叠的检查和修改记录。[查看原尺寸截图](screenshots/verification-chronology.png),或[运行可交互的时序演示](docs/diagnostics.md#修改与验证的先后顺序)。*
*真实界面,合成演示数据,不是上面私人 Codex 案例的截图。展开的面板区分先前测试、与修改重叠的检查和修改记录。[查看原尺寸截图](screenshots/verification-chronology.png),或[运行可交互的时序演示](docs/diagnostics.md#修改与验证的先后顺序)。*

### 可以复现的报告

Expand Down Expand Up @@ -102,7 +118,7 @@ node bin/agentxray.js inspect --platform omp \

**已实现并有测试:**本地浏览、确定性的执行证据规则、离线报告。UI 与 CLI 共用诊断源码;测试覆盖证据行号、保守匹配及先后顺序反例。[CLI 测试](test/inspect.test.js) · [界面与样例验证](docs/diagnostics-verification.md) · [公开数字的复算依据](claims.json)

**尚未证明:**提高真实 Agent 任务完成率、降低费用或节省开发者时间。[首轮合成实验](experiments/effectiveness-pilot/RESULTS.md)中,三组均通过 **12/12** 个任务,未证明 AgentXRay 优于机械摘要。实验代码位于 `experiments/`,不包含在 npm 安装包中,也不是产品默认行为。
**真实日志证据:**上述单个快照中的进程关联和原始行号与独立解析一致。**尚未证明:**提高 Agent 任务完成率、降低模型成本或缩短人工排查时间。受控合成实验尚未证明整体 Agent 性能优势。[首轮实验](experiments/effectiveness-pilot/RESULTS.md) · [完整报告与分层 CLI 对照](experiments/layered-comparison/RESULTS.md) · [无工具与调用策略对照](experiments/invocation-policy/RESULTS.md)。实验仅在源码仓库中,不包含在 npm 包内,也不会默认运行。

**不是完成判官:**不推断根因、不自动修复、不证明测试覆盖,也不实时探测进程。缺少记录只是缺少证据,不能证明某件事没有发生。如果你需要埋点式生产 tracing 或托管团队服务,本机日志查看器不是那类产品。

Expand Down
5 changes: 3 additions & 2 deletions bin/agentxray.js
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
#!/usr/bin/env node
// CLI entry: parse --port/--host, export them, then boot the server.
const argv = process.argv.slice(2);
if (argv[0] === 'inspect') {
void require('./inspect').main(argv.slice(1));
if (['inspect', 'evidence'].includes(argv[0])) {
void require('./inspect').main(argv.slice(1), argv[0]);
} else {
let port = process.env.PORT;
let host = process.env.HOST;
Expand All @@ -21,6 +21,7 @@ if (argv[0] === 'inspect') {
} else if (flag === '--help' || flag === '-h') {
console.log('Usage: agentxray [--port <port>] [--host <host>] [--version]');
console.log('Offline evidence: agentxray inspect --help');
console.log('Explicit raw evidence: agentxray evidence --help');
process.exit(0);
} else {
console.error(`agentxray: unknown option '${arg}'`);
Expand Down
97 changes: 71 additions & 26 deletions bin/inspect.js
Original file line number Diff line number Diff line change
@@ -1,23 +1,54 @@
const HELP = `Usage: agentxray inspect --platform <omp|codex|claude-code> <file.jsonl> [--json] [--fail-on pending-failures]
const HELP = `Usage: agentxray inspect --platform <omp|codex|claude-code> <file.jsonl> [--json] [--summary] [--fail-on pending-failures]
agentxray evidence --platform <omp|codex|claude-code> <file.jsonl> --sha256 <hash> --line <number> [--offset <bytes>] [--max-bytes <4..16384>] [--json]

Read one stable regular UTF-8 JSONL file (maximum 64 MiB), without starting a server.
Reports omit raw logs, paths, commands and IDs; source references are one-based lines/positions.
inspect: full report by default; --summary requires --json and bounds references, not aggregate counts.
evidence: explicit RAW content, potentially sensitive; always JSON, hash-checked, default maximum 4096 content bytes.
References use one-based physical lines. Byte pagination preserves UTF-8; use nextOffset to continue.
Exit 0: complete report, NOT task success. Exit 1: input/runtime/coverage error.
Exit 2: pending failure records found, only when --fail-on pending-failures is requested.
Exit 2: pending failure records found, only when inspect --fail-on pending-failures is requested.
JSON failures: {schemaVersion:1,kind:"error",error:{code,message}} on stdout.
Incomplete coverage retains report data on stdout and reports COVERAGE_INCOMPLETE on stderr.
`;

async function main(args) {
let filename,
platform,
json = false,
policy;
function wantsJson(args, mode) {
if (mode === 'evidence') return true;
const terminator = args.indexOf('--');
return args.slice(0, terminator < 0 ? args.length : terminator).includes('--json');
}

function errorDocument(code, message) {
return { schemaVersion: 1, kind: 'error', error: { code, message } };
}

function emitError(json, mode, code, message, help = false) {
if (json) process.stdout.write(`${JSON.stringify(errorDocument(code, message))}\n`);
process.stderr.write(`agentxray ${mode}: ${message}\n${help && !json ? HELP : ''}`);
process.exitCode = 1;
}

function integer(value) {
if (!/^\d+$/.test(value) || !Number.isSafeInteger(Number(value)))
throw new Error('Expected a nonnegative safe integer.');
return Number(value);
}

async function main(args, mode = 'inspect') {
let filename, platform, policy;
let summary = false;
const json = wantsJson(args, mode);
const evidenceOptions = {};
const seen = new Set();
let positionalOnly = false;
try {
if (args.length === 1 && ['--help', '-h'].includes(args[0])) {
process.stdout.write(HELP);
return;
}
const allowed =
mode === 'evidence'
? ['--platform', '--json', '--sha256', '--line', '--offset', '--max-bytes']
: ['--platform', '--json', '--summary', '--fail-on'];
for (let index = 0; index < args.length; index++) {
const arg = args[index];
if (!positionalOnly && arg === '--') {
Expand All @@ -26,18 +57,20 @@ async function main(args) {
}
if (!positionalOnly && arg.startsWith('-')) {
const [flag, ...inline] = arg.split('=');
if (!['--platform', '--json', '--fail-on'].includes(flag) || seen.has(flag))
throw new Error('Invalid or duplicate option.');
if (!allowed.includes(flag) || seen.has(flag)) throw new Error('Invalid or duplicate option.');
seen.add(flag);
if (flag === '--json') {
if (inline.length) throw new Error('--json does not take a value.');
json = true;
if (['--json', '--summary'].includes(flag)) {
if (inline.length) throw new Error('Boolean flags do not take a value.');
if (flag === '--summary') summary = true;
continue;
}
const value = inline.length ? inline.join('=') : args[++index];
if (!value || value.startsWith('-')) throw new Error('Missing option value.');
if (flag === '--platform') platform = value;
else policy = value;
else if (flag === '--fail-on') policy = value;
else if (flag === '--sha256') evidenceOptions.sha256 = value;
else
evidenceOptions[{ '--line': 'line', '--offset': 'offset', '--max-bytes': 'maxBytes' }[flag]] = integer(value);
} else {
if (filename !== undefined) throw new Error('Provide exactly one input file.');
filename = arg;
Expand All @@ -46,31 +79,43 @@ async function main(args) {
if (!filename || !platform) throw new Error('Explicit --platform and one input file are required.');
if (policy !== undefined && policy !== 'pending-failures')
throw new Error('Supported --fail-on policy: pending-failures.');
if (summary && !json) throw new Error('--summary requires --json.');
if (mode === 'evidence' && (!evidenceOptions.sha256 || !evidenceOptions.line))
throw new Error('Evidence requires --sha256 and --line.');
} catch (error) {
process.stderr.write(`agentxray inspect: ${error.message}\n${HELP}`);
process.exitCode = 1;
emitError(json, mode, 'INVALID_ARGUMENT', error.message, true);
return;
}
let implementation;
let implementation, detail;
try {
implementation = require('../lib/inspect');
if (summary || mode === 'evidence') detail = require('../lib/inspect-detail');
} catch {
process.stderr.write('agentxray inspect: bundled rules unavailable; rebuild or reinstall the package.\n');
process.exitCode = 1;
emitError(json, mode, 'RULES_UNAVAILABLE', 'Bundled rules unavailable; rebuild or reinstall the package.');
return;
}
const { inspectFile, renderText, InspectError } = implementation;
try {
const report = await inspectFile(filename, platform);
const full =
mode === 'evidence'
? await detail.readEvidence(filename, platform, evidenceOptions)
: await inspectFile(filename, platform);
const report = summary ? detail.createSummary(full) : full;
process.stdout.write(json ? `${JSON.stringify(report, null, 2)}\n` : renderText(report));
process.exitCode = !report.complete ? 1 : policy && report.summary.pendingRecords ? 2 : 0;
if (!report.complete)
process.stderr.write('agentxray inspect: adapter coverage is incomplete; see report.coverage.issues.\n');
process.exitCode = !full.complete ? 1 : policy && full.summary.pendingRecords ? 2 : 0;
if (!full.complete) {
const message = 'Adapter coverage is incomplete; use the full inspect report for coverage issues.';
process.stderr.write(
json ? `${JSON.stringify(errorDocument('COVERAGE_INCOMPLETE', message))}\n` : `agentxray ${mode}: ${message}\n`
);
}
} catch (error) {
process.stderr.write(
`agentxray inspect: ${error instanceof InspectError ? error.message : 'Inspection failed; no report generated.'}\n`
emitError(
json,
mode,
error instanceof InspectError ? error.code : 'INSPECTION_FAILED',
error instanceof InspectError ? error.message : 'Inspection failed; no report generated.'
);
process.exitCode = 1;
}
}

Expand Down
Loading
Loading