Skip to content

feat: a wire record of what the agent asked the model, outside the agent - #3

Closed
praxagent wants to merge 3 commits into
mainfrom
feat/wire-record
Closed

praxagent wants to merge 3 commits into
mainfrom
feat/wire-record

Conversation

@praxagent

Copy link
Copy Markdown
Owner

Idea credit: NVIDIA's Open Agent Safety Platform: monitoring "on the node's only path to the model", out of the agent's reach. Assessed in praxagent/prax#246.

Why. An agent's audit log and traces live in the process they audit, so a compromised agent can drop its own entries. This proxy is on the model path and outside that process.

What (PROXY_WIRE_RECORD, default off; secrets_proxy/wire_record.py). The forward proxy appends one line per model response. It parses OpenAI chat, the OpenAI Responses API and Anthropic, as JSON or SSE, and reassembles streamed tool-call arguments. Each line holds caller, host, path, status, model, request hash, size, and the tool calls the model returned as names + argument hashes, never text.

  • Hash-chained: an edit, deletion or reordering breaks the chain (python -m secrets_proxy.wire_record verify FILE).
  • Tamper-evident, not tamper-proof: someone who can rewrite the whole file can rebuild the chain. Keep ./wire writable only by the proxy and anchor the head hash off the box. The README says so.
  • caller (the proxy username) separates instances sharing the proxy, e.g. dev and prod.
  • Recording never breaks a response. Responses over stream_large_bodies (1 MB) are recorded by size only; documented.
  • wire/ is gitignored (runtime data).

Prax's side (scripts/check_wire_record.py, compare the record with Prax's traces) follows in a prax PR. Its chain check was cross-verified against records written by this code.

Tests: tests/test_wire_record.py (11): parsers, no-text guarantee, chaining across restarts, three kinds of tampering, host filtering, the mitmproxy hook. 74 pass; ruff clean.

Idea credit: NVIDIA's Open Agent Safety Platform — monitoring on the node's
only path to the model, out of the agent's reach.

The agent's audit log lives in the process it audits, so a compromised agent
can drop entries. With PROXY_WIRE_RECORD set, the forward proxy appends one
line per model response (OpenAI chat/Responses and Anthropic, JSON or SSE):
host, path, status, model, request hash, size, and the tool calls the model
returned as names + argument hashes — never text. Lines are hash-chained, so
an edit, deletion or reordering breaks the chain (python -m
secrets_proxy.wire_record verify). Tamper-evident, not tamper-proof: keep the
file writable only by the proxy and anchor the head hash off the box.
Recording never breaks a response.
Dev and prod can share the forward proxy; the caller label (the proxy
username) is captured before the credential is stripped and written on each
wire line, so a check runs against one instance's traces.
… after stripping it

Every injection audit line said caller=-. Read the caller once, before the
strip, and use it for both the audit line and the wire record.
@praxagent

Copy link
Copy Markdown
Owner Author

Added a fix found while reading this code: the injection audit line has always logged caller=-, because it read Proxy-Authorization after the header was stripped. That's true on main too. The caller is now read once, before the strip, and used for both the audit line and the wire record. New test test_audit_line_names_the_caller fails without the fix. 75 passed.

@praxagent

Copy link
Copy Markdown
Owner Author

Folded into #8 (stacked PRs re-conflict after every squash merge). Every commit from this branch is in #8 — merge that one.

@praxagent praxagent closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant