Skip to content

docs(research): AutoBenchmark — cross-provider difficulty, and hard-for-all means audit the key - #252

Closed
praxagent wants to merge 5 commits into
mainfrom
docs/autobenchmark
Closed

praxagent wants to merge 5 commits into
mainfrom
docs/autobenchmark

Conversation

@praxagent

Copy link
Copy Markdown
Owner

Assessment of Meta RAM's AutoBenchmark: a research agent builds benchmarks for research agents.

Stacked on #251, which is on #246, because all three edit the same README list and tracker.

Verdict: document + adopt two checks; don't run the loop.

  • The finding. Unaided benchmarks saturate: solvers score above 80, and Opus-5 scores 98.0. A detailed human spec with curated grounding roughly halves solver scores (Rebuttal 84.4 → 43.5). A one-line intent barely helps. This matches our benchmark-saturation finding that resilience comes from expert curation.
  • Adopt 1. Record the difficulty of every LLM-authored case (ARC synthetics, praxbench, battery) on a solver from another provider, and never select on it.
  • Adopt 2. Treat "hard for every solver" as a reason to audit the answer key. The loop keeps the lowest-scoring checkpoint, which also selects wrong reference answers, and every solver fails those alike. The correctness judge is the proposer's own model (Muse-Spark plays proposer, in-loop solver and every judge). The post reports no human check of the keys.
  • Not adopted. Running the loop: there's no code yet, it needs 24 h / 16 GB per task, and our suite already fails the discrimination check (docs(research): assess Anthropic's automated eval design and hillclimbing #245).

Blog-level evidence only: the report is planned, and no code or data is out.

Document + adopt two ideas; confirmation of the Muse-parity containment.
- An out-of-band wire record in the secrets proxy — the tool calls the model
  returned, hashed, append-only, where Prax can't write — checked against
  Prax's own trace, so a compromised Prax can't hide activity by editing it.
- A diff of newly granted access for egress-policy changes and timed grants.
The DPU hardware is out of reach; OpenShell (0.1.x) is a peer to watch.
OpenWorker (Andrew Ng et al., MIT): the closest peer to Prax's governance
stance. Adopt hard floors — a declared set enforced after every rule that can
lower risk; Prax's earned trust can lower two login steps from HIGH to MEDIUM
on self-reported success today — and parked approvals for unattended runs
instead of refusing and losing the work. Plus approval provenance per call.

OpenShell's product page adds per-program network policy and a policy prover
with an access ceiling: queue the ceiling, and a time-boxed evaluation of
OpenShell as prax-sandbox's runtime.
Not ruled out: Prax should be highly competitive with OpenShell. Candidate
routes recorded — per-program proxy identity inside the sandbox, cgroup/eBPF
attribution, or OpenShell's supervisor after the evaluation.
…nnel-held identity

Google's CNCF sandbox application (cncf/sandbox#523). Its egress design is the
closest published match to the secrets proxy. Adopted: never inject into
cleartext (ours did; fixed in prax-secrets-proxy #7). Queued: a trusted tunnel
client holds the sandbox's proxy identity, so the program can't read it. The
platform is a scale non-goal; its DNS bypass matches our documented gap.
…or-all means audit the key

Meta RAM's agent-built benchmarks saturate unaided; detailed human specs halve
solver scores. Adopt: difficulty of LLM-authored cases measured on a solver
from another provider, and a case every solver fails gets its answer key
audited — lowest-score selection also selects wrong keys.
@praxagent

Copy link
Copy Markdown
Owner Author

Folded into #256, which contains every commit from this branch (the five research PRs were a stack, and stacked PRs re-conflict after every squash merge). Merge #256.

@praxagent praxagent closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant