Repository navigation
Conversation
Document + adopt two ideas; confirmation of the Muse-parity containment. - An out-of-band wire record in the secrets proxy — the tool calls the model returned, hashed, append-only, where Prax can't write — checked against Prax's own trace, so a compromised Prax can't hide activity by editing it. - A diff of newly granted access for egress-policy changes and timed grants. The DPU hardware is out of reach; OpenShell (0.1.x) is a peer to watch.
OpenWorker (Andrew Ng et al., MIT): the closest peer to Prax's governance stance. Adopt hard floors — a declared set enforced after every rule that can lower risk; Prax's earned trust can lower two login steps from HIGH to MEDIUM on self-reported success today — and parked approvals for unattended runs instead of refusing and losing the work. Plus approval provenance per call. OpenShell's product page adds per-program network policy and a policy prover with an access ceiling: queue the ceiling, and a time-boxed evaluation of OpenShell as prax-sandbox's runtime.
Not ruled out: Prax should be highly competitive with OpenShell. Candidate routes recorded — per-program proxy identity inside the sandbox, cgroup/eBPF attribution, or OpenShell's supervisor after the evaluation.
…nnel-held identity Google's CNCF sandbox application (cncf/sandbox#523). Its egress design is the closest published match to the secrets proxy. Adopted: never inject into cleartext (ours did; fixed in prax-secrets-proxy #7). Queued: a trusted tunnel client holds the sandbox's proxy identity, so the program can't read it. The platform is a scale non-goal; its DNS bypass matches our documented gap.
…or-all means audit the key Meta RAM's agent-built benchmarks saturate unaided; detailed human specs halve solver scores. Adopt: difficulty of LLM-authored cases measured on a solver from another provider, and a case every solver fails gets its answer key audited — lowest-score selection also selects wrong keys.
praxagent
force-pushed
the
docs/autobenchmark
branch
from
October 2, 2026 05:34
88adfb0 to
444c578
Compare
Owner
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Assessment of Meta RAM's AutoBenchmark: a research agent builds benchmarks for research agents.
Stacked on #251, which is on #246, because all three edit the same README list and tracker.
Verdict: document + adopt two checks; don't run the loop.
Blog-level evidence only: the report is planned, and no code or data is out.