Repository navigation
docs(research): five assessments — NVIDIA/OpenShell/OpenWorker, Agent Substrate, AutoBenchmark, DeepSeek Harness (corrected), Context Language Models - #256
Merged
Conversation
Document + adopt two ideas; confirmation of the Muse-parity containment. - An out-of-band wire record in the secrets proxy — the tool calls the model returned, hashed, append-only, where Prax can't write — checked against Prax's own trace, so a compromised Prax can't hide activity by editing it. - A diff of newly granted access for egress-policy changes and timed grants. The DPU hardware is out of reach; OpenShell (0.1.x) is a peer to watch.
OpenWorker (Andrew Ng et al., MIT): the closest peer to Prax's governance stance. Adopt hard floors — a declared set enforced after every rule that can lower risk; Prax's earned trust can lower two login steps from HIGH to MEDIUM on self-reported success today — and parked approvals for unattended runs instead of refusing and losing the work. Plus approval provenance per call. OpenShell's product page adds per-program network policy and a policy prover with an access ceiling: queue the ceiling, and a time-boxed evaluation of OpenShell as prax-sandbox's runtime.
Not ruled out: Prax should be highly competitive with OpenShell. Candidate routes recorded — per-program proxy identity inside the sandbox, cgroup/eBPF attribution, or OpenShell's supervisor after the evaluation.
…nnel-held identity Google's CNCF sandbox application (cncf/sandbox#523). Its egress design is the closest published match to the secrets proxy. Adopted: never inject into cleartext (ours did; fixed in prax-secrets-proxy #7). Queued: a trusted tunnel client holds the sandbox's proxy identity, so the program can't read it. The platform is a scale non-goal; its DNS bypass matches our documented gap.
…or-all means audit the key Meta RAM's agent-built benchmarks saturate unaided; detailed human specs halve solver scores. Adopt: difficulty of LLM-authored cases measured on a solver from another provider, and a case every solver fails gets its answer key audited — lowest-score selection also selects wrong keys.
dsh sandboxes report full/partial enforcement; Prax's hand-installed containment is never checked from inside. Bank: scrubbed child environments (69 subprocess calls, none sets env=; exposure unverified) and an advisory reminder before the hard turn limits. Monotonic deny-only guards confirm hard floors.
… the proxy token in HTTPS_PROXY
…ing it, and correct it The September README-only note said dsh has no governance layer. The code has approvals, deny-only guards and a process sandbox; what it lacks is network policy and audit. One page, one index entry, the correction stated.
… under Prax's invariants UW/Meta CLMs edit a file mirror of their own context and beat ACM, RLM and summarisation zero-shot. Third sighting with ACM and OptMem, so the queued context_compact/context_recall row moves up — but rebuilt: user turns and approvals pinned, roles and provenance immutable, the record never edited. As published, roles are rewritable and tool results fold into user. Code is CC BY-NC; nothing vendored.
Entries added since were inserted between the Muse entry and its indented continuation, so its Sentinel/authd/verdict bullets rendered under the wrong entry. Pure move; no line changed.
praxagent
force-pushed
the
docs/context-language-models
branch
from
October 2, 2026 05:34
2915144 to
9f69a38
Compare
This was referenced Oct 2, 2026
docs(research): AutoBenchmark — cross-provider difficulty, and hard-for-all means audit the key
#252
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR now carries all five open research assessments. #246, #251, #252 and #253 were a stacked chain, and stacked PRs re-conflict after every squash merge. They are folded in here and closed with a pointer. Merge this one.
docs/research/README.md. Entries had been inserted between the Meta Muse entry and its sub-bullets, so its Sentinel/verdict bullets rendered under the wrong entry. Pure move: no line changed.Original #256 description
Assessment of Context Language Models (UW / Meta; code).
Stacked on #253, the end of the research-doc chain.
Verdict: document + adopt the mechanism, rebuilt under Prax's invariants.
context_compact/context_recallrow moves up.user. Injected text could be laundered into the user's voice. The authors flag this risk but don't evaluate it.context_recall).