Skip to content

feat: hard floors, parked approvals for unattended runs, and why every call was allowed - #249

Merged
praxagent merged 4 commits into
mainfrom
feat/approval-provenance
Oct 3, 2026
Merged

praxagent merged 4 commits into
mainfrom
feat/approval-provenance

Conversation

@praxagent

@praxagent praxagent commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

This PR now carries #247 and #248 as well: it was stacked on them, and stacked PRs re-conflict after every squash merge. Both are closed with a pointer. Merge this one.

#247 Hard floors

Some actions always need a person's decision: credentials and logins, plugin install/activate/write, self-deploy, and GPU power-on.

  • The floors are enforced after every risk-lowering rule: earned trust, auto-approve, the turn latch, and spoke enforce.
  • Timed grants don't count.
  • With approvals off, the user's own message must name the action and its target.
  • Behind HARD_FLOORS_ENABLED; extra floors via HARD_FLOOR_EXTRA_TOOLS. Idea credit: OpenWorker (Andrew Ng et al.).

#248 Parked approvals

An unattended run (scheduler or task runner) parks its TeamWork approval instead of failing.

#249 Approval provenance

Every executed call's audit entry records why it was allowed to run: person:<id>, user_message, auto:user_request, auto:earned_trust, model_reconfirmed, earlier_confirmation, high_risk_not_enforced, none_needed.


Original #249 description

Stacked on #248 (itself on #247). Merge those first; I'll rebase this onto main afterwards.

Pattern credit: OpenWorker (Andrew Ng and contributors): approval provenance on every tool call.

The audit entry of each executed call now carries approval:

value meaning
person:<id> a person approved it in TeamWork
user_message the user's own message named it (hard-floor fallback)
auto:user_request smart auto-approve from the user's message
auto:earned_trust earned trust lowered its risk
model_reconfirmed the model confirmed itself by calling again
earlier_confirmation unlocked by a confirmation earlier in the turn
high_risk_not_enforced HIGH risk in a spoke whose gate is off
none_needed not HIGH risk

The two bold paths were invisible before. Provenance is recorded per call, not on the shared turn state, because parallel tool calls share that. The trifecta gate is covered too; behaviour is unchanged.

Tests: tests/test_approval_provenance.py (8). make ci: green; no secrets-proxy requests during the run.

Pattern credit: OpenWorker (Andrew Ng and contributors) — dangerous
operations are human-only, always.

Every path past Prax's HIGH-risk gate could be lowered or skipped: earned
trust downgrades browser login steps on self-reported success; the fallback
lets the model confirm by calling again; one confirmation can unlock every
HIGH tool for the turn; inside spokes the gate isn't enforced unless
SPOKE_GOVERNANCE_ENABLED (off in production); a timed grant approves on
arrival; and browser_login / browser_credentials, which hand the model a
stored password, weren't classified at all.

HARD_FLOORS_ENABLED (default off): logging in or revealing credentials,
installing or activating code with Prax's authority, and starting a billable
GPU run only on a person's decision about that exact call — checked before
all of the above. The decision is an out-of-band TeamWork approval that a
timed grant did not give (decided_by is now carried through), or, with
approvals off, the user's own message naming the action and its target.
HARD_FLOOR_EXTRA_TOOLS adds floors; built-ins can't be removed.
Pattern credit: OpenWorker (Andrew Ng and contributors) — unattended runs
never self-approve; requests park in an inbox.

A scheduled or task-runner turn that reached an action needing a person
waited APPROVAL_WAIT_SECONDS for someone who wasn't there, was refused, and
lost the work. With PARKED_APPROVALS_ENABLED it creates the TeamWork request
with a long lifetime (PARKED_APPROVAL_HOURS, needs teamwork's
expires_in_seconds), records what would re-run the task and which exact
action is waiting, and ends saying so. A poller re-runs the task when a person
approves, spending that approval on that exact action once; refusal or expiry
is reported as "not done". A different action asks again, up to
PARKED_MAX_RESUMES per task. The store survives restarts.

Also: approval outcomes carry expires_at; hard-floor refusals pass a parked
message through so the model says the task is waiting.
Pattern credit: OpenWorker (Andrew Ng and contributors) — approval
provenance on every tool call.

The audit entry of an executed call now carries approval: person:<id>,
user_message, auto:user_request, auto:earned_trust, model_reconfirmed,
earlier_confirmation, high_risk_not_enforced, or none_needed. The two weak
paths — the model confirming itself, and HIGH-risk tools in unenforced spokes
— were invisible before; now they're named on every call. Recorded per call,
not on the shared turn state, because parallel tool calls share it.
@praxagent praxagent changed the title feat: every executed call records why it was allowed to run feat: hard floors, parked approvals for unattended runs, and why every call was allowed Oct 2, 2026
@praxagent
praxagent merged commit 72a09a1 into main Oct 3, 2026
1 check passed
@praxagent
praxagent deleted the feat/approval-provenance branch October 3, 2026 00:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant