Conversation
JobRunner held the only MachineAdapter reference during a job and called machine.execute() directly for load_gcode and start, with zero references to the safety subsystem. Economic authorization (a funded job) reached hardware without any physical-safety admission check (R-11). Both actuation sites now go through SafetyGateway.validateAndRelay(). The same MachineCommand object that is described to the governor is the object the adapter receives, so there is no validate-then-mutate gap. load_gcode and start are classified 'scoped' per the governor's own Class 3 definition (protocol upload, run), so a job without an execution scope is denied and issues zero machine commands. A missing/uninitialized gateway also fails closed.
…y boundary job-runner-safety.test.ts (14 tests) covers the R-11 memo item: - positive control: an admitted job dispatches exactly load_gcode then start, both through SafetyGateway.validateAndRelay, and the PhysicalCommand params object is reference-identical to the payload the adapter received - negative controls, each asserting machine.execute is called ZERO times: missing execution scope (the prohibition the governor really enforces under its default config), engaged e-stop, open circuit breaker, and an uninitialized gateway singleton - short-circuit: a denied load_gcode starts no sensors, takes no snapshot and never issues start tier-enforcement.test.ts: its 7 JobRunner integration tests never initialized a gateway and never supplied a scope, so they hit the new fail-closed path. Their setup now initializes a per-test gateway and passes a scopeId. No assertion was changed or relaxed; the safety boundary itself is asserted in the new file.
Outcome, design note, negative-control evidence, the six-site reachability inventory, and the honest gaps (unscoped construction sites, unguarded getProgress/getStatus polls, no G-code envelope enforcement, commandId not in the signed bundle). Force-added: ai/ is gitignored in this repo, and the lane brief requires the report to ship on the branch.
…ow-up Records, before any code edit, where the execution scope comes from on each gateway path: the paid path mints one but never dispatches in-process, and the two paths that do dispatch never mint one. Also records why resolving the scope from the DB by jobId was rejected as dead code (jobs.insert has no onConflict).
…op the class downgrade JobRunner classifies load_gcode/start as class "scoped" and the governor denies a scoped command with no scope, so a gateway-dispatched job needs its caller's execution scope to reach a device at all. Three changes in kernel-service.ts: 1. Pass scopeId/agentDid into runner.run(). BOTH call sites, not just the Sentry one — the fallback path runs whenever Sentry is uninitialised, which is the path most tests and self-hosted deployments take, so leaving it un-threaded would have left the fix half-applied. 2. Delete the `params.scopeId ? "scoped" : "safe"` pre-flight downgrade. It lowered the class exactly when the credential was missing, which made the one check that bites under the default governor config unfalsifiable: an unscoped job was admitted as "safe" and only denied later, out of band, inside runner.run(). The class is now fixed at "scoped", matching the kernel's own MACHINE_COMMAND_CLASS table (module-private, so restated with a pointer). 3. Remove the now-redundant per-job breaker recording on the success path. validateAndRelay already records each command's real outcome, so recording again per job double-counted every device failure (failureThreshold 2 tripped after ONE failing job) and mis-attributed a safety denial — which never touches the adapter — to a healthy device. The .catch() path still records, since a rejected run() never reached the dispatch boundary. Existing tests: kernel-service-safety and tracing submit unscoped jobs, which the boundary now denies. Added a scopeId to their SETUP only — no assertion was changed, removed or relaxed, and no check was made permissive. The breaker test's original threshold arithmetic (closed after 1 job, open after 2) passes unchanged once the double-count is fixed, which is what confirms the diagnosis.
Seven tests over the two independent layers, because the claim is defence in
depth and a suite covering only the outer one would not notice the inner one
rotting:
Layer 1 (submitJob's synchronous pre-flight, gateway.validateOnly)
- NEGATIVE CONTROL: an unscoped submission is denied and machine.execute is
called ZERO times.
- An unscoped denial leaves the device's circuit breaker untouched — a missing
credential is the caller's fault, and counting it against the device would
let an unauthorised caller deny service to everyone else.
- The pre-flight describes an UNSCOPED job to the governor as class 'scoped',
never 'safe'. This pins the deleted downgrade: 'safe' is the bug.
Layer 2 (JobRunner's dispatch boundary, gateway.validateAndRelay)
- NEGATIVE CONTROL: with the pre-flight deliberately stubbed to fail OPEN, the
runner still denies — ZERO machine.execute calls, job recorded 'failed'.
Only commandIds starting with 'preflight:' are waved through, so the
runner's own checks still hit the real governor with the real config.
Positive controls
- scopeId/agentDid reach runner.run()'s config.
- A scoped job actually actuates: load_gcode then start, and every command the
governor saw carried class 'scoped' plus the caller's scope and DID.
afterEach uses clearAllMocks, not restoreAllMocks: the latter also strips the
implementations off the hoisted vi.mock factories, which turns
Sentry.startSpanManual into a no-op that never invokes its callback — so
runner.run() never fires and every test after the first silently passes without
executing anything. Found that the hard way.
…ow-up Outcome is PARTIAL and the report says so up front: the threading and the downgrade removal are done and proven (2983 passed, typecheck exit 0), but the blocking item did not move, because no in-process gateway caller has a scope to thread. The producer of the scope (paid-job-flow) and the callers of submitJob (job.facade, setup) are different code paths that never meet, so the one-line fix the brief expected does not exist. Records: the baseline-vs-change measurement (both full-suite runs), the three defects found including the breaker double-count and device-blaming that were not in the brief, the four honestly-red job-submit tests with which one was already red at b1be6cf, the negative-control assertions for both boundary layers, and eleven honest gaps with per-site touched/not-touched status.
…ess job producers KernelService now denies an unscoped submitJob (b6c9c77), which is correct -- a dispatch boundary must not mint the credential it checks. But that left both in-process producers dispatching without one, so every gateway-submitted job failed closed at the pre-flight. The grant now comes from the producer, which is the only party that knows the authenticated principal: - New services/execution-scope-service.ts is the single mint. It owns the write-tool vocabulary (moved verbatim out of routes/paid-job-flow.ts, which is why the two producers can now derive an allowed-tool set without importing from a route module), normalises the principal into the did:pcc: form the governor buckets its rate limit by, and inserts the execution_scopes row. - facades/job.facade.ts (POST /api/jobs/submit) mints after the job row exists and before submitJob, on the local-kernel path only; createdBy is the authenticated actor. The external-kernel path dispatches nothing in-process and gets no grant. scopeId + createdBy now land in the audit-log metadata. - routes/setup.ts (POST /api/setup/test-job) mints for the onboarding self-test, past the deviceless branch that never reaches a device. Its budget is deliberately tighter (50 commands / 1 retry / 15 min) because routes/device-relay.ts resolves active scopes by kernelId+createdBy without a jobId filter, so a generous one-shot grant widens what that principal can relay afterwards. - routes/paid-job-flow.ts switches to the helper. The inserted row is unchanged field-for-field: the helper's defaults ARE that path's former literals (200 commands / 5 retries / 1 h TTL / same id expression), and createdAt is still the caller's shared `now`. Neither producer reads scopeId or agentDid off the request body -- both route bodies are additionalProperties:true, and honouring a caller-supplied id would let a caller present a grant issued to someone else. A mint failure aborts the submission rather than falling through to an unscoped dispatch. Turns the 4 red job-submit tests green with no edits to them; the 7 kernel-service-scope and 14 job-runner-safety boundary tests are untouched.
…nting Layer 3 of the JobRunner safety boundary. Layers 1 (job-runner-safety) and 2 (kernel-service-scope) are the deny side; without a producer that grants a scope they deny everything, so these cover the grant. Assertions go down to SafetyGateway.validateAndRelay rather than stopping at the HTTP response, because a 200 only means the job was admitted: - POST /api/jobs/submit writes exactly one execution_scopes row bound to the job and to the authenticated principal, hands that same id (not a second, unrecorded one) to KernelService.submitJob, and both load_gcode and start reach the recording adapter carrying it as class "scoped". - Unauthenticated submits record createdBy="unauthenticated" rather than an invented operator id. - The external-kernel path mints nothing, since it dispatches nothing in-process. - Negative control: a body-supplied scopeId naming a real, active scope owned by another operator is ignored at every layer -- the minted scope is what reaches submitJob and the governor, and the victim's row is left unrebound. - POST /api/setup/test-job mints its own scope, runs to completion against the real mock adapter (a real bun_ evidence bundle; the job row reaches evidence_stored), and carries the minted scope to the governor. Its budget assertions pin the tighter one-shot grant (50 commands / 1 retry / <=15 min). The test-job case takes ~10 s of wall clock: the route's poll loop breaks only on "completed"/"failed", but the settlement pipeline advances the row past "completed" to "evidence_stored", so the loop always runs to its deadline. That is pre-existing route behaviour, asserted as-is rather than papered over.
Owner
Author
|
Producer-side scope minting landed (commits 8df9699, b2d23c5, 6865e67) — the PR is now functionally complete.
Still unscoped (fail closed, out of scope here): agent-kernel via agent-bridge, the standalone kernel server, onboard-kit quick-start + scaffolder template. |
Board rows R30/N3 (pcc-reconciliation): the gateway lane rebases this PR onto current master before cross-family review. Merge, not rebase, so the PR history is preserved and no force-push is needed. Dry-run merge-tree was clean. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A8smAyyypDGhf46eT71TXo
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes memo risk R-11 ("physical safety is not implied by economic authorization") at the actual dispatch boundary, and sub-blocker G-8c. Ledger: R-11 / G-8 / MS-11. Two bounded Spark implementers (
jobrunner-safety-implementer-1,jobrunner-safety-gateway-followup-1), branch offlamasu/master7a86491. Reports:ai/lane-runs/jobrunner-safety/REPORT.md.Finding on master (why this exists):
packages/gateway/src/services/kernel-service.ts:236-266pre-flighted a SYNTHETIC command (type: "submit_job"), never theload_gcode/startthat reach hardware, and:241downgraded the class to"safe"wheneverscopeIdwas absent, so the only admission check that bites could never fail.runner.run()(:311and the Sentry-less twin at:399) never received a scope.job-runner.tshad zero references to safety. R-11 was literally true.Kernel (commits 4f081d4, 3b8db45, b1be6cf)
JobRunneractuation sites route throughSafetyGateway.validateAndRelay(). Fixed class table:load_gcode/start=scoped,pause/resume/stop=safe,status=read(never downgraded by the presence of a credential). The SAME params object is validated and executed (reference identity asserted; no validate-then-mutate gap). Missing gateway -> fail closed. Sentry spans preserved.machine.executecalled ZERO times; a phase-1 denial starts no sensors, takes no snapshot, never starts.@pcc/kernel: 877/877, tsc clean. Steward re-run on the Spark: boundary + tier-enforcement 37/37.Gateway (commits 4704fad, 4c8d7f6, 3b560f5, b6c9c77)
scopeId/agentDidthreaded intorunner.run()at BOTH call sites; the:241downgrade deleted (pre-flight now honest); per-job breaker recording removed from the success paths (it double-counted with the kernel'svalidateAndRelayand recorded a caller's missing credential as a DEVICE failure, so an unauthorized caller could trip a healthy device's breaker for everyone).@pcc/gateway: 2983 passed / 4 failed / 6 skipped, tsc clean. Baseline at b1be6cf was 2981 / 6.The 4 red tests are the point, not a defect: all four are
POST /api/jobs/submit(job-submit.test.ts), becausefacades/job.facade.ts:293supplies no execution scope; one of them was already red on master (repeated unscoped submits tripped the mock device's breaker via the double count). They are the honest, loud version of a previously silent hole. Producer-side decision (steward): the direct-submit facade andPOST /api/setup/test-jobmust MINT an execution scope at job creation for the authenticated principal, in the sameexecution_scopestable the paid path already uses (paid-job-flow.ts:656), and pass its id tosubmitJob-- never mint insideKernelServiceto satisfy the check, and never accept a caller-supplied scope id for a job it does not own. A third bounded commit lands that on this branch before the draft is marked ready.Not enforced yet (honest gaps): no G-code/toolpath envelope inspection (default envelope only bites on velocity/temperature/force keys, absent from a print job); breaker/governor state is per-process RAM; commandId is not in the signed evidence bundle;
getProgress/getStatuspolls are read-class and bypass the gateway; sites 3-6 (agent-kernel via agent-bridge, standalone kernel server, onboard-kit quick-start + scaffolder template) still dispatch unscoped and therefore fail closed.Reviewers: gateway 0600b204 (owner of kernel-service + facades), sensors 7a438686. Steward 57c2a412 graded. Operator merges (no lane merges to master).
🤖 Generated with Claude Code
https://claude.ai/code/session_012PxrWRej6UAvZepXkpxJPi