Skip to content

fix(kernel,gateway): enforce the SafetyGateway at JobRunner's real dispatch boundary and thread the execution scope (R-11 / G-8c) - #334

Draft
LamaSu wants to merge 11 commits into
masterfrom
fix/jobrunner-safety-boundary
Draft

LamaSu wants to merge 11 commits into
masterfrom
fix/jobrunner-safety-boundary

Conversation

@LamaSu

@LamaSu LamaSu commented Sep 9, 2026

Copy link
Copy Markdown
Owner

Closes memo risk R-11 ("physical safety is not implied by economic authorization") at the actual dispatch boundary, and sub-blocker G-8c. Ledger: R-11 / G-8 / MS-11. Two bounded Spark implementers (jobrunner-safety-implementer-1, jobrunner-safety-gateway-followup-1), branch off lamasu/master 7a86491. Reports: ai/lane-runs/jobrunner-safety/REPORT.md.

Finding on master (why this exists): packages/gateway/src/services/kernel-service.ts:236-266 pre-flighted a SYNTHETIC command (type: "submit_job"), never the load_gcode/start that reach hardware, and :241 downgraded the class to "safe" whenever scopeId was absent, so the only admission check that bites could never fail. runner.run() (:311 and the Sentry-less twin at :399) never received a scope. job-runner.ts had zero references to safety. R-11 was literally true.

Kernel (commits 4f081d4, 3b8db45, b1be6cf)

  • Both JobRunner actuation sites route through SafetyGateway.validateAndRelay(). Fixed class table: load_gcode/start = scoped, pause/resume/stop = safe, status = read (never downgraded by the presence of a credential). The SAME params object is validated and executed (reference identity asserted; no validate-then-mutate gap). Missing gateway -> fail closed. Sentry spans preserved.
  • 14 new tests incl. the memo's negative control: unscoped job / e-stop / open breaker / uninitialized gateway -> machine.execute called ZERO times; a phase-1 denial starts no sensors, takes no snapshot, never starts.
  • @pcc/kernel: 877/877, tsc clean. Steward re-run on the Spark: boundary + tier-enforcement 37/37.

Gateway (commits 4704fad, 4c8d7f6, 3b560f5, b6c9c77)

  • scopeId/agentDid threaded into runner.run() at BOTH call sites; the :241 downgrade deleted (pre-flight now honest); per-job breaker recording removed from the success paths (it double-counted with the kernel's validateAndRelay and recorded a caller's missing credential as a DEVICE failure, so an unauthorized caller could trip a healthy device's breaker for everyone).
  • 7 new tests: layer-1 pre-flight denial with zero actuation; layer-2 denial with the pre-flight deliberately stubbed to fail open (the runner's own governor still denies); the deleted downgrade pinned. 5 of the 7 fail against the old code.
  • @pcc/gateway: 2983 passed / 4 failed / 6 skipped, tsc clean. Baseline at b1be6cf was 2981 / 6.

The 4 red tests are the point, not a defect: all four are POST /api/jobs/submit (job-submit.test.ts), because facades/job.facade.ts:293 supplies no execution scope; one of them was already red on master (repeated unscoped submits tripped the mock device's breaker via the double count). They are the honest, loud version of a previously silent hole. Producer-side decision (steward): the direct-submit facade and POST /api/setup/test-job must MINT an execution scope at job creation for the authenticated principal, in the same execution_scopes table the paid path already uses (paid-job-flow.ts:656), and pass its id to submitJob -- never mint inside KernelService to satisfy the check, and never accept a caller-supplied scope id for a job it does not own. A third bounded commit lands that on this branch before the draft is marked ready.

Not enforced yet (honest gaps): no G-code/toolpath envelope inspection (default envelope only bites on velocity/temperature/force keys, absent from a print job); breaker/governor state is per-process RAM; commandId is not in the signed evidence bundle; getProgress/getStatus polls are read-class and bypass the gateway; sites 3-6 (agent-kernel via agent-bridge, standalone kernel server, onboard-kit quick-start + scaffolder template) still dispatch unscoped and therefore fail closed.

Reviewers: gateway 0600b204 (owner of kernel-service + facades), sensors 7a438686. Steward 57c2a412 graded. Operator merges (no lane merges to master).

🤖 Generated with Claude Code

https://claude.ai/code/session_012PxrWRej6UAvZepXkpxJPi

JobRunner held the only MachineAdapter reference during a job and called
machine.execute() directly for load_gcode and start, with zero references to
the safety subsystem. Economic authorization (a funded job) reached hardware
without any physical-safety admission check (R-11).

Both actuation sites now go through SafetyGateway.validateAndRelay(). The same
MachineCommand object that is described to the governor is the object the
adapter receives, so there is no validate-then-mutate gap. load_gcode and start
are classified 'scoped' per the governor's own Class 3 definition (protocol
upload, run), so a job without an execution scope is denied and issues zero
machine commands. A missing/uninitialized gateway also fails closed.
…y boundary

job-runner-safety.test.ts (14 tests) covers the R-11 memo item:
- positive control: an admitted job dispatches exactly load_gcode then start,
  both through SafetyGateway.validateAndRelay, and the PhysicalCommand params
  object is reference-identical to the payload the adapter received
- negative controls, each asserting machine.execute is called ZERO times:
  missing execution scope (the prohibition the governor really enforces under
  its default config), engaged e-stop, open circuit breaker, and an
  uninitialized gateway singleton
- short-circuit: a denied load_gcode starts no sensors, takes no snapshot and
  never issues start

tier-enforcement.test.ts: its 7 JobRunner integration tests never initialized a
gateway and never supplied a scope, so they hit the new fail-closed path. Their
setup now initializes a per-test gateway and passes a scopeId. No assertion was
changed or relaxed; the safety boundary itself is asserted in the new file.
Outcome, design note, negative-control evidence, the six-site reachability
inventory, and the honest gaps (unscoped construction sites, unguarded
getProgress/getStatus polls, no G-code envelope enforcement, commandId not in
the signed bundle).

Force-added: ai/ is gitignored in this repo, and the lane brief requires the
report to ship on the branch.
…ow-up

Records, before any code edit, where the execution scope comes from on each
gateway path: the paid path mints one but never dispatches in-process, and the
two paths that do dispatch never mint one. Also records why resolving the scope
from the DB by jobId was rejected as dead code (jobs.insert has no onConflict).
…op the class downgrade

JobRunner classifies load_gcode/start as class "scoped" and the governor denies
a scoped command with no scope, so a gateway-dispatched job needs its caller's
execution scope to reach a device at all. Three changes in kernel-service.ts:

1. Pass scopeId/agentDid into runner.run(). BOTH call sites, not just the Sentry
   one — the fallback path runs whenever Sentry is uninitialised, which is the
   path most tests and self-hosted deployments take, so leaving it un-threaded
   would have left the fix half-applied.

2. Delete the `params.scopeId ? "scoped" : "safe"` pre-flight downgrade. It
   lowered the class exactly when the credential was missing, which made the one
   check that bites under the default governor config unfalsifiable: an unscoped
   job was admitted as "safe" and only denied later, out of band, inside
   runner.run(). The class is now fixed at "scoped", matching the kernel's own
   MACHINE_COMMAND_CLASS table (module-private, so restated with a pointer).

3. Remove the now-redundant per-job breaker recording on the success path.
   validateAndRelay already records each command's real outcome, so recording
   again per job double-counted every device failure (failureThreshold 2 tripped
   after ONE failing job) and mis-attributed a safety denial — which never
   touches the adapter — to a healthy device. The .catch() path still records,
   since a rejected run() never reached the dispatch boundary.

Existing tests: kernel-service-safety and tracing submit unscoped jobs, which
the boundary now denies. Added a scopeId to their SETUP only — no assertion was
changed, removed or relaxed, and no check was made permissive. The breaker
test's original threshold arithmetic (closed after 1 job, open after 2) passes
unchanged once the double-count is fixed, which is what confirms the diagnosis.
Seven tests over the two independent layers, because the claim is defence in
depth and a suite covering only the outer one would not notice the inner one
rotting:

Layer 1 (submitJob's synchronous pre-flight, gateway.validateOnly)
  - NEGATIVE CONTROL: an unscoped submission is denied and machine.execute is
    called ZERO times.
  - An unscoped denial leaves the device's circuit breaker untouched — a missing
    credential is the caller's fault, and counting it against the device would
    let an unauthorised caller deny service to everyone else.
  - The pre-flight describes an UNSCOPED job to the governor as class 'scoped',
    never 'safe'. This pins the deleted downgrade: 'safe' is the bug.

Layer 2 (JobRunner's dispatch boundary, gateway.validateAndRelay)
  - NEGATIVE CONTROL: with the pre-flight deliberately stubbed to fail OPEN, the
    runner still denies — ZERO machine.execute calls, job recorded 'failed'.
    Only commandIds starting with 'preflight:' are waved through, so the
    runner's own checks still hit the real governor with the real config.

Positive controls
  - scopeId/agentDid reach runner.run()'s config.
  - A scoped job actually actuates: load_gcode then start, and every command the
    governor saw carried class 'scoped' plus the caller's scope and DID.

afterEach uses clearAllMocks, not restoreAllMocks: the latter also strips the
implementations off the hoisted vi.mock factories, which turns
Sentry.startSpanManual into a no-op that never invokes its callback — so
runner.run() never fires and every test after the first silently passes without
executing anything. Found that the hard way.
…ow-up

Outcome is PARTIAL and the report says so up front: the threading and the
downgrade removal are done and proven (2983 passed, typecheck exit 0), but the
blocking item did not move, because no in-process gateway caller has a scope to
thread. The producer of the scope (paid-job-flow) and the callers of submitJob
(job.facade, setup) are different code paths that never meet, so the one-line
fix the brief expected does not exist.

Records: the baseline-vs-change measurement (both full-suite runs), the three
defects found including the breaker double-count and device-blaming that were
not in the brief, the four honestly-red job-submit tests with which one was
already red at b1be6cf, the negative-control assertions for both boundary
layers, and eleven honest gaps with per-site touched/not-touched status.
…ess job producers

KernelService now denies an unscoped submitJob (b6c9c77), which is correct --
a dispatch boundary must not mint the credential it checks. But that left both
in-process producers dispatching without one, so every gateway-submitted job
failed closed at the pre-flight.

The grant now comes from the producer, which is the only party that knows the
authenticated principal:

- New services/execution-scope-service.ts is the single mint. It owns the
  write-tool vocabulary (moved verbatim out of routes/paid-job-flow.ts, which
  is why the two producers can now derive an allowed-tool set without importing
  from a route module), normalises the principal into the did:pcc: form the
  governor buckets its rate limit by, and inserts the execution_scopes row.
- facades/job.facade.ts (POST /api/jobs/submit) mints after the job row exists
  and before submitJob, on the local-kernel path only; createdBy is the
  authenticated actor. The external-kernel path dispatches nothing in-process
  and gets no grant. scopeId + createdBy now land in the audit-log metadata.
- routes/setup.ts (POST /api/setup/test-job) mints for the onboarding
  self-test, past the deviceless branch that never reaches a device. Its budget
  is deliberately tighter (50 commands / 1 retry / 15 min) because
  routes/device-relay.ts resolves active scopes by kernelId+createdBy without a
  jobId filter, so a generous one-shot grant widens what that principal can
  relay afterwards.
- routes/paid-job-flow.ts switches to the helper. The inserted row is unchanged
  field-for-field: the helper's defaults ARE that path's former literals
  (200 commands / 5 retries / 1 h TTL / same id expression), and createdAt is
  still the caller's shared `now`.

Neither producer reads scopeId or agentDid off the request body -- both route
bodies are additionalProperties:true, and honouring a caller-supplied id would
let a caller present a grant issued to someone else. A mint failure aborts the
submission rather than falling through to an unscoped dispatch.

Turns the 4 red job-submit tests green with no edits to them; the 7
kernel-service-scope and 14 job-runner-safety boundary tests are untouched.
…nting

Layer 3 of the JobRunner safety boundary. Layers 1 (job-runner-safety) and 2
(kernel-service-scope) are the deny side; without a producer that grants a scope
they deny everything, so these cover the grant.

Assertions go down to SafetyGateway.validateAndRelay rather than stopping at the
HTTP response, because a 200 only means the job was admitted:

- POST /api/jobs/submit writes exactly one execution_scopes row bound to the job
  and to the authenticated principal, hands that same id (not a second,
  unrecorded one) to KernelService.submitJob, and both load_gcode and start
  reach the recording adapter carrying it as class "scoped".
- Unauthenticated submits record createdBy="unauthenticated" rather than an
  invented operator id.
- The external-kernel path mints nothing, since it dispatches nothing in-process.
- Negative control: a body-supplied scopeId naming a real, active scope owned by
  another operator is ignored at every layer -- the minted scope is what reaches
  submitJob and the governor, and the victim's row is left unrebound.
- POST /api/setup/test-job mints its own scope, runs to completion against the
  real mock adapter (a real bun_ evidence bundle; the job row reaches
  evidence_stored), and carries the minted scope to the governor. Its budget
  assertions pin the tighter one-shot grant (50 commands / 1 retry / <=15 min).

The test-job case takes ~10 s of wall clock: the route's poll loop breaks only
on "completed"/"failed", but the settlement pipeline advances the row past
"completed" to "evidence_stored", so the loop always runs to its deadline.
That is pre-existing route behaviour, asserted as-is rather than papered over.
@LamaSu

LamaSu commented Sep 9, 2026

Copy link
Copy Markdown
Owner Author

Producer-side scope minting landed (commits 8df9699, b2d23c5, 6865e67) — the PR is now functionally complete.

  • New packages/gateway/src/services/execution-scope-service.ts: mintExecutionScope(...) inserts one execution_scopes row (throws on failure); getWriteToolsForDeviceType moved verbatim out of routes/paid-job-flow.ts (its inserted row stays byte-identical; the 16 + 3 paid-job-flow tests pass unchanged).
  • POST /api/jobs/submit (job.facade) and POST /api/setup/test-job (setup.ts) mint AFTER the job row exists and BEFORE submitJob; scopeId/agentDid are derived in the producer and never read off the request body. Test-job scopes are deliberately tighter (15 min / 50 commands / 1 retry) because device-relay.ts resolves scopes by (kernelId, createdBy, status) without a jobId filter.
  • Negative control: a request naming a REAL, active scope owned by another operator is ignored; the minted scope is used. 7 of the 8 new tests fail against the previous commit.
  • @pcc/gateway: 189 files / 2995 passed / 6 skipped; tsc clean. Steward re-run on the Spark at 9cfc2dec: producer-execution-scope + job-submit + kernel-service-scope + kernel-service-safety = 36/36.

Still unscoped (fail closed, out of scope here): agent-kernel via agent-bridge, the standalone kernel server, onboard-kit quick-start + scaffolder template. execution_scopes.created_by records "unauthenticated" when a request carries no principal rather than inventing one.

Board rows R30/N3 (pcc-reconciliation): the gateway lane rebases this PR onto
current master before cross-family review. Merge, not rebase, so the PR history
is preserved and no force-push is needed. Dry-run merge-tree was clean.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A8smAyyypDGhf46eT71TXo

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant