Repository navigation
Drill self-hosted operator diagnosis across published surfaces #143
Description
Activity
- addedauthority:githubGitHub is the authoritative lifecycle record for this workGitHub is the authoritative lifecycle record for this workkind:cross-repositoryWork spans more than one public repositoryWork spans more than one public repositorypriority:P2Normal-priority product workNormal-priority product workstatus:readyReady for implementationReady for implementationrepo:github-control-planeOwned by the public organization control planeOwned by the public organization control planecompletion:evidence-requiredClose only after all explicit acceptance and operational evidence is publicClose only after all explicit acceptance and operational evidence is public
on Sep 27, 2026 Next qualification tuple
The activity-outcome work in Workflow #575 is published and verified. The next diagnosis drill can freeze:
- Server 2.4.30, image index
sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccf. - Workflow 2.2.21, PHP SDK 2.1.6, Python SDK 2.3.7, Rust SDK 2.1.2, CLI 2.1.2 and Waterline 2.0.7.
Start with the absent-worker and wrong-queue cases in a fresh isolated synthetic stack. Capture a clean reference state and the staged fault, then record published CLI/API/Waterline observations, diagnosis and recovery timings, the safe action selected and verified workflow state after recovery. The remaining lease, activity failure, backend interruption and storage-pressure cases follow the same evidence format. Each ambiguity gets a fix in its owning repository and a repeat against the published artifact.
This is the recorded next action. The fault drill has not yet run.
- Server 2.4.30, image index
First published operator case staged and diagnosed
Used Server 2.4.30 pinned to
sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccfwith its documented single-node SQLite bootstrap/dispatch path, and published CLI 2.1.2dw.pharverified against the release checksum manifest. This is a self-guided drill, not a blind independent-investigator result. Ground truth is recorded separately from the visible diagnostic evidence. No product source or image overlay is used.Started one synthetic
op143.greeterworkflow onop143-orderswithout an application SDK worker. Published commands used:dw workflow:start --type=op143.greeter --task-queue=op143-orders \ --workflow-id=op143-absent-worker --input='["Ada"]' \ --execution-timeout=3600 --json dw workflow:describe op143-absent-worker --json dw task-queue:describe op143-orders --json dw doctor --output=jsonWorkflow describe reports
pendingon the expected queue. Queue describe explicitly reportsno_active_workers, 0 active workers, 1 ready workflow task, 0 leases and an empty poller list. This supplies an actionable diagnosis from published surfaces: Server accepted the work, and no matching application worker can lease it. The safe next action is to start the application's SDK worker for namespacedefaultand queueop143-orders, then verify the original execution and history complete once. Manual task completion would replace application behavior and is not the recovery action for this case.This is diagnostic evidence, not a completed-case pass. Matching-worker recovery, timing and final workflow/history verification remain next, followed by the wrong-queue case and the other four failures. Raw synthetic state is retained for this continuation, and the local stack is stopped between phases. No Cloud runtime or paid host is used.
Absent SDK worker: recovered and verified with published artifacts
The CLI identified the stalled workflow on
op143-orders, one ready workflow task, zero leases and no active worker. The published PHP SDK worker was then started on that queue with the workflow and activity registered. No history, task or database state was manually completed or repaired.The original run
01m3rm6pq6382xq2fec9nmt71xcompleted once with{"greeting":"hello, Ada"}. History contains one each ofStartAccepted,WorkflowStarted,ActivityScheduled,ActivityStarted,ActivityCompletedandWorkflowCompleted. The original run count remains one. The queue subsequently reports one active PHP SDK 2.1.6 worker, zero ready tasks, zero leased tasks and accepting admission.The worker container started at 08:22:16.801 UTC, and the workflow closed at 08:22:18.321 UTC, about 1.52 seconds later. The workflow had been intentionally left pending while independent work continued. Its roughly 40-minute total wait is not a diagnosis or service-recovery benchmark. This is a self-guided CLI/API case, not an independent blind investigator result. Waterline remains to be exercised before the full drill is complete.
Published tuple: Server 2.4.30,
sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccf, Workflow 2.2.21, CLI 2.1.2 with release checksum verification, PHP SDK 2.1.6. SQLite is the Server's documented published-image setup for this initial isolated case.Commands used to investigate and verify:
dw workflow:describe op143-absent-worker --json dw task-queue:describe op143-orders --json dw doctor --output=json dw workflow:history op143-absent-worker 01m3rm6pq6382xq2fec9nmt71x --output=json
Next case is a live worker polling a different queue. The new synthetic run targets
op143-orderswhile that worker pollsop143-other. Investigate through the same published surfaces, correct its queue configuration and verify the original run, without restarting or creating a replacement workflow.Wrong queue: diagnosed and recovered
The published CLI showed the pending workflow targeting
op143-orders. Its queue had one ready task, zero leases and zero active workers. The previous worker registration was explicitly marked stale. In contrast,op143-othershowed one active PHP SDK 2.1.6 worker, zero ready tasks and accepting admission. A healthy worker process therefore did not imply that the workflow's queue was served.Recovery was to stop the worker polling
op143-otherand start the same registered application worker onop143-orders. The original workflowop143-wrong-queue, run01m3rpt75fjmvm9jzptrge19x9, completed at 08:30:34.924 UTC with{"greeting":"hello, Ada"}and run count one. Its history has one activity completion and one workflow completion. No replacement execution or manual task completion was used.This case uses the same exact published tuple as the preceding absent-worker result. The evidence distinguishes configured queue, worker liveness, ready work and active leases. No product defect has been established in these two CLI/API cases. Waterline coverage, expired leases, activity failure, backend interruption and storage pressure remain open. This is still a self-guided drill.
All task containers are stopped between phases. Next action is to exercise the published Waterline diagnostics alongside a bounded expired-lease failure and verify safe lease recovery through the original run's history.
Expired workflow-task lease: observed and recovered
A synthetic diagnostic worker claimed the original task and then stopped sending heartbeats or completion. The task's published lease expired at 08:41:51.068 UTC. The CLI subsequently showed that exact task, owner and expired deadline, one expired lease and one repair candidate on
op143-lease.The published Waterline 2.0.7 browser UI showed that queue as Needs attention, with 1 leased, 1 expired, one repair candidate and an expired-age counter. Its task-transport health check warned about unhealthy lease state. The Workers page and its API both returned 200 with no browser JavaScript errors. Waterline's exact image is
sha256:b93f9a0c4b4b1c3e28311538c44bf0c6cd415907755cc0bad1f440c9145c2299; its packaged PHP SDK is 2.1.0. The application worker remains published PHP SDK 2.1.6, against the frozen Server 2.4.30 tuple.Starting the registered application SDK worker on the same queue resumed the original workflow. Container start 08:44:12.199 UTC, workflow completion 08:44:13.326 UTC, about 1.13 seconds later. Final describe confirms run
01m3rqj7ypkq9vh632ccwrydyg, run count one, completed and{"greeting":"hello, Ada"}. History containsRepairRequestedfollowed by one activity completion and one workflow completion. The lease was not manually overwritten, the workflow was not resubmitted and the synthetic claim was not falsely completed.The drill exposed a real Waterline explanation gap: its summary reported four stale registrations but its empty fleet panel claimed no registrations had ever been observed. Fix and published-image repeat are tracked in durable-workflow/waterline#131 and PR durable-workflow/waterline#132. The UI already exposes the expired lease and repair candidate. This result does not establish that a new independent investigator could infer every safe recovery action without help.
Task containers are stopped between phases. Next: finish the UI patch and published repeat, then failed activity, backend interruption and bounded storage pressure. No provider resource was created.
Failed activity with bounded retry: visible and recovered
The published PHP SDK fixture threw
RuntimeException: Synthetic invoice service unavailableon activity attempt one. Its declared policy allowed two attempts, with a 60-second backoff for the first run and a 180-second backoff for the observed waiting-state repeat.Waterline's browser detail page showed the original run waiting / running, the activity pending, attempt #1 / failed, a live worker on the correct queue and 2 attempts / backoff 180s. Expanding
ActivityRetryScheduledexposed the failure message, retryability, attempt limit and exactretry_available_at. The safe action was to let the declared retry happen with the worker and dependency healthy. Repairing or resubmitting the workflow was unnecessary.The observed repeat
op143-retry-wait, run01m3rs32hprhr87rgjd83b10sb, advertised retry availability 09:10:22.654 UTC and completed at 09:10:29.409 UTC. Its final output is{"greeting":"hello, Ada"}and run count one. History contains one retry schedule, two activity starts, one activity completion and one workflow completion. The earlier 60-second case also completed through its original run. These are recorded observations, not a load or retry-latency guarantee.This exercises a retryable failed activity, not exhaustion of all retries or a permanent application error. The browser pages and detail API returned 200 with no JavaScript errors. Published tuple remains Server 2.4.30, application PHP SDK 2.1.6, CLI 2.1.2 and Waterline 2.0.7 with its packaged SDK 2.1.0.
Four initial cases now have recorded diagnosis and original-run recovery: absent SDK worker, wrong queue, expired workflow-task lease and retryable activity failure. Waterline's stale-registration wording fix and exact published-image repeat remain in #131/PR #132. Dependency qualification in #133/PR #134 precedes that patch release. Backend interruption and bounded storage pressure remain the next failure cases.
Published 2.0.8 verified, September 30
Delivered through #134 and #132. Both merged branches were deleted and their absence verified. Final source passed all 27 successful checks and one permitted skip. The merge tree exactly matches the qualified head.
- Composer release 2.0.8 resolves to source
e8fbdec490e2dc9d63b9989f510139f7596fad62. - Published image index:
durableworkflow/waterline:2.0.8@sha256:a413329030cf2b9ad0501b88897ac332062d830eb2cef35d5b5d0657451bccc8. - Linux amd64 manifest:
sha256:b649d51de2192f869deea8e24b34dbf7ded236f684ea570e6b9d923e5b63ca04. - Linux arm64 manifest:
sha256:8b4339e968314b94c6c0549d134f8f8599c84d40b001172dc9794d737f3f3c5a. - Exact-source publication succeeded, including service smoke and native arm64 qualification.
The downloaded image contains Laravel 12.69.3, Flysystem 3.36.0 and the unchanged SDK 2.1.0. Moment is 2.31.0 in the source lock. Published JavaScript matches the qualified rebuilt asset byte for byte, SHA-256
3777166c6b5e27b6cb5c6557e9436a4a24ad4e3500daabe26df39e3cadd2abe0. Ordinary Composer/npm audit gates passed with no advisories or vulnerabilities.Repeated the actual Workers browser view against the published image and Server 2.4.30. It reported seven synthetic registrations, zero active and seven stale, with No active workers, the process/queue/heartbeat recovery instruction and no never-registered claim. Page and observed API requests returned 200, with no JavaScript errors. Source regression also renders the first-registration state for embedded and service snapshots.
Retained sanitized evidence includes browser screenshot/text, API outcomes, actual package versions, image labels and the method. Archive checksum.
This patch changes UI presentation and dependency versions without changing the workflow protocol. Server and Workflow do not lock Waterline as a runtime dependency. The Sample App locks 2.0.7 and needs an embedded consumer update; managed deployment qualification is tracked separately by its owner. Those rollouts are not implied by publishing a package/image.
- Composer release 2.0.8 resolves to source
Backend interruption and bounded storage pressure, September 30
Executed the remaining two synthetic failure cases using published Server 2.4.30 (
sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccf), Workflow 2.2.21, PHP SDK 2.1.6, CLI 2.1.2, Waterline 2.0.8 and MySQL 8.4.5. The MySQL image is pinned atsha256:679e7e924f38a3cbb62a3d7df32924b83f7321a602d3f9f967c01b3df18495d6. These are self-guided local operator cases. They do not constitute an independent blind diagnosis drill or Python/Rust runtime coverage.Backend interruption: discovery and doctor explain the temporarily unavailable backend, readiness returns 503, and Waterline displays an unavailable health panel with a retry action. Restoring the database recovers the original acknowledged run with
hello, Ada, one activity completion and one workflow completion. A second run interrupted while an SDK worker was active also completed after restoration. The SDK worker stayed running with zero restarts. The dispatcher recovered under the publishedunless-stoppedrestart policy. No workflow was resubmitted. The interrupted retry run retains two activity attempts, one retry and one completion.Storage pressure: an isolated MySQL system tablespace capped at 64 MiB accepted 44,040,192 bytes of synthetic filler before reporting error 1114. A previously acknowledged small workflow was retained. A 96-KiB new start failed its workflow-run insert. Readiness remained readable, Server returned generic 500, and CLI replaced that failure with a missing response-envelope error. Both failed start probes left no workflow instance.
Increasing only the tablespace's autoextension maximum to 128 MiB, preserving the volume and initial size, then resuming the worker recovered the original acknowledged workflow. Its final history has exactly one each of StartAccepted, WorkflowStarted, ActivityScheduled, ActivityStarted, ActivityCompleted and WorkflowCompleted, and the expected result. No host disk was filled, no customer namespace was used and no paid host was created. Task containers are stopped between phases.
Confirmed product findings and follow-through:
- Preserve HTTP failures before checking control-plane response envelopes cli#35: HTTP error preservation is implemented and merged in CLI Verify release-plan supersession approval against GitHub authority #36, with normal PHP 8.2–8.5 gates and live control/worker smoke passing. Publication and the exact published storage repeat remain required.
- Explain durable storage exhaustion through the operator API server#284: storage-specific safe diagnosis and database readiness scope are implemented on a draft PR. Local classification, atomic-start, worker-error, connection-loss and readiness regressions pass together: 119 tests, 6,537 assertions. Normal CI and published PHP/Python/Rust error/recovery qualification remain required.
- Sample App Audit conformance experiments for complete first-party SDK coverage #122 adopted published Waterline 2.0.8. Its post-merge CI, embedded and service smoke, published PHP/Python/Rust smoke and devcontainer publication passed.
All six failure types have now been exercised. This parent stays open for the two published diagnostic fixes and repeats, reproducible fixture/procedure handoff and final sanitized evidence retention. Source tests are not delivery evidence for those remaining artifact requirements.
Published diagnostic fixes and affected-case repeats complete
- Preserve HTTP failures before checking control-plane response envelopes cli#35 is delivered in CLI 2.1.3. The original Server 2.4.30 bounded storage failure now reports actual HTTP 500,
Server Errorand exit 6 instead of a missing response-envelope diagnostic. Published native/PHAR, installer and upgrade verification passes. - Explain durable storage exhaustion through the operator API server#284 is published as Server 2.4.31, image index
sha256:8c736663f9c02dc752328b338dc2dd8b73ba1bf7abd0405c36baa25436a66b34. Its API identifies database storage exhaustion and gives a safe recovery action. Readiness explicitly states its connection-only scope. - The new published image passed bounded-MySQL failed-start and failed-completion recovery with PHP 2.1.6, Python 2.3.7 and Rust 2.1.2. All three original acknowledged runs retained their identities and exact results, with one completion each after capacity restoration. CLI 2.1.3 preserved the typed 503 and exit 6. Waterline 2.0.8 showed completed originals during pressure without browser errors. Additional acknowledged starts also survived the next recovery and completed once.
- Reproducible synthetic evidence and procedure, checksum and published recovery guide are retained. The archive's published checksum was verified. Sample App's tuple adoption is in Adopt published CLI 2.1.3 and Server 2.4.31 in the example tuple sample-app#123 / PR Complete Python and Rust worker-affinity parity #124.
All six initial failure types have recorded observations, and the confirmed CLI, Server and Waterline presentation gaps have published fixes and affected-case repeats. The parent remains open for a focused six-case pass against the final published tuple, with reproducible investigator commands, diagnosis and recovery timing, safe-action evidence and final verified state. Earlier intervals interleaved release work and deliberately stopped workers. They are not diagnosis-speed measurements, and the self-guided investigation is not a blind study.
Next action: finish the Sample App consumer, then run the focused timed pass using Server 2.4.31, CLI 2.1.3, Waterline 2.0.8 and the frozen SDK versions. Use only the published operator surfaces for diagnosis, keep staged-fault ground truth separate, and retain the sanitized final report here. Task stacks are stopped between phases. No paid test host was created.
- Preserve HTTP failures before checking control-plane response envelopes cli#35 is delivered in CLI 2.1.3. The original Server 2.4.30 bounded storage failure now reports actual HTTP 500,
Six-case operator drill completed
The final published-artifact pass ran 2026-09-30 13:05:39.535–13:23:18.859 UTC, with exact phase timestamps retained in the evidence. It used Server 2.4.31, Workflow 2.2.21, CLI 2.1.3, Waterline 2.0.8, application PHP SDK 2.1.6 and MySQL 8.4.5. Server's image index is
sha256:8c736663f9c02dc752328b338dc2dd8b73ba1bf7abd0405c36baa25436a66b34. Compose and SDK locks retain the other image identities and limits.Failure Published evidence and safe recovery Diagnosis seconds Recovery through verification seconds Absent worker Pending original, ready task and no active poller. Start its registered worker on the target queue. 46.084 3.973 Wrong queue Ready work on the target queue, a stale previous registration and an active worker on another queue. Correct queue affinity and restart the worker. 107.245 6.543 Expired lease CLI and Waterline identify the expired task, owner, deadline and repair candidate. Resume a compatible worker and let Server reclaim the original task. 36.460 5.167 Failed activity History and Waterline show the exception, two-attempt policy and advertised retry time. Leave the healthy worker running for the scheduled retry. 44.144 62.310 Backend interruption Doctor identifies an unavailable backend, readiness returns 503, Waterline shows unavailable health with Retry. Restore the existing database and inspect/resume the original run. 85.364 109.732 Storage pressure CLI preserves typed storage 503 and exit 6. Readiness states connection-only scope. Restore capacity on the same volume, then resume the original worker. 28.550 30.850 All six original runs retain their accepted run IDs and run count one, expected decoded result, one activity completion and one workflow completion. The activity case retains two activity starts and one retry schedule. The expired lease retains
RepairRequested. The storage start rejected with 503 leaves zero instances. Its previously acknowledged original recovers an exact 98,311-byte result after increasing only the bounded tablespace maximum from 64 to 128 MiB.This is one self-guided pass, with staged ground truth separated from investigator observations. Timings include actual observation delays, heartbeat expiry, scheduled backoff, a corrected browser selector and the dispatcher's recovery from its existing unhealthy state. They are not a blind-investigator or service-latency claim. The investigator's product diagnosis uses the published surfaces. Docker controls stage/recover the synthetic stack, and Docker health/log inspection confirmed the dispatcher transition during recovery. Python/Rust did not execute these six cases. Their separate exact-published storage recovery execution is recorded in Server #284.
Delivered gaps and repeats:
- Waterline Link cross-language child matrix in conformance runbook #131 / Link executed Rust saga experiment in conformance runbook #132: stale registrations receive a useful active-worker explanation, published and repeated in 2.0.8.
- CLI Select release plans by immutable authority, not mutable Release ordering #35 / Verify release-plan supersession approval against GitHub authority #36: non-success HTTP errors survive envelope validation, published and repeated in 2.1.3 against both old generic 500 and new typed storage 503.
- Server #284 / #285: safe capacity diagnostic, readiness scope and recovery guide, published and repeated in 2.4.31 with PHP/Python/Rust.
- Sample App Withdrawn documentation change #123 / Complete Python and Rust worker-affinity parity #124: adopts the new CLI/Server pins. Post-merge application, embedded/service and published polyglot checks pass, and both published devcontainer architectures are qualified.
The reusable operator guide is delivered through #145. Tested head
2e00da951791562c474ea8685ec1af62cb97f230and mergee786fa4605db19f0b1af11ded3904a8903ffef70have equal trees. Post-merge checks pass. The public guide matches the reviewed file, and the merged remote branch is deleted and verified absent.Retained synthetic drill evidence and reproduction fixture includes published CLI/API responses, browser observations, commands, timestamps, staged/final ground truth, probes/locks and the verifier. Checksum was verified after downloading and extracting the public archive. Running its verifier against that downloaded evidence passes all six cases.
All task containers, networks and eight data volumes from the initial and final phases have been removed, with absence independently checked. No customer namespace or paid test host was used. This completes the issue's acceptance criteria.
Customer outcome
A self-hosted operator can tell what happened, what will happen next and which recovery action is safe using only the published CLI, Server API, Waterline and documentation.
Drill
Done when
All six cases can be diagnosed and safely handled from published operator surfaces, with reproducible drill evidence and clear user-facing recovery guidance. This issue coordinates the cross-repository customer outcome; implementation belongs in the repository for each affected surface.