Skip to content

Drill self-hosted operator diagnosis across published surfaces #143

Description

@rmcdaniel

Customer outcome

A self-hosted operator can tell what happened, what will happen next and which recovery action is safe using only the published CLI, Server API, Waterline and documentation.

Drill

  • Use published artifacts and an isolated synthetic Server stack. Give an investigator the same access and documentation a self-hosted operator has, without private runbooks or source-only knowledge.
  • Stage six realistic cases: absent worker, wrong task queue, expired lease, failed activity, backend interruption and storage pressure. Preserve a clean reference state and capture the actual underlying fault for later comparison.
  • For each case, record the exact commands or API calls used, visible evidence, diagnosis reached, expected next behavior, chosen recovery action and whether that action was safe and effective. Time the diagnosis and recovery. Verify workflow state after the action.
  • Where a case is ambiguous or misleading, fix the owning CLI, Server API, Waterline or documentation surface, then repeat the case using published artifacts. Link the implementation issues and PRs from this record.

Done when

All six cases can be diagnosed and safely handled from published operator surfaces, with reproducible drill evidence and clear user-facing recovery guidance. This issue coordinates the cross-repository customer outcome; implementation belongs in the repository for each affected surface.

Activity

  1. added
    authority:githubGitHub is the authoritative lifecycle record for this work
    kind:cross-repositoryWork spans more than one public repository
    priority:P2Normal-priority product work
    status:readyReady for implementation
    completion:evidence-requiredClose only after all explicit acceptance and operational evidence is public
    on Sep 27, 2026
  2. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Next qualification tuple

    The activity-outcome work in Workflow #575 is published and verified. The next diagnosis drill can freeze:

    • Server 2.4.30, image index sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccf.
    • Workflow 2.2.21, PHP SDK 2.1.6, Python SDK 2.3.7, Rust SDK 2.1.2, CLI 2.1.2 and Waterline 2.0.7.

    Start with the absent-worker and wrong-queue cases in a fresh isolated synthetic stack. Capture a clean reference state and the staged fault, then record published CLI/API/Waterline observations, diagnosis and recovery timings, the safe action selected and verified workflow state after recovery. The remaining lease, activity failure, backend interruption and storage-pressure cases follow the same evidence format. Each ambiguity gets a fix in its owning repository and a repeat against the published artifact.

    This is the recorded next action. The fault drill has not yet run.

  3. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    First published operator case staged and diagnosed

    Used Server 2.4.30 pinned to sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccf with its documented single-node SQLite bootstrap/dispatch path, and published CLI 2.1.2 dw.phar verified against the release checksum manifest. This is a self-guided drill, not a blind independent-investigator result. Ground truth is recorded separately from the visible diagnostic evidence. No product source or image overlay is used.

    Started one synthetic op143.greeter workflow on op143-orders without an application SDK worker. Published commands used:

    dw workflow:start --type=op143.greeter --task-queue=op143-orders \
      --workflow-id=op143-absent-worker --input='["Ada"]' \
      --execution-timeout=3600 --json
    dw workflow:describe op143-absent-worker --json
    dw task-queue:describe op143-orders --json
    dw doctor --output=json

    Workflow describe reports pending on the expected queue. Queue describe explicitly reports no_active_workers, 0 active workers, 1 ready workflow task, 0 leases and an empty poller list. This supplies an actionable diagnosis from published surfaces: Server accepted the work, and no matching application worker can lease it. The safe next action is to start the application's SDK worker for namespace default and queue op143-orders, then verify the original execution and history complete once. Manual task completion would replace application behavior and is not the recovery action for this case.

    This is diagnostic evidence, not a completed-case pass. Matching-worker recovery, timing and final workflow/history verification remain next, followed by the wrong-queue case and the other four failures. Raw synthetic state is retained for this continuation, and the local stack is stopped between phases. No Cloud runtime or paid host is used.

  4. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Absent SDK worker: recovered and verified with published artifacts

    The CLI identified the stalled workflow on op143-orders, one ready workflow task, zero leases and no active worker. The published PHP SDK worker was then started on that queue with the workflow and activity registered. No history, task or database state was manually completed or repaired.

    The original run 01m3rm6pq6382xq2fec9nmt71x completed once with {"greeting":"hello, Ada"}. History contains one each of StartAccepted, WorkflowStarted, ActivityScheduled, ActivityStarted, ActivityCompleted and WorkflowCompleted. The original run count remains one. The queue subsequently reports one active PHP SDK 2.1.6 worker, zero ready tasks, zero leased tasks and accepting admission.

    The worker container started at 08:22:16.801 UTC, and the workflow closed at 08:22:18.321 UTC, about 1.52 seconds later. The workflow had been intentionally left pending while independent work continued. Its roughly 40-minute total wait is not a diagnosis or service-recovery benchmark. This is a self-guided CLI/API case, not an independent blind investigator result. Waterline remains to be exercised before the full drill is complete.

    Published tuple: Server 2.4.30, sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccf, Workflow 2.2.21, CLI 2.1.2 with release checksum verification, PHP SDK 2.1.6. SQLite is the Server's documented published-image setup for this initial isolated case.

    Commands used to investigate and verify:

    dw workflow:describe op143-absent-worker --json
    dw task-queue:describe op143-orders --json
    dw doctor --output=json
    dw workflow:history op143-absent-worker 01m3rm6pq6382xq2fec9nmt71x --output=json

    Next case is a live worker polling a different queue. The new synthetic run targets op143-orders while that worker polls op143-other. Investigate through the same published surfaces, correct its queue configuration and verify the original run, without restarting or creating a replacement workflow.

  5. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Wrong queue: diagnosed and recovered

    The published CLI showed the pending workflow targeting op143-orders. Its queue had one ready task, zero leases and zero active workers. The previous worker registration was explicitly marked stale. In contrast, op143-other showed one active PHP SDK 2.1.6 worker, zero ready tasks and accepting admission. A healthy worker process therefore did not imply that the workflow's queue was served.

    Recovery was to stop the worker polling op143-other and start the same registered application worker on op143-orders. The original workflow op143-wrong-queue, run 01m3rpt75fjmvm9jzptrge19x9, completed at 08:30:34.924 UTC with {"greeting":"hello, Ada"} and run count one. Its history has one activity completion and one workflow completion. No replacement execution or manual task completion was used.

    This case uses the same exact published tuple as the preceding absent-worker result. The evidence distinguishes configured queue, worker liveness, ready work and active leases. No product defect has been established in these two CLI/API cases. Waterline coverage, expired leases, activity failure, backend interruption and storage pressure remain open. This is still a self-guided drill.

    All task containers are stopped between phases. Next action is to exercise the published Waterline diagnostics alongside a bounded expired-lease failure and verify safe lease recovery through the original run's history.

  6. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Expired workflow-task lease: observed and recovered

    A synthetic diagnostic worker claimed the original task and then stopped sending heartbeats or completion. The task's published lease expired at 08:41:51.068 UTC. The CLI subsequently showed that exact task, owner and expired deadline, one expired lease and one repair candidate on op143-lease.

    The published Waterline 2.0.7 browser UI showed that queue as Needs attention, with 1 leased, 1 expired, one repair candidate and an expired-age counter. Its task-transport health check warned about unhealthy lease state. The Workers page and its API both returned 200 with no browser JavaScript errors. Waterline's exact image is sha256:b93f9a0c4b4b1c3e28311538c44bf0c6cd415907755cc0bad1f440c9145c2299; its packaged PHP SDK is 2.1.0. The application worker remains published PHP SDK 2.1.6, against the frozen Server 2.4.30 tuple.

    Starting the registered application SDK worker on the same queue resumed the original workflow. Container start 08:44:12.199 UTC, workflow completion 08:44:13.326 UTC, about 1.13 seconds later. Final describe confirms run 01m3rqj7ypkq9vh632ccwrydyg, run count one, completed and {"greeting":"hello, Ada"}. History contains RepairRequested followed by one activity completion and one workflow completion. The lease was not manually overwritten, the workflow was not resubmitted and the synthetic claim was not falsely completed.

    The drill exposed a real Waterline explanation gap: its summary reported four stale registrations but its empty fleet panel claimed no registrations had ever been observed. Fix and published-image repeat are tracked in durable-workflow/waterline#131 and PR durable-workflow/waterline#132. The UI already exposes the expired lease and repair candidate. This result does not establish that a new independent investigator could infer every safe recovery action without help.

    Task containers are stopped between phases. Next: finish the UI patch and published repeat, then failed activity, backend interruption and bounded storage pressure. No provider resource was created.

  7. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Failed activity with bounded retry: visible and recovered

    The published PHP SDK fixture threw RuntimeException: Synthetic invoice service unavailable on activity attempt one. Its declared policy allowed two attempts, with a 60-second backoff for the first run and a 180-second backoff for the observed waiting-state repeat.

    Waterline's browser detail page showed the original run waiting / running, the activity pending, attempt #1 / failed, a live worker on the correct queue and 2 attempts / backoff 180s. Expanding ActivityRetryScheduled exposed the failure message, retryability, attempt limit and exact retry_available_at. The safe action was to let the declared retry happen with the worker and dependency healthy. Repairing or resubmitting the workflow was unnecessary.

    The observed repeat op143-retry-wait, run 01m3rs32hprhr87rgjd83b10sb, advertised retry availability 09:10:22.654 UTC and completed at 09:10:29.409 UTC. Its final output is {"greeting":"hello, Ada"} and run count one. History contains one retry schedule, two activity starts, one activity completion and one workflow completion. The earlier 60-second case also completed through its original run. These are recorded observations, not a load or retry-latency guarantee.

    This exercises a retryable failed activity, not exhaustion of all retries or a permanent application error. The browser pages and detail API returned 200 with no JavaScript errors. Published tuple remains Server 2.4.30, application PHP SDK 2.1.6, CLI 2.1.2 and Waterline 2.0.7 with its packaged SDK 2.1.0.

    Four initial cases now have recorded diagnosis and original-run recovery: absent SDK worker, wrong queue, expired workflow-task lease and retryable activity failure. Waterline's stale-registration wording fix and exact published-image repeat remain in #131/PR #132. Dependency qualification in #133/PR #134 precedes that patch release. Backend interruption and bounded storage pressure remain the next failure cases.

  8. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Published 2.0.8 verified, September 30

    Delivered through #134 and #132. Both merged branches were deleted and their absence verified. Final source passed all 27 successful checks and one permitted skip. The merge tree exactly matches the qualified head.

    • Composer release 2.0.8 resolves to source e8fbdec490e2dc9d63b9989f510139f7596fad62.
    • Published image index: durableworkflow/waterline:2.0.8@sha256:a413329030cf2b9ad0501b88897ac332062d830eb2cef35d5b5d0657451bccc8.
    • Linux amd64 manifest: sha256:b649d51de2192f869deea8e24b34dbf7ded236f684ea570e6b9d923e5b63ca04.
    • Linux arm64 manifest: sha256:8b4339e968314b94c6c0549d134f8f8599c84d40b001172dc9794d737f3f3c5a.
    • Exact-source publication succeeded, including service smoke and native arm64 qualification.

    The downloaded image contains Laravel 12.69.3, Flysystem 3.36.0 and the unchanged SDK 2.1.0. Moment is 2.31.0 in the source lock. Published JavaScript matches the qualified rebuilt asset byte for byte, SHA-256 3777166c6b5e27b6cb5c6557e9436a4a24ad4e3500daabe26df39e3cadd2abe0. Ordinary Composer/npm audit gates passed with no advisories or vulnerabilities.

    Repeated the actual Workers browser view against the published image and Server 2.4.30. It reported seven synthetic registrations, zero active and seven stale, with No active workers, the process/queue/heartbeat recovery instruction and no never-registered claim. Page and observed API requests returned 200, with no JavaScript errors. Source regression also renders the first-registration state for embedded and service snapshots.

    Retained sanitized evidence includes browser screenshot/text, API outcomes, actual package versions, image labels and the method. Archive checksum.

    This patch changes UI presentation and dependency versions without changing the workflow protocol. Server and Workflow do not lock Waterline as a runtime dependency. The Sample App locks 2.0.7 and needs an embedded consumer update; managed deployment qualification is tracked separately by its owner. Those rollouts are not implied by publishing a package/image.

  9. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Backend interruption and bounded storage pressure, September 30

    Executed the remaining two synthetic failure cases using published Server 2.4.30 (sha256:dc761d6f99053afc26d1a57e261fbab512eca6d99f934e8eb1b38598fed3cccf), Workflow 2.2.21, PHP SDK 2.1.6, CLI 2.1.2, Waterline 2.0.8 and MySQL 8.4.5. The MySQL image is pinned at sha256:679e7e924f38a3cbb62a3d7df32924b83f7321a602d3f9f967c01b3df18495d6. These are self-guided local operator cases. They do not constitute an independent blind diagnosis drill or Python/Rust runtime coverage.

    Backend interruption: discovery and doctor explain the temporarily unavailable backend, readiness returns 503, and Waterline displays an unavailable health panel with a retry action. Restoring the database recovers the original acknowledged run with hello, Ada, one activity completion and one workflow completion. A second run interrupted while an SDK worker was active also completed after restoration. The SDK worker stayed running with zero restarts. The dispatcher recovered under the published unless-stopped restart policy. No workflow was resubmitted. The interrupted retry run retains two activity attempts, one retry and one completion.

    Storage pressure: an isolated MySQL system tablespace capped at 64 MiB accepted 44,040,192 bytes of synthetic filler before reporting error 1114. A previously acknowledged small workflow was retained. A 96-KiB new start failed its workflow-run insert. Readiness remained readable, Server returned generic 500, and CLI replaced that failure with a missing response-envelope error. Both failed start probes left no workflow instance.

    Increasing only the tablespace's autoextension maximum to 128 MiB, preserving the volume and initial size, then resuming the worker recovered the original acknowledged workflow. Its final history has exactly one each of StartAccepted, WorkflowStarted, ActivityScheduled, ActivityStarted, ActivityCompleted and WorkflowCompleted, and the expected result. No host disk was filled, no customer namespace was used and no paid host was created. Task containers are stopped between phases.

    Confirmed product findings and follow-through:

    All six failure types have now been exercised. This parent stays open for the two published diagnostic fixes and repeats, reproducible fixture/procedure handoff and final sanitized evidence retention. Source tests are not delivery evidence for those remaining artifact requirements.

  10. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Published diagnostic fixes and affected-case repeats complete

    All six initial failure types have recorded observations, and the confirmed CLI, Server and Waterline presentation gaps have published fixes and affected-case repeats. The parent remains open for a focused six-case pass against the final published tuple, with reproducible investigator commands, diagnosis and recovery timing, safe-action evidence and final verified state. Earlier intervals interleaved release work and deliberately stopped workers. They are not diagnosis-speed measurements, and the self-guided investigation is not a blind study.

    Next action: finish the Sample App consumer, then run the focused timed pass using Server 2.4.31, CLI 2.1.3, Waterline 2.0.8 and the frozen SDK versions. Use only the published operator surfaces for diagnosis, keep staged-fault ground truth separate, and retain the sanitized final report here. Task stacks are stopped between phases. No paid test host was created.

  11. rmcdaniel commented on Sep 30, 2026

    @rmcdaniel
    MemberAuthor

    Six-case operator drill completed

    The final published-artifact pass ran 2026-09-30 13:05:39.535–13:23:18.859 UTC, with exact phase timestamps retained in the evidence. It used Server 2.4.31, Workflow 2.2.21, CLI 2.1.3, Waterline 2.0.8, application PHP SDK 2.1.6 and MySQL 8.4.5. Server's image index is sha256:8c736663f9c02dc752328b338dc2dd8b73ba1bf7abd0405c36baa25436a66b34. Compose and SDK locks retain the other image identities and limits.

    Failure Published evidence and safe recovery Diagnosis seconds Recovery through verification seconds
    Absent worker Pending original, ready task and no active poller. Start its registered worker on the target queue. 46.084 3.973
    Wrong queue Ready work on the target queue, a stale previous registration and an active worker on another queue. Correct queue affinity and restart the worker. 107.245 6.543
    Expired lease CLI and Waterline identify the expired task, owner, deadline and repair candidate. Resume a compatible worker and let Server reclaim the original task. 36.460 5.167
    Failed activity History and Waterline show the exception, two-attempt policy and advertised retry time. Leave the healthy worker running for the scheduled retry. 44.144 62.310
    Backend interruption Doctor identifies an unavailable backend, readiness returns 503, Waterline shows unavailable health with Retry. Restore the existing database and inspect/resume the original run. 85.364 109.732
    Storage pressure CLI preserves typed storage 503 and exit 6. Readiness states connection-only scope. Restore capacity on the same volume, then resume the original worker. 28.550 30.850

    All six original runs retain their accepted run IDs and run count one, expected decoded result, one activity completion and one workflow completion. The activity case retains two activity starts and one retry schedule. The expired lease retains RepairRequested. The storage start rejected with 503 leaves zero instances. Its previously acknowledged original recovers an exact 98,311-byte result after increasing only the bounded tablespace maximum from 64 to 128 MiB.

    This is one self-guided pass, with staged ground truth separated from investigator observations. Timings include actual observation delays, heartbeat expiry, scheduled backoff, a corrected browser selector and the dispatcher's recovery from its existing unhealthy state. They are not a blind-investigator or service-latency claim. The investigator's product diagnosis uses the published surfaces. Docker controls stage/recover the synthetic stack, and Docker health/log inspection confirmed the dispatcher transition during recovery. Python/Rust did not execute these six cases. Their separate exact-published storage recovery execution is recorded in Server #284.

    Delivered gaps and repeats:

    The reusable operator guide is delivered through #145. Tested head 2e00da951791562c474ea8685ec1af62cb97f230 and merge e786fa4605db19f0b1af11ded3904a8903ffef70 have equal trees. Post-merge checks pass. The public guide matches the reviewed file, and the merged remote branch is deleted and verified absent.

    Retained synthetic drill evidence and reproduction fixture includes published CLI/API responses, browser observations, commands, timestamps, staged/final ground truth, probes/locks and the verifier. Checksum was verified after downloading and extracting the public archive. Running its verifier against that downloaded evidence passes all six cases.

    All task containers, networks and eight data volumes from the initial and final phases have been removed, with absence independently checked. No customer namespace or paid test host was used. This completes the issue's acceptance criteria.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    authority:githubGitHub is the authoritative lifecycle record for this workcompletion:evidence-requiredClose only after all explicit acceptance and operational evidence is publickind:cross-repositoryWork spans more than one public repositorypriority:P2Normal-priority product workrepo:github-control-planeOwned by the public organization control planestatus:readyReady for implementation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions