Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions conformance/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,13 @@ Each runner documents its required exact-version environment variables and
result filename in `--help`. Use a unique result directory and isolated Docker
project for every run.

## Operator experience

The [self-hosted operator diagnosis drill](operator-diagnosis.md) covers absent
workers, queue mismatches, expired leases, failed activities, backend interruption
and storage pressure. It records what an operator can observe and which recovery
actions work. Report its scope separately from the stable protocol experiments.

## Evidence scope

A pinned or successfully installed SDK is not evidence that its worker executed
Expand Down
69 changes: 69 additions & 0 deletions conformance/operator-diagnosis.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# Self-hosted operator diagnosis drill

This drill checks whether an operator can understand a stalled workflow and
recover it using published CLI, Server API, Waterline and documentation.
Use an isolated synthetic stack and published artifacts. Record their exact
versions and image digests. Keep the staged fault and reference state separate
from the investigator's observations.

## Start with the original workflow

Use the published CLI to inspect the workflow and the queue it actually targets:

```sh
dw workflow:describe WORKFLOW_ID --json
dw task-queue:describe TASK_QUEUE --json
dw doctor --output=json
dw workflow:history WORKFLOW_ID RUN_ID --output=json
```

Supply the Server endpoint, namespace and operator credential through the
CLI's documented connection options or profile. Inspect `/api/ready` for the
scope and outcome of its checks. In Waterline, use **Workers** for registrations,
queue affinity, heartbeat freshness and leases, then the original run's detail
page for activity attempts, history and pending recovery paths.

Keep the original run ID before acting. After recovery, verify its final status,
expected result and history. A process that is running, a readable dashboard
or a successful database connection alone does not establish that work is
being completed.

## Six cases

| Case | Evidence to inspect | Safe next action and expected behavior |
| --- | --- | --- |
| Absent application worker | Pending run, ready task, no active poller on its queue | Start a compatible worker with the workflow and activity registered on that queue. The original task should be claimed and the same run should complete. |
| Wrong queue | Run's target queue has ready work while a live worker polls another queue | Correct the worker's queue configuration and restart it. Check the advertised heartbeat window before treating a stopped worker's recent registration as current. Recover the original run. |
| Expired lease | Queue detail identifies the task, owner, expired deadline and repair candidate. Waterline shows expired leases. | Restore a compatible worker on the original queue and let the runtime reclaim the expired task. Inspect repair history and completion. Preserve lease fencing and keep activity side effects idempotent. |
| Retryable failed activity | Attempt failure, exception, retry limit and `retry_available_at` in history and Waterline | Restore the dependency and leave a healthy worker running. Allow the declared retry, then verify its attempt and the original result. Normal scheduled retries do not require a replacement workflow. |
| Backend interruption | Doctor/discovery reports an unavailable backend, readiness returns 503, Waterline cannot read health | Restore connectivity to the existing backend with its data intact. Recheck readiness and discovery, inspect the original run and queue, and resume its worker if needed. Inspect the outcome before retrying an ambiguous producer operation. |
| Database storage pressure | Failed write reports `database_storage_exhausted`, the resource and remediation. Readiness may still pass its connection check. | Stop avoidable writers and restore database or temporary-storage capacity. Preserve data and initial tablespace settings. Inspect the original operation before retrying. Verify acknowledged work and ensure rejected starts left no partial instance. |

The [Server storage recovery guide](https://github.com/durable-workflow/server/blob/main/docs/database-storage-recovery.md)
explains the bounded MySQL procedure and the limits of connection-only readiness.
Use the backend's documented recovery procedure for the actual storage system.
A full filler table can coexist with reusable allocations in other tables, so
record the actual failing durable operation rather than assuming every write
must fail.

## Record and repeat

For each case retain:

1. The clean reference, staged fault, exact published artifacts and resource limits.
2. Investigator commands, visible API/UI evidence, diagnosis and expected next behavior.
3. The chosen recovery action, original run identity, final result and history checks.
4. UTC start/end times for diagnosis and for recovery through verified state.
5. Product gaps, their owning fixes and the affected-case repeat after publication.

Count diagnosis from the first investigation command through the recorded
decision. Count recovery from the action through verification. Include policy
waits, heartbeat expiry and inspection delays. Report these as observations from
the recorded environment. A self-guided pass is not a blind investigator study
or a latency guarantee.

If another task interrupts a case, keep its interval separate from focused
timings. Preserve corrections in the work record. Retain sanitized results on
the owning issue or linked artifact, then remove the experiment's containers,
volumes and other resources. This drill does not replace protocol conformance
or qualify unexecuted SDK/runtime combinations.
Loading