Skip to content

[core] Diagnose residual slot-mode CORRUPTED_EVENT_LOG: draws are not stable under log extension - #3543

Draft
VaguelySerious wants to merge 1 commit into
mainfrom
peter/race-repro-diagnostics
Draft

[core] Diagnose residual slot-mode CORRUPTED_EVENT_LOG: draws are not stable under log extension#3543
VaguelySerious wants to merge 1 commit into
mainfrom
peter/race-repro-diagnostics

Conversation

@VaguelySerious

Copy link
Copy Markdown
Member

Draft / diagnostics — not for merge as-is. This PR carries the offline reproduction and instrumentation for the residual CORRUPTED_EVENT_LOG failures on spec-6 (slot-identity) runs, most recently wrun_41KZYJ92TP0GYBNDKW3FJBWQ3Y (step-storm repro on the #3519 preview, 1/14).

What the data shows

Reconstructed from the run's full event log (staging o11y) plus per-invocation runtime logs:

  • The final round's 8th finalizeStep was created under two different correlation ids by two different invocations (…QMWZ eagerly at slot 630, …QMX0 lazily at slot 631), and both executed (duplicate step execution). A third trajectory (the failing replayer) assigned …QMWZ to a releaseStep, which is the decrypted divergence: step event …QMWZ belongs to "finalizeStep", but the current step consumer is "releaseStep".
  • The writer of slot 630 replayed a dense, gap-free 610-event prefix (eventCount: 610 in its logs) and continued correctly. Nothing it did was wrong given what it loaded.

Offline reproduction (in this PR)

storm-log-replay.test.ts rebuilds the run's exact log shape (slot order, entity kinds, step names, ULID ranks remapped onto the test harness's deterministic sequence) and replays it through a faithful port of stepStormReproWorkflow:

  • Replaying the full 655-event log reproduces the production divergence verbatim, at the same event.
  • Replaying the writer's exact 610-event prefix reproduces the writer's committed binding (rank 198 = finalizeStep) as a clean suspension — the writer was prefix-determined.

storm-log-sweep.test.ts (opt-in via STORM_LOG_SWEEP=1) sweeps prefix lengths:

len<=611: rank197=releaseStep rank198=finalizeStep
len>=612: rank197=releaseStep rank198=releaseStep rank199=finalizeStep
len>=630: deterministic ReplayDivergenceError at the rank-198 create

The flip event (slot 612) is an ordinary finalize step_completed. Note also that at len 610 the settled branch's finalize draw (its waking event is at slot 577) lands after draws woken at slots 588–602.

Diagnosis

Replay is deterministic for a byte-identical log, but draw order is not stable under log extension: a branch's post-Promise.race continuation draws its next correlation id at a point in the microtask schedule that depends on how much log is loaded, not at its waking event's log position. Two honest replayers holding different-length (both dense, both valid) snapshots therefore bind the same ordinal to different steps; each commits creates from its own trajectory; the log ends up carrying mutually inconsistent bindings, and every replayer that loads past the conflicting create fails deterministically → 4 recovery replays → CORRUPTED_EVENT_LOG.

This is the property the delivery-barrier work (step-delivery-ordering.test.ts, step-delivery-hop-count.test.ts) pins for adjacent-event shapes; the storm shape (8-wide Promise.race + finally + interleaved recover chains) escapes it. race-padded-draw-ordering.test.ts (also in this PR) shows the minimal 2-branch race shape is correctly ordered cold+warm, so the escape needs the wider interleaving.

Also included: runtime.ts DIAG probes (error-level array-order check before each pass, per-suspension draw-binding log, array-order dump on divergence) so the preview repro lane produces the same forensics without ClickHouse spelunking. The array-order probe has stayed silent locally — the events array is not the problem.

Fix directions (follow-up, not in this PR)

  1. Pin draw order to delivery order: deliver one waking event at a time and let the VM reach quiescence before delivering the next, so every draw is attributable to the delivery that caused it regardless of log length. Replay-latency cost is in-VM microtasks only.
  2. Call-site-derived correlation ids: removes ordinal renaming entirely; a wrong-trajectory writer then produces duplicate/orphan creates instead of conflicting bindings (needs *_created tolerance to become sound).
  3. Currency-checking the fence (412 when the log moved at all) would also close it but serializes fan-out; rejected before (wfs#724 scoping).

🤖 Generated with Claude Code

…e replay of the corrupted storm log + draw-order probes

Offline replay tests built from the actual corrupted event log of
wrun_41KZYJ92TP0GYBNDKW3FJBWQ3Y (step-storm repro, preview, spec 6):

- storm-log-replay.test.ts: a faithful replay of the full 655-event log
  reproduces the production divergence verbatim; a faithful replay of the
  corrupting writer's exact 610-event prefix reproduces the writer's
  committed binding, proving the writer was prefix-determined and the
  binding conflict is created by log growth, not by a misbehaving writer.
- storm-log-sweep.test.ts (STORM_LOG_SWEEP=1): sweeps prefix lengths and
  finds the flip at slot 612 - a branch's post-Promise.race draw is not
  pinned to its waking event's log position, so the ordinal it draws
  depends on how much log is loaded.
- race-padded-draw-ordering.test.ts: the minimal 2-branch race shape stays
  correctly ordered cold+warm (regression coverage for the barrier fix).
- runtime.ts DIAG probes (array order before each pass, draw bindings per
  suspension, array-order dump on divergence) for the preview repro lane.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-bot Bot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 3e060e3

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 0 packages

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
example-nextjs-workflow-turbopack Ready Ready Preview Aug 14, 2026 1:20am
example-nextjs-workflow-webpack Ready Ready Preview Aug 14, 2026 1:20am
example-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-astro-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-express-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-fastify-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-hono-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-nestjs-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-nitro-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-nuxt-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-python-workflow Error Error Aug 14, 2026 1:20am
workbench-sveltekit-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-tanstack-start-workflow Ready Ready Preview Aug 14, 2026 1:20am
workbench-vite-workflow Ready Ready Preview Aug 14, 2026 1:20am
workflow-docs Ready Ready Preview, v0 Aug 14, 2026 1:20am
workflow-swc-playground Ready Ready Preview Aug 14, 2026 1:20am
workflow-tarballs Ready Ready Preview Aug 14, 2026 1:20am
workflow-web Ready Ready Preview Aug 14, 2026 1:20am

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3466 0 590 4056
✅ 💻 Local Development 3673 0 539 4212
✅ 📦 Local Production 3810 0 558 4368
✅ 🐘 Local Postgres 3810 0 558 4368
✅ 🪟 Windows 312 0 0 312
✅ vercel-multi-region 27 0 0 27
Total 15098 0 2245 17343
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 128 0 28
✅ astro-quickjs 128 0 28
✅ example-node 128 0 28
✅ example-quickjs 128 0 28
✅ express-node 128 0 28
✅ express-quickjs 128 0 28
✅ fastify-node 128 0 28
✅ fastify-quickjs 128 0 28
✅ hono-node 128 0 28
✅ hono-quickjs 128 0 28
✅ nest-node 128 0 28
✅ nest-quickjs 128 0 28
✅ nextjs-turbopack-node 153 0 3
✅ nextjs-turbopack-quickjs 153 0 3
✅ nextjs-webpack-node 153 0 3
✅ nextjs-webpack-quickjs 153 0 3
✅ nitro-node 128 0 28
✅ nitro-quickjs 128 0 28
✅ nuxt-node 128 0 28
✅ nuxt-quickjs 128 0 28
✅ sveltekit-node 147 0 9
✅ sveltekit-quickjs 147 0 9
✅ tanstack-start-node 128 0 28
✅ tanstack-start-quickjs 128 0 28
✅ vite-node 128 0 28
✅ vite-quickjs 128 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 156 0 0
✅ nextjs-turbopack-quickjs 156 0 0

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 3e060e3 · Fri, 14 Aug 2026 01:38:46 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1400 (+267%) 🔻 1462 🔴 (+31%) 🔻 1499 🔴 (+32%) 🔻 1651 🔴 (+7.8%) 30
TTFS stream 427 (-58%) 💚 1462 🔴 (+38%) 🔻 1476 🔴 (+38%) 🔻 1516 🔴 (+37%) 🔻 30
TTFS hook + stream 1601 (+25%) 🔻 1729 🔴 (+25%) 🔻 1770 🔴 (+24%) 🔻 7898 🔴 (+388%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 8828 (-1.0%) 10480 (+5.3%) 10486 (+4.0%) 21261 (+57%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 17523 (-0.8%) 19125 (+1.3%) 20601 (+8.4%) 29694 (+27%) 🔻 10
STSO 1020 steps (inline) 123 (±0%) 155 (-19%) 💚 171 (-25%) 💚 229 (-61%) 💚 1019
WO 1020 steps 152558 (-22%) 💚 152558 (-22%) 💚 152558 (-22%) 💚 152558 (-22%) 💚 1
SL stream latency 78 (-1.3%) 121 🔴 (+10%) 165 🔴 (+28%) 🔻 364 🔴 (+6.1%) 30
SO stream overhead (text) 105 (-5.4%) 172 (-4.4%) 284 (+38%) 🔻 8011 🔴 (+1226%) 🔻 30
SO stream overhead (structured) 102 (+6.3%) 161 (+3.2%) 185 (+11%) 301 (+65%) 🔻 30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 194368ms → this run 152318ms (Δ -42050ms, -22%)

  100-150 ms  ███████░░░░░░░░░░░░░░░░┃  main 180  this 661  +481
  150-200 ms  ███████████┃███████████   main 627  this 324  -303
  200-250 ms  ┃████                     main 134  this  28  -106
  250-300 ms  ┃                         main  29  this   4   -25
  300-350 ms  ┃                         main  15  this   2   -13
  350-400 ms  ┃                         main  11  this   0   -11
  400-450 ms  ┃                         main   4  this   0    -4
  450-500 ms  ┃                         main   5  this   0    -5
  500-550 ms  ┃                         main   3  this   0    -3
  550-600 ms  ┃                         main   1  this   0    -1
  600-650 ms  ┃                         main   5  this   0    -5
  650-700 ms  ┃                         main   1  this   0    -1
  750-800 ms  ┃                         main   1  this   0    -1
  800-850 ms  ┃                         main   1  this   0    -1
1100-1150 ms  ┃                         main   1  this   0    -1
4450-4500 ms  ┃                         main   1  this   0    -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision failed 9 1.0m MISMATCH 1
in-flight-before-decision-counted failed 9 1.0m MISMATCH 1
in-flight-after-decision failed 9 1.0m MISMATCH 1
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-append-only.txt

@VaguelySerious

Copy link
Copy Markdown
Member Author

(AI) Local postgres soak addendum: 120 step-storm attempts with the DIAG probes reproduced the same class in 29/120 runs. Each affected run diverges repeatedly at one fixed low slot (58–84, the round-0/1 boundary) while the log keeps growing (e.g. evnt_…059 at eventCounts 114, 145, 249), i.e. a committed binding conflict, not a transient race. The step-entity table shows the signature directly: mint order …F3<F4<F5<F6<F7 (all recoverStep) committed at slots 64, 66, 62, 60, 59 — last-minted first — and the surrounding ordinals are a scrambled interleaving of recover/finalize/release groups from disagreeing trajectories.

Two operational notes from the soak:

  • These runs surface as stuck (divergence-thrash), not CORRUPTED_EVENT_LOG — the storm's extra wakes keep starting fresh deliveries whose divergence count restarts, so the 3-strike exhaustion rarely triggers locally. Stuck-rate is therefore the number to watch locally, not the corrupted-rate.
  • The array-order probe stayed silent across all 120 runs (16k suspension snapshots): the events array is slot-ordered everywhere; the instability is in draw scheduling, not log assembly. Also, the errorMessage field of the DIAG divergence log is dropped by renderStructuredFields (log-format.ts wellKnown handling) — same gap that hides divergence reasons in production logs; worth fixing alongside.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant