Skip to content

[sim-world] Repro: Date.now() after Promise.race is not deterministic - #3583

Draft
shalabhc wants to merge 5 commits into
mainfrom
sleep-resumeat-repro
Draft

[sim-world] Repro: Date.now() after Promise.race is not deterministic#3583
shalabhc wants to merge 5 commits into
mainfrom
sleep-resumeat-repro

Conversation

@shalabhc

@shalabhc shalabhc commented Aug 17, 2026

Copy link
Copy Markdown
Collaborator

Summary

Repro only: this PR adds one sim-world scenario proving that Date.now() immediately after Promise.race is not deterministic across workflow passes, even with:

  • append-only event positions,
  • complete, strongly-consistent event-log reads,
  • stable Promise.race winner ordering,
  • no precondition fence or missing event.

It removes the earlier sleep() scenarios. Those scenarios combined append-only positioning with the old stale-write 412 fence, but slot-identity writes do not use that fence, so they were not valid evidence for the production mode in scope.

Repro

The workflow races a step against an out-of-band hook, then reads Date.now() and branches around a fixed cutoff:

const winner = await Promise.race([
  fast(documentId).then(() => 'fast'),
  hook.then(() => 'hook'),
]);

const beforeCutoff = Date.now() < cutoff;
await sleep(beforeCutoff ? '2m' : '3m');
return beforeCutoff
  ? await afterClockBeforeCutoff(documentId)
  : await afterClockAfterCutoff(documentId);

The scenario scripts this append-only history:

  1. The step completes at +0ms and wins the race.
  2. The first pass reads the VM clock at +0ms, selects the before-cutoff branch, and proposes a two-minute wait_created. The scenario holds that write before commit.
  3. Virtual time advances one minute and the losing hook payload commits.
  4. The held wait commits after the hook, so commit-time slot order is step_completed, hook_received, wait_created.
  5. The hook extends the log and causes a new pass. EventsConsumer synchronously consumes both the step completion and hook payload before promise continuations run.
  6. Delivery barriers still make the step win Promise.race, but event consumption has already advanced the VM clock to the hook's +1m timestamp.
  7. The new pass selects the after-cutoff branch. The run completes there, and the final log cold-replays cleanly because it records only the later answer.

The scenario has two independent observations:

ok   the first pass read the early clock and selected the two-minute wait
FAIL every pass selected the same Date.now() branch

That isolates the bug: the race winner is deterministic, while the clock value observed immediately after the race depends on how much of the committed log that pass consumed.

Why delivery ordering does not fix it

Delivery barriers order when resolved values reach workflow code. The VM clock advances earlier, in EventsConsumer.onConsumedEvent, while the log is synchronously drained. By the time barriers deliver the stable winner, Date.now() already reflects the newest consumed event, including a losing competitor that was absent from the earlier pass.

Result

The live run completes and the final log passes cold replay. The red signal is the scenario's cross-pass check, because the durable log contains only the later branch and cannot recover the first pass's clock observation.

No runtime fix is included.

Verification

pnpm sim clock-after-race --append-only --report-only --no-color
  1 scenario: 0 passed, 1 failed as expected
  outcome=completed, replay=ok, log=append-only

pnpm --filter @workflow/world-sim test
  8 files passed, 72 tests passed

pnpm exec tsc --noEmit -p workbench/sim-world/tsconfig.json
pnpm exec biome check workbench/sim-world/scenarios/clock-after-race.ts \
  workbench/sim-world/scenarios/index.ts \
  workbench/sim-world/workflows/index.ts

@changeset-bot

changeset-bot Bot commented Aug 17, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: ea54074

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

github-actions Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 18b3b71 · Mon, 17 Aug 2026 16:52:50 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 208 (-8.4%) 1414 🔴 (+23%) 🔻 1501 🔴 (+24%) 🔻 1796 🔴 (+33%) 🔻 30
TTFS stream 214 (-8.9%) 1432 🔴 (+20%) 🔻 1573 🔴 (+24%) 🔻 1596 🔴 (+18%) 🔻 30
TTFS hook + stream 534 (+28%) 🔻 1620 🔴 (+14%) 1647 🔴 (+13%) 1857 🔴 (+20%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 9466 (+613%) 🔻 11066 (+330%) 🔻 11201 (+314%) 🔻 14574 (+355%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 18898 (+74%) 🔻 23289 (-79%) 💚 24851 (-87%) 💚 26726 (-86%) 💚 10
STSO 1020 steps (inline) 142 (-7.8%) 489 (-1.8%) 532 (-4.8%) 719 (-8.5%) 1018
STSO 1020 steps (queue-hop) 3364 3364 3364 3364 1
WO 1020 steps 438586 (+5.8%) 438586 (+5.8%) 438586 (+5.8%) 438586 (+5.8%) 1
CRTT first chunk (pooled) 105 (+8.2%) 175 (+13%) 200 (-16%) 💚 391 (+35%) 🔻 28

Streams

Scenario wr c/s rd c/s wr KiB/s rd KiB/s CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 100 (±0%) 99.6 (-2%) 5 (±0%) 5 (-2%) 139 (+3%) 270 (-9%) 412 (-59%) 1116 (-25%) 269 (+29%) 10
size sweep (100/s, 160B-12KB) 100 (±0%) 100 (+1%) 334 (±0%) 334 (+1%) 127 (+9%) 336 (-2%) 389 (-49%) 1024 (-19%) 683 (+119%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 89.2 (±0%) 89.8 (±0%) 16.2 (±0%) 16.3 (±0%) 182 (+82%) 158 (+13%) 219 (+13%) 831 (-5%) 707 (+95%) 3
replay eve-gpt-5.6-sol-2000t (1x) 54.7 (±0%) 54.8 (±0%) 355 (±0%) 356 (±0%) 157 (-6%) 169 (-5%) 236 (-72%) 567 (-84%) 450 (-77%) 2
replay eve-gpt-5.6-sol-2000t (2x) 109 (±0%) 109 (+1%) 710 (±0%) 710 (±0%) 124 (-2%) 260 (-9%) 374 (-45%) 1089 (-44%) 427 (-58%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 413585ms → this run 433700ms (Δ +20115ms, +5%)

  100-150 ms  ┃                         main   0  this   2    +2
  150-200 ms  ┃██████                   main  69  this  14   -55
  200-250 ms  ███┃███████               main 102  this  34   -68
  250-300 ms  ████┃                     main  52  this  46    -6
  300-350 ms  █████████████┃██          main 154  this 139   -15
  350-400 ms  ██████████████┃█████      main 193  this 144   -49
  400-450 ms  ████████░░░░░░░░░░░░┃     main  80  this 200  +120
  450-500 ms  ████████████░░░░░░░░░░░┃  main 118  this 231  +113
  500-550 ms  █████████████┃            main 135  this 133    -2
  550-600 ms  ███┃██                    main  61  this  37   -24
  600-650 ms  █┃                        main  19  this  22    +3
  650-700 ms  ┃                         main  13  this   4    -9
  700-750 ms  ┃                         main   9  this   4    -5
  750-800 ms  ┃                         main   4  this   4    +0
  800-850 ms  ┃                         main   3  this   1    -2
  850-900 ms  ┃                         main   2  this   0    -2
  900-950 ms  ┃                         main   0  this   1    +1
 950-1000 ms  ┃                         main   1  this   0    -1
1050-1100 ms  ┃                         main   0  this   1    +1
1200-1250 ms  ┃                         main   1  this   0    -1
1350-1400 ms  ┃                         main   0  this   1    +1
3150-3200 ms  ┃                         main   1  this   0    -1
3600-3650 ms  ┃                         main   1  this   0    -1
3650-3700 ms  ┃                         main   1  this   0    -1

1020 steps (queue-hop)

Cumulative STSO time: 3364ms over 1 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

3000-3500 ms  ████████████████████████  steps 1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90          p99     n
control  ······▂█▅▁▁··  202.6 (-17%)   163 (+7%)  412 (-59%)  1116 (-25%)  3000
sweep    ······▂█▅▁▁··   211.7 (+1%)  152 (+16%)  389 (-49%)  1024 (-19%)  3000
gw 1x    ·····▁▃█▂▁···  147.1 (+13%)  126 (+10%)  219 (+13%)    831 (-5%)  5295
eve 1x   ·····▁▄█▂▁···  140.7 (-45%)   118 (+2%)  236 (-72%)   567 (-84%)  5186
eve 2x   ·····▁▁█▅▁▁··  200.3 (-26%)   167 (-4%)  374 (-45%)  1089 (-44%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▁▂▅▄█▅▂▁▅▁  173–257ms
sweep    ▃▇▄▇▇█▂▆▃▁  185–233ms
gw 1x    █▅▂▂▁▂▁▂▂▁  120–238ms
eve 1x   █▂▁▄█▇▇▇▅▄  117–156ms
eve 2x   ▇█▁▁▁▂▄▇▃▁  170–254ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  █▄▅▁▄▅▇  207–216ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▁█▃▅▅▅▄▅▄▄  34–71ms
sweep    ▄█▄▃▆▃▁▇▆▃  43–67ms
gw 1x    █▇▃▁▂▆▃▆▃▂  36–48ms
eve 1x   ▅▂▁▃▇▇▇▄▅█  23–34ms
eve 2x   █▇▂▅▇▃▃▁▂▂  25–41ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • hookWithSleepWorkflow - hook payloads delivered correctly with concurrent sleep (nextjs-webpack)
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 18:10:59Z · abandoned wrun_01M08EPC6S84ZPZK4Y5CEZQ22S

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3321 0 735 4056
✅ 💻 Local Development 3498 0 558 4056
✅ 📦 Local Production 3810 0 558 4368
✅ 🐘 Local Postgres 3810 0 558 4368
✅ 🪟 Windows 312 0 0 312
✅ 🌐 Cross-language Conformance 9 0 128 137
✅ vercel-multi-region 27 0 0 27
Total 14787 0 2537 17324
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 128 0 28
✅ astro-quickjs 128 0 28
✅ example-node 128 0 28
✅ example-quickjs 128 0 28
✅ express-node 128 0 28
✅ express-quickjs 128 0 28
✅ fastify-node 128 0 28
✅ fastify-quickjs 128 0 28
✅ hono-node 128 0 28
✅ hono-quickjs 128 0 28
✅ nest-node 128 0 28
✅ nest-quickjs 128 0 28
✅ nextjs-turbopack-node 153 0 3
✅ nextjs-webpack-node 153 0 3
✅ nextjs-webpack-quickjs 153 0 3
✅ nitro-node 128 0 28
✅ nitro-quickjs 128 0 28
✅ nuxt-node 128 0 28
✅ nuxt-quickjs 128 0 28
✅ python-node 8 0 148
✅ sveltekit-node 147 0 9
✅ sveltekit-quickjs 147 0 9
✅ tanstack-start-node 128 0 28
✅ tanstack-start-quickjs 128 0 28
✅ vite-node 128 0 28
✅ vite-quickjs 128 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 156 0 0
✅ nextjs-turbopack-quickjs 156 0 0

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 9 0 128

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 2 fail of 42 total

fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
clock-after-race completed 14 2.0m ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@vercel

vercel Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
example-nextjs-workflow-turbopack Ready Ready Preview Aug 17, 2026 6:13pm
example-nextjs-workflow-webpack Ready Ready Preview Aug 17, 2026 6:13pm
example-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-astro-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-express-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-fastify-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-hono-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-nestjs-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-nitro-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-nuxt-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-python-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-sveltekit-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-tanstack-start-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workbench-vite-workflow Ready Ready Preview Aug 17, 2026 6:13pm
workflow-docs Ready Ready Preview, v0 Aug 17, 2026 6:13pm
workflow-swc-playground Ready Ready Preview Aug 17, 2026 6:13pm
workflow-tarballs Ready Ready Preview Aug 17, 2026 6:13pm
workflow-web Ready Ready Preview Aug 17, 2026 6:13pm

@vercel
vercel Bot temporarily deployed to Preview – workflow-docs August 17, 2026 16:27 Inactive
@vercel
vercel Bot temporarily deployed to Preview – workflow-docs August 17, 2026 17:17 Inactive
vercel Bot and others added 5 commits August 17, 2026 18:08
`sleep(param)` resolves a duration to an instant by reading the *host* wall
clock: `createSleep` is a host closure installed on the VM global, so
`parseDurationToDate` runs in host scope and never sees the VM's
deterministic clock. The conversion happens where the body reaches the
`sleep()` call, so every pass that gets there without having consumed a
`wait_created` for that wait resolves the duration afresh. Nothing pins the
passes to each other.

Two scenarios, both red, for the two things that costs:

- `sleep-resumeat-recomputed` — the drift. The orchestrator is held inside
  its `wait_created` write, a bystander hook commits behind it so the held
  write's snapshot goes stale, the fence rejects it, and the restarted pass
  reads a clock that has moved. Pass A asks for +30m, pass B for +31m30s,
  and pass B's value is the one that commits because pass A's never landed.
  The run completes and its log replays clean — the committed timestamp is
  self-consistent and nothing in the log records what was originally asked
  for, so no oracle over the log can see that the answer moved.

- `sleep-wait-continuation-stranded` — the same recompute reached from the
  other side, and it costs the run. A wait continuation whose read is
  missing its own `wait_created` replays the body, recomputes the deadline,
  re-creates the wait (rejected as a duplicate, and the suspension writer
  swallows the conflict), then dispatches a continuation under the same
  idempotency key as the delivery it is running inside. The queue dedupes it
  against that in-flight key; the message settles; nothing is queued. The
  wait stays open and the run never wakes up.

The 409 guard the design counts on protects the *log* — one `wait_created`,
one `resumeAt`, replays clean. It does not protect the *run*: swallowing the
conflict and carrying on is what walks into the deduped re-dispatch.

Both fail identically with `--append-only`, which is the point: assigning
log positions at commit fixes six position-ordering scenarios and has
nothing to say about a client-side clock read.

Adds `sleepWithBystanderHookWorkflow` — a plain sleep with an open hook
nobody awaits, so a scenario has an out-of-band writer it can fire without
changing which branch runs.

Book: 38/3/3 -> 38/5/3 mint-ordered, 41/0/0 -> 41/2/0 append-only. No
existing scenario changes colour.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
`sleep-resumeat-recomputed` shows a pass resolving `sleep()` against a moved
clock. Read alone it invites a broader conclusion than the code supports —
that every replay re-reads the clock, so a run could never finish a sleep.

Measure it instead. `sleep-replay-inherits-resumeat` commits the wait, moves
the host clock, and forces a full replay of the body mid-sleep. The replay
does re-read the clock — `createSleep` consults no log first — but the wait
consumer then overwrites `queueItem.resumeAt` with the logged
`wait_created`'s value and sets `hasCreatedEvent`, so the suspension handler
neither re-writes the wait nor counts down from the fresh read. The wait
fires at the originally committed instant, 0ms late.

So the exposure is only a pass reaching `sleep()` with no visible
`wait_created`: the first pass, or one whose predecessor's write never
landed. That same overwrite is why the `ReplayDivergenceError` in `sleep.ts`
is unreachable — the guard passing is the evidence the overwrite happened.

Also name which clock, since "not deterministic" does not. Separating the
two candidates by 90s of idle time, the drifting pass tracks
`hostClock + 30m` off by 0ms while `newestEventCreatedAt + 30m` is off by
90000ms. Recorded as notes rather than checks: both state present behaviour,
so asserting them would invert on the fix.

Book: 44 scenarios, 39/5/3 default and 42/2/0 append-only. The control
passes in both worlds — none of this turns on log-position policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
`sleep-resumeat-recomputed` reads as a story about downtime, because the
Slack thread it answers is one. It is not. The trigger for re-reading the
clock is the optimistic-concurrency fence, which fires whenever anything
commits behind an in-flight suspension write — an out-of-band hook payload,
another branch's step result, an attribute write. That is what a busy run
looks like, not a failure mode, and there is no outage anywhere in the
existing repro either: its rejection is a `PreconditionFailedError` raised by
a bystander hook.

`sleep-resumeat-drift-compounds` makes that explicit and adds the part that
changes the risk assessment. Two bystander payloads, each landing behind a
held `wait_created`, give two fence rejections and so three resolutions of
one `sleep('30m')`: 00:30:00, then 00:30:45, then 00:31:30, and the third is
what commits. The drift is the sum of the restarts, so it scales with the
concurrency the run is under rather than being a one-off a retry settles.
The runtime is behaving correctly at every step — it rejects a stale write
and restarts the replay; what it cannot do is carry the deadline across the
restart, because nothing ever wrote the deadline down.

The 45s gaps are inflated for legibility, not a claim about magnitude. In
production the drift is the restart latency; what matters is the ratio, and
the absolute drift that is noise against `sleep('30m')` is the entire
interval for the short sleeps used as poll intervals and race timeouts.

Book: 45 scenarios, 39/6/3 default and 42/3/0 append-only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com>
@shalabhc
shalabhc force-pushed the sleep-resumeat-repro branch from 45493c0 to ea54074 Compare August 17, 2026 18:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant