[e2e] Shrink fibonacciWorkflow's run tree from fib(6) to fib(5) - #3619
Conversation
|
🧪 E2E Test Results✅ All tests passed 🛠 Infra Events (absorbed by the harness)Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.
E2E Test SummarySummary
Details by Category✅ ▲ Vercel Production
✅ 💻 Local Development
✅ 📦 Local Production
✅ 🐘 Local Postgres
✅ 🪟 Windows
✅ 🌐 Cross-language Conformance
✅ vercel-multi-region
|
📊 Workflow Benchmarkscommit Backend:
Streams
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 253258ms → this run 247440ms (Δ -5818ms, -2%) 📈 CRTT drill-down vs main (RTT distributions & profiles)RTT over stream progress (avg per tenth of stream, bars scaled min→max): RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max): Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max): 📜 Previous results (1)f0b125cTue, 18 Aug 2026 19:46:27 GMT · run logs
Streams
ℹ️ Metric definitions & methodologyStreams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach. The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor. |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 world-sim scenario book — 1 fail of 41 total
Full trace: |
fib(6) spawns a 25-run tree whose ~24 concurrent parent polls saturate the workflow scheduler past the test's 180s budget under a concurrent suite (#2083 measured this as one of the three flake classes blocking e2e concurrency re-enablement). fib(5)'s 15-run tree still exercises the test's actual subject - recursive start() composition with parallel children at every level - with 40% less peak load. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
f0b125c to
e1e2b54
Compare
|
No backport to This is a test-only flake mitigation that shrinks the To override, re-run the Backport to stable workflow manually via |
Serial execution has been the dominant wall-clock cost per matrix entry since concurrency was disabled before conf (78048e0): ~128 tests at ~22 of 24 minutes on the Vercel lanes, and lately the slowest lane cannot finish under its 30-minute job timeout on a slow runner day at all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on the vitest phase) and identified what broke; its blockers are now fixed: world-local writeExclusive is atomic (write-then-link), abort-fetch tests are hermetic (#3618), the fibonacci tree fits the scheduler (#3619), and source-map assertions are positive-only (#3620). What this change adds is concurrency-safe per-test attribution. The harness tracked runs and test names in module globals reset by a beforeEach - under concurrency every test clobbered every other's state, so a failing test dumped an unrelated sibling's diagnostics. vitest's getCurrentTest() cannot substitute: it is a plain module variable, wrong after any await. Instead an auto fixture - the one place that receives the test's own context unambiguously - binds a per-test state (name, tracked runs, the test's own skip) via AsyncLocalStorage around each test body, and trackRun / recordInfraEvent / requireFixture read it ambiently with no call-site changes. The conformance gates skip through the bound state's skip, so a mid-body requireFixture skips the right test. Sequential suites (dev, agent, region) keep setupRunTracking's module-global fallback. Full suite passes 137/137 concurrently against a local dev server in under 2 minutes. A test that genuinely cannot share a deployment can opt out with test.sequential. Builds on VaguelySerious's investigation in #2083. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Serial execution has been the dominant wall-clock cost per matrix entry since concurrency was disabled before conf (78048e0): ~128 tests at ~22 of 24 minutes on the Vercel lanes, and lately the slowest lane cannot finish under its 30-minute job timeout on a slow runner day at all. #2083 measured the concurrent suite at ~3x job wall-clock (4-5x on the vitest phase) and identified what broke; its blockers are now fixed: world-local writeExclusive is atomic (write-then-link), abort-fetch tests are hermetic (#3618), the fibonacci tree fits the scheduler (#3619), and source-map assertions are positive-only (#3620). What this change adds is concurrency-safe per-test attribution. The harness tracked runs and test names in module globals reset by a beforeEach - under concurrency every test clobbered every other's state, so a failing test dumped an unrelated sibling's diagnostics. vitest's getCurrentTest() cannot substitute: it is a plain module variable, wrong after any await. Instead an auto fixture - the one place that receives the test's own context unambiguously - binds a per-test state (name, tracked runs, the test's own skip) via AsyncLocalStorage around each test body, and trackRun / recordInfraEvent / requireFixture read it ambiently with no call-site changes. The conformance gates skip through the bound state's skip, so a mid-body requireFixture skips the right test. Sequential suites (dev, agent, region) keep setupRunTracking's module-global fallback. Full suite passes 137/137 concurrently against a local dev server in under 2 minutes. A test that genuinely cannot share a deployment can opt out with test.sequential. Builds on VaguelySerious's investigation in #2083. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Summary & Motivation
fib(6)'s 25-run tree saturates the scheduler past this test's 180s budget when the e2e suite runs concurrently (#2083). fib(5) keeps the subject — recursive
start()composition with parallel children at every level — at 15 runs.Test Plan
Existing coverage: the modified test runs in CI across the workbench matrix.