Skip to content

[pull] main from oven-sh:main - #11

Merged
pull[bot] merged 5 commits into
Mu-L:mainfrom
oven-sh:main
Mar 29, 2025
Merged

[pull] main from oven-sh:main#11
pull[bot] merged 5 commits into
Mu-L:mainfrom
oven-sh:main

Conversation

@pull

@pull pull Bot commented Mar 29, 2025

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.1)

Can you help keep this open source service alive? 💖 Please sponsor : )

@pull pull Bot added the ⤵️ pull label Mar 29, 2025
@pull
pull Bot merged commit bb9128c into Mu-L:main Mar 29, 2025
pull Bot pushed a commit that referenced this pull request Jul 25, 2025
…ck traces upon crash in CI (oven-sh#21143)

### What does this PR do?

Closes oven-sh#13012

On Linux, when any Bun process spawned by `runner.node.mjs` crashes, we
run GDB in batch mode to print a backtrace from the core file.

And on all platforms, we run a mini `bun.report` server which collects
crashes reported by any Bun process executed during the tests, and after
each test `runner.node.mjs` fetches and prints any new crashes from the
server.

<details>
<summary>example 1</summary>

```
#0  crash_handler.crash () at crash_handler.zig:1513
#1  0x0000000002cf4020 in crash_handler.crashHandler (reason=..., error_return_trace=0x0, begin_addr=...) at crash_handler.zig:479
#2  0x0000000002cefe25 in crash_handler.handleSegfaultPosix (sig=<optimized out>, info=<optimized out>) at crash_handler.zig:800
#3  0x00000000045a1124 in WTF::jscSignalHandler (sig=11, info=0x7ffe044e30b0, ucontext=0x0) at vendor/WebKit/Source/WTF/wtf/threads/Signals.cpp:548
#4  <signal handler called>
#5  JSC::JSCell::type (this=0x0) at vendor/WebKit/Source/JavaScriptCore/runtime/JSCellInlines.h:137
#6  JSC::JSObject::getOwnNonIndexPropertySlot (this=0x150bc914fe18, vm=..., structure=0x150a0102de50, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSObject.h:1348
#7  JSC::JSObject::getPropertySlot<false> (this=0x150bc914fe18, globalObject=0x150b864e0088, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSObject.h:1433
#8  JSC::JSValue::getPropertySlot (this=0x7ffe044e4880, globalObject=0x150b864e0088, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSCJSValueInlines.h:1108
#9  JSC::JSValue::get (this=0x7ffe044e4880, globalObject=0x150b864e0088, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSCJSValueInlines.h:1065
#10 JSC::LLInt::performLLIntGetByID (bytecodeIndex=..., codeBlock=0x150b861e7740, globalObject=0x150b864e0088, baseValue=..., ident=..., metadata=...) at vendor/WebKit/Source/JavaScriptCore/llint/LLIntSlowPaths.cpp:878
#11 0x0000000004d7b055 in llint_slow_path_get_by_id (callFrame=0x7ffe044e4ab0, pc=0x150bc92ea0e7) at vendor/WebKit/Source/JavaScriptCore/llint/LLIntSlowPaths.cpp:946
#12 0x0000000003dd6042 in llint_op_get_by_id ()
#13 0x0000000000000000 in ?? ()
```

</details>

<details>
<summary>example 2</summary>

```
  #0  crash_handler.crash () at crash_handler.zig:1513
  #1  0x0000000002c5db80 in crash_handler.crashHandler (reason=..., error_return_trace=0x0, begin_addr=...) at crash_handler.zig:479
  #2  0x0000000002c59f60 in crash_handler.handleSegfaultPosix (sig=<optimized out>, info=<optimized out>) at crash_handler.zig:800
  #3  0x00000000042ecc88 in WTF::jscSignalHandler (sig=11, info=0xfffff60141b0, ucontext=0xfffff6014230) at vendor/WebKit/Source/WTF/wtf/threads/Signals.cpp:548
  #4  <signal handler called>
  #5  bun.js.api.FFIObject.Reader.u8 (globalObject=0x4000554e0088) at /var/lib/buildkite-agent/builds/ip-172-31-75-92/bun/bun/src/bun.js/api/FFIObject.zig:65
  #6  bun.js.jsc.host_fn.toJSHostCall__anon_1711576 (globalThis=0x4000554e0088, args=...) at /var/lib/buildkite-agent/builds/ip-172-31-75-92/bun/bun/src/bun.js/jsc/host_fn.zig:97
  #7  bun.js.jsc.host_fn.DOMCall("Reader"[0..6],bun.js.api.FFIObject.Reader,"u8"[0..2],.{ .reads = .{ ... }, .writes = .{ ... } }).slowpath (globalObject=0x4000554e0088, thisValue=70370172175040, arguments_ptr=0xfffff6015460, arguments_len=1) at /var/lib/buildkite-agent/builds/ip-172-31-75-92/bun/bun/src/bun.js/jsc/host_fn.zig:490
  #8  0x000040003419003c in ?? ()
  #9  0x0000400055173440 in ?? ()
```

</details>

I used GDB instead of LLDB (as the branch name suggests) because it
seems to produce more useful stack traces with musl libc.

- [x] on linux, use gdb to print from core dump of main bun process
crashed
- [x] on linux, use gdb to print from all new core dumps (so including
bun subprocesses spawned by the test that crashed)
- [x] on all platforms, use a mini bun.report server to print a
self-reported trace (depends on oven-sh/bun.report#15; for now our
package.json points to a commit on the branch of that repo)
- [x] fix trying to fetch stack traces too early on windows
- [x] use output groups so the traces show up alongside the log for the
specific test instead of having to find it in the logs from the entire
run
- [x] get oven-sh/bun.report#15 merged, and point to a bun.report commit
on the main branch instead of the PR branch in package.json

### How did you verify your code works?

Manually, and in CI with a crashing test.

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot pushed a commit that referenced this pull request Aug 8, 2025
<details>

<summary> observed in
https://buildkite.com/bun/bun/builds/22442#annotation-test/js/node/zlib/leak.test.ts
</summary>

```
==5045==ERROR: AddressSanitizer: heap-use-after-free on address 0x5220000243c0 at pc 0x00000dad671b bp 0x14f22d4a4990 sp 0x14f22d4a4988
READ of size 8 at 0x5220000243c0 thread T5 (HeapHelper)
======== Stack trace from GDB for HeapHelper-5045.core: ========
Program terminated with signal SIGABRT, Aborted.
#0  0x000014f2c3672eec in ?? () from /lib/x86_64-linux-gnu/libc.so.6
[Current thread is 1 (Thread 0x14f22d4f46c0 (LWP 5050))]
#0  0x000014f2c3672eec in ?? () from /lib/x86_64-linux-gnu/libc.so.6
#1  0x000014f2c3623fb2 in raise () from /lib/x86_64-linux-gnu/libc.so.6
#2  0x000014f2c360e472 in abort () from /lib/x86_64-linux-gnu/libc.so.6
#3  0x000000000e3b2ae2 in uw_init_context_1[cold] ()
#4  0x000000000e3b29fc in _Unwind_Backtrace ()
#5  0x00000000046a6bab in __sanitizer::BufferedStackTrace::UnwindSlow(unsigned long, unsigned int) ()
#6  0x00000000046a181d in __sanitizer::BufferedStackTrace::Unwind(unsigned int, unsigned long, unsigned long, void*, unsigned long, unsigned long, bool) ()
#7  0x00000000046885bd in __sanitizer::BufferedStackTrace::UnwindImpl(unsigned long, unsigned long, void*, bool, unsigned int) ()
#8  0x0000000004601127 in __asan::ErrorGeneric::Print() ()
#9  0x0000000004683180 in __asan::ScopedInErrorReport::~ScopedInErrorReport() ()
#10 0x0000000004686567 in __asan::ReportGenericError(unsigned long, unsigned long, unsigned long, unsigned long, bool, unsigned long, unsigned int, bool) ()
#11 0x0000000004686d46 in __asan_report_load8 ()
#12 0x000000000dad671b in ZSTD_sizeof_CCtx (cctx=<optimized out>) at ./build/release-asan/zstd/vendor/zstd/lib/compress/zstd_compress.c:210
#13 0x0000000006d2284d in bun.js.node.zlib.NativeZstd.estimatedSize () at /var/lib/buildkite-agent/builds/ip-172-31-72-121/bun/bun/src/bun.js/node/zlib/NativeZstd.zig:57
#14 ZigGeneratedClasses.JSNativeZstd.JavaScriptCoreBindings.NativeZstd__estimatedSize (thisValue=<optimized out>) at /var/lib/buildkite-agent/builds/ip-172-31-72-121/bun/bun/build/release-asan/codegen/ZigGeneratedClasses.zig:11122
#15 0x000000000852803b in WebCore::JSNativeZstd::visitChildrenImpl<JSC::SlotVisitor> (cell=0x14f22e190840, visitor=...) at ./build/release-asan/./build/release-asan/codegen/ZigGeneratedClasses.cpp:30728
#16 WebCore::JSNativeZstd::visitChildren (cell=0x14f22e190840, visitor=...) at ./build/release-asan/./build/release-asan/codegen/ZigGeneratedClasses.cpp:30734
#17 0x000000000aa99d6c in JSC::MethodTable::visitChildren (this=<optimized out>, cell=<optimized out>, visitor=...) at vendor/WebKit/Source/JavaScriptCore/runtime/ClassInfo.h:115
#18 0x000000000aa99d6c in JSC::SlotVisitor::visitChildren (this=0x14f277028300, cell=0x14f22e190840)
#19 JSC::SlotVisitor::drain(WTF::MonotonicTime)::$_0::operator()(JSC::MarkStackArray&) const (this=<optimized out>, stack=...) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitor.cpp:509
#20 0x000000000aa8f130 in JSC::SlotVisitor::forEachMarkStack<JSC::SlotVisitor::drain(WTF::MonotonicTime)::$_0>(JSC::SlotVisitor::drain(WTF::MonotonicTime)::$_0 const&) (this=0x14f277028300, func=...) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitorInlines.h:193
#21 JSC::SlotVisitor::drain (this=this@entry=0x14f277028300, timeout=<error reading variable: That operation is not available on integers of more than 8 bytes.>, timeout@entry=...) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitor.cpp:499
#22 0x000000000aa90590 in JSC::SlotVisitor::drainFromShared (this=0x14f277028300, sharedDrainMode=JSC::SlotVisitor::HelperDrain, timeout=<error reading variable: That operation is not available on integers of more than 8 bytes.>) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitor.cpp:699
#23 0x000000000aa08726 in JSC::Heap::runBeginPhase(JSC::GCConductor)::$_1::operator()() const (this=<optimized out>) at vendor/WebKit/Source/JavaScriptCore/heap/Heap.cpp:1508
#24 WTF::SharedTaskFunctor<void (), JSC::Heap::runBeginPhase(JSC::GCConductor)::$_1>::run() (this=<optimized out>) at .WTF/Headers/wtf/SharedTask.h:91
#25 0x000000000aa3b596 in WTF::ParallelHelperClient::runTask(WTF::RefPtr<WTF::SharedTask<void ()>, WTF::RawPtrTraits<WTF::SharedTask<void ()> >, WTF::DefaultRefDerefTraits<WTF::SharedTask<void ()> > > const&) (this=0x14f22e000428, task=...) at vendor/WebKit/Source/WTF/wtf/ParallelHelperPool.cpp:110
#26 0x000000000aa3d976 in WTF::ParallelHelperPool::Thread::work (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/ParallelHelperPool.cpp:201
#27 0x000000000aa4210d in WTF::AutomaticThread::start(WTF::AbstractLocker const&)::$_0::operator()() const (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/AutomaticThread.cpp:225
#28 WTF::Detail::CallableWrapper<WTF::AutomaticThread::start(WTF::AbstractLocker const&)::$_0, void>::call() (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/Function.h:53
#29 0x0000000008958ada in WTF::Function<void ()>::operator()() const (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/Function.h:82
#30 WTF::Thread::entryPoint (newThreadContext=<optimized out>) at vendor/WebKit/Source/WTF/wtf/Threading.cpp:272
#31 0x0000000008a65689 in WTF::wtfThreadEntryPoint (context=0x13b5) at vendor/WebKit/Source/WTF/wtf/posix/ThreadingPOSIX.cpp:255
#32 0x000000000467d347 in asan_thread_start(void*) ()
#33 0x000014f2c36711f5 in ?? () from /lib/x86_64-linux-gnu/libc.so.6
#34 0x000014f2c36f189c in ?? () from /lib/x86_64-linux-gnu/libc.so.6
```

</details>

`ZSTD_sizeof_CCtx` and `ZSTD_sizeof_DCtx` can not be relied upon to be
thread-safe and estimatedSize may be called from any thread
pull Bot pushed a commit that referenced this pull request Aug 20, 2025
…Worker" (oven-sh#21994)

Reverts oven-sh#21962

`vm.ensureTerminationException` allocates a JSString, which is not safe
to do from a thread that doesn't own the API lock.

```ts
Bun Canary v1.2.21-canary.1 (f706382a) Linux x64 (baseline)
Linux Kernel v6.12.38 | musl
CPU: sse42 popcnt avx avx2 avx512
Args: "/var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/release/bun-linux-x64-musl-baseline-profile/bun-profile" "/var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/test/js/node/worker_threads"...
Features: bunfig http_server jsc tsconfig(3) tsconfig_paths workers_spawned(40) workers_terminated(34)
Builtins: "bun:main" "node:worker_threads"
Elapsed: 362ms | User: 518ms | Sys: 63ms
RSS: 0.34GB | Peak: 100.36MB | Commit: 0.34GB | Faults: 0 | Machine: 8.17GB
 
panic(main thread): Segmentation fault at address 0x0
oh no: Bun has crashed. This indicates a bug in Bun, not your code.
 
To send a redacted crash report to Bun's team,
please file a GitHub issue using the link below:
 
 http://localhost:38809/1.2.21/Ba2f706382wNgkgUu11luEm6yX+lwy+Dgtt+oEurthoD8214mE___07+09DA2AA
 
 
 6 | describe("Worker destruction", () => {
 7 |   const method = ["Bun.connect", "Bun.listen", "fetch"];
 8 |   describe.each(method)("bun when %s is used in a Worker that is terminating", method => {
 9 |     // fetch: ASAN failure
10 |     test.skipIf(isBroken && method == "fetch")("exits cleanly", () => {
11 |       expect([join(import.meta.dir, "worker_thread_check.ts"), method]).toRun();
                                                                             ^
error:
 
Command /var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/test/js/node/worker_threads/worker_thread_check.ts Bun.connect failed:
Spawned 10 workers RSS 79 MB
Spawned 10 workers RSS 87 MB
Spawned 10 workers RSS 90 MB
 
      at <anonymous> (/var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/test/js/node/worker_threads/worker_destruction.test.ts:11:73)
✗ Worker destruction > bun when Bun.connect is used in a Worker that is terminating > exits cleanly [597.56ms]
✓ Worker destruction > bun when Bun.listen is used in a Worker that is terminating > exits cleanly [503.47ms]
» Worker destruction > bun when fetch is used in a Worker that is terminating > exits cleanly
 
 
 1 pass
 1 skip
 1 fail
 2 expect() calls
Ran 3 tests across 1 file. [1125.00ms]
======== Stack trace from GDB for bun-profile-28234.core: ========
Program terminated with signal SIGILL, Illegal instruction.
#0  crash_handler.crash () at crash_handler.zig:1523
[Current thread is 1 (LWP 28234)]
#0  crash_handler.crash () at crash_handler.zig:1523
#1  0x0000000002db77aa in crash_handler.crashHandler (reason=..., error_return_trace=0x0, begin_addr=...) at crash_handler.zig:471
#2  0x0000000002db2b55 in crash_handler.handleSegfaultPosix (sig=<optimized out>, info=<optimized out>) at crash_handler.zig:792
#3  0x0000000004716b58 in WTF::jscSignalHandler (sig=11, info=0x7ffe54051e90, ucontext=0x0) at vendor/WebKit/Source/WTF/wtf/threads/Signals.cpp:548
#4  <signal handler called>
#5  JSC::VM::currentThreadIsHoldingAPILock (this=0x148296c30000) at vendor/WebKit/Source/JavaScriptCore/runtime/VM.h:840
#6  JSC::sanitizeStackForVM (vm=...) at vendor/WebKit/Source/JavaScriptCore/runtime/VM.cpp:1369
#7  0x0000000003f4a060 in JSC::LocalAllocator::allocate(JSC::Heap&, unsigned long, JSC::GCDeferralContext*, JSC::AllocationFailureMode)::{lambda()#1}::operator()() const (this=<optimized out>) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/LocalAllocatorInlines.h:46
#8  JSC::FreeList::allocateWithCellSize<JSC::LocalAllocator::allocate(JSC::Heap&, unsigned long, JSC::GCDeferralContext*, JSC::AllocationFailureMode)::{lambda()#1}>(JSC::LocalAllocator::allocate(JSC::Heap&, unsigned long, JSC::GCDeferralContext*, JSC::AllocationFailureMode)::{lambda()#1} const&, unsigned long) (this=0x148296c38e48, cellSize=16, slowPath=...) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/FreeListInlines.h:46
#9  JSC::LocalAllocator::allocate (this=0x148296c38e30, heap=..., cellSize=16, deferralContext=0x0, failureMode=JSC::AllocationFailureMode::Assert) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/LocalAllocatorInlines.h:44
#10 JSC::GCClient::IsoSubspace::allocate (this=0x148296c38e30, vm=..., cellSize=16, deferralContext=0x0, failureMode=JSC::AllocationFailureMode::Assert) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/IsoSubspaceInlines.h:34
#11 JSC::tryAllocateCellHelper<JSC::JSString, (JSC::AllocationFailureMode)0> (vm=..., size=16, deferralContext=0x0) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/JSCellInlines.h:192
#12 JSC::allocateCell<JSC::JSString> (vm=..., size=16) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/JSCellInlines.h:212
#13 JSC::JSString::create (vm=..., value=...) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/JSString.h:204
#14 0x0000000004479ad1 in JSC::jsNontrivialString (vm=..., s=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSString.h:846
#15 JSC::VM::ensureTerminationException (this=0x148296c30000) at vendor/WebKit/Source/JavaScriptCore/runtime/VM.cpp:627
#16 JSGlobalObject__requestTermination (globalObject=<optimized out>) at ./build/release/./src/bun.js/bindings/ZigGlobalObject.cpp:3979
#17 0x0000000003405ab8 in bun.js.web_worker.notifyNeedTermination (this=0x542904f0d80) at /var/lib/buildkite-agent/builds/ip-172-31-16-28/bun/bun/src/bun.js/web_worker.zig:558
#18 0x0000000004362b6f in WebCore::Worker::terminate (this=0x984c900000000000) at ./src/bun.js/bindings/webcore/Worker.cpp:266
#19 WebCore::jsWorkerPrototypeFunction_terminateBody(JSC::JSGlobalObject*, JSC::CallFrame*, WebCore::JSWorker*)::{lambda()#1}::operator()() const (this=<optimized out>) at ./build/release/./src/bun.js/bindings/webcore/JSWorker.cpp:549
#20 WebCore::toJS<WebCore::IDLUndefined, WebCore::jsWorkerPrototypeFunction_terminateBody(JSC::JSGlobalObject*, JSC::CallFrame*, WebCore::JSWorker*)::{lambda()#1}>(JSC::JSGlobalObject&, JSC::ThrowScope&, WebCore::jsWorkerPrototypeFunction_terminateBody(JSC::JSGlobalObject*, JSC::CallFrame*, WebCore::JSWorker*)::{lambda()#1}&&) (lexicalGlobalObject=..., throwScope=..., valueOrFunctor=...) at ./src/bun.js/bindings/webcore/JSDOMConvertBase.h:174
#21 WebCore::jsWorkerPrototypeFunction_terminateBody (lexicalGlobalObject=<optimized out>, callFrame=<optimized out>, castedThis=<optimized out>) at ./build/release/./src/bun.js/bindings/webcore/JSWorker.cpp:549
#22 WebCore::IDLOperation<WebCore::JSWorker>::call<&WebCore::jsWorkerPrototypeFunction_terminateBody, (WebCore::CastedThisErrorBehavior)0> (lexicalGlobalObject=..., operationName=..., callFrame=...) at ./src/bun.js/bindings/webcore/JSDOMOperation.h:63
#23 WebCore::jsWorkerPrototypeFunction_terminate (lexicalGlobalObject=<optimized out>, callFrame=0x7ffe540536b8) at ./build/release/./src/bun.js/bindings/webcore/JSWorker.cpp:554
#24 0x000014825580c038 in ?? ()
#25 0x00007ffe540537b0 in ?? ()
#26 0x0000148255a626cb in ?? ()
#27 0x0000000000000000 in ?? ()
1 crashes reported during this test
```
pull Bot pushed a commit that referenced this pull request Apr 28, 2026
…n-sh#29800)

## What does this PR do?

Fixes a use-after-free in `selectALPNCallback`
(`src/bun.js/api/bun/socket.zig:17`) when a `node:tls` / `Bun.listen`
TLS server with `ALPNProtocols` handles overlapping handshakes.

`SSL_CTX_set_alpn_select_cb` registers on the **listener-level**
`SSL_CTX`, so its `arg` is shared across every accepted connection. The
old code passed the per-connection `*TLSSocket` as `arg`, so each
`onOpen` overwrote the previous one. When connection A's ClientHello is
processed after connection B's `onOpen` has run — and B has since been
freed — the callback dereferences a dangling pointer and feeds garbage
`protos` into `SSL_select_next_proto`:

```
#9  SSL_select_next_proto (..., peer=0xb10008886384bf6d, peer_len=1431130639,
                           supported="\002h2\bhttp/1.1", supported_len=12)
#10 selectALPNCallback (in="\002h2\bhttp/1.1", inlen=12, arg=<freed>) at socket.zig:23
#11 bssl::ssl_negotiate_alpn
#12 bssl::do_select_parameters
#21 ssl_on_data at openssl.c:505
```

(from `fetch-http2-client.test.ts` under `describe.concurrent` on
`:alpine: 3.23 aarch64`)

**Fix:** store the `*TLSSocket` on the per-connection `SSL` via
`SSL_set_ex_data(ssl, 0, this)` and read it back from the `SSL*`
parameter in the callback, ignoring the CTX-level `arg`. Slot 0 is
otherwise unused in the codebase.

## How did you verify your code works?

- `bun bd test test/js/node/tls/node-tls-server.test.ts
test/js/node/tls/node-tls-connect.test.ts
test/js/node/http2/node-http2.test.js
test/js/web/fetch/fetch-http2-client.test.ts` → 349 pass, 0 fail
- `bun run zig:check-all` → all platforms compile
- The existing `connectionListener should emit the right amount of
times, and with alpnProtocol available` test (50 parallel ALPN
connections) covers the path; the original crash was caught by
`fetch-http2-client.test.ts` on aarch64-musl CI
pull Bot pushed a commit that referenced this pull request May 2, 2026
…es (oven-sh#30077)

## What

When a chunked (or HTTP/3) request body exceeds `maxRequestBodySize`,
`onBufferedBodyChunk` writes the 413 directly on the raw uWS response:

```zig
resp.writeStatus("413 Payload Too Large");
resp.endWithoutBody(comptime !http3);
```

`internalEnd` → `markDone()` nulls `onAborted`, so when the socket
closes no abort ever fires to detach `ctx.resp` or release the base ref.
`this.resp` is left pointing at a completed response whose socket is
about to be freed by `us_internal_free_closed_sockets`.

If the fetch handler returned a pending Promise:

- **resolve**: `handleResolve` → `isAbortedOrEnded()` is false
(`this.resp != null`) → `render()` → `runCorkedWithType` corks the freed
socket → **heap-use-after-free** (ASAN trace below).
- **reject**: `handleReject` reads `resp.hasResponded()` off freed
memory, sees `true`, skips the error handler, and returns without ever
releasing the base ref → **RequestContext leaks**
(`server.pendingRequests` never returns to 0).

## Fix

Route through `this.endWithoutBody()` (the `RequestContext` wrapper)
instead of the raw `resp.endWithoutBody()`. That path does
`detachResponse()` (nulls `this.resp`, clears
`onData`/`onAborted`/`onTimeout`) and `deref()` (releases the base ref),
matching every other end path in this file.

The body promise is rejected with the specific `"Request body exceeded
maxRequestBodySize"` error *before* `endWithoutBody()` so
`endRequestStreaming()` doesn't overwrite it with a generic
`ConnectionClosed`. `has_written_status` is set so any later
`renderMissing`/`renderMetadata` knows the status line is already
committed.

## Repro

```
==ERROR: AddressSanitizer: heap-use-after-free
  #0 us_socket_group socket.c:77
  #1 uWS::AsyncSocket<false>::getLoopData() AsyncSocket.h:69
  #2 uWS::AsyncSocket<false>::isCorked() AsyncSocket.h:141
  #3 uWS::HttpResponse<false>::cork(...) HttpResponse.h:647
  #4 uws_res_cork libuwsockets.cpp:1740
  #5 ...runCorkedWithType Response.zig:299
  #6 ...doRenderBlob RequestContext.zig:1942
  ...
  #11 ...handleResolve RequestContext.zig:220
  #12 ...onResolve RequestContext.zig:154
freed by:
  #1 us_poll_free epoll_kqueue.c:73
  #2 us_internal_free_closed_sockets loop.c:305
```

## Test

`test/js/bun/http/serve-pending-promise-abort-leak.test.ts` — new case
sends a raw `Transfer-Encoding: chunked` POST exceeding
`maxRequestBodySize` with a handler that holds its resolve/reject, waits
for the socket to be reclaimed, then settles the Promise. Asserts
`pendingRequests` returns to 0 for both paths, the body was rejected
with the right message, and a follow-up request still works.

Without the fix: ASAN heap-use-after-free on the resolve path; on
release builds the reject path shows `pendingAfterReject: 1` (leak).

Co-authored-by: robobun <robobun@users.noreply.github.com>
pull Bot pushed a commit that referenced this pull request May 4, 2026
…worker panic, never retry (oven-sh#30216)

## What

`bun test --isolate` / `--parallel` crashes when a test file loads a
native addon whose deferred napi finalizers outlive the file. The
`--parallel` coordinator then silently retries the file once, which
masks the panic and lets the run exit 0.

Fixes oven-sh#30205, oven-sh#30191. Supersedes oven-sh#30214 (same NapiEnv fix, but without
the coordinator change, the `cleanup_hooks` retarget, or a test that
actually reproduces on unpatched `main`).

## Reproduction

```sh
git clone https://github.com/workglow-dev/libs && cd libs
bun i && bun run build:packages
bun test --timeout=30000 --parallel=4 packages/test/src/test/{util,task}/*.test.ts
```

On `main` (d484fd6), 3–4 workers crash per run with either
```
ASSERTION FAILED: isMarked(cell)
  JavaScriptCore/heap/Heap.cpp:1232 : void JSC::Heap::addToRememberedSet(const JSCell *)
```
or (when the slot is already being reallocated)
```
ASSERTION FAILED: m_cellState == CellState::DefinitelyWhite
  JavaScriptCore/JSCellInlines.h:69 : JSC::JSCell::JSCell(VM &, Structure *)
```
and in release builds the segfaults at `0x68` / `0xD0` reported in
oven-sh#30205.

## Root cause

Frame-pointer walk from the assertion:

```
#3  Bun::NapiHandleScope::open(Zig::GlobalObject*, bool)
#4  NapiHandleScope__open
#6  napi.Finalizer.run
#7  napi.NapiFinalizerTask.runOnJSThread
#10 event_loop.tick
#11 event_loop.waitForPromise
#13 VirtualMachine.loadEntryPointForTestRunner   ← next test file
```

`NapiEnv::m_globalObject` is a raw `Zig::GlobalObject*`. For
non-experimental addons (`nm_version != NAPI_VERSION_EXPERIMENTAL`,
which is ~every real-world addon — sharp, better-sqlite3, etc.),
`napi_wrap`/`napi_create_external` finalizers are **deferred** to the
event loop as `NapiFinalizerTask` rather than run inside GC sweep.

Objects rooted on the old global (module graph, `globalThis.*`) only
become collectable when `Zig__GlobalObject__createForTestIsolation` runs
`gcUnprotect(oldGlobal)`. The `DeferGC` from oven-sh#29573 ends at that
function's `}`, so the next GC runs there, collects those objects, and
enqueues their finalizers. Those tasks then run on the very next
`eventLoop().tick()` — inside `loadEntryPointForTestRunner`'s
`waitForPromise` for file N+1. `Finalizer.run` opens a `NapiHandleScope`
via `env->globalObject()`, which reads `NapiHandleScopeImplStructure()`
off the dead cell and writes `m_currentNapiHandleScopeImpl` on it →
write barrier on an unmarked cell.

The `--parallel` coordinator's `reapWorker` then re-queued the file once
(`retries[idx] < 1`) into a fresh worker with no stale `NapiEnv`, which
passed — so the run reported 0 fail despite multiple Bun panics in the
log.

## Fix

**NapiEnv retarget** (`ZigGlobalObject.cpp`, `napi.h`):
`Zig__GlobalObject__createForTestIsolation` now calls
`newGlobal->adoptNapiEnvsForTestIsolation(oldGlobal)` before
`gcUnprotect`. Each `NapiEnv::m_globalObject` is repointed at the new
global and the `Ref<NapiEnv>`s are moved over, so late finalizers open
handle scopes on a live global and the envs stay owned after the old
global is swept. `VirtualMachine.swapGlobalForTestIsolation` also
repoints `rare_data.cleanup_hooks[*].globalThis` so `CleanupHook.eql()`
stays accurate.

**No retry, abort on panic** (`Coordinator.zig`): removed the per-file
retry. A worker that dies mid-file is counted as one failure. If it died
by a fatal signal (SIGILL/SIGTRAP/SIGABRT/SIGBUS/SIGFPE/SIGSEGV/SIGSYS —
Bun's own `@trap()`, a JSC/WTF assertion, or native-addon crash), the
whole run aborts with `error: a test worker process crashed with <SIG>
while running <file>`. `process.exit()` / SIGKILL are still just a
per-file failure and the run continues.

## Verification

- `test/regression/issue/30205.test.ts` — 4 tests. Adds a tiny
non-experimental addon (`isolate_finalizer_addon.c`) and a fixture
pattern (`Bun.gc(true)` + module-scope `await 0` + objects rooted on
`globalThis`) that crashes **8/8** on unpatched `main` and passes 8/8
with this change.
- `workglow-dev/libs` full 201-file unit suite: 3× clean `--parallel=4`
runs (was 3–4 crashes/run).
- Gate: `git stash -- src/ && bun bd test
test/regression/issue/30205.test.ts` → 3/4 fail; with fix → 4/4 pass.
- `test/cli/test/isolation.test.ts`,
`test/regression/issue/29519.test.ts` → pass (one pre-existing unrelated
timeout in isolation.test.ts, same as oven-sh#29573).
- `test/cli/test/parallel.test.ts` → all tests I touched pass; the 3
timing-sensitive scale-up/work-steal tests that fail in this container
fail identically on unmodified `main`.

---------

Co-authored-by: robobun <robobun@users.noreply.github.com>
Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot pushed a commit that referenced this pull request Jun 23, 2026
…e re-enters the event loop (oven-sh#32597)

Sentry BUN-2WJA / BUN-2WKB (~290 events combined, Windows x86_64,
`http_server=True`, bun 1.2.23 through 1.3.14):

```
Segmentation fault at address 0xFFFFFFFFFFFFFFFF
  endWithSink      src/runtime/webcore/Sink.zig:577
  endFromJS        src/runtime/webcore/streams.zig:1200
  finalize         src/runtime/webcore/streams.zig:1301
  clearAndFree     src/collections/baby_list.zig:148
  memset           (fault at 0xFFFFFFFFFFFFFFFF)
```

## Cause

The generated `JSReadable*Controller` `end()` and `close()` host
functions (`src/codegen/generate-jssink.ts`) stash `m_sinkPtr` in a
local, call `controller->detach()`, and only afterward dereference the
stashed pointer via `endWithSink()` / `${name}__close()`:

```cpp
void *ptr = controller->wrapped();
controller->detach();              // runs onClose JS synchronously
return ${name}__endWithSink(ptr, lexicalGlobalObject);  // derefs ptr
```

`detach()` invokes the stored `onClose` callback. For a `type: "direct"`
stream this is `readDirectStream`'s `close(stream, reason)`, which calls
`underlyingSource.cancel()`. That is arbitrary user code running while
`ptr` is still live on the C++ stack.

If the stream's `pull()` promise has already settled,
`RequestContext::on_resolve_stream` is sitting in the microtask queue.
Any path from `cancel()` that drains microtasks (e.g. the server-side
drain points in `on_response` / `do_render_with_body`, or an explicit
`drainMicrotasks()`) runs `handle_resolve_stream`, which calls
`destroy_sink` and frees the `HTTPServerWritable`. `endWithSink(ptr)`
then enters `end_from_js` on the freed allocation; `finalize()` reads
garbage for `pooled_buffer` / `buffer.cap` / `buffer.ptr` and faults in
the `memset` the allocator's free-scrub path performs.

The same ordering appears in the Rust port (`streams.rs` / `Sink.rs`)
unchanged.

## Fix

In `${controller}__end` and `${controller}__close`, finish the native
sink operation before any JS runs:

1. Call `${name}__controllerDetached(ptr, controller)` and null
`m_sinkPtr` up front (so `end_from_js`'s own `signal.close()` stays a
no-op, matching the previous behaviour, and so the later `detach()`
won't touch the native side again).
2. Run `endWithSink(ptr)` / `close(ptr)`.
3. Call `controller->detach()` last. With `m_sinkPtr` already null it
only clears `m_onPull` and fires `onClose`; by now we hold no reference
into the sink, so re-entrant teardown is safe.

## Verification

New ASAN-gated test in
`test/js/bun/http/serve-direct-readable-stream.test.ts` reproduces the
exact UAF deterministically by draining microtasks from the stream's
`cancel()` callback (the test uses
`require("bun:jsc").drainMicrotasks()` to force the drain that the
production crash hits via the server's own drain points).

<details>
<summary>ASAN output on the unfixed build</summary>

```
==22203==ERROR: AddressSanitizer: heap-use-after-free on address 0x6ee5f87602ca
READ of size 1 at 0x6ee5f87602ca thread T0
    #0 HTTPServerWritable::end_from_js src/runtime/webcore/streams.rs:1831
    #2 JSSink::js_end_with_sink src/runtime/webcore/Sink.rs:1107
    #4 WebCore::JSReadableHTTPResponseSinkController__end JSSink.cpp:620

freed by thread T0 here:
    #10 HTTPServerWritable::destroy src/runtime/webcore/streams.rs:1950
    #11 RequestContext::destroy_sink src/runtime/server/RequestContext.rs:1930
    #12 RequestContext::handle_resolve_stream src/runtime/server/RequestContext.rs:2680
    #13 RequestContext::on_resolve_stream src/runtime/server/RequestContext.rs:2716
    ...
    #24 JSC::VM::drainMicrotasks()
```
</details>

With the fix the fixture completes normally. Existing suites
(`serve.test.ts`, `bun-server.test.ts`,
`direct-readable-stream.test.tsx`, `streams.test.js`, the sink leak
tests) show no new failures against the unfixed build.

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot pushed a commit that referenced this pull request Jun 26, 2026
…oven-sh#32742)

### What does this PR do?

Fixes a use-after-free in the HTTP client's CONNECT proxy tunnel, caught
by ASAN:

```
READ of size 8 at 0x61e00001fe80 thread T6
  #0 Option<RefPtr<ProxyTunnel>>::as_ref
  #1 proxy_tunnel::on_close            ProxyTunnel.rs:525
  #2 SSLWrapper::trigger_close_callback uws/lib.rs:833
  #3 SSLWrapper::handle_reading         uws/lib.rs:1053
  ...
freed by thread T6 here (same stack, same `handle_reading` call):
  #5 AsyncHTTP::on_async_http_callback_raw            AsyncHTTP.rs:819
  #7 HTTPClient::send_progress_update_without_stage_check
  #9 proxy_tunnel::on_data              ProxyTunnel.rs:350
  #11 SSLWrapper::trigger_data_callback uws/lib.rs:824
  #12 SSLWrapper::handle_reading        uws/lib.rs:1046
```

`SSLWrapper::handle_reading` flushes pending decrypted bytes to the data
callback, then runs the close callback, guarded only by
`closed_notified`:

1. The flushed data callback completes a keep-alive response through the
tunnel. A fatal TLS record error sets only `fatal_error` — none of the
shutdown flags — so the wrapper passed `tunnel_poolable`'s
`!is_shutdown()` check and the tunnel was handed to the keep-alive pool.
Nothing called `wrapper.shutdown()`, so `closed_notified` was never
latched. Dispatching the final result then freed the
`ThreadlocalAsyncHTTP` that embeds the `HTTPClient`.
2. The guard (`ssl.is_none() || closed_notified()`) passes.
3. `trigger_close_callback()` invokes `on_close(handlers.ctx)` with
`ctx` pointing at the freed client.

The pooling branch is the only terminal path that doesn't go through
`close_proxy_tunnel(true)` → `wrapper.shutdown()` → `closed_notified`,
which is the latch the read loop relies on. `SSLWrapper::shutdown`
already special-cases the *close_notify* flavor of this for exactly that
reason; the fatal-error flavor never reaches `shutdown()`.

The fix is one predicate: a tunnel whose wrapper has a fatal error or
pending unconsumed input/output is not poolable. That routes it through
the orderly teardown that latches `closed_notified`, and the pending-I/O
half closes the same hole for a tunnel pooled from a mid-loop data
callback while more decrypted bytes or queued output remain. Both are
also required for the pool to be correct on its own terms — a poisoned
or dirty TLS session must not be handed to the next request.

### How did you verify your code works?

New regression test in `test/js/bun/http/proxy.test.ts` (next to the
existing close_notify sibling): an HTTPS keep-alive response through a
CONNECT proxy with a corrupt TLS record appended to the same TCP burst,
followed by a second request that can only complete if the HTTP client
thread survived the first.

Against an unfixed ASAN debug build the fixture aborts every run:

```
==20981==ERROR: AddressSanitizer: heap-use-after-free on address 0x61e00001fe80
READ of size 8 at 0x61e00001fe80 thread T6
...
exit=134
```

With this change it prints `4096 200 200` and exits 0 with no ASAN
report. `test/js/bun/http/proxy.test.ts` (49/49),
`fetch-proxy-connect-tunnel-split-envelope.test.ts`,
`fetch-proxy-tls-intern-race.test.ts`, and `fetch-keepalive.test.ts` all
pass.
pull Bot pushed a commit that referenced this pull request Jul 16, 2026
…ry rewrite (oven-sh#34271)

`test/js/bun/util/filesystem_router.test.ts` went red on alpine x64 in
build [73276](https://buildkite.com/bun/bun/builds/73276): the `reload()
while Bun.build() resolves the same directory` subprocess segfaulted in
`bust_dir_cache_recursive`, inlined from `NonNull::new`.

## Cause

`RealFS::entries_at` (`src/resolver/lib.rs`) replaces a cached
`DirEntry` in place when the caller's resolver generation is newer than
the cached listing's. The replacement at `*e_ptr = new_entry` drops the
old `DirEntry`, which drops its `data: StringHashMap<*mut Entry>` and
frees the hashmap's bucket allocation. The function's comment says
`entries_mutex held by caller`, but that is only true on one of the five
paths that reach it: `dir_info_uncached`, when entered from
`dir_info_cached_miss`. The other callers (`finalize_result`,
`handle_esm_resolution`, `load_index_with_extension`,
`Transpiler::run_env_loader`) all reach `entries_at` after
`dir_info_cached_maybe_log` has already returned and released both
`RESOLVER_MUTEX` and `entries_mutex`.

`FileSystemRouter::reload()` and `RouteLoader::load` iterate the same
`DirEntry.data` map under `entries_mutex` (the snapshot pattern oven-sh#33056
introduced for exactly this kind of concurrent rewrite). With
`entries_at`'s rewrite unsynchronized, a `Bun.build()` on the bundler
thread can drop the map while `reload()` on the JS thread is
mid-iteration.

The generation mismatch is what makes `entries_at` enter its rewrite
branch, so the window only opens once the bundle thread has processed at
least one batch (it bumps its own generation after every queue drain);
every subsequent `Bun.build()` then re-reads any directory that
`reload()` just refreshed to generation 0.

ASAN catches it as a heap-use-after-free with the two sides of the race
laid out exactly:

```
READ of size 16 (thread T0):
  #6 HashMap::values
  #7 StringHashMap<*mut Entry>::values                       src/collections/array_hash_map.rs:1864
  #8 FileSystemRouter::bust_dir_cache_recursive              src/runtime/api/filesystem_router.rs:395
  #9 FileSystemRouter::bust_dir_cache                        src/runtime/api/filesystem_router.rs:451
  #10 FileSystemRouter::reload                               src/runtime/api/filesystem_router.rs:476

freed by thread T11 (Bundler):
  #11 drop_in_place<bun_resolver::fs_full::DirEntry>
  #12 bun_resolver::fs::RealFS::entries_at                   src/resolver/lib.rs:1639
  #13 DirInfo::get_entries_ref                               src/resolver/dir_info.rs:266
  #14 Resolver::finalize_result                              src/resolver/resolver.rs:1714
  #15 Resolver::resolve_and_auto_install                     src/resolver/resolver.rs:1485
  ...
  #23 BundleThread::generate_in_new_thread                   src/bundler/BundleThread.rs:276

previously allocated by thread T0:
  #17 HashMap::reserve
  #18 Resolver::dir_info_cached_miss                         src/resolver/resolver.rs:4591
  #19 Resolver::dir_info_cached_maybe_log                    src/resolver/resolver.rs:4201
  #20 Resolver::read_dir_info                                src/resolver/resolver.rs:4118
  #21 FileSystemRouter::reload                               src/runtime/api/filesystem_router.rs:492
```

(The use side is sometimes `RouteLoader::load` at
`src/router/lib.rs:816` instead; same map, same lock.)

This has been the shape of `entries_at` since the Rust port; oven-sh#33056
narrowed the race by snapshotting under the lock but assumed the rewrite
side already held it.

## Fix

`entries_at` now takes `entries_mutex` itself, matching
`read_directory_with_iterator` which already does. The one call path
that reaches it with the lock already held (`dir_info_cached_miss` ->
`dir_info_uncached` -> `parent_.get_entries_ref`) routes through a new
`entries_at_locked` / `get_entries_ref_locked` pair so the non-recursive
mutex is not re-entered. That path is the only one that passes a
non-`None` parent to `dir_info_uncached`; the other caller
(`dir_info_for_resolution`) passes `None`, so the parent branch
containing the accessor never runs there.

## Test

The existing concurrency test now awaits one `Bun.build()` first, so the
bundle thread's generation is already past zero when the concurrent
rounds start, and then runs forty reload/build rounds instead of one.
That is the shape that reaches the stale-generation rewrite at all; the
original single-round fixture usually completes with every build still
on generation 0.

The race is scheduling-dependent. Pinning the fixture to a single core
reproduces the ASAN use-after-free on roughly 3 in 10 runs against an
unpatched debug build and 0 in 15 with this change; with all 16 cores
available the unpatched build reproduces at roughly 1 in 30. The
assertions are otherwise the same as before, so the test continues to
cover the behavior oven-sh#33056 added.

Also ran the full `filesystem_router.test.ts`,
`test/bundler/bun-build-api.test.ts` (including the thousands-of-builds
test that exercises the generation path heavily),
`test/js/bun/resolve/resolve.test.ts`, `test/cli/hot/hot.test.ts`,
`test/cli/watch/watch.test.ts`, `test/bake/framework-router.test.ts`,
and `bun run rust:check-all` (10/10 targets).

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 0 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/bun/util/filesystem_router.test.ts

<!-- robobun:evidence:end -->
pull Bot pushed a commit that referenced this pull request Jul 23, 2026
…ven-sh#35209)

Companion to oven-sh/libuv#11 (merged as `2881ce5` on the `bun` branch).

## What changes for Bun

oven-sh/libuv#11 fixes seven Windows error-translation/propagation bugs
found by audit, scoped to those reachable from Bun's own code. Full
breakdown is in [this comment on the libuv
PR](oven-sh/libuv#11 (comment)).

The one user-visible change: on Windows, an async write to a pipe whose
read end has closed now delivers `{code: "EPIPE"}` to the write callback
instead of `{code: "EOF"}`. Before the fix, `uv__process_pipe_write_req`
used the read-side translator, which maps `ERROR_BROKEN_PIPE` /
`ERROR_NO_DATA` to `UV_EOF` / `UV_EAGAIN`; the write-side translator
maps both to `UV_EPIPE`, matching Node and the synchronous `uv_write()`
path.

This also fixes `shell::IOWriter::on_error`'s broken-pipe detection
(`src/runtime/shell/IOWriter.rs:867`), which keys on `E::PIPE` and so
was not firing when the failure arrived via the async write completion.

## Changes

- `scripts/build/deps/libuv.ts`: bump `LIBUV_COMMIT` to `2881ce5` (the
#11 merge) and add #11 to the provenance comment.
- `test/js/bun/spawn/spawn.test.ts`: new test `stdin.end() rejects with
EPIPE when the child exits before consuming the write`. The child reads
one byte and exits, leaving 16MB queued on the parent's stdin (large
enough to exceed macOS socket autotuning and the 64KB Windows named-pipe
buffer), so `end()` is still draining when the read end closes. Verified
fail-before on Windows (`caught.code === "EOF"` with the current
release) and pass on Linux/macOS.

The existing `child_process` `stdin write failure (EPIPE)` test is left
skipped on Windows: its child does `fs.closeSync(0)`, but libuv's
`fs__close` on Windows is a no-op for fd ≤ 2 (`src/win/fs.c:690`), so
the pipe never breaks and the test hangs. That skip is unrelated to #11.

## Not changed (reachable from Bun, strict improvement, no code change
needed)

- `uv_pipe()` (used by `spawn/process.rs:2000`, `sys/lib.rs:4393`,
`install/security_scanner.rs:1129`): previously returned a raw positive
Win32 error on failure; Bun's `ReturnCode` predicate is `< 0`, so that
leaked through as success with garbage fds. Now returns a negative
`UV_E*`. Bun-side check was already correct.
- `uv_fs_realpath` / `uv_fs_readlink` (used by `fs.realpath*`,
`fs.readlink*`, `fs.cp`, `fs.watch`): OOM during UTF-16 conversion now
reports `ENOMEM` instead of success-with-null or a stale error.
- `fs__filemap_ex_filter` (`UV_FS_O_FILEMAP`, exposed via
`fs.constants`): in-page errors now surface the real disk error instead
of `UNKNOWN`.
- `uv__tty_move_caret` (legacy console ANSI emulation): missing `return
-1` added; bypassed on Windows 10+ where Bun enables VT mode at startup.

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 3 · Platform-specific test-only change;
deferring to CI.

<!-- robobun:evidence:end -->

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot pushed a commit that referenced this pull request Jul 29, 2026
…rier (oven-sh#36337)

`JSNativeStreamSourceAdapter::m_controller` was a
`JSC::Weak<JSReadableStreamDefaultController>`. When the native pull
promise is rejected (socket fault on a fetch body) the adapter is queued
as the `onNativePullRejected` reaction context, which roots the
**adapter** but not the **controller**: the adapter's only edge to it
was the `Weak`. `FetchTasklet` releases both native `Strong<>`s to the
body stream before that microtask drains, so a GC in between can leave
the entire consumer graph (`controller -> stream -> reader -> pipe op ->
destination -> writer -> readyPromise`) white. The subsequent error
cascade then enqueues the pipe's writes-drained shutdown deferral
against a corpse `op`, and `performPipeShutdownAction(AbortDestination)`
dereferences a swept `readyPromise`:

```
ASSERTION FAILED: result   JSObject.h(583) JSGlobalObject *JSC::JSObject::realm() const
#5  JSC::JSObject::realm()
#6  JSC::JSPromise::rejectPromise
#7  JSC::JSPromise::reject
#8  Bun::WebStreams::writableStreamDefaultWriterEnsureReadyPromiseRejected
#9  Bun::WebStreams::writableStreamStartErroring
#10 Bun::WebStreams::writableStreamAbort
#11 WebCore::performPipeShutdownAction (AbortDestination)
#12 WebCore::JSStreamPipeToOperation::onWritesFinishedForShutdown
```

On builds without the assert the same path is a silent write into
freed/reused promise memory.

## Fix

Hold `m_controller` as a visited internal field so a queued adapter
roots the controller directly. The edge is cleared on every terminal
path (`nativeSourcePullRejected`, `nativeSourceCallClose`,
`nativeSourceCancel`); `controller->algorithmContext` is cleared by
`readableStreamDefaultControllerClearAlgorithms`, so the abandoned case
is an ordinary intra-heap cycle mark-sweep collects.
`NewSource::this_jsvalue` is only `Strong` during FileReader I/O, where
pinning the consumer graph is the correct behavior anyway.

With the `Weak` gone the adapter no longer needs a destructor, so it is
now a `JSInternalFieldObjectImpl<5>`: the five JSValue members (handle,
pendingView, closer, drainValue, controller) are internal fields visited
by the base class, with typed accessors at call sites. The scalar
members (chunkSize, flag bitfield, text-decode state) stay as plain
members.

## Verification

`native-source-onclose-leak.test.ts` (the partial-read + `releaseLock`
abandonment tests for Blob/fetch/File sources) continues to pass,
confirming the cycle does not pin. `streams.test.js`,
`pipeTo-signal-leak.test.ts`, `compression.test.ts`, `blob.test.ts` all
pass.

The crash itself is 0/1800 standalone; it reproduces ~1/3 only under a
fault-injected tracer replay. `pipeTo-shutdown-gc.test.ts` exercises the
shape (native body source, socket fault mid-stream, fire-and-forget
`pipeTo` under `collectContinuously`, `AbortDestination` shutdown arm)
as a regression surface.

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 2 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/web/streams/pipeTo-shutdown-gc.test.ts

<!-- robobun:evidence:end -->
pull Bot pushed a commit that referenced this pull request Aug 2, 2026
…during VM shutdown (oven-sh#36750)

Fixes `test/js/bun/http/bun-serve-html-405.test.ts` going red on
x64-asan (build [87498](https://buildkite.com/bun/bun/builds/87498) and
several unrelated PR builds since ~87000).

## Repro

```
BUN_DESTRUCT_VM_ON_EXIT=1 ASAN_OPTIONS=detect_leaks=1 \
LSAN_OPTIONS=suppressions=test/leaksan.supp \
bun-debug test test/js/bun/http/bun-serve-html-405.test.ts
```

```
Indirect leak of 2104 byte(s) in 1 object(s) allocated from:
    ...
    #11 new<bun_runtime::server::NewServer<false, true>>
    #12 init<false, true> src/runtime/server/mod.rs:2009:47
    #13 bun_runtime::api::bun_object::serve src/runtime/api/BunObject.rs:1564:26
    ...
    #21 BunObject_callback_serve src/runtime/api/BunObject.rs:230:25
SUMMARY: AddressSanitizer: 2942 byte(s) leaked in 7 allocation(s).
```

10/10 without this change, 0/10 with it (local debug+ASAN).

## Cause

`using server = Bun.serve({ development: true, routes: { "/": html } })`
disposes via `stop(true)`, which makes `deinit_if_we_can()` downgrade
`js_value` to `Weak` and return. The `NewServer` Box is only freed once
the JS wrapper's `finalize()` fires and `schedule_deinit()` enqueues the
actual `deinit()` as a `ManagedTask`.

When the wrapper survives to `lastChanceToFinalize`
(`BUN_DESTRUCT_VM_ON_EXIT=1`, which the CI runner sets on ASAN lanes),
`global_exit()` has already had its last event-loop tick. The enqueued
task never runs, and `EventLoop::deinit()` drops the task box without a
cleanup (`ManagedTask::new` sets `cleanup: None`). The 2104-byte
`NewServer<false, true>` Box, its `config.static_routes` Vec, the
`html_bundle::Route` it refcounts, and the route's path strings are all
orphaned. `Route.server: Cell<Option<AnyServer>>` points back at the
server so LSan sees a pointer cycle and reports every allocation as
indirect.

The path has always existed, but before oven-sh#35356 the per-tick GC sampler
usually collected the wrapper during the handful of event-loop ticks
between the test body and `global_exit()`, so `schedule_deinit()` ran
while the loop was still live. With only the 1s idle-timer GC, a single
fast test like this one reaches shutdown with the wrapper still alive
more often (about half the PR builds since the merge).

## Fix

`schedule_deinit()` now sets `DEINIT_SCHEDULED` and returns without
enqueueing when `is_shutting_down()`. `finalize()` then frees the Box
synchronously when the server has been fully drained: it unboxes via
`Box::into_raw` first so the dealloc goes through the raw owning pointer
rather than a `&mut self` frame (whose FnEntry protector would make the
dealloc Stacked-Borrows UB, same pattern as `Listener::finalize` /
`UDPSocket::finalize`). Every JSC handle on the Drop chain
(`JSPromiseStrong`, `JsRef`, `UserRouteBuilder.callback: Strong`)
funnels through `Strong::Impl::destroy`, which is a no-op past
`is_shutting_down()`, so freeing here is safe.

The inline free is gated on `TERMINATED`: `NewApp::destroy` runs
`us_socket_group_deinit`, which unlinks the socket group from the loop's
list without closing any sockets still in it. A graceful `stop()` only
closes the listener and leaves keep-alive sockets open in the group;
destroying the app there would orphan them (seen as a 280-byte
`us_poll_t` direct leak on
`vendor/elysia/test/core/before-handle-arrow.test.ts` with an earlier
revision of this PR, and a `US_ASSERT(head_sockets==NULL)` abort on the
debug build). `TERMINATED` is set only once `app.close()` has run, so
the inline free is taken for abruptly-stopped servers (what `using
server` does) and skipped for gracefully-stopped ones, which is
identical to `main`'s behaviour for them.

Other callers that can reach `schedule_deinit()` past shutdown (a last
request draining inside `close_all_socket_groups`, which runs before
`lastChanceToFinalize`) only set the flag and leave the Box, since
`NewApp::destroy` there would delete the uws socket group mid-iteration.

## Verification

Two ASAN-only subprocess tests added:
- abrupt `stop(true)` of a dev server with an HTML route under
`BUN_DESTRUCT_VM_ON_EXIT=1` + `detect_leaks=1`: fails on `main` with the
7-allocation LSan report above; passes with this change.
- graceful `stop()` of a plain server with a keep-alive client
connection: passes on both `main` and this change (asserts
`us_socket_group_deinit`'s `head_sockets==NULL` precondition, which an
earlier revision of this change violated).

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 2 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/bun/http/bun-serve-html-405.test.ts

<!-- robobun:evidence:end -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants