[pull] main from oven-sh:main - #10
Merged
Merged
Conversation
pull Bot
pushed a commit
that referenced
this pull request
Jul 25, 2025
…ck traces upon crash in CI (oven-sh#21143) ### What does this PR do? Closes oven-sh#13012 On Linux, when any Bun process spawned by `runner.node.mjs` crashes, we run GDB in batch mode to print a backtrace from the core file. And on all platforms, we run a mini `bun.report` server which collects crashes reported by any Bun process executed during the tests, and after each test `runner.node.mjs` fetches and prints any new crashes from the server. <details> <summary>example 1</summary> ``` #0 crash_handler.crash () at crash_handler.zig:1513 #1 0x0000000002cf4020 in crash_handler.crashHandler (reason=..., error_return_trace=0x0, begin_addr=...) at crash_handler.zig:479 #2 0x0000000002cefe25 in crash_handler.handleSegfaultPosix (sig=<optimized out>, info=<optimized out>) at crash_handler.zig:800 #3 0x00000000045a1124 in WTF::jscSignalHandler (sig=11, info=0x7ffe044e30b0, ucontext=0x0) at vendor/WebKit/Source/WTF/wtf/threads/Signals.cpp:548 #4 <signal handler called> #5 JSC::JSCell::type (this=0x0) at vendor/WebKit/Source/JavaScriptCore/runtime/JSCellInlines.h:137 #6 JSC::JSObject::getOwnNonIndexPropertySlot (this=0x150bc914fe18, vm=..., structure=0x150a0102de50, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSObject.h:1348 #7 JSC::JSObject::getPropertySlot<false> (this=0x150bc914fe18, globalObject=0x150b864e0088, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSObject.h:1433 #8 JSC::JSValue::getPropertySlot (this=0x7ffe044e4880, globalObject=0x150b864e0088, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSCJSValueInlines.h:1108 #9 JSC::JSValue::get (this=0x7ffe044e4880, globalObject=0x150b864e0088, propertyName=..., slot=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSCJSValueInlines.h:1065 #10 JSC::LLInt::performLLIntGetByID (bytecodeIndex=..., codeBlock=0x150b861e7740, globalObject=0x150b864e0088, baseValue=..., ident=..., metadata=...) at vendor/WebKit/Source/JavaScriptCore/llint/LLIntSlowPaths.cpp:878 #11 0x0000000004d7b055 in llint_slow_path_get_by_id (callFrame=0x7ffe044e4ab0, pc=0x150bc92ea0e7) at vendor/WebKit/Source/JavaScriptCore/llint/LLIntSlowPaths.cpp:946 #12 0x0000000003dd6042 in llint_op_get_by_id () #13 0x0000000000000000 in ?? () ``` </details> <details> <summary>example 2</summary> ``` #0 crash_handler.crash () at crash_handler.zig:1513 #1 0x0000000002c5db80 in crash_handler.crashHandler (reason=..., error_return_trace=0x0, begin_addr=...) at crash_handler.zig:479 #2 0x0000000002c59f60 in crash_handler.handleSegfaultPosix (sig=<optimized out>, info=<optimized out>) at crash_handler.zig:800 #3 0x00000000042ecc88 in WTF::jscSignalHandler (sig=11, info=0xfffff60141b0, ucontext=0xfffff6014230) at vendor/WebKit/Source/WTF/wtf/threads/Signals.cpp:548 #4 <signal handler called> #5 bun.js.api.FFIObject.Reader.u8 (globalObject=0x4000554e0088) at /var/lib/buildkite-agent/builds/ip-172-31-75-92/bun/bun/src/bun.js/api/FFIObject.zig:65 #6 bun.js.jsc.host_fn.toJSHostCall__anon_1711576 (globalThis=0x4000554e0088, args=...) at /var/lib/buildkite-agent/builds/ip-172-31-75-92/bun/bun/src/bun.js/jsc/host_fn.zig:97 #7 bun.js.jsc.host_fn.DOMCall("Reader"[0..6],bun.js.api.FFIObject.Reader,"u8"[0..2],.{ .reads = .{ ... }, .writes = .{ ... } }).slowpath (globalObject=0x4000554e0088, thisValue=70370172175040, arguments_ptr=0xfffff6015460, arguments_len=1) at /var/lib/buildkite-agent/builds/ip-172-31-75-92/bun/bun/src/bun.js/jsc/host_fn.zig:490 #8 0x000040003419003c in ?? () #9 0x0000400055173440 in ?? () ``` </details> I used GDB instead of LLDB (as the branch name suggests) because it seems to produce more useful stack traces with musl libc. - [x] on linux, use gdb to print from core dump of main bun process crashed - [x] on linux, use gdb to print from all new core dumps (so including bun subprocesses spawned by the test that crashed) - [x] on all platforms, use a mini bun.report server to print a self-reported trace (depends on oven-sh/bun.report#15; for now our package.json points to a commit on the branch of that repo) - [x] fix trying to fetch stack traces too early on windows - [x] use output groups so the traces show up alongside the log for the specific test instead of having to find it in the logs from the entire run - [x] get oven-sh/bun.report#15 merged, and point to a bun.report commit on the main branch instead of the PR branch in package.json ### How did you verify your code works? Manually, and in CI with a crashing test. --------- Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot
pushed a commit
that referenced
this pull request
Aug 8, 2025
<details> <summary> observed in https://buildkite.com/bun/bun/builds/22442#annotation-test/js/node/zlib/leak.test.ts </summary> ``` ==5045==ERROR: AddressSanitizer: heap-use-after-free on address 0x5220000243c0 at pc 0x00000dad671b bp 0x14f22d4a4990 sp 0x14f22d4a4988 READ of size 8 at 0x5220000243c0 thread T5 (HeapHelper) ======== Stack trace from GDB for HeapHelper-5045.core: ======== Program terminated with signal SIGABRT, Aborted. #0 0x000014f2c3672eec in ?? () from /lib/x86_64-linux-gnu/libc.so.6 [Current thread is 1 (Thread 0x14f22d4f46c0 (LWP 5050))] #0 0x000014f2c3672eec in ?? () from /lib/x86_64-linux-gnu/libc.so.6 #1 0x000014f2c3623fb2 in raise () from /lib/x86_64-linux-gnu/libc.so.6 #2 0x000014f2c360e472 in abort () from /lib/x86_64-linux-gnu/libc.so.6 #3 0x000000000e3b2ae2 in uw_init_context_1[cold] () #4 0x000000000e3b29fc in _Unwind_Backtrace () #5 0x00000000046a6bab in __sanitizer::BufferedStackTrace::UnwindSlow(unsigned long, unsigned int) () #6 0x00000000046a181d in __sanitizer::BufferedStackTrace::Unwind(unsigned int, unsigned long, unsigned long, void*, unsigned long, unsigned long, bool) () #7 0x00000000046885bd in __sanitizer::BufferedStackTrace::UnwindImpl(unsigned long, unsigned long, void*, bool, unsigned int) () #8 0x0000000004601127 in __asan::ErrorGeneric::Print() () #9 0x0000000004683180 in __asan::ScopedInErrorReport::~ScopedInErrorReport() () #10 0x0000000004686567 in __asan::ReportGenericError(unsigned long, unsigned long, unsigned long, unsigned long, bool, unsigned long, unsigned int, bool) () #11 0x0000000004686d46 in __asan_report_load8 () #12 0x000000000dad671b in ZSTD_sizeof_CCtx (cctx=<optimized out>) at ./build/release-asan/zstd/vendor/zstd/lib/compress/zstd_compress.c:210 #13 0x0000000006d2284d in bun.js.node.zlib.NativeZstd.estimatedSize () at /var/lib/buildkite-agent/builds/ip-172-31-72-121/bun/bun/src/bun.js/node/zlib/NativeZstd.zig:57 #14 ZigGeneratedClasses.JSNativeZstd.JavaScriptCoreBindings.NativeZstd__estimatedSize (thisValue=<optimized out>) at /var/lib/buildkite-agent/builds/ip-172-31-72-121/bun/bun/build/release-asan/codegen/ZigGeneratedClasses.zig:11122 #15 0x000000000852803b in WebCore::JSNativeZstd::visitChildrenImpl<JSC::SlotVisitor> (cell=0x14f22e190840, visitor=...) at ./build/release-asan/./build/release-asan/codegen/ZigGeneratedClasses.cpp:30728 #16 WebCore::JSNativeZstd::visitChildren (cell=0x14f22e190840, visitor=...) at ./build/release-asan/./build/release-asan/codegen/ZigGeneratedClasses.cpp:30734 #17 0x000000000aa99d6c in JSC::MethodTable::visitChildren (this=<optimized out>, cell=<optimized out>, visitor=...) at vendor/WebKit/Source/JavaScriptCore/runtime/ClassInfo.h:115 #18 0x000000000aa99d6c in JSC::SlotVisitor::visitChildren (this=0x14f277028300, cell=0x14f22e190840) #19 JSC::SlotVisitor::drain(WTF::MonotonicTime)::$_0::operator()(JSC::MarkStackArray&) const (this=<optimized out>, stack=...) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitor.cpp:509 #20 0x000000000aa8f130 in JSC::SlotVisitor::forEachMarkStack<JSC::SlotVisitor::drain(WTF::MonotonicTime)::$_0>(JSC::SlotVisitor::drain(WTF::MonotonicTime)::$_0 const&) (this=0x14f277028300, func=...) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitorInlines.h:193 #21 JSC::SlotVisitor::drain (this=this@entry=0x14f277028300, timeout=<error reading variable: That operation is not available on integers of more than 8 bytes.>, timeout@entry=...) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitor.cpp:499 #22 0x000000000aa90590 in JSC::SlotVisitor::drainFromShared (this=0x14f277028300, sharedDrainMode=JSC::SlotVisitor::HelperDrain, timeout=<error reading variable: That operation is not available on integers of more than 8 bytes.>) at vendor/WebKit/Source/JavaScriptCore/heap/SlotVisitor.cpp:699 #23 0x000000000aa08726 in JSC::Heap::runBeginPhase(JSC::GCConductor)::$_1::operator()() const (this=<optimized out>) at vendor/WebKit/Source/JavaScriptCore/heap/Heap.cpp:1508 #24 WTF::SharedTaskFunctor<void (), JSC::Heap::runBeginPhase(JSC::GCConductor)::$_1>::run() (this=<optimized out>) at .WTF/Headers/wtf/SharedTask.h:91 #25 0x000000000aa3b596 in WTF::ParallelHelperClient::runTask(WTF::RefPtr<WTF::SharedTask<void ()>, WTF::RawPtrTraits<WTF::SharedTask<void ()> >, WTF::DefaultRefDerefTraits<WTF::SharedTask<void ()> > > const&) (this=0x14f22e000428, task=...) at vendor/WebKit/Source/WTF/wtf/ParallelHelperPool.cpp:110 #26 0x000000000aa3d976 in WTF::ParallelHelperPool::Thread::work (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/ParallelHelperPool.cpp:201 #27 0x000000000aa4210d in WTF::AutomaticThread::start(WTF::AbstractLocker const&)::$_0::operator()() const (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/AutomaticThread.cpp:225 #28 WTF::Detail::CallableWrapper<WTF::AutomaticThread::start(WTF::AbstractLocker const&)::$_0, void>::call() (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/Function.h:53 #29 0x0000000008958ada in WTF::Function<void ()>::operator()() const (this=<optimized out>) at vendor/WebKit/Source/WTF/wtf/Function.h:82 #30 WTF::Thread::entryPoint (newThreadContext=<optimized out>) at vendor/WebKit/Source/WTF/wtf/Threading.cpp:272 #31 0x0000000008a65689 in WTF::wtfThreadEntryPoint (context=0x13b5) at vendor/WebKit/Source/WTF/wtf/posix/ThreadingPOSIX.cpp:255 #32 0x000000000467d347 in asan_thread_start(void*) () #33 0x000014f2c36711f5 in ?? () from /lib/x86_64-linux-gnu/libc.so.6 #34 0x000014f2c36f189c in ?? () from /lib/x86_64-linux-gnu/libc.so.6 ``` </details> `ZSTD_sizeof_CCtx` and `ZSTD_sizeof_DCtx` can not be relied upon to be thread-safe and estimatedSize may be called from any thread
pull Bot
pushed a commit
that referenced
this pull request
Aug 20, 2025
…Worker" (oven-sh#21994) Reverts oven-sh#21962 `vm.ensureTerminationException` allocates a JSString, which is not safe to do from a thread that doesn't own the API lock. ```ts Bun Canary v1.2.21-canary.1 (f706382a) Linux x64 (baseline) Linux Kernel v6.12.38 | musl CPU: sse42 popcnt avx avx2 avx512 Args: "/var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/release/bun-linux-x64-musl-baseline-profile/bun-profile" "/var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/test/js/node/worker_threads"... Features: bunfig http_server jsc tsconfig(3) tsconfig_paths workers_spawned(40) workers_terminated(34) Builtins: "bun:main" "node:worker_threads" Elapsed: 362ms | User: 518ms | Sys: 63ms RSS: 0.34GB | Peak: 100.36MB | Commit: 0.34GB | Faults: 0 | Machine: 8.17GB panic(main thread): Segmentation fault at address 0x0 oh no: Bun has crashed. This indicates a bug in Bun, not your code. To send a redacted crash report to Bun's team, please file a GitHub issue using the link below: http://localhost:38809/1.2.21/Ba2f706382wNgkgUu11luEm6yX+lwy+Dgtt+oEurthoD8214mE___07+09DA2AA 6 | describe("Worker destruction", () => { 7 | const method = ["Bun.connect", "Bun.listen", "fetch"]; 8 | describe.each(method)("bun when %s is used in a Worker that is terminating", method => { 9 | // fetch: ASAN failure 10 | test.skipIf(isBroken && method == "fetch")("exits cleanly", () => { 11 | expect([join(import.meta.dir, "worker_thread_check.ts"), method]).toRun(); ^ error: Command /var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/test/js/node/worker_threads/worker_thread_check.ts Bun.connect failed: Spawned 10 workers RSS 79 MB Spawned 10 workers RSS 87 MB Spawned 10 workers RSS 90 MB at <anonymous> (/var/lib/buildkite-agent/builds/ip-172-31-38-185/bun/bun/test/js/node/worker_threads/worker_destruction.test.ts:11:73) ✗ Worker destruction > bun when Bun.connect is used in a Worker that is terminating > exits cleanly [597.56ms] ✓ Worker destruction > bun when Bun.listen is used in a Worker that is terminating > exits cleanly [503.47ms] » Worker destruction > bun when fetch is used in a Worker that is terminating > exits cleanly 1 pass 1 skip 1 fail 2 expect() calls Ran 3 tests across 1 file. [1125.00ms] ======== Stack trace from GDB for bun-profile-28234.core: ======== Program terminated with signal SIGILL, Illegal instruction. #0 crash_handler.crash () at crash_handler.zig:1523 [Current thread is 1 (LWP 28234)] #0 crash_handler.crash () at crash_handler.zig:1523 #1 0x0000000002db77aa in crash_handler.crashHandler (reason=..., error_return_trace=0x0, begin_addr=...) at crash_handler.zig:471 #2 0x0000000002db2b55 in crash_handler.handleSegfaultPosix (sig=<optimized out>, info=<optimized out>) at crash_handler.zig:792 #3 0x0000000004716b58 in WTF::jscSignalHandler (sig=11, info=0x7ffe54051e90, ucontext=0x0) at vendor/WebKit/Source/WTF/wtf/threads/Signals.cpp:548 #4 <signal handler called> #5 JSC::VM::currentThreadIsHoldingAPILock (this=0x148296c30000) at vendor/WebKit/Source/JavaScriptCore/runtime/VM.h:840 #6 JSC::sanitizeStackForVM (vm=...) at vendor/WebKit/Source/JavaScriptCore/runtime/VM.cpp:1369 #7 0x0000000003f4a060 in JSC::LocalAllocator::allocate(JSC::Heap&, unsigned long, JSC::GCDeferralContext*, JSC::AllocationFailureMode)::{lambda()#1}::operator()() const (this=<optimized out>) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/LocalAllocatorInlines.h:46 #8 JSC::FreeList::allocateWithCellSize<JSC::LocalAllocator::allocate(JSC::Heap&, unsigned long, JSC::GCDeferralContext*, JSC::AllocationFailureMode)::{lambda()#1}>(JSC::LocalAllocator::allocate(JSC::Heap&, unsigned long, JSC::GCDeferralContext*, JSC::AllocationFailureMode)::{lambda()#1} const&, unsigned long) (this=0x148296c38e48, cellSize=16, slowPath=...) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/FreeListInlines.h:46 #9 JSC::LocalAllocator::allocate (this=0x148296c38e30, heap=..., cellSize=16, deferralContext=0x0, failureMode=JSC::AllocationFailureMode::Assert) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/LocalAllocatorInlines.h:44 #10 JSC::GCClient::IsoSubspace::allocate (this=0x148296c38e30, vm=..., cellSize=16, deferralContext=0x0, failureMode=JSC::AllocationFailureMode::Assert) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/IsoSubspaceInlines.h:34 #11 JSC::tryAllocateCellHelper<JSC::JSString, (JSC::AllocationFailureMode)0> (vm=..., size=16, deferralContext=0x0) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/JSCellInlines.h:192 #12 JSC::allocateCell<JSC::JSString> (vm=..., size=16) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/JSCellInlines.h:212 #13 JSC::JSString::create (vm=..., value=...) at cache/webkit-a73e665a39b281c5/include/JavaScriptCore/JSString.h:204 #14 0x0000000004479ad1 in JSC::jsNontrivialString (vm=..., s=...) at vendor/WebKit/Source/JavaScriptCore/runtime/JSString.h:846 #15 JSC::VM::ensureTerminationException (this=0x148296c30000) at vendor/WebKit/Source/JavaScriptCore/runtime/VM.cpp:627 #16 JSGlobalObject__requestTermination (globalObject=<optimized out>) at ./build/release/./src/bun.js/bindings/ZigGlobalObject.cpp:3979 #17 0x0000000003405ab8 in bun.js.web_worker.notifyNeedTermination (this=0x542904f0d80) at /var/lib/buildkite-agent/builds/ip-172-31-16-28/bun/bun/src/bun.js/web_worker.zig:558 #18 0x0000000004362b6f in WebCore::Worker::terminate (this=0x984c900000000000) at ./src/bun.js/bindings/webcore/Worker.cpp:266 #19 WebCore::jsWorkerPrototypeFunction_terminateBody(JSC::JSGlobalObject*, JSC::CallFrame*, WebCore::JSWorker*)::{lambda()#1}::operator()() const (this=<optimized out>) at ./build/release/./src/bun.js/bindings/webcore/JSWorker.cpp:549 #20 WebCore::toJS<WebCore::IDLUndefined, WebCore::jsWorkerPrototypeFunction_terminateBody(JSC::JSGlobalObject*, JSC::CallFrame*, WebCore::JSWorker*)::{lambda()#1}>(JSC::JSGlobalObject&, JSC::ThrowScope&, WebCore::jsWorkerPrototypeFunction_terminateBody(JSC::JSGlobalObject*, JSC::CallFrame*, WebCore::JSWorker*)::{lambda()#1}&&) (lexicalGlobalObject=..., throwScope=..., valueOrFunctor=...) at ./src/bun.js/bindings/webcore/JSDOMConvertBase.h:174 #21 WebCore::jsWorkerPrototypeFunction_terminateBody (lexicalGlobalObject=<optimized out>, callFrame=<optimized out>, castedThis=<optimized out>) at ./build/release/./src/bun.js/bindings/webcore/JSWorker.cpp:549 #22 WebCore::IDLOperation<WebCore::JSWorker>::call<&WebCore::jsWorkerPrototypeFunction_terminateBody, (WebCore::CastedThisErrorBehavior)0> (lexicalGlobalObject=..., operationName=..., callFrame=...) at ./src/bun.js/bindings/webcore/JSDOMOperation.h:63 #23 WebCore::jsWorkerPrototypeFunction_terminate (lexicalGlobalObject=<optimized out>, callFrame=0x7ffe540536b8) at ./build/release/./src/bun.js/bindings/webcore/JSWorker.cpp:554 #24 0x000014825580c038 in ?? () #25 0x00007ffe540537b0 in ?? () #26 0x0000148255a626cb in ?? () #27 0x0000000000000000 in ?? () 1 crashes reported during this test ```
pull Bot
pushed a commit
that referenced
this pull request
Apr 17, 2026
) ## Problem Fuzzilli hit a flaky SIGSEGV (fingerprint `2519cad1804eace1`) from: ```js const v13 = Bun.jest().vi; try { v13.mock("function f2() {\n const v6 = new ArrayBuffer();\n ...\n}"); } catch (e) {} Bun.gc(true); ``` `JSMock__jsModuleMock` calls `Bun__resolveSyncWithSource` on the specifier before validating the callback, which sends the garbage string through the resolver. The resolver's auto-install gate at `loadNodeModules` only checks `esm_ != null`; `ESModule.Package.parse` accepts anything that doesn't start with `.` or contain `\` / `%`, so the whole function source is treated as a package name. `enqueueDependencyToRoot` then calls `PackageManager.sleepUntil`, which re-enters `EventLoop.tick()` from inside a call that is itself running inside an event-loop tick: ``` #0 ConcurrentTask.PackedNextPtr.atomicLoadPtr #1 UnboundedQueue(ConcurrentTask).popBatch #3 event_loop.tickConcurrentWithCount #7 AnyEventLoop.tick #8 PackageManager.sleepUntil #9 PackageManager.enqueueDependencyToRoot #10 Resolver.resolveAndAutoInstall #16 Bun__resolveSyncWithSource #17 JSMock__jsModuleMock ``` The same path is reachable from `Bun.resolveSync`, `import()`, and `require.resolve` with any user-provided string. ## Fix Gate the auto-install branch on `strings.isNPMPackageName(esm_.?.name)`. That validator already exists and is used by `bun link`, `bun pm view`, and the bundler; it rejects newlines, spaces, braces, and anything else that could never be a registry package. Specifiers failing the check fall straight through to `.not_found` — the same result the registry fetch would eventually produce — without initializing the package manager or ticking the event loop. This is a resolver-level fix, so it covers every entry point (not just `mock.module`). It also avoids spurious network requests for garbage specifiers; on this container a single resolve of a multi-line specifier dropped from ~275ms to ~16ms. ## Tests - `test/js/bun/resolve/resolve-autoinstall-invalid-name.test.ts` stands up a local registry and verifies zero manifest requests for a set of invalid names with `--install=force`, plus a positive control that a valid name still hits the registry. - `test/js/bun/test/mock/mock-module-non-string.test.ts` gains a case for `mock.module` with newline / whitespace / bracket specifiers (with and without a callback). - Existing `test/cli/run/run-autoinstall.test.ts` (11 tests) and `test/js/bun/test/mock/mock-module.test.ts` all pass. Related: oven-sh#28945, oven-sh#28956, oven-sh#28500, oven-sh#28511. Fingerprint: `2519cad1804eace1`
pull Bot
pushed a commit
that referenced
this pull request
Apr 28, 2026
…n-sh#29800) ## What does this PR do? Fixes a use-after-free in `selectALPNCallback` (`src/bun.js/api/bun/socket.zig:17`) when a `node:tls` / `Bun.listen` TLS server with `ALPNProtocols` handles overlapping handshakes. `SSL_CTX_set_alpn_select_cb` registers on the **listener-level** `SSL_CTX`, so its `arg` is shared across every accepted connection. The old code passed the per-connection `*TLSSocket` as `arg`, so each `onOpen` overwrote the previous one. When connection A's ClientHello is processed after connection B's `onOpen` has run — and B has since been freed — the callback dereferences a dangling pointer and feeds garbage `protos` into `SSL_select_next_proto`: ``` #9 SSL_select_next_proto (..., peer=0xb10008886384bf6d, peer_len=1431130639, supported="\002h2\bhttp/1.1", supported_len=12) #10 selectALPNCallback (in="\002h2\bhttp/1.1", inlen=12, arg=<freed>) at socket.zig:23 #11 bssl::ssl_negotiate_alpn #12 bssl::do_select_parameters #21 ssl_on_data at openssl.c:505 ``` (from `fetch-http2-client.test.ts` under `describe.concurrent` on `:alpine: 3.23 aarch64`) **Fix:** store the `*TLSSocket` on the per-connection `SSL` via `SSL_set_ex_data(ssl, 0, this)` and read it back from the `SSL*` parameter in the callback, ignoring the CTX-level `arg`. Slot 0 is otherwise unused in the codebase. ## How did you verify your code works? - `bun bd test test/js/node/tls/node-tls-server.test.ts test/js/node/tls/node-tls-connect.test.ts test/js/node/http2/node-http2.test.js test/js/web/fetch/fetch-http2-client.test.ts` → 349 pass, 0 fail - `bun run zig:check-all` → all platforms compile - The existing `connectionListener should emit the right amount of times, and with alpnProtocol available` test (50 parallel ALPN connections) covers the path; the original crash was caught by `fetch-http2-client.test.ts` on aarch64-musl CI
pull Bot
pushed a commit
that referenced
this pull request
May 3, 2026
…mpty keys (oven-sh#30171) ## Repro ```js await Bun.build({ entrypoints: ["./entry.ts"], loader: { "": "js", ".ts": "ts" }, }); // or new Bun.Transpiler({ define: { "": "1", FOO: '"bar"' } }).transformSync("FOO"); ``` Debug build: ``` ==590==ERROR: AddressSanitizer: SEGV on unknown address ... rdx = 0xaaaaaaaaaaaaaaaa #10 options.stringHashMapFromArrays src/options.zig:46 ``` ## Cause `JSPropertyIterator` with `skip_empty_name = true` advances `.i` to the **property position**, not a dense count of yielded entries. Both `JSBundler.Config.fromJS` (`loader`) and `JSTranspiler.Config.fromJS` (`define`) allocated arrays sized to `iter.len` and wrote at `iter.i`: ```zig var loader_names = try allocator.alloc(string, loader_iter.len); // len = 2 while (try loader_iter.next()) |prop| { loader_names[loader_iter.i] = ...; // i = 1 for ".ts" after "" is skipped } ``` With `{ "": "js", ".ts": "ts" }`, the empty key is skipped by the iterator so `.i == 1` on the only yielded entry, leaving slot 0 uninitialized (0xAA in safe builds). `stringHashMapFromArrays` then hashes that garbage, and in the bundler case `Config.deinit` later frees it. The same gap exists for any property the iterator skips internally (`value == .zero`, dead name). ## Fix Replace the raw `alloc` + `iter.i` indexing with `std.ArrayListUnmanaged` + `ensureTotalCapacityPrecise(iter.len)` + `appendAssumeCapacity` in both call sites, then store `.items` directly. The list's `.items` slice reflects exactly what was appended, so skipped entries leave no gaps. On the error path, `errdefer` frees already-appended extension names and the lists. ## Verification Two subprocess tests (`bun-build-api.test.ts`, `transpiler.test.js`): - **without fix** (`bun bd`, origin/main src): both fail — child process SEGVs in `stringHashMapFromArrays` - **with fix**: `Bun.build` succeeds with `{ success: true, outputs: 1 }`; `Transpiler` correctly replaces `FOO` → `"bar"` --------- Co-authored-by: robobun <robobun@users.noreply.github.com> Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot
pushed a commit
that referenced
this pull request
May 4, 2026
…worker panic, never retry (oven-sh#30216) ## What `bun test --isolate` / `--parallel` crashes when a test file loads a native addon whose deferred napi finalizers outlive the file. The `--parallel` coordinator then silently retries the file once, which masks the panic and lets the run exit 0. Fixes oven-sh#30205, oven-sh#30191. Supersedes oven-sh#30214 (same NapiEnv fix, but without the coordinator change, the `cleanup_hooks` retarget, or a test that actually reproduces on unpatched `main`). ## Reproduction ```sh git clone https://github.com/workglow-dev/libs && cd libs bun i && bun run build:packages bun test --timeout=30000 --parallel=4 packages/test/src/test/{util,task}/*.test.ts ``` On `main` (d484fd6), 3–4 workers crash per run with either ``` ASSERTION FAILED: isMarked(cell) JavaScriptCore/heap/Heap.cpp:1232 : void JSC::Heap::addToRememberedSet(const JSCell *) ``` or (when the slot is already being reallocated) ``` ASSERTION FAILED: m_cellState == CellState::DefinitelyWhite JavaScriptCore/JSCellInlines.h:69 : JSC::JSCell::JSCell(VM &, Structure *) ``` and in release builds the segfaults at `0x68` / `0xD0` reported in oven-sh#30205. ## Root cause Frame-pointer walk from the assertion: ``` #3 Bun::NapiHandleScope::open(Zig::GlobalObject*, bool) #4 NapiHandleScope__open #6 napi.Finalizer.run #7 napi.NapiFinalizerTask.runOnJSThread #10 event_loop.tick #11 event_loop.waitForPromise #13 VirtualMachine.loadEntryPointForTestRunner ← next test file ``` `NapiEnv::m_globalObject` is a raw `Zig::GlobalObject*`. For non-experimental addons (`nm_version != NAPI_VERSION_EXPERIMENTAL`, which is ~every real-world addon — sharp, better-sqlite3, etc.), `napi_wrap`/`napi_create_external` finalizers are **deferred** to the event loop as `NapiFinalizerTask` rather than run inside GC sweep. Objects rooted on the old global (module graph, `globalThis.*`) only become collectable when `Zig__GlobalObject__createForTestIsolation` runs `gcUnprotect(oldGlobal)`. The `DeferGC` from oven-sh#29573 ends at that function's `}`, so the next GC runs there, collects those objects, and enqueues their finalizers. Those tasks then run on the very next `eventLoop().tick()` — inside `loadEntryPointForTestRunner`'s `waitForPromise` for file N+1. `Finalizer.run` opens a `NapiHandleScope` via `env->globalObject()`, which reads `NapiHandleScopeImplStructure()` off the dead cell and writes `m_currentNapiHandleScopeImpl` on it → write barrier on an unmarked cell. The `--parallel` coordinator's `reapWorker` then re-queued the file once (`retries[idx] < 1`) into a fresh worker with no stale `NapiEnv`, which passed — so the run reported 0 fail despite multiple Bun panics in the log. ## Fix **NapiEnv retarget** (`ZigGlobalObject.cpp`, `napi.h`): `Zig__GlobalObject__createForTestIsolation` now calls `newGlobal->adoptNapiEnvsForTestIsolation(oldGlobal)` before `gcUnprotect`. Each `NapiEnv::m_globalObject` is repointed at the new global and the `Ref<NapiEnv>`s are moved over, so late finalizers open handle scopes on a live global and the envs stay owned after the old global is swept. `VirtualMachine.swapGlobalForTestIsolation` also repoints `rare_data.cleanup_hooks[*].globalThis` so `CleanupHook.eql()` stays accurate. **No retry, abort on panic** (`Coordinator.zig`): removed the per-file retry. A worker that dies mid-file is counted as one failure. If it died by a fatal signal (SIGILL/SIGTRAP/SIGABRT/SIGBUS/SIGFPE/SIGSEGV/SIGSYS — Bun's own `@trap()`, a JSC/WTF assertion, or native-addon crash), the whole run aborts with `error: a test worker process crashed with <SIG> while running <file>`. `process.exit()` / SIGKILL are still just a per-file failure and the run continues. ## Verification - `test/regression/issue/30205.test.ts` — 4 tests. Adds a tiny non-experimental addon (`isolate_finalizer_addon.c`) and a fixture pattern (`Bun.gc(true)` + module-scope `await 0` + objects rooted on `globalThis`) that crashes **8/8** on unpatched `main` and passes 8/8 with this change. - `workglow-dev/libs` full 201-file unit suite: 3× clean `--parallel=4` runs (was 3–4 crashes/run). - Gate: `git stash -- src/ && bun bd test test/regression/issue/30205.test.ts` → 3/4 fail; with fix → 4/4 pass. - `test/cli/test/isolation.test.ts`, `test/regression/issue/29519.test.ts` → pass (one pre-existing unrelated timeout in isolation.test.ts, same as oven-sh#29573). - `test/cli/test/parallel.test.ts` → all tests I touched pass; the 3 timing-sensitive scale-up/work-steal tests that fail in this container fail identically on unmodified `main`. --------- Co-authored-by: robobun <robobun@users.noreply.github.com> Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot
pushed a commit
that referenced
this pull request
May 16, 2026
…s, http (oven-sh#30722) Hardens 36 reachable security findings across the runtime, package manager, parsers, HTTP client/server, and SQL drivers. Three auto-applied fixes (#61 SSL exception leak, #68 YAML merge dedup, #104 archive overwrite precheck) were dropped: #61 introduced a use-after-free, #68 stored a non-`'static` byte view in a `'static` field, and #104 added dead gating that did not close the traversal. ### Memory safety / lifetime - #2 — Dangling proxy slice across reentrant JS getter — copy `process.env` proxy href to an owned `Vec` before reentrant getters can free the env map (`Blob.rs`) - #15 — Rollback restores dangling editor name pointer — preserve and restore `name_storage` on `detect_editor` failure (`BunObject.rs`) - #81 — Reentrant reconnect frees live handlers — only free previous handlers when `active_connections == 0` (`Listener.rs`) - #110 — Async randomFill uses stale resizable buffer pointer — fill a worker-owned scratch buffer; copy back on the JS thread after re-validating bounds (`node_crypto_binding.rs`) - #119 — Null zero-length slice UB in DOMJIT fast path — use `ffi::slice` which tolerates `(null, 0)` (`Crypto.rs`) - #67 — Raw serialization reads struct padding bytes — add explicit `_padding_*` fields with `offset_of!` proof asserts (`npm.rs`) - #74 — TLS rejection path leaks websocket refcount — route SSL/auth failures through `self.fail()` which clears `outgoing_websocket` (`websocket_client.rs`) - #108 — FD-backed fetch body leaks duplicated descriptor — close `opened_fd` unconditionally after `read_file` (`fetch.rs`) ### Untrusted-input bounds / panics - #10 — Invalid lockfile tag causes panic DoS — replace `unreachable!()` with logged error + `Tag::Uninitialized` (`dependency.rs`) - #20 — Unchecked lockfile string offsets cause OOB slice — bounds-check non-inline `String` pointers against `ctx.buffer` (`dependency.rs`) - #91 — Panic on unvalidated resolution tag — validate `ResolutionTag` discriminants on lockfile load (`Package.rs`) - #24 — Unwrap panic on unexpected 304 response — return `UnexpectedNotModified` when no cached manifest exists (`npm.rs`) - #44 — UDP port getter unwrap panic on transient state — return `undefined` when `socket` is `None` (`udp_socket.rs`) - #36 — Close reason length mismatch causes panic — clamp `body_len` to 125 and bail on overlong UTF-8 transcode (`websocket_client.rs`) - #100 — Windows pipe name length panic DoS — `debug_assert` → real bounds check (`Listener.rs`) - #60 / #111 — Windows shim stack buffer overflows — bounds-check argument and filename writes against `BUF1_LEN`/`BUF2_U16_LEN` before `copy_nonoverlapping` (`bun_shim_impl.rs`) - #76 / #101 — Unchecked bin name/entry name copies — bounds-check before slicing into `abs_dest_buf` (`bin.rs`) - #79 — `if` keyword misclassification causes parser panic — require a delimiter token before classifying (`shell_parser/parse.rs`) - #32 — Bounds check occurs after UTF-16 write — pre-flight key/value lengths before `convert_utf8_to_utf16_in_buffer` (`env_loader.rs`) - #95 — PBKDF2 digest validation allows panic-only algorithm — reject digests with no `EVP_MD` (`PBKDF2.rs`) ### DoS / resource caps - #17 — Unbounded recursion on deep TOML dotted keys — cap dotted-key segments at 512 (`toml.rs`) - #39 — Unbounded brace expansion preallocation — cap expansion count at 65536 in `Bun.$` and `Bun.braces` (`BunObject.rs`, `Expansion.rs`) - #31 — SCRAM PBKDF2 parameters accepted from server — clamp iteration count to `[4096, 10M]`, salt length to `[1, 1024]` (`PostgresSQLConnection.rs`) ### Auth / injection / traversal - #19 — Cleartext password sent after TLS downgrade — require `TLSStatus::SslOk`, not just `ssl_mode != Disable` (`MySQLConnection.rs`) - #83 — Strict TLS request reuses lax-verified pooled socket — track `established_with_reject_unauthorized` and refuse pool reuse for strict callers (`HTTPContext.rs`, `lib.rs`, `ClientSession.rs`) - #73 — IPv6 loopback prefix auth bypass — exact-match `::1` instead of `starts_with` (`server_body.rs`) - #56 — Unsanitized filename injects response headers — reject `\r`/`\n`/NUL/`"` in `content-disposition` filenames (`RequestContext.rs`) - #43 — Missing CRLF checks for signed host/auth headers — also validate `region`, `access_key_id`, and `host` (`s3_signing/credentials.rs`) - #34 — Bucket slash enables S3 host confusion — reject buckets containing `/` (`s3_signing/credentials.rs`) - #25 — Lexical symlink check permits extraction escape — track created symlinks during extraction and refuse paths that traverse them (`libarchive/lib.rs`) - #71 — bunx executes untrusted temp-cache binary — `lstat` cached binary; refuse symlinks and other-uid files (`bunx_command.rs`) ### Permission hygiene - #6 — Bin target chmod always sets mode 0777 — `0o777 & !umask` instead of `umask | 0o777` (`bin.rs`) - #23 — Process umask cleared and never restored — restore umask after probing it in `ensure_umask` (`bin.rs`) ### Parser correctness - #22 — Sign-prefixed scalar misparsed as infinity — fix Zig→Rust `&&`/`||` precedence transliteration (`yaml.rs`)
pull Bot
pushed a commit
that referenced
this pull request
Jun 23, 2026
…e re-enters the event loop (oven-sh#32597) Sentry BUN-2WJA / BUN-2WKB (~290 events combined, Windows x86_64, `http_server=True`, bun 1.2.23 through 1.3.14): ``` Segmentation fault at address 0xFFFFFFFFFFFFFFFF endWithSink src/runtime/webcore/Sink.zig:577 endFromJS src/runtime/webcore/streams.zig:1200 finalize src/runtime/webcore/streams.zig:1301 clearAndFree src/collections/baby_list.zig:148 memset (fault at 0xFFFFFFFFFFFFFFFF) ``` ## Cause The generated `JSReadable*Controller` `end()` and `close()` host functions (`src/codegen/generate-jssink.ts`) stash `m_sinkPtr` in a local, call `controller->detach()`, and only afterward dereference the stashed pointer via `endWithSink()` / `${name}__close()`: ```cpp void *ptr = controller->wrapped(); controller->detach(); // runs onClose JS synchronously return ${name}__endWithSink(ptr, lexicalGlobalObject); // derefs ptr ``` `detach()` invokes the stored `onClose` callback. For a `type: "direct"` stream this is `readDirectStream`'s `close(stream, reason)`, which calls `underlyingSource.cancel()`. That is arbitrary user code running while `ptr` is still live on the C++ stack. If the stream's `pull()` promise has already settled, `RequestContext::on_resolve_stream` is sitting in the microtask queue. Any path from `cancel()` that drains microtasks (e.g. the server-side drain points in `on_response` / `do_render_with_body`, or an explicit `drainMicrotasks()`) runs `handle_resolve_stream`, which calls `destroy_sink` and frees the `HTTPServerWritable`. `endWithSink(ptr)` then enters `end_from_js` on the freed allocation; `finalize()` reads garbage for `pooled_buffer` / `buffer.cap` / `buffer.ptr` and faults in the `memset` the allocator's free-scrub path performs. The same ordering appears in the Rust port (`streams.rs` / `Sink.rs`) unchanged. ## Fix In `${controller}__end` and `${controller}__close`, finish the native sink operation before any JS runs: 1. Call `${name}__controllerDetached(ptr, controller)` and null `m_sinkPtr` up front (so `end_from_js`'s own `signal.close()` stays a no-op, matching the previous behaviour, and so the later `detach()` won't touch the native side again). 2. Run `endWithSink(ptr)` / `close(ptr)`. 3. Call `controller->detach()` last. With `m_sinkPtr` already null it only clears `m_onPull` and fires `onClose`; by now we hold no reference into the sink, so re-entrant teardown is safe. ## Verification New ASAN-gated test in `test/js/bun/http/serve-direct-readable-stream.test.ts` reproduces the exact UAF deterministically by draining microtasks from the stream's `cancel()` callback (the test uses `require("bun:jsc").drainMicrotasks()` to force the drain that the production crash hits via the server's own drain points). <details> <summary>ASAN output on the unfixed build</summary> ``` ==22203==ERROR: AddressSanitizer: heap-use-after-free on address 0x6ee5f87602ca READ of size 1 at 0x6ee5f87602ca thread T0 #0 HTTPServerWritable::end_from_js src/runtime/webcore/streams.rs:1831 #2 JSSink::js_end_with_sink src/runtime/webcore/Sink.rs:1107 #4 WebCore::JSReadableHTTPResponseSinkController__end JSSink.cpp:620 freed by thread T0 here: #10 HTTPServerWritable::destroy src/runtime/webcore/streams.rs:1950 #11 RequestContext::destroy_sink src/runtime/server/RequestContext.rs:1930 #12 RequestContext::handle_resolve_stream src/runtime/server/RequestContext.rs:2680 #13 RequestContext::on_resolve_stream src/runtime/server/RequestContext.rs:2716 ... #24 JSC::VM::drainMicrotasks() ``` </details> With the fix the fixture completes normally. Existing suites (`serve.test.ts`, `bun-server.test.ts`, `direct-readable-stream.test.tsx`, `streams.test.js`, the sink leak tests) show no new failures against the unfixed build. --------- Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot
pushed a commit
that referenced
this pull request
Jun 26, 2026
…sweep (oven-sh#32729) ### Crash ``` ASSERTION FAILED: vm().currentThreadIsHoldingAPILock() => vm().heap.mutatorState() != MutatorState::Sweeping vendor/WebKit/Source/JavaScriptCore/runtime/JSCell.cpp(179) : bool JSC::JSCell::validateIsNotSweeping() const ``` Backtrace (from a release-asan build with asserts): ``` #3 JSC::JSCell::validateIsNotSweeping() #4 JSC::JSCell::classInfo() const #5 WTF::uncheckedDowncast<WebCore::JSResumableFetchSink>(JSValue const&) #6 ResumableFetchSinkPrototype__ondrainSetCachedValue #7 bun_runtime::webcore::fetch::fetch_tasklet::FetchTasklet::ignore_remaining_response_body #8 JSC::WeakBlock::sweep() <- inside GC sweep (Weak finalizer) #9 JSC::WeakSet::sweep() #10 JSC::PreciseAllocation::sweep() #12 JSC::Heap::finalize() #21 JSC::LocalAllocator::allocateSlowCase #23 JSC::ErrorInstance::create <- ordinary allocation kicked off GC ``` Found by the syscall fault-injection fuzzer's client-side grammar scenario (fetch/node:http with abort + transient errno on the client socket). Reproduces ~4/5 under `BUN_JSC_collectContinuously=1`. ### Cause `FetchTasklet::on_response_finalize` is the `WeakRefOwner<FetchResponse>::finalize` callback and runs inside `WeakBlock::sweep` while `MutatorState == Sweeping`. When the response body is `Locked` without a pending promise or stream it calls `ignore_remaining_response_body()`, which called: - `ResumableSink::detach_js()`: writes the sink wrapper's cached `ondrain` / `oncancel` / `stream` slots via the generated `ResumableFetchSinkPrototype__*SetCachedValue` helpers. Each does `uncheckedDowncast<JSResumableFetchSink>(thisValue)`, which reaches `JSCell::classInfo()` and then issues a write barrier on the wrapper cell. - `clear_stream_handlers()`: reaches `ReadableStreamTag__tagged` -> `object->inherits<JSReadableStream>()` (guarded today, but one boolean away). Calling `classInfo()` on any cell while the mutator is sweeping is forbidden: the cell's `Structure` may already have been swept. Assert builds catch it; release builds corrupt the heap. ### Fix Thread a `from_finalizer` flag through `ignore_remaining_response_body`. When `true` (the `on_response_finalize` caller) skip `detach_js()` and `clear_stream_handlers()`; only native state is touched. The sink's JS-side detach still happens from `clear_sink()` in `FetchTasklet::deinit()`, which runs as an event-loop `ConcurrentTask` outside any sweep, so nothing leaks. The `on_stream_cancelled_callback` caller (reader `.cancel()`, runs from JS on the event loop) passes `false` and keeps the immediate detach. Also corrects the `ResumableSink::detach_js` doc comment that claimed finalizer safety. ### Verification New test at `test/js/web/fetch/fetch-response-finalizer-sweep.test.ts`: a child process under `BUN_JSC_collectContinuously=1` does 12 iterations of `fetch()` with a user-constructed `ReadableStream` body (so the sink takes the JS route with a Strong `js_this`) against a raw TCP server that sends headers + a partial chunked body and never terminates it, then drops the `Response` unconsumed and runs `Bun.gc(true)`. Without the fix (`bun bd`, src/ stashed): ``` exitCode: 134 stderr: ASSERTION FAILED: vm().currentThreadIsHoldingAPILock() => vm().heap.mutatorState() != MutatorState::Sweeping ``` With the fix: `stdout: "ok"`, `exitCode: 0`. `test/js/web/fetch/fetch-backpressure.test.ts` (exercises the `on_stream_cancelled_callback` path) passes unchanged. --------- Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
pull Bot
pushed a commit
that referenced
this pull request
Jul 16, 2026
…ry rewrite (oven-sh#34271) `test/js/bun/util/filesystem_router.test.ts` went red on alpine x64 in build [73276](https://buildkite.com/bun/bun/builds/73276): the `reload() while Bun.build() resolves the same directory` subprocess segfaulted in `bust_dir_cache_recursive`, inlined from `NonNull::new`. ## Cause `RealFS::entries_at` (`src/resolver/lib.rs`) replaces a cached `DirEntry` in place when the caller's resolver generation is newer than the cached listing's. The replacement at `*e_ptr = new_entry` drops the old `DirEntry`, which drops its `data: StringHashMap<*mut Entry>` and frees the hashmap's bucket allocation. The function's comment says `entries_mutex held by caller`, but that is only true on one of the five paths that reach it: `dir_info_uncached`, when entered from `dir_info_cached_miss`. The other callers (`finalize_result`, `handle_esm_resolution`, `load_index_with_extension`, `Transpiler::run_env_loader`) all reach `entries_at` after `dir_info_cached_maybe_log` has already returned and released both `RESOLVER_MUTEX` and `entries_mutex`. `FileSystemRouter::reload()` and `RouteLoader::load` iterate the same `DirEntry.data` map under `entries_mutex` (the snapshot pattern oven-sh#33056 introduced for exactly this kind of concurrent rewrite). With `entries_at`'s rewrite unsynchronized, a `Bun.build()` on the bundler thread can drop the map while `reload()` on the JS thread is mid-iteration. The generation mismatch is what makes `entries_at` enter its rewrite branch, so the window only opens once the bundle thread has processed at least one batch (it bumps its own generation after every queue drain); every subsequent `Bun.build()` then re-reads any directory that `reload()` just refreshed to generation 0. ASAN catches it as a heap-use-after-free with the two sides of the race laid out exactly: ``` READ of size 16 (thread T0): #6 HashMap::values #7 StringHashMap<*mut Entry>::values src/collections/array_hash_map.rs:1864 #8 FileSystemRouter::bust_dir_cache_recursive src/runtime/api/filesystem_router.rs:395 #9 FileSystemRouter::bust_dir_cache src/runtime/api/filesystem_router.rs:451 #10 FileSystemRouter::reload src/runtime/api/filesystem_router.rs:476 freed by thread T11 (Bundler): #11 drop_in_place<bun_resolver::fs_full::DirEntry> #12 bun_resolver::fs::RealFS::entries_at src/resolver/lib.rs:1639 #13 DirInfo::get_entries_ref src/resolver/dir_info.rs:266 #14 Resolver::finalize_result src/resolver/resolver.rs:1714 #15 Resolver::resolve_and_auto_install src/resolver/resolver.rs:1485 ... #23 BundleThread::generate_in_new_thread src/bundler/BundleThread.rs:276 previously allocated by thread T0: #17 HashMap::reserve #18 Resolver::dir_info_cached_miss src/resolver/resolver.rs:4591 #19 Resolver::dir_info_cached_maybe_log src/resolver/resolver.rs:4201 #20 Resolver::read_dir_info src/resolver/resolver.rs:4118 #21 FileSystemRouter::reload src/runtime/api/filesystem_router.rs:492 ``` (The use side is sometimes `RouteLoader::load` at `src/router/lib.rs:816` instead; same map, same lock.) This has been the shape of `entries_at` since the Rust port; oven-sh#33056 narrowed the race by snapshotting under the lock but assumed the rewrite side already held it. ## Fix `entries_at` now takes `entries_mutex` itself, matching `read_directory_with_iterator` which already does. The one call path that reaches it with the lock already held (`dir_info_cached_miss` -> `dir_info_uncached` -> `parent_.get_entries_ref`) routes through a new `entries_at_locked` / `get_entries_ref_locked` pair so the non-recursive mutex is not re-entered. That path is the only one that passes a non-`None` parent to `dir_info_uncached`; the other caller (`dir_info_for_resolution`) passes `None`, so the parent branch containing the accessor never runs there. ## Test The existing concurrency test now awaits one `Bun.build()` first, so the bundle thread's generation is already past zero when the concurrent rounds start, and then runs forty reload/build rounds instead of one. That is the shape that reaches the stale-generation rewrite at all; the original single-round fixture usually completes with every build still on generation 0. The race is scheduling-dependent. Pinning the fixture to a single core reproduces the ASAN use-after-free on roughly 3 in 10 runs against an unpatched debug build and 0 in 15 with this change; with all 16 cores available the unpatched build reproduces at roughly 1 in 30. The assertions are otherwise the same as before, so the test continues to cover the behavior oven-sh#33056 added. Also ran the full `filesystem_router.test.ts`, `test/bundler/bun-build-api.test.ts` (including the thousands-of-builds test that exercises the generation path heavily), `test/js/bun/resolve/resolve.test.ts`, `test/cli/hot/hot.test.ts`, `test/cli/watch/watch.test.ts`, `test/bake/framework-router.test.ts`, and `bun run rust:check-all` (10/10 targets). <!-- robobun:evidence:begin --> --- **no test proof** · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/bun/util/filesystem_router.test.ts <!-- robobun:evidence:end -->
pull Bot
pushed a commit
that referenced
this pull request
Jul 16, 2026
…e drain (oven-sh#34278) ## Problem `test/js/node/test/parallel/test-worker-stdio-flush.js` went red on the `debian 13 x64-asan` lane of [build 73374](https://buildkite.com/bun/bun/builds/73374) with: ``` ==18202==ERROR: LeakSanitizer: detected memory leaks Direct leak of 32 byte(s) in 1 object(s) allocated from: #9 ConcurrentTask::new src/event_loop/ConcurrentTask.rs:305 #10 ConcurrentTask::create src/event_loop/ConcurrentTask.rs:319 #12 bun_jsc::virtual_machine_exports::queue_task_concurrently src/jsc/virtual_machine_exports.rs:140 #13 ScriptExecutionContext::postTaskConcurrently src/jsc/bindings/ScriptExecutionContext.cpp:266 #14 ScriptExecutionContext::postTaskTo src/jsc/bindings/ScriptExecutionContext.cpp:125 #15 MessagePortPipe::scheduleDrain src/jsc/bindings/webcore/MessagePortPipe.cpp:74 #16 MessagePort::postMessage src/jsc/bindings/webcore/MessagePort.cpp:143 ``` The leaked allocation is a `ConcurrentTask` (and the `EventLoopTask` it wraps) left in an exiting worker's `concurrent_tasks` queue after the queue has been drained for the last time. ## Cause `WebWorker::shutdown()` runs `process.on('exit')` handlers, then drains the worker's concurrent queue via `release_queued_tasks_for_shutdown()`, then enters `WebWorker__teardownJSCVM` which (first thing) calls `ctx->markTerminating()`. `ScriptExecutionContext::postTaskTo` already refuses to enqueue onto a terminating context, but between the drain and the flag flip there is a short window where a cross-thread poster still sees `isTerminating() == false` and enqueues. In the failing test the worker writes to `process.stdout` inside its `exit` handler. The parent's captured-stdout reader acks each chunk with `port.postMessage(true)` (`src/js/node/worker_threads.ts` `makePortReadable._read`), which routes through `MessagePortPipe::scheduleDrain` to `postTaskTo(workerCtxId, ...)`. When the ack lands in that window it is pushed onto the worker's `concurrent_tasks`; nothing drains it again, and the worker's VM box is `dealloc`'d raw, so LSan reports the `ConcurrentTask` as a direct leak. The window is a few assignments plus one FFI call wide, so it hits probabilistically; the `release-asan` build is fast enough to line up occasionally, debug essentially never. The ordering was introduced in oven-sh#31216; oven-sh#29917 described the same gap ("a task posted between this drain and `removeFromContextsMap()` inside `teardownJSCVM` still leaks") but left it open. ## Fix - `ScriptExecutionContext::markTerminating()` now takes `allScriptExecutionContextsMapLock`, the same lock `postTaskTo` holds across its `isTerminating()` check and `postTaskConcurrently()` enqueue. That makes the flag flip a proper fence against concurrent posters: any `postTaskTo` critical section either runs entirely before `markTerminating()` (its task is visible to the subsequent drain) or entirely after (it observes `true` and drops). - `WebWorker::shutdown()` calls the new `extern "C" ScriptExecutionContext__markTerminating` immediately before `release_queued_tasks_for_shutdown()`, closing the window. The later `markTerminating()` inside `WebWorker__teardownJSCVM` is now redundant but harmless. No behaviour change for `process.on('exit')` itself: that runs before the new call, so a parent ack posted while the handler is running is still enqueued and then freed by the drain (never executed, same as before). Only posts that would have landed after the drain are now dropped instead of leaked. ## Verification The gap is too narrow to reproduce unassisted against a debug build: 200 iterations of the Node test with the CI LSan env, and 150 worker shutdowns with 64 Atomics-synchronized MessagePorts each, all pass on an unpatched `bun bd`. Widening the gap with a temporary `std::thread::sleep(5ms)` between `release_queued_tasks_for_shutdown()` and `WebWorker__teardownJSCVM` makes it deterministic: | build | `test-worker-stdio-flush.js` under LSan | 200-port Atomics-synchronized probe | | --- | --- | --- | | unpatched + 5 ms sleep | 5/5 leak (`32 byte(s) ConcurrentTask`) | 5/5 leak | | this PR + 5 ms sleep | 10/10 clean | 5/5 clean | | this PR (no sleep) | 50/50 clean | clean | `test/js/node/worker_threads/worker-shutdown-post-leak.test.ts` runs the worker-stdio-on-exit scenario under `detect_leaks=1` as an ASAN-lane guard (in a fresh file so it actually runs; `worker_destruction.test.ts` is ASAN-quarantined via `test/expectations.txt`). The race is not observable on the debug gate without `src/` instrumentation, so the fail-before half will not fire there; `test-worker-stdio-flush.js` on the release-asan lane remains the primary signal. Related: oven-sh#31216 (introduced the ordering), oven-sh#29917 (described but left the remaining window). <!-- robobun:evidence:begin --> --- **no test proof** · iteration 1 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/node/worker_threads/worker-shutdown-post-leak.test.ts <!-- robobun:evidence:end -->
pull Bot
pushed a commit
that referenced
this pull request
Jul 29, 2026
…rier (oven-sh#36337) `JSNativeStreamSourceAdapter::m_controller` was a `JSC::Weak<JSReadableStreamDefaultController>`. When the native pull promise is rejected (socket fault on a fetch body) the adapter is queued as the `onNativePullRejected` reaction context, which roots the **adapter** but not the **controller**: the adapter's only edge to it was the `Weak`. `FetchTasklet` releases both native `Strong<>`s to the body stream before that microtask drains, so a GC in between can leave the entire consumer graph (`controller -> stream -> reader -> pipe op -> destination -> writer -> readyPromise`) white. The subsequent error cascade then enqueues the pipe's writes-drained shutdown deferral against a corpse `op`, and `performPipeShutdownAction(AbortDestination)` dereferences a swept `readyPromise`: ``` ASSERTION FAILED: result JSObject.h(583) JSGlobalObject *JSC::JSObject::realm() const #5 JSC::JSObject::realm() #6 JSC::JSPromise::rejectPromise #7 JSC::JSPromise::reject #8 Bun::WebStreams::writableStreamDefaultWriterEnsureReadyPromiseRejected #9 Bun::WebStreams::writableStreamStartErroring #10 Bun::WebStreams::writableStreamAbort #11 WebCore::performPipeShutdownAction (AbortDestination) #12 WebCore::JSStreamPipeToOperation::onWritesFinishedForShutdown ``` On builds without the assert the same path is a silent write into freed/reused promise memory. ## Fix Hold `m_controller` as a visited internal field so a queued adapter roots the controller directly. The edge is cleared on every terminal path (`nativeSourcePullRejected`, `nativeSourceCallClose`, `nativeSourceCancel`); `controller->algorithmContext` is cleared by `readableStreamDefaultControllerClearAlgorithms`, so the abandoned case is an ordinary intra-heap cycle mark-sweep collects. `NewSource::this_jsvalue` is only `Strong` during FileReader I/O, where pinning the consumer graph is the correct behavior anyway. With the `Weak` gone the adapter no longer needs a destructor, so it is now a `JSInternalFieldObjectImpl<5>`: the five JSValue members (handle, pendingView, closer, drainValue, controller) are internal fields visited by the base class, with typed accessors at call sites. The scalar members (chunkSize, flag bitfield, text-decode state) stay as plain members. ## Verification `native-source-onclose-leak.test.ts` (the partial-read + `releaseLock` abandonment tests for Blob/fetch/File sources) continues to pass, confirming the cycle does not pin. `streams.test.js`, `pipeTo-signal-leak.test.ts`, `compression.test.ts`, `blob.test.ts` all pass. The crash itself is 0/1800 standalone; it reproduces ~1/3 only under a fault-injected tracer replay. `pipeTo-shutdown-gc.test.ts` exercises the shape (native body source, socket fault mid-stream, fire-and-forget `pipeTo` under `collectContinuously`, `AbortDestination` shutdown arm) as a regression surface. <!-- robobun:evidence:begin --> --- **no test proof** · iteration 2 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/web/streams/pipeTo-shutdown-gc.test.ts <!-- robobun:evidence:end -->
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.1)
Can you help keep this open source service alive? 💖 Please sponsor : )