Skip to content

fix(orchestrator): survive SIGBUS from failing disks under mmap'd caches - #3385

Merged
arkamar merged 19 commits into
mainfrom
fix/mmap-sigbus-fault-recovery
Jul 29, 2026
Merged

fix(orchestrator): survive SIGBUS from failing disks under mmap'd caches#3385
arkamar merged 19 commits into
mainfrom
fix/mmap-sigbus-fault-recovery

Conversation

@arkamar

@arkamar arkamar commented Jul 24, 2026

Copy link
Copy Markdown
Member

A single bad 4K sector on a node's local NVMe crashed the whole
orchestrator and killed every sandbox on the node: the kernel cannot
return EIO for a memory access, so a failed page-in from an mmap'd cache
file delivers SIGBUS, which Go turns into an unrecoverable fatal error
inside the Cache.ReadAt memmove.

RunFaultSafe converts such faults into a regular error carrying the
fault address (debug.SetPanicOnFault + recover); other panics still
crash. It guards the in-process mmap copies: build segment reads, NBD
dispatch read/write/write-zeroes and the dedup compare. A bad block now
fails a single read or NBD request instead of the whole orchestrator,
and the affected file is just a local cache of remote storage anyway.
The guard is cheap enough to sit on the hot read path.

Recovered faults are logged and counted in
orchestrator.block.memory_fault, which should stay at zero on a healthy
node.

arkamar added 16 commits July 24, 2026 16:26
…errors

An unrecoverable disk read error (bad sector) under a block cache file
raises SIGBUS when the kernel pages in the mapped range. The memmove in
Cache.ReadAt cannot return an error, so the Go runtime turns the fault
into 'fatal error: unexpected fault address' and the whole orchestrator
dies, killing every sandbox on the node.

Add subprocess-based tests (truncated backing file = deterministic
SIGBUS) demonstrating the crash in the three affected paths:
- RunFaultSafe helper (skeleton, pass-through for now)
- build.readSegments (parallel + sequential)
- NBD dispatch read serving Overlay/Cache mmap

The fault conversion lands in the next commit, turning these red tests
green.
Implement block.RunFaultSafe using debug.SetPanicOnFault + recover: a
memory fault raised on the calling goroutine (SIGBUS paging in an mmap'd
cache file backed by an unreadable disk block) becomes an error wrapping
block.ErrMemoryFault instead of a fatal, unrecoverable runtime throw.
Non-fault panics re-panic, so real bugs stay loud. Faults are detected
by the Addr method only fault-derived runtime.Errors carry.

Guard the in-process mmap copy paths:
- build.readSegments (parallel + sequential): covers sandbox start
  prefetch, NBD template reads and uffd reads through File.ReadAt
- NBD dispatch cmdRead/cmdWrite/cmdWriteZeroes: covers Overlay COW
  cache access; the guest gets an NBD error reply (EIO) per request
- Cache.Dedup compare pass

Syscall-based consumers (pwritev, UFFDIO_COPY, socket writes) already
get EFAULT as an ordinary error and need no guard.

A fault now fails one read/request instead of killing the orchestrator
and every sandbox on the node, and is logged with the cache path as a
disk-health signal. In the motivating incident a single bad 4K sector
under one template's cache file crashed the whole node 20 minutes after
the disk started failing.
Intel Core Ultra 7 268V, linux/amd64, -benchtime=2s:

  size=4KiB/unguarded      68.02 ns/op   60215 MB/s
  size=4KiB/guarded        85.34 ns/op   47995 MB/s
  size=256KiB/unguarded     8158 ns/op   32134 MB/s
  size=256KiB/guarded       7872 ns/op   33299 MB/s

~17 ns per call (SetPanicOnFault bool swap + two defers + recover),
independent of size: noticeable only relative to the smallest possible
read and lost in the noise at typical segment sizes.
Add orchestrator.block.memory_fault, incremented inside RunFaultSafe so
every guarded site (build segment reads, NBD dispatch, dedup) is
covered centrally. A non-zero rate on a node means its local disk is
likely failing and the node should be drained before more blocks rot.

The subprocess fault test asserts the counter increments on a recovered
fault.
Thread the request context through RunFaultSafe so the memory-fault
counter can attach trace exemplars, linking a fault datapoint to the
exact read that hit it. The context does not cancel fn, which runs
synchronously; call sites keep their own contextual error logs.
The runtime discards the signal details when SetPanicOnFault converts a
fault: the panic value's message is the generic 'runtime error: invalid
memory address or nil pointer dereference' for SIGBUS and SIGSEGV
alike, and the faulting address is only exposed via its Addr method.

Return *MemoryFaultError carrying Addr instead of a plain wrap, stdlib
style (fs.ErrNotExist / *fs.PathError): errors.Is(err, ErrMemoryFault)
still matches via the Is method, errors.As extracts the address, and
Unwrap keeps the runtime error in the chain. The address can be
correlated with the mmap layout to find the file offset — and hence the
disk block — that failed; the recovery-point logs now carry it as a
structured fault_addr field.
Keep the SIGBUS/bad-sector rationale once, on ErrMemoryFault and
RunFaultSafe; call sites and tests reference it instead of retelling it.
Names and assertions carry the rest; the disk-failure rationale lives on
ErrMemoryFault and the metric description.
Production code matches *MemoryFaultError with errors.As; the sentinel
and its Is method had no consumers outside tests. One error type is
enough, matching the CacheClosedError convention.
RunFaultSafe converts any memory fault, not just those from mapped
files; the type doc keeps the typical cause.
The counter covers any fault RunFaultSafe recovers; the typical cause
belongs to the type doc, not the metric.
This reverts commit 5c00af3355d7ce89da73ceb999b3545dbaf86879.
Remove churn left by earlier iterations: the error-message formatting
assertion (the typed Addr field already covers it), the clean-return
half of the flag-restore test (only the defer ordering after a panic is
a real contract), and the generic metric-iteration helper (inlined at
its single call site).

Add the missing cmdWrite wiring coverage to the dispatch fault child: a
non-zero-payload write into the truncated mmap faults on the store
(all-zero writes take the punch-hole path and never touch the mapping)
and must come back as an NBD error reply. Verified red with the
cmdWrite guard removed. cmdWriteZeroes stays untested deliberately:
whether it faults depends on the MADV_REMOVE vs clear() fallback, so
the assertion would be flaky by design; the Dedup wiring remains a
known gap.
The subprocess re-exec pattern guarded against an unconverted fault
killing the test binary, but that failure only happens when the guard
itself is broken in development, is unmissable in CI, and the crash
dump points straight at the unguarded copy. Meanwhile the pattern
carried a real trap: the -test.run filter must track the test name, and
a rename would leave a child that runs zero tests and passes forever.

Fold the child bodies into the tests and note in each that a binary
crash with 'unexpected fault address' is the expected failure mode.
gochecknoglobals is disabled repo-wide, so its nolint silenced nothing.
The paralleltest suppression on the flag-restore test was justified by
a false claim: SetPanicOnFault is per-goroutine, not process-wide, so
the test can simply run in parallel. Only the counter-swap suppression
does real work and stays.
The swap helper was designed for the subprocess era, where it ran in an
isolated child process; after the in-process conversion it forced the
test to be non-parallel and carried the branch's last nolint. Testing a
one-line counter Add is not worth a global swap; the counter itself
stays as the operational signal.
@cursor

cursor Bot commented Jul 24, 2026

Copy link
Copy Markdown

PR Summary

Medium Risk
Changes sit on hot NBD and dedup paths site-wide; behavior is mostly fail-fast on disk errors, but unguarded mmap reads (e.g. direct Cache.ReadAt) can still crash the process if a fault happens there.

Overview
Bad sectors under mmap'd local cache files used to raise SIGBUS and kill the orchestrator. This adds RunFaultSafe, which turns those faults into a MemoryFaultError (with fault address and an orchestrator.block.memory_fault counter) instead of a fatal runtime throw, and wraps mmap-heavy work on dedup compare, build segment reads, and NBD read/write/write-zeroes so one bad block fails a single operation or NBD request. Linux integration tests simulate truncated mmap backing files to ensure faults surface as errors and the NBD dispatch loop keeps running.

Reviewed by Cursor Bugbot for commit 79be12f. Bugbot is set up for automated code reviews on this repo. Configure here.

@codecov

codecov Bot commented Jul 24, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 87.01299% with 10 lines in your changes missing coverage. Please review.
✅ All tests successful. No failed tests found.

Files with missing lines Patch % Lines
packages/orchestrator/pkg/sandbox/block/cache.go 53.33% 6 Missing and 1 partial ⚠️
packages/orchestrator/pkg/sandbox/block/fault.go 85.00% 3 Missing ⚠️

📢 Thoughts on this report? Let us know!

Comment thread packages/orchestrator/pkg/sandbox/nbd/dispatch.go

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline finding, I checked one other fault-safety edge case and ruled it out: in nbd/dispatch.go's backing Cache.WriteAtWithoutLock, the dirty-tracker update (c.tracker.SetRange(..., Dirty)) runs only after the copy() into the mmap completes, so a fault during that copy propagates as an error before the block is ever marked Dirty — it can't be silently recorded as valid after a partial/faulted write.

Extended reasoning...

Verified directly in packages/orchestrator/pkg/sandbox/block/cache.go: the flush closure in WriteAtWithoutLock calls copy((*c.mmap)[runStart:runEnd], b[...]) before c.tracker.SetRange(startIdx, endIdx, Dirty). Since RunFaultSafe wraps the whole WriteAt call in the NBD dispatch path, a SIGBUS during that copy() unwinds via panic/recover before the subsequent SetRange line ever executes, so the block is never marked Dirty when the write only partially landed. This rules out the candidate concern that a fault-recovered write could leave an already-Dirty block silently treated as valid data.

Comment on lines 252 to +270
return nil
}

// readSegmentFaultSafe converts memory faults from the diff's mmap'd cache
// into errors, so a bad disk block fails one read, not the process.
func (b *File) readSegmentFaultSafe(ctx context.Context, p []byte, s readSegment) error {
err := block.RunFaultSafe(ctx, func() error { return b.readSegment(ctx, p, s) })
var faultErr *block.MemoryFaultError
if errors.As(err, &faultErr) {
cachePath, _ := s.diff.CachePath(ctx)
logger.L().Error(ctx, "memory fault reading build segment; local disk under the cache is likely failing",
zap.Error(err),
zap.String("cache_path", cachePath),
zap.Uintptr("fault_addr", faultErr.Addr),
zap.Int64("offset", s.srcOff),
zap.Int64("length", s.length),
)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟣 The P2P build-serving path (peerserver/seekable.go Stream()ChunkService.ReadAtBuildSeekable) is not covered by the new RunFaultSafe guard: it returns a zero-copy mmap slice via diff.Slice(...) and hands it straight to sender.Send(...), where gRPC's proto marshaling performs the actual page-in memcpy outside any fault protection. A bad sector under a locally-cached build served to a peer will still raise an unrecovered SIGBUS and crash the whole orchestrator — the exact failure mode this PR fixes elsewhere — but this is pre-existing behavior in code the PR does not touch (peerserver/seekable.go, server/chunks.go), so it's not a regression introduced here.

Extended reasoning...

What the bug is: This PR introduces block.RunFaultSafe (debug.SetPanicOnFault + recover) specifically to convert SIGBUS from bad sectors under mmap'd caches into a normal error instead of a process-killing fatal throw, and applies it to build.File.readSegments, NBD read/write/write-zeroes dispatch, and the dedup compare. However, the P2P build-serving path used to share a locally-built template with peer orchestrators is not wrapped, and it performs the exact same class of raw mmap access.

The code path: Server.ReadAtBuildSeekable (packages/orchestrator/pkg/server/chunks.go) calls src.Stream(ctx, offset, length, &seekableStreamSender{stream}). seekableSource.Stream (packages/orchestrator/pkg/sandbox/template/peerserver/seekable.go) does data, _ := f.diff.Slice(ctx, offset, length, nil) and then sender.Send(data[:take]). diff.Slice resolves through StorageDiff.SliceChunker.Slice, whose cached/fetched fast paths return c.cache.Slice(...)/c.cache.sliceDirect(...), both of which are just (*c.mmap)[off:end] — a zero-copy slice header with no byte access at all. The actual page-in happens later, inside seekableStreamSender.Sendstream.Send(&ReadAtBuildSeekableResponse{Data: data}), where gRPC's proto marshaler copies every byte of that mmap-backed slice into the wire buffer, synchronously on the RPC handler goroutine. That goroutine never calls RunFaultSafe/debug.SetPanicOnFault.

Why existing guards don't help: Wrapping Slice itself (the way one might expect, mirroring readSegmentFaultSafe) would not fix this, since Slice never touches the mapped pages — the fault surfaces later, once the zero-copy slice escapes into gRPC's marshaler. The fix has to guard the consumer, e.g. copy the slice into a freshly allocated buffer inside RunFaultSafe before calling sender.Send.

Impact: P2P build serving is reachable whenever a template has been built locally but not yet uploaded to remote storage (the common case for freshly-built or newly-scheduled templates) — server/chunks.go only falls back to remote-storage serving once uploaded. So this is a live, non-exotic path. A bad NVMe sector under the locally-cached memfile/rootfs mmap, hit while serving a peer's ReadAtBuildSeekable request, raises an unrecovered SIGBUS and crashes the whole orchestrator process, killing every sandbox on the node — precisely the blast radius this PR sets out to eliminate.

Step-by-step proof:

  1. Orchestrator A builds a template; the build's memfile diff is cached locally as a memory-mapped file, not yet uploaded to GCS.
  2. Orchestrator B needs the same build and calls ChunkService.ReadAtBuildSeekable against A for a range.
  3. A's Server.ReadAtBuildSeekable handler resolves the local diff and calls seekableSource.Stream.
  4. Stream calls f.diff.Slice(ctx, offset, length, nil), which returns (*c.mmap)[off:end] — no memory access yet, so even if the backing block is a bad sector, nothing has faulted.
  5. Stream calls sender.Send(data[:take])stream.Send(...); gRPC's codec marshals the Data field, reading every byte of the slice — this read hits the bad sector's mapped page and the kernel delivers SIGBUS instead of EIO (mmap has no error-return path).
  6. Because this goroutine never set debug.SetPanicOnFault(true), the SIGBUS becomes an unrecoverable Go fatal runtime throw, not a catchable panic — the whole orchestrator process dies, taking every sandbox on the node down with it.

Why this is pre-existing rather than a PR regression: the PR does not touch peerserver/seekable.go or server/chunks.go, and this path was equally vulnerable before the change — the PR is still a strict improvement over the prior fully-unguarded state on the paths it does touch. It's flagged here because it's a real, in-theme completeness gap directly undercutting the PR's stated goal ("survive SIGBUS from failing disks under mmap'd caches"), and the author may want to close it in a fast follow-up by copying the slice into a fresh buffer inside RunFaultSafe before Send.

@ValentaTomas

Copy link
Copy Markdown
Member

Good find, though, cannot the same thing happen for the Cache struct that also uses mmap?

@ValentaTomas
ValentaTomas force-pushed the main branch 2 times, most recently from 5aad415 to d71980e Compare July 25, 2026 22:53

@jakubno jakubno left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

I added few related follow up tickets EN-2000 and EN-2003

Comment thread packages/orchestrator/pkg/sandbox/build/build.go Outdated
Co-authored-by: Jakub Rojko <jakub@e2b.dev>
Comment thread packages/orchestrator/pkg/sandbox/build/build.go Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit edd284e. Configure here.

Comment thread packages/orchestrator/pkg/sandbox/block/cache.go
runDedup called dedupCompare without the fault guard Cache.Dedup has,
and it is the default pause path. The compare reads base pages that can
come from a disk-backed mmap, so a bad sector there still killed the
process. A fault now fails the pause instead.
@arkamar
arkamar merged commit 728bba3 into main Jul 29, 2026
98 of 100 checks passed
@arkamar
arkamar deleted the fix/mmap-sigbus-fault-recovery branch July 29, 2026 13:09
charlie-e2b added a commit that referenced this pull request Jul 30, 2026
🤖 I have created a release *beep* *boop*
---


## 0.0.1 (2026-07-30)


### Features

* **api:** add sandbox IAM workload token configuration
([13ddb3d](13ddb3d))
* **api:** add sandbox workload identity permission
([#3319](#3319))
([13ddb3d](13ddb3d))
* **api:** SOCKS5 egress proxy on sandbox network config (BYOP)
([#2642](#2642))
([1fc3820](1fc3820))
* **cfg:** add DISABLE_STARTUP_RECLAIM flag
([#3081](#3081))
([7677ca6](7677ca6))
* **clickhouse:** implement multi-cluster fan-out for events and stats
([#2925](#2925))
([39594c6](39594c6))
* dynamic sandbox log routing and ClickHouse-backed log reads
([#3236](#3236))
([1b19a3b](1b19a3b))
* **envd:** give envd realtime IO priority, reset for user processes
([#2681](#2681))
([f4bd1b2](f4bd1b2))
* **envd:** split collapse stats into real migrations vs already-huge
([#3021](#3021))
([0d77614](0d77614))
* **envd:** support user-defined file metadata via xattrs
([#2732](#2732))
([da8fbe4](da8fbe4))
* **featureflags:** support per-service context providers
([#3100](#3100))
([65297c1](65297c1))
* freeze user cgroup across pause/resume to keep envd /init responsive
([#2688](#2688))
([eceb741](eceb741))
* **metrics:** break down pause-snapshot latency by step
([#3426](#3426))
([657559e](657559e))
* **metrics:** label pause telemetry by fs_only
([#3425](#3425))
([411b63e](411b63e))
* **observability:** add kill_reason to sandbox.lifecycle.killed
([#2833](#2833))
([e45418f](e45418f))
* **observability:** include kill_reason in kill-path structured logs
([#2846](#2846))
([33c49f7](33c49f7))
* **orch:** add envd-version to LaunchDarkly sandbox context
([#3051](#3051))
([37d3b92](37d3b92))
* **orch:** add less, nftables, iputils-ping, and jq to base
provisioning ([#2736](#2736))
([a1e010e](a1e010e))
* **orch:** collapse envd's heap into 2 MiB hugepages before pause to
cut cold-resume faults
([#2997](#2997))
([6677f73](6677f73))
* **orch:** debug a sandbox guest kernel with resume-build -gdb
([#3040](#3040))
([37bb0dc](37bb0dc))
* **orch:** decouple warm resume from memfile dedup
([#3166](#3166))
([77f25a0](77f25a0))
* **orch:** distro-aware template base-image provisioning
([#3411](#3411))
([f8c7b5b](f8c7b5b))
* **orchestrator/cgroup:** list and destroy leaked sandbox cgroups
([#3086](#3086))
([bce1d84](bce1d84))
* **orchestrator/nbd:** inspect and disconnect connected devices
([#3087](#3087))
([4d47148](4d47148))
* **orchestrator/network:** list slot namespaces
([#3089](#3089))
([c23dbc7](c23dbc7))
* **orchestrator/network:** list slot namespaces
([#3090](#3090))
([fbfce25](fbfce25))
* **orchestrator:** add -force-reboot to resume-build to cold-boot
memory-snaphsot builds
([#3208](#3208))
([cf8f15b](cf8f15b))
* **orchestrator:** add allocated resource metrics for sandboxes
([#2943](#2943))
([95cb6d3](95cb6d3))
* **orchestrator:** add dummy orchestrator binary for local API dev
([#2744](#2744))
([ab56e25](ab56e25))
* **orchestrator:** add NetworkAssignHook for sandbox lifecycle
extensions ([#3290](#3290))
([3261963](3261963))
* **orchestrator:** add soft-delete marker label to the check metric
([#3144](#3144))
([1ce64f8](1ce64f8))
* **orchestrator:** add v4HeaderForUncompressed FF bit
([#2669](#2669))
([1f459ee](1f459ee))
* **orchestrator:** always include execution metrics in sandbox webhook
events ([#2852](#2852))
([440edfe](440edfe))
* **orchestrator:** classify envd-init by exit type
([#3139](#3139))
([1e39a4f](1e39a4f))
* **orchestrator:** graceful sandbox drain on shutdown
([#3069](#3069))
([6ce68e3](6ce68e3))
* **orchestrator:** graceful template-build drain on shutdown
([#3079](#3079))
([1b3001c](1b3001c))
* **orchestrator:** improved read-path telemetry
([#3063](#3063))
([bc3fe84](bc3fe84))
* **orchestrator:** LD-gated ClickHouse write fan-out feature flag
([#3152](#3152))
([f046fcf](f046fcf))
* **orchestrator:** make build-reserved-disk-space-mb default 256MB
([#3065](#3065))
([d473f98](d473f98))
* **orchestrator:** record upload compression metrics
([#2761](#2761))
([9092e35](9092e35))
* **orchestrator:** report hugepage metrics to API
([#3182](#3182))
([7735bae](7735bae))
* **orchestrator:** run startup reclaim on boot
([#3123](#3123))
([79b838e](79b838e))
* **orchestrator:** single-instance flock on startup
([#3143](#3143))
([1320d6e](1320d6e))
* **orchestrator:** soft-delete consumer enforcement for storage index
([#3034](#3034))
([fbfc918](fbfc918))
* **orchestrator:** tag envd-init meters with start_type
([#3125](#3125))
([4466b48](4466b48))
* **orchestrator:** track and report last status change timestamp
([#2980](#2980))
([f79be77](f79be77))
* **orchestrator:** track sandbox lifecycles
([#2998](#2998))
([057f20c](057f20c))
* **orchestrator:** write layer sizes (logical/mapped/diff) to object
metadata ([#3122](#3122))
([11869c0](11869c0))
* **orch:** harvest resume-prefetch trace on pause
([#3067](#3067))
([97bd4a5](97bd4a5))
* **orch:** last-cycle memory prefetch on resume
([#3258](#3258))
([b22e820](b22e820))
* **orch:** make resume-build -gdb work on real nodes + add copy-build
-gdb ([#3108](#3108))
([c684bd2](c684bd2))
* **orch:** opt-in DSCP marker for sandbox egress (SANDBOX_EGRESS_DSCP)
([#3039](#3039))
([a98cf2c](a98cf2c))
* **orch:** per-start UFFD startup working-set metric
([#2960](#2960))
([dc386b2](dc386b2))
* **orch:** premade NixOS base-image support
([#3412](#3412))
([776ba39](776ba39))
* **orch:** record envd init duration histogram on failure with success
attribute ([#2749](#2749))
([afa7458](afa7458))
* **orch:** snapshot fragmentation metrics
([#2931](#2931))
([842b007](842b007))
* per-team events TTL limit (tier + addons)
([#3181](#3181))
([f76b2cb](f76b2cb))
* **shared:** add OTEL instrumentation to AWS S3 storage client
([#3172](#3172))
([25b0fd1](25b0fd1))
* **storage:** per-role storage URLs, env-free storage library
([#3246](#3246))
([fcbe909](fcbe909))
* **storage:** stamp provenance custom metadata on uploaded objects
(incl. headers) ([#3033](#3033))
([ba8604e](ba8604e))
* **storage:** write-through compressed templates to NFS on upload
([#2827](#2827))
([57503c1](57503c1))


### Bug Fixes

* added api and orch
([#3454](#3454))
([fda5e45](fda5e45))
* **block:** rephrase misleading error message in pwritevAll
([#2816](#2816))
([1555f1b](1555f1b))
* **cache:** use 512-byte units for stat.Blocks in FileSize
([#2949](#2949))
([0f632a9](0f632a9))
* **clean-nfs-cache:** exclude zombies from delete_age
([#3191](#3191))
([3fa2aeb](3fa2aeb))
* **compression:** correctness findings from compression audit
([#2803](#2803))
([d21a6a9](d21a6a9))
* **copy-build:** resolve compression suffix for build data files
([#2859](#2859))
([8966f7e](8966f7e))
* correct 3 CVES ([#3218](#3218))
([076823b](076823b))
* **envd:** stop freezing socat cgroup across pause/resume
([#2923](#2923))
([8b6f2b9](8b6f2b9))
* **inspect-build:** adapt validate to new Chunker upstream API
([#2989](#2989))
([2e0d3da](2e0d3da))
* **nbd:** adjust status poll sleep from 100ns to 100µs
([02bf51b](02bf51b))
* **nbd:** change NBD status poll sleep from 100ns to 100µs to avoid
useless busy spinning
([#2884](#2884))
([02bf51b](02bf51b))
* **nfsproxy:** deflake TestRoundTrip EADDRINUSE
([#2987](#2987))
([55f4d18](55f4d18))
* **orch:** denormalize upload metric file type
([#2865](#2865))
([b1646ca](b1646ca))
* **orch:** disable the chronyd seccomp filter on Alpine when using PHC
([#3453](#3453))
([e58af28](e58af28))
* **orchestrator:** anchor rsync CWD to root in template file copy
([#2835](#2835))
([7160db9](7160db9))
* **orchestrator:** atomically replace metadata
([#3321](#3321))
([0c4ad6b](0c4ad6b))
* **orchestrator:** avoid serializing upload headers twice
([#2762](#2762))
([9b7b149](9b7b149))
* **orchestrator:** chunk readiness bug in P2P-&gt;compressed
([#3185](#3185))
([74a6e5b](74a6e5b))
* **orchestrator:** deschedule flaky eviction-loop race in TestDiffSto…
([#3173](#3173))
([88ff17c](88ff17c))
* **orchestrator:** discard poisoned nftables conn on firewall errors
([#3008](#3008))
([03f10e0](03f10e0))
* **orchestrator:** drop stale pre-init logs
([#3297](#3297))
([8ec4be5](8ec4be5))
* **orchestrator:** emit compression ratios as fractions, not BP
([#2772](#2772))
([866f4c1](866f4c1))
* **orchestrator:** export dirty-page stall counter from process start
([#2992](#2992))
([badc8ad](badc8ad))
* **orchestrator:** harden Firecracker process shutdown
([#2996](#2996))
([df662e7](df662e7))
* **orchestrator:** harden shutdown network cleanup
([#3000](#3000))
([de2f391](de2f391))
* **orchestrator:** implement Docker COPY merge semantics in template
builds ([#3283](#3283))
([9174104](9174104))
* **orchestrator:** keep dedup empty-pages telemetry scan-only
([#2991](#2991))
([35d0832](35d0832))
* **orchestrator:** let build-cache threshold flag raise above its fal…
([#3175](#3175))
([06393c3](06393c3))
* **orchestrator:** log missing egress proxy in startup reclaim instead
of defaulting silently
([#3116](#3116))
([6ca3163](6ca3163))
* **orchestrator:** make copy-build handle filesystem-only snapshots
([#3299](#3299))
([62add04](62add04))
* **orchestrator:** measure ext4 free space from block groups
([#3282](#3282))
([f18f05f](f18f05f))
* **orchestrator:** normalize upload metric file labels
([#2767](#2767))
([6dec8b3](6dec8b3))
* **orchestrator:** order egress config/firewall updates to close BYOP
enable race ([#3313](#3313))
([7faa59e](7faa59e))
* **orchestrator:** order envd.service after local-fs.target
([#3043](#3043))
([ea2663e](ea2663e))
* **orchestrator:** order envd.service after systemd-tmpfiles-setup
([#3130](#3130))
([9481811](9481811))
* **orchestrator:** pause upload retain retry
([#2993](#2993))
([4f81799](4f81799))
* **orchestrator:** pin tap device host-side MAC address
([#3271](#3271))
([3c786ba](3c786ba))
* **orchestrator:** pin UFFD copy source buffers
([#2745](#2745))
([837fa91](837fa91))
* **orchestrator:** preserve full ENV value across stdout chunks
([#2740](#2740))
([4822e6d](4822e6d))
* **orchestrator:** read V3 ancestors as uncompressed instead of failing
([#2994](#2994))
([c479dd3](c479dd3))
* **orchestrator:** reject standby while draining
([#3325](#3325))
([475a7ee](475a7ee))
* **orchestrator:** report real V4 header compression ratio
([#2771](#2771))
([ecd344e](ecd344e))
* **orchestrator:** resolve remaining P2P/compression/V5 issues
([#3015](#3015))
([1e4379e](1e4379e))
* **orchestrator:** sanitize OCI pull errors
([#3096](#3096))
([a3af6c0](a3af6c0))
* **orchestrator:** scope rootfs hash to provision default
([#3129](#3129))
([475f955](475f955))
* **orchestrator:** stop Checks health-loop leaking
([#2739](#2739))
([17e6e60](17e6e60))
* **orchestrator:** survive SIGBUS from failing disks under mmap'd
caches ([#3385](#3385))
([728bba3](728bba3))
* **orchestrator:** tolerate missing header for legacy templates
([#3026](#3026))
([8a44bfe](8a44bfe))
* **orch:** fall back to ID_LIKE with a warning instead of rejecting
([#3459](#3459))
([7167818](7167818))
* **orch:** prevent NBD dispatch read-loop stall on WRITE_ZEROES (behind
flag) ([#3048](#3048))
([efd3d4d](efd3d4d))
* **orch:** split scheduling base build id per artifact
([#2920](#2920))
([3e35a2a](3e35a2a))
* **orch:** validate copy-build -gdb buckets before the snapshot copy
([#3446](#3446))
([586ad74](586ad74))
* **shared:** never report a failed envd command stream as success
([#3281](#3281))
([69c06b6](69c06b6))
* **storage:** compression upload & cache correctness fixes
([#3231](#3231))
([980748f](980748f))
* **storage:** don't assume V4+ ancestor gaps are uncompressed
([#3447](#3447))
([bfdbb24](bfdbb24))
* **uffd:** dedupe deferred page faults
([#2864](#2864))
([9680a41](9680a41))
* WrapContextAsUserError should not misclassify internal timeouts as
user cancellations
([#3155](#3155))
([8f83959](8f83959))


### Performance Improvements

* **build:** cache resolved Diff per BuildId within File.ReadAt
([#2838](#2838))
([53de07f](53de07f))
* **build:** parallelize fragmented backing reads
([#2872](#2872))
([c7655a7](c7655a7))
* **clean-nfs-cache:** restore dirfd-relative statx
([#2766](#2766))
([6bdbedb](6bdbedb))
* **header:** add V5 columnar varint header format
([#2847](#2847))
([9dd931b](9dd931b))
* **header:** pack cached Header.Mapping into a compact form
([#2844](#2844))
([7f0b13c](7f0b13c))
* **orchestrator:** add memfile dedup density threshold
([#2862](#2862))
([7ccfa02](7ccfa02))
* **orchestrator:** avoid V3-ancestor header refresh
([#2999](#2999))
([cb6aa0b](cb6aa0b))
* **orch:** metrics for dirty page throttling
([#2858](#2858))
([d2aa554](d2aa554))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com>
Co-authored-by: Charlie Wyse <charlie.wyse@e2b.dev>
jakubno added a commit that referenced this pull request Aug 3, 2026
…hes (#3385)

A single bad 4K sector on a node's local NVMe crashed the whole
orchestrator and killed every sandbox on the node: the kernel cannot
return EIO for a memory access, so a failed page-in from an mmap'd cache
file delivers SIGBUS, which Go turns into an unrecoverable fatal error
inside the Cache.ReadAt memmove.

RunFaultSafe converts such faults into a regular error carrying the
fault address (debug.SetPanicOnFault + recover); other panics still
crash. It guards the in-process mmap copies: build segment reads, NBD
dispatch read/write/write-zeroes and the dedup compare. A bad block now
fails a single read or NBD request instead of the whole orchestrator,
and the affected file is just a local cache of remote storage anyway.
The guard is cheap enough to sit on the hot read path.

Recovered faults are logged and counted in
orchestrator.block.memory_fault, which should stay at zero on a healthy
node.

---------

Co-authored-by: Jakub Rojko <jakub@e2b.dev>
jakubno pushed a commit that referenced this pull request Aug 3, 2026
🤖 I have created a release *beep* *boop*
---


## 0.0.1 (2026-07-30)


### Features

* **api:** add sandbox IAM workload token configuration
([13ddb3d](13ddb3d))
* **api:** add sandbox workload identity permission
([#3319](#3319))
([13ddb3d](13ddb3d))
* **api:** SOCKS5 egress proxy on sandbox network config (BYOP)
([#2642](#2642))
([1fc3820](1fc3820))
* **cfg:** add DISABLE_STARTUP_RECLAIM flag
([#3081](#3081))
([7677ca6](7677ca6))
* **clickhouse:** implement multi-cluster fan-out for events and stats
([#2925](#2925))
([39594c6](39594c6))
* dynamic sandbox log routing and ClickHouse-backed log reads
([#3236](#3236))
([1b19a3b](1b19a3b))
* **envd:** give envd realtime IO priority, reset for user processes
([#2681](#2681))
([f4bd1b2](f4bd1b2))
* **envd:** split collapse stats into real migrations vs already-huge
([#3021](#3021))
([0d77614](0d77614))
* **envd:** support user-defined file metadata via xattrs
([#2732](#2732))
([da8fbe4](da8fbe4))
* **featureflags:** support per-service context providers
([#3100](#3100))
([65297c1](65297c1))
* freeze user cgroup across pause/resume to keep envd /init responsive
([#2688](#2688))
([eceb741](eceb741))
* **metrics:** break down pause-snapshot latency by step
([#3426](#3426))
([f551118](f551118))
* **metrics:** label pause telemetry by fs_only
([#3425](#3425))
([4be33ba](4be33ba))
* **observability:** add kill_reason to sandbox.lifecycle.killed
([#2833](#2833))
([e45418f](e45418f))
* **observability:** include kill_reason in kill-path structured logs
([#2846](#2846))
([33c49f7](33c49f7))
* **orch:** add envd-version to LaunchDarkly sandbox context
([#3051](#3051))
([37d3b92](37d3b92))
* **orch:** add less, nftables, iputils-ping, and jq to base
provisioning ([#2736](#2736))
([a1e010e](a1e010e))
* **orch:** collapse envd's heap into 2 MiB hugepages before pause to
cut cold-resume faults
([#2997](#2997))
([6677f73](6677f73))
* **orch:** debug a sandbox guest kernel with resume-build -gdb
([#3040](#3040))
([37bb0dc](37bb0dc))
* **orch:** decouple warm resume from memfile dedup
([#3166](#3166))
([77f25a0](77f25a0))
* **orch:** distro-aware template base-image provisioning
([#3411](#3411))
([1abece1](1abece1))
* **orchestrator/cgroup:** list and destroy leaked sandbox cgroups
([#3086](#3086))
([bce1d84](bce1d84))
* **orchestrator/nbd:** inspect and disconnect connected devices
([#3087](#3087))
([4d47148](4d47148))
* **orchestrator/network:** list slot namespaces
([#3089](#3089))
([c23dbc7](c23dbc7))
* **orchestrator/network:** list slot namespaces
([#3090](#3090))
([fbfce25](fbfce25))
* **orchestrator:** add -force-reboot to resume-build to cold-boot
memory-snaphsot builds
([#3208](#3208))
([cf8f15b](cf8f15b))
* **orchestrator:** add allocated resource metrics for sandboxes
([#2943](#2943))
([95cb6d3](95cb6d3))
* **orchestrator:** add dummy orchestrator binary for local API dev
([#2744](#2744))
([ab56e25](ab56e25))
* **orchestrator:** add NetworkAssignHook for sandbox lifecycle
extensions ([#3290](#3290))
([3261963](3261963))
* **orchestrator:** add soft-delete marker label to the check metric
([#3144](#3144))
([1ce64f8](1ce64f8))
* **orchestrator:** add v4HeaderForUncompressed FF bit
([#2669](#2669))
([1f459ee](1f459ee))
* **orchestrator:** always include execution metrics in sandbox webhook
events ([#2852](#2852))
([440edfe](440edfe))
* **orchestrator:** classify envd-init by exit type
([#3139](#3139))
([1e39a4f](1e39a4f))
* **orchestrator:** graceful sandbox drain on shutdown
([#3069](#3069))
([6ce68e3](6ce68e3))
* **orchestrator:** graceful template-build drain on shutdown
([#3079](#3079))
([1b3001c](1b3001c))
* **orchestrator:** improved read-path telemetry
([#3063](#3063))
([bc3fe84](bc3fe84))
* **orchestrator:** LD-gated ClickHouse write fan-out feature flag
([#3152](#3152))
([f046fcf](f046fcf))
* **orchestrator:** make build-reserved-disk-space-mb default 256MB
([#3065](#3065))
([d473f98](d473f98))
* **orchestrator:** record upload compression metrics
([#2761](#2761))
([9092e35](9092e35))
* **orchestrator:** report hugepage metrics to API
([#3182](#3182))
([7735bae](7735bae))
* **orchestrator:** run startup reclaim on boot
([#3123](#3123))
([79b838e](79b838e))
* **orchestrator:** single-instance flock on startup
([#3143](#3143))
([1320d6e](1320d6e))
* **orchestrator:** soft-delete consumer enforcement for storage index
([#3034](#3034))
([fbfc918](fbfc918))
* **orchestrator:** tag envd-init meters with start_type
([#3125](#3125))
([4466b48](4466b48))
* **orchestrator:** track and report last status change timestamp
([#2980](#2980))
([f79be77](f79be77))
* **orchestrator:** track sandbox lifecycles
([#2998](#2998))
([057f20c](057f20c))
* **orchestrator:** write layer sizes (logical/mapped/diff) to object
metadata ([#3122](#3122))
([11869c0](11869c0))
* **orch:** harvest resume-prefetch trace on pause
([#3067](#3067))
([97bd4a5](97bd4a5))
* **orch:** last-cycle memory prefetch on resume
([#3258](#3258))
([ea94196](ea94196))
* **orch:** make resume-build -gdb work on real nodes + add copy-build
-gdb ([#3108](#3108))
([5385594](5385594))
* **orch:** opt-in DSCP marker for sandbox egress (SANDBOX_EGRESS_DSCP)
([#3039](#3039))
([a98cf2c](a98cf2c))
* **orch:** per-start UFFD startup working-set metric
([#2960](#2960))
([dc386b2](dc386b2))
* **orch:** premade NixOS base-image support
([#3412](#3412))
([4bd42d2](4bd42d2))
* **orch:** record envd init duration histogram on failure with success
attribute ([#2749](#2749))
([afa7458](afa7458))
* **orch:** snapshot fragmentation metrics
([#2931](#2931))
([842b007](842b007))
* per-team events TTL limit (tier + addons)
([#3181](#3181))
([f76b2cb](f76b2cb))
* **shared:** add OTEL instrumentation to AWS S3 storage client
([#3172](#3172))
([25b0fd1](25b0fd1))
* **storage:** per-role storage URLs, env-free storage library
([#3246](#3246))
([fcbe909](fcbe909))
* **storage:** stamp provenance custom metadata on uploaded objects
(incl. headers) ([#3033](#3033))
([ba8604e](ba8604e))
* **storage:** write-through compressed templates to NFS on upload
([#2827](#2827))
([57503c1](57503c1))


### Bug Fixes

* added api and orch
([#3454](#3454))
([d56e0a8](d56e0a8))
* **block:** rephrase misleading error message in pwritevAll
([#2816](#2816))
([1555f1b](1555f1b))
* **cache:** use 512-byte units for stat.Blocks in FileSize
([#2949](#2949))
([0f632a9](0f632a9))
* **clean-nfs-cache:** exclude zombies from delete_age
([#3191](#3191))
([3fa2aeb](3fa2aeb))
* **compression:** correctness findings from compression audit
([#2803](#2803))
([d21a6a9](d21a6a9))
* **copy-build:** resolve compression suffix for build data files
([#2859](#2859))
([8966f7e](8966f7e))
* correct 3 CVES ([#3218](#3218))
([076823b](076823b))
* **envd:** stop freezing socat cgroup across pause/resume
([#2923](#2923))
([8b6f2b9](8b6f2b9))
* **inspect-build:** adapt validate to new Chunker upstream API
([#2989](#2989))
([2e0d3da](2e0d3da))
* **nbd:** adjust status poll sleep from 100ns to 100µs
([02bf51b](02bf51b))
* **nbd:** change NBD status poll sleep from 100ns to 100µs to avoid
useless busy spinning
([#2884](#2884))
([02bf51b](02bf51b))
* **nfsproxy:** deflake TestRoundTrip EADDRINUSE
([#2987](#2987))
([55f4d18](55f4d18))
* **orch:** denormalize upload metric file type
([#2865](#2865))
([b1646ca](b1646ca))
* **orch:** disable the chronyd seccomp filter on Alpine when using PHC
([#3453](#3453))
([dfa9764](dfa9764))
* **orchestrator:** anchor rsync CWD to root in template file copy
([#2835](#2835))
([7160db9](7160db9))
* **orchestrator:** atomically replace metadata
([#3321](#3321))
([0c4ad6b](0c4ad6b))
* **orchestrator:** avoid serializing upload headers twice
([#2762](#2762))
([9b7b149](9b7b149))
* **orchestrator:** chunk readiness bug in P2P-&gt;compressed
([#3185](#3185))
([74a6e5b](74a6e5b))
* **orchestrator:** deschedule flaky eviction-loop race in TestDiffSto…
([#3173](#3173))
([88ff17c](88ff17c))
* **orchestrator:** discard poisoned nftables conn on firewall errors
([#3008](#3008))
([03f10e0](03f10e0))
* **orchestrator:** drop stale pre-init logs
([#3297](#3297))
([8ec4be5](8ec4be5))
* **orchestrator:** emit compression ratios as fractions, not BP
([#2772](#2772))
([866f4c1](866f4c1))
* **orchestrator:** export dirty-page stall counter from process start
([#2992](#2992))
([badc8ad](badc8ad))
* **orchestrator:** harden Firecracker process shutdown
([#2996](#2996))
([df662e7](df662e7))
* **orchestrator:** harden shutdown network cleanup
([#3000](#3000))
([de2f391](de2f391))
* **orchestrator:** implement Docker COPY merge semantics in template
builds ([#3283](#3283))
([9174104](9174104))
* **orchestrator:** keep dedup empty-pages telemetry scan-only
([#2991](#2991))
([35d0832](35d0832))
* **orchestrator:** let build-cache threshold flag raise above its fal…
([#3175](#3175))
([06393c3](06393c3))
* **orchestrator:** log missing egress proxy in startup reclaim instead
of defaulting silently
([#3116](#3116))
([6ca3163](6ca3163))
* **orchestrator:** make copy-build handle filesystem-only snapshots
([#3299](#3299))
([62add04](62add04))
* **orchestrator:** measure ext4 free space from block groups
([#3282](#3282))
([f18f05f](f18f05f))
* **orchestrator:** normalize upload metric file labels
([#2767](#2767))
([6dec8b3](6dec8b3))
* **orchestrator:** order egress config/firewall updates to close BYOP
enable race ([#3313](#3313))
([7faa59e](7faa59e))
* **orchestrator:** order envd.service after local-fs.target
([#3043](#3043))
([ea2663e](ea2663e))
* **orchestrator:** order envd.service after systemd-tmpfiles-setup
([#3130](#3130))
([9481811](9481811))
* **orchestrator:** pause upload retain retry
([#2993](#2993))
([4f81799](4f81799))
* **orchestrator:** pin tap device host-side MAC address
([#3271](#3271))
([3c786ba](3c786ba))
* **orchestrator:** pin UFFD copy source buffers
([#2745](#2745))
([837fa91](837fa91))
* **orchestrator:** preserve full ENV value across stdout chunks
([#2740](#2740))
([4822e6d](4822e6d))
* **orchestrator:** read V3 ancestors as uncompressed instead of failing
([#2994](#2994))
([c479dd3](c479dd3))
* **orchestrator:** reject standby while draining
([#3325](#3325))
([475a7ee](475a7ee))
* **orchestrator:** report real V4 header compression ratio
([#2771](#2771))
([ecd344e](ecd344e))
* **orchestrator:** resolve remaining P2P/compression/V5 issues
([#3015](#3015))
([1e4379e](1e4379e))
* **orchestrator:** sanitize OCI pull errors
([#3096](#3096))
([a3af6c0](a3af6c0))
* **orchestrator:** scope rootfs hash to provision default
([#3129](#3129))
([475f955](475f955))
* **orchestrator:** stop Checks health-loop leaking
([#2739](#2739))
([17e6e60](17e6e60))
* **orchestrator:** survive SIGBUS from failing disks under mmap'd
caches ([#3385](#3385))
([8694d08](8694d08))
* **orchestrator:** tolerate missing header for legacy templates
([#3026](#3026))
([8a44bfe](8a44bfe))
* **orch:** fall back to ID_LIKE with a warning instead of rejecting
([#3459](#3459))
([73399b3](73399b3))
* **orch:** prevent NBD dispatch read-loop stall on WRITE_ZEROES (behind
flag) ([#3048](#3048))
([efd3d4d](efd3d4d))
* **orch:** split scheduling base build id per artifact
([#2920](#2920))
([3e35a2a](3e35a2a))
* **orch:** validate copy-build -gdb buckets before the snapshot copy
([#3446](#3446))
([9be382f](9be382f))
* **shared:** never report a failed envd command stream as success
([#3281](#3281))
([69c06b6](69c06b6))
* **storage:** compression upload & cache correctness fixes
([#3231](#3231))
([980748f](980748f))
* **storage:** don't assume V4+ ancestor gaps are uncompressed
([#3447](#3447))
([f828d12](f828d12))
* **uffd:** dedupe deferred page faults
([#2864](#2864))
([9680a41](9680a41))
* WrapContextAsUserError should not misclassify internal timeouts as
user cancellations
([#3155](#3155))
([8f83959](8f83959))


### Performance Improvements

* **build:** cache resolved Diff per BuildId within File.ReadAt
([#2838](#2838))
([53de07f](53de07f))
* **build:** parallelize fragmented backing reads
([#2872](#2872))
([c7655a7](c7655a7))
* **clean-nfs-cache:** restore dirfd-relative statx
([#2766](#2766))
([6bdbedb](6bdbedb))
* **header:** add V5 columnar varint header format
([#2847](#2847))
([9dd931b](9dd931b))
* **header:** pack cached Header.Mapping into a compact form
([#2844](#2844))
([7f0b13c](7f0b13c))
* **orchestrator:** add memfile dedup density threshold
([#2862](#2862))
([7ccfa02](7ccfa02))
* **orchestrator:** avoid V3-ancestor header refresh
([#2999](#2999))
([cb6aa0b](cb6aa0b))
* **orch:** metrics for dirty page throttling
([#2858](#2858))
([d2aa554](d2aa554))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com>
Co-authored-by: Charlie Wyse <charlie.wyse@e2b.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants