fix(orchestrator): survive SIGBUS from failing disks under mmap'd caches - #3385
Conversation
…errors An unrecoverable disk read error (bad sector) under a block cache file raises SIGBUS when the kernel pages in the mapped range. The memmove in Cache.ReadAt cannot return an error, so the Go runtime turns the fault into 'fatal error: unexpected fault address' and the whole orchestrator dies, killing every sandbox on the node. Add subprocess-based tests (truncated backing file = deterministic SIGBUS) demonstrating the crash in the three affected paths: - RunFaultSafe helper (skeleton, pass-through for now) - build.readSegments (parallel + sequential) - NBD dispatch read serving Overlay/Cache mmap The fault conversion lands in the next commit, turning these red tests green.
Implement block.RunFaultSafe using debug.SetPanicOnFault + recover: a memory fault raised on the calling goroutine (SIGBUS paging in an mmap'd cache file backed by an unreadable disk block) becomes an error wrapping block.ErrMemoryFault instead of a fatal, unrecoverable runtime throw. Non-fault panics re-panic, so real bugs stay loud. Faults are detected by the Addr method only fault-derived runtime.Errors carry. Guard the in-process mmap copy paths: - build.readSegments (parallel + sequential): covers sandbox start prefetch, NBD template reads and uffd reads through File.ReadAt - NBD dispatch cmdRead/cmdWrite/cmdWriteZeroes: covers Overlay COW cache access; the guest gets an NBD error reply (EIO) per request - Cache.Dedup compare pass Syscall-based consumers (pwritev, UFFDIO_COPY, socket writes) already get EFAULT as an ordinary error and need no guard. A fault now fails one read/request instead of killing the orchestrator and every sandbox on the node, and is logged with the cache path as a disk-health signal. In the motivating incident a single bad 4K sector under one template's cache file crashed the whole node 20 minutes after the disk started failing.
Intel Core Ultra 7 268V, linux/amd64, -benchtime=2s: size=4KiB/unguarded 68.02 ns/op 60215 MB/s size=4KiB/guarded 85.34 ns/op 47995 MB/s size=256KiB/unguarded 8158 ns/op 32134 MB/s size=256KiB/guarded 7872 ns/op 33299 MB/s ~17 ns per call (SetPanicOnFault bool swap + two defers + recover), independent of size: noticeable only relative to the smallest possible read and lost in the noise at typical segment sizes.
Add orchestrator.block.memory_fault, incremented inside RunFaultSafe so every guarded site (build segment reads, NBD dispatch, dedup) is covered centrally. A non-zero rate on a node means its local disk is likely failing and the node should be drained before more blocks rot. The subprocess fault test asserts the counter increments on a recovered fault.
Thread the request context through RunFaultSafe so the memory-fault counter can attach trace exemplars, linking a fault datapoint to the exact read that hit it. The context does not cancel fn, which runs synchronously; call sites keep their own contextual error logs.
The runtime discards the signal details when SetPanicOnFault converts a fault: the panic value's message is the generic 'runtime error: invalid memory address or nil pointer dereference' for SIGBUS and SIGSEGV alike, and the faulting address is only exposed via its Addr method. Return *MemoryFaultError carrying Addr instead of a plain wrap, stdlib style (fs.ErrNotExist / *fs.PathError): errors.Is(err, ErrMemoryFault) still matches via the Is method, errors.As extracts the address, and Unwrap keeps the runtime error in the chain. The address can be correlated with the mmap layout to find the file offset — and hence the disk block — that failed; the recovery-point logs now carry it as a structured fault_addr field.
Keep the SIGBUS/bad-sector rationale once, on ErrMemoryFault and RunFaultSafe; call sites and tests reference it instead of retelling it.
Names and assertions carry the rest; the disk-failure rationale lives on ErrMemoryFault and the metric description.
Production code matches *MemoryFaultError with errors.As; the sentinel and its Is method had no consumers outside tests. One error type is enough, matching the CacheClosedError convention.
RunFaultSafe converts any memory fault, not just those from mapped files; the type doc keeps the typical cause.
The counter covers any fault RunFaultSafe recovers; the typical cause belongs to the type doc, not the metric.
This reverts commit 5c00af3355d7ce89da73ceb999b3545dbaf86879.
Remove churn left by earlier iterations: the error-message formatting assertion (the typed Addr field already covers it), the clean-return half of the flag-restore test (only the defer ordering after a panic is a real contract), and the generic metric-iteration helper (inlined at its single call site). Add the missing cmdWrite wiring coverage to the dispatch fault child: a non-zero-payload write into the truncated mmap faults on the store (all-zero writes take the punch-hole path and never touch the mapping) and must come back as an NBD error reply. Verified red with the cmdWrite guard removed. cmdWriteZeroes stays untested deliberately: whether it faults depends on the MADV_REMOVE vs clear() fallback, so the assertion would be flaky by design; the Dedup wiring remains a known gap.
The subprocess re-exec pattern guarded against an unconverted fault killing the test binary, but that failure only happens when the guard itself is broken in development, is unmissable in CI, and the crash dump points straight at the unguarded copy. Meanwhile the pattern carried a real trap: the -test.run filter must track the test name, and a rename would leave a child that runs zero tests and passes forever. Fold the child bodies into the tests and note in each that a binary crash with 'unexpected fault address' is the expected failure mode.
gochecknoglobals is disabled repo-wide, so its nolint silenced nothing. The paralleltest suppression on the flag-restore test was justified by a false claim: SetPanicOnFault is per-goroutine, not process-wide, so the test can simply run in parallel. Only the counter-swap suppression does real work and stays.
The swap helper was designed for the subprocess era, where it ran in an isolated child process; after the in-process conversion it forced the test to be non-parallel and carried the branch's last nolint. Testing a one-line counter Add is not worth a global swap; the counter itself stays as the operational signal.
PR SummaryMedium Risk Overview Reviewed by Cursor Bugbot for commit 79be12f. Bugbot is set up for automated code reviews on this repo. Configure here. |
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
There was a problem hiding this comment.
Beyond the inline finding, I checked one other fault-safety edge case and ruled it out: in nbd/dispatch.go's backing Cache.WriteAtWithoutLock, the dirty-tracker update (c.tracker.SetRange(..., Dirty)) runs only after the copy() into the mmap completes, so a fault during that copy propagates as an error before the block is ever marked Dirty — it can't be silently recorded as valid after a partial/faulted write.
Extended reasoning...
Verified directly in packages/orchestrator/pkg/sandbox/block/cache.go: the flush closure in WriteAtWithoutLock calls copy((*c.mmap)[runStart:runEnd], b[...]) before c.tracker.SetRange(startIdx, endIdx, Dirty). Since RunFaultSafe wraps the whole WriteAt call in the NBD dispatch path, a SIGBUS during that copy() unwinds via panic/recover before the subsequent SetRange line ever executes, so the block is never marked Dirty when the write only partially landed. This rules out the candidate concern that a fault-recovered write could leave an already-Dirty block silently treated as valid data.
| return nil | ||
| } | ||
|
|
||
| // readSegmentFaultSafe converts memory faults from the diff's mmap'd cache | ||
| // into errors, so a bad disk block fails one read, not the process. | ||
| func (b *File) readSegmentFaultSafe(ctx context.Context, p []byte, s readSegment) error { | ||
| err := block.RunFaultSafe(ctx, func() error { return b.readSegment(ctx, p, s) }) | ||
| var faultErr *block.MemoryFaultError | ||
| if errors.As(err, &faultErr) { | ||
| cachePath, _ := s.diff.CachePath(ctx) | ||
| logger.L().Error(ctx, "memory fault reading build segment; local disk under the cache is likely failing", | ||
| zap.Error(err), | ||
| zap.String("cache_path", cachePath), | ||
| zap.Uintptr("fault_addr", faultErr.Addr), | ||
| zap.Int64("offset", s.srcOff), | ||
| zap.Int64("length", s.length), | ||
| ) | ||
| } | ||
|
|
There was a problem hiding this comment.
🟣 The P2P build-serving path (peerserver/seekable.go Stream() → ChunkService.ReadAtBuildSeekable) is not covered by the new RunFaultSafe guard: it returns a zero-copy mmap slice via diff.Slice(...) and hands it straight to sender.Send(...), where gRPC's proto marshaling performs the actual page-in memcpy outside any fault protection. A bad sector under a locally-cached build served to a peer will still raise an unrecovered SIGBUS and crash the whole orchestrator — the exact failure mode this PR fixes elsewhere — but this is pre-existing behavior in code the PR does not touch (peerserver/seekable.go, server/chunks.go), so it's not a regression introduced here.
Extended reasoning...
What the bug is: This PR introduces block.RunFaultSafe (debug.SetPanicOnFault + recover) specifically to convert SIGBUS from bad sectors under mmap'd caches into a normal error instead of a process-killing fatal throw, and applies it to build.File.readSegments, NBD read/write/write-zeroes dispatch, and the dedup compare. However, the P2P build-serving path used to share a locally-built template with peer orchestrators is not wrapped, and it performs the exact same class of raw mmap access.
The code path: Server.ReadAtBuildSeekable (packages/orchestrator/pkg/server/chunks.go) calls src.Stream(ctx, offset, length, &seekableStreamSender{stream}). seekableSource.Stream (packages/orchestrator/pkg/sandbox/template/peerserver/seekable.go) does data, _ := f.diff.Slice(ctx, offset, length, nil) and then sender.Send(data[:take]). diff.Slice resolves through StorageDiff.Slice → Chunker.Slice, whose cached/fetched fast paths return c.cache.Slice(...)/c.cache.sliceDirect(...), both of which are just (*c.mmap)[off:end] — a zero-copy slice header with no byte access at all. The actual page-in happens later, inside seekableStreamSender.Send → stream.Send(&ReadAtBuildSeekableResponse{Data: data}), where gRPC's proto marshaler copies every byte of that mmap-backed slice into the wire buffer, synchronously on the RPC handler goroutine. That goroutine never calls RunFaultSafe/debug.SetPanicOnFault.
Why existing guards don't help: Wrapping Slice itself (the way one might expect, mirroring readSegmentFaultSafe) would not fix this, since Slice never touches the mapped pages — the fault surfaces later, once the zero-copy slice escapes into gRPC's marshaler. The fix has to guard the consumer, e.g. copy the slice into a freshly allocated buffer inside RunFaultSafe before calling sender.Send.
Impact: P2P build serving is reachable whenever a template has been built locally but not yet uploaded to remote storage (the common case for freshly-built or newly-scheduled templates) — server/chunks.go only falls back to remote-storage serving once uploaded. So this is a live, non-exotic path. A bad NVMe sector under the locally-cached memfile/rootfs mmap, hit while serving a peer's ReadAtBuildSeekable request, raises an unrecovered SIGBUS and crashes the whole orchestrator process, killing every sandbox on the node — precisely the blast radius this PR sets out to eliminate.
Step-by-step proof:
- Orchestrator A builds a template; the build's memfile diff is cached locally as a memory-mapped file, not yet uploaded to GCS.
- Orchestrator B needs the same build and calls
ChunkService.ReadAtBuildSeekableagainst A for a range. - A's
Server.ReadAtBuildSeekablehandler resolves the local diff and callsseekableSource.Stream. Streamcallsf.diff.Slice(ctx, offset, length, nil), which returns(*c.mmap)[off:end]— no memory access yet, so even if the backing block is a bad sector, nothing has faulted.Streamcallssender.Send(data[:take])→stream.Send(...); gRPC's codec marshals theDatafield, reading every byte of the slice — this read hits the bad sector's mapped page and the kernel delivers SIGBUS instead of EIO (mmap has no error-return path).- Because this goroutine never set
debug.SetPanicOnFault(true), the SIGBUS becomes an unrecoverable Go fatal runtime throw, not a catchable panic — the whole orchestrator process dies, taking every sandbox on the node down with it.
Why this is pre-existing rather than a PR regression: the PR does not touch peerserver/seekable.go or server/chunks.go, and this path was equally vulnerable before the change — the PR is still a strict improvement over the prior fully-unguarded state on the paths it does touch. It's flagged here because it's a real, in-theme completeness gap directly undercutting the PR's stated goal ("survive SIGBUS from failing disks under mmap'd caches"), and the author may want to close it in a fast follow-up by copying the slice into a fresh buffer inside RunFaultSafe before Send.
|
Good find, though, cannot the same thing happen for the Cache struct that also uses mmap? |
5aad415 to
d71980e
Compare
Co-authored-by: Jakub Rojko <jakub@e2b.dev>
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit edd284e. Configure here.
runDedup called dedupCompare without the fault guard Cache.Dedup has, and it is the default pause path. The compare reads base pages that can come from a disk-backed mmap, so a bad sector there still killed the process. A fault now fails the pause instead.
🤖 I have created a release *beep* *boop* --- ## 0.0.1 (2026-07-30) ### Features * **api:** add sandbox IAM workload token configuration ([13ddb3d](13ddb3d)) * **api:** add sandbox workload identity permission ([#3319](#3319)) ([13ddb3d](13ddb3d)) * **api:** SOCKS5 egress proxy on sandbox network config (BYOP) ([#2642](#2642)) ([1fc3820](1fc3820)) * **cfg:** add DISABLE_STARTUP_RECLAIM flag ([#3081](#3081)) ([7677ca6](7677ca6)) * **clickhouse:** implement multi-cluster fan-out for events and stats ([#2925](#2925)) ([39594c6](39594c6)) * dynamic sandbox log routing and ClickHouse-backed log reads ([#3236](#3236)) ([1b19a3b](1b19a3b)) * **envd:** give envd realtime IO priority, reset for user processes ([#2681](#2681)) ([f4bd1b2](f4bd1b2)) * **envd:** split collapse stats into real migrations vs already-huge ([#3021](#3021)) ([0d77614](0d77614)) * **envd:** support user-defined file metadata via xattrs ([#2732](#2732)) ([da8fbe4](da8fbe4)) * **featureflags:** support per-service context providers ([#3100](#3100)) ([65297c1](65297c1)) * freeze user cgroup across pause/resume to keep envd /init responsive ([#2688](#2688)) ([eceb741](eceb741)) * **metrics:** break down pause-snapshot latency by step ([#3426](#3426)) ([657559e](657559e)) * **metrics:** label pause telemetry by fs_only ([#3425](#3425)) ([411b63e](411b63e)) * **observability:** add kill_reason to sandbox.lifecycle.killed ([#2833](#2833)) ([e45418f](e45418f)) * **observability:** include kill_reason in kill-path structured logs ([#2846](#2846)) ([33c49f7](33c49f7)) * **orch:** add envd-version to LaunchDarkly sandbox context ([#3051](#3051)) ([37d3b92](37d3b92)) * **orch:** add less, nftables, iputils-ping, and jq to base provisioning ([#2736](#2736)) ([a1e010e](a1e010e)) * **orch:** collapse envd's heap into 2 MiB hugepages before pause to cut cold-resume faults ([#2997](#2997)) ([6677f73](6677f73)) * **orch:** debug a sandbox guest kernel with resume-build -gdb ([#3040](#3040)) ([37bb0dc](37bb0dc)) * **orch:** decouple warm resume from memfile dedup ([#3166](#3166)) ([77f25a0](77f25a0)) * **orch:** distro-aware template base-image provisioning ([#3411](#3411)) ([f8c7b5b](f8c7b5b)) * **orchestrator/cgroup:** list and destroy leaked sandbox cgroups ([#3086](#3086)) ([bce1d84](bce1d84)) * **orchestrator/nbd:** inspect and disconnect connected devices ([#3087](#3087)) ([4d47148](4d47148)) * **orchestrator/network:** list slot namespaces ([#3089](#3089)) ([c23dbc7](c23dbc7)) * **orchestrator/network:** list slot namespaces ([#3090](#3090)) ([fbfce25](fbfce25)) * **orchestrator:** add -force-reboot to resume-build to cold-boot memory-snaphsot builds ([#3208](#3208)) ([cf8f15b](cf8f15b)) * **orchestrator:** add allocated resource metrics for sandboxes ([#2943](#2943)) ([95cb6d3](95cb6d3)) * **orchestrator:** add dummy orchestrator binary for local API dev ([#2744](#2744)) ([ab56e25](ab56e25)) * **orchestrator:** add NetworkAssignHook for sandbox lifecycle extensions ([#3290](#3290)) ([3261963](3261963)) * **orchestrator:** add soft-delete marker label to the check metric ([#3144](#3144)) ([1ce64f8](1ce64f8)) * **orchestrator:** add v4HeaderForUncompressed FF bit ([#2669](#2669)) ([1f459ee](1f459ee)) * **orchestrator:** always include execution metrics in sandbox webhook events ([#2852](#2852)) ([440edfe](440edfe)) * **orchestrator:** classify envd-init by exit type ([#3139](#3139)) ([1e39a4f](1e39a4f)) * **orchestrator:** graceful sandbox drain on shutdown ([#3069](#3069)) ([6ce68e3](6ce68e3)) * **orchestrator:** graceful template-build drain on shutdown ([#3079](#3079)) ([1b3001c](1b3001c)) * **orchestrator:** improved read-path telemetry ([#3063](#3063)) ([bc3fe84](bc3fe84)) * **orchestrator:** LD-gated ClickHouse write fan-out feature flag ([#3152](#3152)) ([f046fcf](f046fcf)) * **orchestrator:** make build-reserved-disk-space-mb default 256MB ([#3065](#3065)) ([d473f98](d473f98)) * **orchestrator:** record upload compression metrics ([#2761](#2761)) ([9092e35](9092e35)) * **orchestrator:** report hugepage metrics to API ([#3182](#3182)) ([7735bae](7735bae)) * **orchestrator:** run startup reclaim on boot ([#3123](#3123)) ([79b838e](79b838e)) * **orchestrator:** single-instance flock on startup ([#3143](#3143)) ([1320d6e](1320d6e)) * **orchestrator:** soft-delete consumer enforcement for storage index ([#3034](#3034)) ([fbfc918](fbfc918)) * **orchestrator:** tag envd-init meters with start_type ([#3125](#3125)) ([4466b48](4466b48)) * **orchestrator:** track and report last status change timestamp ([#2980](#2980)) ([f79be77](f79be77)) * **orchestrator:** track sandbox lifecycles ([#2998](#2998)) ([057f20c](057f20c)) * **orchestrator:** write layer sizes (logical/mapped/diff) to object metadata ([#3122](#3122)) ([11869c0](11869c0)) * **orch:** harvest resume-prefetch trace on pause ([#3067](#3067)) ([97bd4a5](97bd4a5)) * **orch:** last-cycle memory prefetch on resume ([#3258](#3258)) ([b22e820](b22e820)) * **orch:** make resume-build -gdb work on real nodes + add copy-build -gdb ([#3108](#3108)) ([c684bd2](c684bd2)) * **orch:** opt-in DSCP marker for sandbox egress (SANDBOX_EGRESS_DSCP) ([#3039](#3039)) ([a98cf2c](a98cf2c)) * **orch:** per-start UFFD startup working-set metric ([#2960](#2960)) ([dc386b2](dc386b2)) * **orch:** premade NixOS base-image support ([#3412](#3412)) ([776ba39](776ba39)) * **orch:** record envd init duration histogram on failure with success attribute ([#2749](#2749)) ([afa7458](afa7458)) * **orch:** snapshot fragmentation metrics ([#2931](#2931)) ([842b007](842b007)) * per-team events TTL limit (tier + addons) ([#3181](#3181)) ([f76b2cb](f76b2cb)) * **shared:** add OTEL instrumentation to AWS S3 storage client ([#3172](#3172)) ([25b0fd1](25b0fd1)) * **storage:** per-role storage URLs, env-free storage library ([#3246](#3246)) ([fcbe909](fcbe909)) * **storage:** stamp provenance custom metadata on uploaded objects (incl. headers) ([#3033](#3033)) ([ba8604e](ba8604e)) * **storage:** write-through compressed templates to NFS on upload ([#2827](#2827)) ([57503c1](57503c1)) ### Bug Fixes * added api and orch ([#3454](#3454)) ([fda5e45](fda5e45)) * **block:** rephrase misleading error message in pwritevAll ([#2816](#2816)) ([1555f1b](1555f1b)) * **cache:** use 512-byte units for stat.Blocks in FileSize ([#2949](#2949)) ([0f632a9](0f632a9)) * **clean-nfs-cache:** exclude zombies from delete_age ([#3191](#3191)) ([3fa2aeb](3fa2aeb)) * **compression:** correctness findings from compression audit ([#2803](#2803)) ([d21a6a9](d21a6a9)) * **copy-build:** resolve compression suffix for build data files ([#2859](#2859)) ([8966f7e](8966f7e)) * correct 3 CVES ([#3218](#3218)) ([076823b](076823b)) * **envd:** stop freezing socat cgroup across pause/resume ([#2923](#2923)) ([8b6f2b9](8b6f2b9)) * **inspect-build:** adapt validate to new Chunker upstream API ([#2989](#2989)) ([2e0d3da](2e0d3da)) * **nbd:** adjust status poll sleep from 100ns to 100µs ([02bf51b](02bf51b)) * **nbd:** change NBD status poll sleep from 100ns to 100µs to avoid useless busy spinning ([#2884](#2884)) ([02bf51b](02bf51b)) * **nfsproxy:** deflake TestRoundTrip EADDRINUSE ([#2987](#2987)) ([55f4d18](55f4d18)) * **orch:** denormalize upload metric file type ([#2865](#2865)) ([b1646ca](b1646ca)) * **orch:** disable the chronyd seccomp filter on Alpine when using PHC ([#3453](#3453)) ([e58af28](e58af28)) * **orchestrator:** anchor rsync CWD to root in template file copy ([#2835](#2835)) ([7160db9](7160db9)) * **orchestrator:** atomically replace metadata ([#3321](#3321)) ([0c4ad6b](0c4ad6b)) * **orchestrator:** avoid serializing upload headers twice ([#2762](#2762)) ([9b7b149](9b7b149)) * **orchestrator:** chunk readiness bug in P2P->compressed ([#3185](#3185)) ([74a6e5b](74a6e5b)) * **orchestrator:** deschedule flaky eviction-loop race in TestDiffSto… ([#3173](#3173)) ([88ff17c](88ff17c)) * **orchestrator:** discard poisoned nftables conn on firewall errors ([#3008](#3008)) ([03f10e0](03f10e0)) * **orchestrator:** drop stale pre-init logs ([#3297](#3297)) ([8ec4be5](8ec4be5)) * **orchestrator:** emit compression ratios as fractions, not BP ([#2772](#2772)) ([866f4c1](866f4c1)) * **orchestrator:** export dirty-page stall counter from process start ([#2992](#2992)) ([badc8ad](badc8ad)) * **orchestrator:** harden Firecracker process shutdown ([#2996](#2996)) ([df662e7](df662e7)) * **orchestrator:** harden shutdown network cleanup ([#3000](#3000)) ([de2f391](de2f391)) * **orchestrator:** implement Docker COPY merge semantics in template builds ([#3283](#3283)) ([9174104](9174104)) * **orchestrator:** keep dedup empty-pages telemetry scan-only ([#2991](#2991)) ([35d0832](35d0832)) * **orchestrator:** let build-cache threshold flag raise above its fal… ([#3175](#3175)) ([06393c3](06393c3)) * **orchestrator:** log missing egress proxy in startup reclaim instead of defaulting silently ([#3116](#3116)) ([6ca3163](6ca3163)) * **orchestrator:** make copy-build handle filesystem-only snapshots ([#3299](#3299)) ([62add04](62add04)) * **orchestrator:** measure ext4 free space from block groups ([#3282](#3282)) ([f18f05f](f18f05f)) * **orchestrator:** normalize upload metric file labels ([#2767](#2767)) ([6dec8b3](6dec8b3)) * **orchestrator:** order egress config/firewall updates to close BYOP enable race ([#3313](#3313)) ([7faa59e](7faa59e)) * **orchestrator:** order envd.service after local-fs.target ([#3043](#3043)) ([ea2663e](ea2663e)) * **orchestrator:** order envd.service after systemd-tmpfiles-setup ([#3130](#3130)) ([9481811](9481811)) * **orchestrator:** pause upload retain retry ([#2993](#2993)) ([4f81799](4f81799)) * **orchestrator:** pin tap device host-side MAC address ([#3271](#3271)) ([3c786ba](3c786ba)) * **orchestrator:** pin UFFD copy source buffers ([#2745](#2745)) ([837fa91](837fa91)) * **orchestrator:** preserve full ENV value across stdout chunks ([#2740](#2740)) ([4822e6d](4822e6d)) * **orchestrator:** read V3 ancestors as uncompressed instead of failing ([#2994](#2994)) ([c479dd3](c479dd3)) * **orchestrator:** reject standby while draining ([#3325](#3325)) ([475a7ee](475a7ee)) * **orchestrator:** report real V4 header compression ratio ([#2771](#2771)) ([ecd344e](ecd344e)) * **orchestrator:** resolve remaining P2P/compression/V5 issues ([#3015](#3015)) ([1e4379e](1e4379e)) * **orchestrator:** sanitize OCI pull errors ([#3096](#3096)) ([a3af6c0](a3af6c0)) * **orchestrator:** scope rootfs hash to provision default ([#3129](#3129)) ([475f955](475f955)) * **orchestrator:** stop Checks health-loop leaking ([#2739](#2739)) ([17e6e60](17e6e60)) * **orchestrator:** survive SIGBUS from failing disks under mmap'd caches ([#3385](#3385)) ([728bba3](728bba3)) * **orchestrator:** tolerate missing header for legacy templates ([#3026](#3026)) ([8a44bfe](8a44bfe)) * **orch:** fall back to ID_LIKE with a warning instead of rejecting ([#3459](#3459)) ([7167818](7167818)) * **orch:** prevent NBD dispatch read-loop stall on WRITE_ZEROES (behind flag) ([#3048](#3048)) ([efd3d4d](efd3d4d)) * **orch:** split scheduling base build id per artifact ([#2920](#2920)) ([3e35a2a](3e35a2a)) * **orch:** validate copy-build -gdb buckets before the snapshot copy ([#3446](#3446)) ([586ad74](586ad74)) * **shared:** never report a failed envd command stream as success ([#3281](#3281)) ([69c06b6](69c06b6)) * **storage:** compression upload & cache correctness fixes ([#3231](#3231)) ([980748f](980748f)) * **storage:** don't assume V4+ ancestor gaps are uncompressed ([#3447](#3447)) ([bfdbb24](bfdbb24)) * **uffd:** dedupe deferred page faults ([#2864](#2864)) ([9680a41](9680a41)) * WrapContextAsUserError should not misclassify internal timeouts as user cancellations ([#3155](#3155)) ([8f83959](8f83959)) ### Performance Improvements * **build:** cache resolved Diff per BuildId within File.ReadAt ([#2838](#2838)) ([53de07f](53de07f)) * **build:** parallelize fragmented backing reads ([#2872](#2872)) ([c7655a7](c7655a7)) * **clean-nfs-cache:** restore dirfd-relative statx ([#2766](#2766)) ([6bdbedb](6bdbedb)) * **header:** add V5 columnar varint header format ([#2847](#2847)) ([9dd931b](9dd931b)) * **header:** pack cached Header.Mapping into a compact form ([#2844](#2844)) ([7f0b13c](7f0b13c)) * **orchestrator:** add memfile dedup density threshold ([#2862](#2862)) ([7ccfa02](7ccfa02)) * **orchestrator:** avoid V3-ancestor header refresh ([#2999](#2999)) ([cb6aa0b](cb6aa0b)) * **orch:** metrics for dirty page throttling ([#2858](#2858)) ([d2aa554](d2aa554)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com> Co-authored-by: Charlie Wyse <charlie.wyse@e2b.dev>
…hes (#3385) A single bad 4K sector on a node's local NVMe crashed the whole orchestrator and killed every sandbox on the node: the kernel cannot return EIO for a memory access, so a failed page-in from an mmap'd cache file delivers SIGBUS, which Go turns into an unrecoverable fatal error inside the Cache.ReadAt memmove. RunFaultSafe converts such faults into a regular error carrying the fault address (debug.SetPanicOnFault + recover); other panics still crash. It guards the in-process mmap copies: build segment reads, NBD dispatch read/write/write-zeroes and the dedup compare. A bad block now fails a single read or NBD request instead of the whole orchestrator, and the affected file is just a local cache of remote storage anyway. The guard is cheap enough to sit on the hot read path. Recovered faults are logged and counted in orchestrator.block.memory_fault, which should stay at zero on a healthy node. --------- Co-authored-by: Jakub Rojko <jakub@e2b.dev>
🤖 I have created a release *beep* *boop* --- ## 0.0.1 (2026-07-30) ### Features * **api:** add sandbox IAM workload token configuration ([13ddb3d](13ddb3d)) * **api:** add sandbox workload identity permission ([#3319](#3319)) ([13ddb3d](13ddb3d)) * **api:** SOCKS5 egress proxy on sandbox network config (BYOP) ([#2642](#2642)) ([1fc3820](1fc3820)) * **cfg:** add DISABLE_STARTUP_RECLAIM flag ([#3081](#3081)) ([7677ca6](7677ca6)) * **clickhouse:** implement multi-cluster fan-out for events and stats ([#2925](#2925)) ([39594c6](39594c6)) * dynamic sandbox log routing and ClickHouse-backed log reads ([#3236](#3236)) ([1b19a3b](1b19a3b)) * **envd:** give envd realtime IO priority, reset for user processes ([#2681](#2681)) ([f4bd1b2](f4bd1b2)) * **envd:** split collapse stats into real migrations vs already-huge ([#3021](#3021)) ([0d77614](0d77614)) * **envd:** support user-defined file metadata via xattrs ([#2732](#2732)) ([da8fbe4](da8fbe4)) * **featureflags:** support per-service context providers ([#3100](#3100)) ([65297c1](65297c1)) * freeze user cgroup across pause/resume to keep envd /init responsive ([#2688](#2688)) ([eceb741](eceb741)) * **metrics:** break down pause-snapshot latency by step ([#3426](#3426)) ([f551118](f551118)) * **metrics:** label pause telemetry by fs_only ([#3425](#3425)) ([4be33ba](4be33ba)) * **observability:** add kill_reason to sandbox.lifecycle.killed ([#2833](#2833)) ([e45418f](e45418f)) * **observability:** include kill_reason in kill-path structured logs ([#2846](#2846)) ([33c49f7](33c49f7)) * **orch:** add envd-version to LaunchDarkly sandbox context ([#3051](#3051)) ([37d3b92](37d3b92)) * **orch:** add less, nftables, iputils-ping, and jq to base provisioning ([#2736](#2736)) ([a1e010e](a1e010e)) * **orch:** collapse envd's heap into 2 MiB hugepages before pause to cut cold-resume faults ([#2997](#2997)) ([6677f73](6677f73)) * **orch:** debug a sandbox guest kernel with resume-build -gdb ([#3040](#3040)) ([37bb0dc](37bb0dc)) * **orch:** decouple warm resume from memfile dedup ([#3166](#3166)) ([77f25a0](77f25a0)) * **orch:** distro-aware template base-image provisioning ([#3411](#3411)) ([1abece1](1abece1)) * **orchestrator/cgroup:** list and destroy leaked sandbox cgroups ([#3086](#3086)) ([bce1d84](bce1d84)) * **orchestrator/nbd:** inspect and disconnect connected devices ([#3087](#3087)) ([4d47148](4d47148)) * **orchestrator/network:** list slot namespaces ([#3089](#3089)) ([c23dbc7](c23dbc7)) * **orchestrator/network:** list slot namespaces ([#3090](#3090)) ([fbfce25](fbfce25)) * **orchestrator:** add -force-reboot to resume-build to cold-boot memory-snaphsot builds ([#3208](#3208)) ([cf8f15b](cf8f15b)) * **orchestrator:** add allocated resource metrics for sandboxes ([#2943](#2943)) ([95cb6d3](95cb6d3)) * **orchestrator:** add dummy orchestrator binary for local API dev ([#2744](#2744)) ([ab56e25](ab56e25)) * **orchestrator:** add NetworkAssignHook for sandbox lifecycle extensions ([#3290](#3290)) ([3261963](3261963)) * **orchestrator:** add soft-delete marker label to the check metric ([#3144](#3144)) ([1ce64f8](1ce64f8)) * **orchestrator:** add v4HeaderForUncompressed FF bit ([#2669](#2669)) ([1f459ee](1f459ee)) * **orchestrator:** always include execution metrics in sandbox webhook events ([#2852](#2852)) ([440edfe](440edfe)) * **orchestrator:** classify envd-init by exit type ([#3139](#3139)) ([1e39a4f](1e39a4f)) * **orchestrator:** graceful sandbox drain on shutdown ([#3069](#3069)) ([6ce68e3](6ce68e3)) * **orchestrator:** graceful template-build drain on shutdown ([#3079](#3079)) ([1b3001c](1b3001c)) * **orchestrator:** improved read-path telemetry ([#3063](#3063)) ([bc3fe84](bc3fe84)) * **orchestrator:** LD-gated ClickHouse write fan-out feature flag ([#3152](#3152)) ([f046fcf](f046fcf)) * **orchestrator:** make build-reserved-disk-space-mb default 256MB ([#3065](#3065)) ([d473f98](d473f98)) * **orchestrator:** record upload compression metrics ([#2761](#2761)) ([9092e35](9092e35)) * **orchestrator:** report hugepage metrics to API ([#3182](#3182)) ([7735bae](7735bae)) * **orchestrator:** run startup reclaim on boot ([#3123](#3123)) ([79b838e](79b838e)) * **orchestrator:** single-instance flock on startup ([#3143](#3143)) ([1320d6e](1320d6e)) * **orchestrator:** soft-delete consumer enforcement for storage index ([#3034](#3034)) ([fbfc918](fbfc918)) * **orchestrator:** tag envd-init meters with start_type ([#3125](#3125)) ([4466b48](4466b48)) * **orchestrator:** track and report last status change timestamp ([#2980](#2980)) ([f79be77](f79be77)) * **orchestrator:** track sandbox lifecycles ([#2998](#2998)) ([057f20c](057f20c)) * **orchestrator:** write layer sizes (logical/mapped/diff) to object metadata ([#3122](#3122)) ([11869c0](11869c0)) * **orch:** harvest resume-prefetch trace on pause ([#3067](#3067)) ([97bd4a5](97bd4a5)) * **orch:** last-cycle memory prefetch on resume ([#3258](#3258)) ([ea94196](ea94196)) * **orch:** make resume-build -gdb work on real nodes + add copy-build -gdb ([#3108](#3108)) ([5385594](5385594)) * **orch:** opt-in DSCP marker for sandbox egress (SANDBOX_EGRESS_DSCP) ([#3039](#3039)) ([a98cf2c](a98cf2c)) * **orch:** per-start UFFD startup working-set metric ([#2960](#2960)) ([dc386b2](dc386b2)) * **orch:** premade NixOS base-image support ([#3412](#3412)) ([4bd42d2](4bd42d2)) * **orch:** record envd init duration histogram on failure with success attribute ([#2749](#2749)) ([afa7458](afa7458)) * **orch:** snapshot fragmentation metrics ([#2931](#2931)) ([842b007](842b007)) * per-team events TTL limit (tier + addons) ([#3181](#3181)) ([f76b2cb](f76b2cb)) * **shared:** add OTEL instrumentation to AWS S3 storage client ([#3172](#3172)) ([25b0fd1](25b0fd1)) * **storage:** per-role storage URLs, env-free storage library ([#3246](#3246)) ([fcbe909](fcbe909)) * **storage:** stamp provenance custom metadata on uploaded objects (incl. headers) ([#3033](#3033)) ([ba8604e](ba8604e)) * **storage:** write-through compressed templates to NFS on upload ([#2827](#2827)) ([57503c1](57503c1)) ### Bug Fixes * added api and orch ([#3454](#3454)) ([d56e0a8](d56e0a8)) * **block:** rephrase misleading error message in pwritevAll ([#2816](#2816)) ([1555f1b](1555f1b)) * **cache:** use 512-byte units for stat.Blocks in FileSize ([#2949](#2949)) ([0f632a9](0f632a9)) * **clean-nfs-cache:** exclude zombies from delete_age ([#3191](#3191)) ([3fa2aeb](3fa2aeb)) * **compression:** correctness findings from compression audit ([#2803](#2803)) ([d21a6a9](d21a6a9)) * **copy-build:** resolve compression suffix for build data files ([#2859](#2859)) ([8966f7e](8966f7e)) * correct 3 CVES ([#3218](#3218)) ([076823b](076823b)) * **envd:** stop freezing socat cgroup across pause/resume ([#2923](#2923)) ([8b6f2b9](8b6f2b9)) * **inspect-build:** adapt validate to new Chunker upstream API ([#2989](#2989)) ([2e0d3da](2e0d3da)) * **nbd:** adjust status poll sleep from 100ns to 100µs ([02bf51b](02bf51b)) * **nbd:** change NBD status poll sleep from 100ns to 100µs to avoid useless busy spinning ([#2884](#2884)) ([02bf51b](02bf51b)) * **nfsproxy:** deflake TestRoundTrip EADDRINUSE ([#2987](#2987)) ([55f4d18](55f4d18)) * **orch:** denormalize upload metric file type ([#2865](#2865)) ([b1646ca](b1646ca)) * **orch:** disable the chronyd seccomp filter on Alpine when using PHC ([#3453](#3453)) ([dfa9764](dfa9764)) * **orchestrator:** anchor rsync CWD to root in template file copy ([#2835](#2835)) ([7160db9](7160db9)) * **orchestrator:** atomically replace metadata ([#3321](#3321)) ([0c4ad6b](0c4ad6b)) * **orchestrator:** avoid serializing upload headers twice ([#2762](#2762)) ([9b7b149](9b7b149)) * **orchestrator:** chunk readiness bug in P2P->compressed ([#3185](#3185)) ([74a6e5b](74a6e5b)) * **orchestrator:** deschedule flaky eviction-loop race in TestDiffSto… ([#3173](#3173)) ([88ff17c](88ff17c)) * **orchestrator:** discard poisoned nftables conn on firewall errors ([#3008](#3008)) ([03f10e0](03f10e0)) * **orchestrator:** drop stale pre-init logs ([#3297](#3297)) ([8ec4be5](8ec4be5)) * **orchestrator:** emit compression ratios as fractions, not BP ([#2772](#2772)) ([866f4c1](866f4c1)) * **orchestrator:** export dirty-page stall counter from process start ([#2992](#2992)) ([badc8ad](badc8ad)) * **orchestrator:** harden Firecracker process shutdown ([#2996](#2996)) ([df662e7](df662e7)) * **orchestrator:** harden shutdown network cleanup ([#3000](#3000)) ([de2f391](de2f391)) * **orchestrator:** implement Docker COPY merge semantics in template builds ([#3283](#3283)) ([9174104](9174104)) * **orchestrator:** keep dedup empty-pages telemetry scan-only ([#2991](#2991)) ([35d0832](35d0832)) * **orchestrator:** let build-cache threshold flag raise above its fal… ([#3175](#3175)) ([06393c3](06393c3)) * **orchestrator:** log missing egress proxy in startup reclaim instead of defaulting silently ([#3116](#3116)) ([6ca3163](6ca3163)) * **orchestrator:** make copy-build handle filesystem-only snapshots ([#3299](#3299)) ([62add04](62add04)) * **orchestrator:** measure ext4 free space from block groups ([#3282](#3282)) ([f18f05f](f18f05f)) * **orchestrator:** normalize upload metric file labels ([#2767](#2767)) ([6dec8b3](6dec8b3)) * **orchestrator:** order egress config/firewall updates to close BYOP enable race ([#3313](#3313)) ([7faa59e](7faa59e)) * **orchestrator:** order envd.service after local-fs.target ([#3043](#3043)) ([ea2663e](ea2663e)) * **orchestrator:** order envd.service after systemd-tmpfiles-setup ([#3130](#3130)) ([9481811](9481811)) * **orchestrator:** pause upload retain retry ([#2993](#2993)) ([4f81799](4f81799)) * **orchestrator:** pin tap device host-side MAC address ([#3271](#3271)) ([3c786ba](3c786ba)) * **orchestrator:** pin UFFD copy source buffers ([#2745](#2745)) ([837fa91](837fa91)) * **orchestrator:** preserve full ENV value across stdout chunks ([#2740](#2740)) ([4822e6d](4822e6d)) * **orchestrator:** read V3 ancestors as uncompressed instead of failing ([#2994](#2994)) ([c479dd3](c479dd3)) * **orchestrator:** reject standby while draining ([#3325](#3325)) ([475a7ee](475a7ee)) * **orchestrator:** report real V4 header compression ratio ([#2771](#2771)) ([ecd344e](ecd344e)) * **orchestrator:** resolve remaining P2P/compression/V5 issues ([#3015](#3015)) ([1e4379e](1e4379e)) * **orchestrator:** sanitize OCI pull errors ([#3096](#3096)) ([a3af6c0](a3af6c0)) * **orchestrator:** scope rootfs hash to provision default ([#3129](#3129)) ([475f955](475f955)) * **orchestrator:** stop Checks health-loop leaking ([#2739](#2739)) ([17e6e60](17e6e60)) * **orchestrator:** survive SIGBUS from failing disks under mmap'd caches ([#3385](#3385)) ([8694d08](8694d08)) * **orchestrator:** tolerate missing header for legacy templates ([#3026](#3026)) ([8a44bfe](8a44bfe)) * **orch:** fall back to ID_LIKE with a warning instead of rejecting ([#3459](#3459)) ([73399b3](73399b3)) * **orch:** prevent NBD dispatch read-loop stall on WRITE_ZEROES (behind flag) ([#3048](#3048)) ([efd3d4d](efd3d4d)) * **orch:** split scheduling base build id per artifact ([#2920](#2920)) ([3e35a2a](3e35a2a)) * **orch:** validate copy-build -gdb buckets before the snapshot copy ([#3446](#3446)) ([9be382f](9be382f)) * **shared:** never report a failed envd command stream as success ([#3281](#3281)) ([69c06b6](69c06b6)) * **storage:** compression upload & cache correctness fixes ([#3231](#3231)) ([980748f](980748f)) * **storage:** don't assume V4+ ancestor gaps are uncompressed ([#3447](#3447)) ([f828d12](f828d12)) * **uffd:** dedupe deferred page faults ([#2864](#2864)) ([9680a41](9680a41)) * WrapContextAsUserError should not misclassify internal timeouts as user cancellations ([#3155](#3155)) ([8f83959](8f83959)) ### Performance Improvements * **build:** cache resolved Diff per BuildId within File.ReadAt ([#2838](#2838)) ([53de07f](53de07f)) * **build:** parallelize fragmented backing reads ([#2872](#2872)) ([c7655a7](c7655a7)) * **clean-nfs-cache:** restore dirfd-relative statx ([#2766](#2766)) ([6bdbedb](6bdbedb)) * **header:** add V5 columnar varint header format ([#2847](#2847)) ([9dd931b](9dd931b)) * **header:** pack cached Header.Mapping into a compact form ([#2844](#2844)) ([7f0b13c](7f0b13c)) * **orchestrator:** add memfile dedup density threshold ([#2862](#2862)) ([7ccfa02](7ccfa02)) * **orchestrator:** avoid V3-ancestor header refresh ([#2999](#2999)) ([cb6aa0b](cb6aa0b)) * **orch:** metrics for dirty page throttling ([#2858](#2858)) ([d2aa554](d2aa554)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: e2b-release-please[bot] <298072688+e2b-release-please[bot]@users.noreply.github.com> Co-authored-by: Charlie Wyse <charlie.wyse@e2b.dev>

A single bad 4K sector on a node's local NVMe crashed the whole
orchestrator and killed every sandbox on the node: the kernel cannot
return EIO for a memory access, so a failed page-in from an mmap'd cache
file delivers SIGBUS, which Go turns into an unrecoverable fatal error
inside the Cache.ReadAt memmove.
RunFaultSafe converts such faults into a regular error carrying the
fault address (debug.SetPanicOnFault + recover); other panics still
crash. It guards the in-process mmap copies: build segment reads, NBD
dispatch read/write/write-zeroes and the dedup compare. A bad block now
fails a single read or NBD request instead of the whole orchestrator,
and the affected file is just a local cache of remote storage anyway.
The guard is cheap enough to sit on the hot read path.
Recovered faults are logged and counted in
orchestrator.block.memory_fault, which should stay at zero on a healthy
node.