Skip to content

v2.2.3

Choose a tag to compare

@github-actions github-actions released this 24 Jul 13:07
· 249 commits to master since this release

Fixed

  • Fixed mitmproxy child process crashing with SIGABRT (V8 heap OOM) on corporate networks after running for ~1-2 hours under high-concurrency HTTPS traffic. Root cause: 7 log.error call sites in packages/mitmproxy/src/lib/proxy/mitmproxy/createRequestHandler.js passed the full rOptions object to jsonApi.stringify2(rOptions). Because rOptions contains the HttpsAgent instance (circular references), JSON.stringify throws, and stringify2's catch fallback (return obj) hands the raw object back to log4js, which then util.inspect-expands the entire agent — including every HttpsAgent.sockets key. On corporate networks each socket key is host:port::<full 141-cert PEM> (because loadExtraCaCerts merges tls.rootCertificates + the corporate root CA into the ca array passed to every agent), so a single error log produces a multi-megabyte string. Under high-concurrency error bursts (e.g. Xray-tunneled chatgpt.com requests failing with EPROTO), dozens of these dumps accumulate in the V8 heap faster than GC can reclaim them, exhausting the 96 MB old-space cap and aborting the process. Added a safeROptionsForLog(rOptions) helper that extracts only scalar fields (protocol, method, hostname, port, path, servername, headers, etc.) and strips agent/socket/ca; all 7 call sites now log jsonApi.stringify2(safeROptionsForLog(rOptions)), which serializes cleanly to a small JSON string with no fallback to the raw object. This keeps MemoryHigh at 512 MB (service template) / ≤300 MB (user constraint) viable without raising the V8 heap cap.

Changed

  • Disabled chromium zygote processes in service mode by adding --no-zygote switch in packages/gui/src/background.js when DEV_SIDECAR_SERVICE_MODE=true. In service mode, no BrowserWindow is created, so there are no renderer processes; the 5 zygote processes (chromium's pre-fork templates for renderers) are pure overhead, consuming ~39MB RSS and contributing shared file cache pages to the cgroup. This is set before app.whenReady().
  • Closed inherited chromium resource file descriptors in the mitmproxy child process entry point (packages/mitmproxy/src/index.js). The mitmproxy child is forked from the Electron main process via child_process.fork(), which inherits all non-CLOEXEC file descriptors opened by Electron — including chrome_100_percent.pak, chrome_200_percent.pak, resources.pak, locales/zh-CN.pak, icudtl.dat, v8_context_snapshot.bin, and /dev/shm/.org.chromium.Chromium.* shared memory segments. These are useless to the mitmproxy process but keeping them open prevents the kernel from reclaiming the corresponding file cache pages under cgroup MemoryHigh pressure. On Linux, the mitmproxy entry now walks /proc/self/fd and closeSync()s any fd pointing to chromium resources, pak files, icudtl, v8 snapshots, or /dev/shm/ — reducing cold-boot file cache pressure.
  • Moved Xray probe processes into an isolated cgroup (/sys/fs/cgroup/system.slice/dev-sidecar-xray-probe.scope) on Linux so their file cache does NOT count against the dev-sidecar service's MemoryHigh limit. On cold boot (after a full machine restart), the system page cache is empty; Xray probe processes read an 800MB+ SQLite cache (nodes_cache.sqlite), pulling ~137MB of file pages into the cgroup. Combined with the service's anon memory, this pushes memory.current to ~230MB against a 280MB MemoryHigh, triggering 3000+ kernel high reclaim events. On warm boot these SQLite pages are already in the system page cache (not charged to any cgroup), which is why warm-boot memory is low. The new moveProcessToIsolatedCgroup() in packages/core/src/modules/plugin/xray/util.cgroup.js creates a sibling cgroup with memory.high=max (no limit) and moves the probe PID into it immediately after spawn(). The probe's file cache is then charged to the isolated cgroup, keeping the service cgroup at warm-boot levels (~158MB) even on cold boot. A xray-probe-cgroup.sh helper script (installed by packages/gui/pkg/linux/postinst) encapsulates the cgroup operations, and a new sudoers rule allows the service user to run it via sudo -n without a password.
  • Added --no-sandbox alongside --no-zygote in the systemd service ExecStart. Electron requires sandbox to be disabled when zygote is disabled; without this, the service crashes with Zygote cannot be disabled if sandbox is enabled.
  • Increased reclaimStartupMemory reclaim amount from 100M to 200M in packages/core/src/expose.js. On cold boot (after a full machine restart, not just a service restart), the system page cache is empty, so the first load of the Electron binary (~150MB), app.asar, chromium runtime, Xray binary, and SQLite cache files produces ~200MB of cgroup file cache. The old 100M reclaim only covered half of this, allowing the cold-boot memory peak to reach the MemoryHigh limit (280M on this deployment), triggering 110 kernel high reclaim events. The new 200M reclaim drops file cache to near zero before the mitmproxy child process is forked, leaving headroom for subsequent Xray Stage3 probing. Added log.info/log.warn confirmation logging (previously the function silently returned true/false, making it impossible to verify from logs whether reclaim executed).
  • Raised V8 old-space cap and stage3 GC threshold for level 1 and level 2 in STAGE3_BATCH_LEVEL_TABLE (packages/core/src/modules/plugin/xray/config.js) and the mirrored STAGE3_MAX_OLD_SPACE_BY_LEVEL in packages/core/src/modules/server/index.js:
    • level 1: maxOldSpaceSizeMB 48→64, stage3GcThresholdMB 32→44
    • level 2: maxOldSpaceSizeMB 80→96, stage3GcThresholdMB 56→68
    • level 3-5: unchanged
  • The old level 2 values (80MB/56MB) were tuned for v2.1.x (HTTP/1.1 fake server). v2.2.0 introduced HTTP/2 fake servers (http2.createSecureServer), which increase per-request V8 heap pressure (Http2Session/Http2Stream JS wrappers). The old 80MB cap caused intermittent SIGABRT (V8 heap OOM) during Stage3 batch probing. The new 96MB cap gives 28MB GC buffer (was 24MB), and the GC threshold 68MB triggers earlier (was 56MB), reducing the chance of a GC-miss OOM.
  • Added LimitCORE=0 to the systemd service template (packages/gui/pkg/linux/dev-sidecar.service) to prevent 100GB+ Linux core dumps on SIGABRT. WSL2 crash dumps are controlled separately by .wslconfig MaxCrashDumpCount=-1.
  • Replaced Electron service-mode entry with a pure-Node entry point (service-entry.cjs) that runs with ELECTRON_RUN_AS_NODE=1, completely eliminating all chromium subprocess overhead in service mode. Previously, even in service mode (no BrowserWindow), the Electron main process still spawned chromium infrastructure: a GPU process, a NetworkService utility process, and (if not disabled) zygote processes. These contributed ~20MB RSS and ~50MB shared file cache (pak files, icudtl, v8 snapshots) to the dev-sidecar cgroup. The new service-entry.cjs is compiled by webpack into a standalone service-entry.js bundle at the asar root (alongside mitmproxy.js), with @docmirror/dev-sidecar and all its JavaScript deps inlined — only native modules (better-sqlite3, fadvise-linux, sysproxy) are loaded from the asar's node_modules at runtime. The systemd service ExecStart now runs /opt/dev-sidecar/@docmirrordev-sidecar-gui /opt/dev-sidecar/resources/app.asar/service-entry.js with Environment=ELECTRON_RUN_AS_NODE=1, which makes the Electron binary behave as a pure Node.js runtime (no chromium initialization at all). This drops the service cgroup memory from ~130MB (warm) / ~230MB (cold) to ~102MB, with only 3 processes total: the Node main process, the mitmproxy fork, and the Xray probe. The --no-zygote and --no-sandbox switches are no longer needed (there is no chromium to disable). The reclaimStartupMemory cgroup memory reclaim is also no longer needed in pure-Node mode (no Electron binary page cache to reclaim), but is retained as a safety net.
  • Added multi-layer cgroup memory reclaim during startup in packages/core/src/expose.js and packages/core/src/modules/plugin/xray/index.js. On cold boot, the mitmproxy fork + gsettings/D-Bus + SQLite cache reads produce ~180MB of cgroup file cache that pushes memory.peak to ~282MB (above the 280M MemoryHigh limit, triggering 260+ kernel high reclaim events). The new reclaim points execute memory.reclaim (via sudo -n /usr/lib/dev-sidecar/reclaim-memory.sh with NOPASSWD sudoers) at three stages: (1) before server.start() (reclaimStartupMemory, 200M), (2) after proxy.start() (dynamic 100-350M based on memory.current), (3) before SQLite cache reads in the Xray startup precheck (dynamic 100-300M). Each reclaim is followed by dropSqliteFileCache (POSIX_FADV_DONTNEED) to hint the kernel. Warm-boot peak dropped from ~280MB to ~245MB; cold-boot peak remains ~282MB (physical limit: cold-boot file cache cannot be fully reclaimed because subsequent reads immediately re-fault the pages, confirmed by workingset_refault_file counter).
  • Fixed stage3-initial-count memory.reclaim silently failing with EACCES in packages/core/src/modules/plugin/xray/index.js. The memory.reclaim cgroup file is --w------- root root, so direct fs.writeFileSync by the service user was silently caught and swallowed. Changed to use xrayCache.reclaimCgroupMemory() which has a sudo -n reclaim-memory.sh fallback (NOPASSWD sudoers rule). This was the root cause of the 280MB cold-boot peak — reclaim never actually executed.
  • Fixed Stage3 cache refresh hanging on corporate networks with SSL interception in packages/core/src/modules/plugin/xray/network_guard.js. The requestLocalNetworkCanary function used https.get with default certificate verification, which fails with UNABLE_TO_GET_ISSUER_CERT_LOCALLY on corporate networks because Node.js's built-in CA store does not include the corporate re-signing CA. This caused Stage3 to loop on "检测到本地网络离线,暂停当前批次" indefinitely. Added rejectUnauthorized: false to the canary request options — canary only checks TCP/TLS connectivity, not certificate validity, so this is safe.
  • Added StartupMemoryHigh=300M to the systemd service template (packages/gui/pkg/linux/dev-sidecar.service). Cold-boot peak (~282MB) slightly exceeds the runtime MemoryHigh=280M, causing kernel throttling during startup. StartupMemoryHigh=300M allows the startup peak through without throttling, then systemd automatically switches to MemoryHigh=280M once the service reaches active state. Also changed the template's default MemoryHigh from 512M to 280M (matching the actual deployment target). The install_devSideCar.sh deployment script's drop-in (10-devsidecar-setup.conf) still overrides MemoryHigh=280M for site-specific environments.
  • Changed mitmproxy child process --max-old-space-size from 5-level dynamic adjustment to a fixed 96MB in packages/core/src/modules/server/index.js. Previously, STAGE3_MAX_OLD_SPACE_BY_LEVEL mapped cacheRefreshBatchLevel (1-5) to V8 heap caps (64-512MB). This only affects the mitmproxy Node.js child process — the Xray probe binary is a Go executable in an isolated cgroup with no V8 heap. Mitmproxy steady-state heap is ~15MB, so 96MB gives 6x headroom for all throughput levels. Removed the STAGE3_MAX_OLD_SPACE_BY_LEVEL lookup table and batchLevel parsing logic.
  • Skip gsettings system proxy setup on headless Linux (no desktop environment) in packages/core/src/shell/scripts/set-system-proxy/index.js. On WSL2 and headless servers, gsettings is useless — no application reads GNOME proxy settings (apt uses /etc/apt/apt.conf.d/, curl uses http_proxy env var, git uses git config, npm uses ~/.npmrc, Docker uses docker.service.d/, Chrome does not read gsettings without a GNOME session). Previously, calling gsettings without a desktop session spawned three D-Bus helper processes (dbus-launch, dbus-daemon, dconf-service) via GLib autolaunch, consuming ~6MB RSS for no benefit. Now the Linux proxy setup checks for /usr/bin/gsettings and the X server socket (/tmp/.X11-unix/X<n>); if either is missing, it logs "跳过 gsettings 系统代理设置" and returns immediately, avoiding the D-Bus process spawn.
  • Removed Environment=DISPLAY=:0 from the systemd service template (packages/gui/pkg/linux/dev-sidecar.service). This hardcoded DISPLAY was only needed for gsettings, which is now skipped on headless systems. On desktop Linux, the login session sets DISPLAY naturally, so hardcoding :0 was both unnecessary and fragile (wrong if the user's display is not :0).
  • Changed sudo to sudo -n in packages/core/src/shell/scripts/setup-ca.js for consistency with the other three sudo call sites (expose.js, cache.js, util.cgroup.js). The -n flag makes sudo fail immediately if the NOPASSWD sudoers rule is missing, instead of hanging waiting for password input (which would deadlock the service since systemd services have no tty).
  • Fixed keep-alive socket reuse causing intermittent 7-second timeouts in packages/mitmproxy/src/lib/proxy/mitmproxy/createRequestHandler.js. The connectionTimer (7s) was cleared only via socket.once('connect'), but when agentkeepalive reuses an already-connected socket, the 'connect' event never fires, so the timer always expires — even for successful requests. This caused git clone and other sequential HTTPS requests to intermittently fail with 连接超时 after the first few requests succeeded. The fix checks socket.writable && !socket.connecting to detect already-connected sockets and clears the timer immediately without waiting for a 'connect' event.
  • Switched better-sqlite3 from the forked npm:@atomlong/better-sqlite3@12.12.0 alias to the official better-sqlite3@13.0.1 in packages/core/package.json and packages/gui/package.json. The @atomlong scope npm registry was experiencing intermittent 500 errors that broke CI builds and local pnpm install. The official 13.0.1 release supports Node.js >=22 and is compatible with the project's Electron 41 (Node 24) runtime. Removed the @atomlong:registry line from .npmrc and updated the rebuild-core-native.js warning message to drop the @atomlong reference.