Releases: atomlong/dev-sidecar
Release list
v2.2.10
Added
- WebUI 备份功能:配置目录备份/恢复到 S3 兼容对象存储(Cloudflare R2 / AWS S3 / MinIO) (fork). 侧边栏新增「备份」页(设置表单 + 上次备份状态 + 立即备份 + 云端备份列表的恢复/下载/删除)与 8 个 API 端点(
GET/POST /api/backup/config、POST /api/backup/test、POST /api/backup/run、GET /api/backup/list、GET /api/backup/download、POST /api/backup/restore、POST /api/backup/delete)。实现为两个新模块:webui/s3.js——最小化 S3 客户端(手写 AWS SigV4 签名 +node:https,仅实现备份所需的 Put/Get/List/Delete/Test 五操作,path-style 寻址,不引入 5MB+ 的 AWS SDK;R2 region 固定auto);webui/backup.js——备份服务。备份设置存<userBasePath>/backup.json(独立于主配置树:凭据不随GET /api/config泄露、不被整树回写流程篡改,也不进备份本体)。备份内容为整个配置目录 tar.gz 打包(自动排除logs/、xray/、running.json、*.pid、*.log、*.bak-*、backup.json本体;GNU tar 与 bsdtar 的./前缀差异用双份排除模式兼容),对象命名<prefix>/<hostname>/<时间戳>.tar.gz——多台机器共用同一 bucket 互不覆盖,保留份数(keepLast,0=不清理)也按主机前缀清理最旧;可选 AES-256-GCM 加密(scrypt 派生密钥,DSBK头含 salt/iv/tag,备份含 CA 私钥强烈建议开启),恢复时校验 gzip 魔数与 tar 完整性、当前config.json先留.bak-restore-<ts>安全副本再解包覆盖,needsRestart: true提示重启生效(前端接既有POST /api/service/restart)。定时备份为插件启动的 10 分钟周期检查(到达intervalHours间隔且配置完整才执行;schedule.enabled每 tick 实时读盘,改设置无需重启)。secret 掩码语义:GET 返回******,POST 回传掩码或不传即保留旧值(口令传空串显式清除);POST /api/backup/test支持「先测试再保存」——body 携带未保存设置,掩码secretAccessKey自动回退已保存值。测试:新增test/backup.test.js21 用例(SigV4 对本地假 S3 HTTP 服务的真实签名请求断言——Authorization 格式/credential scope/signedHeaders/payloadHash/path-style URI 编码/查询串排序/403 错误透传、加解密往返+防篡改、打包排除清单实测、保留份数按主机隔离清理、恢复回环+安全副本、口令缺失/错误提示、非法 key(跨前缀/..)拒绝、调度 tick 到期执行/未到期跳过/关闭后不执行)+webui.test.js备份路由组 9 用例(mock 注入context.backupApi(同 xrayApi 注入模式):脱敏读写、数组 body 400、400BACKUP_NOT_CONFIGURED与 502BACKUP_UPSTREAM_FAILED错误映射、恢复 key 缺失 400、下载流 attachment 头、未知备份路由 404 兜底;期间抓出真 bug——GET /api/backup/list路由漏await异步listBackups()把 Promise 序列化成{})。webui 70 passing、主套件 193 passing。真实 Cloudflare R2 全链路实测打通:控制台创建存储桶(dev-sidecar)→ 创建最小权限 Account API 令牌(对象读和写、仅限该桶)→ WebUI 配置(测试连接 → 保存 → 立即备份)→ R2 控制台与 WebUI 列表双侧确认加密归档上云;操作教程见doc/wiki/备份到Cloudflare R2使用说明.md(含新机灾备与常见问题)。Windows 兼容(三平台发布 CI 直接暴露的坑):GNU tar 会把带盘符的归档路径C:\...误判为远程主机语法(Cannot connect to C: resolve failed),tar 的-f参数统一改用裸文件名 +cwd定位(GNU tar / bsdtar 双兼容),修后三平台 CI 全绿。 - export 节点导出新增
alive=true:live observatory 实时存活过滤(下游反馈) (fork). 下游(CharmHyper_Register)实测:available=true过滤出的 351 个"可用"节点随机抽 60 个全部 SYN 挂起——缓存 delay 来自 Stage3 探测快照、最长落后一天,available只保证"某次探测时活过",不等于"现在活着";同一时间点 observatory 实时探测 20/20 alive 且手动起独立 xray 出口验证通过,但这些实时活节点因缓存 delay 过时沉在 smart 排序(delay 升序)尾部,顺序遍历的消费方根本取不到。新增alive=true查询参数:从 live xray metricsPort 的/debug/vars读 observatory 实时状态,经插件getLiveNodeFingerprints()的 tag→fingerprint 映射(无需解析 outbound)关联缓存行(readCacheEntriesByFingerprints)——meta.delay用 observatory 实时值排序与展示(覆盖过时的缓存 delay),meta新增lastTry(最近探测 unix 秒时间戳,判断存活数据新鲜度);可与available(连败阈值)/maxDelay(按实时 delay 判断)/country/owner/sort/shuffle/limit/offset组合,total= 过滤后存活总数;xray 未运行时 HTTP 200 +total=0+reason=xray_not_running(data为对应 format 的空结构)。消费方可从"拉全池本地 shuffle 逐个实测(平均 25 个命中)"改为alive=true&shuffle=true&limit=30轻量拉取。测试:webui.test.js新增 8 用例(mock/debug/vars+ fingerprint 映射:实时 delay 覆盖与排序、available 连败阈值、maxFailureStreak 放宽、maxDelay 实时过滤、country 过滤、offset/limit 分页、shuffle 组合、xray 未跑空集+reason)。 - WebUI 配置页图形化编辑 + 侧边栏布局:textarea 嵌入 JSON 全部替换为结构化表单 (fork). 规划见
doc/webui-config-refactor-plan.md。顶部标签页改为左侧侧边栏(品牌区 + 导航 + 底部版本号/WS 连接状态点,窄屏回退横条);配置页重构为左侧 14 分区导航 + 右侧表单(加速服务/拦截规则/Xray 插件/DNS/域名白名单/预设 IP/代理排除/梯子/网络检测/Git·Node·Pip/远程配置/WebUI/应用与日志/原始 JSON 只读),控件经data-path路径绑定直写 draft(增删行才重渲染,不打断输入),带未保存更改计数、放弃更改、每分区恢复默认。拦截规则(server.intercepts)做三级结构化编辑器:域名列表(搜索/添加/重命名/删除,metaInfo过滤,config.json 有覆盖的域名带"用户"标记)→ 路径正则行(new RegExp实时校验着色,失焦重命名)→ 动作字段编辑器——proxy(+backup/test)、sni、redirect、abort、success、script、tampermonkeyScript、请求/响应头替换行编辑([remove]哨兵提示)、缓存族数值+单位下拉(读时识别已用键,cacheSeconds/Minutes/Hours/Days/Weeks/Months/Years)、desc/remark,未知键折叠"高级字段"逐键兜底编辑不丢数据;值可为null墓碑(屏蔽默认/远程继承)→ 显示"已屏蔽"标记可恢复继承。全部修改持久化到~/.dev-sidecar/config.json(用户覆盖层,合并优先级最高),语义与 GUIconfigApi.save完全一致——删除远程/默认配置来源的条目时 doDiff 写null墓碑、合并时剥除(已用真实 configApi 在隔离 HOME 验证:删除远程来源域名 → config.json 落 null → 重载后生效配置消失);remote_config_personal.json5不被程序改写,保持手工注释。被 config.json 覆盖的字段/域名显示"用户"来源徽章(新端点GET /api/config/user)。保存后若plugin.xray.*/server.setting.xrayPort变更且 Xray 运行中,自动调POST /api/xray/restart(GUI applyBefore 同款)。新端点:GET /api/config/user(用户覆盖层+远程元信息)、POST /api/config/reset(分区恢复默认,key 白名单正则防任意路径)、POST /api/xray/restart;api()携带 Bearer token + 401 弹框重试(远程 token 访问此前无法保存);routes.js 静态路径加仓库内回退(开发态无需安装)。测试:webui.test.js新增"webui config view / reset / restart routes"组 5 用例 + write-ops 组改造 4 用例(整树/局部树 400、configFromFiles 剥离断言、intercepts 子树替换断言),webui 61 passing;生产 31182 真机全链路验证(PUT 无损回环 config.json 落盘前后完全一致、编辑-放弃回环、来源徽章对真实 config.json 命中、xray restart 端点 8s 恢复 20 节点且 sticky 正确 re-apply)。deb 重新构建部署:electron:build --linux deb→dpkg -i。
Fixed
- 修复 export
shuffle=true在候选池小于 limit 时完全失效(下游反馈) (fork). 下游实测同参数两次调用(间隔 31s,已过 30s 响应缓存窗口)返回顺序完全相同,与shuffle=false行为一致;配合available=true的过时 delay 数据,固定顺序导致顺序遍历的消费方反复在同一批 smart 排序头部的死节点上空转(3 轮 × 30 节点全部失败)。根因:打乱逻辑带result.length > limit守卫——仅当候选数大于 limit 才执行 Fisher-Yates,而消费方场景是 available 池 351 条 <limit=500,打乱被整体跳过,smart 排序固定原序返回(下游猜测的"候选池 ≥ 2N 时抽样被跳过"方向相反:是池子 ≤ limit 时不打乱)。修复:shuffle=true无条件打乱,候选池不足 limit 时同样全量随机排序。真机以消费方原始参数(available=true&maxFailureStreak=5&sort=smart&shuffle=true&limit=500)复现:三次调用首节点各不相同(其中一次恰好随机到反馈文档中的原固定首节点vless://58c1b245-...,证明同一池子仅随机化生效)。测试:webui.test.js回归用例——4 节点池(< limit=100)连续 12 次shuffle=true调用断言出现 >1 种排列(4 节点 24 种排列,12 次全同序概率 ~1e-15)。 - 修复 WebUI 配置保存的删除静默失效:PUT /api/config 由 merge 改整树 save 语义 (fork). 原实现 PUT /api/config →
globalConfig.update(body):lodash.mergeWith无法表达"删除键"——前端删掉的域名在 merge 阶段复活、doDiff 视为无变化不落盘,任何删除操作都静默失效(GUI 不受影响,它直接调configApi.save(整树));PUT /api/intercepts、/api/presetiplist同病(对象子树 merge 删除失效),PUT /api/xray/rules因数组整体替换恰好幸存。修复:PUT /api/config 改configApi.save(整树)(body 必须含app/server/plugin顶层键,局部树 400,防未携带子树被墓碑化误删);其余三者改为 clone 当前树 →lodash.set整体替换子树 → save,删除经子树替换正确墓碑化。配套两处运行态剥离:GET /api/config剥离 Xray 自动注入条目(desc === 'Auto-injected by Xray Plugin')与configFromFiles调试快照(模块加载时烘焙进默认配置的合并副本,不剥离则前端把运行态当用户配置回写产生幽灵 diff);PUT 时恢复configFromFiles快照再 save(真实 doDiff 以含此键的默认配置为基线,缺失会被墓碑化为configFromFiles: null写进 config.json——GUI 的 draft 未剥离故无此问题,WebUI 的 GET 剥离造成的不对称在此补齐)。测试:webui 61 passing、主套件 172 passing 无回归;真机 PUT 无损回环验证 config.json 落盘前后字节级一致。
Changed
- 工作区包版本号 bump(core/gui/mitmproxy/cli 2.2.9 → 2.2.10),
/api/version随之生效 (fork). 下游靠版本号探测 export 能力(v2.2.9 已把/api/version对齐 corepackage.json实际版本,此前运行中服务报旧版本号导致下游误判能力缺失),本次 export 契约新增alive=true与 shuffle 行为修复将随 2.2.10 发布生效;deb 文件名随之变为DevSidecar-2.2.10-amd64.deb。webui 测试套件 74 passing(+9)、主套件 172 passing 无回归。 - WebUI 探测页出站节点详情表"地址/端口"列改为"服务器/出口"列 (fork). 20 个 live 节点实测仅 3 个出口 IP(16 个共享同一 OVH 落地)——入口地址各异的 CF 中转变体让旧表"地址"列看起来每个节点出口都不一样,产生误导。"地址"(入口连接点)与"端口"合并为"服务器"列(
IP:端口形式),腾出的"出口"列直接显示缓存元数据的 egress IP(为空显示-),多节点共享出口一眼可见;表排序键addr/port同步改为server/exitip。出口 IP 数据后端/api/xray/nodes的nodeMetadata本就返回(无需后端改动),仅前端renderStage1Table消费。已部署本机实测:出口列 51.159.163.185/51.15.243.182 交替出现,集中性直观可见。Stage1 卡片"当前选择"右侧同主题新增"出口地址"格,显示当前 balancer 选中节点的 egress IP(随轮询刷新)。
v2.2.9
Changed
- 发布 v2.2.9:工作区包版本号 bump(core/gui/mitmproxy/cli 2.2.8 → 2.2.9),
/api/version随之生效 (fork). 下游反馈:export 等 v2.2.9 新能力已随 deb 部署,但运行中服务/api/version仍报2.2.8,消费方无法据版本号探测"export 可用与否"。本次发布正式 bump 四包版本(/api/version读 corepackage.json),deb 文件名随之变为DevSidecar-2.2.9-amd64.deb。顺带澄清 exporttotal语义(文档webui-api.md):total= 当前过滤条件下的总条数——带available等过滤时为过滤后总数,完全无过滤时为全库节点数(实测 38.6 万);并把"非 2xx 响应不含total/data字段"写入契约——下游实测的total=None(Python)是把429 RATE_LIMITED限流响应体当正常响应解析所致,消费方必须先检查 HTTP 状态码/error字段。 - WebUI 日志页接入 mitmproxy 子进程日志:模块下拉不再只有 core(core + server) (fork). v2.2.9 日志重构的已知边界(当时 CHANGELOG 明示"mitmproxy 子进程的 server.log 不再进 WebUI"):mitmproxy 是 fork 的独立进程,其
getLogger('server')的 log4js 实例在子进程内存里,主进程的环形缓冲天然收不到——服务模式下guicategory 又无人写日志,于是模块下拉永远只有一项、模块过滤形同虚设。本次把 server 日志接进实时缓冲:IPC 转发——core fork mitmproxy 时注入DS_LOG_IPC_FORWARD=1标记 env,util.log-ring.jsappender 在标记+process.send同时存在时不本地存储、直接process.send({ __dsLog: entry })(消息全 JSON 安全字段;断开时静默降级回文件/stdout);主进程server/index.js的既有message监听头部新增__dsLog分支调appendExternal()防御性清洗(非法 level/category/非对象丢弃、超长二次截断)后写入主进程环形缓冲,categories 下拉自动出现server,WS 实时推送同时覆盖。坑:不能只判typeof process.send === 'function'——pnpm/mocha 跑测试时 mocha 本身就是 IPC fork 子进程(首版实现即在此翻车,本地存储被误判成转发),必须以显式 env 标记为准。message监听里__dsLog分支放在既有 debug 序列化之前提前 return,避免每条日志一次JSON.stringify高频开销。测试:logRing.test.js+5 用例(appendExternal 入 ring+订阅推送、防御性丢弃+截断、转发不本地存、IPC 断开静默降级、pnpm fork 有 send 但无标记不转发)。主套件 172 passing、webui 65 passing、mitmproxy 17 passing。 - WebUI 日志页重构:文件轮询改为结构化实时日志(时间/等级/模块/消息) (fork). 原实现每 5s 读日志文件尾巴 dump 进
<pre>,无结构、无过滤。新架构:core 新增 log4js 自定义 appenderutil.log-ring.js——结构化环形缓冲(容量 3000、消息截断 2000 字符、内存上界 ~6MB),挂到全部 category;GET /api/logs改查环形缓冲(level=最低等级、q=消息/模块子串不分大小写、category=精确模块、limit=最新 N 条时间升序,并返回去重模块清单与容量);实时增量经既有 WS/ws新增log频道推送(webui 插件启动时订阅环形缓冲、关闭时退订)。前端日志页重写:四列表格(sticky 表头、等级色标 debug 灰/info 蓝/warn 黄/error 红)、搜索防抖、等级/模块下拉过滤(模块清单自动收集)、实时追尾(贴近底部自动滚、上翻暂停并出现"↓ 底部")、行点击展开全文、清屏、显示 X / 匹配 Y / 缓冲 Z统计;DOM 行数上限 1000 防长会话膨胀。readLogTail文件读取路径整体移除;注意:日志范围从"core/server/gui 三个日志文件"变为"本进程实时缓冲"(service 模式即 core+gui 分类;mitmproxy 子进程的 server.log 不再进 WebUI,仍落盘——已于后续条目经 IPC 转发接回)。测试:logRing.test.js9 用例(结构化捕获含对象/Error 堆栈、容量淘汰、截断限长、level/q/category/limit 过滤、订阅推送/退订/订阅方异常隔离、模块去重)+webui.test.js路由级 4 用例(结构化返回、min-level、q 大小写、category+limit);旧文件式契约 2 用例随端点退役移除、stable code field用例改锚 404。webui 套件 65 passing、主套件 163 passing。 - WebUI 探测页出站节点表 SNI 列改为 ASN 列 (fork). SNI 对节点选择没有实际参考价值(多为空或节点自身伪装域名),换成缓存元数据的 ASN 组织归属(
nodeMetadata.owner,即 ASN mmdb 查得的归属组织,如 "AMAZON.COM, INC."),与allowedOwners筛选语义直接对应、便于核对!cloudflare排除是否生效;列排序同步切换,缓存页"已探测节点"表同语义列头一并由 owner 改为 ASN(仅显示标签,内部字段/DB/API/export 契约的owner命名不动)。 - Stage3 轮次显示改为独立计数
roundNumber(从 1 起),不再复用内部generation代际 (fork). WebUI"轮次"原显示generation——它是进程内失效代际(启动 Stage2 抓取占 docmirror#1、close()也会 +1),导致服务重启后首个 Stage3 探测轮显示 docmirror#2。新增stage3RoundCount只在refreshCacheFromCacheOnly真实入轮时 +1,经getStageStatus().stage3.roundNumber暴露;generation保留为内部失效保护(Stage2/批次守卫的generation === refreshGeneration检查不变)。WebUI 轮次显示切换到roundNumber,webui-api.md示例同步。回归测试:xrayStageGating.test.js新增用例——空缓存连续两轮断言 roundNumber 0→1→2 且每轮后 state='idle'(第二轮顺带复验空轮 isStageRunning 复位)。 - Stage3 节点判死阈值从 3 连败放宽到 4 连败,退避阶梯 [7,30,90] 天完整生效 (fork). 原实现
nextFailureStreak >= 3在第 3 次连续探测失败时即删除节点+写墓碑——continue跳过了退避赋值,导致CACHE_FAILURE_BACKOFF_DAYS的 90 天档位成为不可达死代码(第 3 败永远走不到getFailureBackoffMs(3)),节点从入库到判死最短仅 ~37 天(7+30),"退避最长 3 个月"名不副实。改为>= 4:第 3 败保留并退避 90 天,第 4 败才删除+写墓碑,最短判死周期 ~127 天(7+30+90),死节点翻身窗口完整覆盖整个阶梯。live 池过滤保持 3 不变(maybeRegenerateLiveConfigFromCache的failureStreak < 3保留过滤与批次 rmo 超限裁剪的streak >= 3)——3 连败节点仍会被及时移出 xray 运行时池不扛真实流量,只有缓存保留/墓碑判定放宽。xrayStageGating.test.js三个用例同步更新:third-strike 用例改名 fourth-strike 且 streak 种子 2→3,多轮模拟补第 4 轮(90 天退避后的 2026-09-26)并断言前 3 轮节点均保留。
Fixed
- 修复服务重启后 Xray 报"端口 10801 被占用 (Strict Mode)"且插件永久停摆——孤儿 Xray 进程自愈清理 (fork). 生产实锤:服务 restart 后
Xray 启动失败: 端口 10801 被占用 (Strict Mode),但ss -tlnp显示 10801 被上一代服务残留的 xray(ppid=1孤儿)监听着。根因链:主 Xray 为内存优化被移入隔离 cgroup(dev-sidecar-xray-probe.scope),systemdKillMode=control-group在 stop 时根本杀不到它;上一代服务若被 SIGKILL 兜底(shutdown 超时)或 stop 竞态漏杀,xray 就成孤儿占住本地端口,新服务 Strict 检查发现端口被占直接放弃,且无任何自愈手段(2026-05-23 就出现过一次,当时手动恢复)。修复:process.js主 Xray spawn 成功后写 pidfile(<xrayDir>/xray.pid)、stop()完成后删除;新增cleanupStaleProcess(binPath, configPath)——读 pidfile → 验存活 →/proc/<pid>/cmdline双重身份验证(必须同时含 xray 二进制路径与 live config 路径,pid 复用绝不误杀,非 Linux 无 /proc 不杀) → SIGTERM 等 3s → SIGKILL 兜底 → 清 pidfile;index.jsStrict 分支端口被占时先自愈清理再重查一次(清理不了才报原错误)。测试:新增xrayStaleProcess.test.js4 用例(真实进程:自家进程按 pidfile 被杀+pidfile 删除、cmdline 不匹配的外来进程拒绝杀、死 pid 清 pidfile 返 false、无 pidfile 返 false)。主套件 167 passing。 - 修复启动/轮末候选窗口 SQL 层不应用 country/owner/maxDelay 筛选,轮末重生成把 live 池砍到 1 个节点(用户报 7→1) (fork).
readCacheEntriesForStartup(bootstrap 与轮末maybeRegenerateLiveConfigFromCache回填的共同候选源)只按延迟取probed_node_ids前limit个,country/owner/maxDelay 全留给 JS 后置——生产实锤:宽测轮跑完后存活池 537 个,top-100 低延迟窗口被 cloudflare-owned 的合规国家节点霸榜 71 个(CF WARP 延迟最低),!cloudflare后置过滤后只剩 1 个候选;轮末重生成据此把 live 池里 19 个健康注入节点全部 rmo(liveNodes=1, kept=0, added=0, removed=19),全量合规节点其实有 183 个。bootstrap 同雷:下次重启候选也会饿死到 0-1 个。修复:窗口改为直接 SQL 查询node_runtime_v2,用现成的buildCompactV2FilterClauses在 SQL 层应用 country/owner/maxDelay/stable 筛选后再ORDER BY delay ASC, stable DESC, updated_at DESC LIMIT n(同 partial index,代价与每批维护 probed 列表等价);生产 DB 验证窗口 100/100 全合规。回归测试:xrayCacheOrdering.test.js新增 CF-starvation 用例(低延迟 CF-SG/HK 霸榜 + 高延迟非CF合规节点必须入选 + 无筛选取全局 top + 全筛空返回 []);xrayStageGating.test.js既有用例从"ignores filters"断言更新为"filters 窗口内生效"。主套件 154 passing。 - 修复 WebUI 探测页出站节点表 vmess 协议地址/端口列为空 (fork). 出站节点详情表从
/api/xray/nodes(xray API 的 TypedMessage 序列化形态)提取地址/端口,链路覆盖 trojan/ss 的proxySettings.server、vless 的proxySettings.vnext、ss-2022 扁平字段——但 vmess 在该序列化下的字段是大写Receiver(proxySettings.Receiver.address/port),恰好 live 池 20 节点中仅有的 2 个 vmess(proxy_85/proxy_127)地址端口显示为空。修复:addrOf/portOf补Receiver回退级。已用生产 live API 真实数据验证:旧链空 2 个、新链 20/20 全提取(proxy_127→168.138.43.75:80、proxy_85→165.140.216.142:443)。 - 恢复 Stage3 宽松探测 / Stage1 严格探测的宽严分离设计(用户澄清被 92601af 重构漂移破坏) (fork). 原设计(
4c2efb2f):Stage1 bootstrap 与 live observatory 用observatoryProbeUrl(严格,接近真实目标),Stage3 周期探测用probeUrl(宽松,gstatic 204)粗筛——缓存是共享资产,严格探测会把"能到 gstatic 但到不了 chatgpt.com"的节点判死并经 4 连败阶梯写墓碑,误删对下游 export(按国家/owner 取节点)有价值的节点。92601af3(08-21,固定模板+ado/rmo 重构)在重构时把 Stage3 的常驻探测与逐批探测两处调用点复制成了 bootstrap 的严格表达式(提交未提及探测语义,属无意漂移)。修复:两处 Stage3 调用点恢复cfg.probeUrl || pluginConfig.probeUrl,probeNodesBatch注释同步恢复宽严分离说明。契约锁定:两处 URL 解析抽成具名纯函数resolveStage3ProbeUrl/resolveStrictProbeUrl(全部 5 个调用点改走它们),xrayStageGating.test.js新增宽严分离契约用例 2 个——"Stage3 只认 probeUrl,observatoryProbeUrl 配了也不得渗入"(92601af3 漂移现场反向验证过:人为复刻漂移该用例精确变红)+ "Stage1 严格优先、未配置时回退 probeUrl",未来重构再复制粘贴严格表达式进 Stage3 会直接挂测试。已知权衡:轮内批次注入现以宽松存活为准注入 live 池,严格 observatory 不过的节点占槽但 leastPing 永不选中(流量不受影响);live 池可用节点多样性取决于严格/宽松通过率之差,重启 bootstrap(严格 burst)会净化。 - 修复 Stage3 轮内逐批次"边测边注入"路径完全绕过 allowedCountries/allowedOwners 筛选 (fork). 生产实锤:
allowedCountries=["US","DE","SG","FR","GB"]配置下,live 池 20 节点中 6 个越界(HK/JP/NL×2/FI×2),balancer 甚至把 🇭🇰 节点选为当前出口。根因:国家/owner 筛选只存在于 Stage1 启动 bootstrap 与轮末maybeRegenerateLiveConfigFromCache回填两条路径,而 Stage3 轮内每批次探测成功后的 ado 热注入(缓存周期探测批次已写回→Inject available nodes from this batch)只检查delay>0 && 节点合法,rmo 超限裁剪也只按失败/延迟——任何探测存活的越界节点都会被直接注入 live 并可能凭低延迟挤掉合法节点。修复:新增selectLiveInjectableEntries(复用collectBootstrapCandidateEntries的双重筛选语义,country 未知在白名单非空时同样拦截),批次注入前统一过滤,被筛掉时打注入前国家/owner筛选: X/Y 通过日志。轮末回填对 keptNodes 只查存活不复查国家(节点 country 漂移仍会保留,属已知边界)。回归测试:xrayStageGating.test.js新增 5 用例(HK 白名单外拦截、country 未知拦截、!cloudflareowner 排除、空筛选全放行、空输入)。 - 修复 Stage3 空批次轮永久卡死"运行中",后续所有轮被跳过(用户报 WebUI 显示异常) (fork).
refreshCacheFromCacheOnly在入口置isStageRunning = true后、主路径try/finally(复位所在)开始之前有一个"当前没有到期的可探测节点"早退return——空批次轮从此不复位:WebUI 永显"运行中"、耗时无限增长(实测卡 4h+)、nextRefreshAt过期显示"(已过期)";更严重的是后续所有定时触发都命中"Stage 正在运行,跳过本轮"守卫,Stage3 整体停摆直到服务重启(生产实测:13:22 空轮后 16:22 的下一轮被跳过)。修复:早退前复位isStageRunning = false。回归测试:xrayStageGating.test.js新增空批次轮用例——实例化真实 Plugin 工厂(fake context)+ 临时目录播种零节点缓存,直接调用refreshCacheFromCacheOnly命中早退路径,断言getStageStatus().stage3.state === 'idle';已在撤掉修复行的代码上验证该用例精确复现'running' !== 'idle'卡死症状。 - 修复 WebUI
/api/xray/cache/nodes/export全部过滤参数静默失效 + 默认 sharelink 格式崩溃(下游反馈) (fork). 四个叠加缺陷:①nodeToShareLink是packages/core/src/modules/plugin/xray/index.js的模块顶层函数但从未导出——export 路由require('../xray/index')解构出undefined,默认format=sharelink首次调用即抛 "nodeToShareLink is not a function"(生产包压缩后为e is not a function),被 catch 包装成误导性的CACHE_NOT_READY;② export 路由传给readCacheEntries/countCacheEntries的选项键名(sort/available/maxDelay/countries/owners/protocols/since)与缓存层buildCompactV2FilterClauses认的键(orderBy/maxDelayMs/countryInclude/ownerInclude)完全不匹配——所有过滤静默 no-op,total返回全库数而非过滤后数;③available连败阈值过滤在缓存层不存在;④nodeToShareLink的 ss 分支只读settings.servers[0],ss-2022 扁平结构返回空分享链接。修复:导出nodeToShareLink并给 ss 分支补 ss-2022 扁平回退;键名重映射 + 接availableOnly/maxFailureStreak(delay>0 AND failure_streak < 阈值,默认 3 可调);getCompactV2OrderByClause补delay/stable排序模式(顺带修复前端缓存页sort=delay|stable自 v2.2.8 起即 no-op 的问题);meta 补stable字段。该接口现满足下游按条件取完整可起 xray 的 outbound(如每注册任务按国家过滤随机取 N 节点)的场景,契约见doc/webui-api.md。 - 修复 WebUI
/api/xray/cache/nodes行投影 address/port 对 vless/vmess/ss-2022 为空串(下游反馈) (fork). 行投影只读settings.servers[0](trojan/旧 ss 形态)——vless/vmess 的地址在settings.vnext[0]、ss-2022 在扁平settings.address/port,均取不到(与 v2.2.9 前端addrOf/portOf同型缺口,后端漏改)。修复:按协议结构三级回退提取(settings.server || settings.servers[0] || settings.vnext[0] ||扁平字段)。webui-api.md同步:rows 示例对齐实测 10 字段(原文档误含不存在的total字段)、export 章节按实际契约重写(原文档误写成"后...
v2.2.8
Changed
- WebUI 重构:合并探测页 + 局部刷新 + 国家 emoji (fork).
packages/gui/extra/webui/index.html重写:合并"Xray 节点"和"探测状态"为"探测"页(Stage1/Stage2/Stage3 三卡片),去掉"上轮汇总"。Stage1 卡片显示主 xray 进程状态/当前选中节点(含国家 emoji)/出站节点数,右下角放 Balancer 锁定/解锁控件 + 时长下拉(锁定时禁用),完整 20 节点详情表可折叠展开(列排序)。Stage2 卡片显示开关/间隔/下一轮预计时间/上轮耗时/上轮抓取数。Stage3 卡片显示开关/轮次/批次进度/节点进度/耗时/下一轮时间/本轮可用/本轮失败。仪表盘新增版本号卡(GET /api/version,含 dev-sidecar + Xray core + Node.js 三版本),端口卡补充 31180 HTTP 代理,去掉"重启服务"按钮。日志页加手动刷新按钮 + 自动刷新切换 + 间隔下拉(2s/5s/10s/30s)。所有面板改局部刷新(textContent更新,不重建 DOM),避免 10s 自动刷新时界面抖动。缓存/配置页骨架优先渲染(固定卡片布局立即显示 +-占位 + "加载中..." tbody,API 返回后局部填充)。国家字段用String.fromCodePoint算法生成 emoji 国旗(ISO 2 字母代码),内联TwemojiCountryFlags.woff2字体(base64 data URL)解决 Linux Chromium 不渲染国旗 emoji 问题。导航精简为 5 项(仪表盘/探测/缓存/日志/配置)。 - Stage1 bootstrap 后立即临时锁定出口节点 (fork). 解决"启动后 5 分钟内无可用出口"问题(observatory
probeInterval=300s,第一次探完前 balancer 无延时数据无法 leastPing 选节点)。Stage1 ado 注入完成后立即调用overrideBalancer锁定延时最低的节点,sticky 状态通过setTimeout(probeInterval * 1000)自动到期后removeBalancerOverride解锁,让 observatory 的 leastPing 策略接管。手动解锁会clearTimeout取消自动解锁 timer,与定时器互斥无竞态。 - WebUI Xray 节点字段提取规则 (fork). 从
_TypedMessage_(带尾部下划线)提取协议,从proxySettings.server(trojan)或proxySettings.vnext(vless)提取地址,从streamSettings.protocolName提取传输协议,从securitySettings[0].serverName提取 SNI,从metricsRes.observatory(不是observation)提取延时和存活。 - WebUI Balancer 改为解析式展示 (fork). 不再展示 YAML 原文,解析出当前选中节点 tag 显示;锁定/解锁按钮 + 时长下拉框(5min/10min/30min/1h/24h/永久)右侧布局。
- WebUI 延时显示单位 (fork). 延时直接显示毫秒(不除以 1e6)。
- 远程配置未启用时不删除已有文件 (fork).
packages/core/src/config-api.js的downloadRemoteConfig/doDownloadRemoteConfig原逻辑在远程配置未启用或 URL 为空时调用deleteRemoteConfigFile删除本地文件。用户可能手动放置了remote_config_personal.json5,不应被空 URL 删除。改为只跳过下载,不删除已有文件。 - config.update 用 mergeWith 替代 merge 避免数组合并 (fork).
packages/core/src/config-api.js的update(partConfig)原用lodash.merge合并配置,对数组字段(如plugin.xray.rules)会按索引合并而非整体替换,导致用户想替换整个数组时旧元素残留。改为lodash.mergeWith+ customizer,对源端数组返回srcValue(整体替换),同时cloneDeep目标避免污染原配置。 - WebUI Stage2/Stage3 三态状态 + Stage2 运行实时进度 (fork). 用户反馈 Stage2/Stage3 应区分三种状态(已关闭/空闲/运行中),且 Stage2 运行中看不到进度。
getStageStatus()新增结构化state字段(stage2:off/idle/running,running=远端订阅抓取进行中,新增isStage2Running标志覆盖启动时后台同步与 Stage3 轮末触发两个入口,三处出口(含异常兜底)复位;stage3 同款三态 +enabled字段);删除与state === 'running'完全等价的冗余isRunning字段。Stage2 新增本轮实时数据:progress: {current, total}(正在抓第几个订阅/订阅总数,loadSubscriptionNodes新增onSubscriptionProgress回调,模块级函数经 options 传回插件闭包)、startedAt(本轮开始时间,前端算实时耗时)、fetched(本轮已抓取节点数,实时累计;结束后回落lastSyncFetchedCount)。nextTriggerAt = max(nextSyncAt, 下一轮 Stage3 开始时间)替代原"待触发"显示——Stage2 在 Stage3 轮末按需触发,前端显示"约 HH:MM:SS"。前端 Stage2 卡片:状态三态、新增"进度"格、"上轮耗时"→"耗时"(运行中显示本轮实时耗时并随自动刷新跳动,结束后为本轮总耗时)、"上轮抓取"→"抓取";Stage3 耗时字段同步改用state。
Added
-
WebUI 监控面板 (fork). 新增
packages/gui/extra/webui/index.html单页面 Web UI(端口 31182),用于无显示器服务器场景的 dev-sidecar 状态监控。面板包含:Dashboard(Xray/系统代理/服务器/Stage 状态概览)、Xray 节点(协议/地址/端口/SNI/传输/延时/存活状态 + 列排序)、缓存(统计卡片 + 已探测节点详情 + 国家分布 + 最优节点 + 订阅源)、Stage(探测进度/批次/候选数)、日志(tail 实时滚动)、配置(只读视图)、Balancer(解析式展示当前选中节点 + sticky 锁定/解锁下拉时长 5min/10min/30min/1h/24h/永久)。骨架优先渲染(render(null)占位 → API 返回后填充),避免加载闪烁。每 10 秒自动刷新当前激活面板。 -
/api/xray/nodes节点列表 API (fork).packages/core/src/modules/plugin/webui/routes.js新增节点列表接口,调用xray api lso(ListOutbounds)获取 live xray 进程的 outbound 列表,JSON.parse 后过滤 direct/block/metrics 节点,提取 protocol/address/port/SNI/transport 字段;合并 observatory metrics 数据(从/debug/vars拉取)展示延时和存活状态。 -
/api/xray/cache/nodes分页节点接口 (fork). 新增缓存节点分页接口(page/pageSize/sort参数),sort=smart时按延时升序 + 失效节点置底。 -
/api/xray/cache/stats缓存统计接口 (fork). 返回{ totalNodes, dbSizeBytes, countryDistribution },新增packages/core/src/modules/plugin/xray/cache.js的readCountryDistribution(cacheFilePath, limit)SQLGROUP BY country聚合国家分布。 -
/api/xray/probed-stats已探测节点统计接口 (fork). 读取probed-node-stats.json,返回{ totalProbed, countryDistribution, nodes }。 -
/api/xray/balancer返回 sticky 状态 (fork). balancer 接口返回{ balancer, xrayEnabled, sticky },sticky 来自getStickyStatus()(比解析 balancer 文本更可靠)。 -
/api/xray/stickyPOST/DELETE 锁定/解锁接口 (fork). POSTduration参数(秒),0 表示永久(10 年);DELETE 手动解锁,会clearTimeout取消自动解锁 timer,避免重复触发。 -
getStageStatus扩展 Stage1/2/3 字段 (fork).packages/core/src/modules/plugin/xray/index.js的getStageStatus()从 6 个字段扩展为结构化输出:stage1(processStarted/livePort/apiPort/metricsPort/liveNodes/currentSelectTag)、stage2(enabled/intervalHours/lastSyncAt/lastSyncDurationMs/lastSyncFetchedCount/nextSyncAt/nextSyncOverdue)、stage3(isRunning/generation/roundStartedAt/nextRefreshAt/totalDue/processed/batchIndex/plannedBatchCount/successBatchCount/availableCount/explicitFailureCount/removedCount)。nextRefreshAt改为真实时间戳(替换原Date.now()+1占位符)。新增getLiveNodeFingerprints()方法返回tag→fingerprint反向映射,供/api/xray/nodes关联 country,不暴露到getStageStatus避免响应过大。 -
Stage2 同步耗时/抓取数持久化 (fork).
packages/core/src/modules/plugin/xray/cache.js新增setStage2LastSyncStats(cacheFilePath, durationMs, fetchedCount)和getStage2LastSyncStats(cacheFilePath),持久化到 SQLite cache meta(stage2_last_sync_duration_ms/stage2_last_sync_fetched_count)。index.jsStage2 同步开始时记录stage2SyncStartedAt,完成后写入耗时和subscriptionNodeCount。Stage2 下一轮预计时间 =lastSyncAt + intervalHours*3600*1000(超过当前时间显示"待触发",因 Stage2 在 Stage3 轮末按需触发,非定时调度)。 -
Stage3 轮次计时暴露 (fork).
index.js新增外层变量stage3RoundStartedAt/stage3NextRefreshAt/stage3Progress,在refreshCacheFromCacheOnly入口、批次成功、轮次结束 3 个nextDelay计算点更新。WebUI 实时显示本轮耗时和下一轮触发时间。 -
/api/xray/nodes关联 country/exitIp (fork).routes.js的/api/xray/nodes路由调getLiveNodeFingerprints()拿tag→fingerprint,再用readCacheEntriesByFingerprints查缓存,返回nodeMetadata: {tag: {country, exitIp, owner}},前端展示节点所属国家。 -
mitmproxy 子进程启用 SIGUSR2 堆快照诊断信号 (fork).
packages/core/src/modules/server/index.js的 forkexecArgv新增--heapsnapshot-signal=SIGUSR2和--diagnostic-dir=<userBasePath>/logs:长期运行的 mitmproxy 子进程可随时kill -USR2 <pid>生成.heapsnapshot(Chrome DevTools Memory 面板加载对比,三快照法定位内存泄漏),进程不会被杀死(Node 注册 handler 捕获信号写快照后恢复;未加 flag 的进程收 USR2 会按 POSIX 默认行为终止——运维时只对 mitmproxy PID 发)。平时零开销;触发时 stop-the-world(实测 19MB 堆 ~秒级暂停,代理功能暂停后自动恢复),快照文件几十至几百 MB 需定期清理。未设--report-signal(同信号会冲突);诊断报告可经运行时process.report.writeReport()获取,无需 flag。实测:发 USR2 后进程存活、快照落盘~/.dev-sidecar/logs/、代理恢复 HTTP 200。
Fixed
- 修复 WebUI
readBody解析失败导致 unhandled rejection 崩进程 (fork).packages/core/src/modules/plugin/webui/routes.js的readBody对非法 JSON body reject,而 5 处调用点都在路由的 try 之外 await 它——非法 JSON 请求(如PUT /api/configbody 为not json)使 server 端 async handler 未捕获 reject,客户端永远等不到响应(fetch 挂起 → mocha 用例超时 + after-all 钩子挂起),CI 三平台test packages/core全部挂死在同一用例。修复:readBody解析失败改为 resolve(null),由路由已有的 body 类型校验(isPlainObject/Array.isArray)统一拦截返回 400INVALID_BODY;POST /api/xray/sticky补同款校验(原实现 body 为 null 时body.duration访问会 TypeError 崩进程)。另修复configApiSave.test.js的进程级HOME污染:该文件为隔离配置把process.env.HOME指向临时目录,但 mocha 单进程顺序执行使后续依赖真实 HOME 的测试(versionTest 等发真实网络请求的用例)读到已删除的目录而失败——after 钩子现恢复原 HOME 并configApi.reload()单例,临时目录保留(log appenders 仍指向它,删除会 ENOENT);core 测试脚本加--timeout 10000(CI 冷环境下首次require expose初始化超 mocha 默认 2s)。测试基建(CI 全量挂死的其余根因):(1)xrayGenConfig.test.js的"零节点无 balancer"断言是 v2.2.7 固定模板行为变更前的过时期望,更新为"零节点也总是生成 balancer";(2)webui.test.js是集成式测试(14 处路由 requireexpose加载完整 app),与同进程其他测试文件交互引发 worker 异常退出(mocha parallel 表现为 Workerpool Worker terminated,单进程表现为文件交接处 exit 1 无汇总)——reInjectXrayRules改为支持context.xrayApi注入(xrayApiOverride !== undefined即跳过全局 require,测试 fake context 传xrayApi: null使写用例不再触达真实 expose),webui.test.js 从全量排除(.mocharc.jsonignore)并新增test:webuiscript 独立进程运行(--no-config绕过 ignore);(3) core 测试切换 mocha--parallel --jobs 2(每文件独立 worker 进程,隔离进程级副作用)。全量pnpm test131 passing、pnpm run test:webui50 passing。mitmproxy 测试同治(第三次 CI 失败定位):(1)dnsLookupTest.mjs/dnsTest*.mjs是上游的真实 DNS 网络测试(DoH 到 quad9 等),GitHub runner 被 quad9 返回 HTML 页导致解析失败——.mocharc.json的 spec 限定test/*.js排除 .mjs(保留本地手动node test/dnsLookupTest.mjs运行);(2)wwwAuthenticateTest.js硬编码/home/uif79392/.dev-sidecar绝对路径的 CA 证书——CI runner 必然挂——改为os.homedir()推导(支持DS_CA_CERT/DS_CA_KEY环境变量覆盖),CA 不存在时describe.skip整套件(CI 上 14 passing + 3 pending,本地有 CA 全跑)。(3)tlsUtilsAkiTest.js的"真实落盘 CA"用例存在两处同型地雷:/home/uif79392/硬编码路径 + 箭头函数里this.skip()(this 非 mocha context,CI 无 CA 走该分支时this.skip is not a functionTypeError 崩掉)——改为同款 homedir 推导 + 用例级条件it.skip。 - 修复 Stage3 轮末 Stage2 周期触发是死代码,长期运行服务订阅永不刷新 (fork).
refreshCacheFromCacheOnly入口置isStageRunning = true(v2.2.7 修复"未设置 isStageRunning"时引入),但"Stage3 后触发 Stage2"的轮末检查用的是!isStageRunning守卫——该轮末段运行在refreshCacheFromCacheOnly体内,isStageRunning恒为true,守卫恒 false,周期性 Stage2 订阅刷新从未执行过(全部历史日志 0 次出现该触发)。影响:dev-sidecar 长期不重启时,subscriptionSyncIntervalHours=24h的订阅重抓永不发生,免费订阅节点逐渐失效、节点池枯竭;此前靠服务频繁重启(每次启动 start() 触发一次 Stage2)掩盖。修复:守卫改为!isStage2Running(防与启动时后台 Stage2 并发——isStage2Running是"订阅抓取进行中"的精确信号),generation守卫保留。触发块内的冷却判断(shouldSkipRemoteFetchDueToCooldown)已有双保险,无误触发风险。 - 恢复 Stage1 bootstrap 探测 (fork). v2.2.7 commit 92601af 误删了 Stage1 bootstrap 探测逻辑(
readCacheEntriesForStartup→probeNodesBatch→annotateProbeEntries→ 按 delay/country/owner 过滤 → 排序 → 切片startupNodeLimit),导致启动后无节点注入。已重新实现:bootstrap 探测成功后通过ado注入startupNodeLimit个节点(默认 20)到 live xray 进程。 - 修复
refreshCacheFromCacheOnly未设置isStageRunning(fork).refreshCacheFromCacheOnly函数入口未设置isStageRunning = true,导致 Stage3 运行时 Dashboard 显示"空闲"。修复:入口处设true,finally块设false。 - 修复 WebUI Xray 面板 enabled 读取错误 (fork). 前端从
stage/status读取enabled字段(不存在),导致 Xray 状态显示"关"。修复:改为从status.plugin.xray.enabled读取,支持 3 态显示(关/启动中/开)。 - 修复
xray lso返回 JSON 字符串未解析 (fork)./api/xray/nodes接口直接JSON.stringify了xray lso的 st...
v2.2.7
Fixed
- 修复 Xray 主进程 stop() 不等 close 事件导致重启时端口冲突 (fork).
packages/core/src/modules/plugin/xray/process.js的stop()调用child.kill()(SIGTERM)后立即child = null返回,不等close事件。systemctl restart时主进程在旧 xray 释放端口 10801 前就退出,新实例启动报"端口 10801 被占用 (Strict Mode)"。更严重的是 xray 被移到隔离 cgroup(dev-sidecar-xray-probe.scope),systemdKillMode=control-group杀不到它,旧 xray 成孤儿进程持续占端口。修复:stop()改为await等close事件(3 秒超时 SIGKILL 兜底),确保进程退出、端口释放后才返回。 - 修复 Xray bootstrap probe 全失败时复用旧节点 (fork).
packages/core/src/modules/plugin/xray/index.js在 bootstrap probe 返回 0 个可用节点时,旧逻辑有两层 fallback:(1) 复用上次config.json的节点,(2) 复用缓存数据库里stable=true但未重新 probe 的节点。两者的节点来源都是缓存数据库,既然数据库已确认无可用节点,这些节点大概率也已失效。继续用只会填充 config.json 全是 dead 节点,误导 observatory 反复探测已知失效节点,浪费探测周期。修复:两层 fallback 全部移除,probe 全失败时 config.json 只包含手动预置节点(cfg.nodes),无预置节点则只有 Direct/Block。Stage3 探测出新节点后通过热刷新(API ado/rmo)动态注入 config.json。
Changed
- 移除
config.json.bak备份机制 (fork).config.json.bak原用于 Stage2 读取"上次启动时的节点快照"作为订阅去重比对的来源。随着 Phase 2 热刷新(config.json 不再被 Stage3 同步写回)和 bootstrap 不再 fallback 到旧 config.json 节点,备份已无用途。Stage2 现在直接读config.json。移除了backupFileIfExists()辅助函数、liveConfigBakPath变量/参数、config.json.bak创建逻辑、configSourcePathfallback 逻辑。
v2.2.6
Added
- Xray Stage3 常驻探测子进程 (fork). Stage3 一轮探测的所有批次共用一个常驻 xray 探测子进程,批次间用
xray api ado/rmo动态换节点(固定 tagproxy_0~proxy_127复用),不再每批 spawn 一次性子进程。40000 节点场景下 spawn 次数从 ~313 降到 1(-99.7%)。packages/core/src/modules/plugin/xray/probe.js的isObservationReady/waitForObservatoryMetrics新增expectedTags+minLastTryTime参数,按 tag 集合过滤并区分新旧探测结果(rmo 后 observatory status 永久残留)。packages/core/src/modules/plugin/xray/xray_api.js的addOutbounds改为解析 stdout 返回逐节点结果(ado 遇无效节点会停止后续处理),removeOutbounds改为并行(128 tags 从 2.5s 降到 ~0.6s)。Stage1 bootstrap 仍用一次性 spawn(只探测一批)。 - Xray Stage3 常驻出口探测子进程 (fork). 活节点的出口 IP/country/owner 探测共用一个常驻 xray 子进程(固定 tag
egress_0,通过 ado/rmo 切换节点),不再每节点 spawn 一次性子进程。123 活节点场景下 spawn 次数从 ~123 降到 1(-99.2%),waitForProxyPortReady仅首次调用(端口固定)。resolveEntryEgressMetadata新增egressController参数,annotateProbeEntries在有 controller 时并发降为 1(单进程串行)。Stage1 bootstrap 不传 controller,走 legacy 路径。总 spawn 次数(批次+出口)从 ~436 降到 ~2(-99.5%)。swapNode首次调用时等待 API 端口就绪(10s 超时),失败后回退到一次性 spawn。 - Xray sticky balancer 锁定 (fork). 新增
enableSticky({duration})/disableSticky()/getStickyStatus()API,通过xray api bo(Balancer Override)锁定出口 IP,防止 ChatGPT 注册等场景的ERR_NETWORK_CHANGED。锁定后所有新连接走同一节点,duration秒后自动解锁。xray_api.js新增overrideBalancer/removeBalancerOverride/getBalancerInfo。sticky 操作通过stickyOpChain串行化防止竞态;热刷新 rmo 删除锁定节点时自动解除;xray 重启时重置 sticky 状态。
Fixed
- 修复 Xray 常驻出口探测 API 端口不可达 (fork).
startEgressProbeProcessspawn xray 后立即返回,但 gRPC API 端口需要时间初始化。第一个swapNode调用addOutbounds时 API 还没就绪,导致failed to dial 127.0.0.1:<apiPort>(生产环境 8月14日 170 次、8月15日 419 次失败)。修复:swapNode首次调用时通过waitForProxyPortReady等 API 端口就绪(10s 超时),检查失败后标记portCheckFailed不再重试,resolveEntryEgressMetadata捕获swapNode失败后回退到一次性 spawn。
Changed
- 移除 Xray balancer 的
fallbackTag: 'direct'直连回退 (fork). 所有代理节点不可用时,请求直接失败而非走直连,避免暴露真实 IP。packages/core/src/modules/plugin/xray/gen_config.js不再生成fallbackTag字段。 subscriptionSyncIntervalDays改为subscriptionSyncIntervalHours(fork). 单位从天改为小时,默认 24 小时(原 3 天),最小 1 小时(原 1 天)。packages/core/src/modules/plugin/xray/config.js和index.js同步更新。cacheRefreshInterval改为cacheRefreshIntervalHours(fork). 单位从秒改为小时,默认 6 小时(原 21600 秒),最小 1 小时(原 3 小时)。packages/core/src/modules/plugin/xray/config.js和index.js同步更新。
v2.2.5
Synced upstream docmirror/dev-sidecar master (29 commits). Upstream introduced a CLI rewrite (native ds-cli binary with SEA packaging), single-instance mutex, and Linux/macOS environment-variable proxy support. All fork-specific changes (CA cert passthrough, Xray plugin, keep-alive socket fix, log.debug hot-path demotion, conditional linuxTargets, native module rebuild, CSS variable theme) were preserved. The fork's desktop-detection logic in set-system-proxy was merged with upstream's env-var support so that headless servers now write proxy env vars (previously the early return true skipped them).
Added
- CLI rewritten as native
ds-clibinary (upstream). The CLI (packages/cli/) was rewritten from a plain Node script into a native command-line toolds-cliwith Sea-of-Nodejs (SEA) packaging support. Newpackages/cli/scripts/build.jsperforms incremental builds with SHA256 checksums and parallel downloads, auto-cleans stale build products while preserving thenode-bincache, dynamically fetches the Node.js supported platform list, and displays OS type/version/arch. New.github/workflows/build-cli.ymlCI workflow builds the CLI binary. Newpackages/cli/README.mddocuments SEA packaging and cross-compilation.packages/cli/src/sea-entry.jsis the SEA entry point;packages/cli/.gitignoreignores build products exceptsea-config.json. - CLI new commands (upstream):
proxy on/off(fork worker sets system proxy immediately),plugin start/stop(persists toconfig.json, takes effect on restart),service install/uninstall(boot autostart),help(command formds-cli helponly,--help/-hremoved),status(shows running instance and plugin state),start/stop/restart. Workers split intoproxy-worker.js,plugin-worker.js,free-eye-worker.js.packages/cli/src/user_config.json5removed (no longer shipped as a static template). - CLI/GUI single-instance mutex via
proper-lockfilelong-lock (packages/core/src/modules/instance/index.js, 140 lines).packages/core/src/expose.jsexposesapi.instance. GUIbackground.jsacquires the lock onapp.whenReady()and writes instance info (type: 'gui', pid, command, startTime) torunning.json; if another instance holds the lock, the GUI quits with an error. CLI does the same onstart/proxy on/plugin start.packages/core/test/instanceTest.js(153 lines) covers acquire/release/write/cleanup. - Linux/macOS environment-variable proxy support (upstream).
packages/core/src/shell/scripts/set-system-proxy/index.jsgainedwriteProxyEnvFile/addProxyEnvToShellProfile/removeProxyEnvFromShellProfilehelpers and asetEnvparam. WhensetEnvis true, the proxy env vars (http_proxy,https_proxy,no_proxy) are written to a file and appended to the detected shell profile (~/.bashrc/~/.zshrc), independent ofgsettings— so CLI tools (curl, git, npm) on headless servers pick up the proxy even without a desktop. macOS uses the same env-var file mechanism alongsidenetworksetup. - Plugin status events (upstream).
overwallandpipplugins now fireevent.fire('status', ...)on start/close and log开启/关闭【X】代理成功.overwallgained astatus: { enabled: false }block.packages/core/src/modules/server/index.jsemitsstatus.server.enabledevents.packages/gui/src/view/pages/proxy.vuereads the new status. - CLI test suite (upstream). 60+ new Mocha test cases:
packages/cli/test/{gui,index,plugin,proxy,service,start,status}.test.jscovering proxy on/off, config persistence, shell detection, service/help/version/unknown commands. - Xray Stage3 零中断热刷新(Phase 2) (fork). Stage3 后台热刷新改为基于 Xray HandlerService gRPC API 的动态增删 outbound,不再重写 config.json + 重启 xray 进程,现有连接完全不中断。新增
packages/core/src/modules/plugin/xray/xray_api.js模块,封装xray api ado(AddOutbound,stdin 传 JSON)/xray api rmo(RemoveOutbound,按 tag)/xray api lso(ListOutbounds)三个 CLI 子进程调用(execFile+ 5s 超时)。packages/core/src/modules/plugin/xray/gen_config.js新增apiPort参数和api块生成(listen: 127.0.0.1:<apiPort>,services: [HandlerService, ObservatoryService, RoutingService]),并新增 7 个单元测试(packages/core/test/xrayGenConfig.test.js)。balancer.selector和observatory.subjectSelector从显式 tag 列表["proxy_0",...]改为前缀["proxy_"],通过 Xraystrings.HasPrefix自动包含运行时动态新增的proxy_Ntag。RemoveHandler只从 manager 的 map 删除 tag 引用、不调handler.Close(),已建立连接继续完成、新连接不再路由到该 tag;Observatorybackground()每个探测周期自动发现新 outbound 并探测。packages/core/src/modules/plugin/xray/index.js路径 B 新增 API 热刷新逻辑:维护currentLiveNodeTags(Map: fingerprint→tag)和nextProxyTagIndex计数器,先addOutbounds新节点(observatory 下个周期探测后才可选,正好支持"先 Add 后 Remove"策略),再removeOutbounds坏节点;API 调用失败时回滚 in-memory 状态并 fallback 到重启路径(Phase 1 行为)。兼容 Xray-core v26.3.27+(v26.3.27 observatory 不清理已移除 outbound 状态但无害)。73 个 core 测试全部通过无回归。 - Xray 主进程启用 metrics 端口供运行时调试 (fork). 主进程(live xray)的 3 个
genConfig调用点(冷启动 regen / fallback 重启 / 初始启动)新增metricsPort参数,生成metrics块(expvar/debug/vars)。运维人员可通过curl -s http://127.0.0.1:<metricsPort>/debug/vars | jq '.observatory'查看主进程各节点的 alive/delay/lastErrorReason,无需依赖上游未合并的xray api obs命令。metricsPort和apiPort通过event.fire('status', ...)写入running.json的app.status.plugin.xray,方便随时查看。packages/core/src/modules/instance/index.js的watchStatusEvents过滤逻辑从仅同步*.enabled扩展为同时同步plugin.xray.port/apiPort/metricsPort三个端口字段。复用已有 config 时从metrics.listen提取端口(与api.listen同逻辑)。 - 新增
doc/xray-devops.md(fork). Xray 插件开发与运维文档,面向调试运行时状态或排查 Stage3 热刷新行为的开发者/运维人员。包含:Stage3 热刷新机制(工作原理 + 日志关键词解读)、运行时调试方式(curl/debug/vars查看 observatory 节点延时 [推荐] /xray api lso/obs/bi子命令)、Xray-core v26.3.27 vs v26.7.28 版本兼容性对比表、部署后首次重启注意事项。
Changed
- Xray Stage3 热刷新路径 B 改为 API 动态增删 (fork).
packages/core/src/modules/plugin/xray/index.js的maybeRegenerateLiveConfigFromCache路径 B(热刷新,config.json 已有节点)不再调processApi.restart(stop SIGTERM + 200ms + start,中断流量),改为通过 HandlerService gRPC API 动态增删 outbound。仅在冷启动(路径 A,config.json 无节点)或 API 调用失败 fallback 时才重启 xray 进程。运行时增删纯走 API,config.json 定位保持"上次启动时的快照"语义,不同步写回(避免写文件并发问题,Stage1 冷启动筛选本身可靠不依赖 config.json 实时性)。 - 简化 Stage3 重启判断条件(Phase 1) (fork).
packages/core/src/modules/plugin/xray/index.js的短路条件从keptNodes.length === currentConfigNodes.length && keptNodes.length >= startupNodeLimit改为keptNodes.length >= startupNodeLimit。旧条件要求"可用节点数达标且节点集合未变化"才跳过重启,导致 5 个节点全可用但startupNodeLimit=10时因数量变化触发无谓重启。新条件只要可用节点数达标即跳过,节点集合的变化由 Phase 2 的 API 动态增删处理,无需重启。 - Stage1 候选节点 SQL 层提前过滤 country/owner (fork).
packages/core/src/modules/plugin/xray/index.js的 5 处buildCacheEntryQueryOptions调用(启动主路径/冷启动 regen/热刷新)新增allowedCountries和allowedOwners参数,让 SQLWHERE子句直接过滤country IN (...)和owner NOT LIKE '%cloudflare%',而不是在 JS 层collectBootstrapCandidateEntries事后过滤。之前 LIMIT 100 取出的节点可能大部分不符合 country/owner 条件(如前 100 个低延迟节点大多是 CN/JP),导致 probe 后筛出的可用节点不足startupNodeLimit。修复后 LIMIT 100 取的全是符合 country/owner 条件的节点,probe 后能筛出更多符合maxDelayMs的节点。bootstrapCandidateLimit默认值从 31 提高到 100。 - Stage1 快速复检改用 observatoryProbeUrl 探测 (fork).
packages/core/src/modules/plugin/xray/index.js的probeNodesBatch新增probeUrl参数,Stage1 bootstrap 传cfg.observatoryProbeUrl(严格,如 chatgpt.com),Stage3 周期探测传cfg.probeUrl(宽松,gstatic 204)。之前 Stage1 用probeUrl(宽松)探测,导致通过 gstatic 筛选的节点在主进程 observatory 用observatoryProbeUrl(严格)重新探测时大量 dead(delay=99999999)。修复后 Stage1 和主进程 observatory 用同一探测目标,确保进 config.json 的节点都能通过 observatory 探测。 - Stage1 不再用未 probe 的 stable 节点凑数 (fork).
packages/core/src/modules/plugin/xray/index.js启动节点选择不再将supportedFallbackEntries(只检查格式、未 probe 的 stable 节点)与bootstrapSelectedEntries(probe 验证过的节点)合并凑满startupNodeLimit。现在只用 probe 验证过的节点,仅在 probe 完全失败(0 个节点)时才 fallback 到 stable 节点兜底,避免未验证的 delay=0 节点进入 config.json 导致主进程 observatory 标记为 dead。 - 移除 Stage3/cache 和 bootstrap 探测的人为超时 (fork).
packages/core/src/modules/plugin/xray/index.js的两处probeNodesBatch调用(Stage1 bootstrap + Stage3 cache 周期探测)从传递cacheBatchTimeout/bootstrapBatchTimeout秒数改为直接传timeoutMs: 0。packages/core/src/modules/plugin/xray/probe.js的waitForObservatoryMetrics在timeoutMs <= 0时使用deadline = Infinity,即不设超时上限,让 observatory 自然收集完所有 sample 再返回;探测进程崩溃通过child.exitCode检查捕获。之前cacheBatchTimeout: 120(2 分钟)和 bootstrap 超时会在 v26.3.27 observatory 慢速收集时提前中断探测,导致 "Observatory metrics have not collected N samples yet" 错误和节点遗漏。同时删除了不再需要的getCacheBatchTimeoutSeconds/getBootstrapBatchTimeoutSeconds辅助函数,并从packages/core/src/modules/plugin/xray/config.js移除cacheBatchTimeout配置项。 - 探测从 burst observatory 切换到 regular observatory + 并发探测 (fork).
packages/core/src/modules/plugin/xray/index.js的runSingleProbePass将probeMode从'burst'改为'observatory'。v26.3.27 的 burst observatory 是串行探测的——每个 alive 节点需要timeout × sampling秒(15s × 2 = 30s),128 个 alive 节点需要 64 分钟;之前之所以快是因为大部分节点 dead(连接拒绝毫秒级完成),SQL 过滤后节点质量提升、alive 比例增加后速度暴跌。regular observatory 配合enableConcurrency: true并发探测所有节点,128 节点仅需 5-6 秒。实测 Stage3 第一轮 14332 节点(112 批)从预估 93+ 小时降至 ~11 分钟完成。同时修复了buildCacheEntriesFromObservatory(packages/core/src/modules/plugin/xray/cache.js)不处理 regular observatory 格式的问题——regular observatory 的节点状态是{alive, delay, outbound_tag}(无HealthPing字段),旧代码只处理HealthPing路径、跳过无HealthPing的状态,导致 regular observatory 的 alive 节点全部被丢弃;新增 fallback 路径直接从alive/delay字段构建 cache entry。probe.js的isObservationReady也同步更新:当检测到无HealthPing字段时(regular observatory),改用delay > 0判断所有节点是否已被探测(dead 节点的 delay=99999999 也满足此条件),而非等待healthPing.all >= expectedSamples。 set-system-proxycomplementary merge (fork + upstream). The fork's desktop-detection logic (skipgsettingswhen/usr/bin/gsettingsis absent or the X server socket/tmp/.X11-unix/X*does not exist, to avoid spawning dbus-launch + dbus-daemon + dconf-service on headless servers) was refactored from an earlyreturn trueinto ahasDesktopflag. U...
v2.2.4
Fixed
- Fixed
tunnel://127.0.0.1:0(port 0 placeholder for xray) always failing withECONNREFUSED 127.0.0.1even aftersetting.xrayPortwas set. Root cause:getTunnelAgentinpackages/mitmproxy/src/lib/proxy/common/util.jsused global single-variable caches (httpsOverHttpAgent,httpsOverHttpsAgent,httpOverHttpsAgent) — the first call (before xray started, port still 0) cached a dead agent that was reused forever, ignoring the port 0 → xrayPort replacement done bydoProxy. Replaced with a Map cache keyed by${agentType}:${hostname}:${port}, so port 0 and port 10801 get separate agents. Verified:tunnel://127.0.0.1:0now works correctly, port is auto-replaced bysetting.xrayPortat runtime. - Fixed mitmproxy child process crashing with
SIGABRT(V8 heap OOM) roughly 10 minutes after a machine reboot under high-concurrency HTTPS browsing (e.g. navigatinggithub.comwith many simultaneous image/asset sub-requests). Root cause: 22log.infocall sites in the mitmproxy package printed full per-request data on the hot path — the two biggest offenders werecreateRequestHandler.js(logged the complete request headers JSON, including ~1 KB GitHub_gh_sess/user_sessioncookies) andutil.match.js(logged the fullinterceptOptsJSON for every match, with 3–6 KB of tampermonkey/cache/rule config per request). Under high concurrency these strings (plus log4js's 50 MBmaxLogSizefile buffer) accumulated in the V8 heap faster than GC could reclaim them, exhausting the 96 MB--max-old-space-sizecap set on the mitmproxy child process. The previous v2.2.3safeROptionsForLogfix only covered the error path (stringify2fallback toutil.inspect); this fix covers the normal INFO path. Demoted 57 hot-pathlog.infocalls tolog.debugacross 22 files inpackages/mitmproxy/src/:lib/proxy/mitmproxy/{createRequestHandler,createConnectHandler,createUpgradeHandler,dnsLookup}.js,lib/dns/base.js,lib/proxy/common/util.js,lib/proxy/middleware/overwall.js,lib/interceptor/impl/req/{OPTIONS,abort,cacheRequest,proxy,redirect,requestReplace,sni,success,unVerifySsl}.js,lib/interceptor/impl/res/{AfterOPTIONSHeaders,cacheResponse,responseReplace,script}.js,utils/util.match.js,options.js. Codebase knowledge-graph trace ofrequestHandleroutbound callees was used to verify full hot-path coverage with no omissions (and no over-demotion of startup-once / per-domain-once logs). Verified by codebase-memo MCP. Verified by local deploy: under the same browsing pattern that previously crashed in ~10 min, the new build ran with no SIGABRT.
Added
- Added
observatoryProbeUrlconfig option to the Xray plugin (packages/core/src/modules/plugin/xray/config.js). When set, the Stage1/observatory runtime probe uses this URL instead ofprobeUrl(Stage3 cache probe remains unchanged). This allows using a stricter probe target (e.g.https://chatgpt.com/) for runtime node selection while keeping the lenient probe (gstatic.com/generate_204) for cache-wide screening. xray observatory treats any HTTP response (including 403) as "alive", only timeout/reset marks a node as dead — so a CF-protected site like chatgpt.com filters out junk HTTP proxies that pass simple 204 probes but fail on real targets. Default: empty (falls back toprobeUrl). ThreegenConfig()call sites inpackages/core/src/modules/plugin/xray/index.jswere updated to passcfg.observatoryProbeUrl || cfg.probeUrl.
Changed
- Added mitmproxy child-process auto-respawn in
packages/core/src/modules/server/index.js. Previously, when the mitmproxy child process crashed (SIGABRT/SIGSEGV/non-zero exit), the mainservice-entry.jsprocess stayed alive but the proxy port 31181 was dead — every browser request gotECONNREFUSEDwhilesystemctl status dev-sidecarreportedactive (running)(systemd'sRestart=on-failureonly watches the main PID, not grandchild processes forked viachild_process.fork()). TheserverProcess.on('exit')handler now detects abnormal exits (code !== 0 or signal !== null) and re-forks the mitmproxy child process by callingserverApi.start(...). A 30-second sliding window limits respawn to at most 3 attempts; if exceeded, the handler fires anerrorevent (value: 'respawn_exceeded') and stops retrying to avoid restart storms. Thekill()/close()/restart()paths set anintentionalStopflag so that graceful shutdowns don't trigger self-healing. Verified by local deploy: manuallykill -SIGABRTon the mitmproxy PID triggered respawn within ~23 ms with zero browser-visible downtime (curl -x http://127.0.0.1:31181 https://github.comreturned 200 before and after the kill). - Removed dead
maxOldSpaceSizeMBfield fromSTAGE3_BATCH_LEVEL_TABLEinpackages/core/src/modules/plugin/xray/config.js. This field was a leftover from v2.2.0's per-level dynamic--max-old-space-sizeadjustment, which v2.2.3 replaced with a fixed 96MB inpackages/core/src/modules/server/index.js(and removedSTAGE3_MAX_OLD_SPACE_BY_LEVEL). ThemaxOldSpaceSizeMBfield was never read by any code after v2.2.3 — onlybatchSizeandstage3GcThresholdMBare used. Also fixed stale comments:level=N 对应 batchSize=N*64→batchSize = 64 << (level-1); removed the misleadingmitmproxy 子进程与 Stage3 探测共用同一 fork 路径note (xray probe is a Go binary with no V8 heap). Verified by codebase-memo MCP knowledge-graph trace: confirmedmaxOldSpaceSizeMBhas zero references outsideconfig.js, andstage3GcThresholdMBis read byrefreshCacheFromCacheOnlyinindex.js.
v2.2.3
Fixed
- Fixed mitmproxy child process crashing with
SIGABRT(V8 heap OOM) on corporate networks after running for ~1-2 hours under high-concurrency HTTPS traffic. Root cause: 7log.errorcall sites inpackages/mitmproxy/src/lib/proxy/mitmproxy/createRequestHandler.jspassed the fullrOptionsobject tojsonApi.stringify2(rOptions). BecauserOptionscontains theHttpsAgentinstance (circular references),JSON.stringifythrows, andstringify2's catch fallback (return obj) hands the raw object back to log4js, which thenutil.inspect-expands the entire agent — including everyHttpsAgent.socketskey. On corporate networks each socket key ishost:port::<full 141-cert PEM>(becauseloadExtraCaCertsmergestls.rootCertificates+ the corporate root CA into thecaarray passed to every agent), so a single error log produces a multi-megabyte string. Under high-concurrency error bursts (e.g. Xray-tunneledchatgpt.comrequests failing withEPROTO), dozens of these dumps accumulate in the V8 heap faster than GC can reclaim them, exhausting the 96 MB old-space cap and aborting the process. Added asafeROptionsForLog(rOptions)helper that extracts only scalar fields (protocol,method,hostname,port,path,servername,headers, etc.) and stripsagent/socket/ca; all 7 call sites now logjsonApi.stringify2(safeROptionsForLog(rOptions)), which serializes cleanly to a small JSON string with no fallback to the raw object. This keepsMemoryHighat 512 MB (service template) / ≤300 MB (user constraint) viable without raising the V8 heap cap.
Changed
- Disabled chromium zygote processes in service mode by adding
--no-zygoteswitch inpackages/gui/src/background.jswhenDEV_SIDECAR_SERVICE_MODE=true. In service mode, noBrowserWindowis created, so there are no renderer processes; the 5 zygote processes (chromium's pre-fork templates for renderers) are pure overhead, consuming ~39MB RSS and contributing shared file cache pages to the cgroup. This is set beforeapp.whenReady(). - Closed inherited chromium resource file descriptors in the mitmproxy child process entry point (
packages/mitmproxy/src/index.js). The mitmproxy child is forked from the Electron main process viachild_process.fork(), which inherits all non-CLOEXEC file descriptors opened by Electron — includingchrome_100_percent.pak,chrome_200_percent.pak,resources.pak,locales/zh-CN.pak,icudtl.dat,v8_context_snapshot.bin, and/dev/shm/.org.chromium.Chromium.*shared memory segments. These are useless to the mitmproxy process but keeping them open prevents the kernel from reclaiming the corresponding file cache pages under cgroupMemoryHighpressure. On Linux, the mitmproxy entry now walks/proc/self/fdandcloseSync()s any fd pointing to chromium resources, pak files, icudtl, v8 snapshots, or/dev/shm/— reducing cold-boot file cache pressure. - Moved Xray probe processes into an isolated cgroup (
/sys/fs/cgroup/system.slice/dev-sidecar-xray-probe.scope) on Linux so their file cache does NOT count against the dev-sidecar service'sMemoryHighlimit. On cold boot (after a full machine restart), the system page cache is empty; Xray probe processes read an 800MB+ SQLite cache (nodes_cache.sqlite), pulling ~137MB of file pages into the cgroup. Combined with the service's anon memory, this pushesmemory.currentto ~230MB against a 280MBMemoryHigh, triggering 3000+ kernelhighreclaim events. On warm boot these SQLite pages are already in the system page cache (not charged to any cgroup), which is why warm-boot memory is low. The newmoveProcessToIsolatedCgroup()inpackages/core/src/modules/plugin/xray/util.cgroup.jscreates a sibling cgroup withmemory.high=max(no limit) and moves the probe PID into it immediately afterspawn(). The probe's file cache is then charged to the isolated cgroup, keeping the service cgroup at warm-boot levels (~158MB) even on cold boot. Axray-probe-cgroup.shhelper script (installed bypackages/gui/pkg/linux/postinst) encapsulates the cgroup operations, and a new sudoers rule allows the service user to run it viasudo -nwithout a password. - Added
--no-sandboxalongside--no-zygotein the systemd serviceExecStart. Electron requires sandbox to be disabled when zygote is disabled; without this, the service crashes withZygote cannot be disabled if sandbox is enabled. - Increased
reclaimStartupMemoryreclaim amount from 100M to 200M inpackages/core/src/expose.js. On cold boot (after a full machine restart, not just a service restart), the system page cache is empty, so the first load of the Electron binary (~150MB),app.asar, chromium runtime, Xray binary, and SQLite cache files produces ~200MB of cgroup file cache. The old 100M reclaim only covered half of this, allowing the cold-boot memory peak to reach theMemoryHighlimit (280M on this deployment), triggering 110 kernelhighreclaim events. The new 200M reclaim drops file cache to near zero before the mitmproxy child process is forked, leaving headroom for subsequent Xray Stage3 probing. Addedlog.info/log.warnconfirmation logging (previously the function silently returned true/false, making it impossible to verify from logs whether reclaim executed). - Raised V8 old-space cap and stage3 GC threshold for level 1 and level 2 in
STAGE3_BATCH_LEVEL_TABLE(packages/core/src/modules/plugin/xray/config.js) and the mirroredSTAGE3_MAX_OLD_SPACE_BY_LEVELinpackages/core/src/modules/server/index.js:- level 1:
maxOldSpaceSizeMB48→64,stage3GcThresholdMB32→44 - level 2:
maxOldSpaceSizeMB80→96,stage3GcThresholdMB56→68 - level 3-5: unchanged
- level 1:
- The old level 2 values (80MB/56MB) were tuned for v2.1.x (HTTP/1.1 fake server). v2.2.0 introduced HTTP/2 fake servers (
http2.createSecureServer), which increase per-request V8 heap pressure (Http2Session/Http2Stream JS wrappers). The old 80MB cap caused intermittent SIGABRT (V8 heap OOM) during Stage3 batch probing. The new 96MB cap gives 28MB GC buffer (was 24MB), and the GC threshold 68MB triggers earlier (was 56MB), reducing the chance of a GC-miss OOM. - Added
LimitCORE=0to the systemd service template (packages/gui/pkg/linux/dev-sidecar.service) to prevent 100GB+ Linux core dumps on SIGABRT. WSL2 crash dumps are controlled separately by.wslconfigMaxCrashDumpCount=-1. - Replaced Electron service-mode entry with a pure-Node entry point (
service-entry.cjs) that runs withELECTRON_RUN_AS_NODE=1, completely eliminating all chromium subprocess overhead in service mode. Previously, even in service mode (noBrowserWindow), the Electron main process still spawned chromium infrastructure: a GPU process, a NetworkService utility process, and (if not disabled) zygote processes. These contributed ~20MB RSS and ~50MB shared file cache (pak files, icudtl, v8 snapshots) to the dev-sidecar cgroup. The newservice-entry.cjsis compiled by webpack into a standaloneservice-entry.jsbundle at the asar root (alongsidemitmproxy.js), with@docmirror/dev-sidecarand all its JavaScript deps inlined — only native modules (better-sqlite3, fadvise-linux, sysproxy) are loaded from the asar'snode_modulesat runtime. The systemd serviceExecStartnow runs/opt/dev-sidecar/@docmirrordev-sidecar-gui /opt/dev-sidecar/resources/app.asar/service-entry.jswithEnvironment=ELECTRON_RUN_AS_NODE=1, which makes the Electron binary behave as a pure Node.js runtime (no chromium initialization at all). This drops the service cgroup memory from ~130MB (warm) / ~230MB (cold) to ~102MB, with only 3 processes total: the Node main process, the mitmproxy fork, and the Xray probe. The--no-zygoteand--no-sandboxswitches are no longer needed (there is no chromium to disable). ThereclaimStartupMemorycgroup memory reclaim is also no longer needed in pure-Node mode (no Electron binary page cache to reclaim), but is retained as a safety net. - Added multi-layer cgroup memory reclaim during startup in
packages/core/src/expose.jsandpackages/core/src/modules/plugin/xray/index.js. On cold boot, the mitmproxy fork + gsettings/D-Bus + SQLite cache reads produce ~180MB of cgroup file cache that pushesmemory.peakto ~282MB (above the 280MMemoryHighlimit, triggering 260+ kernelhighreclaim events). The new reclaim points executememory.reclaim(viasudo -n /usr/lib/dev-sidecar/reclaim-memory.shwith NOPASSWD sudoers) at three stages: (1) beforeserver.start()(reclaimStartupMemory, 200M), (2) afterproxy.start()(dynamic 100-350M based onmemory.current), (3) before SQLite cache reads in the Xray startup precheck (dynamic 100-300M). Each reclaim is followed bydropSqliteFileCache(POSIX_FADV_DONTNEED) to hint the kernel. Warm-boot peak dropped from ~280MB to ~245MB; cold-boot peak remains ~282MB (physical limit: cold-boot file cache cannot be fully reclaimed because subsequent reads immediately re-fault the pages, confirmed byworkingset_refault_filecounter). - Fixed
stage3-initial-countmemory.reclaimsilently failing withEACCESinpackages/core/src/modules/plugin/xray/index.js. Thememory.reclaimcgroup file is--w------- root root, so directfs.writeFileSyncby the service user was silently caught and swallowed. Changed to usexrayCache.reclaimCgroupMemory()which has asudo -n reclaim-memory.shfallback (NOPASSWD sudoers rule). This was the root cause of the 280MB cold-boot peak — reclaim never actually executed. - Fixed Stage3 cache refresh hanging on corporate networks with SSL interception in
packages/core/src/modules/plugin/xray/network_guard.js. TherequestLocalNetworkCanaryfunction usedhttps.getwith default certificate verification, which fails withUNABLE_TO_GET_ISSUER_CERT_LOCALLYon corporate networks because Node.js's built-in CA store does not include the corporate re-signing CA. This caused Stage3 to loop on "检测到本地网...
v2.2.2
Fixed
- Restored
tunnel:→http:protocol conversion inpackages/mitmproxy/src/lib/proxy/common/util.js(getTunnelAgent) that was lost when syncing upstream v2.2.0. The v2.1.6 code had an explicitif (protocol === 'tunnel:') { protocol = 'http:' }mapping becausetunnel://127.0.0.1:10801is DevSidecar's pseudo-protocol for forwarding HTTPS traffic to the local Xray HTTP inbound via HTTP CONNECT. Without this conversion,tunnel:did not matchhttp:orhttps:, sogetTunnelAgentfell through tohttpsOverHttps— which attempts a TLS handshake to the Xray HTTP inbound port, causing Xray to reject the TLS ClientHello bytes asmalformed HTTP request. This broke all Xray-tunneled domains (chatgpt.com, openai.com, linux.do) after v2.2.0.
Changed
- Changed the Xray plugin's default
probeUrlfromhttps://www.google.com/generate_204tohttps://www.gstatic.com/generate_204inpackages/core/src/modules/plugin/xray/config.js. The old default (www.google.com) is not directly reachable from mainland China, so CN-based proxy nodes could not complete the observatory probe. The new default (www.gstatic.com) is reachable from mainland China (~0.1s) and uses HTTPS (port 443), which ensures only nodes that actually support 443 CONNECT tunneling are marked as available. Previously, an HTTP probeUrl (e.g.http://connect.rom.miui.com/generate_204) would pass nodes that only support port 80 forwarding, causingECONNRESETwhen users accessed HTTPS sites through the Xray tunnel.
v2.2.1
Fixed
- Restored
NODE_EXTRA_CA_CERTS/SSL_CERT_FILEcertificate loading logic inpackages/mitmproxy/src/lib/proxy/common/util.jsthat was inadvertently removed when syncing upstream v2.2.0. The v2.2.0 sync adopted upstream'sutil.json the assumption that the newREQUEST_CA_BUNDLEenvironment variable approach replaced the fork'sNODE_EXTRA_CA_CERTSworkaround; however, the two solve different problems and are complementary, not substitutes:REQUEST_CA_BUNDLEis read by client applications (curl, Python requests, etc.) to trust DevSidecar's MITM CA — it does not affect DevSidecar's own outbound HTTPS handshakes.NODE_EXTRA_CA_CERTSis read by Node.jsHttpsAgentfor DevSidecar's outbound TLS handshakes to target sites — and the Electron-bundled Node runtime ignores this env var, so the CA list must be read from the PEM file and passed explicitly via thecaoption.- Without this restoration, DevSidecar on corporate networks with TLS interception (SASE devices) fails outbound HTTPS handshakes with
UNABLE_TO_GET_ISSUER_CERT_LOCALLYbecause it does not trust the corporate re-signing CA. The restoredloadExtraCaCertsreads the PEM file, merges the certificates with Node's built-in root certificates, and passes the combined list via thecaoption to bothagentkeepalive'sHttpsAgent(verify andunVerifySslvariants) andtunnel-agent'shttpsOverHttp/httpsOverHttps. Module-level cache ensures the file is read only once per process.
Added
- Added TCP connect timeout retry in
packages/mitmproxy/src/lib/proxy/mitmproxy/createRequestHandler.js. WhenproxyRequestPromise()rejects with a connection timeout (hardcoded 7s TCP connect timeout), the request is retried up to 2 more times (MAX_CONNECT_RETRIES = 2, 3 total attempts). TheRequestCountermechanism switches to the next backup IP on each failure viadoCount(ip, true)→changeNext, so each retry attempt targets a different IP. This complements the existing IP-switch-on-next-request behavior by also recovering the current request instead of letting it fail. Only errors whose message contains连接超时are retried; all other errors propagate immediately.