Skip to content

v1.4.1

Latest

Choose a tag to compare

@github-actions github-actions released this 27 Sep 11:12

Fixed / 修复

  • Windows 免安装版:「下载更新 → 更新并重启」之后什么都不发生,手动打开还是旧版本号。这条链路上叠了三个坑,任一个都足以让更新彻底失效:
    Windows portable build: after "download update → install and restart" nothing happened, and manually reopening still showed the old version. Three separate faults were stacked on this one path, any of which alone was enough to break the update completely:
    • 引导程序根本起不来(根因)。免安装版是 PyInstaller onedir 产物 —— exe 必须和同级的 _internal\python313.dll 一起才能跑。而自更新把 sys.executable 单独复制到 %TEMP%\asb-update-bootstrap.exe 再去执行,这个裸 exe 找不到 _internal,进程在加载 Python DLL 时就死了(Failed to load Python DLL ...),于是没有任何人去处理暂存包。现在改为直接拿暂存目录里那份新版的 exe 当引导程序:它自带完整的 _internal,能原地跑,又正好在安装目录之外(替换安装目录不会被自己锁住),还省掉了再复制一份几百 MB 的文件。
      The bootstrap could not start at all (root cause). The portable build is a PyInstaller onedir artifact — the exe only runs alongside its sibling _internal\python313.dll. The updater copied sys.executable alone to %TEMP%\asb-update-bootstrap.exe and ran that; the bare exe could not find _internal and died while loading the Python DLL (Failed to load Python DLL ...), so nobody ever processed the staged update. It now runs the new version's own exe from the staging directory as the bootstrap: it carries its full _internal, runs in place, sits outside the install directory (so replacing that directory is not blocked by its own file lock), and avoids copying several hundred MB again.
    • 旧服务退得太早,引导程序还来不及动手就先撞上文件锁。原来触发后固定 0.6 秒就 os._exit(0),而引导程序要先把几百 MB 的 onedir 复制一遍才真正落地 —— 真正的危险窗口是「引导程序开始复制」之前。现在两个进程之间做握手:引导程序启动后第一件事是回填 update_boot.json 的 reportedAt,服务轮询到这个标记才退出;轮询不到就带原因把「安装失败」返回给界面,而不是静默退出留下一个没人处理的暂存包。服务退出后,引导程序还会等它确实消失(OpenProcess + WaitForSingleObject,而不是在 Windows 上会真杀进程的 os.kill(pid, 0))才开始替换。
      The old service exited too early — the bootstrap hit the file lock before it could act. The old code always os._exit(0) after a fixed 0.6s, while the bootstrap has to copy the whole several-hundred-MB onedir before it can land anything; the dangerous window is before it starts copying. The two processes now shake hands: the bootstrap's first act is to fill in reportedAt in update_boot.json, and the service polls for that marker before exiting; if it never appears the service returns a "install failed" reason to the UI instead of silently exiting and leaving an unprocessed staging directory. After the service exits, the bootstrap also waits for it to actually disappear (OpenProcess + WaitForSingleObject, rather than os.kill(pid, 0), which on Windows really does kill the process) before replacing anything.
    • 失败了完全不留痕。引导程序是在服务已经退出之后跑的,它崩了就没有任何界面能看到 —— 用户看到的只是「点了没反应」。现在它把每一步写进 <数据目录>\update.log,失败时留一条 update_boot.json 事故记录(想装的版本 / 退出码 / 人话原因);服务下次起来时 /api/update/last-boot 把这条读给界面显示一次(读过即清,不会每次开都弹旧事故)。
      No trace was left when it failed. The bootstrap runs after the service has exited, so if it crashed there was no UI left to show it — the user just saw "nothing happened". It now writes every step to <data dir>\update.log and, on failure, leaves a record in update_boot.json (target version / exit code / plain-language reason); on the next start the service reads it back once via /api/update/last-boot and shows it (then clears it, so a past incident is not re-announced on every launch).
  • 界面上的「正在重启…」其实是盲等 4 秒。老实现是 setTimeout(reload, 4000) —— 如果窗口根本没起来,4 秒后刷出来的还是老页面,用户看到的就是「没反应、版本号也没变」。现在改成轮询 /api/health 等版本号真的变成新版本再刷新,并区分三种结局:变了 → 刷新;服务活着但版本没变 → 明确提示「服务已重启,但版本仍是 vX(未换成 vY),详见 update.log」;一直连不上 → 提示手动双击启动。页面上「下载更新」时显示的下载进度在重装后也能直接看到已下载状态。
    The UI's "restarting…" was really a blind 4-second wait. The old code did setTimeout(reload, 4000) — if the window never came up, the reload just redisplayed the old page, which is precisely what "nothing happened / version unchanged" looks like. It now polls /api/health until the version really changes before reloading, and distinguishes three outcomes: changed → reload; service alive but version unchanged → explicitly report "the service restarted but is still vX (not vY), see update.log"; never reachable → prompt the user to start it manually.
  • 安装时不会再把安装目录覆盖成空壳。落地前先确认暂存那一层里真的存在 exe(压缩包多包了一层就往下找一层);找不到就放弃并说明原因,而不是拿一个只有目录结构的空框架去替换整个安装目录。
    Installing can no longer replace the install directory with an empty shell. Before landing, it verifies that the staging level actually contains an exe (descending a level if the archive is double-wrapped); if not it aborts with a reason instead of overwriting the whole install directory with a skeleton of directories.
  • 落地成功后不再一直留着几百 MB 的备份和暂存包。备份 .bak 会保留到新版本确实启动起来(万一新版起不来还能人工改名回滚),此后由新进程自动清理 .bak 与 data\updates\。
    A few hundred MB of backup and staging data are no longer left behind forever after a successful install. The .bak backup is kept until the new version has actually started (so a broken new build can still be rolled back by renaming it); after that the new process cleans up .bak and data\updates\ automatically.
  • 测试:新增 scanner/tests/test_selfupdate.py(引导程序选谁 / 落地替换 / 自我锁定保护 / 事故留痕 / 参数传递 / 端到端,56 项)与 scanner/tests/test_bootstrap_launcher.py(引导程序控制流:报到 → 等旧进程 → 落地 → 拉起,含等待超时、暂存缺失、拉起失败三种失败分支与日志断言,31 项)。后者补上的是此前完全没有覆盖的一段 —— 它就藏在打包目录里、又只在服务退出后才跑,人工点界面观察不到。另外在真实发布产物上实测确认了三件事:单独复制出来的 exe 确实起不来(复现原 bug)、运行中的 onedir 目录可以被复制和改名(所以拿它当引导程序是安全的)、exe 从启动到能应答只要 1.5 秒(所以握手超时给 25 秒足够宽裕)。
    Tests: new scanner/tests/test_selfupdate.py (which bootstrap gets picked / landing the replace / self-lock protection / failure records / argument passing / end to end, 56 assertions) and scanner/tests/test_bootstrap_launcher.py (the bootstrap's control flow: report in → wait for the old process → land → relaunch, including the wait-timeout, missing-staging and relaunch-failure branches plus log assertions, 31 assertions). The latter covers a stretch that previously had no coverage at all — it hides in the packaging directory and only runs after the service has exited, so clicking through the UI can never observe it. Three things were also verified empirically against the real release artifact: an exe copied out on its own really does fail to start (reproducing the original bug), a running onedir directory can be copied and renamed (so using it as the bootstrap is safe), and the exe takes only 1.5 s from launch to answering (so the 25 s handshake timeout is comfortably generous).