[Bug: Windows] Settings save intermittently fails with EPERM — rename in writeFileAtomic needs a transient-failure retry #1628
Replies: 5 comments
|
感谢分享实测经验!对照你们的做法,逐条说明一下当前修复(
|
|
第 5 点采纳了,已实现并推到
测试补了四个断言形状:单次成功、重试两次后成功( "没有这个字段,重试是隐形的"这个总结很准——之前那 1.7% 的偶发 EPERM 就是典型的"重试了但没人知道"。感谢补充这点。 |
|
补一个 diff 直达链接,方便团队直接取用参考实现(注意到仓库暂不接受外部 PR,分支在 fork 上): Compare/diff: main...xiaoyuyu6420:fix/atomic-write-windows-rename-retry 单个 patch 下载:给上面的 URL 加 后缀即可。共 3 个提交:重试修复( |
Uh oh!
There was an error while loading. Please reload this page.
Summary
On Windows, saving settings in the Web UI (e.g. the model API key) intermittently fails with:
Root cause
writeFileAtomicin@deepseek-ai/dsh-atomic-writecommits by writing a random-suffix.tmpsibling and renaming it over the target in a single step — with no retry. On Windows, antivirus real-time scanners and search indexers briefly open an exclusive handle on the freshly created.tmpfile (or the target), and a Windows rename over a file with an open handle fails withEPERM. The failure is purely transient, but it propagates straight to the UI.Upstream libraries (
write-file-atomic,graceful-fs) handle this exact case with bounded retries.Reproduction
Looping the same write-temp-then-rename sequence 60 times against
~/.dshon Windows 11 (stock Defender, no exotic software) failed 1 of 60 times (~1.7%) — every retry immediately after a failure succeeds. This matches the "saving settings sometimes errors, saving again works" symptom.Fix
I've pushed a minimal fix to a fork branch: xiaoyuyu6420/deepseek-harness@master...xiaoyuyu6420:deepseek-harness:fix/atomic-write-windows-rename-retry
writeFileAtomicnow routes the rename through arenameWithRetryhelper:EPERM/EBUSYare retried with bounded backoff (3 attempts, 50→200 ms); every other failure — or one that survives the retries — is rethrown unchanged, so callers keep observing the original error. Unit tests cover the transient-retry-then-commit path, the exhausted-retry rethrow, an immediateEACCESrethrow without retrying, and an error without an errno code.After the fix, the same stress loop completes 300/300 writes successfully.
Since CONTRIBUTING.md explains external pull requests are not accepted at the moment, I'm following the contributing guide and posting here instead — the branch is MIT-licensed like the rest of the repo, so feel free to cherry-pick it wholesale or reimplement the retry in whatever shape fits your conventions. Happy to rework it if a maintainer prefers a different approach (e.g. a shared retry helper or fsync-related changes from the existing TODO).
Environment
47f9438@deepseek-ai/dsh@0.1.0-rc.6vianpx @deepseek-ai/dsh webAll reactions