[Bug][Windows] skill-filesystem crashes dsh web when a custom skill root contains an inaccessible directory #7537
Replies: 3 comments 2 replies
|
Reproduced — but the fatality is not where the report places it. Thanks for the exact stack, the
The chain, including the escape point
MeasurementsFixture: 16 sibling skill directories, each with one inaccessible grandchild. The reproduction is Windows-shaped by construction — a local copy of chokidar 5.0.0 whose
Instrumenting Two consequences. First, the only difference between dying and surviving is whether a listener is still attached when the late errors land (arm 1 vs arm 2) — the provider's own isolation is not the missing piece. Second, the carrier is the unawaited recursive call: a Why your line is a rejection, not an exceptionAt your commit Workaround until this is fixed
Fix options, with what each one measures
On your three suggestions"Watcher isolates a single directory's EPERM" is already implemented ( One platform note for anyone re-testing: on macOS I could not reproduce this without the shim. A |
|
Your A/B/D arms replicate the mechanism independently, and both open questions in your follow-up have a concrete answer: the boundary condition for the race (it needs two unreadable entries, not one), and where the browser actually enters the picture. Your arm C: one unreadable entry cannot raceArm C is the only one that disagrees with my measurements, and the difference is the entry count. The fatal needs a second failing
So it is not a rare flake: two entries fire it 8 times out of 10 with no artificial latency at all, and the only row that survives is the one where nothing closes the watcher. The window is not a millisecond-scale gap — the close path ( That also explains why the count matters on your machine: you measured 2 of the 6 unreadable directories inside the scan range. Two is exactly the minimum, and if arm C was run against a fixture with a single one — or one where the entries were visited strictly serially — it could not reproduce. Re-running arm C with two depth-1 unreadable directories should now reproduce it reliably; your own ACL-denied directories are enough, no shim required. The browser is the trigger, not the causeYour isolation run is right, and the reason is stronger than a timing correlation: the watcher is demand-driven, and the Web client is one of exactly two things that ever ask for the catalog.
With no client attached, no session scope is born, nobody calls This is also testable in one line: with the client attached, open the skills panel (or type Recap of the fix surfaceUnchanged from the earlier reply, and the near-determinism above is the reason the fix has to be structural rather than a narrower race window: pass |

Uh oh!
There was an error while loading. Please reload this page.
问题描述
在 Windows 上,当
@deepseek-ai/dsh-skill-filesystem配置customSkillDirs后,如果被监视的 Skill 目录中存在当前用户无法访问的子目录,dsh web会因为 Chokidar 调用realpath()返回EPERM而整体退出。我的配置为:
启动命令:
启动过程中会先输出 Web URL:
随后
skill-filesystem扫描 Skill 目录时发生:复现情况
最初出错目录为:
该目录本身存在:
均能正常返回。
但无法读取 ACL:
普通用户直接删除也会失败。
修复权限并删除该
.pytest_cache后,再次启动dsh web,随后在另一个 Skill 中出现完全相同的问题:再次修复并删除:
之后
dsh web可以正常启动。因此目前至少在两个独立 Skill 目录中可以观察到相同的失败模式。
预期行为
customSkillDirs中某个无关子目录不可访问时,不应该导致整个dsh web退出。比较合理的行为可能包括:
EPERM做错误隔离;.pytest_cache、__pycache__、.venv等明显与 Skill 定义无关的生成目录。关于
.pytest_cache权限问题这些无法访问的
.pytest_cache是我 Windows 开发环境中已经存在一段时间的问题,在使用 Codex 开发时也曾多次遇到。目前尚未确定这些目录的 ACL 为什么会变成不可访问状态,因此我不认为有足够证据说明这是 DSH 创建或破坏的。
本报告关注的问题仅是:
环境
v24.14.1c36a83ff6bb95e3f82cf79f9be7c724270a8aa61All reactions