[Bug][Windows] skill-filesystem watcher can hit a high-CPU rename-event storm on a custom Codex skills root #6099
Replies: 3 comments
|
Thanks for the detailed trace — the static claim holds up, and there is one more relevant fact in the tree that sharpens the scope of a fix. Confirmed at source (
The missing piece: the repository already has this concept. For a non- One trap worth avoiding while you are in there. A fix that ignores depth-1 directories that do not currently contain a
Minor evidence note: On attribution: the A/B you describe cannot separate the two hypotheses — closing the in-memory watcher removes both the subscription and the storm, so a 96x CPU drop is equally consistent with an external process generating the notifications. Worth a control before committing to a mitigation: with DSH stopped, attach a bare Scope note: this is provider-side (config schema + watcher construction) inside the first-party package, so a downstream plugin cannot patch it — a plugin cannot inject an |
|
Follow-up from the reporter: I implemented and committed a local non-polling containment patch in the RC.1 source checkout as I deliberately did not extend Local containment behaviorFor Windows native watchers only, the provider now counts Chokidar raw The defaults are configurable as: watchNativeBurstLimit: 256 # 0 disables this guard
watchNativeBurstWindowMs: 1000
watchNativeRecoveryMs: 2000This addresses the expensive part after a storm starts: closing the watcher destroys active I also checked the post-update steady state: a 5-second recursive snapshot of Validation completed:
The initial recovery delay is deliberately fixed and conservative because the observed event source is now quiet. If an immediate reopen proves to retrigger the burst repeatedly, the next focused change will be per-root exponential recovery backoff rather than switching to polling or ignoring |
|
Source update: the tested patch is now pushed to my fork, not only present in the local checkout.
The branch contains the Windows-only raw-rename burst recovery described above, plus its watcher and configuration-validation tests. It deliberately retains native watching and the shared Before the push I reran the repository's full |
Uh oh!
There was an error while loading. Please reload this page.
Summary
On Windows, a DSH Web process consumed significant CPU while idle (no UI interaction and no active model call). I traced it to the
@deepseek-ai/dsh-skill-filesystemChokidar watcher after configuring a custom skills root at:This is not a default DSH skills path; it was added through
customSkillDirs. The direct child\.systememitted a very high rate of native Windowsrenamenotifications, causing repeated Chokidar directory scans andlstatcalls on DSH's main process.Environment
0.1.5-alpha.2(a second0.1.5-alpha.1instance was configured with the same custom root)v24.13.05.0.0dsh webEvidence captured during the spike
CPU
Native watcher-event rate
Instrumenting the affected native
FSWatcher._handle.onchangecallback for 3 seconds recorded 109,487 events (about 36k/s).The events were repeatedly reported as
renamefor:During a short metadata check, the directory's
LastWriteTimeremained unchanged, so this looked like repeated watcher notifications rather than a sustained ordinary file update.V8 CPU profile
The hot stacks were consistent with Chokidar repeatedly rescanning that directory:
Other sampled hot functions were
node:pathnormalization/join/relative and Chokidar's event throttling.A/B mitigation
Why this seems avoidable in DSH
packages/skill/skill-filesystem/src/index.tsopens a depth-1 Chokidar watcher for each root inopenRootWatcher()(currently without anignoredfilter).For this custom root,
.systemis an immediate child, so it is watched. However, DSH's current direct-root discovery looks for either a top-level*.mdfile or<top-level-directory>/SKILL.md; this.systemcontainer has no direct.system/SKILL.md(it contains nested system-skill folders). Therefore it appears to be an irrelevant subtree for the current catalog, while it can still trigger repeated scans.Expected behavior
A custom skills root should not allow unrelated metadata/system directories or pathological native watcher notifications to consume a sustained amount of DSH CPU when the agent is otherwise idle.
Possible fixes / hardening ideas
Any of these would help:
skill-filesystem, and/or ignore non-skill metadata directories such as.git; carefully handle.systemonly where it cannot be a valid direct skill.*.md, and<candidate>/SKILL.md, rather than recursively watching every immediate subtree.readdirpscan.I do not suggest blindly ignoring every hidden directory, since deployments may intentionally use one. The main request is to avoid fully watching a subtree that cannot currently contribute a discoverable direct skill, or to make that behavior configurable.
Current workaround
Use a curated skills directory containing only actual DSH-compatible skill entries instead of pointing
customSkillDirsat a whole externally managed skills tree, or disable live skill watching for that root until a fix is available.I can provide a minimal reproduction / more profiler data if that would be useful.
All reactions