[Performance][Windows][0.1.2-rc.1] Published npm packages in flat ZIP: cold Web startup dominated by synchronous file reads #5807
huangyiyang89
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
On Windows, the first launch from a newly extracted installation can take roughly 20–31 seconds before the
dsh web:ready URL appears. Subsequent processes using the same installation, even with a different emptyDSH_HOMEon every run, take roughly 3 seconds.I originally suspected MCP initialization, following #5129. Transport timing and a no-MCP control instead located most of the reproduced delay before any MCP process was spawned. The stock official Web profile also reproduces the synchronous file-read delay without the downstream desktop shell or plugins.
Entrypoint clarification (2026-09-06)
The measured entrypoint was the published npm package's built CLI,
node_modules/@deepseek-ai/dsh/lib/bin.js, launched by standalone Node.js 24.20.0. The profiling runs additionally used the diagnostic--requirepreload shown below. Pear resolves the npm package's declaredbin.dsh, which points tolib/bin.js.The source repository has a different launcher: its root
package.jsonatdsh-v0.1.2-rc.1definespnpm dsh webasnode --import tsx/esm apps/cli/src/bin.ts web. That source/tsx launch path was not benchmarked here. See the source repository script and CLI package bin declaration.Accordingly, “official CLI control” below means the unmodified published npm CLI in the same flat dependency distribution, started without the Pear Patch or Electron. These measurements establish the bottleneck for that launch path and layout; they do not establish the performance or cause of slow startup through the source repository's
pnpm dsh webcommand.Environment and controls
@deepseek-ai/dshand DSH packages:0.1.2-rc.1; standalone Node.js24.20.0.node_modulestree in a portable desktop ZIP. The official package contents are unchanged. This is a downstream package layout, not an official DSH executable release.DSH_HOMEand working directory. No existing sessions, custom model configuration, or downloaded plugin installation was involved.child_process.spawn, stdio JSON-RPC method timing, and public filesystem APIs. They did not alter official package source, MCP responses, configuration semantics, or application control flow. No tokens, MCP arguments/results, or session contents are included here.MCP on/off results
These downstream launcher measurements end after the CLI ready URL and an HTTP readiness check; they do not measure browser rendering. There is a small instrumentation/settling overhead.
The 30.9-second run had this transport timeline, relative to spawning the DSH Node process:
initializerequests sentinitializereplies receivedtools/listreplies receiveddsh web:URL emittedThus about 28 seconds elapsed before either MCP started. Both completed initialization and tool discovery; no hanging MCP connection was observed. An eight-run warm on/off matrix also completed normally in 2.57–3.31 seconds. Disabling MCP did not eliminate the cold behavior.
File-read profile
For extraction B, with both MCP entries disabled:
fs.readFileSynccallsV8 inspector samples identified these paths:
getSourceSync/readFileSync/ nativeopen.@deepseek-ai/dsh-app-boot:healProfilesModuleFallback→resolveModuleFallbackEntries→readModuleFallbackManifest. About 2.4 seconds of sampled self time was in the native UTF-8 reader under that manifest-reading path in this cold run.@deepseek-ai/dsh-client-modules:initialBundleSnapshot→ synchronous client bundle reads. Some individual large frontend bundle reads took 40–83 ms.These are sample attribution and synchronous-call wall times, not independent stage totals to add together. Timings of asynchronous filesystem promises were deliberately excluded from the table because their completion can be delayed by the blocked event loop.
Official CLI control, without the downstream Patch
A third newly extracted copy was launched directly:
There was no Electron process, Pear Patch, or configured MCP server in this control.
readFileSynccallsThe first control's HTTP assertion initially expected an unauthenticated 200 and received 401 after the ready URL and profile were recorded. That test assertion was corrected to accept the authenticated server's response; the subsequent controls completed. The cold numbers above are specifically the captured preload-to-announcement and synchronous-read measurements, not a claim that the original full smoke assertion passed.
For reproduction, prepare the published dependency tree before timing, extract it to a new program location, and start the command above with a new Home. Stop that process, keep the program location, change to another empty Home, and repeat. Timing
npxdependency installation itself would measure a different operation.A small diagnostic preload sufficient to observe the synchronous-read contrast:
Source-level follow-up (2026-09-06)
I traced the published build against the official
dsh-v0.1.2-rc.1source and analyzed the existing official-CLI cold.cpuprofilefurther. These are additional attribution of the same trace, not new benchmark runs or a measured code fix.In that 18,960 ms profiler window, mutually exclusive sample buckets inside synchronous reads were approximately:
getSourceSyncdefaultLoadImplreadModuleFallbackManifestinitialBundleSnapshotThe remaining roughly 3.75 seconds includes other reads, initialization, JS execution, native loading, and other samples. These buckets attribute time inside the read; they must not be added to the inclusive caller totals from a profiler UI or interpreted as CPU time consumed by file parsing.
The code explains the amplification of cold per-file latency:
runProfile; publishingbin.jsdoes not bundle all runtime dependencies into one file. A separately observed warm launch read set contained 1,558 non-client JS module paths, 483 existing package manifests, 46 client bundles, and 44 other existing JSON/YAML paths. These are unique existing paths in the observed read set; the totalreadFileSynccount is higher because paths can be read repeatedly and some probes fail.@earendil-works/pi-ai@0.84.4rootdist/index.jsre-exportsTypefromtypebox; TypeBox1.3.7root exports link its action/engine/extends/script/type modules. The observed startup set included 660 TypeBox JS paths, 124 pi-ai JS paths, and 94 Zod JS paths. These are module counts, not per-package timing attribution; this does not mean every module executes substantial work or any provider performs a network request. DSH's adapter imports are one actionable place to investigate deferring heavy runtime imports or providing a narrower supported upstream entrypoint.healProfilesModuleFallbackcallsresolveModuleFallbackEntriesfirst, which breadth-first traversesdependenciesandpeerDependenciesand synchronously reads manifests. Only afterward does it callmoduleFallbackCurrent. Existing profile links therefore do not skip that traversal. A process-local visited map avoids duplicates within the traversal, but there is no cross-launch closure cache on this path.initialBundleSnapshotsynchronously reads the client bundle and optional source map; composition builds startup batches and response maps held in process memory. The Web app waits for Loader settlement before announcing the URL consumed by the desktop launcher.EntryGroup.updateusesPromise.allSettled. However, this Node version's normal loader path reachesdefaultLoadSync/getSourceSync/fs.readFileSync. Merely wrapping more plugin starts inPromise.allcannot make those synchronous reads concurrent. This is an explanation of the observed path, not evidence of a Node regression.This narrows the optimization targets to the eagerly linked dependency graph, installation/profile metadata reuse, and frontend preparation. The cold/warm difference is largely the latency of repeated file access: in the earlier control, both runs performed 2,542 synchronous reads, while their accumulated duration changed from 15,204 ms to 239 ms. Windows cache/scanner attribution remains unproven.
Byte-preserving module archive experiment (2026-09-06)
A follow-up proof of concept aggregates unchanged module source bytes into a single archive and serves them through the public Node.js
module.registerHooks({ load })API. Module resolution and original file URLs remain in place; an indexed cache hit supplies the original source bytes and format, while misses usenextLoad. It does not rewrite the official package files or concatenate modules into one JavaScript scope. This source-delivery optimization is separate from the read-only instrumentation in the earlier controls. See the Node load hook API.For the published CLI control, the prebuilt archive contained 1,189 modules / 6,667,958 bytes. Capture compared every recorded source with the original file bytes and found zero differences. On a fresh extracted path, all 1,189 entries hit; reading, validating and indexing the archive took 16 ms.
Two additional directories were freshly extracted from the same ZIP. The archived case additionally received the prebuilt cache. Each directory was then run twice, with a different empty Home for each process. Both cases directly used the published official Web profile without Electron, the Pear Patch or any MCP, and both had the same filesystem/CPU diagnostic instrumentation.
Both cases reached the URL and passed a separate HTTP reachability check (401 expected without the authentication token). These times do not measure browser rendering, and archive construction time is excluded because the cache was generated in advance. This is one paired feasibility experiment, not a reproducible 49% improvement claim for every launch or machine. The same no-reboot/no-cache-flush/Defender-enabled limitations above apply.
A separate integration check exercised the real Electron GUI with the archive enabled, including the existing downstream Patch/plugin assertions, Browser and Office MCP tool checks, and DSH restart. It passed in capture mode and load mode. That check used a separately generated archive for the development pnpm layout: 1,281 modules, zero captured byte differences, and 1,281 hits on both startup and restart. It is behavioral validation of the mechanism, not another cold-start timing sample.
The proof of concept deliberately leaves source formats not explicitly identified by the public hook, package manifests, directly read client bundles and native files on their ordinary loading paths. It currently validates the Node version and archive payload checksum, but production use would additionally need build/dependency identity, update invalidation, and a safe fallback for a missing or invalid archive. It has not been wired into the distributed desktop runtime. The result supports investigating an officially maintained aggregate source/distribution cache to reduce cold per-file overhead without requiring a downstream DSH fork.
Potential optimization and request
An exploratory best-effort pre-read of the observed startup file set, using 16 concurrent filesystem reads before starting the unchanged official CLI, produced 10,447 ms total to readiness on another fresh extraction. That total includes 2,968 ms spent pre-reading. This is only one exploratory run, not a validated fix or a recommended user workaround, and it was not shipped in the desktop application.
Could the official startup path reduce serialized small-file reads—for example through installation-version-aware dependency metadata, bounded parallel preparation, or an officially supported distribution/startup cache—while retaining profile/plugin correctness? Separate timings for dependency closure, module import, client bundle preparation, and MCP readiness would also make these cases much easier to distinguish.
This report establishes a reproducible synchronous-file-loading bottleneck on this installation. It does not establish which Windows filesystem/cache/security component causes each slow read, and it does not dispute the separate MCP readiness problem in #5129.
中文摘要:每次均使用空用户数据,禁用全部 MCP 后首次启动仍约 20 秒。直接运行官方 CLI 也能复现,其中 2,542 次同步读文件累计约 15.2 秒,下一次仅约 0.24 秒。另一次完整启动前 28 秒尚未创建任何 MCP 子进程。希望优化依赖清单、模块和前端 bundle 的串行读取;系统文件缓存与实时扫描的具体占比尚未确认。
All reactions