JavaScript OOM #1897
差不多是使用goal进行翻译工作时,建立了480个子代理的时候遇到的 |
Replies: 4 comments
|
This is consistent with a workflow-capacity failure, but the current evidence does not yet distinguish a leak from legitimate retained state. Two source-backed details matter here:
For a translation workload of this size, I would split the input into restartable batches, begin with 2-4 concurrent children, set Example starting bounds: - id: delegation
name: cordis:group
group: true
isolate:
workflows: true
config:
- id: workflow-worker-thread
name: '@deepseek-ai/dsh-workflow-worker-thread'
config:
provider: spawn
maxConcurrentAgents: 4
maxTotalAgents: 64
- id: tool-workflow
name: '@deepseek-ai/dsh-tool-workflow'Visual operator guide with the evidence boundary and batching route: https://sandbaseai.github.io/deepseek-harness-handbook/subagent-scale-oom.html Handbook source update for review: sandbaseai/deepseek-harness-handbook#19 |
|
Reproduced as persistence backpressure, not a count-only leak.
|
|
This is a this has occurred because "JavaScript heap memory was filled" error. Quick Fix Anyways, with the current evidence it's hard to determine the root cause. If this keeps happening even with increased memory, there might be a memory leak worth investigating. Context: This appears to be a separate issue from the persistence backpressure problem described above—this is a general heap OOM on the dsh web server, not the workflow worker. The --max-old-space-size fix should resolve it for now, but if you're seeing repeated OOMs under normal load, profiling with --inspect or heap snapshots would help identify any leaks. |
|
We found that the fix for this is here: https://tomkornblit.substack.com/p/we-gave-an-agent-write-access-to |
This is consistent with a workflow-capacity failure, but the current evidence does not yet distinguish a leak from legitimate retained state.
Two source-backed details matter here:
maxConcurrentAgentsdefaults to0, which auto-resolves tomin(16, max(1, cores - 2)).maxTotalAgentsdefaults to1000. That is a runaway-call guard, not a memory budget, so a fatal OOM after about 480 children can occur below it.For a translation workload of this size, I would split the input into restartable batches, begin with 2-4 concurrent children, set
maxTotalAgentsclose to the expected batch size, and persist results between batches. Record peak heap, total and settled child counts, duration, failure…