Skip to content
Discussion options

You must be logged in to vote

This is consistent with a workflow-capacity failure, but the current evidence does not yet distinguish a leak from legitimate retained state.

Two source-backed details matter here:

  • maxConcurrentAgents defaults to 0, which auto-resolves to min(16, max(1, cores - 2)).
  • maxTotalAgents defaults to 1000. That is a runaway-call guard, not a memory budget, so a fatal OOM after about 480 children can occur below it.

For a translation workload of this size, I would split the input into restartable batches, begin with 2-4 concurrent children, set maxTotalAgents close to the expected batch size, and persist results between batches. Record peak heap, total and settled child counts, duration, failure…

Replies: 4 comments

Comment options

You must be logged in to vote
0 replies
Answer selected by smter
Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
5 participants