You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Extend splitChunks candidate discovery to include intersections of chunk sets. This allows Rspack to extract shared modules even when no individual module belongs to exactly the desired set of chunks.
The proposed optimization.splitChunks.dedupDepth option controls how many rounds of intersection discovery to perform. The proposed defaults are 1 in production and 0 in development, balancing output size against build and rebuild performance.
Motivation
Large applications can contain substantial duplication across entry points and dynamic imports. Improving shared-module extraction can reduce this duplication while preserving the existing chunking constraints.
Why chunking involves trade-offs
Chunking groups modules into output chunks. Consider three entries:
EntryA: a b
EntryB: a c
EntryC: b c
Keeping these entries separate produces three chunks. Each module appears twice in the output, but each entry loads only the modules it needs.
Extracting every shared module removes duplication:
EntryA -> chunk-a + chunk-b
EntryB -> chunk-a + chunk-c
EntryC -> chunk-b + chunk-c
chunk-a: a
chunk-b: b
chunk-c: c
Assuming the three entry chunks remain, there are now six chunks in total. Each entry needs additional requests, potentially for very small files.
Alternatively, all three modules can be placed in one shared chunk:
EntryA -> chunk-abc
EntryB -> chunk-abc
EntryC -> chunk-abc
chunk-abc: a b c
This produces four chunks under the same assumption. It avoids duplicated modules, but each entry loads one module it does not need: c for EntryA, b for EntryB, and a for EntryC.
These counts describe the entire example's output, not the number of requests for one page load. They illustrate three competing goals:
Reduce duplicated module bytes.
Avoid excessive small chunks and requests.
Avoid loading modules that a particular entry does not need.
The role of the runtime
Chunking based directly on native ESM imports and exports must account for module identity and evaluation order. Keeping a shared module in a single output location avoids duplicating its evaluation, but can lead to many small chunks. Duplicating modules where it is semantically safe can reduce fragmentation, although the opportunities depend on the application. A side-effect-free classification alone does not eliminate every module-identity or evaluation-order concern.
A bundler-controlled runtime separates chunk loading, module registration, and module execution. This gives the bundler more flexibility in choosing where module code is placed and when it executes.
This RFC focuses on improving shared-module extraction within Rspack's existing runtime and splitChunks model. It does not propose a new chunk-loading architecture.
Guide-level explanation
Terminology
Term
Meaning
Chunk set
The set of chunks containing a particular module.
Original chunk set
A distinct chunk set already present before additional intersection discovery.
Candidate
A possible shared-module extraction associated with a set of source chunks and a cache group.
Intersection
Chunks common to two chunk sets.
Discovery round
One pass that finds additional intersections; newly discovered sets become available in the following round.
Grouping by the same source chunk set matters: every module in the group is already used by every chunk in that set. Extracting the group therefore does not make those source chunks load unrelated modules.
These are candidates, not a final splitting plan. Cache-group conditions, minimum size, minimum chunk count, request limits, and the existing selection process still determine which candidates are used. For a cache group with minChunks: 2, the singleton candidates are ineligible.
The missing candidate
Now remove m3:
m1 -> {a, b, c}
m2 -> {a, b, d}
The original chunk sets are now only {a,b,c} and {a,b,d}. Neither is a subset of the other. The existing approach does not invent {a,b}, even though it is their intersection.
If m1 and m2 are individually large enough, they may be extracted separately. But consider this size relationship:
Neither individual extraction meets minSize, so duplication remains. Discovering {a,b} creates a candidate containing both modules:
Before:
a: m1 m2
b: m1 m2
c: m1
d: m2
After extracting the {a,b} candidate:
a -> shared
b -> shared
c: m1
d: m2
shared: m1 m2
Chunks c and d retain their own copies, but one copy of each module is removed from the combined output of a and b. The shared chunk meets minSize, and neither a nor b loads unrelated modules.
This partial deduplication is the key opportunity addressed by this RFC.
Reference-level explanation
Discover additional chunk-set intersections
For the example above:
{a,b,c} intersect {a,b,d} = {a,b}
Associate each discovered set with every original chunk set that contains it, not just the two sets that first produced it. Aggregate eligible modules for the relevant cache groups, then pass the additional candidates through the existing splitChunks selection process.
The objective is to find useful common chunk sets. Selecting only the largest intersection is insufficient, because savings depend on both the number of source chunks and the sizes of the modules that can be grouped.
Ignoring runtime overhead and compression, extracting modules of combined size S from k chunks removes approximately (k - 1) * S bytes: k copies become one shared copy. A smaller chunk set can therefore be more valuable if its eligible modules are sufficiently large. Overlapping candidates also affect one another as extraction proceeds.
Even complete intersection discovery does not guarantee the globally smallest output. Cache-group constraints and greedy candidate selection still apply.
Alternative: complete intersection closure
One approach is to repeatedly intersect newly discovered sets with the original sets until no new intersections appear.
Here, "complete" means the distinct nonempty intersections obtainable from the original sets. It does not mean enumerating every arbitrary subset of every chunk set.
Let:
N be the number of distinct original chunk sets.
D be the number of distinct chunks.
L be the maximum size of an original chunk set.
K be the number of additional distinct intersections.
For sorted sparse sets, one intersection costs O(L). Comparing original pairs costs approximately O(N²L), and comparing each newly discovered intersection with the originals adds O(NKL). The discovery phase is therefore approximately O((N² + NK)L), excluding bookkeeping and downstream candidate processing.
The problem is that K can grow exponentially. For example:
S1 = {b,c,d,e} // all chunks except a
S2 = {a,c,d,e} // all chunks except b
S3 = {a,b,d,e} // all chunks except c
S4 = {a,b,c,e} // all chunks except d
S5 = {a,b,c,d} // all chunks except e
S1 intersect S2 = {c,d,e}
S1 intersect S2 intersect S3 = {d,e}
S1 intersect S3 = {b,d,e}
Choosing different collections of original sets generates exponentially many distinct intersections in this family. There are at most 2^D possible chunk subsets, but that bound is still prohibitive.
Bounded discovery: pairwise intersections
The first round computes only intersections between original sets. Newly discovered intersections do not participate until a subsequent round.
Original sets: {a,b,c,d}, {a,b,c,e}, {a,b,d,e}
First round discovers:
{a,b,c}, {a,b,d}, {a,b,e}
A subsequent round can discover:
{a,b}
A single round is an approximation: it captures opportunities visible in pairs of original sets but can miss intersections that require combining more sets.
For a straightforward sparse-set implementation, first-round discovery costs O(N²L), with at most N(N - 1) / 2 additional candidates before deduplication. Associating discovered sets with all containing originals can still cost O(NKL) in the worst case. Bounding discovery does not remove downstream candidate-processing costs.
dedupDepth is a nonnegative integer controlling the maximum number of discovery rounds:
Value
Behavior
0
Keep existing candidate discovery; do not discover additional intersections.
1
Discover pairwise intersections of the original chunk sets.
2
Perform another round in which intersections from the first round can participate.
Higher values
Continue for up to the specified number of rounds, stopping early if no new intersections are found.
Each round uses the sets available at its start. Sets discovered during that round become available in the next round. Increasing the depth can expose additional candidates, but increases compilation work and does not guarantee monotonically smaller final output.
Proposed defaults:
Production: 1, to improve output size with bounded discovery.
Development: 0, to preserve existing behavior and avoid additional HMR/rebuild cost.
Users can explicitly choose 0 when build performance is more important, or evaluate higher values for workloads where additional deduplication is worthwhile. The option extends candidate discovery; it does not override existing cache-group filters or splitting constraints.
Drawbacks and alternatives
Additional discovery increases compilation time and memory use. More candidates can also increase the cost of cache-group processing and alter chunk boundaries, request counts, and cache behavior. Pairwise discovery intentionally misses some opportunities available to a deeper search.
Keeping the existing algorithm avoids this overhead but leaves useful partial-deduplication opportunities undiscovered. Complete intersection closure finds more candidates but can become prohibitively expensive.
GPU acceleration is another possible direction because set intersections involve many simple operations on numeric data. However, a GPU is not a dependable requirement for build environments, and accelerating intersections alone would not eliminate the cost of subsequent candidate and cache-group processing. It is not part of this proposal.
Questions for discussion
Is one discovery round an appropriate production default, given the trade-off between output size and compilation cost?
Are explicit depth control and the existing splitting constraints sufficient, or should discovery also have a candidate or work budget?
What workloads and page-level measurements should guide the choice of defaults?
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Implementation PR: #15497 — feat(split-chunks): add configurable dedupDepth
Summary
Extend
splitChunkscandidate discovery to include intersections of chunk sets. This allows Rspack to extract shared modules even when no individual module belongs to exactly the desired set of chunks.The proposed
optimization.splitChunks.dedupDepthoption controls how many rounds of intersection discovery to perform. The proposed defaults are1in production and0in development, balancing output size against build and rebuild performance.Motivation
Large applications can contain substantial duplication across entry points and dynamic imports. Improving shared-module extraction can reduce this duplication while preserving the existing chunking constraints.
Why chunking involves trade-offs
Chunking groups modules into output chunks. Consider three entries:
Keeping these entries separate produces three chunks. Each module appears twice in the output, but each entry loads only the modules it needs.
Extracting every shared module removes duplication:
Assuming the three entry chunks remain, there are now six chunks in total. Each entry needs additional requests, potentially for very small files.
Alternatively, all three modules can be placed in one shared chunk:
This produces four chunks under the same assumption. It avoids duplicated modules, but each entry loads one module it does not need:
cfor EntryA,bfor EntryB, andafor EntryC.These counts describe the entire example's output, not the number of requests for one page load. They illustrate three competing goals:
The role of the runtime
Chunking based directly on native ESM imports and exports must account for module identity and evaluation order. Keeping a shared module in a single output location avoids duplicating its evaluation, but can lead to many small chunks. Duplicating modules where it is semantically safe can reduce fragmentation, although the opportunities depend on the application. A side-effect-free classification alone does not eliminate every module-identity or evaluation-order concern.
A bundler-controlled runtime separates chunk loading, module registration, and module execution. This gives the bundler more flexibility in choosing where module code is placed and when it executes.
This RFC focuses on improving shared-module extraction within Rspack's existing runtime and
splitChunksmodel. It does not propose a new chunk-loading architecture.Guide-level explanation
Terminology
How existing candidates are found
Consider four chunks containing three modules:
The existing approach considers original chunk sets that are subsets of a module's chunk set, as well as singleton sets. In this example:
Conceptually, reversing this mapping gives possible module groups:
Grouping by the same source chunk set matters: every module in the group is already used by every chunk in that set. Extracting the group therefore does not make those source chunks load unrelated modules.
These are candidates, not a final splitting plan. Cache-group conditions, minimum size, minimum chunk count, request limits, and the existing selection process still determine which candidates are used. For a cache group with
minChunks: 2, the singleton candidates are ineligible.The missing candidate
Now remove
m3:The original chunk sets are now only
{a,b,c}and{a,b,d}. Neither is a subset of the other. The existing approach does not invent{a,b}, even though it is their intersection.If
m1andm2are individually large enough, they may be extracted separately. But consider this size relationship:Neither individual extraction meets
minSize, so duplication remains. Discovering{a,b}creates a candidate containing both modules:Chunks
canddretain their own copies, but one copy of each module is removed from the combined output ofaandb. The shared chunk meetsminSize, and neitheranorbloads unrelated modules.This partial deduplication is the key opportunity addressed by this RFC.
Reference-level explanation
Discover additional chunk-set intersections
For the example above:
Associate each discovered set with every original chunk set that contains it, not just the two sets that first produced it. Aggregate eligible modules for the relevant cache groups, then pass the additional candidates through the existing
splitChunksselection process.The objective is to find useful common chunk sets. Selecting only the largest intersection is insufficient, because savings depend on both the number of source chunks and the sizes of the modules that can be grouped.
For example:
Ignoring runtime overhead and compression, extracting modules of combined size
Sfromkchunks removes approximately(k - 1) * Sbytes:kcopies become one shared copy. A smaller chunk set can therefore be more valuable if its eligible modules are sufficiently large. Overlapping candidates also affect one another as extraction proceeds.Even complete intersection discovery does not guarantee the globally smallest output. Cache-group constraints and greedy candidate selection still apply.
Alternative: complete intersection closure
One approach is to repeatedly intersect newly discovered sets with the original sets until no new intersections appear.
Here, "complete" means the distinct nonempty intersections obtainable from the original sets. It does not mean enumerating every arbitrary subset of every chunk set.
Let:
Nbe the number of distinct original chunk sets.Dbe the number of distinct chunks.Lbe the maximum size of an original chunk set.Kbe the number of additional distinct intersections.For sorted sparse sets, one intersection costs
O(L). Comparing original pairs costs approximatelyO(N²L), and comparing each newly discovered intersection with the originals addsO(NKL). The discovery phase is therefore approximatelyO((N² + NK)L), excluding bookkeeping and downstream candidate processing.The problem is that
Kcan grow exponentially. For example:Choosing different collections of original sets generates exponentially many distinct intersections in this family. There are at most
2^Dpossible chunk subsets, but that bound is still prohibitive.Bounded discovery: pairwise intersections
The first round computes only intersections between original sets. Newly discovered intersections do not participate until a subsequent round.
A single round is an approximation: it captures opportunities visible in pairs of original sets but can miss intersections that require combining more sets.
For a straightforward sparse-set implementation, first-round discovery costs
O(N²L), with at mostN(N - 1) / 2additional candidates before deduplication. Associating discovered sets with all containing originals can still costO(NKL)in the worst case. Bounding discovery does not remove downstream candidate-processing costs.Configuration
dedupDepthis a nonnegative integer controlling the maximum number of discovery rounds:012Each round uses the sets available at its start. Sets discovered during that round become available in the next round. Increasing the depth can expose additional candidates, but increases compilation work and does not guarantee monotonically smaller final output.
Proposed defaults:
1, to improve output size with bounded discovery.0, to preserve existing behavior and avoid additional HMR/rebuild cost.Users can explicitly choose
0when build performance is more important, or evaluate higher values for workloads where additional deduplication is worthwhile. The option extends candidate discovery; it does not override existing cache-group filters or splitting constraints.Drawbacks and alternatives
Additional discovery increases compilation time and memory use. More candidates can also increase the cost of cache-group processing and alter chunk boundaries, request counts, and cache behavior. Pairwise discovery intentionally misses some opportunities available to a deeper search.
Keeping the existing algorithm avoids this overhead but leaves useful partial-deduplication opportunities undiscovered. Complete intersection closure finds more candidates but can become prohibitively expensive.
GPU acceleration is another possible direction because set intersections involve many simple operations on numeric data. However, a GPU is not a dependable requirement for build environments, and accelerating intersections alone would not eliminate the cost of subsequent candidate and cache-group processing. It is not part of this proposal.
Questions for discussion
All reactions