-
Notifications
You must be signed in to change notification settings - Fork 2
Multi Sample Binning
MAGGIC implements strata-limited multi-sample binning, a computationally efficient approach that improves MAG recovery while keeping alignment counts tractable for large cohorts.
Multi-sample binning aligns reads from multiple samples against every assembled contig set, giving binning tools (VAMB, SemiBin2, MetaBat 2) cross-sample coverage profiles. As shown by Han et al. 2025, this recovers 41-43% more species and strains compared to single-sample binning, including rare taxa and antibiotic resistance gene hosts.
However, traditional all-vs-all alignment produces N² BAM files (50 samples = 2,500 alignments; 100 samples = 10,000 alignments). Coverage diversity saturates around 20-30 samples for most metagenomic experiments; beyond that, marginal improvement occurs at exponentially increasing computational cost (Han et al. 2025; Nissen et al. 2021; Haryono et al. 2022).
MAGGIC selects a subset of strata_size samples (default 15) and aligns their reads to every assembly, producing strata_size × N BAM files instead of N². Selection uses staggered sampling, which distributes selected samples uniformly across the full sorted sample list.
This avoids geographic selection bias. In runs where sample IDs encode spatial information (state prefixes, site codes, collection dates), taking the first N samples concentrates coverage from a single region. Geographic distance is the primary driver of beta diversity in environmental samples (Cheng et al. 2024; Peng et al. 2025).
For example, with 500 samples and strata_size = 15:
| Method | Selected indices | Geographic spread |
|---|---|---|
| First-N (sequential) | 0, 1, 2, 3, ... 14 | 1-2 states if IDs are alphabetically ordered |
| Staggered intervals | 0, 33, 66, 99, ... 483 | ~15 states evenly distributed |
The binning algorithms (VAMB' variational autoencoder, MetaBAT 2' coverage clustering, SemiBin2's graph neural network) learn from these co-abundance patterns (Nissen et al. 2021; Han et al. 2025).
| Parameter | Default | Description |
|---|---|---|
--multi_sample_strata |
true |
Enable strata-limited mode (true by default. Set false for full all-vs-all) |
--strata_size |
15 |
Number of samples whose reads are aligned to each assembly |
| Samples | Full All-vs-All | Strata (15) | Reduction |
|---|---|---|---|
| 30 | 900 BAMs | 450 BAMs | 2x |
| 50 | 2,500 BAMs | 750 BAMs | 3.3x |
| 100 | 10,000 BAMs | 1,500 BAMs | 6.7x |
| 200 | 40,000 BAMs | 3,000 BAMs | 13.3x |
| 500 | 250,000 BAMs | 7,500 BAMs | 33.3x |
The default strata size of 15 balances cross-sample coverage diversity with compute cost. For highly diverse cohorts (e.g., environmental samples from disparate locations), increase --strata_size to 20-30. For homogeneous cohorts (e.g., clinical samples from the same body site), 10-15 is sufficient.
./cpipes \
--pipeline maggic \
--input /path/to/fastq/dir \
--output /path/to/output \
--strata_size 20 \
-profile ahptainer \
-resume