Releases: iLivius/cFMDbench
Release list
cFMDbench v1.0.1
Summary
This patch release improves the reproducibility of the cFMDbench Conda/R setup and fixes issues encountered when rendering the Quarto workflow with current conda-forge and mlr3 packages.
The default installation path now focuses on the standard ranger and xgboost workflow, while optional backends such as R torch/mlr3torch, TabPFN/mlr3extralearners remain explicit opt-ins.
Changes
- Reworked the R package installer to use a Conda-first strategy for compiled dependencies.
- Removed broad
dependencies = TRUEinstallation behavior. - Kept optional learner ecosystems out of the default install path.
- Added explicit install flags for optional backends:
INSTALL_R_TORCH=1INSTALL_TABPFN_R=1
- Hardened the Quarto notebook for server rendering.
- Improved project-root detection during Quarto runs.
- Made package loading deterministic with explicit conflict preferences.
- Disabled global
progressrhandlers during Quarto rendering. - Set
xgboostexplicitly tobooster = "gbtree"for compatibility with currentmlr3learners/paradoxvalidation. - Updated README and Conda setup documentation.
Recommended Default Workflow
For the default ranger and xgboost analysis:
conda env create -f conda/cfmdbench-r.yml
conda activate cfmdbench-r
Rscript conda/install_r_packages.R
quarto render analysis/cFMDbench.qmdcFMDbench v1.0.0
cFMDbench v1.0.0
ML classification on cFMD taxonomic profiles with mlr3 — Quarto workflow, Conda environments, for teaching and testing.
cFMDbench provides a reproducible, config-driven workflow for benchmarking machine-learning classifiers on metagenomic taxonomic profiles using the mlr3 ecosystem and the cFMD food metagenome resource. This is the first public release.
Highlights
Modular, auditable architecture
The codebase is organised as a lightweight Quarto notebook (analysis/cFMDbench.qmd) that orchestrates four dedicated R helper modules handling configuration, data acquisition, model training, and scoring. This separation keeps the analytical narrative readable while making each processing step independently inspectable and testable.
Declarative configuration
All tunable parameters — method selection, tuning budget, cross-validation scheme, preprocessing options, feature-selection thresholds, and visualisation settings — are centralised in a single config.yaml file. A config snapshot, together with the Conda environment specs, fully describes an experiment for reproducibility.
Nine classifiers, including GPU-accelerated learners
The benchmark supports Elastic Net (glmnet), k-NN, LDA, MLP (via mlr3torch), Naive Bayes, Random Forest, SVM, TabPFN, and XGBoost. TabPFN and MLP leverage GPU acceleration when a CUDA device is available. Each learner is wrapped in an mlr3 pipeline with joint hyperparameter tuning over preprocessing and model parameters.
Adaptive preprocessing and feature selection
The workflow includes configurable feature filtering (variance/correlation, information gain), optional SMOTE balancing (with automatic variant selection), and recursive feature elimination (RFE) in both single-learner and ensemble modes. Preprocessing runs inside each cross-validation fold to prevent data leakage.
Reproducible environment management
Two Conda environment specifications are provided: one for the R stack (cfmdbench-r.yml) and one for TabPFN with GPU support (cfmdbench-tabpfn-gpu.yml). A dedicated install script (conda/install_r_packages.R) pre-installs heavy R dependencies, and conda/README.md documents IDE-specific setup for RStudio, Positron, and VS Code. Instructions for pinning exact package versions are included.
SHA-based data caching
Dataset files are fetched from GitHub via the Contents API with SHA-based caching: only files whose hash has changed are re-downloaded. The dataset version is pinned in config.yaml, ensuring that the same configuration always retrieves the same data.
Benchmark results (cFMD v1.3.0)
The workflow was validated on cFMD v1.3.0 (3,252 metagenomes, 4,058 taxa reduced to 146 features via RFE, target: food category). Performance of four learners on the 20% held-out test set:
| Learner | Accuracy | Balanced Accuracy | Logloss |
|---|---|---|---|
| Random Forest | 0.964 | 0.867 | 0.233 |
| Support Vector Machine | 0.927 | 0.786 | 0.282 |
| TabPFN | 0.957 | 0.926 | 0.159 |
| XGBoost | 0.979 | 0.906 | 0.088 |
Model-agnostic SHAP feature importance is computed for the best-performing learner.
Planned for the next release
- LODO resampling — Leave-One-Dataset-Out cross-validation as an alternative outer resampling strategy for stricter generalisation assessment.
- Fairness analysis — Algorithmic fairness evaluation via mlr3fairness, measuring performance equity across groups defined by dataset of origin, geographic region, or food subtype.
Acknowledgements
- MASTER — Microbiome Applications for Sustainable food systems through Technologies and Enterprise.
- DOMINO — Harnessing the potential of fermentation for healthy and sustainable foods.
- FlavourFerm — Unleashing the flavour potential of plant-based foods.
Citation
If you use cFMDbench, please cite:
Antonielli, L. (2026). cFMDbench: Benchmarking ML classifiers on food metagenomic profiles with mlr3. Zenodo. https://doi.org/10.5281/zenodo.19607623
Please also cite the cFMD resource:
Carlino, N., et al. (2024). Unexplored microbial diversity from 2,500 food metagenomes and links with the human microbiome. Cell, 187(20), 5775–5795.e15. https://doi.org/10.1016/j.cell.2024.07.039