release v1.1.0 by FerriolCalvet · Pull Request #440 · bbglab/deepCSA

FerriolCalvet · 2026-04-06T16:21:50Z

This PR includes several updates mainly related to:

Increased efficiency
Updated output structure
Several new QCs
Updated signature analysis
Bug fixes

AI updates

…annotating with no flags."

…eated in create_consensus_panel.py

Implemented parallel processing of VEP annotation through configurable chunking: - Added `panel_sites_chunk_size` parameter (default: 0, no chunking) - When >0, splits sites file into chunks for parallel VEP annotation - Uses bash `split` command for efficient chunking with preserved headers - Modified SITESFROMPOSITIONS module: - Outputs multiple chunk files (*.sites4VEP.chunk*.tsv) instead of single file - Logs chunk configuration and number of chunks created - Chunk size configurable via `ext.chunk_size` in modules.config - Updated CREATE_PANELS workflow: - Flattens chunks with `.transpose()` for parallel processing - Each chunk gets unique ID for VEP tracking - Merges chunks using `collectFile` with header preservation - Added SORT_MERGED_PANEL module: - Sorts merged panels by chromosome and position (genomic order) - Prevents "out of order" errors in downstream BED operations - Applied to both compact and rich annotation outputs - Enhanced logging across chunking pipeline: - SITESFROMPOSITIONS: reports chunk_size and number of chunks created - POSTPROCESS_VEP_ANNOTATION: shows internal chunk_size and expected chunks - CUSTOM_ANNOTATION_PROCESSING: displays chr_chunk_size and processing info Configuration: - `panel_sites_chunk_size`: controls file chunking (0=disabled) - `panel_postprocessing_chunk_size`: internal memory management - `panel_custom_processing_chunk_size`: internal chromosome chunking Benefits: - Parallelizes VEP annotation for large panels - Reduces memory footprint per task - Maintains genomic sort order for downstream tools

…ng configs

* add negative regions to subgenic generation - tested * update naming to complement

* include mut_density_qc compile - not tested * annotate failing omegas * improve functionalities of script - add more plots * handle subgenic omega computations proper plotting generation * remove unnecessary function * minor update to variability plotting - skip clustermap if not enough groups - add synonymous

- following what Fede (federica.brando@irbbarcelona.org) did in the PR #153

- Also adding a new feature to extract flagged regions to a bed file. - Adding wrapper function to clean-up main. e On branch 118-reapply-mask-discarded-mutation

e Your branch is ahead of 'origin/118-reapply-mask-discarded-mutation' by 1 commit.

- add output as pipeline in 'module.config' - add output in `filtermaf` module - add output in the `mutationpreprocessing`subworkflow

* add updated compute mutdensity script * add subworkflow enrichpanels (untested) * fix missing input channel enrichpanels * fix components enrichpanels * fix variable * fix permissions compute_mutdensity.py changed * update expand regions outputs to panel names * remove redundant DNA2PROTEINMAPPING step * updates and fixes in synmutdensity * fix bug subgenic definitions - update naming - update omega container * update omega version definition - add proper CPU definition in estimator * apply review suggestions - not tested * fix mut density bug after review * update previous fix --------- Co-authored-by: rochamorro1 <rocio.chamorro@irbbarcelona.org>

…all species, and add shell injection protection (#393) * Initial plan * Update plotting to work with any available selection data - Modified workflows/deepcsa.nf to allow plotting when ANY selection method is available (changed && to ||) - Updated plot_selectionsideplots.py to check file existence before loading data - Added graceful handling of missing data files with informative warnings - Script now dynamically determines which tracks to plot based on available data - No plots generated if no data files are found Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Add safety checks for empty tracks and None figure - Added check in plot_all_positive_selection to handle n_tracks == 0 - Added check to avoid saving figure when None is returned - Improved robustness when no data is available Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Remove homo_sapiens species requirement from plotting and DNA2PROTEINMAPPING - Removed homo_sapiens check from plotting condition (line 548) - Removed homo_sapiens check from DNA2PROTEINMAPPING (line 232) - DNA2PROTEINMAPPING now runs for all species - Added comprehensive unit tests for plot_selectionsideplots.py - Tests cover file existence checks, dynamic track selection, and edge cases Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * add within gene saturation script * make saturation plots optional * update schema * update qcs and saturation - mutdensityqc works without requiring omega - integrated plotting for saturation of subgenic regions * update plot saturation - works when not all data is available * Add shell injection protection to samplesheet validation - Enhanced sample name validation in check_samplesheet.py - Only allow alphanumeric characters, underscores, hyphens, and dots - Prevent sample names starting with hyphen (flag injection) - Block shell metacharacters: semicolons, dollar signs, backticks, pipes, ampersands, redirects, quotes, slashes - Added comprehensive test suite (15 tests) in test_check_samplesheet.py - Prevents command injection attacks via malicious sample names in samplesheet Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * update tests location --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> Co-authored-by: FerriolCalvet <ferriolcalvet@gmail.com>

- fix remaining linting

@FerriolCalvet

* Initial plan * Add plot_vaf_depth.py script and Nextflow module Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Add documentation for VAF vs depth plots Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Add example script for VAF depth plots Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Add comprehensive implementation summary Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Refactor: Split VAF depth plots into independent modules - Created separate scripts for VAF/mutation density (with hyperbolic curves) and selection metrics (without) - Created independent Nextflow modules for each plot type - Added PLOT_DEPTH_RELATIONSHIPS subworkflow to integrate modules - Connected subworkflow to main workflow with plot_depth_relationships parameter - Removed hyperbolic curves from omega and OncodriveFML plots as requested Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Fix: Correct channel handling and remove old module - Fixed subworkflow to use per-sample mutation density channels - Removed old vaf_depth module and script - Simplified channel joining logic in subworkflow - Extract depth_per_gene file from depths output Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * Update documentation to reflect module split and independence - Updated vaf_depth_plots.md to document two independent modules - Clarified that hyperbolic curves only appear in VAF/mutation density plots - Updated example script to show separate usage of each module - Added clear distinction between modules with and without hyperbolic curves Co-authored-by: FerriolCalvet <38539786+FerriolCalvet@users.noreply.github.com> * working update - still plots to be polished * add plotting QCs at single mutation level * fix bugs after testing * add qc of individual mutations -not tested * add mutation counts per sample plot - not tested * fix bug after testing * Apply suggestion from @FerriolCalvet --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>

@bkohrn

- thanks @bkohrn for reporting it

…pdated

…g in VEP annotation and adjust related schema defaults. Linting fixes

- increase test cpus

…ading

This reverts commit 19239f5.

TESTS: testing with small bladder set

add bbgregressions

…erformance

- Modified nextflow.config to include general reference paths and skip validation for specific parameters. - Increased resource limits for processes to accommodate VEP execution. - Changed panel_sites_chunk_size to 0 and disabled parameter validation. - Added new input_maf.csv file with sample and VCF path data for testing.

…SVEPPANEL Dev chunk optimization postprocessveppanel

- tests working with external files, pending omega

Add minor fixes

migrau and others added 30 commits November 16, 2025 13:10

fix: gene omega error: "No flagged entries found; skipping plots and …

b2f12fd

…annotating with no flags."

fix: Add debug logging and ensure failing_consensus file is always cr…

d4ed3c2

…eated in create_consensus_panel.py

Update naming of sampling method to without replacement

6060612

feat: add parallel_processing_parameters section to schema for chunki…

e52cb76

…ng configs

update dnds genes list

92580ce

make failing consensus optional

19d4210

Add the option to generate complementary subgenic regions (#396)

0bea730

* add negative regions to subgenic generation - tested * update naming to complement

update omega container to v0.2.1

eb551da

refactor: move repetitive variants filter to function (#118)

04e740e

- following what Fede (federica.brando@irbbarcelona.org) did in the PR #153

feat: substitute regressions module by bbgregressions

fc1e24a

feat: update config with predictors info

a608e82

fix: correct syntax to input omega outputs to regressions

d9ea4c7

refactor: follow PR153 code refactoring (#118)

78b182b

- Also adding a new feature to extract flagged regions to a bed file. - Adding wrapper function to clean-up main. e On branch 118-reapply-mask-discarded-mutation

refactor: add docstrings as in PR153 (#118)

9123b0e

e Your branch is ahead of 'origin/118-reapply-mask-discarded-mutation' by 1 commit.

add: output BED of flagged positions from FILTERBATCH step (#118)

71fe982

- add output as pipeline in 'module.config' - add output in `filtermaf` module - add output in the `mutationpreprocessing`subworkflow

update predefined profiles

a124b68

fix linting warnings

f7e3c35

additional fixes

c355494

update configs and panel outputs

94601fc

rename to MUTREADSDENSITY

74c37a5

- fix remaining linting

rename subsetmutations to querymutations

f642b94

rename subsetdepths to querydepths

eec7579

minor renaming of subgenic level info

adfbed2

fix after quick test

8f7f953

fix bug of filename collision

ebe4f67

migrau and others added 21 commits April 2, 2026 17:48

fix: Correct typo in process name for panel comparison configuration

fdbbdb5

fix bug in downsample notebook

4a3bdc5

- thanks @bkohrn for reporting it

Merge remote-tracking branch 'origin/dev' into feat/bbg-regressions_u…

d526c8e

…pdated

fix bug in gene coverage plots

dd196f0

update gene selection and config output

ea6da6f

feat: Update panel_sites_chunk_size to 1,000,000 for improved chunkin…

c503b8e

…g in VEP annotation and adjust related schema defaults. Linting fixes

doc: update env variable in docs

5a8f62c

- increase test cpus

feat: Add Protein_position to schema overrides for VEP output file lo…

3ecfb18

…ading

refactor: default min depth to 30

27916cc

refactor: rename test and update docs

fef7c2c

minor updates to config outputs

19239f5

Revert "minor updates to config outputs"

b12160b

This reverts commit 19239f5.

apply copilot suggestions and linting

9b0aa87

refactor: update snapshot

3013e66

Merge pull request #438 from bbglab/test/add_testing

7097fe4

TESTS: testing with small bladder set

Merge pull request #441 from bbglab/feat/bbg-regressions_updated

2bba6e3

add bbgregressions

feat: Update resource allocation for analysis processes to optimize p…

835656f

…erformance

Merge branch 'dev' into minor-customizations

3a3cf28

Merge branch 'dev' into dev-chunk-optimization-POSTPROCESSVEPPANEL n

a089e9a

Merge pull request #390 from bbglab/dev-chunk-optimization-POSTPROCES…

cfa2a97

…SVEPPANEL Dev chunk optimization postprocessveppanel

FerriolCalvet marked this pull request as ready for review April 10, 2026 13:04

FerriolCalvet and others added 6 commits April 10, 2026 15:06

Merge branch 'dev' into minor-customizations

2c49974

update outputs structure

5889a07

- tests working with external files, pending omega

add missing params to nextflow_schema

9505e82

reintroduce omega content tests with bladder data

7c469b1

update snapshot

96acfaf

Merge pull request #443 from bbglab/minor-customizations

d1de688

Add minor fixes

FerriolCalvet merged commit ea9d4e3 into main Apr 11, 2026

FerriolCalvet deleted the dev branch April 11, 2026 15:40

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

release v1.1.0#440

release v1.1.0#440
FerriolCalvet merged 299 commits intomainfrom
dev

FerriolCalvet commented Apr 6, 2026 •

edited

Loading

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

6 participants

Conversation

FerriolCalvet commented Apr 6, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

AI updates

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

6 participants

FerriolCalvet commented Apr 6, 2026 •

edited

Loading