Repository navigation
0. Snakemake Pipeline
All the processing and analysis scripts for DNA-DNA SPRITE have been implemented into an automated pipeline using the Snakemake workflow management system.
![]()
The rules within the snakefile are written to utilise conda environments, which means the user does not need to install any packages other than conda and snakemake itself. Instructions on how to install conda and snakemake can be found here. Allow the installer to edit your .bashrc file.
The pipeline utilizes Bowtie2 for read alignment. The genome indexes required by Bowtie2 need to be supplied by the user. They can either be created from the genome reference fasta file or downloaded from the Bowtie2 project page.
The contents of the config.yaml file:
#email to which errors will be sent
email: "your_email@domain"
#Location of the config file for barcodeIdentification
bID: "./config.txt"
#Location of the samples json file produced with fastq2json.py script
samples: "./samples.json"
#output directory
output_dir: ""
#Currently "mm10" and "hg38" available
assembly: "hg38"
#Number of barcodes used
num_tags: "5"
#Repeat mask used for filtering DNA contacts
mask:
mm10: "mm10_blacklist.rmsk.milliDivLessThan140.bed"
hg38: "hg38_blacklist_rmsk.milliDivLessThan140.bed"
#Bowtie2 indexes location with prefix
bowtie2_index:
mm10: "add/path/to/bowtie2/mm10/index"
hg38: "add/path/to/bowtie2/hg38/index"
#Setting for mating heatmap matrix
#Plot a chromosome 'chr1' or plot genome wide 'genome'
chromosome:
- genome
# Options 'none', 'n_minus_one', 'n_over_two'
downweighting:
- n_over_two
#Number of ICE iterations
ice_iterations:
- 100
#Contact matrix resolution in bp
resolution:
- 1000000
min_cluster_size:
- 2
max_cluster_size:
- 1000
#Max value for heatmap plotting
max_value:
- 255
Setting which samples will be processed through the pipeline is done using the fastq2json.py script. It is advisable to create a single directory with all the fastq files to be processed. The location, file name and file pairs are summarised in json format by running:
python fastq2json.py --fastq_dir <path_to_fastq_directory>
This json summary is subsequently used to process each sample. The fastq2json.py script searches for fastq files with the following regular expression (.+)_(R[12]).fastq.gz which means file names like example_R1_001.fastq.gz will cause issues. The expected name is example_R1.fastq.gz.
Final sample.json file:
{
"example": {
"R1": [
"fastqs/example_R1.fastq.gz"
],
"R2": [
"fastqs/example_R2.fastq.gz"
]
}
}
The repository includes two shell scripts intended to make running snakemake easier.
run_pipeline_local.sh is intended for execution on a local machine and contains the following run commands:
mkdir -p workup
mkdir -p workup/logs
snakemake \
--snakefile Snakefile \
--use-conda \
--cores 10
run_pipeline.sh is intended for execution on a hpc with slurm. It contains the following commands:
mkdir -p workup
mkdir -p workup/logs
mkdir -p workup/logs/cluster
snakemake \
--snakefile Snakefile \
--use-conda \
-j 32 \
--cluster-config cluster.yaml \
--cluster "sbatch -c {cluster.cpus} \
-t {cluster.time} -N {cluster.nodes} \
--mem {cluster.mem} \
--output {cluster.output} \
--error {cluster.error}"
For running on an hpc, a cluster.yaml configuration file is included that specifies resource allocation for each snakemake rule. Job logs are written to a separate cluster subdirectory found in the main logs directory.