Skip to content

Generating Sample File

ha trang phung edited this page May 3, 2023 · 2 revisions

An example of a correct sample file is provided in the samples.csv file of the Git repository, which contains 4 columns:

  • SampleName: Specifies bulk names, including P1 (for Parent 1), P2 (for Parent 2), R (for Resistance), S (for Sensitivity)
  • mode: Sequencing mode. The valid modes are: PE (Paired-End), SE (Single-End), SPET (Single Primer Enrichment Technology with UMI), and SPETNOUMI (Single Primer Enrichment Technology without UMI)
  • fq1: Forward reads (R1) sequencing data filename(s), which are required for all sequencing modes
  • fq2: Reverse reads (R2) or UMI sequencing data filename(s), which are only required for PE and SPET with UMI modes
SampleName mode fq1 fq2
P1 SPETNOUMI R1-P1.fastq.gz
P2 PE R1-P2.fastq.gz R2-P2.fastq.gz
R SPET R1-R.fastq.gz R2-R.fastq.gz
S SPET R1-S.fastq.gz R2-S.fastq.gz

You can use a text editor, Excel, or Python script to create such a CSV file. For instance, in case of using Python to create the DataFrame in the example above, you can write a Python script named gen_samples.py:

import pandas as pd
import numpy as np

df = pd.DataFrame(columns=["SampleName", "mode", "fq1", "fq2"]) # required columns: don't need to change

df["SampleName"] = ["P1", "P2", "R", "S"]  # bulk names: don't need to change

df["mode"] = ["SPETNOUMI", "PE", "SPET", "SPET"]  # must be pe, se, spet ou spetnoumi (case-insensitive)

df["fq1"] = ["R1-P1.fastq.gz", "R1-P2.fastq.gz", "R1-R.fastq.gz", "R1-S.fastq.gz"]  # list of R1 filename(s)

df["fq2"] = [np.nan, "R2-P2.fastq.gz", "R2-R.fastq.gz", "R2-S.fastq.gz"]  # list of R2 filename(s)

df.to_csv("samples.csv", sep='\t', index=False)  # saving format and filename: don't need to change

and then run the following command:

python gen_samples.py

Note that the QTLseq analysis pipeline is capable of processing multiple input files for each sample simultaneously, if needed. To specify multiple input files for a sample in the sample file, simply separate them with a semicolon (;) in the fq1 and fq2 column. For example:

df["fq1"] = ["R1-P1.fastq.gz", 
             "R1-P2.fastq.gz", 
             "R1-R-1.fastq.gz;R1-R-2.fastq.gz", 
             "R1-S-1.fastq.gz;R1-S-2.fastq.gz"]

df["fq2"] = [np.nan, 
             "R2-P2.fastq.gz", 
             "R2-R-1.fastq.gz;R2-R-2.fastq.gz", 
             "R2-S-1.fastq.gz;R2-S-2.fastq.gz"]

IMPORTANT: Please make sure the order of input files specified in the fq1 column of the sample file need to corresponds correctly with the order of input files in the fq2. If you are working with SE or SPET without UMI, the fq2 column can be left empty.

Clone this wiki locally