Skip to content

Output Description

Ha Trang Phung edited this page Aug 28, 2023 · 1 revision

After running the pipeline, you will find several output files and directories that contain the important results of each step in the analysis.Here is an overview of the main output components:

DAG (Directed Acyclic Graph) Diagram

The pipeline generates a DAG diagram that visualizes the workflow and the dependencies between the different steps. This diagram is useful for understanding the overall analysis flow. It was generated before the pipeline execution. To access the DAG diagram, you can find it in the repository where you work. The DAG diagram is typically stored as a PDF file named dag_rules.pdf or dag_samples.pdf.

Out repository

That contains all the output files for each step of the pipeline after execution.This practice allows for easy access and management of the results, making it simpler to review, analyze, and share the pipeline outcomes. Here's a brief explanation of 4 important sub-repository.

1. multiqc

The "multiqc" sub-repository likely contains the output files generated by running MultiQC, which aggregates the results from multiple QC tools. This can provide a comprehensive summary of quality reads after preprocessing. multiqc_report_fastqc.html file is a interactive HTML reports that you can open in a web browser to explore the quality metrics in a user-friendly manner.

2. mapped

In this sub-repository, you can expect the output files related to the "mapping" step in the pipeline. This may include alignment files (e.g., BAM files) and the de-duplicate summary for each sample.

3. variant

Here, you would find the output files generated during the "variant calling" step in the pipeline. This include VCF files containing identified SNPs and INDELs, as well as variant statistics.

As shown in the dag_rules.pdf DAG (Directed Acyclic Graph) diagram, the variant identification and filtering process is composed of several steps. Here is a summary of the steps involved:

  • Variant Identification: This step involves identifying variants (SNPs and INDELs) using a variant calling tool such as GATK. The output of this step is a VCF file containing all the identified variants, including raw SNPs and INDELs. Depending on the user's preference, the pipeline provides an option to separate SNPs and INDELs into two separate files. This can be done to analyze SNPs and INDELs separately in downstream analyses.

    • gatk_all.raw_indels_snps.vcf.gz: a compressed Variant Call Format (VCF) file that contains the raw SNPs and INDELs identified
    • gatk_all.score_raw_snps.csv: a csv file for facility visualization

    If this option split_SNPs_INDELs in config file is selected, two separate VCF files for SNPs and INDELs are generated.

    • gatk_all.raw_indels.vcf.gz: a VCF file contains the raw SNPs. -gatk_all.raw_snps.vcf.gz: a VCF file contains the raw INDELs.
  • Hard Filtering of SNPs: After the identification of variants, a hard filtering step is applied specifically for SNPs. This filtering process helps to remove low-quality or unreliable SNPs from the analysis based on various quality metrics provided by GATK. The output is a filtered VCF file containing high-confidence SNPs.

    • gatk_all.filtered.vcf.gz: a VCF file that contain the SNP filtered or SNP filtered with INDELs
    • gatk_all.filtered.stats.txt: a statistics file contains information about the number of SNPs retained and the number of SNPs lost after applying the hard filtering by GATK.
  • Filtering for Biallelic Variants: In this step, only biallelic variants are retained, and multi-allelic variants are filtered out. Biallelic variants are those with only two different alleles at a given genomic position, simplifying downstream analyses.

    • gatk_all.keep_biallele.vcf.gz: a VCF file that contain only the bialleic SNPs and INDELs.
    • gatk_all.keep_biallele.stats.txt: a statistics file contains information about the number of SNPs retained and the number of SNPs lost after applying the biallelic filtering criteria.
  • Filtering by Read Depth: Variants with low read depth (i.e., the number of reads covering the variant position) are filtered out. This step helps to ensure that variants with sufficient supporting read information are considered for subsequent analyses.

    • gatk_all.keep_filter_dp.vcf.gz: a VCF file contain the SNPs and INDELs after read depth filtering
    • gatk_all.keep_filter_dp.stats.txt: a statistics file contains information about the number of SNPs retained and the number of SNPs lost after applying the read depth filtering criteria.
  • Keeping Only Sites With Non-missing Genotypes: By removing sites with missing genotype data, the pipeline ensures that the analysis focuses only on variants where genotype information is available for all individuals, avoiding any biases that could arise due to incomplete data.

    • gatk_all.keep_snps_genotyped.vcf.gz: a VCF file that the SNPs/INDELs with missing genotype information are filtered out
    • gatk_all.keep_snps_genotyped.stats.txt: a statistics file contains information about the number of SNPs retained and the number of SNPs lost after applying the no missing genotyped filtering criteria.
  • Keeping Only Parent Polymorphic Sites: Here, the focus is on variants that are polymorphic in the parent samples. Polymorphic variants in the parents are crucial for QTL (Quantitative Trait Loci) analysis, as they can potentially influence trait variations in the offspring.

    • gatk_all.filter_P_snps.vcf.gz: a VCF that contain only SNPs/INDELs that exhibit polymorphism in the parental bulks are kept
    • gatk_all.filter_P_snps.stats.txt: a statistics file contains information about the number of SNPs retained and the number of SNPs lost after applying the polymorphic between the parents filtering criteria.

4. Rqtl

The "Rqtl" sub-repository likely contains the results of the statistical analysis performed using the QTLseqR package. This may include QTL mapping results, plots, and any other relevant outputs.

  • filtered.csv: a file contains a list of SNPs and INDELs that have undergone additional filtering and processing after the QTLseqR filtering step.
  • Takagi.jpg: delta SNP index plot
  • Gp: G' plot.

Clone this wiki locally