Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

455 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Promoter-embedded TE Methylation (PeTEM) analyzer

PeTEM is designed to analyze the association between promoter-embedded TE methylation and neighboring gene expression. It integrates genome annotation, methylome, and transcriptome data to evaluate genome-wide correlations between TE methylation and gene expression, and identify TE–gene pairs showing coordinated methylation and expression changes across conditions.

image

Tutorial

See the tutorial for an example workflow.

Table of Contents

System Requirements

  • Runtime dependencies

    • Bash
    • Perl (v5.22.0+)
    • gzip / gunzip
    • awk / sort / uniq
  • Environment setup

    The following sections show the required environments and packages.
    In installation section, we provide three alternative methods for setting up the environment.

    👉 R version ≥ 4.2 (tested on 4.3.2)
    • optparse
    • dplyr
    • tidyr
    • zoo
    • reshape2
    • stringr
    • ggplot2
    • gplots
    • ggalluvial
    • ggpointdensity
    • ggbreak
    • RColorBrewer
    • viridis
    • rlang
    👉 Python version ≥ 3.8 (tested on 3.8.10)
    • pandas (≥ 1.2.4)
    👉 Bioinformatics tools
    • samtools (tested on 1.10)
    • bedtools (tested on v2.27.1)
    • wigToBigWig (bundled in this repository)
    • bigWigAverageOverBed (bundled in this repository)

Installation

  • Clone repository

    git clone https://github.com/PaoyangLab/PeTEM.git
    cd PeTEM
  • Download example data

    wget https://github.com/PaoyangLab/PeTEM/releases/download/Example_data/PeTEM_data.tar.gz
    tar -xzvf PeTEM_data.tar.gz
  • Set up environment

    Choose one of the following installation methods to set up the PeTEM environment.

    • Conda setup (👍recommended)

      conda env create -f environment.yml
      conda activate petem
      bash ./scripts/env_check.sh ##optional
    • Docker image

      docker build -t petem:local .
      docker run --rm petem:local --help
    • Local setup

      The script installs dependencies using apt-get, pip3 --user, and Rscript, then runs bash env_check.sh.
      If apt-get is not available the script prints the package list to install manually.

      bash ./scripts/setup.sh
      bash ./scripts/env_check.sh ##optional

Input Files

PeTEM integrates genome annotations, DNA methylation data, and expression data as input files. In PeTEM, modules 1 and 2 rely solely on annotation data, whereas the remaining modules additionally require methylation and expression data.

  • Genome Annotation

    • General features annotations file (GFF format)

      The genome annotation file contains genomic feature coordinates and hierarchical annotations.

      Format: (genomic.gff)

      Column Description
      seqid Sequence ID, chromosome, or scaffold name
      source Annotation source
      type Genomic feature type
      start Start coordinate (1-based)
      end End coordinate
      score Annotation score
      strand Strand information (+, -, or .)
      phase CDS phase information
      attributes Feature attributes and hierarchical annotation information

      Example:

      Chr1  Araport11  gene             3631  5899  .  +  .  ID=AT1G01010;Name=AT1G01010;full_name=NAC domain containing protein 1
      Chr1  Araport11  mRNA             3631  5899  .  +  .  ID=AT1G01010.1;Name=AT1G01010.1;Parent=AT1G01010
      Chr1  Araport11  CDS              3760  3913  .  +  0  ID=AT1G01010:CDS:1;Parent=AT1G01010.1
      Chr1  Araport11  exon             3631  3913  .  +  .  ID=AT1G01010:exon:1;Parent=AT1G01010.1
      Chr1  Araport11  five_prime_UTR   3631  3759  .  +  .  ID=AT1G01010:five_prime_UTR:1
      Chr1  Araport11  three_prime_UTR  5631  5899  .  +  .  ID=AT1G01010:three_prime_UTR:1
      
      ℹ️ Click for file sources

    • Genome index file (FASTA index)

      The genome FASTA index file provides the names and lengths of each chromosome. Users can create the file using: samtools faidx genome.fa
      Format: (genome.fa.fai)

      Column Description
      name Chromosome or sequence name
      length Sequence length (bp)
      offset Byte offset of the sequence in the FASTA file
      linebases Number of bases per sequence line
      linewidth Number of bytes per sequence line, including newline characters

      Example:

      Chr1  30427671  74         79  80
      Chr2  19698289  30812981   79  80
      Chr3  23459830  50760691   79  80
      Chr4  18585056  74517556   79  80
      Chr5  26975502  93337941   79  80
      ChrC  154478    120654981  79  80
      ChrM  367808    120811562  70  71
      
      ℹ️ Click for file sources

      FASTA index file is generated from genome fasta file using samtools:

      samtools faidx genome.fa
      

    • Transposable element coordinates

      The transposable element annotation file contains genomic coordinates and classification information for annotated transposable elements (TEs), including TE family and strand orientation.

      Format: (TE.txt)

      Column Description
      TE name Transposable element identifier
      chromosome Chromosome or scaffold name
      start Start coordinate (0-based)
      end End coordinate
      score Annotation score
      strand Strand information (+, -, or .)
      TE family Transposable element family classification

      Example:

      AT1TE00010  Chr1  11897  11976  0  +  LTR/Copia
      AT1TE00020  Chr1  16883  17009  0  -  RC/Helitron
      AT1TE00025  Chr1  17024  18924  0  +  RC/Helitron
      AT1TE00030  Chr1  18331  18642  0  -  DNA/HAT
      
      ℹ️ Click for file sources

      Processed TE annotation files from 18 species are available here

      Download files

      wget https://github.com/PaoyangLab/PeTEM/releases/download/TE_annotation/TE_files.tar.gz
      tar -xzvf TE_files.tar.gz

      Included species

      Kingdom Species File
      Animal Homo sapiens human_TE.txt
      Animal Mus musculus mouse_TE.txt
      Animal Danio rerio zebrafish_TE.txt
      Animal Drosophila melanogaster fruit_fly_TE.txt
      Plant Arabidopsis thaliana Arabidopsis_TE.txt
      Plant Oryza sativa rice_TE.txt
      Plant Zea mays maize_TE.txt
      Plant Glycine max soybean_TE.txt
      Fungi Botrytis cinerea Botrytis_cinerea_TE.txt
      Fungi Blumeria graminis Blumeria_graminis_TE.txt
      Fungi Colletotrichum higginsianum Colletotrichum_higginsianum_TE.txt
      Fungi Leptosphaeria maculans Leptosphaeria_maculans_TE.txt
      Fungi Melampsora larici-populina Melampsora_larici-populina_TE.txt
      Fungi Magnaporthe oryzae Magnaporthe_oryzae_TE.txt
      Fungi Microbotryum violaceum Microbotryum_violaceum_TE.txt
      Fungi Puccinia graminis f. sp. tritici Puccinia_graminis_TE.txt
      Fungi Sclerotinia sclerotiorum Sclerotinia_sclerotiorum_TE.txt
      Fungi Tuber melanosporum Tuber_melanosporum_TE.txt

  • Methylation Data

    CGmap is a file format for storing single-base resolution DNA methylation data.

    Important: CGmap file prefixes must match the condition names in the expression tables. For example, root_01.CGmap.gz and root_02.CGmap.gz correspond to the condition root.

    Format: (*.CGmap.gz files)

    Column Description
    chromosome Chromosome or scaffold name
    nucleotide Cytosine (C) or guanine (G) on the forward or reverse strand
    position Genomic position of the methylation site (1-based)
    context Methylation context (CG, CHG, or CHH)
    dinucleotide Dinucleotide sequence surrounding the methylation site
    methylation level Methylation level ranging from 0 to 1
    #C site Number of reads supporting methylated cytosine
    #C+T site Total sequencing depth at the site

    Example:

    Chr3  C  556  CG   CG  0.877551  43  49
    Chr3  G  557  CG   CG  0.787879  26  33
    Chr3  G  558  CHG  CC  0.405405  15  37
    Chr3  G  560  CHH  CA  0.102564   4  39
    
    ℹ️ Click for file sources

    CGmap files can be generated from WGBS data using:


  • Expression Data

    Expression data consists of gene expression (gene_expression.txt) and transposable element expression (TE_expression.txt) tables.

    Each row represents a gene or TE. The first columns contain the average expression level (RPKM) for each condition, followed by differential expression statistics, including log2 fold change (logFC), p-value, and false discovery rate (FDR) for pairwise condition comparisons.

    Column naming convention

    Column pattern Description Example
    condition_name Average expression level (RPKM) of a condition root, leaf
    logFC_condition1_condition2 Log2 fold change between two conditions logFC_root_leaf
    PValue_condition1_condition2 Statistical significance of differential expression PValue_root_leaf
    FDR_condition1_condition2 Multiple-testing adjusted p-value FDR_root_leaf

    Format:

    Column Description
    Row name Gene/Transposon identifier
    condition_name Average expression level (RPKM) of each condition
    logFC_condition1_condition2 Log2 fold change between two conditions
    PValue_condition1_condition2 Differential expression p-value
    FDR_condition1_condition2 Adjusted p-value (FDR)

    Example: (gene_expression.txt)

                root   leaf  logFC_root_leaf  PValue_root_leaf  FDR_root_leaf
    AT1G01010  11.64  10.93             0.09              0.60              1
    AT1G01020   6.39   5.95             0.10              0.62              1
    AT1G01030   0.65   1.01            -0.63              0.19           0.84
    

    Example: (TE_expression.txt)

                  root    leaf  logFC_root_leaf  PValue_root_leaf  FDR_root_leaf
    AT1TE00010  386.06  240.27             0.63              0.73              1
    AT1TE00020    0.00    0.00             0.00              1.00              1
    AT1TE00025    0.00    0.00             0.00              1.00              1
    
    ℹ️ Click for file sources

    Expression files can be generated from expression/count matrices using edgeR_for_expression.R bundled in this repository

    Rscript differential_expression.R \
      --eg gene_expression.txt \
      --et TE_expression.txt \
      -o results
    

Pipeline Modules

Users must run module 0 first to preprocess the input files before running modules 1–5. Modules 1–5 are independent and not sequential.

image

Module 0. Preprocessing

To prepare the inputs required for all downstream modules, module 0 generates promoter regions (promoter.bed) and integrates methylation and expression data.

  • Inputs:

    • genomic.gff, TE.txt, genome.fa.fai, gene_expression.txt, TE_expression.txt, *.CGmap.gz
  • Outputs:

    • OUTPUT_0_embedded_TE_gene_number.txt, promoter.bed, TE_overlap_promoter.bed, Tab_*.txt
    • PETEM_MODULE0_MANIFEST.json: Module 0 generates a JSON file for subsequent modules to simplify input paths.
  • Usage:

    ./petem --0 \
      -g <ANNOTATION_GFF3> \
      -t <TE_BED> \
      -eg <GENE_EXPRESSION> \
      -et <TE_EXPRESSION> \
      -f <GENOME_FAI> \
      -m <CGMAP_FILES> \
      -o <OUTPUT_DIR>
  • Parameters:

          Required       Description
    --0 Run module 0.
    -g Path to the genome annotation file in GFF3 format.
    -t Path to the transposable element annotation file in BED format.
    -eg Path to the gene expression table.
    -et Path to the transposable element expression table.
    -f Path to the genome FASTA index file (.fai). This is used to automatically generate gene, promoter, intergenic region, and other BED files.
    -m One or more CGmap files, separated by spaces. Compressed files such as .CGmap.gz are supported.
          Optional       Description
    -up Promoter region at the upstream of TSS. (Default: 1500)
    -dn Promoter region at the downstream of TSS. (Default: 500)
    -o Output directory for module 0 results.                                                                                                                 

Module 1. Profile TE genomic distribution

Analyze TE distribution across genomic features.

This reveals whether TEs preferentially accumulate in specific genomic regions.

  • Inputs:

    • genomic.gff, TE.txt, genome.fa.fai
    • promoter.bed (from Module 0)
  • Outputs:

    • OUTPUT_1_TE_distribution_enrichment.png, OUTPUT_1_TE_distribution_enrichment.png
  • Usage:

    ./petem --1 \
      -g <MODULE0_ANNOTATION_DIR> \
      -t <TE_BED> \
      -f <GENOME_FAI> \
      -o <OUTPUT_DIR>
  • Parameters:

          Required       Description
    --1 Run module 1. Equivalent to petem 0_preprocessing.
    --manifest Use manifest from module 0 to simplify input path.                                                                                   
          Optional       Description
    -g Path to the annotation directory generated by module 0 (module_0_annotation).
    -t Path to the transposable element annotation file in BED format.
    -f Path to the genome FASTA index file (.fai).
    -o Output directory for module 1 results.                                                                                                             

Module 2. Identify enriched promoter-embedded TE families

Identify enriched TE families overlapping with promoters.

This identifies TE families enriched in promoters and highlights candidates that may influence host gene expression.

  • Inputs:

    • TE.txt
    • promoter.bed (from Module 0)
  • Outputs:

    • OUTPUT_2_Promoter_embedded_TE_family_enrichment.png, OUTPUT_2_Promoter_embedded_TE_family.txt
  • Usage:

    ./petem --2 \
      -t <TE_BED> \
      -T <TE_FAMILY_TABLE> \
      -i <TE_OVERLAP_PROMOTER_BED> \
      -o <OUTPUT_DIR>
  • Parameters:

          Required       Description
    --2 Run module 2.
    --manifest Use manifest from module 0 to simplify input path.                                                                                   
          Optional       Description
    -t Path to the transposable element annotation file in BED format.
    -T Path to the transposable element family annotation table.
    -i Path to the TE_overlap_promoter.bed file generated by module 0.
    -o Output directory for module 2 results.                                                                                                                         

Module 3. Visualize TE methylation near gene

Visualize distance impact of TE methylation on gene expression.

This provides a genome-wide view of how nearby methylated TEs influence neighboring gene expression.

  • Inputs:

    • TE.txt, gene_expression.txt, TE_expression.txt
    • gene.bed(from Module 0)
  • Outputs:

    • OUTPUT_3_gene_TE_number.txt, OUTPUT_3_gene_proximal_TE_*.png
  • Usage:

    ./petem --3 \
      -g <MODULE0_ANNOTATION_DIR> \
      -t <TE_BED> \
      -eg <GENE_EXPRESSION> \
      -et <TE_EXPRESSION> \
      -m <CGMAP_FILES> \
      -o <OUTPUT_DIR> \
      -d <PROMOTER_DISTANCE> \
      -p <MIN_COVERAGE> \
      -w <WINDOW_SIZE> \
      -c <CONTROL_MODE> \
      -l <CONTROL_TE_CLASS>
  • Parameters:

          Required       Description
    --3 Run module 3.
    --manifest Use manifest from module 0 to simplify input path
    -m One or more CGmap files, separated by spaces. Compressed files such as .CGmap.gz are supported.                                                                                   
          Optional       Description
    -g Path to the annotation directory generated by module 0 (module_0_annotation).
    -t Path to the transposable element annotation file in BED format.
    -eg Path to the gene expression table.
    -et Path to the transposable element expression table.
    -o Output directory for module 3 results.
    -d The total up- and downstream distance (in bp) to consider gene-proximal TEs. (Default 4000 means ±4 kb around genes will be analyzed)
    -p The top/bottom X% of genes will be considered as the highly/lowly expressed genes. X must range from 1 to 50. (Default: 10, means top/bottom 10% of genes will be the highly/lowly expressed genes)
    -w For choosing average. Sliding window size (bp) used to smooth the TE methylation level curve. (Default: 100)
    -l There are several options to show the pattern of gene-proximal TE methylation, including (1) average TE methylation within each window, (2) linear regression line, (3) second-degree polynomial regression line, and (4) local regression line (average, linear, poly2, and poly). Default: average
    -CI To show the 95% CI on the plot or not. Default: no
    -border To show the border that TE methylation near lowly expressed genes is significantly higher than that near highly expressed genes on the plot or not. Default: no
    -unexp To include TEs with zero expression across all samples in the analysis. Default: yes
    -nTE To show the methylation level of non transposon sites around genes. Default: yes

Module 4. Calculate correlation coefficients

Correlate gene expression with TE/promoter methylation, and TE expression with TE methylation.

This module quantifies the genome-wide regulatory effects of TE methylation.

  • Inputs:

    • gene_expression.txt, TE_expression.txt
  • Outputs:

    • OUTPUT_4_geneexp/TEexp_*.png, OUTPUT_4_*_correlation_*.png
  • Usage:

    ./petem --4 \
      -eg <GENE_EXPRESSION> \
      -et <TE_EXPRESSION> \
      --module0-dir <MODULE0_OUTPUT_DIR> \
      -o <OUTPUT_DIR>
  • Parameters:

          Required       Description
    --4 Run module 4.
    --manifest Use manifest from module 0 to simplify input path
          Optional       Description
    -eg Path to the gene expression table.
    -et Path to the transposable element expression table.
    --module0-dir Path to the module 0 output directory. Required for loading methylation and annotation results generated in module 0.
    -o Output directory for module 4 results.
    --smooth Adjust the degree of curve smoothing. Values range from 1 to 5; 1 preserves the original curve shape (no smoothing), while 5 produces the smoothest curve. Default: 3.

Module 5. Identify associated TE and gene pairs

Examine the correlations between changes in TE methylation, TE expression, and gene expression across different conditions.

This identifies TE–gene pairs whose condition-specific expression changes are associated with TE methylation dynamics.

  • Inputs:

    • gene_expression.txt, TE_expression.txt
  • Outputs:

    • OUTPUT_5_*_scatter.png & .txt, OUTPUT_5_Q2/4_boxplot_*.png, OUTPUT_5_Q2/4_*.txt
  • Usage:

    ./petem --5 \
      --DEG <DEG_TABLE> \
      --DETE <DETE_TABLE> \
      --module0-dir <MODULE0_OUTPUT_DIR> \
      -o <OUTPUT_DIR>
  • Parameters:

          Required       Description
    --5 Run module 5.
    --manifest Use manifest from module 0 to simplify input path                                                                                   
          Optional       Description
    --DEG Path to the differentially expressed gene (DEG) table.
    --DETE Path to the differentially expressed transposable element (DETE) table.
    --module0-dir Path to the module 0 output directory. Required for loading methylation and annotation results generated in module 0.
    -stage Two condition names for pairwise comparison. The specified conditions must match the condition names in the expression tables.
    -o Output directory for module 5 results.
    --positive Specify the correlation direction to analyze. Set to yes for positive correlation analysis or no for negative correlation analysis. Default: yes.

About

Promoter-embedded TE Methylation (PeTEM) analyzer

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages