-
Notifications
You must be signed in to change notification settings - Fork 37
Home
The program gffcompare can be used to compare, merge, annotate and estimate accuracy of one or more GFF files (the "query" files), when compared with a reference annotation (also provided as GFF/GTF). A more detailed documentation for the program and its output files can be found here (gffcompare documentation page)
Gffcompare can be used to evaluate and compare the accuracy of transcript assemblers - in terms of their structural correctness (exon/intron coordinates). This assessment can even be performed in case of more generic "transcript discovery" programs like gene finders. The best way to do this would be to use a simulated data set (where the "reference annotation" is also the set of the expressed transcripts being simulated), but for well annotated reference genomes (human, mouse etc.), gffcompare can be used to evaluate and compare the general accuracy of isoform discovery programs on a real data set, using just the known (reference) annotation of that genome.
As a practical example, let's assume we ran both Cufflinks and StringTie on a mouse RNA-Seq sample
and we want to compare the overall accuracy of the two programs. In order to compare the baseline
de novo transcript assembly accuracy, both Cufflinks and StringTie should be run without using any
reference annotation data (i.e. no -G or -g options were used).
Assuming that Cufflinks' transcript assembly output file name is cufflinks_asm.gtf and StringTie's
output is in stringtie_asm.gtf, while the reference annotation would be in a file called mm10.gff,
the gffcompare commands would be:
gffcompare -R -r mm10.gff -o cuffcmp cufflinks_asm.gtf
gffcompare -R -r mm10.gff -o strtcmp stringtie_asm.gtf
The -R option is used here in order to adjust the sensitivity calculation as to only consider the
"expressed" genes, which are those reference genes for which gffcompare found at least one
overlapping transfrag in the given assembly data (*_asm.gtf file). (Of course this option would
not be needed in the case of simulated RNA-Seq experiments where the reference transcripts would be
all "expressed"). Multiple output files will be generated by gffcompare, with the given prefix - in
the example above, for the stringtie assemblies, the output files will be:
strtcmp.combined.gtf
strtcmp.loci
strtcmp.stats
strtcmp.tracking
strtcmp.stringtie_asm.gtf.refmap
strtcmp.stringtie_asm.gtf.tmap
Comparing the transcript assembly accuracy of the two programs is done by looking at the Sensitivity and Precision values in the *.stats output files for each program (strtcmp.stats vs cuffcmp.stats in the example above). That output looks like this (partially):
#= Summary for dataset: stringtie_asm.gtf
# Query mRNAs : 23555 in 17628 loci (17231 multi-exon transcripts)
# (3731 multi-transcript loci, ~1.3 transcripts per locus)
# Reference mRNAs : 16628 in 12062 loci (15850 multi-exon)
# Super-loci w/ reference transcripts: 11552
#-----------------| Sensitivity | Precision |
Base level: 82.4 | 76.5 |
Exon level: 81.2 | 82.9 |
Intron level: 86.1 | 94.8 |
Intron chain level: 56.9 | 52.4 |
Transcript level: 55.2 | 38.9 |
Locus level: 70.1 | 48.0 |
- project on Github
- source package: gffcompare-0.12.10.tar.gz
- precompiled Linux x86_64 binary (statically built on CentOS 5) : gffcompare-0.12.10.Linux_x86_64.tar.gz
- precompiled Apple OS X binary (for OS X 10.7 and higher): gffcompare-0.12.10.OSX_x86_64.tar.gz
Pertea G and Pertea M. GFF Utilities: GffRead and GffCompare. F1000Research 2020, 9:304 DOI:10.12688/f1000research.23297.1