Skip to content

Substitution Models

minh edited this page Nov 9, 2015 · 78 revisions

IQ-TREE supports a wide range of substitution models, including advanced partition and mixture models. This guide gives a detailed information of all available models.

NOTICE: If you do not know which model to use, simply run IQ-TREE with -m TEST option, which automatically determines best-fit model for your data.

DNA models

Base substitution rates

IQ-TREE includes all common DNA models (ordered by complexity):

  • JC or JC69: equal substitution rates and equal base frequencies (Jukes and Cantor, 1969).
  • F81: equal rates but unequal base freq. (Felsenstein, 1981).
  • K80 or K2P: unequal transition/transversion rates and equal base freq. (Kimura, 1980).
  • HKY or HKY85: unequal transition/transversion rates and unequal base freq. (Hasegawa, Kishino and Yano, 1985).
  • TN or TN93: Like HKY but unequal purine/pyrimidine rates (Tamura and Nei, 1993).
  • TNe: Like TN but equal base freq.
  • K81 or K3P: three substitution types model and equal base freq. (Kimura, 1981).
  • K81u: Like K81 but unequal base freq.
  • TPM2: AC=AT, AG=CT, CG=GT and equal base freq.
  • TPM2u: Like TPM2 but unequal base freq.
  • TPM3: AC=CG, AG=CT, AT=GT and equal base freq.
  • TPM3u: Like TPM3 but unequal base freq.
  • TIM: transition model, AC=GT, AT=CG and unequal base freq.
  • TIMe: Like TIM but equal base freq.
  • TIM2: AC=AT, CG=GT and unequal base freq.
  • TIM2e: Like TIM2 but equal base freq.
  • TIM3: AC=CG, AT=GT and unequal base freq.
  • TIM3e: Like TIM3 but equal base freq.
  • TVM: transversion model, AG=CT and unequal base freq.
  • TVMe: Like TVM but equal base freq.
  • SYM: Symmetric model with unequal rates and equal base freq. (Zharkihk, 1994).
  • GTR: General time reversible model with unequal rates and unequal base freq. (Tavare, 1986).

Moreover, IQ-TREE supports arbitrarily restricted DNA model via a 6-digit code. The 6 digits define the equality for 6 nucleotide substitution rates: A-C, A-G, A-T, C-G, C-T and G-T. 010010 means that A-G rate is equal to C-T rate and the remaining four substitution rates are equal. Thus, 010010 is equivalent to K80 or HKY model (depending on whether base frequencies are equal or not). 123450 is equivalent to GTR or SYM model as there is no restriction defined by such 6-digit code.

If users want to fix model parameters, append the model name with a curly bracket {, followed by the comma-separated rate parameters, and a closing curly bracket }. For example, GTR{1.0,2.0,1.5,3.7,2.8,1.0} specifies 6 substitution rates A-C=1.0, A-G=2.0, A-T=1.5, C-G=3.7, C-T=2.8 and G-T=1.0.

Base frequencies

Users can specify three different kinds of base frequencies:

  • +F: empirical base frequencies. This is the default if model has unequal base freq.
  • +FQ: equal base frequencies.
  • +FO: optimized base frequencies by maximum-likelihood.

For example, GTR+FO optimizes base frequencies by ML whereas GTR+F (default) counts base frequencies from directly the alignment.

Finally, users can fix base frequencies with e.g. GTR+F{0.1,0.2,0.3,0.4} to fix the corresponding frequencies of A, C, G and T (must sum up to 1.0).

Protein models

Amino-acid exchange rate matrices

IQ-TREE supports all common empirical amino-acid exchange rate matrices:

If the matrix name does not match the above listed models, IQ-TREE assumes that it is a file containing AA exchange rates and frequencies in PAML format. It contains the lower diagonal part of the matrix and 20 AA frequencies, e.g.:

0.425093 
0.276818 0.751878 
0.395144 0.123954 5.076149 
2.489084 0.534551 0.528768 0.062556 
0.969894 2.807908 1.695752 0.523386 0.084808 
1.038545 0.363970 0.541712 5.243870 0.003499 4.128591 
2.066040 0.390192 1.437645 0.844926 0.569265 0.267959 0.348847 
0.358858 2.426601 4.509238 0.927114 0.640543 4.813505 0.423881 0.311484 
0.149830 0.126991 0.191503 0.010690 0.320627 0.072854 0.044265 0.008705 0.108882 
0.395337 0.301848 0.068427 0.015076 0.594007 0.582457 0.069673 0.044261 0.366317 4.145067 
0.536518 6.326067 2.145078 0.282959 0.013266 3.234294 1.807177 0.296636 0.697264 0.159069 0.137500 
1.124035 0.484133 0.371004 0.025548 0.893680 1.672569 0.173735 0.139538 0.442472 4.273607 6.312358 0.656604 
0.253701 0.052722 0.089525 0.017416 1.105251 0.035855 0.018811 0.089586 0.682139 1.112727 2.592692 0.023918 1.798853 
1.177651 0.332533 0.161787 0.394456 0.075382 0.624294 0.419409 0.196961 0.508851 0.078281 0.249060 0.390322 0.099849 0.094464 
4.727182 0.858151 4.008358 1.240275 2.784478 1.223828 0.611973 1.739990 0.990012 0.064105 0.182287 0.748683 0.346960 0.361819 1.338132 
2.139501 0.578987 2.000679 0.425860 1.143480 1.080136 0.604545 0.129836 0.584262 1.033739 0.302936 1.136863 2.020366 0.165001 0.571468 6.472279 
0.180717 0.593607 0.045376 0.029890 0.670128 0.236199 0.077852 0.268491 0.597054 0.111660 0.619632 0.049906 0.696175 2.457121 0.095131 0.248862 0.140825 
0.218959 0.314440 0.612025 0.135107 1.165532 0.257336 0.120037 0.054679 5.306834 0.232523 0.299648 0.131932 0.481306 7.803902 0.089613 0.400547 0.245841 3.151815 
2.547870 0.170887 0.083688 0.037967 1.959291 0.210332 0.245034 0.076701 0.119013 10.649107 1.702745 0.185202 1.898718 0.654683 0.296501 0.098369 2.188158 0.189510 0.249313 

0.079066 0.055941 0.041977 0.053052 0.012937 0.040767 0.071586 0.057337 0.022355 0.062157 0.099081 0.064600 0.022951 0.042302 0.044040 0.061197 0.053287 0.012066 0.034155 0.069147 

(this is an example of LG matrix taken from PAML package). Note that the amino-acid order in this file is:

 A   R   N   D   C   Q   E   G   H   I   L   K   M   F   P   S   T   W   Y   V
Ala Arg Asn Asp Cys Gln Glu Gly His Ile Leu Lys Met Phe Pro Ser Thr Trp Tyr Val

Amino-acid frequencies

By default, AA frequencies are given by the model. Users can change this with:

  • +F: empirical AA frequencies from the data.
  • +FO: ML optimized AA frequencies from the data.
  • +FQ: Equal AA frequencies.

Users can also specify AA frequencies with, e.g.:

+F{0.079066,0.055941,0.041977,0.053052,0.012937,0.040767,0.071586,0.057337,0.022355,0.062157,0.099081,0.064600,0.022951,0.042302,0.044040,0.061197,0.053287,0.012066,0.034155,0.069147}

(example corresponds to AA frequencies of LG matrix).

Codon models

To use codon model one should use option -st CODON, which implicitly assumes standard genetic code. You can change to other genetic code with following options:

Option Genetic code
-st CODON1 The Standard Code (same as -st CODON)
-st CODON2 The Vertebrate Mitochondrial Code
-st CODON3 The Yeast Mitochondrial Code
-st CODON4 The Mold, Protozoan, and Coelenterate Mitochondrial Code and the Mycoplasma/Spiroplasma Code
-st CODON5 The Invertebrate Mitochondrial Code
-st CODON6 The Ciliate, Dasycladacean and Hexamita Nuclear Code
-st CODON9 The Echinoderm and Flatworm Mitochondrial Code
-st CODON10 The Euplotid Nuclear Code
-st CODON11 The Bacterial, Archaeal and Plant Plastid Code
-st CODON12 The Alternative Yeast Nuclear Code
-st CODON13 The Ascidian Mitochondrial Code
-st CODON14 The Alternative Flatworm Mitochondrial Code
-st CODON16 Chlorophycean Mitochondrial Code
-st CODON21 Trematode Mitochondrial Code
-st CODON22 Scenedesmus obliquus Mitochondrial Code
-st CODON23 Thraustochytrium Mitochondrial Code
-st CODON24 Pterobranchia Mitochondrial Code
-st CODON25 Candidate Division SR1 and Gracilibacteria Code

(the IDs correspond to see specification at http://www.ncbi.nlm.nih.gov/Taxonomy/Utils/wprintgc.cgi).

Codon substitution rates

IQ-TREE supports several codon models:

The last three models (ECMK07, ECMrest or ECMS05) are called empirical codon models, whereas the others are called mechanistic codon models.

IQ-TREE also supports combined empirical-mechanistic codon models, where the name of an empirical model and a mechanistic model are combined with an underscore separator (_). For example:

  • ECMK07_GY2K: The combined ECMK07 and GY2K model, with the rate entries being multiplication of the two corresponding rate matrices.

Thus, there can be many such combinations.

If the model name does not match the above listed models, IQ-TREE assumes that it is a file containing codon exchange rates and frequencies in PAML format. It contains the lower diagonal part of the matrix and codon frequencies. For an example, see http://www.ebi.ac.uk/goldman/ECM/.

NOTICE: Branch lengths under codon models are interpreted as number of nucleotide substitutions per codon site. Thus, they are typically 3 times longer than under DNA models.

Codon frequencies

IQ-TREE supports the following codon frequencies:

  • +F: empirical codon frequencies counted from the data.
  • +FQ: equal codon frequencies.
  • +F1X4: unequal nucleotide frequencies but equal nt frequencies over three codon positions.
  • +F3X4: unequal nucleotide frequencies and unequal nt frequencies over three codon positions.

If not specified, the default codon frequency will be +F3X4 for MG-type models, +F for GY-type models and given by the model for empirical codon models.

Binary and morphological models

The binary alignments should contain state 0 and 1. Whereas for morphological data, the valid states are 0 to 9 and A to Z.

  • JC2: Jukes-Cantor type model for binary data.
  • GTR2: general time reversible model for binary data.
  • MK: Jukes-Cantor type model for morphological data.
  • ORDERED: allowing exchange of neighboring states only.

Except for GTR2 that has unequal state frequencies, all other models have equal state frequencies.

NOTICE: If morphological alignments do not contain constant sites (typically the case), then an ascertainment bias correction model (+ASC) should be applied to correct the branch lengths for the absence of constant sites.

Ascertainment bias correction

An ascertainment bias correction (+ASC) model (Lewis, 2001) should be applied if the alignment does not contain constant sites (such as morphological or SNPs data). For example:

  • MK+ASC: for morphological data.
  • GTR+ASC: for SNPs data.

+ASC will correct the likelihood conditioned on variable sites. Without +ASC, the branch lengths might be overestimated.

Rate heterogeneity across sites

Partition models

Mixture models

Customized models

Clone this wiki locally