Phylogeny-aware gap placement prevents errors in sequence alignment and evolutionary analysis

Science. 2008 Jun 20;320(5883):1632-5. doi: 10.1126/science.1158395.

Abstract

Genetic sequence alignment is the basis of many evolutionary and comparative studies, and errors in alignments lead to errors in the interpretation of evolutionary information in genomes. Traditional multiple sequence alignment methods disregard the phylogenetic implications of gap patterns that they create and infer systematically biased alignments with excess deletions and substitutions, too few insertions, and implausible insertion-deletion-event histories. We present a method that prevents these systematic errors by recognizing insertions and deletions as distinct evolutionary events. We show theoretically and practically that this improves the quality of sequence alignments and downstream analyses over a wide range of realistic alignment problems. These results suggest that insertions and sequence turnover are more common than is currently thought and challenge the conventional picture of sequence evolution and mechanisms of functional and structural changes.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Algorithms*
  • Computer Simulation
  • Evolution, Molecular*
  • HIV / chemistry
  • HIV / classification
  • HIV / genetics
  • HIV Envelope Protein gp120 / chemistry*
  • HIV Envelope Protein gp120 / genetics
  • HIV-1 / chemistry
  • HIV-1 / classification
  • HIV-1 / genetics
  • Membrane Glycoproteins / chemistry*
  • Membrane Glycoproteins / genetics
  • Mutagenesis, Insertional
  • Phylogeny*
  • Sequence Alignment / methods*
  • Sequence Deletion
  • Simian Immunodeficiency Virus / chemistry
  • Simian Immunodeficiency Virus / classification
  • Simian Immunodeficiency Virus / genetics
  • Viral Envelope Proteins / chemistry*
  • Viral Envelope Proteins / genetics

Substances

  • HIV Envelope Protein gp120
  • Membrane Glycoproteins
  • Viral Envelope Proteins
  • gp120 protein, Simian immunodeficiency virus