Comparative proteomics reveals a significant bias toward alternative protein isoforms with conserved structure and function

Mol Biol Evol. 2012 Sep;29(9):2265-83. doi: 10.1093/molbev/mss100. Epub 2012 Mar 22.

Abstract

Advances in high-throughput mass spectrometry are making proteomics an increasingly important tool in genome annotation projects. Peptides detected in mass spectrometry experiments can be used to validate gene models and verify the translation of putative coding sequences (CDSs). Here, we have identified peptides that cover 35% of the genes annotated by the GENCODE consortium for the human genome as part of a comprehensive analysis of experimental spectra from two large publicly available mass spectrometry databases. We detected the translation to protein of "novel" and "putative" protein-coding transcripts as well as transcripts annotated as pseudogenes and nonsense-mediated decay targets. We provide a detailed overview of the population of alternatively spliced protein isoforms that are detectable by peptide identification methods. We found that 150 genes expressed multiple alternative protein isoforms. This constitutes the largest set of reliably confirmed alternatively spliced proteins yet discovered. Three groups of genes were highly overrepresented. We detected alternative isoforms for 10 of the 25 possible heterogeneous nuclear ribonucleoproteins, proteins with a key role in the splicing process. Alternative isoforms generated from interchangeable homologous exons and from short indels were also significantly enriched, both in human experiments and in parallel analyses of mouse and Drosophila proteomics experiments. Our results show that a surprisingly high proportion (almost 25%) of the detected alternative isoforms are only subtly different from their constitutive counterparts. Many of the alternative splicing events that give rise to these alternative isoforms are conserved in mouse. It was striking that very few of these conserved splicing events broke Pfam functional domains or would damage globular protein structures. This evidence of a strong bias toward subtle differences in CDS and likely conserved cellular function and structure is remarkable and strongly suggests that the translation of alternative transcripts may be subject to selective constraints.

Publication types

  • Research Support, N.I.H., Extramural
  • Research Support, Non-U.S. Gov't

MeSH terms

  • Alternative Splicing*
  • Amino Acid Sequence
  • Animals
  • Catalytic Domain
  • Drosophila
  • Genome
  • Humans
  • Mice
  • Models, Molecular
  • Molecular Sequence Annotation
  • Molecular Sequence Data
  • Nonsense Mediated mRNA Decay
  • Peptides / chemistry
  • Peptides / genetics
  • Proteasome Endopeptidase Complex / chemistry
  • Protein Biosynthesis
  • Protein Conformation
  • Protein Interaction Domains and Motifs
  • Protein Isoforms
  • Proteins / chemistry*
  • Proteins / genetics*
  • Proteins / metabolism
  • Proteomics*
  • Sequence Alignment

Substances

  • Peptides
  • Protein Isoforms
  • Proteins
  • Proteasome Endopeptidase Complex