Genomic determinants of pathogenicity in SARS-CoV-2 and other human coronaviruses

Proc Natl Acad Sci U S A. 2020 Jun 30;117(26):15193-15199. doi: 10.1073/pnas.2008176117. Epub 2020 Jun 10.


Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) poses an immediate, major threat to public health across the globe. Here we report an in-depth molecular analysis to reconstruct the evolutionary origins of the enhanced pathogenicity of SARS-CoV-2 and other coronaviruses that are severe human pathogens. Using integrated comparative genomics and machine learning techniques, we identify key genomic features that differentiate SARS-CoV-2 and the viruses behind the two previous deadly coronavirus outbreaks, SARS-CoV and Middle East respiratory syndrome coronavirus (MERS-CoV), from less pathogenic coronaviruses. These features include enhancement of the nuclear localization signals in the nucleocapsid protein and distinct inserts in the spike glycoprotein that appear to be associated with high case fatality rate of these coronaviruses as well as the host switch from animals to humans. The identified features could be crucial contributors to coronavirus pathogenicity and possible targets for diagnostics, prognostication, and interventions.

Keywords: COVID-19; coronaviruses; nucleocapsid; pathogenicity; spike protein.

Publication types

  • Research Support, N.I.H., Extramural
  • Research Support, N.I.H., Intramural
  • Research Support, Non-U.S. Gov't

MeSH terms

  • Animals
  • Betacoronavirus / classification
  • Betacoronavirus / genetics*
  • Betacoronavirus / pathogenicity
  • Evolution, Molecular*
  • Genome, Viral*
  • Host Specificity
  • Humans
  • Machine Learning
  • Middle East Respiratory Syndrome Coronavirus / classification
  • Middle East Respiratory Syndrome Coronavirus / genetics
  • Middle East Respiratory Syndrome Coronavirus / pathogenicity
  • Mutagenesis, Insertional
  • Nuclear Localization Signals / genetics
  • Nucleocapsid Proteins / chemistry
  • Nucleocapsid Proteins / genetics*
  • Phylogeny
  • SARS-CoV-2
  • Sequence Homology
  • Spike Glycoprotein, Coronavirus / chemistry
  • Spike Glycoprotein, Coronavirus / genetics*
  • Virulence / genetics


  • Nuclear Localization Signals
  • Nucleocapsid Proteins
  • Spike Glycoprotein, Coronavirus