Comprehensive analysis of single molecule sequencing-derived complete genome and whole transcriptome of Hyposidra talaca nuclear polyhedrosis virus

Sci Rep. 2018 Jun 12;8(1):8924. doi: 10.1038/s41598-018-27084-y.


We sequenced the Hyposidra talaca NPV (HytaNPV) double stranded circular DNA genome using PacBio single molecule sequencing technology. We found that the HytaNPV genome is 139,089 bp long with a GC content of 39.6%. It encodes 141 open reading frames (ORFs) including the 37 baculovirus core genes, 25 genes conserved among lepidopteran baculoviruses, 72 genes known in baculovirus, and 7 genes unique to the HytaNPV genome. It is a group II alphabaculovirus that codes for the F protein and lacks the gp64 gene found in group I alphabaculovirus viruses. Using RNA-seq, we confirmed the expression of the ORFs identified in the HytaNPV genome. Phylogenetic analysis showed HytaNPV to be closest to BusuNPV, SujuNPV and EcobNPV that infect other tea pests, Buzura suppressaria, Sucra jujuba, and Ectropis oblique, respectively. We identified repeat elements and a conserved non-coding baculovirus element in the genome. Analysis of the putative promoter sequences identified motif consistent with the temporal expression of the genes observed in the RNA-seq data.

MeSH terms

  • Amino Acid Sequence
  • Animals
  • Base Sequence
  • Genes, Viral / genetics
  • Genome, Viral / genetics*
  • Larva / virology
  • Moths / virology*
  • Nucleopolyhedroviruses / classification
  • Nucleopolyhedroviruses / genetics*
  • Nucleopolyhedroviruses / physiology
  • Open Reading Frames / genetics
  • Phylogeny
  • Sequence Homology, Amino Acid
  • Sequence Homology, Nucleic Acid
  • Transcriptome / genetics*
  • Whole Genome Sequencing / methods*