NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Kim D Pruitt; Tatiana Tatusova; Donna R Maglott

doi:10.1093/nar/gkl842

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5. doi: 10.1093/nar/gkl842. Epub 2006 Nov 27.

Authors

Kim D Pruitt¹, Tatiana Tatusova, Donna R Maglott

Affiliation

¹ National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Rm 6An.12J, 45 Center Drive, Bethesda, MD 20892-6510, USA. pruitt@ncbi.nlm.nih.gov

Abstract

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

Publication types

Research Support, N.I.H., Intramural

MeSH terms

Amino Acid Sequence
Base Sequence
Databases, Nucleic Acid*
Databases, Protein*
Genome
Internet
National Library of Medicine (U.S.)
Quality Control
RNA, Messenger / chemistry
Reference Standards
Sequence Analysis, DNA / standards*
Sequence Analysis, Protein / standards*
Sequence Analysis, RNA / standards*
United States
User-Computer Interface

Substances

RNA, Messenger

Grants and funding

Intramural NIH HHS/United States