» Articles » PMID: 17130148

NCBI Reference Sequences (RefSeq): a Curated Non-redundant Sequence Database of Genomes, Transcripts and Proteins

Overview
Specialty Biochemistry
Date 2006 Nov 30
PMID 17130148
Citations 1612
Authors
Affiliations
Soon will be listed here.
Abstract

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

Citing Articles

Genome mining the black-yeast Aureobasidium pullulans NRRL 62031 for biotechnological traits.

Xiao D, Driller M, Stein K, Blank L, Tiso T BMC Genomics. 2025; 26(1):244.

PMID: 40082747 PMC: 11905612. DOI: 10.1186/s12864-025-11395-2.


Genomic insights into the plasmidome of non-tuberculous mycobacteria.

Diricks M, Maurer F, Dreyer V, Barilar I, Utpatel C, Merker M Genome Med. 2025; 17(1):19.

PMID: 40038805 PMC: 11877719. DOI: 10.1186/s13073-025-01443-7.


TransHLA: a Hybrid Transformer model for HLA-presented epitope detection.

Lu T, Wang X, Nie W, Huo M, Li S Gigascience. 2025; 14.

PMID: 40036690 PMC: 11878767. DOI: 10.1093/gigascience/giaf008.


RNA virus diversity highlights the potential biosecurity threat posed by Antarctic krill.

Xu T, Zhao X, Loch T, Zhu J, Wang W, Wang X Mar Life Sci Technol. 2025; 7(1):96-109.

PMID: 40027325 PMC: 11871207. DOI: 10.1007/s42995-024-00270-w.


Understanding the genetic epidemiology of hereditary breast cancer in India using whole genome data from 1029 healthy individuals.

Vatsyayan A, Mathur P, Bhoyar R, Imran M, Senthivel V, Divakar M Cancer Causes Control. 2025; .

PMID: 40024972 DOI: 10.1007/s10552-025-01974-9.


References
1.
Benson D, Lipman D, Ostell J, Rapp B, Wheeler D . GenBank. Nucleic Acids Res. 1999; 28(1):15-8. PMC: 102453. DOI: 10.1093/nar/28.1.15. View

2.
Wheeler D, Barrett T, Benson D, Bryant S, Canese K, Chetvernin V . Database resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2006; 35(Database issue):D5-12. PMC: 1781113. DOI: 10.1093/nar/gkl1031. View

3.
Schuler G, Epstein J, Ohkawa H, Kans J . Entrez: molecular biology database and retrieval system. Methods Enzymol. 1996; 266:141-62. DOI: 10.1016/s0076-6879(96)66012-1. View

4.
Maglott D, Ostell J, Pruitt K, Tatusova T . Entrez Gene: gene-centered information at NCBI. Nucleic Acids Res. 2010; 39(Database issue):D52-7. PMC: 3013746. DOI: 10.1093/nar/gkq1237. View

5.
Tatusova T, Ostell J . Complete genomes in WWW Entrez: data representation and analysis. Bioinformatics. 1999; 15(7-8):536-43. DOI: 10.1093/bioinformatics/15.7.536. View