A systematic approach to modeling, capturing, and disseminating proteomics experimental data

Nat Biotechnol. 2003 Mar;21(3):247-54. doi: 10.1038/nbt0303-247.


Both the generation and the analysis of proteome data are becoming increasingly widespread, and the field of proteomics is moving incrementally toward high-throughput approaches. Techniques are also increasing in complexity as the relevant technologies evolve. A standard representation of both the methods used and the data generated in proteomics experiments, analogous to that of the MIAME (minimum information about a microarray experiment) guidelines for transcriptomics, and the associated MAGE (microarray gene expression) object model and XML (extensible markup language) implementation, has yet to emerge. This hinders the handling, exchange, and dissemination of proteomics data. Here, we present a UML (unified modeling language) approach to proteomics experimental data, describe XML and SQL (structured query language) implementations of that model, and discuss capture, storage, and dissemination strategies. These make explicit what data might be most usefully captured about proteomics experiments and provide complementary routes toward the implementation of a proteome repository.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Database Management Systems*
  • Databases, Protein*
  • Documentation / methods
  • Hypermedia
  • Information Dissemination / methods
  • Information Storage and Retrieval / methods*
  • Models, Molecular
  • Protein Conformation
  • Proteins / chemistry*
  • Proteins / genetics
  • Proteins / metabolism
  • Proteomics / methods*
  • Sequence Analysis, Protein / methods
  • Software
  • Software Design
  • User-Computer Interface


  • Proteins