On the utility of pooling biological samples in microarray experiments

Proc Natl Acad Sci U S A. 2005 Mar 22;102(12):4252-7. doi: 10.1073/pnas.0500607102. Epub 2005 Mar 8.


Over 15% of the data sets catalogued in the Gene Expression Omnibus Database involve RNA samples that have been pooled before hybridization. Pooling affects data quality and inference, but the exact effects are not yet known because pooling has not been systematically studied in the context of microarray experiments. Here we report on the results of an experiment designed to evaluate the utility of pooling and the impact on identifying differentially expressed genes. We find that inference for most genes is not adversely affected by pooling, and we recommend that pooling be done when fewer than three arrays are used in each condition. For larger designs, pooling does not significantly improve inferences if few subjects are pooled. The realized benefits in this case do not outweigh the price paid for loss of individual specific information. Pooling is beneficial when many subjects are pooled, provided that independent samples contribute to multiple pools.

Publication types

  • Comparative Study
  • Research Support, U.S. Gov't, P.H.S.

MeSH terms

  • Analysis of Variance
  • Animals
  • Female
  • Gene Expression Profiling / methods
  • Gene Expression Profiling / statistics & numerical data
  • Nicotinic Acids / pharmacology
  • Oligonucleotide Array Sequence Analysis / methods*
  • Oligonucleotide Array Sequence Analysis / statistics & numerical data
  • RNA / genetics
  • Rats
  • Rats, Inbred WF
  • Retinoid X Receptors / agonists
  • Tetrahydronaphthalenes / pharmacology


  • Nicotinic Acids
  • Retinoid X Receptors
  • Tetrahydronaphthalenes
  • RNA
  • LG 100268

Associated data

  • OMIM/GSE2331