Sampling and Sampling Frames in Big Data Epidemiology

Curr Epidemiol Rep. 2019 Mar;6(1):14-22. doi: 10.1007/s40471-019-0179-y. Epub 2019 Feb 2.


Purpose of review: The 'big data' revolution affords the opportunity to reuse administrative datasets for public health research. While such datasets offer dramatically increased statistical power compared with conventional primary data collection, typically at much lower cost, their use also raises substantial inferential challenges. In particular, it can be difficult to make population inferences because the sampling frames for many administrative datasets are undefined. We reviewed options for accounting for sampling in big data epidemiology.

Recent findings: We identified three common strategies for accounting for sampling when the data available were not collected from a deliberately constructed sample: 1) explicitly reconstruct the sampling frame, 2) test the potential impacts of sampling using sensitivity analyses, and 3) limit inference to sample.

Summary: Inference from big data can be challenging because the impacts of sampling are unclear. Attention to sampling frames can minimize risks of bias.

Keywords: Big Data; Research Methods; Sampling; Sampling Frames; Secondary Data.