A curated mammography data set for use in computer-aided detection and diagnosis research

Sci Data. 2017 Dec 19;4:170177. doi: 10.1038/sdata.2017.177.


Published research results are difficult to replicate due to the lack of a standard evaluation data set in the area of decision support systems in mammography; most computer-aided diagnosis (CADx) and detection (CADe) algorithms for breast cancer in mammography are evaluated on private data sets or on unspecified subsets of public databases. This causes an inability to directly compare the performance of methods or to replicate prior results. We seek to resolve this substantial challenge by releasing an updated and standardized version of the Digital Database for Screening Mammography (DDSM) for evaluation of future CADx and CADe systems (sometimes referred to generally as CAD) research in mammography. Our data set, the CBIS-DDSM (Curated Breast Imaging Subset of DDSM), includes decompressed images, data selection and curation by trained mammographers, updated mass segmentation and bounding boxes, and pathologic diagnosis for training data, formatted similarly to modern computer vision data sets. The data set contains 753 calcification cases and 891 mass cases, providing a data-set size capable of analyzing decision support systems in mammography.

Publication types

  • Dataset
  • Research Support, N.I.H., Extramural

MeSH terms

  • Algorithms
  • Breast Neoplasms* / diagnosis
  • Breast Neoplasms* / prevention & control
  • Databases, Factual
  • Diagnosis, Computer-Assisted*
  • Female
  • Humans
  • Mammography*