Self-tuned healthy homogeneous core: Addressing heterogeneities in biomedical datasets

Artif Intell Med. 2026 Jun 5:180:103464. doi: 10.1016/j.artmed.2026.103464. Online ahead of print.

Abstract

Heterogeneity within healthy biomedical populations can degrade disease-classification performance by distorting the decision boundary between healthy and diseased cohorts. In this study, the healthy cohort is decomposed into two subtypes: borderline healthy samples, defined operationally as healthy-labeled subjects that lie closer to the diseased distribution in the learned feature space, and a representative subset termed the self-tuned Homogeneous Healthy Core (H2C), characterized by high intra-class homogeneity. Using an information-theoretic framework, it is shown that training on H2C yields a tighter classification-error bound than training on the full heterogeneous healthy cohort. To operationalize this idea, a healthy-only kernel density estimation procedure with bootstrap-stability-based self-tuning is developed and evaluated on three biomedical case studies: myocardial infarction classification on PTB-XL+, arrhythmia detection, and an institutional migraine dataset. Compared with the full-healthy baseline and simpler healthy-subset baselines, H2C consistently shows strong performance on the subset-gated evaluation view in terms of F1, precision, recall, and accuracy, with the clearest advantage under the balanced training setup. For example, under the balanced training setup and subset-gated evaluation view, H2C improves mean F1 from 0.899 to 0.952 on PTB-XL+, from 0.693 to 0.854 on Arrhythmia, and from 0.643 to 0.741 on Migraine relative to the full-healthy baseline. These results show that curating a stable healthy reference subset can improve downstream classification behavior and support the construction of smaller, higher-quality healthy reference cohorts for biomedical modeling. More broadly, H2C provides a self-tuned and classifier-agnostic strategy for healthy-cohort curation, an important but often overlooked problem in biomedical machine learning.

Keywords: Biomedical data heterogeneity; Clinical classification; Label-noise mitigation; Sample selection; Self-tuned Homogeneous Healthy Core.