Early identification of people at risk of schistosomiasis infection is critical to interrupting disease transmission. We develop and validate an explainable machine learning prediction model that integrates demographic, behavioral, and environmental factors to identify these individuals. A total of 103,707 individuals were included to train and internally validate the model, and 16,574 individuals were used for external validation. The Random Forest (RF) model demonstrated the best discriminative performance among the five machine learning models evaluated. It accurately predicted schistosomiasis seropositivity in both internal validation (AUC = 0.943, F1 score = 0.809) and external validation (AUC = 0.897, F1 score = 0.770) and has been translated into a practical tool to support real-world application. Feature importance analysis indicated that the most significant predictors of schistosomiasis seropositivity included the presence of schistosomiasis symptoms, history of exposure to infected water, endemicity types of the village, gender, and village risk category. Furthermore, the SHapley Additive exPlanation (SHAP) method was employed to explain how these variables influence the prediction outcomes. This study provides a reference for early identification of high-risk populations and facilitates the translation of theoretical modeling studies into practical work applications.
Keywords: Machine learning; Prediction model; SHAP; Schistosomiasis; Seropositivity.
Copyright © 2025 Australian Society for Parasitology. Published by Elsevier Ltd. All rights reserved.