Improved Machine Learning-Based Predictive Models for Breast Cancer Diagnosis

Abdur Rasool; Chayut Bunterngchit; Luo Tiejian; Md Ruhul Islam; Qiang Qu; Qingshan Jiang

doi:10.3390/ijerph19063211

Improved Machine Learning-Based Predictive Models for Breast Cancer Diagnosis

Int J Environ Res Public Health. 2022 Mar 9;19(6):3211. doi: 10.3390/ijerph19063211.

Authors

Abdur Rasool^{1

2}, Chayut Bunterngchit^{1

3}, Luo Tiejian¹, Md Ruhul Islam⁴, Qiang Qu², Qingshan Jiang²

Affiliations

¹ University of Chinese Academy of Sciences, Beijing 101408, China.
² Shenzhen Key Lab for High Performance Data Mining, Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China.
³ State Key Laboratory of Management and Control for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China.
⁴ Department of Electrical Engineering and Computer Science, University of Stavanger, 4044 Stavanger, Norway.

Abstract

Breast cancer death rates are higher than any other cancer in American women. Machine learning-based predictive models promise earlier detection techniques for breast cancer diagnosis. However, making an evaluation for models that efficiently diagnose cancer is still challenging. In this work, we proposed data exploratory techniques (DET) and developed four different predictive models to improve breast cancer diagnostic accuracy. Prior to models, four-layered essential DET, e.g., feature distribution, correlation, elimination, and hyperparameter optimization, were deep-dived to identify the robust feature classification into malignant and benign classes. These proposed techniques and classifiers were implemented on the Wisconsin Diagnostic Breast Cancer (WDBC) and Breast Cancer Coimbra Dataset (BCCD) datasets. Standard performance metrics, including confusion matrices and K-fold cross-validation techniques, were applied to assess each classifier's efficiency and training time. The models' diagnostic capability improved with our DET, i.e., polynomial SVM gained 99.3%, LR with 98.06%, KNN acquired 97.35%, and EC achieved 97.61% accuracy with the WDBC dataset. We also compared our significant results with previous studies in terms of accuracy. The implementation procedure and findings can guide physicians to adopt an effective model for a practical understanding and prognosis of breast cancer tumors.

Keywords: breast cancer diagnosis; data exploratory techniques; machine learning models; tumors classification.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Algorithms
Breast
Breast Neoplasms* / diagnosis
Female
Humans
Machine Learning
Support Vector Machine