Machine Learning Enables Accurate and Rapid Prediction of Active Molecules Against Breast Cancer Cells

Shuyun He; Duancheng Zhao; Yanle Ling; Hanxuan Cai; Yike Cai; Jiquan Zhang; Ling Wang

doi:10.3389/fphar.2021.796534

Machine Learning Enables Accurate and Rapid Prediction of Active Molecules Against Breast Cancer Cells

Front Pharmacol. 2021 Dec 17:12:796534. doi: 10.3389/fphar.2021.796534. eCollection 2021.

Authors

Shuyun He^{1

2}, Duancheng Zhao^{1

2}, Yanle Ling^{1

2}, Hanxuan Cai^{1

2}, Yike Cai³, Jiquan Zhang⁴, Ling Wang^{1

2}

Affiliations

¹ Guangdong Provincial Key Laboratory of Fermentation and Enzyme Engineering, Guangdong Provincial Engineering and Technology Research Center of Biopharmaceuticals, School of Biology and Biological Engineering, South China University of Technology, Guangzhou, China.
² Joint International Research Laboratory of Synthetic Biology and Medicine, Guangdong Provincial Engineering and Technology Research Center of Biopharmaceuticals, School of Biology and Biological Engineering, South China University of Technology, Guangzhou, China.
³ Center for Certification and Evaluation, Guangdong Drug Administration, Guangzhou, China.
⁴ State Key Laboratory of Functions and Applications of Medicinal Plants, College of Pharmacy, Guizhou Provincial Engineering Technology Research Center for Chemical Drug R&D, Guizhou Medical University, Guiyang, China.

Abstract

Breast cancer (BC) has surpassed lung cancer as the most frequently occurring cancer, and it is the leading cause of cancer-related death in women. Therefore, there is an urgent need to discover or design new drug candidates for BC treatment. In this study, we first collected a series of structurally diverse datasets consisting of 33,757 active and 21,152 inactive compounds for 13 breast cancer cell lines and one normal breast cell line commonly used in in vitro antiproliferative assays. Predictive models were then developed using five conventional machine learning algorithms, including naïve Bayesian, support vector machine, k-Nearest Neighbors, random forest, and extreme gradient boosting, as well as five deep learning algorithms, including deep neural networks, graph convolutional networks, graph attention network, message passing neural networks, and Attentive FP. A total of 476 single models and 112 fusion models were constructed based on three types of molecular representations including molecular descriptors, fingerprints, and graphs. The evaluation results demonstrate that the best model for each BC cell subtype can achieve high predictive accuracy for the test sets with AUC values of 0.689-0.993. Moreover, important structural fragments related to BC cell inhibition were identified and interpreted. To facilitate the use of the model, an online webserver called ChemBC (http://chembc.idruglab.cn/) and its local version software (https://github.com/idruglab/ChemBC) were developed to predict whether compounds have potential inhibitory activity against BC cells.

Keywords: breast cancer; graph neural networks; machine learning; molecular fingerprints; structural fragments.