Beware of the generic machine learning-based scoring functions in structure-based virtual screening

Chao Shen; Ye Hu; Zhe Wang; Xujun Zhang; Jinping Pang; Gaoang Wang; Haiyang Zhong; Lei Xu; Dongsheng Cao; Tingjun Hou

doi:10.1093/bib/bbaa070

Beware of the generic machine learning-based scoring functions in structure-based virtual screening

Brief Bioinform. 2021 May 20;22(3):bbaa070. doi: 10.1093/bib/bbaa070.

Authors

Chao Shen¹, Ye Hu¹, Zhe Wang¹, Xujun Zhang¹, Jinping Pang¹, Gaoang Wang¹, Haiyang Zhong¹, Lei Xu¹, Dongsheng Cao¹, Tingjun Hou¹

Affiliation

¹ Central South University, China.

PMID: 32484221
DOI: 10.1093/bib/bbaa070

Abstract

Machine learning-based scoring functions (MLSFs) have attracted extensive attention recently and are expected to be potential rescoring tools for structure-based virtual screening (SBVS). However, a major concern nowadays is whether MLSFs trained for generic uses rather than a given target can consistently be applicable for VS. In this study, a systematic assessment was carried out to re-evaluate the effectiveness of 14 reported MLSFs in VS. Overall, most of these MLSFs could hardly achieve satisfactory results for any dataset, and they could even not outperform the baseline of classical SFs such as Glide SP. An exception was observed for RFscore-VS trained on the Directory of Useful Decoys-Enhanced dataset, which showed its superiority for most targets. However, in most cases, it clearly illustrated rather limited performance on the targets that were dissimilar to the proteins in the corresponding training sets. We also used the top three docking poses rather than the top one for rescoring and retrained the models with the updated versions of the training set, but only minor improvements were observed. Taken together, generic MLSFs may have poor generalization capabilities to be applicable for the real VS campaigns. Therefore, it should be quite cautious to use this type of methods for VS.

Keywords: machine learning; machine learning-based scoring function; scoring function; virtual screening.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Datasets as Topic
Drug Discovery / methods*
Machine Learning*
Molecular Docking Simulation
Molecular Structure
Protein Binding
User-Computer Interface*