Beware of the generic machine learning-based scoring functions in structure-based virtual screening

descriptionPublicationkeyboard_double_arrow_right Article 02 Jun 2020 English Publisher:Oxford University Press (OUP)Journal:Briefings in Bioinformatics, volume 22 (eissn: 1477-4054,

Copyright policy )

Authors: Chao Shen 0008; Ye Hu; Zhe Wang 0041; Xujun Zhang; Jinping Pang; Gaoang Wang; Haiyang Zhong; +3 Authors

doi: 10.1093/bib/bbaa070

pmid: 32484221

Beware of the generic machine learning-based scoring functions in structure-based virtual screening

- Summary
- Subjects
- Metrics

Abstract

Abstract Machine learning-based scoring functions (MLSFs) have attracted extensive attention recently and are expected to be potential rescoring tools for structure-based virtual screening (SBVS). However, a major concern nowadays is whether MLSFs trained for generic uses rather than a given target can consistently be applicable for VS. In this study, a systematic assessment was carried out to re-evaluate the effectiveness of 14 reported MLSFs in VS. Overall, most of these MLSFs could hardly achieve satisfactory results for any dataset, and they could even not outperform the baseline of classical SFs such as Glide SP. An exception was observed for RFscore-VS trained on the Directory of Useful Decoys-Enhanced dataset, which showed its superiority for most targets. However, in most cases, it clearly illustrated rather limited performance on the targets that were dissimilar to the proteins in the corresponding training sets. We also used the top three docking poses rather than the top one for rescoring and retrained the models with the updated versions of the training set, but only minor improvements were observed. Taken together, generic MLSFs may have poor generalization capabilities to be applicable for the real VS campaigns. Therefore, it should be quite cautious to use this type of methods for VS.

Related Organizations

Central South University
China (People's Republic of)

Keywords

Machine Learning, Molecular Docking Simulation, User-Computer Interface, Molecular Structure, Drug Discovery, Datasets as Topic, Protein Binding

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	63
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 1%

Found an issue? Give us feedback

63

Top 1%

Top 10%

Top 1%

hybrid

Fields of Science (3) View all

medical and health sciences

basic medicine

Fields of Science

medical and health sciences

basic medicine

View all