Retrosynthetic Reaction Prediction Using Neural Sequence-to-Sequence Models

Bowen Liu; Bharath Ramsundar; Prasad Kawthekar; Jade Shi; Joseph Gomes; Quang Luu Nguyen; Stephen Ho; Jack Sloane; Paul Wender; Vijay Pande

doi:10.1021/acscentsci.7b00303

Retrosynthetic Reaction Prediction Using Neural Sequence-to-Sequence Models

ACS Cent Sci. 2017 Oct 25;3(10):1103-1113. doi: 10.1021/acscentsci.7b00303. Epub 2017 Sep 5.

Authors

Bowen Liu¹, Bharath Ramsundar², Prasad Kawthekar², Jade Shi¹, Joseph Gomes¹, Quang Luu Nguyen¹, Stephen Ho¹, Jack Sloane¹, Paul Wender^{1

3}, Vijay Pande^{1

2

4}

Affiliations

¹ Department of Chemistry, Stanford University, Stanford, California 94305, United States.
² Department of Computer Science, Stanford University, Stanford, California 94305, United States.
³ Department of Chemical and Systems Biology, Stanford University, Stanford, California 94305, United States.
⁴ Department of Structural Biology, Stanford University, Stanford, California 94305, United States.

Abstract

We describe a fully data driven model that learns to perform a retrosynthetic reaction prediction task, which is treated as a sequence-to-sequence mapping problem. The end-to-end trained model has an encoder-decoder architecture that consists of two recurrent neural networks, which has previously shown great success in solving other sequence-to-sequence prediction tasks such as machine translation. The model is trained on 50,000 experimental reaction examples from the United States patent literature, which span 10 broad reaction types that are commonly used by medicinal chemists. We find that our model performs comparably with a rule-based expert system baseline model, and also overcomes certain limitations associated with rule-based expert systems and with any machine learning approach that contains a rule-based expert system component. Our model provides an important first step toward solving the challenging problem of computational retrosynthetic analysis.

Abstract

Grants and funding