Tantiangco, Hanz (2026) Augmenting Generative De Novo Drug Design with Synthetic Data. PhD thesis, University of Sheffield.
Abstract
A recent trend in de novo drug design is the use of generative deep learning models, such as recurrent neural networks (RNNs). However, despite the promises of RNNs, they can be outperformed by traditional methods such as the Graph GA (Brown et al., 2019). Recent works have used randomised/enumerated SMILES as a form of data augmentation to expand the training data. This has been shown to improve the performance of RNNs for distribution learning tasks (Arús-Pous, Johansson, et al., 2019; Moret et al., 2020). However, the application of data augmentation in goal-directed generative design remains limited. This thesis describes the development of novel operators for the data augmentation of generative de novo design for goal-directed generation. The first step was to identify whether data augmentation of the primary/pretraining dataset improved model performance. Results indicate that, in general, introducing data augmentation to the primary training dataset improved goal-directed generation performance.
Next, data augmentation was investigated at different points of the fine-tuning loop in an RNN-LSTM. It was identified that subset loop augmentation performed the best. Subsequently, NLP-inspired operators were developed, which include atom-swap, atom-insert, atom-delete, and atom-fusion operators. Although these operators resulted in improvements of the design objectives, it was identified that they had a negative impact on the synthesisability of the generated molecules. Finally, methods to improve the synthesisability of the generated molecules were explored. These methods include chemistry-based operators (graph-fusion and bioisosteric substitution), masked atom-fusion (SAscore based- and functional group-based masking), and finetuning dataset filtering. Results show that bioisosteric substitution may perform the best overall as it is able to significantly increase the MPO score with minimal impact on the SAscore of the generated molecules relative to the other operators. Additionally, filtering of the finetuning dataset generally improved the SAscore of the generated molecules.
Metadata
| Supervisors: | Gillet, Val and Chen, Beining |
|---|---|
| Keywords: | Chemoinformatics, Generative AI, Language Models, De novo Drug Design, Synthetic Data, Data Augmentation |
| Awarding institution: | University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Social Sciences (Sheffield) > Information School (Sheffield) |
| Date Deposited: | 13 Aug 2026 14:14 |
| Last Modified: | 13 Aug 2026 14:14 |
| Open Archives Initiative ID (OAI ID): | oai:etheses.whiterose.ac.uk:39218 |
Download
Final eThesis - complete (pdf)
Filename: Final Thesis - Hanz.pdf
Licence:

This work is licensed under a Creative Commons Attribution NonCommercial NoDerivatives 4.0 International License
Export
Statistics
You do not need to contact us to get a copy of this thesis. Please use the 'Download' link(s) above to get a copy.
You can contact us about this thesis. If you need to make a general enquiry, please see the Contact us page.