A-P.52: A Machine Learning Framework for High-Performance Prediction of Enzyme-Substrate Interactions Using Advanced Protein Embeddings and Molecular Fingerprints
Motivation: Enzymes catalyze biochemical reactions with remarkable specificity and efficiency, yet predicting enzyme–substrate interactions remains challenging due to limited experimental data, incomplete annotations, and the scarcity of validated non-binding pairs.
Traditional computational approaches, such as docking and molecular dynamics, are resource-intensive and poorly scalable, while many machine learning (ML) models struggle to generalize to chemically diverse or previously unseen molecules.
Results: We present a novel ML framework that integrates experimentally validated enzyme–substrate pairs with advanced feature representations.
Our approach combines protein embeddings from the pre-trained language model ESM2 with NPClassifierFP molecular fingerprints to capture functional and structural properties of enzymes and substrates.
A gradient boosting classifier (XGBoost) is trained using rigorous data partitioning with the GraphPart method to prevent data leakage and ensure robust generalization across similarity thresholds. Benchmarking against state-of-the-art models (ESP, ProSmith) and similarity-based approaches (BLASTp + Tanimoto similarity) demonstrates that our method consistently outperforms existing approaches, achieving high F1 scores and Matthews Correlation Coefficients (MCC).
Notably, the model maintains strong predictive performance under stringent protein (40–80% sequence identity) and compound (20–80% Tanimoto similarity) similarity constraints, and generalizes well to chemically distinct, previously unseen molecules.
Co-authors: João Ribeiro, Ana Monteiro, Dick de Ridder, Oscar Dias
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826