B-P.50: Reading TEA-Leaves for de novo protein design
Lisa Brandenburg
Department of Biosystems, Science and Engineering, ETH Zürich
Lorenzo Pantolini
Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics
Basile Wicky
Department of Biosystems, Science and Engineering, ETH Zürich; SIB Swiss Institute of Bioinformatics
Janani Durairaj
Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics
Keywords
de novo protein design, protein language models, Monte Carlo simulations
[sponser-meet-now-chat][/sponser-meet-now-chat]
The field of de novo protein design aims to produce novel protein sequences or even novel structures with improved functional properties for biotechnological applications. The breakthrough in protein folding methods opened ways to evaluate whether sequences fold into intended structures in silico, which catalysed the protein design progress. Orthogonally, we witnessed the appearance of protein language models (pLMs) encoding biological properties in their embeddings. However, pLM-based iterative optimization approaches requiring repeated structure evaluations remain computationally expensive, limiting their practical throughput.
We propose that the structure-informed alphabet derived from pLM embeddings (TEA) could address this bottleneck. Therefore, we present TEA-Leaves method leveraging TEA alphabet to efficiently guide the Markov chain Monte Carlo sampling to generate de novo sequences folding into template structures.
We showcase TEA-Leaves functionality in several scenarios by generating: de novo sequences folding into a diverse set of monomeric structures and sequence-similar but structure-dissimilar pairs of proteins.
In addition to structure prediction confidence metrics from ColabFold, we used several in silico quality metrics to select designs for experimental validation: values of TEA-Leaves losses, solubility, and aggregation scores. Namely, we selected designs with sequence novelty relative to natural proteins to assess the ability of TEA-Leaves to explore sequence space beyond homology adjacency. The quality scoring filters were collected in a pipeline Sieve to prioritize the candidates for the downstream experimental validation of expressibility, solubility, and oligomerization propensity.
Altogether, using pLM-derived structural signals we aim to probe the determinants of protein folding outside of the homology barrier.
Co-authors: Lisa Brandenburg, Lorenzo Pantolini, Basile Wicky, Janani Durairaj
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826