C-P.17: Improving protein structure prediction with structure-aware multiple sequence alignments
Leveraging evolutionary information is central to modern protein structure prediction. As protein structure is more conserved than sequence throughout evolution, searching for distant protein homologs using structural information is a highly appealing strategy. However, until recently, performing high-throughput searches for distant structural homologs with low sequence identity was not feasible without access to large databases of structures. The Embedded Alphabet (TEA), a novel 20-letter 1D structural alphabet, addresses this challenge by enabling fast and sensitive detection of distant homologs without requiring prior structural information. In this work, we investigate whether structure-aware multiple sequence alignments (MSAs) generated with TEA can improve protein structure prediction by enhancing the evolutionary signal captured in the alignments. Structure-aware MSAs were constructed using Search with TEA against Many (STEAM), a tool built on the Foldseek framework and adapted for the TEA alphabet. We evaluated these MSAs on two datasets: the CASP13 dataset, containing query proteins with shallow MSAs, and viral proteins from BFVD, previously characterized by poor-quality MSAs using the default ColabFold DB. To evaluate the impact on structure prediction, we ran AlphaFold2 (AF2) in custom MSA mode using both STEAM-generated structure-aware MSAs and MMseqs2-generated MSAs. We present an evaluation of STEAM-generated MSAs and their effect on AF2 prediction quality across both datasets, with a particular focus on challenging low-homology and viral targets, where sequence-based searches alone may be insufficient.
Co-authors: Celia Ulrich, Lorenzo Pantolini, Janani Durairaj
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826