C-P.41: Generating and evaluating coevolution and phylogeny-aware MSAs
Molecular phylogenetic inference involves inferring evolutionary relationships from a multiple sequence alignment (MSA). Generally, a single phylogenetic tree or tree distribution is inferred for one MSA at a time by methods like maximum likelihood and Bayesian inference. Alternatively, a generalizable, supervised learning approach requires training data in the form of “MSA-ground truth tree†pairs. Phylogenetics lacks actual ground-truths, so training data needs to be generated by simulating evolution along a known tree to yield leaf sequences constituting an MSA. The tree is then by definition the ground truth tree for the generated MSA. However, the choice of the simulation method is crucial for obtaining realistic MSAs from known trees. Conventionally, simulation is performed using models that assume independent evolution across all sites in a sequence, while sites that are functionally coupled tend to coevolve and feature correlations in amino-acid usage. Recognising the importance of accounting for these interactions, we developed a Markov Chain Monte Carlo (MCMC) based simulation method that incorporates interactions flexibly, using Potts models, or single-sequence protein language models (PLMs), or MSA-based PLMs. We compared MSAs obtained from these three model types in terms of their closeness to protein families of interest, diversity within generated sequences and novelty when compared to natural sequences. Results indicate more favourable metrics when using the PLM-based generation methods, with the MSA-based PLM having an edge when considering shallow protein families. We are now using these coevolution-aware MSAs as more realistic training data to develop machine-learning models for tree inference.
Co-authors: Anne-Florence Bitbol
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826