View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

C-P.01: Interpreting the function of germline missense variants for the BAP1 gene using protein centric information.
Track: Proteins and structural biology
  • Pierre Marchal, Department of Medical Oncology, Inselspital, Bern University Hospital, DBMR, GCB, University of Bern, Switzerland., Switzerland
  • Joel Spies, Department of Medical Oncology, Inselspital, Bern University Hospital, DBMR, University of Bern, Switzerland., Switzerland
  • Lena Bürgi, Department of Medical Oncology, Inselspital, Bern University Hospital, University of Bern, Bern, Switzerland., Switzerland
  • Kasun Samarasinghe, SIB Swiss Institute of Bioinformatics and University of Geneva, Switzerland, Switzerland
  • Lydie Lane, SIB Swiss Institute of Bioinformatics and University of Geneva, Switzerland, Switzerland
  • Ferdinando Cerciello, Department of Medical Oncology, Inselspital, Bern University Hospital, DBMR, University of Bern, Switzerland., Switzerland


Presentation Overview: Show

Introduction: BAP1 is a tumor suppressor and deubiquitinating protein, which misfunction is associated with the rare hereditary BAP1 tumor predisposition syndrome. Current variants effect predictors (VEP) tools are capable to predict variants effect but lack specificity for the rare cancers associated with BAP1 mutations. Here, we propose a protein centric approach based on highly curated protein information integrated in the model ProFeaTrOL to predict the effect that missense variants might have on the function of BAP1.

Methods: We selected functional relevant BAP1 features from the knowledgebase neXtProt to establish the Protein Feature Training based for Oncology Likelihood (ProFeaTrOL) model. We trained and validated ProFeaTrOL based on literature available datasets of BAP1 mutations and tested its performance on ClinVar reports. we further verified the model on the BAP1 functionally associated protein BRCA1 and the independent gene TP53.

Results: ProFeaTrOL presented an AUC of 0.84 for discriminating functional and non-functional BAP1 variants of the validation set and of 0.89 in the testing set from ClinVar. As a comparison, PolyPhen-2 and SIFT presented AUC of 0.77 and 0.68 in the validation and 0.84 and 0.78 in the testing. Despite being trained on BAP1, if applied to the functionally related protein BRCA1, the model showed a high AUC of 0.81, while for TP53 it was only modest, confirming the specificity of the features selected for ProFeaTrOL.

Conclusion: The protein centric approach of ProFeaTrOL offers a strategy for the accurate SNVs effect prediction specific for the rare cancer conditions associated with germline variants of BAP1.

C-P.02: Generalizable Protein Language Models for Complex Variant Prediction via Multi-Task Epistatic Learning
Track: Proteins and structural biology
  • Fabio Mazza, University of Trento, Italy
  • Mattia Tenuti, University of Trento, Italy
  • Gianluca Lattanzi, University of Trento, Italy
  • Alessandro Romanel, University of Trento, Italy


Presentation Overview: Show

Understanding the impact of variants on protein stability and function is a core challenge in computational biology. Protein language models (PLMs) have demonstrated strong zero-shot capabilities in variant effect prediction. However, robust supervised fine-tuning across diverse functional readouts remains challenging. Furthermore, epistatic interactions, where the effect of a single mutation depends on another one, are rarely modeled and evaluated explicitly in PLM-based approaches, and multi-mutant training has usually focused on thermodynamic stability.
In this study, we present a multi-tasking fine-tuning framework for ESM-2 that integrates a Deep Mutational Scanning (DMS) pairwise ranking objective across multiple types of DMS assays with an auxiliary MLM objective that preserves ESM-2 language-modeling perplexity, supporting generalization capabilities. Rather than training separate prediction heads from scratch, we reuse the pretrained language-modeling head and learn lightweight task-specific matrices that bias the LM logits for each task. DMS scores are then computed as unmasked pseudo log-likelihoods of the sequences with the corresponding bias. To explicitly capture non-additivity, we introduce an epistatic loss term that evaluates mutational cycles to directly match predicted and experimental epistatic terms.
Benchmarking on held-out DMS assays, including assays containing INDELs, shows that our model achieves the best overall performance among the methods evaluated in ProteinGym. Performance gains are particularly pronounced on multi-mutant assays, where explicit epistasis training substantially improves correlation with experimentally derived specific epistasis effects, and also extend to unseen assay categories, including expression and activity assays.

C-P.03: Do AlphaFold Ensembles Capture Side-Chain Fluctuations? A Large-Scale Benchmark Against MD and PDB Homologs
Track: Proteins and structural biology
  • Alexander Korsunsky, Linkoping University, Sweden
  • Nicholas Pearce, Uppsala University, Sweden
  • Bjorn Wallner, Linkoping University, Sweden


Presentation Overview: Show

Protein side-chain fluctuations are essential for cellular processes such as enzymatic catalysis and molecular recognition. Accurately modeling these dynamics is vital for drug design and interpreting experimental data from NMR or crystallography.

In this work, we evaluate AlphaFold2 (AF2) and variants such as AlphaFold3 (AF3), AFsample2, AFsample3, and AF2χ by generating 1,000 structures for each of 120 proteins to analyze side-chain flexibility. We compare predicted side-chain dihedral (χ₁-χ₅) distributions, using both marginal (single-angle) and joint (multi-angle) analyses, against MD (ATLAS) and ensembles of sequence-homologous PDB structures.

To quantify differences between ensembles, we devised a metric based on optimal transport (Wasserstein distance) that accounts for the geometry of rotamer space while avoiding the bin-count dependency of a standard measure from information theory (Jenson-Shannon Divergence) and enables comparisons of multivariate (joint) distributions.

Our results show that the original structure prediction models AF2 and AF3 produce side-chain distributions with very limited conformational variation, which is consistent with the models being trained primarily on rigid crystal structures. While variants tuned for diversity (AFsample2, AF2χ) show increased variation, their similarity to MD and experimental ensembles remains low.
This study provides a benchmark to determine if ensembles from structure prediction tools can accurately replicate the physical fluctuations and highlights the current limitations of using structure prediction models to capture the full landscape of protein side-chain dynamics.

C-P.04: Benchmarking structure prediction in the time of AI
Track: Proteins and structural biology
  • Xavier Robin, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Gabriel Studer, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Stefan Bienert, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Janani Durairaj, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Peter Å krinjar, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Andrew M. Waterhouse, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Gerardo Tauriello, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Torsten Schwede, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland


Presentation Overview: Show

Recent advances in artificial intelligence have led to major improvements in biomolecular structure prediction, particularly for single‑chain protein targets, while also opening new and more challenging application areas. These include protein‑protein interactions, protein‑ligand complexes, and multi‑component assemblies. Emerging AI‑based methods address these problems, but their performance across different target types and modelling tasks is not yet systematically established, motivating the need for robust and adaptable benchmarking.
We present two complementary resources supporting objective and reproducible evaluation of structure prediction methods. OpenStructure provides a flexible and extensible toolkit for macromolecular structure comparison, enabling evaluation beyond a small set of global accuracy measures. It supports alternative scoring schemes, local and interface‑focused analyses, and dedicated scores for protein‑ligand interactions. The OpenStructure benchmarking framework also constitutes the basis of evaluation in community‑wide assessments such as CASP. Building on this toolkit, CAMEO provides fully automated and continuous benchmarking, with recent extensions enabling blind evaluation of macromolecular complexes, including protein‑protein and protein‑ligand predictions.
Together, CAMEO and OpenStructure provide a framework for benchmarking structure prediction across a broad and evolving range of modelling scenarios. They enable benchmarks to address central questions including how to quantify target difficulty, how to track progress over time, and how to determine when a modelling problem can be considered effectively solved. Separately, the framework supports broader inclusion of modern AI‑based methods alongside traditional prediction servers, ensuring benchmarking remains representative of current practice.

C-P.05: LLM-Assisted Development of Large-Scale Proteomics Infrastructure: An Empirical Study from PRIDE Reanalysis
Track: Proteins and structural biology
  • David Teschner, Institute of Computer Science, Johannes Gutenberg University, Mainz, Germany
  • David Gomez-Zepeda, Helmholtz Institute for Translational Oncology, Mainz, Germany, Germany
  • Tim Maier, Institute of Computer Science, Johannes Gutenberg University, Mainz, Germany
  • Thomas Kemmer, Institute of Computer Science, Johannes Gutenberg University, Mainz, Germany
  • Maximilian Sprang, Universitätsmedizin Mainz, Germany
  • Asis Asis Hallab, Technische Hochschule Bingen, Germany
  • Stefan Tenzer, University Medical Center, Johannes-Gutenberg University, 55131 Mainz, Germany, Germany
  • Andreas Hildebrandt, Institute of Computer Science, Johannes Gutenberg University Mainz, Germany


Presentation Overview: Show

In data-intensive fields such as proteomics, the primary bottleneck is not only algorithm development but also the construction and operation of scalable data infrastructure. Building such systems, including pipelines, data layers, and validation frameworks, typically requires large and specialized teams.

We present a case study based on a fully operational pipeline for the systematic reanalysis of public timsTOF datasets from the PRIDE repository. The system integrates multiple search engines, performs raw signal extraction, and produces a versioned, bias-documented reference layer for downstream analysis and model training. The pipeline was developed by a single researcher using LLM-assisted software development and evolved under real data constraints, providing a unique setting to study infrastructure construction at scale.

We analyze the development process through version history and code evolution and investigate system behavior under increasing data volume to identify characteristic failure modes and validation challenges. Our observations indicate that while LLM-assisted development enables rapid construction of complex systems, it introduces new classes of errors and shifts the primary challenge from implementation to validation and system-level consistency.

These results highlight the need for systematic validation strategies for LLM-assisted scientific software and suggest that such methods may allow individual researchers to build infrastructure that previously required dedicated teams.

C-P.06: Amyloid Landscape Explorer: A Network-Based Platform for Thermodynamic and Structural Comparison of Amyloid Fibrils
Track: Proteins and structural biology
  • Wannes Hermans, VIB Center for AI and Computational Biology, KU Leuven, Belgium
  • Vasilis Kyriazis, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain Institute, University of Texas, United States
  • Laxmikant Gadhe, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain Institute, University of Texas, United States
  • Katerina Konstantoulea, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain Institute, University of Texas, United States
  • Joana Pereira, VIB Center for AI and Computational Biology, KU Leuven, Belgium
  • Robin Bouchez, VIB Center for Neuroscience Leuven, Belgium
  • Rodrigo Gallardo, VIB Center for Neuroscience Leuven, Belgium
  • Nikolaos Louros, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain Institute, University of Texas, United States
  • Joost Schymkowitz, VIB Center for Neuroscience Leuven, KU Leuven, Belgium
  • Frederic Rousseau, VIB Center for Neuroscience Leuven, KU Leuven, Belgium


Presentation Overview: Show

Recent advances in cryo-electron microscopy and solid-state NMR have led to a rapid increase in high-resolution amyloid fibril structures, providing new insights into structural polymorphism, disease specificity, and aggregation behavior. This growing dataset calls for integrative approaches to systematically compare amyloids beyond traditional sequence- or structure-based classifications.

Here, we present the Amyloid Landscape Explorer, an interactive extension of the Amyloid Explorer, a public repository for structural and thermodynamic analysis of full-length amyloid fibril core structures. Within this framework, pairwise similarities between protofilament ΔG profiles are quantified using a custom Needleman-Wunsch-based alignment with a Gaussian similarity function and affine gap penalties, enabling comparison of energetic profiles. These similarities are represented as a weighted network, where nodes correspond to amyloid structures and edges reflect positive-scoring alignments. Leiden clustering identifies thermodynamic clusters, revealing groups of amyloids with similar stability patterns. In parallel, structural similarities are computed using pairwise structural alignments, defining structural clusters that are incorporated as complementary metadata.

The resulting landscape highlights both concordance and divergence between thermodynamic and structural organization. An interactive web interface enables database browsing and filtering by protein, disease, and experimental method, with options to color and filter the network based on thermodynamic or structural clusters. Additional features include a 3D molecular viewer with per-residue energy mapping, and cluster-specific multiple alignment analysis. Together, the Amyloid Landscape Explorer provides a comprehensive resource for exploring the energetic and structural basis of amyloid diversity.

C-P.07: The Encyclopaedia of protein Domains (TED) from AlphaFold2-predicted structures: expansion of CATH domain space and insights into structural diversity
Track: Proteins and structural biology
  • Andy Lau, Department of Computer Science, University College London, London WC1E 6BT, UK; InstaDeep Ltd, London W21AY, UK, United Kingdom
  • Nicola Bordin, Institute of Structural and Molecular Biology, University College London, London WC1E 6BT, UK., United Kingdom
  • Shaun Kandathil, Department of Computer Science, University College London, London WC1E 6BT, UK., United Kingdom
  • Ian Sillitoe, Institute of Structural and Molecular Biology, University College London, London WC1E 6BT, UK., United Kingdom
  • Vaishali Waman, University College London, United Kingdom
  • Jude Wells, University College London, United Kingdom
  • Christine Orengo, University College London, United Kingdom
  • David Jones, University College London, United Kingdom


Presentation Overview: Show

The AlphaFold Structure Database (AFDB) provides predictions for > 214 million three-dimensional structures of full-length proteins. Protein structure is composed of one or more domains i.e. independent folding units that can be found in multiple structural and functional contexts. We present The Encyclopedia of Domains (TED) which is the first structural-based resource to classify domains from the AFDB (https://ted.cathdb.info/).
TED is developed as a collaboration between the Orengo and Jones group. TED combines state-of-the-art deep learning-based domain detection, structure comparison (Foldseek) and fold recognition (FoldClass) algorithms to identify and classify domains across the structures from AFDB. We used novel deep learning-based domain detection methods (Merizo, Chainsaw) and UniDoc, to segment AFDB protein structures into protein domains.
TED contains over 370 million protein domains, providing domain structure coverage to over 1 million taxa/species. Notably, 80% of the TED domains exhibit similarities with known superfamilies in the CATH database, expanding the CATH by over 600-fold. We identified over 6,000 domains with potentially new folds, some of which have unique protein architectures not seen previously (e.g. 11-helix propeller and 11-bladed beta-propeller).
TED as a unique domain resource derived from AFDB, will aid in a multitude of downstream analyses including identification of remote homologues and functional diversity. We demonstrate an analysis of TED data providing insights into substrate-binding specificities in CoA-dependent acyltransferases in the context of pathogens.
TED and CATH (https://www.cathdb.info/) domain annotations are now made available via AlphaFold Database (https://alphafold.ebi.ac.uk/). CATH and TED, paves the way to understand biodiversity using protein domains.

C-P.08: Iterative HMM-Refined Coevolution Analysis Combined with AlphaFold3 for Protein Complex Discovery
Track: Proteins and structural biology
  • Giovanni Merici, University of Parma, Italy
  • Giulia Sassi, University of Parma, Italy
  • Karen Cárdenas Casillas, University of Parma, Mexico
  • Kristian Salvatore, University of Parma, Italy
  • Riccardo Percudani, University of Parma, Italy


Presentation Overview: Show

Understanding protein–protein interactions is essential to elucidate molecular mechanisms, yet identifying novel complexes or interactors of known assemblies remains challenging. Here, we present a computational pipeline integrating coevolution analysis (cotr) with AlphaFold3 (AF3) structural modelling to predict and prioritise candidate interactions supported by both evolutionary and structural evidence.

Orthogroups were first inferred from eukaryotic proteomes using SonicParanoid2. Each orthogroup was then modelled as a hidden Markov model (HMM) and used to iteratively search an expanded dataset of eukaryotic proteomes, with newly identified homologs incorporated at each round to update the HMM and refine protein presence–absence distributions across species. Cotr was subsequently applied to identify statistically significant coevolutionary associations between proteins, yielding coevolving modules. These were combinatorially assembled into putative complexes and evaluated using an optimised AF3 workflow. Predicted assemblies were prioritised using confidence metrics including ipTM, pTM, and PAE.

We applied this framework to the BBSome complex, identifying PDE6D as a potential interactor from coevolutionary clusters. AF3 multimer predictions revealed stable assemblies involving PDE6D with BBS2 and additional subunits. Molecular dynamics simulations further supported the stability of the PDE6D–BBS2 interaction, showing persistent inter-chain contacts and stable interface distances over time. Key interactions include a salt bridge and conserved hydrogen bonds with high occupancy across simulations. Interface energetics and in silico mutagenesis confirmed the importance of these residues, with mutations at the interface leading to measurable destabilisation of the complex.

Overall, this approach enables the systematic generation of experimentally testable hypotheses to discover and prioritise novel protein–protein interactions in disease-relevant complexes.

C-P.09: From 1,000 Spiders to Designed Protein Materials: Decoding the Spider Silkome
Track: Proteins and structural biology
  • Kazuharu Arakawa, Keio University, Japan


Presentation Overview: Show

Spider silks exhibit an unparalleled combination of tensile strength, extensibility, and toughness, defining a unique processing–property space among biopolymers. However, the sequence–property relationships underlying this diversity remain incompletely understood. To systematically explore this space, we established a global “1000 Spiders” initiative, collecting ~1,000 species and comprehensively characterizing their silk genes and mechanical properties. The resulting open resource, the Spider Silkome Database, integrates >10,000 gene sequences and thousands of material property measurements, enabling large-scale genotype–phenotype analyses.

Bioinformatic mining identified key sequence motifs and previously underappreciated components, including MaSp3 and a novel additive-like protein (SpiCE), that significantly influence silk performance. Motif-level engineering validated both positive and negative correlations with properties such as strength, toughness, and supercontraction.

Building on this dataset, we developed a data-driven materials transformation (DxMT) framework coupled with generative AI. Due to the extreme length and repetitive architecture of spidroins, conventional protein language models are insufficient; instead, we implemented a modular strategy that decomposes sequences into N-terminal, repeat, and C-terminal domains, followed by retrieval-based assembly of pseudo–full-length sequences. This enabled the construction of ~10,000 full-length candidates and improved prediction of mechanical properties. Fine-tuned models (e.g., ESM2 + LoRA) further enhanced property prediction accuracy.

Finally, AI-guided motif design was experimentally validated through recombinant silk spinning, demonstrating improved molecular orientation and mechanical performance. This integrated approach establishes a scalable paradigm for decoding and designing high-performance protein materials from evolutionary diversity.

C-P.10: LIGYSIS: a resource for predictor training, benchmarking and functional characterisation of ligand binding sites
Track: Proteins and structural biology
  • Javier Sánchez Utgés, University College London, United Kingdom
  • Stuart MacGowan, University of Dundee, United Kingdom
  • Diane Lee, University of Dundee - current address: University of York, United Kingdom
  • Gopal Sapkota, University of Dundee, United Kingdom
  • Geoff Barton, University of Dundee, United Kingdom


Presentation Overview: Show

Reliable protein-ligand binding site definition and prediction underpin function annotation, variant interpretation and drug discovery, yet many commonly used datasets are built from single asymmetric units, leading to artificial crystal contacts, redundant interfaces, and incomplete site definitions. LIGYSIS is a comprehensive dataset that aggregates unique, biologically relevant protein-ligand interfaces across the biological assemblies of the multiple structures of a protein. The human LIGYSIS component comprises 3448 proteins, 8244 binding sites, and 65,116 ligands, and has been used as a reference set for the largest independent systematic comparative evaluation of ligand site prediction tools to date.

To make these data accessible, we developed LIGYSIS-web – an open resource hosting the entire LIGYSIS set, which includes 64,782 binding sites across 25,003 proteins. Users can query UniProt accession identifiers or upload their own structures for automated site definition, characterisation, interactive 3D visualisation, exploration, and download. Sites are characterised by evolutionary divergence, human missense variation, and solvent accessibility-derived functional scoring, enabling prioritisation of likely functional pockets and residues, as shown in our recent study.

Finally, LIGYSIS is increasingly being adopted to train and test models for the prediction of general binding, allosteric, and cryptic sites, as well as being integrated in new resources. In parallel, we are using the full LIGYSIS set to train a new ligand site predictor that uses protein language models, developing an integration of LIGYSIS-web with Jalview, and extending our work on ligand site prediction evaluation and good practices.

C-P.11: 3DSeqCheck: Identifying discrepancies of structure-based resources such as AlphaFoldDB in relation to the evolving UniProt entries
Track: Proteins and structural biology
  • Ifigenia Tsitsa, Imperial College London, United Kingdom
  • Anja Conev, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Suhail Islam, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Alessia David, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Michael J E Sternberg, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom


Presentation Overview: Show

UniProt is a central repository of protein sequences and annotations, with entries being updated several times a year as new sequencing evidence is collected. By contrast, protein structure resources often evolve at a different pace. The AlphaFold database remained unchanged for four years, until September 2025. In our article in Nature Structural and Molecular Biology in 2025, we documented the discrepancies that have accumulated in AlphaFoldDB as it aged in comparison to UniProt until its latest release in 2025. During that time, nearly 3% of the associated sequences underwent revisions in UniProt.
In a range of bioinformatics tasks, protein structure data is paired with sequence annotations from UniProt. Mapping annotations to outdated structure files can lead to errors in downstream analysis. While this concern has been addressed for experimental structures through SIFTS, efforts for computationally modelled structures are lacking. Here, we present 3DSeqCheck, published in JMB in December 2025, which is a lightweight web tool that enables quick comparison of both computationally modelled and experimental structures to the latest UniProt entries. 3DSeqCheck provides an interactive visual panel of the alignment and the comparison of the residue numbering and can be accessed freely at: https://missense3d.bc.ic.ac.uk/3dseqcheck.

C-P.12: Rewriting protein alphabets with language models
Track: Proteins and structural biology
  • Janani Durairaj, University of Basel, Switzerland
  • Gabriel Studer, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Lorenzo Pantolini, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland
  • Laura Engist, Department of Mathematics and Computer Science, University of Basel, Basel Switzerland, Switzerland
  • Florian Pommerening, Department of Mathematics and Computer Science, University of Basel, Basel Switzerland, Switzerland
  • Ieva PudžiuvelytÄ—, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland
  • Andrew M. Waterhouse, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Stefan Bienert, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Gerardo Tauriello, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
  • Martin Steinegger, School of Biological Sciences, Seoul National University, Seoul, South Korea, Switzerland
  • Torsten Schwede, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland


Presentation Overview: Show

Detecting remote homology with speed and sensitivity is crucial for tasks like function annotation and structure prediction. We introduce a novel approach using contrastive learning to convert protein language model embeddings into a new 20-letter alphabet, The Embedded Alphabet (TEA), enabling highly efficient large-scale protein homology searches.

Searching with our alphabet performs on par with and complements structure-based methods without requiring any structural information, and with the speed of sequence search. This provides a significant advantage for proteins that lack structural data and enables searches at the scale of available sequence data numbering over a billion. The inbuilt entropy metric provides a measure of confidence offering a valuable estimate of prediction reliability for downstream analysis. This work has broad implications, significantly accelerating and broadening the scope of protein sequence analysis. The ability to detect remote homologs more effectively enhances functional annotation transfer and drug discovery efforts. We also provide a user-friendly web-server at https://pickybinders.org/tea enabling conversion of amino acid sequences to TEA sequences and searches on the UniRef50 and Foldseek AlphaFold Clusters databases.

Ultimately, we bring the exciting advances in protein language model representation learning to the plethora of sequence bioinformatics algorithms developed over the past century, offering a powerful new tool for biological discovery.

C-P.13: Predicting changes in protein-protein binding affinity upon mutation with statistical potentials
Track: Proteins and structural biology
  • Gabriel Cia, Université Libre de Bruxelles, Belgium
  • André Ciupitu, Université Libre de Bruxelles (ULB), Belgium
  • Marianne Rooman, Université Libre de Bruxelles, Belgium
  • Fabrizio Pucci, Université Libre de Bruxelles, Belgium


Presentation Overview: Show

In-silico approaches for predicting changes in protein-protein interaction (PPI) binding affinity upon mutation (ΔΔG_b) continue to exhibit critical biases and difficulties generalizing beyond their training datasets. We propose a novel physics-based approach, termed "interface potentials," for the description of PPI interfaces and the prediction of variant effect. These statistical potentials are derived from datasets of known PPI structures using the Boltzmann Law and describe the interfaces in terms of distances, torsion angles, and solvent accessibility of residues. Interface potentials demonstrate great performance on both the SKEMPI2 experimental dataset and Deep Mutational Scanning datasets without requiring additional machine learning or model training, highlighting the power of the physics-based framework. Furthermore, interface potentials can serve as invaluable physics-based features in sophisticated prediction models such as AbMuSiC, our state-of-the-art physics-based ΔΔG_b prediction method for antibody-antigen interactions.

C-P.14: AI Enhanced Viral Structural Phylogenetics
Track: Proteins and structural biology
  • David Moi, University of Lausanne, Switzerland
  • Christophe Dessimoz, University of Lausanne, Department of Computational Biology, Swiss Institute of Bioinfromatics, Switzerland
  • Dongwook Kim, UNIL, Switzerland


Presentation Overview: Show

Recent efforts in structural phylogenetics have highlighted their ability to resolve deeper evolutionary relationships than traditional sequence-based methods. These approaches hold promise for elucidating phylogenetic and taxonomic relationships within the virosphere which tend to escape traditional phylogenetics due to the pace of viral evolution. However, the paucity of available structures for the incredible observable diversity of extant viruses and the computational cost of inferring structures has prevented a systematic structural exploration of the virosphere. We have fine-tuned ESMc to produce structural tokens through Foldseek's 3Di alphabet using the structures available in the AlphaFold2-based structures available in the Big Fantastic Virus Database. We have succeeded in producing a viral-focused protein LLM which allows for rapid conversion of amino acid sequences into 3Di tokens and opens up the voluminous viral sequencing datasets for structurally informed homology searching and phylogenetics through the use of Foldseek. We have also created a pipelines and tools for infering cleavage sites, constructing phylogenies of structurally homologous cleaved products, annotation of functional content and the construction of 'taxonomic' consensus trees centered around the structure enhanced representation of viral proteomes.

C-P.16: ProInterVal-BioXtal: A Web Server to Distinguish Biological Interfaces from Crystal Contacts Using Graph Deep Learning
Track: Proteins and structural biology
  • Defne Alnigenis, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey, Turkey
  • Damla Ovek Baydar, Dept. of Computer Eng., Koç Univ., Istanbul, Turkey & NCMBM, Univ. of Oslo, Norway, Norway
  • Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey, Turkey
  • Attila Gursoy, Department of Computer Engineering, Koç University, Istanbul, 34450, Turkey, Turkey


Presentation Overview: Show

Accurate discrimination between biologically relevant protein–protein interfaces (PPIs) and crystallographic contacts is essential for reliable interpretation of macromolecular assemblies and their cellular functions. While X-ray crystallography remains a primary method for determining protein complex structures, crystallographic interfaces may be formed as a byproduct of the crystal packing. Therefore, robust computational approaches for interface annotation are needed.
This study introduces ProInterVal-BioXtal, a web server that predicts class scores for PPIs, differentiating between biologically relevant and crystallographic interfaces. The method leverages protein representation learning and graph-based deep learning to capture structural and physicochemical features. Interface graphs derived from input complexes are processed by our novel graph-based contrastive model to learn interface representations which are used by a graph neural network for classification.
The model is trained and validated on the MANY benchmark dataset comprising 5739 dimers with a balanced distribution of biological and crystal interfaces and evaluated on the DC benchmark dataset with curated interfaces of similar interface areas.
On the DC test set, ProInterVal-BioXtal achieves 88% accuracy, 88% precision, and 85% F1 score, outperforming state-of-the-art methods including DeepRank-GNN, PRODIGY-CRYSTAL, EPPIC 3, PISA, and QSAlign. On an independent benchmark, the method achieves 83% accuracy and 0.91 AUC, surpassing DeepRank-GNN (AUC = 0.85).
Through a user-friendly web interface, ProInterVal-BioXtal enables rapid predictions from PDB files or IDs, returning probabilistic classification scores. Our tool addresses a fundamental challenge in structural biology by utilizing graph-based protein representation which can detect complex interactions and dependencies. ProInterVal-BioXtal server is freely available at https://3dpath.ku.edu.tr/prointerval-bioxtal/

C-P.17: Improving protein structure prediction with structure-aware multiple sequence alignments
Track: Proteins and structural biology
  • Diana Rapota, Biozentrum, University of Basel, Basel Switzerland; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland
  • Celia Ulrich, Biozentrum, University of Basel, Spitalstrasse, 4056, Basel Switzerland, Switzerland
  • Lorenzo Pantolini, Biozentrum, University of Basel, Basel Switzerland; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland
  • Janani Durairaj, Biozentrum, University of Basel, Basel Switzerland; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland


Presentation Overview: Show

Leveraging evolutionary information is central to modern protein structure prediction. As protein structure is more conserved than sequence throughout evolution, searching for distant protein homologs using structural information is a highly appealing strategy. However, until recently, performing high-throughput searches for distant structural homologs with low sequence identity was not feasible without access to large databases of structures. The Embedded Alphabet (TEA), a novel 20-letter 1D structural alphabet, addresses this challenge by enabling fast and sensitive detection of distant homologs without requiring prior structural information. In this work, we investigate whether structure-aware multiple sequence alignments (MSAs) generated with TEA can improve protein structure prediction by enhancing the evolutionary signal captured in the alignments. Structure-aware MSAs were constructed using Search with TEA against Many (STEAM), a tool built on the Foldseek framework and adapted for the TEA alphabet. We evaluated these MSAs on two datasets: the CASP13 dataset, containing query proteins with shallow MSAs, and viral proteins from BFVD, previously characterized by poor-quality MSAs using the default ColabFold DB. To evaluate the impact on structure prediction, we ran AlphaFold2 (AF2) in custom MSA mode using both STEAM-generated structure-aware MSAs and MMseqs2-generated MSAs. We present an evaluation of STEAM-generated MSAs and their effect on AF2 prediction quality across both datasets, with a particular focus on challenging low-homology and viral targets, where sequence-based searches alone may be insufficient.

C-P.18: Computational prediction of ligand-receptor pairs using transformer-based language models, interaction networks, and structural modelling
Track: Proteins and structural biology
  • Iuliia Trifonova, Department of Biosciences and Medical Biology, Center for Tumor Biology and Immunology, Paris Lodron University Salzburg, Austria
  • Markus Wiederstein, Department of Biosciences and Medical Biology, Paris Lodron University Salzburg, Austria
  • Nikolaus Fortelny, Department of Biosciences and Medical Biology, Center for Tumor Biology and Immunology, Paris Lodron University Salzburg, Austria


Presentation Overview: Show

Ligand-receptor (LR) interactions regulate cellular communication and shape processes in development, immune responses, and disease. However, existing curated databases remain incomplete, and many biologically relevant interactions are described only in primary literature, limiting systematic discovery. We developed a scalable computational framework that integrates literature mining with transformer-based classification and structural validation to identify previously unannotated LR interactions among known protein-protein interactions.
We first harmonised 18 curated LR databases into one reference database of 18,283 unique protein-protein LR pairs and used the high-confidence ConnectomeDB2025 subset (3,550 pairs) as true positives. We evaluated multiple strategies to derive true negatives, finally using topologically distant proteins based on a consensus protein-protein interaction network from the PCNet database. We then extracted sentences mentioning true positive and true negative pairs from Europe PMC and PubMed.
To separate contextual co-mention from mechanistic interaction, we trained a logistic-regression classifier on sentence embeddings from a frozen PubMedBERT encoder. This classifier reached a pair-level ROC AUC of 0.86. We next evaluated high-scoring candidates structurally using AlphaFold3. The strongest previously unannotated candidate, COL2A1-TGM2 (ipTM = 0.80), showed a plausible interface, though experimental work is needed to confirm whether these are true extracellular LR pairs or other PPIs.
Together, these results suggest that literature-guided classification combined with structural filtering can potentially expand LR knowledge beyond curated resources, although distinguishing true LR pairs from other PPIs and improving semantic negation handling remain important next steps.

C-P.19: Generating tailored high-quality datasets for benchmarking structure-based computational drug design tools: a docking assessment
Track: Proteins and structural biology
  • Marine Mathieu, SIB Swiss Institute of Bioinformatics, Switzerland
  • Ute F. Röhrig, SIB Swiss Institute of Bioinformatics, Switzerland
  • Vincent Zoete, UNIL University of Lausanne, Switzerland


Presentation Overview: Show

Structural bioinformatics is critical for drug discovery, as it provides methods and tools to predict, analyze, and validate 3D data of macromolecules. The ELIXIR 3D-BioInfo community is organized around five main activities, one of which aims to create the tools necessary for developing large-scale, high-quality, and sustainable datasets of ligand-protein complexes. These datasets will be used to assess and benchmark structure-based computer-aided drug design (SB-CADD) algorithms, such as docking, virtual screening and binding site detection tools.
We developed three interconnected Nextflow pipelines capable of constructing datasets from complex, protein, or ligand PDB identifiers. Data sources include multiple APIs and dedicated software developed by ELIXIR participating groups for data retrieval and processing.
The pipelines generate multifaceted datasets with comprehensive annotations. Complex analysis involves retrieving structural details including proteins, bound ligands and their interactions, as well as quality metrics such as resolution, missing atoms, crystal contacts, and electronic density support from the Protein Data Bank. Coordinate files are retrieved for complexes and ligands, as well as binding affinity, refined structures, tautomers, and protonation states. The protein characterization pipeline extracts UniProt KB descriptions, associated 3D structures, bound ligands, binding sites, active and inactive molecules, and protein flexibility. For ligands, molecular properties, 3D conformers, partial charges, and associated complex structures are provided.
We present benchmark sets generated from complex, protein, or ligand identifiers, along with an initial docking assessment. Docking with AutoDock Vina, AutoDock4, and SMINA provides a first evaluation of SB-CADD tool performance on the curated, high-quality data.

C-P.20: Integrated CSF-serum proteomic profiling in spinal cord injury
Track: Proteins and structural biology
  • Sasimonthakan Tanarsuwongkul, Spinal Cord Injury Center, Heidelberg University Hospital, Germany
  • Norbert Weidner, Spinal Cord Injury Center, Heidelberg University Hospital, Germany
  • Giada Sandrini, German Cancer Research Center (DKFZ), Germany
  • Katalin Barkovits, Medical Faculty, Ruhr University Bochum, Germany
  • Andreas Hug, Spinal Cord Injury Center, Heidelberg University Hospital, Germany
  • Junyan Lu, Institute for Computational Biomedicine, Heidelberg University, Germany


Presentation Overview: Show

Traumatic spinal cord injury (SCI) leads to lifelong impairment and disability due to the lack of spontaneous nerve regeneration. Neuroregenerative therapies are being investigated in search for effective interventions to improve clinical outcomes for SCI patients. The current patient stratification and neurological outcome assessment in SCI clinical trials involves the use of the American Spinal Injury Association Impairment Scale (AIS). However, AIS grades are unable to distinguish different neurological impairments. Therefore, it often fails to reflect clinically important differences in severity and recovery, especially in people with cervical SCI. To address this, proteome of cerebrospinal fluid and blood serum from NISCI trial (ClinicalTrials.gov, NCT03935321), which tests the effectiveness of a nogo-A antibody in acute SCI, have been evaluated. This longitudinal integrated CSF-serum proteomic profiling identified underlying pathways of SCI severity and recovery and molecular predictors of neuronal improvement, which could be used to stratify patients with SCI.

C-P.21: Exploring the mobile genetic element continuum using PhageProfiler
Track: Proteins and structural biology
  • Jiawei Wang, University of Bath; EMBL-EBI, United Kingdom
  • Jinzheng Ren, University of Bath; EMBL-EBI; Australian National University, United Kingdom
  • Licheng Zong, University of Bath; EMBL-EBI; The Chinese University of Hong Kong, Hong Kong
  • Robert Finn, EMBL-EBI, United Kingdom


Presentation Overview: Show

Mobile genetic elements (MGEs), including plasmids and bacteriophages, are often treated as discrete categories, despite growing evidence for a modular and evolutionary continuum shaped by gene exchange, recombination, and hybrid elements such as phage–plasmids. Here, we present PhageProfiler, an embedding-based framework for profiling MGEs across sequence space. By learning task-specific representations at different biological levels, PhageProfiler captures continuous organization at the replicon level, while preserving discrete structure at taxonomic and functional levels. These representations provide complementary views of MGE organization, linking genome-level organization to both evolutionary relationships and underlying gene content. Using virulent phages as a case study, we demonstrate that these perspectives are mutually consistent and biologically informative.

C-P.22: Exploring protein language model approaches for thermal stability prediction
Track: Proteins and structural biology
  • Gregory Coolen, Université Libre de Bruxelles, 3BIO-BioInfo, Belgium
  • Marianne Rooman, Université Libre de Bruxelles, 3BIO-BioInfo, Belgium
  • Fabrizio Pucci, Université Libre de Bruxelles, 3BIO-BioInfo, Belgium


Presentation Overview: Show

Accurately determining a protein's melting temperature is essential for many biotechnological applications such as enzyme design or understanding organismal adaptation to environmental conditions. While experimental methods are accurate, they are also costly and time-consuming, which justifies the development of computational alternatives. Recently, protein language models (pLM) applied to protein sequences have shown impressive performance across various tasks, including protein structure prediction. These methods have also been explored for thermostability prediction and reported to achieve good performance. In this work, we review existing pLM-based thermostability predictors and benchmark their robustness on new datasets of thermal stability that we collected from the literature. We found that their actual performance remains modest. This can be attributed to the limited availability of annotated data and the fact that melting temperature is sometimes sensitive to minor sequence variations. Finally, to better understand the biophysical principles underlying thermal stability and what pLM-based models capture, we analyze their internal representations using standard interpretability techniques, such as ablation.

C-P.23: A Unified Framework for TCR-pMHC Structural Model Assessment
Track: Proteins and structural biology
  • Miguel Romero-Durana, Barcelona Supercomputing Center, Spain
  • Alex Ascunce-París, Barcelona Supercomputing Center, Spain
  • Alfonso Valencia, Barcelona Supercomputing Center, Spain
  • Roc Farriol-Duran, Barcelona Supercomputing Center, Spain
  • Victor Guallar, Barcelona Supercomputing Center, Spain


Presentation Overview: Show

Structural characterization of T cell receptor (TCR) and peptide-MHC complex (pMHC) interactions is fundamental to understanding adaptive immunity. However, the scarcity of experimentally resolved TCR-pMHC structures constrains progress in this field. Computational modelling tools can help bridge this gap, but their predictions are difficult to evaluate reliably without experimental references.
We present a scalable and interpretable framework for reference-free quality assessment of TCR-pMHC class I structural models.
We benchmarked seven state-of-the-art modelling methods on 265 experimentally determined PDB complexes, covering pre- and post-training-cutoff datasets. AlphaFold3 demonstrated superior generalization, achieving over 90% acceptable-quality predictions for structures absent from its training set.
Then, we integrated multiple confidence metrics into a random forest classifier trained on 1,325 AlphaFold3 models, with quality labels derived from their 265 experimental counterparts. The classifier stratifies models into four tiers (low, acceptable, medium, high), outperforming single-metric approaches. SHAP analysis identified pDockQ2 and iPDE as the strongest predictors of structural quality.
Finally, we applied this framework to a VDJdb/TCRvdb-derived subset of validating and non-validating TCR-pMHC pairs, and to IMMREP23, comprising validated positives and synthetic negatives. True binders were consistently enriched in higher quality tiers, confirming that structural quality is a reliable indicator of functional interaction.
Overall, this work contributes a large-scale dataset of evaluated TCR-pMHC structural models (21,775 models from 4,355 complexes; 265 PDB, 609 TCRvdb, 3,484 IMMREP23), and a robust framework for their quality assessment. This will facilitate the curation of structural TCR-pMHC datasets, and downstream applications in immunology and TCR-based therapeutics.

C-P.24: From Proteome Profiles to Clinical Diagnosis in Children with Bone and Joint Infections
Track: Proteins and structural biology
  • Réka Gonda, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital, Copenhagen, Denmark, Denmark
  • Allan Bybeck Nielsen, Department of Paediatrics and Adolescent Medicine, Copenhagen University Hospital – Rigshospitalet, Copenhagen, Denmark, Denmark
  • Anna Benedetti, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital; Copenhagen Center for Translational Research, Denmark
  • Anna Melidi, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital, Copenhagen, Denmark, Denmark
  • Nicolai Wewer Albrechtsen, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital; Department of Clinical Medicine, UCPH, Denmark
  • Ulrikka Nygaard, Department of Paediatrics and Adolescent Medicine, Rigshospitalet; Department of Clinical Medicine, UCPH, Denmark
  • Annelaura Bach Nielsen, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital; Copenhagen Center for Translational Research, Denmark


Presentation Overview: Show

Background: Bone and joint infections (BJI) in children represent a diagnostic challenge due to their clinical overlap with transient synovitis (TS). Current diagnostic approaches rely on clinical scoring systems, inflammatory markers and imaging, yet none achieve sufficient sensitivity to distinguish the two conditions without invasive procedures. Plasma proteomics offers a non-invasive approach to address this diagnostic gap, as mass spectrometry-based proteomics captures the systemic host response at a molecular resolution beyond what is achievable with single biomarker measurements.
Methods and Results: We performed untargeted liquid chromatography mass spectrometry of plasma samples from 262 pediatric patients across five inflammatory disease groups, including BJI (n=142) and TS (n=83). Following preprocessing, batch correction and imputation, differential abundance analysis identified 118 significantly differentially abundant proteins between BJI and TS (FDR-adjusted p<0.05). The most strongly upregulated proteins in BJI included the acute phase reactants SAA1, SAA2 and CRP, while apolipoproteins were among the most downregulated. Reactome pathway enrichment revealed significant enrichment of Toll-like receptor signalling, innate immune activation and neutrophil degranulation in BJI, alongside suppression of metabolic and lipid transport pathways.
Perspective: Ongoing analyses aim to stratify BJI by microbiological confirmation status to identify pathogen-specific proteomic signatures. We are developing a plasma proteomics-based machine learning model to distinguish BJI from TS, with the long-term perspective of providing a clinical decision support tool guiding clinicians between surgical intervention and conservative management to reduce unnecessary invasive procedures in children with suspected BJI.

C-P.25: PUFFIN: Protein Unit Discovery with Functional Supervision
Track: Proteins and structural biology
  • Gökçe Uludoğan, Bogazici University, Turkey
  • Buse Giledereli, Bogazici University, Turkey
  • Elif Ozkirimli, Roche AG, Switzerland
  • Arzucan Ozgur, Bogazici University, Turkey


Presentation Overview: Show

Motivation:
Proteins carry out biological functions through the coordinated action of groups of residues organized into structural arrangements. These arrangements, which we refer to as protein units, exist at an intermediate scale, being larger than individual residues yet smaller than entire proteins. A deeper understanding of protein function can be achieved by identifying these units and their associations with function. However, existing approaches either focus on residue-level signals, rely on curated annotations, or segment protein structures without incorporating functional information, thereby limiting interpretable analysis of structure-function relationships.
Results:
We introduce PUFFIN, a data-driven framework for discovering protein units by jointly learning structural partitioning and functional supervision. PUFFIN represents proteins as residue-level structure graphs and applies a graph neural network with a structure-aware pooling mechanism that partitions each protein into multi-residue units, with functional supervision that shapes the partition.
We show that the learned units are structurally coherent, exhibit organized associations with molecular function, and show meaningful correspondence with curated InterPro annotations. Together, these results demonstrate that PUFFIN provides an interpretable framework for analyzing structure-function relationships using learned protein units and their statistical function associations.
Availability and Implementation: We made our source code available at https:/github.com/boun-tabi-lifelu/puffin

C-P.26: PPInterface v2: A Dataset and Web Resource for 3D Protein–Protein Interface Analysis and Visualization
Track: Proteins and structural biology
  • Defne Alnigenis, Department of Chemical and Biological Engineering, Koç University, Istanbul, Turkey, Turkey
  • Moaaz Ur Rehman Azhar Khokhar, Department of Computer Engineering, Koç University, Istanbul, 34450, Turkey, Turkey
  • Zeynep Abali, Computational Science and Engineering Graduate Program, Koç University, Istanbul 34450, Turkey, Turkey
  • Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey, Turkey
  • Attila Gursoy, Department of Computer Engineering, Koç University, Istanbul, 34450, Turkey, Turkey


Presentation Overview: Show

Proteins regulate cellular processes through sophisticated interactions formed at protein-protein interfaces (PPIs). Structural characterization of these interfaces is essential for understanding biological functions and molecular mechanisms of proteins. Given the pivotal roles of PPIs, the construction of a well-curated dataset of interface structures holds significant promise for large-scale analysis.
This study presents an updated version of the PPInterface dataset, a comprehensive resource of structural protein–protein interaction data gathered from the Protein Data Bank (PDB). Building on our previously published framework (2024), the updated dataset integrates newly available structures and interaction annotations. The previous release of PPInterface contained 815,082 interfaces extracted from over 215,000 three-dimensional protein structures. The database now contains 252,476 protein complexes and more than 940,000 interfaces, indicating a substantial increase over the previous release. To our knowledge, PPInterface provides one of the most comprehensive datasets of PPIs, allowing users to search, explore, and download data for a wide range of analyses.
In addition to expanded data coverage, the user-friendly web resource of PPInterface enables detailed interpretation of PPIs, including structural, physicochemical, and evolutionary features through efficient exploration of with advanced search, filtering, and visualization. PPInterface, as a valuable resource for the analysis of protein interactions, may be used as a benchmark dataset for developing computational methods, including machine learning-based approaches. This updated database provides an up-to-date and scalable resource to the community for which can be used to support studies in structural bioinformatics, systems biology and PPIs in general.

C-P.27: RECUR: identifying recurrent amino acid substitutions from multiple sequence alignments
Track: Proteins and structural biology
  • Elizabeth Robbins, Ellison Institute of Technology Oxford, United Kingdom
  • Yi Liu, University of Oxford, United Kingdom
  • Steven Kelly, University of Oxford, United Kingdom


Presentation Overview: Show

Identifying recurrent changes in biological sequences is important to multiple aspects of biological research—from understanding the molecular basis of convergent phenotypes, to pinpointing the causative sequence changes that give rise to antibiotic resistance and disease. Here, we present RECUR, a method for identifying recurrent amino acid substitutions from multiple sequence alignments that is fast, easy to use, and scalable to thousands of sequences. We demonstrate that RECUR's recurrence detection achieves 100% accuracy on simulated data with known evolutionary histories. We further show that RECUR is robust to realistic levels of tree inference error. Finally, we apply RECUR to a large set of surface glycoprotein (S) protein sequences from SARS-CoV-2. This analysis identified widespread recurrent evolution throughout the protein with significant enrichment in the exposed receptor-binding S1 subunit and at the interface with the human angiotensin-converting enzyme 2 (hACE2). In contrast, recurrent substitutions were depleted at the trimeric interface of the S protein. In silico modelling showed that recurrent substitutions had no directional effect on stability at either interface, but effects at the hACE2 interface were significantly more variable. Multiple substitutions with large destabilizing effects on hACE2 binding have been linked to immune escape, while others represented reversions back to the reference sequence, suggesting that recurrent evolution at this interface reflects opposing selective pressures balancing receptor binding with immune evasion. A standalone implementation of the algorithm is available under the GPLv3 license at https://github.com/OrthoFinder/RECUR.

C-P.28: Machine Learning Classification and Analysis of Gain-of-Function, Loss-of-Function, and Neutral Mutations in Cancer
Track: Proteins and structural biology
  • Biniam Haile, University Of Sussex, United Kingdom
  • Adnan Cinar, University of Sussex, United Kingdom
  • Frances Pearl, University of Sussex, United Kingdom


Presentation Overview: Show

Missense mutations in cancer can be broadly categorised as gain-of-function (GOF), which enhance or alter protein activity, and loss-of-function (LOF), which impair or abolish protein function. Mutations that contribute to tumour growth by providing a selective growth advantage are considered driver mutations, while others are classified as passenger mutations, having no direct selective advantage for cell growth.

Accurately distinguishing GOF, LOF, and Neutral/passenger mutations is essential for understanding tumour biology, identifying therapeutic targets, and developing precision oncology strategies. In this study, we conduct a comparative analysis of these mutation classes by examining their structural and functional impact profiles, and their roles in cancer progression. We integrate a range of computational tools and curated mutation dataset (GOF, LOF, and Neutral) for structural modelling (SAAP, and NetsurfP), functional consequence prediction, and statistical evaluation.

Our findings were able explore distinct molecular patterns that might be very important to differentiate GOF, LOF, and Neutral mutations, providing insights into their molecular mechanisms in cancer. This comprehensive characterisation underscores the importance of combining structural impact analysis with functional annotations to improve mutation interpretation or classification. That was ultimately used in developing a machine learning model to classify missense mutation, which achieved ROC AUC of 0.92 classifying GOF and LOF mutations, and ROC AUC 0.943 on external validation mutation dataset.

C-P.29: AI-driven Intrabody Design: Representation Learning and Multi-objective Optimization of Nanobody Frameworks
Track: Proteins and structural biology
  • Savelii Komlev, Argentys LLC, France
  • Ilya Mazo, Argentys LLC, United States
  • Aleksei Artemiev, Argentys LLC, Spain


Presentation Overview: Show

Intrabodies are intracellular antibody fragments that enable targeting of proteins inaccessible to conventional biologics, but their development is limited by poor stability in the reducing cytoplasmic environment. As a result, intrabody engineering remains largely empirical and lacks scalable computational approaches.

We present an AI-driven pipeline for intrabody design that integrates protein language model–based representation learning, supervised stability prediction, and generative nanobody optimization. Sequence embeddings derived from pretrained models (ESM2, ProtT5, IgT5) are used to train machine learning classifiers for predicting intracellular stability of nanobodies. Models trained only on framework regions of nanobodies outperform full-sequence models, achieving ROC-AUC > 0.80, indicating that stability-related signals are primarily encoded in framework residues.

To explore sequence space, we implement a hybrid generative pipeline combining consensus-based framework mutagenesis (consensus nanobody sequence based on a set experimentally validated stable intrabodies), language model–guided sequence infilling, and stochastic diversification. More than 5,000 candidate nanobody variants are generated from a small set of parental sequences.

Generated variants are evaluated using independent structure-based metrics not used during model training, including AlphaFold-Multimer interface confidence (ipTM) and FoldX interface energies. Top candidates show a strong shift toward higher predicted interaction confidence (ipTM ~0.55–0.75 vs. ~0.1–0.2 in parental sequences) and more favorable binding energetics (median FoldX ΔG shifts from ~−5 to ~−25 kcal/mol, with reduced variance across top-ranked variants).

Selected candidates are currently undergoing experimental validation. These results demonstrate that combining representation learning, generative design, and structure-based screening enables scalable computational intrabody engineering.

C-P.30: Identification and characterization of Polyurethane-Degrading Enzymes from MGnify Metagenomes
Track: Proteins and structural biology
  • Joel Roca Martinez, UCL, United Kingdom
  • Christine Orengo, UCL, United Kingdom


Presentation Overview: Show

The discovery of enzymes capable of degrading synthetic polymers remains a major challenge due to the vast size and functional diversity of metagenomic sequence space. Here, we present a rational, structure- and sequence-informed pipeline for enzyme discovery, applied to the identification of novel polyurethane-degrading enzymes (PURases) from large-scale metagenomic data. Starting from 2.4 billion protein sequences in MGnify, we performed homology-based filtering against three known PURases followed by conservation analysis of the catalytic triad, yielding ~8,000 high-confidence candidates. To prioritize functionally diverse yet tractable subsets, we clustered sequences into functional families using an embedding-based classifier (eMMA-FunFamer) and identified function-determining positions (FDPs) conserved within families but variable across them. Physicochemical variation at these FDPs was used to construct a sequence similarity network, revealing distinct functional clusters that guided representative selection.
Candidate prioritization further integrated metagenomic biome metadata to enrich for thermostable enzymes and structural features such as loop architecture near the active site, predicted using AlphaFold2. From this rational selection pipeline, 20 diverse enzymes were selected for experimental characterization. Biochemical assays identified 9 enzymes with activity against carbamate substrates, including 2 that also exhibited activity against polyurethane polymers. These results demonstrate that combining functional family analysis, residue tunability metrics, and structural diversity enables efficient navigation of metagenomic sequence space and substantially improves hit rates in enzyme discovery. This pipeline is broadly applicable to other challenging catalytic functions where experimental screening capacity is limited.

C-P.31: EnzymeSifter: a pipeline for identifying and filtering industrial enzymes from metagenomes across multiple predicted biochemical properties
Track: Proteins and structural biology
  • Omar Darawsheh, Northumbria University, United Kingdom
  • Matthew Bashton, Northumbria University, United Kingdom


Presentation Overview: Show

Identifying candidate enzymes from environmental samples for industrial applications requires evaluating multiple biochemical properties simultaneously, including solubility, thermal stability, and pH. While individual predictors for these properties exist, integrating them into a coherent workflow is a manual and error-prone process. Here we present EnzymeSifter, a two-stage Snakemake pipeline that automates enzyme discovery, multi-property characterisation, and generates a ranked shortlist.
In Stage 1, input amino-acid sequences are optionally filtered by user-defined catalytic motifs, Pfam domain annotations, and/or predicted EC number, and optionally clustered at a user-specified sequence identity threshold using MMseqs2, to produce a non-redundant set for structure prediction. In Stage 2, predicted PDB structures are screened with EnzyMM for enzymatic activity, and all confirmed hits are characterised across 5 properties: solubility and usability (NetSolP), optimal pH (pHoptNN), and optimal temperature and melting temperature (Seq2Topt). Predictions are merged, and users may specify threshold or interval filters for any property. A MUSCLE multiple sequence alignment and neighbour-joining tree are constructed. The tree can be partitioned into a user-defined number of clades to select a representative enzyme within each clade. Representatives are enzymes that achieved the best composite score according to the properties defined in the user's input.
EnzymeSifter provides a standardised framework that reduces the effort required to go from large sequence datasets to a shortlist of promising enzymes for experimental validation.

C-P.32: High Diversity Gene Libraries Facilitate Machine Learning Guided Exploration of Fluorescent Protein Sequence Space
Track: Proteins and structural biology
  • Anissa Benabbas, University of Oregon, United States


Presentation Overview: Show

Protein language models (PLMs) are increasingly used for protein design but remain limited by the diversity and structure of available training data. Models trained on natural sequences often operate in an extrapolative regime, reducing reliability when exploring sparsely sampled regions of sequence space. Here, we test whether experimentally expanding sequence diversity can shift this regime toward interpolation. Using large-scale gene synthesis and DNA shuffling, we generate libraries spanning broad regions of fluorescent protein sequence space and identify thousands of functional blue fluorescent variants through high-throughput screening. Fine-tuning
ProtGPT2 on this dataset enables the generation of diverse fluorescent proteins, including variants that extend beyond regions occupied by known natural sequences while retaining function. Together, these results support a strategy in which experimentally expanded diversity improves the ability of machine learning models to explore and design functional proteins. This approach provides a general framework for improving machine learning–guided protein design by experimentally expanding functional sequence diversity.

C-P.33: Deep generative modeling captures maturation-dependent pairing patterns in human antibodies
Track: Proteins and structural biology
  • Lea Brönnimann, University of Bern, Switzerland
  • Thomas Lemmin, University of Bern, Switzerland
  • Chiara Rodella, University of Bern, Switzerland


Presentation Overview: Show

Understanding antibody heavy-light chain pairing is critical for decoding immune repertoire architecture and designing therapeutic antibodies, yet most sequence databases lack paired chain information. To address this gap, we developed a two-stage deep learning framework. Transformer-based language models were first pre-trained on large corpora of unpaired heavy- and light-chain sequences, then integrated into a sequence-to-sequence model to generate light chains from heavy chain input. Although native light chain recovery was moderate, generated sequences exhibited high germline identity, improved structural quality, and broader framework and complementarity-determining region coverage. Heavy chains from memory B cells generated light chains with more restricted V gene usage, reflecting maturation-dependent selection. Generated kappa light chains exhibited a trimodal similarity distribution, indicating distinct functional pairing modes from promiscuous to highly specific. Our approach demonstrates that sequence-to-sequence modeling can uncover inter-chain dependencies and generate plausible antibody pairs, providing a foundation for computational repertoire analysis and therapeutic design.

C-P.34: Enhancing protein-ligand binding site predictions: Integrating protein language models with geometric smoothing and clustering
Track: Proteins and structural biology
  • Vít Škrhák, Charles University, Czechia
  • Lukáš Polák, Department of Software Engineering, Faculty of Mathematics and Physics, Charles University, Prague, Czech Republic, Czechia
  • Marian Novotný, Department of Cell Biology, Faculty of Science, Charles University, Prague, Czech Republic, Czechia
  • David Hoksza, Department of Software Engineering, Faculty of Mathematics and Physics, Charles University, Prague, Czech Republic, Czechia


Presentation Overview: Show

The identification of ligand binding sites (LBS) is a cornerstone of computer-aided drug design. While Protein Language Models (pLMs) have recently demonstrated high performance in classifying binding residues, their primary limitation remains their ""residue-centric"" focus, which frequently produces spatially disjointed or fragmented predictions when mapped onto 3D structures.

We introduce Seq2Pocket, a framework designed to bridge the gap between sequence-based classification and structural pocket continuity. Our approach utilizes a fine-tuned pLM augmented by an embedding-aware smoothing classifier and a clustering algorithm. To mitigate prediction disjointedness, we propose the Pocket Fragmentation Index (PFI), a metric used to optimize the mapping between predicted residues and binding cavities.

Experimental results on the scPDB dataset show that Seq2Pocket achieves state-of-the-art performance, with a significant 11% improvement in DCC recall over current methods. Further validation on the LIGYSIS and CryptoBench benchmarks confirms that our framework not only maintains high performance but also provides the structural coherence necessary for practical downstream drug discovery workflows.

C-P.35: PMGen: From Peptide-MHC Structure Prediction to Peptide Generation
Track: Proteins and structural biology
  • Amir H. Asgary, Quantitative and Computational Biology Group, Max Planck Institute for Multidisciplinary Sciences, Germany
  • Johannes Soeding, Quantitative and Computational Biology Group, Max Planck Institute for Multidisciplinary Sciences, Germany


Presentation Overview: Show

Accurate structural modeling of peptide-MHC (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited.

We introduce PMGen (Peptide-MHC Generator), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies,Initial Guess and Template Engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core C-alpha RMSDs of 0.54 A for MHC-I and 0.33 A for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher-affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 10,216 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.19 to 0.40, highlighting the value of accurate predicted structures for downstream machine learning.

PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

C-P.36: Mapping targetable sites on the human surfaceome for the design of novel binders
Track: Proteins and structural biology
  • Hamed Khakzad, Inria, France


Presentation Overview: Show

The human cell surfaceome, integral to cell communication and disease mechanisms, presents a prime target for therapeutic intervention. De novo protein binder design against these cell surface proteins offers a promising yet underexplored strategy for drug development. However, the vast search space and limited data on natural or competitive binders have historically limited experimental success. In this study, we systematically analyzed the entire human surfaceome, identifying approximately 4,500 targetable sites and introducing potential binding seeds for initiating protein design applications. To validate these seeds, we implemented two experimental approaches (protein scaffolding and peptide cyclization) on three representative targets (FGFR2, IFNAR2, and HER3). Our results revealed a high success rate, showing that seeds provide valuable starting points for binder design against our identified targetable sites, as well as the need for constant improvements of computational protein design pipelines utilizing machine learning and physics-based methods. Additionally, we present SURFACE-Bind, an interactive database offering open access to all generated data. The high-throughput computational design methods and target-specific binder seeds established here pave the way for a new generation of targeted therapeutics for the human surfaceome.

C-P.37: Mitigating Functional Classification Hallucination in Protein Language Models Through Target-Decoy Training
Track: Proteins and structural biology
  • Ho-Jin Gwak, Hankuk University of Foreign Studies, South Korea
  • Xiaofang Jiang, National Institutes of Health, United States
  • Ikbeom Jang, Hankuk University of Foreign Studies, South Korea
  • LeAnn Lindsey, Lawrence Berkeley National Labs, United States


Presentation Overview: Show

Protein language model (pLM)-based classifiers are increasingly used for functional annotation, but their behavior on biologically meaningless sequences remains poorly characterized. Here, we investigate this problem in phage protein function prediction using PHROGs functional categories and three representative pLMs. We show that classifiers trained on pLM embeddings can assign high-confidence functional labels to shuffled or reversed decoy sequences, revealing a hallucination-like false positive problem that cannot be fully resolved by simple confidence thresholding. To address this limitation, we introduce a target-decoy training strategy in which shuffled decoy proteins are incorporated as explicit negative examples during classifier training. This approach enables the resulting our model to reject biologically uninformative inputs, reducing decoy false positive rates to levels comparable to alignment-based methods such as Pharokka and Phold, while largely preserving classification performance on genuine PHROG-annotated proteins. Notably, a model trained only with shuffled decoys also rejects reversed decoys, suggesting that target-decoy training generalizes beyond a specific synthetic artifact. When applied to large-scale phage protein datasets, PPAM-TDM increases annotation coverage in INPHARED and PhageScope relative to alignment-based baselines, while maintaining controlled specificity. These results demonstrate that target-decoy training is a simple and broadly applicable strategy for improving the reliability of pLM-based protein function classifiers in open-world annotation settings.

C-P.38: Assessing the stability and oligomerisation of a β-hairpin through gas-phase molecular dynamic simulation
Track: Proteins and structural biology
  • Stijn De Schepper, Uppsala University, Sweden
  • Erik Marklund, Uppsala University, Sweden


Presentation Overview: Show

Protein self-assembly into supramolecular clusters is involved in a wide range of processes like the amyloid plaque formation, but still not fully understood. Governed by noncovalent interactions, clusters can experience major structural rearrangements upon transitioning into gas phase. As a model of these cluster forming molecules, we study the antimicrobial peptide Protegrin-1 (PG1), which adopts a well-defined beta-hairpin fold, stabilised by two disulfide bridges. Ion mobility mass spectrometry (IM-MS) studies have shown it assembles into mono-, di-, tri- and tetramers, and that reducing the disulfide bridges results in depletion of the oligomeric states and loss of beta-sheet content. Nevertheless, an accurate atomistic understanding is still missing.
Recent advances in the field of gas-phase molecular dynamics simulations, considering the changed electrostatic forces in electrospray ionisation due to the loss of the dielectric medium enable us to dive deeper into the assembly of PG1. The simulations highlight that the reduction of PG1 leads to a lower gas-phase stability but does not seem to lead to a significant increase in collision cross section (CCS). Surprisingly, an increase in β-sheet content is observed however structurally different from the original β-hairpin motif. Additionally, different protonation states in gas phase tremendously influence the simulations and thus complicate the comparison to experimental data from IM-MS and gas-phase IR-spectroscopy.
This work will inform a better understanding of molecular scale changes influencing oligomerization and aid understanding amyloid plaques involved in neurodegenerative diseases.

C-P.39: Breaking the Data Bottleneck: Leveraging Transfer Learning for data-scarce Post-Translational Modifications
Track: Proteins and structural biology
  • Yannick Hartmaring, Hasso Plattner Institute for Digital Engineering, Digital Engineering Faculty, University of Potsdam, Germany
  • Shengbo Wang, European Molecular Biology Laboratory - European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Juan Antonio Vizcaino, European Molecular Biology Laboratory - European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Christoph N. Schlaffner, Hasso Plattner Institute for Digital Engineering, Digital Engineering Faculty, University of Potsdam, Germany
  • Bernhard Y. Renard, Hasso Plattner Institute for Digital Engineering, Digital Engineering Faculty, University of Potsdam, Germany


Presentation Overview: Show

The large functional diversity of the proteome is possible through post-translational modifications (PTMs) resulting in up to one million different proteoforms. This large diversity harbours the challenge of considering all possible combinations of around 300 different PTMs when identifying proteomic mass spectra. To combat this, AHLF, a Deep Learning binary classifier, trained on 10.5 million Phosphorylation spectra can stratify unseen mass spectra, for targeted searches with and without modification. For Phosphorylation as the best studied PTM large amounts of data are readily available. However, other PTMs such as Acetylation and Ubiquitination suffer from drastically fewer studies, and for Deep Learning commonly insufficient data. To combat this problem, we applied a transfer-learning approach utilizing the AHLF model pre-trained on Phosphorylation to build models for Acetylation and Ubiquitination.
Our Ubiquitination set containing around 2 million high quality modified peptide-spectrum matches (PSMs) was curated from 11 publicly available ubi-enriched datasets, while the Acetylation set contained 0.5 million modified PSMs from 9 public enriched datasets. The fine tuned models for Ubiquitination and Acetylation achieve AUCs of 0.90 and 0.87, respectively. Further artificially rescued training sets show that around 28,500 modified PSMs (0.3% of the Phosphorylation dataset) already result in sufficiently converging models. We also show that our fine-tuned models incorporating label free and SILAC labelled spectra outperform fine-tuned models trained on similarly sized datasets stratified by quantification method. This highlights that fine-tuning models with mixed labelling states boosts classification performance and drastically increases the availability of data for training.

C-P.40: Plant Biocuration in UniProtKB/Swiss-Prot
Track: Proteins and structural biology
  • Emmanuel Boutet, SIB - Swiss Institute of Bioinformatics, Switzerland
  • The Uniprot Consortium Uniprot, SIB - Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

The UniProt Knowledgebase (UniProtKB, https://www.uniprot.org) is a comprehensive, freely accessible resource providing high-quality protein sequences and functional data. Its expert-curated UniProtKB/Swiss-Prot section contains approximately 580,000 sequences, including around 42,000 from plants such as Arabidopsis thaliana and Oryza sativa (release 2026_01).

A. thaliana remains a key model organism in plant biology. Building on this importance, a collaborative effort led by The Arabidopsis Information Resource (TAIR) and involving UniProtKB/Swiss-Prot has contributed to an updated genome assembly of the Columbia cultivar. This assembly is now integrated into UniProtKB, ensuring consistency between genomic and protein-level data.

Within this framework, a central focus lies in the annotation of plant enzymes. Biochemical reactions are described using Rhea (https://www.rhea-db.org), which provides standardized, computable descriptions of biochemical reactions. This integration enhances the accuracy and interconnectivity of enzyme function annotations. At present, the database comprises 16,900 manually curated plant enzyme entries, including 5,993 from Arabidopsis thaliana and 3,450 from various species selected to represent diverse biosynthetic pathways.

Together, these efforts produce structured, high-quality data that enable interoperability with other resources and support metabolic modelling, multi-omics integration, and machine learning approaches for predicting enzyme functions and plant biosynthetic pathways.

C-P.41: Generating and evaluating coevolution and phylogeny-aware MSAs
Track: Proteins and structural biology
  • Anamay Samant, 1) Institute of Bioengineering, School of Life Sciences (EPFL) 2) Swiss Institute of Bioinformatics (SIB), Switzerland
  • Anne-Florence Bitbol, 1) Institute of Bioengineering, School of Life Sciences (EPFL) 2) Swiss Institute of Bioinformatics (SIB), Switzerland


Presentation Overview: Show

Molecular phylogenetic inference involves inferring evolutionary relationships from a multiple sequence alignment (MSA). Generally, a single phylogenetic tree or tree distribution is inferred for one MSA at a time by methods like maximum likelihood and Bayesian inference. Alternatively, a generalizable, supervised learning approach requires training data in the form of “MSA-ground truth tree” pairs. Phylogenetics lacks actual ground-truths, so training data needs to be generated by simulating evolution along a known tree to yield leaf sequences constituting an MSA. The tree is then by definition the ground truth tree for the generated MSA. However, the choice of the simulation method is crucial for obtaining realistic MSAs from known trees. Conventionally, simulation is performed using models that assume independent evolution across all sites in a sequence, while sites that are functionally coupled tend to coevolve and feature correlations in amino-acid usage. Recognising the importance of accounting for these interactions, we developed a Markov Chain Monte Carlo (MCMC) based simulation method that incorporates interactions flexibly, using Potts models, or single-sequence protein language models (PLMs), or MSA-based PLMs. We compared MSAs obtained from these three model types in terms of their closeness to protein families of interest, diversity within generated sequences and novelty when compared to natural sequences. Results indicate more favourable metrics when using the PLM-based generation methods, with the MSA-based PLM having an edge when considering shallow protein families. We are now using these coevolution-aware MSAs as more realistic training data to develop machine-learning models for tree inference.

C-P.42: AI-driven structural modeling reveals hidden functional and evolutionary relationships in divergent dsRNA viruses
Track: Proteins and structural biology
  • Edouard De Castro, SIB Swiss Institute of Bioinformatics, Switzerland
  • David Moi, SIB Swiss Institute of Bioinformatics, Switzerland
  • Gerardo Tauriello, SIB Swiss Institute of Bioinformatics, Switzerland
  • Paul Thomas, SIB Swiss Institute of Bioinformatics, Switzerland
  • Jelle Matthijnssens, Rega Institute, Laboratory of Viral Metagenomics, Belgium
  • Houssam Attoui, The National Research Institute for Agriculture, Food and Environment (INRAe), France
  • Philippe Le Mercier, SIB Swiss Institute of Bioinformatics, Switzerland
  • Fauziah Mohd Jaafar, The National Research Institute for Agriculture, Food and Environment (INRAe), France


Presentation Overview: Show

Double-stranded RNA (dsRNA) viruses harbor numerous proteins whose functions remain poorly understood due to extreme sequence divergence that defeats conventional annotation methods. Here, we demonstrate that AI-driven structural modeling systematically overcomes this barrier, enabling functional assignment across highly divergent viral proteomes.
Using AlphaFold3 combined with Foldseek structural similarity searches, we predicted three-dimensional structures of proteins from representative Reovirales and Ghabrivirales members. Structure-based analyses identified multiple previously uncharacterized virion components—inner capsid, outer capsid, and capping enzymes (turret proteins)—providing a near-complete structural map despite negligible sequence similarity to known homologs. Critically, structural modeling enabled the first structure-based phylogenetic reconstruction of capsid proteins, circumventing sequence-based limitations.
Remarkably, structural modeling of Micromonas pusilla reovirus (MpRV) revealed that its outer capsid protein adopts a fold closely related to the birnavirus capsid, despite these viruses being deeply divergent. This discovery suggests ancient horizontal gene transfer and demonstrates that capsid modules can be exchanged between distantly related dsRNA viruses.
Our results demonstrate that AI-guided structural modeling not only expands functional annotation of viral proteomes but provides a powerful framework for exploring deep evolutionary relationships among highly divergent proteins. This approach reveals hidden functional and evolutionary connections within the virosphere that remain inaccessible to traditional genomic methods, opening new avenues for understanding viral evolution and diversity.

C-P.43: The sequence-structure landscape of antibody framework regions
Track: Proteins and structural biology
  • Teodora Christina Purice, Institute of Biochemistry of the Romanian Academy, Romania
  • Anca Iacob, Institute of Biochemistry of the Romanian Academy, Romania
  • Laurentiu Spiridon, Institute of Biochemistry of the Romanian Academy, Romania
  • Andrei-Jose Petrescu, Institute of Biochemistry of the Romanian Academy, Romania


Presentation Overview: Show

Due to the low antibody scaffold stability, aggregation remains a major problem in therapeutic development. State-of-the-art AI-MD aggregation predictors are constrained by limited available framework region (FR) templates, while full-antibody modeling tools primarily focused on hypervariable complementarity-determining regions (CDRs) and hinges can introduce FR inaccuracies, degrading overall stability. Addressing these problems requires a data-driven landscape of the FR sequence and structural space.
We present here results on the development of a pipeline extracting information on >3000 human and mouse immunoglobulin structures (CryoEM, X-Ray) from SabDab stratified by interaction state, chain, and domain (Fv, Fab, Fc, full Ig). AHo-aligned, outlier-filtered FR sequences were profiled (PSSM heatmaps) and clustered (30-90%, MMseqs2), visualized as network graphs, distance heatmaps, and dendrograms. Secondary structure assignment step yields hydrogen-bond maps and similarity clustering; Ramachandran and bond-angle torsion analyses provide torsion-angle distributions, statistical overview, and similarity clustering. Sequence-structure correlations generated a non-redundant template set. Structure prediction benchmarking has been initiated with RaptorX-Single, with ABodyBuilder-3, AlphaFold-3, and Ibex to follow.
Outputs confirm canonical FR features (C23, W43, C106) and composition patterns. Light chain FRs exhibit greater redundancy (20% templates required for full set coverage) than heavy chains (30-34%). Multi-threshold clustering was used to identify substructure families. DSSP, contact maps, internal coordinates analyses and structural similarity clustering have been obtained across the full set. RaptorX-Single predictions have been generated for all Fvs.
This empirical picture of FR sequence and structure space provides a foundation to build and implement a robust modeling workflow for full antibodies and antibody-antigen complexes.

C-P.44: Dynamic Behavior of the RAG2 Acidic Region and Its Possible Functional Significance
Track: Proteins and structural biology
  • Anca-L Iacob, Institute of Biochemistry of the Romanian Academy, Romania
  • Eliza - Cristina Martin, Yale University, United States of America
  • Andrei - Jose Petrescu, Institute of Biochemistry of the Romanian Academy, Romania


Presentation Overview: Show

The RAG recombinase drives V(D)J recombination and adaptive immunity, yet its evolutionary origin from a Transib-family transposon means it retains a latent transposase activity threatening genomic stability and contributing to chromosomal translocations and leukemias. Although molecular domestication has introduced suppressive adaptations, the mechanistic role of the RAG2 acidic hinge (AH), an intrinsically disordered region linking the Kelch and PHD domains remains poorly understood.
The AH was modeled in a fully extended state from the resolved RAG2 core using Modeller, then explored through 70 independent molecular dynamics simulations (100 ns each) with OpenMM, CHARMM36 force field, and implicit solvent at 310K. Trajectories were clustered to identify topological states and mapped onto the RAG tetramer surface.
Five recurrent conformational states provide a structural framework for understanding AH-mediated inhibition. In the most prevalent clusters, the AH localizes above the RAG1 DNA-binding groove, creating a steric and electrostatic barrier blocking target DNA acquisition and preventing the U-shape conformation required for transposition. A distinct cluster shows the AH intercalated at the lateral RAG1/RAG2 interface, acting as an allosteric wedge restricting inter-subunit flexibility essential for transposition. One particularly informative cluster shows the AH extended toward the RAG1 basic N-terminal region.
The acidic hinge thus emerges as a dynamic regulatory element whose electrostatic steering by basic surface patches positions it strategically to inhibit transposition and safeguard genomic stability during V(D)J recombination.

C-P.45: Mapping Evolutionary Switches Driving Functional Diversification in AsnC-like Transcription Factors
Track: Proteins and structural biology
  • Luc Lafrenaye, Institute of Molecular Systems Biology, ETH Zürich, Switzerland
  • Pedro Beltrao, Institute of Molecular Systems Biology, ETH Zürich, Switzerland
  • Julian Trouillon, Institute of Molecular Systems Biology, ETH Zürich, Switzerland


Presentation Overview: Show

Deciphering the transcriptional regulatory code requires understanding how a transcription
factor's amino acid sequence dictates its DNA-binding specificity. AsnC-like transcription
factors (InterPro: IPR019888) form an ancient regulatory superfamily that offers a dense
evolutionary sampling for this purpose. Here, we leverage the phylogeny of the AsnC-like
family to systematically map this functional diversification.
To identify the key changes causing functional diversification, we used the burst after
duplication divergence metric with ancestral sequence reconstructions. This approach
highlights "evolutionary switches"
- residues conserved within subfamilies but different
between them, suggesting strong functional importance. By mapping the high-scoring
residues onto structural models, we can deduce their specific biological roles: switches at
the protein-DNA interface suggest residues that determine specificity, while those at
effector-binding or dimerization sites show the evolution of metabolic sensing and complex
assembly.
Beyond this, we are broadening our analysis to explore the larger factors and timing of these
functional shifts. By characterizing the physicochemical nature of the mutations at switch
sites, we aim to gain understanding of how specific substitutions alter function. By linking the
identified evolutionary switches with known DNA-binding specificities, effector molecules,
and the host species' environments, we can deduce the selective pressures that drive
regulatory adaptation. Mapping these switches across the phylogeny also allows us to
estimate when these adaptations emerged. This evolutionary approach ultimately aims at
clarifying how regulatory domains take on novel functions in the AsnC-like transcription
factor family, offering a sound basis for understanding adaptation through protein evolution.

C-P.46: Residue-Level Attributions in Protein Language Models Do Not Recover Allergen Epitopes
Track: Proteins and structural biology
  • Jianzhou Yao, Swiss Institute of Allergy and Asthma Research, Davos; ETH Zurich, Zurich, Switzerland
  • Anxiong Song, Swiss Institute of Allergy and Asthma Research, Davos; ETH Zurich, Zurich, Switzerland
  • Katja Baerenfaller, Swiss Institute of Allergy and Asthma Research, Davos; Swiss Institute of Bioinformatics, Lausanne, Switzerland
  • Damir Zhakparov, Swiss Institute of Allergy and Asthma Research, Davos; Swiss Institute of Bioinformatics, Lausanne, Switzerland


Presentation Overview: Show

Background. Protein language models (PLMs) achieve state-of-the-art allergenicity classification, yet the molecular basis of these predictions remains uncharacterized. Residue-level attribution methods, including Integrated Gradients (IG), are widely interpreted as localizing immunologically relevant regions, but this claim has only been supported by comparisons to known epitopes and has not been quantitatively assessed against curated immunological ground truth.
Methods. We developed a residue-level benchmark for assessing immunological faithfulness in PLM-based allergenicity prediction models, defined as concordance between residue-level attribution scores and experimentally validated MHC Class II epitopes from the Immune Epitope Database (IEDB). We evaluated attribution scores across ESM-2-based classifiers, including a multi-task architecture with an auxiliary residue-level head trained under direct epitope supervision. To differentiate immunological faithfulness from model faithfulness, we conducted IG-guided masking and in silico saturation mutagenesis at high-attribution positions.
Results. Residue-level attribution scores exhibited near-random concordance with IEDB-annotated epitopes, despite high protein-level performance. Multi-task-learning recovered epitope signal in the auxiliary residue head but did not increase epitope concordance of classifier attributions, indicating epitope-relevant features are not recruited by the classification objective even when available during optimization. Masking high-attribution residues reduced prediction confidence, confirming attribution faithfulness to model decisions. Saturation mutagenesis revealed sensitivity to physicochemical properties and local compositional context rather than epitope-defining substitutions.
Conclusion. Model and immunological faithfulness are distinct: attributions can reflect decision drivers while failing to recover immunologically meaningful sequence properties. Residue-importance maps cannot be treated as biological explanations without epitope-based validation. We propose quantitative faithfulness benchmarking as necessary for interpretability evaluation in allergenicity prediction.

C-P.47: Computational search for Myelin-associated Glycoprotein (MAG) binders that could modulate the neuroprotective properties of oligodendrocytes in neurodegenerative contexts
Track: Proteins and structural biology
  • Carmela Felippa Ambort, Department of Theoretical and Computational Chemistry, School of Chemical Sciences, National University of Córdoba, Argentina
  • Rodrigo Quiroga, Department of Theoretical and Computational Chemistry, School of Chemical Sciences, National University of Córdoba, Argentina


Presentation Overview: Show

Myelin-Associated Glycoprotein (MAG/Siglec-4) is a cell-surface lectin belonging to the Siglec family, which binds sialic acid–containing glycans. MAG is expressed in myelinating oligodendrocytes of the central nervous system and is localized in the periaxonal space. It shows high specificity for the Neu5Acα2-3Galβ1-3GalNAc motif present in gangliosides such as GT1a and GD1b. Structurally, this interaction is mediated by a conserved Arg118 residue in the V domain, which forms a salt bridge with the sialic acid carboxyl group. Upon activation, MAG promotes glutamate reuptake after injury, contributing to neuroprotective and antioxidant effects in oligodendrocytes and neurons. This function is particularly relevant in neurodegenerative conditions such as stroke and multiple sclerosis.
The aim of this study is to identify potential MAG activators or inhibitors with high affinity and selectivity through molecular docking and virtual screening of diverse compound libraries (e.g., ZINC20). For this purpose, the 2VINARDO scoring function, an enhanced version of AutoDock Vina, is employed. Developed by Quiroga and Villareal, 2VINARDO improves the description of non-covalent interactions by expanding atom types and interaction parameters.
Validation of the scoring function was performed using a re-docking protocol with seven crystallographic Siglec–ligand complexes, including MAG. In all cases, ligand poses were accurately predicted (RMSD < 2 Å), preserving key interactions such as hydrogen bonds involving Arg118. Additional analyses include binding site flexibility and correlation with experimental binding affinities (Kd).
These results support the application of 2VINARDO for large-scale virtual screening, aiming to identify selective MAG modulators with potential therapeutic relevance.

C-P.48: A novel analysis pipeline for scFv discovery from long-read platforms
Track: Proteins and structural biology
  • Ozge Gizlenci, AstraZeneca, United Kingdom
  • Gareth Griffin, AstraZeneca, United Kingdom
  • Jurgen Haas, AstraZeneca, United Kingdom


Presentation Overview: Show

Antibody discovery increasingly relies on sequencing to validate candidates and characterise repertoire diversity. While next-generation sequencing enables fast, cost-effective, high-throughput analysis, long-read platforms are becoming especially valuable for single-chain variable fragment (scFv) discovery because they can capture full-length constructs and preserve VH-VL pairing. This is particularly important for candidates designed through machine learning-guided workflows.

We developed a novel platform- and UMI-aware long-read analysis workflow for scFv discovery using the long-read technologies recently implemented in-house, including PacBio HiFi and Oxford Nanopore Technologies (ONT). The pipeline performs preprocessing, platform-specific consensus generation, translation of scFv constructs, domain annotation with ANARCI, and automated CDR3 extraction, producing quantitative summaries and exportable TSV/HTML reports for downstream library characterisation. It consistently recovers key scFv features, supports repertoire-level CDR3 analysis, and provides high-level quality metrics across datasets.

Long-read sequencing offers clear advantages for antibody discovery, but it also presents challenges. Compared with short-read methods, long-read data can be more error-prone, particularly for ONT, and may show variable quality profiles across reads. This can complicate accurate domain identification, CDR parsing, translation, and clonotype quantification. Long-read workflows may also have lower throughput, higher cost per read, and greater computational demands for consensus generation and error correction. In addition, currently available tools rarely provide end-to-end paired-domain annotation and quantitative reporting tailored to scFv long-read datasets.

Next, we plan to extend the pipeline to Illumina's emerging long-read approach, evaluate ANARCI II and benchmark against tools such as RIOT, with the goal of integrating the workflow into a broader modular NGS platform.

C-P.49: Investigating Enzyme Function by Geometric Matching of Catalytic Motifs
Track: Proteins and structural biology
  • Raymund Hackett, Leiden University Medical Center, European Bioinformatics Institute (EMBL-EBI), Netherlands
  • Ioannis Riziotis, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Martin Larralde, Leiden University Medical Center, Netherlands
  • António J. M. Ribeiro, European Bioinformatics Institute (EMBL-EBI), Portugal
  • Georg Zeller, Leiden University Medical Center, Netherlands
  • Janet Thornton, European Bioinformatics Institute (EMBL-EBI), United Kingdom


Presentation Overview: Show

The rapidly growing universe of predicted protein structures offers opportunities for data driven exploration but requires computationally scalable and interpretable tools. We developed a method, Enzyme Motif Miner, to detect catalytic features in protein structures, providing insights into enzyme function and mechanism. A library of 6780 3D coordinate sets describing enzyme catalytic sites, referred to as templates, has been collected from manually curated examples of 762 enzyme catalytic mechanisms described in the Mechanism and Catalytic Site Atlas. We implemented RMSD and residue orientation filters to differentiate catalytically informative matches from spurious ones. We validated this approach on a non-redundant set of high quality experimental (n=3751, <40% amino acid identity) enzyme structures with well annotated catalytic sites as well as predicted structures of the human proteome. We show that matching catalytic templates is more sensitive than sequence- and 3D-structure-based approaches in identifying homology between distantly related enzymes. Since geometric matching does not depend on conserved sequence motifs or even common evolutionary history, we are able to identify examples of structural active site similarity in highly divergent and possibly convergent enzymes. Such examples make interesting case studies into the evolution of enzyme function. Though not intended for characterizing substrate-specific binding pockets, the speed and knowledge-driven interpretability of our method make it well suited for expanding enzyme active-site annotation across large predicted proteomes. Enzyme Motif Miner is available as a python module at https://github.com/rayhackett/enzymm and as a webserver at https://www.ebi.ac.uk/thornton-srv/m-csa/enzymm.

C-P.50: Unveiling Recurrent Binding Sites in H1N1 Nucleoprotein via Ensemble-Based Pocket Clustering
Track: Proteins and structural biology
  • Xinyu Qi, Université Paris Cité, CNRS, Inserm, Unité de Biologie Fonctionnelle et Adaptative, F- 75013 Paris, France, France
  • Inés Sabine Rahali, Université Paris Cité, CNRS, Inserm, Unité de Biologie Fonctionnelle et Adaptative, F- 75013 Paris, France, France
  • Sandie Munier, Institut Pasteur, Université Paris Cité, Lyssavirus Epidemiology and Neuropathology Unit, F-75015 Paris, France, France
  • Anne Badel, Université Paris Cité, CNRS, Inserm, Unité de Biologie Fonctionnelle et Adaptative, F- 75013 Paris, France, France
  • Delphine Flatters, Université Paris Cité, CNRS, Inserm, Unité de Biologie Fonctionnelle et Adaptative, F- 75013 Paris, France, France
  • Anne-Claude Camproux, Université Paris Cité, CNRS, Inserm, Unité de Biologie Fonctionnelle et Adaptative, F- 75013 Paris, France, France


Presentation Overview: Show

Influenza A nucleoprotein (NP) is a conserved multifunctional protein involved in RNA binding and oligomerization, making it an attractive antiviral target. However, static structures provide only a partial view of its ligandable regions, as conformational variability can alter pocket accessibility and predicted druggability.
Here, we characterized the binding site landscape of H1N1 NP using molecular dynamics simulations combined with ensemble-based pocket detection and residue-based clustering. Across 753 conformations, 16,404 detected pockets were clustered by residue composition to reconstruct recurrent binding site candidates and assess their recurrence and predicted druggability. Seven recurrent sites were retained based on recurrence and cluster-level predicted druggability and were further characterized according to residue-level organization, plasticity, and their position relative to NP–RNA and NP–NP interface regions. These sites displayed distinct structural behaviors, included stable compact NP head sites, RNA-interface-associated sites, and a dynamic extended site spanning the inter-domain groove. Comparison with the static reference structure distinguished stable, shifted, and dynamically emerging sites, including regions not apparent in the reference conformation.
Overall, this study provides an ensemble-based structural framework to reconstruct and prioritize recurrent ligandable regions in influenza A NP. By integrating recurrence, predicted druggability, site organization, and structural plasticity, it refines the description of NP binding sites in H1N1 and opens the way to ongoing comparative work with H5N1 NP to determine which recurrent binding sites are conserved, shifted, or subtype-specific across pathogenic influenza A subtypes.

C-P.51: Tokenization-Aware Protein Language Modeling with Evolution-Guided Units and Dynamic Biological Knowledge Fusion
Track: Proteins and structural biology
  • Amirreza Sattarzadeh Khanehbargh, Boğaziçi University, Türkiye
  • Burak Suyunu, Boğaziçi University, Türkiye
  • Özdeniz Dolu, Boğaziçi University, Türkiye
  • Arzucan Özgür, Boğaziçi University, Türkiye


Presentation Overview: Show

Protein language models (PLMs) have become key tools for computational protein analysis, but most still rely on single amino-acid tokens as the default representation unit. While simple and effective, this choice can limit biological expressiveness and may interact in non-trivial ways with downstream adaptation strategies. We present a tokenization-aware framework that compares character-level tokenization, frequency-based subword tokenization such as BPE and Unigram, and PUMA, a mutation-aware tokenizer that organizes sequence patterns into mutation-informed token families.

Using multiple pretrained backbones, we study various strategies for implementing tokenization in protein language modeling, including frozen encoders, full fine-tuning, and parameter-efficient LoRA. Experiments cover representative tasks in computational biology, including subcellular localization, protein function classification, and stability prediction. Beyond accuracy, we assess memory usage, sequence compression, compute cost, and interpretability. We further examine how token-level representations can be enriched by integrating biochemical, functional, and evolutionary priors through task-aware knowledge fusion. All in all, we provide an analysis of the effect of tokenization on protein representation learning performance across computational and biological dimensions.

Our results indicate that PUMA can outperform frequency-based subword tokenization in several settings, while vocabulary size has a substantial effect on both predictive performance and efficiency. Amino-acid tokenization remains a strong baseline overall; however, at certain vocabulary sizes, PUMA can match or surpass amino-acid-level representations, suggesting that biologically informed tokenization can offer practical advantages when appropriately configured.

Overall, this work provides a structured comparison of tokenization choices, adaptation regimes, and representation quality in PLMs. By analyzing their combined effects across multiple biological tasks, our framework offers practical guidance for selecting tokenization and fine-tuning strategies and introduces an extensible benchmark for biologically informed protein sequence modeling.

C-P.52: Computational insights into enhanced deubiquitination by SARS-CoV-2 PLpro K232Q – A structure - function study
Track: Proteins and structural biology
  • Janani Ganesh, Homi Bhabha National Institute, India
  • Rimanshee Arya, Homi Bhabha National Institute, India
  • Vishal Prashar, Homi Bhabha National Institute, India
  • Mukesh Kumar, Homi Bhabha National Institute, India


Presentation Overview: Show

Papain-like protease (PLpro) of SARS-CoV-2, a domain of non-structural protein 3 (Nsp3), plays dual role in its pathogenesis by processing viral polyproteins and modulating host immunity through deubiquitination and deISGylation. To examine how mutations influence these functions, we performed large-scale computational analysis of ~14 million Nsp3 sequences from the GISAID database using custom Python pipelines (Arya, et al., 2023: Microbial Pathogenesis, 185, 106460). This analysis identified five mutations (A145D, P77L, P77S, V187A, and K232Q) with significant global prevalence, associated with multiple variants of concern. Among these, we found that the K232Q mutation significantly increased (~four-fold) deubiquitination activity compared to the wild-type enzyme in our biochemical assays. Sequence comparison with SARS-CoV-1 suggested that this substitution represents a reversion to its ancestral residue, consistent with the higher deubiquitination activity observed in SARS-CoV-1 PLpro. Structural mapping placed residue 232 at the ubiquitin-binding interface, prompting molecular dynamics simulations of the PLpro-ubiquitin complex (PDB: 7RBR) which revealed a shift in interaction behaviour. In the wild-type, K232 alternates between interactions with A46(Ub) and Y207(PLpro), whereas Q232 forms more balanced and persistent contacts with both residues, along with a shorter interaction distance to A46(Ub). Trajectory analysis using MDAnalysis indicated changes in substrate positioning near the catalytic site and buried surface area, while DCCM and MM-PBSA analyses further revealed differences in correlated motions and binding energetics between mutant and wild-type systems. Together, these results provide a structural basis for the enhanced catalytic activity of K232Q PLpro to uncover subtle changes in enzyme–substrate dynamics due to mutation.