View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
C-P.01: Interpreting the function of germline missense variants for the BAP1 gene using protein centric
information.
Track: Proteins and structural biology
-
Pierre Marchal, Department of Medical Oncology, Inselspital, Bern University Hospital, DBMR, GCB,
University of Bern, Switzerland., Switzerland
-
Joel Spies, Department of Medical Oncology, Inselspital, Bern University Hospital, DBMR, University of
Bern, Switzerland., Switzerland
-
Lena Bürgi, Department of Medical Oncology, Inselspital, Bern University Hospital, University of Bern,
Bern, Switzerland., Switzerland
-
Kasun Samarasinghe, SIB Swiss Institute of Bioinformatics and University of Geneva, Switzerland,
Switzerland
-
Lydie Lane, SIB Swiss Institute of Bioinformatics and University of Geneva, Switzerland, Switzerland
-
Ferdinando Cerciello, Department of Medical Oncology, Inselspital, Bern University Hospital, DBMR,
University of Bern, Switzerland., Switzerland
Presentation Overview: Show
Introduction: BAP1 is a tumor suppressor and deubiquitinating protein, which misfunction is associated with
the rare hereditary BAP1 tumor predisposition syndrome. Current variants effect predictors (VEP) tools are
capable to predict variants effect but lack specificity for the rare cancers associated with BAP1 mutations.
Here, we propose a protein centric approach based on highly curated protein information integrated in the
model ProFeaTrOL to predict the effect that missense variants might have on the function of BAP1.
Methods: We selected functional relevant BAP1 features from the knowledgebase neXtProt to establish the
Protein Feature Training based for Oncology Likelihood (ProFeaTrOL) model. We trained and validated ProFeaTrOL
based on literature available datasets of BAP1 mutations and tested its performance on ClinVar reports. we
further verified the model on the BAP1 functionally associated protein BRCA1 and the independent gene TP53.
Results: ProFeaTrOL presented an AUC of 0.84 for discriminating functional and non-functional BAP1 variants of
the validation set and of 0.89 in the testing set from ClinVar. As a comparison, PolyPhen-2 and SIFT presented
AUC of 0.77 and 0.68 in the validation and 0.84 and 0.78 in the testing. Despite being trained on BAP1, if
applied to the functionally related protein BRCA1, the model showed a high AUC of 0.81, while for TP53 it was
only modest, confirming the specificity of the features selected for ProFeaTrOL.
Conclusion: The protein centric approach of ProFeaTrOL offers a strategy for the accurate SNVs effect
prediction specific for the rare cancer conditions associated with germline variants of BAP1.
C-P.02: Generalizable Protein Language Models for Complex Variant Prediction via Multi-Task Epistatic
Learning
Track: Proteins and structural biology
-
Fabio Mazza, University of Trento, Italy
- Mattia Tenuti, University of Trento, Italy
- Gianluca Lattanzi, University of Trento, Italy
- Alessandro Romanel, University of Trento, Italy
Presentation Overview: Show
Understanding the impact of variants on protein stability and function is a core challenge in computational
biology. Protein language models (PLMs) have demonstrated strong zero-shot capabilities in variant effect
prediction. However, robust supervised fine-tuning across diverse functional readouts remains challenging.
Furthermore, epistatic interactions, where the effect of a single mutation depends on another one, are rarely
modeled and evaluated explicitly in PLM-based approaches, and multi-mutant training has usually focused on
thermodynamic stability.
In this study, we present a multi-tasking fine-tuning framework for ESM-2 that integrates a Deep Mutational
Scanning (DMS) pairwise ranking objective across multiple types of DMS assays with an auxiliary MLM objective
that preserves ESM-2 language-modeling perplexity, supporting generalization capabilities. Rather than
training separate prediction heads from scratch, we reuse the pretrained language-modeling head and learn
lightweight task-specific matrices that bias the LM logits for each task. DMS scores are then computed as
unmasked pseudo log-likelihoods of the sequences with the corresponding bias. To explicitly capture
non-additivity, we introduce an epistatic loss term that evaluates mutational cycles to directly match
predicted and experimental epistatic terms.
Benchmarking on held-out DMS assays, including assays containing INDELs, shows that our model achieves the
best overall performance among the methods evaluated in ProteinGym. Performance gains are particularly
pronounced on multi-mutant assays, where explicit epistasis training substantially improves correlation with
experimentally derived specific epistasis effects, and also extend to unseen assay categories, including
expression and activity assays.
C-P.03: Do AlphaFold Ensembles Capture Side-Chain Fluctuations? A Large-Scale Benchmark Against MD and PDB
Homologs
Track: Proteins and structural biology
-
Alexander Korsunsky, Linkoping University, Sweden
- Nicholas Pearce, Uppsala University, Sweden
- Bjorn Wallner, Linkoping University, Sweden
Presentation Overview: Show
Protein side-chain fluctuations are essential for cellular processes such as enzymatic catalysis and molecular
recognition. Accurately modeling these dynamics is vital for drug design and interpreting experimental data
from NMR or crystallography.
In this work, we evaluate AlphaFold2 (AF2) and variants such as AlphaFold3 (AF3), AFsample2, AFsample3, and
AF2χ by generating 1,000 structures for each of 120 proteins to analyze side-chain flexibility. We compare
predicted side-chain dihedral (χ₁-χ₅) distributions, using both marginal (single-angle) and joint
(multi-angle) analyses, against MD (ATLAS) and ensembles of sequence-homologous PDB structures.
To quantify differences between ensembles, we devised a metric based on optimal transport (Wasserstein
distance) that accounts for the geometry of rotamer space while avoiding the bin-count dependency of a
standard measure from information theory (Jenson-Shannon Divergence) and enables comparisons of multivariate
(joint) distributions.
Our results show that the original structure prediction models AF2 and AF3 produce side-chain distributions
with very limited conformational variation, which is consistent with the models being trained primarily on
rigid crystal structures. While variants tuned for diversity (AFsample2, AF2χ) show increased variation, their
similarity to MD and experimental ensembles remains low.
This study provides a benchmark to determine if ensembles from structure prediction tools can accurately
replicate the physical fluctuations and highlights the current limitations of using structure prediction
models to capture the full landscape of protein side-chain dynamics.
C-P.04: Benchmarking structure prediction in the time of AI
Track: Proteins and structural biology
-
Xavier Robin, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel,
Switzerland
-
Gabriel Studer, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
-
Stefan Bienert, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
-
Janani Durairaj, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
-
Peter Å krinjar, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
-
Andrew M. Waterhouse, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel,
Switzerland
-
Gerardo Tauriello, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel,
Switzerland
-
Torsten Schwede, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
Presentation Overview: Show
Recent advances in artificial intelligence have led to major improvements in biomolecular structure
prediction, particularly for single‑chain protein targets, while also opening new and more challenging
application areas. These include protein‑protein interactions, protein‑ligand complexes, and
multi‑component assemblies. Emerging AI‑based methods address these problems, but their performance across
different target types and modelling tasks is not yet systematically established, motivating the need for
robust and adaptable benchmarking.
We present two complementary resources supporting objective and reproducible evaluation of structure
prediction methods. OpenStructure provides a flexible and extensible toolkit for macromolecular structure
comparison, enabling evaluation beyond a small set of global accuracy measures. It supports alternative
scoring schemes, local and interface‑focused analyses, and dedicated scores for protein‑ligand
interactions. The OpenStructure benchmarking framework also constitutes the basis of evaluation in
community‑wide assessments such as CASP. Building on this toolkit, CAMEO provides fully automated and
continuous benchmarking, with recent extensions enabling blind evaluation of macromolecular complexes,
including protein‑protein and protein‑ligand predictions.
Together, CAMEO and OpenStructure provide a framework for benchmarking structure prediction across a broad and
evolving range of modelling scenarios. They enable benchmarks to address central questions including how to
quantify target difficulty, how to track progress over time, and how to determine when a modelling problem can
be considered effectively solved. Separately, the framework supports broader inclusion of modern AI‑based
methods alongside traditional prediction servers, ensuring benchmarking remains representative of current
practice.
C-P.05: LLM-Assisted Development of Large-Scale Proteomics Infrastructure: An Empirical Study from PRIDE
Reanalysis
Track: Proteins and structural biology
-
David Teschner, Institute of Computer Science, Johannes Gutenberg University, Mainz, Germany
-
David Gomez-Zepeda, Helmholtz Institute for Translational Oncology, Mainz, Germany, Germany
-
Tim Maier, Institute of Computer Science, Johannes Gutenberg University, Mainz, Germany
-
Thomas Kemmer, Institute of Computer Science, Johannes Gutenberg University, Mainz, Germany
- Maximilian Sprang, Universitätsmedizin Mainz, Germany
- Asis Asis Hallab, Technische Hochschule Bingen, Germany
-
Stefan Tenzer, University Medical Center, Johannes-Gutenberg University, 55131 Mainz, Germany, Germany
-
Andreas Hildebrandt, Institute of Computer Science, Johannes Gutenberg University Mainz, Germany
Presentation Overview: Show
In data-intensive fields such as proteomics, the primary bottleneck is not only algorithm development but also
the construction and operation of scalable data infrastructure. Building such systems, including pipelines,
data layers, and validation frameworks, typically requires large and specialized teams.
We present a case study based on a fully operational pipeline for the systematic reanalysis of public timsTOF
datasets from the PRIDE repository. The system integrates multiple search engines, performs raw signal
extraction, and produces a versioned, bias-documented reference layer for downstream analysis and model
training. The pipeline was developed by a single researcher using LLM-assisted software development and
evolved under real data constraints, providing a unique setting to study infrastructure construction at
scale.
We analyze the development process through version history and code evolution and investigate system behavior
under increasing data volume to identify characteristic failure modes and validation challenges. Our
observations indicate that while LLM-assisted development enables rapid construction of complex systems, it
introduces new classes of errors and shifts the primary challenge from implementation to validation and
system-level consistency.
These results highlight the need for systematic validation strategies for LLM-assisted scientific software and
suggest that such methods may allow individual researchers to build infrastructure that previously required
dedicated teams.
C-P.06: Amyloid Landscape Explorer: A Network-Based Platform for Thermodynamic and Structural Comparison of
Amyloid Fibrils
Track: Proteins and structural biology
-
Wannes Hermans, VIB Center for AI and Computational Biology, KU Leuven, Belgium
-
Vasilis Kyriazis, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain
Institute, University of Texas, United States
-
Laxmikant Gadhe, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain
Institute, University of Texas, United States
-
Katerina Konstantoulea, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain
Institute, University of Texas, United States
- Joana Pereira, VIB Center for AI and Computational Biology, KU Leuven, Belgium
- Robin Bouchez, VIB Center for Neuroscience Leuven, Belgium
- Rodrigo Gallardo, VIB Center for Neuroscience Leuven, Belgium
-
Nikolaos Louros, Center for Alzheimer's and Neurodegenerative Diseases, Peter O’Donnell Jr. Brain
Institute, University of Texas, United States
- Joost Schymkowitz, VIB Center for Neuroscience Leuven, KU Leuven, Belgium
- Frederic Rousseau, VIB Center for Neuroscience Leuven, KU Leuven, Belgium
Presentation Overview: Show
Recent advances in cryo-electron microscopy and solid-state NMR have led to a rapid increase in
high-resolution amyloid fibril structures, providing new insights into structural polymorphism, disease
specificity, and aggregation behavior. This growing dataset calls for integrative approaches to systematically
compare amyloids beyond traditional sequence- or structure-based classifications.
Here, we present the Amyloid Landscape Explorer, an interactive extension of the Amyloid Explorer, a public
repository for structural and thermodynamic analysis of full-length amyloid fibril core structures. Within
this framework, pairwise similarities between protofilament ΔG profiles are quantified using a custom
Needleman-Wunsch-based alignment with a Gaussian similarity function and affine gap penalties, enabling
comparison of energetic profiles. These similarities are represented as a weighted network, where nodes
correspond to amyloid structures and edges reflect positive-scoring alignments. Leiden clustering identifies
thermodynamic clusters, revealing groups of amyloids with similar stability patterns. In parallel, structural
similarities are computed using pairwise structural alignments, defining structural clusters that are
incorporated as complementary metadata.
The resulting landscape highlights both concordance and divergence between thermodynamic and structural
organization. An interactive web interface enables database browsing and filtering by protein, disease, and
experimental method, with options to color and filter the network based on thermodynamic or structural
clusters. Additional features include a 3D molecular viewer with per-residue energy mapping, and
cluster-specific multiple alignment analysis. Together, the Amyloid Landscape Explorer provides a
comprehensive resource for exploring the energetic and structural basis of amyloid diversity.
C-P.07: The Encyclopaedia of protein Domains (TED) from AlphaFold2-predicted structures: expansion of CATH
domain space and insights into structural diversity
Track: Proteins and structural biology
-
Andy Lau, Department of Computer Science, University College London, London WC1E 6BT, UK; InstaDeep Ltd,
London W21AY, UK, United Kingdom
-
Nicola Bordin, Institute of Structural and Molecular Biology, University College London, London WC1E 6BT,
UK., United Kingdom
-
Shaun Kandathil, Department of Computer Science, University College London, London WC1E 6BT, UK., United
Kingdom
-
Ian Sillitoe, Institute of Structural and Molecular Biology, University College London, London WC1E 6BT,
UK., United Kingdom
- Vaishali Waman, University College London, United Kingdom
- Jude Wells, University College London, United Kingdom
-
Christine Orengo, University College London, United Kingdom
- David Jones, University College London, United Kingdom
Presentation Overview: Show
The AlphaFold Structure Database (AFDB) provides predictions for > 214 million three-dimensional structures
of full-length proteins. Protein structure is composed of one or more domains i.e. independent folding units
that can be found in multiple structural and functional contexts. We present The Encyclopedia of Domains (TED)
which is the first structural-based resource to classify domains from the AFDB (https://ted.cathdb.info/).
TED is developed as a collaboration between the Orengo and Jones group. TED combines state-of-the-art deep
learning-based domain detection, structure comparison (Foldseek) and fold recognition (FoldClass) algorithms
to identify and classify domains across the structures from AFDB. We used novel deep learning-based domain
detection methods (Merizo, Chainsaw) and UniDoc, to segment AFDB protein structures into protein domains.
TED contains over 370 million protein domains, providing domain structure coverage to over 1 million
taxa/species. Notably, 80% of the TED domains exhibit similarities with known superfamilies in the CATH
database, expanding the CATH by over 600-fold. We identified over 6,000 domains with potentially new folds,
some of which have unique protein architectures not seen previously (e.g. 11-helix propeller and 11-bladed
beta-propeller).
TED as a unique domain resource derived from AFDB, will aid in a multitude of downstream analyses including
identification of remote homologues and functional diversity. We demonstrate an analysis of TED data providing
insights into substrate-binding specificities in CoA-dependent acyltransferases in the context of
pathogens.
TED and CATH (https://www.cathdb.info/) domain annotations are now made available via AlphaFold Database
(https://alphafold.ebi.ac.uk/). CATH and TED, paves the way to understand biodiversity using protein
domains.
C-P.08: Iterative HMM-Refined Coevolution Analysis Combined with AlphaFold3 for Protein Complex
Discovery
Track: Proteins and structural biology
-
Giovanni Merici, University of Parma, Italy
- Giulia Sassi, University of Parma, Italy
- Karen Cárdenas Casillas, University of Parma, Mexico
- Kristian Salvatore, University of Parma, Italy
- Riccardo Percudani, University of Parma, Italy
Presentation Overview: Show
Understanding protein–protein interactions is essential to elucidate molecular mechanisms, yet identifying
novel complexes or interactors of known assemblies remains challenging. Here, we present a computational
pipeline integrating coevolution analysis (cotr) with AlphaFold3 (AF3) structural modelling to predict and
prioritise candidate interactions supported by both evolutionary and structural evidence.
Orthogroups were first inferred from eukaryotic proteomes using SonicParanoid2. Each orthogroup was then
modelled as a hidden Markov model (HMM) and used to iteratively search an expanded dataset of eukaryotic
proteomes, with newly identified homologs incorporated at each round to update the HMM and refine protein
presence–absence distributions across species. Cotr was subsequently applied to identify statistically
significant coevolutionary associations between proteins, yielding coevolving modules. These were
combinatorially assembled into putative complexes and evaluated using an optimised AF3 workflow. Predicted
assemblies were prioritised using confidence metrics including ipTM, pTM, and PAE.
We applied this framework to the BBSome complex, identifying PDE6D as a potential interactor from
coevolutionary clusters. AF3 multimer predictions revealed stable assemblies involving PDE6D with BBS2 and
additional subunits. Molecular dynamics simulations further supported the stability of the PDE6D–BBS2
interaction, showing persistent inter-chain contacts and stable interface distances over time. Key
interactions include a salt bridge and conserved hydrogen bonds with high occupancy across simulations.
Interface energetics and in silico mutagenesis confirmed the importance of these residues, with mutations at
the interface leading to measurable destabilisation of the complex.
Overall, this approach enables the systematic generation of experimentally testable hypotheses to discover and
prioritise novel protein–protein interactions in disease-relevant complexes.
C-P.09: From 1,000 Spiders to Designed Protein Materials: Decoding the Spider Silkome
Track: Proteins and structural biology
-
Kazuharu Arakawa, Keio University, Japan
Presentation Overview: Show
Spider silks exhibit an unparalleled combination of tensile strength, extensibility, and toughness, defining a
unique processing–property space among biopolymers. However, the sequence–property relationships
underlying this diversity remain incompletely understood. To systematically explore this space, we established
a global “1000 Spiders†initiative, collecting ~1,000 species and comprehensively characterizing their silk
genes and mechanical properties. The resulting open resource, the Spider Silkome Database, integrates
>10,000 gene sequences and thousands of material property measurements, enabling large-scale
genotype–phenotype analyses.
Bioinformatic mining identified key sequence motifs and previously underappreciated components, including
MaSp3 and a novel additive-like protein (SpiCE), that significantly influence silk performance. Motif-level
engineering validated both positive and negative correlations with properties such as strength, toughness, and
supercontraction.
Building on this dataset, we developed a data-driven materials transformation (DxMT) framework coupled with
generative AI. Due to the extreme length and repetitive architecture of spidroins, conventional protein
language models are insufficient; instead, we implemented a modular strategy that decomposes sequences into
N-terminal, repeat, and C-terminal domains, followed by retrieval-based assembly of pseudo–full-length
sequences. This enabled the construction of ~10,000 full-length candidates and improved prediction of
mechanical properties. Fine-tuned models (e.g., ESM2 + LoRA) further enhanced property prediction accuracy.
Finally, AI-guided motif design was experimentally validated through recombinant silk spinning, demonstrating
improved molecular orientation and mechanical performance. This integrated approach establishes a scalable
paradigm for decoding and designing high-performance protein materials from evolutionary diversity.
C-P.10: LIGYSIS: a resource for predictor training, benchmarking and functional characterisation of ligand
binding sites
Track: Proteins and structural biology
-
Javier Sánchez Utgés, University College London, United Kingdom
- Stuart MacGowan, University of Dundee, United Kingdom
-
Diane Lee, University of Dundee - current address: University of York, United Kingdom
- Gopal Sapkota, University of Dundee, United Kingdom
- Geoff Barton, University of Dundee, United Kingdom
Presentation Overview: Show
Reliable protein-ligand binding site definition and prediction underpin function annotation, variant
interpretation and drug discovery, yet many commonly used datasets are built from single asymmetric units,
leading to artificial crystal contacts, redundant interfaces, and incomplete site definitions. LIGYSIS is a
comprehensive dataset that aggregates unique, biologically relevant protein-ligand interfaces across the
biological assemblies of the multiple structures of a protein. The human LIGYSIS component comprises 3448
proteins, 8244 binding sites, and 65,116 ligands, and has been used as a reference set for the largest
independent systematic comparative evaluation of ligand site prediction tools to date.
To make these data accessible, we developed LIGYSIS-web – an open resource hosting the entire LIGYSIS set,
which includes 64,782 binding sites across 25,003 proteins. Users can query UniProt accession identifiers or
upload their own structures for automated site definition, characterisation, interactive 3D visualisation,
exploration, and download. Sites are characterised by evolutionary divergence, human missense variation, and
solvent accessibility-derived functional scoring, enabling prioritisation of likely functional pockets and
residues, as shown in our recent study.
Finally, LIGYSIS is increasingly being adopted to train and test models for the prediction of general binding,
allosteric, and cryptic sites, as well as being integrated in new resources. In parallel, we are using the
full LIGYSIS set to train a new ligand site predictor that uses protein language models, developing an
integration of LIGYSIS-web with Jalview, and extending our work on ligand site prediction evaluation and good
practices.
C-P.11: 3DSeqCheck: Identifying discrepancies of structure-based resources such as AlphaFoldDB in relation to
the evolving UniProt entries
Track: Proteins and structural biology
-
Ifigenia Tsitsa, Imperial College London, United Kingdom
-
Anja Conev, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Suhail Islam, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Alessia David, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Michael J E Sternberg, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London,
United Kingdom
Presentation Overview: Show
UniProt is a central repository of protein sequences and annotations, with entries being updated several times
a year as new sequencing evidence is collected. By contrast, protein structure resources often evolve at a
different pace. The AlphaFold database remained unchanged for four years, until September 2025. In our article
in Nature Structural and Molecular Biology in 2025, we documented the discrepancies that have accumulated in
AlphaFoldDB as it aged in comparison to UniProt until its latest release in 2025. During that time, nearly 3%
of the associated sequences underwent revisions in UniProt.
In a range of bioinformatics tasks, protein structure data is paired with sequence annotations from UniProt.
Mapping annotations to outdated structure files can lead to errors in downstream analysis. While this concern
has been addressed for experimental structures through SIFTS, efforts for computationally modelled structures
are lacking. Here, we present 3DSeqCheck, published in JMB in December 2025, which is a lightweight web tool
that enables quick comparison of both computationally modelled and experimental structures to the latest
UniProt entries. 3DSeqCheck provides an interactive visual panel of the alignment and the comparison of the
residue numbering and can be accessed freely at: https://missense3d.bc.ic.ac.uk/3dseqcheck.
C-P.12: Rewriting protein alphabets with language models
Track: Proteins and structural biology
-
Janani Durairaj, University of Basel, Switzerland
-
Gabriel Studer, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
-
Lorenzo Pantolini, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel,
Switzerland, Switzerland
-
Laura Engist, Department of Mathematics and Computer Science, University of Basel, Basel Switzerland,
Switzerland
-
Florian Pommerening, Department of Mathematics and Computer Science, University of Basel, Basel
Switzerland, Switzerland
-
Ieva Pudžiuvelytė, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel,
Switzerland, Switzerland
-
Andrew M. Waterhouse, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel,
Switzerland
-
Stefan Bienert, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
-
Gerardo Tauriello, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel,
Switzerland
-
Martin Steinegger, School of Biological Sciences, Seoul National University, Seoul, South Korea,
Switzerland
-
Torsten Schwede, SIB Swiss Institute of Bioinformatics & Biozentrum, University of Basel, Switzerland
Presentation Overview: Show
Detecting remote homology with speed and sensitivity is crucial for tasks like function annotation and
structure prediction. We introduce a novel approach using contrastive learning to convert protein language
model embeddings into a new 20-letter alphabet, The Embedded Alphabet (TEA), enabling highly efficient
large-scale protein homology searches.
Searching with our alphabet performs on par with and complements structure-based methods without requiring any
structural information, and with the speed of sequence search. This provides a significant advantage for
proteins that lack structural data and enables searches at the scale of available sequence data numbering over
a billion. The inbuilt entropy metric provides a measure of confidence offering a valuable estimate of
prediction reliability for downstream analysis. This work has broad implications, significantly accelerating
and broadening the scope of protein sequence analysis. The ability to detect remote homologs more effectively
enhances functional annotation transfer and drug discovery efforts. We also provide a user-friendly web-server
at https://pickybinders.org/tea enabling conversion of amino acid sequences to TEA sequences and searches on
the UniRef50 and Foldseek AlphaFold Clusters databases.
Ultimately, we bring the exciting advances in protein language model representation learning to the plethora
of sequence bioinformatics algorithms developed over the past century, offering a powerful new tool for
biological discovery.
C-P.13: Predicting changes in protein-protein binding affinity upon mutation with statistical
potentials
Track: Proteins and structural biology
-
Gabriel Cia, Université Libre de Bruxelles, Belgium
- André Ciupitu, Université Libre de Bruxelles (ULB), Belgium
- Marianne Rooman, Université Libre de Bruxelles, Belgium
- Fabrizio Pucci, Université Libre de Bruxelles, Belgium
Presentation Overview: Show
In-silico approaches for predicting changes in protein-protein interaction (PPI) binding affinity upon
mutation (ΔΔG_b) continue to exhibit critical biases and difficulties generalizing beyond their training
datasets. We propose a novel physics-based approach, termed "interface potentials," for the description of PPI
interfaces and the prediction of variant effect. These statistical potentials are derived from datasets of
known PPI structures using the Boltzmann Law and describe the interfaces in terms of distances, torsion
angles, and solvent accessibility of residues. Interface potentials demonstrate great performance on both the
SKEMPI2 experimental dataset and Deep Mutational Scanning datasets without requiring additional machine
learning or model training, highlighting the power of the physics-based framework. Furthermore, interface
potentials can serve as invaluable physics-based features in sophisticated prediction models such as AbMuSiC,
our state-of-the-art physics-based ΔΔG_b prediction method for antibody-antigen interactions.
C-P.14: AI Enhanced Viral Structural Phylogenetics
Track: Proteins and structural biology
-
David Moi, University of Lausanne, Switzerland
-
Christophe Dessimoz, University of Lausanne, Department of Computational Biology, Swiss Institute of
Bioinfromatics, Switzerland
- Dongwook Kim, UNIL, Switzerland
Presentation Overview: Show
Recent efforts in structural phylogenetics have highlighted their ability to resolve deeper evolutionary
relationships than traditional sequence-based methods. These approaches hold promise for elucidating
phylogenetic and taxonomic relationships within the virosphere which tend to escape traditional phylogenetics
due to the pace of viral evolution. However, the paucity of available structures for the incredible observable
diversity of extant viruses and the computational cost of inferring structures has prevented a systematic
structural exploration of the virosphere. We have fine-tuned ESMc to produce structural tokens through
Foldseek's 3Di alphabet using the structures available in the AlphaFold2-based structures available in the Big
Fantastic Virus Database. We have succeeded in producing a viral-focused protein LLM which allows for rapid
conversion of amino acid sequences into 3Di tokens and opens up the voluminous viral sequencing datasets for
structurally informed homology searching and phylogenetics through the use of Foldseek. We have also created a
pipelines and tools for infering cleavage sites, constructing phylogenies of structurally homologous cleaved
products, annotation of functional content and the construction of 'taxonomic' consensus trees centered around
the structure enhanced representation of viral proteomes.
C-P.16: ProInterVal-BioXtal: A Web Server to Distinguish Biological Interfaces from Crystal Contacts Using
Graph Deep Learning
Track: Proteins and structural biology
-
Defne Alnigenis, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450,
Turkey, Turkey
-
Damla Ovek Baydar, Dept. of Computer Eng., Koç Univ., Istanbul, Turkey & NCMBM, Univ. of Oslo, Norway,
Norway
-
Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450, Turkey,
Turkey
-
Attila Gursoy, Department of Computer Engineering, Koç University, Istanbul, 34450, Turkey, Turkey
Presentation Overview: Show
Accurate discrimination between biologically relevant protein–protein interfaces (PPIs) and crystallographic
contacts is essential for reliable interpretation of macromolecular assemblies and their cellular functions.
While X-ray crystallography remains a primary method for determining protein complex structures,
crystallographic interfaces may be formed as a byproduct of the crystal packing. Therefore, robust
computational approaches for interface annotation are needed.
This study introduces ProInterVal-BioXtal, a web server that predicts class scores for PPIs, differentiating
between biologically relevant and crystallographic interfaces. The method leverages protein representation
learning and graph-based deep learning to capture structural and physicochemical features. Interface graphs
derived from input complexes are processed by our novel graph-based contrastive model to learn interface
representations which are used by a graph neural network for classification.
The model is trained and validated on the MANY benchmark dataset comprising 5739 dimers with a balanced
distribution of biological and crystal interfaces and evaluated on the DC benchmark dataset with curated
interfaces of similar interface areas.
On the DC test set, ProInterVal-BioXtal achieves 88% accuracy, 88% precision, and 85% F1 score, outperforming
state-of-the-art methods including DeepRank-GNN, PRODIGY-CRYSTAL, EPPIC 3, PISA, and QSAlign. On an
independent benchmark, the method achieves 83% accuracy and 0.91 AUC, surpassing DeepRank-GNN (AUC = 0.85).
Through a user-friendly web interface, ProInterVal-BioXtal enables rapid predictions from PDB files or IDs,
returning probabilistic classification scores. Our tool addresses a fundamental challenge in structural
biology by utilizing graph-based protein representation which can detect complex interactions and
dependencies. ProInterVal-BioXtal server is freely available at https://3dpath.ku.edu.tr/prointerval-bioxtal/
C-P.17: Improving protein structure prediction with structure-aware multiple sequence alignments
Track: Proteins and structural biology
-
Diana Rapota, Biozentrum, University of Basel, Basel Switzerland; SIB Swiss Institute of
Bioinformatics, Basel, Switzerland, Switzerland
-
Celia Ulrich, Biozentrum, University of Basel, Spitalstrasse, 4056, Basel Switzerland, Switzerland
-
Lorenzo Pantolini, Biozentrum, University of Basel, Basel Switzerland; SIB Swiss Institute of
Bioinformatics, Basel, Switzerland, Switzerland
-
Janani Durairaj, Biozentrum, University of Basel, Basel Switzerland; SIB Swiss Institute of
Bioinformatics, Basel, Switzerland, Switzerland
Presentation Overview: Show
Leveraging evolutionary information is central to modern protein structure prediction. As protein structure is
more conserved than sequence throughout evolution, searching for distant protein homologs using structural
information is a highly appealing strategy. However, until recently, performing high-throughput searches for
distant structural homologs with low sequence identity was not feasible without access to large databases of
structures. The Embedded Alphabet (TEA), a novel 20-letter 1D structural alphabet, addresses this challenge by
enabling fast and sensitive detection of distant homologs without requiring prior structural information. In
this work, we investigate whether structure-aware multiple sequence alignments (MSAs) generated with TEA can
improve protein structure prediction by enhancing the evolutionary signal captured in the alignments.
Structure-aware MSAs were constructed using Search with TEA against Many (STEAM), a tool built on the Foldseek
framework and adapted for the TEA alphabet. We evaluated these MSAs on two datasets: the CASP13 dataset,
containing query proteins with shallow MSAs, and viral proteins from BFVD, previously characterized by
poor-quality MSAs using the default ColabFold DB. To evaluate the impact on structure prediction, we ran
AlphaFold2 (AF2) in custom MSA mode using both STEAM-generated structure-aware MSAs and MMseqs2-generated
MSAs. We present an evaluation of STEAM-generated MSAs and their effect on AF2 prediction quality across both
datasets, with a particular focus on challenging low-homology and viral targets, where sequence-based searches
alone may be insufficient.
C-P.18: Computational prediction of ligand-receptor pairs using transformer-based language models,
interaction networks, and structural modelling
Track: Proteins and structural biology
-
Iuliia Trifonova, Department of Biosciences and Medical Biology, Center for Tumor Biology and
Immunology, Paris Lodron University Salzburg, Austria
-
Markus Wiederstein, Department of Biosciences and Medical Biology, Paris Lodron University Salzburg,
Austria
-
Nikolaus Fortelny, Department of Biosciences and Medical Biology, Center for Tumor Biology and Immunology,
Paris Lodron University Salzburg, Austria
Presentation Overview: Show
Ligand-receptor (LR) interactions regulate cellular communication and shape processes in development, immune
responses, and disease. However, existing curated databases remain incomplete, and many biologically relevant
interactions are described only in primary literature, limiting systematic discovery. We developed a scalable
computational framework that integrates literature mining with transformer-based classification and structural
validation to identify previously unannotated LR interactions among known protein-protein interactions.
We first harmonised 18 curated LR databases into one reference database of 18,283 unique protein-protein LR
pairs and used the high-confidence ConnectomeDB2025 subset (3,550 pairs) as true positives. We evaluated
multiple strategies to derive true negatives, finally using topologically distant proteins based on a
consensus protein-protein interaction network from the PCNet database. We then extracted sentences mentioning
true positive and true negative pairs from Europe PMC and PubMed.
To separate contextual co-mention from mechanistic interaction, we trained a logistic-regression classifier on
sentence embeddings from a frozen PubMedBERT encoder. This classifier reached a pair-level ROC AUC of 0.86. We
next evaluated high-scoring candidates structurally using AlphaFold3. The strongest previously unannotated
candidate, COL2A1-TGM2 (ipTM = 0.80), showed a plausible interface, though experimental work is needed to
confirm whether these are true extracellular LR pairs or other PPIs.
Together, these results suggest that literature-guided classification combined with structural filtering can
potentially expand LR knowledge beyond curated resources, although distinguishing true LR pairs from other
PPIs and improving semantic negation handling remain important next steps.
C-P.19: Generating tailored high-quality datasets for benchmarking structure-based computational drug design
tools: a docking assessment
Track: Proteins and structural biology
-
Marine Mathieu, SIB Swiss Institute of Bioinformatics, Switzerland
- Ute F. Röhrig, SIB Swiss Institute of Bioinformatics, Switzerland
- Vincent Zoete, UNIL University of Lausanne, Switzerland
Presentation Overview: Show
Structural bioinformatics is critical for drug discovery, as it provides methods and tools to predict,
analyze, and validate 3D data of macromolecules. The ELIXIR 3D-BioInfo community is organized around five main
activities, one of which aims to create the tools necessary for developing large-scale, high-quality, and
sustainable datasets of ligand-protein complexes. These datasets will be used to assess and benchmark
structure-based computer-aided drug design (SB-CADD) algorithms, such as docking, virtual screening and
binding site detection tools.
We developed three interconnected Nextflow pipelines capable of constructing datasets from complex, protein,
or ligand PDB identifiers. Data sources include multiple APIs and dedicated software developed by ELIXIR
participating groups for data retrieval and processing.
The pipelines generate multifaceted datasets with comprehensive annotations. Complex analysis involves
retrieving structural details including proteins, bound ligands and their interactions, as well as quality
metrics such as resolution, missing atoms, crystal contacts, and electronic density support from the Protein
Data Bank. Coordinate files are retrieved for complexes and ligands, as well as binding affinity, refined
structures, tautomers, and protonation states. The protein characterization pipeline extracts UniProt KB
descriptions, associated 3D structures, bound ligands, binding sites, active and inactive molecules, and
protein flexibility. For ligands, molecular properties, 3D conformers, partial charges, and associated complex
structures are provided.
We present benchmark sets generated from complex, protein, or ligand identifiers, along with an initial
docking assessment. Docking with AutoDock Vina, AutoDock4, and SMINA provides a first evaluation of SB-CADD
tool performance on the curated, high-quality data.
C-P.20: Integrated CSF-serum proteomic profiling in spinal cord injury
Track: Proteins and structural biology
-
Sasimonthakan Tanarsuwongkul, Spinal Cord Injury Center, Heidelberg University Hospital,
Germany
- Norbert Weidner, Spinal Cord Injury Center, Heidelberg University Hospital, Germany
- Giada Sandrini, German Cancer Research Center (DKFZ), Germany
- Katalin Barkovits, Medical Faculty, Ruhr University Bochum, Germany
- Andreas Hug, Spinal Cord Injury Center, Heidelberg University Hospital, Germany
- Junyan Lu, Institute for Computational Biomedicine, Heidelberg University, Germany
Presentation Overview: Show
Traumatic spinal cord injury (SCI) leads to lifelong impairment and disability due to the lack of spontaneous
nerve regeneration. Neuroregenerative therapies are being investigated in search for effective interventions
to improve clinical outcomes for SCI patients. The current patient stratification and neurological outcome
assessment in SCI clinical trials involves the use of the American Spinal Injury Association Impairment Scale
(AIS). However, AIS grades are unable to distinguish different neurological impairments. Therefore, it often
fails to reflect clinically important differences in severity and recovery, especially in people with cervical
SCI. To address this, proteome of cerebrospinal fluid and blood serum from NISCI trial (ClinicalTrials.gov,
NCT03935321), which tests the effectiveness of a nogo-A antibody in acute SCI, have been evaluated. This
longitudinal integrated CSF-serum proteomic profiling identified underlying pathways of SCI severity and
recovery and molecular predictors of neuronal improvement, which could be used to stratify patients with SCI.
C-P.21: Exploring the mobile genetic element continuum using PhageProfiler
Track: Proteins and structural biology
-
Jiawei Wang, University of Bath; EMBL-EBI, United Kingdom
-
Jinzheng Ren, University of Bath; EMBL-EBI; Australian National University, United Kingdom
-
Licheng Zong, University of Bath; EMBL-EBI; The Chinese University of Hong Kong, Hong Kong
- Robert Finn, EMBL-EBI, United Kingdom
Presentation Overview: Show
Mobile genetic elements (MGEs), including plasmids and bacteriophages, are often treated as discrete
categories, despite growing evidence for a modular and evolutionary continuum shaped by gene exchange,
recombination, and hybrid elements such as phage–plasmids. Here, we present PhageProfiler, an
embedding-based framework for profiling MGEs across sequence space. By learning task-specific representations
at different biological levels, PhageProfiler captures continuous organization at the replicon level, while
preserving discrete structure at taxonomic and functional levels. These representations provide complementary
views of MGE organization, linking genome-level organization to both evolutionary relationships and underlying
gene content. Using virulent phages as a case study, we demonstrate that these perspectives are mutually
consistent and biologically informative.
C-P.22: Exploring protein language model approaches for thermal stability prediction
Track: Proteins and structural biology
-
Gregory Coolen, Université Libre de Bruxelles, 3BIO-BioInfo, Belgium
- Marianne Rooman, Université Libre de Bruxelles, 3BIO-BioInfo, Belgium
- Fabrizio Pucci, Université Libre de Bruxelles, 3BIO-BioInfo, Belgium
Presentation Overview: Show
Accurately determining a protein's melting temperature is essential for many biotechnological applications
such as enzyme design or understanding organismal adaptation to environmental conditions. While experimental
methods are accurate, they are also costly and time-consuming, which justifies the development of
computational alternatives. Recently, protein language models (pLM) applied to protein sequences have shown
impressive performance across various tasks, including protein structure prediction. These methods have also
been explored for thermostability prediction and reported to achieve good performance. In this work, we review
existing pLM-based thermostability predictors and benchmark their robustness on new datasets of thermal
stability that we collected from the literature. We found that their actual performance remains modest. This
can be attributed to the limited availability of annotated data and the fact that melting temperature is
sometimes sensitive to minor sequence variations. Finally, to better understand the biophysical principles
underlying thermal stability and what pLM-based models capture, we analyze their internal representations
using standard interpretability techniques, such as ablation.
C-P.23: A Unified Framework for TCR-pMHC Structural Model Assessment
Track: Proteins and structural biology
-
Miguel Romero-Durana, Barcelona Supercomputing Center, Spain
- Alex Ascunce-ParÃs, Barcelona Supercomputing Center, Spain
- Alfonso Valencia, Barcelona Supercomputing Center, Spain
- Roc Farriol-Duran, Barcelona Supercomputing Center, Spain
- Victor Guallar, Barcelona Supercomputing Center, Spain
Presentation Overview: Show
Structural characterization of T cell receptor (TCR) and peptide-MHC complex (pMHC) interactions is
fundamental to understanding adaptive immunity. However, the scarcity of experimentally resolved TCR-pMHC
structures constrains progress in this field. Computational modelling tools can help bridge this gap, but
their predictions are difficult to evaluate reliably without experimental references.
We present a scalable and interpretable framework for reference-free quality assessment of TCR-pMHC class I
structural models.
We benchmarked seven state-of-the-art modelling methods on 265 experimentally determined PDB complexes,
covering pre- and post-training-cutoff datasets. AlphaFold3 demonstrated superior generalization, achieving
over 90% acceptable-quality predictions for structures absent from its training set.
Then, we integrated multiple confidence metrics into a random forest classifier trained on 1,325 AlphaFold3
models, with quality labels derived from their 265 experimental counterparts. The classifier stratifies models
into four tiers (low, acceptable, medium, high), outperforming single-metric approaches. SHAP analysis
identified pDockQ2 and iPDE as the strongest predictors of structural quality.
Finally, we applied this framework to a VDJdb/TCRvdb-derived subset of validating and non-validating TCR-pMHC
pairs, and to IMMREP23, comprising validated positives and synthetic negatives. True binders were consistently
enriched in higher quality tiers, confirming that structural quality is a reliable indicator of functional
interaction.
Overall, this work contributes a large-scale dataset of evaluated TCR-pMHC structural models (21,775 models
from 4,355 complexes; 265 PDB, 609 TCRvdb, 3,484 IMMREP23), and a robust framework for their quality
assessment. This will facilitate the curation of structural TCR-pMHC datasets, and downstream applications in
immunology and TCR-based therapeutics.
C-P.24: From Proteome Profiles to Clinical Diagnosis in Children with Bone and Joint Infections
Track: Proteins and structural biology
-
Réka Gonda, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital, Copenhagen,
Denmark, Denmark
-
Allan Bybeck Nielsen, Department of Paediatrics and Adolescent Medicine, Copenhagen University Hospital
– Rigshospitalet, Copenhagen, Denmark, Denmark
-
Anna Benedetti, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital; Copenhagen
Center for Translational Research, Denmark
-
Anna Melidi, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital, Copenhagen,
Denmark, Denmark
-
Nicolai Wewer Albrechtsen, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital;
Department of Clinical Medicine, UCPH, Denmark
-
Ulrikka Nygaard, Department of Paediatrics and Adolescent Medicine, Rigshospitalet; Department of Clinical
Medicine, UCPH, Denmark
-
Annelaura Bach Nielsen, Department of Clinical Biochemistry, Bispebjerg and Frederiksberg Hospital;
Copenhagen Center for Translational Research, Denmark
Presentation Overview: Show
Background: Bone and joint infections (BJI) in children represent a diagnostic challenge due to their clinical
overlap with transient synovitis (TS). Current diagnostic approaches rely on clinical scoring systems,
inflammatory markers and imaging, yet none achieve sufficient sensitivity to distinguish the two conditions
without invasive procedures. Plasma proteomics offers a non-invasive approach to address this diagnostic gap,
as mass spectrometry-based proteomics captures the systemic host response at a molecular resolution beyond
what is achievable with single biomarker measurements.
Methods and Results: We performed untargeted liquid chromatography mass spectrometry of plasma samples from
262 pediatric patients across five inflammatory disease groups, including BJI (n=142) and TS (n=83). Following
preprocessing, batch correction and imputation, differential abundance analysis identified 118 significantly
differentially abundant proteins between BJI and TS (FDR-adjusted p<0.05). The most strongly upregulated
proteins in BJI included the acute phase reactants SAA1, SAA2 and CRP, while apolipoproteins were among the
most downregulated. Reactome pathway enrichment revealed significant enrichment of Toll-like receptor
signalling, innate immune activation and neutrophil degranulation in BJI, alongside suppression of metabolic
and lipid transport pathways.
Perspective: Ongoing analyses aim to stratify BJI by microbiological confirmation status to identify
pathogen-specific proteomic signatures. We are developing a plasma proteomics-based machine learning model to
distinguish BJI from TS, with the long-term perspective of providing a clinical decision support tool guiding
clinicians between surgical intervention and conservative management to reduce unnecessary invasive procedures
in children with suspected BJI.
C-P.25: PUFFIN: Protein Unit Discovery with Functional Supervision
Track: Proteins and structural biology
-
Gökçe Uludoğan, Bogazici University, Turkey
- Buse Giledereli, Bogazici University, Turkey
- Elif Ozkirimli, Roche AG, Switzerland
- Arzucan Ozgur, Bogazici University, Turkey
Presentation Overview: Show
Motivation:
Proteins carry out biological functions through the coordinated action of groups of residues organized into
structural arrangements. These arrangements, which we refer to as protein units, exist at an intermediate
scale, being larger than individual residues yet smaller than entire proteins. A deeper understanding of
protein function can be achieved by identifying these units and their associations with function. However,
existing approaches either focus on residue-level signals, rely on curated annotations, or segment protein
structures without incorporating functional information, thereby limiting interpretable analysis of
structure-function relationships.
Results:
We introduce PUFFIN, a data-driven framework for discovering protein units by jointly learning structural
partitioning and functional supervision. PUFFIN represents proteins as residue-level structure graphs and
applies a graph neural network with a structure-aware pooling mechanism that partitions each protein into
multi-residue units, with functional supervision that shapes the partition.
We show that the learned units are structurally coherent, exhibit organized associations with molecular
function, and show meaningful correspondence with curated InterPro annotations. Together, these results
demonstrate that PUFFIN provides an interpretable framework for analyzing structure-function relationships
using learned protein units and their statistical function associations.
Availability and Implementation: We made our source code available at
https:/github.com/boun-tabi-lifelu/puffin
C-P.26: PPInterface v2: A Dataset and Web Resource for 3D Protein–Protein Interface Analysis and
Visualization
Track: Proteins and structural biology
-
Defne Alnigenis, Department of Chemical and Biological Engineering, Koç University, Istanbul, Turkey,
Turkey
-
Moaaz Ur Rehman Azhar Khokhar, Department of Computer Engineering, Koç University, Istanbul, 34450,
Turkey, Turkey
-
Zeynep Abali, Computational Science and Engineering Graduate Program, Koç University, Istanbul 34450,
Turkey, Turkey
-
Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Istanbul, 34450,
Turkey, Turkey
-
Attila Gursoy, Department of Computer Engineering, Koç University, Istanbul, 34450, Turkey, Turkey
Presentation Overview: Show
Proteins regulate cellular processes through sophisticated interactions formed at protein-protein interfaces
(PPIs). Structural characterization of these interfaces is essential for understanding biological functions
and molecular mechanisms of proteins. Given the pivotal roles of PPIs, the construction of a well-curated
dataset of interface structures holds significant promise for large-scale analysis.
This study presents an updated version of the PPInterface dataset, a comprehensive resource of structural
protein–protein interaction data gathered from the Protein Data Bank (PDB). Building on our previously
published framework (2024), the updated dataset integrates newly available structures and interaction
annotations. The previous release of PPInterface contained 815,082 interfaces extracted from over 215,000
three-dimensional protein structures. The database now contains 252,476 protein complexes and more than
940,000 interfaces, indicating a substantial increase over the previous release. To our knowledge, PPInterface
provides one of the most comprehensive datasets of PPIs, allowing users to search, explore, and download data
for a wide range of analyses.
In addition to expanded data coverage, the user-friendly web resource of PPInterface enables detailed
interpretation of PPIs, including structural, physicochemical, and evolutionary features through efficient
exploration of with advanced search, filtering, and visualization. PPInterface, as a valuable resource for the
analysis of protein interactions, may be used as a benchmark dataset for developing computational methods,
including machine learning-based approaches. This updated database provides an up-to-date and scalable
resource to the community for which can be used to support studies in structural bioinformatics, systems
biology and PPIs in general.
C-P.27: RECUR: identifying recurrent amino acid substitutions from multiple sequence alignments
Track: Proteins and structural biology
- Elizabeth Robbins, Ellison Institute of Technology Oxford, United Kingdom
-
Yi Liu, University of Oxford, United Kingdom
- Steven Kelly, University of Oxford, United Kingdom
Presentation Overview: Show
Identifying recurrent changes in biological sequences is important to multiple aspects of biological
research—from understanding the molecular basis of convergent phenotypes, to pinpointing the causative
sequence changes that give rise to antibiotic resistance and disease. Here, we present RECUR, a method for
identifying recurrent amino acid substitutions from multiple sequence alignments that is fast, easy to use,
and scalable to thousands of sequences. We demonstrate that RECUR's recurrence detection achieves 100%
accuracy on simulated data with known evolutionary histories. We further show that RECUR is robust to
realistic levels of tree inference error. Finally, we apply RECUR to a large set of surface glycoprotein (S)
protein sequences from SARS-CoV-2. This analysis identified widespread recurrent evolution throughout the
protein with significant enrichment in the exposed receptor-binding S1 subunit and at the interface with the
human angiotensin-converting enzyme 2 (hACE2). In contrast, recurrent substitutions were depleted at the
trimeric interface of the S protein. In silico modelling showed that recurrent substitutions had no
directional effect on stability at either interface, but effects at the hACE2 interface were significantly
more variable. Multiple substitutions with large destabilizing effects on hACE2 binding have been linked to
immune escape, while others represented reversions back to the reference sequence, suggesting that recurrent
evolution at this interface reflects opposing selective pressures balancing receptor binding with immune
evasion. A standalone implementation of the algorithm is available under the GPLv3 license at
https://github.com/OrthoFinder/RECUR.
C-P.28: Machine Learning Classification and Analysis of Gain-of-Function, Loss-of-Function, and Neutral
Mutations in Cancer
Track: Proteins and structural biology
-
Biniam Haile, University Of Sussex, United Kingdom
- Adnan Cinar, University of Sussex, United Kingdom
- Frances Pearl, University of Sussex, United Kingdom
Presentation Overview: Show
Missense mutations in cancer can be broadly categorised as gain-of-function (GOF), which enhance or alter
protein activity, and loss-of-function (LOF), which impair or abolish protein function. Mutations that
contribute to tumour growth by providing a selective growth advantage are considered driver mutations, while
others are classified as passenger mutations, having no direct selective advantage for cell growth.
Accurately distinguishing GOF, LOF, and Neutral/passenger mutations is essential for understanding tumour
biology, identifying therapeutic targets, and developing precision oncology strategies. In this study, we
conduct a comparative analysis of these mutation classes by examining their structural and functional impact
profiles, and their roles in cancer progression. We integrate a range of computational tools and curated
mutation dataset (GOF, LOF, and Neutral) for structural modelling (SAAP, and NetsurfP), functional consequence
prediction, and statistical evaluation.
Our findings were able explore distinct molecular patterns that might be very important to differentiate GOF,
LOF, and Neutral mutations, providing insights into their molecular mechanisms in cancer. This comprehensive
characterisation underscores the importance of combining structural impact analysis with functional
annotations to improve mutation interpretation or classification. That was ultimately used in developing a
machine learning model to classify missense mutation, which achieved ROC AUC of 0.92 classifying GOF and LOF
mutations, and ROC AUC 0.943 on external validation mutation dataset.
C-P.29: AI-driven Intrabody Design: Representation Learning and Multi-objective Optimization of Nanobody
Frameworks
Track: Proteins and structural biology
-
Savelii Komlev, Argentys LLC, France
- Ilya Mazo, Argentys LLC, United States
- Aleksei Artemiev, Argentys LLC, Spain
Presentation Overview: Show
Intrabodies are intracellular antibody fragments that enable targeting of proteins inaccessible to
conventional biologics, but their development is limited by poor stability in the reducing cytoplasmic
environment. As a result, intrabody engineering remains largely empirical and lacks scalable computational
approaches.
We present an AI-driven pipeline for intrabody design that integrates protein language model–based
representation learning, supervised stability prediction, and generative nanobody optimization. Sequence
embeddings derived from pretrained models (ESM2, ProtT5, IgT5) are used to train machine learning classifiers
for predicting intracellular stability of nanobodies. Models trained only on framework regions of nanobodies
outperform full-sequence models, achieving ROC-AUC > 0.80, indicating that stability-related signals are
primarily encoded in framework residues.
To explore sequence space, we implement a hybrid generative pipeline combining consensus-based framework
mutagenesis (consensus nanobody sequence based on a set experimentally validated stable intrabodies), language
model–guided sequence infilling, and stochastic diversification. More than 5,000 candidate nanobody variants
are generated from a small set of parental sequences.
Generated variants are evaluated using independent structure-based metrics not used during model training,
including AlphaFold-Multimer interface confidence (ipTM) and FoldX interface energies. Top candidates show a
strong shift toward higher predicted interaction confidence (ipTM ~0.55–0.75 vs. ~0.1–0.2 in parental
sequences) and more favorable binding energetics (median FoldX ΔG shifts from ~−5 to ~−25 kcal/mol, with
reduced variance across top-ranked variants).
Selected candidates are currently undergoing experimental validation. These results demonstrate that combining
representation learning, generative design, and structure-based screening enables scalable computational
intrabody engineering.
C-P.30: Identification and characterization of Polyurethane-Degrading Enzymes from MGnify Metagenomes
Track: Proteins and structural biology
-
Joel Roca Martinez, UCL, United Kingdom
- Christine Orengo, UCL, United Kingdom
Presentation Overview: Show
The discovery of enzymes capable of degrading synthetic polymers remains a major challenge due to the vast
size and functional diversity of metagenomic sequence space. Here, we present a rational, structure- and
sequence-informed pipeline for enzyme discovery, applied to the identification of novel polyurethane-degrading
enzymes (PURases) from large-scale metagenomic data. Starting from 2.4 billion protein sequences in MGnify, we
performed homology-based filtering against three known PURases followed by conservation analysis of the
catalytic triad, yielding ~8,000 high-confidence candidates. To prioritize functionally diverse yet tractable
subsets, we clustered sequences into functional families using an embedding-based classifier (eMMA-FunFamer)
and identified function-determining positions (FDPs) conserved within families but variable across them.
Physicochemical variation at these FDPs was used to construct a sequence similarity network, revealing
distinct functional clusters that guided representative selection.
Candidate prioritization further integrated metagenomic biome metadata to enrich for thermostable enzymes and
structural features such as loop architecture near the active site, predicted using AlphaFold2. From this
rational selection pipeline, 20 diverse enzymes were selected for experimental characterization. Biochemical
assays identified 9 enzymes with activity against carbamate substrates, including 2 that also exhibited
activity against polyurethane polymers. These results demonstrate that combining functional family analysis,
residue tunability metrics, and structural diversity enables efficient navigation of metagenomic sequence
space and substantially improves hit rates in enzyme discovery. This pipeline is broadly applicable to other
challenging catalytic functions where experimental screening capacity is limited.
C-P.31: EnzymeSifter: a pipeline for identifying and filtering industrial enzymes from metagenomes across
multiple predicted biochemical properties
Track: Proteins and structural biology
-
Omar Darawsheh, Northumbria University, United Kingdom
- Matthew Bashton, Northumbria University, United Kingdom
Presentation Overview: Show
Identifying candidate enzymes from environmental samples for industrial applications requires evaluating
multiple biochemical properties simultaneously, including solubility, thermal stability, and pH. While
individual predictors for these properties exist, integrating them into a coherent workflow is a manual and
error-prone process. Here we present EnzymeSifter, a two-stage Snakemake pipeline that automates enzyme
discovery, multi-property characterisation, and generates a ranked shortlist.
In Stage 1, input amino-acid sequences are optionally filtered by user-defined catalytic motifs, Pfam domain
annotations, and/or predicted EC number, and optionally clustered at a user-specified sequence identity
threshold using MMseqs2, to produce a non-redundant set for structure prediction. In Stage 2, predicted PDB
structures are screened with EnzyMM for enzymatic activity, and all confirmed hits are characterised across 5
properties: solubility and usability (NetSolP), optimal pH (pHoptNN), and optimal temperature and melting
temperature (Seq2Topt). Predictions are merged, and users may specify threshold or interval filters for any
property. A MUSCLE multiple sequence alignment and neighbour-joining tree are constructed. The tree can be
partitioned into a user-defined number of clades to select a representative enzyme within each clade.
Representatives are enzymes that achieved the best composite score according to the properties defined in the
user's input.
EnzymeSifter provides a standardised framework that reduces the effort required to go from large sequence
datasets to a shortlist of promising enzymes for experimental validation.
C-P.32: High Diversity Gene Libraries Facilitate Machine Learning Guided Exploration of Fluorescent Protein
Sequence Space
Track: Proteins and structural biology
-
Anissa Benabbas, University of Oregon, United States
Presentation Overview: Show
Protein language models (PLMs) are increasingly used for protein design but remain limited by the diversity
and structure of available training data. Models trained on natural sequences often operate in an
extrapolative regime, reducing reliability when exploring sparsely sampled regions of sequence space. Here, we
test whether experimentally expanding sequence diversity can shift this regime toward interpolation. Using
large-scale gene synthesis and DNA shuffling, we generate libraries spanning broad regions of fluorescent
protein sequence space and identify thousands of functional blue fluorescent variants through high-throughput
screening. Fine-tuning
ProtGPT2 on this dataset enables the generation of diverse fluorescent proteins, including variants that
extend beyond regions occupied by known natural sequences while retaining function. Together, these results
support a strategy in which experimentally expanded diversity improves the ability of machine learning models
to explore and design functional proteins. This approach provides a general framework for improving machine
learning–guided protein design by experimentally expanding functional sequence diversity.
C-P.33: Deep generative modeling captures maturation-dependent pairing patterns in human antibodies
Track: Proteins and structural biology
-
Lea Brönnimann, University of Bern, Switzerland
- Thomas Lemmin, University of Bern, Switzerland
- Chiara Rodella, University of Bern, Switzerland
Presentation Overview: Show
Understanding antibody heavy-light chain pairing is critical for decoding immune repertoire architecture and
designing therapeutic antibodies, yet most sequence databases lack paired chain information. To address this
gap, we developed a two-stage deep learning framework. Transformer-based language models were first
pre-trained on large corpora of unpaired heavy- and light-chain sequences, then integrated into a
sequence-to-sequence model to generate light chains from heavy chain input. Although native light chain
recovery was moderate, generated sequences exhibited high germline identity, improved structural quality, and
broader framework and complementarity-determining region coverage. Heavy chains from memory B cells generated
light chains with more restricted V gene usage, reflecting maturation-dependent selection. Generated kappa
light chains exhibited a trimodal similarity distribution, indicating distinct functional pairing modes from
promiscuous to highly specific. Our approach demonstrates that sequence-to-sequence modeling can uncover
inter-chain dependencies and generate plausible antibody pairs, providing a foundation for computational
repertoire analysis and therapeutic design.
C-P.34: Enhancing protein-ligand binding site predictions: Integrating protein language models with geometric
smoothing and clustering
Track: Proteins and structural biology
-
Vít Škrhák, Charles University, Czechia
-
Lukáš Polák, Department of Software Engineering, Faculty of Mathematics and Physics, Charles University,
Prague, Czech Republic, Czechia
-
Marian Novotný, Department of Cell Biology, Faculty of Science, Charles University, Prague, Czech
Republic, Czechia
-
David Hoksza, Department of Software Engineering, Faculty of Mathematics and Physics, Charles University,
Prague, Czech Republic, Czechia
Presentation Overview: Show
The identification of ligand binding sites (LBS) is a cornerstone of computer-aided drug design. While Protein
Language Models (pLMs) have recently demonstrated high performance in classifying binding residues, their
primary limitation remains their ""residue-centric"" focus, which frequently produces spatially disjointed or
fragmented predictions when mapped onto 3D structures.
We introduce Seq2Pocket, a framework designed to bridge the gap between sequence-based classification and
structural pocket continuity. Our approach utilizes a fine-tuned pLM augmented by an embedding-aware smoothing
classifier and a clustering algorithm. To mitigate prediction disjointedness, we propose the Pocket
Fragmentation Index (PFI), a metric used to optimize the mapping between predicted residues and binding
cavities.
Experimental results on the scPDB dataset show that Seq2Pocket achieves state-of-the-art performance, with a
significant 11% improvement in DCC recall over current methods. Further validation on the LIGYSIS and
CryptoBench benchmarks confirms that our framework not only maintains high performance but also provides the
structural coherence necessary for practical downstream drug discovery workflows.
C-P.35: PMGen: From Peptide-MHC Structure Prediction to Peptide Generation
Track: Proteins and structural biology
-
Amir H. Asgary, Quantitative and Computational Biology Group, Max Planck Institute for
Multidisciplinary Sciences, Germany
-
Johannes Soeding, Quantitative and Computational Biology Group, Max Planck Institute for Multidisciplinary
Sciences, Germany
Presentation Overview: Show
Accurate structural modeling of peptide-MHC (pMHC) complexes is essential for structure-driven immunotherapy
design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths,
insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope
and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical
insights from pMHC structures largely unexploited.
We introduce PMGen (Peptide-MHC Generator), an integrated framework for structure prediction and
structure-guided design of variable-length peptides across MHC class I and II. PMGen enforces anchor
constraints within AlphaFold2 through two complementary strategies,Initial Guess and Template Engineering,
achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen
outperforms all existing methods, yielding median peptide-core C-alpha RMSDs of 0.54 A for MHC-I and 0.33 A
for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT
scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type
pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show
that ProteinMPNN sampling on PMGen-predicted backbones yields higher-affinity peptides while preserving the
parental 3D conformation. Using PMGen to generate 10,216 high-confidence pMHC structures as training data, we
further improve ProteinMPNN's peptide sequence recovery from 0.19 to 0.40, highlighting the value of accurate
predicted structures for downstream machine learning.
PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at
https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.
C-P.36: Mapping targetable sites on the human surfaceome for the design of novel binders
Track: Proteins and structural biology
-
Hamed Khakzad, Inria, France
Presentation Overview: Show
The human cell surfaceome, integral to cell communication and disease mechanisms, presents a prime target for
therapeutic intervention. De novo protein binder design against these cell surface proteins offers a promising
yet underexplored strategy for drug development. However, the vast search space and limited data on natural or
competitive binders have historically limited experimental success. In this study, we systematically analyzed
the entire human surfaceome, identifying approximately 4,500 targetable sites and introducing potential
binding seeds for initiating protein design applications. To validate these seeds, we implemented two
experimental approaches (protein scaffolding and peptide cyclization) on three representative targets (FGFR2,
IFNAR2, and HER3). Our results revealed a high success rate, showing that seeds provide valuable starting
points for binder design against our identified targetable sites, as well as the need for constant
improvements of computational protein design pipelines utilizing machine learning and physics-based methods.
Additionally, we present SURFACE-Bind, an interactive database offering open access to all generated data. The
high-throughput computational design methods and target-specific binder seeds established here pave the way
for a new generation of targeted therapeutics for the human surfaceome.
C-P.37: Mitigating Functional Classification Hallucination in Protein Language Models Through Target-Decoy
Training
Track: Proteins and structural biology
-
Ho-Jin Gwak, Hankuk University of Foreign Studies, South Korea
- Xiaofang Jiang, National Institutes of Health, United States
- Ikbeom Jang, Hankuk University of Foreign Studies, South Korea
- LeAnn Lindsey, Lawrence Berkeley National Labs, United States
Presentation Overview: Show
Protein language model (pLM)-based classifiers are increasingly used for functional annotation, but their
behavior on biologically meaningless sequences remains poorly characterized. Here, we investigate this problem
in phage protein function prediction using PHROGs functional categories and three representative pLMs. We show
that classifiers trained on pLM embeddings can assign high-confidence functional labels to shuffled or
reversed decoy sequences, revealing a hallucination-like false positive problem that cannot be fully resolved
by simple confidence thresholding. To address this limitation, we introduce a target-decoy training strategy
in which shuffled decoy proteins are incorporated as explicit negative examples during classifier training.
This approach enables the resulting our model to reject biologically uninformative inputs, reducing decoy
false positive rates to levels comparable to alignment-based methods such as Pharokka and Phold, while largely
preserving classification performance on genuine PHROG-annotated proteins. Notably, a model trained only with
shuffled decoys also rejects reversed decoys, suggesting that target-decoy training generalizes beyond a
specific synthetic artifact. When applied to large-scale phage protein datasets, PPAM-TDM increases annotation
coverage in INPHARED and PhageScope relative to alignment-based baselines, while maintaining controlled
specificity. These results demonstrate that target-decoy training is a simple and broadly applicable strategy
for improving the reliability of pLM-based protein function classifiers in open-world annotation settings.
C-P.38: Assessing the stability and oligomerisation of a β-hairpin through gas-phase molecular dynamic
simulation
Track: Proteins and structural biology
-
Stijn De Schepper, Uppsala University, Sweden
- Erik Marklund, Uppsala University, Sweden
Presentation Overview: Show
Protein self-assembly into supramolecular clusters is involved in a wide range of processes like the amyloid
plaque formation, but still not fully understood. Governed by noncovalent interactions, clusters can
experience major structural rearrangements upon transitioning into gas phase. As a model of these cluster
forming molecules, we study the antimicrobial peptide Protegrin-1 (PG1), which adopts a well-defined
beta-hairpin fold, stabilised by two disulfide bridges. Ion mobility mass spectrometry (IM-MS) studies have
shown it assembles into mono-, di-, tri- and tetramers, and that reducing the disulfide bridges results in
depletion of the oligomeric states and loss of beta-sheet content. Nevertheless, an accurate atomistic
understanding is still missing.
Recent advances in the field of gas-phase molecular dynamics simulations, considering the changed
electrostatic forces in electrospray ionisation due to the loss of the dielectric medium enable us to dive
deeper into the assembly of PG1. The simulations highlight that the reduction of PG1 leads to a lower
gas-phase stability but does not seem to lead to a significant increase in collision cross section (CCS).
Surprisingly, an increase in β-sheet content is observed however structurally different from the original
β-hairpin motif. Additionally, different protonation states in gas phase tremendously influence the
simulations and thus complicate the comparison to experimental data from IM-MS and gas-phase
IR-spectroscopy.
This work will inform a better understanding of molecular scale changes influencing oligomerization and aid
understanding amyloid plaques involved in neurodegenerative diseases.
C-P.39: Breaking the Data Bottleneck: Leveraging Transfer Learning for data-scarce Post-Translational
Modifications
Track: Proteins and structural biology
-
Yannick Hartmaring, Hasso Plattner Institute for Digital Engineering, Digital Engineering Faculty,
University of Potsdam, Germany
-
Shengbo Wang, European Molecular Biology Laboratory - European Bioinformatics Institute (EMBL-EBI), United
Kingdom
-
Juan Antonio Vizcaino, European Molecular Biology Laboratory - European Bioinformatics Institute
(EMBL-EBI), United Kingdom
-
Christoph N. Schlaffner, Hasso Plattner Institute for Digital Engineering, Digital Engineering Faculty,
University of Potsdam, Germany
-
Bernhard Y. Renard, Hasso Plattner Institute for Digital Engineering, Digital Engineering Faculty,
University of Potsdam, Germany
Presentation Overview: Show
The large functional diversity of the proteome is possible through post-translational modifications (PTMs)
resulting in up to one million different proteoforms. This large diversity harbours the challenge of
considering all possible combinations of around 300 different PTMs when identifying proteomic mass spectra. To
combat this, AHLF, a Deep Learning binary classifier, trained on 10.5 million Phosphorylation spectra can
stratify unseen mass spectra, for targeted searches with and without modification. For Phosphorylation as the
best studied PTM large amounts of data are readily available. However, other PTMs such as Acetylation and
Ubiquitination suffer from drastically fewer studies, and for Deep Learning commonly insufficient data. To
combat this problem, we applied a transfer-learning approach utilizing the AHLF model pre-trained on
Phosphorylation to build models for Acetylation and Ubiquitination.
Our Ubiquitination set containing around 2 million high quality modified peptide-spectrum matches (PSMs) was
curated from 11 publicly available ubi-enriched datasets, while the Acetylation set contained 0.5 million
modified PSMs from 9 public enriched datasets. The fine tuned models for Ubiquitination and Acetylation
achieve AUCs of 0.90 and 0.87, respectively. Further artificially rescued training sets show that around
28,500 modified PSMs (0.3% of the Phosphorylation dataset) already result in sufficiently converging models.
We also show that our fine-tuned models incorporating label free and SILAC labelled spectra outperform
fine-tuned models trained on similarly sized datasets stratified by quantification method. This highlights
that fine-tuning models with mixed labelling states boosts classification performance and drastically
increases the availability of data for training.
C-P.40: Plant Biocuration in UniProtKB/Swiss-Prot
Track: Proteins and structural biology
-
Emmanuel Boutet, SIB - Swiss Institute of Bioinformatics, Switzerland
- The Uniprot Consortium Uniprot, SIB - Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
The UniProt Knowledgebase (UniProtKB, https://www.uniprot.org) is a comprehensive, freely accessible resource
providing high-quality protein sequences and functional data. Its expert-curated UniProtKB/Swiss-Prot section
contains approximately 580,000 sequences, including around 42,000 from plants such as Arabidopsis thaliana and
Oryza sativa (release 2026_01).
A. thaliana remains a key model organism in plant biology. Building on this importance, a collaborative effort
led by The Arabidopsis Information Resource (TAIR) and involving UniProtKB/Swiss-Prot has contributed to an
updated genome assembly of the Columbia cultivar. This assembly is now integrated into UniProtKB, ensuring
consistency between genomic and protein-level data.
Within this framework, a central focus lies in the annotation of plant enzymes. Biochemical reactions are
described using Rhea (https://www.rhea-db.org), which provides standardized, computable descriptions of
biochemical reactions. This integration enhances the accuracy and interconnectivity of enzyme function
annotations. At present, the database comprises 16,900 manually curated plant enzyme entries, including 5,993
from Arabidopsis thaliana and 3,450 from various species selected to represent diverse biosynthetic
pathways.
Together, these efforts produce structured, high-quality data that enable interoperability with other
resources and support metabolic modelling, multi-omics integration, and machine learning approaches for
predicting enzyme functions and plant biosynthetic pathways.
C-P.41: Generating and evaluating coevolution and phylogeny-aware MSAs
Track: Proteins and structural biology
-
Anamay Samant, 1) Institute of Bioengineering, School of Life Sciences (EPFL) 2) Swiss Institute of
Bioinformatics (SIB), Switzerland
-
Anne-Florence Bitbol, 1) Institute of Bioengineering, School of Life Sciences (EPFL) 2) Swiss Institute of
Bioinformatics (SIB), Switzerland
Presentation Overview: Show
Molecular phylogenetic inference involves inferring evolutionary relationships from a multiple sequence
alignment (MSA). Generally, a single phylogenetic tree or tree distribution is inferred for one MSA at a time
by methods like maximum likelihood and Bayesian inference. Alternatively, a generalizable, supervised learning
approach requires training data in the form of “MSA-ground truth tree†pairs. Phylogenetics lacks actual
ground-truths, so training data needs to be generated by simulating evolution along a known tree to yield leaf
sequences constituting an MSA. The tree is then by definition the ground truth tree for the generated MSA.
However, the choice of the simulation method is crucial for obtaining realistic MSAs from known trees.
Conventionally, simulation is performed using models that assume independent evolution across all sites in a
sequence, while sites that are functionally coupled tend to coevolve and feature correlations in amino-acid
usage. Recognising the importance of accounting for these interactions, we developed a Markov Chain Monte
Carlo (MCMC) based simulation method that incorporates interactions flexibly, using Potts models, or
single-sequence protein language models (PLMs), or MSA-based PLMs. We compared MSAs obtained from these three
model types in terms of their closeness to protein families of interest, diversity within generated sequences
and novelty when compared to natural sequences. Results indicate more favourable metrics when using the
PLM-based generation methods, with the MSA-based PLM having an edge when considering shallow protein families.
We are now using these coevolution-aware MSAs as more realistic training data to develop machine-learning
models for tree inference.
C-P.42: AI-driven structural modeling reveals hidden functional and evolutionary relationships in divergent
dsRNA viruses
Track: Proteins and structural biology
- Edouard De Castro, SIB Swiss Institute of Bioinformatics, Switzerland
- David Moi, SIB Swiss Institute of Bioinformatics, Switzerland
- Gerardo Tauriello, SIB Swiss Institute of Bioinformatics, Switzerland
- Paul Thomas, SIB Swiss Institute of Bioinformatics, Switzerland
- Jelle Matthijnssens, Rega Institute, Laboratory of Viral Metagenomics, Belgium
-
Houssam Attoui, The National Research Institute for Agriculture, Food and Environment (INRAe), France
-
Philippe Le Mercier, SIB Swiss Institute of Bioinformatics, Switzerland
-
Fauziah Mohd Jaafar, The National Research Institute for Agriculture, Food and Environment (INRAe), France
Presentation Overview: Show
Double-stranded RNA (dsRNA) viruses harbor numerous proteins whose functions remain poorly understood due to
extreme sequence divergence that defeats conventional annotation methods. Here, we demonstrate that AI-driven
structural modeling systematically overcomes this barrier, enabling functional assignment across highly
divergent viral proteomes.
Using AlphaFold3 combined with Foldseek structural similarity searches, we predicted three-dimensional
structures of proteins from representative Reovirales and Ghabrivirales members. Structure-based analyses
identified multiple previously uncharacterized virion components—inner capsid, outer capsid, and capping
enzymes (turret proteins)—providing a near-complete structural map despite negligible sequence similarity to
known homologs. Critically, structural modeling enabled the first structure-based phylogenetic reconstruction
of capsid proteins, circumventing sequence-based limitations.
Remarkably, structural modeling of Micromonas pusilla reovirus (MpRV) revealed that its outer capsid protein
adopts a fold closely related to the birnavirus capsid, despite these viruses being deeply divergent. This
discovery suggests ancient horizontal gene transfer and demonstrates that capsid modules can be exchanged
between distantly related dsRNA viruses.
Our results demonstrate that AI-guided structural modeling not only expands functional annotation of viral
proteomes but provides a powerful framework for exploring deep evolutionary relationships among highly
divergent proteins. This approach reveals hidden functional and evolutionary connections within the virosphere
that remain inaccessible to traditional genomic methods, opening new avenues for understanding viral evolution
and diversity.
C-P.43: The sequence-structure landscape of antibody framework regions
Track: Proteins and structural biology
-
Teodora Christina Purice, Institute of Biochemistry of the Romanian Academy, Romania
- Anca Iacob, Institute of Biochemistry of the Romanian Academy, Romania
- Laurentiu Spiridon, Institute of Biochemistry of the Romanian Academy, Romania
- Andrei-Jose Petrescu, Institute of Biochemistry of the Romanian Academy, Romania
Presentation Overview: Show
Due to the low antibody scaffold stability, aggregation remains a major problem in therapeutic development.
State-of-the-art AI-MD aggregation predictors are constrained by limited available framework region (FR)
templates, while full-antibody modeling tools primarily focused on hypervariable complementarity-determining
regions (CDRs) and hinges can introduce FR inaccuracies, degrading overall stability. Addressing these
problems requires a data-driven landscape of the FR sequence and structural space.
We present here results on the development of a pipeline extracting information on >3000 human and mouse
immunoglobulin structures (CryoEM, X-Ray) from SabDab stratified by interaction state, chain, and domain (Fv,
Fab, Fc, full Ig). AHo-aligned, outlier-filtered FR sequences were profiled (PSSM heatmaps) and clustered
(30-90%, MMseqs2), visualized as network graphs, distance heatmaps, and dendrograms. Secondary structure
assignment step yields hydrogen-bond maps and similarity clustering; Ramachandran and bond-angle torsion
analyses provide torsion-angle distributions, statistical overview, and similarity clustering.
Sequence-structure correlations generated a non-redundant template set. Structure prediction benchmarking has
been initiated with RaptorX-Single, with ABodyBuilder-3, AlphaFold-3, and Ibex to follow.
Outputs confirm canonical FR features (C23, W43, C106) and composition patterns. Light chain FRs exhibit
greater redundancy (20% templates required for full set coverage) than heavy chains (30-34%). Multi-threshold
clustering was used to identify substructure families. DSSP, contact maps, internal coordinates analyses and
structural similarity clustering have been obtained across the full set. RaptorX-Single predictions have been
generated for all Fvs.
This empirical picture of FR sequence and structure space provides a foundation to build and implement a
robust modeling workflow for full antibodies and antibody-antigen complexes.
C-P.44: Dynamic Behavior of the RAG2 Acidic Region and Its Possible Functional Significance
Track: Proteins and structural biology
-
Anca-L Iacob, Institute of Biochemistry of the Romanian Academy, Romania
- Eliza - Cristina Martin, Yale University, United States of America
- Andrei - Jose Petrescu, Institute of Biochemistry of the Romanian Academy, Romania
Presentation Overview: Show
The RAG recombinase drives V(D)J recombination and adaptive immunity, yet its evolutionary origin from a
Transib-family transposon means it retains a latent transposase activity threatening genomic stability and
contributing to chromosomal translocations and leukemias. Although molecular domestication has introduced
suppressive adaptations, the mechanistic role of the RAG2 acidic hinge (AH), an intrinsically disordered
region linking the Kelch and PHD domains remains poorly understood.
The AH was modeled in a fully extended state from the resolved RAG2 core using Modeller, then explored through
70 independent molecular dynamics simulations (100 ns each) with OpenMM, CHARMM36 force field, and implicit
solvent at 310K. Trajectories were clustered to identify topological states and mapped onto the RAG tetramer
surface.
Five recurrent conformational states provide a structural framework for understanding AH-mediated inhibition.
In the most prevalent clusters, the AH localizes above the RAG1 DNA-binding groove, creating a steric and
electrostatic barrier blocking target DNA acquisition and preventing the U-shape conformation required for
transposition. A distinct cluster shows the AH intercalated at the lateral RAG1/RAG2 interface, acting as an
allosteric wedge restricting inter-subunit flexibility essential for transposition. One particularly
informative cluster shows the AH extended toward the RAG1 basic N-terminal region.
The acidic hinge thus emerges as a dynamic regulatory element whose electrostatic steering by basic surface
patches positions it strategically to inhibit transposition and safeguard genomic stability during V(D)J
recombination.
C-P.45: Mapping Evolutionary Switches Driving Functional Diversification in AsnC-like Transcription
Factors
Track: Proteins and structural biology
-
Luc Lafrenaye, Institute of Molecular Systems Biology, ETH Zürich, Switzerland
- Pedro Beltrao, Institute of Molecular Systems Biology, ETH Zürich, Switzerland
- Julian Trouillon, Institute of Molecular Systems Biology, ETH Zürich, Switzerland
Presentation Overview: Show
Deciphering the transcriptional regulatory code requires understanding how a transcription
factor's amino acid sequence dictates its DNA-binding specificity. AsnC-like transcription
factors (InterPro: IPR019888) form an ancient regulatory superfamily that offers a dense
evolutionary sampling for this purpose. Here, we leverage the phylogeny of the AsnC-like
family to systematically map this functional diversification.
To identify the key changes causing functional diversification, we used the burst after
duplication divergence metric with ancestral sequence reconstructions. This approach
highlights "evolutionary switches"
- residues conserved within subfamilies but different
between them, suggesting strong functional importance. By mapping the high-scoring
residues onto structural models, we can deduce their specific biological roles: switches at
the protein-DNA interface suggest residues that determine specificity, while those at
effector-binding or dimerization sites show the evolution of metabolic sensing and complex
assembly.
Beyond this, we are broadening our analysis to explore the larger factors and timing of these
functional shifts. By characterizing the physicochemical nature of the mutations at switch
sites, we aim to gain understanding of how specific substitutions alter function. By linking the
identified evolutionary switches with known DNA-binding specificities, effector molecules,
and the host species' environments, we can deduce the selective pressures that drive
regulatory adaptation. Mapping these switches across the phylogeny also allows us to
estimate when these adaptations emerged. This evolutionary approach ultimately aims at
clarifying how regulatory domains take on novel functions in the AsnC-like transcription
factor family, offering a sound basis for understanding adaptation through protein evolution.
C-P.46: Residue-Level Attributions in Protein Language Models Do Not Recover Allergen Epitopes
Track: Proteins and structural biology
-
Jianzhou Yao, Swiss Institute of Allergy and Asthma Research, Davos; ETH Zurich, Zurich,
Switzerland
-
Anxiong Song, Swiss Institute of Allergy and Asthma Research, Davos; ETH Zurich, Zurich, Switzerland
-
Katja Baerenfaller, Swiss Institute of Allergy and Asthma Research, Davos; Swiss Institute of
Bioinformatics, Lausanne, Switzerland
-
Damir Zhakparov, Swiss Institute of Allergy and Asthma Research, Davos; Swiss Institute of Bioinformatics,
Lausanne, Switzerland
Presentation Overview: Show
Background. Protein language models (PLMs) achieve state-of-the-art allergenicity classification, yet the
molecular basis of these predictions remains uncharacterized. Residue-level attribution methods, including
Integrated Gradients (IG), are widely interpreted as localizing immunologically relevant regions, but this
claim has only been supported by comparisons to known epitopes and has not been quantitatively assessed
against curated immunological ground truth.
Methods. We developed a residue-level benchmark for assessing immunological faithfulness in PLM-based
allergenicity prediction models, defined as concordance between residue-level attribution scores and
experimentally validated MHC Class II epitopes from the Immune Epitope Database (IEDB). We evaluated
attribution scores across ESM-2-based classifiers, including a multi-task architecture with an auxiliary
residue-level head trained under direct epitope supervision. To differentiate immunological faithfulness from
model faithfulness, we conducted IG-guided masking and in silico saturation mutagenesis at high-attribution
positions.
Results. Residue-level attribution scores exhibited near-random concordance with IEDB-annotated epitopes,
despite high protein-level performance. Multi-task-learning recovered epitope signal in the auxiliary residue
head but did not increase epitope concordance of classifier attributions, indicating epitope-relevant features
are not recruited by the classification objective even when available during optimization. Masking
high-attribution residues reduced prediction confidence, confirming attribution faithfulness to model
decisions. Saturation mutagenesis revealed sensitivity to physicochemical properties and local compositional
context rather than epitope-defining substitutions.
Conclusion. Model and immunological faithfulness are distinct: attributions can reflect decision drivers while
failing to recover immunologically meaningful sequence properties. Residue-importance maps cannot be treated
as biological explanations without epitope-based validation. We propose quantitative faithfulness benchmarking
as necessary for interpretability evaluation in allergenicity prediction.
C-P.47: Computational search for Myelin-associated Glycoprotein (MAG) binders that could modulate the
neuroprotective properties of oligodendrocytes in neurodegenerative contexts
Track: Proteins and structural biology
-
Carmela Felippa Ambort, Department of Theoretical and Computational Chemistry, School of Chemical
Sciences, National University of Córdoba, Argentina
-
Rodrigo Quiroga, Department of Theoretical and Computational Chemistry, School of Chemical Sciences,
National University of Córdoba, Argentina
Presentation Overview: Show
Myelin-Associated Glycoprotein (MAG/Siglec-4) is a cell-surface lectin belonging to the Siglec family, which
binds sialic acid–containing glycans. MAG is expressed in myelinating oligodendrocytes of the central nervous
system and is localized in the periaxonal space. It shows high specificity for the Neu5Acα2-3Galβ1-3GalNAc
motif present in gangliosides such as GT1a and GD1b. Structurally, this interaction is mediated by a conserved
Arg118 residue in the V domain, which forms a salt bridge with the sialic acid carboxyl group. Upon
activation, MAG promotes glutamate reuptake after injury, contributing to neuroprotective and antioxidant
effects in oligodendrocytes and neurons. This function is particularly relevant in neurodegenerative
conditions such as stroke and multiple sclerosis.
The aim of this study is to identify potential MAG activators or inhibitors with high affinity and selectivity
through molecular docking and virtual screening of diverse compound libraries (e.g., ZINC20). For this
purpose, the 2VINARDO scoring function, an enhanced version of AutoDock Vina, is employed. Developed by
Quiroga and Villareal, 2VINARDO improves the description of non-covalent interactions by expanding atom types
and interaction parameters.
Validation of the scoring function was performed using a re-docking protocol with seven crystallographic
Siglec–ligand complexes, including MAG. In all cases, ligand poses were accurately predicted (RMSD < 2 Å),
preserving key interactions such as hydrogen bonds involving Arg118. Additional analyses include binding site
flexibility and correlation with experimental binding affinities (Kd).
These results support the application of 2VINARDO for large-scale virtual screening, aiming to identify
selective MAG modulators with potential therapeutic relevance.
C-P.48: A novel analysis pipeline for scFv discovery from long-read platforms
Track: Proteins and structural biology
-
Ozge Gizlenci, AstraZeneca, United Kingdom
- Gareth Griffin, AstraZeneca, United Kingdom
- Jurgen Haas, AstraZeneca, United Kingdom
Presentation Overview: Show
Antibody discovery increasingly relies on sequencing to validate candidates and characterise repertoire
diversity. While next-generation sequencing enables fast, cost-effective, high-throughput analysis, long-read
platforms are becoming especially valuable for single-chain variable fragment (scFv) discovery because they
can capture full-length constructs and preserve VH-VL pairing. This is particularly important for candidates
designed through machine learning-guided workflows.
We developed a novel platform- and UMI-aware long-read analysis workflow for scFv discovery using the
long-read technologies recently implemented in-house, including PacBio HiFi and Oxford Nanopore Technologies
(ONT). The pipeline performs preprocessing, platform-specific consensus generation, translation of scFv
constructs, domain annotation with ANARCI, and automated CDR3 extraction, producing quantitative summaries and
exportable TSV/HTML reports for downstream library characterisation. It consistently recovers key scFv
features, supports repertoire-level CDR3 analysis, and provides high-level quality metrics across datasets.
Long-read sequencing offers clear advantages for antibody discovery, but it also presents challenges. Compared
with short-read methods, long-read data can be more error-prone, particularly for ONT, and may show variable
quality profiles across reads. This can complicate accurate domain identification, CDR parsing, translation,
and clonotype quantification. Long-read workflows may also have lower throughput, higher cost per read, and
greater computational demands for consensus generation and error correction. In addition, currently available
tools rarely provide end-to-end paired-domain annotation and quantitative reporting tailored to scFv long-read
datasets.
Next, we plan to extend the pipeline to Illumina's emerging long-read approach, evaluate ANARCI II and
benchmark against tools such as RIOT, with the goal of integrating the workflow into a broader modular NGS
platform.
C-P.49: Investigating Enzyme Function by Geometric Matching of Catalytic Motifs
Track: Proteins and structural biology
-
Raymund Hackett, Leiden University Medical Center, European Bioinformatics Institute (EMBL-EBI),
Netherlands
- Ioannis Riziotis, European Bioinformatics Institute (EMBL-EBI), United Kingdom
- Martin Larralde, Leiden University Medical Center, Netherlands
- António J. M. Ribeiro, European Bioinformatics Institute (EMBL-EBI), Portugal
- Georg Zeller, Leiden University Medical Center, Netherlands
- Janet Thornton, European Bioinformatics Institute (EMBL-EBI), United Kingdom
Presentation Overview: Show
The rapidly growing universe of predicted protein structures offers opportunities for data driven exploration
but requires computationally scalable and interpretable tools. We developed a method, Enzyme Motif Miner, to
detect catalytic features in protein structures, providing insights into enzyme function and mechanism. A
library of 6780 3D coordinate sets describing enzyme catalytic sites, referred to as templates, has been
collected from manually curated examples of 762 enzyme catalytic mechanisms described in the Mechanism and
Catalytic Site Atlas. We implemented RMSD and residue orientation filters to differentiate catalytically
informative matches from spurious ones. We validated this approach on a non-redundant set of high quality
experimental (n=3751, <40% amino acid identity) enzyme structures with well annotated catalytic sites as
well as predicted structures of the human proteome. We show that matching catalytic templates is more
sensitive than sequence- and 3D-structure-based approaches in identifying homology between distantly related
enzymes. Since geometric matching does not depend on conserved sequence motifs or even common evolutionary
history, we are able to identify examples of structural active site similarity in highly divergent and
possibly convergent enzymes. Such examples make interesting case studies into the evolution of enzyme
function. Though not intended for characterizing substrate-specific binding pockets, the speed and
knowledge-driven interpretability of our method make it well suited for expanding enzyme active-site
annotation across large predicted proteomes. Enzyme Motif Miner is available as a python module at
https://github.com/rayhackett/enzymm and as a webserver at https://www.ebi.ac.uk/thornton-srv/m-csa/enzymm.
C-P.50: Unveiling Recurrent Binding Sites in H1N1 Nucleoprotein via Ensemble-Based Pocket Clustering
Track: Proteins and structural biology
-
Xinyu Qi, UniversiteÌ Paris CiteÌ, CNRS, Inserm, UniteÌ de Biologie Fonctionnelle et Adaptative, F-
75013 Paris, France, France
-
Inés Sabine Rahali, UniversiteÌ Paris CiteÌ, CNRS, Inserm, UniteÌ de Biologie Fonctionnelle et Adaptative,
F- 75013 Paris, France, France
-
Sandie Munier, Institut Pasteur, Université Paris Cité, Lyssavirus Epidemiology and Neuropathology Unit,
F-75015 Paris, France, France
-
Anne Badel, UniversiteÌ Paris CiteÌ, CNRS, Inserm, UniteÌ de Biologie Fonctionnelle et Adaptative, F-
75013 Paris, France, France
-
Delphine Flatters, UniversiteÌ Paris CiteÌ, CNRS, Inserm, UniteÌ de Biologie Fonctionnelle et Adaptative,
F- 75013 Paris, France, France
-
Anne-Claude Camproux, UniversiteÌ Paris CiteÌ, CNRS, Inserm, UniteÌ de Biologie Fonctionnelle et
Adaptative, F- 75013 Paris, France, France
Presentation Overview: Show
Influenza A nucleoprotein (NP) is a conserved multifunctional protein involved in RNA binding and
oligomerization, making it an attractive antiviral target. However, static structures provide only a partial
view of its ligandable regions, as conformational variability can alter pocket accessibility and predicted
druggability.
Here, we characterized the binding site landscape of H1N1 NP using molecular dynamics simulations combined
with ensemble-based pocket detection and residue-based clustering. Across 753 conformations, 16,404 detected
pockets were clustered by residue composition to reconstruct recurrent binding site candidates and assess
their recurrence and predicted druggability. Seven recurrent sites were retained based on recurrence and
cluster-level predicted druggability and were further characterized according to residue-level organization,
plasticity, and their position relative to NP–RNA and NP–NP interface regions. These sites displayed
distinct structural behaviors, included stable compact NP head sites, RNA-interface-associated sites, and a
dynamic extended site spanning the inter-domain groove. Comparison with the static reference structure
distinguished stable, shifted, and dynamically emerging sites, including regions not apparent in the reference
conformation.
Overall, this study provides an ensemble-based structural framework to reconstruct and prioritize recurrent
ligandable regions in influenza A NP. By integrating recurrence, predicted druggability, site organization,
and structural plasticity, it refines the description of NP binding sites in H1N1 and opens the way to ongoing
comparative work with H5N1 NP to determine which recurrent binding sites are conserved, shifted, or
subtype-specific across pathogenic influenza A subtypes.
C-P.51: Tokenization-Aware Protein Language Modeling with Evolution-Guided Units and Dynamic Biological
Knowledge Fusion
Track: Proteins and structural biology
-
Amirreza Sattarzadeh Khanehbargh, Boğaziçi University, Türkiye
- Burak Suyunu, Boğaziçi University, Türkiye
- Özdeniz Dolu, Boğaziçi University, Türkiye
- Arzucan Özgür, Boğaziçi University, Türkiye
Presentation Overview: Show
Protein language models (PLMs) have become key tools for computational protein analysis, but most still rely
on single amino-acid tokens as the default representation unit. While simple and effective, this choice can
limit biological expressiveness and may interact in non-trivial ways with downstream adaptation strategies. We
present a tokenization-aware framework that compares character-level tokenization, frequency-based subword
tokenization such as BPE and Unigram, and PUMA, a mutation-aware tokenizer that organizes sequence patterns
into mutation-informed token families.
Using multiple pretrained backbones, we study various strategies for implementing tokenization in protein
language modeling, including frozen encoders, full fine-tuning, and parameter-efficient LoRA. Experiments
cover representative tasks in computational biology, including subcellular localization, protein function
classification, and stability prediction. Beyond accuracy, we assess memory usage, sequence compression,
compute cost, and interpretability. We further examine how token-level representations can be enriched by
integrating biochemical, functional, and evolutionary priors through task-aware knowledge fusion. All in all,
we provide an analysis of the effect of tokenization on protein representation learning performance across
computational and biological dimensions.
Our results indicate that PUMA can outperform frequency-based subword tokenization in several settings, while
vocabulary size has a substantial effect on both predictive performance and efficiency. Amino-acid
tokenization remains a strong baseline overall; however, at certain vocabulary sizes, PUMA can match or
surpass amino-acid-level representations, suggesting that biologically informed tokenization can offer
practical advantages when appropriately configured.
Overall, this work provides a structured comparison of tokenization choices, adaptation regimes, and
representation quality in PLMs. By analyzing their combined effects across multiple biological tasks, our
framework offers practical guidance for selecting tokenization and fine-tuning strategies and introduces an
extensible benchmark for biologically informed protein sequence modeling.
C-P.52: Computational insights into enhanced deubiquitination by SARS-CoV-2 PLpro K232Q – A structure -
function study
Track: Proteins and structural biology
-
Janani Ganesh, Homi Bhabha National Institute, India
- Rimanshee Arya, Homi Bhabha National Institute, India
- Vishal Prashar, Homi Bhabha National Institute, India
- Mukesh Kumar, Homi Bhabha National Institute, India
Presentation Overview: Show
Papain-like protease (PLpro) of SARS-CoV-2, a domain of non-structural protein 3 (Nsp3), plays dual role in
its pathogenesis by processing viral polyproteins and modulating host immunity through deubiquitination and
deISGylation. To examine how mutations influence these functions, we performed large-scale computational
analysis of ~14 million Nsp3 sequences from the GISAID database using custom Python pipelines (Arya, et al.,
2023: Microbial Pathogenesis, 185, 106460). This analysis identified five mutations (A145D, P77L, P77S, V187A,
and K232Q) with significant global prevalence, associated with multiple variants of concern. Among these, we
found that the K232Q mutation significantly increased (~four-fold) deubiquitination activity compared to the
wild-type enzyme in our biochemical assays. Sequence comparison with SARS-CoV-1 suggested that this
substitution represents a reversion to its ancestral residue, consistent with the higher deubiquitination
activity observed in SARS-CoV-1 PLpro. Structural mapping placed residue 232 at the ubiquitin-binding
interface, prompting molecular dynamics simulations of the PLpro-ubiquitin complex (PDB: 7RBR) which revealed
a shift in interaction behaviour. In the wild-type, K232 alternates between interactions with A46(Ub) and
Y207(PLpro), whereas Q232 forms more balanced and persistent contacts with both residues, along with a shorter
interaction distance to A46(Ub). Trajectory analysis using MDAnalysis indicated changes in substrate
positioning near the catalytic site and buried surface area, while DCCM and MM-PBSA analyses further revealed
differences in correlated motions and binding energetics between mutant and wild-type systems. Together, these
results provide a structural basis for the enhanced catalytic activity of K232Q PLpro to uncover subtle
changes in enzyme–substrate dynamics due to mutation.