View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
A-P.01: PLM-eXplain: Divide and Conquer the Protein Embedding Space
Track: Proteins and structural biology
- Jan van Eck, Utrecht University, Netherlands
- Dea Gogishvili, Utrecht University, Netherlands
-
Wilson Silva, Utrecht University, Netherlands
- Sanne Abeln, Utrecht University, Netherlands
Presentation Overview: Show
Motivation: Protein language models (PLMs) have revolutionised computational biology through their ability to
generate powerful sequence representations for diverse prediction tasks. However, their black-box nature
limits biological interpretation and translation to actionable insights. Bridging this gap requires approaches
that maintain predictive performance while providing interpretable explanations of model behaviour.
Results: We present PLM-eXplain (PLM-X), an explainable adapter layer that bridges this gap by factoring PLM
embeddings into two complementary components: an interpretable subspace based on established biochemical
features, and a residual subspace that retains predictive, non-interpretable information. Using embeddings
from ESM2 and ProtBert, PLM-X incorporates well-established properties, including secondary structure and
hydropathy, while maintaining high predictive performance. We demonstrate the effectiveness of our approach
across three biologically relevant classification tasks: extracellular vesicle association, transmembrane
helix prediction, and aggregation propensity prediction. PLM-X enables biological interpretation of model
decisions without sacrificing accuracy, offering a generalisable solution for enhancing PLM interpretability
across various downstream applications.
A-P.02: A self-assembled protein β-helix as a self-contained biofunctional motif
Track: Proteins and structural biology
-
Javier Garcia-Ruiz, National Physical Laboratory, United Kingdom
- Camilla Dondi, National Physical Laboratory, United Kingdom
- Maxim G. Ryadnov, National Physical Laboratory, United Kingdom
Presentation Overview: Show
Nature constructs matter by employing protein folding motifs, many of which have been synthetically
reconstituted to exploit function. A less understood motif whose structure-function relationships remain
unexploited is formed by parallel beta-strands arranged in a helical repetitive pattern, termed a beta-helix.
Herein we reconstitute a protein beta-helix by design and endow it with biological function. Unlike
beta-helical proteins, which are contiguous covalent structures, this beta-helix self-assembles from an
elementary sequence of 18 amino acids. Using a combination of experimental and computational methods, we
demonstrate that the resulting assemblies are discrete cylindrical structures exhibiting conserved dimensions
at the nanoscale. We provide evidence for the structures to form a carpet-like three-dimensional scaffold
promoting and inhibiting the growth of human and bacterial cells, respectively, while being able to mediate
intracellular gene delivery. The study introduces a self-assembled beta-helix as a self-contained bio- and
multi-functional motif for exploring and exploiting mechanistic biology.
A-P.03: Structure-Guided RNA Design via Multinomial Diffusion
Track: Proteins and structural biology
-
Georg Back, Weihenstephan-Triesdorf University of Applied Sciences, Department of Sustainable
Agriculture and Energy Systems, Germany
-
Dirk Walther, Max Planck Institute for Molecular Plant Physiology, Department of Bioinformatics, Germany
-
Florian Haselbeck, Weihenstephan-Triesdorf University of Applied Sciences, Department of Sustainable
Agriculture and Energy Systems, Germany
Presentation Overview: Show
The structure of RNA plays an essential role in its biological function, with the secondary structure being
the primary determinant. While significant advances have been made in RNA folding models, the inverse problem,
i.e., generating sequences that fold into a target structure has received less attention. Existing approaches
commonly rely on iterative optimization, often guided by a folding oracle, which can limit scalability and
risk converging to local optima. Diffusion models have shown a superior performance in many application areas,
but their noising process is not directly applicable to discrete data such as nucleotide sequences. In this
paper, we adapt the multinomial diffusion framework, which introduces discrete noise via categorical
transitions, to the RNA sequence space. Furthermore, we incorporate secondary structure information to guide
the generative process toward sequences that fold into a desired target structure. For this purpose, we
evaluate two architectural variants: (i) modeling the problem either by projecting the sequence embedding into
a 2D pairwise representation that reflects the structure of a contact map, or (ii) by restructuring the
architecture as a purely 1D sequence model. Our results show that both approaches generate sequences that fold
with high accuracy while maintaining substantial variability among generated sequence samples. We further
evaluated constrained generation via sequence inpainting, resulting in performance comparable to unconstrained
generation. Even though the performance on the human-designed Eterna100 benchmark was lower than that of the
state-of-the-art method LEARNA, our results indicate that diffusion-based models provide a promising and
scalable framework for conditional nucleotide sequence generation.
A-P.04: Protein language models improve detection of divergent beta-lactamases
Track: Proteins and structural biology
-
Mateusz Wlodarski, McMaster University, Canada
- Andrew McArthur, McMaster University, Canada
Presentation Overview: Show
Accurate detection of β-lactamase genes is critical for antimicrobial resistance surveillance, yet highly
divergent sequences with little similarity to curated enzymes often escape detection. Although protein
language models (PLMs) show promise for functional inference, their ability to detect remote resistance genes
and the biological basis of their predictions remain unclear.
We evaluated transformer-based PLMs for β-lactamase detection across controlled sequence divergence bins.
Using the Comprehensive Antibiotic Resistance Database (CARD) as a curated reference space, we stratified
withheld β-lactamase sequences by identity to their nearest CARD homolog and compared performance against
sequence similarity and Hidden Markov model -based approaches. PLM classifiers maintained high recall in
low-identity bins where traditional methods failed, while larger representations such as Evolutionary Scale
Modeling (ESM-2) also improved false positive rates on hydrolase negatives. Embedding analyses revealed Ambler
class-consistent structure, and motif-centered perturbation analyses demonstrated that conserved catalytic
regions contribute disproportionately to model predictions, supporting the biological relevance of learned
features. Together, these results demonstrate that PLMs can extend β-lactamase detection beyond the limits of
traditional sequence similarity methods and highlight the benefits of larger capacity learned representations.
A-P.05: TRPM8 Virtual Screening: The Essential Role of Target-Specific, Inactive-Enriched Machine-Learning
Scoring Functions
Track: Proteins and structural biology
-
Nivya James, Imperial College London, United Kingdom
- Pedro Ballester, Imperial College London, United Kingdom
Presentation Overview: Show
Motivation: The Transient Receptor Potential Melastatin 8 (TRPM8) ion channel is an emerging therapeutic
target implicated in pain, inflammation, and other disorders. While structure-based virtual screening (VS) is
widely used in drug discovery, the effectiveness of generic and target-specific machine-learning (ML) scoring
functions (SFs) for ion channels such as TRPM8 remains to be investigated.
Results: We established the first comprehensive structure-based VS benchmark for TRPM8. This evaluates
classical docking SFs, generic ML SFs and target-specific ML models retrospectively across multiple TRPM8
protein conformations and docking protocols. The evaluation revealed that generic ML rescoring at most
provided modest and conformation-dependent improvements over classical docking on this target. Target-specific
ML models based only on protein-ligand interaction fingerprints improved early enrichment on test sets when
regression algorithms were employed. However, performance remained sensitive to protein conformation choice
and chemical dissimilar-ity. Systematic evaluation of feature representations showed that combining
structure-derived interaction features with ligand-based descriptors enhanced enrichment under selected
docking tool-protein conformation combinations. Critically, presenting the learning algorithm with many more
negative training instances, via inactive-enriched training, strongly reduced false positives. This also
resulted in markedly improved generalization to chemically dissimilar test sets for structure-based models. By
contrast, ligand-only QSAR models proved sensitive to more class-imbalanced training sets. Under
inactive-enriched training, PLEC-based support vector regression (SVR) models achieved robust and strong VS
performance, highlighting the critical role of learning the vast diversity of inactive molecules better during
model training in developing reliable target-specific ML SFs for TRPM8.
A-P.06: A pretrained model for RNA inverse folding with formal grammar representations
Track: Proteins and structural biology
- Kentaro Watanabe, Keio University, Japan
- Manato Akiyama, Kitasato University, Japan
-
Yasubumi Sakakibara, Kitasato University, Japan
Presentation Overview: Show
Motivation: RNA inverse folding designs nucleotide sequences that fold into a given RNA secondary structure,
enabling the creation and modification of functional RNAs. State-of-the-art solvers are largely search-based,
but the search space grows exponentially with sequence length and structural complexity, and it is non-trivial
to incorporate additional sequence-level constraints such as target GC content or family-specific conserved
features.
Results: We propose a generative model based on a Transformer-based conditional variational au-toencoder
(CVAE) that takes a context-free grammar (CFG) parse-tree representation of the target secondary structure as
input and generates sequences that fold into the specified secondary struc-ture. The grammar-based tree makes
the hierarchical organization and base-pair correspondences explicit. By combining self-refinement learning
and latent-space optimization, we substantially im-prove the recovery of high-fidelity solutions. On the
EteRNA100 benchmark, our model alone achieves competitive accuracy, and the generated sequences consistently
improve the success rate of downstream search-based solvers when used as warm starts. We further demonstrate
controllable generation under GC-content constraints and improved family consistency through large-scale
pre-training on natural RNAs.
A-P.07: Energy-Matched Diffusion for Antibody Generation
Track: Proteins and structural biology
- Vasanth Durvasula, Nanyang Technological University, Singapore
- Tiara Natasha Binte Sayuti, Nanyang Technological University, Singapore
-
Jagath Rajapakse, Nanyang Technological University, Singapore
Presentation Overview: Show
Motivation: Diffusion models have been used for antibody design, which enable joint generation of
complementarity-determining region (CDR) sequence and structure, but their denoising objectives are typically
optimized locally and do not explicitly enforce global geometric consistency. As a result, generated
structures exhibit accumulated geometric drift despite locally accurate predictions.
Results:
We propose an energy-matched diffusion framework that learns an antibody-specific scalar energy function
jointly with the denoising network. During training, an energy-matching objective aligns the gradient of the
learned energy with diffusion-consistent denoising targets, constraining the denoising vector field to arise
from a single scalar potential. This introduces a global integrability bias that promotes smoother, more
consistent denoising trajectories and reduces structural instability during generation. To avoid hand-tuned
multi-term loss weights, we introduce an exponential-moving-average controller that adaptively scales the
energy-matching term relative to a reconstruction anchor throughout training.
On antibody–antigen complexes from the SAbDab benchmark, the proposed method improves both sequence quality
and backbone accuracy, achieving lower RMSD than an unguided diffusion baseline. Through our results, we show
that energy matching serves as an effective training-time regularizer for improving geometric fidelity in
antibody CDR generation without requiring external energy supervision or manual hyperparameter search.
A-P.08: Gated Modulation Protein Language Model for Prediction of Antimicrobial Peptide Activity
Track: Proteins and structural biology
- Prem Singh Bist, Nanyang Technological University, Singapore
-
Jagath Rajapakse, Nanyang Technological University, Singapore
Presentation Overview: Show
Motivation: Antibiotic resistance poses a critical global health threat, demanding new strategies for rapid
therapeutic discovery. Antimicrobial peptides (AMPs) are promising candidates due to their broad-spectrum of
activity and design flexibility; however, experimental screening of peptide activity remains costly and
time-intensive.
Results: We introduce a novel Gated Modulation Protein language model, GMProt, that integrates pretrained
sequence representations with adaptive gating to predict peptide functional activity, including Minimum
Inhibitory Concentration (MIC). The framework enhances activity estimation while enabling reliable candidate
ranking prior to wet-lab validation. The proposed framework demonstrates strong predictive accuracy and
ranking consistency, outperforming eight baseline models. Specifically, it reduces activity estimation error
(RMSE) from 0.524 to 0.493 and increases ranking performance (Kendall's correlation) from 0.462 to 0.487
compared with the current state-of-the-art approach. These findings demonstrate the potential of the proposed
approach to enable data-driven peptide prioritization, offering a scalable route to accelerate AMP discovery.
A-P.09: Integrating Graph Encoders and Protein Language Models for Antibody–Antigen Binding Free Energy
Prediction
Track: Proteins and structural biology
- Tiara Natasha Binte Sayuti, Nanyang Technological University, Singapore
- Vasanth Durvasula, Nanyang Technological University, Singapore
- Aishwarya Anand, Nanyang Technological University, Singapore
-
Jagath Rajapakse, Nanyang Technological University, Singapore
Presentation Overview: Show
Motivation: Antibody–antigen affinity prediction is central to antibody engineering and lead optimization,
where accurate estimation of binding free energy (Delta G) enables ranking and prioritization of variants.
Despite rapid advances in protein language models and graph neural networks, most studies emphasize binary
interaction prediction or build regression models under random splits that risk antigen-level leakage. Such
settings may overestimate generalization, particularly when closely related antigen families appear across
training and test partitions. Robust $\Delta G$ prediction under cluster-aware evaluation remains
underexplored.
{Results: We introduce GEPBind, a multimodal regression framework that integrates ESM-2 sequence
representations with structure-aware graph encoding via a GraphGPS backbone.
%combining GINE and Performer layers.
Antibody–antigen complexes are represented using residue-level C_alpha$ radius graphs with physicochemical
and positional features, and modality-specific predictions are fused through a learned gated residual
calibration mechanism. Under a strict 75% antigen-cluster split with grouped cross-validation and an
independent holdout benchmark, GEPBind achieves state-of-the-art performance, with an RMSE of 1.569 and a
Pearson correlation of 0.575, consistently surpassing sequence-based, structure-based variants, and
established baselines. Ablation studies demonstrate that multimodal fusion, GraphGPS structure and consistent
ESM-2 encoding across chains collectively contribute to robust out-of-cluster generalization.
A-P.10: CAPRINI-M: An AI-curated Cardiac-Specific Atlas of Protein Interactions in Mice
Track: Proteins and structural biology
-
Enio Gjerga, University Hospital Heidelberg, Germany
- Philipp Wiesenbach, University Hospital Heidelberg, Germany
-
Chris-Andris Görner, Faculty of Informatics, Heilbronn University of Applied Sciences​, Germany
-
Ying Zhang, Faculty of Mathematics and Computer Science, Heidelberg University, Germany
- Konstantin Pelz, Technical University of Munich, Germany
- Markus List, Technical University of Munich, Germany
- Christoph Dieterich, University Hospital Heidelberg, Germany
Presentation Overview: Show
Motivation: Protein–protein interactions are central to cardiovascular disease, yet relevant information is
scattered across the literature and general databases, making it difficult to curate. Protein–protein
interactions are fundamental to cardiovascular disease biology, but the corresponding knowledge is dispersed
across the literature and heterogeneous databases, making systematic curation time-consuming. Moreover, many
existing PPI resources may be biased and lack detailed information on structural interaction interfaces or
associated thermodynamic parameters.
Results: We present CAPRINI-M (CArdiac PRotein INteractions In Mice), a web-based tool hosting an AI-curated
atlas of cardiac protein interactions. We mined 9,105 cardiobiology manuscripts and used open-source LLMs
(LLaMA-3.3 70b) to extract 11,189 protein–protein interactions. These were analysed with AlphaFold3 to infer
interaction interfaces and thermodynamic properties linked to complex stability, and to estimate the
likelihood that each protein pair forms a complex. Benchmarking showed CAPRINI-M outperformed general-purpose
PPI resources in hypertrophy-focused edge-topology enrichment. Predicted interaction favourability also agreed
with published experimental evidence, with lower predicted Gibbs free energy associated with experimentally
preferred binding partners. Overall, CAPRINI-M provides more complete, mechanistically informative access to
cardiovascular-disease–relevant protein-protein interactions by integrating literature evidence with
structural interface and stability-aware annotations.
Availability: The CAPRINI-M web application is available at https://shiny.dieterichlab.org/app/caprinim. The
source code used in this study is linked in the Availability section of the manuscripts.
Contact: E.Gjerga@uni-heidelberg.de
Supplementary information: Supplementary data are available at Bioinformatics online
A-P.11: A comparison of machine learning strategies for classification and clustering of transmembrane
protein-protein interactions
Track: Proteins and structural biology
-
Lisa Allmesberger-Riegler, Department of Biotechnology and Food Science, University of Natural
Resources and Life Sciences, Vienna, Austria
-
Fabian Frommelt, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria,
Austria
-
Brianda Lopez Santini, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences,
Austria, Austria
-
Evandro Ferrada, Instituto de Neurociencia, Facultad de Ciencias, Universidad de ValparaıÌso,
ValparaıÌso, Chile, Chile
-
Giulio Superti-Furga, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences,
Austria, Austria
-
Peter Sykacek, Department of Biotechnology and Food Science, University of Natural Resources and Life
Sciences, Vienna, Austria
Presentation Overview: Show
Motivation: Protein-protein interactions (PPIs) are fundamental for cellular signaling, structural integrity,
and the regulation of cellular homeostasis,
yet their diversity remains poorly understood. Aggregated physicochemical features are interpretable but
ignore spatial context, whereas graph
neural networks (GNNs) preserve interaction topology at the cost of transparency. We systematically compare
these representation strategies to test
whether GNN embeddings outperform interpretable physicochemical features for classifying PPI interfaces and to
identify the key physicochemical
features driving TM interactions.
Results: We compare the embeddings for classifying TM PPI interfaces in 2 042 complexes which are
experimentally confirmed, predicted using
AlphaFold (AF) and refined by molecular dynamics (MD) simulations. The aggregated physicochemical features and
GNN embeddings achieved
comparable predictive accuracy, suggesting that preserving interaction topology provides no additional
predictive power. Feature importance
analysis revealed a hierarchy dominated by amino acid composition, particularly hydrophobic and charged
residue proportions, consistent with the
physicochemical constraints of membrane insertion. Energetic and structural features are secondary
contributors. Unsupervised clustering revealed
that aggregated features capture richer biological structure than GNN embeddings. GNN-based clustering
recovered the axis consistent with the
supervised training objective, while aggregated features additionally resolved interface size, energetics, and
contact chemistry as distinct sources of
variation beyond the classification target.
A-P.12: Multi-Modal Protein Representation Learning with CLASP
Track: Proteins and structural biology
-
Nicolas Bolouri, McGill University, Canada
- Joseph Szymborski, McGill University, Canada
- Amin Emad, McGill University, Canada
Presentation Overview: Show
Effectively integrating data modalities pertaining to proteins' amino acid sequences, three-dimensional
structures, and curated text-based descriptions of their biochemical and functional properties can lead to
informative representations capturing different views of proteins. Here, we introduce CLASP, a unified
tri-modal framework that combines the strengths of geometric deep learning, natural large language models
(LLMs), protein language models (pLMs), and contrastive learning to learn informative protein representations
based on their structure, amino acid sequence, and text-based biochemical and functional descriptions. We show
that CLASP enables accurate zero-shot classification and retrieval tasks, such as matching a protein structure
to its sequence or description, outperforming state-of-the-art baselines. CLASP embeddings also exhibit
superior clustering by protein family, and ablation studies confirm that all three modalities contribute
synergistically to performance. Our results highlight the power of integrating structural, sequential, and
textual signals in a single model, establishing CLASP as a general-purpose embedding framework for protein
understanding.
A-P.13: Structure-based virtual screening for TRPM8 modulators
Track: Proteins and structural biology
-
Nivya James, Imperial College London, United Kingdom
- Pedro Ballester, Imperial College London, United Kingdom
Presentation Overview: Show
Since its discovery as the primary cold-sensing ion channel, the Transient Receptor Potential Cation Channel
Melastatin Member 8 (TRPM8) has emerged as an attractive therapeutic target for cancer, pain, respiratory, and
immune disorders. Although the underlying molecular mechanisms of temperature sensing and ligand recognition
remain elusive, high-resolution structures of TRPM8 in agonist- and antagonist-bound conformations have been
determined, enabling structure-based drug discovery approaches such as virtual screening (VS). However,
systematic benchmarking of TRPM8 structural models and docking protocols has not yet been performed. Here, we
establish a reproducible VS benchmark for TRPM8 by evaluating two high-resolution structures representing
agonist-bound (TRPM87WRE) and antagonist-bound (TRPM89B6G) conformations using the open-source docking tools
Smina and rDock with experimentally reported TRPM8 inhibitors and property-matched decoys. In addition, we
investigated how different docking strategies can be combined to improve early hit identification.
rDock achieved higher hit rates and ranked true actives first among their corresponding decoy sets, whereas
Smina showed stronger dependence on the target structure but provided better overall ranking quality across
both conformations. Both docking tools displayed considerable overlap between active and decoy score
distributions, indicating the limited discriminatory power of docking scores alone. When prioritizing a small
subset of top-ranked compounds, integrated screening approaches improved recovery of true actives. Notably,
the hierarchical protocol achieved performance comparable to consensus protocol while requiring substantially
lower computational cost. Together, this work establishes the first systematic VS benchmark for TRPM8 and
highlights the importance of integrated docking workflows for scalable hit discovery.
A-P.14: Beyond Lookup? Quantifying What LLMs Actually Contribute to Enzyme Function Prediction.
Track: Proteins and structural biology
-
Abishek Gnanasekaran, Ghent University Global Campus, South Korea
- Ho-min Park, Ghent University Global Campus, South Korea
- Sunjai Hwang, Ghent University Global Campus, South Korea
- Wesley De Neve, Ghent University Global Campus, South Korea
- Joris Vankerschaver, Ghent University Global Campus, South Korea
Presentation Overview: Show
Large language models are increasingly evaluated as tools for predicting protein function, raising the
question of whether they can complement or replace established sequence-analysis methods. However, it remains
unclear whether these models genuinely reason about biochemical evidence or simply retrieve memorized
associations from their training data. Here we show that large language models perform substantially below
classical baselines on enzyme function prediction and exhibit no detectable reasoning when challenged with
counterfactual evidence. We benchmarked three models against BLASTp and CLEAN on 1,000 enzymes, finding that
the best-performing model achieved 53% accuracy compared with 82% for BLASTp, and that all models collapsed
further on a holdout set of 321 enzymes lacking close database homologs. A signal detection analysis of
counterfactual perturbations — in which biological evidence was swapped between enzymes — yielded sensitivity
indices near zero for all models, with Bayes Factor analysis providing moderate evidence that the models
cannot distinguish genuine evidence changes from neutral reordering. These findings indicate that current
large language models function as stochastic lookup systems rather than biochemical reasoners, and that a
simple sequence-similarity search remains the more reliable tool for enzyme classification.
A-P.15: Bridging Scales: A Multi-Level Graph Neural Network for Protein Function Prediction
Track: Proteins and structural biology
-
Antoine Toffano, LIRMM, Univ. Montpellier, CNRS, Montpellier, France, France
- Pierre Larmande, DIADE, Univ. Montpellier, IRD, CIRAD, Montpellier, France, France
- Jérôme Azé, LIRMM, Univ. Montpellier, CNRS, Montpellier, France, France
Presentation Overview: Show
The exponential growth of protein sequence data has outpaced experimental functional characterization,
resulting in a widening annotation gap. Addressing this requires both the discovery of novel roles and the
refinement of existing Gene Ontology annotations. Current computational approaches to this problem operate
either at the atomic residue scale or the systemic network scale, which prevents a holistic understanding of
protein functional roles. To address this limitation, we propose MS-GNN, a graph neural network framework that
bridges these scales. At the protein level, we construct spatial graphs of amino acids through language model
embeddings and AlphaFold-derived structural contact maps. At the systemic level, these representations are
integrated into a global protein–protein association network. This unified architecture achieves
state-of-the-art performance across all three sub-ontologies. Beyond architectural innovation, we demonstrate
that complementing strict experimental labels with broader, curated functional data drastically improves
performance, showing that annotation sparsity, rather than algorithmic capacity, is the primary bottleneck in
model improvement.
Finally, ablation studies confirm the necessity network-level information, particularly functional features,
highlighting the necessity of multi-scale integration.
A-P.16: Ensembling Structure Model Outputs for Nearly Free Performance Gains in TCR-pMHC Interaction
Prediction
Track: Proteins and structural biology
-
Fredo Guan, Arizona State University, United States
- Heewook Lee, Arizona State University, United States
Presentation Overview: Show
TCR-pMHC interactions play a central role in adaptive immune recognition. Much recent work has focused on
modeling these interactions using deep neural networks, trained to classify a TCR-pMHC pair as binding or
non-binding. However, benchmark studies show that these models fail to generalize to unseen epitopes,
exhibiting near-random performance. Conversely, several recent works leveraging protein structure models show
that simple 0-shot predictors using confidence metrics or Rosetta binding energy can achieve higher
performance on the binding prediction task, motivating further exploration. Here, we assess both zero-shot
structure model outputs and ensembles of confidence metrics and Rosetta binding energy values on three
different TCR-pMHC interaction prediction tasks, with ensembles achieving superior performance over zero-shot
outputs. We further identify strong baseline ensembles for each task and demonstrate their superior
performance on TCR-pMHC binding prediction for unseen epitopes compared to supervised predictive models.
A-P.17: A Unified Computational Framework for the Integration of AI Models in Structure-Based Drug
Design
Track: Proteins and structural biology
-
Luna Pianesi, University of Bielefeld, Germany
- Alexander Schoenhuth, University of bielefeld, Germany
Presentation Overview: Show
The rapid proliferation of sophisticated computational models for drug discovery has created unprecedented
opportunities for innovation, yet the field lacks comprehensive frameworks to systematically integrate these
diverse tools into coherent workflows. Although these models demonstrate remarkable individual capabilities,
researchers are forced to navigate fragmented toolsets requiring extensive computational expertise, limiting
the practical impact of these advances and creating an accessibility problem. This fragmentation is
particularly pronounced at the intersection of molecular generation, molecular docking, and binding affinity
prediction, where no single tool currently bridges all three stages in a unified manner. In this paper, we
present a comprehensive and modular computational drug discovery pipeline that provides the first systematic
framework for integrating diverse state-of-the-art models into an accessible unified drug discovery workflow.
Our pipeline is designed with modularity as a core principle, enabling researchers to swap individual
components as the field evolves without disrupting the broader workflow. The workflow is based on the
integration of state-of-the-art generative and docking models, with a special focus on ensuring the synthetic
accessibility and real world scenario plausibility of the proposed molecules. By lowering the barrier to entry
for computational drug discovery, we aim to empower a broader community of researchers to leverage these
powerful tools in pursuit of novel therapeutics.
A-P.18: BepiCon: A Geometric Deep Learning Framework for Conformational B Cell Epitope Prediction
Track: Proteins and structural biology
-
Bünyamin Şen, Dept. of Bioinformatics, Graduate School of Health Sciences, Hacettepe University,
Ankara, Turkey, Turkey
-
Tunca Doğan, Dept. of Bioinformatics, Graduate School of Health Sciences, Hacettepe University, Ankara,
Turkey, Turkey
Presentation Overview: Show
Accurate and reliable prediction of B cell epitopes holds critical importance in immunology and vaccine
development. While traditional experimental methods offer high accuracy in identifying epitope regions, they
are often laborious, time-consuming, and costly. Therefore, attempts are made to increase the efficiency of
experimental characterization processes by using computational approaches. Since approximately 90\% of
epitopes are conformational, the prediction processes must account for three-dimensional protein structures
and the geometric details of antigen-antibody interactions. In response to these requirements, our study
introduces BepiCon, a two-stage geometric deep learning framework that models antigen proteins as graph
structures, incorporating structural and physicochemical properties and protein language model embeddings to
predict epitope regions on antigen proteins. In the first stage, the model was trained using a graph
contrastive learning approach to learn high-quality representations of epitope and non-epitope residues. In
the second stage, the pre-trained model was fine-tuned using supervised learning to perform conformational
epitope prediction. The developed framework has demonstrated effective and generalizable performance when
applied to both experimentally determined protein structures and predicted structures. Comparative analysis
revealed that our approach distinguishes itself from existing B cell epitope prediction methods by exhibiting
a lower false-positive rate and generating more reliable predictions. Our work contributes significantly to
scientific research and therapeutic design processes by showcasing the advantages of geometric deep-learning
approaches in B-cell epitope prediction.
A-P.19: Leakage-controlled evaluation of single-to-multi mutation generalization in fungal azole resistance
prediction
Track: Proteins and structural biology
-
Yuan Fei, Shanghai Jiao Tong University School of Medicine, China
- Thanh Thien Le, VinUniversity, Viet Nam
- Io Hong Cheong, Shanghai Jiao Tong University School of Medicine, China
- Zisis Kozlakidis, World Health Organization, France
- Jiaqi Yin, Northwestern Polytechnical University, China
- Yin Wu, Shanghai Jiao Tong University School of Medicine, China
- Xiaoqi Zheng, Shanghai Jiao Tong University School of Medicine, China
- Yang Yang, Shanghai Jiao Tong University School of Medicine, China
Presentation Overview: Show
Motivation: Clinical azole resistance in fungi often involves combinations of mutations, whereas labeled
higher-order mutants (k ≥ 2 substitutions) remain scarce. This creates a realistic setting in which models
trained on single mutants for azole resistance prediction must generalize to higher-order mutants. A central
question is whether apparent generalization reflects genuine transfer or is driven by overlap between
substitutions seen in single mutants and those present in held-out higher-order mutants.
Methods: Using FungAMR, we evaluated PLM-based and structured feature representations with aggregation
baselines, tabular machine-learning models, an MLP, and gated Two-Tower models under three protocols:
single-to-multi evaluation, a leakage-controlled variant that removes constituent substitutions overlapping
with held-out higher-order mutants, and cross-species transfer from C. albicans to non-C. albicans fungi.
Results: We established a single-to-multi evaluation protocol in which models are trained on single mutants
and evaluated on 51 held-out higher-order mutants. Ranking performance was strong, with the best configuration
reaching PR-AUC 0.944. To test whether this performance was inflated by constituent overlap, we introduced a
leakage-controlled protocol that removed single-mutant substitutions appearing in held-out higher-order
mutants from the training-validation pool. Representative end-to-end models remained stable under this
control, with PR-AUC shifts ≤ 0.01, whereas aggregation baselines deteriorated markedly. Cross-species
evaluation showed measurable zero-shot transfer from C. albicans to non-C. albicans fungi, with pooled PR-AUC
0.711. Together, these results suggest that leakage-controlled evaluation helps distinguish genuine
generalization to higher-order mutants from constituent-overlap shortcuts in fungal resistance prediction.
Availability and implementation: Source code is available at https://github.com/asunafy/fungamr-paper.
A-P.20: PUMA: Discovery of Protein Units via Mutation-Aware Merging
Track: Proteins and structural biology
-
Burak Suyunu, Computer Engineering, BoÄŸaziçi University, Bebek, 34342, İstanbul, Türkiye,
Turkey
-
Özdeniz Dolu, Computer Engineering, BoÄŸaziçi University, Bebek, 34342, İstanbul, Türkiye, Turkey
-
Ibukunoluwa Abigail Olaosebikan, C. Eugene Bennett Department of Chemistry, West Virginia University,
Morgantown, West Virginia 26505, United States, United States
-
Hacer Karatas Bristow, C. Eugene Bennett Department of Chemistry, West Virginia University, Morgantown,
West Virginia 26505, United States, United States
-
Arzucan Ozgur, Computer Engineering, BoÄŸaziçi University, Bebek, 34342, İstanbul, Türkiye, Turkey
Presentation Overview: Show
Motivation: Proteins are the essential drivers of biological processes. At the molecular level, they are
chains of amino acids that can be viewed through a linguistic lens where the twenty standard residues serve as
an alphabet combining to form a complex language, referred to as the language of life. To understand this
language, we must first identify its fundamental units. Analogous to words, these units are hypothesized to
represent an intermediate layer between single residues and larger domains. Crucially, just as protein
diversity arises from evolution, these units should inherently reflect evolutionary relationships. We
introduce PUMA (Protein Units via Mutation-Aware Merging) to discover these meaningful units. PUMA employs an
iterative merging algorithm guided by substitution matrices to identify protein units and organize them into
families linked by plausible mutations. This process creates a genealogy where units and their mutational
variants coexist, simultaneously producing a unit vocabulary and a hierarchy connecting them.
Results: We validate PUMA's biological relevance through three key findings. First, mutations occurring within
a PUMA family are significantly more likely to be clinically benign than pathogenic. Second, by leveraging its
genealogy, PUMA surpasses frequency-based baselines in mapping protein units to functional annotations.
Finally, case studies demonstrate that PUMA families preserve essential physicochemical properties and
functional motifs. By explicitly linking mutational variants, PUMA provides an evolutionarily grounded and
interpretable vocabulary for deciphering the language of life.
Availability and implementation: The source code is available at https://github.com/boun-tabi-lifelu/PUMA.
A-P.21: SpeciefAI: Multi-species mRNA-level Antibody Framework Generation using Transformers
Track: Proteins and structural biology
-
Dominik Grabarczyk, University of Edinburgh, United Kingdom
- Mikołaj Kocikowski, Independent Researcher, Poland
- Maciej Parys, University of Edinburgh, United Kingdom
- Shay Cohen, University of Edinburgh, United Kingdom
- Javier Alfaro, University of Calgary, Canada
Presentation Overview: Show
Motivation: Encoding antibodies (Abs) and nanobodies (Nbs) as mRNA enables in vivo production of therapeutic
proteins. However, this approach requires meeting two species-dependent requirements: the mRNA encoding must
support efficient expression in the host species, and the encoded protein sequence must resemble the natural
Ab repertoire of the recipient species to minimize immunogenicity. These requirements motivate
species-conditioned generative models for joint mRNA and protein design.
Results: We propose SpeciefAI a transformer-based model for multi-species Ab and Nb species
sequence-harmonisation by generation of novel Framework Regions (FRs) tailored to input
Complementarity-Determining Regions (CDRs). Our model works directly in the mRNA space and learns the
correspondence between FRs and CDRs in six species. The model is capable of generating sequences with a highly
similar distribution to natural sequences and a mean absolute difference in codon adaptation index (CAI) of
0.013 and 0.033 for humans and dogs respectively. We show that the generated human sequences are highly human
(0.95 T20 score) and canine sequences highly canine (0.95 cT20 score). We furthermore demonstrate that we can
generate diverse candidate sequences using our method.
Availability and Implementation: Source code is available on https://github.com/Dominko/SpeciefAI. OAS and
COGNANO data are publicly available on https://opig.stats.ox.ac.uk/webapps/oas/ and
https://cognanous.com/datasets/vhh-corpus (preprocessed versions available upon request). Canine data is
available on https://zenodo.org/records/18301526
A-P.22: WACA-DTA: Water-Aware Geometric Biases for Structure-Conditioned Drug-Target Affinity
Prediction
Track: Proteins and structural biology
-
Kehan Huang, China Pharmaceutical University, China
- Chang Li, Tsinghua University, China
- Jing Ji, China Pharmaceutical University, China
- Zimo Tang, China Pharmaceutical University, China
Presentation Overview: Show
Drug-target affinity (DTA) prediction is central to computational drug discovery, yet many structure-aware
models still weakly constrain geometric
locality and solvent mediation at the binding interface. We present WACA-DTA, a structure-conditioned affinity
model that reformulates interface
matching as pose-conditioned atom-residue cross-attention with factorized interaction logits. Each
cross-interface score is decomposed into
direct, geometric, and hydration-mediated terms, with geometry and hydration injected as additive logit-level
priors rather than post hoc feature
concatenations. Under a fixed pair-pose protocol and identical structural preprocessing, WACA-DTA improves
over a matched pocket-aware
baseline on Davis and KIBA across drug, target, and pair affinity cold-start splits, while controlled
ablations across Davis, KIBA, and PDBbind
show that the gains are most consistent when pair-specific geometric and hydration cues are injected directly
into the interaction logits. Our
claims are restricted to input-matched comparisons and within-dataset evaluation, and the hydration
interpretation is cross-checked on PDBbind
complexes with retained crystallographic waters. Code, split definitions, and preprocessing manifests are
available at https://github.com/khan1
14514/WACA-DTA2.
A-P.23: Swiss-PO. Advancing Cancer Mutation and Structural analysis for Precision Oncology with the Latest
Release
Track: Proteins and structural biology
-
Fanny Krebs, University of Lausanne, Ludwig Institute for Cancer Research Lausanne, SIB Swiss
Institute of Bioinformatics, Switzerland
- Olivier Michielin, HUG, SIB Swiss Institute of Bioinformatics, Switzerland
-
Vincent Zoete, SIB Swiss Institute of Bioinformatics, University of Lausanne, Ludwig Institute for Cancer
Research Lausanne, Switzerland
Presentation Overview: Show
The rapid expansion of precision oncology has led to a substantial increase in identified oncodriver genes and
associated variants. This growth has also increased the number of mutations with unknown functional impact,
highlighting the need for molecular modeling tools to support variant interpretation. Swiss-PO has been
expanded and redesigned to integrate large-scale oncogenomic data with structure- and sequence-based
analytical tools. The platform combines oncodriver gene data with multiple sequence alignments (MSAs) for
conservation analysis across orthologs and gene families, as well as visualization of experimental and
predicted three-dimensional (3D) protein structures to assess the potential effects of mutations. Additional
modules include a BRAF kinase mutation classification model and a dedicated page for ligands targeting
proteins, alongside integrated links to external databases to enable multidimensional analyses.
The updated Swiss-PO platform integrates data for nearly 1,500 oncodriver genes, covering more than 3 million
mutations and post-translational modification annotations, over 26,000 experimental and predicted 3D protein
structures, more than 4,000 MSAs, and information on over 200,000 protein ligands. The platform also includes
a BRAF kinase mutation class predictor to support therapeutic decision-making. With these enhancements,
Swiss-PO provides a comprehensive resource for oncologists, bioinformaticians, and molecular biologists
involved in interpreting cancer-associated mutations and advancing precision medicine. The website is
available at https://swiss-po.ch/
A-P.25: FusionPath: Gene fusion pathogenicity prediction using protein structural data and contextual protein
embeddings
Track: Proteins and structural biology
-
Nadine Sina Kurz, University Medical Center Göttingen, Germany
- Irem Berna Güven, University Medical Center Göttingen, Germany
- Tim Beissbarth, University Medical Center Göttingen, Germany
- Jürgen Dönitz, University Medical Center Göttingen, Germany
Presentation Overview: Show
Accurate prediction of gene fusion pathogenicity is critical for understanding oncogenic mechanisms and
advancing precision oncology. While existing computational methods provide valuable insights, their
performance remains limited by incomplete integration of multi-scale biological features and insufficient
model interpretability for clinical translation. We present FusionPath, a novel deep learning framework for
gene fusion pathogenicity prediction that addresses these limitations through multimodal integration of
complementary biological data. FusionPath uniquely integrates embeddings from multiple pretrained protein
language models, including FusON-pLM and ProtBERT, alongside retained protein domains and Gene Ontology (GO)
functional annotations. The model was trained and validated on a rigorously curated dataset of more than
70,000 gene fusions derived from FusionPDB, ChimerDB4.0, and 27 RNA-seq datasets of normal tissues. FusionPath
significantly outperformed state-of-the-art methods and provides interpretable insights through SHAP analysis,
revealing cancer-type-specific pathogenicity patterns and identifying protein kinase domains as key
determinants of oncogenic potential. By synergistically leveraging sequence, structural, and functional
information with explicit modeling of wild-type sequence context, FusionPath yields biologically grounded
pathogenicity scores with mechanistic insights.
A-P.26: Elucidating the Structural Dynamics and Binding Mechanism of ProQ-raiZ Using an ML based Deep-TDA
Approach for Enhanced Sampling Simulations: Implications for Novel Antimicrobial Targeting
Track: Proteins and structural biology
-
Shilpi Singh, Indian Institute of Technology Delhi, India
- Nisha Kumari, Indian Institute of Technology Delhi, India
- Tarak Karmakar, Indian Institute of Technology Delhi, India
- Tanmay Dutta, Indian Institute of Technology Delhi, India
Presentation Overview: Show
ProQ, a bacterial RNA chaperone, plays a crucial function in post-transcriptional gene regulation which makes
it a potential target for combating antibiotic resistance. The N-terminal domain (NTD) is recognised for its
role in RNA binding; however, the structural dynamics and energetic landscape of its interaction with small
RNAs that play roles in virulence, such as raiZ, remain poorly explored. This research uses an advanced
computational framework to delineate the ProQ-raiZ complex dynamics, establishing a structural foundation for
novel therapeutic strategies.
We employed molecular dynamics simulations, combining spontaneous binding analyses, to identify the
protein-RNA interface. We utilised an advanced enhanced sampling technique like On-the-fly
Probability-Enhanced Sampling (OPES) to get a free energy landscape for the binding mechanism. We utilised
Deep- Targeted Discriminant Analysis (Deep-TDA), a machine learning methodology to find non-linear features
inside the conformational space and to formulate appropriate Collective Variables (CVs) for more effective
sampling.
Our findings prove that raiZ maintains a semi-folded conformation upon interacting with the concave surface of
ProQ, reinforced by a network of electrostatic contacts throughout the pocket in NTD. In silico mutagenesis
revealed key residues critical for complex stability; their alteration results in raiZ dissociation,
suggesting these locations as high-priority ""hotspots"" for small-molecule inhibition. This study integrates
machine learning with physics-based Molecular dynamics simulations to provide a high-resolution picture of
ProQ dynamics. These insights provide a framework for engineering compounds that can interfere with RNA-based
regulatory pathways, exemplifying a strong computational approach for novel drug discovery.
A-P.27: Assessing Long-distance Epistatic Interactions in Enzymes
Track: Proteins and structural biology
-
Mykolas Malevicius, Institute of Biomedical Informatics, Technical University, Graz, Austria
-
Jasmin Zuson, Institue of Molecular Biotechnology, Technical University, Graz, Austria
-
Gerhard G. Thallinger, Institute of Biomedical Informatics, Technical University, Graz, Austria
-
Leila Taher, CUBiDA, Medical Center for Information and Communication Technology, Universitätsklinikum
Erlangen, Germany
Presentation Overview: Show
Deep mutational scanning is a technique that uses next-generation sequencing to study the functional impact of
thousands of mutations in an enzyme in a single experiment. Due to epistasis, the functional impact of
multiple mutations is non-additive, their effects must be examined together. To capture such effects, the
entire molecule must be sequenced as a single read, which necessitates long-read sequencing approaches such as
that provided by Oxford Nanopore Technologies (ONT).
While ONT enables long-read sequencing, its inherently high error rate of approximately 2% hinders exact
variant identification. Additionally, multiple PCR rounds during library preparation can introduce chimeras,
artificial fusions of distinct DNA molecules, that further complicate data analysis. To combat these
discrepancies, we designed and implemented a computational workflow that leverages unique molecular
identifiers to construct consensus sequences (CSs) and correct errors.
We analysed two datasets, comprising of ~5M and ~10M reads, and found that chimeras can be prominent, reaching
33.4% in one of the datasets. Furthermore, data pre-processing resulted in read losses of 39% and 54%, coupled
with per-CS error rates of 9.32, 0.38, 0.09 and 0.04 in sequences constructed from 2, 3, 4 and 5 reads,
respectively. Therefore, at least five reads are necessary to build a reliable CS. Moreover, more than half of
the molecules were only sequenced once, increasing the number of uninformative sequences. Currently, we are
designing a deep learning approach to classify the substructures of the amplicons at signal-level, which will
enable categorisation in real time.
A-P.28: Computational Design of Anti-VEGF Binders: From Scaffold Optimization to De Novo Design
Track: Proteins and structural biology
-
Georg Kuenze, Institute for Drug Discovery, Leipzig University, Germany
- Abibe Useini, Institute for Drug Discovery, Leipzig University, Germany
-
Carsten Geist, Fraunhofer Institute for Cell Therapy and Immunology - IZI, Leipzig, Germany
-
Stefan Kalkhof, Fraunhofer Institute for Cell Therapy and Immunology - IZI, Leipzig, Germany
Presentation Overview: Show
Anti-VEGF inhibitors are central to the treatment of angiogenesis-driven diseases such as cancer and
age-related macular degeneration; however, current antibody-based therapies are limited by their size,
production cost, and delivery constraints. Miniproteins represent a promising alternative, combining high
target specificity with enhanced stability and manufacturability. Here, we report the computational design and
experimental validation of anti-VEGF miniprotein binders across two development stages.
Our methodology integrates (i) structure-guided redesign of the Z12 miniprotein scaffold using ProteinMPNN to
improve stability and VEGF inhibition, (ii) de novo binder generation using BindCraft to explore novel
structural solutions, and (iii) experimental validation using surface plasmon resonance and a cell-based VEGF
receptor bioassay.
Optimized Z12-derived variants achieved an IC50 of 0.15 µM, corresponding to an approximately 10-fold
improvement in inhibitory potency compared to the parent scaffold, while also exhibiting enhanced thermal
stability with Tm increases of 10-17°C. In parallel, BindCraft-designed de novo binders reached sub-micromolar
binding affinities and demonstrated strong functional inhibition, with top candidates reducing VEGF activity
to ~4% at 1 µM, indicative of potent antagonistic activity. X-ray crystallographic analysis of the
top-performing binders confirmed close agreement between the computational models and experimentally
determined structures.
Overall, this work demonstrates that combining scaffold optimization with de novo design enables the rapid
development of potent and stable anti-VEGF miniproteins for anti-angiogenic applications. These findings
establish a foundation for next-generation biologics targeting angiogenesis-related diseases and highlight the
potential of AI-driven protein design in therapeutic discovery.
A-P.29: CCD2MD: A suite of Packages for Preparing Co-Folded Outputs for Molecular Dynamics
Simulations
Track: Proteins and structural biology
-
Katarina Blow, University of Warwick, United Kingdom
- Matyas Parrag, University of Warwick, United Kingdom
- Phillip Stansfeld, University of Warwick, United Kingdom
Presentation Overview: Show
Protein/lipid interactions play a crucial role in the stability and function of membrane proteins. While
experimental approaches to characterise these interactions in a native-like membrane environment can be
challenging, computational techniques offer a powerful alternative for identifying and analysing potential
binding sites. Recent advances in co-folding methods (such as AlphaFold[1]) now enable the prediction of holo
protein structures, capturing conformational changes that may occur upon lipid binding and thereby improving
the accuracy of binding site characterisation. However, the outputs from these methods often require
post-processing to ensure compatibility with widely used molecular dynamics force fields. Here, we present
CCD2MD[2], a modular toolkit designed to convert co-folding outputs into simulation-ready systems. CCD2MD
supports both atomistic and coarse-grained representations, with optional membrane embedding facilitated via
MemPrO[3]. The modular design of CCD2MD allows for straightforward adaptation to other co-folded biomolecular
assemblies, incorporating complexes with nucleic acids, small molecules, carbohydrates, or metal ions, thereby
enabling a variety of simulation setups across multiple scales. We also discuss recent updates and
improvements to the code.
A-P.30: Computational Mutational Scanning of DHFR for Mutation-Dependent Inhibitor Preference
Track: Proteins and structural biology
-
Busra Tayhan, Sabanci University, Turkey
- Ebru Cetin, Sabanci University, Turkey
- Tandac Furkan Guclu, Istinye University, Turkey
-
Muhammed Sadik Yildiz, University of Texas Southwestern Medical Center, United States
- Erdal Toprak, University of Texas Southwestern Medical Center, United States
- Ali Rana Atilgan, Sabanci University, Turkey
- Canan Atilgan, Sabanci University, Turkey
Presentation Overview: Show
Antibiotic resistance is a public health challenge driven by mutations that enable bacteria to maintain
function while escaping drug pressure. One enzyme extensively studied in this context is dihydrofolate
reductase (DHFR), an essential catalyst in folate metabolism and nucleotide biosynthesis. This protein is
suitable for studying antibiotic resistance because resistance often arises from single amino-acid
substitutions that preserve function while altering inhibitor binding. In this work, we investigated the
mutational landscape of E. coli DHFR in the presence of competitive inhibitors, trimethoprim(TMP) and its
derivative 4′-deoxytrimethoprim(4′-DTMP). Because TMP and 4′-DTMP are closely related chemically,
comparing them allows us to assess whether a structural modification can reshape mutation-dependent inhibitor
preference and alter high-impact resistance trajectories. To characterize mutation-dependent energetic
effects, we generated a deep mutational scanning library. Large-scale free-energy perturbation(FEP)
simulations were performed for apo, TMP-bound, and 4′-DTMP-bound DHFR to estimate mutational free energies
and relative binding preferences. Single amino-acid substitutions were examined across DHFR positions with
triplicate FEP calculations. Although TMP and 4′-DTMP showed similar behavior across much of the scanned
landscape, the strongest differences were concentrated at important DHFR regions, including cryptic-site and
allosteric regions, demonstrating that even a subtle chemical modification can reshape inhibitor preference at
sites most relevant to resistance dynamics and function. These findings provide a structural-energetic map of
mutation-dependent inhibitor preference and highlight the potential of alchemical calculations for
mutation-resolved studies of antibiotic resistance. Results were interpreted alongside experimental relative
fitness data, with explicit incorporation of uncertainty from both computational and experimental sources.
A-P.31: An efficient integer programming model for RNA secondary structure prediction
Track: Proteins and structural biology
-
Olga Karelkina, Systems Research Institute Polish Academy of Sciences, Poland
Presentation Overview: Show
RNA secondary structure prediction problem is commonly addressed through dynamic programming, context-free
grammar, and, more recently, machine learning techniques. We present an alternative integer programming (IP)
model, based on a classical hydrodynamics hypothesis. As an energy-directed method, it relies on Turner's
nearest-neighbor rules to evaluate the free energy of candidate structures. To predict secondary structure,
the model explores all feasible configurations of two-dimensional structural motifs, defined through binary
variables and linear constraints. Furthermore, the proposed IP model can be extended with additional
structural constraints to generate suboptimal solutions.
While the IP approach offers high modeling flexibility, it is time-consuming, which can limit its practical
applicability. We introduce a compact formulation with a dedicated set of constraints to model multibranch
loops, which significantly reduces complexity and computation time, enabling analysis of longer and more
challenging sequences.
The model is benchmarked on the Archive II dataset and compared against leading dynamic programming methods.
The results demonstrate competitive overall accuracy, and, for certain sequences, the model produces solutions
that more closely match native structures than those obtained by traditional methods.
A-P.32: Novel universal domain-centric method for protein classification
Track: Proteins and structural biology
-
Shakiba Fadaei, Student, university of Lausanne, Switzerland
- Fanny Krebs, University of Lausanne / SIB, Switzerland
- Vincent Zoete, University of Lausanne / SIB, Switzerland
Presentation Overview: Show
Human protein kinases constitute a large superfamily of about 500 genes, historically classified into
subfamilies based on phylogenetic relationship. However, many kinases remain unclassified. Phylogeny is
typically based on multiple sequence alignments, and neglects the physico-chemical properties of residues at
each position of the sequence. By incorporating these properties, we can gain deeper insights beyond basic
alignments. Here we use, for the first time, a detailed physico-chemical description of kinases to identify
class-specific structural regions, supporting an unsupervised classification method capable of classifying
previously unlabeled kinases. This novel approach aligns with existing phylogeny-based classifications while
offering refinements and enhanced accuracy. Ultimately, we use machine learning techniques to classify
unlabeled kinases, validated by analyzing class-specific structural regions. This new classification approach
goes beyond current rankings and can be applied to any type of protein, such as immunoglobulins and G
protein-coupled receptors.
A-P.33: BrightDB : Brightness Protein Database with fast accession to millions of proteins and an interactive
dashboard
Track: Proteins and structural biology
-
Damien Legros, VIB.AI/KULeuven, Belgium
- Joana Pereira, VIB.AI/KULeuven, Belgium
Presentation Overview: Show
Storing data efficiently is a fundamental requirement for performing large-scale bioinformatics analysis.
However, as the genomic and proteomic data grows exponentially, traditional relational databases often
struggle with multi-modal information and random access. To address these persistent bottlenecks, we present
BrightDB, an innovative local aggregator database designed to centralize and manage data from multiple public
databases (such as UniProt and InterPro), results from models, and experiments findings by leveraging the
cutting-edge capabilities of modern data lakehouse architecture.
BrightDB bridges the gap between the flexibility of unstructured data lakes and the management of features in
traditional databases. By utilizing high-performance and columnar-based novel Python libraries, BrightDB
centralizes multi-modal biological data into a unified, optimized file format. This columnar approach allows
for substantial data compression and significant speed improvements in data retrieval, directly facilitating
the generation of datasets for model training, statistical analysis, and complex visualization.
Beyond its backend efficiency, BrightDB also offers a real-time interactive dashboard. This local interface
allows users to perform near-instantaneous searches across millions of entries, providing immediate access to
the available information for each protein. By removing the technical difficulties between data storage and
visualization, BrightDB significantly reduces the time-intensive data management. Ultimately, this
architecture provides a scalable, future-proof local aggregator database solution for the proteomics
community, ensuring that all types of data modalities from legacy formats to emerging ones remain accessible
inside a unified storage environment.
A-P.34: Reference-Free Ranking Method for RNA 3D Models
Track: Proteins and structural biology
-
Jan Pielesiak, Poznan University of Technology, Poland
- Maciej Antczak, Poznan University of Technology, Polish Academy of Sciences, Poland
- Marta Szachniuk, Poznan University of Technology, Polish Academy of Sciences, Poland
- Tomasz Zok, Poznan University of Technology, Poland
Presentation Overview: Show
In current times, researchers have started to notice the importance of understanding the structure and
function of RNA. This prompts the growth of the number of attempts in RNA 3D structure prediction. As the
knowledge of RNA molecules grows, we can use the advancements made in protein structure prediction to improve
our prognoses in the RNA field; however, it faces a challenge – produced models need to have their quality
assessed. Multiple or sometimes even thousands of models can be generated from a single input. The traditional
method for determining the quality of the model is based on energy terms calculation with force fields or
coarse-grained statistical potentials. The main issue with such an approach is that the energy landscape
usually contains many local minima, which leads to inconclusive results.
Therefore, we propose a different approach for ranking multiple 3D models of the same RNA sequence. The basis
for this algorithm is the analysis of the base pairs and stacking interactions within them. A consensus
secondary structure is built from the extracted data, and each model has its interaction network ranked
against the aforementioned consensus to provide a final ranking.
Our method has been benchmarked on datasets used in RNA 3D modeling to verify its quality in comparison to the
state-of-the-art energy-based evaluations. The entire system has been made publicly available at
https://rnative.cs.put.poznan.pl/, and published in Bioinformatics
(https://doi.org/10.1093/bioinformatics/btaf601)
A-P.35: TmProt 1.0: ML-based Tool for Protein Melting Temperature Prediction with Cross-Method Validation on
Biophysical Data
Track: Proteins and structural biology
-
Karen Pailozian, Masaryk univevrsity, Czechia
Presentation Overview: Show
Protein melting temperature (Tm) prediction accelerates the discovery of thermostable enzymes crucial for
industrial biotechnology, where proteins must endure harsh reaction conditions. Experimental determination of
Tm remains labour-intensive and varies across techniques, motivating the development of in silico predictors.
Recent advances in high-throughput mass spectrometry have led to large-scale proteomics-based Tm datasets
enabling effective training of machine learning models. However, the generalisability of such tools across
diverse proteomics- and biophysics-based datasets remains an open question.
This study aims to (i) assemble comprehensive Tm datasets including high-quality biophysics sources for
independent evaluation, (ii) evaluate generalisability of state-of-the-art approaches based on deep learning,
ESM-2 sequence embeddings, and parameter-efficient low-rank adaptation (LoRA), and (iii) explore ESM-3
structural embeddings as a counterpart to sequence-based approaches.
We assembled the ProMelt dataset (45,441 proteins) from Meltome Atlas and ProThermDB and trained baseline and
advanced embedding-based predictors. Models were evaluated on five independent biophysics-based datasets
derived from BRENDA, FireProtDB, and literature.
Our analysis revealed substantial inconsistencies in reported Tm values between proteomics- and
biophysics-based measurements, highlighting the need for generalizable predictors. Fine-tuned embedding-based
models showed competitive performance compared to DeepSTABp, TemBERTure, and SaProt and achieved superior
performance in binary classification of thermostable proteins (Tm ≥ 60 °C): ESM2-LoRA achieved AUC = 0.75,
while ESM3-MLP achieved AUC = 0.77.
The best-performing model is deployed as TmProt, a user-friendly web server on Hugging Face.
Together, these results show that LoRA-adapted ESM-2 and ESM-3 structural embeddings improve thermostability
prediction and emphasize the importance of rigorous evaluation across diverse experimental methodologies.
A-P.36: Accelerating Multiple Myeloma Diagnosis with Proteomics & Machine Learning
Track: Proteins and structural biology
-
Anna Melidi, Bispebjerg Hospital, Denmark
Presentation Overview: Show
Importance: Multiple myeloma (MM), a blood cancer arising from malignant plasma cells in the bone marrow, is
complex and time-consuming to diagnose, requiring several laboratory tests, which can delay treatment
increasing progression risk.
Objective: To develop a diagnostic test based on untargeted mass spectrometry(MS)-based proteomics and machine
learning (ML) to diagnose MM already at the first point of screening.
Design, Setting, and Participants: This is a prospective cohort study involving 2,501 samples (serum,
peripheral plasma, bone marrow plasma) obtained from patients diagnosed with MM, monoclonal gammopathy of
undetermined significance (MGUS), smoldering MM (SMM), patients referred for M-component screening, and
healthy controls.
Methods: Proteomes were quantified by LC-MS and analyzed with supervised ML to identify diagnostic signatures.
Performance was evaluated by cross-validation and external validation, with correlation to disease severity,
survival, and progression assessed.
Results: The model achieved AUCs of 0.85 (CV-test), 0.91 (national validation cohort), and 0.79 (international
validation cohort). The predicted disease score showed a continuous dis-tribution correlating with disease
severity from controls to MGUS to SMM to MM. Correla-tion between the LC-MS-derived estimate of monoclonality
and conventional M-spike quan-tification was 0.78 (p=2×10â»â¶Â¹). Classification performance was 95% in
healthy peripheral plasma and 95% (peripheral) and 83% (bone marrow) in newly diagnosed MM. We found no
significant survival differences between upper and lower quartiles of the disease score in newly diagnosed
MM.
Conclusions and Relevance: Combining proteomics with ML allows diagnostic support and severity assessment in
MM and potentially other medical conditions.
A-P.37: Alpha&ESMhFolds: An updated database for the comparison and functional annotation of predicted
structural models for the human proteome
Track: Proteins and structural biology
-
Matteo Manfredi, Biocomputing Group, University of Bologna, Italy
- Gabriele Vazzana, Biocomputing Group - University of Bologna, Italy
- Castrense Savojardo, Biocomputing Group, University of Bologna, Italy
- Pier Luigi Martelli, Biocomputing Group, University of Bologna, Italy
- Rita Casadio, University of Bologna, Italy
Presentation Overview: Show
The adoption of predicted models of protein structures generated by methods such as AlphaFold and ESMFold is
ever-increasing. We release an updated version of a public database available at
https://alpha-esmhfolds.biocomp.unibo.it/, storing pairs of models for 48,815 human proteins enriched with
functional characterization.
This release, synchronized with the latest updates of resources like UniProt, PDB, AlphaFold DB, and Pfam,
introduces new functionalities. We extract for each protein Pfam annotations and known pathogenic variants.
Both are mapped on the predicted structures and can be visualized from the web server. Moreover, we complement
the structural superimposition of AlphaFold2 and ESMFold models with an external validation performed by three
state-of-the-art Quality Assessment tools, providing a consensus to suggest the best model for each
protein.
Exploiting those functionalities, we perform large-scale analysis to identify two interesting trends. First,
both AlphaFold2 and ESMFold are consistently better at predicting regions of the proteins covered by Pfam
entries. Not only do both methods exhibit higher average pLDDT scores in those regions when compared to the
average of the full models, but their agreement, measured by the TM-score upon structural superimposition,
also increases from 0.58 to 0.88.
Second, focusing on the subset of proteins for which the two methods produce diverging models, we observe that
the consensus of external QA tools assigns better scores to AlphaFold for only 51% of the cases, supporting
the usefulness of having different models available to exploit the strengths of each method, particularly when
the models differ.
A-P.38: Exploring the Proteomics Dark Matter: Non-Canonical Peptide Identification and the Quantification of
all Precursor Ions
Track: Proteins and structural biology
-
Dominik Lux, Ruhr University of Bochum / Medical Proteome Center, Germany
- Svitlana Rozanova, Ruhr University of Bochum / Medical Proteome Center, Germany
-
Britt Mollenhauer, University Medical Center Goettingen / Department of Neurology / Kassel /
Paracelsus-Elena Klinik, Germany
- Katalin Barkovits, Ruhr University of Bochum / Medical Proteome Center, Germany
- Julian Uszkoreit, Ruhr University of Bochum / Medical Bioinformatics, Germany
- Katrin Marcus, Ruhr University of Bochum / Medical Proteome Center, Germany
- Martin Eisenacher, Ruhr University of Bochum / Medical Proteome Center, Germany
Presentation Overview: Show
Most analysis workflows for LCMS proteomics data captured with data-dependent acquisition use search engine
identification to match MS2 spectra against known protein databases. These identification-first approaches
then utilize precursor ions of identified MS2 spectra to retrieve quantitative values for downstream analysis.
While these workflows are well-established, most precursor ions are not or cannot be identified and are
systematically not reported. We define this unused set as the proteomics dark matter. In this work, we
demonstrate that some unidentified precursors can be identified and all of them quantified regardless of their
identification status. To address this, we developed two workflows: ProtGraph and UnbeQuant. ProtGraph
utilizes protein graphs considering protein variations, mutations and specific cleavage points to identify
non-canonical peptides. Complementary, UnbeQuant enables the quantification of precursor ions regardless of
identification status, allowing for quantitative comparisons across measurements. Using a cerebrospinal fluid
dataset across two groups and three time points, we demonstrate that downstream analysis with the proteomics
dark matter is possible by highlighting un-/identified and identified non-canonical precursor ions between the
groups and time points. Further, by using un-/identified precursor ions, we show that unconsidered
post-translational modifications can be discovered. Finally, we demonstrate identification ratios across
different sample types including blood, cell culture, or stool to illustrate the potential of the proteomics
dark matter in different datasets. Ultimately, we show that identification-first approaches may only report
limited results of data at hand and encourage researchers to explore the proteomics dark matter for a more
comprehensive analysis of their data.
A-P.39: Characterization of SLC25A38 as a Mitochondrial Pyridoxal 5′-Phosphate Transporter
Track: Proteins and structural biology
-
Miriana Quaranta, Sapienza University of Rome, Italy
- Stefano Pascarella, Sapienza University of Rome, Italy
Presentation Overview: Show
Pyridoxal 5′-phosphate (PLP), the active form of vitamin B6, is an essential cofactor for mitochondrial
enzymes involved in amino acid metabolism and heme biosynthesis. Despite its central role in cellular
homeostasis, the mechanisms regulating PLP import into mitochondria remain poorly understood. The human
mitochondrial carrier SLC25A38, previously annotated as a glycine transporter, has been proposed to mediate
PLP transport. Indeed, silencing of SLC25A38 selectively reduces mitochondrial PLP levels (Pena et al., 2025).
Moreover, several point mutations are associated with pyridoxine-refractory sideroblastic anemia (SIDBA2).
We present a computational study to structurally characterize SLC25A38 as a mitochondrial PLP transporter.
Structural model was generated using AlphaFold3. Binding pocket was identified using PrankWeb and validated by
comparison with homologous known proteins. Ligand binding was predicted through in silico molecular docking of
PLP and glycine using GNINA The apo and holo complexes were embedded in a phospholipid bilayer representative
of the inner mitochondrial membrane and subjected to 1 μs molecular dynamics simulations in triplicate.
Across all replicates, glycine rapidly dissociated from the transporter within 10 ns, indicating low binding
stability. In contrast, PLP showed stable binding, supported by MM/PBSA analysis with an average binding
energy of -68.63 ± 0.34 kcal/mol. Per-residue energy decomposition identified R96, R187, and K242 as key
contributors to PLP stabilization (from -14 to -16 kcal/mol), whereas these residues showed unfavorable
contributions in glycine simulations (~1 kcal/mol).
Our results support the hypothesis that SLC25A38 may function as a mitochondrial PLP transporter and provide
molecular insights into its proposed role.
A-P.40: Uncovering the Dark Side of the Immunopeptidome
Track: Proteins and structural biology
-
Steffen Lemke, Department of Peptide-based Immunotherapy, Institute of Immunology, University and
University Hospital Tübingen, Germany
-
Marcel Wacker, Department of Peptide-based Immunotherapy, Institute of Immunology, University and
University Hospital Tübingen, Germany
-
Jens Bauer, Department of Peptide-based Immunotherapy, Institute of Immunology, University and University
Hospital Tübingen, Germany
-
Jonas Scheid, Department of Peptide-based Immunotherapy, Institute of Immunology, University and
University Hospital Tübingen, Germany
-
Annika Nelde, Department of Peptide-based Immunotherapy, Institute of Immunology, University and
University Hospital Tübingen, Germany
- Sven Nahnsen, Quantitative Biology Center (QBiC), University of Tübingen, Germany
-
Juliane S. Walz, Department of Peptide-based Immunotherapy, Institute of Immunology, University and
University Hospital Tübingen, Germany
Presentation Overview: Show
The immune system targets cancer cells via T cell-mediated recognition of antigenic peptides presented on
human leucocyte antigen (HLA) molecules. Mass spectrometry (MS)-based immunopeptidomics represents the
standard method for direct identification of naturally presented tumor peptides, essential for developing T
cell-based immunotherapies. However, a large fraction of proteins and protein regions lack detectable
HLA-presented peptides, forming immunopeptidomic ‘dark spots' whose biological and technical origins remain
uncharacterized. Here, we used the PCI-DB, a uniquely large database comprising >10 million HLA-presented
peptides, for comprehensive mapping of the immunopeptidome dark spot landscape. While HLA binding predictions
suggest a median proteome coverage of 87% per person across UK Biobank donors, experimental immunopeptidome
data have captured only 30% of the human proteome. HLA peptide coverage scales with gene expression, with the
11% of proteins lacking any detectable peptides showing the lowest expression. Analysis of cancer-related
variants from the COSMIC database revealed that driver mutations are more frequent in light spots, whereas
passenger mutations occur more frequently in dark spots. To investigate structural and post-translational
contributors to dark spot formation, transmembrane annotations (TOPdb) and glycosylation sites (GlyGen) were
compared to dark and light spots, revealing both to be overrepresented in dark areas. Reprocessing PCI-DB data
to include glycopeptides increased glycosylation site coverage by 58%. Broader post-translational modification
searches resolved ~1.4% of dark spots, with adapted MS acquisition methods expected to increase this further.
Overall, characterizing dark spots as either technical or biological in origin will enhance epitope selection
accuracy and expand the repertoire for T-cell-based immunotherapy.
A-P.41: Stability-informed low-N learning of protein fitness landscape
Track: Proteins and structural biology
-
Shannon Zhang, Georgia Institute of Technology, United States
- Yunan Luo, Georgia Institute of Technology, United States
Presentation Overview: Show
Predicting how protein sequence mutations affect fitness is central to protein engineering and variant effect
interpretation, yet most experimental assays yield only limited labeled variants, making low-N learning a
major challenge.
We present PsiFit, a protein stability-informed framework for low-N protein fitness prediction. PsiFit is
motivated by the biological principle that proteins must maintain sufficient stability to function. PsiFit
explicitly incorporates this constraint into learning the sequence–fitness landscape. Building on our prior
work in mutation stability prediction (SPURS) and low-N fitness modeling (ConFit), PsiFit introduces a unified
framework that integrates stability-aware biophysical priors with protein language model adaptation.
Specifically, PsiFit leverages predicted mutation-induced stability changes as an explicit signal to guide
fitness learning, while employing a contrastive fine-tuning strategy that preserves the general biological
knowledge encoded during pretraining and aligns the model with experimentally measured fitness. By embedding
stability constraints directly into the learning process, PsiFit improves sample efficiency and mitigates
overfitting in small and noisy datasets.
Across more than 100 deep mutational scanning datasets in ProteinGym, PsiFit consistently outperformed
multiple state-of-the-art low-N fitness prediction methods. The gains were observed across diverse proteins
and assay types, indicating that stability provides a broadly useful inductive bias for learning protein
fitness landscapes from sparse data. These results show that integrating biophysical stability priors with
protein language models offers an effective and general strategy for data-efficient protein fitness
prediction, with implications for ML-guided protein engineering and variant effect interpretation.
A-P.42: Integrative Proteomic Analysis of Inner Ear and Systemic Fluids in Noise-Induced Hearing Loss
Track: Proteins and structural biology
-
Motahare Khorrami, School of Engineering, Macquarie University, Sydney, New South Wales, Australia,
Australia
-
Robert Gay, Cochlear Limited, 1 University Avenue, Macquarie University, North Ryde, NSW 2109, Australia,
Australia
-
Ya Lang Enke, Cochlear Limited, 1 University Avenue, Macquarie University, North Ryde, NSW 2109,
Australia, Australia
-
Mohsen Asadnia, School of Engineering, Faculty of Science and Engineering, Macquarie University, North
Ryde, NSW 2109, Australia, Australia
-
Paul A. Haynes, School of Natural Sciences, Macquarie University, North Ryde, NSW 2109, Australia,
Australia
-
Christopher Pastras, School of Engineering, Faculty of Science and Engineering, Macquarie University,
North Ryde, NSW 2109, Australia, Australia
Presentation Overview: Show
Noise-induced hearing loss (NIHL) arises from prolonged exposure to loud sounds in the environment. Early
diagnosis is challenging because some individuals may not exhibit significant threshold shifts on standard
audiometry and other functional tests. Characterising the cellular changes underlying NIHL may aid early
intervention and the development of targeted therapies. This study provides the first comparative proteomic
analysis of perilymph, cerebrospinal fluid (CSF), whole blood, and plasma from Sprague-Dawley rats after acute
noise exposure. We identify shared molecular signatures of NIHL relative to controls, and potentially relevant
systemic biomarkers.
Guinea pigs were exposed to acute noise (110 dB, 90 min) under anaesthesia, resulting in a ~40 dB click-evoked
threshold shift, quantified via cochlear nerve compound action potential measurement. Proteins from perilymph,
CSF, blood, and plasma were analysed by nanoLC–MS/MS in data-independent acquisition mode. Differential
expression analysis was conducted using the Mass Spectrometry Downstream Analysis Pipeline (MS-DAP).
We identified a total of 2,268 proteins in perilymph, 2,548 in CSF, 3,351 in whole blood, and 993 in plasma
across 10 animals. Using the MS-EmpiRe model, differential expression analysis revealed 388 significantly
altered proteins in perilymph, 396 in CSF, 80 in blood, and 95 in plasma following noise exposure. Exposure to
intense noise triggers a systemic surge in Serum amyloid A protein, Isocitrate dehydrogenase,
Fructose-bisphosphate aldolase, and Glutathione S-transferase B across the samples. This synchronized
elevation across multiple fluid compartments reflects a complex physiological response involving acute
inflammation, altered energy metabolism, and heightened oxidative stress defense in response to acoustic
trauma.
A-P.43: Unveiling the Diversity of Catalytically Inactive Long-B Prokaryotic Argonaute Proteins
Track: Proteins and structural biology
-
Simonas Ašmontas, Institute of Biotechnology, Life Sciences Center, Vilnius University,
Lithuania
-
ÄŒeslovas Venclovas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
Presentation Overview: Show
Argonaute proteins are present in all domains of life. Prokaryotic Argonaute proteins (pAgos) are known to act
as defense systems against foreign genetic elements. Historically, pAgos have been divided into two major
groups, Short and Long pAgos, the latter being subdivided into A and B clades. Most pAgos in the Long-A clade
have an intact canonical Argonaute nuclease active site, whereas Long-B pAgos have lost this feature. Short
pAgos,being catalytically inactive like the Long-B pAgos, compensate for the lack of an active site by
employing fused effector domains. However, most Long-B pAgos do not feature such fusions. Instead, in some
cases it has been shown that Long-B pAgos function together with associated effector proteins. Yet, the
diversity and distribution of putative Long-B pAgo effectors so far have not been thoroughly assessed.
Our bioinformatic study presents a comprehensive picture of the associations of Long-B pAgos and their
effectors. Using a large set of Long-B pAgos we discovered three major groups of effector proteins associated
with Long-B pAgos, and a minor group representing pAgos fused with effector domains. The members of thefirst
major group of effector proteins exhibit modular organization, consisting of a conserved adaptor domain and a
fused variable effector, including PDEXK nuclease domains, putative transmembrane segments, and other toxic
domains. The second major group is made up of single-domain SIR2 effectors. The last major group corresponds
to a mysterious protein family whose possible functions are yet unclear, but whose ability to interact with
pAgos is supported by AlphaFold3 modeling experiments.
A-P.44: Characterization of a highly diverged Cas7 homologs in Felix phages
Track: Proteins and structural biology
-
Jyotika Pachauri, Institute of Biotechnology, Life Sciences Center, Vilnius University,
Lithuania
-
Irmantas Mogila, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
-
Lidija Truncaitė, Institute of Biochemistry, Life Sciences Center, Vilnius University, Lithuania
-
Jonas Juozapaitis, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
-
Aistė Skorupskaitė, LSC-EMBL Partnership Institute for Genome Editing Technologies, Life Sciences Center,
Vilnius University, Lithuania
-
Lukas Valančauskas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
-
Praneet Prabhanjan, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
-
Monika Šimoliūnienė, Institute of Biochemistry, Life Sciences Center, Vilnius University, Lithuania
-
Kristupas Užkurnys, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
-
Tomas Šinkūnas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
-
Česlovas Venclovas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
-
Patrick Pausch, LSC-EMBL Partnership Institute for Genome Editing Technologies, Life Sciences Center,
Vilnius University, Lithuania
-
Darius Kazlauskas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
Presentation Overview: Show
Cas7 proteins are essential structural components of Class 1 CRISPR-Cas effector complexes. They form a
backbone that binds CRISPR RNA (crRNA) through an RNA Recognition Motif (RRM) fold to facilitate target
recognition and interference. Apart from their role in the CRISPR-Cas system, no independent functions have
been described. We identified Gp87, a highly diverged Cas7 homolog encoded by Felix phage VpaE1 outside any
recognized CRISPR-Cas genomic context. Despite having minimal sequence conservation, structural models
indicate that Gp87 shares the canonical Cas7 fold but lacks the typical active site and possesses two unique
alpha-helical insertions. Using comparative genomics, biochemical experiments, structural modeling, and deep
learning models we characterized Gp87 and its distribution across phages, archaea, and bacteria.
We found that Gp87 is co-expressed with Gp86, a helix-turn-helix (HTH) protein, and specifically binds to a
flanking non-coding RNA (ncRNA) element containing GGTNN repeats. Structural models show that this interaction
is mediated by Gp87's helical insertions which binds with the GG-rich repeats, an RNA-recognition process
different from that of canonical Cas7 proteins. Furthermore, we identified that the architecture, comprising
Gp87, Gp86, and the flanking ncRNA elements, is conserved not only across all known Felix phages, but also
among diverse bacterial clades that maintain conserved GG-repeat regions and, in some groups, additionally
encode flanking HNH endonucleases. This work reveals a novel, widespread Cas7-like system that operates
independently of the CRISPR-Cas machinery, with a distinct RNA-binding mechanism and a transposon-associated
genomic context.
A-P.45: Identification of new P-loop NTPase-like families
Track: Proteins and structural biology
-
Shreya Tapaswi, Department of Bioinformatics, Institute of Biochemistry and Biophysics, Polish
Academy of Sciences, Poland
-
Kamil Steczkiewicz, Department of Bioinformatics, Institute of Biochemistry and Biophysics, Polish Academy
of Sciences, Poland
Presentation Overview: Show
P-loop NTPases, constitute a major portion of any organism's proteome; they generally bind or hydrolyze
nucleoside triphosphates (NTPs) for chemo-mechanical energy transduction to drive biochemical processes like,
gene expression regulation, signal transduction, chromosome partitioning, DNA repair and recombination,
intracellular and membrane transport, etc. They retain αβα sandwich structural core of central β-sheet of
4-6 parallel β-strands surrounded by α-helices (CATH Superfamily 3.40.50.300 P-loop containing nucleotide
triphosphate hydrolases, SCOP fold: P-loop containing nucleoside triphosphate hydrolases) and possess two
major highly conserved sequence motifs, namely, Walker A and Walker B. Despite the common fold, P-loop
NTPase-like proteins display considerable diversity on both sequence and structure levels. Using remote
homology detection methods we identified seven new families of unknown function (Domains of Unknown Functions,
DUFs), to be classified as P-loop NTPases, and described their diversity including variability in either the
number of the core structural elements or composition of functionally important Walker motifs. Based on
predicted cellular localization, presence of transmembrane helices, genomic neighborhood and domain fusion
analysis, we have hypothesized on identified families' functions which include phospholipid biosynthesis,
stress response, etc. We hope our results will inspire further experimental efforts to unravel detailed
functions of these seven newly annotated families, especially in the context of developing drugs against
pathogenic species.
A-P.46: Machine learning approaches for charting the functional and structural landscape of protein
ubiquitination
Track: Proteins and structural biology
-
Julian van Gerwen, ETH Zurich, Switzerland
- Pedro Beltrao, ETH Zurich, Switzerland
Presentation Overview: Show
Protein ubiquitination regulates cell biology through diverse avenues, from quality control-linked protein
degradation to regulatory functions such as modulating protein-protein interactions and protein conformations.
Mass spectrometry-based proteomics has allowed proteome-scale quantification of hundreds of thousands of
ubiquitination sites (ubi-sites), however the functional importance and molecular mechanisms of most ubi-sites
remain undefined. By integrating multi-species proteomics data we found that regulatory ubi-sites may be more
important than degradation-linked sites due to higher evolutionary conservation. To further prioritize
regulatory ubi-sites performing cell-critical functions, we developed a machine learning-based ubi-site
functional score by integrating evolutionary, proteomics, and structural features. Analysis of clinical
genomics data and experiments with chemical genetics and genetic code expansion confirmed that our score
pinpoints ubi-sites important for cellular fitness. Next, to pave the way for mechanistic interpretation of
regulatory ubi-sites we developed an approach for modeling protein-protein covalent bonds - such as
ubiquitination - in the deep-learning structural predictor AlphaFold3, which is not natively possible. We
benchmarked this approach by re-predicting 338 experimental structures of ubiquitinated proteins in the
Protein Data Bank, and achieved moderate-to-high accuracy across a range of protein types, including
mono-ubiquitinated proteins, polyubiquitin chains, and catalytic intermediates of ubiquitin-processing
enzymes. Combined with our ubi-site functional score, AlphaFold3 confirmed the regulatory potential of diverse
ubi-sites by revealing ubiquitination-induced structural changes connected to protein functional states.
Overall, we have leveraged machine learning through multiple avenues to chart the functional significance and
structural consequences of ubiquitination at proteome-scale.
A-P.47: Multi-property Evaluation of Nanobodies Designed by an AI-Driven Pipeline for Binding to Specified
Sites with Experimental Validation
Track: Proteins and structural biology
-
Yiran Wang, University College London, United Kingdom
- Paul Dalby, University College London, United Kingdom
Presentation Overview: Show
Obtaining nanobodies with superior overall performance remains a major challenge in the field of protein
engineering. Traditional approaches often yield nanobodies that bind only to the most optimal site, making it
difficult to explore other regions of the antigen surface or to recognize specific epitopes. Although
researchers have gradually introduced AI-driven tools into protein design, current designs remain restricted
to known epitopes, with limited exploration of novel binding sites. Moreover, most studies emphasize binding
affinity as the primary objective while overlooking other properties of the designed nanobodies. Therefore,
fully leveraging the potential of AI-driven tools to identify a broader range of binding sites, designing
nanobodies capable of interacting with user-defined hotspots, and systematically evaluating their properties
will be a key future direction.
Based on this motivation, we developed an AI-driven nanobody design-build-test pipeline that integrates both
structure-based and sequence-based design strategies to generate candidates targeting specified epitopes. A
multi-attribute evaluation framework was then used to assess physicochemical properties. We designed a series
of nanobodies and experimentally validated their performances, including expression, stability, and binding
affinity. Our results demonstrate that some variants achieve high-level expression, many exhibit satisfactory
stability, and some show favorable binding activity. Further comparative analyses highlight the contributions
of key pipeline components to overall design performance, as well as the variability in nanobody behavior
across different designs.
In summary, this work provides a general framework for AI-based nanobody design with experimentally validated
functionality, enabling systematic exploration of nanobody-antigen interactions.
A-P.48: SELECTIVE MODULATION OF CHEMOKINE-LIKE RECEPTOR 2 BY SMALL MOLECULES
Track: Proteins and structural biology
-
Jonas Golling, Institute of Biochemistry, Faculty of Life Sciences, Leipzig University, 04103
Leipzig, Germany, Germany
-
Tina Schermeng, Institute of Biochemistry, Faculty of Life Sciences, Leipzig University, 04103 Leipzig,
Germany, Germany
-
Emre Duman, Institute of Pharmaceutical Chemistry, Goethe University Frankfurt, 60438 Frankfurt am Main,
Germany, Germany
-
Jens Meiler, Institute for Drug Discovery, Leipzig University, Leipzig 04103, Germany, Germany
-
Ewgenij Proschak, Institute of Pharmaceutical Chemistry, Goethe University Frankfurt, 60438 Frankfurt am
Main, Germany, Germany
-
Annette Beck-Sickinger, Institute of Biochemistry, Faculty of Life Sciences, Leipzig University, 04103
Leipzig, Germany, Germany
Presentation Overview: Show
The closely related chemokine-like receptors 1 and 2 (CMKLR1 and CMKLR2), both activated by the adipokine
chemerin, play an important role in inflammation and metabolism. While CMKLR1 drives chemotaxis and
adipogenesis, CMKLR2 has been considered as an atypical scavenger receptor for a long time. Knockout studies
of CMKLR2 demonstrated a reduction in insulin release and glucose uptake, while showing no impact on immune
cells. This suggests a distinct role for CMKLR2 in glucose homeostasis. Thus, selective chemical probes
targeting CMKLR2 could further elucidate the chemerin system and serve as a foundation for novel therapeutics
targeting metabolic diseases.
Originally developed as a ligand for G protein-coupled receptor 132 (GPR132), the small molecule T-10418 was
identified to induce arrestin-3 recruitment at CMKLR2 during GPCR-selectivity screening. Subsequent
pharmacological characterization confirmed its selectivity for CMKLR2 over CMKLR1. Through systematic
ligand-based lead optimization of the partial agonist T-10418, the maximum efficacy was enhanced from 50% to
100% in comparison to the reference peptide chemerin-9. Additionally, a group of derivatives exhibiting
antagonistic properties was identified. By integrating computational docking and molecular dynamics
simulations with biochemical validation, the binding mode of the most potent agonist within the orthosteric
pocket was elucidated. Based on these models, a structural hypothesis for the observed CMKLR2 selectivity and
activation mechanism is derived. With the identification of these selective probes, the base for a precise
pharmacological interrogation of CMKLR2 in metabolic health has been established.
A-P.49: Multiscale Computational Dissection of CCRL2-Mediated Chemerin Presentation
Track: Proteins and structural biology
-
Arianna Migliorini, Sapienza University of Rome, Italy
-
Domenico Raimondo, Department of Molecular Medicine, Laboratory Affiliated to Istituto Pasteur Italia,
Sapienza University of Rome, Italy
Presentation Overview: Show
CCRL2 is an atypical, non-signaling G-protein-coupled receptor (GPCR) that concentrates chemerin on the
surface of expressing cells, facilitating its presentation to the functional chemerin receptor CMKLR1 and
enabling recruitment of innate immune cells under inflammatory conditions. Despite this important biological
role, the structural basis of CCRL2–chemerin recognition has remained poorly characterized, with no
experimental structure available for CCRL2 to date. We present a comprehensive multiscale computational study
integrating coarse-grained molecular dynamics (CG-MD) and all-atom molecular dynamics (AA-MD) simulations with
structural modeling to dissect the mechanism of CCRL2–chemerin binding. Starting from AlphaFold2-predicted
structures, we performed 26 independent CG-MD simulations totaling approximately 80 microseconds,
demonstrating spontaneous chemerin association with CCRL2 following a two-step binding mechanism analogous to
canonical chemokine receptors. Free energy landscape analysis identified three representative bound
conformations, which were further refined through 45 microseconds of CG stable-binding simulations and
subsequently characterized at atomic resolution via all-atom MD. Our results reveal a stable binding interface
primarily mediated by chemerin's β1 strand and CCRL2's extracellular loop 2 (ECL2), with additional
electrostatic anchoring between CCRL2's N-terminal domain and chemerin's loop 3. Critically, chemerin's
C-terminal region remains solvent-accessible, consistent with a presentation-competent, nonsignaling binding
mode. A modeled ternary CCRL2–chemerin–CMKLR1 complex provides a structural framework for chemerin handoff
to CMKLR1. Notably, given the established role of CCRL2 in controlling NK cell homing in non-small cell lung
cancer, these structural insights may inform therapeutic strategies aimed at enhancing innate immune
recruitment within the tumor microenvironment.
A-P.50: Biological meaning in protein embedding space is resolution-dependent
Track: Proteins and structural biology
-
Licheng Zong, EMBL-EBI; University of Bath; The Chinese University of Hong Kong, Hong Kong
-
Jinzheng Ren, EMBL-EBI; University of Bath; Australian National University, United Kingdom
- Yu Li, The Chinese University of Hong Kong, Hong Kong
- Robert Finn, EMBL-EBI, United Kingdom
-
Jiawei Wang, EMBL-EBI; University of Bath, United Kingdom
Presentation Overview: Show
Protein language model embeddings are increasingly used to organise biological sequences, yet how biological
meaning is encoded within embedding neighbourhoods remains poorly understood. Using two independent
hierarchical enzyme systems, carbohydrate-active enzymes and peptidases, we investigated how biological
interpretation changes across embedding organisations aligned to different levels of biological hierarchy.
Different embedding organisations give rise to distinct neighbourhood semantics. When aligned to
membership-boundary resolution, embeddings robustly separated artefacts and unrelated proteins from members of
the target category. However, embeddings aligned to functional-grouping resolution maintained compositional
neighbourhood structure for multi-domain proteins spanning more than one functional or catalytic group.
Finally, embeddings aligned to local-family resolution recovered compact family-like neighbourhoods, including
families withheld from training, while weakening broader membership-boundary and functional-grouping
relationships. Moreover, embeddings optimised toward the same level of biological organisation retain
different biological relationships depending on optimisation trajectory employed. Together, our results show
that proximity in protein embedding space has no fixed biological interpretation. Instead, biological meaning
emerges across embedding resolutions through selective preservation of different forms of biological
organisation.
A-P.51: DeepAlloWeb: A Web Server for Allosteric Pocket Prediction and Attention Based Interpretation
Track: Proteins and structural biology
-
Moaaz Ur Rehman Azhar Khokhar, Koç University, Turkey
- Ozlem Keskin, Koç University, Turkey
- Attila Gursoy, Koç University, Turkey
Presentation Overview: Show
Allostery plays a central role in protein function and is an important direction in drug discovery because
allosteric drugs can be more selective and may lead to fewer side effects than orthosteric drugs. Here, we
present DeepAlloWeb, an interactive web server built on our previous method, DeepAllo, for allosteric pocket
prediction and interpretation. DeepAllo combines FPocket features with a fine-tuned ProtBERT-BFD protein
language model to rank candidate allosteric pockets. DeepAlloWeb extends this framework by allowing users to
visualize residue-level attention and inspect how the model relates residues within and beyond predicted
pockets. To study this more systematically, we performed a dataset-wide analysis across all 30 layers and 16
heads of the model. We found that the most informative patterns were concentrated in a smaller group of mainly
late-layer heads, especially layer 30, and that these heads often highlighted residues far from the selected
pocket residue, consistent with longer-range allosteric communication. For example, layer 30 head 11 achieved
a mean average precision of 0.393. In a CheY case study, attention from D57 consistently highlighted known
communication residues including T87 and Y106 across multiple heads. Together, these results show that
DeepAlloWeb provides both prediction and interpretation in a single resource for studying allosteric
regulation. The webserver may be accessed via: https://3dpath.ku.edu.tr/DeepAllo/.
A-P.52: A Machine Learning Framework for High-Performance Prediction of Enzyme-Substrate Interactions Using
Advanced Protein Embeddings and Molecular Fingerprints
Track: Proteins and structural biology
- João Ribeiro, Centre of Biological Engineering, Portugal
- Ana Monteiro, Centre of Biological Engineering, Portugal
-
Dick de Ridder, Bioinformatics Group, Department of Plant Sciences}, \orgname{Wageningen University and
Research, Netherlands
-
Oscar Dias, Centre of Biological Engineering; LABBELS - Associate Laboratory, Portugal
-
Miguel Rocha, Centre of Biological Engineering; LABBELS - Associate Laboratory, Portugal
Presentation Overview: Show
Motivation: Enzymes catalyze biochemical reactions with remarkable specificity and efficiency, yet predicting
enzyme–substrate interactions remains challenging due to limited experimental data, incomplete annotations,
and the scarcity of validated non-binding pairs.
Traditional computational approaches, such as docking and molecular dynamics, are resource-intensive and
poorly scalable, while many machine learning (ML) models struggle to generalize to chemically diverse or
previously unseen molecules.
Results: We present a novel ML framework that integrates experimentally validated enzyme–substrate pairs
with advanced feature representations.
Our approach combines protein embeddings from the pre-trained language model ESM2 with NPClassifierFP
molecular fingerprints to capture functional and structural properties of enzymes and substrates.
A gradient boosting classifier (XGBoost) is trained using rigorous data partitioning with the GraphPart method
to prevent data leakage and ensure robust generalization across similarity thresholds. Benchmarking against
state-of-the-art models (ESP, ProSmith) and similarity-based approaches (BLASTp + Tanimoto similarity)
demonstrates that our method consistently outperforms existing approaches, achieving high F1 scores and
Matthews Correlation Coefficients (MCC).
Notably, the model maintains strong predictive performance under stringent protein (40–80% sequence
identity) and compound (20–80% Tanimoto similarity) similarity constraints, and generalizes well to
chemically distinct, previously unseen molecules.