View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

A-P.01: PLM-eXplain: Divide and Conquer the Protein Embedding Space
Track: Proteins and structural biology
  • Jan van Eck, Utrecht University, Netherlands
  • Dea Gogishvili, Utrecht University, Netherlands
  • Wilson Silva, Utrecht University, Netherlands
  • Sanne Abeln, Utrecht University, Netherlands


Presentation Overview: Show

Motivation: Protein language models (PLMs) have revolutionised computational biology through their ability to generate powerful sequence representations for diverse prediction tasks. However, their black-box nature limits biological interpretation and translation to actionable insights. Bridging this gap requires approaches that maintain predictive performance while providing interpretable explanations of model behaviour.

Results: We present PLM-eXplain (PLM-X), an explainable adapter layer that bridges this gap by factoring PLM embeddings into two complementary components: an interpretable subspace based on established biochemical features, and a residual subspace that retains predictive, non-interpretable information. Using embeddings from ESM2 and ProtBert, PLM-X incorporates well-established properties, including secondary structure and hydropathy, while maintaining high predictive performance. We demonstrate the effectiveness of our approach across three biologically relevant classification tasks: extracellular vesicle association, transmembrane helix prediction, and aggregation propensity prediction. PLM-X enables biological interpretation of model decisions without sacrificing accuracy, offering a generalisable solution for enhancing PLM interpretability across various downstream applications.

A-P.02: A self-assembled protein β-helix as a self-contained biofunctional motif
Track: Proteins and structural biology
  • Javier Garcia-Ruiz, National Physical Laboratory, United Kingdom
  • Camilla Dondi, National Physical Laboratory, United Kingdom
  • Maxim G. Ryadnov, National Physical Laboratory, United Kingdom


Presentation Overview: Show

Nature constructs matter by employing protein folding motifs, many of which have been synthetically reconstituted to exploit function. A less understood motif whose structure-function relationships remain unexploited is formed by parallel beta-strands arranged in a helical repetitive pattern, termed a beta-helix. Herein we reconstitute a protein beta-helix by design and endow it with biological function. Unlike beta-helical proteins, which are contiguous covalent structures, this beta-helix self-assembles from an elementary sequence of 18 amino acids. Using a combination of experimental and computational methods, we demonstrate that the resulting assemblies are discrete cylindrical structures exhibiting conserved dimensions at the nanoscale. We provide evidence for the structures to form a carpet-like three-dimensional scaffold promoting and inhibiting the growth of human and bacterial cells, respectively, while being able to mediate intracellular gene delivery. The study introduces a self-assembled beta-helix as a self-contained bio- and multi-functional motif for exploring and exploiting mechanistic biology.

A-P.03: Structure-Guided RNA Design via Multinomial Diffusion
Track: Proteins and structural biology
  • Georg Back, Weihenstephan-Triesdorf University of Applied Sciences, Department of Sustainable Agriculture and Energy Systems, Germany
  • Dirk Walther, Max Planck Institute for Molecular Plant Physiology, Department of Bioinformatics, Germany
  • Florian Haselbeck, Weihenstephan-Triesdorf University of Applied Sciences, Department of Sustainable Agriculture and Energy Systems, Germany


Presentation Overview: Show

The structure of RNA plays an essential role in its biological function, with the secondary structure being the primary determinant. While significant advances have been made in RNA folding models, the inverse problem, i.e., generating sequences that fold into a target structure has received less attention. Existing approaches commonly rely on iterative optimization, often guided by a folding oracle, which can limit scalability and risk converging to local optima. Diffusion models have shown a superior performance in many application areas, but their noising process is not directly applicable to discrete data such as nucleotide sequences. In this paper, we adapt the multinomial diffusion framework, which introduces discrete noise via categorical transitions, to the RNA sequence space. Furthermore, we incorporate secondary structure information to guide the generative process toward sequences that fold into a desired target structure. For this purpose, we evaluate two architectural variants: (i) modeling the problem either by projecting the sequence embedding into a 2D pairwise representation that reflects the structure of a contact map, or (ii) by restructuring the architecture as a purely 1D sequence model. Our results show that both approaches generate sequences that fold with high accuracy while maintaining substantial variability among generated sequence samples. We further evaluated constrained generation via sequence inpainting, resulting in performance comparable to unconstrained generation. Even though the performance on the human-designed Eterna100 benchmark was lower than that of the state-of-the-art method LEARNA, our results indicate that diffusion-based models provide a promising and scalable framework for conditional nucleotide sequence generation.

A-P.04: Protein language models improve detection of divergent beta-lactamases
Track: Proteins and structural biology
  • Mateusz Wlodarski, McMaster University, Canada
  • Andrew McArthur, McMaster University, Canada


Presentation Overview: Show

Accurate detection of β-lactamase genes is critical for antimicrobial resistance surveillance, yet highly divergent sequences with little similarity to curated enzymes often escape detection. Although protein language models (PLMs) show promise for functional inference, their ability to detect remote resistance genes and the biological basis of their predictions remain unclear.
We evaluated transformer-based PLMs for β-lactamase detection across controlled sequence divergence bins. Using the Comprehensive Antibiotic Resistance Database (CARD) as a curated reference space, we stratified withheld β-lactamase sequences by identity to their nearest CARD homolog and compared performance against sequence similarity and Hidden Markov model -based approaches. PLM classifiers maintained high recall in low-identity bins where traditional methods failed, while larger representations such as Evolutionary Scale Modeling (ESM-2) also improved false positive rates on hydrolase negatives. Embedding analyses revealed Ambler class-consistent structure, and motif-centered perturbation analyses demonstrated that conserved catalytic regions contribute disproportionately to model predictions, supporting the biological relevance of learned features. Together, these results demonstrate that PLMs can extend β-lactamase detection beyond the limits of traditional sequence similarity methods and highlight the benefits of larger capacity learned representations.

A-P.05: TRPM8 Virtual Screening: The Essential Role of Target-Specific, Inactive-Enriched Machine-Learning Scoring Functions
Track: Proteins and structural biology
  • Nivya James, Imperial College London, United Kingdom
  • Pedro Ballester, Imperial College London, United Kingdom


Presentation Overview: Show

Motivation: The Transient Receptor Potential Melastatin 8 (TRPM8) ion channel is an emerging therapeutic target implicated in pain, inflammation, and other disorders. While structure-based virtual screening (VS) is widely used in drug discovery, the effectiveness of generic and target-specific machine-learning (ML) scoring functions (SFs) for ion channels such as TRPM8 remains to be investigated.

Results: We established the first comprehensive structure-based VS benchmark for TRPM8. This evaluates classical docking SFs, generic ML SFs and target-specific ML models retrospectively across multiple TRPM8 protein conformations and docking protocols. The evaluation revealed that generic ML rescoring at most provided modest and conformation-dependent improvements over classical docking on this target. Target-specific ML models based only on protein-ligand interaction fingerprints improved early enrichment on test sets when regression algorithms were employed. However, performance remained sensitive to protein conformation choice and chemical dissimilar-ity. Systematic evaluation of feature representations showed that combining structure-derived interaction features with ligand-based descriptors enhanced enrichment under selected docking tool-protein conformation combinations. Critically, presenting the learning algorithm with many more negative training instances, via inactive-enriched training, strongly reduced false positives. This also resulted in markedly improved generalization to chemically dissimilar test sets for structure-based models. By contrast, ligand-only QSAR models proved sensitive to more class-imbalanced training sets. Under inactive-enriched training, PLEC-based support vector regression (SVR) models achieved robust and strong VS performance, highlighting the critical role of learning the vast diversity of inactive molecules better during model training in developing reliable target-specific ML SFs for TRPM8.

A-P.06: A pretrained model for RNA inverse folding with formal grammar representations
Track: Proteins and structural biology
  • Kentaro Watanabe, Keio University, Japan
  • Manato Akiyama, Kitasato University, Japan
  • Yasubumi Sakakibara, Kitasato University, Japan


Presentation Overview: Show

Motivation: RNA inverse folding designs nucleotide sequences that fold into a given RNA secondary structure, enabling the creation and modification of functional RNAs. State-of-the-art solvers are largely search-based, but the search space grows exponentially with sequence length and structural complexity, and it is non-trivial to incorporate additional sequence-level constraints such as target GC content or family-specific conserved features.
Results: We propose a generative model based on a Transformer-based conditional variational au-toencoder (CVAE) that takes a context-free grammar (CFG) parse-tree representation of the target secondary structure as input and generates sequences that fold into the specified secondary struc-ture. The grammar-based tree makes the hierarchical organization and base-pair correspondences explicit. By combining self-refinement learning and latent-space optimization, we substantially im-prove the recovery of high-fidelity solutions. On the EteRNA100 benchmark, our model alone achieves competitive accuracy, and the generated sequences consistently improve the success rate of downstream search-based solvers when used as warm starts. We further demonstrate controllable generation under GC-content constraints and improved family consistency through large-scale pre-training on natural RNAs.

A-P.07: Energy-Matched Diffusion for Antibody Generation
Track: Proteins and structural biology
  • Vasanth Durvasula, Nanyang Technological University, Singapore
  • Tiara Natasha Binte Sayuti, Nanyang Technological University, Singapore
  • Jagath Rajapakse, Nanyang Technological University, Singapore


Presentation Overview: Show

Motivation: Diffusion models have been used for antibody design, which enable joint generation of complementarity-determining region (CDR) sequence and structure, but their denoising objectives are typically optimized locally and do not explicitly enforce global geometric consistency. As a result, generated structures exhibit accumulated geometric drift despite locally accurate predictions.

Results:
We propose an energy-matched diffusion framework that learns an antibody-specific scalar energy function jointly with the denoising network. During training, an energy-matching objective aligns the gradient of the learned energy with diffusion-consistent denoising targets, constraining the denoising vector field to arise from a single scalar potential. This introduces a global integrability bias that promotes smoother, more consistent denoising trajectories and reduces structural instability during generation. To avoid hand-tuned multi-term loss weights, we introduce an exponential-moving-average controller that adaptively scales the energy-matching term relative to a reconstruction anchor throughout training.
On antibody–antigen complexes from the SAbDab benchmark, the proposed method improves both sequence quality and backbone accuracy, achieving lower RMSD than an unguided diffusion baseline. Through our results, we show that energy matching serves as an effective training-time regularizer for improving geometric fidelity in antibody CDR generation without requiring external energy supervision or manual hyperparameter search.

A-P.08: Gated Modulation Protein Language Model for Prediction of Antimicrobial Peptide Activity
Track: Proteins and structural biology
  • Prem Singh Bist, Nanyang Technological University, Singapore
  • Jagath Rajapakse, Nanyang Technological University, Singapore


Presentation Overview: Show

Motivation: Antibiotic resistance poses a critical global health threat, demanding new strategies for rapid therapeutic discovery. Antimicrobial peptides (AMPs) are promising candidates due to their broad-spectrum of activity and design flexibility; however, experimental screening of peptide activity remains costly and time-intensive.

Results: We introduce a novel Gated Modulation Protein language model, GMProt, that integrates pretrained sequence representations with adaptive gating to predict peptide functional activity, including Minimum Inhibitory Concentration (MIC). The framework enhances activity estimation while enabling reliable candidate ranking prior to wet-lab validation. The proposed framework demonstrates strong predictive accuracy and ranking consistency, outperforming eight baseline models. Specifically, it reduces activity estimation error (RMSE) from 0.524 to 0.493 and increases ranking performance (Kendall's correlation) from 0.462 to 0.487 compared with the current state-of-the-art approach. These findings demonstrate the potential of the proposed approach to enable data-driven peptide prioritization, offering a scalable route to accelerate AMP discovery.

A-P.09: Integrating Graph Encoders and Protein Language Models for Antibody–Antigen Binding Free Energy Prediction
Track: Proteins and structural biology
  • Tiara Natasha Binte Sayuti, Nanyang Technological University, Singapore
  • Vasanth Durvasula, Nanyang Technological University, Singapore
  • Aishwarya Anand, Nanyang Technological University, Singapore
  • Jagath Rajapakse, Nanyang Technological University, Singapore


Presentation Overview: Show

Motivation: Antibody–antigen affinity prediction is central to antibody engineering and lead optimization, where accurate estimation of binding free energy (Delta G) enables ranking and prioritization of variants. Despite rapid advances in protein language models and graph neural networks, most studies emphasize binary interaction prediction or build regression models under random splits that risk antigen-level leakage. Such settings may overestimate generalization, particularly when closely related antigen families appear across training and test partitions. Robust $\Delta G$ prediction under cluster-aware evaluation remains underexplored.
{Results: We introduce GEPBind, a multimodal regression framework that integrates ESM-2 sequence representations with structure-aware graph encoding via a GraphGPS backbone.
%combining GINE and Performer layers.
Antibody–antigen complexes are represented using residue-level C_alpha$ radius graphs with physicochemical and positional features, and modality-specific predictions are fused through a learned gated residual calibration mechanism. Under a strict 75% antigen-cluster split with grouped cross-validation and an independent holdout benchmark, GEPBind achieves state-of-the-art performance, with an RMSE of 1.569 and a Pearson correlation of 0.575, consistently surpassing sequence-based, structure-based variants, and established baselines. Ablation studies demonstrate that multimodal fusion, GraphGPS structure and consistent ESM-2 encoding across chains collectively contribute to robust out-of-cluster generalization.

A-P.10: CAPRINI-M: An AI-curated Cardiac-Specific Atlas of Protein Interactions in Mice
Track: Proteins and structural biology
  • Enio Gjerga, University Hospital Heidelberg, Germany
  • Philipp Wiesenbach, University Hospital Heidelberg, Germany
  • Chris-Andris Görner, Faculty of Informatics, Heilbronn University of Applied Sciences​, Germany
  • Ying Zhang, Faculty of Mathematics and Computer Science, Heidelberg University, Germany
  • Konstantin Pelz, Technical University of Munich, Germany
  • Markus List, Technical University of Munich, Germany
  • Christoph Dieterich, University Hospital Heidelberg, Germany


Presentation Overview: Show

Motivation: Protein–protein interactions are central to cardiovascular disease, yet relevant information is scattered across the literature and general databases, making it difficult to curate. Protein–protein interactions are fundamental to cardiovascular disease biology, but the corresponding knowledge is dispersed across the literature and heterogeneous databases, making systematic curation time-consuming. Moreover, many existing PPI resources may be biased and lack detailed information on structural interaction interfaces or associated thermodynamic parameters.
Results: We present CAPRINI-M (CArdiac PRotein INteractions In Mice), a web-based tool hosting an AI-curated atlas of cardiac protein interactions. We mined 9,105 cardiobiology manuscripts and used open-source LLMs (LLaMA-3.3 70b) to extract 11,189 protein–protein interactions. These were analysed with AlphaFold3 to infer interaction interfaces and thermodynamic properties linked to complex stability, and to estimate the likelihood that each protein pair forms a complex. Benchmarking showed CAPRINI-M outperformed general-purpose PPI resources in hypertrophy-focused edge-topology enrichment. Predicted interaction favourability also agreed with published experimental evidence, with lower predicted Gibbs free energy associated with experimentally preferred binding partners. Overall, CAPRINI-M provides more complete, mechanistically informative access to cardiovascular-disease–relevant protein-protein interactions by integrating literature evidence with structural interface and stability-aware annotations.
Availability: The CAPRINI-M web application is available at https://shiny.dieterichlab.org/app/caprinim. The source code used in this study is linked in the Availability section of the manuscripts.
Contact: E.Gjerga@uni-heidelberg.de
Supplementary information: Supplementary data are available at Bioinformatics online

A-P.11: A comparison of machine learning strategies for classification and clustering of transmembrane protein-protein interactions
Track: Proteins and structural biology
  • Lisa Allmesberger-Riegler, Department of Biotechnology and Food Science, University of Natural Resources and Life Sciences, Vienna, Austria
  • Fabian Frommelt, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria, Austria
  • Brianda Lopez Santini, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria, Austria
  • Evandro Ferrada, Instituto de Neurociencia, Facultad de Ciencias, Universidad de Valparaı́so, Valparaı́so, Chile, Chile
  • Giulio Superti-Furga, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria, Austria
  • Peter Sykacek, Department of Biotechnology and Food Science, University of Natural Resources and Life Sciences, Vienna, Austria


Presentation Overview: Show

Motivation: Protein-protein interactions (PPIs) are fundamental for cellular signaling, structural integrity, and the regulation of cellular homeostasis,
yet their diversity remains poorly understood. Aggregated physicochemical features are interpretable but ignore spatial context, whereas graph
neural networks (GNNs) preserve interaction topology at the cost of transparency. We systematically compare these representation strategies to test
whether GNN embeddings outperform interpretable physicochemical features for classifying PPI interfaces and to identify the key physicochemical
features driving TM interactions.
Results: We compare the embeddings for classifying TM PPI interfaces in 2 042 complexes which are experimentally confirmed, predicted using
AlphaFold (AF) and refined by molecular dynamics (MD) simulations. The aggregated physicochemical features and GNN embeddings achieved
comparable predictive accuracy, suggesting that preserving interaction topology provides no additional predictive power. Feature importance
analysis revealed a hierarchy dominated by amino acid composition, particularly hydrophobic and charged residue proportions, consistent with the
physicochemical constraints of membrane insertion. Energetic and structural features are secondary contributors. Unsupervised clustering revealed
that aggregated features capture richer biological structure than GNN embeddings. GNN-based clustering recovered the axis consistent with the
supervised training objective, while aggregated features additionally resolved interface size, energetics, and contact chemistry as distinct sources of
variation beyond the classification target.

A-P.12: Multi-Modal Protein Representation Learning with CLASP
Track: Proteins and structural biology
  • Nicolas Bolouri, McGill University, Canada
  • Joseph Szymborski, McGill University, Canada
  • Amin Emad, McGill University, Canada


Presentation Overview: Show

Effectively integrating data modalities pertaining to proteins' amino acid sequences, three-dimensional structures, and curated text-based descriptions of their biochemical and functional properties can lead to informative representations capturing different views of proteins. Here, we introduce CLASP, a unified tri-modal framework that combines the strengths of geometric deep learning, natural large language models (LLMs), protein language models (pLMs), and contrastive learning to learn informative protein representations based on their structure, amino acid sequence, and text-based biochemical and functional descriptions. We show that CLASP enables accurate zero-shot classification and retrieval tasks, such as matching a protein structure to its sequence or description, outperforming state-of-the-art baselines. CLASP embeddings also exhibit superior clustering by protein family, and ablation studies confirm that all three modalities contribute synergistically to performance. Our results highlight the power of integrating structural, sequential, and textual signals in a single model, establishing CLASP as a general-purpose embedding framework for protein understanding.

A-P.13: Structure-based virtual screening for TRPM8 modulators
Track: Proteins and structural biology
  • Nivya James, Imperial College London, United Kingdom
  • Pedro Ballester, Imperial College London, United Kingdom


Presentation Overview: Show

Since its discovery as the primary cold-sensing ion channel, the Transient Receptor Potential Cation Channel Melastatin Member 8 (TRPM8) has emerged as an attractive therapeutic target for cancer, pain, respiratory, and immune disorders. Although the underlying molecular mechanisms of temperature sensing and ligand recognition remain elusive, high-resolution structures of TRPM8 in agonist- and antagonist-bound conformations have been determined, enabling structure-based drug discovery approaches such as virtual screening (VS). However, systematic benchmarking of TRPM8 structural models and docking protocols has not yet been performed. Here, we establish a reproducible VS benchmark for TRPM8 by evaluating two high-resolution structures representing agonist-bound (TRPM87WRE) and antagonist-bound (TRPM89B6G) conformations using the open-source docking tools Smina and rDock with experimentally reported TRPM8 inhibitors and property-matched decoys. In addition, we investigated how different docking strategies can be combined to improve early hit identification.

rDock achieved higher hit rates and ranked true actives first among their corresponding decoy sets, whereas Smina showed stronger dependence on the target structure but provided better overall ranking quality across both conformations. Both docking tools displayed considerable overlap between active and decoy score distributions, indicating the limited discriminatory power of docking scores alone. When prioritizing a small subset of top-ranked compounds, integrated screening approaches improved recovery of true actives. Notably, the hierarchical protocol achieved performance comparable to consensus protocol while requiring substantially lower computational cost. Together, this work establishes the first systematic VS benchmark for TRPM8 and highlights the importance of integrated docking workflows for scalable hit discovery.

A-P.14: Beyond Lookup? Quantifying What LLMs Actually Contribute to Enzyme Function Prediction.
Track: Proteins and structural biology
  • Abishek Gnanasekaran, Ghent University Global Campus, South Korea
  • Ho-min Park, Ghent University Global Campus, South Korea
  • Sunjai Hwang, Ghent University Global Campus, South Korea
  • Wesley De Neve, Ghent University Global Campus, South Korea
  • Joris Vankerschaver, Ghent University Global Campus, South Korea


Presentation Overview: Show

Large language models are increasingly evaluated as tools for predicting protein function, raising the question of whether they can complement or replace established sequence-analysis methods. However, it remains unclear whether these models genuinely reason about biochemical evidence or simply retrieve memorized associations from their training data. Here we show that large language models perform substantially below classical baselines on enzyme function prediction and exhibit no detectable reasoning when challenged with counterfactual evidence. We benchmarked three models against BLASTp and CLEAN on 1,000 enzymes, finding that the best-performing model achieved 53% accuracy compared with 82% for BLASTp, and that all models collapsed further on a holdout set of 321 enzymes lacking close database homologs. A signal detection analysis of counterfactual perturbations — in which biological evidence was swapped between enzymes — yielded sensitivity indices near zero for all models, with Bayes Factor analysis providing moderate evidence that the models cannot distinguish genuine evidence changes from neutral reordering. These findings indicate that current large language models function as stochastic lookup systems rather than biochemical reasoners, and that a simple sequence-similarity search remains the more reliable tool for enzyme classification.

A-P.15: Bridging Scales: A Multi-Level Graph Neural Network for Protein Function Prediction
Track: Proteins and structural biology
  • Antoine Toffano, LIRMM, Univ. Montpellier, CNRS, Montpellier, France, France
  • Pierre Larmande, DIADE, Univ. Montpellier, IRD, CIRAD, Montpellier, France, France
  • Jérôme Azé, LIRMM, Univ. Montpellier, CNRS, Montpellier, France, France


Presentation Overview: Show

The exponential growth of protein sequence data has outpaced experimental functional characterization, resulting in a widening annotation gap. Addressing this requires both the discovery of novel roles and the refinement of existing Gene Ontology annotations. Current computational approaches to this problem operate either at the atomic residue scale or the systemic network scale, which prevents a holistic understanding of protein functional roles. To address this limitation, we propose MS-GNN, a graph neural network framework that bridges these scales. At the protein level, we construct spatial graphs of amino acids through language model embeddings and AlphaFold-derived structural contact maps. At the systemic level, these representations are integrated into a global protein–protein association network. This unified architecture achieves state-of-the-art performance across all three sub-ontologies. Beyond architectural innovation, we demonstrate that complementing strict experimental labels with broader, curated functional data drastically improves performance, showing that annotation sparsity, rather than algorithmic capacity, is the primary bottleneck in model improvement.
Finally, ablation studies confirm the necessity network-level information, particularly functional features, highlighting the necessity of multi-scale integration.

A-P.16: Ensembling Structure Model Outputs for Nearly Free Performance Gains in TCR-pMHC Interaction Prediction
Track: Proteins and structural biology
  • Fredo Guan, Arizona State University, United States
  • Heewook Lee, Arizona State University, United States


Presentation Overview: Show

TCR-pMHC interactions play a central role in adaptive immune recognition. Much recent work has focused on modeling these interactions using deep neural networks, trained to classify a TCR-pMHC pair as binding or non-binding. However, benchmark studies show that these models fail to generalize to unseen epitopes, exhibiting near-random performance. Conversely, several recent works leveraging protein structure models show that simple 0-shot predictors using confidence metrics or Rosetta binding energy can achieve higher performance on the binding prediction task, motivating further exploration. Here, we assess both zero-shot structure model outputs and ensembles of confidence metrics and Rosetta binding energy values on three different TCR-pMHC interaction prediction tasks, with ensembles achieving superior performance over zero-shot outputs. We further identify strong baseline ensembles for each task and demonstrate their superior performance on TCR-pMHC binding prediction for unseen epitopes compared to supervised predictive models.

A-P.17: A Unified Computational Framework for the Integration of AI Models in Structure-Based Drug Design
Track: Proteins and structural biology
  • Luna Pianesi, University of Bielefeld, Germany
  • Alexander Schoenhuth, University of bielefeld, Germany


Presentation Overview: Show

The rapid proliferation of sophisticated computational models for drug discovery has created unprecedented opportunities for innovation, yet the field lacks comprehensive frameworks to systematically integrate these diverse tools into coherent workflows. Although these models demonstrate remarkable individual capabilities, researchers are forced to navigate fragmented toolsets requiring extensive computational expertise, limiting the practical impact of these advances and creating an accessibility problem. This fragmentation is particularly pronounced at the intersection of molecular generation, molecular docking, and binding affinity prediction, where no single tool currently bridges all three stages in a unified manner. In this paper, we present a comprehensive and modular computational drug discovery pipeline that provides the first systematic framework for integrating diverse state-of-the-art models into an accessible unified drug discovery workflow. Our pipeline is designed with modularity as a core principle, enabling researchers to swap individual components as the field evolves without disrupting the broader workflow. The workflow is based on the integration of state-of-the-art generative and docking models, with a special focus on ensuring the synthetic accessibility and real world scenario plausibility of the proposed molecules. By lowering the barrier to entry for computational drug discovery, we aim to empower a broader community of researchers to leverage these powerful tools in pursuit of novel therapeutics.

A-P.18: BepiCon: A Geometric Deep Learning Framework for Conformational B Cell Epitope Prediction
Track: Proteins and structural biology
  • Bünyamin Şen, Dept. of Bioinformatics, Graduate School of Health Sciences, Hacettepe University, Ankara, Turkey, Turkey
  • Tunca Doğan, Dept. of Bioinformatics, Graduate School of Health Sciences, Hacettepe University, Ankara, Turkey, Turkey


Presentation Overview: Show

Accurate and reliable prediction of B cell epitopes holds critical importance in immunology and vaccine development. While traditional experimental methods offer high accuracy in identifying epitope regions, they are often laborious, time-consuming, and costly. Therefore, attempts are made to increase the efficiency of experimental characterization processes by using computational approaches. Since approximately 90\% of epitopes are conformational, the prediction processes must account for three-dimensional protein structures and the geometric details of antigen-antibody interactions. In response to these requirements, our study introduces BepiCon, a two-stage geometric deep learning framework that models antigen proteins as graph structures, incorporating structural and physicochemical properties and protein language model embeddings to predict epitope regions on antigen proteins. In the first stage, the model was trained using a graph contrastive learning approach to learn high-quality representations of epitope and non-epitope residues. In the second stage, the pre-trained model was fine-tuned using supervised learning to perform conformational epitope prediction. The developed framework has demonstrated effective and generalizable performance when applied to both experimentally determined protein structures and predicted structures. Comparative analysis revealed that our approach distinguishes itself from existing B cell epitope prediction methods by exhibiting a lower false-positive rate and generating more reliable predictions. Our work contributes significantly to scientific research and therapeutic design processes by showcasing the advantages of geometric deep-learning approaches in B-cell epitope prediction.

A-P.19: Leakage-controlled evaluation of single-to-multi mutation generalization in fungal azole resistance prediction
Track: Proteins and structural biology
  • Yuan Fei, Shanghai Jiao Tong University School of Medicine, China
  • Thanh Thien Le, VinUniversity, Viet Nam
  • Io Hong Cheong, Shanghai Jiao Tong University School of Medicine, China
  • Zisis Kozlakidis, World Health Organization, France
  • Jiaqi Yin, Northwestern Polytechnical University, China
  • Yin Wu, Shanghai Jiao Tong University School of Medicine, China
  • Xiaoqi Zheng, Shanghai Jiao Tong University School of Medicine, China
  • Yang Yang, Shanghai Jiao Tong University School of Medicine, China


Presentation Overview: Show

Motivation: Clinical azole resistance in fungi often involves combinations of mutations, whereas labeled higher-order mutants (k ≥ 2 substitutions) remain scarce. This creates a realistic setting in which models trained on single mutants for azole resistance prediction must generalize to higher-order mutants. A central question is whether apparent generalization reflects genuine transfer or is driven by overlap between substitutions seen in single mutants and those present in held-out higher-order mutants.
Methods: Using FungAMR, we evaluated PLM-based and structured feature representations with aggregation baselines, tabular machine-learning models, an MLP, and gated Two-Tower models under three protocols: single-to-multi evaluation, a leakage-controlled variant that removes constituent substitutions overlapping with held-out higher-order mutants, and cross-species transfer from C. albicans to non-C. albicans fungi.
Results: We established a single-to-multi evaluation protocol in which models are trained on single mutants and evaluated on 51 held-out higher-order mutants. Ranking performance was strong, with the best configuration reaching PR-AUC 0.944. To test whether this performance was inflated by constituent overlap, we introduced a leakage-controlled protocol that removed single-mutant substitutions appearing in held-out higher-order mutants from the training-validation pool. Representative end-to-end models remained stable under this control, with PR-AUC shifts ≤ 0.01, whereas aggregation baselines deteriorated markedly. Cross-species evaluation showed measurable zero-shot transfer from C. albicans to non-C. albicans fungi, with pooled PR-AUC 0.711. Together, these results suggest that leakage-controlled evaluation helps distinguish genuine generalization to higher-order mutants from constituent-overlap shortcuts in fungal resistance prediction.
Availability and implementation: Source code is available at https://github.com/asunafy/fungamr-paper.

A-P.20: PUMA: Discovery of Protein Units via Mutation-Aware Merging
Track: Proteins and structural biology
  • Burak Suyunu, Computer Engineering, BoÄŸaziçi University, Bebek, 34342, İstanbul, Türkiye, Turkey
  • Özdeniz Dolu, Computer Engineering, BoÄŸaziçi University, Bebek, 34342, İstanbul, Türkiye, Turkey
  • Ibukunoluwa Abigail Olaosebikan, C. Eugene Bennett Department of Chemistry, West Virginia University, Morgantown, West Virginia 26505, United States, United States
  • Hacer Karatas Bristow, C. Eugene Bennett Department of Chemistry, West Virginia University, Morgantown, West Virginia 26505, United States, United States
  • Arzucan Ozgur, Computer Engineering, BoÄŸaziçi University, Bebek, 34342, İstanbul, Türkiye, Turkey


Presentation Overview: Show

Motivation: Proteins are the essential drivers of biological processes. At the molecular level, they are chains of amino acids that can be viewed through a linguistic lens where the twenty standard residues serve as an alphabet combining to form a complex language, referred to as the language of life. To understand this language, we must first identify its fundamental units. Analogous to words, these units are hypothesized to represent an intermediate layer between single residues and larger domains. Crucially, just as protein diversity arises from evolution, these units should inherently reflect evolutionary relationships. We introduce PUMA (Protein Units via Mutation-Aware Merging) to discover these meaningful units. PUMA employs an iterative merging algorithm guided by substitution matrices to identify protein units and organize them into families linked by plausible mutations. This process creates a genealogy where units and their mutational variants coexist, simultaneously producing a unit vocabulary and a hierarchy connecting them.

Results: We validate PUMA's biological relevance through three key findings. First, mutations occurring within a PUMA family are significantly more likely to be clinically benign than pathogenic. Second, by leveraging its genealogy, PUMA surpasses frequency-based baselines in mapping protein units to functional annotations. Finally, case studies demonstrate that PUMA families preserve essential physicochemical properties and functional motifs. By explicitly linking mutational variants, PUMA provides an evolutionarily grounded and interpretable vocabulary for deciphering the language of life.

Availability and implementation: The source code is available at https://github.com/boun-tabi-lifelu/PUMA.

A-P.21: SpeciefAI: Multi-species mRNA-level Antibody Framework Generation using Transformers
Track: Proteins and structural biology
  • Dominik Grabarczyk, University of Edinburgh, United Kingdom
  • MikoÅ‚aj Kocikowski, Independent Researcher, Poland
  • Maciej Parys, University of Edinburgh, United Kingdom
  • Shay Cohen, University of Edinburgh, United Kingdom
  • Javier Alfaro, University of Calgary, Canada


Presentation Overview: Show

Motivation: Encoding antibodies (Abs) and nanobodies (Nbs) as mRNA enables in vivo production of therapeutic proteins. However, this approach requires meeting two species-dependent requirements: the mRNA encoding must support efficient expression in the host species, and the encoded protein sequence must resemble the natural Ab repertoire of the recipient species to minimize immunogenicity. These requirements motivate species-conditioned generative models for joint mRNA and protein design.

Results: We propose SpeciefAI a transformer-based model for multi-species Ab and Nb species sequence-harmonisation by generation of novel Framework Regions (FRs) tailored to input Complementarity-Determining Regions (CDRs). Our model works directly in the mRNA space and learns the correspondence between FRs and CDRs in six species. The model is capable of generating sequences with a highly similar distribution to natural sequences and a mean absolute difference in codon adaptation index (CAI) of 0.013 and 0.033 for humans and dogs respectively. We show that the generated human sequences are highly human (0.95 T20 score) and canine sequences highly canine (0.95 cT20 score). We furthermore demonstrate that we can generate diverse candidate sequences using our method.

Availability and Implementation: Source code is available on https://github.com/Dominko/SpeciefAI. OAS and COGNANO data are publicly available on https://opig.stats.ox.ac.uk/webapps/oas/ and https://cognanous.com/datasets/vhh-corpus (preprocessed versions available upon request). Canine data is available on https://zenodo.org/records/18301526

A-P.22: WACA-DTA: Water-Aware Geometric Biases for Structure-Conditioned Drug-Target Affinity Prediction
Track: Proteins and structural biology
  • Kehan Huang, China Pharmaceutical University, China
  • Chang Li, Tsinghua University, China
  • Jing Ji, China Pharmaceutical University, China
  • Zimo Tang, China Pharmaceutical University, China


Presentation Overview: Show

Drug-target affinity (DTA) prediction is central to computational drug discovery, yet many structure-aware models still weakly constrain geometric
locality and solvent mediation at the binding interface. We present WACA-DTA, a structure-conditioned affinity model that reformulates interface
matching as pose-conditioned atom-residue cross-attention with factorized interaction logits. Each cross-interface score is decomposed into
direct, geometric, and hydration-mediated terms, with geometry and hydration injected as additive logit-level priors rather than post hoc feature
concatenations. Under a fixed pair-pose protocol and identical structural preprocessing, WACA-DTA improves over a matched pocket-aware
baseline on Davis and KIBA across drug, target, and pair affinity cold-start splits, while controlled ablations across Davis, KIBA, and PDBbind
show that the gains are most consistent when pair-specific geometric and hydration cues are injected directly into the interaction logits. Our
claims are restricted to input-matched comparisons and within-dataset evaluation, and the hydration interpretation is cross-checked on PDBbind
complexes with retained crystallographic waters. Code, split definitions, and preprocessing manifests are available at https://github.com/khan1
14514/WACA-DTA2.

A-P.23: Swiss-PO. Advancing Cancer Mutation and Structural analysis for Precision Oncology with the Latest Release
Track: Proteins and structural biology
  • Fanny Krebs, University of Lausanne, Ludwig Institute for Cancer Research Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
  • Olivier Michielin, HUG, SIB Swiss Institute of Bioinformatics, Switzerland
  • Vincent Zoete, SIB Swiss Institute of Bioinformatics, University of Lausanne, Ludwig Institute for Cancer Research Lausanne, Switzerland


Presentation Overview: Show

The rapid expansion of precision oncology has led to a substantial increase in identified oncodriver genes and associated variants. This growth has also increased the number of mutations with unknown functional impact, highlighting the need for molecular modeling tools to support variant interpretation. Swiss-PO has been expanded and redesigned to integrate large-scale oncogenomic data with structure- and sequence-based analytical tools. The platform combines oncodriver gene data with multiple sequence alignments (MSAs) for conservation analysis across orthologs and gene families, as well as visualization of experimental and predicted three-dimensional (3D) protein structures to assess the potential effects of mutations. Additional modules include a BRAF kinase mutation classification model and a dedicated page for ligands targeting proteins, alongside integrated links to external databases to enable multidimensional analyses.

The updated Swiss-PO platform integrates data for nearly 1,500 oncodriver genes, covering more than 3 million mutations and post-translational modification annotations, over 26,000 experimental and predicted 3D protein structures, more than 4,000 MSAs, and information on over 200,000 protein ligands. The platform also includes a BRAF kinase mutation class predictor to support therapeutic decision-making. With these enhancements, Swiss-PO provides a comprehensive resource for oncologists, bioinformaticians, and molecular biologists involved in interpreting cancer-associated mutations and advancing precision medicine. The website is available at https://swiss-po.ch/

A-P.25: FusionPath: Gene fusion pathogenicity prediction using protein structural data and contextual protein embeddings
Track: Proteins and structural biology
  • Nadine Sina Kurz, University Medical Center Göttingen, Germany
  • Irem Berna Güven, University Medical Center Göttingen, Germany
  • Tim Beissbarth, University Medical Center Göttingen, Germany
  • Jürgen Dönitz, University Medical Center Göttingen, Germany


Presentation Overview: Show

Accurate prediction of gene fusion pathogenicity is critical for understanding oncogenic mechanisms and advancing precision oncology. While existing computational methods provide valuable insights, their performance remains limited by incomplete integration of multi-scale biological features and insufficient model interpretability for clinical translation. We present FusionPath, a novel deep learning framework for gene fusion pathogenicity prediction that addresses these limitations through multimodal integration of complementary biological data. FusionPath uniquely integrates embeddings from multiple pretrained protein language models, including FusON-pLM and ProtBERT, alongside retained protein domains and Gene Ontology (GO) functional annotations. The model was trained and validated on a rigorously curated dataset of more than 70,000 gene fusions derived from FusionPDB, ChimerDB4.0, and 27 RNA-seq datasets of normal tissues. FusionPath significantly outperformed state-of-the-art methods and provides interpretable insights through SHAP analysis, revealing cancer-type-specific pathogenicity patterns and identifying protein kinase domains as key determinants of oncogenic potential. By synergistically leveraging sequence, structural, and functional information with explicit modeling of wild-type sequence context, FusionPath yields biologically grounded pathogenicity scores with mechanistic insights.

A-P.26: Elucidating the Structural Dynamics and Binding Mechanism of ProQ-raiZ Using an ML based Deep-TDA Approach for Enhanced Sampling Simulations: Implications for Novel Antimicrobial Targeting
Track: Proteins and structural biology
  • Shilpi Singh, Indian Institute of Technology Delhi, India
  • Nisha Kumari, Indian Institute of Technology Delhi, India
  • Tarak Karmakar, Indian Institute of Technology Delhi, India
  • Tanmay Dutta, Indian Institute of Technology Delhi, India


Presentation Overview: Show

ProQ, a bacterial RNA chaperone, plays a crucial function in post-transcriptional gene regulation which makes it a potential target for combating antibiotic resistance. The N-terminal domain (NTD) is recognised for its role in RNA binding; however, the structural dynamics and energetic landscape of its interaction with small RNAs that play roles in virulence, such as raiZ, remain poorly explored. This research uses an advanced computational framework to delineate the ProQ-raiZ complex dynamics, establishing a structural foundation for novel therapeutic strategies.

We employed molecular dynamics simulations, combining spontaneous binding analyses, to identify the protein-RNA interface. We utilised an advanced enhanced sampling technique like On-the-fly Probability-Enhanced Sampling (OPES) to get a free energy landscape for the binding mechanism. We utilised Deep- Targeted Discriminant Analysis (Deep-TDA), a machine learning methodology to find non-linear features inside the conformational space and to formulate appropriate Collective Variables (CVs) for more effective sampling.

Our findings prove that raiZ maintains a semi-folded conformation upon interacting with the concave surface of ProQ, reinforced by a network of electrostatic contacts throughout the pocket in NTD. In silico mutagenesis revealed key residues critical for complex stability; their alteration results in raiZ dissociation, suggesting these locations as high-priority ""hotspots"" for small-molecule inhibition. This study integrates machine learning with physics-based Molecular dynamics simulations to provide a high-resolution picture of ProQ dynamics. These insights provide a framework for engineering compounds that can interfere with RNA-based regulatory pathways, exemplifying a strong computational approach for novel drug discovery.

A-P.27: Assessing Long-distance Epistatic Interactions in Enzymes
Track: Proteins and structural biology
  • Mykolas Malevicius, Institute of Biomedical Informatics, Technical University, Graz, Austria
  • Jasmin Zuson, Institue of Molecular Biotechnology, Technical University, Graz, Austria
  • Gerhard G. Thallinger, Institute of Biomedical Informatics, Technical University, Graz, Austria
  • Leila Taher, CUBiDA, Medical Center for Information and Communication Technology, Universitätsklinikum Erlangen, Germany


Presentation Overview: Show

Deep mutational scanning is a technique that uses next-generation sequencing to study the functional impact of thousands of mutations in an enzyme in a single experiment. Due to epistasis, the functional impact of multiple mutations is non-additive, their effects must be examined together. To capture such effects, the entire molecule must be sequenced as a single read, which necessitates long-read sequencing approaches such as that provided by Oxford Nanopore Technologies (ONT).
While ONT enables long-read sequencing, its inherently high error rate of approximately 2% hinders exact variant identification. Additionally, multiple PCR rounds during library preparation can introduce chimeras, artificial fusions of distinct DNA molecules, that further complicate data analysis. To combat these discrepancies, we designed and implemented a computational workflow that leverages unique molecular identifiers to construct consensus sequences (CSs) and correct errors.
We analysed two datasets, comprising of ~5M and ~10M reads, and found that chimeras can be prominent, reaching 33.4% in one of the datasets. Furthermore, data pre-processing resulted in read losses of 39% and 54%, coupled with per-CS error rates of 9.32, 0.38, 0.09 and 0.04 in sequences constructed from 2, 3, 4 and 5 reads, respectively. Therefore, at least five reads are necessary to build a reliable CS. Moreover, more than half of the molecules were only sequenced once, increasing the number of uninformative sequences. Currently, we are designing a deep learning approach to classify the substructures of the amplicons at signal-level, which will enable categorisation in real time.

A-P.28: Computational Design of Anti-VEGF Binders: From Scaffold Optimization to De Novo Design
Track: Proteins and structural biology
  • Georg Kuenze, Institute for Drug Discovery, Leipzig University, Germany
  • Abibe Useini, Institute for Drug Discovery, Leipzig University, Germany
  • Carsten Geist, Fraunhofer Institute for Cell Therapy and Immunology - IZI, Leipzig, Germany
  • Stefan Kalkhof, Fraunhofer Institute for Cell Therapy and Immunology - IZI, Leipzig, Germany


Presentation Overview: Show

Anti-VEGF inhibitors are central to the treatment of angiogenesis-driven diseases such as cancer and age-related macular degeneration; however, current antibody-based therapies are limited by their size, production cost, and delivery constraints. Miniproteins represent a promising alternative, combining high target specificity with enhanced stability and manufacturability. Here, we report the computational design and experimental validation of anti-VEGF miniprotein binders across two development stages.
Our methodology integrates (i) structure-guided redesign of the Z12 miniprotein scaffold using ProteinMPNN to improve stability and VEGF inhibition, (ii) de novo binder generation using BindCraft to explore novel structural solutions, and (iii) experimental validation using surface plasmon resonance and a cell-based VEGF receptor bioassay.
Optimized Z12-derived variants achieved an IC50 of 0.15 µM, corresponding to an approximately 10-fold improvement in inhibitory potency compared to the parent scaffold, while also exhibiting enhanced thermal stability with Tm increases of 10-17°C. In parallel, BindCraft-designed de novo binders reached sub-micromolar binding affinities and demonstrated strong functional inhibition, with top candidates reducing VEGF activity to ~4% at 1 µM, indicative of potent antagonistic activity. X-ray crystallographic analysis of the top-performing binders confirmed close agreement between the computational models and experimentally determined structures.
Overall, this work demonstrates that combining scaffold optimization with de novo design enables the rapid development of potent and stable anti-VEGF miniproteins for anti-angiogenic applications. These findings establish a foundation for next-generation biologics targeting angiogenesis-related diseases and highlight the potential of AI-driven protein design in therapeutic discovery.

A-P.29: CCD2MD: A suite of Packages for Preparing Co-Folded Outputs for Molecular Dynamics Simulations
Track: Proteins and structural biology
  • Katarina Blow, University of Warwick, United Kingdom
  • Matyas Parrag, University of Warwick, United Kingdom
  • Phillip Stansfeld, University of Warwick, United Kingdom


Presentation Overview: Show

Protein/lipid interactions play a crucial role in the stability and function of membrane proteins. While experimental approaches to characterise these interactions in a native-like membrane environment can be challenging, computational techniques offer a powerful alternative for identifying and analysing potential binding sites. Recent advances in co-folding methods (such as AlphaFold[1]) now enable the prediction of holo protein structures, capturing conformational changes that may occur upon lipid binding and thereby improving the accuracy of binding site characterisation. However, the outputs from these methods often require post-processing to ensure compatibility with widely used molecular dynamics force fields. Here, we present CCD2MD[2], a modular toolkit designed to convert co-folding outputs into simulation-ready systems. CCD2MD supports both atomistic and coarse-grained representations, with optional membrane embedding facilitated via MemPrO[3]. The modular design of CCD2MD allows for straightforward adaptation to other co-folded biomolecular assemblies, incorporating complexes with nucleic acids, small molecules, carbohydrates, or metal ions, thereby enabling a variety of simulation setups across multiple scales. We also discuss recent updates and improvements to the code.

A-P.30: Computational Mutational Scanning of DHFR for Mutation-Dependent Inhibitor Preference
Track: Proteins and structural biology
  • Busra Tayhan, Sabanci University, Turkey
  • Ebru Cetin, Sabanci University, Turkey
  • Tandac Furkan Guclu, Istinye University, Turkey
  • Muhammed Sadik Yildiz, University of Texas Southwestern Medical Center, United States
  • Erdal Toprak, University of Texas Southwestern Medical Center, United States
  • Ali Rana Atilgan, Sabanci University, Turkey
  • Canan Atilgan, Sabanci University, Turkey


Presentation Overview: Show

Antibiotic resistance is a public health challenge driven by mutations that enable bacteria to maintain function while escaping drug pressure. One enzyme extensively studied in this context is dihydrofolate reductase (DHFR), an essential catalyst in folate metabolism and nucleotide biosynthesis. This protein is suitable for studying antibiotic resistance because resistance often arises from single amino-acid substitutions that preserve function while altering inhibitor binding. In this work, we investigated the mutational landscape of E. coli DHFR in the presence of competitive inhibitors, trimethoprim(TMP) and its derivative 4′-deoxytrimethoprim(4′-DTMP). Because TMP and 4′-DTMP are closely related chemically, comparing them allows us to assess whether a structural modification can reshape mutation-dependent inhibitor preference and alter high-impact resistance trajectories. To characterize mutation-dependent energetic effects, we generated a deep mutational scanning library. Large-scale free-energy perturbation(FEP) simulations were performed for apo, TMP-bound, and 4′-DTMP-bound DHFR to estimate mutational free energies and relative binding preferences. Single amino-acid substitutions were examined across DHFR positions with triplicate FEP calculations. Although TMP and 4′-DTMP showed similar behavior across much of the scanned landscape, the strongest differences were concentrated at important DHFR regions, including cryptic-site and allosteric regions, demonstrating that even a subtle chemical modification can reshape inhibitor preference at sites most relevant to resistance dynamics and function. These findings provide a structural-energetic map of mutation-dependent inhibitor preference and highlight the potential of alchemical calculations for mutation-resolved studies of antibiotic resistance. Results were interpreted alongside experimental relative fitness data, with explicit incorporation of uncertainty from both computational and experimental sources.

A-P.31: An efficient integer programming model for RNA secondary structure prediction
Track: Proteins and structural biology
  • Olga Karelkina, Systems Research Institute Polish Academy of Sciences, Poland


Presentation Overview: Show

RNA secondary structure prediction problem is commonly addressed through dynamic programming, context-free grammar, and, more recently, machine learning techniques. We present an alternative integer programming (IP) model, based on a classical hydrodynamics hypothesis. As an energy-directed method, it relies on Turner's nearest-neighbor rules to evaluate the free energy of candidate structures. To predict secondary structure, the model explores all feasible configurations of two-dimensional structural motifs, defined through binary variables and linear constraints. Furthermore, the proposed IP model can be extended with additional structural constraints to generate suboptimal solutions.

While the IP approach offers high modeling flexibility, it is time-consuming, which can limit its practical applicability. We introduce a compact formulation with a dedicated set of constraints to model multibranch loops, which significantly reduces complexity and computation time, enabling analysis of longer and more challenging sequences.

The model is benchmarked on the Archive II dataset and compared against leading dynamic programming methods. The results demonstrate competitive overall accuracy, and, for certain sequences, the model produces solutions that more closely match native structures than those obtained by traditional methods.

A-P.32: Novel universal domain-centric method for protein classification
Track: Proteins and structural biology
  • Shakiba Fadaei, Student, university of Lausanne, Switzerland
  • Fanny Krebs, University of Lausanne / SIB, Switzerland
  • Vincent Zoete, University of Lausanne / SIB, Switzerland


Presentation Overview: Show

Human protein kinases constitute a large superfamily of about 500 genes, historically classified into subfamilies based on phylogenetic relationship. However, many kinases remain unclassified. Phylogeny is typically based on multiple sequence alignments, and neglects the physico-chemical properties of residues at each position of the sequence. By incorporating these properties, we can gain deeper insights beyond basic alignments. Here we use, for the first time, a detailed physico-chemical description of kinases to identify class-specific structural regions, supporting an unsupervised classification method capable of classifying previously unlabeled kinases. This novel approach aligns with existing phylogeny-based classifications while offering refinements and enhanced accuracy. Ultimately, we use machine learning techniques to classify unlabeled kinases, validated by analyzing class-specific structural regions. This new classification approach goes beyond current rankings and can be applied to any type of protein, such as immunoglobulins and G protein-coupled receptors.

A-P.33: BrightDB : Brightness Protein Database with fast accession to millions of proteins and an interactive dashboard
Track: Proteins and structural biology
  • Damien Legros, VIB.AI/KULeuven, Belgium
  • Joana Pereira, VIB.AI/KULeuven, Belgium


Presentation Overview: Show

Storing data efficiently is a fundamental requirement for performing large-scale bioinformatics analysis. However, as the genomic and proteomic data grows exponentially, traditional relational databases often struggle with multi-modal information and random access. To address these persistent bottlenecks, we present BrightDB, an innovative local aggregator database designed to centralize and manage data from multiple public databases (such as UniProt and InterPro), results from models, and experiments findings by leveraging the cutting-edge capabilities of modern data lakehouse architecture.
BrightDB bridges the gap between the flexibility of unstructured data lakes and the management of features in traditional databases. By utilizing high-performance and columnar-based novel Python libraries, BrightDB centralizes multi-modal biological data into a unified, optimized file format. This columnar approach allows for substantial data compression and significant speed improvements in data retrieval, directly facilitating the generation of datasets for model training, statistical analysis, and complex visualization.
Beyond its backend efficiency, BrightDB also offers a real-time interactive dashboard. This local interface allows users to perform near-instantaneous searches across millions of entries, providing immediate access to the available information for each protein. By removing the technical difficulties between data storage and visualization, BrightDB significantly reduces the time-intensive data management. Ultimately, this architecture provides a scalable, future-proof local aggregator database solution for the proteomics community, ensuring that all types of data modalities from legacy formats to emerging ones remain accessible inside a unified storage environment.

A-P.34: Reference-Free Ranking Method for RNA 3D Models
Track: Proteins and structural biology
  • Jan Pielesiak, Poznan University of Technology, Poland
  • Maciej Antczak, Poznan University of Technology, Polish Academy of Sciences, Poland
  • Marta Szachniuk, Poznan University of Technology, Polish Academy of Sciences, Poland
  • Tomasz Zok, Poznan University of Technology, Poland


Presentation Overview: Show

In current times, researchers have started to notice the importance of understanding the structure and function of RNA. This prompts the growth of the number of attempts in RNA 3D structure prediction. As the knowledge of RNA molecules grows, we can use the advancements made in protein structure prediction to improve our prognoses in the RNA field; however, it faces a challenge – produced models need to have their quality assessed. Multiple or sometimes even thousands of models can be generated from a single input. The traditional method for determining the quality of the model is based on energy terms calculation with force fields or coarse-grained statistical potentials. The main issue with such an approach is that the energy landscape usually contains many local minima, which leads to inconclusive results.
Therefore, we propose a different approach for ranking multiple 3D models of the same RNA sequence. The basis for this algorithm is the analysis of the base pairs and stacking interactions within them. A consensus secondary structure is built from the extracted data, and each model has its interaction network ranked against the aforementioned consensus to provide a final ranking.
Our method has been benchmarked on datasets used in RNA 3D modeling to verify its quality in comparison to the state-of-the-art energy-based evaluations. The entire system has been made publicly available at https://rnative.cs.put.poznan.pl/, and published in Bioinformatics (https://doi.org/10.1093/bioinformatics/btaf601)

A-P.35: TmProt 1.0: ML-based Tool for Protein Melting Temperature Prediction with Cross-Method Validation on Biophysical Data
Track: Proteins and structural biology
  • Karen Pailozian, Masaryk univevrsity, Czechia


Presentation Overview: Show

Protein melting temperature (Tm) prediction accelerates the discovery of thermostable enzymes crucial for industrial biotechnology, where proteins must endure harsh reaction conditions. Experimental determination of Tm remains labour-intensive and varies across techniques, motivating the development of in silico predictors. Recent advances in high-throughput mass spectrometry have led to large-scale proteomics-based Tm datasets enabling effective training of machine learning models. However, the generalisability of such tools across diverse proteomics- and biophysics-based datasets remains an open question.
This study aims to (i) assemble comprehensive Tm datasets including high-quality biophysics sources for independent evaluation, (ii) evaluate generalisability of state-of-the-art approaches based on deep learning, ESM-2 sequence embeddings, and parameter-efficient low-rank adaptation (LoRA), and (iii) explore ESM-3 structural embeddings as a counterpart to sequence-based approaches.
We assembled the ProMelt dataset (45,441 proteins) from Meltome Atlas and ProThermDB and trained baseline and advanced embedding-based predictors. Models were evaluated on five independent biophysics-based datasets derived from BRENDA, FireProtDB, and literature.
Our analysis revealed substantial inconsistencies in reported Tm values between proteomics- and biophysics-based measurements, highlighting the need for generalizable predictors. Fine-tuned embedding-based models showed competitive performance compared to DeepSTABp, TemBERTure, and SaProt and achieved superior performance in binary classification of thermostable proteins (Tm ≥ 60 °C): ESM2-LoRA achieved AUC = 0.75, while ESM3-MLP achieved AUC = 0.77.
The best-performing model is deployed as TmProt, a user-friendly web server on Hugging Face.
Together, these results show that LoRA-adapted ESM-2 and ESM-3 structural embeddings improve thermostability prediction and emphasize the importance of rigorous evaluation across diverse experimental methodologies.

A-P.36: Accelerating Multiple Myeloma Diagnosis with Proteomics & Machine Learning
Track: Proteins and structural biology
  • Anna Melidi, Bispebjerg Hospital, Denmark


Presentation Overview: Show

Importance: Multiple myeloma (MM), a blood cancer arising from malignant plasma cells in the bone marrow, is complex and time-consuming to diagnose, requiring several laboratory tests, which can delay treatment increasing progression risk.

Objective: To develop a diagnostic test based on untargeted mass spectrometry(MS)-based proteomics and machine learning (ML) to diagnose MM already at the first point of screening.

Design, Setting, and Participants: This is a prospective cohort study involving 2,501 samples (serum, peripheral plasma, bone marrow plasma) obtained from patients diagnosed with MM, monoclonal gammopathy of undetermined significance (MGUS), smoldering MM (SMM), patients referred for M-component screening, and healthy controls.

Methods: Proteomes were quantified by LC-MS and analyzed with supervised ML to identify diagnostic signatures. Performance was evaluated by cross-validation and external validation, with correlation to disease severity, survival, and progression assessed.

Results: The model achieved AUCs of 0.85 (CV-test), 0.91 (national validation cohort), and 0.79 (international validation cohort). The predicted disease score showed a continuous dis-tribution correlating with disease severity from controls to MGUS to SMM to MM. Correla-tion between the LC-MS-derived estimate of monoclonality and conventional M-spike quan-tification was 0.78 (p=2×10⁻⁶¹). Classification performance was 95% in healthy peripheral plasma and 95% (peripheral) and 83% (bone marrow) in newly diagnosed MM. We found no significant survival differences between upper and lower quartiles of the disease score in newly diagnosed MM.

Conclusions and Relevance: Combining proteomics with ML allows diagnostic support and severity assessment in MM and potentially other medical conditions.

A-P.37: Alpha&ESMhFolds: An updated database for the comparison and functional annotation of predicted structural models for the human proteome
Track: Proteins and structural biology
  • Matteo Manfredi, Biocomputing Group, University of Bologna, Italy
  • Gabriele Vazzana, Biocomputing Group - University of Bologna, Italy
  • Castrense Savojardo, Biocomputing Group, University of Bologna, Italy
  • Pier Luigi Martelli, Biocomputing Group, University of Bologna, Italy
  • Rita Casadio, University of Bologna, Italy


Presentation Overview: Show

The adoption of predicted models of protein structures generated by methods such as AlphaFold and ESMFold is ever-increasing. We release an updated version of a public database available at https://alpha-esmhfolds.biocomp.unibo.it/, storing pairs of models for 48,815 human proteins enriched with functional characterization.

This release, synchronized with the latest updates of resources like UniProt, PDB, AlphaFold DB, and Pfam, introduces new functionalities. We extract for each protein Pfam annotations and known pathogenic variants. Both are mapped on the predicted structures and can be visualized from the web server. Moreover, we complement the structural superimposition of AlphaFold2 and ESMFold models with an external validation performed by three state-of-the-art Quality Assessment tools, providing a consensus to suggest the best model for each protein.

Exploiting those functionalities, we perform large-scale analysis to identify two interesting trends. First, both AlphaFold2 and ESMFold are consistently better at predicting regions of the proteins covered by Pfam entries. Not only do both methods exhibit higher average pLDDT scores in those regions when compared to the average of the full models, but their agreement, measured by the TM-score upon structural superimposition, also increases from 0.58 to 0.88.

Second, focusing on the subset of proteins for which the two methods produce diverging models, we observe that the consensus of external QA tools assigns better scores to AlphaFold for only 51% of the cases, supporting the usefulness of having different models available to exploit the strengths of each method, particularly when the models differ.

A-P.38: Exploring the Proteomics Dark Matter: Non-Canonical Peptide Identification and the Quantification of all Precursor Ions
Track: Proteins and structural biology
  • Dominik Lux, Ruhr University of Bochum / Medical Proteome Center, Germany
  • Svitlana Rozanova, Ruhr University of Bochum / Medical Proteome Center, Germany
  • Britt Mollenhauer, University Medical Center Goettingen / Department of Neurology / Kassel / Paracelsus-Elena Klinik, Germany
  • Katalin Barkovits, Ruhr University of Bochum / Medical Proteome Center, Germany
  • Julian Uszkoreit, Ruhr University of Bochum / Medical Bioinformatics, Germany
  • Katrin Marcus, Ruhr University of Bochum / Medical Proteome Center, Germany
  • Martin Eisenacher, Ruhr University of Bochum / Medical Proteome Center, Germany


Presentation Overview: Show

Most analysis workflows for LCMS proteomics data captured with data-dependent acquisition use search engine identification to match MS2 spectra against known protein databases. These identification-first approaches then utilize precursor ions of identified MS2 spectra to retrieve quantitative values for downstream analysis. While these workflows are well-established, most precursor ions are not or cannot be identified and are systematically not reported. We define this unused set as the proteomics dark matter. In this work, we demonstrate that some unidentified precursors can be identified and all of them quantified regardless of their identification status. To address this, we developed two workflows: ProtGraph and UnbeQuant. ProtGraph utilizes protein graphs considering protein variations, mutations and specific cleavage points to identify non-canonical peptides. Complementary, UnbeQuant enables the quantification of precursor ions regardless of identification status, allowing for quantitative comparisons across measurements. Using a cerebrospinal fluid dataset across two groups and three time points, we demonstrate that downstream analysis with the proteomics dark matter is possible by highlighting un-/identified and identified non-canonical precursor ions between the groups and time points. Further, by using un-/identified precursor ions, we show that unconsidered post-translational modifications can be discovered. Finally, we demonstrate identification ratios across different sample types including blood, cell culture, or stool to illustrate the potential of the proteomics dark matter in different datasets. Ultimately, we show that identification-first approaches may only report limited results of data at hand and encourage researchers to explore the proteomics dark matter for a more comprehensive analysis of their data.

A-P.39: Characterization of SLC25A38 as a Mitochondrial Pyridoxal 5′-Phosphate Transporter
Track: Proteins and structural biology
  • Miriana Quaranta, Sapienza University of Rome, Italy
  • Stefano Pascarella, Sapienza University of Rome, Italy


Presentation Overview: Show

Pyridoxal 5′-phosphate (PLP), the active form of vitamin B6, is an essential cofactor for mitochondrial enzymes involved in amino acid metabolism and heme biosynthesis. Despite its central role in cellular homeostasis, the mechanisms regulating PLP import into mitochondria remain poorly understood. The human mitochondrial carrier SLC25A38, previously annotated as a glycine transporter, has been proposed to mediate PLP transport. Indeed, silencing of SLC25A38 selectively reduces mitochondrial PLP levels (Pena et al., 2025). Moreover, several point mutations are associated with pyridoxine-refractory sideroblastic anemia (SIDBA2).

We present a computational study to structurally characterize SLC25A38 as a mitochondrial PLP transporter. Structural model was generated using AlphaFold3. Binding pocket was identified using PrankWeb and validated by comparison with homologous known proteins. Ligand binding was predicted through in silico molecular docking of PLP and glycine using GNINA The apo and holo complexes were embedded in a phospholipid bilayer representative of the inner mitochondrial membrane and subjected to 1 μs molecular dynamics simulations in triplicate.

Across all replicates, glycine rapidly dissociated from the transporter within 10 ns, indicating low binding stability. In contrast, PLP showed stable binding, supported by MM/PBSA analysis with an average binding energy of -68.63 ± 0.34 kcal/mol. Per-residue energy decomposition identified R96, R187, and K242 as key contributors to PLP stabilization (from -14 to -16 kcal/mol), whereas these residues showed unfavorable contributions in glycine simulations (~1 kcal/mol).

Our results support the hypothesis that SLC25A38 may function as a mitochondrial PLP transporter and provide molecular insights into its proposed role.

A-P.40: Uncovering the Dark Side of the Immunopeptidome
Track: Proteins and structural biology
  • Steffen Lemke, Department of Peptide-based Immunotherapy, Institute of Immunology, University and University Hospital Tübingen, Germany
  • Marcel Wacker, Department of Peptide-based Immunotherapy, Institute of Immunology, University and University Hospital Tübingen, Germany
  • Jens Bauer, Department of Peptide-based Immunotherapy, Institute of Immunology, University and University Hospital Tübingen, Germany
  • Jonas Scheid, Department of Peptide-based Immunotherapy, Institute of Immunology, University and University Hospital Tübingen, Germany
  • Annika Nelde, Department of Peptide-based Immunotherapy, Institute of Immunology, University and University Hospital Tübingen, Germany
  • Sven Nahnsen, Quantitative Biology Center (QBiC), University of Tübingen, Germany
  • Juliane S. Walz, Department of Peptide-based Immunotherapy, Institute of Immunology, University and University Hospital Tübingen, Germany


Presentation Overview: Show

The immune system targets cancer cells via T cell-mediated recognition of antigenic peptides presented on human leucocyte antigen (HLA) molecules. Mass spectrometry (MS)-based immunopeptidomics represents the standard method for direct identification of naturally presented tumor peptides, essential for developing T cell-based immunotherapies. However, a large fraction of proteins and protein regions lack detectable HLA-presented peptides, forming immunopeptidomic ‘dark spots' whose biological and technical origins remain uncharacterized. Here, we used the PCI-DB, a uniquely large database comprising >10 million HLA-presented peptides, for comprehensive mapping of the immunopeptidome dark spot landscape. While HLA binding predictions suggest a median proteome coverage of 87% per person across UK Biobank donors, experimental immunopeptidome data have captured only 30% of the human proteome. HLA peptide coverage scales with gene expression, with the 11% of proteins lacking any detectable peptides showing the lowest expression. Analysis of cancer-related variants from the COSMIC database revealed that driver mutations are more frequent in light spots, whereas passenger mutations occur more frequently in dark spots. To investigate structural and post-translational contributors to dark spot formation, transmembrane annotations (TOPdb) and glycosylation sites (GlyGen) were compared to dark and light spots, revealing both to be overrepresented in dark areas. Reprocessing PCI-DB data to include glycopeptides increased glycosylation site coverage by 58%. Broader post-translational modification searches resolved ~1.4% of dark spots, with adapted MS acquisition methods expected to increase this further. Overall, characterizing dark spots as either technical or biological in origin will enhance epitope selection accuracy and expand the repertoire for T-cell-based immunotherapy.

A-P.41: Stability-informed low-N learning of protein fitness landscape
Track: Proteins and structural biology
  • Shannon Zhang, Georgia Institute of Technology, United States
  • Yunan Luo, Georgia Institute of Technology, United States


Presentation Overview: Show

Predicting how protein sequence mutations affect fitness is central to protein engineering and variant effect interpretation, yet most experimental assays yield only limited labeled variants, making low-N learning a major challenge.
We present PsiFit, a protein stability-informed framework for low-N protein fitness prediction. PsiFit is motivated by the biological principle that proteins must maintain sufficient stability to function. PsiFit explicitly incorporates this constraint into learning the sequence–fitness landscape. Building on our prior work in mutation stability prediction (SPURS) and low-N fitness modeling (ConFit), PsiFit introduces a unified framework that integrates stability-aware biophysical priors with protein language model adaptation. Specifically, PsiFit leverages predicted mutation-induced stability changes as an explicit signal to guide fitness learning, while employing a contrastive fine-tuning strategy that preserves the general biological knowledge encoded during pretraining and aligns the model with experimentally measured fitness. By embedding stability constraints directly into the learning process, PsiFit improves sample efficiency and mitigates overfitting in small and noisy datasets.
Across more than 100 deep mutational scanning datasets in ProteinGym, PsiFit consistently outperformed multiple state-of-the-art low-N fitness prediction methods. The gains were observed across diverse proteins and assay types, indicating that stability provides a broadly useful inductive bias for learning protein fitness landscapes from sparse data. These results show that integrating biophysical stability priors with protein language models offers an effective and general strategy for data-efficient protein fitness prediction, with implications for ML-guided protein engineering and variant effect interpretation.

A-P.42: Integrative Proteomic Analysis of Inner Ear and Systemic Fluids in Noise-Induced Hearing Loss
Track: Proteins and structural biology
  • Motahare Khorrami, School of Engineering, Macquarie University, Sydney, New South Wales, Australia, Australia
  • Robert Gay, Cochlear Limited, 1 University Avenue, Macquarie University, North Ryde, NSW 2109, Australia, Australia
  • Ya Lang Enke, Cochlear Limited, 1 University Avenue, Macquarie University, North Ryde, NSW 2109, Australia, Australia
  • Mohsen Asadnia, School of Engineering, Faculty of Science and Engineering, Macquarie University, North Ryde, NSW 2109, Australia, Australia
  • Paul A. Haynes, School of Natural Sciences, Macquarie University, North Ryde, NSW 2109, Australia, Australia
  • Christopher Pastras, School of Engineering, Faculty of Science and Engineering, Macquarie University, North Ryde, NSW 2109, Australia, Australia


Presentation Overview: Show

Noise-induced hearing loss (NIHL) arises from prolonged exposure to loud sounds in the environment. Early diagnosis is challenging because some individuals may not exhibit significant threshold shifts on standard audiometry and other functional tests. Characterising the cellular changes underlying NIHL may aid early intervention and the development of targeted therapies. This study provides the first comparative proteomic analysis of perilymph, cerebrospinal fluid (CSF), whole blood, and plasma from Sprague-Dawley rats after acute noise exposure. We identify shared molecular signatures of NIHL relative to controls, and potentially relevant systemic biomarkers.
Guinea pigs were exposed to acute noise (110 dB, 90 min) under anaesthesia, resulting in a ~40 dB click-evoked threshold shift, quantified via cochlear nerve compound action potential measurement. Proteins from perilymph, CSF, blood, and plasma were analysed by nanoLC–MS/MS in data-independent acquisition mode. Differential expression analysis was conducted using the Mass Spectrometry Downstream Analysis Pipeline (MS-DAP).
We identified a total of 2,268 proteins in perilymph, 2,548 in CSF, 3,351 in whole blood, and 993 in plasma across 10 animals. Using the MS-EmpiRe model, differential expression analysis revealed 388 significantly altered proteins in perilymph, 396 in CSF, 80 in blood, and 95 in plasma following noise exposure. Exposure to intense noise triggers a systemic surge in Serum amyloid A protein, Isocitrate dehydrogenase, Fructose-bisphosphate aldolase, and Glutathione S-transferase B across the samples. This synchronized elevation across multiple fluid compartments reflects a complex physiological response involving acute inflammation, altered energy metabolism, and heightened oxidative stress defense in response to acoustic trauma.

A-P.43: Unveiling the Diversity of Catalytically Inactive Long-B Prokaryotic Argonaute Proteins
Track: Proteins and structural biology
  • Simonas AÅ¡montas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • ÄŒeslovas Venclovas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania


Presentation Overview: Show

Argonaute proteins are present in all domains of life. Prokaryotic Argonaute proteins (pAgos) are known to act as defense systems against foreign genetic elements. Historically, pAgos have been divided into two major groups, Short and Long pAgos, the latter being subdivided into A and B clades. Most pAgos in the Long-A clade have an intact canonical Argonaute nuclease active site, whereas Long-B pAgos have lost this feature. Short pAgos,being catalytically inactive like the Long-B pAgos, compensate for the lack of an active site by employing fused effector domains. However, most Long-B pAgos do not feature such fusions. Instead, in some cases it has been shown that Long-B pAgos function together with associated effector proteins. Yet, the diversity and distribution of putative Long-B pAgo effectors so far have not been thoroughly assessed.

Our bioinformatic study presents a comprehensive picture of the associations of Long-B pAgos and their effectors. Using a large set of Long-B pAgos we discovered three major groups of effector proteins associated with Long-B pAgos, and a minor group representing pAgos fused with effector domains. The members of thefirst major group of effector proteins exhibit modular organization, consisting of a conserved adaptor domain and a fused variable effector, including PDEXK nuclease domains, putative transmembrane segments, and other toxic domains. The second major group is made up of single-domain SIR2 effectors. The last major group corresponds to a mysterious protein family whose possible functions are yet unclear, but whose ability to interact with pAgos is supported by AlphaFold3 modeling experiments.

A-P.44: Characterization of a highly diverged Cas7 homologs in Felix phages
Track: Proteins and structural biology
  • Jyotika Pachauri, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Irmantas Mogila, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Lidija Truncaitė, Institute of Biochemistry, Life Sciences Center, Vilnius University, Lithuania
  • Jonas Juozapaitis, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Aistė Skorupskaitė, LSC-EMBL Partnership Institute for Genome Editing Technologies, Life Sciences Center, Vilnius University, Lithuania
  • Lukas Valančauskas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Praneet Prabhanjan, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Monika Šimoliūnienė, Institute of Biochemistry, Life Sciences Center, Vilnius University, Lithuania
  • Kristupas Užkurnys, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Tomas Šinkūnas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Česlovas Venclovas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania
  • Patrick Pausch, LSC-EMBL Partnership Institute for Genome Editing Technologies, Life Sciences Center, Vilnius University, Lithuania
  • Darius Kazlauskas, Institute of Biotechnology, Life Sciences Center, Vilnius University, Lithuania


Presentation Overview: Show

Cas7 proteins are essential structural components of Class 1 CRISPR-Cas effector complexes. They form a backbone that binds CRISPR RNA (crRNA) through an RNA Recognition Motif (RRM) fold to facilitate target recognition and interference. Apart from their role in the CRISPR-Cas system, no independent functions have been described. We identified Gp87, a highly diverged Cas7 homolog encoded by Felix phage VpaE1 outside any recognized CRISPR-Cas genomic context. Despite having minimal sequence conservation, structural models indicate that Gp87 shares the canonical Cas7 fold but lacks the typical active site and possesses two unique alpha-helical insertions. Using comparative genomics, biochemical experiments, structural modeling, and deep learning models we characterized Gp87 and its distribution across phages, archaea, and bacteria.

We found that Gp87 is co-expressed with Gp86, a helix-turn-helix (HTH) protein, and specifically binds to a flanking non-coding RNA (ncRNA) element containing GGTNN repeats. Structural models show that this interaction is mediated by Gp87's helical insertions which binds with the GG-rich repeats, an RNA-recognition process different from that of canonical Cas7 proteins. Furthermore, we identified that the architecture, comprising Gp87, Gp86, and the flanking ncRNA elements, is conserved not only across all known Felix phages, but also among diverse bacterial clades that maintain conserved GG-repeat regions and, in some groups, additionally encode flanking HNH endonucleases. This work reveals a novel, widespread Cas7-like system that operates independently of the CRISPR-Cas machinery, with a distinct RNA-binding mechanism and a transposon-associated genomic context.

A-P.45: Identification of new P-loop NTPase-like families
Track: Proteins and structural biology
  • Shreya Tapaswi, Department of Bioinformatics, Institute of Biochemistry and Biophysics, Polish Academy of Sciences, Poland
  • Kamil Steczkiewicz, Department of Bioinformatics, Institute of Biochemistry and Biophysics, Polish Academy of Sciences, Poland


Presentation Overview: Show

P-loop NTPases, constitute a major portion of any organism's proteome; they generally bind or hydrolyze nucleoside triphosphates (NTPs) for chemo-mechanical energy transduction to drive biochemical processes like, gene expression regulation, signal transduction, chromosome partitioning, DNA repair and recombination, intracellular and membrane transport, etc. They retain αβα sandwich structural core of central β-sheet of 4-6 parallel β-strands surrounded by α-helices (CATH Superfamily 3.40.50.300 P-loop containing nucleotide triphosphate hydrolases, SCOP fold: P-loop containing nucleoside triphosphate hydrolases) and possess two major highly conserved sequence motifs, namely, Walker A and Walker B. Despite the common fold, P-loop NTPase-like proteins display considerable diversity on both sequence and structure levels. Using remote homology detection methods we identified seven new families of unknown function (Domains of Unknown Functions, DUFs), to be classified as P-loop NTPases, and described their diversity including variability in either the number of the core structural elements or composition of functionally important Walker motifs. Based on predicted cellular localization, presence of transmembrane helices, genomic neighborhood and domain fusion analysis, we have hypothesized on identified families' functions which include phospholipid biosynthesis, stress response, etc. We hope our results will inspire further experimental efforts to unravel detailed functions of these seven newly annotated families, especially in the context of developing drugs against pathogenic species.

A-P.46: Machine learning approaches for charting the functional and structural landscape of protein ubiquitination
Track: Proteins and structural biology
  • Julian van Gerwen, ETH Zurich, Switzerland
  • Pedro Beltrao, ETH Zurich, Switzerland


Presentation Overview: Show

Protein ubiquitination regulates cell biology through diverse avenues, from quality control-linked protein degradation to regulatory functions such as modulating protein-protein interactions and protein conformations. Mass spectrometry-based proteomics has allowed proteome-scale quantification of hundreds of thousands of ubiquitination sites (ubi-sites), however the functional importance and molecular mechanisms of most ubi-sites remain undefined. By integrating multi-species proteomics data we found that regulatory ubi-sites may be more important than degradation-linked sites due to higher evolutionary conservation. To further prioritize regulatory ubi-sites performing cell-critical functions, we developed a machine learning-based ubi-site functional score by integrating evolutionary, proteomics, and structural features. Analysis of clinical genomics data and experiments with chemical genetics and genetic code expansion confirmed that our score pinpoints ubi-sites important for cellular fitness. Next, to pave the way for mechanistic interpretation of regulatory ubi-sites we developed an approach for modeling protein-protein covalent bonds - such as ubiquitination - in the deep-learning structural predictor AlphaFold3, which is not natively possible. We benchmarked this approach by re-predicting 338 experimental structures of ubiquitinated proteins in the Protein Data Bank, and achieved moderate-to-high accuracy across a range of protein types, including mono-ubiquitinated proteins, polyubiquitin chains, and catalytic intermediates of ubiquitin-processing enzymes. Combined with our ubi-site functional score, AlphaFold3 confirmed the regulatory potential of diverse ubi-sites by revealing ubiquitination-induced structural changes connected to protein functional states. Overall, we have leveraged machine learning through multiple avenues to chart the functional significance and structural consequences of ubiquitination at proteome-scale.

A-P.47: Multi-property Evaluation of Nanobodies Designed by an AI-Driven Pipeline for Binding to Specified Sites with Experimental Validation
Track: Proteins and structural biology
  • Yiran Wang, University College London, United Kingdom
  • Paul Dalby, University College London, United Kingdom


Presentation Overview: Show

Obtaining nanobodies with superior overall performance remains a major challenge in the field of protein engineering. Traditional approaches often yield nanobodies that bind only to the most optimal site, making it difficult to explore other regions of the antigen surface or to recognize specific epitopes. Although researchers have gradually introduced AI-driven tools into protein design, current designs remain restricted to known epitopes, with limited exploration of novel binding sites. Moreover, most studies emphasize binding affinity as the primary objective while overlooking other properties of the designed nanobodies. Therefore, fully leveraging the potential of AI-driven tools to identify a broader range of binding sites, designing nanobodies capable of interacting with user-defined hotspots, and systematically evaluating their properties will be a key future direction.

Based on this motivation, we developed an AI-driven nanobody design-build-test pipeline that integrates both structure-based and sequence-based design strategies to generate candidates targeting specified epitopes. A multi-attribute evaluation framework was then used to assess physicochemical properties. We designed a series of nanobodies and experimentally validated their performances, including expression, stability, and binding affinity. Our results demonstrate that some variants achieve high-level expression, many exhibit satisfactory stability, and some show favorable binding activity. Further comparative analyses highlight the contributions of key pipeline components to overall design performance, as well as the variability in nanobody behavior across different designs.

In summary, this work provides a general framework for AI-based nanobody design with experimentally validated functionality, enabling systematic exploration of nanobody-antigen interactions.

A-P.48: SELECTIVE MODULATION OF CHEMOKINE-LIKE RECEPTOR 2 BY SMALL MOLECULES
Track: Proteins and structural biology
  • Jonas Golling, Institute of Biochemistry, Faculty of Life Sciences, Leipzig University, 04103 Leipzig, Germany, Germany
  • Tina Schermeng, Institute of Biochemistry, Faculty of Life Sciences, Leipzig University, 04103 Leipzig, Germany, Germany
  • Emre Duman, Institute of Pharmaceutical Chemistry, Goethe University Frankfurt, 60438 Frankfurt am Main, Germany, Germany
  • Jens Meiler, Institute for Drug Discovery, Leipzig University, Leipzig 04103, Germany, Germany
  • Ewgenij Proschak, Institute of Pharmaceutical Chemistry, Goethe University Frankfurt, 60438 Frankfurt am Main, Germany, Germany
  • Annette Beck-Sickinger, Institute of Biochemistry, Faculty of Life Sciences, Leipzig University, 04103 Leipzig, Germany, Germany


Presentation Overview: Show

The closely related chemokine-like receptors 1 and 2 (CMKLR1 and CMKLR2), both activated by the adipokine chemerin, play an important role in inflammation and metabolism. While CMKLR1 drives chemotaxis and adipogenesis, CMKLR2 has been considered as an atypical scavenger receptor for a long time. Knockout studies of CMKLR2 demonstrated a reduction in insulin release and glucose uptake, while showing no impact on immune cells. This suggests a distinct role for CMKLR2 in glucose homeostasis. Thus, selective chemical probes targeting CMKLR2 could further elucidate the chemerin system and serve as a foundation for novel therapeutics targeting metabolic diseases.

Originally developed as a ligand for G protein-coupled receptor 132 (GPR132), the small molecule T-10418 was identified to induce arrestin-3 recruitment at CMKLR2 during GPCR-selectivity screening. Subsequent pharmacological characterization confirmed its selectivity for CMKLR2 over CMKLR1. Through systematic ligand-based lead optimization of the partial agonist T-10418, the maximum efficacy was enhanced from 50% to 100% in comparison to the reference peptide chemerin-9. Additionally, a group of derivatives exhibiting antagonistic properties was identified. By integrating computational docking and molecular dynamics simulations with biochemical validation, the binding mode of the most potent agonist within the orthosteric pocket was elucidated. Based on these models, a structural hypothesis for the observed CMKLR2 selectivity and activation mechanism is derived. With the identification of these selective probes, the base for a precise pharmacological interrogation of CMKLR2 in metabolic health has been established.

A-P.49: Multiscale Computational Dissection of CCRL2-Mediated Chemerin Presentation
Track: Proteins and structural biology
  • Arianna Migliorini, Sapienza University of Rome, Italy
  • Domenico Raimondo, Department of Molecular Medicine, Laboratory Affiliated to Istituto Pasteur Italia, Sapienza University of Rome, Italy


Presentation Overview: Show

CCRL2 is an atypical, non-signaling G-protein-coupled receptor (GPCR) that concentrates chemerin on the surface of expressing cells, facilitating its presentation to the functional chemerin receptor CMKLR1 and enabling recruitment of innate immune cells under inflammatory conditions. Despite this important biological role, the structural basis of CCRL2–chemerin recognition has remained poorly characterized, with no experimental structure available for CCRL2 to date. We present a comprehensive multiscale computational study integrating coarse-grained molecular dynamics (CG-MD) and all-atom molecular dynamics (AA-MD) simulations with structural modeling to dissect the mechanism of CCRL2–chemerin binding. Starting from AlphaFold2-predicted structures, we performed 26 independent CG-MD simulations totaling approximately 80 microseconds, demonstrating spontaneous chemerin association with CCRL2 following a two-step binding mechanism analogous to canonical chemokine receptors. Free energy landscape analysis identified three representative bound conformations, which were further refined through 45 microseconds of CG stable-binding simulations and subsequently characterized at atomic resolution via all-atom MD. Our results reveal a stable binding interface primarily mediated by chemerin's β1 strand and CCRL2's extracellular loop 2 (ECL2), with additional electrostatic anchoring between CCRL2's N-terminal domain and chemerin's loop 3. Critically, chemerin's C-terminal region remains solvent-accessible, consistent with a presentation-competent, nonsignaling binding mode. A modeled ternary CCRL2–chemerin–CMKLR1 complex provides a structural framework for chemerin handoff to CMKLR1. Notably, given the established role of CCRL2 in controlling NK cell homing in non-small cell lung cancer, these structural insights may inform therapeutic strategies aimed at enhancing innate immune recruitment within the tumor microenvironment.

A-P.50: Biological meaning in protein embedding space is resolution-dependent
Track: Proteins and structural biology
  • Licheng Zong, EMBL-EBI; University of Bath; The Chinese University of Hong Kong, Hong Kong
  • Jinzheng Ren, EMBL-EBI; University of Bath; Australian National University, United Kingdom
  • Yu Li, The Chinese University of Hong Kong, Hong Kong
  • Robert Finn, EMBL-EBI, United Kingdom
  • Jiawei Wang, EMBL-EBI; University of Bath, United Kingdom


Presentation Overview: Show

Protein language model embeddings are increasingly used to organise biological sequences, yet how biological meaning is encoded within embedding neighbourhoods remains poorly understood. Using two independent hierarchical enzyme systems, carbohydrate-active enzymes and peptidases, we investigated how biological interpretation changes across embedding organisations aligned to different levels of biological hierarchy. Different embedding organisations give rise to distinct neighbourhood semantics. When aligned to membership-boundary resolution, embeddings robustly separated artefacts and unrelated proteins from members of the target category. However, embeddings aligned to functional-grouping resolution maintained compositional neighbourhood structure for multi-domain proteins spanning more than one functional or catalytic group. Finally, embeddings aligned to local-family resolution recovered compact family-like neighbourhoods, including families withheld from training, while weakening broader membership-boundary and functional-grouping relationships. Moreover, embeddings optimised toward the same level of biological organisation retain different biological relationships depending on optimisation trajectory employed. Together, our results show that proximity in protein embedding space has no fixed biological interpretation. Instead, biological meaning emerges across embedding resolutions through selective preservation of different forms of biological organisation.

A-P.51: DeepAlloWeb: A Web Server for Allosteric Pocket Prediction and Attention Based Interpretation
Track: Proteins and structural biology
  • Moaaz Ur Rehman Azhar Khokhar, Koç University, Turkey
  • Ozlem Keskin, Koç University, Turkey
  • Attila Gursoy, Koç University, Turkey


Presentation Overview: Show

Allostery plays a central role in protein function and is an important direction in drug discovery because allosteric drugs can be more selective and may lead to fewer side effects than orthosteric drugs. Here, we present DeepAlloWeb, an interactive web server built on our previous method, DeepAllo, for allosteric pocket prediction and interpretation. DeepAllo combines FPocket features with a fine-tuned ProtBERT-BFD protein language model to rank candidate allosteric pockets. DeepAlloWeb extends this framework by allowing users to visualize residue-level attention and inspect how the model relates residues within and beyond predicted pockets. To study this more systematically, we performed a dataset-wide analysis across all 30 layers and 16 heads of the model. We found that the most informative patterns were concentrated in a smaller group of mainly late-layer heads, especially layer 30, and that these heads often highlighted residues far from the selected pocket residue, consistent with longer-range allosteric communication. For example, layer 30 head 11 achieved a mean average precision of 0.393. In a CheY case study, attention from D57 consistently highlighted known communication residues including T87 and Y106 across multiple heads. Together, these results show that DeepAlloWeb provides both prediction and interpretation in a single resource for studying allosteric regulation. The webserver may be accessed via: https://3dpath.ku.edu.tr/DeepAllo/.

A-P.52: A Machine Learning Framework for High-Performance Prediction of Enzyme-Substrate Interactions Using Advanced Protein Embeddings and Molecular Fingerprints
Track: Proteins and structural biology
  • João Ribeiro, Centre of Biological Engineering, Portugal
  • Ana Monteiro, Centre of Biological Engineering, Portugal
  • Dick de Ridder, Bioinformatics Group, Department of Plant Sciences}, \orgname{Wageningen University and Research, Netherlands
  • Oscar Dias, Centre of Biological Engineering; LABBELS - Associate Laboratory, Portugal
  • Miguel Rocha, Centre of Biological Engineering; LABBELS - Associate Laboratory, Portugal


Presentation Overview: Show

Motivation: Enzymes catalyze biochemical reactions with remarkable specificity and efficiency, yet predicting enzyme–substrate interactions remains challenging due to limited experimental data, incomplete annotations, and the scarcity of validated non-binding pairs.
Traditional computational approaches, such as docking and molecular dynamics, are resource-intensive and poorly scalable, while many machine learning (ML) models struggle to generalize to chemically diverse or previously unseen molecules.
Results: We present a novel ML framework that integrates experimentally validated enzyme–substrate pairs with advanced feature representations.
Our approach combines protein embeddings from the pre-trained language model ESM2 with NPClassifierFP molecular fingerprints to capture functional and structural properties of enzymes and substrates.
A gradient boosting classifier (XGBoost) is trained using rigorous data partitioning with the GraphPart method to prevent data leakage and ensure robust generalization across similarity thresholds. Benchmarking against state-of-the-art models (ESP, ProSmith) and similarity-based approaches (BLASTp + Tanimoto similarity) demonstrates that our method consistently outperforms existing approaches, achieving high F1 scores and Matthews Correlation Coefficients (MCC).
Notably, the model maintains strong predictive performance under stringent protein (40–80% sequence identity) and compound (20–80% Tanimoto similarity) similarity constraints, and generalizes well to chemically distinct, previously unseen molecules.