View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

B-P.01: STAR-GO: Improving Protein Function Prediction by Learning to Hierarchically Integrate Ontology-Informed Semantic Embeddings
Track: Proteins and structural biology
  • Mehmet Efe Akça, BoÄŸaziçi Üniversitesi, Turkey
  • Gökçe UludoÄŸan, Bogazici University, Turkey
  • Arzucan Ozgur, Bogazici University, Turkey
  • Inci BaytaÅŸ, Bogazici University, Turkey


Presentation Overview: Show

Motivation: Accurate prediction of protein function is essential for elucidating molecular mechanisms and advancing biological and therapeutic discovery. Yet experimental annotation lags far behind the rapid growth of protein sequence data. Computational approaches address this gap by associating proteins with Gene Ontology (GO) terms, which encode functional knowledge through hierarchical relations and textual definitions. However, existing models often emphasize one modality over the other, limiting their ability to generalize, particularly to unseen or newly introduced GO terms that frequently arise as the ontology evolves, and making the previously trained models outdated.
Results: We present STAR-GO, a Transformer-based framework that jointly models the semantic and structural characteristics of GO terms to enhance zero-shot protein function prediction. STAR-GO integrates textual definitions with ontology graph structure to learn unified GO representations, which are processed in hierarchical order to propagate information from general to specific terms. These representations are then aligned with protein sequence embeddings to capture sequence–function relationships. STAR-GO achieves state-of-the-art performance and superior zero-shot generalization, demonstrating the utility of integrating semantics and structure for robust and adaptable protein function prediction.
Availability: Code and pre-trained models are available at https://github.com/boun-tabi-lifelu/stargo

B-P.02: Promiscuous Mitochondrial Targeting Under Proteotoxic Stress Rewires Cytosolic Proteostasis
Track: Proteins and structural biology
  • Shubham Goyal, CSIR-Institute of Genomics and Integrative Biology, India
  • Soumen Kundu, CSIR-Institute of Genomics and Integrative Biology, India
  • Rishab Singh, CSIR-Institute of Genomics and Integrative Biology, India
  • Kausik Chakraborty, CSIR-Institute of Genomics and Integrative Biology, India


Presentation Overview: Show

Maintaining cytosolic proteostasis is essential for cellular viability, and its disruption is a characteristic feature of neurodegenerative proteinopathies. Although quality control is traditionally associated with the proteasome and autophagy, the MAGIC (Mitochondria As Guardian In Cytosol) pathway posits that mitochondria function as crucial inter-organellar hubs. In this study, we elucidate the molecular principles that govern the stress-induced recruitment of cytosolic proteins to the mitochondrial surface. Through data integration and reanalysis of ribo-seq profiling datasets, we identify a signal-independent targeting mechanism activated by proteotoxic stress. While the translatome follows a Gaussian distribution, the mitochondrial interactome displays a distal shift in ribosomal density. This recruitment is strictly linked to translational progress; both mitochondrial and cytosolic proteins within the interactome exhibit significantly longer sequences, particularly for nascent chains that have been translated beyond a critical threshold. Importantly, biophysical profiling and protein structure analysis differentiate canonical from stress-induced targeting. Resident mitochondrial proteins undergo rigorous physicochemical filtering for enhanced stability and structural complexity. In contrast, recruited cytosolic proteins bypass these constraints, displaying biophysical instability and reduced core strength. These altered protein-protein interactions at the mitochondrial surface are experimentally validated in yeast; cycloheximide treatment and mitochondrial isolation reveal a time-dependent accumulation of K48-linked polyubiquitinated proteins. Furthermore, modulating import capacity through receptor deletion or downregulation mitigates fitness defects associated with global protein misfolding. Our findings establish mitochondria as dynamic buffers of proteostasis, underscoring a systems-level trade-off between organelle biogenesis and global protein quality control.

B-P.03: Nucleic acid 3D structure search and alignment with GTalign
Track: Proteins and structural biology
  • Mindaugas Margelevicius, Vilnius University, Lithuania
  • Shikhar Rana, vilnius university, Lithuania


Presentation Overview: Show

Structural comparison of nucleic acids, particularly RNA, is essential for identifying evolutionary and functional relationships that are often not apparent from sequence alone. However, efficient methods for large-scale search and alignment of nucleic acid three-dimensional (3D) structures remain scarce. Here we present an extension of GTalign, a high-performance structural alignment method, to support nucleic acid structures.

The proposed approach enables unified search and alignment across diverse macromolecules while maintaining high computational efficiency. We evaluated the method on a diverse non-redundant dataset of RNA structures in an all-against-all benchmark. GTalign achieves improved alignment accuracy compared to existing methods, as reflected by higher TM-scores and an increased number of significant matches. At the same time, it operates at substantially lower computational cost, providing considerable speedups across different parameter settings.

These results demonstrate that GTalign effectively balances sensitivity and efficiency for structurally diverse RNA molecules. The method supports scalable structural comparison and database search, facilitating large-scale analyses and the identification of conserved structural motifs in RNAs and other nucleic acids.

B-P.04: PUCAR : Protein unit discovery using community detection on pLM attention-weighted residue graphs
Track: Proteins and structural biology
  • Nuriye Ozlem Ozcan Simsek, Bogazici University, Turkey
  • Burak Suyunu, Bogazici University, Turkey
  • Ozdeniz Dolu, Bogazici University, Turkey
  • Enes Taylan, Bogazici University, Turkey
  • Arzucan Ozgur, Bogazici University, Turkey


Presentation Overview: Show

Motivation: Identifying functional units within protein sequences remains a central challenge in computational biology. Transformer attention in protein language models (pLMs) captures residue-to-residue relationships along the sequence. These relationships may reflect how residues organize into groups, thus enabling the discovery of biologically meaningful protein units (PUs) directly from sequence. Here, we present PUCAR, a method that converts pLM attention scores into weighted residue graphs and partitions the entire protein sequence into protein units by assigning each residue to a community through community detection.
Results: Using domain and motif annotations as reference protein units, we show that PUCAR achieves competitive or improved performance relative to baseline approaches. The method includes a correlation-based calibration step that identifies attention heads most strongly associated with the target annotations, and ablation analyses demonstrate that this biologically informed head selection improves sequence partitioning quality. On independent held-out test sets, PUCAR successfully recovered annotated protein regions despite being calibrated with limited annotation data. In addition to recovering curated regions, PUCAR enables sequence-wide decomposition of proteins into units, extending PU discovery beyond curated annotations. Collectively, these findings establish a link between transformer attention and biologically meaningful protein sequence organization.

B-P.05: 3Dswappred2: Enhanced Sequence-Based Prediction of Protein Domain Swapping
Track: Proteins and structural biology
  • Dheemanth Regati, National Centre for Biological Sciences, India
  • Sarthak Ghatkar, National Centre for Biological Sciences, India
  • Sowdhamini Ramanathan, National Centre for Biological Sciences, India


Presentation Overview: Show

Domain swapping is a unique form of protein-protein interaction where proteins that are in a monomeric form interact with each other to form an oligomeric structure by exchanging domains with each other. Domain swapping is linked to functional regulation and protein aggregation. We present 3Dswappred2, an updated machine learning framework that builds on our lab's previous work to predict the propensity for domain swapping from sequence. Using a curated dataset of approximately 4,500 proteins, we have implemented a CNN-based model that currently achieves ~85% accuracy, representing a significant performance gain over our previous tools (Shameer et al. 2011, Upadhyay, A. K., & Sowdhamini, R. (2016).

To further improve predictive power, we are exploring additional architectures and integrating specific biophysical triggers. A key focus is the prediction of hinge regions, which are the flexible segments essential for conformational exchange (Shingate et al. 2012), the prediction of the extent of swapping, and the modelling of pH-dependence. Since pH shifts often act as critical switches for domain swapping in vivo, incorporating these environmental factors allows for a more physiologically relevant assessment of protein stability and misfolding (Shingate et al. 2015).

3Dswappred2 aims to provide a robust resource by combining sequence-derived propensity with localised structural insights. We discuss the current model performance and the ongoing integration of environmental modulators that influence the swapping equilibrium.

B-P.06: Evaluating Confidence Scores in AlphaFold for Protein Interaction Prediction
Track: Proteins and structural biology
  • Arthur Valentin, VIB AI / KU Leuven, Belgium
  • Damien Legros, VIB AI / KU Leuven, Belgium
  • Iker Reinares Zapata, UAM, Spain
  • Joana Pereira, KU Leuven - VIB.AI, Belgium


Presentation Overview: Show

Accurate prediction of protein protein interactions remains a central challenge in computational biology, particularly in the context of scalable, high-throughput analyses. Recent advances in structure prediction models have introduced confidence metrics, such as interface predicted TM-score (ipTM), that are increasingly used as proxies for interaction likelihood in protein complexes. However, the extent to which such metrics reliably capture true interaction signals remains insufficiently understood.
In this work, we investigate the determinants and limitations of ipTM as a scoring function for protein protein interactions. We analyze how residue-level properties, such as structural confidence and spatial relationships, contribute to the resulting scores, with the aim of clarifying the signals that drive scores predictions. We further explore how broader contextual factors, such as the composition of protein assemblies, influence these metrics, and assess their robustness under varying modeling conditions.
In addition, we examine practical considerations related to the computational cost of large-scale interaction prediction, exploring avenues to make such approaches more efficient and broadly applicable. We also consider potential directions for improving current scoring strategies, with the goal of enabling more reliable assessment of protein protein interactions.
Together, our findings contribute to a deeper understanding of confidence metrics in protein complex prediction.

B-P.07: Scaling and democratising structure-based protein function prediction with metagenomic-deepFRI
Track: Proteins and structural biology
  • Valentyn Bezshapkin, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich, Switzerland
  • Filip Schymik, AGH University of Krakow, Sano Centre for Computational Medicine, Krakow, Poland, Poland
  • Piotr Kucharski, Aiformatics, Krakow, Poland, Poland
  • Pawel Szczerbiak, AGH University of Krakow, Sano Centre for Computational Medicine, Krakow, Poland, Poland
  • Jakub Wojciechowski, Sano Centre for Computational Medicine, Krakow, Poland, Poland
  • Lukasz Szydlowski, Sano Centre for Computational Medicine, Krakow, Poland, Poland
  • Tomasz Kosciolek, AGH University of Krakow, Sano Centre for Computational Medicine, Krakow, Poland, Poland


Presentation Overview: Show

Proteins drive the functioning of all living matter, from cell structural integrity to complex metabolic reactions. While high-throughput sequencing has revealed an large diversity of proteins, driving an exponential expansion of public databases, functional characterization lags behind sequence acquisition. To aid the experimental characterization, numerous computational approaches were developed. One example is deepFRI, a GO-term prediction tool that leverages both sequence and structure of a protein. The inclusion of structure improves both the confidence and specificity of deepFRI's prediction. For large scale metagenomic experiments however, obtaining high quality protein structures for all discovered sequences may be computationally challenging. Metagenomic-deepFRI is an extension of the original pipeline, that allows for automatic retrieval of homologous protein structures from the reference databases such as AlphaFold database or ESM atlas. It leverages lightweight storage of reference protein structures, MMSeqs2 search for scalable retrieval of proteins similar to the query, a quick global-alignment between the query and the matched protein and contact-map construction as a representation of protein structure. Current limitation of metagenomic-deepFRI is that it only predicts Gene Ontology terms, which might be insufficient for thorough investigation of a set of proteins, for example, coming from an environmental sample. Many other modalities could be explored, such as genomic context, oligomeric state of the protein or possible interactions. Metagenomic-deepFRI is therefore under ongoing development, that aims to transform it into a comprehensive pipeline for both bioinformaticians running large scale computations and wet-lab oriented scientists looking for familiarity and ease of use.

B-P.08: Phyre2.2: A web server to predict protein structure and protein/ligand complexes
Track: Proteins and structural biology
  • Harold R Powell, Imperial College London, United Kingdom
  • Suhail Islam, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Anja Conev, Imperial College London, United Kingdom
  • Eleanor Stevens, Imperial College, United Kingdom
  • Alessia David, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Michael Sternberg, Imperial College London, United Kingdom


Presentation Overview: Show

Template-based protein structure prediction remains a powerful complementary approach to the recent machine-learning algorithms such as AlphaFold and Boltz. Indeed, our Phyre server continues to be widely used with over 30,000 unique users running over 200,000 jobs in 2025.

Our new release, Phyre2.2, enables a user to input a sequence and the program then identifies the closest AlphaFold2 model which then acts as a template for model prediction. This is in addition to the traditional Phyre2 approach of basing the model on an experimental Protein Data Bank (PDB) structure. Using an AlphaFold structure as a template will be particularly useful when a new proteome has been sequenced and models are not yet available to the community.

We are also launching Phyre2.2 Ligand which enables a user to obtain a model for a ligand located within a Phyre2.2 predicted structure. There are two modes for Phyre2.2 Ligand. In both, the first step cavities are identified in the predicted structure. Then, in the first mode, the coordinates of a ligand in the template are transplanted into the predicted model to generate a downloadable protein/ligand complex. In the second mode Phyre2.2 ligand will enable a user to dock a selected ligand into the predicted structure using AutoDock Vina. The ligand can be from the PDB template, from a UniProt entry or user defined via a SMILE string.

Phyre2.2, an ELIXIR resource, is freely available to all users, including commercial users, at https://www.sbg.bio.ic.ac.uk/phyre2/ .

B-P.09: UniProt Pan Proteomes: a scalable resource for comparative proteome analyses and exploring species diversity
Track: Proteins and structural biology
  • Tanushree Tunstall, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Giuseppe Insana, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Stephanie Lo, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • John Lees, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Maria Martin, martin@ebi.ac.uk, United Kingdom


Presentation Overview: Show

The rapid growth of sequence data requires scalable frameworks that capture species-level diversity beyond single reference proteomes. We present UniProt Pan Proteomes (PP), a resource integrating conserved (core-like) and variable (accessory) proteins to support comparative proteomics, pathogenicity and resistance studies, gene essentiality analyses, and population-aware target discovery.

For each species with >=3 eligible proteomes, sequences are clustered with MMseqs2 (90% identity, 50% coverage) to generate a non-redundant PP. The best annotated protein per cluster is selected using a hierarchical strategy prioritising reference proteomes and reviewed entries. Each species's PP dataset includes three outputs: FASTA entries with protein frequency; PP matrix, akin to the pan-genome (PG) presence-absence matrix; a statistics file summarising proteome composition, clustering, and frequencies.

The upcoming 2026_02 release includes ~3,200 species, created from ~70,000 proteomes, and ~270 million clustered proteins. Preliminary benchmarking in Escherichia coli shows strong PP and PG concordance in core-like subsets (96% bidirectional matching), while full-set PP-PG comparisons retain expected accessory diversity. Inter-species PP comparisons (e.g. E. coli vs. Shigella) identify conserved and divergent sequences. The Streptococcus pneumoniae PP, built from ~156 proteomes, spans ~40 pneumococcal lineages showcasing robust coverage of species diversity. It comprises 6,411 protein clusters consistent with recent PG studies. To demonstrate translational use, we mapped 64 S. pneumoniae vaccine antigens to within- and across-species PP clusters to identify homologues with functional annotations for target prioritisation. UniProt PP facilitates cross- species comparative analyses, supporting diverse research use cases. We invite community feedback to shape future releases.

B-P.10: Probing multi-omic datasets to unravel multi-stability mechanisms in CAZymes from extreme biomass-rich environments.
Track: Proteins and structural biology
  • Carlos Huertas Dí­az, KTH Royal Institute of Technology, Sweden
  • Lauren Sara McKee, KTH Royal Institute of Technology, Sweden
  • Johan Larsbrink, Chalmers University, Sweden
  • Pakinee Thianheng, KTH Royal Institute of Technology, Sweden


Presentation Overview: Show

Carbohydrate-active enzymes (CAZymes) are important biocatalysts for industrial processes where harsh conditions are present, such as high temperatures, extreme pH, high shear stress, or high solids loading, but their stability mechanisms remain poorly understood. This is especially true when multiple stressors are relevant at the same time. In this project, we mine metagenomes from extreme biomass-rich environments to identify robust CAZyme candidates with predicted favourable stability across a range of parameters. We use dbCAN3 for accurate CAZyme annotation, combined with domain and sub-family assignment, a structural feature analysis, and in silico stability parameter predictions to spotlight enzymes with a likely tolerance to multi-stress biocatalysis conditions. The use of these tools will focus on glycoside hydrolases (GHs) that depolymerize complex polysaccharides. Candidate proteins are further evaluated for the presence of carbohydrate-binding modules (CBMs), and other sequence or structural features associated with thermostability, pH range, aggregation propensity, and secretion signals, used in a robustness scoring system. Selected hits will then be cloned, expressed, and experimentally characterized for enzymatic activity and stability parameters. This integrative pipeline aims to generate a curated library of multi-stable CAZymes from metagenomic sources to improve the understandings of the molecular features of CAZyme robustness. The resulting candidates may serve to validate and improve the accuracy of the stability predictors, as well as representing robust new biocatalysts for industrial processes.

B-P.11: Computational analysis of protein-protein interactions in biocondensates
Track: Proteins and structural biology
  • Enrique Alanis Dominguez, CABD/CSIC, Spain
  • Luis Ángel Rodríguez Lumbreras, ICVV-CSIC, Spain
  • Ana Rojas, CABD/CSIC, Spain
  • Juan Fernandez-Recio, ICVV/CSIC, Spain


Presentation Overview: Show

Cellular biocondensates are membraneless compartments that are increasingly recognized as key regulators of cell physiology, especially under stress and during development. They allow for a rapid response and improve the spatiotemporal control of biochemical reactions. However, their study has traditionally been limited by the availability of experimental data in the form of high-throughtput experiments and only a few structural models. In this work, we propose an innovative approximation to the study of biocondensates based on the large-scale biophysical analysis of protein-protein interaction structural models generated with AlphaFold2. Our objective is to identify differential features on protein interactions that typically appear in different contexts: Inside biocondensates vs out of biocondensates, or in one biocondensante vs other biocondensates. We are also interested in comparing interaction binding energetics between each pair of proteins to other interactions for each of the proteins. By deepening the analysis of protein-protein interactions using biophysical features, we close a knowledge gap that could not be properly assessed before the recent AI revolution on protein complex modelling. Finally, this project pioneers a new protein-protein interaction analysis strategy that can be exploided for the study of other molecular systems.

B-P.12: AbMuSiC: a physics-based method for predicting the change in binding affinity upon mutation in antibody-antigen complexes
Track: Proteins and structural biology
  • Andre Ciupitu, Université Libre de Bruxelles, Belgium
  • Gabriel Cia, Université Libre de Bruxelles, Belgium
  • Marianne Rooman, Université Libre de Bruxelles, Belgium
  • Fabrizio Pucci, Université Libre de Bruxelles, Belgium


Presentation Overview: Show

Antibodies have become indispensable in biomedical research and are rapidly becoming an important platform for the development of next generation therapeutics. In this context, computational tools can be invaluable to accelerate the rational optimization of initial antibody candidates and minimize experimental screening. Here we introduce AbMuSiC, a structure-based method to predict the change in binding affinity upon mutation (DDG_b) that is specifically designed for antibody-antigen interfaces. AbMuSiC is a physics-based model that linearly combines coarse-grain statistical potentials derived from experimental protein structures. Additionally, the model includes a clash term and an amino acid volume term, which are crucial for the accurate prediction of mutations that fill interface cavities. AbMuSiC takes as input a 3D structure of the wildtype antibody-antigen complex, and can predict the effect on the binding affinity of both single and multiple interface mutations. When evaluated in strict cross-validation on all antibody-antigen mutations from the SKEMPIv2 experimental dataset, AbMuSiC reaches a Pearson correlation of 0.55 and a standard deviation of 1.70 kcal/mol. On AbAgym, our recently published antibody-antigen specific benchmark that contains 35k data points from 68 deep mutational scanning experiments that capture the effect of interface mutations on antibody-antigen binding, AbMuSiC is the best performing method. In conclusion, AbMuSiC is a state-of-the-art physics-based DDG_b prediction method that can help in the rational optimization and design of antibody-antigen interfaces. It will be made freely available for academic use as a Python package.

B-P.13: Identifying Candidate Precursors of Cyclic Repetitive Antibody Targets with the Peptidic Antigen Sequence Complexity Analyzer (PASCA)
Track: Proteins and structural biology
  • Salvador Eugenio Caoili, University of the Philippines Manila, Philippines


Presentation Overview: Show

Peptide sequences are potentially useful antibody targets, notably among synthetic constructs that can replace more complex, labile and/or hazardous antigens (e.g., pathogen virulence factors) in the manufacture of biomedical products (e.g., vaccines, immunodiagnostics and prophylactic/therapeutic antibodies). However, product development may fail due to structural differences between linear peptides and the intended final antibody targets, particularly where the peptide amino- and/or carboxy-termini correspond to internal protein-sequence positions in those targets. Yet, such structural mismatching can conceivably be avoided by using repetitive peptide sequences (e.g., tandem repeats) in cyclic rather than linear form. Hence, the present work provides the Peptidic Antigen Sequence Complexity Analyzer (PASCA) for identifying candidate precursors of cyclic repetitive antibody targets in FASTA-formatted input data comprising linear B-cell epitopes (LBCEs). PASCA characterizes peptidic (e.g., peptide and protein) antigens with regard to repetitive-sequence content quantified as paratope-binding degeneracy z, itself estimated as a lower bound on the relative perplexity of paratope-binding modes among circularly permuted sequence segments representing LBCEs and defined by a sliding window, z thus being unity for homopolymers and the reciprocal of sequence length for nondegenerate sequences. PASCA forms part of the PASCA User Toolkit (PUT), which also comprises the PASCA-associated Utility for Matrix Assignment (PUMA) and the PASCA-associated Utility for Sequence Selection (PUSS). PUMA provides for user-defined alternatives to the default residue-similarity matrix of Boltzmann weights used by PASCA to evaluate z, whereas PUSS facilitates preparation of PASCA input sequence data (e.g., from the Immune Epitope Database [https://iedb.org/]). All of PUT is freely accessible online (https://freeshell.de/badong/put.htm).

B-P.14: InterProScan 6: a modern large-scale protein function annotation pipeline
Track: Proteins and structural biology
  • Matthias Blum, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Emma Hobbs, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Laise Cavalcanti Florentino, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • Alex Bateman, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), United Kingdom


Presentation Overview: Show

InterProScan integrates predictive models from the InterPro consortium to annotate protein sequences with domains, families, and functional sites and is widely used in genome and metagenome annotation pipelines. It underpins large-scale annotation efforts at resources such as UniProt, Ensembl, and MGnify and is routinely applied in comparative genomics and functional annotation workflows. While InterProScan 5 has been extensively adopted over more than a decade, its architecture increasingly limited scalability, portability, and integration with modern workflow systems.

InterProScan 6 addresses these limitations through a complete reimplementation as a Nextflow-based workflow. The pipeline supports execution across local systems, high-performance computing (HPC) clusters, and cloud platforms with native container support via Docker, Singularity, and Apptainer. Pipeline logic is decoupled from signature data, allowing multiple InterPro releases to coexist within shared installations and enabling on-demand retrieval of required datasets. A redesigned Matches API provides access to precomputed InterPro annotations for UniParc sequences, allowing transparent reuse of existing results during execution.

Benchmarking across nine reference proteomes, from bacteria to complex eukaryotes, shows consistent wall-clock runtime reductions relative to InterProScan 5, with approximately two-fold speedups on large eukaryotic proteomes. When annotation reuse via the Matches API is available, end-to-end runtimes fall to minutes. Concordance analysis across the full Swiss-Prot dataset shows that InterProScan 6 reproduces InterProScan 5 results with near-identical precision and sensitivity across all InterPro member databases.

By combining workflow-native design, flexible data management, and reuse of precomputed annotations, InterProScan 6 enables efficient, reproducible, and scalable protein function annotation for growing genomic and metagenomic datasets.

B-P.15: Integrated Computational and Experimental Characterization of a Novel Metagenome-Derived Tryptophan Indole-Lyase
Track: Proteins and structural biology
  • Hovsep Aganyants, The Scientific and Production Center "Armbiotechnology", National Academy of Sciences of Armenia, Armenia
  • Artur Hambardzumyan, Scientific and Production Center “Armbiotechnology”, National Academy of Sciences of Armenia, Armenia
  • Tigran Soghomonyan, Scientific and Production Center “Armbiotechnology”, National Academy of Sciences of Armenia, Armenia
  • Marina Paronyan, Scientific and Production Center “Armbiotechnology”, National Academy of Sciences of Armenia, Armenia
  • Anichka Hovsepyan, Scientific and Production Center “Armbiotechnology”, National Academy of Sciences of Armenia, Armenia
  • Vladimir Vukic, University of Novi Sad, Faculty of Technology, Serbia
  • Haykanush Koloyan, Scientific and Production Center “Armbiotechnology”, National Academy of Sciences of Armenia, Armenia


Presentation Overview: Show

Tryptophan indole-lyase (TIL) is a pyridoxal 5'-phosphate-dependent enzyme that catalyzes the reversible β-elimination of L-tryptophan into indole, pyruvate, and ammonia. This enzyme has significant potential for the synthesis of L-tryptophan and its derivatives, which are valuable in pharmaceutical, agricultural, and fine chemical applications. Here, a metagenome derived from chicken manure and straw compost was used to mine for a TIL. We combined molecular cloning and computational modeling to identify and characterize a novel TIL, aiming to develop a technology for the production of L-tryptophan and its derivatives. The TIL gene was cloned into the pET-24a(+) vector using Gibson assembly and expressed in E. coli BL21 Star cells. We demonstrated the TIL capability for both tryptophan synthesis and degradation, with specific activities of 0.022 U/mg and 0.69 U/mg, respectively. Using AlphaFold 2, a 3D model of TIL was generated and evaluated. Further, L-tryptophan and a panel of substituted analogs were analyzed by molecular docking. Among tested analogs 7-Aza-L-tryptophan showed the highest affinity (-8.37 kcal/mol) towards the enzyme. This enzyme-ligand complex and ligand-free enzyme were subjected to 100 ns MD simulation and compared. Structural stability, flexibility, compactness, and solvent exposure were evaluated for both systems, revealing stability throughout the simulation. Protein-ligand hydrogen bond formation showed stable interactions across the simulation. MM-PBSA analysis demonstrated an energetically favorable interaction (-37.35 kcal/mol); additionally, individual residue contributions were analyzed. Together, in vitro and in silico analyses confirm TIL functionality and catalytic potential, supporting further biochemical characterization and engineering for sustainable biocatalysis.

B-P.16: In Silico Reverse Vaccinology Approach for Multi-Epitope Vaccine Design Against Stenotrophomonas maltophilia
Track: Proteins and structural biology
  • Sumithra B, Chaitanya Bharathi Institute of Technology, Department of Biotechnology, Hyderabad, India, India
  • Uma Laasya Mantena, Chaitanya Bharathi Institute of Technology, Department of Biotechnology, Hyderabad, India, India
  • Nusaybah Mohammed Saleem, Chaitanya Bharathi Institute of Technology, Department of Biotechnology, Hyderabad, India, India


Presentation Overview: Show

The global escalation of antimicrobial resistance poses a critical challenge to modern healthcare, necessitating computational approaches to identify novel vaccine targets against multidrug-resistant pathogens. Stenotrophomonas maltophilia, an opportunistic Gram-negative bacterium associated with hospital-acquired infections, is known for its intrinsic resistance to multiple antibiotic classes and currently lacks an approved vaccine. This study applied a reverse vaccinology-driven immunoinformatics framework to design a potential multi-epitope vaccine candidate against S. maltophilia. The complete proteome was systematically screened to identify surface-exposed and virulence-associated proteins, followed by antigenicity assessment to prioritize suitable targets. Selected proteins were subjected to B-cell and T-cell epitope prediction, and the predicted epitopes were further filtered based on antigenicity, non-allergenicity, non-toxicity, and immunogenic potential. To improve broad applicability, conserved epitope analysis across multiple strains was incorporated. The shortlisted epitopes were assembled into a multi-epitope construct using appropriate linkers and an adjuvant to enhance immune response. Molecular docking and immune simulation analyses were performed to evaluate the interaction of the vaccine construct with immune receptors and to assess its immunogenic response. Population coverage analysis was performed to evaluate the global applicability of the selected epitopes. The final construct exhibited favorable immunogenicity and safety profiles based on immunoinformatics criteria. This study highlights the utility of integrating reverse vaccinology with advanced immunoinformatics approaches to address antimicrobial resistance and provides a foundation for subsequent experimental validation.

B-P.17: Benchmarking HERMES for TCR-pMHC Binding Prediction Across Diverse Systems Using AlphaFold3-Predicted Structures
Track: Proteins and structural biology
  • Max Hoffmann, Institute of Medical Data Science, Faculty of Medicine, Otto-von-Guericke-University, Magdeburg, Germany
  • Johannes Steffen, Department of Hematology, Oncology, Cell and Radiotherapy, Faculty of Medicine, Otto-von-Guericke-University, Magdeburg, Germany
  • Matthias Leisegang, Department of Hematology, Oncology, Cell and Radiotherapy, Faculty of Medicine, Otto-von-Guericke-University, Magdeburg, Germany
  • Julian Varghese, Institute of Medical Data Science, Faculty of Medicine, Otto-von-Guericke-University, Magdeburg, Germany
  • Sarah Sandmann, Institute of Medical Data Science, Faculty of Medicine, Otto-von-Guericke-University, Magdeburg, Germany


Presentation Overview: Show

Predicting T-cell receptor (TCR) binding to peptide major histocompatibility complexes (pMHCs) is highly relevant to vaccine development and immunotherapy design. Compared to sequence-only approaches, structure-based machine learning models, such as HERMES, appear especially promising. However, while HERMES has shown competitive performance on a limited number of systems, its reliability across diverse systems without experimental structures has not yet been assessed.
To fill this gap, we benchmarked HERMES (model: fixed) against BATCAVE, a literature-derived altered peptide ligand dataset filtered to MHC-I. Eighty-four of 104 TCR-pMHC complexes (53 human, 31 mouse) were suitable for analysis. For each complex an average of 162 mutated epitopes were mapped to measured TCR activation data. The structures containing the index epitope were predicted with AlphaFold3. Finally, HERMES computes peptide energy, a proxy for binding, for all mutated epitopes.
Performance varied substantially across TCR-pMHC complexes, with Spearman correlations between predicted peptide energy scores and measured TCR activation ranging from -0.31 to +0.57. Mouse systems performed poorly, with a median correlation of 0.
Although human systems performed better, median correlations were still low (cancer antigens: 0.09; neoantigens: 0.16; viral antigens: 0.27). AlphaFold3 confidence only explained a small portion of the variability in correlation, with correlation in high confidence structures ranging from 0 to 0.49.
Our results suggest that while HERMES can be useful when applied to high-quality experimental structures, its utility on predicted structures is limited, highlighting the need for improved structure prediction tools before broader application.

B-P.18: Rhea, a FAIR resource of expert curated biochemical and transport reactions
Track: Proteins and structural biology
  • Elisabeth Coudert, Swiss Institute of Bioinformatics, Switzerland
  • Kristian B. Axelsen, SIB Swiss Institute of Bioinformatics, Switzerland
  • Lucila Aimo, SIB Swiss Institute of Bioinformatics, Switzerland
  • Nevila Hyka-Nouspikel, SIB Swiss Institute of Bioinformatics, Switzerland
  • Parit Bansal, SIB Swiss Institute of Bioinformatics, Switzerland
  • Teresa Batista Neto, SIB Swiss Institute of Bioinformatics, Switzerland
  • Edouard de Castro, SIB Swiss Institute of Bioinformatics, Switzerland
  • Nicole Redaschi, SIB Swiss Institute of Bioinformatics, Switzerland
  • Alan Bridge, SIB Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

Despite the central role of biochemical reactions in life, reaction knowledge remains fragmented across databases, inconsistently described, and often not interoperable. This limits accurate enzyme annotation, integration across resources, and downstream applications such as pathway reconstruction, enzyme function prediction, and metabolic network analysis.

Rhea (www.rhea-db.org) addresses this challenge by providing an expert-curated, FAIR resource of biochemical and transport reactions from the literature that is grounded in the ChEBI ontology and fully mapped to UniProtKB. Rhea includes over 18,500 reactions covering primary and secondary metabolism of a broad range of taxa, including reactions of the enzyme classification of the IUBMB and thousands more. Rhea has been adopted as the reference vocabulary for the annotation of enzymes and transporters in UniProtKB, covering over 25 million protein sequences, and provides reaction chemistry for resources such as the Gene Ontology (GO), Reactome, Model Organism Databases (MODs), SwissLipids, and the metabolic modeling platform MetaNetX. All Rhea data is freely available to query and download from our website and APIs, including our SPARQL endpoint, and is now widely used to develop methods for enzyme function prediction, protein design, and pathway reconstruction, and the annotation of metabolic networks, metagenomes, and environmental biotransformations.

Here we present recent developments in Rhea, including improvements in reaction coverage, integration with the GO, curation workflows that integrate AI to accelerate coverage of novel enzyme functions and "de-orphan" enzymes, and new visualizations of reaction chemistry for experts and non-experts alike.

B-P.19: Structure-guided embeddings predict CYP substrate specificity: engineering citrus flavors for climate resilience
Track: Proteins and structural biology
  • Eli Draizen, Barcelona Supercomputing Center, Spain
  • Sara Tolosa-Alarcón, Barcelona Supercomputing Center, Spain
  • Emre Cicekyurt, Barcelona Supercomputing Center, Spain
  • Miguel Romero-Durana, Barcelona Supercomputing Center, Spain
  • Alfonso Valencia, Barcelona Supercomputing Center, Spain


Presentation Overview: Show

Extreme climates may threaten crop-growing soil, limiting access to plant metabolites used for pharmaceuticals, flavors, and fragrances. Microbial cell factories could be a viable replacement by scaling up metabolite production, reducing land use and avoiding polluting extraction, but engineering them requires understanding the plant enzymes that produce these metabolites. Many such compounds are synthesized through Cytochrome P450 enzymes (CYPs), which catalyze substrate oxidation. We focus on a CYP subfamily that converts valencene to nootkatone, a popular grapefruit-flavored compound. However, only nine true positive and eight true negative valencene-catalyzing CYPs are known, limiting our understanding of which CYPs work and why. This motivates the search for additional natural valencene-catalyzing CYPs.

First, we modeled all true positives with heme and valencene. The negative set and 180,000 plant CYP candidates were modeled with heme only, superimposing valencene from their closest true positive homolog. Next, seven sequence- and structure-based embedding methods were compared across full proteins and binding site regions, using cosine similarity, embedding-based Needleman-Wunsch alignment, and optimal transport. Finally, logistic regression predicted catalysis using six features: best, top-3 mean, and median similarity to true positives and negatives for each embedding and similarity measure.

Frame2seq, a structure-guided masked language model, produced the most informative embeddings, achieving the highest three-fold cross-validation AUC closest to leave-one-out performance despite limited training data. We subjected 180,000 plant CYPs to the model, generating a ranked candidate list which we sent to collaborators for experimental validation. This will expand our dataset and deepen our understanding of valencene catalysis.

B-P.20: Target-Specific De Novo Design of Drug Candidate Molecules with Graph-Transformer-Based Generative Adversarial Networks
Track: Proteins and structural biology
  • Atabey Ünlü, Hacettepe University, Turkey
  • Elif Çevrim, Hacettepe University, Turkey
  • Melih Gökay YiÄŸit, Dept. of Computer Engineering, Middle East Technical University, 06800, Ankara, Turkey, Turkey
  • Ahmet Sarıgün, Middle East Technical University, Turkey
  • Hayriye Çelikbilek, Hacettepe University, Turkey
  • Osman Bayram, Bahcesehir University, Turkey
  • Deniz Cansen Kahraman, Middle East Technical University, Germany
  • Abdurrahman Olgac, Gazi University, Turkey
  • Ahmet Süreyya RifaioÄŸlu, Heidelberg University, Turkey
  • Erden Banoglu, Gazi University, Turkey
  • Tunca Dogan, Hacettepe University, Turkey


Presentation Overview: Show

Discovering novel small-molecule drug candidates that specifically interact with a protein target remains a central challenge in rational drug design. While deep generative models have demonstrated capacity to sample chemically valid molecules, most prior work focuses on property-optimized generation. We present DrugGEN, an end-to-end generative framework for target-centric de novo molecular design, combining generative adversarial networks (GANs) with graph representation learning. DrugGEN represents molecules as graphs encoding atom types and bond connectivity. A key architectural contribution is the integration of graph transformer encoder blocks into both generator and discriminator modules, incorporating a modified attention mechanism that explicitly amplifies atomic bond information to capture both local and long-range interatomic dependencies. To our knowledge, this is the first study to employ molecular graph transformers within a GAN pipeline. The framework operates at realistic drug-sized molecular scales (~45 heavy atoms), a complexity that most prior generative models do not address. Trained on ChEMBL-derived datasets and evaluated against 15 competing generative models, DrugGEN ranked first across both targeted and non-targeted benchmarks. In AKT1-targeted molecular docking, DrugGEN achieved 99.98% of real inhibitor binding performance, outperforming all comparators. Deep learning-based bioactivity prediction identified 748 high-confidence AKT1-active de novo molecules. Five synthesized candidates were experimentally validated via in vitro enzymatic assays, with two demonstrating AKT1 inhibition at low-micromolar doses. Attention map analysis further revealed that the model correctly identifies binding-critical atoms without explicit interaction supervision, providing mechanistic interpretability. DrugGEN is openly available at github.com/HUBioDataLab/DrugGEN, enabling the community to retrain the system for any druggable protein target.

B-P.21: Joint-Embedding Predictive Representation Learning for Protein-Ligand Pose Generation
Track: Proteins and structural biology
  • Atabey Ünlü, Hacettepe University, Turkey
  • Tunca Doğan, Hacettepe University, Turkey


Presentation Overview: Show

Accurate prediction of protein-ligand binding poses remains challenging, as models must capture interaction geometry while accommodating diverse, multimodal pose hypotheses. The Joint-Embedding Predictive Architecture (JEPA) structures latent space to enable more tractable learning of such multimodal pose distributions. Here, we propose a framework that integrates retrieval-style latent supervision with pose generation to achieve accurate and controllable protein-ligand docking, presented here as ongoing work. The two-stage framework combines joint-embedding predictive representation learning with SE(3)-equivariant explicit pose decoding. In stage one, protein pocket features and coordinates are encoded alongside a ligand graph and a randomly initialized query conformer. These representations are fused through bidirectional cross-graph interaction layers to produce a protein-conditioned invariant latent representation. A target-side bound-pose encoder maps the experimentally observed bound pose and local unbound-pocket context into a shared latent space, and the model is trained with a bidirectional InfoNCE objective to align these representations without exposing bound ligand geometry in the conditioning path. In stage two, the learned representation is frozen and used to condition an anchor-based rigid-body and torsion pose decoder that iteratively predicts translational, rotational, and torsional updates with local cross-attention and recycling-based refinement, with planned extensions to multi-start pose generation and flow-based decoding. Initial experiments on PDBBind show stable training, increasing cosine agreement between paired latent representations, and strong top-1 retrieval of the matched target pose, indicating that the learned space captures pose-relevant interaction signals. Ongoing work evaluates full-docking accuracy, pose diversity, and iterative-refinement quality relative to existing generative docking baselines.

B-P.22: Co-folding based prediction of ligand interaction sites in cytochromes P450
Track: Proteins and structural biology
  • Karoli­na Plankova, Department of Physical chemistry, Faculty of Science, Palacký University Olomouc, Czechia
  • Karel Berka, Department of Physical chemistry, Faculty of Science, Palacký University Olomouc, Czechia
  • Vaclav Bazgier, Department of Physical chemistry, Faculty of Science, Palacký University Olomouc, Czechia


Presentation Overview: Show

Accurate prediction of enzymatic reaction sites remains a central challenge in computational chemistry and metabolic modelling. Cytochromes P450 (CYPs), particularly the CYP2C9 isoform, are critical enzymes involved in the metabolism of approximately 20% of marketed drugs, xenobiotics, and endogenous compounds. Identifying precise ligand reaction centres is essential for predicting metabolic pathways and supporting rational drug design.
In this study, we evaluated the performance of recent co-folding frameworks, specifically Boltz and AlphaFold 3 in predicting ligand reaction sites. Co-folding allows for simultaneous prediction of protein and ligand conformations, accounting for induced-fit effects. A dataset of 178 unique reactants corresponding to 310 metabolic reactions catalysed by CYP2C9 was analyzed. Predicted ligand poses and their respective reaction sites were evaluated based on their spatial proximity to the heme cofactor and compared with experimental data from the DrugBank database.
Both approaches indicate that co-folding methods can capture key aspects of enzyme ligand interactions. Boltz models achieved a success rate of 51% in identifying correct reaction sites, slightly outperforming AlphaFold 3 (43%). Successful predictions showed a clear correlation with proximity to the heme iron, typically peaking at distances of 4 to 5 Ã…. To further contextualize these findings, the performance of co-folding is now being compared with traditional molecular docking. While no single structural features consistently determine the prediction success, these results highlight co-folding as a promising approach for improving the accuracy of metabolic site prediction and enhancing computational drug development tools.

B-P.23: ISS-Priority: Priority-guided graph propagation for indirect protein structure retrieval
Track: Proteins and structural biology
  • Hao Liu, University of Helsinki, Finland
  • Liisa Holm, University of Helsinki, Finland


Presentation Overview: Show

Structural relationships between proteins help infer function and evolutionary history. Fast structural search tools such as Foldseek enable large-scale retrieval but often miss remote structural homologs. We address this gap in two stages.

First, we learn a compact retrieval embedding from ESM2 representations using DALI-derived structural similarity as contrastive supervision. This injects structural signal during training while keeping inference sequence-only, and yields a direct retrieval baseline of 0.5918 pooled AUPRC on a SCOPe fold-level benchmark against the AlphaFold Database v2 (AFDB2), substantially above Foldseek (0.3828). Second, we introduce ISS-Priority, a propagation-based retrieval method that searches the learned embedding space as a graph and uses selective DALI validation to guide frontier expansion. Because runtime is controlled by the DALI validation budget, users can trade speed for recall. By recovering indirect structural hits through validated intermediate neighbors, ISS-Priority raises pooled AUPRC to 0.6128.

Together, these results show that combining DALI-supervised representation learning with propagation-based retrieval extends structure-aware search beyond direct retrieval alone.

B-P.24: Mutation-Specific ∆∆G Modeling for Alanine Substitutions Improves Stability Prediction and Reveals Distinct Constraint Regimes
Track: Proteins and structural biology
  • Maya Czeneszew, Sorbonne Université, CNRS, IBPS, Laboratoire de Biologie Computationnelle et Quantitative, France
  • Alessandra Carbone, Sorbonne Université, CNRS, IBPS, LBCQ; Institut Universitaire de France (IUF), France


Presentation Overview: Show

Alanine scanning emerged as a widely used methodology to assess protein stability by systematically mutating amino acid residues to alanine, whose small, neutral side chain isolates the loss of native side-chain interactions. Despite its utility, to our knowledge, no dedicated deep learning model exists for predicting alanine substitution ∆∆G values trained exclusively on single-residue substitutions to alanine.

We introduce a model that fine-tunes ESM-2 650M via low-rank adaptation (LoRA) with a combined prediction head including a light-attention and a transformer-based embedding difference module. The model is trained on a dataset of ~2,100 experimental ∆∆G measurements for alanine substitutions, augmented with ~2,100 FoldX-computed values, resulting in a curated alanine scanning dataset.

On the ProteinGym stability benchmark (66 assays), the model achieves a Spearman ρ of 0.73 on alanine substitutions, outperforming ESCOTT (ρ = 0.53), a mutational effects predictor modeling evolutionary conservation, structure, and epistasis, and VespaG (ρ = 0.46), the top single-sequence method on the ProteinGym leaderboard. Comparison of ESCOTT and predicted ∆∆G at the single-mutation level across representative proteins identifies regimes where stability and evolution decouple, highlighting highly constrained residues with minimal predicted stability impact that coincide with previously described spandrel-like sites characterized by persistent local frustration.

Alanine-specific ∆∆G modeling sharpens our ability to disentangle structural stability from evolutionary constraint at single-residue resolution. By exposing regimes where these signals diverge, the model provides a clearer lens on mutation effects and opens new avenues for precise variant interpretation and rational protein design.

B-P.25: Decoding protein functional determinants: A simple adapter-based module with interpretable attention
Track: Proteins and structural biology
  • Chujun Lyu, Sorbonne Université, CNRS, IBPS, Department of Computational, Quantitative and Synthetic Biology (CQSB), Paris, France, France
  • Alessandra Carbone, Sorbonne Université, CNRS, IBPS, Department of Computational, Quantitative and Synthetic Biology (CQSB), France


Presentation Overview: Show

Proteins orchestrate nearly all biological processes, yet predicting their functional determinants—key amino acids responsible for a desired function—directly from sequence remains a major challenge. Protein language models (PLMs) capture evolutionary and structural information from primary sequences, but leveraging their embeddings for residue-level functional annotation demands specialized architectures. Here, we introduce a novel approach that integrates LoRA adaptations into a PLM, using lightweight attention networks to efficiently model residue-level dependencies, followed by a linear classifier for binary token-level prediction. The PLM backbone is initialized from a model fine-tuned on soft disorder information. Our framework incorporates two complementary label types: experimental interface annotations from the Protein Data Bank and predicted post-translational modification labels from DeepMVP. Extensive validation on diverse benchmarks demonstrates strong performance. On the CAID3 Binding-IDR challenge (52 proteins), competing with tools designed for the task, our method achieves a probability prediction AUROC of 0.578, ranking 8th overall. On VenusMutHub (796 proteins), it exhibits heightened sensitivity to selectivity and stability mutations. In a threshold‑dependent manner, our method can outperform AlphaMissense on gain‑of‑function mutations (13 proteins, 52.2% vs. 42.4%) and match PRESCOTT on Finnish disease heritage mutations (22 missense, 55.5% vs. 54.5%). Over specific proteins, including two copies of TRX in Chlamydomonas, it can identify functional differences at the level of the single protein copies in the same genome. These results highlight the efficacy of lightweight attention mechanisms built upon PLM embeddings for capturing nuanced functional determinants across structured and disordered proteomes, aiding functional tree reconstruction and protein annotation.

B-P.26: Target-Conditioned Molecule Generation using Protein and Chemical Language Models with Bioactivity-Guided Post-Training
Track: Proteins and structural biology
  • Ahmet Furkan Öztürk, Middle East Technical University, Turkey
  • Atabey Ünlü, Hacettepe University, Turkey
  • Elif Çevrim, Hacettepe University, Turkey
  • Tunca Doğan, Hacettepe University, Turkey


Presentation Overview: Show

Drug discovery increasingly leverages deep generative models to accelerate the design of bioactive molecules. In this context, molecule design can be formulated as a sequence-to-sequence problem, where protein sequences act as conditioning inputs and their ligands are generated in a candidate space. We present Prot2Mol, a transformer-based autoregressive framework for de novo, target-specific molecule generation that integrates protein and chemical language modeling. Prot2Mol employs SELFIES representations and incorporates multiple protein language models: ProtT5, ESM2, and the structure-aware SaProt to capture sequence and structural signals. The model follows an encoder–decoder architecture, where protein embeddings are co-trained and used to condition a GPT-2–based molecular decoder via cross-attention. To further optimize generation, we extend the framework with a pairwise reward model that scores protein–molecule compatibility. The reward model encodes protein sequences using the same protein encoder architecture and molecules with a SELFIES-based encoder, and is trained separately from the generative model while jointly learning protein–molecule interactions to produce a ranking score and a binary activity prediction. Instead of direct regression of bioactivity values, we adopt pairwise ranking with binary classification, mitigating biases from dataset-specific protein priors and encouraging interaction-dependent signals. The reward function is integrated into a GRPO-based reinforcement learning pipeline to guide molecule generation toward higher bioactivity. Experiments on the Papyrus dataset show that Prot2Mol generates valid, diverse, and drug-like molecules while maintaining target specificity. The framework establishes a unified pipeline combining generative modeling with reward-driven optimization for target-informed molecular design. The tool is publicly available at https://github.com/HUBioDataLab/Prot2Mol.

B-P.27: Apo, Holo, and Everything in Between: What State Are AlphaFold Binding Pockets In?
Track: Proteins and structural biology
  • Christos Feidakis, Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Radoslav Krivak, Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Derrick Agwora, Department of Cell Biology, Faculty of Science, Charles University, Prague, Czech Republic, Czechia
  • Chloé Weiler, Faculty of Engineering Sciences, Heidelberg University, Heidelberg, Germany, Germany
  • David Hoksza, Department of Software Engineering, Faculty of Mathematics and Physics, Charles University, Prague, Czech Republic, Czechia
  • Jiri Vondrasek, Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Marian Novotny, Department of Cell Biology, Faculty of Science, Charles University, Prague, Czech Republic, Czechia


Presentation Overview: Show

AlphaFold DB models are widely used for ligand-binding analysis, but a practical question remains: do their binding pockets look more like ligand-free apo states, ligand-bound holo states, or something in between? We addressed this using AHoJ-DB, which links experimentally observed alternative pocket conformations for the same protein regions. From 462,413 query-pocket entries with AlphaFold classifications, after filtering for usable classifications and at least one apo comparator, we retained 222,055 pockets from 11,629 UniProt accessions. At first glance, AlphaFold pockets appeared more holo-like than apo-like, but this imbalance moved toward equilibrium once overlapping pockets were merged into unified binding sites, showing that redundancy matters. More importantly, experimental apo/holo evidence balance did not produce a hard switch between classes; instead, AlphaFold pockets shifted gradually along an apo-holo continuum with the holo-to-apo experimental evidence ratio. We then tested whether this local pocket state mattered in two downstream analyses. In same-UniProt AlphaFill matches, ligand transplants fit similarly in low-separation pockets, but fit worsened for apo-like AlphaFold pockets when the experimental apo and holo states were more distinct. P2Rank showed higher binding-site detection success for holo-like pockets and for pockets supported by higher holo evidence ratios. These results suggest that AlphaFold pockets are best interpreted as occupying a measurable structural spectrum, and that this local pocket state can help flag when ligand placement and pocket prediction are likely to be straightforward or require extra caution.

B-P.28: Structure-Function Analysis of Arabidopsis Cellulose Synthases and Small Molecule Inhibitors Using Computational and In Vivo Approaches
Track: Proteins and structural biology
  • Carlo Perolo, University of Toronto, Canada
  • Nicholas Provart, University of Toronto, Canada
  • Heather McFarlane, University of Toronto, Canada


Presentation Overview: Show

The main structural component of plant cell wall is cellulose, an unbranched polymer of β-1,4-glucose monomers synthesised by enzymes called Cellulose Synthases (CESA). In Arabidopsis thaliana, primary cell wall CESAs are hypothesized to work as hexamers of heterotrimers moving along microtubule tracks to deposit cellulose in a controlled way. Some of the most valuable tools to study plant CESAs are Cellulose Biosynthesis Inhibitors (CBIs), small molecules able to directly affect their function and localization in a dose-dependent manner. Despite significant similarity in structure and sequence of the three A. thaliana primary cell wall CESAs (CESA1, CESA3, and CESA6) resistance-conferring mutation screens reveal a surprising specificity of several CBIs for one monomer over the others. This provides us with an ideal framework to study their mechanism of inhibition, especially with the recent advances in protein structure prediction tools like AlphaFold. We performed Molecular Docking simulations on multiple CESA resolved and predicted structures and identified several CBIs' potential binding pocket. To validate our prediction, an extensive cross-resistance screen was conducted, and finally site-directed mutagenesis was used to introduce mutations targeting residues within the proposed binding pocket on CESA1. The resulting mutant plants show specific resistance to CBIs, supporting our predictions. Overall, our work furthers the understanding of CBI mechanism of inhibition by proposing a binding pocket and highlights the predictive power of computational structural approaches in understanding protein-ligand interactions. Lastly, we are currently implementing our pipeline as a publicly available tool on the Bio-Analytic Resource for Plant Biology (BAR) website.

B-P.29: Integrated computational and experimental analysis of a matricellular protein-peptide interaction enables rational antifibrotic binder design
Track: Proteins and structural biology
  • Jennifer Faúndez-Contreras, Facultad de Medicina, Universidad San Sebastián, Santiago, Chile. Fundación Ciencia & Vida, Santiago, Chile., Chile
  • Enrique Brandan, Facultad de Medicina, Universidad San Sebastián, Santiago, Chile. Fundación Ciencia & Vida, Santiago, Chile., Chile
  • Sebastián Bazaes, Centro Científico y Tecnológico de Excelencia Ciencia & Vida, Santiago, Chile., Chile
  • Raul Araya-Secchi, Facultad de Ingenieria. Universidad San Sebastián, Valdivia, Chile; Centro de Estudios Cientificos CECs, Valdivia., Chile
  • Felipe García-Olave, Programa de doctorado en Biologia Computacional. Universidad San Sebastian, Santiago, Chile, Chile
  • Tiaren Ruiz, Facultad de Ingenieria, Universidad San Sebastián, Santiago, Chile., Chile


Presentation Overview: Show

Introduction. Fibrosis is characterized by excessive extracellular matrix (ECM)
deposition leading to organ dysfunction. Connective Tissue Growth Factor (CTGF)
is a key mediator of fibrosis, promoting ECM production and fibroblast activation.
Its inhibition can attenuate fibrotic progression. Structurally, CTGF comprises four
conserved domains organized into N-terminal (domains 1-2) and C-terminal
(domains 3-4) regions, connected by a protease-sensitive hinge. The C-terminal
region has been associated with profibrotic activity. We previously identified a
peptide capable of inhibiting CTGF, however, its mechanism remains unclear. This
study aimed to define the CTGF-peptide interaction to guide the rational design of
improved CTGF-targeting binders.
Methods. Structural models of CTGF alone and in complex with the peptide were
generated using AlphaFold2 and AlphaFold-Multimer, followed by 1.2 μs molecular
dynamics simulations. Binding contributions were assessed by alanine scanning
using Rosetta Flex ddG. Experimental validation included ELISA alanine scanning
and co-immunoprecipitation assays with full-length and truncated CTGF. Functional
effects were assessed in fibroblasts via fibronectin deposition and stress fibers
formation. Statistical analysis was performed using two-way ANOVA.
Results. Computational models suggested dual peptide interaction with both
CTGF regions, promoting conformational compaction and reduced proteolytic
accessibility. Experimental data revealed predominant binding to the C-terminal
region and identified key peptide residues critical for binding and functional activity.
Conclusion. These findings define a CTGF-peptide interaction interface and
support the rational design of targeted antifibrotic binders.
Acknowledgements. This work is supported by FONDECYT-N°1230054,
Proyecto CCTE Ciencia y Vida Basal-FB210008, FONDEF-ID25I10016, and
ANID/BECA DOCTORADO NACIONAL-21241714.

B-P.30: Coordinated pMHC Dynamics Reveal TCR-Recognizable States for Neoepitope Prioritization in Personalized Cancer Vaccine Design
Track: Proteins and structural biology
  • Jelang Muhammmad Dirgantara, Laboratory of In Silico Design, National Institutes of Biomedical Innovation, Health and Nutrition, Japan, Japan
  • Kazuma Kiyotani, Laboratory of Immunogenomics, National Institute of Biomedical Innovation, Health and Nutrition, Japan, Japan
  • Takuto Nogimori, Laboratory of Precision Immunology, National Institutes of Biomedical Innovation, Health and Nutrition Japan, Japan
  • Shokichi Takahama, Laboratory of Precision Immunology, National Institutes of Biomedical Innovation, Health and Nutrition Japan, Japan
  • Takashi Morisaki, Fukuoka General Cancer Clinic, Japan, Japan
  • Takuya Yamamoto, Laboratory of Precision Immunology, National Institutes of Biomedical Innovation, Health and Nutrition Japan, Japan
  • Suyong Re, Laboratory of In Silico Design, National Institutes of Biomedical Innovation, Health and Nutrition, Japan, Japan


Presentation Overview: Show

Accurate identification of immunogenic peptide-MHC (pMHC) complexes remains a central challenge for immunotherapy, as binding affinity alone poorly predicts T cell activation. To elucidate structural determinants of T cell receptor (TCR) recognition, we analyzed experimentally resolved pMHC and TCR-pMHC crystal structures across multiple HLA class I alleles in Protein Data Bank. We show that peptide backbone bulge topology at TCR-facing residues is largely pre-organized prior to TCR engagement, supporting a conformational selection mechanism. Notably, this indicates that key structural features relevant to TCR recognition are encoded within the pMHC state, enabling inference of TCR-interacting properties without explicit TCR information. This is particularly important because, in personalized cancer vaccine design, TCR information is typically unavailable (neoepitope identification and HLA typing are performed without knowledge of cognate TCRs). Guided by these insights, we performed 100-ns all-atom molecular dynamics (MD) simulations on 945 pMHC complexes derived from the largest Japanese cancer patient neoepitope cohort with ELISpot immunogenicity labels, totaling 94.5 microseconds of simulation. We quantified anchor residue stability within MHC pockets and flexibility of TCR-facing residues. While individual dynamic features partially enrich immunogenic epitopes, they fail to provide a consistent or biophysically interpretable framework across HLA alleles. To address this, we developed a Consensus Scoring Function (CSF) integrating anchor stability and bulging dynamics into a unified measure of pMHC fitness for TCR recognition. CSF improves immunogenic hit rates by up to 8% while reducing experimental burden by half, providing a scalable framework for neoepitope prioritization in personalized cancer vaccine design.

B-P.31: ZFP-CanPred: Predicting the Effect of Mutations in Zinc-Finger Proteins in Cancers Using Protein Language Models
Track: Proteins and structural biology
  • Amit Phogat, Department of Biotechnology, Indian Institute of Technology Madras, Chennai, India, India
  • Sowmya Ramaswamy Krishnan, Department of Biotechnology, Indian Institute of Technology Madras, Chennai, India, India
  • Medha Pandey, Department of Biotechnology, Indian Institute of Technology Madras, Chennai, India, India
  • M. Michael Gromiha, Department of Biotechnology, Indian Institute of Technology Madras, Chennai, India, India


Presentation Overview: Show

Zinc-finger proteins (ZNFs) are the largest group of transcription factors and are involved in many important cellular functions. Missense mutations in ZNFs can disrupt protein-DNA interactions and may contribute to the development of different cancers. In this study, we introduce ZFP-CanPred, a deep learning-based model designed to predict cancer-associated driver mutations in ZNFs. The model uses representations obtained from protein language models (PLMs), focusing on the structural neighbourhood around mutation sites, to distinguish between cancer-causing and neutral mutations. ZFP-CanPred achieved strong performance on an independent test set, with an accuracy of 0.72, F1-score of 0.79, and area under the ROC curve (AUC) of 0.74.In comparison with 11 existing prediction tools on a curated dataset of 331 mutations, ZFP-CanPred showed the highest AUROC of 0.74, outperforming both general and cancer-specific methods. Its balanced sensitivity and specificity help address a key limitation seen in current approaches. The source code and related materials are available at: https://github.com/amitphogat/ZFP-CanPred.git. This work may support a better understanding of cancer mechanisms and assist in the development of targeted therapies.

B-P.32: Functional aggregation-prone regions mediate carbohydrate-binding and harbour disease-associated mutations
Track: Proteins and structural biology
  • Lekshmi S, Indian Institute of Technology Madras, Chennai, India, India
  • Siva Shanmugam N R, Department of Food Science and Technology, University of Nebraska - Lincoln, Lincoln, USA, India
  • Prabakaran R, Department of Biology, Emory University, Atlanta, GA, USA, India
  • Rawat P, Department of Immunology, University of Oslo, Oslo, Norway, India
  • Michael Gromiha M, Indian Institute of Technology Madras, Chennai, India, India


Presentation Overview: Show

Carbohydrate-protein interactions play key roles in cellular processes such as cell-cell communication, differentiation, and immune response, and their disruption by mutations is implicated in several diseases. Emerging evidence suggests that aggregation-prone regions (APRs), which are associated with disease-related protein aggregation, occur near functional sites. However, the role of APRs in carbohydrate recognition is largely unexplored. We address this gap using a systematic sequence- and structure-based analysis of APRs in protein-carbohydrate interfaces using a curated dataset of 1119 carbohydrate-binding protein chains. We identified a distinct subset of residues in APRs that bind to carbohydrates, termed functional aggregation-prone regions (fAPRs). Our results show that ~40% of carbohydrate-binding residues overlap with APRs, indicating that aggregation propensity is tolerated within functional sites. The non-overlapping carbohydrate-binding residues are in spatial proximity to the APRs, contributing to intermolecular interactions. Notably, ~85% of fAPRs harbour at least one likely pathogenic missense mutation, highlighting their utility in the interpretation of variant pathogenicity. The fAPRs are enriched with aromatic residues, consistent with known carbohydrate-binding mechanisms. Despite the functional role, aggregation propensity and carbohydrate-binding affinity appear to be uncoupled, suggesting that aggregation and binding can be independently modulated. Overall, this study reveals fAPRs as important structural and functional elements at the protein-carbohydrate interface, offering new insights into carbohydrate-recognition and mutant effect prediction.

B-P.33: Construction and Evaluation of a Multimodal Fusion-Based Model for Compound Protein Interaction Prediction
Track: Proteins and structural biology
  • Si Zheng, Chinese Academy of Medical Sciences and Peking Union Medical College, China
  • Jiao Li, Chinese Academy of Medical Sciences and Peking Union Medical College, China


Presentation Overview: Show

Compound protein interaction (CPI) prediction is an important task in lead discovery and drug repositioning. Although in vitro assays are reliable, their high cost and low throughput limit their application in large-scale screening. Recent deep learning methods have provided efficient computational alternatives for CPI prediction; however, their generalizability in realistic virtual screening settings remains limited, particularly under cold-start scenarios involving unseen compounds or proteins.
Here, we present MoCL-CPI, a multimodal framework for CPI prediction. MoCL-CPI integrates protein sequences, molecular structures, physicochemical descriptors, multiple sequence alignment features, and embeddings derived from a large-scale heterogeneous biomedical network to represent compounds and proteins from complementary structural, functional, topological, and evolutionary perspectives. In particular, the biomedical network-based modality captures global associations and functional context that are not directly available from local structural representations, thereby improving model generalization in challenging screening settings.
We evaluated MoCL-CPI on three public benchmark datasets, including C. elegans, BindingDB, and BioSNAP. Experimental results show that MoCL-CPI achieves improved robustness and predictive performance across diverse screening scenarios. Notably, under inductive data splits that more closely resemble cold-start settings, the proposed framework shows more consistent advantages, suggesting that multimodal collaborative learning provides more stable support for interaction prediction involving unseen entities. These results suggest that multimodal representation learning can be an effective strategy for mitigating the cold-start problem in CPI prediction and may provide a useful framework for computational compound screening.

B-P.34: Order–disorder interfaces in viral pRb inactivation: molecular dynamics insights into SLiM-mediated recognition and implications for interaction databases
Track: Proteins and structural biology
  • Carla Padilla Franzotti, Department of Science and Technology, National University of Quilmes, CONICET, Argentina
  • Nicolas Palopoli, Department of Science and Technology, National University of Quilmes, CONICET, Argentina
  • Gustavo Pierdominici-Sottile, Department of Science and Technology, National University of Quilmes, CONICET, Argentina
  • Henning Hermjakob, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), United Kingdom


Presentation Overview: Show

Intrinsically disordered regions (IDRs) mediate protein interactions through poorly understood mechanisms. We studied the retinoblastoma protein (pRb) and its interaction with the SV40 Large T antigen (LTSV40), a viral oncoprotein that displaces E2F factors.
Using molecular dynamics and umbrella sampling, we show the LTSV40 LXCXE motif is part of a conserved Order–Motif–IDR architecture. The ordered N-terminal region drives initial pRb recognition via induced folding, adding over 6 kcal/mol to affinity. Simultaneously, the C-terminal IDR undergoes a bent-to-extended transition, sterically occluding the pRb AB cleft to prevent E2F binding. These coupled phenomena are evolutionarily conserved across 14 polyomaviruses, as confirmed by AlphaFold and MobiDB, suggesting a common pRb inactivation strategy.
These results highlight a broader challenge for the field: functionally decisive IDR behaviors such as binding-induced folding and steric occlusion are not yet representable within current interaction data models, even when complementary evidence exists in resources such as DisProt or MobiDB. Bridging this gap through systematic integration of IDR annotations into databases such as IntAct and Complex Portal, supported by projection onto AlphaFold3-predicted complex structures, would enable community-scale identification of complexes where disorder is mechanistically decisive.

B-P.35: Statistical modelling of TCR sequences provides robust quality control for repertoire datasets
Track: Proteins and structural biology
  • Dana Moreno, Department of Fundamental Oncology, Ludwig Institute for Cancer Research, University of Lausanne, Lausanne, Switzerland, Switzerland
  • Giancarlo Croce, Department of Fundamental Oncology, Ludwig Institute for Cancer Research, University of Lausanne, Lausanne, Switzerland, Switzerland
  • David Gfeller, Department of Fundamental Oncology, Ludwig Institute for Cancer Research, University of Lausanne, Lausanne, Switzerland, Switzerland


Presentation Overview: Show

T-Cell Receptors (TCRs) show extensive sequence diversity across T cells. This diversity arises from different choices of V and J genes and from insertions and deletions at the V(D)J junction within the Complementary-Determining Region 3 (CDR3) loop. Here, we quantify how V and J gene usage shapes CDR3 length and amino acid composition. In repertoires of TCRs with either unknown or known specificity, we show that CDR3 length is strongly influenced by the number of germline-encoded CDR3 residues in V and J genes, and that, on average, 80% of CDR3α and 65% of CDR3β residues are determined by V and J gene usage. We further show that inconsistencies between V and J gene annotations and CDR3 sequences can be leveraged to identify potential issues in multiple TCR repertoire datasets. Overall, our study quantifies the impact of V and J gene usage on CDR3 length and amino acid composition and provides a robust framework for quality control of TCR repertoire sequencing data.

B-P.36: Uncovering Hidden Proteomic Landscapes: Isoform-Resolved Drug Response Analysis with NOMAD
Track: Proteins and structural biology
  • Jesse Angelis, Computational Mass Spectrometry, Technical University of Munich, Freising, Germany, Germany
  • Mathias Wilhelm, Computational Mass Spectrometry, TUM, Freising, Germany; Munich Data Science Institute (MDSI), TUM, Garching, Germany, Germany


Presentation Overview: Show

Modern dose resolved proteomics platforms provide unprecedented maps of drug induced expression changes. These experiments measure protein abundance across drug concentrations to record cellular perturbations, identifying drug targets and modes of action. While invaluable, these datasets are restricted to protein groups (clusters of proteins sharing identified peptides) - an inherent shortcoming of bottom-up proteomics. By aggregating molecular signals, these approaches overlook the functional diversity of individual isoforms, potentially obscuring critical regulatory events.
We hypothesised that drug induced perturbations often target specific isoforms, leading to ""signal washing"" at the protein group level. To resolve this, we present NOMAD, an analytical framework that deconvolves isoform specific trajectories from dose response data. Utilizing Non-negative Matrix Factorization (NMF), NOMAD separates the relative contributions of individual isoforms to reveal latent regulatory profiles that remain invisible to standard aggregation methods.
Applying NOMAD to proteomic measurements of dose-resolved drug treatments of three drugs in Jurkat cells, we identified widespread isoform specific effects. In an analysis of 3,939 shared protein groups, NOMAD suggests substantial regulations missed by global analysis. Notably, 81.7% of NOMAD exclusive findings were attributed to ""Internal Friction,"" where divergent isoform trends neutralize the global signal. Furthermore, 16.2% of shared significant proteins exhibited ""Isoform Flipping,"" where a sub variant responds in opposition to the canonical trend. While these initial results require further experimental validation, they suggest that isoform resolution is necessary to understand the true molecular complexity of drug cell interactions otherwise missed by protein group analysis.

B-P.37: Whole-Exome Sequencing and Structural Modeling Identify Candidate Drivers in a Multiplex Multiple Sclerosis Family
Track: Proteins and structural biology
  • Simone Bonora, Department of Chemistry and Biology A. Zambelli, University of Salerno, Fisciano (SA), Italy, Italy
  • Anna Marabotti, Department of Chemistry and Biology A. Zambelli, University of Salerno, Fisciano (SA), Italy, Italy
  • Carla Lintas, Research Unit of Medical Genetics, Department of Medicine, Campus Bio-Medico University of Rome, Rome, Italy, Italy
  • Claudio Tabolacci, National Center for Rare Diseases, Istituto Superiore di Sanità, Italy
  • Maria Luisa Scattoni, National Center for Rare Diseases, Istituto Superiore di Sanità, Italy
  • Fioravante Capone, Operative Research Unit of Neurology, Fondazione Policlinico Universitario Campus Bio-Medico, Rome, Italy, Italy
  • Mariagrazia Rossi, Operative Research Unit of Neurology, Fondazione Policlinico Universitario Campus Bio-Medico, Rome, Italy, Italy
  • Vincenzo Di Lazzaro, Operative Research Unit of Neurology, Fondazione Policlinico Universitario Campus Bio-Medico, Rome, Italy, Italy
  • Fiorella Gurrieri, Research Unit of Medical Genetics, Department of Medicine, Campus Bio-Medico University of Rome, Rome, Italy, Italy


Presentation Overview: Show

Multiple sclerosis (MS) is a chronic inflammatory and neurodegenerative disorder of the central nervous system whose heritability is only partly explained by common-risk loci. To investigate the contribution of rare coding variants in familial disease, we applied whole-exome sequencing (WES) to a multigenerational Italian multiplex family and integrated segregation analysis with structural modeling and comparative structural analysis.
WES identified 47 rare co-segregating variants, which were refined to three top candidates after evaluation in unaffected relatives: RTN4 p.Pro148Leu, JAK2 p.Phe560Val, and DUOX2 p.Tyr1150Cys. For JAK2, we selected the PDB structure 6BS0, whereas for DUOX2 we used the AlphaFold model AF-Q9NRD8-F1, since no experimental structure is available and the region containing p.Tyr1150Cys is predicted with high local confidence. Mutant models were generated using the Mutate Model procedure derived from MODELLER. Comparative structural analysis was performed with RING 4.0, DSSP, NACCESS, DynaMut2, DUET, and INPS-MD.
In JAK2, p.Phe560Val disrupts a π–π stacking interaction with Phe547, alters local solvent accessibility and van der Waals contacts, and may perturb a nearby ligand-binding pocket, overall supporting a destabilizing effect. In DUOX2, p.Tyr1150Cys abolishes π–π interactions with Phe591/Phe598 and hydrogen bonds with Ser1153/Val1154 within the ferric reductase-like transmembrane region, again supporting local structural destabilization. By contrast, RTN4 could not be interpreted structurally with confidence because the variant maps to a predominantly disordered region. Overall, this integrative WES-to-structure workflow supports an oligogenic model of MS susceptibility and prioritizes JAK2 and DUOX2 for functional follow-up.

B-P.38: autoeval: A robust toolkit for standardized ad-hoc protein language model evaluation
Track: Proteins and structural biology
  • Sebastian Franz, Technical University of Munich, Germany
  • Aleena Siji, Helmholtz AI, Munich, Germany
  • Lisa M. Spindler, Technical University of Munich, Germany
  • Michael Heinzinger, Institute of Computational Biology, Helmholtz Munich, Germany
  • Burkhard Rost, Technical University of Munich (TUM), Germany


Presentation Overview: Show

Protein Language Models (pLMs) became an essential cornerstone when working with proteins over the last years. For the development of new pLMs, however, it is essential to put the performance of the new model into perspective of existing ones. To complement existing evaluation frameworks which tend to focus on variant effect prediction (VEP), we are presenting autoeval, a robust toolkit for standardized ad-hoc pLM evaluation. Autoeval checks the representation quality of pLM by performing linear probing on several curated, established supervised prediction tasks. We provide validation splits to allow direct integration of autoeval into model development workflows without risking information leakage. Once hyperparameters are set and pLM pre-training is finished, autoeval can be switched into evaluation mode to assess the final model quality on established test sets, putting the new model's performance directly into perspective of existing ones. For simplifying comparison to existing work, we also support automated zero-shot evaluation on de-facto standard VEP benchmarking with ProteinGym as well as contact map prediction. To simplify usability, autoeval can be used via a simple Python API that directly integrates with a comprehensive web platform to analyze the results. The toolkit is available publicly at https://autoeval.biocentral.cloud.

B-P.39: Interpreting Structural Representations and Attention in Pairformer-Based Protein Folding Models for Protein–Protein Interactions
Track: Proteins and structural biology
  • Wout Keymis, VIB center for AI and Computational Biology, KU Leuven, Belgium
  • Wannes Hermans, VIB center for AI and Computational Biology, KU Leuven, Belgium
  • Joana Pereira, VIB center for AI and Computational Biology, KU Leuven, Belgium


Presentation Overview: Show

Understanding how modern protein structure prediction models encode structural information remains a challenge in computational biology. This study investigates how structural features emerge across layers in a Pairformer-based folding model (Boltz1) applied to protein–protein complexes.

Intermediate single-sequence and pair embeddings are extracted from each layer and analyzed in relation to the model's attention mechanisms. To disentangle these embeddings, layer-wise sparse autoencoders (SAEs) with top-k sparsity are trained, yielding interpretable and selectively activated latent features.

Individual latent dimensions are then probed for structural information using linear classifiers, evaluating feature separability (ROC-AUC) and activation-conditioned presence (F1). This enables layer-by-layer tracking of how structural signals are encoded and refined.

In parallel, attention patterns are analyzed, focusing on triangle self-attention and single-sequence attention with pair bias. Attention patterns are evaluated using two complementary analyses: statistical enrichment of structural features and SAE activations in high-attention regions, and co-occurrence between attention and these signals.

Sparse latent factors are found to align with meaningful structural properties, while specific attention heads exhibit strong specialization and evolve across layers. For example, early layers often emphasize residue pairs that are both sequentially and spatially proximal, whereas later layers tend to shift toward predominantly spatial proximity with reduced dependence on sequence adjacency. Similarly, early attention shows broad sensitivity to cysteine residues, while deeper layers refine this signal to a more targeted focus on disulfide bridges. In addition, single-sequence attention heads can highlight binding interface hotspots.

These results provide an initial view of how Pairformer-based models represent protein structure.

B-P.40: Structure based scoring of AI-predicted peptide-GPCR complexes for receptor annotation in a non-model insect genome
Track: Proteins and structural biology
  • Febrina Margaretha, Genetics Program, Graduate Institute for Advanced Studies, SOKENDAI, Japan
  • Mika Sakamoto, Genome Informatics Laboratory, National Institute of Genetics, Japan
  • Hitomi Seike, Graduate School of Frontier Sciences, University of Tokyo, Japan
  • Shinji Nagata, Graduate School of Frontier Sciences, University of Tokyo, Japan
  • Yasukazu Nakamura, Genome Informatics Laboratory, National Institute of Genetics, Japan
  • Takako Mochizuki, Genome Informatics Laboratory, National Institute of Genetics, Japan


Presentation Overview: Show

Genome annotation is essential for linking DNA sequences to biological functions. Neuropeptides are neuromodulatory molecules regulating diverse physiological processes and behaviors, including development, feeding, and courtship, thereby making accurate annotation of the genes encoding them and their receptors critical. Neuropeptides exert their effects by binding to specific receptors, most often G protein-coupled receptors (GPCRs). In non-model insects like Gryllus bimaculatus, annotation of neuropeptides and their receptors remains limited. Our previous work has improved neuropeptide gene functional annotation in this species through de novo genome assembly and homology-based searches against insect databases. However, receptor annotation remains incomplete as sequence‑based homology approaches struggle to assign precise functional annotations despite high sequence similarity among GPCRs and limited availability of reference sequences.

In this study, we explore a structure‑based strategy to assign candidate neuropeptide receptors in the G. bimaculatus by assessing predicted peptide–receptor complexes using AlphaFold3 and Boltz‑2. Focusing on the major classes of insect neuropeptide receptors, rhodopsin‑like and secretin‑like GPCR families, we systematically assign annotated neuropeptides with candidate receptor structures and predict their complexes. To evaluate these combination ligand and receptor models, we apply interface‑focused confidence metrics: interface pLDDT, which measures the local reliability of residues at the binding interface, and ipSAE that summarizes conformational plausibility of the interface geometry and alignment features, to score the structural validity of each peptide–receptor interaction. By ranking structure‑aware confidence measures for G. bimaculatus neuropeptide putative receptors, we construct a list of prioritized candidates that can guide subsequent experimental validation and improve genome annotation in non-model organisms.

B-P.41: AMP-DiT: Antimicrobial Peptide Design with AMP-classifier Conditional Diffusion Transformers
Track: Proteins and structural biology
  • Alireza Noroozi, Koc university, Turkey
  • Ozlem Keskin, Professor of Chemical and Biological Engineering, Koc Univeristy, Turkey
  • Attila Gürsoy, Professor of Computer Science, Koc University, Turkey


Presentation Overview: Show

Antimicrobial resistance is a global health threat. Antimicrobial peptides (AMP) can help for their potent and promising ability to fight resistant pathogens. While AI is being employed to advance AMP discovery, deep learning methods remain limited by poor controllability, suboptimal sequence representations, and low experimental hit rates. We introduce AMP-DiT, a conditional discrete diffusion framework that generates AMPs directly in sequence space, avoiding reliance on continuous latent embeddings and large protein language model priors.
Our approach leverages a denoising diffusion transformer architecture with classifier guidance, enabling high fitness antimicrobial peptide properties during generation. Crucially, we condition the generative process on biologically meaningful signals derived from external AMP predictor, Macrel, explicitly biasing sampling toward sequences with high predicted antimicrobial activity. Since AMP activity is the primary determinant of downstream success, this conditioning serves as a strong inductive bias for generating functional peptides.
We observe that conditioning on a single AMP predictor not only improves scores on that model but also generalizes across other independent AMP prediction frameworks, suggesting that the model captures underlying antimicrobial features rather than overfitting to a specific predictor.
Unlike prior approaches that depend on protein language models trained on proteins, AMP-DiT is trained on peptide-specific data with guidance signals, reducing representation mismatch and avoiding biases inherited from global protein distributions. Overall, AMP-DiT establishes a framework for AMP generation that outperforms existing AMP design methods across key evaluation metrics. As we got 10 percent lower MIC in an AMP prediction model, meanwhile keeping diversity.

B-P.42: AMP-BindGen: Guided Protein Design for Antimicrobial Peptides and Targeted Inhibitors Against AMR
Track: Proteins and structural biology
  • Alireza Noroozi, Graduate Sience and Engineering school of Koc university, Turkey
  • Ani Akpinar, KUISCID, Molecular Biology and Genetics, Koc university, Turkey
  • Erkmen Erken, Computer Science, Koc university, Turkey
  • Eylül Sevinç Sevinç, Computer Science, Koc university, Turkey
  • Fusun Can, School of Medicine, Koc university, Turkey
  • Önder Ergönül, School of Medicine, Koc university, Turkey
  • Ozlem Keskin, Professor of Chemical and Biological Engineering, Koc Univeristy, Turkey
  • Attila Gürsoy, Professor of Computer Science, Koc University, Turkey


Presentation Overview: Show

Antimicrobial resistance (AMR) is rapidly outpacing antibiotic development, posing a critical global health threat. Many resistance mechanisms are mediated by specific bacterial proteins, suggesting that designing binders to inhibit these targets, alongside discovering novel antimicrobial peptides (AMPs) offers a promising therapeutic strategy. Here, we introduce AMP-BindGen, a framework that explores the ability of modern protein design methods to jointly address targeted binding and antimicrobial activity. Building on Boltzgen, we guide sequence generation using an external AMP predictor, Macrel, effectively coupling structural/biophysical design with functional antimicrobial constraints. Our approach first generates diverse peptide candidates using a flexible generative prior trained on AMP-like sequences. We then incorporate activity-aware guidance through Macrel-based scoring, applying rejection sampling, soft reweighting, and diversity-preserving selection to approximate the conditional distribution of peptides with high antimicrobial potential. This allows us to explicitly bias generation toward functional AMPs while retaining sequence diversity. Importantly, we extend this paradigm beyond unconstrained AMP generation by evaluating whether binder design frameworks can be repurposed to produce AMPs that also target key AMR-related proteins, bridging two traditionally separate design objectives: binder design and antimicrobial function. In silico results show that AMP-BindGen significantly enriches sequences with high predicted antimicrobial activity compared to unguided generation, while maintaining broad physicochemical diversity. Our findings suggest that integrating binder design principles with AMP activity guidance provides a powerful and practical route for generating multifunctional peptide candidates, opening new directions for combating antimicrobial resistance. Adding Guidance doubled the generated binders' AMP-likeliness while maintaining binding metrics.

B-P.43: Missense 3D portal, enabling investigation and interpretation of missense variants in their structural context
Track: Proteins and structural biology
  • Ryan Pye, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Ifigenia Tsitsa, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Gordon Hanna, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Anja Conev, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Haotian Zhao, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Grace Walker, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Wendy Tran, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Olivia Simmonds, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Suhail Islam, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Michael J E Sternberg, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom
  • Alessia David, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United Kingdom


Presentation Overview: Show

The expansion of the three-dimensional coverage of the human proteome and those of other organisms, achieved with AlphaFold, offers the unprecedented opportunity to assess the structural impact of millions of missense variants.

The Missense3D Portal, freely available at https://missense3d.bc.ic.ac.uk/, is a suite of algorithms and databases that leverage the availability of experimental and modelled 3D structure, including protein complexes, to provide mechanistic explanations for the damaging effect of missense variants, such as disruption of chemical bonds and steric clashes.

Structural predictions for >9 million missense variants, produced by our core algorithm, Missense3D, are freely available through the Missense3d-αDB database and will soon be integrated into ProtVar at EBI. Furthermore, structural predictions for >4 million human variants, generated using experimental structures and 3D models produced with our Phyre2 homology modelling tool, are available from the Missense3D-DB database and integrated into DECIPHER at EBI.

Missense3D-PTM, published in JMB December 2025, is the latest addition to the Missense3D Portal. Developed as part of the UK Human Functional Genomics Initiative, it is a “one-stop-shop” for dynamically investigating missense variants in the context of post-translational modifications (PTMs). Its user-friendly web interface provides mapping of 11,490,257 naturally-occurring human missense variants and 334,255 PTM sites (spanning >50 PTM types, including phosphorylation, glycosylation, and sumoylation) along with their neighbouring residues in sequence and 3D structure space, using AlphaFold models for 20,431 proteins. Furthermore, precalculated distances in 3D space between any PTM and any residue, enable visualization and exploration of novel variants currently not included in the Missense3D-PTM database.   

B-P.44: Designing Novel Enzymes with a Transformer-Guided Generative AI Framework
Track: Proteins and structural biology
  • Ufuk Tanrıverdi, Hacettepe University, Turkey
  • Tunca Doğan, Hacettepe University, Turkey


Presentation Overview: Show

Designing functional proteins is central to therapeutics and biotechnology, but remains difficult due to the astronomical sequence space and the loose mapping between sequence, structure, and function. Generative AI offers a principled way to explore this space by learning the patterns underlying functional proteins. We propose a generative adversarial network (GAN)-based framework for de novo protein design where both generator and critic are initialized from a masked protein language model, ProtBERT, and fine-tuned in two-stages, first on a large enzyme corpus, then on subclass-specific data, using dynamic masking. The generator performs iterative partial decoding, retaining only the top-confidence ~10% of tokens per step while remasking the rest until convergence. We evaluate two generation modes: generation from 90%-masked real sequences and fully blind generation from entirely masked inputs. As a case study, we apply the framework to DNA methyltransferases (DNMTs), particularly DNMT3A, which regulate gene expression with epigenetic modification and are essential for genome stability. Its dysregulation is linked to developmental disorders and cancer, making novel enzyme design valuable. Generated sequences undergo a multi-stage evaluation pipeline: ESMFold pLDDT, self-consistency via ProteinMPNN, DNMT-like structural similarity via Progres, and pairwise TM-score diversity serve as initial filters. Top candidates proceed to AlphaFold3 co-folding with DNA and SAM, with DNA-SAM binding distance as the final functional criterion. Our best candidate exhibited secondary structure organization resembling reference DNMT-3a with spatially similar alpha-helical and beta-sheet elements, and retained the hydrophilic interactions necessary for SAM-mediated DNA methylation, despite being a completely novel sequence.

B-P.45: Benchmarking Classical and Modern Structural Alignment Algorithms for Protein–Protein Interfaces in Template-Based Docking
Track: Proteins and structural biology
  • Reza Mohammadian Mazraeh Shadi, Graduate School of Sciences and Engineering, Koç University, Turkey
  • Fatma Cankara, Graduate School of Sciences and Engineering, Koç University, Turkey
  • Nurcan Tuncbag, Department of Chemical and Biological Engineering, and School of Medicine, Koç University, Turkey
  • Attila Gursoy, Department of Computer Engineering, Koç University, Turkey
  • Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Turkey


Presentation Overview: Show

Protein–protein interactions are pivotal for many cellular functions, and understanding their underlying mechanisms is key to elucidating biological processes. Among several approaches for identifying protein–protein interactions, template-based docking is particularly powerful, with performance strongly governed by template library selection and structural alignment quality. In this study, we systematically assess structural alignment algorithms for detecting structurally similar protein–protein interfaces in template-based docking.
Using a curated subset of our PiFace interface dataset, representing non-contiguous, sequence-order–independent binding regions (9,775 interfaces in 78 clusters), together with the MALISAM and MALIDUP motif/domain benchmarks and their non-sequential variants (742 pairs total), we compare TM-align with more recent algorithms, including US-align, Multiprot, KPAX and KPAX-flex, GTalign, Foldseek, and DeepAlign. Performance is evaluated in docking-relevant interface-to-global and interface-to-interface scenarios using TM-score, root-mean-square deviation (RMSD), and alignment length.
Across more than 2.6 million interface-centered docking case comparisons, TM-align, US-align, and GTalign achieve consistently strong performance, with mean TM-scores of ~0.67–0.80 in docking-relevant cases, mean RMSDs of ~1.3–1.6 Å, and normalized alignment coverage of ~0.83–1.00. In positive controls, TM-scores approach 1.0 with near-complete coverage. KPAX-flex often attains low RMSDs with shorter alignments. GTalign approach matches TM-align's accuracy while providing substantial GPU-enabled speed-ups for enabling large-scale template screening. Protein language model–based tools, such as Foldseek, benefit from geometric re-scoring for rapid template retrieval, although they may over-extend alignments. Overall, our results delineate trade-offs among accuracy, coverage, and computational cost, and provide practical guidance for interface-driven, template-based docking pipelines.

B-P.46: The Role of Local Energetic Frustration Across Conformational Ensembles of Disordered Proteins
Track: Proteins and structural biology
  • Franco L Simonetti, Barcelona Supercomputing Center, Spain
  • Edgar P. Chacón, Barcelona Supercomputing Center, Spain
  • Maria I. Freiberger, Sorbonne Université, CNRS, IBPS, Laboratory of Computational and Quantitative Biology, France
  • Alexander M. Monzón, Department of Information Engineering, University of Padova, Padova, Italy, Italy
  • Leandro Radusky, Mimark Diagnostics S.L., Parc Cientific Barcelona, Spain
  • Diego U. Ferreiro, Protein Physiology Lab, Facultad de Ciencias Exactas y Naturales, Universidad de Buenos Aires, Argentina
  • R. Gonzalo Parra, Barcelona Supercomputing Center, Spain


Presentation Overview: Show

Globular proteins satisfy the minimum frustration principle, i.e. their folding landscapes are strongly funneled towards their native state with residual energetic conflicts often associated with function. In contrast, intrinsically disordered proteins populate heterogeneous conformational ensembles as a consequence of an excess of competing interactions and therefore flatter energy landscapes. Still, the specific energetic signatures that lead to protein disorder are poorly understood. Here, we characterized the local frustration patterns of 2,964 proteins and 54,766 conformers spanning different flavours of intrinsic disorder.
We found that in the case of direct residue-residue interactions, except for a minor contribution of specific interaction types, disordered residues are not excessively highly frustrated; instead, they are enriched in neutrally frustrated interactions and a reduced prevalence of minimally frustrated ones. Furthermore, we found that the most distinguishable feature between ordered and disordered residues comes from long range interactions where disordered regions are depleted of stabilizing interactions and enriched in energetic conflicts. When disordered regions are analysed in complex with known protein partners, their energetic profiles resemble those of ordered residues.
We conclude that intrinsic disorder is not a consequence of excessive energetic conflicts as believed for the case of statistical random heteropolymers. Instead, disordered proteins exist at the edge of foldability, where a small number of stabilizing interactions with other proteins or substrates can shift their equilibrium towards more structured, functional conformations.

B-P.47: A Replicate-Based Molecular Dynamics Framework for Mapping Variant-Induced Structural and Network Perturbations in Drug-Metabolising NAT2 Enzyme
Track: Proteins and structural biology
  • Wayde Veldman, Rhodes University, South Africa
  • Ozlem Tastan Bishop, Rhodes University, South Africa


Presentation Overview: Show

Investigating how residue variations alter enzyme structure and dynamic communication remains essential for understanding inter-individual variation in drug metabolism. We performed all-atom molecular dynamics simulations of human arylamine N-acetyltransferase 2 (NAT2), the enzyme responsible for metabolising the tuberculosis drug isoniazid, to determine how sequence variation drives the transition from rapid to slow acetylation phenotypes.

We developed a replicate-based comparative framework in which replicate-averaged rapid acetylator simulations define a reference state. Slow acetylator deviations were quantified relative to the reference mean and standard deviation across hydrogen bonding patterns, residue flexibility, and dynamic residue network centrality metrics which are then mapped to enzyme structure. This approach enables systematic identification of structurally and functionally meaningful perturbations rather than relying on single-trajectory comparisons.

Slow acetylator variants exhibited destabilisation of the active conformation characterised by altered hydrogen bonding, increased residue flexibility, and disrupted residue communication networks. In the R64Q+K268R variant for example, loss of hydrogen bonding at residue 64 propagated through the structural network, reducing betweenness centrality of catalytic residue D122 and putative isoniazid-binding residues S125 and F217. On the other hand, eigenvector centrality was reduced for N72, D122, and active-site loop G124, indicating reduced residue communication. In line with these changes, altered active-site geometry was observed via shortening of distances between binding residue F217 and catalytic residues H107 and D122.

Together, these results provide mechanistic insight into how distal residue variations propagate through dynamic residue networks to impair catalytic organisation, offering a transferable framework for studying structure/function relationships in pharmacogenetically relevant enzymes.

B-P.48: Benchmarking Deconvolution Tools for Proteomics Data
Track: Proteins and structural biology
  • Tobias Scheithauer, Institute for Machine Learning, ETH Zurich, Switzerland
  • Ludovica Sibilia, Institute for Machine Learning, ETH Zurich, Switzerland
  • Mira Herold, Institute for Machine Learning, ETH Zurich, Switzerland
  • Samuel Gair, Institute for Machine Learning, ETH Zurich, Switzerland
  • Marta Pinto Carbo, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
  • Lopamudra Chatterjee, Swiss Institute of Allergy and Asthma Research (SIAF), University of Zurich, Switzerland
  • Nadia Djerbi, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
  • Dorothea Rutishauser, Department of Pathology and Molecular Pathology, University Hospital Zurich, Switzerland
  • Rosary Yao, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
  • Thi Huong Lan Do, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
  • Marco Bühler, Department of Pathology and Molecular Pathology, University Hospital Zurich, Switzerland
  • Christoph Messner, Swiss Institute of Allergy and Asthma Research (SIAF), University of Zurich, Switzerland
  • Thorsten Zenz, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
  • Valentina Boeva, Institute for Machine Learning, ETH Zurich, Switzerland


Presentation Overview: Show

The cellular composition of the tumor microenvironment is essential for diagnosis, prognosis, and treatment selection in precision oncology. Computational deconvolution of bulk omics data can infer cell type proportions from mixed samples without physical cell sorting. While numerous deconvolution tools exist for transcriptomics, very few target bulk proteomics data. The distinct nature of proteomic data raises the question whether transcriptomics-derived tools can be reliably applied to proteomics without adaptation.

We conducted a multi-dimensional benchmarking study to evaluate the applicability of existing deconvolution methods to bulk proteomics data. Using two public immune cell proteomics, we generated synthetic bulk mixtures with known cell type proportions in different mixture composition settings. Our benchmark evaluates six deconvolution algorithms using different imputation strategies for missing values and various signature matrix construction methods.

Our benchmarking reveals a substantial performance degradation of all algorithms when moving from idealistic to more realistic simulation settings, indicating that published benchmarks may overestimate real-world accuracy. Second, the optimal algorithm choice depends on the combination of preprocessing and experimental scenario. Third, we observe systematic cell-type-specific biases showing as consistent over- or underestimation of individual cell types. Fourth, small signatures based on well-characterized marker proteins may be more robust than larger data-driven signature matrices. Finally, unsupervised methods demonstrate limited reliability across the range of tested conditions.

Our findings motivate both the development of dedicated, proteomics-aware deconvolution tools and the creation of experimentally validated benchmark datasets to enable robust immune cell profiling in precision oncology.

B-P.49: Why Think Multimodal? Data Treatment for Integrated Structural Biology
Track: Proteins and structural biology
  • Wojciech Potrzebowski, SciLifeLab Lund, Centre for Research Infrastructure in Health and Sciences, Lund University, Lund, Sweden, Sweden
  • Filip Arman, Centre for Research Infrastructure in Health and Sciences, Lund University, Lund, Sweden, Sweden
  • Joao Figueira, Swedish NMR Centre, Department of Chemistry, UmeÃ¥ University, UmeÃ¥, Sweden, Sweden
  • Sayyed Jalil Mahdizadeh, Swedish NMR Centre, Department of chemistry and molecular biology, University of Gothenburg, Sweden., Sweden
  • Anton Sellerberg, Centre for Research Infrastructure in Health and Sciences, Lund University, Lund, Sweden, Sweden
  • Simon Ekstrom, Centre for Research Infrastructure in Health and Sciences, Lund University, Lund, Sweden, Sweden
  • Hanna Kultima, Department of Immunology, Genetics, and Pathology, Uppsala University, Uppsala, Sweden, Sweden
  • Johan Rung, Department of Immunology, Genetics, and Pathology, Uppsala University, Uppsala, Sweden, Sweden
  • Nicholas Pearce, Department of Cell and Molecular Biology, SciLifeLab, Uppsala University, Sweden, Sweden


Presentation Overview: Show

Integrated structural biology increasingly depends on combining data from multiple experimental and computational approaches to answer biological questions that cannot be resolved by any single technique alone. However, multimodal projects are still often organized around individual methods, which creates fragmentation in data handling, metadata capture, quality control, and downstream reuse. This slows scientific progress, makes cross-technique interpretation more difficult, and increases the risk that valuable intermediate data, context, and decisions are lost before publication or deposition.

At the SciLifeLab Integrated Structural Biology (ISB) platform, multimodal data treatment is being developed as a way to shift focus from techniques to scientific questions. By treating structural biology data as connected evidence streams rather than isolated outputs, it becomes possible to use one modality to support, validate, or troubleshoot another, and to follow the sample and data journey from production to integrated interpretation. Such an approach is also essential for making data FAIR, reusable, and increasingly AI-ready, with sufficiently rich metadata, validation, and provenance to support future method development and cross-study analysis.
Here we describe the rationale, challenges, and practical principles for multimodal data treatment at the SciLifeLab ISB platform, with a focus on data integration across techniques, preservation of experimental context, and preparation of high-quality datasets for sharing, reuse, and AI-driven analysis.

B-P.50: Reading TEA-Leaves for de novo protein design
Track: Proteins and structural biology
  • Ieva Pudžiuvelytė, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland
  • Lisa Brandenburg, Department of Biosystems, Science and Engineering, ETH Zürich, Basel, Switzerland, Switzerland
  • Lorenzo Pantolini, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland
  • Basile Wicky, Department of Biosystems, Science and Engineering, ETH Zürich, Basel, Switzerland, Switzerland
  • Janani Durairaj, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland


Presentation Overview: Show

The field of de novo protein design aims to produce novel protein sequences or even novel structures with improved functional properties for biotechnological applications. The breakthrough in protein folding methods opened ways to evaluate whether sequences fold into intended structures in silico, which catalysed the protein design progress. Orthogonally, we witnessed the appearance of protein language models (pLMs) encoding biological properties in their embeddings. However, pLM-based iterative optimization approaches requiring repeated structure evaluations remain computationally expensive, limiting their practical throughput.

We propose that the structure-informed alphabet derived from pLM embeddings (TEA) could address this bottleneck. Therefore, we present TEA-Leaves method leveraging TEA alphabet to efficiently guide the Markov chain Monte Carlo sampling to generate de novo sequences folding into template structures.

We showcase TEA-Leaves functionality in several scenarios by generating: de novo sequences folding into a diverse set of monomeric structures and sequence-similar but structure-dissimilar pairs of proteins.

In addition to structure prediction confidence metrics from ColabFold, we used several in silico quality metrics to select designs for experimental validation: values of TEA-Leaves losses, solubility, and aggregation scores. Namely, we selected designs with sequence novelty relative to natural proteins to assess the ability of TEA-Leaves to explore sequence space beyond homology adjacency. The quality scoring filters were collected in a pipeline Sieve to prioritize the candidates for the downstream experimental validation of expressibility, solubility, and oligomerization propensity.

Altogether, using pLM-derived structural signals we aim to probe the determinants of protein folding outside of the homology barrier.

B-P.51: Identification of Highly Conserved Conformational Epitopes on the SARS-CoV-2 Spike Glycoprotein: A Computational Foundation for Broad-spectrum Antibody Design
Track: Proteins and structural biology
  • Cheng-Wei Cheng, Kaohsiung Medical University, Taiwan


Presentation Overview: Show

The continuous emergence of SARS-CoV-2 variants underscores the critical need for broad-spectrum therapeutic antibodies. Identifying highly conserved conformational epitopes on the spike glycoprotein, which represent regions that remain invariant across diverse lineages, is essential for developing treatments resistant to viral escape. This study utilizes an integrated computational pipeline to pinpoint these critical structural motifs.
Our framework utilized reconstructed high-resolution structures of the SARS-CoV-2 spike glycoprotein, which were computationally refined to restore missing residues and fully glycosylated to reflect native physiological conditions. By calculating the Relative Solvent Accessibility (RSA) of each residue, we categorized amino acids as exposed, shielded, or buried. This step ensures that candidate epitopes are not only conserved but also physically accessible to antibodies. Simultaneously, we performed an evolutionary analysis using the covSPECTRUM platform to track recent mutation trajectories, identifying residues with exceptionally high conservation across the latest global variants.
By synthesizing RSA data with evolutionary conservation scores, we identified and validated several highly conserved conformational epitopes. These spatial clusters represent ""vulnerability patches"" that are evolutionarily constrained and structurally accessible. The preliminary results demonstrate that these validated regions are distinct from highly variable domains prone to mutation. These findings highlight that these highly conserved conformational epitopes serve as high-priority targets for antibody design. Our study provides a strategic roadmap for engineering next-generation biologics capable of maintaining long-term potency against the evolving landscape of SARS-CoV-2 and future sarbecoviruses.

B-P.52: BiochemicalAlgorithms.jl for Stochastic Protein Energy Minimization
Track: Proteins and structural biology
  • Jennifer Leclaire, Institute for Computer Science, JGU Mainz, Germany
  • Thomas Kemmer, Institute of Computer Science, Johannes Gutenberg University Mainz, Germany
  • Andreas Hildebrandt, Institute of Computer Science, Johannes Gutenberg University Mainz, Germany


Presentation Overview: Show

Computational methods for structural biology often begin with rapid prototyping in Python. Performance demands, however, frequently lead to implementations in C or C++, resulting in software packages with C++ cores and Python user interfaces. The Biochemical Algorithms Library (BALL) is a representative example of this approach: a well-established C++ framework initiated in 1996, with Python bindings added in 2010. BALL was once the largest open-source library of its kind, offering broad functionality for molecular structure analysis.
This motivated the redesign of the library and led to our Julia-centric framework, BiochemicalAlgorithms.jl. It provides core data structures, file I/O, preprocessing, molecular mechanics, optimization algorithms, and Julia ecosystem integration across the structural bioinformatics pipeline. Moving development to Julia has simplified the implementation of key design goals, particularly rapid application development.
Building on this recently published framework, we will demonstrate its practical value through a stochastic gradient descent-based protein energy minimization scheme. While classical energy minimization is often performed with deterministic methods such as conjugate gradient, stochastic optimization, as used in deep learning, introduces controlled randomness that can support broader exploration of the protein energy landscape and help avoid premature trapping in local minima. In this sense, the example serves both as an optimization strategy and as a demonstration of how quickly new algorithmic ideas can be implemented within BiochemicalAlgorithms.jl. More broadly, this work highlights how the framework supports efficient method development and provides a practical environment for building computational tools in structural bioinformatics.