View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
B-P.01: STAR-GO: Improving Protein Function Prediction by Learning to Hierarchically Integrate
Ontology-Informed Semantic Embeddings
Track: Proteins and structural biology
-
Mehmet Efe Akça, BoÄŸaziçi Üniversitesi, Turkey
- Gökçe UludoÄŸan, Bogazici University, Turkey
- Arzucan Ozgur, Bogazici University, Turkey
- Inci BaytaÅŸ, Bogazici University, Turkey
Presentation Overview: Show
Motivation: Accurate prediction of protein function is essential for elucidating molecular mechanisms and
advancing biological and therapeutic discovery. Yet experimental annotation lags far behind the rapid growth
of protein sequence data. Computational approaches address this gap by associating proteins with Gene Ontology
(GO) terms, which encode functional knowledge through hierarchical relations and textual definitions. However,
existing models often emphasize one modality over the other, limiting their ability to generalize,
particularly to unseen or newly introduced GO terms that frequently arise as the ontology evolves, and making
the previously trained models outdated.
Results: We present STAR-GO, a Transformer-based framework that jointly models the semantic and structural
characteristics of GO terms to enhance zero-shot protein function prediction. STAR-GO integrates textual
definitions with ontology graph structure to learn unified GO representations, which are processed in
hierarchical order to propagate information from general to specific terms. These representations are then
aligned with protein sequence embeddings to capture sequence–function relationships. STAR-GO achieves
state-of-the-art performance and superior zero-shot generalization, demonstrating the utility of integrating
semantics and structure for robust and adaptable protein function prediction.
Availability: Code and pre-trained models are available at https://github.com/boun-tabi-lifelu/stargo
B-P.02: Promiscuous Mitochondrial Targeting Under Proteotoxic Stress Rewires Cytosolic Proteostasis
Track: Proteins and structural biology
-
Shubham Goyal, CSIR-Institute of Genomics and Integrative Biology, India
- Soumen Kundu, CSIR-Institute of Genomics and Integrative Biology, India
- Rishab Singh, CSIR-Institute of Genomics and Integrative Biology, India
- Kausik Chakraborty, CSIR-Institute of Genomics and Integrative Biology, India
Presentation Overview: Show
Maintaining cytosolic proteostasis is essential for cellular viability, and its disruption is a characteristic
feature of neurodegenerative proteinopathies. Although quality control is traditionally associated with the
proteasome and autophagy, the MAGIC (Mitochondria As Guardian In Cytosol) pathway posits that mitochondria
function as crucial inter-organellar hubs. In this study, we elucidate the molecular principles that govern
the stress-induced recruitment of cytosolic proteins to the mitochondrial surface. Through data integration
and reanalysis of ribo-seq profiling datasets, we identify a signal-independent targeting mechanism activated
by proteotoxic stress. While the translatome follows a Gaussian distribution, the mitochondrial interactome
displays a distal shift in ribosomal density. This recruitment is strictly linked to translational progress;
both mitochondrial and cytosolic proteins within the interactome exhibit significantly longer sequences,
particularly for nascent chains that have been translated beyond a critical threshold. Importantly,
biophysical profiling and protein structure analysis differentiate canonical from stress-induced targeting.
Resident mitochondrial proteins undergo rigorous physicochemical filtering for enhanced stability and
structural complexity. In contrast, recruited cytosolic proteins bypass these constraints, displaying
biophysical instability and reduced core strength. These altered protein-protein interactions at the
mitochondrial surface are experimentally validated in yeast; cycloheximide treatment and mitochondrial
isolation reveal a time-dependent accumulation of K48-linked polyubiquitinated proteins. Furthermore,
modulating import capacity through receptor deletion or downregulation mitigates fitness defects associated
with global protein misfolding. Our findings establish mitochondria as dynamic buffers of proteostasis,
underscoring a systems-level trade-off between organelle biogenesis and global protein quality control.
B-P.03: Nucleic acid 3D structure search and alignment with GTalign
Track: Proteins and structural biology
- Mindaugas Margelevicius, Vilnius University, Lithuania
-
Shikhar Rana, vilnius university, Lithuania
Presentation Overview: Show
Structural comparison of nucleic acids, particularly RNA, is essential for identifying evolutionary and
functional relationships that are often not apparent from sequence alone. However, efficient methods for
large-scale search and alignment of nucleic acid three-dimensional (3D) structures remain scarce. Here we
present an extension of GTalign, a high-performance structural alignment method, to support nucleic acid
structures.
The proposed approach enables unified search and alignment across diverse macromolecules while maintaining
high computational efficiency. We evaluated the method on a diverse non-redundant dataset of RNA structures in
an all-against-all benchmark. GTalign achieves improved alignment accuracy compared to existing methods, as
reflected by higher TM-scores and an increased number of significant matches. At the same time, it operates at
substantially lower computational cost, providing considerable speedups across different parameter
settings.
These results demonstrate that GTalign effectively balances sensitivity and efficiency for structurally
diverse RNA molecules. The method supports scalable structural comparison and database search, facilitating
large-scale analyses and the identification of conserved structural motifs in RNAs and other nucleic acids.
B-P.04: PUCAR : Protein unit discovery using community detection on pLM attention-weighted residue
graphs
Track: Proteins and structural biology
-
Nuriye Ozlem Ozcan Simsek, Bogazici University, Turkey
- Burak Suyunu, Bogazici University, Turkey
- Ozdeniz Dolu, Bogazici University, Turkey
- Enes Taylan, Bogazici University, Turkey
- Arzucan Ozgur, Bogazici University, Turkey
Presentation Overview: Show
Motivation: Identifying functional units within protein sequences remains a central challenge in computational
biology. Transformer attention in protein language models (pLMs) captures residue-to-residue relationships
along the sequence. These relationships may reflect how residues organize into groups, thus enabling the
discovery of biologically meaningful protein units (PUs) directly from sequence. Here, we present PUCAR, a
method that converts pLM attention scores into weighted residue graphs and partitions the entire protein
sequence into protein units by assigning each residue to a community through community detection.
Results: Using domain and motif annotations as reference protein units, we show that PUCAR achieves
competitive or improved performance relative to baseline approaches. The method includes a correlation-based
calibration step that identifies attention heads most strongly associated with the target annotations, and
ablation analyses demonstrate that this biologically informed head selection improves sequence partitioning
quality. On independent held-out test sets, PUCAR successfully recovered annotated protein regions despite
being calibrated with limited annotation data. In addition to recovering curated regions, PUCAR enables
sequence-wide decomposition of proteins into units, extending PU discovery beyond curated annotations.
Collectively, these findings establish a link between transformer attention and biologically meaningful
protein sequence organization.
B-P.05: 3Dswappred2: Enhanced Sequence-Based Prediction of Protein Domain Swapping
Track: Proteins and structural biology
-
Dheemanth Regati, National Centre for Biological Sciences, India
- Sarthak Ghatkar, National Centre for Biological Sciences, India
- Sowdhamini Ramanathan, National Centre for Biological Sciences, India
Presentation Overview: Show
Domain swapping is a unique form of protein-protein interaction where proteins that are in a monomeric form
interact with each other to form an oligomeric structure by exchanging domains with each other. Domain
swapping is linked to functional regulation and protein aggregation. We present 3Dswappred2, an updated
machine learning framework that builds on our lab's previous work to predict the propensity for domain
swapping from sequence. Using a curated dataset of approximately 4,500 proteins, we have implemented a
CNN-based model that currently achieves ~85% accuracy, representing a significant performance gain over our
previous tools (Shameer et al. 2011, Upadhyay, A. K., & Sowdhamini, R. (2016).
To further improve predictive power, we are exploring additional architectures and integrating specific
biophysical triggers. A key focus is the prediction of hinge regions, which are the flexible segments
essential for conformational exchange (Shingate et al. 2012), the prediction of the extent of swapping, and
the modelling of pH-dependence. Since pH shifts often act as critical switches for domain swapping in vivo,
incorporating these environmental factors allows for a more physiologically relevant assessment of protein
stability and misfolding (Shingate et al. 2015).
3Dswappred2 aims to provide a robust resource by combining sequence-derived propensity with localised
structural insights. We discuss the current model performance and the ongoing integration of environmental
modulators that influence the swapping equilibrium.
B-P.06: Evaluating Confidence Scores in AlphaFold for Protein Interaction Prediction
Track: Proteins and structural biology
-
Arthur Valentin, VIB AI / KU Leuven, Belgium
- Damien Legros, VIB AI / KU Leuven, Belgium
- Iker Reinares Zapata, UAM, Spain
- Joana Pereira, KU Leuven - VIB.AI, Belgium
Presentation Overview: Show
Accurate prediction of protein protein interactions remains a central challenge in computational biology,
particularly in the context of scalable, high-throughput analyses. Recent advances in structure prediction
models have introduced confidence metrics, such as interface predicted TM-score (ipTM), that are increasingly
used as proxies for interaction likelihood in protein complexes. However, the extent to which such metrics
reliably capture true interaction signals remains insufficiently understood.
In this work, we investigate the determinants and limitations of ipTM as a scoring function for protein
protein interactions. We analyze how residue-level properties, such as structural confidence and spatial
relationships, contribute to the resulting scores, with the aim of clarifying the signals that drive scores
predictions. We further explore how broader contextual factors, such as the composition of protein assemblies,
influence these metrics, and assess their robustness under varying modeling conditions.
In addition, we examine practical considerations related to the computational cost of large-scale interaction
prediction, exploring avenues to make such approaches more efficient and broadly applicable. We also consider
potential directions for improving current scoring strategies, with the goal of enabling more reliable
assessment of protein protein interactions.
Together, our findings contribute to a deeper understanding of confidence metrics in protein complex
prediction.
B-P.07: Scaling and democratising structure-based protein function prediction with
metagenomic-deepFRI
Track: Proteins and structural biology
-
Valentyn Bezshapkin, Institute of Microbiology and Swiss Institute of Bioinformatics, ETH Zurich,
Switzerland
-
Filip Schymik, AGH University of Krakow, Sano Centre for Computational Medicine, Krakow, Poland,
Poland
- Piotr Kucharski, Aiformatics, Krakow, Poland, Poland
-
Pawel Szczerbiak, AGH University of Krakow, Sano Centre for Computational Medicine, Krakow, Poland, Poland
- Jakub Wojciechowski, Sano Centre for Computational Medicine, Krakow, Poland, Poland
- Lukasz Szydlowski, Sano Centre for Computational Medicine, Krakow, Poland, Poland
-
Tomasz Kosciolek, AGH University of Krakow, Sano Centre for Computational Medicine, Krakow, Poland, Poland
Presentation Overview: Show
Proteins drive the functioning of all living matter, from cell structural integrity to complex metabolic
reactions. While high-throughput sequencing has revealed an large diversity of proteins, driving an
exponential expansion of public databases, functional characterization lags behind sequence acquisition. To
aid the experimental characterization, numerous computational approaches were developed. One example is
deepFRI, a GO-term prediction tool that leverages both sequence and structure of a protein. The inclusion of
structure improves both the confidence and specificity of deepFRI's prediction. For large scale metagenomic
experiments however, obtaining high quality protein structures for all discovered sequences may be
computationally challenging. Metagenomic-deepFRI is an extension of the original pipeline, that allows for
automatic retrieval of homologous protein structures from the reference databases such as AlphaFold database
or ESM atlas. It leverages lightweight storage of reference protein structures, MMSeqs2 search for scalable
retrieval of proteins similar to the query, a quick global-alignment between the query and the matched protein
and contact-map construction as a representation of protein structure. Current limitation of
metagenomic-deepFRI is that it only predicts Gene Ontology terms, which might be insufficient for thorough
investigation of a set of proteins, for example, coming from an environmental sample. Many other modalities
could be explored, such as genomic context, oligomeric state of the protein or possible interactions.
Metagenomic-deepFRI is therefore under ongoing development, that aims to transform it into a comprehensive
pipeline for both bioinformaticians running large scale computations and wet-lab oriented scientists looking
for familiarity and ease of use.
B-P.08: Phyre2.2: A web server to predict protein structure and protein/ligand complexes
Track: Proteins and structural biology
- Harold R Powell, Imperial College London, United Kingdom
-
Suhail Islam, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Anja Conev, Imperial College London, United Kingdom
- Eleanor Stevens, Imperial College, United Kingdom
-
Alessia David, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
- Michael Sternberg, Imperial College London, United Kingdom
Presentation Overview: Show
Template-based protein structure prediction remains a powerful complementary approach to the recent
machine-learning algorithms such as AlphaFold and Boltz. Indeed, our Phyre server continues to be widely used
with over 30,000 unique users running over 200,000 jobs in 2025.
Our new release, Phyre2.2, enables a user to input a sequence and the program then identifies the closest
AlphaFold2 model which then acts as a template for model prediction. This is in addition to the traditional
Phyre2 approach of basing the model on an experimental Protein Data Bank (PDB) structure. Using an AlphaFold
structure as a template will be particularly useful when a new proteome has been sequenced and models are not
yet available to the community.
We are also launching Phyre2.2 Ligand which enables a user to obtain a model for a ligand located within a
Phyre2.2 predicted structure. There are two modes for Phyre2.2 Ligand. In both, the first step cavities are
identified in the predicted structure. Then, in the first mode, the coordinates of a ligand in the template
are transplanted into the predicted model to generate a downloadable protein/ligand complex. In the second
mode Phyre2.2 ligand will enable a user to dock a selected ligand into the predicted structure using AutoDock
Vina. The ligand can be from the PDB template, from a UniProt entry or user defined via a SMILE string.
Phyre2.2, an ELIXIR resource, is freely available to all users, including commercial users, at
https://www.sbg.bio.ic.ac.uk/phyre2/ .
B-P.09: UniProt Pan Proteomes: a scalable resource for comparative proteome analyses and exploring species
diversity
Track: Proteins and structural biology
-
Tanushree Tunstall, European Bioinformatics Institute (EMBL-EBI), United Kingdom
- Giuseppe Insana, European Bioinformatics Institute (EMBL-EBI), United Kingdom
- Stephanie Lo, European Bioinformatics Institute (EMBL-EBI), United Kingdom
- John Lees, European Bioinformatics Institute (EMBL-EBI), United Kingdom
- Maria Martin, martin@ebi.ac.uk, United Kingdom
Presentation Overview: Show
The rapid growth of sequence data requires scalable frameworks that capture species-level diversity beyond
single reference proteomes. We present UniProt Pan Proteomes (PP), a resource integrating conserved
(core-like) and variable (accessory) proteins to support comparative proteomics, pathogenicity and resistance
studies, gene essentiality analyses, and population-aware target discovery.
For each species with >=3 eligible proteomes, sequences are clustered with MMseqs2 (90% identity, 50%
coverage) to generate a non-redundant PP. The best annotated protein per cluster is selected using a
hierarchical strategy prioritising reference proteomes and reviewed entries. Each species's PP dataset
includes three outputs: FASTA entries with protein frequency; PP matrix, akin to the pan-genome (PG)
presence-absence matrix; a statistics file summarising proteome composition, clustering, and frequencies.
The upcoming 2026_02 release includes ~3,200 species, created from ~70,000 proteomes, and ~270 million
clustered proteins. Preliminary benchmarking in Escherichia coli shows strong PP and PG concordance in
core-like subsets (96% bidirectional matching), while full-set PP-PG comparisons retain expected accessory
diversity. Inter-species PP comparisons (e.g. E. coli vs. Shigella) identify conserved and divergent
sequences. The Streptococcus pneumoniae PP, built from ~156 proteomes, spans ~40 pneumococcal lineages
showcasing robust coverage of species diversity. It comprises 6,411 protein clusters consistent with recent PG
studies. To demonstrate translational use, we mapped 64 S. pneumoniae vaccine antigens to within- and
across-species PP clusters to identify homologues with functional annotations for target prioritisation.
UniProt PP facilitates cross- species comparative analyses, supporting diverse research use cases. We invite
community feedback to shape future releases.
B-P.10: Probing multi-omic datasets to unravel multi-stability mechanisms in CAZymes from extreme
biomass-rich environments.
Track: Proteins and structural biology
-
Carlos Huertas Díaz, KTH Royal Institute of Technology, Sweden
- Lauren Sara McKee, KTH Royal Institute of Technology, Sweden
- Johan Larsbrink, Chalmers University, Sweden
- Pakinee Thianheng, KTH Royal Institute of Technology, Sweden
Presentation Overview: Show
Carbohydrate-active enzymes (CAZymes) are important biocatalysts for industrial processes where harsh
conditions are present, such as high temperatures, extreme pH, high shear stress, or high solids loading, but
their stability mechanisms remain poorly understood. This is especially true when multiple stressors are
relevant at the same time. In this project, we mine metagenomes from extreme biomass-rich environments to
identify robust CAZyme candidates with predicted favourable stability across a range of parameters. We use
dbCAN3 for accurate CAZyme annotation, combined with domain and sub-family assignment, a structural feature
analysis, and in silico stability parameter predictions to spotlight enzymes with a likely tolerance to
multi-stress biocatalysis conditions. The use of these tools will focus on glycoside hydrolases (GHs) that
depolymerize complex polysaccharides. Candidate proteins are further evaluated for the presence of
carbohydrate-binding modules (CBMs), and other sequence or structural features associated with
thermostability, pH range, aggregation propensity, and secretion signals, used in a robustness scoring system.
Selected hits will then be cloned, expressed, and experimentally characterized for enzymatic activity and
stability parameters. This integrative pipeline aims to generate a curated library of multi-stable CAZymes
from metagenomic sources to improve the understandings of the molecular features of CAZyme robustness. The
resulting candidates may serve to validate and improve the accuracy of the stability predictors, as well as
representing robust new biocatalysts for industrial processes.
B-P.11: Computational analysis of protein-protein interactions in biocondensates
Track: Proteins and structural biology
-
Enrique Alanis Dominguez, CABD/CSIC, Spain
- Luis Ãngel RodrÃguez Lumbreras, ICVV-CSIC, Spain
- Ana Rojas, CABD/CSIC, Spain
- Juan Fernandez-Recio, ICVV/CSIC, Spain
Presentation Overview: Show
Cellular biocondensates are membraneless compartments that are increasingly recognized as key regulators of
cell physiology, especially under stress and during development. They allow for a rapid response and improve
the spatiotemporal control of biochemical reactions. However, their study has traditionally been limited by
the availability of experimental data in the form of high-throughtput experiments and only a few structural
models. In this work, we propose an innovative approximation to the study of biocondensates based on the
large-scale biophysical analysis of protein-protein interaction structural models generated with AlphaFold2.
Our objective is to identify differential features on protein interactions that typically appear in different
contexts: Inside biocondensates vs out of biocondensates, or in one biocondensante vs other biocondensates. We
are also interested in comparing interaction binding energetics between each pair of proteins to other
interactions for each of the proteins. By deepening the analysis of protein-protein interactions using
biophysical features, we close a knowledge gap that could not be properly assessed before the recent AI
revolution on protein complex modelling. Finally, this project pioneers a new protein-protein interaction
analysis strategy that can be exploided for the study of other molecular systems.
B-P.12: AbMuSiC: a physics-based method for predicting the change in binding affinity upon mutation in
antibody-antigen complexes
Track: Proteins and structural biology
-
Andre Ciupitu, Université Libre de Bruxelles, Belgium
- Gabriel Cia, Université Libre de Bruxelles, Belgium
- Marianne Rooman, Université Libre de Bruxelles, Belgium
- Fabrizio Pucci, Université Libre de Bruxelles, Belgium
Presentation Overview: Show
Antibodies have become indispensable in biomedical research and are rapidly becoming an important platform for
the development of next generation therapeutics. In this context, computational tools can be invaluable to
accelerate the rational optimization of initial antibody candidates and minimize experimental screening. Here
we introduce AbMuSiC, a structure-based method to predict the change in binding affinity upon mutation (DDG_b)
that is specifically designed for antibody-antigen interfaces. AbMuSiC is a physics-based model that linearly
combines coarse-grain statistical potentials derived from experimental protein structures. Additionally, the
model includes a clash term and an amino acid volume term, which are crucial for the accurate prediction of
mutations that fill interface cavities. AbMuSiC takes as input a 3D structure of the wildtype antibody-antigen
complex, and can predict the effect on the binding affinity of both single and multiple interface mutations.
When evaluated in strict cross-validation on all antibody-antigen mutations from the SKEMPIv2 experimental
dataset, AbMuSiC reaches a Pearson correlation of 0.55 and a standard deviation of 1.70 kcal/mol. On AbAgym,
our recently published antibody-antigen specific benchmark that contains 35k data points from 68 deep
mutational scanning experiments that capture the effect of interface mutations on antibody-antigen binding,
AbMuSiC is the best performing method. In conclusion, AbMuSiC is a state-of-the-art physics-based DDG_b
prediction method that can help in the rational optimization and design of antibody-antigen interfaces. It
will be made freely available for academic use as a Python package.
B-P.13: Identifying Candidate Precursors of Cyclic Repetitive Antibody Targets with the Peptidic Antigen
Sequence Complexity Analyzer (PASCA)
Track: Proteins and structural biology
-
Salvador Eugenio Caoili, University of the Philippines Manila, Philippines
Presentation Overview: Show
Peptide sequences are potentially useful antibody targets, notably among synthetic constructs that can replace
more complex, labile and/or hazardous antigens (e.g., pathogen virulence factors) in the manufacture of
biomedical products (e.g., vaccines, immunodiagnostics and prophylactic/therapeutic antibodies). However,
product development may fail due to structural differences between linear peptides and the intended final
antibody targets, particularly where the peptide amino- and/or carboxy-termini correspond to internal
protein-sequence positions in those targets. Yet, such structural mismatching can conceivably be avoided by
using repetitive peptide sequences (e.g., tandem repeats) in cyclic rather than linear form. Hence, the
present work provides the Peptidic Antigen Sequence Complexity Analyzer (PASCA) for identifying candidate
precursors of cyclic repetitive antibody targets in FASTA-formatted input data comprising linear B-cell
epitopes (LBCEs). PASCA characterizes peptidic (e.g., peptide and protein) antigens with regard to
repetitive-sequence content quantified as paratope-binding degeneracy z, itself estimated as a lower bound on
the relative perplexity of paratope-binding modes among circularly permuted sequence segments representing
LBCEs and defined by a sliding window, z thus being unity for homopolymers and the reciprocal of sequence
length for nondegenerate sequences. PASCA forms part of the PASCA User Toolkit (PUT), which also comprises the
PASCA-associated Utility for Matrix Assignment (PUMA) and the PASCA-associated Utility for Sequence Selection
(PUSS). PUMA provides for user-defined alternatives to the default residue-similarity matrix of Boltzmann
weights used by PASCA to evaluate z, whereas PUSS facilitates preparation of PASCA input sequence data (e.g.,
from the Immune Epitope Database [https://iedb.org/]). All of PUT is freely accessible online
(https://freeshell.de/badong/put.htm).
B-P.14: InterProScan 6: a modern large-scale protein function annotation pipeline
Track: Proteins and structural biology
-
Matthias Blum, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI),
United Kingdom
-
Emma Hobbs, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), United
Kingdom
-
Laise Cavalcanti Florentino, European Molecular Biology Laboratory, European Bioinformatics Institute
(EMBL-EBI), United Kingdom
-
Alex Bateman, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), United
Kingdom
Presentation Overview: Show
InterProScan integrates predictive models from the InterPro consortium to annotate protein sequences with
domains, families, and functional sites and is widely used in genome and metagenome annotation pipelines. It
underpins large-scale annotation efforts at resources such as UniProt, Ensembl, and MGnify and is routinely
applied in comparative genomics and functional annotation workflows. While InterProScan 5 has been extensively
adopted over more than a decade, its architecture increasingly limited scalability, portability, and
integration with modern workflow systems.
InterProScan 6 addresses these limitations through a complete reimplementation as a Nextflow-based workflow.
The pipeline supports execution across local systems, high-performance computing (HPC) clusters, and cloud
platforms with native container support via Docker, Singularity, and Apptainer. Pipeline logic is decoupled
from signature data, allowing multiple InterPro releases to coexist within shared installations and enabling
on-demand retrieval of required datasets. A redesigned Matches API provides access to precomputed InterPro
annotations for UniParc sequences, allowing transparent reuse of existing results during execution.
Benchmarking across nine reference proteomes, from bacteria to complex eukaryotes, shows consistent wall-clock
runtime reductions relative to InterProScan 5, with approximately two-fold speedups on large eukaryotic
proteomes. When annotation reuse via the Matches API is available, end-to-end runtimes fall to minutes.
Concordance analysis across the full Swiss-Prot dataset shows that InterProScan 6 reproduces InterProScan 5
results with near-identical precision and sensitivity across all InterPro member databases.
By combining workflow-native design, flexible data management, and reuse of precomputed annotations,
InterProScan 6 enables efficient, reproducible, and scalable protein function annotation for growing genomic
and metagenomic datasets.
B-P.15: Integrated Computational and Experimental Characterization of a Novel Metagenome-Derived Tryptophan
Indole-Lyase
Track: Proteins and structural biology
-
Hovsep Aganyants, The Scientific and Production Center "Armbiotechnology", National Academy of
Sciences of Armenia, Armenia
-
Artur Hambardzumyan, Scientific and Production Center “Armbiotechnologyâ€, National Academy of Sciences
of Armenia, Armenia
-
Tigran Soghomonyan, Scientific and Production Center “Armbiotechnologyâ€, National Academy of Sciences
of Armenia, Armenia
-
Marina Paronyan, Scientific and Production Center “Armbiotechnologyâ€, National Academy of Sciences of
Armenia, Armenia
-
Anichka Hovsepyan, Scientific and Production Center “Armbiotechnologyâ€, National Academy of Sciences of
Armenia, Armenia
- Vladimir Vukic, University of Novi Sad, Faculty of Technology, Serbia
-
Haykanush Koloyan, Scientific and Production Center “Armbiotechnologyâ€, National Academy of Sciences of
Armenia, Armenia
Presentation Overview: Show
Tryptophan indole-lyase (TIL) is a pyridoxal 5'-phosphate-dependent enzyme that catalyzes the reversible
β-elimination of L-tryptophan into indole, pyruvate, and ammonia. This enzyme has significant potential for
the synthesis of L-tryptophan and its derivatives, which are valuable in pharmaceutical, agricultural, and
fine chemical applications. Here, a metagenome derived from chicken manure and straw compost was used to mine
for a TIL. We combined molecular cloning and computational modeling to identify and characterize a novel TIL,
aiming to develop a technology for the production of L-tryptophan and its derivatives. The TIL gene was cloned
into the pET-24a(+) vector using Gibson assembly and expressed in E. coli BL21 Star cells. We demonstrated the
TIL capability for both tryptophan synthesis and degradation, with specific activities of 0.022 U/mg and 0.69
U/mg, respectively. Using AlphaFold 2, a 3D model of TIL was generated and evaluated. Further, L-tryptophan
and a panel of substituted analogs were analyzed by molecular docking. Among tested analogs 7-Aza-L-tryptophan
showed the highest affinity (-8.37 kcal/mol) towards the enzyme. This enzyme-ligand complex and ligand-free
enzyme were subjected to 100 ns MD simulation and compared. Structural stability, flexibility, compactness,
and solvent exposure were evaluated for both systems, revealing stability throughout the simulation.
Protein-ligand hydrogen bond formation showed stable interactions across the simulation. MM-PBSA analysis
demonstrated an energetically favorable interaction (-37.35 kcal/mol); additionally, individual residue
contributions were analyzed. Together, in vitro and in silico analyses confirm TIL functionality and catalytic
potential, supporting further biochemical characterization and engineering for sustainable biocatalysis.
B-P.16: In Silico Reverse Vaccinology Approach for Multi-Epitope Vaccine Design Against Stenotrophomonas
maltophilia
Track: Proteins and structural biology
-
Sumithra B, Chaitanya Bharathi Institute of Technology, Department of Biotechnology, Hyderabad, India,
India
-
Uma Laasya Mantena, Chaitanya Bharathi Institute of Technology, Department of Biotechnology,
Hyderabad, India, India
-
Nusaybah Mohammed Saleem, Chaitanya Bharathi Institute of Technology, Department of Biotechnology,
Hyderabad, India, India
Presentation Overview: Show
The global escalation of antimicrobial resistance poses a critical challenge to modern healthcare,
necessitating computational approaches to identify novel vaccine targets against multidrug-resistant
pathogens. Stenotrophomonas maltophilia, an opportunistic Gram-negative bacterium associated with
hospital-acquired infections, is known for its intrinsic resistance to multiple antibiotic classes and
currently lacks an approved vaccine. This study applied a reverse vaccinology-driven immunoinformatics
framework to design a potential multi-epitope vaccine candidate against S. maltophilia. The complete proteome
was systematically screened to identify surface-exposed and virulence-associated proteins, followed by
antigenicity assessment to prioritize suitable targets. Selected proteins were subjected to B-cell and T-cell
epitope prediction, and the predicted epitopes were further filtered based on antigenicity, non-allergenicity,
non-toxicity, and immunogenic potential. To improve broad applicability, conserved epitope analysis across
multiple strains was incorporated. The shortlisted epitopes were assembled into a multi-epitope construct
using appropriate linkers and an adjuvant to enhance immune response. Molecular docking and immune simulation
analyses were performed to evaluate the interaction of the vaccine construct with immune receptors and to
assess its immunogenic response. Population coverage analysis was performed to evaluate the global
applicability of the selected epitopes. The final construct exhibited favorable immunogenicity and safety
profiles based on immunoinformatics criteria. This study highlights the utility of integrating reverse
vaccinology with advanced immunoinformatics approaches to address antimicrobial resistance and provides a
foundation for subsequent experimental validation.
B-P.17: Benchmarking HERMES for TCR-pMHC Binding Prediction Across Diverse Systems Using AlphaFold3-Predicted
Structures
Track: Proteins and structural biology
-
Max Hoffmann, Institute of Medical Data Science, Faculty of Medicine, Otto-von-Guericke-University,
Magdeburg, Germany
-
Johannes Steffen, Department of Hematology, Oncology, Cell and Radiotherapy, Faculty of Medicine,
Otto-von-Guericke-University, Magdeburg, Germany
-
Matthias Leisegang, Department of Hematology, Oncology, Cell and Radiotherapy, Faculty of Medicine,
Otto-von-Guericke-University, Magdeburg, Germany
-
Julian Varghese, Institute of Medical Data Science, Faculty of Medicine, Otto-von-Guericke-University,
Magdeburg, Germany
-
Sarah Sandmann, Institute of Medical Data Science, Faculty of Medicine, Otto-von-Guericke-University,
Magdeburg, Germany
Presentation Overview: Show
Predicting T-cell receptor (TCR) binding to peptide major histocompatibility complexes (pMHCs) is highly
relevant to vaccine development and immunotherapy design. Compared to sequence-only approaches,
structure-based machine learning models, such as HERMES, appear especially promising. However, while HERMES
has shown competitive performance on a limited number of systems, its reliability across diverse systems
without experimental structures has not yet been assessed.
To fill this gap, we benchmarked HERMES (model: fixed) against BATCAVE, a literature-derived altered peptide
ligand dataset filtered to MHC-I. Eighty-four of 104 TCR-pMHC complexes (53 human, 31 mouse) were suitable for
analysis. For each complex an average of 162 mutated epitopes were mapped to measured TCR activation data. The
structures containing the index epitope were predicted with AlphaFold3. Finally, HERMES computes peptide
energy, a proxy for binding, for all mutated epitopes.
Performance varied substantially across TCR-pMHC complexes, with Spearman correlations between predicted
peptide energy scores and measured TCR activation ranging from -0.31 to +0.57. Mouse systems performed poorly,
with a median correlation of 0.
Although human systems performed better, median correlations were still low (cancer antigens: 0.09;
neoantigens: 0.16; viral antigens: 0.27). AlphaFold3 confidence only explained a small portion of the
variability in correlation, with correlation in high confidence structures ranging from 0 to 0.49.
Our results suggest that while HERMES can be useful when applied to high-quality experimental structures, its
utility on predicted structures is limited, highlighting the need for improved structure prediction tools
before broader application.
B-P.18: Rhea, a FAIR resource of expert curated biochemical and transport reactions
Track: Proteins and structural biology
-
Elisabeth Coudert, Swiss Institute of Bioinformatics, Switzerland
- Kristian B. Axelsen, SIB Swiss Institute of Bioinformatics, Switzerland
- Lucila Aimo, SIB Swiss Institute of Bioinformatics, Switzerland
- Nevila Hyka-Nouspikel, SIB Swiss Institute of Bioinformatics, Switzerland
- Parit Bansal, SIB Swiss Institute of Bioinformatics, Switzerland
- Teresa Batista Neto, SIB Swiss Institute of Bioinformatics, Switzerland
- Edouard de Castro, SIB Swiss Institute of Bioinformatics, Switzerland
- Nicole Redaschi, SIB Swiss Institute of Bioinformatics, Switzerland
- Alan Bridge, SIB Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Despite the central role of biochemical reactions in life, reaction knowledge remains fragmented across
databases, inconsistently described, and often not interoperable. This limits accurate enzyme annotation,
integration across resources, and downstream applications such as pathway reconstruction, enzyme function
prediction, and metabolic network analysis.
Rhea (www.rhea-db.org) addresses this challenge by providing an expert-curated, FAIR resource of biochemical
and transport reactions from the literature that is grounded in the ChEBI ontology and fully mapped to
UniProtKB. Rhea includes over 18,500 reactions covering primary and secondary metabolism of a broad range of
taxa, including reactions of the enzyme classification of the IUBMB and thousands more. Rhea has been adopted
as the reference vocabulary for the annotation of enzymes and transporters in UniProtKB, covering over 25
million protein sequences, and provides reaction chemistry for resources such as the Gene Ontology (GO),
Reactome, Model Organism Databases (MODs), SwissLipids, and the metabolic modeling platform MetaNetX. All Rhea
data is freely available to query and download from our website and APIs, including our SPARQL endpoint, and
is now widely used to develop methods for enzyme function prediction, protein design, and pathway
reconstruction, and the annotation of metabolic networks, metagenomes, and environmental biotransformations.
Here we present recent developments in Rhea, including improvements in reaction coverage, integration with the
GO, curation workflows that integrate AI to accelerate coverage of novel enzyme functions and "de-orphan"
enzymes, and new visualizations of reaction chemistry for experts and non-experts alike.
B-P.19: Structure-guided embeddings predict CYP substrate specificity: engineering citrus flavors for climate
resilience
Track: Proteins and structural biology
-
Eli Draizen, Barcelona Supercomputing Center, Spain
- Sara Tolosa-Alarcón, Barcelona Supercomputing Center, Spain
- Emre Cicekyurt, Barcelona Supercomputing Center, Spain
- Miguel Romero-Durana, Barcelona Supercomputing Center, Spain
- Alfonso Valencia, Barcelona Supercomputing Center, Spain
Presentation Overview: Show
Extreme climates may threaten crop-growing soil, limiting access to plant metabolites used for
pharmaceuticals, flavors, and fragrances. Microbial cell factories could be a viable replacement by scaling up
metabolite production, reducing land use and avoiding polluting extraction, but engineering them requires
understanding the plant enzymes that produce these metabolites. Many such compounds are synthesized through
Cytochrome P450 enzymes (CYPs), which catalyze substrate oxidation. We focus on a CYP subfamily that converts
valencene to nootkatone, a popular grapefruit-flavored compound. However, only nine true positive and eight
true negative valencene-catalyzing CYPs are known, limiting our understanding of which CYPs work and why. This
motivates the search for additional natural valencene-catalyzing CYPs.
First, we modeled all true positives with heme and valencene. The negative set and 180,000 plant CYP
candidates were modeled with heme only, superimposing valencene from their closest true positive homolog.
Next, seven sequence- and structure-based embedding methods were compared across full proteins and binding
site regions, using cosine similarity, embedding-based Needleman-Wunsch alignment, and optimal transport.
Finally, logistic regression predicted catalysis using six features: best, top-3 mean, and median similarity
to true positives and negatives for each embedding and similarity measure.
Frame2seq, a structure-guided masked language model, produced the most informative embeddings, achieving the
highest three-fold cross-validation AUC closest to leave-one-out performance despite limited training data. We
subjected 180,000 plant CYPs to the model, generating a ranked candidate list which we sent to collaborators
for experimental validation. This will expand our dataset and deepen our understanding of valencene catalysis.
B-P.20: Target-Specific De Novo Design of Drug Candidate Molecules with Graph-Transformer-Based Generative
Adversarial Networks
Track: Proteins and structural biology
- Atabey Ünlü, Hacettepe University, Turkey
- Elif Çevrim, Hacettepe University, Turkey
-
Melih Gökay YiÄŸit, Dept. of Computer Engineering, Middle East Technical University, 06800, Ankara,
Turkey, Turkey
- Ahmet Sarıgün, Middle East Technical University, Turkey
- Hayriye Çelikbilek, Hacettepe University, Turkey
- Osman Bayram, Bahcesehir University, Turkey
- Deniz Cansen Kahraman, Middle East Technical University, Germany
- Abdurrahman Olgac, Gazi University, Turkey
- Ahmet Süreyya RifaioÄŸlu, Heidelberg University, Turkey
- Erden Banoglu, Gazi University, Turkey
-
Tunca Dogan, Hacettepe University, Turkey
Presentation Overview: Show
Discovering novel small-molecule drug candidates that specifically interact with a protein target remains a
central challenge in rational drug design. While deep generative models have demonstrated capacity to sample
chemically valid molecules, most prior work focuses on property-optimized generation. We present DrugGEN, an
end-to-end generative framework for target-centric de novo molecular design, combining generative adversarial
networks (GANs) with graph representation learning. DrugGEN represents molecules as graphs encoding atom types
and bond connectivity. A key architectural contribution is the integration of graph transformer encoder blocks
into both generator and discriminator modules, incorporating a modified attention mechanism that explicitly
amplifies atomic bond information to capture both local and long-range interatomic dependencies. To our
knowledge, this is the first study to employ molecular graph transformers within a GAN pipeline. The framework
operates at realistic drug-sized molecular scales (~45 heavy atoms), a complexity that most prior generative
models do not address. Trained on ChEMBL-derived datasets and evaluated against 15 competing generative
models, DrugGEN ranked first across both targeted and non-targeted benchmarks. In AKT1-targeted molecular
docking, DrugGEN achieved 99.98% of real inhibitor binding performance, outperforming all comparators. Deep
learning-based bioactivity prediction identified 748 high-confidence AKT1-active de novo molecules. Five
synthesized candidates were experimentally validated via in vitro enzymatic assays, with two demonstrating
AKT1 inhibition at low-micromolar doses. Attention map analysis further revealed that the model correctly
identifies binding-critical atoms without explicit interaction supervision, providing mechanistic
interpretability. DrugGEN is openly available at github.com/HUBioDataLab/DrugGEN, enabling the community to
retrain the system for any druggable protein target.
B-P.21: Joint-Embedding Predictive Representation Learning for Protein-Ligand Pose Generation
Track: Proteins and structural biology
-
Atabey Ünlü, Hacettepe University, Turkey
- Tunca Doğan, Hacettepe University, Turkey
Presentation Overview: Show
Accurate prediction of protein-ligand binding poses remains challenging, as models must capture interaction
geometry while accommodating diverse, multimodal pose hypotheses. The Joint-Embedding Predictive Architecture
(JEPA) structures latent space to enable more tractable learning of such multimodal pose distributions. Here,
we propose a framework that integrates retrieval-style latent supervision with pose generation to achieve
accurate and controllable protein-ligand docking, presented here as ongoing work. The two-stage framework
combines joint-embedding predictive representation learning with SE(3)-equivariant explicit pose decoding. In
stage one, protein pocket features and coordinates are encoded alongside a ligand graph and a randomly
initialized query conformer. These representations are fused through bidirectional cross-graph interaction
layers to produce a protein-conditioned invariant latent representation. A target-side bound-pose encoder maps
the experimentally observed bound pose and local unbound-pocket context into a shared latent space, and the
model is trained with a bidirectional InfoNCE objective to align these representations without exposing bound
ligand geometry in the conditioning path. In stage two, the learned representation is frozen and used to
condition an anchor-based rigid-body and torsion pose decoder that iteratively predicts translational,
rotational, and torsional updates with local cross-attention and recycling-based refinement, with planned
extensions to multi-start pose generation and flow-based decoding. Initial experiments on PDBBind show stable
training, increasing cosine agreement between paired latent representations, and strong top-1 retrieval of the
matched target pose, indicating that the learned space captures pose-relevant interaction signals. Ongoing
work evaluates full-docking accuracy, pose diversity, and iterative-refinement quality relative to existing
generative docking baselines.
B-P.22: Co-folding based prediction of ligand interaction sites in cytochromes P450
Track: Proteins and structural biology
-
Karolina Plankova, Department of Physical chemistry, Faculty of Science, Palacký University
Olomouc, Czechia
-
Karel Berka, Department of Physical chemistry, Faculty of Science, Palacký University Olomouc, Czechia
-
Vaclav Bazgier, Department of Physical chemistry, Faculty of Science, Palacký University Olomouc, Czechia
Presentation Overview: Show
Accurate prediction of enzymatic reaction sites remains a central challenge in computational chemistry and
metabolic modelling. Cytochromes P450 (CYPs), particularly the CYP2C9 isoform, are critical enzymes involved
in the metabolism of approximately 20% of marketed drugs, xenobiotics, and endogenous compounds. Identifying
precise ligand reaction centres is essential for predicting metabolic pathways and supporting rational drug
design.
In this study, we evaluated the performance of recent co-folding frameworks, specifically Boltz and AlphaFold
3 in predicting ligand reaction sites. Co-folding allows for simultaneous prediction of protein and ligand
conformations, accounting for induced-fit effects. A dataset of 178 unique reactants corresponding to 310
metabolic reactions catalysed by CYP2C9 was analyzed. Predicted ligand poses and their respective reaction
sites were evaluated based on their spatial proximity to the heme cofactor and compared with experimental data
from the DrugBank database.
Both approaches indicate that co-folding methods can capture key aspects of enzyme ligand interactions. Boltz
models achieved a success rate of 51% in identifying correct reaction sites, slightly outperforming AlphaFold
3 (43%). Successful predictions showed a clear correlation with proximity to the heme iron, typically peaking
at distances of 4 to 5 Ã…. To further contextualize these findings, the performance of co-folding is now being
compared with traditional molecular docking. While no single structural features consistently determine the
prediction success, these results highlight co-folding as a promising approach for improving the accuracy of
metabolic site prediction and enhancing computational drug development tools.
B-P.23: ISS-Priority: Priority-guided graph propagation for indirect protein structure retrieval
Track: Proteins and structural biology
-
Hao Liu, University of Helsinki, Finland
- Liisa Holm, University of Helsinki, Finland
Presentation Overview: Show
Structural relationships between proteins help infer function and evolutionary history. Fast structural search
tools such as Foldseek enable large-scale retrieval but often miss remote structural homologs. We address this
gap in two stages.
First, we learn a compact retrieval embedding from ESM2 representations using DALI-derived structural
similarity as contrastive supervision. This injects structural signal during training while keeping inference
sequence-only, and yields a direct retrieval baseline of 0.5918 pooled AUPRC on a SCOPe fold-level benchmark
against the AlphaFold Database v2 (AFDB2), substantially above Foldseek (0.3828). Second, we introduce
ISS-Priority, a propagation-based retrieval method that searches the learned embedding space as a graph and
uses selective DALI validation to guide frontier expansion. Because runtime is controlled by the DALI
validation budget, users can trade speed for recall. By recovering indirect structural hits through validated
intermediate neighbors, ISS-Priority raises pooled AUPRC to 0.6128.
Together, these results show that combining DALI-supervised representation learning with propagation-based
retrieval extends structure-aware search beyond direct retrieval alone.
B-P.24: Mutation-Specific ∆∆G Modeling for Alanine Substitutions Improves Stability Prediction and Reveals
Distinct Constraint Regimes
Track: Proteins and structural biology
-
Maya Czeneszew, Sorbonne Université, CNRS, IBPS, Laboratoire de Biologie Computationnelle et
Quantitative, France
-
Alessandra Carbone, Sorbonne Université, CNRS, IBPS, LBCQ; Institut Universitaire de France (IUF), France
Presentation Overview: Show
Alanine scanning emerged as a widely used methodology to assess protein stability by systematically mutating
amino acid residues to alanine, whose small, neutral side chain isolates the loss of native side-chain
interactions. Despite its utility, to our knowledge, no dedicated deep learning model exists for predicting
alanine substitution ∆∆G values trained exclusively on single-residue substitutions to alanine.
We introduce a model that fine-tunes ESM-2 650M via low-rank adaptation (LoRA) with a combined prediction head
including a light-attention and a transformer-based embedding difference module. The model is trained on a
dataset of ~2,100 experimental ∆∆G measurements for alanine substitutions, augmented with ~2,100
FoldX-computed values, resulting in a curated alanine scanning dataset.
On the ProteinGym stability benchmark (66 assays), the model achieves a Spearman ρ of 0.73 on alanine
substitutions, outperforming ESCOTT (ρ = 0.53), a mutational effects predictor modeling evolutionary
conservation, structure, and epistasis, and VespaG (ρ = 0.46), the top single-sequence method on the
ProteinGym leaderboard. Comparison of ESCOTT and predicted ∆∆G at the single-mutation level across
representative proteins identifies regimes where stability and evolution decouple, highlighting highly
constrained residues with minimal predicted stability impact that coincide with previously described
spandrel-like sites characterized by persistent local frustration.
Alanine-specific ∆∆G modeling sharpens our ability to disentangle structural stability from evolutionary
constraint at single-residue resolution. By exposing regimes where these signals diverge, the model provides a
clearer lens on mutation effects and opens new avenues for precise variant interpretation and rational protein
design.
B-P.25: Decoding protein functional determinants: A simple adapter-based module with interpretable
attention
Track: Proteins and structural biology
-
Chujun Lyu, Sorbonne Université, CNRS, IBPS, Department of Computational, Quantitative and Synthetic
Biology (CQSB), Paris, France, France
-
Alessandra Carbone, Sorbonne Université, CNRS, IBPS, Department of Computational, Quantitative and
Synthetic Biology (CQSB), France
Presentation Overview: Show
Proteins orchestrate nearly all biological processes, yet predicting their functional determinants—key amino
acids responsible for a desired function—directly from sequence remains a major challenge. Protein language
models (PLMs) capture evolutionary and structural information from primary sequences, but leveraging their
embeddings for residue-level functional annotation demands specialized architectures. Here, we introduce a
novel approach that integrates LoRA adaptations into a PLM, using lightweight attention networks to
efficiently model residue-level dependencies, followed by a linear classifier for binary token-level
prediction. The PLM backbone is initialized from a model fine-tuned on soft disorder information. Our
framework incorporates two complementary label types: experimental interface annotations from the Protein Data
Bank and predicted post-translational modification labels from DeepMVP. Extensive validation on diverse
benchmarks demonstrates strong performance. On the CAID3 Binding-IDR challenge (52 proteins), competing with
tools designed for the task, our method achieves a probability prediction AUROC of 0.578, ranking 8th overall.
On VenusMutHub (796 proteins), it exhibits heightened sensitivity to selectivity and stability mutations. In a
threshold‑dependent manner, our method can outperform AlphaMissense on gain‑of‑function mutations (13
proteins, 52.2% vs. 42.4%) and match PRESCOTT on Finnish disease heritage mutations (22 missense, 55.5% vs.
54.5%). Over specific proteins, including two copies of TRX in Chlamydomonas, it can identify functional
differences at the level of the single protein copies in the same genome. These results highlight the efficacy
of lightweight attention mechanisms built upon PLM embeddings for capturing nuanced functional determinants
across structured and disordered proteomes, aiding functional tree reconstruction and protein annotation.
B-P.26: Target-Conditioned Molecule Generation using Protein and Chemical Language Models with
Bioactivity-Guided Post-Training
Track: Proteins and structural biology
-
Ahmet Furkan Öztürk, Middle East Technical University, Turkey
- Atabey Ünlü, Hacettepe University, Turkey
- Elif Çevrim, Hacettepe University, Turkey
- Tunca Doğan, Hacettepe University, Turkey
Presentation Overview: Show
Drug discovery increasingly leverages deep generative models to accelerate the design of bioactive molecules.
In this context, molecule design can be formulated as a sequence-to-sequence problem, where protein sequences
act as conditioning inputs and their ligands are generated in a candidate space. We present Prot2Mol, a
transformer-based autoregressive framework for de novo, target-specific molecule generation that integrates
protein and chemical language modeling. Prot2Mol employs SELFIES representations and incorporates multiple
protein language models: ProtT5, ESM2, and the structure-aware SaProt to capture sequence and structural
signals. The model follows an encoder–decoder architecture, where protein embeddings are co-trained and used
to condition a GPT-2–based molecular decoder via cross-attention. To further optimize generation, we extend
the framework with a pairwise reward model that scores protein–molecule compatibility. The reward model
encodes protein sequences using the same protein encoder architecture and molecules with a SELFIES-based
encoder, and is trained separately from the generative model while jointly learning protein–molecule
interactions to produce a ranking score and a binary activity prediction. Instead of direct regression of
bioactivity values, we adopt pairwise ranking with binary classification, mitigating biases from
dataset-specific protein priors and encouraging interaction-dependent signals. The reward function is
integrated into a GRPO-based reinforcement learning pipeline to guide molecule generation toward higher
bioactivity. Experiments on the Papyrus dataset show that Prot2Mol generates valid, diverse, and drug-like
molecules while maintaining target specificity. The framework establishes a unified pipeline combining
generative modeling with reward-driven optimization for target-informed molecular design. The tool is publicly
available at https://github.com/HUBioDataLab/Prot2Mol.
B-P.27: Apo, Holo, and Everything in Between: What State Are AlphaFold Binding Pockets In?
Track: Proteins and structural biology
-
Christos Feidakis, Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences,
Prague, Czech Republic, Czechia
-
Radoslav Krivak, Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences, Prague, Czech
Republic, Czechia
-
Derrick Agwora, Department of Cell Biology, Faculty of Science, Charles University, Prague, Czech
Republic, Czechia
-
Chloé Weiler, Faculty of Engineering Sciences, Heidelberg University, Heidelberg, Germany, Germany
-
David Hoksza, Department of Software Engineering, Faculty of Mathematics and Physics, Charles University,
Prague, Czech Republic, Czechia
-
Jiri Vondrasek, Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences, Prague, Czech
Republic, Czechia
-
Marian Novotny, Department of Cell Biology, Faculty of Science, Charles University, Prague, Czech
Republic, Czechia
Presentation Overview: Show
AlphaFold DB models are widely used for ligand-binding analysis, but a practical question remains: do their
binding pockets look more like ligand-free apo states, ligand-bound holo states, or something in between? We
addressed this using AHoJ-DB, which links experimentally observed alternative pocket conformations for the
same protein regions. From 462,413 query-pocket entries with AlphaFold classifications, after filtering for
usable classifications and at least one apo comparator, we retained 222,055 pockets from 11,629 UniProt
accessions. At first glance, AlphaFold pockets appeared more holo-like than apo-like, but this imbalance moved
toward equilibrium once overlapping pockets were merged into unified binding sites, showing that redundancy
matters. More importantly, experimental apo/holo evidence balance did not produce a hard switch between
classes; instead, AlphaFold pockets shifted gradually along an apo-holo continuum with the holo-to-apo
experimental evidence ratio. We then tested whether this local pocket state mattered in two downstream
analyses. In same-UniProt AlphaFill matches, ligand transplants fit similarly in low-separation pockets, but
fit worsened for apo-like AlphaFold pockets when the experimental apo and holo states were more distinct.
P2Rank showed higher binding-site detection success for holo-like pockets and for pockets supported by higher
holo evidence ratios. These results suggest that AlphaFold pockets are best interpreted as occupying a
measurable structural spectrum, and that this local pocket state can help flag when ligand placement and
pocket prediction are likely to be straightforward or require extra caution.
B-P.28: Structure-Function Analysis of Arabidopsis Cellulose Synthases and Small Molecule Inhibitors Using
Computational and In Vivo Approaches
Track: Proteins and structural biology
-
Carlo Perolo, University of Toronto, Canada
- Nicholas Provart, University of Toronto, Canada
- Heather McFarlane, University of Toronto, Canada
Presentation Overview: Show
The main structural component of plant cell wall is cellulose, an unbranched polymer of β-1,4-glucose
monomers synthesised by enzymes called Cellulose Synthases (CESA). In Arabidopsis thaliana, primary cell wall
CESAs are hypothesized to work as hexamers of heterotrimers moving along microtubule tracks to deposit
cellulose in a controlled way. Some of the most valuable tools to study plant CESAs are Cellulose Biosynthesis
Inhibitors (CBIs), small molecules able to directly affect their function and localization in a dose-dependent
manner. Despite significant similarity in structure and sequence of the three A. thaliana primary cell wall
CESAs (CESA1, CESA3, and CESA6) resistance-conferring mutation screens reveal a surprising specificity of
several CBIs for one monomer over the others. This provides us with an ideal framework to study their
mechanism of inhibition, especially with the recent advances in protein structure prediction tools like
AlphaFold. We performed Molecular Docking simulations on multiple CESA resolved and predicted structures and
identified several CBIs' potential binding pocket. To validate our prediction, an extensive cross-resistance
screen was conducted, and finally site-directed mutagenesis was used to introduce mutations targeting residues
within the proposed binding pocket on CESA1. The resulting mutant plants show specific resistance to CBIs,
supporting our predictions. Overall, our work furthers the understanding of CBI mechanism of inhibition by
proposing a binding pocket and highlights the predictive power of computational structural approaches in
understanding protein-ligand interactions. Lastly, we are currently implementing our pipeline as a publicly
available tool on the Bio-Analytic Resource for Plant Biology (BAR) website.
B-P.29: Integrated computational and experimental analysis of a matricellular protein-peptide interaction
enables rational antifibrotic binder design
Track: Proteins and structural biology
-
Jennifer Faúndez-Contreras, Facultad de Medicina, Universidad San Sebastián, Santiago, Chile.
Fundación Ciencia & Vida, Santiago, Chile., Chile
-
Enrique Brandan, Facultad de Medicina, Universidad San Sebastián, Santiago, Chile. Fundación Ciencia
& Vida, Santiago, Chile., Chile
-
Sebastián Bazaes, Centro CientÃfico y Tecnológico de Excelencia Ciencia & Vida, Santiago, Chile.,
Chile
-
Raul Araya-Secchi, Facultad de Ingenieria. Universidad San Sebastián, Valdivia, Chile; Centro de Estudios
Cientificos CECs, Valdivia., Chile
-
Felipe GarcÃa-Olave, Programa de doctorado en Biologia Computacional. Universidad San Sebastian,
Santiago, Chile, Chile
-
Tiaren Ruiz, Facultad de Ingenieria, Universidad San Sebastián, Santiago, Chile., Chile
Presentation Overview: Show
Introduction. Fibrosis is characterized by excessive extracellular matrix (ECM)
deposition leading to organ dysfunction. Connective Tissue Growth Factor (CTGF)
is a key mediator of fibrosis, promoting ECM production and fibroblast activation.
Its inhibition can attenuate fibrotic progression. Structurally, CTGF comprises four
conserved domains organized into N-terminal (domains 1-2) and C-terminal
(domains 3-4) regions, connected by a protease-sensitive hinge. The C-terminal
region has been associated with profibrotic activity. We previously identified a
peptide capable of inhibiting CTGF, however, its mechanism remains unclear. This
study aimed to define the CTGF-peptide interaction to guide the rational design of
improved CTGF-targeting binders.
Methods. Structural models of CTGF alone and in complex with the peptide were
generated using AlphaFold2 and AlphaFold-Multimer, followed by 1.2 μs molecular
dynamics simulations. Binding contributions were assessed by alanine scanning
using Rosetta Flex ddG. Experimental validation included ELISA alanine scanning
and co-immunoprecipitation assays with full-length and truncated CTGF. Functional
effects were assessed in fibroblasts via fibronectin deposition and stress fibers
formation. Statistical analysis was performed using two-way ANOVA.
Results. Computational models suggested dual peptide interaction with both
CTGF regions, promoting conformational compaction and reduced proteolytic
accessibility. Experimental data revealed predominant binding to the C-terminal
region and identified key peptide residues critical for binding and functional activity.
Conclusion. These findings define a CTGF-peptide interaction interface and
support the rational design of targeted antifibrotic binders.
Acknowledgements. This work is supported by FONDECYT-N°1230054,
Proyecto CCTE Ciencia y Vida Basal-FB210008, FONDEF-ID25I10016, and
ANID/BECA DOCTORADO NACIONAL-21241714.
B-P.30: Coordinated pMHC Dynamics Reveal TCR-Recognizable States for Neoepitope Prioritization in
Personalized Cancer Vaccine Design
Track: Proteins and structural biology
-
Jelang Muhammmad Dirgantara, Laboratory of In Silico Design, National Institutes of Biomedical
Innovation, Health and Nutrition, Japan, Japan
-
Kazuma Kiyotani, Laboratory of Immunogenomics, National Institute of Biomedical Innovation, Health and
Nutrition, Japan, Japan
-
Takuto Nogimori, Laboratory of Precision Immunology, National Institutes of Biomedical Innovation, Health
and Nutrition Japan, Japan
-
Shokichi Takahama, Laboratory of Precision Immunology, National Institutes of Biomedical Innovation,
Health and Nutrition Japan, Japan
- Takashi Morisaki, Fukuoka General Cancer Clinic, Japan, Japan
-
Takuya Yamamoto, Laboratory of Precision Immunology, National Institutes of Biomedical Innovation, Health
and Nutrition Japan, Japan
-
Suyong Re, Laboratory of In Silico Design, National Institutes of Biomedical Innovation, Health and
Nutrition, Japan, Japan
Presentation Overview: Show
Accurate identification of immunogenic peptide-MHC (pMHC) complexes remains a central challenge for
immunotherapy, as binding affinity alone poorly predicts T cell activation. To elucidate structural
determinants of T cell receptor (TCR) recognition, we analyzed experimentally resolved pMHC and TCR-pMHC
crystal structures across multiple HLA class I alleles in Protein Data Bank. We show that peptide backbone
bulge topology at TCR-facing residues is largely pre-organized prior to TCR engagement, supporting a
conformational selection mechanism. Notably, this indicates that key structural features relevant to TCR
recognition are encoded within the pMHC state, enabling inference of TCR-interacting properties without
explicit TCR information. This is particularly important because, in personalized cancer vaccine design, TCR
information is typically unavailable (neoepitope identification and HLA typing are performed without knowledge
of cognate TCRs). Guided by these insights, we performed 100-ns all-atom molecular dynamics (MD) simulations
on 945 pMHC complexes derived from the largest Japanese cancer patient neoepitope cohort with ELISpot
immunogenicity labels, totaling 94.5 microseconds of simulation. We quantified anchor residue stability within
MHC pockets and flexibility of TCR-facing residues. While individual dynamic features partially enrich
immunogenic epitopes, they fail to provide a consistent or biophysically interpretable framework across HLA
alleles. To address this, we developed a Consensus Scoring Function (CSF) integrating anchor stability and
bulging dynamics into a unified measure of pMHC fitness for TCR recognition. CSF improves immunogenic hit
rates by up to 8% while reducing experimental burden by half, providing a scalable framework for neoepitope
prioritization in personalized cancer vaccine design.
B-P.31: ZFP-CanPred: Predicting the Effect of Mutations in Zinc-Finger Proteins in Cancers Using Protein
Language Models
Track: Proteins and structural biology
-
Amit Phogat, Department of Biotechnology, Indian Institute of Technology Madras, Chennai, India,
India
-
Sowmya Ramaswamy Krishnan, Department of Biotechnology, Indian Institute of Technology Madras, Chennai,
India, India
-
Medha Pandey, Department of Biotechnology, Indian Institute of Technology Madras, Chennai, India, India
-
M. Michael Gromiha, Department of Biotechnology, Indian Institute of Technology Madras, Chennai, India,
India
Presentation Overview: Show
Zinc-finger proteins (ZNFs) are the largest group of transcription factors and are involved in many important
cellular functions. Missense mutations in ZNFs can disrupt protein-DNA interactions and may contribute to the
development of different cancers. In this study, we introduce ZFP-CanPred, a deep learning-based model
designed to predict cancer-associated driver mutations in ZNFs. The model uses representations obtained from
protein language models (PLMs), focusing on the structural neighbourhood around mutation sites, to distinguish
between cancer-causing and neutral mutations. ZFP-CanPred achieved strong performance on an independent test
set, with an accuracy of 0.72, F1-score of 0.79, and area under the ROC curve (AUC) of 0.74.In comparison with
11 existing prediction tools on a curated dataset of 331 mutations, ZFP-CanPred showed the highest AUROC of
0.74, outperforming both general and cancer-specific methods. Its balanced sensitivity and specificity help
address a key limitation seen in current approaches. The source code and related materials are available
at:Â https://github.com/amitphogat/ZFP-CanPred.git. This work may support a better understanding of cancer
mechanisms and assist in the development of targeted therapies.
B-P.32: Functional aggregation-prone regions mediate carbohydrate-binding and harbour disease-associated
mutations
Track: Proteins and structural biology
-
Lekshmi S, Indian Institute of Technology Madras, Chennai, India, India
-
Siva Shanmugam N R, Department of Food Science and Technology, University of Nebraska - Lincoln, Lincoln,
USA, India
- Prabakaran R, Department of Biology, Emory University, Atlanta, GA, USA, India
- Rawat P, Department of Immunology, University of Oslo, Oslo, Norway, India
- Michael Gromiha M, Indian Institute of Technology Madras, Chennai, India, India
Presentation Overview: Show
Carbohydrate-protein interactions play key roles in cellular processes such as cell-cell communication,
differentiation, and immune response, and their disruption by mutations is implicated in several diseases.
Emerging evidence suggests that aggregation-prone regions (APRs), which are associated with disease-related
protein aggregation, occur near functional sites. However, the role of APRs in carbohydrate recognition is
largely unexplored. We address this gap using a systematic sequence- and structure-based analysis of APRs in
protein-carbohydrate interfaces using a curated dataset of 1119 carbohydrate-binding protein chains. We
identified a distinct subset of residues in APRs that bind to carbohydrates, termed functional
aggregation-prone regions (fAPRs). Our results show that ~40% of carbohydrate-binding residues overlap with
APRs, indicating that aggregation propensity is tolerated within functional sites. The non-overlapping
carbohydrate-binding residues are in spatial proximity to the APRs, contributing to intermolecular
interactions. Notably, ~85% of fAPRs harbour at least one likely pathogenic missense mutation, highlighting
their utility in the interpretation of variant pathogenicity. The fAPRs are enriched with aromatic residues,
consistent with known carbohydrate-binding mechanisms. Despite the functional role, aggregation propensity and
carbohydrate-binding affinity appear to be uncoupled, suggesting that aggregation and binding can be
independently modulated. Overall, this study reveals fAPRs as important structural and functional elements at
the protein-carbohydrate interface, offering new insights into carbohydrate-recognition and mutant effect
prediction.
B-P.33: Construction and Evaluation of a Multimodal Fusion-Based Model for Compound Protein Interaction
Prediction
Track: Proteins and structural biology
-
Si Zheng, Chinese Academy of Medical Sciences and Peking Union Medical College, China
- Jiao Li, Chinese Academy of Medical Sciences and Peking Union Medical College, China
Presentation Overview: Show
Compound protein interaction (CPI) prediction is an important task in lead discovery and drug repositioning.
Although in vitro assays are reliable, their high cost and low throughput limit their application in
large-scale screening. Recent deep learning methods have provided efficient computational alternatives for CPI
prediction; however, their generalizability in realistic virtual screening settings remains limited,
particularly under cold-start scenarios involving unseen compounds or proteins.
Here, we present MoCL-CPI, a multimodal framework for CPI prediction. MoCL-CPI integrates protein sequences,
molecular structures, physicochemical descriptors, multiple sequence alignment features, and embeddings
derived from a large-scale heterogeneous biomedical network to represent compounds and proteins from
complementary structural, functional, topological, and evolutionary perspectives. In particular, the
biomedical network-based modality captures global associations and functional context that are not directly
available from local structural representations, thereby improving model generalization in challenging
screening settings.
We evaluated MoCL-CPI on three public benchmark datasets, including C. elegans, BindingDB, and BioSNAP.
Experimental results show that MoCL-CPI achieves improved robustness and predictive performance across diverse
screening scenarios. Notably, under inductive data splits that more closely resemble cold-start settings, the
proposed framework shows more consistent advantages, suggesting that multimodal collaborative learning
provides more stable support for interaction prediction involving unseen entities. These results suggest that
multimodal representation learning can be an effective strategy for mitigating the cold-start problem in CPI
prediction and may provide a useful framework for computational compound screening.
B-P.34: Order–disorder interfaces in viral pRb inactivation: molecular dynamics insights into SLiM-mediated
recognition and implications for interaction databases
Track: Proteins and structural biology
-
Carla Padilla Franzotti, Department of Science and Technology, National University of Quilmes,
CONICET, Argentina
-
Nicolas Palopoli, Department of Science and Technology, National University of Quilmes, CONICET, Argentina
-
Gustavo Pierdominici-Sottile, Department of Science and Technology, National University of Quilmes,
CONICET, Argentina
-
Henning Hermjakob, European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI),
United Kingdom
Presentation Overview: Show
Intrinsically disordered regions (IDRs) mediate protein interactions through poorly understood mechanisms. We
studied the retinoblastoma protein (pRb) and its interaction with the SV40 Large T antigen (LTSV40), a viral
oncoprotein that displaces E2F factors.
Using molecular dynamics and umbrella sampling, we show the LTSV40 LXCXE motif is part of a conserved
Order–Motif–IDR architecture. The ordered N-terminal region drives initial pRb recognition via induced
folding, adding over 6 kcal/mol to affinity. Simultaneously, the C-terminal IDR undergoes a bent-to-extended
transition, sterically occluding the pRb AB cleft to prevent E2F binding. These coupled phenomena are
evolutionarily conserved across 14 polyomaviruses, as confirmed by AlphaFold and MobiDB, suggesting a common
pRb inactivation strategy.
These results highlight a broader challenge for the field: functionally decisive IDR behaviors such as
binding-induced folding and steric occlusion are not yet representable within current interaction data models,
even when complementary evidence exists in resources such as DisProt or MobiDB. Bridging this gap through
systematic integration of IDR annotations into databases such as IntAct and Complex Portal, supported by
projection onto AlphaFold3-predicted complex structures, would enable community-scale identification of
complexes where disorder is mechanistically decisive.
B-P.35: Statistical modelling of TCR sequences provides robust quality control for repertoire
datasets
Track: Proteins and structural biology
-
Dana Moreno, Department of Fundamental Oncology, Ludwig Institute for Cancer Research, University of
Lausanne, Lausanne, Switzerland, Switzerland
-
Giancarlo Croce, Department of Fundamental Oncology, Ludwig Institute for Cancer Research, University of
Lausanne, Lausanne, Switzerland, Switzerland
-
David Gfeller, Department of Fundamental Oncology, Ludwig Institute for Cancer Research, University of
Lausanne, Lausanne, Switzerland, Switzerland
Presentation Overview: Show
T-Cell Receptors (TCRs) show extensive sequence diversity across T cells. This diversity arises from different
choices of V and J genes and from insertions and deletions at the V(D)J junction within the
Complementary-Determining Region 3 (CDR3) loop. Here, we quantify how V and J gene usage shapes CDR3 length
and amino acid composition. In repertoires of TCRs with either unknown or known specificity, we show that CDR3
length is strongly influenced by the number of germline-encoded CDR3 residues in V and J genes, and that, on
average, 80% of CDR3α and 65% of CDR3β residues are determined by V and J gene usage. We further show that
inconsistencies between V and J gene annotations and CDR3 sequences can be leveraged to identify potential
issues in multiple TCR repertoire datasets. Overall, our study quantifies the impact of V and J gene usage on
CDR3 length and amino acid composition and provides a robust framework for quality control of TCR repertoire
sequencing data.
B-P.36: Uncovering Hidden Proteomic Landscapes: Isoform-Resolved Drug Response Analysis with NOMAD
Track: Proteins and structural biology
-
Jesse Angelis, Computational Mass Spectrometry, Technical University of Munich, Freising, Germany,
Germany
-
Mathias Wilhelm, Computational Mass Spectrometry, TUM, Freising, Germany; Munich Data Science Institute
(MDSI), TUM, Garching, Germany, Germany
Presentation Overview: Show
Modern dose resolved proteomics platforms provide unprecedented maps of drug induced expression changes. These
experiments measure protein abundance across drug concentrations to record cellular perturbations, identifying
drug targets and modes of action. While invaluable, these datasets are restricted to protein groups (clusters
of proteins sharing identified peptides) - an inherent shortcoming of bottom-up proteomics. By aggregating
molecular signals, these approaches overlook the functional diversity of individual isoforms, potentially
obscuring critical regulatory events.
We hypothesised that drug induced perturbations often target specific isoforms, leading to ""signal washing""
at the protein group level. To resolve this, we present NOMAD, an analytical framework that deconvolves
isoform specific trajectories from dose response data. Utilizing Non-negative Matrix Factorization (NMF),
NOMAD separates the relative contributions of individual isoforms to reveal latent regulatory profiles that
remain invisible to standard aggregation methods.
Applying NOMAD to proteomic measurements of dose-resolved drug treatments of three drugs in Jurkat cells, we
identified widespread isoform specific effects. In an analysis of 3,939 shared protein groups, NOMAD suggests
substantial regulations missed by global analysis. Notably, 81.7% of NOMAD exclusive findings were attributed
to ""Internal Friction,"" where divergent isoform trends neutralize the global signal. Furthermore, 16.2% of
shared significant proteins exhibited ""Isoform Flipping,"" where a sub variant responds in opposition to the
canonical trend. While these initial results require further experimental validation, they suggest that
isoform resolution is necessary to understand the true molecular complexity of drug cell interactions
otherwise missed by protein group analysis.
B-P.37: Whole-Exome Sequencing and Structural Modeling Identify Candidate Drivers in a Multiplex Multiple
Sclerosis Family
Track: Proteins and structural biology
-
Simone Bonora, Department of Chemistry and Biology A. Zambelli, University of Salerno, Fisciano (SA),
Italy, Italy
-
Anna Marabotti, Department of Chemistry and Biology A. Zambelli, University of Salerno, Fisciano (SA),
Italy, Italy
-
Carla Lintas, Research Unit of Medical Genetics, Department of Medicine, Campus Bio-Medico University of
Rome, Rome, Italy, Italy
-
Claudio Tabolacci, National Center for Rare Diseases, Istituto Superiore di Sanità, Italy
-
Maria Luisa Scattoni, National Center for Rare Diseases, Istituto Superiore di Sanità, Italy
-
Fioravante Capone, Operative Research Unit of Neurology, Fondazione Policlinico Universitario Campus
Bio-Medico, Rome, Italy, Italy
-
Mariagrazia Rossi, Operative Research Unit of Neurology, Fondazione Policlinico Universitario Campus
Bio-Medico, Rome, Italy, Italy
-
Vincenzo Di Lazzaro, Operative Research Unit of Neurology, Fondazione Policlinico Universitario Campus
Bio-Medico, Rome, Italy, Italy
-
Fiorella Gurrieri, Research Unit of Medical Genetics, Department of Medicine, Campus Bio-Medico University
of Rome, Rome, Italy, Italy
Presentation Overview: Show
Multiple sclerosis (MS) is a chronic inflammatory and neurodegenerative disorder of the central nervous system
whose heritability is only partly explained by common-risk loci. To investigate the contribution of rare
coding variants in familial disease, we applied whole-exome sequencing (WES) to a multigenerational Italian
multiplex family and integrated segregation analysis with structural modeling and comparative structural
analysis.
WES identified 47 rare co-segregating variants, which were refined to three top candidates after evaluation in
unaffected relatives: RTN4 p.Pro148Leu, JAK2 p.Phe560Val, and DUOX2 p.Tyr1150Cys. For JAK2, we selected the
PDB structure 6BS0, whereas for DUOX2 we used the AlphaFold model AF-Q9NRD8-F1, since no experimental
structure is available and the region containing p.Tyr1150Cys is predicted with high local confidence. Mutant
models were generated using the Mutate Model procedure derived from MODELLER. Comparative structural analysis
was performed with RING 4.0, DSSP, NACCESS, DynaMut2, DUET, and INPS-MD.
In JAK2, p.Phe560Val disrupts a π–π stacking interaction with Phe547, alters local solvent accessibility and
van der Waals contacts, and may perturb a nearby ligand-binding pocket, overall supporting a destabilizing
effect. In DUOX2, p.Tyr1150Cys abolishes π–π interactions with Phe591/Phe598 and hydrogen bonds with
Ser1153/Val1154 within the ferric reductase-like transmembrane region, again supporting local structural
destabilization. By contrast, RTN4 could not be interpreted structurally with confidence because the variant
maps to a predominantly disordered region. Overall, this integrative WES-to-structure workflow supports an
oligogenic model of MS susceptibility and prioritizes JAK2 and DUOX2 for functional follow-up.
B-P.38: autoeval: A robust toolkit for standardized ad-hoc protein language model evaluation
Track: Proteins and structural biology
-
Sebastian Franz, Technical University of Munich, Germany
- Aleena Siji, Helmholtz AI, Munich, Germany
- Lisa M. Spindler, Technical University of Munich, Germany
- Michael Heinzinger, Institute of Computational Biology, Helmholtz Munich, Germany
- Burkhard Rost, Technical University of Munich (TUM), Germany
Presentation Overview: Show
Protein Language Models (pLMs) became an essential cornerstone when working with proteins over the last years.
For the development of new pLMs, however, it is essential to put the performance of the new model into
perspective of existing ones. To complement existing evaluation frameworks which tend to focus on variant
effect prediction (VEP), we are presenting autoeval, a robust toolkit for standardized ad-hoc pLM evaluation.
Autoeval checks the representation quality of pLM by performing linear probing on several curated, established
supervised prediction tasks. We provide validation splits to allow direct integration of autoeval into model
development workflows without risking information leakage. Once hyperparameters are set and pLM pre-training
is finished, autoeval can be switched into evaluation mode to assess the final model quality on established
test sets, putting the new model's performance directly into perspective of existing ones. For simplifying
comparison to existing work, we also support automated zero-shot evaluation on de-facto standard VEP
benchmarking with ProteinGym as well as contact map prediction. To simplify usability, autoeval can be used
via a simple Python API that directly integrates with a comprehensive web platform to analyze the results. The
toolkit is available publicly at https://autoeval.biocentral.cloud.
B-P.39: Interpreting Structural Representations and Attention in Pairformer-Based Protein Folding Models for
Protein–Protein Interactions
Track: Proteins and structural biology
-
Wout Keymis, VIB center for AI and Computational Biology, KU Leuven, Belgium
- Wannes Hermans, VIB center for AI and Computational Biology, KU Leuven, Belgium
- Joana Pereira, VIB center for AI and Computational Biology, KU Leuven, Belgium
Presentation Overview: Show
Understanding how modern protein structure prediction models encode structural information remains a challenge
in computational biology. This study investigates how structural features emerge across layers in a
Pairformer-based folding model (Boltz1) applied to protein–protein complexes.
Intermediate single-sequence and pair embeddings are extracted from each layer and analyzed in relation to the
model's attention mechanisms. To disentangle these embeddings, layer-wise sparse autoencoders (SAEs) with
top-k sparsity are trained, yielding interpretable and selectively activated latent features.
Individual latent dimensions are then probed for structural information using linear classifiers, evaluating
feature separability (ROC-AUC) and activation-conditioned presence (F1). This enables layer-by-layer tracking
of how structural signals are encoded and refined.
In parallel, attention patterns are analyzed, focusing on triangle self-attention and single-sequence
attention with pair bias. Attention patterns are evaluated using two complementary analyses: statistical
enrichment of structural features and SAE activations in high-attention regions, and co-occurrence between
attention and these signals.
Sparse latent factors are found to align with meaningful structural properties, while specific attention heads
exhibit strong specialization and evolve across layers. For example, early layers often emphasize residue
pairs that are both sequentially and spatially proximal, whereas later layers tend to shift toward
predominantly spatial proximity with reduced dependence on sequence adjacency. Similarly, early attention
shows broad sensitivity to cysteine residues, while deeper layers refine this signal to a more targeted focus
on disulfide bridges. In addition, single-sequence attention heads can highlight binding interface
hotspots.
These results provide an initial view of how Pairformer-based models represent protein structure.
B-P.40: Structure based scoring of AI-predicted peptide-GPCR complexes for receptor annotation in a non-model
insect genome
Track: Proteins and structural biology
-
Febrina Margaretha, Genetics Program, Graduate Institute for Advanced Studies, SOKENDAI,
Japan
- Mika Sakamoto, Genome Informatics Laboratory, National Institute of Genetics, Japan
- Hitomi Seike, Graduate School of Frontier Sciences, University of Tokyo, Japan
- Shinji Nagata, Graduate School of Frontier Sciences, University of Tokyo, Japan
-
Yasukazu Nakamura, Genome Informatics Laboratory, National Institute of Genetics, Japan
-
Takako Mochizuki, Genome Informatics Laboratory, National Institute of Genetics, Japan
Presentation Overview: Show
Genome annotation is essential for linking DNA sequences to biological functions. Neuropeptides are
neuromodulatory molecules regulating diverse physiological processes and behaviors, including development,
feeding, and courtship, thereby making accurate annotation of the genes encoding them and their receptors
critical. Neuropeptides exert their effects by binding to specific receptors, most often G protein-coupled
receptors (GPCRs). In non-model insects like Gryllus bimaculatus, annotation of neuropeptides and their
receptors remains limited. Our previous work has improved neuropeptide gene functional annotation in this
species through de novo genome assembly and homology-based searches against insect databases. However,
receptor annotation remains incomplete as sequence‑based homology approaches struggle to assign precise
functional annotations despite high sequence similarity among GPCRs and limited availability of reference
sequences.
In this study, we explore a structure‑based strategy to assign candidate neuropeptide receptors in the G.
bimaculatus by assessing predicted peptide–receptor complexes using AlphaFold3 and Boltz‑2. Focusing on
the major classes of insect neuropeptide receptors, rhodopsin‑like and secretin‑like GPCR families, we
systematically assign annotated neuropeptides with candidate receptor structures and predict their complexes.
To evaluate these combination ligand and receptor models, we apply interface‑focused confidence metrics:
interface pLDDT, which measures the local reliability of residues at the binding interface, and ipSAE that
summarizes conformational plausibility of the interface geometry and alignment features, to score the
structural validity of each peptide–receptor interaction. By ranking structure‑aware confidence measures
for G. bimaculatus neuropeptide putative receptors, we construct a list of prioritized candidates that can
guide subsequent experimental validation and improve genome annotation in non-model organisms.
B-P.41: AMP-DiT: Antimicrobial Peptide Design with AMP-classifier Conditional Diffusion Transformers
Track: Proteins and structural biology
-
Alireza Noroozi, Koc university, Turkey
-
Ozlem Keskin, Professor of Chemical and Biological Engineering, Koc Univeristy, Turkey
- Attila Gürsoy, Professor of Computer Science, Koc University, Turkey
Presentation Overview: Show
Antimicrobial resistance is a global health threat. Antimicrobial peptides (AMP) can help for their potent and
promising ability to fight resistant pathogens. While AI is being employed to advance AMP discovery, deep
learning methods remain limited by poor controllability, suboptimal sequence representations, and low
experimental hit rates. We introduce AMP-DiT, a conditional discrete diffusion framework that generates AMPs
directly in sequence space, avoiding reliance on continuous latent embeddings and large protein language model
priors.
Our approach leverages a denoising diffusion transformer architecture with classifier guidance, enabling high
fitness antimicrobial peptide properties during generation. Crucially, we condition the generative process on
biologically meaningful signals derived from external AMP predictor, Macrel, explicitly biasing sampling
toward sequences with high predicted antimicrobial activity. Since AMP activity is the primary determinant of
downstream success, this conditioning serves as a strong inductive bias for generating functional peptides.
We observe that conditioning on a single AMP predictor not only improves scores on that model but also
generalizes across other independent AMP prediction frameworks, suggesting that the model captures underlying
antimicrobial features rather than overfitting to a specific predictor.
Unlike prior approaches that depend on protein language models trained on proteins, AMP-DiT is trained on
peptide-specific data with guidance signals, reducing representation mismatch and avoiding biases inherited
from global protein distributions. Overall, AMP-DiT establishes a framework for AMP generation that
outperforms existing AMP design methods across key evaluation metrics. As we got 10 percent lower MIC in an
AMP prediction model, meanwhile keeping diversity.
B-P.42: AMP-BindGen: Guided Protein Design for Antimicrobial Peptides and Targeted Inhibitors Against
AMR
Track: Proteins and structural biology
-
Alireza Noroozi, Graduate Sience and Engineering school of Koc university, Turkey
- Ani Akpinar, KUISCID, Molecular Biology and Genetics, Koc university, Turkey
- Erkmen Erken, Computer Science, Koc university, Turkey
- Eylül Sevinç Sevinç, Computer Science, Koc university, Turkey
- Fusun Can, School of Medicine, Koc university, Turkey
- Önder Ergönül, School of Medicine, Koc university, Turkey
-
Ozlem Keskin, Professor of Chemical and Biological Engineering, Koc Univeristy, Turkey
- Attila Gürsoy, Professor of Computer Science, Koc University, Turkey
Presentation Overview: Show
Antimicrobial resistance (AMR) is rapidly outpacing antibiotic development, posing a critical global health
threat. Many resistance mechanisms are mediated by specific bacterial proteins, suggesting that designing
binders to inhibit these targets, alongside discovering novel antimicrobial peptides (AMPs) offers a promising
therapeutic strategy. Here, we introduce AMP-BindGen, a framework that explores the ability of modern protein
design methods to jointly address targeted binding and antimicrobial activity. Building on Boltzgen, we guide
sequence generation using an external AMP predictor, Macrel, effectively coupling structural/biophysical
design with functional antimicrobial constraints. Our approach first generates diverse peptide candidates
using a flexible generative prior trained on AMP-like sequences. We then incorporate activity-aware guidance
through Macrel-based scoring, applying rejection sampling, soft reweighting, and diversity-preserving
selection to approximate the conditional distribution of peptides with high antimicrobial potential. This
allows us to explicitly bias generation toward functional AMPs while retaining sequence diversity.
Importantly, we extend this paradigm beyond unconstrained AMP generation by evaluating whether binder design
frameworks can be repurposed to produce AMPs that also target key AMR-related proteins, bridging two
traditionally separate design objectives: binder design and antimicrobial function. In silico results show
that AMP-BindGen significantly enriches sequences with high predicted antimicrobial activity compared to
unguided generation, while maintaining broad physicochemical diversity. Our findings suggest that integrating
binder design principles with AMP activity guidance provides a powerful and practical route for generating
multifunctional peptide candidates, opening new directions for combating antimicrobial resistance. Adding
Guidance doubled the generated binders' AMP-likeliness while maintaining binding metrics.
B-P.43: Missense 3D portal, enabling investigation and interpretation of missense variants in their
structural context
Track: Proteins and structural biology
-
Ryan Pye, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Ifigenia Tsitsa, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London,
United Kingdom
-
Gordon Hanna, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Anja Conev, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Haotian Zhao, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Grace Walker, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Wendy Tran, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Olivia Simmonds, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London,
United Kingdom
-
Suhail Islam, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
-
Michael J E Sternberg, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London,
United Kingdom
-
Alessia David, Centre for Integrative Systems Biology and Bioinformatics, Imperial College London, United
Kingdom
Presentation Overview: Show
The expansion of the three-dimensional coverage of the human proteome and those of other organisms, achieved
with AlphaFold, offers the unprecedented opportunity to assess the structural impact of millions of missense
variants.
The Missense3D Portal, freely available at https://missense3d.bc.ic.ac.uk/, is a suite of algorithms and
databases that leverage the availability of experimental and modelled 3D structure, including protein
complexes, to provide mechanistic explanations for the damaging effect of missense variants, such as
disruption of chemical bonds and steric clashes.
Structural predictions for >9 million missense variants, produced by our core algorithm, Missense3D, are
freely available through the Missense3d-αDB database and will soon be integrated into ProtVar at EBI.
Furthermore, structural predictions for >4 million human variants, generated using experimental structures
and 3D models produced with our Phyre2 homology modelling tool, are available from the Missense3D-DB database
and integrated into DECIPHER at EBI.
Missense3D-PTM, published in JMB December 2025, is the latest addition to the Missense3D Portal. Developed as
part of the UK Human Functional Genomics Initiative, it is a “one-stop-shop†for dynamically investigating
missense variants in the context of post-translational modifications (PTMs). Its user-friendly web interface
provides mapping of 11,490,257 naturally-occurring human missense variants and 334,255 PTM sites
(spanning >50 PTM types, including phosphorylation, glycosylation, and sumoylation) along with their
neighbouring residues in sequence and 3D structure space, using AlphaFold models for 20,431
proteins. Furthermore, precalculated distances in 3D space between any PTM and any residue, enable
visualization and exploration of novel variants currently not included in the Missense3D-PTM
database.   
B-P.44: Designing Novel Enzymes with a Transformer-Guided Generative AI Framework
Track: Proteins and structural biology
-
Ufuk Tanrıverdi, Hacettepe University, Turkey
- Tunca Doğan, Hacettepe University, Turkey
Presentation Overview: Show
Designing functional proteins is central to therapeutics and biotechnology, but remains difficult due to the
astronomical sequence space and the loose mapping between sequence, structure, and function. Generative AI
offers a principled way to explore this space by learning the patterns underlying functional proteins. We
propose a generative adversarial network (GAN)-based framework for de novo protein design where both generator
and critic are initialized from a masked protein language model, ProtBERT, and fine-tuned in two-stages, first
on a large enzyme corpus, then on subclass-specific data, using dynamic masking. The generator performs
iterative partial decoding, retaining only the top-confidence ~10% of tokens per step while remasking the rest
until convergence. We evaluate two generation modes: generation from 90%-masked real sequences and fully blind
generation from entirely masked inputs. As a case study, we apply the framework to DNA methyltransferases
(DNMTs), particularly DNMT3A, which regulate gene expression with epigenetic modification and are essential
for genome stability. Its dysregulation is linked to developmental disorders and cancer, making novel enzyme
design valuable. Generated sequences undergo a multi-stage evaluation pipeline: ESMFold pLDDT,
self-consistency via ProteinMPNN, DNMT-like structural similarity via Progres, and pairwise TM-score diversity
serve as initial filters. Top candidates proceed to AlphaFold3 co-folding with DNA and SAM, with DNA-SAM
binding distance as the final functional criterion. Our best candidate exhibited secondary structure
organization resembling reference DNMT-3a with spatially similar alpha-helical and beta-sheet elements, and
retained the hydrophilic interactions necessary for SAM-mediated DNA methylation, despite being a completely
novel sequence.
B-P.45: Benchmarking Classical and Modern Structural Alignment Algorithms for Protein–Protein Interfaces in
Template-Based Docking
Track: Proteins and structural biology
-
Reza Mohammadian Mazraeh Shadi, Graduate School of Sciences and Engineering, Koç University,
Turkey
- Fatma Cankara, Graduate School of Sciences and Engineering, Koç University, Turkey
-
Nurcan Tuncbag, Department of Chemical and Biological Engineering, and School of Medicine, Koç
University, Turkey
- Attila Gursoy, Department of Computer Engineering, Koç University, Turkey
-
Ozlem Keskin, Department of Chemical and Biological Engineering, Koç University, Turkey
Presentation Overview: Show
Protein–protein interactions are pivotal for many cellular functions, and understanding their underlying
mechanisms is key to elucidating biological processes. Among several approaches for identifying
protein–protein interactions, template-based docking is particularly powerful, with performance strongly
governed by template library selection and structural alignment quality. In this study, we systematically
assess structural alignment algorithms for detecting structurally similar protein–protein interfaces in
template-based docking.
Using a curated subset of our PiFace interface dataset, representing non-contiguous,
sequence-order–independent binding regions (9,775 interfaces in 78 clusters), together with the MALISAM and
MALIDUP motif/domain benchmarks and their non-sequential variants (742 pairs total), we compare TM-align with
more recent algorithms, including US-align, Multiprot, KPAX and KPAX-flex, GTalign, Foldseek, and DeepAlign.
Performance is evaluated in docking-relevant interface-to-global and interface-to-interface scenarios using
TM-score, root-mean-square deviation (RMSD), and alignment length.
Across more than 2.6 million interface-centered docking case comparisons, TM-align, US-align, and GTalign
achieve consistently strong performance, with mean TM-scores of ~0.67–0.80 in docking-relevant cases, mean
RMSDs of ~1.3–1.6 Å, and normalized alignment coverage of ~0.83–1.00. In positive controls, TM-scores
approach 1.0 with near-complete coverage. KPAX-flex often attains low RMSDs with shorter alignments. GTalign
approach matches TM-align's accuracy while providing substantial GPU-enabled speed-ups for enabling
large-scale template screening. Protein language model–based tools, such as Foldseek, benefit from geometric
re-scoring for rapid template retrieval, although they may over-extend alignments. Overall, our results
delineate trade-offs among accuracy, coverage, and computational cost, and provide practical guidance for
interface-driven, template-based docking pipelines.
B-P.46: The Role of Local Energetic Frustration Across Conformational Ensembles of Disordered
Proteins
Track: Proteins and structural biology
-
Franco L Simonetti, Barcelona Supercomputing Center, Spain
- Edgar P. Chacón, Barcelona Supercomputing Center, Spain
-
Maria I. Freiberger, Sorbonne Université, CNRS, IBPS, Laboratory of Computational and Quantitative
Biology, France
-
Alexander M. Monzón, Department of Information Engineering, University of Padova, Padova, Italy, Italy
- Leandro Radusky, Mimark Diagnostics S.L., Parc Cientific Barcelona, Spain
-
Diego U. Ferreiro, Protein Physiology Lab, Facultad de Ciencias Exactas y Naturales, Universidad de Buenos
Aires, Argentina
- R. Gonzalo Parra, Barcelona Supercomputing Center, Spain
Presentation Overview: Show
Globular proteins satisfy the minimum frustration principle, i.e. their folding landscapes are strongly
funneled towards their native state with residual energetic conflicts often associated with function. In
contrast, intrinsically disordered proteins populate heterogeneous conformational ensembles as a consequence
of an excess of competing interactions and therefore flatter energy landscapes. Still, the specific energetic
signatures that lead to protein disorder are poorly understood. Here, we characterized the local frustration
patterns of 2,964 proteins and 54,766 conformers spanning different flavours of intrinsic disorder.
We found that in the case of direct residue-residue interactions, except for a minor contribution of specific
interaction types, disordered residues are not excessively highly frustrated; instead, they are enriched in
neutrally frustrated interactions and a reduced prevalence of minimally frustrated ones. Furthermore, we found
that the most distinguishable feature between ordered and disordered residues comes from long range
interactions where disordered regions are depleted of stabilizing interactions and enriched in energetic
conflicts. When disordered regions are analysed in complex with known protein partners, their energetic
profiles resemble those of ordered residues.
We conclude that intrinsic disorder is not a consequence of excessive energetic conflicts as believed for the
case of statistical random heteropolymers. Instead, disordered proteins exist at the edge of foldability,
where a small number of stabilizing interactions with other proteins or substrates can shift their equilibrium
towards more structured, functional conformations.
B-P.47: A Replicate-Based Molecular Dynamics Framework for Mapping Variant-Induced Structural and Network
Perturbations in Drug-Metabolising NAT2 Enzyme
Track: Proteins and structural biology
-
Wayde Veldman, Rhodes University, South Africa
- Ozlem Tastan Bishop, Rhodes University, South Africa
Presentation Overview: Show
Investigating how residue variations alter enzyme structure and dynamic communication remains essential for
understanding inter-individual variation in drug metabolism. We performed all-atom molecular dynamics
simulations of human arylamine N-acetyltransferase 2 (NAT2), the enzyme responsible for metabolising the
tuberculosis drug isoniazid, to determine how sequence variation drives the transition from rapid to slow
acetylation phenotypes.
We developed a replicate-based comparative framework in which replicate-averaged rapid acetylator simulations
define a reference state. Slow acetylator deviations were quantified relative to the reference mean and
standard deviation across hydrogen bonding patterns, residue flexibility, and dynamic residue network
centrality metrics which are then mapped to enzyme structure. This approach enables systematic identification
of structurally and functionally meaningful perturbations rather than relying on single-trajectory
comparisons.
Slow acetylator variants exhibited destabilisation of the active conformation characterised by altered
hydrogen bonding, increased residue flexibility, and disrupted residue communication networks. In the
R64Q+K268R variant for example, loss of hydrogen bonding at residue 64 propagated through the structural
network, reducing betweenness centrality of catalytic residue D122 and putative isoniazid-binding residues
S125 and F217. On the other hand, eigenvector centrality was reduced for N72, D122, and active-site loop G124,
indicating reduced residue communication. In line with these changes, altered active-site geometry was
observed via shortening of distances between binding residue F217 and catalytic residues H107 and D122.
Together, these results provide mechanistic insight into how distal residue variations propagate through
dynamic residue networks to impair catalytic organisation, offering a transferable framework for studying
structure/function relationships in pharmacogenetically relevant enzymes.
B-P.48: Benchmarking Deconvolution Tools for Proteomics Data
Track: Proteins and structural biology
-
Tobias Scheithauer, Institute for Machine Learning, ETH Zurich, Switzerland
- Ludovica Sibilia, Institute for Machine Learning, ETH Zurich, Switzerland
- Mira Herold, Institute for Machine Learning, ETH Zurich, Switzerland
- Samuel Gair, Institute for Machine Learning, ETH Zurich, Switzerland
-
Marta Pinto Carbo, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
-
Lopamudra Chatterjee, Swiss Institute of Allergy and Asthma Research (SIAF), University of Zurich,
Switzerland
-
Nadia Djerbi, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
-
Dorothea Rutishauser, Department of Pathology and Molecular Pathology, University Hospital Zurich,
Switzerland
-
Rosary Yao, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
-
Thi Huong Lan Do, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
-
Marco Bühler, Department of Pathology and Molecular Pathology, University Hospital Zurich, Switzerland
-
Christoph Messner, Swiss Institute of Allergy and Asthma Research (SIAF), University of Zurich,
Switzerland
-
Thorsten Zenz, Department of Medical Oncology and Hematology, University Hospital Zurich, Switzerland
- Valentina Boeva, Institute for Machine Learning, ETH Zurich, Switzerland
Presentation Overview: Show
The cellular composition of the tumor microenvironment is essential for diagnosis, prognosis, and treatment
selection in precision oncology. Computational deconvolution of bulk omics data can infer cell type
proportions from mixed samples without physical cell sorting. While numerous deconvolution tools exist for
transcriptomics, very few target bulk proteomics data. The distinct nature of proteomic data raises the
question whether transcriptomics-derived tools can be reliably applied to proteomics without adaptation.
We conducted a multi-dimensional benchmarking study to evaluate the applicability of existing deconvolution
methods to bulk proteomics data. Using two public immune cell proteomics, we generated synthetic bulk mixtures
with known cell type proportions in different mixture composition settings. Our benchmark evaluates six
deconvolution algorithms using different imputation strategies for missing values and various signature matrix
construction methods.
Our benchmarking reveals a substantial performance degradation of all algorithms when moving from idealistic
to more realistic simulation settings, indicating that published benchmarks may overestimate real-world
accuracy. Second, the optimal algorithm choice depends on the combination of preprocessing and experimental
scenario. Third, we observe systematic cell-type-specific biases showing as consistent over- or
underestimation of individual cell types. Fourth, small signatures based on well-characterized marker proteins
may be more robust than larger data-driven signature matrices. Finally, unsupervised methods demonstrate
limited reliability across the range of tested conditions.
Our findings motivate both the development of dedicated, proteomics-aware deconvolution tools and the creation
of experimentally validated benchmark datasets to enable robust immune cell profiling in precision oncology.
B-P.49: Why Think Multimodal? Data Treatment for Integrated Structural Biology
Track: Proteins and structural biology
-
Wojciech Potrzebowski, SciLifeLab Lund, Centre for Research Infrastructure in Health and Sciences,
Lund University, Lund, Sweden, Sweden
-
Filip Arman, Centre for Research Infrastructure in Health and Sciences, Lund University, Lund, Sweden,
Sweden
-
Joao Figueira, Swedish NMR Centre, Department of Chemistry, Umeå University, Umeå, Sweden, Sweden
-
Sayyed Jalil Mahdizadeh, Swedish NMR Centre, Department of chemistry and molecular biology, University of
Gothenburg, Sweden., Sweden
-
Anton Sellerberg, Centre for Research Infrastructure in Health and Sciences, Lund University, Lund,
Sweden, Sweden
-
Simon Ekstrom, Centre for Research Infrastructure in Health and Sciences, Lund University, Lund, Sweden,
Sweden
-
Hanna Kultima, Department of Immunology, Genetics, and Pathology, Uppsala University, Uppsala, Sweden,
Sweden
-
Johan Rung, Department of Immunology, Genetics, and Pathology, Uppsala University, Uppsala, Sweden, Sweden
-
Nicholas Pearce, Department of Cell and Molecular Biology, SciLifeLab, Uppsala University, Sweden, Sweden
Presentation Overview: Show
Integrated structural biology increasingly depends on combining data from multiple experimental and
computational approaches to answer biological questions that cannot be resolved by any single technique alone.
However, multimodal projects are still often organized around individual methods, which creates fragmentation
in data handling, metadata capture, quality control, and downstream reuse. This slows scientific progress,
makes cross-technique interpretation more difficult, and increases the risk that valuable intermediate data,
context, and decisions are lost before publication or deposition.
At the SciLifeLab Integrated Structural Biology (ISB) platform, multimodal data treatment is being developed
as a way to shift focus from techniques to scientific questions. By treating structural biology data as
connected evidence streams rather than isolated outputs, it becomes possible to use one modality to support,
validate, or troubleshoot another, and to follow the sample and data journey from production to integrated
interpretation. Such an approach is also essential for making data FAIR, reusable, and increasingly AI-ready,
with sufficiently rich metadata, validation, and provenance to support future method development and
cross-study analysis.
Here we describe the rationale, challenges, and practical principles for multimodal data treatment at the
SciLifeLab ISB platform, with a focus on data integration across techniques, preservation of experimental
context, and preparation of high-quality datasets for sharing, reuse, and AI-driven analysis.
B-P.50: Reading TEA-Leaves for de novo protein design
Track: Proteins and structural biology
-
Ieva Pudžiuvelytė, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel,
Switzerland, Switzerland
-
Lisa Brandenburg, Department of Biosystems, Science and Engineering, ETH Zürich, Basel, Switzerland,
Switzerland
-
Lorenzo Pantolini, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel,
Switzerland, Switzerland
-
Basile Wicky, Department of Biosystems, Science and Engineering, ETH Zürich, Basel, Switzerland,
Switzerland
-
Janani Durairaj, Biozentrum, University of Basel; SIB Swiss Institute of Bioinformatics, Basel,
Switzerland, Switzerland
Presentation Overview: Show
The field of de novo protein design aims to produce novel protein sequences or even novel structures with
improved functional properties for biotechnological applications. The breakthrough in protein folding methods
opened ways to evaluate whether sequences fold into intended structures in silico, which catalysed the protein
design progress. Orthogonally, we witnessed the appearance of protein language models (pLMs) encoding
biological properties in their embeddings. However, pLM-based iterative optimization approaches requiring
repeated structure evaluations remain computationally expensive, limiting their practical throughput.
We propose that the structure-informed alphabet derived from pLM embeddings (TEA) could address this
bottleneck. Therefore, we present TEA-Leaves method leveraging TEA alphabet to efficiently guide the Markov
chain Monte Carlo sampling to generate de novo sequences folding into template structures.
We showcase TEA-Leaves functionality in several scenarios by generating: de novo sequences folding into a
diverse set of monomeric structures and sequence-similar but structure-dissimilar pairs of proteins.
In addition to structure prediction confidence metrics from ColabFold, we used several in silico quality
metrics to select designs for experimental validation: values of TEA-Leaves losses, solubility, and
aggregation scores. Namely, we selected designs with sequence novelty relative to natural proteins to assess
the ability of TEA-Leaves to explore sequence space beyond homology adjacency. The quality scoring filters
were collected in a pipeline Sieve to prioritize the candidates for the downstream experimental validation of
expressibility, solubility, and oligomerization propensity.
Altogether, using pLM-derived structural signals we aim to probe the determinants of protein folding outside
of the homology barrier.
B-P.51: Identification of Highly Conserved Conformational Epitopes on the SARS-CoV-2 Spike Glycoprotein: A
Computational Foundation for Broad-spectrum Antibody Design
Track: Proteins and structural biology
-
Cheng-Wei Cheng, Kaohsiung Medical University, Taiwan
Presentation Overview: Show
The continuous emergence of SARS-CoV-2 variants underscores the critical need for broad-spectrum therapeutic
antibodies. Identifying highly conserved conformational epitopes on the spike glycoprotein, which represent
regions that remain invariant across diverse lineages, is essential for developing treatments resistant to
viral escape. This study utilizes an integrated computational pipeline to pinpoint these critical structural
motifs.
Our framework utilized reconstructed high-resolution structures of the SARS-CoV-2 spike glycoprotein, which
were computationally refined to restore missing residues and fully glycosylated to reflect native
physiological conditions. By calculating the Relative Solvent Accessibility (RSA) of each residue, we
categorized amino acids as exposed, shielded, or buried. This step ensures that candidate epitopes are not
only conserved but also physically accessible to antibodies. Simultaneously, we performed an evolutionary
analysis using the covSPECTRUM platform to track recent mutation trajectories, identifying residues with
exceptionally high conservation across the latest global variants.
By synthesizing RSA data with evolutionary conservation scores, we identified and validated several highly
conserved conformational epitopes. These spatial clusters represent ""vulnerability patches"" that are
evolutionarily constrained and structurally accessible. The preliminary results demonstrate that these
validated regions are distinct from highly variable domains prone to mutation. These findings highlight that
these highly conserved conformational epitopes serve as high-priority targets for antibody design. Our study
provides a strategic roadmap for engineering next-generation biologics capable of maintaining long-term
potency against the evolving landscape of SARS-CoV-2 and future sarbecoviruses.
B-P.52: BiochemicalAlgorithms.jl for Stochastic Protein Energy Minimization
Track: Proteins and structural biology
-
Jennifer Leclaire, Institute for Computer Science, JGU Mainz, Germany
-
Thomas Kemmer, Institute of Computer Science, Johannes Gutenberg University Mainz, Germany
-
Andreas Hildebrandt, Institute of Computer Science, Johannes Gutenberg University Mainz, Germany
Presentation Overview: Show
Computational methods for structural biology often begin with rapid prototyping in Python. Performance
demands, however, frequently lead to implementations in C or C++, resulting in software packages with C++
cores and Python user interfaces. The Biochemical Algorithms Library (BALL) is a representative example of
this approach: a well-established C++ framework initiated in 1996, with Python bindings added in 2010. BALL
was once the largest open-source library of its kind, offering broad functionality for molecular structure
analysis.
This motivated the redesign of the library and led to our Julia-centric framework, BiochemicalAlgorithms.jl.
It provides core data structures, file I/O, preprocessing, molecular mechanics, optimization algorithms, and
Julia ecosystem integration across the structural bioinformatics pipeline. Moving development to Julia has
simplified the implementation of key design goals, particularly rapid application development.
Building on this recently published framework, we will demonstrate its practical value through a stochastic
gradient descent-based protein energy minimization scheme. While classical energy minimization is often
performed with deterministic methods such as conjugate gradient, stochastic optimization, as used in deep
learning, introduces controlled randomness that can support broader exploration of the protein energy
landscape and help avoid premature trapping in local minima. In this sense, the example serves both as an
optimization strategy and as a demonstration of how quickly new algorithmic ideas can be implemented within
BiochemicalAlgorithms.jl. More broadly, this work highlights how the framework supports efficient method
development and provides a practical environment for building computational tools in structural
bioinformatics.