View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
C-S.B.01: Notation variance in chemical language models: effects of inconsistent SMILES on representation
stability and benchmark evaluation
Track: Systems biology, multi-omics integration, modeling
-
Tadahaya Mizuno, The University of Tokyo / The Institute of Mathematical Statistics, Japan
- Yosuke Kikuchi, The University of Tokyo, Japan
- Yasuhiro Yoshikai, The University of Tokyo, Japan
- Shumpei Nemoto, The University of Tokyo, Japan
- Ayako Furuhama, National Institute of Health Sciences, Japan
- Takashi Yamada, National Institute of Health Sciences, Japan
- Hiroyuki Kusuhara, The University of Tokyo, Japan
Presentation Overview: Show
Chemical language models (CLMs) use molecular strings such as SMILES as direct model inputs. However, the same
molecule can be represented by multiple valid strings, and even “canonical†SMILES are not uniquely defined
across software toolkits. This raises the possibility that notation variance, when not harmonized, affects
learned representations and benchmark evaluation.
We investigated how inconsistent SMILES influence CLM behavior. First, in a survey of 264 CLM-related papers
indexed in PubMed, about half did not explicitly report their canonicalization procedure, indicating limited
transparency in handling notation variance. We then examined public benchmark datasets and found substantial
heterogeneity in molecular notation, including redundant aromatic forms and frequent omission of
stereochemical annotations. Using a molecular translation model trained on RDKit-standardized SMILES, we
compared dataset-provided Raw SMILES with standardized inputs. Translation accuracy was consistently lower for
Raw SMILES, indicating impaired structural recovery. In addition, Levenshtein distance between paired input
strings positively correlated with L2 distance between their latent vectors, showing that character-level
differences in notation were associated with divergence in representation. In several property-prediction
tasks, performance degradation was limited, suggesting that downstream models can partially absorb unstable
latent components. However, in ClinTox and BBBP, Raw SMILES yielded higher AUROC, and further analyses
indicated that this apparent improvement was driven by notation-dependent confounding rather than improved
chemical learning.
These results show that unharmonized notation variance can influence learned representations and, in some
cases, distort benchmark evaluation in CLMs. Careful control and explicit reporting of molecular notation are
therefore important for reproducible molecular machine learning.
C-S.B.02: Integrative multi-omics analysis identifies a phenol-associated microbe-metabolite-host interaction
in MASLD
Track: Systems biology, multi-omics integration, modeling
-
Yuan Wang, Imperial College London, United Kingdom
- Manyi Jia, Imperial College London, United Kingdom
- Kanta Chechi, Imperial College London, United Kingdom
-
Zhaojie Wang, European Genomics Institute for Diabetes, Institut Pasteur de Lille, Lille University
Hospital, University of Lille, France
- Fiona Newberry, Nottingham Trent University, United Kingdom
- Marina Cardellini, Tor Vergata University of Rome, Italy
- Rossella Menghini, Tor Vergata University of Rome, Italy
- José María Moreno-Navarrete, IDIBGI, Instituto de Salud Carlos III, Spain
- Jordi Mayneris-Perxachs, IDIBGI, Instituto de Salud Carlos III, Spain
-
Ulrike Löber, Max Delbrück Center for Molecular Medicine in the Helmholtz Association (Max Delbrück
Center), Germany
-
Sofia Forslund, Max Delbrück Center for Molecular Medicine in the Helmholtz Association (Max Delbrück
Center), Germany
- Rémy Burcelin, Université de Toulouse, France
- Julian Marchesi, Imperial College London, United Kingdom
- Miriam Moffatt, Imperial College London, United Kingdom
- Lesley Hoyles, Nottingham Trent University, United Kingdom
- Jose Manuel Fernández-Real, IDIBGI, Instituto de Salud Carlos III, Spain
- Massimo Federici, Tor Vergata University of Rome, Italy
- Marc-Emmanuel Dumas, Imperial College London, United Kingdom
Presentation Overview: Show
Metabolic dysfunction-associated steatotic liver disease (MASLD) is a complex metabolic disease associated
with multiple systemic comorbidities and poses a major public health challenge. Although growing evidence
suggests alterations in the gut microbiome and its derived metabolites contribute to MASLD development, the
mechanisms underlying microbe-metabolite-host interactions remain incompletely understood. Here, we present an
integrative multi-omics workflow for linking gut microbial composition and functional potential with microbial
metabolites and disease-related host phenotypes.
We applied this workflow to the FLORINASH cohort, comprising 662 non-diabetic individuals with MASLD from
Spain and Italy, spanning a range of obesity levels and liver disease severity. For each participant, matched
shotgun metagenomic sequencing, untargeted UHPLC-MS metabolomics, and clinical phenotyping data were
available. Metagenomic data were used to generate species-level taxonomic profiles and KEGG Orthology-based
functional annotations, while gutSMASH was applied to predict microbial metabolic gene clusters across the
identified gut species. These results were then integrated with metabolomic and clinical data to characterise
microbe-metabolite-host associations.
Using this workflow, we identified a candidate phenol-associated host-microbiome interaction in MASLD.
Phascolarctobacterium was implicated in hydroxybenzoate-to-phenol metabolism, with its abundance positively
associated with marker metabolites of phenol metabolism. Mediation analysis further suggested that both
Phascolarctobacterium abundance and related KEGG Orthology gene counts were linked to kidney function-related
clinical features via phenol-associated metabolites in MASLD.
Together, these findings highlight the utility of our multi-omics workflow for uncovering
microbe-metabolite-host interactions and suggest its potential for biomarker discovery and microbiome-targeted
therapeutic strategies.
C-S.B.03: Data-Driven Disentanglement of Confounding Factors in Sjögren's Syndrome
Track: Systems biology, multi-omics integration, modeling
-
Kristina Lacasta Lopez, University of Seville, Spain
- Angela Gandara Alvarez, University of Seville, Spain
- MarÃa Jiménez Rus, University of Seville, Spain
- Cristiane Cantiga Silva, University of Seville, Spain
- Virginia Moreira Navarrete, Virgen Macarena University Hospital, Spain
- Carmen Domínguez Quesada, Virgen Macarena University Hospital, Spain
- Jose Javier Perez Venegas, Virgen Macarena University Hospital, Spain
- Juan Antonio Ortega, University of Seville, Spain
- Aurea Simon-Soro, University of Seville, Spain
Presentation Overview: Show
Clinical machine learning models are highly vulnerable to confounding bias, often exploiting systemic
variables rather than intrinsic pathological signals. Sjögren's Syndrome (SS) exemplifies this challenge as
physiological aging and polypharmacy can create a confounding profile, rendering elderly controls clinically
similar to autoimmune patients. We hypothesized that systematic computational disentanglement could expose
this bias and recover a biologically interpretable disease-related signature. In this study, we propose a
confounding-aware multi-stage computational framework and apply it to an observational cohort of 196 women.
The framework integrates Principal Component Analysis (PCA), Random Forest with SHAP, and multivariable
regression. Unsupervised PCA revealed a latent metabolic and pharmacological phenotype in controls that
simulates autoimmune xerostomia. Strikingly, standard classifiers (AUC 0.92) relied heavily on these
comorbidities, indicating shortcut learning. By ablating this systemic signal, we isolated a de-confounded
oral model that retained substantial diagnostic performance (AUC 0.82), primarily driven by cumulative dental
damage and stimulated salivary flow. Multivariable regression showed that SS diagnosis and xerogenic
medication use were independently associated with lower stimulated salivary flow, whereas age showed a smaller
negative association. The magnitude of the disease effect was comparable to 28 years of age-related decline,
although this comparison should be interpreted as a coefficient-based approximation. Age-stratified analysis
revealed reduced diagnostic performance in patients older than 70 years, suggesting increasing phenotypic
overlap between physiological aging and disease. Overall, unadjusted algorithms may overestimate diagnostic
performance in complex clinical settings. These findings highlight the need for interpretable,
confounding-aware workflows to capture disease-related signals and support reliable, age-aware clinical
decision-making.
C-S.B.04: AI-driven classification of signaling proteins: histidine kinases as a case study
Track: Systems biology, multi-omics integration, modeling
-
Louison Silly, Laboratoire de Biométrie et de Biologie Évolutive, Lyon 1 - BIAM, CEA de Cadarache,
France
- Guy Perrière, Laboratoire de Biométrie et Biologie Évolutive, UMR CNRS 5558, France
-
Philippe Ortet, Institute of Bioscience and Biotechnology of Aix-Marseille, UMR CEA/CNRS/AMU 7265, France
Presentation Overview: Show
In prokaryotes, especially bacteria, signal transduction is often carried out by two components systems (TCS).
The classical TCS is made of an histidine kinase (HK) and a response regulator (RR). Upon signal recognition,
the HK autophosphorylate on a specific histidine residue and then transfers its phosphate group to its partner
RR, who will regulate gene expression through various means. It has been shown that TCS are involved in
response to many environmental stimuli, like light sensing or biochemical changes. Studying these systems can
help understand the different types of signals a cell can perceive and how it will adapt to them. Several
classifications of HKs have been proposed, based on conserved motif in their C-terminal end. One such motif is
the H-Box that contains the phosphorylable histidine. These classifications where made in the early 2000s and
does not include the diversity of sequenced organisms (and their TCS) that are available nowadays. We propose
here a new way of classifying HKs, based on the neighborhood of the phosphorylable histidine, using protein
Language Models (pLM). We use Kernel Principal Component Analysis (Kernel PCA) to reduce the dimension of the
embeddings produced by the pLMs, followed by an unsupervised clustering. With this approach we are able to
group HKs accordingly to the existing classification schemes and to identify clusters of HKs that could form
new groups previously unseen. We plan to extend this classification method to RRs, opening new perspectives
for motif-centered exploration of signaling proteins and beyond.
C-S.B.05: Towards Comparative QTLomics
Track: Systems biology, multi-omics integration, modeling
- Alex Warwick Vesztrocy, BioSoft Research UK, United Kingdom
-
Natasha Glover, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
- Christophe Dessimoz, University of Lausanne, Switzerland
-
Irene Julca, Aarhus univeristy, Denmark
Presentation Overview: Show
New plant breeding technologies rely on identifying genes associated with key agronomic traits, such as yield,
fruit size, or disease resistance. Quantitative trait loci (QTL) studies have been important in linking
genomic regions to these complex traits. However, QTL regions can contain hundreds of genes, which hinders
further experimental validation. With QTL data available for over 400 species and more than 1,000 plant
genomes, there is an unprecedented opportunity to integrate cross-species information through Comparative
QTLomics. Here, we present a phylogeny-aware tool, currently under development, that integrates QTL data with
phylogenomic relationships and functional evidence to improve candidate gene prioritisation. As a proof of
concept, we applied this approach to the trait fruit size using QTL data from tomato, bell pepper, melon, and
watermelon. Preliminary results show that Comparative QTLomics can refine candidate gene lists, recovering
genes with known functional roles while also identifying novel candidates for future validation. This work
further emphasises the importance of using a consistent ontology to standardise QTL data and complementary
functional information to improve gene prioritisation. Overall, the framework provides an alternative strategy
to use existing QTL information to accelerate the identification of genes underlying complex traits, including
those lacking prior annotation.
C-S.B.06: TopOmics: Topic Modelling for all -Omics
Track: Systems biology, multi-omics integration, modeling
-
Federico Caretti, Scuola Internazionale Superiore di Studi Avanzati, Italy
- Nour El Kazwini, Scuola Internazionale Superiore di Studi Avanzati, Italy
- Guido Sanguinetti, Scuola Internazionale Superiore di Studi Avanzati, Italy
Presentation Overview: Show
Topic models have emerged as a popular paradigm to analyse and interpret complex single-cell and spatial data.
Yet, current implementations are usually data-type specific and rely on different modelling and estimation
approaches, hindering usability and interoperability. In this work we introduce TopOmics, a library to perform
efficient and flexible topic modeling with any combination of -omics data, including novel spatial multi-omic
data. The framework leverages standard libraries of the Python ecosystem, guaranteeing seamless integration
with existing pipelines, and shows competitive performance against state-of-the-art methods while preserving
interpretability. We provide several examples of TopOmics on diverse data sets, demonstrating the usefulness
of a unified framework for modern bioinformatic analyses.
C-S.B.07: MetaNetX: enhancing metabolomics data integration through comprehensive reconciliation
Track: Systems biology, multi-omics integration, modeling
-
Sebastien Moretti, SIB Swiss Institute of Bioinformatics, Switzerland
- Anne Niknejad, Vital-IT Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Marco Pagni, Vital-IT Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Florence Mehl, Vital-IT Group, SIB Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Integrating metabolomics with other omics and computational models
provides a holistic view of biological systems. MetaNetX
(https://www.metanetx.org) plays a key role by reconciling metabolites
and biochemical reactions across major databases (ChEBI, HMDB,
KEGG, MetaCyc, Reactome, SwissLipids, LIPID MAPS, etc.), creating a
unified namespace for annotation and cross-referencing.
Despite challenges such as incomplete stereochemistry and database
inconsistencies, MetaNetX uses molecular structures and reaction
context to improve accuracy, with manual curation for critical cases.
The resource is freely available as raw files, via SPARQL, and through an
ID mapping tool, supporting systems biology and metabolomics
research.
C-S.B.08: Integrating tumour evolutionary patterns with genomic and transcriptomic signatures improves
patient stratification in metastatic prostate cancer.
Track: Systems biology, multi-omics integration, modeling
-
Richard Norris, Vall d'Hebron Institute of Oncology, Spain
- Joaquin Mateo, Vall d'Hebron Institute of Oncology, Spain
Presentation Overview: Show
Background
Prostate cancer (PC) is one of the most common cancers in men. Understanding how genomic profiles map to
transcriptional programs and clinical phenotypes is central to patient stratification. Homologous
recombination deficiency (HRD) defines a clinically relevant subset of metastatic (m) PC; however, DNA-based
HRD markers (including BRCA1/2 alterations) have limited discriminatory power.
Methods
Matched WGS and RNA-seq data from >300 mPC patients (Hartwig Medical Foundation) were analysed using
differential expression and GSEA to test whether HRD can be resolved through integrated genomic and
transcriptomic features. To investigate whether similar transcriptional profiles arise via distinct
evolutionary pathways, we developed a computational framework to estimate tumour clonality by integrating
somatic SNVs, SVs, and CNVs.
Results
We identified 31 HRD cases, including 19 with BRCA2 alterations. GSEA revealed enrichment of an RB1-loss
signature in HRD tumours; however, this signal was also present in TP53-altered tumours (n = 135), limiting
specificity. HRD tumours showed limited subclonal diversity (clonality >0.9; 1 = fully clonal), consistent
with early genomic instability and clonal fixation. In contrast, non-HRD tumours exhibited subclonal
heterogeneity (<0.9), indicative of ongoing evolutionary diversification. TP53 and RB1 alterations were
associated with increased subclonal burden. Differences remained significant after adjusting for tumour purity
and treatment exposure.
Conclusions
HRD tumours show restricted subclonal diversity, whereas non-HRD tumours with TP53 or RB1 alterations exhibit
greater subclonal complexity, consistent with poor prognosis. Integrating subclonality with driver alterations
and transcriptomic profiles improves patient stratification beyond genomic alterations alone.
C-S.B.09: A unified single-cell atlas of mouse tissue damage across vascular disease models
Track: Systems biology, multi-omics integration, modeling
-
Shamim Ashrafiyan, Goethe University Frankfurt, Germany
- Carolin Becker, Goethe University Frankfurt, Germany
- Iaroslav Kosaretskii, Goethe University Frankfurt, Germany
- Marcel H Schulz, Goethe University Frankfurt, Germany
Presentation Overview: Show
Single-cell RNA-sequencing (scRNA-seq) offers powerful insights into cellular responses across tissues and
disease states. We present an integrated atlas of 390,000 mouse cells across multiple tissues (heart, carotid,
lung, and brain) and disease models, including heart disorder, stroke, and lung injury. To harmonize technical
and biological variation across datasets, we applied scVI to integrate raw read counts and effectively removed
technical batch effects.
Cell type annotation was performed using curated marker genes from the literature. We compared
disease-associated gene signatures with gene–disease associations from DisGeNET, enabling the identification
of overlapping and condition-specific genes.
The atlas provides a versatile framework for downstream investigations.
It enables exploration of shared and unique transcriptional programs across diseases, identification of
disease-specific cellular states, and analysis of conserved patterns across conditions. Importantly, the
learned latent representations from this atlas can be applied to spatial transcriptomics data, allowing
inference of disease presence. Another application is the inference of dynamic transcription factor (TF)
networks from time-series gene expression data, allowing the identification of stage-specific regulators
across disease progression.
To ensure broad accessibility, we developed a user-friendly web-based platform, the VDA app, which enables
intuitive exploration and analysis without requiring programming expertise. Users can explore the atlas and
its rich metadata, perform differential expression analyses, and investigate gene functions across specific
cell types and disease conditions.
C-S.B.10: Evaluating and Improving Optimal Transport for Temporal Single-Cell RNA-Seq Data.
Track: Systems biology, multi-omics integration, modeling
-
Shashank Tiwari, Max Delbrück Center for Molecular Medicine, Germany
- Dr. Jana Wolf, Max Delbrück Center for Molecular Medicine, Germany
- Dr. Laleh Haghverdi, Max Delbrück Center for Molecular Medicine, Germany
- Dr. Bjoern Goldenbogen, Max Delbrück Center for Molecular Medicine, Germany
Presentation Overview: Show
Single-cell sequencing (scRNA seq) has transformed our understanding of biological systems by revealing
extensive cellular heterogeneity and diverse cell states across time. However, these measurements only provide
us snapshot observations of the system. Mapping these discrete snapshots onto continuous trajectories requires
computational modeling. Optimal Transport (OT) has emerged as a useful framework to predict cell-state
transitions by creating a coupling between two distributions across time-points. In particular, by using the
unbalanced OT formulation it is possible to account for cell growth and death by applying a penalty for the
imbalance of mass between time-points.
However, the impact of this relaxation and its related parameters have not been systematically analyzed. We
created a simulated dataset with known ground truth to investigate how modeling assumptions, including prior
estimate of birth-death rate, distance between cells and relaxation parameters for OT, impact the resulting
transition maps. Our computational experiments revealed that in developing systems with evolving cell-type
composition, OT successfully identifies cell-type transitions. However, in a stationary system where cell-type
composition stays the same across time-points, OT performs well in capturing cell growth and death but often
fails to reliably capture transitions to other cell types.
To address this problem we introduced biologically informed weights using a Reference Measure. This results in
better disentanglement between cell-type transitions and growth dynamics while preserving the convex structure
of the problem. Our findings provide a critical evaluation of the current OT approach and offer practical
solution to improve the performance of OT in trajectory inference.
C-S.B.12: A modular RDF schema for multi-source biomedical data integration: application to Crohn's
disease
Track: Systems biology, multi-omics integration, modeling
-
Domenico Palladino, University of Salerno, Via Giovanni Paolo II, 132, 84084 Fisciano (SA), Italy,
Italy
-
Anna Marabotti, University of Salerno, Via Giovanni Paolo II, 132, 84084 Fisciano (SA), Italy, Italy
-
Olivier Dameron, Univ Rennes, Inria, CNRS, IRISA - UMR 6074, F-35000 Rennes, France, France
-
Myriam Bontonou, Univ Rennes, Inria, CNRS, IRISA - UMR 6074, F-35000 Rennes, France, France
Presentation Overview: Show
Translational research routinely requires integrating clinical records, follow-up measurements, treatment
histories and molecular data. These layers are typically stored as independent tables with implicit temporal
links, limiting cross-domain querying and forcing study-specific schemas to be rebuilt for each new cohort.
Currently, although each type of data has well established representation formats, there is no standard data
schema supporting their integration, which leads to redundant engineering developments.
We are designing a generic, modular RDF schema that decouples the core data structure from domain-specific
ontology choices. The backbone models four reusable entities: subjects, clinical events, biological samples
and molecular measurements, linked by relations that make the time ordering of events explicit and queryable.
Disease-specific vocabularies are kept in a separate ontology layer, allowing the same backbone to be reused
across pathologies and data types.
TSV files structured according to the schema are imported into AskOmics, enabling visual SPARQL querying. The
framework is being instantiated on a longitudinal Crohn's disease cohort, integrating four clinical tables
spanning nearly 3,000 patients and over 40,000 follow-up records. Microbiome-derived functional data are being
added through the same modular mechanism.
The resulting knowledge graph enables queries combining semantic and temporal constraints, for example
retrieving samples collected within defined treatment windows and stratifying them by disease activity, which
would require complex multi-table joins in a relational approach. Decoupling the data backbone from domain
ontologies reduces schema redesign effort for new cohorts, supports integrative queries across clinical and
molecular layers and favours reproducibility across translational studies.
C-S.B.13: A Large-Scale Resource for Single-Cell H&E and Spatial Transcriptomics Reveals the Importance
of Tissue-Specific Learning
Track: Systems biology, multi-omics integration, modeling
-
Marc Glettig, ETH Zürich, Switzerland
- Aurélien Cormorèche, ETH Zürich, Switzerland
- Valentina Boeva, ETH Zürich, Switzerland
Presentation Overview: Show
Learning meaningful representations of single-cell resolution histology (H&E) images remains limited by
small, heterogeneous datasets and inconsistent annotations, constraining the development of robust image-based
cell representations. Recent work has proposed diverse architectures for training models on histology paired
with spatial transcriptomics data. We address these challenges by constructing a large-scale, standardized
resource and systematically assessing how training strategies impact downstream performance.
We curated and harmonized public datasets to assemble over 20 million individual cells with paired H&E
image patches and spatial transcriptomic profiles. Cells were annotated with hierarchical cell type labels at
multiple levels of granularity. We trained models using contrastive learning objectives to align image and
transcriptomic embeddings and benchmarked them against established cell-level image embedding approaches
across tasks including cell type classification and gene expression prediction.
While patch-level histology foundation models such as UNI2 and Virchow2 generalize well across samples,
specialized models such as CellViT exhibit strong sample-specific effects. Contrary to expectations,
integrating data across diverse tissue types did not improve general-purpose single-cell image embeddings and
in some cases reduced performance. In contrast, tissue-specific training consistently improved cell type
classification (+0.45 ARI) and gene expression prediction (+0.1 PCC), indicating that
morphological–molecular relationships are context-dependent.
These findings highlight the importance of tissue context for learning effective image embeddings and provide
a foundation for scalable, tissue-aware multimodal learning in computational pathology.
C-S.B.14: Optimized Clustering of Patients with Type 2 Diabetes (T2D): Mining Data from Population-Based
Cohort Studies
Track: Systems biology, multi-omics integration, modeling
-
Alina Skrylnik, Free University of Bozen-Bolzano / Eurac Research, Italy
- Giuseppe Tallini, Free University of Bozen-Bolzano, Italy
- Agathe Vasseur, University of Technology of Compiègne, France
- Anton Dignös, Free University of Bozen-Bolzano, Italy
- Christian Fuchsberger, Eurac Research, Italy
- Johann Gamper, Free University of Bozen-Bolzano, Italy
Presentation Overview: Show
Type 2 diabetes (T2D) is a highly heterogeneous disease that requires a personalised approach to treatment for
each patient. Several attempts have been made to identify subgroups within T2D; one of the most common is
based on clinical data. However, to enhance subgroup analysis, additional molecular layers need to be
incorporated. Multi-omics data, which combines genomics, proteomics and metabolomics data, is available from
cohort studies and may provide insights into T2D subgroups that are not captured by clinical data alone. In
this study, our aim is to identify T2D subgroups using multi-omics data from a population-based cohort, and to
compare the performance of different integrative multi-omics clustering methods. We applied three algorithms:
Similarity Network Fusion (SNF), Multi-Omics Factor Analysis (MOFA2) and iCluster+. For MOFA2, latent factors
were further clustered using k-means clustering. The clustering results were evaluated based on structure
(number and size of clusters), quality metrics such as the silhouette score, and computational performance,
including runtime. Additionally, we assessed the biological interpretability of the results by identifying
features associated with each cluster.
C-S.B.15: Systematic capture of human receptor-ligand interactions as Gene Ontology Causal Activity
Models
Track: Systems biology, multi-omics integration, modeling
-
Patrick Masson, Swiss Institute of Bioinformatics, Switzerland
- Lionel Breuza, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Cristina Casals-Casas, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Guilaine Argoud-Puy, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Nadine Gruaz, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Lucille Pourcel, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Sylvain Poux, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Livia Famiglietti, Swiss Institute of Bioinformatics. Swiss-Prot group., Switzerland
- Pascale Gaudet, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Alan Bridge, University of Southern California, United States
- Paul D. Thomas, University of Southern California, United States
-
The Uniprot Consortium, SIB Swiss Institute of Bioinformatics; European Bioinformatics Institute (EBI);
Protein Information Resource (PIR), Switzerland
Presentation Overview: Show
Cell-cell communication is crucial for the development of complex multicellular organisms. In systems biology,
computational biology, and biomedical research, elucidating cell-cell interactions requires a comprehensive
reference database of robust, physiologically relevant ligand-receptor interactions. However, existing
biological resources suffer from fragmented coverage, lack pathway and cellular context, and show poor
cross-resource consistency. To address this, we have begun a targeted biocuration effort using the Gene
Ontology Causal Activity Model (GO-CAM) framework. GO-CAMs integrate gene product activities into causally
connected, machine-readable models by using three complementary ontologies: Molecular Function (MF),
Biological Process (BP), and Cellular Component (CC) from the Gene Ontology (GO). Applied to receptor-ligand
biology, GO-CAM connects interactions to downstream signalling pathways, specifies the cell types in which
they occur, and captures microenvironmental context such as subcellular localisation and co-interacting
proteins. Of the more than 1,600 human receptors catalogued in UniProtKB/Swiss-Prot, those with known ligand
interactions will be systematically represented as GO-CAM models, using LLM-assisted literature screening to
identify and extract the relevant experimental evidence. To date, approximately 400 receptor-ligand pairs have
been curated and integrated into the GO-CAM framework.
C-S.B.16: Protein abundance inference at single-cell resolution
Track: Systems biology, multi-omics integration, modeling
-
Marc Zimmerli, University of Bern, Department for BioMedical Research, Urology Lab,
Switzerland
-
Panagiotis Chouvardas, Department of Urology, Inselspital, Bern University Hospital, University of Bern,
Bern, 3010, Switzerland, Switzerland
-
Albert Widjaja, University of Bern, Department for BioMedical Research, Urology Lab, Switzerland
-
Kristin Olsen, University of Bern, Department for BioMedical Research, Urology Lab, Switzerland
-
Beat Roth, Department of Urology, Inselspital, Bern University Hospital, University of Bern, Bern, 3010,
Switzerland, Switzerland
-
Marianna Kruithof-de Julio, Department of Urology, Inselspital, Bern University Hospital, University of
Bern, Bern, 3010, Switzerland, Switzerland
Presentation Overview: Show
Single-cell transcriptomics are scalable and widely used, but proteins remain closer to cellular function and
phenotype. This project aims to bridge that gap by inferring protein abundance directly from RNA data,
approximating protein-level insights whilst significantly reducing the associated experimental costs.
We present the second generation of our previously published method, named scLinear2, that translates
transcriptomic profiles into predicted protein signatures and evaluate its ability to capture biologically
meaningful variation. ScLinear2 performs at state-of-the-art accuracy levels across different datasets, while
employing an efficient and interpretable linear regression approach. We demonstrate that predicted protein
profiles can support downstream analysis tasks such as automated cell-type annotation by training tissue
specific models. Moreover, we utilize the interpretable nature of scLinear2 to characterize the most
informative genes for each protein across datasets, revealing interesting mechanistic insights. Of note,
ScLinear2 is developed as both R and Python package and is directly compatible with the most widely used
single-cell analysis workflows, allowing for its easier adaptation.
Taken together, we present scLinear2: an efficient and interpretable cross-modality prediction method. Our
results highlight the potential downstream applications of conversion between different omics modalities
whilst conserving much of the functional differences between cells. Future work includes the application of
scLinear2 in clinical data, which will allow the translational exploration of inferred protein levels.
C-S.B.17: Identifying Functional ROIs for Spatial Transcriptomics from H&E Images via Pathology
Foundation Models
Track: Systems biology, multi-omics integration, modeling
-
Kota Adachi, Department of Data-Driven Biology, Graduate School of Medicine, Nagoya University,
Japan
-
Chikara Mizukoshi, Medical Research Laboratory, Institute of Integrated Research, Institute of Science
Tokyo, Japan
-
Teppei Shimamura, Medical Research Laboratory, Institute of Integrated Research, Institute of Science
Tokyo, Japan
Presentation Overview: Show
Spatial omics enables the analysis of gene expression and cell states while preserving tissue architecture,
but high costs make comprehensive profiling across whole tissue sections difficult. In practice, regions of
interest (ROIs) for spatial assays are often selected by visual inspection of H&E slides. However, some
functionally relevant regions, including those associated with intratumor heterogeneity and the tumor immune
microenvironment, cannot be identified from morphology alone. Here, we propose a framework that integrates
pathology foundation models with supervised regression to identify and recommend biologically informative
functional ROIs for spatial transcriptomics from H&E images. Using paired H&E and spatial
transcriptomics data from HEST-1k, we define transcriptomics-derived functional scores and train models to
predict regional scores from H&E embeddings extracted by OmiCLIP. In our experiments, our method showed
higher correlation with transcriptomics-derived targets than a gene-query-based OmiCLIP zero-shot baseline.
Moreover, ROI selection based on predicted scores recovered regions close to the optimal regions defined from
ground-truth scores. These results suggest that the proposed framework identifies functionally informative
hotspots beyond morphological appearance and may improve ROI design for costly spatial omics experiments.
C-S.B.18: jsPCA: fast, scalable, and interpretable identification of spatial domains and variable genes
across multi-slice and multi-sample spatial transcriptomics data
Track: Systems biology, multi-omics integration, modeling
-
Ines Assali, Aix-Marseille Université, MMG, Inserm U1251, Turing Centre for Living systems,
Marseille, France, France
-
Paul Escande, Institut de Mathématiques de Toulouse; UMR 5219, Université de Toulouse, CNRS ; UPS, F-31062
Toulouse Cedex 9, France, France
-
Franck Picard, Université de Lyon, ENS de Lyon, Université Claude Bernard, CNRS UMR 5239, INSERM U1210,
Lyon, France, France
-
Paul Villoutreix, Aix-Marseille Université, MMG, Inserm U1251, Turing Centre for Living systems,
Marseille, France, France
Presentation Overview: Show
Spatial transcriptomics technologies record genome-wide measurements of gene expression with high spatial
resolution. These technologies generate large and high dimensional datasets requiring efficient automated
methods for their analysis. In this study we introduce joint spatial PCA (jsPCA), a novel, fast, scalable and
interpretable method for the automatic identification of spatial domains and variable genes in multi-slice and
multi-sample spatial transcriptomics data. jsPCA relies on a simple mathematical formulation of a spatial
covariance defined as the product of the gene expression covariance with the spatial autocorrelation. The
principal components of this spatial covariance yield a biologically meaningful low-dimensional
representation. From this representation, we can derive spatial domains by simple clustering. In addition,
spatially variable genes can be identified directly from the principal components coefficients. Moreover, this
approach enables the joint representation of multiple slices and samples, a frequent experimental setting.
This joint representation is obtained without spatial alignment by computing common principal components via
joint diagonalization of the set of spatial covariance matrices obtained for each slice. By leveraging data
sparsity and non-convex optimization on manifold, jsPCA leads to computing time in the order of seconds to
minutes, substantially outperforming state-of-the-art approaches. We benchmarked jsPCA on the Visium 10x
dataset of human dorsolateral prefrontal cortex and the Stereo-seq MOSTA dataset of mouse embryonic
development against 10 state-of-the-art methods. Our approach demonstrated excellent performances, comparable
or better than state-of-the-art methods, such as SpatialPCA, BASS, GraphPCA or Stagate, while being much
faster, interpretable, and scalable to very large datasets.
C-S.B.19: Deciphering context-specific protein interactomes to uncover protein functions and therapeutic
targets
Track: Systems biology, multi-omics integration, modeling
-
Cédric Vincent-Cuaz, University of Bern, Department of Biomedical research, Switzerland
- Alois Thomas, University of Bern, Department of Biomedical research, Switzerland
- Lisa Fournier, University of Bern, Department of Biomedical research, Switzerland
- Vincent Jung, EPFL, LTS4, Switzerland
- Pascal Frossard, EPFL, LTS4, Switzerland
-
Raphaëlle Luisier, University of Bern, Department of Biomedical research, Switzerland
Presentation Overview: Show
Understanding molecular function requires accurate modeling of interactomes within specific cellular contexts,
yet experimental observations of context-dependent interactions remain sparse and difficult to generate at
scale. Computational approaches are therefore essential for inferring missing interactions and uncovering the
principles that shape context-specific protein–protein interactions (PPIs). Here, we introduce ProtScape, a
novel multi-scale framework for context-specific representation learning across proteins, cells, and tissues.
We first construct contextwise PPI networks for 207 cell types, using a biologically grounded gene-selection
strategy that integrates transcriptomic profiles from diverse human tissues, neuronal populations, and
specific cellular states in amyotrophic lateral sclerosis (ALS). ProtScape then learns representations across
contexts and scales, combining context-free protein foundation model embeddings, a novel hierarchical graph
neural architecture tailored to the resulting heterophilic graphs, and a masked graph autoencoding learning
strategy. To reliably evaluate the biological information captured by our representations despite the absence
of context-specific protein annotations, we introduce a weakly supervised multiple-instance learning framework
for protein contexts, rather than assuming that protein-level labels are uniformly valid across cellular
contexts.
ProtScape achieves substantial improvements in PPI prediction, while revealing context-dependent interaction
rewiring, including in RNA-binding proteins implicated in ALS. Moreover, our curated downstream evaluation
protocol shows that ProtScape captures strong biological signals related to protein complexes and therapeutic
targets, improving the prioritization of disease-relevant candidates as validated in Parkinson’s, while
enabling robust identification of relevant cellular contexts. while enabling robust identification of relevant
cellular contexts. Together, these results position ProtScape as a scalable foundation for context-aware
interactome modeling. Future extensions toward multimodal integration, including RNA-protein interactions,
will further enable the study of regulatory mechanisms underlying complex diseases such as neurodegeneration.
C-S.B.20: Deciphering Disease Transcriptional Logic by Revealing TF Composite Modules
Track: Systems biology, multi-omics integration, modeling
-
Aleksandr Kalmykov, geneXplain GmbH, Germany
- Alexander Kel, geneXplain GmbH, Germany
- Jochen Prehn, Royal College of Surgeons in Ireland (RCSI), Ireland
- Mohannad Dabbour, Royal College of Surgeons in Ireland (RCSI), Ireland
Presentation Overview: Show
Background: Predicting transcription factor (TF) interactions in biological systems presents a major
computational challenge critical for deciphering regulatory logic. This is particularly important in contexts
like aggressive brain tumors (e.g., Glioblastoma, GBM), where tumor subtype-specific regulatory networks
orchestrate therapeutic resistance. Advanced computational modeling is needed to accurately predict
cooperative TF-DNA binding in specific cellular contexts.This study details an enhanced computational tool to
model cooperative binding of multiple TFs to target regulatory regions (promoters and enhancers).
Materials and methods: We significantly enhanced the Composite Module Analyst (CMA) algorithm, a method
originally designed to predict TF binding compositions to promoter regions. The primary computational
enhancement involves a redesigned fitness function incorporating gene expression levels (RNA-seq) and modeling
factor-factor interaction dynamics. The new CMA was trained on a specific GBM dataset to identify tumor
subtype-specific TF composite modules, which serves as a proof-of-concept application for its ability to
correlate computational predictions with RNA-seq data.
Results: Our enhanced CMA demonstrated a marked ability for identifying cooperating TFs binding to their
composite sites. Application of this robust model to the GBM subtype data successfully revealed distinct sets
of key TFs (e.g., SP100, TP53, MITF, MeCP2) that are active in different tumor microenvironment contexts.
Conclusions: This work presents an advanced computational tool for modeling complex transcription regulatory
landscapes. Our model provides a novel, mechanistic approach to identifying key cooperating TFs in
subtype-specific regulatory networks . This advancement is applicable for dissecting the transcriptional logic
of any disease context and highlights the identified cooperating TFs.
C-S.B.21: KidsCan Analytical Pipelines - A Multi-Omics Framework for Pediatric Precision Oncology in
Switzerland
Track: Systems biology, multi-omics integration, modeling
-
Babih Velazquez De Burnay, Universitäts-Kinderspital Zürich, Switzerland
- Fabio Steffen, Universitäts-Kinderspital Zürich, Switzerland
- Raphael Johannes Morscher, Universitäts-Kinderspital Zürich, Switzerland
- Jean-Pierre Bourquin, Universitäts-Kinderspital Zürich, Switzerland
- Ana Sofia Guerreiro Stücklin, Universitäts-Kinderspital Zürich, Switzerland
- Linda Grob, Universitäts-Kinderspital Zürich, Switzerland
Presentation Overview: Show
Despite advances in treatments, cancer remains the leading cause of mortality among children in Switzerland.
The implementation of precision oncology for pediatric malignancies has been hindered by the lack of
standardized analytical frameworks to process high-throughput multi-omics data within actionable timeframes.
Here we present the tumor profiling and analytical infrastructure of KidsCan, a national pediatric precision
oncology initiative.
Within the KidsCan-01: Swiss Personalized Oncology for Children flagship study, we develop reproducible
bioinformatics pipelines to profile pediatric tumors and provide decision support. Our framework integrates a
whole-genome sequencing pipeline to prioritize oncogenic single nucleotide variants, insertions/deletions,
structural variants, and copy number alterations. This is paired with an bulk RNA-sequencing pipeline for gene
expression quantification, tumor-type classification, and fusion detection.
To ensure reproducibility required for clinical decision support, we implement version control and
containerize all pipeline modules using Docker, guaranteeing consistent execution across systems. Within this
robust framework, we implemented a unified quality-control pipeline that evaluates sample integrity and
analytical performance across sequencing modalities. We validated and refined our pipelines using a series of
benchmarking experiments with multi-omic data and tumor profiling reports processed by the INFORM initiative.
Furthermore, to facilitate translational application, we engineer a reporting pipeline that aggregates
multi-omics findings for the pediatric molecular tumor board, streamlining interpretation of molecular
vulnerabilities.
By harmonizing molecular outputs with structured clinical metadata via FHIR standards, KidsCan not only
optimizes clinical decision tools for immediate patient care but also pioneers a scalable, interoperable data
ecosystem to propel future pediatric cancer research in Switzerland.
C-S.B.22: Multimodal Deep Learning for Predicting of RNA Subcellular Localization
Track: Systems biology, multi-omics integration, modeling
-
Hayato Ishii, Institute of Science Tokyo, Japan
- Chikara Mizukoshi, Institute of Science Tokyo, Japan
- Haruhiko Morita, Institute of Science Tokyo, Japan
- Teppei Shimamura, Institute of Science Tokyo, Japan
Presentation Overview: Show
RNA subcellular localization is a fundamental biological phenomenon that underlies processes such as cell
polarity formation and differentiation. Importantly, localization patterns change in response to intracellular
and extracellular environments as well as cellular states, and their disruption is known to contribute to
neurological disorders and cancer. Therefore, elucidating the mechanisms of RNA localization is an important
challenge for understanding complex biological regulatory networks and overcoming disease. To date, methods
for predicting RNA localization have primarily relied on RNA sequence information. However, these approaches
have mainly focused on predicting population-averaged (bulk-level) RNA localization across multiple cells,
making single-cell-level prediction difficult. As a result, they have been unable to fully capture
cell-state-dependent variations in RNA localization. In this study, we developed a deep learning method to
predict RNA localization at the single-cell level by integrating RNA sequence information, cellular imaging
data from Cell Painting, a multiplexed morphological profiling assay, and spatial transcriptomics data.
Compared with a baseline model that outputs the average localization proportion in the training data, our
method predicted RNA localization proportions more accurately. In addition, the model predicted with high
correlation the localization proportions of individual cells averaged across the gene dimension, suggesting
that it can capture cell-specific average RNA localization tendencies reflective of cellular state. This study
demonstrates the potential of integrating RNA sequence and cellular state information to enable
single-cell-level prediction of RNA localization and contribute to a better understanding of
cell-state-dependent mechanisms of RNA localization control.
C-S.B.23: A Deep Generative Framework for Joint Modeling of Single-Cell Isoform Expression and Full-Length
Transcript Sequences
Track: Systems biology, multi-omics integration, modeling
-
Taichi Oso, Department of Computational and Systems Biology, Medical Research Laboratory, Institute
of Science Tokyo, Japan
-
Chikara Mizukoshi, Department of Computational and Systems Biology, Medical Research Laboratory, Institute
of Science Tokyo, Japan
-
Teppei Shimamura, Department of Computational and Systems Biology, Medical Research Laboratory, Institute
of Science Tokyo, Japan
Presentation Overview: Show
Single-cell RNA sequencing (scRNA-seq) has greatly advanced our understanding of cellular heterogeneity, and
recent advances in long-read scRNA-seq have made it possible to resolve isoform expression at single-cell
resolution. In parallel, deep generative models have provided a powerful framework for learning latent
cell-state representations from high-dimensional single-cell data. Because transcript isoforms can alter
coding potential, RNA stability, subcellular localization, and regulatory interactions, isoform diversity is a
key layer of gene regulation underlying cell identity, differentiation, and disease. However, existing methods
still lack a unified framework for integrating isoform-level expression variation and sequence differences,
making it difficult to determine how sequence-defined isoform usage contributes to cell-state-specific
regulation and biological function. Here, we present a deep generative framework that jointly learns
cell-specific isoform expression and sequence embeddings of full-length transcripts derived from a nucleotide
language model. By linking latent cell-state representations with isoform sequence embeddings, our method
enables integrated analysis of sequence-associated isoform variation across cells. Applied to long-read
scRNA-seq data, our framework learned biologically meaningful cell-state representations while enabling
estimation of isoform heterogeneity linked to sequence features. Furthermore, gradient-based attribution
analysis and in silico mutagenesis showed that the model could identify nucleotide sequence regions important
for predicting isoform variation. These results provide a computational foundation for linking
isoform-defining sequence features to cell-state-dependent regulation at single-cell resolution.
C-S.B.24: A spatially informed self-supervised framework reveals the morphological landscape of astrocyte
substates and its dissociation from transcriptional variation in ALS
Track: Systems biology, multi-omics integration, modeling
-
Elisa Messori, Swiss Institute of Bioinformatics; EPFL; University of Bern; Idiap,
Switzerland
- Doaa Taha, The Francis Crick Institute; University College London, United Kingdom
-
Lisa Fournier, Swiss Institute of Bioinformatics; University of Bern; Hôpitaux Universitaires de Genève,
Switzerland
- Anna Foix Romero, EMBL-EBI, Hinxton, UK, United Kingdom
- Virginie Uhlman, BioVisionCenter Universität Zürich, Switzerland, Switzerland
- Pascal Frossard, EPFL, Switzerland
- Cédric Vincent-Cuaz, University of Bern, Switzerland
-
Rickie Patani, The Francis Crick Institute; University College London; National University of Singapore,
Singapore
-
Raphaëlle Luisier, Swiss Institute of Bioinformatics; University of Bern, Switzerland
Presentation Overview: Show
Astrocyte reactivity is a central contributor to neurodegeneration, yet the morphological heterogeneity of
reactive states and their relationship to transcriptional programs remain poorly understood. A key challenge
is that biological conditions rarely correspond to a single, homogeneous phenotype, a complexity obscured in
aggregate population analysis.
Here, we address this by generating a multimodal dataset of human iPSC-derived astrocytes from ALS patients
and controls, profiled by high-content fluorescence imaging and bulk RNA sequencing under basal and controlled
pro-inflammatory conditions. We develop SI-SimCLR, a spatially informed contrastive learning framework that
learns biologically meaningful representations from microscopy images, without segmentation or predefined
labels. SI-SimCLR outperforms standard baselines in capturing disease- and inflammation-associated
morphological variation across experimental batches.
Unsupervised analysis of SI-SimCLR embeddings revealed a structured morphological landscape composed of twelve
distinct substates. Using Optimal Transport to construct a morphological transition graph, we found VCP-mutant
astrocytes occupy a constrained genotype-specific region under basal conditions, while high-dose inflammatory
stimulation partially shifts these states toward control-like configurations. Substate-resolved analysis
further highlighted untreated VCP-mutant astrocytes as cell-autonomous morphological states overlapping with
inflammation-induced reactive phenotypes.
Integration with bulk RNA sequencing revealed a striking dissociation: while inflammatory stimulation
dominates transcriptional variation, the ALS mutation primarily drives morphological organization. This
indicates morphological and transcriptional responses to disease represent partially independent axes of
astrocyte dysfunction.
Together, these results establish a scalable, annotation-free framework for high-resolution characterization
of phenotypic heterogeneity, providing a principled foundation for substate-resolved analysis of astrocyte
biology in neurodegeneration and beyond.
C-S.B.25: To Translate or Not: Multistable Decision-Making in Translation Initiation under Normal Conditions
and Integrated Stress Response
Track: Systems biology, multi-omics integration, modeling
-
Harika G L, Center for Computational Biology, Department of Computational Biology, IIIT-Delhi, INDIA,
India
-
Sriram K, Department of Computational Biology, Center for Computational Biology and Centre for AI,
IIIT-Delhi, New Delhi, India, India
Presentation Overview: Show
Translation initiation is an essential regulatory stage of protein synthesis, preceding elongation. This stage
operates through two mechanisms, primary and secondary, involving complex interactions among multiple
eukaryotic initiation factors (eIFs). In the primary mechanism, eIF2-GDP, eIF2B, and eIF5 function analogously
to clutch, brake, and accelerator, regulating GDP-GTP exchange and the formation of the active eIF2-GTP
complex. Further, the binding of Met-tRNA forms the ternary complex (TC), which serves as the clutch in the
secondary mechanism. Understanding the operation of this clutch-brake-accelerator molecular system under the
integrated stress response (ISR) is essential for studying the translational control. Despite extensive
biochemical insights, it remains unclear how these interactions generate decision-making dynamics under both
normal and stress conditions. Here, we develop a mechanistic mathematical model based on experimental
observation to investigate how transitions between translation initiation and termination occur at the
initiation stage. The model incorporates phosphorylation-dephosphorylation (PdP) reactions to model ISR as a
fail-safe mechanism. Because the reaction network contains several unknown kinetic parameters, we use Chemical
Reaction Network Theory (CRNT) to perform structural analysis and identify hidden positive feedback loops
embedded in the initiation mechanism. Bifurcation analysis reveals ultrasensitivity and bistability under
normal conditions. Under ISR, the system exhibits both bistability and tristability for physiologically
relevant parameter regimes. We associate bistability with switching between initiation and termination, and
tristability with recovery and attenuation during stress. Together, these results demonstrate that translation
initiation operates as a threshold-driven switching system rather than a graded process, enabling rapid,
robust decisions across varying cellular conditions.
C-S.B.26: Structured dimensionality reduction for refining coding-agent-authored single-cell embedding
models
Track: Systems biology, multi-omics integration, modeling
-
Niklas Brunn, Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical
Center – University of Freiburg, Germany
-
Sonia Maria Krissmer, Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical
Center – University of Freiburg, Germany
-
Maximilian Frosch, Institute of Neuropathology, Faculty of Medicine, University of Freiburg, Germany
-
Marco Prinz, Institute of Neuropathology, Faculty of Medicine, University of Freiburg, Germany
-
Harald Binder, Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical Center
– University of Freiburg, Germany
Presentation Overview: Show
Coding agents can author interpretable single-cell embedding models, here called blueprints, directly from the
scientific literature, without ever training a computational model or being exposed to gene expression data.
Such a blueprint is auditable by construction, but being purely prior-driven it captures only what the
literature already describes and cannot, on its own, correct mismatches with a given dataset or surface
structure the literature has not named.
Here, we let a coding agent refine the blueprint against data with structured feedback from a boosting
autoencoder that performs structured dimensionality reduction with implicit feature selection. Its latent
dimensions are interpretable, each linked to a small gene module, matching the gene-selected, named modules of
the agent-authored blueprint. Trained on the data with the blueprint's axes as a prior, it learns additional
axes that capture complementary structure the blueprint does not explain. Because the newly learned axes share
the blueprint's gene-module vocabulary, contrasting them with the blueprint-anchored axes yields quantitative
metrics and qualitative descriptors of cell-group gene programs and latent-space topology.
We first show that, across multiple datasets, agent-authored blueprints yield embeddings whose named axes
faithfully discriminate the cell types they name and reach quality competitive with conventional,
foundation-model, and program-informed baselines, while remaining batch-robust by construction. Building on
this, we demonstrate the refinement on mouse cortex data, where the blueprint misses cell subtypes. The
boosting autoencoder surfaces these as structure beyond the literature prior, and a coding agent folds them
back into the blueprint as new axes that recover the missed subtypes.
C-S.B.27: Specifying single-cell data simulation tasks for LLM agents to contrast literature knowledge with
real data
Track: Systems biology, multi-omics integration, modeling
-
Sonia Maria Krißmer, Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and
Medical Center – University of Freiburg, Germany
-
Niklas Brunn, Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical Center
– University of Freiburg, Germany
-
Maximilian Frosch, Institute of Neuropathology, Faculty of Medicine, University of Freiburg, Germany
-
Marco Prinz, Institute of Neuropathology, Faculty of Medicine, University of Freiburg, Germany
-
Harald Binder, Institute of Medical Biometry and Statistics (IMBI), Faculty of Medicine and Medical Center
- University of Freiburg, Germany
Presentation Overview: Show
LLMs encode large amounts of gene expression literature knowledge, offering a promising resource for
downstream inference and the integration of single-cell RNA sequencing data. However, leveraging this
knowledge remains challenging due to instability and limited interpretability in prompt-based LLM approaches.
We suggest that providing a simulation task for synthetic single-cell data forces an agentic LLM system to
encode structural relationships from the literature into quantitative models, thereby improving stability and
interpretability. These models enable generation of synthetic data that can be contrasted with real-data
patterns.
Specifically, our approach builds on an agent interacting with a modular simulation script in a two-step
process. The simulation script contains a cell blueprint, defining the action space of the agent, and a
statistical model generating synthetic cells represented as unranked lists of highly expressed genes. The
agent is first tasked with updating the blueprint to reflect biological knowledge from retrieved literature.
Next, the agent updates the blueprint based on patterns from real data. The resulting blueprints and their
differences are directly interpretable, exposing modules, gene-level probabilities, temporal trends, and
knockout effects.
We demonstrate the framework on a simulation task for microglia in the developing mouse brain with a HexB
knockout. This simulation task allows to extract concise descriptions of literature knowledge. Assessing the
variability of the derived blueprints across multiple agent runs highlights how uncertainty can be decreased
by restricting the action space of the agent. Moreover, the variability quantifies how real data contains
additional or diverging patterns compared to the literature.
C-S.B.28: From trial-and-error to precision therapy: A multi-omics knowledge-graph framework for primary
immune regulatory disorders
Track: Systems biology, multi-omics integration, modeling
-
Chaimae El Houjjaji, School of Life Sciences, Ecole Polytechnique Fédérale de Lausanne, Lausanne,
Switzerland, Switzerland
-
Ali Saadat, School of Life Sciences, Ecole Polytechnique Fédérale de Lausanne, Lausanne, Switzerland,
Switzerland
-
Mariam Ait Oumelloul, School of Life Sciences, Ecole Polytechnique Fédérale de Lausanne, Lausanne,
Switzerland, Switzerland
-
Jacques Fellay, School of Life Sciences, Ecole Polytechnique Fédérale de Lausanne, Lausanne, Switzerland,
Switzerland
Presentation Overview: Show
Primary Immune Regulatory Disorders (PIRDs), which includes Common Variable Immunodeficiency (CVID) and
Combined Immunodeficiency (CID), affect about 77,500 patients across Europe. They are characterized by immune
dysregulation. Although 25% of cases have a monogenic cause, the majority lack a molecular diagnosis, leaving
clinicians to rely on empirical, trial-and-error treatment selection that prolongs disease activity and
increases healthcare burden.
We propose a multi-omics integration framework that combines genomics, transcriptomics, proteomics,
epigenomics, microbiome, and clinical data for patient stratification and therapy prioritization. This
approach uses trained models, like HyenaDNA and RNABERT to encode each type of data. Then it combines them
through a fusion module to produce a patient-level embedding and contextualizes this embedding with a
biomedical knowledge graph to output ranked drug candidates.
Classical multi-omics integration methods such as MOFA or DIABLO identify patterns in the data but are not
designed to use external biological knowledge, which is crucial when patient cohorts are small. Coupling
modality-specific embeddings with a knowledge graph helps compensating for this scarcity by injecting curated
priors on gene-gene, gene-drug, and pathway relationships while also capturing interactions across immune
layers that no single omic can reveal on its own.
By coupling multi-omics patient embeddings with a biomedical knowledge graph, this design aims to leverage
scarce rare-disease data using structured biological priors. The approach is intended to enable precision
medicine in PIRDs by generating interpretable, patient-specific therapy rankings.
C-S.B.29: Phenotype-driven parallel embedding for microbiome multi-omic data integration
Track: Systems biology, multi-omics integration, modeling
-
Tal Bamberger, Gray Faculty of Medical & Health Sciences, Tel Aviv University, Tel Aviv, Israel,
Israel
- Dap Consortium, University of washington, United States
-
Elhanan Borenstein, Blavatnik School of Computer Science and AI, Tel Aviv University, Tel Aviv, Israel,
Israel
Presentation Overview: Show
The human microbiome is a key determinant of health and disease, but most reported associations remain
descriptive and lack a system-level view. Multi-omic profiling can provide such a view, yet integration is
challenging because omics differ in scale, structure, and biological meaning. Existing embedding-based
approaches often either have limited predictive power or collapse all omics into a single latent space, losing
omic-specific information.
We introduce PAPRICA (Phenotype-Aware Parallel Representation for Integrative omiC Analysis), a parallel
encoder-decoder framework that embeds each omic in its own latent space while jointly modeling cross-omic and
phenotype-associated relationships. PAPRICA is trained to reconstruct each omic, align samples across
omic-specific latent spaces, and capture variation related to continuous phenotypes such as fecal calprotectin
in IBD.
Across three multi-omic microbiome datasets, PAPRICA outperformed alternative integration models in both
cross-omic prediction and phenotype prediction. These results suggest that PAPRICA effectively preserves
omic-specific signals while capturing shared and phenotype-relevant structure, providing a flexible framework
for microbiome multi-omic integration.
C-S.B.30: MUSE enables cross-species multi-omics integration that incorporates transcriptional regulatory
modules
Track: Systems biology, multi-omics integration, modeling
-
Fuka Nakae, Department of Computational and Systems Biology, Medical Research Laboratory, Institute
of Science Tokyo, Japan
-
Shintaro Yuki, Department of Computational and Systems Biology, Medical Research Laboratory, Institute of
Science Tokyo, Japan
-
Zhenan Liu, Department of Computational and Systems Biology, Medical Research Laboratory, Institute of
Science Tokyo, Japan
-
Chikara Mizukoshi, Department of Computational and Systems Biology, Medical Research Laboratory, Institute
of Science Tokyo, Japan
-
Shuto Hayashi, Department of Computational and Systems Biology, Medical Research Laboratory, Institute of
Science Tokyo, Japan
-
Teppei Shimamura, Department of Computational and Systems Biology, Medical Research Laboratory, Institute
of Science Tokyo, Japan
-
Hiroshi Yadohisa, Department of Culture and Information Science, Doshisha University, Japan
Presentation Overview: Show
Recent advances in cross-species RNA-seq analysis have enabled the identification of conserved and
species-specific features, contributing to both evolutionary biology and biomedical research. However,
alignments based solely on RNA expression are limited in their ability to capture regulatory logic. With the
emergence of multi-omics data, particularly chromatin accessibility measured by ATAC-seq, it has become
possible to investigate upstream regulatory states and transcription factor activity. Integrating RNA and ATAC
enables more direct characterization of transcriptional regulatory modules, which are often more conserved
than gene expression patterns across species. Despite these advantages, multi-omics integration faces several
challenges, including technical heterogeneity and the limited availability of comparable datasets. To address
these issues, various methods have been proposed to enable integration across modalities that are not measured
in the same cells (diagonal integration). However, these approaches do not adequately model species-specific
differences or evolutionary changes in regulatory relationships. In this study, we propose a novel framework,
Multi-omics Unified embedding across Species (MUSE), which constructs a cross-species graph by introducing
inter-species edges based on orthologous relationships and protein sequence similarity, while preserving the
regulatory graph structure within each species. Based on this integrated graph, MUSE aligns omics measurements
across species, enabling biologically meaningful cross-species comparisons at both the omics and regulatory
module levels.
C-S.B.31: MultiOmics Centre: a driver of innovation in food, agriculture and health
Track: Systems biology, multi-omics integration, modeling
-
Emma Busarello, Eurac Research, Italy
Presentation Overview: Show
Multi-omics technologies are essential for advancing systems-level understanding of biological processes
through the integration of genomics, transcriptomics, proteomics, and metabolomics data. However, their
adoption in research and industry remains limited by insufficient computational infrastructure, lack of
standardized workflows, and limited bioinformatics expertise. In addition, the growing scale and complexity of
multi-omics datasets require robust, reproducible, and scalable analytical solutions.
The MultiOmics Centre (MOC), a joint initiative of Laimburg Research Centre and Eurac Research, addresses
these challenges by combining advanced experimental and computational capabilities. The centre provides the
laboratory infrastructure and trained personnel for omics data generation, alongside a dedicated computational
environment. This includes high-performance computing, data management systems, and standardized pipelines
that support reproducible and FAIR-compliant multi-omics analyses.
To support genomic analyses, I developed a web-based application that streamlines and standardizes the
execution of the in-house Nextflow pipeline for genome-wide association studies (GWAS). The application
provides a user-friendly interface to specify input data and analysis settings, automatically handling
configuration and pipeline execution. This reduces manual intervention, minimizes errors, and ensures
consistent, reproducible analyses across studies. The underlying pipeline, nf-gwas, is designed for
biobank-scale GWAS analyses and it automates pre- and post-processing steps, integrates regression modelling
via the REGENIE package, and supports single-variant, gene-based, and interaction testing. It also provides
comprehensive reporting, enabling the exploration of results across thousands of phenotypes.
By streamlining the use of the nf-gwas pipeline, the web application enables fast, scalable, and reproducible
GWAS analyses, contributing to the MOC's efforts to accelerate the adoption of multi-omics approaches.
C-S.B.32: OXidative Stress PREDictor: A Supervised Learning Approach for Annotating Cellular Oxidative Stress
States in Inflammatory Cells
Track: Systems biology, multi-omics integration, modeling
-
Po-Yuan Chen, Nationa Yang Mingl Chiao Tung University / Academia Sinica, Taiwan
- Tai-Ming Ko, Nationa Yang Mingl Chiao Tung University / Academia Sinica, Taiwan
Presentation Overview: Show
Oxidative stress, characterized by an imbalance between reactive oxygen species (ROS) and antioxidants, plays
a pivotal role in inflammatory responses associated with both chronic diseases and acute injuries. In this
study, OXidative Stress PREDictor (OxSpred), a supervised learning model tailored to accurately annotate the
oxidative stress state of innate immune cells at the single-cell level, is introduced. Compared to the
traditional gene-set-variation-analysis-based enrichment method, OxSpred demonstrates superior accuracy with
an area under the receiver operating characteristic curve of 0.89 and offers interpretable embeddings with
significant biological relevance. Using the predicted ROS states, precise elucidation and interpretation of
the roles of novel innate immune cell subtypes can be achieved. Overall, OxSpred enhances the utility of
single-cell transcriptomic datasets by providing a robust in silico method for determining intracellular
oxidative stress states, thereby enriching the understanding of innate immune cell functions during
inflammation.
C-S.B.33: Deciphering signal propagation in the context of an activating cyclin dependent kinase 4
mutation
Track: Systems biology, multi-omics integration, modeling
-
Abel Szkalisity, Department of Anatomy and Stem Cells and Metabolism Research Program, Faculty of
Medicine, University of Helsinki, Finland
-
Maarit Holtta, Department of Anatomy and Stem Cells and Metabolism Research Program, Faculty of Medicine,
University of Helsinki, Finland
-
Kari Moisio, Institute of Biotechnology and Helsinki Institute of Life Science, University of Helsinki,
Finland
-
Liisa Sarkio, Department of Anatomy and Stem Cells and Metabolism Research Program, Faculty of Medicine,
University of Helsinki, Finland
-
Ville Hietakangas, Institute of Biotechnology and Helsinki Institute of Life Science, University of
Helsinki, Finland
-
Elina Ikonen, Department of Anatomy and Stem Cells and Metabolism Research Program, Faculty of Medicine,
University of Helsinki, Finland
Presentation Overview: Show
We have found an activating mutation (M207T) in the early cell-cycle regulator cyclin dependent kinase 4
(CDK4) in a patient with severely disturbed adipose tissue distribution and metabolic syndrome. To understand
the underlying mechanisms, we overexpressed the mutant CDK4 or pharmacologically inhibited the kinase, and
conducted proteomic, phosphoproteomic and metabolic analyses during the first two days of adipogenic
differentiation of human mesenchymal stem cells. Together, our data covering over 8000 proteins, 20000
phosphosites and 400 metabolites suggest that the mutant CDK4 induces major changes in the flux of glucose,
including altered use of the TCA cycle. Here, we present two distinct computational approaches to dissect how
CDK4 activity controls signal propagation during early adipogenesis. First, we portray conventional
multi-omics integration methods, including matrix factorization and constrained based reconstruction to derive
the downstream effectors of active CDK4. Second, we identify possible signal propagation pathways that explain
how the observed metabolic changes are controlled by CDK4, by utilizing the SIGnaling Network Open Resource
(SIGNOR) database to search for potential shortest paths starting from an activated CDK4 and leading to the
observed effectors. We follow the binary activation/deactivation rules of the database to identify the
expected activation of proteins along a path and match these expectations to our observed data at both protein
and phosphorylation levels. Our approach shows the potential of integrating longitudinal multi-omics data with
curated databases to reveal dynamic, intracellular signal propagation pathways that control the early steps of
cellular differentiation.
C-S.B.34: Sensitivity and correlation analysis of Calvin-Benson cycle models identifies key parameters and
reduced effective dimensionality
Track: Systems biology, multi-omics integration, modeling
-
Salma Tariq, University of Potsdam, Germany
- Anika Küken, University of Potsdam, Germany
Presentation Overview: Show
Kinetic models of the Calvin-Benson cycle (CBC) are used to study photosynthesis, but their behavior depends
strongly on parameter values that are often uncertain. In this study, we performed a sensitivity analysis on
several published CBC models to assess how changes in parameters influence steady-state metabolite levels.
Each kinetic parameter and metabolite concentration was perturbed individually by ±5% to ±25%, and the
resulting effects on the system were quantified. In addition, selected pairwise perturbations were analyzed to
examine interaction effects. Across all models, we observed that metabolic control is not evenly distributed.
Instead, a relatively small number of parameters consistently exert a stronger influence on system behavior,
while most parameters affect only one or a few metabolites. Key processes affecting model simulations include
phosphate balance, adenosine triphosphate (ATP) synthesis, and specific Calvin cycle reactions such as
3-phosphoglycerate (PGA) reduction, sedoheptulose-1,7-bisphosphatase (SBPase) activity, and
ribulose-5-phosphate kinase (Ru5P kinase), although the extent of this control differs between models.
Overall, these results suggest that only a subset of parameters is critical for determining system behavior.
Comparison of correlation structures across models reveals clear model-specific differences, with certain
parameter pairs showing positive correlations in one model but weak, absent, or negative correlations in
others. Correlation analysis further shows that parameters are not independent but organized into coordinated
groups, reducing the effective dimensionality of the parameter space. Identifying these key parameters and
dependencies can help simplify models and improve parameter estimation when experimental data are limited.
C-S.B.35: Comprehensive reannotation of human enzymes and transporters in Swiss-Prot using Rhea
Track: Systems biology, multi-omics integration, modeling
- Nadine Gruaz, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
The Uniprot Consortium, SIB Swiss Institute of Bioinformatics; European Bioinformatics Institute (EBI);
Protein Information Resource (PIR), Switzerland
- Paul D. Thomas, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Alan Bridge, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Shyamala Sundaram, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Lucille Pourcel, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Nevila Hyka-Nouspikel, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Anne Morgat, SIB Swiss Institute of Bioinformatics, Switzerland
- Patrick Masson, Swiss Institute of Bioinformatics, Switzerland
-
Cristina Casals-Casas, Swiss Institute of Bioinformatics. Swiss-Prot group., Switzerland
- Arnaud Gos, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Livia Famiglietti, Swiss Institute of Bioinformatics. Swiss-Prot group., Switzerland
-
Elisabeth Coudert, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Lionel Breuza, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Kristian B. Axelsen, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
-
Guilaine Argoud-Puy, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
- Lucila Aimo, Swiss-Prot Group, SIB Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
The UniProt Knowledgebase (UniProtKB) is a central reference resource for protein sequences and functional
annotation supporting multiomics analyses. To improve integration of protein and small molecule knowledge, we
completed a first-pass systematic reannotation of all human enzymes and transporters in UniProtKB/Swiss-Prot
using Rhea, a curated resource of biochemical reactions structured on the ChEBI ontology.
All human enzymes and transporters were systematically reassessed and updated, resulting in reaction-based
annotations for more than 4,100 human proteins. These annotations include balanced reactions where all
molecules are linked to the ChEBI ontology, and atoms are mapped between reactants and products. This first
phase establishes the foundation for comprehensive reaction-based annotation of enzyme and transporter
functions and will support expanded analyses of metabolic processes and multiomics data integration in future
releases.
C-S.B.36: Phototransduction in the visual system as an example of biological system-based curation of protein
functions in Swiss-Prot
Track: Systems biology, multi-omics integration, modeling
-
Lucille Pourcel, SIB, Switzerland
- The Uniprot Consortium, SIB, Switzerland
Presentation Overview: Show
* The UniProt Knowledgebase (UniProtKB; www.uniprot.org) is a central resource for protein sequences and
functional annotations. However, protein annotation remains largely protein-centric, often limiting integrated
and system-level pathway representations required for mechanistic understanding of biological processes.
* Here, we present a pathway-centric curation approach for the complete functional annotation of biological
processes, using the visual phototransduction pathway as a model.
* Using a combination of UniProtKB queries, AI-based search strategies, and the identification of proteins
associated with vision disorders, we identified 71 proteins mapped to the visual pathway network, spanning
phototransduction in photoreceptors to retinal processing. Among these, 54 proteins are associated with vision
disease variants. We are currently annotating each protein in the phototransduction signaling pathway with
up-to-date protein functions, including standardized biochemical reactions (Rhea) within UniProtKB, and
disease associated variants. The Gene Ontology Causal Activity Models (GO-CAM) framework is used to formally
represent causal relationships between molecular activities, enabling machine-readable and computable pathway
models.
* This work provides a standardized, computable representation of the phototransduction pathway and
establishes a scalable approach for complete pathway annotation, supporting integrative analysis of molecular
mechanisms and disease phenotypes.
C-S.B.37: Decoding spatial niche architecture of dedifferentiation in aggressive thyroid cancer
Track: Systems biology, multi-omics integration, modeling
-
Han Sai Lee, Department of Molecular Medicine and Biopharmaceutical Sciences, Seoul National
University, South Korea
-
Hongyoon Choi, Department of Nuclear Medicine, Seoul National University College of Medicine, South Korea
-
Young Shin Song, Department of Internal Medicine, Seoul Metropolitan Government Seoul National University
Boramae Medical Center, South Korea
-
Young Joo Park, Department of Internal Medicine, Seoul National University College of Medicine, South
Korea
Presentation Overview: Show
Purpose: Anaplastic thyroid cancer (ATC) is among the deadliest malignancies, yet the spatial co-evolution of
dedifferentiation and microenvironment remodeling during progression remains poorly understood.
Methods: We integrated bulk RNA-seq (n=1,634), scRNA-seq (n=106), and spatial transcriptomics (Xenium: 82
cores, 2.35M cells; Visium: 48 slides, ~94k spots). A 25D embedding of cell-type fractions and functional
signatures was resolved by spatial Leiden micro-communities then KMeans (k=7), transferred to Visium via
Elastic Net, and linked to genotype, IHC, and survival with patient-grouped cross-validation.
Results: Seven niches emerged — Perivascular, Quiescent, Immune effector, Immune suppressed, ECM remodeling,
TAN infiltration, EMT transition — arranged on a monotonic low-to-high thyrocyte ATC-score axis.
Perivascular and Quiescent showed mutual boundary enrichment (O/E 2.8–4.0); at the opposite pole, TAN
infiltration enveloped EMT transition at the highest adjacency (O/E≈8.2, FDR<0.001), with cores ~88% ATC.
Immune suppressed was 3.3-fold enriched in BRAF V600E and correlated with PD-L1 IHC (p<0.05); NRAS favored
EMT/TAN. TAN and Immune suppressed conferred the highest Cox hazards (HR=6.71, FDR<10â»â¶); an Elastic Net
classifier reached AUROC=0.95 (patient-grouped CV) for ATC vs others. Within Quiescent (Q), dedifferentiation
was clonal (Moran's I 0.15–0.39, FDR<0.05), driven by TACSTD2/MUC1 gain and EPCAM/CDH1 loss. TROP2/MUC1
retained epithelial identity (PTC>PDTC>ATC) and topped Q intra-heterogeneity drivers (Stouffer Z>17);
their bulk composite stratified DFS in DTC (HR≈2.2, p=0.003) but not ATC, flagging DTC recurrence risk.
Conclusion: Niche-resolved mapping yields actionable stratification: Immune suppressed identifies PD-L1-high
BRAF V600E potential ICI candidates; EMT/TAN mark NRAS-associated aggression for TKI or neutrophil targeting;
Perivascular/Quiescent identify indolent disease with TROP2/MUC1 surveillance.
C-S.B.38: Data-driven subgrouping of individuals on Alzheimer's continuum based on amyloid-β
aggregation
Track: Systems biology, multi-omics integration, modeling
-
Arina Tagmazian, Institute for Molecular Medicine Finland, HiLIFE, University of Helsinki,
Finland
-
Eero Vuoksimaa, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Finland
-
Esa Pitkänen, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Finland
Presentation Overview: Show
Amyloid positron emission tomography (PET) is commonly used to classify individuals as amyloid-β (Aβ)
negative or positive, but binary approach may overlook intermediate stages of pathology. Here, we aimed to
identify data-driven subgroups along the Alzheimer's disease (AD) continuum using Aβ PET imaging to enable a
more nuanced characterization of amyloid accumulation.
We analyzed 3,110 Aβ PET scans from the ADNI and A4 cohorts using petVAE, a variational autoencoder trained
to reconstruct two-dimensional PET slices without diagnostic labels or predefined regions of interest. The
model generated 11,648-dimensional latent representations per scan, which were used for exploratory analyses
and clustering across the AD spectrum.
We identified four clusters that differed significantly in standardized uptake value ratio (p <
1.64×10â»â¸) and cerebrospinal fluid (CSF) Aβ levels (p < 0.02), indicating that petVAE effectively
positions scans along a continuous Aβ trajectory. Two clusters (Aβ−, Aβ−+) were largely
amyloid-negative, while two (Aβ+, Aβ++) were predominantly amyloid-positive. The extreme clusters (Aβ−,
Aβ++) aligned with conventional classifications and showed marked differences in cognition, APOE ε4 carrier
frequency, and CSF Aβ, Tau, and phosphorylated Tau (p < 3×10â»â¶). Intermediate clusters exhibited
increased odds of APOE ε4 carriership (p < 0.026). Individuals in Aβ+ and Aβ++ clusters had a higher
risk of progression to AD over 6 years (hazard ratios 2.42 and 9.43; p < 1.17×10â»â·).
Overall, petVAE reconstructs PET images accurately while learning biologically meaningful representations of
Aβ pathology. This data-driven framework enables detection of subtle disease stages and supports
investigation of preclinical AD.
C-S.B.39: MorphoMapper: 3D Segmentation-Driven Feature Mapping of Stress-Perturbed Mitochondrial
Remodeling
Track: Systems biology, multi-omics integration, modeling
-
Subasini Thangamani, Computational Systems Biology, RPTU Kaiserslautern-Landau, Germany
- Sadia S. Tamanna, Molecular botany, RPTU Kaiserslautern-Landau, Germany
- Simon Foellinger, Computational Systems Biology, RPTU Kaiserslautern-Landau, Germany
- Sophie Pompejus, Molecular botany, RPTU Kaiserslautern-Landau, Germany
- Stefanie Mueller-Schuessele, Molecular botany, RPTU Kaiserslautern-Landau, Germany
- David Zimmer, Computational Systems Biology, RPTU Kaiserslautern-Landau, Germany
- Timo Muehlhaus, Computational Systems Biology, RPTU Kaiserslautern-Landau, Germany
Presentation Overview: Show
Organelle morphology provides a sensitive and integrative readout of cellular state, reflecting changes in
metabolism, signaling, and environmental perturbations such as stress. While three-dimensional (3D)
fluorescence microscopy enables detailed observation of organelle structure, translating volumetric image data
into biologically meaningful conclusions remains challenging. Existing approaches often address only
individual analysis steps and rely on summary statistics of morphometric descriptors, limiting their ability
to capture population-level heterogeneity and perturbation-dependent remodeling.
Here, we present Morphomapper, a FAIR end-to-end workflow for quantitative analysis of organelle morphology
from 3D confocal fluorescence imaging. Morphomapper integrates automated deep learning-based volumetric
segmentation, extraction of geometric features, dimensionality reduction, and statistical inference into a
reproducible and FAIR-compliant pipeline. We benchmark automated 3D segmentation against expert human
annotations, demonstrating robust reconstruction of complex organelle shapes. Using the resulting
segmentations, we systematically compare two-dimensional and three-dimensional morphometric descriptors and
show that 3D features substantially improve the discrimination of stress perturbations. Importantly,
Morphomapper treats organelles as populations rather than isolated objects. We introduce an optimal
transport-based framework to compare distributions of organelle morphologies across perturbations, enabling
principled quantification of stress-induced shifts in population structure. By operating on full morphological
distributions instead of mean descriptors alone, Morphomapper captures heterogeneity as a primary signal and
links structural remodeling to environmental perturbations. This framework bridges the gap between volumetric
segmentation and biologically actionable inference, providing a scalable approach for identifying
morphology-based biomarkers of stress perturbation states from 3D imaging data.
C-S.B.40: Interactive Interfaces for Comprehensive Genetic Evidence in Target Identification: The Open
Targets Platform
Track: Systems biology, multi-omics integration, modeling
-
Ricardo Esteban Martinez Osorio, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United
Kingdom
-
Ellen M McDonagh, Open Targets, Wellcome Sanger Institute, Cambridge, United Kingdom, United Kingdom
-
David G Hulcoop, Open Targets, Wellcome Sanger Institute, Cambridge, United Kingdom, United Kingdom
-
Yakov Tsepilov, Open Targets, Wellcome Sanger Institute, Cambridge, United Kingdom, United Kingdom
-
Szymon Szyszkowski, Open Targets, Wellcome Sanger Institute, Cambridge, United Kingdom, United Kingdom
-
Xiangyu Jack Ge, Open Targets, Wellcome Sanger Institute, Cambridge, United Kingdom, United Kingdom
-
Daniel Considine, Open Targets, Wellcome Sanger Institute, Cambridge, United Kingdom, United Kingdom
-
Wei Wen Vivien Ho, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Tobi Alegbe, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Polina Rusina, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Carlos Cruz-Castillo, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United
Kingdom
-
James D Hayhurst, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Graham McNeill, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Irene Lopez, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Helena Cornu, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Javier Ferrer, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
David Ochoa, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Daniel Suveges, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
-
Annalisa Buniello, Open Targets, EMBL-EBI, Hinxton, Cambridge, United Kingdom, United Kingdom
Presentation Overview: Show
Drug discovery faces significant challenges — 90% of drugs entering Phase 1 clinical trials never reach the
market. Open Targets, a public-private partnership, integrates human genetics and genomics data to
systematically address these challenges through the Open Targets Platform (platform.opentargets.org).
The Platform integrates data from over 20 diverse sources to build and score gene-disease associations,
supporting evidence-based target prioritisation for drug discovery. In its 10th year, the Platform underwent a
major expansion incorporating the most comprehensive set of genetic associations to date, requiring
significant advances in both data infrastructure and user-facing interfaces.
This update includes ancestry-specific fine-mapping and colocalisation of GWAS and molecular QTL data from the
GWAS Catalog, eQTL Catalogue, FinnGen, and UK Biobank Pharma Proteomics Project. Over 2.6 million credible
sets were generated, yielding more than 400,000 gene-disease associations supported by over 1 million evidence
items.
To surface this complexity, new interactive interfaces were designed and built: variant-centric pages display
credible sets, colocalisation results, and clinical and pharmacogenomic annotations for over 6.5 million
variants. New study and credible set pages allow researchers to explore GWAS and QTL evidence in detail.
Enhanced machine learning–based Locus-to-Gene (L2G) assignments further refine gene prioritisation,
identifying more than 358,000 credible sets with L2G score >0.5.
These frontend-driven enhancements unify evidence from common and rare variant studies within a single
explorable interface, enabling researchers to build stronger causal links between genes and diseases and
supporting data-driven drug target identification.
C-S.B.41: Using predictive multiplicity in biologically informed neural networks to uncover disease
heterogeneity
Track: Systems biology, multi-omics integration, modeling
-
Dennis Gankin, ETH Zurich, Switzerland
- Pedro Beltrao, ETH Zurich, Switzerland
Presentation Overview: Show
Biologically informed neural networks (BINNs) embed pathway, ontology, or protein interaction structure
directly into neural networks, promising interpretable disease prediction in which hidden nodes map to named
biological entities. Yet BINNs have been difficult to train at biobank scale, and the reliability of their
biological interpretations remains largely untested.
Here, we present a fast BINN implementation to train on UK Biobank genotype and plasma proteomics data from
~450,000 individuals across six common diseases. BINNs achieve competitive predictive performance, but we
uncover two major limits to their interpretability. First, attribution scores are strongly biased by graph
topology, so node degree and layer position explain most of the variance in the scores. A simple normalization
reduces this bias, but can weaken enrichment for known disease genes. Second, BINNs exhibit substantial
predictive multiplicity: independently trained models with identical architectures and data reach similarly
accurate solutions while prioritizing different genes and pathways. We show that this multiplicity persists in
highly predictive models even after reducing model parameters, network complexity and input correlation.
Although this multiplicity makes single-model explanations unstable, the range of different interpretations
can also reveal complex disease biology. In 100 replicate BINNs for type 2 diabetes prediction from proteomics
data, we find distinct solution clusters that prioritize either inflammatory or hepatic-metabolic pathways,
mirroring known disease heterogeneity. Thus, training and analyzing large ensembles of BINNs can turn
multiplicity into a tool for studying complex disease mechanisms in human cohorts.
C-S.B.42: Assessing the relative contributions of mosaic and regulatory developmental modes from single-cell
trajectories
Track: Systems biology, multi-omics integration, modeling
-
Solene Song, Aix-Marseille Université, MMG, Inserm U1251, Turing Centre for Living systems, Marseille,
France, France
-
Paul Villoutreix, Aix-Marseille Université, MMG, Inserm U1251, Turing Centre for Living systems,
Marseille, France, France
Presentation Overview: Show
Development is a complex process driven by coordinated cell proliferation, differentiation, and spatial
organization. Classically, two ways to specify cell types during development have been hypothesized: the
mosaic and regulative modes. In the mosaic mode, the fate of a cell rely on lineage-inherited factors. In
contrast, in the regulative mode, the fate of a cell depends on space-dependent factors. The relative
contributions of both modes remain poorly quantified. We present a novel approach to measure these
contributions from single-cell data from C. elegans development. The invariant lineage of C. elegans allows
the integration of spatial positions, lineage relationships, and protein expression data. Using single-cell
protein expression profiles as a readout of cell state, we define two quantifiable metrics: 1) a proxy for the
contribution of the mosaic mode, computed as the strength of the relationship between the cell-cell lineage
distance and the cell-cell expression distance, 2) a proxy for the contribution of the regulative mode,
computed as the strength of the relationship between the cell-cell context distance - capturing spatial
neighborhood similarity - and the cell-cell expression distance. To validate these metrics, we compared
empirical results from C. elegans to artificial models with defined developmental rules. Our analysis reveals
the coexistence of mosaic and regulative modes, with their relative contributions varying across tissues and
developmental stages. For example, in skin tissue, the mosaic mode dominates in early development, while the
regulative mode prevails later. Our approach offers a quantitative, unbiased, and perturbation-free method to
study fundamental principles of developmental biology.
C-S.B.43: LGTM: Gaussian Process Modulated Neural Topic Modeling for Longitudinal Microbiome
Track: Systems biology, multi-omics integration, modeling
-
Xiao Yuan, University of Helsinki, Finland
- Ádám Arany, KU Leuven, Belgium
- András Formanek, KU Leuven, Belgium
- Yves Moreau, KU Leuven, Belgium
- Harri Lähdesmäki, Aalto University, Finland
- Tommi Vatanen, University of Helsinki, Finland
Presentation Overview: Show
Longitudinal microbiome data are key to understanding the dynamics of microbial communities and their
relationships with the host and environment. However, analysis of such data is challenging due to high
dimensionality, compositionality, irregular sampling and temporal dependencies on external covariates.
Existing analytical approaches typically address only subsets of these challenges, limiting their ability to
yield biologically interpretable insights. We introduce LGTM, a probabilistic modeling framework that combines
flexible non-linear longitudinal modeling with interpretable topic-based representations of the microbiome.
LGTM simultaneously discovers coherent microbial subcommunities (""topics"") and models how their abundances
change over time and in relation to host and environmental covariates. Using multiple longitudinal human gut
microbiome datasets, we demonstrate that LGTM identifies diverse and stable microbial topics while achieving
competitive performance in imputation and forecasting tasks. A key strength of the framework is its
interpretability: LGTM discovers biologically coherent microbial topics and directly quantifies associations
between covariates and microbial dynamics. LGTM is available at https://github.com/yuanx749/lgtm.
C-S.B.44: GPU-accelerated WGCNA for scalable gene co-expression analysis
Track: Systems biology, multi-omics integration, modeling
-
Felix Jung, RPTU Kaiserslautern-Landau, Germany
- David Zimmer, RPTU Kaiserslautern-Landau, Germany
- Timo Muehlhaus, RPTU Kaiserslautern-Landau, Germany
Presentation Overview: Show
Transcriptomic profiling provides a global, top-down view of cellular responses to environmental conditions,
perturbations, and disease. Weighted gene co-expression network analysis (WGCNA) is among the most widely used
methods to model these responses, grouping co-expressed genes into modules that can be related to sample
traits and summarized by hub genes. However, because modern transcriptomic datasets quantify tens of thousands
of genes, WGCNA’s quadratic scaling with gene number becomes prohibitive. This is especially problematic for
resampling-based workflows, which rerun WGCNA across many bootstrapped datasets to assess module robustness or
stabilize recovered modules.
A common way to reduce this cost is blockwise WGCNA, which partitions the gene set and analyzes it in chunks.
Across the Tabula Muris Senis dataset, we find that blockwise WGCNA reproduces full-dataset WGCNA in its
module–trait correlations while differing substantially in module gene content. With a block size of half the
genes, the average Jaccard similarity between corresponding modules is only 0.6.
Here, we present tensor-wgcna, a GPU-accelerated PyTorch/Triton implementation that, on a single consumer GPU
(RTX 4090), runs up to 10 times faster than the R implementation and uses roughly half the peak memory, while
retaining a module Jaccard similarity of 0.9-1.0 across the Tabula Muris Senis and GTEx adult datasets.
Lastly, we demonstrate the practical relevance of tensor-wgcna on a hand-curated cold-stress dataset from
Arabidopsis thaliana leaves (90 samples, 23,668 genes). Using bootstrap-based module stabilization, we not
only recover functional modules but also align them with physiologically relevant phases of the cold
response.
C-S.B.45: Random-walk multi-omics integration for rare-disease gene prioritization
Track: Systems biology, multi-omics integration, modeling
-
Mariam Ait Oumelloul, GHI, Switzerland
- Barry Ryan, École Polytechnique Fédérale de Lausanne, Switzerland
- Daphné Chopard, Department of Computer Science, ETH Zurich, Switzerland
-
Vito Zanotelli, Division of Metabolism and Children's Research Center, University Children's Hospital
Zürich, University of Zürich, Switzerland
- Chaimae El Houjjaji, École Polytechnique Fédérale de Lausanne, Switzerland
-
Ali Saadat, School of Life Sciences, Ecole Polytechnique Fédérale de Lausanne (EPFL), Switzerland
-
Sean Froese, Division of Metabolism and Children's Research Center, University Children's Hospital
Zürich, University of Zürich, Switzerland
-
Jacques Fellay, School of Life Sciences, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland,
Switzerland
Presentation Overview: Show
Whole-genome sequencing (WGS) has improved rare-disease diagnosis, yet 50–60% of patients remain
undiagnosed, underscoring persistent challenges in causal variant prioritization. Single-omics outlier
analyses, particularly RNA-seq, can increase diagnostic yield, but are often performed separately for each
omics layer or combined only through late integration. We hypothesized that integrating phenotypic and
multi-omics evidence within a unified network framework would improve causal gene prioritization.
We developed an integration framework that jointly models HPO-encoded phenotypes and patient-specific omics
signals to prioritize causal genes. The method constructs a heterogeneous, multi-layer knowledge graph linking
diseases, phenotypes, genes, and biological processes, with optional patient-specific omics subnetworks
connected through gene nodes. For each patient, random walk with restart (RWR) is initiated from phenotype
seed nodes to compute gene-prioritization scores. We first evaluated whether applying RWR on the
knowledge-network backbone alone, without omics profiles, could serve as an effective gene-prioritization
strategy. Performance was benchmarked against Phen2Gene and GADO using MyGene2 (N=146) and a subset of the
Deciphering Developmental Disorders cohort (DDD; N=164). We then assessed a multi-omics extension in the
SwissPedHealth cohort (N=37).
Our RWR framework improved causal-gene prioritization, ranking the causal gene in MyGene2 within the top 10
for 26.0% of cases, compared with 17.4% for Phen2Gene, and 2.1% for GADO. Performance was higher in the DDD
subset (RWR: 86.6% top 10, Phen2gene: 73.0%, GADO: 39%). Adding multiomics evidence in SwissPedHealth further
improved ranking beyond phenotype-only inference, highlighting the potential of RWR on heterogeneous graphs
for propagation of phenotype and multi-omics signals in patient-specific gene prioritization.
C-S.B.46: Integrating LINCS L1000 transcriptomic repurposing with graph convolutional drug-target interaction
prediction identifies subtype-specific therapeutic candidates and novel targets in IBD
Track: Systems biology, multi-omics integration, modeling
-
Mohamed Farh, AI-Bio Convergence Research Institute, Soongsil University, Seoul, Republic of Korea,
South Korea
-
Seoyoung Jung, Department of Biomedical Systems, Soongsil University, 369 Sangdo-ro, Dongjak-gu, Seoul
06978, Republic of Korea, South Korea
-
Jae Yong Ryu, Department of Biomedical Systems, Soongsil University, 369 Sangdo-ro, Dongjak-gu, Seoul
06978, Republic of Korea, South Korea
Presentation Overview: Show
Inflammatory bowel disease (IBD) remains a therapeutically challenging, molecularly heterogeneous condition
for which most approved agents have been developed as pan-disease therapies, overlooking the distinct
pathogenic drivers of ulcerative colitis (UC) and Crohn's disease (CD). To uncover subtype-specific
therapeutic opportunities, we integrated bulk RNA-seq from seven independent cohorts — three CD cohorts (317
patients vs. 91 controls) and four UC cohorts (175 patients vs. 94 controls) — into an enhanced LINCS
L1000–based drug repurposing pipeline. This analysis nominated two distinct, subtype-selective mechanisms of
action: cholesterol-uptake inhibition (ezetimibe) for UC and histone deacetylase (HDAC) inhibition (vorinostat
and pyroxamide) for CD. To systematically map the on- and off-target landscape of these candidates, we
developed a graph convolutional network (GCN) model for drug–target interaction prediction, achieving
area-under-the-precision-recall curves (AUPRC) of 0.89 and 0.95 on validation and test sets, and 74–80.5%
top-100 accuracy when benchmarked against the Pabon reference dataset. Among the top 30 GCN-predicted targets
for each candidate, 13 (ezetimibe), 6 (vorinostat), and 14 (pyroxamide) emerged as novel — neither annotated
as direct drug targets in pharmacological databases nor previously linked to the corresponding IBD subtype.
Together, these results nominate pharmacologically tractable, subtype-resolved therapeutic candidates and
uncover previously uncharacterized targets supporting a stratified-medicine framework for IBD. Experimental
validation of the prioritized drug–target pairs is ongoing.
C-S.B.47: MicroKnow: Quantifying the Mechanistic Evidence Gap in Clinical Microbe-Disease
Associations
Track: Systems biology, multi-omics integration, modeling
-
Shahad Qathan, University of Siegen, Germany
- Florian Centler, University of Siegen, Germany
Presentation Overview: Show
Clinical microbiome databases document thousands of microbe–disease associations, yet most lack a curated
molecular mechanism, and the evidence needed to trace one is scattered across resources with incompatible
identifiers. Researchers have no automated way to determine which reported associations already have
mechanistic support.
We present MicroKnow, a Neo4j knowledge graph integrating clinical microbe-disease evidence with mechanistic
evidence across microbes, metabolites, host genes, and diseases (46,911 nodes, 738,096 edges), designed as a
hypothesis-generation tool for experimental microbiome research. For each of 12,902 reported microbe-disease
pairs, we ask whether the same microbe reaches the same disease through a metabolite it produces and a host
gene that metabolite regulates.
Of 12,902 pairs, 17.4% recover such a path, a 1.11-fold enrichment over a degree-preserving null (z = 10.4, p
< 0.001). However, restricting the comparison to pairs where a path could form shows reported and
unreported pairs are equally likely to carry a mechanism (62%, p = 0.45), indicating the overlap reflects
annotation density rather than biological concentration. MicroKnow maps the gap bidirectionally: 10,661
reported pairs lack any mechanism, while 13,394 mechanism-grounded pairs lack clinical validation. Only 5.8%
of metabolites carry any documented gene interaction, making metabolite-to-gene curation the main
bottleneck.
A Dash web interface (microknow.de) supports pair- and disease-level queries and returns a ranked shortlist of
clinically unvalidated microbe–disease pairs as targets for experimental follow-up.
C-S.B.48: Delineation of signaling routes that underlie differences in macrophage phenotypic states
Track: Systems biology, multi-omics integration, modeling
-
Marija Buljan, Empa, Switzerland
- Katharina Sribike, Empa, ETH Zurich, Switzerland
- Lukas Haeuser, Empa, ETH Zurich, Switzerland
- Tiberiu Totu, Empa, ETH Zurich, SIB, Switzerland
- Jonas Bossart, Empa, ETH Zurich, SIB, Switzerland
- Elana Caire, Empa, ETH Zurich, SIB, Switzerland
- Vanesa Ayala-Nunez, Empa, Switzerland
- Bettina Sobottka, UniversityHospitalZurichandUniversityofZurich, Switzerland
- Markus Rottmar, Empa, Switzerland
Presentation Overview: Show
Macrophages represent a major immune cell type in tumor microenvironments, they exist in multiple functional
states and are of strong interest for therapeutic reprogramming. While signaling cascades defining
proinflammatory macrophages are better characterized, pathways that drive polarization in immunosuppressive
macrophages are incompletely mapped. We exposed primary human macrophages to a range of stimuli, which is
abundant in the tumor microenvironment, and profiled the induced transcriptome, proteome and phosphoproteome
changes. We rank-normalized the gene expression levels and aligned the in vitro transcriptomes to a pan-cancer
macrophage atlas built from single-cell RNA sequencing profiles collected from over 30 patient studies. This
showed that in vitro macrophages were able to recapitulate different functional aspects of the in vivo states.
For instance, exposure to adenosine resulted in the upregulation of a dozen of metallothionein genes and
downregulation of the MHC-II antigen presentation system, which strongly resembled transcriptional state of
Metallo Macrophages – poorly studied and tumor associated macrophages linked to bad prognosis. Furthermore,
we mapped several under-appreciated kinase signaling cascades that play a role in the establishment of
immunosuppressive states, such as those involving the PAK2 kinase, and we integrated the generated multi-omics
profiles by mapping the significant hits to a knowledge-based interaction network, which allowed us to find
modules of highly connected elements. The latter approach is available through the NOODAI web platform.
Overall, this study contributes to in-depth multi-omics characterizations of macrophage phenotypic landscapes,
which can be of relevance for assisting future interventions that aim to therapeutically alter immune cell
compartments.
C-S.B.49: MetaboViz: A Zero-Install Three-Tier Browser Platform for Interactive Genome-Scale Metabolic
Modelling
Track: Systems biology, multi-omics integration, modeling
- Tamoghna Das, Loughborough University, United Kingdom
-
M. Ahsanul Islam, Loughborough University, United Kingdom
Presentation Overview: Show
Constraint-based metabolic modelling tools — COBRApy, COBRA Toolbox, KBase — uniformly require local
installation, cloud accounts, or queued HPC jobs, creating friction that impedes exploratory analysis and
classroom use. Browser-based alternatives such as Escher-FBA and Fluxer exist but rely on server-side solvers
or GLPK.js executing on the main UI thread, blocking interaction for models beyond ~500 reactions.
We present MetaboViz, an open-source React application implementing a three-tier compute architecture that
automatically routes analyses to the fastest available solver: (i) a HiGHS WebAssembly Worker running
off-thread in the browser, requiring zero installation; (ii) a local Python kernel (pip install
metaboviz-kernel) communicating via WebSocket JSON-RPC, which loads the model once and keeps it resident
across repeated interactive solves; and (iii) a stateless FastAPI edge service as fallback.
We benchmark all tiers on four BiGG models spanning 95–3,942 reactions. The local kernel achieves FBA solve
times of 0.4–15 ms after a one-time model load (19 ms–1.1 s), and FVA of 19 ms–4.8 s using parallel
dual-simplex. The edge tier is dominated by HTTP serialisation rather than solver time: iJO1366 (2,879 KB
payload) incurs 1.0 s wall time versus 82 ms solver time, confirming that model caching in the kernel tier is
essential for interactive use. Results are presented as Jupyter-style notebook cells with per-reaction flux
visualisation and tier provenance.
MetaboViz enables sub-second FBA on genome-scale models in any browser, with no login or installation
required, filling a gap between lightweight educational tools and full platforms such as KBase.
C-S.B.50: Multimodal integration reveals immune mechanisms of an effector consortium
Track: Systems biology, multi-omics integration, modeling
-
Erika Kvalem, Biocenter, Institute of Bioinformatics, Medical University of Innsbruck, Innsbruck,
Austria., Austria
-
Nina Boeck, Biocenter, Institute of Bioinformatics, Medical University of Innsbruck, Innsbruck, Austria.,
Austria
-
Gregor Sturm, Biocenter, Institute of Bioinformatics, Medical University of Innsbruck, Innsbruck,
Austria., Austria
-
Christina Plattner, Biocenter, Institute of Bioinformatics, Medical University of Innsbruck, Innsbruck,
Austria., Austria
-
Alexander Kirchmaier, Biocenter, Institute of Bioinformatics, Medical University of Innsbruck, Innsbruck,
Austria., Austria
-
Georgios Fotakis, Biocenter, Institute of Bioinformatics, Medical University of Innsbruck, Innsbruck,
Austria., Austria
-
Takeshi Tanoue, Department of Microbiology and Immunology, Keio University School of Medicine, Tokyo,
Japan, Japan
-
Kenya Honda, Department of Microbiology and Immunology, Keio University School of Medicine, Tokyo, Japan,
Japan
-
Lorenzo Galuzzi, Department of Radiation Oncology, Weill Cornell Medical College,New York, NY, United
States, Austria
-
Zlatko Trajanoski, Biocenter, Institute of Bioinformatics, Medical University of Innsbruck, Innsbruck,
Austria., Austria
Presentation Overview: Show
The mechanisms by which a defined effector commensal bacterial consortium promotes CD8+ T cell-mediated
anti-cancer immunity remain unknown. We developed a multimodal immune profiling and data integration
framework, coupled with a bacterial antigen prediction pipeline, to resolve how microbial signals are
transmitted across intestinal and systemic compartments, supported by complementary in vitro and in vivo
validation.
Two mechanistic axes were investigated: (1) molecular mimicry between bacterial peptides and tumor
neoantigens, and (2) epithelial reprogramming by bacterial-derived products. For the first we developed a
computational antigen prediction pipeline that interrogates genomic sequences of the effector consortium to
identify cross-reactive bacterial peptides, followed by in vivo validation in mouse models. For the second, we
integrated bulk RNA sequencing and FC of intestinal organoids exposed to bacterial supernatants. Systemic
immune responses were further characterized by serum cytokine profiling and single-cell RNA sequencing of CD8
T cells from ICI-treated mice, integrating transcriptomic, TCR, and protein data.
While in vivo studies did not support molecular mimicry, pathway and cytokine modeling identified IFN-gamma
signaling and responses to bacterial-derived molecules as dominant features of the effector condition. We
observed epithelial-derived cytokines propagating systemically, together with consistent upregulation of Cxcl9
and Cxcl10 in serum, and increased expression of their receptor Cxcr3 in CD8 T cells. These findings define a
coordinated chemokine axis linking microbial stimulation to effector T cell recruitment.
This integrative multimodal approach provides a better understanding of the biology of the complex
host-microbiome system and establishes a mechanistic basis for enhanced anti-tumor immunity.
C-S.B.51: Microglia as transcriptional pacemakers of neuronal aging heterogeneity across individuals
Track: Systems biology, multi-omics integration, modeling
-
Christine Lim, Yusuf Hamied Department of Chemistry, University of Cambridge, United Kingdom
Presentation Overview: Show
Neuronal aging pace varies markedly between individuals, but what drives this variation remains unknown. Using
cell-type-specific transcriptomic clocks applied to single-nucleus RNA sequencing data from 226 adults (ages
20-90), we quantified neuronal aging residuals as a donor- dominant phenotype. Variance decomposition revealed
that microglial transcriptional programs predict inter-individual variation in neuronal aging residuals, a
directional asymmetry consistent with a non-cell-autonomous relationship between microglial states and
neuronal aging trajectories. This asymmetry is accompanied by an age-dependent shift from homeostatic to
inflammatory microglial dominance beginning in midlife, with inflammatory dominance probability rising from
26% at age 35 to 92% by age 65, replicated in an independent cohort. IFNγ signaling emerges as the dominant
microglial program associated with accelerated neuronal aging in late adulthood. Candidate regulators of
microglial IFNγ activity are computationally prioritized as intervention targets warranting functional
validation.
C-S.B.52: A Computational Framework for Cross-Species Single-Cell Atlas Integration Reveals Conserved and
Divergent Transcriptional Programs in Tongue Pain Circuits
Track: Systems biology, multi-omics integration, modeling
-
Cristal Villalba, University of Texas Health San Antonio, Texas, USA, United States
-
Jaclyn Merlo, Department of Endodontics, School of Dentistry, University of Texas Health San Antonio,
Texas, USA, United States
-
Sergey Shein, Department of Microbiology, Immunology and Molecular Genetics, University of Texas Health
San Antonio, Texas, USA, United States
-
Zhao Lai, Greehey Children's Cancer Institute, University of Texas Health San Antonio, San Antonio, USA,
United States
-
Yidong Chen, Greehey Children's Cancer Institute, University of Texas Health San Antonio, San Antonio,
USA., United States
-
Shivani Ruparel, Department of Endodontics, School of Dentistry, University of Texas Health San Antonio,
Texas, USA, United States
Presentation Overview: Show
Animal models are essential for studying nociception, yet differences in cellular and transcriptional
architecture complicate translation to human pain biology. To address this gap, we developed a cross-species
single-cell transcriptomic integration pipeline to build a comprehensive tongue atlas from mouse, rat,
marmoset, and human datasets. Tongue tissue from naïve C57BL/6 mice and common marmosets was collected and
processed for scRNA-seq using 10x Genomics platforms. Human datasets were obtained from the Tabula Sapiens
Consortium, and rat datasets from NCBI GEO. A unified human gene space was constructed using species-specific
ortholog mapping strategies — Ensembl-derived tables for mouse and rat, and NCBI Gene E-utilities for
marmoset — yielding ~17,000 mapped genes. Species-aware batch correction was applied using Harmony with
conservative parameterization to preserve biological divergence. The final atlas comprises 135,736 cells
spanning 7 major cell types. To systematically quantify cross-species conservation, we implemented a
multi-step pipeline using FindConservedMarkers (log2FC > 0.6, Bonferroni P < 0.05 across all species),
identifying 1,342 conserved genes from 4,026 candidates. Normalized rank-based standard deviation analysis
classified 51.8% of conserved genes as stable, including 206 highly stable genes. Epithelial cells showed the
strongest transcriptional conservation (Pearson r = 0.83–0.86), while Schwann cells exhibited the highest
variability. Validation against canonical markers confirmed conservation in 17 of 24 genes. This atlas and
accompanying pipeline provide a scalable computational framework for comparative and translational studies of
pain-relevant biology.
C-S.B.54: Multimodal Siamese Network for Parkinson's Disease Diagnosis and Structure-based Drug Discovery for
Therapeutic Development
Track: Systems biology, multi-omics integration, modeling
-
Allison Huang, East Brunswick High School, United States
- Shaurya Gandhi, East Brunswick High School, United States
-
Yuqi Zhang, Department of Computer Science; Lewis-Sigler Institute of Integrative Genomics, Princeton
University, United States
Presentation Overview: Show
Parkinson’s Disease (PD) is often diagnosed after irreversible neurodegeneration due to reliance on physical
symptoms for diagnosis. Current machine learning (ML) diagnosis models often rely on a single data modality,
which limits predictive performance and disease mechanisms identification. Integrating multiple molecular and
imaging modalities offers an opportunity to improve diagnostic accuracy while enabling systems-level insights
into PD pathogenesis. We developed a multimodal ML framework combining MRI, biomarkers, and miRNA from
Parkinson's Precision Medicine Initiative (PPMI). We used a Siamese Network with triplet loss to overcome data
scarcity and model complex patient-similarity representations. We used model feature importances to create
therapeutic hypotheses for targeting PD, leading to the identification of IL-6 as a gene driving PD
inflammation. We validated IL-6 as a target through differential gene expression and gene set enrichment
analyses. We applied structure-based drug design to inhibit IL-6 by first generating fragments within
computationally identified binding pockets. These fragments were developed into final candidates through
iterative optimization based on random mutation and selection of high-scoring structures, mimicking an
evolutionary model. Our model achieves an ROC AUC of 0.94, leading to performance superior to single modality
or naive data integration baseline methods. Using CNN-based docking and scoring, we designed drug candidates
with binding affinities of -11.2 and -11.5 kcal/mol. This study offers a unique end-to-end pipeline linking
multi-omics integration with structure-based drug design, demonstrating how interpretable multimodal machine
learning can improve PD diagnosis and uncover disease mechanisms while identifying therapeutic targets.
C-S.B.55: AgroLD: a knowledge graph for the plant sciences
Track: Systems biology, multi-omics integration, modeling
-
Pierre Larmande, IRD, France
- Bill Gates Happi Happi, IRD, France
- Bertrand Pitollat, CIRAD, France
- Ndomassi Tando, IRD, FRANCE
Presentation Overview: Show
The demand for food is expected to grow substantially in the coming years. To address this challenge,
especially in the context of climate change, a deeper understanding of genotype-phenotype relationships is
crucial for improving crop yields. Recent advances in high-throughput technologies have transformed the
landscape of plant science research. However, there is an urgent need to integrate and consolidate
complementary data to understand the biological system.
We introduce AgroLD, a knowledge graph that uses Semantic Web technologies to seamlessly integrate plant
science data. AgroLD is designed to facilitate hypothesis formulation and validation within the scientific
community. With approximately 1 billion triples, it integrates and annotates data from more than 151 datasets
across 19 distinct sources.
The overarching goal is to provide a specialized knowledge platform addressing complex biological questions in
the plant sciences, including gene participation in plant disease resistance and adaptive responses to climate
change.
C-S.B.56: Core genome translational genomics - A framework to combine multi species data for the
understanding of complex traits
Track: Systems biology, multi-omics integration, modeling
-
Sonia Eynard, INRAE - GenPhySE, France
- Juliette Riquet, INRAE - GenPhySE, France
- Julie Demars, INRAE - GenPhySE, France
Presentation Overview: Show
In the current genomic data jungle it is tempting to integrate data not only at a multi omics level but also
across multiple species. Major research fields focus on comparative genomics, where knowledge on multiple
species is compared to build stronger evidence for biological mechanisms and on the concept of translational
genomics, also known as precision medicine where one uses model organisms to understand specific human
diseases and adapted treatments. In addition, recent developments, such as pan genomics, bring new technical
opportunities to describe similarities between population and species.
In this context we focus on the understanding of the biological mechanisms underlying feed efficiency, the
capacity to convert feed into production gains, using experimental populations of pig and rabbit. We first
looked into orthogroups between the two species and ambitioned to detect orthoSNPs, genomic positions
localised in the same genome functional unit and showing variability in both species. Based on this knowledge
we will perform genome wide association studies and selection signature detection on each species separately
and combining genomic and phenotypic information coming from the experimental designs on both species. We hope
to identify species specific biology pathways and responses to selection for feed efficiency and using the
combined set up common, universal, biological pathways explaining variability in feed efficiency.
With this study we intend to build a modelling framework that could be extended to other phenotypes and
species, always considering species evolutionary history as a keep criteria to combine study designs.
C-S.B.57: Harnessing AI to Identify Key Microbial Drivers of Stable State Microbiomes in Low Emitting
Ruminants
Track: Systems biology, multi-omics integration, modeling
-
James Barnard, Queen's University Belfast, United Kingdom
- Christopher Creevey, Queen's University Belfast, United Kingdom
- Robert Atkinson, University of Strathclyde,
Presentation Overview: Show
The rumen microbiome plays a critical role in livestock productivity and environmental sustainability, yet our
understanding of the ecological drivers governing stable-state microbial communities remains limited.
Host-associated microbial communities influence digestion and health while simultaneously producing greenhouse
gas emissions. The rumen represents a dynamic ecosystem, characterised by temporally distinct microbial
enzymatic action during feed digestion. Temporal succession, driven by niche specialisation facilitates
development of stable microbial communities. This research aims to develop and apply AI techniques to uncover
the key ecological drivers of stable-state rumen microbial communities and their functional outcomes. Analysis
of microbiome data is hindered by batch effects and data heterogeneity across studies. By integrating machine
learning models with microbiome data, this project will identify key descriptors of niche specialisation and
microbial community dynamics that influence both productivity and environmental impact. Traditional supervised
methods can fail to capture universal patterns which limits their application to large scale datasets in
diverse contexts. Our approach involves developing rumen-specific foundation models capable of capturing the
complexity and temporal dynamics of rumen microbiome data while maintaining biological interpretability.
Transformer-based language models employ self-attention and masked-language modelling to learn contextual
representations of microbial community composition, enabling transfer learning in downstream tasks. Leveraging
metagenomic datasets of thousands of samples collected over several years, this model will reveal which
microbial profiles and community interactions drive desirable outcomes, such as reduced methane emissions and
improved feed efficiency. This work addresses a gap where microbial ecology, livestock productivity, and
computational biology intersect. The anticipated outcomes include AI-driven tools that can inform livestock
management strategies, enabling targeted interventions to optimise rumen function while reducing greenhouse
gas emissions.
C-S.B.58: ToxCast Evidence Graph: Portable Semantic Access to Bioactivity and Product-Use Context
Track: Systems biology, multi-omics integration, modeling
-
Arif Dönmez, IUF – Leibniz Research Institute for Environmental Medicine, Germany
- Oleksiy Nosov, DNTOX GmbH, Germany
- Kristina Heck, DNTOX GmbH, Germany
-
Ellen Fritsche, SCAHT – Swiss Centre for Applied Human Toxicology / DNTOX GmbH, Switzerland
- Axel Mosig, Ruhr University Bochum / DNTOX GmbH, Germany
-
Katharina Koch, IUF – Leibniz Research Institute for Environmental Medicine / DNTOX GmbH, Germany
Presentation Overview: Show
High-throughput screening (HTS) resources like ToxCast are vital for regulatory research but are often
hindered by their scale (~100 GB) and complex database architectures. We present the ToxCast Evidence Graph, a
lightweight semantic layer that transforms massive screening data into a portable, queryable system. By
filtering the database into assay-scoped SQLite builds, we reduced the footprint to ~3 GB for specific use
cases (e.g., developmental neurotoxicity) while preserving concentration-response data, fitted models,
biological endpoints, and quality flags.
A custom RDF projection exposes these elements as linked entities in GraphDB, enabling structured queries
across chemicals, assays, and potency parameters (AC50). To bridge bioactivity with exposure context, we
integrated CPDat v4.0 functional-use and product-category records. While the graph manages semantic
relationships, dense curve data remains in SQLite, retrieved on-demand for drill-down analysis.
A Streamlit prototype facilitates exploration via a locally deployed, lightweight LLM that translates natural
language into SPARQL. Strictly grounded by a versioned RDF schema, the LLM acts as a query generator rather
than a knowledge source, ensuring data sovereignty and preventing hallucinations.
This architecture provides a transparent, auditable blueprint for regulatory agencies (e.g., BfR) where
traceability and expert review are paramount. By providing a portable ""evidence-in-a-box,"" we demonstrate
how semantic technologies and local AI can democratize access to complex toxicological big data.
C-S.B.59: Ecoli-GEM, an updated and standardized consensus genome-scale metabolic model, shows that E. coli
is proton-saturated
Track: Systems biology, multi-omics integration, modeling
-
Claudia de Buck, Bioprocess Engineering, Laboratory of Systems & Synthetic Biology, Wageningen
University & Research, Netherlands
-
Kennet Lindquist, Department of Biology and Biological Engineering, Chalmers University of Technology,
Sweden
-
Mihail Anton, Department of Biology and Biological Engineering, Chalmers University of Technology, Sweden
-
Mark Bisschops, Bioprocess Engineering, Wageningen University & Research, Netherlands
-
Ruud Weusthuis, Bioprocess Engineering, Wageningen University & Research, Netherlands
-
Maria Suarez Diez, Laboratory of Systems & Synthetic Biology, Wageningen University & Research,
Netherlands
Presentation Overview: Show
Genome-scale metabolic models (GEMs) provide valuable insight into microbial systems. The first GEM of the
model organism Escherichia coli was developed in 2000. The model has been updated several times, arriving at
the most recent version iML1515 (2017). Over the years, several studies published improvements to the
iML1515-model, but these improvements have not been integrated, leading to several parallel versions of
iML1515. We have, therefore, combined previously reported modifications of the iML1515-model to arrive to a
consensus model: Ecoli-GEM.
This consensus model is distributed in a standardized way called Standard-GEM, which is a defined github
repository structure. As such, we have created an efficient way to collaborate on improving the E. coli model,
as parallel versions and model duplications have been merged. This increases the model quality since all
current improvements have been verified, and future alterations can be openly discussed in the repository.
We have used the Ecoli-GEM (and its enzyme-constrained version) to study the biological role of protons in E.
coli, as proton fluxes are linked to the energy metabolism. We show that E. coli is not only nutrient-limited,
but also energy-limited due to proton-saturation as maintaining a neutral cytosolic pH places an energetic
burden on the cell. We have quantified this energetic burden for different carbon sources, in oxic and anoxic
conditions, and we show that the recent insights in the dynamics of formate transporter (FocA) greatly impact
the simulated growth rate. These insights in the dispersion and biological relevance of GEMs will improve
their applicability in metabolic engineering studies.
C-S.B.60: Inferring somatic gene expression evolution from single-cell data
Track: Systems biology, multi-omics integration, modeling
-
Laura Tomás, Department of Biosystems Science and Engineering, ETH Zurich, Basel, Switzerland,
Switzerland
-
Antoine Zwaans, Department of Biosystems Science and Engineering, ETH Zurich, Basel, Switzerland,
Switzerland
-
Daniele Silvestro, Department of Biosystems Science and Engineering, ETH Zurich, Basel, Switzerland,
Switzerland
-
Tanja Stadler, Department of Biosystems Science and Engineering, ETH Zurich, Basel, Switzerland,
Switzerland
Presentation Overview: Show
Understanding how gene expression evolves along somatic cell lineages is central to processes such as
development and tumor progression. Recent single-cell technologies now enable joint measurement of gene
expression and lineage relationships, opening new opportunities to study temporal dynamics. Continuous trait
models, such as Brownian Motion and Ornstein-Uhlenbeck processes, provide a natural framework, yet their
applicability in single-cell contexts remains largely unexplored.
Here, we evaluate these models under realistic scenarios of somatic evolution and introduce an observation
model that explicitly accounts for scRNA-seq noise. Using simulations based on age-dependent branching
processes, we generated lineage trees across biologically motivated scenarios and assessed model recovery
under different selective regimes.
We show that inference is sensitive to transition diagram structure, cell type uncertainty, and errors in tree
reconstruction, often leading to overestimation of selection. Critically, methods that ignore the observation
process are strongly biased by expression noise, often overestimating the effect of selection.
Overall, our results highlight the need for caution when interpreting the results of these models in
single-cell settings, where technical noise, model misspecification, and uncertainty in lineage reconstruction
can strongly influence inferred evolutionary dynamics. Explicitly modeling the observation process is
essential for robust phylogenetic inference from single-cell gene expression data.
C-S.B.61: Microneedle-Based Salivary Lactate Monitoring for Model-Assisted Assessment of Exercise Intensity
and Overtraining Risk
Track: Systems biology, multi-omics integration, modeling
-
Ling Ding, Waseda University, Japan
- XIru Li, Waseda University, Japan
- Kameoka Jun, Waseda University, Japan
Presentation Overview: Show
Non-invasive biomarker monitoring can provide quantitative physiological data for model-assisted health
assessment. Lactate is an important metabolic indicator of exercise intensity, fatigue, and training load.
However, conventional blood lactate testing requires invasive sampling and is unsuitable for frequent
monitoring during exercise and recovery. In this study, we developed a microneedle-based electrochemical
lactate sensor for rapid salivary lactate testing and evaluated its potential for lactate-based exercise
intensity assessment. The working electrode was modified with Prussian Blue, MWCNT-COOH and lactate oxidase
immobilized through EDC/NHS crosslinking to enable enzymatic electrochemical lactate detection. The sensor was
characterized in buffer and human saliva, including calibration performance, detection limit, response
behavior, reproducibility, saliva matrix effects, and signal correction. Sensor-derived salivary lactate
changes were further mapped into a lactate-based interpretation framework to classify low, moderate, and high
exercise-load responses and to identify potential overreaching or overtraining-risk patterns based on elevated
lactate responses and delayed recovery trends. This approach links non-invasive salivary biomarker sensing
with physiological-state classification, supporting rapid point-of-care assessment of exercise intensity and
recovery status. Although larger participant cohorts and additional physiological markers are required for
further validation, this work demonstrates the potential of salivary lactate dynamics as an accessible
biomarker input for model-assisted training-load assessment and personalized health monitoring.
C-S.B.62: A Probabilistic Digital Twin Framework for Uro-Oncology: In Silico Modeling and Virtual
Experimentation of the FGFR3 Pathway
Track: Systems biology, multi-omics integration, modeling
-
Sinan Kolukısaoğlu, Ankara University Faculty of Medicine, Turkey
Presentation Overview: Show
Background:
Adopting a Molecular Pathological Epidemiology perspective, we present a probabilistic in silico Digital Twin
of the FGFR3 pathway. Bridging pathologic epidemiological data with micro-cellular multi-omics, this framework
establishes a dynamic, predictive modeling foundation for uro-oncology to execute virtual experiments.
Methods:
Developed in Python via a ""Grey Box"" network inference approach, the multi-pathway FGFR3 architecture (IP3,
MAPK, JAK/STAT, mTOR) was modeled as a probabilistic electrical circuit to overcome data sparsity. Integrating
systems biology and multimodal data, input voltages were derived from normalized transcriptomic expression of
FGFR3. Signal routing via stochastic Markov Transition Matrices was governed by pathway conductance and
endogenous inhibitors. Epidemiological constraints utilized cBioPortal mutation frequencies (NMIBC cohort),
while final current outputs were predicted by formula and quantified via statistical transcriptomic
footprinting (decoupleR).
Results:
The dynamic model successfully mapped signal fluxes across parallel cascades, calculating probabilistic gene
activations. Virtual experiments simulating targeted interventions computed realistic signal rerouting. At
higher resistances, flux diverted to PLCγ/STAT survival pathways — mathematically recapitulating adaptive
bypass resistance observed clinically upon MAPK/PI3K inhibition in FGFR-driven urothelial carcinoma. These
simulated flux diversions and bypass resistance mechanisms (BRAF-MEK/ERK) demonstrated strong biological
validity in FGFR-targeting drug resistance, aligning with established literature of cancer genetics.
Conclusions:
This biologic framework translates multi-omics and mutational data into dynamic, predictive models. By
rendering the FGFR3 network as an electrical circuit, we provide a literature-consistent proof-of-concept for
systems biology in uro-oncology, offering a computational platform for anticipating bypass resistance and
screening novel therapeutics in silico.
C-S.B.63: Automated lesion segmentation enables large-scale analysis of heterogeneity in multiple
sclerosis
Track: Systems biology, multi-omics integration, modeling
-
Shivam Kumar, University Medical Center Groningen, Netherlands
- Abel Koffeman, University of Groningen, Netherlands
-
Inge Holtman, Department of Biomedical Sciences, University Medical Center Groningen; The Netherlands
Brain Bank, Amsterdam, Netherlands
Presentation Overview: Show
Multiple sclerosis (MS) is characterized by heterogeneous white matter lesions that vary in immune activity,
myelin damage, and spatial organization. This heterogeneity complicates the systematic characterization of
lesion subtypes and their role in disease progression. To address this, our group is generating a large-scale
neuropathological imaging dataset in close collaboration with the Netherlands Brain Bank (NBB), a unique
resource providing high-quality post-mortem brain tissue with minimal post-mortem delay. These high-resolution
histological images capture the full spectrum of lesion variability; however, their analysis remains limited
by the need for expert annotation.
Our core technical contributions are threefold: (1) a patch-based processing strategy that enables efficient
analysis of gigapixel whole-slide images; (2) the use of a vision foundation model, whose rich and
generalizable representations - learned from large-scale image data - transfer effectively to this specialized
histological domain; and (3) a label-efficient pipeline that leverages these pre-trained features to achieve
strong performance using only 100 annotated images from a cohort of ~2,000. Notably, the foundation model
embeddings were sufficiently expressive to capture pathologically relevant structures. The model achieved 85%
accuracy in lesion segmentation. Overall, this embedding-based framework provides a scalable and
data-efficient approach for large-scale analysis and can be extended to other neuropathological
applications.
C-S.B.64: Ion Channel Degeneracy Underlies Pacemaker Identity: Large-Scale Computational Evidence from a
Biophysically Constrained Hodgkin-Huxley Framework
Track: Systems biology, multi-omics integration, modeling
-
Batuhan Safa Kar, Ankara University Faculty of Medicine, Turkey
Presentation Overview: Show
Pacemaker neuronal identity is classically attributed to specific ion channels — most notably HCN and T-type
calcium channels. Here we challenge this view through large-scale computational sampling of a 21-channel,
single-compartment Hodgkin-Huxley model incorporating dynamic Nernst potentials, full calcium and sodium
homeostasis, and biologically constrained parameter bounds.
From 1.5 million simulated cells, 390,022 were validated as pacemakers. Three principal findings emerge.
First, calcium channel subtypes show near-zero inter-channel correlations (mean |r| = 0.004), with coefficient
of variation exceeding 0.93 within any frequency band — demonstrating that calcium channel composition is
degenerate with respect to pacemaker identity. Second, 29.4% of all pacemaker cells fire without any Kv1
conductance, achieving stable rhythmic output through alternative repolarization mechanisms. Third, HCN
conductance correlates with firing frequency at r = 0.097 across 390,022 cells, indicating that HCN is neither
necessary nor sufficient for pacemaker function.
These findings are cross-validated against a human substantia nigra single-nucleus RNA-seq dataset (n = 22,048
dopaminergic neurons; Kamath et al., 2022), where HCN2 expression variance and calcium channel co-expression
patterns are consistent with predicted degeneracy profiles.
Together, these results reframe pacemaker identity as an emergent property of ion channel composition space
— with direct implications for understanding differential neuronal vulnerability in Parkinson's disease and
for future therapeutic target selection.
C-S.B.65: Cluster-Aware Functional Principal Component Analysis for Imputing Missing Values in Longitudinal
Microbiome Data
Track: Systems biology, multi-omics integration, modeling
-
Alireza Dostmohammadi, Department of Bioinformatics and Computational Biophysics, University of
Duisburg-Essen, Essen, Germany, Germany
-
Mohammad Darbalaei, Department of Bioinformatics and Computational Biophysics, University of
Duisburg-Essen, Essen, Germany, Germany
-
Daniel Hoffmann, Department of Bioinformatics and Computational Biophysics, University of Duisburg-Essen,
Essen, Germany, Germany
-
Farnoush Farahpour, Department of Bioinformatics and Computational Biophysics, University of
Duisburg-Essen, Essen, Germany, Germany
Presentation Overview: Show
Missing observations arising from irregular sampling and participant attrition are a major analytical
bottleneck in longitudinal microbiome studies. Existing imputation methods either treat temporal observations
through discrete multivariate representations, without leveraging the underlying continuity of biological
processes, or rely on deep generative models such as GANs and diffusion architectures, which require
substantial training data and can be unstable or difficult to interpret in studies with limited replicates or
sparse time points.
We propose an imputation framework that models each taxon trajectory as a smooth latent function on the
centered log-ratio scale, estimated through Functional Principal Component Analysis (FPCA) in the Principal
Analysis by Conditional Expectation (PACE) framework to handle sparse and irregular sampling. Imputation pools
information from observed time points within each trajectory and from other replicates through the shared FPCA
covariance structure. To address heterogeneity across replicates, an adaptive clustering step based on FPCA
scores is performed locally for each missing observation, so that imputation borrows strength only from
trajectories with comparable temporal patterns. Robustness is further enhanced through Fraiman–Muniz
functional depth, which down-weights anomalous trajectories that would otherwise distort covariance
estimation. Uncertainty is quantified through both analytic and nonparametric bootstrap confidence
intervals.
We evaluate the method on two longitudinal datasets across multiple taxonomic resolutions and missingness
rates ranging from 10% to 70%, under MCAR, MAR, and MNAR mechanisms, benchmarking against DeepMicroGen. The
framework achieves lower mean absolute error at substantially reduced computational cost, with improvements
most pronounced at high missingness rates.
C-S.B.66: Characterizing the Clinical and Functional Correlates of Age of Onset in 5729 Mendelian
Diseases
Track: Systems biology, multi-omics integration, modeling
-
Aybuge Altay, Berlin Institute of Health at Charité (BIH), Germany
- Peter Robinson, Berlin Institute of Health at Charité (BIH), Germany
- Peter Hansen, Berlin Institute of Health at Charité (BIH), Germany
- Kyran Wissink, University of Amsterdam, Netherlands
Presentation Overview: Show
The age at which Mendelian diseases first manifest varies widely, ranging from prenatal to adult presentation,
yet the biological and clinical determinants of this variability remain poorly understood. Although individual
disorders have provided valuable clues, a systematic analysis across the Mendelian landscape has been
lacking.
We conducted a comprehensive analysis using curated age-of-onset annotations from the Human Phenotype Ontology
(HPO). Mendelian diseases were grouped into four onset categories: congenital, neonatal, pediatric, and adult.
We assessed the enrichment of 827 phenotypic features (HPO terms), 596 gene functions (Gene Ontology), and
modes of inheritance across these categories. Furthermore, we trained a Random Forest classifier to predict
disease onset based on clinical phenotypes, achieving an F1-score of 0.71 across the four categories.
Our results reveal distinctive phenotypic signatures: congenital onset was uniquely characterized by
developmental anomalies such as Ventricular septal defect, while adult-onset diseases were enriched for
reproductive features and late-stage neurological manifestations. Statistical analysis showed that adult-onset
diseases were substantially enriched for autosomal recessive inheritance, whereas autosomal dominant
inheritance was comparatively underrepresented. The supervised model highlighted that morphological and
developmental abnormalities served as the most informative features for onset prediction.
These results establish HPO-curated onset annotations as a valuable resource for elucidating relationships
among disease manifestation, clinical phenotype, inheritance patterns, and gene function. Together, our
findings provide a unified framework for understanding temporal variation in Mendelian disease presentation
and may support future efforts in automated disease characterization.
C-S.B.67: A Biochemical Reaction Network Model of Autophagy-Apoptosis Crosstalk Under Metabolic
Stress
Track: Systems biology, multi-omics integration, modeling
-
Krisztian Szuppinger, Pazmany Peter Catholic University, Hungary
- Zita Ruszinko, Pazmany Peter Catholic University, Hungary
- Bence Hajdu, Institut Curie, Hungary
- Tibor Nagy, HUN-REN Research Centre for Natural Sciences, Hungary
- Orsolya Kapuy, Semmelweis University, Hungary
Presentation Overview: Show
Cell death mechanisms, such as autophagy and apoptosis, play critical roles in cellular homeostasis and
disease. Autophagy is an evolutionarily conserved mechanism that remains active under homeostatic conditions
and supports cell survival by degrading damaged or obsolete components. It is closely linked to programmed
cell death (apoptosis), and disruption of their crosstalk is implicated in neurodegeneration, inflammation,
cancer, and aging. This work provides a computational framework to study regulatory interactions between these
processes.
We introduce two new inputs to our existing, basally accurate biochemical reaction network model: glucose
starvation and treatment with the autophagy inducer rapamycin. Based on mass-action kinetics, the model
describes 175 interactions among 114 species using ordinary differential equations implemented in Cantera,
with parameter optimization via the FOCTOPUS algorithm in the Optima++ environment. Simulations are validated
against experimental data.
The starvation module reproduces glucose dynamics accurately and incorporates mechanistic and structural
features of the energy sensor AMPK, enabling differential activation by distinct adenosine nucleotides. The
model predicts two-fold higher ADP-mediated than AMP-mediated activation during early starvation, a
distinction often not captured by comparable models. The rapamycin module implements passive transport and
accounts for protein-protein interactions using literature-derived reaction rates. The refined model
reproduces the dynamics between AMPK and the autophagy suppressor mTORC1 and initiator ULK1.
Overall, the new modules accurately simulate glucose starvation and rapamycin treatment, facilitating
inference of real-world effects from in silico simulations. Our work therefore bridges computational systems
biology and biomedical research, providing a foundation for patient-specific simulations, supporting future
applications in personalized medicine.
C-S.B.68: Mutation-driven reorganization of proteomic covariation networks
Track: Systems biology, multi-omics integration, modeling
-
Seokjin Ham, Spanish National Cancer Research Center (CNIO), Spain
- Solip Park, Spanish National Cancer Research Center (CNIO), Spain
Presentation Overview: Show
Background. Cellular proteomes encode regulatory relationships not only through physical interactions but also
via coordinated variation in protein abundance. While positive covariation reflects cooperative processes,
negative covariation may capture inhibitory or compensatory constraints. However, how somatic mutations
reshape these network-level dependencies remains largely unexplored.
Approach. We integrated genomic and proteomic profiles from 1,067 CPTAC tumors and analyzed
mutation-associated changes in protein covariation across more than 20,000 proteins. Using regression models
with mutation–interaction terms, we quantified how missense variants alter correlation structures and
assessed the utility of covariation patterns for scalable AI-based interaction prediction.
Results. Positive associations were widespread (~69,000 pairs at |r|>0.5, adj. p<0.05), whereas negative
associations were comparatively rare (~580 pairs). Somatic mutations perturbed both types of relationships,
with negative dependencies disproportionately affected. Pathogenic variants induced a stronger attenuation of
correlations (median −0.30) than benign variants (−0.16). Notably, pathogenic mutations accounted for the
largest fraction of significant interaction changes among negatively correlated pairs (~21%), suggesting
selective vulnerability of inhibitory links.
Conclusion. These results indicate that cancer involves a mutation-driven disruption of proteomic regulatory
balance, particularly affecting inhibitory interactions. The resulting covariation signatures provide
mechanistic signals that can inform next-generation AI models for inferring interaction states from
large-scale proteomic data.
C-S.B.69: Joint Representation Learning and Graph Construction for Multimodal Patient Similarity
Networks
Track: Systems biology, multi-omics integration, modeling
-
Emilia Agasi, School of Informatics, The University of Edinburgh, United Kingdom
-
Charlie Gourley, The University of Edinburgh, Institute of Genetics and Cancer,, United Kingdom
- Ian Simpson, School of Informatics, The University of Edinburgh, United Kingdom
Presentation Overview: Show
High-grade serous ovarian carcinoma (HGSOC) presents substantial inter-patient heterogeneity at morphological,
molecular and clinical levels. No current method delivers reliable patient stratification for prognosis or
treatment selection. Patient similarity networks (PSNs) combined with graph-based learning offer a way to
model relational structure across patients. However, their effectiveness depends on the quality of the
underlying patient representations and the strategies used to construct the graphs. Representations from
generic pre-trained models often fail to capture task-relevant features, and commonly used similarity metrics
do not reflect cross-modal relationships. We present a framework for multimodal PSN construction that
addresses these limitations. We refine feature representations from high-dimensional modalities, including
histopathology, to better capture domain-specific structure, and integrate them with complementary clinical
and molecular modalities. Modalities are combined using similarity network fusion alongside alternative
data-driven fusion strategies. We systematically compare similarity measures and graph construction techniques
to assess their effect on downstream learning, and use the resulting networks to train graph-based models for
survival prediction. We expect that improved representations and more informed graph construction will yield
PSNs that better capture clinically meaningful relationships between patients, improving survival prediction
in HGSOC. By moving away from task-agnostic features and heuristic similarity definitions, our approach aims
to generalise to other heterogeneous disease settings where patient stratification remains a challenge.
C-S.B.70: Deep Learning for BioImaging: What Are We learning ?
Track: Systems biology, multi-omics integration, modeling
- Ivan Svatko, Ecole Normale Supérieure Paris Science et Lettres, Ukraine
-
Maxime Sanchez, Ecole Normale Supérieure Paris Science et Lettres - Institut Curie, France
- Ihab Bendidi, Ecole Normale Supérieure Paris Science et Lettres, France
- Auguste Genovesio, Ecole Normale Supérieure Paris Science et Lettres, France
Presentation Overview: Show
Recent advances in representation learning have transformed natural image analysis, yet their impact on
biological microscopy remains poorly understood. In this work, we systematically investigate what deep
learning models actually learn from large-scale bioimaging data, focusing on two key biological contexts: cell
culture imaging and tissue histology.
We benchmark a wide range of representations, including state-of-the-art pretrained vision models,
domain-specific foundation models, randomly initialized networks, and simple handcrafted features. Across
multiple datasets and tasks, we show that surprisingly simple or untrained representations can achieve
performance comparable to advanced pretrained models. This suggests that current benchmarks may not reliably
capture biologically meaningful representations, but instead exploit low-level visual cues or dataset-specific
biases.
To better understand these behaviors, we analyze representation structure through dimensionality, layer-wise
contributions, and robustness across biological tasks. We find that high-performing models often rely on
low-dimensional signals and that shallow features can outperform deeper representations, challenging common
assumptions derived from natural image domains. Furthermore, we demonstrate that biologically interpretable
structure-only baselines, constructed without cell-level information, can remain competitive, reinforcing
concerns about the validity of existing evaluation protocols.
Overall, our results highlight a critical gap between benchmark performance and biological relevance in
bioimaging. We advocate for more rigorous evaluation strategies and stronger baselines to ensure that learned
representations capture meaningful biological mechanisms rather than spurious correlations. This work provides
practical guidelines for the development and assessment of deep learning models in computational biology and
high-content imaging.
C-S.B.71: β-Cell State Transitions Reveal Metabolic Rewiring and Adaptive Stress Programs in Type 2
Diabetes
Track: Systems biology, multi-omics integration, modeling
-
Malvika Sudhakar, Novo Nordisk Foundation Center for Basic Metabolic Research, University of
Copenhagen, Denmark, Denmark
-
Maria Fernandes, Novo Nordisk Foundation Center for Basic Metabolic Research, University of Copenhagen,
Denmark, Denmark
-
Jordi Merino, Novo Nordisk Foundation Center for Basic Metabolic Research, University of Copenhagen,
Denmark, Denmark
Presentation Overview: Show
Pancreatic β-cell dysfunction, characteristic of type 2 diabetes (T2D), involves transitions across distinct
cellular states preceding failure, yet the mechanisms driving these transitions remain poorly understood.
Although single-cell RNA sequencing (scRNA-seq) studies have enabled the identification of -cell
subpopulations, they lack large sample sizes and prediabetes representation, constraining insight into disease
progression across the spectrum of dysglycemia. To address this gap, we integrated five scRNA-seq datasets to
(i) map β-cell state dynamics across non-diabetic, pre-T2D, and T2D, and (ii) to define molecular programs
underlying state transitions. We identified eight transcriptionally distinct β-cell clusters spanning three
functional axes related to insulin (INS) secretion, amylin (IAPP) secretion, and cellular stress responses.
INS-low clusters were enriched in T2D, while a metalloprotein cluster was enriched in non-diabetic, consistent
with early protective or compensatory activity. In addition, we identified two distinct IAPP-expressing
clusters aligned with either INS-high or INS-low states, delineating divergent programs of functional
maintenance versus stress adaptation and metabolic rewiring. Trajectory inference revealed a continuum from
functional to dysfunctional β-cell states, highlighting progressive changes in transcriptional programs
related to calcium ion channel activity and GABAergic signaling pathways. These findings suggest adaptive
mechanisms that may support β-cell survival despite functional decline. Integration with genome-wide
association study data showed that T2D-enriched β-cell clusters are significantly associated with genetic risk
for T2D. In summary, we uncover a novel spectrum of β-cell states linking functional identity, stress
adaptation, and genetic risk, refining our understanding of coordinated metabolic and signaling rewiring
underlying β-cell dysfunction and persistence in T2D.
C-S.B.72: Critical Assessment of Metagenome Interpretation: Round Three of Metagenomic Software Benchmarking
Challenges
Track: Systems biology, multi-omics integration, modeling
-
Fernando Meyer, Helmholtz Centre for Infection Research, Braunschweig, Germany, Germany
-
Zhi-Luo Deng, Helmholtz Centre for Infection Research, Braunschweig, Germany, Germany
-
Philipp Muench, Helmholtz Centre for Infection Research, Braunschweig, Germany, Germany
-
Hesham Almessady, Helmholtz Centre for Infection Research, Braunschweig, Germany, Germany
-
Gary Robertson, Helmholtz Centre for Infection Research, Braunschweig, Germany, Germany
- Liren Huang, Bielefeld University, Bielefeld, Germany, Germany
- David Koslicki, Penn State University, University Park, PA, USA, United States
- Alexander Sczyrba, Bielefeld University, Bielefeld, Germany, Germany
-
Alice C. McHardy, Helmholtz Centre for Infection Research, Braunschweig, Germany, Germany
Presentation Overview: Show
Selecting appropriate software and parameter settings for processing shotgun metagenomic data is essential for
accurate analysis and interpretation. This remains challenging given the growing number of available
bioinformatics tools. The Critical Assessment of Metagenome Interpretation (CAMI) is a community-driven
initiative that provides standardized, unbiased evaluations of metagenomic methods, including assemblers,
taxonomic profilers, binners, and pathogen detection tools. Community participation is driven by the provision
of benchmarking challenges, in which participants apply their methods on novel datasets and submit their
results for independent evaluation. Previous work by the community established widely adopted datasets,
metrics, and formats, guiding both method applications for data analysis and their development. In 2026, CAMI
III challenges will be launched, covering datasets from different high and low microbial biomass body sites,
reflecting the latest sequencing technologies and introducing new evaluation categories. As before, microbial
communities will cover multiple domains, i.e., Bacteria, Eukaryotes, Archaea, and Viruses, and also include
plasmids. To support method development, testing of scalability, and familiarization with CAMI formats ahead
of the upcoming benchmarking challenges, a preparatory ("toy") human gut dataset comprising paired short- and
long-read shotgun metagenomes was recently released. These data were modeled to represent 20 longitudinal
samples across 10 individuals. In contrast to the toy dataset, challenge datasets will include unpublished
data, which is essential for informative benchmarking of reference-based methods. Further information is
available at https://cami-challenge.org/.
C-S.B.73: Module Graph: A Web-Based Framework for Module-Centred Metabolic Network Analysis—Bridging KEGG
Orthologs and Functional Metabolic Organisation
Track: Systems biology, multi-omics integration, modeling
- Run Jie Xia, University Ca Foscari of Venice, Italy
- Victoria Grosu, University Ca Foscari of Venice, Italy
-
Mariana Reyes-Prieto, Foundation for the Promotion of Sanitary and Biomedical Research of the Valencia
Region(FISABIO), Spain
-
Merce Llabres, University of the Balearic Islands, Spain
- Marta Simeoni, University Ca Foscari of Venice, Italy
Presentation Overview: Show
Metabolism comprises the network of chemical reactions that sustains life, including nutrient processing,
energy production, and the synthesis of essential cellular components. Due to its complexity, metabolism is
typically represented as a network of pathways, where reactions are grouped into functional units that perform
specific biological tasks. A major resource for studying metabolism is the Kyoto Encyclopedia of Genes and
Genomes (KEGG), which organises metabolic knowledge into pathway maps linking genes—represented as KEGG
Orthologs (KOs)—through the reactions they catalyse. Although KO annotations derived from sequencing data
are essential for identifying the functional potential of organisms, they are often difficult to interpret in
isolation and do not directly reveal which metabolic functions are active.
Between pathways and individual KOs, KEGG defines modules as intermediate functional units. These modules
correspond to smaller, well-defined sub-pathways composed of specific combinations of KOs and associated
reactions. Owing to their finer granularity and reusability, modules provide a promising framework for
representing metabolism as a set of functional building blocks. However, modules remain underutilised, as
current KEGG tools are largely pathway-centric and present them mainly as static diagrams rather than
structured, network-based representations.
To address this limitation, this work introduces Module Graph, a web-based tool that generates module graphs
from a set of KOs. In this representation, modules are nodes connected when they share active compounds,
enabling the visualisation, analysis, and comparison of metabolic organisation across organisms and
metagenomic samples while preserving biological interpretability.
C-S.B.74: Risk-averse optimization of genetic circuits under uncertainty
Track: Systems biology, multi-omics integration, modeling
-
Michal Kobiela, University of Edinburgh, United Kingdom
- Diego A. Oyarzun, The University of Edinburgh, United Kingdom
- Michael U. Gutmann, The University of Edinburgh, United Kingdom
Presentation Overview: Show
Engineering biological systems with specified functions requires navigating an extensive design space, which
is challenging to achieve with wet-lab experiments alone. To expedite the design process, mathematical
modeling is typically employed to predict circuit function in silico ahead of implementation, which, when
coupled with computational optimization, can be used to automatically identify promising designs. However,
circuit models are inherently inaccurate, which can result in suboptimal or non-functional in vivo
performance. To mitigate this, we propose combining Bayesian inference, Thompson sampling, and risk management
to find optimal circuit designs. Our approach employs data from non-functional designs to estimate the
distribution of model parameters and then employs risk-averse optimization to select design parameters that
are expected to perform well, given parameter uncertainty and biomolecular noise. We illustrate the approach
by designing adaptation circuits and genetic oscillators using real and simulated data, with models of varied
complexity.
C-S.B.75: From Code to Response: A Modular Data Infrastructure for Longitudinal SARS-CoV-2 Immune
Surveillance
Track: Systems biology, multi-omics integration, modeling
-
Ceilidh Welsh, Bioinformatics and Biostatistics, The Francis Crick Institute, London; Department of
Zoology, University of Cambridge, United Kingdom
-
David Greenwood, Bioinformatics and Biostatistics, The Francis Crick Institute, 1 Midland Road, London NW1
1AT, United Kingdom
-
Mary Y Wu, Viral & Immune Surveillance Platform, The Francis Crick Institute, 1 Midland Road, London,
NW1 1AT, United Kingdom
-
David Lv Bauer, The Francis Crick Institute, 1 Midland Road, London NW1 1AT; Genotype-to-Phenotype 2
Consortium (G2P2-UK), United Kingdom
-
Edward J Carr, The Francis Crick Institute; UCL Centre for Kidney and Bladder Health, Division of
Medicine, UCL, United Kingdom
-
Emma Wall, The Francis Crick Institute; NIHR BRC, UCLH; Centre for Immunobiology and Infection, Blizard
Institute, QMUL, United Kingdom
Presentation Overview: Show
The rapid generation of heterogeneous biomedical research data during the SARS-CoV-2 pandemic demonstrated
some of the challenges integrating multi-modal datasets required for rapid-response decision-making. In
particular, it highlighted the need for research data systems capable of harmonising longitudinal clinical,
serological, and experimental assay data to support viral and immune surveillance.
In response, we developed a modular data infrastructure that underpins our ongoing virology research for the
UCLH Crick Legacy study (REC20/HRA/4717). This observational cohort began in January 2021 to characterise
SARS-CoV-2 infection- and vaccine-induced immune responses. The infrastructure supports the ingestion and
integration of participant-level data spanning serological measurements, clinical metadata, recorded
vaccination, and experimental assay outputs conducted repeatedly over years.
Our data model enforces harmonisation across modalities, enabling linkage of longitudinal records at the
participant and sample levels. Time-indexing and data alignment is performed using our publicly-available R
package, chronogram. Data integration workflows standardise assay outputs and clinical variables into a
unified schema, supporting version-controlled updates as new data are acquired.
Our system provides structured data outputs to support downstream analyses that investigate the ongoing
effects of the pandemic on immunity, reducing the gap between hypothesis generation and computational
implementation. As a result, it has supported analyses within the Legacy study and collaborating consortia,
including seasonal vaccine monitoring and real-time estimates of immune response against SARS-CoV-2
variants.
Beyond Legacy, this research infrastructure continues to develop with pandemic preparedness in mind, and
presents a framework for developing modular data architectures that can support future infectious disease
surveillance and response.
C-S.B.76: Systematic Discovery of Alternative Splicing–Derived Neoepitopes Using Integrative MHC
Presentation and Immunogenicity Prediction
Track: Systems biology, multi-omics integration, modeling
-
GülÅŸen Eymen Dediler, FHNW - SIB Member, Switzerland
- Abdullah Kahraman, FHNW-SIB Group Leader, Switzerland
Presentation Overview: Show
Neoepitope prediction is central to cancer immunotherapy, yet current computational pipelines often generate
large candidate sets with limited biological interpretability, largely due to the neglect of transcript
isoform diversity and redundancy across highly similar gene families. Consequently, the biologically relevant
neoepitope landscape remains obscured by substantial noise.
Here, we present a transcript- and isoform-aware framework for neoepitope discovery that integrates
predictions from MHCflurry, NetMHCpan, and PRIME across globally prevalent HLA class I alleles. Using peptide
repertoires spanning 8–11 amino acids derived from the Genotype-Tissue Expression (GTEx) Project dataset, we
constructed a large-scale, transcript-aware neoepitope database and applied a consensus-based prioritization
strategy.
Initial integration yielded over 53,000 candidate peptide–HLA pairs. Incorporating most-dominant transcript
(MDT) information as part of the prioritization pipeline reduced this set to 58 candidates, reflecting the
combined effect of transcript-level filtering and multi-model consensus. These high-confidence candidates
formed a consistent, non-redundant core, predominantly originating from a single gene, CYFIP2.
In contrast, the broader candidate space was dominated by immune receptor gene families, including KIR and
immunoglobulin loci, which inflated apparent neoepitope diversity through sequence redundancy and
multi-transcript mapping. After removing this redundancy, we identified 3,380 uniquely mapped candidate
peptides from biologically relevant genes.
Collectively, our results define a three-layer structure of the predicted neoepitope landscape: a large
redundant background, an intermediate set of uniquely mapped candidates, and a minimal high-confidence core
defined by transcript dominance. This framework provides a principled strategy for reducing false positives
and improving the biological relevance of neoepitope prioritization.
C-S.B.77: A leakage-aware benchmark for gene prioritization across feature representations and integration
strategies
Track: Systems biology, multi-omics integration, modeling
-
Ziyu Zhang, Department of Informatics, University of Oslo, Norway
-
Fatemeh Ghorbani, School of Electrical and Computer Engineering, College of Engineering, University of
Tehran, Iran
- Ole Christian Lingjærde, Department of Informatics, University of Oslo, Norway
- Pooya Zakeri, Department of Informatics, University of Oslo, Norway
Presentation Overview: Show
Gene prioritization (GP) methods are increasingly used to rank candidate disease genes from heterogeneous
biological data, but reported performance remains difficult to compare across studies because methods are
evaluated using different feature sets, training protocols, and metrics, and may also be affected by
information leakage. We present a unified, time-aware benchmarking study that evaluates classical and modern
machine-learning approaches for GP, including kernel methods and deep neural networks. We further benchmark
multiple gene-level feature modalities under a prospective evaluation protocol and compare data integration
strategies at multiple levels across both deep-learning and kernel-based settings. To better reflect practical
discovery settings, we evaluate performance using metrics that emphasize both global ranking and early
retrieval, and introduce BioSim, a biological enrichment-based similarity metric that assesses the functional
coherence between top-ranked unlabeled genes and known disease genes.
Across 48 diseases from 11 ICD-10 categories, our benchmark shows that classical methods such as SVMs remain
highly competitive with deep neural networks, and that feature type and feature quality strongly influence
model performance. Early fusion can produce mixed results, whereas intermediate fusion in both deep-learning
and kernel-based settings, particularly geometric kernel fusion, yields the most consistent gains across
ranking metrics. For example, in Diseases of the Blood and Certain Disorders, intermediate fusion
substantially improved early retrieval, more than doubling Recall@150 relative to the best corresponding
non-fused setting. Late fusion, meanwhile, more often gives stronger BioSim performance. Overall, this work
provides a leakage-aware benchmark and practical guidance for evaluating GP methods in realistic discovery
scenarios.
C-S.B.78: Novel Software to Analyze Long-term Potentiation Recordings
Track: Systems biology, multi-omics integration, modeling
-
Mohammadamin Beheshti Dehkordi, University of Eastern Finland, Finland
-
Mireia Gómez-Budia, A. I. Virtanen Institute for Molecular Sciences, University of Eastern Finland,
Kuopio. , Finland
-
Anssi Pelkonen, A. I. Virtanen Institute for Molecular Sciences, University of Eastern Finland, Kuopio. ,
Finland
-
Tarja Malm, A. I. Virtanen Institute for Molecular Sciences, University of Eastern Finland, Kuopio.,
Finland
-
Luca Giudice, A. I. Virtanen Institute for Molecular Sciences, University of Eastern Finland, Kuopio.,
Finland
Presentation Overview: Show
Right at the moment while we are learning or forming new information, our brain generates fascinating synaptic
signals called Long-Term Potentiation (LTP), a neurophysiological process underlying learning and memory
formation. During this process, the brain generates synaptic signals that can be recorded as time-series
waveforms using techniques such as Microelectrode Arrays (MEA). These signals, particularly the first field
excitatory postsynaptic potential (fEPSP), rapidly change in response to stimulation, pharmacological
intervention, or subtle pathological alterations, thereby capturing biologically meaningful information
regarding synaptic strength and temporal dynamics.
However, current LTP analysis methods suffer from major limitations. Most studies rely on oversimplified
analytical approaches without validating whether waveform changes truly represent synaptic activity or
drug-specific effects. Synaptic regions are often manually selected, while statistical methods are frequently
applied without verifying assumptions such as data normality. Furthermore, standardized quality control
procedures for detecting artifacts, noise, waveform heterogeneity, or batch effects are often lacking.
To address these challenges, we developed LTP Analysis Software, a standardized platform for LTP waveform
analysis. The platform performs waveform preprocessing, synaptic region detection, noise removal, artifact
detection, unsupervised clustering, feature extraction, and interpretable supervised machine learning.
We evaluated the software using electrophysiological recordings from over 50 idiopathic Normal Pressure
Hydrocephalus (iNPH) patients, a cohort in which approximately 50% of cortical biopsies exhibit early
Alzheimer's Disease (AD)-related pathology. Using Random Forest classification with group cross-validation,
the framework achieved 92% accuracy in pathology prediction, suggesting that different pathological conditions
are associated with distinct LTP waveform signatures.