View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

B-G.C.01: Computational Insights into the pH-Dependent Behavior of Ipilimumab-CTLA-4
Track: General computational biology
  • Wanda Destiarani, Doctoral Program in Biology, University of Tsukuba, Japan
  • Kowit Hengphasatporn, Center for Computational Sciences, University of Tsukuba, Japan
  • Yasuteru Shigeta, Center for Computational Sciences, University of Tsukuba, Japan
  • Ryuhei Harada, Center for Computational Sciences, University of Tsukuba, Japan


Presentation Overview: Show

Therapeutic antibodies face on-target/off-tumor toxicity, causing adverse effects such as cardiotoxicity, skin rashes, and organ inflammation. To mitigate these challenges, pH-dependent antibodies have been engineered to preferentially bind in the acidic tumor microenvironment while reducing interactions under physiological pH conditions. Building on experimental work that generated Ipilimumab variants (Ipi95, Ipi105, Ipi106) through charged amino acid substitutions in complementarity-determining regions, we employed molecular dynamics simulations to examine their interactions with CTLA-4 in physiological and acidic conditions. All variants exhibited enhanced binding affinity at acidic pH, with a reasonable agreement between computational and experimental binding free energies (R² = 0.7736; Pearson's r = 0.8795, p = 0.0039; Spearman's ρ = 0.8333, p = 0.0102). Statistical analysis revealed notable differences across conditions, most notably for Ipi95, which demonstrated the highest degree of pH sensitivity. Although no major global structural changes were observed between conditions, our simulations revealed distinct local energetic rearrangements and residue-level interaction changes at the binding interface. Decomposition analysis on binding energy further indicated that the overall antigen-binding mode was maintained, whereas the introduced charged residues modulated local interaction strengths. These results provide mechanistic insights into how targeted mutations modulate pH-dependent recognition, offering a framework for the rational design of safer therapeutic antibodies.

B-G.C.02: termal: Interactive Exploration of Multiple Sequence Alignments in the Terminal
Track: General computational biology
  • Thomas Junier, Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

Multiple sequence alignments (MSAs) are central to several subfields of
bioinformatics, including phylogenetics, comparative genomics, and sequence
conservation analysis. Such analyses are often carried out on HPC clusters
accessed through SSH sessions, where graphical tools can be unavailable or
impractical, making interactive inspection of alignments difficult.

To address this issue, we developed Termal, an interactive MSA viewer designed
for use in terminals. Termal is entirely keyboard-driven and, like graphical
alignment viewers, provides scrolling, residue colouring, consensus display,
conservation plots, and sequence sorting by properties such as similarity to the
consensus or ungapped length. It also supports zoomed-out views for rapid
inspection of global alignment structure, as well as regular-expression searches
in both headers and sequences.

Termal is implemented in Rust and efficiently handles large alignments. For
example, an alignment of more than 15,000 sequences x 1,500 columns loads in
well under a second on an standard laptop hardware.

Termal is distributed as standalone binaries for Linux, macOS, and Windows. The
program has no external runtime dependencies and integrates naturally into
existing bioinformatics workflows.

By combining interactivity, portability, and performance in a terminal-native
interface, termal provides a practical solution for researchers working on
remote or resource-constrained systems.

B-G.C.03: A Comparative Study of QSPR Methods on a Unique Multitask PAMPA dataset
Track: General computational biology
  • Adam Arany, ESAT-STADIUS, KU Leuven, 3001 Leuven, Belgium, Belgium
  • Andras Formanek, ESAT-STADIUS, KU Leuven, 3001 Leuven, Belgium, Belgium
  • Anna Vincze, Dept. Pharm. Chem., Semmelweis University, HÅ‘gyes Endre Street 7-9, H-1092 Budapest, Hungary, Hungary
  • Richard Bicsak, Depart. Chem. and Environ. Proc. Engineering, BUTE, Műegyetem rkp. 3, H-1111 Budapest, Hungary, Hungary
  • Gyorgy Tibor Balogh, Dept. Pharm. Chem., Semmelweis University, HÅ‘gyes Endre Street 7-9, H-1092 Budapest, Hungary, Hungary
  • Yves Moreau, ESAT-STADIUS, KU Leuven, 3001 Leuven, Belgium, Belgium


Presentation Overview: Show

We present a unique, multitask dataset comprising 143 drug and drug candidate molecules,
each evaluated on in vitro, parallel artificial-membrane permeability assays (PAMPA) using six different model membranes.
Using this resource, we systematically assess the effectiveness of various molecular descriptors and regression models in predicting passive membrane permeability.
The studied models range from simple linear regression to a modern pre-trained transformer architecture.
Particular attention is given to the trade-off between predictive performance and model interpretability, highlighting the challenges introduced by machine learning approaches.
To our knowledge, this is the most comprehensive study on simultaneous modeling of multiple organ-specific PAMPA membranes to date,
offering novel insights into membrane-specific permeability profiles.

We found that expert-designed physico-chemical property descriptors are more fitting for a limited sample size permeabilty study than deep learning based representations.

B-G.C.04: Repurposing blood whole-genome sequencing to study leaky gut
Track: General computational biology
  • Ulas Isildak, Leibniz Institute on Aging - Fritz Lipmann Institute (FLI), Germany
  • Handan Melike Donertas, Leibniz Institute on Aging - Fritz Lipmann Institute (FLI), Germany


Presentation Overview: Show

Intestinal barrier dysfunction (""leaky gut"") has been linked to ageing and chronic disease, but evidence of its prevalence at the population scale remains scarce. Human whole-genome sequencing (WGS) routinely yields reads that fail to map to the host genome and are otherwise discarded. We asked whether these unmapped reads can be repurposed as a signal of gut leakiness. To assess the sensitivity of our approach, we spiked sterile blood with a mock microbial community and identified a detection limit of roughly 100 cells/mL for kraken2-based classification, with measurable background in negative controls. In a preliminary analysis of a pilot cohort from the UK Biobank, following normalisation and filtration informed by these sensitivity estimates, we detected putative gut-associated microbial presence in approximately 37% of individuals. Having established the workflow and characterised its detection sensitivity, we aim to investigate whether microbial content in blood WGS is associated with markers of systemic inflammation, frailty indices, and age-related phenotypes and diseases. Together, these results suggest that unmapped reads from existing human WGS data may offer a useful proxy for microbial translocation.

B-G.C.05: Classification of Metabolic Reaction Graphs using Graph Transformers
Track: General computational biology
  • Aniello Di Vaio, Universitá di Bologna, Italy
  • Mercè Llabrés, University of the Balearic Islands, Spain
  • Jairo Rocha, University of the Balearic Islands, Spain


Presentation Overview: Show

We show the effectiveness of Graphormer, a graph transformer, for the classification of biological reaction graphs associated with different species. We compare the results with graph kernels, a more traditional network analyzer.
Reaction graphs derived from metabolic processes compiled by KEGG provide a compact and structured representation of the biochemical activity of an organism. These graphs often exhibit species-specific patterns, making them a valuable resource for
comparative analysis and classification tasks in bioinformatics.
Traditional graph-based machine learning approaches, such as graph kernels or graph neural networks, have been used to analyse these structures, but they may struggle to capture long-range dependencies and complex global patterns.
Recently, transformer-based architectures have emerged as powerful models capable of learning rich representations from structured data. Graph Transformers extend the self-attention mechanism to graph domains, enabling the integration of both local and global structural information.
In this context, graph transformers offer a promising approach for encoding reaction graphs into meaningful vector representations that can be used for downstream tasks such as classification. In this study, we use Graphormer as a fixed encoder and combined it with a species classification neural network head. The reaction graphs are subsampled when they exceed the usual input size for Graphormer.
This work seeks to provide insights into the suitability of graph transformers for bioinformatics applications involving structured biological data. The results demonstrate that these models can outperform or complement more traditional graph-based approaches, thereby contributing to the broader use of advanced AI techniques in biological network analysis.

B-G.C.06: A Tissue-Independent Landscape of Macrophage Dynamics in Regeneration
Track: General computational biology
  • Hanane Moha Ouchane, Institute of Molecular Medicine and Experimental Immunology, University Hospital Bonn, Germany
  • Ali Enver Bilecen, German Centre for Neurodegenerative Diseases, University Hospital Bonn, Germany
  • Ozgun Gokce, German Centre for Neurodegenerative Diseases, University Hospital Bonn, Germany
  • Zeinab Abdullah, Institute of Molecular Medicine and Experimental Immunology, University Hospital Bonn, Germany


Presentation Overview: Show

Macrophages are key regulators of tissue regeneration across organs and species. Experimental depletion of macrophages consistently results in severely impaired regeneration across diverse biological contexts. The macrophage regenerative response is commonly described as a two-phase response: an early pro-inflammatory M1 phase that promotes debris clearance, progenitor cell recruitment, and proliferation, followed by a later anti-inflammatory M2 phase that supports differentiation, angiogenesis, and extracellular matrix remodeling.
However, this M1/M2 polarization framework is an oversimplification that fails to capture the dynamic spectrum of macrophage states observed in vivo. Moreover, current understanding of macrophage activation dynamics is largely based on independent studies of individual organs and species, making it difficult to identify generalizable patterns. Identifying such common patterns across tissues and species is key to revealing the macrophage programs that drive successful regeneration and could help guide strategies to improve outcomes in tissues with limited regenerative capacity.
Here, we aim to characterize macrophage regenerative programs and pathological states that impair regeneration in an organ- and species-independent manner. We first construct a mouse regeneration atlas by integrating publicly available and in-house single-cell and single-nucleus RNA sequencing datasets spanning 11 organs, capturing the temporal dynamics of regeneration and including both successful and unsuccessful regenerative outcomes. We then apply disentanglement learning to generate a tissue-independent latent embedding of cell states, thereby separating organ-specific transcriptional signatures from conserved regenerative programs and shared pathological states across tissues. Finally, we aim to extend this framework toward a cross-species characterization of macrophage responses to identify conserved principles of regeneration.

B-G.C.07: Semantic annotation for artifact detection in spatial proteomics using DINOv3 as a vision foundation model
Track: General computational biology
  • Sviatoslav Kharuk, European Molecular Biology Laboratory (EMBL), Germany
  • Matthias Meyer-Bender, European Molecular Biology Laboratory (EMBL), Germany
  • Wolfgang Huber, European Molecular Biology Laboratory (EMBL), Germany


Presentation Overview: Show

Our understanding of tissue biology has been transformed in recent years by highly multiplexed imaging that can capture the spatial organization of proteins in tissues. This has also made it possible to collect tissue images at a resolution that was previously unattainable. This kind of data can be produced using a number of technologies, including CyCIF, MIBI, CODEX, and IMC. However, tissue folds, detritus, and antibody aggregates can occur as a result of improper sample handling during sample preparation. These artifacts can affect downstream analysis and make it more complicated. I use the vision foundational model DINOv3 to learn representations from highly multiplexed images. I examine whether this model can effectively learn to identify and detect artefacts in highly multiplexed images providing automated quality control.

B-G.C.08: BGC2NP-CLIP: Contrastive Cross-Modal Retrieval Between Biosynthetic Gene Clusters and Natural Products
Track: General computational biology
  • Anastasiia Kolchina, Saarland University, Germany
  • Olga Kalinina, Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Germany
  • Dietrich Klakow, Saarland University, Germany


Presentation Overview: Show

Biosynthetic gene clusters (BGCs) encode enzyme repertoires that synthesize natural products (NPs), yet linking genomic loci to their chemical products remains a major bottleneck in natural product discovery. We present BGC2NP-CLIP, a lightweight contrastive model for learning a shared embedding space between BGCs and their reported products using experimentally validated associations from MIBiG. BGCs are encoded using simple sequence-derived one-hot features aggregated across their constituent protein sequences, while NPs are described by structure-based molecular fingerprints. A CLIP-style objective aligns matched BGC–NP pairs and separates the mismatched ones, enabling independent encoding of each modality and efficient bidirectional retrieval without relying on foundation-model embeddings or task-specific pre-training.

We evaluate BGC2NP-CLIP using bidirectional cross-modal retrieval with multi-positive ground truth, reflecting that a single cluster may be associated with multiple reported products. The model achieves promising retrieval performance across retrieval directions, consistently ranking matched products among top candidate hits. In addition, we assess whether the learned embeddings form reusable representations for downstream prediction. BGC embeddings from the trained model support biosynthetic class prediction, while NP embeddings retain predictive signal for chemical properties, including molecular weight and compound origin type. Overall, BGC2NP-CLIP provides a fast and reproducible model framework for scalable BGC–NP retrieval and downstream representation learning, showing that simple features combined with contrastive supervision can capture meaningful biosynthetic–chemical structure.

B-G.C.09: DEEPScreen++: A Modular Image-Based Deep Learning Framework for Drug–Target Interaction Prediction
Track: General computational biology
  • Atabey Ünlü, Hacettepe University, Turkey
  • Mehmet Furkan Çalışkan, Hacettepe University, Turkey
  • Furkan Necati İnan, Hacettepe University, Turkey
  • Kemal Örer, Hacettepe University, Turkey
  • Kerem Örer, Hacettepe University, Turkey
  • Tunca DoÄŸan, Hacettepe University, Turkey


Presentation Overview: Show

Computational drug-target interaction (DTI) prediction can reduce the cost and time of drug discovery, yet existing approaches require complex preprocessing, specialized molecular representations, or structure-based workflows, limiting rapid and scalable deployment. Convolutional neural networks offer an alternative by learning discriminative patterns directly from 2D image representations, bypassing extensive feature engineering. Here, we present DEEPScreen++, a modular, open-source framework that formulates DTI prediction as an image-based learning task using RDKit-generated 2D molecular images. DEEPScreen++ provides an end-to-end pipeline for automated data curation from ChEMBL, MoleculeNet, and Therapeutic Data Commons (TDC); configurable 36-fold rotational data augmentation; and efficient model training and inference. The framework supports multiple modern backbones, including convolutional neural networks (CNNs), SwinV2 vision transformers (ViTs), and YOLOv11. With optimized training procedures, target-specific models converge rapidly and can be trained within hours on widely available consumer-grade GPUs. Across public benchmarks from TDC and MoleculeNet, DEEPScreen++ achieves competitive performance relative to existing methods. Integrated SHAP- and saliency-based interpretability modules highlight atom-level features driving model predictions. We demonstrate the practical utility of DEEPScreen++ in two independent studies: drug repurposing for monkeypox virus and de novo screening for AKT1 kinase, both of which identify experimentally confirmed active molecules. Designed for reproducible benchmarking and straightforward extension to new targets, DEEPScreen++ requires minimal preprocessing and accessible computational resources, positioning it as a practical alternative to traditional descriptor-based pipelines. The tool is openly available at https://github.com/HUBioDataLab/DEEPScreen2.

B-G.C.10: Phylogenomic surveillance of outbreaks and antimicrobial resistance across German hospitals
Track: General computational biology
  • Victoria Cepeda Espinoza, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Germany
  • Stephan Ossowski, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Germany


Presentation Overview: Show

Background: Tracking antimicrobial resistance (AMR) in hospital networks requires centralized infrastructure integrating genomic sequencing, standardized bioinformatics, and predictive analytics. We established GenSurv+, a multi-institutional platform with a centralized data hub that collects and analyzes bacterial genome sequences from multiple technologies. These data are linked to phenotypic resistance profiles, enabling comprehensive AMR surveillance across Germany.
Methods: A total of 456 Gram-negative clinical isolates from seven hospitals, sequenced using Illumina, Oxford Nanopore, or PacBio technologies, were analyzed. Analyses included quality control, genome assembly, annotation, core-SNP phylogenetics, AMR profiling with mobile genetic element annotation, and plasmid mobility typing. For isolates with antimicrobial susceptibility testing (AST), a statistical framework combining univariate and multivariate analyses, temporal and stratified methods, network inference, and machine learning was applied to identify genetic drivers of resistance.
Results: Four inter-hospital clonal transmission clusters were identified, comprising 39 Klebsiella pneumoniae isolates across all hospitals. Dominant sequence types were ST15/ST101 (K. pneumoniae) and ST131 (E. coli). AMR genes were widespread, with beta-lactamases detected in 88% of isolates and carbapenemases in 12%. The latter were mainly OXA-48 and NDM-1 in K. pneumoniae and VIM in P. aeruginosa, often located on mobilizable plasmids associated with IS26 transposons. Multidrug resistance was observed in 78% of AST-tested isolates. Machine learning showed high predictive performance for tobramycin, ceftolozane-tazobactam, and ciprofloxacin (AUC 0.972, 0.957, 0.92).
Conclusion: GenSurv+ demonstrates that centralized integration of sequencing, plasmid analysis, statistics, and machine learning enables robust transmission detection and resistance prediction, supporting effective AMR surveillance and intervention strategies.

B-G.C.11: From Sequence to Significance: The LCRPlatform Webserver for Similarity-Driven Functional Annotation Transfer in Low-Complexity Protein Regions
Track: General computational biology
  • Sylwia Morawiec, Silesian Unviersity of Technology, Poland
  • Patryk Jarnot, Silesian University of Technology, Poland
  • Joanna Ziemska-LegiÄ™cka, Institute of Biochemistry and Biophysics, Polish Academy of Sciences, Poland
  • Urszula Ziemska, University of Warsaw, Warsaw, Poland, Poland
  • Marcin Grynberg, Institute of Biochemistry and Biophysics PAS, Poland
  • Aleksandra Gruca, Silesian University of Technology, Poland


Presentation Overview: Show

Low-complexity regions (LCRs) are vital to protein function and evolution, yet they remain notoriously difficult to characterize functionally. We here present the latest update to LCRPlatform, a comprehensive metaserver designed for the precise identification, functional annotation, and evolutionary analysis of LCRs.

The server supports submission of protein IDs and FASTA-formatted sequences, including custom sequences absent from public databases. The platform integrates five state-of-the art LCR detection methods: SEG, CATS, fLPS, SIMPLE, and GBSC through a single, customizable interface. Results are processed across five specialized modules: Annotations, which synthesizes metadata from 12 external repositories via LCRAnnotationsDB; Domains, which maps LCRs against domain databases such as Pfam and SMART; GBSC, for clustering and functional annotation transfer; and Tree View, for phylogenetic visualization.

A key new feature of the latest version of LCRPlatform is a similarity-based annotation transfer, enabling functional information from LCRAnnotationsDB to be projected onto user-provided sequences. Specifically, the platform detects LCRs and utilizes GBSC to identify clusters sharing compositional patterns at customizable coverage thresholds. The GBSC score is then applied to assess the compositional resemblance between a query LCR and the cluster's STR pattern, supporting similarity-based hypothesis generation for previously unannotated regions. This granular approach, for the first time, shifts the analytical focus from whole proteins to the specific functional context of individual LCRs.

Supported by a new infrastructure that utilizes local processing for increased throughput and reliability, LCRPlatform remains a unique and essential resource for functional characterization of LCRs. The server is avialable at https://lcrplatform.lcr-lab.org/

B-G.C.12: Segmenting Organoids: Evaluating Classical, Specialised, and Generalist AI Approaches for Brightfield Microscopy
Track: General computational biology
  • Kerry Heffernan, Univeristy of New South Wales, Australia
  • Shafagh Waters, Univeristy of New South Wales, Australia
  • Fatemeh Vafaee, Univeristy of New South Wales, Australia


Presentation Overview: Show

Organoids are three-dimensional stem cell–derived cultures that replicate key spatial architecture and functional behaviours of their tissue of origin, making them valuable disease models. Patient-derived organoids preserve the individual's unique characteristics and phenotypes, enabling personalised assessment of treatment efficacy. However, robustly extracting quantitative data from brightfield microscopy remains a significant challenge, as no analysis platform has been widely adopted to overcome key hurdles, including accurate segmentation of individual organoids in dense, confluent cultures and the quantitation of complex phenotypes such as lumen cavity dynamics. This study evaluates five diverse image segmentation approaches to establish a robust method for organoid analysis, using in-house and public, intestinal organoid data. These include Gaussian-blur based adaptive thresholding (Classical), layered adaptive thresholding from a published platform (OrganoSeg2), the generalist transformer-based Segment Anything Model (SAM), and published deep-learning based tools (OrganoID and POST). Algorithm performance was compared on a class (semantic) level and at the individual organoid (instance) segmentation level. Semantically, the performance of the five algorithms were similar (F1≈0.8) except SAM which showed oversegmentation (F1=0.65). The instance level performance decreased globally by ~0.2 F1 compared to semantic. POST (sens=0.84; prec=0.45), OrganoSeg2 (sens=0.76; prec=0.51) and SAM (sens=0.92; prec=0.34) had high sensitivity with lower precision, OrganoID (sens=0.52; prec=0.83) showed the opposite trend whereas Classical (sens=0.58; prec=0.56) was more balanced. These results highlight fundamental trade-offs between sensitivity and precision, demonstrating that current methods show reduced performance for instance-level segmentation in dense organoid cultures motivating the development of more robust approaches.

B-G.C.13: Interpretable Single-Cell Transcriptomics Analysis using Generalized Additive Variational Autoencoders
Track: General computational biology
  • Azad Sadr, University of Trieste, Italy
  • Giulio Caravagna, University of Trieste, Italy


Presentation Overview: Show

Deep generative models, particularly Variational Autoencoders (VAEs), have emerged as powerful tools for modeling the high-dimensional complexities of single-cell RNA-seq data. However, canonical deep VAEs operate as opaque black boxes, severely obscuring the precise biological mechanisms and gene regulatory networks driving cellular heterogeneity. To bridge the critical gap between high predictive power and biological interpretability, we introduce a novel, biologically grounded VAE architecture that replaces the standard dense neural network decoder with a Generalized Additive Model (GAM).

Guided by a domain-specific binary prior matrix mapping individual genes to known biological pathways or gene programs, our model enforces a structurally sparse decoder. In this framework, each latent space dimension explicitly represents a distinct, interpretable gene program. We model the expected expression of each gene as the sum of independent, non-linear contributions from its specifically associated programs. To achieve computational efficiency suitable for large-scale atlases, we parameterize these univariate GAM functions using uniform B-spline basis expansions. A structural sparsity mask mathematically guarantees that unassociated latent programs exert strictly zero influence on a gene, thereby preserving exact additive interpretability.

By enforcing structural independence through this additive decoder and a scaled Kullback-Leibler (KL) divergence loss, our framework learns a highly disentangled latent representation. This architecture empowers researchers to uncover complex, non-linear gene regulatory dynamics without sacrificing the transparency of linear models, providing a highly scalable, interpretable solution for advanced mechanistic single-cell analysis.

B-G.C.14: UshEffect-3D: Uncovering the Pathogenicity Landscape within USH2A VUS through Structure-based Machine Learning
Track: General computational biology
  • Diya Prabhuram, The University of Queensland, Australia
  • Stephanie Portelli, The University of Queensland, Australia
  • David Ascher, The University of Queensland, Australia


Presentation Overview: Show

Variants of uncertain significance (VUS) represent a major interpretive bottleneck in the clinical management of inherited retinal diseases (IRDs). USH2A, encoding the large extracellular matrix protein Usherin, is among the most frequently mutated genes among IRDs, yet over 70% of its ClinVar submissions lack a definitive clinical classification. We present UshEffect-3D, a gene-specific machine learning framework that integrates protein structural descriptors, evolutionary conservation metrics, and local biochemical environment features derived from AlphaFold2 structures to predict the pathogenicity of USH2A missense variants. Of eleven classifiers trained on 545 curated variants - Random Forest achieved the highest performance, with an MCC of 0.87, precision of 0.97, sensitivity of 0.92, and specificity of 0.95 on a blind test set, substantially outperforming five general-purpose variant effect predictors including PolyPhen-2, AlphaMissense, and ESM-1b. SHAP analysis identified evolutionary constraint features as the primary driver of model predictions, while structural stability and local residue environment features provided complementary contributions. Applied to 2,639 USH2A VUS in ClinVar, the model prioritised 886 (33.6%) as likely pathogenic, with predicted pathogenic variants enriched within structured domains, particularly the Laminin N-terminal and Laminin G-like regions. These findings demonstrate that gene-specific, structure-informed modeling can meaningfully improve variant interpretation in a clinically high-stakes setting and provide a tractable prioritisation resource for USH2A-associated retinal disease. UshEffect-3D is freely accessible via an interactive web server.

B-G.C.15: An extension of over-representation methods based on metabolic distance: application to the interpretation of metabolic perturbation in relation to key events
Track: General computational biology
  • Maxime Lecomte, INRAE, France
  • Fabien Jourdan, INRAE / MetaboHUB-MetaToul, France
  • Louison Fresnais, L'Oréal, France
  • Kahina Abed, L'Oréal, France
  • Mickael Le Balch, L'Oréal, France
  • Romain Grall, L'Oréal, France
  • Gladys Ouedraogo, L'Oréal, France
  • Nathalie Poupin, INRAe, France


Presentation Overview: Show

Over-representation methods are widely used to interpret omics analyses and identify enriched gene sets or pathways. These approaches assume that all relevant entities are included in predefined gene or pathway sets. To tackle this limitation, we investigated scenarios in which elements of interest do not belong to any existing set. We leveraged the topology of the human metabolic network to compute metrics based on the metabolic distance between entities and a predefined set.

To illustrate our purpose, we considered sets of metabolic reactions identified as modulated under chemical exposure using a modelling pipeline. To better capture the mechanism of action of a chemical, modulated reactions were linked to key events (KE), defined as measurable biological changes leading to adverse outcomes. Relevant metabolic genes were identified combining knowledge graph resources and expert curation. These genes were mapped to metabolic reactions using metabolic network annotations, yielding KE-associated reaction sets. To compute proposed metrics, relying on statistical and graph-neighborhood exploration methods, a refined reaction distance matrix of the human metabolic network was constructed using the Met4j library. This metrics quantifies both proximity and specificity of modulated reactions relative to cluster of KE-associated reactions.

This approach enables the identification of reactions that are highly specific to particular KE clusters. More broadly, it extends traditional enrichment analyses by leveraging network topology, moving beyond strict list-based methods. As a result, it provides a more flexible and functional interpretation of omics data, adaptable to different biological questions.

B-G.C.16: The IDERHA Platform: Federated data governanace and analysis for cross-institutional health research
Track: General computational biology
  • Hanna Ćwiek-Kupczyńska, Luxembourg Centre for Systems Biomedicine, University of Luxembourg, Luxembourg
  • Mostafa Kamal Mallick, Fraunhofer ISST, Dortmund, Germany, Germany
  • Anja Burmann, Fraunhofer ISST, Dortmund, Germany, Germany
  • Venkata Pardhasaradhi Satagopam, Luxembourg Centre For Systems Biomedicine, University of Luxembourg, Luxembourg


Presentation Overview: Show

Secondary use of health data is hindered by a fundamental tension: research requires diverse multi-source datasets while regulations demand that sensitive data remain under institutional control. Federated learning platforms address the computational challenge but data governance (how to control what code can run, manage access across studies, or audit compliance) frequently remains a concern.

We present the IDERHA Platform — a federated infrastructure that integrates data governance and federated learning in a single, standards-based framework, ensuring security, sovereignty and control over data to their providers, while offering a unified interface to data for data users. It supports the research data lifecycle: dataset discovery, access request and permit issuance, and federated analysis, aligning with FAIR principles, GDPR and the spirit of the EHDS regulation.

The platform's architecture integrates established technologies and standards to enforce data sovereignty, so that data providers retain control at every stage: access applications, algorithm approval, results release. Each institution advertises their data via HealthDCAT-AP-based federated catalogue, and implements local policy-based access control via Eclipse Dataspace Components, with ODRL policies governing data release. GA4GH Passports and Visas provide access authorization with fine-grained data minimization. Approved analysis code executes in isolated Trusted Research Environments; only derived outputs are released. Federated analysis is orchestrated via NVFlare, extended with custom governance plugins.

The platform is validated through lung cancer research use cases, integrating multi-modal clinical and imaging data (OMOP CDM, DICOM) from European hospitals.

The IDERHA project is supported by the IHI JU under grant agreement No 101112135.

B-G.C.17: Receptor signaling architecture links pharmacology to subjective psychedelic experience across chemical classes
Track: General computational biology
  • Tereza Kubatova, University of Copenhagen, Denmark
  • Alexander Hauser, University of Copenhagen, Denmark


Presentation Overview: Show

Psychedelics produce a wide range of phenomenological and physiological effects, some beneficial for treating mental health disorders and others associated with substantial adverse risk. Here, we present a multimodal framework linking molecular structure, receptor-level functional pharmacology, and trip-report semantics. Using experimental potencies for 41 psychedelics across 25 GPCRs, we used ligand-based, proteochemometric, and co-folding models to predict receptor activity profiles for additional compounds lacking experimental data. To capture subjective experience, we analysed ~9,000 single-substance trip reports using topic modelling to derive phenomenological embeddings. Across compounds with full experimental data, pharmacological similarity correlated with experiential similarity,  but this relationship disappeared when using predicted rather than experimental potencies, suggesting that accumulated prediction error obscures the true structure–pharmacology–phenomenology relationship.

B-G.C.18: GExMix: Enhancing Molecular Representation with Stratified Data Augmentation for Drug Response Prediction
Track: General computational biology
  • Daksh Pratap Singh Pamar, University of Melbourne, Australia
  • Diyuan Lu, Helmholtz Zentrum München, Germany
  • Ginte Kutkaite, Ludwig-Maximilians-Universität München, Helmholtz Zentrum München, Germany
  • Alexander Ohnmacht, Ludwig-Maximilians-Universität München, Helmholtz Zentrum München, Germany
  • Francesco Paolo Casale, Helmholtz Zentrum München, Germany
  • Michael Menden, University of Melbourne, Helmholtz Zentrum München, Australia


Presentation Overview: Show

Clinical drug response prediction (DRP) is fundamentally constrained by small patient cohorts, characterised by high biological heterogeneity and pronounced response imbalance. While substantial progress has been made through model-centric innovations, data-centric strategies for augmenting transcriptomic and molecular representations remain underexplored.

We introduce GExMix, a model-agnostic data augmentation framework that generates high-utility patient representations via controlled, label-aware interpolation in molecular feature space. GExMix is designed to preserve biological structure while enriching underrepresented response patterns. We systematically evaluate GExMix across three clinical cohorts, five anticancer drugs, and six predictive models, including state-of-the-art DRP architectures.

GExMix consistently outperforms existing augmentation approaches, achieving average relative improvements of 5.7% in ROC-AUC and 6.3% in PR-AUC. Notably, it reduces the performance gap between simple linear models and complex deep learning approaches, highlighting its broad applicability. Mechanistically, GExMix preserves the underlying biological manifold and enhances signal detection, recovering clinically relevant oncogenic and resistance-associated features such as HMGA2 and CD40, which are often obscured in sparse datasets.

Furthermore, GExMix enables seamless integration of auxiliary preclinical data, including cell-line models, yielding an additional 4-6% performance gain by mitigating domain shift between in vitro and in vivo systems.

Together, these results position principled, stratified data augmentation as a critical and previously underutilised dimension in precision oncology. GExMix provides a robust and generalisable framework to improve clinical DRP, particularly in data-limited settings.

B-G.C.19: Evaluating the reliability of drug response prediction across transcriptomic domains
Track: General computational biology
  • Juho Mikkonen, University of Eastern Finland, Finland
  • Teemu Rintala, University of Eastern Finland, Finland
  • Vittorio Fortino, University of Eastern Finland, Finland


Presentation Overview: Show

Deep learning (DL) has become central to modern bioinformatics and health sciences, enabling applications ranging from gene expression -based diagnostics and patient stratification to cell type classification, and multi-omics data integration. Although these models have the potential to substantially improve clinical decision-making and patient outcomes, they often fail to generalize when applied to new cohorts. Without rigorous external validation using independent datasets, DL-based models risk overestimating performance by capturing dataset-specific artifacts rather than underlying biological signals. Biological and technical heterogeneity further exacerbate this issue by introducing substantial distribution shifts between datasets. Such shifts arise from differences in sequencing platforms, library preparation protocols, and patient demographics, leading to unreliable predictions outside the model's original training domain. This challenge is particularly pronounced in transfer learning and domain adaptation settings involving gene expression data, where the goal is often to transfer drug sensitivity predictions from in vitro cancer cell line models to patient tumors. We recently developed a tool, MODAE, which addresses this task by learning transferable representations using data from CCLE, GDSC, CTRP, and TCGA. However, validating the reliability of such models on truly external datasets remains a major challenge.
Here, we propose a diagnostic framework for assessing domain adaptation in deep learning models. This diagnostic was evaluated using MODAE and is designed to support the development of robust validation protocols for deep learning -based de-confounding autoencoders applied to external datasets. Ultimately, this approach provides a principled means to assess whether predictions generated on external cohorts can be considered trustworthy.

B-G.C.20: A Reproducible R Workflow for Flow Cytometry Data Analysis
Track: General computational biology
  • Anastasiya Boersch, University of Basel, Basel, Switzerland; Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland
  • Robert Ivanek, University of Basel, Basel, Switzerland; Swiss Institute of Bioinformatics, Basel, Switzerland, Switzerland


Presentation Overview: Show

Over the past decade, flow cytometry technologies have advanced significantly, now enabling the simultaneous measurement of up to 50 parameters per cell and offering a great analytical potential. However, it also presents substantial challenges, especially when working with large numbers of samples, making the traditional analysis tool FlowJo less suited for handling such complexity. In contrast, Bioconductor provides multiple R packages and workflows that support comprehensive flow cytometry data analysis, including automated data-driven gating and rich visualization capabilities. Nonetheless, the heterogeneity of data structures and often limited documentation across packages can hinder seamless integration and usability.
To address these challenges, we developed an R workflow that streamlines key steps of flow cytometry data analysis and ensures the reproducibility of obtained results: compensation (optional), transformation, quality control, batch correction (optional), gating, and the extraction of gating statistics. This workflow is designed to run locally, making advanced cytometric analysis more accessible and customizable for researchers.

B-G.C.21: LitSABER-chem: LLM-driven Extraction and Prioritization of Enzymatic Reactions from Scientific Literature for Rhea
Track: General computational biology
  • Edouard de Castro, SIB Swiss Institute of Bioinformatics, Switzerland
  • Elisabeth Coudert, Swiss Institute of Bioinformatics, Switzerland
  • Nicole Redaschi, SIB Swiss Institute of Bioinformatics, Switzerland
  • Alan Bridge, SIB Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

Rhea (www.rhea-db.org) is an expert-curated knowledgebase of biochemical reactions built on the chemical ontology ChEBI. As the primary reaction vocabulary of UniProtKB, Rhea provides functional annotations for millions of enzymes across all domains of life, linking chemistry to protein function. Together Rhea and UniProt provide an essential foundation for enzyme function prediction, enzyme design, pathway reconstruction methods, and metabolic network annotation.
Despite its broad utility, expert curation from the rapidly expanding literature remains a major bottleneck, leaving substantial annotation gaps. Large language models (LLMs) now offer an opportunity to address this challenge at scale.
We present LitSABER-chem, an LLM-assisted pipeline for extracting and prioritizing enzymatic reactions from the scientific literature to support Rhea curation. The system performs document triage, applies zero-shot LLM-based named entity recognition and reaction extraction, normalizes entities to ChEBI via retrieval-augmented generation, and grounds extracted reactions against Rhea via SPARQL queries. A ranking module prioritizes reactant-product pairs for expert review.
Evaluated on over 15,000 enzymology abstracts linked to UniProtKB/TrEMBL, the pipeline found thousands of novel candidate reactions absent from Rhea and proposed protein associations for numerous orphan reactions. Within a human-in-the-loop framework, LitSABER-chem is now used to produce curator-validated Rhea reaction entries and associated UniProt annotations, providing a practical demonstration that LLM-augmented literature mining can target and accelerate expert-driven biochemical curation.

B-G.C.22: Content-Based Retrieval for High-Dimensional Biological Data with Multi-Resolution Hierarchical Similarity
Track: General computational biology
  • Arina Surko, Østfold University of Applied Sciences, Norway
  • Hasan Oğul, Østfold University of Applied Sciences, Norway


Presentation Overview: Show

Advances in sequencing technologies provide valuable insights into biology and medicine, such as information about gene expression variation across conditions, between cell types, or over time. The increasing application of these technologies leads to the expansion of the associated databases containing sparse high-dimensional data, which creates a need for effective information retrieval strategies. Content-based search is particularly challenging in such repositories, especially when datasets lack a shared feature alignment.
In this work, we present a workflow for content-based retrieval of high-dimensional biological data. Our framework involves representing each gene expression profile with a hierarchical gene clustering. The similarity between pairs of profiles is computed with the multi-resolution adjusted Rand index - a novel similarity metric incorporating multiple levels of hierarchy into the similarity calculation. We demonstrate our database search workflow for single-cell RNA sequencing data and time-series microarrays. Our results indicate that the framework is flexible, interpretable, and applicable for content-based search across diverse types of high-dimensional data.

B-G.C.23: TRUSTA: Taxon-Aware Error Thresholding for Reliable Species-Level Assignment from Full-Length 16S rRNA Sequencing
Track: General computational biology
  • Mariachiara Vardeu, PhD National Programme in One Health approaches to infectious diseases, University of Pavia, University of Padova, Italy
  • Elena Ghiorzi, PhD National Programme in One Health approaches to infectious diseases, University of Pavia, University of Padova, Italy
  • Alessandro Vezzi, University of Padova, Italy
  • Giulia Bernabè, University of Padova, Italy
  • Massimo Bellato, University of Padova, Italy
  • Stefano Toppo, University of Padova, Italy
  • Cristiano Salata, University of Padova, Italy
  • Ignazio Castagliuolo, University of Padova, Italy
  • Enrico Lavezzo, University of Padova, Italy


Presentation Overview: Show

Full-length 16S rRNA gene sequencing is an increasingly adopted and cost-effective approach for microbial community profiling, providing improved taxonomic resolution compared to short-read amplicon sequencing. However, accurate species-level assignment remains challenging due to high sequence similarity among closely related taxa and the impact of sequencing errors on alignment accuracy, limiting the effectiveness of standard classification approaches based on fixed similarity thresholds.
Here, we present TRUSTA (Taxon-aware Reliable Using Species-level Thresholding for Assignment), a taxon-aware framework for classification of full-length 16S data. Unlike conventional methods that rely on universal cutoffs, TRUSTA adapts assignment criteria to both sequencing error rates and intrinsic taxon-specific resolution limits. It defines taxon-specific reliability thresholds across taxonomic ranks, enabling adaptive confidence estimation and more accurate, interpretable microbial profiles.

To quantify the combined effects of sequencing error and intrinsic sequence similarity, we generated synthetic full-length 16S reads from the ITGDB-16S database using primer-trimmed reference sequences. Reads were simulated with Nanopore-like error rates (1–5%) and aligned using minimap2. Ground truth labels enabled direct evaluation of precision and recall at species and genus levels. Reference clustering identified groups of species with indistinguishable sequences; in these cases, TRUSTA reports multiple candidate species. Performance was also evaluated on a Nanopore-sequenced mock community. Near-perfect accuracy was observed at low error rates (1–2%), with species-specific differences in robustness to increasing error. Overall, TRUSTA shows that incorporating taxon-specific resolution limits and error-aware thresholds improves the reliability of species-level assignment from full-length 16S data.

B-G.C.24: Multimodal Machine Learning for Lung Function Prediction in Interstitial Lung Disease
Track: General computational biology
  • Emma Chanut, Hochschule Hannover - University of Applied Sciences and Arts, Germany
  • Simone Buchholz, Boehringer Ingelheim, Germany
  • Volker Ahlers, Hochschule Hannover - University of Applied Sciences and Arts, Germany


Presentation Overview: Show

This work evaluates the feasibility of predicting one-year follow-up forced vital capacity (FVC) in patients with interstitial lung disease (ILD) using multimodal machine learning models based on three-dimensional high-resolution computed tomography (HRCT) scans and baseline clinical variables. A curated subset of the OSIC cohort was used to compare tabular-only, image-only, and multimodal late-fusion approaches under consistent preprocessing, training, and evaluation protocols, with image-based models employing axial-coronal-sagittal convolutions.
Across all experiments, models relying solely on baseline clinical variables achieved the strongest and most stable predictive performance. Baseline FVC was the dominant predictor, explaining most of the variance in one-year lung function outcomes. In contrast, image-only models trained on full three-dimensional HRCT volumes showed significantly weaker performance, indicating limited prognostic value of baseline imaging for short-term functional decline. Multimodal late-fusion models improved upon image-only approaches but did not outperform tabular baselines, suggesting only modest incremental benefit from HRCT-derived features when strong clinical predictors are available.
These findings highlight a key distinction between diagnostic and prognostic imaging tasks. Although HRCT is essential for ILD diagnosis and disease characterization, its utility for predicting short-term pulmonary function changes from a single baseline scan appears limited under moderate data and computational constraints. Radiomics-based models performed competitively relative to end-to-end deep learning approaches, underscoring the relevance of feature-based methods in data-limited settings. Overall, this work emphasizes aligning prediction targets with the information content of available data and motivates future studies using larger, longitudinal, and more harmonized datasets, as well as more data-efficient modeling strategies.

B-G.C.25: FENNEC: Fine-Tuned Ensemble Neural Networks Accelerate Chemically Modified siRNA Design and Screening
Track: General computational biology
  • Alexander Larsen, Helmholtz / Roche, Germany
  • Daniel Butnaru, Roche, Switzerland
  • Johannes Braun, Roche, Switzerland
  • Rachapun Rottattanadumrong, Roche, Switzerland
  • Philipp Berninger, Roche, Switzerland
  • Dimitar Yonchev, Roche, Switzerland
  • Julien Gagneur, Technical University of Munich, Germany
  • Annalisa Marsico, Helmholtz Munich, Germany


Presentation Overview: Show

Small interfering RNAs (siRNAs) represent a clinically validated therapeutic modality, yet designing potent, chemically modified sequences remains a costly, iterative process constrained by limited public data. Computational prediction of siRNA efficacy is essential for rational design, but few current methods effectively account for chemical modifications, which are critical for therapeutic stability and potency.

We present FENNEC (Fine-Tuned Ensemble of Neural Networks for siRNA Efficiency Characterization), a machine-learning framework designed to predict siRNA activity across chemically diverse design spaces. To overcome data scarcity, we curated the largest dataset of chemically modified siRNAs to date from 49 patents using OCR-based extraction and stringent quality filtering. FENNEC's architecture integrates temporal convolutional networks with thermodynamic descriptors, experimental covariates, and embeddings from RNA foundation models (e.g., RNA-FM, RiNALMo). This allows the model to capture both local chemical rules and broader target-context information.

In benchmarking, FENNEC significantly outperformed classical machine-learning baselines and state-of-the-art deep learning models, demonstrating robust generalization to unseen chemical space. It achieved Spearman correlations of 0.52 under gene-level cross-validation and 0.76 across unique base-siRNA scaffolds. Prospective experimental validation on a novel AHSA1-targeting set further confirmed its utility, yielding a Spearman correlation of 0.68 at 6.33 nM in U251 cells across 94 modified siRNAs. Model interpretation successfully recovered established design principles, including position-specific effects of glycol nucleic acids and phosphorothioate backbones. Overall, FENNEC provides a chemistry-aware deep-learning framework that accelerates the discovery and optimization of modified therapeutic siRNAs.

B-G.C.26: Learning latent symmetries from related bioassays for data-efficient toxicology
Track: General computational biology
  • Arif Dönmez, IUF – Leibniz Research Institute for Environmental Medicine / DNTOX GmbH, Germany
  • Katharina Koch, IUF – Leibniz Research Institute for Environmental Medicine / DNTOX GmbH, Germany
  • Kristina Heck, DNTOX GmbH, Germany
  • Axel Mosig, Ruhr University Bochum / DNTOX GmbH, Germany
  • Ellen Fritsche, SCAHT – Swiss Centre for Applied Human Toxicology / DNTOX GmbH, Switzerland


Presentation Overview: Show

Many toxicological prediction tasks are constrained by scarce labeled data, heterogeneous assay contexts, and fragmented biological evidence. We present a representation-learning framework for this setting that learns latent symmetries during pretraining on biologically related in vitro bioassays and transfers them to downstream toxicology tasks. The central idea is that related assays share structured biological regularities that can be captured before endpoint-specific fine-tuning, yielding more stable and data-efficient molecular representations.
We evaluate this concept in two case studies. First, we consider PFAS, a class of persistent environmental contaminants of major health concern, and predict activity at a single thyroid-axis assay endpoint from a published PFAS screening panel. This panel covers thyroid-axis molecular initiating events relevant to mechanistically informed screening. Second, we examine developmental neurotoxicity (DNT) endpoints to test whether the same pretraining principle generalizes to a broader toxicological context.
Across both case studies, structured pretraining on related bioassays improves performance over models trained only on the target endpoint, supporting the hypothesis that transferable latent symmetry structure can be learned from partially related assay collections. These results position latent-symmetry pretraining as a practical strategy for computational toxicology, with potential value for environmentally relevant chemical screening, endpoint prioritization, and mechanistically informed risk assessment.

B-G.C.27: Modelling gradual and saltational evolutionary dynamics of chromosomal instability across cancers
Track: General computational biology
  • Sara Cocomello, Department of Mathematics, Informatics and Geosciences, University of Trieste, Italy, Italy
  • Giovanni Santacatterina, Department of Mathematics, Informatics and Geosciences, University of Trieste, Italy, Italy
  • Alice Antonello, Department of Mathematics, Informatics and Geosciences, University of Trieste, Italy, Italy
  • Gabriele Oliveto, Department of Experimental Oncology, IEO European Institute of Oncology IRCCS, Milan, 20139, Italy, Italy
  • Martin Schaefer, Department of Experimental Oncology, IEO European Institute of Oncology IRCCS, Milan, 20139, Italy, Italy
  • Giulio Caravagna, Department of Mathematics, Informatics and Geosciences, University of Trieste, Italy, Italy


Presentation Overview: Show

Chromosomal instability (CIN) is a hallmark of solid cancers, encompassing copy number alterations (CNAs) and structural variants that reshape genome content. Recurrent, cancer-type-specific patterns of chromosomal gains and losses suggest that tumour subtypes follow distinct, temporally ordered evolutionary trajectories. Extreme events such as chromothripsis and whole-genome doubling can generate bursts of alterations consistent with punctuated modes of evolution, yet their prevalence and biological impact on disease progression remain only partially understood.
To systematically characterize CIN evolution from whole-genome sequencing data, we introduce TickTack, a hierarchical Bayesian mixture model that uses bulk WGS data from a single timepoint to simultaneously reconstruct the temporal ordering of CNAs and identify clusters of co-occurring events across the genome. The model distinguishes whether alterations accumulate gradually or arise in rapid succession, directly informing our understanding of gradual versus saltational tumour evolution. TickTack was validated on synthetic data and benchmarked against competing approaches.
We applied TickTack to 5,885 tumours from the Pan-Cancer Analysis of Whole Genomes (PCAWG) and the Genomics England 100,000 Genomes Project, recovering cancer-type-specific co-occurring CNA patterns and recurrent evolutionary trajectories across cohorts. We identified a subset of tumours, which we term hopeful monsters, in which a remarkable fraction of the genome was altered at a single time point, consistent with saltational dynamics and linked to elevated genomic instability. By examining how signatures of positive and negative selection distribute across early and late CNA events, we further illuminate the selective forces shaping tumour evolution, with potential implications for therapeutic target identification and prognosis.

B-G.C.28: MariNET: A Linear Mixed Model Framework for Network Analysis of Longitudinal EHR Data
Track: General computational biology
  • Marina Vargas-Fernandez, UGR, Spain
  • Jordi Martorell-Marugán, Foundation for the Promotion of Health and Biomedical Research in the Valencian Region (FISABIO), Spain
  • Pedro Carmona-Saez, UGR, Spain


Presentation Overview: Show

The rapid expansion of healthcare data stored in electronic health records (EHRs) has created new opportunities for biomedical research, enabling the study of complex, heterogeneous, and longitudinal clinical information. However, extracting meaningful relationships among clinical variables remains challenging due to high dimensionality, temporal dependence, and the presence of confounding factors. Traditional approaches, such as Gaussian Graphical Models and Vector Autoregression, are often limited by strong assumptions, including independence and stationarity, which restrict their applicability to real-world EHR data.

To address these challenges, we present MariNET, an R package designed for the construction and analysis of networks derived from longitudinal clinical data. MariNET is based on linear mixed models, allowing the inference of interactions between clinical variables while explicitly accounting for repeated measures, subject-level variability, and confounding effects. The package provides a flexible and scalable framework for modeling dynamic relationships, supporting both continuous and categorical variables, as well as unbalanced and incomplete datasets commonly found in EHRs.

MariNET includes a comprehensive set of functionalities for data preprocessing, model fitting, and network construction, enabling users to estimate associations between variables over time and represent them as interpretable network structures. The package further facilitates customization of model specifications, including the incorporation of fixed and random effects, and offers tools for assessing statistical significance and robustness of inferred interactions. Designed with usability in mind, MariNET integrates seamlessly within the R ecosystem, providing an accessible interface for researchers without requiring advanced expertise in statistical modeling.

B-G.C.29: Shifting the spotlight to unexpected guests: A workflow for the taxonomic and functional profiling of non-target microorganisms in transcriptomics datasets
Track: General computational biology
  • Laura Carmen Terron Camero, Institute of Parasitology and Biomedicine “López-Neyra” (IPBLN), Spanish National Research Council (CSIC), Spain
  • Eduardo Andres-Leon, Institute of Parasitology and Biomedicine “López-Neyra” (IPBLN), Spanish National Research Council (CSIC), Spain
  • Hubert Rehrauer, Functional Genomics Center Zurich, ETH Zurich/University of Zurich, Zurich, Switzerland, Switzerland
  • Jose Luis Ruiz Rodriguez, Functional Genomics Center Zurich, ETH Zurich/University of Zurich, Zurich, Switzerland, Switzerland


Presentation Overview: Show

Transcriptomics has become a cornerstone omics technique in modern biological research. While novel single-cell and spatial approaches offer unprecedented layers of resolution, bulk RNA sequencing remains an essential tool for addressing diverse biological questions. Consequently, an ever-growing wealth of transcriptomics data is populating publicly accessible databases. In reference-based transcriptomics, there is typically a fraction of sequencing reads that fails to align to the target organism and is routinely discarded as noise or contamination. However, this unmapped fraction often harbours valuable microbial information. Furthermore, while many microbiome studies rely on DNA-based approaches for precise taxonomic identification, applying meta-transcriptomics to leverage incidental microbial transcripts within larger samples can provide a critical proxy for investigating molecular functions and biological pathways that are active in situ.
Here, conforming to the reusability aspect of the FAIR principles, we implemented a workflow to mine both novel and previously published transcriptomics datasets that did not originally target microorganisms. As a proof of concept, we aimed to uncover active, unexpected microbiota that remain undiscovered but hold biological relevance. We also benchmarked k-mer- and marker gene-based approaches, demonstrating the need to carefully select tools and databases while fine-tuning default parameters.
Despite technical limitations inherent to original experimental designs, RNA extraction, and library preparation methods optimized for other purposes, we consistently uncovered valuable microbiota-related insights, particularly regarding host-pathogen interactions and vector-borne diseases. This work highlights the immense potential of data reusability to drive meta-transcriptomic discovery across a variety of biological systems and fields where bulk transcriptomics is already routinely applied.