View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
B-G.C.01: Computational Insights into the pH-Dependent Behavior of Ipilimumab-CTLA-4
Track: General computational biology
-
Wanda Destiarani, Doctoral Program in Biology, University of Tsukuba, Japan
-
Kowit Hengphasatporn, Center for Computational Sciences, University of Tsukuba, Japan
- Yasuteru Shigeta, Center for Computational Sciences, University of Tsukuba, Japan
- Ryuhei Harada, Center for Computational Sciences, University of Tsukuba, Japan
Presentation Overview: Show
Therapeutic antibodies face on-target/off-tumor toxicity, causing adverse effects such as cardiotoxicity, skin
rashes, and organ inflammation. To mitigate these challenges, pH-dependent antibodies have been engineered to
preferentially bind in the acidic tumor microenvironment while reducing interactions under physiological pH
conditions. Building on experimental work that generated Ipilimumab variants (Ipi95, Ipi105, Ipi106) through
charged amino acid substitutions in complementarity-determining regions, we employed molecular dynamics
simulations to examine their interactions with CTLA-4 in physiological and acidic conditions. All variants
exhibited enhanced binding affinity at acidic pH, with a reasonable agreement between computational and
experimental binding free energies (R² = 0.7736; Pearson's r = 0.8795, p = 0.0039; Spearman's ρ = 0.8333, p =
0.0102). Statistical analysis revealed notable differences across conditions, most notably for Ipi95, which
demonstrated the highest degree of pH sensitivity. Although no major global structural changes were observed
between conditions, our simulations revealed distinct local energetic rearrangements and residue-level
interaction changes at the binding interface. Decomposition analysis on binding energy further indicated that
the overall antigen-binding mode was maintained, whereas the introduced charged residues modulated local
interaction strengths. These results provide mechanistic insights into how targeted mutations modulate
pH-dependent recognition, offering a framework for the rational design of safer therapeutic antibodies.
B-G.C.02: termal: Interactive Exploration of Multiple Sequence Alignments in the Terminal
Track: General computational biology
-
Thomas Junier, Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Multiple sequence alignments (MSAs) are central to several subfields of
bioinformatics, including phylogenetics, comparative genomics, and sequence
conservation analysis. Such analyses are often carried out on HPC clusters
accessed through SSH sessions, where graphical tools can be unavailable or
impractical, making interactive inspection of alignments difficult.
To address this issue, we developed Termal, an interactive MSA viewer designed
for use in terminals. Termal is entirely keyboard-driven and, like graphical
alignment viewers, provides scrolling, residue colouring, consensus display,
conservation plots, and sequence sorting by properties such as similarity to the
consensus or ungapped length. It also supports zoomed-out views for rapid
inspection of global alignment structure, as well as regular-expression searches
in both headers and sequences.
Termal is implemented in Rust and efficiently handles large alignments. For
example, an alignment of more than 15,000 sequences x 1,500 columns loads in
well under a second on an standard laptop hardware.
Termal is distributed as standalone binaries for Linux, macOS, and Windows. The
program has no external runtime dependencies and integrates naturally into
existing bioinformatics workflows.
By combining interactivity, portability, and performance in a terminal-native
interface, termal provides a practical solution for researchers working on
remote or resource-constrained systems.
B-G.C.03: A Comparative Study of QSPR Methods on a Unique Multitask PAMPA dataset
Track: General computational biology
-
Adam Arany, ESAT-STADIUS, KU Leuven, 3001 Leuven, Belgium, Belgium
- Andras Formanek, ESAT-STADIUS, KU Leuven, 3001 Leuven, Belgium, Belgium
-
Anna Vincze, Dept. Pharm. Chem., Semmelweis University, Hőgyes Endre Street 7-9, H-1092 Budapest,
Hungary, Hungary
-
Richard Bicsak, Depart. Chem. and Environ. Proc. Engineering, BUTE, Műegyetem rkp. 3, H-1111 Budapest,
Hungary, Hungary
-
Gyorgy Tibor Balogh, Dept. Pharm. Chem., Semmelweis University, Hőgyes Endre Street 7-9, H-1092 Budapest,
Hungary, Hungary
- Yves Moreau, ESAT-STADIUS, KU Leuven, 3001 Leuven, Belgium, Belgium
Presentation Overview: Show
We present a unique, multitask dataset comprising 143 drug and drug candidate molecules,
each evaluated on in vitro, parallel artificial-membrane permeability assays (PAMPA) using six different model
membranes.
Using this resource, we systematically assess the effectiveness of various molecular descriptors and
regression models in predicting passive membrane permeability.
The studied models range from simple linear regression to a modern pre-trained transformer architecture.
Particular attention is given to the trade-off between predictive performance and model interpretability,
highlighting the challenges introduced by machine learning approaches.
To our knowledge, this is the most comprehensive study on simultaneous modeling of multiple organ-specific
PAMPA membranes to date,
offering novel insights into membrane-specific permeability profiles.
We found that expert-designed physico-chemical property descriptors are more fitting for a limited sample size
permeabilty study than deep learning based representations.
B-G.C.04: Repurposing blood whole-genome sequencing to study leaky gut
Track: General computational biology
-
Ulas Isildak, Leibniz Institute on Aging - Fritz Lipmann Institute (FLI), Germany
-
Handan Melike Donertas, Leibniz Institute on Aging - Fritz Lipmann Institute (FLI), Germany
Presentation Overview: Show
Intestinal barrier dysfunction (""leaky gut"") has been linked to ageing and chronic disease, but evidence of
its prevalence at the population scale remains scarce. Human whole-genome sequencing (WGS) routinely yields
reads that fail to map to the host genome and are otherwise discarded. We asked whether these unmapped reads
can be repurposed as a signal of gut leakiness. To assess the sensitivity of our approach, we spiked sterile
blood with a mock microbial community and identified a detection limit of roughly 100 cells/mL for
kraken2-based classification, with measurable background in negative controls. In a preliminary analysis of a
pilot cohort from the UK Biobank, following normalisation and filtration informed by these sensitivity
estimates, we detected putative gut-associated microbial presence in approximately 37% of individuals. Having
established the workflow and characterised its detection sensitivity, we aim to investigate whether microbial
content in blood WGS is associated with markers of systemic inflammation, frailty indices, and age-related
phenotypes and diseases. Together, these results suggest that unmapped reads from existing human WGS data may
offer a useful proxy for microbial translocation.
B-G.C.05: Classification of Metabolic Reaction Graphs using Graph Transformers
Track: General computational biology
- Aniello Di Vaio, Universitá di Bologna, Italy
- Mercè Llabrés, University of the Balearic Islands, Spain
-
Jairo Rocha, University of the Balearic Islands, Spain
Presentation Overview: Show
We show the effectiveness of Graphormer, a graph transformer, for the classification of biological reaction
graphs associated with different species. We compare the results with graph kernels, a more traditional
network analyzer.
Reaction graphs derived from metabolic processes compiled by KEGG provide a compact and structured
representation of the biochemical activity of an organism. These graphs often exhibit species-specific
patterns, making them a valuable resource for
comparative analysis and classification tasks in bioinformatics.
Traditional graph-based machine learning approaches, such as graph kernels or graph neural networks, have been
used to analyse these structures, but they may struggle to capture long-range dependencies and complex global
patterns.
Recently, transformer-based architectures have emerged as powerful models capable of learning rich
representations from structured data. Graph Transformers extend the self-attention mechanism to graph domains,
enabling the integration of both local and global structural information.
In this context, graph transformers offer a promising approach for encoding reaction graphs into meaningful
vector representations that can be used for downstream tasks such as classification. In this study, we use
Graphormer as a fixed encoder and combined it with a species classification neural network head. The reaction
graphs are subsampled when they exceed the usual input size for Graphormer.
This work seeks to provide insights into the suitability of graph transformers for bioinformatics applications
involving structured biological data. The results demonstrate that these models can outperform or complement
more traditional graph-based approaches, thereby contributing to the broader use of advanced AI techniques in
biological network analysis.
B-G.C.06: A Tissue-Independent Landscape of Macrophage Dynamics in Regeneration
Track: General computational biology
-
Hanane Moha Ouchane, Institute of Molecular Medicine and Experimental Immunology, University Hospital
Bonn, Germany
-
Ali Enver Bilecen, German Centre for Neurodegenerative Diseases, University Hospital Bonn, Germany
-
Ozgun Gokce, German Centre for Neurodegenerative Diseases, University Hospital Bonn, Germany
-
Zeinab Abdullah, Institute of Molecular Medicine and Experimental Immunology, University Hospital Bonn,
Germany
Presentation Overview: Show
Macrophages are key regulators of tissue regeneration across organs and species. Experimental depletion of
macrophages consistently results in severely impaired regeneration across diverse biological contexts. The
macrophage regenerative response is commonly described as a two-phase response: an early pro-inflammatory M1
phase that promotes debris clearance, progenitor cell recruitment, and proliferation, followed by a later
anti-inflammatory M2 phase that supports differentiation, angiogenesis, and extracellular matrix remodeling.
However, this M1/M2 polarization framework is an oversimplification that fails to capture the dynamic spectrum
of macrophage states observed in vivo. Moreover, current understanding of macrophage activation dynamics is
largely based on independent studies of individual organs and species, making it difficult to identify
generalizable patterns. Identifying such common patterns across tissues and species is key to revealing the
macrophage programs that drive successful regeneration and could help guide strategies to improve outcomes in
tissues with limited regenerative capacity.
Here, we aim to characterize macrophage regenerative programs and pathological states that impair regeneration
in an organ- and species-independent manner. We first construct a mouse regeneration atlas by integrating
publicly available and in-house single-cell and single-nucleus RNA sequencing datasets spanning 11 organs,
capturing the temporal dynamics of regeneration and including both successful and unsuccessful regenerative
outcomes. We then apply disentanglement learning to generate a tissue-independent latent embedding of cell
states, thereby separating organ-specific transcriptional signatures from conserved regenerative programs and
shared pathological states across tissues. Finally, we aim to extend this framework toward a cross-species
characterization of macrophage responses to identify conserved principles of regeneration.
B-G.C.07: Semantic annotation for artifact detection in spatial proteomics using DINOv3 as a vision
foundation model
Track: General computational biology
-
Sviatoslav Kharuk, European Molecular Biology Laboratory (EMBL), Germany
- Matthias Meyer-Bender, European Molecular Biology Laboratory (EMBL), Germany
- Wolfgang Huber, European Molecular Biology Laboratory (EMBL), Germany
Presentation Overview: Show
Our understanding of tissue biology has been transformed in recent years by highly multiplexed imaging that
can capture the spatial organization of proteins in tissues. This has also made it possible to collect tissue
images at a resolution that was previously unattainable. This kind of data can be produced using a number of
technologies, including CyCIF, MIBI, CODEX, and IMC. However, tissue folds, detritus, and antibody aggregates
can occur as a result of improper sample handling during sample preparation. These artifacts can affect
downstream analysis and make it more complicated. I use the vision foundational model DINOv3 to learn
representations from highly multiplexed images. I examine whether this model can effectively learn to identify
and detect artefacts in highly multiplexed images providing automated quality control.
B-G.C.08: BGC2NP-CLIP: Contrastive Cross-Modal Retrieval Between Biosynthetic Gene Clusters and Natural
Products
Track: General computational biology
-
Anastasiia Kolchina, Saarland University, Germany
-
Olga Kalinina, Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Germany
- Dietrich Klakow, Saarland University, Germany
Presentation Overview: Show
Biosynthetic gene clusters (BGCs) encode enzyme repertoires that synthesize natural products (NPs), yet
linking genomic loci to their chemical products remains a major bottleneck in natural product discovery. We
present BGC2NP-CLIP, a lightweight contrastive model for learning a shared embedding space between BGCs and
their reported products using experimentally validated associations from MIBiG. BGCs are encoded using simple
sequence-derived one-hot features aggregated across their constituent protein sequences, while NPs are
described by structure-based molecular fingerprints. A CLIP-style objective aligns matched BGC–NP pairs and
separates the mismatched ones, enabling independent encoding of each modality and efficient bidirectional
retrieval without relying on foundation-model embeddings or task-specific pre-training.
We evaluate BGC2NP-CLIP using bidirectional cross-modal retrieval with multi-positive ground truth, reflecting
that a single cluster may be associated with multiple reported products. The model achieves promising
retrieval performance across retrieval directions, consistently ranking matched products among top candidate
hits. In addition, we assess whether the learned embeddings form reusable representations for downstream
prediction. BGC embeddings from the trained model support biosynthetic class prediction, while NP embeddings
retain predictive signal for chemical properties, including molecular weight and compound origin type.
Overall, BGC2NP-CLIP provides a fast and reproducible model framework for scalable BGC–NP retrieval and
downstream representation learning, showing that simple features combined with contrastive supervision can
capture meaningful biosynthetic–chemical structure.
B-G.C.09: DEEPScreen++: A Modular Image-Based Deep Learning Framework for Drug–Target Interaction
Prediction
Track: General computational biology
- Atabey Ünlü, Hacettepe University, Turkey
- Mehmet Furkan Çalışkan, Hacettepe University, Turkey
- Furkan Necati İnan, Hacettepe University, Turkey
- Kemal Örer, Hacettepe University, Turkey
- Kerem Örer, Hacettepe University, Turkey
-
Tunca DoÄŸan, Hacettepe University, Turkey
Presentation Overview: Show
Computational drug-target interaction (DTI) prediction can reduce the cost and time of drug discovery, yet
existing approaches require complex preprocessing, specialized molecular representations, or structure-based
workflows, limiting rapid and scalable deployment. Convolutional neural networks offer an alternative by
learning discriminative patterns directly from 2D image representations, bypassing extensive feature
engineering. Here, we present DEEPScreen++, a modular, open-source framework that formulates DTI prediction as
an image-based learning task using RDKit-generated 2D molecular images. DEEPScreen++ provides an end-to-end
pipeline for automated data curation from ChEMBL, MoleculeNet, and Therapeutic Data Commons (TDC);
configurable 36-fold rotational data augmentation; and efficient model training and inference. The framework
supports multiple modern backbones, including convolutional neural networks (CNNs), SwinV2 vision transformers
(ViTs), and YOLOv11. With optimized training procedures, target-specific models converge rapidly and can be
trained within hours on widely available consumer-grade GPUs. Across public benchmarks from TDC and
MoleculeNet, DEEPScreen++ achieves competitive performance relative to existing methods. Integrated SHAP- and
saliency-based interpretability modules highlight atom-level features driving model predictions. We
demonstrate the practical utility of DEEPScreen++ in two independent studies: drug repurposing for monkeypox
virus and de novo screening for AKT1 kinase, both of which identify experimentally confirmed active molecules.
Designed for reproducible benchmarking and straightforward extension to new targets, DEEPScreen++ requires
minimal preprocessing and accessible computational resources, positioning it as a practical alternative to
traditional descriptor-based pipelines. The tool is openly available at
https://github.com/HUBioDataLab/DEEPScreen2.
B-G.C.10: Phylogenomic surveillance of outbreaks and antimicrobial resistance across German hospitals
Track: General computational biology
-
Victoria Cepeda Espinoza, Institute of Medical Genetics and Applied Genomics, University of Tübingen,
Germany
-
Stephan Ossowski, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Germany
Presentation Overview: Show
Background: Tracking antimicrobial resistance (AMR) in hospital networks requires centralized infrastructure
integrating genomic sequencing, standardized bioinformatics, and predictive analytics. We established
GenSurv+, a multi-institutional platform with a centralized data hub that collects and analyzes bacterial
genome sequences from multiple technologies. These data are linked to phenotypic resistance profiles, enabling
comprehensive AMR surveillance across Germany.
Methods: A total of 456 Gram-negative clinical isolates from seven hospitals, sequenced using Illumina, Oxford
Nanopore, or PacBio technologies, were analyzed. Analyses included quality control, genome assembly,
annotation, core-SNP phylogenetics, AMR profiling with mobile genetic element annotation, and plasmid mobility
typing. For isolates with antimicrobial susceptibility testing (AST), a statistical framework combining
univariate and multivariate analyses, temporal and stratified methods, network inference, and machine learning
was applied to identify genetic drivers of resistance.
Results: Four inter-hospital clonal transmission clusters were identified, comprising 39 Klebsiella pneumoniae
isolates across all hospitals. Dominant sequence types were ST15/ST101 (K. pneumoniae) and ST131 (E. coli).
AMR genes were widespread, with beta-lactamases detected in 88% of isolates and carbapenemases in 12%. The
latter were mainly OXA-48 and NDM-1 in K. pneumoniae and VIM in P. aeruginosa, often located on mobilizable
plasmids associated with IS26 transposons. Multidrug resistance was observed in 78% of AST-tested isolates.
Machine learning showed high predictive performance for tobramycin, ceftolozane-tazobactam, and ciprofloxacin
(AUC 0.972, 0.957, 0.92).
Conclusion: GenSurv+ demonstrates that centralized integration of sequencing, plasmid analysis, statistics,
and machine learning enables robust transmission detection and resistance prediction, supporting effective AMR
surveillance and intervention strategies.
B-G.C.11: From Sequence to Significance: The LCRPlatform Webserver for Similarity-Driven Functional
Annotation Transfer in Low-Complexity Protein Regions
Track: General computational biology
- Sylwia Morawiec, Silesian Unviersity of Technology, Poland
- Patryk Jarnot, Silesian University of Technology, Poland
-
Joanna Ziemska-Legięcka, Institute of Biochemistry and Biophysics, Polish Academy of Sciences, Poland
- Urszula Ziemska, University of Warsaw, Warsaw, Poland, Poland
- Marcin Grynberg, Institute of Biochemistry and Biophysics PAS, Poland
-
Aleksandra Gruca, Silesian University of Technology, Poland
Presentation Overview: Show
Low-complexity regions (LCRs) are vital to protein function and evolution, yet they remain notoriously
difficult to characterize functionally. We here present the latest update to LCRPlatform, a comprehensive
metaserver designed for the precise identification, functional annotation, and evolutionary analysis of LCRs.
The server supports submission of protein IDs and FASTA-formatted sequences, including custom sequences absent
from public databases. The platform integrates five state-of-the art LCR detection methods: SEG, CATS, fLPS,
SIMPLE, and GBSC through a single, customizable interface. Results are processed across five specialized
modules: Annotations, which synthesizes metadata from 12 external repositories via LCRAnnotationsDB; Domains,
which maps LCRs against domain databases such as Pfam and SMART; GBSC, for clustering and functional
annotation transfer; and Tree View, for phylogenetic visualization.
A key new feature of the latest version of LCRPlatform is a similarity-based annotation transfer, enabling
functional information from LCRAnnotationsDB to be projected onto user-provided sequences. Specifically, the
platform detects LCRs and utilizes GBSC to identify clusters sharing compositional patterns at customizable
coverage thresholds. The GBSC score is then applied to assess the compositional resemblance between a query
LCR and the cluster's STR pattern, supporting similarity-based hypothesis generation for previously
unannotated regions. This granular approach, for the first time, shifts the analytical focus from whole
proteins to the specific functional context of individual LCRs.
Supported by a new infrastructure that utilizes local processing for increased throughput and reliability,
LCRPlatform remains a unique and essential resource for functional characterization of LCRs. The server is
avialable at https://lcrplatform.lcr-lab.org/
B-G.C.12: Segmenting Organoids: Evaluating Classical, Specialised, and Generalist AI Approaches for
Brightfield Microscopy
Track: General computational biology
-
Kerry Heffernan, Univeristy of New South Wales, Australia
- Shafagh Waters, Univeristy of New South Wales, Australia
- Fatemeh Vafaee, Univeristy of New South Wales, Australia
Presentation Overview: Show
Organoids are three-dimensional stem cell–derived cultures that replicate key spatial architecture and
functional behaviours of their tissue of origin, making them valuable disease models. Patient-derived
organoids preserve the individual's unique characteristics and phenotypes, enabling personalised assessment of
treatment efficacy. However, robustly extracting quantitative data from brightfield microscopy remains a
significant challenge, as no analysis platform has been widely adopted to overcome key hurdles, including
accurate segmentation of individual organoids in dense, confluent cultures and the quantitation of complex
phenotypes such as lumen cavity dynamics. This study evaluates five diverse image segmentation approaches to
establish a robust method for organoid analysis, using in-house and public, intestinal organoid data. These
include Gaussian-blur based adaptive thresholding (Classical), layered adaptive thresholding from a published
platform (OrganoSeg2), the generalist transformer-based Segment Anything Model (SAM), and published
deep-learning based tools (OrganoID and POST). Algorithm performance was compared on a class (semantic) level
and at the individual organoid (instance) segmentation level. Semantically, the performance of the five
algorithms were similar (F1≈0.8) except SAM which showed oversegmentation (F1=0.65). The instance level
performance decreased globally by ~0.2 F1 compared to semantic. POST (sens=0.84; prec=0.45), OrganoSeg2
(sens=0.76; prec=0.51) and SAM (sens=0.92; prec=0.34) had high sensitivity with lower precision, OrganoID
(sens=0.52; prec=0.83) showed the opposite trend whereas Classical (sens=0.58; prec=0.56) was more balanced.
These results highlight fundamental trade-offs between sensitivity and precision, demonstrating that current
methods show reduced performance for instance-level segmentation in dense organoid cultures motivating the
development of more robust approaches.
B-G.C.13: Interpretable Single-Cell Transcriptomics Analysis using Generalized Additive Variational
Autoencoders
Track: General computational biology
-
Azad Sadr, University of Trieste, Italy
- Giulio Caravagna, University of Trieste, Italy
Presentation Overview: Show
Deep generative models, particularly Variational Autoencoders (VAEs), have emerged as powerful tools for
modeling the high-dimensional complexities of single-cell RNA-seq data. However, canonical deep VAEs operate
as opaque black boxes, severely obscuring the precise biological mechanisms and gene regulatory networks
driving cellular heterogeneity. To bridge the critical gap between high predictive power and biological
interpretability, we introduce a novel, biologically grounded VAE architecture that replaces the standard
dense neural network decoder with a Generalized Additive Model (GAM).
Guided by a domain-specific binary prior matrix mapping individual genes to known biological pathways or gene
programs, our model enforces a structurally sparse decoder. In this framework, each latent space dimension
explicitly represents a distinct, interpretable gene program. We model the expected expression of each gene as
the sum of independent, non-linear contributions from its specifically associated programs. To achieve
computational efficiency suitable for large-scale atlases, we parameterize these univariate GAM functions
using uniform B-spline basis expansions. A structural sparsity mask mathematically guarantees that
unassociated latent programs exert strictly zero influence on a gene, thereby preserving exact additive
interpretability.
By enforcing structural independence through this additive decoder and a scaled Kullback-Leibler (KL)
divergence loss, our framework learns a highly disentangled latent representation. This architecture empowers
researchers to uncover complex, non-linear gene regulatory dynamics without sacrificing the transparency of
linear models, providing a highly scalable, interpretable solution for advanced mechanistic single-cell
analysis.
B-G.C.14: UshEffect-3D: Uncovering the Pathogenicity Landscape within USH2A VUS through Structure-based
Machine Learning
Track: General computational biology
-
Diya Prabhuram, The University of Queensland, Australia
- Stephanie Portelli, The University of Queensland, Australia
- David Ascher, The University of Queensland, Australia
Presentation Overview: Show
Variants of uncertain significance (VUS) represent a major interpretive bottleneck in the clinical management
of inherited retinal diseases (IRDs). USH2A, encoding the large extracellular matrix protein Usherin, is among
the most frequently mutated genes among IRDs, yet over 70% of its ClinVar submissions lack a definitive
clinical classification. We present UshEffect-3D, a gene-specific machine learning framework that integrates
protein structural descriptors, evolutionary conservation metrics, and local biochemical environment features
derived from AlphaFold2 structures to predict the pathogenicity of USH2A missense variants. Of eleven
classifiers trained on 545 curated variants - Random Forest achieved the highest performance, with an MCC of
0.87, precision of 0.97, sensitivity of 0.92, and specificity of 0.95 on a blind test set, substantially
outperforming five general-purpose variant effect predictors including PolyPhen-2, AlphaMissense, and ESM-1b.
SHAP analysis identified evolutionary constraint features as the primary driver of model predictions, while
structural stability and local residue environment features provided complementary contributions. Applied to
2,639 USH2A VUS in ClinVar, the model prioritised 886 (33.6%) as likely pathogenic, with predicted pathogenic
variants enriched within structured domains, particularly the Laminin N-terminal and Laminin G-like regions.
These findings demonstrate that gene-specific, structure-informed modeling can meaningfully improve variant
interpretation in a clinically high-stakes setting and provide a tractable prioritisation resource for
USH2A-associated retinal disease. UshEffect-3D is freely accessible via an interactive web server.
B-G.C.15: An extension of over-representation methods based on metabolic distance: application to the
interpretation of metabolic perturbation in relation to key events
Track: General computational biology
-
Maxime Lecomte, INRAE, France
- Fabien Jourdan, INRAE / MetaboHUB-MetaToul, France
- Louison Fresnais, L'Oréal, France
- Kahina Abed, L'Oréal, France
- Mickael Le Balch, L'Oréal, France
- Romain Grall, L'Oréal, France
- Gladys Ouedraogo, L'Oréal, France
- Nathalie Poupin, INRAe, France
Presentation Overview: Show
Over-representation methods are widely used to interpret omics analyses and identify enriched gene sets or
pathways. These approaches assume that all relevant entities are included in predefined gene or pathway sets.
To tackle this limitation, we investigated scenarios in which elements of interest do not belong to any
existing set. We leveraged the topology of the human metabolic network to compute metrics based on the
metabolic distance between entities and a predefined set.
To illustrate our purpose, we considered sets of metabolic reactions identified as modulated under chemical
exposure using a modelling pipeline. To better capture the mechanism of action of a chemical, modulated
reactions were linked to key events (KE), defined as measurable biological changes leading to adverse
outcomes. Relevant metabolic genes were identified combining knowledge graph resources and expert curation.
These genes were mapped to metabolic reactions using metabolic network annotations, yielding KE-associated
reaction sets. To compute proposed metrics, relying on statistical and graph-neighborhood exploration methods,
a refined reaction distance matrix of the human metabolic network was constructed using the Met4j library.
This metrics quantifies both proximity and specificity of modulated reactions relative to cluster of
KE-associated reactions.
This approach enables the identification of reactions that are highly specific to particular KE clusters. More
broadly, it extends traditional enrichment analyses by leveraging network topology, moving beyond strict
list-based methods. As a result, it provides a more flexible and functional interpretation of omics data,
adaptable to different biological questions.
B-G.C.16: The IDERHA Platform: Federated data governanace and analysis for cross-institutional health
research
Track: General computational biology
-
Hanna Ćwiek-Kupczyńska, Luxembourg Centre for Systems Biomedicine, University of Luxembourg,
Luxembourg
- Mostafa Kamal Mallick, Fraunhofer ISST, Dortmund, Germany, Germany
- Anja Burmann, Fraunhofer ISST, Dortmund, Germany, Germany
-
Venkata Pardhasaradhi Satagopam, Luxembourg Centre For Systems Biomedicine, University of Luxembourg,
Luxembourg
Presentation Overview: Show
Secondary use of health data is hindered by a fundamental tension: research requires diverse multi-source
datasets while regulations demand that sensitive data remain under institutional control. Federated learning
platforms address the computational challenge but data governance (how to control what code can run, manage
access across studies, or audit compliance) frequently remains a concern.
We present the IDERHA Platform — a federated infrastructure that integrates data governance and federated
learning in a single, standards-based framework, ensuring security, sovereignty and control over data to their
providers, while offering a unified interface to data for data users. It supports the research data lifecycle:
dataset discovery, access request and permit issuance, and federated analysis, aligning with FAIR principles,
GDPR and the spirit of the EHDS regulation.
The platform's architecture integrates established technologies and standards to enforce data sovereignty, so
that data providers retain control at every stage: access applications, algorithm approval, results release.
Each institution advertises their data via HealthDCAT-AP-based federated catalogue, and implements local
policy-based access control via Eclipse Dataspace Components, with ODRL policies governing data release. GA4GH
Passports and Visas provide access authorization with fine-grained data minimization. Approved analysis code
executes in isolated Trusted Research Environments; only derived outputs are released. Federated analysis is
orchestrated via NVFlare, extended with custom governance plugins.
The platform is validated through lung cancer research use cases, integrating multi-modal clinical and imaging
data (OMOP CDM, DICOM) from European hospitals.
The IDERHA project is supported by the IHI JU under grant agreement No 101112135.
B-G.C.17: Receptor signaling architecture links pharmacology to subjective psychedelic experience across
chemical classes
Track: General computational biology
-
Tereza Kubatova, University of Copenhagen, Denmark
- Alexander Hauser, University of Copenhagen, Denmark
Presentation Overview: Show
Psychedelics produce a wide range of phenomenological and physiological effects, some beneficial for treating
mental health disorders and others associated with substantial adverse risk. Here, we present a multimodal
framework linking molecular structure, receptor-level functional pharmacology, and trip-report semantics.
Using experimental potencies for 41 psychedelics across 25 GPCRs, we used ligand-based, proteochemometric, and
co-folding models to predict receptor activity profiles for additional compounds lacking experimental data. To
capture subjective experience, we analysed ~9,000 single-substance trip reports using topic modelling to
derive phenomenological embeddings. Across compounds with full experimental data, pharmacological similarity
correlated with experiential similarity, but this relationship disappeared when using predicted rather than
experimental potencies, suggesting that accumulated prediction error obscures the true
structure–pharmacology–phenomenology relationship.
B-G.C.18: GExMix: Enhancing Molecular Representation with Stratified Data Augmentation for Drug Response
Prediction
Track: General computational biology
-
Daksh Pratap Singh Pamar, University of Melbourne, Australia
- Diyuan Lu, Helmholtz Zentrum München, Germany
-
Ginte Kutkaite, Ludwig-Maximilians-Universität München, Helmholtz Zentrum München, Germany
-
Alexander Ohnmacht, Ludwig-Maximilians-Universität München, Helmholtz Zentrum München, Germany
- Francesco Paolo Casale, Helmholtz Zentrum München, Germany
- Michael Menden, University of Melbourne, Helmholtz Zentrum München, Australia
Presentation Overview: Show
Clinical drug response prediction (DRP) is fundamentally constrained by small patient cohorts, characterised
by high biological heterogeneity and pronounced response imbalance. While substantial progress has been made
through model-centric innovations, data-centric strategies for augmenting transcriptomic and molecular
representations remain underexplored.
We introduce GExMix, a model-agnostic data augmentation framework that generates high-utility patient
representations via controlled, label-aware interpolation in molecular feature space. GExMix is designed to
preserve biological structure while enriching underrepresented response patterns. We systematically evaluate
GExMix across three clinical cohorts, five anticancer drugs, and six predictive models, including
state-of-the-art DRP architectures.
GExMix consistently outperforms existing augmentation approaches, achieving average relative improvements of
5.7% in ROC-AUC and 6.3% in PR-AUC. Notably, it reduces the performance gap between simple linear models and
complex deep learning approaches, highlighting its broad applicability. Mechanistically, GExMix preserves the
underlying biological manifold and enhances signal detection, recovering clinically relevant oncogenic and
resistance-associated features such as HMGA2 and CD40, which are often obscured in sparse datasets.
Furthermore, GExMix enables seamless integration of auxiliary preclinical data, including cell-line models,
yielding an additional 4-6% performance gain by mitigating domain shift between in vitro and in vivo
systems.
Together, these results position principled, stratified data augmentation as a critical and previously
underutilised dimension in precision oncology. GExMix provides a robust and generalisable framework to improve
clinical DRP, particularly in data-limited settings.
B-G.C.19: Evaluating the reliability of drug response prediction across transcriptomic domains
Track: General computational biology
-
Juho Mikkonen, University of Eastern Finland, Finland
- Teemu Rintala, University of Eastern Finland, Finland
- Vittorio Fortino, University of Eastern Finland, Finland
Presentation Overview: Show
Deep learning (DL) has become central to modern bioinformatics and health sciences, enabling applications
ranging from gene expression -based diagnostics and patient stratification to cell type classification, and
multi-omics data integration. Although these models have the potential to substantially improve clinical
decision-making and patient outcomes, they often fail to generalize when applied to new cohorts. Without
rigorous external validation using independent datasets, DL-based models risk overestimating performance by
capturing dataset-specific artifacts rather than underlying biological signals. Biological and technical
heterogeneity further exacerbate this issue by introducing substantial distribution shifts between datasets.
Such shifts arise from differences in sequencing platforms, library preparation protocols, and patient
demographics, leading to unreliable predictions outside the model's original training domain. This challenge
is particularly pronounced in transfer learning and domain adaptation settings involving gene expression data,
where the goal is often to transfer drug sensitivity predictions from in vitro cancer cell line models to
patient tumors. We recently developed a tool, MODAE, which addresses this task by learning transferable
representations using data from CCLE, GDSC, CTRP, and TCGA. However, validating the reliability of such models
on truly external datasets remains a major challenge.
Here, we propose a diagnostic framework for assessing domain adaptation in deep learning models. This
diagnostic was evaluated using MODAE and is designed to support the development of robust validation protocols
for deep learning -based de-confounding autoencoders applied to external datasets. Ultimately, this approach
provides a principled means to assess whether predictions generated on external cohorts can be considered
trustworthy.
B-G.C.20: A Reproducible R Workflow for Flow Cytometry Data Analysis
Track: General computational biology
-
Anastasiya Boersch, University of Basel, Basel, Switzerland; Swiss Institute of Bioinformatics,
Basel, Switzerland, Switzerland
-
Robert Ivanek, University of Basel, Basel, Switzerland; Swiss Institute of Bioinformatics, Basel,
Switzerland, Switzerland
Presentation Overview: Show
Over the past decade, flow cytometry technologies have advanced significantly, now enabling the simultaneous
measurement of up to 50 parameters per cell and offering a great analytical potential. However, it also
presents substantial challenges, especially when working with large numbers of samples, making the traditional
analysis tool FlowJo less suited for handling such complexity. In contrast, Bioconductor provides multiple R
packages and workflows that support comprehensive flow cytometry data analysis, including automated
data-driven gating and rich visualization capabilities. Nonetheless, the heterogeneity of data structures and
often limited documentation across packages can hinder seamless integration and usability.
To address these challenges, we developed an R workflow that streamlines key steps of flow cytometry data
analysis and ensures the reproducibility of obtained results: compensation (optional), transformation, quality
control, batch correction (optional), gating, and the extraction of gating statistics. This workflow is
designed to run locally, making advanced cytometric analysis more accessible and customizable for researchers.
B-G.C.21: LitSABER-chem: LLM-driven Extraction and Prioritization of Enzymatic Reactions from Scientific
Literature for Rhea
Track: General computational biology
-
Edouard de Castro, SIB Swiss Institute of Bioinformatics, Switzerland
- Elisabeth Coudert, Swiss Institute of Bioinformatics, Switzerland
- Nicole Redaschi, SIB Swiss Institute of Bioinformatics, Switzerland
- Alan Bridge, SIB Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Rhea (www.rhea-db.org) is an expert-curated knowledgebase of biochemical reactions built on the chemical
ontology ChEBI. As the primary reaction vocabulary of UniProtKB, Rhea provides functional annotations for
millions of enzymes across all domains of life, linking chemistry to protein function. Together Rhea and
UniProt provide an essential foundation for enzyme function prediction, enzyme design, pathway reconstruction
methods, and metabolic network annotation.
Despite its broad utility, expert curation from the rapidly expanding literature remains a major bottleneck,
leaving substantial annotation gaps. Large language models (LLMs) now offer an opportunity to address this
challenge at scale.
We present LitSABER-chem, an LLM-assisted pipeline for extracting and prioritizing enzymatic reactions from
the scientific literature to support Rhea curation. The system performs document triage, applies zero-shot
LLM-based named entity recognition and reaction extraction, normalizes entities to ChEBI via
retrieval-augmented generation, and grounds extracted reactions against Rhea via SPARQL queries. A ranking
module prioritizes reactant-product pairs for expert review.
Evaluated on over 15,000 enzymology abstracts linked to UniProtKB/TrEMBL, the pipeline found thousands of
novel candidate reactions absent from Rhea and proposed protein associations for numerous orphan reactions.
Within a human-in-the-loop framework, LitSABER-chem is now used to produce curator-validated Rhea reaction
entries and associated UniProt annotations, providing a practical demonstration that LLM-augmented literature
mining can target and accelerate expert-driven biochemical curation.
B-G.C.22: Content-Based Retrieval for High-Dimensional Biological Data with Multi-Resolution Hierarchical
Similarity
Track: General computational biology
-
Arina Surko, Østfold University of Applied Sciences, Norway
- Hasan Oğul, Østfold University of Applied Sciences, Norway
Presentation Overview: Show
Advances in sequencing technologies provide valuable insights into biology and medicine, such as information
about gene expression variation across conditions, between cell types, or over time. The increasing
application of these technologies leads to the expansion of the associated databases containing sparse
high-dimensional data, which creates a need for effective information retrieval strategies. Content-based
search is particularly challenging in such repositories, especially when datasets lack a shared feature
alignment.
In this work, we present a workflow for content-based retrieval of high-dimensional biological data. Our
framework involves representing each gene expression profile with a hierarchical gene clustering. The
similarity between pairs of profiles is computed with the multi-resolution adjusted Rand index - a novel
similarity metric incorporating multiple levels of hierarchy into the similarity calculation. We demonstrate
our database search workflow for single-cell RNA sequencing data and time-series microarrays. Our results
indicate that the framework is flexible, interpretable, and applicable for content-based search across diverse
types of high-dimensional data.
B-G.C.23: TRUSTA: Taxon-Aware Error Thresholding for Reliable Species-Level Assignment from Full-Length 16S
rRNA Sequencing
Track: General computational biology
-
Mariachiara Vardeu, PhD National Programme in One Health approaches to infectious diseases,
University of Pavia, University of Padova, Italy
-
Elena Ghiorzi, PhD National Programme in One Health approaches to infectious diseases, University of
Pavia, University of Padova, Italy
- Alessandro Vezzi, University of Padova, Italy
- Giulia Bernabè, University of Padova, Italy
- Massimo Bellato, University of Padova, Italy
- Stefano Toppo, University of Padova, Italy
- Cristiano Salata, University of Padova, Italy
- Ignazio Castagliuolo, University of Padova, Italy
- Enrico Lavezzo, University of Padova, Italy
Presentation Overview: Show
Full-length 16S rRNA gene sequencing is an increasingly adopted and cost-effective approach for microbial
community profiling, providing improved taxonomic resolution compared to short-read amplicon sequencing.
However, accurate species-level assignment remains challenging due to high sequence similarity among closely
related taxa and the impact of sequencing errors on alignment accuracy, limiting the effectiveness of standard
classification approaches based on fixed similarity thresholds.
Here, we present TRUSTA (Taxon-aware Reliable Using Species-level Thresholding for Assignment), a taxon-aware
framework for classification of full-length 16S data. Unlike conventional methods that rely on universal
cutoffs, TRUSTA adapts assignment criteria to both sequencing error rates and intrinsic taxon-specific
resolution limits. It defines taxon-specific reliability thresholds across taxonomic ranks, enabling adaptive
confidence estimation and more accurate, interpretable microbial profiles.
To quantify the combined effects of sequencing error and intrinsic sequence similarity, we generated synthetic
full-length 16S reads from the ITGDB-16S database using primer-trimmed reference sequences. Reads were
simulated with Nanopore-like error rates (1–5%) and aligned using minimap2. Ground truth labels enabled
direct evaluation of precision and recall at species and genus levels. Reference clustering identified groups
of species with indistinguishable sequences; in these cases, TRUSTA reports multiple candidate species.
Performance was also evaluated on a Nanopore-sequenced mock community. Near-perfect accuracy was observed at
low error rates (1–2%), with species-specific differences in robustness to increasing error. Overall, TRUSTA
shows that incorporating taxon-specific resolution limits and error-aware thresholds improves the reliability
of species-level assignment from full-length 16S data.
B-G.C.24: Multimodal Machine Learning for Lung Function Prediction in Interstitial Lung Disease
Track: General computational biology
-
Emma Chanut, Hochschule Hannover - University of Applied Sciences and Arts, Germany
- Simone Buchholz, Boehringer Ingelheim, Germany
-
Volker Ahlers, Hochschule Hannover - University of Applied Sciences and Arts, Germany
Presentation Overview: Show
This work evaluates the feasibility of predicting one-year follow-up forced vital capacity (FVC) in patients
with interstitial lung disease (ILD) using multimodal machine learning models based on three-dimensional
high-resolution computed tomography (HRCT) scans and baseline clinical variables. A curated subset of the OSIC
cohort was used to compare tabular-only, image-only, and multimodal late-fusion approaches under consistent
preprocessing, training, and evaluation protocols, with image-based models employing axial-coronal-sagittal
convolutions.
Across all experiments, models relying solely on baseline clinical variables achieved the strongest and most
stable predictive performance. Baseline FVC was the dominant predictor, explaining most of the variance in
one-year lung function outcomes. In contrast, image-only models trained on full three-dimensional HRCT volumes
showed significantly weaker performance, indicating limited prognostic value of baseline imaging for
short-term functional decline. Multimodal late-fusion models improved upon image-only approaches but did not
outperform tabular baselines, suggesting only modest incremental benefit from HRCT-derived features when
strong clinical predictors are available.
These findings highlight a key distinction between diagnostic and prognostic imaging tasks. Although HRCT is
essential for ILD diagnosis and disease characterization, its utility for predicting short-term pulmonary
function changes from a single baseline scan appears limited under moderate data and computational
constraints. Radiomics-based models performed competitively relative to end-to-end deep learning approaches,
underscoring the relevance of feature-based methods in data-limited settings. Overall, this work emphasizes
aligning prediction targets with the information content of available data and motivates future studies using
larger, longitudinal, and more harmonized datasets, as well as more data-efficient modeling strategies.
B-G.C.25: FENNEC: Fine-Tuned Ensemble Neural Networks Accelerate Chemically Modified siRNA Design and
Screening
Track: General computational biology
-
Alexander Larsen, Helmholtz / Roche, Germany
- Daniel Butnaru, Roche, Switzerland
- Johannes Braun, Roche, Switzerland
- Rachapun Rottattanadumrong, Roche, Switzerland
- Philipp Berninger, Roche, Switzerland
- Dimitar Yonchev, Roche, Switzerland
- Julien Gagneur, Technical University of Munich, Germany
- Annalisa Marsico, Helmholtz Munich, Germany
Presentation Overview: Show
Small interfering RNAs (siRNAs) represent a clinically validated therapeutic modality, yet designing potent,
chemically modified sequences remains a costly, iterative process constrained by limited public data.
Computational prediction of siRNA efficacy is essential for rational design, but few current methods
effectively account for chemical modifications, which are critical for therapeutic stability and potency.
We present FENNEC (Fine-Tuned Ensemble of Neural Networks for siRNA Efficiency Characterization), a
machine-learning framework designed to predict siRNA activity across chemically diverse design spaces. To
overcome data scarcity, we curated the largest dataset of chemically modified siRNAs to date from 49 patents
using OCR-based extraction and stringent quality filtering. FENNEC's architecture integrates temporal
convolutional networks with thermodynamic descriptors, experimental covariates, and embeddings from RNA
foundation models (e.g., RNA-FM, RiNALMo). This allows the model to capture both local chemical rules and
broader target-context information.
In benchmarking, FENNEC significantly outperformed classical machine-learning baselines and state-of-the-art
deep learning models, demonstrating robust generalization to unseen chemical space. It achieved Spearman
correlations of 0.52 under gene-level cross-validation and 0.76 across unique base-siRNA scaffolds.
Prospective experimental validation on a novel AHSA1-targeting set further confirmed its utility, yielding a
Spearman correlation of 0.68 at 6.33 nM in U251 cells across 94 modified siRNAs. Model interpretation
successfully recovered established design principles, including position-specific effects of glycol nucleic
acids and phosphorothioate backbones. Overall, FENNEC provides a chemistry-aware deep-learning framework that
accelerates the discovery and optimization of modified therapeutic siRNAs.
B-G.C.26: Learning latent symmetries from related bioassays for data-efficient toxicology
Track: General computational biology
-
Arif Dönmez, IUF – Leibniz Research Institute for Environmental Medicine / DNTOX GmbH,
Germany
-
Katharina Koch, IUF – Leibniz Research Institute for Environmental Medicine / DNTOX GmbH, Germany
- Kristina Heck, DNTOX GmbH, Germany
- Axel Mosig, Ruhr University Bochum / DNTOX GmbH, Germany
-
Ellen Fritsche, SCAHT – Swiss Centre for Applied Human Toxicology / DNTOX GmbH, Switzerland
Presentation Overview: Show
Many toxicological prediction tasks are constrained by scarce labeled data, heterogeneous assay contexts, and
fragmented biological evidence. We present a representation-learning framework for this setting that learns
latent symmetries during pretraining on biologically related in vitro bioassays and transfers them to
downstream toxicology tasks. The central idea is that related assays share structured biological regularities
that can be captured before endpoint-specific fine-tuning, yielding more stable and data-efficient molecular
representations.
We evaluate this concept in two case studies. First, we consider PFAS, a class of persistent environmental
contaminants of major health concern, and predict activity at a single thyroid-axis assay endpoint from a
published PFAS screening panel. This panel covers thyroid-axis molecular initiating events relevant to
mechanistically informed screening. Second, we examine developmental neurotoxicity (DNT) endpoints to test
whether the same pretraining principle generalizes to a broader toxicological context.
Across both case studies, structured pretraining on related bioassays improves performance over models trained
only on the target endpoint, supporting the hypothesis that transferable latent symmetry structure can be
learned from partially related assay collections. These results position latent-symmetry pretraining as a
practical strategy for computational toxicology, with potential value for environmentally relevant chemical
screening, endpoint prioritization, and mechanistically informed risk assessment.
B-G.C.27: Modelling gradual and saltational evolutionary dynamics of chromosomal instability across
cancers
Track: General computational biology
-
Sara Cocomello, Department of Mathematics, Informatics and Geosciences, University of Trieste, Italy,
Italy
-
Giovanni Santacatterina, Department of Mathematics, Informatics and Geosciences, University of Trieste,
Italy, Italy
-
Alice Antonello, Department of Mathematics, Informatics and Geosciences, University of Trieste, Italy,
Italy
-
Gabriele Oliveto, Department of Experimental Oncology, IEO European Institute of Oncology IRCCS, Milan,
20139, Italy, Italy
-
Martin Schaefer, Department of Experimental Oncology, IEO European Institute of Oncology IRCCS, Milan,
20139, Italy, Italy
-
Giulio Caravagna, Department of Mathematics, Informatics and Geosciences, University of Trieste, Italy,
Italy
Presentation Overview: Show
Chromosomal instability (CIN) is a hallmark of solid cancers, encompassing copy number alterations (CNAs) and
structural variants that reshape genome content. Recurrent, cancer-type-specific patterns of chromosomal gains
and losses suggest that tumour subtypes follow distinct, temporally ordered evolutionary trajectories. Extreme
events such as chromothripsis and whole-genome doubling can generate bursts of alterations consistent with
punctuated modes of evolution, yet their prevalence and biological impact on disease progression remain only
partially understood.
To systematically characterize CIN evolution from whole-genome sequencing data, we introduce TickTack, a
hierarchical Bayesian mixture model that uses bulk WGS data from a single timepoint to simultaneously
reconstruct the temporal ordering of CNAs and identify clusters of co-occurring events across the genome. The
model distinguishes whether alterations accumulate gradually or arise in rapid succession, directly informing
our understanding of gradual versus saltational tumour evolution. TickTack was validated on synthetic data and
benchmarked against competing approaches.
We applied TickTack to 5,885 tumours from the Pan-Cancer Analysis of Whole Genomes (PCAWG) and the Genomics
England 100,000 Genomes Project, recovering cancer-type-specific co-occurring CNA patterns and recurrent
evolutionary trajectories across cohorts. We identified a subset of tumours, which we term hopeful monsters,
in which a remarkable fraction of the genome was altered at a single time point, consistent with saltational
dynamics and linked to elevated genomic instability. By examining how signatures of positive and negative
selection distribute across early and late CNA events, we further illuminate the selective forces shaping
tumour evolution, with potential implications for therapeutic target identification and prognosis.
B-G.C.28: MariNET: A Linear Mixed Model Framework for Network Analysis of Longitudinal EHR Data
Track: General computational biology
- Marina Vargas-Fernandez, UGR, Spain
-
Jordi Martorell-Marugán, Foundation for the Promotion of Health and Biomedical Research in the Valencian
Region (FISABIO), Spain
-
Pedro Carmona-Saez, UGR, Spain
Presentation Overview: Show
The rapid expansion of healthcare data stored in electronic health records (EHRs) has created new
opportunities for biomedical research, enabling the study of complex, heterogeneous, and longitudinal clinical
information. However, extracting meaningful relationships among clinical variables remains challenging due to
high dimensionality, temporal dependence, and the presence of confounding factors. Traditional approaches,
such as Gaussian Graphical Models and Vector Autoregression, are often limited by strong assumptions,
including independence and stationarity, which restrict their applicability to real-world EHR data.
To address these challenges, we present MariNET, an R package designed for the construction and analysis of
networks derived from longitudinal clinical data. MariNET is based on linear mixed models, allowing the
inference of interactions between clinical variables while explicitly accounting for repeated measures,
subject-level variability, and confounding effects. The package provides a flexible and scalable framework for
modeling dynamic relationships, supporting both continuous and categorical variables, as well as unbalanced
and incomplete datasets commonly found in EHRs.
MariNET includes a comprehensive set of functionalities for data preprocessing, model fitting, and network
construction, enabling users to estimate associations between variables over time and represent them as
interpretable network structures. The package further facilitates customization of model specifications,
including the incorporation of fixed and random effects, and offers tools for assessing statistical
significance and robustness of inferred interactions. Designed with usability in mind, MariNET integrates
seamlessly within the R ecosystem, providing an accessible interface for researchers without requiring
advanced expertise in statistical modeling.
B-G.C.29: Shifting the spotlight to unexpected guests: A workflow for the taxonomic and functional profiling
of non-target microorganisms in transcriptomics datasets
Track: General computational biology
-
Laura Carmen Terron Camero, Institute of Parasitology and Biomedicine “López-Neyra†(IPBLN), Spanish
National Research Council (CSIC), Spain
-
Eduardo Andres-Leon, Institute of Parasitology and Biomedicine “López-Neyra†(IPBLN), Spanish National
Research Council (CSIC), Spain
-
Hubert Rehrauer, Functional Genomics Center Zurich, ETH Zurich/University of Zurich, Zurich, Switzerland,
Switzerland
-
Jose Luis Ruiz Rodriguez, Functional Genomics Center Zurich, ETH Zurich/University of Zurich, Zurich,
Switzerland, Switzerland
Presentation Overview: Show
Transcriptomics has become a cornerstone omics technique in modern biological research. While novel
single-cell and spatial approaches offer unprecedented layers of resolution, bulk RNA sequencing remains an
essential tool for addressing diverse biological questions. Consequently, an ever-growing wealth of
transcriptomics data is populating publicly accessible databases. In reference-based transcriptomics, there is
typically a fraction of sequencing reads that fails to align to the target organism and is routinely discarded
as noise or contamination. However, this unmapped fraction often harbours valuable microbial information.
Furthermore, while many microbiome studies rely on DNA-based approaches for precise taxonomic identification,
applying meta-transcriptomics to leverage incidental microbial transcripts within larger samples can provide a
critical proxy for investigating molecular functions and biological pathways that are active in situ.
Here, conforming to the reusability aspect of the FAIR principles, we implemented a workflow to mine both
novel and previously published transcriptomics datasets that did not originally target microorganisms. As a
proof of concept, we aimed to uncover active, unexpected microbiota that remain undiscovered but hold
biological relevance. We also benchmarked k-mer- and marker gene-based approaches, demonstrating the need to
carefully select tools and databases while fine-tuning default parameters.
Despite technical limitations inherent to original experimental designs, RNA extraction, and library
preparation methods optimized for other purposes, we consistently uncovered valuable microbiota-related
insights, particularly regarding host-pathogen interactions and vector-borne diseases. This work highlights
the immense potential of data reusability to drive meta-transcriptomic discovery across a variety of
biological systems and fields where bulk transcriptomics is already routinely applied.