View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

A-G.C.01: DifFracTion: Systematic normalization and differential interaction identification for cross-sample comparison of Hi-C datasets
Track: General computational biology
  • J. Carlos Angel, Department of Biomedical Informatics, Columbia University, New York, NY, 10032, United States, United States
  • Gamze Gürsoy, Department of Applied Mathematics and Theoretical Physics, University of Cambridge, United Kingdom


Presentation Overview: Show

The three-dimensional (3D) organization of the genome plays a critical role in gene regulation, and disruptions to this architecture have been implicated in a wide range of diseases. Hi-C enables genome-wide measurement of chromatin interactions, motivating the development of methods for comparative analysis across biological conditions. However, most existing comparative approaches treat chromatin interactions as independent events, ignoring the strong distance-dependent decay of interaction frequencies, which limits their ability to detect differential chromatin interactions. Moreover, these methods fail when the Hi-C matrices being compared are sequenced at different depths. Here, we introduce DifFracTion, an algorithm that addresses these limitations through an iterative normalization approach that explicitly accounts for distance-dependent interaction frequency decay while correcting for sequencing depth imbalances, followed by a subsampling-based framework for estimating statistical significance. Under a strict null scenario with no biological differences, DifFracTion controls false positives across multiple resolutions, achieving zero false positive rate when sequencing-depth differences are modest and maintaining low FPR even under extreme imbalance. Additionally, DifFracTion accurately identifies true differential chromatin interactions across a wide range of fold-change magnitudes, achieving perfect accuracy and specificity, and exhibits a controllable trade-off between sensitivity and precision while outperforming both single- and multi-replicate approaches.

A-G.C.03: ROTS 2.0: A reproducibility-driven framework for complex statistical models in high-throughput omics
Track: General computational biology
  • Tomi Suomi, Turku Bioscience, University of Turku, Finland
  • Laura Elo, Turku Bioscience, University of Turku, Finland


Presentation Overview: Show

Reproducibility is a fundamental requirement for generating reliable and impactful scientific findings, particularly in the context of high-dimensional omics data such as transcriptomics and proteomics. A central task in these studies is differential expression analysis, which aims to identify molecular features whose expression levels differ across biological conditions. The datasets are often complex, characterized by a large number of features measured in relatively small sample sizes, technical noise, and sampling variability. Ensuring that detected differentially expressed features are reproducible across independent studies and experimental conditions is therefore critical.

The reproducibility-optimized test statistic (ROTS) framework was developed to address these challenges by prioritizing features that demonstrate high reproducibility. ROTS accommodates linear models, linear mixed-effects models, two-group comparisons, multi-group comparisons, and survival analysis. These modeling approaches enable robust identification of differentially expressed features in longitudinal studies, time-to-event analyses, and other complex experimental designs commonly encountered in clinical and systems biology research. The incorporation of model flexibility allows researchers to assess differential expression while accounting for confounding variables, repeated measures, and time-dependent effects, improving the biological relevance of the findings.

The reproducibility-optimization framework is validated using both simulated datasets and real-world omics studies, demonstrating that it consistently outperforms conventional statistical methods in identifying reproducible and biologically meaningful features, even under challenging conditions with noise and variability. Unlike the conventional approaches that focus on statistical significance, ROTS explicitly optimizes the statistic for reproducibility and generally offers a more robust criterion for feature selection in high-throughput data analysis.

A-G.C.04: Emergent Operon Structure in Genomic Language Models
Track: General computational biology
  • Antoine Lambert, KU Leuven - VIB.AI, Belgium
  • Arthur Valentin, KU Leuven - VIB.AI, Belgium
  • Luna Ceyssens, KU Leuven, Belgium
  • Joana Pereira, KU Leuven - VIB.AI, Belgium


Presentation Overview: Show

Recent genomic foundation models (gLMs) such as Nucleotide Transformer, DNABERT-2, and EVO2 have demonstrated strong performance across sequence-based tasks. However, it remains unclear to what extent these models encode higher-order genomic organization, such as operon structure, within their latent representations.

In this work, we benchmark state-of-the-art gLMs for operon detection in prokaryotic genomes, focusing on two complementary strategies: zero-shot inference and task-specific fine-tuning. These strategies help us investigate how operon relationships can be recovered from the embedding space of gLMs.

Beyond predictive performance, we aim to understand what these models learn about genomic organization. To this end, we introduce a series of controlled perturbation experiments, including gene duplication and sequence-level modifications outside annotated operons. These experiments probe whether models rely on local sequence features, positional context, or broader genomic signals to encode operon membership, and assess the stability of these representations under controlled genomic alterations. We hypothesize that gLMs encode operon structure as an emergent property of contextualized representations, such that functionally related genes exhibit coherent organization in latent space even in the absence of explicit supervision.

This work provides a quantitative and mechanistic assessment of how genomic foundation models organize nucleotide information in latent space. More broadly, it establishes a foundation for interpreting whether and how regulatory architecture is implicitly captured by large-scale genomic models, and whether these representations generalize across species.

A-G.C.05: From fragmented MAGs to complete chromosomes: refining functional diversity in uncultured Patescibacteria
Track: General computational biology
  • Sanchita Kamath, Helmholtz Centre for Environmental Research - UFZ, Germany


Presentation Overview: Show

Most metagenome-assembled genomes (MAGs) derived from short-read sequencing data remain fragmented, limiting structural and functional interpretation. This challenge is particularly evident in Patescibacteria, an abundant and uncultured bacterial phylum with ultra-small genomes, reduced metabolic capacity, and putative symbiotic lifestyles. Fragmentation frequently results in incomplete or misinterpreted genomic features, obscuring genome architecture, replication dynamics, and metabolic potential. In contrast, long-read sequencing approaches remain cost-prohibitive and difficult to implement at scale.
JORG, an iterative short-read assembly and binning framework, was applied to circularize Patescibacterial MAGs from marine and terrestrial environments. Based on data-driven prioritization criteria (≤10 contigs; 25–300× coverage), 128 high-quality MAGs were selected, and 60 (46.9%) were successfully circularized into single-chromosome assemblies.
Circularization significantly increased genome completeness in 76.6% of cases (paired t-test, p < 0.001) without increasing contamination, demonstrating robust structural refinement. Circularized genomes exhibited improved structural coherence and interpretability. Structural alignments revealed insertions and deletions in 86% of genomes, consistent with resolution of low-confidence regions and assembly gaps. Circularization refined near-threshold functional assignments, including KEGG modules linked to central carbon metabolism and metabolite transport, strengthening inference of metabolic dependencies and potential host interactions, and enabled unambiguous identification and chromosomal contextualization of replication origins.
These improvements were observed in genomes already classified as high quality, indicating that circularization enhances biological resolution beyond standard completeness metrics. Together, these findings demonstrate that automated genome circularization substantially improves ecological and evolutionary interpretation of uncultured microbial lineages such as Patescibacteria across diverse environments.

A-G.C.06: Ocrelizumab causes transient changes in the composition of intestinal bacteria and modulates the immune response depending on the response to treatment
Track: General computational biology
  • Miloslav Kverka, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Eva Kubala Havrdova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech Republic, Czechia
  • Helena Tlaskalova-Hogenova, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Jakub Kreisinger, Faculty of Science, Department of Zoology, Charles University, Prague, Czech Republic, Czechia
  • Ivana Kovarova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech Republic, Czechia
  • Jana Lizrova Preiningerova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech Republic, Czechia
  • Pavlina Kleinova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech Republic, Czechia
  • Miluse Pavelcova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech Republic, Czechia
  • Veronika Ticha, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech Republic, Czechia
  • Stepan Coufal, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Tomas Hrncir, Institute of Microbiology of the Czech Academy of Sciences, Novy Hradek, Czech Republic, Czechia
  • Ruth Tachezy, Faculty of Science, Charles University, BIOCEV, Vestec, Czech Republic, Czechia
  • Martina Salakova, Faculty of Science, Charles University, BIOCEV, Vestec, Czech Republic, Czechia
  • Dominika Kadleckova, Faculty of Science, Charles University, BIOCEV, Vestec, Czech Republic, Czechia
  • Radka Roubalova, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Tomas Thon, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic, Czechia
  • Zuzana Jiraskova Zakostelska, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic, Czechia


Presentation Overview: Show

Multiple sclerosis (MS) is an autoimmune disease that leads to the loss of myelin and atrophy of the central nervous system. The role of gut microbiota dysbiosis has been implicated in MS pathogenesis and may also influence treatment outcomes.

In our study, we included 25 newly diagnosed persons with MS (PwMS) with clinically isolated syndrome (CIS), which were treatment-naïve. The eighty-one healthy control subjects were also recruited. Stool and serum in study groups were collected before the treatment and every 3 months for a minimum of 12 months.

We identified changes in the gut microbiota that are already present in CIS persons who are naive to MS treatment. Gut bacteria alterations were transient during the first 12 months of anti-CD20 therapy. After 12 months, responders showed increased gut microbiota alpha diversity approaching healthy control levels, while non-responders showed a significant decline. Key changes involved Parabacteroides spp., producers of short-chain fatty acids that support gut barrier function and have anti-inflammatory potential. We detected altered gut barrier biomarkers and antibodies against common commensals in MS patients, which were modulated by anti-CD20 treatment. Notably, lipopolysaccharide-binding protein and mannose-binding lectin decreased only in responders.

These findings suggest that intestinal barrier damage contributes to immune responses linked to microbial translocation, MS pathogenesis, and treatment outcomes.

This research was supported by grants from the Ministry of Education, Youth and Sports of the Czech Republic grant Talking microbes-understanding microbial interactions within One Health framework (CZ.02.01.01/00/22_008/0004597).

A-G.C.07: Efficient Population Structure-Adjusted Logistic Regression for Genome-wide Association Interaction Studies via Clustered Covariates
Track: General computational biology
  • Volker Neff, Institute of Clinical Molecular Biology (IKMB), Kiel University, Germany
  • Lars Wienbrandt, Institute of Clinical Molecular Biology, Kiel University, Germany
  • David Ellinghaus, Institute of Clinical Molecular Biology, Kiel University, Germany


Presentation Overview: Show

Genome-wide association interaction studies (GWAIS) provide a robust framework for detecting epistasis, yet they present significant computational challenges due to the extensive search space. Logistic regression remains the standard method for testing statistical interactions; however, covariates that account for population structure, typically derived from principal component analysis (PCA), are often excluded to improve computational efficiency. This exclusion may result in inflated test statistics and biased effect estimates.

This work introduces a novel approximation that combines clustered proxy-covariates with contingency tables to substantially reduce the computational complexity of covariate-adjusted logistic regression for epistasis detection. Rather than including per-sample PCA covariates, population structure is summarized using a limited number of proxy-covariates. This approach reduces runtime complexity from O(N I) to O(N + IK), where N denotes the number of samples, I the number of regression iterations, and K the number of proxy-covariates, while maintaining statistical accuracy relative to the full covariate-adjusted model, as measured by mean relative error (MRE).

The proposed method was evaluated on a real-world case-control genome-wide association study (GWAS) dataset with hospitalized COVID-19 patients and healthy controls. The clustered proxy-covariates reduced the MRE by up to 68% compared to logistic regression without covariates, and achieved speedups of up to 92-fold with 5 clusters relative to logistic regression with per-sample PCA covariates. These findings indicate that covariate clustering enables rapid and statistically meaningful GWAIS analyses, thereby making population-structure-adjusted interaction testing feasible at the genome-wide scale.

A-G.C.08: The SIB RDF knowledge graph network: FAIR in practice
Track: General computational biology
  • Saadia Ismail, SciCore, University of Basel, Switzerland
  • Frédérique Lisacek, Swiss Institute of Bioinformatics, Switzerland
  • Kasun Samarasinghe, Swiss Institute of Bioinformatics, Switzerland
  • Nicole Redaschi, Swiss Institute of Bioinformatics, Switzerland
  • Panayiotis Smeros, Swiss Institute of Bioinformatics, Switzerland
  • Paul Thomas, Swiss Institute of Bioinformatics, Switzerland
  • Parit Bansal, Swiss Institute of Bioinformatics, Switzerland
  • Pierre-Andre Michel, Swiss Institute of Bioinformatics, Switzerland
  • Robin Engler, Swiss Institute of Bioinformatics, Switzerland
  • Frédéric Burdet, Swiss Institute of Bioinformatics, Switzerland
  • Sébastien Gehant, Swiss Institute of Bioinformatics, Switzerland
  • Sébastien Moretti, Swiss Institute of Bioinformatics, Switzerland
  • Séverine Duvaud, Swiss Institute of Bioinformatics, Switzerland
  • Shubham Kapoor, Swiss Institute of Bioinformatics, Switzerland
  • Silvano Alda, Swiss Institute of Bioinformatics, Switzerland
  • Sofia Georgakopoulou, SciCore, University of Basel, Switzerland
  • Vassilios Ioannidis, Swiss Institute of Bioinformatics, Switzerland
  • Vasundra Toure, Swiss Institute of Bioinformatics, Switzerland
  • Tarcisio Mendes, Swiss Institute of Bioinformatics, Switzerland
  • Ana Claudia Sima, Swiss Institute of Bioinformatics, Switzerland
  • Marco Pagni, Swiss Institute of Bioinformatics, Switzerland
  • Vincent Emonet, Swiss Institute of Bioinformatics, Switzerland
  • Deepak Unni, Swiss Institute of Bioinformatics, Switzerland
  • Sabine Oesterle, Swiss Institute of Bioinformatics, Switzerland
  • Dmitry Kuznetsov, Swiss Institute of Bioinformatics, Switzerland
  • Orlin Topalov, Swiss Institute of Bioinformatics, Switzerland
  • Ruijie Wang, Swiss Institute of Bioinformatics, Switzerland
  • Jerven Bolleman, Swiss Institute of Bioinformatics, Switzerland
  • Florence Mehl, Swiss Insitute of Bioinformatics, Switzerland
  • Monique Zahn, Swiss Institute of Bioinformatics, Switzerland
  • Adrian Altenhoff, Swiss Institute of Bioinformatics, Switzerland
  • Alan Bridge, Swiss Institute of Bioinformatics, Switzerland
  • Davide Chiarugi, Swiss Institute of Bioinformatics, Switzerland
  • Delphine Baratin, Swiss Institute of Bioinformatics, Switzerland
  • Ekaterina Stepanova, Swiss Institute of Bioinformatics, Switzerland
  • Florent Tassy, Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

The Resource Description Framework (RDF) aligns with the FAIR principles by design, with its emphasis on durable data access via persistent identifiers (PIDs) and on interoperability through SPARQL (a standardized query language for knowledge graphs). The SIB Swiss Institute of Bioinformatics has fostered a broad bottom-up network of RDF knowledge graph providers who maintain a rich collection of resources that are publicly accessible via SPARQL endpoints, including several ELIXIR Core Data Resources. They cover areas such as orthology, gene expression, proteins, reactions, chemicals, lipids, glycans, cell lines and clinical care data. The SIB also maintains internal SPARQL endpoints that support quality assurance processes or contain sensitive data.

To facilitate the use of the SIB's public SPARQL endpoints, we publish a curated set of competency questions at https://sib-swiss.github.io/sparql-examples/, focusing in particular on federated queries across multiple resources. This set of questions is also used to guide large language models (LLMs) in constructing SPARQL queries. Major commercial LLM services are able to interact with our endpoints directly via a Model Context Protocol (MCP) server (https://github.com/sib-swiss/sparql-llm), thus allowing users who are unfamiliar with the SPARQL query language and the structure of our RDF knowledge graphs to pose questions in natural language. A demo chat client aimed at developers who want to discover SPARQL and our RDF resources is available at https://www.expasy.org/chat.

A-G.C.09: Detecting Antibiotic Resistance in Drug-Free Conditions: An AI-Powered Morphological Approach in Klebsiella pneumoniae
Track: General computational biology
  • Amr Mostafa, Berliner Hochschule für Technik (BHT), Germany
  • Mario Koddenbrock, Hochschule für Technik und Wirtschaft (HTW), Germany
  • Erik Rodner, Hochschule für Technik und Wirtschaft (HTW), Germany
  • Elisabeth Grohmann, Berliner Hochschule für Technik (BHT), Germany


Presentation Overview: Show

Antimicrobial resistance is one of the most pressing challenges in clinical microbiology, yet current susceptibility testing remains slow and dependent on antibiotic exposure of viable cultures, delaying treatment decisions by 24-48 hours. We ask whether resistance can instead be detected directly from the intrinsic morphology of bacterial cells, without antibiotic exposure. We present a computational platform that couples high-resolution fluorescence microscopy with machine learning to identify morphological signatures of resistance in Klebsiella pneumoniae.
Our approach combines two complementary strain panels: laboratory-evolved resistant strains, providing a controlled genetic background with reduced biological noise, and clinical and environmental isolates, capturing real-world resistance mechanisms and clinical relevance. Fluorescent imaging of these populations feeds an end-to-end pipeline of automated preprocessing, segmentation, and multi-parameter morphological feature extraction (area, perimeter, axis lengths, circularity, Feret diameters), producing quantitative single-cell profiles at scale.
These profiles train AI classifiers that contrast resistant and susceptible populations and compare patterns between laboratory-evolved and naturally occurring resistant strains, with a validation step targeting predictive accuracy and clinical applicability. Preliminary results confirm robust extraction of discriminative morphological feature distributions across strain panels, establishing the basis for the resistance-classification models now under development.
By reframing susceptibility testing as a computer-vision problem on drug-free cells, this work outlines a path toward rapid, culture-light resistance prediction relevant to diagnostic microbiology and the broader integration of AI into clinical decision-making.

A-G.C.10: SwissLipids2.0: Towards a Semantically Integrated Lipid Knowledge Resource
Track: General computational biology
  • Jerven Bolleman, Swiss Institute of Bioinformatics, Switzerland
  • Lucila Aimo, SIB Swiss Institute of Bioinformatics, Switzerland
  • Teresa Neto, SIB, Swiss Institute of Bioinformatics, Switzerland
  • Nicole Redaschi, Swiss Institute of Bioinformatics, Switzerland
  • Alan Bridge, Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

SwissLipids is a comprehensive reference knowledgebase for lipids and lipidomics, housing approximately 600,000 theoretically feasible lipid structures, generated computationally through SMILES enumeration using lipid classes and fatty acyl groups that are mapped to ChEBI and curated in Rhea and UniProtKB.

In this presentation we describe the development of SwissLipids v2.0, which aims at streamlining the underlying database architecture, enriching lipid annotation, and deepening integration with the broader biochemical knowledge ecosystem.

SwissLipids v2.0 exploits a new RDF data model for lipids, and the in-built mapping to ChEBI created during lipid structure enumeration, to source up-to-date annotations from resources using ChEBI through federated SPARQL queries.

This turns SwissLipids from a static resource into a lightweight, integrative layer continuously enriched by the ongoing curation efforts of UniProtKB/Swiss-Prot, Rhea, and other resources using ChEBI such as the GO, Reactome, and MetaboLights, and marks the transition of SwissLipids toward a scalable, FAIR, and fully integrated component of the biochemical data ecosystem.

A-G.C.11: Exploring Metabolite-Based Cluster Patterns Associated with Severity in Inflammatory Bowel Disease
Track: General computational biology
  • Cezar-Fabian Moise, Aalborg University, Denmark
  • Malene Revsbech Christiansen, Aalborg University, Denmark
  • Marie Vibeke Vestergaard, Aalborg University, Denmark
  • Tine Jess, Aalborg University, Denmark
  • Filip Ottosson, Aalborg University, Denmark


Presentation Overview: Show

Inflammatory bowel disease (IBD), encompassing Crohn’s disease (CD) and ulcerative colitis (UC), is a chronic immune-mediated condition with a heterogeneous clinical course that is increasingly being explored through advances in omics, particularly metabolomics. This study investigates whether untargeted serum metabolite profiles, acquired via LC-MS processing, are associated with IBD severity in 183 patients sampled two years before to one year after diagnosis. Unsupervised clustering was applied to non-linear metabolite projections, enabling a metabolite-based approach followed by exploratory machine learning on the resulting clusters. More specifically, five machine learning models were evaluated to identify associations between metabolite clusters and disease severity, with Random Forest and Support Vector Classifier showing the best performance for CD and UC, respectively. Notably, clusters formed around chemically homogeneous groups, with molecules such as hippurate and indole-propionic acid linked to non-severe CD cases, whereas glycerophosphocholines were associated with severe CD but non-severe UC. Our findings underscore the need for subtype-specific modeling approaches and demonstrate that well-designed metabolite-based analyses in IBD can serve as valuable tools for uncovering or validating severity-related patterns, particularly when combined with mechanistic and clinical interpretation.

A-G.C.12: Sex-Specific Analyses Reveal Differences in Genetic Architecture and Cardiometabolic Associations of the Retinal Microvasculature
Track: General computational biology
  • Leah Bottger, University of Lausanne, Switzerland
  • Dennis Bontempi , University of Lausanne, Switzerland
  • Olga Trofimova , University of Lausanne, Switzerland
  • Sacha Bors , University of Lausanne, Switzerland
  • Ilaria Iuliani , University of Lausanne, Switzerland
  • Ian Quintas, University of Lausanne, Switzerland
  • David Presby , University of Lausanne, Switzerland
  • Sven Bergmann , University of Lausanne, Switzerland


Presentation Overview: Show

Sex as a biological variable (SABV) is increasingly recognized in complex trait genetics, yet its role in shaping retinal microvascular phenotypes and their genetic architecture remains largely unexplored. Retinal vascular traits are emerging as non-invasive biomarkers for ocular, systemic vascular, and neurodegenerative diseases. Leveraging the UK Biobank (N=68k), we conducted sex-stratified and sex-interaction genome-wide association studies (GWAS), including the X chromosome, to dissect sex-differential genetic architecture of quantitative retinal vascular features. These analyses revealed suggestive trends toward more associated variants and higher heritability in females compared to males. Statistical gene analyses identified 12 unique genes; 3 significant in females and 9 in males, implicating distinct pathways in microvascular development and remodeling, and highlighting sex-specific genetic contributions. Notably, we report the first X-chromosome GWAS associations for retinal vascular morphology, supporting sex-chromosomal influences on microvascular structure. Sex-differential patterns also emerged in cardiometabolic trait associations and prognostic Cox models for mortality and cardiovascular events, underscoring limitations of sex-agnostic approaches. Our findings demonstrate the value of SABV-aware imaging-genetics for revealing nuanced sex effects, with implications for computational phenotyping and precision vascular risk assessment.

A-G.C.13: From rosettes to cells: multi-scale shape plasticity in Arabidopsis thaliana under different temperatures
Track: General computational biology
  • Adeleh Dehghani Nazhvani, University of Potsdam, Germany
  • Prabal Das, Max Planck Institute of Molecular Plant Physiology, Germany
  • Jacqueline Nowak, University of Potsdam, Max Planck Institute of Molecular Plant Physiology, Germany
  • Arun Sampathkumar, Max Planck Institute of Molecular Plant Physiology, Germany


Presentation Overview: Show

Understanding how plant morphology adapts to environmental conditions is central to uncovering mechanisms of phenotypic plasticity. Temperature, as a key environmental factor, influences growth across multiple biological scales; however, how these effects propagate from whole-plant to cellular organization remains insufficiently characterized.
Here, we analyze and compare shape plasticity across 24 Arabidopsis thaliana accessions grown at 17°C and 27°C, integrating morphological traits at both the tissue scale (rosette and leaf) and the cellular scale. We quantify both common descriptors (e.g., area and circularity) and level-specific metrics, including the spatial distribution of leaf growth, rosette symmetry, leaf elongation, and cellular shape complexity (e.g., area and lobeyness).
Our multi-scale framework reveals substantial morphological plasticity across accessions in response to temperature variation.
These findings highlight the extent of morphological plasticity and demonstrate the impact of temperature on multi-level trait organization. providing insight into genotype-specific strategies of structural adaptation under different thermal conditions.

A-G.C.14: Automatic differentiation enables efficient 13C isotopically non-stationary metabolic flux analysis
Track: General computational biology
  • Fayaz Soleymani, University of Potsdam, Germany
  • Zoran Nikoloski, Max Planck Institute of Molecular Plant Physiology, Germany


Presentation Overview: Show

Metabolic flux analysis (MFA) is essential for quantifying intracellular metabolic fluxes. Isotopically non-stationary MFA (INST-MFA) with 13C-labeled CO2 enables the estimation of key parameters–metabolic fluxes and intracellular metabolite pool sizes in autotrophic organisms, where isotopic steady-state labeling patterns are non-informative of intracellular fluxes. However, applications of INST-MFA remains limited by high computational demands and challenges in parameter estimation.
We developed a computational framework for 13C INST-MFA that leverages automatic differentiation to improve optimization efficiency. The approach uses the IPOPT optimizer to integrate automatic differentiation within a system of ordinary differential equations (ODEs), describing the incorporation of 13C label in metabolic pools, and to estimate fluxes and metabolite pool sizes from time-resolved isotopic labeling data. The framework was implemented in Python and evaluated on a metabolic model of central carbon metabolism of Synechocystis sp. PCC 6803 with 60 reactions and 31 metabolites with synthetic and real-world labeling data.
Compared to conventional approaches, our method accurately estimates metabolic fluxes and metabolite pool sizes, together with 90% confidence intervals, while reducing computational time by at least two-fold and improving goodness-of-fit, measured by the reduced X2 statistic. These results demonstrate the potential of automatic differentiation to enhance the scalability and robustness of 13C INST-MFA workflows.

A-G.C.15: How Phosphorylation Affects Peptide Interaction with Adaptor Domains
Track: General computational biology
  • Debarshee Sengupta, Saarland University, Germany
  • Volkhard Helms, Saarland University, Germany


Presentation Overview: Show

Protein-protein interactions (PPIs) are of fundamental relevance to numerous cellular processes and hence the focus of considerable experimental and computational efforts. They are often modulated by post-translational modifications such as phosphorylation, which can either favor or disfavor the formation of a particular complex. Molecular dynamics (MD) simulations offer a powerful framework for investigating the binding affinities, association kinetics, and structural pathways of protein complexes, including how mutations or phosphorylation events modulate these interactions. As a model system, we are studying PDZ domains in complex with phosphorylated vs. non-phosphorylated C-terminal peptides. Using available X-ray structures of Scribble and DLG PDZ domains, we computed relative binding free energy differences between phosphorylated and non-phosphorylated peptides. In the case of Scribble PDZ1, MD simulations with the CHARMM forcefield yielded free energy differences that closely matched experimental values, validating the applicability of this approach for studying phospho/non-phosphopeptide interactions. Can this be extended to systems lacking experimental support? We found that Alpha-Fold predicted PDZ-peptide systems that lacked NMR or X-ray data provided reliable starting points for MD simulations and free energy calculations. Together with improved water models (e.g., TIP4P) and adjusted phosphate group parameters, our results indicated stable interactions and binding free energies within the range of experimental observations. These results highlight the potential of combining structure prediction tools with MD simulations to study phosphorylation-dependent interactions.

A-G.C.16: Navigating Career Pathways with the EMBL-EBI Competency Hub
Track: General computational biology
  • Daria Sokolova, EMBL-EBI, United Kingdom
  • Kim Gurwitz, EMBL-EBI, United Kingdom
  • Catherine Brooksbank, EMBL-EBI, United Kingdom
  • Vera Matser, The Alan Turing Institute, United Kingdom
  • Denise Bianco, The Alan Turing Institute, United Kingdom
  • Giulia Tomba, The Alan Turing Institute, United Kingdom
  • Emma Karoune, The Alan Turing Institute, United Kingdom


Presentation Overview: Show

In today's skills-based economy, professional success depends on demonstrating measurable competencies rather than relying on static job titles. A competency-based approach provides a structured ""common language"" that clarifies career paths and helps organisations build resilient teams through better role definition and targeted development.
The Competency Hub from EMBL's European Bioinformatics Institute (EMBL-EBI) is a specialised web infrastructure designed to operationalise these frameworks. As a centralised, open-access repository, it empowers users to assess their attributes (KSAs/KSBs) against recognised frameworks, explore professional personas, and identify training resources mapped to specific skills gaps.
Access to such a repository is critical in biomedical data science, where the convergence of biology, medicine, and computation requires a highly adaptable workforce. The Hub's utility is demonstrated by the Advancing Biomedical Data Science Careers (ABDC) project led by The Alan Turing Institute and EMBL-EBI. We're utilising the Hub to host a comprehensive competency mapping that synthesises existing frameworks to identify 12 Minimum Standard Requirements and specific personas. Following this mapping, a gap analysis will identify training needs to be reflected in the Hub.
Furthermore, the Hub supports granular professional extensions, such as the recently published career pathway for bioinformatics core facility scientists, serving as a blueprint for other specialists, such as biocurators. By hosting these outputs, the Competency Hub translates complex ecosystem analyses into practical, live templates for organisations to standardise roles and support inclusive workforces. Additionally, we developed the ""10 Simple Rules for producing FAIR competency frameworks"" to foster a connected global training ecosystem.

A-G.C.17: AmpliPhy improves gene trees by adding homologous sequences without affecting alignments
Track: General computational biology
  • Dongwook Kim, University of Lausanne, Switzerland
  • Manuel Gil, Zürich University of Applied Sciences, Switzerland
  • Kazutaka Katoh, University of Tokyo, Japan
  • Christophe Dessimoz, University of Lausanne, Switzerland


Presentation Overview: Show

In phylogenomics, gene tree reconstruction depends on multiple sequence alignment (MSA) and tree inference, and ongoing work continues to improve inference quality. Denser taxon sampling has been associated with improved gene tree inference, suggesting that adding homologs could be a practical route to higher accuracy as sequence databases continue to expand. However, adding sequences can influence multiple steps of typical inference pipelines, and little is known on its specific effect on the multiple sequence alignment, tree reconstruction, and rooting steps. We performed a large-scale empirical and simulated benchmarks to quantify how homolog enrichment affects alignment and phylogenetic inference. Using an enrichment-impoverishment design and a measure of tree accuracy based on taxonomic congruence, we found that enrichment consistently improves tree inference quality, while effects on alignment quality are marginal. We show that this improvement is associated with, but not restricted to accurate root placement on enriched trees when sensitive homolog search is accompanied. Notably, much of the benefit can be retained with relatively compact alignments produced by sequence addition. Building on these observations, we provide a tool, AmpliPhy, which efficiently improves phylogenetic reconstruction of protein families through homolog enrichment. The AmpliPhy open-source pipeline software is available at https://github.com/DessimozLab/ampliphy.

A-G.C.18: Towards FAIR and Privacy-Preserving Synthetic Health Data: A Reproducible Pipeline for Generation and Evaluation
Track: General computational biology
  • Georgy Lepsaya, Integrative Bioinformatics Group, RÄ«ga Stradiņš University, Riga, Latvia, Latvia
  • Maksims Ivanovs, Institute of Electronics and Computer Science (EDI), Riga, Latvia, Latvia
  • Baiba Vilne, Integrative Bioinformatics Group, RÄ«ga Stradiņš University, Riga, Latvia, Latvia


Presentation Overview: Show

Machine learning (ML) in health research depends on large, shareable datasets, yet frameworks such as the General Data Protection Regulation (GDPR) and the Health Insurance Portability and Accountability Act (HIPAA) restrict access to individual-level data, creating a fundamental bottleneck for model development, validation, and reproducibility. Synthetic tabular data generation has emerged as a promising solution, but two challenges remain unresolved: the privacy–utility tradeoff in generative models and the lack of standardised workflows for publishing synthetic datasets in a FAIR-compliant manner.

We present a reproducible Nextflow pipeline that addresses both challenges. The pipeline integrates hyperparameter-optimised synthetic data generation, systematic evaluation, and automated metadata annotation using the Croissant schema extended with a custom provenance vocabulary capturing model configurations, hyperparameters, and quality metrics. The output is a publication-ready research object suitable for direct deposition on standard data platforms.

To quantify the privacy–utility tradeoff, we compared a privacy-focused generative model (ADS-GAN) with a utility-oriented baseline (CTGAN) across several tabular health datasets from diverse disease domains. ADS-GAN consistently achieved stronger privacy protection at the cost of reduced predictive utility, empirically confirming the fundamental tradeoff in synthetic data generation.

By integrating generation, evaluation, and FAIR-compliant packaging into a single portable workflow, the proposed pipeline enables both responsible sharing of synthetic health data and their effective use for developing and validating ML models, supporting reproducible and Open Science–compliant research. Future work will extend benchmarking to larger and more heterogeneous datasets and explore models that better balance privacy and utility.

A-G.C.19: Conformal Inference for Gene-Specific Cell Subsets in Single-Cell Differential Expression
Track: General computational biology
  • Justine Leclerc, University of Zürich/University Hospital of Zürich, Switzerland
  • Christof Seiler, University of Zürich/University Hospital of Zürich, Switzerland


Presentation Overview: Show

Differential expression analysis in single-cell data requires predefining cell groups in which expression is assumed to be homogeneous. Standard methods cluster cells using nearest-neighbor graphs before testing for differential expression. An alternative is to infer structure after differential analysis and identify groups with shared responses, without pre-defined cluster labels. This is relevant in cancer, where tumor heterogeneity leads to subpopulations with shared treatment responses. Pre-clustering can miss such signals.

We build on the lemur R-package. This method avoids prior clustering and models gene expression through a low-dimensional latent representation of cells with covariate-specific structure. It estimates individual treatment effects (ITEs) at the single-cell level. For each gene, it identifies a cell subset forming a region in latent space with coherent differential signal. These neighborhoods can capture patterns which might be missed by clustering before differential testing.

We extend lemur in two directions. First, we build prediction intervals for lemur ITEs estimates using conformal prediction. This method turns point predictions into sets with finite-sample, distribution-free coverage under exchangeability. Second, we construct prediction sets for each cells membership within a gene neighborhood. This quantifies uncertainty for cluster assignments. Singleton sets indicate low uncertainty and sets of size two reflect higher uncertainty by retaining multiple plausible outcomes under the same coverage guarantee.

We provide visualizations, theoretical guarantees, and empirical validation on simulated and real datasets. We combine latent-space differential modeling with conformal uncertainty quantification to reveal gene-specific cell structure potentially missed by standard clustering.

A-G.C.20: A haplotype reference panel constructed from 490,085 UK Biobank genomes improves genotype imputation for global populations
Track: General computational biology
  • Lars Wienbrandt, Institute of Clinical Molecular Biology, Kiel University, Germany
  • Volker Neff, Institute of Clinical Molecular Biology, Kiel University, Germany
  • Maria Gretsova, Institute of Clinical Molecular Biology, Kiel University, Germany
  • Guillermo Torres, Institute of Clinical Molecular Biology, Kiel University, Germany
  • Eike Matthias Wacker, Institute of Clinical Molecular Biology, Kiel University, Germany
  • Andre Franke, Institute of Clinical Molecular Biology, Kiel University, Germany
  • David Ellinghaus, Institute of Clinical Molecular Biology, Kiel University, Germany


Presentation Overview: Show

We constructed a genotype imputation reference panel from whole-genome sequence data of 490,085 UK Biobank participants and implemented cloud-based phasing and imputation through EagleImp-RAP within the UK Biobank Research Analysis Platform (UKB-RAP). We benchmarked the panel using 1000 Genomes Project data and four real-world COVID-19 GWAS datasets from Italy, Germany, Spain and Norway.

Compared with the widely used TOPMed r3 reference panel (133,597 individuals), the UK Biobank reference panel achieved lower phasing switch error rates for all global superpopulations except the Admixed American (AMR) and for 25 of 26 1000 Genomes Project subpopulations.

Imputation accuracy, measured by mean absolute error, improved for four of five superpopulations and 12 subpopulations for variants with minor allele frequencies down to 0.01%. Estimated squared correlation (R²) further improved across all five superpopulations and 22 subpopulations, with the largest gains in East and South Asian populations. Re-imputation of the COVID-19 GWAS datasets showed that common-variant imputation is not yet saturated, improving the discovery of genome-wide significant loci and fine-mapping resolution.

EagleImp-RAP with the UKB reference panel is currently being integrated into UKB-RAP as a UKB-RAP application for general use by UKB-approved projects.

A-G.C.21: Building Regional AI Capacity in Biosciences: The BiotrAIn Training Model for Latin America
Track: General computational biology
  • Jose Arturo Molina-Mora, Universidad de Costa Rica, Costa Rica
  • Kim Gurtwitz, EMBL-EBI, United Kingdom
  • Rebeca Campos-Sánchez, Universidad de Costa Rica, Costa Rica
  • Lizzie Divala, EMBL-EBI, United Kingdom
  • Juanita Riveros, EMBL-EBI, United Kingdom
  • Cindy Aguilar-Bartels, Universidad de Costa Rica, Costa Rica
  • Cath Brooksbank, EMBL-EBI, United Kingdom


Presentation Overview: Show

The rapid expansion of AI in biosciences presents significant opportunities, yet Latin America faces persistent structural barriers to equitable participation, including limited infrastructure, fragmented research ecosystems, and insufficient access to contextually relevant training. Addressing this gap requires a scalable, sustainable model for how training is designed, coordinated, and delivered across diverse regional contexts.
The BiotrAIn project, a partnership between EMBL-EBI, the University of Costa Rica, and CABANAnet, funded by the Chan Zuckerberg Initiative, was built around three phases: knowledge exchange, collaborative curriculum development, and structured training delivery. The model was validated through a pilot course reaching 25 participants, establishing proof of concept for a regionally adapted, open-access AI curriculum covering core concepts, and practical bioscience applications developed by Latin American faculty and early-career scientists.
BiotrAIn has since scaled to a hub-and-spoke delivery model spanning seven sites and reaching 165 participants. This structure enables localised participation while maintaining curriculum coherence across the region, and embeds a train-the-trainer track to ensure sustainability.The effectiveness of this approach is demonstrated by pilot participants now serving as hosts across South America, embedding local expertise and driving regional ownership of the programme
This work describes the coordination model underpinning this growth, how the programme was scoped, partnerships structured, and delivery sequenced to scale from pilot to multi-site programme, offering transferable lessons for teams facing comparable barriers elsewhere. The freely available curriculum, adaptable to different teaching and learning environments, provides a practical foundation for computational biologists seeking to implement similar initiatives.

A-G.C.22: Graph-based Machine Learning approaches for predicting microbiome composition
Track: General computational biology
  • Fatemeh Rajaei Nesheli, Queen's University Belfast, United Kingdom
  • Chris Creevey, Queen's University Belfast, United Kingdom
  • Huiru Zheng, Ulster University, United Kingdom
  • John-Paul Wilkins, Queen's University Belfast, United Kingdom


Presentation Overview: Show

Current microbiome research focuses on describing community composition and function within individual studies, but it remains limited in its ability to predict likely microbial communities and their functional potential given the presence of specific taxa or environmental changes. This limitation is particularly evident in studies based on 16S rRNA sequencing, which is widely used because of its cost-effectiveness but provides limited taxonomic resolution and only indirect functional insight.
Taxonomy-free frameworks such as Life Identification Numbers (LINs) provide a stable, genome-based representation of microbial diversity, where hierarchical relationships between lineages capture similarity at multiple resolutions, but their use remains underexplored in predictive microbiome modelling.
Recently the Creevey lab have developed a LIN-like taxonomy-free classification approach for 16S microbiome data for the entire Greengenes2 database. Here, we propose a computational framework that integrates this with machine learning for predictive microbiome analysis. First, we model microbial community structure by learning lineage-level co-occurrence patterns using graph-based approaches that leverage the hierarchical structure of LINs to predict coexisting lineages given a target lineage. Second, we extend this framework to functional inference by linking LINs to pangenome-derived gene content. Using graph-based propagation across related lineages, we estimate gene presence–absence profiles and aggregate these into community-level functional predictions mapped to metabolic pathways.
This framework provides a scalable and consistent approach for linking microbiome composition to function across heterogeneous environments and supports a shift from descriptive to predictive microbiome analysis using widely available 16S data.

A-G.C.23: CNETML2: A Multi-scale Model for Copy Number Alteration Evolution in Cancer
Track: General computational biology
  • Ruolin Wu, University of Surrey, United Kingdom
  • Bingxin Lu, University of Surrey, United Kingdom


Presentation Overview: Show

Somatic copy number alterations (CNAs) are pervasive in cancer, particularly in tumours driven by chromosomal instability. Their strong association with disease progression and adverse clinical outcomes makes understanding how CNAs accumulate and change over time essential for early detection, prognosis, and personalized therapy.
Although numerous studies have characterized CNA patterns and relative timing, few have estimated CNA rates in absolute chronological time. We previously developed CNETML, a maximum-likelihood framework for reconstructing CNA-based phylogenetic trees and estimating CNA evolutionary rates from shallow whole-genome sequencing data. Applications of CNETML to cancer datasets have enabled quantitative reconstruction of tumour evolutionary histories. However, existing models typically assume a single class of CNA events and constant evolutionary dynamics, whereas cancer genomes often undergo multiple types of large-scale alterations that shape copy number evolution.
Here we extend CNETML with a multi-scale model that explicitly incorporates three classes of genomic events: segment-level duplications and deletions, chromosome-level gains and losses, and whole-genome doubling. These processes are represented by three transition matrices describing independent Markov chains for each alteration scale. The likelihood is computed dynamically by combining these matrices during phylogenetic inference, allowing different CNA processes to contribute jointly to genome evolution. To improve computational efficiency, the model prunes redundant haplotype copy-number states that are incompatible with the constraining maximum copy number boundaries, substantially reducing the size of the state space. Implemented within the CNETML framework, the new model enables more efficient and biologically realistic inference of tumour phylogenies and CNA evolutionary rates from cancer sequencing data.

A-G.C.24: Metadata Accelerator: Improving scientific data descriptions with Natural Language Processing methods (NLP) and Feedback
Track: General computational biology
  • Maria Juliana Rodriguez Cubillos, Centre for Engineering Biology, School of Biological Sciences and School of Informatics, University of Edinburgh, UK, United Kingdom
  • Tomasz ZieliÅ„ski, Centre for Engineering Biology, School of Biological Sciences, University of Edinburgh, Edinburgh EH9 3BF, UK, United Kingdom
  • Jason R. Swedlow, Divisions of Molecular Cell and Developmental Biology, and Computational Biology, University of Dundee, Dundee, UK, United Kingdom
  • T. Ian Simpson, School of Informatics, University of Edinburgh, 10 Crichton Street, Edinburgh EH8 9AB, UK, United Kingdom
  • Andrew J. Millar, Centre for Engineering Biology, School of Biological Sciences, University of Edinburgh, Edinburgh EH9 3BF, UK, United Kingdom


Presentation Overview: Show

Ensuring the availability and accessibility of data is fundamental to advancing knowledge. This idea has been codified as the FAIR principles (Findable, Accessible, Interoperable, and Reusable) in scientific data management. Accurate documentation of studies, commonly known as metadata, is indispensable for achieving this goal. Regrettably, scientific records often fall short, providing inadequate, repetitive, or incomplete descriptions that hinder the seamless flow of knowledge. This project addresses the metadata challenge by identifying common failure points in descriptions and generating better-structured metadata using NLP and LLMs, with a focus on sustainability.

The Metadata Accelerator is a pipeline that aims to predict metadata using information from free-text descriptions. Based on initial user inputs, the pipeline suggests improvements that fit template-based schemas and provides feedback using quality metrics. Users can then refine the suggested changes before export, allowing us to engage users and explore whether richer interaction improves metadata quality.

The pilot version of the pipeline predicts “species” and “study type” for each entry from user-provided descriptions using a combination of NLP methods and LLMs, including Llama 3.2 and BioBERT. The Image Data Resource (IDR) and BioDare2, a repository for circadian and biological data, were selected to develop a pipeline that suggests metadata terms from existing free-text descriptions. For IDR, the system identified 97 species compared with 110 in a manual review and matched 130 of 132 study-type entries.

Ongoing work will evaluate performance on BioDare2 and other repositories, improve efficiency, and expand to additional categories.

A-G.C.25: Discovering Novel Ciliary Genes Through Machine Learning
Track: General computational biology
  • Emilia Torriglia, Institut Pasteur de Montevideo, Uruguay
  • Florencia Irigoin, Institut Pasteur de Montevideo, Uruguay
  • Laura Romanelli-Cedrez, Institut Pasteur de Montevideo, Uruguay
  • Flavio Pazos Obregón, Institut Pasteur de Montevideo, Uruguay


Presentation Overview: Show

The primary cilium is an evolutionarily conserved organelle present in many eukaryotic cells, where it plays a central role in cellular signaling and environmental sensing. Despite its importance, the full set of genes required for its structure and function remains incomplete. Caenorhabditis elegans provides a powerful system to study ciliary biology due to its well-characterized nervous system and ease of genetic and experimental manipulation, enabling the identification of conserved mechanisms across species. We hypothesized that genes associated with ciliary function share expression-based features and other biological characteristics, independent of DNA sequence similarity, that can be used for their identification.

To investigate this, transcriptomic datasets were constructed from publicly available single-cell RNA sequencing data (Packer et al. 2019; Taylor et al. 2021). These datasets include neuronal cell types across developmental stages, comprising both ciliated and non-ciliated populations, and were integrated prior to analysis. Genes were filtered based on expression across different cell types and variability, and annotated ciliary genes were split into enrichment analysis and evaluation sets. Expression values were standardized, and agglomerative hierarchical clustering was applied across multiple distance thresholds. Enrichment analysis identified clusters overrepresented in ciliary genes, from which candidate genes were defined. Candidates were compared across conditions and stages, and the method was evaluated by quantifying the enrichment of evaluation-set genes within enriched clusters.

In future work, additional biological features independent of DNA sequence, such as genomic localization of functional groups and structural motifs, will be incorporated, alongside supervised machine learning approaches to improve candidate prioritization.

A-G.C.26: COMPAS: COntrastive Multimodal Polypharmacology-Aware multi-target drug-target interaction Screener
Track: General computational biology
  • Gamze Deprem, Hacettepe University, Turkey
  • Tunca Doğan, Hacettepe University, Turkey


Presentation Overview: Show

Complex diseases such as cancer, diabetes, and neurodegenerative disorders are driven by multiple interconnected biological mechanisms, limiting the effectiveness of traditional one-drug-one-target strategies. Controlled polypharmacology, in which a single molecule modulates multiple disease-relevant targets in a selective and coordinated manner, has therefore emerged as a promising therapeutic paradigm. However, current computational drug repurposing and discovery approaches remain largely focused on single-target interactions, use biomolecular modalities in limited ways, and do not directly model multi-target interaction patterns, thereby restricting generalizability. In this study, COMPAS (Contrastive Multimodal Polypharmacology-Aware Screener) is proposed as a new machine learning model for directly predicting interactions between small molecules and multiple targets. COMPAS utilizes molecular embeddings obtained from pretrained multimodal models using SMILES and graph representations, while protein embeddings are derived from pretrained multimodal protein language models using sequence, structure, text modalities. These embeddings are projected into a shared latent space through trainable projection layers. A molecule-centered, multi-positive supervised contrastive learning strategy is then applied to organize this space, so that each molecule is positioned closer to its interacting proteins and farther from irrelevant targets. In this way, coordinated multi-target interaction profiles are learned within a unified multimodal setting. The proposed model is developed using data from public resources: ChEMBL, PubChem, ZINC, and UniProt. Alzheimer's disease is adopted as a use case focusing on AChE and BACE-1, and predicted candidates are intended for molecular docking analysis. COMPAS is expected to improve the identification of biologically relevant multi-target drug candidates and support scalable drug discovery for complex diseases.

A-G.C.27: Improving comparative genomics with orthology uncertainty metrics
Track: General computational biology
  • Veronica Sondervan, VIB-UGent Center for Plant Systems Biology, Belgium
  • Yves Van de Peer, VIB-UGent Center for Plant Systems Biology, Belgium
  • Zhen Li, VIB-UGent Center for Plant Systems Biology, Belgium


Presentation Overview: Show

Accurate identification of orthologs across species is critical for avoiding spurious gene-loss identification, inflated family expansions, and unreliable selection inferences in comparative genomics. Studies have shown that current orthology methods can vary widely in their orthology calls, but there remains no quantified method for determining the uncertainty of an ortholog group. Here, we develop an orthology uncertainty scoring system by integrating metrics including sequence similarity, synteny, and phylogenetic support. We also evaluate whether protein language model embeddings can provide complementary uncertainty signal. We benchmark our system with Dendrobium orchids, a genus known for morphological differences associated with gene loss, including transitions between lithophytic v.s. epiphytic lifestyles and self-compatibility v.s. incompatibility. Applying established orthology methods, we show that our uncertainty scoring scheme aids in identifying and confirming gene losses and spotting poorly supported orthology calls. Our work represents an important step forward in increasing the precision and certainty of evolutionary investigations into species history and adaptation.

A-G.C.28: FLKit: A Community-Driven Onboarding Toolkit for Federated Analytics and Learning in Health and Life Sciences
Track: General computational biology
  • Ashkan Pirmani, KU Leuven - Uhasselt, Belgium
  • Ilse Vermeulen, Uhasselt, Belgium
  • Goran Vinterhalter, KU Leuven, Belgium
  • Lotte Geys, UHasselt, Belgium
  • Axel Faes, UHasselt, Belgium
  • Muhammad Quamber Ali, KU Leuven, Belgium
  • Nshkala Sattanathan, UAntwerpen, Belgium
  • Geert Vandeweyer, UAntwerpen, Belgium
  • Yves Moreau, KU Leuven, Belgium
  • Liesbet Peeters, UHasselt, Belgium


Presentation Overview: Show

Federated learning (FL) allows institutions to collaboratively train models without sharing raw sensitive data, making it a natural fit for health and life sciences research bound by strict data protection regulations. Despite growing technical maturity, most researchers, clinicians, data stewards, and legal advisors entering this space encounter a fragmented landscape of frameworks, governance requirements, and terminology with no structured starting point tailored to their role.
FLKit addresses this gap as an open, community-maintained onboarding toolkit that guides multidisciplinary teams through the full federated learning lifecycle. Inspired by the ELIXIR Research Data Management Kit (RDMkit) and developed using a design science approach, FLKit is organised around four lifecycle stages (Governance, Infrastructure, Wrangling, and Analysis), 11 role-specific entry points spanning clinical researchers, data stewards, legal advisors, and research software engineers, a curated glossary, a FAIR-aligned FL Story template for documenting real-world use cases, and a directory of tools, frameworks, and communities.
Since its demo launch in December 2024, FLKit has grown to 39 pages across eight content sections. Seven FL Stories document completed and ongoing projects across multiple sclerosis disability prediction, inflammatory bowel disease, genomics, and ECoG-based brain-computer interfaces.
FLKit is openly available at https://uhasselt-biomedicaldatasciences.github.io/federated-learning-toolkit/ and is actively inviting community contributions from across the life sciences.

A-G.C.29: HOROSCOPE: Decoding human centromere architecture from short reads using k-mer signatures
Track: General computational biology
  • Carsten Hain, EMBL, Germany
  • Tobias Rausch, EMBL, Germany
  • Jan Korbel, EMBL, Germany


Presentation Overview: Show

By directing kinetochore formation and chromosome segregation, centromeres safeguard genome integrity. Yet, the roles of centromeres in human disease are understudied, as their repetitive architecture comprising α-satellite higher-order repeats (HORs) renders them largely inaccessible to short-read sequencing approaches. Here we develop HOROSCOPE (Higher-Order Repeat organization and Size of Centromeres using Oligonucleotide Profiles for Estimation), a computational framework for k-mer-based inference of centromere structure and length from short-read data. Based on a reference atlas of 11,836 human centromeres extracted from completely assembled (telomere-to-telomere) haplotypes, we systematically interrogate chromosome-specific centromere architectures, deriving architecture-specific k-mer signatures as well as centromeric length-informative k-mers from common to rare centromere architectures. Leveraging these diagnostic k-mers, HOROSCOPE achieves a precision of 99.3% and a recall of 99.5% in classifying chromosome-specific centromere architectures from short-read de Bruijn graphs. We perform a population-scale analysis of global centromeric architectures in 4,029 human samples with ancestry from 80 human populations sequenced with short reads, uncovering continental haplotype structure for centromeric regions and highlighting African-enriched rare centromeric architectures. Furthermore, by analyzing 1,359 cancer genomes, we link graded HOR truncation events to arm-level copy-number alterations, and uncover a general dependency of chromosomal rearrangement locations on the position of the centromere dip region (CDR) defining the kinetochore attachment site. HOROSCOPE enables large-scale centromere genomics directly from short-read sample cohorts.