View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
A-G.C.01: DifFracTion: Systematic normalization and differential interaction identification for cross-sample
comparison of Hi-C datasets
Track: General computational biology
-
J. Carlos Angel, Department of Biomedical Informatics, Columbia University, New York, NY, 10032,
United States, United States
-
Gamze Gürsoy, Department of Applied Mathematics and Theoretical Physics, University of Cambridge, United
Kingdom
Presentation Overview: Show
The three-dimensional (3D) organization of the genome plays a critical role in gene regulation, and
disruptions to this architecture have been implicated in a wide range of diseases. Hi-C enables genome-wide
measurement of chromatin interactions, motivating the development of methods for comparative analysis across
biological conditions. However, most existing comparative approaches treat chromatin interactions as
independent events, ignoring the strong distance-dependent decay of interaction frequencies, which limits
their ability to detect differential chromatin interactions. Moreover, these methods fail when the Hi-C
matrices being compared are sequenced at different depths. Here, we introduce DifFracTion, an algorithm that
addresses these limitations through an iterative normalization approach that explicitly accounts for
distance-dependent interaction frequency decay while correcting for sequencing depth imbalances, followed by a
subsampling-based framework for estimating statistical significance. Under a strict null scenario with no
biological differences, DifFracTion controls false positives across multiple resolutions, achieving zero false
positive rate when sequencing-depth differences are modest and maintaining low FPR even under extreme
imbalance. Additionally, DifFracTion accurately identifies true differential chromatin interactions across a
wide range of fold-change magnitudes, achieving perfect accuracy and specificity, and exhibits a controllable
trade-off between sensitivity and precision while outperforming both single- and multi-replicate approaches.
A-G.C.03: ROTS 2.0: A reproducibility-driven framework for complex statistical models in high-throughput
omics
Track: General computational biology
-
Tomi Suomi, Turku Bioscience, University of Turku, Finland
- Laura Elo, Turku Bioscience, University of Turku, Finland
Presentation Overview: Show
Reproducibility is a fundamental requirement for generating reliable and impactful scientific findings,
particularly in the context of high-dimensional omics data such as transcriptomics and proteomics. A central
task in these studies is differential expression analysis, which aims to identify molecular features whose
expression levels differ across biological conditions. The datasets are often complex, characterized by a
large number of features measured in relatively small sample sizes, technical noise, and sampling variability.
Ensuring that detected differentially expressed features are reproducible across independent studies and
experimental conditions is therefore critical.
The reproducibility-optimized test statistic (ROTS) framework was developed to address these challenges by
prioritizing features that demonstrate high reproducibility. ROTS accommodates linear models, linear
mixed-effects models, two-group comparisons, multi-group comparisons, and survival analysis. These modeling
approaches enable robust identification of differentially expressed features in longitudinal studies,
time-to-event analyses, and other complex experimental designs commonly encountered in clinical and systems
biology research. The incorporation of model flexibility allows researchers to assess differential expression
while accounting for confounding variables, repeated measures, and time-dependent effects, improving the
biological relevance of the findings.
The reproducibility-optimization framework is validated using both simulated datasets and real-world omics
studies, demonstrating that it consistently outperforms conventional statistical methods in identifying
reproducible and biologically meaningful features, even under challenging conditions with noise and
variability. Unlike the conventional approaches that focus on statistical significance, ROTS explicitly
optimizes the statistic for reproducibility and generally offers a more robust criterion for feature selection
in high-throughput data analysis.
A-G.C.04: Emergent Operon Structure in Genomic Language Models
Track: General computational biology
-
Antoine Lambert, KU Leuven - VIB.AI, Belgium
- Arthur Valentin, KU Leuven - VIB.AI, Belgium
- Luna Ceyssens, KU Leuven, Belgium
- Joana Pereira, KU Leuven - VIB.AI, Belgium
Presentation Overview: Show
Recent genomic foundation models (gLMs) such as Nucleotide Transformer, DNABERT-2, and EVO2 have demonstrated
strong performance across sequence-based tasks. However, it remains unclear to what extent these models encode
higher-order genomic organization, such as operon structure, within their latent representations.
In this work, we benchmark state-of-the-art gLMs for operon detection in prokaryotic genomes, focusing on two
complementary strategies: zero-shot inference and task-specific fine-tuning. These strategies help us
investigate how operon relationships can be recovered from the embedding space of gLMs.
Beyond predictive performance, we aim to understand what these models learn about genomic organization. To
this end, we introduce a series of controlled perturbation experiments, including gene duplication and
sequence-level modifications outside annotated operons. These experiments probe whether models rely on local
sequence features, positional context, or broader genomic signals to encode operon membership, and assess the
stability of these representations under controlled genomic alterations. We hypothesize that gLMs encode
operon structure as an emergent property of contextualized representations, such that functionally related
genes exhibit coherent organization in latent space even in the absence of explicit supervision.
This work provides a quantitative and mechanistic assessment of how genomic foundation models organize
nucleotide information in latent space. More broadly, it establishes a foundation for interpreting whether and
how regulatory architecture is implicitly captured by large-scale genomic models, and whether these
representations generalize across species.
A-G.C.05: From fragmented MAGs to complete chromosomes: refining functional diversity in uncultured
Patescibacteria
Track: General computational biology
-
Sanchita Kamath, Helmholtz Centre for Environmental Research - UFZ, Germany
Presentation Overview: Show
Most metagenome-assembled genomes (MAGs) derived from short-read sequencing data remain fragmented, limiting
structural and functional interpretation. This challenge is particularly evident in Patescibacteria, an
abundant and uncultured bacterial phylum with ultra-small genomes, reduced metabolic capacity, and putative
symbiotic lifestyles. Fragmentation frequently results in incomplete or misinterpreted genomic features,
obscuring genome architecture, replication dynamics, and metabolic potential. In contrast, long-read
sequencing approaches remain cost-prohibitive and difficult to implement at scale.
JORG, an iterative short-read assembly and binning framework, was applied to circularize Patescibacterial MAGs
from marine and terrestrial environments. Based on data-driven prioritization criteria (≤10 contigs;
25–300× coverage), 128 high-quality MAGs were selected, and 60 (46.9%) were successfully circularized into
single-chromosome assemblies.
Circularization significantly increased genome completeness in 76.6% of cases (paired t-test, p < 0.001)
without increasing contamination, demonstrating robust structural refinement. Circularized genomes exhibited
improved structural coherence and interpretability. Structural alignments revealed insertions and deletions in
86% of genomes, consistent with resolution of low-confidence regions and assembly gaps. Circularization
refined near-threshold functional assignments, including KEGG modules linked to central carbon metabolism and
metabolite transport, strengthening inference of metabolic dependencies and potential host interactions, and
enabled unambiguous identification and chromosomal contextualization of replication origins.
These improvements were observed in genomes already classified as high quality, indicating that
circularization enhances biological resolution beyond standard completeness metrics. Together, these findings
demonstrate that automated genome circularization substantially improves ecological and evolutionary
interpretation of uncultured microbial lineages such as Patescibacteria across diverse environments.
A-G.C.06: Ocrelizumab causes transient changes in the composition of intestinal bacteria and modulates the
immune response depending on the response to treatment
Track: General computational biology
-
Miloslav Kverka, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic,
Czechia
-
Eva Kubala Havrdova, First Medical Faculty, Charles University and General Medical Hospital in Prague,
Czech Republic, Czechia
-
Helena Tlaskalova-Hogenova, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech
Republic, Czechia
-
Jakub Kreisinger, Faculty of Science, Department of Zoology, Charles University, Prague, Czech Republic,
Czechia
-
Ivana Kovarova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech
Republic, Czechia
-
Jana Lizrova Preiningerova, First Medical Faculty, Charles University and General Medical Hospital in
Prague, Czech Republic, Czechia
-
Pavlina Kleinova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech
Republic, Czechia
-
Miluse Pavelcova, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech
Republic, Czechia
-
Veronika Ticha, First Medical Faculty, Charles University and General Medical Hospital in Prague, Czech
Republic, Czechia
-
Stepan Coufal, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic,
Czechia
-
Tomas Hrncir, Institute of Microbiology of the Czech Academy of Sciences, Novy Hradek, Czech Republic,
Czechia
-
Ruth Tachezy, Faculty of Science, Charles University, BIOCEV, Vestec, Czech Republic, Czechia
-
Martina Salakova, Faculty of Science, Charles University, BIOCEV, Vestec, Czech Republic, Czechia
-
Dominika Kadleckova, Faculty of Science, Charles University, BIOCEV, Vestec, Czech Republic, Czechia
-
Radka Roubalova, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic,
Czechia
-
Tomas Thon, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech Republic, Czechia
-
Zuzana Jiraskova Zakostelska, Institute of Microbiology of the Czech Academy of Sciences, Prague, Czech
Republic, Czechia
Presentation Overview: Show
Multiple sclerosis (MS) is an autoimmune disease that leads to the loss of myelin and atrophy of the central
nervous system. The role of gut microbiota dysbiosis has been implicated in MS pathogenesis and may also
influence treatment outcomes.
In our study, we included 25 newly diagnosed persons with MS (PwMS) with clinically isolated syndrome (CIS),
which were treatment-naïve. The eighty-one healthy control subjects were also recruited. Stool and serum in
study groups were collected before the treatment and every 3 months for a minimum of 12 months.
We identified changes in the gut microbiota that are already present in CIS persons who are naive to MS
treatment. Gut bacteria alterations were transient during the first 12 months of anti-CD20 therapy. After 12
months, responders showed increased gut microbiota alpha diversity approaching healthy control levels, while
non-responders showed a significant decline. Key changes involved Parabacteroides spp., producers of
short-chain fatty acids that support gut barrier function and have anti-inflammatory potential. We detected
altered gut barrier biomarkers and antibodies against common commensals in MS patients, which were modulated
by anti-CD20 treatment. Notably, lipopolysaccharide-binding protein and mannose-binding lectin decreased only
in responders.
These findings suggest that intestinal barrier damage contributes to immune responses linked to microbial
translocation, MS pathogenesis, and treatment outcomes.
This research was supported by grants from the Ministry of Education, Youth and Sports of the Czech Republic
grant Talking microbes-understanding microbial interactions within One Health framework
(CZ.02.01.01/00/22_008/0004597).
A-G.C.07: Efficient Population Structure-Adjusted Logistic Regression for Genome-wide Association Interaction
Studies via Clustered Covariates
Track: General computational biology
-
Volker Neff, Institute of Clinical Molecular Biology (IKMB), Kiel University, Germany
- Lars Wienbrandt, Institute of Clinical Molecular Biology, Kiel University, Germany
- David Ellinghaus, Institute of Clinical Molecular Biology, Kiel University, Germany
Presentation Overview: Show
Genome-wide association interaction studies (GWAIS) provide a robust framework for detecting epistasis, yet
they present significant computational challenges due to the extensive search space. Logistic regression
remains the standard method for testing statistical interactions; however, covariates that account for
population structure, typically derived from principal component analysis (PCA), are often excluded to improve
computational efficiency. This exclusion may result in inflated test statistics and biased effect
estimates.
This work introduces a novel approximation that combines clustered proxy-covariates with contingency tables to
substantially reduce the computational complexity of covariate-adjusted logistic regression for epistasis
detection. Rather than including per-sample PCA covariates, population structure is summarized using a limited
number of proxy-covariates. This approach reduces runtime complexity from O(N I) to O(N + IK), where N denotes
the number of samples, I the number of regression iterations, and K the number of proxy-covariates, while
maintaining statistical accuracy relative to the full covariate-adjusted model, as measured by mean relative
error (MRE).
The proposed method was evaluated on a real-world case-control genome-wide association study (GWAS) dataset
with hospitalized COVID-19 patients and healthy controls. The clustered proxy-covariates reduced the MRE by up
to 68% compared to logistic regression without covariates, and achieved speedups of up to 92-fold with 5
clusters relative to logistic regression with per-sample PCA covariates. These findings indicate that
covariate clustering enables rapid and statistically meaningful GWAIS analyses, thereby making
population-structure-adjusted interaction testing feasible at the genome-wide scale.
A-G.C.08: The SIB RDF knowledge graph network: FAIR in practice
Track: General computational biology
- Saadia Ismail, SciCore, University of Basel, Switzerland
- Frédérique Lisacek, Swiss Institute of Bioinformatics, Switzerland
- Kasun Samarasinghe, Swiss Institute of Bioinformatics, Switzerland
- Nicole Redaschi, Swiss Institute of Bioinformatics, Switzerland
- Panayiotis Smeros, Swiss Institute of Bioinformatics, Switzerland
- Paul Thomas, Swiss Institute of Bioinformatics, Switzerland
- Parit Bansal, Swiss Institute of Bioinformatics, Switzerland
- Pierre-Andre Michel, Swiss Institute of Bioinformatics, Switzerland
- Robin Engler, Swiss Institute of Bioinformatics, Switzerland
- Frédéric Burdet, Swiss Institute of Bioinformatics, Switzerland
- Sébastien Gehant, Swiss Institute of Bioinformatics, Switzerland
- Sébastien Moretti, Swiss Institute of Bioinformatics, Switzerland
- Séverine Duvaud, Swiss Institute of Bioinformatics, Switzerland
- Shubham Kapoor, Swiss Institute of Bioinformatics, Switzerland
- Silvano Alda, Swiss Institute of Bioinformatics, Switzerland
- Sofia Georgakopoulou, SciCore, University of Basel, Switzerland
- Vassilios Ioannidis, Swiss Institute of Bioinformatics, Switzerland
- Vasundra Toure, Swiss Institute of Bioinformatics, Switzerland
- Tarcisio Mendes, Swiss Institute of Bioinformatics, Switzerland
- Ana Claudia Sima, Swiss Institute of Bioinformatics, Switzerland
- Marco Pagni, Swiss Institute of Bioinformatics, Switzerland
- Vincent Emonet, Swiss Institute of Bioinformatics, Switzerland
- Deepak Unni, Swiss Institute of Bioinformatics, Switzerland
- Sabine Oesterle, Swiss Institute of Bioinformatics, Switzerland
- Dmitry Kuznetsov, Swiss Institute of Bioinformatics, Switzerland
- Orlin Topalov, Swiss Institute of Bioinformatics, Switzerland
- Ruijie Wang, Swiss Institute of Bioinformatics, Switzerland
-
Jerven Bolleman, Swiss Institute of Bioinformatics, Switzerland
- Florence Mehl, Swiss Insitute of Bioinformatics, Switzerland
- Monique Zahn, Swiss Institute of Bioinformatics, Switzerland
- Adrian Altenhoff, Swiss Institute of Bioinformatics, Switzerland
- Alan Bridge, Swiss Institute of Bioinformatics, Switzerland
- Davide Chiarugi, Swiss Institute of Bioinformatics, Switzerland
- Delphine Baratin, Swiss Institute of Bioinformatics, Switzerland
- Ekaterina Stepanova, Swiss Institute of Bioinformatics, Switzerland
- Florent Tassy, Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
The Resource Description Framework (RDF) aligns with the FAIR principles by design, with its emphasis on
durable data access via persistent identifiers (PIDs) and on interoperability through SPARQL (a standardized
query language for knowledge graphs). The SIB Swiss Institute of Bioinformatics has fostered a broad bottom-up
network of RDF knowledge graph providers who maintain a rich collection of resources that are publicly
accessible via SPARQL endpoints, including several ELIXIR Core Data Resources. They cover areas such as
orthology, gene expression, proteins, reactions, chemicals, lipids, glycans, cell lines and clinical care
data. The SIB also maintains internal SPARQL endpoints that support quality assurance processes or contain
sensitive data.
To facilitate the use of the SIB's public SPARQL endpoints, we publish a curated set of competency questions
at https://sib-swiss.github.io/sparql-examples/, focusing in particular on federated queries across multiple
resources. This set of questions is also used to guide large language models (LLMs) in constructing SPARQL
queries. Major commercial LLM services are able to interact with our endpoints directly via a Model Context
Protocol (MCP) server (https://github.com/sib-swiss/sparql-llm), thus allowing users who are unfamiliar with
the SPARQL query language and the structure of our RDF knowledge graphs to pose questions in natural language.
A demo chat client aimed at developers who want to discover SPARQL and our RDF resources is available at
https://www.expasy.org/chat.
A-G.C.09: Detecting Antibiotic Resistance in Drug-Free Conditions: An AI-Powered Morphological Approach in
Klebsiella pneumoniae
Track: General computational biology
-
Amr Mostafa, Berliner Hochschule für Technik (BHT), Germany
- Mario Koddenbrock, Hochschule für Technik und Wirtschaft (HTW), Germany
- Erik Rodner, Hochschule für Technik und Wirtschaft (HTW), Germany
- Elisabeth Grohmann, Berliner Hochschule für Technik (BHT), Germany
Presentation Overview: Show
Antimicrobial resistance is one of the most pressing challenges in clinical microbiology, yet current
susceptibility testing remains slow and dependent on antibiotic exposure of viable cultures, delaying
treatment decisions by 24-48 hours. We ask whether resistance can instead be detected directly from the
intrinsic morphology of bacterial cells, without antibiotic exposure. We present a computational platform that
couples high-resolution fluorescence microscopy with machine learning to identify morphological signatures of
resistance in Klebsiella pneumoniae.
Our approach combines two complementary strain panels: laboratory-evolved resistant strains, providing a
controlled genetic background with reduced biological noise, and clinical and environmental isolates,
capturing real-world resistance mechanisms and clinical relevance. Fluorescent imaging of these populations
feeds an end-to-end pipeline of automated preprocessing, segmentation, and multi-parameter morphological
feature extraction (area, perimeter, axis lengths, circularity, Feret diameters), producing quantitative
single-cell profiles at scale.
These profiles train AI classifiers that contrast resistant and susceptible populations and compare patterns
between laboratory-evolved and naturally occurring resistant strains, with a validation step targeting
predictive accuracy and clinical applicability. Preliminary results confirm robust extraction of
discriminative morphological feature distributions across strain panels, establishing the basis for the
resistance-classification models now under development.
By reframing susceptibility testing as a computer-vision problem on drug-free cells, this work outlines a path
toward rapid, culture-light resistance prediction relevant to diagnostic microbiology and the broader
integration of AI into clinical decision-making.
A-G.C.10: SwissLipids2.0: Towards a Semantically Integrated Lipid Knowledge Resource
Track: General computational biology
-
Jerven Bolleman, Swiss Institute of Bioinformatics, Switzerland
- Lucila Aimo, SIB Swiss Institute of Bioinformatics, Switzerland
- Teresa Neto, SIB, Swiss Institute of Bioinformatics, Switzerland
- Nicole Redaschi, Swiss Institute of Bioinformatics, Switzerland
- Alan Bridge, Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
SwissLipids is a comprehensive reference knowledgebase for lipids and lipidomics, housing approximately
600,000 theoretically feasible lipid structures, generated computationally through SMILES enumeration using
lipid classes and fatty acyl groups that are mapped to ChEBI and curated in Rhea and UniProtKB.
In this presentation we describe the development of SwissLipids v2.0, which aims at streamlining the
underlying database architecture, enriching lipid annotation, and deepening integration with the broader
biochemical knowledge ecosystem.
SwissLipids v2.0 exploits a new RDF data model for lipids, and the in-built mapping to ChEBI created during
lipid structure enumeration, to source up-to-date annotations from resources using ChEBI through federated
SPARQL queries.
This turns SwissLipids from a static resource into a lightweight, integrative layer continuously enriched by
the ongoing curation efforts of UniProtKB/Swiss-Prot, Rhea, and other resources using ChEBI such as the GO,
Reactome, and MetaboLights, and marks the transition of SwissLipids toward a scalable, FAIR, and fully
integrated component of the biochemical data ecosystem.
A-G.C.11: Exploring Metabolite-Based Cluster Patterns Associated with Severity in Inflammatory Bowel
Disease
Track: General computational biology
-
Cezar-Fabian Moise, Aalborg University, Denmark
- Malene Revsbech Christiansen, Aalborg University, Denmark
- Marie Vibeke Vestergaard, Aalborg University, Denmark
- Tine Jess, Aalborg University, Denmark
- Filip Ottosson, Aalborg University, Denmark
Presentation Overview: Show
Inflammatory bowel disease (IBD), encompassing Crohn’s disease (CD) and ulcerative colitis (UC), is a chronic
immune-mediated condition with a heterogeneous clinical course that is increasingly being explored through
advances in omics, particularly metabolomics. This study investigates whether untargeted serum metabolite
profiles, acquired via LC-MS processing, are associated with IBD severity in 183 patients sampled two years
before to one year after diagnosis. Unsupervised clustering was applied to non-linear metabolite projections,
enabling a metabolite-based approach followed by exploratory machine learning on the resulting clusters. More
specifically, five machine learning models were evaluated to identify associations between metabolite clusters
and disease severity, with Random Forest and Support Vector Classifier showing the best performance for CD and
UC, respectively. Notably, clusters formed around chemically homogeneous groups, with molecules such as
hippurate and indole-propionic acid linked to non-severe CD cases, whereas glycerophosphocholines were
associated with severe CD but non-severe UC. Our findings underscore the need for subtype-specific modeling
approaches and demonstrate that well-designed metabolite-based analyses in IBD can serve as valuable tools for
uncovering or validating severity-related patterns, particularly when combined with mechanistic and clinical
interpretation.
A-G.C.12: Sex-Specific Analyses Reveal Differences in Genetic Architecture and Cardiometabolic Associations
of the Retinal Microvasculature
Track: General computational biology
-
Leah Bottger, University of Lausanne, Switzerland
- Dennis Bontempi , University of Lausanne, Switzerland
- Olga Trofimova , University of Lausanne, Switzerland
- Sacha Bors , University of Lausanne, Switzerland
- Ilaria Iuliani , University of Lausanne, Switzerland
- Ian Quintas, University of Lausanne, Switzerland
- David Presby , University of Lausanne, Switzerland
- Sven Bergmann , University of Lausanne, Switzerland
Presentation Overview: Show
Sex as a biological variable (SABV) is increasingly recognized in complex trait genetics, yet its role in
shaping retinal microvascular phenotypes and their genetic architecture remains largely unexplored. Retinal
vascular traits are emerging as non-invasive biomarkers for ocular, systemic vascular, and neurodegenerative
diseases. Leveraging the UK Biobank (N=68k), we conducted sex-stratified and sex-interaction genome-wide
association studies (GWAS), including the X chromosome, to dissect sex-differential genetic architecture of
quantitative retinal vascular features. These analyses revealed suggestive trends toward more associated
variants and higher heritability in females compared to males. Statistical gene analyses identified 12 unique
genes; 3 significant in females and 9 in males, implicating distinct pathways in microvascular development and
remodeling, and highlighting sex-specific genetic contributions. Notably, we report the first X-chromosome
GWAS associations for retinal vascular morphology, supporting sex-chromosomal influences on microvascular
structure. Sex-differential patterns also emerged in cardiometabolic trait associations and prognostic Cox
models for mortality and cardiovascular events, underscoring limitations of sex-agnostic approaches. Our
findings demonstrate the value of SABV-aware imaging-genetics for revealing nuanced sex effects, with
implications for computational phenotyping and precision vascular risk assessment.
A-G.C.13: From rosettes to cells: multi-scale shape plasticity in Arabidopsis thaliana under different
temperatures
Track: General computational biology
-
Adeleh Dehghani Nazhvani, University of Potsdam, Germany
- Prabal Das, Max Planck Institute of Molecular Plant Physiology, Germany
-
Jacqueline Nowak, University of Potsdam, Max Planck Institute of Molecular Plant Physiology, Germany
- Arun Sampathkumar, Max Planck Institute of Molecular Plant Physiology, Germany
Presentation Overview: Show
Understanding how plant morphology adapts to environmental conditions is central to uncovering mechanisms of
phenotypic plasticity. Temperature, as a key environmental factor, influences growth across multiple
biological scales; however, how these effects propagate from whole-plant to cellular organization remains
insufficiently characterized.
Here, we analyze and compare shape plasticity across 24 Arabidopsis thaliana accessions grown at 17°C and
27°C, integrating morphological traits at both the tissue scale (rosette and leaf) and the cellular scale. We
quantify both common descriptors (e.g., area and circularity) and level-specific metrics, including the
spatial distribution of leaf growth, rosette symmetry, leaf elongation, and cellular shape complexity (e.g.,
area and lobeyness).
Our multi-scale framework reveals substantial morphological plasticity across accessions in response to
temperature variation.
These findings highlight the extent of morphological plasticity and demonstrate the impact of temperature on
multi-level trait organization. providing insight into genotype-specific strategies of structural adaptation
under different thermal conditions.
A-G.C.14: Automatic differentiation enables efficient 13C isotopically non-stationary metabolic flux
analysis
Track: General computational biology
-
Fayaz Soleymani, University of Potsdam, Germany
- Zoran Nikoloski, Max Planck Institute of Molecular Plant Physiology, Germany
Presentation Overview: Show
Metabolic flux analysis (MFA) is essential for quantifying intracellular metabolic fluxes. Isotopically
non-stationary MFA (INST-MFA) with 13C-labeled CO2 enables the estimation of key parameters–metabolic fluxes
and intracellular metabolite pool sizes in autotrophic organisms, where isotopic steady-state labeling
patterns are non-informative of intracellular fluxes. However, applications of INST-MFA remains limited by
high computational demands and challenges in parameter estimation.
We developed a computational framework for 13C INST-MFA that leverages automatic differentiation to improve
optimization efficiency. The approach uses the IPOPT optimizer to integrate automatic differentiation within a
system of ordinary differential equations (ODEs), describing the incorporation of 13C label in metabolic
pools, and to estimate fluxes and metabolite pool sizes from time-resolved isotopic labeling data. The
framework was implemented in Python and evaluated on a metabolic model of central carbon metabolism of
Synechocystis sp. PCC 6803 with 60 reactions and 31 metabolites with synthetic and real-world labeling
data.
Compared to conventional approaches, our method accurately estimates metabolic fluxes and metabolite pool
sizes, together with 90% confidence intervals, while reducing computational time by at least two-fold and
improving goodness-of-fit, measured by the reduced X2 statistic. These results demonstrate the potential of
automatic differentiation to enhance the scalability and robustness of 13C INST-MFA workflows.
A-G.C.15: How Phosphorylation Affects Peptide Interaction with Adaptor Domains
Track: General computational biology
-
Debarshee Sengupta, Saarland University, Germany
- Volkhard Helms, Saarland University, Germany
Presentation Overview: Show
Protein-protein interactions (PPIs) are of fundamental relevance to numerous cellular processes and hence the
focus of considerable experimental and computational efforts. They are often modulated by post-translational
modifications such as phosphorylation, which can either favor or disfavor the formation of a particular
complex. Molecular dynamics (MD) simulations offer a powerful framework for investigating the binding
affinities, association kinetics, and structural pathways of protein complexes, including how mutations or
phosphorylation events modulate these interactions. As a model system, we are studying PDZ domains in complex
with phosphorylated vs. non-phosphorylated C-terminal peptides. Using available X-ray structures of Scribble
and DLG PDZ domains, we computed relative binding free energy differences between phosphorylated and
non-phosphorylated peptides. In the case of Scribble PDZ1, MD simulations with the CHARMM forcefield yielded
free energy differences that closely matched experimental values, validating the applicability of this
approach for studying phospho/non-phosphopeptide interactions. Can this be extended to systems lacking
experimental support? We found that Alpha-Fold predicted PDZ-peptide systems that lacked NMR or X-ray data
provided reliable starting points for MD simulations and free energy calculations. Together with improved
water models (e.g., TIP4P) and adjusted phosphate group parameters, our results indicated stable interactions
and binding free energies within the range of experimental observations. These results highlight the potential
of combining structure prediction tools with MD simulations to study phosphorylation-dependent interactions.
A-G.C.16: Navigating Career Pathways with the EMBL-EBI Competency Hub
Track: General computational biology
-
Daria Sokolova, EMBL-EBI, United Kingdom
- Kim Gurwitz, EMBL-EBI, United Kingdom
- Catherine Brooksbank, EMBL-EBI, United Kingdom
- Vera Matser, The Alan Turing Institute, United Kingdom
- Denise Bianco, The Alan Turing Institute, United Kingdom
- Giulia Tomba, The Alan Turing Institute, United Kingdom
- Emma Karoune, The Alan Turing Institute, United Kingdom
Presentation Overview: Show
In today's skills-based economy, professional success depends on demonstrating measurable competencies rather
than relying on static job titles. A competency-based approach provides a structured ""common language"" that
clarifies career paths and helps organisations build resilient teams through better role definition and
targeted development.
The Competency Hub from EMBL's European Bioinformatics Institute (EMBL-EBI) is a specialised web
infrastructure designed to operationalise these frameworks. As a centralised, open-access repository, it
empowers users to assess their attributes (KSAs/KSBs) against recognised frameworks, explore professional
personas, and identify training resources mapped to specific skills gaps.
Access to such a repository is critical in biomedical data science, where the convergence of biology,
medicine, and computation requires a highly adaptable workforce. The Hub's utility is demonstrated by the
Advancing Biomedical Data Science Careers (ABDC) project led by The Alan Turing Institute and EMBL-EBI. We're
utilising the Hub to host a comprehensive competency mapping that synthesises existing frameworks to identify
12 Minimum Standard Requirements and specific personas. Following this mapping, a gap analysis will identify
training needs to be reflected in the Hub.
Furthermore, the Hub supports granular professional extensions, such as the recently published career pathway
for bioinformatics core facility scientists, serving as a blueprint for other specialists, such as
biocurators. By hosting these outputs, the Competency Hub translates complex ecosystem analyses into
practical, live templates for organisations to standardise roles and support inclusive workforces.
Additionally, we developed the ""10 Simple Rules for producing FAIR competency frameworks"" to foster a
connected global training ecosystem.
A-G.C.17: AmpliPhy improves gene trees by adding homologous sequences without affecting alignments
Track: General computational biology
-
Dongwook Kim, University of Lausanne, Switzerland
- Manuel Gil, Zürich University of Applied Sciences, Switzerland
- Kazutaka Katoh, University of Tokyo, Japan
- Christophe Dessimoz, University of Lausanne, Switzerland
Presentation Overview: Show
In phylogenomics, gene tree reconstruction depends on multiple sequence alignment (MSA) and tree inference,
and ongoing work continues to improve inference quality. Denser taxon sampling has been associated with
improved gene tree inference, suggesting that adding homologs could be a practical route to higher accuracy as
sequence databases continue to expand. However, adding sequences can influence multiple steps of typical
inference pipelines, and little is known on its specific effect on the multiple sequence alignment, tree
reconstruction, and rooting steps. We performed a large-scale empirical and simulated benchmarks to quantify
how homolog enrichment affects alignment and phylogenetic inference. Using an enrichment-impoverishment design
and a measure of tree accuracy based on taxonomic congruence, we found that enrichment consistently improves
tree inference quality, while effects on alignment quality are marginal. We show that this improvement is
associated with, but not restricted to accurate root placement on enriched trees when sensitive homolog search
is accompanied. Notably, much of the benefit can be retained with relatively compact alignments produced by
sequence addition. Building on these observations, we provide a tool, AmpliPhy, which efficiently improves
phylogenetic reconstruction of protein families through homolog enrichment. The AmpliPhy open-source pipeline
software is available at https://github.com/DessimozLab/ampliphy.
A-G.C.18: Towards FAIR and Privacy-Preserving Synthetic Health Data: A Reproducible Pipeline for Generation
and Evaluation
Track: General computational biology
-
Georgy Lepsaya, Integrative Bioinformatics Group, Rīga Stradiņš University, Riga, Latvia,
Latvia
-
Maksims Ivanovs, Institute of Electronics and Computer Science (EDI), Riga, Latvia, Latvia
-
Baiba Vilne, Integrative Bioinformatics Group, Rīga Stradiņš University, Riga, Latvia, Latvia
Presentation Overview: Show
Machine learning (ML) in health research depends on large, shareable datasets, yet frameworks such as the
General Data Protection Regulation (GDPR) and the Health Insurance Portability and Accountability Act (HIPAA)
restrict access to individual-level data, creating a fundamental bottleneck for model development, validation,
and reproducibility. Synthetic tabular data generation has emerged as a promising solution, but two challenges
remain unresolved: the privacy–utility tradeoff in generative models and the lack of standardised workflows
for publishing synthetic datasets in a FAIR-compliant manner.
We present a reproducible Nextflow pipeline that addresses both challenges. The pipeline integrates
hyperparameter-optimised synthetic data generation, systematic evaluation, and automated metadata annotation
using the Croissant schema extended with a custom provenance vocabulary capturing model configurations,
hyperparameters, and quality metrics. The output is a publication-ready research object suitable for direct
deposition on standard data platforms.
To quantify the privacy–utility tradeoff, we compared a privacy-focused generative model (ADS-GAN) with a
utility-oriented baseline (CTGAN) across several tabular health datasets from diverse disease domains. ADS-GAN
consistently achieved stronger privacy protection at the cost of reduced predictive utility, empirically
confirming the fundamental tradeoff in synthetic data generation.
By integrating generation, evaluation, and FAIR-compliant packaging into a single portable workflow, the
proposed pipeline enables both responsible sharing of synthetic health data and their effective use for
developing and validating ML models, supporting reproducible and Open Science–compliant research. Future
work will extend benchmarking to larger and more heterogeneous datasets and explore models that better balance
privacy and utility.
A-G.C.19: Conformal Inference for Gene-Specific Cell Subsets in Single-Cell Differential Expression
Track: General computational biology
-
Justine Leclerc, University of Zürich/University Hospital of Zürich, Switzerland
- Christof Seiler, University of Zürich/University Hospital of Zürich, Switzerland
Presentation Overview: Show
Differential expression analysis in single-cell data requires predefining cell groups in which expression is
assumed to be homogeneous. Standard methods cluster cells using nearest-neighbor graphs before testing for
differential expression. An alternative is to infer structure after differential analysis and identify groups
with shared responses, without pre-defined cluster labels. This is relevant in cancer, where tumor
heterogeneity leads to subpopulations with shared treatment responses. Pre-clustering can miss such
signals.
We build on the lemur R-package. This method avoids prior clustering and models gene expression through a
low-dimensional latent representation of cells with covariate-specific structure. It estimates individual
treatment effects (ITEs) at the single-cell level. For each gene, it identifies a cell subset forming a region
in latent space with coherent differential signal. These neighborhoods can capture patterns which might be
missed by clustering before differential testing.
We extend lemur in two directions. First, we build prediction intervals for lemur ITEs estimates using
conformal prediction. This method turns point predictions into sets with finite-sample, distribution-free
coverage under exchangeability. Second, we construct prediction sets for each cells membership within a gene
neighborhood. This quantifies uncertainty for cluster assignments. Singleton sets indicate low uncertainty and
sets of size two reflect higher uncertainty by retaining multiple plausible outcomes under the same coverage
guarantee.
We provide visualizations, theoretical guarantees, and empirical validation on simulated and real datasets. We
combine latent-space differential modeling with conformal uncertainty quantification to reveal gene-specific
cell structure potentially missed by standard clustering.
A-G.C.20: A haplotype reference panel constructed from 490,085 UK Biobank genomes improves genotype
imputation for global populations
Track: General computational biology
-
Lars Wienbrandt, Institute of Clinical Molecular Biology, Kiel University, Germany
- Volker Neff, Institute of Clinical Molecular Biology, Kiel University, Germany
- Maria Gretsova, Institute of Clinical Molecular Biology, Kiel University, Germany
- Guillermo Torres, Institute of Clinical Molecular Biology, Kiel University, Germany
-
Eike Matthias Wacker, Institute of Clinical Molecular Biology, Kiel University, Germany
- Andre Franke, Institute of Clinical Molecular Biology, Kiel University, Germany
- David Ellinghaus, Institute of Clinical Molecular Biology, Kiel University, Germany
Presentation Overview: Show
We constructed a genotype imputation reference panel from whole-genome sequence data of 490,085 UK Biobank
participants and implemented cloud-based phasing and imputation through EagleImp-RAP within the UK Biobank
Research Analysis Platform (UKB-RAP). We benchmarked the panel using 1000 Genomes Project data and four
real-world COVID-19 GWAS datasets from Italy, Germany, Spain and Norway.
Compared with the widely used TOPMed r3 reference panel (133,597 individuals), the UK Biobank reference panel
achieved lower phasing switch error rates for all global superpopulations except the Admixed American (AMR)
and for 25 of 26 1000 Genomes Project subpopulations.
Imputation accuracy, measured by mean absolute error, improved for four of five superpopulations and 12
subpopulations for variants with minor allele frequencies down to 0.01%. Estimated squared correlation (R²)
further improved across all five superpopulations and 22 subpopulations, with the largest gains in East and
South Asian populations. Re-imputation of the COVID-19 GWAS datasets showed that common-variant imputation is
not yet saturated, improving the discovery of genome-wide significant loci and fine-mapping resolution.
EagleImp-RAP with the UKB reference panel is currently being integrated into UKB-RAP as a UKB-RAP application
for general use by UKB-approved projects.
A-G.C.21: Building Regional AI Capacity in Biosciences: The BiotrAIn Training Model for Latin America
Track: General computational biology
- Jose Arturo Molina-Mora, Universidad de Costa Rica, Costa Rica
- Kim Gurtwitz, EMBL-EBI, United Kingdom
- Rebeca Campos-Sánchez, Universidad de Costa Rica, Costa Rica
-
Lizzie Divala, EMBL-EBI, United Kingdom
- Juanita Riveros, EMBL-EBI, United Kingdom
- Cindy Aguilar-Bartels, Universidad de Costa Rica, Costa Rica
- Cath Brooksbank, EMBL-EBI, United Kingdom
Presentation Overview: Show
The rapid expansion of AI in biosciences presents significant opportunities, yet Latin America faces
persistent structural barriers to equitable participation, including limited infrastructure, fragmented
research ecosystems, and insufficient access to contextually relevant training. Addressing this gap requires a
scalable, sustainable model for how training is designed, coordinated, and delivered across diverse regional
contexts.
The BiotrAIn project, a partnership between EMBL-EBI, the University of Costa Rica, and CABANAnet, funded by
the Chan Zuckerberg Initiative, was built around three phases: knowledge exchange, collaborative curriculum
development, and structured training delivery. The model was validated through a pilot course reaching 25
participants, establishing proof of concept for a regionally adapted, open-access AI curriculum covering core
concepts, and practical bioscience applications developed by Latin American faculty and early-career
scientists.
BiotrAIn has since scaled to a hub-and-spoke delivery model spanning seven sites and reaching 165
participants. This structure enables localised participation while maintaining curriculum coherence across the
region, and embeds a train-the-trainer track to ensure sustainability.The effectiveness of this approach is
demonstrated by pilot participants now serving as hosts across South America, embedding local expertise and
driving regional ownership of the programme
This work describes the coordination model underpinning this growth, how the programme was scoped,
partnerships structured, and delivery sequenced to scale from pilot to multi-site programme, offering
transferable lessons for teams facing comparable barriers elsewhere. The freely available curriculum,
adaptable to different teaching and learning environments, provides a practical foundation for computational
biologists seeking to implement similar initiatives.
A-G.C.22: Graph-based Machine Learning approaches for predicting microbiome composition
Track: General computational biology
-
Fatemeh Rajaei Nesheli, Queen's University Belfast, United Kingdom
- Chris Creevey, Queen's University Belfast, United Kingdom
- Huiru Zheng, Ulster University, United Kingdom
- John-Paul Wilkins, Queen's University Belfast, United Kingdom
Presentation Overview: Show
Current microbiome research focuses on describing community composition and function within individual
studies, but it remains limited in its ability to predict likely microbial communities and their functional
potential given the presence of specific taxa or environmental changes. This limitation is particularly
evident in studies based on 16S rRNA sequencing, which is widely used because of its cost-effectiveness but
provides limited taxonomic resolution and only indirect functional insight.
Taxonomy-free frameworks such as Life Identification Numbers (LINs) provide a stable, genome-based
representation of microbial diversity, where hierarchical relationships between lineages capture similarity at
multiple resolutions, but their use remains underexplored in predictive microbiome modelling.
Recently the Creevey lab have developed a LIN-like taxonomy-free classification approach for 16S microbiome
data for the entire Greengenes2 database. Here, we propose a computational framework that integrates this with
machine learning for predictive microbiome analysis. First, we model microbial community structure by learning
lineage-level co-occurrence patterns using graph-based approaches that leverage the hierarchical structure of
LINs to predict coexisting lineages given a target lineage. Second, we extend this framework to functional
inference by linking LINs to pangenome-derived gene content. Using graph-based propagation across related
lineages, we estimate gene presence–absence profiles and aggregate these into community-level functional
predictions mapped to metabolic pathways.
This framework provides a scalable and consistent approach for linking microbiome composition to function
across heterogeneous environments and supports a shift from descriptive to predictive microbiome analysis
using widely available 16S data.
A-G.C.23: CNETML2: A Multi-scale Model for Copy Number Alteration Evolution in Cancer
Track: General computational biology
-
Ruolin Wu, University of Surrey, United Kingdom
- Bingxin Lu, University of Surrey, United Kingdom
Presentation Overview: Show
Somatic copy number alterations (CNAs) are pervasive in cancer, particularly in tumours driven by chromosomal
instability. Their strong association with disease progression and adverse clinical outcomes makes
understanding how CNAs accumulate and change over time essential for early detection, prognosis, and
personalized therapy.
Although numerous studies have characterized CNA patterns and relative timing, few have estimated CNA rates in
absolute chronological time. We previously developed CNETML, a maximum-likelihood framework for reconstructing
CNA-based phylogenetic trees and estimating CNA evolutionary rates from shallow whole-genome sequencing data.
Applications of CNETML to cancer datasets have enabled quantitative reconstruction of tumour evolutionary
histories. However, existing models typically assume a single class of CNA events and constant evolutionary
dynamics, whereas cancer genomes often undergo multiple types of large-scale alterations that shape copy
number evolution.
Here we extend CNETML with a multi-scale model that explicitly incorporates three classes of genomic events:
segment-level duplications and deletions, chromosome-level gains and losses, and whole-genome doubling. These
processes are represented by three transition matrices describing independent Markov chains for each
alteration scale. The likelihood is computed dynamically by combining these matrices during phylogenetic
inference, allowing different CNA processes to contribute jointly to genome evolution. To improve
computational efficiency, the model prunes redundant haplotype copy-number states that are incompatible with
the constraining maximum copy number boundaries, substantially reducing the size of the state space.
Implemented within the CNETML framework, the new model enables more efficient and biologically realistic
inference of tumour phylogenies and CNA evolutionary rates from cancer sequencing data.
A-G.C.24: Metadata Accelerator: Improving scientific data descriptions with Natural Language Processing
methods (NLP) and Feedback
Track: General computational biology
-
Maria Juliana Rodriguez Cubillos, Centre for Engineering Biology, School of Biological Sciences and
School of Informatics, University of Edinburgh, UK, United Kingdom
-
Tomasz Zieliński, Centre for Engineering Biology, School of Biological Sciences, University of Edinburgh,
Edinburgh EH9 3BF, UK, United Kingdom
-
Jason R. Swedlow, Divisions of Molecular Cell and Developmental Biology, and Computational Biology,
University of Dundee, Dundee, UK, United Kingdom
-
T. Ian Simpson, School of Informatics, University of Edinburgh, 10 Crichton Street, Edinburgh EH8 9AB, UK,
United Kingdom
-
Andrew J. Millar, Centre for Engineering Biology, School of Biological Sciences, University of Edinburgh,
Edinburgh EH9 3BF, UK, United Kingdom
Presentation Overview: Show
Ensuring the availability and accessibility of data is fundamental to advancing knowledge. This idea has been
codified as the FAIR principles (Findable, Accessible, Interoperable, and Reusable) in scientific data
management. Accurate documentation of studies, commonly known as metadata, is indispensable for achieving this
goal. Regrettably, scientific records often fall short, providing inadequate, repetitive, or incomplete
descriptions that hinder the seamless flow of knowledge. This project addresses the metadata challenge by
identifying common failure points in descriptions and generating better-structured metadata using NLP and
LLMs, with a focus on sustainability.
The Metadata Accelerator is a pipeline that aims to predict metadata using information from free-text
descriptions. Based on initial user inputs, the pipeline suggests improvements that fit template-based schemas
and provides feedback using quality metrics. Users can then refine the suggested changes before export,
allowing us to engage users and explore whether richer interaction improves metadata quality.
The pilot version of the pipeline predicts “species†and “study type†for each entry from user-provided
descriptions using a combination of NLP methods and LLMs, including Llama 3.2 and BioBERT. The Image Data
Resource (IDR) and BioDare2, a repository for circadian and biological data, were selected to develop a
pipeline that suggests metadata terms from existing free-text descriptions. For IDR, the system identified 97
species compared with 110 in a manual review and matched 130 of 132 study-type entries.
Ongoing work will evaluate performance on BioDare2 and other repositories, improve efficiency, and expand to
additional categories.
A-G.C.25: Discovering Novel Ciliary Genes Through Machine Learning
Track: General computational biology
-
Emilia Torriglia, Institut Pasteur de Montevideo, Uruguay
- Florencia Irigoin, Institut Pasteur de Montevideo, Uruguay
- Laura Romanelli-Cedrez, Institut Pasteur de Montevideo, Uruguay
- Flavio Pazos Obregón, Institut Pasteur de Montevideo, Uruguay
Presentation Overview: Show
The primary cilium is an evolutionarily conserved organelle present in many eukaryotic cells, where it plays a
central role in cellular signaling and environmental sensing. Despite its importance, the full set of genes
required for its structure and function remains incomplete. Caenorhabditis elegans provides a powerful system
to study ciliary biology due to its well-characterized nervous system and ease of genetic and experimental
manipulation, enabling the identification of conserved mechanisms across species. We hypothesized that genes
associated with ciliary function share expression-based features and other biological characteristics,
independent of DNA sequence similarity, that can be used for their identification.
To investigate this, transcriptomic datasets were constructed from publicly available single-cell RNA
sequencing data (Packer et al. 2019; Taylor et al. 2021). These datasets include neuronal cell types across
developmental stages, comprising both ciliated and non-ciliated populations, and were integrated prior to
analysis. Genes were filtered based on expression across different cell types and variability, and annotated
ciliary genes were split into enrichment analysis and evaluation sets. Expression values were standardized,
and agglomerative hierarchical clustering was applied across multiple distance thresholds. Enrichment analysis
identified clusters overrepresented in ciliary genes, from which candidate genes were defined. Candidates were
compared across conditions and stages, and the method was evaluated by quantifying the enrichment of
evaluation-set genes within enriched clusters.
In future work, additional biological features independent of DNA sequence, such as genomic localization of
functional groups and structural motifs, will be incorporated, alongside supervised machine learning
approaches to improve candidate prioritization.
A-G.C.26: COMPAS: COntrastive Multimodal Polypharmacology-Aware multi-target drug-target interaction
Screener
Track: General computational biology
-
Gamze Deprem, Hacettepe University, Turkey
- Tunca Doğan, Hacettepe University, Turkey
Presentation Overview: Show
Complex diseases such as cancer, diabetes, and neurodegenerative disorders are driven by multiple
interconnected biological mechanisms, limiting the effectiveness of traditional one-drug-one-target
strategies. Controlled polypharmacology, in which a single molecule modulates multiple disease-relevant
targets in a selective and coordinated manner, has therefore emerged as a promising therapeutic paradigm.
However, current computational drug repurposing and discovery approaches remain largely focused on
single-target interactions, use biomolecular modalities in limited ways, and do not directly model
multi-target interaction patterns, thereby restricting generalizability. In this study, COMPAS (Contrastive
Multimodal Polypharmacology-Aware Screener) is proposed as a new machine learning model for directly
predicting interactions between small molecules and multiple targets. COMPAS utilizes molecular embeddings
obtained from pretrained multimodal models using SMILES and graph representations, while protein embeddings
are derived from pretrained multimodal protein language models using sequence, structure, text modalities.
These embeddings are projected into a shared latent space through trainable projection layers. A
molecule-centered, multi-positive supervised contrastive learning strategy is then applied to organize this
space, so that each molecule is positioned closer to its interacting proteins and farther from irrelevant
targets. In this way, coordinated multi-target interaction profiles are learned within a unified multimodal
setting. The proposed model is developed using data from public resources: ChEMBL, PubChem, ZINC, and UniProt.
Alzheimer's disease is adopted as a use case focusing on AChE and BACE-1, and predicted candidates are
intended for molecular docking analysis. COMPAS is expected to improve the identification of biologically
relevant multi-target drug candidates and support scalable drug discovery for complex diseases.
A-G.C.27: Improving comparative genomics with orthology uncertainty metrics
Track: General computational biology
-
Veronica Sondervan, VIB-UGent Center for Plant Systems Biology, Belgium
- Yves Van de Peer, VIB-UGent Center for Plant Systems Biology, Belgium
- Zhen Li, VIB-UGent Center for Plant Systems Biology, Belgium
Presentation Overview: Show
Accurate identification of orthologs across species is critical for avoiding spurious gene-loss
identification, inflated family expansions, and unreliable selection inferences in comparative genomics.
Studies have shown that current orthology methods can vary widely in their orthology calls, but there remains
no quantified method for determining the uncertainty of an ortholog group. Here, we develop an orthology
uncertainty scoring system by integrating metrics including sequence similarity, synteny, and phylogenetic
support. We also evaluate whether protein language model embeddings can provide complementary uncertainty
signal. We benchmark our system with Dendrobium orchids, a genus known for morphological differences
associated with gene loss, including transitions between lithophytic v.s. epiphytic lifestyles and
self-compatibility v.s. incompatibility. Applying established orthology methods, we show that our uncertainty
scoring scheme aids in identifying and confirming gene losses and spotting poorly supported orthology calls.
Our work represents an important step forward in increasing the precision and certainty of evolutionary
investigations into species history and adaptation.
A-G.C.28: FLKit: A Community-Driven Onboarding Toolkit for Federated Analytics and Learning in Health and
Life Sciences
Track: General computational biology
-
Ashkan Pirmani, KU Leuven - Uhasselt, Belgium
- Ilse Vermeulen, Uhasselt, Belgium
- Goran Vinterhalter, KU Leuven, Belgium
- Lotte Geys, UHasselt, Belgium
- Axel Faes, UHasselt, Belgium
- Muhammad Quamber Ali, KU Leuven, Belgium
- Nshkala Sattanathan, UAntwerpen, Belgium
- Geert Vandeweyer, UAntwerpen, Belgium
- Yves Moreau, KU Leuven, Belgium
- Liesbet Peeters, UHasselt, Belgium
Presentation Overview: Show
Federated learning (FL) allows institutions to collaboratively train models without sharing raw sensitive
data, making it a natural fit for health and life sciences research bound by strict data protection
regulations. Despite growing technical maturity, most researchers, clinicians, data stewards, and legal
advisors entering this space encounter a fragmented landscape of frameworks, governance requirements, and
terminology with no structured starting point tailored to their role.
FLKit addresses this gap as an open, community-maintained onboarding toolkit that guides multidisciplinary
teams through the full federated learning lifecycle. Inspired by the ELIXIR Research Data Management Kit
(RDMkit) and developed using a design science approach, FLKit is organised around four lifecycle stages
(Governance, Infrastructure, Wrangling, and Analysis), 11 role-specific entry points spanning clinical
researchers, data stewards, legal advisors, and research software engineers, a curated glossary, a
FAIR-aligned FL Story template for documenting real-world use cases, and a directory of tools, frameworks, and
communities.
Since its demo launch in December 2024, FLKit has grown to 39 pages across eight content sections. Seven FL
Stories document completed and ongoing projects across multiple sclerosis disability prediction, inflammatory
bowel disease, genomics, and ECoG-based brain-computer interfaces.
FLKit is openly available at https://uhasselt-biomedicaldatasciences.github.io/federated-learning-toolkit/ and
is actively inviting community contributions from across the life sciences.
A-G.C.29: HOROSCOPE: Decoding human centromere architecture from short reads using k-mer signatures
Track: General computational biology
-
Carsten Hain, EMBL, Germany
- Tobias Rausch, EMBL, Germany
- Jan Korbel, EMBL, Germany
Presentation Overview: Show
By directing kinetochore formation and chromosome segregation, centromeres safeguard genome integrity. Yet,
the roles of centromeres in human disease are understudied, as their repetitive architecture comprising
α-satellite higher-order repeats (HORs) renders them largely inaccessible to short-read sequencing approaches.
Here we develop HOROSCOPE (Higher-Order Repeat organization and Size of Centromeres using Oligonucleotide
Profiles for Estimation), a computational framework for k-mer-based inference of centromere structure and
length from short-read data. Based on a reference atlas of 11,836 human centromeres extracted from completely
assembled (telomere-to-telomere) haplotypes, we systematically interrogate chromosome-specific centromere
architectures, deriving architecture-specific k-mer signatures as well as centromeric length-informative
k-mers from common to rare centromere architectures. Leveraging these diagnostic k-mers, HOROSCOPE achieves a
precision of 99.3% and a recall of 99.5% in classifying chromosome-specific centromere architectures from
short-read de Bruijn graphs. We perform a population-scale analysis of global centromeric architectures in
4,029 human samples with ancestry from 80 human populations sequenced with short reads, uncovering continental
haplotype structure for centromeric regions and highlighting African-enriched rare centromeric architectures.
Furthermore, by analyzing 1,359 cancer genomes, we link graded HOR truncation events to arm-level copy-number
alterations, and uncover a general dependency of chromosomal rearrangement locations on the position of the
centromere dip region (CDR) defining the kinetochore attachment site. HOROSCOPE enables large-scale centromere
genomics directly from short-read sample cohorts.