View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
C-G.01: Adaptive Dynamics of HIV-1 Populations Over a Six-Year-Long Experimental Evolution: A Genomic
Perspective
Track: Genomics, epigenomics, and genome editing
-
Ali Movasati, Department of Infectious Diseases and Hospital Epidemiology, Universitätsspital Zürich,
Switzerland, Switzerland
-
Christine Leemann, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital
Zürich, Switzerland, Switzerland
-
Kathrin Neumann, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital Zürich,
Switzerland, Switzerland
-
Rongfeng Chen, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital Zürich,
Switzerland, Switzerland
-
Lygeri Sakellaridi, University of Würzburg, Institute for Virology and Immunobiology, Germany, Germany
-
Karin Metzner, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital Zürich,
Switzerland, Switzerland
-
Roland Regoes, Institute of Integrative Biology, ETH Zürich, Switzerland, Switzerland
Presentation Overview: Show
Numerous experimental evolution studies have suggested that adaptation rate of microbial populations evolving
in stable environments decline over time. To investigate the characteristics of adaptation deceleration in a
fast-evolving virus, we propagated HIV-1 in two human T-cell lines (MT-2 and MT-4) for approximately 4.8 years
and tracked its genome evolution through NGS. The sequencing data can be explored via LTEEviz, an interactive
web application. Time-resolved sequencing data indicated that despite constant fixation rate of 0.085 (MT-2)
and 0.042 (MT-4) mutations per generation, the fixation kinetics of adaptive mutations changed considerably
over time. The rate of fixation of adaptive parallel mutations decreased by 44% per 300 generations, while
their conferred fitness gain diminished by 27% (MT-2) and 18% (MT-4) per every added adaptive mutation in
their genetic background. Furthermore, we identified unique yet consistent patterns of sequence evolution
among different regions of the HIV-1 genome. In particular, nef and vpr accessory genes demonstrated patterns
of random evolution, expected in the absence of selection. Condordantly, the evolving populations of HIV-1
acquired and fixed multiple loss-of-function mutations in nef and vpr. Additionally, nef gene repeatedly
underwent large deletions, leading to the removal of approximately 400 bp (equivalent to 4.3% of the HIV-1
genome). These large nef deletions increased in frequency faster than expected under neutrality and in a
length-dependent manner. Together, our results confirm that HIV-1 genomic evolution is characterized by a
swift and substantial deceleration of adaptation, while highlighting progressive genome shrinkage as one of
the underlying adaptive mechanisms.
C-G.02: Decoding the noncoding: a computational and CRISPR-enabled framework from long noncoding RNA
discovery to therapeutic targeting
Track: Genomics, epigenomics, and genome editing
-
Maina Bitar, QIMRB, Australia
- Stacey Edwards, QIMRB, Australia
- Juliet French, QIMRB, Australia
- Haran Sivakumaran, QIMRB, Australia
Presentation Overview: Show
Long noncoding RNAs (lncRNAs) represent a vast and largely unexplored layer of cancer biology, harbouring the
majority of somatic mutations in cancer. Here, we present an integrated computational and functional genomics
framework to systematically discover, characterise and prioritise lncRNAs for therapeutic targeting in breast
and ovarian cancers.
Over the past four years, we developed a framework powered by a range of computational tools. We started
developing in silico metatranscriptome assembly strategies to uncover thousands of previously unannotated
lncRNAs from relevant cell and tissue samples, initially using short-read data (ShROOM) and now extended to
hybrid long- and short-read integration (HyDRA). These tools enabled the redefinition of normal breast
epithelial cell populations and the construction of pseudo-longitudinal models of ovarian cancer progression
and chemoresistance, comprehensively mapping lncRNA involvement in disease.
To prioritise candidates from this expanded gene set, we computationally integrate lncRNA discovery with
genetic association signals, including genome-wide association studies. Functional interrogation is achieved
through a suite of CRISPR-based technologies. We pioneered RNA-targeting CRISPR–Cas13 screens, establishing
the first platform for transcript perturbation at scale, and are extending this approach to single-cell
readouts (CROP-seq). We are also developing high-throughput CRISPR prime editing to assess the functional
impact of thousands of lncRNA mutations.
This framework enables the identification of clinically actionable lncRNAs, including functional drivers like
BRRIAR, a strong candidate for oestrogen receptor–positive breast cancer therapy, the highlight of this
talk. Our work establishes a scalable path from discovery to therapeutic targeting, positioning lncRNAs as
central figures in precision oncology.
C-G.03: Predicting Genomic Determinants of Chromatin Compaction Using Machine Learning
Track: Genomics, epigenomics, and genome editing
- Ryan Burke, CEZAMAT, Warsaw University of Technology, Poland
- Ewelina Holm Bidstrup, CEZAMAT, Warsaw University of Technology, Poland
- Shuting Liu, University of Illinois at Urbana-Champaign, United States
- Ilaria Lupi, CEZAMAT, Warsaw University of Technology, Poland
- Monika Staniszewska, CEZAMAT, Warsaw University of Technology, Poland
- Andrew Belmont, University of Illinois at Urbana-Champaign, United States
-
Teresa Szczepinska, CEZAMAT, Warsaw University of Technology, Poland
Presentation Overview: Show
The hierarchical 3D organization of chromatin in eukaryotic cells plays a crucial role in gene regulation and
adapts during development, in response to stimuli, and in disease. While large-scale chromatin domains at
scales larger than TADs (up to hundreds of nanometers) have been observed using microscopy, recent advances
such as PCC-seq enable high-throughput measurement of chromatin compaction at kilobase resolution, revealing
local variability, including those around transcription start sites (TSS).
Here, we applied machine learning approaches, including Logistic Regression, Random Forest, and
Histogram-based Gradient Boosting, to classify genomic regions according to chromatin compaction levels.
Models were trained for regions of TSS (±1 kb, ±5 kb), genome-wide (1 kb, 100 kb), considering 2, 3, and 5
compaction classes. Predictions were based on 519 genomic features, including proteins and histone
modifications ChIP-seq data, chromatin accessibility (ATAC-seq, DNA-seq), and transcription (BRU-seq,
RNA-seq). As a baseline comparison, an analogous classification was performed for ATAC-seq signal in TSS ±1 kb
regions.
Using permutation feature importance and single-feature ROC AUC scores, we ranked genomic marks associated
with chromatin compaction. In TSS ±1 kb regions, EP400, ELF4, POLR2H, H2AFZ, and ATAC-seq signals were
associated with decompaction, while H3K27me3 and STAG1 correlated with compaction. Moreover, in TSS ±5 kb
regions, H3K27ac and PCBP1 were linked to decompaction, whereas MCM3 correlated with compaction. At whole
genome scales, in 100 kb bins THRAP3 and MBD1 were associated with decompaction, while ZBTB33 and H3K9me3
correlated with compaction. At 1 kb resolution, SETDB1, ZC3H4, and ZNF263 were associated with increased
compaction.
C-G.04: GPxLMM: Gaussian Process-Augmented Linear Mixed Models for Genotype-by-Environment Interaction
Analysis
Track: Genomics, epigenomics, and genome editing
-
Bibiana Mailyn Horn, Hasso Plattner Institute, University of Potsdam, Germany
-
Zoran Nikoloski, University of Potsdam; Max Planck Institute of Molecular Plant Physiology, Germany
-
Christoph Lippert, Hasso Plattner Institute, University of Potsdam; Icahn School of Medicine at Mount
Sinai, Germany
Presentation Overview: Show
Multivariate genome-wide association studies are essential for understanding phenotypic plasticity and
identifying genetic variants driven by genotype-by-environment (GxE) interactions. Existing methods, however,
rely on analytically derived gradients and fixed covariance assumptions, which can lead to model
misspecification and reduce power when GxE effects vary smoothly across continuous environmental gradients. To
address this, we introduce GPxLMM, a framework integrating Gaussian process covariance learning with linear
mixed model inference for fixed effects. By leveraging automatic differentiation in GPyTorch, new covariance
structures can be incorporated without deriving analytical gradients. Alongside the standard index kernel for
discrete phenotypes, we introduce a reaction norm kernel to model the genetic covariance of function-valued
phenotypes. By capturing continuous phenotypic variation across changing environments, this approach enables
phenotype prediction and genotype ranking in unobserved environments. In simulations using real Arabidopsis
thaliana genome data, GPxLMM accurately detects rescaling and heterogeneous GxE effects and produces
well-calibrated variance decompositions. For discrete phenotypes, it matches LIMIX in statistical power while
retaining the flexibility to incorporate alternative covariances. In leave-one-environment-out
cross-validation across 21 DROPS environments, the reaction norm kernel achieved a mean predictive correlation
of ~0.70. Furthermore, when selecting the top 20 genotypes, the model demonstrated on average 75% selection
efficiency. To conclude, GPxLMM provides an extensible framework for integrating flexible covariance learning
into linear mixed models, advancing genomic prediction and quantitative genetics beyond fixed structural
assumptions.
C-G.05: nf-core/raredisease: Modernizing Clinical Genomics by Transitioning to Community-Driven
Workflows
Track: Genomics, epigenomics, and genome editing
-
Ramprasad Neethiraj, KTH Royal Institute of Technology, Stockholm, Sweden, Sweden
-
Anders Jemt, Genomics Medicine Centre Karolinska, Karolinska University Hospital, Stockholm, Sweden,
Sweden
-
Henrik Stranneheim, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm,
Sweden, Sweden
-
Peter Pruisscher, Genomics Medicine Centre Karolinska, Karolinska University Hospital, Stockholm, Sweden,
Sweden
-
Daniel Nilsson, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm,
Sweden, Sweden
-
Chiara Rasi, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm,
Sweden, Sweden
-
Valtteri Wirta, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm,
Sweden, Sweden
Presentation Overview: Show
The clinical diagnosis of rare genetic disorders requires the seamless integration of single-nucleotide
variants (SNVs), structural variants (SVs), and repeat expansions. For over a decade, the Stockholm healthcare
region relied on the in-house Mutation Identification Pipeline (MIP) to process 10,000+ clinical samples. To
meet increasing demands for scalability, portability, and long-term sustainability, we developed and
transitioned to nf-core/raredisease as its official clinical successor.
nf-core/raredisease leverages the collective logic and diagnostic experience gained from MIP while adopting
the modern Nextflow framework and nf-core best practices. The pipeline orchestrates state-of-the-art tools for
whole-genome and exome sequencing into a containerized, FAIR-compliant workflow. Key features include
integrated mitochondrial analysis, rank-based variant prioritization, and comprehensive detection of complex
variants.
Having recently replaced MIP in routine clinical production at Karolinska University Hospital,
nf-core/raredisease demonstrates how legacy clinical expertise can be successfully ported to a
community-driven model. This transition has significantly improved maintenance efficiency and reproducibility
without sacrificing the diagnostic rigor established over years of service. This poster details the pipeline's
architecture, the technical challenges of the migration process, and the benefits of moving from a bespoke
legacy system to a standardized open-source framework. By sharing this transition, we provide a blueprint for
other clinical centers looking to modernize their genomic infrastructure through collaborative, peer-reviewed
software development.
C-G.06: Role of histone H2A lysine119 ubiquitination in gene regulation in BAP1-positive and BAP1-negative
uveal melanoma eye cancer.
Track: Genomics, epigenomics, and genome editing
-
Anneke Brummer, University of Lausanne, Switzerland
- Adeline Berger, Hopital ophtalmique Jules-Gonin, Switzerland
- Nicolas Guex, University of Lausanne, Switzerland
- Alexandre Moulin, Hopital ophtalmique Jules-Gonin, Switzerland
Presentation Overview: Show
Uveal melanoma (UM), though a rare cancer affecting only about 5 per million adults, is the most common
primary eye tumor. BAP1 (BRCA1-associated protein 1) inactivation is a critical mutation associated with
metastasis and poor prognosis. BAP1 deubiquitinates histone H2A lysine 119 (H2AK119ub, known for gene
silencing) and thereby activates chromatin and gene expression. Here, we analyse data from 11 BAP1-positive
and 12 BAP1-negative UM patient tumor samples, integrating genetic, epigenetic, and transcriptomic
information, to better understand the role of K119ub in gene expression regulation in these tumor subgroups.
We first confirmed that transcriptomes of BAP1-positive and BAP1-negative subgroups were clearly separated.
This was mostly due to differential gene regulation, and not to gene copy number variations. Overall, more
genes had higher expression levels in BAP1-positive tumors, in agreement with an activating function of BAP1.
However, notably many genes were also more expressed in BAP1-negative samples. We next sought to relate gene
expression with chromatin state differences considering 5 histone modifications (H2AK119 ubiquitination, H3K27
tri-methylation, H3K27 acetylation, H3K4 mono-methylation and H3K4 tri-methylation). Inferred chromatin states
largely agreed with measured chromatin accessibility by ATAC-Seq. Surprisingly, around transcription start
sites, repressive chromatin, characterized by K27me3 and K119ub, was more abundant in BAP1-positive than
BAP1-negative samples. In contrast, in BAP1-negative tumors, K119ub more often co-localized with K4me1 in this
region, including in active chromatin states together with K27ac and K4me3. Active chromatin states without
K119ub were more abundant in BAP1-positive samples, as expected. Overall, chromatin states agreed with gene
expression in BAP1 subgroups.
C-G.07: Imputation of Bacterial Genomes based on cgMLST Profiles
Track: Genomics, epigenomics, and genome editing
-
Christina Kirschbaum, Robert Koch Institute, Germany
- Torsten Houwaart, Robert Koch Institute, Germany
- Vladimir Bajić, Robert Koch Institute, Germany
- Simon H. Tausch, Robert Koch Institute, Germany
- Hugues Richard, Robert Koch Institute, Germany
Presentation Overview: Show
Accurate pathogen identification is a crucial step for effective treatment, outbreak management and risk
characterization. Nowadays a common genotyping methodology for bacteria is core genome multi locus sequence
typing (cgMLST), which characterizes samples by a set of genes commonly found in the species of interest. In
metagenomics, new methods with shallow sequencing are on the rise and a tool to account for missing genes in
genotyping methods is needed. We propose a method that performs genotype imputation similar to approaches for
human genomes.
We first assessed linkage between genes, and showed that by grouping highly correlated loci, we could reach
compression levels up to 29% for Listeria monocytogenes and over 50% with Mycobacterium tuberculosis given
canonical typing schemes. We then implemented a Markov model (MM) on the core genome that estimates allele
succession on adjacent genes. On Listeria monocytogenes, MM always outperformed the baseline imputation method
using maximum frequency allocation when tested on different thresholds of masked alleles. Masking 15% of the
alleles, MM correctly reconstructs 67% of the alleles (baseline: 26%) when applied to the reference data, and
59% (baseline: 24%) on a larger pathogen.watch and BigsDB dataset. Preliminary results on Mycobacterium
tuberculosis are more mitigated (64%/69% MM, 74%/84% baseline). The performance is likely compounded by a
smaller reference set and a low allelic variability leading to overfitting.
We are aiming towards an improved model considering haplotype blocs which would be tested on an extended set
of reconstruction scenarios.
C-G.08: A systematic investigation into the robustness of mutational signature fitting to unknown
signatures
Track: Genomics, epigenomics, and genome editing
-
Maria Katsantoni, Department for BioMedical Research, Inselspital, Bern University Hospital and
University of Bern, Bern, Switzerland, Switzerland
-
Qixuan Wang, Department for BioMedical Research, Inselspital, Bern University Hospital and University of
Bern, Bern, Switzerland, Switzerland
-
Matúš Medo, Department for BioMedical Research, Inselspital, Bern University Hospital and University of
Bern, Bern, Switzerland, Switzerland
Presentation Overview: Show
Accurate deconvolution of mutational signatures leads to a better understanding of tumor etiology, yet the
mathematical stability of signature attribution remains a significant challenge in the field. As highlighted
in previous benchmark studies, existing algorithms used to fit mutational signatures frequently struggle with
two intertwined phenomena: the misallocation of stochastic noise as biological signal (overfitting) and the
"displacement" of mutations from signatures absent in reference catalogs onto known ones. The latter
represents a critical limitation where a model's incomplete basis (underfitting) directly causes the spurious
inflation of existing signatures (overfitting), leading to misleading clinical interpretations.
Building upon previous software for benchmarking signature analysis tools, we have tested a broad range of
synthetic samples to characterize the exact conditions, such as low mutational burden and high cosine
similarity between signatures, under which existing tools deviate from biological ground truth. We developed a
new algorithm based on regularized deconvolution and residual analysis to isolate 'unknown' signals even when
the underlying signatures are undefined, thereby preventing the erroneous mapping of mutations to the
reference catalog.
To ensure the portability and reproducibility of these findings, we have implemented the entire analytical
pipeline as a modular Snakemake workflow. Preliminary results suggest that regularized deconvolution can
achieve more robust and transparent signature assignments. Our work provides a diagnostic perspective on the
limitations of current fitting paradigms and offers a scalable path toward more reliable genomic reporting.
C-G.09: Exploration of twin genes with DupyliCate and biological implications
Track: Genomics, epigenomics, and genome editing
-
Shakunthala Natarajan, University of Bonn, Germany
- Claudia Sterling, University of Bonn, Germany
- Boas Pucker, University of Bonn, Germany
Presentation Overview: Show
Paralogs, copies of a gene, form an important basis for novelty during evolution. Analysis of such gene
duplications is important to understand the emergence of novel evolutionary traits. DupyliCate is a Python
tool that has been developed for the identification, classification, and characterization of gene copies. With
the ability to process multiple datasets concurrently, flexible features, and parameters to set
species-specific thresholds, DupyliCate offers a high-throughput method for gene duplicate array
identification. It also facilitates downstream gene expression divergence analysis of the identified
duplicates, enabling their fate prediction. DupyliCate was applied on the flavonoid synthase (FLS) gene family
in Brassicales and subgroup 7 myeloblastosis (MYB) transcription factors (SG7 MYB) across a diverse range of
plant species to understand the duplication dynamics of these important players in flavonoid biosynthesis.
This helped uncover a potential radiation of FLS genes in the Brassicaceae and a deep duplication of the SG7
MYB lineage in some dicots. Further, DupyliCate was also used to identify gene duplications in other key genes
of the flavonoid biosynthesis. A downstream expression analysis of these mined duplicates using the tool,
combined with a systematic analysis of their promoter sequences helped ascertain the genetic triggers
explaining the observed divergence and redundancy.
DupyliCate is available at: https://github.com/ShakNat/DupyliCate
C-G.10: Cracking the animal venom code using phylogenetic big data
Track: Genomics, epigenomics, and genome editing
-
Athina Gavriilidou, University of Lausanne, Switzerland
- Giulia Zancolli, University of Lausanne, Switzerland
- Christophe Dessimoz, University of Lausanne, Switzerland
- Natasha Glover, Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Venom systems are among nature's most sophisticated biological innovations, yet their evolution and functional
integration remain incompletely understood. We conducted a large-scale comparative analysis of venom-related
genes across animals using phylogenetic and genomic approaches.
We compiled a dataset of venom-associated protein families from public databases, spanning diverse taxa to
capture broad evolutionary patterns. To identify coevolving proteins, we applied phylogenetic profiling,
representing each gene as a binary vector encoding its presence or absence across species. By comparing these
profiles, we detected protein families with similar evolutionary trajectories. Using the Metazoa dataset from
the OMA database, which uses hierarchical orthologous groups (HOGs), we employed the HogProf algorithm to
identify HOGs with high similarity to venom-related profiles. This enabled systematic detection of candidate
genes potentially involved in venom function or self-resistance.
Preliminary results show that proteins within the same venom cocktail often share highly similar phylogenetic
profiles, supporting coordinated evolution. We also observed a correlation between expression patterns and
phylogenetic similarity, suggesting functional links. Additionally, we identified protein families with
venom-like evolutionary signatures not currently associated with venom, highlighting candidates for resistance
mechanisms.
A central aim is to understand how venomous species avoid self-intoxication through biochemical adaptations.
Investigating these mechanisms may clarify the evolutionary interplay between toxin production and resistance,
and provide broader insight into the evolution of toxin resistance across species.
C-G.11: MosCoverY: A method to estimate mosaic loss of Y chromosome from sequencing coverage data
Track: Genomics, epigenomics, and genome editing
-
Valeriia Timonina, School of Life Sciences, École Polytechnique Fédérale de Lausanne,
Switzerland
- Astrid Marchal, Imagine Institute, Université Paris Cité, France
-
Laurent Abel, Hôpital Necker Enfants Malades; Imagine Institute, Université Paris Cité, France
-
Aurélie Cobat, Hôpital Necker Enfants Malades; Imagine Institute, Université Paris Cité, France
-
Jacques Fellay, School of Life Sciences, École Polytechnique Fédérale de Lausanne, Switzerland
Presentation Overview: Show
Mosaic loss of the Y chromosome (mLOY) is the most common somatic genomic event in men, accumulating with age
and associated with diverse health outcomes, including all-cause mortality, cardiovascular disease, and
cancer. Despite its clinical and biological relevance, detection of mLOY relies on DNA genotyping arrays,
limiting its assessment in large cohorts where only sequencing data are available. Here, we present MosCoverY,
a computational method for estimating mLOY directly from exome or whole-genome sequencing data.
MosCoverY addresses the challenges of the Y chromosome's structure by restricting analysis to single-copy
genes and normalizing their sequencing coverage against autosomal exons matched by GC content and length. This
design yields a robust individual-level estimate of normalized chromosome Y coverage, from which a binary mLOY
and a continuous cell fraction harbouring mLOY can be derived.
We validated MosCoverY in 212,062 male participants from the UK Biobank, benchmarking it against two
established methods based on genotyping arrays and whole-genome sequencing. MosCoverY identified mLOY in 5.6%
of men, showed a strong correlation with other methods, and comparable performance in replicated associations
with age, tobacco smoking, all-cause mortality, and germline genetic loci, yielding the strongest effect
estimates in several cases. We further demonstrated method robustness at reduced sequencing depth and in
single-sample analyses without population-level data. Finally, applying MosCoverY to The Cancer Genome Atlas
confirmed its utility for detecting variable mLOY in tumors.
MosCoverY expands the toolkit for somatic genomic research, enabling mLOY detection in the growing body of
sequencing-based population and clinical cohorts.
C-G.12: Computational Detection of Methylated Bases in Bacteria: PacBio vs. ONT
Track: Genomics, epigenomics, and genome editing
-
Mohammad Umair, Brno University of Technology, Czechia
-
Marketa Jakubickova, Department of Biomedical Engineering, Brno University of Technology, Czechia
- Iva Buchtikova, Brno University of Technology, Czechia
- Matej Bezdicek, Masaryk University, Czechia
- Stanislav Obruca, Brno University of Technology, Czechia
- Helena Vitkova, Brno University of Technology, Czechia
-
Karel Sedlar, Department of Biomedical Engineering, Brno University of Technology, Czechia
Presentation Overview: Show
DNA methylation is a widespread epigenetic modification in bacterial genomes, playing key roles in processes
like restriction-modification systems and gene expression regulation. The advent of long-read sequencing has
enabled detection of these modifications directly from native DNA; However, cross-platform comparability of
the DNA methylation profiles remains insufficiently understood. In this study, we performed a cross-platform
comparison of bacterial DNA methylation detection using Pacific Biosciences (PacBio) and Oxford Nanopore
Technologies (ONT) sequencing, leveraging data from phylogenetically diverse bacterial strains.
using the Standardised Analysis Framework, we evaluated all three major prokaryotic DNA methylation types
(6mA,4mC, and 5mC) across multiple analytical levels. These included site-level overlap, methylation fraction
estimate, motif discovery, genomic distribution, and the functional categorisation of genes associated with
methylated regions. Beyond assessing overlap in detected modified sites, this approach enabled us to examine
whether both platforms capture consistent methylation patterns at broader genomic and functional scales. While
PacBio and ONT revealed partially overlapping methylomes, their outputs were clearly non-identical.
Concordance was highest for adenine methylation, whereas cytosine methylation showed greater platform
dependency, particularly at the site level and in estimated modification fraction. Nevertheless, both
technologies demonstrated similar trends in motif detection and preferred genomic and functional contexts.
Overall, our findings indicate that PacBio and ONT are not interchangeable, as each captures overlapping yet
distinct features of the bacterial methylation landscape. This comparison provides a practical framework for
selecting appropriate long-read sequencing strategies and highlights the value of integrative approaches for
achieving a more comprehensive view of the bacterial genome.
C-G.13: A coverage-focused workflow for high-confidence structural variant calling in cattle without
validated truth sets
Track: Genomics, epigenomics, and genome editing
-
Laura Dekker, Animal Genomics ETH Zürich, Switzerland
- Alexander Leonard, Animal Genomics ETH Zürich, Switzerland
- Hubert Pausch, Animal Genomics ETH Zürich, Switzerland
Presentation Overview: Show
Structural variants are variants in the genome that span 50 bases or more. Advances in long read sequencing
make it increasingly possible to investigate structural variant diversity in large cohorts. In many non-model
species, differentiating between true and erroneous structural variant calls is difficult due to the lack of a
truth set. This project aims at establishing a workflow to obtain a high-confidence set of structural variant
genotypes in cattle which does not rely on a truth set of validated variants. We aligned PacBio HiFi reads
from 57 cattle that had at least 15x coverage to the Bos taurus ARS-UCD2.0 reference genome to identify
structural variants using Sniffles2 and Sawfish. A set of increased confidence structural variants was
constructed by standard filtering of variants with low quality scores, missing genotypes and consensus length
below 45 bp. Additionally, genomic areas with significantly outlying coverage values were included in a
‘blacklist' of coordinates from which to exclude structural variant calls. Results show that the structural
variants highlighted by the latter method overlap largely with structural variants filtered out by the set of
initial filters but exclude an additional 200-500 structural variants depending on the caller. These findings
suggest that incorporating coverage-based exclusion criteria can improve the reliability of structural variant
datasets in the absence of validated truth sets.
C-G.14: Neural posterior estimation for population genetics
Track: Genomics, epigenomics, and genome editing
- Jiseon Min, Institute of Ecology and Evolution, University of Oregon, United States
-
Yuxin Ning, Quantitative Biology Center (QBiC), university of Tübingen, Germany
-
Nathaniel Pope, Institute of Ecology and Evolution, University of Oregon, United States
- Franz Baumdicker, Justus Liebig University Giessen, Germany
- Andrew Kern, Institute of Ecology and Evolution, University of Oregon, United States
Presentation Overview: Show
Simulation-based inference methods are increasingly being used in population genetics due to their flexibility
and ability to be applied in settings where likelihood-based methods are intractable. One of the best known
such method is Approximate Bayesian Computation (ABC). However, its popularity is offset by its shortcomings
which include computational expense and an unfortunate inability to efficiently fit models to high-dimensional
summaries of the data. An alternative approach that solves these issues is supervised machine learning (ML),
but ML methods generally do not yield Bayesian uncertainty estimates of the quantities they predict.
Here, we apply a recently introduced method, neural posterior estimation (NPE), that combines the best facets
of ABC and supervised ML by training a neural network to estimate the posterior distribution of a population
genetics model. We first compare neural posterior estimation with other inference methods for a variety of
population genetic tasks and show that neural posterior estimators yield posterior distributions with high
accuracy and efficiency. We compare learned posterior distributions given raw genotypes and various summary
statistics as input data. Additionally, we apply neural posterior estimation for demographic inference for
simple and more complex models to highlight its application, including an analysis of demographic history in
Drosophila melanogaster. Finally, we provide a user-friendly Snakemake workflow that enables others to perform
neural posterior estimation on their own genetic data.
C-G.15: Extending the megSAP pipeline for medical genetics to long-read sequencing
Track: Genomics, epigenomics, and genome editing
-
Marc Sturm, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen,
Germany, Germany
-
Tobias Haack, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen,
Germany, Germany
-
Stephan Ossowski, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen,
Germany, Germany
-
Leon Schütz, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen,
Germany, Germany
Presentation Overview: Show
megSAP is an open-source sequence data analysis pipeline for rare disease and oncology. Since its initial
development in 2016, the pipeline has been continuously refined to incorporate advances in short-read analysis
for Illumina sequencing data.
Since 2024, functionality has been progressively added to support long-read technologies from Oxford Nanopore
Technologies and PacBio.
This includes technology-specific mapping and variant-calling tools, as well as support for methylation
analysis and haplotype phasing based on long-read data.
Consistent with its short-read workflow, megSAP enables comprehensive detection of major variant classes,
including single-nucleotide variants and small indels, copy number variants, structural variants, and repeat
expansions.
The resulting variant lists are extensively annotated and optimized for downstream interpretation in a
clinical genetics setting.
megSAP also supports multi-sample analyses, for example in trio settings. This functionality has been extended
to allow the combined analysis of short-read and long-read variants for small variants and copy-number
variants.
Support for mixed analyses of structural variants is not yet available.
megSAP is freely available at https://github.com/imgag/megSAP
C-G.16: MtDNA Variant and Heteroplasmy Profiling in Patients with COPD
Track: Genomics, epigenomics, and genome editing
-
Daria Borodko, Institute of General Pathology and Pathophysiology, Russia
- Vasily Sukhorukov, Institute of General Pathology and Pathophysiology, Russia
- Andrey Omelchenko, Institute of General Pathology and Pathophysiology, Russia
Presentation Overview: Show
Chronic obstructive pulmonary disease (COPD) is associated with systemic inflammation and mitochondrial
impairment, yet the contribution of mitochondrial DNA (mtDNA) heteroplasmy to disease biology remains
incompletely understood. Here, we analysed mtDNA variants in the blood of 13 COPD patients.
MtDNA sequencing was performed using the Oxford Nanopore R2C2 protocol. Reads were aligned with bwa mem, and
heteroplasmy levels were quantified using mtDNA-Server2 in a Docker container.
Across all individuals, we identified two recurrent heteroplasmic variants, m.902G>C and m.912T>A,
localised to the MT-RNR1 gene. These mutations exhibited mean VAFs of 0.0798 and 0.409, respectively,
suggesting consistent low-to-intermediate heteroplasmic states across the cohort. MT-RNR1 encodes the
mitochondrial 12S rRNA and has been associated with increased cytokine levels in senescent cells, a pattern we
have observed in patient blood samples.
In addition, we detected three protein-coding variants: m.4732A>G, m.10609T>C, and
m.12406G>A—mapping to MT-ND2, MT-ND4L, and MT-ND5, respectively. These variants exhibited partial
co-occurrence patterns (pairwise Jaccard similarity 0.33-0.78). Functional annotation indicated that these
substitutions affect OXPHOS Complex I genes and include non-synonymous changes with predicted moderate
pathogenic potential.
Overall, our findings demonstrate a recurrent MT-RNR1 heteroplasmic signature across COPD patients and
identify additional heteroplasmic protein-coding variants with variable co-occurrence patterns in OXPHOS
genes. These results support a model in which mtDNA heteroplasmy may contribute to metabolic and mitochondrial
dysfunction in COPD, warranting further investigation in larger cohorts and functional studies.
This study was supported by RSF grant #24-65-00027.
C-G.17: Not All Out-of-Distribution Is Created Equal: Reasoning Demand in Perturbation Prediction
Track: Genomics, epigenomics, and genome editing
-
Anna Kalygina, ETHZ, D-INFK, Switzerland
- Alexander Theus, ETHZ, D-INFK, Switzerland
- Marina Medina Esteban, ETHZ, D-INFK, Switzerland
- Valentina Boeva, ETHZ, D-INFK, Switzerland
Presentation Overview: Show
Whether deep learning models truly outperform simple baselines for single-cell perturbation prediction remains
an open question. We argue that a key variable is missing from this debate: the rule demand imposed by the
evaluation task itself. Prediction tasks are not equally demanding - they differ in whether test split
requires learning beyond simple, interpretable rules. When mean-based or additive baselines already explain
most of the evaluable signal, even highly expressive architectures have little opportunity to demonstrate a
meaningful advantage
We introduce Rule Demand (RD), a meta-metric that quantifies the fraction of calibrated signal left
unexplained by simple baselines. Across six scenarios derived from the same synthetic gene regulatory network,
we show that splits commonly presented as “hard†out-of-distribution settings (cell-type extrapolation,
dose extrapolation, and epistatic knockout prediction) can be almost fully explained by mean-based or additive
baselines (RD ≈ 0), whereas unseen-perturbation splits preserve substantial headroom (RD → 1). We find
that this illusion of difficulty also appears in real perturbation datasets, including Norman19, Wessels23,
and XAtlas-Orion. Rule Demand offers, for the first time, an explanatory axis for why published studies report
inconsistent improvements over naive rules. Across all scenarios and datasets, model gains over simple
baselines are strongly predicted by RD.
In this work, we present a calibrated benchmarking suite that extends the Dynamic Range Fraction (DRF)
framework with signal-magnitude-aware normalization, more than 60 performance metrics, and systematic control
diagnostics for both metric choice and dataset selection.
C-G.18: Automated cell type prediction in Mass Cytometry data using machine learning with interactive
web-based annotation review
Track: Genomics, epigenomics, and genome editing
-
Gábor Beke, Institute of Molecular Biology, Slovak Academy of Sciences, Bratislava, Slovakia
-
Milan Hucko, Institute of Molecular Biology, Slovak Academy of Sciences, Bratislava, Slovakia
-
Lubos Klucar, Institute of Molecular Biology, Slovak Academy of Sciences, Bratislava, Slovakia
-
Dana Cholujova, Cancer Research Institute, Biomedical Research Center, Slovak Academy of Sciences,
Bratislava, Slovakia
-
Jana Jakubikova, Cancer Research Institute, Biomedical Research Center, Slovak Academy of Sciences,
Bratislava, Slovakia
Presentation Overview: Show
Mass cytometry (Cytometry by Time-Of-Flight - CyTOF), is an advanced technology widely used in immunology,
cancer research, drug discovery and systems biology. Compared to traditional flow cytometry, which is limited
by spectral overlap and the availability of fluorophores, mass cytometry uses heavy metal isotopes as labels.
This enables the simultaneous measurement of a greater number of markers on individual cells, resulting in
high-dimensional datasets. Analyzing CyTOF data demands robust and scalable computational approaches.
Clustering algorithms such as SPADE exist, however incorporating new samples into existing workflows typically
requires reprocessing the entire dataset. While tools such as Seurat's sketching workflow have addressed
analogous scalability challenges in single-cell RNA sequencing, there is no equivalent framework currently
available for CyTOF data. To address this limitation, we used our manually annotated SPADE results to train an
XGBoost-based classifier capable of predicting cell populations in CyTOF data with accuracy 85% - 90%,
improving upon our previous XGBoost model (~75% accuracy). This approach allows new samples to be annotated
without reprocessing the original dataset, greatly reducing computational time. To visually inspect and
validate the results, we developed a web-based visualization portal using Python and Flask. The portal allows
researchers to explore predicted cell type annotations, compare them against t-SNE and UMAP projections, and
manually review or overwrite classifications where necessary, ensuring the final annotations remain
biologically accurate. This work was supported by grants APVV-23-0482 and APVV-24-0471.
C-G.19: Management of Large Variant Datasets
Track: Genomics, epigenomics, and genome editing
-
Mohamed Abouelhoda, KFSH&RC (King Faisal Specialist Hospital & Research Center), Saudi
Arabia
- Mohamed El-Kalioby, KACST-KFSHRC, Saudi Arabia
- Saudi Arabia
Presentation Overview: Show
Currently, Next Generation Sequencing (NGS) has become a widely used technology to identify variations
associated with the disease for research and diagnostic. The variant analysis workflow on NGS data yields text
files in VCF format. This format is inefficient to query large cohorts. To solve this problem, one uses one of
the following technologies: 1) ready-to-use variant management systems (e.g., GTRAC, GenomicsDB, Gemini); 2)
native relational database management systems (e.g., MySQL) or 3) NoSQL database systems such as Clickhouse or
MongoDB.
In this poster we compare the performance of different systems in these categories (GTRAC, GenomicsDB, Gemini,
MySQL, and Clickhouse) and present best practices and recommendations.
We used 1000 Genome Project dataset (1092 VCFs), each includes ~39.7 million variants. Total size of the
database is about 76 GB. We used a server with 24-CPUs, 128GB RAM. The query set included: 1D range query to
look for variants in a genomic range (chr:start_pos-end_pos), and 2D range queries to search for variants in a
range and in group of samples.
Results: GTRAC has the best compression, but it lacks major functions for frequent queries. Clickhouse and
GenomicsDB have the least insertion time per sample. Gemini and MySQL require to build indices after
populating the tables and after each insertion. All tools in general have acceptable query/retrieval time.
MySQL has best query time due to best indexing, but its space consumption and the insertion time is
prohibitive for huge datasets. Comparing Clickhouse, Gemini and GenomicsDB, we observe that Clickhouse
performs slightly better.
C-G.20: DNA methylation changes in genomic regulatory blocks in head and neck cancer
Track: Genomics, epigenomics, and genome editing
-
Katarina Mandić, Ruđer Bošković Institute, Croatia
- Anja Barešić, Ruđer Bošković Institute, Croatia
Presentation Overview: Show
Genomic regulatory blocks (GRBs) are large genomic domains enriched in highly conserved non-coding elements
that control the expression of a single target gene, often across megabase distances. These regions are
thought to act as long-range regulatory units, integrating structural and regulatory information to maintain
precise expression patterns of the target gene. Target genes are usually transcription factors involved in
embryonic development and differentiation, and their promoters often overlap extended CpG islands that extend
into the gene body, suggesting that DNA methylation of CpGs may contribute to their regulation. However, the
methylation landscape across entire GRBs remains poorly understood, particularly in cancer. Here, we
investigate whether differential methylation within GRBs is associated with deregulation of their target genes
in head and neck cancer patients compared with controls. By integrating regional methylation analysis with
gene expression or target-gene annotation, we aim to determine whether altered CpG methylation in GRBs
reflects disrupted long-range regulatory control in cancer. Understanding these patterns could provide new
insight into the role of methylation in GRBs and how cancer alters the expression of target genes.
C-G.21: Decoding Heritable Breast Cancer Predisposition through Integrative Gene-Based Pathway
Analysis
Track: Genomics, epigenomics, and genome editing
- Shirel Schreiber, The Hebrew University of Jerusalem, Israel
- Roei Zucker, hebrew university of jerusalem, Israel
- Amos Stern, The Hebrew University of Jerusalem, Israel
-
Michal Linial, The Hebrew University of Jerusalem, Israel
Presentation Overview: Show
Heritable breast cancer (BC) risk is driven by well-established high-penetrance genes such as BRCA1, BRCA2,
PALB2, and CHEK2, yet the contribution of numerous moderate- and low-penetrance genes remains incompletely
understood. To address this gap, we implemented a gene-centric, integrative framework across large
multi-ethnic genomic resources, including the UK Biobank (UKB) and FinnGen, combining complementary
association strategies: genome-, transcriptome-, and proteome-wide association studies (GWAS, TWAS, and PWAS).
By consolidating variant-level signals into gene-level evidence and rigorously filtering likely false
positives, we defined a conservative set of 38 high-confidence BC predisposition genes. Notably, only
approximately 50% of these genes are currently represented in clinical diagnostic/prognostic panels. The set
includes established DNA repair genes alongside emerging candidates such as APOBEC3A, TNS1, and PEX14, each
supported by independent lines of evidence. PWAS further highlighted genes with potential recessive effects
that are typically overlooked by standard GWAS, underscoring the value of incorporating proteome-level
information. In parallel, a complementary family-based strategy focusing on BC-affected pedigrees leveraged
rare and ultra-rare variant analyses to delineate core biological pathways and identify shared low-frequency
contributors. Replication across cohorts demonstrated robust and consistent signals in populations of European
ancestry, with more limited transferability to other groups. Importantly, excluding individuals with non-BC
cancers (10.6%) from controls in UKB (15.7k cases) preserved risk estimates. Overall, this study provides a
stringent and interpretable framework for refining gene prioritization in heritable BC and guiding future
functional and clinical investigation.
C-G.22: SynVar: extending variant query expansion to GA4GH-compliant variant normalisation for literature
annotation
Track: Genomics, epigenomics, and genome editing
-
Anais Mottaz, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics,
Switzerland
-
Emilie Pasche, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics,
Switzerland
-
Alexandre Flament, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics,
Switzerland
-
Luc Mottin, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics,
Switzerland
-
Patrick Ruch, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics,
Switzerland
Presentation Overview: Show
SynVar is the variant query expander behind Variomes, a high-recall biomedical literature search engine
(variomes.sibils.org). SynVar recognises variants in free text as well as rsIDs and HGVS expressions across
protein, coding and genomic levels, and automatically resolves implicit reference sequences through a gene
name or chromosome. Unlike LitVar2, SynVar does not require pre-existing database entries for variant
resolution and normalisation. It resolves variants using UniProt cross-reference maps, with sequence-level
validation and coordinate mapping delegated to VariantValidator and protein back-translation to Mutalyzer.
Here, we extend SynVar beyond substitutions to deletions, insertions, duplications and delins, moving from
query expansion to full variant normalisation for literature annotation. SynVar now also accepts SPDI, VCF,
LRG and Ensembl inputs. Its normalisation mode returns canonical HGVS on GRCh38 and GRCh37, coding and protein
HGVS on the MANE Select transcript, SPDI, VCF representation, dbSNP rsID, ClinGen CAID and GA4GH VRS Allele.
For non-SNP variants, normalisation includes protein indel back-translation, 3'-shifting under HGVS rules,
left-aligned VCF output, insertion-duplication equivalence and HGVS repeat notation.
We evaluated 59 pathogenic ClinVar variants selected for class coverage and rewritten into up to 13 input
representations each (653 test cases). About 95% were parsed and 92-96% matched ClinVar across output fields.
Mismatches arose from ambiguous protein-only frameshift inputs and non-canonical NCBI-only RefSeq isoforms not
cross-referenced by UniProt. Integration into the SIBiLS annotation pipeline will enable evaluation on
biomedical literature and identification of variants not yet catalogued in existing databases. SynVar is
available at synvar.sibils.org.
C-G.23: A Bi-partite Graph Neural Network for Polygenic Risk Scoring under the Omnigenic Model
Track: Genomics, epigenomics, and genome editing
-
Ilaria Looser, Institute of AI for Health, Computational Health Center, Helmholtz Zentrum Muenchen,
Germany
-
Sergey Vilov, Institute of Computational Biology, Computational Health Center, Helmholtz Zentrum Muenchen,
Germany
-
Carsten Marr, Institute of AI for Health, Helmholtz Zentrum Muenchen / Department of Medicine III, LMU
Hospital / DKTK, Germany
-
Matthias Heinig, Institute of Computational Biology, Helmholtz Zentrum Muenchen / Department of Computer
Science, TUM / DZHK, Germany
Presentation Overview: Show
Genome-wide association studies (GWAS) identify single nucleotide polymorphisms (SNPs) that are correlated
with complex traits by computing the marginal effect size of each SNP on the trait. Subsequently, by
aggregating the individual SNP contributions into a single genetic risk estimate per individual, we can
compute the polygenic risk score (PRS), serving a clinical utility especially for early disease screening.
Current PRS models explain only a fraction of trait heritability and provide limited mechanistic insight.
While GWAS identifies SNP-trait associations, linking these signals to causal genes and pathways typically
requires separate fine-mapping. Standard PRS approaches inherit this limitation by aggregating SNP effects
without modelling relationships between variants and genes, resulting in a lack of integrated
interpretability. In contrast, the omnigenic model suggests that genetic effects propagate through gene
regulatory networks, where peripheral genes influence core disease pathways.
Here we introduce OmniGRS, a graph neural network that models PRS under the omnigenic hypothesis. OmniGRS
constructs a per-individual bipartite graph connecting SNPs to genes and extends it with protein-protein
interaction networks (STRING) to propagate regulatory context. We incorporate functional annotations (CADD)
and SNP-gene relationships (e.g., genomic proximity or eQTLs) as node features and edge weights, enabling
meaningful variant prioritisation.
OmniGRS outperforms baselines PRS models in both classification (AUC) and regression (Pearson correlation)
tasks. Attention-weighted pooling over SNP and gene nodes yields interpretable importance scores that
accurately recover causal variants in simulation.
Overall, OmniGRS provides a biologically informed framework for modelling non-additive genetic effects,
improving risk prediction and interpretability in complex traits.
C-G.24: Domain-wide Mapping of Peer-reviewed Literature for Genetic Developmental Disorders using Machine
Learning and Gene2Phenotype
Track: Genomics, epigenomics, and genome editing
-
Michael Yates, University of Edinburgh, United Kingdom
- Sarah E Hunt, EMBL-EBI, United Kingdom
- Diana Lemos, EMBL-EBI, United Kingdom
- Seeta Ramaraju Pericherla, EMBL-EBI, United Kingdom
- Elena Cibrian Uhalte, EMBL-EBI, United Kingdom
- Morad Ansari, South East Scotland Genetic Service, United Kingdom
- Louise Thompson, South East Scotland Genetic Service, United Kingdom
- Caroline Wright, University of Exeter, United Kingdom
- Helen V Firth, Addenbrooke's Hospital Cambridge University Hospitals, United Kingdom
- Ian Simpson, University of Edinburgh, United Kingdom
Presentation Overview: Show
Genetically-determined developmental disorders (GDD) are rare conditions whose diagnosis increasingly depends
on synthesis of dispersed genotype–phenotype evidence. Manual literature curation is labour-intensive and
difficult to scale. We present an automated pipeline that identifies PubMed abstracts describing human case
reports/series and maps them to GDD in Gene2Phenotype (G2P). LitDD combines a BERT abstract classifier to
detect GDD-relevant abstracts and a cross-encoder to rank candidate diseases, each fine-tuned on 13,738
annotated title–abstract pairs, with a large language model for final mapping adjudication. LitDD BERT
achieved precision 0.83 and recall 0.94; the cross-encoder achieved top-5 recall of 0.99; and the full
ensemble precision 0.89 and recall 0.82. Applied PubMed-wide, the pipeline identified 69,200 manuscripts
mapped to G2P gene-disease entities, with ~70% retrieval against independent manually curated datasets,
supporting generalisation across diverse mechanisms, inheritance patterns, and curation standards. The LitDD
corpus is incorporated into routine G2P biocuration and we provide examples where this has enabled rapid
upgrade of diseases to clinically reportable status.
To evaluate clinical utility of LitDD-derived disease models, we integrated them into Exomiser
phenotype-driven disease prioritisation and benchmarked against 7,406 DD phenopackets. Incorporating G2P
literature-derived HPO terms improved diagnostic AUC from 0.891 to 0.919. G2P disease models outperformed
baseline for 25% of diseases, particularly those lacking rich phenotype annotations. Diagnostic performance
correlated with phenotype specificity (Spearman r=0.13, p<0.001).
The LitDD corpus is openly accessible via G2P. This enables scalable, updatable literature surveillance and
faster diagnostic evidence review in genomic medicine.
C-G.25: AIOMICS4CARE WGS: AI enabled Annotation System for Whole Genome Sequencing Data
Track: Genomics, epigenomics, and genome editing
- Hanin Omer, KFSH&RC, Saudi Arabia
- Mohamed Bamajboor, KFSH&RC, Saudi Arabia
- Wedad Albalawi, KFSH&RC, Saudi Arabia
- Turki Alzahrani, KFSH&RC, Saudi Arabia
- Azza Althagafi, KFSH&RC, Saudi Arabia
- Monther Alhamdoosh, KFSH&RC, Saudi Arabia
- Abdullah Alfalah, KFSH&RC, Saudi Arabia
- Ahmad Alfares, KFSH&RC, Saudi Arabia
-
Mohamed Abouelhoda, KFSH&RC, Saudi Arabia
Presentation Overview: Show
AIOMICS4CARE WGS is a whole-genome sequencing (WGS) workflow augmented with two AI modules to improve
diagnostic yield, accelerate reporting, and support timely clinical decision-making in a high-throughput
genomic center processing approximately 400 samples per month, including urgent rapid-WGS cases with direct
patient impact. The first module, pheno priori, integrates a genotype-derived KFSH score with the patient's
clinical phenotype to produce a phenotype-informed prioritization score. The second module, ai_consilium,
classifies variants into Primary, Secondary, or Carrier findings to streamline interpretation and
reporting.
This workflow was designed to address several persistent challenges in clinical genomics, including limited
population-specific allele frequency data, incomplete representation of local variation, and the risk of
missing clinically important variants. To overcome these limitations, we incorporated a locally deployed WGS
allele-frequency database, a curated local masked database, and public reference data from gnomAD. These
resources were combined to generate a prioritized and manageable list of variants for scientist review, while
preserving sensitivity for locally relevant findings.
The workflow was implemented in a reproducible and configurable environment using Conda. Retrospective
evaluation of AI consilium on 899 samples demonstrated a 74% precision for Primary Findings when deployed on
the locally hosted QUINN model.
Overall, this workflow introduces a practical and scalable AI-assisted framework for WGS interpretation that
leverages local population knowledge, improves prioritization, and supports faster and more accurate genomic
reporting in clinical settings.
C-G.26: Searching for bacterial capsule at scale
Track: Genomics, epigenomics, and genome editing
-
Matthew Russell, EMBL-EBI, United Kingdom
- Teodora Mateeva, Wellcome Sanger Institute, United Kingdom
- Samuel Horsfield, University of Neuchâtel, Switzerland
- Stephanie Lo, EMBL-EBL, United Kingdom
- Stephen Bentley, Wellcome Sanger Institute, United Kingdom
- John Lees, EMBL-EBI, United Kingdom
Presentation Overview: Show
Capsular polysaccharides are key determinants of bacterial virulence, environmental persistence, and host
immune evasion, yet their distribution across the bacterial tree remains incompletely characterised. This is
largely due to the diversity of capsule biosynthesis systems and the limitations of existing annotation
approaches when applied at scale. Here, we present a framework for identifying capsule-producing bacteria
within large, uniformly assembled genome collections such as AllTheBacteria. We combine two complementary
search strategies. First, we implement gene locus models based on sequence homology to known capsule
biosynthesis clusters, extending concepts from tools like CapsuleFinder. Such models capsure conserved gene
content and locus structure while remaining flexibile enough to account for much of the diversity within a
capsule production system. Second, we incorporate structural homology searches for capsular proteins, enabling
detection of functionally conserved components that may evade sequence-based methods due to high divergence.
By integrating sequence- and structure-based evidence, we aim to improve sensitivity and robustness of capsule
system detection across phylogenetically distant taxa. Projecting these results onto a marker-gene-based tree
of all samples in AllTheBacteria, enables evolutionary analysis such as within-lineage capsule gain/loss and
horizontal gene transfer. Our method enables consistent annotation of capsule biosynthesis production systems
across millions of genomes. This provides a comprehensive view of capsule distribution, diversity, and
evolutionary patterns. The framework is scalable, modular, and adaptable to expanding reference datasets,
offering a deeper understanding of bacterial surface biology at ecosystem scale.
C-G.27: Benchmarking Pathogenicity Estimation Methods with Real-World Clinical Data in Inherited Heart
Disease
Track: Genomics, epigenomics, and genome editing
-
Nooshin Bayat, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
-
Angel Bernabe Garcia, Cardiogenetics Lab, IMIB,Instituto Murciano de Investigación Biosanitaria, Murcia,
Spain, Spain
-
Nuria Garcia-Santa, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
-
Laura Martinez Gomez, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
-
Kristina Ibaanez, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
- Abe Shuya, Fujitsu Laboratory, Fujitsu Ltd, Kawasaki, Japan, Japan
- Fuji Masaru, Fujitsu Laboratory, Fujitsu Ltd, Kawasaki, Japan, Japan
-
Raul Valin, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
-
Juan Ramon Gimeno, Inherited Cardiomyopathy Unit. CSUR. ERN. Hospital Virgen de la Arrixaca, Murcia,
Spain, Spain
-
Maria Sabater Molina, Legal Medicine Department, University of Murcia, Murcia, Spain., Spain
Presentation Overview: Show
Reliable classification of genetic variants, especially variants of uncertain significance (VUS), remains a
major bottleneck in clinical genomics. Pathogenicity estimation is a key step in variant classification;
however, most prediction tools are limited to specific variant types and lack full coverage.
FAI is a LightGBM-based predictor trained on ClinVar data (Oct 2023). It integrates variant-level features and
neighbouring variant information for binary classification.
We analysed 815 variants from inherited heart disease cohort (392 Pathogenic, 70 Benign, 329 VUS, 24 None) and
systematically benchmarked FAI against leading dbNSFP meta-predictors (BayesDel, MetaRNN, ClinPred,
AlphaMissense, REVEL, CADD) using balanced accuracy and coverage. Against high-confidence ClinVar variants
(>2 stars, Mar 2026, n=157), FAI achieved the best overall performance (~100% coverage, ~0.98 balanced
accuracy). In a real-world setting using hospital classifications (n=462), performance decreased across all
methods; however, FAI maintained the most favourable balance (~100% coverage, ~0.80 balanced accuracy).
Re-analysis of clinically challenging cases demonstrated translational value. FAI showed 86% concordance with
ClinGen classifications. Independently, FAI-driven reclassification, when consistent with ClinVar evidence,
downgraded five hospital-classified pathogenic variants to benign, reclassified three VUS as pathogenic, and
proposed pathogenic classification for three previously unannotated variants with clinical impact, as
confirmed by expert reconsideration. Discrepancies remained in 24% of cases, mostly due to clinical
evidence.
These results demonstrate that benchmarking must consider performance and coverage to ensure generalizability.
Coverage gaps may arise from missing annotations in resources such as dbNSFP, potentially biasing comparisons.
Importantly, variant interpretation is disease-specific and requires integration of clinical context to ensure
accurate classification.
C-G.28: Identification of conserved gene clusters across diverse prokaryotic pangenomes
Track: Genomics, epigenomics, and genome editing
-
Ikuo Uchiyama, National Institute for Basic Biology, National Institutes of Natural Sciences,
Japan
Presentation Overview: Show
Prokaryotes, including bacteria and archaea, possess a diverse collective gene repertoire despite their
relatively small individual genome sizes. This diversity is evident even within species; for instance, a
pangenome can be substantially larger than any single genome. Horizontal gene transfer (HGT) is a primary
mechanism for maintaining such diverse pangenomes, enabling adaptation to various environments. In this study,
we identified conserved gene orders within various pangenomes and compared them to identify conserved gene
clusters (CGCs) across distantly related species. We classified conserved gene orders into two types: "core
genomes," where gene orders are conserved across almost all strains, and "islands," which are conserved in
only a limited number of strains. Using the Microbial Genome Database for Comparative Analysis (MBGD), which
organizes microbial genomes into hierarchical ortholog groups (species, genus, and top levels), we employed
CoreAligner to identify core genomes. A modified version of this algorithm was then applied to the remaining
genomic regions to identify islands. After representing these regions as strings of top-level ortholog
identifiers, we used HomologyTeams to identify gene clusters in which all adjacent gene pairs are located
within a specified interval across different species, followed by a clustering algorithm to define CGCs. This
pipeline was applied to 488 species (each having at least six different strains in MBGD), resulting in the
identification of over 1,000 CGCs. We are currently analyzing these CGCs, focusing on their core/non-core
status across species as potential indicators of their propagation, including HGT events.
C-G.29: Epigenomic Instability - Is Epigenetic Age Acceleration a Survival Indicator in Lung Cancer?
Track: Genomics, epigenomics, and genome editing
-
Alexandra Anke Baumann, Department of Systems Biology and Bioinformatics, University of Rostock,
Germany
-
Michael Seifert, Institute for Medical Informatics and Biometry (IMB), TUD Dresden University of
Technology, Germany
-
Zholdas Buribayev, Department of Computer Science, Faculty of Information Technologies, Al-Farabi Kazakh
National University, Kazakhstan
-
Olaf Wolkenhauer, Department of Systems Biology and Bioinformatics, University of Rostock, Germany
-
Markus Wolfien, Institute for Medical Informatics and Biometry (IMB), TUD Dresden University of
Technology, Germany
Presentation Overview: Show
Cancer genomes are characterized by a considerable amount of modifications on various multi-omics levels,
often caused by genomic and epigenomic instability. This cancer hallmark represents, e.g., epigenetic
alterations in DNA methylation, histone remodeling, and non-coding RNA regulation. Understanding these
patterns can give insights into tumor evolution, heterogeneity, and therapeutic resistance. However, it
requires both robust data infrastructure and mechanistic insight.
To address the computational challenge of handling large-scale cancer genomic data, we developed
TCGADownloadHelper. This streamlined pipeline simplifies data retrieval from The Cancer Genome Atlas (TCGA)
via the GDC portal. Thus, transparent preprocessing of multimodal, patient-linked datasets is possible. This
tool formed the foundation for our downstream analyses of DNA methylation data across TCGA lung cancer
cohorts.
Building on these resources, we investigated epigenetic age acceleration (EAA) across the lung cancer cohorts
TCGA-LUAD and TCGA-LUSC. Illumina 450K DNA methylation data was obtained from the GDC portal via the
TCGADownloadHelper. The deviation between biological (epigenetic) and chronological age was estimated from DNA
methylation patterns using established epigenetic clocks (Horvath, Zhang2019,...). We demonstrate that
approximately two thirds of tumors exhibited accelerated epigenetic aging, while one third showed a
""rejuvenating"" effect. Interestingly, patients with rejuvenating EAA showed significantly lower overall
survival in Kaplan-Meier analyses. EAA varied by sex, tumor stage, and cancer type, highlighting the
importance of patient-specific factors in the epigenetic age.
Together, these findings introduce epigenetic aging signatures as a promising stratification biomarker and
motivate joined mechanistic studies on genomic and epigenomic instability to characterize actionable targets
in precision oncology.
C-G.30: Unveiling the functional fate of duplicated genes through expression profiling and structural
analysis
Track: Genomics, epigenomics, and genome editing
-
Alex Warwick Vesztrocy, BioSoft Research UK, United Kingdom
-
Natasha Glover, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
-
Paul D Thomas, University of Southern California, SIB Swiss Institute of Bioinformatics, United States
-
Christophe Dessimoz, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
- Irene Julca, Aarhus univeristy, Denmark
Presentation Overview: Show
Gene duplication is a major evolutionary source of functional innovation. Following duplication events, gene
copies (paralogs) may undergo various fates, including retention with functional modifications (such as
subfunctionalization or neofunctionalization) or loss. When paralogs are retained, this results in complex
orthology relationships, including one-to-many or many-to-many. In such cases, determining which one-to-one
pair is more likely to have conserved functions can be challenging. It has been proposed that, following gene
duplication, the copy that diverges more slowly in sequence is more likely to maintain the ancestral function
referred to here as the least diverged ortholog (LDO) conjecture. This study explores this conjecture, using a
novel method to identify asymmetric evolution of paralogs and applying it to all gene families across the Tree
of Life in the PANTHER database. Structural data for over 1 million proteins and expression data for 16
animals and 20 plants are used to investigate functional divergence following duplication. This analysis, the
most comprehensive to date, reveals that, whereas the majority of paralogs display similar rates of sequence
evolution, significant differences in branch lengths following gene duplication can be correlated with
functional divergence. Overall, the results support the least diverged ortholog conjecture, suggesting that
the least diverged ortholog tends to retain the ancestral function, whereas the most diverged ortholog (MDO)
may acquire a new, potentially specialized role.
C-G.31: Integrative gene network analysis of genome-wide association data in myalgic encephalomyelitis /
chronic fatigue syndrome
Track: Genomics, epigenomics, and genome editing
-
Ekaterina Antonenko, CBIO Mines Paris PSL, France
- Giann Karlo Aguirre-SambonÃ, CBIO Mines Paris PSL, France
- Florian Massip, CBIO Mines Paris PSL, France
- Chloé-Agathe Azencott, CBIO Mines Paris PSL, France
Presentation Overview: Show
Myalgic encephalomyelitis / chronic fatigue syndrome (ME/CFS) is a common though poorly understood disease
affecting millions of people worldwide. The biological mechanisms underlying ME/CFS remain largely unclear, no
effective treatments currently exist, and the disease disproportionately affects females and is frequently
triggered by acute infection. However, no satisfactory mechanistic explanation for either factor has been
established.
In the DecodeME study, the first genome-wide association studies (GWAS) were performed on a large cohort of
cases (15,579) and controls (259,909) with European genetic ancestry. Eight loci were reported to be
significantly associated with ME/CFS, three of which are proximate to genes involved in the response to viral
or bacterial infection, consistent with the known infection trigger. The initial findings also suggest that
both immunological and neurological processes contribute to the genetic risk of ME/CFS.
However, GWAS is by design limited to individual SNP-phenotype associations, and gene-gene interactions are
largely overlooked by this framework. Polygenic diseases such as ME/CFS are likely shaped by the coordinated
activity of multiple genes within shared biological pathways, rather than by isolated variants alone. Gene
network methods integrating pathway and interaction data have therefore been developed to refine and enrich
classical GWAS signals, and combining multiple such methods has been shown to improve statistical power and
interpretability, with successful applications in breast cancer and psoriasis.
In the present study, we re-analyse the DecodeME GWAS summary statistics and apply a combination of gene
network methods across curated pathway databases and experimentally derived protein interaction networks. This
integrative approach yields a robust consensus of genes potentially involved in ME/CFS pathogenesis. The
network analysis partly recovers the genes from the original study and additionally identifies multiple
previously unreported pathways and candidate genes related to immune regulation and neurological function.
C-G.32: Evolutionary Conservation and Functional Constraints of TP53 Mutation Hotspots Across Mammalian
Species
Track: Genomics, epigenomics, and genome editing
-
Ritika Rawat, University of Mumbai, India
- Sermarani Nadar, University of Mumbai, India
- Gursimran Kaur Uppal, University of Mumbai, India
Presentation Overview: Show
The tumor suppressor gene TP53 plays a central role in maintaining genomic stability and is one of the most
frequently mutated genes in human cancers. Investigating the evolutionary conservation of TP53 mutation
hotspots across species can provide insights into their functional importance and selective constraints.
In this study, we performed a comparative genomic analysis of TP53 across multiple mammalian species to
identify conserved regions and mutation hotspots. TP53 sequences were retrieved from publicly available
databases (NCBI) and subjected to multiple sequence alignment and phylogenetic analysis. Conservation scoring
was used to identify functionally constrained regions, and known human mutation hotspots were mapped onto
conserved domains.
Our analysis reveals that several mutation hotspots in human TP53 coincide with highly conserved regions
across mammals, indicating strong evolutionary pressure to maintain their functional integrity. These regions
predominantly correspond to DNA-binding domains essential for transcriptional regulation and tumor
suppression. In contrast, less conserved regions exhibit variability suggestive of species-specific
adaptations.
Overall, this study highlights the evolutionary significance of TP53 mutation hotspots and provides insights
into their functional constraints across species. These findings contribute to a better understanding of
cancer-associated mutations within an evolutionary framework.
C-G.33: PanTEon: a cross-kingdom framework to guide the design of transposable element classifiers
Track: Genomics, epigenomics, and genome editing
-
Simon Orozco-Arias, Life Science Department, Barcelona Supercomputing Center, BSC-CNS, 08032
Barcelona, Spain, Spain
- Iamil Ferrer-Pomer, Universitat Oberta de Catalunya, 08018 Barcelona, Spain, Spain
-
Fabiana Rodrigues de Goes, The Rosalind Franklin Institute, OX11 0QX Didcot, United Kingdom, United
Kingdom
-
Simon Gaviria-Orrego, Department of Computer Science, Universidad Autónoma de Manizales, 170001
Manizales, Colombia, Colombia
-
Juan Gómiz-Fernández, Universitat Oberta de Catalunya, 08018 Barcelona, Spain, Spain
- Jordi Llatser-Torres, Universitat Oberta de Catalunya, 08018 Barcelona, Spain, Spain
-
Alexandre R. Paschoal, The Rosalind Franklin Institute, OX11 0QX Didcot, United Kingdom, United Kingdom
-
Romain Guyot, UMR DIADE, IRD, CIRAD, Université de Montpellier, Montpellier, France, France
-
Toni Gabaldón, Life Science Department, Barcelona Supercomputing Center, BSC-CNS, 08032 Barcelona, Spain,
Spain
Presentation Overview: Show
Transposable elements (TEs) are fundamental drivers of genome evolution, yet their annotation and
classification remain inconsistent, fragmented, and difficult to reproduce across species. This challenge
arises from sequence divergence, lineage-specific innovations, and heterogeneous taxonomies across databases
and computational tools, ultimately limiting large-scale comparative analyses. Here, we present PanTEon, a
cross-kingdom deep learning framework designed to enable reproducible, scalable, and standardized TE
classification. PanTEon integrates two key components: (i) the PanTEon Database, an automatically curated
repository comprising ~240,000 structurally validated TE sequences from 2,790 species across animals, plants,
and fungi, and (ii) a modular benchmarking and training platform that supports parallel evaluation and
deployment of multiple ML and DL models. The PanTEon framework enables unified training, inference, and
comparison across nine ML/DL architectures, while remaining fully extensible to user-defined models. Using
this standardized environment, we benchmark seven state-of-the-art TE classifiers and demonstrate that
classification performance is strongly influenced by taxonomic origin and TE superfamily, revealing
significant biases and limitations in current approaches. Furthermore, we leverage the PanTEon training module
to systematically retrain and compare nine architectures on a harmonized task involving 30 TE superfamilies.
Our results show that ensemble strategies and taxon-specific models substantially improve predictive
performance, while cross-kingdom generalization remains a key unresolved challenge. Additionally, we
demonstrate the flexibility of the framework by addressing auxiliary tasks such as false positive detection in
TE libraries. Overall, PanTEon establishes the first cross-kingdom, deep learning-driven ecosystem for TE
classification, providing a robust foundation for benchmarking, model development, and large-scale genomic
analyses.
C-G.34: DeepCAST-GWAS: Improving the Discovery of Genetic Associations Using Deep Learning-Based Regulatory
SNP Prioritization
Track: Genomics, epigenomics, and genome editing
-
Lovro Rabuzin, ETH Zurich, Switzerland
- Konstantin Heep, ETH Zurich, Switzerland
- Sophie Sigfstead, University of Alberta, Canada
- Valentina Boeva, ETH Zurich, Switzerland
Presentation Overview: Show
Genome-wide association studies (GWAS) have uncovered numerous variants linked to complex traits, yet power
remains limited by the large multiple testing burden and the inclusion of many variants with minimal
regulatory impact. We present Deep learning-based Chromatin Accessibility SNP Targeting for GWAS
(DeepCAST-GWAS), a framework that integrates functional annotations derived from deep learning models to
improve both the yield and the reliability of GWAS findings. DeepCAST-GWAS uses SNP Activity Difference (SAD)
scores from in silico mutagenesis with the Enformer model to estimate the predicted effect of each variant on
chromatin accessibility across tissues, allowing statistical testing to focus on variants with stronger
regulatory evidence. Using conservative family-wise error rate (FWER) control, DeepCAST-FWER produces fewer
associations than existing power-boosting approaches, but the associations it reports replicate in larger
cohort GWAS at substantially higher rates. For applications where discovery count is more important,
DeepCAST-sFDR increases the number of genome-wide significant findings above baseline GWAS by using the
Enformer SAD scores for stratified False Discovery Rate (sFDR) control. DeepCAST-sFDR achieves performance
comparable to the strongest competing method, while maintaining reliability on par with a standard GWAS.
Subsampling analyses across a wide range of traits confirm these improvements in both sensitivity and
replicability. DeepCAST-GWAS offers a principled way to incorporate sequence-based regulatory predictions into
population-scale association testing, demonstrating that chromatin accessibility activity scores can improve
the stability of GWAS discoveries.
C-G.35: Causal thinking in a correlated world: improving genomic prediction through SNP selection in the AI
era
Track: Genomics, epigenomics, and genome editing
-
Thomas Crow, The University of Queensland, ARC CoE for Plant Success in Nature and Agriculture,
Australia
-
David Kainer, The University of Queensland, ARC CoE for Plant Success in Nature and Agriculture, Australia
Presentation Overview: Show
Many plant breeding programs have embraced genomic prediction to accelerate crop improvement. A central
challenge in genomic prediction is the number of genetic markers (eg. SNPs) greatly exceed the number of
phenotyped individuals, reducing accuracy.
Dimensional reduction methods like linkage disequilibrium and minor allele frequency filtering can reduce this
imbalance, but they may filter out informative SNPs.
Clearly, our genomic prediction models would benefit from selecting only SNPs which contribute to phenotype
variation. Such causal SNPs would be invaluable, but the true genetic architecture of a trait is typically
unknown.
To what extent can informed SNP selection identify or approximate these causal SNPs and how does doing so
affect genomic prediction accuracy?
This talk presents a conceptual framework for SNP selection in genomic prediction, discussing how different
strategies attempt to filter for causal SNPs. We examine existing approaches spanning statistical selection
using Bayesian shrinkage and machine learning models, as well as biological selection methods informed by GWAS
signals, gene expression, functional annotations, and variant effect prediction.
We also discuss emerging directions in AI-augmented SNP selection, where DNA foundation models, large language
models, and graph-based methods can be used to integrate many biological studies into a single SNP selection
model. Finally, we present a new approach which uses biological knowledge networks for SNP selection.
Genomic prediction accuracy will not improve by simply adding more markers, but by using SNP selection to
better reflect the underlying biology of traits.
C-G.36: Managing workflow executions with WESkit
Track: Genomics, epigenomics, and genome editing
-
Landfried Kraatz, Berlin Institute of Health at Charité Universitätsmedizin Berlin, Germany
-
Valentin Schneider-Lunitz, Berlin Institute of Health at Charité Universitätsmedizin Berlin, Germany
-
Sven Twardziok, Berlin Institute of Health at Charité Universitätsmedizin Berlin, Germany
Presentation Overview: Show
Managing computational workflows across diverse biomedical projects, each with its own parameters, tools, and
execution environments, poses persistent challenges for scalability, reproducibility, and collaborative
research. We introduce WESkit, a robust implementation of the Global Alliance for Genomics and Health (GA4GH)
Workflow Execution Service (WES) specification that unifies the execution, monitoring, and documentation of
data‑processing workflows. By supporting both Snakemake and Nextflow, WESkit enables consistent automation
and centralized oversight across large numbers of heterogeneous workflow runs. This design empowers research
groups and service units to maintain long‑term reproducibility, streamline multi‑project operations, and
scale computational efforts with confidence. Seamless integration with cloud infrastructures further positions
WESkit as a practical contributor to the GA4GH cloud ecosystem and a valuable tool for modern, collaborative
biomedical data analysis.
C-G.37: Quantitative Modeling of Clone-Specific Treatment Resistance via Joint Bayesian Inference of
Compositional and Population Size Data
Track: Genomics, epigenomics, and genome editing
-
Mohammad Darbalaei, Bioinformatics and Computational Biophysics, University of Duisburg-Essen, Essen,
Germany, Germany
-
Julia Zummack, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis, Univ.
Duisburg-Essen, Germany, Germany
-
Thomas Mühlenberg, German Cancer Consortium (DKTK), partner site Essen/Düsseldorf, DKFZ and Univ.
Duisburg-Essen, Germany, Germany
-
Philip Dujardin, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis,
Univ. Duisburg-Essen, Germany, Germany
-
Susanne Grunewald, German Cancer Consortium (DKTK), partner site Essen/Düsseldorf, DKFZ and Univ.
Duisburg-Essen, Germany, Germany
-
Patricia Munteanu, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis,
Univ. Duisburg-Essen, Germany, Germany
-
Marina Martinez Cruz, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and
Metastasis, Univ. Duisburg-Essen, Germany, Germany
-
Madeleine Dorsch, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis,
Univ. Duisburg-Essen, Germany, Germany
-
Alexander Schramm, West German Cancer Center, Dept. of Medical Oncology, Molecular Oncology, Univ.
Duisburg-Essen, Germany, Germany
-
Sebastian Bauer, German Cancer Consortium (DKTK), partner site Essen/Düsseldorf, DKFZ and Univ.
Duisburg-Essen, Germany, Germany
-
Barbara M. Grüner, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis,
Univ. Duisburg-Essen, Germany, Germany
-
Daniel Hoffmann, Bioinformatics and Computational Biophysics, University of Duisburg-Essen, Essen,
Germany, Germany
Presentation Overview: Show
Multiplexed assays such as DNA-barcoded cell line mixtures generate compositional sequencing data that capture
only relative clone abundances. This inherent constraint complicates the inference of absolute, clone-specific
treatment effects, as changes in relative abundance do not necessarily reflect true growth behavior.
Here, we develop a hierarchical Bayesian modeling framework to infer clone-specific treatment responses by
jointly integrating compositional sequencing data with independent measurements of total population size, such
as confluency in vitro or tumor volume in vivo. Barcode count data are modeled using a Dirichlet–multinomial
likelihood to capture sampling variability and biological overdispersion. Tumor volume and confluency
measurements are modeled using log-normal and beta likelihoods, respectively, enabling inference of absolute
growth. By combining these data sources, the framework reconstructs clone-specific changes in absolute
abundance under treatment relative to control. A hierarchical structure captures variability across replicates
and treatment conditions, enabling partial pooling and stabilizing estimates in settings with limited
observations. Inference is performed in a fully probabilistic manner, allowing uncertainty from both
sequencing and population-level measurements to be propagated to the final estimates. This yields a
quantitative measure of treatment response for each clone with associated uncertainty, resolving ambiguities
inherent to compositional data.
We validated the method by comparison of experimental data from treatment responses in cancer cell line
mixtures and individual cell lines. This approach provides a general framework for multiplexed drug screening
assays and reduces the number of required animal experiments through pooled experimental designs.
C-G.38: PansimNuc: Rapid nucleotide-level simulation of selection and genetic element mobility in eukaryote
pangenomes
Track: Genomics, epigenomics, and genome editing
-
Samuel Horsfield, University of Neuchâtel, Switzerland
- Tobias Baril, University of Neuchâtel, Switzerland
- Jigisha Jigisha, University of Neuchâtel, Switzerland
- Daniel Croll, University of Neuchâtel, Switzerland
Presentation Overview: Show
Many eukaryotic species possess huge inter-individual genomic diversity, known as a "pangenome", made up of
small variants at a single nucleotide level, and large variants consisting of gain or loss of entire
chromosomes. Pangenome diversity drives specie's adaptability and evolution, however, the evolutionary
mechanisms structuring pangenomes remain unknown.
To understand how pangenomes arise and what mechanisms underpin their persistence, we developed PansimNuc, a
rapid nucleotide-level population simulator, written in Rust. PansimNuc simulates small variation at the level
of single nucleotide polymorphisms and indels, and large scale chromosomal rearrangements through genetic
element mobility, including transposable element (TE) dynamics and inter-genome recombination. PansimNuc also
simulates demographic dynamics, including splitting and admixture of populations, and migration events.
Finally, PansimNuc simulates selection at a nucleotide-level, enabling the impact of fitness effects in many
complex evolutionary scenarios to be explored.
We show simulated populations generated by PansimNuc can reproduce complex dynamics including joint effects of
genetic drift, migration and selective sweeps. Additionally, PansimNuc can recapitulate genome-wide TE copy
number expansions expected if active TEs are present in a population. Tracking of haplotype characteristics
over time across a variety of evolutionary scenarios provides a powerful tool to explore drivers of eukaryotic
pangenome diversity at high resolution.
We foresee that PansimNuc has a large range of applications, including genome annotation tool benchmarking and
training of statistical models, such as machine-learning and deep-learning approaches, to identify complex
genome dynamics in genome sequence data.
C-G.39: Expanding Gene Ontology coverage of the human functionome in the PAN-GO project
Track: Genomics, epigenomics, and genome editing
-
Marc Feuermann, SIB Swiss Institute for Bioinformatics, Geneva, Switzerland
-
Huaiyu Mi, Department of Population and Public Health Sciences, University of Southern California, Los
Angeles CA, United States
- Li Ni, The Jackson Laboratory for Mammalian Genomics, Bar Harbor, ME, United States
- Pascale Gaudet, SIB Swiss Institute for Bioinformatics, Geneva, Switzerland
-
Anushya Muruganujan, Department of Population and Public Health Sciences, University of Southern
California, Los Angeles CA, United States
-
Dustin Ebert, Department of Population and Public Health Sciences, University of Southern California, Los
Angeles CA, United States
-
Tremayne Mushayahama, Department of Population and Public Health Sciences, University of Southern
California, Los Angeles CA, United States
-
Paul D. Thomas, SIB Swiss Institute for Bioinformatics, Geneva / University of Southern California, Los
Angeles CA, Switzerland
Presentation Overview: Show
Understanding the functions of protein-coding genes in the human genome has long been a central goal of
biomedical research. Over the past two decades, significant progress has been made, driven by improved genome
annotation and major advances in the experimental characterization of human genes and their homologs in
well-studied model organisms. Within the Gene Ontology knowledgebase, functional information has been
systematically integrated using an evolutionary framework based on phylogenetic trees. This approach enabled
the construction of curated models of function evolution and supported the initial assignment (version 1.0,
February 2025) of at least one functional characteristic to 82% of human protein-coding genes. Importantly,
each annotation can be traced back to experimental evidence from human and/or non-human model systems.
Here, we describe the second phase of this effort, which focuses on improving existing annotations and
expanding coverage by addressing annotation gaps through both manual curation and AI-guided approaches. In
version 2.0 of the PAN-GO functionome (https://functionome.geneontology.org/), coverage of human
protein-coding genes has increased to 84% (an increase of 400 genes compared to version 1.0). Additionally,
annotations were improved for thousands of genes, and 47% of genes now have Gene Ontology annotations across
all three aspects - Molecular Function (MF), Biological Process (BP), and Cellular Component (CC) -
representing a 5% increase (1000 genes) compared to the previous version. Further improvements have already
been made in preparation for a third version.
C-G.40: From DMPs to DMRs: CMEnt, Characterization of methylation using positional entanglement
Track: Genomics, epigenomics, and genome editing
-
Vasileios Lemonidis, UAntwerp - KU Leuven, Belgium
- Ken Op de Beek, UAntwerp, Belgium
- Joris Vermeesch, KU Leuven - UZ Leuven, Belgium
- Timon Vandamme, UZ Antwerp, Belgium
- Guy Van Camp, UAntwerp, Belgium
- Joe Ibrahim, UAntwerp, Belgium
Presentation Overview: Show
DNA Cytosine methylation is a naturally occurring point modification closely linked to chromatin state.
Recruitment of de novo methyltransferases is facilitated by histone modifications, while DNA methylation
itself can promote heterochromatin formation. This reciprocal relationship creates structured local
dependencies among methylated cytosines, including correlation patterns that may inform the identification of
differentially methylated regions (DMRs). Here, we leverage this property in a model-light framework for
regional methylation analysis. We present CMEnt (Characterization of Methylation through positional
Entanglement), an R package for statistically guided DMR identification from significant differentially
methylated positions (DMPs). CMEnt starts from user-defined seed loci, typically DMPs, and tests within-group
correlation among proximal seeds to assemble them into connected clusters. These clusters are then extended to
neighbouring significantly correlated cytosines and merged when overlapping, yielding candidate DMRs. This
design provides a coherent transition from DMP-level discoveries to region-level interpretation, enabling
linked downstream analyses at both resolutions. CMEnt supports multiple input formats, including methylation
microarray beta-value matrices indexed by genomic position, BED-like genomic coordinate files, and BSseq
objects generated from short- or long-read sequencing assays such as whole-genome bisulfite sequencing (WGBS)
and Oxford Nanopore Technologies (ONT). Across real-world datasets, CMEnt produces highly specific regions,
while simulations show strong agreement with known ground truth when benchmarked against established DMR
detection method. CMEnt therefore provides a competitive and complementary strategy for differential
methylation analysis across diverse experimental platforms.
C-G.41: Computel 2.0: A Scalable Platform for Comprehensive Telomere Profiling and TRV Phenotyping
Track: Genomics, epigenomics, and genome editing
-
Davit Tarverdyan, Armenian Bioinformatics Institute, Institue of Molecular Biology NAS,
Armenia
-
Anahit Yeghiazaryan, Armenian Bioinformatics Institute, Institue of Molecular Biology NAS, Armenia
-
Tatevik Jalatyan, Armenian Bioinformatics Institute, Institue of Molecular Biology NAS, Armenia
- Lilit Nersisyan, Armenian Bioinformatics Institute, Armenia
Presentation Overview: Show
Telomeres cap chromosome ends, and their shortening and dysfunction drive genome instability in aging and
cancer. Telomere length, repeat-variant composition (Telomeric Repeat Variants, TRVs), and fusion events are
informative readouts of telomere state, yet existing tools resolve only parts of it: Mean Telomere Length
(MTL) estimates assume genome-wide average coverage reflects a normal diploid genome, an assumption that
collapses under the copy number alterations (CNAs) and aneuploidy pervasive in tumors, while TRV detection is
limited to simple substitutions, fusions are rarely detected, and these capabilities stay split across
disconnected tools.
To address these limitations, we present Computel 2.0, a unified, Python-based workflow for comprehensive
telomere analysis. This updated platform systematically integrates CNA-aware MTL estimation, the profiling of
TRVs, and the detection of telomeric fusions directly from short-read sequencing data (FASTQ/BAM). Computel
2.0 introduces a dynamic normalization algorithm that adjusts for coverage distortions, ensuring highly
accurate MTL calculations even in highly aberrant cancer genomes. Furthermore, the platform achieves advanced
TRV phenotyping by extracting a wider range of TRVs, including insertions and deletions.
We have applied Computel 2.0 across diverse sequencing modalities, demonstrating its scalability and precision
on bulk, single-cell, and cell-free DNA (cfDNA) datasets. To streamline data exploration, the tool generates
interactive HTML reports visualizing MTL distributions, TRV composition, and read-specific patterns. By
effectively resolving telomere dynamics and characterizing TRV phenotypes, Computel 2.0 serves as a
high-precision, scalable platform for investigating telomere biology.
C-G.42: From ageing clocks to human digital twins in personalising healthcare through biological age
analysis
Track: Genomics, epigenomics, and genome editing
-
Murih Pusparum, Hasselt University and VITO NV, Belgium
- Gokhan Ertaylan, VITO NV, Belgium
- Olivier Thas, Hasselt University, Belgium
-
Simone Ecker, UCL Cancer Institute, University College London, London, United Kingdom
-
Stephan Beck, UCL Cancer Institute, University College London, London, United Kingdom
Presentation Overview: Show
Age is the most important risk factor for the majority human diseases, yet chronological age alone poorly
captures the biological diversity shaping individual health trajectories. Biological age (BA) predictors
derived from epigenomic, proteomic, metabolomic, and clinical biochemistry data offer a promising way to
quantify ageing as a dynamic, measurable process. In this study, we demonstrate the value of BA analysis
within the IAM Frontier cohort, a deeply phenotyped 13-month longitudinal study of 30 healthy adults aged
45–59 years. Across repeated time points, we computed BA and health-related predictions using 29 epigenetic,
4 clinical-biochemistry, 2 proteomic, and 3 metabolomic clocks, alongside gold-standard clinical risk
indicators.
Our findings show that ageing signatures differ markedly between individuals while remaining relatively stable
within individuals, with epigenetic clocks providing the most consistent long-term signals and
clinical/proteomic clocks showing greater sensitivity to short-term physiological change. Subject-level
analyses revealed that multi-omics BA profiles could highlight subtle deviations, including smoking-related
risk, abnormal lipid profiles, shortened telomere predictions, and immune-cell composition changes, several of
which aligned with clinical or self-reported health indicators.
These results position BA predictors not merely as retrospective ageing measures, but as actionable biomarkers
for preventive and personalised medicine. Integrated into human digital twin frameworks, longitudinal BA
measurements could anchor real-time models of individual health, detect early departures from expected
trajectories, and support simulation of lifestyle or therapeutic interventions. This work underscores the
potential of multi-omics ageing clocks to complement routine clinical testing and advance scalable, adaptive,
and biologically informed digital twins for precision healthcare.
C-G.43: Annotating Eukaryotic Genomes by Combining Deep Learning with Extrinsic Evidence
Track: Genomics, epigenomics, and genome editing
- Lars Gabriel, University of Greifswald, Germany
-
Katharina Jasmin Hoff, University of Greifswald, Germany
Presentation Overview: Show
The accuracy of ab initio gene prediction has been significantly advanced by Tiberius, a deep-learning gene
finder that utilizes convolutional and LSTM layers paired with a differentiable Hidden Markov Model (HMM).
However, even state-of-the-art deep learning architectures can benefit from the integration of extrinsic
biological data to capture alternative splicing. We present a scalable evidence processing pipeline designed
to complement the ab initio output of Tiberius by incorporating extrinsic evidence from transcriptomic and
proteomic sources.
The pipeline operates as an integration framework that independently processes short-read and long-read RNA
sequencing data, alongside large-scale protein database alignments. This allows for: (1) the addition of
alternative isoforms supported by RNA-seq/Iso-Seq evidence that the ab initio model may not prioritize, (2)
the refinement of gene boundaries using protein homology to validate coding sequences, (3) increased
sensitivity in regions where genomic signals are weak but extrinsic support is robust.
We provide a modular path to incorporate experimental data into gene prediction with Tiberius. This pipeline
improves transcript-level accuracy across eukaryotic datasets, offering a comprehensive solution for
high-quality genome annotation. Tiberius and its associated evidence workflow are available at
https://github.com/Gaius-Augustus/Tiberius.
C-G.44: Convergent niches, divergent genomes: a four-species pan-GWAS of Aspergillus pathogenicity and
domestication
Track: Genomics, epigenomics, and genome editing
-
Minji Kim, BRIGHT, Technical University of Denmark, Denmark
- Eduard Kerkhoven, Chalmers University of Technology, Sweden
- Patrick Phaneuf, BRIGHT, Technical University of Denmark, Denmark
Presentation Overview: Show
The genus Aspergillus contains the filamentous fungi of substantial combined clinical and industrial
importance: A. fumigatus and A. flavus are the principal agents of invasive aspergillosis, while A. niger and
A. oryzae drive global industries in enzyme production and food fermentation. Whether these convergent
phenotypes reflect shared genomic adaptations has not been tested at scale across the genus.
To address this, we constructed per-species pangenomes for four Aspergillus species of 211 ANI-verified,
high-quality genomes spanning A. fumigatus (n=89), A. flavus (n=70), A. niger (n=19), and A. oryzae (n=33),
together with a genus-level pangenome of 15,163 orthogroups. In the accessory genome, phenotype-labelled
pan-genome-wide association studies (pan-GWAS) identified 2-117 significant orthogroup associations per
species-phenotype contrast (BH-FDR<0.05), yet cross-species convergence testing revealed no shared
evolutionary tactics. Even literature-curated virulence and industrial gene panels in the core genome offered
no gene-content explanation: 76% of 226 anchored trait genes were core in all four species, and none of 124
virulence genes were pathogen-specific-indicating that lifestyle is not encoded by the presence or absence of
shared genomic toolkits.
Instead, niche adaptation operates through lineage-specific rare gene compartments. Human-pathogenic strains
showed significant rare gene expansion in A. fumigatus (p=8.1×10⁻⁹) and A. flavus (p=0.014), while industrial
strains carried fewer lineage-specific genes than environmental strains (pooled, p=0.018). The rare gene
compartment, often discarded as noise, represents the primary evolutionary substrate for clinical and
biotechnological adaptation in this genus.
C-G.45: Gemsparcl: Rapid and consistent clustering of millions of genomes highlights the diversity of
prokaryotic life
Track: Genomics, epigenomics, and genome editing
-
Johanna von Wachsmann, European Bioinformatics Institute, University of Cambridge, United
Kingdom
- John A. Lees, European Bioinformatics Institute, United Kingdom
- Robert D. Finn, European Bioinformatics Institute, United Kingdom
Presentation Overview: Show
Bacterial genome databases collectively approach ten million assembled genomes, yet redundancy and limited
scalability of existing tools create bottlenecks for comprehensive, tree-of-life-scale genomic analyses. One
widely used approach involves dereplication to remove redundancy, but methods struggle beyond a few thousand
genomes, making global organisation of public genomes computationally infeasible.
Here we present gemsparcl, a tool that clusters bacterial genomes into genomically coherent units (GCUs),
orders of magnitude faster than existing methods. Central to gemsparcl is sketchlib.rust, a highly efficient
one-permutation MinHash implementation with an auxiliary inverted index that substantially accelerates
all-versus-all genome comparisons. gemsparcl further applies statistical correction for incomplete
metagenome-assembled genomes (MAGs) and uses network-based quality filtering to remove edges weakly connecting
distinct GCUs.
We clustered 5.6 million high-quality bacterial genomes (2.88 million isolates and 2.77 million MAGs) into
92,954 GCUs in approximately 14 hours using 48 cores and less than 16.5 GB of memory, achieving 99.94% cluster
purity. Delving into these results revealed 2,582 GCUs, each comprising more than 50 genomes, with clusters
formed exclusively by MAGs, highlighting priorities for future isolation efforts. More detailed analysis of
the GCU networks provides insights into genomic diversity, with work ongoing to routinely extract this
information from the clusters.
The scalability of gemsparcl transforms what was previously intractable: routine and up-to-date organisation
of all public bacterial genomes and large-scale biodiversity assessment.
C-G.46: Whole-organism sequence-to-function modelling stratifies the cis-regulatory code of Drosophila into
enhancer and locus grammars
Track: Genomics, epigenomics, and genome editing
-
Eren Can Eksi, VIB-KU Leuven, Belgium
- Anton De Brabandere, VIB-KU Leuven, Belgium
- Swann Floc'Hlay, VIB-KU Leuven, Belgium
- Berfin Dag, VIB-KU Leuven, Belgium
- Valerie Christiaens, VIB-KU Leuven, Belgium
- Katina Spanier, VIB-KU Leuven, Belgium
- Stein Aerts, VIB-KU Leuven, Belgium
Presentation Overview: Show
The canonical model of gene regulation is based on enhancer-promoter interactions to determine spatiotemporal
properties of gene expression. So far, most studies on gene regulation in complex multicellular organisms have
focused on particular tissues or cell types, having limited power in investigating possible higher
hierarchical level gene regulatory principles. Therefore, we set out to investigate the rules of gene
regulation and enhancer grammar in the context of a whole organism using sequence-to-function models. For
this, we generated a whole adult fly scATAC-seq atlas containing around 700,000 cells and 150,000 peaks that
are accessible in at least one cell type. We used topic modelling to cluster the cells in the atlas and
transferred cell type annotations from the scRNA-seq atlas at two hierarchical levels (class and specific).
This resulted in a pseudo-multiome atlas with matched scATAC-seq and scRNA-seq profiles across more than 100
cell types. We stratified scATAC-seq peaks according to their accessibility across cell types into ubiquitous,
cell-type-specific, and putative cell-class-specific types. Next, we trained a suite of local and global
sequence-to-function models to understand the grammars of the three classes of regions we identified and to
integrate locus level information to predict gene regulation. Using these models, we identified and validated
cell-type-specific enhancers for different cell types, investigated the promoters of genes and the relation of
promoter-related motifs to gene expression patterns, and investigated the grammar and gene expression effects
of a novel putative class of regions that we call 'class-level enhancers'.
C-G.47: MoleMap: fast alignment-free molecule mapping for long-read and linked-read sequencing data
Track: Genomics, epigenomics, and genome editing
-
Richard Lüpken, Leibniz Institute for Immunotherapy, Germany
- Thomas Krannich, German Cancer Research Center, Germany
- Markus Schuelke, Charite at Universitätsmedizin Berlin, Germany
- Birte Kehr, Medizinische Hochschule Hannover, Germany
Presentation Overview: Show
With sequencing costs continuing to decline, the computational expenses of read alignment are becoming an
increasingly important consideration. When only portions of the genome are of interest --as in most clinical
genome analyses-- or base-pair precise alignment is not required --as for most local assembly approaches--,
performing full read alignment on whole-genome sequencing (WGS) data imposes unnecessary computational
costs.
We introduce MoleMap, a “molecule mapping†method for mapping long reads and linked-read molecules. MoleMap
uses a minimized open-addressing k-mer index of the reference genome and a fast k-mer clustering procedure to
map reads or molecules. In benchmarks, MoleMap agrees with minimap2 alignments of PacBio HiFi and ONT long
read data on 99.68% and 98.03% of mappings respectively, outside of centromere and satellite repeat regions,
while running 3-8x faster than other mappers and 10-60x faster compared to aligners on 32 threads. In
addition, its low memory footprint allows whole-genome processing on a standard laptop computer. We apply
MoleMap to filter reads for local assembly of a known variant region including non-reference sequence and
showcase its benefits in diagnosing a patient with a rare disease.
MoleMap is an efficient mapping tool for all applications that do not require base pair precise alignment.
Further, MoleMap can serve as a pre-processing tool for WGS data, accelerating targeted analyses by selecting
reads from loci of interest before costly alignment. In a clinical sequencing setting, MoleMap can
substantially reduce computational costs and diagnostic turnaround times.
C-G.48: Federated Genomic Variant Discovery & Analysis Using GA4GH Standards
Track: Genomics, epigenomics, and genome editing
-
Anais Mottaz, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics,
Switzerland
- Michael Baudis, Swiss Institute of Bioinformatics, Switzerland
- Valerie Barbie, Swiss Institute of Bioinformatics, Switzerland
- Tim Beck, University of Nottingham, United Kingdom
- Melissa Cline, UC Santa Cruz Genomics Institute, United States
- Friederike Ehrhart, Maastricht University., Netherlands
- Hindrik Kerstens, Princess Maxima Center, Netherlands
- Worawich Phornsiricharoenphant, Swiss Institute of Bioinformatics, Switzerland
- Jordi Rambla, Centre de Regulacio Genomica, Spain
- David Salgado, Institut Francais de Bioinformatique (IFB-Core), France
- Venkata Satagopam, Luxembourg Centre For Systems Biomedicine (LCSB), Luxembourg
- Sergi Beltran, Centro Nacional de Analisis Genomico (CNAG), Spain
- Emidio Capriotti, University of Bologna, Italy
Presentation Overview: Show
Genomic data are increasingly distributed across clinical, research and national infrastructures, making
interoperability a key challenge for their effective reuse. This is particularly critical for rare disease,
cancer and population genomics, where variant interpretation requires integrating both observational data and
interpretative knowledge across sources.
GA4GH standards provide a common framework to address these challenges. The Variant Representation
Specification (VRS) enables consistent, computable representations of genomic variants across formats and
sources, while Categorical VRS (Cat-VRS) extends this to groups of related variants. The Beacon protocol
supports federated discovery across distributed datasets without centralising sensitive data, with controlled
access and adjustable granularity of query responses. Downstream analyses can be performed within Trusted
Research Environments (TREs), where standards such as the Task Execution Service (TES) support federated
execution by bringing computation to the data.
This work illustrates how these components can be articulated into a federated analysis pipeline, using a rare
disease-oriented scenario as an illustrative example. Starting from a candidate variant, consistent
representations enable linking observational occurrences with interpretative evidence from databases and
literature, including related variants within the same functional class. Federated discovery through Beacon
identifies relevant datasets, while subsequent analyses within TREs enable comparison of phenotypic profiles
across cohorts.
The work highlights current efforts and remaining challenges in interoperability, including variant
representation harmonisation, integration of heterogeneous data types and alignment between discovery and
analysis layers.
C-G.49: Large-scale genetic analysis of healthcare expenditure in 1.4 million individuals: The GenCost
Consortium
Track: Genomics, epigenomics, and genome editing
- Finngen, University of Helsinki, Finland
- Andrea Ganna, University of Helsinki, Finland
- Jakob German, University of Helsinki, Finland
- Tatiana Cajuso Pons, University of Helsinki, Finland
- Mykyta Artomov, University of Helsinki, Finland
- Nikita Kolosov, University of Helsinki, Finland
- Padraig Dixon, Nuffield Department of Primary Care Health Sciences, United Kingdom
-
Søren Brunak, Department of Public Health, Novo Nordisk Foundation Center for Protein Research, Denmark
- Sarah E. Medland, QIMR Berghofer, Australia
- Neil Davies, University College London, United Kingdom
- Nicholas G Martin, QIMR Berghofer, Australia
- Riccardo Marioni, University of Edinburgh, United Kingdom
- Dorret Boomsma, Vrije Universiteit Amsterdam, Netherlands
- Hamdi Mbarek, Qatar Precision Health Institute, Qatar
- David van Heel, Queen Mary University of London, United Kingdom
-
Sebastian May-Wilson, University of Helsinki, Finland
- Pradeep Natarajan, Broad Institute of MIT and Harvard, United States
- Zhiyu Yang, University of Helsinki, Finland
- Kristina Zguro, University of Helsinki, Finland
- Erik Abner, University of Tartu, Estonia
- Arne Kukkonen, University of Tartu, Estonia
- Patrick Fahr, University of Oxford, United Kingdom
- Stavroula Kanoni, Queen Mary University of London, United Kingdom
- Camiel van der Laan, Vrije Universiteit Amsterdam, Netherlands
- Chadi Saad, Qatar Precision Health Institute, Qatar
- Anne Richmond, University of Edinburgh, United Kingdom
- Penelope A. Lind, QIMR Berghofer, Australia
-
Ioannis Louloudis, Department of Public Health, Novo Nordisk Foundation Center for Protein Research,
Denmark
- Jiwoo Lee, Broad Institute of MIT and Harvard, United States
- Tomoko Nakanishi, University of Helsinki, Finland
Presentation Overview: Show
We performed the largest genome-wide association study of healthcare costs to date, aiming to define the
genetic architecture underlying genetic variation leading to increased healthcare expenditure. We analysed
four major phenotypes: inpatient, inpatient plus outpatient, primary care, and prescription drug costs in up
to 1.4 million individuals from 11 studies across 7 countries, using harmonized cost phenotypes derived from
electronic health records and administrative databases. The cost phenotypes were analysed via GWAS
meta-analysis, followed by conditional analyses, HLA fine-mapping, rare variant burden testing,
colocalization, and polygenic score evaluation.
We identify 380 conditionally independent common variant associations across 248 loci, with the strongest and
most pleiotropic signals located in the HLA region. Individual common variants had modest effects, typically
altering annual costs by approximately 1–2% per allele, whereas rare deleterious variants in clinically
actionable cancer predisposition genes, including BRCA1, BRCA2, MSH2, and APC, were associated with
substantially larger increases in annual inpatient costs. Colocalization analyses linked cost-associated loci
to autoimmune, cardiometabolic, pain-related, and psychiatric traits, indicating that healthcare expenditure
reflects a broad spectrum of underlying disease biology.
Polygenic scores for healthcare costs predicted expenditure in independent cohorts and retained significant
effects in within-family analyses, supporting largely direct genetic influences. Established polygenic scores
for diseases and risk factors explained even greater variance in costs than cost-derived scores themselves.
These findings show that healthcare expenditure has a measurable polygenic and rare variant architecture,
providing a foundation for integrating human genetics into health economics, preventive strategies, and
population-level screening.
C-G.50: High-Resolution Characterization of Repeat Expansions in Friedreich's Ataxia by Targeted Long-Read
Sequencing
Track: Genomics, epigenomics, and genome editing
-
Christina Matlok, Dept. Molecular Neurology, FAU, Erlangen, Germany; ZSEER University Hospital Erlangen,
Erlangen, Germany, Germany
-
Isabell Cordts, Department of Neurology, TUM University Hospital, Munich, Germany, Germany
-
Angela Abicht, Medical Genetics Center (MGZ), Munich, Germany; Dept. of Neurology, LMU, FBI, LMU München,
Munich, Germany, Germany
-
Anna Anna Benet-Pagès, Medical Genetics Center (MGZ), Munich, Germany; Inst. of Neurogenomics, Helmholtz
Center Munich, Neuherberg, Germany, Germany
-
David Brenner, Dept. of Neurology, University Hospital, Ulm, Germany; DZNE, Ulm, Germany; Center for Rare
Diseases (ZSE), Ulm, Germany, Germany
-
Thomas Klopstock, FBI, Department of Neurology; DZNE, Munich, Germany; SynNergy, Munich, Germany, Germany
-
Martin Regensburger, Dept. of Molecular Neurology, FAU, Erlangen, Germany; ZSEER, University Hospital
Erlangen, Erlangen, Germany, Germany
-
Annekathrin Rödiger, Department of Neurology, UKJ, Jena, Germany; Center for Rare Diseases, UKJ, Jena,
Germany, Germany
-
Hayrettin Tumani, Department of Neurology, University Hospital Ulm, Ulm, Germany, Germany
-
Herbert Schreiber, Neurological Practice Center, Neuropoint Academy & NTD, Ulm, Germany, Germany
-
Benjamin Vlad, Department of Neurology, Jena University Hospital, Jena, Germany, Germany
-
Andreas Hauser, Medical Genetics Center (MGZ) Munich, Munich, Germany; Department of Neurology, TUM
University Hospital, Munich, Germany, Germany
-
Franziska Bachhuber, Department of Neurology, University Hospital Ulm, Ulm, Germany, Germany
-
Boriana Büchner, Friedrich-Baur-Institute, Department of Neurology, LMU University Hospital, Munich,
Germany, Germany
-
Almut Bischoff, Friedrich-Baur-Institute, Department of Neurology, LMU University Hospital, Munich,
Germany, Germany
-
Jasper Hesebeck-Brinckmann, Department of Neurology, University Hospital Ulm, Ulm, Germany, Germany
-
Finja Grimm, Department of Neurology, TUM University Hospital, Munich, Germany, Germany
-
Marie Hackenberg, Medical Genetics Center (MGZ) Munich, Munich, Germany; Department of Neurology, TUM
University Hospital, Munich, Germany, Germany
- Thomas Risch, Medical Genetics Center (MGZ) Munich, Munich, Germany, Germany
- Vitus Prokosch, Medical Genetics Center (MGZ) Munich, Munich, Germany, Germany
-
Ricarda von Heynitz, Department of Neurology, TUM University Hospital, Munich, Germany, Germany
- Florentine Scharf, Medical Genetics Center (MGZ) Munich, Munich, Germany, Germany
Presentation Overview: Show
Friedreich's Ataxia (FRDA) is an autosomal recessive neurodegenerative disorder caused by pathological GAA
repeat expansions within intron 1 of the FXN gene, with the shorter of two expanded alleles being the primary
determinant of disease severity and age of onset. Non-GAA interruptions within the repeat are suspected to
influence clinical manifestation but remain incompletely characterized by conventional methods. Long-read
sequencing now enables comprehensive repeat characterization of repeat length and sequence.
We investigated GAA repeat characteristics in 39 FRDA patients from 5 centers in Germany using PureTarget
Cas9-based enrichment with PacBio HiFi long-read sequencing. HiFi reads were processed using an in-house
pipeline combining TRGT-based repeat length quantification, benchmarked against orthogonal methods (PCR,
Southern Blot), including custom modules for signal processing-driven repeat interruption detection and motif
composition determination.
Across the cohort, we obtained a mean of 257 repeat-spanning reads for the shorter allele. Repeat length
estimates showed concordance with orthogonal methods (r=0.75), with shorter repeat lengths confirming the
known inverse correlation with the age of onset (r=-0.48). Interruptions in short repeat alleles and
non-canonical motif composition were identified in 17 and 2 patients, respectively findings inaccessible to
conventional methods.
Targeted PacBio HiFi sequencing enables high-resolution characterization of GAA repeat expansions, including
sequence features inaccessible to conventional approaches. This work contributes to deepening our
understanding of the structural complexity and variability of repeat expansions in FRDA. Future efforts aim to
explore additional molecular features, including methylation, and further evaluate long-read sequencing as a
comprehensive diagnostic tool in clinical diagnostics.
C-G.51: Design and Validation of Deep Learning Models for Decoding Cis-Regulatory Logic in Drosophila
melanogaster
Track: Genomics, epigenomics, and genome editing
-
Berfin Dag, VIB Center for AI and Computational Biology & KU Leuven, Belgium
-
Valerie Christiaens, VIB Center for AI and Computational Biology & KU Leuven, Belgium
- Lukas Mahieu, VIB Center for AI and Computational Biology & KU Leuven, Belgium
-
Anton De Brabandere, VIB Center for AI and Computational Biology & KU Leuven, Belgium
- Julie De Man, VIB Center for AI and Computational Biology & KU Leuven, Belgium
-
Gert Hulselmans, VIB Center for AI and Computational Biology & KU Leuven, Belgium
- Stein Aerts, VIB Center for AI and Computational Biology & KU Leuven, Belgium
Presentation Overview: Show
Gene expression is controlled with remarkable precision across cell types and developmental stages through
cis-regulatory elements, which integrate information about cell identity, expression magnitude, timing, and
signalling inputs at the level of DNA sequence. Current deep learning models have shown that the sequence
contains sufficient information to predict cell-type-specific enhancer activity, yet this represents only a
fraction of the regulatory information embedded in the genome. Whether models can extract higher-order
regulatory features and design functionally equivalent synthetic elements remains an open question.
Here we present a computational and experimental framework to decode cis-regulatory grammar in Drosophila
melanogaster, using eye development as a model system where all cell types and regulatory factors are well
characterized. We generated the largest single-cell multiome atlas of Drosophila eye development to date,
integrating newly generated and public scATAC-seq and scRNA-seq datasets comprising over 100,000 cells across
all major cell types. Using this atlas, we trained two complementary models: CREsted, which addresses the
challenge of predicting cell-type-specific enhancer activity from local sequence context and enables design of
synthetic enhancers with defined expression pattern; and a fine-tuned Fly-Borzoi model, which predicts how
combinations of enhancers, their genomic context, and long-range interactions jointly determine gene
expression level across a ~131kb receptive field.
To functionally validate model predictions, we have established an experimental framework based on precise
genome editing at genomic loci, enabling quantitative transcriptional readout and phenotypic assessment. This
pipeline is designed to link sequence-level predictions to organismal outcomes, providing a rigorous in vivo
benchmark for sequence-to-function modeling.
C-G.52: Comparative benchmarking of DNA methylation imputation methods in the human placenta
Track: Genomics, epigenomics, and genome editing
-
Chen Zhang, University of Cambridge, United Kingdom
- Dafina Angelova, University of Cambridge, United Kingdom
- Gordon Smith, University of Cambridge, United Kingdom
- Steve Charnock-Jones, University of Cambridge, United Kingdom
- Sung Sam Gong, University of Cambridge, United Kingdom
Presentation Overview: Show
DNA modifications such as 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC), are key epigenic
regulatory mechanisms. It is now possible to study DNA methylation at single-base resolution. However, if the
coverage is less than the desired depth, the level of methylation of a given CpG is uncertain and is normally
given as 'missing'. Imputation approaches have shown potential for addressing this missingness in the DNA
methylation data, but comprehensive benchmarking studies that consider genomic regulatory context remain
limited. Here, we benchmarked 10 imputation methods and compared their performance using the human placenta
methylation data from the Pregnancy Outcome Prediction Study (POPS). We obtained the methylation data from 119
placenta samples using the Biomodal evoC platform that measures both 5hmC from 5mC. We generated ground truth
datasets across four genomic regions (promoters, CpG islands, gene bodies and 5000bp tiling arrays) including
only CpG sites with at least 10X— coverage across all samples. Imputation performance was evaluated separately
for 5mC and 5hmC using Root Mean Square Error (RMSE). Our findings indicate that mean imputation method
consistently outperformed other approaches evaluated, such as boostme, missForest and imputePCA. For example,
within gene bodies, the RMSE of the mean method was the smallest for both 5mC (0.102) and 5hmC (0.051), while
KNN and softImpute produced the highest RMSE values (0.176 at 5mC and 0.0746 at 5hmC) relative to the other
methods. In summary, the mean approach offers a robust and accurate strategy for handling missing values in
placental mC and hmC sequencing data.
C-G.53: Polygenic scores derived from breed-level obesity traits predict food motivation in individual
dogs
Track: Genomics, epigenomics, and genome editing
-
Enoch Alex, University of Cambridge, United Kingdom
- Jade Scardham, University Of Cambridge, United Kingdom
- Anna Morros-Nuevo, University of Cambridge, United Kingdom
- Natalie Wallis, University of Liverpool, United Kingdom
- Alyce McClellan, University College London, United Kingdom
- Eleanor Raffan, University of Cambridge, United Kingdom
Presentation Overview: Show
Human polygenic risk scores can predict complex disease liability, but their performance often weakens when
applied across populations with different allele frequencies, linkage disequilibrium and environments.
Domestic dogs offer a comparative version of the same problem, intensified by breed structure, where breeds
differ in demography, haplotype background, obesity susceptibility and appetite related behaviour. We asked
whether obesity related genetic effects estimated at breed level retain predictive value when applied to
individual dogs.
Using breed average GWAS summary statistics for obesity probability and food motivation, we built polygenic
scores using clumping and thresholding, LDpred2 and lassosum2. Scores were tuned in 325 dogs and evaluated in
487 independent multibreed dogs, with individual food motivation as the target and covariate adjustment for
age, sex, neuter status and genetic principal components.
The food motivation derived score predicted individual food motivation with beta = 0.338, p = 5.64e-05 and
incremental R² = 0.0318. The food motivation derived score stratified high food motivation, with top quintile
dogs showing 3.91 fold higher adjusted odds than bottom quintile dogs, p = 0.007. Signals were retained in
mixed breed dogs. External transfer was cohort dependent, with prediction of body condition in Labradors, beta
= -0.217 and p = 0.007, but not Golden Retrievers.
These results show that breed level obesity genetics can retain modest individual level signal for appetite
related behaviour, while highlighting phenotype alignment and population structure as shared constraints on
polygenic prediction across human and canine genetics.
C-G.54: Genomic insights into racing camels: inbreeding levels and positive selection linked to athletic
traits
Track: Genomics, epigenomics, and genome editing
-
Hussain Bahbahani, Kuwait University, Faculty of Science, Department of Biological sciences,
Kuwait
Presentation Overview: Show
Racing dromedary camels are widely distributed across the Arabian Peninsula, with the highest concentrations
in its northern and southeastern regions. In this study, whole‑genome sequences from 34 racing camels were
analyzed to evaluate their genetic relationships with non‑racing populations, estimate inbreeding levels,
calculate Weir and Cockerham's fixation index (Fst), assess effective population size (Ne), and identify
genomic regions exhibiting signatures of positive selection.
Both racing and non‑racing camels showed comparable levels of genomic inbreeding (FROH = 0.21), and no
significant genetic differentiation was detected between the two groups. The estimated Fst values further
confirmed minimal population structure. Ne estimates revealed a declining trend in both groups over the past
5,000 years, with racing camels exhibiting slightly lower recent Ne values compared to their non‑racing
counterparts.
Signatures of positive selection in racing camel genomes were identified using two haplotype‑based
statistics—the integrated haplotype homozygosity score (iHS) and the between‑population extended haplotype
homozygosity test (Rsb)—alongside a runs‑of‑homozygosity (ROH) analysis. A total of 33 candidate regions
were detected using iHS, 19 using Rsb, and 24 through ROH. These regions overlapped with genes involved in
biological pathways potentially linked to athletic performance, including musculoskeletal development, lipid
metabolism, stress response, bone integrity, endurance, and power.
Overall, these findings explore the racing dromedary genome, with the aim of identifying variants and
haplotypes associated with athletic traits. Such insights could support the development of genetically
informed breeding programs designed to enhance specialized racing lines.