View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

C-G.01: Adaptive Dynamics of HIV-1 Populations Over a Six-Year-Long Experimental Evolution: A Genomic Perspective
Track: Genomics, epigenomics, and genome editing
  • Ali Movasati, Department of Infectious Diseases and Hospital Epidemiology, Universitätsspital Zürich, Switzerland, Switzerland
  • Christine Leemann, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital Zürich, Switzerland, Switzerland
  • Kathrin Neumann, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital Zürich, Switzerland, Switzerland
  • Rongfeng Chen, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital Zürich, Switzerland, Switzerland
  • Lygeri Sakellaridi, University of Würzburg, Institute for Virology and Immunobiology, Germany, Germany
  • Karin Metzner, Department of Infectious Diseases and Hospital Epidemiology, Universtitätsspital Zürich, Switzerland, Switzerland
  • Roland Regoes, Institute of Integrative Biology, ETH Zürich, Switzerland, Switzerland


Presentation Overview: Show

Numerous experimental evolution studies have suggested that adaptation rate of microbial populations evolving in stable environments decline over time. To investigate the characteristics of adaptation deceleration in a fast-evolving virus, we propagated HIV-1 in two human T-cell lines (MT-2 and MT-4) for approximately 4.8 years and tracked its genome evolution through NGS. The sequencing data can be explored via LTEEviz, an interactive web application. Time-resolved sequencing data indicated that despite constant fixation rate of 0.085 (MT-2) and 0.042 (MT-4) mutations per generation, the fixation kinetics of adaptive mutations changed considerably over time. The rate of fixation of adaptive parallel mutations decreased by 44% per 300 generations, while their conferred fitness gain diminished by 27% (MT-2) and 18% (MT-4) per every added adaptive mutation in their genetic background. Furthermore, we identified unique yet consistent patterns of sequence evolution among different regions of the HIV-1 genome. In particular, nef and vpr accessory genes demonstrated patterns of random evolution, expected in the absence of selection. Condordantly, the evolving populations of HIV-1 acquired and fixed multiple loss-of-function mutations in nef and vpr. Additionally, nef gene repeatedly underwent large deletions, leading to the removal of approximately 400 bp (equivalent to 4.3% of the HIV-1 genome). These large nef deletions increased in frequency faster than expected under neutrality and in a length-dependent manner. Together, our results confirm that HIV-1 genomic evolution is characterized by a swift and substantial deceleration of adaptation, while highlighting progressive genome shrinkage as one of the underlying adaptive mechanisms.

C-G.02: Decoding the noncoding: a computational and CRISPR-enabled framework from long noncoding RNA discovery to therapeutic targeting
Track: Genomics, epigenomics, and genome editing
  • Maina Bitar, QIMRB, Australia
  • Stacey Edwards, QIMRB, Australia
  • Juliet French, QIMRB, Australia
  • Haran Sivakumaran, QIMRB, Australia


Presentation Overview: Show

Long noncoding RNAs (lncRNAs) represent a vast and largely unexplored layer of cancer biology, harbouring the majority of somatic mutations in cancer. Here, we present an integrated computational and functional genomics framework to systematically discover, characterise and prioritise lncRNAs for therapeutic targeting in breast and ovarian cancers.

Over the past four years, we developed a framework powered by a range of computational tools. We started developing in silico metatranscriptome assembly strategies to uncover thousands of previously unannotated lncRNAs from relevant cell and tissue samples, initially using short-read data (ShROOM) and now extended to hybrid long- and short-read integration (HyDRA). These tools enabled the redefinition of normal breast epithelial cell populations and the construction of pseudo-longitudinal models of ovarian cancer progression and chemoresistance, comprehensively mapping lncRNA involvement in disease.

To prioritise candidates from this expanded gene set, we computationally integrate lncRNA discovery with genetic association signals, including genome-wide association studies. Functional interrogation is achieved through a suite of CRISPR-based technologies. We pioneered RNA-targeting CRISPR–Cas13 screens, establishing the first platform for transcript perturbation at scale, and are extending this approach to single-cell readouts (CROP-seq). We are also developing high-throughput CRISPR prime editing to assess the functional impact of thousands of lncRNA mutations.

This framework enables the identification of clinically actionable lncRNAs, including functional drivers like BRRIAR, a strong candidate for oestrogen receptor–positive breast cancer therapy, the highlight of this talk. Our work establishes a scalable path from discovery to therapeutic targeting, positioning lncRNAs as central figures in precision oncology.

C-G.03: Predicting Genomic Determinants of Chromatin Compaction Using Machine Learning
Track: Genomics, epigenomics, and genome editing
  • Ryan Burke, CEZAMAT, Warsaw University of Technology, Poland
  • Ewelina Holm Bidstrup, CEZAMAT, Warsaw University of Technology, Poland
  • Shuting Liu, University of Illinois at Urbana-Champaign, United States
  • Ilaria Lupi, CEZAMAT, Warsaw University of Technology, Poland
  • Monika Staniszewska, CEZAMAT, Warsaw University of Technology, Poland
  • Andrew Belmont, University of Illinois at Urbana-Champaign, United States
  • Teresa Szczepinska, CEZAMAT, Warsaw University of Technology, Poland


Presentation Overview: Show

The hierarchical 3D organization of chromatin in eukaryotic cells plays a crucial role in gene regulation and adapts during development, in response to stimuli, and in disease. While large-scale chromatin domains at scales larger than TADs (up to hundreds of nanometers) have been observed using microscopy, recent advances such as PCC-seq enable high-throughput measurement of chromatin compaction at kilobase resolution, revealing local variability, including those around transcription start sites (TSS).
Here, we applied machine learning approaches, including Logistic Regression, Random Forest, and Histogram-based Gradient Boosting, to classify genomic regions according to chromatin compaction levels. Models were trained for regions of TSS (±1 kb, ±5 kb), genome-wide (1 kb, 100 kb), considering 2, 3, and 5 compaction classes. Predictions were based on 519 genomic features, including proteins and histone modifications ChIP-seq data, chromatin accessibility (ATAC-seq, DNA-seq), and transcription (BRU-seq, RNA-seq). As a baseline comparison, an analogous classification was performed for ATAC-seq signal in TSS ±1 kb regions.
Using permutation feature importance and single-feature ROC AUC scores, we ranked genomic marks associated with chromatin compaction. In TSS ±1 kb regions, EP400, ELF4, POLR2H, H2AFZ, and ATAC-seq signals were associated with decompaction, while H3K27me3 and STAG1 correlated with compaction. Moreover, in TSS ±5 kb regions, H3K27ac and PCBP1 were linked to decompaction, whereas MCM3 correlated with compaction. At whole genome scales, in 100 kb bins THRAP3 and MBD1 were associated with decompaction, while ZBTB33 and H3K9me3 correlated with compaction. At 1 kb resolution, SETDB1, ZC3H4, and ZNF263 were associated with increased compaction.

C-G.04: GPxLMM: Gaussian Process-Augmented Linear Mixed Models for Genotype-by-Environment Interaction Analysis
Track: Genomics, epigenomics, and genome editing
  • Bibiana Mailyn Horn, Hasso Plattner Institute, University of Potsdam, Germany
  • Zoran Nikoloski, University of Potsdam; Max Planck Institute of Molecular Plant Physiology, Germany
  • Christoph Lippert, Hasso Plattner Institute, University of Potsdam; Icahn School of Medicine at Mount Sinai, Germany


Presentation Overview: Show

Multivariate genome-wide association studies are essential for understanding phenotypic plasticity and identifying genetic variants driven by genotype-by-environment (GxE) interactions. Existing methods, however, rely on analytically derived gradients and fixed covariance assumptions, which can lead to model misspecification and reduce power when GxE effects vary smoothly across continuous environmental gradients. To address this, we introduce GPxLMM, a framework integrating Gaussian process covariance learning with linear mixed model inference for fixed effects. By leveraging automatic differentiation in GPyTorch, new covariance structures can be incorporated without deriving analytical gradients. Alongside the standard index kernel for discrete phenotypes, we introduce a reaction norm kernel to model the genetic covariance of function-valued phenotypes. By capturing continuous phenotypic variation across changing environments, this approach enables phenotype prediction and genotype ranking in unobserved environments. In simulations using real Arabidopsis thaliana genome data, GPxLMM accurately detects rescaling and heterogeneous GxE effects and produces well-calibrated variance decompositions. For discrete phenotypes, it matches LIMIX in statistical power while retaining the flexibility to incorporate alternative covariances. In leave-one-environment-out cross-validation across 21 DROPS environments, the reaction norm kernel achieved a mean predictive correlation of ~0.70. Furthermore, when selecting the top 20 genotypes, the model demonstrated on average 75% selection efficiency. To conclude, GPxLMM provides an extensible framework for integrating flexible covariance learning into linear mixed models, advancing genomic prediction and quantitative genetics beyond fixed structural assumptions.

C-G.05: nf-core/raredisease: Modernizing Clinical Genomics by Transitioning to Community-Driven Workflows
Track: Genomics, epigenomics, and genome editing
  • Ramprasad Neethiraj, KTH Royal Institute of Technology, Stockholm, Sweden, Sweden
  • Anders Jemt, Genomics Medicine Centre Karolinska, Karolinska University Hospital, Stockholm, Sweden, Sweden
  • Henrik Stranneheim, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm, Sweden, Sweden
  • Peter Pruisscher, Genomics Medicine Centre Karolinska, Karolinska University Hospital, Stockholm, Sweden, Sweden
  • Daniel Nilsson, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm, Sweden, Sweden
  • Chiara Rasi, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm, Sweden, Sweden
  • Valtteri Wirta, Department of Microbiology, Tumour and Cell Biology, Karolinska Institutet, Stockholm, Sweden, Sweden


Presentation Overview: Show

The clinical diagnosis of rare genetic disorders requires the seamless integration of single-nucleotide variants (SNVs), structural variants (SVs), and repeat expansions. For over a decade, the Stockholm healthcare region relied on the in-house Mutation Identification Pipeline (MIP) to process 10,000+ clinical samples. To meet increasing demands for scalability, portability, and long-term sustainability, we developed and transitioned to nf-core/raredisease as its official clinical successor.
nf-core/raredisease leverages the collective logic and diagnostic experience gained from MIP while adopting the modern Nextflow framework and nf-core best practices. The pipeline orchestrates state-of-the-art tools for whole-genome and exome sequencing into a containerized, FAIR-compliant workflow. Key features include integrated mitochondrial analysis, rank-based variant prioritization, and comprehensive detection of complex variants.
Having recently replaced MIP in routine clinical production at Karolinska University Hospital, nf-core/raredisease demonstrates how legacy clinical expertise can be successfully ported to a community-driven model. This transition has significantly improved maintenance efficiency and reproducibility without sacrificing the diagnostic rigor established over years of service. This poster details the pipeline's architecture, the technical challenges of the migration process, and the benefits of moving from a bespoke legacy system to a standardized open-source framework. By sharing this transition, we provide a blueprint for other clinical centers looking to modernize their genomic infrastructure through collaborative, peer-reviewed software development.

C-G.06: Role of histone H2A lysine119 ubiquitination in gene regulation in BAP1-positive and BAP1-negative uveal melanoma eye cancer.
Track: Genomics, epigenomics, and genome editing
  • Anneke Brummer, University of Lausanne, Switzerland
  • Adeline Berger, Hopital ophtalmique Jules-Gonin, Switzerland
  • Nicolas Guex, University of Lausanne, Switzerland
  • Alexandre Moulin, Hopital ophtalmique Jules-Gonin, Switzerland


Presentation Overview: Show

Uveal melanoma (UM), though a rare cancer affecting only about 5 per million adults, is the most common primary eye tumor. BAP1 (BRCA1-associated protein 1) inactivation is a critical mutation associated with metastasis and poor prognosis. BAP1 deubiquitinates histone H2A lysine 119 (H2AK119ub, known for gene silencing) and thereby activates chromatin and gene expression. Here, we analyse data from 11 BAP1-positive and 12 BAP1-negative UM patient tumor samples, integrating genetic, epigenetic, and transcriptomic information, to better understand the role of K119ub in gene expression regulation in these tumor subgroups. We first confirmed that transcriptomes of BAP1-positive and BAP1-negative subgroups were clearly separated. This was mostly due to differential gene regulation, and not to gene copy number variations. Overall, more genes had higher expression levels in BAP1-positive tumors, in agreement with an activating function of BAP1. However, notably many genes were also more expressed in BAP1-negative samples. We next sought to relate gene expression with chromatin state differences considering 5 histone modifications (H2AK119 ubiquitination, H3K27 tri-methylation, H3K27 acetylation, H3K4 mono-methylation and H3K4 tri-methylation). Inferred chromatin states largely agreed with measured chromatin accessibility by ATAC-Seq. Surprisingly, around transcription start sites, repressive chromatin, characterized by K27me3 and K119ub, was more abundant in BAP1-positive than BAP1-negative samples. In contrast, in BAP1-negative tumors, K119ub more often co-localized with K4me1 in this region, including in active chromatin states together with K27ac and K4me3. Active chromatin states without K119ub were more abundant in BAP1-positive samples, as expected. Overall, chromatin states agreed with gene expression in BAP1 subgroups.

C-G.07: Imputation of Bacterial Genomes based on cgMLST Profiles
Track: Genomics, epigenomics, and genome editing
  • Christina Kirschbaum, Robert Koch Institute, Germany
  • Torsten Houwaart, Robert Koch Institute, Germany
  • Vladimir Bajić, Robert Koch Institute, Germany
  • Simon H. Tausch, Robert Koch Institute, Germany
  • Hugues Richard, Robert Koch Institute, Germany


Presentation Overview: Show

Accurate pathogen identification is a crucial step for effective treatment, outbreak management and risk characterization. Nowadays a common genotyping methodology for bacteria is core genome multi locus sequence typing (cgMLST), which characterizes samples by a set of genes commonly found in the species of interest. In metagenomics, new methods with shallow sequencing are on the rise and a tool to account for missing genes in genotyping methods is needed. We propose a method that performs genotype imputation similar to approaches for human genomes.
We first assessed linkage between genes, and showed that by grouping highly correlated loci, we could reach compression levels up to 29% for Listeria monocytogenes and over 50% with Mycobacterium tuberculosis given canonical typing schemes. We then implemented a Markov model (MM) on the core genome that estimates allele succession on adjacent genes. On Listeria monocytogenes, MM always outperformed the baseline imputation method using maximum frequency allocation when tested on different thresholds of masked alleles. Masking 15% of the alleles, MM correctly reconstructs 67% of the alleles (baseline: 26%) when applied to the reference data, and 59% (baseline: 24%) on a larger pathogen.watch and BigsDB dataset. Preliminary results on Mycobacterium tuberculosis are more mitigated (64%/69% MM, 74%/84% baseline). The performance is likely compounded by a smaller reference set and a low allelic variability leading to overfitting.
We are aiming towards an improved model considering haplotype blocs which would be tested on an extended set of reconstruction scenarios.

C-G.08: A systematic investigation into the robustness of mutational signature fitting to unknown signatures
Track: Genomics, epigenomics, and genome editing
  • Maria Katsantoni, Department for BioMedical Research, Inselspital, Bern University Hospital and University of Bern, Bern, Switzerland, Switzerland
  • Qixuan Wang, Department for BioMedical Research, Inselspital, Bern University Hospital and University of Bern, Bern, Switzerland, Switzerland
  • Matúš Medo, Department for BioMedical Research, Inselspital, Bern University Hospital and University of Bern, Bern, Switzerland, Switzerland


Presentation Overview: Show

Accurate deconvolution of mutational signatures leads to a better understanding of tumor etiology, yet the mathematical stability of signature attribution remains a significant challenge in the field. As highlighted in previous benchmark studies, existing algorithms used to fit mutational signatures frequently struggle with two intertwined phenomena: the misallocation of stochastic noise as biological signal (overfitting) and the "displacement" of mutations from signatures absent in reference catalogs onto known ones. The latter represents a critical limitation where a model's incomplete basis (underfitting) directly causes the spurious inflation of existing signatures (overfitting), leading to misleading clinical interpretations.

Building upon previous software for benchmarking signature analysis tools, we have tested a broad range of synthetic samples to characterize the exact conditions, such as low mutational burden and high cosine similarity between signatures, under which existing tools deviate from biological ground truth. We developed a new algorithm based on regularized deconvolution and residual analysis to isolate 'unknown' signals even when the underlying signatures are undefined, thereby preventing the erroneous mapping of mutations to the reference catalog.

To ensure the portability and reproducibility of these findings, we have implemented the entire analytical pipeline as a modular Snakemake workflow. Preliminary results suggest that regularized deconvolution can achieve more robust and transparent signature assignments. Our work provides a diagnostic perspective on the limitations of current fitting paradigms and offers a scalable path toward more reliable genomic reporting.

C-G.09: Exploration of twin genes with DupyliCate and biological implications
Track: Genomics, epigenomics, and genome editing
  • Shakunthala Natarajan, University of Bonn, Germany
  • Claudia Sterling, University of Bonn, Germany
  • Boas Pucker, University of Bonn, Germany


Presentation Overview: Show

Paralogs, copies of a gene, form an important basis for novelty during evolution. Analysis of such gene duplications is important to understand the emergence of novel evolutionary traits. DupyliCate is a Python tool that has been developed for the identification, classification, and characterization of gene copies. With the ability to process multiple datasets concurrently, flexible features, and parameters to set species-specific thresholds, DupyliCate offers a high-throughput method for gene duplicate array identification. It also facilitates downstream gene expression divergence analysis of the identified duplicates, enabling their fate prediction. DupyliCate was applied on the flavonoid synthase (FLS) gene family in Brassicales and subgroup 7 myeloblastosis (MYB) transcription factors (SG7 MYB) across a diverse range of plant species to understand the duplication dynamics of these important players in flavonoid biosynthesis. This helped uncover a potential radiation of FLS genes in the Brassicaceae and a deep duplication of the SG7 MYB lineage in some dicots. Further, DupyliCate was also used to identify gene duplications in other key genes of the flavonoid biosynthesis. A downstream expression analysis of these mined duplicates using the tool, combined with a systematic analysis of their promoter sequences helped ascertain the genetic triggers explaining the observed divergence and redundancy.
DupyliCate is available at: https://github.com/ShakNat/DupyliCate

C-G.10: Cracking the animal venom code using phylogenetic big data
Track: Genomics, epigenomics, and genome editing
  • Athina Gavriilidou, University of Lausanne, Switzerland
  • Giulia Zancolli, University of Lausanne, Switzerland
  • Christophe Dessimoz, University of Lausanne, Switzerland
  • Natasha Glover, Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

Venom systems are among nature's most sophisticated biological innovations, yet their evolution and functional integration remain incompletely understood. We conducted a large-scale comparative analysis of venom-related genes across animals using phylogenetic and genomic approaches.
We compiled a dataset of venom-associated protein families from public databases, spanning diverse taxa to capture broad evolutionary patterns. To identify coevolving proteins, we applied phylogenetic profiling, representing each gene as a binary vector encoding its presence or absence across species. By comparing these profiles, we detected protein families with similar evolutionary trajectories. Using the Metazoa dataset from the OMA database, which uses hierarchical orthologous groups (HOGs), we employed the HogProf algorithm to identify HOGs with high similarity to venom-related profiles. This enabled systematic detection of candidate genes potentially involved in venom function or self-resistance.
Preliminary results show that proteins within the same venom cocktail often share highly similar phylogenetic profiles, supporting coordinated evolution. We also observed a correlation between expression patterns and phylogenetic similarity, suggesting functional links. Additionally, we identified protein families with venom-like evolutionary signatures not currently associated with venom, highlighting candidates for resistance mechanisms.
A central aim is to understand how venomous species avoid self-intoxication through biochemical adaptations. Investigating these mechanisms may clarify the evolutionary interplay between toxin production and resistance, and provide broader insight into the evolution of toxin resistance across species.

C-G.11: MosCoverY: A method to estimate mosaic loss of Y chromosome from sequencing coverage data
Track: Genomics, epigenomics, and genome editing
  • Valeriia Timonina, School of Life Sciences, École Polytechnique Fédérale de Lausanne, Switzerland
  • Astrid Marchal, Imagine Institute, Université Paris Cité, France
  • Laurent Abel, Hôpital Necker Enfants Malades; Imagine Institute, Université Paris Cité, France
  • Aurélie Cobat, Hôpital Necker Enfants Malades; Imagine Institute, Université Paris Cité, France
  • Jacques Fellay, School of Life Sciences, École Polytechnique Fédérale de Lausanne, Switzerland


Presentation Overview: Show

Mosaic loss of the Y chromosome (mLOY) is the most common somatic genomic event in men, accumulating with age and associated with diverse health outcomes, including all-cause mortality, cardiovascular disease, and cancer. Despite its clinical and biological relevance, detection of mLOY relies on DNA genotyping arrays, limiting its assessment in large cohorts where only sequencing data are available. Here, we present MosCoverY, a computational method for estimating mLOY directly from exome or whole-genome sequencing data.
MosCoverY addresses the challenges of the Y chromosome's structure by restricting analysis to single-copy genes and normalizing their sequencing coverage against autosomal exons matched by GC content and length. This design yields a robust individual-level estimate of normalized chromosome Y coverage, from which a binary mLOY and a continuous cell fraction harbouring mLOY can be derived.
We validated MosCoverY in 212,062 male participants from the UK Biobank, benchmarking it against two established methods based on genotyping arrays and whole-genome sequencing. MosCoverY identified mLOY in 5.6% of men, showed a strong correlation with other methods, and comparable performance in replicated associations with age, tobacco smoking, all-cause mortality, and germline genetic loci, yielding the strongest effect estimates in several cases. We further demonstrated method robustness at reduced sequencing depth and in single-sample analyses without population-level data. Finally, applying MosCoverY to The Cancer Genome Atlas confirmed its utility for detecting variable mLOY in tumors.
MosCoverY expands the toolkit for somatic genomic research, enabling mLOY detection in the growing body of sequencing-based population and clinical cohorts.

C-G.12: Computational Detection of Methylated Bases in Bacteria: PacBio vs. ONT
Track: Genomics, epigenomics, and genome editing
  • Mohammad Umair, Brno University of Technology, Czechia
  • Marketa Jakubickova, Department of Biomedical Engineering, Brno University of Technology, Czechia
  • Iva Buchtikova, Brno University of Technology, Czechia
  • Matej Bezdicek, Masaryk University, Czechia
  • Stanislav Obruca, Brno University of Technology, Czechia
  • Helena Vitkova, Brno University of Technology, Czechia
  • Karel Sedlar, Department of Biomedical Engineering, Brno University of Technology, Czechia


Presentation Overview: Show

DNA methylation is a widespread epigenetic modification in bacterial genomes, playing key roles in processes like restriction-modification systems and gene expression regulation. The advent of long-read sequencing has enabled detection of these modifications directly from native DNA; However, cross-platform comparability of the DNA methylation profiles remains insufficiently understood. In this study, we performed a cross-platform comparison of bacterial DNA methylation detection using Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) sequencing, leveraging data from phylogenetically diverse bacterial strains.
using the Standardised Analysis Framework, we evaluated all three major prokaryotic DNA methylation types (6mA,4mC, and 5mC) across multiple analytical levels. These included site-level overlap, methylation fraction estimate, motif discovery, genomic distribution, and the functional categorisation of genes associated with methylated regions. Beyond assessing overlap in detected modified sites, this approach enabled us to examine whether both platforms capture consistent methylation patterns at broader genomic and functional scales. While PacBio and ONT revealed partially overlapping methylomes, their outputs were clearly non-identical. Concordance was highest for adenine methylation, whereas cytosine methylation showed greater platform dependency, particularly at the site level and in estimated modification fraction. Nevertheless, both technologies demonstrated similar trends in motif detection and preferred genomic and functional contexts.
Overall, our findings indicate that PacBio and ONT are not interchangeable, as each captures overlapping yet distinct features of the bacterial methylation landscape. This comparison provides a practical framework for selecting appropriate long-read sequencing strategies and highlights the value of integrative approaches for achieving a more comprehensive view of the bacterial genome.

C-G.13: A coverage-focused workflow for high-confidence structural variant calling in cattle without validated truth sets
Track: Genomics, epigenomics, and genome editing
  • Laura Dekker, Animal Genomics ETH Zürich, Switzerland
  • Alexander Leonard, Animal Genomics ETH Zürich, Switzerland
  • Hubert Pausch, Animal Genomics ETH Zürich, Switzerland


Presentation Overview: Show

Structural variants are variants in the genome that span 50 bases or more. Advances in long read sequencing make it increasingly possible to investigate structural variant diversity in large cohorts. In many non-model species, differentiating between true and erroneous structural variant calls is difficult due to the lack of a truth set. This project aims at establishing a workflow to obtain a high-confidence set of structural variant genotypes in cattle which does not rely on a truth set of validated variants. We aligned PacBio HiFi reads from 57 cattle that had at least 15x coverage to the Bos taurus ARS-UCD2.0 reference genome to identify structural variants using Sniffles2 and Sawfish. A set of increased confidence structural variants was constructed by standard filtering of variants with low quality scores, missing genotypes and consensus length below 45 bp. Additionally, genomic areas with significantly outlying coverage values were included in a ‘blacklist' of coordinates from which to exclude structural variant calls. Results show that the structural variants highlighted by the latter method overlap largely with structural variants filtered out by the set of initial filters but exclude an additional 200-500 structural variants depending on the caller. These findings suggest that incorporating coverage-based exclusion criteria can improve the reliability of structural variant datasets in the absence of validated truth sets.

C-G.14: Neural posterior estimation for population genetics
Track: Genomics, epigenomics, and genome editing
  • Jiseon Min, Institute of Ecology and Evolution, University of Oregon, United States
  • Yuxin Ning, Quantitative Biology Center (QBiC), university of Tübingen, Germany
  • Nathaniel Pope, Institute of Ecology and Evolution, University of Oregon, United States
  • Franz Baumdicker, Justus Liebig University Giessen, Germany
  • Andrew Kern, Institute of Ecology and Evolution, University of Oregon, United States


Presentation Overview: Show

Simulation-based inference methods are increasingly being used in population genetics due to their flexibility and ability to be applied in settings where likelihood-based methods are intractable. One of the best known such method is Approximate Bayesian Computation (ABC). However, its popularity is offset by its shortcomings which include computational expense and an unfortunate inability to efficiently fit models to high-dimensional summaries of the data. An alternative approach that solves these issues is supervised machine learning (ML), but ML methods generally do not yield Bayesian uncertainty estimates of the quantities they predict.

Here, we apply a recently introduced method, neural posterior estimation (NPE), that combines the best facets of ABC and supervised ML by training a neural network to estimate the posterior distribution of a population genetics model. We first compare neural posterior estimation with other inference methods for a variety of population genetic tasks and show that neural posterior estimators yield posterior distributions with high accuracy and efficiency. We compare learned posterior distributions given raw genotypes and various summary statistics as input data. Additionally, we apply neural posterior estimation for demographic inference for simple and more complex models to highlight its application, including an analysis of demographic history in Drosophila melanogaster. Finally, we provide a user-friendly Snakemake workflow that enables others to perform neural posterior estimation on their own genetic data.

C-G.15: Extending the megSAP pipeline for medical genetics to long-read sequencing
Track: Genomics, epigenomics, and genome editing
  • Marc Sturm, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen, Germany, Germany
  • Tobias Haack, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen, Germany, Germany
  • Stephan Ossowski, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen, Germany, Germany
  • Leon Schütz, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen, Germany, Germany


Presentation Overview: Show

megSAP is an open-source sequence data analysis pipeline for rare disease and oncology. Since its initial development in 2016, the pipeline has been continuously refined to incorporate advances in short-read analysis for Illumina sequencing data.

Since 2024, functionality has been progressively added to support long-read technologies from Oxford Nanopore Technologies and PacBio.
This includes technology-specific mapping and variant-calling tools, as well as support for methylation analysis and haplotype phasing based on long-read data.

Consistent with its short-read workflow, megSAP enables comprehensive detection of major variant classes, including single-nucleotide variants and small indels, copy number variants, structural variants, and repeat expansions.

The resulting variant lists are extensively annotated and optimized for downstream interpretation in a clinical genetics setting.

megSAP also supports multi-sample analyses, for example in trio settings. This functionality has been extended to allow the combined analysis of short-read and long-read variants for small variants and copy-number variants.
Support for mixed analyses of structural variants is not yet available.

megSAP is freely available at https://github.com/imgag/megSAP

C-G.16: MtDNA Variant and Heteroplasmy Profiling in Patients with COPD
Track: Genomics, epigenomics, and genome editing
  • Daria Borodko, Institute of General Pathology and Pathophysiology, Russia
  • Vasily Sukhorukov, Institute of General Pathology and Pathophysiology, Russia
  • Andrey Omelchenko, Institute of General Pathology and Pathophysiology, Russia


Presentation Overview: Show

Chronic obstructive pulmonary disease (COPD) is associated with systemic inflammation and mitochondrial impairment, yet the contribution of mitochondrial DNA (mtDNA) heteroplasmy to disease biology remains incompletely understood. Here, we analysed mtDNA variants in the blood of 13 COPD patients.
MtDNA sequencing was performed using the Oxford Nanopore R2C2 protocol. Reads were aligned with bwa mem, and heteroplasmy levels were quantified using mtDNA-Server2 in a Docker container.
Across all individuals, we identified two recurrent heteroplasmic variants, m.902G>C and m.912T>A, localised to the MT-RNR1 gene. These mutations exhibited mean VAFs of 0.0798 and 0.409, respectively, suggesting consistent low-to-intermediate heteroplasmic states across the cohort. MT-RNR1 encodes the mitochondrial 12S rRNA and has been associated with increased cytokine levels in senescent cells, a pattern we have observed in patient blood samples.
In addition, we detected three protein-coding variants: m.4732A>G, m.10609T>C, and m.12406G>A—mapping to MT-ND2, MT-ND4L, and MT-ND5, respectively. These variants exhibited partial co-occurrence patterns (pairwise Jaccard similarity 0.33-0.78). Functional annotation indicated that these substitutions affect OXPHOS Complex I genes and include non-synonymous changes with predicted moderate pathogenic potential.
Overall, our findings demonstrate a recurrent MT-RNR1 heteroplasmic signature across COPD patients and identify additional heteroplasmic protein-coding variants with variable co-occurrence patterns in OXPHOS genes. These results support a model in which mtDNA heteroplasmy may contribute to metabolic and mitochondrial dysfunction in COPD, warranting further investigation in larger cohorts and functional studies.
This study was supported by RSF grant #24-65-00027.

C-G.17: Not All Out-of-Distribution Is Created Equal: Reasoning Demand in Perturbation Prediction
Track: Genomics, epigenomics, and genome editing
  • Anna Kalygina, ETHZ, D-INFK, Switzerland
  • Alexander Theus, ETHZ, D-INFK, Switzerland
  • Marina Medina Esteban, ETHZ, D-INFK, Switzerland
  • Valentina Boeva, ETHZ, D-INFK, Switzerland


Presentation Overview: Show

Whether deep learning models truly outperform simple baselines for single-cell perturbation prediction remains an open question. We argue that a key variable is missing from this debate: the rule demand imposed by the evaluation task itself. Prediction tasks are not equally demanding - they differ in whether test split requires learning beyond simple, interpretable rules. When mean-based or additive baselines already explain most of the evaluable signal, even highly expressive architectures have little opportunity to demonstrate a meaningful advantage

We introduce Rule Demand (RD), a meta-metric that quantifies the fraction of calibrated signal left unexplained by simple baselines. Across six scenarios derived from the same synthetic gene regulatory network, we show that splits commonly presented as “hard” out-of-distribution settings (cell-type extrapolation, dose extrapolation, and epistatic knockout prediction) can be almost fully explained by mean-based or additive baselines (RD ≈ 0), whereas unseen-perturbation splits preserve substantial headroom (RD → 1). We find that this illusion of difficulty also appears in real perturbation datasets, including Norman19, Wessels23, and XAtlas-Orion. Rule Demand offers, for the first time, an explanatory axis for why published studies report inconsistent improvements over naive rules. Across all scenarios and datasets, model gains over simple baselines are strongly predicted by RD.

In this work, we present a calibrated benchmarking suite that extends the Dynamic Range Fraction (DRF) framework with signal-magnitude-aware normalization, more than 60 performance metrics, and systematic control diagnostics for both metric choice and dataset selection.

C-G.18: Automated cell type prediction in Mass Cytometry data using machine learning with interactive web-based annotation review
Track: Genomics, epigenomics, and genome editing
  • Gábor Beke, Institute of Molecular Biology, Slovak Academy of Sciences, Bratislava, Slovakia
  • Milan Hucko, Institute of Molecular Biology, Slovak Academy of Sciences, Bratislava, Slovakia
  • Lubos Klucar, Institute of Molecular Biology, Slovak Academy of Sciences, Bratislava, Slovakia
  • Dana Cholujova, Cancer Research Institute, Biomedical Research Center, Slovak Academy of Sciences, Bratislava, Slovakia
  • Jana Jakubikova, Cancer Research Institute, Biomedical Research Center, Slovak Academy of Sciences, Bratislava, Slovakia


Presentation Overview: Show

Mass cytometry (Cytometry by Time-Of-Flight - CyTOF), is an advanced technology widely used in immunology, cancer research, drug discovery and systems biology. Compared to traditional flow cytometry, which is limited by spectral overlap and the availability of fluorophores, mass cytometry uses heavy metal isotopes as labels. This enables the simultaneous measurement of a greater number of markers on individual cells, resulting in high-dimensional datasets. Analyzing CyTOF data demands robust and scalable computational approaches. Clustering algorithms such as SPADE exist, however incorporating new samples into existing workflows typically requires reprocessing the entire dataset. While tools such as Seurat's sketching workflow have addressed analogous scalability challenges in single-cell RNA sequencing, there is no equivalent framework currently available for CyTOF data. To address this limitation, we used our manually annotated SPADE results to train an XGBoost-based classifier capable of predicting cell populations in CyTOF data with accuracy 85% - 90%, improving upon our previous XGBoost model (~75% accuracy). This approach allows new samples to be annotated without reprocessing the original dataset, greatly reducing computational time. To visually inspect and validate the results, we developed a web-based visualization portal using Python and Flask. The portal allows researchers to explore predicted cell type annotations, compare them against t-SNE and UMAP projections, and manually review or overwrite classifications where necessary, ensuring the final annotations remain biologically accurate. This work was supported by grants APVV-23-0482 and APVV-24-0471.

C-G.19: Management of Large Variant Datasets
Track: Genomics, epigenomics, and genome editing
  • Mohamed Abouelhoda, KFSH&RC (King Faisal Specialist Hospital & Research Center), Saudi Arabia
  • Mohamed El-Kalioby, KACST-KFSHRC, Saudi Arabia
  • Saudi Arabia


Presentation Overview: Show

Currently, Next Generation Sequencing (NGS) has become a widely used technology to identify variations associated with the disease for research and diagnostic. The variant analysis workflow on NGS data yields text files in VCF format. This format is inefficient to query large cohorts. To solve this problem, one uses one of the following technologies: 1) ready-to-use variant management systems (e.g., GTRAC, GenomicsDB, Gemini); 2) native relational database management systems (e.g., MySQL) or 3) NoSQL database systems such as Clickhouse or MongoDB.

In this poster we compare the performance of different systems in these categories (GTRAC, GenomicsDB, Gemini, MySQL, and Clickhouse) and present best practices and recommendations.

We used 1000 Genome Project dataset (1092 VCFs), each includes ~39.7 million variants. Total size of the database is about 76 GB. We used a server with 24-CPUs, 128GB RAM. The query set included: 1D range query to look for variants in a genomic range (chr:start_pos-end_pos), and 2D range queries to search for variants in a range and in group of samples.

Results: GTRAC has the best compression, but it lacks major functions for frequent queries. Clickhouse and GenomicsDB have the least insertion time per sample. Gemini and MySQL require to build indices after populating the tables and after each insertion. All tools in general have acceptable query/retrieval time. MySQL has best query time due to best indexing, but its space consumption and the insertion time is prohibitive for huge datasets. Comparing Clickhouse, Gemini and GenomicsDB, we observe that Clickhouse performs slightly better.

C-G.20: DNA methylation changes in genomic regulatory blocks in head and neck cancer
Track: Genomics, epigenomics, and genome editing
  • Katarina Mandić, Ruđer Bošković Institute, Croatia
  • Anja Barešić, Ruđer Bošković Institute, Croatia


Presentation Overview: Show

Genomic regulatory blocks (GRBs) are large genomic domains enriched in highly conserved non-coding elements that control the expression of a single target gene, often across megabase distances. These regions are thought to act as long-range regulatory units, integrating structural and regulatory information to maintain precise expression patterns of the target gene. Target genes are usually transcription factors involved in embryonic development and differentiation, and their promoters often overlap extended CpG islands that extend into the gene body, suggesting that DNA methylation of CpGs may contribute to their regulation. However, the methylation landscape across entire GRBs remains poorly understood, particularly in cancer. Here, we investigate whether differential methylation within GRBs is associated with deregulation of their target genes in head and neck cancer patients compared with controls. By integrating regional methylation analysis with gene expression or target-gene annotation, we aim to determine whether altered CpG methylation in GRBs reflects disrupted long-range regulatory control in cancer. Understanding these patterns could provide new insight into the role of methylation in GRBs and how cancer alters the expression of target genes.

C-G.21: Decoding Heritable Breast Cancer Predisposition through Integrative Gene-Based Pathway Analysis
Track: Genomics, epigenomics, and genome editing
  • Shirel Schreiber, The Hebrew University of Jerusalem, Israel
  • Roei Zucker, hebrew university of jerusalem, Israel
  • Amos Stern, The Hebrew University of Jerusalem, Israel
  • Michal Linial, The Hebrew University of Jerusalem, Israel


Presentation Overview: Show

Heritable breast cancer (BC) risk is driven by well-established high-penetrance genes such as BRCA1, BRCA2, PALB2, and CHEK2, yet the contribution of numerous moderate- and low-penetrance genes remains incompletely understood. To address this gap, we implemented a gene-centric, integrative framework across large multi-ethnic genomic resources, including the UK Biobank (UKB) and FinnGen, combining complementary association strategies: genome-, transcriptome-, and proteome-wide association studies (GWAS, TWAS, and PWAS). By consolidating variant-level signals into gene-level evidence and rigorously filtering likely false positives, we defined a conservative set of 38 high-confidence BC predisposition genes. Notably, only approximately 50% of these genes are currently represented in clinical diagnostic/prognostic panels. The set includes established DNA repair genes alongside emerging candidates such as APOBEC3A, TNS1, and PEX14, each supported by independent lines of evidence. PWAS further highlighted genes with potential recessive effects that are typically overlooked by standard GWAS, underscoring the value of incorporating proteome-level information. In parallel, a complementary family-based strategy focusing on BC-affected pedigrees leveraged rare and ultra-rare variant analyses to delineate core biological pathways and identify shared low-frequency contributors. Replication across cohorts demonstrated robust and consistent signals in populations of European ancestry, with more limited transferability to other groups. Importantly, excluding individuals with non-BC cancers (10.6%) from controls in UKB (15.7k cases) preserved risk estimates. Overall, this study provides a stringent and interpretable framework for refining gene prioritization in heritable BC and guiding future functional and clinical investigation.

C-G.22: SynVar: extending variant query expansion to GA4GH-compliant variant normalisation for literature annotation
Track: Genomics, epigenomics, and genome editing
  • Anais Mottaz, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics, Switzerland
  • Emilie Pasche, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics, Switzerland
  • Alexandre Flament, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics, Switzerland
  • Luc Mottin, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics, Switzerland
  • Patrick Ruch, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

SynVar is the variant query expander behind Variomes, a high-recall biomedical literature search engine (variomes.sibils.org). SynVar recognises variants in free text as well as rsIDs and HGVS expressions across protein, coding and genomic levels, and automatically resolves implicit reference sequences through a gene name or chromosome. Unlike LitVar2, SynVar does not require pre-existing database entries for variant resolution and normalisation. It resolves variants using UniProt cross-reference maps, with sequence-level validation and coordinate mapping delegated to VariantValidator and protein back-translation to Mutalyzer.
Here, we extend SynVar beyond substitutions to deletions, insertions, duplications and delins, moving from query expansion to full variant normalisation for literature annotation. SynVar now also accepts SPDI, VCF, LRG and Ensembl inputs. Its normalisation mode returns canonical HGVS on GRCh38 and GRCh37, coding and protein HGVS on the MANE Select transcript, SPDI, VCF representation, dbSNP rsID, ClinGen CAID and GA4GH VRS Allele. For non-SNP variants, normalisation includes protein indel back-translation, 3'-shifting under HGVS rules, left-aligned VCF output, insertion-duplication equivalence and HGVS repeat notation.
We evaluated 59 pathogenic ClinVar variants selected for class coverage and rewritten into up to 13 input representations each (653 test cases). About 95% were parsed and 92-96% matched ClinVar across output fields. Mismatches arose from ambiguous protein-only frameshift inputs and non-canonical NCBI-only RefSeq isoforms not cross-referenced by UniProt. Integration into the SIBiLS annotation pipeline will enable evaluation on biomedical literature and identification of variants not yet catalogued in existing databases. SynVar is available at synvar.sibils.org.

C-G.23: A Bi-partite Graph Neural Network for Polygenic Risk Scoring under the Omnigenic Model
Track: Genomics, epigenomics, and genome editing
  • Ilaria Looser, Institute of AI for Health, Computational Health Center, Helmholtz Zentrum Muenchen, Germany
  • Sergey Vilov, Institute of Computational Biology, Computational Health Center, Helmholtz Zentrum Muenchen, Germany
  • Carsten Marr, Institute of AI for Health, Helmholtz Zentrum Muenchen / Department of Medicine III, LMU Hospital / DKTK, Germany
  • Matthias Heinig, Institute of Computational Biology, Helmholtz Zentrum Muenchen / Department of Computer Science, TUM / DZHK, Germany


Presentation Overview: Show

Genome-wide association studies (GWAS) identify single nucleotide polymorphisms (SNPs) that are correlated with complex traits by computing the marginal effect size of each SNP on the trait. Subsequently, by aggregating the individual SNP contributions into a single genetic risk estimate per individual, we can compute the polygenic risk score (PRS), serving a clinical utility especially for early disease screening.
Current PRS models explain only a fraction of trait heritability and provide limited mechanistic insight. While GWAS identifies SNP-trait associations, linking these signals to causal genes and pathways typically requires separate fine-mapping. Standard PRS approaches inherit this limitation by aggregating SNP effects without modelling relationships between variants and genes, resulting in a lack of integrated interpretability. In contrast, the omnigenic model suggests that genetic effects propagate through gene regulatory networks, where peripheral genes influence core disease pathways.
Here we introduce OmniGRS, a graph neural network that models PRS under the omnigenic hypothesis. OmniGRS constructs a per-individual bipartite graph connecting SNPs to genes and extends it with protein-protein interaction networks (STRING) to propagate regulatory context. We incorporate functional annotations (CADD) and SNP-gene relationships (e.g., genomic proximity or eQTLs) as node features and edge weights, enabling meaningful variant prioritisation.
OmniGRS outperforms baselines PRS models in both classification (AUC) and regression (Pearson correlation) tasks. Attention-weighted pooling over SNP and gene nodes yields interpretable importance scores that accurately recover causal variants in simulation.
Overall, OmniGRS provides a biologically informed framework for modelling non-additive genetic effects, improving risk prediction and interpretability in complex traits.

C-G.24: Domain-wide Mapping of Peer-reviewed Literature for Genetic Developmental Disorders using Machine Learning and Gene2Phenotype
Track: Genomics, epigenomics, and genome editing
  • Michael Yates, University of Edinburgh, United Kingdom
  • Sarah E Hunt, EMBL-EBI, United Kingdom
  • Diana Lemos, EMBL-EBI, United Kingdom
  • Seeta Ramaraju Pericherla, EMBL-EBI, United Kingdom
  • Elena Cibrian Uhalte, EMBL-EBI, United Kingdom
  • Morad Ansari, South East Scotland Genetic Service, United Kingdom
  • Louise Thompson, South East Scotland Genetic Service, United Kingdom
  • Caroline Wright, University of Exeter, United Kingdom
  • Helen V Firth, Addenbrooke's Hospital Cambridge University Hospitals, United Kingdom
  • Ian Simpson, University of Edinburgh, United Kingdom


Presentation Overview: Show

Genetically-determined developmental disorders (GDD) are rare conditions whose diagnosis increasingly depends on synthesis of dispersed genotype–phenotype evidence. Manual literature curation is labour-intensive and difficult to scale. We present an automated pipeline that identifies PubMed abstracts describing human case reports/series and maps them to GDD in Gene2Phenotype (G2P). LitDD combines a BERT abstract classifier to detect GDD-relevant abstracts and a cross-encoder to rank candidate diseases, each fine-tuned on 13,738 annotated title–abstract pairs, with a large language model for final mapping adjudication. LitDD BERT achieved precision 0.83 and recall 0.94; the cross-encoder achieved top-5 recall of 0.99; and the full ensemble precision 0.89 and recall 0.82. Applied PubMed-wide, the pipeline identified 69,200 manuscripts mapped to G2P gene-disease entities, with ~70% retrieval against independent manually curated datasets, supporting generalisation across diverse mechanisms, inheritance patterns, and curation standards. The LitDD corpus is incorporated into routine G2P biocuration and we provide examples where this has enabled rapid upgrade of diseases to clinically reportable status.

To evaluate clinical utility of LitDD-derived disease models, we integrated them into Exomiser phenotype-driven disease prioritisation and benchmarked against 7,406 DD phenopackets. Incorporating G2P literature-derived HPO terms improved diagnostic AUC from 0.891 to 0.919. G2P disease models outperformed baseline for 25% of diseases, particularly those lacking rich phenotype annotations. Diagnostic performance correlated with phenotype specificity (Spearman r=0.13, p<0.001).

The LitDD corpus is openly accessible via G2P. This enables scalable, updatable literature surveillance and faster diagnostic evidence review in genomic medicine.

C-G.25: AIOMICS4CARE WGS: AI enabled Annotation System for Whole Genome Sequencing Data
Track: Genomics, epigenomics, and genome editing
  • Hanin Omer, KFSH&RC, Saudi Arabia
  • Mohamed Bamajboor, KFSH&RC, Saudi Arabia
  • Wedad Albalawi, KFSH&RC, Saudi Arabia
  • Turki Alzahrani, KFSH&RC, Saudi Arabia
  • Azza Althagafi, KFSH&RC, Saudi Arabia
  • Monther Alhamdoosh, KFSH&RC, Saudi Arabia
  • Abdullah Alfalah, KFSH&RC, Saudi Arabia
  • Ahmad Alfares, KFSH&RC, Saudi Arabia
  • Mohamed Abouelhoda, KFSH&RC, Saudi Arabia


Presentation Overview: Show

AIOMICS4CARE WGS is a whole-genome sequencing (WGS) workflow augmented with two AI modules to improve diagnostic yield, accelerate reporting, and support timely clinical decision-making in a high-throughput genomic center processing approximately 400 samples per month, including urgent rapid-WGS cases with direct patient impact. The first module, pheno priori, integrates a genotype-derived KFSH score with the patient's clinical phenotype to produce a phenotype-informed prioritization score. The second module, ai_consilium, classifies variants into Primary, Secondary, or Carrier findings to streamline interpretation and reporting.
This workflow was designed to address several persistent challenges in clinical genomics, including limited population-specific allele frequency data, incomplete representation of local variation, and the risk of missing clinically important variants. To overcome these limitations, we incorporated a locally deployed WGS allele-frequency database, a curated local masked database, and public reference data from gnomAD. These resources were combined to generate a prioritized and manageable list of variants for scientist review, while preserving sensitivity for locally relevant findings.
The workflow was implemented in a reproducible and configurable environment using Conda. Retrospective evaluation of AI consilium on 899 samples demonstrated a 74% precision for Primary Findings when deployed on the locally hosted QUINN model.
Overall, this workflow introduces a practical and scalable AI-assisted framework for WGS interpretation that leverages local population knowledge, improves prioritization, and supports faster and more accurate genomic reporting in clinical settings.

C-G.26: Searching for bacterial capsule at scale
Track: Genomics, epigenomics, and genome editing
  • Matthew Russell, EMBL-EBI, United Kingdom
  • Teodora Mateeva, Wellcome Sanger Institute, United Kingdom
  • Samuel Horsfield, University of Neuchâtel, Switzerland
  • Stephanie Lo, EMBL-EBL, United Kingdom
  • Stephen Bentley, Wellcome Sanger Institute, United Kingdom
  • John Lees, EMBL-EBI, United Kingdom


Presentation Overview: Show

Capsular polysaccharides are key determinants of bacterial virulence, environmental persistence, and host immune evasion, yet their distribution across the bacterial tree remains incompletely characterised. This is largely due to the diversity of capsule biosynthesis systems and the limitations of existing annotation approaches when applied at scale. Here, we present a framework for identifying capsule-producing bacteria within large, uniformly assembled genome collections such as AllTheBacteria. We combine two complementary search strategies. First, we implement gene locus models based on sequence homology to known capsule biosynthesis clusters, extending concepts from tools like CapsuleFinder. Such models capsure conserved gene content and locus structure while remaining flexibile enough to account for much of the diversity within a capsule production system. Second, we incorporate structural homology searches for capsular proteins, enabling detection of functionally conserved components that may evade sequence-based methods due to high divergence. By integrating sequence- and structure-based evidence, we aim to improve sensitivity and robustness of capsule system detection across phylogenetically distant taxa. Projecting these results onto a marker-gene-based tree of all samples in AllTheBacteria, enables evolutionary analysis such as within-lineage capsule gain/loss and horizontal gene transfer. Our method enables consistent annotation of capsule biosynthesis production systems across millions of genomes. This provides a comprehensive view of capsule distribution, diversity, and evolutionary patterns. The framework is scalable, modular, and adaptable to expanding reference datasets, offering a deeper understanding of bacterial surface biology at ecosystem scale.

C-G.27: Benchmarking Pathogenicity Estimation Methods with Real-World Clinical Data in Inherited Heart Disease
Track: Genomics, epigenomics, and genome editing
  • Nooshin Bayat, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
  • Angel Bernabe Garcia, Cardiogenetics Lab, IMIB,Instituto Murciano de Investigación Biosanitaria, Murcia, Spain, Spain
  • Nuria Garci­a-Santa, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
  • Laura Martinez Gomez, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
  • Kristina Ibaanez, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
  • Abe Shuya, Fujitsu Laboratory, Fujitsu Ltd, Kawasaki, Japan, Japan
  • Fuji Masaru, Fujitsu Laboratory, Fujitsu Ltd, Kawasaki, Japan, Japan
  • Raul Valin, Fujitsu Research Of Europe Limited Sucursal En España, Madrid, Spain, Spain
  • Juan Ramon Gimeno, Inherited Cardiomyopathy Unit. CSUR. ERN. Hospital Virgen de la Arrixaca, Murcia, Spain, Spain
  • Maria Sabater Molina, Legal Medicine Department, University of Murcia, Murcia, Spain., Spain


Presentation Overview: Show

Reliable classification of genetic variants, especially variants of uncertain significance (VUS), remains a major bottleneck in clinical genomics. Pathogenicity estimation is a key step in variant classification; however, most prediction tools are limited to specific variant types and lack full coverage.
FAI is a LightGBM-based predictor trained on ClinVar data (Oct 2023). It integrates variant-level features and neighbouring variant information for binary classification.
We analysed 815 variants from inherited heart disease cohort (392 Pathogenic, 70 Benign, 329 VUS, 24 None) and systematically benchmarked FAI against leading dbNSFP meta-predictors (BayesDel, MetaRNN, ClinPred, AlphaMissense, REVEL, CADD) using balanced accuracy and coverage. Against high-confidence ClinVar variants (>2 stars, Mar 2026, n=157), FAI achieved the best overall performance (~100% coverage, ~0.98 balanced accuracy). In a real-world setting using hospital classifications (n=462), performance decreased across all methods; however, FAI maintained the most favourable balance (~100% coverage, ~0.80 balanced accuracy).
Re-analysis of clinically challenging cases demonstrated translational value. FAI showed 86% concordance with ClinGen classifications. Independently, FAI-driven reclassification, when consistent with ClinVar evidence, downgraded five hospital-classified pathogenic variants to benign, reclassified three VUS as pathogenic, and proposed pathogenic classification for three previously unannotated variants with clinical impact, as confirmed by expert reconsideration. Discrepancies remained in 24% of cases, mostly due to clinical evidence.
These results demonstrate that benchmarking must consider performance and coverage to ensure generalizability. Coverage gaps may arise from missing annotations in resources such as dbNSFP, potentially biasing comparisons. Importantly, variant interpretation is disease-specific and requires integration of clinical context to ensure accurate classification.

C-G.28: Identification of conserved gene clusters across diverse prokaryotic pangenomes
Track: Genomics, epigenomics, and genome editing
  • Ikuo Uchiyama, National Institute for Basic Biology, National Institutes of Natural Sciences, Japan


Presentation Overview: Show

Prokaryotes, including bacteria and archaea, possess a diverse collective gene repertoire despite their relatively small individual genome sizes. This diversity is evident even within species; for instance, a pangenome can be substantially larger than any single genome. Horizontal gene transfer (HGT) is a primary mechanism for maintaining such diverse pangenomes, enabling adaptation to various environments. In this study, we identified conserved gene orders within various pangenomes and compared them to identify conserved gene clusters (CGCs) across distantly related species. We classified conserved gene orders into two types: "core genomes," where gene orders are conserved across almost all strains, and "islands," which are conserved in only a limited number of strains. Using the Microbial Genome Database for Comparative Analysis (MBGD), which organizes microbial genomes into hierarchical ortholog groups (species, genus, and top levels), we employed CoreAligner to identify core genomes. A modified version of this algorithm was then applied to the remaining genomic regions to identify islands. After representing these regions as strings of top-level ortholog identifiers, we used HomologyTeams to identify gene clusters in which all adjacent gene pairs are located within a specified interval across different species, followed by a clustering algorithm to define CGCs. This pipeline was applied to 488 species (each having at least six different strains in MBGD), resulting in the identification of over 1,000 CGCs. We are currently analyzing these CGCs, focusing on their core/non-core status across species as potential indicators of their propagation, including HGT events.

C-G.29: Epigenomic Instability - Is Epigenetic Age Acceleration a Survival Indicator in Lung Cancer?
Track: Genomics, epigenomics, and genome editing
  • Alexandra Anke Baumann, Department of Systems Biology and Bioinformatics, University of Rostock, Germany
  • Michael Seifert, Institute for Medical Informatics and Biometry (IMB), TUD Dresden University of Technology, Germany
  • Zholdas Buribayev, Department of Computer Science, Faculty of Information Technologies, Al-Farabi Kazakh National University, Kazakhstan
  • Olaf Wolkenhauer, Department of Systems Biology and Bioinformatics, University of Rostock, Germany
  • Markus Wolfien, Institute for Medical Informatics and Biometry (IMB), TUD Dresden University of Technology, Germany


Presentation Overview: Show

Cancer genomes are characterized by a considerable amount of modifications on various multi-omics levels, often caused by genomic and epigenomic instability. This cancer hallmark represents, e.g., epigenetic alterations in DNA methylation, histone remodeling, and non-coding RNA regulation. Understanding these patterns can give insights into tumor evolution, heterogeneity, and therapeutic resistance. However, it requires both robust data infrastructure and mechanistic insight.
To address the computational challenge of handling large-scale cancer genomic data, we developed TCGADownloadHelper. This streamlined pipeline simplifies data retrieval from The Cancer Genome Atlas (TCGA) via the GDC portal. Thus, transparent preprocessing of multimodal, patient-linked datasets is possible. This tool formed the foundation for our downstream analyses of DNA methylation data across TCGA lung cancer cohorts.
Building on these resources, we investigated epigenetic age acceleration (EAA) across the lung cancer cohorts TCGA-LUAD and TCGA-LUSC. Illumina 450K DNA methylation data was obtained from the GDC portal via the TCGADownloadHelper. The deviation between biological (epigenetic) and chronological age was estimated from DNA methylation patterns using established epigenetic clocks (Horvath, Zhang2019,...). We demonstrate that approximately two thirds of tumors exhibited accelerated epigenetic aging, while one third showed a ""rejuvenating"" effect. Interestingly, patients with rejuvenating EAA showed significantly lower overall survival in Kaplan-Meier analyses. EAA varied by sex, tumor stage, and cancer type, highlighting the importance of patient-specific factors in the epigenetic age.
Together, these findings introduce epigenetic aging signatures as a promising stratification biomarker and motivate joined mechanistic studies on genomic and epigenomic instability to characterize actionable targets in precision oncology.

C-G.30: Unveiling the functional fate of duplicated genes through expression profiling and structural analysis
Track: Genomics, epigenomics, and genome editing
  • Alex Warwick Vesztrocy, BioSoft Research UK, United Kingdom
  • Natasha Glover, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
  • Paul D Thomas, University of Southern California, SIB Swiss Institute of Bioinformatics, United States
  • Christophe Dessimoz, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
  • Irene Julca, Aarhus univeristy, Denmark


Presentation Overview: Show

Gene duplication is a major evolutionary source of functional innovation. Following duplication events, gene copies (paralogs) may undergo various fates, including retention with functional modifications (such as subfunctionalization or neofunctionalization) or loss. When paralogs are retained, this results in complex orthology relationships, including one-to-many or many-to-many. In such cases, determining which one-to-one pair is more likely to have conserved functions can be challenging. It has been proposed that, following gene duplication, the copy that diverges more slowly in sequence is more likely to maintain the ancestral function referred to here as the least diverged ortholog (LDO) conjecture. This study explores this conjecture, using a novel method to identify asymmetric evolution of paralogs and applying it to all gene families across the Tree of Life in the PANTHER database. Structural data for over 1 million proteins and expression data for 16 animals and 20 plants are used to investigate functional divergence following duplication. This analysis, the most comprehensive to date, reveals that, whereas the majority of paralogs display similar rates of sequence evolution, significant differences in branch lengths following gene duplication can be correlated with functional divergence. Overall, the results support the least diverged ortholog conjecture, suggesting that the least diverged ortholog tends to retain the ancestral function, whereas the most diverged ortholog (MDO) may acquire a new, potentially specialized role.

C-G.31: Integrative gene network analysis of genome-wide association data in myalgic encephalomyelitis / chronic fatigue syndrome
Track: Genomics, epigenomics, and genome editing
  • Ekaterina Antonenko, CBIO Mines Paris PSL, France
  • Giann Karlo Aguirre-Samboní, CBIO Mines Paris PSL, France
  • Florian Massip, CBIO Mines Paris PSL, France
  • Chloé-Agathe Azencott, CBIO Mines Paris PSL, France


Presentation Overview: Show

Myalgic encephalomyelitis / chronic fatigue syndrome (ME/CFS) is a common though poorly understood disease affecting millions of people worldwide. The biological mechanisms underlying ME/CFS remain largely unclear, no effective treatments currently exist, and the disease disproportionately affects females and is frequently triggered by acute infection. However, no satisfactory mechanistic explanation for either factor has been established.

In the DecodeME study, the first genome-wide association studies (GWAS) were performed on a large cohort of cases (15,579) and controls (259,909) with European genetic ancestry. Eight loci were reported to be significantly associated with ME/CFS, three of which are proximate to genes involved in the response to viral or bacterial infection, consistent with the known infection trigger. The initial findings also suggest that both immunological and neurological processes contribute to the genetic risk of ME/CFS.

However, GWAS is by design limited to individual SNP-phenotype associations, and gene-gene interactions are largely overlooked by this framework. Polygenic diseases such as ME/CFS are likely shaped by the coordinated activity of multiple genes within shared biological pathways, rather than by isolated variants alone. Gene network methods integrating pathway and interaction data have therefore been developed to refine and enrich classical GWAS signals, and combining multiple such methods has been shown to improve statistical power and interpretability, with successful applications in breast cancer and psoriasis.

In the present study, we re-analyse the DecodeME GWAS summary statistics and apply a combination of gene network methods across curated pathway databases and experimentally derived protein interaction networks. This integrative approach yields a robust consensus of genes potentially involved in ME/CFS pathogenesis. The network analysis partly recovers the genes from the original study and additionally identifies multiple previously unreported pathways and candidate genes related to immune regulation and neurological function.

C-G.32: Evolutionary Conservation and Functional Constraints of TP53 Mutation Hotspots Across Mammalian Species
Track: Genomics, epigenomics, and genome editing
  • Ritika Rawat, University of Mumbai, India
  • Sermarani Nadar, University of Mumbai, India
  • Gursimran Kaur Uppal, University of Mumbai, India


Presentation Overview: Show

The tumor suppressor gene TP53 plays a central role in maintaining genomic stability and is one of the most frequently mutated genes in human cancers. Investigating the evolutionary conservation of TP53 mutation hotspots across species can provide insights into their functional importance and selective constraints.

In this study, we performed a comparative genomic analysis of TP53 across multiple mammalian species to identify conserved regions and mutation hotspots. TP53 sequences were retrieved from publicly available databases (NCBI) and subjected to multiple sequence alignment and phylogenetic analysis. Conservation scoring was used to identify functionally constrained regions, and known human mutation hotspots were mapped onto conserved domains.

Our analysis reveals that several mutation hotspots in human TP53 coincide with highly conserved regions across mammals, indicating strong evolutionary pressure to maintain their functional integrity. These regions predominantly correspond to DNA-binding domains essential for transcriptional regulation and tumor suppression. In contrast, less conserved regions exhibit variability suggestive of species-specific adaptations.

Overall, this study highlights the evolutionary significance of TP53 mutation hotspots and provides insights into their functional constraints across species. These findings contribute to a better understanding of cancer-associated mutations within an evolutionary framework.

C-G.33: PanTEon: a cross-kingdom framework to guide the design of transposable element classifiers
Track: Genomics, epigenomics, and genome editing
  • Simon Orozco-Arias, Life Science Department, Barcelona Supercomputing Center, BSC-CNS, 08032 Barcelona, Spain, Spain
  • Iamil Ferrer-Pomer, Universitat Oberta de Catalunya, 08018 Barcelona, Spain, Spain
  • Fabiana Rodrigues de Goes, The Rosalind Franklin Institute, OX11 0QX Didcot, United Kingdom, United Kingdom
  • Simon Gaviria-Orrego, Department of Computer Science, Universidad Autónoma de Manizales, 170001 Manizales, Colombia, Colombia
  • Juan Gómiz-Fernández, Universitat Oberta de Catalunya, 08018 Barcelona, Spain, Spain
  • Jordi Llatser-Torres, Universitat Oberta de Catalunya, 08018 Barcelona, Spain, Spain
  • Alexandre R. Paschoal, The Rosalind Franklin Institute, OX11 0QX Didcot, United Kingdom, United Kingdom
  • Romain Guyot, UMR DIADE, IRD, CIRAD, Université de Montpellier, Montpellier, France, France
  • Toni Gabaldón, Life Science Department, Barcelona Supercomputing Center, BSC-CNS, 08032 Barcelona, Spain, Spain


Presentation Overview: Show

Transposable elements (TEs) are fundamental drivers of genome evolution, yet their annotation and classification remain inconsistent, fragmented, and difficult to reproduce across species. This challenge arises from sequence divergence, lineage-specific innovations, and heterogeneous taxonomies across databases and computational tools, ultimately limiting large-scale comparative analyses. Here, we present PanTEon, a cross-kingdom deep learning framework designed to enable reproducible, scalable, and standardized TE classification. PanTEon integrates two key components: (i) the PanTEon Database, an automatically curated repository comprising ~240,000 structurally validated TE sequences from 2,790 species across animals, plants, and fungi, and (ii) a modular benchmarking and training platform that supports parallel evaluation and deployment of multiple ML and DL models. The PanTEon framework enables unified training, inference, and comparison across nine ML/DL architectures, while remaining fully extensible to user-defined models. Using this standardized environment, we benchmark seven state-of-the-art TE classifiers and demonstrate that classification performance is strongly influenced by taxonomic origin and TE superfamily, revealing significant biases and limitations in current approaches. Furthermore, we leverage the PanTEon training module to systematically retrain and compare nine architectures on a harmonized task involving 30 TE superfamilies. Our results show that ensemble strategies and taxon-specific models substantially improve predictive performance, while cross-kingdom generalization remains a key unresolved challenge. Additionally, we demonstrate the flexibility of the framework by addressing auxiliary tasks such as false positive detection in TE libraries. Overall, PanTEon establishes the first cross-kingdom, deep learning-driven ecosystem for TE classification, providing a robust foundation for benchmarking, model development, and large-scale genomic analyses.

C-G.34: DeepCAST-GWAS: Improving the Discovery of Genetic Associations Using Deep Learning-Based Regulatory SNP Prioritization
Track: Genomics, epigenomics, and genome editing
  • Lovro Rabuzin, ETH Zurich, Switzerland
  • Konstantin Heep, ETH Zurich, Switzerland
  • Sophie Sigfstead, University of Alberta, Canada
  • Valentina Boeva, ETH Zurich, Switzerland


Presentation Overview: Show

Genome-wide association studies (GWAS) have uncovered numerous variants linked to complex traits, yet power remains limited by the large multiple testing burden and the inclusion of many variants with minimal regulatory impact. We present Deep learning-based Chromatin Accessibility SNP Targeting for GWAS (DeepCAST-GWAS), a framework that integrates functional annotations derived from deep learning models to improve both the yield and the reliability of GWAS findings. DeepCAST-GWAS uses SNP Activity Difference (SAD) scores from in silico mutagenesis with the Enformer model to estimate the predicted effect of each variant on chromatin accessibility across tissues, allowing statistical testing to focus on variants with stronger regulatory evidence. Using conservative family-wise error rate (FWER) control, DeepCAST-FWER produces fewer associations than existing power-boosting approaches, but the associations it reports replicate in larger cohort GWAS at substantially higher rates. For applications where discovery count is more important, DeepCAST-sFDR increases the number of genome-wide significant findings above baseline GWAS by using the Enformer SAD scores for stratified False Discovery Rate (sFDR) control. DeepCAST-sFDR achieves performance comparable to the strongest competing method, while maintaining reliability on par with a standard GWAS. Subsampling analyses across a wide range of traits confirm these improvements in both sensitivity and replicability. DeepCAST-GWAS offers a principled way to incorporate sequence-based regulatory predictions into population-scale association testing, demonstrating that chromatin accessibility activity scores can improve the stability of GWAS discoveries.

C-G.35: Causal thinking in a correlated world: improving genomic prediction through SNP selection in the AI era
Track: Genomics, epigenomics, and genome editing
  • Thomas Crow, The University of Queensland, ARC CoE for Plant Success in Nature and Agriculture, Australia
  • David Kainer, The University of Queensland, ARC CoE for Plant Success in Nature and Agriculture, Australia


Presentation Overview: Show

Many plant breeding programs have embraced genomic prediction to accelerate crop improvement. A central challenge in genomic prediction is the number of genetic markers (eg. SNPs) greatly exceed the number of phenotyped individuals, reducing accuracy.
Dimensional reduction methods like linkage disequilibrium and minor allele frequency filtering can reduce this imbalance, but they may filter out informative SNPs.
Clearly, our genomic prediction models would benefit from selecting only SNPs which contribute to phenotype variation. Such causal SNPs would be invaluable, but the true genetic architecture of a trait is typically unknown.
To what extent can informed SNP selection identify or approximate these causal SNPs and how does doing so affect genomic prediction accuracy?
This talk presents a conceptual framework for SNP selection in genomic prediction, discussing how different strategies attempt to filter for causal SNPs. We examine existing approaches spanning statistical selection using Bayesian shrinkage and machine learning models, as well as biological selection methods informed by GWAS signals, gene expression, functional annotations, and variant effect prediction.
We also discuss emerging directions in AI-augmented SNP selection, where DNA foundation models, large language models, and graph-based methods can be used to integrate many biological studies into a single SNP selection model. Finally, we present a new approach which uses biological knowledge networks for SNP selection.
Genomic prediction accuracy will not improve by simply adding more markers, but by using SNP selection to better reflect the underlying biology of traits.

C-G.36: Managing workflow executions with WESkit
Track: Genomics, epigenomics, and genome editing
  • Landfried Kraatz, Berlin Institute of Health at Charité Universitätsmedizin Berlin, Germany
  • Valentin Schneider-Lunitz, Berlin Institute of Health at Charité Universitätsmedizin Berlin, Germany
  • Sven Twardziok, Berlin Institute of Health at Charité Universitätsmedizin Berlin, Germany


Presentation Overview: Show

Managing computational workflows across diverse biomedical projects, each with its own parameters, tools, and execution environments, poses persistent challenges for scalability, reproducibility, and collaborative research. We introduce WESkit, a robust implementation of the Global Alliance for Genomics and Health (GA4GH) Workflow Execution Service (WES) specification that unifies the execution, monitoring, and documentation of data‑processing workflows. By supporting both Snakemake and Nextflow, WESkit enables consistent automation and centralized oversight across large numbers of heterogeneous workflow runs. This design empowers research groups and service units to maintain long‑term reproducibility, streamline multi‑project operations, and scale computational efforts with confidence. Seamless integration with cloud infrastructures further positions WESkit as a practical contributor to the GA4GH cloud ecosystem and a valuable tool for modern, collaborative biomedical data analysis.

C-G.37: Quantitative Modeling of Clone-Specific Treatment Resistance via Joint Bayesian Inference of Compositional and Population Size Data
Track: Genomics, epigenomics, and genome editing
  • Mohammad Darbalaei, Bioinformatics and Computational Biophysics, University of Duisburg-Essen, Essen, Germany, Germany
  • Julia Zummack, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis, Univ. Duisburg-Essen, Germany, Germany
  • Thomas Mühlenberg, German Cancer Consortium (DKTK), partner site Essen/Düsseldorf, DKFZ and Univ. Duisburg-Essen, Germany, Germany
  • Philip Dujardin, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis, Univ. Duisburg-Essen, Germany, Germany
  • Susanne Grunewald, German Cancer Consortium (DKTK), partner site Essen/Düsseldorf, DKFZ and Univ. Duisburg-Essen, Germany, Germany
  • Patricia Munteanu, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis, Univ. Duisburg-Essen, Germany, Germany
  • Marina Martinez Cruz, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis, Univ. Duisburg-Essen, Germany, Germany
  • Madeleine Dorsch, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis, Univ. Duisburg-Essen, Germany, Germany
  • Alexander Schramm, West German Cancer Center, Dept. of Medical Oncology, Molecular Oncology, Univ. Duisburg-Essen, Germany, Germany
  • Sebastian Bauer, German Cancer Consortium (DKTK), partner site Essen/Düsseldorf, DKFZ and Univ. Duisburg-Essen, Germany, Germany
  • Barbara M. Grüner, West German Cancer Center, Dept. of Medical Oncology, Cell Plasticity and Metastasis, Univ. Duisburg-Essen, Germany, Germany
  • Daniel Hoffmann, Bioinformatics and Computational Biophysics, University of Duisburg-Essen, Essen, Germany, Germany


Presentation Overview: Show

Multiplexed assays such as DNA-barcoded cell line mixtures generate compositional sequencing data that capture only relative clone abundances. This inherent constraint complicates the inference of absolute, clone-specific treatment effects, as changes in relative abundance do not necessarily reflect true growth behavior.
Here, we develop a hierarchical Bayesian modeling framework to infer clone-specific treatment responses by jointly integrating compositional sequencing data with independent measurements of total population size, such as confluency in vitro or tumor volume in vivo. Barcode count data are modeled using a Dirichlet–multinomial likelihood to capture sampling variability and biological overdispersion. Tumor volume and confluency measurements are modeled using log-normal and beta likelihoods, respectively, enabling inference of absolute growth. By combining these data sources, the framework reconstructs clone-specific changes in absolute abundance under treatment relative to control. A hierarchical structure captures variability across replicates and treatment conditions, enabling partial pooling and stabilizing estimates in settings with limited observations. Inference is performed in a fully probabilistic manner, allowing uncertainty from both sequencing and population-level measurements to be propagated to the final estimates. This yields a quantitative measure of treatment response for each clone with associated uncertainty, resolving ambiguities inherent to compositional data.
We validated the method by comparison of experimental data from treatment responses in cancer cell line mixtures and individual cell lines. This approach provides a general framework for multiplexed drug screening assays and reduces the number of required animal experiments through pooled experimental designs.

C-G.38: PansimNuc: Rapid nucleotide-level simulation of selection and genetic element mobility in eukaryote pangenomes
Track: Genomics, epigenomics, and genome editing
  • Samuel Horsfield, University of Neuchâtel, Switzerland
  • Tobias Baril, University of Neuchâtel, Switzerland
  • Jigisha Jigisha, University of Neuchâtel, Switzerland
  • Daniel Croll, University of Neuchâtel, Switzerland


Presentation Overview: Show

Many eukaryotic species possess huge inter-individual genomic diversity, known as a "pangenome", made up of small variants at a single nucleotide level, and large variants consisting of gain or loss of entire chromosomes. Pangenome diversity drives specie's adaptability and evolution, however, the evolutionary mechanisms structuring pangenomes remain unknown.

To understand how pangenomes arise and what mechanisms underpin their persistence, we developed PansimNuc, a rapid nucleotide-level population simulator, written in Rust. PansimNuc simulates small variation at the level of single nucleotide polymorphisms and indels, and large scale chromosomal rearrangements through genetic element mobility, including transposable element (TE) dynamics and inter-genome recombination. PansimNuc also simulates demographic dynamics, including splitting and admixture of populations, and migration events. Finally, PansimNuc simulates selection at a nucleotide-level, enabling the impact of fitness effects in many complex evolutionary scenarios to be explored.

We show simulated populations generated by PansimNuc can reproduce complex dynamics including joint effects of genetic drift, migration and selective sweeps. Additionally, PansimNuc can recapitulate genome-wide TE copy number expansions expected if active TEs are present in a population. Tracking of haplotype characteristics over time across a variety of evolutionary scenarios provides a powerful tool to explore drivers of eukaryotic pangenome diversity at high resolution.

We foresee that PansimNuc has a large range of applications, including genome annotation tool benchmarking and training of statistical models, such as machine-learning and deep-learning approaches, to identify complex genome dynamics in genome sequence data.

C-G.39: Expanding Gene Ontology coverage of the human functionome in the PAN-GO project
Track: Genomics, epigenomics, and genome editing
  • Marc Feuermann, SIB Swiss Institute for Bioinformatics, Geneva, Switzerland
  • Huaiyu Mi, Department of Population and Public Health Sciences, University of Southern California, Los Angeles CA, United States
  • Li Ni, The Jackson Laboratory for Mammalian Genomics, Bar Harbor, ME, United States
  • Pascale Gaudet, SIB Swiss Institute for Bioinformatics, Geneva, Switzerland
  • Anushya Muruganujan, Department of Population and Public Health Sciences, University of Southern California, Los Angeles CA, United States
  • Dustin Ebert, Department of Population and Public Health Sciences, University of Southern California, Los Angeles CA, United States
  • Tremayne Mushayahama, Department of Population and Public Health Sciences, University of Southern California, Los Angeles CA, United States
  • Paul D. Thomas, SIB Swiss Institute for Bioinformatics, Geneva / University of Southern California, Los Angeles CA, Switzerland


Presentation Overview: Show

Understanding the functions of protein-coding genes in the human genome has long been a central goal of biomedical research. Over the past two decades, significant progress has been made, driven by improved genome annotation and major advances in the experimental characterization of human genes and their homologs in well-studied model organisms. Within the Gene Ontology knowledgebase, functional information has been systematically integrated using an evolutionary framework based on phylogenetic trees. This approach enabled the construction of curated models of function evolution and supported the initial assignment (version 1.0, February 2025) of at least one functional characteristic to 82% of human protein-coding genes. Importantly, each annotation can be traced back to experimental evidence from human and/or non-human model systems.
Here, we describe the second phase of this effort, which focuses on improving existing annotations and expanding coverage by addressing annotation gaps through both manual curation and AI-guided approaches. In version 2.0 of the PAN-GO functionome (https://functionome.geneontology.org/), coverage of human protein-coding genes has increased to 84% (an increase of 400 genes compared to version 1.0). Additionally, annotations were improved for thousands of genes, and 47% of genes now have Gene Ontology annotations across all three aspects - Molecular Function (MF), Biological Process (BP), and Cellular Component (CC) - representing a 5% increase (1000 genes) compared to the previous version. Further improvements have already been made in preparation for a third version.

C-G.40: From DMPs to DMRs: CMEnt, Characterization of methylation using positional entanglement
Track: Genomics, epigenomics, and genome editing
  • Vasileios Lemonidis, UAntwerp - KU Leuven, Belgium
  • Ken Op de Beek, UAntwerp, Belgium
  • Joris Vermeesch, KU Leuven - UZ Leuven, Belgium
  • Timon Vandamme, UZ Antwerp, Belgium
  • Guy Van Camp, UAntwerp, Belgium
  • Joe Ibrahim, UAntwerp, Belgium


Presentation Overview: Show

DNA Cytosine methylation is a naturally occurring point modification closely linked to chromatin state. Recruitment of de novo methyltransferases is facilitated by histone modifications, while DNA methylation itself can promote heterochromatin formation. This reciprocal relationship creates structured local dependencies among methylated cytosines, including correlation patterns that may inform the identification of differentially methylated regions (DMRs). Here, we leverage this property in a model-light framework for regional methylation analysis. We present CMEnt (Characterization of Methylation through positional Entanglement), an R package for statistically guided DMR identification from significant differentially methylated positions (DMPs). CMEnt starts from user-defined seed loci, typically DMPs, and tests within-group correlation among proximal seeds to assemble them into connected clusters. These clusters are then extended to neighbouring significantly correlated cytosines and merged when overlapping, yielding candidate DMRs. This design provides a coherent transition from DMP-level discoveries to region-level interpretation, enabling linked downstream analyses at both resolutions. CMEnt supports multiple input formats, including methylation microarray beta-value matrices indexed by genomic position, BED-like genomic coordinate files, and BSseq objects generated from short- or long-read sequencing assays such as whole-genome bisulfite sequencing (WGBS) and Oxford Nanopore Technologies (ONT). Across real-world datasets, CMEnt produces highly specific regions, while simulations show strong agreement with known ground truth when benchmarked against established DMR detection method. CMEnt therefore provides a competitive and complementary strategy for differential methylation analysis across diverse experimental platforms.

C-G.41: Computel 2.0: A Scalable Platform for Comprehensive Telomere Profiling and TRV Phenotyping
Track: Genomics, epigenomics, and genome editing
  • Davit Tarverdyan, Armenian Bioinformatics Institute, Institue of Molecular Biology NAS, Armenia
  • Anahit Yeghiazaryan, Armenian Bioinformatics Institute, Institue of Molecular Biology NAS, Armenia
  • Tatevik Jalatyan, Armenian Bioinformatics Institute, Institue of Molecular Biology NAS, Armenia
  • Lilit Nersisyan, Armenian Bioinformatics Institute, Armenia


Presentation Overview: Show

Telomeres cap chromosome ends, and their shortening and dysfunction drive genome instability in aging and cancer. Telomere length, repeat-variant composition (Telomeric Repeat Variants, TRVs), and fusion events are informative readouts of telomere state, yet existing tools resolve only parts of it: Mean Telomere Length (MTL) estimates assume genome-wide average coverage reflects a normal diploid genome, an assumption that collapses under the copy number alterations (CNAs) and aneuploidy pervasive in tumors, while TRV detection is limited to simple substitutions, fusions are rarely detected, and these capabilities stay split across disconnected tools.

To address these limitations, we present Computel 2.0, a unified, Python-based workflow for comprehensive telomere analysis. This updated platform systematically integrates CNA-aware MTL estimation, the profiling of TRVs, and the detection of telomeric fusions directly from short-read sequencing data (FASTQ/BAM). Computel 2.0 introduces a dynamic normalization algorithm that adjusts for coverage distortions, ensuring highly accurate MTL calculations even in highly aberrant cancer genomes. Furthermore, the platform achieves advanced TRV phenotyping by extracting a wider range of TRVs, including insertions and deletions.

We have applied Computel 2.0 across diverse sequencing modalities, demonstrating its scalability and precision on bulk, single-cell, and cell-free DNA (cfDNA) datasets. To streamline data exploration, the tool generates interactive HTML reports visualizing MTL distributions, TRV composition, and read-specific patterns. By effectively resolving telomere dynamics and characterizing TRV phenotypes, Computel 2.0 serves as a high-precision, scalable platform for investigating telomere biology.

C-G.42: From ageing clocks to human digital twins in personalising healthcare through biological age analysis
Track: Genomics, epigenomics, and genome editing
  • Murih Pusparum, Hasselt University and VITO NV, Belgium
  • Gokhan Ertaylan, VITO NV, Belgium
  • Olivier Thas, Hasselt University, Belgium
  • Simone Ecker, UCL Cancer Institute, University College London, London, United Kingdom
  • Stephan Beck, UCL Cancer Institute, University College London, London, United Kingdom


Presentation Overview: Show

Age is the most important risk factor for the majority human diseases, yet chronological age alone poorly captures the biological diversity shaping individual health trajectories. Biological age (BA) predictors derived from epigenomic, proteomic, metabolomic, and clinical biochemistry data offer a promising way to quantify ageing as a dynamic, measurable process. In this study, we demonstrate the value of BA analysis within the IAM Frontier cohort, a deeply phenotyped 13-month longitudinal study of 30 healthy adults aged 45–59 years. Across repeated time points, we computed BA and health-related predictions using 29 epigenetic, 4 clinical-biochemistry, 2 proteomic, and 3 metabolomic clocks, alongside gold-standard clinical risk indicators.
Our findings show that ageing signatures differ markedly between individuals while remaining relatively stable within individuals, with epigenetic clocks providing the most consistent long-term signals and clinical/proteomic clocks showing greater sensitivity to short-term physiological change. Subject-level analyses revealed that multi-omics BA profiles could highlight subtle deviations, including smoking-related risk, abnormal lipid profiles, shortened telomere predictions, and immune-cell composition changes, several of which aligned with clinical or self-reported health indicators.
These results position BA predictors not merely as retrospective ageing measures, but as actionable biomarkers for preventive and personalised medicine. Integrated into human digital twin frameworks, longitudinal BA measurements could anchor real-time models of individual health, detect early departures from expected trajectories, and support simulation of lifestyle or therapeutic interventions. This work underscores the potential of multi-omics ageing clocks to complement routine clinical testing and advance scalable, adaptive, and biologically informed digital twins for precision healthcare.

C-G.43: Annotating Eukaryotic Genomes by Combining Deep Learning with Extrinsic Evidence
Track: Genomics, epigenomics, and genome editing
  • Lars Gabriel, University of Greifswald, Germany
  • Katharina Jasmin Hoff, University of Greifswald, Germany


Presentation Overview: Show

The accuracy of ab initio gene prediction has been significantly advanced by Tiberius, a deep-learning gene finder that utilizes convolutional and LSTM layers paired with a differentiable Hidden Markov Model (HMM). However, even state-of-the-art deep learning architectures can benefit from the integration of extrinsic biological data to capture alternative splicing. We present a scalable evidence processing pipeline designed to complement the ab initio output of Tiberius by incorporating extrinsic evidence from transcriptomic and proteomic sources.

The pipeline operates as an integration framework that independently processes short-read and long-read RNA sequencing data, alongside large-scale protein database alignments. This allows for: (1) the addition of alternative isoforms supported by RNA-seq/Iso-Seq evidence that the ab initio model may not prioritize, (2) the refinement of gene boundaries using protein homology to validate coding sequences, (3) increased sensitivity in regions where genomic signals are weak but extrinsic support is robust.

We provide a modular path to incorporate experimental data into gene prediction with Tiberius. This pipeline improves transcript-level accuracy across eukaryotic datasets, offering a comprehensive solution for high-quality genome annotation. Tiberius and its associated evidence workflow are available at https://github.com/Gaius-Augustus/Tiberius.

C-G.44: Convergent niches, divergent genomes: a four-species pan-GWAS of Aspergillus pathogenicity and domestication
Track: Genomics, epigenomics, and genome editing
  • Minji Kim, BRIGHT, Technical University of Denmark, Denmark
  • Eduard Kerkhoven, Chalmers University of Technology, Sweden
  • Patrick Phaneuf, BRIGHT, Technical University of Denmark, Denmark


Presentation Overview: Show

The genus Aspergillus contains the filamentous fungi of substantial combined clinical and industrial importance: A. fumigatus and A. flavus are the principal agents of invasive aspergillosis, while A. niger and A. oryzae drive global industries in enzyme production and food fermentation. Whether these convergent phenotypes reflect shared genomic adaptations has not been tested at scale across the genus.
To address this, we constructed per-species pangenomes for four Aspergillus species of 211 ANI-verified, high-quality genomes spanning A. fumigatus (n=89), A. flavus (n=70), A. niger (n=19), and A. oryzae (n=33), together with a genus-level pangenome of 15,163 orthogroups. In the accessory genome, phenotype-labelled pan-genome-wide association studies (pan-GWAS) identified 2-117 significant orthogroup associations per species-phenotype contrast (BH-FDR<0.05), yet cross-species convergence testing revealed no shared evolutionary tactics. Even literature-curated virulence and industrial gene panels in the core genome offered no gene-content explanation: 76% of 226 anchored trait genes were core in all four species, and none of 124 virulence genes were pathogen-specific-indicating that lifestyle is not encoded by the presence or absence of shared genomic toolkits.
Instead, niche adaptation operates through lineage-specific rare gene compartments. Human-pathogenic strains showed significant rare gene expansion in A. fumigatus (p=8.1×10⁻⁹) and A. flavus (p=0.014), while industrial strains carried fewer lineage-specific genes than environmental strains (pooled, p=0.018). The rare gene compartment, often discarded as noise, represents the primary evolutionary substrate for clinical and biotechnological adaptation in this genus.

C-G.45: Gemsparcl: Rapid and consistent clustering of millions of genomes highlights the diversity of prokaryotic life
Track: Genomics, epigenomics, and genome editing
  • Johanna von Wachsmann, European Bioinformatics Institute, University of Cambridge, United Kingdom
  • John A. Lees, European Bioinformatics Institute, United Kingdom
  • Robert D. Finn, European Bioinformatics Institute, United Kingdom


Presentation Overview: Show

Bacterial genome databases collectively approach ten million assembled genomes, yet redundancy and limited scalability of existing tools create bottlenecks for comprehensive, tree-of-life-scale genomic analyses. One widely used approach involves dereplication to remove redundancy, but methods struggle beyond a few thousand genomes, making global organisation of public genomes computationally infeasible.
Here we present gemsparcl, a tool that clusters bacterial genomes into genomically coherent units (GCUs), orders of magnitude faster than existing methods. Central to gemsparcl is sketchlib.rust, a highly efficient one-permutation MinHash implementation with an auxiliary inverted index that substantially accelerates all-versus-all genome comparisons. gemsparcl further applies statistical correction for incomplete metagenome-assembled genomes (MAGs) and uses network-based quality filtering to remove edges weakly connecting distinct GCUs.
We clustered 5.6 million high-quality bacterial genomes (2.88 million isolates and 2.77 million MAGs) into 92,954 GCUs in approximately 14 hours using 48 cores and less than 16.5 GB of memory, achieving 99.94% cluster purity. Delving into these results revealed 2,582 GCUs, each comprising more than 50 genomes, with clusters formed exclusively by MAGs, highlighting priorities for future isolation efforts. More detailed analysis of the GCU networks provides insights into genomic diversity, with work ongoing to routinely extract this information from the clusters.
The scalability of gemsparcl transforms what was previously intractable: routine and up-to-date organisation of all public bacterial genomes and large-scale biodiversity assessment.

C-G.46: Whole-organism sequence-to-function modelling stratifies the cis-regulatory code of Drosophila into enhancer and locus grammars
Track: Genomics, epigenomics, and genome editing
  • Eren Can Eksi, VIB-KU Leuven, Belgium
  • Anton De Brabandere, VIB-KU Leuven, Belgium
  • Swann Floc'Hlay, VIB-KU Leuven, Belgium
  • Berfin Dag, VIB-KU Leuven, Belgium
  • Valerie Christiaens, VIB-KU Leuven, Belgium
  • Katina Spanier, VIB-KU Leuven, Belgium
  • Stein Aerts, VIB-KU Leuven, Belgium


Presentation Overview: Show

The canonical model of gene regulation is based on enhancer-promoter interactions to determine spatiotemporal properties of gene expression. So far, most studies on gene regulation in complex multicellular organisms have focused on particular tissues or cell types, having limited power in investigating possible higher hierarchical level gene regulatory principles. Therefore, we set out to investigate the rules of gene regulation and enhancer grammar in the context of a whole organism using sequence-to-function models. For this, we generated a whole adult fly scATAC-seq atlas containing around 700,000 cells and 150,000 peaks that are accessible in at least one cell type. We used topic modelling to cluster the cells in the atlas and transferred cell type annotations from the scRNA-seq atlas at two hierarchical levels (class and specific). This resulted in a pseudo-multiome atlas with matched scATAC-seq and scRNA-seq profiles across more than 100 cell types. We stratified scATAC-seq peaks according to their accessibility across cell types into ubiquitous, cell-type-specific, and putative cell-class-specific types. Next, we trained a suite of local and global sequence-to-function models to understand the grammars of the three classes of regions we identified and to integrate locus level information to predict gene regulation. Using these models, we identified and validated cell-type-specific enhancers for different cell types, investigated the promoters of genes and the relation of promoter-related motifs to gene expression patterns, and investigated the grammar and gene expression effects of a novel putative class of regions that we call 'class-level enhancers'.

C-G.47: MoleMap: fast alignment-free molecule mapping for long-read and linked-read sequencing data
Track: Genomics, epigenomics, and genome editing
  • Richard Lüpken, Leibniz Institute for Immunotherapy, Germany
  • Thomas Krannich, German Cancer Research Center, Germany
  • Markus Schuelke, Charite at Universitätsmedizin Berlin, Germany
  • Birte Kehr, Medizinische Hochschule Hannover, Germany


Presentation Overview: Show

With sequencing costs continuing to decline, the computational expenses of read alignment are becoming an increasingly important consideration. When only portions of the genome are of interest --as in most clinical genome analyses-- or base-pair precise alignment is not required --as for most local assembly approaches--, performing full read alignment on whole-genome sequencing (WGS) data imposes unnecessary computational costs.
We introduce MoleMap, a “molecule mapping” method for mapping long reads and linked-read molecules. MoleMap uses a minimized open-addressing k-mer index of the reference genome and a fast k-mer clustering procedure to map reads or molecules. In benchmarks, MoleMap agrees with minimap2 alignments of PacBio HiFi and ONT long read data on 99.68% and 98.03% of mappings respectively, outside of centromere and satellite repeat regions, while running 3-8x faster than other mappers and 10-60x faster compared to aligners on 32 threads. In addition, its low memory footprint allows whole-genome processing on a standard laptop computer. We apply MoleMap to filter reads for local assembly of a known variant region including non-reference sequence and showcase its benefits in diagnosing a patient with a rare disease.
MoleMap is an efficient mapping tool for all applications that do not require base pair precise alignment. Further, MoleMap can serve as a pre-processing tool for WGS data, accelerating targeted analyses by selecting reads from loci of interest before costly alignment. In a clinical sequencing setting, MoleMap can substantially reduce computational costs and diagnostic turnaround times.

C-G.48: Federated Genomic Variant Discovery & Analysis Using GA4GH Standards
Track: Genomics, epigenomics, and genome editing
  • Anais Mottaz, University of Applied Sciences Geneva HES-SO & Swiss Institute of Bioinformatics, Switzerland
  • Michael Baudis, Swiss Institute of Bioinformatics, Switzerland
  • Valerie Barbie, Swiss Institute of Bioinformatics, Switzerland
  • Tim Beck, University of Nottingham, United Kingdom
  • Melissa Cline, UC Santa Cruz Genomics Institute, United States
  • Friederike Ehrhart, Maastricht University., Netherlands
  • Hindrik Kerstens, Princess Maxima Center, Netherlands
  • Worawich Phornsiricharoenphant, Swiss Institute of Bioinformatics, Switzerland
  • Jordi Rambla, Centre de Regulacio Genomica, Spain
  • David Salgado, Institut Francais de Bioinformatique (IFB-Core), France
  • Venkata Satagopam, Luxembourg Centre For Systems Biomedicine (LCSB), Luxembourg
  • Sergi Beltran, Centro Nacional de Analisis Genomico (CNAG), Spain
  • Emidio Capriotti, University of Bologna, Italy


Presentation Overview: Show

Genomic data are increasingly distributed across clinical, research and national infrastructures, making interoperability a key challenge for their effective reuse. This is particularly critical for rare disease, cancer and population genomics, where variant interpretation requires integrating both observational data and interpretative knowledge across sources.
GA4GH standards provide a common framework to address these challenges. The Variant Representation Specification (VRS) enables consistent, computable representations of genomic variants across formats and sources, while Categorical VRS (Cat-VRS) extends this to groups of related variants. The Beacon protocol supports federated discovery across distributed datasets without centralising sensitive data, with controlled access and adjustable granularity of query responses. Downstream analyses can be performed within Trusted Research Environments (TREs), where standards such as the Task Execution Service (TES) support federated execution by bringing computation to the data.
This work illustrates how these components can be articulated into a federated analysis pipeline, using a rare disease-oriented scenario as an illustrative example. Starting from a candidate variant, consistent representations enable linking observational occurrences with interpretative evidence from databases and literature, including related variants within the same functional class. Federated discovery through Beacon identifies relevant datasets, while subsequent analyses within TREs enable comparison of phenotypic profiles across cohorts.
The work highlights current efforts and remaining challenges in interoperability, including variant representation harmonisation, integration of heterogeneous data types and alignment between discovery and analysis layers.

C-G.49: Large-scale genetic analysis of healthcare expenditure in 1.4 million individuals: The GenCost Consortium
Track: Genomics, epigenomics, and genome editing
  • Finngen, University of Helsinki, Finland
  • Andrea Ganna, University of Helsinki, Finland
  • Jakob German, University of Helsinki, Finland
  • Tatiana Cajuso Pons, University of Helsinki, Finland
  • Mykyta Artomov, University of Helsinki, Finland
  • Nikita Kolosov, University of Helsinki, Finland
  • Padraig Dixon, Nuffield Department of Primary Care Health Sciences, United Kingdom
  • Søren Brunak, Department of Public Health, Novo Nordisk Foundation Center for Protein Research, Denmark
  • Sarah E. Medland, QIMR Berghofer, Australia
  • Neil Davies, University College London, United Kingdom
  • Nicholas G Martin, QIMR Berghofer, Australia
  • Riccardo Marioni, University of Edinburgh, United Kingdom
  • Dorret Boomsma, Vrije Universiteit Amsterdam, Netherlands
  • Hamdi Mbarek, Qatar Precision Health Institute, Qatar
  • David van Heel, Queen Mary University of London, United Kingdom
  • Sebastian May-Wilson, University of Helsinki, Finland
  • Pradeep Natarajan, Broad Institute of MIT and Harvard, United States
  • Zhiyu Yang, University of Helsinki, Finland
  • Kristina Zguro, University of Helsinki, Finland
  • Erik Abner, University of Tartu, Estonia
  • Arne Kukkonen, University of Tartu, Estonia
  • Patrick Fahr, University of Oxford, United Kingdom
  • Stavroula Kanoni, Queen Mary University of London, United Kingdom
  • Camiel van der Laan, Vrije Universiteit Amsterdam, Netherlands
  • Chadi Saad, Qatar Precision Health Institute, Qatar
  • Anne Richmond, University of Edinburgh, United Kingdom
  • Penelope A. Lind, QIMR Berghofer, Australia
  • Ioannis Louloudis, Department of Public Health, Novo Nordisk Foundation Center for Protein Research, Denmark
  • Jiwoo Lee, Broad Institute of MIT and Harvard, United States
  • Tomoko Nakanishi, University of Helsinki, Finland


Presentation Overview: Show

We performed the largest genome-wide association study of healthcare costs to date, aiming to define the genetic architecture underlying genetic variation leading to increased healthcare expenditure. We analysed four major phenotypes: inpatient, inpatient plus outpatient, primary care, and prescription drug costs in up to 1.4 million individuals from 11 studies across 7 countries, using harmonized cost phenotypes derived from electronic health records and administrative databases. The cost phenotypes were analysed via GWAS meta-analysis, followed by conditional analyses, HLA fine-mapping, rare variant burden testing, colocalization, and polygenic score evaluation.
We identify 380 conditionally independent common variant associations across 248 loci, with the strongest and most pleiotropic signals located in the HLA region. Individual common variants had modest effects, typically altering annual costs by approximately 1–2% per allele, whereas rare deleterious variants in clinically actionable cancer predisposition genes, including BRCA1, BRCA2, MSH2, and APC, were associated with substantially larger increases in annual inpatient costs. Colocalization analyses linked cost-associated loci to autoimmune, cardiometabolic, pain-related, and psychiatric traits, indicating that healthcare expenditure reflects a broad spectrum of underlying disease biology.
Polygenic scores for healthcare costs predicted expenditure in independent cohorts and retained significant effects in within-family analyses, supporting largely direct genetic influences. Established polygenic scores for diseases and risk factors explained even greater variance in costs than cost-derived scores themselves. These findings show that healthcare expenditure has a measurable polygenic and rare variant architecture, providing a foundation for integrating human genetics into health economics, preventive strategies, and population-level screening.

C-G.50: High-Resolution Characterization of Repeat Expansions in Friedreich's Ataxia by Targeted Long-Read Sequencing
Track: Genomics, epigenomics, and genome editing
  • Christina Matlok, Dept. Molecular Neurology, FAU, Erlangen, Germany; ZSEER University Hospital Erlangen, Erlangen, Germany, Germany
  • Isabell Cordts, Department of Neurology, TUM University Hospital, Munich, Germany, Germany
  • Angela Abicht, Medical Genetics Center (MGZ), Munich, Germany; Dept. of Neurology, LMU, FBI, LMU München, Munich, Germany, Germany
  • Anna Anna Benet-Pagès, Medical Genetics Center (MGZ), Munich, Germany; Inst. of Neurogenomics, Helmholtz Center Munich, Neuherberg, Germany, Germany
  • David Brenner, Dept. of Neurology, University Hospital, Ulm, Germany; DZNE, Ulm, Germany; Center for Rare Diseases (ZSE), Ulm, Germany, Germany
  • Thomas Klopstock, FBI, Department of Neurology; DZNE, Munich, Germany; SynNergy, Munich, Germany, Germany
  • Martin Regensburger, Dept. of Molecular Neurology, FAU, Erlangen, Germany; ZSEER, University Hospital Erlangen, Erlangen, Germany, Germany
  • Annekathrin Rödiger, Department of Neurology, UKJ, Jena, Germany; Center for Rare Diseases, UKJ, Jena, Germany, Germany
  • Hayrettin Tumani, Department of Neurology, University Hospital Ulm, Ulm, Germany, Germany
  • Herbert Schreiber, Neurological Practice Center, Neuropoint Academy & NTD, Ulm, Germany, Germany
  • Benjamin Vlad, Department of Neurology, Jena University Hospital, Jena, Germany, Germany
  • Andreas Hauser, Medical Genetics Center (MGZ) Munich, Munich, Germany; Department of Neurology, TUM University Hospital, Munich, Germany, Germany
  • Franziska Bachhuber, Department of Neurology, University Hospital Ulm, Ulm, Germany, Germany
  • Boriana Büchner, Friedrich-Baur-Institute, Department of Neurology, LMU University Hospital, Munich, Germany, Germany
  • Almut Bischoff, Friedrich-Baur-Institute, Department of Neurology, LMU University Hospital, Munich, Germany, Germany
  • Jasper Hesebeck-Brinckmann, Department of Neurology, University Hospital Ulm, Ulm, Germany, Germany
  • Finja Grimm, Department of Neurology, TUM University Hospital, Munich, Germany, Germany
  • Marie Hackenberg, Medical Genetics Center (MGZ) Munich, Munich, Germany; Department of Neurology, TUM University Hospital, Munich, Germany, Germany
  • Thomas Risch, Medical Genetics Center (MGZ) Munich, Munich, Germany, Germany
  • Vitus Prokosch, Medical Genetics Center (MGZ) Munich, Munich, Germany, Germany
  • Ricarda von Heynitz, Department of Neurology, TUM University Hospital, Munich, Germany, Germany
  • Florentine Scharf, Medical Genetics Center (MGZ) Munich, Munich, Germany, Germany


Presentation Overview: Show

Friedreich's Ataxia (FRDA) is an autosomal recessive neurodegenerative disorder caused by pathological GAA repeat expansions within intron 1 of the FXN gene, with the shorter of two expanded alleles being the primary determinant of disease severity and age of onset. Non-GAA interruptions within the repeat are suspected to influence clinical manifestation but remain incompletely characterized by conventional methods. Long-read sequencing now enables comprehensive repeat characterization of repeat length and sequence.

We investigated GAA repeat characteristics in 39 FRDA patients from 5 centers in Germany using PureTarget Cas9-based enrichment with PacBio HiFi long-read sequencing. HiFi reads were processed using an in-house pipeline combining TRGT-based repeat length quantification, benchmarked against orthogonal methods (PCR, Southern Blot), including custom modules for signal processing-driven repeat interruption detection and motif composition determination.

Across the cohort, we obtained a mean of 257 repeat-spanning reads for the shorter allele. Repeat length estimates showed concordance with orthogonal methods (r=0.75), with shorter repeat lengths confirming the known inverse correlation with the age of onset (r=-0.48). Interruptions in short repeat alleles and non-canonical motif composition were identified in 17 and 2 patients, respectively findings inaccessible to conventional methods.

Targeted PacBio HiFi sequencing enables high-resolution characterization of GAA repeat expansions, including sequence features inaccessible to conventional approaches. This work contributes to deepening our understanding of the structural complexity and variability of repeat expansions in FRDA. Future efforts aim to explore additional molecular features, including methylation, and further evaluate long-read sequencing as a comprehensive diagnostic tool in clinical diagnostics.

C-G.51: Design and Validation of Deep Learning Models for Decoding Cis-Regulatory Logic in Drosophila melanogaster
Track: Genomics, epigenomics, and genome editing
  • Berfin Dag, VIB Center for AI and Computational Biology & KU Leuven, Belgium
  • Valerie Christiaens, VIB Center for AI and Computational Biology & KU Leuven, Belgium
  • Lukas Mahieu, VIB Center for AI and Computational Biology & KU Leuven, Belgium
  • Anton De Brabandere, VIB Center for AI and Computational Biology & KU Leuven, Belgium
  • Julie De Man, VIB Center for AI and Computational Biology & KU Leuven, Belgium
  • Gert Hulselmans, VIB Center for AI and Computational Biology & KU Leuven, Belgium
  • Stein Aerts, VIB Center for AI and Computational Biology & KU Leuven, Belgium


Presentation Overview: Show

Gene expression is controlled with remarkable precision across cell types and developmental stages through cis-regulatory elements, which integrate information about cell identity, expression magnitude, timing, and signalling inputs at the level of DNA sequence. Current deep learning models have shown that the sequence contains sufficient information to predict cell-type-specific enhancer activity, yet this represents only a fraction of the regulatory information embedded in the genome. Whether models can extract higher-order regulatory features and design functionally equivalent synthetic elements remains an open question.

Here we present a computational and experimental framework to decode cis-regulatory grammar in Drosophila melanogaster, using eye development as a model system where all cell types and regulatory factors are well characterized. We generated the largest single-cell multiome atlas of Drosophila eye development to date, integrating newly generated and public scATAC-seq and scRNA-seq datasets comprising over 100,000 cells across all major cell types. Using this atlas, we trained two complementary models: CREsted, which addresses the challenge of predicting cell-type-specific enhancer activity from local sequence context and enables design of synthetic enhancers with defined expression pattern; and a fine-tuned Fly-Borzoi model, which predicts how combinations of enhancers, their genomic context, and long-range interactions jointly determine gene expression level across a ~131kb receptive field.

To functionally validate model predictions, we have established an experimental framework based on precise genome editing at genomic loci, enabling quantitative transcriptional readout and phenotypic assessment. This pipeline is designed to link sequence-level predictions to organismal outcomes, providing a rigorous in vivo benchmark for sequence-to-function modeling.

C-G.52: Comparative benchmarking of DNA methylation imputation methods in the human placenta
Track: Genomics, epigenomics, and genome editing
  • Chen Zhang, University of Cambridge, United Kingdom
  • Dafina Angelova, University of Cambridge, United Kingdom
  • Gordon Smith, University of Cambridge, United Kingdom
  • Steve Charnock-Jones, University of Cambridge, United Kingdom
  • Sung Sam Gong, University of Cambridge, United Kingdom


Presentation Overview: Show

DNA modifications such as 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC), are key epigenic regulatory mechanisms. It is now possible to study DNA methylation at single-base resolution. However, if the coverage is less than the desired depth, the level of methylation of a given CpG is uncertain and is normally given as 'missing'. Imputation approaches have shown potential for addressing this missingness in the DNA methylation data, but comprehensive benchmarking studies that consider genomic regulatory context remain limited. Here, we benchmarked 10 imputation methods and compared their performance using the human placenta methylation data from the Pregnancy Outcome Prediction Study (POPS). We obtained the methylation data from 119 placenta samples using the Biomodal evoC platform that measures both 5hmC from 5mC. We generated ground truth datasets across four genomic regions (promoters, CpG islands, gene bodies and 5000bp tiling arrays) including only CpG sites with at least 10X— coverage across all samples. Imputation performance was evaluated separately for 5mC and 5hmC using Root Mean Square Error (RMSE). Our findings indicate that mean imputation method consistently outperformed other approaches evaluated, such as boostme, missForest and imputePCA. For example, within gene bodies, the RMSE of the mean method was the smallest for both 5mC (0.102) and 5hmC (0.051), while KNN and softImpute produced the highest RMSE values (0.176 at 5mC and 0.0746 at 5hmC) relative to the other methods. In summary, the mean approach offers a robust and accurate strategy for handling missing values in placental mC and hmC sequencing data.

C-G.53: Polygenic scores derived from breed-level obesity traits predict food motivation in individual dogs
Track: Genomics, epigenomics, and genome editing
  • Enoch Alex, University of Cambridge, United Kingdom
  • Jade Scardham, University Of Cambridge, United Kingdom
  • Anna Morros-Nuevo, University of Cambridge, United Kingdom
  • Natalie Wallis, University of Liverpool, United Kingdom
  • Alyce McClellan, University College London, United Kingdom
  • Eleanor Raffan, University of Cambridge, United Kingdom


Presentation Overview: Show

Human polygenic risk scores can predict complex disease liability, but their performance often weakens when applied across populations with different allele frequencies, linkage disequilibrium and environments. Domestic dogs offer a comparative version of the same problem, intensified by breed structure, where breeds differ in demography, haplotype background, obesity susceptibility and appetite related behaviour. We asked whether obesity related genetic effects estimated at breed level retain predictive value when applied to individual dogs.

Using breed average GWAS summary statistics for obesity probability and food motivation, we built polygenic scores using clumping and thresholding, LDpred2 and lassosum2. Scores were tuned in 325 dogs and evaluated in 487 independent multibreed dogs, with individual food motivation as the target and covariate adjustment for age, sex, neuter status and genetic principal components.

The food motivation derived score predicted individual food motivation with beta = 0.338, p = 5.64e-05 and incremental R² = 0.0318. The food motivation derived score stratified high food motivation, with top quintile dogs showing 3.91 fold higher adjusted odds than bottom quintile dogs, p = 0.007. Signals were retained in mixed breed dogs. External transfer was cohort dependent, with prediction of body condition in Labradors, beta = -0.217 and p = 0.007, but not Golden Retrievers.
These results show that breed level obesity genetics can retain modest individual level signal for appetite related behaviour, while highlighting phenotype alignment and population structure as shared constraints on polygenic prediction across human and canine genetics.

C-G.54: Genomic insights into racing camels: inbreeding levels and positive selection linked to athletic traits
Track: Genomics, epigenomics, and genome editing
  • Hussain Bahbahani, Kuwait University, Faculty of Science, Department of Biological sciences, Kuwait


Presentation Overview: Show

Racing dromedary camels are widely distributed across the Arabian Peninsula, with the highest concentrations in its northern and southeastern regions. In this study, whole‑genome sequences from 34 racing camels were analyzed to evaluate their genetic relationships with non‑racing populations, estimate inbreeding levels, calculate Weir and Cockerham's fixation index (Fst), assess effective population size (Ne), and identify genomic regions exhibiting signatures of positive selection.
Both racing and non‑racing camels showed comparable levels of genomic inbreeding (FROH = 0.21), and no significant genetic differentiation was detected between the two groups. The estimated Fst values further confirmed minimal population structure. Ne estimates revealed a declining trend in both groups over the past 5,000 years, with racing camels exhibiting slightly lower recent Ne values compared to their non‑racing counterparts.
Signatures of positive selection in racing camel genomes were identified using two haplotype‑based statistics—the integrated haplotype homozygosity score (iHS) and the between‑population extended haplotype homozygosity test (Rsb)—alongside a runs‑of‑homozygosity (ROH) analysis. A total of 33 candidate regions were detected using iHS, 19 using Rsb, and 24 through ROH. These regions overlapped with genes involved in biological pathways potentially linked to athletic performance, including musculoskeletal development, lipid metabolism, stress response, bone integrity, endurance, and power.
Overall, these findings explore the racing dromedary genome, with the aim of identifying variants and haplotypes associated with athletic traits. Such insights could support the development of genetically informed breeding programs designed to enhance specialized racing lines.