View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

A-G.01: esloco: simulation-based estimation of local coverage in long-read DNA sequencing
Track: Genomics, epigenomics, and genome editing
  • Adrian Weich, Friedrich-Alexander-Universität (FAU) Erlangen-Nürnberg and Uniklinikum Erlangen, Erlangen, Germany, Germany
  • Christopher Lischer, Friedrich-Alexander-Universität (FAU) Erlangen-Nürnberg and Uniklinikum Erlangen, Erlangen, Germany, Germany
  • Julio Vera, Friedrich-Alexander-Universität (FAU) Erlangen-Nürnberg and Uniklinikum Erlangen, Erlangen, Germany, Germany


Presentation Overview: Show

Motivation
Long-read DNA sequencing is increasingly used for whole-genome and targeted sequencing applications due to its ability to resolve complex structural variants and its reduced coverage requirements for variant detection. However, planning long-read sequencing experiments often lacks reliable a priori estimates of target region coverage, leading to costly and time-consuming pilot studies and biological replicates to meet expectations.

Results

We present esloco, a Monte Carlo-based simulation framework for estimating local coverage in long-read sequencing experiments. The tool supports scenarios with unknown target regions (e.g. viral integration sites, CRISPR-Cas9 edits) or PCR-free designs (e.g. base modifications). By modeling coverage as a function of sequencing depth and read length distribution, esloco enables quantitative predictions of local sequencing outcomes.

Benchmarking on a 45-gene panel demonstrated close agreement between simulated and empirical coverage profiles for both Oxford Nanopore and PacBio data. Predictive accuracy further improved when incorporating weighted masking to account for sequencing biases. These results establish esloco as a practical framework for simulation-guided experimental design in long-read sequencing studies.

Availability and Implementation
esloco is implemented in Python and is freely available on PyPI and Github as an open-source package.

A-G.02: Analytic Integration over Tree Space: Efficient Inference of Tumor Evolution Dynamics
Track: Genomics, epigenomics, and genome editing
  • Tobias Dieselhorst, Department of Biosystems Science and Engineering, ETH Zurich, Switzerland
  • Marcus Overwater, Department of Biosystems Science and Engineering, ETH Zurich, Switzerland
  • François Bienvenu, Laboratoire de mathématiques de Besançon, Université de Franche-Comté, France
  • Tanja Stadler, Department of Biosystems Science and Engineering, ETH Zurich, Switzerland


Presentation Overview: Show

The joint inference of phylogenetic trees and underlying population dynamics from sequence alignments is a powerful but computationally expensive procedure. Currently, sampling trees via Markov Chain Monte Carlo (MCMC) algorithms represents the major bottleneck in these schemes. In applications such as tumor evolution, however, the topology of the phylogenetic tree is often not of primary interest; consequently, the MCMC merely serves as a costly numerical integration over the tree space.
We present mathematical results on the properties of summary statistics for sequence alignments. In the limit of large trees, we show that summary statistics - specifically the distribution of mutations per sample and the site frequency spectrum - converge to the probability mass functions of statistically independent samples. The latter can be computed efficiently and correspond to an analytic integration over the tree space. We exploit this property to infer birth-death processes directly from site frequency spectra, which are commonly measured from bulk tumor samples but do not allow for tree reconstruction. We further discuss applications for testing substitution model adequacy based on the distribution of mutation counts.

A-G.03: Insights into the genomes of Mycobacterium abscessus: Decoding the arsenal shaping pathogenesis
Track: Genomics, epigenomics, and genome editing
  • Saubashya Sur, Ramananda College, India
  • Mistu Karmakar, Ramananda College, India


Presentation Overview: Show

Mycobacterium abscessus is an exteremely antibiotic-resistant non-tuberculous mycobacterium responsible for lung disease and extrapulmonary infections. Individuals suffering from cystic fibrosis, tuberculosis, bronchiectasis, AIDS, and chronic obstructive pulmonary disease are vulnerable. M. abscessus subsp. abscessus, M. abscessus subsp. massiliense, and M. abscessus subsp. bolletii are associated with diverse clinical manifestations and outbreaks. With no available vaccine and limited treatment options, comprehending the pathogenic arsenal from their genomes is crucial. The objective of this computational analysis was to explore the genomic island components; virulence factors, antibiotic resistance genes, anti-phage defence systems, and integrated prophage regions in M. abscessus genomes. While, GIPSY identified genomic islands, BRIG facilitated comparative visualization, VFDB and CARD identified the virulence and resistance factors. PADLOC and PHASTEST explored the repertoire of anti-phage defence systems, and integrated prophage regions. The genomes displayed variability. Pathogenicity islands were enriched in ABC transporters, efflux pumps, ISxac3 transposase, and TetR regulators. The resistance islands had notable prescence of acriflavine resistance protein B, AmpR, aminoglycoside acetyltransferases, and HAE1 efflux systems. In the metabolic islands, 4-hydroxybenzoate polyprenyltransferase, acetyl-CoA acetyltransferase, cytochrome P450, and ABC-type transporters were prominent. Virulent proteins like fbpA/C, espI/R, mmpS4/L4, mce family proteins, mprA-B, phoP/R, prrA-B, were dominant. Enriched antibiotic resiatance gene families were: class A β-lactamase, 16S rRNA mutations conferring aminoglycoside resistance, and 23S rRNA mutations conferring macrolide resistance. Notable anti-phage systems were PD-T4-6 and MTase_II. The prophage regions associated with virulence were identified. This intregrative analysis showcased the formidable arsenal shaping M. abscessus pathogenesis, underlining several novel therapaetic targets against this pathogen.

A-G.04: Comparative Analysis of Sugar Transporters in Yeast Strains for Improved Xylose Utilisation in Biofuel Applications
Track: Genomics, epigenomics, and genome editing
  • Kellie Ashton, University of Newcastle, Australia, Australia
  • Ian Grainge, University of Newcastle, Australia
  • Peter Lewis, Ethanol Technologies, Australia


Presentation Overview: Show

Efficient co-consumption of glucose and xylose is a major bottleneck in yeast-based second-generation biofuel production and is strongly influenced by sugar transporter specificity. While many industrial yeast strains transport glucose efficiently, xylose uptake remains limiting. Non-conventional yeasts with native xylose utilisation represent promising alternatives, however, their sugar transporter repertoires are often poorly characterised or inconsistently annotated.

In this study, a comparative bioinformatic analysis of the major facilitator superfamily (MFS) sugar transporters was performed in a proprietary yeast strain with native xylose utilisation. All predicted MFS transporters were identified from the genome and compared to characterised hexose transporters from Saccharomyces cerevisiae and known xylose transporters from various microorganisms to examine sequence conservation, transmembrane architecture, and residues implicated in substrate binding and translocation.

Multiple sequence alignments, phylogenetic tree construction, transmembrane domain prediction, and conserved motif analysis identified conserved MFS motifs as well as variable regions that may have functional significance. Several transporters shared structural features and conserved residues with characterised xylose transporters, while others clustered with hexose transporters, allowing us to predict which transporters may be utilising pentose sugars as substrates.

Comparative analysis clarified the relationship between closely related Hxt6-like transporters, despite multiple similar sequences and inconsistent nomenclature across datasets. These findings support the prioritisation of specific transporters for further functional investigation.

This analysis provides a framework for examining sugar transporter diversity in non-conventional yeasts and informs the rational selection of transporter candidates for engineering improved xylose uptake and glucose/xylose co-consumption in yeast.

A-G.05: ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data
Track: Genomics, epigenomics, and genome editing
  • Huda Ahmad, University of Birmingham, United Kingdom
  • Hannah Doherty, University of Cologne, Germany
  • Samuel Benedict, Newcastle University, United Kingdom
  • James Haycocks, Newcastle University, United Kingdom
  • Ge Zhou, King Abdullah University of Science and Technology, Saudi Arabia
  • Patrick Moynihan, Western University, Canada
  • Danesh Moradigaravand, King Abdullah University of Science and Technology, Saudi Arabia
  • Manuel Banzhaf, Newcastle University, United Kingdom


Presentation Overview: Show

Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes. However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive tools for visualization and analysis. Existing workflows often require multiple independent software tools, limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables efficient exploration and sharing of chemical genomics data.

We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic profiles, assess gene–gene and condition–condition correlations, perform GO and KEGG enrichment analysis, and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances reproducibility, and promotes collaboration within the research community. ChemGenXplore is available at https://chemgenxplore.kaust.edu.sa/.

A-G.06: ROQ: a new measure for more accurately filtering mapped reads
Track: Genomics, epigenomics, and genome editing
  • Nayoung Park, Konkuk University, South Korea
  • Minji Gu, Konkuk University, South Korea
  • Younhee Ko, Hankuk University of Foreign Studies, South Korea
  • Jaebum Kim, Konkuk University, South Korea


Presentation Overview: Show

Motivation: Next-generation sequencing (NGS) has transformed genomic and multiomic research, making accurate read alignment essential for reliable downstream analyses such as variant detection and gene expression quantification. Widely used alignment quality metrics, such as Mapping Quality (MAPQ), do not explicitly account for the true genomic origin of sequencing reads and are computed inconsistently across alignment tools. As a result, misaligned reads may receive high confidence scores, while correctly aligned reads can be undervalued, particularly in complex or repetitive genomic regions. Although machine learning-based recalibration approaches have demonstrated the feasibility of improving alignment quality assessment, no existing metric directly models the spatial concordance between a read's mapped position and its true genomic origin, and its generalizability across diverse aligners, species, and sequencing platforms has not been systematically evaluated.
Results: We introduce Read Overlapping Quality (ROQ), a machine learning-based metric designed to predict the Read Overlap Ratio (ROR), which quantifies the degree of spatial overlap between a read's true genomic origin and its mapped position. Trained on thirteen alignment-derived features using the XGBoost algorithm, ROQ provides a more origin-aware and comprehensive assessment of alignment quality than MAPQ alone. ROQ consistently outperforms MAPQ in correlation with the true read origin across diverse genomic contexts, multiple species, and alignment tools, and demonstrates cross-platform robustness on independent long-read PacBio data. When applied to real whole-genome sequencing data, ROQ-based filtering reduces false-positive variant calls while maintaining sensitivity, supporting its utility as a complementary metric for genomic analysis pipelines.

A-G.07: Sorting Bacterial Genomes Below the Species Rank
Track: Genomics, epigenomics, and genome editing
  • Beatriz Vieira Mourato, Max-Planck-Institute for Evolutionary Biology, Germany
  • Sara-Lena Welk, Max-Planck-Institute for Evolutionary Biology, Germany
  • Fabian Kloetzl, None, United Kingdom
  • Bernhard Haubold, Max-Planck-Institute for Evolutionary Biology, Germany


Presentation Overview: Show

Motivation: Bacterial genomes are classified as belonging to
the same species if their average nucleotide identity, ANI, is at
least 96. The resulting species include organisms
like \emph{Escherichia coli} that comprise both commensals and
pathogens, that is, they harbor important taxonomic diversity below
the species rank. This diversity is less well characterized than at
the species rank and above because the required comparison between
taxonomic and phylogenetic information is more difficult to automate
than applying an ANI threshold. As a result, the majority of bacterial
genomes are only classified at the species rank even if there taxa
below that rank.
Results: Here we quantify the distribution of
bacterial genomes between and within species to show that over 3/4 of
genomes derive from species with below species structure. We establish
a workflow for sorting genomes within species by constructing a genome
phylogeny of all genomes in the species and querying it for the target
taxon. This workflow requires a tool for ANI computation and tools for
querying phylogenies, which we describe. We demonstrate this approach
for the pathogen \emph{E. coli} O111:H8, which we analyze in the
context of the over 7,000 complete genome sequences of \emph{E. coli}
to find 26 genomes in addition to its original 22. We use these 48
genomes to construct candidate diagnostic PCR primers
for \emph{E. coli} O111:H8.
Availability: Our workflow is
distributed as a literate program via the companion website for this
paper, \ty{github.com/evolbioinf/o111h8}.

A-G.08: How Private Are DNA Embeddings? Inverting Foundation Model Representations of Genomic Sequences
Track: Genomics, epigenomics, and genome editing
  • Sofiane Ouaari, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
  • Jules Kreuer, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
  • Nico Pfeifer, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany


Presentation Overview: Show

DNA foundation models have become transformative tools in bioinformatics and healthcare applications. Trained on vast genomic datasets, these models can be used to generate sequence embeddings, dense vector representations that capture complex genomic information. These embeddings are increasingly being shared via Embeddings-as-a-Service (EaaS) frameworks to facilitate downstream tasks, while supposedly protecting the privacy of the underlying raw sequences. However, as this practice becomes more prevalent, the security of these representations is being called into question. This study evaluates the resilience of DNA foundation models to model inversion attacks, whereby adversaries attempt to reconstruct sensitive training data from model outputs. In our study, the model's output for reconstructing the DNA sequence is a zero-shot embedding, which is then fed to a decoder. We evaluated the privacy of three DNA foundation models: DNABERT-2, Evo 2, and Nucleotide Transformer v2 (NTv2). Our results show that per-token embeddings allow near-perfect sequence reconstruction across all models. For mean-pooled embeddings, reconstruction quality degrades as sequence length increases, though it remains substantially above random baselines. Evo 2 and NTv2 prove to be most vulnerable, especially for shorter sequences with reconstruction similarities > 90\%, while DNABERT-2's BPE tokenization provides the greatest resilience. We found that the correlation between embedding similarity and sequence similarity was a key predictor of reconstruction success. Our findings emphasize the urgent need for privacy-aware design in genomic foundation models prior to their widespread deployment in EaaS settings. Training code, model weights and evaluation pipeline are released on: https://github.com/not-a-feature/DNA-Embedding-Inversion.

A-G.09: TabPFN-Wide: Continued Pre-Training for Extreme Feature Counts
Track: Genomics, epigenomics, and genome editing
  • Christopher Kolberg, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
  • Jules Kreuer, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
  • Jonas Huurdeman, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
  • Sofiane Ouaari, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
  • Katharina Eggensperger, Lamarr Institute for Machine Learning and Artificial Intelligence, TU Dortmund, Germany
  • Nico Pfeifer, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany


Presentation Overview: Show

Revealing novel insights from the relationship between molecular measurements and pathology remains a very impactful application of machine learning in biomedicine. Data in this domain typically contain only a few observations but thousands of potentially noisy features, posing challenges for conventional tabular machine learning approaches. While prior-data fitted networks emerge as foundation models for predictive tabular data tasks, they are currently not suited to handle large feature counts (>500). Although feature reduction enables their application, it hinders feature importance analysis. We propose a strategy that extends existing models through continued pre-training on synthetic data sampled from a customized prior. The resulting model, TabPFN-Wide, matches or exceeds its base model's performance, while exhibiting improved robustness to noise. It seamlessly scales beyond 30,000 categorical and continuous features, regardless of noise levels, while maintaining inherent interpretability, which is critical for biomedical applications. Our results demonstrate that prior-informed adaptation is suitable to enhance the capability of foundation models for high-dimensional data. On real-world omics datasets, we show that many of the most relevant features identified by the model overlap with previous biological findings, while others propose potential starting points for future studies.

A-G.10: Short-Context gLM Training for Long-Range Variant Effect Prediction
Track: Genomics, epigenomics, and genome editing
  • Megha Hegde, Kingston University London, United Kingdom
  • Jean-Christophe Nebel, Kingston University London, United Kingdom
  • Farzana Rahman, Kingston University London, United Kingdom


Presentation Overview: Show

Disease susceptibility in humans is frequently driven by mutations within the DNA. Contemporary research has shown that many such mutations lie within the non-coding regions of the genome, several thousands of base-pairs (bp) from the transcription start sites (TSS) of their target genes. In the age of artificial intelligence, Transformer-based genomic language models (gLMs) are commonly used to interpret variant effects. Their ability to exploit the large datasets produced by next-generation sequencing, and their aptitude for modelling long-range interactions within sequences, makes them ideal for the task. However, the quadratic scaling of the attention mechanism with context length results in high computational resource consumption, and inefficiency when training gLMs on very long (10000bp+) sequences.
DroPE (Dropping the Positional Embeddings of LMs after training), originally developed for generative text-based large language models (LLMs), provides a method for extending the context of pretrained LLMs without the need for long-context fine-tuning. This paper adapts and applies DroPE to BERT-based gLMs, and demonstrates that it can extend the context of models pretrained on short genomic sequences. Furthermore, models incorporating DroPE demonstrate an enhanced ability to model long-range context, achieving competitive performance on variants far (>35kbp) from the TSS without the need for computationally expensive long-context pretraining and fine-tuning. Code and fine-tuned models are available at: https://github.com/meghegde/gLM-DroPE.

A-G.11: NCRP: enhancing long-read classification through neighborhood-consistency refinement and propagation in the overlap graph
Track: Genomics, epigenomics, and genome editing
  • Xun Ding, SHANDONG UNIVERSITY, China
  • Lianrong Pu, SHANDONG UNIVERSITY, China
  • Haitao Jiang, SHANDONG UNIVERSITY, China
  • Zhu Daming, SHANDONG UNIVERSITY, China


Presentation Overview: Show

Taxonomic classification of metagenomic sequencing reads is a crucial task in metagenome analysis, with significant implications for research on health, diet, and drug responses. With the advancement of sequencing technologies, several methods have been developed for long-read taxonomic classification, including Kraken2, Centrifuge, and CLARK. While these classifiers achieve high precision, they often suffer from low sensitivity, leaving a substantial proportion of reads unclassified, especially for highly diverse and rapidly evolving organisms such as viruses. ClassGraph addressed this challenge by introducing a label propagation strategy, significantly improving sensitivity at the cost of precision.To further address this issue, we propose NCRP, a Neighborhood-Consistency Refinement and Propagation algorithm, designed to enhance both precision and sensitivity in long-read classification. NCRP builds upon an initial classifier and a read overlap graph. It first identifies and removes ambiguous classifications that are inconsistent with neighboring reads in the overlap graph, then propagates confident taxonomic assignments from classified reads to unclassified ones. Evaluation on simulated and mock datasets demonstrates that NCRP consistently outperforms existing classifiers in both precision and sensitivity, achieving the highest F1 score across all tests. We further benchmarked the ability of taxonomic classifiers to detect microbial species. NCRP combined with Kraken2 or Centrifuge achieved the best performance, whereas CLARK showed relatively poor performance on this task.NCRP is open-source and can be accessed at https://github.com/SDU-ACG-Lab/NCRP.

A-G.13: Compounding of rare copy-number variants and polygenic risk: A genetic signature of assortative mating
Track: Genomics, epigenomics, and genome editing
  • Caterina Cevallos, University of Lausanne, Switzerland
  • Chiara Auwerx, University of Lausanne, Switzerland
  • Robin Hofmeister, University of Lausanne, Switzerland
  • Théo Cavinato, University of Lausanne, Switzerland
  • Tabea Schoeler, University of Lausanne, Switzerland
  • Zoltán Kutalik, University of Lausanne, Switzerland
  • Alexandre Reymond, University of Lausanne, Switzerland


Presentation Overview: Show

Background:
Large copy-number variants (CNVs) are linked to a broad spectrum of outcomes, with carriers of the same CNV exhibiting variable disease severity. Although additional rare and common variants have been proposed as modifiers of CNVs' expressivity, their interplay still remains largely understudied.

Material and Methods:
We explored the impact of polygenic scores (PGS) on shaping CNV carriers' heterogeneity in the UK Biobank, focusing on 119 established CNV-trait associations comprising 43 traits and 27 CNVs. Linear regressions assessed the individual, joint, and synergistic contributions of CNVs and PGSs. Due to participation bias in population biobanks, we expected trait-increasing CNV carriers to have lower PGSs.

Results:
We demonstrate additive contribution of PGS and CNV for 45 (38%) CNV-trait pairs, as well as two interactions between the 22q11.23 duplication and PGSs for grip strength and gamma-glutamyltransferase levels. Strikingly, CNVs and PGSs exhibited a widespread positive correlation, revealing a tendency for PGSs to exacerbate CNV effects—a pattern that could be explained by linkage disequilibrium only for a single CNV-trait pair. Given a non-null inheritance rate for all 17 testable CNVs, we investigated whether assortative mating could account for this phenomenon. We found strong agreement between the empirical CNV-PGS correlation and the one predicted by assortment (r=0.45, p=3.9e-7). Similar trends of positive correlation were observed between PGSs and genome-wide burden of CNVs or loss-of-function variants.

Conclusion:
PGSs improve stratification of CNV carriers at risk of developing clinically-relevant comorbidities, compounding the impact of rare damaging variants through assortative mating.

A-G.14: ReVSeq Visuals: A Rapid, Modular Environment for Viral Sequence Visualization
Track: Genomics, epigenomics, and genome editing
  • Charlyne Bürki, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland
  • Mara Neacsu, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland
  • Louis du Plessis, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland
  • Tanja Stadler, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland


Presentation Overview: Show

Motivation: Rapidly investigating and understanding the evolution of co-circulating respiratory pathogens is paramount for robust surveillance systems.
However, bottlenecks exist in implementing bioinformatics pipelines and interpreting their outputs for non-domain specialized users such as public
health professionals and clinicians, especially in environments requiring regular data updates or studies involving multi-center collections.
Results: We developed ReVSeq Visuals, a framework to build customizable dashboards to obtain instantaneous insights from sequencing data in an
accessible and intuitive manner. ReVSeq Visuals consists of a backend module to curate the output of viral genomic bioinformatics pipelines and an
interactive visualization platform to derive insights from sequencing data. Together, with limited bioinformatics configurations required by the user,
these two components form the dashboard-building framework that visualize insights based on viral sequencing data together with its associated
spatiotemporal metadata. The dashboard integrates existing public health tools, allowing to fully customize the studied viral strains. We showcase
the dashboard's functionalities using a set of diverse respiratory viruses collected in Switzerland from 2023 to 2025, and illustrate derivable insights
for RSV-A/B co-circulation.
Availability and Implementation: All code (predominantly Python) is available at https://github.com/charlynebuerki/ReVSeq-dashboard and
the dashboard is available at https://revseq.charlynebuerki.com/.

A-G.15: Genome wide eccDNA hotspot propensity estimation from confounder aware features
Track: Genomics, epigenomics, and genome editing
  • Gerardo Benevento, University of Salerno, Italy
  • Delfina Malandrino, University of Salerno, Italy
  • Daniele Salerno, University of Salerno, Italy
  • Alessia Ture, University of Salerno, Italy
  • Rocco Zaccagnino, University of Salerno, Italy


Presentation Overview: Show

Studying extrachromosomal circular DNA (eccDNA and ecDNA) can reveal genome instability and regulatory change, especially in cancer
disease. Interpreting genome wide recurrence is difficult because apparent hotspots can be inflated by technical and genomic confounders,
including low mappability, repeats, and segmental duplications. We therefore move from classifying individual circles to a locus propensity view
and learn to rank genomic windows using interpretable context features, while using technical tracks only for filtering and diagnostics. We score
recurrence across public experiments from CircleBase v2 in fixed 25 kb bins. For each bin we compute replicate support, defined as the number
of independent experiments (unique PubMed ID and assay) that report at least one overlapping eccDNA interval. To define a fair baseline,
we keep the number and sizes of eccDNA intervals from each experiment fixed, but randomly shuffle their genomic positions within the same
chromosome many times. We then measure how often each bin would be supported under these randomized placements and use the average
as the expected background support. Comparing observed support to this baseline provides a measure of the continuous enrichment score used
for regression and genome wide ranking tasks. The proposed approach was evaluated along two axes: (i) recovery and tail prioritization of
the continuous enrichment signal (random-split = 0.711 ± 0.004, Lift@1% 38, chromosome holdout 0.51, with telomere/centromere
ablations probing positional reliance) and (ii) matched comparison to sequence-only methods via the DeepCircle 1 kb balanced classification
protocol, where we obtain competitive performance.

A-G.16: SPA-C: an hybrid tool to accurately scaffold genomes using Hi-C and Deep-Learning
Track: Genomics, epigenomics, and genome editing
  • Alexis Mergez, University of Toulouse, INRAE, France
  • Raphael Mourad, University of Toulouse, France
  • Guillermina Hernandez-Raquet, University of Toulouse, TBI, INRAE, France
  • Matthias Zytnicki, University of Toulouse, INRAE, France


Presentation Overview: Show

Genome assembly is a computational pipeline designed to reconstruct chromosomes from small sequencing reads. Following their assembly, contiguous sequences (contigs) are arranged into chromosome-long sequences during scaffolding. Hi-C, a long-range linkage information between regions of the genome widely used in recent large sequencing projects, is often required to correctly order contigs. Several tools have been developed to automate this task following either statistical or deep-learning approaches. Statistical approaches summarise 2D Hi-C matrices into contact densities across sequences, thus ignoring informative visual patterns. The sole existing deep-learning tool uses a transformer-based computer vision model to correct the assembly. It has been trained on several species and uses Hi-C matrices directly. Yet it comes as a supplementary step in the scaffolding process, introducing extra computation time, and has been trained on a dataset that might contain labelling errors, that could limit its utilisation.
We propose SPA-C, an hybrid pipeline combining the strengths of both approaches. Linkage prediction is handled with a frugal CNN-based model and a graph-solving algorithm is used to generate the scaffolds. Through our input's design, the model is able to both correct errors within assemblies and link contigs, leveraging small, local Hi-C contact matrices. We handled low-complexity regions that might induce erroneous predictions using an external tool, improving the overall accuracy of generated assemblies. On a benchmark of six various genomes and four standard metrics, SPA-C outperformed four out of four state-of-the-art methods while achieving comparable start-to-end computation time.
Python and Bash scripts are available at zenodo (doi: 10.5281/zenodo.19000362).

A-G.17: Kente: A Graph-based Pangenomic Approach for Horizontal Gene Transfer Detection in Microbiomes.
Track: Genomics, epigenomics, and genome editing
  • Natalie Kokroko, Rice University, United States
  • Richa Jayanti, Rice University, United States
  • Nicolae Sapoval, Rice University, United States
  • Michael Nute, Rice University, United States
  • Luay Nakhleh, Rice University, United States
  • Todd Treangen, Rice University, United States


Presentation Overview: Show

Motivation: Horizontal gene transfer (HGT) shapes bacterial evolution and microbial ecosystems, yet detecting HGT within microbiomes remains a challenge due to fragmented metagenomic assemblies, reference bias, reliance on gene boundaries, and limited ability to model structural mosaicism and patterns across genomes.
Methods: We present Kente, a novel pangenome graph-based framework designed for HGT detection that aligns
metagenomic assembly contigs to a curated database of >600 genus-level bacterial pangenome graphs constructed
using minigraph. Kente infers local taxonomic composition along contigs using alignment evidence and classifies
candidate transfers using structured clade-transition topologies (e.g., A-B-A sandwich, open tips, and mosaic
patterns). A complementary intra-genus module detects inter-species transfers within a single genus graph using
segment-level clade annotations.
Results: Across simulated intra- and inter-genus transfer scenarios, Kente achieves higher precision and comparable
recall relative to existing gene-centric microbiome HGT detection approaches while reducing false positives from
fragmented assemblies. Application to real human gut metagenomes (HMP2, n = 26) demonstrates Kente's ability
to detect candidate cross-lineage transfer regions in complex microbial communities. Runtime profiling shows near-linear
scaling with input size, enabling efficient analysis of large metagenomic assemblies.
Availability and Implementation: https://github.com/treangenlab/Kente

A-G.18: A context-aware framework leveraging LLM-estimated variant effects for individualized variant-to-phenotype interpretation
Track: Genomics, epigenomics, and genome editing
  • Nayoung Park, Department of Biomedical Science and Engineering, Konkuk University, South Korea
  • Dayeon Kim, Division of Biomedical Engineering, Hankuk University of Foreign Studies, South Korea
  • Jaebum Kim, Department of Biomedical Science and Engineering, Konkuk University, South Korea
  • Younhee Ko, Division of Biomedical Engineering, Hankuk University of Foreign Studies, South Korea


Presentation Overview: Show

Motivation: Advances in sequencing technologies have uncovered millions of human genetic variants, yet accurately estimating pathogenic risk at the sample level remains a major challenge. Polygenic risk scores (PRS), built on genome-wide association studies (GWAS), provide a useful framework for aggregating population-level effects, but they have important limitations. Because PRS depends on cohort-specific effect sizes, its performance often does not generalize well across populations. In addition, PRS is not well-suited to capturing rare or context-dependent variant effects that may contribute substantially to disease risk in individual genomes. To address these limitations, we developed the Variant Probability Ratio-based score (VPRscore), a context-aware framework that leverages probability ratios derived from a large language model trained on genomic sequences. For each variant, the model estimates the probability of the alternative allele relative to the reference allele within its native sequence context, yielding a variant probability ratio (VPR) that reflects potential functional impact.

Results: We integrated these variant-level VPR signals with CADD to construct a unified sample-level score, VPRscore, that provides a biologically grounded estimate of disease risk informed by variant pathogenicity. In benchmark analyses, VPRscore robustly distinguished pathogenic from benign variants in ClinVar and accurately tracked increasing proportions of pathogenic variants across simulated cohorts. The method achieved high discriminative performance, with AUC values reaching 0.989. By deriving risk estimates directly from DNA sequence context rather than relying on population-level association statistics, VPRscore provides an interpretable and potentially more generalizable foundation for individualized genomic risk assessment.

A-G.19: Read-Level Classification of HPV Viral Metagenomic Data Using Transformer Architectures
Track: Genomics, epigenomics, and genome editing
  • Simone Rancati, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
  • Sakshi Pandey, Department of Computer and Information Science and Engineering, University of Florida, United States
  • Micheal Sy, Department of Epidemiology, College of Public Health and Health Professions, University of Florida, United States
  • Pablo Arozarena Donelli, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
  • Giovanna Nicora, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
  • Conrad Testagrose, Department of Computer and Information Science and Engineering, University of Florida, United States
  • Christina Boucher, Department of Computer and Information Science and Engineering, University of Florida, United States
  • Riccardo Bellazzi, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
  • Marco Salemi, Emerging Pathogens Institute, University of Florida, United States
  • Enea Parimbelli, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
  • Simone Marini, Department of Epidemiology, College of Public Health and Health Professions, University of Florida, United States


Presentation Overview: Show

Motivation: Accurate detection of viral reads within overwhelming human and bacterial background is essential for viral metagenomics, outbreak surveillance, and precision diagnostics. This is also relevant in cancer and, more broadly, in shotgun sequencing workflows, where accurate read-level discrimination can improve downstream genotyping, assembly, and interpretation. Alignment-free tools such as Kraken-2 are widely
used, but their reliance on exact or near-exact k-mer matches may reduce sensitivity to divergent or low-abundance viral reads. Transformer encoders provide contextual nucleotide representations that may overcome these limitations, yet their performance on realistic HPV-scale short reads remains poorly explored. Results: We compared two viral transformers, ViBE and XVir, against Kraken-2 on 150-bp Illumina-like reads simulated from 184 viral genomes, 147 bacterial genomes, and 62 human assemblies, using human papillomavirus (HPV) as the pathogen of interest. Across four tasks (Virus vs Human, Virus vs Bacterial, Human vs Bacterial, and Virus vs Human vs Bacterial), ViBE consistently generated the most informative embedding space and the best downstream classification, achieving 95–98% balanced accuracy in binary tasks and 91–92% in the three-class setting, often with only the top 100 embedding features. XVir, despite its HPV-focused design, performed well
mainly in Virus vs Human discrimination and generalized less effectively to bacterial background. Kraken-2 remained strongest in Human vs Bacterial separation but showed lower performance on HPV-centered tasks. Availability and implementation: https://github.com/simoRancati/single-read-viral-transformers-Read-Level-Classification-of-Viral-Metagenomic-Data-Using-Transforme

A-G.20: Characterizing the temporal dynamics of Cancer Hallmarks and somatic mutations across 80,000+ tumors
Track: Genomics, epigenomics, and genome editing
  • Federica Malighetti, University of Milano Bicocca, Italy
  • Alberto Maria Villa, University of Milano Bicocca, Italy
  • Ivan Civettini, San Raffaele Scientific Institute, Milan, Italy, Italy
  • Luis Zapata, Institute of Cancer Research, London, UK, United Kingdom
  • Andrea Aroldi, University of Milano Bicocca, Italy
  • Luca Mologni, University of Milano Bicocca, Italy
  • Daniele Ramazzotti, University of Milano Bicocca, Italy


Presentation Overview: Show

Understanding cancer evolution across diverse tumor types remains challenging due to its inherent heterogeneity. However, uncovering evolutionary pathways common across tumors might elucidate fundamental principles of cancer progression. This study extends previous efforts to addresses this challenge by integrating recurrent mutations with the temporal acquisition of Hallmarks of Cancer across a diverse cohort of over 80,000 tumors.
By focusing on Hallmarks dynamics, we discover distinct patterns of tumor progression with clinical implications. Leveraging the ASCETIC framework to characterize cancer evolution, we comprehensively analyze over 80,000 tumors spanning most cancer types. We map Hallmarks enrichment across the different phases of tumor progression and track the enrichment of Hallmarks at various stages of tumor evolution, revealing which Hallmarks exert significant influence during cancer progression in distinct phases, and their impact on prognosis and metastasis. Key findings include the identification of conserved mutation trajectories, such as TP53, ARID1A, and RNF43 as early drivers, and the sequential acquisition of mutations in ErbB receptors like EGFR and ERBB family members. These findings significantly advance our understanding of cancer evolution and might have significant implications for the development of broad-spectrum targeted therapies.

A-G.21: Interpreting Structural Variations in Breast Cancer Cell Lines Using 3D Genome Data
Track: Genomics, epigenomics, and genome editing
  • Jingyu Hao, The Hong Kong University of Science and Technology, Hong Kong
  • Yongyi Luo, The Chinese University of Hong Kong, Hong Kong
  • Jiandong Shi, The Chinese University of Hong Kong, Hong Kong
  • Depeng Wang, GrandOmics Inc, China
  • Shu Wang, Peking University People’s Hospital, China
  • Xiaodan Fan, The Chinese University of Hong Kong, Hong Kong
  • Weichuan Yu, The Hong Kong University of Science and Technology, Hong Kong


Presentation Overview: Show

Structural variations (SVs) are a major source of genomic alterations in cancer, yet the majority of SVs occur in non-coding regions and remain difficult to interpret using linear genome annotations alone. 3D genome data provides additional spatial information for understanding potential regulatory impacts of SVs, but its practical contributions and limitations remain unclear. Here, we perform an integrative analysis of SVs and 3D genome organization in three breast cancer cell lines. Although the majority of SVs occur in non-coding regions, approximately 70% of them are associated with distal gene promoters via high-frequency 3D contacts, revealing potential regulatory relationships beyond linear genome annotations. In the HCC1937 tumor–normal pair under a controlled genetic setting, we observed widespread 3D genome reorganization, including A/B compartment switching affecting roughly one-third of the genome and pronounced insulation score changes at about 30% of somatic SV loci. Detailed analyses of representative somatic SVs further revealed local alterations in topologically associating domain (TAD) structure at SV loci, including the formation of new boundaries or the disruption of pre-existing ones. These SV loci also form high-frequency contacts with distal cancer-related gene loci, such as SLC2A1, HGF, and STAG2, with similar interactions observed for some of these genes in SK-BR-3, suggesting that these spatial associations are common across breast cancer cell lines. These observations indicate that SVs that appear functionally silent under conventional annotation may influence gene regulation through their 3D genome, providing a potential mechanism for prioritizing non-coding variants for further functional investigation.

A-G.22: GENA-Web - GENomic Annotations Web Inference using DNA language models
Track: Genomics, epigenomics, and genome editing
  • Aleksei Shmelev, AXXX, Moscow, Russia; HSE University, Moscow, Russia, Russia
  • Maxim Petrov, AXXX, Moscow, Russia, Russia
  • Dmitry Penzar, Vavilov Institute of General Genetics, Russian Academy of Sciences, Moscow, Russia, Russia
  • Nikolay Akhmetyanov, AXXX, Moscow, Russia, Russia
  • Maksim Tavritskiy, AXXX, Moscow, Russia, Russia
  • Stepan Mamontov, AXXX, Moscow, Russia, Russia
  • Yuri Kuratov, AXXX, Moscow, Russia; Moscow Independent Research Institute of Artificial Intelligence, Moscow, Russia, Russia
  • Mikhail Burtsev, London Institute for Mathematical Sciences, London, United Kingdom, United Kingdom
  • Olga Kardymon, Quantori, Belgrade, Serbia, Serbia
  • Veniamin Fishman, AXXX, Moscow, Russia; Sirius University, Sochi, Russia; Institute of Cytology and Genetics, Novosibirsk, Russia, Russia


Presentation Overview: Show

The advent of advanced sequencing technologies has significantly reduced the cost and increased the feasibility of assembling high-quality genomes. Yet, the annotation of genomic elements remains a complex challenge. Even for species with comprehensively annotated reference genomes, the functional assessment of individual genetic variants is not straightforward. In response to these challenges, recent breakthroughs in machine learning have led to the development of DNA language models. These transformer-based architectures are designed to tackle a wide array of genomic tasks with enhanced efficiency and accuracy. In this context, we introduce GENA-Web, a web-based platform that consolidates a suite of genome annotation tools powered by DNA language models. The version of GENA-Web presented here encompasses a diverse set of models trained on human data, including the prediction of promoter activity, annotation of splice sites, determination of various chromatin features, and a model for scoring of enhancer activity in \textit{Drosophila}. GENA-Web is accessible online at https://dnalm.airi.net/

A-G.23: Mapping and improving the limited sensitivity of SigProfilerAssignment to flat mutational signatures
Track: Genomics, epigenomics, and genome editing
  • Qixuan Wang, Inselspital, Bern University Hospital and University of Bern, Switzerland, Switzerland
  • Maria Katsantoni, Inselspital, Bern University Hospital and University of Bern, Switzerland, Switzerland
  • Matúš Medo, Inselspital, Bern University Hospital and University of Bern, Switzerland, Switzerland


Presentation Overview: Show

SigProfilerAssignment is a popular and well-performing tool for estimating the activity of mutational signatures in sequenced samples. Although signature analysis generally becomes more accurate with increasing mutation burden in the analyzed samples, we report that this is not always the case for SigProfilerAssignment. In particular, the tool's sensitivity deteriorates in the range of mutation burden typical for whole-genome sequencing when multiple signatures with flat mutational profiles are simultaneously active. We propose a modified algorithm, FSPA, that achieves a higher sensitivity without sacrificing the high precision of SigProfilerAssignment. This modification enhances the reliability of mutational signature analysis in cancer genome studies, thereby ensuring more accurate downstream biological interpretations.

A-G.24: Biopsy-Aware Phylogenetic Reconstruction of Cancer Evolution Using Single-Cell Copy Number Profiles
Track: Genomics, epigenomics, and genome editing
  • Jarosław Paszek, Faculty of Mathematics, Informatics and Mechanics University of Warsaw, Poland
  • Agnieszka Mykowiecka, Faculty of Mathematics, Informatics and Mechanics, University of Warsaw, Poland
  • Krzysztof Gogolewski, Institute of Informatics, University of Warsaw, Poland


Presentation Overview: Show

Reconstructing tumour evolution from single-cell copy-number profiles (CNPs) collected across multiple biopsy time points requires methods that explicitly exploit the temporal ordering of biopsies. We introduce a biopsy-aware framework for simulating, reconstructing, and benchmarking copy-number phylogenies under realistic models of genome evolution. The framework treats biopsies as explicit temporal constraints and supports two complementary inference strategies. The first, biopsy-guided ancestry inference, enforces time-consistent parent–child relations and resolves the earliest branching structure using an auxiliary NJ-like step. The second, pairwise NJ-anticentral inference, iteratively selects cell pairs and designates an ancestor by combining anticentrality, plausibility, and parsimony criteria, yielding fully labelled CNP-trees. We evaluate reconstruction accuracy using ancestor–descendant recovery (AD-F1) and Generalized Robinson–Foulds (GRF) distances. Our framework enables not only the reconstruction of copy-number phylogenies but also a controlled assessment of their reliability. By simulating full evolutionary histories, we can compare inferred and true distances, quantify how tumour heterogeneity, biopsy sparsity, and multi-locus events degrade signal, and delineate the limits these factors impose on any method. Within this framework we introduced new NJ-like strategies and demonstrated, across diverse regimes, that principled use of biopsy structure together with informed pairwise ancestry decisions yields substantial improvements in reconstruction accuracy.

A-G.25: Evaluating local assembly strategies for long-read structural variant calling
Track: Genomics, epigenomics, and genome editing
  • Tim Mirus, Leibniz Institute for Immunotherapy, Germany
  • Richard Lüpken, Leibniz Institute for Immunotherapy, Germany
  • Birte Kehr, Leibniz Institute for Immunotherapy, Germany


Presentation Overview: Show

Long-read sequencing enables detection of structural variants (SVs) with far greater accuracy and completeness than was previously possible. This advance has spurred the development of many SV callers that implement diverse algorithmic strategies. Commonly, SV calling proceeds in several processing steps. Evaluations of SV callers, however, typically focus on overall performance and conceal which steps drive observed differences. This hinders targeted, evidence-based improvements to existing SV calling approaches. Here, we implemented five consensus computation strategies similar to those employed by popular SV callers that perform a local assembly. We compare their outputs on extensive simulated data and reveal that some strategies are more effective for particular input data types. For example, the strategy employed by PacBio's tool Sawfish works much better on our simulated HiFi reads than on the more noisy ONT reads. While the read error profiles and the read pre-processing has a substantial influence on some strategies, other strategies are more robust overall. These findings guide algorithmic decisions for the development of improved future SV callers.

A-G.26: Multilateration-Based Indexing and Navigation for Error-Tolerant Read Mapping
Track: Genomics, epigenomics, and genome editing
  • Luting Zhou, Shanghai Jiao Tong University, China
  • Minghao Fang, Shanghai Jiao Tong University, China
  • Yiru He, Shanghai Jiao Tong University, China
  • Cheng Wang, Shanghai Jiao Tong University, China
  • Jinpu Cai, Shanghai Jiao Tong University, China
  • Yuxuan Wang, Shanghai Jiao Tong University, China
  • Zhenwei Huang, Shanghai Jiao Tong University, China
  • Shiyang Yu, Shanghai Sixth People's Hospital Affiliated to Shanghai Jiao Tong University School of Medicine, China
  • Shiyang Ma, Shanghai Jiao Tong University, China
  • Yan Li, State Key Laboratory of Genome and Multi-omics Technologies, BGI Research, Shenzhen 518083, China, China
  • Hongyi Xin, Shanghai Jiao Tong University, China


Presentation Overview: Show

Indexing is a fundamental step in read mapping.
To maximize efficiency, ideal indexers must utilize long seeds while remaining robust to sequencing errors and genetic variations.
Existing schemes typically rely on probabilistic error tolerance at fixed sub-locations or approximate edit distances via Euclidean embeddings.
However, because edit-distance space is a metric space but not a normed space, approximations using inner-product spaces (like Euclidean space) are inherently lossy.
Here, we demonstrate that high-precision positioning is attainable in coordinate-less string space through multilateration, provided the discrete space is covered by a meticulously selected set of beacon strings.
We show that for any query, a comprehensive set of high-similarity candidates can be efficiently localized using a small subset of proximal beacons.
Leveraging these insights, we propose NavigaMer, a multi-tiered navigation system for string indexing.
NavigaMer retrieves candidate reference sequences by progressively narrowing the search space within the sequence universe, guided by beacon sets ranging from coarse-to-fine resolutions.
Experiments on simulated and real genomes demonstrate that NavigaMer is fully error-tolerant and highly accurate.
It identifies all reference candidates within a specified edit distance with minimal false positives, providing a high-sensitivity foundation for future read mappers.

A-G.27: Pathoplexus: Building a new kind of pathogen database
Track: Genomics, epigenomics, and genome editing
  • Emma Hodcroft, Swiss TPH, Allschwil; University of Basel, Basel; Swiss Bioinformatics Institute, Lausanne, Switzerland
  • Chaoran Chen, ETH Zürich, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne, Switzerland, Switzerland
  • Cornelius Roemer, University of Basel, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne, Switzerland, Switzerland
  • Anna Parker, Swiss TPH, Allschwil; University of Basel, Basel; Swiss Bioinformatics Institute, Lausanne, Switzerland
  • Anderson Brito, Instituto Todos pela Saúde, São Paulo, Brazil, Brazil
  • George Githinji, KEMRI-Wellcome Trust Research Programme, Kilifi, Kenya, Kenya
  • Senjuti Saha, Child Health Research Foundation, Dhaka, Bangladesh, Bangladesh
  • Richard Neher, University of Basel, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne, Switzerland, Switzerland
  • Tanja Stadler, ETH Zürich, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne, Switzerland, Switzerland
  • Theo Sanderson, London School of Hygiene and Tropical Medicine, London, UK, United Kingdom


Presentation Overview: Show

Sharing viral genomic data is essential for advancing scientific research and informing public health responses. Although platforms such as the International Nucleotide Sequence Database Collaboration (INSDC) and GISAID facilitate data sharing, they do not fully address several needs of the pathogen genomics community. Persistent concerns about data misuse, "scooping," and restrictions on reuse from protected repositories highlight the need for more flexible and equitable data-sharing systems.
Pathoplexus is a community-driven viral genomics database designed to balance data accessibility with submitter autonomy. Built using open-source technologies and guided by transparent governance, the platform allows data submitters to temporarily control how their data are used while still enabling rapid access for researchers and public health officials. All Pathoplexus submissions are automatically uploaded to INSDC - either immediately if fully open, or after one year if initially shared under a restricted-use option.
Pathoplexus leverages Loculus, an open-source platform for managing viral sequence databases. Its web interface and API support both interactive exploration and automated analyses. The project operates as a non-profit association governed by an international Executive Board and includes members from 14 countries across five continents, reflecting a commitment to equity and diverse public health priorities.
Despite launching only recently, Pathoplexus has already made a measurable impact. As of March 2026, the database contains more than 10,400 directly submitted sequences. It hosts the first available Sudan ebolavirus sequence from the February 2025 outbreak in Uganda. In 2025, Pathoplexus also received more mpox sequences than any other database.

A-G.28: A Fluorescent Probe Under the Microscope Showing Dual Recognition of B-DNA and G-Quadruplex DNA
Track: Genomics, epigenomics, and genome editing
  • Richard López-Corbalán, Universidad de Alcalá, Spain
  • Lorenzo Gramolini, Universidad de Alcalá, Spain
  • Cristina Garcí­a-Iriepa, Universidad de Alcalá, Spain
  • Marco Marazzi, Universidad de Alcalá, Spain


Presentation Overview: Show

The use of small molecules as fluorescent DNA markers has been extensively explored, yet designing selective probes remains challenging due to the missing relationship between probe structure and DNA conformation. To address this gap, we present a novel multistep computational protocol applied to the recently proposed fluorescent marker QCy(MeBT)3, which is capable of simultaneously recognizing both B-DNA and G-quadruplex (G4) DNA through distinct emission signatures.
Our methodology integrates molecular docking, molecular dynamics with a newly parameterized force field to capture conformational flexibility, and hybrid QM/MM simulations to compute absorption and fluorescence spectra. Our protocol successfully identifies the specific molecular conformations that enable selective probe-DNA binding following a lock-and-key recognition mechanism. We predict a remarkably high binding affinity for both DNA targets, driven by diverse energetic interactions primarily electrostatic in B-DNA minor groove binding, and a combination of electrostatic and Van der Waals interactions in G4 top-stacking.
Crucially, our simulations reveal that the distinct absorption and fluorescence shifts observed experimentally originate predominantly from the intrinsic conformational properties of the probe in water, rather than solely from specific interactions with the DNA environment. Overall, this work establishes a fundamental principle for predicting the performance of fluorescent probes in complex biological environments: the structural and electronic properties of specific molecular conformations dictate their photobiophysical fate upon binding to a given DNA sequence, thereby governing their ability to recognize specific DNA topologies.

A-G.29: SVlog: Unlocking Structural Variation in Rare Diseases through an Extensible Logic Programming Framework
Track: Genomics, epigenomics, and genome editing
  • Mikhail Gudkov, Garvan Institute of Medical Research, Australia
  • André Luiz Martins Reis, Garvan Institute of Medical Research, Australia
  • Meutia Kumaheri, Garvan Institute of Medical Research, Australia
  • Ira Deveson, Garvan Institute of Medical Research, Australia


Presentation Overview: Show

Structural variants (SVs) are a diverse group of genetic variants defined by a minimum size of 50 base pairs. SVs account for the majority of all variant bases in a person's genome, and have been frequently implicated in inherited disease and cancer. However, SV analysis is challenging due to imprecise breakpoints, variation in type and size, involvement of repetitive sequences, and general complexity of the induced changes to genomic elements. Despite recent advances in the detection and characterisation of SVs, it remains difficult to assess SVs beyond basic annotations and comparisons.

Here we introduce SVlog, a transparent and extensible meta-programming framework for analysing SVs. By utilising the logic programming language Soufflé, SVlog provides an algorithm-free, declarative ontology defining relationships among SVs and other elements. Genomic datasets are converted into relational facts, to which SVlog applies composable deterministic rules to assess SVs without relying on stochastic "black box" approaches.

Despite the compact codebase of the SVlog library, it currently evaluates more than 50 input predicates to generate over 60 informative output predicates, enabling SV annotation, comparison and prioritisation. Our tiered filtering strategy efficiently streamlines the identification of candidate pathogenic SVs in patients with rare inherited disease. Applied to our disease cohort, SVlog successfully prioritised all previously known pathogenic events, while also identifying novel candidates in several unsolved patients.

By focusing on explainability and modularity, SVlog offers a fast, reliable library for SV analysis and is a powerful deterministic alternative to traditional bioinformatics pipelines for clinical variant curation.

A-G.30: Predicting gene expression profiles and drug targets from DNA methylation in gliomas
Track: Genomics, epigenomics, and genome editing
  • Marcel Weinberg, Goethe University Frankfurt, Germany
  • Dennis Hecker, Goethe University Frankfurt, Germany
  • Nikoletta Katsaouni, Goethe University Frankfurt, Germany
  • Luca Malena Berger, Goethe University Frankfurt, Germany
  • Katharina Weber, Goethe University Frankfurt, Germany
  • Marcel Schulz, Goethe University Frankfurt, Germany


Presentation Overview: Show

Deregulation of DNA methylation is a hallmark of cancer. Studies that compare methylation profiles across groups of samples, often referred to as epigenome-wide association studies (EWAS), generate growing evidence that changes in methylation sites are of relevance as biomarkers and can help in developing disease treatments. Especially in glioma, DNA methylation profiling is routinely performed as part of the clinical diagnosis. The systematic interpretation of methylation sites can be challenging, especially for sites that are not located in the promoter region of a gene or in expressed genes despite hypermethylated promoters. Here, we reimplemented TDImpute (10.1093/gigascience/giaa076), a neural network predicting gene expression from methylation profiles trained on The Cancer Genome Atlas. We applied the model on genome-wide methylation data of glioma patients to predict expression profiles. This enabled us to identify genes whose expression patterns differ significantly between patient groups and to find candidates that associate with cancer sub-types. In addition, we ran the tool RANKOR, a machine learning model trained on drug-induced gene expression changes that ranks potential candidate drugs for each sample based on the patient expression signatures. RANKOR creates a latent space for the transcriptome profiles and the drugs' chemical structures, which makes it applicable even to unseen compounds. With this work, we aim to deepen the understanding of tumor biology by imputing gene expression profiles from DNA methylation data and translating this knowledge into potential drug targets to foster individualized treatment.

A-G.31: Unbiased deconvolution of Oxford Nanopore Technologies long reads to reconstruct methylomes at cell type level
Track: Genomics, epigenomics, and genome editing
  • Fanny Mollandin, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology, Barcelona, Spain
  • Laura Moutard, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology, Barcelona, Spain
  • Merce Planas-Felix, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology, Barcelona, Spain
  • Jakob Admard, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen, Germany
  • Eva Bru-Tari, Department of Genetic Medicine and Development, iGE3 and Centre facultaire du diabète, University of Geneva, Geneva, Switzerland
  • Javier Garci­a-Hurtado, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology, Barcelona, Spain
  • Pere Santamaria, Institut D'Investigacions Biomediques August Pi i Sunyer, Barcelona, Spain
  • Pedro Herrera, Department of Genetic Medicine and Development, iGE3 and Centre facultaire du diabète, University of Geneva, Geneva, Switzerland
  • Stephan Ossowski, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen, Germany
  • Jorge Ferrer, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology, Barcelona, Spain


Presentation Overview: Show

CpG methylation is a pivotal component of epigenetic landscapes that shape cell-specific transcriptional programs. Oxford Nanopore Technologies (ONT) long read sequencing has emerged as a powerful method to simultaneously profile 5-methylcytosine (5mC) as well as 5-hydroxymethylcytosine (5hmC), an intermediate of the active demethylation pathway. Current methods that profile 5mC and 5hmC are largely restricted to bulk tissue sequencing, often hiding cell-specific signals in heterogenous tissues.
To tackle this limitation, we developed a novel genome-wide unbiased deconvolution framework that clusters reads based on their 5mC and 5hmC patterns, annotates cell-type specific clusters, and infers cell type-specific methylome and hydroxymethylome. We evaluated the algorithm on in-silico mixtures of purified cells and highly heterogeneous samples. We further applied the algorithm on human pancreatic islets, and investigated cell specific changes of 5mC and 5hmC across age, disease and exposure to variable glucose concentrations. In sum, this ONT-based methodology offers an opportunity to assess cell-specific 5mC and 5hmC methylomes in complex tissue samples.

A-G.32: SIEVE: Sparse Interpretable Exome Variant Explainer
Track: Genomics, epigenomics, and genome editing
  • Davide Bagordo, Department of Biology and Biotechnology "L. Spallanzani", University of Pavia, Italy
  • Cezar Grigorean, Department of Biology and Biotechnology "L. Spallanzani", University of Pavia, Italy
  • Francesco Lescai, Department of Biology and Biotechnology "L. Spallanzani", University of Pavia, Italy


Presentation Overview: Show

Deep learning methods for case-control variant discovery depend on functional annotations, treat variants as unordered sets, and usually restrict the analysis to either rare or common variants. These choices leave fundamental questions unaddressed: would the model discover different biology without annotations? Does genomic position carry signal that permutation-invariant architectures discard? What is missed when common and rare variants are not modelled jointly? We present SIEVE, a framework addressing five methodological gaps: (1) sinusoidal positional encoding enables learning of position-dependent relationships; (2) attribution regularisation encourages sparse, stable variant rankings during training; (3) an annotation-ablation protocol trains the same architecture at four different annotation levels comparing discoveries systematically; (4) frequency-agnostic processing analyses the full allele-frequency; (5) epistatic interactions are intrinsically estimated from the attention layer at both the individual-level and variant-level, as well as by collapsing variant attributions at gene-level. We applied SIEVE to a coronary artery disease whole-exome cohort from the Ottawa Heart Genomics Study. Classification performance is consistent with expectations from exome variants, and null-baseline permutation confirms that attributions exceed chance. Pairwise Jaccard overlap across ablation-annotation levels is low (0.03–0.07), confirming that each level discovers genuinely different biology. Annotation-free models identify genes implicated in TNF signalling, ubiquitin-proteasome regulation, and endothelial apoptosis, i.e. pathways with established cardiovascular role. Adding positional encoding alone shifts top-ranked genes toward HIF-1 signalling, FOXO pathway, and mitochondrial import. Top-ranked attributions span the full MAF spectrum, confirming joint modelling of rare and common variation. The workflow is implemented as a reproducible open-source Nextflow pipeline.

A-G.33: Impact of AI-assisted workflow standardization in a bioinformatics core facility
Track: Genomics, epigenomics, and genome editing
  • Jan Meier-Kolthoff, University of Augsburg, Augsburg Bioinformatics Core Facility, Germany


Presentation Overview: Show

Bioinformatics core facilities often face a growing demand for reproducible, scalable, and rapidly deployable data analysis solutions across diverse project types. At the same time, many analyses are still implemented in an ad hoc manner, leading to inconsistencies in workflow structure, documentation, and long-term maintainability. Here, we describe the impact of AI-assisted workflow standardization in a university bioinformatics core facility, with a focus on the use of Nextflow as a unifying workflow framework and large language model (LLM)-based tools such as Claude Code and ChatGPT for rapid script and workflow generation.

We evaluated how AI-assisted development supported the design, refactoring, and documentation of standardized pipelines for common bioinformatics tasks, including data preprocessing, quality control, variant analysis, and downstream summarization. LLM-based tools were particularly useful for accelerating the generation of boilerplate workflow components, helper scripts, configuration templates, and documentation drafts, while human experts retained full responsibility for validation, benchmarking, and domain-specific adaptation. The combination of Nextflow-based standardization and AI-assisted implementation reduced development time, improved consistency across projects, and lowered the barrier for converting one-off analyses into reusable workflows.

Our experience suggests that AI-assisted workflow engineering can substantially improve operational efficiency in bioinformatics service environments when embedded in a controlled expert-driven framework. We argue that this approach is especially valuable for core facilities, where reproducibility, maintainability, and rapid turnaround are equally important.

A-G.34: NANO-SCOPE: Read-level tumor detection from cfDNA methylomes by Nanopore sequencing
Track: Genomics, epigenomics, and genome editing
  • Danial Tabbakh, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
  • Francesca Maria Stefanizzi, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
  • Camille Cialini, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
  • Petr Nazarov, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
  • Katrin Frauenknecht, Institute for Neuropathology, University Medical Center of the Johannes Gutenberg University Mainz, Germany, Germany
  • Vladimir Despotovic, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
  • Anna Golebiewska, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
  • Urszula Lawrynowicz, Medical University of Gdańsk, Poland, Poland
  • Reka Toth, Luxembourg Institute of Health, Strassen, Luxembourg,


Presentation Overview: Show

Nanopore sequencing enables direct, conversion-free detection of DNA methylation from low-input circulating cell-free DNA (cfDNA), making it a promising platform for liquid biopsy analysis. In cancer, cfDNA contains a complex mixture of tumor-derived and non-tumor fragments, where tumor-associated methylation patterns reflect tissue of origin and molecular subtype. Robust inference from these data remains challenging due to signal sparsity and the heterogeneous composition of individual samples. To address this, we developed NANO-SCOPE, a deep learning framework for distinguishing tumor-derived from normal-origin cfDNA fragments at the read level. The method combines two complementary components. First, a methylation-aware transformer based on DNABERT-2 generates contextual read embeddings by integrating sequence, methylation, and genomic position information. Second, a one-class autoencoder trained exclusively on non-tumor cfDNA learns the manifold of normal fragments and assigns each read a reconstruction score (RS) that quantifies deviation from this normal profile. An attention-based aggregation module then integrates read embeddings and RS values for robust sample-level predictions of tumor presence.
Applied to Nanopore cfDNA datasets from patients with lung cancer, NANO-SCOPE distinguished cancer cases from healthy controls and enabled origin-aware filtering of tumor-associated reads. This filtering enriched the tumor-related methylation signal and improved the interpretability of downstream analyses, including tissue-of-origin inference and molecular subtype prediction. NANO-SCOPE provides a scalable framework for sensitive liquid biopsy analysis with potential applications in early detection, relapse monitoring, and molecular stratification.

A-G.35: Modeling Longitudinal Cancer Dynamics from Omics Data: a Systematic Review of Current Limitations and the Road Toward Temporal Generation
Track: Genomics, epigenomics, and genome editing
  • Guillermo Prol Castelo, Barcelona Supercomputing Center (BSC-CNS), Spain
  • Davide Cirillo, Barcelona Supercomputing Center, Spain
  • Alfonso Valencia, Barcelona Supercomputing Centre BSC, Spain


Presentation Overview: Show

The variational autoencoder (VAE) has become a prominent tool in cancer omics research, valued for its ability to learn structured embeddings of high-dimensional, heterogeneous biological data. Yet a systematic understanding of how VAEs and related deep representation learning (DRL) methods engage with the temporal aspects of cancer in omics data to capture longitudinal molecular dynamics of tumor progression has been lacking. We present such an assessment. Screening 440 publications from 2014 to 2024, 21 directly relevant studies were identified. DRL methods are predominantly applied to cancer omics data for subtyping, diagnosis, and prognosis. However, these tasks do not explicitly leverage the longitudinal molecular dimension of disease progression. Longitudinal omics studies of primary cancer are scarce, constrained by practical, ethical, and biological barriers, including the destructive nature of sequencing technologies and inter-patient heterogeneity. Temporal dimension in the omics data are most commonly represented through pseudo-time inference with single-cell data, not necessarily reflecting longitudinal molecular dynamics. Alternatively, cancer stages may be used as a proxy time axis. Against this landscape, the VAE emerges as uniquely suited to bridge the gap between available cross-sectional omics data and the need for longitudinal molecular modeling. Its latent space is directly operable: sample representations can be interpolated and fed to the decoder to produce realistic intermediate omics profiles. We discuss the methodological requirements needed to unlock this potential, and briefly present an application focusing on generating and forecasting synthetic cancer stage trajectories from transcriptomic data in renal cell carcinoma.

A-G.36: TOGA2 delivers scalable and precise gene annotation and ortholog identification across vertebrate genomes
Track: Genomics, epigenomics, and genome editing
  • Yury Malovichko, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Bernhard Bein, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Alejandro Gonzales-Irribarren, Senckenberg Research Institute; Department of Computer Science, Leipzig University, Germany
  • Evgeny Leushkin, Senckenberg Research Institute, Germany
  • Leon Hilgers, Senckenberg Research Institute, Germany
  • Amy Stephen, Senckenberg Research Institute, Germany
  • Xueling Yi, Senckenberg Research Institute, Germany
  • Michele Albertini, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Tim Stadager, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Markus Zumpt, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Luca Hoppach, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Felix Götz, Senckenberg Research Institute, Germany
  • Niklas Himstedt, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Lucas Koch, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
  • Michael Hiller, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany


Presentation Overview: Show

Accurate annotation of coding genes and inference of orthology relationships in newly sequenced genomes remain central challenges in modern genomics. We present TOGA2, the next generation of the TOGA (Tool to infer Orthologs from Genome Alignments) framework for scalable reference-based orthologous gene annotation. By introducing an exon-wise gene annotation approach, TOGA2 achieves a 513-fold memory reduction and a 6-fold runtime reduction compared to its predecessor, while improving exon-level annotation precision. We demonstrate that deep learning models for splice-site prediction trained on human data remain effective across diverse vertebrates. Incorporating these predictions further enhances exon boundary annotation precision in TOGA2 and enables detection of evolutionary changes affecting gene exon–intron structure, such as splice-site shifts or intron gains and losses. Multiple false-positive prediction filters and a new gene tree–based reconciliation step further improve ortholog inference, specifically identifying new 1:1 ortholog pairs. We further show how features integrated into TOGA2 improve annotation of immunoglobulin and T-cell receptor gene segments, processed pseudogenes, and functional retrogenes, and how synteny information from inferred orthologs can be leveraged for ancestral chromosome reconstruction and phylogenomic analyses. To demonstrate scalability across multiple genomes, we provide a comprehensive comparative genomics resource for >900 mammal and >680 bird assemblies, including gene annotations, ortholog sets, retrogene candidates, and codon alignments.

A-G.37: GAP-MS: Automated validation of gene predictions using integrated mass spectrometry evidence
Track: Genomics, epigenomics, and genome editing
  • Qussai Abbas, Technical University of Munich, Germany
  • Mathias Wilhelm, Technical University of Munich, Germany
  • Bernhard Kuster, Technical University of Munich, Germany
  • Dimitri Frishman, Technical University of Munich, Germany


Presentation Overview: Show

Accurate genome annotation is fundamental to modern biology, yet distinguishing authentic protein-coding sequences from prediction artifacts remains challenging, particularly in complex plant genomes. We present GAP-MS, an automated proteogenomic pipeline that leverages mass spectrometry evidence to systematically validate the protein-level accuracy of predicted gene models. Applied across 9 major crop species, GAP-MS consistently improved prediction precision for four widely used tools, with substantial gains of up to 32%. Beyond filtering artifacts, the pipeline identified over 9000 novel peptide-supported gene models absent from current RefSeq annotations. Furthermore, GAP-MS successfully corrected structural errors by identifying independent translation initiation and termination sites via specific N- and C-terminal peptides. These results demonstrate that direct proteomic evidence provides a robust framework for resolving annotation ambiguities and defining high-confidence reference proteomes.

A-G.38: Embedding-based statistical framework enables quantitative gene function representation and hypothesis testing with large language models
Track: Genomics, epigenomics, and genome editing
  • Yanhao Tan, UPMC Hillman Cancer Center, University of Pittsburgh, United States
  • Li-Ju Wang, UPMC Hillman Cancer Center, University of Pittsburgh, United States
  • Yu-Chiao Chiu, UPMC Hillman Cancer Center, University of Pittsburgh, United States


Presentation Overview: Show

Accurately delineating gene function is central to interpreting high-throughput genomic data, yet most existing methods depend on predefined gene sets and largely qualitative interpretations. Large language models (LLMs) provide a promising alternative by extracting functional relationships from biological text, though their ability to support quantitative analysis has not been fully established. We introduce a statistical framework that leverages LLM-derived embeddings of genes and biological functions to enable quantitative assessment of gene-gene and gene-function relationships across diverse biological settings. We evaluated seven leading embedding models using both curated gene annotations and literature-derived descriptions. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 showed the strongest performance, recovering gene-gene relationships in up to 98% of Gene Ontology biological processes and approximately 99% of canonical pathways. Gene-function association analyses further demonstrated high sensitivity (95-98%) and specificity (73-84%). Importantly, this framework enables rigorous testing of functional hypotheses generated from gene lists without requiring predefined annotations. It consistently differentiates biologically coherent gene sets from noise, surpassing both confidence-based LLM outputs and traditional enrichment methods. Application to drug response data uncovered candidate pathways linked to cancer immune sensitization and supported systematic exploration of drug mechanisms. Overall, our results position LLM-based embeddings as a scalable quantitative tool for functional genomics.

A-G.39: Unmasking the Trypanosoma cruzi Genome: A Comprehensive Map of the MASP Superfamily.
Track: Genomics, epigenomics, and genome editing
  • Aldana Alexandra Cepeda Dean, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET, Argentina
  • Luisa Berná, Laboratorio de Genómica Evolutiva, Facultad de Ciencias, Universidad de la República, Montevideo, Uruguay
  • Javier De Gaudenzi, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET, Argentina
  • Virginia Balouz, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET, Argentina
  • Carlos Andrés Buscaglia, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET, Argentina


Presentation Overview: Show

Trypanosoma cruzi, the causative agent of Chagas disease, harbors one of the most complex eukaryotic genomes, dominated by massive expansion and diversification of multigene families encoding surface virulence factors. Among them, mucin-associated surface proteins (MASPs) represent a paradigmatic case of genomic plasticity, yet their true diversity has remained obscured by fragmented assemblies and the limitations of standard annotation pipelines.
We developed a dedicated bioinformatic framework combining a curated sequence database, HMM-based motif detection, and clustering with CD-HIT at 80% similarity to enable de novo annotation and classification of MASP sequences across 13 T. cruzi strains spanning all major evolutionary lineages. This approach uncovered ~1000 canonical MASP genes, ~30 chimeric variants, and >500 pseudogenes depending on lineage, outperforming conventional pipelines. Comparative analyses revealed marked differences in MASP repertoires, with lineage-specific innovations, tandem duplication-driven expansions, and a conserved core of MASP sequences shared across strains.
Integration of RNA-seq data showed that over 70% of the MASP repertoire is transcriptionally active in human-infective stages. This proportion is unexpectedly high, indicating that nearly the entire family contributes to antigenic diversity and immune evasion. Distinct transcriptomic signatures further separate canonical and chimeric variants, as well as MASP genes from pseudogenes.
Our findings highlight MASPs as key contributors to T. cruzi immune evasion, while demonstrating how tailored computational strategies can resolve highly repetitive, structurally complex genomes. The framework is generalizable for dissecting large gene families in pathogens, bridging genotype to phenotype and advancing our understanding of genome evolution, plasticity, and host-pathogen interaction.

A-G.40: NeoEvo-AIS: Rethinking Tumor Immunogenicity at the Tumor Level
Track: Genomics, epigenomics, and genome editing
  • Ankita Singh, Center for Applied and Translational Genomics (CATG) . MBRU, Dubai Health, Dubai, UAE, United Arab Emirates
  • Youssef Ahmed Elkenawi, Center for Applied and Translational Genomics (CATG) . MBRU, Dubai Health, Dubai, UAE, United Arab Emirates
  • Costerwell Khyriem, University of Freiburg, Germany
  • Sven Hauns, University of Freiburg, Germany
  • Filippo Castiglione, Institute for Applied Computing (IAC) National Research Council of Italy (CNR),Rome, Italy
  • Rolf Backofen, University of Freiburg, Germany
  • Omer S. Alkhnbashi, Mohammed Bin Rashid University of Medicine and Health Sciences (MBRU), United Arab Emirates


Presentation Overview: Show

Tumor mutation burden (TMB) and predicted neoantigen counts are often used as measures to estimate tumor immunogenicity; however, they simplify a complex biological landscape into a single metric. By mainly focusing on mutation quantity, these measures often miss important aspects such as clonality, antigen quality, and intratumoral diversity, which can lead to an inability to distinguish tumors with similar mutation burdens but significantly different immune responses.

To overcome these limitations, we have developed NeoEvo-AIS, a computational framework designed to provide a more comprehensive view of tumor antigenic properties. Central to this framework is the Antigenic Instability Score (AIS-STATIC), a tumor-level metric that combines biologically relevant features. Using pan-cancer somatic mutation data from The Cancer Genome Atlas (TCGA) MC3 harmonized call set, we assessed mutation architecture and inferred clonality by incorporating variant allele frequency alongside estimates of copy number and tumor purity. Candidate neoantigen peptides were predicted with NetMHCpan-4.1 and MHCflurry, while immunogenicity was evaluated through deep learning models that consider peptide foreignness and structural features important for T-cell recognition. Antigenic diversity was measured using entropy-based metrics across mutated genes and inferred clonal populations.

By integrating these features into a single interpretable score, AIS-STATIC effectively captures major differences in tumor antigenic architecture that are not solely reflected by TMB. Across various tumor types, those with higher AIS-STATIC scores show increased immune activity and greater immune cell infiltration. Therefore, NeoEvo-AIS offers a practical framework for studying tumor-immune interactions and sets the stage for future research into how antigenic properties evolve over time.

A-G.41: TFClassPredict: A Deep Learning Framework for Transcription Factor Binding Site Analysis Using Evolutionarily Conserved DNA-Binding Domain Annotations
Track: Genomics, epigenomics, and genome editing
  • Christian Ickes, Institute of Cardiovascular Physiology, University Medical Center Göttingen, Germany, Germany
  • Cigdem Hazal Timucin, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
  • Bendix Christian Harms, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
  • Umut Akgül, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
  • Inigo Vincente Hernandez, Department of Gastroenterology, University Medical Center Göttingen, Germany, Germany
  • Ivan Bogeski, Institute of Cardiovascular Physiology, University Medical Center Göttingen, Germany, Germany
  • Tim Beißbarth, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
  • Martin Haubrock, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany


Presentation Overview: Show

Transcription factors (TFs), the core proteins of transcription initiating processes, regulate gene expression by binding to short genomic sequences, known as transcription factor binding site (TFBS), through defined DNA-binding domains (DBDs). These interactions are central to understand gene regulation, which underlies many cellular processes. Although many models exist for predicting transcription factor binding, none provide a comprehensive framework that accounts for the similarity of binding characteristics among TFs within the same DBD class. In fact, several of these families are severely under represented in most current analysis pipelines.

Our model, TFClassPredict, introduces a novel approach to identify transcription factor binding sites based on structural annotations of evolutionarily conserved DBDs. By leveraging canonical binding patterns, TFClassPredict provides high-confidence predictions essential for gene regulation analysis. Fine-tuned from the DNABERT model, TFClassPredict classifies DNA sequences across 23 classes of DBDs. TFClassPredict achieved strong performances, naturally aggregates predicted sites by DBD class, improving both interpretability and the balanced representation of all DBD classes. These results demonstrate that TFClassPredict constitutes a reliable, family aware framework for uncovering regulatory differences and for advancing comprehensive TF binding analyses.

A-G.42: End-to-end deep learning methods for genetic risk prediction of Schizophrenia
Track: Genomics, epigenomics, and genome editing
  • Nora Verplaetse, KULeuven, Belgium
  • Yves Moreau, Katholieke Universiteit Leuven, Belgium
  • Daniele Raimondi, Institut de Génétique Moléculaire de Montpellier, France


Presentation Overview: Show

Schizophrenia is a highly heritable psychiatric disorder with a complex genetic basis. While recent Genome Wide Association Studies (GWAS) and Whole Exome Sequencing (WES) have identified numerous risk loci, existing clinical prediction models rely primarily on linear assumptions, limiting their ability to capture complex, nonlinear genetic effects such as epistasis.

In this study, we apply a new paradigm of end-to-end Genome Interpretation (GI) Neural Network (NN) models to predict Schizophrenia risk from WES data in a large cohort of 6,135 cases and 6,245 controls. We show that nonlinear NNs significantly outperform conventional additive models, when sufficient sample size is available. These findings further support the fact that high-order genetic interactions between alleles and variants should be considered by clinical and quantitative genetics models.

To investigate the decision process our models follow, we integrate Explainable AI (XAI) techniques and biological priors into our models, using them to identify predictive genes and pathways. Our approach recovers both known Schizophrenia risk genes and recommends BASP1 as a potential understudied Schizophrenia gene involved in neuronal development.

A-G.43: Contrastive Learning of Multiple Sequence Alignments for Phylogenetic Inference
Track: Genomics, epigenomics, and genome editing
  • Jens-Uwe Ulrich, Robert Koch Institute, Germany
  • Denise Kühnert, Robert Koch Institute, Germany


Presentation Overview: Show

Phylogenetic inference from large-scale multiple sequence alignments (MSAs) is essential for understanding evolutionary relationships, tracking pathogen outbreak dynamics, and quantifying biodiversity patterns across the tree of life. Traditional distance-based and likelihood-based methods struggle to scale efficiently while capturing the complex hierarchical relationships inherent in evolutionary data. Thus, low dimensional representations of MSAs are needed to infer phylogenies from large MSAs. We present a self-supervised contrastive learning framework that learns compact, phylogenetically informative representations of MSA patches without requiring explicit phylogenetic labels.
Our architecture combines a multi-scale convolutional encoder with biologically motivated data augmentations tailored to MSA characteristics. The encoder processes local alignment patches and projects them into a normalized latent space optimized via an NT-Xent loss. Crucially, we designed domain-specific augmentations that reflect biological invariances: row shuffling captures sequence order independence, column masking simulates alignment uncertainty and missing data, row subsampling models varying taxonomic sampling, and reverse complementation accounts for strand ambiguity. This augmentation strategy enforces robust feature learning while maintaining phylogenetically relevant sequence information.
We evaluated our framework assessing multiple aspects of representation quality: hierarchical clustering fidelity through cophenetic correlation, discrete evolutionary group recovery via k-nearest neighbor classification, continuous sequence similarity preservation through correlation analysis, and latent space geometry metrics including uniformity and intrinsic dimensionality. Our analysis revealed fundamental trade-offs between learning discrete phylogenetic groupings versus preserving fine-grained evolutionary distances.
This work establishes contrastive learning as a viable approach for phylogenetic representation learning, opening avenues for scalable phylogenetic placement, distance matrix approximation, and integration with downstream evolutionary analyses.

A-G.44: PanelClone: signature-aware subclonal reconstruction that scales with mutation count
Track: Genomics, epigenomics, and genome editing
  • Sungjin Park, Institute of Human Behavior & Genetics, College of Medicine, Korea University, Seoul, Republic of Korea, South Korea
  • Sangwon Um, Institute of Human Behavior & Genetics, College of Medicine, Korea University, South Korea
  • Taehun Kim, AI Center, Advanced Medical Imaging Institute, Anam Hospital, College of Medicine, Korea University, South Korea
  • Hyowon Lee, DigitalBio R&D Center, Korea University Anam Hospital, South Korea
  • Cheol Soon Lee, Institute of Human Behavior & Genetics, College of Medicine, Korea University, South Korea


Presentation Overview: Show

Subclonal reconstruction is essential for understanding tumor evolution and predicting treatment resistance, yet reliable deconvolution has required the hundreds-to-thousands of somatic mutations available only from whole-genome/exome sequencing—leaving clinical deep-panel assays (300–500 genes; 5–50 mutations/sample) underserved. We present PanelClone, a single-sample framework that maximizes subclonal resolution by jointly exploiting two orthogonal signals—variant frequency and mutational signature—and by adapting inference stringency to the available mutation count. PanelClone combines three components. (1) Multiplicity-aware CCF inference with explicit neutral-tail modeling: a beta-binomial mixture corrects for purity and local copy number, marginalizes mutation multiplicity, and models neutral passenger mutations as a truncated 1/f power-law tail, suppressing spurious (“ghost”) subclones. (2) A dual-axis ghost gate: because distinct subclones often arise under distinct mutational processes, PanelClone treats each mutation’s trinucleotide signature as an axis orthogonal to frequency, letting it separate subclones that overlap in CCF and rescue signature-distinct subclones buried in the neutral tail—cases that frequency-only methods (PyClone-VI, MOBSTER) merge or miss and that signature-only methods (CloneSig) over-call. (3) Tiered, mutation-count-adaptive selection: sparse Bayesian (MAP-EM) clustering with automatic cluster-number determination that degrades gracefully from full clustering to binary clonal/subclonal classification to individual-CCF reporting as counts fall. In simulation (BAMSurgeon-style data with injected copy-number, purity and trinucleotide-signature structure), evaluated with SMC-Het metrics (V-measure, CCF MAD, co-clustering), the dual-axis design recovers subclones that frequency-only baselines miss and avoids the over-segmentation of signature-only approaches, with the advantage concentrated where subclones differ in mutational process. Extension to clinical panel scale—including population-based haplotype phasing across panel targets and paired bulk RNA-seq for genome-wide copy-number and allele-specific-expression validation—and benchmarking against matched multi-region cohorts are underway. PanelClone will be made publicly available upon publication.

A-G.45: Large cohorts analysis to identify pathogenic digenic interactions in autoinflammatory disorders
Track: Genomics, epigenomics, and genome editing
  • Jaume Reig-Palou, Universitat de Barcelona, Spain
  • Diego Garrido-Martin, Universitat de Barcelona, Spain
  • Llorenc Villalonga-Gonyalons, Universitat Pompeu Fabra, Spain
  • Alex Velaza-Gil, Universitat de Barcelona, Spain
  • Nerea Moreno-Ruiz, Universitat Pompeu Fabra, Spain
  • Roderic Guigo, Centre for Genomic Regulation, Spain
  • Hafid Laayouni, Universitat Pompeu Fabra, Spain
  • Juan Ignacio Arostegui, Hospital Clínic de Barcelona, Spain
  • Ferran Casals, Universitat de Barcelona, Spain


Presentation Overview: Show

The role of the digenic model in rare disorders has been relatively under explored, mostly due to both the difficulties of statistically supporting these associations to disease and the limitations of generating experimental models. We have developed an approach based on the analysis of genomic data from hundreds of thousands of individuals available from large biomedical projects to detect or validate pathogenic digenic interactions.

We analyzed genotyping data for 70 autoinflammatory disorder genes in UK Biobank for 467,000 European ancestry individuals based on PCA analyses. We compared the expected and observed frequencies of co-carriers for all possible pairs of variants from different genes.

The analysis revealed candidate digenic interactions among AID genes, evidenced by deviations from neutrality of the genotypes distribution. One of the interactions, validated in an independent patient cohort, may provide insight into the incomplete penetrance of Familial Mediterranean Fever (FMF) where heterozygous individuals carrying pathogenic variants in the MEFV gene may or may not develop the disease. Specifically, we suggest that a genetic variant in a functionally related gene acts as a modifier, influencing disease expression by modifying the action of the primary pathogenic variant in MEFV.

Our approach can identify digenic and modifier interactions in rare diseases. We have identified a candidate modifier gene that may explain incomplete penetrance in FMF.

A-G.46: PlanMine: An Interactive Database for Exploring Planarian Genomics and Regeneration
Track: Genomics, epigenomics, and genome editing
  • Muhammad Rizwan Riaz, Max Planck Institute for Multidisciplinary Sciences, Germany
  • Hongkee Moon, Max Planck Institute of Molecular Cell Biology and Genetics Dresden, Germany
  • Jeremias Brand, Max Planck Institute for Multidisciplinary Sciences Göttingen, Germany
  • Martin Leandro Paleico, Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen, Germany
  • Julian Kunkel, Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG), Germany
  • Jochen Rink, Max Planck Institute for Multidisciplinary Sciences Göttingen, Germany


Presentation Overview: Show

Schmidtea mediterranea is a free-living freshwater planarian and a powerful model for studying regeneration, stem cell biology, and the evolution of complex traits. As an early-diverging bilaterian, it combines simple morphology with features of complex organisms, including centralized brain, nervous system, and coordinated behaviors. Over the past decade, high-throughput sequencing has generated extensive molecular resources, including high-quality genomes, transcriptomes, and expression datasets. These developments underscore the need for an integrated data exploration platform for the planarian research community. PlanMine has been developed by our lab to address this need and to evolve alongside the advancements in the model system.

PlanMine is built on the InterMine framework, enabling the integration of diverse datasets along with flexible tools for querying and visualization. It uses a PostgreSQL backend, incorporates custom BLAST databases via SequenceServer, and supports genome navigation through JBrowse and the UCSC Genome Browser. Together, these components form a unified, extensible platform for the planarian research community.

PlanMine offers a comprehensive, user-friendly interface for accessing and analyzing planarian molecular data. The updated version (v4.0- projected release: October 2026) incorporates the latest chromosome-scale genome assembly and annotations of the model species, S. mediterranea, along with additional planarian transcriptomes. Users can explore gene models between the new and legacy annotations, homologues, domain content, pathways, and bulk and single-cell expression patterns derived from experimental, computational, and curated sources. By centralizing data, advanced search and visualization, PlanMine enables integrative analyses and hypothesis generation and thus valuable services for the further development of planarians as model taxon.

A-G.47: Genome Enhancer reveals an IGF-centered regulatory program underlying shared epigenetic mechanisms in Kabuki syndrome subtypes
Track: Genomics, epigenomics, and genome editing
  • Daria Stelmashenko, geneXplain GmbH, Wolfenbuettel, Germany; Center for Integrative Biology (CIBIO), University of Trento, Trento, Italy, Germany
  • Kel Alexander, geneXplain GmbH, Wolfenbuettel, Germany; Royal College of Surgeons in Ireland (RCSI), Dublin, Ireland, Germany
  • Giuseppe Merla, Department of Molecular Medicine and Medical Biotechnology, University of Naples Federico II, Naples, Italy, Italy
  • Emile Danvin, Department of Molecular Medicine and Medical Biotechnology, University of Naples Federico II, Naples, Italy, Italy
  • Antonio Ammendola, Department of Molecular Medicine and Medical Biotechnology, University of Naples Federico II, Naples, Italy, Italy


Presentation Overview: Show

Genome-wide DNA methylation data provide an indirect readout of regulatory activity, requiring translation of CpG-level changes into mechanistic models. Here, we present Genome Enhancer, an automated pipeline that converts epigenomic profiles into causal networks and identifies master regulators by integrating cgID-to-gene signature conversion, TFBS enrichment, composite module detection, and upstream network reconstruction, thereby linking epigenetic changes to transcription factors, signaling pathways, and candidate targets; its application is demonstrated in uncovering shared regulatory mechanisms in Kabuki syndrome subtypes.

Kabuki syndrome is a rare developmental disorder caused by LOF mutations in KMT2D (KS1) and KDM6A (KS2). By independently analyzing KS1 and KS2 blood DNA methylation data against paired healthy controls, we identified 65 hyper- and 49 hypomethylated genes shared between subtypes. These shared signatures were analyzed with Genome Enhancer using TRANSFAC-based TFBS enrichment, CompositeModuleAnalyst for combinatorial modules, and TRANSPATH for upstream network reconstruction. The analysis revealed a common IGF-centered upstream regulatory program in which IGF1R, IGF1,IGF2, and IGFBP4 activate downstream modules mediated by SMAD3,TEAD1, and RAD21, linking growth factor signaling to enhancer-driven gene regulation. In the context of KMT2D/KDM6A dysfunction, this machinery is likely impaired, leading to reduced transmission of IGF-dependent developmental signals into stable transcriptional activation and resulting in coordinated hypermethylation and silencing of key developmental genes. This connects methylation findings directly to one of the core clinical dimensions of Kabuki: abnormal growth and developmental signaling, suggesting that KS1 and KS2 may converge not only at the phenotype level, but also at the level of a common growth-factor-dependent transcriptional control architecture.

A-G.48: Improved reconstruction of transcripts and coding sequences from RNA-seq data
Track: Genomics, epigenomics, and genome editing
  • Jan Grau, Institute of Computer Science, Martin Luther University Halle-Wittenberg, Germany
  • Deborah Weise, Institute of Computer Science, Martin Luther University Halle-Wittenberg, Germany
  • Marika Panster, Institute of Biology, Martin Luther University Halle-Wittenberg, Germany
  • Martin H Schattat, Institute of Biology, Martin Luther University Halle-Wittenberg, Germany
  • Jens Keilwagen, Julius Kühn-Institut (JKI) - Federal Research Centre for Cultivated Plants, Germany


Presentation Overview: Show

Accurate annotation of gene and transcript models in newly sequenced genomes is a pivotal requirement for many subsequent analyses. We present GeMoSeq, a novel approach for RNA-seq-based gene prediction. GeMoSeq shall complement homology-based (e.g., GeMoMa) predictions and, hence, focuses on the prediction of protein-coding genes.
Starting from genomic mappings of RNA-seq reads, GeMoSeq partitions the genome into covered regions, builds a read graph with basepair resolution connecting positions that are adjacent in mapped reads. For each connected component of the read graph, GeMoSeq then merges consecutive positions without alternative edges into a splicing graph, and combinatorially enumerates candidate transcripts. Candidate transcripts may still be chimeras of multiple transcripts, which are split considering coverage, CDS prediction and splice site orientation. Resulting potential transcripts are quantified based on RNA-seq evidence in an EM-like algorithm, filtered by abundance, and finally merged to genes.
We benchmark GeMoSeq against state-of-the-art tools on a large collection of RNA-seq libraries of seven species using the respective reference annotations as ground truth. For the F1 measure on the level of CDSs, which have been in the focus of GeMoSeq development, we observe that GeMoSeq yields better predictions than all previous approaches for almost all data sets, where the improvement is specifically pronounced for S. cerevisiae, C. elegans and A. thaliana.
We finally combine RNA-seq-based prediction of GeMoSeq with homology-based prediction of GeMoMa to reannotate two recently sequenced genomes of N. benthamiana lab strains and validate several predictions experimentally.

A-G.49: Chromatin-informed sparsification of neural networks improves gene expression prediction and regulatory interaction recovery
Track: Genomics, epigenomics, and genome editing
  • Maxim Evsioukov, Goethe University Frankfurt, Germany
  • Dennis Hecker, Goethe University Frankfurt, Germany
  • Shamim Ashrafiyan, Goethe University Frankfurt, Germany
  • Marcel Schulz, Goethe University Frankfurt, Germany


Presentation Overview: Show

Gene expression prediction from epigenomic data links regulatory elements to transcriptional output. Multilayer perceptron (MLP) models rely primarily on histone modification signals at candidate cis-regulatory elements (CREs) and often ignore genomic and 3D chromatin context. These models are often overparameterized and may interpret correlated signals as causal regulatory effects. These limitations motivate the incorporation of biological priors.

We introduce a biologically guided pruning framework to systematically evaluate sparsification strategies in gene expression prediction models. Using H3K27ac signals at CREs, we train MLPs on bulk RNA-seq data and impose sparsity via three approaches: data-driven pruning, pruning guided by linear genomic distance between CREs and transcription start sites (TSSs) of target genes, and pruning informed by Hi-C chromatin interaction frequencies between CREs and TSSs.

On a benchmark of 200 genes, Hi-C-guided pruning achieves the strongest predictive performance, improving median correlation and reducing mean squared error. Stratifying genes by predictability (threshold r = 0.7) reveals heterogeneous effects, with larger gains in low-correlation genes and smaller improvements in high-correlation genes.

We then apply this approach to 1,574 genes from an enhancer-gene interaction map and evaluate biological validity. Hi-C-pruned models improve recovery of functional CRE-gene regulatory relationships, increasing AUC from 0.159 to 0.180 for CRISPRi-validated links and showing stronger enrichment of eQTL-supported associations across multiple tissues from the GTEx Consortium, relative to the unpruned model.

These results demonstrate that chromatin-informed sparsification improves predictive performance and recovery of experimentally supported regulatory interactions.

A-G.50: Ancestry calibrated polygenic risk score for breast cancer risk stratification in an admixed Colombian cohort
Track: Genomics, epigenomics, and genome editing
  • Yina Tatiana Zambrano, Biosciences - SURA Colombia, Colombia
  • Danny Styvens Cardona Pineda, Biosciences - SURA Colombia, Colombia
  • Harvy Mauricio Velzco, Biosciences - SURA Colombia, Colombia


Presentation Overview: Show

Breast cancer (BC) polygenic risk scores (PRS) derived from European GWAS often lack transferability to admixed Latin Americans due to tri-ancestral genomic backgrounds, European (EUR), Native American (NAM), and African (AFR), which alter allele frequencies and linkage disequilibrium. Ancestry-unaware models risk systematic miscalibration, hindering equitable clinical implementation.

We enrolled 1,997 Colombian women (ages 40–65) in a translational precision medicine program at Biosciences unit Grupo SURA (2022–2024). Genotyping (Illumina GSA v3) was followed by phasing (Eagle v2.4.1) and imputation (TOPMed Freeze 8). Ancestry proportions were inferred via ADMIXTURE (K=3) using the HGDP panel. Risk modeling integrated: (i) a genome-wide PRS from established BC loci, (ii) principal components as calibration covariates, and (iii) modifiable clinical variables. To date, 32 incident BC cases have been identified among the cohort, with ~700 baseline controls under prospective follow-up.

Ancestry inference revealed a tri-ethnic profile (EUR: 0.59, NAM: 0.30, AFR: 0.11). The integrated genomic-clinical model achieved an AUC of 77.6% (OR per SD: 1.87; 95% CI: 1.68–2.07), with 72.6% accuracy, 67.7% sensitivity, and 74.3% specificity. Notably, the model yielded a Net Reclassification Improvement (NRI) of 48% and an NPV of 86.9%.

Integrating ancestry calibrated PRS with clinical factors substantially enhances BC risk discrimination in admixed populations. With 32 prospectively ascertained cases, this cohort serves as a critical resource for validating polygenic risk in Latin America, underscoring the necessity of diverse genomic data to achieve precision oncology.

A-G.51: Targeted long-read sequencing for high-resolution repeat profiling in myotonic dystrophy type 1
Track: Genomics, epigenomics, and genome editing
  • Yoojung Han, Interdisciplinary Program in Bioinformatics, Seoul National University, South Korea
  • Ja-Hyun Jang, Department of Laboratory Medicine and Genetics, Samsung Medical Center, Sungkyunkwan University School of Medicine, South Korea
  • Hyeshik Chang, School of Biological Sciences, Seoul National University, South Korea


Presentation Overview: Show

Tandem repeat expansion disorders can be difficult to diagnose when expansions exceed 200 repeats, as standard methods (for example, Southern blot and modified PCR) often fail. We present a Cas9-targeted nanopore sequencing workflow and an automated analysis pipeline, RepeatLab, for accurate repeat-length estimation, structure assessment, and high-resolution methylation profiling. Validated on 13 myotonic dystrophy type 1 samples, 4 healthy controls, and 4 cell lines, this approach demonstrates improved sensitivity and accuracy for large expansions. Key refinements include an alternative basecalling strategy for extended repeats and a repeat-length calling algorithm that remains robust at lower sequencing throughput. The platform also automatically reports methylation near the DMPK repeat region, including five CpG site groups that could inform more nuanced clinical evaluations. This integrated workflow offers a rapid, cost-effective diagnostic solution with a turnaround time under 24 h and costs comparable to standard assays. Its compatibility with readily available computational resources enhances accessibility and scalability.

A-G.52: Pattern-based segmentation of the methylome reveals chemotherapy-associated changes in ovarian cancer
Track: Genomics, epigenomics, and genome editing
  • Susanna Holmström, University of Helsinki, Finland
  • Giovanni Marchi, University of Helsinki, Finland
  • Jaana Oikkonen, University of Helsinki, Finland
  • Sampsa Hautaniemi, University of Helsinki, Finland


Presentation Overview: Show

Selective pressures, including chemotherapy, shape tumor evolution by favoring phenotypes with increased fitness. In ovarian high-grade serous carcinoma (HGSC), a subset of patients develops chemoresistant tumors, markedly reducing survival. Despite extensive genomic characterization, the mechanisms underlying chemoresistance remain only partially understood. Consequently, recent research has expanded toward alternative regulatory layers, including DNA methylation (DNAm).

DNAm regulates transcription through activation or silencing of regulatory elements. However, the specific regulatory regions affected precisely by DNAm remain challenging to be defined, and most studies focus on DNAm within pre-annotated promoters and enhancers, thereby overlooking large portions of the epigenome where aberrant regulation may occur in cancer. To address this limitation, we developed FUSE (implemented in the R package methFuse), a computational framework that identifies candidate functional regulatory elements genome-wide based on DNAm patterns. In whole-genome bisulfite sequencing (WGBS) data from healthy donors obtained via the ENCODE portal, highly stable FUSE segments were enriched in known promoters, enhancers, and repetitive elements, supporting their biological relevance.

We then applied FUSE to 191 WGBS samples from 62 HGSC patients enrolled in the DECIDER trial (ClinicalTrials.gov identifier: NCT04846933), integrating matched RNA-seq data to identify candidate regulatory regions through correlation with gene expression. Comparing pre- and post-treatment samples revealed that methylation changes in pre-annotated promoter regions provided limited insight, whereas alterations in FUSE-defined regions outside canonical promoters highlighted chemotherapy-associated effects on cell-cycle regulation and Rho GTPase signaling processes. These results demonstrate that our DNAm-based segmentation captures chemotherapy-associated changes not detectable through promoter-focused analyses.

A-G.53: Evaluating portability of polygenic scores across ancestries with precision-weighted meta-analysis reveals trait-specific rather than universal poor portability
Track: Genomics, epigenomics, and genome editing
  • Natalia Nunes, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg, Salzburg, Austria, Brazil
  • Daniel Katzlberger, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg, Salzburg, Austria, Austria
  • Iuliia Trifonova, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg, Salzburg, Austria, Austria
  • Arne Bathke, Department of Artificial Intelligence and Human Interfaces, Paris Lodron University Salzburg, Austria
  • Georg Zimmermann, Department of Artificial Intelligence and Human Interfaces, Paris Lodron University Salzburg, Austria
  • Nikolaus Fortelny, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg, Salzburg, Austria, Austria


Presentation Overview: Show

Polygenic scores (PGS) are increasingly proposed for clinical risk stratification, yet poor portability is often reported when models trained in Europeans are applied to non-European populations. Most comparisons, however, compare large European cohorts against much smaller non-European cohorts, generating wide confidence intervals in non-Europeans and masking whether observed performance gaps reflect genuine biology or sampling noise.
We developed a series of computational approaches to systematically and rigorously evaluate the portability of PGS across ancestries. We first compared methods to obtain confidence interval on data from the Parkinson's Progression Markers Initiative and then developed an inverse-variance weighted meta-analysis framework, which pools AUCs across test cohorts per PGS, weighting by precision and excluding scores with substantial heterogeneity (I² > 80%), and then quantified portability as a ΔAUC per trait and ancestry. This approach enabled us to assess PGS portability across ~3,900 evaluations from the PGS Catalog. We then validated these observations in the All of Us Research Program, an NIH-funded biobank linking genotypes and electronic health records across ~245,000 diverse participants, using bootstrap confidence intervals.
In summary, we find systematic baseline genetic differences between ancestries. Yet, the mean ΔAUC in All of Us was only −0.02, and clear evidence of poor portability was confined to a restricted set of traits: asthma in African and East Asian ancestries, type 2 diabetes and coronary heart disease in African ancestry. Our results suggest that concerns about poor portability is better framed as trait-dependent rather than universal, emphasizing the value of imbalance-aware evaluation before clinical deployment.

A-G.54: Microsynteny and functional relatedness: insights into eukaryotic genome evolution
Track: Genomics, epigenomics, and genome editing
  • Silvia Prieto Banos, University of Lausanne, Switzerland
  • Alex Warwick Vesztrocy, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
  • Natasha Glover, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
  • Christophe Dessimoz, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

Genome rearrangements have shaped eukaryotic evolution, resulting in extensive variation in genome organization and limited conservation of gene order across distant species. While prokaryotic operons and eukaryotic gene clusters suggest that neighboring genes are often functionally related, the relationship between gene function and gene order conservation (microsynteny) remains poorly understood in eukaryotes.

Here, we investigate the evolutionary conservation of gene microsynteny and its association with functional similarity across a diverse set of eukaryotic genomes. Using Gene Ontology (GO) similarity measures, we show that adjacent genes are significantly more functionally related than expected by chance. Leveraging the unprecedented wealth of publicly available genomes and edgeHOG, a method for reconstructing gene order across the eukaryotic tree of life, we systematically quantify the conservation of gene adjacencies in both extant and ancestral genomes. This enables us to date the emergence of conserved adjacencies and assess how their evolutionary conservation relates to functional coupling between genes.

We further examine the influence of intergenic distance and transcriptional orientation,and explore the role of gene expression and regulatory elements in gene order conservation. Our work contributes to better understanding general trends in eukaryotic genome organization, providing more insights into the subtle interplay between genome rearrangements and gene function.