View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
A-G.01: esloco: simulation-based estimation of local coverage in long-read DNA sequencing
Track: Genomics, epigenomics, and genome editing
-
Adrian Weich, Friedrich-Alexander-Universität (FAU) Erlangen-Nürnberg and Uniklinikum Erlangen,
Erlangen, Germany, Germany
-
Christopher Lischer, Friedrich-Alexander-Universität (FAU) Erlangen-Nürnberg and Uniklinikum Erlangen,
Erlangen, Germany, Germany
-
Julio Vera, Friedrich-Alexander-Universität (FAU) Erlangen-Nürnberg and Uniklinikum Erlangen, Erlangen,
Germany, Germany
Presentation Overview: Show
Motivation
Long-read DNA sequencing is increasingly used for whole-genome and targeted sequencing applications due to its
ability to resolve complex structural variants and its reduced coverage requirements for variant detection.
However, planning long-read sequencing experiments often lacks reliable a priori estimates of target region
coverage, leading to costly and time-consuming pilot studies and biological replicates to meet expectations.
Results
We present esloco, a Monte Carlo-based simulation framework for estimating local coverage in long-read
sequencing experiments. The tool supports scenarios with unknown target regions (e.g. viral integration sites,
CRISPR-Cas9 edits) or PCR-free designs (e.g. base modifications). By modeling coverage as a function of
sequencing depth and read length distribution, esloco enables quantitative predictions of local sequencing
outcomes.
Benchmarking on a 45-gene panel demonstrated close agreement between simulated and empirical coverage profiles
for both Oxford Nanopore and PacBio data. Predictive accuracy further improved when incorporating weighted
masking to account for sequencing biases. These results establish esloco as a practical framework for
simulation-guided experimental design in long-read sequencing studies.
Availability and Implementation
esloco is implemented in Python and is freely available on PyPI and Github as an open-source package.
A-G.02: Analytic Integration over Tree Space: Efficient Inference of Tumor Evolution Dynamics
Track: Genomics, epigenomics, and genome editing
-
Tobias Dieselhorst, Department of Biosystems Science and Engineering, ETH Zurich, Switzerland
-
Marcus Overwater, Department of Biosystems Science and Engineering, ETH Zurich, Switzerland
-
François Bienvenu, Laboratoire de mathématiques de Besançon, Université de Franche-Comté, France
-
Tanja Stadler, Department of Biosystems Science and Engineering, ETH Zurich, Switzerland
Presentation Overview: Show
The joint inference of phylogenetic trees and underlying population dynamics from sequence alignments is a
powerful but computationally expensive procedure. Currently, sampling trees via Markov Chain Monte Carlo
(MCMC) algorithms represents the major bottleneck in these schemes. In applications such as tumor evolution,
however, the topology of the phylogenetic tree is often not of primary interest; consequently, the MCMC merely
serves as a costly numerical integration over the tree space.
We present mathematical results on the properties of summary statistics for sequence alignments. In the limit
of large trees, we show that summary statistics - specifically the distribution of mutations per sample and
the site frequency spectrum - converge to the probability mass functions of statistically independent samples.
The latter can be computed efficiently and correspond to an analytic integration over the tree space. We
exploit this property to infer birth-death processes directly from site frequency spectra, which are commonly
measured from bulk tumor samples but do not allow for tree reconstruction. We further discuss applications for
testing substitution model adequacy based on the distribution of mutation counts.
A-G.03: Insights into the genomes of Mycobacterium abscessus: Decoding the arsenal shaping
pathogenesis
Track: Genomics, epigenomics, and genome editing
-
Saubashya Sur, Ramananda College, India
- Mistu Karmakar, Ramananda College, India
Presentation Overview: Show
Mycobacterium abscessus is an exteremely antibiotic-resistant non-tuberculous mycobacterium responsible for
lung disease and extrapulmonary infections. Individuals suffering from cystic fibrosis, tuberculosis,
bronchiectasis, AIDS, and chronic obstructive pulmonary disease are vulnerable. M. abscessus subsp. abscessus,
M. abscessus subsp. massiliense, and M. abscessus subsp. bolletii are associated with diverse clinical
manifestations and outbreaks. With no available vaccine and limited treatment options, comprehending the
pathogenic arsenal from their genomes is crucial. The objective of this computational analysis was to explore
the genomic island components; virulence factors, antibiotic resistance genes, anti-phage defence systems, and
integrated prophage regions in M. abscessus genomes. While, GIPSY identified genomic islands, BRIG facilitated
comparative visualization, VFDB and CARD identified the virulence and resistance factors. PADLOC and PHASTEST
explored the repertoire of anti-phage defence systems, and integrated prophage regions. The genomes displayed
variability. Pathogenicity islands were enriched in ABC transporters, efflux pumps, ISxac3 transposase, and
TetR regulators. The resistance islands had notable prescence of acriflavine resistance protein B, AmpR,
aminoglycoside acetyltransferases, and HAE1 efflux systems. In the metabolic islands, 4-hydroxybenzoate
polyprenyltransferase, acetyl-CoA acetyltransferase, cytochrome P450, and ABC-type transporters were
prominent. Virulent proteins like fbpA/C, espI/R, mmpS4/L4, mce family proteins, mprA-B, phoP/R, prrA-B, were
dominant. Enriched antibiotic resiatance gene families were: class A β-lactamase, 16S rRNA mutations
conferring aminoglycoside resistance, and 23S rRNA mutations conferring macrolide resistance. Notable
anti-phage systems were PD-T4-6 and MTase_II. The prophage regions associated with virulence were identified.
This intregrative analysis showcased the formidable arsenal shaping M. abscessus pathogenesis, underlining
several novel therapaetic targets against this pathogen.
A-G.04: Comparative Analysis of Sugar Transporters in Yeast Strains for Improved Xylose Utilisation in
Biofuel Applications
Track: Genomics, epigenomics, and genome editing
-
Kellie Ashton, University of Newcastle, Australia, Australia
- Ian Grainge, University of Newcastle, Australia
- Peter Lewis, Ethanol Technologies, Australia
Presentation Overview: Show
Efficient co-consumption of glucose and xylose is a major bottleneck in yeast-based second-generation biofuel
production and is strongly influenced by sugar transporter specificity. While many industrial yeast strains
transport glucose efficiently, xylose uptake remains limiting. Non-conventional yeasts with native xylose
utilisation represent promising alternatives, however, their sugar transporter repertoires are often poorly
characterised or inconsistently annotated.
In this study, a comparative bioinformatic analysis of the major facilitator superfamily (MFS) sugar
transporters was performed in a proprietary yeast strain with native xylose utilisation. All predicted MFS
transporters were identified from the genome and compared to characterised hexose transporters from
Saccharomyces cerevisiae and known xylose transporters from various microorganisms to examine sequence
conservation, transmembrane architecture, and residues implicated in substrate binding and translocation.
Multiple sequence alignments, phylogenetic tree construction, transmembrane domain prediction, and conserved
motif analysis identified conserved MFS motifs as well as variable regions that may have functional
significance. Several transporters shared structural features and conserved residues with characterised xylose
transporters, while others clustered with hexose transporters, allowing us to predict which transporters may
be utilising pentose sugars as substrates.
Comparative analysis clarified the relationship between closely related Hxt6-like transporters, despite
multiple similar sequences and inconsistent nomenclature across datasets. These findings support the
prioritisation of specific transporters for further functional investigation.
This analysis provides a framework for examining sugar transporter diversity in non-conventional yeasts and
informs the rational selection of transporter candidates for engineering improved xylose uptake and
glucose/xylose co-consumption in yeast.
A-G.05: ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data
Track: Genomics, epigenomics, and genome editing
-
Huda Ahmad, University of Birmingham, United Kingdom
- Hannah Doherty, University of Cologne, Germany
- Samuel Benedict, Newcastle University, United Kingdom
- James Haycocks, Newcastle University, United Kingdom
- Ge Zhou, King Abdullah University of Science and Technology, Saudi Arabia
- Patrick Moynihan, Western University, Canada
-
Danesh Moradigaravand, King Abdullah University of Science and Technology, Saudi Arabia
- Manuel Banzhaf, Newcastle University, United Kingdom
Presentation Overview: Show
Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes.
However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive
tools for visualization and analysis. Existing workflows often require multiple independent software tools,
limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables
efficient exploration and sharing of chemical genomics data.
We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and
analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented
datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic
profiles, assess gene–gene and condition–condition correlations, perform GO and KEGG enrichment analysis,
and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also
facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these
features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances
reproducibility, and promotes collaboration within the research community. ChemGenXplore is available at
https://chemgenxplore.kaust.edu.sa/.
A-G.06: ROQ: a new measure for more accurately filtering mapped reads
Track: Genomics, epigenomics, and genome editing
-
Nayoung Park, Konkuk University, South Korea
- Minji Gu, Konkuk University, South Korea
- Younhee Ko, Hankuk University of Foreign Studies, South Korea
- Jaebum Kim, Konkuk University, South Korea
Presentation Overview: Show
Motivation: Next-generation sequencing (NGS) has transformed genomic and multiomic research, making accurate
read alignment essential for reliable downstream analyses such as variant detection and gene expression
quantification. Widely used alignment quality metrics, such as Mapping Quality (MAPQ), do not explicitly
account for the true genomic origin of sequencing reads and are computed inconsistently across alignment
tools. As a result, misaligned reads may receive high confidence scores, while correctly aligned reads can be
undervalued, particularly in complex or repetitive genomic regions. Although machine learning-based
recalibration approaches have demonstrated the feasibility of improving alignment quality assessment, no
existing metric directly models the spatial concordance between a read's mapped position and its true genomic
origin, and its generalizability across diverse aligners, species, and sequencing platforms has not been
systematically evaluated.
Results: We introduce Read Overlapping Quality (ROQ), a machine learning-based metric designed to predict the
Read Overlap Ratio (ROR), which quantifies the degree of spatial overlap between a read's true genomic origin
and its mapped position. Trained on thirteen alignment-derived features using the XGBoost algorithm, ROQ
provides a more origin-aware and comprehensive assessment of alignment quality than MAPQ alone. ROQ
consistently outperforms MAPQ in correlation with the true read origin across diverse genomic contexts,
multiple species, and alignment tools, and demonstrates cross-platform robustness on independent long-read
PacBio data. When applied to real whole-genome sequencing data, ROQ-based filtering reduces false-positive
variant calls while maintaining sensitivity, supporting its utility as a complementary metric for genomic
analysis pipelines.
A-G.07: Sorting Bacterial Genomes Below the Species Rank
Track: Genomics, epigenomics, and genome editing
-
Beatriz Vieira Mourato, Max-Planck-Institute for Evolutionary Biology, Germany
- Sara-Lena Welk, Max-Planck-Institute for Evolutionary Biology, Germany
- Fabian Kloetzl, None, United Kingdom
- Bernhard Haubold, Max-Planck-Institute for Evolutionary Biology, Germany
Presentation Overview: Show
Motivation: Bacterial genomes are classified as belonging to
the same species if their average nucleotide identity, ANI, is at
least 96. The resulting species include organisms
like \emph{Escherichia coli} that comprise both commensals and
pathogens, that is, they harbor important taxonomic diversity below
the species rank. This diversity is less well characterized than at
the species rank and above because the required comparison between
taxonomic and phylogenetic information is more difficult to automate
than applying an ANI threshold. As a result, the majority of bacterial
genomes are only classified at the species rank even if there taxa
below that rank.
Results: Here we quantify the distribution of
bacterial genomes between and within species to show that over 3/4 of
genomes derive from species with below species structure. We establish
a workflow for sorting genomes within species by constructing a genome
phylogeny of all genomes in the species and querying it for the target
taxon. This workflow requires a tool for ANI computation and tools for
querying phylogenies, which we describe. We demonstrate this approach
for the pathogen \emph{E. coli} O111:H8, which we analyze in the
context of the over 7,000 complete genome sequences of \emph{E. coli}
to find 26 genomes in addition to its original 22. We use these 48
genomes to construct candidate diagnostic PCR primers
for \emph{E. coli} O111:H8.
Availability: Our workflow is
distributed as a literate program via the companion website for this
paper, \ty{github.com/evolbioinf/o111h8}.
A-G.08: How Private Are DNA Embeddings? Inverting Foundation Model Representations of Genomic
Sequences
Track: Genomics, epigenomics, and genome editing
-
Sofiane Ouaari, Institute for Bioinformatics and Medical Informatics, University of Tübingen,
Germany
-
Jules Kreuer, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
-
Nico Pfeifer, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
Presentation Overview: Show
DNA foundation models have become transformative tools in bioinformatics and healthcare applications. Trained
on vast genomic datasets, these models can be used to generate sequence embeddings, dense vector
representations that capture complex genomic information. These embeddings are increasingly being shared via
Embeddings-as-a-Service (EaaS) frameworks to facilitate downstream tasks, while supposedly protecting the
privacy of the underlying raw sequences. However, as this practice becomes more prevalent, the security of
these representations is being called into question. This study evaluates the resilience of DNA foundation
models to model inversion attacks, whereby adversaries attempt to reconstruct sensitive training data from
model outputs. In our study, the model's output for reconstructing the DNA sequence is a zero-shot embedding,
which is then fed to a decoder. We evaluated the privacy of three DNA foundation models: DNABERT-2, Evo 2, and
Nucleotide Transformer v2 (NTv2). Our results show that per-token embeddings allow near-perfect sequence
reconstruction across all models. For mean-pooled embeddings, reconstruction quality degrades as sequence
length increases, though it remains substantially above random baselines. Evo 2 and NTv2 prove to be most
vulnerable, especially for shorter sequences with reconstruction similarities > 90\%, while DNABERT-2's BPE
tokenization provides the greatest resilience. We found that the correlation between embedding similarity and
sequence similarity was a key predictor of reconstruction success. Our findings emphasize the urgent need for
privacy-aware design in genomic foundation models prior to their widespread deployment in EaaS settings.
Training code, model weights and evaluation pipeline are released on:
https://github.com/not-a-feature/DNA-Embedding-Inversion.
A-G.09: TabPFN-Wide: Continued Pre-Training for Extreme Feature Counts
Track: Genomics, epigenomics, and genome editing
-
Christopher Kolberg, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
-
Jules Kreuer, Institute for Bioinformatics and Medical Informatics, University of Tübingen,
Germany
-
Jonas Huurdeman, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
-
Sofiane Ouaari, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
-
Katharina Eggensperger, Lamarr Institute for Machine Learning and Artificial Intelligence, TU Dortmund,
Germany
-
Nico Pfeifer, Institute for Bioinformatics and Medical Informatics, University of Tübingen, Germany
Presentation Overview: Show
Revealing novel insights from the relationship between molecular measurements and pathology remains a very
impactful application of machine learning in biomedicine. Data in this domain typically contain only a few
observations but thousands of potentially noisy features, posing challenges for conventional tabular machine
learning approaches. While prior-data fitted networks emerge as foundation models for predictive tabular data
tasks, they are currently not suited to handle large feature counts (>500). Although feature reduction
enables their application, it hinders feature importance analysis. We propose a strategy that extends existing
models through continued pre-training on synthetic data sampled from a customized prior. The resulting model,
TabPFN-Wide, matches or exceeds its base model's performance, while exhibiting improved robustness to noise.
It seamlessly scales beyond 30,000 categorical and continuous features, regardless of noise levels, while
maintaining inherent interpretability, which is critical for biomedical applications. Our results demonstrate
that prior-informed adaptation is suitable to enhance the capability of foundation models for high-dimensional
data. On real-world omics datasets, we show that many of the most relevant features identified by the model
overlap with previous biological findings, while others propose potential starting points for future studies.
A-G.10: Short-Context gLM Training for Long-Range Variant Effect Prediction
Track: Genomics, epigenomics, and genome editing
-
Megha Hegde, Kingston University London, United Kingdom
- Jean-Christophe Nebel, Kingston University London, United Kingdom
- Farzana Rahman, Kingston University London, United Kingdom
Presentation Overview: Show
Disease susceptibility in humans is frequently driven by mutations within the DNA. Contemporary research has
shown that many such mutations lie within the non-coding regions of the genome, several thousands of
base-pairs (bp) from the transcription start sites (TSS) of their target genes. In the age of artificial
intelligence, Transformer-based genomic language models (gLMs) are commonly used to interpret variant effects.
Their ability to exploit the large datasets produced by next-generation sequencing, and their aptitude for
modelling long-range interactions within sequences, makes them ideal for the task. However, the quadratic
scaling of the attention mechanism with context length results in high computational resource consumption, and
inefficiency when training gLMs on very long (10000bp+) sequences.
DroPE (Dropping the Positional Embeddings of LMs after training), originally developed for generative
text-based large language models (LLMs), provides a method for extending the context of pretrained LLMs
without the need for long-context fine-tuning. This paper adapts and applies DroPE to BERT-based gLMs, and
demonstrates that it can extend the context of models pretrained on short genomic sequences. Furthermore,
models incorporating DroPE demonstrate an enhanced ability to model long-range context, achieving competitive
performance on variants far (>35kbp) from the TSS without the need for computationally expensive
long-context pretraining and fine-tuning. Code and fine-tuned models are available at:
https://github.com/meghegde/gLM-DroPE.
A-G.11: NCRP: enhancing long-read classification through neighborhood-consistency refinement and propagation
in the overlap graph
Track: Genomics, epigenomics, and genome editing
- Xun Ding, SHANDONG UNIVERSITY, China
-
Lianrong Pu, SHANDONG UNIVERSITY, China
- Haitao Jiang, SHANDONG UNIVERSITY, China
- Zhu Daming, SHANDONG UNIVERSITY, China
Presentation Overview: Show
Taxonomic classification of metagenomic sequencing reads is a crucial task in metagenome analysis, with
significant implications for research on health, diet, and drug responses. With the advancement of sequencing
technologies, several methods have been developed for long-read taxonomic classification, including Kraken2,
Centrifuge, and CLARK. While these classifiers achieve high precision, they often suffer from low sensitivity,
leaving a substantial proportion of reads unclassified, especially for highly diverse and rapidly evolving
organisms such as viruses. ClassGraph addressed this challenge by introducing a label propagation strategy,
significantly improving sensitivity at the cost of precision.To further address this issue, we propose NCRP, a
Neighborhood-Consistency Refinement and Propagation algorithm, designed to enhance both precision and
sensitivity in long-read classification. NCRP builds upon an initial classifier and a read overlap graph. It
first identifies and removes ambiguous classifications that are inconsistent with neighboring reads in the
overlap graph, then propagates confident taxonomic assignments from classified reads to unclassified ones.
Evaluation on simulated and mock datasets demonstrates that NCRP consistently outperforms existing classifiers
in both precision and sensitivity, achieving the highest F1 score across all tests. We further benchmarked the
ability of taxonomic classifiers to detect microbial species. NCRP combined with Kraken2 or Centrifuge
achieved the best performance, whereas CLARK showed relatively poor performance on this task.NCRP is
open-source and can be accessed at https://github.com/SDU-ACG-Lab/NCRP.
A-G.13: Compounding of rare copy-number variants and polygenic risk: A genetic signature of assortative
mating
Track: Genomics, epigenomics, and genome editing
-
Caterina Cevallos, University of Lausanne, Switzerland
- Chiara Auwerx, University of Lausanne, Switzerland
- Robin Hofmeister, University of Lausanne, Switzerland
- Théo Cavinato, University of Lausanne, Switzerland
- Tabea Schoeler, University of Lausanne, Switzerland
- Zoltán Kutalik, University of Lausanne, Switzerland
- Alexandre Reymond, University of Lausanne, Switzerland
Presentation Overview: Show
Background:
Large copy-number variants (CNVs) are linked to a broad spectrum of outcomes, with carriers of the same CNV
exhibiting variable disease severity. Although additional rare and common variants have been proposed as
modifiers of CNVs' expressivity, their interplay still remains largely understudied.
Material and Methods:
We explored the impact of polygenic scores (PGS) on shaping CNV carriers' heterogeneity in the UK Biobank,
focusing on 119 established CNV-trait associations comprising 43 traits and 27 CNVs. Linear regressions
assessed the individual, joint, and synergistic contributions of CNVs and PGSs. Due to participation bias in
population biobanks, we expected trait-increasing CNV carriers to have lower PGSs.
Results:
We demonstrate additive contribution of PGS and CNV for 45 (38%) CNV-trait pairs, as well as two interactions
between the 22q11.23 duplication and PGSs for grip strength and gamma-glutamyltransferase levels. Strikingly,
CNVs and PGSs exhibited a widespread positive correlation, revealing a tendency for PGSs to exacerbate CNV
effects—a pattern that could be explained by linkage disequilibrium only for a single CNV-trait pair. Given
a non-null inheritance rate for all 17 testable CNVs, we investigated whether assortative mating could account
for this phenomenon. We found strong agreement between the empirical CNV-PGS correlation and the one predicted
by assortment (r=0.45, p=3.9e-7). Similar trends of positive correlation were observed between PGSs and
genome-wide burden of CNVs or loss-of-function variants.
Conclusion:
PGSs improve stratification of CNV carriers at risk of developing clinically-relevant comorbidities,
compounding the impact of rare damaging variants through assortative mating.
A-G.14: ReVSeq Visuals: A Rapid, Modular Environment for Viral Sequence Visualization
Track: Genomics, epigenomics, and genome editing
-
Charlyne Bürki, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland
-
Mara Neacsu, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland
-
Louis du Plessis, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland
-
Tanja Stadler, Department of Biosystems Science and Engineering, ETH Zürich, Switzerland
Presentation Overview: Show
Motivation: Rapidly investigating and understanding the evolution of co-circulating respiratory pathogens is
paramount for robust surveillance systems.
However, bottlenecks exist in implementing bioinformatics pipelines and interpreting their outputs for
non-domain specialized users such as public
health professionals and clinicians, especially in environments requiring regular data updates or studies
involving multi-center collections.
Results: We developed ReVSeq Visuals, a framework to build customizable dashboards to obtain instantaneous
insights from sequencing data in an
accessible and intuitive manner. ReVSeq Visuals consists of a backend module to curate the output of viral
genomic bioinformatics pipelines and an
interactive visualization platform to derive insights from sequencing data. Together, with limited
bioinformatics configurations required by the user,
these two components form the dashboard-building framework that visualize insights based on viral sequencing
data together with its associated
spatiotemporal metadata. The dashboard integrates existing public health tools, allowing to fully customize
the studied viral strains. We showcase
the dashboard's functionalities using a set of diverse respiratory viruses collected in Switzerland from 2023
to 2025, and illustrate derivable insights
for RSV-A/B co-circulation.
Availability and Implementation: All code (predominantly Python) is available at
https://github.com/charlynebuerki/ReVSeq-dashboard and
the dashboard is available at https://revseq.charlynebuerki.com/.
A-G.15: Genome wide eccDNA hotspot propensity estimation from confounder aware features
Track: Genomics, epigenomics, and genome editing
-
Gerardo Benevento, University of Salerno, Italy
- Delfina Malandrino, University of Salerno, Italy
- Daniele Salerno, University of Salerno, Italy
- Alessia Ture, University of Salerno, Italy
- Rocco Zaccagnino, University of Salerno, Italy
Presentation Overview: Show
Studying extrachromosomal circular DNA (eccDNA and ecDNA) can reveal genome instability and regulatory change,
especially in cancer
disease. Interpreting genome wide recurrence is difficult because apparent hotspots can be inflated by
technical and genomic confounders,
including low mappability, repeats, and segmental duplications. We therefore move from classifying individual
circles to a locus propensity view
and learn to rank genomic windows using interpretable context features, while using technical tracks only for
filtering and diagnostics. We score
recurrence across public experiments from CircleBase v2 in fixed 25 kb bins. For each bin we compute replicate
support, defined as the number
of independent experiments (unique PubMed ID and assay) that report at least one overlapping eccDNA interval.
To define a fair baseline,
we keep the number and sizes of eccDNA intervals from each experiment fixed, but randomly shuffle their
genomic positions within the same
chromosome many times. We then measure how often each bin would be supported under these randomized placements
and use the average
as the expected background support. Comparing observed support to this baseline provides a measure of the
continuous enrichment score used
for regression and genome wide ranking tasks. The proposed approach was evaluated along two axes: (i) recovery
and tail prioritization of
the continuous enrichment signal (random-split = 0.711 ± 0.004, Lift@1% 38, chromosome holdout 0.51, with
telomere/centromere
ablations probing positional reliance) and (ii) matched comparison to sequence-only methods via the DeepCircle
1 kb balanced classification
protocol, where we obtain competitive performance.
A-G.16: SPA-C: an hybrid tool to accurately scaffold genomes using Hi-C and Deep-Learning
Track: Genomics, epigenomics, and genome editing
-
Alexis Mergez, University of Toulouse, INRAE, France
- Raphael Mourad, University of Toulouse, France
- Guillermina Hernandez-Raquet, University of Toulouse, TBI, INRAE, France
- Matthias Zytnicki, University of Toulouse, INRAE, France
Presentation Overview: Show
Genome assembly is a computational pipeline designed to reconstruct chromosomes from small sequencing reads.
Following their assembly, contiguous sequences (contigs) are arranged into chromosome-long sequences during
scaffolding. Hi-C, a long-range linkage information between regions of the genome widely used in recent large
sequencing projects, is often required to correctly order contigs. Several tools have been developed to
automate this task following either statistical or deep-learning approaches. Statistical approaches summarise
2D Hi-C matrices into contact densities across sequences, thus ignoring informative visual patterns. The sole
existing deep-learning tool uses a transformer-based computer vision model to correct the assembly. It has
been trained on several species and uses Hi-C matrices directly. Yet it comes as a supplementary step in the
scaffolding process, introducing extra computation time, and has been trained on a dataset that might contain
labelling errors, that could limit its utilisation.
We propose SPA-C, an hybrid pipeline combining the strengths of both approaches. Linkage prediction is handled
with a frugal CNN-based model and a graph-solving algorithm is used to generate the scaffolds. Through our
input's design, the model is able to both correct errors within assemblies and link contigs, leveraging small,
local Hi-C contact matrices. We handled low-complexity regions that might induce erroneous predictions using
an external tool, improving the overall accuracy of generated assemblies. On a benchmark of six various
genomes and four standard metrics, SPA-C outperformed four out of four state-of-the-art methods while
achieving comparable start-to-end computation time.
Python and Bash scripts are available at zenodo (doi: 10.5281/zenodo.19000362).
A-G.17: Kente: A Graph-based Pangenomic Approach for Horizontal Gene Transfer Detection in
Microbiomes.
Track: Genomics, epigenomics, and genome editing
-
Natalie Kokroko, Rice University, United States
- Richa Jayanti, Rice University, United States
- Nicolae Sapoval, Rice University, United States
- Michael Nute, Rice University, United States
- Luay Nakhleh, Rice University, United States
- Todd Treangen, Rice University, United States
Presentation Overview: Show
Motivation: Horizontal gene transfer (HGT) shapes bacterial evolution and microbial ecosystems, yet detecting
HGT within microbiomes remains a challenge due to fragmented metagenomic assemblies, reference bias, reliance
on gene boundaries, and limited ability to model structural mosaicism and patterns across genomes.
Methods: We present Kente, a novel pangenome graph-based framework designed for HGT detection that aligns
metagenomic assembly contigs to a curated database of >600 genus-level bacterial pangenome graphs
constructed
using minigraph. Kente infers local taxonomic composition along contigs using alignment evidence and
classifies
candidate transfers using structured clade-transition topologies (e.g., A-B-A sandwich, open tips, and
mosaic
patterns). A complementary intra-genus module detects inter-species transfers within a single genus graph
using
segment-level clade annotations.
Results: Across simulated intra- and inter-genus transfer scenarios, Kente achieves higher precision and
comparable
recall relative to existing gene-centric microbiome HGT detection approaches while reducing false positives
from
fragmented assemblies. Application to real human gut metagenomes (HMP2, n = 26) demonstrates Kente's
ability
to detect candidate cross-lineage transfer regions in complex microbial communities. Runtime profiling shows
near-linear
scaling with input size, enabling efficient analysis of large metagenomic assemblies.
Availability and Implementation: https://github.com/treangenlab/Kente
A-G.18: A context-aware framework leveraging LLM-estimated variant effects for individualized
variant-to-phenotype interpretation
Track: Genomics, epigenomics, and genome editing
-
Nayoung Park, Department of Biomedical Science and Engineering, Konkuk University, South
Korea
-
Dayeon Kim, Division of Biomedical Engineering, Hankuk University of Foreign Studies, South Korea
-
Jaebum Kim, Department of Biomedical Science and Engineering, Konkuk University, South Korea
-
Younhee Ko, Division of Biomedical Engineering, Hankuk University of Foreign Studies, South Korea
Presentation Overview: Show
Motivation: Advances in sequencing technologies have uncovered millions of human genetic variants, yet
accurately estimating pathogenic risk at the sample level remains a major challenge. Polygenic risk scores
(PRS), built on genome-wide association studies (GWAS), provide a useful framework for aggregating
population-level effects, but they have important limitations. Because PRS depends on cohort-specific effect
sizes, its performance often does not generalize well across populations. In addition, PRS is not well-suited
to capturing rare or context-dependent variant effects that may contribute substantially to disease risk in
individual genomes. To address these limitations, we developed the Variant Probability Ratio-based score
(VPRscore), a context-aware framework that leverages probability ratios derived from a large language model
trained on genomic sequences. For each variant, the model estimates the probability of the alternative allele
relative to the reference allele within its native sequence context, yielding a variant probability ratio
(VPR) that reflects potential functional impact.
Results: We integrated these variant-level VPR signals with CADD to construct a unified sample-level score,
VPRscore, that provides a biologically grounded estimate of disease risk informed by variant pathogenicity. In
benchmark analyses, VPRscore robustly distinguished pathogenic from benign variants in ClinVar and accurately
tracked increasing proportions of pathogenic variants across simulated cohorts. The method achieved high
discriminative performance, with AUC values reaching 0.989. By deriving risk estimates directly from DNA
sequence context rather than relying on population-level association statistics, VPRscore provides an
interpretable and potentially more generalizable foundation for individualized genomic risk assessment.
A-G.19: Read-Level Classification of HPV Viral Metagenomic Data Using Transformer Architectures
Track: Genomics, epigenomics, and genome editing
-
Simone Rancati, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
-
Sakshi Pandey, Department of Computer and Information Science and Engineering, University of Florida,
United States
-
Micheal Sy, Department of Epidemiology, College of Public Health and Health Professions, University of
Florida, United States
-
Pablo Arozarena Donelli, Department of Electrical, Computer and Biomedical Engineering, University of
Pavia, Italy
-
Giovanna Nicora, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
-
Conrad Testagrose, Department of Computer and Information Science and Engineering, University of Florida,
United States
-
Christina Boucher, Department of Computer and Information Science and Engineering, University of Florida,
United States
-
Riccardo Bellazzi, Department of Electrical, Computer and Biomedical Engineering, University of Pavia,
Italy
- Marco Salemi, Emerging Pathogens Institute, University of Florida, United States
-
Enea Parimbelli, Department of Electrical, Computer and Biomedical Engineering, University of Pavia, Italy
-
Simone Marini, Department of Epidemiology, College of Public Health and Health Professions,
University of Florida, United States
Presentation Overview: Show
Motivation: Accurate detection of viral reads within overwhelming human and bacterial background is essential
for viral metagenomics, outbreak surveillance, and precision diagnostics. This is also relevant in cancer and,
more broadly, in shotgun sequencing workflows, where accurate read-level discrimination can improve downstream
genotyping, assembly, and interpretation. Alignment-free tools such as Kraken-2 are widely
used, but their reliance on exact or near-exact k-mer matches may reduce sensitivity to divergent or
low-abundance viral reads. Transformer encoders provide contextual nucleotide representations that may
overcome these limitations, yet their performance on realistic HPV-scale short reads remains poorly explored.
Results: We compared two viral transformers, ViBE and XVir, against Kraken-2 on 150-bp Illumina-like reads
simulated from 184 viral genomes, 147 bacterial genomes, and 62 human assemblies, using human papillomavirus
(HPV) as the pathogen of interest. Across four tasks (Virus vs Human, Virus vs Bacterial, Human vs Bacterial,
and Virus vs Human vs Bacterial), ViBE consistently generated the most informative embedding space and the
best downstream classification, achieving 95–98% balanced accuracy in binary tasks and 91–92% in the
three-class setting, often with only the top 100 embedding features. XVir, despite its HPV-focused design,
performed well
mainly in Virus vs Human discrimination and generalized less effectively to bacterial background. Kraken-2
remained strongest in Human vs Bacterial separation but showed lower performance on HPV-centered tasks.
Availability and implementation:
https://github.com/simoRancati/single-read-viral-transformers-Read-Level-Classification-of-Viral-Metagenomic-Data-Using-Transforme
A-G.20: Characterizing the temporal dynamics of Cancer Hallmarks and somatic mutations across 80,000+
tumors
Track: Genomics, epigenomics, and genome editing
- Federica Malighetti, University of Milano Bicocca, Italy
-
Alberto Maria Villa, University of Milano Bicocca, Italy
- Ivan Civettini, San Raffaele Scientific Institute, Milan, Italy, Italy
- Luis Zapata, Institute of Cancer Research, London, UK, United Kingdom
- Andrea Aroldi, University of Milano Bicocca, Italy
- Luca Mologni, University of Milano Bicocca, Italy
- Daniele Ramazzotti, University of Milano Bicocca, Italy
Presentation Overview: Show
Understanding cancer evolution across diverse tumor types remains challenging due to its inherent
heterogeneity. However, uncovering evolutionary pathways common across tumors might elucidate fundamental
principles of cancer progression. This study extends previous efforts to addresses this challenge by
integrating recurrent mutations with the temporal acquisition of Hallmarks of Cancer across a diverse cohort
of over 80,000 tumors.
By focusing on Hallmarks dynamics, we discover distinct patterns of tumor progression with clinical
implications. Leveraging the ASCETIC framework to characterize cancer evolution, we comprehensively analyze
over 80,000 tumors spanning most cancer types. We map Hallmarks enrichment across the different phases of
tumor progression and track the enrichment of Hallmarks at various stages of tumor evolution, revealing which
Hallmarks exert significant influence during cancer progression in distinct phases, and their impact on
prognosis and metastasis. Key findings include the identification of conserved mutation trajectories, such as
TP53, ARID1A, and RNF43 as early drivers, and the sequential acquisition of mutations in ErbB receptors like
EGFR and ERBB family members. These findings significantly advance our understanding of cancer evolution and
might have significant implications for the development of broad-spectrum targeted therapies.
A-G.21: Interpreting Structural Variations in Breast Cancer Cell Lines Using 3D Genome Data
Track: Genomics, epigenomics, and genome editing
-
Jingyu Hao, The Hong Kong University of Science and Technology, Hong Kong
- Yongyi Luo, The Chinese University of Hong Kong, Hong Kong
- Jiandong Shi, The Chinese University of Hong Kong, Hong Kong
- Depeng Wang, GrandOmics Inc, China
- Shu Wang, Peking University People’s Hospital, China
- Xiaodan Fan, The Chinese University of Hong Kong, Hong Kong
- Weichuan Yu, The Hong Kong University of Science and Technology, Hong Kong
Presentation Overview: Show
Structural variations (SVs) are a major source of genomic alterations in cancer, yet the majority of SVs occur
in non-coding regions and remain difficult to interpret using linear genome annotations alone. 3D genome data
provides additional spatial information for understanding potential regulatory impacts of SVs, but its
practical contributions and limitations remain unclear. Here, we perform an integrative analysis of SVs and 3D
genome organization in three breast cancer cell lines. Although the majority of SVs occur in non-coding
regions, approximately 70% of them are associated with distal gene promoters via high-frequency 3D contacts,
revealing potential regulatory relationships beyond linear genome annotations. In the HCC1937 tumor–normal
pair under a controlled genetic setting, we observed widespread 3D genome reorganization, including A/B
compartment switching affecting roughly one-third of the genome and pronounced insulation score changes at
about 30% of somatic SV loci. Detailed analyses of representative somatic SVs further revealed local
alterations in topologically associating domain (TAD) structure at SV loci, including the formation of new
boundaries or the disruption of pre-existing ones. These SV loci also form high-frequency contacts with distal
cancer-related gene loci, such as SLC2A1, HGF, and STAG2, with similar interactions observed for some of these
genes in SK-BR-3, suggesting that these spatial associations are common across breast cancer cell lines. These
observations indicate that SVs that appear functionally silent under conventional annotation may influence
gene regulation through their 3D genome, providing a potential mechanism for prioritizing non-coding variants
for further functional investigation.
A-G.22: GENA-Web - GENomic Annotations Web Inference using DNA language models
Track: Genomics, epigenomics, and genome editing
- Aleksei Shmelev, AXXX, Moscow, Russia; HSE University, Moscow, Russia, Russia
- Maxim Petrov, AXXX, Moscow, Russia, Russia
-
Dmitry Penzar, Vavilov Institute of General Genetics, Russian Academy of Sciences, Moscow, Russia,
Russia
- Nikolay Akhmetyanov, AXXX, Moscow, Russia, Russia
- Maksim Tavritskiy, AXXX, Moscow, Russia, Russia
- Stepan Mamontov, AXXX, Moscow, Russia, Russia
-
Yuri Kuratov, AXXX, Moscow, Russia; Moscow Independent Research Institute of Artificial Intelligence,
Moscow, Russia, Russia
-
Mikhail Burtsev, London Institute for Mathematical Sciences, London, United Kingdom, United Kingdom
- Olga Kardymon, Quantori, Belgrade, Serbia, Serbia
-
Veniamin Fishman, AXXX, Moscow, Russia; Sirius University, Sochi, Russia; Institute of Cytology and
Genetics, Novosibirsk, Russia, Russia
Presentation Overview: Show
The advent of advanced sequencing technologies has significantly reduced the cost and increased the
feasibility of assembling high-quality genomes. Yet, the annotation of genomic elements remains a complex
challenge. Even for species with comprehensively annotated reference genomes, the functional assessment of
individual genetic variants is not straightforward. In response to these challenges, recent breakthroughs in
machine learning have led to the development of DNA language models. These transformer-based architectures are
designed to tackle a wide array of genomic tasks with enhanced efficiency and accuracy. In this context, we
introduce GENA-Web, a web-based platform that consolidates a suite of genome annotation tools powered by DNA
language models. The version of GENA-Web presented here encompasses a diverse set of models trained on human
data, including the prediction of promoter activity, annotation of splice sites, determination of various
chromatin features, and a model for scoring of enhancer activity in \textit{Drosophila}. GENA-Web is
accessible online at https://dnalm.airi.net/
A-G.23: Mapping and improving the limited sensitivity of SigProfilerAssignment to flat mutational
signatures
Track: Genomics, epigenomics, and genome editing
-
Qixuan Wang, Inselspital, Bern University Hospital and University of Bern, Switzerland,
Switzerland
-
Maria Katsantoni, Inselspital, Bern University Hospital and University of Bern, Switzerland, Switzerland
-
Matúš Medo, Inselspital, Bern University Hospital and University of Bern, Switzerland, Switzerland
Presentation Overview: Show
SigProfilerAssignment is a popular and well-performing tool for estimating the activity of mutational
signatures in sequenced samples. Although signature analysis generally becomes more accurate with increasing
mutation burden in the analyzed samples, we report that this is not always the case for SigProfilerAssignment.
In particular, the tool's sensitivity deteriorates in the range of mutation burden typical for whole-genome
sequencing when multiple signatures with flat mutational profiles are simultaneously active. We propose a
modified algorithm, FSPA, that achieves a higher sensitivity without sacrificing the high precision of
SigProfilerAssignment. This modification enhances the reliability of mutational signature analysis in cancer
genome studies, thereby ensuring more accurate downstream biological interpretations.
A-G.24: Biopsy-Aware Phylogenetic Reconstruction of Cancer Evolution Using Single-Cell Copy Number
Profiles
Track: Genomics, epigenomics, and genome editing
-
Jarosław Paszek, Faculty of Mathematics, Informatics and Mechanics University of Warsaw,
Poland
-
Agnieszka Mykowiecka, Faculty of Mathematics, Informatics and Mechanics, University of Warsaw, Poland
- Krzysztof Gogolewski, Institute of Informatics, University of Warsaw, Poland
Presentation Overview: Show
Reconstructing tumour evolution from single-cell copy-number profiles (CNPs) collected across multiple biopsy
time points requires methods that explicitly exploit the temporal ordering of biopsies. We introduce a
biopsy-aware framework for simulating, reconstructing, and benchmarking copy-number phylogenies under
realistic models of genome evolution. The framework treats biopsies as explicit temporal constraints and
supports two complementary inference strategies. The first, biopsy-guided ancestry inference, enforces
time-consistent parent–child relations and resolves the earliest branching structure using an auxiliary
NJ-like step. The second, pairwise NJ-anticentral inference, iteratively selects cell pairs and designates an
ancestor by combining anticentrality, plausibility, and parsimony criteria, yielding fully labelled CNP-trees.
We evaluate reconstruction accuracy using ancestor–descendant recovery (AD-F1) and Generalized
Robinson–Foulds (GRF) distances. Our framework enables not only the reconstruction of copy-number
phylogenies but also a controlled assessment of their reliability. By simulating full evolutionary histories,
we can compare inferred and true distances, quantify how tumour heterogeneity, biopsy sparsity, and
multi-locus events degrade signal, and delineate the limits these factors impose on any method. Within this
framework we introduced new NJ-like strategies and demonstrated, across diverse regimes, that principled use
of biopsy structure together with informed pairwise ancestry decisions yields substantial improvements in
reconstruction accuracy.
A-G.25: Evaluating local assembly strategies for long-read structural variant calling
Track: Genomics, epigenomics, and genome editing
-
Tim Mirus, Leibniz Institute for Immunotherapy, Germany
- Richard Lüpken, Leibniz Institute for Immunotherapy, Germany
- Birte Kehr, Leibniz Institute for Immunotherapy, Germany
Presentation Overview: Show
Long-read sequencing enables detection of structural variants (SVs) with far greater accuracy and completeness
than was previously possible. This advance has spurred the development of many SV callers that implement
diverse algorithmic strategies. Commonly, SV calling proceeds in several processing steps. Evaluations of SV
callers, however, typically focus on overall performance and conceal which steps drive observed differences.
This hinders targeted, evidence-based improvements to existing SV calling approaches. Here, we implemented
five consensus computation strategies similar to those employed by popular SV callers that perform a local
assembly. We compare their outputs on extensive simulated data and reveal that some strategies are more
effective for particular input data types. For example, the strategy employed by PacBio's tool Sawfish works
much better on our simulated HiFi reads than on the more noisy ONT reads. While the read error profiles and
the read pre-processing has a substantial influence on some strategies, other strategies are more robust
overall. These findings guide algorithmic decisions for the development of improved future SV callers.
A-G.26: Multilateration-Based Indexing and Navigation for Error-Tolerant Read Mapping
Track: Genomics, epigenomics, and genome editing
-
Luting Zhou, Shanghai Jiao Tong University, China
- Minghao Fang, Shanghai Jiao Tong University, China
- Yiru He, Shanghai Jiao Tong University, China
- Cheng Wang, Shanghai Jiao Tong University, China
- Jinpu Cai, Shanghai Jiao Tong University, China
- Yuxuan Wang, Shanghai Jiao Tong University, China
- Zhenwei Huang, Shanghai Jiao Tong University, China
-
Shiyang Yu, Shanghai Sixth People's Hospital Affiliated to Shanghai Jiao Tong University School of
Medicine, China
- Shiyang Ma, Shanghai Jiao Tong University, China
-
Yan Li, State Key Laboratory of Genome and Multi-omics Technologies, BGI Research, Shenzhen 518083, China,
China
- Hongyi Xin, Shanghai Jiao Tong University, China
Presentation Overview: Show
Indexing is a fundamental step in read mapping.
To maximize efficiency, ideal indexers must utilize long seeds while remaining robust to sequencing errors and
genetic variations.
Existing schemes typically rely on probabilistic error tolerance at fixed sub-locations or approximate edit
distances via Euclidean embeddings.
However, because edit-distance space is a metric space but not a normed space, approximations using
inner-product spaces (like Euclidean space) are inherently lossy.
Here, we demonstrate that high-precision positioning is attainable in coordinate-less string space through
multilateration, provided the discrete space is covered by a meticulously selected set of beacon strings.
We show that for any query, a comprehensive set of high-similarity candidates can be efficiently localized
using a small subset of proximal beacons.
Leveraging these insights, we propose NavigaMer, a multi-tiered navigation system for string indexing.
NavigaMer retrieves candidate reference sequences by progressively narrowing the search space within the
sequence universe, guided by beacon sets ranging from coarse-to-fine resolutions.
Experiments on simulated and real genomes demonstrate that NavigaMer is fully error-tolerant and highly
accurate.
It identifies all reference candidates within a specified edit distance with minimal false positives,
providing a high-sensitivity foundation for future read mappers.
A-G.27: Pathoplexus: Building a new kind of pathogen database
Track: Genomics, epigenomics, and genome editing
-
Emma Hodcroft, Swiss TPH, Allschwil; University of Basel, Basel; Swiss Bioinformatics Institute,
Lausanne, Switzerland
-
Chaoran Chen, ETH Zürich, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne, Switzerland,
Switzerland
-
Cornelius Roemer, University of Basel, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne,
Switzerland, Switzerland
-
Anna Parker, Swiss TPH, Allschwil; University of Basel, Basel; Swiss Bioinformatics Institute, Lausanne,
Switzerland
- Anderson Brito, Instituto Todos pela Saúde, São Paulo, Brazil, Brazil
- George Githinji, KEMRI-Wellcome Trust Research Programme, Kilifi, Kenya, Kenya
- Senjuti Saha, Child Health Research Foundation, Dhaka, Bangladesh, Bangladesh
-
Richard Neher, University of Basel, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne,
Switzerland, Switzerland
-
Tanja Stadler, ETH Zürich, Basel, Switzerland; Swiss Bioinformatics Institute, Lausanne, Switzerland,
Switzerland
-
Theo Sanderson, London School of Hygiene and Tropical Medicine, London, UK, United Kingdom
Presentation Overview: Show
Sharing viral genomic data is essential for advancing scientific research and informing public health
responses. Although platforms such as the International Nucleotide Sequence Database Collaboration (INSDC) and
GISAID facilitate data sharing, they do not fully address several needs of the pathogen genomics community.
Persistent concerns about data misuse, "scooping," and restrictions on reuse from protected repositories
highlight the need for more flexible and equitable data-sharing systems.
Pathoplexus is a community-driven viral genomics database designed to balance data accessibility with
submitter autonomy. Built using open-source technologies and guided by transparent governance, the platform
allows data submitters to temporarily control how their data are used while still enabling rapid access for
researchers and public health officials. All Pathoplexus submissions are automatically uploaded to INSDC -
either immediately if fully open, or after one year if initially shared under a restricted-use option.
Pathoplexus leverages Loculus, an open-source platform for managing viral sequence databases. Its web
interface and API support both interactive exploration and automated analyses. The project operates as a
non-profit association governed by an international Executive Board and includes members from 14 countries
across five continents, reflecting a commitment to equity and diverse public health priorities.
Despite launching only recently, Pathoplexus has already made a measurable impact. As of March 2026, the
database contains more than 10,400 directly submitted sequences. It hosts the first available Sudan ebolavirus
sequence from the February 2025 outbreak in Uganda. In 2025, Pathoplexus also received more mpox sequences
than any other database.
A-G.28: A Fluorescent Probe Under the Microscope Showing Dual Recognition of B-DNA and G-Quadruplex
DNA
Track: Genomics, epigenomics, and genome editing
-
Richard López-Corbalán, Universidad de Alcalá, Spain
- Lorenzo Gramolini, Universidad de Alcalá, Spain
- Cristina García-Iriepa, Universidad de Alcalá, Spain
- Marco Marazzi, Universidad de Alcalá, Spain
Presentation Overview: Show
The use of small molecules as fluorescent DNA markers has been extensively explored, yet designing selective
probes remains challenging due to the missing relationship between probe structure and DNA conformation. To
address this gap, we present a novel multistep computational protocol applied to the recently proposed
fluorescent marker QCy(MeBT)3, which is capable of simultaneously recognizing both B-DNA and G-quadruplex (G4)
DNA through distinct emission signatures.
Our methodology integrates molecular docking, molecular dynamics with a newly parameterized force field to
capture conformational flexibility, and hybrid QM/MM simulations to compute absorption and fluorescence
spectra. Our protocol successfully identifies the specific molecular conformations that enable selective
probe-DNA binding following a lock-and-key recognition mechanism. We predict a remarkably high binding
affinity for both DNA targets, driven by diverse energetic interactions primarily electrostatic in B-DNA minor
groove binding, and a combination of electrostatic and Van der Waals interactions in G4 top-stacking.
Crucially, our simulations reveal that the distinct absorption and fluorescence shifts observed experimentally
originate predominantly from the intrinsic conformational properties of the probe in water, rather than solely
from specific interactions with the DNA environment. Overall, this work establishes a fundamental principle
for predicting the performance of fluorescent probes in complex biological environments: the structural and
electronic properties of specific molecular conformations dictate their photobiophysical fate upon binding to
a given DNA sequence, thereby governing their ability to recognize specific DNA topologies.
A-G.29: SVlog: Unlocking Structural Variation in Rare Diseases through an Extensible Logic Programming
Framework
Track: Genomics, epigenomics, and genome editing
-
Mikhail Gudkov, Garvan Institute of Medical Research, Australia
- André Luiz Martins Reis, Garvan Institute of Medical Research, Australia
- Meutia Kumaheri, Garvan Institute of Medical Research, Australia
- Ira Deveson, Garvan Institute of Medical Research, Australia
Presentation Overview: Show
Structural variants (SVs) are a diverse group of genetic variants defined by a minimum size of 50 base pairs.
SVs account for the majority of all variant bases in a person's genome, and have been frequently implicated in
inherited disease and cancer. However, SV analysis is challenging due to imprecise breakpoints, variation in
type and size, involvement of repetitive sequences, and general complexity of the induced changes to genomic
elements. Despite recent advances in the detection and characterisation of SVs, it remains difficult to assess
SVs beyond basic annotations and comparisons.
Here we introduce SVlog, a transparent and extensible meta-programming framework for analysing SVs. By
utilising the logic programming language Soufflé, SVlog provides an algorithm-free, declarative ontology
defining relationships among SVs and other elements. Genomic datasets are converted into relational facts, to
which SVlog applies composable deterministic rules to assess SVs without relying on stochastic "black box"
approaches.
Despite the compact codebase of the SVlog library, it currently evaluates more than 50 input predicates to
generate over 60 informative output predicates, enabling SV annotation, comparison and prioritisation. Our
tiered filtering strategy efficiently streamlines the identification of candidate pathogenic SVs in patients
with rare inherited disease. Applied to our disease cohort, SVlog successfully prioritised all previously
known pathogenic events, while also identifying novel candidates in several unsolved patients.
By focusing on explainability and modularity, SVlog offers a fast, reliable library for SV analysis and is a
powerful deterministic alternative to traditional bioinformatics pipelines for clinical variant curation.
A-G.30: Predicting gene expression profiles and drug targets from DNA methylation in gliomas
Track: Genomics, epigenomics, and genome editing
- Marcel Weinberg, Goethe University Frankfurt, Germany
-
Dennis Hecker, Goethe University Frankfurt, Germany
- Nikoletta Katsaouni, Goethe University Frankfurt, Germany
- Luca Malena Berger, Goethe University Frankfurt, Germany
- Katharina Weber, Goethe University Frankfurt, Germany
- Marcel Schulz, Goethe University Frankfurt, Germany
Presentation Overview: Show
Deregulation of DNA methylation is a hallmark of cancer. Studies that compare methylation profiles across
groups of samples, often referred to as epigenome-wide association studies (EWAS), generate growing evidence
that changes in methylation sites are of relevance as biomarkers and can help in developing disease
treatments. Especially in glioma, DNA methylation profiling is routinely performed as part of the clinical
diagnosis. The systematic interpretation of methylation sites can be challenging, especially for sites that
are not located in the promoter region of a gene or in expressed genes despite hypermethylated promoters.
Here, we reimplemented TDImpute (10.1093/gigascience/giaa076), a neural network predicting gene expression
from methylation profiles trained on The Cancer Genome Atlas. We applied the model on genome-wide methylation
data of glioma patients to predict expression profiles. This enabled us to identify genes whose expression
patterns differ significantly between patient groups and to find candidates that associate with cancer
sub-types. In addition, we ran the tool RANKOR, a machine learning model trained on drug-induced gene
expression changes that ranks potential candidate drugs for each sample based on the patient expression
signatures. RANKOR creates a latent space for the transcriptome profiles and the drugs' chemical structures,
which makes it applicable even to unseen compounds. With this work, we aim to deepen the understanding of
tumor biology by imputing gene expression profiles from DNA methylation data and translating this knowledge
into potential drug targets to foster individualized treatment.
A-G.31: Unbiased deconvolution of Oxford Nanopore Technologies long reads to reconstruct methylomes at cell
type level
Track: Genomics, epigenomics, and genome editing
-
Fanny Mollandin, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology,
Barcelona, Spain
-
Laura Moutard, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology,
Barcelona, Spain
-
Merce Planas-Felix, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology,
Barcelona, Spain
-
Jakob Admard, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen,
Germany
-
Eva Bru-Tari, Department of Genetic Medicine and Development, iGE3 and Centre facultaire du diabète,
University of Geneva, Geneva, Switzerland
-
Javier Garcia-Hurtado, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology,
Barcelona, Spain
-
Pere Santamaria, Institut D'Investigacions Biomediques August Pi i Sunyer, Barcelona, Spain
-
Pedro Herrera, Department of Genetic Medicine and Development, iGE3 and Centre facultaire du diabète,
University of Geneva, Geneva, Switzerland
-
Stephan Ossowski, Institute of Medical Genetics and Applied Genomics, University of Tübingen, Tübingen,
Germany
-
Jorge Ferrer, Centre for Genomic Regulation, The Barcelona Institute of Science and Technology, Barcelona,
Spain
Presentation Overview: Show
CpG methylation is a pivotal component of epigenetic landscapes that shape cell-specific transcriptional
programs. Oxford Nanopore Technologies (ONT) long read sequencing has emerged as a powerful method to
simultaneously profile 5-methylcytosine (5mC) as well as 5-hydroxymethylcytosine (5hmC), an intermediate of
the active demethylation pathway. Current methods that profile 5mC and 5hmC are largely restricted to bulk
tissue sequencing, often hiding cell-specific signals in heterogenous tissues.
To tackle this limitation, we developed a novel genome-wide unbiased deconvolution framework that clusters
reads based on their 5mC and 5hmC patterns, annotates cell-type specific clusters, and infers cell
type-specific methylome and hydroxymethylome. We evaluated the algorithm on in-silico mixtures of purified
cells and highly heterogeneous samples. We further applied the algorithm on human pancreatic islets, and
investigated cell specific changes of 5mC and 5hmC across age, disease and exposure to variable glucose
concentrations. In sum, this ONT-based methodology offers an opportunity to assess cell-specific 5mC and 5hmC
methylomes in complex tissue samples.
A-G.32: SIEVE: Sparse Interpretable Exome Variant Explainer
Track: Genomics, epigenomics, and genome editing
-
Davide Bagordo, Department of Biology and Biotechnology "L. Spallanzani", University of Pavia, Italy
-
Cezar Grigorean, Department of Biology and Biotechnology "L. Spallanzani", University of Pavia, Italy
-
Francesco Lescai, Department of Biology and Biotechnology "L. Spallanzani", University of Pavia,
Italy
Presentation Overview: Show
Deep learning methods for case-control variant discovery depend on functional annotations, treat variants as
unordered sets, and usually restrict the analysis to either rare or common variants. These choices leave
fundamental questions unaddressed: would the model discover different biology without annotations? Does
genomic position carry signal that permutation-invariant architectures discard? What is missed when common and
rare variants are not modelled jointly? We present SIEVE, a framework addressing five methodological gaps: (1)
sinusoidal positional encoding enables learning of position-dependent relationships; (2) attribution
regularisation encourages sparse, stable variant rankings during training; (3) an annotation-ablation protocol
trains the same architecture at four different annotation levels comparing discoveries systematically; (4)
frequency-agnostic processing analyses the full allele-frequency; (5) epistatic interactions are intrinsically
estimated from the attention layer at both the individual-level and variant-level, as well as by collapsing
variant attributions at gene-level. We applied SIEVE to a coronary artery disease whole-exome cohort from the
Ottawa Heart Genomics Study. Classification performance is consistent with expectations from exome variants,
and null-baseline permutation confirms that attributions exceed chance. Pairwise Jaccard overlap across
ablation-annotation levels is low (0.03–0.07), confirming that each level discovers genuinely different
biology. Annotation-free models identify genes implicated in TNF signalling, ubiquitin-proteasome regulation,
and endothelial apoptosis, i.e. pathways with established cardiovascular role. Adding positional encoding
alone shifts top-ranked genes toward HIF-1 signalling, FOXO pathway, and mitochondrial import. Top-ranked
attributions span the full MAF spectrum, confirming joint modelling of rare and common variation. The workflow
is implemented as a reproducible open-source Nextflow pipeline.
A-G.33: Impact of AI-assisted workflow standardization in a bioinformatics core facility
Track: Genomics, epigenomics, and genome editing
-
Jan Meier-Kolthoff, University of Augsburg, Augsburg Bioinformatics Core Facility, Germany
Presentation Overview: Show
Bioinformatics core facilities often face a growing demand for reproducible, scalable, and rapidly deployable
data analysis solutions across diverse project types. At the same time, many analyses are still implemented in
an ad hoc manner, leading to inconsistencies in workflow structure, documentation, and long-term
maintainability. Here, we describe the impact of AI-assisted workflow standardization in a university
bioinformatics core facility, with a focus on the use of Nextflow as a unifying workflow framework and large
language model (LLM)-based tools such as Claude Code and ChatGPT for rapid script and workflow generation.
We evaluated how AI-assisted development supported the design, refactoring, and documentation of standardized
pipelines for common bioinformatics tasks, including data preprocessing, quality control, variant analysis,
and downstream summarization. LLM-based tools were particularly useful for accelerating the generation of
boilerplate workflow components, helper scripts, configuration templates, and documentation drafts, while
human experts retained full responsibility for validation, benchmarking, and domain-specific adaptation. The
combination of Nextflow-based standardization and AI-assisted implementation reduced development time,
improved consistency across projects, and lowered the barrier for converting one-off analyses into reusable
workflows.
Our experience suggests that AI-assisted workflow engineering can substantially improve operational efficiency
in bioinformatics service environments when embedded in a controlled expert-driven framework. We argue that
this approach is especially valuable for core facilities, where reproducibility, maintainability, and rapid
turnaround are equally important.
A-G.34: NANO-SCOPE: Read-level tumor detection from cfDNA methylomes by Nanopore sequencing
Track: Genomics, epigenomics, and genome editing
-
Danial Tabbakh, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
-
Francesca Maria Stefanizzi, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
- Camille Cialini, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
- Petr Nazarov, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
-
Katrin Frauenknecht, Institute for Neuropathology, University Medical Center of the Johannes Gutenberg
University Mainz, Germany, Germany
-
Vladimir Despotovic, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
- Anna Golebiewska, Luxembourg Institute of Health, Strassen, Luxembourg, Luxembourg
- Urszula Lawrynowicz, Medical University of Gdańsk, Poland, Poland
- Reka Toth, Luxembourg Institute of Health, Strassen, Luxembourg,
Presentation Overview: Show
Nanopore sequencing enables direct, conversion-free detection of DNA methylation from low-input circulating
cell-free DNA (cfDNA), making it a promising platform for liquid biopsy analysis. In cancer, cfDNA contains a
complex mixture of tumor-derived and non-tumor fragments, where tumor-associated methylation patterns reflect
tissue of origin and molecular subtype. Robust inference from these data remains challenging due to signal
sparsity and the heterogeneous composition of individual samples. To address this, we developed NANO-SCOPE, a
deep learning framework for distinguishing tumor-derived from normal-origin cfDNA fragments at the read level.
The method combines two complementary components. First, a methylation-aware transformer based on DNABERT-2
generates contextual read embeddings by integrating sequence, methylation, and genomic position information.
Second, a one-class autoencoder trained exclusively on non-tumor cfDNA learns the manifold of normal fragments
and assigns each read a reconstruction score (RS) that quantifies deviation from this normal profile. An
attention-based aggregation module then integrates read embeddings and RS values for robust sample-level
predictions of tumor presence.
Applied to Nanopore cfDNA datasets from patients with lung cancer, NANO-SCOPE distinguished cancer cases from
healthy controls and enabled origin-aware filtering of tumor-associated reads. This filtering enriched the
tumor-related methylation signal and improved the interpretability of downstream analyses, including
tissue-of-origin inference and molecular subtype prediction. NANO-SCOPE provides a scalable framework for
sensitive liquid biopsy analysis with potential applications in early detection, relapse monitoring, and
molecular stratification.
A-G.35: Modeling Longitudinal Cancer Dynamics from Omics Data: a Systematic Review of Current Limitations and
the Road Toward Temporal Generation
Track: Genomics, epigenomics, and genome editing
-
Guillermo Prol Castelo, Barcelona Supercomputing Center (BSC-CNS), Spain
- Davide Cirillo, Barcelona Supercomputing Center, Spain
- Alfonso Valencia, Barcelona Supercomputing Centre BSC, Spain
Presentation Overview: Show
The variational autoencoder (VAE) has become a prominent tool in cancer omics research, valued for its ability
to learn structured embeddings of high-dimensional, heterogeneous biological data. Yet a systematic
understanding of how VAEs and related deep representation learning (DRL) methods engage with the temporal
aspects of cancer in omics data to capture longitudinal molecular dynamics of tumor progression has been
lacking. We present such an assessment. Screening 440 publications from 2014 to 2024, 21 directly relevant
studies were identified. DRL methods are predominantly applied to cancer omics data for subtyping, diagnosis,
and prognosis. However, these tasks do not explicitly leverage the longitudinal molecular dimension of disease
progression. Longitudinal omics studies of primary cancer are scarce, constrained by practical, ethical, and
biological barriers, including the destructive nature of sequencing technologies and inter-patient
heterogeneity. Temporal dimension in the omics data are most commonly represented through pseudo-time
inference with single-cell data, not necessarily reflecting longitudinal molecular dynamics. Alternatively,
cancer stages may be used as a proxy time axis. Against this landscape, the VAE emerges as uniquely suited to
bridge the gap between available cross-sectional omics data and the need for longitudinal molecular modeling.
Its latent space is directly operable: sample representations can be interpolated and fed to the decoder to
produce realistic intermediate omics profiles. We discuss the methodological requirements needed to unlock
this potential, and briefly present an application focusing on generating and forecasting synthetic cancer
stage trajectories from transcriptomic data in renal cell carcinoma.
A-G.36: TOGA2 delivers scalable and precise gene annotation and ortholog identification across vertebrate
genomes
Track: Genomics, epigenomics, and genome editing
-
Yury Malovichko, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe
University Frankfurt, Germany
-
Bernhard Bein, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe
University Frankfurt, Germany
-
Alejandro Gonzales-Irribarren, Senckenberg Research Institute; Department of Computer Science, Leipzig
University, Germany
- Evgeny Leushkin, Senckenberg Research Institute, Germany
- Leon Hilgers, Senckenberg Research Institute, Germany
- Amy Stephen, Senckenberg Research Institute, Germany
- Xueling Yi, Senckenberg Research Institute, Germany
-
Michele Albertini, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe
University Frankfurt, Germany
-
Tim Stadager, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
-
Markus Zumpt, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
-
Luca Hoppach, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
- Felix Götz, Senckenberg Research Institute, Germany
-
Niklas Himstedt, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
-
Lucas Koch, Institute of Cell Biology and Neuroscience, Goethe University Frankfurt, Germany
-
Michael Hiller, Senckenberg Research Institute; Institute of Cell Biology and Neuroscience, Goethe
University Frankfurt, Germany
Presentation Overview: Show
Accurate annotation of coding genes and inference of orthology relationships in newly sequenced genomes remain
central challenges in modern genomics. We present TOGA2, the next generation of the TOGA (Tool to infer
Orthologs from Genome Alignments) framework for scalable reference-based orthologous gene annotation. By
introducing an exon-wise gene annotation approach, TOGA2 achieves a 513-fold memory reduction and a 6-fold
runtime reduction compared to its predecessor, while improving exon-level annotation precision. We demonstrate
that deep learning models for splice-site prediction trained on human data remain effective across diverse
vertebrates. Incorporating these predictions further enhances exon boundary annotation precision in TOGA2 and
enables detection of evolutionary changes affecting gene exon–intron structure, such as splice-site shifts
or intron gains and losses. Multiple false-positive prediction filters and a new gene tree–based
reconciliation step further improve ortholog inference, specifically identifying new 1:1 ortholog pairs. We
further show how features integrated into TOGA2 improve annotation of immunoglobulin and T-cell receptor gene
segments, processed pseudogenes, and functional retrogenes, and how synteny information from inferred
orthologs can be leveraged for ancestral chromosome reconstruction and phylogenomic analyses. To demonstrate
scalability across multiple genomes, we provide a comprehensive comparative genomics resource for >900
mammal and >680 bird assemblies, including gene annotations, ortholog sets, retrogene candidates, and codon
alignments.
A-G.37: GAP-MS: Automated validation of gene predictions using integrated mass spectrometry evidence
Track: Genomics, epigenomics, and genome editing
-
Qussai Abbas, Technical University of Munich, Germany
- Mathias Wilhelm, Technical University of Munich, Germany
- Bernhard Kuster, Technical University of Munich, Germany
- Dimitri Frishman, Technical University of Munich, Germany
Presentation Overview: Show
Accurate genome annotation is fundamental to modern biology, yet distinguishing authentic protein-coding
sequences from prediction artifacts remains challenging, particularly in complex plant genomes. We present
GAP-MS, an automated proteogenomic pipeline that leverages mass spectrometry evidence to systematically
validate the protein-level accuracy of predicted gene models. Applied across 9 major crop species, GAP-MS
consistently improved prediction precision for four widely used tools, with substantial gains of up to 32%.
Beyond filtering artifacts, the pipeline identified over 9000 novel peptide-supported gene models absent from
current RefSeq annotations. Furthermore, GAP-MS successfully corrected structural errors by identifying
independent translation initiation and termination sites via specific N- and C-terminal peptides. These
results demonstrate that direct proteomic evidence provides a robust framework for resolving annotation
ambiguities and defining high-confidence reference proteomes.
A-G.38: Embedding-based statistical framework enables quantitative gene function representation and
hypothesis testing with large language models
Track: Genomics, epigenomics, and genome editing
- Yanhao Tan, UPMC Hillman Cancer Center, University of Pittsburgh, United States
- Li-Ju Wang, UPMC Hillman Cancer Center, University of Pittsburgh, United States
-
Yu-Chiao Chiu, UPMC Hillman Cancer Center, University of Pittsburgh, United States
Presentation Overview: Show
Accurately delineating gene function is central to interpreting high-throughput genomic data, yet most
existing methods depend on predefined gene sets and largely qualitative interpretations. Large language models
(LLMs) provide a promising alternative by extracting functional relationships from biological text, though
their ability to support quantitative analysis has not been fully established. We introduce a statistical
framework that leverages LLM-derived embeddings of genes and biological functions to enable quantitative
assessment of gene-gene and gene-function relationships across diverse biological settings. We evaluated seven
leading embedding models using both curated gene annotations and literature-derived descriptions. OpenAI's
text-embedding-3-large and Google's gemini-embedding-001 showed the strongest performance, recovering
gene-gene relationships in up to 98% of Gene Ontology biological processes and approximately 99% of canonical
pathways. Gene-function association analyses further demonstrated high sensitivity (95-98%) and specificity
(73-84%). Importantly, this framework enables rigorous testing of functional hypotheses generated from gene
lists without requiring predefined annotations. It consistently differentiates biologically coherent gene sets
from noise, surpassing both confidence-based LLM outputs and traditional enrichment methods. Application to
drug response data uncovered candidate pathways linked to cancer immune sensitization and supported systematic
exploration of drug mechanisms. Overall, our results position LLM-based embeddings as a scalable quantitative
tool for functional genomics.
A-G.39: Unmasking the Trypanosoma cruzi Genome: A Comprehensive Map of the MASP Superfamily.
Track: Genomics, epigenomics, and genome editing
-
Aldana Alexandra Cepeda Dean, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET,
Argentina
-
Luisa Berná, Laboratorio de Genómica Evolutiva, Facultad de Ciencias, Universidad de la República,
Montevideo, Uruguay
-
Javier De Gaudenzi, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET, Argentina
-
Virginia Balouz, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET, Argentina
-
Carlos Andrés Buscaglia, Instituto de Investigaciones Biotecnologicas, EBYN-UNSAM, CONICET, Argentina
Presentation Overview: Show
Trypanosoma cruzi, the causative agent of Chagas disease, harbors one of the most complex eukaryotic genomes,
dominated by massive expansion and diversification of multigene families encoding surface virulence factors.
Among them, mucin-associated surface proteins (MASPs) represent a paradigmatic case of genomic plasticity, yet
their true diversity has remained obscured by fragmented assemblies and the limitations of standard annotation
pipelines.
We developed a dedicated bioinformatic framework combining a curated sequence database, HMM-based motif
detection, and clustering with CD-HIT at 80% similarity to enable de novo annotation and classification of
MASP sequences across 13 T. cruzi strains spanning all major evolutionary lineages. This approach uncovered
~1000 canonical MASP genes, ~30 chimeric variants, and >500 pseudogenes depending on lineage, outperforming
conventional pipelines. Comparative analyses revealed marked differences in MASP repertoires, with
lineage-specific innovations, tandem duplication-driven expansions, and a conserved core of MASP sequences
shared across strains.
Integration of RNA-seq data showed that over 70% of the MASP repertoire is transcriptionally active in
human-infective stages. This proportion is unexpectedly high, indicating that nearly the entire family
contributes to antigenic diversity and immune evasion. Distinct transcriptomic signatures further separate
canonical and chimeric variants, as well as MASP genes from pseudogenes.
Our findings highlight MASPs as key contributors to T. cruzi immune evasion, while demonstrating how tailored
computational strategies can resolve highly repetitive, structurally complex genomes. The framework is
generalizable for dissecting large gene families in pathogens, bridging genotype to phenotype and advancing
our understanding of genome evolution, plasticity, and host-pathogen interaction.
A-G.40: NeoEvo-AIS: Rethinking Tumor Immunogenicity at the Tumor Level
Track: Genomics, epigenomics, and genome editing
-
Ankita Singh, Center for Applied and Translational Genomics (CATG) . MBRU, Dubai Health, Dubai, UAE,
United Arab Emirates
-
Youssef Ahmed Elkenawi, Center for Applied and Translational Genomics (CATG) . MBRU, Dubai Health, Dubai,
UAE, United Arab Emirates
- Costerwell Khyriem, University of Freiburg, Germany
- Sven Hauns, University of Freiburg, Germany
-
Filippo Castiglione, Institute for Applied Computing (IAC) National Research Council of Italy (CNR),Rome,
Italy
- Rolf Backofen, University of Freiburg, Germany
-
Omer S. Alkhnbashi, Mohammed Bin Rashid University of Medicine and Health Sciences (MBRU), United Arab
Emirates
Presentation Overview: Show
Tumor mutation burden (TMB) and predicted neoantigen counts are often used as measures to estimate tumor
immunogenicity; however, they simplify a complex biological landscape into a single metric. By mainly focusing
on mutation quantity, these measures often miss important aspects such as clonality, antigen quality, and
intratumoral diversity, which can lead to an inability to distinguish tumors with similar mutation burdens but
significantly different immune responses.
To overcome these limitations, we have developed NeoEvo-AIS, a computational framework designed to provide a
more comprehensive view of tumor antigenic properties. Central to this framework is the Antigenic Instability
Score (AIS-STATIC), a tumor-level metric that combines biologically relevant features. Using pan-cancer
somatic mutation data from The Cancer Genome Atlas (TCGA) MC3 harmonized call set, we assessed mutation
architecture and inferred clonality by incorporating variant allele frequency alongside estimates of copy
number and tumor purity. Candidate neoantigen peptides were predicted with NetMHCpan-4.1 and MHCflurry, while
immunogenicity was evaluated through deep learning models that consider peptide foreignness and structural
features important for T-cell recognition. Antigenic diversity was measured using entropy-based metrics across
mutated genes and inferred clonal populations.
By integrating these features into a single interpretable score, AIS-STATIC effectively captures major
differences in tumor antigenic architecture that are not solely reflected by TMB. Across various tumor types,
those with higher AIS-STATIC scores show increased immune activity and greater immune cell infiltration.
Therefore, NeoEvo-AIS offers a practical framework for studying tumor-immune interactions and sets the stage
for future research into how antigenic properties evolve over time.
A-G.41: TFClassPredict: A Deep Learning Framework for Transcription Factor Binding Site Analysis Using
Evolutionarily Conserved DNA-Binding Domain Annotations
Track: Genomics, epigenomics, and genome editing
-
Christian Ickes, Institute of Cardiovascular Physiology, University Medical Center Göttingen,
Germany, Germany
-
Cigdem Hazal Timucin, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
-
Bendix Christian Harms, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
-
Umut Akgül, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
-
Inigo Vincente Hernandez, Department of Gastroenterology, University Medical Center Göttingen, Germany,
Germany
-
Ivan Bogeski, Institute of Cardiovascular Physiology, University Medical Center Göttingen, Germany,
Germany
-
Tim Beißbarth, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
-
Martin Haubrock, Medical Bioinformatics, University Medical Center Göttingen, Germany, Germany
Presentation Overview: Show
Transcription factors (TFs), the core proteins of transcription initiating processes, regulate gene expression
by binding to short genomic sequences, known as transcription factor binding site (TFBS), through defined
DNA-binding domains (DBDs). These interactions are central to understand gene regulation, which underlies many
cellular processes. Although many models exist for predicting transcription factor binding, none provide a
comprehensive framework that accounts for the similarity of binding characteristics among TFs within the same
DBD class. In fact, several of these families are severely under represented in most current analysis
pipelines.
Our model, TFClassPredict, introduces a novel approach to identify transcription factor binding sites based on
structural annotations of evolutionarily conserved DBDs. By leveraging canonical binding patterns,
TFClassPredict provides high-confidence predictions essential for gene regulation analysis. Fine-tuned from
the DNABERT model, TFClassPredict classifies DNA sequences across 23 classes of DBDs. TFClassPredict achieved
strong performances, naturally aggregates predicted sites by DBD class, improving both interpretability and
the balanced representation of all DBD classes. These results demonstrate that TFClassPredict constitutes a
reliable, family aware framework for uncovering regulatory differences and for advancing comprehensive TF
binding analyses.
A-G.42: End-to-end deep learning methods for genetic risk prediction of Schizophrenia
Track: Genomics, epigenomics, and genome editing
-
Nora Verplaetse, KULeuven, Belgium
- Yves Moreau, Katholieke Universiteit Leuven, Belgium
- Daniele Raimondi, Institut de Génétique Moléculaire de Montpellier, France
Presentation Overview: Show
Schizophrenia is a highly heritable psychiatric disorder with a complex genetic basis. While recent Genome
Wide Association Studies (GWAS) and Whole Exome Sequencing (WES) have identified numerous risk loci, existing
clinical prediction models rely primarily on linear assumptions, limiting their ability to capture complex,
nonlinear genetic effects such as epistasis.
In this study, we apply a new paradigm of end-to-end Genome Interpretation (GI) Neural Network (NN) models to
predict Schizophrenia risk from WES data in a large cohort of 6,135 cases and 6,245 controls. We show that
nonlinear NNs significantly outperform conventional additive models, when sufficient sample size is available.
These findings further support the fact that high-order genetic interactions between alleles and variants
should be considered by clinical and quantitative genetics models.
To investigate the decision process our models follow, we integrate Explainable AI (XAI) techniques and
biological priors into our models, using them to identify predictive genes and pathways. Our approach recovers
both known Schizophrenia risk genes and recommends BASP1 as a potential understudied Schizophrenia gene
involved in neuronal development.
A-G.43: Contrastive Learning of Multiple Sequence Alignments for Phylogenetic Inference
Track: Genomics, epigenomics, and genome editing
-
Jens-Uwe Ulrich, Robert Koch Institute, Germany
- Denise Kühnert, Robert Koch Institute, Germany
Presentation Overview: Show
Phylogenetic inference from large-scale multiple sequence alignments (MSAs) is essential for understanding
evolutionary relationships, tracking pathogen outbreak dynamics, and quantifying biodiversity patterns across
the tree of life. Traditional distance-based and likelihood-based methods struggle to scale efficiently while
capturing the complex hierarchical relationships inherent in evolutionary data. Thus, low dimensional
representations of MSAs are needed to infer phylogenies from large MSAs. We present a self-supervised
contrastive learning framework that learns compact, phylogenetically informative representations of MSA
patches without requiring explicit phylogenetic labels.
Our architecture combines a multi-scale convolutional encoder with biologically motivated data augmentations
tailored to MSA characteristics. The encoder processes local alignment patches and projects them into a
normalized latent space optimized via an NT-Xent loss. Crucially, we designed domain-specific augmentations
that reflect biological invariances: row shuffling captures sequence order independence, column masking
simulates alignment uncertainty and missing data, row subsampling models varying taxonomic sampling, and
reverse complementation accounts for strand ambiguity. This augmentation strategy enforces robust feature
learning while maintaining phylogenetically relevant sequence information.
We evaluated our framework assessing multiple aspects of representation quality: hierarchical clustering
fidelity through cophenetic correlation, discrete evolutionary group recovery via k-nearest neighbor
classification, continuous sequence similarity preservation through correlation analysis, and latent space
geometry metrics including uniformity and intrinsic dimensionality. Our analysis revealed fundamental
trade-offs between learning discrete phylogenetic groupings versus preserving fine-grained evolutionary
distances.
This work establishes contrastive learning as a viable approach for phylogenetic representation learning,
opening avenues for scalable phylogenetic placement, distance matrix approximation, and integration with
downstream evolutionary analyses.
A-G.44: PanelClone: signature-aware subclonal reconstruction that scales with mutation count
Track: Genomics, epigenomics, and genome editing
-
Sungjin Park, Institute of Human Behavior & Genetics, College of Medicine, Korea University,
Seoul, Republic of Korea, South Korea
-
Sangwon Um, Institute of Human Behavior & Genetics, College of Medicine, Korea University, South Korea
-
Taehun Kim, AI Center, Advanced Medical Imaging Institute, Anam Hospital, College of Medicine, Korea
University, South Korea
- Hyowon Lee, DigitalBio R&D Center, Korea University Anam Hospital, South Korea
-
Cheol Soon Lee, Institute of Human Behavior & Genetics, College of Medicine, Korea University, South
Korea
Presentation Overview: Show
Subclonal reconstruction is essential for understanding tumor evolution and predicting treatment resistance,
yet reliable deconvolution has required the hundreds-to-thousands of somatic mutations available only from
whole-genome/exome sequencing—leaving clinical deep-panel assays (300–500 genes; 5–50 mutations/sample)
underserved. We present PanelClone, a single-sample framework that maximizes subclonal resolution by jointly
exploiting two orthogonal signals—variant frequency and mutational signature—and by adapting inference
stringency to the available mutation count. PanelClone combines three components. (1) Multiplicity-aware CCF
inference with explicit neutral-tail modeling: a beta-binomial mixture corrects for purity and local copy
number, marginalizes mutation multiplicity, and models neutral passenger mutations as a truncated 1/f
power-law tail, suppressing spurious (“ghost”) subclones. (2) A dual-axis ghost gate: because distinct
subclones often arise under distinct mutational processes, PanelClone treats each mutation’s trinucleotide
signature as an axis orthogonal to frequency, letting it separate subclones that overlap in CCF and rescue
signature-distinct subclones buried in the neutral tail—cases that frequency-only methods (PyClone-VI,
MOBSTER) merge or miss and that signature-only methods (CloneSig) over-call. (3) Tiered,
mutation-count-adaptive selection: sparse Bayesian (MAP-EM) clustering with automatic cluster-number
determination that degrades gracefully from full clustering to binary clonal/subclonal classification to
individual-CCF reporting as counts fall. In simulation (BAMSurgeon-style data with injected copy-number,
purity and trinucleotide-signature structure), evaluated with SMC-Het metrics (V-measure, CCF MAD,
co-clustering), the dual-axis design recovers subclones that frequency-only baselines miss and avoids the
over-segmentation of signature-only approaches, with the advantage concentrated where subclones differ in
mutational process. Extension to clinical panel scale—including population-based haplotype phasing across
panel targets and paired bulk RNA-seq for genome-wide copy-number and allele-specific-expression
validation—and benchmarking against matched multi-region cohorts are underway. PanelClone will be made
publicly available upon publication.
A-G.45: Large cohorts analysis to identify pathogenic digenic interactions in autoinflammatory
disorders
Track: Genomics, epigenomics, and genome editing
-
Jaume Reig-Palou, Universitat de Barcelona, Spain
- Diego Garrido-Martin, Universitat de Barcelona, Spain
- Llorenc Villalonga-Gonyalons, Universitat Pompeu Fabra, Spain
- Alex Velaza-Gil, Universitat de Barcelona, Spain
- Nerea Moreno-Ruiz, Universitat Pompeu Fabra, Spain
- Roderic Guigo, Centre for Genomic Regulation, Spain
- Hafid Laayouni, Universitat Pompeu Fabra, Spain
- Juan Ignacio Arostegui, Hospital ClÃnic de Barcelona, Spain
- Ferran Casals, Universitat de Barcelona, Spain
Presentation Overview: Show
The role of the digenic model in rare disorders has been relatively under explored, mostly due to both the
difficulties of statistically supporting these associations to disease and the limitations of generating
experimental models. We have developed an approach based on the analysis of genomic data from hundreds of
thousands of individuals available from large biomedical projects to detect or validate pathogenic digenic
interactions.
We analyzed genotyping data for 70 autoinflammatory disorder genes in UK Biobank for 467,000 European ancestry
individuals based on PCA analyses. We compared the expected and observed frequencies of co-carriers for all
possible pairs of variants from different genes.
The analysis revealed candidate digenic interactions among AID genes, evidenced by deviations from neutrality
of the genotypes distribution. One of the interactions, validated in an independent patient cohort, may
provide insight into the incomplete penetrance of Familial Mediterranean Fever (FMF) where heterozygous
individuals carrying pathogenic variants in the MEFV gene may or may not develop the disease. Specifically, we
suggest that a genetic variant in a functionally related gene acts as a modifier, influencing disease
expression by modifying the action of the primary pathogenic variant in MEFV.
Our approach can identify digenic and modifier interactions in rare diseases. We have identified a candidate
modifier gene that may explain incomplete penetrance in FMF.
A-G.46: PlanMine: An Interactive Database for Exploring Planarian Genomics and Regeneration
Track: Genomics, epigenomics, and genome editing
-
Muhammad Rizwan Riaz, Max Planck Institute for Multidisciplinary Sciences, Germany
-
Hongkee Moon, Max Planck Institute of Molecular Cell Biology and Genetics Dresden, Germany
-
Jeremias Brand, Max Planck Institute for Multidisciplinary Sciences Göttingen, Germany
-
Martin Leandro Paleico, Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen, Germany
-
Julian Kunkel, Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen (GWDG), Germany
- Jochen Rink, Max Planck Institute for Multidisciplinary Sciences Göttingen, Germany
Presentation Overview: Show
Schmidtea mediterranea is a free-living freshwater planarian and a powerful model for studying regeneration,
stem cell biology, and the evolution of complex traits. As an early-diverging bilaterian, it combines simple
morphology with features of complex organisms, including centralized brain, nervous system, and coordinated
behaviors. Over the past decade, high-throughput sequencing has generated extensive molecular resources,
including high-quality genomes, transcriptomes, and expression datasets. These developments underscore the
need for an integrated data exploration platform for the planarian research community. PlanMine has been
developed by our lab to address this need and to evolve alongside the advancements in the model system.
PlanMine is built on the InterMine framework, enabling the integration of diverse datasets along with flexible
tools for querying and visualization. It uses a PostgreSQL backend, incorporates custom BLAST databases via
SequenceServer, and supports genome navigation through JBrowse and the UCSC Genome Browser. Together, these
components form a unified, extensible platform for the planarian research community.
PlanMine offers a comprehensive, user-friendly interface for accessing and analyzing planarian molecular data.
The updated version (v4.0- projected release: October 2026) incorporates the latest chromosome-scale genome
assembly and annotations of the model species, S. mediterranea, along with additional planarian
transcriptomes. Users can explore gene models between the new and legacy annotations, homologues, domain
content, pathways, and bulk and single-cell expression patterns derived from experimental, computational, and
curated sources. By centralizing data, advanced search and visualization, PlanMine enables integrative
analyses and hypothesis generation and thus valuable services for the further development of planarians as
model taxon.
A-G.47: Genome Enhancer reveals an IGF-centered regulatory program underlying shared epigenetic mechanisms in
Kabuki syndrome subtypes
Track: Genomics, epigenomics, and genome editing
-
Daria Stelmashenko, geneXplain GmbH, Wolfenbuettel, Germany; Center for Integrative Biology (CIBIO),
University of Trento, Trento, Italy, Germany
-
Kel Alexander, geneXplain GmbH, Wolfenbuettel, Germany; Royal College of Surgeons in Ireland (RCSI),
Dublin, Ireland, Germany
-
Giuseppe Merla, Department of Molecular Medicine and Medical Biotechnology, University of Naples Federico
II, Naples, Italy, Italy
-
Emile Danvin, Department of Molecular Medicine and Medical Biotechnology, University of Naples Federico
II, Naples, Italy, Italy
-
Antonio Ammendola, Department of Molecular Medicine and Medical Biotechnology, University of Naples
Federico II, Naples, Italy, Italy
Presentation Overview: Show
Genome-wide DNA methylation data provide an indirect readout of regulatory activity, requiring translation of
CpG-level changes into mechanistic models. Here, we present Genome Enhancer, an automated pipeline that
converts epigenomic profiles into causal networks and identifies master regulators by integrating cgID-to-gene
signature conversion, TFBS enrichment, composite module detection, and upstream network reconstruction,
thereby linking epigenetic changes to transcription factors, signaling pathways, and candidate targets; its
application is demonstrated in uncovering shared regulatory mechanisms in Kabuki syndrome subtypes.
Kabuki syndrome is a rare developmental disorder caused by LOF mutations in KMT2D (KS1) and KDM6A (KS2). By
independently analyzing KS1 and KS2 blood DNA methylation data against paired healthy controls, we identified
65 hyper- and 49 hypomethylated genes shared between subtypes. These shared signatures were analyzed with
Genome Enhancer using TRANSFAC-based TFBS enrichment, CompositeModuleAnalyst for combinatorial modules, and
TRANSPATH for upstream network reconstruction. The analysis revealed a common IGF-centered upstream regulatory
program in which IGF1R, IGF1,IGF2, and IGFBP4 activate downstream modules mediated by SMAD3,TEAD1, and RAD21,
linking growth factor signaling to enhancer-driven gene regulation. In the context of KMT2D/KDM6A dysfunction,
this machinery is likely impaired, leading to reduced transmission of IGF-dependent developmental signals into
stable transcriptional activation and resulting in coordinated hypermethylation and silencing of key
developmental genes. This connects methylation findings directly to one of the core clinical dimensions of
Kabuki: abnormal growth and developmental signaling, suggesting that KS1 and KS2 may converge not only at the
phenotype level, but also at the level of a common growth-factor-dependent transcriptional control
architecture.
A-G.48: Improved reconstruction of transcripts and coding sequences from RNA-seq data
Track: Genomics, epigenomics, and genome editing
-
Jan Grau, Institute of Computer Science, Martin Luther University Halle-Wittenberg, Germany
-
Deborah Weise, Institute of Computer Science, Martin Luther University Halle-Wittenberg, Germany
-
Marika Panster, Institute of Biology, Martin Luther University Halle-Wittenberg, Germany
-
Martin H Schattat, Institute of Biology, Martin Luther University Halle-Wittenberg, Germany
-
Jens Keilwagen, Julius Kühn-Institut (JKI) - Federal Research Centre for Cultivated Plants, Germany
Presentation Overview: Show
Accurate annotation of gene and transcript models in newly sequenced genomes is a pivotal requirement for many
subsequent analyses. We present GeMoSeq, a novel approach for RNA-seq-based gene prediction. GeMoSeq shall
complement homology-based (e.g., GeMoMa) predictions and, hence, focuses on the prediction of protein-coding
genes.
Starting from genomic mappings of RNA-seq reads, GeMoSeq partitions the genome into covered regions, builds a
read graph with basepair resolution connecting positions that are adjacent in mapped reads. For each connected
component of the read graph, GeMoSeq then merges consecutive positions without alternative edges into a
splicing graph, and combinatorially enumerates candidate transcripts. Candidate transcripts may still be
chimeras of multiple transcripts, which are split considering coverage, CDS prediction and splice site
orientation. Resulting potential transcripts are quantified based on RNA-seq evidence in an EM-like algorithm,
filtered by abundance, and finally merged to genes.
We benchmark GeMoSeq against state-of-the-art tools on a large collection of RNA-seq libraries of seven
species using the respective reference annotations as ground truth. For the F1 measure on the level of CDSs,
which have been in the focus of GeMoSeq development, we observe that GeMoSeq yields better predictions than
all previous approaches for almost all data sets, where the improvement is specifically pronounced for S.
cerevisiae, C. elegans and A. thaliana.
We finally combine RNA-seq-based prediction of GeMoSeq with homology-based prediction of GeMoMa to reannotate
two recently sequenced genomes of N. benthamiana lab strains and validate several predictions experimentally.
A-G.49: Chromatin-informed sparsification of neural networks improves gene expression prediction and
regulatory interaction recovery
Track: Genomics, epigenomics, and genome editing
-
Maxim Evsioukov, Goethe University Frankfurt, Germany
- Dennis Hecker, Goethe University Frankfurt, Germany
- Shamim Ashrafiyan, Goethe University Frankfurt, Germany
- Marcel Schulz, Goethe University Frankfurt, Germany
Presentation Overview: Show
Gene expression prediction from epigenomic data links regulatory elements to transcriptional output.
Multilayer perceptron (MLP) models rely primarily on histone modification signals at candidate cis-regulatory
elements (CREs) and often ignore genomic and 3D chromatin context. These models are often overparameterized
and may interpret correlated signals as causal regulatory effects. These limitations motivate the
incorporation of biological priors.
We introduce a biologically guided pruning framework to systematically evaluate sparsification strategies in
gene expression prediction models. Using H3K27ac signals at CREs, we train MLPs on bulk RNA-seq data and
impose sparsity via three approaches: data-driven pruning, pruning guided by linear genomic distance between
CREs and transcription start sites (TSSs) of target genes, and pruning informed by Hi-C chromatin interaction
frequencies between CREs and TSSs.
On a benchmark of 200 genes, Hi-C-guided pruning achieves the strongest predictive performance, improving
median correlation and reducing mean squared error. Stratifying genes by predictability (threshold r = 0.7)
reveals heterogeneous effects, with larger gains in low-correlation genes and smaller improvements in
high-correlation genes.
We then apply this approach to 1,574 genes from an enhancer-gene interaction map and evaluate biological
validity. Hi-C-pruned models improve recovery of functional CRE-gene regulatory relationships, increasing AUC
from 0.159 to 0.180 for CRISPRi-validated links and showing stronger enrichment of eQTL-supported associations
across multiple tissues from the GTEx Consortium, relative to the unpruned model.
These results demonstrate that chromatin-informed sparsification improves predictive performance and recovery
of experimentally supported regulatory interactions.
A-G.50: Ancestry calibrated polygenic risk score for breast cancer risk stratification in an admixed
Colombian cohort
Track: Genomics, epigenomics, and genome editing
-
Yina Tatiana Zambrano, Biosciences - SURA Colombia, Colombia
- Danny Styvens Cardona Pineda, Biosciences - SURA Colombia, Colombia
- Harvy Mauricio Velzco, Biosciences - SURA Colombia, Colombia
Presentation Overview: Show
Breast cancer (BC) polygenic risk scores (PRS) derived from European GWAS often lack transferability to
admixed Latin Americans due to tri-ancestral genomic backgrounds, European (EUR), Native American (NAM), and
African (AFR), which alter allele frequencies and linkage disequilibrium. Ancestry-unaware models risk
systematic miscalibration, hindering equitable clinical implementation.
We enrolled 1,997 Colombian women (ages 40–65) in a translational precision medicine program at Biosciences
unit Grupo SURA (2022–2024). Genotyping (Illumina GSA v3) was followed by phasing (Eagle v2.4.1) and
imputation (TOPMed Freeze 8). Ancestry proportions were inferred via ADMIXTURE (K=3) using the HGDP panel.
Risk modeling integrated: (i) a genome-wide PRS from established BC loci, (ii) principal components as
calibration covariates, and (iii) modifiable clinical variables. To date, 32 incident BC cases have been
identified among the cohort, with ~700 baseline controls under prospective follow-up.
Ancestry inference revealed a tri-ethnic profile (EUR: 0.59, NAM: 0.30, AFR: 0.11). The integrated
genomic-clinical model achieved an AUC of 77.6% (OR per SD: 1.87; 95% CI: 1.68–2.07), with 72.6% accuracy,
67.7% sensitivity, and 74.3% specificity. Notably, the model yielded a Net Reclassification Improvement (NRI)
of 48% and an NPV of 86.9%.
Integrating ancestry calibrated PRS with clinical factors substantially enhances BC risk discrimination in
admixed populations. With 32 prospectively ascertained cases, this cohort serves as a critical resource for
validating polygenic risk in Latin America, underscoring the necessity of diverse genomic data to achieve
precision oncology.
A-G.51: Targeted long-read sequencing for high-resolution repeat profiling in myotonic dystrophy type
1
Track: Genomics, epigenomics, and genome editing
-
Yoojung Han, Interdisciplinary Program in Bioinformatics, Seoul National University, South
Korea
-
Ja-Hyun Jang, Department of Laboratory Medicine and Genetics, Samsung Medical Center, Sungkyunkwan
University School of Medicine, South Korea
- Hyeshik Chang, School of Biological Sciences, Seoul National University, South Korea
Presentation Overview: Show
Tandem repeat expansion disorders can be difficult to diagnose when expansions exceed 200 repeats, as standard
methods (for example, Southern blot and modified PCR) often fail. We present a Cas9-targeted nanopore
sequencing workflow and an automated analysis pipeline, RepeatLab, for accurate repeat-length estimation,
structure assessment, and high-resolution methylation profiling. Validated on 13 myotonic dystrophy type 1
samples, 4 healthy controls, and 4 cell lines, this approach demonstrates improved sensitivity and accuracy
for large expansions. Key refinements include an alternative basecalling strategy for extended repeats and a
repeat-length calling algorithm that remains robust at lower sequencing throughput. The platform also
automatically reports methylation near the DMPK repeat region, including five CpG site groups that could
inform more nuanced clinical evaluations. This integrated workflow offers a rapid, cost-effective diagnostic
solution with a turnaround time under 24 h and costs comparable to standard assays. Its compatibility with
readily available computational resources enhances accessibility and scalability.
A-G.52: Pattern-based segmentation of the methylome reveals chemotherapy-associated changes in ovarian
cancer
Track: Genomics, epigenomics, and genome editing
-
Susanna Holmström, University of Helsinki, Finland
- Giovanni Marchi, University of Helsinki, Finland
- Jaana Oikkonen, University of Helsinki, Finland
- Sampsa Hautaniemi, University of Helsinki, Finland
Presentation Overview: Show
Selective pressures, including chemotherapy, shape tumor evolution by favoring phenotypes with increased
fitness. In ovarian high-grade serous carcinoma (HGSC), a subset of patients develops chemoresistant tumors,
markedly reducing survival. Despite extensive genomic characterization, the mechanisms underlying
chemoresistance remain only partially understood. Consequently, recent research has expanded toward
alternative regulatory layers, including DNA methylation (DNAm).
DNAm regulates transcription through activation or silencing of regulatory elements. However, the specific
regulatory regions affected precisely by DNAm remain challenging to be defined, and most studies focus on DNAm
within pre-annotated promoters and enhancers, thereby overlooking large portions of the epigenome where
aberrant regulation may occur in cancer. To address this limitation, we developed FUSE (implemented in the R
package methFuse), a computational framework that identifies candidate functional regulatory elements
genome-wide based on DNAm patterns. In whole-genome bisulfite sequencing (WGBS) data from healthy donors
obtained via the ENCODE portal, highly stable FUSE segments were enriched in known promoters, enhancers, and
repetitive elements, supporting their biological relevance.
We then applied FUSE to 191 WGBS samples from 62 HGSC patients enrolled in the DECIDER trial
(ClinicalTrials.gov identifier: NCT04846933), integrating matched RNA-seq data to identify candidate
regulatory regions through correlation with gene expression. Comparing pre- and post-treatment samples
revealed that methylation changes in pre-annotated promoter regions provided limited insight, whereas
alterations in FUSE-defined regions outside canonical promoters highlighted chemotherapy-associated effects on
cell-cycle regulation and Rho GTPase signaling processes. These results demonstrate that our DNAm-based
segmentation captures chemotherapy-associated changes not detectable through promoter-focused analyses.
A-G.53: Evaluating portability of polygenic scores across ancestries with precision-weighted meta-analysis
reveals trait-specific rather than universal poor portability
Track: Genomics, epigenomics, and genome editing
-
Natalia Nunes, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg,
Salzburg, Austria, Brazil
-
Daniel Katzlberger, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg,
Salzburg, Austria, Austria
-
Iuliia Trifonova, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg,
Salzburg, Austria, Austria
-
Arne Bathke, Department of Artificial Intelligence and Human Interfaces, Paris Lodron University Salzburg,
Austria
-
Georg Zimmermann, Department of Artificial Intelligence and Human Interfaces, Paris Lodron University
Salzburg, Austria
-
Nikolaus Fortelny, Department of Biosciences and Medical Biology, Paris-Lodron University Salzburg,
Salzburg, Austria, Austria
Presentation Overview: Show
Polygenic scores (PGS) are increasingly proposed for clinical risk stratification, yet poor portability is
often reported when models trained in Europeans are applied to non-European populations. Most comparisons,
however, compare large European cohorts against much smaller non-European cohorts, generating wide confidence
intervals in non-Europeans and masking whether observed performance gaps reflect genuine biology or sampling
noise.
We developed a series of computational approaches to systematically and rigorously evaluate the portability of
PGS across ancestries. We first compared methods to obtain confidence interval on data from the Parkinson's
Progression Markers Initiative and then developed an inverse-variance weighted meta-analysis framework, which
pools AUCs across test cohorts per PGS, weighting by precision and excluding scores with substantial
heterogeneity (I² > 80%), and then quantified portability as a ΔAUC per trait and ancestry. This approach
enabled us to assess PGS portability across ~3,900 evaluations from the PGS Catalog. We then validated these
observations in the All of Us Research Program, an NIH-funded biobank linking genotypes and electronic health
records across ~245,000 diverse participants, using bootstrap confidence intervals.
In summary, we find systematic baseline genetic differences between ancestries. Yet, the mean ΔAUC in All of
Us was only −0.02, and clear evidence of poor portability was confined to a restricted set of traits: asthma
in African and East Asian ancestries, type 2 diabetes and coronary heart disease in African ancestry. Our
results suggest that concerns about poor portability is better framed as trait-dependent rather than
universal, emphasizing the value of imbalance-aware evaluation before clinical deployment.
A-G.54: Microsynteny and functional relatedness: insights into eukaryotic genome evolution
Track: Genomics, epigenomics, and genome editing
-
Silvia Prieto Banos, University of Lausanne, Switzerland
-
Alex Warwick Vesztrocy, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
-
Natasha Glover, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
-
Christophe Dessimoz, University of Lausanne, SIB Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Genome rearrangements have shaped eukaryotic evolution, resulting in extensive variation in genome
organization and limited conservation of gene order across distant species. While prokaryotic operons and
eukaryotic gene clusters suggest that neighboring genes are often functionally related, the relationship
between gene function and gene order conservation (microsynteny) remains poorly understood in eukaryotes.
Here, we investigate the evolutionary conservation of gene microsynteny and its association with functional
similarity across a diverse set of eukaryotic genomes. Using Gene Ontology (GO) similarity measures, we show
that adjacent genes are significantly more functionally related than expected by chance. Leveraging the
unprecedented wealth of publicly available genomes and edgeHOG, a method for reconstructing gene order across
the eukaryotic tree of life, we systematically quantify the conservation of gene adjacencies in both extant
and ancestral genomes. This enables us to date the emergence of conserved adjacencies and assess how their
evolutionary conservation relates to functional coupling between genes.
We further examine the influence of intergenic distance and transcriptional orientation,and explore the role
of gene expression and regulatory elements in gene order conservation. Our work contributes to better
understanding general trends in eukaryotic genome organization, providing more insights into the subtle
interplay between genome rearrangements and gene function.