Attention Presenters - please review the Speaker Information Page available here
Schedule subject to change
All times listed are in CET
Monday 31 August
13:30-14:30
Session: From Signals to Risk: Scalable Genotyping and Genetic Inference.
Proceedings Presentation: Leveraging ONT move table values for signal aware variant calling
Confirmed Presenter: Xian Yu, School of Computing and Data Science, University of Hong Kong, Hong Kong, China, Hong Kong

Room: Room AD
Moderator(s): Shilpa Garg; Knut Reinert


Authors List: Show

  • Xian Yu, School of Computing and Data Science, University of Hong Kong, Hong Kong, China, Hong Kong
  • Zhenxian Zheng, School of Computing and Data Science, University of Hong Kong, Hong Kong, China, Hong Kong
  • Lei Chen, School of Computing and Data Science, University of Hong Kong, Hong Kong, China, Hong Kong
  • Zilan Qin, School of Computing and Data Science, University of Hong Kong, Hong Kong, China, Hong Kong
  • Minggao He, School of Computing and Data Science, University of Hong Kong, Hong Kong, China, Hong Kong
  • Ruibang Luo, School of Computing and Data Science, University of Hong Kong, Hong Kong, China, Hong Kong

Presentation Overview: Show

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table—a lightweight byproduct of basecalling that maps signal events to nucleotide positions—to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10× depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Proceedings Presentation: GUANinE v1.1 Reveals Complementarity of Supervised and Genomic Language Models
Confirmed Presenter: Eyes Robson, UC Berkeley, United States

Room: Room AD
Moderator(s): Shilpa Garg; Knut Reinert


Authors List: Show

  • Eyes Robson, UC Berkeley, United States
  • Nilah Ioannidis, UC Berkeley, United States

Presentation Overview: Show

Summary: There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models. We identify uniquely performant models, and we conclude by showing hybrid or post-trained language models may define the next era of machine learning in genomics.

Locityper enables biobank-scale genotyping of structurally complex loci
Confirmed Presenter: Timofey Prodanov, Institute for Medical Biometry and Bioinformatics, Medical Faculty, Heinrich-Heine-Universität Düsseldorf, Germany

Room: Room AD
Moderator(s): Shilpa Garg; Knut Reinert


Authors List: Show

  • Timofey Prodanov, Institute for Medical Biometry and Bioinformatics, Medical Faculty, Heinrich-Heine-Universität Düsseldorf, Germany
  • Tobias Marschall, Institute for Medical Biometry and Bioinformatics, Medical Faculty, Heinrich Heine University, Düsseldorf, Germany, Germany

Presentation Overview: Show

The human genome harbors many structural variants (SVs)–long and complex genomic differences that are difficult to characterize and genotype, with >50% of all SVs missed by short read callers. SV-rich loci generate numerous diverse haplotypes, poorly represented in a single reference genome, making read alignments ambiguous and unreliable, and hundreds of disease-associated genes inaccessible to accurate variant calling.

To tackle this problem we developed Locityper (Nature Genetics, Oct. 2025), a targeted genotyper that searches pangenome assemblies for two locus haplotypes that best explain a given whole genome sequencing dataset. Specifically, Locityper recruits and maps reads to all haplotypes and identifies the haplotype pair with optimal alignment likelihood and coverage profile. Leveraging an ever-growing collection of high-quality genome assemblies, Locityper predicts diploid locus sequence more accurately than state-of-the-art phased variant calling pipelines across 265 challenging medically relevant genes, raising median QV (Phred-scaled sequence divergence) from 24.4 to 35.3 (short reads) and 35.2 to 36.9 (long reads).

Subsequently, Locityper was used for genotyping tandem repeat-rich mucin genes; for evaluation of the new HPRC2 and HGSVC3 pangenomes; and for validation of putative disease-relevant SVs in the All-of-Us cohort. Most recently, we genotyped 50 SVs across 424k UK Biobank short read datasets and identified disease associations for 4 of them (FDR-corrected p-value <0.05). To facilitate this analysis we extended and optimized the Locityper pipeline, reducing compute cost to just 0.007¢/sample/SV. This low genotyping cost, together with high accuracy, underscores Locityper's value for large-scale diagnostic and discovery analyses across SV-rich loci.

14:45-15:45
Session: Sequence Data structures and search
Proceedings Presentation: RSHash: A fast and space-efficient hash table for k-mers
Confirmed Presenter: Jonas Schulte-Mattler, Algorithmic Bioinformatics, Germany

Room: Room BC
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Jonas Schulte-Mattler, Algorithmic Bioinformatics, Germany
  • Knut Reinert, Algorithmic Bioinformatics, Germany

Presentation Overview: Show

Large genomic data collections can be viewed as a continuous string of DNA characters. The essential operations for data structures indexing the $k$-mer content of such a string are {\tt lookup} and {\tt locate}. {\tt Lookup} determines whether a query $k$-mer $q$ exists in the string and {\tt locate} returns all locations in the string where $q$ is present. High-throughput DNA sequencing generates very many k-mer sets of size exceeding billions of characters. In such scenarios, memory consumption and query efficiency pose significant challenges to a data structure supporting the above mentioned queries.

To address this problem, we describe a simple, compressed, static data structure for k-mers that answers {\tt lookup} and can be extended for supporting {\tt locate}. The general scheme follows the use of minimizers like the state-of-the art SSHash. However, instead of using minimum perfect hash functions our solution (RSHash for Rank-Select Hash) relies on bitvectors with rank and select support, a multiple layered minimizer scheme, and a clever buffering strategy. We can show that RSHash is on average $40\%$ and in some cases up to two times faster than SSHash while having the same memory requirements. Indeed we can go as low as $8$~bits per canonical $31$-mer on a human dataset.

Ultrafast and accurate sequence alignment and clustering of viral genomes
Confirmed Presenter: Adam Gudys, Silesian University of Technology, Poland

Room: Room BC
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Andrzej Zielezinski, Adam Mickiewicz University, Poland
  • Adam Gudys, Silesian University of Technology, Poland
  • Jakub Barylski, Adam Mickiewicz University, Poland
  • Krzysztof Siminski, Silesian University of Technology, Poland
  • Piotr Rozwalak, Adam Mickiewicz University, Poland
  • Bas E. Dutilh, Utrecht University, Netherlands
  • Sebastian Deorowicz, Silesian University of Technology, Poland

Presentation Overview: Show

An essential challenge in clustering massive datasets of viral genomes is efficient calculation of average nucleotide identity (ANI). Since gold-standard, alignment-based approaches for calculating ANI (e.g., BLAST) are computationally unsuitable for large-scale analyses, more efficient k-mer-based methods employing sketching (FastANI, RabbitTClust) or approximate alignment (skani) have been developed. These tools, however, have lower accuracy than alignment-based methods and/or do not provide clustering capabilities.

We present Vclust, a package for clustering millions of viral genomes. The method consists of three steps. The first one is a k-mer-based filtering of sequence pairs performed with Kmer-db 2, an updated version of Kmer-db, equipped with algorithms for producing sparse distance matrices. This is followed by an accurate estimation of ANI for selected sequence pairs using LZ-ANI, an alignment method grounded in Lempel-Ziv parsing. The final step is clustering genomes with Clusty, a package providing efficient implementations of multiple clustering algorithms suited for sparse distance matrices.

When assigning 90,000 contigs from IMG/VR database to virus OTUs, Vclust exhibited the highest agreement with BLAST (99%) compared to MegaBLAST (97%), skani (96%) and FastANI (93%), being noticeably faster than the competitors. MMSeqs2, the only method comparable in terms of execution times, was significantly less accurate (70% BLAST agreement). The ability to process the entire IMG/VR database (over 15 million contigs) in 2.4 hours and 92GB of RAM, confirms Vclust to be ready for future challenges in the viral genomics. With over 18 thousand Bioconda/GitHub downloads, Vclust has already proven its usefulness in the community.

k-mer manifold approximation and projection for visualizing DNA sequences
Confirmed Presenter: Lu Cheng, University of Eastern Finland, Finland

Room: Room BC
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Chengbo Fu, Aalto University, Finland
  • Einari Niskanen, University of Eastern Finland, Finland
  • Gong Hong Wei, University of Oulu, Finland
  • Zhirong Yang, Norwegian University of Science and Technology (NTNU), Norway
  • Marta Sanvicente-García, Universitat Pompeu Fabra, Spain
  • Marc Güell, Universitat Pompeu Fabra, Spain
  • Lu Cheng, University of Eastern Finland, Finland

Presentation Overview: Show

Identifying and illustrating patterns in DNA sequences are crucial tasks in various biological data analyses. In this task, patterns are often represented by sets of k-mers, the fundamental building blocks of DNA sequences. To visually unveil these patterns, one could project each k-mer onto a point in two-dimensional (2D) space. However, this projection poses challenges owing to the high-dimensional nature of k-mers and their unique mathematical properties. Here, we establish a mathematical system to address the peculiarities of the k-mer manifold. Leveraging this k-mer manifold theory, we develop a statistical method named KMAP for detecting k-mer patterns and visualizing them in 2D space. We applied KMAP to three distinct data sets to showcase its utility. KMAP achieves a comparable performance to the classical method MEME, with ∼90% similarity in motif discovery from HT-SELEX data. In the analysis of H3K27ac ChIP-seq data from Ewing sarcoma (EWS), we find that BACH1, OTX2, and KNCH2 might affect EWS prognosis by binding to promoter and enhancer regions across the genome. We also observe potential colocalization of BACH1, OTX2, and the motif CCCAGGCTGGAGTGC in ∼70 bp windows in the enhancer regions. Furthermore, we find that FLI1 binds to the enhancer regions after ETV6 degradation, indicating competitive binding between ETV6 and FLI1. Moreover, KMAP identifies four prevalent patterns in gene editing data of the AAVS1 locus, aligning with findings reported in the literature. These applications underscore that KMAP can be a valuable tool across various biological contexts.

Tuesday 1 September
11:45-12:45
Session: RNA regulation and chromatin context
Improved RNA-DNA interaction calling suggests RNA-based gene regulation of phenotypic transitions
Confirmed Presenter: Timothy Warwick, Goethe University Frankfurt, Germany

Room: Room AD
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Simonida Zehr, Goethe University Frankfurt, Germany
  • Ralf Brandes, Goethe University Frankfurt, Germany
  • Marcel Schulz, Goethe University Frankfurt, Germany
  • Timothy Warwick, Goethe University Frankfurt, Germany

Presentation Overview: Show

Chromatin-localized RNAs play diverse roles in gene regulation and nuclear architecture. Mapping genome-wide RNA–DNA interactions is possible using a variety of molecular methods, including using bridging oligonucleotides to ligate RNA and DNA in proximity. While molecular methods have progressed, a robust computational method for calling biologically meaningful RNA–DNA interactions from these data is lacking. Herein, we present RADIAnT, a reads-to-interactions pipeline for analyzing RNA–DNA ligation data. RADIAnT calls interactions against a dataset-specific, unified background, which considers RNA binding site–TSS distance and genomic region bias, and outperforms previously proposed methods in the accurate recall of genome-wide RNA–DNA interactions. Accurate RNA–DNA interaction calling enables the analysis of gene regulatory RNAs in dynamic biological contexts. Here, dynamically chromatin-associated RNAs were identified in the physiologically- and pathologically relevant process of endothelial-to-mesenchymal transition. By depleting candidate chromatin-associated lncRNAs, their gene regulatory behavior at bound target genes important to endothelial phenotype maintenance could be validated. These data demonstrate how effective RNA–DNA interaction calling can help to place RNAs at key points in gene regulatory networks governing cellular behavior.

From Transcripts to Cells: Dissecting Sensitivity, Signal Contamination, and Specificity in Xenium Spatial Transcriptomics
Confirmed Presenter: Mariia Bilous, Biomedical Data Science Center, CHUV, Switzerland

Room: Room AD
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Mariia Bilous, Biomedical Data Science Center, CHUV, Switzerland
  • Daria Buszta, Biomedical Data Science Center, CHUV; Department of Oncology, CHUV; Swiss Cancer Center Leman, Switzerland
  • Jonathan Bac, Biomedical Data Science Center, CHUV, Switzerland
  • Senbai Kang, Biomedical Data Science Center, CHUV, Switzerland
  • Yixing Dong, Biomedical Data Science Center, CHUV, Switzerland
  • Stephanie Tissot, Department of Oncology, CHUV; Swiss Cancer Center Leman; Ludwig Institute for Cancer Research, Switzerland
  • Sylvie Andre, Department of Oncology, CHUV; Swiss Cancer Center Leman; Ludwig Institute for Cancer Research, Switzerland
  • Marina Alexandre-Gaveta, Department of Oncology, CHUV; Swiss Cancer Center Leman; Ludwig Institute for Cancer Research, Switzerland
  • Christel Voize, Department of Oncology, CHUV; Swiss Cancer Center Leman; Ludwig Institute for Cancer Research, Switzerland
  • Solange Peters, Department of Oncology, Lausanne University Hospital; Swiss Cancer Center Leman, Lausanne, Switzerland
  • Krisztian Homicsko, Department of Oncology, CHUV; Swiss Cancer Center Leman; Ludwig Institute for Cancer Research, Switzerland
  • Raphael Gottardo, Biomedical Data Science Center, CHUV; Swiss Institute of Bioinformatics; FBM, UNIL; School of Life Sciences, EPFL, Switzerland

Presentation Overview: Show

Spatial transcriptomics enables high-resolution gene expression mapping in intact tissues, yet the technical properties and limitations of widely adopted platforms remain poorly characterized. We present one of the most comprehensive Xenium datasets to date, encompassing over 40 breast and lung tumor sections profiled across diverse gene panels. Leveraging this resource, we systematically dissect technical noise - including transcript spillover - along with assay specificity, panel performance, and segmentation strategies, providing a critical benchmark for Xenium performance in complex tissues.

To our knowledge, we are the first to systematically characterize and quantify transcript spillover from neighboring cells at scale, identifying it as a dominant source of noise that generates mixed signals, reduces specificity, and obscures biologically meaningful programs, particularly within immune populations in cancer tissues. We demonstrate that single-nucleus RNA-seq enables precise quantification of this phenomenon, establishing a rigorous framework for its assessment.

Building on these insights, we introduce SPLIT (Spatial Purification of Layered Intracellular Transcripts), a novel computational method that models and corrects transcript spillover using cell type deconvolution outputs within an interpretable linear modeling framework. SPLIT is implemented as an open-source R package, scales to large spatial datasets, and is applicable in principle to other spatial omics modalities where signal mixing occurs.

SPLIT substantially improves background correction and cell-type resolution, enabling detection of T-cell exhaustion signatures associated with malignant cell colocalization - signals otherwise obscured in uncorrected data - demonstrating its value for spatially informed biological inference.

CREsted: modeling genomic and synthetic cell-type-specific enhancers across tissues and species
Confirmed Presenter: Niklas Kempynck, VIB-KU Leuven, Belgium

Room: Room AD
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Niklas Kempynck, VIB-KU Leuven, Belgium
  • Seppe De Winter, VIB-KU Leuven, Belgium
  • Cas Blaauw, VIB-KU Leuven, Belgium
  • Vasilieios Konstantakos, VIB-KU Leuven, Belgium
  • Eren Can EkÅŸi, VIB-KU Leuven, Belgium
  • Sam Dieltiens, VIB-KU Leuven, Belgium
  • Darina Abaffyová, VIB-KU Leuven, Belgium
  • Valérie Bercier, VIB-KU Leuven, Belgium
  • Ibrahim Taskiran, Illumina, United States
  • Gert Hulselmans, VIB-KU Leuven, Belgium
  • Valerie Christiaens, VIB-KU Leuven, Belgium
  • Ludo Van Den Bosch, VIB-KU Leuven, Belgium
  • Lukas Mahieu, VIB-KU Leuven, Belgium
  • Stein Aerts, VIB-KU Leuven, Belgium

Presentation Overview: Show

Sequence-based deep learning models have become the state of the art for analyzing the genomic regulatory code. Particularly for enhancers, these models excel at deciphering sequence grammar that underlies their activity. To enable end-to-end enhancer modeling and design, we developed a software package called CREsted (cis-regulatory element sequence training, explanation and design). It combines preprocessing and analysis of single-cell assay for transposase-accessible chromatin using sequencing data, modeling chromatin accessibility from sequence, sequence design and downstream analysis to decipher enhancer grammar. We demonstrate CREsted's functionality on a mouse cortex and a human peripheral blood mononuclear cell dataset. Additionally, we use CREsted to compare mesenchymal-like cancer cell states between tumor types, and we investigate different fine-tuning strategies of genomic foundation models within CREsted. Finally, we train a model on a zebrafish development atlas and use this to design and in vivo validate cell-type-specific enhancers. For varying datasets, we demonstrate that CREsted facilitates efficient training and analyses, enabling scrutinization of the enhancer logic and design of synthetic enhancers across tissues and species.

15:15-16:15
Session: Cancer Biomarkers and Immune targets
Proceedings Presentation: Liquid biopsies reveal dual compartments of cancer risk from tumor and host-derived mutations
Confirmed Presenter: Federica Malighetti, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy, Italy

Room: Room AD
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Federica Malighetti, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy, Italy
  • Ivan Civettini, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy, Italy
  • Andrea Aroldi, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy,, Italy
  • Alberto Maria Villa, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy, Italy
  • Matteo Villa, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy, Italy
  • Luca Sala, Fondazione IRCCS San Gerardo dei Tintori di Monza, Monza, Italy, Italy
  • Nicoletta Cordani, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy, Italy
  • Massimiliano Cadamuro, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy;, Italy
  • Diego Cortinovis, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy;, Italy
  • Rocco Piazza, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy;, Italy
  • Luca Mologni, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy;, Italy
  • Daniele Ramazzotti, Department of Medicine and Surgery, University of Milano-Bicocca, Milano, Italy,, Italy

Presentation Overview: Show

Circulating tumor DNA (ctDNA) and clonal hematopoiesis of indeterminate potential (CHIP) are two biologically distinct sources of somatic mutations detectable in blood. While ctDNA captures tumor-intrinsic alterations, CHIP arises from age-related hematopoietic clones and is often considered back-ground noise. Here, we conduct a large-scale, tumor-type-resolved analysis of over 9,000 patients with CHIP data and 1,500 patients with ctDNA data across solid tumors profiled at Memorial Sloan Kettering Cancer Center.
Our results reveal that CHIP and ctDNA mutations exhibit non-overlapping, clinically meaningful signals. CHIP mutations, particularly in DNA damage response and epigenetic regulators (e.g., PPM1D, CHEK2, ATM, TP53, ASXL1), are associated with worse overall survival, increased metastatic potential, and site-specific dissemination. ctDNA mutations in canonical oncogenic drivers (e.g., TP53, EGFR, KRAS, STK11) reflect tumor aggressiveness and correlate with poor prognosis and metastasis across multiple cancer types. Joint modeling in lung adenocarcinoma confirms the independent prognostic contributions of both compartments. Additionally, longitudinal clonal analysis links specific CHIP mutations to the emergence of hematologic malignancies under therapeutic pressure. These findings support a dual-compartment model of liquid biopsy, in which tumor- and host-derived mutations jointly inform on cancer risk, progression, and metastatic behavior. Integrating both compartments may enhance the clinical utility of blood-based biomarkers in oncology.

Proceedings Presentation: Micropeptides encoded by lncRNAs associated with cancer progression reveal novel immunogenic epitopes
Confirmed Presenter: Stav Zok, The Hebrew University of Jerusalem, Israel

Room: Room AD
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Stav Zok, The Hebrew University of Jerusalem, Israel
  • Michal Linial, The Hebrew University of Jerusalem, Israel

Presentation Overview: Show

Motivation: Long non-coding RNAs (lncRNAs) regulate gene expression, chromatin organization, and
cellular signaling. Recent studies indicate that ~20% of the ~36,000 human lncRNA genes harbor small
open reading frames (sORFs) capable of producing micropeptides (MPs), whose functions remain largely
unknown. Whether these peptides contribute to the cancer immunopeptidome is largely unexplored.
Results: We systematically analyzed lncRNAs with strong experimental and computational evidence of
MP-encoding potential (~13% of the initial MP collection). Using The Cancer Genome Atlas (TCGA), we
identified 2,606 high-confidence lncRNA-derived MPs encoded by 647 genes across 16 cancer types.
We then focused on 501 MPs from 124 lncRNA genes whose expression changes significantly across
tumor stages and metastatic transitions, representing cancer transitional lncRNAs (Tr-lncRNAs). Dipeptide composition and conservation analyses showed that these MPs differ from a size-matched human coding proteome, supporting their potential as neoantigens. All possible 9-mer peptides were evaluated
for predicted binding to prevalent European HLA class I alleles. Approximately 60% of Tr-lncRNA genes
and 184 (37%) of derived peptides exhibited strong predicted HLA binding. Peptides from XIST, PCAT7,
PVT1, HAND2-AS1 showed broad HLA coverage. Notably, TTN-AS1, encoded an MP (79 aa) generated
33 predicted distinct epitopes spanning all 27 HLA alleles. Our analysis identifies lncRNA-derived MPs
as a previously underexplored source of potential cancer neoantigens, highlighting their promise as biomarkers and targets for immunotherapy.

Proceedings Presentation: Modeling time-varying genetic effects on binary disease risk via functional Mendelian Randomization
Confirmed Presenter: Nicole Fontana, Politecnico di Milano, Department of Mathematics / Human Technopole, Health Data Science Center, Italy

Room: Room AD
Moderator(s): Maria Anisimova; Knut Reinert


Authors List: Show

  • Nicole Fontana, Politecnico di Milano, Department of Mathematics / Human Technopole, Health Data Science Center, Italy
  • Piercesare Secchi, Politecnico di Milano, Department of Mathematics, Italy
  • Emanuele Di Angelantonio, Human Technopole, Health Data Science Centre, Italy
  • Francesca Ieva, Politecnico di Milano, Department of Mathematics / Human Technopole, Health Data Science Center, Italy

Presentation Overview: Show

Motivation: Genome-wide association studies have identified thousands of genetic variants associated with complex traits, establishing Mendelian Randomization (MR) as a powerful framework for causal inference using variants as natural experiments. However, existing MR methods treat causal effects as static, rely on cross-sectional exposure measurements, and ignore how genetic predispositions to disease operate dynamically across the life course. Recovering age-specific causal effect functions from longitudinal data requires combining functional data representations of exposure trajectories with instrumental variable estimation strategies suitable for binary disease endpoints, a methodological gap that has remained unaddressed.
Results: We develop a functional MR framework for binary outcomes that integrates Functional Principal Component Analysis with Two-Stage Residual Inclusion (2SRI), ensuring consistent estimation under the nonlinear logistic link function that renders standard instrumental variable estimators inconsistent. Simulations across different causal effect trajectory shapes, varying measurement densities, and varying instrument strengths demonstrate accurate recovery of time-varying genetically predicted effects with minimal bias. Applied to UK Biobank data, the framework identifies an age-specific causal effect of genetically predicted body mass index on type 2 diabetes risk concentrated in early mid-adulthood and progressively attenuating thereafter. Concordance between the proposed 2SRI estimator applied to type 2 diabetes and the established continuous-outcome functional MR estimator applied to the paired glycated haemoglobin marker in the same cohort provides indirect empirical support for the validity of the proposed approach.
Availability and implementation: The method is implemented in the R package mvfmr, with a full tutorial vignette.

Wednesday 2 September
10:30-11:30
Session: Interpretable genomics and model trust
Proceedings Presentation: BaGGLS: A Bayesian Shrinkage Framework for Interpretable Modeling of Interactions in High-Dimensional Biological Data
Confirmed Presenter: Marta Lemanczyk, Hasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, Germany

Room: Room BC
Moderator(s): Shilpa Garg; Knut Reinert


Authors List: Show

  • Marta Lemanczyk, Hasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, Germany
  • Lucas Kock, Department of Statistics and Data Science, National University of Singapore, Singapore
  • Johanna Schlimme, Hasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, Germany
  • Nadja Klein, Scientific Computing Center, Karlsruhe Institute of Technology, Germany
  • Bernhard Renard, Hasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, Germany

Presentation Overview: Show

Biological data is often high-dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features are typically active. While statistical approaches often result in more interpretable models, deep learning models are effective at modeling complex interactions and accurate predictions, yet their black-box nature limits interpretability.
We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure of interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods in interaction detection and is much faster than MCMC sampling under a horseshoe prior. We also demonstrate the usefulness of BaGGLS for interaction discovery from motif scanner outputs (e.g., FIMO) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across various high-dimensional tasks in computational biology.

Proceedings Presentation: PRISM-G: an interpretable privacy scoring framework for assessing risk in synthetic human genome data
Confirmed Presenter: Alejandro Correa Rojo, KU Leuven, Belgium

Room: Room BC
Moderator(s): Shilpa Garg; Knut Reinert


Authors List: Show

  • Alejandro Correa Rojo, KU Leuven, Belgium
  • Yves Moreau, KU Leuven, Belgium
  • Gökhan Ertaylan, Flemish Institute for Technological Research (VITO), Belgium

Presentation Overview: Show

Synthetic genomic data promises broader data access, but unresolved privacy risks remain a major concern. Existing evaluations often rely on similarity-based metrics that measure proximity between real and synthetic genomes, overlooking additional mechanisms through which genomic information may leak. We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genomic data across three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure through rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0–100 PRISM-G score. By pairing PRISM-G with downstream utility metrics, the framework also enables analysis of privacy–utility trade-offs across generative models. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT solver (Genomator). Our results show that privacy vulnerabilities arise along different axes across models and marker densities, demonstrating that a single similarity-based metric is insufficient to characterize genomic privacy risk. The source code of PRISM-G is available at https://github.com/alejocrojo09/prismg
.

Proceedings Presentation: Knowledge-Guided Learning with Curated Prior Genetic Biomarkers for Robust Model Interpretation
Confirmed Presenter: Beomsu Baek, University of Nevada, Las Vegas, United States

Room: Room BC
Moderator(s): Shilpa Garg; Knut Reinert


Authors List: Show

  • Beomsu Baek, University of Nevada, Las Vegas, United States
  • Eunyoung Jang, University of Nevada, Las Vegas, United States
  • Sai Phani Parsa, University of Nevada, Las Vegas, United States
  • Youngsoon Kim, Gyeongsang National University, South Korea
  • Mingon Kang, University of Nevada, Las Vegas, United States

Presentation Overview: Show

Motivation: Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms.

Results: In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures.

Availability and implementation: The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

13:15-14:15
Session: Evolutionary graphs and Biosynthetic Inference
Proceedings Presentation: Nerpa 2: probabilistic linking of biosynthetic gene clusters to nonribosomal peptides
Confirmed Presenter: Ilia Olkhovskii, Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Germany

Room: Room BC
Moderator(s): Maria Anisimova; Shilpa Garg


Authors List: Show

  • Ilia Olkhovskii, Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Germany
  • Aleksandra Kushnareva, Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Germany
  • Azat Tagirdzhanov, Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Germany
  • Alexey Gurevich, Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Germany

Presentation Overview: Show

Motivation: Nonribosomal peptides (NRPs) are bioactive microbial metabolites with high pharmaceutical potential. Although genome mining enables large-scale detection of biosynthetic gene clusters (BGCs) predicted to encode NRPs, reliably linking these clusters to their chemical products remains challenging due to the flexible and heterogeneous organization of NRP assembly pathways.

Results: We present Nerpa 2, a probabilistic framework for accurate and scalable linking of NRP BGCs to candidate chemical structures. The method represents assembly lines as hidden Markov models (HMMs) that capture uncertainty and alternative biosynthetic routes. On curated datasets of experimentally validated BGC–product pairs, our tool outperforms existing methods in linking accuracy and pathway reconstruction. When applied to large genome mining datasets, Nerpa 2 efficiently identifies BGCs likely associated with known compounds and highlights potential producers of novel chemistry.

Availability and implementation: Nerpa 2 is freely available at https://github.com/gurevichlab/nerpa.

Proceedings Presentation: Episode Clustering in Phylogenetic Networks
Confirmed Presenter: Pawel Gorecki, University of Warsaw, Poland

Room: Room BC
Moderator(s): Maria Anisimova; Shilpa Garg


Authors List: Show

  • Pawel Gorecki, University of Warsaw, Poland
  • Agnieszka Mykowiecka, Faculty of Mathematics, Informatics and Mechanics, University of Warsaw, Poland
  • JarosÅ‚aw Paszek, Faculty of Mathematics, Informatics and Mechanics University of Warsaw, Poland

Presentation Overview: Show

The classical duplication episode clustering (EC) model introduced by Guigo et al. in the 1990s provides a foundational approach for inferring genomic duplication events crucial to understanding genome evolution. This model clusters single gene duplications from a collection of gene trees at locations in the species tree to minimize the total number of such locations, called duplication episodes.
Here, we introduce NetEC, a novel extension of this problem to phylogenetic networks. To solve NetEC, we first develop a polynomial-time dynamic programming (DP) algorithm for testing whether a given set of network nodes can serve as episode locations. We then propose a main inference algorithm that utilizes this DP component to optimize the episode count; while the feasibility test runs in polynomial time, the full optimization has exponential worst-case complexity, and an optional heuristic mode is provided for larger instances. We also propose an extended episode analysis procedure that identifies additional genomic duplication candidates below reticulation nodes, complementing the main algorithm by resolving potential upward clustering of duplications induced by reticulation. We evaluate our method on simulated data and on an empirical Pandanales dataset comprising over 29,000 gene trees, demonstrating exact and accurate inference of genomic duplication events even in the presence of multiple reticulations.

Proceedings Presentation: ARGformer: learning on ancestral recombination graphs with transformers
Confirmed Presenter: David Bonet, University of California, Santa Cruz, United States

Room: Room BC
Moderator(s): Maria Anisimova; Shilpa Garg


Authors List: Show

  • David Bonet, University of California, Santa Cruz, United States
  • Cole Shanks, University of California, Santa Cruz, United States
  • Marçal Comajoan Cara, University of California, Berkeley, United States
  • Jordi Abante, Stanford University, United States
  • Alexander Ioannidis, Stanford University, United States

Presentation Overview: Show

Recent advances in inference of the ancestral recombination graph (ARG), which describes how segments of chromosomes trace back through recombination and shared lineages, have made it possible to reconstruct genome-wide genealogies for large cohorts, but it remains difficult to summarize and use this information for population genetic analyses. We present ARGformer, an encoder-only transformer that learns context-dependent embeddings with a self-supervised masked objective finetuned with contrastive learning for downstream retrieval tasks. We train ARGformer on genealogies from coalescent simulations and on genealogies inferred from ancient and present-day Homo sapiens genomes. Using only these learned embeddings, without access to genotype matrices, ARGformer captures patterns of global population structure and supports ancestry inference through clustering and nearest-neighbor retrieval. On genealogies that include archaic hominins, ARGformer can highlight Denisovan-derived segments in Oceanian genomes and reveals Oceanian-like ancestry in South American Indigenous populations.