View posters by category

Scroll down to view Results

Session A: Monday 31 August 12:00-13:30
Session B: Tuesday 1 September 16:15-17:45
Session C: Wednesday 2 September 11:30-13:00

Results

B-G.01: Circulating DNA reveals nucleosome occupancy patterns that are associated with nucleosome-DNA affinity and are affected in cancer
Track: Genomics, epigenomics, and genome editing
  • Marianne Richaud, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University, Montpellier, France
  • Ekaterina Pisareva, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University, Montpellier, France
  • Alain Thierry, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University, Montpellier, France
  • Jacques Colinge, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University, Montpellier, France


Presentation Overview: Show

The study of cell-free circulating DNA (cirDNA) fragments (fragmentomics) from liquid biopsies has received increasing attention. By mapping a large ensemble of well-positioned nucleosomes (WPNs), we found that nucleosome occupancy was associated with histone-DNA affinity, as evidenced by codon usage bias and differences in cirDNA fragment sizes. Moreover, nucleosome occupancy was different in healthy and cancer samples, thus allowing developing a high-performance machine learning approach for cancer detection (specificity and sensitivity >0.95 for seven cancer types in Cristiano et al., 2016, data). Cancer influenced nucleosome occupancy in a global manner, although distinct cancer types retained specific features. WPN occupancy at transcription factor binding sites revealed shared, pan-cancer regulation of transcriptional programs involved in hematopoietic cell differentiation and neutrophil biology, the main cirDNA sources. This work provides new fundamental insights into cirDNA and DNA sequence using cirDNA as a physical readout. It also bares translational significance by disclosing a new high-performance strategy for cancer detection from liquid biopsies.

B-G.02: Genome-wide association analysis of primary response to anti-TNF therapy in inflammatory bowel disease
Track: Genomics, epigenomics, and genome editing
  • Maria Gretsova, Institute of Clinical Molecular Biology, Kiel University, Germany
  • Johan Burisch, Gastro Unit, Medical Division, University Hospital Copenhagen, Amager and Hvidovre Hospital, Denmark
  • Vibeke Andersen, Institute of Molecular Medicine, University of Southern Denmark, Denmark
  • Walter Reinisch, Department of Internal Medicine, Medical University of Vienna, Austria
  • David Ellinghaus, Institute of Clinical Molecular Biology, Kiel University, Germany


Presentation Overview: Show

Inflammatory bowel disease (IBD), comprising Crohn's disease (CD) and ulcerative colitis (UC), is a chronic relapsing inflammatory disorder with rising global prevalence. Anti-TNF therapy is widely used as a first-line biologic treatment, yet a substantial proportion of patients fail to achieve early clinical remission. Predictive biomarkers for treatment response are therefore of considerable clinical interest, but robust genetic markers have remained largely undefined.
To investigate genetic determinants of anti-TNF response, we analyzed nine discovery cohorts from Denmark, Austria, and Germany including IBD patients treated with anti-TNF as their first biologic therapy. Two independent cohorts from Germany and the United States were used for replication. For each IBD subtype, we performed genome-wide single-variant association analyses, followed by fine-mapping and colocalization, and complemented these analyses with gene-based association testing. We additionally applied cross-phenotype meta-analysis to explore shared genetic signals across CD and UC. To place loci emerging from the genetic analyses into a disease-relevant biological context, we integrated publicly available single-cell RNA-seq data from the scIBD database and assessed expression patterns across healthy and inflamed intestinal tissue states.
This study establishes an analytical framework for the discovery and biological contextualization of genetic factors associated with anti-TNF treatment response in IBD.

B-G.03: Exploring plasmids within human gut microbiomes via Micro-C metagenomics
Track: Genomics, epigenomics, and genome editing
  • Marcos Bermejo-Ruiz, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany, Germany
  • Carolin Wilhelm, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany, Germany
  • Heike Budde, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany, Germany
  • Ruth Ley, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany, Germany
  • Alexander Tyakht, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany, Germany


Presentation Overview: Show

Bacterial plasmids are known to be present in the human gut and to confer diverse traits to their hosts, such as antibiotic resistance or metabolite absorption. However, the diversity, specificity, and dynamics of plasmid-bacteria associations remain to be fully characterized. Due to the limited reference data and the narrow scope of existing databases and signature genes, conventional metagenomics approaches struggle to assign bacterial hosts to plasmids. To overcome these caveats, Hi-C metagenomics has been employed to obtain spatial information about DNA. However, this method has shown limited performance in human microbiomes, and its reliance on restriction enzymes further constrains its effectiveness. In our study, we introduce Micro-C metagenomics. Micro-C, similar to Hi-C but independent of restriction enzymes, provided the spatial information needed to link sequences of plasmids and bacterial chromosomes, obtained using short- and long-read metagenomics. Tested on a synthetic community and validated on a real-life stool sample, this approach enabled the construction of high-resolution contact maps of the sequences. Additionally, the method successfully recovered the expected plasmid-bacteria links in the synthetic community, and subsequently identified links also in the stool sample. Thus, Micro-C metagenomics can effectively retrieve connections between plasmid sequences and their bacterial hosts. This enables the construction of interaction networks and allows the tracking of plasmid transmission, which facilitates further study of plasmid roles in the human gut microecology.

B-G.04: Improved tumor-only variant calling and mutation burden estimation with VarNet-T
Track: Genomics, epigenomics, and genome editing
  • Kiran Krishnamachari, Genome Institute of Singapore, A*STAR, Singapore
  • Anders Skanderup, Genome Institute of Singapore, A*STAR, Singapore


Presentation Overview: Show

Somatic variant calling algorithms typically detect mutations in cancer genomes by comparing sequence data from a tumor sample against a matched normal sample. However, matched normal samples are often unavailable in clinical diagnostics or retrospective analyses of archival tumor samples in biobanks, compromising variant calling accuracy due to the difficulty in distinguishing somatic mutations from germline mutations or sequencing artifacts. Here, we introduce VarNet-T, an end-to-end weakly supervised deep learning framework for accurately identifying somatic variants from aligned tumor reads without a matched normal sample. VarNet-T is trained using millions of high-confidence variants and benchmarked using public datasets, demonstrating 20-33% performance improvement over existing methods. We assess the accuracy of tumor mutation burden (TMB) estimation on 1000 tumor samples spanning 10 solid cancer types. Compared to existing methods, VarNet-T demonstrates >3x higher accuracy in TMB-high status classification, suggesting significant potential to improve patient selection for immunotherapy. Overall, the improved accuracy of VarNet-T has the potential to enhance the utility of tumor-only sequencing in cancer research and clinical molecular diagnostics.

B-G.05: Gene2Phenotype (G2P): a database of detailed, structured gene-disease associations
Track: Genomics, epigenomics, and genome editing
  • Diana Lemos, EMBL-EBI, United Kingdom
  • Seeta Ramaraju Pericherla, EMBL-EBI, United Kingdom
  • Sarah E Hunt, EMBL-EBI, United Kingdom
  • Elena Cibrian Uhalte, EMBL-EBI, United Kingdom
  • Michael Yates, University of Edinburgh, United Kingdom
  • Ian Simpson, University of Edinburgh, United Kingdom
  • Helen V Firth, Addenbrooke's Hospital Cambridge University Hospitals, United Kingdom
  • Mallory Freeberg, EMBL-EBI, United Kingdom


Presentation Overview: Show

Up-to-date gene-disease association information and tools are needed to identify plausibly disease-associated variants from the large numbers generated in diagnostic genome sequencing. The Gene2Phenotype (G2P) system was designed to meet this need. It combines expert-curated gene-disease association information from the literature with a genotype filtering method built around the Ensembl Variant Effect Predictor (VEP) to enable robust identification of genotypes for prioritisation.
We present a redesigned G2P web interface and a new REST API to enable both interactive exploration and programmatic access. To support more accurate diagnosis, the research and development of novel therapies, the updated platform disseminates more detailed information on disease mechanisms. The website provides an enhanced user experience with an intuitive interface and search tool and includes links to publications supporting each association. Additionally, each G2P record has a stable unique identifier enabling tracking and direct links to external pages. The API allows flexible programmatic querying of G2P records by commonly used identifiers or full data download, facilitating the integration of G2P data into other systems.
We are now utilising machine learning to support manual curation while improving data quality. We have developed a pipeline which identifies relevant publications and maps them to existing G2P gene-disease records. To improve robustness, we perform cross-model validation using an AI model to assess the relevance of candidate publications with the corresponding G2P records. These methods will accelerate evidence acquisition and assessment, making new knowledge accessible for research and diagnostic purposes more rapidly.

B-G.06: Longitudinal cfDNA methylation profiling by EM-seq and TAPS+ for early breast cancer biomarker discovery
Track: Genomics, epigenomics, and genome editing
  • Daniyar Karabayev, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki, Finland, Finland
  • Rodos Rodosthenous, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki, Finland, Finland
  • Maare Arffman, Applied Tumor Genomics, Research Programs Unit, Faculty of Medicine, University of Helsinki, Helsinki, Finland, Finland
  • Ican, iCAN Digital Precision Cancer Medicine Flagship, Helsinki, Finland, Finland
  • Sirpa Leppä, Applied Tumor Genomics, Research Programs Unit, Faculty of Medicine, University of Helsinki, Helsinki, Finland, Finland
  • Andrea Ganna, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki, Finland, Finland
  • Esa Pitkänen, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki, Finland, Finland


Presentation Overview: Show

Cell-free DNA (cfDNA) methylation profiling has emerged as a promising non-invasive approach for early cancer detection, yet breast cancer remains particularly difficult to identify via liquid biopsy. Targeted methylation-based approaches have shown strong performance across cancer types, but breast cancer-specific signal remains elusive. Integrating whole-genome cfDNA methylation with matched tumor multi-omics data offers an opportunity to identify more sensitive and specific epigenetic biomarkers for early detection.
Plasma cfDNA was obtained from breast cancer patients through the Helsinki Biobank, matched to tumor samples available in the iCAN Digital Precision Cancer Medicine Flagship. Two complementary whole-genome (WG) cfDNA methylation sequencing approaches were applied. WG EM-seq was performed on 52 plasma samples from 26 individuals, each with a pre-diagnostic (mean 21 months prior; range 6–47 months) and a post-diagnostic timepoint (<2 months after diagnosis). WG TAPS+ is being applied to an expanded cohort of 181 samples from unique individuals with single timepoints post-diagnosis. Sequencing reads were processed using DRAGEN and nf-core/methylseq and integrated with matched tumor multi-omics data on the iCAN Discovery Platform.
cfDNA concentrations in the EM-seq cohort ranged from 36 to 1,170 ng/mL (mean 198 ng/mL) at a median coverage depth of 6.5× and a 95% alignment rate. Initial exploratory analysis of genome-wide methylation profiles revealed preliminary clustering patterns distinguishing breast cancer samples from external healthy controls at both pre- and post-diagnostic timepoints, suggesting a detectable epigenetic signal before diagnosis. Copy number profiles called using ichorCNA were largely copy-number neutral across the cohort, with a single case exhibiting detectable alterations.

B-G.07: Methylation profile analysis of 'Candidatus Phytoplasma mali' infected apple leaves
Track: Genomics, epigenomics, and genome editing
  • Sara Bortolini, Laimburg Research Centre, Bozen-Bolzano, Italy and University of Siena, Siena, Italy, Italy
  • Daniela Tarau, Laimburg Research Centre, Italy
  • Luca Fontanesi, Laimburg Research Centre, Corporate System Improvement Kerakoll Group, Italy
  • Mattia Tabarelli, Laimburg Research Centre, Italy
  • Cameron Cullinan, Laimburg Research Centre and University of Bolzano, Italy
  • Cecilia Mittelberger, Laimburg Research Centre, Italy
  • Katrin Janik, Laimburg Research Centre, Italy


Presentation Overview: Show

Apple proliferation (AP) disease, caused by 'Candidatus Phytoplasma mali', represents one of the most significant threats to apple (Malus domestica) cultivation. The pathogen is a phloem-limited, cell wall–less bacterium transmitted primarily by psyllid vectors. AP disease presents itself as a wide range of symptoms including the characteristic rosettes, witches' brooms, and enlarged stipules, as well as more general symptoms like leaf yellowing or reddening, reduced growth, and undesirable fruits.
In this study, we focused on the methylation profile of the host plant, M. domestica, to better understand host responses to infection. DNA methylation studies have recently emerged as an important field for advancing our understanding of plant–pathogen interactions. A set of infected apple trees was selected, and both symptomatic and asymptomatic leaves were collected from each plant. Infection was confirmed by PCR analysis of genomic DNA extracted from root tissues.
Subsequently, DNA samples were subjected to bisulfite treatment, sequencing libraries were prepared, and high-throughput Next Generation Sequencing was performed. Here, we present a bioinformatic pipeline for the analysis of DNA methylation profiles in infected apple plants, providing insights into epigenetic changes associated with 'Ca. P. mali' infection.
The analysis workflow comprised quality control, read trimming, and mapping to the reference genome and corresponding annotation. Furthermore, statistical analysis of the different genomic elements of M. domestica were conducted on CpG, CHG and CHH methylation. This approach should help to understand the influence of DNA methylation patterns on the manifestation of symptoms in apple trees.

B-G.08: NanoCanvas: A Unified Interactive Browser for Multimodal Nanopore Sequencing Data
Track: Genomics, epigenomics, and genome editing
  • Guangzhao Cheng, FIMM, HiLIFE & Applied Tumor Genomics Research Program, University of Helsinki, Finland
  • Esa Pitkänen, FIMM, HiLIFE & Applied Tumor Genomics Research Program, University of Helsinki; iCAN Flagship, Helsinki, Finland, Finland


Presentation Overview: Show

Motivation: Nanopore sequencing has become a transformative technology for genomic and epigenomic research, generating multimodal data including raw electrical signals (POD5/FAST5), basecalled alignments (BAM), resquiggled signal events (EventAlign TSV), and base modification calls (BED). However, the analytical ecosystem remains fragmented: researchers must juggle incompatible tools and formats, struggle with datasets spanning tens to hundreds of gigabytes, and manually script ad-hoc comparisons across cohorts (e.g., wild-type vs. knockout). Existing visualization tools generate static images and crash under interactive queries at this scale.

Results: We present NanoCanvas, an open-source, web-based platform that unifies nanopore data exploration through interactive multimodal visualization. A FastAPI backend ingests POD5/FAST5, BAM, EventAlign TSV, BED, and reference FASTA into a tiered indexing layer that retrieves any single read from multi-gigabyte POD5 datasets in approximately 50 ms, scaling near-constantly with cohort size, and delivers cohort-level QC and signal aggregation, reducing exploratory analyses from minutes of static plotting (e.g., NanoPlot) to sub-second interactive responses. The frontend integrates five workspaces (quality control, single-read inspection, k-mer context model, transcript view, and gene/locus view) that synchronize raw signals, basecalls, and modifications within a single physical coordinate system. Native multi-cohort overlays enable comparison of QC metrics, signal distributions, and modification patterns across groups, eliminating custom plotting scripts. By unifying genomic context, basecalls, modifications, and raw signal in a single interactive environment, NanoCanvas lowers the barrier to exploratory nanopore analysis beyond specialist users.

B-G.09: Discovery of disease genes by combining transcription factor, epigenome and DNA variation data
Track: Genomics, epigenomics, and genome editing
  • Nina Baumgarten, Goethe University Frankfurt, Germany
  • Marcel H. Schulz, Goethe University Frankfurt, Germany


Presentation Overview: Show

Genome wide association studies (GWAS) identified thousands of genetic variants, such as Single Nucleotide Variants (SNVs) associated to traits or diseases. A substantial amount of these variants is non-coding and might be located within cis-regulatory elements (CREs). They can affect Transcription Factor (TF) binding sites and alter the target gene expression. Identifying those genes is a crucial step in understanding the molecular mechanisms underlying a trait or disease.
Previously developed method for disease gene identification primarily focuses on DNA variation and/or epigenome data, neglecting the regulatory aspect. Especially for non-coding SNVs considering altering TF binding is important to understand the regulatory mechanisms of disease genes.
We present an approach for the interpretation of non-coding variants, by identifying regulatory SNVs (rSNVs), which are predicted to affect TF binding sides. Furthermore, we integrate epigenome data to link rSNVs to target genes and use statistical methods to aggregate information of TFs and genes. To priorities disease genes with high confidence, we model for each gene the expected number of CREs overlapping with rSNVs using the Poisson-binomial distribution, while considering the LD structure. Using GWAS from different complex diseases, we showcase that our approach can identify disease relevant genes, whether protein-coding or non-coding RNA and highlight relevant genes not found with a well-established method. Further, we illustrate the usefulness of our approach using epigenome and variation data from the epiATLAS to predict disease-specific changes in TF binding sites across neuropsychiatric disorders.

B-G.10: Assessing SNV and SV Callers using HiFi Long Read WGS to Find Causal Variants for Mendelian Traits in Goats
Track: Genomics, epigenomics, and genome editing
  • Laura Voitl, Institute of Genetics, University of Bern; Interfaculty Bioinformatics Unit, University of Bern, Switzerland
  • Rémy Bruggmann, Interfaculty Bioinformatics Unit, University of Bern; Swiss Institute of Bioinformatics, Switzerland
  • Cord Drögemüller, Institute of Genetics, University of Bern, Switzerland
  • Anna Letko, Institute of Genetics, University of Bern; Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

Long-read whole-genome sequencing (LR-WGS) has the advantage of enabling detection of large and repetitive structural variants (SV) compared to short-read (SR-WGS). Here, we sequenced 20 Swiss goat genomes representing 10 local breeds using PacBio Revio. To broaden the diversity of our dataset, we included 12 publicly available genomes from the European Nucleotide Archive. We aligned all samples using pbmm2 to the T2T-goat1.0 Inner Mongolia cashmere goat reference available on NCBI.
We subsequently constructed variant catalogs based on the aligned genomes using a reference-based approach. We used two tools for SNV calling (DeepVariant and clair3) and SV calling (Sniffles2 and Sawfish2) respectively. SNV calls were also compared to SR-WGS data of the Swiss cohort. Our aim was to compare the output of each caller by investigating which variants were called by each tool. We also tested for the presence of known caprine variants for Mendelian traits from the Online Mendelian Inheritance in Animals (OMIA) database.
Both shared and unique variants were detected across the different callers. Known variants of interest included copy number variants affecting coat color, a complex SV on chromosome 1 causing polled intersex syndrome, and variants present in CSN1S1, which affect milk protein. We confirmed the expected presence of these different variants in the Swiss goats and evaluated their occurrence in the public dataset.
This work shows the benefit of using LR-WGS to detect the full spectrum of genetic variation in goats. The variant catalogs provide a valuable and sustainable resource for small ruminant genomics.

B-G.11: The nationwide genomic characteristics and phylogenetic evolution of ST23-K1 hypervirulent Klebsiella pneumoniae in relation to virulence and antimicrobial resistance acquisition
Track: Genomics, epigenomics, and genome editing
  • Qiucheng Shi, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Jingyi Zhu, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Rui Weng, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Guanhong Lu, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Yan Pan, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Ping Zhang, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Jingjing Quan, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Dongdong Zhao, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Yunsong Yu, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
  • Zhengan Wang, Zhejiang University, China
  • Yan Jiang, Zhejiang University, China


Presentation Overview: Show

Objectives: Hypervirulent Klebsiella pneumoniae (hvKp) ST23-K1 poses a global health threat due to its high virulence and increasing antimicrobial resistance. This study aimed to characterise the genomic feature and phylogenetic evolution of ST23-K1 in China.
Methods: K1 isolates from a nationwide epidemiological surveillance project underwent whole-genome sequencing. Virulence was assessed using hypermucoviscosity phenotyping and a murine infection model. For ST23-K1 carrying acquired antimicrobial resistance genes (ARGs), the CRISPR/Cas system, protospacers, anti-CRISPR (Acr) genes, and plasmidome were characterised. Time-resolved phylogenetic analysis was performed using integrated locally generated and publicly available data.
Results: Among 400 K1 isolates, ST23 was the most prevalent sequence type, and its effective population size increased following CG23-I divergence. The CG23-I sub-lineage was widely distributed nationwide with limited evidence of clonal transmission. Isolates with an incomplete cps locus exhibited significantly reduced virulence compared with those carrying an intact locus. The prevalence of extended-spectrum β-lactamase-positive ST23-K1 isolates increased over time, whereas carbapenemase-producing isolates remained stable. Among acquired ARGs-positive ST23-K1 isolates, a conserved protospacer corresponding to a prevalent spacer was identified. This protospacer, together with AcrIE genes, was frequently co-located on IncFII-type plasmids.
Conclusion: ST23-K1 remains a hypervirulent lineage undergoing ongoing evolutionary expansion. The presence of acquired ARGs in ST23-K1 may be associated with AcrIE-harbouring IncFII plasmids, and functional validation is required to clarify the underlying mechanisms. Continuous genomic surveillance is essential to monitor the evolution and antimicrobial resistance trends of ST23-K1.

B-G.12: STOAT : a tool for pangenome-guided genome wide association studies
Track: Genomics, epigenomics, and genome editing
  • Xian Chang, MIAT, INRAE, France
  • Matis Alias-Bagarre, IRSD, INSERM, France
  • Jean Monlong, IRSD, INSERM, France
  • Matthias Zytnicki, MIAT, INRAE, France


Presentation Overview: Show

Genome wide association studies (GWAS) are a fundamental tool for finding associations between genotypes and phenotypes, but standard GWAS tools are limited by their reliance on the reference genome. Such studies are generally limited in their ability to detect large and complex structural variants, and are incapable of detecting potentially significant “off-reference” variants that do not touch the reference genome, such as a SNP that is inside of an insertion. Pangenomics is a powerful emerging paradigm that uses a collection of genomes as a reference, rather than a single haplotype reference genome. Pangenomes facilitate the identification of off-reference variants, including large and complex structural variants.
We present a new tool for using pangenomes to perform GWAS. Our tool, STOAT (snarl tree-orchestrated association test), uses a pangenome to identify and characterize variants and the hierarchical relationship among them, including off-reference variants that may be missed by traditional methods. Using either the samples in the pangenome graph or variant calls from a pangenomic genotyping pipeline, we can test all variants identified in the pangenome and in particular, we are able to independently test nested variants, ignoring their parent or child variant. We show on simulated and real data that STOAT is able to detect the same variants deemed significant by traditional methods, and that it can test variants that these tools would miss.

B-G.13: A reference-free pangenomic pipeline to uncover transposable element dynamics across genomes
Track: Genomics, epigenomics, and genome editing
  • Johann Confais, INRAE / URGI, France
  • Hadi Quesneville, URGI-INRAE, France
  • Somia Saidi, URGI-INRAE, France


Presentation Overview: Show

Transposable elements (TEs) are major drivers of genome evolution and adaptation, yet their intraspecific dynamics remain difficult to characterize due to methodological biases associated with reference-based approaches. With the increasing availability of high-quality de novo genome assemblies, there is a need for scalable tools that can accurately capture TE diversity at the pangenome level.

Here, we present panREPET, a novel reference-free computational pipeline designed to detect and characterize shared TE insertions across multiple genomes. By performing pairwise comparisons of independently annotated assemblies, panREPET identifies homologous TE copies based on sequence similarity and genomic context, enabling the reconstruction of TE insertion histories without reliance on a reference genome. The method provides precise coordinates and full sequences of TE copies in each genome, and classifies them into core, shared, and singleton insertions.

We applied panREPET to 42 Brachypodium distachyon genomes, identifying over 18,000 shared TE insertions and revealing that the majority of TE diversity is driven by recent, lineage-specific events. Using a SNP-based dating framework, we uncovered four major bursts of TE activity, including two associated with key climatic transitions—the Last Glacial Maximum (~22 kya) and the Holocene (~10 kya)—suggesting a link between environmental stress and TE mobilization. Comparative benchmarking shows that panREPET improves specificity and resolution over existing reference-based and structural-variant-based methods.

Overall, panREPET enables a comprehensive and unbiased exploration of TE dynamics at the species level, providing new insights into the role of mobile elements in genome evolution and environmental adaptation.

B-G.14: Multi-feature predictive modeling of novel cancer predisposition genes using pan-cancer data
Track: Genomics, epigenomics, and genome editing
  • Jaejun Lee, Spanish national cancer research center, Spain
  • A-Reum Nam, Spanish national cancer research center, Spain
  • Solip Park, Spanish national cancer research center, Spain


Presentation Overview: Show

Cancer predisposition genes (CPGs) are defined as those inherited variants that increase the risk of tumorigenesis. Despite their functional and clinical role in tumorigenesis, the number of identified CPGs remains limited. Due to the rarity of pathogenic germline variants and the limitations of traditional case-control approaches, the systematic approach to identify novel cancer predisposition genes is demanding. We developed a machine learning-based predictive model integrating 13 various biological features, including germline genomic mutation status, gene expression, somatic second hit mutation, and general characteristics of human genes, to identify novel candidates. While single features possessed moderate predictive performance, our integrative model outperformed all single-feature models (AUC =0.74 vs. 0.53-0.58), representing the value of multi-feature integration. We further applied the model across individual cancer types with sufficient sample sizes. Overall, we identified the top 8 high-confidence CPG candidates at both cancer-type-specific and pan-cancer levels. Several of the candidates were validated using an independent cancer cohort. This approach provides a generalizable and scalable framework for uncovering novel CPGs and offers a new direction for understanding the genetic basis of cancer susceptibility.

B-G.15: Genomic scars of chemotherapy mark the emergence of resistance in childhood cancer
Track: Genomics, epigenomics, and genome editing
  • Laura Wheaton, Queen's University, Canada
  • Marie Wong-Erasmus, The University of New South Wales, Australia
  • Max F. Levine, Memorial Sloan Kettering Cancer Center, United States
  • Katherine E. Miller, Nationwide Children's Hospital, United States
  • Neerav N. Shukla, Memorial Sloan Kettering Cancer Center, United States
  • Michael D. Kinnaman, Memorial Sloan Kettering Cancer Center, United States
  • Dominik Glodzik, Harvard Medical School, United States
  • Carol Portwine, McMaster Children's Hospital, Canada
  • Sabrina Millson, McMaster Children's Hospital, Canada
  • Alexandra Zorzi, Children’s Hospital at London Health Sciences Centre, Canada
  • Mariam Mikhail, Children’s Hospital at London Health Sciences Centre, Canada
  • Conrad Fernandez, IWK Health Centre, Canada
  • Noemi A. Fuentes-Bolanos, The University of New South Wales, Australia
  • Gunes Gundem, Memorial Sloan Kettering Cancer Center, United States
  • Andrew L. Kung, Memorial Sloan Kettering Cancer Center, United States
  • Uri Tabori, The Hospital for Sick Children, Canada
  • Chelsea Mayoh, The University of New South Wales, Australia
  • Elli Papaemmanuil, Memorial Sloan Kettering Cancer Center, United States
  • Mark J. Cowley, The University of New South Wales, Australia
  • David Malkin, The Hospital for Sick Children, Canada
  • Anita Villani, The Hospital for Sick Children, Canada
  • Ludmil B. Alexandrov, University of California San Diego, United States
  • Adam Shlien, The Hospital for Sick Children, Canada
  • Timmy Wen, The Hospital for Sick Children, Canada
  • Marcos Díaz-Gay, Spanish National Cancer Research Center, Spain
  • Eric N. Bergstrom, University of California San Diego, United States
  • Mathepan J. Mahendralingam, The Hospital for Sick Children, Canada
  • Nicholas Light, The Hospital for Sick Children, Canada
  • Sasha Blay, The Hospital for Sick Children, Canada
  • Joshua O. Nash, The Hospital for Sick Children, Canada
  • Nathaniel D. Anderson, The Hospital for Sick Children, Canada
  • Jessica N. Au, University of California San Diego, United States
  • Scott Davidson, The Hospital for Sick Children, Canada
  • Pedro L. Ballester, The Hospital for Sick Children, Canada
  • Mehdi Layeghifard, The Hospital for Sick Children, Canada
  • Syed Kashif Daud, The Hospital for Sick Children, Canada
  • Lisa-Monique Edward, The Hospital for Sick Children, Canada
  • S.M. Ashiqul Islam, University at Albany, United States
  • Azhar Khandekar, University of California San Diego, United States
  • Burcak Otlu, Middle East Technical University, Turkey
  • Ledia Brunga, The Hospital for Sick Children, Canada
  • Rawan Hammad, King Abdulaziz University, Saudi Arabia
  • Shimaa Nassif, The Hospital for Sick Children, Canada
  • Nirav H. Thacker, The Hospital for Sick Children, Canada
  • Tara Feltham, The Hospital for Sick Children, Canada


Presentation Overview: Show

Childhood cancer survival now exceeds 80%, yet therapy resistance remains a defining clinical challenge, affecting approximately one-third of patients and driving poor long-term outcomes. A central unanswered question is when resistance emerges and whether it can be detected genomically before it becomes clinically manifest. Chemotherapy leaves characteristic mutational footprints in tumor genomes, and because these signatures require cancer cells to survive drug exposure, they represent a direct molecular record of resistance. In this large-scale multi-omics study of hard-to-treat pediatric tumors from three independent precision medicine programs, linked to curated exposure data for 86 therapies, we establish therapy-induced mutational signatures as markers of resistance emergence. Comprehensive signature analysis identified novel signatures exclusive to treated tumors. Platinum drugs were the dominant mutagen, accounting for 10% of all single-base substitutions. Critically, platinum signatures were detectable as early as 91 days after treatment initiation, and their accumulation mirrored resistance kinetics: 35% of platinum-treated patients showed measurable signatures within one year, rising to ~50% by 18 months. Signature-positive tumors exhibited significant overexpression of platinum resistance genes and worse outcomes upon re-exposure, functionally validating signatures as resistance biomarkers. Subclonal analysis further identified hidden platinum signatures in tumors lacking dominant resistance clones, suggesting early detection of emerging resistance before clinical failure. An ensemble machine learning model revealed additional platinum-associated genomic features beyond known signatures. These findings establish a genomic framework for monitoring resistance in real time and informing therapy de-escalation in children with cancer.

B-G.16: LongcallD: joint calling and phasing of small, structural and mosaic variants from long reads
Track: Genomics, epigenomics, and genome editing
  • Yan Gao, Harbin Institute of Technology, China


Presentation Overview: Show

Long-read sequencing is a powerful technique capturing multiple variants within single continuous reads. This length allows individual reads to bridge small and structural variants while carrying crucial phasing information. However, current computational tools treat small variant calling, structural variant (SV) detection and phasing as largely disconnected problems, failing to unleash the full potential of long reads. Here, we present longcallD, a unified framework utilizing local multiple-sequence alignment to simultaneously call and phase small and structural variants. By integrating germline phasing and retrotransposition hallmarks, longcallD also identifies low-fraction mosaic variants and detects mobile element insertions supported by a single read. Compared to existing methods, our unified approach substantially improves SV discovery and mosaic variants accuracy while maintaining competitive small variant calling. We anticipate that longcallD will provide a robust foundation for resolving complex genetic architectures in clinical and evolutionary applications.

B-G.17: Integration of Chromatin Interaction Maps with GWAS Identifies TBKBP1 as a Novel Chronic HBV Susceptibility Gene
Track: Genomics, epigenomics, and genome editing
  • Qiao Ye, Laboratory of Clinical Medicine, Air Force Medical Center, Air Force Medical University, China
  • Xiangyi Zheng, Laboratory of Clinical Medicine, Air Force Medical Center, Air Force Medical University, China
  • Guangyun Wang, Laboratory of Clinical Medicine, Air Force Medical Center, Air Force Medical University, China


Presentation Overview: Show

Background: Genome-wide association studies (GWAS) have identified numerous non-coding single nucleotide polymorphisms (SNPs) associated with susceptibility to chronic hepatitis B virus (HBV) infection. However, assigning them to functional genes is challenging, as conventional nearest-gene annotation overlooks long-range chromatin regulatory interactions.
Methods: We conducted a meta-analysis across two East Asian GWAS cohorts, and built liver- and blood-specific chromatin interaction maps (Activity-by-Contact (ABC) and High-throughput Chromosome Conformation Capture (Hi-C)). Using these tissue-resolved regulatory maps, we linked non-coding variants to candidate genes and performed gene-based association tests.
Results: Among genome-wide significant SNPs identified in our meta-analysis, 95.3% were located in non-coding region. Using SNP–gene regulatory pairs from ABC and Hi-C maps to perform gene-based association testing, we identified 187 genes associated with chronic HBV infection (FDR < 0.05), representing a 3.2-fold increase over conventional nearest-gene annotation (45 genes). These 187 candidate genes were enriched in innate immune pathways, including Toll-like receptor (TLR), retinoic acid-inducible gene I (RIG-I), Janus kinase-signal transducer and activator of transcription (JAK-STAT) signaling. We further identified a novel candidate gene, TBKBP1, whose protective variant (rs67919208) was associated with reduced TBKBP1 expression in blood monocytes and liver (expression quantitative trait locus (eQTL) P = 9.35 × 10⁻¹² and 1.06 × 10⁻¹⁵; colocalization posterior probability for shared causal variant (PP4) > 0.8 for both). In vitro assays further demonstrated that TBKBP1 suppresses antiviral immunity, thereby increasing susceptibility to chronic HBV infection.
Conclusion: Integrating chromatin interaction maps with GWAS uncovers novel susceptibility genes for chronic HBV infection, including TBKBP1.

B-G.18: Pangenome graph annotation
Track: Genomics, epigenomics, and genome editing
  • Nina Marthe, IRD, France
  • Matthias Zytnicki, INRAE, France
  • Francois Sabot, IRD, France


Presentation Overview: Show

The increasing availability of genome sequences has highlighted the limitations of using a single reference genome to represent the diversity within a species. Pangenomes, encompassing the genomic information from multiple genomes, offer thus a more comprehensive representation of intraspecific diversity. However, pangenomes in form of graph often lack annotation information, which limits their utility for forward analyses.

The tool GrAnnoT was designed to transfer annotation in such graphs efficiently and reliably, by projecting existing annotations from a single source genome to the graph, and subsequently to other embedded genomes. It provides informative outputs, such as presence-absence matrices for genes, and alignments of transferred features between source and target genomes, aiding in the study of genomic variations and evolution.
GrAnnoT was published last year in PCJ : 10.24072/pcjournal.651

GrAnnoT was then improved to handle multiple source genomes for the annotation, giving a more complete graph annotation. This also allows to compare the annotations between the genomes embedded in the graph, and explore the variations within the genes.

We also developped a method to annotate de novo the parts of the graph that are not covered by annotated genomes, ensuring annotation information is available for the whole pangenome.

B-G.19: Improved 16S rRNA-focused computational methods for bacterial strain identification from long read datasets
Track: Genomics, epigenomics, and genome editing
  • Laura Tingley, UKHSA, United Kingdom
  • Jo Dicks, Culture Collections, UK Health Security Agency, 61 Colindale Avenue, London, NW9 5EQ, UK, United Kingdom
  • Katharina T. Huber, University of East Anglia, Norwich Research Park, Norwich, NR4 7TJ, UK, United Kingdom


Presentation Overview: Show

The 16S region of the ribosomal DNA (rDNA) has for decades served as a vital resource in the identification and classification of bacterial species, first in taxonomic and phylogenetic studies and more recently in metagenomic analysis. However, bacterial genomes contain between 1 and ~15 copies of the 16S, with unquantified levels of sequence variation between them. This poses several questions. How much 16S variation exists within a genome? Do the 16S sequences of distinct species overlap? Can intra-genome 16S variation alter identification outcomes?

We used long PacBio reads from UKHSA's NCTC3000 dataset, which have the capacity to span the 16S gene and beyond the rDNA operon, to assess sequence variation between distinct 16S sequence copies in Escherichia coli, Escherichia marmotae and Shigella genomes. We developed a computational pipeline to extract operon copy-specific reads and create consensus sequences for each rDNA operon. Extensive analysis was then conducted to provide insight into the breadth of 16S alleles and variants within and between both strains and species.

Extensive 16S variation was observed between copies, with 168 distinct 16S alleles amongst the 53 genomes analysed. Despite this high level of variation, almost all alleles were found within a single species, with only three 16S sequences shared between E. coli and Shigella genomes and no overlap between the other genome pairs. This remarkable finding suggests our new pipeline could enable rapid, enhanced bacterial identification capable of discriminating between phylogenetically similar but distinct bacterial species using a multi-allelic 16S method facilitated by long read analysis.

B-G.20: Ontology-Driven Prioritization of Candidate Genes in Neurodevelopmental Disorders
Track: Genomics, epigenomics, and genome editing
  • Chiara Rivi, BioFolD Unit, Department of Pharmacy and Biotechnology, University of Bologna, Via gobetti 85, 40126 Bologna (Italy), Italy
  • Emidio Capriotti, BioFolD Unit, Department of Pharmacy and Biotechnology, University of Bologna, Via gobetti 85, 40126 Bologna (Italy), Italy


Presentation Overview: Show

Neurodevelopmental disorders (NDDs) represent a complex group of conditions characterised by impairments in brain functions affecting social, motor, and cognitive abilities to varying degrees. Their genetic heterogeneity and broad phenotypic spectrum hinder the understanding of the molecular mechanisms underlying them, slowing the advancement of diagnostic processes and therapeutic strategies.
This work involved a multi-step approach to identify genes associated with NDDs. We first integrated HGNC genes related to NDDs from several databases, including: Orphanet, SFARI, GeneTrek and a recent publication. Then, a score based on the consistency of annotations across the different sources was used to divide it into different sets with varying degree of association to NDDs. The obtained dataset was then subjected to a gene enrichment analysis, in order to pinpoint overrepresented functions, processes, and pathways. A Multi-Ontology Enrichment (MOE) score was hence developed. This approach prioritizes genes based on functional annotation overlap across different Gene Ontology domains and two pathway databases.
We retrieved 9,495 genes related to NDDs, of which 2,639 showed high-confidence associations. The gene enrichment analysis identified several enriched biological processes and pathways related to nervous system development, synaptic processes, energy metabolism, mitochondrial dysfunction, tRNA aminoacylation and notably, cancer-related pathways. Our MOE scoring system highlighted genes with convergent functional annotations, underlying in particular 534 genes common across all the domains tested, hence obtaining maximum MOE score. This work provides new insights into the complex genetic landscape of NDDs, by advancing the understanding of NDD-associated genes, offering a novel way to prioritize candidate genes.

B-G.21: Sincei: A toolkit for exploring single-cell epigenomics data
Track: Genomics, epigenomics, and genome editing
  • Fernando Sancho Gómez, Utrecht University, Netherlands
  • Soufiane Mourragui, Ensocell, United Kingdom
  • Vivek Bhardwaj, Utrecht University, Netherlands


Presentation Overview: Show

Emerging single-cell sequencing protocols allow researchers to study multiple layers of epigenetic regulation while resolving tissue heterogeneity. However, despite the rising popularity of such single-cell epigenomic assays, a lack of user-friendly computational tools for flexible quality control and genome-wide data exploration hinders their broad adoption. We introduce the Single-Cell Informatics (sincei) toolkit, a command-line interface for the exploration of data from a wide range of single-cell (epi)genomics protocols directly from aligned reads stored in BAM format. Sincei provides tools for preprocessing, feature identification, signal aggregation, dimensionality reduction, clustering and visualization. Results are stored in the AnnData format for seamless compatibility with a wide variety of tools. Additionally, sincei includes a Python API that enables advanced use cases, such as generating cell embeddings using Generalized Principal Component Analysis (GLM-PCA). We show that sincei resolves cellular heterogeneity and improves interpretation of genomic regions by using single-cell histone modification signal across species, tissues and protocols.

B-G.22: Methylation-aware read representation for metagenomic classification and host decontamination
Track: Genomics, epigenomics, and genome editing
  • Valentina Galeone, Robert Koch Institute, Germany
  • Akiyama Manato, Kitasato University, Japan
  • Yasubumi Sakakibara, Kitasato University, Japan
  • Martin Hölzer, Robert Koch Institute, Germany


Presentation Overview: Show

Bacterial methylation profiles offer a largely underexploited layer of biological information. Generated by strain-specific restriction-modification systems, methylation patterns vary substantially even within a species and carry a signal that is orthogonal to sequence composition. Crucially, these signals are best preserved at the read level, before assembly erases molecule-level information. Until recently, technical noise in modification calling and limited ground truth have held back read-level exploitation of this signal, limitations that long-read nanopore sequencing is beginning to overcome.
Here, we propose a methylation-aware representation of sequencing reads to improve metagenomic classification and host decontamination, treating reads not only as nucleotide sequences but also as carriers of modification signals.
To this end, we extend standard k-mer frequency profiles from a 4-symbol nucleotide alphabet to a 7-symbol alphabet incorporating per-base calls for 5mC, 4mC, and 6mA from Oxford Nanopore sequencing. We show that methylation-aware k-mers consistently improve classification accuracy over sequence-only baselines, with especially strong gains for host decontamination, where the dense CpG methylation signature of mammalian genomes provides a categorical signal distinct from any bacterial pattern with the potential of transferring readily across eukaryotic hosts. Beyond host decontamination, we investigate whether methylation-aware profiles can resolve strain-level substructure and address metagenomic challenges where sequence composition alone reaches a ceiling, including the long-standing problem of linking mobile genetic elements to their host genomes.
To address these and broader metagenomic challenges, we develop methods that integrate sequence and methylation signals to capture richer organism-specific signatures.

B-G.23: A Synthetic Data Framework for Evaluating Structural Variant Callers in Tumor-Normal Long-Read Sequencing
Track: Genomics, epigenomics, and genome editing
  • Francisco José Villena González, Bioinformatics Unit, Spanish National Cancer Research Centre (CNIO), Madrid, Spain, Spain
  • Fátima Di Domenico Al-Shahrour, Bioinformatics Unit, Spanish National Cancer Research Centre (CNIO), Madrid, Spain, Spain
  • Tomás Di Domenico, Bioinformatics Unit, Spanish National Cancer Research Centre (CNIO), Madrid, Spain, Spain


Presentation Overview: Show

Structural variants (SVs) are genomic alterations encompassing deletions, insertions, and segment rearrangements, ranging from kilobases to entire chromosomes. Yet they remain understudied compared to single nucleotide variants, largely because short-read sequencing technologies struggle to resolve complex genomic regions. The emergence of long-read sequencing has transformed this landscape, dramatically improving SV detection capabilities.

Despite these advancements, benchmarking SV callers in tumor-normal settings presents a major challenge: real biological datasets require costly, time-consuming experimental procedures and often lack a reliable ground truth. In silico approaches offer a powerful alternative, allowing controlled, reproducible evaluation where the true variant landscape is fully known.

In this work, we developed a pipeline to generate synthetic tumor-normal long-read sequencing datasets with defined SV profiles, and used them to systematically benchmark long-read SV calling tools. Notably, no consensus has yet emerged on a gold-standard method for somatic SV detection, making rigorous and reproducible benchmarking efforts particularly critical. Analyses were focused on large SVs inspired by the genomic landscape of Multiple Myeloma, enabling our results to directly inform variant detection decisions in an ongoing real-world research study on this disease.

This framework provides a scalable, bias-controlled solution for evaluating SV callers, with direct implications for the design of somatic variant detection pipelines in cancer genomics.

B-G.24: polars-bio: fast, scalable, and out-of-core genomic intervals and format I/O for Python DataFrames
Track: Genomics, epigenomics, and genome editing
  • Marek Wiewiorka, Institute of Computer Science, Warsaw University of Technology, Poland
  • Tomasz Gambin, Institute of Computer Science, Warsaw University of Technology, Poland


Presentation Overview: Show

Motivation. Operations on genomic intervals, such as overlap, nearest, coverage, and count_overlaps, are foundational to bioinformatics pipelines, from variant annotation to 3D chromatin analysis. Yet widely used Python libraries (Pybedtools, PyRanges, Bioframe, GenomicRanges) struggle past a few million intervals, lack multi-threaded out-of-core execution, and tie peak memory to dataset size. The I/O layer feeding them, such as pysam, is an equally severe bottleneck at the biobank scale.

Results. We present polars-bio, a Python/Rust library on Apache DataFusion, Arrow, and Polars unifying genomic format readers (BAM/CRAM/VCF/FASTQ/FASTA/GFF/GTF/BED) and multiple interval operations in a single vectorized, streaming, multi-threaded engine (Wiewiórka et al., Bioinformatics 2025). In single-threaded benchmarks (AIList, 10^7 vs 1.2×10^6 intervals), polars-bio outpaces Bioframe by 6.5× (overlap), 15.5× (nearest), 38× (count_overlaps), and 15× (coverage), using up to 90× less peak memory in streaming mode. Recent releases extend the operation set to eight (adding cluster, complement, merge, and subtract) and, still single-threaded, is the fastest library in all operations on the largest tested dataset, ahead of PyRanges1,GenomicRanges, and Bioframe. polars-bio also delivers fast, out-of-core, multi-threaded I/O across many popular file formats with range predicates pushdown via file indexes (CSI/BAI/CRAI), column pruning, and limit optimizations, and write support for native formats. It enables efficient file-format transformations with near-linear thread scalability on both interval operations and genomic readers. Federated SQL queries over cloud storages extend these capabilities to datasets exceeding local memory and storage. Recent benchmark results are documented in recent project blog posts (https://biodatageeks.org/polars-bio/blog/2026/02/14/benchmarking-genomic-format-readers-in-python-with-polars/; https://biodatageeks.org/polars-bio/blog/2026/02/20/interval-operations-benchmark--update-february-2026/)

Availability. Open source (Apache 2.0)
https://biodatageeks.org/polars-bio/

B-G.25: vepyr: a fast, scalable, and composable Apache DataFusion-based variant annotation engine
Track: Genomics, epigenomics, and genome editing
  • Tomasz Gambin, Institute of Computer Science, Warsaw University of Technology, Poland
  • Marek Wiewiórka, Warsaw University of Technology, Poland


Presentation Overview: Show

Motivation. Ensembl VEP is the de facto standard for variant annotation, yet its Perl codebase and plugin architecture (CADD, gnomAD, SpliceAI via tabix) scale poorly: whole-genome annotation of a single WGS VCF takes tens of minutes to hours, and biobank-scale cohorts multiply the cost linearly. There are faster alternatives each address one piece ( bcftools/csq for consequence calling, Illumina Connected Annotations as a C# batch binary), but none combines Rust-native performance, a composable query-engine interface, pluggable cache format and full Ensemble VEP concordance.

Results. We present vepyr, a Rust/Python variant annotation engine built on Apache DataFusion and Apache Arrow, integrated with the polars-bio ecosystem. vepyr exposes annotation as DataFusion table-valued functions that can be integrated with any SQL query plan: (i) lookup_variants() for co-located variant and population-frequency lookup, and (ii) annotate_variants() for consequence prediction, HGVS, and IMPACT classification. The novel cache layer offers two interchangeable paths: a full Parquet cache consumed through any DataFusion TableProvider for scan-heavy workloads, and an optimized fjall LSM-tree store tuned for fast point lookups. Predicate and projection pushdown with columnar file format ensure only requested annotation columns are read from disk, eliminating VEP's full-bundle-per-variant overhead, and streaming execution enables out-of-core annotation of population-scale cohorts. A plugin framework for sources such as SpliceAI, AlphaMissense, and CADD is in progress, as well as multi-threaded execution. On the GIAB HG002 v4.2.1 benchmark (4.2M variants, GRCh38), preliminary results show up to 50× speedup over Ensembl VEP when run with --everything and --hgvsc flags.

Availability. Open source (Apache 2.0); https://biodatageeks.org/vepyr/

B-G.26: MSClust: de novo clustering of single-cell methylome sequencing data
Track: Genomics, epigenomics, and genome editing
  • Johnathan Wong, The University of British Columbia, Canada
  • Lauren Coombe, BC Cancer Research Institute, Canada
  • Parham Kazemi, BC Cancer Research Institute, Canada
  • Rene Warren, BC Cancer Research Institute, Canada
  • Inanc Birol, The University of British Columbia, Canada


Presentation Overview: Show

DNA methylation is a crucial epigenetic modification, playing a central role in regulating gene expression. To detect methylation at single-base resolution, researchers commonly rely on bisulfite sequencing, which converts unmethylated cytosines to uracil. Because conventional bulk sequencing masks methylation heterogeneity between cell types, single-cell approaches are required; however, single-cell data are characterized by sparsity. To generate comprehensive cell type methylation profiles, current methods must map these sparse reads to a reference before clustering cells based on shared epigenetic signals. Yet, the reduced genomic complexity inherent in bisulfite conversion leaves 40–60% of reads unmapped potentially overlooking subpopulations defined by these regions.

We introduce MSClust, a novel reference-free clustering methodology capable of using methylation information from all reads. For each cell, MSClust records the methylation state of CG sites within a two-tiered Bloom filter data structure. These high-dimensional profiles are then processed through UMAP dimensionality reduction and spectral clustering to identify cell types. When validated on a dataset of 32 mouse embryonic stem cells (~14X), MSClust achieved a 1:1 match with the experimental ground truth while demonstrating a 10-fold and 3-fold decrease in run time and memory consumption, respectively, compared to traditional methods (Run time: 3h, Memory: 22GB). On a larger dataset of 1,390 human neurons (~259X), the method maintained high concordance (Adjusted Rand Index = 0.76) with experimental ground truth. These results suggest that MSClust provides a scalable, reference-independent framework that enables discovery and study of novel biomarkers that are otherwise obscured by reference-mapping limitations.

B-G.27: Benchmarking knowledge graph embedding models for the prediction of oligogenic combinations
Track: Genomics, epigenomics, and genome editing
  • Inas Bosch, Université Libre de Bruxelles, Vrije Universiteit Brussels, Belgium
  • Barbara Gravel, Université Libre de Bruxelles, Vrije Universiteit Brussels, Belgium
  • Alexandre Renaux, Universit´ Libre de Bruxelles, Vrije Universiteit Brussels, Belgium
  • Ann Nowé, Vrije Universiteit Brussels, Belgium
  • Maris Laan, University of Tartu, Estonia
  • Tom Lenaerts, Université Libre de Bruxelles, Vrije Universiteit Brussels, Belgium


Presentation Overview: Show

Identifying the oligogenic causes of rare diseases remains a challenge, notwithstanding the advancements made in the last decade. While a variety of predictive and ranking approaches have been proposed, their precision remains limited, as it remains difficult to know which features may be most relevant for the design of new predictors. We hypothesize that structured biological information, which provides an integration of relevant biological networks and ontologies in a single heterogeneous knowledge graph, can make a difference as it allows for learning a relevant genetic representation through KGE methods. An exhaustive benchmarking is performed wherein we assess the performance of various state-of-the-art embedding models for the task of identifying potentially pathogenic gene pairs. The results obtained show that these KGEs provide highly accurate predictions, leading to an AUC PR of up to 0.93, representing a significant advancement over previous approaches. We show nonetheless that care needs to be taken in the cross-validation when using embeddings, as data leakage between folds will reveal overly optimistic results. The further evaluation of the methods on a holdout set and on a group of new male infertility cases show that three Translational Distance models (TransE, MurE, RotatE) and two Semantic Matching models (DistMult, QuatE) provide better results. The analysis is concluded by comparing all known gene combinations for these top-ranking models, examining their similarities and differences. Overall, KGEs provide a predictive advancement but new steps will need to be taken to generate explanations as to why the pairs are relevant for oligogenic diseases.

B-G.28: hicVerse: A modular platform for scalable exploration of harmonized Hi-C datasets across human samples
Track: Genomics, epigenomics, and genome editing
  • Karol Piera, Univeristy of Lausanne, Switzerland
  • Gian Marco Franceschini, University of Lausanne, Switzerland
  • Giovanni Ciriello, Universtiy of Lausanne, Switzerland


Presentation Overview: Show

High-throughput chromosome conformation capture (Hi-C) enables genome-wide interrogation of three-dimensional chromatin organization; however, the growing volume and heterogeneity of public datasets limit systematic cross-sample analysis. Here, we present *hicVerse*, a modular platform for the curation, harmonization, and exploration of Hi-C data across diverse human samples spanning healthy and disease conditions.

hicVerse integrates approximately 1000 publicly available Hi-C datasets into a standardized, curated, multi-resolution resource. All datasets are processed through a unified pipeline, including normalization, quality control, and subcompartment assignment using Calder.

To support both interactive exploration and large-scale analysis, hicVerse adopts a dual-layer architecture: a tiled visualization backend, powered by HiGlass, for real-time rendering of contact maps, and a columnar query engine enabling efficient retrieval of genomic interactions, features, and metadata across the entire atlas. Crucially, this architecture enables fast cross-dataset queries over arbitrary genomic loci, which would otherwise require substantial computational resources and extensive preprocessing.

The platform comprises a metadata-rich FastAPI backend (chromAPI), an Angular-based web interface, and a HiGlass-powered visualization layer. By decoupling visualization from analytical querying, hicVerse enables high-performance interactive browsing alongside scalable integrative analyses.

hicVerse provides a unified resource for investigating chromatin architecture across biological contexts, enabling large-scale hypothesis generation and comparative analyses in genome regulation, development, and disease. The processing pipeline and platform components are designed for reproducible deployment and integration of new data, and will be made publicly available.

B-G.29: Introducing the Y-chromosomal ancestral-like reference sequence - Improving the capture of human evolutionary information
Track: Genomics, epigenomics, and genome editing
  • Zehra Köksal, Department of Biomedical and Clinical Sciences, Linköping University, Linköping, Sweden, Sweden
  • Annina Preussner, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki, Finland, Finland
  • Jaakko Leinonen, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki, Finland
  • Taru Tukiainen, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki, Finland


Presentation Overview: Show

An essential part of reproducible genetic analyses is the use of reference sequences. However, the widely used human reference sequences represent evolutionarily young sequences of modern men of mostly European ancestry. For the human Y chromosome (chrY) that is widely used in evolutionary studies, this can result in misleading variant calling. We address this problem by reconstructing the Y-chromosomal ancestral-like reference sequence (Y-ARS) for unambiguous variant calling of evolutionary relevance.

We applied a weighted maximum parsimony approach to human and primate chrY sequencing data to construct the Y-ARS. For Y-ARS benchmarking, we aligned 40 chrY short-read sequences from diverse haplogroups to existing references GRCh37, GRCh38 and T2T-CHM13. Alignment to the Y-ARS yielded the largest and most consistent number of variants across sample (mean=1400; SD=77). Alternative references yielded on average fewer variants with greater variability across samples (mean=866–968; SD=457–531) depending on their phylogenetic distance from the reference. Among the variants called after alignment to alternative references, an average of 46% carried the ancestral allele, while alignments to the Y-ARS resulted in calling solely variants with evolutionarily derived alleles.

Here, we show that the existing human reference sequences cannot capture the full range of evolutionary information on the chrY. The Y-ARS improves presenting evolutionary information on the chrY, which makes it a valuable resource for evolutionary applications, such as sample age estimations (e.g., TMRCA) and phylogenetic analyses. Finally, we provide a publicly available tool, polaryzer, to annotate variants as ancestral or derived in pre-aligned chrY data (vcf files).

B-G.30: Sparrowhawk: running bacterial bioinformatics analyses anywhere, locally, with WebAssembly
Track: Genomics, epigenomics, and genome editing
  • Víctor Rodríguez Bouza, EMBL's European Bioinformatics Institute (EMBL-EBI), United Kingdom
  • John Lees, EMBL's European Bioinformatics Institute (EMBL-EBI), United Kingdom


Presentation Overview: Show

The rapid expansion of public sequencing repositories is transforming the computational demands of genomic studies. Most datasets are shared as raw reads, yet downstream analyses, especially in pathogen surveillance and outbreak response, demand fast, accessible, and privacy-preserving tools. To meet this need, we developed Sparrowhawk, a unified toolkit of basic bioinformatic methods compiled from Rust into WebAssembly and delivered as a single, user-friendly website.

Sparrowhawk integrates a lightweight bacterial genome assembler with additional WebAssembly methods for common infectious-disease workflows, including taxonomic identification, mapping and alignment, transmission clustering, gene calling, and host depletion. This unified, maintenance-light (as everything runs locally) interface lowers the barrier to genomic analysis for users without command-line expertise and supports settings with unreliable connectivity or sensitive patient data.

Benchmarking across six bacterial species with both simulated and real datasets shows the assembler's performance is comparable to Minia, while requiring reduced computational resources. The toolkit runs entirely in the user's browser, eliminating data transfers and enabling secure, offline-friendly analyses ideal for clinical and field environments.

Ongoing work includes general optimisations, GPU-accelerated components, and the integration of antimicrobial resistance detection. Sparrowhawk aims to make high-quality genomic analysis more accessible and sustainable for infectious-disease research and genomics in general.

B-G.31: Adipose core genes drive cardiovascular risk via EMT and adipogenesis pathways
Track: Genomics, epigenomics, and genome editing
  • Amos Romer, Technical University of Munich, Germany
  • Sebastian Doetsch, Technical University of Munich, Germany
  • Anastasiia Diagel, Technical University of Munich, Germany
  • Shuangyue Li, Technical University of Munich, Germany
  • Ling Li, Technical University of Munich, Germany
  • Moritz von Scheidt, Technical University of Munich, Germany
  • Daniel Tews, Ulm University Medical Center, Germany
  • Martin Wabitsch, Ulm University Medical Center, Germany
  • Heribert Schunkert, Technical University of Munich, Germany
  • Matthias Heinig, Technical University of Munich, Germany
  • Zhifen Chen, Technical University of Munich, Germany


Presentation Overview: Show

Genetic loci for complex traits such as coronary artery disease (CAD) are distributed widely across the genome, often mapping near genes with unclear connections to disease biology. The omnigenic model proposes that gene regulatory networks interconnect these loci in disease-relevant cells, allowing peripheral genes to influence a limited set of core effector pathways. To functionally investigate this architecture in adipose tissue, we applied a single-cell CRISPR perturbation platform to genetically prioritized CAD genes in human adipocytes. Perturbation of multiple candidate loci revealed structured transcriptional convergence on a program characterized by suppression of adipogenesis and activation of extracellular matrix remodeling. Network analysis identified a shared set of convergent core genes enriched for cardiometabolic pathways and druggable targets. Integration of perturbation results with human genetic analyses linked several core genes to lipid and metabolic traits, indicating that these intermediate phenotypes may mediate the effects of core genes' function in adipose remodeling on cardiometabolic disease risk.

B-G.32: Determinants of functional burden pleiotropy and gene dosage responses across human traits
Track: Genomics, epigenomics, and genome editing
  • Omar Shanta, Department of Psychiatry, University of California San Diego, La Jolla, CA, USA, United States
  • Sebastien Jacquemont, CHU Sainte-Justine, Canada
  • Guillaume Dumas, CHU Sainte-Justine, Canada
  • David Glahn, Boston Children's Hospital/Harvard Medical School, United States
  • Jonathan Sebat, Department of Psychiatry, University of California San Diego, La Jolla, CA, USA, United States
  • Almasy Laura, Children's Hospital of Philadelphia, United States
  • Stephen W. Scherer, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, Ontario, Canada, Canada
  • Celia M. T. Greenwood, Lady Davis Institute for Medical Research, Jewish General Hospital, Montreal, QC, Canada, Canada
  • Jeffrey R. MacDonald, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, Ontario, Canada, Canada
  • Bhooma Thiruvahindrapuram, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, Ontario, Canada, Canada
  • Sayeh Kazem, Universiry Of Montreal, Canada
  • Worrawat Engchuan, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, Ontario, Canada, Canada
  • Emma E.M Knowles, Harvard Medical School, Department of Psychiatry, United States
  • Laura M. Schultz, Department of Biomedical and Health Informatics, The Children’s Hospital of Philadelphia, Philadelphia, PA, United States
  • Thomas Renne, Montreal University, Canada
  • Josephine Mollon, Harvard Medical School, Department of Psychiatry, United States
  • Guillaume Huguet, Université de Montréal, CHU Sainte Justine, Canada
  • Florian Benitiere, CHU Sainte-Justine, Canada
  • Jane Yang, CHU Sainte-Justine, Canada
  • Kuldeep Kumar, CHU Sainte-Justine, Canada


Presentation Overview: Show

Gene dosage alterations, such as rare copy-number variants (CNVs), are major drivers of disease risk and whole-body multimorbidity. Their rarity makes deciphering this biological footprint challenging. To overcome this statistical limitation, we developed Functional Burden analysis (FunBurd), aggregating protein-coding CNVs across 172 transcriptomic-derived tissue and cell-type networks. Applying this to 43 complex traits in ~500,000 UK Biobank participants, we captured widespread functional associations missed by single-gene approaches.
Mapping this architecture revealed a unifying principle: CNV pleiotropy is fundamentally restricted by evolutionary constraint and overwhelmingly concentrated within brain-specific functions. Crucially, mediation analysis proved these pleiotropic links are driven overwhelmingly (~84%) by direct CNV effects. This architectural constraint simultaneously governs gene dosage responses. Highly conserved brain functions, intolerant to dosage deviation, exhibited predominantly non-monotonic (same-direction) effects, whereas less constrained non-brain traits followed traditional monotonic rules. Within these brain functions, we uncovered profound sex differences: females exhibited a significantly higher proportion of deletion-driven associations for mental health traits.
We successfully replicated these associations in an independent, ancestrally diverse cohort of ~500,000 All of Us participants. Furthermore, deletion burden demonstrated positive effect size correlations (up to r=0.78) with rare pLoF SNVs across 87% of traits, confirming that the effects captured by the functional burden test are consistent with an independent class of loss-of-function variants.
Our results highlight the key role of genetic constraint and brain-specific mechanisms in shaping CNV-driven pleiotropy and monotonic gene-dosage response, providing a mechanistic basis for the whole-body multimorbidity observed in neurodevelopmental and psychiatric conditions.

B-G.33: Multi-scale characterization of epigenomic alterations in pediatric brain tumors
Track: Genomics, epigenomics, and genome editing
  • Neda Shokraneh Kenari, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada
  • Alejandro Mejía García, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada
  • Marco Gallo, Arnie Charbonneau Cancer Institute, Cumming School of Medicine, University of Calgary, Calgary, AB T2N 4N1, Canada, Canada
  • Verónica Rendo, Department of Immunology, Genetics, and Pathology, Uppsala University, Uppsala, Sweden, Sweden
  • Bing Ren, Department of Cellular and Molecular Medicine, University of California, San Diego School of Medicine, La Jolla, CA, USA, United States
  • Mathieu Blanchette, School of Computer Science, McGill University, Montreal, QCH3A 2A7, Canada, Canada
  • Audrey Baguette, Quantitative Life Sciences, McGill University, Montreal, Quebec H3A 2A7, Canada, Canada
  • Bhavyaa Chandarana, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada
  • Steven Hébert, Lady Davis Research Institute, Jewish General Hospital, Montreal, QC H3T 1E2, Canada, Canada
  • Nathan Zemke, Department of Cellular and Molecular Medicine, University of California, San Diego School of Medicine, La Jolla, CA, USA, United States
  • Michael Taylor, Department of Pediatrics, Baylor College of Medicine, Houston, TX, 77030, USA, United States
  • Nada Jabado, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada
  • Claudia L. Kleinman, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada


Presentation Overview: Show

Many lethal pediatric brain tumors are driven by aberrant epigenetic landscapes that inhibit the normal differentiation of neural progenitor cells and transform them into a malignant state. The epigenetic landscape is multilayered, including DNA methylation, histone modifications, and three-dimensional (3D) genome organization, and these layers collectively shape cell identity.

To investigate how these layers contribute to tumor maintenance, we simultaneously profiled the 3D genome and DNA methylation from two tumor types using snm3C-seq. Leveraging both modalities, we inferred structural variants, identified intratumoral genetic heterogeneity, and defined clones within each sample. DNA methylation-based embedding and clustering stratified malignant cells into groups concordant with structural variant-defined clones. Differential methylation analysis across these clones uncovered distinct partially methylated domain patterns, probably linked to variation in proliferative states among malignant cells.

3D genome-based embedding, in turn, revealed concordant heterogeneity, identifying clone-specific 3D genome features. Understanding the role of these layers and their coordination in initiating and maintaining malignancy in these lethal tumor types may enable the development of novel therapeutic targets.

B-G.34: ROS-Driven Somatic Mutation Landscape and Tumorigenesis in Prx1 Knockout Mice
Track: Genomics, epigenomics, and genome editing
  • Yukyung Jun, Division of National Supercomputing, Korea Institute of Science and Technology Information, Daejeon 34141, Korea, South Korea
  • Jiheon Shin, Department of Life Science, Ewha Womans University, Seoul 03760, Korea, South Korea
  • Sang Won Kang, Department of Life Science, Ewha Womans University, Seoul 03760, Korea, South Korea
  • Sanghyuk Lee, Department of Life Science, Ewha Womans University, Seoul 03760, Korea, South Korea


Presentation Overview: Show

Peroxiredoxin 1 (Prx1) is a key antioxidant enzyme that maintains redox homeostasis by eliminating reactive oxygen species (ROS). Prx1-deficient mice exhibit reduced lifespan, hemolytic anemia, and age-dependent tumor formation, suggesting a role of ROS in tumorigenesis. To investigate the impact of chronic oxidative stress on somatic mutations, we performed whole-exome sequencing of liver and spleen tissues from 15-month-old Prx1 knockout mice. Primary cells from these mice showed elevated nuclear ROS levels and increased DNA damage, indicating ROS-driven genomic instability. Gene ontology analysis revealed enrichment in aging-associated pathways, including DNA damage response and cognitive processes, as well as oxidoreductase activity and DNA binding functions. Mutational signature analysis showed an age-related increase in SBS40, a signature linked to aging and human cancers. Comparative analysis of shared mutation profiles identified 16 candidate genes. Among them, Nek4 harbored an age-dependent stop-gain mutation, suggesting its potential role in ROS-associated tumorigenesis. These results characterize the mutational landscape under chronic oxidative stress and provide insight into the link between ROS accumulation, aging-related mutations, and cancer development.

B-G.35: tfClone: Accurate Determination of Haplotype Specific Clonal Copy Number Profiles in Circulating Tumour DNA
Track: Genomics, epigenomics, and genome editing
  • Emilia Hurtado, The University of British Columbia, Canada
  • Andrew Roth, University of British Columbia, Canada


Presentation Overview: Show

In cancer, tumour heterogeneity is driven by mutations that result in the development of genetically distinct subpopulations of cells called clones. These clones may respond differentially to treatment, leading to selective survivorship and proliferation of treatment resistant cell populations. Copy number variation and copy number variants (CNVs) represent a significant source of signal in separating and defining clonal profiles, and as such several methods have been developed to infer clonal copy number profiles from both bulk and single-cell whole genome sequencing. Recent tissue sequencing methods have demonstrated the potential of integrating phasing information when inferring clonal copy number and clonal prevalence. We propose tfClone (tissue free clone), a computational method that uses a phase-informed Bayesian hierarchical model to characterise the clonal composition and copy number profiles of a cancer from circulating tumour DNA. tfClone models clonal copy number profiles using a factorial hidden Markov model, where each hidden chain is used to infer the copy number profile of a single clone, while per-sample clonal prevalence is modeled as a latent variable for clonal mixing proportions. Inference is performed using a combination of Gibbs sampling, forward-filtering backwards sampling, slice sampling, and adaptive MCMC. We demonstrate the performance of tfClone's improved sensitivity in tumour fraction and clonal prevalence detection using both synthetic and real datasets.

B-G.36: Advancing copy-number phylogenetics through updates to MEDICC2
Track: Genomics, epigenomics, and genome editing
  • Chenxi Nie, Institute for Computational Cancer Biology, Germany
  • Tom L. Kaufmann, Institute for Computational Cancer Biology, CIO, CCCE, Faculty of Medicine and University Hospital Cologne, Germany, Germany
  • Alexander Nicolay, Institute for Computational Cancer Biology, University Hospital Cologne, Germany
  • Roland F. Schwarz, Cancer Research Center Cologne Essen (CCCE), University Hospital and University of Cologne, Germany


Presentation Overview: Show

Somatic copy-number alterations (SCNAs) and chromosomal instability are ubiquitous in cancer and contribute to genome plasticity and intratumor heterogeneity. Accurate phylogenetic reconstruction from SCNA data is vital for understanding tumor progression, yet phylogenetic inference from SCNAs remains challenging: their large size and propensity for overlapping events violate the infinite sites assumption, rendering many standard phylogenetic approaches unsuitable.
MEDICC2 addresses these challenges through the Minimum Event Distance (MED) framework, which uses Finite State Transducers (FSTs) to model copy-number evolution without the independent bin assumption. Here, we present algorithmic and methodological advances to MEDICC2 that improve both accuracy and computational efficiency. Specifically, we introduce an optimized MED calculation algorithm with improved runtime performance, a Metropolis-Hastings MCMC scheme for tree topology exploration, and redesigned FSTs that incorporate event length for improved ancestral reconstruction.
We evaluate these improvements through a comprehensive benchmarking pipeline comparing our improvements against the original MEDICC2 version across simulated datasets spanning a range of tree sizes, CNA overlap levels, and mutation rates. Our results demonstrate notable gains in reconstruction accuracy, particularly in challenging settings with highly overlapped CNA events, as well as substantial improvements in runtime performance. Together, these advances further establish MEDICC2 as the leading toolkit for phylogenetic inference from SCNA data, offering a powerful and more efficient tool for studying tumor evolution.

B-G.37: W-ASAP: Rapid Interactive Wastewater-based Viral Variant Detection
Track: Genomics, epigenomics, and genome editing
  • Alexander Taepper, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
  • Gordon Koehn, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
  • Felix Hennig, Independent, Switzerland
  • Ivan Topolsky, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
  • Chaoran Chen, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
  • Tanja Stadler, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
  • Niko Beerenwinkel, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland


Presentation Overview: Show

Wastewater-based genomic surveillance enables population-level monitoring of viral diversity, with its utility for early variant detection and tracking resistance mutations having been demonstrated in various studies. However, the large size of the sequencing data have made analysis workflows technically demanding and slow. Existing dashboards and reports typically address a limited set of predefined questions but do not allow users to explore the data independently.

The W-ASAP platform addresses these limitations by providing an interactive, browser-based system for real-time exploration of read-level wastewater sequencing data. Built on the genomic query engines LAPIS and SILO, which index hundreds of millions of reads, it enables querying of large datasets within milliseconds. This allows rapid identification and analysis of emerging variants, reducing turnaround times from days to minutes. The platform currently supports SARS-CoV-2 and RSV, providing public dashboards and APIs.

B-G.38: Automatic reanalysis of genetic data from patients with Primary Ciliary Dyskinesia
Track: Genomics, epigenomics, and genome editing
  • Anna-Lena Katzke, Department of Human Genetics, Hannover Medical School, Hannover, Germany
  • Paul Siek, Department of Human Genetics, Hannover Medical School, Hannover, Germany
  • Martin Wetzke, Department of Pediatrics, Pediatric Pulmonology, Allergology and Neonatology, Hannover Medical School, Hannover, Germany
  • Felix C. Ringshausen, Department of Respiratory Medicine and Infectious Diseases, Hannover Medical School, Hannover, Germany
  • Ben Ole Staar, Hannover; Department of Respiratory Medicine and Infectious Diseases, Hannover Medical School, Hannover, Germany
  • Gunnar Schmidt, Deparment of Human Genetics, Hannover Medical School, Hannover, Germany
  • Bernd Auber, Deparment of Human Genetics, Hannover Medical School, Hannover, Germany
  • Sandra V. Hardenberg, Deparment of Human Genetics, Hannover Medical School, Hannover, Germany


Presentation Overview: Show

Background:
Primary ciliary dyskinesia (PCD) is characterized by motile cilia dysfunction leading to chronic lung disease. PCD is inherited predominantly in an autosomal recessive manner, caused by pathogenic variants in more than 50 genes. 20-30% of patients with a well-defined PCD phenotype remain without molecular diagnosis, highlighting the need for periodic reanalysis. As variant interpretation has become the bottleneck in genetic diagnostics, we aim to investigate automated variant classification capability for the reanalysis of PCD patients.
Material and Methods:
A validation (n=30), a reanalysis (n=159), and a first-analysis cohort (n=66) of patients with a PCD phenotype were analyzed using the automated classification tool HerediClassify, screening 908 genes. Based on clinical records, phenotypic features were assessed using the PICADAR score (PrImary CiliARy DyskinesiA Rule) in children and the modified PICADAR score in adults. To all PICADAR-positive cases, the PP4 criterion for disease-specific phenotype was applied by HerediClassify.
Results:
HerediClassify showed a sensitivity of 0.67 and specificity of 1 in the validation cohort. In the reanalysis cohort, which included patients who had remained without a genetic diagnosis following routine diagnostic workup, six additional patients were genetically diagnosed. 71% of patients with a manual diagnosis were correctly identified by HerediClassify in the first-analysis cohort.
Conclusion:
Automated variant classification can enhance reanalysis efficiently by identifying high-priority variants. The systematic inclusion of phenotype-specific information, such as the PICADAR score, substantially improves the interpretation of genetic data and should be considered standard practice.

B-G.39: Optimizing genetic association tests for small cohorts
Track: Genomics, epigenomics, and genome editing
  • Piotr Suszyński, Warsaw University of Technology, Poland
  • Tomasz Gambin, Warsaw University of Technology, Poland


Presentation Overview: Show

Running genetic association tests on small cohorts, which is common for rare diseases, present unique challenges. We have to be especially careful during data filtering stages, commonly executed before actual tests. It also makes it more challenging to correct for population stratification. For such scenarios, when parameter values selected using researcher experience and generally recommended default values are not precise enough, we propose a data-driven approach to select optimal settings, maximizing efficacy and providing trust in the results. We built a comprehensive, high performance Nextflow workflow for genetic association testing, which includes all commonly performed data filtering and preparation steps, controlled by a rich set of modifiable parameters. We also built another workflow, which generates synthetic testing datasets with realistic characteristics, and a versatile tool that orchestrates the execution of mentioned workflows for the purpose of optimizing their parameters. We have run our software on Thousand Genomes project data, executing more than 7000 association testing workflow runs with various parameters and datasets, and we obtained an optimized set of parameters. Results proved that the parameters are highly interdependent and rules for selecting them in isolation fail to capture that complexity. We also found that using dosage is superior to traditional hard genotypes association tests. To enhance the interpretability of our results we applied the Accumulated Local Effects XAI technique. Our tools can also be used to tune the parameters to an individual dataset characteristics, enabling robust association testing for small cohorts.

B-G.40: Transposable element diversity and its potential role in symbiosis regulation explored using new chromosome-scale assemblies of Medicago truncatula ecotypes
Track: Genomics, epigenomics, and genome editing
  • Paulina Poniatowska-Rynkiewicz, Institute of Bioorganic Chemistry PAS, Poland
  • Paweł Wojciechowski, Poznań University of Technology, Poland
  • Agnieszka Żmieńko, Institute of Bioorganic Chemistry PAS, Poland


Presentation Overview: Show

Legumes play a central ecological and agricultural role due to their ability to establish symbiotic nitrogen fixation with rhizobia. Medicago truncatula is a widely used model for studying legume biology, including nodulation and symbiotic genome regulation. To expand genomic resources for this model legume, we generated chromosome-scale assemblies for three additional geographically distinct M. truncatula ecotypes using long-read sequencing technologies. By literally doubling the number of available high-quality assemblies for this species, we establish a comparative genomic resource for investigating its structural variation and repetitive sequence diversity. Using a multi-stage de novo transposable element (TE) annotation pipeline we curated accession-specific TE libraries and characterized TE landscapes across all assemblies. Importantly, we investigate how TE diversity intersects with genes involved in symbiotic nitrogen fixation. In M. truncatula, nodulation involves activation of gene clusters located within symbiotic islands, regions known to undergo dynamic epigenetic remodeling. Previous studies have shown that these regions undergo DNA demethylation during nodule development, while adjacent TEs may become transiently activated before being re-silenced in the mature nodules. We aim to establish whether natural variation in TE copy number and genomic positioning may influence the epigenetic landscape and transcriptional regulation of symbiosis-related genes.
Our ongoing comparative analyses focus on TE diversity both genome-wide and within symbiotic islands, aiming to link structural variation with potential regulatory consequences. These newly generated assemblies and curated TE annotations establish a foundation for studying TE-driven genome evolution, epigenetic regulation, and environmental adaptation in legumes.

B-G.41: DNA damage assessment from panel sequencing data and its use in the study of survivability
Track: Genomics, epigenomics, and genome editing
  • Thomas Minotto, University of Montpellier, France
  • Luka Pavageau, Institut Universitaire du Cancer de Toulouse, France
  • Mehmet Samur, Dana-Farber Cancer Institute, United States
  • Jill Corre, Institut Universitaire du Cancer de Toulouse, France
  • Sophie Lebre, University of Montpellier, France
  • Alice Cleynen, University of Montpellier, France


Presentation Overview: Show

Assessing DNA damage from sequencing data is crucial for the diagnosis of cancer patients. Several computational tools have been developed for that purpose. One of them is the Genomic Scar Score (GSS), which evaluates copy number variations in whole genome sequencing data via a segmentation algorithm to retrieve gains and deletions. This allows the detection of losses of heterozygosity, large-scale transitions, and telomeric allelic imbalances, which are summed up into a final score value. In multiple myeloma, a low GSS has been associated with superior outcome for patients. Yet whole genome sequencing is costly to scale up and not widely used by practitioners.

Here, we adapt existing tools for GSS computation to targeted sequencing (panel) data, a sequencing technique which is more accessible. We optimize the tools FACETS, CNVkit and PureCN, on a clinical dataset to compensate for the lower information density, and we also study how the new score correlates with survival information.

We find that most DNA damage is still detectable from panel data, including gain and deletion events specific to multiple myeloma, provided that the coverage in these locations is sufficient. However, some events are missed when genomic coverage is too low, which happens in our data with telomeric allelic imbalances. Resulting copy number profiles can be used to predict patient survival. This new usage of panel data confirms its efficacy for routine patient diagnosis and offers guidance on optimal probes positioning and coverage strategies for adequate screening of significant genomic regions.

B-G.42: Genomic Instability Pattern Analysis in Advanced Stages of High-Grade Serous Ovarian Cancer
Track: Genomics, epigenomics, and genome editing
  • Sara Potente, Department of Biology, Unversity of Padova, Italy, Italy
  • Luca Beltrame, IRCCS Humanitas Research Hospital, Milan, Italy
  • Sonia Ismari, IRCCS Humanitas Research Hospital, Milan, Italy
  • Federica Cetti, IRCCS Humanitas Research Hospital, Milan, Italy
  • Lara Paracchini, IRCCS Humanitas Research Hospital, Milan, Italy
  • Maurizio D'Incalci, IRCCS Humanitas Research Hospital, Milan, Italy
  • Sergio Marchini, IRCCS Humanitas Research Hospital, Milan, Italy
  • Chiara Romualdi, Department of Biology, University of Padova, Italy


Presentation Overview: Show

High-grade serous ovarian cancer (HGSOC) is the most common and lethal ovarian cancer histotype, with most patients diagnosed at advanced stages (III–IV) and a 5-year survival rate below 30%. It is characterized by near-universal TP53 mutations, defects in homologous recombination (HR) DNA repair, and pervasive somatic copy number alterations (SCNA), which collectively contribute to extensive genomic rearrangements. However, the landscape of chromosomal instability processes in advanced-stage disease remains poorly characterized, and whether instability subgroups identified in early-stage HGSOC are conserved across stages and tumor sites is unknown.
In this preliminary study, 116 formalin-fixed paraffin-embedded (FFPE) samples from 40 Stage III–IV HGSOC patients were analyzed, including primary tumors and matched metastatic lesions from ovary, omentum, peritoneum, and other sites. Samples were profiled by shallow whole-genome sequencing (sWGS), and SCNA profiles and chromosomal instability (CIN) signatures were derived using the SAMURAI bioinformatics pipeline. Each sample was classified as unstable (U) or highly unstable (HU) based on the fraction of altered genome and number of breakpoints.
Classification identified 63 HU and 53 U samples. Clustering revealed distinct subgroups in both primary and metastatic lesions across instability classes. CX1 and CX3 signatures — related to chromosome missegregation and replication stress, respectively — were dominant across all clusters. The overall instability landscape in advanced HGSOC closely resembles that of early-stage disease, suggesting conserved processes across stages. Notably, 13 patients showed a shift in instability pattern between primary and metastatic sites, supporting dynamic genomic evolution during metastatic progression.

B-G.43: Phylogenetic tree inference from single-cell RNA sequencing data
Track: Genomics, epigenomics, and genome editing
  • Norio Zimmermann, Department of Biosystems Science and Engineering, ETH Zurich, 4056 Basel, Switzerland, Switzerland
  • Xiaoyu Sun, None, Germany
  • Joanna Hård, KTH Royal Institute of Technology, 100 44 Stockholm, Sweden, Sweden
  • Jack Kuipers, ETH Zurich, D-BSSE, Computational Biology Group, Switzerland
  • Niko Beerenwinkel, ETH Zurich, Switzerland


Presentation Overview: Show

Single-cell RNA sequencing technologies enable the large-scale measurement of gene expression profiles at the individual cell level to assess cellular diversity and function. In oncology, leveraging these single-cell transcriptomic data to reconstruct the phylogenetic relationships among cancer cells can provide insights into tumor evolution, metastasis formation, and the development of treatment resistance. However, phylogenetic inference from single-cell RNA sequencing is challenging due to sparse and noisy data and large dataset sizes. We present a novel tree inference method designed for such data that takes reference and alternative read counts of single-nucleotide variants and reconstructs a phylogenetic tree of the sequenced cells via maximum likelihood using a random-scan greedy search.
To overcome local optima in the search, our algorithm alternates between two different tree representations: cell lineage trees, where cells are represented by nodes and mutations are attached to edges, and mutation trees, where mutation nodes encode the mutational events and cells are attached to them. Because a local optimum in one tree space generally does not correspond to a local optimum in the other space, we maximize the likelihood by switching between the two tree spaces until convergence is achieved in both. We demonstrate superior performance on simulated data compared to existing methods. Furthermore, we show the applicability of our approach to cancer single-cell RNA sequencing data, where it allows us to link evolutionary trajectories of cells to their gene expression profiles.

B-G.44: An Integrated Framework for Multi-Omics Long-Read Analysis at Single-Molecule Resolution
Track: Genomics, epigenomics, and genome editing
  • Isabell Wienpahl, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany, Germany
  • Henrik Köppke, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany; Humboldt-Universität zu Berlin, Germany
  • Scott Lacadie, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany, Germany
  • Roseen Musallam, Humboldt-Universität zu Berlin; Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany, Germany
  • Uwe Ohler, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany; Humboldt-Universität zu Berlin, Germany


Presentation Overview: Show

The epigenetic landscapes of genes define their transcriptional activity. Multiple epigenetic marks crosstalk with each other and consequently shape cell fate in development and disease. We have developed a multi-omics assay that jointly measures chromatin accessibility, DNA methylation and 3D genome organization on the /same/ DNA molecule. Combined with third-generation sequencing, this approach preserves higher-order concatemers and thus enables us to capture multi-way chromatin contacts in addition to pairwise contacts. However, current software typically addresses either long-read epigenomics or concatemer-based 3D genome analysis, and no comprehensive solution exists for an integrated interpretation of these readouts at single-molecule resolution. We therefore aim to develop a modular toolbox that addresses all modalities of multi-omics long read information. The toolbox is designed to analyze the full assay output within one workflow and includes reference-based deconvolution for mixed samples, enabling haplotype-aware assignment of molecules and cell-type deconvolution from informative loci. By unifying single-molecule epigenetic and 3D genome analysis, our framework enables systematic characterization of epigenetic states together with their spatial context. This work provides a computational foundation for studying allele-specific regulation, cellular heterogeneity and genome organization from native long-read multi-omics data.

B-G.45: A reproducible workflow for HEK293 methylome profiling during adaptation to suspension growth
Track: Genomics, epigenomics, and genome editing
  • Alexander Karl Molin, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess Science and Engineering, Austria
  • Nikolaus Virgolini, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess Science and Engineering, Austria
  • Georg Smesnik, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess Science and Engineering, Austria
  • Astrid Duerauer, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess Science and Engineering, Austria
  • Nicole Borth, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess Science and Engineering, Austria


Presentation Overview: Show

Biomanufacturing for transient protein expression and recombinant Adeno-associated virus (rAAV) production often relies on empirical process development. Transitioning towards knowledge-driven optimization requires a fundamental understanding of host cell biology, which in the case of Human Embryonic Kidney 293 (HEK293) cells, the main expression system for rAAV production, is still limited. As such, this study presents a computational workflow for in-depth characterization of the epigenetic landscape of HEK293 cells and will be used in further investigations.
An adherent HEK293 cell line was adapted to suspension growth using four distinct commercial media formulations, with HEK293-6E serving as a commercially available suspension-adapted reference. Post-adaptation, DNA samples were taken and whole-methylome sequencing was performed using an enzymatic methyl-conversion kit and Illumina sequencing (150 bp paired-end, 37x coverage). The resulting data was evaluated, implementing a reproducible Snakemake-based processing and analysis pipeline. Raw data was processed and aligned to the human reference genome augmented with the Adenovirus5 region, prior to methylation calling via the Bismark program.
Next a downstream analysis pipeline was implemented, utilizing MethylKit and custom scripts to identify conserved methylomic signatures associated with suspension growth. To assess cell regulation, a chromatin state model was generated integrating publicly available datasets with ChromHMM. Mapping the experimental data against this model revealed characteristic mammalian patterns, specifically strong methylation within transcribed regions and lower levels at active transcription start sites and promoters. This work establishes a reproducible data analysis framework that provides critical insights into the HEK293 epigenome, offering a foundation for targeted host cell engineering.

B-G.46: SNooPy: a statistical framework for long-read metagenomic variant calling
Track: Genomics, epigenomics, and genome editing
  • Roland Faure, Institut Pasteur, France
  • Ulysse Faure, Dep. of Mathematics, ETH Zürich, Switzerland, Switzerland
  • Tam Truong, University Rennes, Inria, CNRS, IRISA - UMR 6074, Rennes, France,, France
  • Alessandro Derzelle, Service Evolution Biologique et Ecologie, Universit´e libre de Bruxelles (ULB), Brussels, Belgium, Belgium
  • Dominique Lavenier, University Rennes, Inria, CNRS, IRISA - UMR 6074, Rennes, France,, France
  • Jean-François Flot, Service Evolution Biologique et Ecologie, Universit´e libre de Bruxelles (ULB), Brussels, Belgium, Belgium
  • Christopher Quince, Organisms and Ecosystems, Earlham Institute, Norwich, UK, United Kingdom


Presentation Overview: Show

Current long-read single-nucleotide variant callers were designed primarily for genomic data—particularly human genomes. While some have been used on metagenomic data, their underlying assumptions and training procedures fail to account for the inherent complexity of metagenomic samples. To date, no long-read variant caller has been purpose-built for metagenomic applications. To address this gap, we present SNooPy, a SNP-calling tool that implements a new statistical framework tailored to long-read metagenomic data. Unlike previous genomic methods, our approach makes no assumptions about the number of haplotypes present, their evolutionary relationships, or their sequence divergence. We demonstrate that SNooPy outperforms both traditional statistical and deep learning–based SNP callers. Our results suggest that future integration of this framework with deep learning approaches could further enhance variant calling performance. SNooPy is freely available on github.com/rolandfaure/snoopy.

B-G.47: Impact of Sperm Methylome on In Vitro Produced Embryo Methylome, Transcriptome and Development
Track: Genomics, epigenomics, and genome editing
  • Pratik P. Pathade, Aarhus University, Denmark
  • Aurélie Bonnet, ELIANCE, France
  • Aurélie Chaulot Talmon, INRAE, France
  • Valentin Costes, ELIANCE, France
  • Marie-Christine Deloche, ELIANCE, France
  • Anne Frambourg, INRAE, France
  • Udayaraja Gk, Aarhus University, Denmark
  • Laurent Schibler, INRAE, France
  • Hélène Kiefer, INRAE, France
  • Véronique Duranthon, INRAE, France
  • Haja N. Kadarmideen, Aarhus University, Denmark


Presentation Overview: Show

Elite bull semen is widely used in cattle artificial insemination and reproductive technologies. Beyond genetic variation, sperm epigenetic changes can be transmitted to embryos, potentially affecting embryo methylome, tran-scriptome, development, and calf phenotypic performance. This study evaluated the impact of sperm DNA methyl-ation on in vitro produced (IVP) embryo methylome and transcriptome using a paired experimental design. Ten replicates used oocyte pools from five cows fertilized with sperm from two bulls with extreme CpG methylation differences across ~115,000 sites, generating a total of 20 pools of Day-7 expanded blastocysts (10 pools per bull) for RRBS and RNA-seq analysis (ARS-UCD2.0).
RRBS identified 5,077 hypermethylated and 4,929 hypomethylated DMCs, predominantly in intronic regions, followed by exonic, intergenic, and promoter regions. GO enrichment revealed hypermethylated DMCs enriched for tissue development, neurogenesis, and DNA-binding transcription factors, while hypomethylated DMCs were associated with transcription regulation and RNA biosynthesis. RNA-seq detected 38,001 genes, yielding 44 dif-ferentially expressed genes (FDR<0.05). Weighted Gene Co-expression Networks (WGCNA) with top 25% varia-ble genes and power=18, R²≥0.70, identified 29 modules, three of which were significantly correlated with bull origin (|r|≥0.60, adj.P≤0.05).
These findings suggest that contrasting sperm methylomes are associated with broad embryo epigenetic repro-gramming, affecting both individual gene expression and co-expression networks, establishing a framework for studying paternal epigenetic transmission in cattle IVP embryos.

B-G.48: Dysregulating genomes - mapping how breaking DNA-lamina interactions drives muscular dystrophy
Track: Genomics, epigenomics, and genome editing
  • Anna Antonatou Papaioannou Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany, Germany
  • Julia Torres Rivera, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany, Germany
  • Narasimha Swamy Telugu, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany, Germany
  • Ines Lahmann, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany, Germany
  • Werner Stenzel, Department of Neuropathology, Charité-Universitätsmedizin, Berlin, Germany, Germany
  • Sebastian Diecke, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany, Germany
  • Mina Gouti, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany, Germany
  • Michael Ian Robson, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin, Germany, Germany


Presentation Overview: Show

A key question in genome biology is how 3D genome organization regulate gene expression in health and disease. Central to this organization is the nuclear lamina, which tethers compacted heterochromatin to the nuclear periphery in lamina-associated domains (LADs) that safeguard gene repression and cell identity. Mutations in lamina-tethering proteins that disrupt LAD organization are proposed to cause diverse, untreatable laminopathy diseases, including Congenital Muscular Dystrophy (CMD). However, how LAD disruption impacts heterochromatin and gene regulation to drive pathology remains unknown. Here, we address this by combining high-throughput electron microscopy (EM) and single-cell multiomics to systematically profile heterochromatin defects and their functional consequences in CMD. EM imaging of hundreds of muscle nuclei reveals heterochromatin undergoes extensive decondensation at the nuclear periphery in CMD. However, heterochromatin also unexpectedly decondenses in the nuclear interior, suggesting that lamina-tethering is required to globally maintain heterochromatin function throughout the nucleus. To test this, we have established a robust experimental and computational pipeline for joint single-cell epigenome and transcriptome profiling in tissues. With this, we are now mapping loci and cell types where lamina-interactions and heterochromatin states are disrupted, and when these disruptions cause ectopic gene activation. Combined, this will reveal mechanisms driving laminopathies like CMD, revealing both the functions of lamina-directed genome organization and potential targets for therapeutic intervention. More generally, our combed EM and single-cell toolkit provides a generalizable framework to study epigenome dysregulation in complex tissue in other disease contexts.

B-G.49: A lesson of grammar: unraveling Transcription Factors DNA-binding syntax rules in the model plant Arabidopsis thaliana
Track: Genomics, epigenomics, and genome editing
  • Alice Jegou, CEA - LPCV, France
  • Jérémy Lucas, CEA - LPCV, France
  • Marianne Dreuillet, CEA - LPCV, France
  • François Parcy, CEA - LPCV, France
  • Romain Blanc-Mathieu, CEA - LPCV, France


Presentation Overview: Show

Transcription factors (TFs) are master regulators of gene expression in eukaryotes, binding to specific DNA motifs within cis-regulatory regions. Yet, beyond motif recognition, the parameters that shape TF-DNA binding remain poorly understood, especially in plants. Studies suggest that TFs motifs follow grammar rules: spatial arrangements that dictate binding and thus their spatial occupation. While such rules have been described for specific families like Auxin Response Factors or MADS-TF, resulting in the formation of regulatory protein complexes, their prevalence remains not fully explored.
In the model plant Arabidopsis thaliana, we classified 1,627 TFs into 56 families based on their DNA-binding domains (DBDs). Then we built a comprehensive atlas of TFs binding sites (TFBS) in this model plant, integrating 681 genome-wide binding experiments (ChIP-seq, DAP-seq, ampDAP-seq) spanning 40 TF families. By modeling these TFBSs, we uncovered that syntax rules are not exceptions but a widespread phenomenon: nearly every family exhibits unique, family-dependent homotypic grammar. Strikingly, comparing in vitro and in vivo conditions revealed dynamic shifts in these rules, pointing to complex regulatory mechanisms.
These findings underscore the critical role of TFs binding grammar and suggest the formation of higher-order TFs complexes or DNA mediated cooperative interactions. Investigating further, we would like to combine heteromeric syntax analysis with single-cell RNA-seq and chromatin accessibility data to predict potential TF complexes in silico.
This work provides an essential resource to the plant biology community to decode the regulatory language of TFs.

B-G.50: Epigenomic Rewiring of Enterocyte Chromatin During Post natal Salmonella Typhimurium Infection
Track: Genomics, epigenomics, and genome editing
  • Caterina Alfano, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
  • Anna-Lena Ullrich, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
  • Johannes Schöneich, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
  • Matthias A. Schmitz, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
  • Samuel Ward, Institute of Experimental and Clinical Pharmacology and Toxicology, University of Freiburg, Germany
  • Sebastian Preissl, Institute of Experimental and Clinical Pharmacology and Toxicology, University of Freiburg and I.P.W., Graz University, Germany
  • Aline Dupont, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
  • Mathias W. Hornef, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany


Presentation Overview: Show

Salmonella Typhimurium (STm) is a food-borne pathogen that causes salmonellosis. Following ingestion, STm reaches the gastrointestinal tract, where it attaches and invades the intestinal epithelium allowing intracellular bacterial replication and colonization. The interaction with the intestinal epithelium is thus a crucial step for the establishment of the infection. Our goal is to uncover epigenetic modifications in the intestinal epithelium induced by STm infection during postnatal development, a stage in which the epithelium matures to fulfill the changing requirements in digestion, absorption, antimicrobial activity and barrier formation. For this purpose, we generated and analyzed scATAC-seq profiles of 5-day-old mice (two controls and two samples inoculated on day-1). By focusing on enterocytes, the primary target cells of STm, we identified more than 15.000 differentially accessible regions, ~25% of which are annotated as promoters. Regions more open in the infected samples are enriched in negative regulation of T cells activation and display a strong enrichment of AP 1 family motifs, known to play a role in the host's immune response to bacterial infection. Finally, network-based co-accessibility analysis revealed striking differences: regions' co-accessibility is extremely more frequent in infected samples, especially between regions annotated to different genes. Large communities of co accessible regions are unique to the infected network, and are mostly made of promoters, signaling an infection driven reorganization of transcriptional regulatory mechanisms. Our findings highlight a wide-spread epigenetic reprogramming of the intestinal epithelium driven by the interaction with the pathogen. Integration with scRNA-seq experimental data will allow further characterizations of these processes.

B-G.51: Detection of haplotype-specific contacts in phased chromatin contact maps with Genome Architecture Mapping
Track: Genomics, epigenomics, and genome editing
  • Claudia Robens, Institute for Computational Cancer Biology (ICCB), University Hospital Cologne, University of Cologne, Germany, Germany
  • Julia Markowski, Max-Delbrück-Center for Molecular Medicine, Berlin Institute for Medical Systems Biology, Humboldt University of Berlin, Germany
  • Alexander Kukalev, Max-Delbrück-Center for Molecular Medicine, Berlin Institute for Medical Systems Biology, Germany, Germany
  • Christoph J. Thieme, Max-Delbrück-Center for Molecular Medicine, Berlin Institute for Medical Systems Biology, Germany, Germany
  • Adam Streck, Institute for Computational Cancer Biology (ICCB), University Hospital Cologne, University of Cologne, Germany, Germany
  • Ana Pombo, Max-Delbrück-Center for Molecular Medicine, Berlin, Germany; Johns Hopkins University, Baltimore, USA, Germany
  • Roland F. Schwarz, Institute for Computational Cancer Biology (ICCB), University Hospital Cologne, University of Cologne, Germany, Germany


Presentation Overview: Show

Chromatin structure is key to the orchestration of gene regulation and its misfolding is involved in many diseases including developmental diseases and cancer. Genome Architecture Mapping (GAM) is a ligation-free method for determining chromatin conformation from minimal input material. Understanding allele-specific regulation requires haplotype-resolved chromatin organization, which depends on direct phasing, where sequencing reads are assigned to haplotypes based on single-nucleotide variants. The sparse variant density in the human genome poses a challenge to direct phasing efficiency and thus limits the generation of haplotype-specific chromatin contact maps.
We here present our novel read phasing strategy ‘CoPhasing', which leverages local haplotype information of GAM data to correctly assign sequencing reads to their haplotype of origin, even without overlapping heterozygous variants. CoPhasing allows for haplotype-specific analysis of chromatin folding and detection of haplotype-specific chromatin contacts in variant-sparse genomes. We introduce a permutation test-based algorithm for detecting significant haplotype-specific contacts in cophased GAM data. The method marks haplotype-specific interactions in contact maps and indicates which of the two homologous chromosomes exhibits stronger interactions. By leveraging differential contacts identified through the permutation test, we detect genomic regions enriched for haplotype-specific interactions, highlighting loci with pronounced allelic differences in chromatin organization.
Together, our algorithms will enable novel insights into 3D chromatin architecture in health and disease, adding to the shifting focus beyond genomic mutations in coding regions towards epigenetic gene-regulatory mechanisms.

B-G.52: Somatic variant detection using a personalized pangenome reference: a study of a paediatric acute myeloid leukaemia patient
Track: Genomics, epigenomics, and genome editing
  • Elizaveta Kulaeva, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Netherlands
  • Mark van Roosmalen, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Netherlands
  • Rico Hagelaar, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Netherlands
  • Eline Bertrums, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Netherlands
  • Andrea Guarracino, Bioinnovation and Genome Sciences, Translational Genomics Research Institute (TGen), part of City of Hope, United States
  • Ruben van Boxtel, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Utrecht, The Netherlands, Netherlands


Presentation Overview: Show

Current clinical pipelines using linear reference genomes cause reads carrying mutations to be misaligned or lost, which is particularly problematic for low-coverage assays like single-cell whole-genome sequencing (scWGS), where technical noise can interfere with true somatic events. To overcome this, we applied a haplotype-resolved pangenome reference to a paediatric acute myeloid leukaemia (pAML) case and assessed the clinical utility of this approach.

We constructed a personalized pangenome from long-read Nanopore sequencing of healthy haematopoietic stem and progenitor cells, incorporating two haplotype-resolved assemblies with Minigraph-Cactus. This added 36.3 Mb of novel sequence, 42% of which represented variation absent from the standard hg38 reference. We compared somatic variant calling across bulk, clonally expanded, and single-cell WGS data of healthy and cancerous cells. Reference bias in pangenome alignments was quantified using Biastools and found to be significantly reduced compared to linear alignments.

The personalized pangenome reduced reads with mates mapped to different chromosomes up to 8,000-fold which allowed more accurate split-read event detection and remapping of previously ambiguous alignments. It also rescued SNVs previously filtered out by mapping quality thresholds and eliminated up to 80% of short indels in low-complexity regions. Driver analysis revealed that frameshift insertion artifacts in scWGS data were largely removed, and the medically relevant and computationally challenging PRSS1 locus was substantially disentangled due to reduced misalignment.

Our findings provide systematic evidence that personalized pangenome improves somatic variant detection in cancer genomics with the most benefit for scWGS data and complex genomic regions previously misaligned due to linear reference bias.

B-G.53: Prioritizing non-coding variants to discover mutations governing agronomic traits in plants
Track: Genomics, epigenomics, and genome editing
  • Mathis Pochon, CEA, France
  • Romain Blanc-Mathieu, CEA, France
  • Francois Parcy, CNRS, France


Presentation Overview: Show

During the development of agriculture over centuries, humans have selected plants with advantageous agronomic traits. This process has led to genomic diversification between varieties within a species. A striking example is Brassica oleracea, a highly diversified species that includes cabbages, cauliflowers, and broccoli.
Association studies often highlight cis-regulatory regions as targets of selection. However, identifying causal mutations remains a major challenge, as genomic regions associated with a trait can harbor thousands of variants in strong linkage disequilibrium.
To address this challenge, we developed a tool dedicated to prioritizing non-coding variants (NCVs) with potential regulatory functions. This tool is based on a framework originally developed for human disease studies and distinguishes likely functional variants from non-functional ones by leveraging transcription factor binding sites (TFBS) and modules of TFBS that show higher-than-expected phylogenetic conservation.
As a case study, we applied this tool to curd formation in cauliflower, a trait known to result from deregulation of the floral development gene regulatory network. Using kmer variants from several accessions of B. oleracea, we first performed an association study that highlighted both coding and non-coding regions. We are now using our tool to specifically prioritize NCVs potentially involved in curd formation.
These results with our current knowledge of transcription factors and the gene regulatory network underlying floral development will allow us to identify candidate causal mutations responsible for cauliflower head formation, which can then be experimentally validated. Beyond this example, our tool provides a general framework to identify causal NCVs associated with other agronomic traits.

B-G.54: Plasma cfDNA fragmentomics provides insight into Preeclampsia pathophysiology
Track: Genomics, epigenomics, and genome editing
  • Irene D'Onofrio, University of Zurich, Switzerland
  • Zsolt Balázs, University of Zurich, Switzerland
  • Elisabetta Ranieri, University Hospital of Zurich, Switzerland
  • Todor Gitchev, University of Zurich, Switzerland
  • Michael Krauthammer, University of Zurich, Switzerland
  • Tilo Burkhardt, University Hospital of Zurich, Switzerland


Presentation Overview: Show

Introduction
Preeclampsia (PE) is a multisystem hypertensive disorder of pregnancy characterized by placental dysfunction, maternal endothelial and end-organ injury. Despite affecting 3-8% pregnancies, PE etiology remains incompletely understood, partly due to limited access to placental tissue during pregnancy. However, cell-free DNA is emerging as a non-invasive biomarker for PE, with potential to provide insights into its pathophysiology.
Methods
Plasma cfDNA from 22 pregnant individuals with preeclampsia (7 mild and 15 severe) and 16 healthy controls was analyzed by low-coverage whole genome sequencing. We compared cfDNA fragment length profiles and applied non-negative matrix factorization to identify group-specific fragmentation patterns. Transcription factor (TF) profiling was performed to identify PE-associated regulators and cell-type contribution analysis was used to estimate the cellular origin of plasma cfDNA.
Results
PE samples showed a more prominent mononucleosomal peak, while healthy controls showed longer dinucleosomal fragments, suggesting differences in cfDNA fragmentation and clearance dynamics. TF profiling identified 18 differential accessible TFs, with a monotonic trend from healthy to mild and severe PE. Several of these TFs are linked to immune regulation and placental function, including RORC, a regulator of Th17 cells, whose imbalance is known in PE, and GRHL3, linked to PE-associated preterm birth. Cell-type contribution analysis revealed increased endothelial and extravillous trophoblast cfDNA signals in PE, in line with placental hypoxia and disrupted trophoblast turnover.
Conclusion
Low-coverage cfDNA sequencing identified fragmentomic and epigenetic signatures associated with PE. These results support the use of cfDNA as a non-invasive tool for disease biology investigation and biomarker discovery in PE.