View posters by category
Scroll down to view Results
Session A: Monday 31 August 12:00-13:30
|
Session B: Tuesday 1 September 16:15-17:45
|
|
|
|
Session C: Wednesday 2 September 11:30-13:00
|
|
|
|
Results
B-G.01: Circulating DNA reveals nucleosome occupancy patterns that are associated with nucleosome-DNA
affinity and are affected in cancer
Track: Genomics, epigenomics, and genome editing
-
Marianne Richaud, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University,
Montpellier, France
-
Ekaterina Pisareva, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University,
Montpellier, France
-
Alain Thierry, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University,
Montpellier, France
-
Jacques Colinge, IRCM, Montpellier Cancer Research Institute, INSERM U1194, Montpellier University,
Montpellier, France
Presentation Overview: Show
The study of cell-free circulating DNA (cirDNA) fragments (fragmentomics) from liquid biopsies has received
increasing attention. By mapping a large ensemble of well-positioned nucleosomes (WPNs), we found that
nucleosome occupancy was associated with histone-DNA affinity, as evidenced by codon usage bias and
differences in cirDNA fragment sizes. Moreover, nucleosome occupancy was different in healthy and cancer
samples, thus allowing developing a high-performance machine learning approach for cancer detection
(specificity and sensitivity >0.95 for seven cancer types in Cristiano et al., 2016, data). Cancer
influenced nucleosome occupancy in a global manner, although distinct cancer types retained specific features.
WPN occupancy at transcription factor binding sites revealed shared, pan-cancer regulation of transcriptional
programs involved in hematopoietic cell differentiation and neutrophil biology, the main cirDNA sources. This
work provides new fundamental insights into cirDNA and DNA sequence using cirDNA as a physical readout. It
also bares translational significance by disclosing a new high-performance strategy for cancer detection from
liquid biopsies.
B-G.02: Genome-wide association analysis of primary response to anti-TNF therapy in inflammatory bowel
disease
Track: Genomics, epigenomics, and genome editing
-
Maria Gretsova, Institute of Clinical Molecular Biology, Kiel University, Germany
-
Johan Burisch, Gastro Unit, Medical Division, University Hospital Copenhagen, Amager and Hvidovre
Hospital, Denmark
-
Vibeke Andersen, Institute of Molecular Medicine, University of Southern Denmark, Denmark
-
Walter Reinisch, Department of Internal Medicine, Medical University of Vienna, Austria
- David Ellinghaus, Institute of Clinical Molecular Biology, Kiel University, Germany
Presentation Overview: Show
Inflammatory bowel disease (IBD), comprising Crohn's disease (CD) and ulcerative colitis (UC), is a chronic
relapsing inflammatory disorder with rising global prevalence. Anti-TNF therapy is widely used as a first-line
biologic treatment, yet a substantial proportion of patients fail to achieve early clinical remission.
Predictive biomarkers for treatment response are therefore of considerable clinical interest, but robust
genetic markers have remained largely undefined.
To investigate genetic determinants of anti-TNF response, we analyzed nine discovery cohorts from Denmark,
Austria, and Germany including IBD patients treated with anti-TNF as their first biologic therapy. Two
independent cohorts from Germany and the United States were used for replication. For each IBD subtype, we
performed genome-wide single-variant association analyses, followed by fine-mapping and colocalization, and
complemented these analyses with gene-based association testing. We additionally applied cross-phenotype
meta-analysis to explore shared genetic signals across CD and UC. To place loci emerging from the genetic
analyses into a disease-relevant biological context, we integrated publicly available single-cell RNA-seq data
from the scIBD database and assessed expression patterns across healthy and inflamed intestinal tissue
states.
This study establishes an analytical framework for the discovery and biological contextualization of genetic
factors associated with anti-TNF treatment response in IBD.
B-G.03: Exploring plasmids within human gut microbiomes via Micro-C metagenomics
Track: Genomics, epigenomics, and genome editing
-
Marcos Bermejo-Ruiz, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen,
Germany, Germany
-
Carolin Wilhelm, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany,
Germany
-
Heike Budde, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany,
Germany
-
Ruth Ley, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany, Germany
-
Alexander Tyakht, Department of Microbiome Science, Max Planck Institute for Biology, Tübingen, Germany,
Germany
Presentation Overview: Show
Bacterial plasmids are known to be present in the human gut and to confer diverse traits to their hosts, such
as antibiotic resistance or metabolite absorption. However, the diversity, specificity, and dynamics of
plasmid-bacteria associations remain to be fully characterized. Due to the limited reference data and the
narrow scope of existing databases and signature genes, conventional metagenomics approaches struggle to
assign bacterial hosts to plasmids. To overcome these caveats, Hi-C metagenomics has been employed to obtain
spatial information about DNA. However, this method has shown limited performance in human microbiomes, and
its reliance on restriction enzymes further constrains its effectiveness. In our study, we introduce Micro-C
metagenomics. Micro-C, similar to Hi-C but independent of restriction enzymes, provided the spatial
information needed to link sequences of plasmids and bacterial chromosomes, obtained using short- and
long-read metagenomics. Tested on a synthetic community and validated on a real-life stool sample, this
approach enabled the construction of high-resolution contact maps of the sequences. Additionally, the method
successfully recovered the expected plasmid-bacteria links in the synthetic community, and subsequently
identified links also in the stool sample. Thus, Micro-C metagenomics can effectively retrieve connections
between plasmid sequences and their bacterial hosts. This enables the construction of interaction networks and
allows the tracking of plasmid transmission, which facilitates further study of plasmid roles in the human gut
microecology.
B-G.04: Improved tumor-only variant calling and mutation burden estimation with VarNet-T
Track: Genomics, epigenomics, and genome editing
-
Kiran Krishnamachari, Genome Institute of Singapore, A*STAR, Singapore
- Anders Skanderup, Genome Institute of Singapore, A*STAR, Singapore
Presentation Overview: Show
Somatic variant calling algorithms typically detect mutations in cancer genomes by comparing sequence data
from a tumor sample against a matched normal sample. However, matched normal samples are often unavailable in
clinical diagnostics or retrospective analyses of archival tumor samples in biobanks, compromising variant
calling accuracy due to the difficulty in distinguishing somatic mutations from germline mutations or
sequencing artifacts. Here, we introduce VarNet-T, an end-to-end weakly supervised deep learning framework for
accurately identifying somatic variants from aligned tumor reads without a matched normal sample. VarNet-T is
trained using millions of high-confidence variants and benchmarked using public datasets, demonstrating 20-33%
performance improvement over existing methods. We assess the accuracy of tumor mutation burden (TMB)
estimation on 1000 tumor samples spanning 10 solid cancer types. Compared to existing methods, VarNet-T
demonstrates >3x higher accuracy in TMB-high status classification, suggesting significant potential to
improve patient selection for immunotherapy. Overall, the improved accuracy of VarNet-T has the potential to
enhance the utility of tumor-only sequencing in cancer research and clinical molecular diagnostics.
B-G.05: Gene2Phenotype (G2P): a database of detailed, structured gene-disease associations
Track: Genomics, epigenomics, and genome editing
-
Diana Lemos, EMBL-EBI, United Kingdom
- Seeta Ramaraju Pericherla, EMBL-EBI, United Kingdom
- Sarah E Hunt, EMBL-EBI, United Kingdom
- Elena Cibrian Uhalte, EMBL-EBI, United Kingdom
- Michael Yates, University of Edinburgh, United Kingdom
- Ian Simpson, University of Edinburgh, United Kingdom
- Helen V Firth, Addenbrooke's Hospital Cambridge University Hospitals, United Kingdom
- Mallory Freeberg, EMBL-EBI, United Kingdom
Presentation Overview: Show
Up-to-date gene-disease association information and tools are needed to identify plausibly disease-associated
variants from the large numbers generated in diagnostic genome sequencing. The Gene2Phenotype (G2P) system was
designed to meet this need. It combines expert-curated gene-disease association information from the
literature with a genotype filtering method built around the Ensembl Variant Effect Predictor (VEP) to enable
robust identification of genotypes for prioritisation.
We present a redesigned G2P web interface and a new REST API to enable both interactive exploration and
programmatic access. To support more accurate diagnosis, the research and development of novel therapies, the
updated platform disseminates more detailed information on disease mechanisms. The website provides an
enhanced user experience with an intuitive interface and search tool and includes links to publications
supporting each association. Additionally, each G2P record has a stable unique identifier enabling tracking
and direct links to external pages. The API allows flexible programmatic querying of G2P records by commonly
used identifiers or full data download, facilitating the integration of G2P data into other systems.
We are now utilising machine learning to support manual curation while improving data quality. We have
developed a pipeline which identifies relevant publications and maps them to existing G2P gene-disease
records. To improve robustness, we perform cross-model validation using an AI model to assess the relevance of
candidate publications with the corresponding G2P records. These methods will accelerate evidence acquisition
and assessment, making new knowledge accessible for research and diagnostic purposes more rapidly.
B-G.06: Longitudinal cfDNA methylation profiling by EM-seq and TAPS+ for early breast cancer biomarker
discovery
Track: Genomics, epigenomics, and genome editing
-
Daniyar Karabayev, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki,
Helsinki, Finland, Finland
-
Rodos Rodosthenous, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki,
Helsinki, Finland, Finland
-
Maare Arffman, Applied Tumor Genomics, Research Programs Unit, Faculty of Medicine, University of
Helsinki, Helsinki, Finland, Finland
- Ican, iCAN Digital Precision Cancer Medicine Flagship, Helsinki, Finland, Finland
-
Sirpa Leppä, Applied Tumor Genomics, Research Programs Unit, Faculty of Medicine, University of Helsinki,
Helsinki, Finland, Finland
-
Andrea Ganna, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki,
Finland, Finland
-
Esa Pitkänen, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki,
Finland, Finland
Presentation Overview: Show
Cell-free DNA (cfDNA) methylation profiling has emerged as a promising non-invasive approach for early cancer
detection, yet breast cancer remains particularly difficult to identify via liquid biopsy. Targeted
methylation-based approaches have shown strong performance across cancer types, but breast cancer-specific
signal remains elusive. Integrating whole-genome cfDNA methylation with matched tumor multi-omics data offers
an opportunity to identify more sensitive and specific epigenetic biomarkers for early detection.
Plasma cfDNA was obtained from breast cancer patients through the Helsinki Biobank, matched to tumor samples
available in the iCAN Digital Precision Cancer Medicine Flagship. Two complementary whole-genome (WG) cfDNA
methylation sequencing approaches were applied. WG EM-seq was performed on 52 plasma samples from 26
individuals, each with a pre-diagnostic (mean 21 months prior; range 6–47 months) and a post-diagnostic
timepoint (<2 months after diagnosis). WG TAPS+ is being applied to an expanded cohort of 181 samples from
unique individuals with single timepoints post-diagnosis. Sequencing reads were processed using DRAGEN and
nf-core/methylseq and integrated with matched tumor multi-omics data on the iCAN Discovery Platform.
cfDNA concentrations in the EM-seq cohort ranged from 36 to 1,170 ng/mL (mean 198 ng/mL) at a median coverage
depth of 6.5× and a 95% alignment rate. Initial exploratory analysis of genome-wide methylation profiles
revealed preliminary clustering patterns distinguishing breast cancer samples from external healthy controls
at both pre- and post-diagnostic timepoints, suggesting a detectable epigenetic signal before diagnosis. Copy
number profiles called using ichorCNA were largely copy-number neutral across the cohort, with a single case
exhibiting detectable alterations.
B-G.07: Methylation profile analysis of 'Candidatus Phytoplasma mali' infected apple leaves
Track: Genomics, epigenomics, and genome editing
-
Sara Bortolini, Laimburg Research Centre, Bozen-Bolzano, Italy and University of Siena, Siena, Italy,
Italy
-
Daniela Tarau, Laimburg Research Centre, Italy
-
Luca Fontanesi, Laimburg Research Centre, Corporate System Improvement Kerakoll Group, Italy
- Mattia Tabarelli, Laimburg Research Centre, Italy
- Cameron Cullinan, Laimburg Research Centre and University of Bolzano, Italy
- Cecilia Mittelberger, Laimburg Research Centre, Italy
- Katrin Janik, Laimburg Research Centre, Italy
Presentation Overview: Show
Apple proliferation (AP) disease, caused by 'Candidatus Phytoplasma mali', represents one of the most
significant threats to apple (Malus domestica) cultivation. The pathogen is a phloem-limited, cell wall–less
bacterium transmitted primarily by psyllid vectors. AP disease presents itself as a wide range of symptoms
including the characteristic rosettes, witches' brooms, and enlarged stipules, as well as more general
symptoms like leaf yellowing or reddening, reduced growth, and undesirable fruits.
In this study, we focused on the methylation profile of the host plant, M. domestica, to better understand
host responses to infection. DNA methylation studies have recently emerged as an important field for advancing
our understanding of plant–pathogen interactions. A set of infected apple trees was selected, and both
symptomatic and asymptomatic leaves were collected from each plant. Infection was confirmed by PCR analysis of
genomic DNA extracted from root tissues.
Subsequently, DNA samples were subjected to bisulfite treatment, sequencing libraries were prepared, and
high-throughput Next Generation Sequencing was performed. Here, we present a bioinformatic pipeline for the
analysis of DNA methylation profiles in infected apple plants, providing insights into epigenetic changes
associated with 'Ca. P. mali' infection.
The analysis workflow comprised quality control, read trimming, and mapping to the reference genome and
corresponding annotation. Furthermore, statistical analysis of the different genomic elements of M. domestica
were conducted on CpG, CHG and CHH methylation. This approach should help to understand the influence of DNA
methylation patterns on the manifestation of symptoms in apple trees.
B-G.08: NanoCanvas: A Unified Interactive Browser for Multimodal Nanopore Sequencing Data
Track: Genomics, epigenomics, and genome editing
-
Guangzhao Cheng, FIMM, HiLIFE & Applied Tumor Genomics Research Program, University of Helsinki,
Finland
-
Esa Pitkänen, FIMM, HiLIFE & Applied Tumor Genomics Research Program, University of Helsinki; iCAN
Flagship, Helsinki, Finland, Finland
Presentation Overview: Show
Motivation: Nanopore sequencing has become a transformative technology for genomic and epigenomic research,
generating multimodal data including raw electrical signals (POD5/FAST5), basecalled alignments (BAM),
resquiggled signal events (EventAlign TSV), and base modification calls (BED). However, the analytical
ecosystem remains fragmented: researchers must juggle incompatible tools and formats, struggle with datasets
spanning tens to hundreds of gigabytes, and manually script ad-hoc comparisons across cohorts (e.g., wild-type
vs. knockout). Existing visualization tools generate static images and crash under interactive queries at this
scale.
Results: We present NanoCanvas, an open-source, web-based platform that unifies nanopore data exploration
through interactive multimodal visualization. A FastAPI backend ingests POD5/FAST5, BAM, EventAlign TSV, BED,
and reference FASTA into a tiered indexing layer that retrieves any single read from multi-gigabyte POD5
datasets in approximately 50 ms, scaling near-constantly with cohort size, and delivers cohort-level QC and
signal aggregation, reducing exploratory analyses from minutes of static plotting (e.g., NanoPlot) to
sub-second interactive responses. The frontend integrates five workspaces (quality control, single-read
inspection, k-mer context model, transcript view, and gene/locus view) that synchronize raw signals,
basecalls, and modifications within a single physical coordinate system. Native multi-cohort overlays enable
comparison of QC metrics, signal distributions, and modification patterns across groups, eliminating custom
plotting scripts. By unifying genomic context, basecalls, modifications, and raw signal in a single
interactive environment, NanoCanvas lowers the barrier to exploratory nanopore analysis beyond specialist
users.
B-G.09: Discovery of disease genes by combining transcription factor, epigenome and DNA variation
data
Track: Genomics, epigenomics, and genome editing
-
Nina Baumgarten, Goethe University Frankfurt, Germany
- Marcel H. Schulz, Goethe University Frankfurt, Germany
Presentation Overview: Show
Genome wide association studies (GWAS) identified thousands of genetic variants, such as Single Nucleotide
Variants (SNVs) associated to traits or diseases. A substantial amount of these variants is non-coding and
might be located within cis-regulatory elements (CREs). They can affect Transcription Factor (TF) binding
sites and alter the target gene expression. Identifying those genes is a crucial step in understanding the
molecular mechanisms underlying a trait or disease.
Previously developed method for disease gene identification primarily focuses on DNA variation and/or
epigenome data, neglecting the regulatory aspect. Especially for non-coding SNVs considering altering TF
binding is important to understand the regulatory mechanisms of disease genes.
We present an approach for the interpretation of non-coding variants, by identifying regulatory SNVs (rSNVs),
which are predicted to affect TF binding sides. Furthermore, we integrate epigenome data to link rSNVs to
target genes and use statistical methods to aggregate information of TFs and genes. To priorities disease
genes with high confidence, we model for each gene the expected number of CREs overlapping with rSNVs using
the Poisson-binomial distribution, while considering the LD structure. Using GWAS from different complex
diseases, we showcase that our approach can identify disease relevant genes, whether protein-coding or
non-coding RNA and highlight relevant genes not found with a well-established method. Further, we illustrate
the usefulness of our approach using epigenome and variation data from the epiATLAS to predict
disease-specific changes in TF binding sites across neuropsychiatric disorders.
B-G.10: Assessing SNV and SV Callers using HiFi Long Read WGS to Find Causal Variants for Mendelian Traits in
Goats
Track: Genomics, epigenomics, and genome editing
-
Laura Voitl, Institute of Genetics, University of Bern; Interfaculty Bioinformatics Unit, University
of Bern, Switzerland
-
Rémy Bruggmann, Interfaculty Bioinformatics Unit, University of Bern; Swiss Institute of Bioinformatics,
Switzerland
- Cord Drögemüller, Institute of Genetics, University of Bern, Switzerland
-
Anna Letko, Institute of Genetics, University of Bern; Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Long-read whole-genome sequencing (LR-WGS) has the advantage of enabling detection of large and repetitive
structural variants (SV) compared to short-read (SR-WGS). Here, we sequenced 20 Swiss goat genomes
representing 10 local breeds using PacBio Revio. To broaden the diversity of our dataset, we included 12
publicly available genomes from the European Nucleotide Archive. We aligned all samples using pbmm2 to the
T2T-goat1.0 Inner Mongolia cashmere goat reference available on NCBI.
We subsequently constructed variant catalogs based on the aligned genomes using a reference-based approach. We
used two tools for SNV calling (DeepVariant and clair3) and SV calling (Sniffles2 and Sawfish2) respectively.
SNV calls were also compared to SR-WGS data of the Swiss cohort. Our aim was to compare the output of each
caller by investigating which variants were called by each tool. We also tested for the presence of known
caprine variants for Mendelian traits from the Online Mendelian Inheritance in Animals (OMIA) database.
Both shared and unique variants were detected across the different callers. Known variants of interest
included copy number variants affecting coat color, a complex SV on chromosome 1 causing polled intersex
syndrome, and variants present in CSN1S1, which affect milk protein. We confirmed the expected presence of
these different variants in the Swiss goats and evaluated their occurrence in the public dataset.
This work shows the benefit of using LR-WGS to detect the full spectrum of genetic variation in goats. The
variant catalogs provide a valuable and sustainable resource for small ruminant genomics.
B-G.11: The nationwide genomic characteristics and phylogenetic evolution of ST23-K1 hypervirulent Klebsiella
pneumoniae in relation to virulence and antimicrobial resistance acquisition
Track: Genomics, epigenomics, and genome editing
-
Qiucheng Shi, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
- Jingyi Zhu, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
- Rui Weng, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
-
Guanhong Lu, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
- Yan Pan, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
- Ping Zhang, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
-
Jingjing Quan, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
-
Dongdong Zhao, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
- Yunsong Yu, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, China
- Zhengan Wang, Zhejiang University, China
-
Yan Jiang, Zhejiang University, China
Presentation Overview: Show
Objectives: Hypervirulent Klebsiella pneumoniae (hvKp) ST23-K1 poses a global health threat due to its high
virulence and increasing antimicrobial resistance. This study aimed to characterise the genomic feature and
phylogenetic evolution of ST23-K1 in China.
Methods: K1 isolates from a nationwide epidemiological surveillance project underwent whole-genome sequencing.
Virulence was assessed using hypermucoviscosity phenotyping and a murine infection model. For ST23-K1 carrying
acquired antimicrobial resistance genes (ARGs), the CRISPR/Cas system, protospacers, anti-CRISPR (Acr) genes,
and plasmidome were characterised. Time-resolved phylogenetic analysis was performed using integrated locally
generated and publicly available data.
Results: Among 400 K1 isolates, ST23 was the most prevalent sequence type, and its effective population size
increased following CG23-I divergence. The CG23-I sub-lineage was widely distributed nationwide with limited
evidence of clonal transmission. Isolates with an incomplete cps locus exhibited significantly reduced
virulence compared with those carrying an intact locus. The prevalence of extended-spectrum
β-lactamase-positive ST23-K1 isolates increased over time, whereas carbapenemase-producing isolates remained
stable. Among acquired ARGs-positive ST23-K1 isolates, a conserved protospacer corresponding to a prevalent
spacer was identified. This protospacer, together with AcrIE genes, was frequently co-located on IncFII-type
plasmids.
Conclusion: ST23-K1 remains a hypervirulent lineage undergoing ongoing evolutionary expansion. The presence of
acquired ARGs in ST23-K1 may be associated with AcrIE-harbouring IncFII plasmids, and functional validation is
required to clarify the underlying mechanisms. Continuous genomic surveillance is essential to monitor the
evolution and antimicrobial resistance trends of ST23-K1.
B-G.12: STOAT : a tool for pangenome-guided genome wide association studies
Track: Genomics, epigenomics, and genome editing
-
Xian Chang, MIAT, INRAE, France
- Matis Alias-Bagarre, IRSD, INSERM, France
- Jean Monlong, IRSD, INSERM, France
- Matthias Zytnicki, MIAT, INRAE, France
Presentation Overview: Show
Genome wide association studies (GWAS) are a fundamental tool for finding associations between genotypes and
phenotypes, but standard GWAS tools are limited by their reliance on the reference genome. Such studies are
generally limited in their ability to detect large and complex structural variants, and are incapable of
detecting potentially significant “off-reference†variants that do not touch the reference genome, such as
a SNP that is inside of an insertion. Pangenomics is a powerful emerging paradigm that uses a collection of
genomes as a reference, rather than a single haplotype reference genome. Pangenomes facilitate the
identification of off-reference variants, including large and complex structural variants.
We present a new tool for using pangenomes to perform GWAS. Our tool, STOAT (snarl tree-orchestrated
association test), uses a pangenome to identify and characterize variants and the hierarchical relationship
among them, including off-reference variants that may be missed by traditional methods. Using either the
samples in the pangenome graph or variant calls from a pangenomic genotyping pipeline, we can test all
variants identified in the pangenome and in particular, we are able to independently test nested variants,
ignoring their parent or child variant. We show on simulated and real data that STOAT is able to detect the
same variants deemed significant by traditional methods, and that it can test variants that these tools would
miss.
B-G.13: A reference-free pangenomic pipeline to uncover transposable element dynamics across genomes
Track: Genomics, epigenomics, and genome editing
-
Johann Confais, INRAE / URGI, France
- Hadi Quesneville, URGI-INRAE, France
- Somia Saidi, URGI-INRAE, France
Presentation Overview: Show
Transposable elements (TEs) are major drivers of genome evolution and adaptation, yet their intraspecific
dynamics remain difficult to characterize due to methodological biases associated with reference-based
approaches. With the increasing availability of high-quality de novo genome assemblies, there is a need for
scalable tools that can accurately capture TE diversity at the pangenome level.
Here, we present panREPET, a novel reference-free computational pipeline designed to detect and characterize
shared TE insertions across multiple genomes. By performing pairwise comparisons of independently annotated
assemblies, panREPET identifies homologous TE copies based on sequence similarity and genomic context,
enabling the reconstruction of TE insertion histories without reliance on a reference genome. The method
provides precise coordinates and full sequences of TE copies in each genome, and classifies them into core,
shared, and singleton insertions.
We applied panREPET to 42 Brachypodium distachyon genomes, identifying over 18,000 shared TE insertions and
revealing that the majority of TE diversity is driven by recent, lineage-specific events. Using a SNP-based
dating framework, we uncovered four major bursts of TE activity, including two associated with key climatic
transitions—the Last Glacial Maximum (~22 kya) and the Holocene (~10 kya)—suggesting a link between
environmental stress and TE mobilization. Comparative benchmarking shows that panREPET improves specificity
and resolution over existing reference-based and structural-variant-based methods.
Overall, panREPET enables a comprehensive and unbiased exploration of TE dynamics at the species level,
providing new insights into the role of mobile elements in genome evolution and environmental adaptation.
B-G.14: Multi-feature predictive modeling of novel cancer predisposition genes using pan-cancer data
Track: Genomics, epigenomics, and genome editing
-
Jaejun Lee, Spanish national cancer research center, Spain
- A-Reum Nam, Spanish national cancer research center, Spain
- Solip Park, Spanish national cancer research center, Spain
Presentation Overview: Show
Cancer predisposition genes (CPGs) are defined as those inherited variants that increase the risk of
tumorigenesis. Despite their functional and clinical role in tumorigenesis, the number of identified CPGs
remains limited. Due to the rarity of pathogenic germline variants and the limitations of traditional
case-control approaches, the systematic approach to identify novel cancer predisposition genes is demanding.
We developed a machine learning-based predictive model integrating 13 various biological features, including
germline genomic mutation status, gene expression, somatic second hit mutation, and general characteristics of
human genes, to identify novel candidates. While single features possessed moderate predictive performance,
our integrative model outperformed all single-feature models (AUC =0.74 vs. 0.53-0.58), representing the value
of multi-feature integration. We further applied the model across individual cancer types with sufficient
sample sizes. Overall, we identified the top 8 high-confidence CPG candidates at both cancer-type-specific and
pan-cancer levels. Several of the candidates were validated using an independent cancer cohort. This approach
provides a generalizable and scalable framework for uncovering novel CPGs and offers a new direction for
understanding the genetic basis of cancer susceptibility.
B-G.15: Genomic scars of chemotherapy mark the emergence of resistance in childhood cancer
Track: Genomics, epigenomics, and genome editing
- Laura Wheaton, Queen's University, Canada
- Marie Wong-Erasmus, The University of New South Wales, Australia
- Max F. Levine, Memorial Sloan Kettering Cancer Center, United States
- Katherine E. Miller, Nationwide Children's Hospital, United States
- Neerav N. Shukla, Memorial Sloan Kettering Cancer Center, United States
- Michael D. Kinnaman, Memorial Sloan Kettering Cancer Center, United States
- Dominik Glodzik, Harvard Medical School, United States
- Carol Portwine, McMaster Children's Hospital, Canada
- Sabrina Millson, McMaster Children's Hospital, Canada
- Alexandra Zorzi, Children’s Hospital at London Health Sciences Centre, Canada
- Mariam Mikhail, Children’s Hospital at London Health Sciences Centre, Canada
- Conrad Fernandez, IWK Health Centre, Canada
- Noemi A. Fuentes-Bolanos, The University of New South Wales, Australia
- Gunes Gundem, Memorial Sloan Kettering Cancer Center, United States
- Andrew L. Kung, Memorial Sloan Kettering Cancer Center, United States
- Uri Tabori, The Hospital for Sick Children, Canada
- Chelsea Mayoh, The University of New South Wales, Australia
- Elli Papaemmanuil, Memorial Sloan Kettering Cancer Center, United States
- Mark J. Cowley, The University of New South Wales, Australia
- David Malkin, The Hospital for Sick Children, Canada
- Anita Villani, The Hospital for Sick Children, Canada
- Ludmil B. Alexandrov, University of California San Diego, United States
- Adam Shlien, The Hospital for Sick Children, Canada
- Timmy Wen, The Hospital for Sick Children, Canada
- Marcos DÃaz-Gay, Spanish National Cancer Research Center, Spain
- Eric N. Bergstrom, University of California San Diego, United States
- Mathepan J. Mahendralingam, The Hospital for Sick Children, Canada
- Nicholas Light, The Hospital for Sick Children, Canada
- Sasha Blay, The Hospital for Sick Children, Canada
- Joshua O. Nash, The Hospital for Sick Children, Canada
- Nathaniel D. Anderson, The Hospital for Sick Children, Canada
- Jessica N. Au, University of California San Diego, United States
- Scott Davidson, The Hospital for Sick Children, Canada
- Pedro L. Ballester, The Hospital for Sick Children, Canada
-
Mehdi Layeghifard, The Hospital for Sick Children, Canada
- Syed Kashif Daud, The Hospital for Sick Children, Canada
- Lisa-Monique Edward, The Hospital for Sick Children, Canada
- S.M. Ashiqul Islam, University at Albany, United States
- Azhar Khandekar, University of California San Diego, United States
- Burcak Otlu, Middle East Technical University, Turkey
- Ledia Brunga, The Hospital for Sick Children, Canada
- Rawan Hammad, King Abdulaziz University, Saudi Arabia
- Shimaa Nassif, The Hospital for Sick Children, Canada
- Nirav H. Thacker, The Hospital for Sick Children, Canada
- Tara Feltham, The Hospital for Sick Children, Canada
Presentation Overview: Show
Childhood cancer survival now exceeds 80%, yet therapy resistance remains a defining clinical challenge,
affecting approximately one-third of patients and driving poor long-term outcomes. A central unanswered
question is when resistance emerges and whether it can be detected genomically before it becomes clinically
manifest. Chemotherapy leaves characteristic mutational footprints in tumor genomes, and because these
signatures require cancer cells to survive drug exposure, they represent a direct molecular record of
resistance. In this large-scale multi-omics study of hard-to-treat pediatric tumors from three independent
precision medicine programs, linked to curated exposure data for 86 therapies, we establish therapy-induced
mutational signatures as markers of resistance emergence. Comprehensive signature analysis identified novel
signatures exclusive to treated tumors. Platinum drugs were the dominant mutagen, accounting for 10% of all
single-base substitutions. Critically, platinum signatures were detectable as early as 91 days after treatment
initiation, and their accumulation mirrored resistance kinetics: 35% of platinum-treated patients showed
measurable signatures within one year, rising to ~50% by 18 months. Signature-positive tumors exhibited
significant overexpression of platinum resistance genes and worse outcomes upon re-exposure, functionally
validating signatures as resistance biomarkers. Subclonal analysis further identified hidden platinum
signatures in tumors lacking dominant resistance clones, suggesting early detection of emerging resistance
before clinical failure. An ensemble machine learning model revealed additional platinum-associated genomic
features beyond known signatures. These findings establish a genomic framework for monitoring resistance in
real time and informing therapy de-escalation in children with cancer.
B-G.16: LongcallD: joint calling and phasing of small, structural and mosaic variants from long reads
Track: Genomics, epigenomics, and genome editing
-
Yan Gao, Harbin Institute of Technology, China
Presentation Overview: Show
Long-read sequencing is a powerful technique capturing multiple variants within single continuous reads. This
length allows individual reads to bridge small and structural variants while carrying crucial phasing
information. However, current computational tools treat small variant calling, structural variant (SV)
detection and phasing as largely disconnected problems, failing to unleash the full potential of long reads.
Here, we present longcallD, a unified framework utilizing local multiple-sequence alignment to simultaneously
call and phase small and structural variants. By integrating germline phasing and retrotransposition
hallmarks, longcallD also identifies low-fraction mosaic variants and detects mobile element insertions
supported by a single read. Compared to existing methods, our unified approach substantially improves SV
discovery and mosaic variants accuracy while maintaining competitive small variant calling. We anticipate that
longcallD will provide a robust foundation for resolving complex genetic architectures in clinical and
evolutionary applications.
B-G.17: Integration of Chromatin Interaction Maps with GWAS Identifies TBKBP1 as a Novel Chronic HBV
Susceptibility Gene
Track: Genomics, epigenomics, and genome editing
-
Qiao Ye, Laboratory of Clinical Medicine, Air Force Medical Center, Air Force Medical University,
China
-
Xiangyi Zheng, Laboratory of Clinical Medicine, Air Force Medical Center, Air Force Medical University,
China
-
Guangyun Wang, Laboratory of Clinical Medicine, Air Force Medical Center, Air Force Medical University,
China
Presentation Overview: Show
Background: Genome-wide association studies (GWAS) have identified numerous non-coding single nucleotide
polymorphisms (SNPs) associated with susceptibility to chronic hepatitis B virus (HBV) infection. However,
assigning them to functional genes is challenging, as conventional nearest-gene annotation overlooks
long-range chromatin regulatory interactions.
Methods: We conducted a meta-analysis across two East Asian GWAS cohorts, and built liver- and blood-specific
chromatin interaction maps (Activity-by-Contact (ABC) and High-throughput Chromosome Conformation Capture
(Hi-C)). Using these tissue-resolved regulatory maps, we linked non-coding variants to candidate genes and
performed gene-based association tests.
Results: Among genome-wide significant SNPs identified in our meta-analysis, 95.3% were located in non-coding
region. Using SNP–gene regulatory pairs from ABC and Hi-C maps to perform gene-based association testing, we
identified 187 genes associated with chronic HBV infection (FDR < 0.05), representing a 3.2-fold increase
over conventional nearest-gene annotation (45 genes). These 187 candidate genes were enriched in innate immune
pathways, including Toll-like receptor (TLR), retinoic acid-inducible gene I (RIG-I), Janus kinase-signal
transducer and activator of transcription (JAK-STAT) signaling. We further identified a novel candidate gene,
TBKBP1, whose protective variant (rs67919208) was associated with reduced TBKBP1 expression in blood monocytes
and liver (expression quantitative trait locus (eQTL) P = 9.35 × 10⁻¹² and 1.06 × 10⁻¹⁵; colocalization
posterior probability for shared causal variant (PP4) > 0.8 for both). In vitro assays further demonstrated
that TBKBP1 suppresses antiviral immunity, thereby increasing susceptibility to chronic HBV infection.
Conclusion: Integrating chromatin interaction maps with GWAS uncovers novel susceptibility genes for chronic
HBV infection, including TBKBP1.
B-G.18: Pangenome graph annotation
Track: Genomics, epigenomics, and genome editing
-
Nina Marthe, IRD, France
- Matthias Zytnicki, INRAE, France
- Francois Sabot, IRD, France
Presentation Overview: Show
The increasing availability of genome sequences has highlighted the limitations of using a single reference
genome to represent the diversity within a species. Pangenomes, encompassing the genomic information from
multiple genomes, offer thus a more comprehensive representation of intraspecific diversity. However,
pangenomes in form of graph often lack annotation information, which limits their utility for forward
analyses.
The tool GrAnnoT was designed to transfer annotation in such graphs efficiently and reliably, by projecting
existing annotations from a single source genome to the graph, and subsequently to other embedded genomes. It
provides informative outputs, such as presence-absence matrices for genes, and alignments of transferred
features between source and target genomes, aiding in the study of genomic variations and evolution.
GrAnnoT was published last year in PCJ : 10.24072/pcjournal.651
GrAnnoT was then improved to handle multiple source genomes for the annotation, giving a more complete graph
annotation. This also allows to compare the annotations between the genomes embedded in the graph, and explore
the variations within the genes.
We also developped a method to annotate de novo the parts of the graph that are not covered by annotated
genomes, ensuring annotation information is available for the whole pangenome.
B-G.19: Improved 16S rRNA-focused computational methods for bacterial strain identification from long read
datasets
Track: Genomics, epigenomics, and genome editing
-
Laura Tingley, UKHSA, United Kingdom
-
Jo Dicks, Culture Collections, UK Health Security Agency, 61 Colindale Avenue, London, NW9 5EQ, UK, United
Kingdom
-
Katharina T. Huber, University of East Anglia, Norwich Research Park, Norwich, NR4 7TJ, UK, United Kingdom
Presentation Overview: Show
The 16S region of the ribosomal DNA (rDNA) has for decades served as a vital resource in the identification
and classification of bacterial species, first in taxonomic and phylogenetic studies and more recently in
metagenomic analysis. However, bacterial genomes contain between 1 and ~15 copies of the 16S, with
unquantified levels of sequence variation between them. This poses several questions. How much 16S variation
exists within a genome? Do the 16S sequences of distinct species overlap? Can intra-genome 16S variation alter
identification outcomes?
We used long PacBio reads from UKHSA's NCTC3000 dataset, which have the capacity to span the 16S gene and
beyond the rDNA operon, to assess sequence variation between distinct 16S sequence copies in Escherichia coli,
Escherichia marmotae and Shigella genomes. We developed a computational pipeline to extract operon
copy-specific reads and create consensus sequences for each rDNA operon. Extensive analysis was then conducted
to provide insight into the breadth of 16S alleles and variants within and between both strains and
species.
Extensive 16S variation was observed between copies, with 168 distinct 16S alleles amongst the 53 genomes
analysed. Despite this high level of variation, almost all alleles were found within a single species, with
only three 16S sequences shared between E. coli and Shigella genomes and no overlap between the other genome
pairs. This remarkable finding suggests our new pipeline could enable rapid, enhanced bacterial identification
capable of discriminating between phylogenetically similar but distinct bacterial species using a
multi-allelic 16S method facilitated by long read analysis.
B-G.20: Ontology-Driven Prioritization of Candidate Genes in Neurodevelopmental Disorders
Track: Genomics, epigenomics, and genome editing
-
Chiara Rivi, BioFolD Unit, Department of Pharmacy and Biotechnology, University of Bologna, Via
gobetti 85, 40126 Bologna (Italy), Italy
-
Emidio Capriotti, BioFolD Unit, Department of Pharmacy and Biotechnology, University of Bologna, Via
gobetti 85, 40126 Bologna (Italy), Italy
Presentation Overview: Show
Neurodevelopmental disorders (NDDs) represent a complex group of conditions characterised by impairments in
brain functions affecting social, motor, and cognitive abilities to varying degrees. Their genetic
heterogeneity and broad phenotypic spectrum hinder the understanding of the molecular mechanisms underlying
them, slowing the advancement of diagnostic processes and therapeutic strategies.
This work involved a multi-step approach to identify genes associated with NDDs. We first integrated HGNC
genes related to NDDs from several databases, including: Orphanet, SFARI, GeneTrek and a recent publication.
Then, a score based on the consistency of annotations across the different sources was used to divide it into
different sets with varying degree of association to NDDs. The obtained dataset was then subjected to a gene
enrichment analysis, in order to pinpoint overrepresented functions, processes, and pathways. A Multi-Ontology
Enrichment (MOE) score was hence developed. This approach prioritizes genes based on functional annotation
overlap across different Gene Ontology domains and two pathway databases.
We retrieved 9,495 genes related to NDDs, of which 2,639 showed high-confidence associations. The gene
enrichment analysis identified several enriched biological processes and pathways related to nervous system
development, synaptic processes, energy metabolism, mitochondrial dysfunction, tRNA aminoacylation and
notably, cancer-related pathways. Our MOE scoring system highlighted genes with convergent functional
annotations, underlying in particular 534 genes common across all the domains tested, hence obtaining maximum
MOE score. This work provides new insights into the complex genetic landscape of NDDs, by advancing the
understanding of NDD-associated genes, offering a novel way to prioritize candidate genes.
B-G.21: Sincei: A toolkit for exploring single-cell epigenomics data
Track: Genomics, epigenomics, and genome editing
- Fernando Sancho Gómez, Utrecht University, Netherlands
- Soufiane Mourragui, Ensocell, United Kingdom
-
Vivek Bhardwaj, Utrecht University, Netherlands
Presentation Overview: Show
Emerging single-cell sequencing protocols allow researchers to study multiple layers of epigenetic regulation
while resolving tissue heterogeneity. However, despite the rising popularity of such single-cell epigenomic
assays, a lack of user-friendly computational tools for flexible quality control and genome-wide data
exploration hinders their broad adoption. We introduce the Single-Cell Informatics (sincei) toolkit, a
command-line interface for the exploration of data from a wide range of single-cell (epi)genomics protocols
directly from aligned reads stored in BAM format. Sincei provides tools for preprocessing, feature
identification, signal aggregation, dimensionality reduction, clustering and visualization. Results are stored
in the AnnData format for seamless compatibility with a wide variety of tools. Additionally, sincei includes a
Python API that enables advanced use cases, such as generating cell embeddings using Generalized Principal
Component Analysis (GLM-PCA). We show that sincei resolves cellular heterogeneity and improves interpretation
of genomic regions by using single-cell histone modification signal across species, tissues and protocols.
B-G.22: Methylation-aware read representation for metagenomic classification and host decontamination
Track: Genomics, epigenomics, and genome editing
-
Valentina Galeone, Robert Koch Institute, Germany
- Akiyama Manato, Kitasato University, Japan
- Yasubumi Sakakibara, Kitasato University, Japan
- Martin Hölzer, Robert Koch Institute, Germany
Presentation Overview: Show
Bacterial methylation profiles offer a largely underexploited layer of biological information. Generated by
strain-specific restriction-modification systems, methylation patterns vary substantially even within a
species and carry a signal that is orthogonal to sequence composition. Crucially, these signals are best
preserved at the read level, before assembly erases molecule-level information. Until recently, technical
noise in modification calling and limited ground truth have held back read-level exploitation of this signal,
limitations that long-read nanopore sequencing is beginning to overcome.
Here, we propose a methylation-aware representation of sequencing reads to improve metagenomic classification
and host decontamination, treating reads not only as nucleotide sequences but also as carriers of modification
signals.
To this end, we extend standard k-mer frequency profiles from a 4-symbol nucleotide alphabet to a 7-symbol
alphabet incorporating per-base calls for 5mC, 4mC, and 6mA from Oxford Nanopore sequencing. We show that
methylation-aware k-mers consistently improve classification accuracy over sequence-only baselines, with
especially strong gains for host decontamination, where the dense CpG methylation signature of mammalian
genomes provides a categorical signal distinct from any bacterial pattern with the potential of transferring
readily across eukaryotic hosts. Beyond host decontamination, we investigate whether methylation-aware
profiles can resolve strain-level substructure and address metagenomic challenges where sequence composition
alone reaches a ceiling, including the long-standing problem of linking mobile genetic elements to their host
genomes.
To address these and broader metagenomic challenges, we develop methods that integrate sequence and
methylation signals to capture richer organism-specific signatures.
B-G.23: A Synthetic Data Framework for Evaluating Structural Variant Callers in Tumor-Normal Long-Read
Sequencing
Track: Genomics, epigenomics, and genome editing
-
Francisco José Villena González, Bioinformatics Unit, Spanish National Cancer Research Centre (CNIO),
Madrid, Spain, Spain
-
Fátima Di Domenico Al-Shahrour, Bioinformatics Unit, Spanish National Cancer Research Centre (CNIO),
Madrid, Spain, Spain
-
Tomás Di Domenico, Bioinformatics Unit, Spanish National Cancer Research Centre (CNIO), Madrid, Spain,
Spain
Presentation Overview: Show
Structural variants (SVs) are genomic alterations encompassing deletions, insertions, and segment
rearrangements, ranging from kilobases to entire chromosomes. Yet they remain understudied compared to single
nucleotide variants, largely because short-read sequencing technologies struggle to resolve complex genomic
regions. The emergence of long-read sequencing has transformed this landscape, dramatically improving SV
detection capabilities.
Despite these advancements, benchmarking SV callers in tumor-normal settings presents a major challenge: real
biological datasets require costly, time-consuming experimental procedures and often lack a reliable ground
truth. In silico approaches offer a powerful alternative, allowing controlled, reproducible evaluation where
the true variant landscape is fully known.
In this work, we developed a pipeline to generate synthetic tumor-normal long-read sequencing datasets with
defined SV profiles, and used them to systematically benchmark long-read SV calling tools. Notably, no
consensus has yet emerged on a gold-standard method for somatic SV detection, making rigorous and reproducible
benchmarking efforts particularly critical. Analyses were focused on large SVs inspired by the genomic
landscape of Multiple Myeloma, enabling our results to directly inform variant detection decisions in an
ongoing real-world research study on this disease.
This framework provides a scalable, bias-controlled solution for evaluating SV callers, with direct
implications for the design of somatic variant detection pipelines in cancer genomics.
B-G.24: polars-bio: fast, scalable, and out-of-core genomic intervals and format I/O for Python
DataFrames
Track: Genomics, epigenomics, and genome editing
-
Marek Wiewiorka, Institute of Computer Science, Warsaw University of Technology, Poland
-
Tomasz Gambin, Institute of Computer Science, Warsaw University of Technology, Poland
Presentation Overview: Show
Motivation. Operations on genomic intervals, such as overlap, nearest, coverage, and count_overlaps, are
foundational to bioinformatics pipelines, from variant annotation to 3D chromatin analysis. Yet widely used
Python libraries (Pybedtools, PyRanges, Bioframe, GenomicRanges) struggle past a few million intervals, lack
multi-threaded out-of-core execution, and tie peak memory to dataset size. The I/O layer feeding them, such as
pysam, is an equally severe bottleneck at the biobank scale.
Results. We present polars-bio, a Python/Rust library on Apache DataFusion, Arrow, and Polars unifying genomic
format readers (BAM/CRAM/VCF/FASTQ/FASTA/GFF/GTF/BED) and multiple interval operations in a single vectorized,
streaming, multi-threaded engine (Wiewiórka et al., Bioinformatics 2025). In single-threaded benchmarks
(AIList, 10^7 vs 1.2×10^6 intervals), polars-bio outpaces Bioframe by 6.5× (overlap), 15.5× (nearest), 38×
(count_overlaps), and 15× (coverage), using up to 90× less peak memory in streaming mode. Recent releases
extend the operation set to eight (adding cluster, complement, merge, and subtract) and, still
single-threaded, is the fastest library in all operations on the largest tested dataset, ahead of
PyRanges1,GenomicRanges, and Bioframe. polars-bio also delivers fast, out-of-core, multi-threaded I/O across
many popular file formats with range predicates pushdown via file indexes (CSI/BAI/CRAI), column pruning, and
limit optimizations, and write support for native formats. It enables efficient file-format transformations
with near-linear thread scalability on both interval operations and genomic readers. Federated SQL queries
over cloud storages extend these capabilities to datasets exceeding local memory and storage. Recent benchmark
results are documented in recent project blog posts
(https://biodatageeks.org/polars-bio/blog/2026/02/14/benchmarking-genomic-format-readers-in-python-with-polars/;
https://biodatageeks.org/polars-bio/blog/2026/02/20/interval-operations-benchmark--update-february-2026/)
Availability. Open source (Apache 2.0)
https://biodatageeks.org/polars-bio/
B-G.25: vepyr: a fast, scalable, and composable Apache DataFusion-based variant annotation engine
Track: Genomics, epigenomics, and genome editing
-
Tomasz Gambin, Institute of Computer Science, Warsaw University of Technology, Poland
- Marek Wiewiórka, Warsaw University of Technology, Poland
Presentation Overview: Show
Motivation. Ensembl VEP is the de facto standard for variant annotation, yet its Perl codebase and plugin
architecture (CADD, gnomAD, SpliceAI via tabix) scale poorly: whole-genome annotation of a single WGS VCF
takes tens of minutes to hours, and biobank-scale cohorts multiply the cost linearly. There are faster
alternatives each address one piece ( bcftools/csq for consequence calling, Illumina Connected Annotations as
a C# batch binary), but none combines Rust-native performance, a composable query-engine interface, pluggable
cache format and full Ensemble VEP concordance.
Results. We present vepyr, a Rust/Python variant annotation engine built on Apache DataFusion and Apache
Arrow, integrated with the polars-bio ecosystem. vepyr exposes annotation as DataFusion table-valued functions
that can be integrated with any SQL query plan: (i) lookup_variants() for co-located variant and
population-frequency lookup, and (ii) annotate_variants() for consequence prediction, HGVS, and IMPACT
classification. The novel cache layer offers two interchangeable paths: a full Parquet cache consumed through
any DataFusion TableProvider for scan-heavy workloads, and an optimized fjall LSM-tree store tuned for fast
point lookups. Predicate and projection pushdown with columnar file format ensure only requested annotation
columns are read from disk, eliminating VEP's full-bundle-per-variant overhead, and streaming execution
enables out-of-core annotation of population-scale cohorts. A plugin framework for sources such as SpliceAI,
AlphaMissense, and CADD is in progress, as well as multi-threaded execution. On the GIAB HG002 v4.2.1
benchmark (4.2M variants, GRCh38), preliminary results show up to 50× speedup over Ensembl VEP when run with
--everything and --hgvsc flags.
Availability. Open source (Apache 2.0); https://biodatageeks.org/vepyr/
B-G.26: MSClust: de novo clustering of single-cell methylome sequencing data
Track: Genomics, epigenomics, and genome editing
-
Johnathan Wong, The University of British Columbia, Canada
- Lauren Coombe, BC Cancer Research Institute, Canada
- Parham Kazemi, BC Cancer Research Institute, Canada
- Rene Warren, BC Cancer Research Institute, Canada
- Inanc Birol, The University of British Columbia, Canada
Presentation Overview: Show
DNA methylation is a crucial epigenetic modification, playing a central role in regulating gene expression. To
detect methylation at single-base resolution, researchers commonly rely on bisulfite sequencing, which
converts unmethylated cytosines to uracil. Because conventional bulk sequencing masks methylation
heterogeneity between cell types, single-cell approaches are required; however, single-cell data are
characterized by sparsity. To generate comprehensive cell type methylation profiles, current methods must map
these sparse reads to a reference before clustering cells based on shared epigenetic signals. Yet, the reduced
genomic complexity inherent in bisulfite conversion leaves 40–60% of reads unmapped potentially overlooking
subpopulations defined by these regions.
We introduce MSClust, a novel reference-free clustering methodology capable of using methylation information
from all reads. For each cell, MSClust records the methylation state of CG sites within a two-tiered Bloom
filter data structure. These high-dimensional profiles are then processed through UMAP dimensionality
reduction and spectral clustering to identify cell types. When validated on a dataset of 32 mouse embryonic
stem cells (~14X), MSClust achieved a 1:1 match with the experimental ground truth while demonstrating a
10-fold and 3-fold decrease in run time and memory consumption, respectively, compared to traditional methods
(Run time: 3h, Memory: 22GB). On a larger dataset of 1,390 human neurons (~259X), the method maintained high
concordance (Adjusted Rand Index = 0.76) with experimental ground truth. These results suggest that MSClust
provides a scalable, reference-independent framework that enables discovery and study of novel biomarkers that
are otherwise obscured by reference-mapping limitations.
B-G.27: Benchmarking knowledge graph embedding models for the prediction of oligogenic combinations
Track: Genomics, epigenomics, and genome editing
-
Inas Bosch, Université Libre de Bruxelles, Vrije Universiteit Brussels, Belgium
- Barbara Gravel, Université Libre de Bruxelles, Vrije Universiteit Brussels, Belgium
-
Alexandre Renaux, Universit´ Libre de Bruxelles, Vrije Universiteit Brussels, Belgium
- Ann Nowé, Vrije Universiteit Brussels, Belgium
- Maris Laan, University of Tartu, Estonia
- Tom Lenaerts, Université Libre de Bruxelles, Vrije Universiteit Brussels, Belgium
Presentation Overview: Show
Identifying the oligogenic causes of rare diseases remains a challenge, notwithstanding the advancements made
in the last decade. While a variety of predictive and ranking approaches have been proposed, their precision
remains limited, as it remains difficult to know which features may be most relevant for the design of new
predictors. We hypothesize that structured biological information, which provides an integration of relevant
biological networks and ontologies in a single heterogeneous knowledge graph, can make a difference as it
allows for learning a relevant genetic representation through KGE methods. An exhaustive benchmarking is
performed wherein we assess the performance of various state-of-the-art embedding models for the task of
identifying potentially pathogenic gene pairs. The results obtained show that these KGEs provide highly
accurate predictions, leading to an AUC PR of up to 0.93, representing a significant advancement over previous
approaches. We show nonetheless that care needs to be taken in the cross-validation when using embeddings, as
data leakage between folds will reveal overly optimistic results. The further evaluation of the methods on a
holdout set and on a group of new male infertility cases show that three Translational Distance models
(TransE, MurE, RotatE) and two Semantic Matching models (DistMult, QuatE) provide better results. The analysis
is concluded by comparing all known gene combinations for these top-ranking models, examining their
similarities and differences. Overall, KGEs provide a predictive advancement but new steps will need to be
taken to generate explanations as to why the pairs are relevant for oligogenic diseases.
B-G.28: hicVerse: A modular platform for scalable exploration of harmonized Hi-C datasets across human
samples
Track: Genomics, epigenomics, and genome editing
-
Karol Piera, Univeristy of Lausanne, Switzerland
- Gian Marco Franceschini, University of Lausanne, Switzerland
- Giovanni Ciriello, Universtiy of Lausanne, Switzerland
Presentation Overview: Show
High-throughput chromosome conformation capture (Hi-C) enables genome-wide interrogation of three-dimensional
chromatin organization; however, the growing volume and heterogeneity of public datasets limit systematic
cross-sample analysis. Here, we present *hicVerse*, a modular platform for the curation, harmonization, and
exploration of Hi-C data across diverse human samples spanning healthy and disease conditions.
hicVerse integrates approximately 1000 publicly available Hi-C datasets into a standardized, curated,
multi-resolution resource. All datasets are processed through a unified pipeline, including normalization,
quality control, and subcompartment assignment using Calder.
To support both interactive exploration and large-scale analysis, hicVerse adopts a dual-layer architecture: a
tiled visualization backend, powered by HiGlass, for real-time rendering of contact maps, and a columnar query
engine enabling efficient retrieval of genomic interactions, features, and metadata across the entire atlas.
Crucially, this architecture enables fast cross-dataset queries over arbitrary genomic loci, which would
otherwise require substantial computational resources and extensive preprocessing.
The platform comprises a metadata-rich FastAPI backend (chromAPI), an Angular-based web interface, and a
HiGlass-powered visualization layer. By decoupling visualization from analytical querying, hicVerse enables
high-performance interactive browsing alongside scalable integrative analyses.
hicVerse provides a unified resource for investigating chromatin architecture across biological contexts,
enabling large-scale hypothesis generation and comparative analyses in genome regulation, development, and
disease. The processing pipeline and platform components are designed for reproducible deployment and
integration of new data, and will be made publicly available.
B-G.29: Introducing the Y-chromosomal ancestral-like reference sequence - Improving the capture of human
evolutionary information
Track: Genomics, epigenomics, and genome editing
-
Zehra Köksal, Department of Biomedical and Clinical Sciences, Linköping University, Linköping,
Sweden, Sweden
-
Annina Preussner, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki,
Helsinki, Finland, Finland
-
Jaakko Leinonen, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki,
Helsinki, Finland
-
Taru Tukiainen, Institute for Molecular Medicine Finland (FIMM), HiLIFE, University of Helsinki, Helsinki,
Finland
Presentation Overview: Show
An essential part of reproducible genetic analyses is the use of reference sequences. However, the widely used
human reference sequences represent evolutionarily young sequences of modern men of mostly European ancestry.
For the human Y chromosome (chrY) that is widely used in evolutionary studies, this can result in misleading
variant calling. We address this problem by reconstructing the Y-chromosomal ancestral-like reference sequence
(Y-ARS) for unambiguous variant calling of evolutionary relevance.
We applied a weighted maximum parsimony approach to human and primate chrY sequencing data to construct the
Y-ARS. For Y-ARS benchmarking, we aligned 40 chrY short-read sequences from diverse haplogroups to existing
references GRCh37, GRCh38 and T2T-CHM13. Alignment to the Y-ARS yielded the largest and most consistent number
of variants across sample (mean=1400; SD=77). Alternative references yielded on average fewer variants with
greater variability across samples (mean=866–968; SD=457–531) depending on their phylogenetic distance
from the reference. Among the variants called after alignment to alternative references, an average of 46%
carried the ancestral allele, while alignments to the Y-ARS resulted in calling solely variants with
evolutionarily derived alleles.
Here, we show that the existing human reference sequences cannot capture the full range of evolutionary
information on the chrY. The Y-ARS improves presenting evolutionary information on the chrY, which makes it a
valuable resource for evolutionary applications, such as sample age estimations (e.g., TMRCA) and phylogenetic
analyses. Finally, we provide a publicly available tool, polaryzer, to annotate variants as ancestral or
derived in pre-aligned chrY data (vcf files).
B-G.30: Sparrowhawk: running bacterial bioinformatics analyses anywhere, locally, with WebAssembly
Track: Genomics, epigenomics, and genome editing
-
Víctor Rodríguez Bouza, EMBL's European Bioinformatics Institute (EMBL-EBI), United Kingdom
- John Lees, EMBL's European Bioinformatics Institute (EMBL-EBI), United Kingdom
Presentation Overview: Show
The rapid expansion of public sequencing repositories is transforming the computational demands of genomic
studies. Most datasets are shared as raw reads, yet downstream analyses, especially in pathogen surveillance
and outbreak response, demand fast, accessible, and privacy-preserving tools. To meet this need, we developed
Sparrowhawk, a unified toolkit of basic bioinformatic methods compiled from Rust into WebAssembly and
delivered as a single, user-friendly website.
Sparrowhawk integrates a lightweight bacterial genome assembler with additional WebAssembly methods for common
infectious-disease workflows, including taxonomic identification, mapping and alignment, transmission
clustering, gene calling, and host depletion. This unified, maintenance-light (as everything runs locally)
interface lowers the barrier to genomic analysis for users without command-line expertise and supports
settings with unreliable connectivity or sensitive patient data.
Benchmarking across six bacterial species with both simulated and real datasets shows the assembler's
performance is comparable to Minia, while requiring reduced computational resources. The toolkit runs entirely
in the user's browser, eliminating data transfers and enabling secure, offline-friendly analyses ideal for
clinical and field environments.
Ongoing work includes general optimisations, GPU-accelerated components, and the integration of antimicrobial
resistance detection. Sparrowhawk aims to make high-quality genomic analysis more accessible and sustainable
for infectious-disease research and genomics in general.
B-G.31: Adipose core genes drive cardiovascular risk via EMT and adipogenesis pathways
Track: Genomics, epigenomics, and genome editing
- Amos Romer, Technical University of Munich, Germany
- Sebastian Doetsch, Technical University of Munich, Germany
-
Anastasiia Diagel, Technical University of Munich, Germany
- Shuangyue Li, Technical University of Munich, Germany
- Ling Li, Technical University of Munich, Germany
- Moritz von Scheidt, Technical University of Munich, Germany
- Daniel Tews, Ulm University Medical Center, Germany
- Martin Wabitsch, Ulm University Medical Center, Germany
- Heribert Schunkert, Technical University of Munich, Germany
- Matthias Heinig, Technical University of Munich, Germany
- Zhifen Chen, Technical University of Munich, Germany
Presentation Overview: Show
Genetic loci for complex traits such as coronary artery disease (CAD) are distributed widely across the
genome, often mapping near genes with unclear connections to disease biology. The omnigenic model proposes
that gene regulatory networks interconnect these loci in disease-relevant cells, allowing peripheral genes to
influence a limited set of core effector pathways. To functionally investigate this architecture in adipose
tissue, we applied a single-cell CRISPR perturbation platform to genetically prioritized CAD genes in human
adipocytes. Perturbation of multiple candidate loci revealed structured transcriptional convergence on a
program characterized by suppression of adipogenesis and activation of extracellular matrix remodeling.
Network analysis identified a shared set of convergent core genes enriched for cardiometabolic pathways and
druggable targets. Integration of perturbation results with human genetic analyses linked several core genes
to lipid and metabolic traits, indicating that these intermediate phenotypes may mediate the effects of core
genes' function in adipose remodeling on cardiometabolic disease risk.
B-G.32: Determinants of functional burden pleiotropy and gene dosage responses across human traits
Track: Genomics, epigenomics, and genome editing
-
Omar Shanta, Department of Psychiatry, University of California San Diego, La Jolla, CA, USA, United
States
- Sebastien Jacquemont, CHU Sainte-Justine, Canada
- Guillaume Dumas, CHU Sainte-Justine, Canada
- David Glahn, Boston Children's Hospital/Harvard Medical School, United States
-
Jonathan Sebat, Department of Psychiatry, University of California San Diego, La Jolla, CA, USA, United
States
- Almasy Laura, Children's Hospital of Philadelphia, United States
-
Stephen W. Scherer, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, Ontario,
Canada, Canada
-
Celia M. T. Greenwood, Lady Davis Institute for Medical Research, Jewish General Hospital, Montreal, QC,
Canada, Canada
-
Jeffrey R. MacDonald, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, Ontario,
Canada, Canada
-
Bhooma Thiruvahindrapuram, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto,
Ontario, Canada, Canada
-
Sayeh Kazem, Universiry Of Montreal, Canada
-
Worrawat Engchuan, The Centre for Applied Genomics, The Hospital for Sick Children, Toronto, Ontario,
Canada, Canada
- Emma E.M Knowles, Harvard Medical School, Department of Psychiatry, United States
-
Laura M. Schultz, Department of Biomedical and Health Informatics, The Children’s Hospital of
Philadelphia, Philadelphia, PA, United States
- Thomas Renne, Montreal University, Canada
- Josephine Mollon, Harvard Medical School, Department of Psychiatry, United States
- Guillaume Huguet, Université de Montréal, CHU Sainte Justine, Canada
- Florian Benitiere, CHU Sainte-Justine, Canada
- Jane Yang, CHU Sainte-Justine, Canada
- Kuldeep Kumar, CHU Sainte-Justine, Canada
Presentation Overview: Show
Gene dosage alterations, such as rare copy-number variants (CNVs), are major drivers of disease risk and
whole-body multimorbidity. Their rarity makes deciphering this biological footprint challenging. To overcome
this statistical limitation, we developed Functional Burden analysis (FunBurd), aggregating protein-coding
CNVs across 172 transcriptomic-derived tissue and cell-type networks. Applying this to 43 complex traits in
~500,000 UK Biobank participants, we captured widespread functional associations missed by single-gene
approaches.
Mapping this architecture revealed a unifying principle: CNV pleiotropy is fundamentally restricted by
evolutionary constraint and overwhelmingly concentrated within brain-specific functions. Crucially, mediation
analysis proved these pleiotropic links are driven overwhelmingly (~84%) by direct CNV effects. This
architectural constraint simultaneously governs gene dosage responses. Highly conserved brain functions,
intolerant to dosage deviation, exhibited predominantly non-monotonic (same-direction) effects, whereas less
constrained non-brain traits followed traditional monotonic rules. Within these brain functions, we uncovered
profound sex differences: females exhibited a significantly higher proportion of deletion-driven associations
for mental health traits.
We successfully replicated these associations in an independent, ancestrally diverse cohort of ~500,000 All of
Us participants. Furthermore, deletion burden demonstrated positive effect size correlations (up to r=0.78)
with rare pLoF SNVs across 87% of traits, confirming that the effects captured by the functional burden test
are consistent with an independent class of loss-of-function variants.
Our results highlight the key role of genetic constraint and brain-specific mechanisms in shaping CNV-driven
pleiotropy and monotonic gene-dosage response, providing a mechanistic basis for the whole-body multimorbidity
observed in neurodevelopmental and psychiatric conditions.
B-G.33: Multi-scale characterization of epigenomic alterations in pediatric brain tumors
Track: Genomics, epigenomics, and genome editing
-
Neda Shokraneh Kenari, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada,
Canada
-
Alejandro Mejía García, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada,
Canada
-
Marco Gallo, Arnie Charbonneau Cancer Institute, Cumming School of Medicine, University of Calgary,
Calgary, AB T2N 4N1, Canada, Canada
-
Verónica Rendo, Department of Immunology, Genetics, and Pathology, Uppsala University, Uppsala, Sweden,
Sweden
-
Bing Ren, Department of Cellular and Molecular Medicine, University of California, San Diego School of
Medicine, La Jolla, CA, USA, United States
-
Mathieu Blanchette, School of Computer Science, McGill University, Montreal, QCH3A 2A7, Canada, Canada
-
Audrey Baguette, Quantitative Life Sciences, McGill University, Montreal, Quebec H3A 2A7, Canada, Canada
-
Bhavyaa Chandarana, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada
-
Steven Hébert, Lady Davis Research Institute, Jewish General Hospital, Montreal, QC H3T 1E2, Canada,
Canada
-
Nathan Zemke, Department of Cellular and Molecular Medicine, University of California, San Diego School of
Medicine, La Jolla, CA, USA, United States
-
Michael Taylor, Department of Pediatrics, Baylor College of Medicine, Houston, TX, 77030, USA, United
States
-
Nada Jabado, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada
-
Claudia L. Kleinman, Department of Human Genetics, McGill University, Montreal, QC H3A 0C7, Canada, Canada
Presentation Overview: Show
Many lethal pediatric brain tumors are driven by aberrant epigenetic landscapes that inhibit the normal
differentiation of neural progenitor cells and transform them into a malignant state. The epigenetic landscape
is multilayered, including DNA methylation, histone modifications, and three-dimensional (3D) genome
organization, and these layers collectively shape cell identity.
To investigate how these layers contribute to tumor maintenance, we simultaneously profiled the 3D genome and
DNA methylation from two tumor types using snm3C-seq. Leveraging both modalities, we inferred structural
variants, identified intratumoral genetic heterogeneity, and defined clones within each sample. DNA
methylation-based embedding and clustering stratified malignant cells into groups concordant with structural
variant-defined clones. Differential methylation analysis across these clones uncovered distinct partially
methylated domain patterns, probably linked to variation in proliferative states among malignant cells.
3D genome-based embedding, in turn, revealed concordant heterogeneity, identifying clone-specific 3D genome
features. Understanding the role of these layers and their coordination in initiating and maintaining
malignancy in these lethal tumor types may enable the development of novel therapeutic targets.
B-G.34: ROS-Driven Somatic Mutation Landscape and Tumorigenesis in Prx1 Knockout Mice
Track: Genomics, epigenomics, and genome editing
-
Yukyung Jun, Division of National Supercomputing, Korea Institute of Science and Technology
Information, Daejeon 34141, Korea, South Korea
-
Jiheon Shin, Department of Life Science, Ewha Womans University, Seoul 03760, Korea, South Korea
-
Sang Won Kang, Department of Life Science, Ewha Womans University, Seoul 03760, Korea, South Korea
-
Sanghyuk Lee, Department of Life Science, Ewha Womans University, Seoul 03760, Korea, South Korea
Presentation Overview: Show
Peroxiredoxin 1 (Prx1) is a key antioxidant enzyme that maintains redox homeostasis by eliminating reactive
oxygen species (ROS). Prx1-deficient mice exhibit reduced lifespan, hemolytic anemia, and age-dependent tumor
formation, suggesting a role of ROS in tumorigenesis. To investigate the impact of chronic oxidative stress on
somatic mutations, we performed whole-exome sequencing of liver and spleen tissues from 15-month-old Prx1
knockout mice. Primary cells from these mice showed elevated nuclear ROS levels and increased DNA damage,
indicating ROS-driven genomic instability. Gene ontology analysis revealed enrichment in aging-associated
pathways, including DNA damage response and cognitive processes, as well as oxidoreductase activity and DNA
binding functions. Mutational signature analysis showed an age-related increase in SBS40, a signature linked
to aging and human cancers. Comparative analysis of shared mutation profiles identified 16 candidate genes.
Among them, Nek4 harbored an age-dependent stop-gain mutation, suggesting its potential role in ROS-associated
tumorigenesis. These results characterize the mutational landscape under chronic oxidative stress and provide
insight into the link between ROS accumulation, aging-related mutations, and cancer development.
B-G.35: tfClone: Accurate Determination of Haplotype Specific Clonal Copy Number Profiles in Circulating
Tumour DNA
Track: Genomics, epigenomics, and genome editing
-
Emilia Hurtado, The University of British Columbia, Canada
- Andrew Roth, University of British Columbia, Canada
Presentation Overview: Show
In cancer, tumour heterogeneity is driven by mutations that result in the development of genetically distinct
subpopulations of cells called clones. These clones may respond differentially to treatment, leading to
selective survivorship and proliferation of treatment resistant cell populations. Copy number variation and
copy number variants (CNVs) represent a significant source of signal in separating and defining clonal
profiles, and as such several methods have been developed to infer clonal copy number profiles from both bulk
and single-cell whole genome sequencing. Recent tissue sequencing methods have demonstrated the potential of
integrating phasing information when inferring clonal copy number and clonal prevalence. We propose tfClone
(tissue free clone), a computational method that uses a phase-informed Bayesian hierarchical model to
characterise the clonal composition and copy number profiles of a cancer from circulating tumour DNA. tfClone
models clonal copy number profiles using a factorial hidden Markov model, where each hidden chain is used to
infer the copy number profile of a single clone, while per-sample clonal prevalence is modeled as a latent
variable for clonal mixing proportions. Inference is performed using a combination of Gibbs sampling,
forward-filtering backwards sampling, slice sampling, and adaptive MCMC. We demonstrate the performance of
tfClone's improved sensitivity in tumour fraction and clonal prevalence detection using both synthetic and
real datasets.
B-G.36: Advancing copy-number phylogenetics through updates to MEDICC2
Track: Genomics, epigenomics, and genome editing
-
Chenxi Nie, Institute for Computational Cancer Biology, Germany
-
Tom L. Kaufmann, Institute for Computational Cancer Biology, CIO, CCCE, Faculty of Medicine and University
Hospital Cologne, Germany, Germany
-
Alexander Nicolay, Institute for Computational Cancer Biology, University Hospital Cologne, Germany
-
Roland F. Schwarz, Cancer Research Center Cologne Essen (CCCE), University Hospital and University of
Cologne, Germany
Presentation Overview: Show
Somatic copy-number alterations (SCNAs) and chromosomal instability are ubiquitous in cancer and contribute to
genome plasticity and intratumor heterogeneity. Accurate phylogenetic reconstruction from SCNA data is vital
for understanding tumor progression, yet phylogenetic inference from SCNAs remains challenging: their large
size and propensity for overlapping events violate the infinite sites assumption, rendering many standard
phylogenetic approaches unsuitable.
MEDICC2 addresses these challenges through the Minimum Event Distance (MED) framework, which uses Finite State
Transducers (FSTs) to model copy-number evolution without the independent bin assumption. Here, we present
algorithmic and methodological advances to MEDICC2 that improve both accuracy and computational efficiency.
Specifically, we introduce an optimized MED calculation algorithm with improved runtime performance, a
Metropolis-Hastings MCMC scheme for tree topology exploration, and redesigned FSTs that incorporate event
length for improved ancestral reconstruction.
We evaluate these improvements through a comprehensive benchmarking pipeline comparing our improvements
against the original MEDICC2 version across simulated datasets spanning a range of tree sizes, CNA overlap
levels, and mutation rates. Our results demonstrate notable gains in reconstruction accuracy, particularly in
challenging settings with highly overlapped CNA events, as well as substantial improvements in runtime
performance. Together, these advances further establish MEDICC2 as the leading toolkit for phylogenetic
inference from SCNA data, offering a powerful and more efficient tool for studying tumor evolution.
B-G.37: W-ASAP: Rapid Interactive Wastewater-based Viral Variant Detection
Track: Genomics, epigenomics, and genome editing
-
Alexander Taepper, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
- Gordon Koehn, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
- Felix Hennig, Independent, Switzerland
- Ivan Topolsky, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
- Chaoran Chen, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
- Tanja Stadler, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
- Niko Beerenwinkel, ETH Zurich; SIB Swiss Institute of Bioinformatics, Switzerland
Presentation Overview: Show
Wastewater-based genomic surveillance enables population-level monitoring of viral diversity, with its utility
for early variant detection and tracking resistance mutations having been demonstrated in various studies.
However, the large size of the sequencing data have made analysis workflows technically demanding and slow.
Existing dashboards and reports typically address a limited set of predefined questions but do not allow users
to explore the data independently.
The W-ASAP platform addresses these limitations by providing an interactive, browser-based system for
real-time exploration of read-level wastewater sequencing data. Built on the genomic query engines LAPIS and
SILO, which index hundreds of millions of reads, it enables querying of large datasets within milliseconds.
This allows rapid identification and analysis of emerging variants, reducing turnaround times from days to
minutes. The platform currently supports SARS-CoV-2 and RSV, providing public dashboards and APIs.
B-G.38: Automatic reanalysis of genetic data from patients with Primary Ciliary Dyskinesia
Track: Genomics, epigenomics, and genome editing
-
Anna-Lena Katzke, Department of Human Genetics, Hannover Medical School, Hannover, Germany
- Paul Siek, Department of Human Genetics, Hannover Medical School, Hannover, Germany
-
Martin Wetzke, Department of Pediatrics, Pediatric Pulmonology, Allergology and Neonatology, Hannover
Medical School, Hannover, Germany
-
Felix C. Ringshausen, Department of Respiratory Medicine and Infectious Diseases, Hannover Medical School,
Hannover, Germany
-
Ben Ole Staar, Hannover; Department of Respiratory Medicine and Infectious Diseases, Hannover Medical
School, Hannover, Germany
-
Gunnar Schmidt, Deparment of Human Genetics, Hannover Medical School, Hannover, Germany
- Bernd Auber, Deparment of Human Genetics, Hannover Medical School, Hannover, Germany
-
Sandra V. Hardenberg, Deparment of Human Genetics, Hannover Medical School, Hannover, Germany
Presentation Overview: Show
Background:
Primary ciliary dyskinesia (PCD) is characterized by motile cilia dysfunction leading to chronic lung disease.
PCD is inherited predominantly in an autosomal recessive manner, caused by pathogenic variants in more than 50
genes. 20-30% of patients with a well-defined PCD phenotype remain without molecular diagnosis, highlighting
the need for periodic reanalysis. As variant interpretation has become the bottleneck in genetic diagnostics,
we aim to investigate automated variant classification capability for the reanalysis of PCD patients.
Material and Methods:
A validation (n=30), a reanalysis (n=159), and a first-analysis cohort (n=66) of patients with a PCD phenotype
were analyzed using the automated classification tool HerediClassify, screening 908 genes. Based on clinical
records, phenotypic features were assessed using the PICADAR score (PrImary CiliARy DyskinesiA Rule) in
children and the modified PICADAR score in adults. To all PICADAR-positive cases, the PP4 criterion for
disease-specific phenotype was applied by HerediClassify.
Results:
HerediClassify showed a sensitivity of 0.67 and specificity of 1 in the validation cohort. In the reanalysis
cohort, which included patients who had remained without a genetic diagnosis following routine diagnostic
workup, six additional patients were genetically diagnosed. 71% of patients with a manual diagnosis were
correctly identified by HerediClassify in the first-analysis cohort.
Conclusion:
Automated variant classification can enhance reanalysis efficiently by identifying high-priority variants. The
systematic inclusion of phenotype-specific information, such as the PICADAR score, substantially improves the
interpretation of genetic data and should be considered standard practice.
B-G.39: Optimizing genetic association tests for small cohorts
Track: Genomics, epigenomics, and genome editing
-
Piotr Suszyński, Warsaw University of Technology, Poland
- Tomasz Gambin, Warsaw University of Technology, Poland
Presentation Overview: Show
Running genetic association tests on small cohorts, which is common for rare diseases, present unique
challenges. We have to be especially careful during data filtering stages, commonly executed before actual
tests. It also makes it more challenging to correct for population stratification. For such scenarios, when
parameter values selected using researcher experience and generally recommended default values are not precise
enough, we propose a data-driven approach to select optimal settings, maximizing efficacy and providing trust
in the results. We built a comprehensive, high performance Nextflow workflow for genetic association testing,
which includes all commonly performed data filtering and preparation steps, controlled by a rich set of
modifiable parameters. We also built another workflow, which generates synthetic testing datasets with
realistic characteristics, and a versatile tool that orchestrates the execution of mentioned workflows for the
purpose of optimizing their parameters. We have run our software on Thousand Genomes project data, executing
more than 7000 association testing workflow runs with various parameters and datasets, and we obtained an
optimized set of parameters. Results proved that the parameters are highly interdependent and rules for
selecting them in isolation fail to capture that complexity. We also found that using dosage is superior to
traditional hard genotypes association tests. To enhance the interpretability of our results we applied the
Accumulated Local Effects XAI technique. Our tools can also be used to tune the parameters to an individual
dataset characteristics, enabling robust association testing for small cohorts.
B-G.40: Transposable element diversity and its potential role in symbiosis regulation explored using new
chromosome-scale assemblies of Medicago truncatula ecotypes
Track: Genomics, epigenomics, and genome editing
-
Paulina Poniatowska-Rynkiewicz, Institute of Bioorganic Chemistry PAS, Poland
- Paweł Wojciechowski, Poznań University of Technology, Poland
- Agnieszka Żmieńko, Institute of Bioorganic Chemistry PAS, Poland
Presentation Overview: Show
Legumes play a central ecological and agricultural role due to their ability to establish symbiotic nitrogen
fixation with rhizobia. Medicago truncatula is a widely used model for studying legume biology, including
nodulation and symbiotic genome regulation. To expand genomic resources for this model legume, we generated
chromosome-scale assemblies for three additional geographically distinct M. truncatula ecotypes using
long-read sequencing technologies. By literally doubling the number of available high-quality assemblies for
this species, we establish a comparative genomic resource for investigating its structural variation and
repetitive sequence diversity. Using a multi-stage de novo transposable element (TE) annotation pipeline we
curated accession-specific TE libraries and characterized TE landscapes across all assemblies. Importantly, we
investigate how TE diversity intersects with genes involved in symbiotic nitrogen fixation. In M. truncatula,
nodulation involves activation of gene clusters located within symbiotic islands, regions known to undergo
dynamic epigenetic remodeling. Previous studies have shown that these regions undergo DNA demethylation during
nodule development, while adjacent TEs may become transiently activated before being re-silenced in the mature
nodules. We aim to establish whether natural variation in TE copy number and genomic positioning may influence
the epigenetic landscape and transcriptional regulation of symbiosis-related genes.
Our ongoing comparative analyses focus on TE diversity both genome-wide and within symbiotic islands, aiming
to link structural variation with potential regulatory consequences. These newly generated assemblies and
curated TE annotations establish a foundation for studying TE-driven genome evolution, epigenetic regulation,
and environmental adaptation in legumes.
B-G.41: DNA damage assessment from panel sequencing data and its use in the study of survivability
Track: Genomics, epigenomics, and genome editing
-
Thomas Minotto, University of Montpellier, France
- Luka Pavageau, Institut Universitaire du Cancer de Toulouse, France
- Mehmet Samur, Dana-Farber Cancer Institute, United States
- Jill Corre, Institut Universitaire du Cancer de Toulouse, France
- Sophie Lebre, University of Montpellier, France
- Alice Cleynen, University of Montpellier, France
Presentation Overview: Show
Assessing DNA damage from sequencing data is crucial for the diagnosis of cancer patients. Several
computational tools have been developed for that purpose. One of them is the Genomic Scar Score (GSS), which
evaluates copy number variations in whole genome sequencing data via a segmentation algorithm to retrieve
gains and deletions. This allows the detection of losses of heterozygosity, large-scale transitions, and
telomeric allelic imbalances, which are summed up into a final score value. In multiple myeloma, a low GSS has
been associated with superior outcome for patients. Yet whole genome sequencing is costly to scale up and not
widely used by practitioners.
Here, we adapt existing tools for GSS computation to targeted sequencing (panel) data, a sequencing technique
which is more accessible. We optimize the tools FACETS, CNVkit and PureCN, on a clinical dataset to compensate
for the lower information density, and we also study how the new score correlates with survival
information.
We find that most DNA damage is still detectable from panel data, including gain and deletion events specific
to multiple myeloma, provided that the coverage in these locations is sufficient. However, some events are
missed when genomic coverage is too low, which happens in our data with telomeric allelic imbalances.
Resulting copy number profiles can be used to predict patient survival. This new usage of panel data confirms
its efficacy for routine patient diagnosis and offers guidance on optimal probes positioning and coverage
strategies for adequate screening of significant genomic regions.
B-G.42: Genomic Instability Pattern Analysis in Advanced Stages of High-Grade Serous Ovarian Cancer
Track: Genomics, epigenomics, and genome editing
-
Sara Potente, Department of Biology, Unversity of Padova, Italy, Italy
- Luca Beltrame, IRCCS Humanitas Research Hospital, Milan, Italy
- Sonia Ismari, IRCCS Humanitas Research Hospital, Milan, Italy
- Federica Cetti, IRCCS Humanitas Research Hospital, Milan, Italy
- Lara Paracchini, IRCCS Humanitas Research Hospital, Milan, Italy
- Maurizio D'Incalci, IRCCS Humanitas Research Hospital, Milan, Italy
- Sergio Marchini, IRCCS Humanitas Research Hospital, Milan, Italy
- Chiara Romualdi, Department of Biology, University of Padova, Italy
Presentation Overview: Show
High-grade serous ovarian cancer (HGSOC) is the most common and lethal ovarian cancer histotype, with most
patients diagnosed at advanced stages (III–IV) and a 5-year survival rate below 30%. It is characterized by
near-universal TP53 mutations, defects in homologous recombination (HR) DNA repair, and pervasive somatic copy
number alterations (SCNA), which collectively contribute to extensive genomic rearrangements. However, the
landscape of chromosomal instability processes in advanced-stage disease remains poorly characterized, and
whether instability subgroups identified in early-stage HGSOC are conserved across stages and tumor sites is
unknown.
In this preliminary study, 116 formalin-fixed paraffin-embedded (FFPE) samples from 40 Stage III–IV HGSOC
patients were analyzed, including primary tumors and matched metastatic lesions from ovary, omentum,
peritoneum, and other sites. Samples were profiled by shallow whole-genome sequencing (sWGS), and SCNA
profiles and chromosomal instability (CIN) signatures were derived using the SAMURAI bioinformatics pipeline.
Each sample was classified as unstable (U) or highly unstable (HU) based on the fraction of altered genome and
number of breakpoints.
Classification identified 63 HU and 53 U samples. Clustering revealed distinct subgroups in both primary and
metastatic lesions across instability classes. CX1 and CX3 signatures — related to chromosome missegregation
and replication stress, respectively — were dominant across all clusters. The overall instability landscape
in advanced HGSOC closely resembles that of early-stage disease, suggesting conserved processes across stages.
Notably, 13 patients showed a shift in instability pattern between primary and metastatic sites, supporting
dynamic genomic evolution during metastatic progression.
B-G.43: Phylogenetic tree inference from single-cell RNA sequencing data
Track: Genomics, epigenomics, and genome editing
-
Norio Zimmermann, Department of Biosystems Science and Engineering, ETH Zurich, 4056 Basel,
Switzerland, Switzerland
- Xiaoyu Sun, None, Germany
- Joanna Hård, KTH Royal Institute of Technology, 100 44 Stockholm, Sweden, Sweden
- Jack Kuipers, ETH Zurich, D-BSSE, Computational Biology Group, Switzerland
- Niko Beerenwinkel, ETH Zurich, Switzerland
Presentation Overview: Show
Single-cell RNA sequencing technologies enable the large-scale measurement of gene expression profiles at the
individual cell level to assess cellular diversity and function. In oncology, leveraging these single-cell
transcriptomic data to reconstruct the phylogenetic relationships among cancer cells can provide insights into
tumor evolution, metastasis formation, and the development of treatment resistance. However, phylogenetic
inference from single-cell RNA sequencing is challenging due to sparse and noisy data and large dataset sizes.
We present a novel tree inference method designed for such data that takes reference and alternative read
counts of single-nucleotide variants and reconstructs a phylogenetic tree of the sequenced cells via maximum
likelihood using a random-scan greedy search.
To overcome local optima in the search, our algorithm alternates between two different tree representations:
cell lineage trees, where cells are represented by nodes and mutations are attached to edges, and mutation
trees, where mutation nodes encode the mutational events and cells are attached to them. Because a local
optimum in one tree space generally does not correspond to a local optimum in the other space, we maximize the
likelihood by switching between the two tree spaces until convergence is achieved in both. We demonstrate
superior performance on simulated data compared to existing methods. Furthermore, we show the applicability of
our approach to cancer single-cell RNA sequencing data, where it allows us to link evolutionary trajectories
of cells to their gene expression profiles.
B-G.44: An Integrated Framework for Multi-Omics Long-Read Analysis at Single-Molecule Resolution
Track: Genomics, epigenomics, and genome editing
-
Isabell Wienpahl, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin,
Germany, Germany
-
Henrik Köppke, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany;
Humboldt-Universität zu Berlin, Germany
-
Scott Lacadie, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany,
Germany
-
Roseen Musallam, Humboldt-Universität zu Berlin; Max Delbrück Center for Molecular Medicine in the
Helmholtz Association, Berlin, Germany, Germany
-
Uwe Ohler, Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany;
Humboldt-Universität zu Berlin, Germany
Presentation Overview: Show
The epigenetic landscapes of genes define their transcriptional activity. Multiple epigenetic marks crosstalk
with each other and consequently shape cell fate in development and disease. We have developed a multi-omics
assay that jointly measures chromatin accessibility, DNA methylation and 3D genome organization on the /same/
DNA molecule. Combined with third-generation sequencing, this approach preserves higher-order concatemers and
thus enables us to capture multi-way chromatin contacts in addition to pairwise contacts. However, current
software typically addresses either long-read epigenomics or concatemer-based 3D genome analysis, and no
comprehensive solution exists for an integrated interpretation of these readouts at single-molecule
resolution. We therefore aim to develop a modular toolbox that addresses all modalities of multi-omics long
read information. The toolbox is designed to analyze the full assay output within one workflow and includes
reference-based deconvolution for mixed samples, enabling haplotype-aware assignment of molecules and
cell-type deconvolution from informative loci. By unifying single-molecule epigenetic and 3D genome analysis,
our framework enables systematic characterization of epigenetic states together with their spatial context.
This work provides a computational foundation for studying allele-specific regulation, cellular heterogeneity
and genome organization from native long-read multi-omics data.
B-G.45: A reproducible workflow for HEK293 methylome profiling during adaptation to suspension growth
Track: Genomics, epigenomics, and genome editing
-
Alexander Karl Molin, BOKU University, Department of Biotechnology and Food Science, Institute of
Bioprocess Science and Engineering, Austria
-
Nikolaus Virgolini, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess
Science and Engineering, Austria
-
Georg Smesnik, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess
Science and Engineering, Austria
-
Astrid Duerauer, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess
Science and Engineering, Austria
-
Nicole Borth, BOKU University, Department of Biotechnology and Food Science, Institute of Bioprocess
Science and Engineering, Austria
Presentation Overview: Show
Biomanufacturing for transient protein expression and recombinant Adeno-associated virus (rAAV) production
often relies on empirical process development. Transitioning towards knowledge-driven optimization requires a
fundamental understanding of host cell biology, which in the case of Human Embryonic Kidney 293 (HEK293)
cells, the main expression system for rAAV production, is still limited. As such, this study presents a
computational workflow for in-depth characterization of the epigenetic landscape of HEK293 cells and will be
used in further investigations.
An adherent HEK293 cell line was adapted to suspension growth using four distinct commercial media
formulations, with HEK293-6E serving as a commercially available suspension-adapted reference.
Post-adaptation, DNA samples were taken and whole-methylome sequencing was performed using an enzymatic
methyl-conversion kit and Illumina sequencing (150 bp paired-end, 37x coverage). The resulting data was
evaluated, implementing a reproducible Snakemake-based processing and analysis pipeline. Raw data was
processed and aligned to the human reference genome augmented with the Adenovirus5 region, prior to
methylation calling via the Bismark program.
Next a downstream analysis pipeline was implemented, utilizing MethylKit and custom scripts to identify
conserved methylomic signatures associated with suspension growth. To assess cell regulation, a chromatin
state model was generated integrating publicly available datasets with ChromHMM. Mapping the experimental data
against this model revealed characteristic mammalian patterns, specifically strong methylation within
transcribed regions and lower levels at active transcription start sites and promoters. This work establishes
a reproducible data analysis framework that provides critical insights into the HEK293 epigenome, offering a
foundation for targeted host cell engineering.
B-G.46: SNooPy: a statistical framework for long-read metagenomic variant calling
Track: Genomics, epigenomics, and genome editing
-
Roland Faure, Institut Pasteur, France
- Ulysse Faure, Dep. of Mathematics, ETH Zürich, Switzerland, Switzerland
-
Tam Truong, University Rennes, Inria, CNRS, IRISA - UMR 6074, Rennes, France,, France
-
Alessandro Derzelle, Service Evolution Biologique et Ecologie, Universit´e libre de Bruxelles (ULB),
Brussels, Belgium, Belgium
-
Dominique Lavenier, University Rennes, Inria, CNRS, IRISA - UMR 6074, Rennes, France,, France
-
Jean-François Flot, Service Evolution Biologique et Ecologie, Universit´e libre de Bruxelles (ULB),
Brussels, Belgium, Belgium
-
Christopher Quince, Organisms and Ecosystems, Earlham Institute, Norwich, UK, United Kingdom
Presentation Overview: Show
Current long-read single-nucleotide variant callers were designed primarily for genomic data—particularly
human genomes. While some have been used on metagenomic data, their underlying assumptions and training
procedures fail to account for the inherent complexity of metagenomic samples. To date, no long-read variant
caller has been purpose-built for metagenomic applications. To address this gap, we present SNooPy, a
SNP-calling tool that implements a new statistical framework tailored to long-read metagenomic data. Unlike
previous genomic methods, our approach makes no assumptions about the number of haplotypes present, their
evolutionary relationships, or their sequence divergence. We demonstrate that SNooPy outperforms both
traditional statistical and deep learning–based SNP callers. Our results suggest that future integration of
this framework with deep learning approaches could further enhance variant calling performance. SNooPy is
freely available on github.com/rolandfaure/snoopy.
B-G.47: Impact of Sperm Methylome on In Vitro Produced Embryo Methylome, Transcriptome and
Development
Track: Genomics, epigenomics, and genome editing
-
Pratik P. Pathade, Aarhus University, Denmark
- Aurélie Bonnet, ELIANCE, France
- Aurélie Chaulot Talmon, INRAE, France
- Valentin Costes, ELIANCE, France
- Marie-Christine Deloche, ELIANCE, France
- Anne Frambourg, INRAE, France
- Udayaraja Gk, Aarhus University, Denmark
- Laurent Schibler, INRAE, France
- Hélène Kiefer, INRAE, France
- Véronique Duranthon, INRAE, France
- Haja N. Kadarmideen, Aarhus University, Denmark
Presentation Overview: Show
Elite bull semen is widely used in cattle artificial insemination and reproductive technologies. Beyond
genetic variation, sperm epigenetic changes can be transmitted to embryos, potentially affecting embryo
methylome, tran-scriptome, development, and calf phenotypic performance. This study evaluated the impact of
sperm DNA methyl-ation on in vitro produced (IVP) embryo methylome and transcriptome using a paired
experimental design. Ten replicates used oocyte pools from five cows fertilized with sperm from two bulls with
extreme CpG methylation differences across ~115,000 sites, generating a total of 20 pools of Day-7 expanded
blastocysts (10 pools per bull) for RRBS and RNA-seq analysis (ARS-UCD2.0).
RRBS identified 5,077 hypermethylated and 4,929 hypomethylated DMCs, predominantly in intronic regions,
followed by exonic, intergenic, and promoter regions. GO enrichment revealed hypermethylated DMCs enriched for
tissue development, neurogenesis, and DNA-binding transcription factors, while hypomethylated DMCs were
associated with transcription regulation and RNA biosynthesis. RNA-seq detected 38,001 genes, yielding 44
dif-ferentially expressed genes (FDR<0.05). Weighted Gene Co-expression Networks (WGCNA) with top 25%
varia-ble genes and power=18, R²≥0.70, identified 29 modules, three of which were significantly correlated
with bull origin (|r|≥0.60, adj.P≤0.05).
These findings suggest that contrasting sperm methylomes are associated with broad embryo epigenetic
repro-gramming, affecting both individual gene expression and co-expression networks, establishing a framework
for studying paternal epigenetic transmission in cattle IVP embryos.
B-G.48: Dysregulating genomes - mapping how breaking DNA-lamina interactions drives muscular
dystrophy
Track: Genomics, epigenomics, and genome editing
-
Anna Antonatou Papaioannou Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association
(MDC), Berlin, Germany, Germany
-
Julia Torres Rivera, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC),
Berlin, Germany, Germany
-
Narasimha Swamy Telugu, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC),
Berlin, Germany, Germany
-
Ines Lahmann, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin,
Germany, Germany
-
Werner Stenzel, Department of Neuropathology, Charité-Universitätsmedizin, Berlin, Germany, Germany
-
Sebastian Diecke, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC),
Berlin, Germany, Germany
-
Mina Gouti, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC), Berlin,
Germany, Germany
-
Michael Ian Robson, Max-Delbrück-Centrum for Molecular Medicine in the Helmholtz Association (MDC),
Berlin, Germany, Germany
Presentation Overview: Show
A key question in genome biology is how 3D genome organization regulate gene expression in health and disease.
Central to this organization is the nuclear lamina, which tethers compacted heterochromatin to the nuclear
periphery in lamina-associated domains (LADs) that safeguard gene repression and cell identity. Mutations in
lamina-tethering proteins that disrupt LAD organization are proposed to cause diverse, untreatable laminopathy
diseases, including Congenital Muscular Dystrophy (CMD). However, how LAD disruption impacts heterochromatin
and gene regulation to drive pathology remains unknown. Here, we address this by combining high-throughput
electron microscopy (EM) and single-cell multiomics to systematically profile heterochromatin defects and
their functional consequences in CMD. EM imaging of hundreds of muscle nuclei reveals heterochromatin
undergoes extensive decondensation at the nuclear periphery in CMD. However, heterochromatin also unexpectedly
decondenses in the nuclear interior, suggesting that lamina-tethering is required to globally maintain
heterochromatin function throughout the nucleus. To test this, we have established a robust experimental and
computational pipeline for joint single-cell epigenome and transcriptome profiling in tissues. With this, we
are now mapping loci and cell types where lamina-interactions and heterochromatin states are disrupted, and
when these disruptions cause ectopic gene activation. Combined, this will reveal mechanisms driving
laminopathies like CMD, revealing both the functions of lamina-directed genome organization and potential
targets for therapeutic intervention. More generally, our combed EM and single-cell toolkit provides a
generalizable framework to study epigenome dysregulation in complex tissue in other disease contexts.
B-G.49: A lesson of grammar: unraveling Transcription Factors DNA-binding syntax rules in the model plant
Arabidopsis thaliana
Track: Genomics, epigenomics, and genome editing
-
Alice Jegou, CEA - LPCV, France
- Jérémy Lucas, CEA - LPCV, France
- Marianne Dreuillet, CEA - LPCV, France
- François Parcy, CEA - LPCV, France
- Romain Blanc-Mathieu, CEA - LPCV, France
Presentation Overview: Show
Transcription factors (TFs) are master regulators of gene expression in eukaryotes, binding to specific DNA
motifs within cis-regulatory regions. Yet, beyond motif recognition, the parameters that shape TF-DNA binding
remain poorly understood, especially in plants. Studies suggest that TFs motifs follow grammar rules: spatial
arrangements that dictate binding and thus their spatial occupation. While such rules have been described for
specific families like Auxin Response Factors or MADS-TF, resulting in the formation of regulatory protein
complexes, their prevalence remains not fully explored.
In the model plant Arabidopsis thaliana, we classified 1,627 TFs into 56 families based on their DNA-binding
domains (DBDs). Then we built a comprehensive atlas of TFs binding sites (TFBS) in this model plant,
integrating 681 genome-wide binding experiments (ChIP-seq, DAP-seq, ampDAP-seq) spanning 40 TF families. By
modeling these TFBSs, we uncovered that syntax rules are not exceptions but a widespread phenomenon: nearly
every family exhibits unique, family-dependent homotypic grammar. Strikingly, comparing in vitro and in vivo
conditions revealed dynamic shifts in these rules, pointing to complex regulatory mechanisms.
These findings underscore the critical role of TFs binding grammar and suggest the formation of higher-order
TFs complexes or DNA mediated cooperative interactions. Investigating further, we would like to combine
heteromeric syntax analysis with single-cell RNA-seq and chromatin accessibility data to predict potential TF
complexes in silico.
This work provides an essential resource to the plant biology community to decode the regulatory language of
TFs.
B-G.50: Epigenomic Rewiring of Enterocyte Chromatin During Post natal Salmonella Typhimurium
Infection
Track: Genomics, epigenomics, and genome editing
-
Caterina Alfano, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
-
Anna-Lena Ullrich, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
-
Johannes Schöneich, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
-
Matthias A. Schmitz, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
-
Samuel Ward, Institute of Experimental and Clinical Pharmacology and Toxicology, University of Freiburg,
Germany
-
Sebastian Preissl, Institute of Experimental and Clinical Pharmacology and Toxicology, University of
Freiburg and I.P.W., Graz University, Germany
-
Aline Dupont, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
-
Mathias W. Hornef, Institute of Medical Microbiology, University Hospital RWTH Aachen, Germany
Presentation Overview: Show
Salmonella Typhimurium (STm) is a food-borne pathogen that causes salmonellosis. Following ingestion, STm
reaches the gastrointestinal tract, where it attaches and invades the intestinal epithelium allowing
intracellular bacterial replication and colonization. The interaction with the intestinal epithelium is thus a
crucial step for the establishment of the infection. Our goal is to uncover epigenetic modifications in the
intestinal epithelium induced by STm infection during postnatal development, a stage in which the epithelium
matures to fulfill the changing requirements in digestion, absorption, antimicrobial activity and barrier
formation. For this purpose, we generated and analyzed scATAC-seq profiles of 5-day-old mice (two controls and
two samples inoculated on day-1). By focusing on enterocytes, the primary target cells of STm, we identified
more than 15.000 differentially accessible regions, ~25% of which are annotated as promoters. Regions more
open in the infected samples are enriched in negative regulation of T cells activation and display a strong
enrichment of AP 1 family motifs, known to play a role in the host's immune response to bacterial infection.
Finally, network-based co-accessibility analysis revealed striking differences: regions' co-accessibility is
extremely more frequent in infected samples, especially between regions annotated to different genes. Large
communities of co accessible regions are unique to the infected network, and are mostly made of promoters,
signaling an infection driven reorganization of transcriptional regulatory mechanisms. Our findings highlight
a wide-spread epigenetic reprogramming of the intestinal epithelium driven by the interaction with the
pathogen. Integration with scRNA-seq experimental data will allow further characterizations of these
processes.
B-G.51: Detection of haplotype-specific contacts in phased chromatin contact maps with Genome Architecture
Mapping
Track: Genomics, epigenomics, and genome editing
-
Claudia Robens, Institute for Computational Cancer Biology (ICCB), University Hospital Cologne,
University of Cologne, Germany, Germany
-
Julia Markowski, Max-Delbrück-Center for Molecular Medicine, Berlin Institute for Medical Systems
Biology, Humboldt University of Berlin, Germany
-
Alexander Kukalev, Max-Delbrück-Center for Molecular Medicine, Berlin Institute for Medical Systems
Biology, Germany, Germany
-
Christoph J. Thieme, Max-Delbrück-Center for Molecular Medicine, Berlin Institute for Medical Systems
Biology, Germany, Germany
-
Adam Streck, Institute for Computational Cancer Biology (ICCB), University Hospital Cologne, University of
Cologne, Germany, Germany
-
Ana Pombo, Max-Delbrück-Center for Molecular Medicine, Berlin, Germany; Johns Hopkins University,
Baltimore, USA, Germany
-
Roland F. Schwarz, Institute for Computational Cancer Biology (ICCB), University Hospital Cologne,
University of Cologne, Germany, Germany
Presentation Overview: Show
Chromatin structure is key to the orchestration of gene regulation and its misfolding is involved in many
diseases including developmental diseases and cancer. Genome Architecture Mapping (GAM) is a ligation-free
method for determining chromatin conformation from minimal input material. Understanding allele-specific
regulation requires haplotype-resolved chromatin organization, which depends on direct phasing, where
sequencing reads are assigned to haplotypes based on single-nucleotide variants. The sparse variant density in
the human genome poses a challenge to direct phasing efficiency and thus limits the generation of
haplotype-specific chromatin contact maps.
We here present our novel read phasing strategy ‘CoPhasing', which leverages local haplotype information of
GAM data to correctly assign sequencing reads to their haplotype of origin, even without overlapping
heterozygous variants. CoPhasing allows for haplotype-specific analysis of chromatin folding and detection of
haplotype-specific chromatin contacts in variant-sparse genomes. We introduce a permutation test-based
algorithm for detecting significant haplotype-specific contacts in cophased GAM data. The method marks
haplotype-specific interactions in contact maps and indicates which of the two homologous chromosomes exhibits
stronger interactions. By leveraging differential contacts identified through the permutation test, we detect
genomic regions enriched for haplotype-specific interactions, highlighting loci with pronounced allelic
differences in chromatin organization.
Together, our algorithms will enable novel insights into 3D chromatin architecture in health and disease,
adding to the shifting focus beyond genomic mutations in coding regions towards epigenetic gene-regulatory
mechanisms.
B-G.52: Somatic variant detection using a personalized pangenome reference: a study of a paediatric acute
myeloid leukaemia patient
Track: Genomics, epigenomics, and genome editing
-
Elizaveta Kulaeva, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute,
Netherlands
-
Mark van Roosmalen, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Netherlands
-
Rico Hagelaar, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Netherlands
-
Eline Bertrums, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Netherlands
-
Andrea Guarracino, Bioinnovation and Genome Sciences, Translational Genomics Research Institute (TGen),
part of City of Hope, United States
-
Ruben van Boxtel, Prinses Máxima Centrum for Pediatric Oncology and Oncode Institute, Utrecht, The
Netherlands, Netherlands
Presentation Overview: Show
Current clinical pipelines using linear reference genomes cause reads carrying mutations to be misaligned or
lost, which is particularly problematic for low-coverage assays like single-cell whole-genome sequencing
(scWGS), where technical noise can interfere with true somatic events. To overcome this, we applied a
haplotype-resolved pangenome reference to a paediatric acute myeloid leukaemia (pAML) case and assessed the
clinical utility of this approach.
We constructed a personalized pangenome from long-read Nanopore sequencing of healthy haematopoietic stem and
progenitor cells, incorporating two haplotype-resolved assemblies with Minigraph-Cactus. This added 36.3 Mb of
novel sequence, 42% of which represented variation absent from the standard hg38 reference. We compared
somatic variant calling across bulk, clonally expanded, and single-cell WGS data of healthy and cancerous
cells. Reference bias in pangenome alignments was quantified using Biastools and found to be significantly
reduced compared to linear alignments.
The personalized pangenome reduced reads with mates mapped to different chromosomes up to 8,000-fold which
allowed more accurate split-read event detection and remapping of previously ambiguous alignments. It also
rescued SNVs previously filtered out by mapping quality thresholds and eliminated up to 80% of short indels in
low-complexity regions. Driver analysis revealed that frameshift insertion artifacts in scWGS data were
largely removed, and the medically relevant and computationally challenging PRSS1 locus was substantially
disentangled due to reduced misalignment.
Our findings provide systematic evidence that personalized pangenome improves somatic variant detection in
cancer genomics with the most benefit for scWGS data and complex genomic regions previously misaligned due to
linear reference bias.
B-G.53: Prioritizing non-coding variants to discover mutations governing agronomic traits in plants
Track: Genomics, epigenomics, and genome editing
-
Mathis Pochon, CEA, France
- Romain Blanc-Mathieu, CEA, France
- Francois Parcy, CNRS, France
Presentation Overview: Show
During the development of agriculture over centuries, humans have selected plants with advantageous agronomic
traits. This process has led to genomic diversification between varieties within a species. A striking example
is Brassica oleracea, a highly diversified species that includes cabbages, cauliflowers, and broccoli.
Association studies often highlight cis-regulatory regions as targets of selection. However, identifying
causal mutations remains a major challenge, as genomic regions associated with a trait can harbor thousands of
variants in strong linkage disequilibrium.
To address this challenge, we developed a tool dedicated to prioritizing non-coding variants (NCVs) with
potential regulatory functions. This tool is based on a framework originally developed for human disease
studies and distinguishes likely functional variants from non-functional ones by leveraging transcription
factor binding sites (TFBS) and modules of TFBS that show higher-than-expected phylogenetic conservation.
As a case study, we applied this tool to curd formation in cauliflower, a trait known to result from
deregulation of the floral development gene regulatory network. Using kmer variants from several accessions of
B. oleracea, we first performed an association study that highlighted both coding and non-coding regions. We
are now using our tool to specifically prioritize NCVs potentially involved in curd formation.
These results with our current knowledge of transcription factors and the gene regulatory network underlying
floral development will allow us to identify candidate causal mutations responsible for cauliflower head
formation, which can then be experimentally validated. Beyond this example, our tool provides a general
framework to identify causal NCVs associated with other agronomic traits.
B-G.54: Plasma cfDNA fragmentomics provides insight into Preeclampsia pathophysiology
Track: Genomics, epigenomics, and genome editing
-
Irene D'Onofrio, University of Zurich, Switzerland
- Zsolt Balázs, University of Zurich, Switzerland
- Elisabetta Ranieri, University Hospital of Zurich, Switzerland
- Todor Gitchev, University of Zurich, Switzerland
- Michael Krauthammer, University of Zurich, Switzerland
- Tilo Burkhardt, University Hospital of Zurich, Switzerland
Presentation Overview: Show
Introduction
Preeclampsia (PE) is a multisystem hypertensive disorder of pregnancy characterized by placental dysfunction,
maternal endothelial and end-organ injury. Despite affecting 3-8% pregnancies, PE etiology remains
incompletely understood, partly due to limited access to placental tissue during pregnancy. However, cell-free
DNA is emerging as a non-invasive biomarker for PE, with potential to provide insights into its
pathophysiology.
Methods
Plasma cfDNA from 22 pregnant individuals with preeclampsia (7 mild and 15 severe) and 16 healthy controls was
analyzed by low-coverage whole genome sequencing. We compared cfDNA fragment length profiles and applied
non-negative matrix factorization to identify group-specific fragmentation patterns. Transcription factor (TF)
profiling was performed to identify PE-associated regulators and cell-type contribution analysis was used to
estimate the cellular origin of plasma cfDNA.
Results
PE samples showed a more prominent mononucleosomal peak, while healthy controls showed longer dinucleosomal
fragments, suggesting differences in cfDNA fragmentation and clearance dynamics. TF profiling identified 18
differential accessible TFs, with a monotonic trend from healthy to mild and severe PE. Several of these TFs
are linked to immune regulation and placental function, including RORC, a regulator of Th17 cells, whose
imbalance is known in PE, and GRHL3, linked to PE-associated preterm birth. Cell-type contribution analysis
revealed increased endothelial and extravillous trophoblast cfDNA signals in PE, in line with placental
hypoxia and disrupted trophoblast turnover.
Conclusion
Low-coverage cfDNA sequencing identified fragmentomic and epigenetic signatures associated with PE. These
results support the use of cfDNA as a non-invasive tool for disease biology investigation and biomarker
discovery in PE.