B-G.26: MSClust: de novo clustering of single-cell methylome sequencing data
Authors
Lauren Coombe
BC Cancer Research Institute
Parham Kazemi
The University of British Columbia
Rene Warren
BC Cancer Research Institute
Inanc Birol
The University of British Columbia
Keywords
methylome, single-cell, 5mc, clustering, de novo, bisulfite, scBS-seq, k-mer,
[sponser-meet-now-chat][/sponser-meet-now-chat]
DNA methylation is a crucial epigenetic modification, playing a central role in regulating gene expression. To detect methylation at single-base resolution, researchers commonly rely on bisulfite sequencing, which converts unmethylated cytosines to uracil. Because conventional bulk sequencing masks methylation heterogeneity between cell types, single-cell approaches are required; however, single-cell data are characterized by sparsity. To generate comprehensive cell type methylation profiles, current methods must map these sparse reads to a reference before clustering cells based on shared epigenetic signals. Yet, the reduced genomic complexity inherent in bisulfite conversion leaves 40–60% of reads unmapped potentially overlooking subpopulations defined by these regions.
We introduce MSClust, a novel reference-free clustering methodology capable of using methylation information from all reads. For each cell, MSClust records the methylation state of CG sites within a two-tiered Bloom filter data structure. These high-dimensional profiles are then processed through UMAP dimensionality reduction and spectral clustering to identify cell types. When validated on a dataset of 32 mouse embryonic stem cells (~14X), MSClust achieved a 1:1 match with the experimental ground truth while demonstrating a 10-fold and 3-fold decrease in run time and memory consumption, respectively, compared to traditional methods (Run time: 3h, Memory: 22GB). On a larger dataset of 1,390 human neurons (~259X), the method maintained high concordance (Adjusted Rand Index = 0.76) with experimental ground truth. These results suggest that MSClust provides a scalable, reference-independent framework that enables discovery and study of novel biomarkers that are otherwise obscured by reference-mapping limitations.
Co-authors: Lauren Coombe, Parham Kazemi, Rene Warren, Inanc Birol
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826