B-S.B.27: Towards Stable Clustering in Hierarchical Variational Autoencoder Models for Multi-Omics Cancer Data Integration
Cancer is characterized by intra-tumor heterogeneity, with distinct subpopulations driven by diverse molecular programs influencing disease progression and treatment response. Multi-omics profiling provides complementary views of tumor biology, yet many computational approaches rely on shared representations that obscure modality-specific signals and limit disentangling of regulatory contributions.
We previously introduced CAVACHON, a hierarchical multi-omics framework using a variational ladder autoencoder (VAE) that leverages prior biological knowledge through a directed acyclic graph (DAG) to encode dependencies between modalities. It preserves hierarchical relationships in latent space, enabling identification of regulatory programs and subpopulations, with prior validation in single-cell multi-omics integration.
Here, we extend CAVACHON to improve clustering stability and interpretability, focusing on the reconstruction-regularization balance, a key challenge in variational models. We introduce three methodological improvements: (i) continuous Kullback-Leibler divergence annealing, enabling smooth transition from vanilla VAE to Gaussian Mixture Model-based clustering, (ii) K-means++ prior initialization, and (iii) per-sample dispersion modelling to separate data variability from regularization.
We evaluated these improvements on large-scale bulk RNA-seq datasets (TCGA), showing reliable performance on bulk data, the most abundant type in cancer research. We assessed clustering stability across 12 cancer types over 10 independent runs, demonstrating strong reproducibility (mean pairwise ARI = 0.75, std = 0.05). Across cancers, 6/12 showed high consistency (≥70%), including 4/12 with perfect reproducibility (std = 0). The remainder consistently formed single or multi-cluster patterns, while cancers with established molecular subtypes (BRCA, KIRC, KIRP, LGG) exhibited multi-cluster organization, indicating that the approach provides stable, reproducible clustering while preserving biologically meaningful heterogeneity.
Co-authors: Ping-Han Hsieh, Tatiana Belova, Mariike Kuijjer
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826