WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.
WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.

C-S.B.65: Cluster-Aware Functional Principal Component Analysis for Imputing Missing Values in Longitudinal Microbiome Data

Authors

University of Duisburg-Essen
Mohammad Darbalaei1
University of Duisburg-Essen
Daniel Hoffmann
University of Duisburg-Essen
Farnoush Farahpour
University of Duisburg-Essen

Keywords

microbiome, longitudinal microbiome imputation, microbiome time series, trajectory modelling, functional data analysis, missing data
[sponser-meet-now-chat][/sponser-meet-now-chat]

Missing observations arising from irregular sampling and participant attrition are a major analytical bottleneck in longitudinal microbiome studies. Existing imputation methods either treat temporal observations through discrete multivariate representations, without leveraging the underlying continuity of biological processes, or rely on deep generative models such as GANs and diffusion architectures, which require substantial training data and can be unstable or difficult to interpret in studies with limited replicates or sparse time points.

We propose an imputation framework that models each taxon trajectory as a smooth latent function on the centered log-ratio scale, estimated through Functional Principal Component Analysis (FPCA) in the Principal Analysis by Conditional Expectation (PACE) framework to handle sparse and irregular sampling. Imputation pools information from observed time points within each trajectory and from other replicates through the shared FPCA covariance structure. To address heterogeneity across replicates, an adaptive clustering step based on FPCA scores is performed locally for each missing observation, so that imputation borrows strength only from trajectories with comparable temporal patterns. Robustness is further enhanced through Fraiman-Muniz functional depth, which down-weights anomalous trajectories that would otherwise distort covariance estimation. Uncertainty is quantified through both analytic and nonparametric bootstrap confidence intervals.

We evaluate the method on two longitudinal datasets across multiple taxonomic resolutions and missingness rates ranging from 10% to 60%, under MCAR, MAR, and MNAR mechanisms, benchmarking against DeepMicroGen. The framework achieves lower mean absolute error at substantially reduced computational cost, with improvements most pronounced at high missingness rates.

Co-authors: Mohammad Darbalaei, Daniel Hoffmann, Farnoush Farahpour

Please login to see details