WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.
WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.

A-S.B.05: Genomic Feature Transformation for Multi-Omics Integration in Clinical Prediction Models

Author

Keywords

multi-omics, AI, Genomics
[sponser-meet-now-chat][/sponser-meet-now-chat]

Multi-omics integration in clinical prediction models relies on capturing cross-modal interactions to improve prediction beyond single-omics baselines. Incorporating genomics into such models remains non-trivial: genomic data is extremely high-dimensional, sparse in per-variant information content, and typically requires far larger cohorts than other omics layers to yield informative features. As a result, methods for integrating genomics with complementary omics data remain poorly standardised. Using the All of Us (AoU) and Swiss HIV Cohort Study (SHCS) datasets, we trained linear and deep-learning classifiers in single-omic and multi-omic settings to predict Coronary Artery Disease (CAD) and Chronic Kidney Disease (CKD), two conditions with established genetic and multi-omic associations. Genomics data was represented in four ways: (i) raw SNP genotype matrices, (ii) Principal Component Analysis (PCA), (iii) Polygenic Risk Scores (PRS), and (iv) AlphaGenome-derived gene-level functional impact scores. Each representation was evaluated individually and integrated with complementary omics via feature concatenation or a deep-learning encoder. Performance was estimated using nested three-fold cross-validation, with identical patient splits held constant across all genomic representations and their combinations; we report mean macro-averaged F1 score and standard deviation, with paired significance testing against raw SNP matrices. Genomic data alone was consistently outperformed by other individual omics layers. When integrated into multi-omics models, raw SNP matrices and PCA representations frequently degraded performance relative to genomics-free multi-omics models, whereas PRS matched or improved performance across most cohort-disease combinations and significantly outperformed raw SNPs for CAD prediction in the SHCS. Simple feature concatenation performed as well as, or better than, more complex encoder architectures, suggesting that genomic signal is captured in a largely additive rather than complementary manner at current sample sizes. These findings indicate that biologically informed transformations such as PRS, applied prior to integration, offer a practical strategy for incorporating genomics into multi-omics prediction models without requiring larger cohorts.

Please login to see details