WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.
WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.

C-S.B.77: A leakage-aware benchmark for gene prioritization across feature representations and integration strategies

Author

[sponser-meet-now-chat][/sponser-meet-now-chat]
Gene prioritization (GP) methods are increasingly used to rank candidate disease genes from heterogeneous biological data, but reported performance remains difficult to compare across studies because methods are evaluated using different feature sets, training protocols, and metrics, and may also be affected by information leakage. We present a unified, time-aware benchmarking study that evaluates classical and modern machine-learning approaches for GP, including kernel methods and deep neural networks. We further benchmark multiple gene-level feature modalities under a prospective evaluation protocol and compare data integration strategies at multiple levels across both deep-learning and kernel-based settings. To better reflect practical discovery settings, we evaluate performance using metrics that emphasize both global ranking and early retrieval, and introduce BioSim, a biological enrichment-based similarity metric that assesses the functional coherence between top-ranked unlabeled genes and known disease genes. Across 48 diseases from 11 ICD-10 categories, our benchmark shows that classical methods such as SVMs remain highly competitive with deep neural networks, and that feature type and feature quality strongly influence model performance. Early fusion can produce mixed results, whereas intermediate fusion in both deep-learning and kernel-based settings, particularly geometric kernel fusion, yields the most consistent gains across ranking metrics. For example, in Diseases of the Blood and Certain Disorders, intermediate fusion substantially improved early retrieval, more than doubling Recall@150 relative to the best corresponding non-fused setting. Late fusion, meanwhile, more often gives stronger BioSim performance. Overall, this work provides a leakage-aware benchmark and practical guidance for evaluating GP methods in realistic discovery scenarios. Co-authors: Fatemeh Ghorbani, Ole Christian Lingjærde, Pooya Zakeri

Please login to see details