B-T.23: Robust Explainable Regression for Noisy Omics via a Rank-Based Algorithm
Authors
Mario Lauria
Department of Mathematics, University of Trento, Trento, Italy. Fondazione The Microsoft Research – University of Trento Centre for Computational and Systems Biology (COSBI), Rovereto (TN), Italy.
Carlo Alberto Rossi
Department of Cellular, Computational and Integrative Biology (CIBIO), University of Trento, Trento, Italy. Institute of Neuroscience, National Research Council, Padua, Italy.
Roberto Bizzotto
Institute of Neuroscience, National Research Council, Padua, Italy.
Luca Marchetti
Department of Cellular, Computational and Integrative Biology (CIBIO), University of Trento, Trento, Italy.
Keywords
Machine Learning, Explainable AI, Transcriptomic Data Analysis. Biomarker Extraction
[sponser-meet-now-chat][/sponser-meet-now-chat]
In omics analysis, traditional regression struggles with high dimensionality and noise. We introduce BI-SCUDO Regression, an explainable rank-based algorithm for predicting continuous clinical variables. It utilizes a genetic algorithm to extract patient-specific signatures compiled into a biomarker, whose statistical significance is validated by permutation testing (n = 10000), generating distance-weighted predictions.
Performance was compared against LASSO, Random Forest Regression (RFR), and XGBoost, by considering increasingly noisy synthetic datasets. In addition, TCGA transcriptomics (n=3464) was used to test the methods in a real-world scenario, trying to predict overall survival across 32 cancer projects stratified by cancer stage (whole dataset, stages I-IV, II-IV, and III-IV). Predictive power was evaluated using 10-fold cross-validation R² and the Ratio of Performance to Deviation (RPD). Feature selection was assessed via Fisher's exact test for cancer-gene enrichment.
In synthetic datasets, BI-SCUDO demonstrated superior robustness (p≤0.0003), maintaining a median R² of 0.78 at maximum noise, whereas LASSO collapsed (R²≈0) and RFR and XGBoost degraded to 0.70 and 0.69. On TCGA data, BI-SCUDO achieved significant biomarkers in 95% of the datasets (p < 0.05) and was the only method yielding positive median R² values (0.23–0.25) across all stratifications. XGBoost was the only method identifying an equal or greater number of significantly cancer-gene-enriched biomarkers compared to BI-SCUDO. However, given that XGBoost biomarkers are on average twice as long as those from BI-SCUDO and XGBoost's poor predictive performance, this result supports BI-SCUDO as the method that best balances biological relevance with predictive stability.
Co-authors: Mario Lauria, Carlo Alberto Rossi, Roberto Bizzotto, Luca Marchetti
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826