WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.
WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.

B-S.B.15: Assessing the reliability of viral host prediction models: A comparison of temporal and random data partitioning

Author

[sponser-meet-now-chat][/sponser-meet-now-chat]
Rapid viral host identification is critical for public health, yet discovery often relies on high-throughput sequencing reads or partial contigs. While numerous machine learning (ML) prediction tools exist, performance is typically evaluated using random cross-validation splits of training and test data. These methods fail to address real world challenges, such as temporal viral evolution and potential host shifts, where models must predict outcomes for future sequences using only historical data. Furthermore, many existing methods lack transparency regarding database curation and feature engineering, impeding reproducibility and standardization. This study investigates the performance of host prediction models by comparing traditional random partitioning against a temporal split strategy that more accurately reflects realistic scenarios. We evaluated the performances of random versus temporal split models using viral genome sequences from family Coronaviridae. K-mer frequencies were extracted using simulated NGS reads and refined via distance-based dimensionality reduction. We assessed the effects of different normalization schemes and centroid vectors on feature robustness, subsequently, benchmarking multiple supervised ML classifiers to determine the optimal virus-host prediction framework. Preliminary analyses compare predictive accuracy of ML models across two strategies: a random, accession- and host-stratified split and a temporal split. We anticipate that the temporal split will provide a more rigorous assessment of the model generalizability on future unseen data, whereas the random split will identify patterns within known evolutionary clusters. These results will identify the optimal combination of normalization, centroid selection, and model architecture to determine the ideal framework for a hierarchical classification pipeline. Co-authors: Sergej Ruff, Martin Ludlow, Klaus Jung

Please login to see details