WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.
WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.

B-S.B.51: CROssBARv2: A Unified Biomedical Knowledge Graph for Heterogeneous Data Representation and LLM-Driven Exploration

[sponser-meet-now-chat][/sponser-meet-now-chat]

Biomedical knowledge required for understanding disease mechanisms and developing effective therapeutics is dispersed across hundreds of databases, ontologies, and publications in heterogeneous and non-standardised formats. Knowledge graphs (KGs) offer a principled framework for integrating such data; however, most existing biomedical KGs remain constrained by limited scope, lack of systematic update pipelines, sparse metadata, and interfaces that are inaccessible to non-technical users. We present CROssBARv2, a large-scale heterogeneous biomedical KG system designed to support systems biology research and drug discovery. CROssBARv2 integrates data from 34 curated sources into a Neo4j graph database comprising 2,709,502 nodes and 12,688,124 relationships across 14 biologically meaningful node types, including proteins, genes, drugs, compounds, diseases, pathways, phenotypes, GO-terms, and more. The system features fully automated pipelines for regular data retrieval and standardisation, ensuring long-term maintainability and scalability. Node embeddings encoding biological features are incorporated to enable vector-based similarity searches and downstream predictive tasks. A key contribution is CROssBAR-LLM, a natural language interface that translates user-submitted biomedical questions into Cypher graph-db queries, executes them against the KG, and returns contextually accurate responses, enabling graph exploration without programming expertise. Systematic benchmarking across multiple datasets demonstrated that CROssBAR-LLM substantially outperforms web-search-augmented LLMs in biomedical question-answering accuracy, effectively eliminating hallucinations through grounding in structured data. Deep learning models trained on CROssBARv2 for protein function prediction achieved state-of-the-art performance, further demonstrating the KG's utility for biological inference. CROssBARv2 is publicly accessible at https://crossbarv2.hubiodatalab.com/llm, with full source code and datasets available at https://github.com/HUBioDataLab/CROssBARv2.

Co-authors: Erva Ulusoy, Melih Darcan, Mert Ergün, Sebastian Lobentanzer, Ahmet Süreyya Rifaioğlu, Dénes Türei, Julio Saez-Rodriguez, Tunca Dogan

Please login to see details