A-G.38: Embedding-based statistical framework enables quantitative gene function representation and hypothesis testing with large language models
Accurately delineating gene function is central to interpreting high-throughput genomic data, yet most existing methods depend on predefined gene sets and largely qualitative interpretations. Large language models (LLMs) provide a promising alternative by extracting functional relationships from biological text, though their ability to support quantitative analysis has not been fully established. We introduce a statistical framework that leverages LLM-derived embeddings of genes and biological functions to enable quantitative assessment of gene-gene and gene-function relationships across diverse biological settings. We evaluated seven leading embedding models using both curated gene annotations and literature-derived descriptions. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 showed the strongest performance, recovering gene-gene relationships in up to 98% of Gene Ontology biological processes and approximately 99% of canonical pathways. Gene-function association analyses further demonstrated high sensitivity (95-98%) and specificity (73-84%). Importantly, this framework enables rigorous testing of functional hypotheses generated from gene lists without requiring predefined annotations. It consistently differentiates biologically coherent gene sets from noise, surpassing both confidence-based LLM outputs and traditional enrichment methods. Application to drug response data uncovered candidate pathways linked to cancer immune sensitization and supported systematic exploration of drug mechanisms. Overall, our results position LLM-based embeddings as a scalable quantitative tool for functional genomics.
Co-authors: Yanhao Tan, Li-Ju Wang
Contact Attendee
Warning: Attempt to read property "user_email" on string in /home/1276969.cloudwaysapps.com/ydbgzhdjeq/public_html/wp-content/plugins/my-conference-now/functions.php on line 2826