WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.
WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.

A-G.10: Short-Context gLM Training for Long-Range Variant Effect Prediction

Author

Kingston University London
[sponser-meet-now-chat][/sponser-meet-now-chat]
Disease susceptibility in humans is frequently driven by mutations within the DNA. Contemporary research has shown that many such mutations lie within the non-coding regions of the genome, several thousands of base-pairs (bp) from the transcription start sites (TSS) of their target genes. In the age of artificial intelligence, Transformer-based genomic language models (gLMs) are commonly used to interpret variant effects. Their ability to exploit the large datasets produced by next-generation sequencing, and their aptitude for modelling long-range interactions within sequences, makes them ideal for the task. However, the quadratic scaling of the attention mechanism with context length results in high computational resource consumption, and inefficiency when training gLMs on very long (10000bp+) sequences. DroPE (Dropping the Positional Embeddings of LMs after training), originally developed for generative text-based large language models (LLMs), provides a method for extending the context of pretrained LLMs without the need for long-context fine-tuning. This paper adapts and applies DroPE to BERT-based gLMs, and demonstrates that it can extend the context of models pretrained on short genomic sequences. Furthermore, models incorporating DroPE demonstrate an enhanced ability to model long-range context, achieving competitive performance on variants far (>35kbp) from the TSS without the need for computationally expensive long-context pretraining and fine-tuning. Code and fine-tuned models are available at: https://github.com/meghegde/gLM-DroPE. Co-authors: Jean-Christophe Nebel, Farzana Rahman

Please login to see details