WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.
WAWABILITY July 11–12, 2025 Washington DC. Big ideas. Bold Progress. Global Impact. Powered by TDIforAccess.

C-P.51: Tokenization-Aware Protein Language Modeling with Evolution-Guided Units and Dynamic Biological Knowledge Fusion

[sponser-meet-now-chat][/sponser-meet-now-chat]

Protein language models (PLMs) have become key tools for computational protein analysis, but most still rely on single amino-acid tokens as the default representation unit. While simple and effective, this choice can limit biological expressiveness and may interact in non-trivial ways with downstream adaptation strategies. We present a tokenization-aware framework that compares character-level tokenization, frequency-based subword tokenization such as BPE and Unigram, and PUMA, a mutation-aware tokenizer that organizes sequence patterns into mutation-informed token families.

Using multiple pretrained backbones, we study various strategies for implementing tokenization in protein language modeling, including frozen encoders, full fine-tuning, and parameter-efficient LoRA. Experiments cover representative tasks in computational biology, including subcellular localization, protein function classification, and stability prediction. Beyond accuracy, we assess memory usage, sequence compression, compute cost, and interpretability. We further examine how token-level representations can be enriched by integrating biochemical, functional, and evolutionary priors through task-aware knowledge fusion. All in all, we provide an analysis of the effect of tokenization on protein representation learning performance across computational and biological dimensions.

Our results indicate that PUMA can outperform frequency-based subword tokenization in several settings, while vocabulary size has a substantial effect on both predictive performance and efficiency. Amino-acid tokenization remains a strong baseline overall; however, at certain vocabulary sizes, PUMA can match or surpass amino-acid-level representations, suggesting that biologically informed tokenization can offer practical advantages when appropriately configured.

Overall, this work provides a structured comparison of tokenization choices, adaptation regimes, and representation quality in PLMs. By analyzing their combined effects across multiple biological tasks, our framework offers practical guidance for selecting tokenization and fine-tuning strategies and introduces an extensible benchmark for biologically informed protein sequence modeling.

Co-authors: Burak Suyunu, Özdeniz Dolu, Arzucan Özgür

Please login to see details