Abstract
Large Language Models (LLMs) are increasingly applied to genomic tasks, yet core challenges remain concerning tokenization, evaluation, and data scarcity. This study focuses on promoter classification and systematically evaluates four tokenization methods: non-overlapping 6-mer, overlapping 6-mer, Byte Pair Encoding (BPE), and WordPiece (WPC). We show that the commonly used k-mer approach, specifically the non-overlapping variant, outperforms BPE and WPC across eight organisms, challenging assumptions derived from natural language processing. To ensure robustness, we evaluated performance under two distinct negative data strategies: positive-promoter-shuffled and random-non-promoter-fragments. Using a positional SHAP framework, we demonstrate that the model learns biologically plausible positional patterns rather than exploiting artifacts from these negative data generation processes. Furthermore, evolutionary-informed transfer learning experiments and external validation on an unseen organism reveal that training on phylogenetically related species significantly improves performance, particularly in low-data regimes. These findings underscore the significant impact of tokenization and negative data design, providing practical guidance for refining genomic classifiers.
| Original language | English (US) |
|---|---|
| Article number | lqag025 |
| Journal | NAR Genomics and Bioinformatics |
| Volume | 8 |
| Issue number | 1 |
| DOIs | |
| State | Published - Mar 1 2026 |
| Externally published | Yes |
All Science Journal Classification (ASJC) codes
- Structural Biology
- Molecular Biology
- Genetics
- Computer Science Applications
- Applied Mathematics
Fingerprint
Dive into the research topics of 'Optimizing genomic language models for promoter prediction: a comparative study of tokenization and cross-species learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver