An eleven-slide breakdown of DNABERT-2: why k-mer tokenization held DNA models back, how byte-pair encoding cuts sequence length about 5x and avoids the leakage problem, and how a 117-million-parameter model matched the 2.5-billion Nucleotide Transformer on GUE at roughly 92x less GPU time. It also ships GUE, a standardized multi-species benchmark of 36 datasets across 9 task types.
Read the paper: DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genomes (ICLR, 2024).