A twelve-slide breakdown of the Nucleotide Transformer, the origin of the NT line: how a masked language model reads DNA in 6-mer tokens, what it learns about introns, coding regions and variants with no supervision, why it matched SpliceAI on splice-site prediction, and how version 2 reached the top benchmark score with a model ten times smaller than the 2.5-billion one.
Read the paper: Nucleotide Transformer: building and evaluating robust foundation models for human genomics (Nature Methods, 2025).