Cheat sheet

gLM cheat sheet: Evo 2

One reference card for Evo 2: architecture, tokenization, size, context, training data, what it was shown to do and what it does not do. One card in a set that covers one genomic language model at a time.

GenomicsLLMMachine Learning

Genomic language model · Cheat sheet

Evo 2

Trained across all domains of life, with a million-token context and generation at genome scale

First public
bioRxiv, February 2025
Published
Nature, March 2026
Lab
Arc Institute, Stanford, NVIDIA
Architecture
StripedHyena 2. Three kinds of convolution that adapt to the input, plus attention, with each layer working at its own range.
Tokenization
Single nucleotide, byte-level tokenizer
Parameters
40B
Context
1,000,000 tokens at single-nucleotide resolution.
Training data
OpenGenome2. 9.3 trillion tokens from bacteria, archaea, eukaryotes and bacteriophage. Eukaryote-infecting viruses were excluded.
Used for
Prediction and generation, from single variants to genome scale

Headline result

Predicts the impact of many classes of genetic change with no task-specific training, and beat every model tested on noncoding variants in the cancer gene BRCA1. Generated mitochondrial, bacterial and eukaryotic sequences at genome scale, nearly 70% of the genes in the bacterial ones matching a known protein family.

What it does not do

Generated genomes are not shown to work. The paper notes they lack some essential genes and does not claim they are functional or able to copy themselves.

Commonly confused

The 0.95 AUROC often quoted for BRCA1 comes from a separate classifier trained on Evo 2's features, not from Evo 2 alone.

Download this card PNG, 1080 x 1350, sized for a post or a slide