Cheat sheet

gLM cheat sheet: Evo

One reference card for Evo: architecture, tokenization, size, context, training data, what it was shown to do and what it does not do. One card in a set that covers one genomic language model at a time.

GenomicsLLMMachine Learning

Genomic language model · Cheat sheet

Evo

The first genomic language model whose designed sequences were built and tested in a lab

First public
bioRxiv, February 2024
Published
Science, November 2024
Lab
Arc Institute, Stanford, TogetherAI
Architecture
StripedHyena. Hyena convolutions with a few attention layers. First genomic model to use it; the architecture itself is Poli et al. 2023.
Tokenization
Single nucleotide, byte-level tokenizer
Parameters
7B
Context
131,072 tokens. Pretrained at 8,192, then extended in a second stage.
Training data
OpenGenome. 300 billion nucleotides from prokaryotic genomes, phage and plasmid sequences.
Used for
Prediction and generation, from molecular to genome scale

Headline result

Zero-shot function prediction competitive with domain-specific language models, without task-specific training. Generated CRISPR-Cas and IS200/IS605 transposon systems whose functional activity was experimentally validated: a synthesized Evo-generated sgRNA showed in vitro cleavage activity comparable to SpCas9.

What it does not do

Prokaryotes only. No eukaryotic genomes in training, which is exactly what Evo 2 changed.

Commonly confused

Often mixed up with Evo 2. Evo holds 131,072 tokens in context; the 1 million token context belongs to Evo 2. Evo does generate sequences over 1 megabase, but those were assessed computationally, not built.

Download this card PNG, 1080 x 1350, sized for a post or a slide