Cheat sheet

gLM cheat sheet: Caduceus

One reference card for Caduceus: architecture, tokenization, size, context, training data, what it was shown to do and what it does not do. One card in a set that covers one genomic language model at a time.

GenomicsLLMMachine Learning

Genomic language model · Cheat sheet

Caduceus

The first DNA language model to build strand symmetry and both reading directions into pretraining

First public
arXiv, March 2024
Published
ICML, July 2024
Lab
Cornell, Princeton, Carnegie Mellon
Architecture
MambaDNA. Bi-directional Mamba blocks (BiMamba) made reverse-complement equivariant. No attention, so no quadratic cost in length.
Tokenization
Single nucleotide, character-level
Parameters
7.7M
Context
131,072 bases. The longest sequence length it was pretrained at.
Training data
The human reference genome (HG38). One species.
Used for
Prediction and classification, especially over long ranges

Headline result

Outperforms 10x larger Transformer-based models on many tasks, especially long-range ones. On long-range variant effect prediction it exceeds Nucleotide Transformer v2 (500M), and past 100,000 bases from the gene start it beats Enformer. Best top-1 accuracy on all eight Genomics Benchmark tasks.

What it does not do

Human only, and not a generator. Pretrained on one genome, with a masked objective that fills in bases rather than writing new sequence.

Commonly confused

"10x larger" is the paper's own phrase and refers to Transformer-based models. The abstract ties it to long-range variant effect prediction, against models that lack both bi-directionality and equivariance. It is framing, not a measured ratio: the 500M baseline is far more than 10x this size.

Download this card PNG, 1080 x 1350, sized for a post or a slide