Every model in the landscape figure is one dot with a tooltip. That is enough to compare them and not enough to use one. These cards are the other half: the same ten fields for every model, in the same order, so you can read one without having read any of the others.
The colour on the tokenization row is the same colour the model carries on the landscape chart, which makes each card a zoom-in on that figure rather than a separate thing to learn.
Two of the fields matter more than the rest. Headline result gives the metric and the condition it was measured under, because the number on its own does not tell you much. What it does not do is the field summaries usually leave out, and it is often the one that tells you whether the model fits your problem.
On Evo 2 specifically, one number gets misattributed constantly and it is the one on the card’s correction row. The 0.95 AUROC quoted for BRCA1 variant classification comes from a supervised logistic regression trained on Evo 2 embeddings, not from Evo 2 running zero-shot. The paper is explicit that the supervised model outperformed zero-shot Evo 2 on that test set. The zero-shot claim that does hold is narrower and still strong: on BRCA1 noncoding variants, Evo 2 outperformed every other model tested.
The generation results need their hedge kept too. Evo 2 produced mitochondrial, prokaryotic and eukaryotic sequences at genome scale, including 580,000-base generations prompted from a fragment of the M. genitalium genome, and nearly 70% of the generated genes carried a significant Pfam hit. None of it was built. The paper notes the generations lack some essential genes and does not claim they are functional or able to replicate.
The chromatin experiment is the one that was tested in cells. Guided during generation by two separate predictors, Evo 2 wrote sequences whose accessibility pattern spelled EVO2 in Morse code in mouse embryonic stem cells. That is a pattern of open and closed chromatin, not letters written into the DNA itself.
Against Evo 1, the changes are scope and scale: all domains of life instead of prokaryotes only, 1 million tokens of context instead of 131,072, and StripedHyena 2 instead of StripedHyena.
Numbers are from the Nature paper. Sizes are the largest reported checkpoint, and release is the first public preprint, matching the conventions used on the landscape figure.