A twelve-slide breakdown of GROVER: how byte-pair encoding lets the genome pick its own 601-token vocabulary, why next-k-mer prediction chooses the best one, how masked token prediction teaches it sequence context, and what the trained model reveals about genes, repeats, chromatin and lexical ambiguity, unsupervised, from the human genome alone.
Read the paper: DNA language model GROVER learns sequence context in the human genome (Nature Machine Intelligence, 2024).