On chromosome 19, at position 852,326, a single base change from A to T is enough to cause cyclic neutropenia, a disorder in which the body periodically runs short of the white blood cells it needs to fight infection. The mutation lands inside a Kozak sequence, the short stretch of RNA that tells a ribosome exactly where to start translating a gene, and it replaces a purine sitting three letters upstream of the start codon with a pyrimidine. Two of the field’s standard conservation-scoring tools, PhyloP and PhastCons, both in routine use for nearly two decades, rate this change as close to harmless. A newer tool, built specifically to model the DNA of primates rather than the whole animal kingdom, gets it right.
That primate-only model is one of three versions of a genomic language model called GPN-Star, short for genomic pretrained network with species tree and alignment representations, published this week in Nature1 by a team at UC Berkeley led by Yun Song. The assumption behind PhyloP and PhastCons, and behind most of the field before them, is that more evolutionary distance means more signal: compare a piece of human DNA against hundreds of species reaching back hundreds of millions of years, and whatever never changed is whatever evolution refused to let change. GPN-Star was built to test that assumption directly, training separate copies of the same model on three different slices of the tree of life, one spanning roughly 600 million years of vertebrate evolution, one spanning about 100 million years of mammals, and one confined to other primates. For a surprising amount of human biology, the shortest of those three windows carried the most information, not the longest.
Primates share a common ancestor with humans a little under 65 million years back, roughly a tenth as deep as the vertebrate alignment GPN-Star’s broadest version was trained on.
The human genome was fully sequenced more than twenty years ago, three billion base pairs read end to end, and most of what those letters mean is still unclear. An estimated one to two percent of the sequence codes for protein. The rest is a mix of evolutionary debris, sequence that no longer does anything, and regulatory elements, the switches that decide when a gene turns on, in which tissue, and how strongly. Sorting the debris from the switches, one base pair at a time, is the problem that determines whether a clinician can tell a patient what a mutation in their own DNA actually means, and it is the problem GPN-Star was built to attack.
Song, a professor of computer science and statistics at Berkeley and an investigator at the Innovative Genomics Institute, framed the goal in practical terms. Laboratories have built increasingly creative ways to test what a given mutation does, he said, but they still “cannot experimentally test every single variant in the genome,” so the model’s job is to tell biologists which of the millions of candidates are worth testing first. GPN-Star approaches that goal by training on data that has already been through a phylogenetic sorting process, a whole-genome alignment mapped onto a species tree, rather than on raw, unaligned sequence. Competing models trained the harder way include Evo 2, built on more than 100,000 species spanning every domain of life using 2,000 NVIDIA H100 processors running for months, and Nucleotide Transformer, a 2.5-billion-parameter model trained on 128 processors for a month. GPN-Star, at 200 million parameters, trains in days on eight NVIDIA A100 GPUs.










