DeepMind precomputed all nine billion single-letter mutations in the human genome
AlphaGenome Atlas is a petabyte of predictions, thirty times the AlphaFold Database, free for academic use and carrying a line saying it is not approved for any clinical use.

Google DeepMind published AlphaGenome Atlas on September 8, a database of predicted molecular effects for all 9 billion possible single-letter changes in the human genome. It runs to a petabyte, more than thirty times the size of the AlphaFold Database, and it is free to use for academic research through a web portal, an API and a skill in Google Antigravity.

We’re launching AlphaGenome Atlas: an AI-powered searchable database mapping the predicted impact of all 9 billion possible single-letter DNA changes.
Here’s how it could help researchers better understand our biology 🧵
Nine billion is not a sample. Do the arithmetic and it is three alternative letters at each of roughly three billion positions, which is the entire space. Nobody had to choose which mutations were worth predicting.
Because they all were.
But the piece that will matter most is a single number. The AlphaGenome Variant Impact score folds AlphaGenome's regulatory predictions and AlphaMissense's protein predictions into one figure per variant, so a researcher can rank nine billion mutations by how much damage they probably do. DeepMind says the AVI performs best in class across variant pathogenicity and rare disease benchmarks, and pairs each score with attributions showing which process — splicing, chromatin accessibility, conservation — drove it.
And it has already found things. Working with the GREGoR Consortium, researchers at the Broad Institute used AVI to prioritise variants in unsolved rare disease cases and surfaced one in DNM1, a gene tied to epileptic encephalopathy, which turned out to create an incorrect splice site that extended the resulting protein. Experimental screens confirmed it (which is the sentence that separates this from a demo). At Exeter, Gareth Hawkes ran the Atlas against whole-genome data from more than 54,000 UK Biobank participants and pulled out 22 percent more non-coding associations than the statistics alone would surface, including regulatory variants driving PLA2G7 and EGLN1.
That 22 percent is the number to hold onto. Most trait-associated variation sits in the 98 percent of the genome that does not code for protein, where the signal drowns in harmless noise, and a ranking that lifts a fifth more real associations out of that noise is a working instrument rather than a demo.
Now the sentence at the bottom of the page. "AlphaGenome has not been validated for, and is not approved for, any clinical use."
Our read is that the disclaimer will not survive contact with the AVI score, and that DeepMind knows it. A single interpretable number attached to every possible mutation is precisely the artifact that leaks out of research and into decisions, the way a p-value did. The Atlas gives you 111 kilobytes of prediction per variant if you divide the petabyte by the count, and exactly one digit that anyone will actually read.
Would you want your variant of uncertain significance ranked by a model that has never been near a clinical trial? We would want to know the false-negative rate first, and that is not on the page.
It rarely is.
One petabyte is also a fact about access, not just size. Nobody is downloading this. Every researcher who uses the Atlas queries it where Google keeps it, non-commercially for now and on Google Cloud for commercial use soon, which makes the most comprehensive map of human genetic variation a hosted service. So AlphaFold's database went the same way, and the field decided it was fine. This one is thirty times larger and predicts function rather than shape, and function is what people act on.
