Google DeepMind publishes AlphaGenome Atlas, predictions for 9 billion single-letter DNA variants
- Google DeepMind released AlphaGenome Atlas, a searchable database of predicted molecular effects for every possible single-nucleotide variant in the human genome.
- The AlphaGenome AI model pre-calculated regulatory impacts for 9 billion single-letter DNA changes, producing a dataset of about 1 petabyte.
- The Atlas assigns an AlphaGenome Variant Impact (AVI) score that combines coding and non-coding predictions to rank variants for follow-up research.
- Broad Institute researchers used AVI to prioritize a DNM1 variant in an unsolved rare-disease case, where the model predicted creation of an incorrect splice site.
- Using data from more than 54,000 UK Biobank participants, Gareth Hawkes reported 22% more non-coding associations and found 19 BMI-linked regions among the top 1% of predicted-impact variants.
Hacker News opinions
I do not know the field well enough to judge its utility, but the non-commercial terms make me wonder whether DeepMind plans to sell access to pharma. My understanding is that Isomorphic Labs may already be the commercial route.
I think individual SNP predictions have little direct drug-discovery value. Pharma might license this mainly because it does not want to miss something, while diagnostics look like the more plausible use.
The announcement reads like a cache release, while the underlying AlphaGenome model and whether its predictions deserve trust are barely discussed. The January Nature paper is where I would look for that.
I cannot tell whether this is a major biology result or a PR-driven convenience layer. If it matters, the field should make that clear within a few months.
This may not be new science, but it makes several Google and DeepMind resources less painful to use. I can program, yet molecular biologists who cannot may benefit, and they may also avoid treating Claude's bad output as reliable.
I do not see this as advanced biology by itself. It is a prediction system that may improve the accuracy or accessibility of biological inference, which is a different claim.
I would like to know whether it accepts an indel VCF file. The announcement only talks about single-nucleotide variants.
I would not expect this to turn a 23andMe result into a pathogenic-mutation screen. Those services assay a small selected set of SNPs rather than sequence an entire genome, and their findings often need clinical confirmation.
Katie Pollard's ISMB talk argued that existing human variation does not supply enough context to infer variant effects. Comparative species data and much larger mutagenesis experiments may be necessary.
Variant effects also do not generally occur in isolation. GWAS mainly captures small additive effects because cohorts of even one million people are marginal for studying epistasis.
The promoter question is missing from the post. Non-coding coverage should include promoters, but I want to know whether the atlas can query joint probabilities for promoter sequences and target-protein occurrence in human genomes.
I do not think Demis Hassabis was central to this work. Žiga Avsec developed Enformer, one of the first useful sequence-to-function models, and then AlphaGenome.
A mutagenesis study of a simple virus reportedly tested what this atlas predicts in silico, and several specialized AI models predicted the outcomes poorly. If effects are hard to predict for that virus, human variant predictions need careful validation.
I am concerned that an ad company controls this kind of scientific resource. Claims of access also ring hollow if the database is gated to elite institutions and private businesses.