DeepMind Releases Petabyte-Scale AlphaGenome Atlas for Genetic Variant Research
The Google subsidiary has published precomputed predictions for all 9 billion possible single-letter mutations in the human genome, with commercial access planned via Google Cloud.

Google subsidiary DeepMind has published the AlphaGenome Atlas, a massive one-petabyte dataset containing predictions for every single-letter variation across the human genome. Comprising predictions for all 9 billion potential point mutations, the repository has been made available at no cost for academic research.
The release integrates predictions from two models—AlphaGenome and AlphaMissense—to generate a consolidated metric known as the AlphaGenome Variant Impact score. This framework assesses both the 2% of human DNA responsible for coding proteins and the remaining 98% of non-coding genetic material. In addition, the release includes an index featuring more than 2,500 recurring DNA motifs. The overall dataset expands on previous computational biology efforts, coming in at more than 30 times the scale of the AlphaFold database, which originally mapped over 200 million protein structures from roughly 190,000 experimental baselines.
Initial research trials have already yielded results using the dataset. Gareth Hawkes, a Medical Research Council fellow located at the University of Exeter, applied the score to whole-genome sequences from over 54,000 subjects in the UK Biobank. His analysis uncovered 22% more non-coding genetic links than previously recognized. By isolating the top 1% of variants identified as most influential by the Atlas, Hawkes discovered 19 specific genomic locations connected to body mass index.
Additional validation occurred in the United States, where a research team led by Laura Covill at the Broad Institute utilized the metric to evaluate mutations in the DNM1 gene. The model flagged a specific single-letter edit that generates an abnormal splice site, an outcome that was subsequently validated through biological testing in the laboratory. The DNM1 gene plays a key role in severe neurological conditions, specifically epileptic encephalopathy.
While the data is currently accessible to non-commercial researchers without fee, DeepMind plans to roll out commercial availability through Google Cloud, which already hosts a portfolio of specialized scientific artificial intelligence tools. Pricing details for enterprise deployments have not yet been disclosed, as first reported by The Next Web.
The enterprise rollout highlights ongoing discussions around regulatory compliance for scientific machine learning models. Under the European Union's AI Act, models created exclusively for scientific research and development are granted exemptions. However, an analysis published in npj Digital Medicine in January highlighted the complexities of applying this rule, noting that project objectives frequently shift over time and that non-profit academic studies regularly intersect with commercial initiatives.
By bringing the dataset to Google Cloud, DeepMind follows a deployment structure established with earlier biological platforms. The commercialization strategy positions Google's cloud infrastructure as a central distributor for high-volume computational biology resources as enterprise life sciences firms evaluate AI-driven research workflows.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



