Identifying the genetic variant that drives a disease, and the gene it acts through, sits at the foundation of most drug target selection decisions. For most trait-associated variants, that identification requires interrogating noncoding DNA: the 98 percent of the genome that regulates gene activity but does not encode proteins. Running a computational model across every candidate variant at that scale demands graphics processing unit (GPU) infrastructure that most academic groups and small biotechs cannot access on demand.
Google DeepMind released AlphaGenome Atlas to address that bottleneck. Built on AlphaGenome, the DNA sequence model described in Nature in January 2026, the Atlas precomputes molecular effect predictions for all 9 billion possible single-nucleotide variants (SNVs) in the human genome and serves them through a no-code web portal, free for non-commercial research.
AlphaGenome Atlas: a precomputed noncoding variant database for drug discovery
AlphaGenome Atlas contains precomputed molecular effect predictions for all 9 billion possible SNVs in the human genome, plus over 100 million short insertions and deletions drawn from gnomAD, UK Biobank, and the All of Us Research Program. For each variant, it stores an average of 27,000 experiment-specific scalar predictions spanning hundreds of human and mouse cell types and tissues. The full dataset runs to approximately 1 petabyte, more than 30 times the size of the AlphaFold Database.
Before the Atlas, noncoding variant prioritization required either running AlphaGenome on demand against each candidate, outsourcing to a specialist bioinformatics group, or accepting the coarser signals of older tools. Drug discovery teams working from whole-genome sequencing data typically flag thousands of candidates; narrowing that list is where the compute bottleneck bites. The Atlas converts that inference step into a free lookup with no coding required.
The Atlas also includes a genome-wide catalogue of over 2,500 recurrent DNA sequence motifs, the short sequence elements that transcription factors bind, mapped to their positions across the full genome. Available resources include:
- AVI score: a single combined impact number per variant, covering both coding and noncoding regions
- AVI feature attributions breaking down contributions from splicing, gene expression, chromatin accessibility, conservation, and protein impact
- Access via web portal (no coding required), AlphaGenome API, and as a skill in Google Antigravity
What does the AlphaGenome Variant Impact score measure?
The AlphaGenome Variant Impact (AVI) score condenses per-variant predictions from AlphaGenome, AlphaMissense (DeepMind's tool for protein-altering variants), evolutionary conservation data, and protein loss-of-function signals into a single number. It follows a PHRED-like scale: a score of 10 puts a variant in the top 10 percent of predicted impact, while a score of 30 reaches the top 0.1 percent. AlphaMissense covers only missense changes in protein-coding genes; the AVI score extends that interpretive reach across the full genome, coding and noncoding alike.
Each AVI score carries feature attributions that identify the most disrupted process: RNA splicing, gene expression, chromatin accessibility, or protein function. For a team generating a target hypothesis around a regulatory element, that breakdown turns a ranking into an actionable mechanistic lead.
AVI scores outperform existing tools across rare-disease benchmarks
On solved cases from the Genomics Research to Elucidate the Genetics of Rare diseases, the GREGoR Consortium, a US National Institutes of Health programme, the AVI score placed the known causal variant among the top 50 candidates in 29.5 percent of cases. The most widely used genome-wide scoring tool, Combined Annotation Dependent Depletion (CADD), reached 12.5 percent on the same benchmark. That roughly 2.4-fold difference matters at the front end of any target identification workflow.
| Approach | Genome coverage | Noncoding variants | Access model |
|---|---|---|---|
| Integrated multi-signal scoring (AVI) | Full genome (coding + noncoding) | Yes | Free non-commercial; commercial planned |
| Protein-coding pathogenicity tools | Coding regions only | No | Varies; most free for research |
| Genome-wide deleteriousness scores (CADD) | Full genome | Yes | Free |
| Splice-site-specific predictors | Splice-site regions only | Partial | Varies |
| Laboratory-developed variant pipelines | Depends on institution | Varies | Internal only |
At the University of Exeter, statistical geneticist Gareth Hawkes applied Atlas predictions to 54,000 whole genomes from the UK Biobank, grouping rare variants by their predicted molecular effect to surface noncoding associations with blood protein levels. The analysis uncovered 22 percent more noncoding genetic associations than prior methods found in the same data. Noncoding variant-to-phenotype signals at that scale, generated without per-query GPU inference, are exactly what complex disease target discovery requires.
At the Broad Institute of MIT and Harvard, Laura Covill and Anne O'Donnell-Luria applied AVI scores to unsolved rare disease cases and identified a deep intronic variant in DNM1, a gene associated with epileptic encephalopathy. The model flagged the variant as creating an aberrant splice site; experimental validation later supported that call. The Atlas makes computational ranking followed by targeted functional validation practical at a scale most labs could not previously attempt.
AlphaGenome Atlas has known gaps in oncology and clinical contexts
DeepMind has not validated AlphaGenome Atlas for clinical use; the accompanying preprint positions it as one input within a broader diagnostic evidence chain, not a standalone basis for clinical decisions. Training data gaps include missing cell types and non-polyadenylated RNAs, and tissue labels reflect broad ontological categories rather than the specific cellular states where regulatory elements can behave differently.
The gaps are most consequential in oncology. AlphaGenome trained on germline reference sequences from general population data, so tumor-specific regulatory contexts shaped by somatic variation and clonal evolution fall outside its scope. Near the TAL1 oncogene in T-cell leukemia, predictions diverged from observed RNA and protein patterns because a short oncogenic isoform was absent from the training annotation.
Kristian Helin, CEO of the Institute of Cancer Research in London, described AlphaGenome as a major advance in computational genomics in response to the January 2026 Nature paper, while noting that capturing cell-type-specific regulation remains an important challenge. The Atlas research remains a preprint pending peer review, and DeepMind describes the current release as a first-generation baseline with broader cell-type coverage targeted for future iterations. Used as a prioritization signal feeding into experimental follow-up rather than a standalone endpoint, AVI scores already represent a meaningful step forward for labs that previously had no practical route to genome-wide noncoding variant ranking.
Commercial access via Google Cloud will determine Atlas adoption in pharma
DeepMind makes AlphaGenome Atlas free for non-commercial research today and plans commercial access through Google Cloud, with no pricing or timeline announced. A queryable commercial tier would let pharma and biotech teams embed Atlas lookups directly in existing cloud informatics pipelines, sidestepping the need for bespoke variant scoring infrastructure. Isomorphic Labs, DeepMind's drug-discovery sister company, plans to draw on the Atlas in its own pipeline, and DeepMind describes the current release as a first-generation baseline with broader cell-type coverage and improved rare-isoform representation targeted for future iterations.
This article is based on a press release issued by Google DeepMind and was produced under Drug Discovery News' AI Editorial Guidelines.










