Predicting traits from DNA
We cannot yet reliably say how a person's DNA turns into their traits and disease risks.
Open in the interactive tree →Genome-wide studies link thousands of DNA variants to traits, but each has a tiny effect, and genes act through cell type, environment and chance. Reading a genome and predicting the person is the central open problem of genetics.
As of October 2026
Polygenic risk scores reach modest accuracy for many diseases (AUC about 0.63 to 0.64 from genes alone for breast cancer and coronary disease in European-ancestry studies) and perform worse for people of non-European ancestry; embryo screening with them has been offered since 2019 and is controversial. DeepMind's AlphaGenome (launched 25 June 2025, published in Nature on 28 January 2026) predicts how variants change gene regulation from up to one million DNA letters, but it is weaker for effects over 100,000 letters away and was trained only on human and mouse data. An 'AlphaGenome Atlas' with predicted effects for nearly all possible single-letter changes in the human genome was released in September 2026.
What is missing
- Genomes plus detailed health data from millions of people of all ancestries
- Experimental measurement of what each variant does in the right cell type (massively parallel assays, perturbation screens)
- Models of gene-regulation networks that capture interactions between variants and with the environment
- Data on the non-coding genome and on epigenetic state over a lifetime
- Causal methods that separate real effects from mere correlations
Becomes possible once solved
- Accurate disease-risk prediction from birth
- Designing gene therapies for the exact variant that causes a disease
- Digital twins of individual patients
Open steps
- Effects of non-coding variants High AI leveragePredict what a single-letter DNA change does to gene regulation in each cell type, including enhancers over 100,000 letters from their gene.
- Lab data for variant effects Medium AI leverageMassively parallel assays and perturbation screens measure what each variant does in the right cell type, but cover only a small share of variants.
- Risk scores for all ancestries Medium AI leveragePolygenic scores lose accuracy in people of non-European ancestry because most studies enrolled Europeans.
- Causal variants versus correlation Medium AI leverageSeparate variants that cause a trait from those that merely travel with them, and link each to a gene and a cell type.
- Gene and environment together Low AI leverageModels of how variants interact with each other, with epigenetic state and with the environment over a lifetime are largely missing.
Where AI could help
Medium AI leverage. AI leads on non-coding variant effects, but lab data, diverse cohorts, causality and gene-environment effects still limit trait prediction.
- Sequence-to-function models that score non-coding variants in each cell type
- Better polygenic scores that combine deep-learning variant effects with biobank data
- Active learning to choose which variants to test in massively parallel assays
- Causal fine-mapping and cell-type annotation across large biobanks
Shown so far
- In September 2023 Google DeepMind's AlphaMissense classified 89% of all 71 million possible missense variants as likely pathogenic or likely benign at a 90% precision threshold (Science). source
- In January 2026 Nature published AlphaGenome, which predicts variant effects on gene regulation from 1 million DNA letters and matched or beat the best external models in 25 of 26 variant-effect evaluations. source
Prerequisites
- Modern evolutionary synthesis1942Fisher's quantitative genetics is the basis of predicting traits from genes
- Genome-wide association studies2005Genome-wide association studies produced the variant-to-trait links the problem is built on
- Deep Learning2012
- Single-cell atlases2016
- Ancient DNA & Human Origins2022Ancient genomes show which variants were selected over time, clues to what they do
- Complete genome and pangenome2022