# AlphaGenome author roundtable

Source: https://www.youtube.com/watch?v=V8lhUqKqzUc
Recap page: https://rapidrecap.app/video/V8lhUqKqzUc
Generated: 2026-01-28T12:32:14.346+00:00

---
## Quick Overview

The AlphaGenome team discusses the motivation, model architecture, and future direction of their recently released tool, which uses AI to predict the functional impact of genetic variants across millions of base pairs and multiple species, emphasizing the need for high-resolution, multi-modal predictions to better understand gene regulation.

**Key Points:**
- AlphaGenome is a unified DNA sequence to function prediction model released in Nature that predicts the functional impact of genetic variants across millions of base pairs and millions of species.
- The team specifically focused on predicting effects in non-coding regions (the 98% of the genome) which was historically difficult due to complexity and sparse data.
- The model was built using a multi-modal approach, combining sequence data with contact maps and splicing information to capture intricate regulatory interactions.
- A key technical challenge overcome was efficiently processing long DNA sequences (up to 100,000 base pairs) by using a 1D representation that allowed for parallel processing across TPUs.
- The evaluation strategy involved two main parts: checking performance against known experimental data and ensuring the model could accurately predict outcomes for unseen data (like novel splice sites or promoter effects).
- The team is excited about future applications, including predicting the effect of variants on gene expression across different cell types and tissues, and making the model more accessible via an API.
- The development process involved close collaboration, with different team members specializing in different modalities (e.g., splicing, enhancers) to integrate them effectively.

![Screenshot at 00:37: The title card for Part 1, "The Big Picture: Why AlphaGenome?", sets the stage for the discussion by framing the foundational motivation for developing the predictive genomic AI model.](https://ss.rapidrecap.app/screens/V8lhUqKqzUc/00-00-37.jpg)

**Context:** This video is an author roundtable discussion following the release of the AlphaGenome paper, likely published in Nature, detailing a new AI model developed by Google DeepMind to predict the functional consequences of genetic variation. The panel consists of five members of the research team: Dhavi Hariharan (Product Manager), Žiga Avsec (Research Scientist), Tom Ward (Software Engineer), Natasha Latysheva (Research Engineer), and Jun Cheng (Research Scientist), who discuss the scientific motivation, the technical challenges of building a model capable of handling massive genomic data, and the future trajectory of the project.

## Detailed Analysis

The roundtable discussion covers the genesis, model design, and future of AlphaGenome. Dhavi Hariharan introduced AlphaGenome as a unified DNA sequence to function prediction model that predicts the functional impact of genetic variants, particularly in the 98% of the genome that is non-coding. Žiga Avsec explained the mission: to build an AI system capable of deciphering the genome's source code, which has profound implications for health, especially for rare genetic diseases where diagnosis remains challenging. Jun Cheng elaborated on his background in predicting genetic mutations, noting that the complexity arises from the need to model interactions between different regulatory elements across long sequences, often requiring trade-offs between sequence length and resolution. Tom Ward discussed the engineering challenge of training the model on massive datasets (40-50 gigabytes per sample) and the realization that they needed to move beyond 1D representations, leading to the development of a model that incorporates 2D modalities like contact maps and splicing information. Natasha Latysheva highlighted that the model's ability to accurately predict effects across different cell types and the impact of variants on complex biological processes like gene regulation (e.g., enhancers and insulators) is a significant advancement over previous single-task models. The team emphasized the rigorous evaluation, which included checking performance on held-out data and ensuring the predictions were biologically plausible across different cell types. Looking forward, the team plans to continue improving predictive accuracy, especially for rare variants, and making the tool more accessible via an API for the scientific community.

### Part 1 The Big Picture

- Why AlphaGenome?: Dhavi Hariharan introduced AlphaGenome as a unified DNA sequence to function prediction model that predicts the functional impact of genetic variants in Nature
- Žiga Avsec emphasized the goal of deciphering the genome's source code for health benefits, especially for rare genetic diseases
- Jun Cheng mentioned his background in predicting genetic mutations and the complexity of modeling interactions across long sequences.

### Part 2 The Model

- What's New and How We Built It: Tom Ward detailed engineering challenges like handling long sequences (100k base pairs) and the need for multi-modal data (contact maps, splicing)
- Natasha Latysheva explained the model's ability to capture complex regulatory interactions and predict effects across cell types, outperforming older models
- The team evaluated the model on held-out data, finding strong performance in predicting outcomes for unseen variants and cell types.

### Part 3 From Model to API

- Opening It Up: Dhavi Hariharan discussed how the team is happy with the model's current state and future steps
- Žiga Avsec mentioned the model's ability to score variants across different tissues and the continuous effort to improve accuracy and generalization
- Tom Ward noted the importance of making the API accessible for large-scale analysis and experimentation.

![Screenshot at 00:05: Dhavi Hariharan, Product Manager at Google DeepMind, introduces the AlphaGenome project.](https://ss.rapidrecap.app/screens/V8lhUqKqzUc/00-00-05.jpg)
![Screenshot at 00:38: Title card for Part 1: "The Big Picture: Why AlphaGenome?"](https://ss.rapidrecap.app/screens/V8lhUqKqzUc/00-00-38.jpg)
![Screenshot at 01:04: The five team members—Tom Ward, Žiga Avsec, Dhavi Hariharan, Natasha Latysheva, and Jun Cheng—seated for the roundtable discussion.](https://ss.rapidrecap.app/screens/V8lhUqKqzUc/00-01-04.jpg)
![Screenshot at 01:16: Tom Ward explains the complexity of predicting effects from DNA sequence and the need for a single model to handle multiple modalities.](https://ss.rapidrecap.app/screens/V8lhUqKqzUc/00-01-16.jpg)
![Screenshot at 06:27: Title card for Part 2: "Part 2 The Model: What's New and How We Built It".](https://ss.rapidrecap.app/screens/V8lhUqKqzUc/00-06-27.jpg)
