# Session 5: Data in the Age of Generative AI

Source: https://www.youtube.com/watch?v=CAcCAmp_RAw
Recap page: https://rapidrecap.app/video/CAcCAmp_RAw
Generated: 2025-10-30T16:37:41.812+00:00

---
## Quick Overview

The presentation, titled "Data in the Age of Generative AI," covered research from Stanford HAI focusing on two main areas: data created by AI (synthetic data) and data used by AI (training data), projecting that the stock of human-generated public text will be exhausted by 2028 due to LLM consumption. Key takeaways include the successful development of a framework to rigorously evaluate synthetic data, where 83% of experts could not distinguish synthetic from real data, and the immediate next steps involve deploying synthetic data in population health sciences, scaling data attribution for LLMs, and designing mechanisms for potential future GenAI data markets.

**Key Points:**
- The available stock of human-generated public text, estimated to be around $10^{16}$ tokens, will likely be fully utilized by generative AI models by the median date of 2028.
- Recent advances in LLMs and diffusion models are heavily driven by consuming this finite human-generated data, leading to concerns about data scarcity and quality.
- The research team developed a framework to rigorously evaluate synthetic data, showing that 83% of experts could not distinguish between real and synthetic texts in a survey.
- The synthetic data generated by their methods (like the Type 2 Diabetes patient notes example) maintain complex relationships and structure found in real data, scoring a median $r^2$ of 0.9 between theory and experiment on accuracy.
- Next steps include deploying synthetic data with Population Health Sciences, scaling data attribution methods for LLMs, and designing mechanisms for potential GenAI data markets.
- Data attribution methods are crucial for understanding the sources of model 'creativity' and for attributing value (and potentially pricing) to different data sources.

![Screenshot at 00:37: Slide showing the projection that the estimated stock of human-generated public text will be fully used by LLMs around the median date of 2028, highlighting the critical nature of data scarcity.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-00-37.png)

**Context:** This presentation, given by members of the Stanford Human-Centered Artificial Intelligence (HAI) institute, addresses the rapidly growing data demands of large language models (LLMs) and generative AI. The speakers, including James Zou, Daniel E. Ho, and Surya Ganguli (listed as PIs/Co-PIs), discuss the impending exhaustion of high-quality public text data and explore solutions like synthetic data generation and data attribution to manage future AI development responsibly and effectively.

## Detailed Analysis

The presentation began by establishing that data is the fuel for generative AI, noting that recent advances in LLMs and diffusion models are consuming public text data at an exponential rate. A key projection shown in a graph indicates that the estimated stock of human-generated public text will be exhausted by 2028, creating a critical need for new data strategies. The research focuses on two fronts: data created by AI (synthetic data) and data used by AI (training data). Regarding synthetic data, the team developed a new framework to rigorously evaluate its quality, finding that 83% of experts could not distinguish synthetic text from real text, and 92% would use synthetic data in grant proposals. The evaluation of GPT-2 pre-training on the LAMBADA benchmark showed a high correlation ($r^2=0.9$) between theoretical predictions and empirical results, suggesting their methods are effective. The speaker also highlighted that the creativity in diffusion models might stem from their *failure* to perfectly reverse the diffusion process, leading to novel, yet sometimes hallucinatory, outputs. Finally, the team outlined next steps: deploying synthetic data in population health sciences, scaling data attribution techniques for LLMs, designing mechanisms for GenAI data markets, and broadening the analysis of legal/economic implications.

### Introduction and Data Scarcity

- Data is the fuel for AI
- Recent advances driven by consuming public data
- Estimated stock of human-generated public text exhausted by 2028 (median date)
- Data usage by LLMs is accelerating across all scientific domains.

### Synthetic Data Promise and Evaluation

- Synthetic data offers promise for privacy and data-efficient learning
- 83% of surveyed experts could not distinguish real vs. synthetic text
- 92% would use synthetic data for grant proposals
- New framework developed to rigorously develop and evaluate synth data for pretraining.

### Analytic Theory for Creativity/Hallucination

- Diffusion model creativity stems from the *failure* to perfectly reverse diffusion (locality and equivariance biases)
- This leads to patch mosaics—new creative images built from training set patches
- Theory explains hallucinations like 3-legged pants because local patches don't know global context.

### Data Attribution and Marginal Value

- Data attribution quantifies the value of training data by comparing model performance with and without specific data sources (e.g., NYTimes)
- This is hard due to the sheer size of models (e.g., 3B parameters) and high computational cost for counterfactuals.

### Next Steps

- Deploy synthetic data with Population Health Sciences
- Scale up data attribution for LLMs
- Mechanism design for GenAI data markets
- Broader analysis of legal, policy, and economic considerations.

![Screenshot at 00:04: Introduction slide listing the main PIs and Co-PIs for the Hoffman-Yee project on Data in the Age of Generative AI.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-00-04.png)
![Screenshot at 00:37: Graph projecting the exhaustion of human-generated text data stock by LLMs around 2028, illustrating the data scarcity problem.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-00-37.png)
![Screenshot at 01:41: News headlines highlighting recent high-profile copyright lawsuits involving Anthropic and The New York Times, contextualizing the legal challenges of training data.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-01-41.png)
![Screenshot at 02:27: Slide outlining the two core questions of the project: how data shapes GenAI and how GenAI reshapes data, divided into 'Data created by AI' and 'Data used by AI'.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-02-27.png)
![Screenshot at 04:25: Chart showing the sharp post-ChatGPT increase in the frequency of words like 'commendable' and 'notable' in scientific papers, suggesting LLM writing influence.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-04-25.png)
![Screenshot at 06:28: Four-panel chart showing broad LLM adoption across consumer complaints \(A\), press releases \(B, C\), and LinkedIn job postings \(D\) following the ChatGPT launch.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-06-28.png)
![Screenshot at 08:53: Diagram illustrating the secure, HIPAA-compliant infrastructure used by the Stanford Center for Population Health Sciences for processing sensitive data.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-08-53.png)
![Screenshot at 10:36: Slide summarizing survey results: 83% couldn't distinguish real vs. synthetic data, 92% would use synthetic data for grant proposals, and the data looks very similar.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-10-36.png)
![Screenshot at 12:25: Slide detailing the three-step process for creating diverse synthetic data using knowledge graphs and LLMs, showing performance gains on a graph.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-12-25.png)
![Screenshot at 18:25: Slide titled 'A puzzle: perfect diffusion models cannot be creative,' explaining that perfect reversal leads to memorization, not creativity, implying creativity comes from failure to learn perfectly.](https://ss.rapidrecap.app/screens/CAcCAmp_RAw/00-18-25.png)
