Session 5: Data in the Age of Generative AI

Quick Overview

The presentation, titled "Data in the Age of Generative AI," covered research from Stanford HAI focusing on two main areas: data created by AI (synthetic data) and data used by AI (training data), projecting that the stock of human-generated public text will be exhausted by 2028 due to LLM consumption. Key takeaways include the successful development of a framework to rigorously evaluate synthetic data, where 83% of experts could not distinguish synthetic from real data, and the immediate next steps involve deploying synthetic data in population health sciences, scaling data attribution for LLMs, and designing mechanisms for potential future GenAI data markets.

Key Points: The available stock of human-generated public text, estimated to be around $10^{16}$ tokens, will likely be fully utilized by generative AI models by the median date of 2028. Recent advances in LLMs and diffusion models are heavily driven by consuming this finite human-generated data, leading to concerns about data scarcity and quality. The research team developed a framework to rigorously evaluate synthetic data, showing that 83% of experts could not distinguish between real and synthetic texts in a survey. The synthetic data generated by their methods (like the Type 2 Diabetes patient notes example) maintain complex relationships and structure found in real data, scoring a median $r^2$ of 0.9 between theory and experiment on accuracy. Next steps include deploying synthetic data with Population Health Sciences, scaling data attribution methods for LLMs, and designing mechanisms for potential GenAI data markets. Data attribution methods are crucial for understanding the sources of model 'creativity' and for attributing value (and potentially pricing) to different data sources.

Context: This presentation, given by members of the Stanford Human-Centered Artificial Intelligence (HAI) institute, addresses the rapidly growing data demands of large language models (LLMs) and generative AI. The speakers, including James Zou, Daniel E. Ho, and Surya Ganguli (listed as PIs/Co-PIs), discuss the impending exhaustion of high-quality public text data and explore solutions like synthetic data generation and data attribution to manage future AI development responsibly and effectively.

Raw markdown version of this recap