Measuring Agents With Interactive Evaluations
Quick Overview
The presentation introduces ARC-AGI-3, a new benchmark for evaluating general artificial intelligence, which focuses on interactive evaluations rather than static benchmarks to measure an agent's ability to explore, learn rules, plan actions, and achieve goals efficiently in novel environments, demonstrating that current frontier AI models still exhibit a significant gap compared to average human performance in these complex, interactive tasks.
Key Points: Interactive evaluations are presented as the key to measuring progress towards Artificial General Intelligence (AGI), contrasting with traditional static benchmarks. The ARC-AGI-3 benchmark features over 150 novel video game environments designed to be easy for humans but hard for AI, testing exploration, rule learning, planning, and efficiency. Human performance in ARC-AGI-3 requires significantly more actions per level than the current frontier AI models, demonstrating a clear Human-to-AI Gap in action efficiency (0:13:57, 12:48). The speaker highlights François Chollet's definition of intelligence as 'Skill Acquisition Efficiency' (1:28), emphasizing the importance of learning efficiency in novel scenarios. The agent criteria for success in ARC-AGI-3 include navigating unseen environments, learning rules, executing a plan, and matching or surpassing human-level action efficiency (19:14). The presentation concludes with a call to action for developers to 'Play Games' or 'Build Agents' using the ARC-AGI-3 environments, with resources available at three.arcprize.org (20:22).
![Screenshot at 0:10: The speaker, Greg Kamradt, President of the ARC Prize Foundation, introduces the topic "Measuring Agents With Interactive Evaluations" to the audience at OpenAI DevDay \[2025\].](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-00-10.png)
Context: This presentation, given by Greg Kamradt, President of the ARC Prize Foundation, at the OpenAI DevDay [2025], introduces the ARC-AGI-3 benchmark. This benchmark is designed to move beyond traditional static AI evaluations by focusing on interactive environments, specifically video games, to better measure an agent's general intelligence capabilities, including exploration, rule inference, planning, and efficient action execution.