# Measuring Agents With Interactive Evaluations

Source: https://www.youtube.com/watch?v=TK9MN22q6E0
Recap page: https://rapidrecap.app/video/TK9MN22q6E0
Generated: 2025-10-08T17:35:50.259+00:00

---
## Quick Overview

The presentation introduces ARC-AGI-3, a new benchmark for evaluating general artificial intelligence, which focuses on interactive evaluations rather than static benchmarks to measure an agent's ability to explore, learn rules, plan actions, and achieve goals efficiently in novel environments, demonstrating that current frontier AI models still exhibit a significant gap compared to average human performance in these complex, interactive tasks.

**Key Points:**
- Interactive evaluations are presented as the key to measuring progress towards Artificial General Intelligence (AGI), contrasting with traditional static benchmarks.
- The ARC-AGI-3 benchmark features over 150 novel video game environments designed to be easy for humans but hard for AI, testing exploration, rule learning, planning, and efficiency.
- Human performance in ARC-AGI-3 requires significantly more actions per level than the current frontier AI models, demonstrating a clear Human-to-AI Gap in action efficiency (0:13:57, 12:48).
- The speaker highlights François Chollet's definition of intelligence as 'Skill Acquisition Efficiency' (1:28), emphasizing the importance of learning efficiency in novel scenarios.
- The agent criteria for success in ARC-AGI-3 include navigating unseen environments, learning rules, executing a plan, and matching or surpassing human-level action efficiency (19:14).
- The presentation concludes with a call to action for developers to 'Play Games' or 'Build Agents' using the ARC-AGI-3 environments, with resources available at three.arcprize.org (20:22).

![Screenshot at 0:10: The speaker, Greg Kamradt, President of the ARC Prize Foundation, introduces the topic "Measuring Agents With Interactive Evaluations" to the audience at OpenAI DevDay \[2025\].](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-00-10.png)

**Context:** This presentation, given by Greg Kamradt, President of the ARC Prize Foundation, at the OpenAI DevDay [2025], introduces the ARC-AGI-3 benchmark. This benchmark is designed to move beyond traditional static AI evaluations by focusing on interactive environments, specifically video games, to better measure an agent's general intelligence capabilities, including exploration, rule inference, planning, and efficient action execution.

## Detailed Analysis

Greg Kamradt introduces the concept of using interactive evaluations as a crucial signal for measuring progress toward AGI, arguing that static benchmarks alone are insufficient because they fail to capture essential abilities like exploration, memory, goal acquisition, and alignment. He emphasizes François Chollet's definition of intelligence as 'Skill Acquisition Efficiency' (1:28), which requires measuring how efficiently an agent can learn new skills in an environment. The ARC-AGI-3 benchmark is detailed as a collection of over 150 novel, open-source video game environments, split into public training and private evaluation sets (0:07:06). A live demonstration of an agent playing Pokémon Crystal shows that while the agent can learn, it struggles with efficiency, often taking many more actions than a human to achieve a goal (1:51). A key graph comparing human performance against frontier AI performance (12:44) shows that humans are significantly more action-efficient, closing the 'Human-to-AI Gap' being the goal. The criteria for an agent to be considered successful in ARC-AGI-3 are to navigate unseen environments, learn the rules, execute a plan, and match or surpass human action efficiency (19:14). The presentation concludes by inviting the audience to get involved by playing the games or building agents via three.arcprize.org (20:22).

### Introduction to Interactive Evaluations

- Interactive benchmarks are the key to measuring progress towards AGI
- Intelligence is inherently interactive, requiring perception, planning, action, memory, goal acquisition, and alignment
- Static benchmarks fail to capture these dynamic capabilities.

### ARC-AGI-3 Benchmark Details

- Features 150+ novel, open-source video game environments
- Environments are split into Public Training and Private Evaluation sets
- The goal is to test an agent's ability to generalize to unseen environments (2:44).

### Defining Intelligence

- Cites François Chollet's 2019 definition: 'Skill Acquisition Efficiency' (1:28)
- The metric measures how efficiently an agent learns new skills, rather than just raw intelligence.

### Performance Comparison

- A chart shows human performance requires far more actions per level than frontier AI (Blue Line) (13:57)
- The Human-AI Gap is quantified by action efficiency, where humans perform better on average (18:38).

### Criteria for an Agent

- Successful agents must navigate unseen novel environments, learn environmental rules, execute a plan to achieve the goal, and match or surpass human-level action efficiency (19:14).

### Call to Action

- Get involved by playing ARC-AGI-3 games as a human or building adversarial agents to shape the architecture
- Visit three.arcprize.org to participate (20:22).

![Screenshot at 0:08: Greg Kamradt introducing the topic "Measuring Agents With Interactive Evaluations" to the audience.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-00-08.png)
![Screenshot at 0:27: Slide showing François Chollet's definition of intelligence as 'Skill Acquisition Efficiency' \(1:28\).](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-00-27.png)
![Screenshot at 0:40: Slide summarizing the ARC Prize Foundation's achievements, including 1.4K teams and 17K submissions in the 2024 competition.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-00-40.png)
![Screenshot at 2:50: The screen displays the core principle: "Intelligence is interactive," emphasizing feedback and action loops.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-02-50.png)
![Screenshot at 3:20: Demonstration of GPT-5 playing Pokémon Crystal, showing the complex interface the agent must process.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-03-20.png)
![Screenshot at 3:43: Slide listing the five key capabilities interactive environments can test: Exploration, Percept -\> Plan -\> Action, Memory, Goal acquisition, and Alignment.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-03-43.png)
![Screenshot at 5:18: A slide illustrating the two axes of evaluation: "Did you complete the goal?" and "How efficiently did you do it?"](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-05-18.png)
![Screenshot at 10:19: Slide summarizing the agent criteria: Navigating unseen environments, learning rules, executing plans, and matching/surpassing human efficiency.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-10-19.png)
![Screenshot at 13:04: A graph comparing human performance \(multiple gray/orange lines\) against Frontier AI \(blue line\) in terms of actions per level, showing a large gap.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-13-04.png)
![Screenshot at 20:22: Final slide encouraging involvement via playing games or building agents at three.arcprize.org.](https://ss.rapidrecap.app/screens/TK9MN22q6E0/00-20-22.png)
