# Butter-Bench: Evaluating LLM-Controlled Robots for Practical Intelligence

Source: https://www.youtube.com/watch?v=_4rmO86P6Ko
Recap page: https://rapidrecap.app/video/_4rmO86P6Ko
Generated: 2025-11-11T08:10:13.643+00:00

---
## Quick Overview

The "Butter-Bench" paper reveals that Large Language Models (LLMs) struggle significantly with practical intelligence tasks, scoring only 40% completion on tasks requiring physical manipulation and showing a critical gap compared to human common sense, as demonstrated by the failure of advanced models like GPT-4 and Gemini 2.5 Pro on simple tasks like identifying butter in a picture or performing safety checks.

**Key Points:**
- The best-performing LLM tested, Claude Opus, only managed 40% completion on tasks involving physical manipulation.
- When performing real-world tasks like identifying butter in a picture or checking status, LLMs failed to exhibit common sense reasoning.
- The study found a significant gap between LLM analytical/coding skills (scoring 95%) and practical reasoning skills (scoring 40% or lower).
- Specific models tested, including GPT-4 and Gemini 2.5 Pro, exhibited similar shortcomings in practical reasoning, with Gemini 2.5 Pro scoring only 27% on robotics data tasks.
- Failures often stemmed from an inability to handle ambiguity, social cues, or the complexity of real-world physics, rather than just lacking basic skills.
- The paper suggests that current fine-tuning methods, which often rely on internal monologue or self-correction, do not effectively bridge the gap to real-world physical intelligence.

![Screenshot at 00:50: The quantitative results showing the best tested LLM achieving only 40% completion on tasks related to physical manipulation, highlighting the core performance deficit this paper addresses.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-00-50.png)

**Context:** This video summarizes findings from the research paper "Butter-Bench," which evaluates the practical intelligence of Large Language Models (LLMs) when controlling robots. The evaluation specifically contrasts the LLMs' high performance in analytical and coding tasks against their performance in tasks requiring common sense reasoning about the physical world, such as navigating environments or interacting with objects like butter.

## Detailed Analysis

The discussion centers on the "Butter-Bench" paper, which evaluates LLM-controlled robots across five main categories: tool use, movement driving/turning, kinematics control, environmental perception, and social interaction. The paper finds a major disconnect between LLMs' high analytical abilities (like coding) and their capacity for practical intelligence. For instance, while LLMs excel at abstract reasoning, they struggle profoundly with real-world tasks. The best model, Claude Opus, scored only 40% on physical manipulation tasks. When tested on specific failures, like identifying butter in a picture or assessing if a robot could carry an item on a tray, advanced models like GPT-4 and Gemini 2.5 Pro performed poorly, sometimes scoring as low as 27% on certain robotics data sets. A key finding is that models often fail when encountering ambiguity, social cues, or unexpected physical situations, suggesting that current fine-tuning methods that emphasize internal monologue do not adequately teach the practical common sense needed for real-world deployment.

### LLM Performance Comparison

- Claude Opus scored 40% on physical tasks, while humans scored 95%; Gemini 2.5 Pro scored 27% on specific robotics data.

### Failure Categories

- LLMs struggle with ambiguity, social cues (like waiting for confirmation), and complex physical tasks, unlike simple go-fetch commands.

### Specific Task Failures

- Models failed to identify butter in a picture or determine if a robot could carry an object on a tray, showing a lack of common sense.

### Implications for Deployment

- The gap between analytical skill and practical intelligence poses a critical safety risk, as models fail in unexpected real-world scenarios.

### Training Methodology Critique

- The paper implies that current fine-tuning, even with self-correction loops, is insufficient for instilling practical, real-world common sense.

![Screenshot at 00:28: The hosts introduce the paper "Butter-Bench" which evaluates LLMs controlling robots for practical intelligence.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-00-28.png)
![Screenshot at 00:49: A visual representation of the poor quantitative results, noting the best LLM achieved only 40% completion on physical tasks.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-00-49.png)
![Screenshot at 01:25: The speaker delineates the two types of intelligence: analytical \(problem-solving/logic\) versus practical \(street smarts\).](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-01-25.png)
![Screenshot at 02:24: The speaker mentions the VLA \(Visual Language Action\) model, which translates planning into movement.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-02-24.png)
![Screenshot at 04:48: The speaker discusses the robot's failure when asked to identify butter in a picture, showing it couldn't infer context.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-04-48.png)
![Screenshot at 05:57: The discussion highlights that LLM failures often result from the complexity of real-world messiness rather than simple execution errors.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-05-57.png)
![Screenshot at 07:12: The speaker uses the example of the Grok 4 model rushing a task instead of waiting for confirmation.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-07-12.png)
![Screenshot at 11:48: The speaker notes that models often generate lengthy internal monologues but still fail in simple physical tasks, indicating a lack of true understanding.](https://ss.rapidrecap.app/screens/_4rmO86P6Ko/00-11-48.png)
