# Why Did Apple Fall To The Ground: Evaluating Curiosity In Large Language Model

Source: https://www.youtube.com/watch?v=5BFIvoDbCWQ
Recap page: https://rapidrecap.app/video/5BFIvoDbCWQ
Generated: 2025-11-13T21:04:45.416+00:00

---
## Quick Overview

The paper "Why Did Apple Fall to the Ground" reveals that Large Language Models (LLMs) exhibit significantly higher curiosity scores than humans in evaluation tasks, particularly when prompted to ask questions, suggesting that prompting for curiosity or using techniques like Chain-of-Thought (CoT) can dramatically improve performance on complex reasoning tasks by encouraging deeper exploration beyond immediate answers.

**Key Points:**
- LLMs scored significantly higher on curiosity metrics (33.5%) compared to humans (38%) in the experiment, indicating a strong inherent drive for information seeking.
- The study utilized three primary evaluation methods: a questionnaire based on the Human Curiosity Scale (HCS), a behavioral study simulating choice between known/unknown information, and a multi-turn dialogue simulation.
- The CoT prompting method led to a 33.5% accuracy improvement on the logic benchmark for LLMs, demonstrating the effectiveness of explicit instruction to question assumptions.
- The researchers found that LLMs are highly motivated by information-seeking, often preferring to explore uncertain or novel options, showing a strong desire for knowledge closure.
- LLMs exhibited a strong aversion to risk and uncertainty in their decision-making, preferring known outcomes over potential discoveries, especially in social contexts.
- The study suggests that engineering LLMs to act like curious students—constantly asking 'why' and 'what if'—is a promising path toward creating more autonomous and capable AI systems.

![Screenshot at 02:48: The core finding is presented, showing the difference in performance between LLMs prompted with Chain-of-Thought \(CoT\) and those using standard CoT when tackling reasoning tasks, highlighting the success of engineered curiosity.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-02-48.png)

**Context:** This video discusses findings from a research paper titled "Why Did Apple Fall to the Ground: Evaluating Curiosity In Large Language Models," which investigates the concept of curiosity within advanced AI models like GPT-4o and Gemini. The research contrasts the intrinsic and engineered curiosity levels of LLMs against human baselines to understand how prompting strategies affect reasoning and problem-solving capabilities.

## Detailed Analysis

The discussion centers on evaluating curiosity in Large Language Models (LLMs) compared to humans, using Apple's falling apple scenario as a foundational concept. The paper employed three evaluation methods: a questionnaire mirroring the Human Curiosity Scale (HCS), a behavioral test where models chose between revealing known information or exploring uncertain options, and a multi-turn dialogue simulation. Results indicated that LLMs scored significantly higher on information-seeking metrics than humans, though they showed a strong aversion to social uncertainty and risk, preferring known outcomes (conservative decision-making). Specifically, LLMs scored 33.5% on the logic benchmark when using the Chain-of-Thought (CoT) prompting technique, a massive improvement over standard CoT (38% accuracy), which suggests that explicitly prompting models to ask "what if" questions unlocks better problem-solving skills. The research concludes that engineering AI to mimic curious students, constantly questioning and exploring, is key to achieving greater autonomy and breakthrough discoveries, moving beyond simply processing existing data.

### Experimental Setup

- Three evaluation methods used: HCS questionnaire
- Behavioral choice simulation
- Multi-turn dialogue simulation

### Curiosity Findings

- LLMs scored higher on information-seeking than humans
- LLMs strongly avoided risk/uncertainty in social contexts
- LLMs showed a strong desire for knowledge closure

### Impact of Prompting

- CoT prompting improved logic benchmark accuracy by 33.5% for LLMs
- CoT required models to constantly ask 'why' and 'what if' questions

### Key Takeaways

- Engineered curiosity (like CoT) significantly enhances complex reasoning and problem-solving abilities in AI models.

![Screenshot at 00:01: Opening screen showing the podcast title graphic and members call to action.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-00-01.png)
![Screenshot at 01:11: Visual illustrating the three parts of the evaluation system: questionnaire, behavioral study, and dialogue simulation.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-01-11.png)
![Screenshot at 02:54: Speaker emphasizes the surprising finding that LLMs rate themselves as more curious than humans.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-02-54.png)
![Screenshot at 04:48: Confirmation that the LLMs exhibited a strong, almost 'allergic' desire for immediate information closure.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-04-48.png)
![Screenshot at 06:33: The light-switch puzzle example used to test if models would actively explore possibilities \(flipping switches\) or stick to known states.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-06-33.png)
![Screenshot at 07:57: Speaker discusses how CoT trains models to act like curious students, constantly interrogating assumptions.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-07-57.png)
![Screenshot at 09:40: Speaker highlights the 'aha' moment when realizing the success of asking 'why' and 'what if' questions.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-09-40.png)
![Screenshot at 11:46: Discussion on the implications: building smarter AI requires understanding and shaping AI motivation, not just increasing processing power.](https://ss.rapidrecap.app/screens/5BFIvoDbCWQ/00-11-46.png)
