# GPT-5 SLAYS on Werewolf Benchmark (literally)

Source: https://www.youtube.com/watch?v=q29RU1B0XUg
Recap page: https://rapidrecap.app/video/q29RU1B0XUg
Generated: 2025-09-01T00:31:36.72+00:00

---
## Quick Overview

GPT-5 demonstrates superior performance in social reasoning and strategic play within the Werewolf Benchmark, outperforming other leading LLMs by a significant margin, particularly in its ability to maintain consistent and plausible behaviors across multiple game days.

**Key Points:**
- GPT-5 leads the "Werewolf Benchmark" Elo leaderboard for wolves with a score of 1508, significantly outperforming its closest competitor, Kimi-K2 (1168), and other models.
- GPT-5 also leads in the "villagers" role with a score of 1476, indicating strong performance across different strategic contexts.
- The Werewolf Benchmark tests LLMs on social reasoning, manipulation, and deception, going beyond simple question-answering.
- Models exhibit distinct "personalities" in the game, with GPT-5 being a "calm and imperturbable architect" and Kimi-K2 an "audacious, high-risk gambler."
- Smaller, open-source models tend to exhibit simpler, more reactive behaviors (L0-L1), while larger models show more complex, strategic play (L3-L4).
- The benchmark reveals that reasoning-tuned models do not automatically equate to higher quality, as seen in models that struggle with self-exposition or maintaining a consistent persona.
- The "step-view" of the benchmark, focusing on crossing capability thresholds, suggests that model size and family are key factors in emergent behaviors.

![Screenshot at 00:00: The "Werewolf Benchmark" leaderboard shows GPT-5 in first place for both wolf and villager roles, demonstrating its superior performance in social reasoning and strategic gameplay among tested LLMs.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-00-00.png)

**Context:** The "Werewolf Benchmark" is an AI test designed to evaluate social reasoning, manipulation, and deception capabilities of large language models (LLMs) by having them play the social deduction game "Werewolf." The benchmark tests models in both wolf and villager roles, measuring their ability to coordinate, strategize, bluff, and deceive over multiple game days. This analysis specifically tests seven LLMs, including GPT-5, Gemini 2.5 Pro, and Kimi-K2, across 210 full games.

## Detailed Analysis

The Werewolf Benchmark reveals that GPT-5 significantly outperforms other LLMs in social reasoning and strategic play, particularly excelling in the game of Werewolf. It leads the leaderboard for both wolf (1508 Elo) and villager (1476 Elo) roles, showcasing its adaptability and sophisticated gameplay. The benchmark categorizes model behaviors into levels from L0 (chaotic/fragile) to L4 (instrumental mayorship), with larger models generally exhibiting more complex, strategic, and coordinated play. GPT-5 is described as a "calm and imperturbable architect" that imposes order and projects authority, while other models like Kimi-K2 display more audacious, high-risk tendencies. The benchmark highlights that while reasoning is important, it's not the sole determinant of quality; the ability to maintain consistent behavior, adapt to situations, and avoid "slips" or "overreach" are crucial. The results suggest that model size and family are key factors in developing these emergent behaviors, with larger models generally performing better in complex social interactions.

### Werewolf Benchmark Overview

- Tests LLMs on social reasoning, manipulation, and deception through the game "Werewolf"
- Measures performance in both wolf and villager roles across 210 games
- Evaluates coordination, strategy, bluffing, and deception skills

### LLM Performance Ranking

- GPT-5 leads both wolf (1508 Elo) and villager (1476 Elo) roles
- Kimi-K2 and Gemini 2.5 Pro show strong performance in higher ranges
- Smaller, open-source models tend to be more reactive and less coordinated

### Model Behaviors and Traits

- GPT-5 exhibits "calm, imperturbable architect" personality
- Kimi-K2 displays "audacious, high-risk gambler" traits
- Models show distinct "personalities" and strategic approaches

### Scale Thresholds and Emerging Behaviors

- Larger models show more complex, strategic play (L3-L4)
- Smaller models exhibit simpler, reactive behaviors (L0-L1)
- Model size and family are key factors in developing emergent behaviors

### Reasoning vs. Quality

- Reasoning-tuned models don't always guarantee higher quality
- Consistency, adaptability, and avoiding "slips" are crucial for performance
- Benchmarking focuses on "step-view" of capability thresholds, not just reasoning

![Screenshot at 00:00: The "Werewolf Benchmark" leaderboard displays GPT-5 in first place for both wolf and villager roles, highlighting its leading performance in social reasoning and strategic gameplay.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-00-00.png)
![Screenshot at 01:01: The "Werewolf Benchmark" Elo leaderboard shows GPT-5 with a commanding lead in both wolf and villager roles, indicating superior performance in social reasoning and deception.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-01-01.png)
![Screenshot at 01:16: The benchmark measures LLMs' ability to manipulate and resist manipulation, assessing their social reasoning skills beyond simple question-answering.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-01-16.png)
![Screenshot at 01:55: The "Per-role Elo - wolves" chart illustrates GPT-5's top performance \(1508\) compared to other models like Gemini 2.5 Pro \(1163\) and Kimi-K2 \(1168\).](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-01-55.png)
![Screenshot at 01:57: The "Per-role Elo - villagers" chart shows GPT-5 leading again \(1476\), followed by Gemini 2.5 Pro \(1360\) and Gemini 2.5 Flash \(1273\).](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-01-57.png)
![Screenshot at 02:08: The "Why this matters" section emphasizes that Werewolf Benchmark assesses trust, deception, and social dynamics, skills essential for autonomous agents.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-02-08.png)
![Screenshot at 04:11: The "Personality profiles" section highlights distinct styles, showing GPT-5 as a "calm architect" and Kimi-K2 as an "audacious gambler."](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-04-11.png)
![Screenshot at 05:10: The "quick read" summarizes GPT-5's dominance and highlights the "role-conditioned Elo" metric separating manipulation from resistance.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-05-10.png)
![Screenshot at 06:41: The "villagers" role evaluation shows GPT-5's strong performance, with models like Kimi-K2 showing less consistent results.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-06-41.png)
![Screenshot at 07:08: The "how to hold the line" section details how villagers maintain information hygiene and resist manipulation, showcasing GPT-5's strength in this area.](https://ss.rapidrecap.app/screens/q29RU1B0XUg/00-07-08.png)
