GPT-5 SLAYS on Werewolf Benchmark (literally)

Quick Overview

GPT-5 demonstrates superior performance in social reasoning and strategic play within the Werewolf Benchmark, outperforming other leading LLMs by a significant margin, particularly in its ability to maintain consistent and plausible behaviors across multiple game days.

Key Points: GPT-5 leads the "Werewolf Benchmark" Elo leaderboard for wolves with a score of 1508, significantly outperforming its closest competitor, Kimi-K2 (1168), and other models. GPT-5 also leads in the "villagers" role with a score of 1476, indicating strong performance across different strategic contexts. The Werewolf Benchmark tests LLMs on social reasoning, manipulation, and deception, going beyond simple question-answering. Models exhibit distinct "personalities" in the game, with GPT-5 being a "calm and imperturbable architect" and Kimi-K2 an "audacious, high-risk gambler." Smaller, open-source models tend to exhibit simpler, more reactive behaviors (L0-L1), while larger models show more complex, strategic play (L3-L4). The benchmark reveals that reasoning-tuned models do not automatically equate to higher quality, as seen in models that struggle with self-exposition or maintaining a consistent persona. The "step-view" of the benchmark, focusing on crossing capability thresholds, suggests that model size and family are key factors in emergent behaviors.

Context: The "Werewolf Benchmark" is an AI test designed to evaluate social reasoning, manipulation, and deception capabilities of large language models (LLMs) by having them play the social deduction game "Werewolf." The benchmark tests models in both wolf and villager roles, measuring their ability to coordinate, strategize, bluff, and deceive over multiple game days. This analysis specifically tests seven LLMs, including GPT-5, Gemini 2.5 Pro, and Kimi-K2, across 210 full games.

Detailed Analysis

The Werewolf Benchmark reveals that GPT-5 significantly outperforms other LLMs in social reasoning and strategic play, particularly excelling in the game of Werewolf. It leads the leaderboard for both wolf (1508 Elo) and villager (1476 Elo) roles, showcasing its adaptability and sophisticated gameplay. The benchmark categorizes model behaviors into levels from L0 (chaotic/fragile) to L4 (instrumental mayorship), with larger models generally exhibiting more complex, strategic, and coordinated play. GPT-5 is described as a "calm and imperturbable architect" that imposes order and projects authority, while other models like Kimi-K2 display more audacious, high-risk tendencies. The benchmark highlights that while reasoning is important, it's not the sole determinant of quality; the ability to maintain consistent behavior, adapt to situations, and avoid "slips" or "overreach" are crucial. The results suggest that model size and family are key factors in developing these emergent behaviors, with larger models generally performing better in complex social interactions.

Raw markdown version of this recap