# Divergent Creativity in Humans and Large Language Models

Source: https://www.youtube.com/watch?v=U8VJY3rO1RE
Recap page: https://rapidrecap.app/video/U8VJY3rO1RE
Generated: 2026-01-26T23:31:41.044+00:00

---
## Quick Overview

A study published in Nature Scientific Reports in 2026 found that while Large Language Models (LLMs) like GPT-4 exhibit high scores on creativity metrics (like generating diverse word lists), their creativity is often mathematically consistent and lacks the genuine semantic divergence seen in human writing, suggesting LLMs excel at pattern following but struggle with true novelty, especially when temperature settings are low.

**Key Points:**
- A 2026 study in Nature Scientific Reports compared creativity metrics between LLMs (GPT-4, Claude 3, Gemini) and humans on creative tasks like poetry generation.
- LLMs, particularly GPT-4, scored significantly higher than the average human on the quantitative creativity metric (Divergent Association Test - DAT).
- However, when testing for semantic divergence, humans showed a distinct advantage, scoring higher than the models across various creative writing tasks.
- The study found that LLMs tended to cluster their responses in a small region of semantic space, often defaulting to high-probability, low-complexity words (like 'ocean' when prompted for 'sea').
- When the temperature setting was low (making the model more deterministic), the LLMs showed even greater redundancy and less diversity.
- The researchers concluded that LLMs are excellent at following prescribed patterns (like the 5-7-5 structure of a Haiku) but lack the human ability to break constraints or make semantically distant but contextually meaningful leaps.

![Screenshot at 00:24: The video displays the central conflict of the study: one camp claims LLMs have surpassed humans in creativity \(e.g., GPT-4 outperforming the average person in DAT scores\), while the other argues LLMs only mimic patterns, leading to a discussion on the nuance between quantitative scores and genuine semantic novelty.](https://ss.rapidrecap.app/screens/U8VJY3rO1RE/00-00-24.jpg)

**Context:** The video discusses the findings of a scientific study comparing the creative capabilities of state-of-the-art Large Language Models (LLMs), such as GPT-4, against human performance. The core debate centers on whether LLMs are truly creative or merely masters of pattern replication, using specific metrics like the Divergent Association Test (DAT) and semantic distance analysis to quantify the differences in creative output.

## Detailed Analysis

The discussion revolves around a 2026 study from Nature Scientific Reports investigating divergent creativity in LLMs versus humans. The initial findings showed that models like GPT-4, Claude 3, and Gemini outperformed the average human on the Divergent Association Test (DAT), achieving a benchmark score equivalent to 100,000 human participants. However, the researchers dug deeper, finding that this high score was misleading. When testing creative writing tasks, such as generating 10 semantically different words or writing a Haiku, humans demonstrated superior semantic divergence. The study revealed that LLMs tend to operate within a narrow, predictable semantic space, often defaulting to high-probability words (e.g., 'ocean' instead of 'sea' when prompted for a synonym) even when attempting creative tasks. When the model's temperature setting was lowered, this tendency toward predictable, low-complexity output was exacerbated, leading to lower scores. The researchers suggest that while LLMs are excellent at following rules (like the 5-7-5 structure of a Haiku), they lack the genuine, unpredictable spark of human creativity, which often involves making semantically distant yet meaningful connections. The paper flags that simply optimizing for the mathematical score (like the DAT) risks missing the real goal: meaningful novelty.

### Study Context and Models

- Study published in Nature Scientific Reports (2026)
- Compared GPT-4, Claude 3, Gemini vs. 100,000 human participants
- Initial DAT scores showed LLMs significantly higher than average humans

### Creativity Metrics and Findings

- DAT measures ability to generate diverse word lists
- Humans scored higher on semantic divergence for creative tasks
- LLMs cluster responses, favoring high-probability, low-complexity words (e.g., 'ocean' over 'sea')

### Impact of Temperature Setting

- Low temperature (deterministic setting) forces models toward predictable, safe answers
- High temperature (chaotic setting) forces models to take risks, potentially leading to hallucinations

### Creative Writing Tasks

- LLMs excelled at following structural rules (Haiku)
- Humans showed greater semantic distance and novel connections in story/synopsis generation
- LLMs often stick to safe, low-complexity word choices

![Screenshot at 00:00: Introductory screen with podcast graphic and 'Become a Member Today!' CTA set against an audio waveform display.](https://ss.rapidrecap.app/screens/U8VJY3rO1RE/00-00-00.jpg)
![Screenshot at 00:16: The speakers introduce the topic: a study trying to quantify creativity, which is usually considered impossible to measure.](https://ss.rapidrecap.app/screens/U8VJY3rO1RE/00-00-16.jpg)
![Screenshot at 00:53: The title of the paper being discussed: 'Divergent Creativity in Humans and Large Language Models' by Bellmare Peppen and colleagues.](https://ss.rapidrecap.app/screens/U8VJY3rO1RE/00-00-53.jpg)
![Screenshot at 01:37: A visual representation of the core finding: GPT-4 hits a mathematical ceiling but struggles to reach the extreme creativity levels of the best human writers.](https://ss.rapidrecap.app/screens/U8VJY3rO1RE/00-01-37.jpg)
![Screenshot at 04:50: The speaker notes that GPT-4 generated 'ocean' in over 90% of its responses when prompted for a word semantically distant from 'car', suggesting redundancy.](https://ss.rapidrecap.app/screens/U8VJY3rO1RE/00-04-50.jpg)
