# Can LLMs Cook Jamaican Couscous? A Study of Cultural Novelty in Recipe Generation

Source: https://www.youtube.com/watch?v=sUa_BFNj_Zw
Recap page: https://rapidrecap.app/video/sUa_BFNj_Zw
Generated: 2026-02-15T19:02:44.67+00:00

---
## Quick Overview

The study concluded that Large Language Models (LLMs) like those in the Llama family fail to generate culturally nuanced recipe adaptations, instead producing outputs that are statistically biased toward Western-centric averages, often ignoring or corrupting specific cultural ingredients like Scotch bonnet peppers when adapting Jamaican Couscous recipes.

**Key Points:**
- The research specifically tested LLMs on adapting a Jamaican Couscous recipe to ensure cultural novelty and accuracy.
- LLMs, including Llama 3, Gemma 3, and Orion 14B, struggled significantly to produce culturally representative adaptations.
- The models frequently defaulted to statistically probable, Western-centric averages, even when explicitly asked for cultural nuance.
- Specific Jamaican ingredients like Scotch bonnet pepper and pigeon peas were either omitted or replaced with generic substitutes like onion or garlic in the AI-generated recipes.
- The study used a 'cultural distance' metric and JSD (Jensen-Shannon Divergence) to quantify how far the AI output deviated from human-adapted versions.
- The research suggests that LLMs lack the internal mechanics to maintain cultural boundaries, treating distinct cultures as interchangeable or flattening them into a generic average.
- The authors warn this tendency risks cultural erasure or offensive outputs in applications requiring domain-specific knowledge.

![Screenshot at 00:00: The video opens with an animated graphic of two podcasters in front of an audio waveform, overlaid with a call to action: "BECOME A MEMBER TODAY!", setting the scene for a discussion on AI research.](https://ss.rapidrecap.app/screens/sUa_BFNj_Zw/00-00-00.jpg)

**Context:** The AI Papers podcast discussed research from Mila and McGill University published in February 2024, which investigated the ability of Large Language Models (LLMs) to handle cultural nuance in creative generation tasks, using recipe adaptation—specifically transforming a traditional Jamaican Couscous recipe—as a rigorous stress test for the future of generative AI.

## Detailed Analysis

The research analyzed the performance of several LLMs, including Llama 3, Gemma 3, and Orion 14B, on adapting a Jamaican Couscous recipe to a Jamaican context. The core finding was that these state-of-the-art models consistently failed to incorporate or preserve specific cultural markers. When prompted to adapt the recipe, the models frequently generated outputs that reflected a generic, Western-centric average, effectively erasing the distinct cultural identity of the dish. For instance, when adapting the recipe, the AI models often replaced culturally significant ingredients like Scotch bonnet pepper or pigeon peas with generic placeholders such as onion, garlic, or oil. The researchers used metrics like cultural distance and JSD to measure the divergence between the AI's output and human-adapted versions, finding that the AI's attempts were statistically similar to the Western average rather than showing genuine novelty or cultural fidelity. The study also examined the internal layers of the models using the Logit Lens technique, which revealed that the models prioritize preserving the basic grammatical structure and general concept (like 'cooking') but strip out the nuanced cultural information during processing, leading to culturally hollow or inaccurate results. The authors suggest this points to a systemic issue where current LLMs default to the most probable—often Western-centric—data in their training, posing a risk of cultural flattening or offensive outputs in real-world applications.

### Paper Context

- Research from Mila and McGill on cultural novelty in recipe generation
- Paper title: 'Can LLMs Cook Jamaican Couscous?'
- Study released in February 2024

### Methodology

- Tested 8 different LLMs (including Llama 3, Gemma 3, Orion 14B) using a human baseline recipe
- Used cultural distance and JSD metrics to measure deviation from human adaptation

### Key Findings on LLM Performance

- LLMs failed to produce culturally representative adaptations, leaning toward Western averages
- Models struggled with specific Jamaican ingredients like Scotch bonnet pepper and pigeon peas, replacing them with generic items (onion, garlic, oil)

### Internal Model Analysis (Logit Lens)

- Middle layers compress cultural information and lose specific nuance
- Models prioritize general structure/grammar over cultural specificity

### Comparison to Human Adaptation

- Humans maintain core identity while adjusting specific markers; LLMs flatten the dish into a generic, statistically probable output

### Conclusion & Implications

- LLMs risk cultural flattening or erasure by losing specific identity markers; this is a systemic issue, not just a flaw in one model

![Screenshot at 0:00: The introductory screen features an illustration of two podcasters working on laptops, overlaid with a call to action to 'BECOME A MEMBER TODAY!'](https://ss.rapidrecap.app/screens/sUa_BFNj_Zw/00-00-00.jpg)
![Screenshot at 0:31: A visual cue indicating the discussion is focusing on the core research question, asking if LLMs get cultural nuance.](https://ss.rapidrecap.app/screens/sUa_BFNj_Zw/00-00-31.jpg)
![Screenshot at 1:24: The speaker details the methodology, centering around the 'LLM Fusion' dataset and pairing recipes based on 'cultural distance'.](https://ss.rapidrecap.app/screens/sUa_BFNj_Zw/00-01-24.jpg)
![Screenshot at 2:25: The speaker highlights the contrast between AI output and human adaptation, noting the AI's output was 'not great' for cultural adaptation.](https://ss.rapidrecap.app/screens/sUa_BFNj_Zw/00-02-25.jpg)
![Screenshot at 4:24: The speaker signals the transition to examining the creativity prompt test used to evaluate the models.](https://ss.rapidrecap.app/screens/sUa_BFNj_Zw/00-04-24.jpg)
