Can LLMs Cook Jamaican Couscous? A Study of Cultural Novelty in Recipe Generation
Quick Overview
The study concluded that Large Language Models (LLMs) like those in the Llama family fail to generate culturally nuanced recipe adaptations, instead producing outputs that are statistically biased toward Western-centric averages, often ignoring or corrupting specific cultural ingredients like Scotch bonnet peppers when adapting Jamaican Couscous recipes.
Key Points: The research specifically tested LLMs on adapting a Jamaican Couscous recipe to ensure cultural novelty and accuracy. LLMs, including Llama 3, Gemma 3, and Orion 14B, struggled significantly to produce culturally representative adaptations. The models frequently defaulted to statistically probable, Western-centric averages, even when explicitly asked for cultural nuance. Specific Jamaican ingredients like Scotch bonnet pepper and pigeon peas were either omitted or replaced with generic substitutes like onion or garlic in the AI-generated recipes. The study used a 'cultural distance' metric and JSD (Jensen-Shannon Divergence) to quantify how far the AI output deviated from human-adapted versions. The research suggests that LLMs lack the internal mechanics to maintain cultural boundaries, treating distinct cultures as interchangeable or flattening them into a generic average. The authors warn this tendency risks cultural erasure or offensive outputs in applications requiring domain-specific knowledge.
Context: The AI Papers podcast discussed research from Mila and McGill University published in February 2024, which investigated the ability of Large Language Models (LLMs) to handle cultural nuance in creative generation tasks, using recipe adaptation—specifically transforming a traditional Jamaican Couscous recipe—as a rigorous stress test for the future of generative AI.
Detailed Analysis
The research analyzed the performance of several LLMs, including Llama 3, Gemma 3, and Orion 14B, on adapting a Jamaican Couscous recipe to a Jamaican context. The core finding was that these state-of-the-art models consistently failed to incorporate or preserve specific cultural markers. When prompted to adapt the recipe, the models frequently generated outputs that reflected a generic, Western-centric average, effectively erasing the distinct cultural identity of the dish. For instance, when adapting the recipe, the AI models often replaced culturally significant ingredients like Scotch bonnet pepper or pigeon peas with generic placeholders such as onion, garlic, or oil. The researchers used metrics like cultural distance and JSD to measure the divergence between the AI's output and human-adapted versions, finding that the AI's attempts were statistically similar to the Western average rather than showing genuine novelty or cultural fidelity. The study also examined the internal layers of the models using the Logit Lens technique, which revealed that the models prioritize preserving the basic grammatical structure and general concept (like 'cooking') but strip out the nuanced cultural information during processing, leading to culturally hollow or inaccurate results. The authors suggest this points to a systemic issue where current LLMs default to the most probable—often Western-centric—data in their training, posing a risk of cultural flattening or offensive outputs in real-world applications.