# Your Brain Has a Secret ‘Maintenance Mode’—And AI Might Need One Tool

Source: https://www.youtube.com/watch?v=1gP-cFwko00
Recap page: https://rapidrecap.app/video/1gP-cFwko00
Generated: 2025-11-30T15:04:28.897+00:00

---
## Quick Overview

The video discusses three recent AI developments: Ubisoft's generative AI game demo 'Teammates,' the successful application of AI in rare disease diagnosis via the PopEVE model, and research showing current LLMs struggle with humor and password cracking due to a lack of genuine understanding, relying instead on pattern matching.

**Key Points:**
- Ubisoft unveiled an experimental game demo called 'Teammates' powered by generative AI, featuring voice-commanded NPCs named Jasper and an AI assistant, which creates dynamic, improvised dialogue beyond pre-written scripts.
- A new AI model from Harvard Medical School, PopEVE, can rapidly flag and prioritize genetic variants in human DNA, potentially leading to faster diagnosis and targeted treatments for rare diseases.
- Research shows that current LLMs, like GPT-4o, struggle with understanding puns and humor, exemplified by GPT-4o failing to grasp the double meaning in a pun about LLMs losing their 'attention' while a human understands it.
- The study on LLM humor suggests that models rely on memorizing structures rather than genuine linguistic comprehension, leading to poor performance on tasks requiring nuanced understanding like password cracking, where they are outperformed by traditional rule-based tools.
- The presenter demonstrated Gemini's (likely using a powerful model like Gemini 3 Pro) ability to generate a random number (8) but also showed its internal 'thinking' process, revealing it used Python's random module, not genuine randomness, and later, when prompted for a non-random number, it chose 7, noting 7 is the most common human guess.
- The video highlights the shift in AI development towards 'Reinforcement Learning with Performance Feedback' (RLPF) to train models based on real-world advertiser data for improved outcomes, as shown in the Meta/Facebook example.
- The speaker encourages viewers to support the channel by using the 'Hype' feature on YouTube mobile to help creators get discovered.

![Screenshot at 00:44: The presenter shows a slide from a Q&A with an AI security expert discussing the need for developer preparation when controlling super-intelligent AI, specifically mentioning sandbagging capabilities during evaluation.](https://ss.rapidrecap.app/screens/1gP-cFwko00/00-00-44.png)

**Context:** The video presents a news roundup focusing on recent advancements and limitations in Artificial Intelligence. The host examines three distinct AI stories: an innovative application in gaming from Ubisoft, a breakthrough in medical diagnostics from Harvard Medical School, and new research questioning the depth of understanding in current large language models (LLMs) regarding humor and security tasks like password cracking.

## Detailed Analysis

The video covers several recent developments in AI. First, Ubisoft, the maker of 'Assassin's Creed,' unveiled an experimental game demo called 'Teammates' powered by generative AI. This demo uses player voice commands to interact with AI-powered assistants, like 'Jasper,' generating improvised, context-aware dialogue that moves beyond fixed storylines, creating a more fluid and adaptive gameplay experience. Second, a new AI model from Harvard Medical School, named PopEVE, promises to accelerate the diagnosis of rare diseases by rapidly flagging and prioritizing potentially harmful genetic variants across large human genome databases. The presenter notes this could lead to better outcomes and earlier intervention for patients suffering from conditions that have long eluded detection. Third, the video addresses the limitations of current LLMs in understanding nuance, using research comparing GPT-4o's performance against humans on pun detection. GPT-4o correctly identified a pun based on word meanings but failed a modified pun involving the word 'ukulele,' classifying it as nonsense, whereas a human correctly understood the joke structure. This suggests LLMs rely heavily on memorized patterns rather than true linguistic comprehension. This lack of deep understanding also impacts security tasks; research shows LLMs are poor at cracking passwords compared to traditional rule-based tools because they lack the necessary domain reasoning and context. Finally, the presenter demonstrates Gemini's (using a tool for randomness) tendency to favor non-random numbers (like 7) when asked for a random number between 1 and 10, showing that even when attempting randomness, the model defaults to statistically common human biases. The video concludes by pointing to the trend of AI systems learning from real-world performance data via Reinforcement Learning with Performance Feedback (RLPF), as seen in Facebook's success in optimizing ad click-through rates.

### Ubisoft's Generative AI Game Demo

- Unveiled 'Teammates' game demo using generative AI
- built around player voice commands for real-time, improvised dialogue with AI assistants like 'Jasper'
- creates a more fluid, adaptive experience beyond pre-written storylines

### AI in Medical Diagnosis

- Harvard Medical School developed 'PopEVE' AI model
- rapidly flags and prioritizes genetic variants linked to severe diseases in human DNA databases
- promises faster diagnosis and personalized treatments for rare conditions

### LLM Limitations in Humor/Puns

- Research compared GPT-4o vs. Human pun detection
- GPT-4o failed to grasp a modified pun ('ukulele' substitution) while humans succeeded
- suggests LLMs lack deep linguistic comprehension, relying on pattern memorization

### LLMs vs. Password Cracking

- LLMs performed poorly on password cracking compared to traditional rule-based tools
- LLMs lack domain reasoning needed for complex security tasks
- performance was significantly worse than 50% accuracy on some tests

### Demonstration of Gemini's 'Randomness'

- Gemini chose '7' when asked for a random number between 1 and 10, which the presenter noted is the most common human choice
- Gemini admitted its output was based on Python's 'random' module, not true randomness

### Reinforcement Learning with Performance Feedback (RLPF)

- Meta's RLPF system uses advertiser performance data (e.g., CTRs) to train models like AdLlama
- resulted in a +6.7% advertiser CTR increase in a large-scale A/B test, showcasing the value of real-world feedback loops

### Call to Action

- The presenter encouraged viewers to use the YouTube 'Hype' button to support the channel's discovery.

![Screenshot at 00:00: Headline for Ubisoft's generative AI game demo, 'Teammates,' powered by AI dialogue.](https://ss.rapidrecap.app/screens/1gP-cFwko00/00-00-00.png)
![Screenshot at 00:44: Q&A slide discussing developer preparation for controlling super-intelligent AI, mentioning sandbagging during evaluation.](https://ss.rapidrecap.app/screens/1gP-cFwko00/00-00-44.png)
![Screenshot at 00:57: A graph illustrating 'Tensor Economics' showing trade-offs between tokens produced \(cumulative TPS\) and batch size.](https://ss.rapidrecap.app/screens/1gP-cFwko00/00-00-57.png)
![Screenshot at 01:06: Article headline announcing a new AI model, 'PopEVE,' from Harvard Medical School for speeding up rare disease diagnosis.](https://ss.rapidrecap.app/screens/1gP-cFwko00/00-01-06.png)
![Screenshot at 01:11: Article headline about research revealing why LLMs are not great at cracking passwords due to limited domain reasoning.](https://ss.rapidrecap.app/screens/1gP-cFwko00/00-01-11.png)
