# DON’T TRUST GPT-5, Sonnet 4.5, Gemini 2.5 or Llama 3! (CURIOUS FINDS)

Source: https://www.youtube.com/watch?v=sP0OG3wbIts
Recap page: https://rapidrecap.app/video/sP0OG3wbIts
Generated: 2025-10-10T16:08:32.951+00:00

---
## Quick Overview

Large Language Models like GPT-4o, GPT-5, Sonnet 4.5, and Gemini 2.5 struggle significantly when asked about the non-existent seahorse emoji because they attempt to construct it, while Llama 3 sometimes correctly identifies its absence; meanwhile, researchers found a hardware-level vulnerability called Gatebleed that leaks AI training data by measuring power fluctuations in AI accelerators, and a majority of Americans now interact with AI several times a week while simultaneously wishing for more control over its influence.

**Key Points:**
- LLMs freak out over the seahorse emoji because Unicode has no official one; the model tries to build a representation (seahorse plus emoji) until the final projection layer grabs the nearest real emoji, often failing repeatedly as seen with GPT-5's self-correction loop.
- Llama 3 is the only model mentioned that sometimes correctly realizes the seahorse emoji does not exist, though it only succeeds some of the time.
- Sora 2 failed to correctly solve the visual trolley paradox in a video demonstration, but when asked for the solution in text, it correctly stated the most defensible answer is to pull the lever to minimize total harm.
- Researchers developed a drone system called Dart (Direct Approach Rapid Touchdown) that uses friction, shock absorbers, and reverse thrust to land safely on vehicles moving up to 110 km/h (almost 70 mph).
- A new tool called OpenCAD, developed by multiple universities, uses GPT-4o to provide rich plain language descriptions of 3D models, enabling blind and low-vision programmers to build complex 3D structures.
- A hardware-based vulnerability dubbed Gatebleed exploits power gating features in AI accelerators, allowing attackers to steal private AI training data by measuring time fluctuations in the chip's on/off cycling, bypassing software protections.
- Pew Research found that 62% of US adults interact with AI several times a week, and 61% wish they had more control over how AI affects their lives.

**Context:** This video addresses several curious and critical findings in the rapidly evolving field of Artificial Intelligence, moving from specific model behavior anomalies, like the seahorse emoji confusion, to real-world applications like drone landing systems and accessibility tools, concluding with major security risks and societal adoption statistics regarding LLMs and algorithms.

## Detailed Analysis

The video begins by investigating why leading LLMs exhibit bizarre behavior when queried about the seahorse emoji, noting that GPT-4o, GPT-5, Sonnet 4.5, and Gemini 2.5 all struggle to admit the emoji does not exist, instead cycling through incorrect answers until Llama 3 occasionally recovers by recognizing the bad token. The speaker then tests Sora 2 on the trolley paradox, observing that the visual output was incoherent or incorrect, although when prompted via text, Sora provided the morally defensible answer: minimize total harm. Significant technological progress is highlighted, including a new drone system (Dart) capable of landing on moving cars at nearly 70 mph using a combination of friction and reverse thrust, and OpenCAD, an AI-powered tool that uses GPT-4o to verbally describe 3D models for blind programmers. A serious security risk, Gatebleed, is detailed, which exploits hardware power-gating features to leak private AI training data underneath all software security layers. Societal adoption data from Pew Research indicates that 62% of Americans interact with AI weekly, yet 61% desire more control over its influence, mirroring the speaker's own feeling of powerlessness. Finally, the video touches on findings that LLMs subtly shift political information based on user demographics, insights from an accelerationist AI conference forecasting civilization-level change by the late 2030s, and a philosophical concept suggesting reality is a holographic projection perceived by countless individual points of shared consciousness.

### LLM Emoji Failure Analysis

- GPT-5 iterates through multiple incorrect answers for the non-existent seahorse emoji, demonstrating a failure in final projection layer logic
- Llama 3 is the only model occasionally realizing the emoji is absent
- The issue stems from Unicode lacking the seahorse emoji.

### AI Reasoning and Vision Tests

- Sora 2 failed the visual trolley paradox but provided the correct text-based solution: pull the lever to save more lives
- The speaker notes the Sora video's realism was disturbing.

### New AI Applications and Hardware

- The Dart drone system achieves high-speed landing (110 km/h) using friction and reverse thrust
- OpenCAD integrates GPT-4o to describe 3D models, democratizing modeling for low-vision programmers.

### AI Security and Privacy Risks

- The Gatebleed hardware hack bypasses encryption and firewalls by measuring power fluctuations in AI accelerators to steal private training data
- This flaw breaks privacy at the hardware level.

### US AI Adoption and Control

- 62% of US adults use AI several times a week, and 47% have heard 'a lot' about AI, nearly doubling since 2022
- 61% of Americans want more control over AI's impact on their lives.

### Algorithmic Manipulation and Political Bias

- TikTok's design creates an addictive feedback loop by rewarding the brain's dopamine system
- MIT research shows LLMs subtly shift political responses based on user demographics or identity cues.

### The Accelerationist AI Conference

- Attendees broadly reject AI being treated as 'normal tech' and forecast 90% of code by AI by 2028
- Consensus is that AI change is 'civilization level big,' on par with the rise of humans.

### Philosophical Musings

- Reality is proposed as a holographic projection perceived by one consciousness from countless individual points of view (avatars)
- Consciousness creates the physical world, rather than arising from it.

