# GROK 4 STUNNING New Ability? Emerging "Fluid Intelligence" in AI Models?

Source: https://www.youtube.com/watch?v=XTqEOt1EI84
Recap page: https://rapidrecap.app/video/XTqEOt1EI84
Generated: 2025-07-17T04:33:24.919+00:00

---
## Quick Overview

Grok 4 demonstrates a stunning new ability by exhibiting "nonzero levels of fluid intelligence," allowing it to adapt and solve novel problems efficiently, a capability previously lacking in large language models. It achieved a 16% accuracy on the ARC AGI benchmark, significantly outperforming competitors, and also dominated the "vending machine" business simulation by nearly tenfold.

**Key Points:**
- Grok 4 achieved a 16% accuracy on the ARC AGI benchmark, a test for fluid intelligence, significantly surpassing the previous top score of 8% and outperforming purpose-built solutions.
- Greg Comrade, President of ARC AGI, confirmed Grok 4 shows "nonzero levels of fluid intelligence," indicating its ability to learn new skills and adapt to novel situations.
- Grok 4 demonstrated exceptional business acumen in a "vending machine" simulation, turning an initial $500 into approximately $4700, nearly ten times the human baseline of $844.
- Elon Musk's XAI achieved this leading position by investing heavily in compute, utilizing 100,000 H100 Nvidia GPUs and planning to scale to 200,000, including shipping an "overseas power plant" to Memphis.
- The model's success is attributed to a 10x increase in reinforcement learning (RL) compute for Grok 4's reasoning compared to Grok 3, suggesting that scaling RL compute is key to emerging abilities.
- Grok 4 is currently the number one public model on ARC AGI and the New York Times Connections benchmark, though its dedicated coding model is still weeks away from release.
- Despite Grok 4's current lead, upcoming models like Google's Gemini 3.0 Pro and OpenAI's GPT-5 (rumored to be "a tad over Grok 4 heavy" on internal evals) are expected to challenge its position soon.

**Context:** The video discusses the emergence of Grok 4, Elon Musk's latest AI model from XAI, and its surprising performance against established competitors like OpenAI's GPT models and Google DeepMind's Gemini. A central theme is the distinction between "crystallized intelligence" (drawing on vast knowledge) and "fluid intelligence" (solving novel problems and adapting), with the latter being a significant challenge for large language models. The ARC AGI benchmark, created by Francois Chollet, specifically measures fluid intelligence by assessing a model's efficiency in acquiring new skills on unknown tasks, rather than its proficiency in specific, pre-trained skills.

## Detailed Analysis

Grok 4, Elon Musk's latest AI model, has made a significant impact by demonstrating what appears to be "fluid intelligence," a critical ability for solving novel problems and adapting to new situations, which traditional large language models typically lack. This was evidenced by its unprecedented 16% accuracy on the ARC AGI benchmark, a test designed to measure skill acquisition on unknown tasks, far surpassing previous top scores of 8% and even outperforming custom-built solutions. Grok 4's success is largely attributed to XAI's massive investment in compute, deploying 100,000 Nvidia H100 GPUs and planning to double that, alongside a 10x increase in reinforcement learning (RL) compute for reasoning compared to its predecessor. This suggests that scaling RL compute is a key factor in unlocking new AI capabilities. Beyond benchmarks, Grok 4 showcased remarkable practical intelligence by turning $500 into approximately $4700 in a "vending machine" business simulation, outperforming human baselines and reigning champions like Claude Opus. While Grok 4 currently leads in reasoning and specific tasks like New York Times Connections, its dedicated coding model is still pending release, and strong competition from Google's Gemini 3.0 Pro and OpenAI's GPT-5 (rumored to slightly exceed Grok 4 Heavy in internal evaluations) is anticipated soon. The video concludes that the idea of scaling hitting a wall is misleading, as continued investment in compute, particularly RL, appears to foster the emergence of new, adaptive abilities in AI models.

### Grok 4's Benchmark Dominance

- Grok 4 and Grok 4 Heavy are "head and shoulders above the competition" on "humanity's last exam"
- It achieved a 16% accuracy on the ARC AGI benchmark, significantly higher than the previous top score of 8%
- Grok 4 is the top-performing publicly available model on ARC AGI, even outperforming purpose-built solutions on Kaggle
- It is also the number one model for the New York Times Connections benchmark.

### Emergence of Fluid Intelligence

- Fluid intelligence is defined as the ability to solve new problems and adapt to novel situations, distinct from crystallized intelligence (drawing on past knowledge)
- Large language models typically have high crystallized but low fluid intelligence
- Greg Comrade, President of ARC AGI, states Grok 4 shows "nonzero levels of fluid intelligence"
- The ARC AGI benchmark specifically tests fluid intelligence, focusing on "the efficiency of skill acquisition on unknown tasks."

### The Role of Compute in Grok 4's Success

- Elon Musk's XAI achieved its leading position by deploying "100,000 H100's Nvidia's GPUs" and plans to scale to 200,000
- XAI increased pre-training compute 10x from Grok 2 to Grok 3, and reinforcement learning (RL) compute 10x for Grok 4 reasoning compared to Grok 3
- The speaker suggests that scaling RL compute is crucial for emerging abilities, potentially "completely dwarfing the pre-training compute" in the future
- Elon Musk even ordered an "overseas power plant" to be shipped to Memphis to support compute needs.

### Grok 4's Practical Application and Business Acumen

- In a "vending machine" business experiment, Grok 4 turned an initial $500 into approximately $4700, nearly a 10x return
- This performance "wipes the floor with everybody," including previous champions like Claude 3.5 Sonnet and Claude Opus
- The experiment highlights Grok 4's ability to make a profit and manage a business, unlike models trained solely for helpfulness.

### Upcoming Competition and Future Outlook

- Google DeepMind's Gemini 3.0 Pro is expected soon, with its CEO and Google's CEO congratulating Elon Musk on Grok 4's release
- OpenAI's GPT-5 is also rumored to be released very soon, with internal evaluations suggesting it is "a tad over Grok 4 heavy" on some tests
- Grok 4's dedicated coding model is expected within weeks (around August), which will be the "ultimate test" for its coding capabilities
- The speaker believes that "scaling is not hitting a wall" and that throwing more compute at training and reinforcement learning continues to yield new abilities.

