# xAI's new model is insane...

Source: https://www.youtube.com/watch?v=wIR6tRlxgp4
Recap page: https://rapidrecap.app/video/wIR6tRlxgp4
Generated: 2025-11-18T03:32:14.04+00:00

---
## Quick Overview

Grok 4.1 achieved peak performance post-training, ranking #1 on the LMSYS Chatbot Arena leaderboard and setting a new standard in emotional intelligence benchmarks like EQ-Bench3, while also significantly reducing factual hallucinations compared to previous versions.

**Key Points:**
- Grok 4.1 achieved the #1 overall position on the LM Arena Text leaderboard with a 1483 Elo score for its thinking mode.
- Grok 4.1's non-reasoning mode achieved a 1585 Elo score, beating the next non-xAI model by a 31-point margin.
- The development used a massive 10x increase in RL compute for reasoning tasks compared to Grok 4, leveraging infrastructure originally built for Grok 4.
- Grok 4.1 demonstrated significant improvements in emotional intelligence, scoring highly on the EQ-Bench3 benchmark, where it significantly outperformed Grok 4.
- The non-reasoning model showed a marked reduction in factual hallucinations, dropping from 12.09% (Grok 4 Fast) to 4.22% (Grok 4.1 Non-Reasoning) on the FActScore benchmark.
- The new model can be explicitly selected as "Grok 4.1" immediately in Auto mode on grok.com, X, and mobile apps.
- Elon Musk suggested that Grok 5, which will use a 6-trillion parameter model, has a non-zero chance of achieving Artificial General Intelligence (AGI) by Q1.

![Screenshot at 00:00: Elon Musk participating in a remote interview, discussing the significance of the Grok 5 model's potential for achieving AGI, setting the stage for the announcement of Grok 4.1.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-00-00.png)

**Context:** This video details the release and performance improvements of xAI's new large language model, Grok 4.1, which was announced on November 17, 2025. The presentation features Elon Musk discussing the technical advancements, particularly in reinforcement learning (RL) and reasoning capabilities, alongside commentary from a third-party analyst reviewing the performance data from benchmarks like LM Arena and EQ-Bench3, as well as Elon Musk's earlier speculation about AGI potential with Grok 5.

## Detailed Analysis

The video announces the release of Grok 4.1, available immediately on grok.com, X, and mobile apps, with explicit selection available in the model picker. Elon Musk claims Grok 5 will have a non-zero chance of AGI by Q1, noting that the massive leap in reasoning compute (10x more RL compute than Grok 4) was crucial. Grok 4.1 shows significant improvements across general domains, including emotional intelligence and reduced hallucinations. On the LM Arena Text leaderboard, Grok 4.1 Thinking (code name quasar-flux) ranks #1 with an Elo of 1586, beating the next non-xAI model, Polaris Alpha (early GPT 5.1), by 24 points, and beating Grok 4 by 380 points. The non-reasoning version (code name tensor) ranks #2 overall at 1465 Elo. Regarding hallucinations, Grok 4.1's non-reasoning mode achieved a 4.22% hallucination rate, a significant drop from Grok 4 Fast's 12.09%, and a superior FActScore of 2.97% compared to Grok 4 Fast's 9.89%. The improvements were achieved by applying the large-scale RL infrastructure used for Grok 4 to optimize style, personality, helpfulness, and alignment of the new model, utilizing frontier agentic reasoning models as reward models. The model also shows better EQ performance (1585 Elo on EQ-Bench) and improved custom instruction handling compared to previous versions.

### Grok 4.1 Release & Availability

- Grok 4.1 is available immediately on grok.com, X, and mobile apps, selectable explicitly in Auto mode
- Grok 5 predicted to have a non-zero chance of AGI by Q1 based on 6-trillion parameter model.

### Reasoning & Compute Leap

- Grok 4.1 used 10x more RL compute than Grok 4 for reasoning, applied to the same large-scale RL infrastructure used for Grok 4.

### LM Arena Text Performance

- Grok 4.1 Thinking scores 1586 Elo (#1), Grok 4.1 (non-reasoning) scores 1465 Elo (#2), significantly ahead of competitors like Gemini 2.5 Pro (1452 Elo).

### Factual Grounding (Reduced Hallucinations)

- Non-reasoning hallucination rate dropped from 12.09% (Grok 4 Fast) to 2.97% (Grok 4.1 Non-Reasoning); FActScore improved from 9.89% to 2.97%.

### Emotional Intelligence (EQ-Bench3)

- Grok 4.1 scored 1585 Elo, significantly better than Grok 4 (1206 Elo), demonstrating better handling of emotional scenarios.

### Customization & Memory

- Improved custom instructions functionality; models show better personality adherence (e.g., avoiding overly enthusiastic emojis) across multi-turn conversations.

### Solar Power Analogy

- Elon Musk used the calculation that powering a 1GW data center in space requires roughly 2.44 sq km of solar panels, which is far less than the 14,250 sq km needed on Earth, illustrating efficiency gains.

![Screenshot at 00:00: Elon Musk in a video conference discussing the significance of the Grok 5 timeline for AGI.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-00-00.png)
![Screenshot at 00:21: Screen displaying the official announcement for Grok 4.1 availability across grok.com, X, and mobile apps.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-00-21.png)
![Screenshot at 00:39: The presenter reviewing the Hostinger sponsorship details and offering a discount code.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-00-39.png)
![Screenshot at 00:54: Elon Musk comparing the anticipated capabilities of Grok 5 to achieving AGI, citing a 10% chance.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-00-54.png)
![Screenshot at 03:41: Demonstration of setting up an n8n workflow on Hostinger using their AI assistant, Kodee.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-03-41.png)
![Screenshot at 06:56: Chart illustrating the 'Ludicrous rate of progress' in compute usage across Grok versions, highlighting the massive increase for Grok 4 reasoning.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-06-56.png)
![Screenshot at 07:46: Tweet showing Grok 4 leading the Kimi K2 Thinking benchmark, with Grok 4.1 Thinking now at #1 \(1586 Elo\) in the LM Arena Text leaderboard.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-07-46.png)
![Screenshot at 10:01: Hand-drawn diagram illustrating the concept of Reinforcement Learning with Human Feedback \(RLHF\) leading to RLAF \(Reinforcement Learning with AI Feedback\).](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-10-01.png)
![Screenshot at 15:16: LM Arena Creative Writing v3 leaderboard showing Grok 4.1 Thinking at #2 \(1721.9\) and Grok 4.1 at #3 \(1706.6\), below Polaris Alpha \(1756.2\).](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-15-16.png)
![Screenshot at 16:32: Bar chart comparing hallucination rates, showing Grok 4.1 Non-Reasoning at 4.22% versus Grok 4 Fast Non-Reasoning at 12.09%.](https://ss.rapidrecap.app/screens/wIR6tRlxgp4/00-16-32.png)
