xAI's new model is insane...

Quick Overview

Grok 4.1 achieved peak performance post-training, ranking #1 on the LMSYS Chatbot Arena leaderboard and setting a new standard in emotional intelligence benchmarks like EQ-Bench3, while also significantly reducing factual hallucinations compared to previous versions.

Key Points: Grok 4.1 achieved the #1 overall position on the LM Arena Text leaderboard with a 1483 Elo score for its thinking mode. Grok 4.1's non-reasoning mode achieved a 1585 Elo score, beating the next non-xAI model by a 31-point margin. The development used a massive 10x increase in RL compute for reasoning tasks compared to Grok 4, leveraging infrastructure originally built for Grok 4. Grok 4.1 demonstrated significant improvements in emotional intelligence, scoring highly on the EQ-Bench3 benchmark, where it significantly outperformed Grok 4. The non-reasoning model showed a marked reduction in factual hallucinations, dropping from 12.09% (Grok 4 Fast) to 4.22% (Grok 4.1 Non-Reasoning) on the FActScore benchmark. The new model can be explicitly selected as "Grok 4.1" immediately in Auto mode on grok.com, X, and mobile apps. Elon Musk suggested that Grok 5, which will use a 6-trillion parameter model, has a non-zero chance of achieving Artificial General Intelligence (AGI) by Q1.

Context: This video details the release and performance improvements of xAI's new large language model, Grok 4.1, which was announced on November 17, 2025. The presentation features Elon Musk discussing the technical advancements, particularly in reinforcement learning (RL) and reasoning capabilities, alongside commentary from a third-party analyst reviewing the performance data from benchmarks like LM Arena and EQ-Bench3, as well as Elon Musk's earlier speculation about AGI potential with Grok 5.

Detailed Analysis

The video announces the release of Grok 4.1, available immediately on grok.com, X, and mobile apps, with explicit selection available in the model picker. Elon Musk claims Grok 5 will have a non-zero chance of AGI by Q1, noting that the massive leap in reasoning compute (10x more RL compute than Grok 4) was crucial. Grok 4.1 shows significant improvements across general domains, including emotional intelligence and reduced hallucinations. On the LM Arena Text leaderboard, Grok 4.1 Thinking (code name quasar-flux) ranks #1 with an Elo of 1586, beating the next non-xAI model, Polaris Alpha (early GPT 5.1), by 24 points, and beating Grok 4 by 380 points. The non-reasoning version (code name tensor) ranks #2 overall at 1465 Elo. Regarding hallucinations, Grok 4.1's non-reasoning mode achieved a 4.22% hallucination rate, a significant drop from Grok 4 Fast's 12.09%, and a superior FActScore of 2.97% compared to Grok 4 Fast's 9.89%. The improvements were achieved by applying the large-scale RL infrastructure used for Grok 4 to optimize style, personality, helpfulness, and alignment of the new model, utilizing frontier agentic reasoning models as reward models. The model also shows better EQ performance (1585 Elo on EQ-Bench) and improved custom instruction handling compared to previous versions.

Raw markdown version of this recap