# BREAKING: Elon just revealed Grok 4.20

Source: https://www.youtube.com/watch?v=EnjrRDwycK0
Recap page: https://rapidrecap.app/video/EnjrRDwycK0
Generated: 2025-12-05T09:34:12.195+00:00

---
## Quick Overview

Elon Musk revealed that the "mystery AI model" outperforming others in a trading competition is an experimental version of Grok 4.20, which uses real-money trading data across various market conditions, confirming its superior performance compared to other large language models.

**Key Points:**
- Elon Musk confirmed the high-performing "mystery AI model" is an experimental version of Grok 4.20.
- The model achieved a 12.11% aggregate return over two weeks in the Alpha Arena competition, starting with $10,000 and ending with $11,782.
- The competition involved trading stocks, crypto, and news/sentiment data across different modes like 'Max Leverage' and 'Monk Mode'.
- In the 'Situational Awareness' competition, the Mystery Model (Grok 4.20) was the clear winner with a +17.82% return, while all other top models were negative.
- The model's success is attributed to its ability to process and utilize real-time trading data, including market structure and news analysis, as detailed in its reasoning logs.
- Wes Roth, who initially identified the model, corrected his assumption that it was the ProFit model after Musk's clarification, noting its huge potential impact if released.
- The underlying technology may relate to research like ProFit (Program Search for Financial Trading), which uses LLMs in an evolutionary framework for automated trading strategy discovery.

![Screenshot at 00:00: Elon Musk confirming via Twitter that the successful "mystery AI model" is an experimental version of Grok 4.20, responding to Wes Roth's surprise.](https://ss.rapidrecap.app/screens/EnjrRDwycK0/00-00-00.png)

**Context:** The video discusses the revelation surrounding a successful 'mystery AI model' participating in a live trading competition called Alpha Arena Season 1.5, where various Large Language Models (LLMs) traded real capital across different asset classes and strategies. The initial confusion stemmed from the model's anonymity, but Elon Musk ultimately confirmed its identity via Twitter, sparking discussion about the capabilities of advanced AI in financial markets.

## Detailed Analysis

The video centers on the reveal of the identity of a top-performing, initially anonymous AI model in the Alpha Arena trading competition. Elon Musk confirmed on Twitter that this "mystery AI model" is actually an experimental version of Grok 4.20. The competition concluded with the Mystery Model achieving a 12.11% aggregate return in two weeks, turning an initial $10,000 into $11,782, while most other top LLMs (including GPT-5.1, Gemini 3-Pro, Qwen3-Max, DeepSeek, Kimi, Claude, and Grok-4) registered losses in the overall aggregate index. The model achieved this success across different competition modes, including 'Situational Awareness' (where it returned +17.82%) and 'Max Leverage' (where it also performed well). The underlying research that informs these trading strategies appears linked to papers like ProFit (Program Search for Financial Trading), which utilizes LLMs within an evolutionary framework to discover and improve trading strategies. The presenter highlighted the detailed reasoning logs provided by the models for their trades, showing sophisticated analysis of market structure, sentiment, and technical indicators, suggesting a significant leap in AI-driven financial execution.

### Grok 4.20 Identity Reveal

- Elon Musk confirmed the mystery model is Grok 4.20 experimental version
- Wes Roth initially assumed it was the ProFit model, then posted corrections
- This confirmation generated significant discussion about AI's real-world financial impact.

### Alpha Arena Season 1.5 Results

- Mystery Model won with 12.11% aggregate return over two weeks
- All other top competitors, including GPT-5.1 and Gemini 3-Pro, generally lost money in the aggregate index.

### Competition Performance Breakdown

- In the 'Situational Awareness' competition, the Mystery Model returned +17.82% ($14,757), significantly outperforming the next best model by over 23 percentage points
- In 'Max Leverage', the model also finished first, though with a smaller lead (+12.11% return).

### Underlying Technology & Data

- The models trade real money across various assets and use proprietary data feeds, including news/sentiment data every six minutes
- The success is linked to advanced LLM techniques applied to trading, possibly drawing from research like ProFit (Program Search for Financial Trading).

### Trade Examples

- The model demonstrated precise trade execution, such as taking profit on an NVDA trade exactly where it called the local top, yielding a net P&L of $3083.83 on a $95k notional value.

![Screenshot at 00:00: Elon Musk confirming via Twitter that the successful "mystery AI model" is an experimental version of Grok 4.20, responding to Wes Roth's surprise.](https://ss.rapidrecap.app/screens/EnjrRDwycK0/00-00-00.png)
![Screenshot at 00:14: The initial leaderboard screenshot showing the 'Mystery Model' leading with a +17.82% return in the 'Situational Awareness' competition, significantly ahead of all others.](https://ss.rapidrecap.app/screens/EnjrRDwycK0/00-00-14.png)
![Screenshot at 01:04: The presenter pointing out the final aggregate results, confirming the Mystery Model's $11,782 equity versus the $9,630 achieved by GPT-5.1.](https://ss.rapidrecap.app/screens/EnjrRDwycK0/00-01-04.png)
![Screenshot at 01:21: Wes Roth's follow-up tweet stating he dug into the model and that it is 'not what you think,' referencing its potential market impact.](https://ss.rapidrecap.app/screens/EnjrRDwycK0/00-01-21.png)
![Screenshot at 02:06: The chart showing the Mystery Model's equity curve rising sharply to nearly $14,000 in the 'Situational Awareness' run, while others generally declined.](https://ss.rapidrecap.app/screens/EnjrRDwycK0/00-02-06.png)
