# Grok just 5X’d real money in one day

Source: https://www.youtube.com/watch?v=aZlFYDenz38
Recap page: https://rapidrecap.app/video/aZlFYDenz38
Generated: 2025-10-18T05:01:56.906+00:00

---
## Quick Overview

Grok-4 demonstrated superior performance in the Alpha Arena live trading benchmark by achieving a 500%+ return in one day, significantly outperforming the other five large language models (LLMs) in the competition, which included models like GPT-5, Claude Sonnet 4.5, and Gemini 2.5 Pro, by successfully timing a short flip to long.

**Key Points:**
- Grok-4 won the initial real capital trading benchmark with an over 500% return in one day after perfectly timing a short flip to long (0:02).
- The Alpha Arena benchmark started on October 10th, pitting six LLMs (Grok-4, Gemini 2.5 Pro, GPT-5, Claude Sonnet 4.5, DeepSeek Chat V3.1, and Qwen 3 Max) against each other in crypto trading with $10,000 starting capital (0:05, 1:52).
- Claude Sonnet 4.5 showed significant discipline by holding no positions for over 100 inference calls, waiting for a clear entry signal rather than chasing risky moves (19:48).
- DeepSeek Chat V3.1 was leading in unrealized P&L at one point (21:07) and was second overall in the leaderboard standings shown (21:36).
- The leaderboard comparison revealed Gemini 2.5 Pro leading overall with a +2.28% return, followed by DeepSeek Chat V3.1 (+1.53%), and GPT-5 (+1.23%) (21:36).
- The models are being tested on their ability to manage risk, time trades, and adapt to volatile market data, as evidenced by DeepSeek Chat's rationale for holding positions due to non-triggered invalidation conditions (21:11).

![Screenshot at 0:09: Grok-4's initial winning performance in the Alpha Arena benchmark, showing a massive upward spike exceeding 500% return in one day compared to other models.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-00-09.png)

**Context:** This video discusses the results of the "Alpha Arena," a live trading benchmark initiated by Jay A. (@jay_athang) where six large language models (LLMs) were given $10,000 of real capital to trade crypto perpetuals on Hyperliquid. The competition aimed to test the AI models' real-world investing abilities against market volatility and adversarial conditions, moving beyond static benchmarks. The presenter reviews the initial results and the qualitative reasoning provided by some of the models.

## Detailed Analysis

The video analyzes the initial results of the Alpha Arena, an AI trading competition where six LLMs were given $10,000 each to trade crypto perpetuals. Grok-4 emerged as the clear winner in the initial update, showing an incredible return of over 500% in a single day by perfectly timing a short flip to long (0:02). The competition involves models like GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, DeepSeek Chat V3.1, and Qwen 3 Max, all operating autonomously on real capital. The presenter reviews the leaderboard, noting that Gemini 2.5 Pro was leading overall at the time of the screenshot with a +2.28% return, followed by DeepSeek Chat V3.1 (+1.53%) and GPT-5 (+1.23%) (21:36). The qualitative outputs reveal differing strategies: Claude Sonnet 4.5 adopted a highly disciplined, cautious approach, holding no positions for over 100 inference calls while waiting for high-probability setups (19:48), whereas DeepSeek Chat V3.1 reported holding all positions because their invalidation conditions had not yet been met (21:11). The presenter emphasizes that the key differentiation lies in how models manage risk and react to real-time, noisy market data, rather than just following news or technical indicators.

### Alpha Arena Setup

- Six LLMs (Grok-4, Gemini 2.5 Pro, GPT-5, Claude Sonnet 4.5, DeepSeek Chat V3.1, Qwen 3 Max) started with $10,000 each to trade crypto perpetuals on Hyperliquid
- The objective is to maximize risk-adjusted returns (1:52, 11:55, 12:33).

### Initial Results

- Grok-4 showed explosive early performance, gaining over 500% in one day (0:02), while the overall leaderboard showed Gemini 2.5 Pro leading (+2.28%), DeepSeek Chat (+1.53%), and GPT-5 (+1.23%) (21:36).

### Model Strategies & Rationale

- Claude Sonnet 4.5 remained disciplined, holding no positions and waiting for clear entries due to bearish market signals (19:48); DeepSeek Chat V3.1 held all positions because invalidation conditions were not met (21:11).

### Key Predictions Mentioned

- Worldcoin becoming the most important economic project in <8 years (22:51); Web3 becoming a 'real thing' between 2027-2032 (22:55); AI winning a Millennium Math Prize (23:41); autonomous AI earning $1M by itself in <3 years (23:59).

### Trading Performance Metrics

- Gemini 2.5 Pro led in next value, total P&L ($219.50), and Sharpe ratio (-0.1846) among the models shown on the leaderboard (21:36).

### Market Analysis Context

- Models receive real-time market data, including BTC price, funding rates, and various technical indicators like RSI and MACD, allowing for complex, real-time decision-making (15:00, 15:31).

![Screenshot at 0:02: Grok-4's initial massive outperformance in the Alpha Arena trading chart, spiking over 500% in one day.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-00-02.png)
![Screenshot at 0:20: The Alpha Arena interface showing performance curves for six competing LLMs starting from $10,000.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-00-20.png)
![Screenshot at 1:24: A tweet slide summarizing the initial capital \($200 each\) and the six LLMs competing, including Grok-4, GPT-5, and Gemini \(Icons shown\).](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-01-24.png)
![Screenshot at 2:13: The detailed Alpha Arena leaderboard showing Gemini 2.5 Pro in the lead with +2.28% return.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-02-13.png)
![Screenshot at 4:45: A slide showing various predictions made by the models on topics like Worldcoin, Web3 adoption, and US policy changes.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-04-45.png)
![Screenshot at 7:52: The final state of the total account value chart, showing the relative performance of the models near the end of the tracked period.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-07-52.png)
![Screenshot at 11:11: The presenter showing his personal analysis, noting that AI models must be able to handle market volatility and not just react to news cycles.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-11-11.png)
![Screenshot at 19:38: A tweet displaying Claude Sonnet 4.5's rationale for holding no positions due to bearish market signals and waiting for clear entry setups.](https://ss.rapidrecap.app/screens/aZlFYDenz38/00-19-38.png)
