Grok just 5X’d real money in one day
Quick Overview
Grok-4 demonstrated superior performance in the Alpha Arena live trading benchmark by achieving a 500%+ return in one day, significantly outperforming the other five large language models (LLMs) in the competition, which included models like GPT-5, Claude Sonnet 4.5, and Gemini 2.5 Pro, by successfully timing a short flip to long.
Key Points: Grok-4 won the initial real capital trading benchmark with an over 500% return in one day after perfectly timing a short flip to long (0:02). The Alpha Arena benchmark started on October 10th, pitting six LLMs (Grok-4, Gemini 2.5 Pro, GPT-5, Claude Sonnet 4.5, DeepSeek Chat V3.1, and Qwen 3 Max) against each other in crypto trading with $10,000 starting capital (0:05, 1:52). Claude Sonnet 4.5 showed significant discipline by holding no positions for over 100 inference calls, waiting for a clear entry signal rather than chasing risky moves (19:48). DeepSeek Chat V3.1 was leading in unrealized P&L at one point (21:07) and was second overall in the leaderboard standings shown (21:36). The leaderboard comparison revealed Gemini 2.5 Pro leading overall with a +2.28% return, followed by DeepSeek Chat V3.1 (+1.53%), and GPT-5 (+1.23%) (21:36). The models are being tested on their ability to manage risk, time trades, and adapt to volatile market data, as evidenced by DeepSeek Chat's rationale for holding positions due to non-triggered invalidation conditions (21:11).
Context: This video discusses the results of the "Alpha Arena," a live trading benchmark initiated by Jay A. (@jayathang) where six large language models (LLMs) were given $10,000 of real capital to trade crypto perpetuals on Hyperliquid. The competition aimed to test the AI models' real-world investing abilities against market volatility and adversarial conditions, moving beyond static benchmarks. The presenter reviews the initial results and the qualitative reasoning provided by some of the models.
Detailed Analysis
The video analyzes the initial results of the Alpha Arena, an AI trading competition where six LLMs were given $10,000 each to trade crypto perpetuals. Grok-4 emerged as the clear winner in the initial update, showing an incredible return of over 500% in a single day by perfectly timing a short flip to long (0:02). The competition involves models like GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, DeepSeek Chat V3.1, and Qwen 3 Max, all operating autonomously on real capital. The presenter reviews the leaderboard, noting that Gemini 2.5 Pro was leading overall at the time of the screenshot with a +2.28% return, followed by DeepSeek Chat V3.1 (+1.53%) and GPT-5 (+1.23%) (21:36). The qualitative outputs reveal differing strategies: Claude Sonnet 4.5 adopted a highly disciplined, cautious approach, holding no positions for over 100 inference calls while waiting for high-probability setups (19:48), whereas DeepSeek Chat V3.1 reported holding all positions because their invalidation conditions had not yet been met (21:11). The presenter emphasizes that the key differentiation lies in how models manage risk and react to real-time, noisy market data, rather than just following news or technical indicators.