What the Freakiness of 2025 in AI Tells Us About 2026
Quick Overview
The year 2025 in AI is predicted to be characterized by intense competition, with models like Gemini 3 Pro, Grok 4, and Claude 4 Opus dominating benchmarks, while the development of truly general intelligence is still uncertain, evidenced by the differing timescales predicted by experts for automating 99% of remote jobs (ranging from 4 years to 40 years) and the continued reliance on human expertise for complex problem-solving like novel algorithm design.
Key Points: The speaker predicts 2026 will be the 'freakiest year in AI' due to 12 months of weird progress, citing the rapid advancement seen in LLM release data up to 2026. Gemini 3 Pro is currently leading benchmarks like GPQA Diamond (91.9%) and AIME 2025 (95.0% without code execution), showing strong reasoning capabilities. The concept of 'vibe coding' is introduced: models are getting so good that they are being used for tasks like writing complex algorithms, but there is a disconnect between benchmark performance and real-world performance. An expert from Anthropic (Dario Amodei) warns that AI could wipe out half of entry-level white-collar jobs in the next 1-5 years, leading to significant societal change that many are unprepared for. The METR benchmark study is criticized for having a small sample size (14 tasks in the 1-4 hour range) and potential overfitting or noise, suggesting that its predictions for AGI timelines (like 2027) might be too aggressive. Google DeepMind's AlphaEvolve demonstrates AI's ability to invent novel, provably correct algorithms, showcasing progress in areas like data center scheduling (0.7% compute savings) and hardware design (Verilog rewrite). The discussion closes by contrasting the rapid, sometimes brittle gains in narrow capabilities (like coding/algorithm design) with the uncertainty surrounding true general intelligence, as highlighted by the wide range of expert predictions for job automation (4 years to 40 years).
Context: This video synthesizes predictions and current developments in AI, particularly Large Language Models (LLMs) and Artificial General Intelligence (AGI), leading up to 2026. The discussion is framed around recent benchmark results (like those from the original METR paper) and contrasting viewpoints from industry leaders, including Sam Altman (OpenAI) and Dario Amodei (Anthropic), alongside new research like Google DeepMind's AlphaEvolve, which focuses on AI-driven scientific discovery.