# What the Freakiness of 2025 in AI Tells Us About 2026

Source: https://www.youtube.com/watch?v=FMMpUO1uAYk
Recap page: https://rapidrecap.app/video/FMMpUO1uAYk
Generated: 2025-12-23T18:03:47.129+00:00

---
## Quick Overview

The year 2025 in AI is predicted to be characterized by intense competition, with models like Gemini 3 Pro, Grok 4, and Claude 4 Opus dominating benchmarks, while the development of truly general intelligence is still uncertain, evidenced by the differing timescales predicted by experts for automating 99% of remote jobs (ranging from 4 years to 40 years) and the continued reliance on human expertise for complex problem-solving like novel algorithm design.

**Key Points:**
- The speaker predicts 2026 will be the 'freakiest year in AI' due to 12 months of weird progress, citing the rapid advancement seen in LLM release data up to 2026.
- Gemini 3 Pro is currently leading benchmarks like GPQA Diamond (91.9%) and AIME 2025 (95.0% without code execution), showing strong reasoning capabilities.
- The concept of 'vibe coding' is introduced: models are getting so good that they are being used for tasks like writing complex algorithms, but there is a disconnect between benchmark performance and real-world performance.
- An expert from Anthropic (Dario Amodei) warns that AI could wipe out half of entry-level white-collar jobs in the next 1-5 years, leading to significant societal change that many are unprepared for.
- The METR benchmark study is criticized for having a small sample size (14 tasks in the 1-4 hour range) and potential overfitting or noise, suggesting that its predictions for AGI timelines (like 2027) might be too aggressive.
- Google DeepMind's AlphaEvolve demonstrates AI's ability to invent novel, provably correct algorithms, showcasing progress in areas like data center scheduling (0.7% compute savings) and hardware design (Verilog rewrite).
- The discussion closes by contrasting the rapid, sometimes brittle gains in narrow capabilities (like coding/algorithm design) with the uncertainty surrounding true general intelligence, as highlighted by the wide range of expert predictions for job automation (4 years to 40 years).

![Screenshot at 00:00: The title slide sets the stage for the video's central theme: '2026 and the freakiest year in AI,' featuring a graph showing the rapid, accelerating release schedule of LLMs \(GPT-3, Claude, Gemini, Grok\) over time.](https://ss.rapidrecap.app/screens/FMMpUO1uAYk/00-00-00.jpg)

**Context:** This video synthesizes predictions and current developments in AI, particularly Large Language Models (LLMs) and Artificial General Intelligence (AGI), leading up to 2026. The discussion is framed around recent benchmark results (like those from the original METR paper) and contrasting viewpoints from industry leaders, including Sam Altman (OpenAI) and Dario Amodei (Anthropic), alongside new research like Google DeepMind's AlphaEvolve, which focuses on AI-driven scientific discovery.

## Detailed Analysis

The speaker forecasts that 2026 will be an exceptionally 'freaky' year in AI progress, based on the accelerating pace of LLM releases shown in a chart (0:00). Current benchmarks show Gemini 3 Pro leading in reasoning tasks, achieving 91.9% on GPQA Diamond and 95.0% on AIME 2025 (0:33). However, the speaker notes a disconnect between benchmark performance and real-world utility, referring to the models' ability to create complex outputs like novel algorithms while still making basic errors, which he terms 'vibe coding' (1:39, 2:08). The danger of current models is exemplified by an AI-generated text suggesting a person stopped taking medication because they believed radio signals caused their issues (0:01, 9:51), showcasing a lack of robust reasoning. The discussion shifts to broader implications, referencing Dario Amodei's warning that AI could automate half of entry-level white-collar jobs in 1-5 years (3:47). The speaker then discusses the METR benchmark, noting that while it shows exponential progress, the data is limited (only 14 samples in the 1-4 hour range), suggesting that extrapolations for AGI timelines might be overly optimistic (9:05, 13:33). The speaker's final key takeaway is that the field is currently in a confusing phase where many experts disagree on AGI timelines (2:27, 3:50), with job automation predictions ranging wildly from 4 years (Daniel) to 40 years (Ege) (23:31). Finally, Google DeepMind's AlphaEvolve paper (30:43) showcases AI's ability to discover novel algorithms and optimize data center scheduling (0:7:55), suggesting that while specific feats are impressive, the path to general intelligence remains unclear.

### LLM Benchmark Performance

- Gemini 3 Pro leads on GPQA Diamond (91.9%) and AIME 2025 (95.0%) without code execution
- Grok 4 ranks 6th with 60.5% score
- Claude Opus 4.5 shows high performance but large error bars on the time-horizon chart (13:09)

### The AI Timeline Debate

- Experts disagree on job automation timelines, with predictions for 99% automation ranging from 4 years (Daniel) to 40 years (Ege) (23:31)

### The Problem with Current Benchmarks

- The METR plot's small sample size (14 tasks in the 1-4 hr range) makes it easy to game, and the performance gains might not generalize to real-world tasks (13:33, 13:53)

### AI in Practical Application (AlphaEvolve)

- Google DeepMind's AlphaEvolve autonomously finds novel algorithms, improving data center scheduling by 0.7% of Google's worldwide compute resources (30:30, 30:45)

### Future Trajectory

- The AI 2027 chart shows a branch point around late 2027/early 2028, where alignment solutions could determine a slower or faster path to AGI (20:27)

### General Intelligence vs. Specialization

- The debate centers on whether current models are truly general or just highly specialized, as evidenced by the hallucination example where the model invented a story about radio signals (9:51, 16:22)

![Screenshot at 0:01: Title slide highlighting the projection that 2026 will be the 'freakiest year in AI' based on LLM release acceleration.](https://ss.rapidrecap.app/screens/FMMpUO1uAYk/00-00-01.jpg)
![Screenshot at 0:33: A leaderboard comparing recent LLMs \(Gemini 3 Pro, Grok 4, Claude 4 Opus\) across various benchmarks, showing Gemini 3 Pro in the top ranks.](https://ss.rapidrecap.app/screens/FMMpUO1uAYk/00-00-33.jpg)
![Screenshot at 1:19: A paper figure illustrating how RLHR training paths \(RLVR Model\) differ from the Base Model, suggesting more efficient sampling of reasoning paths.](https://ss.rapidrecap.app/screens/FMMpUO1uAYk/00-01-19.jpg)
![Screenshot at 2:52: A clip from a video showing an AI-generated world based on a text prompt, illustrating generative capabilities.](https://ss.rapidrecap.app/screens/FMMpUO1uAYk/00-02-52.jpg)
![Screenshot at 10:40: A slide illustrating OpenAI's Compute Flywheel: More Compute leads to Better Models/Products, which leads to More Revenue, which funds More Compute, showing compute scaling at 3x per year.](https://ss.rapidrecap.app/screens/FMMpUO1uAYk/00-10-40.jpg)
