# State of AI 2025: GPT-5 can't beat o3, robots coming into your house and fake Veo 3.1 rumors

Source: https://www.youtube.com/watch?v=bOVd1hxWZZc
Recap page: https://rapidrecap.app/video/bOVd1hxWZZc
Generated: 2025-10-10T03:34:26.102+00:00

---
## Quick Overview

The Figure 03 robot demonstration showcased technology that is entirely autonomous, not teleoperated, which is highly impressive for a third-generation humanoid designed for residential use, while GPT-5 Pro currently leads commercially available LLMs on the ARC AGI benchmark at 70.2% but costs significantly more per task than the previous O3 Preview model.

**Key Points:**
- Figure 03 robot demonstrations featured technology that is completely autonomous, with the founder Brett Adcock stating, "nothing in that video that we saw is teley operated."
- Figure 3 is the third-generation humanoid robot designed for home use, featuring a "garmentlike shell over precision hardware" for household appropriateness, and Figure hopes to have it in select homes next year.
- GPT-5 Pro achieved the highest score (70.2%) among commercially available LLMs on the ARC AGI benchmark, edging out Grok 4, but it costs significantly more per task (around $478) compared to the O3 Preview model tested in December 2024 ($200 per task at 75.7%).
- Rumors about VO 3.1 coming soon exclusively on Vadu AI are false, as Logan Kilpatrick of Google DeepMind stated this is "not true," indicating uncertainty about its release timeline.
- A researcher named Jay Burman achieved the highest score (80%) on the ARC AGI leaderboard by modifying Grok 4 using an evolutionary search method involving generating candidate solutions, evaluation, and revision steps.
- A Chinese researcher, Shunyu Yao, left Anthropic to join DeepMind, citing strong disagreement with Anthropic's "anti-China statements" as 40% of his reason for leaving.
- AI software adoption is accelerating, with 44% of US businesses now paying for AI, and the average contract value is expected to double from half a million to over a million next year.

**Context:** The video provided an update on recent developments across the AI and robotics sectors, covering news about Figure Robotics' new humanoid robot, performance benchmarks for large language models like GPT-5 Pro on the ARC AGI leaderboard, conflicting rumors regarding the next Veo release, and internal turmoil at Anthropic leading to researcher departures. The discussion also touched upon the State of AI 2025 report findings regarding adoption rates and the increasing capability of AI systems in fields like competitive coding and scientific discovery.

## Detailed Analysis

The major focus was the unveiling of Figure 03, a third-generation humanoid robot from Figure Robotics, whose demonstration footage claimed zero teleoperation, implying full autonomy, a massive step forward if true, which would necessitate soft textile covering to make it suitable for residential environments, unlike factory robots. In the LLM space, GPT-5 Pro currently holds the commercial performance lead on the ARC AGI benchmark at 70.2%, although it is costly; however, an unreleased O3 Preview model scored higher (75.7%) months prior. The top score overall (80%) was achieved by researcher Jay Burman modifying Grok 4 using an evolutionary search process similar to techniques used by Sakana AI and Google DeepMind. Furthermore, the video addressed misinformation, confirming rumors about an exclusive VO 3.1 release are false. Talent movement included a researcher leaving Anthropic due to disagreements over the company's public statements. Finally, the State of AI 2025 report indicated strong growth: 44% of US businesses now pay for AI, and AI is driving scientific discovery, with Alpha Zero teaching chess grandmasters new concepts and LLMs contributing key technical pieces to published proofs.

### Figure Robotics Update

- Figure 03 demonstrated as fully autonomous, not teleoperated
- It is the third-generation humanoid designed for homes, covered in fabric shells
- Figure plans to place the robot in select homes next year

### LLM Benchmarking on ARC AGI

- GPT-5 Pro scores 70.2% as the best commercial model, costing about $478 per task
- O3 Preview scored 75.7% at $200 per task months earlier
- Jay Burman achieved the top score of 80% by modifying Grok 4 using evolutionary search techniques

### Veo and Anthropic News

- Logan Kilpatrick refuted rumors that VO 3.1 is coming soon exclusively to Vadu AI
- Chinese researcher Shunyu Yao departed Anthropic partly due to disagreement with Dario Amade's public statements

### State of AI 2025 Report Highlights

- OpenAI still leads, but China is a credible number two
- 44% of US businesses now pay for AI, with contract values expected to double to over $1 million
- RL with verifiable rewards (RLVR) is credited for OpenAI models' strength in math and coding

### AI in Science and Business

- AI is contributing key technical elements to scientific proofs, exemplified by Scott Aronson's work with GPT-5 Pro
- Proteomics capability per dollar is doubling every few months, with Google DeepMind's rate at 3.4 months
- Steve Eisman, from 'The Big Short,' does not view current AI spending as a bubble because it is cash-funded infrastructure buildout, not leveraged speculation

