State of AI 2025: GPT-5 can't beat o3, robots coming into your house and fake Veo 3.1 rumors

Quick Overview

The Figure 03 robot demonstration showcased technology that is entirely autonomous, not teleoperated, which is highly impressive for a third-generation humanoid designed for residential use, while GPT-5 Pro currently leads commercially available LLMs on the ARC AGI benchmark at 70.2% but costs significantly more per task than the previous O3 Preview model.

Key Points: Figure 03 robot demonstrations featured technology that is completely autonomous, with the founder Brett Adcock stating, "nothing in that video that we saw is teley operated." Figure 3 is the third-generation humanoid robot designed for home use, featuring a "garmentlike shell over precision hardware" for household appropriateness, and Figure hopes to have it in select homes next year. GPT-5 Pro achieved the highest score (70.2%) among commercially available LLMs on the ARC AGI benchmark, edging out Grok 4, but it costs significantly more per task (around $478) compared to the O3 Preview model tested in December 2024 ($200 per task at 75.7%). Rumors about VO 3.1 coming soon exclusively on Vadu AI are false, as Logan Kilpatrick of Google DeepMind stated this is "not true," indicating uncertainty about its release timeline. A researcher named Jay Burman achieved the highest score (80%) on the ARC AGI leaderboard by modifying Grok 4 using an evolutionary search method involving generating candidate solutions, evaluation, and revision steps. A Chinese researcher, Shunyu Yao, left Anthropic to join DeepMind, citing strong disagreement with Anthropic's "anti-China statements" as 40% of his reason for leaving. AI software adoption is accelerating, with 44% of US businesses now paying for AI, and the average contract value is expected to double from half a million to over a million next year.

Context: The video provided an update on recent developments across the AI and robotics sectors, covering news about Figure Robotics' new humanoid robot, performance benchmarks for large language models like GPT-5 Pro on the ARC AGI leaderboard, conflicting rumors regarding the next Veo release, and internal turmoil at Anthropic leading to researcher departures. The discussion also touched upon the State of AI 2025 report findings regarding adoption rates and the increasing capability of AI systems in fields like competitive coding and scientific discovery.

Raw markdown version of this recap