# Real World AI Evaluations

Source: https://www.youtube.com/watch?v=DLppD1aqn0A
Recap page: https://rapidrecap.app/video/DLppD1aqn0A
Generated: 2025-12-15T14:32:23.927+00:00

---
## Quick Overview

Artificial Analysis released GDPval-AA, a new leaderboard and evaluation harness for comparing OpenAI's GDPval dataset of real-world knowledge work tasks, with Anthropic's Claude Opus 4.5 leading the initial findings, followed by GPT-5, Claude Sonnet 4.5, and a tie between DeepSeek V3.2 and Gemini 3 Pro, while noting that while benchmarks are flawed, GDPval-AA attempts to measure real-world utility across 44 occupations.

**Key Points:**
- Artificial Analysis introduced GDPval-AA, a new leaderboard and evaluation harness designed to compare language models on OpenAI's GDPval dataset of real-world knowledge work tasks across 44 occupations.
- Key findings place Claude Opus 4.5 as the leader in GDPval-AA performance, followed by GPT-5 (not GPT-5.11), Claude Sonnet 4.5, and a tie between DeepSeek V3.2 and Gemini 3 Pro for the fifth spot.
- The evaluation system uses human 'graders'—experienced professionals—who blindly compare model outputs without knowing the source (AI vs. human) to rank deliverables.
- The evaluation also incorporates an 'automated grader,' an AI system trained to predict human expert ratings, used to quickly predict output preferences and supplement human review.
- The speaker expresses skepticism about benchmarks generally but praises GDPval-AA for testing models on economically valuable tasks, noting that GPT-5.1 performed slightly worse than GPT-5 in this specific test.
- Claude Opus 4.5 achieved its top performance while using significantly fewer tokens (half the amount) compared to GPT-5.1, suggesting better efficiency.
- The presentation transitioned to news about ChatGPT nearing 900 million weekly active users and a report claiming DeepSeek used banned Nvidia Blackwell chips, which Nvidia is disputing.

![Screenshot at 0:05: The initial tweet from Artificial Analysis announcing the GDPval-AA leaderboard, showing key findings where Claude Opus 4.5 is listed as the leader.](https://ss.rapidrecap.app/screens/DLppD1aqn0A/00-00-05.png)

**Context:** The video discusses the release of a new AI evaluation framework called GDPval-AA by Artificial Analysis, intended to measure how well Large Language Models (LLMs) perform on real-world knowledge work tasks relevant to various occupations, moving beyond traditional, often criticized, academic benchmarks. The speaker contrasts this new evaluation method with previous benchmarks and then pivots to recent news concerning OpenAI's user growth and allegations against the Chinese AI lab DeepSeek regarding the use of restricted hardware.

## Detailed Analysis

The video introduces GDPval-AA, a new leaderboard and evaluation harness created by Artificial Analysis to test LLMs on OpenAI's GDPval dataset, focusing on real-world knowledge work tasks across 44 occupations. The initial results highlight Claude Opus 4.5 as the leader, followed by GPT-5, Claude Sonnet 4.5, and a tie between DeepSeek V3.2 and Gemini 3 Pro. The grading methodology relies on experienced human professionals ('graders') who evaluate outputs blindly, alongside an 'automated grader' AI for efficiency. The speaker notes that while benchmarks can be flawed, this one targets economically valuable tasks, and points out that GPT-5.1 performed slightly worse than GPT-5, despite using twice the number of tokens, suggesting efficiency trade-offs. Following the evaluation discussion, the video covers two news items: ChatGPT approaching 900 million weekly active users, challenging Google's Gemini, and a report alleging that Chinese AI startup DeepSeek used banned Nvidia Blackwell chips for training, a claim Nvidia publicly rebutted by stating they saw no substantiation.

### GDPval-AA Evaluation Introduction

- GDPval-AA announced as a leaderboard for comparing models on OpenAI's GDPval dataset of real-world knowledge work tasks
- The evaluation uses human graders who blindly compare model outputs
- An automated grader is also used to predict human scores quickly.

### GDPval-AA Key Findings

- Claude Opus 4.5 leads, followed by GPT-5 and Claude Sonnet 4.5
- DeepSeek V3.2 and Gemini 3 Pro tied for fifth place
- GPT-5.1 performed slightly worse than GPT-5 despite using more tokens, highlighting efficiency trade-offs.

### AI News Headlines

- ChatGPT nears 900 million weekly active users, suggesting Gemini is catching up
- The Information reported DeepSeek used banned Nvidia Blackwell chips for training
- Nvidia formally rebutted the report, stating they saw no substantiation.

### Oracle Earnings Report

- Oracle stock dropped 11% after reporting disappointing cloud sales
- Cloud sales grew 34% to $7.98 billion, but infrastructure revenue growth of 68% to $4.08 billion fell short of analyst estimates
- Oracle raised CapEx forecast to $50 billion for the fiscal year.

![Screenshot at 0:05: The initial tweet from Artificial Analysis announcing the GDPval-AA leaderboard, showing key findings where Claude Opus 4.5 is listed as the leader.](https://ss.rapidrecap.app/screens/DLppD1aqn0A/00-00-05.png)
![Screenshot at 0:44: The GDPval-AA Leaderboard chart clearly displaying the relative rankings of various LLMs, with Claude Opus 4.5 at the top.](https://ss.rapidrecap.app/screens/DLppD1aqn0A/00-00-44.png)
![Screenshot at 2:48: A screenshot of an article headline from The Information stating, "ChatGPT Nears 900 Million Weekly Active Users But Gemini is Catching Up."](https://ss.rapidrecap.app/screens/DLppD1aqn0A/00-02-48.png)
![Screenshot at 3:00: The Information article headline about DeepSeek using banned Nvidia chips to build its next model, featuring a whale image amidst falling Nvidia logos.](https://ss.rapidrecap.app/screens/DLppD1aqn0A/00-03-00.png)
![Screenshot at 4:04: The CNBC article headline showing Nvidia responding to reports that DeepSeek is using banned Blackwell AI chips.](https://ss.rapidrecap.app/screens/DLppD1aqn0A/00-04-04.png)
