Real World AI Evaluations
Quick Overview
Artificial Analysis released GDPval-AA, a new leaderboard and evaluation harness for comparing OpenAI's GDPval dataset of real-world knowledge work tasks, with Anthropic's Claude Opus 4.5 leading the initial findings, followed by GPT-5, Claude Sonnet 4.5, and a tie between DeepSeek V3.2 and Gemini 3 Pro, while noting that while benchmarks are flawed, GDPval-AA attempts to measure real-world utility across 44 occupations.
Key Points: Artificial Analysis introduced GDPval-AA, a new leaderboard and evaluation harness designed to compare language models on OpenAI's GDPval dataset of real-world knowledge work tasks across 44 occupations. Key findings place Claude Opus 4.5 as the leader in GDPval-AA performance, followed by GPT-5 (not GPT-5.11), Claude Sonnet 4.5, and a tie between DeepSeek V3.2 and Gemini 3 Pro for the fifth spot. The evaluation system uses human 'graders'—experienced professionals—who blindly compare model outputs without knowing the source (AI vs. human) to rank deliverables. The evaluation also incorporates an 'automated grader,' an AI system trained to predict human expert ratings, used to quickly predict output preferences and supplement human review. The speaker expresses skepticism about benchmarks generally but praises GDPval-AA for testing models on economically valuable tasks, noting that GPT-5.1 performed slightly worse than GPT-5 in this specific test. Claude Opus 4.5 achieved its top performance while using significantly fewer tokens (half the amount) compared to GPT-5.1, suggesting better efficiency. The presentation transitioned to news about ChatGPT nearing 900 million weekly active users and a report claiming DeepSeek used banned Nvidia Blackwell chips, which Nvidia is disputing.
Context: The video discusses the release of a new AI evaluation framework called GDPval-AA by Artificial Analysis, intended to measure how well Large Language Models (LLMs) perform on real-world knowledge work tasks relevant to various occupations, moving beyond traditional, often criticized, academic benchmarks. The speaker contrasts this new evaluation method with previous benchmarks and then pivots to recent news concerning OpenAI's user growth and allegations against the Chinese AI lab DeepSeek regarding the use of restricted hardware.