Evals in Action: From Frontier Research to Production Applications

Quick Overview

OpenAI's Evals framework, exemplified by GDPval, allows developers to rigorously evaluate frontier AI models on real-world, economically valuable tasks using human expert grading, leading to the creation of better, more reliable products with improved integration, automation, and ease of use.

Key Points: GDPval is a new evaluation framework that measures model performance economically valuable real-world tasks across 9 occupations, starting with 1000+ tasks covering real estate, CAD design, retail, and more. The evaluation process uses pairwise expert grading, where human experts compare a model's output against a human-gold deliverable for a given prompt, determining a win rate. GPT-4o achieved a win rate of approximately 17% (wins and ties) on the AIME 2025 competition math benchmark without tools, compared to 42.1% without thinking. The GDPval framework is designed to be easy to use, allowing developers to define custom graders (like the 'Scoring grader' for the 'Fundamentalanalyst' agent) and run evaluations on their own data sets. The speakers encourage a development philosophy of 'Start simple' (few inputs + simple grader), 'Use human data' (real user inputs, not hypotheticals), and 'Annotate & align' (aligning graders and adding SMEs). The tool supports third-party models via OpenRouter and enterprise features like Zero Data Retention and Enterprise Key Management. The development team is working on automating prompt optimization based on trace analysis to further accelerate iteration cycles and focus developer effort on high-value tasks.

Context: The video presents Tejal Patwardhan and Henry Scott-Green from OpenAI detailing their Evals framework, specifically highlighting the GDPval (Gross Domestic Product value) evaluation system. This system moves beyond academic benchmarks like the AIME math competition to measure AI model performance on tasks relevant to economically valuable real-world work, such as financial analysis and real estate brochure design, using expert human pairwise grading to guide model improvement.

Raw markdown version of this recap