Ethan Mollick: Giving Your AI a Job Interview
Quick Overview
The comparison between general standardized tests (like SAT scores) and rigorous, job-specific evaluations for AI models reveals that while general tests measure memorization, they fail to capture crucial real-world utility like complex problem-solving, judgment, and nuanced understanding, leading to a significant gap in assessing true AI capability for high-stakes tasks.
Key Points: Standardized tests like SAT scores score AI models highly on memorization but fail to capture real-world utility like complex problem-solving. Models like GPT-5 and Claude 4.5 show significant variance in performance across different real-world tasks, indicating a lack of uniform capability. The best models (like GPT-5) outperform human experts on standardized tests (e.g., 85% correct on MMLU vs. 84% for humans) but struggle with nuanced tasks like writing creative prose or judging risk. The study used a three-step process: task creation by domain experts, testing against various AIs, and evaluation using both standardized scores and subjective 'vibe checks.' The author suggests that organizations must shift testing to be more like hiring decisions, focusing on job-specific, complex scenarios rather than relying solely on general scores. A key finding is that models like GPT-5 scored significantly lower on nuanced tasks (like writing a story about an otter on a plane) compared to their high scores on objective assessments.
Context: This video discusses the limitations of current standardized testing methodologies, such as the SAT, when evaluating the real-world capabilities of large language models (LLMs) like GPT-4, GPT-5, and Claude. The speaker argues that while these models excel at rote memorization and scoring well on objective tests, these metrics often fail to reflect their true utility in complex, nuanced business or professional scenarios that require judgment, creativity, or handling subjective risk.
Detailed Analysis
The speaker argues that relying solely on standardized tests (like SAT or MMLU scores) to evaluate AI models creates a misleading picture of their true capabilities, especially for complex, high-stakes jobs. While models like GPT-5 score highly on these objective tests—sometimes even surpassing human experts (85% vs. 84% on MMLU)—they often fail when tested on tasks requiring nuance, judgment, or risk assessment, such as financial advice or creative writing. The speaker details a three-step evaluation process involving task creation by domain experts, testing multiple AIs (including GPT-5 and Claude 4.5), and evaluating results using both standardized scores and subjective 'vibe checks.' The main takeaway is that organizations must move beyond simple scores and implement rigorous, job-specific evaluations that account for the AI's inherent biases and its ability to handle complexity and nuance, rather than just pattern-matching.