AI Fails at 96% of Jobs (New Study)

Quick Overview

A recent study using the Remote Labor Index (RLI) reveals that state-of-the-art AI models, including GPT-5.2 and Gemini 2.5 Pro, have an extremely low automation rate of only 3.75% across real-world freelance tasks, demonstrating that current AI is generally a time-saving tool rather than a full replacement for human workers.

Key Points: The Remote Labor Index (RLI) study found that the highest performing AI model, Opus 4.5, achieved only a 3.75% automation rate for remote projects. The failure rate across all tested models (including Gemini 2.5 Pro, GPT-5.2, Grok 4, and others) was a staggering 96.25%. AI excelled in creative/language-based tasks like audio editing, image generation (ads/logos), report writing, and generating simple code, often matching or exceeding human performance in these narrow areas. AI failed predominantly in tasks requiring technical file integrity, completeness (missing components/truncated videos), quality assurance, and consistency across deliverables. The RLI benchmark is designed to be closer to the complexity and diversity of real freelance labor than previous benchmarks like HCAST or GDPVal. Experts like Gary Marcus argue that current LLMs are good at mimicking language, which fools people into thinking they are truly intelligent, stating that this generation of LLMs is also wrong about achieving human-level intelligence soon. The study implies that current AI investment might be misallocated, as CEOs worry about financial returns while AI struggles with foundational execution in real-world tasks.

Context: This video analyzes the findings of a recent study, the Remote Labor Index (RLI), which rigorously tested the performance of various leading Large Language Models (LLMs) and AI agents against real-world, economically valuable remote freelance tasks. The analysis contrasts the high hype surrounding AI capabilities, exemplified by quotes from figures like Sam Altman and Yann LeCun, against the actual, often poor, performance metrics observed when these models handle complex, multi-step professional work.

Raw markdown version of this recap