# AI Fails at 96% of Jobs (New Study)

Source: https://www.youtube.com/watch?v=z3kaLM8Oj4o
Recap page: https://rapidrecap.app/video/z3kaLM8Oj4o
Generated: 2026-02-13T15:34:08.703+00:00

---
## Quick Overview

A recent study using the Remote Labor Index (RLI) reveals that state-of-the-art AI models, including GPT-5.2 and Gemini 2.5 Pro, have an extremely low automation rate of only 3.75% across real-world freelance tasks, demonstrating that current AI is generally a time-saving tool rather than a full replacement for human workers.

**Key Points:**
- The Remote Labor Index (RLI) study found that the highest performing AI model, Opus 4.5, achieved only a 3.75% automation rate for remote projects.
- The failure rate across all tested models (including Gemini 2.5 Pro, GPT-5.2, Grok 4, and others) was a staggering 96.25%.
- AI excelled in creative/language-based tasks like audio editing, image generation (ads/logos), report writing, and generating simple code, often matching or exceeding human performance in these narrow areas.
- AI failed predominantly in tasks requiring technical file integrity, completeness (missing components/truncated videos), quality assurance, and consistency across deliverables.
- The RLI benchmark is designed to be closer to the complexity and diversity of real freelance labor than previous benchmarks like HCAST or GDPVal.
- Experts like Gary Marcus argue that current LLMs are good at mimicking language, which fools people into thinking they are truly intelligent, stating that this generation of LLMs is also wrong about achieving human-level intelligence soon.
- The study implies that current AI investment might be misallocated, as CEOs worry about financial returns while AI struggles with foundational execution in real-world tasks.

![Screenshot at 0:29: A bar chart displaying the Remote Labor Index \(RLI\) automation rates, where Opus 4.5 leads with 3.75% automation, while the lowest performers like Gemini 2.5 Pro score only 0.83%.](https://ss.rapidrecap.app/screens/z3kaLM8Oj4o/00-00-29.jpg)

**Context:** This video analyzes the findings of a recent study, the Remote Labor Index (RLI), which rigorously tested the performance of various leading Large Language Models (LLMs) and AI agents against real-world, economically valuable remote freelance tasks. The analysis contrasts the high hype surrounding AI capabilities, exemplified by quotes from figures like Sam Altman and Yann LeCun, against the actual, often poor, performance metrics observed when these models handle complex, multi-step professional work.

## Detailed Analysis

The video details the disappointing performance of current AI models on real-world remote work tasks, as measured by the Remote Labor Index (RLI). The key finding is that the maximum automation rate achieved by any tested model (Opus 4.5) was only 3.75%, resulting in a 96.25% failure rate across all models. The study highlights four primary failure modes: technical/file integrity issues (producing corrupt or unusable files), incomplete or malformed deliverables (missing components, truncated videos), poor quality work that doesn't meet professional standards, and inconsistencies between files. Conversely, AI showed success in predominantly creative tasks involving audio editing, image generation (like ads and logos), report writing, and generating basic code for data visualization, where performance matched or exceeded human benchmarks. Experts like Gary Marcus argue that the ability of LLMs to manipulate language creates an illusion of intelligence, noting that past predictions of human-level AI within a decade were wrong, and this generation of LLMs is similarly overhyped regarding general intelligence. The video concludes that AI is currently a time-saving tool, not a replacement, and warns that the significant investment flowing into AI might be misallocated if companies expect immediate financial returns based on flawed benchmarks that don't reflect real-world task complexity.

### AI Performance Metrics

- Opus 4.5 achieved the highest automation rate at 3.75%; Gemini 2.5 Pro scored the lowest at 0.83%
- Overall failure rate across tested models was 96.25%.

### Failure Modes

- Rejections clustered around four categories: Technical/File Integrity Issues (corrupt/empty files)
- Incomplete/Malformed Deliverables (missing assets)
- Quality Issues (not meeting professional standards)
- Inconsistencies between files.

### AI Success Areas

- AI performed comparably or better than humans in creative projects, specifically audio editing, image generation (ads/logos), report writing, and simple code generation for data visualization.

### Benchmark Critique

- The RLI is designed to be closer to the complexity/diversity of real freelance labor compared to benchmarks like HCAST or GDPVal, which focus heavily on software engineering.

### Expert Commentary (Gary Marcus)

- LLMs mimic language well, fooling people into thinking they are intelligent; previous predictions of human-level AI within 10 years were wrong, and current hype is similarly misleading.

### Future Outlook

- AI is seen as a time-saving tool, not a replacement, and there is a risk of misallocating billions of dollars if the industry ignores the current limitations shown by the RLI.

![Screenshot at 0:29: A bar chart showing the automation rates of various AI models on the RLI, with Opus 4.5 at 3.75% and Gemini 2.5 Pro at 0.83%.](https://ss.rapidrecap.app/screens/z3kaLM8Oj4o/00-00-29.jpg)
![Screenshot at 0:39: A movie clip depicting chaos, used to illustrate the potential negative outcomes or the need for human oversight.](https://ss.rapidrecap.app/screens/z3kaLM8Oj4o/00-00-39.jpg)
![Screenshot at 1:42: Diagram of the RLI Evaluation Pipeline, showing the loop of Question, AI Agent Output, Human Output, Compare, Answer \(Accept/Reject\), Brief, and Understand steps.](https://ss.rapidrecap.app/screens/z3kaLM8Oj4o/00-01-42.jpg)
![Screenshot at 2:49: A bar chart comparing the performance of various models \(Opus 4.5, GPT-5.2, Gemini 3 Pro, etc.\) on the RLI, highlighting the low automation rates.](https://ss.rapidrecap.app/screens/z3kaLM8Oj4o/00-02-49.jpg)
![Screenshot at 4:04: A comparison slide showing a 'Human Deliverable' \(detailed 2D animation\) versus an 'AI Deliverable' \(a poorly rendered document\) for an animation task, illustrating failure in quality and completeness.](https://ss.rapidrecap.app/screens/z3kaLM8Oj4o/00-04-04.jpg)
