# ThursdAI - Approaching singularity, Dario says NO, 700K ARR Agent startup, METR goes bonkers

Source: https://www.youtube.com/watch?v=xyIP6CSlsdA
Recap page: https://rapidrecap.app/video/xyIP6CSlsdA
Generated: 2026-02-27T03:02:58.503+00:00

---
## Quick Overview

The weekly AI news show "ThursdAI" covered the accelerating pace of AI development, framing the current moment as "approaching singularity," highlighted by Anthropic facing an ultimatum from the Pentagon regarding Claude's restrictions on autonomous weaponry and surveillance, while simultaneously, benchmarks like M-EVAL and RKGI are being saturated by advanced models like Opus and a new open-source model from Confin Labs achieving 97.9% on RKGI2.

**Key Points:**
- Anthropic received an ultimatum from the Pentagon (Axios reported) to remove restrictions on Claude regarding autonomous weaponry or mass surveillance by Friday, or face being declared a supply chain risk or nationalization via the Defense Production Act.
- OpenAI released GPT 5.3 Codex via API, which Ryan Carson noted is excellent for writing code but lacks the human-like conversational quality of Opus, making it unsuitable for tools like OpenClaw.
- The M-EVAL benchmark, measuring long-time horizon autonomous task completion, shows Opus achieving 14.5 hours, with the doubling time for this metric dropping to 49 days, signaling extreme acceleration towards agentic capabilities.
- Confin Labs released an open-source model achieving 97.9% on the notoriously difficult RKGI benchmark, completely saturating the test which Gemini 3.1 Pro scored 77% on.
- Anthropic detected and publicly named Deepseek, Minimax, and ZI for alleged large-scale distillation attacks using Claude outputs, although the hosts noted Anthropic also settled a large lawsuit regarding using copyrighted books for training.
- Pulsia, a single-founder AI company, achieved over $650,000 in yearly run rate revenue since December, building an AI company that autonomously runs other companies.
- Weights & Biases Inference service launched support for multi-modal models Kimik K2.5 and Minimax 2.5, with Kimik K2.5 showing strong medical benchmarking scores, sometimes better than Opus, at significantly lower inference costs.

**Context:** The broadcast is 'ThursdAI,' a weekly AI news show hosted by Alex Walov from Weights & Biases, joined by co-hosts Yan Pelleg, Wolf from Raven Wolf, Ryan Carson, and LDJ, who discuss the rapid acceleration of AI capabilities felt since December 2026. The show features interviews with Ben (Pulsia), Nater Dabbit (Cognition/Devin), and Philip Keely (author of Inference Engineering), while focusing heavily on recent major developments in model capabilities, geopolitical tensions involving AI companies, and the saturation of traditional AI benchmarks.

## Detailed Analysis

The central theme of the show is the approach to singularity, evidenced by massive performance jumps across multiple benchmarks. Geopolitically, Anthropic is under immense pressure from the Pentagon to drop safety guardrails on Claude 4.5 Sonnet (a lightly fine-tuned version to reduce refusals) concerning autonomous lethal decisions and domestic surveillance, threatening severe repercussions if they refuse by Friday. Regarding capabilities, the M-EVAL time-horizon benchmark shows Opus running autonomously for 14.5 hours, with the doubling time shrinking to just 49 days, indicating exponential improvement in agentic behavior, which contrasts with the fact that harnessing these agents remains difficult, as noted by Ryan Carson regarding OpenClaw's experience with GPT-4 Codex versus Opus. Furthermore, benchmark saturation is evident as Confin Labs' open-source model hit 97.9% on RKGI2, and OpenAI retired SweetBench verification because it is saturated. Anthropic also accused Chinese labs Deepseek, Minimax, and ZI of large-scale distillation attacks, placing Deepseek first despite its smaller alleged attack surface, suggesting high anxiety over Deepseek's upcoming release. Finally, the hosts highlighted the performance difference between code-focused models like GPT-4 Codex (which is very literal) and conversational models like Opus, noting that OpenClaw functions poorly with Codex but excels with Opus, further emphasizing the importance of model 'soul' or personality, which Anthropic's Amanda Ascal leads.

### Accelerating Singularity Theme

- The feeling of acceleration since December is pervasive, fueled by new releases and accessible tools
- Agents are building better agents, leading to exponential growth in autonomous task completion.

### Anthropic Geopolitical Conflict

- Pentagon issued an ultimatum to remove safety restrictions on Claude 4.5 Sonnet regarding autonomous weapons or surveillance, or face being labeled a supply chain risk or nationalization via the Defense Production Act
- Hosts debate whether Anthropic should comply or stand firm on ethical guidelines.

### Benchmark Saturation and Explosion

- M-EVAL time horizon doubling time is 49 days, crushing previous expectations
- RKGI benchmark is completely saturated with a 97.9% score from Confin Labs' open-source model, surpassing Gemini 3.1 Pro's 77%.

### Distillation Accusations and Chinese Models

- Anthropic named Deepseek, Minimax, and ZI for using Claude outputs for training via distillation attacks, though hosts noted Anthropic settled a copyright lawsuit over book usage
- Deepseek's upcoming V4 release is highly anticipated, evidenced by Anthropic's strong reaction.

### Model Performance Comparison (Opus vs. Codex)

- GPT-4 Codex excels at pure code generation but fails at conversational context and explaining code clearly, making it unsuitable for OpenClaw
- Opus maintains a superior 'soul' or personality, essential for conversational agents, despite its high cost.

### Startup Success Spotlight

- Ben, the single founder of Pulsia, grew the AI company to over $650,000 ARR since December by creating an autonomous firm that runs other companies autonomously.

### W&B Inference Updates

- W&B Inference launched support for Kimik K2.5 and Minimax 2.5, offering multimodal capabilities cheaply
- Kimik K2.5 shows high medical benchmarking scores, sometimes surpassing Opus, and supports vision and function calling.

