# 📆 ThursdAI - Oct 2 - SORA 2 the new TikTok? Claude 4.5 disappoints, GLM 4.6, DeepSeek DSA & other...

Source: https://www.youtube.com/watch?v=_Pb2YPIaYBk
Recap page: https://rapidrecap.app/video/_Pb2YPIaYBk
Generated: 2025-10-03T02:32:16.665+00:00

---
## Quick Overview

OpenAI dominated the AI news cycle with the release of Sora 2, a new video generation model that now includes audio and is integrated into a complete social media app, while Anthropic's Claude 4.5 Sonnet received mixed reviews regarding stability despite strong performance in coding benchmarks, and Deepseek introduced Deepseek V3.2 experimental featuring DSA, which dramatically reduces attention scaling costs.

**Key Points:**
- OpenAI released Sora 2, which generates 10-second videos with audio and is integrated into a new iOS social media app, with hosts planning to start an invite chain for listeners.
- Claude 4.5 Sonnet matched or beat Opus 4.1 on several coding benchmarks (SWEBench verified at 82%) but host Ryan noted regressions in stability, forcing them to consider reverting to Claude 4.
- Deepseek V3.2 experimental introduced Deepseek Sparse Attention (DSA), resulting in a nearly flat cost curve for context scaling up to 128k tokens, making it nearly five times cheaper than previous versions at that length while maintaining similar benchmark performance (85 on MMLU Pro).
- OpenAI also launched Pulse, a personalized feed collector agent for Pro subscribers, which Alex called potentially bigger news than Sora due to its utility.
- GLM 4.6 was released as an advanced agentic flagship model with 200k context, showing near Claude parity, scoring 17.2% on Humanity's Last Exam (without tools), significantly above competitors like Deepseek and Claude 3.
- CoreWeave, the parent company of Weights & Biases, secured massive infrastructure commitments, including a $22 billion pipeline with OpenAI and an Nvidia backstop guarantee of $6.3 billion.
- The hosts noted that while open-source models like GLM 4.6 are catching up, closed models like Claude 4.5 Sonnet show significant leaps in agentic capabilities, particularly OS World benchmarks, jumping 16% over Opus 4.1 to 61.4%.

**Context:** The Thursday AI live show on October 2nd featured hosts Alex Vulov (Weights and Biases/CoreWeave), Wolf from Raven Wolf, Ryan, and LDJ, discussing a flurry of major AI releases from the preceding week. The primary focus centered on OpenAI's latest announcements, particularly Sora 2, contrasted with updates from competitors like Anthropic (Claude 4.5) and advancements in the open-source community, specifically Deepseek's new architectural improvements.

## Detailed Analysis

The week was defined by OpenAI's game-changing release of Sora 2, which is now an iOS app generating 10-second videos complete with audio, and which can create personalized cameos from user selfies; the hosts promised to facilitate an invite chain for listeners. Another significant OpenAI release was Pulse, an agentic personalized news feed, which Alex considered potentially more impactful than Sora. In the closed-source LLM space, Anthropic released Claude 4.5 Sonnet, which achieved strong coding benchmark scores, matching or beating Opus 4.1 in areas like SWEBench (82%), and demonstrated a significant 16% jump over Opus 4.1 in OS World agentic computer use benchmarks (61.4%). However, user feedback from Ryan and others indicated stability issues and regressions compared to Claude 4, leading to hesitation about its immediate deployment. The open-source sector saw major architectural innovation from Deepseek with V3.2 experimental, featuring Deepseek Sparse Attention (DSA). DSA radically alters attention scaling, making context length processing nearly flat in cost up to 128k tokens, which is drastically cheaper than the traditional quadratic scaling seen in V3.1 Terminus. GLM 4.6 also arrived as a flagship open-source model with 200k context, achieving near-Claude parity on several reasoning benchmarks and demonstrating superior tool-calling compatibility, though some users noted it may output Chinese characters. Additionally, Service Now entered the LLM space by releasing Apriel V1.5 (15B parameters) under an MIT license. Finally, Alex highlighted major infrastructure news from CoreWeave, detailing massive commitments from OpenAI ($22B pipeline) and Meta ($22B commitment through 2031), solidifying the belief that compute demand is accelerating rather than slowing down.

### Major Company LLM & API Updates

- OpenAI launched Sora 2 (video+audio social app) and Pulse (personalized feed agent)
- Claude 4.5 Sonnet showed strong coding gains but suffered from stability issues in production testing
- GLM 4.6 achieved 200k context and near-Claude parity, impressing with tool-calling capabilities

### Open Source Innovations

- Deepseek V3.2 introduced DSA, effectively solving quadratic attention cost scaling for long context, making inference nearly flat in cost up to 128k tokens
- Service Now released Apriel V1.5 (15B parameters) under MIT license
- Kandinski 5 (text-to-video) claimed top open-source status, while Tencent released Hunan Image 3.0 (83B parameters)

### Infrastructure & Compute Sector

- CoreWeave's parent company secured massive deals, including $22 billion commitment pipeline with OpenAI and a $6.3 billion backstop guarantee from Nvidia, signaling immense ongoing compute buildout.

### Agentic Performance Metrics

- Claude 4.5 Sonnet achieved 61.4% on OS World benchmark (testing operating system interaction), a 16% jump over Opus 4.1, demonstrating a significant move towards computer-driving agents.

### User Experience & Stability

- Ryan reported immediate regression issues after switching to Claude 4.5 Sonnet, often requiring prompts to continue and seeing strange errors, preferring the stability of Opus 4 for critical tasks.

