# 📆 ThursdAI - Sep 25 - Grok Fast, OAI/NVIDIA $100B deal, Qwen VL/Omni, Wanimate, Kling 4.5, Moondr...

Source: https://www.youtube.com/watch?v=A_QrHAJTzFU
Recap page: https://rapidrecap.app/video/A_QrHAJTzFU
Generated: 2025-09-26T14:33:32.029+00:00

---
## Quick Overview

This week's AI news roundup highlights significant releases from Alibaba's Tongi lab, including the Qwen 3 VL multimodal model and Qwen 3 Omni, alongside Meta's Code World Gen model for agentic code reasoning. OpenAI's $100 billion compute deal with Nvidia and new evaluations like GDP eval and Gemini Robotics ER showcase advancements in large-scale infrastructure and real-world AI applications. Moonream 3 also debuted as a highly accurate, small, open-source vision-language model.

**Key Points:**
- Alibaba's Tongi lab released multiple models, notably Qwen 3 VL, a vision-enabled multimodal model, and Qwen 3 Omni, a 30B parameter model with 3B active parameters capable of processing text, image, audio, and video.
- Meta released Code World Gen, a 32 billion parameter Python-based model for agentic code reasoning, designed to understand code more like a compiler.
- OpenAI is reportedly in a $100 billion compute deal with Nvidia, signifying a massive investment in AI infrastructure.
- New evaluations were introduced, including OpenAI's GDP eval, measuring AI performance on economically valuable tasks, and Google's Gemini Robotics ER for embodied reasoning in robots.
- Moonream 3, an open-source vision-language model, was previewed, featuring a 9B mixture of experts model with 2B active parameters, known for high accuracy and grounding capabilities, with a full release expected in a couple of weeks.
- Other notable releases include DeepSeek V3.1 Terminus, an update focusing on bug fixes and enhanced agentic capabilities, and XAI's Grok 4 Fast, a cost-efficient multimodal model.
- The discussion also touched on the challenges and advancements in multimodal AI, agentic behavior, and the importance of efficient, smaller models for production deployment.

**Context:** This episode of ThursdAI covers a dense week of AI releases and announcements, hosted by Alex Walov, Ryan Carson, and Yan Pelleg, with guests Vic Coropathy and Jay from Moonream. The discussion spans open-source models, large-scale infrastructure deals, and new evaluation benchmarks, reflecting the rapid pace of development in the AI field. Key organizations like Alibaba, Meta, OpenAI, Nvidia, and XAI are featured for their recent contributions.

## Detailed Analysis

This week's AI landscape was dominated by significant releases and strategic investments. Alibaba's Tongi lab continued its prolific output with the Qwen 3 VL, a vision-enabled multimodal model, and Qwen 3 Omni, a compact yet powerful 30B parameter model with only 3B active parameters, capable of processing text, image, audio, and video. Meta contributed with Code World Gen, a 32B parameter Python model aimed at improving agentic code reasoning by mimicking compiler-like understanding. On the infrastructure front, OpenAI is reportedly in a monumental $100 billion compute deal with Nvidia. New evaluation frameworks emerged, including OpenAI's GDP eval for measuring AI's economic value and Google's Gemini Robotics ER, designed for embodied reasoning in robotic systems. Moonream previewed its third iteration, Moonream 3, an open-source vision-language model emphasizing accuracy and grounding with a small parameter count (2B active parameters), with a full release anticipated soon. Updates also included DeepSeek V3.1 Terminus, focusing on bug fixes and agentic improvements, and XAI's Grok 4 Fast, a cost-efficient multimodal model. The conversation highlighted the growing importance of multimodal capabilities, agentic behaviors, and the ongoing quest for efficient, deployable AI models, contrasting them with large-scale, general-purpose models.

### Open Source Releases

- Qwen 3 VL (vision-enabled multimodal)
- Qwen 3 Omni (30B/3B active, text/image/audio/video)
- Meta Code World Gen (32B, Python-based code reasoning)
- DeepSeek V3.1 Terminus (bug fixes, agentic capabilities)
- Moonream 3 (9B MoE/2B active, vision-language, accurate, grounding)
- XAI Grok 4 Fast (cost-efficient multimodal)

### Big Tech & Infrastructure

- OpenAI/Nvidia $100B compute deal
- OpenAI's Stargate initiative data center preview
- Alibaba's Quentry Max flagship LLM and roadmap

### Evaluations & Benchmarks

- Meta/Hugging Face Gaia benchmark for agent evaluation (GPT-5 high, Kimik A2 open source lead)
- Scale AI Swebench Hard
- Among Us AI benchmark (GPT-5 mischievous)
- OpenAI GDP eval (measures AI on economically valuable tasks, Claude Opus 4.1 leads)
- Google Gemini Robotics ER (embodied reasoning for robots)

### Vision & Video

- Moonream 3 preview (small, accurate, grounding VLM)
- Alibaba Wani-animate (open-source lip sync, character swap)
- Kling 2.5 Turbo (30% cheaper, Prograde AI with sound)
- Suno V5 (AI music model)

### Multimodal & Audio

- Qwen 3 Omni (audio, text, image, video input/output)
- Qwen 3 TTS Flash (API only)
- Suno V5 (AI music, redefined audio quality)

### Interviews & Guests

- Vic Coropathy & Jay from Moonream discuss Moonream 3 capabilities, targeting agentic uses, cost-efficiency, and open-source advantages
- Discussion on the 'last mile problem' in vision AI and RL applications

### Breaking AI News

- Google Gemini Robotics ER released
- OpenAI introduces GDP eval for real-world economic tasks

