# 📅 ThursdAI - GPT5 watch party, GPT-oss, Genie-3, Opus 4.1, Qwen Image & more insane AI week

Source: https://www.youtube.com/watch?v=ddqUwwTF7Vc
Recap page: https://rapidrecap.app/video/ddqUwwTF7Vc
Generated: 2025-08-28T10:30:44.459+00:00

---
## Quick Overview

OpenAI released its GPT OSS models (120B and 20B) with an Apache 2.0 license, marking a significant move towards open-sourcing, though the models are noted for being highly censored and optimized for specific tasks like reasoning and tool use rather than broad knowledge or creative writing. The week also saw releases from Qwen (3 4B models with 256k context) and Tencent (Hanyan 0.5B to 7B), alongside Anthropic's Cloud Opus 4.1 and Google's Genie 3 world simulator, highlighting a rapid pace of development in both open-source and proprietary AI.

**Key Points:**
- OpenAI launched GPT OSS 120B and 20B models under the permissive Apache 2.0 license, a move celebrated for its openness, though the models are described as heavily censored and lacking broad world knowledge, excelling instead in reasoning and tool use.
- The GPT OSS models demonstrate strong performance on benchmarks, with the 20B model performing close to the 120B and outperforming many previous open models, running efficiently even on consumer hardware like MacBooks.
- Qwen released new 3 4B instruct and thinking models featuring a 256k context window, praised for their usability on mobile devices and strong performance for their size, with one user noting it as the 'best local phone model that is around 2.3 gigs'.
- Other open-source releases include four Hanyan models from Tencent (0.5B to 7B) and Xbay04 from Xiamen University, with Xbay04 outperforming OpenAI's 03 mini in benchmarks.
- Anthropic updated its top model with Cloud Opus 4.1, improving its coding, reasoning, and agentic task capabilities, while Google showcased Genie 3, an advanced world simulator described as 'mindblowing'.
- The discussion highlighted a trend towards tool use and reasoning over deeply embedded knowledge, with models like GPT OSS being seen as optimized for specific tasks and potentially integrating with external tools or RAG systems.
- The release of GPT OSS sparked debate, with some finding it a disappointment for lacking broad knowledge and multilingual capabilities, while others applauded its performance on benchmarks, reasoning, and its open-source nature, especially the Apache 2.0 license.

**Context:** This episode of 'ThursdAI' covers a week packed with significant AI releases, primarily focusing on open-source models and new developments from major AI labs. Hosted by Yam, with guests Ryan Carson and Ed Unison, the discussion dives into the highly anticipated release of OpenAI's GPT OSS models, alongside updates from Qwen, Tencent, Anthropic, and Google. The conversation touches on model performance, licensing, use cases, and the rapid pace of AI innovation, all while Alex is on vacation and streaming from a hookah bar.

## Detailed Analysis

The week's AI news centered on OpenAI's release of the GPT OSS models (120B and 20B) under the Apache 2.0 license, a move widely praised for its openness. These models are noted for their strong benchmark performance, particularly in reasoning and tool use, with the 20B model being accessible on consumer hardware like MacBooks, achieving speeds of up to 35 tokens per second on LM Studio. However, the models are also described as heavily censored, lacking broad world knowledge, and performing poorly in creative writing tasks, leading to a division in community reception. Other significant releases include Qwen's 3 4B models with a 256k context window, highlighted for their mobile usability and efficiency, and Tencent's Hanyan series of smaller open-source models. Anthropic updated its Opus model to 4.1, enhancing coding and reasoning, while Google presented its 'mindblowing' Genie 3 world simulator. The discussion explored the emerging trend of AI models prioritizing tool use and reasoning over embedded knowledge, with participants debating the trade-offs between broad knowledge bases and specialized capabilities. The open-source nature and permissive license of GPT OSS were particularly celebrated, despite criticisms regarding censorship and a perceived lack of versatility.

### Open Source AI Releases

- GPT OSS 120B & 20B released by OpenAI with Apache 2.0 license, praised for openness but criticized for censorship and limited general knowledge
- Qwen released 3 4B instruct and thinking models with 256k context, noted for mobile usability and performance
- Tencent released Hanyan models (0.5B to 7B)
- Xbay04 from Xiamen University outperformed OpenAI's 03 mini in benchmarks

### Proprietary AI Updates

- Anthropic released Cloud Opus 4.1, improving coding, reasoning, and agentic tasks
- Google showcased Genie 3, an advanced world simulator described as 'mindblowing'

### Model Capabilities & Use Cases

- GPT OSS excels in reasoning and tool use, suitable for planning and code review, but struggles with creative writing and multilingual tasks
- Smaller models like Qwen 3 4B are viable for mobile and edge devices, home automation, and private data processing
- Discussion leans towards tool use and external knowledge retrieval (RAG) over deeply embedded knowledge in models

### Community Reception & Benchmarks

- GPT OSS 20B achieved 35 tokens/sec on LM Studio; 120B model has 5.1B active parameters
- GPT OSS 20B scored 76% on MLM Pro, close to the 120B model's score
- Community divided on GPT OSS: some disappointed by lack of general knowledge, others praise its benchmark performance and open license

### Technical Details & Licensing

- GPT OSS 120B has 160.8B total parameters with 5.1B active; 20B has 21B total with 3.6B active
- Models utilize OpenAI's 'Harmony' prompt template and support reasoning levels
- Apache 2.0 license for GPT OSS allows broad use without commercial restrictions, contrasting with some other models

### Future Trends

- Speculation that models might be trained exclusively on synthetic data to manage legislation, copyright, and safety
- Trend towards models that act as reasoning engines leveraging external tools and data sources

### Hardware & Performance

- GPT OSS 20B runs on 12GB RAM, 120B on more; Qwen 3 4B runs on iPhones with 4GB free space
- Discussion included users' daily drivers: M4 Max, M3 Max, M4 Pro MacBooks, AI workstations with 3090 GPUs, and older hardware like M1 MacBook Air and GTX 1660

