📅 ThursdAI - GPT5 watch party, GPT-oss, Genie-3, Opus 4.1, Qwen Image & more insane AI week

Quick Overview

OpenAI released its GPT OSS models (120B and 20B) with an Apache 2.0 license, marking a significant move towards open-sourcing, though the models are noted for being highly censored and optimized for specific tasks like reasoning and tool use rather than broad knowledge or creative writing. The week also saw releases from Qwen (3 4B models with 256k context) and Tencent (Hanyan 0.5B to 7B), alongside Anthropic's Cloud Opus 4.1 and Google's Genie 3 world simulator, highlighting a rapid pace of development in both open-source and proprietary AI.

Key Points: OpenAI launched GPT OSS 120B and 20B models under the permissive Apache 2.0 license, a move celebrated for its openness, though the models are described as heavily censored and lacking broad world knowledge, excelling instead in reasoning and tool use. The GPT OSS models demonstrate strong performance on benchmarks, with the 20B model performing close to the 120B and outperforming many previous open models, running efficiently even on consumer hardware like MacBooks. Qwen released new 3 4B instruct and thinking models featuring a 256k context window, praised for their usability on mobile devices and strong performance for their size, with one user noting it as the 'best local phone model that is around 2.3 gigs'. Other open-source releases include four Hanyan models from Tencent (0.5B to 7B) and Xbay04 from Xiamen University, with Xbay04 outperforming OpenAI's 03 mini in benchmarks. Anthropic updated its top model with Cloud Opus 4.1, improving its coding, reasoning, and agentic task capabilities, while Google showcased Genie 3, an advanced world simulator described as 'mindblowing'. The discussion highlighted a trend towards tool use and reasoning over deeply embedded knowledge, with models like GPT OSS being seen as optimized for specific tasks and potentially integrating with external tools or RAG systems. The release of GPT OSS sparked debate, with some finding it a disappointment for lacking broad knowledge and multilingual capabilities, while others applauded its performance on benchmarks, reasoning, and its open-source nature, especially the Apache 2.0 license.

Raw markdown version of this recap