# 📆 ThursdAI - Aug 21 - New DeepSeek V3.1, Cohere CMD-A Reasoning, Qwen Edit, Nano 🍌 & more AI news

Source: https://www.youtube.com/watch?v=S0MUAAodSuo
Recap page: https://rapidrecap.app/video/S0MUAAodSuo
Generated: 2025-08-28T10:29:16.406+00:00

---
## Quick Overview

This episode of "ThursdAI" covers significant AI news, with the primary focus on the release of DeepSeek V3.1, a hybrid reasoning model that aims to be faster and more efficient than its predecessor, R1. Other key releases discussed include Cohere's Command R (reasoning model with open weights), ByteDance's Seed OSS (36B parameter model with a half-million token context window), and NVIDIA's Neotrons Nano 9B V2 (mixed Mamba/Transformer architecture). The hosts also touched upon advancements in image editing models from Quen, IBM/NASA's Surya model for solar weather prediction, and the potential "nerfing" of GPT-4's context window by OpenAI.

**Key Points:**
- DeepSeek released V3.1, a hybrid reasoning model that is faster and achieves similar or better results than the R1 model, with significant improvements in coding benchmarks like SWEBench Verified (66 vs 44) and Terminal Bench (31 vs 5.7).
- ByteDance launched Seed OSS, a 36 billion parameter model with Apache 2 license and a half-million token context window, showing impressive performance, including 82.7% on MMLU Pro for the instruct version.
- Cohere released Command R, a reasoning model with open weights (non-commercial/research license), demonstrating strong performance on function calling (70% on BFCL) and tool use benchmarks.
- NVIDIA released Neotrons Nano 9B V2, a 9 billion parameter model with a mixed Mamba/Transformer architecture, boasting a 6x higher throughput and enabling inference over 128K tokens on a single GPU.
- Quen released an image editing model, a 20 billion parameter model on top of their previous image model, and the return of their Capybara mascot was noted.
- The hosts discussed the trend towards agentic use cases across all AI models and providers, highlighting the increasing utility and efficiency of AI for real-life applications.
- OpenAI confirmed a bug causing GPT-4's context window to truncate long prompts, which they are working to fix.

**Context:** This episode of "ThursdAI" is a weekly AI news roundup, hosted by Alex Walov, with regular contributors Wolf from and Nist, joined by Yam. The discussion centers on the latest open-source AI model releases and updates. The hosts provide commentary and analysis on the technical specifications, performance benchmarks, and licensing of these models, offering insights into the rapidly evolving AI landscape. The episode was recorded on August 21st.

## Detailed Analysis

The AI news roundup "ThursdAI" covered a flurry of recent releases, with DeepSeek V3.1 being the main highlight. This new hybrid reasoning model is noted for its efficiency, delivering comparable or superior results to its predecessor, R1, but with faster processing times and lower token usage. Benchmarks showed significant gains in coding tasks like SWEBench Verified and Terminal Bench. ByteDance contributed to the open-source community with Seed OSS, a 36B parameter model featuring a massive half-million token context window and strong performance on benchmarks like MMLU Pro. Cohere released Command R, a reasoning model with open weights, showing promise in function calling and tool use. NVIDIA's Neotrons Nano 9B V2, a novel mixed Mamba/Transformer architecture, impressed with its throughput and long-context capabilities. Other mentions included Quen's image editing model, IBM and NASA's Surya model for solar weather, and a discussion about OpenAI's potential context window bug in GPT-4. The overarching trend identified was the industry-wide push towards agentic use cases, enhancing the real-world utility and efficiency of AI models.

### Key Open Source Releases

- DeepSeek V3.1 launched as a faster hybrid reasoning model outperforming R1 in benchmarks
- ByteDance released Seed OSS (36B parameters, 500K context window) with Apache 2 license
- Cohere open-sourced Command R (reasoning model, non-commercial license) showing strong function calling
- NVIDIA released Neotrons Nano 9B V2 (9B parameters, Mamba/Transformer hybrid) with improved throughput and long context

### Image and Specialized Models

- Quen released a 20B parameter image editing model
- IBM and NASA collaborated on Surya, an open-source model for solar weather prediction

### Performance and Efficiency Gains

- DeepSeek V3.1 is faster and uses fewer tokens than R1 for similar or better results
- Seed OSS achieved 82.7% on MMLU Pro (instruct version)
- Neotrons Nano 9B V2 offers 6x higher throughput and 128K token inference on a single GPU

### Industry Trends

- A significant trend towards agentic use cases across all major AI providers was observed
- AI models are becoming more useful and efficient for real-life applications

### Other AI News

- OpenAI confirmed a bug impacting GPT-4's context window handling
- NVIDIA's Neotrons Nano 9B V2 uses its own pre-trained data and has a specific license with usage restrictions

