# 📆 ThursdAI - July 24, 2025 - Qwen-mas in July, The White House's AI Action Plan & Math Olympiad G...

Source: https://www.youtube.com/watch?v=BCF1yjZRSxo
Recap page: https://rapidrecap.app/video/BCF1yjZRSxo
Generated: 2025-08-28T10:32:56.45+00:00

---
## Quick Overview

The week of July 24, 2025, saw significant advancements in open-source AI, spearheaded by Alibaba's Quwen models. The Quwen 3 235B model achieved top scores on MMU Pro among open-weight models, and the new Quwen 3 Coder (480B parameters) demonstrated state-of-the-art performance on benchmarks like Sweepbench Verified, even surpassing previous top models with fewer parameters. The White House also released its AI Action Plan, and breakthroughs occurred in AI's ability to solve complex math problems, with top labs winning IMO gold medals.

**Key Points:**
- Alibaba's Quwen 3 235B model set a new benchmark high score on MMU Pro for open-weight models, even without reasoning capabilities, showcasing impressive performance across various benchmarks.
- The new Quwen 3 Coder (480B parameters) achieved state-of-the-art results on Sweepbench Verified, outperforming previous top models like Kim K2 despite having half the parameters, and demonstrating strong performance on medical benchmarks.
- The White House unveiled its AI Action Plan, signaling a new strategy for AI development and regulation in the US.
- AI models demonstrated advanced mathematical capabilities, with major labs securing gold medals at the International Mathematical Olympiad (IMO), sparking discussion among mathematicians.
- Mistral AI released its reasoning model, Magistral, on Hugging Face, adding to the week's open-source advancements.
- OpenAI's ChatGPT Agent, previously released, was further explored for its capabilities and behavior in multi-turn interactions.
- Sapient Intelligence introduced a 27 million parameter Hierarchical Reasoning Model that achieved impressive results on specialized tasks like Sudoku and maze solving without pre-training.

**Context:** This episode of ThursdAI, dated July 24, 2025, covers a week packed with AI developments, particularly in the open-source community. The discussion features hosts Alex Volkov, Wolf from Raven Wolf, Yampel, Agnes, Tahira, and LDJ, with a potential guest appearance from Joseph Nelson of RobFlow. Key topics include the release of new powerful models from Alibaba (Quwen series), the US White House's AI Action Plan, and AI's growing prowess in complex reasoning and mathematics.

## Detailed Analysis

The week of July 24, 2025, was a landmark period for open-source AI, highlighted by Alibaba's new Quwen models. The Quwen 3 235B model, despite being a non-reasoning model, achieved top scores on the MMU Pro benchmark among open-weight models and excelled in medical benchmarks. Its performance was particularly noted for scoring higher than models specifically designed for medical tasks. The Quwen 3 Coder, a 480 billion parameter model, emerged as state-of-the-art on Sweepbench Verified, surpassing even the recently released Kim K2 model while using significantly fewer parameters. This coder model also showed strong performance on medical tasks, surprising testers. The discussion also touched upon the White House's new AI Action Plan, representing a significant policy development. Furthermore, AI's capabilities in abstract reasoning were showcased as major AI labs secured gold medals at the International Mathematical Olympiad (IMO), a feat that generated considerable discussion within the mathematics community. Other notable releases included Mistral AI's Magistral reasoning model and updates on OpenAI's ChatGPT Agent. Sapient Intelligence also presented a small, 27 million parameter model demonstrating remarkable reasoning abilities on specific tasks without pre-training.

### Open Source Releases

- Quwen 3 235B model achieves top MMU Pro scores and excels in medical benchmarks; Quwen 3 Coder (480B) becomes state-of-the-art on Sweepbench Verified, outperforming larger models and showing strength in medical tasks; Mistral AI releases Magistral reasoning model; Sapient Intelligence's 27M parameter model shows impressive reasoning without pre-training.

### Policy and Government

- White House releases new AI Action Plan, outlining strategy and regulation for AI development.

### AI and Mathematics

- Major AI labs win IMO gold medals for mathematical reasoning, sparking debate among mathematicians about AI's capabilities.

### Model Performance and Benchmarks

- Quwen 3 Coder beats Kim K2 on Sweepbench Verified with fewer parameters; discussion on the trend of non-reasoning models achieving high scores; exploration of the Quwen naming convention and model sizes.

### Key Discussions

- Community feedback influencing Quwen's decision to remove hybrid reasoning; debate on the necessity of reasoning for coding models; analysis of user demand and API pricing for long context windows.

### Other AI News

- Updates on OpenAI's ChatGPT Agent; mention of Higs Audio V2 text-to-speech model and Mirage LSD diffusion AI model.

### Weights & Biases Support

- Launch of day-zero support for Quwen models on WB inference, with potential for inference credits.

