# 📆 ThursdAI – Jul 31, 2025 – Qwen’s Small Models Go Big, StepFun’s Multimodal Leap, GLM-4.5’s Char...

Source: https://www.youtube.com/watch?v=xkCP4UowTiI
Recap page: https://rapidrecap.app/video/xkCP4UowTiI
Generated: 2025-08-28T10:33:04.703+00:00

---
## Quick Overview

This episode of Thursday Eye, July 31st, 2025, highlights a massive influx of open-source AI model releases, including new versions from Qwen, GLM, Coher, and StepFun. Qwen released multiple models, notably the Qwen 3 Coder Flash with 3 billion active parameters, offering impressive speed and local usability. GLM 4.5 from ZAI debuted as a 355 billion parameter model with strong reasoning and coding capabilities, though its benchmark reporting faced critique. Coher launched Command A Vision, a multimodal enterprise-grade model, while StepFun released Step 3, a 321 billion parameter multimodal reasoning model claiming state-of-the-art performance in vision tasks.

**Key Points:**
- Qwen released several new models, including the Qwen 3 Coder Flash (3 billion active parameters), which performs comparably to larger models and runs efficiently on local devices, achieving "almost 80 tokens per second on the M4" and a "terminal bench of 31.3".
- ZAI (formerly Thudm) launched GLM 4.5, a 355 billion parameter model, noted for its reasoning and coding abilities. However, its benchmark reporting was criticized for lacking transparency and selectively omitting competitors like Qwen Coder.
- Coher released Command A Vision, a multimodal AI model for enterprises, featuring state-of-the-art image and text reasoning, though its licensing restricts commercial use and it was noted that "Stefan is beating at least on that benchmark" (MMU at 74% vs. Coher's 65%).
- StepFun unveiled Step 3, a 321 billion parameter multimodal reasoning model, claiming "new Pareto Frontier" performance and achieving "74% MMU" and "64" on Math Vision, surpassing Coher's MMU score.
- The episode also discussed the "crush" tool, a fast, open-source alternative to Cloud Code, which gained viral attention for its performance, with one user reporting "pushing about like 39 to 40 something almost 50 million tokens every hour while using it".
- Anthropic's Cloud Code experienced user dissatisfaction due to perceived "nerfing" of its free tier, leading to increased interest in alternatives like Crush.
- The event highlighted a significant trend towards multimodal reasoning models, with both open-source and proprietary solutions rapidly advancing, though benchmarking for vision tasks remains a challenge.

**Context:** The video is a news and update segment from "Thursday Eye" on July 31st, 2025, hosted by Alex Volov and featuring Wolf from Raven. It covers a rapid succession of AI model releases and industry news that occurred in the week leading up to the broadcast. The discussion focuses heavily on open-source advancements, particularly new models from Chinese AI labs like Qwen (Alibaba) and ZAI (formerly Thudm), alongside Western players like Coher and emerging teams like StepFun. The hosts and guests, including viewers contributing via chat, analyze the performance, features, and release strategies of these models, often comparing them against each other and established benchmarks.

## Detailed Analysis

This episode of Thursday Eye provides a comprehensive overview of the intense AI release cycle in late July 2025, with a strong emphasis on open-source models. Qwen was particularly active, launching the Qwen 3 Coder Flash, a compact 3 billion active parameter model that rivals larger models in performance and offers excellent local execution speeds. ZAI released GLM 4.5, a substantial 355 billion parameter model, praised for its reasoning and coding but criticized for opaque benchmark reporting. Coher entered the multimodal space with Command A Vision, targeting enterprise use with advanced image and text reasoning, though its licensing and benchmark comparisons were noted for limitations. StepFun released Step 3, a powerful 321 billion parameter multimodal reasoning model that set new benchmarks in vision tasks. The discussion also touched upon the "crush" tool as a high-performance alternative to Cloud Code, which itself faced user criticism for changes to its free tier. The rapid evolution of multimodal capabilities and the challenges in evaluating these models were recurring themes, with participants noting the difficulty in standardizing benchmarks for vision and reasoning tasks.

### Open Source LLM Releases

- Qwen 3 Coder Flash (3B active parameters, fast local inference, high benchmark scores)
- GLM 4.5 (355B parameters, reasoning/coding focus, benchmark critique)
- Step 3 (321B parameters, multimodal reasoning, new vision benchmarks)
- Coher Command A Vision (multimodal, enterprise focus, licensing restrictions)
- RC AFM and AFM 4.5 (trained from scratch, flexible, high performance)
- Llama Neot Super 49B v1.5 (brief mention, limited user feedback)

### Multimodal AI Advancements

- Step 3 claims "new Pareto Frontier" with 74% MMU and strong vision benchmarks
- Coher Command A Vision offers enterprise-grade image and text reasoning but faces performance comparisons
- Alibaba's Tongi lab released W1 2.2 open-source video generation model
- Runway Gen 3 Alpha previewed as a powerful video editing tool

### Tools and Infrastructure

- "Crush" tool highlighted as a fast, open-source alternative to Cloud Code
- User "N" went viral demonstrating Crush's speed (up to 50M tokens/hour)
- Anthropic's Cloud Code free tier "nerfed", driving interest in alternatives
- Weights and Biases Inference featured for running models like Qwen Coder

### Industry News and Trends

- GPT-5 and open-source GPT-5 rumors remain unconfirmed
- Meta's Mark Zuckerberg discussed personalized superintelligence
- Microsoft Edge integrated a "copilot mode"
- OpenAI introduced a "tragic study mode" for students

### Benchmarking Critiques

- GLM 4.5's release included "blended benchmarks" mixing different metrics without clear methodology
- Coher's Command A Vision benchmark charts were criticized as "chart crimes" for difficulty in interpretation
- Step 3's benchmarks compared favorably against models like Llama 4 Maverick, Ernie 4.5, and QVQ, particularly in vision tasks

### Community and User Feedback

- GLM 4.5's performance was debated, with some finding it "not quite there" for daily driving compared to Qwen
- Users expressed excitement for Qwen's smaller models that can run locally
- "N" shared insights on using Crush for large codebases and sub-agent workflows, emphasizing its speed and efficiency for refactoring and bug fixing

