📆 ThursdAI – Jul 31, 2025 – Qwen’s Small Models Go Big, StepFun’s Multimodal Leap, GLM-4.5’s Char...

Quick Overview

This episode of Thursday Eye, July 31st, 2025, highlights a massive influx of open-source AI model releases, including new versions from Qwen, GLM, Coher, and StepFun. Qwen released multiple models, notably the Qwen 3 Coder Flash with 3 billion active parameters, offering impressive speed and local usability. GLM 4.5 from ZAI debuted as a 355 billion parameter model with strong reasoning and coding capabilities, though its benchmark reporting faced critique. Coher launched Command A Vision, a multimodal enterprise-grade model, while StepFun released Step 3, a 321 billion parameter multimodal reasoning model claiming state-of-the-art performance in vision tasks.

Key Points: Qwen released several new models, including the Qwen 3 Coder Flash (3 billion active parameters), which performs comparably to larger models and runs efficiently on local devices, achieving "almost 80 tokens per second on the M4" and a "terminal bench of 31.3". ZAI (formerly Thudm) launched GLM 4.5, a 355 billion parameter model, noted for its reasoning and coding abilities. However, its benchmark reporting was criticized for lacking transparency and selectively omitting competitors like Qwen Coder. Coher released Command A Vision, a multimodal AI model for enterprises, featuring state-of-the-art image and text reasoning, though its licensing restricts commercial use and it was noted that "Stefan is beating at least on that benchmark" (MMU at 74% vs. Coher's 65%). StepFun unveiled Step 3, a 321 billion parameter multimodal reasoning model, claiming "new Pareto Frontier" performance and achieving "74% MMU" and "64" on Math Vision, surpassing Coher's MMU score. The episode also discussed the "crush" tool, a fast, open-source alternative to Cloud Code, which gained viral attention for its performance, with one user reporting "pushing about like 39 to 40 something almost 50 million tokens every hour while using it". Anthropic's Cloud Code experienced user dissatisfaction due to perceived "nerfing" of its free tier, leading to increased interest in alternatives like Crush. The event highlighted a significant trend towards multimodal reasoning models, with both open-source and proprietary solutions rapidly advancing, though benchmarking for vision tasks remains a challenge.

Raw markdown version of this recap