📆 ThursdAI - Qwen3-Coder & new A22B, 🇺🇸 AI action plan, LLMs win IMO, Sapient HRM & more AI news
Quick Overview
The "ThursdAI" episode for July 2024 highlights significant open-source AI model releases, particularly Alibaba's Qwen 3 and Qwen 3 Coder, which demonstrate state-of-the-art performance in various benchmarks, including coding and medical Q&A. The discussion also covers the US "Win AI Race" action plan, advancements in text-to-speech and diffusion models, and the surprising success of AI in math Olympiads, with OpenAI and DeepMind achieving gold medals.
Key Points: Alibaba released Qwen 3 (235B, 22B active parameters) and Qwen 3 Coder (480B, 35B active parameters), with the latter achieving state-of-the-art results on coding benchmarks like Sweetbench Verified, outperforming previous open-source models and rivaling proprietary ones. The Qwen 3 (235B) model achieved the highest score on medical benchmarks (MedQA) among open-source models tested by the hosts, scoring 79.2 on Arena Hard and outperforming models like DeepSeek and Claude Opus in specific evaluations. Alibaba's decision to release a non-reasoning Qwen 3 model, based on community feedback, yielded impressive results, demonstrating that models without explicit chain-of-thought reasoning can still achieve top-tier performance, even on tasks requiring reasoning. The US White House released a "Win AI Race" action plan, focusing on deregulation and promoting AI development, which Joseph Nelson from RobFlow discussed in detail. AI models from OpenAI and DeepMind achieved gold medals in the International Mathematical Olympiad (IMO), showcasing advanced reasoning capabilities that surprised mathematicians and highlighted the growing prowess of AI in complex problem-solving. New advancements were noted in text-to-speech with Higs Audio V2 from Bzone AI, and in real-time diffusion with Mirage LSD from Deart AI, alongside research on subliminal learning in LLMs and Apple's multi-token prediction for faster inference. The episode also covered the release of Mistral's Magistral model on Hugging Face and the Sapien Intelligence hierarchical reasoning model, a 27-million-parameter model showing significant results on tasks like Sudoku and mazes without pre-training.