📆 ThursdAI - Aug 14 - A week with GPT5, OSS world models, VLMs in OSS, Tiny Gemma & more AI news
Quick Overview
The "ThursdAI" episode for August 14th covered significant AI news, including early reactions to GPT-5, the emergence of open-source world models like Skyworks' Matrix and Hunan's Gamecraft, and the launch of the "Jan" desktop application with its "Gen V1" model that excels at simple QA. Key developments also included Mistral's "Mistral Small 3.1," Gemini's added memory features, Claude's memory updates, and Hugging Face's "GPT-OSS" model, with discussions on the growing trend of models incorporating memory and personalization.
Key Points: OpenAI's GPT-5 saw mixed reactions after a week of testing, with some users cancelling AGI plans and others speeding them up, and OpenAI released a prompting guide and optimizer to help users, humorously suggesting users were "holding it wrong." The AI landscape experienced an "explosion in world models," with Google's "Genie 3" announced last week, followed by two open-source world models: Skyworks' "Matrix" and Hunan's "Gamecraft," both capable of real-time interactive and high-dynamic game video generation. Jan, a desktop application, launched with its "Gen V1" model, a 4 billion parameter model fine-tuned from Llama that achieved 91% performance on the "Symbol QA" benchmark, outperforming Perplexity Pro on simple QA and offering local, private AI capabilities. New open-source vision language models (VLMs) emerged, including Liquid AI's "LFM2VL" (440 million parameters), which boasts fast inference on CPUs and GPUs, and GLM's "GLM-106B," described as state-of-the-art for full-spectrum visual reasoning. Major LLM providers updated their offerings with enhanced memory and personalization features: Gemini Advanced added memory, and Claude also added past memory features to its web interface. Google released "Gemma 3 270M," a highly energy-efficient and instruction-following model that can be fine-tuned on Google Colab in minutes, requiring only 128-200 megabytes of RAM when quantized. Meta released "Dino V3," a state-of-the-art computer vision model trained with self-supervised learning, capable of image segmentation and generating heat maps, which enables semantic segmentation and object detection without human-labeled data.