# 📆 ThursdAI - Aug 14 - A week with GPT5, OSS world models, VLMs in OSS, Tiny Gemma & more AI news

Source: https://www.youtube.com/watch?v=00Lr5Y8Z1iM
Recap page: https://rapidrecap.app/video/00Lr5Y8Z1iM
Generated: 2025-08-28T10:29:29.571+00:00

---
## Quick Overview

The "ThursdAI" episode for August 14th covered significant AI news, including early reactions to GPT-5, the emergence of open-source world models like Skyworks' Matrix and Hunan's Gamecraft, and the launch of the "Jan" desktop application with its "Gen V1" model that excels at simple QA. Key developments also included Mistral's "Mistral Small 3.1," Gemini's added memory features, Claude's memory updates, and Hugging Face's "GPT-OSS" model, with discussions on the growing trend of models incorporating memory and personalization.

**Key Points:**
- OpenAI's GPT-5 saw mixed reactions after a week of testing, with some users cancelling AGI plans and others speeding them up, and OpenAI released a prompting guide and optimizer to help users, humorously suggesting users were "holding it wrong."
- The AI landscape experienced an "explosion in world models," with Google's "Genie 3" announced last week, followed by two open-source world models: Skyworks' "Matrix" and Hunan's "Gamecraft," both capable of real-time interactive and high-dynamic game video generation.
- Jan, a desktop application, launched with its "Gen V1" model, a 4 billion parameter model fine-tuned from Llama that achieved 91% performance on the "Symbol QA" benchmark, outperforming Perplexity Pro on simple QA and offering local, private AI capabilities.
- New open-source vision language models (VLMs) emerged, including Liquid AI's "LFM2VL" (440 million parameters), which boasts fast inference on CPUs and GPUs, and GLM's "GLM-106B," described as state-of-the-art for full-spectrum visual reasoning.
- Major LLM providers updated their offerings with enhanced memory and personalization features: Gemini Advanced added memory, and Claude also added past memory features to its web interface.
- Google released "Gemma 3 270M," a highly energy-efficient and instruction-following model that can be fine-tuned on Google Colab in minutes, requiring only 128-200 megabytes of RAM when quantized.
- Meta released "Dino V3," a state-of-the-art computer vision model trained with self-supervised learning, capable of image segmentation and generating heat maps, which enables semantic segmentation and object detection without human-labeled data.

**Context:** The "ThursdAI" podcast episode on August 14th featured hosts Alex Vulov, Wolf from Raven Wolf, and LDJ, discussing the latest advancements in artificial intelligence. The discussion covered a week of AI news, following up on the previous week's significant GPT-5 announcement, and highlighted a surge in open-source developments, particularly in world models and vision language models, alongside updates from major AI labs.

## Detailed Analysis

The "ThursdAI" episode for August 14th provided a comprehensive overview of the week's AI news, starting with a follow-up on OpenAI's GPT-5, noting that after a week of testing, user reactions were mixed, with OpenAI releasing a prompting guide to assist users. The show then delved into a significant increase in open-source world models, including Skyworks' "Matrix" and Hunan's "Gamecraft," both focused on real-time video generation. In the open-source space, Liquid AI released "LFM2VL," a small vision language model, and the "Jan" application debuted its "Gen V1" model, a 4 billion parameter model that excels at simple QA and offers local, private AI capabilities, which was elaborated upon in an interview with Jan's Alan Dao. Google's "Gemma 3 270M" was highlighted for its efficiency and instruction-following, and Meta released "Dino V3," a self-supervised vision model for segmentation. Updates from major players included Mistral's "Mistral Small 3.1" and enhancements to Gemini and Claude with added memory and personalization features. The program also touched upon the "GPT-OSS" project, which reversed an instruct model back into its base form, and the LM Marina leaderboard, ranking open-source models.

### News Recap

- GPT-5 reactions and OpenAI's prompting guide
- Mistral Small 3.1 launch
- Gemini and Claude add memory features
- Google's Gemma 3 270M release
- Meta's Dino V3 computer vision model
- Hunan's Gamecraft and Skyworks' Matrix world models

### Open Source Highlights

- Liquid AI's LFM2VL vision language model
- Jan's Gen V1 model and desktop application interview
- Jack Morris's GPT-OSS base model reversal
- LM Marina leaderboard updates: Quen 3 leads, GPT-OSS at seventh

### Key Interviews

- Alan Dao from Jan discusses Gen V1 model and application, focusing on its simple QA performance and local capabilities

### Model Capabilities Discussed

- World models for real-time video generation
- Vision language models (VLMs) for image understanding
- AI agents with memory and personalization features
- Self-supervised learning for computer vision

### Technical Deep Dive

- Fine-tuning strategies for Gen V1 using Quen
- Model quantization and energy efficiency of Gemma 3 270M
- Self-supervised learning techniques in Dino V3 for image segmentation

