# 📆 ThursdAI - Aug 14 - a week with GPT-5, OSS world models, Jan-v1 interview & more AI news

Source: https://www.youtube.com/watch?v=D6x24USYdcg
Recap page: https://rapidrecap.app/video/D6x24USYdcg
Generated: 2025-08-28T10:29:34.026+00:00

---
## Quick Overview

This episode of ThursdAI on August 14th covers a week of AI news, with a deep dive into GPT-5's post-release adjustments, including OpenAI's rollback on model selection and the release of a prompting guide. The show also highlights the rapid development of open-source world models like GameCraft and Skyworks Matrix, alongside an interview with Alan Dao from Menlo Research discussing their local AI assistant Jan and its fine-tuned Quen V1 model, which excels at simple QA tasks. Other key topics include updates on Gemini's memory features, Claude's expanded context window, and new open-source vision and language models.

**Key Points:**
- OpenAI released a prompting guide for GPT-5 after users reported issues, and the company reversed its decision to hide legacy models, reintroducing a dropdown selection.
- Two open-source world models, GameCraft from Tencent and Skyworks Matrix (Game 2), were released for real-time interactive video generation.
- Menlo Research released Jan V1, a 4 billion parameter model fine-tuned from Quen, achieving 91% performance on simple QA tasks and outperforming GPT-OSS 20B on web retrieval.
- Gemini Advanced added memory and personalization features, while Claude also enhanced its web UI with memory capabilities.
- Liquid AI released LFM2 VL, a small, fast vision-language model with 440 million and 1.6 billion parameter versions, boasting improved inference speed.
- Google released Gemma 3 270M, a small model trainable in minutes, with strong instruction following and a 4-bit quantized version requiring minimal RAM.
- Claude's Sonnet 4 model now offers a 1 million token context window via API, joining Gemini's 1 million token capacity.

**Context:** The weekly AI news show "ThursdAI" for August 14th featured hosts Alex Walov, Wolf from Raven Wolf, and LDJ discussing the latest developments in artificial intelligence. The episode provided an update on GPT-5 following its initial release, explored advancements in open-source AI models, and included an interview with Alan Dao from Menlo Research about their new AI application and model. The discussion also touched upon updates from major AI labs like Google and Anthropic, and the burgeoning field of world models.

## Detailed Analysis

The episode of ThursdAI on August 14th began by recapping the week's AI news, starting with a detailed look at GPT-5. OpenAI faced community feedback regarding its model selection, leading to the reintroduction of legacy models and the release of a new prompting guide, with hosts noting a bug discovered during initial testing. The show then shifted to open-source developments, highlighting Liquid AI's LFM2 VL, a fast vision-language model, and Stepwise's prover model. Jack Morris's GPT-OSS 20 billion base model was also mentioned. A significant portion of the show was dedicated to an interview with Alan Dao from Menlo Research about their local AI assistant, Jan, and its fine-tuned Quen V1 model. Jan V1, a 4 billion parameter model, demonstrated strong performance on simple QA tasks and web retrieval, outperforming other models in its class. Dao explained the fine-tuning methodology, which involved optimizing the model's reasoning token length and using a novel reward function within an RL framework. The conversation also covered the application's ability to integrate with various search APIs for offline use. In big tech news, Gemini Advanced introduced memory and personalization features, mirroring similar updates from Claude. The show also noted Claude's Sonnet 4 model offering a 1 million token context window, matching Google's Gemini. The rapid progress in world models was a key theme, with the announcement of Google's Genie 3 last week and the subsequent release of two open-source world models this week: GameCraft from Tencent and Skyworks Matrix (Game 2). These models focus on real-time, high-dynamic video generation. The episode also touched upon other tools like Gro's video generation capabilities and a potential movie about Sam Altman.

### Interview with Jan (Menlo Research)

- Discussion with Alan Dao about Jan V1, a 4B parameter fine-tuned Quen model excelling at simple QA and web retrieval
- Explanation of the 'Lucy Bab' methodology focusing on reasoning token optimization and novel reward functions for improved performance
- Jan app features as a local AI assistant supporting various search APIs for offline use
- Jan V1's performance metrics: 91% on Simple QA, outperforming GPT-OSS 20B on web retrieval
- Discussion on fine-tuning strategies and the importance of offline AI solutions

### GPT-5 Updates

- OpenAI's response to user feedback including reintroducing legacy models and releasing a prompting guide
- Mention of a discovered bug affecting initial GPT-5 tests and prompt optimization techniques
- Discussion on GPT-5's routing mechanism and its performance issues
- User feedback on GPT-5 performance and experiences

### Open Source Models & Tools

- Liquid AI's LFM2 VL vision-language model (440M, 1.6B parameters) with fast inference
- Stepwise's prover model for formal math proofs
- Jack Morris's GPT-OSS 20B base model
- Google's Gemma 3 270M with strong instruction following and 4-bit quantization
- Gro's video generation capabilities
- Tencent's Hunan Large Vision and GLM's VLM for visual reasoning

### Big Tech LLMs & APIs

- Gemini Advanced adding memory and personalization features
- Claude enhancing web UI with memory capabilities
- Claude Sonnet 4 and Gemini offering 1 million token context windows
- Rumors about Deep Sea Guard 2 release

### World Models

- Google's Genie 3 for video generation and interactive environments
- Release of two open-source world models: GameCraft (Tencent) for real-time video generation and Skyworks Matrix (Game 2) for real-time interactive world modeling
- Discussion on the implications of advanced world models, potentially leading to simulation debates

