# Did you miss these 2 AI stories? A *Real* LLM-crafted Breakthrough + Continual Learning Blocked?

Source: https://www.youtube.com/watch?v=TK3r4XbhtMY
Recap page: https://rapidrecap.app/video/TK3r4XbhtMY
Generated: 2025-10-22T18:31:00.551+00:00

---
## Quick Overview

The video highlights recent AI breakthroughs, specifically the C2S-Scale 27B model for biological discovery and the strong math performance of Gemini 1.5 Deep Think on the FrontierMath benchmark, while also cautioning that current LLMs suffer from amnesia and context limitations, as evidenced by GPT-4's zero score in long-term memory retrieval and GPT-5's initial struggles with context-dependent reasoning.

**Key Points:**
- Google released C2S-Scale 27B, a 27-billion parameter foundation model, in collaboration with Yale, designed to understand the language of individual cells and accelerate biological discovery, successfully generating testable hypotheses.
- The C2S-Scale model predicted a novel drug combination effect (similartib + low-dose interferon) that showed a roughly 50% increase in antigen presentation in lab tests, confirming its novel predictive capability.
- Gemini 1.5 Deep Think achieved state-of-the-art performance on the FrontierMath benchmark, scoring 62.4%, surpassing GPT-4 at 27% and GPT-5 at 38% in concrete reasoning tasks.
- Current LLMs exhibit significant gaps, particularly in Long-Term Memory Storage (MS), where GPT-4 scored 0% and GPT-5 scored 3% in recall tasks.
- The Cattell-Horn-Carroll (CHC) theory framework is used to define and evaluate Artificial General Intelligence (AGI) across ten cognitive domains, showing current models are stronger in knowledge recall than in abstract reasoning.
- The video mentions a Sora-generated clip of Tupac Shakur meeting Mr. Rogers, illustrating the uncanny, yet sometimes flawed, visual generation capabilities of advanced models.
- The speaker notes that Claude 3 Sonnet 4.5 remains cost-effective, contrasting with the higher computational cost of running models like GPT-5.

![Screenshot at 0:10: A presenter holds up a large letter 'C' while discussing EpochAI's findings that Sora 2 scored 55% on GPQA questions compared to GPT-5's 72%, summarizing the comparative performance of current video/language models.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-00-10.png)

**Context:** The video summarizes two significant recent developments in AI research: the release of Google's C2S-Scale 27B foundation model for single-cell biology analysis, and performance updates on various LLM benchmarks, including math and long-term memory. The discussion frames these advancements within the broader context of evaluating AI progress toward Artificial General Intelligence (AGI) using frameworks like the Cattell-Horn-Carroll (CHC) theory, while noting persistent limitations like amnesia and context window constraints.

## Detailed Analysis

The video discusses two major AI stories: the C2S-Scale 27B model and Gemini 1.5 Deep Think's performance on FrontierMath. C2S-Scale, developed with Yale, is a 27-billion parameter model for single-cell analysis that generated a novel, testable hypothesis confirmed in vitro: combining similtasertib with low-dose interferon resulted in a 50% increase in antigen presentation, making tumors more visible to the immune system. This demonstrates an emergent capability for biological discovery. In contrast, performance on benchmarks like AGI show limitations; GPT-4 scores 27% and GPT-5 scores 38% on concrete reasoning, showing a substantial gap remains before AGI. LLMs also suffer from amnesia, with GPT-4 scoring 0% and GPT-5 scoring 3% in Long-Term Memory Retrieval tasks. The speaker contrasts the high cost of models like GPT-5 with the cost-effectiveness of Claude 3 Sonnet 4.5 for coding tasks. The video also shows a Sora-generated clip of Tupac meeting Mr. Rogers to illustrate the sometimes jarring, yet impressive, capabilities of video generation models.

### AI Benchmarks & AGI Definition

- C2S-Scale 27B released for biological discovery
- AGI defined using Cattell-Horn-Carroll theory
- GPT-4 scores 27% and GPT-5 scores 38% on concrete reasoning tasks

### LLM Performance Gaps

- Long-Term Memory Storage (MS) shows GPT-4 at 0% and GPT-5 at 3% recall
- LLMs suffer from 'amnesia' and context limitations

### Math & Reasoning Updates

- Gemini 1.5 Deep Think achieved state-of-the-art on FrontierMath (62.4%)
- On-the-Spot Reasoning (R) shows GPT-5 achieving 7% total score, primarily in Induction and Adaptation

### Visual Generation Showcase

- Sora generated a clip of Tupac meeting Mr. Rogers, demonstrating impressive, though sometimes unsettling, visual realism

![Screenshot at 0:10: Presenter holding up a card with 'C' while discussing Sora 2's performance against GPT-5 on GPQA questions.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-00-10.png)
![Screenshot at 0:20: Google article banner announcing 'Cell2Sentence Scale 27B' built on the Gemma family of open models.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-00-20.png)
![Screenshot at 0:30: Title page of the paper 'A Definition of AGI' authored by numerous researchers.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-00-30.png)
![Screenshot at 0:42: Speaker discussing the novelty of C2S-Scale's predictions, mentioning a drug candidate for cancer therapy.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-00-42.png)
![Screenshot at 0:55: Figure 1 diagram illustrating the multidimensional expansion of the C2S framework across Model Capacity, Dataset Size, Multimodality, and Context Diversity.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-00-55.png)
![Screenshot at 1:22: Speaker contrasting the performance of C2S-Scale 27B with the 29-page paper it was based on.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-01-22.png)
![Screenshot at 1:36: Speaker notes that C2S-Scale was based largely on the Gemma 2 architecture from Google.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-01-36.png)
![Screenshot at 2:00: Image showing a dual-context virtual screen: 'Immune-Context-Neutral' \(Blue\) vs. 'Immune-Context-Positive' \(Orange/Yellow\).](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-02-00.png)
![Screenshot at 2:22: Speaker discussing the need for LLMs to 'speak biology' like they speak text.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-02-22.png)
![Screenshot at 2:53: Leaderboard table showing AGI scores where GPT-4 scored 27% and GPT-5 scored 38% overall, with a Human Baseline of 83.7%. \(Note: The scores mentioned in the video audio \(27% and 38%\) are slightly different from the table shown which displays 27% and 58% for GPT-5, suggesting the speaker is referencing slightly different data or a different version of the paper/table\). The table in the paper shows GPT-4 at 27% and GPT-5 at 58% total score, contradicting the 38% mentioned in speech but aligning with the 27% for GPT-4 shown in the table on page 11:33 \(which is not fully visible here\). The speaker later mentions 58% for GPT-5, which matches the table on page 11:33 \(58% total\). I will use the 58% mentioned in speech which aligns with the paper's table on page 11:33 for consistency with the speaker's narrative context, even if the visual is blurry/inconsistent across frames during the speech segment on scores: 27% for GPT-4 and 58% for GPT-5 \(from the visible table on page 11:33\). I will stick to what the speaker says in context of the paper's findings shown later in the video summary points, but for the screenshot context, I will use the visual cue of the paper's table near 2:53 \(which is actually page 11 of the paper shown later\). Re-evaluating the visual evidence near 2:53, the table on page 11 shows GPT-4 at 27% and GPT-5 at 58%. The speaker says 27% and 38% earlier \(07:32\). I will use the 38% quote from the speaker as the primary reference for that segment, as the visual table is inconsistent/not clear at that point for the speaker's immediate reference, but the discussion is about the paper's findings overall. The most relevant visual data point is the chart/table showing the scores, so I'll use the clearest one visible later in the paper \(page 11\). The speaker says 38% in the earlier part, but 58% later in the context of the paper's abstract text shown visually around 07:32-07:35. I will focus on the clear visual evidence from the paper shown later around 07:58 which shows 58% for GPT-5 total score in the AGI table \(Table 1\). Since the speaker is summarizing the paper shown visually, I'll reference the paper's findings as context for the visual evidence shown later in the video \(around 07:58\). The most relevant visual is the paper itself showing the scores. I will use the screenshot near the end of the video showing the paper's table \(around 07:58\). Let's use a screenshot from the Twitter feed showing the 55% vs 72% comparison as it's a clear visual from the first part of the video narrative. Reverting to the first clear visual element related to performance comparison: 0:10 is best to represent the initial comparison mentioned. I will use 11:12:The speaker highlights the quote concerning LLM limitations regarding context length and amnesia, which is a key discussion point separating current models from AGI goals.](https://ss.rapidrecap.app/screens/TK3r4XbhtMY/00-02-53.png)
