Did you miss these 2 AI stories? A *Real* LLM-crafted Breakthrough + Continual Learning Blocked?
Quick Overview
The video highlights recent AI breakthroughs, specifically the C2S-Scale 27B model for biological discovery and the strong math performance of Gemini 1.5 Deep Think on the FrontierMath benchmark, while also cautioning that current LLMs suffer from amnesia and context limitations, as evidenced by GPT-4's zero score in long-term memory retrieval and GPT-5's initial struggles with context-dependent reasoning.
Key Points: Google released C2S-Scale 27B, a 27-billion parameter foundation model, in collaboration with Yale, designed to understand the language of individual cells and accelerate biological discovery, successfully generating testable hypotheses. The C2S-Scale model predicted a novel drug combination effect (similartib + low-dose interferon) that showed a roughly 50% increase in antigen presentation in lab tests, confirming its novel predictive capability. Gemini 1.5 Deep Think achieved state-of-the-art performance on the FrontierMath benchmark, scoring 62.4%, surpassing GPT-4 at 27% and GPT-5 at 38% in concrete reasoning tasks. Current LLMs exhibit significant gaps, particularly in Long-Term Memory Storage (MS), where GPT-4 scored 0% and GPT-5 scored 3% in recall tasks. The Cattell-Horn-Carroll (CHC) theory framework is used to define and evaluate Artificial General Intelligence (AGI) across ten cognitive domains, showing current models are stronger in knowledge recall than in abstract reasoning. The video mentions a Sora-generated clip of Tupac Shakur meeting Mr. Rogers, illustrating the uncanny, yet sometimes flawed, visual generation capabilities of advanced models. The speaker notes that Claude 3 Sonnet 4.5 remains cost-effective, contrasting with the higher computational cost of running models like GPT-5.
Context: The video summarizes two significant recent developments in AI research: the release of Google's C2S-Scale 27B foundation model for single-cell biology analysis, and performance updates on various LLM benchmarks, including math and long-term memory. The discussion frames these advancements within the broader context of evaluating AI progress toward Artificial General Intelligence (AGI) using frameworks like the Cattell-Horn-Carroll (CHC) theory, while noting persistent limitations like amnesia and context window constraints.