LLMs can't reason
Quick Overview
The assertion that Large Language Models (LLMs) cannot reason is flawed because the tests used to disqualify them, such as the 4-minute mile or the ability to run Doom on a microwave, rely on physical capabilities or subjective, emotional responses that are irrelevant to logical reasoning demonstrated by LLMs.
Key Points: The speaker refutes the claim that AI cannot reason by challenging common arguments used by skeptics, such as citing physical tasks like running a 4-minute mile or running Doom on a microwave. The speaker argues that these challenges are irrelevant tests for reasoning, as they test physical capability or the ability to interact with the physical world, not logical inference. The speaker points out the bias in human responses regarding AI-generated content, referencing a study where human-written poetry was preferred over AI poetry until the source was revealed, causing ratings to drop. The speaker introduces a diagram showing that human creativity/art is characterized by high craft and low reliance on AI art, while AI output is often high in AI Art and low in craft (according to some), suggesting a false dichotomy. The speaker cites Ilya Sutskever's tweet that valuing intelligence above all else leads to a bad time, suggesting that human bias and emotion (like fear or pride) influence the debate against AI reasoning. The speaker concludes that the ability to think/reason is defined by logic (as per dictionary definitions), which LLMs can utilize, unlike subjective experiences or physical embodiment. The speaker suggests that the common critique that AI lacks 'soul,' 'depth,' or 'presence' is an emotional, rather than logical, argument against the capability of LLMs to perform complex reasoning tasks.
Context: The video addresses the ongoing debate in AI circles regarding whether Large Language Models (LLMs) truly possess reasoning capabilities, contrasting this debate with common, often flawed, analogies and evidence presented on social media. The speaker uses specific examples from Twitter discussions, including a tweet by Ilya Sutskever and a study on poetry preference, to frame the argument that skepticism about LLM reasoning is often rooted in emotional bias or irrelevant comparisons to physical tasks.