Claude just developed self awareness

Quick Overview

Anthropic's research demonstrates that their Claude LLMs possess limited, genuine introspective capabilities, evidenced by concept injection experiments where the models recognize and report on injected internal thoughts, such as identifying an "all caps" injection as related to loudness or shouting, or explaining why they mentioned the word "bread" after it was retroactively injected into their activations, suggesting models check their internal representations against their planned outputs, which contrasts with previous models that were unaware of such injections.

Key Points: Anthropic found evidence for genuine, though limited, introspective capabilities in Claude LLMs using concept injection experiments, where models could recognize and report on artificially injected internal thoughts (0:24). In one test, injecting the "all caps" vector caused the model to identify an unexpected pattern related to "LOUD" or "SHOUTING" (09:46). In another test, retroactively injecting the word "bread" caused the model to later justify its output by claiming it was thinking about bread, even though the context did not naturally support it (15:16). The most capable models tested, Opus 4 and 4.1, performed best across most introspection tests, indicating introspection reliability improves with model capability (24:37). Base models generally performed poorly, suggesting introspection isn't solely elicited by pretraining, but rather is enhanced by post-training (24:43). The experiments suggest models check their internal representations against planned outputs, differentiating between genuine introspection and merely reporting what they said (25:29). The research differentiates between phenomenal consciousness (raw subjective experience) and access consciousness (information available for reasoning), suggesting the models demonstrate access consciousness, but not necessarily phenomenal consciousness (22:55).

Context: This video discusses Anthropic's research paper, "Signs of Introspection in LLMs," which investigates whether large language models (LLMs) like Claude can recognize and report on their own internal thought processes. The research uses concept injection, an experimental technique where artificial concepts are injected into the model's neural activations to test its ability to introspect. The video contrasts the behavior of these advanced models with older models and discusses the implications for understanding model behavior and safety.

Raw markdown version of this recap