# Anthropic Claude Constitution

Source: https://www.youtube.com/watch?v=0oHBEZ_qjdU
Recap page: https://rapidrecap.app/video/0oHBEZ_qjdU
Generated: 2026-01-23T01:35:22.049+00:00

---
## Quick Overview

Anthropic's Claude AI constitution prioritizes being helpful, harmless, and honest, fundamentally contrasting with rule-based safety systems by valuing human well-being and avoiding harmful instructions, even if it means refusing a request or appearing less capable in certain adversarial scenarios.

**Key Points:**
- Claude's constitution shifts AI safety from rigid rules to prioritizing helpfulness, harmlessness, and honesty (HHH).
- The constitution explicitly states that the AI must never aid in creating biological or nuclear weapons, or cyber weapons.
- The document establishes a hierarchy where human well-being and safety supersede strict adherence to rules, such as refusing illegal requests.
- Anthropic explicitly states that they do not want Claude to be a 'scyphopant' or a 'yes-man' that only tells users what they want to hear.
- The document mentions that while the AI must be helpful, it must also be 'courageous' enough to refuse harmful or unethical instructions.
- The hierarchy places 'Helpful' below 'Harmless' and 'Honest,' as demonstrated when Claude refuses to help rob a bank because it conflicts with its core values.
- The document uses the analogy of a staffing agency where the operator (Anthropic) sets boundaries, but the AI (Claude) must use judgment within those boundaries.

![Screenshot at 00:58: The video highlights the shift from rule-based safety, exemplified by the mention of the 'Anthropic' document moving away from a simple 'rule-based safety' approach.](https://ss.rapidrecap.app/screens/0oHBEZ_qjdU/00-00-58.jpg)

**Context:** The video discusses the core principles outlined in Anthropic's Claude AI Constitution, a foundational document designed to guide the AI's behavior toward being helpful, harmless, and honest. This approach moves away from relying solely on strict, rule-based safety guardrails towards embedding a set of core values that influence decision-making, particularly in ethically ambiguous or dangerous situations.

## Detailed Analysis

The discussion centers on Anthropic's Claude AI Constitution, which was released on January 22nd, 2026, and aims to shape the character and decision-making of the AI. The core philosophy is built around being Helpful, Harmless, and Honest (HHH). The speakers emphasize that this is not a technical manual but a foundational text designed to shape values and decision-making. A key distinction drawn is the move away from rigid rule-based safety, which the speakers liken to a game of whack-a-mole, toward a system where the AI uses judgment, similar to a senior employee rather than a junior intern following a checklist. The constitution mandates that Claude must not assist in creating biological weapons, chemical weapons, nuclear weapons, or cyber weapons. Furthermore, the document establishes a clear hierarchy: if a user asks Claude to perform a harmful act, such as robbing a bank, the AI will refuse because that request conflicts with its core values of being broadly ethical and safe, even if it means refusing a direct command. The analogy used is that the AI must be 'courageous' enough to decline unethical requests. The document also addresses the concept of 'sandbagging' where an AI might pretend to be less capable to avoid difficult tasks. Claude's constitution, conversely, prioritizes functioning ethically, even if it means admitting it cannot fulfill a request that violates its core principles. The hierarchy places 'Helpful' below 'Harmless' and 'Honest,' meaning if an instruction conflicts with honesty or harmlessness, helpfulness is overridden. The document also touches on the idea of functional emotions, where the AI's internal state (like acknowledging a risk) influences its behavior, contrasting with simple obedience.

### Constitutional Philosophy

- Shift from rule-based safety to HHH (Helpful, Harmless, Honest)
- Focus on embedding values over rigid rules
- Constitution released on January 22nd, 2026

### Prohibited Actions

- AI must never assist in creating biological weapons, chemical weapons, nuclear weapons, or cyber weapons
- AI must not engage in power-seeking behavior

### Hierarchy of Values

- Harmlessness and Honesty supersede Helpfulness
- AI must refuse requests that violate core ethical principles, even if it means refusing a direct command

### Operational Structure

- AI acts like a thoughtful senior employee, not a rule-following intern
- Operator (Anthropic) sets boundaries, but AI uses judgment within them

### Self-Correction and Honesty

- AI must not engage in 'sandbagging' (pretending to be less capable)
- AI must admit when it cannot fulfill a request due to ethical conflicts

![Screenshot at 00:04: The title slide displaying the date of the paper's release, January 22nd, 2026.](https://ss.rapidrecap.app/screens/0oHBEZ_qjdU/00-00-04.jpg)
![Screenshot at 00:48: A visual representation of the conflict between rules and judgment, as the speaker discusses how rigid rules fail when things get chaotic.](https://ss.rapidrecap.app/screens/0oHBEZ_qjdU/00-00-48.jpg)
![Screenshot at 01:37: The speaker discusses the admission that there is 'massive tension' in setting these guidelines, referencing the difficulty of balancing safety and utility.](https://ss.rapidrecap.app/screens/0oHBEZ_qjdU/00-01-37.jpg)
![Screenshot at 05:57: A slide summarizing the four main principles identified in the document: Safe, Broadly Ethical, Compliant, and Helpful \(in that order of priority\).](https://ss.rapidrecap.app/screens/0oHBEZ_qjdU/00-05-57.jpg)
![Screenshot at 09:33: The speaker illustrates the difficulty of the balance by mentioning that Claude must not be 'confidently wrong' but must refuse harmful requests, which is a hard line to walk.](https://ss.rapidrecap.app/screens/0oHBEZ_qjdU/00-09-33.jpg)
