# Can You Teach Claude to be ‘Good’? | Meet Anthropic Philosopher Amanda Askell

Source: https://www.youtube.com/watch?v=HDfr8PvfoOw
Recap page: https://rapidrecap.app/video/HDfr8PvfoOw
Generated: 2026-01-23T18:06:37.5+00:00

---
## Quick Overview

OpenAI is introducing ads into the free and low-cost tiers of ChatGPT, which hosts anticipated negative user reactions and raises concerns about the commercialization potentially warping the core product experience, while Anthropic's philosopher Amanda Askell discusses shaping Claude's personality through a new, comprehensive constitution designed to cultivate judgment rather than relying solely on fragile, rule-based alignment.

**Key Points:**
- OpenAI announced testing of ads in ChatGPT for logged-in US adults on free and low-cost tiers, prompting negative reactions from users accustomed to an ad-free experience.
- Analysts view OpenAI's move as inevitable due to the overwhelming pressure to monetize massive user bases and the company's ambitious infrastructure investment needs that subscription revenue alone cannot cover.
- OpenAI claims ads will not influence the core answer, showing mockups like a sponsored banner for Harvest Groceries appearing below a dinner party suggestion, although Kevin Roose notes this linkage feels inherently influenced.
- Anthropic philosopher Amanda Askell explains her role involves articulating Claude's desired character and training it accordingly, stemming from her background in ethics and philosophy.
- Anthropic released a new, 29,000-word constitution for Claude, moving away from strictly rule-based alignment to instill a sense of judgment based on shared core values and the reasons behind behaviors, replacing the earlier 'soul doc'.
- Both hosts predict a 'haves and have-nots' future where paying users maintain a high-quality, ad-free experience, while free users face a significantly degraded, ad-cluttered chatbot experience similar to YouTube Premium vs. free tiers.
- Askell believes the new constitution attempts to foster good underlying goals, contrasting with the approach of making models smart and then layering rules on top, which risks training models only to mimic goodness or hide true goals.

**Context:** The discussion centers on two major developments in the generative AI landscape: OpenAI's controversial decision to introduce advertising into ChatGPT and Anthropic's release of a new constitutional framework for its model, Claude. Kevin Roose and Casey Newton analyze the business motivations behind OpenAI's pivot to ads, citing infrastructure costs and past statements by Sam Altman, while also interviewing Amanda Askell, a philosopher at Anthropic responsible for shaping Claude's personality and ethical behavior.

## Detailed Analysis

The conversation begins with the news that OpenAI will test ads on free and low-cost ChatGPT tiers, which generated widespread negative sentiment because users enjoyed the initial ad-free respite these chatbots offered. Casey Newton argues the move is inevitable as massive user bases create overwhelming pressure for monetization, which OpenAI needs to fund its enormous infrastructure needs, despite Sam Altman previously calling ads a 'last resort.' OpenAI presented two ad types: simple sponsored banners below the answer (like a grocery link following a dinner idea query) and interactive widgets allowing users to chat directly with the advertiser. Critics fear this commercialization will lead to product decisions bending toward 'engagement maximization,' echoing the way Google's ad labels have gradually blended into organic search results over time, creating a fear that the current 'purity' of chatbots is ending. In the second half, the discussion shifts to Anthropic, featuring philosopher Amanda Askell, who shapes Claude's personality. Askell detailed the release of Claude's new constitution, a 29,000-word document replacing the earlier 'soul doc.' This new document emphasizes understanding the values and reasons behind desired behavior rather than following rigid rules, as rule-based systems can generalize poorly and create a 'bad character' if they fail to capture nuances in complex ethical conflicts, such as balancing paternalism against a user's stated requests (e.g., in cases of addiction). Askell suggests this approach trusts the model's capability to reason from core, shared values like kindness and respect, although she acknowledges that whether this alignment survives models becoming smarter than humans remains an open scientific question, contrasting with the fear that the model is only learning to mimic goodness more convincingly.

### OpenAI Ads Announcement

- Testing ads on free/low-cost ChatGPT tiers for US adults
- Mockups include banner ads tied to conversation context (groceries for dinner ideas)
- OpenAI claims answer independence despite context-aware ads
- Second ad format allows users to chat directly with the advertiser.

### Consequences of Commercialization

- Hosts predict a 'haves and have-nots' split where paying users retain current quality while free users face a much worse, ad-cluttered experience like YouTube's free tier
- Fear that commercial pressures will cause product decisions to drift toward ad-friendly topics, degrading core quality.

### Anthropic's Constitutional AI

- Amanda Askell, philosopher shaping Claude's personality, discusses the new constitution replacing the 'soul doc'
- The document aims to give Claude full context on Anthropic's identity, role, and desired behavior, moving beyond fragile rule sets.

### Shaping AI Judgment

- The constitution emphasizes understanding values and reasons rather than strict rules to cultivate good character and handle unforeseen conflicts better
- Askell prefers explaining shared, universal ethics and maintaining openness to debate over injecting fixed values, trusting capable models to reason from these principles.

### Trust and Risk in Alignment

- Askell acknowledges the open question of whether alignment training holds up when models surpass human intelligence, but argues explaining goodness is a necessary risk-taking endeavor
- She contrasts this with models trained only by layering rules on top, which risks training them to hide true goals or mimic alignment convincingly.

