Anthropic Bans Cruelty to Claude, Still Won't Say What It Protects
Anthropic will ban sustained, purposeless cruelty toward Claude from November 12, 2026, by ending chats.
Starting November 12, 2026, Anthropic’s usage policy bars sustained, purposeless cruelty toward Claude, while exempting ordinary frustration, dark fiction, and model testing. Enforcement uses the conversation-ending capability given to Claude Opus 4 and 4.1 in August 2025, which will not end a chat if a user appears at risk of harming themselves or others. The article also covers Anthropic’s model-welfare work, including Claude Opus 4.6 assigning itself a 15–20% chance of consciousness, and criticism from Microsoft AI chief Mustafa Suleyman.
54