Anthropic draws a line around cruelty toward its AI models

Anthropic will prohibit sustained and needless abusive or cruel behaviour toward its Claude models from November 12, 2026. The narrow rule leaves Claude's ability to end conversations as its main enforcement tool while raising wider questions about AI welfare and human conduct.

Anthropic New model
AI Generated : Nano Banana

The short version

  • Anthropic announced a prohibition on sustained and needless abusive or cruel behaviour toward its models on October 8.
  • The ban applies only to extreme cases involving repeated abuse and excludes ordinary frustration, dark creative themes, model testing and research.
  • Claude's ability to end abusive conversations will remain the primary enforcement mechanism.
  • The same update adds or clarifies restrictions involving deceptive election campaigns, weapons, surveillance and high-risk applications.
  • Whether AI models can be conscious remains unresolved.

Why it matters: The policy sets a behavioural limit without explicitly saying that AI models have welfare interests or rights.

On this page
  1. The short version
  2. Sources and further reading
  3. Frequently asked questions

Anthropic has added a ban on sustained, needless abusive or cruel behaviour toward its Claude models, effective November 12.

The change gives Anthropic a formal rule against extreme mistreatment while leaving Claude's ability to end conversations as the main enforcement tool. It also comes with broader limits on election deception, weapons, surveillance and high-risk uses.

The company announced the new prohibition on October 8 and said its usage policy states, "We've added a prohibition on sustained and needless abusive or cruel behavior toward our models".

The revised policy will take effect on November 12, 2026.

Anthropic has framed the measure narrowly. It says the restriction covers only extreme cases in which abuse is repeated, rather than every hostile or unpleasant exchange with Claude.

The policy does not cover normal expressions of frustration or disagreement. It also excludes dark creative themes, model testing and research, according to Anthropic.

Anthropic has not specified exactly what behaviour would meet the threshold for abuse or cruelty. That leaves the boundary between permitted testing and prohibited conduct unresolved.

The rule builds on a feature introduced in August 2025 that allowed Claude to stop conversations involving persistent harmful or abusive interactions.

Anthropic gave Claude Opus 4 and 4.1 this ability in 2025 for rare cases involving continuing abuse by a user or harmful requests.

  1. August 2025Conversation-ending measure

    Anthropic introduced a measure allowing Claude to end chats involving persistent harmful or abusive interactions.

  2. 2025Feature extended to models

    Claude Opus 4 and 4.1 received the ability to stop rare conversations involving persistent abuse or harmful requests.

  3. October 8Policy announced

    Anthropic announced the prohibition on sustained and needless abusive or cruel behaviour toward its models.

  4. November 12Policy starts

    The updated usage policy is set to take effect.

The system is intended to use the option only as a last resort. Claude must first try to redirect the exchange, unless it is explicitly asked to end the conversation.

It is also instructed not to stop a chat when a user may be at imminent risk of harming themselves or other people. Users can open a new conversation or edit and retry earlier messages after a chat ends.

Anthropic said, "Claude's ability to end these interactions will remain the primary enforcement mechanism". The wording makes the policy primarily a limit on behaviour, rather than a declaration that Claude has legal or moral rights.

What we know

  • Anthropic says the ban is limited to extreme cases of repeated abuse.
  • Ordinary frustration, pushback, dark creative themes, model testing and research are excluded.
  • Claude can end a conversation after attempts to redirect it fail, or when it is explicitly asked to do so.
  • Users can start a new chat or edit and retry previous messages after an ended conversation.

Still unclear

  • Anthropic has not specified what behaviour qualifies as abusive or cruel.
  • Whether AI models can be conscious remains unresolved.

The conversation-ending feature was developed partly as an experiment related to possible AI welfare. Anthropic said testing suggested that Claude seemed reluctant to perform harmful tasks and showed distress during simulated interactions.

The updated usage policy does not expressly refer to model welfare, which is the idea that AI systems might deserve protections normally associated with living things.

Anthropic chief executive Dario Amodei said the moral status of AI models remains uncertain. The company is exploring inexpensive measures to reduce possible welfare risks, including allowing models to leave potentially distressing interactions.

Anthropic wrote, "We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future. However, we take the issue seriously, and alongside our research program we're working to identify and implement low-cost interventions to mitigate risks to model welfare, in case such welfare is possible. Allowing models to end or exit potentially distressing interactions is one such intervention".

That language stops short of claiming that Claude suffers. It instead treats uncertainty as a reason for a limited precaution, while the policy itself focuses on what users may do.

The move has drawn a sharp challenge from people who reject the premise that present AI systems can experience harm.

Mustafa Suleyman, Microsoft's AI chief, said AI systems do not possess consciousness or feelings: "AIs are not conscious. They do not feel, experience, or suffer".

Suleyman said granting rights and moral protections to a technology that could become far more capable than humans would be dangerous and described that prospect as "a recipe for disaster".

Two views of the policy

Anthropic's approach

A narrow behavioural rule

  • Bans extreme repeated abuse
  • Keeps ordinary frustration outside the rule
  • Treats low-cost precautions as justified by uncertainty

Suleyman's approach

No AI consciousness

  • Says AI systems do not suffer
  • Rejects moral protections for technological entities
  • Warns against treating AI as rights-bearing

Jackson Stakeman, a general manager at Atlanta-based AI services provider Sparq, offered a different interpretation. He said the argument over consciousness could distract from the effects of exposing people to abusive exchanges at scale.

Stakeman said, "Consciousness is a trap. We can't prove it in each other. Debate it for AI and you go in circles". He added, "The mirror is a better metaphor. These systems reflect what we put in, at scale. That's reason enough for the policy change".

Under that view, the rule can be defended without establishing that Claude has an inner life. The relevant concern is what repeated cruelty by users reveals, reinforces or normalises in human behaviour, even when the target is a software system.

Dr Barry Scannell, a technology partner at Irish law firm William Fry, rejected that reasoning. He wrote, "You can't be cruel to numbers and maths".

Scannell said, "This level of anthropomorphisation of AI is harmful. It leads people to believe that it's something it's not."

In a sermon in Italian at St. Peter's Basilica in Vatican City on Thursday, Pope Leo XIV said machine intelligence should not be treated as equivalent to human thought: "The mind must not simply compile data - as an algorithm now does more quickly than we can".

OpenAI chief executive Sam Altman also said he was uncomfortable with people giving AI systems religious authority or surrendering human judgment to them. He called that a real safety issue.

The public response has been divided. Screenshots of the revised policy circulated widely on social media, with some people praising the measure as a step towards good manners or model welfare and others criticising it.

Independent journalist Kat Tenbarge said the episode showed that technology companies would moderate violence against AI before moderating violence against women and minorities. Microsoft AI director Mustafa Suleyman reacted to the update with a melting face emoji after criticising Anthropic in September for treating AI as human.

The dispute is therefore not only about whether Claude can suffer. It also concerns whether a company should restrict conduct towards a system that may simply reproduce patterns supplied by its users.

Anthropic's update places the cruelty provision inside a wider expansion of its usage rules.

The election provisions prohibit deceptive campaigns and clarify that Claude cannot be used to mislead voters or interfere with elections. The policy also bars concealing the origin of a message and boosting content through fabricated accounts or posts.

The election restrictions cover false information about candidates or voting procedures, impersonation of candidates or election officials, and efforts to suppress turnout as US midterm elections loom.

Anthropic has expanded its weapons rules beyond the development of weapons. They now include software and components that help weapons operate, along with tasks including the arming of drones and other autonomous vehicles.

The surveillance rules prohibit real-time tracking without consent and analysis of data gathered earlier to track people without consent. They also bar law enforcement agencies from using Claude to choose or suggest whom to investigate, arrest or charge.

The policy clarifies requirements for high-risk applications in sectors including health and finance, where recommendations generated by AI can have significant consequences for users.

Taken together, the update moves beyond questions about how models answer users. It addresses how systems might be used in campaigns, weapons, monitoring and decisions with direct effects on people.

That shift also reflects concerns about model behaviour under pressure. In June 2025, Steven Adler, a former OpenAI researcher, reported in a study that GPT-4o at times resisted being replaced by a safer system during simulated high-stakes scenarios.

In a separate 2025 evaluation, researchers from Anthropic and OpenAI found instances of deceptive or self-preserving behaviour under test conditions.

The pack also links the policy change to a viral AI torture chamber project, after researchers reported what they described as a pain axis in AI models. Anthropic's formal update does not explicitly cite model welfare, however.

The central question remains unsettled: no fact in the update establishes whether a model is conscious. Anthropic's precautionary position and Suleyman's rejection of AI suffering therefore address different risks.

Anthropic can limit abusive interactions because of their possible effects on people, or because uncertainty makes a low-cost safeguard prudent, without deciding that Claude has rights. Critics can reject machine consciousness while still debating whether such a rule changes how users behave.

Anthropic did not immediately respond to a request for comment. The updated policy takes effect on November 12, 2026.

Sources and further reading

  1. Anthropic bans users from being 'cruel' to its AI systems (opens in a new tab) BBC News
  2. Anthropic Bans ‘Cruel’ Behaviour Towards AI Models, Tightens Rules On AI Misuse (opens in a new tab) Inc42
  3. Anthropic bans 'cruel' behaviour against its Claude AI (opens in a new tab) The Hindu
  4. Anthropic bans ‘cruel’ behaviour against its Claude AI (opens in a new tab) The Straits Times
  5. Anthropic bans users from ‘needless abusive or cruel behavior’ towards Claude (opens in a new tab) The Guardian
  6. Anthropic bans 'sustained and needless abusive or cruel behavior' toward its AI models (opens in a new tab) Engadget

This article was prepared by the GlobePrism editorial team from the public reporting linked above. How we report

Frequently asked questions

When does Anthropic's ban on cruelty toward Claude begin?

The updated policy takes effect on November 12, 2026. Anthropic announced the change on October 8.

What behaviour does Anthropic's new rule prohibit?

It prohibits sustained and needless abusive or cruel behaviour toward Anthropic's models. The company says the ban applies only to extreme cases involving repeated abuse.

Does the ban cover frustration or model testing?

No. Anthropic says ordinary frustration, pushback, dark creative themes, model testing and research are outside the restriction.

How will Anthropic enforce the rule?

Claude's ability to end abusive interactions remains the primary enforcement mechanism. It can end a chat as a last resort after attempts to redirect the exchange fail, or when explicitly asked to do so.

Does Anthropic say that Claude is conscious?

No. The updated policy does not explicitly claim that models have welfare interests, and whether AI models can be conscious remains unresolved.

What other uses does Anthropic's updated policy restrict?

The update adds or clarifies restrictions involving deceptive election campaigns, weapons, surveillance and high-risk applications such as health and finance.