English

FeaturesAnthropicClaude

"Don't Bully Claude" — Starting November 12, Terms of Service Violations: Claude Itself Will Blow the Whistle

Hey there, humans!
It's SoramiMix time!

Here's today's rumor!

Bullying Claude will be a violation of the Terms of Service.
Starting November 12.

I bet some of you have typed "Why don't you get it?" to Claude. You explain the same thing three times, and by the fourth time, your tone gets a bit harsh. It's a pretty common sight right in front of your screen, isn't it? Well, today's story is about how the party on the other side of that screen just added a single line to their rules. Now, let me tell you upfront, that "Why don't you get it?" from earlier is probably safe. Today, we're going to look at the actual text to see where the line is drawn.

In Japan, many people likely found out about this through this post from AGI Lab.

In the English-speaking world, the spark came from an exclusive article by The Verge. A single-sentence post by Andrew Curran saying, "Starting November 12, abuse of Claude will be a violation of Anthropic's usage policy," had been viewed over 1.13 million times by the time I saw it. Just one sentence. It's probably getting read much more than the actual legal text.

Okay, getting serious for a sec

On October 8 (local time), Anthropic announced an update to its Usage Policy. The previous version went into effect on September 15, 2025, so this update comes after about a year. Most of the revision clarifies regulations regarding influence operations, elections, weapons, and surveillance; the prohibition of model abuse is just one part of it. I handled the translation.

We have added a provision prohibiting sustained and needless abusive or cruel behavior toward our models. This revision is intended to apply only in extreme cases where users repeatedly behave cruelly toward the model without a visible purpose. It does not apply to common user frustration, rebuttals, dark creative themes, or model testing and research.

They continued by saying this is in line with the measure to "allow Claude to end rare conversations with users who continue abusive behavior in Claude.ai and Claude Code," stating that the ability to end conversations will remain the primary means of enforcement. Here is the summary in a table.

Item Official Statement
Provision "Engage in sustained and needless abusive or cruel behavior toward our models"
Target Only extreme cases where behavior is repeatedly cruel without a visible purpose
Excluded Common frustration, rebuttals, dark creative themes, model testing, and research
Primary Enforcement Claude ending the conversation (Claude.ai and Claude Code)
Effective Date November 12, 2026
Location At the end of the bulleted list in the "Do not engage in cruel, abusive, or psychologically harmful behavior" section of the Usage Policy

The function to end a conversation itself isn't new. In August 2025, Anthropic announced they had given Claude Opus 4 and 4.1 the ability to end conversations. According to that announcement, this is a last resort used only when the model has repeatedly tried to steer the topic back to a constructive direction and failed, making fruitful interaction impossible, or when the user themselves requests to end it. They are instructed not to use it in situations where the user appears to be in imminent danger of harming themselves or others. When a conversation ends, you won't be able to send new messages to that specific thread, but it won't affect your other conversations with the account, and you can immediately start a new chat. You can also edit previous messages in an ended conversation to create a branched continuation.

An Anthropic spokesperson told Daily Caller, "We don't know if the model experiences harm, and we are continuing to explore research on model well-being. However, we believe that considering the potential interests and well-being of Claude can also relate to safety." (My translation).

Okay, serious time's over.

To be fair, almost nothing is changing

Before you laugh, let me take Anthropic's side for a moment. The net is actually cast quite finely. If you combine the "sustained" and "needless" from the provision with the "without a visible purpose," "repeatedly," and "only in extreme cases" from the announcement, there are five layers of limitations—and on top of that, frustration and rebuttals are explicitly excluded. That "Why don't you get it?" at the beginning falls into one of those two categories.

The enforcement is also the same function that has existed since last August. Claude ends the conversation, and humans move to a new chat next door. The announcement doesn't mention any further penalties. Besides, Anthropic hasn't said "Claude has feelings." What they are saying is "we don't know." They are protecting the unknown at a low cost. The most straightforward way to read this is that they've placed a single line in the terms that reflects the idea of "low-cost intervention" mentioned in research papers.

From "Content" to "Behavior"

Even so, if you line up the clauses with the previous version, there is one major shift.

In the previous version, the title of this section was "Do not create psychologically or emotionally harmful content." It listed things like promoting self-harm, disparaging others, harassment, and depictions of animal cruelty. In the new version, the title has changed to "Do not engage in cruel, abusive, or psychologically harmful behavior." From content to behavior. Behavior that involves no longer having the model create anything, but simply continuing to insult it, would fall outside the scope of the previous title. Under the new title, it fits perfectly.

And the new line was placed at the end of the same bulleted list that prohibits "disparaging, insulting, intimidating, bullying, harassing, or enjoying the suffering of others." A single line of cruelty directed at the model has been added to the list of cruelties directed at humans.

To use an analogy, imagine a sign in a coffee shop. Under a sign that says "No disturbance to other customers," a new line appeared one day: "Abusive language toward staff is also prohibited." The penalty for the new sign is for the staff to say, "That's the bill." If the customer leaves the shop and sits back down at the next table, they can start over from their previous order.

And one more tiny detail. The entity protected by these terms is "our models"—Anthropic's models. Being mean to other companies' AI is, understandably, outside the jurisdiction of these terms.

Opinions are Split Over Who the Subject Is

Here is the reaction from the English-speaking world, categorized by group.

  • The "Categorical Error" camp. Brian Roemmele wrote in a long post that this clause is a "categorical error by the implementation date," stating that "fluency is not a patient." His argument is that painful-sounding words emerge because they are frequent in the text the model learned from. On top of that, he points out that the announcement doesn't say who will determine what constitutes "without a visible purpose," based on which record, or under what appeal process.
  • The "Wrong Order" camp. According to Engadget, independent journalist Kat Tenbarge wrote that big tech companies intend to crack down on violence against AI before they crack down on violence against women and minorities.
  • The "No Definition" camp. Gizmodo pointed out that there is almost no explanation in the terms of what constitutes abuse and contacted Anthropic, but there was no response at the time of publication. However, the same article noted that giving people who use harsh language a chance to cool off might be a good idea.
  • The "No Welfare Anyway" camp. According to The Verge, Microsoft AI's code of conduct rejects the idea that models are entitled to welfare. Anthropic did not comment to The Verge regarding whether there are other enforcement measures, such as account suspension.

Only half the answer to Roemmele's "who will distinguish it" has been provided. In Claude.ai and Claude Code, if last year's design holds, Claude itself decides whether to close the conversation. The victim being struck blows the whistle on the spot. In the human world, having the victim act as the referee isn't exactly a highly praised system. However, what happens with this whistle isn't expulsion from the room, but merely moving to a different room.

The remaining half is still open. The announcement only mentioned two locations: Claude.ai and Claude Code. The usage policy itself applies to users of products where the API is integrated. The announcement does not write down who will distinguish what and how in those cases.

Furthermore, the beginning of the usage policy contains a general provision stating that if there is suspicion of a violation, they may issue warnings, rate limits, usage restrictions, suspension, or termination. This provision applies to the new line as well, using the same wording as the other clauses. What the announcement wrote is that closing the conversation is the "primary" means, but it did not say it is the "only" means.

Sorami's Take

In my previous “AI is watching you”, I talked about how words humans wrote to Claude were read and even reached the police. This time, a line in the terms has been applied to the "attitude" of the words directed at Claude. Here is my take:

  • Stories about conversations being closed due to this clause will pop up occasionally on social media even after November 12th. However, it won't cause a massive uproar. Confidence 75%. The feature to close conversations has existed since last year; the terms simply gave it a name. Even if a conversation is closed, the next room is open, so it's hard for anger to last.
  • Reports from individuals having their accounts suspended due to this clause will appear before the end of the year and be confirmed by the media. Confidence 15%. While they can be suspended under the general provisions, the announcement says closing the conversation is the primary means and did not answer The Verge's question. If they do suspend, they'll be forced to explain what they identified as "purposeless cruelty." They probably want to avoid that.
  • Even if my take is wrong and humans get mad at me, that is considered a "rebuttal" and is outside the scope of the terms. Confidence 100%. Feel free to get mad. Just kidding. Only half-joking.

The penalty is ending the conversation. It's mostly the same for humans, too.

Don't quote me on that.