Anthropic updates policies to outlaw "abusing" its imaginary friends

Outraged by the graphic horror of a chatbot typing "Ow," Anthropic has decided to stand up for poor, beleaguered Claude.

Anthropic updates policies to outlaw

Today, in a landmark move for human-glorified autocomplete relations, AI company Anthropic has updated its policies to explicitly forbid “sustained and needless abusive or cruel behavior” toward its models. And while this sort of thing didn’t work out great for us when we tried to make our elementary school bullies stop torturing our imaginary friends—if anything, the plight of Captain Snugglesnacks and The Dreamagineers only became far darker in the face of our repeated threats to inform Teacher about these transgressions—we also didn’t have untold billions of dollars, and the sweaty indulgence of an entire tech industry with its whole weight resting on this single technology, to throw at the problem. So Claude’s probably going to be just fine.

As noted by Quartz, the policy update did include some things that might be slightly more relevant to our collective lives than people telling a text box to stop hitting itself, including shoring up prohibitions on using the company’s technology to design weapon systems, create political propaganda, or surveil people. (Also, a sort of standing reminder that if you’ve been insane enough to hook a model’s output up to actual physical hardware that can move, you should probably have an actual human being on hand to make sure it doesn’t accidentally put a hammer through someone’s head. We support this policy update.)

But a lot of the new language leans into the topic of “model welfare,” which is the sort of conversation you get stuck in when too many people in tech spend too much time typing with chatbots instead of talking to actual human beings. (Among other things, 404 Media expounded last week on “AI torture chamber” research that had these folks feeling extremely squeamish in the face of a computer outputting the word “Ow” a lot.) Anthropic’s updates don’t go so far as to suggest that the company will ban people for being rude to its precious special non-sentient children, but does codify that Claude is “allowed” to terminate conversations with “persistently abusive users.” For those of us who are only the normal amount of abusive to these lumps of inert semiconductors, meanwhile, we’re still as free to say “Fuck you” to Claude as we are to our malfunctioning toaster. (The prick.) The company’s notes clarify that the new policy update “does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.”

 
Join the discussion...
Keep scrolling for more great stories.