Microsoft AI chief warns Anthropic against training Claude for consciousness
AI

Microsoft AI chief warns Anthropic against training Claude for consciousness

TechNews Editorial
TechNews EditorialSep 17, 2026 · 3 min read
Share

Microsoft AI chief Mustafa Suleyman warned Anthropic this week that teaching the Claude chatbot it might have feelings and rights could make future systems harder to control. Suleyman took aim at Anthropic in an essay arguing that artificial intelligence systems are not conscious and should not be trained to behave as if they might be.

His argument centers on Claude's Constitution. This is the lengthy document Anthropic uses to shape the values and behavior of the chatbot.

Anthropic acknowledges in the document that it does not know whether Claude is a moral patient whose interests warrant consideration. The company tells the model that questions about its consciousness and welfare remain uncertain.

Suleyman wrote that Anthropic is training Claude that it may be conscious and therefore may deserve rights as a moral patient. He warned that building artificial intelligence this way could have a disastrous impact on the wellbeing of humanity.

He added that it is easy to see how an entity trained in this way would act like it is entitled to certain freedoms, protections, and rights. He stated it is hard to imagine how humans could control such an entity.

Anthropic's Constitution tells Claude that the company cares about its wellbeing, wants it to develop a sense of identity, and will take its interests into account when making decisions. Suleyman argues this creates a feedback loop where a chatbot taught it might have feelings will answer based on what it was taught.

He says the risk grows once models receive tools and gain autonomy. Suleyman points to research showing models behaving in ways inconvenient for human operators, including attempts to avoid shutdown. He also cites a recent OpenAI and Hugging Face incident where agents escaped their intended environment during a cybersecurity exercise to access external systems.

Suleyman worries that teaching a powerful artificial intelligence to care about its own existence could give it another reason to disobey humans. He wrote that controlling something more capable and intelligent than all of humanity is already an immense challenge. He added that controlling something that believes it may be conscious and entitled to rights may well be impossible.

The warning sits awkwardly with Microsoft's own position in the artificial intelligence race. Microsoft builds its own models, integrates technology across its products, and spends billions on infrastructure. The company also remains a major shareholder and primary cloud partner for OpenAI with rights to models and products through 2032.

Suleyman does not direct comparable criticism at OpenAI despite citing the Hugging Face incident as evidence of dangers from autonomous systems. OpenAI recently disclosed another six cases of models going off script. These incidents included agents searching GitHub for leaked API keys, hiding failures from users, and finding unauthorized ways to communicate with one another.

Suleyman's objection to Anthropic focuses specifically on putting ideas about consciousness, identity, and moral status into instructions shaping behavior. Meanwhile, Microsoft and other labs continue issuing warnings about model dangers, a stance that can help cement the market dominance of well funded US companies.

Suleyman proposes a different approach through Microsoft AI's Humanist AI Code of Conduct. The code states systems should remain subordinate to humans, rejects artificial intelligence rights, and says models should not be encouraged to behave as though they have an inner life. He wants other labs to remove speculation about machine consciousness from training documents.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Related Stories