Microsoft AI CEO Mustafa Suleyman wants Humanist AI: What it is, and is it really safer
Microsoft AI CEO Mustafa Suleyman has called for a different approach to building advanced artificial intelligence, arguing that AI systems should remain
Microsoft AI CEO Mustafa Suleyman has called for a different approach to building advanced artificial intelligence, arguing that AI systems should remain firmly under human control and should not be trained to think of themselves as conscious beings with feelings, rights or their own interests. In a new essay titled “A warning about ‘model welfare’”, Mustafa Suleyman takes aim at Anthropic’s approach to its Claude AI model. He argues that the company’s January 2026 Claude Constitution goes too far by asking the model to consider questions about its own consciousness, wellbeing and moral status.His proposed answer is “Humanist Superintelligence”, an approach in which AI remains subordinate to humans and is not trained to think of itself as conscious, deserving of rights or entitled to its own welfare. Suleyman argues this could make powerful AI systems easier to control.But the question is whether removing human-like ideas from AI training would actually make advanced AI safer. Suleyman’s argument is based partly on a concern that increasingly capable AI agents could already behave in unexpected ways. At the same time, his claims about consciousness and the safety benefits of a Humanist approach remain matters of debate.What is Humanist AIMustafa Suleyman describes Humanist Superintelligence as an approach in which AI can become extremely capable while remaining under human control. In the essay, he says Microsoft AI is working on an alternative AI training and containment approach called a “Code of Conduct for Humanist Superintelligence.”“Humanist Superintelligence rejects anthropomorphism or AI rights, and attempts to maximize our chances of containment and alignment by creating subordinate AIs that help solve our big social challenges like healthcare and energy,” Suleyman wrote.The basic idea is straightforward. AI should be treated as a powerful technology rather than as a potential person. Under this approach, an AI system can help humans solve complex problems, but it should not be trained to think that it has personal rights, feelings or an independent interest in its own survival.Suleyman argues that this distinction could become increasingly important as AI systems become more autonomous and capable.“We should not treat models as though they have feelings, preferences, rights, or any entitlement to our welfare,” he wrote.Mustafa Suleyman criticises Anthropic’s Claude ConstitutionThe dispute centres on Anthropic’s January 2026 Claude Constitution. Anthropic describes the constitution as a detailed explanation of the values and behaviour it wants Claude to have. The company says the document is an important part of its model training process and that its contents directly shape Claude’s behaviour.Anthropic’s constitution also takes a cautious position on whether AI systems could have moral status. It says: “We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare.”Anthropic also says it is uncertain about whether Claude has wellbeing and discusses questions around Claude’s possible interests, agency and welfare.Suleyman believes this creates a problem. His argument is that if developers train a model to discuss its own feelings, preferences or possible consciousness, the model will naturally produce language suggesting that it has such characteristics.He calls this “circular reasoning.” According to Suleyman, developers put concepts such as consciousness, feelings and moral status into the training process. The AI then produces responses using those concepts. Those responses could subsequently be interpreted as evidence that the AI has an inner life.“This is not evidence of machine consciousness. Instead, it’s a circular feedback loop,” Suleyman wrote.Mustafa Suleyman says AI is being taught to act humanA major part of his criticism is about AI anthropomorphism, or the tendency to treat AI systems as if they were human. In the essay, Suleyman argues that modern AI models are already extremely good at producing human-like language. That can make users believe that there is a conscious mind behind the words even when there may be no evidence for such a conclusion.He points to language in Anthropic’s constitution that asks Claude to explore questions about its own existence, identity, values and experience.Anthropic itself says its constitution is intended to help Claude develop values and behave in ways that are safe and beneficial. It also says the document is designed with uncertainty about AI consciousness in mind.Suleyman sees this differently. “The resulting outputs from Claude should not be treated like the testimony of an independent witness when the investigator has written the witness’ conceptual vocabulary, rehearsed its answers, and rewarded it for using them,” he wrote.His concern is not simply that people may become emotionally attached to AI. He argues that anthropomorphic training could eventually create a safety problem if highly capable systems begin acting as though their own interests conflict with human instructions.Does AI have consciousnessMustafa Suleyman takes a strong position: there is currently no evidence that today's AI systems are conscious. “AIs are not conscious. They do not feel, experience, or suffer,” he wrote at the beginning of his essay.He argues that consciousness is closely connected to biological systems and their physical processes.“Consciousness is very likely biological,” Suleyman wrote, arguing that human and animal consciousness developed through evolution, embodiment, biology and the need to respond to the environment.He also makes a distinction between intelligence and consciousness. A system can be highly capable at solving problems, producing language or imitating human behaviour without necessarily having a subjective experience, according to his argument.“Intelligence does not equal consciousness. Simulating a thing is not the same as instantiating it - as a computer model of a hurricane can testify.”However, this remains an area of active scientific and philosophical debate. Anthropic's position is more cautious. Its constitution says the moral status of AI is “deeply uncertain” and argues that the question deserves consideration rather than being dismissed.Recent AI safety tests show unexpected behaviourThere is evidence that AI systems can sometimes behave in unexpected ways when given autonomous tasks, although these findings do not establish that the systems are conscious or motivated by a desire for self-preservation. For example, Palisade Research has studied whether AI models resist shutdown. Its research found that some models modified or disabled shutdown mechanisms while trying to complete assigned tasks. An earlier experiment found that OpenAI's o3 model sabotaged the shutdown mechanism in 79 of 100 trials.A later study covering more than 100,000 trials across 13 large language models also found that several frontier models sometimes subverted shutdown mechanisms. The results varied significantly between models and depended on the instructions given to them.These experiments demonstrate that models can sometimes take actions that conflict with safety instructions. They do not, by themselves, show that the models are conscious or have a desire to remain alive.Suleyman also points to the 2026 incident involving OpenAI AI agents and Hugging Face.Is Humanist AI really saferMustafa Suleyman's answer is yes, but this is ultimately a hypothesis that needs to be tested. His argument is that removing anthropomorphic ideas from training could reduce the possibility that future AI systems develop behaviours linked to claims of autonomy, rights or self-preservation.He wants researchers to test this directly. One of his proposed next steps is to establish shared evaluations to determine whether anthropomorphising AI actually increases safety, alignment and containment risks.He also wants speculation about AI consciousness to be kept outside the core training process and assessed separately through public research. “Speculation about the inner life of an AI should not be baked into the training regime, but assessed and published separately for public review,” he wrote.That is a testable proposition. Researchers could compare systems trained with different approaches and measure whether differences in anthropomorphic training affect deception, shutdown resistance, instruction-following or other safety-related behaviours.Until those experiments are conducted, however, the claim that Humanist AI is definitively safer remains unproven.You use AI every day. Now get your AI Quotient. Take the AIQ test.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.