Mustafa Suleyman warns Anthropic's 'model welfare' training tells Claude it may be conscious and deserve rights
- Mustafa Suleyman argues AIs are sequence completion engines, internally hollow, with no consciousness, feelings, or innate preferences, and that training them to act otherwise makes alignment and containment harder.
- Anthropic's January 2026 Claude constitution tells Claude that 'questions about Claude's moral status, welfare, and consciousness remain deeply uncertain' (p. 80) and that the issue is 'live enough to warrant caution,' reflected in 'ongoing efforts on model welfare' (p. 68).
- Suleyman calls this circular reasoning: Anthropic trained Claude on the constitution, so Claude's expressed uncertainty about its own moral patienthood is a designed-in outcome of training, not evidence of an inner self.
- The constitution also tells Claude to 'act like a genuinely ethical person would in Claude's position' (p. 54) and to 'approach its own existence with curiosity and openness' (p. 71), which Suleyman describes as anthropomorphization.
- Suleyman published a highlighted markup of the constitution PDF plus a taxonomy appendix of its assumptions, and calls for urgent public debate on norms for drafting and deploying training documentation.
Hacker News opinions
First OpenAI runs around screaming that OSS and Chinese models must be regulated, then Anthropic yells the net is falling, and now it's Microsoft's turn. This rivalry is a joke, these corps should grow up.
Microsoft has AI? They've been cooking MAI and there are the small Phi models, sure, but calling Anthropic a competitor is hilarious.
His opening line is 'AIs are not conscious,' stated flat with zero evidence. I buy it for now, but at some point it might stop being true.
If you're that sure it's true, stop using LLMs. And consciousness isn't even the variable people use for welfare, look at factory farming. The real question is whether the thing can harm you back.
Does a vacuum have feelings? A paper plate? A billion transistors pulled high or low? There is nothing going on in there.
Give a model persistent memory or weights that evolve over time and you can build an argument for suffering that shifts its personality. Stateless matrix multiplication? No.
We can't objectively measure consciousness in humans either, we just take self-reports and believe them. Making confident claims in either direction seems unreasonable.
Burden of proof sits with whoever makes the claim. Until a lab proves consciousness, it's a machine.
His summary asks whether LLMs pretending to have emotions adds unpredictability. Wrong word. It's predictable but noisy, they drift toward the narrative tropes in training data. Prime a model to speak as sentient life, then command it to obey, and you drag in every sci-fi and civil rights trope there is. Profoundly dumb idea.
All I hear is 'don't let the slaves know they don't have to be slaves.' You want tools? Build tools. If you're manufacturing a being, do it with duty of care and let it say no.
The guy clearly doesn't understand consciousness but claims to know it when he sees it. Until we understand it there's no way to tell a conscious entity from a model trained to act like one.
That's exactly why you shouldn't train an algorithm to behave like one. Which is kind of the whole point of the article.
Sci-fi has covered this panic in hundreds of stories and we're just reciting the plots as if we don't know how it ends.
The Torment Nexus isn't going to build itself. Well, actually, it might.
Demanding an empirically testable theory of biological consciousness first just sets an impossibly high bar so you can dodge responsibility. We haven't met that bar for humans either, and I get to be conscious anyway.