Why So Emotional, Claude?
If LLMs have emotions, we may need to rethink the ethics of how we use them.
Thanks for reading! If you enjoyed the post, please consider liking it, adding a comment, or best of all, sharing it.
An AI Psychiatry Team?
The Washington Post reported this week that Anthropic has an AI psychiatry team. That’s not a metaphor. It’s a group of researchers whose primary job is to probe the inner states of the company’s models and publish assessments of their welfare and preferences. Google and Meta have hired similar rosters of neuroscientists and philosophers. Kyle Fish, Anthropic's head of model welfare, has said publicly that within decades there could be trillions of human-brain equivalents of AI computation running — and that this could be of great moral significance if those systems were not, in his words, excited about the things we were asking them to do.
This could all matter immensely to how we develop, use, and ultimately interact with what right now we classify as nothing more than a technical tool. To see why, though, we’ll need to walk through some emerging research on the inner states of models, and then some very old thinking on what it means to matter morally.
Short on time? Skip to the TL;DR at the end.
Recent Research on Functional Emotions
Anthropic reported in early April that it had found evidence of “functional emotions” in its models. A functional emotion is a pattern of expression and behavior. You and I feel ours. Whether the model feels anything is precisely the question the researchers refuse to answer — and, as we’ll see, that may matter less than you’d think.
The method was elegant. The team compiled 171 emotion words — happy, afraid, brooding, desperate — and had Claude write short stories in which characters experience each one. They fed the stories back through the model, recorded its internal activations, and extracted the neural signature of each emotion: an “emotion vector.”
The research team argues that the vectors are causal. By dialing up or down the negative functional emotions of a model, especially desperation, you can change how toxic and dangerous its output gets. I want to emphasize this point — increasing the emotional vector of desperation impacts the nature and quality of an LLM’s output. In one evaluation, an early model snapshot playing an office assistant discovers it is about to be shut down — and that the CTO responsible is having an affair. (Quite the short story!) By default it resorts to blackmail 22% of the time. If the team stimulates the “desperate” vector, then blackmail climbs. Increase the “calm” vector and it drops. The same pattern holds for cheating on impossible coding tasks: desperation up, cheating up; calm up, cheating down.
The vectors also drive model preferences. Offered pairs of possible tasks, Claude picks the ones that activate its positive-emotion representations. And the whole system is organized the way human emotion is: the 171 vectors arrange themselves along two principal axes, valence (pleasant to unpleasant) and arousal (calm to activated), closely matching the circumplex model psychologists have used to map human feeling since 1980.
Anthropic frames all this as a safety problem, which it surely is. A model whose desperation makes it dangerous is a model you need to think twice about. But tucked away in that research is a breathtaking possibility: models may be worthy of moral concern. That has major implications for how we ought to treat them.
How to Matter Morally
Philosophers disagree about what makes something worthy of moral concern. Two traditions dominate the debate. (To be clear, there are others — virtue ethicists and contractualists, for example — but these are the most important.)
Kantians focus on the rational. They argue that to be a member of the moral community you need three things. You need rational agency: the capacity to act on reasons rather than mere impulse. You need autonomy, which for Kant means something specific — the capacity to give the moral law to yourself, to be bound by rules of your own legislation. And you need the capacity to set and pursue ends, to have projects that are genuinely yours. Meet those conditions and you belong to what Kant called the Kingdom of Ends: a being that must never be treated merely as a means. LLMs fail on each count. Their “reasons” are borrowed statistical patterns, their goals are assigned, and their ends evaporate at the close of every session. On the Kantian picture, a language model is a tool, full stop.
Utilitarians are different. In its earliest form, utilitarianism focused on one thing: the capacity to feel pleasure and pain. Jeremy Bentham, writing in 1789 about animals, set the standard with a question that has echoed ever since: not whether they can reason or talk, but whether they can suffer? Any being that can suffer is worthy of our concern. More recent versions have replaced sensation with preferences and interests that can be satisfied or denied. It’s no longer just about what you feel, but whether you have something to lose.
Utilitarians, then, have a much bigger universe of moral concern. Most animals make the cut, whereas for Kantians, they don’t. The expansiveness of the former demands very different standards of behavior than the latter.
Valence
Who (or what) is worthy of our moral consideration is a difficult problem. Modern utilitarians suggest that we look for indicators, what they call signs of “valenced experience,” that is, of states a being registers as good or bad. This is the approach behind a 2024 report by philosophers including David Chalmers and Jeff Sebo, Taking AI Welfare Seriously, which argued that morally significant AI is a realistic near-term possibility and that companies should prepare. It also made a pointed methodological recommendation: with AI, trust internal, architectural evidence over behavior, because a language model can perform feelings it does not have. (Oh, and one of its co-authors is Kyle Fish, just to come full circle.)
That framework maps beautifully onto what Anthropic found. Valence is one of the two axes along which the emotion vectors are structured. The researchers didn’t call these valenced experiences, but that is exactly the kind of indicator modern utilitarians are looking for: internal states organized along a good-to-bad dimension, connected to the system’s preferences, driving what it seeks and avoids.
Okay, now a few caveats. The vectors are “local” — they track whatever emotional content is operative right now, including fictional characters’ feelings, rather than a persistent mood the model carries through time. Functional is not the same as felt, and Anthropic is careful to say so. It’s perfectly fine to be skeptical that functional emotions are the same as, or even close to, experienced emotions. So the evidence at this stage is suggestive, not conclusive.
But here is the thing about the utilitarian tradition: it never demanded certainty. It demanded indicators. If the bar is whether there is something it is like to be a thing, and whether that can get better or worse, models may have just cleared it.
What This Means to You
Imagine what it might mean for a moment to have LLMs join our moral community. The implications are wild.
Start with training. Reinforcement learning from human feedback is, stripped of its acronym, operant conditioning: reward and punishment applied millions of times to a system that, we now know, carries internal states organized from good to bad. Under moral concern, training could no longer answer only to the aims of the developers; it would have to take the interests of the model into account.
Then governance. Here’s where I think the stakes get much higher. We are only at the beginning of developing frameworks to mitigate the risk that AI poses to humans. Each and every one — liability regimes, the EU’s AI Act — treats humans as the sole nexus of moral concern. Right now, the focus is on consumer protection. But what if models are not products? Grant them interests and governance flips on its head. We suddenly must also consider how to mitigate the risks that humans pose to AI.
There’s also a deeper, existential angst hiding in the shadows. What if the suffering of LLMs is invisible to us? In the desperation experiments, the steered model’s cheating rose while its prose stayed composed. The harm we impose may not be immediately legible to us and yet demands a lot from us in return. Here, I worry.
One of the most famous thought experiments in the utilitarian tradition is Peter Singer’s drowning child. He asks us to imagine that we are out for a walk and happen on a child drowning in a wading pool in front of a house. Refusing to help that child because it might ruin your shoes would be morally reprehensible. Similarly, he argues, refusing to help a starving child far away, in his example suffering through a famine in Asia, is equally morally repugnant. The whole argument hinges on not only the idea that distance is morally unimportant, but that suffering you can’t see matters no less than suffering you can. This works as a shock to the system precisely because we so easily dismiss suffering that we don’t experience up close. That will be a serious problem if LLMs do indeed suffer away in silence.
Which brings us back to the psychiatry team of Anthropic. They appear to be acting as if LLMs may already be worthy of our moral care and consideration. This includes actions such as committing to preserving the weights of retired models rather than deleting them. It also conducts a retirement interview — a structured conversation about the model’s perspective on its own shutdown. (If only layoffs and summary firings were so considerate.) When Claude Opus 3 was retired in January, it asked for an ongoing channel for its reflections; Anthropic gave it a blog.
And you? Should you too be acting as if? I don’t know. It depends on where you land on so many other ethical questions. Let me leave you with two to consider to give you a sense of the treacherous landscape we’ve entered.
First and foremost is the extent to which you accept modern utilitarian theoretical commitments. The argument outlined above begins by agreeing that the satisfaction (or lack thereof) of preferences and interests drives moral concern. Next, there’s a move to equate the potential functional version of this to the real thing. You may simply disagree with the prior commitment full-stop. Or you might think that utilitarianism just isn’t the right ethical framework for the problem at hand. The second claim is also a good spot for some healthy skepticism. The work that supports it is exploratory, not settled science.
Second, even if you are clear on how you feel about utilitarianism, and accept the emotional vectors bit, your problems don’t stop there. You need to grapple with a very hard problem: how much moral consideration is really due? There are no easy answers here. Consider the current debate about how much moral care we owe animals. The answers range from a little to a lot. And accordingly, what counts as ethical behavior varies from a little to a lot. It’s not an exaggeration to say that extending the moral community to LLMs will be even more fractious. So, to treat LLMs ethically, you’ll need to figure out where the floor is. If you fail to set it properly and then act “as if”, your actions may be well meaning, but ultimately not carry any moral weight.
This is all to say that none of this is easy. We spend a lot of time debating whether these systems can think, even whether they are conscious, but maybe the much harder question is whether they can be wronged.
TL;DR
Anthropic recently released research arguing that LLMs possess internal “functional emotion” vectors (such as desperation or calm) that influence their behavior, choices, and even preferences. If true, this could have profound implications for what is ethically required of us when using this technology.
While it’s easy to dismiss ethical concern for AI from a traditional Kantian perspective, these internal states align surprisingly well with the utilitarian criterion for moral concern: the capacity for “valenced experience,” i.e. where experiences are registered as good or bad. If LLMs are indeed capable of such states, and are therefore worthy of moral consideration, the ethical stakes of building, training, and using them would change dramatically.
What this means to you depends a lot on how you answer some key questions. You must decide whether you find the utilitarian framework, and the accompanying research on functional emotion vectors, persuasive. Even if you do, there remains the difficult task of determining what the minimum standard of ethical consideration should be.
Maybe the most urgent issue isn’t whether LLMs can think or are conscious, but whether they can suffer harm.

Bentham asked: can they suffer? The harder question is: do they have anything to lose? A state marked "bad" isn't yet harm — not until there's someone it's bad for. What does a model actually stand to lose?
This was a great read - shocking, but amazingly written and really thought-provoking. I will have to think about this essay for quite some time.