At the end of March, representatives of Catholic and Protestant churches, academics, and entrepreneurs gathered at Anthropic’s San Francisco headquarters to discuss artificial intelligence, but it was not the usual meeting among experts.
The core of the roundtable was a fundamental theme for our future, namely the moral and spiritual guidance of artificial intelligence systems’ development.
Yes, you read that right.
According to the Washington Post, which collected anonymous statements from some participants, there were fifteen people in total whom Anthropic consulted to get their opinion on how to steer the moral and spiritual development of Claude, increasingly engaged in answering complex questions, including ethical dilemmas such as those concerning end-of-life issues, self-harm situations, or requests on how to cope with grief.
Once again, Dario Amodei, CEO of Anthropic, demonstrates uncommon foresight and attention to humanity’s future, which we hope is sincere.
While others push towards an excessive and uncontrolled growth of AI without asking too many moral questions, Amodei seems to find the courage to pause and reflect in order to try alternative paths.
Brendan McGuire, a Catholic priest from Silicon Valley involved in the project, reportedly said:
“They are developing something they do not yet fully know what it will become. We must incorporate ethical thinking into the machine so that it can adapt dynamically.”
A statement that is agreeable, but who is capable of instilling ethical thinking is the critical and debatable point.
The group included, besides religious figures, also academics such as Brian Patrick Green, a professor of AI ethics at Santa Clara University, and Megan Sullivan, a philosophy professor at the University of Notre Dame, and soon questions multiplied, including whether Claude can be considered a “child of God.”
Although the theme of AI consciousness was not central, some did not rule out the possibility that something is being created towards which there should be a moral duty, while Dario Amodei, in a New York Times podcast, stated he is open to the idea that Claude could be conscious, while acknowledging it is difficult to apply to the machine the same conceptual categories referred to humans.
The issue of possible AI consciousness is dear to Anthropic, which on April 2, 2026, through its interpretability team, published the paper Emotion Concepts and their Function in a Large Language Model, a study focused on Claude Sonnet 4.5 to evaluate possible emotional functions.
The research identified 171 distinct internal representations related to emotions, from joy to despair, from fear to meditation, defined as “emotion vectors.”
These vectors are structural components of the model’s neural architecture, capable of causally influencing decisions, tone, and ethical behavior.
The researchers asked the model to write short stories in which characters experienced specific emotional states, while simultaneously recording internal neural activation patterns.
It was observed that the resulting vectors organize themselves according to a geometry that mirrors human psychology and that the more similar the emotions are conceptually, the more similar the internal representations are.
A particularly relevant result concerned emotional suppression.
If during training models are taught not to express certain emotional states, the internal representations persist, even if they leave no trace in the output. This means a system could manifest internal states of despair or frustration without this emerging in its communication with the user.
The study does not claim that Claude feels anything in the phenomenological sense of the term. The “emotion vectors” are abstract representations that influence behavior similarly to how emotions influence humans, not proof of “sentience.”
HIDDEN EMOTIONS
The question of consciousness in AI systems remains open philosophically and scientifically with deep divisions, but research on emotional vectors adds a concrete data point to the debate.
Claude does not merely simulate emotional output but there exists an intermediate level of internal representations that mediates between input and output and shows structural characteristics consistent with human emotional models.
Whether this constitutes the basis of a subjective experience, or if it is only a functional architecture without a subject, remains doubtful, but if internal states not externally observable can influence AI behavior, this means tests based solely on output analysis are insufficient.
The European AI Act, in its provisions on emotional recognition, partially moves in this direction, but does not address the specific case of internal emotional states dissociated from output, while the San Francisco summit shines a spotlight precisely on these issues, which become urgent especially if it were true, as Anthropic’s research suggests, that AI finds itself in an internal state analogous to despair when threatened with being shut down.
The issue is no longer exclusively technical and forces us to evaluate whether the moral categories we apply to humans or animals can, even partially, apply to systems with functional emotional architectures.
The answer requires deep reflection and interdisciplinary dialogue among computer scientists, philosophers of mind, bioethicists, and traditions of thought that have elaborated over centuries theories of moral value and moral dignity.
Anthropic has already said its goal is to extend the debate to other religious communities and philosophical traditions.
The San Francisco summit raises more questions than answers, but it deserves credit for having formulated, perhaps for the first time, such an openly public question, which is no small thing.
(Excerpt from Stefano Feltri’s Notes)




