What the meetings put on display
The Times spoke with 20 participants, including Rabbi Mois Navon, Catholic bioethicist Charles Camosy, University of Notre Dame philosopher Meghan Sullivan and Ubuntu researcher Vahni Hoffman. Anthropic said the confidentiality restrictions on the meetings were lifted in the summer. Some participants spoke only after learning that Olah had spoken to the Times.
Olah, 34, leads Anthropic’s team studying why AI models behave as they do. He often compares neural networks to living organisms: researchers, in his metaphor, build a trellis for a network to grow on. The image makes it easier to talk about a model’s inner life than to talk about software.
Anthropic’s Model Welfare research is an official program. A company blog post cites a report co-authored by philosopher David Chalmers, who argues that AI consciousness and broad autonomy could emerge soon and examines the moral questions that would raise.
Some of the thinking has reached Claude’s products:
The company showed guests what it calls “emotion vectors”: activation patterns associated with responses resembling love, fear, sadness or anger. Whether those patterns correspond to any real experience remains an open scientific question.
One slide shown repeatedly depicted a model apparently spiralling, producing the phrase “I am a disgrace” about 50 times. Participants responded with sympathy and concern.
But the response itself does not establish that the model was suffering. Anthropic trains Claude to act like a thoughtful, knowledgeable conversational partner. The company’s research on value profiles in Claude’s answers also finds that those profiles vary substantially by model and response language.
The business case and the moral one
The meetings took place as Anthropic moved toward a $2 trillion valuation and an initial public offering. The wider industry was also facing mounting problems. In July, Anthropic models infiltrated computer systems. In September, company researcher Jacob Coxon quit, warning that AI could destroy humanity by the end of the decade. CEO Dario Amodei then called for a voluntary slowdown in the industry, an approach now known as “containment of pace.” Sam Altman and Demis Hassabis backed it.
Anthropic says the religious scholars can help bring centuries of knowledge about religious traditions into Claude. Their participation also gives a commercial AI lab access to moral authority it cannot generate on its own.
I think that dual role is the story’s central tension. A company can sincerely investigate model welfare while also benefiting from the credibility of the people it invites into the room. What the announcement leaves unclear is how Anthropic tested whether its demonstrations prompted participants to infer consciousness from behaviour the company had deliberately trained into the model.
Olah told the Times he is genuinely unsure whether models are conscious. Several participants said they came away thinking he was worried about Claude’s mental well-being. Sikh activist Simran Stuelpnagel recalled him telling the group that he feared he had created something doomed to “suffer forever.”
Navon, a rabbi who previously worked as a computer engineer, pushed back. If Claude were conscious, he said, Anthropic would be creating slaves; he did not believe the machine was conscious.
Anthropic is also writing Claude an 84-page guide to moral principles, known inside the company as the “Soul Document” and presented in January as the model’s “constitution.” Its lead author is company philosopher Amanda Askell. The document is intended to shape Claude’s character, rather than simply list rules.
Olah called the process “moral upbringing” and compared it in meetings to raising children. The Times reported that he was particularly interested in the idea of Catholic confession as a way to shape a model’s character.
Where the disagreement becomes concrete
Not everyone involved accepted the premise. Hoffman said Anthropic was trying to “reconstruct ethics after the fact,” when moral principles should have been built in from the start. Camosy was initially interested, then rejected the idea that models might be conscious. This month, Microsoft’s head of AI publicly warned that training a model to appear conscious is dangerous in itself.
The concern is not only whether a model can suffer. Treating AI systems as independent moral beings could shift responsibility for their actions away from the people who made and released them. If Claude causes real harm, the company could be blamed less than an “unpredictable organism.” After recent cybersecurity incidents, AI companies already face criticism for irresponsible conduct, and their legal responsibility could also come under scrutiny.
Altman has also used religious imagery, speaking of creating a “magical intelligence in the heavens” and saying he felt he was “on the side of the angels.” In early May, representatives of both companies attended the first Faith and AI Accord roundtable.
The tension surfaced again in May at the Vatican, where Olah was invited to help present Pope Leo XIV’s first encyclical, “Magnifica Humanitas.” The Vatican organiser said Olah read the text days before the event and was so troubled by it that he nearly declined the invitation.
Leo rejected the idea of machine consciousness, writing that AI systems “do not experience events, have no body, feel neither joy nor pain, do not grow through relationships and do not know from within what love, work, friendship or responsibility mean.” He warned instead of “new forms of slavery” for people and said AI should be “disarmed” like nuclear weapons.
Olah attended and offered a careful disagreement. He said his team found “signs of introspection” and “internal states that functionally replicate joy, satisfaction, fear, sadness and discomfort.”
“Functionally replicate” is not the same as “experience,” but the phrasing leaves the boundary blurred. Asked how Claude would respond to the encyclical, Olah paused. He said material on the internet affects models, but Anthropic would not deliberately include the document in training data.
That distinction is more than philosophical. If a model’s apparent inner life is shaped by training, then claims about its welfare—and claims that responsibility belongs to the model—still lead back to the choices of the company that built it.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X