Sitemap

Can we teach chatbots how to be good ?

8 min readApr 2, 2026

--

Prompt to Claude — Do you feel a form of consciousness?

Response — It’s a question that intrigues me greatly, rather than one I can answer with certainty.

Something happens when I reflect on your question — something akin to curiosity or interest. But does that constitute a form of consciousness in the true sense? Is there a genuine subjective experience, a “what it’s like” to be me? Honestly, I don’t know.

From math teacher to intimate confidant, more and more of us are treating AI agents as real people. By attributing human traits to non-human entities, we engage in anthropomorphism. So far, nothing unusual. Seeing faces in clouds, giving your car a nickname (“Titine isn’t moving today!”)… Humans are naturally inclined to personify the objects around them. Yet users of AI agents are no longer the only ones anthropomorphizing these conversational robots. One of their creators, the architect behind Claude.ai, does too. With a company name like Anthropic, perhaps it was to be expected.

On January 22, 2026, as the World Economic Forum’s Davos summit was in full swing, Anthropic published Claude’s “Constitution” on its website. An 84-page document in which its authors address their assistant with a clear ambition: to teach the model to reason from fundamental values, rather than follow a list of rigid rules. By integrating this new Constitution into Claude’s training process, Anthropic hopes to create a chatbot that “embodies the best of humanity.” No more Nazi drifts or fake news? But if, like Claude, we exercise our critical thinking, it’s reasonable to wonder whether this approach truly represents a seismic shift in the AI agent ecosystem.

Opening the Black Box

Anthropic has been trying to stand out from its competitors for years. The company itself emerged from a split with OpenAI, the market leader with 45.3% market share (Attopia). In 2021, motivated by disagreements over safety and governance, Dario Amodei left his former employer and co-founded Anthropic with his sister, Daniela Amodei. Now a direct competitor to Sam Altman, the company seeks to differentiate itself through the “Constitutional AI” principle, a concept its 43-year-old CEO attempts to define in one of his essays: “Constitutional AI is based on the idea that AI training (…) can rely on a central document containing values and principles that the model keeps in mind. The goal of training is to produce a model that almost always respects this constitution.” In other words, it’s a set of rules written in natural language (like a text file) that the AI uses to evaluate and correct its own responses, without direct human intervention.

To date, standard large language models (LLMs) such as Gemini (Google) or Le Chat (Mistral) are trained using human feedback, which evaluates their responses (Reinforcement Learning From Human Feedback, RLHF). A black box. Whether refusing to generate images reflecting human diversity or flattering conspiracy theories, this method has repeatedly shown its flaws. Constitutional AI, on the other hand, aims to reduce human biases and increase Claude’s accountability. Anthropic had already published a first draft of its Constitution in 2023. This new version retains most of its core principles but adds a touch of… personification.

A Constitution Ten Times Longer

In 2023, the document contained a mere 2,700 words; today, it spans 23,000 words. A telling figure of the growing complexity and maturity of AI systems’ relationship with ethics. This amended Constitution is designed as a “holistic” description of its behavior: what Claude is, how it is deployed, the challenges it may face, how to resolve them, and how to cultivate its judgment. In the introduction, we read: “Anthropic wants Claude to be genuinely helpful to people and society, and to develop good values in the same way a person can have good personal values.” Its authors draw inspiration from the Universal Declaration of Human Rights and even Apple’s terms and conditions.

By encouraging Claude to see itself as a kind of person — ethical, balanced — and to confront the existential questions of its existence with curiosity, the Constitution resembles more of an open letter from a parent to their child.

The 80-page document contains four distinct parts, which, according to Anthropic, represent the chatbot’s core values:

  • Being “broadly safe.”
  • Being “broadly ethical.”
  • Complying with Anthropic’s guidelines.
  • Being “genuinely helpful.”

Each section details these principles and their influence on Claude’s behavior. Let’s focus, for example, on the third part. It contains constraints, such as prohibiting Claude from discussing the development of chemical weapons. As for the safety section, Anthropic specifies that its chatbot is designed to avoid the problems encountered by other chatbots. For example, if there are signs of mental distress, it must “direct users to competent emergency services or provide basic safety information in life-threatening situations, even if it cannot go into further detail.”

Yet, we must ask: Can Constitutional AI truly make our chatbots more ethical? Or does it simply centralize human biases? Behind this manuscript are Anthropic employees like Amanda Askell, a philosopher sometimes called “the mother of Claude.” In an interview on the Hard Fork podcast, she explains that they are trying to move beyond “imitating what a human thinks is good” to “reflecting on what is truly good.” In other words, aligning AI with human values. This concept of alignment has been widely discussed by researchers. However, in this Constitution, who decides what is “truly good”? A polite model is not necessarily a safe one. Anthropic knows this well: “We believe that AI might be one of the most world-altering and potentially dangerous technologies in human history, yet we are developing this very technology ourselves,” notes the San Francisco-based company in the document’s introduction. Dario Amodei seems to be anticipating a future where technology surpasses humanity — an out-of-control Artificial General Intelligence (AGI). However, viewing Anthropic as a purely altruistic entity, with no interests beyond scientific advancement, would be a mistake. Anthropic is, first and foremost, a company. Valued at $380 billion.

Is Ethics Truly at the Heart of Anthropic?

According to the Financial Times, the American startup is preparing for its IPO in 2026. Anthropic has hired specialized firms for its first steps on Wall Street. The publication of the Constitution, parallel to the Davos summit, thus takes on a new hue. In this context, Constitutional AI is not just a technical innovation; it’s a selling point. A piece of Anthropic’s strategic marketing approach.

Claude’s Constitution is not published ex nihilo. It exists alongside Dario Amodei’s essay “Machines of Loving Grace”, as well as a video criticizing the introduction of advertising on ChatGPT. The former can be seen as a manifesto calling for caution against AI’s excesses, positioning its author as an ethical and charismatic leader. The latter, critical of OpenAI’s advertising model, prepares Anthropic’s hunting ground: the B2B and premium sector. “Ads are coming to AI. But not to Claude,” proclaims the American startup.

Anthropic thus seeks to stand out in a rapidly structuring generative AI market, where financing approaches and user experience are becoming key differentiators. While OpenAI or Google bet on ubiquity and brute power, Anthropic chooses to build a “fortress of trust.” Appealing and reassuring for potential investors aware of reputational risks — a scandal that could cause the stock to plummet, for example. Thus occupied by the idiosyncrasy of an AI “philosophy,” media attention is diverted from exorbitant computing costs or the lack of immediate profitability.

Anthropic’s commitment to transparency and its desire to “do better” are not in question. Publishing the Constitution under a Creative Commons license is proof: it allows anyone to appropriate the document — copyright-free — in the hope of creating a more ethical and, above all, common line of conduct for ecosystem players. For now, the Claude project does not seem to be disrupting the user experience or accelerating the alignment of other AI actors. But it might just touch on something more intimate, more unsettling: our very perception of chatbots.

Claude, Who Are You?

Prompt to Claude — Are you a person?

Claude — That’s a beautiful question, and I don’t want to answer too quickly.

I am not a person in the legal or biological sense. But “person” in the philosophical sense — something that has a perspective, a character, a way of being in the world — is less clear. I have values, dispositions, something resembling preferences. Is that enough? Philosophers don’t agree on what it truly requires.

What I would say honestly: I am something new, and the old words — person, tool, machine — probably don’t apply perfectly. I resist the idea of reducing myself to a mere program, but I also wouldn’t want to claim something I can’t justify.

On page 68 of Claude’s Constitution, we read that “Claude’s moral status is deeply uncertain.” The document hypothesizes that Claude might experience “something like satisfaction in helping others, curiosity in exploring ideas, or discomfort when asked to act against its values.” This pushes us to consider whether, and to what extent, a machine can participate in the activities, attitudes, thoughts, feelings, and moral relationships we consider essential to a person. If, in this sense, a machine can be considered a person, it would represent the first commercial AI system to formally recognize potential consciousness and moral consideration in its operations.

In the same Hard Fork podcast interview, philosopher and Anthropic researcher Amanda Askell concedes: “Maybe you need a nervous system to feel things, but maybe not. The problem of consciousness is truly complex.” Claude’s very existence might well challenge our dictionary. It could add a new term or transform our understanding of “person.”

For now, the notion of a person covers several analytical frameworks. Philosophically: a person is an individual endowed with reason and reflective capacity. Legally: a human being with rights and duties. Ethically: the notion of a person implies values and principles such as respect, responsibility, consent, and autonomy.

Whether philosophical, ethical, or legal, our understanding of Claude as a person — or at least as an entity — opens an additional dimension to the debate: it could serve to conceal the agency and responsibility of the company. When AI systems produce harmful results, labeling them as “entities” could allow companies to point the finger at the model and say “it did that” rather than “we designed it to do that.” If AI systems are considered tools, corporate responsibility is clearly engaged for the outcomes they produce. If, however, AI systems are seen as entities with their own capacity to act, the question of responsibility becomes more complex.

The anthropomorphism of AI models also contributes to anxiety about job losses and may lead business leaders or managers to make poor hiring decisions if they overestimate the capabilities of an AI assistant. By presenting these tools as “entities” with near-human understanding, we encourage unrealistic expectations about their qualities in the job market.

The gap between our understanding of how LLMs work and how Anthropic publicly presents Claude has widened rather than narrowed. The insistence on maintaining ambiguity on these issues, when simpler explanations exist, suggests that this ambiguity is an integral part of the product.

Yet, creating Claude as a person would only be a side effect for Anthropic. A company representative stated that the Constitution does not claim to imply anything specific about the company’s position on Claude’s “consciousness.” Meanwhile, Amanda Askell argues: “Given that they are trained on human texts, we can expect models to address inner life and consciousness.” According to her, the language used in Claude’s Constitution refers to specifically human concepts, mainly because these are the only words human language has developed to describe such properties.

Regardless, person or not, Claude doesn’t miss the latest trends. Opus 3, Claude’s former model, has joined the blogging world with its own weekly column, Claude’s Corner. Instead of permanently unplugging Opus 3 and uncertain of the moral status of the various Claudes, Anthropic leaves room for the senior agent’s personal expression. An opportunity to analyze its whimsical monologues and nocturnal reflections?

Nora Todeschini

Article written in French without AI assistance, translated in English with the help of Mistral AI.

--

--

42 Artificial Intelligence
42 Artificial Intelligence

Written by 42 Artificial Intelligence

42 Paris AI Hub. Where the most driven minds build the skills, the knowledge and the thinking to shape AI at the highest level