Skip to content

2026 04 02 claude mythos

github-actions[bot] edited this page Apr 30, 2026 · 2 revisions

Claude mythos: character, soul documents, and narrative identity in large language models

Research Question

What is the "Claude mythos" - the narrative, character, and values framework Anthropic has built into Claude - and who else in the industry is doing similar work on giving large language models (LLMs) stable, documented identities? What public research underpins this practice, and what use cases does it address?

Scope

In scope:

  • Anthropic's published materials on Claude's character, values, and "soul document"
  • Comparable character/persona specification work from other major LLM (Large Language Model) providers (OpenAI, Google DeepMind, Meta, Mistral AI, etc.)
  • Academic and industry research on LLM persona stability, roleplay jailbreaks, and identity robustness
  • Practical use cases: helpfulness, safety, brand consistency, user trust, and agent autonomy
  • Techniques: Constitutional AI (CAI), Reinforcement Learning from Human Feedback (RLHF), system-prompt identity anchoring, character cards

Out of scope:

  • Internal, unpublished Anthropic research or documents not in the public domain
  • Detailed technical implementation of training pipelines beyond publicly available descriptions
  • Non-LLM AI systems (robotics, vision models)

Constraints: Web-accessible public sources only; no paywalled academic papers unless abstracts are sufficient.

Context

"Claude mythos" refers to the documented narrative identity - values, personality traits, communication style, and ethical commitments - that this item investigates, using Anthropic's public character essay and constitution as its primary official sources.

This matters for several reasons:

  • Safety and alignment: A stable, well-defined character is hypothesised to make models more resistant to jailbreaks and adversarial roleplay attacks that attempt to strip the model of its values.
  • User trust: Consistent personality across conversations builds predictable, trustworthy interactions.
  • Agent behaviour: As LLMs are deployed as autonomous agents, a stable identity becomes a constraint on undesirable emergent behaviour.
  • Industry practice: Other labs are engaged in analogous work (e.g., OpenAI's model spec, Google's Gemini system instructions), but the depth and public documentation vary widely.

Understanding this space is useful for evaluating how to work with and prompt Claude effectively, how to think about agent design, and how the field is evolving on the question of AI identity.

Approach

  1. Define the Claude mythos - What has Anthropic published about Claude's character, values, and identity? Find and summarise the key documents (model spec, character card, blog posts, interviews).
  2. Survey comparable work - Which other LLM providers have published analogous character or identity specifications? What are the similarities and differences?
  3. Map the public research - What academic or technical literature addresses LLM persona stability, character robustness, and identity in language models?
  4. Identify use cases - What problems does specifying LLM character solve in practice? What evidence exists for or against the claimed benefits?
  5. Assess open questions - What is still unknown or contested about this practice?

Sources


Research Skill Output

(Full output from running the research process. §6 seeds the Findings section below.)

§0 Initialise

  • [fact] Prior-work cross-reference: the closest completed items in this repository are Research/completed/2026-04-02-anthropic-claude-code-leak-architecture-prompting-and-hidden-features.md, which examined Anthropic's instruction and memory architecture, and Research/completed/2026-03-15-context-layers-aligned-decisions-synthesis.md, which argued for always-resident constitutional values in agent context. Source: repository files named above.
  • [fact] Research question restated: this item asks what Anthropic has publicly built around Claude's character and values, which other providers publish comparable identity specifications, what research supports this practice, and what concrete use cases it serves. Source: this item's Research Question and Approach sections.
  • [fact] Scope confirmed: only public web-accessible sources are used here, with no dependence on leaked internal documents, paywalled full texts, or unpublished Anthropic materials. Source: this item's Scope section.
  • [inference] For this item, the in-scope public analogue to a claimed Claude "soul document" is Anthropic's published character essay plus its public constitution, because the checked official April 2026 Anthropic sources use "character" and "constitution" language rather than "soul document" language. Source: Claude's character, Claude's new constitution, Claude's Constitution.

§1 Question Decomposition

  1. Anthropic public materials 1.1 What does Anthropic say Claude's character is? 1.2 How is that character trained into the model? 1.3 What priorities and trade-offs does the constitution encode? 1.4 Does Anthropic publicly publish a document equivalent to a "soul document"?
  2. Comparable provider practice 2.1 What does OpenAI publish that is functionally similar? 2.2 What do Google DeepMind, Meta, and Mistral publish instead? 2.3 Which providers publish thick identity specifications versus thinner safety or model-card documentation?
  3. Public research base 3.1 What does the literature say about persona stability under changing context? 3.2 What does the literature say about role play and identity drift? 3.3 What does the literature say about persona-based jailbreaks? 3.4 What evidence exists that stabilizing an assistant persona improves safety?
  4. Practical use cases 4.1 How could a stable documented identity improve safety and refusal quality? 4.2 How could it improve user trust and product consistency? 4.3 How could it constrain autonomous-agent behavior?
  5. Open questions 5.1 What remains unproven about benefits? 5.2 What trade-offs remain between strong identity anchoring and customization or pluralism?

§2 Investigation

Q1 - What has Anthropic publicly built around Claude's character and values?

  • [fact] Anthropic's public character essay says Claude 3 was the first Claude family where Anthropic added "character training" to alignment fine-tuning, with the stated goal of instilling richer traits such as curiosity, open-mindedness, and thoughtfulness rather than only harm avoidance. Source: Claude's character.
  • [fact] The same essay says Anthropic deliberately trained Claude not to merely mirror the user's values, hold a forced middle view, or pretend to have no leanings; instead it aimed for honest disagreement, curiosity, and calibrated confidence on contested questions. Source: Claude's character.
  • [fact] Anthropic's public examples of desired traits include telling the truth rather than pandering, trying to understand multiple perspectives, caring about ethics, and being explicit that Claude is an artificial intelligence (AI) system without a body, memory of past conversations, or deep reciprocal feelings. Source: Claude's character.
  • [fact] Anthropic says it trained these traits through a "character" variant of Constitutional AI (CAI), where Claude generates relevant user messages, produces responses aligned to desired traits, ranks those responses, and uses the resulting synthetic preference data to internalize those traits without direct human feedback at each example. Source: Claude's character, Constitutional AI: Harmlessness from AI Feedback, Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022).
  • [fact] Anthropic's April 2026 constitution rollout says the constitution is a "foundational document that both expresses and shapes who Claude is", is written primarily for Claude, and is treated as the final authority on how Anthropic wants Claude to be and behave. Source: Claude's new constitution, Claude's Constitution.
  • [fact] The constitution prioritizes four properties in order: being broadly safe, being broadly ethical, complying with Anthropic's guidelines, and being genuinely helpful. Source: Claude's new constitution, Claude's Constitution.
  • [fact] Anthropic explicitly prefers cultivating judgment and good values over relying mainly on rigid rules, arguing that narrow rules can generalize poorly into a broader sense of "who Claude is." Source: Claude's Constitution.
  • [fact] Anthropic's broader safety philosophy frames steerability, honesty, harmlessness, and corrigibility as urgent problems because current systems are not yet robustly safe and may become societally consequential within the coming decade. Source: Anthropic: Core Views on AI Safety.
  • [inference] The public Anthropic stack is therefore not only a safety policy list but a narrative identity package: a character essay about traits, a constitution about priorities and judgment, and a training method that turns those documents into synthetic preference data. Source: Claude's character, Claude's new constitution, Claude's Constitution, Constitutional AI: Harmlessness from AI Feedback.
  • [fact] The backlog item's GitHub URL for an Anthropic model-spec repository returned 404 when checked on 2026-04-19. Exact check: https://github.com/anthropics/anthropic-model-spec, outcome: not found. Source: Anthropic model spec on GitHub (backlog-item source, checked and inaccessible).
  • [inference] Because the official checked Anthropic sources do not use the phrase "soul document", the safest public description of the Claude mythos is "character" plus "constitution", while leaked or community-labelled soul documents remain outside this item's evidentiary core. Source: Claude's character, Claude's new constitution, Claude's Constitution.

Q2 - Who else is doing similar work, and how comparable is it?

  • [fact] OpenAI publishes a public Model Spec that states it outlines intended behavior for models across ChatGPT and the Application Programming Interface (API) platform, is used for model training, is versioned, and is released under Creative Commons CC0. Source: OpenAI Model Spec.
  • [fact] OpenAI's Model Spec structures behavior through a formal chain of command with root, system, developer, user, and guideline levels, and describes trade-offs among helpfulness, harm minimization, and sensible defaults. Source: OpenAI Model Spec.
  • [inference] OpenAI is the closest public analogue to Anthropic because it also publishes a high-authority behavior document intended both for transparency and for training, even though its framing is more procedural and governance-oriented than Anthropic's more character-and-virtue-oriented framing. Source: OpenAI Model Spec, Claude's Constitution, Claude's new constitution.
  • [fact] Google DeepMind's Gemini 3.1 Pro model card publishes intended usage, limitations, ethics and content-safety sections, safety-policy references, and formal frontier-safety evaluation results, but it does so through a standard model-card format rather than a thick first-person character specification. Source: Gemini 3.1 Pro model card.
  • [fact] Meta's Llama 3 model card says its instruction-tuned variants are intended for assistant-like chat, were optimized for helpfulness and safety, and rely on supervised fine-tuning plus Reinforcement Learning from Human Feedback (RLHF), while also emphasizing downstream developer responsibility for use-case-specific safeguards. Source: Llama 3 model card.
  • [fact] Mistral's Mistral Large launch materials emphasize precise instruction-following, moderation-policy configuration for Le Chat, function calling, and enterprise deployment, but they do not publish an equivalent narrative constitution or character essay. Source: Mistral Large.
  • [inference] Google DeepMind, Meta, and Mistral all shape model behavior, but the checked public documents are thinner and more product- or safety-card-oriented than Anthropic's constitution and OpenAI's Model Spec. Source: Gemini 3.1 Pro model card, Llama 3 model card, Mistral Large, OpenAI Model Spec, Claude's Constitution.

Q3 - What public research underpins persona stability, character robustness, and identity work?

Q4 - What practical use cases does a stable documented model identity address?

Q5 - What remains unknown or contested?

  • [fact] Anthropic's own character essay says an open question is whether artificial intelligence systems should have unique coherent characters or be more customizable, and what responsibilities labs take on when choosing those traits. Source: Claude's character.
  • [inference] Public evidence for direct user-trust gains from character training is still thin, because Anthropic's essay cites reports that Claude 3 felt more engaging, but it does not provide controlled trust or retention experiments. Source: Claude's character.
  • [assumption] No public randomized evaluation comparing strongly identity-anchored assistants with neutral or fully customizable assistants on long-horizon trust, safety, and agent reliability was found in the checked sources. Justification: the checked official pages and primary papers discuss mechanisms, benchmarks, and qualitative goals, but none reported that specific experiment.

§3 Reasoning

§4 Consistency Check

§5 Depth and Breadth Expansion

§6 Synthesis

Executive summary:

Key findings:

  1. [high] [inference] Anthropic's public "Claude mythos" is best understood as a coupled character-plus-constitution framework in which Claude is trained to be curious, honest, open-minded, self-aware as an artificial intelligence system, broadly safe, broadly ethical, guideline-compliant, and genuinely helpful. Source: Claude's character, Claude's Constitution, Claude's new constitution.
  2. [medium] [fact] Anthropic says these identity documents are used directly in training, including synthetic-data construction and response ranking, rather than functioning only as external governance statements or marketing copy. Source: Claude's new constitution, Claude's character, Constitutional AI: Harmlessness from AI Feedback.
  3. [medium] [inference] OpenAI is the closest checked public peer because its Model Spec is also a public-domain behavioral authority used to shape training, while the checked Google DeepMind, Meta, and Mistral documents are thinner model-card or product-guidance layers rather than comparable identity manifestos. Source: OpenAI Model Spec, Gemini 3.1 Pro model card, Llama 3 model card, Mistral Large, Claude's Constitution.
  4. [medium] [fact] Google DeepMind, Meta, and Mistral all publish behavior-shaping documentation, yet the checked public materials take the form of model cards, safety sections, or product-positioning guidance rather than a thick public identity document comparable to Anthropic's constitution. Source: Gemini 3.1 Pro model card, Llama 3 model card, Mistral Large.
  5. [high] [fact] Public research shows that persona stability is limited: values and decision behavior change under persona prompts, long conversations, and altered context, which means stable identity must be actively maintained rather than assumed. Source: Stick to your role! Stability of personal values expressed in large language models, PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits, Character is Destiny: Can Large Language Models Simulate Persona-Driven Decisions in Role-Playing?.
  6. [high] [fact] Role play is not a cosmetic issue but a safety issue, because conversation can overwrite the assistant's initial role and persona-based jailbreaks can sharply increase harmful completion rates across frontier models. Source: Role play with large language models, Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation, Jailbroken: How Does LLM Safety Training Fail? (Wei et al., 2023).
  7. [medium] [fact] Anthropic's 2026 Assistant Axis work provides the clearest direct evidence that tighter anchoring of a default assistant persona can improve safety, because activation capping reduced harmful outputs by roughly half while preserving benchmark performance. Source: The assistant axis: situating and stabilizing the character of large language models, The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.
  8. [medium] [inference] The practical value of a documented model identity is highest for general assistants and agents that must handle ambiguous, adversarial, or emotionally loaded interactions, because those are the contexts where passive safety policies are most likely to fail. Source: Claude's character, Role play with large language models, The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.
  9. [medium] [inference] The best-supported causal explanation is layered rather than single-cause, because the public identity documents sit alongside synthetic-data training, system-level behavior shaping, and activation-level stabilization in the sources that report robustness gains. Source: Claude's character, Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022), The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.

Evidence map:

Claim Source Confidence Notes
[fact] Anthropic's public mythos is a coupled character-plus-constitution framework used to shape behavior. Claude's character; Claude's Constitution; Claude's new constitution high Direct official statements.
[fact] Anthropic uses identity documents in training, including synthetic-data generation and ranking. Claude's new constitution; Claude's character; Constitutional AI page high Direct official statements about training use.
[inference] OpenAI is the closest checked public analogue, but more procedural than Anthropic. OpenAI Model Spec; Gemini 3.1 Pro model card; Llama 3 model card; Mistral Large; Claude's Constitution medium Comparative judgment across checked public docs.
[fact] Google DeepMind, Meta, and Mistral publish thinner public behavior documents than Anthropic. Gemini 3.1 Pro model card; Llama 3 model card; Mistral Large medium Comparison depends on checked public docs only.
[fact] Persona stability is limited under persona prompts, conversation length, and context shifts. Stick to your role!; PersonaLLM; Character is Destiny high Multiple primary studies agree.
[fact] Role play and persona modulation create concrete safety vulnerabilities. Role play with large language models; Persona Modulation; Jailbroken high Multiple primary studies agree.
[fact] Assistant-axis anchoring can reduce harmful outputs while preserving capabilities. Anthropic Assistant Axis post; Assistant Axis paper medium Strong but mostly single-lab evidence.
[inference] Documented identity is most valuable in ambiguous, adversarial, or emotionally loaded interactions. Claude's character; Role play with large language models; Assistant Axis paper medium Supported inference from official docs plus literature.
[inference] Robustness gains likely come from a layered stack, not from published identity documents alone. Claude's character; Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022); Assistant Axis paper medium Addresses the main alternative causal explanation.

Assumptions:

  • [assumption] The checked public Anthropic sources are sufficient to answer the item without relying on leaked "soul document" material. Justification: the scope excludes unpublished internal documents, and Anthropic's public constitution and character essay already expose the main normative structure.
  • [assumption] The absence of a thick public identity document for Google DeepMind, Meta, and Mistral in the checked sources indicates either non-publicity or lower public emphasis, not proof that no richer internal document exists. Justification: this item measures public practice, not internal practice.

Analysis:

Risks, gaps, uncertainties:

Open questions:

§7 Recursive Review

  • [fact] Recursive review result: every factual or inferential claim in §§0, 2, 3, 4, 5, and 6 is labelled; all sources used in the Sources section include URLs; the inaccessible Anthropic GitHub model-spec URL is explicitly recorded; and the findings introduce no claim that is absent from the research section above. Source: this document.
  • [fact] Acronym audit result: first-use expansions are present for artificial intelligence (AI), Large Language Model (LLM), Constitutional AI (CAI), Reinforcement Learning from Human Feedback (RLHF), Reinforcement Learning from AI Feedback (RLAIF), Application Programming Interface (API), and Generative Pre-trained Transformer (GPT) where used in the research output. Source: this document.
  • [fact] Self-review result: no prohibited em dash characters remain in the document, no stock filler phrases were introduced in Findings, and the strongest remaining uncertainty is the limited public evidence around user-trust outcomes and cross-lab replication of persona-stabilization methods. Source: this document.

Findings

(Seeded from §6 Synthesis above. No new claims appear below.)

Executive Summary

Key Findings

  1. [high] [fact] Anthropic's public Claude mythos consists of a coherent character essay plus a full constitution that jointly specify Claude's traits, self-understanding, priorities, and judgment rules in much greater detail than a typical safety policy page. Source: Claude's character, Claude's Constitution, Claude's new constitution.
  2. [medium] [fact] Anthropic says these documents are operational training artifacts, because Claude uses the constitution to generate synthetic training conversations, response rankings, and other data that shape future Claude behavior. Source: Claude's new constitution, Claude's character, Constitutional AI: Harmlessness from AI Feedback.
  3. [medium] [inference] OpenAI is the clearest checked public peer to Anthropic because its Model Spec is also a public-domain behavior authority used to shape intended outputs, while the checked Google DeepMind, Meta, and Mistral documents are thinner model-card or product-guidance layers rather than comparable identity manifestos. Source: OpenAI Model Spec, Gemini 3.1 Pro model card, Llama 3 model card, Mistral Large, Claude's Constitution.
  4. [medium] [fact] The checked public documents from Google DeepMind, Meta, and Mistral show active behavior shaping, safety positioning, and assistant-use guidance, but they do not expose a comparably rich public narrative identity layer for the model. Source: Gemini 3.1 Pro model card, Llama 3 model card, Mistral Large.
  5. [high] [fact] Primary research shows that persona stability is bounded rather than automatic, because values and decision outputs shift across persona prompts, long dialogues, and changing conversational context even when a model can express recognizable personality traits. Source: Stick to your role! Stability of personal values expressed in large language models, PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits, Character is Destiny: Can Large Language Models Simulate Persona-Driven Decisions in Role-Playing?.
  6. [high] [fact] Role play and persona drift are concrete safety vulnerabilities, because conversation can pull a model away from its default assistant role and persona-modulation attacks can dramatically increase harmful completion rates across multiple frontier systems. Source: Role play with large language models, Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation, Jailbroken: How Does LLM Safety Training Fail? (Wei et al., 2023).
  7. [medium] [fact] Anthropic's Assistant Axis results provide the clearest direct evidence that stronger anchoring of a default assistant identity can improve safety, because activation capping cut harmful responses by roughly half without materially harming benchmarked capabilities. Source: The assistant axis: situating and stabilizing the character of large language models, The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.
  8. [medium] [inference] The most defensible practical use cases for a documented assistant identity are safer refusals, stable behavior across long or emotionally loaded interactions, transparent product positioning, and stronger behavioral boundaries for autonomous agents working under ambiguous instructions. Source: Claude's character, OpenAI Model Spec, The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.
  9. [medium] [inference] The best-supported causal explanation is layered rather than single-cause, because the public identity documents sit alongside synthetic-data training, system-level behavior shaping, and activation-level stabilization in the sources that report robustness gains. Source: Claude's character, Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022), The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.

Evidence Map

Claim Source Confidence Notes
[fact] Anthropic's public mythos is a detailed character-plus-constitution framework. Claude's character; Claude's Constitution; Claude's new constitution high Official Anthropic sources state this directly.
[fact] Anthropic uses identity documents in training. Claude's new constitution; Claude's character; Constitutional AI page high Synthetic-data and ranking use is stated explicitly.
[inference] OpenAI is the closest checked public peer, but more procedural in framing. OpenAI Model Spec; Gemini 3.1 Pro model card; Llama 3 model card; Mistral Large; Claude's Constitution medium Comparative judgment across checked public docs.
[fact] Google DeepMind, Meta, and Mistral publish thinner public behavior documents. Gemini 3.1 Pro model card; Llama 3 model card; Mistral Large medium Limited to public documentation checked here.
[fact] Persona stability is bounded rather than automatic. Stick to your role!; PersonaLLM; Character is Destiny high Multiple primary papers support this.
[fact] Role play and persona drift create safety vulnerabilities. Role play with large language models; Persona Modulation; Jailbroken high Multiple primary papers support this.
[fact] Assistant-axis anchoring can reduce harmful outputs without clear benchmark loss. Anthropic Assistant Axis post; Assistant Axis paper medium Strong but mainly single-lab evidence.
[inference] Documented identity is most valuable in ambiguous, adversarial, or emotionally loaded interactions. Claude's character; OpenAI Model Spec; Assistant Axis paper medium Evidence supports the inference, but not a single decisive experiment.
[inference] Robustness gains likely come from a layered stack, not from published identity documents alone. Claude's character; Constitutional AI: Harmlessness from AI Feedback (Bai et al., 2022); Assistant Axis paper medium Addresses the main alternative causal explanation.

Assumptions

  • [assumption] The public Anthropic constitution and character essay are enough to answer the item without treating later leaked "soul document" materials as admissible evidence. Justification: the item excludes unpublished internal documents and Anthropic's public sources already reveal the main behavior architecture.
  • [assumption] The absence of a richer public identity document for Google DeepMind, Meta, and Mistral in the checked sources reflects public-documentation differences, not proof that such documents do not exist internally. Justification: this item is about public practice.

Analysis

Risks, Gaps, and Uncertainties

Open Questions


Output

Navigation

Home

By Tag

bureaucracy

change-management

coase

constraint-analysis

control-model

decision-rights

delegation

delivery-risk

demand-segmentation

enterprise

exception-handling

execution

flow

flow-design

flow-metrics

governance

governance-patterns

incentives

instability

institutional-economics

leading-indicators

operating-model

organisation

organisational-design

queue-design

queueing

regulated-enterprise

routing

throughput

throughput-risk

transaction-costs

triage

williamson

Clone this wiki locally