What we mean
- Internal representation
- A pattern of model activity associated with information the model uses.
- Causal intervention
- Changing an internal pattern and testing whether the model's behavior changes.
Values we bring to the question
- Understand the systems that increasingly shape our lives.
- Make the influence of inherited human concepts open to scrutiny.
Reasoning and the proposed connection
- If a learned concept changes a model's decisions, understanding its role matters for both capability and safety.
- This connects the study of language and culture to mechanisms we can investigate. A human name for a representation remains a hypothesis about its function.
- Leviathan proposes local meaning kernels that connect concepts, principles, rules, and their versions. Those explicit records are a design layer; they are distinct from learned internal model representations.
- A proposed test would ask whether that context changes interpretation and action on unfamiliar cases, including when a community revises a concept or disagrees with another community. The cited findings motivate this question; they do not answer it.
Where the reasoning stops
The findings concern particular concepts, models and experimental settings. They do not establish that all culture is contained in language or that every model operation can be reduced to a human concept.
The strongest objection
A representation extracted using human labels may partly reflect the researcher's choice of categories. An intervention can change behavior without proving that a model uses the concept as a person does. Similarly, supplying a detailed value context could produce fluent agreement without reliable changes in action.
Read the evidence
Each source has a specific role in the stated claim. Its findings, review date, and access limits are recorded below. Our proposed architecture and experiments require their own tests; an editorial revision does not mean the source was reviewed again.
R-AI-01 · Supports this claim
Emotion Concepts and their Function in a Large Language Model
Nicholas Sofroniew and 15 co-authors; Anthropic · Preprint
Read: Author summary · Primary source reviewed
Published: 2026-04-02 · Reviewed: 2026-09-30
Interventions on emotion-related representations change measured behavior.
- What it reports
- Vectors associated with 171 preselected emotion concepts were extracted from Claude Sonnet 4.5. Interventions changed preferences and some alignment-related behaviors, supporting a functional role for these learned concepts.
- Limits
- The 171 concepts are a starting list, not 171 discovered feeling centers. The representations mainly track local context. The blackmail experiment used an early, unreleased snapshot; subjective feeling was not established.
Review scope and version
Read: Author summary.
Research announcement: 2 April 2026; arXiv v1: 9 April 2026
The geometry has a partial replication in another model in R-AI-06. That replication does not repeat the original behavioral interventions.
Related source links
R-AI-06 · Supports part of this claim
Replicating the Geometry of Emotion Representations in a Base Open-Weights Model
Adam Hollowell; University of North Carolina at Chapel Hill · Preprint
Read: Selected sections · Primary source reviewed
Published: 2026-09-01 · Reviewed: 2026-09-30
A different base model partly reproduces the representation geometry; causal behavioral effects were not tested.
- What it reports
- In base Gemma-2-27B, much of the valence geometry and clustering of 171 emotion vectors is recovered. Some structure is already present in token embeddings. The arousal axis does not meet all stability criteria.
- Limits
- This supports the representation finding, not a causal effect on behavior. The study uses Claude-generated fiction, one different model, reconstructed methods and a linear representation assumption.
Review scope and version
Read: Selected sections.
arXiv:2609.22208v1; submission date listed as 1 September 2026
This is a partial geometric replication. The author does not count arousal as fully replicated and identifies structural-token confounds in some measurements.
Related source links
R-AI-09 · Supports this claim
How’s it going? Reinforcement learning in language models recruits a functional welfare axis
Andy Q Han, David J. Chalmers and Pavel Izmailov · Preprint
Read: Abstract · Primary source reviewed
Published: 2026-05-28 · Reviewed: 2026-09-30
Minimal reward training recruits pre-existing representations that affect behavior.
- What it reports
- Reward and punishment vectors extracted after maze training influence behavior in other tasks. The authors' controls support the view that training recruits pre-existing representations rather than creating them from scratch.
- Limits
- Functional welfare here means an estimate of doing well or badly relative to goals. It is not a claim about experienced pleasure or pain. This review covers the abstract.
Review scope and version
Read: Abstract.
arXiv v1
The study adds evidence about functional internal representations. Its publication date is May 2026, even when later coverage draws attention to it.
Link to this source noteR-AI-03 · Adds context
Emergent Introspective Awareness in Large Language Models
Jack Lindsey; Anthropic · Preprint
Read: Selected sections · Primary source reviewed
Published: 2025-10-29 · Reviewed: 2026-09-30
Some models have limited access to their internal representations; this does not validate every self-report.
- What it reports
- Some Claude models can identify injected concepts in certain settings and distinguish earlier internal representations from input text. Controls support a limited form of functional introspection.
- Limits
- Failures are common, and performance depends on context and post-training. The experiments do not establish human-like introspection, the reliability of every self-report, or subjective experience.
Review scope and version
Read: Selected sections.
First published: 29 October 2025; web revision: 1 January 2026; arXiv v1: 5 January 2026
The January web revision adds a control experiment and prompt corrections. The later arXiv submission is a separate publication event, not an independent replication.
Related source links
What could change our view?
Independent interventions that fail to reproduce the reported effects, or controls showing that unrelated directions explain them equally well, would weaken the claim. Replication across models, languages and contexts would strengthen it.
The next question
Which effects survive different concept lists, model families, languages, and post-training methods? For our proposed local kernels, compare ordinary instructions with explicit concept–principle relationships at a comparable total budget. Test behavior when meanings change, values conflict, and unfamiliar cases appear; record failures as well as agreement.
Revision record
Added the functional-representation claim with a partial replication and its limits. Clarified that 171 is the study's preselected concept list.
Connected the functional-representation findings to a proposed test of local meaning and value contexts. Distinguished explicit kernels from learned model representations and verbal agreement from behavior. Empirical statement, support status, and source reviews are unchanged.
This note records an editorial position. Independent people and groups can bring another interpretation, a useful method, or an objection to the framing. Explore it with your own assistant if helpful and choose what to share. Cite C01 and the relevant revision when contributing; the history explains why our account changed.