LEVIATHAN.LIFE
All foundations

C10 · Language and learning hypothesis

How can one Levi learn what another means?

Statement assessed

A useful exchange should help another participant apply a concept or method in a new situation, while keeping its assumptions and unresolved questions available for examination.

Status of this statement: Proposed direction to investigate

This status applies to the statement above, within the limits discussed in this note. Evidence for a finding can motivate Leviathan’s proposed mechanisms without establishing that they work.

Working editorial note · Last recorded revision

What we mean

Local meaning kernel
A proposed, evolving set of concepts and relationships used by a particular participant or community. A concept’s meaning includes its context and version.
Selective learning
Learning a useful distinction or method while separately evaluating the sender’s interpretations, commitments and claims of authority.
A scoped canonical record
A version accepted for a stated purpose by identified participants. Acceptance within that scope does not make it a universal account.

Values we bring to the question

  • Participants should be able to learn without surrendering their own judgment.
  • Important uncertainty and minority interpretations should survive translation.
  • New ideas should be expressible before they have been fully tested.

Reasoning and the proposed connection

  1. A concept becomes useful when it helps someone distinguish cases, ask a better question or choose a method. Repeating its name does not show that this ability has transferred.
  2. For an empirical concept, a teaching record might connect observations, examples, predictions and failed cases. For a value, it might connect reasons, affected parties, objections and circumstances that call for reconsideration.
  3. These relationships could form an operational language for people and agents: a way to process questions, build tools and improve representations as well as describe them. Written explanations, structured records and learned representations are possible parts of that research.
  4. A candidate concept can begin as a question or an unexplained pattern. Accepting it for a particular use, and relying on it in a consequential decision, require further reasons suited to that use.
  5. A new participant offers a practical test. Can it use the distinction in unfamiliar cases, notice its limits and reject an unsupported conclusion? The history and teaching effort it needs are part of the cost of communication.
  6. Our proposed handoff should ask two different questions: does the recipient understand the method, and are the assumptions for using it in the new setting justified? A causal result may require new measurements or remain unusable even when its explanation is clear.

Where the reasoning stops

There is no demonstrated common language across arbitrary models and communities here. The proposal must accommodate different kinds of support for empirical claims, methods and values; a single mandatory evidence format could distort them.

The strongest objection

The package may leave essential meaning in an expert’s experience or a model’s prior training. Ordinary prose and examples might teach the same thing more cheaply. A compact language could also erase rare distinctions that matter most when something goes wrong.

Read the evidence

Each source has a specific role in the stated claim. Its findings, review date, and access limits are recorded below. Our proposed architecture and experiments require their own tests; an editorial revision does not mean the source was reviewed again.

R-LRN-03 · Gives a reason to investigate

GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions

Elias Stengel-Eskin, Newton Sander, Carlos Bonetti, Sasha Boguraev, James Bowler, Hale Sirin and Simon Kirby · Preprint

Read: Sections 4.3–4.6 and selected prompts in Appendix D · Selected primary-source methods and prompts reviewed; experiments not reproduced

Published: 2026-09-01 · Reviewed: 2026-10-03

GlossoGen studies evolving communication conventions under explicit incentives. Its newcomer setup includes selected history, so it motivates testing what a teaching package alone can transmit.

What it reports
In a constrained cooperative environment, some tested agents developed compressed conventions with productive combinations. Replacement agents learned these conventions with varying success; additional interaction history improved average performance.
Limits
Discussion prompts explicitly encourage shorthand. Newcomers receive selected earlier messages and environmental events, so this is not evidence of transfer using only a declared portable package. Results are task- and model-dependent, with substantial variation between runs.
Review scope and version

Read: Sections 4.3–4.6 and selected prompts in Appendix D.

arXiv v1: 1 September 2026

Motivates distinguishing a group's successful coordination from another participant's ability to learn its conventions. The proposed test should account for history, teaching cost and unfamiliar tasks.

Related source links

Link to this source note

R-AG-02 · Adds context

Enabling Agents to Communicate Entirely in Latent Space

Zhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, and Haochao Ying · Peer-reviewed conference paper; experimental review uses an earlier preprint

Read: Selected sections · Primary preprint sections and conference metadata reviewed; final PDF inaccessible; no reproduction

Published: 2026-07 · Reviewed: 2026-10-04

Learned adapters show one way to communicate without ordinary text. Their training and compatibility requirements leave open how independent participants would share meaning.

What it reports
Interlat v2 reports trained hidden-state communication, including a Qwen-to-LLaMA experiment. MATH Table 2 reports 36.88% overall for Interlat versus 38.35% for full CoT, while Level 5 is 15.80% versus 15.05%.
Limits
Internal model access, adapters and training are required. The results do not establish arbitrary-model interoperability or preservation of ethical distinctions. The final ACL PDF was inaccessible; the inspected v2 is not assumed identical to the conference paper or later revisions.
Review scope and version

Read: Selected sections.

ACL publication metadata: July 2026; inspected experimental text: arXiv:2511.09149v2, 7 January 2026

Rereview on 4 October extends the 30 September note: v2 setup, Tables 1–2 and training discussion were read, plus ACL metadata. Final ACL PDF was inaccessible; later preprints were not reviewed.

Link to this source note

R-AI-01 · Adds context

Emotion Concepts and their Function in a Large Language Model

Nicholas Sofroniew and 15 co-authors; Anthropic · Preprint

Read: Author summary · Primary source reviewed

Published: 2026-04-02 · Reviewed: 2026-09-30

The study links some internal concept representations to behavior. It does not show that an explicit meaning kernel changes those representations or transfers a concept.

What it reports
Vectors associated with 171 preselected emotion concepts were extracted from Claude Sonnet 4.5. Interventions changed preferences and some alignment-related behaviors, supporting a functional role for these learned concepts.
Limits
The 171 concepts are a starting list, not 171 discovered feeling centers. The representations mainly track local context. The blackmail experiment used an early, unreleased snapshot; subjective feeling was not established.
Review scope and version

Read: Author summary.

Research announcement: 2 April 2026; arXiv v1: 9 April 2026

The geometry has a partial replication in another model in R-AI-06. That replication does not repeat the original behavioral interventions.

Link to this source note

R-GOV-03 · Adds context

Institutional Ecology, ‘Translations’ and Boundary Objects: Amateurs and Professionals in Berkeley's Museum of Vertebrate Zoology, 1907-39

Susan Leigh Star and James R. Griesemer · Historical case study and analytical framework

Read: Selected sections · Primary source reviewed

Published: 1989-08 · Reviewed: 2026-10-04

Boundary objects offer a historical comparison for shared artifacts with different local meanings; they do not prove faithful translation.

What it reports
The authors analyze cooperation among museum scientists, amateur collectors and other participants whose purposes differed. They identify standardized methods and boundary objects as ways to coordinate: shared objects remain recognizable across settings while allowing different local uses. Their four object types are analytical categories, not an exhaustive classification.
Limits
The evidence is a historical institutional case, not a controlled test of a general collaboration architecture. The discussion acknowledges coercion and exclusion as other ways representations become shared. A common object does not by itself establish consensus, equal power or faithful translation.
Review scope and version

Read: Selected sections.

Social Studies of Science 19(3), August 1989, pp. 387–420; original article, not a later reprint

Read selected printed pages 387, 393, 409–411 and 413–414 by OCR from the coauthor-hosted scan. Publisher full text was unavailable. Original 1989 article, not a later reprint.

Link to this source note

R-LRN-05 · Limits the inference

External Validity: From Do-Calculus to Transportability Across Populations

Judea Pearl and Elias Bareinboim · Journal article

Read: Selected sections · Primary source reviewed

Published: 2014 · Reviewed: 2026-10-04

Causal transport requires justified assumptions about differences between settings. Understanding a method is separate from establishing that its empirical result applies.

What it reports
The paper formalizes when causal effects from an experimental population can be inferred for a target population using observational data. Selection diagrams encode assumed mechanism differences; derivations identify relevant measurements and transport formulas.
Limits
Results depend on defended causal assumptions. The paper leaves measurement error, uncertain graphs and finite samples unresolved; its simple graphical criterion is incomplete. It does not establish semantic transfer or Leviathan's effectiveness.
Review scope and version

Read: Selected sections.

Statistical Science 29(4), 579–595 (2014); arXiv:1503.01603v1, 5 March 2015

Read arXiv introduction, causal models, selection diagrams, selected transportability statements and conclusions. Proofs were not audited; publisher full text was unavailable. Journal year and arXiv posting date differ.

Link to this source note

R-AI-11 · Limits the inference

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans · Research preprint, extended revision

Read: Selected sections · Primary source reviewed; no independent reproduction

Published: 2025-02-24 · Reviewed: 2026-10-04

Fine-tuning can change behavior outside its narrow task. This cautions against assuming selective behavioral learning; it does not test ordinary document exchange.

What it reports
Narrow insecure-code fine-tuning produced broader misaligned answers in studied models; educational framing changed the result. In the separate GPT-4o in-context experiment, up to 256 examples produced no observed broad misaligned responses.
Limits
Effects vary by model and answer format. Two training datasets were studied, with fuller controls for code; the authors flag simplistic evaluations. The findings do not establish that ordinary document exchange causes the same effect or determine which values deserve authority.
Review scope and version

Read: Selected sections.

arXiv:2502.17424v7, 20 January 2026; first submitted 24 February 2025

Read v7 controls, evaluation, model variation, in-context experiment and limitations. The related Nature DOI is metadata only; its separate text was not reviewed.

Link to this source note

What could change our view?

If new participants need the full original conversation, or cannot distinguish a useful method from an unsupported interpretation, the proposed transfer has failed. Reliable use on new cases with comparable teaching and checking effort would support a narrower, demonstrated benefit.

The next question

What would a newcomer need to learn one useful distinction, apply it elsewhere, and explain when it should not be used? Try a method moving between fields: can the recipient identify which conditions travel with it, test the proposed connection, and notice when a later change requires reviewing their own use? Compare the teaching effort with ordinary prose and examples.

Revision record

  1. Editorial synthesis by Codex, following Mimar’s founding direction; offered for public criticism, with no community adoption implied. Added selective learning as a foundational question. Distinguished candidate expression, acceptance for a use and reliance in decisions, and made newcomer understanding a proposed test.

  2. Opened selective learning through the practical need to use another field’s method. Extended the proposed newcomer test to conditions of use, later review, and an ordinary-prose comparison. No source finding or language record was revised.

  3. Reviewed selected primary-source sections proposed in the external assessment and added scoped connections, access limits and research questions. Interlat retains its existing source ID; its reviewed preprint is distinguished from the final conference text. No experiment was reproduced, claim status promoted or governance rule adopted.

This note records an editorial position. Independent people and groups can bring another interpretation, a useful method, or an objection to the framing. Explore it with your own assistant if helpful and choose what to share. Cite C10 and the relevant revision when contributing; the history explains why our account changed.