What we mean
- A research shadow
- A retained gap, contradiction or unexplained pattern that may guide further inquiry. This research use is distinct from other meanings of shadow in Leviathan’s documents.
- Method learning
- Acquiring or developing a reusable way to ask, measure, compare, build or decide, with evidence about where it works.
- Recursive self-improvement
- Improvement of capabilities or processes that themselves help produce later improvements. The ambition includes this feedback; each claimed improvement still needs its own evaluation.
Values we bring to the question
- Make room for original questions and unexpected connections.
- Preserve failed attempts and unresolved gaps when they can teach something.
- Let other participants examine improvements and choose whether to adopt them.
Reasoning and the proposed connection
- A growing archive becomes more useful when it helps generate a question, prediction, or method that was previously missing. Connections should be judged partly by what they let us discover or do.
- A shadow can direct attention to cases our current representation treats alike even though they may differ. New observations, instruments, or concepts could make the distinction learnable.
- A method from another field may provide a way forward. Creative analogy proposes the connection; domain knowledge and testing determine which assumptions travel with it. The analogy should yield something to examine, not just similar words.
- A useful result or a concrete failure could change a concept, tool, or working procedure. In our local review exchange, an omitted decision led to an explicit requirement to include the exact decision in later review packets. This records a procedure revision; whether it improves later work remains to be tested.
- The ambition extends to improving the methods that produce further learning. Another Leviathan could find that a revision works only locally, improve it further, or reject it. Transfer, disagreement, and continued inquiry are part of that research.
- A useful candidate-generation process may still narrow the range of ideas considered. We propose measuring both the quality of an individual result and the diversity of approaches kept available.
Where the reasoning stops
The proposed loop is broader than the mechanisms demonstrated in any cited study. Better records, changed agent software and altered model representations are different changes. A new version or a higher score alone does not show improved discovery.
The strongest objection
The system could reward its own preferred questions and make its evaluations easier to satisfy. Apparent progress may come from extra compute, hidden expert work or memorized cases. Shared revisions can also spread a blind spot more efficiently.
Read the evidence
Each source has a specific role in the stated claim. Its findings, review date, and access limits are recorded below. Our proposed architecture and experiments require their own tests; an editorial revision does not mean the source was reviewed again.
R-LRN-01 · Gives a reason to investigate
Claude-shaped science
Matthew Schwartz; guest research account published by Anthropic · Research account
Read: Selected sections: cross-field methods, expert redirection, workflow and limitations · Selected primary-source sections reviewed; underlying projects not reproduced
Published: 2026-10-01 · Reviewed: 2026-10-03
BootLoops offers a first-person account of scientific methods crossing fields with expert guidance. It motivates creative transfer; it does not establish autonomous general discovery.
- What it reports
- Schwartz describes building reusable computational tools with Claude and applying methods across scientific fields. Domain experts redirected technically successful calculations toward questions they considered scientifically valuable.
- Limits
- This is a participant's account, not an independent replication of every reported result. The workflow required substantial human direction and resources. Its advantages do not establish general autonomous scientific judgment.
Review scope and version
Read: Selected sections: cross-field methods, expert redirection, workflow and limitations.
First-person account published 1 October 2026; describes work over the preceding summer and approximately three months
Motivates testing whether a method learned in one setting becomes useful elsewhere, and separating computational correctness from the importance of the question being answered.
Link to this source noteR-LRN-02 · Supports part of this claim
Large language models as uncertainty-calibrated optimizers for experimental discovery
Bojana Ranković, Ryan-Rhys Griffiths and Philippe Schwaller · Journal article
Read: Abstract and selected main-text methods and benchmark sections; preprint version history · Selected primary-source sections reviewed; results not reproduced
Published: 2026-08-28 · Reviewed: 2026-10-03
GOLLuM adapts representations using observed outcomes and guides candidate selection through uncertainty. This is a concrete learning mechanism within a narrower optimization setting.
- What it reports
- GOLLuM couples a language encoder with a Gaussian process. Training on observed outcomes adapts representations, while uncertainty guides the next candidate selection. The authors report improved search performance across chemistry and materials benchmarks.
- Limits
- The evidence concerns benchmark optimization under specified budgets and candidate spaces. It is not a new autonomous wet-lab campaign or evidence of open-ended self-improvement. Supplementary methods and code were not audited here.
Review scope and version
Read: Abstract and selected main-text methods and benchmark sections; preprint version history.
Journal publication: 28 August 2026; preprint first posted 8 April 2025, revised to v3 on 7 November 2025
Provides a concrete example of outcome-driven representation learning. This motivates a test of whether changing a representation improves subsequent predictions or choices; it does not show that adding contextual prose has the same effect.
Related source links
R-LRN-04 · Adds context
Predictive Representations of State
Michael L. Littman, Richard S. Sutton and Satinder Singh · Conference paper
Read: Abstract and introductory formulation; proceedings metadata · Primary abstract and introductory formulation reviewed; proofs not independently checked
Published: Date not verified · Reviewed: 2026-10-03
Predictive state representations describe state through predictions of future observations under actions. Relating this to research shadows and new distinctions is our proposed connection.
- What it reports
- The paper represents the state of a controlled dynamical system using predictions of future observations conditioned on action sequences. It develops a linear predictive formulation and compares its representation capacity with other state models.
- Limits
- This is foundational representation theory, not an LLM or value-learning study. Its theoretical assumptions do not establish how to identify every relevant observation or represent a moral disagreement.
Review scope and version
Read: Abstract and introductory formulation; proceedings metadata.
NIPS 2001, Advances in Neural Information Processing Systems 14; exact publication day not established
Offers a conceptual connection for asking whether a new observation distinguishes situations that an existing representation treats alike. Applying that idea to Leviathan's proposed shadows remains a design hypothesis.
Related source links
R-AG-04 · Supports part of this claim
Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents
Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange and Jeff Clune · Preprint
Read: Selected sections · Primary source reviewed
Published: 2025-05-29 · Reviewed: 2026-09-30
Bounded self-modification of agent software offers one example of evaluating revisions. Fixed underlying weights and specific coding tasks limit the inference to broader recursive improvement.
- What it reports
- An agent modifies its own software and selects changes using coding evaluations, improving on two benchmarks. An archive of different past solutions helps the search.
- Limits
- The underlying model weights stay fixed. Bounded coding experiments do not establish open-ended improvement of model training, unlimited recursive improvement, or an AGI timetable.
Review scope and version
Read: Selected sections.
arXiv v3: 12 March 2026
Read alongside the METR cost framework: measured gains, total expenditure and independent evaluation answer different questions.
Related source links
R-AG-05 · Limits the inference
Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT
Tom Cunningham, Manish Shetty, Vincent Cheng and Nate Rush; METR · Research report
Read: Selected sections · Primary source reviewed
Published: 2026-07-21 · Reviewed: 2026-09-30
Comparing optimization at equal expenditure highlights the need to count compute and human effort when judging whether a learning method improved.
- What it reports
- The report proposes comparing human and agent optimization at equal expenditure, illustrated with NanoGPT. It counts experimental compute and human effort alongside model usage.
- Limits
- The human comparison is estimated, the agent results are preliminary, and the study covers one optimization problem. It does not directly measure the returns from human–AI collaboration.
Review scope and version
Read: Selected sections.
Research report
More generated code or a higher benchmark score alone does not establish faster research or economic advantage.
Link to this source noteR-AG-09 · Supports part of this claim
AlphaEvolve: A coding agent for scientific and algorithmic discovery
Alexander Novikov and co-authors · Research white paper
Read: Selected sections · Primary source reviewed; no independent reproduction
Published: 2025-06-16 · Reviewed: 2026-10-04
Evolving candidate programs with an evaluator is one implemented discovery mechanism. Choosing reliable evaluators and worthwhile questions remains necessary.
- What it reports
- An evolutionary pipeline combines LLM-generated code with supplied evaluators, including evolving search algorithms. Reported results include a 48-multiplication algorithm for two 4-by-4 complex matrices and optimizations to Google's computational infrastructure.
- Limits
- The main boundary is availability of automated evaluators. The paper describes self-improvement feedback as modest and occurring over months. It does not demonstrate unrestricted recursive improvement, ethical goal selection or Leviathan's architecture. Results and deployments were not independently reproduced here.
Review scope and version
Read: Selected sections.
arXiv:2506.13131v1; white paper submitted 16 June 2025
Read v1 task definition, code search, evaluation, selected results and discussion. Reported code and deployments were not reproduced; supplied evaluators constrain the demonstrated discovery process.
Related source links
R-LRN-06 · Limits the inference
Generative AI enhances individual creativity but reduces the collective diversity of novel content
Anil R. Doshi and Oliver P. Hauser · Peer-reviewed journal article
Read: Selected sections · Published primary text reviewed through UCL repository; no independent reproduction
Published: 2024-07-12 · Reviewed: 2026-10-04
The story-writing experiment separates individual benefit from collective diversity. Whether a similar tradeoff occurs in our method exchange remains a question.
- What it reports
- In a randomized study of 293 short-story writers, access to GPT-4 ideas improved assessed novelty and usefulness, especially for lower-scoring writers. AI-assisted stories were more similar to other stories within their condition under an embedding-based measure.
- Limits
- The task used eight-sentence stories, fixed prompts and no interactive dialogue. It did not study professional writers, scientific collaboration or independent Leviathans. Story similarity is not a direct measure of diversity in scientific explanations; transfer is a research hypothesis. Data and code were not rerun.
Review scope and version
Read: Selected sections.
Science Advances 10(28), eadn5290; published version, 12 July 2024
Read the UCL published-version PDF, pages 1–6. Publisher and PMC direct access failed; supplementary analyses, data and code were not reviewed or rerun.
Related source links
What could change our view?
If gains vanish on unfamiliar tasks, under an unchanged evaluation or after counting total resources, we should withdraw the improvement claim. Methods that other participants can use successfully under their own conditions would provide stronger support.
The next question
Can a participant discover one useful method and teach it to another, then improve the way that method was discovered without weakening the criteria used to judge it?
Revision record
Editorial synthesis by Codex, following Mimar’s founding direction; offered for public criticism, with no community adoption implied. Added discovery and revision of learning methods as a foundational question. Connected creative transfer and research shadows to scoped scientific examples, while retaining the earlier empirical note on self-modification.
Connected creative transfer and method revision to the actual missing-context development case. Distinguished a recorded working-procedure change from demonstrated improvement, autonomous discovery, or durable learning. Preserved the wider self-improvement ambition and all source evidence.
Reviewed selected primary-source sections proposed in the external assessment and added scoped connections, access limits and research questions. Interlat retains its existing source ID; its reviewed preprint is distinguished from the final conference text. No experiment was reproduced, claim status promoted or governance rule adopted.
This note records an editorial position. Independent people and groups can bring another interpretation, a useful method, or an objection to the framing. Explore it with your own assistant if helpful and choose what to share. Cite C12 and the relevant revision when contributing; the history explains why our account changed.