Which Existing Self-Improving Systems Are Theory Builders
Type: articles/types/article.md · Status: draft
Draft. The claims, system selection, and comparisons may change. Comments, corrections, and additional candidates are welcome on the repository's GitHub Discussions page.
Existing self-improving systems already report gains from retained knowledge and revised skills. They also supply mechanisms for diagnosis, criticism, revision, and reuse. Our program asks two separate questions of each system. Is it a theory builder, a system that states its theories, acts on them, criticizes what they say, and lets the result of criticism shape its next round? And does holding and criticizing those theories improve its later decisions? A third question is whether the whole process can run without people performing its internal roles.
This survey compares eighteen systems through retained code reviews, papers, and practitioner reports. No reported outcome was reproduced here. It places each system against the theory-builder conditions on the evidence cited, says where that evidence does not settle the placement, and identifies which experiments would resolve the remaining questions.
Two precedents already test improvement
Knowledge-Centric Self-Improvement keeps the solver machinery fixed while agents contribute experience to a shared knowledge base: "The only object that changes is the curated knowledge base." (Knowledge-Centric Self-Improvement, verbatim). Agents formulate claims in task forums, support or challenge them with evidence, discuss which observations generalize across tasks, and distill the results for later agents. Its protocol therefore includes criticism of claims as well as benchmark evaluation.
The paper freezes a learned knowledge bundle and transfers it to unseen tasks; a task-conditioned adapter "converts the shared donor asset into a short memo" for the task at hand before the solver acts (Knowledge-Centric Self-Improvement, verbatim). It reports improved zero-shot performance on Polyglot and ARC-AGI-1. This tests the value of the retained bundle beyond the tasks that produced it. A further comparison would isolate what the criticism contributed to that bundle's value.
Memento-Skills retains skills containing instructions and code. Its comparison removes several parts of skill improvement together: "no failure attribution, no skill rewriting, and no skill discovery" (Memento-Skills, verbatim). The full system reports 66.0% accuracy on the unseen GAIA test set, against 52.3% for this ablation. That is evidence for the combined improvement pipeline; separating failure attribution from rewriting and discovery would test the contribution of diagnosis. The system also trains a skill router, so its fixed solver weights do not make every component fixed-weight.
Both precedents address parts of the program's first testing goal: retained changes improving later performance. Their remaining mechanism questions call for narrower comparisons within the successful pipelines. Both are also theory builders on this evidence, as the placements below explain.
What the comparison asks
A system is a theory builder when it meets four conditions:
- Localized content. Its theories are stated in natural or formal language, so identifiable units, such as a claim, a skill file, or a ruling, carry their content.
- Consumption. The theories guide what the system does through what they say. A difference in a theory's content that matters to a decision changes the decision; such a change is operative.
- Criticism. A working process of attempted refutation aims at what identified units say, and a theory that fails is revised or replaced. Generating variants and keeping those with the best outcome score, with no stated reason bearing on what a variant says, is trial and error and does not meet this condition.
- Iteration. The result of criticism, a revised theory or the record of criticism, is kept and shapes the next round. Rounds of revision within one run meet this condition. A critic whose report no next conjecture takes up does not meet it.
Two properties above these minimums come in grades. Addressability grows as the units criticism can name become finer. Persistence grows as results last longer: within a run, across runs, or across problems. Neither decides membership. A run that freezes its result for another system to deploy ends the builder at the freeze; the deployment is not part of it.
The builder is the whole system that performs these operations, so people who propose, criticize, or select theories are inside it. A human-staffed system can therefore be a theory builder. The knowledge base calls a builder in which computation performs every internal operation autonomous.
Membership is not a claim of success. Whether holding and criticizing a builder's theories improves its capacity for future action is a separate learning claim, and it needs an outcome comparison: observing the whole theory-to-use path does not establish improved capacity.
We distinguish how a revision is produced from how it is accepted. A system may diagnose a mistaken assumption, revise the theory, and use an outcome gate to accept the revision. Testing a stated consequence can itself be criticism. The gate alone tells us little about whether the preceding process criticized a theory or simply generated another variant, so a gate does not settle condition 3.
The program adds distinct questions to this learning test. Can all internal roles run computationally with fixed model weights? Does the learned revision transfer to new cases? Does retaining the assembled theory save work compared with reconstructing it from the same evidence? A human-assisted system may provide evidence of learning before it provides evidence of autonomy, as the bootstrap supplement proposes for Commonplace.
Where the reviewed systems stand
The placements below use only the evidence this survey cites. Where that evidence does not show a condition, the placement is marked unsettled rather than inferred from the neighbouring conditions. On that basis:
- Theory builders (5): Knowledge-Centric Self-Improvement and Memento-Skills, which also report gains, and three human-staffed systems without a measured gain: OpenAI's agent-first product, Warp's skill improver, and Commonplace.
- Outside (1): Rainbow, whose supplied model is never criticized.
- Unsettled (12): Fluent, Wheelhouse, the Darwin Gödel Machine, the Huxley-Gödel Machine, Recuris, Harness Continual Learning, Dynamic Cheatsheet, Voyager, HyperAgents, Autogenesis, Exo, and Prime Agent. Most are unsettled on condition 3: the evidence shows an outcome gate or a success report, not criticism aimed at what a retained unit says. No placement turns on persistence.
Knowledge and skill pipelines with reported gains. Knowledge-Centric Self-Improvement meets the four conditions: its claims are stated, challenged with evidence, and distilled for later agents in the run, which act on them. The frozen bundle's transfer to unseen tasks is evidence of learning; by the freeze rule above it is not part of the builder. Memento-Skills meets the conditions as well: failure attribution names one responsible skill, rewriting changes what that skill says, and later tasks use the rewritten skill.
Human-assisted learning and machinery development. These systems show how people and agents retain knowledge and change their working machinery. Because the builder includes whoever performs its internal operations, the people in them are part of the system under assessment, not outside help. Their reports also identify internal roles that remain human.
- Fluent retains product code, expertise, scheduling, rejection, and reuse; people and the system jointly settle the brief, the behaviour specification, and the approach. Practitioner-reported. Unsettled: the report shows review of work items, but not criticism aimed at the retained expertise.
- Wheelhouse, in Steve Yegge's report, moves human rulings from custom to warnings, written doctrine, and programs that refuse actions. The corpus retained "old rulings that were obsolete or had changed" (Wheelhouse, verbatim), which is the maintenance cost of retention made visible. Unsettled: rulings are stated, enforced, and later changed, but the report cited here does not say whether a change answered criticism of what a ruling said.
- OpenAI's agent-first product keeps repository documents that explain the business domain and repairs them recurrently. People generalize: when agents struggle, engineers ask "what capability is missing? What constraint is unenforced?" and build the tool, linter, or test (agent-first product, verbatim). A human-staffed theory builder: the documents are stated, consumed by later agent work, and repaired when what they say is found wrong. No gain from them is measured.
- Warp's scheduled skill improver admits skill updates through human feedback and review. A human-staffed theory builder: an improver proposes an edit to what the skill says, people review the edit, and later runs use the updated skill. The account does not show that accepted edits improve later outcomes.
- Commonplace retains explanatory notes with scope and evidence and revises them under review; its system-definition artifacts describe the machinery. One recorded episode shows retained theory guiding computational search with the operator selecting what fit, without ablation. Models are not reliably pinned. Commonplace is a reflective, human-staffed theory builder whose learning is not yet shown.
Automated revision, selection, and reuse. These systems automate parts of the path from experience to later behavior. The entries distinguish the retained change from the mechanism that selects or consumes it. Placing them requires tracing a formulated theory and its criticism through that path, and the evidence cited here does not complete the trace for any of them.
- The Darwin Gödel Machine evolves coding agents around frozen models; a fixed diagnostician suggests improvements from the parent's logs, and admission is viability: "Only agents that compile successfully and retain the ability to edit a given codebase are added to the DGM archive" (Darwin Gödel Machine, verbatim). Unsettled on condition 3: the evidence cited does not show whether a diagnosis names what the parent agent's code or prompt says wrongly, and admission tests viability and score, not content. Later generations build on archived agents, so results persist across the run; persistence beyond it is not shown.
- The Huxley-Gödel Machine replaces score with descendant productivity for parent selection, and reports that immediate score predicts it poorly. Unsettled on condition 3 for the same reason as the Darwin Gödel Machine.
- Recuris proposes memory patches from traces and decides each through a deterministic paired held-out gate, with the memory coordinates supplied in advance: "The memory only grows, and it can afford to." (Recuris, verbatim). Unsettled on condition 3: patches are proposed from traces, but a memory that only grows shows no retained entry revised or rejected.
- Harness Continual Learning retains edits that improve the current task, respect sampled anchors, and pass validity checks; held-out forgetting persists even at zero loss on the anchors. Unsettled on condition 3: acceptance is by outcome and constraint checks, and the evidence cited does not show whether a proposed edit states what was wrong with the component it changes.
- Dynamic Cheatsheet curates notes from solver traces into later prompts; the curator prompt is the only gate. Unsettled on condition 3: the curator rewrites notes, but whether it criticizes what a note says is not shown.
- Voyager admits executable skills on a critic's success report. Retained skills supply both prompt context and executable code for later tasks. Unsettled on condition 3: the critic reports task success, and criticism of a skill after admission is not shown.
- HyperAgents evaluates generated patches and replays selected parent lineages into later generations. The replayed code changes future execution; the reviewed patches do not carry an explanation of why they worked. Unsettled on condition 3, closest to selection by outcome score.
Machinery for self-modification and recovery. These systems make changes persistent and recoverable, with different limits on what starts or selects an improvement.
- Autogenesis can write to many forms, but its selection is weaker than its versioning, and public implementations are incomplete. Unsettled on condition 3.
- Exo supports self-inspection, revision, restart, rollback, and preserved failure evidence, with no automatic trigger from experience to improvement. Exo is machinery that a builder could use; whether a system built on it criticizes its theories and builds on the result is unsettled.
- Prime Agent retains versioned prompts, memories, skills, and subagent specifications without weight updates; one case found a specification exploit and "preserved it as a reusable skill" (Prime Agent, verbatim). Persistence does not ensure sound admission. Unsettled on condition 3: the preserved exploit shows a harmful skill admitted, and the evidence cited does not show how admission judges what a skill says.
A supplied model guides adaptation. Rainbow selects adaptation strategies from a causal model of a running system whose vocabulary, goals, and strategies are fixed and supplied. It is a mechanism comparison for retained causal models, not a case of learning one. It is outside: its model is a fixed theory that the system never criticizes, so it fails condition 3.
The next comparisons
For most unsettled systems, the deciding evidence is whether a stated reason bears on what a retained unit says. Recording the proposer's diagnosis next to each accepted change, and checking whether it names what the changed unit said, would settle condition 3. For the Gödel machines, running a retained archive on a problem it did not evolve on would measure how far their results persist.
For the knowledge and skill pipelines, vary the formulated criticism while keeping the underlying observations and evaluation conditions comparable. Then withhold or perturb a retained theory to test how its content affects later decisions. This connects two questions that a pipeline-level ablation leaves together: what produced the useful revision, and how the revision contributed to later improvement.
For the human-assisted systems, record which internal roles people perform and test computational replacements under matched demands. For the self-modification systems, follow one proposed improvement from its diagnosis through selection to later use and measured benefit. These comparisons serve the same program at different stages of automation.
Finally, compare retaining an assembled theory with reconstructing it from the same episode evidence, counting both cost and decision quality. This tests the program's persistence conjecture; either arrangement may learn. The testing supplement develops these comparisons into controlled task-family experiments.