Real self-improving systems occupy combinations no single rung captures
Type: kb/types/note.md · Tags: foundations, self-improving-systems
The argument that improvement-pathway properties do not entail one another is made where each property is defined. This note supplies the other half: the combinations are not merely permitted by the ontology, they are the ones actually occupied by the systems it has to place. No canonical ordering follows from the profile alone — a scale over these thirteen would have to rank a proof-governed rewriter against a randomized relay bank against a repository of human-reviewed natural-language, and any total order it produced would have to import priorities the descriptive fields do not contain. A declared objective can supply those priorities; the profile cannot. Even then the supply is indirect, since an objective stated over outcomes ranks no architecture without a claim connecting structure to outcome.
Every row is a reading under a declared boundary. Change the boundary and the reading changes — most visibly for Commonplace, which is a member at all only under a frame that includes its maintainers, since attributions are elliptical until their parameters are named.
The placements
Selected profile fields, not the whole profile: the governance dimension appears here only through its evidential half, and what each methodology settles is left to the per-system accounts. The last column reads differently by update architecture. For the proposal-selection rows it is oracle domain — what the gate can warrant accepting. For the direct rows there is no gate to hand over, so it names what bounds the update rule's trustworthiness instead, since warranted autonomy is scoped to pathways with an evaluation to hand over.
| System | Update architecture | Reflective | Cumulative | Allocation | Evidential limit |
|---|---|---|---|---|---|
| Ashby's Homeostat | direct, viability-driven | no | no | computational | nothing — retention is negative |
| Parametric self-improvers | direct, gradient | no | yes | computational | training-time evaluation |
| Self-Improving Algorithms | direct, staged | no | yes | computational | the declared input distribution |
| Continual Harness | direct, refiner-mediated | yes | yes | computational | the Refiner's own judgment — end metrics observed, never gating |
| DreamCoder | proposal-selection | partly | yes | computational | statistical program fit |
| Gödel machine | proposal-selection | yes | yes | computational | what its proof system establishes |
| Knowledge-Centric Self-Improvement | proposal-selection | partly | yes | computational | benchmark oracles; debate for transfer |
| Self-Harness | proposal-selection | yes | yes | computational | two-split regression pass counts, the "held-out" split reused as selection data |
| Autogenesis | proposal-selection | yes | yes | computational | score against a fixed objective and safety invariants; acceptance rule and learnability mask sit outside the loop |
| Accumulated Behavioral Rules | proposal-selection | yes | yes | joint — capture human, consumption computational | one engineer's capture-time generalizability judgment; no per-rule outcome test |
| Darwin Gödel Machine | proposal-selection | yes | yes | computational | compile-and-viability check; benchmark score steers only parent sampling |
| Exo | proposal-selection | yes | yes | computational | build, test, immediate behaviour |
| Commonplace | proposal-selection | yes | yes | joint, by decision | tests and validators; human judgment |
What the hard cases teach
The Homeostat is the floor, and it is not a low rung. Operative, computationally autonomous, and non-cumulative at once: its retained setting steers behavior and determines whether reorganization fires, yet the successor comes from a random table and carries nothing of the incumbent. Any scale that reads autonomy as maturity puts a randomized relay bank above a human-reviewed repository.
Parametric learners break the equation of reflection with accumulation. They accumulate reliably through weights nothing inside them can read. This is the deployed default rather than a corner case, which is why an ontology that required reflection for membership would fail on the field's central systems.
Ailon et al. show cumulativity without either reflection or a gate. Its staged training phase is where the accumulation sits: a retained snapshot of a typical instance is built first, and the auxiliary search structures are then constructed against it. The stationary regime that follows retains those structures as the operative basis for later inputs without further improving them. Its objective is expected running time under a declared input distribution, and distribution shift is the boundary where the retained structure stops being warranted.
DreamCoder and the Gödel machine differ in gate kind, not gate strength. Both run reject-capable loops; one accepts on statistical program fit, the other only on proof. DreamCoder is also split internally — an inspectable symbolic library alongside an opaque recognition network — so its reflective coverage has to be reported per component rather than as a verdict about the system.
Knowledge-Centric Self-Improvement is the strongest external case for addressability. Its appendix traces a claim cited by id, challenged, split into two scoped claims, with the falsified branch retained as a rejection — the read-criticize-revise operations exercised computationally, not just structurally available. Its warrant splits: benchmark oracles are strong for pass/fail, while transfer-worthiness rests only on model debate.
The 2025–26 cohort clusters in a region no earlier row occupies. Self-Harness, Continual Harness, Autogenesis, Accumulated Behavioral Rules, and the Darwin Gödel Machine are all reflective and cumulative, and — except the rules pipeline's human capture — computationally allocated, yet their evidential limits are the thinnest in their block: pass counts partly reused as selection data, a fixed direct-update rule, a fixed objective the loop cannot revise, a single capture-time judgment, a bare viability check. Computationally autonomous operation over readable artifacts with thin acceptance warrant is the combination these five occupy. The six-path evidence inventory extends this cohort with HyperAgents, whose meta-agent revision and transfer evidence are not yet profiled in this thirteen-row casebook. The per-function casebook reads all six at finer grain. Two original placements deserve their own note: Continual Harness lands in the direct block for the Homeostat's reason, rejection collapsed into generation; and the Darwin Gödel Machine, carrying the Gödel machine's name, sits far from its namesake's row — acceptance on viability rather than proof, with score demoted to a search signal.
Exo and Commonplace differ most visibly in allocation. Both are reflective, cumulative proposal-selection pathways; Exo's self-representation is unusually literal, the source tree it edits being the organization that determines its behavior, with rebuild-and-restart as the wire from artifact to behavior. Exo is computational throughout, Commonplace joint and varying by decision. On the coarse fields reported here that is the sharpest difference between them — their warrant cells differ too, and finer readings would separate their governance, search, and protected kernels — and it is invisible to any measure that scores both as "self-improving."
Scope
- Placements are readings, not measurements. Each depends on a declared boundary and horizon, and several rest on a single published description rather than on independent inspection.
- "Partly" in the reflective column marks per-component coverage, not a midpoint on a scale — the point is that the verdict does not apply to the system as a whole.
- The casebook establishes that the combinations occur. It does not establish that any of these systems improved, which is a separate question about outcomes against a declared objective.
Relevant Notes:
- Self-improving systems — see-also: the curated head listing the four dimensions these placements are read against
- Self-improvement is relative to a declared objective — grounds: why each row names its frame, and why no ordering follows from the profile alone
- Accumulation counts dependence through the retained result, not through the evidence it caused — grounds: the cumulativity column's criterion
- Reflection buys addressability — grounds: what the reflective column is worth, and what parametric accumulation does without it
- Warranted autonomy is bounded by oracle domain — grounds: the evidential-limit column for the proposal-selection rows, why computational allocation does not fill it, and the scope that keeps it off the direct ones
- Measuring autonomy well enough to see it improve is an open problem — extends: why these rows still cannot be ordered even after the profile is fixed
- Ashby, Design for a Brain — ultrastability — evidenced-by: the operative, non-cumulative, non-reflective floor
- Self-Improving Algorithms — evidenced-by: cumulative retention with no representation and no gate
- DreamCoder — evidenced-by: a statistical reject-capable gate, with coverage split across a symbolic library and an opaque network
- Knowledge-Centric Self-Improvement — evidenced-by: addressability operations exercised computationally, with warrant split by question
- An omitted improvement-loop function and a frozen one need different repairs — contrasts: reads the 2025–26 rows at finer grain, by omitted or frozen loop function rather than by pathway profile
- Six reported self-improvement paths expose bounded redesign surfaces within supplied methods — evidenced-by: extends the detailed five-system reading with HyperAgents while this profile casebook remains at thirteen rows
- Ingest: Self-Harness — evidenced-by: computational proposal over harness surfaces with a partly self-consulted regression gate
- Ingest: Continual Harness — evidenced-by: refiner-mediated direct update with rejection collapsed into generation
- Ingest: Autogenesis — evidenced-by: versioned, rollback-capable retention under a fixed objective and learnability mask
- Ingest: Accumulated Behavioral Rules — evidenced-by: human capture-time judgment as the entire gate over an append-friendly rule file with in-place refinement and unresolved removal policy
- Ingest: Darwin Gödel Machine — evidenced-by: viability-gated acceptance with benchmark score demoted to a search signal, far from its namesake's row
- Exo — evidenced-by: reflective, cumulative, and computationally autonomous at once, over a literal source-tree self-representation
- Commonplace as a reflective system — evidenced-by: the human-inclusive joint-allocation reading