Six reported self-improvement paths expose bounded redesign surfaces within supplied methods

Type: kb/types/note.md · Tags: foundations, self-improving-systems

Six reports from 2025–26 expose paths that revise readable harness, rule, resource, or agent-code artifacts while consequential parts of the method remain outside the reported update surface. Five close a causal path from evidence through installation to later use for at least one kind of change. Accumulated Behavioral Rules establishes installation and loading, but not isolated behavioral dependence. The result is not a ranking of whole systems: it is a bounded comparison of what each reported path demonstrably changes and what it continues to take as supplied.

The distinction requires three evidence levels that the reports do not always separate. Closed means objective-bearing evidence changed represented organization and a later operation depended on the installed result. Unestablished means the record leaves at least one of those causal edges open. Supplied means authority-bearing machinery used by the path sits outside that path's update surface. A declared editable surface is therefore not yet a demonstrated redesign, and retention or loading is not yet later behavioral dependence.

The reported paths

System and architecture Evidence, installation, and later use Demonstrated redesign and unestablished targets Supplied machinery
Self-Harness — proposal selection Held-in failures become structured signatures; bounded candidates face a fixed two-split pass-count gate; accepted edits are merged into the next harness and exercised in later evaluations. Closed for accepted edits. Across three model runs, accepted changes include instruction, runtime-control, artifact-handling, and tool-handling edits. In the Qwen run specifically, sub-agent and skill branches were discarded while tool-error middleware was retained. Model, DeepAgent control architecture, declared edit surface, failure-signature representation, objective and evaluator, task subset and splits, and acceptance rule.
Continual Harness — direct update A Refiner reads recent trajectory windows and applies CRUD changes available at the next game step. Prompt changes and update-to-later-invocation edges for repaired skills close the path; many authored skills never run and memory consultation is sparse, so closure is artifact-specific. Closed for prompt changes and exercised skill repairs. Later use is also measured in aggregate for some sub-agents and inherited artifacts, but not for every edit. The mixed human–agent GPP precursor is not evidence for the automated Refiner path. Game representation and action basis, four-part harness partition, meta-tools and interfaces, Refiner rule and schedule, reward design, model and teacher choices, milestones, and task family.
Autogenesis — proposal selection Typed Reflect, Select, Improve, Evaluate, and Commit operators conditionally install versioned resources consumed in later rounds or tasks. Closed for committed prompt and agent revisions reported as reused. Agent-Evo reports revisions to prompts, tool implementations, and agent code. Participant replacement is an architectural affordance rather than a demonstrated transition; Environment and Memory lack independent evaluation. Solution-Evo changes task output rather than the assessed system's organization. Five-resource ontology, variable-lifting and learnability-mask policy, bus and specialist organization, trace schema, objective, evaluator and acceptance rule, model backbones, task interfaces, and round budget.
Accumulated Behavioral Rules — proposal selection An engineer generalizes accepted review feedback into a standing rule file loaded by two agent interfaces. Installation and loading are established. Behavioral dependence is unestablished despite zero reported recurrence over 74 later exposures, because there is no no-rule or static-rule baseline and no per-rule isolation. The demonstrated redesign is the installed rule layer and its contents, not a causally isolated behavioral effect of any rule. Human generalization and abstraction judgment, rule representation and loading scheme, source taxonomy, and review process. Retirement is described inconsistently and remains unresolved.
Darwin Gödel Machine — proposal selection A fixed o1 diagnostician frames an edit; a parent changes descendant code; a viability criterion admits the child; benchmark score later biases parent sampling. The path is closed only for admitted descendants later sampled and run, not for archive admission alone. Candidate generation plus an FM ranker is a demonstrated installed architectural change. The record does not isolate later-descendant dependence on that exact redesign. o1 diagnostician, viability criterion, archive policy, parent-selection rule and hyperparameters, benchmark objective and evaluator, model weights, iteration budget, task distribution, and sandbox.
HyperAgents — proposal selection A hyperagent patches its unified task/meta-agent program; evaluators score the descendant; selected valid lineages are replayed before later modification and evaluation. Closed for selected executable descendants. Transferred hyperagents whose meta-agent components were held fixed later generated better agents on unseen math grading, supporting one cross-domain compounding contribution at the whole-hyperagent level. Task logic and the procedure that generates later agents are demonstrated targets. The appendix also makes parent selection operative and editable, but its gain over random is not significant and it remains below the handcrafted selector. Feature-level attribution and sustained compounding are unestablished. Main-run task distribution, objectives and evaluators, parent-selection rule, archive and phase controller, model and tool dependencies, resource budget, and sandbox. The appendix moves parent selection inside the editable program, but not evaluation or the outer archive loop.

What the comparison establishes

First, editability, installation, and dependence are different thresholds. Self-Harness merges accepted edits and later evaluates the resulting harness. Continual Harness traces some repairs into later invocations while also creating a long unused tail. Autogenesis reports reuse for only part of its typed surface. Accumulated Behavioral Rules always loads the rule file without isolating its effect. The Darwin Gödel Machine closes the later-use edge only for descendants sampled again. HyperAgents goes further: selected lineages revise both task and meta-agent code, including embedded prompt text, and transferred hyperagents later execute that retained improvement procedure in a new domain. Some successful descendants also install natural-language files of synthesized insights and plans, but selection promotes the bundled executable lineage and the paper does not isolate those artifacts' contribution.

Second, update architecture does not determine redesign breadth. Continual Harness updates directly; the other five expose rejectable candidates. Both architectures can alter consequential organization while leaving objectives, representations, interfaces, evaluators, or controllers supplied. The operative question is which proposed change entered a behavioral path and was later used, not whether the pathway carries a direct or proposal-selection label.

Third, the supplied boundary is a placement, not a verdict that every fixed choice is defective. A fixed evaluator can be protective, a fixed controller can make search affordable, and a fixed interface can make an experiment interpretable. HyperAgents makes the distinction especially clear: moving the agent-generation procedure inside the editable program materially broadens redesign, while the main experiments still supply parent selection, evaluation, and the outer archive process. The evidence establishes where the reported path stopped, not that every supplied choice was wrong or unavailable to the research team.

Lifecycle evidence is uneven too. Self-Harness reports promotion but no criterion-driven retirement path. Continual Harness permits deletion and demotion but also accumulates unused artifacts. Autogenesis supplies version lineage and rollback, which makes changes reversible without retiring them. Accumulated Behavioral Rules supports refinement while describing removal inconsistently. The Darwin Gödel Machine and HyperAgents retain archive lineages by design because apparently weak descendants can be useful stepping stones.

Scope

  • These are readings of six published descriptions through their KB ingests, not independent reproductions. Autogenesis and HyperAgents also have separate code-grounded reviews (Autogenesis, HyperAgents); the Darwin Gödel Machine implementation has not been inspected here. The HyperAgents paper and code review agree on executable lineage, while the checked-in initial meta-agent does not reliably receive prior evaluation history through its prompt.
  • The sample is a narrow research neighbourhood: recent systems that mostly revise readable artifacts around fixed model backbones. It does not estimate how common these placements are across self-improving systems.
  • Closing a causal path establishes operative redesign over the declared horizon. It does not establish net improvement, continuity of the revision path, or a contribution to compounding.

Relevant Notes: