A proposal-selection improvement loop requires search, evaluation, and operative retention

Type: kb/types/note.md · Tags: foundations, computational-model, self-improving-systems

A proposal-selection improvement loop is the architecture of improvement in which candidate changes are generated, evaluated with a possibility of non-adoption, and accepted changes are made operative. It is a named subtype, not the whole of the phenomenon: a self-improving system needs its changes to be responsive to evidence bearing on an improvement objective, and evidence may instead determine an update directly — gradient-, reward-, error-, or viability-driven — with no candidate ever standing to be rejected. What follows is the anatomy of the subtype, and it applies with full force exactly there.

A proposal-selection loop requires three functions: search brings a candidate change into consideration, evaluation supplies grounds for accepting or rejecting it, and operative retention preserves an accepted change with behavioral authority. Remove any one and the loop does not close — a change nobody proposed, nobody could reject, or nobody will ever act on.

The loop is therefore narrower than self-modification. A blind, accidental, or unconditional rewrite may change later behavior without applying any criterion; a transient rewrite may fail to preserve the result. Both can count as self-modification, but neither closes a proposal-selection loop. Conversely, the three functions can close the loop in a system that is not reflective at all.

A terminology note: the concept descends from Ashby's adaptation — his ultrastable system, examined below, is the conceptual ancestor even though it classifies outside the subtype — but it is named for what the loop aims at rather than by his word for it. Everyday adaptation is transient compensation, an eye adjusting to the dark, and retains nothing; retention is one of the three requirements. Where this note says adaptation or adaptive, it means Ashby's phenomenon. The architecture described here is named proposal-selection throughout.

A reflective system supplies one possible causal path into this loop. Through intercession — an operation that changes the system through its causally connected self-representation — it can modify a represented aspect of itself. Making that path available does not itself provide search, evaluation, or retention.

The independence runs both ways. A directly determined update can land on a self-representation as readily as on an opaque substrate — evidence can revise an explicit policy or a recorded lesson with nothing rejectable anywhere in the path — so neither architecture is the general form of reflective improvement.

Search determines what enters consideration

Search brings an unrealized change under consideration. It may include:

  • detecting a problem, opportunity, or adaptation signal;
  • selecting the aspect and operation to change;
  • generating one or more candidates;
  • allocating effort and deciding when to stop or escalate.

At minimum, search must produce a candidate from a space in which other possible changes remain unrealized. It need not compare several candidates at once or operate autonomously. A maintainer may choose the problem, a model may draft a candidate, and a script may enumerate alternatives within one declared socio-technical loop. Assigning those functions establishes the loop's boundary; it does not make the loop reflective.

Search range and evaluation strength are independent limits:

Evaluation cannot select a candidate that search never reaches.

A strong verifier can improve judgments within a narrow generator's range, but it cannot expand that range. Automating KB learning is an open problem gives one concrete search space—extract, split, synthesize, relink, regroup, reformulate, retire—whose judgment-heavy parts remain substantially human-driven.

Evaluation determines which changes may remain operative

Evaluation applies criteria to a proposed or already actualized change. Its result must be able to affect selection, rollback, or continued retention. Evaluation is non-vacuous only if some possible result permits rejection: an unconditional trigger is not an evaluator merely because it precedes a transition, and a conditional trigger whose only effect is to launch the next variation is not one either. The verdict must control an operation distinct from producing the next candidate — select, discard, block, roll back — so that rejecting a change and merely changing again are different events in the mechanism.

Oracle is shorthand for the component or procedure that supplies the evidence or judgment. It may be a proof system, test, validator, empirical measurement, rubric, model evaluator, human review, or some combination. The oracle-strength spectrum grades these mechanisms, while the boundary of automation is the boundary of verification explains why constructing an adequate oracle is often harder than generating candidates.

Any judgment remains scoped to what the check establishes. An oracle may accept a candidate under specified criteria without establishing that the change is globally beneficial. Search and evaluation may be performed by the same person or process, but they fail in different ways and improve by different means. They are analytically separable rather than independent: automating one changes the load on the other.

Operative retention makes the change consequential

Acceptance alone does not make a change consequential. Operative retention combines persistence with an authority path through which the retained result can affect later behavior. In behavioral authority terms, the change needs a consumer, a channel, and a force.

  • A reviewed note that no future reader or prompt-assembly step loads has no consumer.
  • An approved patch that is never merged has no channel.
  • A generated validator that no command invokes has no force.

In each case, search ran and evaluation passed, but the proposal-selection loop remained open: the artifact exists without becoming behaviorally consequential.

Artifact labels do not decide whether retention is operative. A knowledge artifact consumed as evidence or advice can affect later behavior, while a nominal system-definition artifact with no consumer cannot. The test is the behavioral-authority path: consumer, channel, and force relative to the objective and declared horizon.

For self-improvement, the accepted change must reach the system's own behavior-determining organization. Promotion into instruction, enforcement, or configuration is one way to strengthen that path, and may itself run as another proposal-selection instance — the two-layer execution system develops that promotion architecture, with recurrence as the trigger, pre-promotion verification as the gate, and methodology growth plus a coverage-test update as retention — but it is not universally required for reflective or operative change.

Repetition does not establish cumulativity

A proposal-selection loop can repeat on a timer or fresh request without using anything retained by an earlier iteration. Whether later improvement consumes or preserves earlier improvement-relevant information is cumulativity, whose criterion and counterexamples belong to the informational-dependence test on the retained result. Retained rationale can provide that dependence when later search or evaluation actually consumes it. Design rationale management in Commonplace documents the retention surfaces from which such a path could draw, not a demonstrated consumption path.

Boundary cases clarify the claim

Cybernetician W. Ross Ashby's ultrastable system marks the subtype's edge from just outside it, and its exclusion follows from the evaluation criterion above, not from a missing component. Ashby's primary account says an unsuccessful trial changes the way of behaving while a successful one is retained. In the electromechanical Homeostat, an out-of-limit current makes the uniselector move to new values while an in-limit current leaves it in place, and the next values have no special relation to the incumbent or the problem. On this note's analysis, that transition both discards the incumbent configuration and produces its successor — rejection is not an operation distinct from generation — and a configuration that restores viability persists through equilibrium, with nothing whose function is to accept it. The functions collapse into one trigger, so under the definitions here the machine is a non-reflective, direct viability-driven self-improving system, not an instance of this subtype.

What the Homeostat does admit is a functional variation–selection–retention reading: configurations vary, viability determines whether variation continues, and the survivor persists through non-displacement. That reading is an analyst's reconstruction, not architecture, and its value is to mark the floor of each function — search as a draw from a random-number table bearing no relation to the problem, evaluation as a one-bit viability boundary that ranks nothing, retention as equilibrium, a configuration surviving because nothing is left to displace it. Read this way, the Homeostat is the cheapest demonstration of what a stronger generator and a real oracle actually buy. Reflection is still not a premise of the decomposition. An evolutionary strategy supplies the genuine non-reflective instance: it runs an explicit generate-and-select loop over parameters nothing inside it can read.

The Homeostat's contrast with a gated system is architectural, not merely a difference of gate strength. It stands on the excluded side of the boundary just drawn. Online gradient descent supplies another direct-update boundary case: after receiving each cost function, Zinkevich's Greedy Projection computes the next vector directly; that update step contains no candidate-adoption decision. Gödel machines sit inside the subtype at its formal extreme, a proof-mediated gate rather than none at all; that architecture is developed in their own note.

Reflection is a separate axis from this exclusion: the Homeostat is also non-reflective, and what that costs is addressability, not category membership — evidence-responsive operative change to the system's own organization, with or without a self-representation and with or without a gate, is what makes a self-improving system.

What the decomposition claims

The three functions are analytically separable, not architecturally separate. One process may perform several of them — a maintainer who notices a problem, drafts the fix, and merges it performs all three — and evaluation may run before a candidate becomes operative or after. Co-location has a floor, though: the functions must remain causally distinguishable even when one process performs them — rejection, in particular, must be an event distinct from the arrival of the next candidate. Where they collapse into a single evidence-triggered transition, as in the Homeostat, the loop is not weakly present; it is absent, and the pathway is direct. The decomposition specifies what the loop must accomplish, not a sequence, a component diagram, or a division of labour. Its use is diagnostic: when a loop stalls, ask which of the three is missing rather than which component failed.

The neighboring self-adaptive-systems literature supplies a comparison, not this vocabulary. Kephart and Chess describe an autonomic manager coupled to a managed element, monitoring it and its environment and analyzing, planning, and executing in response: "an autonomic element will typically consist of one or more managed elements coupled with a single autonomic manager" (The Vision of Autonomic Computing, IEEE Computer p. 44, verbatim). Weyns calls MAPE-K a reference model for a managing system: "MAPE-K's power is its intuitive structure of the different functions that are involved in realising the feedback control loop in a self-adaptive system" (Software Engineering of Self-Adaptive Systems, Section 3.1, verbatim). A later systematic review reports that MAPE-K's component semantics remain underspecified — "a more specific semantics of these two components is still missing" (Defining Self-adaptive Systems, Section IV-A, verbatim) — and that the field still lacks unified definitions. These sources establish an engineering feedback-loop tradition, but they do not establish the search–evaluation–operative-retention decomposition or its reject-capable boundary.

The proposal-selection decomposition is therefore Commonplace's conceptual model of one architecture, with the category membership question settled elsewhere. It does not claim to define every form of adaptation or self-improvement, and its three functions are not presented as established field vocabulary.

Open Questions

  • Whether search range can be measured or bounded for a socio-technical loop in the way oracle strength can be graded.
  • Whether a fallible evaluator can govern changes to its own acceptance criteria without either an external criterion or the axiomatization that buys formal closure.

Relevant Notes: