Case packet
Neutral case identifier: case-04324ff043d41a
The possible directed relationship from Artifact A to Artifact B is under review.
Artifact A
Methodological and computational closure track different changes
An improvement pathway can stop depending on improvised judgment without stopping its dependence on a human actor, and it can stop depending on a human actor while continuing to improvise. Those are different architectural changes and need different readings of closure.
Methodological closure asks whether the retained methodology settles the consequential decisions that the pathway raises. A method is less closed where it merely says “use judgment,” names an approver, or leaves a meta-decision to be reconstructed from scratch.
Computational closure asks who supplies the decision. A function is computationally closed when its execution needs no human decision; a whole pathway is computationally closed only when every required function meets that condition.
Computational closure and machine autonomy therefore read the same actor allocation: human, computational, or joint for each pathway function. “More computationally autonomous” describes movement in that allocation; “more computationally closed” describes the resulting reduction in functions that still require a human decision.
Neither reading is the cybernetic sense. Organizational closure — the recursive regeneration of a network of component interactions in the autopoiesis tradition — is a different property, already excluded from this cluster's vocabulary in the [reflective-system exclusions]; nothing here asserts or requires it.
Human-inclusive boundaries make allocation load-bearing
A [reflective system] may include established human processes. Put a maintainer with a standing causal role inside the boundary of a maintained system with readable source, and reflective attribution becomes cheap: the maintainer inspects the source as a representation, edits it, and the build carries the edit into operation. The attribution can be true while saying little about machine performance.
Actor allocation restores the missing discrimination. Under a fixed human-inclusive boundary, report each consequential function as human, computational, or joint; computational closure is the no-human endpoint of that profile. Do not replace the profile with a percentage: functions differ in decomposition, authority, and stakes, and cross-system comparison remains [an open measurement problem].
The form is inherited rather than invented. [Parasuraman, Sheridan, and Wickens] report automation per function — information acquisition, analysis, decision and action selection, action implementation — and hold that an allocation is judged by its performance consequences, its reliability, and the cost of the consequences it admits, not by how much of the work the machine has taken over. That shape is what carries across, with three departures. The functions allocated here are the improvement pathway's own — search, evaluation, and retention where the pathway is proposal-selection — rather than task-performance stages. Their within-function ten-level scale is not inherited: the paper's validation is strongest for decision selection, and a graded level per function would reintroduce the percentage this profile refuses. And allocation still establishes nothing about warrant.
Four concrete combinations
| Improvement decision | Methodologically closed? | Computationally closed? | Why |
|---|---|---|---|
| A maintainer manually applies an exact checklist before accepting a patch | Yes | No | The criterion is settled, but a human supplies the verdict. |
| A validator accepts an artifact only when an exact structural predicate holds | Yes | Yes | The criterion and its execution are both explicit and computational. |
| An unattended coding agent is told to inspect failures and “improve the repository” using its own judgment | No | Yes | No human intervenes, but consequential choices remain improvised. |
| A maintainer and agent jointly judge a theory note against “is this good?” | No | No | The criterion is unsettled and a human participates in the verdict. |
Stable but tacit expertise does not count as retained methodology. A maintainer may apply a repeatable internal criterion that was never externalized — settled in practice, unsettled in representation — but methodological closure reads the representation, and the reading has one ground rather than a human-specific rule, [since only explicit retention is currently durable, writable, and addressable at once]: a criterion that cannot be retrieved, cited, criticized, or selectively revised is not available to the pathway as methodology, however consistently it is applied — it is available only as the human actor. The state deserves its own name instead of a closure grade: stable-but-unexternalized practice is a promotion candidate, noticeable by recurrence and convertible by externalization. The last row therefore stands even when the joint judgment is secretly consistent.
The third row needs a named exclusion, not a stronger definition. Computational closure reads actor allocation within the declared frame: a hosted model is a computational actor wherever it runs, so a pathway can be computationally closed while depending on inference infrastructure and a provider outside the selected subsystem. That dependency is real, but it is a boundary and coverage fact — in profile terms, selection-grade coverage of a sealed parametric component, [as reflective coverage is graded across representational forms] — not an actor fact. Widening closure to swallow substrate dependency would leave almost no model-mediated function ever computationally closed and destroy the discrimination the table exists to provide, the same reason the organizational-closure sense is excluded above.
When the two changes advance together
A recurring human decision becomes easier to allocate computationally after its inputs, criterion, and failure response have been made explicit. The conversion usually has three parts:
- Representation — the relevant inputs and commitments become available to the deciding process, [since reflection buys addressability].
- Settlement — the methodology supplies the criterion or determines the result instead of merely naming a decider, [since a methodology governs its own extension only as far as it settles the meta-decisions it raises].
- Warranted execution — a computational procedure or oracle implements the criterion with evidence adequate to the case, [since warranted autonomy is bounded by oracle domain].
The order is forced, not conventional: externalization is allocation's transport. A computational actor can receive a criterion only through an explicit representation — under a selection-only parametric profile nothing else inside the boundary is both writable and durable, and even where fine-tuning adds a write channel the transfer is unaddressable, escaping governance at the moment it succeeds ([only explicit retention is currently durable, writable, and addressable at once]).
These are engineering dependencies, not definitions of one another. A settled gate can remain human-executed; an agent can read explicit commitments yet improvise how to apply them; and a computational procedure can encode a poor proxy. Moving evaluation to a model changes allocation without establishing that its acceptances are trustworthy.
The [Commonplace reference case] applies this conversion to ADR 026 and keeps the trace-specific facts in one place.
Reflection is a separate question
Reflectivity does not require methodological closure. It requires a causally connected representation of the system's own behavior that processes inside the declared frame can read and change. A reflective pathway may expose its rules for criticism while leaving the next revision to open-ended judgment. Conversely, a fixed pipeline may settle every operational choice without representing or revising itself.
The properties reinforce each other when the represented object is the improvement methodology itself: an addressable criterion can be revised, then a settled and warranted version can be executed computationally. That is a trajectory through a [multi-part profile], not one scale of reflectivity or closure.
Scope
- Both closure readings are per decision and per pathway, so mixed profiles are normal: exact validators can coexist with joint review, and settled acceptance rules with improvised objective-setting.
- A loop instance completes when search, evaluation, and operative retention occur. Calling that event closure would conflate completion with architecture.
- Both readings require a declared frame. A whole-system closure claim without named decisions and pathways hides the mixed architecture.
- Comparing allocation profiles across releases or systems inherits the open commensurability problem: [measuring autonomy well enough to see it improve is an open problem].
Open Questions
- When an initial human instruction makes a downstream agent-performed function joint rather than computational; counting every instruction hides agent performance, while counting none hides decision content supplied up front.
- Whether objective-setting can become methodologically closed without freezing the improvement objective rather than improving it.
- How much representational explicitness computational internalization requires when learned components can execute a decision without exposing its criterion.
- How to distinguish a computational implementation of a settled method from a proxy that silently changes what the method decides.
Relevant Notes:
Artifact B
A proposal-selection improvement loop requires search, evaluation, and operative retention
A proposal-selection improvement loop is the architecture of improvement in which candidate changes are generated, evaluated with a possibility of non-adoption, and selectively made operative. It is a named subtype, not the whole of the phenomenon: a [self-improving system] needs its changes to be responsive to evidence bearing on an improvement objective, and evidence may instead determine an update directly — gradient-, reward-, error-, or viability-driven — with no candidate ever standing to be rejected. What follows is the anatomy of the subtype, and it applies with full force exactly there.
A proposal-selection loop requires three functions: search brings a candidate change into consideration, evaluation supplies grounds for accepting or rejecting it, and operative retention preserves an accepted change with behavioral authority. Remove any one and the loop does not close — a change nobody proposed, nobody could reject, or nobody will ever act on.
The loop is therefore narrower than self-modification. A blind, accidental, or unconditional rewrite may change later behavior without applying any criterion; a transient rewrite may fail to preserve the result. Both can count as self-modification, but neither closes a proposal-selection loop. Conversely, the three functions can close the loop in a system that is not reflective at all.
A terminology note: the concept descends from Ashby's adaptation — his ultrastable system, examined below, is the conceptual ancestor even though it classifies outside the subtype — but it is named for what the loop aims at rather than by his word for it. Everyday adaptation is transient compensation, an eye adjusting to the dark, and retains nothing; retention is one of the three requirements. Where this note says adaptation or adaptive, it means Ashby's phenomenon. The architecture described here is named proposal-selection throughout.
A [reflective system] supplies one possible causal path into this loop. Through intercession — an operation that changes the system through its causally connected self-representation — it can modify a represented aspect of itself. Making that path available does not itself provide search, evaluation, or retention.
The independence runs both ways. A directly determined update can land on a self-representation as readily as on an opaque substrate — evidence can revise an explicit policy or a recorded lesson with nothing rejectable anywhere in the path — so neither architecture is the general form of reflective improvement.
Search determines what enters consideration
Search brings an unrealized change under consideration. It may include:
- detecting a problem, opportunity, or adaptation signal;
- selecting the aspect and operation to change;
- generating one or more candidates;
- allocating effort and deciding when to stop or escalate.
At minimum, search must produce a candidate from a space in which other possible changes remain unrealized. It need not compare several candidates at once or operate autonomously. A maintainer may choose the problem, a model may draft a candidate, and a script may enumerate alternatives within one declared socio-technical loop. Assigning those functions establishes the loop's boundary; it does not make the loop reflective.
Search range and evaluation strength are independent limits:
Evaluation cannot select a candidate that search never reaches.
A strong verifier can improve judgments within a narrow generator's range, but it cannot expand that range. [Automating KB learning is an open problem] gives one concrete search space—extract, split, synthesize, relink, regroup, reformulate, retire—whose judgment-heavy parts remain substantially human-driven.
Evaluation determines which changes may remain operative
Evaluation applies criteria to a proposed or already actualized change. Its result must be able to affect selection, rollback, or continued retention. Evaluation is non-vacuous only if some possible result permits rejection: an unconditional trigger is not an evaluator merely because it precedes a transition, and a conditional trigger whose only effect is to launch the next variation is not one either. The verdict must control an operation distinct from producing the next candidate — select, discard, block, roll back — so that rejecting a change and merely changing again are different events in the mechanism.
Oracle is shorthand for the component or procedure that supplies the evidence or judgment. It may be a proof system, test, validator, empirical measurement, rubric, model evaluator, human review, or some combination. The [oracle-strength spectrum] grades these mechanisms, while [the boundary of automation is the boundary of verification] explains why constructing an adequate oracle is often harder than generating candidates.
Any judgment remains scoped to what the check establishes. An oracle may accept a candidate under specified criteria without establishing that the change is globally beneficial. Search and evaluation may be performed by the same person or process, but they fail in different ways and improve by different means. They are analytically separable rather than independent: automating one changes the load on the other.
Operative retention makes the change consequential
Acceptance alone does not make a change consequential. Operative retention combines persistence with an authority path through which the retained result can affect later behavior. In [behavioral authority] terms, the change needs a consumer, a channel, and a force.
- A reviewed note that no future reader or prompt-assembly step loads has no consumer.
- An approved patch that is never merged has no channel.
- A generated validator that no command invokes has no force.
In each case, search ran and evaluation passed, but the proposal-selection loop remained open: the artifact exists without becoming behaviorally consequential.
Artifact labels do not decide whether retention is operative. A knowledge artifact consumed as evidence or advice can affect later behavior, while a nominal system-definition artifact with no consumer cannot. The test is the [behavioral-authority] path: consumer, channel, and force relative to the objective and declared horizon.
For self-improvement, the accepted change must reach the system's own [behavior-determining organization]. Promotion into instruction, enforcement, or configuration is one way to strengthen that path, and may itself run as another proposal-selection instance — [the two-layer execution system] develops that promotion architecture, with recurrence as the trigger, pre-promotion verification as the gate, and methodology growth plus a coverage-test update as retention — but it is not universally required for reflective or operative change.
Repetition does not establish cumulativity
A proposal-selection loop can repeat on a timer or fresh request without using anything retained by an earlier iteration. Whether later improvement consumes or preserves earlier improvement-relevant information is cumulativity, whose criterion and counterexamples belong to [the informational-dependence test on the retained result]. Retained rationale can provide that dependence when later search or evaluation actually consumes it; [design rationale management in Commonplace] documents that path.
Boundary cases clarify the claim
Cybernetician W. Ross Ashby's ultrastable system marks the subtype's edge from just outside it, and its exclusion follows from the evaluation criterion above, not from a missing component. The electromechanical Homeostat has exactly one evidence-responsive transition: when essential variables leave viable bounds, the parameters jump to new random values ([Ashby 1960, chapters 7–8]). That single jump both discards the incumbent configuration and produces its successor — rejection is not an operation distinct from generation — and a configuration that restores viability persists through equilibrium, with nothing whose function is to accept it. The functions collapse into one trigger, so under the definitions here the machine is a non-reflective, direct viability-driven [self-improving system], not an instance of this subtype.
What the Homeostat does admit is a functional variation–selection–retention reading: configurations vary, viability determines whether variation continues, and the survivor persists through non-displacement. That reading is an analyst's reconstruction, not architecture, and its value is to mark the floor of each function — search as a draw from a random-number table bearing no relation to the problem, evaluation as a one-bit viability boundary that ranks nothing, retention as equilibrium, a configuration surviving because nothing is left to displace it. Read this way, the Homeostat is the cheapest demonstration of what a stronger generator and a real oracle actually buy. Reflection is still not a premise of the decomposition. An evolutionary strategy supplies the genuine non-reflective instance: it runs an explicit generate-and-select loop over parameters nothing inside it can read.
The Homeostat's contrast with a gated system is architectural, not a difference of gate strength: a gradient learner has no evaluator either — [online gradient descent] adopts every step the revealed cost dictates, with no accept/reject anywhere (Zinkevich 2003) — and Zinkevich's Greedy Projection/GIGA result supplies the technical counterexample to treating an acceptance gate as universal. The Homeostat stands with it, on the excluded side of the boundary just drawn. [Gödel machines] sit inside the subtype at its formal extreme, a proof-mediated gate rather than none at all; that architecture is developed in their own note.
Reflection is a separate axis from this exclusion: the Homeostat is also non-reflective, and [what that costs is addressability, not category membership] — evidence-responsive operative change to the system's own organization, with or without a self-representation and with or without a gate, is what makes a [self-improving system].
What the decomposition claims
The three functions are analytically separable, not architecturally separate. One process may perform several of them — a maintainer who notices a problem, drafts the fix, and merges it performs all three — and evaluation may run before a candidate becomes operative or after. Co-location has a floor, though: the functions must remain causally distinguishable even when one process performs them — rejection, in particular, must be an event distinct from the arrival of the next candidate. Where they collapse into a single evidence-triggered transition, as in the Homeostat, the loop is not weakly present; it is absent, and the pathway is direct. The decomposition specifies what the loop must accomplish, not a sequence, a component diagram, or a division of labour. Its use is diagnostic: when a loop stalls, ask which of the three is missing rather than which component failed.
The status claimed here matches how the neighboring self-adaptive-systems field treats its own loop models: MAPE-K — introduced in [Kephart and Chess's autonomic-computing vision], which itself supplies no membership test — and its relatives are presented as reference models for engineering adaptation, not as the definition of it ([Weyns, Software Engineering of Self-Adaptive Systems]), and a systematic review of that literature finds no settled formal definition from which any single loop architecture would follow ([Petrovska, Erjiage, and Kugele 2025]). The proposal-selection decomposition is offered in the same spirit — a conceptual model of one architecture, with the [category membership question] settled elsewhere.
Open Questions
- Whether search range can be measured or bounded for a socio-technical loop in the way oracle strength can be graded.
- Whether a fallible evaluator can govern changes to its own acceptance criteria without either an external criterion or the axiomatization that buys formal closure.
Relevant Notes:
Under-review context phrase
supplies the pathway functions over which both properties are reported