Self-improving systems
Type: kb/types/tag-readme.md
Selective head: the membership definition, the update architectures, and the four-part pathway profile, with the load-bearing note per claim. Full tag membership comes from the by-tag sweep (kb/reference/navigation.md).
Membership
A self-improving system makes operative, evidence-responsive changes to its own behavior-determining organization. Read every attribution against a declared frame of boundary, horizon, and objective: including maintainers can make a development system a human-inclusive member, and self-improvement is relative to a declared objective — indexed in the attribution, antecedent in the pathway. When the objective itself changes, the change is improvement only against a level outside it. Below the objective sits the target level, where a structural property is pursued because it is held to serve the objective and is checked for achievement rather than for warrant — the profile dimensions below are read that way whenever they are treated as goals. Membership settles only the category.
Update architecture
Evidence may directly determine an update that is always adopted, as gradients and viability triggers do, or flow through the proposal-selection subtype, where search, reject-capable evaluation, and operative retention let a candidate be rejected first — and where false-positive acceptance becomes operative. A pathway may compose both.
Pathway profile
After membership and update architecture, profile the pathway across four parts rather than placing it on a ladder. The profile is descriptive and selects no order by itself; any comparison between pathways is indexed to a declared objective. (Not the text-contract sense of "profile", which is a normative bundle a collection adopts.)
- Reflective structure — coverage of represented aspects and forms, plus the separate addressability profile over retained commitments.
- Improvement dynamics — cumulativity: later dependence through the retained result; accumulation versus compounding, with compounding tested in later improvement.
- Governance — what the methodology settles, and which of those decisions are warranted.
- Actor allocation — human, computational, or joint per function; allocation carries the comparison, and computational closure is its no-human endpoint, not a grade of reflectivity.
The four do not determine, subsume, or form a monotone progression through one another: each property's note states the entailments it does not license, and the placements below show which combinations actually occur. That is non-entailment, not full independence — coverage is structurally required for some addressability operations, and explicit criteria are what carry a decision to a computational actor.
What reflection adds
A self-improving pathway is reflective when it routes objective-bearing evidence into a change to the system's behavior-determining organization through a causally connected self-representation; later operation must depend on that change. This causal structure permits direct updates, proposal selection, and compositions of both. Authority family — evidence, advice, instruction, enforcement — does not decide reflection.
- Reflection buys addressability — retention later rounds can read, criticize, and selectively revise.
- Repeatable operative revision — complete addressability covers governing machinery; continuity keeps its revision path usable.
- Reflection makes retained lessons second-order — an addressable lesson can reject or rescope a represented prior commitment.
- Retrieval failure is reflection failure — the standing discount on a lesson that never surfaces.
- Payoff hypotheses, still open: theory-mediated sample efficiency, and selective revision needing a faithful rationale.
Governance and computational allocation
- Methodological and computational closure track different changes — settled method can be human-executed; an unattended model can improvise.
- Computationally directed self-improvement is a fixed-boundary reallocation ending in contraction — the transition worth studying is intra-category, and its endpoint is whether the boundary can be contracted to exclude the humans.
- Increasing computational autonomy relocates human effort to the frontier — measure improvements per human judgment, not hours.
- Only explicit retention is durable, writable, and addressable — no tacit channel carries settled methodology.
Placements
Evidence: Commonplace paths show broad addressability; completeness remains open. Six external paths map supplied machinery; thirteen cases map profile combinations. Payoff remains untested.
Base vocabulary and boundary cases
- Behavior-determining organization, operative change, evidence bearing on an improvement objective — the definition's three base terms. All three rest on behavioral authority: consumer, channel, force.
- The definition classifies its boundary cases without ad hoc exceptions — ten cases, from gradient learning to accidental self-modification.
- Measuring autonomy well enough to see it improve is an open problem — a per-function profile locates a system but cannot yet show it becoming more autonomous.
Related Tags
- foundations — the broader core theory this sits inside
- constraining — methodological closure is a constraining property of methodology-as-input
- computational-model — reflection and intercession as computational concepts generalized to socio-technical boundaries
Other tagged notes
- A consumption channel delivers force without the history that earned it - A consumption path can promote content into a higher-force role without checking whether an authorization covers that content, version, and use
- An omitted improvement-loop function and a frozen one need different repairs - Five proposal-selection systems expose frozen functions, while a direct-update contrast shows why absence of a gate is not omission; HyperAgents supplies a preliminary partial unfreezing
- Formal symbolic systems assess explanatory-reach only through causal and proof obligations - Formal symbolic systems assess explanatory-reach only after claimed generality is translated into causal or proof obligations inside a warranted model
- Gödel machines are a proof-governed case of reflective self-modification - The Gödel machine realizes reflective self-modification with a proof-gated acceptance rule, gaining model-relative rigor at the cost of excluding useful changes it cannot prove
- Improving an agentic system crosses the natural-language/symbolic boundary - The error-correction asymmetry sorts agentic behavior between natural-language and code, so reliability-improving changes cross the boundary; reflective coverage of one form cannot carry them
- Machinery persists by warrant, not position, in a reflective loop - Sutton's build-mode assumes a meta-method outside the learned system, exempt from selection by position. A reflective loop has no outside: machinery is artifacts in loop scope, the boundary moves per artifact, and persistence must be earned
- Reach-assessment - Definition — judging whether a commitment's claimed explanatory-reach is genuine across natural-language, symbolic, and distributed-parametric forms
- Stale self-description conceals its own staleness - What artifact drift adds when it is reflexive: the process that would detect it consults the artifact that drifted, the trigger has no edit event to hook, and synchronization load scales with autonomy
- Weakly discriminated qualities tend to be underselected - Conjecture separating available model capability from selection: qualities weakly distinguished by the actual acceptance oracle lose to strongly verified objectives
- World models assess explanatory-reach through action-conditioned prediction - Learned world models can assess explanatory-reach when action-conditioned predictions are tested across the interventions or shifts a commitment claims