Self-improving systems
Type: types/tag-readme.md
A self-improving system makes operative, evidence-responsive changes to its own behavior-determining organization, read against a declared frame of boundary, horizon, and improvement objective. Assign this tag to work on whether or how systems make such changes: their definition and classification, mechanisms, limits, or evidence of improvement. Ordinary task execution or automation is insufficient without that question. Tagging an analysis does not certify that the analyzed system improves. The child areas below cover the specific mechanisms and tests. Deploy-time-learning covers the post-release demand for change, including changes made by human maintainers; learning-theory covers learning more broadly, whether or not the learner changes its own organization.
Child areas
- theory-builder — the theory-builder definition and its companions: tentative and addressable theory, retained theory guiding search, theory fit and warrant, the Popperian precedents; the research program Commonplace runs
- software-factory — factories and houses: family-specific production machinery, factory development and learning, universality, task versus product families
- improvement-loop — the proposal-selection architecture: search control, what to learn, oracle accumulation, diagnostic richness, false-positive filtering, omitted versus frozen loop functions
- reflection — reflective systems: addressability, second-order lessons, graded coverage, retrieval misses, Gödel machines, and the Commonplace reflective-trace evidence
- warranted-autonomy — which decisions a computational actor is warranted to take over from humans, the oracle domain that bounds it, closure, and measuring autonomy
- continual-learning — how a deployed system keeps learning outside model weights: the readable-artifact loop, form coevolution, governing writes, and the Bitter Lesson defense of that bet
Definition and objective
- Operative change — a change counts when it affects later operation over the declared horizon through a behavioral-authority path
- Evidence bearing on an improvement objective — gradients, rewards, errors, viability signals, tests, judgments: what carries information about the criterion
- The definition classifies its boundary cases without ad hoc exceptions — ten cases run against the criteria; the stress falls on declaring the boundary, not on the definition
- Self-improvement is relative to a declared objective — the objective is a declared parameter beside boundary and horizon, indexed by the analyst and antecedent in the pathway
- Revising an improvement objective is licensed from outside it — objective change is improvement only against a level outside both objectives
- A proximate target is checked for achievement, not for warrant — between objective and oracle sits a target level whose linking claim no check in the loop tests
Accumulation and compounding
- Accumulation counts dependence through the retained result — cumulativity is later dependence on what was retained, not on the evidence it caused
- Improvements can accumulate without compounding — compounding needs an earlier benefit to counterfactually improve a later episode
- Compounding is tested in later improvement, not by the accepting metric — displaced productivity measures and causal traces, not the metric that accepted the change
Casebooks and records
- Real self-improving systems occupy combinations no single rung captures — thirteen placements from the Homeostat to Commonplace on the profile fields
- Six reported self-improvement paths expose bounded redesign surfaces — operative redesign separated from revision of governing machinery and from declared editability
- The declared Commonplace frame — the boundary under which Commonplace's own attributions are assessed
- Where change candidates come from in Commonplace — problem-noticing and candidate-drafting beyond a maintainer's judgment
- External systems read as whole systems: Exo (a rewritable executor over a protected substrate), Autogenesis (versioned mutation with rollback), Compound Engineering (a compounding claim separated into its parts), Agno AgentOS (a control plane, not a learner), and AI Agents in Depth (book-level convergence and divergence)
Related Tags
- deploy-time-learning — the phenomenon: deployment forces post-release change, historically the maintainers' work
- learning-theory — the parent area for how systems learn; the six children above are also its routes into self-revision
- computational-model — the execution substrate these systems run on
Other tagged notes
- A benchmark that holds the client fixed exports the least-warrantable decisions by design - A fixed-client benchmark measures worker capability; it leaves broader closure untested when the client supplies internal production decisions, while ordinary user requirements and acceptance may remain external
- A better-factory claim compares operative states under an antecedent assessment relation - The improvement claim's relata are predecessor and operative-successor states and its relation is declared before the development it judges; evaluator location is a separate declaration from the learner boundary
- A claim without external assessment carries three obligations - Without external assessment a claim needs its own contradiction-and-support rule, a comparison level for objective change, and a performance measure it does not grade itself, plus attribution when it asserts a cause
- A claim's warrant does not determine its fit in a working theory - Independent warrant and fit in a working theory answer different questions: a warranted claim may fit poorly, while apparent fit may be produced by an unwarranted or already-assumed claim
- A complete theory path does not establish improved capacity - Mediation, empirical contact, response to criticism, and recurrent mediation support different claims; none alone establishes improved capacity, and retained addressable theory is one realization
- A consumption channel delivers force without the history that earned it - A consumption path can promote content into a higher-force role without checking whether an authorization covers that content, version, and use
- A failure explanation becomes search control only when it changes a later branch decision - An explanation of a failed branch becomes operative search control only when its retention changes a later choice about scope, priority, probing, continuation, or abandonment
- A fixed-model house must retain missing procedures for theory use - With models pinned, newly acquired theory-use procedures must persist outside their parameters; existing general machinery may already supply them, while code can make specified steps cheaper and more reliable
- A hand-crafted bootstrap fits the Bitter Lesson only if learning can outgrow it - A hand-crafted starting state fits the Bitter Lesson only if scalable learning displaces the task- and family-specific production knowledge it supplies as claimed reach widens
- A method's ceiling bounds the method, not the transfer it already made - Separates envelope expansion, where a responsibility leaves the residual human work, from performance gains inside a fixed envelope, so a bounded method reaching its ceiling does not retract the transfer it already made
- A methodology governs its own extension only as far as it settles the meta-decisions it raises - A retained methodology governs the consequential extension decisions it supplies or imports; actor competence can carry the process further without making those choices settled by the method
- A proposal-selection improvement loop requires search, evaluation, and operative retention - A proposal-selection improvement loop — candidates generated, evaluated with possible non-adoption, and accepted changes made operative — requires search, reject-capable evaluation, and operative retention
- A repeatable operative path keeps a redesign class open to revision - Operationalizes repeatable operative revision for a named redesign class as a causal path through representation, evidence-bearing determination, admission, installation, dependence, and continuity
- A retained-theory intervention isolates one surface, not the whole program theory - An intervention on retained theory estimates that surface's causal contribution under matched conditions; influence, explanatory guidance, acquisition, and whole-system theory possession remain different claims
- A retrieval miss is a local reflective-path failure - A missed relevant artifact leaves its represented aspect inert for the affected task and discovery route, while other loading paths and reflective aspects can remain causally connected
- A search controller is tested by what it brings to stronger evaluation - A search controller should be evaluated by the branches and probes it routes into stronger evaluation, not by treating every provisional judgment as an acceptance claim
- A theory's prototype standing is its revision cost: external binding plus lost investment - A theory's prototype standing is its expected revision cost — external binding plus the investment a revision discards — so natural-language versus symbolic form determines neither component and acceptance status is a separate axis
- A vibe-noting trace shows persistence enables revision, not certification - Evidence from one Commonplace note history: persistence enabled later semantic development while review exposed omitted risks, attribution drift, a link error, and an unresolved authority boundary
- Addressable theory - Definition — an addressable theory is a theory formulated in language whose assumptions, scope, and parts can be inspected and revised individually; a graded structural property, separate from tentative status
- An addressable theory can coordinate heterogeneous factory development - Tentative natural-language project theory may provide an addressable way to coordinate heterogeneous factory development while search, testing, and backtracking construct and revise it
- An agentic substrate becomes a software factory through family-specific production machinery - Maps the bounded-call agentic substrate to Greenfield's software-factory ontology without calling every generic harness or generated program a factory
- An omitted improvement-loop function and a frozen one need different repairs - Five proposal-selection systems expose frozen functions, while a direct-update contrast shows why absence of a gate is not omission; HyperAgents supplies a preliminary partial unfreezing
- An open-domain theory builder becomes a software house when new domains require production-machinery changes - A persistent automated theory builder for external users becomes a software house when genuinely new domains require it to revise the software that performs theory production rather than only the theories produced
- An optimal long-run learning strategy invests in its own machinery - A machinery improvement is paid for once and reused by every later learning episode, so over a long horizon its return can exceed immediate learning; where it does, an optimal strategy diverts effort to the machinery.
- Automating KB learning is an open problem - The KB already learns through manual improvement; automating judgment-heavy mutations needs oracles for connections, groupings, and synthesis we cannot yet manufacture
- Backtracking keeps lightweight search control provisional - Backtracking preserves the provisional status of a heuristic branch choice by restoring an earlier usable state and redirecting search after contrary evidence
- Broad software demands create pressure for agentic factory development - Broad software demands make exhaustive predefinition of useful family-specific production machinery practically implausible, motivating agentic factory development without ruling out a fixed universal substrate in principle
- Causal and proof obligations are two formal routes to assessing explanatory-reach - Causal and proof obligations demonstrate two ways formal symbolic systems can assess explanatory-reach inside a warranted model
- Choosing what to learn requires both validity and learning-value gates - Separates two promotion checks for learning loops: whether a candidate is trustworthy enough to learn from, and whether learning it would improve the current system.
- Citing retained theory at the decision point is a mediation trace - A decision record that cites the theory it followed supplies cheap, checkable evidence that the theory was consumed — necessary for a record-based mediation claim, but short of showing correct or load-bearing use
- Commonplace as a reflective self-improving system - Commonplace witnesses that a human-inclusive KB can be reflectively self-improving on one pathway despite uneven coverage and human-gated design judgment
- Commonplace builds a theory builder and tests whether it learns - Commonplace builds a system meeting the theory-builder conditions and tests whether it learns; criticism of content is compared against trial and error, and fine-grained addressability and high persistence against builders with less of each
- Competing causal theories can guide distinguishing experiments - Why observationally equivalent mechanisms can tell a theory builder what to test next: a noisy binary example separates evidence acquisition from choosing or verifying an explanation.
- Computationally directed self-improvement is a fixed-boundary reallocation ending in contraction - The progress question for self-improving systems is not category membership but which decision-bearing functions humans still supply; the endpoint test is whether the boundary can be contracted to exclude them
- Constraining during deployment is continuous learning - Continuous learning can happen outside of weights; constraining is one symbolic-artifact form where prompts, schemas, tools, and tests accumulate durable adaptive capacity during deployment
- Continual learning requires governing behaviour-changing writes, not just storing content - For deployed systems, persistence is insufficient; continual learning must select, validate, authorize, and coordinate behaviour-changing updates across the representational forms a system can change
- Cost-sensitive formalisms for tentative theory search - Exploratory map of backtracking, learning, and complexity models that expose budgets relevant to search guided by tentative theories
- Diagnostic richness constrains outer-loop learning quality - Outer-loop learning depends on inspectable failure evidence, not only on the oracle used to select winning candidates
- Discarding all experience-dependent state prevents cross-run accumulation - Discarding an intermediate artifact loses that artifact's reuse path, not all learning; cross-run accumulation fails only when no experience-dependent state survives to affect later work
- Disconnected witnesses do not establish a full causal path through theory - Evidence of recurrent learning through theory must identify the joins of one causal path; connected use and criticism still need separate evidence of improved capacity
- Distinct residue classes require distinct functions in a self-improving architecture - Different reasons for an untransferred decision identify different missing functions; a single process can supply several, and the current carrier split is not a permanent requirement
- Evaluation automation is phase-gated by comprehension - Optimization loops need diagnostic error analysis and demonstrated judge discrimination before automation can improve behavior rather than just score
- Explicit retention provides direct targets for selective revision - Explicit artifacts give a learner direct targets for inspecting and revising commitments; durability, writability, and effective addressability still depend on the boundary and available operations
- Factory construction is not evidence of production-knowledge acquisition - Recursive software-factory construction is prior art, but the demonstrated constructors receive the family definitions, metamodels, mappings, and expertise that determine the produced factory
- Factory development - Definition — factory development constructs or revises reusable family-level production machinery rather than one product's lifecycle state
- Factory learning is experience-responsive retention that improves the factory - Experience-responsive retention: production experience determines a retained change to reusable family machinery that later production depends on; factory-level learning is retention that improves the factory relative to a declared objective
- Factory-learning mechanisms should be compared on the same causal job - Compares factory-learning mechanisms on their shared causal job — experience-responsive retention — while separating update mechanisms from the project-theory function needed for open-ended coherent modification
- False-positive generation is filtered; false-positive acceptance becomes operative - False-positive generation faces evaluation before retention, while false-positive acceptance becomes operative and can compound
- Gödel machines are a proof-governed case of reflective self-modification - A Gödel machine admits self-rewrites through proof under its current formalization; this restricts admission without establishing how many useful changes are reachable or how reliably they are found
- Holding a program theory means sustaining coherent search under delayed feedback - Holding a program's theory is tested by whether a partial, tentative account of what the program is for keeps modification search, backtracking, and recovery coherent until delayed evidence arrives, not by whether the first change is right
- Improvements outside the admitted formal language need a pre-formal stage somewhere - An improvement whose concepts have no expression in a loop's admitted formal language is reached only through a pre-formal stage, inside the loop or fixed at design time in the choice of language; translation relocates that stage
- Increasing computational autonomy relocates human effort to the frontier instead of reducing it - In an open-ended system, increasing computational autonomy need not cut total human hours — attention moves to the frontier — so measure improvements per human judgment, not human time
- Instantiation alone cannot model agent learning across sessions - The class/instance analogy captures session startup but omits the retained update relation that can revise later agent definitions and reusable-content placement
- Learning inside a fixed decomposition inherits its mistakes - Why optimization cannot repair consequential distinctions, responses, or mappings outside the effective update space of a fixed task decomposition
- Lightweight search control allocates further search without licensing adoption - A search judgment is lightweight when its authority stops at allocating further investigation, probing, continuation, suspension, or abandonment rather than licensing an operative change
- LLM-executed methodologies are metacircular interpreters, not compilers - Self-hosting LLM methodologies are closer to metacircular interpreters than compilers: agents re-interpret natural-language rules each session, while stable paths codify into validators and commands
- Localized retention pays when sparse changes have bounded impact in a matching decomposition - Addressable retention localizes a sparse change when units match its decomposition; total adaptation stays local only when the affected units also have a small, explicit impact closure
- Machinery persists by warrant, not position, in a reflective loop - Reflection makes selected production machinery challengeable, but placement alone neither warrants nor requires revision; fixed general machinery may persist when its role and scope are earned
- Measuring autonomy well enough to see it improve is an open problem - Autonomy is reported per function rather than scored as a percentage, but that profile does not yet support comparison across systems or time
- Mechanistic constraints make Popperian KB recommendations actionable - Bounded context and underspecification don't just permit conjecture-and-refutation — they require it; derives three concrete practices (falsifier blocks, contradiction-first connection, rejected-interpretation capture) from KB mechanics.
- Methodological and computational closure track different changes - Methodological closure tracks what a retained method settles; computational closure tracks the absence of human decisions during the assessed operation, without requiring every judgment to have explicit criteria
- Methodology with incomplete coverage and its live theory fallback form a two-layer execution system - In open or incompletely covered domains, the theory-derived fast path and live theory fallback co-execute while methodology-native content follows a separate maintenance regime
- Missing rationale does not exclude a theory builder; weight-only retention excludes one across runs - The August 2026 Prime Agent, Recuris, and Apodex reports: retained rules without a recorded rationale leave theory-builder membership open, while Apodex carries only weights across runs, where no unit says anything, so no builder spans its runs
- Moving the interpretation–enforcement boundary requires cross-form coverage - Moving responsibility between model-interpreted rules and formal enforcement crosses natural-language and symbolic forms, so governing the transfer requires coverage of both and their mapping
- Natural-language project state may specialize weight-resident search heuristics - The natural-language part of project state may specialize general search heuristics already represented in an LLM's weights by supplying current intent, theory, branch history, and constraints
- Naur's human-only conclusion needs more than the absence of explicit criteria - Naur's human-only conclusion needs a further premise connecting unformulated judgment to computational inability; this reading preserves his functional tests without claiming that learned criteria are inexpressible
- Open-ended construction builds an object and a theory of it - Proposes that construction which must discover and revise an object's organization produces project-specific understanding beyond the object, using programs and theories as its two main cases
- Open-ended improvement must allocate search before decisive evaluation is available - Open-ended improvement must choose which questions, candidates, experiments, or proof paths to develop before decisive evidence about them is available; even a Gödel machine's proof gate retains this prior search problem
- Open-ended theory learning and factory learning close the same reflective loop - In Commonplace's arrangement, theory learning and software-factory learning require one connected reflective path; proof-governed switching alone does not settle criticism
- Oracle accumulation improves selection for later candidates in its maintained domain - A failure retained as a lesson helps tasks that retrieve it; retained as a maintained check it improves selection for later candidates in its domain and amortizes validation
- Preferential codification concentrates less predictable work at the agent boundary - Explains the negative-selection mechanism by which preferential codification changes the composition of work retained at an agent boundary
- Project-theory possession requires comparing new demands with existing organization - For open-ended modification, project-theory possession includes relating a new demand to existing responsibilities before parallel structure becomes the default; an explicit assimilation branch may counter additive coding-agent patches
- Reach-assessment - Definition — judging whether a commitment's claimed explanatory-reach is genuine across natural-language, symbolic, and distributed-parametric forms
- Reflection buys addressability - Self-improvement can accumulate without reflection — parametric learners do — but non-reflective retention gives only indirect handles; reflective retention makes the changed object addressable
- Reflection makes retained lessons second-order: a lesson can reject or rescope a prior commitment - Reflection lets a retained lesson target a prior commitment explicitly — rejecting, revising, or rescoping it — while non-reflective correction acts indirectly through the substrate
- Reflective coverage is graded across representational forms - Reflective coverage is stated per represented form and operation profile; control of an external dependency does not make that dependency part of the system's reflective coverage
- Reflective system - Definition — a system is reflective relative to selected aspects when an internal process uses a causally connected self-representation of them in its operation
- Retained system-definition artifacts enable persistent deployment-time adaptation - Retaining evaluated changes to behavior-shaping prompts, rules, tools, and tests gives deployed systems a persistent adaptation path outside model-weight updates
- Retained theories may improve sample efficiency under structured shifts - Conjecture: retained theories may reduce target observations under structured shifts; a useful theory's reuse benefit is separate from selecting it by estimated explanatory-reach
- Revision guided by rationale needs faithfulness, not just legibility - When revision of an addressable theory relies on rationale to locate a failed premise, misleading rationale can direct repair to the wrong part; rationale is one optional repair aid
- Six Commonplace paths establish broad addressability, not completeness - A six-path Commonplace audit establishes broad path-relative addressability without establishing completeness, while exposing separate admission and model-realization gaps in the broader revision affordance
- Software factory - Definition — in the Greenfield lineage, a software factory is a configured family-specific software-production environment
- Software house - Definition — a software house is the complete persistent system responsible for developing and evolving software for external users
- Stale self-description conceals its own staleness - What artifact drift adds when it is reflexive: the process that would detect it consults the artifact that drifted, the trigger has no edit event to hook, and synchronization load scales with autonomy
- System use is an initial selection environment when theory fit lacks a fixed oracle - When no complete fixed oracle decides whether a claim belongs in a working theory, distributed consequences of live system use can provide an initial selection environment
- System use provides evidence of theory fit and causal usefulness, not independent warrant - Consequences of using a claim in a live system can test its integration and causal usefulness, but independent factual, formal, source, or scope evidence is still needed for its warrant
- Task families and product families classify different things - Task families group obligations or evaluations; software product families group products through declared commonality, variability, and reusable production scope
- Tentative theory - Definition — a tentative theory is a theory proposed as a solution to a problem, which stays open to criticism and revision however well it has survived; Popper's term with nothing added, a status of every theory
- The 2026-08-30 Commonplace revision used retained theory to guide computational search - A 2026-08-30 Commonplace revision shows retained project theory guiding computational search while the operator supplied decisive global-fit selection
- The Bitter Lesson defense portfolio has one load-bearing member for the form-only rebuttal - The production-method versus representational-form distinction answers only a narrow weights-only inference; theory-guided bootstrapping is a provisional first strategy under incomplete global evaluation, not a defense of continuing hand production
- The bitter lesson selects against unearned reach, not against structure - The lesson selects against claims whose reach was asserted rather than earned by a refuting test, not against structure or origin — theory search in readable forms is its own method; earned reach protects the claim, not its carrier
- The bitter lesson selects production methods, not representational forms - The lesson's axis is production method — hand-crafted versus search-and-learning — not representational form. Learned localized forms are therefore a coherent scaling hypothesis, with cross-artifact credit assignment as the decisive open problem
- The deployed system, not the model alone, is the unit of learning - Because prompts, retrieval, tools, and runtime policy jointly determine deployed behavior, model-only learning leaves consequential system choices fixed
- The readable-artifact loop is the tractable unit for continual learning - Identifies the natural-language-plus-symbolic pair as the tractable first loop for representational-form coevolution because it shares context, operates at current tempos, and already has a codification boundary
- The tag-readme change as an observed causal-connection trace - One bounded Commonplace trace establishes that operative self-representation can exert causal force in both directions; it does not establish system-wide reflective coverage
- Theory builder - Definition — a theory builder applies conjecture and refutation to stated theories: it acts on them, criticizes what they say, and lets the result shape the next round; addressability and persistence are graded
- Theory building and capacity building make the same kind of fallible commitment - Theory building and capacity building both retain resolutions their evidence does not entail; an explanatory commitment stays answerable to the object it describes while a constructive commitment changes the object, so retraction differs in kind
- Theory building has distinct epistemic, structural, and implementation precedents - Conjecture and criticism, causal self-representation, and persistent artifact editing supply different precedents for a theory builder's operations; similarity on one does not establish the others
- Tool usefulness, computational autonomy, warrant, and system power are separate dimensions - Tool usefulness, computational autonomy, warrant, and system power move independently in a human-agent system, so a progress claim has to say which one moved and autonomy gains do not license power claims
- Treat continual learning as representational-form coevolution - Behaviour change spans distributed-parametric, natural-language, and symbolic forms, so the question is how their improvement loops relate — not which is the real locus of learning
- Universal software factory needs a declared universality axis - Universal software factory is ambiguous unless the universality axis, covered class, supplied inputs, adequacy relation, and resource bounds are declared
- Warranted autonomy is bounded by oracle domain - Bare autonomy is free, but warranted evaluation autonomy extends only to the candidates an oracle can assess with the required confidence
- Warranted transfer out of the human cut leaves people the hardest-to-warrant decisions - When a system preferentially transfers decisions whose premises, criteria, and checks are available, the remaining human decisions become harder to warrant per decision; this predicts a residue composition, not structural computational openness
- Weakly discriminated qualities tend to be underselected - Statistical conjecture: under named proposal-selection conditions, unequal oracle discrimination yields unequal enrichment; absolute degradation needs an additional directional mechanism
- World models assess explanatory-reach through action-conditioned prediction - Learned world models can assess explanatory-reach when action-conditioned predictions are tested across the interventions or shifts a commitment claims