Open-ended theory learning and factory learning close the same reflective loop

Type: kb/types/note.md · Tags: foundations, self-improving-systems, learning-theory

TODO! remove or update to use software house instead of software factory

Two research directions in this knowledge base look like separate programs. One asks how a system acquires, tests, and revises explanatory theories about its own organization. The other asks how a software factory learns from its production experience. They are the same loop reached from opposite ends: each, pushed to where it stops being satisfiable on its own terms, requires what the other supplies.

The shared loop is the causally co-indexed path that theory-mediated self-improvement needs interpretation, retention, and independent read-back already names: a theory about the system's own behavior-determining organization is posited, interpreted into a change, made operative, exposed to consequences its author did not write, read back against the same retained object, revised, and consumed again. The convergence claim is that neither starting direction can stop short of this whole path.

From theory learning down to the factory

Run the discovery lifecycle — observe, conjecture, derive consequences, test, accept, integrate — on conjectures whose object is the conjecturing system's own organization. Two of its phases then have only one affordable realization.

Testing needs operative retention. For a theory about an external domain, the discriminating consequences can be gathered by observing that domain without touching the observer. For a theory about the system's own organization, the consequences that bear on it are largely the consequences of acting on it: what a change guided by the theory costs when a later demand arrives. So the theory must already be operative—retained, and acted on by later work—before the evidence that could defeat it exists. The test phase does not precede operative retention; it depends on it.

Integration is a machinery change. The lifecycle's final phase reconnects prior evidence under the accepted claim and updates the artifacts that use it. For a self-directed theory those artifacts are the system's own instructions, validators, evaluators, and tools. Integration is therefore a retained change to reusable machinery on which later work depends — the definition of experience-responsive retention.

The self-directed discovery lifecycle already contains a production-and-revision cycle over machinery. It is not adjacent to factory learning; it is factory learning with the theory made explicit.

From the factory up to theory learning

Experience-responsive retention as such asks only that production experience cause a retained machinery change that later production depends on. Shallow patching satisfies that letter: a failing case yields a rule, the rule is retained, later runs load it. The requirement bites at coherence, because holding a program theory means sustaining coherent search under delayed feedback — a modification must preserve purposes and organization that the immediate acceptance tests capture only partly, and the evidence that would expose the damage arrives after the change is already in the machinery. Something has to allocate search and interpret failures before that evidence exists, since open-ended improvement must allocate search before decisive evaluation is available.

What plays that role is a held theory of the factory's own purposes and organization. Such a theory cannot be a summary of the production record, because commitment, not derivation, creates new ground truth: building it fixes content the experience does not determine — a conjectured mechanism, a resolution adopted because production needed some choice. Once committed, that content cannot be re-derived from the record; it can only be revised.

A commitment that exceeds its evidence, whose consequences arrive later, and that must be revised when those consequences contradict it, is an ampliative conjecture under test. Factory learning pushed past patching therefore runs the discovery lifecycle over the factory's theory of itself.

What the convergence claims

Neither direction is a special case of the other. Each supplies the requirement the other cannot generate internally: theory learning supplies the search allocation and the coherence criterion that decide which machinery change is worth making; factory learning supplies the operative retention and the independent later consequence that decide whether the theory was right. A system holding one half fails in a direction the missing half predicts. A knowledge base that conjectures without making theories operative accumulates untested claims, because it never buys the consequences that would defeat them. A factory that patches without holding a theory accumulates local special cases, because nothing in it recognizes which existing organization a new demand should have gone through.

The claim is refutable at these failure modes. A system that sustains long-run coherent factory improvement while holding no revisable theory of its own organization refutes the second derivation directly, and plain retention and retrieval of the raw production record is the specific rival that would do it.

Two axes place the Gödel machine

The convergence holds for a region of the design space, not everywhere in it, and two independent axes locate the region.

Transition licensing — what makes a change to the system's organization admissible: a proof under stated axioms at one end, bounded empirical evaluation at the other.

Theory provenance — where the theory the loop reasons from comes from: supplied and fixed at one end, acquired and revisable at the other.

The axes vary independently. A formal theory-refinement system proof-licenses changes to a theory it also revises. A benchmark-gated coding agent empirically licenses changes inside a task decomposition nobody revises, and learning inside a fixed decomposition inherits its mistakes states what that costs.

Gödel machines are a proof-governed case of reflective self-modification, and that note develops the licensing axis. The second axis places the same construction differently: its axioms describing the machine, its hardware, its environment, and its utility function are premises of every rewrite and are revised by none of them. Its provenance is pinned to supplied-and-fixed.

The Gödel machine therefore enters as a contrast case, not a maturity endpoint. It closes the proposal-selection improvement loop completely — search, reject-capable evaluation, and operative retention are all present and mechanized — while leaving the theory-learning loop empty. Its self-representation is a premise, not a candidate: the machine can improve indefinitely while never revising its account of its own organization.

That separability makes the convergence claim contentful rather than definitional: pin provenance to supplied-and-fixed and the two loops come apart cleanly, so the convergence is asserted only for the corner where the theory is acquired and revisable. Provenance also decides where the pre-formal work sits, since improvements outside the admitted formal language need a pre-formal stage somewhere: acquired provenance puts that stage inside the loop, while supplied provenance fixes it at design time, in whoever chose the axioms. The same placement explains why the machine's utility function is unrevisable from within — revising an improvement objective is licensed from outside it or is not improvement, and a construction whose highest level is its own supplied objective has no outside.

Scope

  • The convergence is over the loop's functional requirements, not shared substrate, tempo, artifact granularity, or identity of research agendas. The two directions still differ in what they measure and what counts as a result.
  • Reflective membership is boundary-relative. Open-ended theory learning whose object lies outside the learner's own behavior-determining organization is ordinary empirical inquiry and does not converge with factory learning.
  • The derivations establish what the loop requires, not that any current system closes it. Both directions currently rely on human judgment at the acceptance and read-back steps.
  • The second derivation assumes delayed and partial acceptance evidence. A production setting whose tests fully capture the purposes a change could damage would not force a held theory, because the immediate gate would carry the coherence burden.

Open Questions

  • Is there a constructible system at the proof-licensed, acquired-provenance corner, or does proof licensing force supplied provenance in practice by requiring the theory to be axiomatized before it can license anything?
  • What is the weakest evidence that a system revised its own theory rather than regenerating a different one, given that only revision preserves the addressability the loop depends on?
  • Can either derivation be run with a human removed from the acceptance step without collapsing to the fixed-provenance corner?

Relevant Notes: