A complete theory path does not establish improved capacity

Type: types/note.md · Tags: self-improving-systems, warranted-autonomy

A theory can guide a useful change without the system learning from criticism of that theory. Evidence of theory use, a relevant outcome, a response to criticism, and later use support successively more complete accounts of the process. A theory builder states its theories, acts on them, criticizes what they say, and lets the result of criticism shape the next round. The definition does not require that this improves anything. Learning, in the sense of improved capacity for future action from the process of conjecture and criticism, is a further claim. A complete observed sequence can still fail to improve that capacity.

The reflective case adds a separate condition. The theory represents selected aspects of the system itself, and that self-representation has a two-way causal connection to those aspects inside the declared boundary. As the reflective-system definition states, this is an architectural capacity; it need not already have been exercised. For a theory builder, the reflective qualifier places this connection in consumption and criticism: the builder's operations consume its method texts, and criticism tests those texts against records of its own operation. The connection is not kept up automatically, so text and operation can diverge until criticism finds the gap. Evidence for a particular reflective learning episode must connect the claimed theory use and criticism to that same system path.

Evidence forms a ladder

Four claims distinguish the links of an empirically tested path:

  1. Mediation. Changing or withholding the theory changes a proposal, evaluation, recovery step, or realized intervention because of what the theory says.
  2. Empirical contact. The intervention produces an outcome that bears on the theory rather than merely accompanying it.
  3. Response to criticism. A formulated criticism of the theory uses the outcome to revise its content, scope, assessed support, or operational role. Rejection can qualify. So can survival of an attempted refutation when the result changes how the system would rely on the theory or choose further tests. Recording a pass alone does not establish this effect; surviving does not change the theory's tentative status.
  4. Recurrent mediation. The result of that criticism affects a later operation on the same path, through a revised theory, changed reliance, or reconstruction from retained criticism.

A contemporaneous citation at the decision point is a mediation trace. It identifies the theory the process claims to have used; it does not establish that the content was load-bearing. Withholding, replacing, or perturbing the theory and observing a changed decision is stronger evidence, bounded by the contrast actually run.

The ladder separates claims about a process. It is not a universal sequence that every instance of learning must traverse: criticism can proceed by argument without an empirical intervention. Nor does the fourth level certify learning. Changed decisions can be worse. The improvement claim needs evidence about the capacity at issue, the criterion of improvement, and why the effect is attributable to the process of conjecture and criticism.

Capacity can improve before an occasion for exercising it arises. Observed later action can establish recurrence and provide evidence of capacity, but actual exercise is not a condition of learning. A claim of capacity at a later time does require the improvement to persist to that time. Later loss does not cancel an earlier improvement. Where improvement is unestablished, report the observed process at its supported strength.

The functions share one path, not one substrate

For a claim of recurrent learning through theory, interpretation, criticism, persistence of its effect, and later use must belong to one connected causal path. Disconnected witnesses do not establish the joins. A theory used in one decision, an unrelated criticism, and a later successful operation cannot be combined merely because they occurred in the same project.

The functions can use different substrates. A model can interpret a theory, a test or later demand can challenge it, an artifact can preserve the result, and a symbolic runtime can carry it into later operation. The effect can also persist through criticism from which a theory is reconstructed. Neither a single carrier nor retention of the assembled theory is required.

Interpretation is not reach-assessment

An interpreter can understand and apply a false theory. Deriving the consequences it claims does not establish that they hold or that its stated scope is sound. Interpretation makes content usable; reach-assessment judges whether the claimed generality is warranted. Collapsing them makes a model appear to warrant whatever it can explain.

Correction requires an opportunity for something the theory says to be challenged. Mechanical tests, decorrelated critics, held-out tasks, and later operational failures offer different strengths of correction. Their independence is an evidence question: can the candidate's rationale be overturned? It is not a requirement that another actor supply criticism. One model can propose and criticize its own theories; decorrelating proposer and critic concerns how well criticism works.

The current LLM-plus-artifact realization

Commonplace studies retained theories as a research arrangement. An LLM can interpret natural-language theories before their relevant concepts have been fully formalized. An artifact can preserve the theory across bounded calls. Addressability lets criticism and revision target particular assumptions, scope conditions, or parts. These are separable functions, and none supplies correction or improved capacity by itself.

This arrangement takes the high end of graded persistence: it can accumulate targeted changes without rebuilding the theory each episode. Its advantage is conjectured against lower grades: a builder that reconstructs the theory from retained criticisms, and the persistence baseline, which rebuilds from records of inputs and outcomes that carry nothing criticism produced. A theory criticized and replaced whole can still support learning; a stored theory that no process would consume cannot. Private linguistic formulation and criticism are not excluded, although opacity may leave their presence or effects unestablished.

Retained addressable theory must earn its retrieval, maintenance, and consistency costs. A hand-crafted bootstrap fits the Bitter Lesson only when learning can outgrow it. A future model or learned module may supply functions now divided between model and artifact. That possibility challenges the chosen realization's value, not the need to establish causal use, openness to criticism, and improved capacity.

What the retained examples establish

Commonplace's human-agent pathway shows operative self-representations being interpreted and changed, including changes to instructions, validators, schemas, and code. In its tag-readme case, the validator pass establishes consistency of the marks with the declared criterion; the improvement remains a claim. Humans supply decisive assessment and adoption judgments. The evidence establishes a bounded connected pathway, not demonstrated improvement in capacity, independent computational theory possession, or a general advantage for the arrangement.

The retained Exo review, pinned to its inspected checkout, describes editing prompts, tools, and executor code, then rebuilding and restarting. That supplies a reflective revision surface, retention, and continuation machinery. Build success, tests, and post-restart behavior can reject broken candidates. The record does not establish independent assessment of a self-theory's semantic reach, a result of criticism mediating a later episode, or improved capacity through that process. The review did not run a live instance.

The Gödel-machine construction shows in specification how a formal self-representation and proof-governed switch can participate in self-modification. Its warrant depends on its formalized assumptions and objective; it does not cover concepts absent from that formalization. Proof-governed switching alone establishes neither the presence nor absence of criticism of theory content in a complete system. The construction is not an observed recurrent learning result.

Scope

  • Reflection is relative to a boundary and represented aspects. Learning about an external target need not be reflective unless that target helps determine the modifying system's own behavior; the two-way causal connection must still hold.
  • Iteration is a theory-builder condition, and reconstruction from retained criticism meets it. Retaining the assembled theory across problems with fine-grained addressability is a design commitment of the chosen arrangement, not a condition of being a theory builder.
  • Criticism can be delayed. A later demand or maintenance failure may provide the relevant challenge, and claims must stay within what it tested.
  • Identity across records helps establish a connected path; it does not independently establish causation or improvement.

Open Questions

  • What interventions best distinguish load-bearing theory use from plausible post-hoc citation?
  • When does addressing particular parts improve capacity or cost compared with whole replacement or reconstruction?
  • How much does decorrelating criticism improve assessment of self-theories?
  • Can a computational system sustain recurrent learning across novel demands and delayed evidence without a person supplying decisive assessment?

Relevant Notes: