Reflective theory refinement needs interpretation, retention, and independent read-back

Type: kb/types/note.md · Tags: foundations, self-improving-systems

Theory refinement may improve sample efficiency under structured shifts regardless of what its theories describe. Reflective theory refinement is the narrower case in which the retained theory describes behavior-determining organization inside the declared system boundary and participates in changing that organization.

The reflective case needs several functions that are easy to collapse into one. They should be kept separate because they fail separately:

  1. Reflective membership. The theory concerns organization that helps determine the system's behavior and lies inside the declared revision path.
  2. Semantic interpretation. Some process can state what the theory claims, apply it to a case, derive consequences, and use it to guide a proposal, diagnosis, or recovery step.
  3. Addressable retention. The theory persists with enough structure that its content, assumptions, scope, confidence, or status can be revised rather than only regenerated or discarded wholesale.
  4. Independent exposure and read-back. Consequences not authored only by the candidate can contradict, qualify, or support the theory and are read back against the same retained object.
  5. Continuation. The resulting theory state affects a later operation on the same path.

The functions share one path, not one substrate

These functions need one causally integrated and co-indexed path. They do not need one substrate. A retained artifact may supply addressability, a language model may supply semantic interpretation, tests or later demands may supply exposure, and a symbolic runtime may supply continuity.

Interpretation is not reach-assessment

An interpreter can understand and apply a false theory. It can derive the consequences the theory claims without having independent grounds for deciding whether those consequences hold in the world or whether the theory's scope is genuine. Interpretation is therefore a semantic function. Reach-assessment is an epistemic function supplied by evidence, comparison, criticism, or an oracle capable of defeating the theory.

Collapsing the two makes a language model appear to warrant whatever it can explain. It also hides the evaluator problem inside the word “interpretation.” The system needs both: interpretation to make a theory operative and independent exposure to correct its use.

The independence requirement is graded rather than absolute. A mechanical test, a decorrelated critic, a held-out task, and a later operational failure provide different strengths of correction. The relevant question is whether the candidate's own rationale can be overturned, not whether every check comes from outside the technical system.

Evidence forms a ladder

A complete recurrent loop is the strongest evidence, but it should not be used as the minimum definition of every improvement a theory mediated. Four claims can be distinguished:

  1. Mediation. Changing or withholding the retained theory changes a proposal, evaluation, recovery step, or realized intervention.
  2. Empirical contact. The intervention produces an outcome that bears on the theory rather than merely accompanying it.
  3. Theory refinement. The outcome changes the theory's content, scope, confidence, status, or operational role. Explicit rejection or principled retention after a refuting opportunity also counts as a theory-state change.
  4. Recurrent mediation. The refined theory state mediates a later operation on the same behavior-determining path.

A contemporaneous citation at the decision point is a mediation trace. It identifies the theory the process claims to have used, but it does not show that the theory was load-bearing. Withholding, replacing, or perturbing the theory and observing a changed decision is stronger evidence.

A useful change may reach the first or second level without reaching the fourth. That is still evidence about theory-mediated operation. It should not be reported as recurrent self-improvement until the later-use link exists.

The current LLM-plus-artifact realization

Natural-language theories often arrive before a formal language, variables, and acceptance test exist. An LLM can interpret those theories across cases that no symbolic procedure already covers. A retained artifact gives the theory a stable, inspectable address across bounded calls. The pair is therefore a practical current realization of semantic interpretation plus addressable retention.

Neither half is sufficient. A model without retained theory re-derives an account each episode and cannot reliably accumulate targeted rescoping. A retained document that nothing interprets or retrieves is inert. But the pair also does not supply independent correction or continuity by itself. Those functions must be connected separately.

This is a current engineering claim, not a theorem that natural-language, parametric, and symbolic carriers must remain separate. Another substrate could supply the same functions, and learned systems may absorb current boundaries. What must survive is the causal role and its independently analysable failure surface, not the carrier.

What current examples establish

Commonplace is a reflective human-agent system: retained theories are interpreted, revised, and sometimes turned into operative instructions, validators, schemas, or code. Humans still supply much of the independent assessment, blame assignment, acceptance, and continuity across ambiguous cases. It therefore gives evidence for the mechanism and for useful human-inclusive theory work, not independent computational theory possession.

Exo provides a stronger technical revision surface: it can edit prompts, tools, and executor code, rebuild, and restart. That shows reflective membership, operative retention, and continuity over some changes. Build success, tests, and post-restart behavior can reject broken candidates, but the retained record does not establish that the system independently assesses the semantic reach of a self-theory or that a revised theory changes a later episode. The missing result is read-back at the strength claimed, not mere ability to edit itself.

A formal proof-governed system could supply the functions inside a formal language, as the Gödel-machine case shows in specification. That route pays for mechanical acceptance with a fixed formal vocabulary and axiomatized objective. It does not cover self-theories whose relevant concepts have not yet been formalized.

The Bitter Lesson applies to the realization

The present LLM-plus-artifact arrangement receives no permanent exemption. Retained theory earns its place only where persistence, selective rescoping, inspection, and cross-episode use improve the learning path enough to pay for retrieval, maintenance, and consistency costs. A hand-crafted bootstrap fits the Bitter Lesson only when learning can outgrow it.

A future model may perform some present theory search internally or move the same functions into learned modules. That would replace the carrier without refuting the functional requirements above. The research claim is that explicit theory is useful working state in the current bootstrap and may remain useful where sparse, named revision and external inspection matter. Its persistence is empirical.

Scope

  • Reflective membership is boundary-relative. Refining a theory of an external target is not the reflective case, unless that target helps determine the modifying system's own behavior.
  • Addressable retention need not mean one document or perfectly atomic claims. It means that the revision operation claimed by the experiment has a stable target.
  • Independent read-back can be delayed. For coherent program modification, a later demand or maintenance failure may be the strongest available oracle.
  • An unchanged theory after confirming evidence can still have mediated a useful improvement. It establishes less than theory refinement unless the record shows a deliberate theory-state judgment.

Open Questions

  • What intervention best distinguishes load-bearing theory use from a plausible post-hoc citation?
  • How much structural addressability is needed for selective rescoping rather than whole-document replacement?
  • What kinds of decorrelated criticism are strong enough to count as independent read-back for self-directed theories?
  • Can a computational composite sustain the full recurrent loop across novel modification demands and delayed evidence without exporting the decisive interpretation or acceptance to a person?

Relevant Notes: