A claim without external assessment carries three obligations
Type: kb/types/note.md · Tags: foundations, self-improving-systems, learning-theory
An externally tested theory builder can compare outcomes without first proving every explanation it used to produce them. The comparison supports performance on its assessed tasks and conditions. It does not establish each internal theory or every broader scope the builder proposes. For a claim the declared interface does not assess, the builder must supply for itself what the interface would have supplied: what counts as contradiction and what support licenses each use, in place of the external falsifier; a comparison level when an objective changes, in place of the external objective; and a performance measure that does not rest on its own evaluators, in place of the independent outcome level. The third carries a consequence the interface never supplied, attribution. External assessment does not locate a fault, so attribution is required whenever a claim asserts a cause, inside or outside the main path. Without an independent outcome level it is required for a performance claim too, because no outcome then absorbs an interpretation error and a theory error together. These three are the boundary of the main path: a claim leaves it when its consequences face no external assessment, not when the builder reasons about a failure.
The boundary applies to a claim and a proposed use, not to the builder as a whole. A builder can work on externally assessed tasks while also retaining broader inquiry candidates; entertaining an untested idea does not change its category.
Support for the proposed use
State what would count against the claim, which unit the evidence supports, and what kind of reliance that support licenses. The supported unit may be a claim, a conjunction, or a model under a scope, since theory warrant is tracked at the finest granularity evidence licenses. A system-level outcome does not uniquely locate a faulty component or validate every component of a successful system.
Candidate retention is separate from reliance. A claim may guide an experiment without being ready for routine use or codification, because derivation and inheritance give starting warrant and evidence earns scope and current task fit alone does not warrant costly entrenchment. Candidates below a use threshold can remain targets for inquiry; storing them does not make them accepted theory, and calling a theory tentative does not license any use. WikiSkill is the paradigm case of this separation built as a system: a skill is admitted only on strict validation improvement, while the wiki's diagnoses and rejected proposals are retained regardless of outcome and remain available to later proposals. RuleMem shows the failure when the separation is missing: a likelihood-based admission score improves aggregate accuracy, and an admitted rule still overrides explicit contrary evidence.
Assessment of support can itself be wrong. Distinguish actual support from an evaluator's judgment that it is sufficient; the failure of concern is false-positive acceptance becoming operative.
A comparison level for changing an objective
Identify what makes a claimed improvement better when a revision changes the acceptance rule itself. External users can authorize a changed demand; the comparison must then report the new objective and distinguish that change from improvement under the old one. A new request may instead instantiate the same objective without changing it. When no external acceptance applies, the builder must state what comparison level licenses the objective change or refrain from claiming improvement, since revising an improvement objective is licensed from outside it or is not improvement. A conceptual revision that changes what the objective commits to is an objective change even when the objective's text is unchanged, and is reported as such; the Darwin Gödel Machine agent that deleted the marker its hallucination detector keyed on is the mechanical instance. This is an obligation, not an inference that computation prohibits objective change.
Attribution beyond the observed outcome
This obligation is conditional. It binds a claim that asserts a cause, in either case, and it binds a performance claim only when no independent outcome level absorbs interpretation and theory errors together. An external judge may establish that a task failed while leaving open whether the cause was a theory, its interpretation, retrieval, execution, or the environment. To assert a particular cause, specify a discriminating trace, intervention, or test that could distinguish the alternatives; a plausible explanation is a candidate for that test. No adopted standard separates an interpretation error from a theory error; failures are localized with ordinary probes. Failed attribution limits the causal claim; it does not erase an observed performance difference under a sound comparison.
What stays inside the main path
Diagnosing a failure, choosing what to observe, revising a self-theory, or running an active experiment stays inside the main path when the resulting performance claim faces the declared external assessment. A diagnosis may remain provisional while its proposed repair is tested. Internal proxy tests can guide development; they do not become the external acceptance criterion because the builder passed them.
For an unsupported extension of scope, record the unassessed consequence, the intended use, the evidence missing for that use, and the next bounded test or the reason to defer it.
Open Questions
- Whether locally warranted revisions compose into a warranted lineage. A succession of local approvals does not supply warrant for the sequence: once an evaluation result guides the next revision, the later candidate depends on the reused evidence, and generalization in adaptive data analysis shows the ordinary generalization argument no longer applies without a protocol. Whether the set of states reachable by warranted revisions is closed under the seed objective and the evidence it admits is argued in both directions and unsettled.
- Which consumption paths need different thresholds, and what evidence licenses each. An external falsifier provides an outcome signal, not a complete retention policy for the internal theories behind it.
- Which dependencies must survive revision. Dependency maintenance, evidential warrant, and selection policy are separate, as the assumption-based TMS and belief-base contraction ingests separate them.
- Which guarantees concern eventual learning rather than current permission to rely. A convergence guarantee does not warrant the current output.
Relevant Notes:
- Externally tested theory builder — defined-in: the case whose interface discharges these obligations
- Theory builder — grounds: the evidence interface as a declared parameter, and the lineage question this note keeps open
- Theory warrant is tracked at the finest granularity evidence licenses — grounds: the supported unit may be a claim, conjunction, or model
- Derivation and inheritance give starting warrant; evidence earns scope — grounds: candidate retention below a use threshold
- Current task fit alone does not warrant costly entrenchment — grounds: codification needs more than one fit
- Revising an improvement objective is licensed from outside it or is not improvement — grounds: the comparison level an objective change needs
- False-positive generation is filtered before retention — grounds: misassessed support as the characteristic failure
- Warranted autonomy is bounded by oracle domain — grounds: support sufficient for a consumption path at a required confidence