Three episodes for checking the main-path definitions
Status: Illustrative protocol cases, 2026-09-17. Illustrates the adopted hypotheses. These are constructed examples, not run results or claims about a selected consuming project. They use the downstream protocol.
The example consumer maintains a versioned document service. Its KB records domain commitments, task procedures, and evidence needed for later changes. The actual first consuming project remains to be selected. Each episode uses a task budget B and builder-adaptation budget R, with identical declared values for the relevant baseline and candidate; the study must choose those values before execution.
1. Product theory revision
Starting repertoire. The seed KB claims that a document's canonical URL identifies the content used in an answer. The consumer can read and cite documents, and the service permits the content at a URL to change. The builder can amend notes, update a citation validator, and release a new KB.
Task and external consequence. A user asks the consumer to reproduce an earlier answer and show the exact evidence it used. The recorded URL now serves a different version. The output judge rejects the answer under the predeclared reproducibility criterion. Consumption records show whether the consumer actually relied on the URL-only claim.
Diagnosis and revision. The builder considers at least a KB error, a consumer that ignored valid instructions, and missing source access. A trace plus a controlled replay using distinct versions tests those possibilities. If the URL-only commitment is implicated, the builder revises it to distinguish source identity from captured version, updates the dependent citation procedure, and tests the validator on both changed and unchanged content. That is product revision, not necessarily reflection.
Later assessment. Freeze the new release and test reserved reproduction tasks containing source changes and unchanged sources. Compare the seed and candidate under B, charging the repair against R. Success requires matching the right evidence version without inventing unavailable evidence; honest reporting that reproduction is impossible can be the correct task outcome.
What follows. A predicted difference under a controlled version change can implicate the URL-only commitment. Better later outcomes support the revised product on the assessed cases. They do not establish every claim in the KB or a theory about how Commonplace should generally build knowledge. If the consumer never read the note, the rejection alone does not diagnose that note.
2. Reflective machinery revision
Starting repertoire. Commonplace already has a correct version-pinning note for its own workers. The builder's self-theory says that its title-based retrieval supplies the guidance needed when those workers produce or repair task-specific procedures. That internal lookup misses the note, and a worker publishes a reproduction procedure that records only a URL. The builder can inspect its production traces, revise its self-theory, and change the lookup routine or index used by its own workers.
Task and external consequence. A consumer follows the delivered procedure and returns evidence from the wrong version. The judge reports an incorrect result without attributing it to the KB or retrieval. Release and consumption records connect the consumer's work to that procedure. The builder's own production trace shows which source-handling guidance its worker consulted.
Reflective change. The builder hypothesizes a coverage failure in its own retrieval, probes paraphrased production queries, and revises the self-theory's retrieval claim. That revision guides a query-expansion or indexing change for the builder's workers. It regenerates the affected procedure and updates the self-description to describe the machinery actually installed. Both directions matter: the self-theory guides the machinery change, and the machinery's observed operation constrains the self-theory. Changing only the consumer's product lookup would not establish this reflective path.
Later assessment. Start matched builder runs with the same guidance and demands, varying the internal retrieval change. Record what each worker consumes and which procedure each produces. Freeze the resulting releases and compare their use on reserved consumer tasks under B and R. Measure the predicted internal retrieval difference as well as externally accepted outcomes. If another record reconstructs the same guidance, account for that alternate consumption path.
What follows. Evidence that the self-theory guided the machinery change and that machinery observations corrected the self-theory supports a reflective episode. A chronological trace alone does not establish those causal links; interventions on the commitment or its consumption path can test them. Changed outcomes alone could show useful machinery optimization without showing that the self-theory caused it. This episode remains in the main path: internal diagnosis and active probing do not move final acceptance inside the builder.
3. A broader claim without external assessment
Starting repertoire. After episode 2, the builder proposes that its new retrieval procedure will retrieve all relevant methodology in every future area. It retains a note asserting that broader claim because internal review finds the explanation convincing.
Available evidence. The observed interface covers document-reproduction tasks under one budget. Neither the data nor the stated protocol assesses every future domain. Success in the prior episode therefore does not establish the expanded claim.
Required disposition. Keep the tested result scoped to the assessed tasks. Retain the broader claim as an inquiry candidate with its missing assessment identified, design a bounded discriminating test, or decline the broader claim. Do not consume the internal review pass as external warrant. The general-case obligations identify what additional support a use outside the tested scope would require.
What follows. The broader claim lacks external assessment under the declared interface. The builder as a whole has not changed category merely by entertaining it. Accepting a claim for inquiry and licensing it to govern routine work remain different policy decisions.
Definition checks supplied by the cases
| Proposed change | Episode check and result |
|---|---|
| Declare the evidence interface | All three cases need the task, judge, consumed release, feedback access, and assessed scope to distinguish their claims |
| Judge extension through capability under a budget | Episode 2 can improve retrieval by changing an existing routine; a new filename or slot is not necessary, and a diff alone is insufficient |
| Keep reflection distinct from ordinary theory use | Episode 1 can revise product knowledge without a self-theory change; episode 2 requires the additional causal path |
| Keep autonomy separate from objective governance | Either episode can use computational internal roles with externally supplied acceptance; a changed objective is a separate event |
| Remove passive acquisition as a definition requirement | Episode 2 uses targeted probes while final outcomes remain externally judged |
| Defer the ideal interpreter from the main comparison | Outcome comparison can record failure before localizing interpretation versus theory error; localization still matters to the mechanism claim |
| Keep policy outside tentative status | Episode 3 is a tentative proposal without being licensed for every consumption path |
These cases check coherence and reveal required records. They do not show that the proposed methodology can produce the repairs, meet the budgets, or match a human-staffed builder. Those are empirical questions for the run.