What makes human-inclusive self-revision non-trivial?

Type: kb/articles/types/article.md · Status: draft

Draft. This article is circulating for comments; its claims, structure, and even its central thesis may still change. Comments are welcome below.

Imagine a constitutional order with its authoritative text stored in a writable file. Someone with access can change the file, but that alone does not constitute an amendment. A formal amendment gains authority through the roles and procedures the order recognizes. It enters the authoritative record and affects practice when institutions apply it. If the order permits changes to its amendment rule, the same process can replace part of its own revision machinery.

People propose, approve, interpret, and apply amendments in roles the order recognizes. The order need not encode their private reasoning.

Commonplace's human-inclusive claim follows the same pattern: its declared boundary includes designated maintainers, whereas the systems compared in the reflective self-improvement article draw their reported boundaries around agent loops. Including their research teams would make those systems broadly editable too. The claim under examination is not that maintainers can change everything. It is that Commonplace gives their changes an explicit path into operation and keeps that path available for revising what they install.

The audit tested a more precise version:

Commonplace gives its human–agent operator a reusable path to revise any repository-defined artifact or relation through which behavioral authority or revision governance is exercised. The path supports identifying and criticizing the incumbent, developing a successor, warranting the transition, installing the successor in a live authority path, and later revising that successor.

Behavioral authority describes how a retained artifact shapes operation: who consumes it, through which channel, and with what force. Complete addressability of behavioral authority means that the human–agent process can inspect, criticize, and selectively revise every repository-defined artifact and relation in those paths, including the machinery governing revision. It does not require access to a maintainer's unarticulated judgment.

A general revision affordance combines complete addressability with an applicable revision method, warrant for adoption, operative installation, and continuity of the revision path.

Meta remains a role within an episode, not a permanent layer. The revision process can repeat: each episode can apply the same addressability requirement to the authority-bearing arrangement left by the previous one.

The six-path audit found broad but incomplete support for this hypothesis on the paths it examined. The claimed affordance remains relative to Commonplace's declared system boundary and does not by itself show that improvements compound.

Why editability is not enough

Including humans makes a weak reflection claim easy: a maintainer can model the system and act on that model. A stronger claim requires a self-representation to affect operation. The practical benefit of such a representation is addressability: a represented commitment can be inspected, criticized, selectively revised, and reused.

Revision likewise requires more than editability. A general revision affordance has six obligations:

Requirement What it rules out
Declared boundary, authority path, and purpose Treating an unspecified ability to edit as "self-improvement" without saying what is changing, how it shapes behavior, or why one result is better.
Complete addressability of behavioral authority A repository-defined artifact or relation that shapes operation but is exempt from inspection and selective revision, or whose role and encoding must be rediscovered.
Applicable revision method An external designer inventing the whole approach afresh for this target. The method may route to open-ended theoretical reasoning rather than prescribe fixed steps.
Warrant for the transition Changing a governing rule merely because it is editable. The successor must earn adoption over retaining the incumbent; uncertain changes may instead earn a bounded experiment.
Operative installation A good proposal or saved rationale that never reaches a consumer with behavioral force.
Continuity across successors A revision that leaves its artifacts inspectable but disables the determination, admission, installation, or later-use path needed for another operative change.

Together, these obligations define the system-level affordance. The narrower repeatable-path test traces representation, determination, admission, installation, dependence, and continuity for one redesign class.

These are logical obligations, not required components. A proposal-selection path can compare an explicit candidate with the incumbent and reject it. A direct update may instead use evidence under a rule that warrants the successor. The general affordance does not require one search algorithm, evaluator, or design lifecycle for every change.

The successor need not be a one-for-one replacement artifact. A better arrangement may narrow a rule, split one concept into several, reroute its useful work, or remove an artifact whose role is no longer warranted. What matters is comparing the successor with continued use of the incumbent arrangement.

Complete addressability, like reflective coverage, is relative to a declared boundary and an operation profile. Commonplace may strongly support inspection and revision of natural-language theory, types, and validators while having weaker paths for objectives, evaluator validity, authority arrangements, or the realization of its internal binding requests at an external model dependency.

How the revision path works

Commonplace combines general theory, specialized methods, and operative artifacts rather than one universal revision algorithm. The audit found the following parts of the general affordance on the paths it examined.

  • Authority-bearing organization is represented. Types, collection contracts, routing rules, instructions, review criteria, ADRs, schemas, validators, configuration, and code expose much of the repository's organization as inspectable artifacts. Contracts, configuration, and code also expose many of the consumer, channel, and force relations through which those artifacts shape operation.
  • Open problems have a place to develop. The workshop layer holds investigations until their results are ready for the library. A claim that extends current knowledge can move through the discovery lifecycle: anomaly, conjecture, derived consequences, testing, acceptance, and integration. A proposal followed by an ADR is one path from a mature design question to an installed decision, not the definition of revision itself.
  • Theory remains a fallback. No finite procedure anticipates every new kind of change. Commonplace's two-layer execution model keeps theory available when the fast methodology does not cover a case. If fallback reasoning recurs, it can be retained as a method. Within the declared boundary, the human operator supplies semantic judgment where the methodology has not yet been codified.
  • Evaluation follows the claim and its form. Natural-language theory needs criticism and semantic judgment; symbolic commitments can use tests, schemas, invariants, or proofs; mixed changes combine them. Explanatory-reach is a central quality criterion for transferable theory, not a universal oracle for every artifact.
  • Accepted changes can become operative. Instructions, contracts, code, configuration, and validators give a decision a consumer, channel, and force. The operative-change test rejects revisions that are merely written down.
  • Rationale and history support another pass. ADRs retain accepted design reasoning; canonical source artifacts retain the result; version control supports diff review, rollback, attribution, and reconstruction. History alone is not the obligatory read path, but it helps a later challenge recover what happened.

Agents and maintainers operate through contracts, skills, review criteria, definitions, and code that they can also revise. These artifacts exercise behavioral authority when agents and maintainers use them to govern their work. This self-application structurally resembles a metacircular interpreter: the rules are artifacts in the system they govern. When those rules do not settle the next choice, live theory and human judgment provide the fallback. Retaining the result in operative artifacts turns the intervention into part of the system's revision machinery rather than an undocumented action outside its description.

A governing criterion revised in practice

Explanatory-reach shapes operation in three places: the root vocabulary names it, the notes collection uses it as a quality goal, and a semantic review gate applies it. It therefore helps Commonplace decide which theoretical claims deserve retention and reuse.

The criterion itself has been revised. The anchor theory originally drew a sharper contrast between adaptive and explanatory claims. A later revision recast that contrast as a polarity, added a test against rival practices, and required observed fit to discipline the explanation. That revision carried the revised test into the notes collection contract and the recurring explanatory-reach review. The later reach-assessment definition reused all four parts, while ADR 055 made the technical name unambiguous across the corpus.

This is substantive theory revision, not just proof that the files are editable. The audit confirms that the semantic gate is routinely invoked. But the gate omits the revised test's observed-fit requirement, and no retained review shows later dependence on all four parts. The evidence therefore shows that the criterion can be revised and reused, but not that later evaluation depends on every part of the revision.

The same path could challenge the core criterion, not only refine or rename it. If counterexamples showed that explanatory-reach rejects useful explanations or rewards a rhetorical shape rather than real transfer, a successor would need to explain that failure and preserve what the old criterion got right. It would then need to survive stated tests, be installed across its authority paths, and govern later review. Until a better successor earns adoption, retaining the incumbent is the normal result of comparing change with continuation.

Other revision cases

Commonplace has already changed several kinds of load-bearing machinery:

  • ADR 042 replaced the claimed exhaustive three-register taxonomy after a worked dialectical collection supplied a counterexample. The successor kept the useful theoretical, descriptive, and prescriptive profiles inside an open text-contract model. This revision changed the definition and root vocabulary and later made the article collection's editorial profile possible.
  • ADR 053 retired a load-bearing theory term after a 464-occurrence audit showed that it merged operations with opposite maintenance requirements. Rather than preserve the term under a new name, the revision moved its useful work to the two-layer theory, explicit lineage relations, and the discovery lifecycle. When use exposed a missing relation, ADR 054 revised the successor arrangement instead of defending the first repair.
  • The tag-README redesign changed instructions, schema, validation, and rendering. Later validation used the new check and exposed a defect in the associated search recipe.
  • ADR 056 revised the proposal and ADR lifecycle. ADR 057 then used its new alternatives requirement, and ADR 063 later challenged and revised the installed article lifecycle.

Together these cases are more informative than a declaration that every file may be edited. They show theory replacement, vocabulary retirement, symbolic enforcement, revision of design machinery, and later dependence on installed successors. They leave open how evenly the affordance covers different authority paths and whether it improves outcomes.

Where the affordance remains incomplete

The audit examined global goals, the explanatory-reach criterion, tag-README validation, the revision lifecycle, model bindings, and maintainer admission. Goals, contracts, criteria, procedures, validator rules, and model requests were inspectable and selectively revisable. In the validator and lifecycle cases, installed machinery was later reused and revised. This is stronger than repository writability and weaker than complete addressability.

The strongest gap in the broader revision affordance is generic maintainer admission. In the constitutional analogy, Commonplace resembles an order with several amendment procedures that refer to designated officials. But no general record states who holds office, what the grant covers, which conditions admit a proposed change, or which approval authorized the operative text. The installed artifact can still exercise behavioral authority; the generic admission and authorization path remains outside the represented revision affordance.

Model binding exposes a different gap. The repository makes the requested model, alias, and freshness partition addressable. But retained evidence shows a divergence between requested or recorded model identity and actual execution. This is an operative-realization gap, not an addressability failure of the request.

A gate's target cohort can change while its consumer, channel, and force stay fixed. A validator's target type and invocation trigger can change under the same conditions. The current three-part decomposition therefore does not fully identify an authority path. The live decomposition proposal leaves one choice open: whether applicability should be a fourth field or a required qualifier. Authorization, runtime realization, and dependency closure remain separate questions.

The audit covers six paths within Commonplace's declared frame. Provider weights, inference infrastructure, and hosting lie outside that frame. Wider coverage may reveal another barrier.

Warrant and compounding are separate questions

The theoretical Gödel machine provides a contrast in warrant: its incumbent formalization must prove that switching is better than continuing, even when the switch revises governing machinery. Commonplace instead uses fallible semantic and empirical warrant. It can therefore admit useful changes that the proof gate cannot certify, but also bad ones. This contrast concerns which changes may be admitted, not whether improvement compounds.

Even complete coverage would not establish compounding. That requires evidence that retained benefits help produce later improvements, directly or through reinvested savings, and that this feedback persists across episodes. The main article develops both the compounding test and the Gödel-machine comparison.