Operationalizing the self-improving-systems cluster

Goal

Operationalize the self-improving-systems cluster as an ontology and a Commonplace-specific overlay on established improvement methods: wire it into the decisions that change Commonplace's behavior-determining organization, and distill only what the host methods cannot supply — the operativity test, the reflective/addressability profile, and the warrant boundary. Do not construct a standalone improvement methodology (delegation stance adopted 2026-07-22; external-delegation-assessment §3a). Keeping the cluster good research and applying it are the same program — every ambiguity or contradiction left in the theory surfaces as a failure at application time.

The trigger finding (full-cluster review, 2026-07-21): the cluster currently fails its own operativity test. In its own vocabulary, it has no reliable consumer, channel, or force into actual change decisions — nothing loads it when someone modifies a skill, validator, type spec, or collection contract, and the shipped review system (the repo's most developed evaluation machinery) is never read through the cluster's loop vocabulary. Hygiene defects found in the same review were fixed directly (world-models tagging gap, misattributed profile citation, explanatory-reach link, cumulativity-test exclusion left implicit, slug-length failure).

Three authority paths

The cluster can reach change-time behavior along three paths, and the cluster's own theory grades them. The workshop treats all three as real design targets, not just the first:

  1. Wired — instructions, review gates, and AGENTS.md routing load the (distilled) theory whenever Commonplace's behavior-determining organization is being changed. The strongest wire; enumeration-like.
  2. User-invoked — the maintainer tells the agent to consult the cluster for a given change. No engineering needed; the human is the retrieval wire, so it does not scale past the maintainer's own noticing.
  3. Link-mediated discovery — the agent, navigating tags, links, and descriptions, decides on its own that the change at hand needs the cluster. Best-effort by construction — retrieval failure is reflection failure applies to the cluster itself — so this path puts its load on the cluster's findability: descriptions that match change-time queries, links from the artifacts agents actually touch when changing the system.

Paths 2 and 3 already exist in weak form; path 1 does not exist at all. Which mix to build is a design decision this workshop owns, not a foregone conclusion in favor of maximal wiring — and the framework below reframes the choice: 2 and 3 are the fallback route of a two-layer system, not merely weaker wiring.

Framework: this is a two-layer build

The library already holds the theory of what this workshop is doing. Methodology with incomplete coverage and its live theory fallback form a two-layer execution system supplies the architecture, and the operationalization should be built as an instance of it:

  • The change-time digest planned below is a derived fast path — action-shaped, cheaper, and strictly narrower than the cluster.
  • The cluster is the generator layer and stays live: a change situation the digest does not cover drops back to the theory. Authority paths 2 and 3 are that fallback route.
  • Recurrence is the promotion signal: fallback reasoning that repeats is what earns a place in the digest. The digest's content is therefore not designed upfront — the phase-1 audit and the instrumented future changes are the first samples of fallback traffic, and they decide what gets promoted.
  • The digest must declare its coverage, so an out-of-coverage change routes to the cluster instead of being mis-served by a checklist that silently does not apply.

Wiring and maintenance are settled vocabulary, not new design: the digest is judgment-dependent derived prose, so it takes the managed-staleness regime, and the link grammar already authorizes exactly this edge — operationalized-from for the kb/notes/ methodology → kb/instructions/ procedure pairing, recorded at the source in an Operationalized into: footer, a source change flagging a judgmental recheck. That is also the phase-0 discipline going forward: theory revisions create no derived-artifact debt today (no derived layer exists yet), but once the digest exists, every later cluster revision owes it the structure note's correspondence check. The eventual mechanical support is already named too: where change candidates come from flags theory-to-implementation lineage as what a wider freshness substrate would need to cover.

External-method delegation (adopted 2026-07-22)

The delegation stance in the Goal is dispositioned item by item in external-delegation-assessment. Current state: host sources ingested (Bainbridge, Kephart & Chess, and the three Moen/Norman PDSA papers); prospective non-duplication in force — audit dispositions and the future digest cite PDSA/MAPE-K instead of restating them; host binding deferred until a worked case (one repository change through a PDSA overlay; MAPE-K only if a computational runtime pathway comes to exist); assurance-case/GSN inheritance deferred as a ledger candidate. Literature-offload-analysis assesses what the ingested sources let us delegate — existing notes shed almost nothing (they already practice conservative extension); the offload lands on text not yet written. The Moen/Norman snapshots are provenance for PDSA's history and logic, not an operating procedure: a worked case that binds PDSA as host should bind a canonical actionable specification (e.g. the Model for Improvement), snapshotted at that point.

Authority-path direction (adopted 2026-07-24)

The maintainer chose a distributed wire over an AGENTS.md-loaded digest: per-surface obligations live in the type specs and collection contracts that already load at authoring time, keeping the always-loaded AGENTS.md budget untouched. The enforcement chain still roots in AGENTS.md at zero marginal cost — its existing "read the target collection's COLLECTION.md before writing" rule is what loads the wired surfaces. kb/types/tag-readme.md (ADR 026's marks regime) is the precedent: cluster-shaped maintenance obligations in a type spec.

First increment, applied 2026-07-24:

  • kb/types/instruction.md — new Operativity section: name the consumer/channel/force path before writing or editing; description as retrieval wire matched to fire-time queries; route-or-record-the-gap for conditional instructions. Carries F1 for the most-edited surface.
  • kb/types/review-gate.md — new Force and warrant section: gates are problem-noticing, not reject-capable evaluation (audit F2 framing); no force-raising before labelled-fixture calibration (F3); judgment-changing prompt/process edits outside the freshness hash owe a deliberate corpus re-review (F4).
  • kb/instructions/COLLECTION.md — instruction-duality paragraph extended to state the operativity test at edit time and the inert-instruction failure mode.

Declared coverage and residue: the wire covers artifacts governed by a wired type spec or collection contract. Out of coverage — edits to COLLECTION.md files themselves, AGENTS.md, and src/commonplace/ — remain on the theory-fallback paths (user invocation, link discovery), per the two-layer framework. This direction does not foreclose a later digest; the type-spec paragraphs are the first promoted digest content, placed in distributed slots, and further promotion still follows recurrence in fallback traffic.

Maintenance: the paragraphs carry rationale edges into the cluster rather than operationalized-from lineage — that edge is authorized only for the kb/notes/kb/instructions/ procedure pairing, and kb/types/ is not in it. Until lineage or freshness covers this dependency, the correspondence obligation is recorded here: a revision to operative change, behavioral authority, retrieval failure is reflection failure, or the audit's F2–F4 dispositions owes these three surfaces a judgmental recheck.

Second increment, applied 2026-07-24: the ADR type contract (kb/reference/types/adr.md) requires decisions dated 2026-07-24 or later to name their operativity path (consumer, channel, force), and the proposals contract (kb/reference/proposals/README.md) requires the same per option, plus the oracle-warrant question for options adding automated evaluation — both dated forward, so the existing ADR and proposal corpora carry no retroactive conformance debt. This wire covers decided and proposed changes where the first increment covers authored artifacts, and instrumented ADRs are the ongoing evidence stream the cluster lacked.

Still open in this decision: whether the residue justifies extending the link grammar's operationalized-from authorization or a thin AGENTS.md line, and rationale backlinks from individual promoted skills. Record the completed mix as an ADR at workshop closure.

The Force-and-warrant section's re-review obligation was exercised on its own introducing change (2026-07-24): the type-spec edit changes the criterion side of every gate's type-conformance pair, so all 38 gates under kb/instructions/review-gates/ were re-read against the new contract. Outcome: 38/38 conform — the corpus was already detector-shaped, with no blocking, acceptance, or retention-authority language; no gate edits required. Three borderline fix-imperative phrasings noted and cleared (prose/proportion-mismatch, prose/source-residue, semantic/load-bearing-qualifiers) — revisit only if the contract is later tightened to forbid imperative fix directives inside tests. No freshness migration applied: this checkout carries no commonplace-store.sqlite, so there are no baselines to acknowledge (audit evidence boundary).

Sequencing (fixed by the maintainer)

  • Phase 0 — close ambiguities and contradictions. Work the ledger below. Items may be resolved in the library notes directly, parked explicitly as open questions, or split out; what they may not do is stay silently ambiguous into phase 1.
  • Phase 1 — DONE (2026-07-23). The operational-artifact audit covers kb/instructions/, src/commonplace/, validators, kb/types/, all collection contracts, and AGENTS.md. It dispositions every suggested change, repairs the fix pipeline's missing rejection vocabulary, corrects authoring/retrieval defects, and records the checkout's empirical evidence boundary.
  • Later phases (order now constrained by the audit): ~~decide the authority-path mix~~ (direction adopted 2026-07-24, first increment applied — see Authority-path direction above; residue routing and the closure ADR remain); distill a change-time instruction digest into kb/instructions/ as the two-layer fast path described under Framework, carrying operationalized-from lineage, a declared coverage region, and deliberate re-review for judgment-changing review-prompt/process edits; write up the two-layer promotion mechanism as the cluster's second worked case; address or reject the promoted-skill portability proposal — now subsumed as option A of the library-grade channel-compiled instruction artifacts proposal (2026-07-25), which reframes F8 from a portability defect to a build-time-resolution gap: channel is knowable at install time, so carrying it as prose branches is runtime parameterisation of the kind generate-instructions-at-build-time already rejects. Disposition is now a choice within that proposal, not a standalone accept/reject; ~~give the declared Commonplace frame a single citable anchor~~ (done 2026-07-22, ahead of sequence — see Resolved); instrument subsequent ADRs and significant changes as the ongoing evidence stream the cluster currently lacks. The audit has already completed the planned review-system mapping: gates are search/problem-noticing, downstream disposition is evaluation, commit/merge is retention, and freshness baselines are evidence bookkeeping.
  • Parasuraman inheritance — DONE (2026-07-21). Both steps executed by the maintainer: the paper is ingested, and the closure note's actor-allocation section carries the form-inheritance paragraph with all three departures (no within-function ladder, improvement functions rather than task-performance stages, allocation does not establish warrant). The ingest's independent read corroborates the stance — it summarizes the paper as "a multidimensional allocation profile, not a scalar autonomy ladder," and its Limitations note that the ten-level scale is validated mainly for decision selection strengthens the no-within-function-ladder departure. The self-aware-computing coverage comparison (external-theory evaluation, finding 7's other half) stays separate and undecided.

Ambiguity ledger (phase 0)

Questions that still need theory work are explicitly parked in their owning notes rather than silently carried by this workshop. One post-audit routing question remains open:

  • Reflective-systems axis factoring — the self-improving-systems tag head dropped its complete mark and went selective (2026-07-23; the spec's default exit after crossing the weight gate). The phase-1 audit did not disposition whether the reflection cluster's traffic justifies an independent reflective-systems tag with its own head. Decide this from routing evidence, not size pressure: reflectivity is an independent axis, so this would be axis factoring with co-occurrence expressing the intersection, not a child split.

Resolved

  • warn outcome vs. the loop's evaluation criterion (2026-07-23) — resolution A adopted in warn-outcome-vs-loop-evaluation: gates are search/problem-noticing, downstream fix/reject/defer disposition is evaluation, commit/merge is retention, and freshness baselines are evidence bookkeeping. The fix-report contract now represents rejected separately from deferred.
  • Closure-note TODOs (2026-07-22) — both resolved in the closure note on one shared ground rather than two special rules. Tacit-but-stable expertise is not retained methodology — it is a promotion candidate (stable-but-unexternalized practice, noticeable by recurrence, convertible by externalization). Row 3's substrate dependency is a boundary/coverage fact (selection-grade coverage of a sealed parametric component), name-and-excluded from the actor-allocation reading the same way organizational closure is. The shared ground is a new library note distilled from a maintainer brainstorm: only explicit retention is currently durable, writable, and addressable at once, with a companion, retaining the episode keeps a distilled rule re-derivable, extending it to the episode/rule retention choice. Externalization-as-transport also landed in the closure note's conversion section.
  • Operation-depth ontology (2026-07-22) — resolved in the coverage note: the second coverage dimension is an operation profile — the set of operations that hold per component — not a depth ladder. Both scalarizations were considered and rejected in the note itself: ordering by reach fails on the note's own cases (selection without observation, configuration without selection), and ordering by count of covered operations (maintainer's candidate, dismissed as not a good measure) collapses capabilities that differ in kind. The profile sense aligns with the cluster's existing anti-ladder moves (pathway profile, allocation profile); rank-talk survives only in its negative use — naming the operation a lever lacks. Vocabulary swept across the live consumers ("operation depth" → "operation profile", "modification-depth" → "modification-grade"); frozen error-catching fixtures left untouched by convention.
  • Cumulativity's environment-mediated exclusion (2026-07-23) — verified in validation-freshness-and-code: freshness baselines and generated retrieval marks directly consume retained prior state; an ordinary edit that only changes the later environment does not qualify by that fact alone.
  • "Promotion" vocabulary watch (2026-07-23) — no harmful collision found. Case promotion, artifact promotion, and authority promotion share a movement-toward-durability sense while their objects and channels disambiguate the operational decision. Reopen only if a write-time case makes two readings actionable at once.
  • Stub note (2026-07-22) — folded: "Reflection puts a system's own organization inside its action environment" is deleted; its claim and three open questions now live in the coverage note's Open Questions, and the five inbound links are repointed. If the phase-1 audit shows a live need for the self-ontology investigation, recreate a narrower note from evidence.
  • Canonical frame (2026-07-22) — the declared Commonplace frame now states the boundary once, citable from any application; Commonplace as a reflective system defers to it instead of carrying the only copy. The theoretical anchor (frame-indexed predicate) was already in the definition's Provenance; this closes the descriptive half.
  • Membership tense (2026-07-21) — resolved in the definition: membership is read over a declared assessment horizon, mirroring operativity; the dispositional attribution (a standing pathway, exercised or not) stays available but must be marked as a different claim. A matching misuse case guards the dormant-pathway reading. From external-theory evaluation, finding 2.
  • "Closure" collision with the autopoietic sense (2026-07-21) — contrast paragraph added to the closure note pointing at the reflective-system exclusions, per the write-time collision rule. From external-theory evaluation, finding 5.

  • Three senses of "reach" (2026-07-21) — explanatory-reach and reach-assessment registered as rare compounds carrying the main sense; registration paragraph lives in reach-assessment (retitled to the hyphenated form). The search sense renamed to search range; the one oracle-side leak ("proof reach") renamed to proof surface. Deliberate residue: the two route-note titles and the first-principles title keep spaced "reach" (title quotes lag by convention; rename only if a title-level collision ever bites), and frozen fixtures under kb/work/error-catching/ plus source-ingest link texts were left untouched.

What closes the workshop

  • The phase 1 audit executed, with every suggested change dispositioned: applied, turned into a proposal, or rejected with reasons recorded.
  • An explicit decision on the authority-path mix (which of the three paths carries the cluster, and what was built or deliberately not built for each).
  • The ledger empty — each item resolved in a library note or explicitly parked there as an open question.
  • Durable outputs extracted to the library (instruction digest, reference mapping, ADRs as warranted) and this directory deleted.

Bookkeeping

Findings and drafts live as files in this directory; phase 1 per-artifact audit results go under audit/. Resolutions land in the library notes they concern, with only pointers kept here.


Links:


Complete file listing (generated at build time)