Automation-boundary retrospective: how this theory was found
The generator theory in this workshop was produced by a human–agent process worth examining with the same discipline as the experiments themselves: linking is core technology in this KB, so the search was reflective — the system theorizing about its own connective tissue — and the episode is a worked case of a theory search with explanatory reach as the goal. This file separates what the process record shows from interpretation, and asks what could be automated versus what still required human input.
The arc (established record)
- (Amended 2026-07-29, from maintainer recollection.) The maintainer noticed a problem with linking — the specific observation is no longer recallable and was not recorded — and requested a review of the linking machinery. The audit below was commissioned, not spontaneous.
- The commissioned cross-collection audit found contradictions among linking contracts, shared vocabulary, and procedures (the consistency workshop's trigger).
- Reconciliation ran through four label reviews (evidence, rationale, grounds, mechanism). Each adjudication needed principles that did not exist; the mechanism review made that unavoidable, and this workshop opened.
- The maintainer identified the missing piece: the inherited Ars Contexta theory justified having typed links but was not generative — it could not produce the vocabulary appropriate to a given collection, a limitation known since the vocabulary failed for the reference collection (repaired by ADR 019).
- The brainstorm produced competing models and the seed-then-harvest generator sketch; six experiments (three retrodiction variants, an A/B, two corpus checks plus a control) tested it in one day, recorded in the retrodiction run.
The episode maps onto the discovery lifecycle almost stage for stage: observe (contradictions, migration data) → conjecture (competing models) → derive consequences (retrodiction predictions) → test (variants, A/B, corpus checks, discrimination control) → accept/integrate (still pending: the handoffs in the run file's Next section).
Division of labor observed
Done by agents, effectively automatically (the maintainer's inputs were one-line "run it" messages, each selecting an item the agent had already proposed):
- synthesizing ~700 classified edges and four ADRs into the repeated observations;
- generating the competing models and the generator sketch;
- designing and executing every experiment: blind protocols, deterministic leak checks, control arms, scoring rubrics, k-sampling;
- self-refutation: the A/B killed the agent's own revision-consequence hypothesis; the portfolio finding corrected the agent's own "three tries, three families" narrative — error-correction inside the loop needed no human push;
- forward prediction (the three gap candidates) and the discrimination control that validated the forward mode.
Required the human, on the evidence of this session:
- the arc's step 0: noticing the initial anomaly and commissioning the review at all — the query that launched everything was human-formed;
- the original noticing that contradictions pointed at a missing theory rather than at more reconciliation;
- supplying the load-bearing historical fact — that the Ars Contexta inheritance was central and non-generative — and its emphasis, via one corrective message after the first brainstorm draft underweighted it;
- adoption authority, deliberately withheld throughout: no label registered, no contract changed;
- pacing and stopping.
The asymmetry is stark: the agent's contributions were voluminous and fast; the human's were two or three sentences — but the sentences were ones no agent produced.
The central puzzle: why agents did not discover the inheritance's role
The maintainer's question: linking is core technology here, the Ars Contexta inheritance is its origin story — why did the agents not surface its crucial role themselves?
Layer 1 — a capture gap, now verified. The durable corpus does not contain the fact. ADR 009 records the adoption; ADR 019 records the repair (collection-owned vocabularies) and even concedes in passing that new collections must design vocabularies by hand; linking-theory.md asks only an open question about the borrowed types. Nowhere does any artifact state the claim itself: the inherited theory is not generative; attempts to derive per-collection vocabularies from it failed. A search for generativity language across those documents returns nothing. The failed derivation attempt — the event that defines this workshop's problem statement — was never an ADR (no decision), never a note (nobody wrote the claim), never a log entry. It survived only in the maintainer's episodic memory, entering the corpus for the first time as a grounding bullet the maintainer dictated into the brainstorm brief. The KB records decisions and repairs; it does not record failed theory searches — and failed theory searches are exactly the problem statements of future ones.
Layer 2 — a valuation gap, not a retrieval gap. Even with the fact in context (the brief's bullet), the first brainstorm draft cited it, built one conclusion around the missing process, and still treated it as one grounding observation among ten rather than as the central explanandum. The brief presented its grounding as a flat list, and flat lists erase importance gradients; the agent weighted what was dense, recent, quantitative, and locally verifiable (the 700-edge migration record) over what was sparse, old, narrative, and counterfactual (a failure that left no artifact). Recognizing the inheritance as a completed natural experiment with a readable result — adopted universally, worked where endpoints were propositions, broke where they were not — requires assembling significance across ADR 009, ADR 019, and an absent artifact. Nothing in the corpus or the procedure prompted that historian's move.
Capability was not the limit. One corrective sentence from the maintainer ("the rules we inherited... we could not create a generator out of it") produced, within a single turn, the full diagnosis — the propositional vocabulary's hidden endpoint-kind parameter — and the two-stage generator that the experiments then validated. The agent could do everything except decide, unprompted, that this was the thing to explain.
A reflexive symmetry
The experiments concluded: the seed generates the semantic skeleton; authorization is selection, and selection belongs to the corpus record and the humans who own it. The meta-process that produced the theory has the same shape: agents generated the models, experiments, corrections, and candidate conclusions; the human supplied significance and holds adoption. The boundary the theory located inside link vocabularies — generated semantics versus selected authorization — recurses onto the theory-building process itself. A second reflexive loop: the normalized footer grammar (protected by brainstorm conclusion 2) is what made the migration data enumerable, which is what gave the agents anything to theorize from — the KB's self-correction capacity was not just claimed in this workshop but exercised by it.
Candidate remedies (generated, not adopted)
- Record framework failures as first-class claims. When work repairs around an inherited framework (as ADR 019 did), the framework's limitation should be captured as a claim-note at that moment — "X is not generative for Y" — not left implicit in the repair. This is a capture rule; it would have put the load-bearing fact within
rg's reach years before this workshop. Route: a note plus possibly one line in ADR-writing guidance. - Rank grounding by explanatory demand. Brainstorm briefs/procedures could require, before modeling: rank the grounding observations by "which is the biggest unexplained fact?", and address the top-ranked first. This is an automatable valuation repair for the flat-list problem. Route: an addition to whatever brainstorm instruction gets promoted from this workshop. (Tested 2026-07-29 — see "Remedy 2 tested" below: the step works, with a refinement — it succeeds by extracting the subsumption structure among the grounding facts, not by elevating the gap-statement bullet, which all rankers correctly demote.)
- Treat maintainer memory as an un-ingested source. At workshop opening, elicit explicitly: what did we already try that failed, and where is that recorded? Anything answered from memory rather than from an artifact is a capture debt to pay before theorizing. Route: workshop-opening checklist.
- Scheduled random audits automate the commissioning half of query formation (added 2026-07-29, maintainer-proposed). The arc's step 0 separates into noticing (ambient, human, during use) and commissioning (launching the review) — and this episode shows the discovery needed only the second: once commissioned, the agent found the contradictions unaided. For anomaly classes discoverable by audit — contract contradictions, drifted claims, stale assumptions, vocabulary incoherence — a periodic, randomly targeted, subsystem-scoped open-ended review substitutes for the noticing that never happens. The review system's open-ended assays are the existing mechanism one level down; QA sampling and chaos engineering are the inherited ontology (random probing where failure locations are unpredictable and checks are expensive). It fills the wire gap between exhaustive validators (cheap checks, every commit) and human noticing (ambient): questions too judgment-heavy to ask always, too important to ask never. What it does not automate: valuation of findings and adoption. Earning criterion: the hit rate of pursued findings against this arc's one datum (commissioned audit → contradictions → validated theory). Route: a design proposal once the workshop closes; the sampling policy is itself machinery needing warrant.
- What stays human, for now: deciding which unexplained fact the project should care about (goal-level significance), adoption/authorization, and stopping. The efficient design this episode demonstrates is not "automate the human away" but the comparative-advantage split: agents generate broadly and self-correct cheaply; the human spends a few sentences of rare, high-leverage steering — provided the capture rules above stop those sentences from being the only place critical history lives.
Correction to Layer 1 (2026-07-29, maintainer-prompted, verified against the ADRs)
Layer 1 as written overclaims a capture gap. The durable corpus does contain the load-bearing fact, distributed across the two most findable artifact types the KB has: ADR 009's description line states the inheritance outright ("borrowed from arscontexta and adapted"), and ADR 019's Context records the failure in substance — "label theory is weak… a single universal vocabulary commits every collection to the same reader-need taxonomy regardless of what its readers actually want" — with the Harder section conceding that new collections cannot fall back on a default and must design vocabularies up front. Inheritance, failure mode, repair, and the hand-design concession were all recorded. What the corpus lacks is only the named conclusion ("the inherited theory is not generative"), not the facts that entail it.
The original Layer 1 verification method exemplifies the failure it diagnosed: "a search for generativity language returns nothing" is an rg for the conclusion's vocabulary — a term nobody had coined yet — where assembly from recorded premises ("where did our link vocabulary come from and what did its failures teach?") answers the question from ADR 009 + 019 directly.
Reclassification: the failure was search-side, not capture-side, and it decomposes the retrieval pipeline into stages — query formation (the historian's question was never asked; motivation, not capability) → execution → assembly → in-context valuation (Layer 2, which stands unchanged). The episode failed at formation and valuation; capture, execution, and capability were fine.
This re-sorts the remedies into a per-stage repair kit, which is a cleaner structure than the original framing: remedy 3 (elicit "what did we already try and where is it recorded?" at workshop opening) repairs query formation and is the load-bearing one; remedy 1 sharpens from "record framework failures" to "name the conclusions your repairs imply" — a findability repair that puts assembled conclusions within lexical reach (frontloading applied to conclusions, per minimum-viable-vocabulary); remedy 2 repairs valuation. Consequently the promotion-worthy claim "failed theory searches are systematically under-captured" loses this episode as its instance and needs a genuine capture case. Two candidates exist: the scaffolding-relaxation workshop's recedes-then-reappears distinction, recorded nowhere until flagged at closure and independently re-arrived at externally on 2026-07-28; and this arc's own step 0 — the triggering linking anomaly was never recorded and the maintainer can no longer reconstruct it, so the arc's origin exists only as the fact that a review was commissioned. Triggering observations join failed searches in the under-captured class. The valuation claim gains a second, closely parallel episode from the same day: the Meta-Harness ablation audit, where the winning arm's retained-reports regime sat fully recoverable in released code and its significance moved only on a maintainer's push.
Remedy 2 tested (2026-07-29): the ranking step retrodicts the missed synthesis
Five fresh predictors were given only an isolated copy of the brainstorm brief — no corpus access, no workshop files — plus one preliminary instruction: enumerate the grounding observations, rank them by explanatory demand using the subsumption test ("which observation, if explained, would explain the most of the others?"), name the single central explanandum, and stop before modeling.
Tally across the five runs:
| grounding observation | ranks (runs 1–5) |
|---|---|
| endpoint-kind ambiguity (documents / claims / systems) | 1, 1, 3, 1, 1 |
| borrowed vocabulary worked for theory, failed for reference | 3, 2, 2, 3, 2 |
| one inherited label covered several assertions | 2, 3, 1, 2, 3 |
| Ars Contexta supplies no generative process | 6, 4, 8, 4, 5 |
Four of five runs produced, as their central explanandum, essentially the endpoint-kind diagnosis — the insight that in the live episode arrived only after the maintainer's corrective message. One run stated the hidden-parameter argument nearly verbatim: "theory notes' endpoints are mostly claims, reference notes' endpoints are documents/systems, so one propositional vocabulary can't cover both." Each run cost under a minute; the live episode needed a human turn to reach the same place.
Interpretation, with one refinement to the remedy's original claim. The remedy as first written predicted the inheritance story would "top such a ranking." Its gap-statement bullet (no generative process) did not — every ranker demoted it to mid-tier, correctly classifying it as a statement of missing method rather than an unexplained fact. What tops the ranking instead is the explanation-shaped conjunction: endpoint ambiguity first, with the transfer failure and label decomposition as its top-3 evidence in all five runs. The subsumption test is what does the work: it forces assembly of the flat list into a dependency structure, and the synthesis the live episode was missing falls out of that assembly. So the fix repairs valuation exactly as the corrected Layer-1 analysis locates it — and it repairs valuation by demanding structure, not emphasis.
Bounds. (a) Placement is load-bearing and untested in the hard direction: these rankers saw only the brief — the hypothesized distractor (the dense migration record) was absent, so the step's survival under full corpus immersion is unestablished; run it before opening the corpus. (b) The step composes with, and cannot substitute for, the capture-side remedies: this brief's grounding was written by the maintainer after the problem was already half-understood, and ranking can only elevate what the grounding contains — a brief missing the transfer-failure bullet has nothing for the subsumption test to find. (c) Five runs, one brief, one model; and the rankers' unanimity partly reflects that this brief's grounding really does have one dominant subsumption structure — a brief with two rival explananda would test the step harder. (d) It repairs valuation only; query formation (remedy 3's territory) is untouched.
Status
Analysis of one episode; the remedies are candidates for the workshop's closure outputs, alongside the theory itself. The promotion-worthy claims flagged here — subject to the correction above: conclusions of repairs go un-named rather than facts un-captured; agent theory-building fails on query formation and valuation before it fails on retrieval execution or capability; the generate/select boundary recurses from the object theory onto the process that built it — should be tested against at least one more episode before any becomes a note. The valuation claim now has its second episode; the naming/capture claim still needs one. Remedy 2 has a positive first test (see above): the valuation failure is repairable by an automatable procedure step, which upgrades "fails on valuation" from a diagnosis to an engineering target.