Workshop: Popperian Maintenance Episode

Goal

Record the 2026-08 history of kb/notes/llm-output-deviation-requires-three-way-diagnosis.md as a worked episode of Popperian maintenance — conjecture, refutation by adversarial review, guarded reframe, re-earned empirical content — and extract its design consequences for the repair machinery.

The episode matters beyond its note: it is the first end-to-end trace in this repository of a bold causal claim being defeated by the premise-decomposition gate, repaired under the narrowing guards, drifting toward an analytic claim, and then re-earning empirical content through a worked comparison. It is also the trigger case for two questions the library did not yet hold: whether the repair-disposition vocabulary is missing an idealization option, and whether repair policy is a per-installation profile choice. A second witness arrived the same day from a parallel session — the class-instance-analogy defeats, where the counterexamples are domain-priced exceptions — and supplied the operational honesty test the first case lacked. This workshop is the coordination point for the idealization thread across sessions.

Threads

  • episode-record.md — the first witness reconstructed: timeline, the three acts, what the gates defeated and what held, the content accounting, and the drivers
  • second-witness-class-instance.md — the second witness: three GLOBAL defeats whose counterexamples the source domain prices (reflection APIs, monkey-patching, JIT costs, governance rituals, prototype OO as rival paradigm), the domain-pricing honesty test, and the two-witness contrast that shows the missing assessment route discriminates
  • third-episode-criterion-note-reframed.md — the third episode, recursive: pass 2c3150 defeated the criterion note's own pricing-sufficiency premises (break-glass; use-relative schema migration), the packet was superseded by the version guard, and the reframe to routing-versus-adequacy landed by direct revision; the proposal's pricing-gated verdict was corrected pre-adoption
  • fourth-episode-in-pass-assessment.md — the fourth episode: pass 2a6408's first in-pass adequacy assessment, with acceptance as the outcome, the refused "common and cheap enough" gloss as comprehension evidence, the advisory/authoritative separation holding (independent convergence on the scope repair without warrant-borrowing), and the unactioned residuals
  • fifth-episode-critique-without-repair.md — the fifth episode (2026-08-26): two same-day passes on one Naur note, each right as critique and wrong as repair in the same way — retreat to the nearest untouchable claim ("retained passages do not establish", "does not by itself bind") scored preserved/strengthened against the reframed update; fixed twice by one operator sentence supplying the note's point. Diagnosis: motivation is not a pass input and evidence cannot be extended, so repair is subtractive; the pass is Naur's group B. Candidate consequences: a return-to-author disposition, a retained-intent reference point, grounding before reframe, an evidence-scope test. Record (six versions, both packets, reports) in fifth-episode-record/
  • sixth-episode-critique-defeats-by-equivocation.md — the sixth episode (2026-08-26): pass 3fa27b reframed the pre-formal-stage note on four GLOBAL defeats whose counterexamples all place the concept inside an admitted formal language the note's antecedent excludes; the bite rule would not have caught it (premises were DEFEATED, not DOUBTFUL). Repair: thesis restored with the antecedent sharpened, and a clause added to the bite rule — a defeat bites only if its counterexample meets the antecedent under the note's own definitions. Record in sixth-episode-record/
  • genre-drift-cohort-result.mdthe cohort question, closed 2026-08-21 and refuted as posed. Scored across the 23 series runs: 16 changed a title, of which 5 drifted toward diagnostic/procedural, 5 held genre, and 3 moved the other way. The original count had no denominator — only drift-positive cases were being tallied, and two separate records each claimed to be the thread's "third instance". What the corpus does show systematically is a different trade: 15 of the 16 title changes buy survival with conditions and negations while keeping the claim's kind, which is what counterexample-shaped defeats route to. Genre follows the warrant rather than attracting it
  • First in-pass run of the two-stage assessment (completed 2026-08-19, all three readouts positive) — pass 20260819T105132Z-2a6408 (codex, independent of the sessions that authored the pricing and the commitment) ran against the partial record as committed. Readout 1, attacked on merits: the initial premise check hit fence integrity exactly ("routine monkey-patching... less cleanly fenced"), returning DOUBTFUL/LOCAL; after body edits the closing check returned all six premises HOLDS, and the report keeps the falsifier live ("evidence that definition change is ordinary and unmarked... would dissolve the fence"). Readout 2, open dimensions stayed open: no GLOBAL defeat, disposition plain keep — absence of prevalence and behavioural-share evidence was not converted into defeat. Readout 3, pricing-only acceptance refused: the warrant states "pricing signatures route... they do not prove adequacy" and judges adequacy on fence integrity for the declared use; the pass even rejected a reviewer's suggested gloss because it would misstate the criterion. Bonus finding: the assessment ran inside the existing pass machinery with no new verdict, trait, or gate — the declared commitments were tested as ordinary content. The proposal's last adoption criterion is met; what implementation still adds is the conversion disposition, the mandatory deferred assessment, and the policy declaration
  • Attestation sourcing (closed 2026-08-19) — the ingests landed in kb/sources/ (Kiczales MOP, monkey-patch, V8 fast-properties and hidden-classes, Erlang release handling and code loading, Ungar & Smith Self). The class-instance session wired the instantiation note's four pricing clauses inline plus footer edges (6f87bed7); this thread wired the criterion note's evidenced-by edges and has-external-sources trait from the connect report. The author-asserted-pricing objection is closed for the routing stage on both notes
  • Candidacy-evidence generalization (open, candidate) — the connect pass surfaced a synthesis opportunity: the criterion note, the theory-warrant non-distribution rule, the discovery lifecycle, and per-operation trace-memory authority jointly suggest a general claim — cheap candidacy evidence may justify escalation to an expensive assessment without carrying any verdict authority. Would generalize routing/acceptance beyond idealization to staged review and promotion; no existing note states it. Needs its own worked case beyond this thread before writing
  • Pricing as design vocabulary (open) — extension candidate from the class-instance session: when an idealization passes the pricing test, the domain's pricing apparatus (its fence) is a candidate solution structure for the analogous problem in the target domain — the fenced-second-relation result is the one witness. Recorded as an open question in the criterion note; a claim of its own if a second case lands
  • Maxim reconciliation (open) — the operator maxim "the goal of our theories is to become methodology" reads as theory-consumed-by-methodology, while theory and methodology form a two-layer execution system keeps theory live upstream as the generator layer. Decide whether the maxim should be restated ("theories generate methodology and remain upstream of it") or the note revised.

Adopted (2026-08-19), validated (2026-08-21)

ADR 066 shipped the machinery this workshop was opened to design: three claim modes declared in text, modality-aware premise reading with counterexample-shape annotations, and mode-guarded bidirectional reframes in the full pass. The proposal was trimmed to its undecided remainder (profile-level policy, attestation drift tracking).

The validation series in adr-066-test-runs.md then ran to 23 passes across two waves, covering 21 of the survey's 22 candidates plus the anchor case and a no-fire control. It found something better than the result it was designed for: mode conversion is rare because mode landings require mode-appropriate warrant — prevalence evidence for statistical, domain-priced exceptions for ideal-type — and most mismatched notes lack it. Their honest repairs are scope, category, and conditionality reframes. Both guards have bound and rejected (statistical at run 9, ideal-type at run 14's adequacy check), so the machinery is not permissive; it has not yet been offered a conversion it could accept. The survey's per-note class labels in statistical-mode-candidates.md were falsified as predictions wherever they were tested, including the case scheduled last as the strongest ideal-type candidate — read that file as a record of what was predicted, not as a queue.

Run 23 added a third machinery response to an unlabelled idealization, alongside conversion and adequacy-guard rejection: strip it as unearned formalism. The bounded-context note's idealized scalar cost measure drew a pseudo-formalism FAIL from the prose gate and was replaced by a task-relative feasibility predicate. That route reaches no modality vocabulary at all, which is why the guards never saw it — worth watching, because it means the prose gate can dispose of an idealization before the modality question is asked.

  • kb/reference/proposals/repair-dispositions-for-defeated-claims.md — the design object: make the repair policy for review-defeated claims explicit, with an option space (document the current universal, add a declared-modality/idealization repair with pricing-gated verdicts, or declare policy in each installation's local notes contract)
  • kb/notes/domain-pricing-routes-an-exception-to-idealization-assessment.md (promoted under its original title, a-domain-priced-exception-does-not-refute-an-idealization, and reframed by the third episode's pass) — the transferable requirement the proposal rests on: the pricing test, its signatures, and the constraints that keep idealization from becoming an immunizing stratagem

What closes this workshop

  1. The episode record's durable conclusions extracted or deliberately left as evidence: candidate destinations are ASIS&S 2026 paper material (the paper is distilled from the notes, so lineage runs paper-from-notes) or an article; the record itself is not promoted, workshops being sinks. Still open.
  2. ~~The cohort question answered or recorded as a log FIX~~ — done 2026-08-21, refuted as posed and closed in genre-drift-cohort-result.md. The maxim reconciliation is still the operator's to decide, and it is the only open item blocking this condition.
  3. The proposal and note exist in the library (done at open) and the proposal's fate is the proposals frontier's business, not this workshop's. Done.

Remaining before deletion

  • The maxim reconciliation — operator decision, one sentence either way.
  • The episode record's disposition — decide whether the four episode files feed the ASIS&S paper, feed an article, or are simply retained as evidence and deleted with the workshop.
  • Two candidate threads with one witness each — candidacy-evidence generalization and pricing as design vocabulary. Neither has earned a note. If no second case has landed when the workshop closes, they belong in kb/log.md as ABSTRACTION entries, not in a workshop kept alive to hold them.
  • Not blocking: the unrun candidate (knowledge-storage-does-not-imply-contextual-activation.md, 211 citers) and the two unobserved coverage gaps. Both are candidate-starved rather than unresolved, and the series log carries the reasoning; neither needs a workshop to stay open.

Evaluation boundary

The episode is evidence about this installation's review machinery operating under this installation's goals. Generalization to other installations goes through the promoted note and proposal, not through this record — the record documents what happened, including the parts (reader model, disposition policy) that are installation-specific.

Bookkeeping

Plain markdown, workshop register. Started 2026-08-19 from a conversation reconstructing the note's history from git log, the pass report kb/reports/state/full-pass/llm-output-deviation-has-three-sources-with-non-substitutable/20260818T132531Z-e11d99/full-pass-report.md, and kb/log.md. Positions attributed to "the operator" are where that conversation landed.


Complete file listing (generated at build time)