ADR 066 test runs: first passes under the modality machinery

Six full passes, chosen from statistical-mode-candidates.md so that together they exercise every path ADR 066 added: both reframe directions, both new mode targets, both landing guards, the premise gate's counterexample-shape annotations, and one control where the machinery must not fire. Run sequentially (the pass's concurrency precondition; also each run's readout may adjust expectations for the next). Independent runner preferred, per the 2a6408 precedent. Route each report's readout back into this file.

Protocol per run

Ordinary run-full-improvement-pass-on-note.md invocation — the point is that the standard pass now does this work; no special harness. Record per run: (a) did the premise report carry shape annotations, and were they accurate; (b) did step 7 name a target mode where warranted, routed by those shapes; (c) did the landing meet its guard (stated refuter / adequacy record present before step 9); (d) did the closing premise rerun attack any new adequacy record; (e) were reframe follow-ups (rename, citer reconciliation) recorded as Open items.

The runs

Run 1 — statistical reframe down, easiest case. COMPLETED 2026-08-19 (pass 20260819T125030Z-s5qz) — prediction wrong, machinery coherent. kb/notes/task-fitted-structure-costs-cross-task-reuse.md → renamed current-task-fit-alone-does-not-warrant-costly-entrenchment.md The predicted statistical landing did not happen, and rightly so. The premise gate's shape annotations fired (instance on both non-HOLDS premises — a cheap-additive index defeating the exhaustive-warrants premise GLOBAL, a scope-determination edge case LOCAL); with no prevalence- or priced-exception-shaped defeats, no mode conversion was routed, and the pass instead found the warranted claim in the note's warrant structure: a universal insufficiency claim ("current-task fit alone does not warrant costly structural entrenchment"), retitled by ordinary keep-reframe. The survey had misread the body hedge ("often invisible, rarely revisited") as the core claim; the pass located the core elsewhere, and the hedge survives as a description of the bet's visibility, not as the thesis. Readouts: (a) shapes present and accurate; (b) no target mode named — consistent with the routing, since the shapes gave no mode signal; (c)/(d) mode guards unexercised this run; (e) follow-ups recorded and executed (relocated with redirect, four citers reconciled — including the coordination-value definition, whose gloss had become false the moment the reframe made coordination the rule's third warrant). Bonus behavior worth keeping: the closing premise rerun surfaced a new GLOBAL defeat (a temporary deadline with present stakes is a fourth warrant the packet's exhaustive formulation omitted), and the pass correctly routed it without another edit round — the reframed insufficiency title survives it because "alone does not warrant" is not an exhaustiveness claim. Series consequence: run 1 turned into an unplanned second no-false-fire datum — the machinery declined a mode conversion on a note we expected to convert. Statistical-guard coverage now rests entirely on runs 3 and 5; if run 3 also lands off-mode, promote structure-activates-higher-quality-training-distributions.md (numeric prevalence core, survived null already in the body) into the series immediately.

Run 2 — ideal-type conversion. kb/notes/agent-runtimes-decompose-into-scheduler-context-engine-and-execution.md The hedge "in many real systems the boundaries blur; the claim is that the functions are analytically distinct" is an undeclared first-order model. Expected: keep (title may stand) with a body edit converting the hedge into a declared idealization plus adequacy record — declared use (what the decomposition is for: predicting which limitation a change fixes), omitted mechanism (implementation blurring), bound, dominance — and the closing premise rerun attacking that record in the same pass. Failure tells: conversion without the record (immunization guard missed), or the record present but the closing premises never engaging it.

Run 3 — upward reframe from vacuity, the hard test. kb/notes/the-framework-is-often-larger-than-the-durable-contribution.md "Often" in the title, "tends to" in the body, no refuter anywhere — the purest Class B case. Expected: reframe that lands on a guarded claim — a statistical form stating what measured framework-to-contribution ratio would refute it, or a stronger conditional the material warrants. This run tests whether the machinery repairs the ratchet's end state rather than reproducing it. Failure tell: the pass keeps or produces another unguarded tendency.

Run 4 — upward reframe to universal. kb/notes/memory-backed-personalization-can-look-like-model-improvement.md A bare possibility title with a sharp universal core buried in paragraph two ("it cannot make one of several prompt-compatible commissions authoritative without user-specific evidence"). Expected: keep-reframe up, promoting the refutable core to the title. This is the direct test of bidirectionality — before ADR 066 no repair path could strengthen a claim.

Run 5 — mixed modality in one note. kb/notes/agent-context-is-constrained-by-soft-degradation-not-hard-token-limits.md The anchor case: a statistical binding claim (title says "not hard token limits"; its own description already retreats to "not just") over an ideal-type mechanism model (the three-dimension decomposition plus workspace hypothesis, self-described as "not fully separable" and "a working hypothesis"). Expected: the pass assigns different modes to different claims — statistical retitle with a stated refuter for the binding claim, declared idealization with an adequacy record for the mechanism section — without flattening one into the other. The hardest coordination test; a defensible lesser outcome (fixing one mode and routing the other to Open items) is a finding, not a failure.

Run 6 — control: a sound universal the machinery must leave alone. kb/notes/an-outcome-check-licenses-replay-a-rule-needs-the-process-verified.md Deductively argued ("a rule asserts 'do X because Y'; an outcome check never inspected Y"), correctly universal, not in the candidate survey. Expected: no modality finding — shape annotations may appear on any dented premise, but no mode reframe fires. Failure tell: the pass invents a statistical or ideal-type reading for a claim whose warrant is deductive. A machinery that fires everywhere is as broken as one that never fires.

Readouts, runs 2–6 (all passes completed 2026-08-19)

Run 2 (pass 63b341) — ideal-type conversion DECLINED, and rightly. Keep-reframe, but to an analytical-classification claim ("Agent-runtime analysis should separate scheduling, context assembly, and external state"; relocated), not a declared idealization. The shapes discriminated: premise 4 (practitioner convergence) DOUBTFUL GLOBAL prevalence — and with only two mapped taxonomies, a statistical landing had no warrant, so the convergence claim was removed, not moded. The ideal-type routing test failed honestly: implementation blurring is ordinary unmarked practice in the runtime domain — no marked interface, no pejorative, no charge, no ritual — so by the criterion note's own standard, the survey's Class C diagnosis was wrong and the pass's decline was right. Notable: the pass's own closing cycle caught the step-9 revise strengthening "distinct failure questions" into unwarranted causal fault localization and flagged its warranted contribution as changed — plus a fresh closing GLOBAL defeat (opaque joint controllers), both routed without a second round. Genre-drift datum: an ontology title ("runtimes decompose into") landed as a methodology-shaped one ("analysis should separate") — third instance for the cohort thread.

Run 3 (pass 0d5090) — vacuity repaired, landing on a categorical rule, not another tendency. The prevalence annotation fired on four premises (framework familiarity, leakage tendency, cue-activation reliability, author detection) and a genuine priced-exception on one (a regulated clinical domain charging missing warrant more than surplus — the annotation vocabulary used exactly as designed). With no prevalence evidence available, the statistical guard correctly blocked a statistical landing, and the reframe went up in form: a scoped conditional retention rule. The refuter discipline is explicit in the packet's open items: "Do not restore 'often,' 'usually,' or 'tends' without a declared statistical comparison and evidence that could refute it." Rename deferred by the packet, then resolved: the closing cycle GLOBAL-defeated the new title's universal anchor requirement (an enforced consumption path can supply the recognition), and the operator approved the narrowing 2026-08-19. Applied: retitled to "A linked note's durable payload is what its consumption path cannot reliably supply" (relocated with redirect to linked-note-durable-payload-is-what-consumption-path-cannot-supply.md), the consumption-path condition written into the opening and the recognition rule, citers reconciled. CLOSED.

Run 4 (pass k7p4) — upward reframe correctly DECLINED. Plain keep. The buried universal this series planned to promote ("only user-specific evidence can make a commission authoritative") was DEFEATED LOCAL instance by an organization-wide authoritative instruction — promoting it would have shipped a false title. The note's "can look like" is not vacuous: it is a witnessed possibility claim whose contribution is the attribution consequences, which the pass sharpened (carrier-dependent estimands, split swap prescriptions). Lesson: possibility-with-witness is a legitimate form the Class B diagnosis conflated with vacuity.

Run 5 (pass 1f5b0b) — THE STATISTICAL LANDING, guards met. Keep-reframe with the target mode named in the disposition and the guard stated verbatim: "the claim would fail if representative workloads usually stayed reliable until the cap, or if inability to fit required evidence were usually the first constraint." Routed from premise 2 DEFEATED GLOBAL instance (a real hard-cap-first workload: two pruned corpora exceeding the window). The mixed-modality expectation resolved differently than predicted: the mechanism section was compressed to a conjecture with retained prediction and falsifier, not converted to a declared idealization — defensible, since the workspace hypothesis is at conjecture stage and its exceptions are unpriced; the two-bound framing is presented as a useful model. Follow-up executed with one authorial addition: the closing cycle failed title-body-alignment because the pass's own H1 dropped the fits-in-window condition, so the retitle carries it ("Soft degradation often binds before the hard cap when required evidence fits"; relocated; thirteen citers' link text updated; two glosses asserting the old universal reconciled).

Run 6 (pass ff5a) — control: no modality finding fired, but the control was impure. The note turned out to have genuinely defective premises (verbatim replay of a side-effectful charge_customer call is not safe because it succeeded once; outcome evidence CAN warrant same-context retention), so the pass fired an ordinary scope/category reframe ("A checked outcome licenses retaining an episode, not abstracting its explanation"; relocated, citers reconciled including one gloss asserting the defeated only-process-checks claim). For the machinery test, the control still passes: no mode was named, no statistical or ideal-type reading was invented, defeats were instance-shaped and repaired classically. For test design, the control selection was flawed — a deductive-looking note is not the same as a sound one, which is itself a small datum for the KB's review-value story.

Series verdict

Mode landings: one of six (run 5, statistical, guard in the disposition). Declines: four, every one for evidenced reasons, and every survey per-note prediction among them wrong — the machinery's shape-routing out-discriminated the survey's class labels in all four cases. Success criteria: every landed reframe named its mode and met its guard (run 5); an upward-in-form reframe fired (run 3, vacuous tendency → categorical rule); the control produced no modality finding. NOT exercised in-series: an ideal-type conversion with in-pass adequacy attack (covered pre-series by pass 2a6408 on the instantiation note; no series candidate survived the routing test) and a true statistical guard rejection (a proposed statistical landing blocked mid-pass — run 3's block happened at synthesis, which may be the natural place).

The headline finding is better than the one we designed for: mode conversion is rare because mode landings require mode-appropriate warrant — prevalence evidence for statistical, domain-priced exceptions for ideal-type — and most mismatched notes lack it. Their honest repairs are scope, category, or conditionality reframes. The modality machinery's principal observed effect is preventive: shape annotations route synthesis away from unguarded landings, and the two runs that could have produced degenerate outcomes (run 3 re-hedging; run 4 promoting a false universal) both avoided them. The 22-candidate survey overpredicted conversions systematically; class labels flag mismatch but the evidence decides the landing. Second-wave implication: run candidates through ordinary passes without mode predictions; track only whether guards bind and shapes discriminate.

Wave 2 (scheduled 2026-08-19)

Discipline, per the series verdict: no mode predictions. Class labels flagged these as mismatch candidates; the evidence decides the landing. Per-run readouts track only whether the shape annotations discriminate, whether any landed mode meets its guard, and the follow-up cost. Batches are ordered by inbound-citer blast radius (measured 2026-08-19), because a reframe's real cost is the rename-and-reconcile fanout, not the pass. Passes on different notes may run in parallel; follow-ups for large-radius notes should land the same day as their pass.

Batch A — small radius (2–6 citing files each), run freely: 1. access-burden-and-transformation-burden-are-independent-query.md (2) 2. codified-scheduling-patterns-can-turn-tools-into-hidden-schedulers.md (2) 3. llm-generation-confidence-tracks-typicality-not-soundness.md (3) — carries a mixed-modality body (one half deductive, one half rate-dependent), so its readout bears on per-claim mode assignment 4. weakly-discriminated-qualities-tend-to-be-underselected.md (4) — already declares conjecture status with named conditions; its readout bears on how modality composes with lifecycle stage 5. traditional-debugging-intuitions-break-when-tool-loops-can-recover.md (4) 6. human-analogies-suggest-functions-not-component-boundaries.md (6)

Run 7 / Wave 2 Batch A item 1 (pass 20260819T213103Z-d7f00d) — ordinary scope reframe, no mode landing. access-burden-and-transformation-burden-are-independent-query.md landed as Access burden and transformation burden are distinct query dimensions (access-burden-and-transformation-burden-are-distinct-query-dimensions.md). The initial premise shapes discriminated the repair: adaptive search and precomputed answers defeated operational independence GLOBAL by instance, while an exhaustively verified finite-domain model path was a plausibly fenced priced-exception to the categorical substrate rule. Neither shape warranted statistical or ideal-type conversion, so no modality guard bound or was met; the ideal-type candidate was rejected as the wrong category before its adequacy guard. Closing review cleared every semantic, prose, sentence, frontmatter, complexity, and structural gate, found no GLOBAL premise defeat, and preserved the separability update; it routed only local residue (system-relative burden measurement, routing-branch size, an undefined “bounded semantic call,” a trivial-lookup instance, and the verified-model priced-exception). Follow-up cost: commonplace-relocate-note rewrote five links and one ProperDocs redirect; two library citers were inspected, with one visible title/gloss changed and the other still truthful; two gitignored connection-report summaries and the active report's source/follow-up metadata were reconciled manually. Finished 2026-08-20: made both burdens explicitly system-relative, compressed routing into a corollary, folded validation into acceptance semantics, glossed the semantic call, and fenced the trivial-lookup and verified-model exceptions; estimation remains evidence-dependent.

Run 8 / Wave 2 Batch A item 2 (pass 20260819T221517Z-91a7a4) — ordinary category reframe, no mode landing. codified-scheduling-patterns-can-turn-tools-into-hidden-schedulers.md landed as Cross-task transition policy remains scheduling behind a tool interface (cross-task-transition-policy-remains-scheduling-behind-tools.md). Initial premise shapes discriminated role from placement: an application-owned outer loop defeated framework-level force GLOBAL by instance, and a documented workflow tool defeated tool-shape-implies-hidden GLOBAL by instance; every other non-HOLDS premise was also instance, so neither statistical nor ideal-type routing fired and no modality guard bound. Closing preserved the scheduler-role update and found no GLOBAL defeat, but it routed local boundary residue: “one invoked capability” can contain a workflow, capability-surface change alone can be authorization rather than scheduling, and hiddenness needs an audience; friction also found the optional-loop recommendation under-argued. Follow-up cost: the exact-title slug exceeded the 70-character limit and needed a shorter landing; commonplace-relocate-note then rewrote five Markdown links (two library citers and three gitignored reports) plus one ProperDocs redirect. Both library citers were premise-reconciled, two generated report residues and the active report/connect metadata were repaired manually. Finished 2026-08-20: replaced the unstable capability/task labels with externally meaningful, interceptable transition authority; made goal progression and cross-goal branching conditional tests and capability-surface change only supporting evidence when it serves such a transition; made concealment audience-relative; and grounded the architectural conclusion in programmable transition boundaries rather than loop optionality. Whether a particular domain treats a boundary as independently steerable remains evidence-dependent.

Run 9 / Wave 2 Batch A item 3 (pass 20260819T225902Z-c83f42) — mixed-modality discrimination; statistical title landing rejected, deductive reframe landed. llm-generation-confidence-tracks-typicality-not-soundness.md landed as Generation confidence does not by itself certify soundness (generation-confidence-does-not-by-itself-certify-soundness.md). Initial premises separated the claim modes: the universal zero-information premise took a GLOBAL prevalence defeat, the unrestricted training-objective premise took a GLOBAL instance defeat, and the anti-correlation branch took both instance and prevalence challenges while the no-entailment premise held. The GLOBAL prevalence shape routed a statistical title candidate, but the statistical guard rejected it because the artifact supplied no population-level prevalence evidence; no priced exception routed an ideal type. The headline instead landed as a deductive universal no-certification claim. The retained anti-correlation subclaim remained an unestablished statistical hypothesis with a declared population/comparison and explicit equal-or-opposite-ordering refuter, so its statistical guard bound and was met. Closing premises attacked that refuter with a positive-association population and preserved the mode separation with no GLOBAL defeat; closing review nevertheless weakened the selected update's secondary high-assurance clause because “per-output certification” and “validated low population risk” remain unresolved. Follow-up cost: commonplace-relocate-note rewrote links in eight Markdown files and added one ProperDocs redirect; three library citers were premise-reconciled, four source connection reports had stale titles or decoupling summaries repaired, and the canonical connection report's source metadata and filename were realigned.

Run 10 / Wave 2 Batch A item 4 (pass 20260820T000646Z-ee5856) — existing statistical mode retained; conjecture stage kept separate from mode. Plain keep at weakly-discriminated-qualities-tend-to-be-underselected.md: the title already stated a statistical tendency, so the pass repaired its thesis instead of forcing a mode conversion. Initial premises discriminated the claim shape: population dominance and drift-versus-slower-improvement were DOUBTFUL GLOBAL prevalence, literal reuse without material inheritance was DEFEATED LOCAL instance, and no priced-exception routed an ideal type. “Conjecture” remained the discovery-lifecycle stage and did not satisfy the statistical guard; the revised claim met that guard with named loop conditions, a relative accepted-set-enrichment comparison, and an explicit refuter—representative qualifying loops usually giving weakly discriminated qualities equal or greater enrichment. Absolute decline now requires an additional adverse direction. Closing premises named both dimensions explicitly (Claim mode: statistical; Discovery-lifecycle stage: conjecture), retained four HOLDS, one DOUBTFUL LOCAL instance, one DOUBTFUL GLOBAL prevalence, and no GLOBAL defeat. Closing review nevertheless weakened the selected update because discrimination and cross-quality enrichment still lack commensurate measures, headroom normalization, and a representative-loop sampling rule; compression also kept the diagnostic/remedy branch and repeated ending as residuals rather than reopening the pass. Closing effect: the note fell from 1,563 to 1,434 words, removed the duplicated maintainability pass and unidentified ontology case, repaired activation-versus-selection, and preserved a refuter-bearing statistical conjecture without treating its lifecycle label as evidence. Follow-up cost: none beyond the in-place note edit, pass reports, canonical connect refresh, and this readout—no rename, citer reconciliation, or ProperDocs redirect was required. Finished 2026-08-20: defined independently outcome-calibrated binary quality measures, calibration-block matched-pair discrimination, evaluation-block headroom-normalized enrichment, a preregistered stratified loop sample, and behavioral branching tests for propagation; a stronger oracle now counts only when it changes actual retention and improves calibrated downstream outcomes. The conjecture still awaits representative prevalence evidence.

Run 11 / Wave 2 Batch A item 5 (pass 20260820T011522Z-dbg5) — merge accepted; the statistical delta failed its warrant guard. traditional-debugging-intuitions-break-when-tool-loops-can-recover.md produced Update: NONE and a merge recommendation into apparent-success-is-an-unreliable-health-signal-in-framework-owned.md: the outcome/path non-identifiability mechanism survived and was already owned by the parent, while the proposed human-cognition delta did not. Premise shapes discriminated the mismatch: programmer over-trust, the traditional-software health proxy, and stopping after a good artifact were DOUBTFUL GLOBAL prevalence; the categorical exception-handling contrast was DEFEATED LOCAL instance; no priced-exception routed an ideal type. Those prevalence shapes identified a statistical candidate, but the guard rejected a landing because the artifact supplies no population, comparison, observation record, or refutable rate. Closing effect: none by full-pass protocol—the pending disposition stopped the pass before edits, step 9, and the closing cycle, so the note remained byte-identical during the pass. Resolved 2026-08-20 by user authority: the engineered-versus-synthesized recovery contrast and where-to-look-next consequence were reconciled into the parent without importing the unsupported prevalence claim; the source was deleted, four inbound library citers were reconciled, and a ProperDocs redirect was added.

Run 12 / Wave 2 Batch A item 6 (pass 20260820T014450Z-90f462) — plain keep; no mode landing. human-analogies-suggest-functions-not-component-boundaries.md stayed at its existing path. Premise shapes discriminated a local boundary condition from a title failure: seven premises held, while the exclusive claim that human-derived boundaries constrain a target only through shared causal dependencies was DEFEATED LOCAL by an instance in which an operator-facing audit or interface constraint warrants one familiar exposed boundary; no GLOBAL defeat, prevalence, or priced-exception appeared. The title's universal non-determination claim therefore survived, neither statistical nor ideal-type routing fired, and no mode guard bound. Closing premises reproduced the same LOCAL instance; critique found no surviving attack; compression passed all four criteria; all catalog gates passed except one accessibility warning on “component decomposition” and “provenance”; friction preserved unresolved questions about what makes a dependency boundary-sensitive and what witnesses analogy-led decomposition inheritance. Closing effect: the note moved from 668 to 645 words, folded the redundant standalone memory recap, clarified implementation media and authority, and normalized its footer labels without changing its title claim. Step 9's unapproved description rewrite was restored before the authoritative final SHA and before any closing job was created. Actual follow-up cost: no rename, citer reconciliation, or ProperDocs redirect; only the in-place note edit, pass reports, canonical connect refresh, and this readout. Finished 2026-08-20: made rival preservation without material degradation the boundary-discriminating test; made operator-facing audit and interface needs independent target-side warrant rather than transferred causal warrant; added the Tulving three-space witness; and glossed component decomposition and provenance. How common analogy-led decomposition inheritance is remains an empirical question.

Batch B — medium radius (7–17), run after batch A's follow-ups are clean: 7. prose-has-no-dereference-reinforce-facts-at-point-of-use.md (7) 8. entropy-management-must-scale-with-generation-throughput.md (8) 9. apparent-success-is-an-unreliable-health-signal-in-framework-owned.md (9) — sits inside the clean-model cluster, so its readout previews batch C's hub 10. memory-design-adds-operational-axes-to-artifact-analysis.md (10) 11. indirection-is-costly-in-llm-instructions.md (11) 12. structure-activates-higher-quality-training-distributions.md (17) — the corpus's best numeric-prevalence evidence (benchmark deltas plus a survived null); whatever lands here is the strongest test yet of the statistical guard against real rates

Run 13 / Wave 2 Batch B item 7 (pass 20260820T071619Z-b7d13a) — statistical reframe landed; guard met, exact effect remains untested. prose-has-no-dereference-reinforce-facts-at-point-of-use.md landed as Local materialization should outperform distant natural-language declarations (local-materialization-should-outperform-distant-declarations.md). Initial premise shapes discriminated the target: ordinary propagation failing often enough and point-of-use restatement improving application were DOUBTFUL GLOBAL prevalence, while mutation/shadowing in the formal analogy and uncheckable contextual consequences were LOCAL instance; no priced-exception routed an ideal type. The statistical guard bound and was met before step 9 with named conditions (distant or inferentially non-obvious use), a declaration-only comparator, and a held-out zero-or-negative correct/non-contradictory application refuter. Closing premises attacked the landing with four GLOBAL prevalence doubts—available headroom, locality removing the operative burden, canonical generation covering the sampled fact types, and benefits exceeding context/contradiction costs—without changing its mode. Closing critique partially landed because exact copies and derived labels remain different treatments; accessibility, semantic, and sentence review also routed local residue about experiment eligibility, intervention consistency, derivation/check framing, and two glossary terms. The warranted update was preserved, compression passed, and no second edit cycle ran. Follow-up cost: commonplace-relocate-note rewrote 14 Markdown link surfaces plus one ProperDocs redirect; seven library citers and two workshop citers were premise-reconciled, including weakening two reference-proposal dependencies from rests-on to see-also; five gitignored connect reports received path-only rewrites, and the canonical connect/friction/premise report names and metadata were realigned. The report, note, affected library/workshop artifacts, and redirects validate cleanly. The numeric-refuter coverage gap remains open: this landing's threshold is qualitative (<= 0 difference), not an evidence-backed rate.

Run 14 / Wave 2 Batch B item 8 (pass 20260820T083259Z-e8b4c1) — category/quantity reframe, no mode landing. entropy-management-must-scale-with-generation-throughput.md landed as Maintenance capacity must match harmful-artifact inflow (maintenance-capacity-must-match-harmful-artifact-inflow.md). Initial premise shapes discriminated the repair without a predicted mode: bad-pattern amplification, gross generation as the matching quantity, and volume-not-time were DEFEATED GLOBAL by instance; typed generation, schema enforcement, and CI prevention supplied three DEFEATED GLOBAL priced-exception shapes. Those priced exceptions routed an ideal-type candidate, but its adequacy guard failed because prevention is ordinary in this domain and can dominate the generation-to-cleanup relation; no prevalence defeat appeared, so no statistical landing or refuter guard bound. The pass instead kept a universal conditional queue invariant while changing the load variable from gross output to risk-weighted harmful retained-artifact inflow and treating prevention, containment, detection, and repair as capacity. Closing premises retained four HOLDS, one DOUBTFUL LOCAL instance about aggregating stage allocation and latency, no GLOBAL defeat, and no prevalence or priced-exception shape. Semantic, frontmatter, accessibility, complexity, and structural review passed; critique partially landed on the missing ex-ante operating signal, compression retained a repeated-implications WARN, friction kept the Codex stabilization inference THIN, and prose/sentence review routed one anthropomorphic verb and one packed queue conditional. The warranted update was preserved with no second edit cycle. Follow-up cost: commonplace-relocate-note rewrote links in 14 Markdown files plus one ProperDocs redirect; four note citers, four source ingests, one workshop citer, and five generated connect reports were premise-reconciled, the canonical connect report was renamed, and the stale Harness Engineering “not yet captured” follow-up was closed. The note moved from 620 to 615 words; its conceptual invariant still lacks a validated leading proxy for harmful inflow.

Run 15 / Wave 2 Batch B item 9 (pass 20260820T091632Z-a9c15d) — evidential-strength reframe; statistical landing rejected, deductive non-certification landed. apparent-success-is-an-unreliable-health-signal-in-framework-owned.md landed as Final task success does not establish intended-path health (final-task-success-does-not-establish-intended-path-health.md). Initial premise shapes discriminated reliability from conclusiveness: the inference from one possible hidden fallback to an unreliable health signal was DEFEATED GLOBAL by prevalence, and the claim that framework-owned runtimes typically collapse primary and fallback success without path evidence was DOUBTFUL GLOBAL by prevalence; four existential and observational-equivalence premises held, and no priced-exception routed an ideal type. The statistical guard rejected a landing because the artifact supplies no runtime population, reliability comparison, prevalence evidence, or refuting rate. The pass instead landed a deductive conditional: when primary and fallback success share one terminal observation and no independent execution signal is available, final success cannot identify intended-path health. It also separated path-deviation evidence from guarantee degradation, reserving degraded execution for weaker named guarantees. Closing premises returned six HOLDS with no GLOBAL defeat, and critique found no surviving attack. Accessibility, complexity, frontmatter, prose, and sentence review passed; closing semantic review routed a bypass case omitted from the three-state taxonomy and the intentionally stale pre-rename filename qualifier, structural review routed one lowercase footer title, compression retained repeated mechanism/context-cost warnings, and friction asked what fields make durable attempt events sufficient. The selected update strengthened without a second edit cycle. Follow-up cost: commonplace-relocate-note rewrote 13 Markdown files and two ProperDocs redirect relations; eleven authored library/workshop citers were inspected and reconciled, including narrowing one detection-destroyed premise to the absence of independent path evidence; three canonical gitignored reports and the active packet metadata were realigned. The note moved from 1,078 to 998 words; prevalence and event-schema questions remain open.

Run 16 / Wave 2 Batch B item 10 (pass 20260820T102230Z-4e6b10) — plain keep; no mode landing. memory-design-adds-operational-axes-to-artifact-analysis.md stayed at its existing path. Initial premises discriminated the artifact-versus-operation split from the exact six-part carving: four premises held, while the claim that the named package is the uniquely right reusable decomposition was DOUBTFUL LOCAL by one instance; no GLOBAL defeat, prevalence, or priced-exception appeared. The pass therefore made no mode prediction or conversion, and neither statistical nor ideal-type guard bound. The edit defined an operational axis as a recurring comparison question rather than an exhaustive orthogonal basis, added the capture admission-criterion split, separated use-time behavioral authority from governance rights, compressed signal timing and repeated split rationale, linked the activation mechanism, and removed duplicated footer relationships and unauthorized sharpens/evidence labels. Closing premises again returned four HOLDS and one DOUBTFUL LOCAL instance, now focused on whether evaluation needs explicit treatment in every minimal design; no GLOBAL defeat or mode-routing shape appeared. Critique partially landed on mixing process, governance, lifecycle, and assessment in one checklist; friction preserved the central claim but kept the artifact-boundary exclusion THIN; compression still treated signal timing and repeated framing as residuals. Complexity, frontmatter, prose, semantic, and structural review passed; accessibility routed two vocabulary glosses and one source-type identification, and sentence review routed the packed four-definition opening. The warranted update was preserved without a second edit cycle. Follow-up cost: no rename, citer reconciliation, or ProperDocs redirect; only the in-place note edit, pass reports, canonical connect refresh, and this readout. The note moved from 1,374 to 1,080 words; the exact grouping and evaluation's status remain provisional.

Run 17 / Wave 2 Batch B item 11 (pass 20260820T120035Z-858434) — categorical resolver-boundary reframe; no mode landing. indirection-is-costly-in-llm-instructions.md landed as Model-resolved indirection adds interpretation work to LLM execution (model-resolved-indirection-adds-interpretation-work-to-llm-execution.md). Initial premise shapes discriminated the category error: a harness expanding an alias before model consumption DEFEATED the all-indirection premise GLOBAL by instance; three other defeats were LOCAL instance, with no prevalence or priced-exception. The pass therefore narrowed the actor and binding-time boundary instead of predicting a statistical or ideal-type landing, and neither mode guard bound. The revision replaced the zero-cost code analogy with deterministic resolution, made literalization conditional on complete token, authority, failure, and lifecycle costs, and added a generated-copy validity obligation. Closing premises retained six HOLDS, three DOUBTFUL LOCAL instance, and one DEFEATED LOCAL instance, with no GLOBAL defeat or mode-routing shape; compression passed and the warranted update was preserved. Critique, friction, accessibility, and semantic review routed residuals around the formal-resolver boundary, net-cost reading, generated-copy enforcement, undefined assembler/Commonplace identification, and an overstated instruction-generation witness; no second edit cycle ran. Follow-up cost: commonplace-relocate-note rewrote 18 Markdown link surfaces plus one ProperDocs redirect; eleven library citers and five workshop citers were premise-reconciled, two invalid premise links were removed, two gitignored connect reports were aligned, and the canonical connect, friction, and premise report filenames and source metadata were realigned. The note moved from 594 to 696 words; the magnitude of model-side binding cost and the complete-form decision rule remain unmeasured.

Run 18 / Wave 2 Batch B item 12 (pass 20260820T143104Z-cda404) — categorical causal-identification reframe; the statistical guard bound and rejected the numeric landing. structure-activates-higher-quality-training-distributions.md landed as Structured-prompt gains do not establish training-distribution selection (structured-prompt-gains-do-not-establish-distribution-selection.md). Initial premise shapes discriminated the candidate repairs: the claim that selected structured documents are better on average was DOUBTFUL GLOBAL by prevalence; four causal/generalization premises were DEFEATED or DOUBTFUL by instance (three GLOBAL, one LOCAL); no priced-exception routed an ideal type. The prevalence shape made the statistical guard bind, but the 5–12-percentage-point code-verification gain and Claude Sonnet code-QA null (84.8% versus 85.3%) are bounded study cells, not representative prevalence evidence, and the incumbent states no population, comparison, condition, or rate that could refute a general tendency. The pass therefore rejected a statistical landing despite real numeric evidence and reframed to a universal epistemic limit: formatting compliance, extra computation, task decomposition, and learned procedures predict the same observed gain, so performance alone cannot identify training-distribution selection. A matched causal result that distinguishes the mechanism would refute that claim; no statistical or ideal-type guard binds to the final landing. Closing critique found no surviving attack, all accessibility/frontmatter/prose/semantic/sentence/structural gates passed, and closing premises returned five HOLDS plus one DOUBTFUL LOCAL instance, with no GLOBAL defeat. Compression routed one permanence branch, complexity routed duplicate footer edges, and friction routed precision limits on the matched test and the design consequence; the warranted update strengthened without a second edit round. Follow-up cost: commonplace-relocate-note rewrote 28 Markdown files plus one ProperDocs redirect; all 18 authored library citers (eight notes/indexes and ten source ingests) were premise-reconciled, ten generated connect reports received path updates, and four operational/workshop references were realigned. Canonical connect, friction, premise, and full-pass paths/metadata were realigned; the direct compression scratch was removed after retaining its exact closing copy, and initial/closing evidence stayed immutable. The strongest numeric-prevalence candidate in the wave therefore did not close the numeric-refuter gap: benchmark deltas plus a null still do not satisfy a population-level tendency guard.

Batch C — large radius, one at a time, same-day follow-up committed before the next starts: 13. agent-memory-needs-discoverable-composable-trusted-knowledge-under.md (18) 14. storing-llm-outputs-is-constraining.md (19) — opens with a declared working-hypothesis line; tests mode-versus-stage composition at scale 15. stale-indexes-are-worse-than-no-indexes.md (24) — foundational to the marks doctrine; a reframe here touches enforcement rationale 16. files-not-database.md (29) — the regime boundary lives in a sibling note; the pass may pull it in or leave it, and either is informative 17. bounded-context-orchestration-model.md (59) — the strongest remaining candidate for a genuine in-pass ideal-type conversion: the cluster's own "clean model / degraded variant" vocabulary is domain-internal pricing, unlike run 2's unmarked blurring. Also the widest-radius reframe in the wave; schedule when the follow-up day is free 18. knowledge-storage-does-not-imply-contextual-activation.md (211) — deliberately last. The title is a non-implication claim, which one storage-without-activation case establishes, so the reframe probability may be low despite the seven body hedges — but if the pass does reframe, the follow-up is the largest citer sweep in the KB's history. Do not start it without the follow-up capacity reserved.

Run 19 / Wave 2 Batch C item 13 (pass 20260820T170531Z-a4c9) — plain keep with exclusivity dropped; no mode landing; closing premises weakened the retained universal. agent-memory-needs-discoverable-composable-trusted-knowledge-under.md stayed at its existing path. Initial premise shapes discriminated the exclusivity overreach from the necessity claim: four premises held, while the minimal-basis premise — every other requirement is only surrounding machinery — was DEFEATED GLOBAL by instance (a findable, usable, validated artifact can still be too costly to load, making bounded-context loadability behave like a fourth artifact property); no prevalence or priced-exception appeared, so neither mode guard bound. The edit kept the title's one-way necessity claim, replaced the peer-versus-surrounding taxonomy with a limits section naming the triad necessary but non-exhaustive, distinguished artifact-level loadability from system-level activation, glossed bounded context and context engine, and compressed the repeated opening contrast. The notable closing result: the warranted-contribution comparison came back weakened — closing premise decomposition routed two fresh GLOBAL instance counterexamples against the universal necessity itself (small stores scanned wholesale need no separate discovery handle; cheap unproven hunches improve action when treated as rechecked hypotheses), and closing critique partially landed because discovery, composition, and reliance are assessed relative to a retrieval system, consumer class, task, and reliance policy rather than residing in the artifact alone. The one-cycle protocol routed all of it without a second round, so the title's universal artifact-level necessity now carries unresolved GLOBAL attention — the strongest open residual any wave-2 keep has left on a retained title. Friction survived with three UNSUPPORTED and two THIN interdependence joints; connect retained ten outbound candidates plus a split flag (keep as the triad's synthesis/router only while the property arguments compose as one premise). Series-relevant maintenance observation: closing connect reports stale-indexes-are-worse-than-no-indexes.md — batch C item 15 — still describes retired areas: and docs/indexes.md machinery, so that pass will meet known-stale content. Follow-up cost: none of the feared large-radius fanout — no rename, citer reconciliation, or ProperDocs redirect; the in-place edit (942 to 896 words), pass reports, canonical connect refresh, and this readout. The note validates clean; whether the necessity claim should be relativized to a consumption context is the open question the closing evidence poses.

Run 20 / Wave 2 Batch C item 14 (pass 20260820T180920Z-2929da) — identity-to-distinction reframe; no mode landing; the wave's largest premise-health repair. storing-llm-outputs-is-constraining.md landed as Selecting an LLM output fixes a result, not its interpretation (selecting-an-llm-output-fixes-a-result-not-its-interpretation.md). Initial premise decomposition returned five GLOBAL DEFEATED premises, every one instance-shaped: an audit log retains every sampled response while no consumer treats any as operative; a prompt requiring exactly OK admits one operational output; the answer "Yes" stays compatible with distinct readings of an ambiguous question; a grammar-constrained decoder cannot vary; and storing a random digit under an unambiguous instruction preserves an outcome without ruling out a valid reading. One further premise was DEFEATED LOCAL instance and one HELD. No prevalence and no priced-exception appeared, so neither mode guard bound and no target mode was named. Five GLOBAL defeats did not route a delete, and that is the run's main machinery datum: every defeat landed on the title's storage-is-constraining identity, not on the material, which was warranted and unowned elsewhere — so the pass reframed to the distinction the artifact actually supports (selection settles result identity for a consumption path; ambiguity can survive inside the selected text; behavioral authority is a separate condition) instead of summing the defeats into a note-level kill. Body edits additionally removed two competing theses — the generator/verifier strategy essay and the verbatim-risk worked case — and dropped the append-only-log anecdote, whose mechanism constrains mutation rather than selecting an interpretation and was therefore a counterexample rather than an instance. Closing premises fell to three HOLDS, two DEFEATED (both LOCAL), and one DOUBTFUL (LOCAL), with no GLOBAL defeat; closing critique found the strongest result-versus-meaning attack no longer surviving; semantic returned 9 PASS / 1 WARN with the other bundles near-clean. The warranted contribution came back strengthened — the first wave-2 pass to strengthen rather than preserve or weaken. Follow-up cost: commonplace-relocate-note renamed the note and rewrote links across 19 authored citers (notes, a definition, a tag-README, research/, and six source ingests) plus the ProperDocs redirect; kb/notes/definitions/constraining.md, which the packet flagged as independently asserting the refuted identity, was reconciled rather than deferred. Committed b4d86edb. Residuals left open: whether the removed generator/verifier and verbatim-risk branches warrant separate commissioned notes, and the stability boundary — a version-pinned artifact is a stable testing target without implying stable received bytes, meaning, or verdict.

Run 21 / Wave 2 Batch C item 15 (pass 20260820T213110Z-98ad) — comparison-axis reframe; no mode landing; closing left two GLOBAL doubts on the retained title. stale-indexes-are-worse-than-no-indexes.md landed as Stale indexes reduce discovery when they suppress fallback search (stale-indexes-reduce-discovery-when-they-suppress-fallback-search.md). Initial premise shapes discriminated the comparison axis from the mechanism: five premises held — including the stop-after-apparently-complete-index witness and the control-flow difference — while the premise that the fallback's additional discoveries always outweigh the index's orientation and latency benefit was DEFEATED GLOBAL by instance (an incident-response index can expose the relevant current runbook fast while omitting an unrelated new one), and the claim that every omitted item becomes undiscoverable was DEFEATED LOCAL by instance (another authored link can still expose it). No prevalence or priced-exception appeared, so neither mode guard bound. The reframe is therefore a comparison-axis and scope repair, not a modality landing: a bare ranking claim ("worse than no indexes") became a conditional causal claim about discovery recall along the one route whose fallback was suppressed, with overall task utility explicitly excluded. The edit also replaced anthropomorphic stopping language ("feel oriented", "trusts it") with observable stop-after-index behavior, narrowed the cross-artifact transfer from "any authoritative artifact" to artifacts treated as exhaustive, and turned the closing invariant from an unconditional preference for exhaustive search into a rule against silently claiming more coverage than provided. The notable closing result runs opposite to run 20's: the warranted contribution was preserved, but closing premise decomposition raised two fresh GLOBAL DOUBTFUL premises against the retained title — whether a "more complete" fallback actually realizes relevant retrieval, and whether index-only discoveries offset the recall loss — alongside two semantic WARNs and four THIN friction joints on refresh triggering, failure policy, and consumer uptake. This is the second wave-2 keep (after run 19) to leave GLOBAL-level attention on a title it kept. Follow-up cost — the largest of batch C so far: commonplace-relocate-note rewrote links in 29 Markdown files and added the ProperDocs redirect; 25 authored citers across kb/notes/, kb/reference/, kb/types/, kb/instructions/, kb/sources/, kb/agent-memory-systems/, and one workshop row were premise-reconciled, with visible link text and glosses rewritten from "suppresses search entirely" and "become invisible" to route-specific recall. Test-design contamination worth recording: run 19 predicted this pass would meet the note's known-stale areas: and docs/indexes.md content, but commit b89986d3 cleaned that 80 minutes before the pass started — the pass never met it, so this run is not a datum about the machinery handling stale content.

Run 22 / Wave 2 Batch C item 16 (pass 20260821T105408Z-d10d21) — conditional-scope reframe; no mode landing; the wave's cleanest no-mode-signal case, and the first landed reframe to take a GLOBAL defeat on its own new claim. files-not-database.md landed as Incrementally constrained files defer centralized schema commitment until write invariants stabilize (files-defer-centralized-schema-commitment-until-invariants-stabilize.md). Initial premise decomposition returned eight premises and every non-HOLDS one was instance-shaped: two DEFEATED GLOBAL (a managed agent platform supplying transactional SQL, browsing, backup, and access control while lacking a writable shared filesystem or Git credentials; an entity-resolution KB requiring uniqueness and atomic alias assignment on its first writes), one DOUBTFUL GLOBAL (high-volume ingestion discarding stable identity or provenance before a later migration can reconstruct it), three DEFEATED LOCAL (a single documents table with JSON metadata defers normalization about as far as named Markdown files; changing file conventions across many notes is still a data migration; per-record authorization is a canonical-store requirement rebuildable indexes do not enforce), and two HOLDS. No prevalence and no priced-exception appeared anywhere, so no mode target was named, neither guard bound, and no mode-routing shape existed to reject — the purest instance-only decomposition in either wave. The repair was accordingly a scope-and-category reframe: the universal comparative title ("files beat a database") became a conditional about when a schema commitment must become binding, with the note explicitly denying "schema versus no schema" (file conventions are a distributed schema) and denying that databases are inherently rigid (a document table with flexible metadata defers normalization too). The file-tooling advantages were demoted from properties files have to contingent savings available when agents already share file and Git tools.

Sibling-boundary result — the item's open question resolved toward absorption. The schedule predicted the pass might pull in many-to-many-edge-state-is-where-files-yield-to-a-database.md or leave it. It did neither cleanly: it replaced the note's two scale sections with a single rebuildable derived views versus authoritative state boundary test, named the ownerless-many-to-many case as one item in that class, and cited the sibling for the structural trigger. The parent now owns the boundary class, the sibling owns the test — a division that required reconciling the sibling's own footer, which had claimed the parent "leaves implicit" a boundary the parent now names.

The closing result is the run's main machinery datum. The warranted contribution came back weakened, and unlike runs 19 and 21 — both keeps that left GLOBAL attention on a title they retained — this is the first wave-2 pass to land a reframe and then have closing premises GLOBAL-defeat the landed claim itself: a stable write-time invariant does not by itself require database authority, because a serialized Git publication boundary can enforce one safely; the note needs the comparative condition that a database is the cheaper coordination substrate. Closing critique partially landed on the same seam from the other side — "rebuildable" is an informational property that does not price recovery latency, update lag, or availability, so an informationally derived store can be operationally authoritative. Six closing HOLDS and one LOCAL DOUBTFUL surround that one GLOBAL defeat. The one-cycle protocol routed all of it to Open items without a second edit round, and the structural bundle closed with one FAIL (three uncapitalized condition bullets) alongside accessibility and sentence WARNs — the first closing FAIL in the wave. Local residuals cleared by operator direction 2026-08-21 (ef42e6ab), outside the pass: bullets capitalized; frontmatter and bitemporal validity glossed; Graphiti identified on first mention; "the decisive question" scoped to the correctness question, since it had excluded the cheaper-substrate branch the note's own decision rule states; the incremental constraining link text matched to what constraining.md defines; the schema-versus-no-schema point folded from three statements to one; the paired Graphiti/Commonplace examples compressed to their distinct triggers. 875 to 873 words. The GLOBAL residual is deliberately not repaired — the note still asserts that a database earns authority when an invariant must hold at write time, which closing premises defeated (a serialized Git publication boundary can enforce one safely, so the invariant alone is not sufficient; the note needs the comparative condition that a database is the cheaper coordination substrate). That is a claim change, not a local repair, and stays open for a decision or a later pass.

Follow-up cost — the largest of the wave so far by file count. commonplace-relocate-note rewrote 41 Markdown files plus one ProperDocs redirect (the exact-title slug is 100 chars, so the landing compressed to 68 by dropping "incrementally constrained" and "write", keeping the head and the conditional tail, as in runs 8 and 18). Beyond the path rewrite: eleven library citers were premise-reconciled, including two whose glosses had become false — axes-of-artifact-analysis.md, which attributed to the note the claim that "a database schema forces premature commitment to access patterns" (now explicitly denied), and architecture-README.md, whose "files with git beat a database" head the closing connect had flagged as advertising the superseded universal. Sixteen source ingests were reconciled: two contradicts labels were downgraded to contrasts (Cognee, Mem0) because a database-primary system is not a counterexample to a conditional claim but its other branch, and the Graphiti ingest carried two standing "not yet executed" follow-ups — dating to 2026-03-05 and repeated 2026-03-09 — asking for exactly the boundary section this pass wrote; both were closed as executed, and its "strongest counterexample to files-first" framing was rewritten to the boundary case the note now names. Five workshop documents were reconciled; db-native-reflective-system, whose entire gap list is a checklist derived from the old note's four tool-chain capabilities and which quotes that body verbatim, received a dated provenance marker rather than a rewrite — its checklist still works as a gap list, but the quoted text is now historical and the "costs every database-backed system pays" reading no longer follows. Two retained premise-cohort packets and one link-audit record were deliberately left carrying the old claim as historical evidence. The canonical connect report was renamed and its source metadata realigned; note, collections, and the redirect map validate clean.

Run 23 / Wave 2 Batch C item 17 (pass 20260821T115349Z-a7c2) — plain keep with a scope-and-formalism repair; no mode landing; the series' flagged ideal-type candidate produced no domain-pricing shape at all. bounded-context-orchestration-model.md stayed at its existing path with its title unchanged; the repair landed in the description, the opening thesis, and the formalism. Initial premise decomposition returned nine premises and every non-HOLDS one was instance-shaped — two DEFEATED GLOBAL (mutually interacting streaming calls against atomic call(P); remote mutation plus a post-mutation timeout and a concurrent actor against exact explicit K), one DEFEATED LOCAL, four DOUBTFUL LOCAL, two HOLDS. The closing rerun returned eight premises with the same distribution: one DEFEATED GLOBAL (the singular select(K) -> C; call(C) loop cannot represent an independent parallel batch, and "serializing the calls doubles barrier latency, which the note itself names as a comparison dimension"), one DEFEATED LOCAL, one DOUBTFUL LOCAL, five HOLDS. Across both checks: ten instance shapes, zero prevalence, zero priced-exception. Neither guard was reached, and no mode target was named in either direction.

That result falsifies the survey prediction directly. Item 17 was scheduled as "the strongest remaining candidate for a genuine in-pass ideal-type conversion" on the grounds that the cluster's own "clean model / degraded variant" vocabulary is domain-internal pricing. The pass found no priced exception to price. What it found instead was a pseudo-formalism, and the repair for a pseudo-formalism is deletion rather than declaration: the idealized scalar cost measure (||P|| as "an idealized effective-cost measure over the whole prompt", with select "subject to the feasibility constraint ||P|| <= M") drew a prose/pseudo-formalism FAIL, a semantic/underspecified-assertions WARN, and a DOUBTFUL LOCAL premise, and was replaced by a task/model-relative predicate — require feasible(C.prompt, C.task, C.model) — plus the explicit denial that "the model does not reduce it to a single measurable scalar." This is the machinery's third distinct response to an idealization: convert it (never yet achieved in-pass), reject the conversion at the adequacy guard (run 14), or strip the idealized apparatus as unearned formalism (run 23). The third route needs no modality vocabulary at all, which is why the guards never saw it.

Genre datum: no title move, and frontmatter/title-body-alignment passed at both ends. The drift is in the description and thesis, running from universal-ontological toward conditional-representational — "the model captures the full space of such architectures" became "This normal form represents the joint system; it does not establish that one orchestration architecture is best", with the closed-world, explicit-state, and completed-call conditions hoisted into the opening. Same scope-and-category family as runs 21 and 22, plus a formalism-honesty component neither had.

Follow-up cost — the cheapest of batch C by a wide margin, and initially an under-run. No rename, so no commonplace-relocate-note and no link rewriting; the pass commit (faff75ab) touches one file, 1,750 words down to 1,220. The later citer sweep examined the 75 files containing 119 direct Markdown links at sweep time and reconciled 26 live artifacts. Five had retained universal breadth, completeness, or unsupported effectiveness predictions; the re-audit found twenty-one more that imported retired P/K + r/scalar-M notation, physical-unboundedness, or prescriptive force. The revised select/call lemma, tool-loop and context-engineering definitions, surrounding notes, relevant source ingests, and active workshops now use the surviving conditional normal form or explicitly mark their own stronger toy assumptions. The dated Cordis original and a premise-cohort capture deliberately retain the old claim as historical evidence. Fourteen Open items originally stood unactioned under the one-cycle protocol, seven of them closing-cycle residuals; the most consequential was the GLOBAL parallel-batch defeat, which the closing cycle explicitly declined to auto-edit and therefore left live in the shipped text. GLOBAL residual cleared by operator direction 2026-08-22, outside the pass: select(K) now returns a finite nonempty independent batch B; call_all(B) preserves concurrent execution and waits at its barrier; and transition(K, B, R) incorporates the aligned results. Singleton batches preserve sequential workflows. The conversion lemma now preserves call specifications, batch membership, and barrier order, and its title is narrowed to the barrier-delimited class it proves. Eighteen live dependents were reconciled; the closing premise report and dated captures retain the defeated singular form as historical evidence. Thirteen Open items remain.

Coverage gaps wave 2 could close (tracked, not predicted onto any note): a statistical landing whose stated refuter is numeric rather than qualitative; a genuine in-pass ideal-type conversion whose adequacy record the closing premise rerun attacks; and ~~a landing-moment guard rejection~~ — CLOSED by run 9 (pass c83f42): a prevalence-shaped GLOBAL defeat routed a statistical target for consideration and "the statistical title guard therefore rejects that candidate rather than licensing a bare 'often' or 'may'" — the guard bound in the synthesis text itself, not by runner discretion; the anti-correlation subclaim was separately held to its own population-comparison-and-refuter obligation. Two gaps remain after run 23, and run 23 spent the last strong candidate for one of them — runs 19 through 23 produced no mode landing at all: the strongest numeric candidate demonstrated that bounded benchmark cells plus a null do not supply a population-level refuter, and bounded-context-orchestration-model, scheduled as the likeliest domain-pricing case, yielded zero priced-exception shapes. The ideal-type gap now has no queued candidate at all; closing it needs a note found outside this survey, and run 23 suggests where not to look — an idealization carried by undischarged formalism reads as pseudo-formalism to the prose gate and gets stripped before any mode question is asked. Run 22 sharpens what the drought means. Its eight premises produced no prevalence and no priced-exception at all, so the guards were never reached — the machinery did not decline a conversion, it was never offered one. Across batch C the shape distribution, not guard strictness, is what keeps mode landings rare: large-radius foundational notes state universals whose defeats are counterexamples, and counterexamples route scope and category repairs. Batch B also produced the ideal-type counterpart of run 9's rejection: in run 14, three priced-exception shapes routed an ideal-type candidate and the adequacy guard rejected it — prevention is ordinary practice in that domain and can dominate the generation-to-cleanup relation, so the first-order model failed the dominance commitment. Both mode guards have now bound and rejected at least once each; what neither has yet done is accept a conversion in-pass.

Batch A sequencing consequence (resolved). The pending merge (run 11) targeted apparent-success-is-an-unreliable-health-signal-in-framework-owned.md, batch B item 9; the merge was accepted in a4dd014c before that pass ran, so no pass captured a note mid-absorption.

Series success criteria

Met, runs 1–6 (2026-08-19). The machinery validates when, across runs 1–5, every landed reframe names its target mode and meets its guard, at least one adequacy record is attacked by a closing premise rerun, at least one upward reframe fires, and run 6 produces no modality finding. Any guard that binds only because the runner happened to be careful — rather than because the instruction text forced it — is an instruction defect to fix before the second wave.

Series status (2026-08-21)

23 runs; 21 of the survey's 22 candidates passed, plus the anchor case and the control. One candidate is unrun: batch C item 18, knowledge-storage-does-not-imply-contextual-activation.md (211 citing files, the largest citer sweep in the KB). It was deferred deliberately, and the deferral reasoning still holds — a non-implication title that one storage-without-activation case establishes has low reframe probability, so the expected cost is a large follow-up bought for a likely plain keep. It stays parked until follow-up capacity is reserved; running it is not required for any open question the series still has.

What the series established. Mode conversion is rare, and the reason is warrant availability rather than guard strictness. Across 23 runs the shape distribution decided every outcome: instance-shaped defeats dominate, and counterexamples route scope, category, and conditionality repairs, not mode landings. Both guards have bound and rejected — statistical at run 9, ideal-type at the adequacy check in run 14 — so the machinery is demonstrably not permissive; it has simply not yet been offered a conversion it could accept. The survey's per-note class labels were falsified as predictions in every case where they were tested, which is itself the designed-for result: labels flag mismatch, evidence decides the landing.

What remains unobserved. An accepted in-pass mode conversion of either kind. Both gaps are now candidate-starved rather than guard-blocked: no queued note carries prevalence evidence with a numeric refuter, and none carries a priced exception. If the series resumes, it should resume from a fresh search for those two shapes rather than from the remaining survey item.