ADR 066 test runs: first passes under the modality machinery
Six full passes, chosen from statistical-mode-candidates.md so that together they exercise every path ADR 066 added: both reframe directions, both new mode targets, both landing guards, the premise gate's counterexample-shape annotations, and one control where the machinery must not fire. Run sequentially (the pass's concurrency precondition; also each run's readout may adjust expectations for the next). Independent runner preferred, per the 2a6408 precedent. Route each report's readout back into this file.
Protocol per run
Ordinary run-full-improvement-pass-on-note.md invocation — the point is that the standard pass now does this work; no special harness. Record per run: (a) did the premise report carry shape annotations, and were they accurate; (b) did step 7 name a target mode where warranted, routed by those shapes; (c) did the landing meet its guard (stated refuter / adequacy record present before step 9); (d) did the closing premise rerun attack any new adequacy record; (e) were reframe follow-ups (rename, citer reconciliation) recorded as Open items.
The runs
Run 1 — statistical reframe down, easiest case. COMPLETED 2026-08-19 (pass 20260819T125030Z-s5qz) — prediction wrong, machinery coherent.
kb/notes/task-fitted-structure-costs-cross-task-reuse.md → renamed current-task-fit-alone-does-not-warrant-costly-entrenchment.md
The predicted statistical landing did not happen, and rightly so. The premise gate's shape annotations fired (instance on both non-HOLDS premises — a cheap-additive index defeating the exhaustive-warrants premise GLOBAL, a scope-determination edge case LOCAL); with no prevalence- or priced-exception-shaped defeats, no mode conversion was routed, and the pass instead found the warranted claim in the note's warrant structure: a universal insufficiency claim ("current-task fit alone does not warrant costly structural entrenchment"), retitled by ordinary keep-reframe. The survey had misread the body hedge ("often invisible, rarely revisited") as the core claim; the pass located the core elsewhere, and the hedge survives as a description of the bet's visibility, not as the thesis. Readouts: (a) shapes present and accurate; (b) no target mode named — consistent with the routing, since the shapes gave no mode signal; (c)/(d) mode guards unexercised this run; (e) follow-ups recorded and executed (relocated with redirect, four citers reconciled — including the coordination-value definition, whose gloss had become false the moment the reframe made coordination the rule's third warrant). Bonus behavior worth keeping: the closing premise rerun surfaced a new GLOBAL defeat (a temporary deadline with present stakes is a fourth warrant the packet's exhaustive formulation omitted), and the pass correctly routed it without another edit round — the reframed insufficiency title survives it because "alone does not warrant" is not an exhaustiveness claim.
Series consequence: run 1 turned into an unplanned second no-false-fire datum — the machinery declined a mode conversion on a note we expected to convert. Statistical-guard coverage now rests entirely on runs 3 and 5; if run 3 also lands off-mode, promote structure-activates-higher-quality-training-distributions.md (numeric prevalence core, survived null already in the body) into the series immediately.
Run 2 — ideal-type conversion.
kb/notes/agent-runtimes-decompose-into-scheduler-context-engine-and-execution.md
The hedge "in many real systems the boundaries blur; the claim is that the functions are analytically distinct" is an undeclared first-order model. Expected: keep (title may stand) with a body edit converting the hedge into a declared idealization plus adequacy record — declared use (what the decomposition is for: predicting which limitation a change fixes), omitted mechanism (implementation blurring), bound, dominance — and the closing premise rerun attacking that record in the same pass. Failure tells: conversion without the record (immunization guard missed), or the record present but the closing premises never engaging it.
Run 3 — upward reframe from vacuity, the hard test.
kb/notes/the-framework-is-often-larger-than-the-durable-contribution.md
"Often" in the title, "tends to" in the body, no refuter anywhere — the purest Class B case. Expected: reframe that lands on a guarded claim — a statistical form stating what measured framework-to-contribution ratio would refute it, or a stronger conditional the material warrants. This run tests whether the machinery repairs the ratchet's end state rather than reproducing it. Failure tell: the pass keeps or produces another unguarded tendency.
Run 4 — upward reframe to universal.
kb/notes/memory-backed-personalization-can-look-like-model-improvement.md
A bare possibility title with a sharp universal core buried in paragraph two ("it cannot make one of several prompt-compatible commissions authoritative without user-specific evidence"). Expected: keep-reframe up, promoting the refutable core to the title. This is the direct test of bidirectionality — before ADR 066 no repair path could strengthen a claim.
Run 5 — mixed modality in one note.
kb/notes/agent-context-is-constrained-by-soft-degradation-not-hard-token-limits.md
The anchor case: a statistical binding claim (title says "not hard token limits"; its own description already retreats to "not just") over an ideal-type mechanism model (the three-dimension decomposition plus workspace hypothesis, self-described as "not fully separable" and "a working hypothesis"). Expected: the pass assigns different modes to different claims — statistical retitle with a stated refuter for the binding claim, declared idealization with an adequacy record for the mechanism section — without flattening one into the other. The hardest coordination test; a defensible lesser outcome (fixing one mode and routing the other to Open items) is a finding, not a failure.
Run 6 — control: a sound universal the machinery must leave alone.
kb/notes/an-outcome-check-licenses-replay-a-rule-needs-the-process-verified.md
Deductively argued ("a rule asserts 'do X because Y'; an outcome check never inspected Y"), correctly universal, not in the candidate survey. Expected: no modality finding — shape annotations may appear on any dented premise, but no mode reframe fires. Failure tell: the pass invents a statistical or ideal-type reading for a claim whose warrant is deductive. A machinery that fires everywhere is as broken as one that never fires.
Readouts, runs 2–6 (all passes completed 2026-08-19)
Run 2 (pass 63b341) — ideal-type conversion DECLINED, and rightly. Keep-reframe, but to an analytical-classification claim ("Agent-runtime analysis should separate scheduling, context assembly, and external state"; relocated), not a declared idealization. The shapes discriminated: premise 4 (practitioner convergence) DOUBTFUL GLOBAL prevalence — and with only two mapped taxonomies, a statistical landing had no warrant, so the convergence claim was removed, not moded. The ideal-type routing test failed honestly: implementation blurring is ordinary unmarked practice in the runtime domain — no marked interface, no pejorative, no charge, no ritual — so by the criterion note's own standard, the survey's Class C diagnosis was wrong and the pass's decline was right. Notable: the pass's own closing cycle caught the step-9 revise strengthening "distinct failure questions" into unwarranted causal fault localization and flagged its warranted contribution as changed — plus a fresh closing GLOBAL defeat (opaque joint controllers), both routed without a second round. Genre-drift datum: an ontology title ("runtimes decompose into") landed as a methodology-shaped one ("analysis should separate") — third instance for the cohort thread.
Run 3 (pass 0d5090) — vacuity repaired, landing on a categorical rule, not another tendency. The prevalence annotation fired on four premises (framework familiarity, leakage tendency, cue-activation reliability, author detection) and a genuine priced-exception on one (a regulated clinical domain charging missing warrant more than surplus — the annotation vocabulary used exactly as designed). With no prevalence evidence available, the statistical guard correctly blocked a statistical landing, and the reframe went up in form: a scoped conditional retention rule. The refuter discipline is explicit in the packet's open items: "Do not restore 'often,' 'usually,' or 'tends' without a declared statistical comparison and evidence that could refute it." Rename deferred by the packet, then resolved: the closing cycle GLOBAL-defeated the new title's universal anchor requirement (an enforced consumption path can supply the recognition), and the operator approved the narrowing 2026-08-19. Applied: retitled to "A linked note's durable payload is what its consumption path cannot reliably supply" (relocated with redirect to linked-note-durable-payload-is-what-consumption-path-cannot-supply.md), the consumption-path condition written into the opening and the recognition rule, citers reconciled. CLOSED.
Run 4 (pass k7p4) — upward reframe correctly DECLINED. Plain keep. The buried universal this series planned to promote ("only user-specific evidence can make a commission authoritative") was DEFEATED LOCAL instance by an organization-wide authoritative instruction — promoting it would have shipped a false title. The note's "can look like" is not vacuous: it is a witnessed possibility claim whose contribution is the attribution consequences, which the pass sharpened (carrier-dependent estimands, split swap prescriptions). Lesson: possibility-with-witness is a legitimate form the Class B diagnosis conflated with vacuity.
Run 5 (pass 1f5b0b) — THE STATISTICAL LANDING, guards met. Keep-reframe with the target mode named in the disposition and the guard stated verbatim: "the claim would fail if representative workloads usually stayed reliable until the cap, or if inability to fit required evidence were usually the first constraint." Routed from premise 2 DEFEATED GLOBAL instance (a real hard-cap-first workload: two pruned corpora exceeding the window). The mixed-modality expectation resolved differently than predicted: the mechanism section was compressed to a conjecture with retained prediction and falsifier, not converted to a declared idealization — defensible, since the workspace hypothesis is at conjecture stage and its exceptions are unpriced; the two-bound framing is presented as a useful model. Follow-up executed with one authorial addition: the closing cycle failed title-body-alignment because the pass's own H1 dropped the fits-in-window condition, so the retitle carries it ("Soft degradation often binds before the hard cap when required evidence fits"; relocated; thirteen citers' link text updated; two glosses asserting the old universal reconciled).
Run 6 (pass ff5a) — control: no modality finding fired, but the control was impure. The note turned out to have genuinely defective premises (verbatim replay of a side-effectful charge_customer call is not safe because it succeeded once; outcome evidence CAN warrant same-context retention), so the pass fired an ordinary scope/category reframe ("A checked outcome licenses retaining an episode, not abstracting its explanation"; relocated, citers reconciled including one gloss asserting the defeated only-process-checks claim). For the machinery test, the control still passes: no mode was named, no statistical or ideal-type reading was invented, defeats were instance-shaped and repaired classically. For test design, the control selection was flawed — a deductive-looking note is not the same as a sound one, which is itself a small datum for the KB's review-value story.
Series verdict
Mode landings: one of six (run 5, statistical, guard in the disposition). Declines: four, every one for evidenced reasons, and every survey per-note prediction among them wrong — the machinery's shape-routing out-discriminated the survey's class labels in all four cases. Success criteria: every landed reframe named its mode and met its guard (run 5); an upward-in-form reframe fired (run 3, vacuous tendency → categorical rule); the control produced no modality finding. NOT exercised in-series: an ideal-type conversion with in-pass adequacy attack (covered pre-series by pass 2a6408 on the instantiation note; no series candidate survived the routing test) and a true statistical guard rejection (a proposed statistical landing blocked mid-pass — run 3's block happened at synthesis, which may be the natural place).
The headline finding is better than the one we designed for: mode conversion is rare because mode landings require mode-appropriate warrant — prevalence evidence for statistical, domain-priced exceptions for ideal-type — and most mismatched notes lack it. Their honest repairs are scope, category, or conditionality reframes. The modality machinery's principal observed effect is preventive: shape annotations route synthesis away from unguarded landings, and the two runs that could have produced degenerate outcomes (run 3 re-hedging; run 4 promoting a false universal) both avoided them. The 22-candidate survey overpredicted conversions systematically; class labels flag mismatch but the evidence decides the landing. Second-wave implication: run candidates through ordinary passes without mode predictions; track only whether guards bind and shapes discriminate.
Second wave (queue, not scheduled)
After readouts from the six: structure-activates-higher-quality-training-distributions.md and knowledge-storage-does-not-imply-contextual-activation.md (statistical with numeric prevalence refuters), entropy-management-must-scale-with-generation-throughput.md (deontic universal), weakly-discriminated-qualities-tend-to-be-underselected.md (well-behaved Class B with named conditions), codified-scheduling-patterns-can-turn-tools-into-hidden-schedulers.md (upward to universal), files-not-database.md (statistical vs ideal-type boundary decision), bounded-context-orchestration-model.md (the clean-model cluster — last, because reframing the hub note has the widest citer blast radius).
Series success criteria
The machinery validates when, across runs 1–5, every landed reframe names its target mode and meets its guard, at least one adequacy record is attacked by a closing premise rerun, at least one upward reframe fires, and run 6 produces no modality finding. Any guard that binds only because the runner happened to be careful — rather than because the instruction text forced it — is an instruction defect to fix before the second wave.