Competing models of formal links

Output of the formal-link brainstorming brief. Everything in "What the corpus now shows" is established observation from completed reviews; everything from "Four competing models" onward is generated during this brainstorm and is not adopted. No contract, label, or corpus edge changes here.

What the corpus now shows

The four completed label reviews (evidence, rationale, grounds, mechanism) classified over 700 active edges row by row. Beyond the per-label results, they produced four repeated observations that any theory of formal links must explain:

  1. Every inherited label decomposed. evidence split into two directions (207 rows); rationale (134 rows) dissolved into five successors; grounds (283 rows) into seven successors plus a deferred cohort; the mechanism surface (128 rows) into six classes. No audited label survived as a single relation.
  2. The successors converge on a small set. All ~700 rows landed in roughly a dozen relations: premised-on, rests-on, evidenced-by/is-evidence-for, extends, exemplifies, defined-in, implements, compares-with, see-also, the candidates explained-by/operates-through, and an unresolved prerequisite family. Decomposition did not open an unbounded ontology; it drained into the same basin every time.
  3. Only the row-level assertion classifies an edge. Origin label did not predict semantic class (the old mechanism rows split 19 explanatory / 51 operational; the deferred grounds rows split 22 / 21). Same-file co-occurrence did not either: all nine files authoring both labels used them inconsistently. Grammatical direction alone did not settle semantics; the assertion had to be read in context.
  4. Adjudication used two tests in tandem. Every accepted split was justified by a distinct reader follow/skip decision and a distinct revision consequence. premised-on vs rests-on: rejection of the target reopens the source's truth vs triggers reconsideration of a design. explained-by vs operates-through: the reader may skip a known explanation but not the operational path; a target change triggers causal-argument review vs interface review. Candidates failing both tests (is-grounded-in, generic depends-on) were rejected even though they were propositionally coherent.

Separately established: the footer surface is the only mechanically enumerable link surface. Every review above was possible because footers have a normalized, rg-findable grammar; the inline surface carries richer assertions but cannot be swept, counted, or migrated.

The inheritance record

The Ars Contexta inheritance is itself a completed experiment, and its record is precise enough to constrain theory:

  • What was inherited, and with what theory. ADR 009 adopted extends, grounds, contradicts, enables, exemplifies as a universal mandate ("every link in the KB must articulate the relationship using one of these types"), drawn via Ars Contexta from concept-mapping research. The theory: the difference between mind mapping ("these relate somehow") and concept mapping ("this extends that because…") is propositional articulation — articulated links retain relationship detail, a constrained vocabulary makes recurring distinctions queryable and forces the author beyond "related," typed links are falsifiable where untyped links are not, and prose can still express what the small vocabulary cannot.
  • Where it worked. The vocabulary held up in the theory collection for years and still supplies most of its live labels. All five inherited labels are relations between propositions — and the theory collection's title-as-claim convention makes (nearly) every endpoint there a proposition. Traversal-as-reasoning ("since [claim]…") composed exactly as ADR 009 promised.
  • Where it broke. Reference artifacts denote systems, components, decisions, versions, and procedures. The relations the reference collection needed — part-of, implements, supersedes, procedure, compares-with, rests-on — are structural, versioning, and theory-dependence relations between described things, which an inferential vocabulary between propositions cannot express. ADR 009 had even flagged the coverage gap ("temporal succession, mutual dependency") as edge cases to be absorbed by clarifying phrases; for reference they were the main cases.
  • What the repair did and did not provide. ADR 019 moved authorization into collection contracts with per-destination blocks and demoted the catalogue to a palette. But it explicitly did not supply a generator: its own consequences section concedes that a new collection "must declare its outbound destinations explicitly," cushioned only by suggested starter sets. Ars Contexta gave reasons to formalize; neither it nor the repair gave a procedure that takes a collection and produces the vocabulary appropriate to it.

One more observation the reviews left lying in plain sight: the shared catalogue's labels cluster by register-of-origin — theoretical-shaped, descriptive-shaped, prescriptive-shaped, lineage-shaped. Nobody designed that grouping as a theory; the labels sorted themselves that way as collections accumulated what they needed. That clustering is corpus evidence about where vocabularies come from.

Four competing models

Model 1 — Propositional edge

The registered identifier is a first-class assertion about subject matter: source <label> target is a mini-proposition, and the footer surface is a queryable claim database over the KB's content.

Explains: the source-as-subject grammar (ADR 058 makes no sense unless the identifier asserts something); why classification must precede renaming (you must recover what proposition each edge actually asserts); why endpoint level-ambiguity is a genuine defect (propositions need typed arguments, and an artifact endpoint can stand for a document, a claim, or a described system).

Fails or leaves open: it cannot predict which propositions get registered. The adjudications rejected propositionally coherent candidates (is-grounded-in, depends-on) for consumer reasons, not truth reasons, and retain see-also, which asserts nothing. Taken strictly, the model demands an endpoint-role signature the corpus does not record — it diagnoses the mechanism ambiguity but offers only an expensive fix (a role ontology) with no demonstrated consumer.

Model 2 — Reader-need router (the incumbent)

Labels name recurring reader needs; link quality is navigation uncertainty reduced per token of context (linking theory, conditional possibilities).

Explains: mandatory context phrases and the articulation test; position-encodes-strength; collection-owned vocabularies (different collections serve different readers, which is exactly why the Ars Contexta vocabulary failed to transfer); the catalogue's reader-need column.

Fails or leaves open: reader need alone did not adjudicate the migrations. premised-on and explained-by can serve the same felt need ("understand why this holds") yet were separated by what rejection of the target propagates. The model gives no reason to care about grammatical direction — a router doesn't need a grammatical subject. And it doesn't explain why the footer must be normalized: an in-context reader gains nothing from mechanical uniformity; only sweeps do.

A live repair: widen "reader" to include the maintenance-time reader — the agent that arrives at the edge because the target changed and must decide what to recheck. Then the revision consequence is a reader need, just for a reader who arrives mechanically rather than mid-argument. Whether this widening is a natural extension or a refutation dressed as one is a genuine open question of the workshop.

Model 3 — Maintenance dependency

A formal edge is authored exactly when a target change should trigger reconsideration at the source; the label names the kind of reconsideration that propagates.

Explains: the premised-on/rests-on boundary (stated in ADR 060 entirely in rejection-consequence terms); the explained-by/operates-through boundary tests (causal-argument review vs interface review); the lineage family, whose four labels differ precisely in maintenance regime (recheck vs re-derive vs re-examine); why weak associative edges should not be authored (no revision consequence, nothing to propagate).

Fails or leaves open: navigation-only labels (see-also, defined-in, compares-with, contains) carry no revision consequence yet earn registration. is-evidence-for explicitly asserts no target-side uptake, so its maintenance consequence lives at the endpoint that did not author it — the model predicts that edge shouldn't exist, yet the corpus demonstrated a real inverse reader journey. And no maintenance consumer currently exists: nothing today walks labels to route revision, so the model's central benefit is still speculative.

Model 4 — Lossy projection onto consumer dimensions (candidate synthesis)

The inline prose assertion is the full meaning of a relationship. The registered footer identifier is a lossy projection of that assertion onto the dimensions recurring consumers actually use — currently two: the follow/skip decision (in-context reader) and the revision route (maintenance-time reader). Labels are index keys for those consumers, not the meaning itself, and the vocabulary is harvested, not designed: a distinction earns an identifier when a recurring cluster of assertions differs on a consumer dimension, and a registered label decomposes when its cluster is discovered to mix consumer behaviors.

Explains: observation 1 (inherited labels decomposed because the old projection lumped edges whose consumer behavior differed); observation 2 (convergence to a small basin — the consumer behaviors are few, even though the assertions are many); observation 4 (the two adjudication tests are the projection dimensions); the transfer failure (reference readers make different follow/skip and maintenance decisions than theory readers, so the projection has to be re-derived per collection — collection-owned vocabulary falls out as a theorem rather than a workaround); the unique footer value (normalization serves the mechanical consumer; prose serves the human-order one); Ars Contexta's own caveat (prose expresses what the vocabulary cannot, because the projection is lossy by design).

It also dissolves rather than answers the duplication question: an inline link and a footer link to the same target are not duplicates, because they serve different consumers — the inline occurrence is the argument, the footer entry is its index record. Duplication is warranted exactly when both consumers exist for that relationship, and suspect otherwise.

Fails or leaves open: the dimension count is asserted, not derived — is associative activation a third consumer dimension or forever a derived view? is deterministic validation a fourth? The migrations applied the projection ex post, over a settled corpus; whether authors can apply it reliably ex ante, at writing time, is untested — and if they cannot, the model quietly degenerates into "audit and re-project periodically," which is a very different operating regime. Level ambiguity is not fixed, only bounded: the projection makes artifact/claim/phenomenon slippage harmless where consumer behavior is level-invariant, and the mechanism case shows it is not always invariant.

How the models relate

Model 4 subsumes the other three as partial views: propositional content lives in prose and claim titles (Model 1's territory), and the two projection axes are Model 2's and Model 3's respective primitives. But the subsumption is cheap to claim and hard to test, and Model 2-widened (maintenance as a species of reader need) is observationally almost equivalent to Model 4. The honest discriminator between them is whether the normalization of the footer surface does independent work — Model 4 says the footer exists for consumers who arrive mechanically, Model 2-widened says it is merely a tidy habit. The migration program itself is evidence for Model 4 here: the KB could only audit and repair its own linking practice because that practice had a normalized surface. Self-correction was purchased by the grammar, not by the vocabulary.

The generator problem

The inheritance record poses the question none of the models answers by itself: is there a generator — a procedure that takes a collection and produces the link vocabulary appropriate to it? Ars Contexta could not be turned into one. This brainstorm's candidate explanation of why: its theory (articulation, queryability, falsifiability) justifies having typed links but is silent on which types, because its vocabulary was implicitly indexed to one endpoint kind. The five inherited labels are relations between propositions. The theory collection satisfied that hidden precondition through title-as-claim; the reference collection violated it; the mandate was universal, so it broke exactly where the precondition failed. A vocabulary was presented as general that was actually the projection of one artifact ontology.

If that diagnosis is right, the generator's input is not "a collection" but what the collection's endpoints denote, and the sketch has two stages:

Stage 1 — seed from the text contract. The contract fixes what an artifact stands for, and the endpoint kind licenses the relation families before any practice accumulates:

endpoints denote licensed relation family corpus witnesses
claims inferential (premised-on, extends, contradicts, evidenced-by, candidate explained-by) theory collection
described systems, decisions, versions structural/mereological/versioning (part-of, implements, supersedes, compares-with) reference collection
procedures and steps control-flow (composition, precondition, invokes, applies-when) instructions collection
captured observations evidential/lineage (is-evidence-for, adapted-from, abstracted-from) sources collection

Cross-kind relations then have type signatures, which is why they arise per destination pairing: rests-on (design → claim), operates-on (procedure → system), evidenced-by (claim → observation). ADR 019's per-destination blocks — adopted as a pragmatic repair — fall out of the sketch as a prediction: authorization must be per-pairing because relations are typed by both endpoint kinds. And the register-of-origin clustering of the catalogue is what the sketch retrodicts: labels sorted by register because register tracks endpoint kind.

Stage 2 — correct by harvest. The seed will be wrong in detail; the observed label lifecycle (drift or overload → row-level classification → two-consumer adjudication → mechanical migration) is the correction loop. Stage 1 without stage 2 is Ars Contexta's mistake in new clothes; stage 2 without stage 1 leaves every new collection guessing, which is the ADR 019 status quo.

The sketch also explains the mechanism ambiguity as its own failure mode rather than an anomaly: collections are only approximately homogeneous in endpoint kind. The theory collection contains notes that describe systems and processes (the orchestration model, the tool-loop note, definitions/codification.md), so its endpoint ontology forks inside one collection — and the EX/OP split is exactly that fork surfacing. When a claim-titled note is read as a proposition, the licensed relation is inferential (explained-by); when read as the description of a process, the licensed relation is operational (operates-through). The 128-row classification was, on this view, endpoint-kind recovery performed ex post, row by row. This predicts where level-ambiguity will bite next: any relation whose two readings straddle the claim/description line, and no relation whose readings stay on one side.

All of this is a candidate explanation with one collection-sized data point per row of the table. It would be refuted if a collection's harvested vocabulary turned out not to track what its artifacts denote, or if the next new collection seeded this way needed mostly relations the seed could not have predicted.

Candidate practical conclusions

Generated, not adopted. Each names its adoption route; none should be executed from this file.

  1. Make the two-consumer test the explicit registration bar. A distinction earns a registered identifier iff it changes the follow/skip decision or the revision consequence; every catalogue entry states both (reader need is already there; revision consequence is not). This is already the de facto test of ADRs 058/060 and the mechanism review — writing it down is codification of existing practice, not a new commitment. Route: catalogue edit + a short ADR or an amendment folded into the eventual mechanism ADR.
  2. Protect grammar strictness independently of vocabulary looseness. The footer grammar (one link, one authorized identifier, one context phrase, rg-findable) is what made every audit and migration possible; the vocabulary can stay collection-owned and loose because the grammar is strict. Treat proposals that weaken the grammar (free-form labels, prose-embedded identifiers) as attacks on the KB's self-correction capacity, not as flexibility. Route: this is arguably already policy; worth one sentence in the link-vocabulary approach section.
  3. Resolve the duplication question per-consumer, not per-relationship. Author inline when the relationship does argumentative work in the body; add the footer entry when the relationship should be visible to sweeps, maintenance, and the typed adjacency scan; both when both. Route: needs worked cases first — trial the rule during normal authoring before contract language.
  4. Codify the two-stage generator as the process Ars Contexta lacks. Stage 1: seed a new collection's vocabulary from its text contract — determine what its artifacts denote, take the relation families that endpoint kind licenses, and derive cross-kind labels per destination pairing from the type signatures. Stage 2: correct by the observed, twice-repeated label lifecycle — author under a loose local contract → drift or overload observed → full-corpus row-level classification → two-consumer adjudication → mechanical migration with tuple conservation. That is the answer to "how should a collection's vocabulary arise": seeded from what the collection is about, then harvested into shape from articulated practice. Route: a note capturing the mechanism plus a pointer from the migration instruction; the seeding stage could later become a paragraph in the link-vocabulary authoring guidance, after the retrodiction test below.
  5. Handle level ambiguity with per-identifier boundary tests, not a role ontology. The mechanism review's boundary tests ("the target is a process/component through which…" vs "the target is an explanatory account answering why…") disambiguated endpoint levels well enough to classify 128 rows. Registering a full source-role/target-role signature would formalize what boundary tests get cheaply. Adopt signatures only for a relation whose consumer behavior demonstrably changes across levels. Route: keep the candidate signature as an evaluation device; add boundary tests to catalogue entries as labels are migrated.
  6. Expect vocabulary closure. ~700 classified edges drained into ~a dozen relations. Treat the successor set as approaching stability: a proposed new label should be presumed to be one of the existing relations misread, until a corpus cluster with distinct consumer behavior is shown. Route: stance, not rule; belongs in the eventual theory note.

Discriminating evidence

Ordered by how much each would settle, cheapest first.

  • Consumer census. Enumerate what actually consumes the footer surface today: navigating agents, cp-skill-connect, the migrations themselves, validation, anything else. If no maintenance-time consumer exists or is planned, Model 3's axis (and half of Model 4) rests on a hypothetical, and conclusion 1 should register reader need only. Cheap: a survey, not an experiment.
  • Generator retrodiction. Apply stage 1 blind: from each collection's contract alone (what its artifacts denote, which destinations it links), predict its relation families, then compare with the live authorized vocabulary. Strong overlap plus explicable residue supports the generator; a collection whose harvested vocabulary doesn't track its endpoint kinds refutes it. Cheap and fully in-corpus. Run 2026-07-29, three variants — see the retrodiction run. Palette-anchored: family structure recovered blind (81% exact outside two systematic miss classes), but stage 1 alone cannot write a label table. Palette-free: removing the catalogue left the semantics intact — relation-level matching 46% vs 48% with every identifier coined blind, and six coined identifiers reproduced live labels character-for-character (rests-on, supersedes, invokes, extends, defined-in, evidenced-by). Family-invention: removing the family inventory too, all nine harvested families were independently re-derived (9/9 census), evidence direction stayed unflipped across all three variants, the contested types→reference pairing resolved correctly once predictors had to state each relation's revision consequence, and one credible novel relation (elaborated-by, detail routing) surfaced as a candidate vocabulary gap. Combined verdict: everything semantic (families, directions, gists, largely names) is generated from endpoint kinds; everything authorizational (completeness, policy constants, lineage-regime coverage, deliberate under-commitments) belongs to the policy/harvest layer — coverage fell 48%→46%→41% as scaffolding was removed while semantic precision held or improved. Supports seed-then-harvest; refutes seed-alone. A controlled A/B (k=5 per arm on the types contract) then refuted the suggestive lead that the revision-consequence requirement disambiguates contested pairings: the correct dependence family emerged in 10/10 runs in both arms, and the "contested pairing" dissolved into a stable over-complete portfolio (dependence + enforcement + consumer relations) from which the contract selects — single-run family assignments had been portfolio sampling, not family confusion. The seed emits a superset portfolio per pairing; authorization is selection from it, and selection is what only the corpus record determines. A subsequent corpus check then confirmed the generator's forward mode: all three relations it coined repeatedly without authorization (elaborated-by, is-enforced-by, is-consumed-by) turned out to name real, recurring relationships currently forced into weaker labels (contains used backwards, enforcement filed as depends-on, consumption as see-also) or lost to unlabeled prose — the same seed that retrodicts the vocabulary predicts where it is incomplete. A control check on a deliberately weak-link pairing (the system-review collections' analogue relation toward reference) then failed as discrimination requires — ~1,280 comparison mentions, essentially zero doc-targeted instances, and contracts that verbatim pre-assign the analogue need to see-also — so the forward mode distinguishes real gaps from over-generation rather than firing everywhere; the control also exposed why the seed over-fires there: the KB mediates cross-system correspondence through shared theory claims, an architecture choice endpoint kinds cannot see. The prospective version — seed the next genuinely new collection and count how many harvested labels the seed failed to predict — remains the real test.
  • Ex-ante projection test. During normal authoring over a few weeks, have the author record label choices with a one-line justification, then blind-reclassify a sample later (or by a second agent). Low agreement on EX/OP-like boundaries kills fine distinctions at authoring time and pushes toward coarse labels + periodic re-projection. This is also exactly the reversal evidence the mechanism review requested.
  • Follow/skip trace study. Same navigation tasks, labels shown vs hidden (titles and context phrases kept). If choices don't change, labels do no navigation work beyond titles — Model 2 loses its primary axis and the footer's value collapses onto the mechanical consumer, which would itself be a decisive (and ironic) finding for Model 4.
  • Revision-propagation trial. Revise one heavily-linked target note; compare the review set selected by walking labeled inbound edges against plain backlinks. If labels don't select a better review set, the revision-consequence axis is decoration and conclusion 1 should drop it.

Tensions preserved

  • Whether "reader need" legitimately covers the maintenance-time reader, or whether that widening concedes Model 3's point while keeping Model 2's name.
  • Whether authors can project ex ante at all — every successful classification so far was ex post over a settled corpus.
  • The prerequisite family remains directionally unresolved (enables vs precondition), and its 10 rows must not be swept into any mechanism outcome.
  • is-evidence-for places the maintenance consequence at the endpoint that didn't author the edge; no model above handles this cleanly, and it may be the strongest argument that a derived inbound view is a real consumer, not a convenience.
  • Whether "what an artifact denotes" is stable enough to be a contract-level fact. The generator sketch needs endpoint kind to be readable off the collection, but the mechanism case shows the theory collection forking between claim and description readings note by note — the generator's input may only exist at the granularity where the ambiguity already lives.
  • Whether associative/weak links ever earn authoring, or whether search and derived views dominate them permanently.