Case packet

Neutral case identifier: case-cbbd1d7667e81e

The possible directed relationship from Artifact A to Artifact B is under review.

Artifact A

Only derivation and inheritance warrant a decomposition's scope claim; discriminating use earns it

A decomposition claims more than it displays. Its categories are visible; its implied scope is not. By using those categories, the designer asserts that they will continue to cut the domain at consequential boundaries, including in cases the designer has not seen. That assertion is an [explanatory-reach] claim. Because [use tests a decomposition only locally], merely operating a system that contains the decomposition does not test its broader scope.

Two judgments about that claim are easy to conflate. Its provenance determines whether it begins with any scope warrant, what kind of warrant that is, and what remains to be tested. Whether the claim is earned is settled the only way [reach-assessment] recognizes: by surviving evidence that could have refuted it. Provenance determines the starting scope warrant; refutation-capable use earns the claim.

Three labels describe that provenance. They are not mutually exclusive bins applied once to an entire decomposition. They may differ by axis or boundary: derivation marks what independently supported constraints fix, inheritance marks what is borrowed together with evidence from a source domain, and free choice marks the residue that neither source fixes. One decomposition may therefore be derived and inherited in some respects while remaining free in others.

Derivation from independently supported constraints gives conditional warrant. Once the axes are fixed, the cells follow, including the empty ones. This derivation warrants the conditional if these axes remain operative, these cells remain exhaustive; it does not establish that the axes track consequential distinctions. Axes selected only because they produce the desired split remain free choices, even if their cells follow mechanically. When independently supported causal or operational constraints fix the axes, that support supplies the conditional's antecedent. The derivation may still be valid even when those constraints cease to matter in the next case: its warrant is conditional, not universal.

Inheritance of a tested ontology gives transferred warrant only when the source testing was discriminating. Repeated survival across cases that could have broken the ontology is the extensional route to a scope claim; borrowing the ontology acquires that evidence secondhand. Mere maturity or repeated use does not. Moreover, the evidence still concerns the source domain. It licenses the decomposition in a target domain only while the constraints that made its boundaries consequential at the source also hold at the target.

Free choice gives no starting warrant for the scope claim. A boundary underdetermined by constraints or evidence may have a pragmatic reason for adoption, such as convenience or preference, but there is no reason to expect it to remain consequential. It should therefore remain cheap to replace.

Direct mechanistic or predictive support is not a fourth provenance. If independent measurement fixes the axes, the decomposition is derived from that constraint. If the decomposition survives a held-out comparison, intervention, or risky prediction, that encounter is already discriminating use, even when it occurs before deployment. A theory-guided proposal that is neither fixed by independently supported constraints nor exposed to refutation-capable evidence remains free wherever it is underdetermined.

Derivation and inheritance relocate the test; use runs it

A provenance label alone does not earn the target-domain scope claim. Derivation supplies no encounter with evidence that could have broken the decomposition, while inheritance rests on evidence from the source domain rather than the target. Both instead relocate the remaining test to somewhere statable. Derivation moves it to the antecedent: vary or intervene on the named constraints and ask whether the boundaries change as predicted. Inheritance moves it to the bridge: establish that the source testing was discriminating, then check whether the source's operative constraints still hold at the target. Free choice supplies no comparably targeted rationale. A single successful local use still establishes only local sufficiency, whereas an accumulating sample of varied, refutation-capable contexts can test transfer extensionally.

Earning is therefore open to every provenance. A decomposition that began as free choice can earn scope one context that could have broken it at a time. This is the route [an automated theory-search loop] must take for each proposal: a discriminating acceptance test contributes earned reach, while a merely confirmatory one manufactures an unearned claim.

[What scale replaces is a generalization whose scope was asserted rather than tested], and a decomposition makes exactly such a generalization when it asserts that its boundaries will continue to matter. Starting warrant merely makes the remaining test statable. A free choice is not necessarily sloppy, but it should stay replaceable. Encoding it in a directory layout, schema field, or widely cited vocabulary term adds revision cost without adding warrant or evidence.

The two triples relate but do not coincide

Retained rationale classifies each boundary by what it answers: an inherited constraint, a local requirement, or a free choice. Provenance classifies why a boundary or decomposition begins with any scope warrant: derivation, inheritance, or neither. The two classifications overlap but do not map one-to-one.

Derivation can serve either of the first two rationale slots. A boundary may be fixed by an inherited constraint or by a local requirement. For example, [decomposing user stories into their step-by-step context needs] derives boundaries from one application's local requirements without borrowing them. Its conditional warrant extends only to problems that share those requirements.

Inheritance describes where the rationale and its evidence came from, not which rationale slot each boundary occupies. When borrowing is undocumented or lossy, source-wide constraints and source-local requirements arrive bundled, making every boundary appear equally load-bearing. That hidden cost comes from the missing provenance record, not from inheritance itself. The distinction also explains why a technique can demonstrate transfer by working in the target while an ontology transfers only if it continues to cut at consequential places there. [Closeness of fields] is evidence for that bridge, not a substitute for it.

Two recorded instances

[Representational form], which classifies how retained content is encoded and consumed, is derived from the axes of consequence-assignment and localization. Those axes generate three occupied cells, an explicit empty fourth cell, and the read/test/probe rule. The starting warrant is conditional on the axes continuing to matter; whether they identify consequential distinctions remains the test.

A [reflective system], which represents and acts through aspects of itself, is inherited. Causal connection, self-representation, and theory-relativity come from Maes's 1988 account and Smith's 1984 lineage. The definition separates retrieval-as-causal-connection as Commonplace's own extension. Because the source evidence concerns interpreters and metaobject protocols, applying those criteria when the causal wire is best-effort discovery creates a separate bridge claim.

Scope

  • The quality of the source testing is the weak joint of inheritance. Nothing here provides a criterion for judging it, and longevity is a poor proxy. An ontology can survive simply because it was never probed, reproducing the failure of a reach claim that no test could refute. The transferred warrant is only as good as the testing behind it, which a borrower often cannot audit.
  • The provenances may combine. Re-deriving a borrowed decomposition from local constraints offers the strongest available position, but also the most expensive. It can be difficult to distinguish from rationalizing the borrowed split after the fact. An unfaithful rationale then inherits the [worse-than-none asymmetry]: it directs tests toward the wrong premise.
  • The classification of direct support is analytic rather than demonstrated. A case whose axes are neither derived from constraints nor inherited, yet which has genuine starting scope warrant before any refutation-capable encounter, would require another provenance or a weaker title.
  • Commonplace has not demonstrated that this claim transfers. The two instances were authored inside the system that states the claim and assessed by nobody outside it. They establish only that both warrants can be recorded and that the record makes departures auditable. They do not show that recording warrants produces better decompositions or that the discipline survives contact with a designer committed to a decomposition they already had.

Open Questions

  • What would make "already tested" checkable for a borrower who cannot rerun the source field's cases? Is there any evidence short of finding a case that should have broken the original decomposition but did not?
  • Can a re-derivation of a borrowed decomposition be distinguished from a rationalization of it by anything other than intervention on the stated constraint?
  • Can "stay replaceable" be operationalized as a forbidden position or a cost ceiling for a free-choice decomposition, or will it remain advice that loses force as soon as another artifact cites the decomposition?
  • What makes a sample of contexts discriminating rather than merely accumulated? The extensional route earns scope only as fast as its sample could refute, and nothing here determines when varied use crosses that line.

Relevant Notes:

Artifact B

The bitter lesson selects against unearned reach, not against structure

The [bitter lesson] is usually compressed to "hand-built structure loses to scale." On that reading any system that discovers, names, and retains explicit theories is building the thing scale is about to eat.

The compression is wrong at a specific point. What loses is not structure and not human origin — it is a generalization whose claimed scope was asserted rather than tested. Human-produced exact specifications, tests, interfaces, and measurement systems are frequently what make scaling possible, and calculators and validators do not become bad because learned systems got better; [exactness and proxyhood attach to an artifact's requirement chain, not to the artifact alone], and only the conjectured links in that chain are exposed.

The sharper statement is about [explanatory-reach]. A theory claims a scope. Where that claim was earned — the structure it names really does hold across the range asserted — a scalable search eventually finds the same structure, and finding it is agreement rather than replacement. Where the claim was not earned — it fit the cases that produced it and its scope was asserted on the strength of that fit — a method with more compute and a better signal replaces it. Low-reach adaptive fit is what loses, and human authorship is merely the most common way to produce it.

Claiming reach is not earning it

The tempting converse is that high-reach methods resist being bitter-lessoned. That is false as stated, and the KB holds the case that refutes it.

[DomainBed] evaluated nine domain-generalization algorithms against carefully tuned empirical risk minimization across seven multi-domain datasets under a declared model-selection protocol. Every one of those algorithms makes an explicit reach claim — that it captures structure surviving a change of environment, which is exactly a claim to operate beyond the distribution that trained it. ERM matched or beat all of them. Reach was claimed in every case; what was absent was any test separating the claim from an artifact of an undeclared selection procedure, and declaring that procedure dissolved the advantage.

Formalizing the claim does not rescue it either. [Rosenfeld, Ravikumar, and Risteski] construct a predictor that discharges the invariant risk minimization objective and is indistinguishable from the invariant predictor on training data, while reverting to ERM once the test environment drifts. The obligation is satisfied and the commitment recovered is still the wrong one.

So a reach claim can be explicit, formal, and checked against an obligation, and still be unearned. What separates the cases is whether anything tested the claim against evidence that could have refuted it — which is [reach-assessment], and which the bitter lesson is best read as measuring in retrospect.

Automation moves who supplies the structure, not whether it was earned

This bears directly on automated theory search. A system that searches theory space, derives consequences, and tests them is running search and learning — the side of the ledger the bitter lesson endorses — and its retained theories are not hand-supplied priors. That much is a real answer to the objection.

But it is an answer only if the acceptance test earns the reach rather than confirming the fit. A loop whose gate is "does this theory account for the cases that produced it" is a machine for manufacturing unearned reach claims faster than a human could, and the lesson applies to its output exactly as it applied to the hand-built version. Automating the search relocates the labor; it does not by itself change the property that determines the outcome.

That is the same failure DomainBed found, arrived at automatically. Nine research groups each ran a search, each retained a theory, and the selection variable that would have tested the claims went undeclared.

What this does and does not predict

It does not restore foresight. [Which side of the boundary a component sits on is not identifiable until scale tests it], and nothing here changes that — estimating whether a theory's reach was genuinely earned, before a shift tests it, is the same open problem under a different name. What this claim supplies is an account of what the test is testing, which turns the lesson from a prophecy about structure into a statement about a property structure can have or lack.

The prediction it carries: components that get bitter-lessoned should be the ones whose scope was asserted from source-case fit, and components that survive scaling should be the ones whose scope was checked against cases that could have broken it. A survey of superseded hand-built components that found no such difference would count against the claim.

Open Questions

  • Whether "earned" can be operationalized ahead of the test, or whether it is only ever assigned in retrospect — in which case the claim explains outcomes without guiding decisions.
  • Whether exact specs are a third category or the limiting case of earned reach, where the claimed scope is the whole problem and there is nothing left to be wrong about.
  • Whether the account survives cases where a well-tested theory loses anyway because the general method found a different and better structure, rather than the same one — agreement and replacement may not exhaust the outcomes.
  • Whether [oracle strength] tracks earnedness, since a hard oracle is what lets a claim be tested against refuting cases in the first place.

Relevant Notes:

Under-review context phrase

identifies asserted-versus-tested scope as the property on which scale selects, and supplies the search loop through which free-born decompositions must earn scope extensionally